AI City Challenge 2026 — Track 5, Generative Traffic Video Forecasting
▲ The submitted pipeline. The 64B Cosmos3-Super base stays frozen; only a rank-16 LoRA is trained, on 15,673 phase-centered clips built from WTS and BDD_PC_5K.
Task
From K observed frames plus paired pedestrian- and vehicle-behavior descriptions, render exactly N future 1280×720 frames — dashcam and fixed-overhead cameras, history 10–224 frames, horizon 51–120.
Approach
- The public 64B Cosmos3-Super backbone stays frozen.
- Rank-16 LoRA on the generation-path attention projections — 39.85M trainable params, 0.062% of the backbone.
- 15,673 phase-level clips from the official WTS and BDD_PC_5K only.
▲ Left column is the last observed frame; the rest are generated under the paired descriptions. One sample per case, eight denoising steps.
Result
- 76.0385 full test / 74.0750 public subset — 2nd place.
- Oral at the ECCV 2026 AI City Challenge Workshop (Malmö, Sept 2026).
Links
Official challenge site — 10th AI City Challenge, ECCV 2026 · Code