← Jungyoon Lee

AI City Challenge 2026 — Track 5, Generative Traffic Video Forecasting

2nd place · 76.0385 on the full test set
Team SSUPER · Team lead · ECCV 2026 AI City Challenge Workshop · Jul 2026

Pipeline — observed history and behavior descriptions into a LoRA-adapted Cosmos3-Super, producing N future frames

▲ The submitted pipeline. The 64B Cosmos3-Super base stays frozen; only a rank-16 LoRA is trained, on 15,673 phase-centered clips built from WTS and BDD_PC_5K.

Task

From K observed frames plus paired pedestrian- and vehicle-behavior descriptions, render exactly N future 1280×720 frames — dashcam and fixed-overhead cameras, history 10–224 frames, horizon 51–120.

Approach

Generated future frames for a fixed-overhead and a dashcam scene, conditioned on pedestrian and vehicle descriptions

▲ Left column is the last observed frame; the rest are generated under the paired descriptions. One sample per case, eight denoising steps.

Result

Links

Official challenge site — 10th AI City Challenge, ECCV 2026  ·  Code