VLALight: A Vision-Language-Action Model for Traffic Signal Control
Quick summary
arXiv:2609.36934v1 Announce Type: new Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight dire
Key takeaways
- arXiv:2609.36934v1 Announce Type: new Abstract: Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion.
- Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions.
- We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos.
Why it matters
“VLALight: A Vision-Language-Action Model for Traffic Signal Control” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments