Lightweight Vision-Language Incident QA for Smart-City Traffic Monitoring with Uncertainty-Calibrated Explanations
- Elena Zhou
- EIRA Journal of Multidisciplinary Research and Development (EIRAJMRD)
- https://doi.org/10.5281/zenodo.21609762
Published:
Sunday, 26 July 2026
Volume:
Volume 2, Issue 4 (2026)
Section:
Articles
Abstract
Smart-city traffic monitoring increasingly requires systems that answer operational questions about unusual road-user behavior while exposing uncertainty. This paper presents LVLIQA, a lightweight incident question-answering framework for active-transportation monitoring. Rather than retaining raw camera frames, the system consumes privacy-preserving sensor-hour slices containing location, direction, mode, count, and seasonal context. A robust seasonal residual encoder converts each slice into compact visual-state tokens; a question-answering layer maps the state to normal, traffic-volume surge, or traffic-volume drop; and isotonic calibration converts residual margins into probabilities used in concise, evidence-linked explanations. The evaluation uses complete 2025 California Department of Transportation pedestrian and bicycle hourly count streams. After duplicate resolution, 1,735,171 sensor-hour slices were divided chronologically into January–August training, September–October validation, and November–December testing. On 346,253 test slices, LVLIQA achieved 0.868 binary F1, 0.997 AUROC, 0.911 AUPRC, and 0.0024 expected calibration error. Event-type question answering reached 0.9916 exact-match accuracy and 0.919 macro-F1. A compact decision tree produced the highest binary F1 (0.927), whereas LVLIQA supplied the most direct calibrated answer-and-evidence pathway. Because event labels are defined from the same seasonal-residual family used by the detectors, these results measure reproducible, rule-consistent anomaly screening rather than independently verified crash detection. LVLIQA should therefore be interpreted as an auditable triage layer that prioritizes count anomalies for subsequent operational review.
Keywords: smart-city traffic monitoring, active transportation, incident question answering, uncertainty calibration, anomaly detection
How to cite this work: Elena Zhou. (2026). Lightweight Vision-Language Incident QA for Smart-City Traffic Monitoring with Uncertainty-Calibrated Explanations. EIRA Journal of Multidisciplinary Research and Development (EIRAJMRD), 2(4), 27–39. https://doi.org/10.5281/zenodo.21609762
