daohang fenxiangbox searchbox qikanlogonew daohangnew searchboxnew navrightzone footerzone paper

SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding AITranslate

1.Department of Civil and Environmental Engineering, University of Wisconsin–Madison, Madison WI 53706, USA
2.Lyles School of Civil and Construction Engineering, Purdue University, West Lafayette IN 47907, USA
3.Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette IN 47907, USA
4.Google, Sunnyvale CA 94089, USA
AITranslate
Publisher: Youke Publish Co., Ltd.
Share Citation Information Add to Favorites Download PDF

    Scan to share on WeChat or Moments

Use WeChat scan.
Share with WeChat friends or Moments

Abstract AITranslate

Multimodal Large Language Models (MLLMs) have achieved remarkable progress across a range of vision-language tasks and demonstrate strong potential for traffic accident understanding. However, existing MLLMs in this domain primarily focus on coarse-grained image-level or video-level comprehension and often struggle to handle fine-grained visual details or localized scene components, limiting their applicability in complex accident scenarios. To address these limitations, we propose SafePLUG, a novel framework that empowers MLLMs with both pixel-level understanding and temporal grounding for comprehensive traffic accident analysis. SafePLUG supports both arbitrary-shaped visual prompts for region-aware question answering and pixel-level segmentation based on language instructions, while also enabling the recognition of temporally anchored events in traffic accident scenarios. To advance the development of MLLMs for traffic accident understanding, we curate a new dataset, SafePLUG-Bench, which contains diverse multimodal question–answer pairs with detailed pixel-level annotations and temporal event boundaries across a wide range of accident scenarios. Experimental results show that SafePLUG achieves strong performance on multiple tasks, including region-based question answering, pixel-level segmentation, temporal event localization, and accident event understanding. These capabilities lay a foundation for a fine-grained understanding of complex traffic scenes, with the potential to improve driving safety and enhance situational awareness in smart transportation systems.

KeyWords AITranslate

multimodal large language models safety-critical perception traffic accident understanding transportation safety

1.S. Yin, C. Fu, S. Zhao, et al., "A Survey on Multimodal Large Language Models," National Science Review 11, no. 12 (2024): 1–20, https://doi.org/10.1093/nsr/nwae403.

2.D. Caffagni, F. Cocchi, L. Barsellotti, et al., "The Revolution of Multimodal Large Language Models: A Survey," paper presented at the Findings of the Association for Computational Linguistics, Bangkok, Thailand, August 11–16, 2024, https://doi.org/10.18653/v1/2024.findings-acl.807.

3.H. Liu, C. Li, Y. Li, and Y. J. Lee, "Improved Baselines with Visual Instruction Tuning," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02484.

4.Z. Li, Z. Cui, H. Liao, et al., "Steering the Future: Redefining Intelligent Transportation Systems with Foundation Models," CHAIN 1, no. 1 (2024): 46–53, https://doi.org/10.23919/CHAIN.2024.100003.

5.Z. Huang, Z. Sheng, Y. Qu, J. You, and S. Chen, "VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving," Transportation Research Part C: Emerging Technologies 180 (2025): 105321, https://doi.org/10.1016/j.trc.2025.105321.

6.Z. Sheng, Z. Huang, Y. Qu, Y. Leng, and S. Chen, "Talk2Traffic: Interactive and Editable Traffic Scenario Generation for Autonomous Driving with Multimodal Large Language Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPRW67362.2025.00364.

7.Z. Sheng, Z. Huang, Y. Qu, et al., "CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models," arXiv preprint, arXiv: 2502.15119, February 21, 2025, https://doi.org/10.48550/arXiv.2502.15119.

8.Z. Huang, Z. Sheng, and S. Chen, "PE-RLHF: Reinforcement Learning with Human Feedback and physics knowledge for safe and trustworthy autonomous driving," Transportation Research Part C: Emerging Technologies 179 (2025): 105262, https://doi.org/10.1016/j.trc.2025.105262.

9.C. Parikh, D. Rawat, R. T. Rakshitha, T. Ghosh, and R. K. Sarvadevabhatla, "RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01770.

10.Z. Xing, H. Chen, B. Xie, et al., "EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01779.

11.Y. Zhou, L. Bai, S. Cai, et al., "TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident," paper presented at the Thirteenth International Conference on Learning Representations, Singapore, April 24, 2025.

12.M. M. Karim, Y. Shi, S. Zhang, et al., "Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review," Artificial Intelligence for Transportation 1 (2025): 100004, https://doi.org/10.1016/j.ait.2025.100004.

13.R. Zhang, B. Wang, J. Zhang, et al., "When Language and Vision Meet Road Safety: Leveraging Multimodal Large Language Models for Video-Based Traffic Accident Analysis," Accident Analysis & Prevention 129 (2025): 108077, https://doi.org/10.1016/j.aap.2025.108077.

14.Y. Yan, Y. Liao, G. Xu, et al., "Large Language Models for Traffic and Transportation Research: Methodologies, State of the Art, and Future Opportunities," arXiv preprint, arXiv: 2503.21330, March 27, 2025, https://doi.org/10.48550/arXiv.2503.21330.

15.S. Jiang, Z. Huang, K. Qian, et al., "A Survey on Vision-Language-Action Models for Autonomous Driving," arXiv preprint, arXiv: 2506.24044, June 30, 2025, https://doi.org/10.48550/arXiv.2506.24044.

16.L. Xu, H. Huang, and J. Liu, "SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, https://doi.org/10.1109/CVPR46437.2021.00975.

17.J. Fang, L. Li, J. Zhou, et al., "Abductive Ego-View Accident Video Understanding for Safe Driving Perception," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02080.

18.Z. Huang, Z. Sheng, Z. Wan, et al., "Sky-Drive: A Distributed Multi-Agent Simulation Platform for Socially-Aware and Human-AI Collaborative Future Transportation," Journal of Intelligent and Connected Vehicles (2026): 1–16, https://doi.org/10.26599/JICV.2026.9210070.

19.Z. Huang, Z. Sheng, C. Ma, and S. Chen, "Human as AI Mentor: Enhanced Human-in-The-Loop Reinforcement Learning for Safe and Efficient Autonomous Driving," Communications in Transportation Research 4 (2024): 100127, https://doi.org/10.1016/j.commtr.2024.100127.

20.K. Long, Z. Sheng, H. Shi, et al., "Physical Enhanced Residual Learning (PERL) Framework for Vehicle Trajectory Prediction," Communications in Transportation Research 5 (2025): 100166, https://doi.org/10.1016/j.commtr.2025.100166.

21.Z. Sheng, Z. Huang, and S. Chen, "Traffic Expertise Meets Residual RL: Knowledge-Informed Model-Based Residual Reinforcement Learning for CAV Trajectory Control," Communications in Transportation Research 4 (2024): 100142, https://doi.org/10.1016/j.commtr.2024.100142.

22.X. Lai, Z. Tian, Y. Chen, et al., "LISA: Reasoning Segmentation via Large Language Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.00915.

23.H. Yuan, X. Li, T. Zhang, et al., "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos," arXiv preprint, arXiv: 2501.04001, November 3, 2025, https://doi.org/10.48550/arXiv.2501.04001.

24.Z. Ren, Z. Huang, Y. Wei, et al., "PixelLM: Pixel Reasoning with Large Multimodal Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02491.

25.W. Lin, X. Wei, R. An, et al., "Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want," paper presented at the International Conference on Learning Representations, Singapore, April 24, 2025.

26.M. Cai, H. Liu, D. Park, et al., "ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01227.

27.T. Zhang, X. Li, H. Fei, et al., "OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding," paper presented at the Conference on Neural Information Processing Systems, Vancouver, Canada, December 10–15, 2024, https://doi.org/10.52202/079017-2291.

28.Y. Yao, X. Wang, M. Xu, et al., "DoTA: Unsupervised Detection of Traffic Anomaly in Driving Videos," IEEE Transactions on Pattern Analysis and Machine Intelligence 45, no. 1 (2023): 444–459, https://doi.org/10.1109/TPAMI.2022.3150763.

29.B. Lin, Y. Ye, B. Zhu, et al., "Video-LLaVA: Learning United Visual Representation by Alignment Before Projection," in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, edited by Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, 5971–5984, Association for Computational Linguistics, 2024, https://doi.org/10.18653/v1/2024.emnlp-main.342.

30.B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, "VTimeLLM: Empower LLM to Grasp Video Moments," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01353.

31.S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, "TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01357.

32.Y. Wu, X. Hu, Y. Sun, et al., "Number it: Temporal Grounding Videos like Flipping Manga," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01284.

33.A. Kirillov, E. Mintun, N. Ravi, et al., "Segment Anything," in Proceedings of the IEEE/CVF Conference on International Conference on Computer Vision, 2023, https://doi.org/10.1109/ICCV51070.2023.00371.

34.Y. Ling and Z. Ma, "Pedestrian Crossing Intention Prediction in the Wild: A Survey," CHAIN 1, no. 4 (2024): 263–279, https://doi.org/10.23919/CHAIN.2024.000008.

35.T. You and B. Han, "Traffic Accident Benchmark for Causality Recognition," paper presented at the European Conference on Computer Vision, Glasgow, UK, August 23–28, 2020, https://doi.org/10.1007/978-3-030-58571-6_32.

36.Y. Qu, Z. Xu, Z. Huang, et al., "MetaSSC: Enhancing 3D semantic scene completion for autonomous driving through meta-learning and long-sequence modeling," Communications in Transportation Research 5 (2025): 100184, https://doi.org/10.1016/j.commtr.2025.100184.

37.J. Fang, J. Qiao, J. Xue, and Z. Li, "Vision-Based Traffic Accident Detection and Anticipation: A Survey," IEEE Transactions on Circuits and Systems for Video Technology 34, no. 4 (2024): 1983–1999, https://doi.org/10.1109/TCSVT.2023.3307655.

38.K. Yin, Y. Li, X. Li, and H. Zhao, "Human Motion Intention Recognition via sEMG and Joint Kinematics Fusion Using MPSO-SVM for Intelligent Transportation Systems," CHAIN 2, no. 2 (2025): 198–209, https://doi.org/10.23919/CHAIN.2025.000012.

39.Y. Guan, H. Liao, C. Wang, et al., "Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation," Communications in Transportation Research 5 (2025): 100214, https://doi.org/10.1016/j.commtr.2025.100214.

40.F. H. Chan, Y. T. Chen, Y. Xiang, and M. Sun, "Anticipating Accidents in Dashcam Videos," paper presented at the Asian Conference on Computer Vision, Taiwan, China, November 20–24, 2016, https://doi.org/10.1007/978-3-319-54190-7_9.

41.H. Lv, C. Zhou, Z. Cui, et al., "Localizing Anomalies From Weakly-Labeled Videos," IEEE Transactions on Image Processing 30 (2021): 4505–4515, https://doi.org/10.1109/TIP.2021.3072863.

42.Y. Yao, M. Xu, Y. Wang, D. J. Crandall, and E. M. Atkins, "Unsupervised Traffic Accident Detection in First-Person Videos," paper presented at the IEEE/RSJ International Conference on Intelligent Robots and Systems, Macau, China, November 3–8, 2019, https://doi.org/10.1109/IROS40897.2019.8967556.

43.W. Bao, Q. Yu, and Y. Kong, "Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational Learning," in Proceedings of the 28th ACM International Conference on Multimedia, 2682–2690, ACM Digital Library, 2020, https://doi.org/10.1145/3394171.3413827.

44.J. Fang, D. Yan, J. Qiao, J. Xue, and H. Yu, "DADA: Driver Attention Prediction in Driving Accident Scenarios," IEEE Transactions on Intelligent Transportation Systems 23, no. 6 (2022): 4959–4971, https://doi.org/10.1109/TITS.2020.3044678.

45.J. Zhang, K. Yang, and R. Stiefelhagen, "Exploring Event-Driven Dynamic Context for Accident Scene Segmentation," IEEE Transactions on Intelligent Transportation Systems 23, no. 3 (2022): 2606–2622, https://doi.org/10.1109/TITS.2021.3134828.

46.K. Grauman, A. Westbury, E. Byrne, et al., "Ego4D: Around the World in 3,000 Hours of Egocentric Video," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, https://doi.org/10.1109/CVPR52688.2022.01842.

47.H. Chen, Y. Hou, C. Qu, et al., "360+x: A Panoptic Multi-modal Scene Understanding Dataset," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01833.

48.B. Zhu, B. Lin, M. Ning, et al., "LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment," arXiv preprint, arXiv: 2310.01852, January 22, 2024, https://doi.org/10.48550/arXiv.2310.01852.

49.X. Huang, L. Shen, J. Liu, et al., "Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine," in Proceedings of the AAAI Conference on Artificial Intelligence, edited by Toby Walsh, Julie Shah, Zico Kolter, 3782–3790, AAAI Press, 2025, https://doi.org/10.1609/aaai.v39i4.32394.

50.R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, "Adaptive Mixtures of Local Experts," Neural Computation 3, no. 1 (1991): 79–87, https://doi.org/10.1162/neco.1991.3.1.79.

51.W. Fedus, B. Zoph, and N. Shazeer, "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," Journal of Machine Learning Research 23, no. 120 (2022): 1–39.

52.Z. Sheng, Z. Huang, and S. Chen, "Kinematics-Aware Multigraph Attention Network with Residual Learning for Heterogeneous Trajectory Prediction," Journal of Intelligent and Connected Vehicles 7, no. 2 (2024): 138–150, https://doi.org/10.26599/JICV.2023.9210036.

53.E. J. Hu, Y. Shen, P. Wallis, et al., "LoRA: Low-Rank Adaptation of Large Language Models," paper presented at the International Conference on Learning Representations, April 25–29, 2022.

54.Z. Chen, J. Wu, W. Wang, et al., "InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02283.

55.A. Yang, B. Yang, B. Zhang, et al., "Qwen2.5 Technical Report," arXiv preprint, arXiv: 2412.15115, January 3, 2025, https://doi.org/10.48550/arXiv.2412.15115.

56.The Vicuna Team, "Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality," Accessed January 14, 2026. https://lmsys.org/blog/2023-03-30-vicuna.

57.K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, "BLEU: A Method for Automatic Evaluation of Machine Translation," in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, edited by Pierre Isabelle, Eugene Charniak, and Dekang Lin, 311–318, ACL, 2002, https://doi.org/10.3115/1073083.1073135.

58.C. Y. Lin, "ROUGE: A Package for Automatic Evaluation of Summaries," paper presented at the Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, July 25–26, 2004.

59.T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, "BERTScore: Evaluating Text Generation with BERT," arXiv preprint, arXiv: 1904.09675, Februray 24, 2020, https://doi.org/10.48550/arXiv.1904.09675.

60.H. Liu, C. Li, Q. Wu, and Y. J. Lee, "Visual Instruction Tuning," paper presented at the Thirty-seventh Conference on Neural Information Processing Systems, New Orleans, USA, December 10–16, 2023.

61.Z. Li, Q. Xu, D. Zhang, et al., "GroundingGPT: Language Enhanced Multi-modal Grounding Model," arXiv preprint, arXiv: 2401.06071, March 5, 2024, https://doi.org/10.48550/arXiv.2401.06071.

62.G. Bertasius, H. Wang, and L. Torresani, "Is Space-Time Attention All You Need for Video Understanding?" in Proceedings of the 38th International Conference on Machine Learning, edited by Marina Meila, Tong Zhang, 813–824, PMLR, 2021.

63.H. Zhang, X. Li, and L. Bing, "Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding," arXiv preprint, arXiv: 2306.02858, October 25, 2023, https://doi.org/10.48550/arXiv.2306.02858.

64.C. Gu, C. Sun, D. A. Ross, et al., "AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions," arXiv preprint, arXiv: 1705.08421, April 30, 2018, https://doi.org/10.48550/arXiv.1705.08421.

65.R. Xu, H. Lin, W. Jeon, et al., "WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios," arXiv preprint, arXiv: 2510.26125, November 13, 2025, https://doi.org/10.48550/arXiv.2510.26125.

Basic Information:

DOI:10.23919/CHAIN.2026.000005

Chinese Library Classification Number:

Citation Information:

Multimodal Large Language Models (MLLMs) have achieved remarkable progress across a range of vision-language tasks and demonstrate strong potential for traffic accident understanding. However, existing MLLMs in this domain primarily focus on coarse-grained image-level or video-level comprehension and often struggle to handle fine-grained visual details or localized scene components, limiting their applicability in complex accident scenarios. To address these limitations, we propose SafePLUG, a novel framework that empowers MLLMs with both pixel-level understanding and temporal grounding for comprehensive traffic accident analysis. SafePLUG supports both arbitrary-shaped visual prompts for region-aware question answering and pixel-level segmentation based on language instructions, while also enabling the recognition of temporally anchored events in traffic accident scenarios. To advance the development of MLLMs for traffic accident understanding, we curate a new dataset, SafePLUG-Bench, which contains diverse multimodal question–answer pairs with detailed pixel-level annotations and temporal event boundaries across a wide range of accident scenarios. Experimental results show that SafePLUG achieves strong performance on multiple tasks, including region-based question answering, pixel-level segmentation, temporal event localization, and accident event understanding. These capabilities lay a foundation for a fine-grained understanding of complex traffic scenes, with the potential to improve driving safety and enhance situational awareness in smart transportation systems.

quote

GB/T 7714-2015 [1] Zihao Sheng, Zilin Huang, Yansong Qu, et al. SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding[J]. Chain, 2026, 3(1): 53-72. DOI:10.23919/CHAIN.2026.000005.
MLA [1] Zihao Sheng, et al., "SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding." Chain, vol. 3, no. 1, 2026, pp. 53-72, https://doi.org/10.23919/CHAIN.2026.000005.
APA [1] Zihao Sheng, Zilin Huang, Yansong Qu, Jiancong Chen, Yuhao Luo, Yen-Jung Chen, Yue Leng, & Sikai Chen. (2026). SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding. Chain, 3(1), 53-72. https://doi.org/10.23919/CHAIN.2026.000005
IEEE [1] Zihao Sheng, Zilin Huang, Yansong Qu, Jiancong Chen, Yuhao Luo, Yen-Jung Chen, Yue Leng, and Sikai Chen, "SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding," Chain, vol. 3, no. 1, pp. 53-72, 2026, doi: 10.23919/CHAIN.2026.000005. keywords: {multimodal large language models;safety-critical perception;traffic accident understanding;transportation safety}