ContentsFigures & Tables
1 Introduction

1 Introduction

2 Related Work

2 Related Work

2.1 Traffic accident understanding methods

2.1 Traffic accident understanding methods

2.2 Traffic accident understanding datasets

2.2 Traffic accident understanding datasets

2.3 Egocentric and multi-view video understanding

2.3 Egocentric and multi-view video understanding

3 Method

3 Method

3.1 Multimodal input encoding

3.1 Multimodal input encoding

3.1.1 Video encoder with number prompts

3.1.1 Video encoder with number prompts

3.1.2 Visual prompt encoder

3.1.2 Visual prompt encoder

3.1.3 Image and pixel encoder

3.1.3 Image and pixel encoder

3.2 Multimodal fusion via LLM

3.2 Multimodal fusion via LLM

3.3 Decoders for textual and pixel output

3.3 Decoders for textual and pixel output

3.4 Training strategy

3.4 Training strategy

3.5 SafePLUG-Bench

3.5 SafePLUG-Bench

3.5.1 Dataset construction

3.5.1 Dataset construction

3.5.2 Dataset statistics

3.5.2 Dataset statistics

4 Experimental

4 Experimental

4.1 Experimental setting

4.1 Experimental setting

4.1.1 Models and training configuration

4.1.1 Models and training configuration

4.1.2 Evaluation metrics

4.1.2 Evaluation metrics

4.2 Performance evaluation

4.2 Performance evaluation

4.2.1 Region-level question answering

4.2.1 Region-level question answering

4.2.2 Pixel grounding

4.2.2 Pixel grounding

4.2.3 Accident description

4.2.3 Accident description

4.2.4 Temporal localization

4.2.4 Temporal localization

4.3 Qualitative analysis

4.3 Qualitative analysis

4.3.1 Accident description

4.3.1 Accident description

4.3.2 Temporal localization

4.3.2 Temporal localization

4.3.3 Region-level question answering

4.3.3 Region-level question answering

4.3.4 Region-level segmentation

4.3.4 Region-level segmentation

4.4 Ablation study

4.4 Ablation study

4.4.1 Effectiveness of dual-LoRA training

4.4.1 Effectiveness of dual-LoRA training

4.4.2 Effectiveness of model components

4.4.2 Effectiveness of model components

4.5 Failure cases analysis

4.5 Failure cases analysis

5 Conclusions

5 Conclusions

References

References

SafePLUG: empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding

Zihao Sheng1Zilin Huang1Yansong Qu2Jiancong Chen2Yuhao Luo1Yen-Jung Chen3Yue Leng4Sikai Chen1
1. Department of Civil and Environmental Engineering, University of Wisconsin–Madison, Madison WI 53706, USA
2. Lyles School of Civil and Construction Engineering, Purdue University, West Lafayette IN 47907, USA
3. Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette IN 47907, USA
4. Google, Sunnyvale CA 94089, USA
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress across a range of vision-language tasks and demonstrate strong potential for traffic accident understanding. However, existing MLLMs in this domain primarily focus on coarse-grained image-level or video-level comprehension and often struggle to handle fine-grained visual details or localized scene components, limiting their applicability in complex accident scenarios. To address these limitations, we propose SafePLUG, a novel framework that empowers MLLMs with both pixel-level understanding and temporal grounding for comprehensive traffic accident analysis. SafePLUG supports both arbitrary-shaped visual prompts for region-aware question answering and pixel-level segmentation based on language instructions, while also enabling the recognition of temporally anchored events in traffic accident scenarios. To advance the development of MLLMs for traffic accident understanding, we curate a new dataset, SafePLUG-Bench, which contains diverse multimodal question–answer pairs with detailed pixel-level annotations and temporal event boundaries across a wide range of accident scenarios. Experimental results show that SafePLUG achieves strong performance on multiple tasks, including region-based question answering, pixel-level segmentation, temporal event localization, and accident event understanding. These capabilities lay a foundation for a fine-grained understanding of complex traffic scenes, with the potential to improve driving safety and enhance situational awareness in smart transportation systems.
Keywords: multimodal large language models; safety-critical perception; traffic accident understanding; transportation safety
Received: 2025-11-06

1 Introduction

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in understanding and reasoning over visual and linguistic information, enabling a wide range of applications from visual question answering (QA) to video analysis [1–3]. Building on these successes, researchers have increasingly explored the potential of MLLMs within intelligent transportation systems and autonomous driving [4–8], particularly in advancing traffic accident understanding [9–11]. By jointly processing information across multiple modalities, MLLMs offer a promising paradigm for analyzing traffic incidents and answering complex queries [12–15]. These capabilities can be valuable in a variety of real-world traffic scenarios. For example, drivers may benefit from real-time accident interpretation and warning feedback, while analysts and planners can use them to assist in post-accident review, liability assessment, and identifying common failure patterns [16–21].

Understanding traffic accidents often requires fine-grained, pixel-level comprehension to ensure accurate identification of critical objects, spatial relationships, and impact regions. However, existing MLLMs in this domain, such as RoadSocial [9] and EchoTraffic [10], primarily operate at a coarse granularity, focusing on global scene understanding at the image or video level while lacking the ability to localize and reason about specific regions involved in an accident. This coarse granularity hinders their ability to capture nuanced visual cues that are essential for accurate accident interpretation. In contrast, pixel-level MLLMs such as LISA [22], Sa2VA [23], and PixelLM [24] are capable of processing fine-grained visual details, with the potential to support more accurate segmentation of collision areas and involved agents. However, these methods primarily focus on general-purpose image or video segmentation and lack temporal grounding capabilities essential for traffic accident analysis. Furthermore, by leveraging arbitrary-shaped pixel-level visual prompts as input, the model can be better guided to attend to semantically and contextually relevant areas, enhancing its ability to filter out irrelevant background and improving accuracy on region-sensitive tasks [25–27].

Another critical aspect of traffic accident understanding is temporal grounding, which refers to identifying the start and end times of specific events within a video. In traffic accident understanding, knowing exactly when the accident occurs is essential for supporting fine-grained accident phase analysis. By distinguishing between pre-, during-, and post-accident phases, the model can separate normal driving behavior from abnormal actions, thereby enabling effective warnings [17, 28]. While recent video-based MLLMs like Video-LlaVA [29] have made substantial progress in recognizing what happens in a scene, they often struggle to determine when it happens [30, 31]. This limitation arises because most models are trained to align visual content with language, focusing on semantic understanding rather than temporal localization [32]. The RoadSocial benchmark [9] also highlights this gap by evaluating models on predicting event boundaries in traffic videos, and finds that even strong MLLMs often produce implausible time spans. These findings underscore the importance of equipping MLLMs with robust temporal grounding abilities to ensure reliable accident interpretation.

To bridge these gaps and advance the application of MLLMs in traffic accident understanding, we propose SafePLUG, a novel framework that empowers MLLMs with both Pixel-Level Understanding and temporal Grounding capabilities. To the best of our knowledge, SafePLUG is among the first framework that jointly supports pixel-level segmentation, region-aware QA, and temporal event localization within a single architecture for traffic accident understanding. For pixel-level understanding, SafePLUG incorporates a visual prompt encoder that extracts region-aware features from arbitrary-shaped visual prompts and aligns them with the language instructions. Unlike prior methods that rely solely on bounding boxes, SafePLUG is among the first to introduce arbitrary-shaped masks as visual prompts for traffic accident analysis. We further extend the LLM vocabulary with a special <SEG> token, whose hidden embedding is utilized by an off-the-shelf Segment Anything Model (SAM)-based decoder [33] to produce pixel-wise segmentation masks. For temporal grounding, we incorporate a lightweight number prompt mechanism inspired by [32], in which unique numeric indicators are overlaid on video frames to convey temporal positions. By treating these numbers as visual cues, the model is guided to associate semantic events with specific temporal segments. Importantly, number prompts integrate seamlessly into the video input without modifying the model architecture or requiring additional training objectives. As illustrated in Figure 1, SafePLUG exhibits remarkable capabilities across four key tasks: accident description, temporal localization, region-level QA, and pixel-level grounding. To support the development and evaluation of such models, we construct SafePLUG-Bench, among the first benchmark datasets for traffic accident understanding that supports both region-level QA and pixel-level grounding QA, along with temporal localization and accident description tasks.

Figure 1 SafePLUG supports both image/video-level and pixel-level understanding through accident description, temporal localization, region-level question answering, and pixel-level grounding, enabling comprehensive traffic accident analysis

In summary, our contributions are as follows:

  • We propose SafePLUG, a novel framework that equips MLLMs with both pixel-level understanding and temporal grounding capabilities, enabling fine-grained reasoning over complex traffic accident scenarios through the integration of visual and number prompts.

  • We curate SafePLUG-Bench, a new benchmark dataset for traffic accident understanding. To the best of our knowledge, this dataset is among the first in this domain to support both region-level QA and pixel-level grounding QA.

  • Extensive experiments across multiple tasks, including region-level QA, pixel-level segmentation, temporal event localization, and accident event understanding, demonstrate the superior performance of SafePLUG. All code, dataset, and model checkpoints will be released to facilitate future research.

2 Related Work

2.1 Traffic accident understanding methods

Traffic accident understanding involves identifying key agents, detecting anomalies, and interpreting causal and temporal dynamics in complex driving scenarios [28, 34–36]. Earlier approaches primarily used Convolutional Neural Network (CNN)-based models to classify accidents or detect behavioral phases from visual inputs [37, 38]. However, these models lack high-level semantic reasoning and cannot answer open-ended questions, such as "Analyze why the accident happened?".

Recently, MLLMs have been introduced to enhance traffic accident understanding. EchoTraffic [10] incorporates audio cues to improve the anomaly reasoning capabilities of MLLMs. RoadSocial [9] demonstrates that fine-tuning general video MLLMs on their proposed dataset improves model performance in road event comprehension, but it does not incorporate explicit mechanisms for pixel-level grounding or temporal localization. TABot [11] combines functional and instruction tuning for MLLMs and leverages bounding box information to provide spatial grounding of accident regions and involved agents. However, it is limited to rectangular regions and does not support arbitrary-shaped visual prompts or pixel-level segmentation. Guan et al. [39] further introduce a domain-enhanced dual-branch framework that integrates multimodal features through large models such as GPT-4o and Long-CLIP. Their model jointly learns visual–temporal representations and domain knowledge embeddings, enabling efficient and interpretable accident anticipation.

In contrast, SafePLUG extends MLLMs to support diverse input and output modalities, including arbitrary-shaped visual prompts, number prompts for temporal grounding, and pixel-level segmentation masks. These capabilities enable fine-grained spatial reasoning and temporally grounded accident analysis that prior methods do not offer.

2.2 Traffic accident understanding datasets

Early traffic accident datasets are primarily designed to support tasks such as accident detection, accident type classification, and identification of involved objects [40, 41]. The A3D dataset [42] provides annotations for accident categories, bounding boxes of involved objects, and timestamps indicating when accidents are identified. DoTA [28] extends A3D by incorporating more videos and richer annotations, including anomaly types, related objects, and tracking IDs. The CCD dataset [43] further offers accident causes for each video sequence, while DADA [44] explores the role of driver attention in traffic accident prediction by collecting eye-gaze data. Based on this, DADA-Seg [45] refines a subset of 313 video sequences with fine-grained segmentation masks for semantic objects. Although these datasets have significantly advanced visual-based traffic accident analysis, they primarily support coarse-grained tasks and lack detailed language annotations.

In recent years, an increasing number of datasets have emerged to support traffic accident understanding through language-based QA. SUTD-TrafficQA [16] is among the first large-scale benchmark in this domain, offering six types of video-QA pairs, such as accident description, forecasting, and reasoning. MM-AU [17] provides textual annotations that cover three key aspects of traffic accidents: causality, prevention strategies, and accident types. TAU-106K [11] advances this direction with questions requiring temporal localization and spatial grounding, where textual answers include timestamps and bounding box coordinates. The RoadSocial dataset [9] further broadens the task scope with diverse video QAs for general road events. Meanwhile, AV-TAU [10] enriches the multimodal context of traffic accident scenarios by incorporating audio signals.

Our dataset, SafePLUG-Bench, further advances the field by uniquely supporting both region-level QA and pixel-level grounding QA. It is densely annotated with segmentation masks and includes over 220 K high-quality multimodal QA pairs across diverse accident scenarios. A comprehensive comparison with existing datasets is summarized in Table 1.

Table 1 Comparison of existing traffic accident understanding datasets with ours
Dataset Year Frame Bbox Mask TG Region QA PGQA QA Pairs
A3D [42] 2019 208 K √ – √ – – –
CCD [43] 2020 75 K √ – √ – – –
DADA [44] 2021 658 K √ √ √ – – –
DADA-Seg [45] 2021 12 K – – √ – – –
SUTD-TrafficQA [16] 2021 1.90 M – – √ – – 62 K
DoTA [28] 2022 732 K √ – √ – – –
MM-AU [17] 2024 2.19 M √ – √ – – 58 K
TAU-106K [11] 2025 – √ – √ – – 332 K
RoadSocial [9] 2025 14 M – – √ – – 260 K
AV-TAU [10] 2025 3.16 M – – √ – – 149 K
SafePLUG-Bench (Ours) 2026 2.26 M √ √ √ √ √ 220 K

2.3 Egocentric and multi-view video understanding

Recent large-scale datasets such as Ego4D [46] and 360 + X [47] have significantly advanced egocentric and panoramic video understanding by emphasizing long-term temporal modeling and fine-grained spatial–temporal grounding. Ego4D focuses on broad daily-life activities captured from a first-person perspective. In contrast, 360 + X explores panoramic and multi-view settings, highlighting the benefits of holistic scene coverage and multi-modal alignment for general action understanding.

SafePLUG is complementary to existing egocentric and panoramic video understanding efforts but targets a distinct safety-critical domain. While also operating on egocentric video, SafePLUG focuses on traffic accident understanding, which requires not only recognizing actions or objects, but also reasoning about causal factors, risk regions, and the temporal evolution of events. SafePLUG emphasizes diversity within the driving domain, and thus opens new directions for safety-aware multimodal reasoning in first-person driving videos.

3 Method

Figure 2 illustrates the overall architecture of the proposed SafePLUG framework, which extends MLLMs with enhanced pixel-level understanding and temporal grounding abilities for comprehensive traffic accident analysis. Unlike prior approaches that focus primarily on global scene comprehension, SafePLUG introduces fine-grained spatial reasoning and temporally aligned perception to capture subtle accident dynamics. In the following subsections, we elaborate on each module in detail.

Figure 2 Overview of SafePLUG. The model takes as input multiple modalities, including video frames with number prompts, image-level context, and user-defined visual prompts, and unifies them with language prompts through an LLM backbone. The features are then decoded into either natural language answers or binary segmentation masks

3.1 Multimodal input encoding

3.1.1 Video encoder with number prompts

Traffic accident analysis critically depends on identifying precise temporal boundaries, including the separation of pre-accident, impact, and post-accident phases. General video MLLMs typically perform temporal reasoning in a descriptive manner without access to explicit frame identity, leading to imprecise temporal predictions. To address this challenge, SafePLUG incorporates an explicit temporal encoding mechanism built upon a pretrained video encoder, enhanced with number prompts that provide direct frame-level identity cues.

To represent the temporal evolution of traffic scenarios, we denote a video sequence as V∈RT×H×W×s, where T is the number of uniformly sampled frames from the raw footage, H and W are the height and width of each video frame. The entire sequence V is fed into a pretrained video encoder εvid from LanguageBind [48] to extract spatiotemporal representations: v = ℰ vid ( V ) ,   V ∈ R T × L v × d v  (1)where Lv denotes the number of tokens per frame, and dv is the hidden dimension of the video encoder. The resulting features V are then transformed by a projection layer Pvid(v) to align with the input space of the language model. To further strengthen the model's temporal grounding capability, we introduce a lightweight yet effective enhancement, i.e., number prompts [32]. During preprocessing, a numerical indicator corresponding to the frame index is visually overlaid onto each frame, providing cues for the model to identify the temporal order of events. The numerical markers are strategically placed (e.g., top-right corner) to preserve visual content integrity. This simple augmentation does not alter the model architecture or objective but has been shown to guide the model in accurately associating semantic events with their temporal boundaries.

3.1.2 Visual prompt encoder

Accident understanding usually requires focusing on specific agents or interaction zones rather than the full scene. In contrast to general MLLMs that rely on implicit attention over global visual tokens, SafePLUG treats localized regions as explicit inputs to the language model. To enable fine-grained spatial reasoning within complex traffic scenes, SafePLUG incorporates a visual prompt encoder that encodes user-defined regions of interest (ROIs) such as polygons or arbitrary-shaped masks. Given an image I∈RH×W×3 and a corresponding binary mask M∈{0, 1}H×W that specifies the visual prompt region, we extract intermediate feature maps F from the pretrained image encoder εimg. The region-aware feature representation is then obtained through masked average pooling: r = 1 | M ′ | 1 ∑ x , y M ′ x , y   F x , y (2)where M' denotes the downsampled version of M aligned with F. The embedding r is projected into the LLM input space through a learnable projection layer Proi. This representation encodes localized visual semantics conditioned on the prompt region, allowing the model to attend to contextually relevant areas selectively. By combining explicit spatial prompts with language cues, the visual prompt encoder provides a flexible and effective mechanism for grounding linguistic instructions in specific image regions.

3.1.3 Image and pixel encoder

To capture both holistic scene context and fine-grained spatial cues, SafePLUG integrates two complementary visual encoders. First, a pretrained image encoder εimg extracts visual features from the image I: i = ℰ img ( I ) ,   i ∈ R L i × d i  (3)where Li denotes the number of tokens, and di is the hidden dimension of the image encoder. These high-level embeddings encode the global semantics of the scene and serve as the contextual foundation for downstream reasoning tasks such as region-level question answering.

Second, a pixel-level encoder εpix built upon the SAM [33] extracts dense spatial features from I: p = ℰ pix ( I ) ,   p ∈ R H p × W p × d p  (4)where p denotes the pixel-level embeddings capturing fine-grained spatial cues, Hp, Wp and dp are the height, width and channel number of p. These features are subsequently utilized by the SAM-based decoder Dpix to generate segmentation masks aligned with language instructions. Together, the image encoder εimg and pixel encoder εpix enable SafePLUG to reason over global semantics and local spatial details, ensuring both contextual understanding and precise localization in complex traffic scenes.

3.2 Multimodal fusion via LLM

SafePLUG integrates multimodal representations within a unified large language model to enable joint spatial–temporal reasoning. The visual features extracted from the video, image, and region encoders are each projected into the language embedding space through their respective projectors Pvid, Pimg, and Proi. During tokenization, we insert special placeholder tokens (<video>, <image>, and <region>) into the input text sequence. The projected visual embeddings then replace these placeholders to form a multimodal input sequence: X = T text | 〈 video 〉 → P vid ( v ) , 〈 image 〉 → P img ( i ) , 〈 region 〉 → P roi ( r ) (5)

The interleaved sequence X is then processed by the language model fllm, which jointly attends to textual and visual information for cross-modal reasoning across spatial regions, temporal sequences, and linguistic context. This fusion design allows SafePLUG to achieve coherent, context-aware interpretation of traffic scenes from multiple modalities.

3.3 Decoders for textual and pixel output

Traffic accidents often involve fine-grained spatial interactions, such as partial vehicle overlap or contact regions, which cannot be reliably captured by image-level descriptions or bounding boxes. To enable precise localization of accident-relevant regions, SafePLUG produces two types of outputs (i.e., natural language responses and pixel-level segmentation masks) within a unified decoding framework. For textual tasks, the language model directly generates responses through its language head, producing token probabilities conditioned on the multimodal input sequence. For pixel-level tasks, following prior works [22, 23, 49], we extend the LLM vocabulary with a special token <SEG>, whose hidden representation h<seg> serves as a query vector for segmentation. After processing the multimodal input, h<seg> is transformed by a learnable projector Pseg to form a segmentation prompt, which is jointly provided with the dense spatial features p extracted by the SAM pixel encoder εpix as inputs to the SAM decoder Dpix. The decoder then produces the final binary mask aligned with the textual instruction. This design allows SafePLUG to seamlessly switch between generating descriptive explanations and grounded visual outputs.

3.4 Training strategy

Accident-related tasks require both high-level causal explanation and low-level spatial precision, which are often competing objectives when optimized jointly. General MLLMs typically rely on a single adaptation pathway, causing interference between heterogeneous supervision signals. To address this challenge, SafePLUG adopts a modular training strategy inspired by the Mixture-of-Experts (MoE) paradigm [50–52], where two specialized sets of Low-Rank Adaptation (LoRA) parameters [53] are optimized to handle distinct output modalities. Rather than relying on a single adaptation pathway for all tasks, SafePLUG maintains a shared multimodal backbone and two lightweight, task-specific LoRA branches: one dedicated to natural language generation and the other to pixel-level segmentation. This design allows each branch to specialize in its respective modality while preserving shared visual–linguistic representations through a unified backbone.

The first LoRA branch focuses on text-oriented tasks, including accident description, region-level QA, and temporal localization. During this phase, the LoRA adapters are integrated throughout both the attention and feed-forward layers of the language model to enable parameter-efficient fine-tuning across the entire transformer structure. A small subset of projection and normalization layers connecting visual and textual representations is also softly updated to enhance cross-modal alignment. The model is optimized using a standard cross-entropy loss for text generation: ℒ text = − ∑ i = 1 T log P ( y i | y < i , X )    (6)where yi denotes the i-th target token and X represents the multimodal input sequence.

The second branch is tailored for mask generation and spatial grounding. Building upon the pretrained backbone from the language branch, LoRA adapters are selectively applied only to the feed-forward layers to refine spatial reasoning while keeping the attention structure intact. In parallel, a compact set of segmentation-specific modules that map hidden language features to dense pixel embeddings is fine-tuned. The optimization objective combines binary cross-entropy (BCE) and Dice losses to balance pixel coverage accuracy and boundary precision: ℒ seg = λ BCE ℒ BCE + λ DICE  ℒ DICE (7)where λBCE and λDICE are weighting coefficients to balance pixel coverage and boundary precision.

After training, the two LoRA branches can be independently loaded or jointly combined depending on the target task. The language-oriented branch governs open-ended textual reasoning and explanation generation, whereas the mask-oriented branch provides pixel-level grounding and segmentation capabilities. Since both branches share the same multimodal encoders and tokenization pipeline, SafePLUG can dynamically activate the corresponding LoRA module according to the task type, enabling seamless switching between linguistic and spatial reasoning without performance interference or retraining.

3.5 SafePLUG-Bench

Existing datasets for traffic accident understanding primarily focus on high-level visual QA but often lack fine-grained annotations required for tasks such as region-level QA and pixel-level grounding QA. These two capabilities are essential for enabling models to reason about specific regions involved in an incident and to localize accident-related semantics at the pixel level. To bridge this gap, we construct SafePLUG-Bench, a new benchmark dataset that supports both region QA and pixel-level grounding QA, in addition to standard accident description and temporal grounding tasks.

3.5.1 Dataset construction

SafePLUG-Bench builds upon two existing benchmarks: DoTA [28] and MM-AU [17]. As illustrated in Figure 3, we adopt a semi-automated annotation pipeline that combines state-of-the-art multimodal foundation models with expert human verification to ensure both scalability and annotation fidelity.

Figure 3 Semi-automated data annotation pipeline leveraging MLLMs and SAM for generating region-level descriptions, accident narratives, and segmentation masks

For region-level QA, we utilize the bounding box annotations from the original datasets. For each region, we generate two independent candidate descriptions: one by feeding the full image with overlaid bounding boxes to InternVL3-78B [54], and another by inputting the cropped region to Qwen2.5-VL-72B [55]. The resulting descriptions are then cross-verified by Qwen2.5-72B [55] to ensure semantic consistency, and inconsistent pairs are discarded. To maintain stylistic consistency, only the InternVL3-78B descriptions are retained as final annotations, while the Qwen2.5-VL-72B outputs serve purely for verification purposes. This process produces region-level answers that emphasize spatial context and local object relationships.

For accident description, we employ InternVL3-78B [54] to generate comprehensive narrative explanations of traffic incidents. Each generation process integrates visual information from video frames with auxiliary textual cues, i.e., annotated accident causes extracted from the original datasets. For DoTA, which provides detailed bounding boxes of involved agents, we further incorporate their verified region descriptions as additional contextual inputs. This multimodal prompting design allows the model to reason over both visual and causal dimensions, producing more coherent and causally grounded accident narratives. To ensure reliability and linguistic consistency, a subset of generated outputs was manually examined, based on which we iteratively refined the input prompt before finalizing the full dataset annotation pipeline.

For temporal localization, we treat each annotated accident cause as a textual query and its corresponding timestamp as the answer, establishing explicit temporal boundaries for each event. For pixel-level grounding QA, we employ SAM to automatically generate segmentation masks based on bounding box annotations, followed by human filtering to remove coarse or fragmented masks. Each valid region is then paired with its verified region description as the question and the corresponding mask as the answer. To enable holistic reasoning, all accident-related masks within a single frame are also merged into an aggregated "accident-region" mask that identifies the full spatial extent of the incident. Additional details on the data annotation pipeline, prompt design, and human verification procedures are provided in Supporting Information A.

3.5.2 Dataset statistics

In total, SafePLUG-Bench comprises over 220 K high-quality multimodal QA pairs covering diverse real-world accident scenarios. To ensure fair evaluation across all tasks, we randomly sample 500 QA pairs for each of the four subtasks. Figure 4 summarizes the overall statistical characteristics of the dataset.

Figure 4 Data statistics of SafePLUG-Bench. (A) Length distribution in accident description. (B) Distribution of video lengths. (C) Ratio of accident time window. (D) Clustered bounding boxes across all regions. (E) Object type distribution across all regions. (F) Accident type distribution

Figure 4A shows the word length distribution of answers in the accident description task. Most responses contain between 150 and 300 words. This indicates that the generated descriptions provide sufficient detail to capture causal reasoning and contextual information. Figure 4B illustrates the distribution of video lengths in frames. Most clips contain between 80 and 160 frames, providing moderate-length driving scenes that include both normal and accident phases. The ratio of the annotated accident window to the total video duration is shown in Figure 4C. The majority of accident windows occupy 30%–60% of the full sequence, such that each sequence includes adequate pre- and post-event context.

Figure 4D visualizes the spatial distribution of bounding boxes in the region QA task. We cluster all bounding box coordinates and then randomly sample 150 boxes for visualization. The spatial density map shows that the annotated regions are distributed broadly across the frame, with a mild concentration in central and lower areas corresponding to vehicles and road surfaces. Each frame contains an average of 5.6 bounding boxes. The object category distribution within all annotated bounding boxes is shown in Figure 4E. The majority of annotated objects are cars, followed by trucks and pedestrians, which together account for the vast majority of traffic entities in urban scenes.

Finally, Figure 4F summarizes the distribution of accident types across all 220 K QA pairs. We follow the accident category introduced in the DoTA [28]. The dataset covers eight major accident types: (1) ST: collision with a vehicle that starts, stops, or is stationary; (2) AH: collision with a vehicle moving ahead or waiting; (3) LA: collision with a vehicle moving laterally in the same direction; (4) OC: collision with an oncoming vehicle; (5) TC: collision with a vehicle turning into or crossing a road; (6) VP: collision with a pedestrian; (7) VO: collision with an obstacle; and (8) OO: loss of control. Notably, non-accident samples (NM) account for 43.3% of the dataset. These samples may include various normal driving scenarios such as routine traffic flow, emergency vehicle operations without collisions, or traffic control situations. This substantial proportion of non-accident samples is crucial for training the model to distinguish between genuine collision-based accidents and routine driving scenarios.

4 Experimental

4.1 Experimental setting

4.1.1 Models and training configuration

We employ the pretrained video and image encoders from LanguageBind [48], which remain frozen during training to preserve their multimodal alignment. The projection layers used for mapping visual embeddings into the language space are initialized from Video-LlaVA [29]. For each video, we uniformly sample 8 frames as input to the video encoder to capture temporal context with balanced computational efficiency. The backbone language model is Vicuna-7B v1.5 [56], which serves as the main reasoning component for multimodal fusion and response generation. For pixel-level segmentation, we adopt the SAM ViT-H [33] model, where the pixel encoder is frozen, and only the decoder is fine-tuned to learn safety-critical spatial segmentation. Training is conducted on 8xA100 GPUs (80 GB each). Additional training implementation details are provided in Supporting Information B.

4.1.2 Evaluation metrics

For text-based tasks, we adopt BLEU-1 [57], ROUGE-1 [58], and BERTScore F1 [59] to measure linguistic similarity against ground-truth annotations. Following prior work [9, 10], we further prompt GPT-3.5 to assess the generated responses in terms of consistency, reasonableness, and level of detail. More details of the GPT-3.5 prompting template are provided in Supporting Information B. For pixel-level grounding and temporal localization tasks, we report the Average Precision (AP) at thresholds of 30, 50, and 70 (denoted as AP@30, AP@50, and AP@70), together with mean Intersection over Union (mIoU). All metric scores are linearly scaled to a range of 0–100 for consistent comparison across tasks.

4.2 Performance evaluation

We evaluate SafePLUG across four key tasks: region-level QA, pixel-level grounding, accident description, and temporal localization. All experiments are conducted on the test split of our proposed SafePLUG-Bench. Tables 2, 3 present quantitative comparisons against a diverse set of existing MLLMs. These baselines include both general-purpose vision-language models and specialized frameworks designed for spatiotemporal understanding. Given the substantial computational requirements of large multimodal models, all reported results are obtained from a single evaluation run using a fixed random seed to ensure reproducibility and fairness.

Table 2 Performance comparison on region QA and pixel grounding. All metric scores range from 0 to 100, with the best performance highlighted in bold
Method Param. Region QA Pixel grounding
BLEU Rouge BERT GPT AP@30 AP@50 AP@70 mIoU
Qwen2.5-VL [55] 72 B 18.46 27.91 82.83 51.84 51.60 47.50 40.80 44.17
InternVL3 [54] 78 B 19.89 27.87 82.05 71.26 5.10 3.90 2.90 4.17
LLaVA [60] 7 B 5.02 13.84 81.54 26.02 23.30 16.80 13.60 18.07
GroundingGPT [61] 7 B 0.01 5.90 80.82 28.46 13.50 12.30 10.40 11.95
LISA [22] 7 B 2.99 10.85 78.93 13.80 21.00 16.50 14.60 17.61
Sa2VA [23] 8 B 0.54 3.90 78.55 13.20 68.80 63.50 56.20 58.74
SafePLUG (Ours) 7 B 34.54 40.15 86.09 65.13 74.30 68.10 59.30 64.07
Table 3 Performance comparison on accident description and temporal localization. All metric scores range from 0 to 100, with the best performance highlighted in bold
Method Param. Accident description Temporal localization
BLEU Rouge BERT GPT AP@30 AP@50 AP@70 mIoU
Qwen2.5-VL [55] 72 B 15.94 29.98 83.37 47.11 19.40 11.80 3.00 11.24
InternVL3 [54] 78 B 1.98 8.38 80.11 19.20 1.40 0.60 0.00 2.44
Video-LLaVA [29] 7 B 3.98 17.11 81.78 19.44 44.00 17.80 3.00 25.93
GroundingGPT [61] 7 B 3.66 13.74 81.28 19.12 3.00 0.20 0.00 2.85
TimeChat [31] 7 B 0.63 9.78 80.93 17.75 1.60 0.00 0.00 1.44
RoadSocial [9] 8 B 0.04 11.39 82.02 30.39 7.00 2.20 0.40 5.66
SafePLUG (Ours) 7 B 30.29 38.31 85.49 66.47 65.60 45.40 19.60 43.18

4.2.1 Region-level question answering

For models that do not support region-specific visual prompts, we incorporate bounding box coordinates into the input prompt as additional textual descriptions. As shown in Table 2, SafePLUG achieves the best overall performance across all textual metrics (BLEU, ROUGE, and BERTScore) and ranks competitively on GPT-based evaluation. Compared with general-purpose vision–language models such as Qwen2.5-VL and LLaVA, SafePLUG yields absolute gains of over 15–20 points on BLEU and ROUGE, indicating stronger alignment between textual responses and reference region descriptions. These improvements stem from SafePLUG's ability to directly attend to arbitrary-shaped visual prompts, which guide the language model to focus on semantically relevant subregions rather than relying on global scene context. Notably, SafePLUG attains a GPT score of 65.13, approaching the best 71.26 achieved by InternVL3 while using an order-of-magnitude smaller model (7 B vs. 78 B). These results suggest that incorporating explicit region-level visual grounding is more effective than simply scaling model size or relying solely on textual heuristics.

4.2.2 Pixel grounding

Among all evaluated baselines, only LISA and Sa2VA natively support mask-level outputs. For fair comparison, we prompt other baseline models to output bounding box coordinates and convert them into masks using SAM. As shown in Table 2, both LISA and Sa2VA exhibit limited generalization to complex accident scenes, achieving modest mean IoU values of 17.61 and 58.74, respectively. In contrast, SafePLUG achieves the highest performance across all metrics, with AP@50 = 68.10 and mIoU = 64.07, outperforming Sa2VA by over 5 points.

These improvements stem from SafePLUG's design that tightly couples pixel-level reasoning with language understanding. By introducing the special <SEG> token into the LLM vocabulary and projecting its hidden representation to the SAM decoder input, SafePLUG enables the language model to produce segmentation-aware embeddings that align linguistic cues with spatial details. This explicit coupling allows SafePLUG to delineate fine-grained accident-related regions, such as overlapping vehicles and collision impact zones, which existing MLLMs often miss due to their coarse spatial grounding.

Furthermore, the consistent gains across all AP thresholds (AP@30–70) demonstrate that SafePLUG not only improves overall segmentation accuracy but also enhances robustness under stricter localization criteria. These results confirm that pixel-level understanding in traffic accident scenarios remains challenging even for advanced MLLMs, and that SafePLUG-Bench provides a valuable benchmark for fine-grained spatial grounding under realistic traffic conditions.

4.2.3 Accident description

In the accident description task, SafePLUG demonstrates clear advantages in generating coherent and causally grounded narratives of traffic incidents. As shown in Table 3, SafePLUG surpasses all baselines across BLEU, ROUGE, BERTScore, and GPT-based metrics, achieving a GPT score of 66.47, which is more than double that of RoadSocial and significantly higher than large-scale models such as Qwen2.5-VL and InternVL3. The strong performance under both lexical (BLEU and ROUGE) and semantic (BERTScore and GPT) measures indicates that SafePLUG not only produces text closely aligned with reference annotations but also captures deeper causal and contextual semantics.

These improvements primarily stem from two design aspects of SafePLUG. First, our LoRA-based training pipeline employs a language-oriented adapter configuration that strengthens cross-modal alignment for causal narrative generation, enhancing coherence and factual grounding in accident descriptions. Second, the proposed SafePLUG-Bench provides richer causal and contextual supervision than prior datasets, incorporating diverse accident causes, object interactions, and environmental cues. This diversity enables SafePLUG to generalize beyond surface-level descriptions and generate detailed explanations that correctly attribute accident responsibility or identify precipitating maneuvers. The consistent improvements across all evaluation metrics validate the role of SafePLUG-Bench as a challenging and comprehensive benchmark for measuring causal and contextual understanding in multimodal accident analysis.

4.2.4 Temporal localization

SafePLUG exhibits strong temporal localization capability, outperforming all baselines on SafePLUG-Bench by a considerable margin. As shown in Table 3, SafePLUG achieves an AP@50 of 45.40 and an mIoU of 43.18, substantially exceeding the best competing model, Video-LLaVA. These results confirm that SafePLUG can accurately identify the onset and offset of accident events, maintaining robustness even under stricter temporal thresholds (AP@70 = 19.60). The key factor behind this improvement is the use of number prompts, which provide effective temporal cues without modifying the model architecture. By embedding numerical indicators directly into video frames, the model learns to associate language expressions such as "when the collision occurred" with specific temporal indices. This alignment between temporal order and semantic content enables the language model to infer accident boundaries with higher precision and temporal coherence.

Although prior approaches such as GroundingGPT and TimeChat are designed for temporal grounding, they underperform on SafePLUG-Bench, likely due to limited exposure to complex traffic accident scenarios during training. In contrast, SafePLUG benefits from the fine-grained temporal annotations and diverse event patterns in SafePLUG-Bench, allowing it to generalize across heterogeneous accident types and accurately capture transitions between pre-accident, impact, and post-accident phases.

4.3 Qualitative analysis

Beyond the quantitative results, we provide qualitative analysis in Figures 1, S8–12 to visually illustrate the strengths of SafePLUG across all four tasks. These case studies demonstrate how SafePLUG achieves a fine-grained and contextually consistent understanding of complex traffic accident scenes, accurately identifying involved agents, spatial regions, and temporal boundaries that baseline models often overlook or misinterpret.

4.3.1 Accident description

Figure S8 presents a qualitative comparison of accident description outputs from baseline models and SafePLUG. We select Qwen2.5-VL and RoadSocial for comparison, as they achieve the highest GPT-based evaluation scores among all baselines. As shown in the figure, baseline models tend to produce scene-level summaries that omit critical causal relationships and misidentify the sequence of key events. For instance, Qwen2.5-VL focuses primarily on general road context and the near-miss interaction but fails to capture the actual collision and its underlying cause, while RoadSocial provides only a generic mention of vehicle damage without any causal interpretation.

In contrast, SafePLUG generates a concise yet causally grounded narrative that correctly identifies the roles of the involved agents and their interactions. Specifically, it infers that Object 1 attempted to turn into or cross the path of Object 2 without ensuring a clear path, and attributes the resulting collision to poor road conditions (icy and slippery surface) that compromised vehicle control and stopping distance. This explanation aligns closely with the ground-truth annotation, reflecting SafePLUG's ability to integrate spatial cues, contextual reasoning, and linguistic coherence within a unified multimodal framework.

4.3.2 Temporal localization

As illustrated in Figure S9, we visualize the accident phase boundaries predicted by SafePLUG and baseline models. Compared to TimeChat, GroundingGPT, and RoadSocial, which often produce overly short or misaligned temporal spans, SafePLUG consistently identifies both the onset and offset of accident events with high precision. Qwen2.5-VL provides closer predictions in some cases but still exhibits delayed or truncated intervals relative to the ground truth. This improvement stems from the integration of number prompts, which embed lightweight numerical indicators into video frames to provide effective temporal anchors. By learning to associate these indicators with semantic cues in the language query, SafePLUG can precisely map described events to their corresponding frame intervals.

4.3.3 Region-level question answering

Figure S10 shows qualitative examples of region QA. For fair comparison, baseline models such as Qwen2.5-VL and InternVL3 are provided with additional bounding box coordinates to specify the target region, whereas SafePLUG directly attends to the highlighted visual region via visual prompts.

Qwen2.5-VL generates a response focusing primarily on road surface markings, but fails to identify the key object (i.e., a white vehicle) within the region. InternVL3 provides a more detailed response using a chain-of-thought style, including contextual cues about traffic signs and road infrastructure. However, its output is overly verbose and occasionally redundant. In contrast, SafePLUG produces concise, spatially accurate, and semantically complete descriptions. It correctly identifies the object type, position, and relational context (e.g., a white car traveling ahead on a multi-lane urban road), aligning closely with the ground truth. This demonstrates the model's capability to integrate localized visual cues with contextual reasoning to achieve fine-grained scene understanding.

To further assess the robustness of SafePLUG under diverse visual prompt shapes and positions, Figure S11 shows two cases where distinct free-form masks are used to mark different spatial regions within the same scene. The left example highlights a police vehicle with flashing emergency lights, and the right example highlights a parked red SUV. Across both examples, SafePLUG accurately interprets each highlighted region and recognizes the corresponding objects. These results indicate that SafePLUG generalizes effectively to arbitrary prompt shapes and spatial placements, confirming the flexibility and precision of its visual prompt encoder.

4.3.4 Region-level segmentation

Figure S12 presents qualitative comparisons of pixel-level segmentation results for SafePLUG, Qwen2.5-V, and Sa2VA across five representative scenarios. The segmentation queries range from simple object-level references (e.g., "a green minivan") to complex causal descriptions involving multiple agents (e.g., "a collision between two vehicles traveling in the same direction"). Qwen2.5-VL generates bounding boxes that are post-processed into masks using SAM, whereas both Sa2VA and SafePLUG natively produce pixel-level mask generation.

Across all examples, SafePLUG yields segmentation masks that are spatially precise and semantically aligned with the described entities. It accurately delineates vehicle contours and captures subtle spatial relationships such as overlapping agents, even in cluttered or low-visibility scenes. By contrast, Sa2VA occasionally produces incomplete masks, while Qwen2.5-VL often highlights irrelevant background regions.

4.4 Ablation study

We conduct ablation studies to assess the contribution of the multi-LoRA training strategy and key model components in SafePLUG. We report BLEU-1 scores for region QA and accident description, and mIoU for pixel-level grounding and temporal localization, as summarized in Tables 4, 5.

Table 4 Effect of using different LoRA branches
Configuration Text LoRA Mask LoRA Region QA Pixel grounding Accident description Temporal localization Mean score
(a) × × 12.16 0.00 2.95 2.14 4.31
(b) √ × 35.14 0.05 30.97 41.56 26.93
(c) × √ 0.02 64.12 0.03 0.00 16.04
(d) √ √ 34.54 64.07 30.29 43.18 43.02
Table 5 Effect of the different modules
Configuration Model Region QA Pixel grounding Accident description Temporal localization Mean score
(a) W/o NP 34.89 64.18 30.06 28.33 39.37
(b) W/o VP 18.75 63.93 30.89 42.11 38.92
(c) W/o PD 35.32 20.46 31.08 41.55 32.10
(d) SafePLUG 34.54 64.07 30.29 43.18 43.02

4.4.1 Effectiveness of dual-LoRA training

To examine the effectiveness of the proposed dual-LoRA design, we conduct an ablation analysis in which each LoRA branch is selectively enabled or disabled during training and inference. As summarized in Table 4, we compare four configurations: (a) a frozen backbone without any LoRA adaptation, (b) the text-oriented LoRA branch only, (c) the mask-oriented LoRA branch only, and (d) both LoRA branches jointly applied.

The results show the specialization of the two LoRA branches. Without any LoRA adaptation, the model fails to generalize across either textual or spatial modalities, indicating that the frozen backbone alone is insufficient to support downstream tasks. Enabling only the text-oriented LoRA substantially improves performance on language-based tasks such as region-level QA and accident description. This improvement demonstrates that the language-oriented adapters effectively enhance semantic understanding and cross-modal reasoning. However, since this configuration lacks the mask-oriented adaptation, its performance on pixel-level grounding still remains limited.

Conversely, activating only the mask-oriented LoRA yields the opposite trend: strong segmentation capability but severely degraded text reasoning. The model accurately localizes visual regions and produces high-quality binary masks, which confirms that the mask-oriented LoRA successfully refines dense spatial representations. Nonetheless, the absence of the text branch leads to weaker language grounding and insufficient semantic context, resulting in poor alignment between visual and textual modalities. This asymmetry highlights the necessity of having both LoRA branches to handle heterogeneous supervision signals.

When both branches are jointly loaded, SafePLUG achieves consistently strong results across all evaluation tasks. The combined configuration inherits the linguistic expressiveness of the text-oriented LoRA and the spatial precision of the mask-oriented LoRA, leading to the best overall performance. Importantly, the modular design allows each LoRA to specialize independently while remaining harmonized through the shared multimodal backbone. These findings validate that the dual-LoRA architecture enables SafePLUG to flexibly switch between textual reasoning and pixel-level understanding according to task requirements without retraining or compromising performance.

4.4.2 Effectiveness of model components

We further conduct ablation experiments to assess the contribution of each key component within SafePLUG, including the number prompt, visual prompt, and pixel decoder. These components correspond to temporal grounding, region QA, and pixel-level segmentation, respectively.

As shown in Table 5 configuration (a), removing the number prompt isolates its effect and leads to a substantial degradation in temporal localization performance. Without explicit numeric cues embedded in video frames, the model struggles to reliably associate semantic events with specific frame indices. This result helps explain why the number prompt provides a large improvement despite the inherent temporal order encoded in video backbones like TimeSformer [62] and Video-LLaMA [63]. While temporal order may be implicitly preserved, accident-related temporal grounding requires index-aware and referenceable anchors to precisely identify event onset and duration. By exposing frame indices directly to the model, the number prompt enables more stable and interpretable temporal grounding.

Comparing configurations (b) vs. (d), the visual prompt contributes +15.79 BLEU points to region QA while maintaining comparable performance on other tasks. This improvement reveals two key insights. First, the visual prompt encoder successfully extracts region-aware features through masked average pooling, which allows the model to selectively attend to semantically relevant areas rather than relying on global scene context. Second, the modest impact on pixel grounding (63.93 vs. 64.07) suggests that the visual prompt primarily enhances linguistic grounding of spatial regions rather than dense pixel-level segmentation itself. This distinction is important: the visual prompt guides the LLM's attention for question answering, while the pixel decoder handles the dense mask prediction. The complementary nature of these two components explains why both are necessary for comprehensive spatial understanding.

The dramatic performance drop in configuration (c) (–43.61 mIoU in pixel grounding) provides strong evidence that the SAM-based decoder is not merely a post-processing module but a critical component for translating language-conditioned queries into precise spatial outputs. Notably, removing the pixel decoder has minimal impact on region QA and accident description, indicating that the LLM backbone retains its semantic reasoning capability even without mask generation. However, the LLM's hidden representations alone lack the spatial resolution needed to delineate fine object boundaries. The pixel decoder addresses this gap by leveraging dense feature maps from the frozen SAM encoder and converting the LLM's <SEG> token embedding into pixel-wise predictions. This design demonstrates that effective pixel-level understanding requires specialized spatial decoders beyond standard language model architectures.

An important observation from Table 5 is the minimal performance degradation across unrelated tasks when removing specific modules. For instance, removing the number prompt (a) reduces temporal localization by 14.85 points but affects region QA by only 0.35 points. Similarly, removing the visual prompt (b) primarily impacts region QA while preserving temporal localization performance (42.11 vs. 43.18). This minimal interference validates our modular design principle: each component targets a specific capability without introducing negative transfer to other tasks.

Finally, the complete model (d) achieves the highest mean performance across all tasks, indicating that each component contributes complementary benefits to SafePLUG's overall capability. The number prompt enhances temporal precision, the visual prompt provides spatial selectivity, and the pixel decoder ensures pixel-level fidelity. Together, these modules enable SafePLUG to perform coherent multimodal reasoning that unifies spatial, temporal, and semantic understanding within a single framework.

4.5 Failure cases analysis

Inspired by failure analysis practices in Ego4D [46] and AVA [64], we provide a failure case analysis to help readers understand the model's limitations. Figure S13 presents representative failure cases across three challenging conditions.

The first row illustrates a case where a truck on the opposite lane collides with a green belt. Due to the occlusion caused by the green belt and roadside vegetation, SafePLUG incorrectly segments surrounding vehicles. This occurs because the pixel encoder struggles to extract reliable spatial features when object boundaries are obscured by foreground obstacles, leading to incorrect mask predictions.

The second row shows a failure case involving a collision with an obstacle (VO), which accounts for only 0.4% of the training data, as shown in Figure 4F. The model produces fragmented and over-extended segmentation masks that incorrectly include background regions such as trees and fences. This suggests that the limited training samples for rare accident categories prevent the model from learning robust visual patterns for these events.

The third row demonstrates a challenging nighttime scenario where visibility is severely limited. SafePLUG generates an over-segmented mask that covers nearly the entire image rather than precisely localizing the accident-involved vehicle. The lack of discriminative visual features under low-light conditions causes the model to output noisy and imprecise predictions.

These failure cases reveal that SafePLUG's performance degrades under conditions of visual ambiguity, data scarcity, and poor illumination. Addressing these limitations through techniques such as data augmentation for rare categories, low-light image enhancement, and occlusion-aware training represents promising directions for future work.

5 Conclusions

In this work, we propose SafePLUG, a novel framework that empowers multimodal large language models with both pixel-level understanding and temporal grounding capabilities for comprehensive traffic accident understanding. By integrating visual prompts for region-aware reasoning, number prompts for implicit temporal cues, and SAM for fine-grained segmentation, SafePLUG enables detailed spatial and temporal analysis across diverse accident scenarios. We further construct a large-scale benchmark dataset, SafePLUG-Bench, to support region QA, pixel-level grounding, accident description, and temporal localization. Extensive experiments demonstrate that SafePLUG outperforms strong baselines across all tasks while maintaining a lightweight architecture. Ablation studies confirm the effectiveness of our dual-LoRA training and modular design.

In future work, several promising directions could be explored to enhance scenario diversity and model capabilities. First, inspired by multi-view egocentric frameworks like 360 + X and recent autonomous driving datasets such as WOD-E2E [65], one potential direction is to extend SafePLUG to multi-camera setups with 360° surround-view coverage to capture blind spots and enable holistic scene understanding in challenging long-tail scenarios. Second, expanding scenario diversity through adverse weather conditions (fog, heavy rain, and snow), low-light/nighttime driving, and diverse road types (rural roads, complex intersections, and school zones) could improve model robustness. Third, integrating richer multimodal contexts beyond vision, such as audio signals and sensor data, represents another avenue for comprehensive accident understanding.

 Author Contributions

Zihao Sheng: Original draft; review & editing; visualization; validation; methodology; investigation; formal analysis; conceptualization. Zilin Huang: Review; validation; methodology; investigation. Yansong Qu: Review; validation; methodology; investigation. Jiancong Chen: Review; validation; methodology; investigation. Yuhao Luo: Review; validation; methodology; investigation. Yen-Jung Chen: Review; validation; methodology; investigation. Yue Leng: Review; validation; methodology; investigation; supervision. Sikai Chen: Original draft; review & drafting; validation; methodology; investigation; supervision; funding acquisition.

 Acknowledgments

Acknowledgements

This work was supported by the Center for Connected and Automated Transportation (CCAT), the USDOT Region 5 University Transportation Center funded by the U.S. Department of Transportation (Grant No. 69A3552348305). The contents of this paper reflect the views of the authors, who are responsible for the facts and the accuracy of the data presented herein, and do not necessarily reflect the official views or policies of the sponsoring organization.

 Conflict of Interests Statement

The authors declare that they have no conflict of interest.

 Data Availability Statement

The code, data, and models that support the findings of this study will be publicly available at https://zihaosheng.github.io/SafePLUG/.

 Supporting Information

Additional supporting information can be found online in the Supporting Information section.

References

1. 

S. Yin, C. Fu, S. Zhao, et al., "A Survey on Multimodal Large Language Models," National Science Review 11, no. 12 (2024): 1–20, https://doi.org/10.1093/nsr/nwae403.

2. 

D. Caffagni, F. Cocchi, L. Barsellotti, et al., "The Revolution of Multimodal Large Language Models: A Survey," paper presented at the Findings of the Association for Computational Linguistics, Bangkok, Thailand, August 11–16, 2024, https://doi.org/10.18653/v1/2024.findings-acl.807.

3. 

H. Liu, C. Li, Y. Li, and Y. J. Lee, "Improved Baselines with Visual Instruction Tuning," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02484.

4. 

Z. Li, Z. Cui, H. Liao, et al., "Steering the Future: Redefining Intelligent Transportation Systems with Foundation Models," CHAIN 1, no. 1 (2024): 46–53, https://doi.org/10.23919/CHAIN.2024.100003.

5. 

Z. Huang, Z. Sheng, Y. Qu, J. You, and S. Chen, "VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving," Transportation Research Part C: Emerging Technologies 180 (2025): 105321, https://doi.org/10.1016/j.trc.2025.105321.

6. 

Z. Sheng, Z. Huang, Y. Qu, Y. Leng, and S. Chen, "Talk2Traffic: Interactive and Editable Traffic Scenario Generation for Autonomous Driving with Multimodal Large Language Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPRW67362.2025.00364.

7. 

Z. Sheng, Z. Huang, Y. Qu, et al., "CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models," arXiv preprint, arXiv: 2502.15119, February 21, 2025, https://doi.org/10.48550/arXiv.2502.15119.

8. 

Z. Huang, Z. Sheng, and S. Chen, "PE-RLHF: Reinforcement Learning with Human Feedback and physics knowledge for safe and trustworthy autonomous driving," Transportation Research Part C: Emerging Technologies 179 (2025): 105262, https://doi.org/10.1016/j.trc.2025.105262.

9. 

C. Parikh, D. Rawat, R. T. Rakshitha, T. Ghosh, and R. K. Sarvadevabhatla, "RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01770.

10. 

Z. Xing, H. Chen, B. Xie, et al., "EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01779.

11. 

Y. Zhou, L. Bai, S. Cai, et al., "TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident," paper presented at the Thirteenth International Conference on Learning Representations, Singapore, April 24, 2025.

12. 

M. M. Karim, Y. Shi, S. Zhang, et al., "Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review," Artificial Intelligence for Transportation 1 (2025): 100004, https://doi.org/10.1016/j.ait.2025.100004.

13. 

R. Zhang, B. Wang, J. Zhang, et al., "When Language and Vision Meet Road Safety: Leveraging Multimodal Large Language Models for Video-Based Traffic Accident Analysis," Accident Analysis & Prevention 129 (2025): 108077, https://doi.org/10.1016/j.aap.2025.108077.

14. 

Y. Yan, Y. Liao, G. Xu, et al., "Large Language Models for Traffic and Transportation Research: Methodologies, State of the Art, and Future Opportunities," arXiv preprint, arXiv: 2503.21330, March 27, 2025, https://doi.org/10.48550/arXiv.2503.21330.

15. 

S. Jiang, Z. Huang, K. Qian, et al., "A Survey on Vision-Language-Action Models for Autonomous Driving," arXiv preprint, arXiv: 2506.24044, June 30, 2025, https://doi.org/10.48550/arXiv.2506.24044.

16. 

L. Xu, H. Huang, and J. Liu, "SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning over Traffic Events," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, https://doi.org/10.1109/CVPR46437.2021.00975.

17. 

J. Fang, L. Li, J. Zhou, et al., "Abductive Ego-View Accident Video Understanding for Safe Driving Perception," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02080.

18. 

Z. Huang, Z. Sheng, Z. Wan, et al., "Sky-Drive: A Distributed Multi-Agent Simulation Platform for Socially-Aware and Human-AI Collaborative Future Transportation," Journal of Intelligent and Connected Vehicles (2026): 1–16, https://doi.org/10.26599/JICV.2026.9210070.

19. 

Z. Huang, Z. Sheng, C. Ma, and S. Chen, "Human as AI Mentor: Enhanced Human-in-The-Loop Reinforcement Learning for Safe and Efficient Autonomous Driving," Communications in Transportation Research 4 (2024): 100127, https://doi.org/10.1016/j.commtr.2024.100127.

20. 

K. Long, Z. Sheng, H. Shi, et al., "Physical Enhanced Residual Learning (PERL) Framework for Vehicle Trajectory Prediction," Communications in Transportation Research 5 (2025): 100166, https://doi.org/10.1016/j.commtr.2025.100166.

21. 

Z. Sheng, Z. Huang, and S. Chen, "Traffic Expertise Meets Residual RL: Knowledge-Informed Model-Based Residual Reinforcement Learning for CAV Trajectory Control," Communications in Transportation Research 4 (2024): 100142, https://doi.org/10.1016/j.commtr.2024.100142.

22. 

X. Lai, Z. Tian, Y. Chen, et al., "LISA: Reasoning Segmentation via Large Language Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.00915.

23. 

H. Yuan, X. Li, T. Zhang, et al., "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos," arXiv preprint, arXiv: 2501.04001, November 3, 2025, https://doi.org/10.48550/arXiv.2501.04001.

24. 

Z. Ren, Z. Huang, Y. Wei, et al., "PixelLM: Pixel Reasoning with Large Multimodal Model," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02491.

25. 

W. Lin, X. Wei, R. An, et al., "Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want," paper presented at the International Conference on Learning Representations, Singapore, April 24, 2025.

26. 

M. Cai, H. Liu, D. Park, et al., "ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01227.

27. 

T. Zhang, X. Li, H. Fei, et al., "OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding," paper presented at the Conference on Neural Information Processing Systems, Vancouver, Canada, December 10–15, 2024, https://doi.org/10.52202/079017-2291.

28. 

Y. Yao, X. Wang, M. Xu, et al., "DoTA: Unsupervised Detection of Traffic Anomaly in Driving Videos," IEEE Transactions on Pattern Analysis and Machine Intelligence 45, no. 1 (2023): 444–459, https://doi.org/10.1109/TPAMI.2022.3150763.

29. 

B. Lin, Y. Ye, B. Zhu, et al., "Video-LLaVA: Learning United Visual Representation by Alignment Before Projection," in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, edited by Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, 5971–5984, Association for Computational Linguistics, 2024, https://doi.org/10.18653/v1/2024.emnlp-main.342.

30. 

B. Huang, X. Wang, H. Chen, Z. Song, and W. Zhu, "VTimeLLM: Empower LLM to Grasp Video Moments," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01353.

31. 

S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, "TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01357.

32. 

Y. Wu, X. Hu, Y. Sun, et al., "Number it: Temporal Grounding Videos like Flipping Manga," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025, https://doi.org/10.1109/CVPR52734.2025.01284.

33. 

A. Kirillov, E. Mintun, N. Ravi, et al., "Segment Anything," in Proceedings of the IEEE/CVF Conference on International Conference on Computer Vision, 2023, https://doi.org/10.1109/ICCV51070.2023.00371.

34. 

Y. Ling and Z. Ma, "Pedestrian Crossing Intention Prediction in the Wild: A Survey," CHAIN 1, no. 4 (2024): 263–279, https://doi.org/10.23919/CHAIN.2024.000008.

35. 

T. You and B. Han, "Traffic Accident Benchmark for Causality Recognition," paper presented at the European Conference on Computer Vision, Glasgow, UK, August 23–28, 2020, https://doi.org/10.1007/978-3-030-58571-6_32.

36. 

Y. Qu, Z. Xu, Z. Huang, et al., "MetaSSC: Enhancing 3D semantic scene completion for autonomous driving through meta-learning and long-sequence modeling," Communications in Transportation Research 5 (2025): 100184, https://doi.org/10.1016/j.commtr.2025.100184.

37. 

J. Fang, J. Qiao, J. Xue, and Z. Li, "Vision-Based Traffic Accident Detection and Anticipation: A Survey," IEEE Transactions on Circuits and Systems for Video Technology 34, no. 4 (2024): 1983–1999, https://doi.org/10.1109/TCSVT.2023.3307655.

38. 

K. Yin, Y. Li, X. Li, and H. Zhao, "Human Motion Intention Recognition via sEMG and Joint Kinematics Fusion Using MPSO-SVM for Intelligent Transportation Systems," CHAIN 2, no. 2 (2025): 198–209, https://doi.org/10.23919/CHAIN.2025.000012.

39. 

Y. Guan, H. Liao, C. Wang, et al., "Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation," Communications in Transportation Research 5 (2025): 100214, https://doi.org/10.1016/j.commtr.2025.100214.

40. 

F. H. Chan, Y. T. Chen, Y. Xiang, and M. Sun, "Anticipating Accidents in Dashcam Videos," paper presented at the Asian Conference on Computer Vision, Taiwan, China, November 20–24, 2016, https://doi.org/10.1007/978-3-319-54190-7_9.

41. 

H. Lv, C. Zhou, Z. Cui, et al., "Localizing Anomalies From Weakly-Labeled Videos," IEEE Transactions on Image Processing 30 (2021): 4505–4515, https://doi.org/10.1109/TIP.2021.3072863.

42. 

Y. Yao, M. Xu, Y. Wang, D. J. Crandall, and E. M. Atkins, "Unsupervised Traffic Accident Detection in First-Person Videos," paper presented at the IEEE/RSJ International Conference on Intelligent Robots and Systems, Macau, China, November 3–8, 2019, https://doi.org/10.1109/IROS40897.2019.8967556.

43. 

W. Bao, Q. Yu, and Y. Kong, "Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational Learning," in Proceedings of the 28th ACM International Conference on Multimedia, 2682–2690, ACM Digital Library, 2020, https://doi.org/10.1145/3394171.3413827.

44. 

J. Fang, D. Yan, J. Qiao, J. Xue, and H. Yu, "DADA: Driver Attention Prediction in Driving Accident Scenarios," IEEE Transactions on Intelligent Transportation Systems 23, no. 6 (2022): 4959–4971, https://doi.org/10.1109/TITS.2020.3044678.

45. 

J. Zhang, K. Yang, and R. Stiefelhagen, "Exploring Event-Driven Dynamic Context for Accident Scene Segmentation," IEEE Transactions on Intelligent Transportation Systems 23, no. 3 (2022): 2606–2622, https://doi.org/10.1109/TITS.2021.3134828.

46. 

K. Grauman, A. Westbury, E. Byrne, et al., "Ego4D: Around the World in 3,000 Hours of Egocentric Video," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, https://doi.org/10.1109/CVPR52688.2022.01842.

47. 

H. Chen, Y. Hou, C. Qu, et al., "360+x: A Panoptic Multi-modal Scene Understanding Dataset," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.01833.

48. 

B. Zhu, B. Lin, M. Ning, et al., "LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment," arXiv preprint, arXiv: 2310.01852, January 22, 2024, https://doi.org/10.48550/arXiv.2310.01852.

49. 

X. Huang, L. Shen, J. Liu, et al., "Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine," in Proceedings of the AAAI Conference on Artificial Intelligence, edited by Toby Walsh, Julie Shah, Zico Kolter, 3782–3790, AAAI Press, 2025, https://doi.org/10.1609/aaai.v39i4.32394.

50. 

R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, "Adaptive Mixtures of Local Experts," Neural Computation 3, no. 1 (1991): 79–87, https://doi.org/10.1162/neco.1991.3.1.79.

51. 

W. Fedus, B. Zoph, and N. Shazeer, "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," Journal of Machine Learning Research 23, no. 120 (2022): 1–39.

52. 

Z. Sheng, Z. Huang, and S. Chen, "Kinematics-Aware Multigraph Attention Network with Residual Learning for Heterogeneous Trajectory Prediction," Journal of Intelligent and Connected Vehicles 7, no. 2 (2024): 138–150, https://doi.org/10.26599/JICV.2023.9210036.

53. 

E. J. Hu, Y. Shen, P. Wallis, et al., "LoRA: Low-Rank Adaptation of Large Language Models," paper presented at the International Conference on Learning Representations, April 25–29, 2022.

54. 

Z. Chen, J. Wu, W. Wang, et al., "InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, https://doi.org/10.1109/CVPR52733.2024.02283.

55. 

A. Yang, B. Yang, B. Zhang, et al., "Qwen2.5 Technical Report," arXiv preprint, arXiv: 2412.15115, January 3, 2025, https://doi.org/10.48550/arXiv.2412.15115.

56. 

The Vicuna Team, "Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality," Accessed January 14, 2026. https://lmsys.org/blog/2023-03-30-vicuna.

57. 

K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, "BLEU: A Method for Automatic Evaluation of Machine Translation," in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, edited by Pierre Isabelle, Eugene Charniak, and Dekang Lin, 311–318, ACL, 2002, https://doi.org/10.3115/1073083.1073135.

58. 

C. Y. Lin, "ROUGE: A Package for Automatic Evaluation of Summaries," paper presented at the Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, July 25–26, 2004.

59. 

T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, "BERTScore: Evaluating Text Generation with BERT," arXiv preprint, arXiv: 1904.09675, Februray 24, 2020, https://doi.org/10.48550/arXiv.1904.09675.

60. 

H. Liu, C. Li, Q. Wu, and Y. J. Lee, "Visual Instruction Tuning," paper presented at the Thirty-seventh Conference on Neural Information Processing Systems, New Orleans, USA, December 10–16, 2023.

61. 

Z. Li, Q. Xu, D. Zhang, et al., "GroundingGPT: Language Enhanced Multi-modal Grounding Model," arXiv preprint, arXiv: 2401.06071, March 5, 2024, https://doi.org/10.48550/arXiv.2401.06071.

62. 

G. Bertasius, H. Wang, and L. Torresani, "Is Space-Time Attention All You Need for Video Understanding?" in Proceedings of the 38th International Conference on Machine Learning, edited by Marina Meila, Tong Zhang, 813–824, PMLR, 2021.

63. 

H. Zhang, X. Li, and L. Bing, "Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding," arXiv preprint, arXiv: 2306.02858, October 25, 2023, https://doi.org/10.48550/arXiv.2306.02858.

64. 

C. Gu, C. Sun, D. A. Ross, et al., "AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions," arXiv preprint, arXiv: 1705.08421, April 30, 2018, https://doi.org/10.48550/arXiv.1705.08421.

65. 

R. Xu, H. Lin, W. Jeon, et al., "WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios," arXiv preprint, arXiv: 2510.26125, November 13, 2025, https://doi.org/10.48550/arXiv.2510.26125.

Top