ContentsFigures & Tables
1 Introduction

1 Introduction

2 Related work

2 Related work

3 Method

3 Method

3.1 Image enhancement module

3.1 Image enhancement module

3.2 Vehicle detection module

3.2 Vehicle detection module

3.2.1 Increasing secondary innovation with CDFA attention

3.2.1 Increasing secondary innovation with CDFA attention

3.2.2 Replacing residual blocks and feature pyramids

3.2.2 Replacing residual blocks and feature pyramids

3.2.3 Knowledge distillation

3.2.3 Knowledge distillation

3.3 Target tracking module

3.3 Target tracking module

3.3.1 Lightweight network improvements

3.3.1 Lightweight network improvements

3.3.2 Improvements to the loss function

3.3.2 Improvements to the loss function

3.4 Monocular camera ranging method

3.4 Monocular camera ranging method

3.5 Vehicle safety distance warning system

3.5 Vehicle safety distance warning system

4 Experimental verification and analysis

4 Experimental verification and analysis

4.1 Experimental environment and dataset

4.1 Experimental environment and dataset

4.2 Evaluation indicators

4.2 Evaluation indicators

4.3 Ablation experiment and comparative experiment

4.3 Ablation experiment and comparative experiment

4.4 Real-vehicle verification

4.4 Real-vehicle verification

5 Conclusion

5 Conclusion

References

References

Overcoming low light and missed detection: A real-time vehicle cooperative perception and early warning method based on deep learning

Hengkai Wang1Zhe Zhang2,4Xiao Luo1Wei Zhong2Hong Qi3Dong Zhang5
1. State Key Laboratory of Advanced Vehicle Integration and Control, China FAW Group Co., LTD., Changchun 130013, China
2. State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China
3. College of Computer Science and Technology, Jilin University, Changchun 130012, China
4. Liaoning University of Technology, Jinzhou 121000, China
5. Department of Mechanical and Aerospace Engineering, Brunel University, London UB8 3PH, UK
Abstract: In traffic scenarios, the dynamic characteristics and random behaviors of vehicles are the main reasons for the frequent occurrence of collision accidents. Traditional early warning systems restrict traffic safety due to low detection accuracy and poor tracking effect. Research has proposed a vehicle safety distance early warning system based on deep learning to enhance traffic safety. Innovations include: adopting the self-calibrated illumination (SCI) algorithm to overcome light interference; the YOLOv11 algorithm is improved by introducing a secondary innovative cross-domain feature attention (CDFA) mechanism, reconstructing the feature pyramid, and integrating knowledge distillation to balance detection accuracy and real-time performance. The DeepSORT algorithm is improved by applying group convolution to reduce the number of parameters and replacing Intersection over Union (IoU) with MPDIoU to enhance tracking accuracy. The distance between vehicles is calculated by using the monocular vision ranging method. The detection, tracking, and ranging modules are integrated into a vehicle safety distance early warning system. Experimental evaluation demonstrates a marked improvement in the system's performance. On the public dataset, the detection model exhibits a gain of 3.89% in mAP@0.5 and 2.76% in mAP@0.5:0.95, while the tracking model achieves a 0.9% increase in multiple object tracking accuracy (MOTA). Furthermore, real-world vehicle validation confirms that the synergistic operation of the detection and tracking modules effectively mitigates the miss rate, thereby substantiating a tangible enhancement in overall system safety.
Keywords: real-time vehicle detection; multi-target tracking; deep learning; traffic safety warning; attention mechanism
Received: 2025-10-20

1 Introduction

With the coordinated development of vehicle–cloud integration and intelligent connected vehicle technology, autonomous driving has become a cornerstone of future intelligent transportation systems [1]. Reliable trajectory prediction and proactive safety warnings require efficient, high-precision, and lightweight solutions for vehicle detection, multi-target tracking, and distance estimation. This requirement becomes particularly critical when vehicle-to-everything (V2X) communications are degraded or unavailable. In such scenarios, low-latency on-vehicle perception must independently ensure the safety of closed-loop control. However, visual perception degrades severely under low-light conditions, which has motivated extensive research into enhancement techniques. Recent approaches have explored deep learning-based Retinex decomposition [2], zero-reference curve estimation methods [3], and multispectral fusion strategies for nighttime scenarios. While these techniques demonstrate promising enhancement capabilities, they introduce substantial computational burdens that compromise real-time processing requirements. Moreover, current enhancement frameworks operate independently from detection systems, lacking unified optimization across the perception pipeline. The computational overhead and architectural decoupling present significant barriers to practical deployment, highlighting the critical need for lightweight enhancement solutions that can be seamlessly integrated into end-to-end detection frameworks without sacrificing processing efficiency.

Resource constraints in autonomous driving systems necessitate efficient architectures for edge deployment. GhostNet, MobileNetV3, and ShuffleNetV2 [4] employ efficient convolution operations to reduce computational costs. Knowledge distillation techniques transfer capabilities from large models to compact ones. Network attached storage (NAS) based methods like EfficientDet [5] and DetNAS automatically search for optimal architectures. Quantization techniques compress models by reducing numerical precision. Despite progress in individual components, current research exhibits several critical gaps. First, existing frameworks lack integration that jointly optimizes detection, tracking, and ranging in a unified pipeline. Second, efficient low-light enhancement solutions that avoid computational overhead remain unavailable. Third, achieving an optimal balance between real-time performance and model lightweightness for edge devices remains unresolved. Fourth, current methods inadequately address robustness when handling occlusions, crowded scenes, and appearance changes simultaneously. These methods typically focus on single aspects rather than holistic optimization. Consequently, they inadequately balance accuracy and compactness for practical autonomous driving scenarios. These limitations motivate this systematic investigation of real-time vehicle detection, multi-target tracking, and ranging in complex road environments. This research aims to develop a comprehensive end-to-end solution that addresses these multifaceted challenges effectively. The principal contributions are as follows:

(1) A vehicle safety distance early warning system has been developed to achieve coordinated dispatching of vehicle detection, tracking, and distance measurement, and to view the real-time positions of surrounding vehicles and vehicle collision early warning information.

(2) In terms of detection, firstly, the low-light enhancement algorithm self-calibrated illumination (SCI) was introduced. Secondly, a novel cross-domain feature attention (CDFA) module was proposed to reduce the problem of vehicle detection accuracy decline caused by environmental occlusion. Then, in order to fully leverage the capabilities of CDFA, a novel MANet-FasterCGLU module was proposed. Finally, to further improve the accuracy without increasing the number of parameters, the knowledge distillation method was adopted.

(3) In terms of tracking, grouped convolution is adopted to replace the convolution in the DeepSORT network to reduce model parameters, and the Intersection over Union (IoU) for tracking matching is replaced with MPDIoU.

2 Related work

Modern object detection has evolved through three major paradigms: anchor-based, anchor-free, and transformer-based architectures. In anchor-based methods, Tian et al. [6] enhanced YOLOv8 with PrFuFPN and GhostConv for Unmanned Aerial Vehicle (UAV) detection. Xie et al. [7] proposed the ADD-CGLU structure in YOLOv10. Anchor-free detectors such as Fully Convolutional One-Stage Object Detection (FCOS) [8], CornerNet [9], and YOLOX [10] eliminate anchor design complexities. Transformer-based methods including Detection Transformer (DETR) [11], Deformable DETR, and Swin Transformer achieve higher detection accuracy. However, these advanced architectures require substantial computational resources, which limits their deployment on edge devices. For multi-object tracking, DeepSORT [12] dominates the tracking-by-detection paradigm by integrating appearance features with Kalman filtering. Recent advances have also explored transformer architectures such as TrackFormer and TransTrack. ByteTrack [13] improves association by utilizing low-score detections. Joint frameworks like JDEdwards EnterpriseOne (JDE) and FairMOT unify detection and re-identification in end-to-end pipelines. Nevertheless, maintaining stable IDs under severe occlusions and in crowded scenes remains a significant challenge. Collaborative perception further enhances robustness through multimodal fusion. Methods such as PointPainting, MVX-Net, and deep continuous fusion integrate LiDAR and camera data. V2X approaches including F-Cooper, V2X-ViT, and V2VNet enable cooperative perception across vehicles. These methods require complex calibration procedures and introduce communication latency. Moreover, they fail when communication infrastructure degrades or becomes unavailable.

To further address the contradiction between the number of parameters and the real-time performance of vehicle detection models, as well as the problems of feature competition and missed detection in multi-object detection and the limited generalization ability in dynamic environments, a vehicle safety distance early warning system is developed to achieve the collaborative scheduling of vehicle detection, tracking, and distance measurement. The multi-algorithm collaborative framework is shown in Fig. 1. The specific process is as follows. First, the camera is loaded to collect real-time video. Second, after enhancing image quality through the SCI algorithm, target recognition and feature extraction are performed. Then consecutive frame targets are associated based on the current frame results and the historical tracking status. Finally, the distance module tracks the relative distance to the target. Users can view the real-time locations and warning information of surrounding vehicles, and a red alert is triggered when the distance between vehicles falls below a predefined threshold.

Figure 1 Multi-algorithm collaborative framework.

3 Method

A vehicle safety distance early warning system has been researched and developed to achieve vehicle detection, tracking, and distance measurement. The main innovation first applies the SCI image enhancement algorithm and the similar triangle distance measurement method. Second, improvements are made to the object detection algorithm, including proposing a new CDFA attention mechanism, a new MANet-FasterCGLU module, and adopting a knowledge distillation method. Finally, innovations in tracking include achieving lightweight design by using group convolution and replacing IoU in the matching mechanism with MPDIoU.

3.1 Image enhancement module

Significant changes in illumination in complex traffic scenarios greatly affect vehicle detection performance. Therefore, it is crucial to select an optimal low-light enhancement method. This study selects representative algorithms from 2020 to 2023 (EnlightenGAN [14], RUAS [15], SCI [16], URetinex-Net [17], and Zero-DCE [18]) and conducts joint experiments of image enhancement and object detection based on the BDD100K low-light dataset. The results are shown in Fig. 2. The comparison results show that although EnlightenGAN improves clarity, it reduces the number of effectively detected vehicles to two; RUAS has overly strong color difference restoration, making it sensitive to light sources and prone to missed detections; URetinex-Net blurs light sources, causing overall blurriness and false detections; Zero-DCE has insufficient clarity, leading to missed detections; whereas the SCI algorithm improves clarity while maintaining moderate sensitivity to light sources, successfully detecting five vehicles with excellent recognition performance. Based on these findings, SCI is selected as the low-light enhancement module.

Figure 2 Image enhancement comparison chart.

3.2 Vehicle detection module

As a representative single-stage object detection method, the YOLO algorithm has become a mainstream framework in this field due to its fast detection speed and simple architecture. Given the requirements of vehicle detection tasks for model lightweight design and real-time performance while maintaining detection accuracy, this study adopts YOLOv11 [19] as the basic detection network. To address common issues in vehicle detection models, such as a large number of parameters, limited real-time performance, and suboptimal accuracy, this study employs its lightweight variant YOLOv11n as the baseline model and introduces targeted improvements to mitigate accuracy degradation.

3.2.1 Increasing secondary innovation with CDFA attention

To alleviate the decline in vehicle detection accuracy caused by environmental occlusion, this study improves the CDFA mechanism. This mechanism was initially proposed by Lei et al. [20] for medical image segmentation. It monitors the representation of target and background features through an auxiliary detection head, generates a Softmax weight map via convolution, and fuses it with the original features to enhance target discriminability.

To better meet the requirements of vehicle detection, this study introduces a feature enhancement strategy based on wavelet transform by embedding the WaveletConv module [21] into the auxiliary detection head. Relying on its multi-resolution analysis capability, the module simultaneously extracts low-frequency components (background-dominant information) and high-frequency components (vehicle detail information) of the image. This frequency-domain feature decomposition enables more accurate capture of discriminative information between vehicles and backgrounds, thereby significantly improving the detection performance of the model under occlusion scenarios. The improved structure is illustrated in Fig. 3.

Figure 3 CDFA attention structure diagram.

3.2.2 Replacing residual blocks and feature pyramids

The design of the MANet-FasterCGLU module draws inspiration from the mixed aggregation network (MANet) structure proposed by Tsinghua University in Hyper-YOLO [22]. This integrated design of heterogeneous convolutional operations significantly enhances the feature extraction efficiency of the backbone network. The improved network structure, HH-YOLOv11, is illustrated in Fig. 4, where the purple regions denote the CDFA module, the red regions denote the MANet-FasterCGLU module, and the yellow regions denote the WFU module. The original MANet structure is shown in Fig. 5.

Figure 4 HH-YOLOv11 network structure diagram.
Figure 5 MANet structure diagram.

To fully exploit the advantages of the CDFA mechanism in target–background separation and the WaveletConv module in extracting discriminative features via wavelet transformation, this study introduces innovative improvements to the ConvNeck module in MANet. The theoretical foundation of this improvement lies in the synergistic enhancement between multi-scale frequency decomposition and attention selection. Specifically, the frequency-domain decomposition of WaveletConv explicitly separates discriminative high-frequency edge features from redundant low-frequency background information through multi-scale frequency components (low-low(LL), low-high(LH), high-low(HL), and high-high(HH) sub-bands), enabling CDFA to weight different frequency bands independently rather than process mixed features. This frequency–semantic correspondence mechanism provides three key advantages. (1) Information entropy reduction: decomposed frequency components carry purer semantic information with lower entropy, which produces sharper attention gradients with higher variance concentrated on target regions. (2) Multi-scale complementarity: low-frequency LL components preserve global vehicle shapes under severe occlusion, while high-frequency HH components capture fine-grained textures such as license plates, thus providing redundant discriminative cues. (3) Adaptive weighting: CDFA learns to dynamically emphasize shape-based attention (higher LL weights) in occluded scenarios and texture-based attention (higher HH weights) in clear conditions. To maximize this frequency–attention synergy, the original ConvNeck module requires two key improvements: Suppressing information redundancy in residual connections that dilutes enhanced frequency components, and introducing adaptive gating mechanisms to prioritize discriminative multi-scale frequency features.

The improved MANet-FasterCGLU module structure is shown in Fig. 6. The specific improvement measures are as follows:

Figure 6 MANet-FasterCGLU structure diagram.

(1) Architectural redesign

To achieve optimal synergy between frequency decomposition and attention mechanisms, the ConvNeck module is redesigned using the FasterNetBlock architecture [23]. This design choice leverages FasterNetBlock's Partial Convolution operations to reduce computational redundancy while processing multi-scale frequency components, and its depth-wise separable structure naturally aligns with independent frequency band processing requirements, enabling more efficient feature extraction and information aggregation.

(2) Redundancy suppression and information enhancement

To address information redundancy in residual connections that dilutes enhanced frequency components, the convolutional gate linear unit (CGLU) attention aggregation mechanism [24] is introduced. This gating mechanism selectively suppresses redundant features while amplifying discriminative information through parallel pathways, effectively filtering background noise and preserving critical edge features, thereby improving information integration efficiency and ensuring the network focuses on discriminative information.

3.2.3 Knowledge distillation

To enhance the vehicle detection network's accuracy, this study introduces a knowledge distillation method while maintaining the number of model parameters, computational complexity and reasoning speed unchanged. As an efficient means of model compression and transfer, this technology can effectively transfer the knowledge of a complex teacher model to streamlined student models, achieving a performance leap without increasing the burden of computing power. The research adopted YOLOv11x equipped with the self-developed LSDECD module to construct the teacher model and the lightweight HH-YOLOv11 was used as the student model.

This study innovatively designed the LSDECD detection head based on the YOLOv11x baseline model (architecture shown in Fig. 7), which includes three core improvements:

Figure 7 LSDECD structure diagram.

(1) Group normalization (GN) [25] is used to replace traditional batch normalization to solve the batch size bottleneck problem caused by memory limitations in large model training. Existing studies such as Fully Convolutional One-Stage Object Detection (FCOS) [26] have confirmed that it can improve the accuracy of target localization and classification.

(2) By integrating the original independent feature extraction operations through the parameter sharing mechanism and combining with the learnable scale module to achieve adaptive feature processing, the computational load and the number of parameters are significantly reduced, enhancing the deployment capability of edge devices.

(3) Detail-enhanced convolution (DEConv) [27] is adopted to strengthen the extraction of discriminative detail features. The collaborative optimization of the three significantly enhances the detection performance while maintaining lightweight.

3.3 Target tracking module

Target tracking is based on video frame analysis. It uses detection algorithms to obtain the category and position of the target, and matches and associates adjacent frames with the same target according to motion and appearance features to obtain the trajectory. Since the industrial camera captures real-time video when the vehicle is in motion, a real-time multitarget tracking algorithm must be selected. Therefore, in this section, the online algorithm DeepSORT is adopted. The improved HH-YOLOv11 from the previous section is used as its detector, and the algorithm is optimized to enhance the model's real-time performance and tracking efficiency.

3.3.1 Lightweight network improvements

Since the DeepSORT algorithm [28] must work in coordination with the image enhancement module and the HH-YOLOv11 detector, the overall model inevitably has a large number of parameters and high computational complexity. Although HH-YOLOv11 has reduced its own complexity by choosing the YOLOv11n model, further lightweight improvements to the DeepSORT algorithm are still needed to enhance the detection efficiency and response speed of the object detection and tracking system. For this purpose, this paper improves the feature extraction network of DeepSORT, replacing the traditional convolution with grouped convolution to achieve lightweighting. The structure of the DeepSORT deep feature extraction network with grouped convolution is shown in Table 1.

Table 1 Improved DeepSORT deep feature extraction network structure.
Network structure name Patch size/step size Groups Output size(channels×height×width)
Conv1 3×3/1 1 64×128×64
MaxPool 3×3/2 - 64×64×32
ResBlock1 3×3/1 4 64×64×32
ResBlock2 3×3/1 4 64×64×32
ResBlock3 3×3/2 8 128×32×16
ResBlock4 3×3/1 8 128×32×16
ResBlock5 3×3/2 16 256×16×8
ResBlock6 3×3/1 16 256×16×8
ResBlock7 3×3/2 32 512×8×4
ResBlock8 3×3/1 32 512×8×4
AvgPool 8×4/1 - 512×1×1
FC1(Linear) - - 256

3.3.2 Improvements to the loss function

In the vehicle detection scenario, although the DeepSORT algorithm adopts a dual strategy of cascade matching and IoU matching for target association, it faces challenges: the cascade matching stage fails due to sudden changes in vehicle appearance features caused by traffic flow occlusion and similarity below the threshold. The IoU mechanism that is instead relied on has inherent flaws. When the detection boxes do not overlap, the IoU is zero, which cannot reflect the spatial positional relationship and is prone to mismatches and missed detections in complex traffic scenarios. To this end, this paper introduces MPDIoU [29] to replace the traditional IoU matching mechanism. MPDIoU can effectively overcome the inherent limitations of traditional IoU to provide more accurate matching metrics and demonstrate superior generalization performance. Therefore, this paper introduces MPDIoU into the DeepSORT algorithm to effectively mitigate the problem of target detection failures caused by mutual occlusion between targets during multitarget tracking, thereby improving the stability and accuracy of tracking. The improved DeepSORT algorithm association process is shown in Fig. 8.

Figure 8 Modified DeepSORT association flowchart.

3.4 Monocular camera ranging method

This section adopts a similar triangle ranging algorithm to efficiently obtain vehicle distance information: based on the camera focal length, the size of the target pixel, and its actual physical size, the distance between the target and the camera is calculated using the principle of geometric similarity. As shown in Fig. 9, given the pixel width of the image of the vehicle in front (or the reference object), the focal length of the camera and its actual width, the target distance can be derived through a similar triangular relationship.

Figure 9 Principle of similar triangle distance measurement algorithm.

In Fig. 9, H represents the actual height of the vehicle, h represents its pixel height on the image plane, and f represents the camera focal length. The specific values of these parameters can be determined through subsequent camera internal parameter calibration. The longitudinal distance between the target vehicle and the camera can be calculated based on the following Eq. (1). f d = h H → d = H f h (1)

3.5 Vehicle safety distance warning system

This section presents the vehicle safety distance warning system (VSDWS), as shown in Fig. 10, which is an intelligent platform based on deep learning, integrating technologies such as low-light enhancement, high precision detection, multi-target tracking and monocular distance measurement. The core is to identify vehicles in images/videos in real time and calculate the relative distance from one's vehicle. When the distance is lower than the safety threshold, a red box is used to warn of the risk of collision. The interface is divided into three zones: the video on the left shows the warning effect; the lower left interaction area includes loading/start-stop detection and image enhancement switches, supporting the comparison of the effects before and after enhancement; the parameter area on the right displays the detection statistics (total number of vehicles/number of warnings) and adjustable safety distance parameters, and it monitors the operation status and logs.

Figure 10 Vehicle distance warning interface.

4 Experimental verification and analysis

This section uses public datasets for testing and real vehicle verification to examine the feasibility and practicality of the improved methods. First, explain the selection of the dataset and the evaluation parameters. Then, the comparison and ablation experiment results of the improved object detection and tracking algorithm with the baseline model on the public set will be presented. Finally, conduct real vehicle verification.

4.1 Experimental environment and dataset

The object detection task utilizes 13,274 images selected from the KITTI [30] and BDD100K [31] datasets, with low-light and normal scenes split at a 5∶5 ratio. These images were partitioned into training, validation, and test sets at a 7∶2∶1 ratio. The tracking task employed the UA-DETRAC [32] video dataset, from which 10,000 video frames were selected and split using the same 7∶2∶1 ratio. All experiments were conducted on a Linux system equipped with a GeForce RTX 4060 GPU and 16GB RAM, implemented using PyTorch 1.18.0 with CUDA 11.7 and Python 3.10. The training process consisted of 300 epochs with a batch size of 8 and a base learning rate of 0.001. The default input resolution was set to 640 × 640 pixels. To enhance model generalization capability, multiple data augmentation techniques were applied during training, including random rotation, random scaling, and random flipping.

4.2 Evaluation indicators

In the evaluation methods of the target detection algorithm, four types of evaluation indicators, including F1-score, frames per second (FPS), giga floating point operations per second (GFLOPs), and mean average precision (mAP), were selected to measure its quality. In the evaluation methods of the target tracking algorithm, multiple object tracking accuracy(MOTA), FPS, and GFLOPs were selected to measure its quality.

(1) Target detection evaluation indicators

Object detection performance is often comprehensively evaluated by precision, recall rate and mAP (average precision of intersection and union of detection frames) [33]. Because there is a trade-off relationship between the two, the F1-score is introduced for harmonic averaging and its formula is shown in Eq. (2). F 1 = 2 × P × R P + R (2)

Among these, P represents accuracy and R represents recall rate. The higher the F1-score value, the better the model's overall performance in terms of accuracy and recall rate.

As mentioned earlier, the IoU of the detection box represents the coincidence ratio between the detection box and the true value box. Therefore, regardless of the target object being detected, the detection effect can be evaluated by a set IoU threshold confidence level, that is, the average detection accuracy. Its calculation formula is shown in Eqs.(3) and (4): A P = ∫ 0 1 P R d r (3) m A P = 1 c ∑ i =1 c A P i (4)

Among these, AP represents detection accuracy, P represents precision, R represents recall rate, and r represents the range of recall rate values. It integrates all possible recall rate values from 0 to 1 to consider the model's precision performance at different recall rates. mAP represents average detection accuracy and c represents the number of categories. The higher the mAP value, the better the model's overall detection accuracy.

(2) Target tracking evaluation indicators

In the study of multi-target tracking algorithms, the MOTA and the number of model parameters are usually used to measure the tracking situation. The calculation formula of MOTA is shown in Eq. (5). The higher its accuracy, the better the tracking effect. m = 1 − ∑ t ( F N t + F P t + I D S W t ) ∑ t G T t (5)

Among them, m is the multi-target tracking accuracy, t is the frame number, FNt is the number of missed detections in frame t, FPt is the number of false detections in framet, IDSWt is the number of incorrect matches in frame t and GTt is the number of real objects in frame t.

4.3 Ablation experiment and comparative experiment

To verify the effectiveness of the improved target detection scheme, ablation experiments were conducted on the KITTI and BDD100K standard datasets (the results are shown in Table 2 and the training curves are shown in Fig. 11). The data shows that the introduction of each improved module has significantly increased the mAP and F1 scores, verifying the effectiveness and synergy of the modules. From the training curve, it can be seen that the improved model is always higher than the original one, which indirectly proves the feasibility of the improved algorithm.

Table 2 Ablation experiments of the object detection module.
Model F1-score mAP50/% mAP50-95/% FPS GFLOPs
YOLOv11n 52.28 51.01 34.06 160.5 6.3
YOLOv11n-CDFA 54.03 53.20 35.57 143.6 8.6
YOLOv11n-MANet-FasterCGLU 64.65 52.65 34.6 138.4 8.9
YOLOv11n-all 54.34 52.99 35.01 121.7 11.2
YOLOv11x 61.78 60.41 41.13 31.2 231.2
YOLOv11x-LSDECD 63.38 61.57 41.73 33.8 228.4
ours 55.23 54.90 36.82 121.7 11.2
Figure 11 Training change comparison curve.

To comprehensively validate the effectiveness of the CDFA and MANet-FasterCGLU fusion model (referred to as YOLOv11n-all in Table 3), feature maps from multiple neck layers were systematically extracted and visualized. As illustrated in the Fig. 12, the rows represent feature activations from shallow (top) to deep (bottom) layers, with the left column showing the fusion model (CDFA+MANet-FasterCGLU) and the right column showing CDFA alone.

Table 3 Comparison test of target detection modules.
Model F1-score mAP50/% mAP50-95/% FPS GFLOPs
RT-DETR-X 34.46 31.68 20.08 25.2 222.5
RT-DETR-r50 36.51 33.15 20.68 31.3 125.6
Faster-R-CNN 43.75 41.95 30.24 18.5 134.3
SSD 37.5 38.55 27.54 42.7 62.8
CenterNet 41.23 40.18 28.65 28.3 96.5
Transformer 38.92 37.84 26.73 15.2 168.9
YOLO-NAS 51.47 50.32 33.58 89.4 28.6
YOLOv5s 52.84 49.66 32.27 193.4 18.7
YOLOv6s 50.63 49.80 33.10 143.3 42.8
YOLOv7 53.28 52.04 34.21 118.6 10.1
YOLOv8s 54.10 52.07 34.29 195.2 23.4
YOLOv9s 55.53 53.70 35.17 83.6 22.1
YOLOv10s 48.63 49.40 32.34 187.6 21.4
ours 55.23 54.90 36.82 121.7 11.2
Figure 12 Training change comparison curve.

The visualization results indicate that CDFA alone (right column) exhibits limited capability in effectively integrating multi-scale information across different feature depths, and shows weaker discrimination between vehicle targets and background regions. In contrast, the fusion model (left column) presents significantly enhanced feature representations: Shallow layers better preserve fine-grained spatial details and vehicle boundaries, while deep layers maintain strong high-level semantic discrimination. This complementary multi-scale fusion mechanism enables superior target–background separation by jointly exploiting geometric cues from shallow layers and semantic features from deep layers.

In addition, the performance of the improved model was compared with that of mainstream detectors on the same dataset. The experimental results are detailed in Table 3.

Experiments show that the improved algorithm outperforms the mainstream object detection methods in both mAP and F1-score on public datasets. It has the advantages of strong adaptability to vehicle detection and lightweight, and is suitable for deployment on mobile/embedded platforms.

The following presents the ablation and comparison experiments of the target tracking algorithm to verify the effectiveness of the improved tracking algorithm (the results are shown in Table 4, and the training curves are shown in Figs. 13 and 14).

Table 4 Comparative experiments of multi-object tracking algorithms.
Model MOTA/% GFLOPs FPS
SORT 49.0 - 110.2
BOTSORT 61.1 - 110.2
ByteTrack 59.0 - 110.2
DeepSORT 61.3 11,448,768 90.1
DeepSORT+Group conv 61.7 951,744 100.3
DeepSORT+MPDIoU 62.9 11,448,768 90.1
DeepSORT+all 62.1 951,744 100.3
Figure 13 Loss change curve during the training process.
Figure 14 Top-1 error change curve during the training process.

Experiments show that all the improved modules have achieved the design goals: grouped convolution significantly reduces the number of parameters and computational complexity to enhance real-time performance; MPDIoU enhances target matching accuracy and optimizes tracking performance—the two works together to achieve comprehensive optimization. The proposed multi-object tracking algorithm has an advantage in MOTA while reducing the number of parameters. Among them, the MOTA is improved by 0.8% compared with the benchmark DeepSORT algorithm, verifying the effectiveness of the improved scheme. It can also be seen from the training curve that the improved model is superior to the original model in terms of Top1err error and loss.

4.4 Real-vehicle verification

To verify the proposed solution's applicability in actual traffic scenarios, this study used a professional data collection vehicle equipped with an onboard camera to conduct road tests and collect real traffic environment data (covering low-light driving scenarios such as nighttime highways, night-time busy streets, rainy days, dusk, and tree-shaded areas). The data collection vehicle is shown in Fig. 15.

Figure 15 Image enhancement comparison chart.

(1) Effectiveness verification of vehicle detection

By comparing the Eigen-CAM heatmap [34] (as shown in Fig. 16), the improved algorithm demonstrates a stronger vehicle positioning capability: compared with the baseline algorithm where the activation area is scattered and the heat map shows scattered aggregation (Fig. 16(a)), the improved algorithm forms a focused and continuous heat distribution around the target vehicle (Fig. 16(b)), which intuitively demonstrates its enhanced feature discrimination ability and more accurate positioning and recognition effect, verifying the optimization of detection performance.

Figure 16 Data collection vehicle demonstration: (a) original; (b) improved.

To validate the practicality of the proposed vehicle detection algorithm under real-world nighttime conditions, high-traffic video streams were collected using on-road vehicles and divided into two parts for analysis. The upper section evaluates detection performance with image enhancement, while the lower section examines the improved algorithm under the same settings [35].

Figure 17(a, b) compares the original and enhanced detection results, respectively, showing that image enhancement significantly improves vehicle detection accuracy in low-light environments. To further assess performance in complex conditions, Fig. 17(c, d) presents intersection scenarios, where the improved algorithm effectively reduces false positives and false negatives, demonstrating superior robustness and reliability across diverse nighttime traffic conditions.

Figure 17 Comparison image of image enhancement and YOLO improvement. (a) Detection results without image enhancement algorithms. (b) Increase the detection effectiveness of image enhancement. (c) YOLOv11 and image enhancement detection results. (d) HH-YOLOv11 and image enhancement detection results.

(2) Validation of vehicle tracking performance

To verify the practicality of the proposed tracking improvement method in real-world scenarios, we collected a 20-min video using an actual driving vehicle for real-world scenario analysis. Figure 18 shows a short segment of the recorded video, comparing the detection and tracking performance of the HH-YOLOv11 model with the original and improved versions of DeepSORT. The upper side represents the original scheme, which exhibits significant identity switching and trajectory loss, while the lower side corresponds to the improved method, showing stable identity association and continuous trajectory tracking. These results effectively confirm that the improved method has enhanced stability and robustness in multi-target tracking tasks under real-world traffic conditions [36].

Figure 18 Comparison of actual vehicle detection and tracking results. (a)–(c) Detection and tracking effects of the combination of HH-YOLOv11 with the original DeepSORT algorithm. (d)–(f) Comparison of actual vehicle detection and tracking results.

5 Conclusion

This study proposes an innovative integration of a lightweight multi-target tracking algorithm that significantly enhances detection accuracy and tracking stability in complex driving environments. Building upon this foundation, a real-time vehicle distance early warning system with dynamic safety thresholds is developed. Both public dataset evaluations and real-vehicle experiments verify that the proposed method effectively improves detection precision, increases ID retention stability, and reduces model complexity, thereby offering a reliable and efficient perception module for intelligent connected vehicles.

Beyond its technical contributions, the proposed approach demonstrates strong potential for application in intelligent connected vehicle systems, particularly in low-light perception, real-time active safety warning, and lightweight on-board deployment. Moreover, it provides a feasible pathway toward vehicle–road cooperative perception, supporting the seamless integration of sensing and decision-making in complex traffic environments.

Future research will focus on enhancing image enhancement and multi-modal data fusion under challenging conditions such as rain, fog, and nighttime illumination, improving long-term tracking stability under occlusion, and exploring lightweight deployment strategies for resource-constrained embedded platforms to further improve real-time performance and energy efficiency. These directions will strengthen the practical value and engineering applicability of the proposed system in next-generation intelligent transportation and cooperative vehicle networks.

 Acknowledgements

Acknowledgements

This work was supported in part by the National Natural Science Foundation of China (No. 52172389).

 Conflict of interests

The authors declare no conflict of interest.

 Author contributions

Conceptualization, Zhe Zhang and Wei Zhong; methodology, Hengkai Wang and Xiao Luo; software, Hengkai Wang and Zhe Zhang; validation, Hengkai Wang, Zhe Zhang, Xiao Luo, Wei Zhong, Hong Qi, and Dong Zhang; formal analysis, Xiao Luo and Hong Qi; investigation, Zhe Zhang; resources, Wei Zhong and Hengkai Wang; data curation, Xiao Luo and Zhe Zhang; writing—original draft preparation, Xiao Luo; writing—review and editing, Zhe Zhang and Wei Zhong; visualization, Xiao Luo and Dong Zhang; supervision, Wei Zhong; project administration, Wei Zhong; funding acquisition, Wei Zhong. All authors have read and agreed to the published version of the manuscript.

References

[1] 

W. Yue, C. Li, P. Duan, and F.-R. Yu, "Revolution on wheels: A survey on the positive and negative impacts of connected and automated vehicles in the era of mixed autonomy," IEEE Internet Things Journal, vol. 10, no. 24, pp. 21820–21835, 2023.

[2] 

J. Xu, Y. Hou, D. Ren, L. Liu, F. Zhu, M. Yu, H. Wang, and L. Shao, "STAR: A structure and texture aware retinex model," IEEE Transactions on Image Processing, vol. 29, pp. 5022–5037, 2020.

[3] 

C. Li, C. Guo, and C. C. Loy, "Learning to enhance low-light image via zero-reference deep curve estimation," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 8, pp. 4225–4238, 2022.

[4] 

H. Zhao, Y. Gao, and W. Deng, "Defect detection using shuffle Net-CA-SSD lightweight network for turbine blades in IoT," IEEE Internet of Things Journal, vol. 11, no. 20, pp. 32804–32812, 2024.

[5] 

S. Yang, Y. Li, J. Nie, S. Ercisli, and T. R. Gadekallu, "Enhanced brain tumor detection using DCGAN augmentation and optimized efficientDet in IoT-based healthcare industry 5.0," IEEE Internet of Things Journal, vol. 12, no. 22, pp. 46093–46103, 2025.

[6] 

Z. Tian, H. Liu, J. Wu, W. Chen, R. Zheng, and Z. Wang, "PrFu-YOLO: A lightweight network model for UAV-assisted real-time vehicle detection toward an IoT underlayer," IEEE Internet Things Journal, vol. 11, no. 23, pp. 37536–37549, 2024.

[7] 

Y. Xie, D. Du, and M. Bi, "YOLO-ACE: A vehicle and pedestrian detection algorithm for autonomous driving scenarios based on knowledge distillation of YOLOv10," IEEE Internet Things Journal, vol. 12, no. 15, pp. 30086–30097, 2025.

[8] 

E. Yıldız and M. Keskinöz, "Efficient FCOS model for fire and smoke detection," 2025 9th International Symposium on Innovative Approaches in Smart Technologies (ISAS), pp. 1–5, 2025.

[9] 

X. Wu and Q. Xue, "An improved corner net-lite method for pedestrian detection of unmanned aerial vehicle images," 2021 China Automation Congress (CAC), pp. 2322–2327, 2021.

[10] 

Q. Gu, H. Huang, Z. Han, Q. Fan, and Y. Li, "GLFE-YOLOX: Global and local feature enhanced YOLOX for remote sensing images," IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1–12, 2024.

[11] 

F. Li, H. Zhang, S. Liu, J. Guo, L. M. Ni, and L. Zhang, "DN-DETR: Accelerate DETR training by introducing query denoising," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 4, pp. 2239–2251, 2024.

[12] 

H. Li, Q. Wang, and Z. Li, "Investigating walking characteristics of passengers in subway corridors by video processing based on YOLOX and DeepSORT algorithm," Proc. China Autom. Congr. (CAC), pp. 5766–5771, 2022.

[13] 

Y. Liu, F. Liu, Y. Huang, J. Hu, W. Zhang, and Y. Hou, "The real-time pavement distress detection system based on edge-cloud collaborative computing," IEEE Trans. Intell. Transp. Syst., vol. 26, no. 7, pp. 10512–10522, 2025.

[14] 

Y. Jiang, X. Gong, D. Liu, Y. Cheng, C. Fang, and X. Shen, "EnlightenGAN: Deep light enhancement without paired supervision," IEEE Trans. Image Process., vol. 30, pp. 2340–2349, 2021.

[15] 

R. Liu, L. Ma, J. Zhang, X. Fan, and Z. Luo, "Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement," Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 10556–10565, 2021.

[16] 

L. Ma, T. Ma, R. Liu, X. Fan, and Z. Luo, "Toward fast, flexible, and robust low-light image enhancement," Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 5627–5636, 2022.

[17] 

W. Wu, J. Weng, P. Zhang, X. Wang, W. Yang, and J. Jiang, "Interpretable optimization-inspired unfolding network for low-light image enhancement," IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 4, pp. 2545–2562, 2025.

[18] 

D. C. Khrisne, M. Sudarma, I. A. Dwi Giriantari, and D. Made Wiharta, "Nighttime traffic surveillance using glare reduction and Zero-DCE-based low-light image enhancement," Proc. Int. Conf. Smart-Green Technol. Electr. Inf. Syst. (ICSGTEIS), pp. 161–165, 2023.

[19] 

J. Shi, S. Ruan, Y. Tao, Y. Rui, J. Deng, P. Liao, and P. Mei, "Improved YOLO algorithm based on multi-scale object detection in haze weather scenarios," CHAIN, vol. 2, no. 2, pp. 183–197, 2025.

[20] 

M. Lei, H. Wu, X. Lv, and X. Wang, "CondSeg: A general medical image segmentation framework via contrast-driven feature enhancement," Proc. AAAI Conf. Artif. Intell. (AAAI), vol. 39, no. 5, pp. 4571–4579, 2025.

[21] 

W. Li, H. Guo, X. Liu, K. Liang, J. Hu, Z. Ma, and J. Guo, "Efficient face super-resolution via wavelet-based feature enhancement network," Proc. 32nd ACM Int. Conf. Multimedia (ACM MM), pp. 4515–4523, 2024.

[22] 

Y. Feng, J. Huang, S. Du, S. Ying, J. Yong, and Y. Li, "Hyper-YOLO: When visual object detection meets hypergraph computation," IEEE Trans. Pattern Anal. Mach. Intell., vol. 47, no. 4, pp. 2388–2401, 2025.

[23] 

J. Chen, S. Kao, H. He, W. Zhou, S. Wen, and C. LEE, "Run, don’t walk: Chasing higher FLOPS for faster neural networks," Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 12021–12031, 2023.

[24] 

D. Shi, "TransNeXt: Robust foveal visual perception for vision transformers," Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 17773–17783, 2024.

[25] 

Y. Wu and K. He, "Group normalization," Proc. Eur. Conf. Comput. Vis. (ECCV), pp. 3–19, 2018.

[26] 

Z. Tian, C. Shen, H. Chen, and T. He, "FCOS: A simple and strong anchor-free object detector," IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 4, pp. 1922–1933, 2022.

[27] 

Z. Chen, Z. He, and Z.-M. Lu, "DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attention," IEEE Trans. Image Process., vol. 33, pp. 1002–1015, 2024.

[28] 

F. Zhang, H. Liu, Z. Chen, L. Wang, Q. Zhang, and L. Guo, "Collision risk warning model for construction vehicles based on YOLO and DeepSORT algorithms," J. Constr. Eng. Manag., vol. 150, no. 6, 2024.

[29] 

J. Xu, Y. Sun, S. Zhang, D. Sun, and Z. Xiang, "Wheat-YOLO: A real-time and high precision object detection for wheat," Proc. IEEE Int. Conf. Syst., Man, Cybern. (SMC), pp. 2165–2170, 2024.

[30] 

K. J. de Jesus, M. O. K. Pereira, L. R. Emmendorfer, and D. F. T. Gamarra, "A comparison of visual SLAM algorithms ORB-SLAM3 and DynaSLAM on KITTI and TUM monocular datasets," Proc. 36th SIBGRAPI Conf. Graph., Patterns Images (SIBGRAPI), pp. 109–114, 2023.

[31] 

F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, and F. Liu, "BDD100K: A diverse driving dataset for heterogeneous multitask learning," Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 2633–2642, 2020.

[32] 

S. Rujikietgumjorn and N. Watcharapinchai, "Vehicle detection with sub-class training using R-CNN for the UA-DETRAC benchmark," Proc. 14th IEEE Int. Conf. Adv. Video Signal Based Surveillance (AVSS), pp. 1–5, 2017.

[33] 

X. Xiong, M. He, T. Li, G. Zheng, W. Xu, and X. Fan, "Adaptive feature fusion and improved attention mechanism-based small object detection for UAV target tracking," IEEE Internet Things Journal, vol. 11, no. 12, pp. 21239–21249, 2024.

[34] 

M. B. Muhammad and M. Yeasin, "Eigen-CAM: Class activation map using principal components," Proc. Int. Joint Conf. Neural Netw. (IJCNN), pp. 1–7, 2020.

[35] 

B. Gao, R. Mei, Y. Lu, D. Chu, S. Zhang, Z. Wang, and K. Li, "Predictive lane-changing control for platoon based on cloud control system in highway scenarios," CHAIN, vol. 1, no. 1, pp. 75–98, 2024.

[36] 

J. Huang, K. Wan, J. Chen, J. Zhou, W. Zhong, and B. Gao, "A forward-looking sequential asynchronous lane-change strategy for vehicle platoon under cloud-vehicle-road integrated architecture," CHAIN, vol. 2, no. 1, pp. 43–56, 2025.

Top