ContentsFigures & Tables
1 Introduction

1 Introduction

2 Problem definition and workflow

2 Problem definition and workflow

2.1 Problem definition

2.1 Problem definition

2.2 Pedestrian crossing intention prediction workflow

2.2 Pedestrian crossing intention prediction workflow

3 Data perception

3 Data perception

3.1 Image perception

3.1 Image perception

3.2 Vehicle speed perception

3.2 Vehicle speed perception

4 Input data representation

4 Input data representation

4.1 Input data representation

4.1 Input data representation

4.1.1 Pedestrian Information

4.1.1 Pedestrian Information

4.1.2 Ego-vehicle information

4.1.2 Ego-vehicle information

4.1.3 Environment information

4.1.3 Environment information

4.1.4 Other information

4.1.4 Other information

4.2 Different model input data representation for pedestrian crossing intention

4.2 Different model input data representation for pedestrian crossing intention

5 Model

5 Model

5.1 Prediction model

5.1 Prediction model

5.1.1 CNN-based model

5.1.1 CNN-based model

5.1.2 RNN-based model

5.1.2 RNN-based model

5.1.3 GCN-based model

5.1.3 GCN-based model

5.1.4 Transformer-based model

5.1.4 Transformer-based model

5.1.5 Reinforcement learning-based model

5.1.5 Reinforcement learning-based model

5.1.6 Fusion structure-based model

5.1.6 Fusion structure-based model

5.2 The model complexity and prediction efficiency

5.2 The model complexity and prediction efficiency

5.3 The model applications

5.3 The model applications

6 Model validation and testing

6 Model validation and testing

6.1 Dataset

6.1 Dataset

6.2 Evaluation metrics

6.2 Evaluation metrics

6.3 Model validation

6.3 Model validation

7 Challenges and future directions

7 Challenges and future directions

7.1 Model input data representation

7.1 Model input data representation

7.2 Model generalization ability

7.2 Model generalization ability

7.3 The uncertainty for the pedestrian crossing intention prediction

7.3 The uncertainty for the pedestrian crossing intention prediction

7.4 The data privacy and ethical considerations

7.4 The data privacy and ethical considerations

8 Conclusions

8 Conclusions

References

References

Pedestrian crossing intention prediction in the wild: A survey

Yancheng Ling1Zhenliang Ma1
KTH Royal Institute of Technology Ringgold Standard Institution, Brinellvägen 23, Stockholm, Stockholm 11428, Sweden
Abstract: In real-world driving scenarios, understanding the intentions of pedestrians in real-time is critical for the built environment safety when operating intelligent vehicles on the roads. Pedestrians crossing the street is a common behavior that can easily lead to accidents. This paper presents a comprehensive review of the prediction of pedestrian crossing intentions, focusing on data, model structure, data representation, information extraction, prediction function, and associated models and challenges. The review highlights that data types, model generalization ability, and prediction uncertainty are key challenges on pedestrian crossing intention prediction. It identifies open challenges and opportunities for future research in pedestrian crossing intention prediction.
Keywords: pedestrian crossing intention prediction; data representation; information extraction; prediction uncertainty
Received: 2024-06-07

1 Introduction

In traffic scenarios, pedestrians are an important component. Due to the lack of adequate protection and limited physical impact resistance and risk avoidance capabilities, they are one of the vulnerable groups in the traffic system, and their safety is also a key focus of traffic safety research. Additionally, pedestrian behavior exhibits significant randomness and unpredictability. For example, pedestrians may walk at different speeds, suddenly change direction, or suddenly stop. Therefore, predicting pedestrian intentions is crucial for safe driving, a cornerstone of the automatic decision-making process in autonomous vehicles [1–3].

Many studies have been conducted on pedestrian intentionprediction, encompassing areas such as pedestrian gaze detection [4–6], trajectory prediction [7–9], pedestrian crossing intention prediction [1, 10, 11], and collision avoidance [12]. Among these, pedestrian crossing intention prediction has garnered significant attention. Pedestrian crossing, a common road activity, represents a critical interaction between vehicles and pedestrians and is a focal point for traffic accidents. Therefore, developing the ability of autonomous vehicles to predict pedestrian crossing intentions is essential for safe driving.

In data collection for driving scenes, existing works on pedestrian crossing prediction include: (1) ego-view: capturing video or data points from a first-person perspective. (2) Bird's eye view (BEV): observing the scene from a top-down perspective, commonly aligned with the world coordinate system. Due to the efficiency in response that the ego-view provides for autonomous vehicles, most existing works on pedestrian crossing intention prediction are based on this perspective. These works analyze pedestrian pose, movement, and the surrounding traffic environment for pedestrian crossing intention prediction, as shown in Fig. 1.

Figure 1 Pedestrian crossing intention prediction in an ego-view context is based on the analysis of pedestrian pose, movement, and the surrounding traffic environment.

From a data perspective, Daimler [13], JAAD [14], PIE [15], STIP [16], BPI [17], PSI [18], PePScenes [19], and some self-constructed datasets [20, 21] are commonly used for model validation. For fair comparison, most proposed models are tested on the JAAD and PIE datasets. Both the JAAD and PIE datasets provide annotated information, including bounding boxes, vehicle speed, and pedestrian intentions (crossing or not-crossing). Additionally, more information can be extracted from these datasets, offering richer input for pedestrian crossing intention prediction.

Many survey studies have been conducted on pedestrian intention prediction. However, most of these works involve various datasets or road agents (pedestrians, vehicles, or human drivers), complicating direct comparisons between model inputs and models themselves. The model input representation and model structure are crucial for the practical performance of the proposed methods. Ensuring robustness, stability, and effectiveness in real-world scenarios depends on these aspects. Therefore, conducting a survey on pedestrian intention prediction that focuses on input data representation and model architecture is essential. Complementary to these surveys, we concentrate on the latest progress in pedestrian intention prediction on JAAD and PIE dataset. Our main objective is to draw inspiration for pedestrian intention prediction research from input data representation, model structure and challenges.

2 Problem definition and workflow

2.1 Problem definition

As is shown in Fig. 2, pedestrian crossing intention prediction can be viewed as a binary classification task [22]. In mathematical terms, the goal is to predict the crossing intention I∈{0, 1} of a pedestrian i at a future time t + n, where n∈{30, 60} (approximately 1 to 2 seconds). For instance, this involves predicting whether a pedestrian will cross the street within the next 1 to 2 seconds by analyzing 16 frames of video images.

Figure 2 Illustration for pedestrian crossing intention prediction. The variable m represents the observation length. For instance, when m is equal to 16, it indicates an observation length of 16 frames used to predict whether pedestrians are crossing the street in the future 1 to 2 s.

2.2 Pedestrian crossing intention prediction workflow

As shown in Fig. 3, the main workflow for a pedestrian crossing intention prediction system, which contains five key steps.

Figure 3 The workflow of pedestrian crossing intention prediction includes the following steps: perception, raw data, input data representation, model, and result.

(1) Perception: It uses sensors to perceive the surrounding environment. For example, it uses an onboard camera to capture the ego-view image and an onboard sensor to obtain the vehicle speed.

(2) Raw data: This data is collected and transmitted from sensors.

(3) Input data representation: It is the processed data derived from raw data that can be directly fed to the model for training and testing. For example, it includes input data such as bounding boxes, vehicle speed, pedestrian image crops, etc.

(4) Model: It is the model for pedestrian crossing intention prediction, which can be classified into different types in terms of the based model structure, such as CNN-based and RNN-based models.

(5) Result: The final prediction results indicate whether the pedestrian intends to cross or not.

3 Data perception

Data Perception is a crucial step for pedestrian crossing intention prediction. Among the various types of Perception data, images and vehicle speed are the most commonly used.

3.1 Image perception

The image is primarily captured by the camera and contains rich scene information that can be used to extract pedestrian information regarding actions and movements. Additionally, it is useful for understanding traffic environment information. The quality of the image and the number of perspectives are important factors for accurate prediction.

The higher image quality can capture more detailed texture information over long distances and improve prediction accuracy. However, image quality is influenced by factors such as the quality of the camera and weather conditions. Higher-quality cameras are typically more expensive, and image quality can be easily compromised on rainy days. Generally, the most commonly used camera resolution in public datasets is 1920 × 1080 with a frame rate of 30 FPS [14, 15].

The number of perspectives is also an important factor for prediction. The most commonly used perspective is the front view, but some datasets also provide left and right perspectives [23]. Compared to a single perspective, using multiple perspectives can increase the perception range. However, it requires the design of more complex algorithms to integrate information from different perspectives.

3.2 Vehicle speed perception

Vehicle speed is important sensing data for pedestrian crossing intention prediction, as it provides information about the movement of the vehicle. Additionally, it can indirectly provide other information, such as traffic flow conditions, since vehicles tend to maintain similar speeds within traffic flow to prevent collisions.

The vehicle speed can be obtained from onboard sensors or GPS. Compared to GPS data, onboard sensors can provide accurate speed information without delay. Therefore, onboard speed is the most commonly used data for prediction.

4 Input data representation

Different types of the model input data representation significantly influence the performance of the proposed methods. This section will introduce the various data model input data representation.

4.1 Input data representation

The input information for the model can be categorized into four types: pedestrian information, ego-vehicle information, environment information, and other information. Some example are shown in Fig. 4. The main preprocessing process and the technologies used to extract model input data representation from the raw data are shown in Fig. 5.

Figure 4 Some examples of model input data types include skeleton data, bounding boxes, vehicle speed, image crops, global images, and semantic segmentation maps.
Figure 5 The model input data representation preprocessing process and the technologies used from the raw data.

4.1.1 Pedestrian Information

This helps the model capture details related to the pedestrian's pose and movements. The primary input data representation include image crops of the pedestrian, pedestrian skeleton data, the state of the pedestrian (e.g., walking or not), the intent of the pedestrian (e.g., looking or not), the pedestrian's observed motion information using bounding boxes, the 3d orientation of head,etc.

Bounding box (B): The bounding box serves as a valuable tool for capturing the pedestrian's trajectory information across time, as well as providing insights into the pedestrian's distance based on the size of the box. The bounding box for a pedestrian can be extracted from the raw global image data using an object detection algorithm [24–26]. The primary method to extract features from the bounding box is to treat it as discrete sequence data and use an RNN model for feature extraction [27]. The second approach involves creating a graph based on the physical structure of the bounding box and extracting features from that graph based on the GCN model [28]. Compared to treating it as discrete sequence data, using a bounding box graph can better preserve the spatial structure and enhance the spatial features.

Image crop (I): Image cropping can provide appearance and pose information of a pedestrian. This information is extracted from the raw image using the bounding box and the global image data. The primary methods to extract features include combining a 2D CNN-based model with an RNN-based model [27], using a 2D CNN model to extract features from the image and then fusing them into a GCN model with an attention module [29], or using the 2D CNN to extract features from the image data and embedding them into a Transformer-based model [30]. Additionally, a 3D CNN-based model can also be used to extract features from the cropped pedestrian image [31].

Pedestrian skeleton (K): The skeleton data can provide detailed information about the pose and actions of a pedestrian. It is extracted from the pedestrian image crop using a human key point detection algorithm [32–35]. Compared to raw image crops, the skeleton data is more robust against variations in image quality because the human key points have intrinsic geometric constraints [36]. These constraints help predict the location of the key points even when parts of the body are obstructed. This robustness makes skeleton data highly valuable for applications that require reliable human pose and action recognition, even in challenging visual conditions. The main methods to extract features from skeleton data include: (1) Treating the skeleton data as discrete data and using an RNN-based model as the backbone, which can process sequences of data, making them suitable for modeling temporal dependencies in the movement of joints over time [27]. (2) Using a GCN model as the backbone: GCN are particularly effective for non-euclidean data, such as skeleton data. GCNs can preserve and exploit the geometric relationships between key points, making them highly suitable for extracting features from this type of data [37, 38]. (3) Using a transformer model as the backbone: Transformer models can also be applied to skeleton data [30]. They are known for their ability to handle long-range dependencies and have been success- fully used in various sequence modeling tasks. Among these methods, the GCN model is especially advantageous for non-euclidean data and excels at preserving geometric information, making it a powerful choice for extracting features from skeleton data.

Looking or not (L): The intent data provides information about whether a pedestrian is paying attention to a vehicle or not. This data is extracted from the pedestrian crop image using an eye detection model [4–6, 39]. A fully connected layer can then be used to extract features from this data [14]. By analyzing the direction and focus of the pedestrian's gaze, the system can determine their level of attention to the vehicle, which is crucial for pedestrian crossing intention prediction.

Walking or not (W): The intent data provides information about whether a pedestrian is standing or walk. This data is extracted from the pedestrian crop image using an walking detection model. A fully connected layer can then be used to extract features from this data [14].

Bounding box center (CB): The bounding box center sequence data can provide motion information about the pedestrian. This data is extracted from the bounding box itself, utilizing the transform backbone as the primary backbone [40].

Bounding box rate (RB): The ratio of bounding box data represents the proportion between the width and height of the bounding box. This ratio provides information about the pedestrian's size and shape. It is extracted from the bounding box, utilizing the transform backbone as the primary backbone [40].

Bounding box area (AB): The area of bounding box data represents the area of the bounding box. It is extracted from the bounding box, utilizing the transform backbone as the primary backbone [40].

3D head orientation (HO): The 3D head orientation provides information about the pedestrian's head pose. This data is extracted from a cropped image of the pedestrian using a face pose estimation model [41–44], which employs the transform as the backbone to extract features.

Local motion of pedestrian (LM): The local motion of a pedestrian can provide detailed movement information. This data is extracted from the pedestrian's cropped image using the optical flow approach [45]. It combines a 2D CNN model and a GRU model to extract features from the local motion image [23].

4.1.2 Ego-vehicle information

This helps the model capture the vehicle's movements, particularly its speed. The main input data representation for this category is the vehicle speed.

Vehicle speed (S): Vehicle speed can provide information about the vehicle's movement and traffic flow. This data can be obtained from GPS or a vehicle speed sensor. Compared to GPS, the vehicle speed sensor offers more precise speed information, making it more widely used. The main methods for extracting features from vehicle speed data include:

(1) RNN-based model: This method uses an RNN-based model as the backbone to extract features from sequence data, treating the speed data as discrete data [27].

(2) Transform-based model: This approach employs a transform-based model to extract features. A motion embedding layer and a multi-layer perceptron (MLP) are used to embed the speed data before feeding it into the transform model [30].

(3) GCN-based model: There are two primary ways to use a GCN-based model as the backbone for feature extraction: The first method uses a 1D convolution layer to extract features, followed by an attention mechanism to embed these features into the GCN network [29]. The second method involves building a vehicle speed graph and using the GCN model to extract features directly from the speed data. In previous work, a virtual node representing pedestrian movement speed was introduced to build a speed graph that includes vehicle speed [28].

4.1.3 Environment information

This allows the model to capture details about the surroundings of the pedestrian, such as traffic flow and crosswalk conditions. The main input data representation contains the global image, semantic segmentation map, and the surrounding environment image around the pedestrian.

Global image (G): The global image sequence provides information about the traffic environment and is obtained from an onboard camera. The main methods to extract features from it include:

(1) Combining a 2D CNN-based model with an RNN-based model: This method extracts features from the sequence data by combining a 2D CNN and an RNN [27].

(2) Combining a 2D CNN-based model with a GCN-based model: This approach uses a 2D CNN to extract features from the global image and then employs an attention mechanism to embed these features into the GCN backbone [29].

(3) Using a 3D CNN-based model: This method extracts features directly using a 3D CNN [31].

(4) Using a transformer model: This approach employs a transformer to extract features from the sequence data [30].

Semantic segmentation map (SGM): The semantic segmentation map provides detailed information about the traffic environment. Compared to the global image, it offers a more precise classification for each element, such as lanes, vehicles, pedestrians, and traffic signs. The semantic segmentation map is derived from the global image [46–49], and the methods to extract it are the same as those used for the global image.

Local image crop (LI): The local image crop provides detailed information about the pedestrian and the surrounding traffic environment. It is derived from the global image using a bounding box, and the methods to extract features from it are the same as those used for the global image [27].

4.1.4 Other information

Other information is required to capture details related to pedestrian crossing intention prediction. The main input data representation includes the relative distance between the target pedestrian and the ego-vehicle, the depth map, and other relevant metrics.

Depth map (DM): The depth map provides distance information for traffic elements. It is obtained from the global image using a depth estimation model [50, 51]. The main methods to extract features from the depth map are similar to those used for the global image.

4.2 Different model input data representation for pedestrian crossing intention

As is shown in Table 1, in terms of the number of input features, the model input data representation can be classified as either single-modality data or multi-modality data. single-modality data primarily relies on pedestrian information. Many early studies used image crops of the pedestrian as inputs [14, 52]. However, the performance of these models was limited by the image quality. To overcome this limitation, more recent models use skeleton data as inputs, which has shown significant improvement compared to using single image data [37, 53]. Compared to the image data, skeleton data can rely on the inherent geometric relationships between nodes to enhance the accuracy of skeleton prediction, thereby improving pose recognition accuracy [54, 55]. However, the single modality data could not capture the vehicle motion information and traffic environment information, which limits their performance.

Table 1 The results comparing model performance on the PIE and JAAD datasets are shown. JAADbeh refers to a subset of the JAAD dataset that includes behavioral labels (only pedestrians who interact with the ego-vehicle), while JAADall encompasses all detected pedestrians. "Acc" stands for accuracy.
Ref. Year Book title Model variant Input data PIE JAADbeh JAADall
Acc Acc Acc
Simonyan et al. [52] 2014 NIPS VGG16 I,OF 64 56 60
Shi et al. [59] 2015 NIPS VGG19+LSTM I 58 49 63
NIPS ResNet50+LSTM I 54 59 70
Yue et al. [60] 2015 CVPR GRU I,OF 82 60 79
Du et al. [61] 2015 CVPR GRU K 82 53 80
Tran et al. [62] 2015 ICCV 3DConv I 77 61 84
Carreira et al. [31] 2017 ICCV 3DConv I 80 62 81
CVPR Opticflow+3DConv I,OF 81 62 84
Bhattacharyya et al. [63] 2018 CVPR GRU - 83 61 79
Cadena et al. [53] 2019 ITSC GCN K 76 62 80
Kotseruba et al. [56] 2020 IV GRU I,B,S,Int 83 58 65
IV LSTM I,B,S,Int 81 51 78
Rasouli et al. [64] 2020 BMVC GRU I,G,K,B,S 82 51 84
Piccoli et al. [65] 2020 ACSSC DenseNet B,K - 59 60
Kotseruba et al. [22] 2021 WACV 3DConv+RNN I,B,K,S 86 50 70
Yang et al. [27] 2021 TIV VGG+GRU I,B,K,S,SGM - 62 83
Gesnouin et al. [66] 2021 FG GRU B,K,S,ED 88 64 85
Cadena et al. [29] 2022 TITS Conv+GCN I,K,S,SGM 89 70 86
Zhang et al. [38] 2022 TITS ST-GCN K - 63 -
Zhou et al. [30] 2023 TITS Transformer I,K,S,B,G 91 70 87
Ling et al. [37]. 2023 ITSC STA-GCN K - 69 -
Dong et al. [57] 2023 ACML GRU+attention K,S,EM,SGM,RD,LP,B 91 67 90
Yang et al. [67] 2023 TITS GCN+CNN K,G 91 71 89
Wang et al. [58] 2023 APSIPA ASC Transformer K,B,S,I,G,HO,SE 91 68 89
Ling et al. [28] 2024 TITS STA-GCN K,S,B 91 69 89
Azarmi et al. [23] 2024 arXiv GRU+attention K,S,B,I,LM,SGM,CD 91 - -
Zhang et al. [40] 2024 AAAI Transformer B,CB,AB,RB 93 - 91
Yang et al. [68] 2024 TTE GCN+GRU K,B,S 94 - 89
Chen et al. [69] 2024 TITS Transformer B,S 93 71 88
Lv et al. [70] 2024 IEEE Sensors Letters GRU K,B,S,I,G,SE 90 - -

To achieve better performance, more studies use multi-modality data. Kotseruba et al. used image crops of the pedestrian, bounding boxes, vehicle speed, and pedestrian intention (looking or not) as inputs [56]. In later work, they also incorporated skeleton data to obtain more robust pedestrian pose and motion information. Segmentation maps and global images are popular inputs for capturing environmental information [22].

Many researchers are designing novel model data representation to capture richer features. For instance, Dong et al. calculated the relative distance between the pedestrian and the vehicle, the location of the pedestrian, and pedestrian motion information [57]. They combined these inputs with skeleton data, speed, image crops of the pedestrian, and segmentation maps. Additionally, Wang et al. used the 3D head orientation of the pedestrian as input [58]. Azarmi incorporated additional local motion of the pedestrian and depth estimation images as inputs. Zhang et al. used the bounding box, the centre of the bounding box, the area of the bounding box, and the rate of the bounding box between width and height as input [40]. Their methods achieve the most competitive performance in PIE and JAADall dataset.

5 Model

Different model structure significantly influence the performance of the proposed methods. This section will introduce the various model structure used in these methods and their applications.

5.1 Prediction model

The prediction model can be classified as: CNN-based, RNN-based, GCN-based, Transformer-based, Reinforcement Learning based model, and Fusion structure based model. Some examples are shown in Fig. 6.

Figure 6 Some examples of different models used for predicting pedestrian crossing intentions include: (a) an RNN-based model [27], (b) a GCN-based model [28], (c) a transformer-based model [30], and (d) a fusion structure-based model [68].

5.1.1 CNN-based model

The CNN-based model utilizes various CNN backbones, including VGG16 [71], 3DConv [62], and ResNet50 [72], to extract features. Early research used CNN-based models to extract features from input data. Rasouli et al. separately utilized VGG16 and ResNet50 as backbones to extract features from pedestrian image crops and global images [14]. They also included pedestrian intention information (walking or not) as input. However, since the model used only the last frame as input, it could not capture temporal information, which limited its performance. To capture temporal information, 3DCNNs are used to extract spatial-temporal features from image sequences [62]. Compared to previous methods, this approach significantly improves model performance. However, most CNN-based models can only use image-related data as input, which limits their overall performance.

5.1.2 RNN-based model

The RNN-based model utilizes RNN backbones, including LSTM [73] and GRU [74], to extract features. Many of these models also use attention modules to fuse multi-modality features from different information flows. To extract spatial-temporal features and utilize modality data as input, RNN-based models have been introduced for pedestrian crossing prediction. Shi et al. used CNNs, specifically VGG19 and ResNet50, to extract features from pedestrian image crops and employed LSTM to capture the spatial information from the sequence of CNN output features [59]. Kotseruba et al. used 2D bounding boxes, pedestrian image crops, vehicle speed, and intention information as inputs [56]. They employed LSTM or GRU as the backbone network to extract spatial-temporal features and used a fully connected layer based on the last hidden state for pedestrian crossing intention prediction. To improve performance, Kotseruba et al. later included pedestrian image crops, bounding boxes, skeleton data, and vehicle speed as inputs [22]. They used 3DCNN to extract features from image sequences and RNN models to extract features from bounding boxes, skeletons, and vehicle speed. Finally, they applied modality attention to fuse the different types of features for making predictions. To capture environmental information, Yang et al. added segmentation maps as input [27]. They used GRU as the backbone to extract spatial-temporal features from input data and compared different combinations of data type streams. However, RNN-based models treat the skeleton as discrete data and lose its spatial geometric information. Additionally, these models contain a large number of parameters, which limits their real-time performance.

5.1.3 GCN-based model

The GCN-based models use GCN backbones [75] to extract spatial-temporal features from the input data. Compare to the CNN and RNN model, the GCN model is good at non euclidean data and can preserve geometric features of data [76]. Cadena et al. used a GCN model to extract pose information from skeleton data and employed a 2-layer GCN model for pedestrian crossing intention prediction [53]. To better extract spatial-temporal features, Zhang et al. introduced the spatial-temporal GCN as the backbone for pedestrian crossing intention prediction [38]. Ling et al. further enhanced the spatial-temporal GCN by incorporating spatial attention, temporal attention, and channel attention into the spatial-temporal GCN to extract features from sequence data more effectively [37]. However, using the pure skeleton data as input could not capture the vehicle movement and environmental information, which limited the performance of the model. Cadena et al. used pedestrian image crops, skeleton data, speed, and segmentation maps as inputs [29]. They employed a GCN model as the main backbone to extract features from the skeleton data. For the other data, they used a CNN model to extract features as a second information flow and then utilized an attention mechanism module to embed the extracted features into the GCN model. Additionally, they introduced a pose prediction model [77] to obtain the skeleton data. Ling et al. proposed a pure GCN model for pedestrian prediction [28]. They used skeleton data, bounding boxes, and vehicle speed as inputs and introduced an attention mechanism into the spatial-temporal GCN as the backbone. They designed a bounding box graph and a speed graph to enable the GCN model to directly handle bounding box and speed data. They also introduced modality attention to dynamically fuse features from different data types with a weighting mechanism.

5.1.4 Transformer-based model

Transformer-based models use transformer backbones [78] to extract spatial-temporal features from input data representations. Increasingly, research is utilizing transformers as the backbone for pedestrian crossing intention prediction. Zhou et al. proposed a Transformer-based model for this purpose, using bounding boxes, pedestrian image crops, skeleton data, ego-vehicle speed, and global images as inputs [30]. They proposed a temporal fusion block and a self-attention mechanism to collaboratively and progressively capture the dynamic spatio-temporal interactions between the pedestrian, vehicle, and environment. Wang et al. proposed a multi-modal pedestrian crossing intention prediction framework that used vehicle speed, global images, 3D head orientation, bounding boxes, skeleton data, pedestrian image crops, and surrounding environment data as inputs [58]. They developed a feature pre-processing module to identify traffic lights, road signs, and crosswalks from global images, representing traffic environment information in a novel way that allows efficient exploitation. Zhang et al. proposed a novel Transformer-based prediction model [40]. They used bounding boxes, the center of the bounding box, the area of the bounding box, and the aspect ratio of the bounding box as inputs. In addition, they designed a transformer module to examine the temporal relationships among input features in pedestrian video sequences and implemented a deep evidential learning model to address AI uncertainty in complex scene scenarios.

5.1.5 Reinforcement learning-based model

The Reinforcment Learning-based model utilizes Reinforcement Learning to enhance prediction accuracy and efficiency. Dai et al. [79] employ Reinforcement Learning to generate soft labels for the training dataset, aiming to address observed data uncertainty and improve the model's performance and reliability. The results demonstrate that existing models perform better when predictive uncertainty is incorporated to learn more informative soft labels.

5.1.6 Fusion structure-based model

The Fusion structure-be applied in insurance and legal contexts, including accident based models use at least two types of backbones to extract analysis and insurance risk assessment. spatial-temporal features from the input data. Compared to single structure-based models, Fusion structure-based models use different backbones for spatial-temporal feature extraction for pedestrian crossing intention prediction. Early research combined CNN and RNN models as the backbone, with the most used dataset for pedestrian cross intention predic-CNNs extracting features from visual sequences and RNNs processing non-visual sequences [22]. GCN models have also been combined with other modules as the backbone. Yang et al used bounding boxes vehicle speed, and skeleton data as inputs[68]. They employed GRU to extract features from the bounding box and vehicle speed data, while the GCN model extracted features from the skeleton data. They proposed an efficient feature fusion module to explore the complementary multi-modal inputs.

5.2 The model complexity and prediction efficiency

Table 2 shows the complexity and prediction efficiency of various state-of-the-art (SOTA) models. The comparison includes model prediction time (M) and the total execution time, which encompasses both data preparation and inference (M+D). Compared to RNN-based models [22, 27, 65], GCN-based [28, 29, 53] and transformer-based models [30] generally offer better performance, with reduced prediction time and smaller model sizes, primarily due to their lighter structures and enhanced feature extraction capabilities. Additionally, data preprocessing time significantly impacts inference performance, highlighting the importance of selecting appropriate input data.

Table 2 A comparison of the profiles of various SOTA models is presented. "M" refers to the inference time of the model alone, while "M+D" represents the inference time for both the model and the input data.
Model JAADall
Acc Inference time (ms) Size (MB)
M M+D
PCPA [22] 70.93 38.6 194.6 118.8
Global PCPA [27] 83.49 70.83 416.83 374.2
FUSSI-net [65] 60.96 34.92 190.92 8.4
Pedestrian graph [53] 80.08 29.01 185.01 0.22
Pedestrian graph + [29] 86.97 5.47 351.47 0.27
PedAST-GCN [28] 88.01 10 166 10.4
PIT-base [30] 87.00 4.8 166 21.2
Faster-PCPNet [68] 89.00 0.5 166 0.14

5.3 The model applications

The pedestrian crossing intention model can be applied in several areas, especially in the context of autonomous vehicles and advanced driver assistance systems (ADAS). It can be utilized in autonomous vehicles and ADAS for predictive decision-making and safety enhancements. Additionally, it can be implemented in traffic management systems, such as smart traffic lights and urban planning. Moreover, the model can be applied in insurance and legal contexts, including accident analysis and insurance risk assessment.

6 Model validation and testing

6.1 Dataset

The most used dataset for pedestrian cross intention prediction includes the JAAD and PIE dataset.

The JAAD dataset [14] is specifically designed for autonomous driving, featuring 346 video clips, each with a duration of 5 to 10 seconds. The JAAD behavioral subset (JAADbeh) includes data on 495 pedestrians who are crossing and 191 pedestrians who intend to cross. The complete JAAD dataset (JAADall) adds 2,100 other visible pedestrians who are far from the road and do not intend to cross. The official split for training and testing, as used in Ref. [22], includes 177 videos for training, 29 videos for validation, and 117 for testing.

PIE The PIE dataset [15] is a real-world dataset for predicting pedestrian crossing intentions. It contains 1,322 non-crossing and 512 crossing instances. The PIE dataset includes all pedestrians close to the road, regardless of whether they are hesitant to cross. The official split for training and testing, as suggested in Ref. [22], uses set01, set02, and set04 for training, set05 and set06 for validation, and set03 for testing.

6.2 Evaluation metrics

There are five evaluation metrics (Accuracy, AUC, F1 score, Precision, Recall) to evaluate the performance of models, as proposed in Ref. [22]. The definitions of these evaluation metrics are as follows: A c c u r a c y = T N + T P T N + T P + F N + F P (1) P r e c i s i o n = T P T P + F P (2) R e c a l l = T P T P + F N (3) F 1     s c o r e = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l (4)In this context, TP indicates true positives, TN represents true negatives, FP refers to false positives, and FN stands for false negatives. AUC is the area under the ROC curve.

6.3 Model validation

It is important to compare the proposed methods with existing methods to verify the effectiveness of the proposed model. The main comparative indicators include accuracy, precision, recall, F1-score, and ROC. Achieving better results across more indicators compared to the most advanced methods indicates the stronger effectiveness of the proposed method.

The ablation study is primarily used to assess the impact of each proposed contribution. The main ablation experiments include variations in the model, input data type, input observation length, weather conditions, noisy data, etc.

The model variant experiment explores different aspects of the model, such as using various versions (Base, Tiny, Medium, Large) [30], incorporating specific modules (e.g., with or without attention) [28, 68], and testing different combinations within the model (e.g., late fusion, early fusion, or hierarchical fusion) [27]. These experiments aim to verify the effectiveness of the proposed model in predicting pedestrian crossing intentions.

The type of input data representation significantly influences the prediction of pedestrian crossing intentions [67]. While incorporating more input types can improve accuracy, it also increases model complexity and affects inference speed. Therefore, selecting the appropriate input data representation types is crucial. Additionally, different input combinations can help us understand the contribution of each type to the prediction results.

The observation length significantly influences the performance of the model [28, 38]. A longer observation length can provide richer information, enhancing the model's ability to make accurate predictions. However, a shorter observation length allows for more training samples, potentially improving the model's performance. Therefore, varying the observation length (e.g., 4, 8, 16, 32, 40, etc.) can help verify the model's robustness.

Weather conditions significantly influence the performance of the model [29]. Verifying the model's performance under different weather conditions, such as cloudy, clear, rainy, or foggy, can validate its robustness.

Noisy data is inevitable in real-life scenarios. For example, sensor signal loss can lead to a loss of speed information, and pedestrian detection failures can result in the loss of pedestrian information [28]. Verifying the model's robustness in different noise scenarios is beneficial for ensuring its reliability.

Model profiling involves comparing the model with others in terms of accuracy, computational complexity, number of parameters, size, etc Refs. [29, 37]. This comparison is crucial for verifying the model's performance in practical applications. Also, some examples of comparisons with other models can be visualized to highlight the superiority of the proposed model [28, 68].

7 Challenges and future directions

The review investigates pedestrian crossing intention prediction studies in terms of its model input data (representation) and model structure (feature extraction and prediction). Some challenges and future directions are discussed below.

7.1 Model input data representation

Model input data representation has important influence on the performance of the proposed method. The modality data representation could help the model to obtain more information interacting with pedestrians, environment, and vehicle movement.

As is discussed in Ref. [68], K, CB represents the pose and movement of the pedestrian, S represents the vehicle speed of the vehicle, and P represents the interaction between environment and pedestrians. The performance of the model is limited when only use the skeleton as input. The model has the best performance when use the K, CB, S, and P as input, which include the information of pedestrians, environment, and vehicle movement. The similar results are also in other research [28, 29]. It shows that combine different input data representation could help the model capture more information in term of pedestrians, environment, and vehicle.

However, incorporating additional input data representation can increase the cost of data preprocessing and computing resources, thus reducing the practicality of the model. For instance, many studies introduce segmentation maps as input to extract environmental information, which requires significant time and computing resources to obtain from raw images [27, 29]. Additionally, some input data representation overlap; for example, we can use either pedestrian image crops or skeleton data to extract pose information, but using both simultaneously may be redundant. Furthermore, certain information can be obtained directly or indirectly. For instance, traffic environment information can be acquired through panoramic images or inferred by analyzing the vehicle's speed. Vehicles on roads typically maintain speeds similar to traffic flow to avoid conflicts.

As shown in Ref. [28], "PedAST-GCN(skeleton)" signifies the prediction results generated by the PedAST-GCN model using only the skeleton as input. "Global PCPA" denotes prediction results from the Global PCPA, which employs RNN as a backbone and utilizes inputs such as the skeleton, vehicle speed, bounding box, pedestrian image crop, and semantic segmentation image. "PedAST-GCN" represents prediction results from the PedAST-GCN model, which utilizes the skeleton, vehicle speed, and bounding box as input. Compared to PedAST-GCN (skeleton), PedAST-GCN and Global PCPA exhibit better performance in crowded scenes. This is because PedAST-GCN (skeleton) solely uses skeleton data as input, which performs poorly when the pedestrian's pose is occluded. However, despite containing additional surrounding environment and segmentation maps as input, Global PCPA's performance is inferior to PedAST-GCN. PedAST-GCN demonstrates superior ability in feature extraction and fusion compared to Global PCPA. Additionally, PedAST-GCN achieves faster inference speeds than Global PCPA. Therefore, utilizing suitable input data representation and combining useful backbones are crucial for model performance.

7.2 Model generalization ability

The generalization ability of a model is crucial as it ensures good performance in real-world scenarios. As is discussed in Ref. [80], the model exhibits poorer performance when trained and tested on the different dataset compared to when tested and trained on same datasets in most scenarios. This highlights the importance of focusing on the model's generalization ability. Improvements in both model structure and normalization of input data are necessary to enhance performance.

7.3 The uncertainty for the pedestrian crossing intention prediction

Uncertainty is important for model performance. For example, from the perspective of human recognition of pedestrian crossing intention, different drivers have varying judgments about a pedestrian's true intention, especially in ambiguous situations. However, most existing models do not incorporate uncertainty like humans do. As shown in Ref. [40]), the proposed model "TrEP" introduced AI uncertainty for pedestrian crossing intention prediction. It is observed that the model's performance significantly improves when it accounts for uncertainty.

From the perspective of model prediction, model structure uncertainty and model input uncertainty are two different sources of uncertainty [81]. Once training is completed, the model has fixed parameters, which limits its ability to simulate the decision-making process of different drivers and thus restricts its performance. Different initial parameters can result in models with varying parameters, introducing uncertainty into the model. Additionally, input data uncertainty, such as missing or incorrect data, can affect the model's performance. Therefore, considering uncertainty in both the model and input data is crucial for improving overall performance.

7.4 The data privacy and ethical considerations

Data privacy and ethical considerations are paramount in the realm of AI and autonomous vehicles [82, 83], particularly in the context of pedestrian crossing intention models. Regarding the collection of pedestrian information, safeguarding data privacy is essential. For example, when employing car cameras to capture footage, it is prudent to ensure that the process minimizes the inadvertent capture of personal information belonging to pedestrians, respecting their privacy rights. Moreover, the predictive model tasked with analyzing this data must uphold the principle of unbiased data processing. For example, this entails ensuring that the model's predictions are not tainted by discriminatory biases based on attributes such as age, gender, race, or any other protected characteristic. The elimination of such biases is critical to fostering a fair and equitable transportation environment. Furthermore, the decision-making processes inherent in these predictive models must embody transparency and interpretability. Lastly, both the developers and users of these predictive models carry a weighty responsibility. In this way, AI and autonomous vehicles can contribute positively to society while upholding the highest standards of privacy, fairness, and ethics.

This commitment extends beyond the current work and necessitates a continued focus on data privacy and ethical considerations in future endeavors. By doing so, we can harness the power of AI and autonomous vehicles to enhance society while preserving the values that are essential to our collective well-being.

8 Conclusions

Pedestrian crossing intention is critical for vehicle intelligence. This paper presents a comprehensive review of pedestrian crossing intention prediction studies, investigating data perception, model input data representation, model structures (feature extraction and prediction), and challenges. Key factors for the prediction performance and challenges are discussed. We conclude the challenges are for model input data inputs representation, model generalization ability, and prediction uncertainty. The survey can contribute to promote future research to improve the usability and trust ability of the pedestrian crossing intention prediction in the wild.

References

[1] 

J. Fang, F. Wang, J. Xue, and T. -S. Chua, "Behavioral intention prediction in driving scenes: A survey," IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 8, pp. 8334–8355, 2024.

[2] 

N. Sharma, C. Dhiman, and S. Indu, "Pedestrian intention prediction for autonomous vehicles: A comprehensive survey," Neurocomputing, vol. 508, pp. 120–152, 2022.

[3] 

H. Razali, T. Mordan, and A. Alahi, "Pedestrian intention prediction: A convolutional bottom-up multi-task approach," Transportation Research Part C: Emerging Technologies, vol. 130, p. 103259, 2021.

[4] 

Y. Ling, Z. Ma, B. Xie, Q. Zhang, and X. Weng, "SA-BiGCN: Bi-stream graph convolution networks with spatial attentions for the eye contact detection in the wild," IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 2, pp. 2089–2100, 2023.

[5] 

Y. Belkada, L. Bertoni, R. Caristan, T. Mordan, and A. Alahi, "Do pedestrians pay attention? eye contact detection in the wild," arXiv preprint arXiv:2112.04212, 2021.

[6] 

V. Onkhar, P. Bazilinskyy, J. C. Stapel, D. Dodou, D. Gavrila, and J. C. de Winter, "Towards the detection of driver–pedestrian eye contact," Pervasive and Mobile Computing, vol. 76, p. 101455, 2021.

[7] 

R. Korbmacher and A. Tordeux, "Review of pedestrian trajectory prediction methods: Comparing deep learning and knowledge-based approaches," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 12, pp. 24126–24144, 2022.

[8] 

X. Song, K. Chen, X. Li, J. Sun, B. Hou, Y. Cui, B. Zhang, G. Xiong, and Z. Wang, "Pedestrian trajectory prediction based on deep convolutional LSTM network," IEEE Transactions on Intelligent Transportation Systems, vol. 22, no. 6, pp. 3285–3302, 2020.

[9] 

M. Golchoubian, M. Ghafurian, K. Dautenhahn, and N. L. Azad, "Pedestrian trajectory prediction in pedestrian-vehicle mixed environments: A systematic review," IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 11, pp. 11544–11567, 2023.

[10] 

S. Zhang, M. Abdel-Aty, Y. Wu, and O. Zheng, "Pedestrian crossing intention prediction at red-light using pose estimation," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 3, pp. 2331–2339, 2021.

[11] 

T. Chen, R. Tian, and Z. Ding, "Visual reasoning using graph convolutional networks for predicting pedestrian crossing intention,"2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada, 11–17 October, 2021, pp. 3103–3109.

[12] 

H. Kim, K. Lee, G. Hwang, and C. Suh, "Crash to not crash: Learn to identify dangerous vehicles using a simulator,"in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 978–985.

[13] 

J. F. P. Kooij, N. Schneider, F. Flohr, and D. M. Gavrila, "Contextbased pedestrian path prediction,"Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6–12, 2014.

[14] 

A. Rasouli, I. Kotseruba, and J. K. Tsotsos, "Are they going to cross? A benchmark dataset and baseline for pedestrian crosswalk behavior,"in Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy, 22–29 October, 2017, pp. 206–213.

[15] 

A. Rasouli, I. Kotseruba, T. Kunic, and J. K. Tsotsos, "Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction,"in Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), 27 October–02 November, 2019, pp. 6262–6271.

[16] 

B. Liu, E. Adeli, Z. Cao, K.-H. Lee, A. Shenoi, A. Gaidon, and J. C. Niebles, "Spatiotemporal relationship reasoning for pedestrian intent prediction," IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3485–3492, 2020.

[17] 

H. Wu, L. Wang, S. Zheng, Q. Xu, and J. Wang, "Crossing-road pedestrian trajectory prediction based on intention and behavior identification,"in 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), Rhodes, Greece, 20–23 September, 2020, pp. 1–6.

[18] 

T. Chen, T. Jing, R. Tian, Y. Chen, J. Domeyer, H. Toyoda, R. Sherony, and Z. Ding, "PSI: A pedestrian behavior dataset for socially intelligent autonomous car," arXiv preprint arXiv:2112.02604, 2021.

[19] 

A. Rasouli, T. Yau, P. Lakner, S. Malekmohammadi, M. Rohani, and J. Luo, "PePScenes: A novel dataset and baseline for pedestrian action prediction in 3d," arXiv preprint arXiv:2012.07773, 2020.

[20] 

B. Völz, K. Behrendt, H. Mielenz, I. Gilitschenski, R. Siegwart, and J. Nieto, "A data-driven approach for pedestrian intention estimation,"2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC). Rio de Janeiro, Brazil, 1–4 November, 2016, pp. 2607–2612.

[21] 

C. Zhang, A. H. Kalantari, Y. Yang, Z. Ni, G. Markkula, N. Merat, and C. Berger, "Cross or wait? predicting pedestrian interaction outcomes at unsignalized crossings,"in 2023 IEEE Intelligent Vehicles Symposium (IV). Rio de Janeiro, Brazil, 01–04 November, 2016, 2023, pp. 1–8.

[22] 

I. Kotseruba, A. Rasouli, and J. K. Tsotsos, "Benchmark for evaluating pedestrian action prediction,"in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 03–08 January, 2021, pp. 1258–1268.

[23] 

M. Azarmi, M. Rezaei, H. Wang, and S. Glaser, "PIP-Net: Pedestrian intention prediction in the wild," arXiv preprint arXiv:2402.12810, 2024.

[24] 

J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, Las Vegas, NV, USA, June 27–June 30 2016, pp. 779–788.

[25] 

W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C. -Y. Fu, and A. C. Berg, "SSD: Single shot multibox detector," Computer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016.

[26] 

R. Girshick, "Fast R-CNN,"in Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 07–13 December, 2015, pp. 1440–1448.

[27] 

D. Yang, H. Zhang, E. Yurtsever, K. A. Redmill, and Ü. Özgüner, "Predicting pedestrian crossing intention with feature fusion and spatiotemporal attention," IEEE Transactions on Intelligent Vehicles, vol. 7, no. 2, pp. 221–230, 2022.

[28] 

Y. Ling, Z. Ma, Q. Zhang, B. Xie, and X. Weng, "PedAST-GCN: Fast pedestrian crossing intention prediction using spatial–temporal attention graph convolution networks," IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 10, pp. 13277–13290, 2024.

[29] 

P. R. G. Cadena, Y. Qian, C. Wang, and M. Yang, "Pedestrian graph+: A fast pedestrian crossing prediction model based on graph convolutional networks," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 11, pp. 21050–21061, 2022.

[30] 

Y. Zhou, G. Tan, R. Zhong, Y. Li, and C. Gou, "PIT: Progressive interaction transformer for pedestrian crossing intention prediction," IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 12, pp. 14213–14225, 2023.

[31] 

J. Carreira and A. Zisserman, "Quo vadis, action recognition? A new model and the kinetics dataset,"in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, Honolulu, HI, USA, 21–26 July 2017, pp. 6299–6308.

[32] 

Z. Cao, T. Simon, S. -E. Wei, and Y. Sheikh, "Realtime multi-person 2d pose estimation using part affinity fields,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, Honolulu, HI, USA, 21–26 July, 2017, pp. 7291–7299.

[33] 

G. Hidalgo, Y. Raaj, H. Idrees, D. Xiang, H. Joo, T. Simon, and Y. Sheikh, "Single-network whole-body pose estimation,"in Proceedings of the IEEE/CVF International Conference on Computer Vision, 27 Oct.–2 Nov., 2019, Seoul, Korea (South), 2019, pp. 6982–6991.

[34] 

K. Sun, B. Xiao, D. Liu, and J. Wang, "Deep high-resolution representation learning for human pose estimation,"in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, Seoul, Korea (South), 27 Oct.–2 Nov., 2019, pp. 5693–5703.

[35] 

S. Kreiss, L. Bertoni, and A. Alahi, "PifPaf: Composite fields for human pose estimation,"in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, Seattle, WA, USA, 13–19 June, 2020, pp. 11977–11986.

[36] 

A. Bucksch, R. Lindenbergh, and M. Menenti, "SkelTre: Robust skeleton extraction from imperfect point clouds," The Visual Computer, vol. 26, pp. 1283–1300, 2010.

[37] 

Y. Ling, Q. Zhang, X. Weng, and Z. Ma, "STMA-GCN pedcross: Skeleton based spatial-temporal graph convolution networks with multiple attentions for fast pedestrian crossing intention prediction,"in 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC), Bilbao, Spain, 13 February, 2024, pp. 500–506.

[38] 

X. Zhang, P. Angeloudis, and Y. Demiris, "ST crossingpose: A spatial-temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 11, pp. 20773–20782, 2022.

[39] 

T. Mordan, M. Cord, P. Pérez, and A. Alahi, "Detecting 32 pedestrian attributes for autonomous vehicles," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 8, pp. 11823–11835, 2021.

[40] 

Z. Zhang, R. Tian, and Z. Ding, "TrEP: Transformer-based evidential prediction for pedestrian intention with uncertainty,"in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, pp. 3534–3542, 2023.

[41] 

X. Guo, S. Li, J. Yu, J. Zhang, J. Ma, L. Ma, W. Liu, and H. Ling, "PFLD: A practical facial landmark detector," arXiv preprint arXiv:1902.10859, 2019.

[42] 

W. Wu, C. Qian, S. Yang, Q. Wang, Y. Cai, and Q. Zhou, "Look at boundary: A boundary-aware face alignment algorithm,"in Proceedings of 2018 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, June 18–23, 2018, pp. 2129–2138.

[43] 

H. Jin, S. Liao, and L. Shao, "Pixel-in-pixel net: Towards efficient facial landmark detection in the wild," International Journal of Computer Vision, vol. 129, no. 12, pp. 3174–3194, 2021.

[44] 

X. Dong, Y. Yan, W. Ouyang, and Y. Yang, "Style aggregated network for facial landmark detection,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, Salt Lake City, UT, USA, June 18–23, 2018, pp. 379–388.

[45] 

E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox, "Flownet 2.0: Evolution of optical flow estimation with deep networks,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, Honolulu, HI, USA, 21–26 July 2017, pp. 2462–2470.

[46] 

L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, "Rethinking atrous convolution for semantic image segmentation," arXiv preprint arXiv:1706.05587, 2017.

[47] 

O. Ronneberger, P. Fischer, and T. Brox, "U-Net: Convolutional networks for biomedical image segmentation,"in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5–9, 2015.

[48] 

T.-Y. Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, "Feature pyramid networks for object detection,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2117–2125.

[49] 

W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, "Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,"in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 568–578.

[50] 

D. Eigen, C. Puhrsch, and R. Fergus, "Depth map prediction from a single image using a multi-scale deep network," Advances in Neural Information Processing Systems, vol. 3, no. 9, pp. 2366–2374, 2014.

[51] 

J. Watson, O. Mac Aodha, V. Prisacariu, G. Brostow, and M. Fir-man, "The temporal opportunist: Self-supervised multi-frame monocular depth,"in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June, 2021, pp. 1164–1174.

[52] 

K. Simonyan and A. Zisserman, "Two-stream convolutional networks for action recognition in videos," Advances in Neural Information Processing Systems, vol. 27, pp. 568–576, 2014.

[53] 

P. R. G. Cadena, M. Yang, Y. Qian, and C. Wang, "Pedestrian graph: Pedestrian crossing prediction based on 2d pose estimation and graph convolutional networks,"in 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, New Zealand, 27–30 October, 2019, pp. 2000–2005.

[54] 

F. U. H. A. Mattoo, U. S. Khan, T. Nawaz, and N. Rashid, "Deep learning-based feature fusion for action recognition using skeleton information,"in 2023 International Conference on Robotics and Automation in Industry (ICRAI), Peshawar, Pakistan, 03–05 March, 2023, pp. 1–6.

[55] 

Y. Ito, Q. Kong, K. Morita, and T. Yoshinaga, "Efficient and accurate skeleton-based two-person interaction recognition using inter-and intrabody graphs,"in 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 16–19 October, 2022, pp. 231–235.

[56] 

I. Kotseruba, A. Rasouli, and J. K. Tsotsos, "Do they want to cross? understanding pedestrian intention for behavior prediction,"in 2020 IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA, 19 October–November, 2020, pp. 1688–1693.

[57] 

M. Dong, "Pedestrian cross forecasting with hybrid feature fusion,"in Asian Conference on Machine Learning. ACML, 2024, Hanoi, Vietnam, December 5–8, 2024, pp. 327–342.

[58] 

T. W. Wang and S.-H. Lai, "Pedestrian crossing intention prediction with multi-modal transformer-based model,"in 2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Taipei, Taiwan, 20 November, 2023, pp. 1349–1356.

[59] 

X. Shi, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-C. Woo, "Convolutional lstm network: A machine learning approach for precipitation nowcasting," NIPS'15: Proceedings of the 28th International Conference on Neural Information Processing Systems, vol. 1, pp. 802–810.

[60] 

J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, "Beyond short snippets: Deep networks for video classification,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, arXiv preprint arXiv:1503.08909, 2015.

[61] 

Y. Du, W. Wang, and L. Wang, "Hierarchical recurrent neural network for skeleton based action recognition,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 07–12 June, 2015, pp. 1110– 1118.

[62] 

D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, "Learning spatiotemporal features with 3d convolutional networks,"in Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 18 February, 2016, pp. 2380–7504.

[63] 

A. Bhattacharyya, M. Fritz, and B. Schiele, "Long-term on-board prediction of people in traffic scenes under uncertainty,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June, 2018, pp. 4194–4202.

[64] 

A. Rasouli, I. Kotseruba, and J. K. Tsotsos, "Pedestrian action anticipation using contextual feature fusion in stacked RNNs," arXiv preprint arXiv:2005.06582, 2020.

[65] 

F. Piccoli, R. Balakrishnan, M. J. Perez, M. Sachdeo, C. Nunez, M. Tang, K. Andreasson, K. Bjurek, R. D. Raj, E. Davidsson, et al., "Fussi-Net: Fusion of spatio-temporal skeletons for intention prediction network,"in 2020 54th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 01–04, November, 2020, pp. 68–72.

[66] 

J. Gesnouin, S. Pechberti, B. Stanciulcscu, and F. Moutarde, "Trouspinet: Spatio-temporal attention on parallel atrous convolutions and ugrus for skeletal pedestrian crossing prediction,"in 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), Jodhpur, India, 15–18 December, 2021.

[67] 

B. Yang, Z. Wei, H. Hu, R. Wang, C. Yang, and R. Ni, "Dpcian: A novel dual-channel pedestrian crossing intention anticipation network," IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 6, pp. 6023–6034, 2023.

[68] 

B. Yang, J. Zhu, C. Hu, Z. Yu, H. Hu, and R. Ni, "Faster pedestrian crossing intention prediction based on efficient fusion of diverse intention influencing factors," IEEE Transactions on Transportation Electrification, Early Access, 2024, DOI: 10.1109/TTE.2024.3360966.

[69] 

X. Chen, S. Zhang, J. Li, and J. Yang, "Pedestrian crossing intention prediction based on cross-modal transformer and uncertainty-aware multi-task learning for autonomous driving," IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 9, 12538–12549, 2024.

[70] 

N. Lv, Y. Huang, H. Zhang, and F. Wu, "Pedestrian crossing prediction with pathwise feature fusion and stacked gate recurrent unit," IEEE Sensors Letters, vol. 8, no. 2, pp. 2475–1472, 2024.

[71] 

K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," arXiv preprint arXiv:1409.1556, 2014.

[72] 

K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition,"in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 12 December, 2016, pp. 770–778.

[73] 

A. Graves and A. Graves, "Long short-term memory," Supervised Sequence Labelling with Recurrent Neural Networks, Springer, 2012, pp. 37–45.

[74] 

J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, "Empirical evaluation of gated recurrent neural networks on sequence modeling," arXiv preprint arXiv:1412.3555, 2014.

[75] 

T. N. Kipf and M. Welling, "Semi-supervised classification with graph convolutional networks," arXiv preprint arXiv:1609.02907, 2016.

[76] 

X. Wei, S. Yao, C. Zhao, D. Hu, H. Luo, and Y. Lu, "Graph convolutional networks (GCN)-based lightweight detection model for dangerous driving behavior,"in International Conference on Wireless Algorithms, Systems, and Applications. Springer, 2022, pp. 27–39.

[77] 

T. Sofianos, A. Sampieri, L. Franco, and F. Galasso, "Space-time separable graph convolutional network for pose forecasting,"in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, Montreal, QC, Canada, Oct. 10–Oct. 17 2021, pp. 11209–11218.

[78] 

A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," arXiv preprint arXiv:1706.03762, 2017.

[79] 

S. Dai, J. Liu and N.-M. Cheung, "Uncertainty-aware pedestrian crossing prediction via reinforcement learning," IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 10, pp. 9540–9549, 2024.

[80] 

J. Gesnouin, S. Pechberti, B. Stanciulescu, and F. Moutarde, "Assessing cross-dataset generalization of pedestrian crossing predictors,"in 2022 IEEE Intelligent Vehicles Symposium (IV), Aachen, Germany, 04–09 June, 2022, pp. 419–426.

[81] 

J. Gawlikowski, C. R. N. Tassi, M. Ali, J. Lee, M. Humt, J. Feng, A. Kruspe, R. Triebel, P. Jung, R. Roscher, et al., "A survey of uncertainty in deep neural networks," Artificial Intelligence Review, vol. 56, no. Suppl. 1, pp. 1513–1589, 2023.

[82] 

K. Terzidou, I. Krontiris, K. Grammenou, M. Zacharopoulou, M. Tsikintikou, F. Baladima, C. Sakellari, and K. Kaouras, "Autonomous vehicles: Data protection and ethical considerations," Journal of the Association for Computing Machinery, 2020.

[83] 

I. Krontiris, K. Grammenou, K. Terzidou, M. Zacharopoulou, M. Tsikintikou, F. Baladima, C. Sakellari, and K. Kaouras, "Autonomous vehicles: Data protection and ethical considerations,"in Proceedings of the 4th ACM Computer Science in Cars Symposium, Feldkirchen Germany, 2 December, 2020, pp. 1–10.

Top