Yanjun Qin

dblp:210/3849 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs
abstract
While Vision-Language Models (VLMs) have garnered increasing attention in the AI community due to their promising practical applications, they exhibit persistent hallucination issues, generating outputs misaligned with visual inputs. Recent studies attribute these hallucinations to VLMs' over-reliance on linguistic priors and insufficient visual feature integration, proposing heuristic decoding calibration strategies to mitigate them. However, the non-trainable nature of these strategies inherently limits their optimization potential. To this end, we propose an adversarial parametric editing framework for Hallucination mitigation in VLMs, which follows an Activate-Locate-Edit Adversarially paradigm. Specifically, we first construct an activation dataset that comprises grounded responses (positive samples attentively anchored in visual features) and hallucinatory responses (negative samples reflecting LLM prior bias and internal knowledge artifacts). Next, we identify critical hallucination-prone parameter clusters by analyzing differential hidden states of response pairs. Then, these clusters are fine-tuned using prompts injected with adversarial prefixes optimized via prompt tuning to maximize visual neglect, thereby forcing the model to prioritize visual evidence over inherent parametric biases. Evaluations on both generative and discriminative VLM tasks demonstrate the significant effectiveness of ALEAHallu in alleviating hallucinations.
Beibei Li 0001, Jiangwei Xia, Yanjun Qin, Zhongshi He
AAAI4
2026 Adaptive Frequency Pathways for Spatiotemporal Forecasting
abstract
Spatiotemporal forecasting is a fundamental task in areas such as traffic flow prediction, environmental sensing, and urban planning. Recent advances have shown that decomposing temporal signals into multiple frequencies and modeling them jointly with spatial structures can significantly enhance forecasting performance. However, existing multifrequency forecasting models still face two critical limitations. First, the importance of different temporal frequencies evolves over time, yet most models assume fixed or static frequency contributions. Second, spatial dependencies are inherently frequency-sensitive. For instance, low-frequency components often align with global spatial patterns, while highfrequency components tend to correspond to localized interactions. However, current approaches typically use a shared spatial information across all frequencies, introducing spatiotemporal inconsistency. To address these challenges, we propose a novel Adaptive Frequency Pathways (AdaFre) for spatiotemporal forecasting, which adaptively captures both dynamic frequency relevance and frequency-aligned spatial structures. AdaFre employs a multi-frequency routing mechanism to dynamically select and aggregate the most informative temporal frequency components, while associating each with its corresponding spatial representation derived from frequency-aware embeddings. Spatiotemporal backbones are then used to model each path independently before final aggregation. Extensive experiments on several real-world datasets demonstrate that AdaFre significantly outperforms state-of-the-art baselines.
Yanjun Qin, Yuchen Fang 0001, Xinke Jiang, Hao Miao 0001, Xiaoming Tao 0001
AAAI1
2026 STEP: Stable Gradient Projection for Continual Learning
abstract
Continual learning (CL) aims to enable networks to learn continuously from sequentially arriving task streams while avoiding catastrophic forgetting (CF) of previously learned tasks. In recent years, Orthogonal gradient projection (OGP)-based CL methods have garnered significant attention from the research community due to their remarkable performance. However, existing OGP approaches overlook two critical issues: (1) representation matrices are typically constructed via random sampling, which introduces misclassified and class-imbalanced samples into the projection basis, contaminating important gradient directions and degrading stability; and (2) task-specific output scale variations induce domain drift, resulting in projection bias that weakens orthogonal constraints across tasks. To address these limitations, we propose Stable Gradient Projection for Continual Learning (STEP), a plug-and-play enhancement framework for OGP-based CL that integrates Correctness-aware Balanced Sampling (CBS) to construct purified and class-balanced projection subspaces using only correctly classified samples, and Sigmoid Attention Constraint (SAC) to enforce consistent output scaling via a sigmoid-based gating mechanism, thereby mitigating scale-induced projection bias. Extensive experiments on Split CIFAR-100, CIFAR-100 Superclass, and 5-Datasets demonstrate that STEP consistently improves state-of-the-art OGP methods, achieving up to +1.4% average accuracy (ACC) gains on Split CIFAR-100, improving backward transfer (BWT) from − 0.37 to − 0.09 for GPM and from − 1.06 to − 0.73 for SGP, and attaining 93.28% ACC with positive BWT (0.17) on 5-Datasets. These results validate STEP as a simple yet effective strategy for enhancing stability–plasticity balance in OGP-based CL.
Longlong Zhai, Jiao Tian, Yanjun Qin, Shaochen Jiang, Chong Peng 0001, Panpan Zheng
ICMR5
2026 Assisted refinement network based on channel information interaction for camouflaged object detection
Kuan Wang 0007, Yanjun Qin, Mengge Lu, Xiaoming Tao 0001
Expert Syst. Appl.2
2025 HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
abstract
Video Anomaly Detection (VAD) aims to identify and locate deviations from normal patterns in video sequences. Traditional methods often struggle with substantial computational demands and a reliance on extensive labeled datasets, thereby restricting their practical applicability. To address these constraints, we propose HiProbe-VAD, a novel framework that leverages pre-trained Multimodal Large Language Models (MLLMs) for VAD without requiring fine-tuning. In this paper, we discover that the intermediate hidden states of MLLMs contain information-rich representations, exhibiting higher sensitivity and linear separability for anomalies compared to the output layer. To capitalize on this, we propose a Dynamic Layer Saliency Probing (DLSP) mechanism that intelligently identifies and extracts the most informative hidden states from the optimal intermediate layer during the MLLMs reasoning. Then a lightweight anomaly scorer and temporal localization module efficiently detects anomalies using these extracted hidden states and finally generate explanations. Experiments on the UCF-Crime and XD-Violence datasets demonstrate that HiProbe-VAD outperforms existing training-free and most traditional approaches. Furthermore, our framework exhibits remarkable cross-model generalization capabilities in different MLLMs without any tuning, unlocking the potential of pre-trained MLLMs for video anomaly detection and paving the way for more practical and scalable solutions.
Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Yanjun Qin
ACM Multimedia4
2025 CLIP-LMFA: Few-Shot Anomaly Detection via Large Language Model-Driven Hybrid Prompts and Multi-scale Adaptive Fusion
Shengchang Wang, Yanjun Qin, Yongke Li, Zhaoru Guo, Haoxiang Huang, Tiquan Gu, Panpan Zheng
PRICAI (5)3
2025 A diffusion-based feature enhancement approach for driving behavior classification with EEG data
Yanjun Qin, Shanghang Zhang, Xiaoming Tao 0001
Adv. Eng. Informatics2
2025 EEG-Driven Classification of Driver Mental Workload in Diverse Environments: A Dual-Branch Network for Efficient In-Vehicle Applications
abstract
The mental load of drivers can profoundly affect their driving performance, to the extent that it affects traffic safety. Therefore, monitoring mental workload has become a crucial aspect of sensor-based driver monitoring systems, especially in the context of the Industrial Internet of Things (IIoT), where driver status information can be exchanged between vehicles to enhance safety. However, the substantial energy consumption and transmission latency associated with traditional central-server-based IoT systems are prominent issues that necessitate the development of lighter algorithms for edge computing in individual vehicles. In this article, we focus on the impact of external traffic events and environmental changes on the mental load of drivers, as well as effective classification algorithms applied in monitoring systems. To analyze the physiological responses of drivers to road events and non driving related tasks under different weather conditions, we proposed a dual branch model, DMW-Net, based on attention mechanism branches and graph attention modules to discriminate the mental load level of drivers from physiological signals. The proposed method was validated on the manD dataset and achieved an accuracy of 90.07% in physiological signals of three different load levels, which is higher than the comparison models. This study provides innovative methods for driver monitoring systems, contributing to advanced driving assistance systems (ADAS) and traffic safety.
Yanjun Qin, Shanghang Zhang, Yiping Duan, Xiaoming Tao 0001
IEEE Internet Things J.2
2025 Empowering Corner Case Detection in Autonomous Vehicles With Multimodal Large Language Models
abstract
Object detection powered by deep learning is an essential component in the realm of self-driving vehicles. However, the model may be affected by corner cases, which are rare or unusual objects and scenarios, and can significantly impact the reliability of object detection systems. In this paper, we applied a Multimodal Large Language Model (MLLM) to address the challenge of corner cases in autonomous driving systems. The MLLM consists of an image encoder, a text tokenizer, a modal alignment layer, and a pre-trained large language model, enabling the model to understand multimodal semantic information. We added text descriptions on the basis of corner case dataset CODA and constructed the CODA-REC dataset. This dataset is then used to perform instruction fine-tuning on the MLLM to adapt it to the object detection task. The proposed method leverages the extensive knowledge and zero-shot learning capabilities of LLMs to enhance the semantic understanding of text and images, enabling the detection and appropriate response to corner cases that were previously difficult to handle. The experimental results show that MLLM achieved better performance than baseline models, with an improvement of about 10% in mAR and mAP metrics compared to most closed-set models, and an improvement of 10% mAP compared to open set models. We hope that our work can inspire the application of MLLMs in the field of autonomous driving, contributing to more advanced intelligent transportation systems.
Yanjun Qin, Shanghang Zhang, Xiaoming Tao 0001
IEEE Signal Process. Lett.2
2025 Diversifying Latent Flows for Safety-Critical Scenarios Generation With CARLA Simulator
abstract
The likelihood of encountering scenarios that lead to accidents, namely safety-critical scenarios, is minimal compared to long-term safe driving environments. The generation of repeatable and scalable safety-critical scenarios is essential for the advancement of human and autonomous driving capabilities. Compared with the high complexity and low practicality of existing scenario generation methods, in this paper we propose a real-time approach to automatically generate challenging scenarios and instantiate them in a CARLA-based simulator. First, the safety-critical scenario is decomposed into a perturbed and optimized vehicle trajectory and the remaining reusable Unreal Engine assets based on a hierarchical model. Second, a model that is based on a graph conditional variational autoencoder (VAE) is employed to predict future trajectories and head angles based on past information. Third, the safety-critical scene generation model is used to enhance the diversity of the scene by diversifying the latent variables over a pre-trained trajectory representation model. Finally, the trajectories of real-world vehicles are placed into the simulator by adapting them to enable the generation of safety-critical scenes in a three-dimensional environment. The results demonstrate that the proposed approach generates scenarios that are more plausible than those generated by the baselines, with a performance improvement of over 10% in collision metrics for scenario generation. The research facilitates the simplification of the long-tail scenario construction process for autonomous vehicles, which in turn facilitates the optimization of algorithms such as autonomous trajectory planning.
Dingcheng Gao, Yanjun Qin, Xiaoming Tao 0001, Jianhua Lu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Cross-Scenario Vigilance Detection Based on EEG Analysis for Safety Driving in Autonomous
abstract
Safety driver vigilance is a prerequisite for the safe operation of autonomous vehicles. In contrast to vehicle behavioral trajectory detection, which suffers from high latency and low accuracy, vigilance detection based on physiological signals is currently the most reliable and accurate method. While vigilance monitoring methods using electroencephalograms (EEG) have made considerable progress in experimental scenarios, they remain a challenging problem in scenario-constrained conditions, such as high-speed moving autonomous vehicles. This is due to the low signal-to-noise ratio in EEG signal acquisition and the difficulty of real-time processing. Moreover, cumbersome data acquisition processes and the challenges of labeling have hindered progress in this area. Given the successful use of EEG for monitoring in experimental settings, we believe that the transfer of knowledge learned from these scenarios to new contexts is reasonably feasible. Thus, this work aims to bridge the domain gap between experimental and real-world scenarios while balancing the number of channels and accuracy. Specifically, we propose a framework for EEG vigilance detection capable ofCross-scenario,Cross-subject, andCross-device, calledCCC. The proposed framework leverages the standard montage structure of EEG channels, reducing the number of channels by considering the common regions of EEG channels across different scenarios. The results show that our proposed model achieves an average accuracy of 86.20% on the SEED-VIG dataset with 12 subjects, which is higher than the 82.21% achieved by state-of-the-art deep learning approaches. Finally, we investigate the role of the attention mechanism and transfer learning, and further attempt to explain the advantages of our proposed approach from a visualization perspective.
Dingcheng Gao, Xiaoming Tao 0001, Xia Wu 0001, Yanjun Qin, Jianhua Lu
IEEE Trans. Intell. Transp. Syst.5
2024 DMGSTCN: Dynamic Multigraph Spatio-Temporal Convolution Network for Traffic Forecasting
abstract
Traffic forecasting belongs to intelligent transportation systems and is helpful for public property and life safety. Therefore, to forecast traffic accurately, researchers pay great attention to dealing with complex problems by mining intricate spatial and temporal dependencies of the traffic. However, some challenges still hold back traffic forecasting: 1) Most studies mainly focus on modeling correlations of traffic time series of close distances on the road network and ignore correlations of remote but similar traffic time series; 2) Previous static graph-based methods failed to reflect the dynamic changed spatial relations of multiple time series in the evolving traffic system. To tackle the above issues, we design a new dynamic multi-graph spatio-temporal convolution network (DMGSTCN) in this paper, which utilizes the gated causal convolution with the dynamic multi-graph convolution network (DMGCN) to simultaneously extract spatial and temporal information. Specifically, DMGCN uses not only distance-based graphs but also structure-based graphs to obtain spatial information from nearby and remote but similar traffic time series, respectively. Moreover, to dynamically model spatial correlations, DMGCN first splits neighbors of each traffic time series into different regions according to relative position relationships. Then DMGCN assigns different weights to different regions at different time slices. Empirical evaluations on four traffic forecasting benchmarks reveal that DMGSTCN outperforms existing methods.
Yanjun Qin, Xiaoming Tao 0001, Yuchen Fang 0001, Haiyong Luo, Fang Zhao 0003, Chenxing Wang 0001
IEEE Internet Things J.1
2024 Decoding Brain-Controlled Intention for UAVs and IVs Based on Lightweight Network
abstract
Brain-Computer-Interface (BCI) plays an important role in the Internet of Things (IoT). With the development of electroencephalogram (EEG) signal processing and deep learning, researchers are beginning to decode intention for controlling smart movable devices by EEG signals, such as unmanned aerial vehicles (UAVs) and intelligent vehicles (IVs). However, the current related studies less consider the generic decoding of the two, and the performance of the adopted models needs to be improved in terms of both lightness and accuracy. In this paper, we are dedicated to the study of a generic brain-controlled intention decoding task serving UAVs and IVs. We adopt three different datasets, encompassing different BCI paradigms. A lightweight self-attention enhancement model is proposed, which incorporates self-attention into the model for feature enhancement. Experimental results show that our method outperforms the baselines on brain-controlled intention decoding for all three datasets. This research is useful for the work in the field of brain-controlled intention decoding, and also provides new ideas for lightweight network-based control of UAVs and IVs.
Wenqi Zhang 0003, Yanjun Qin, Xiaoming Tao 0001
IEEE Internet Things J.2
2024 MMPHGCN: A Hypergraph Convolutional Network for Detection of Driver Intention on Multimodal Physiological Signals
abstract
Detecting driving behavior is crucial as it is closely related to driving safety. With the development of electroencephalogram (EEG) signal processing, researchers have started studying driving intentions through EEG signals. However, current research mainly focuses on a single intention, and EEG signals suffer from low spatial resolution, leading to less accurate detection. This paper aims to investigate a set of driving intentions, targeting a classification task based on multimodal physiological signals. We propose a model that incorporates hypergraph convolution for feature extraction. Experimental results demonstrate that our approach outperforms the baselines in detecting various types of driving intentions with 74.40% accuracy, 5.89% higher than the best baseline. This research contributes to both traffic safety and brain-computer interfaces.
Wenqi Zhang 0003, Yanjun Qin, Xiaoming Tao 0001
IEEE Signal Process. Lett.2
2024 STWave$^+$+: A Multi-Scale Efficient Spectral Graph Attention Network With Long-Term Trends for Disentangled Traffic Flow Forecasting
abstract
Traffic forecasting is crucial for public safety and resource optimization, yet is very challenging due to the temporal changes and the dynamic spatial correlations. To capture these intricate dependencies, spatio-temporal networks, such as recurrent neural networks with graph convolution networks, are applied. However, traffic forecasting is still a non-trivial task because of three major challenges: 1) Previous spatio-temporal networks are based on end-to-end training and thus fail to handle the distribution shift in the non-stationary traffic time series. 2) Existing methods always utilize the one-hour input to forecast future traffic and the long-term historical trend knowledge is ignored. 3) The efficient and effective algorithm for modeling multi-scale spatial correlations is still lacking in prior networks. Therefore, in this paper, rather than proposing yet another end-to-end model, we provide a novel disentangle-fusion framework STWave+to mitigate the distribution shift issue. The framework first decouples the complex one-hour traffic data into stable trends and fluctuating events, followed by a dual-channel spatio-temporal network to model trends and events, respectively. Moreover, long-term trends are used as a self-supervised signal in STWave+to teach overall temporal information into one-hour trends through a contrastive loss. Finally, reasonable future traffic can be predicted through the adaptive fusion of one-hour trends and events. Additionally, we incorporate a novel query sampling strategy and multi-scale graph wavelet positional encoding into the full graph attention network to efficiently and effectively model dynamic hierarchical spatial correlations. Extensive experiments on four traffic datasets show the superiority of our approach,i.e., the higher forecasting accuracy with lower computational cost.
Yuchen Fang 0001, Yanjun Qin, Haiyong Luo, Fang Zhao 0003, Kai Zheng 0001
IEEE Trans. Knowl. Data Eng.2
2023 When Spatio-Temporal Meet Wavelets: Disentangled Traffic Forecasting via Efficient Spectral Graph Attention Networks
abstract
Traffic forecasting is crucial for public safety and resource optimization, yet is very challenging due to the temporal changes and the dynamic spatial correlations of the traffic data. To capture these intricate dependencies, spatio-temporal networks, such as recurrent neural networks with graph convolution networks, graph convolution networks with temporal convolution networks, and temporal attention networks with full graph attention networks, are applied. However, previous spatio-temporal networks are based on end-to-end training and thus fail to handle the distribution shift in the non-stationary traffic time series. On the other hand, the efficient and effective algorithm for modeling spatial correlations is still lacking in prior networks.In this paper, rather than proposing yet another end-to-end model, we aim to provide a novel disentangle-fusion framework STWave to mitigate the distribution shift issue. The framework first decouples the complex traffic data into stable trends and fluctuating events, followed by a dual-channel spatio-temporal network to model trends and events, respectively. Finally, reasonable future traffic can be predicted through the fusion of trends and events. Besides, we incorporate a novel query sampling strategy and graph wavelet-based graph positional encoding into the full graph attention network to efficiently and effectively model dynamic spatial correlations. Extensive experiments on six traffic datasets show the superiority of our approach, i.e., the higher forecasting accuracy with lower computational cost.
Yuchen Fang 0001, Yanjun Qin, Haiyong Luo, Fang Zhao 0003, Bingbing Xu 0001, Liang Zeng 0002, Chenxing Wang 0001
ICDE2
2023 A Hybrid Approach for Driving Behavior Recognition: Integration of CNN and Transformer-Encoder with EEG data
abstract
Human factors are considered as one of the main causes affecting road traffic safety. Therefore, it is highly necessary to establish a driving behavior model for predicting driver behaviors and states, which can be used for risk monitoring. As an objective indicator that can accurately measure human cognitive and emotional states, electroencephalography (EEG) has attracted widespread attention from researchers due to its suitability and high temporal resolution for driver state and behavior measurement. Our study proposes a novel EEG-based driving behavior recognition algorithm that combines CNN and Transformer-Encoder modules. We deploy CNN to extract spatial and temporal features from EEG signals, capturing local dependencies between different time points. Subsequently, the encoder module of Transformer is utilized to handle long-term dependencies, enhancing the correlation between different time steps and improving classification accuracy. Experimental results demonstrate that our proposed model outperforms the baseline model on a self-constructed driving behavior classification dataset, providing a foundation for further in-depth research into real-world road traffic safety.
Yanjun Qin, Xiaoming Tao 0001
VTC Fall3
2023 Spatio-temporal hierarchical MLP network for traffic forecasting
Yanjun Qin, Haiyong Luo, Fang Zhao 0003, Yuchen Fang 0001, Xiaoming Tao 0001, Chenxing Wang 0001
Inf. Sci.1
2022 Next Point-of-Interest Recommendation with Auto-Correlation Enhanced Multi-Modal Transformer Network
abstract
Next Point-of-Interest (POI) recommendation is a pivotal issue for researchers in the field of location-based social networks. While many recent efforts show the effectiveness of recurrent neural network-based next POI recommendation algorithms, several important challenges have not been well addressed yet: (i) The majority of previous models only consider the dependence of consecutive visits, while ignoring the intricate dependencies of POIs in traces; (ii) The nature of hierarchical and the matching of sub-sequence in POI sequences are hardly model in prior methods; (iii) Most of the existing solutions neglect the interactions between two modals of POI and the density category. To tackle the above challenges, we propose an auto-correlation enhanced multi-modal Transformer network (AutoMTN) for the next POI recommendation. Particularly, AutoMTN uses the Transformer network to explicitly exploits connections of all the POIs along the trace. Besides, to discover the dependencies at the sub-sequence level and attend to cross-modal interactions between POI and category sequences, we replace self-attention in Transformer with the auto-correlation mechanism and design a multi-modal network. Experiments results on two real-world datasets demonstrate the ascendancy of AutoMTN contra state-of-the-art methods in the next POI recommendation.
Yanjun Qin, Yuchen Fang 0001, Haiyong Luo, Fang Zhao 0003, Chenxing Wang 0001
SIGIR1
2022 Memory attention enhanced graph convolution long short-term memory network for traffic forecasting
abstract
In recent years, traffic forecasting has gradually attracted attention in data mining because of the increasing availability of large-scale traffic data. However, it faces substantial challenges of complex temporal-spatial correlations in traffic. Recent studies mainly focus on modeling the local spatial correlations by utilizing graph neural networks and neglect the influence of long-distance spatial correlations. Besides, most existing works utilize recurrent neural networks-based encoder–decoder architecture to forecast multistep traffic volume and suffer from accumulative errors in recurrent neural networks. To deal with these issues, we propose the memory attention (MA) enhanced graph convolution long short-term memory network (MAEGCLSTM), a novel deep learning model for traffic forecasting. Specifically, MAEGCLSTM combines the MA and the vanilla graph convolution long short-term memory to capture global and local spatio-temporal dependencies, respectively. Then MAEGCLSTM utilizes a simplified GCLSTM to effectively fuse the global and local information. Moreover, we integrate the MAEGCLSTM into an encoder–decoder architecture to forecast multistep traffic volume. Besides MAEGCLSTM, we add the convolution neural network and encoder–decoder attention into the decoder to ease accumulative errors caused by iterative prediction and gain whole historical information from the encoder. Experiments on four real-world traffic data sets show that our model significantly outperforms by up to 6.07 % $6.07 \% $ improvement in L 1 $L1$ measure over 14 baselines.
Yanjun Qin, Fang Zhao 0003, Yuchen Fang 0001, Haiyong Luo, Chenxing Wang 0001
Int. J. Intell. Syst.1
2022 An abnormal driving behavior recognition algorithm based on the temporal convolutional network and soft thresholding
abstract
Most traffic accidents are caused by bad driving habits. Online monitoring of the abnormal driving behaviors of drivers can help reduce traffic accidents. Recently, abnormal driving behavior recognition based on the sensors' data embedded in commodity smartphones has attracted much attention. Though much progress has been made about driving behavior recognition, the existing works cannot achieve high recognition accuracy and show poor robustness. To improve the driving behaviors recognition accuracy and robustness, we propose an algorithm based on Soft Thresholding and Temporal Convolutional Network (S-TCN) for driving behavior recognition. In this algorithm, we first introduce a soft attention mechanism to learn the importance of different sensors. The TCN has the advantages of small memory requirement and high computational efficiency. And the soft thresholding can further filter the redundant features and extract the main features. So, we fuse the TCN and soft thresholding to improve the model's stability and accuracy. Our proposed model is extensively evaluated on four real public data sets. The experimental results show that our proposed model outperforms best state-of-the-art baselines by 2.24%.
Yunyun Zhao, Hongwei Jia, Haiyong Luo, Fang Zhao 0003, Yanjun Qin
Int. J. Intell. Syst.5
2022 Learning All Dynamics: Traffic Forecasting via Locality-Aware Spatio-Temporal Joint Transformer
abstract
Forecasting traffic flow and speed in the urban is important for many applications, ranging from the intelligent navigation of map applications to congestion relief of city management systems. Therefore, mining the complex spatio-temporal correlations in the traffic data to accurately predict traffic is essential for the community. However, previous studies that combined the graph convolution network or self-attention mechanism with deep time series models (e.g., the recurrent neural network) can only capture spatial dependencies in each time slot and temporal dependencies in each sensor, ignoring the spatial and temporal correlations across different time slots and sensors. Besides, the state-of-the-art Transformer architecture used in previous methods is insensitive to local spatio-temporal contexts, which is hard to suit with traffic forecasting. To solve the above two issues, we propose a novel deep learning model for traffic forecasting, named Locality-aware spatio-temporal joint Transformer (Lastjormer), which elaborately designs a spatio-temporal joint attention in the Transformer architecture to capture all dynamic dependencies in the traffic data. Specifically, our model utilizes the dot-product self-attention on sensors across many time slots to extract correlations among them and introduces the linear and convolution self-attention mechanism to reduce the computation needs and incorporate local spatio-temporal information. Experiments on three real-world traffic datasets, England, METR-LA, and PEMS-BAY, demonstrate that our Lastjormer achieves state-of-the-art performances on a variety of challenging traffic forecasting benchmarks.
Yuchen Fang 0001, Fang Zhao 0003, Yanjun Qin, Haiyong Luo, Chenxing Wang 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Fine-Grained Trajectory-Based Travel Time Estimation for Multi-City Scenarios Based on Deep Meta-Learning
abstract
Travel Time Estimation (TTE) is indispensable in intelligent transportation system (ITS). It is significant to achieve the fine-grained Trajectory-based Travel Time Estimation (TTTE) for multi-city scenarios, namely to accurately estimate travel time of the given trajectory for multiple city scenarios. However, it faces great challenges due to complex factors including dynamic temporal dependencies and fine-grained spatial dependencies. To tackle these challenges, we propose a meta learning based framework, MetaTTE, to continuously provide accurate travel time estimation over time by leveraging well-designed deep neural network model called DED, which consists of Data preprocessing module and Encoder-Decoder network module. By introducing meta learning techniques, the generalization ability of MetaTTE is enhanced using small amount of examples, which opens up new opportunities to increase the potential of achieving consistent performance on TTTE when traffic conditions and road networks change over time in the future. The DED model adopts an encoder-decoder network to capture fine-grained spatial and temporal representations. Extensive experiments on two real-world datasets are conducted to confirm that our MetaTTE outperforms nine state-of-art baselines, and improve 29.35% and 25.93% accuracy than the best baseline on Chengdu and Porto datasets, respectively.
Chenxing Wang 0001, Fang Zhao 0003, Haiyong Luo, Yanjun Qin, Yuchen Fang 0001
IEEE Trans. Intell. Transp. Syst.5
2021 Combining Residual and LSTM Recurrent Networks for Transportation Mode Detection Using Multimodal Sensors Integrated in Smartphones
abstract
In recent years, with the rapid development of public transportation, the ways people travel has become more diversified and complicated. Transportation mode detection, as a significant branch of human activity recognition (HAR), is of great importance in analyzing human travel patterns, traffic prediction and planning. Though many works have been devoted to transportation mode detection, there remains challenge for accurate and robust transportation pattern identification. In this paper, we propose a residual and LSTM recurrent networks-based transportation mode detection algorithm using multiple light-weight sensors integrated in commodity smartphones. Feature representation learning is adopted separately on multiple preprocessed sensor data using deep residual and LSTM network, which can enhance the identification accuracy and support one or more sensors. Residual units are introduced to accelerate the learning speed and enhance the accuracy of transportation mode detection. Furthermore, we also leverage the attention model to learn the significance of different features and different timesteps to enhance the recognition accuracy. Extensive experimental results on three datasets indicate that using our proposed model can achieve the best recognition accuracy for eight transportation modes including being stationary, walking, running, cycling, taking a car, taking a bus, taking a subway and taking a train, which outperforms other benchmark algorithms.
Chenxing Wang 0001, Haiyong Luo, Fang Zhao 0003, Yanjun Qin
IEEE Trans. Intell. Transp. Syst.4