Tao Deng 0002

dblp:69/6013-2 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0001-5094-5879ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DDL-Net: A Task-Balanced Panoptic Perception Network with Channel-Reorganized Attention in Autonomous Driving
abstract
Multi-task perception in traffic scenes requires unified representations for object recognition, region understanding, and structural reasoning. However, heterogeneous tasks such as detection and segmentation impose conflicting optimization objectives on shared features, leading to representation imbalance and reduced robustness. We propose DDL-Net, a task-balanced panoptic perception network that addresses this issue through coordinated representation learning and task-aware decoding, introducing a Channel-Rearranging Attention Interaction (CRAI) module built upon Channel Rearranging Boosting Attention (CRBA) units to reduce attention redundancy and enhance informative channel responses, improving representation of small and degraded targets. In addition, a task-specific decoding mechanism combining the Task-Guided Feature Balancing Module (TGBM) and Frequency-Aware Task-Interaction Guided Decoder (FTIG-Decoder) aligns shared features with task semantics, enabling balanced performance across detection and segmentation tasks. Experiments on BDD100K validate the effectiveness and robustness of the proposed design.
Bowei Fang, Yunyi Tang, Wentao Mu, Wenbo Liu 0006, Fei Yan 0006, Tao Deng 0002
ICMR6
2026 TranSpikformer: Enhanced transformer-based spiking neural networks with uniform attention augmentation and depthwise convolutional positional encoding
Tao Deng 0002, Chengfan Yang, Wenbo Liu 0006, Yi Huang 0022
Knowl. Based Syst.1
2026 Precise positioning of ultrasound-guided fine-needle aspiration biopsy of thyroid nodule
Tao Deng 0002, Shengqi Chen 0004, Chengfan Yang, Yi Huang 0022, Buyun Ma, Yang Chen 0060
Pattern Recognit.1
2025 SalM²: An Extremely Lightweight Saliency Mamba Model for Real-Time Cognitive Awareness of Driver Attention
abstract
Driver attention recognition in driving scenarios is a popular direction in traffic scene perception technology. It aims to understand human driver attention to focus on specific targets/objects in the driving scene. However, traffic scenes contain not only a large amount of visual information but also semantic information related to driving tasks. Existing methods lack attention to the actual semantic information present in driving scenes. Additionally, the traffic scene is a complex and dynamic process that requires constant attention to objects related to the current driving task. Existing models, influenced by their foundational frameworks, tend to have large parameter counts and complex structures. Therefore, this paper proposes a real-time saliency Mamba network based on the latest Mamba framework. As shown in Figure 1, our model uses very few parameters (0.08M, only 0.09~11.16% of other models), while maintaining SOTA performance or achieving over 98% of the SOTA model's performance.
Wentao Mu, Wenbo Liu 0006, Fei Yan 0006, Tao Deng 0002
AAAI6
2025 VP-YOLO: Robust Vehicle-Pedestrian Detection in Challenging Traffic Scenarios via A Human Visual Perception-Inspired Network
abstract
Intelligent vehicles need to provide rational driving strategies for assisted driving systems based on driving scenarios. Since pedestrians and vehicles are the main players in these scenarios, accurate detection and localization of pedestrians and vehicles are crucial for intelligent driving systems to make reliable decisions in dynamic environments. However, existing pedestrian and vehicle detection models often lack robustness under dynamic and complex traffic conditions, resulting in missed detections and false alarms, which pose significant safety risks. To address this problem, we categorize complex traffic scenarios into three typical challenges: long-distance, truncation, and occlusion, and focus on designing a novel enhancement stage to make the model more robust to these challenges. In this enhancement stage, inspired by human visual perception, we design a Visual Attention Module (VAM). This module can gather high-quality horizontal and vertical spatial features and efficiently interact between horizontal and vertical spatial features, enhancing the model’s perceptual ability by mimicking optic chiasm. Additionally, we use a Feature Reconstruction Module (FRM) to reduce redundant information in the feature maps and enhance the model’s inference ability. We conduct comprehensive experiments on the KITTI benchmark and Cityscapes dataset, and the experimental results demonstrate that our algorithm achieves state-of-the-art performance in various challenging scenarios.
Wenbo Liu 0006, Tao Deng 0002, Fei Yan 0006
ICASSP2
2025 HID-NAS: A Novel Neural Architecture Search Pipeline for High Information Density Data
abstract
Neural Architecture Search (NAS) is a core component of automated machine learning, enabling the automatic discovery of task-specific architectures. However, directly employing raw data for NAS faces challenges in terms of large data volume, low search efficiency, and potential data privacy concerns. To address these limitations, we condense large-scale target datasets into high information-density proxy datasets. And we propose a novel pipeline, HID-NAS, which skillfully exploits High Information Density(HID) data for efficient neural architecture search. At the algorithmic level, our algorithm focuses on the inherent problem of DARTS-derived algorithms, i.e., the discrepancy between trained supernet and selected architectures. We pioneer the phenomenon of "Selection Error" that occurs during architecture selection and introduce an attention mechanism that focuses on candidate edges to improve the accuracy of pipeline-selected architectures. Furthermore, we also formulate a regularization mechanism that employs the high information density data of the proxy dataset to transform the architectural parameter updates into an adversarial game and adjust the strengths through a normalization mechanism. These three key components – attention, regularization, and normalization – allow our pipeline to efficiently identify high-quality models using the distilled dataset. Experimental results demonstrate a significant reduction in search time, achieving high-quality models within just two minutes.
Wenbo Liu 0006, Tao Deng 0002, Fei Yan 0006
ICASSP2
2025 Ultra-Lightweight Thyroid Puncture Positioning Detection Guided by Nodule Location
Shengqi Chen 0004, Yi Huang 0022, Chengfan Yang, Buyun Ma, Fei Yan 0006, Yang Chen 0060, Tao Deng 0002
PRCV (13)7
2025 Driving Fixation Prediction for Clear-to-Adverse Weather Scenes via Adversarial Unsupervised Domain Adaptation
Wentao Mu, Wenbo Liu 0006, Yanghua Zhang, Fei Yan 0006, Tao Deng 0002
PRCV (12)7
2025 Improving vehicle detection accuracy in complex traffic scenes through context attention and multi-scale feature fusion module
Wenbo Liu 0006, Binglin Zhao, Tao Deng 0002, Fei Yan 0006
Appl. Intell.4
2025 Continuous-Discrete Alignment Optimization for efficient differentiable neural architecture search
abstract
Differential Architecture Search (DARTS) has become a prominent technique for neural architecture search in recent years. Despite its merits, the issue of discretization discrepancy within DARTS still necessitates further exploration, as it can degrade in performance. In this paper, we introduce a novel algorithm termed Continuous–Discrete Alignment Optimization (DARTS-CDAO), designed to address the discretization discrepancy and thereby enhance the robustness and generalization capabilities of the discovered neural architectures. Our proposed DARTS-CDAO algorithm seamlessly integrates the discretization process into the training phase of the architecture parameters, thereby bolstering the search algorithm’s adaptability to the inherent discretization processes. Specifically, our methodology commences by formalizing the process of architecture parameter discretization. Subsequently, we introduce a coarse gradient weighting algorithm that is employed to update the architecture parameters, effectively minimizing the divergence between the representation of continuous and discrete parameters. Rigorous theoretical analysis, coupled with extensive experimental outcomes, substantiates that our proposed approach can elevate the performance of the searched models. Notably, this enhancement is achieved without incurring additional search time, rendering DARTS more robust and endowed with a heightened capacity for generalization.
Wenbo Liu 0006, Jia Wu 0005, Tao Deng 0002, Fei Yan 0006
Eng. Appl. Artif. Intell.3
2025 VP-YOLO: A human visual perception-inspired robust vehicle-pedestrian detection model for complex traffic scenarios
Wenbo Liu 0006, Xiaoyun Qiao, Tao Deng 0002, Fei Yan 0006
Expert Syst. Appl.4
2025 A new pipeline with ultimate search efficiency for neural architecture search
Wenbo Liu 0006, Xiaoyun Qiao, Tao Deng 0002, Fei Yan 0006
Neural Networks4
2025 Parallel Multi-Path Network for Ocular Disease Detection Inspired by Visual Cognition Mechanism
abstract
Various ocular diseases such as cataracts, glaucoma, and diabetic retinopathy have become several major factors causing non-congenital visual impairment, which seriously threatens people's vision health. The shortage of ophthalmic medical resources has brought huge obstacles to large-scale ocular disease screening. Therefore, it is necessary to use computer-aided diagnosis (CAD) technology to achieve large-scale screening and diagnosis of ocular diseases. In this work, inspired by the human visual cognition mechanism, we propose a parallel multi-path network for multiple ocular diseases detection, called PMP-OD, which integrates the detection of multiple common ocular diseases, including cataracts, glaucoma, diabetic retinopathy, and pathological myopia. The bottom-up features of the fundus image are extracted by a common convolutional module, the Low-level Feature Extraction module, which simulates the non-selective pathway. Simultaneously, the top-down vessel and other lesion features are extracted by the High-level Feature Extraction module that simulates the selective pathway. The retinal vessel and lesion features can be regarded as task-driven high-level semantic information in the physician's disease diagnosis process. Then, the features are fused by a feature fusion module based on the attention mechanism. Finally, the disease classifier gives prediction results according to the integrated multi-features. The experimental results indicate that our PMP-OD model outperforms other state-of-the-art (SOTA) models on an ocular disease dataset reconstructed from ODIR-5K, APTOS-2019, ORIGA-light, and Kaggle.
Tao Deng 0002, Yi Huang 0022, Chengfan Yang
IEEE J. Biomed. Health Informatics1
2025 VP2Net: Visual Perception-Inspired Network for Exploring the Causes of Drivers' Attention Shift
abstract
With the rapid development of autonomous driving technology, the recognition/understanding of driving events has become increasingly important for improving road safety. Existing methods for recognizing driving events rely solely on the inherent features of driving scenes, lacking real-time modeling of driver attention and the integration of driver attention for understanding driving events. Research has shown that understanding driver attention will be beneficial for subsequent analysis of driving events. We propose the attention-based driving event dataset (ADED), which includes rich driving scenes, eye movement data, reasons for attention shifts, and event time windows. It enables the use of prior information about driver attention to guide the recognition of driving events. Based on our dataset, we propose a visual dual-perception network, named VP2Net, to explore the reasons behind driver attention shifts. The goal of VP2Net is to use driver attention to guide the recognition of driving events. Inspired by the human visual dual cognition process mechanism, we build a bottom-up sequential information encoding branch for extracting spatio-temporal low-level information in the driving scene. Additionally, we establish a top-down attention perceptual encoding branch that simulates the driver’s high-level visual cognitive process. It not only captures the driver’s spatial attention allocation (“where to focus”) but also performs a temporal dimensional perceptual enhancement (“when to focus”), allowing us to extract the driver’s spatial attention enhancement information. We use the driver’s spatial attention enhancement information to guide the fusion of spatio-temporal information of the driving scene and selectively highlight the core objects/areas in the current driving task/event. Finally, we compare our proposed model with other SOTA networks and visualize the results of the key components of the model. Our code is available athttps://github.com/zhao-chunyu/VP2Net
Tao Deng 0002, Pengcheng Du, Wenbo Liu 0006, Yi Huang 0022, Fei Yan 0006
IEEE Trans. Intell. Transp. Syst.2
2024 PHANet: Progressive Hybrid Attention Network for Enhanced Video Deraining
Tao Deng 0002, Chengfan Yang, Yi Huang 0022, Wenbo Liu 0006
PRCV (8)2
2024 DARTS-CGW: Research on Differentiable Neural Architecture Search Algorithm Based on Coarse Gradient Weighting
Wenbo Liu 0006, Tao Deng 0002, Rui An, Fei Yan 0006
PRCV (3)2
2024 Driving Visual Saliency Prediction of Dynamic Night Scenes via a Spatio-Temporal Dual-Encoder Network
abstract
Driving at night is more challenging and dangerous than driving during the day. Modeling driver eye movement and attention allocation during night driving can help guide unmanned intelligent vehicles and improve safety during similar situations. However, until now, few studies have modeled a drivers’ true fixations and attention allocation in specific night circumstance. Therefore, we collected an eye tracking dataset from 30 experienced drivers while they viewed night driving videos under a hypothetical driving condition, termed Driver Fixation Dataset in night (DrFixD(night)). Based on DrFixD(night) which includes multiple drivers’ attention allocation, we proposed a spatio-temporal dual-encoder network model, named as STDE-Net, to improve saliency detection in night driving condition. The model includes three modules: i) spatio-temporal dual encoding module, ii) fusion module based on attention mechanism, and iii) decoding module. A convolutional LSTM is employed to learn the time connection of video sequences, and a convolution neural network combined pyramid dilated convolution is adopted to extract spatial features in the spatio-temporal dual encoding module. The attention mechanism is exploited to fuse the temporal and spatial features together and selectively highlight the significant features in night traffic scene. We compared the proposed model with other traditional methods and deep learning models, both qualitatively and quantitatively, and found that the proposed model can predict driver’s fixation more accurately. Specifically, the proposed model not only predicts the main goals, but also predicts the important sub goals, such as pedestrians, bicycles and so on, showing excellent prediction of dimly lit targets at night.
Tao Deng 0002, Lianfang Jiang, Yi Shi 0012, Jiang Wu 0016, Zhangbi Wu, Shun Yan
IEEE Trans. Intell. Transp. Syst.1
2023 What Causes a Driver's Attention Shift? A Driver's Attention-Guided Driving Event Recognition Model
abstract
Despite much effort to try to research driver's spatial attention allocation in driving situations, the computer vision community rarely focuses on what causes driver's attention shifts. In this paper, we built an attention-based driving event dataset (ADED) constructed from the attention distributed on the traffic participants or elements and proposed a model using driver's attention as the guidance to better recognize the events that lead driver's attention shifts. We relabeled and redivided BDD-A, a driver attention dataset in critical traffic situations, into six different semantic categories of driving events. The new dataset is introduced for driving event recognition. In addition, we proposed a special driving event recognition model (called DER-Net) with driver's attention guidance to recognize the event causes a driver's attention shift. In DER-Net, a driver's attention-guided (DAG) branch is constructed to consider the driver's spatiotemporal attention information. The proposed model achieves a superior performance compared to other state-of-the-art models in action recognition. In ablation study, many experiments are conducted to discuss proper length of the sequence to input the model, the optimal criterion to find out the frame when driver's attention shifts and the appropriate function to quantify different attention maps in neural network.
Pengcheng Du, Tao Deng 0002, Fei Yan 0006
IJCNN2
2022 TARConvGRU: A Cross-dimension Spatiotemporal Model for Lane Detection
abstract
Spatiotemporal information plays a critical role in the autonomous driving environment. In the task of lane detection, the features extracted by the existing spatiotemporal work from a video or consecutive frames contain a lot of irrelevant information, which reduces the performance of the model to a certain extent. To address the trouble, we proposed a novel lane detection model with the TARConvGRU module. This module can help our model focus on the feature representation and feature location of lane lines by channel and spatial operations and consciously guides some computing resources tend to the most likely lane line features by capturing the cross-dimension information interaction. Furthermore, we demonstrate the effectiveness of the proposed model on three popular lane detection benchmarks: TuSimple, Unsupervised Labeled Lane Markers (Unsupervised LLAMAS), and Dynamic Vision Sensor Dataset (DET). The experiments show that our model achieves competitive results compared with other state-of-the-art models. In addition, we discuss some significant problems encountered in the designing process in detail in the ablation study.
Tao Deng 0002, Fei Yan 0006
IJCNN2
2022 Fastest containment control of discrete-time multi-agent systems using static linear feedback protocol
Fei Yan 0006, Tao Feng 0006, Tao Deng 0002, Yue Zhao 0004
Inf. Sci.4
2022 ID-YOLO: Real-Time Salient Object Detection Based on the Driver's Fixation Region
abstract
Object detection is an important task for self-driving vehicles or advanced driver assistant systems (ADASs). Additionally, visual selective attention is a crucial neural mechanism in a driver’s vision system that can rapidly filter out unnecessary visual information in a driving scene. Some existing models detect all objects in driving scenes from the aspect of computer vision. However, in a rapidly changing driving environment, detecting salient or critical objects appearing in drivers’ interested or safety-relevant areas is more useful for ADASs. In this paper, we managed to detect salient and critical objects based on drivers’ fixation regions. To this end, we built an augmented eye tracking object detection (ETOD) dataset based on driving videos with multiple drivers’ eye movement collected by Denget al.Furthermore, we proposed a real-time salient object detection network named increase-decrease YOLO (ID-YOLO) to discriminate the critical objects within the drivers’ fixation region. The proposed ID-YOLO shows excellent detection of major objects that drivers are concerned about during driving. Compared with the present object detection models in autonomous and assisted driving systems, our object detection framework simulates the selective attention mechanism of drivers. Thus, it does not detect all of the objects appearing in the driving scenes but only detects the most relevant ones for driving safety. It can largely reduce the interference of irrelevant scene information, showing potential practical applications in intelligent or assisted driving systems.
Long Qin 0002, Yi Shi 0012, Yahui He, Junrui Zhang 0010, Yongjie Li 0001, Tao Deng 0002
IEEE Trans. Intell. Transp. Syst.7
2022 Lane Detection Model Based on Spatio-Temporal Network With Double Convolutional Gated Recurrent Units
abstract
Lane detection is one of the indispensable and key elements of self-driving environmental perception. Many lane detection models have been proposed, solving lane detection under challenging conditions, including intersection merging and splitting, curves, boundaries, occlusions and combinations of scene types. Nevertheless, lane detection will remain an open problem for some time to come. The ability to cope well with those challenging scenes impacts greatly the applications of lane detection on advanced driver assistance systems (ADASs). In this paper, a spatio-temporal network with double Convolutional Gated Recurrent Units (ConvGRUs) is proposed to address lane detection in challenging scenes. Both of ConvGRUs have the same structures, but different locations and functions in our network. One is used to extract the information of the most likely low-level features of lane markings. The extracted features are input into the next layer of the end-to-end network after concatenating them with the outputs of some blocks. The other one takes some continuous frames as its input to process the spatio-temporal driving information. Extensive experiments on the large-scale TuSimple lane marking challenge dataset and Unsupervised LLAMAS dataset demonstrate that the proposed model can effectively detect lanes in the challenging driving scenes. Our model can outperform the state-of-the-art lane detection models.
Tao Deng 0002, Fei Yan 0006, Wenbo Liu 0006
IEEE Trans. Intell. Transp. Syst.2
2021 Driving Video Fixation Prediction Model Via Spatio-Temporal Networks and Attention Gates
abstract
Driving fixation prediction is becoming an essential research problem in human-like driving systems or advanced driver assistance systems (ADAS) in a dynamic driving environment. However, it is still a lack of driving video fixation prediction models that can dynamically predict drivers’ fixational locations. In this work, we propose a driving video fixation prediction model via spatio-temporal networks and attention gates method, named as DSTANet, to predict drivers’ attention in the dynamic driving videos. The spatial and temporal driving information are both considered in DSTANet by convolutional long short-term memory (ConvLSTM). In addition, the human attention mechanism is designed to filter some driving-irrelevant information via the attention gates (AGs). The experimental results indicate that the proposed DSTANet outperforms the state-of-the-art saliency models and predicts drivers’ attentional spatial locations more accurately. Furthermore, the time-series prediction results show that DSTANet is more robust and coherent than others and includes more temporal information.
Tao Deng 0002, Fei Yan 0006
ICME1
2020 How Do Drivers Allocate Their Potential Attention? Driving Fixation Prediction via Convolutional Neural Networks
abstract
The traffic driving environment is a complex and dynamic changing scene in which drivers have to pay close attention to salient and important targets or regions for safe driving. Modeling drivers' eye movements and attention allocation in traffic driving can also help guiding unmanned intelligent vehicles. However, until now, few studies have modeled drivers' true fixations and allocations while driving. To this end, we collect an eye tracking dataset from a total of 28 experienced drivers viewing 16 traffic driving videos. Based on the multiple drivers' attention allocation dataset, we propose a convolutional-deconvolutional neural network (CDNN) to predict the drivers' eye fixations. The experimental results indicate that the proposed CDNN outperforms the state-of-the-art saliency models and predicts drivers' attentional locations more accurately. The proposed CDNN can predict the major fixation location and shows excellent detection of secondary important information or regions that cannot be ignored during driving if they exist. Compared with the present object detection models in autonomous and assisted driving systems, our human-like driving model does not detect all of the objects appearing in the driving scenes, but it provides the most relevant regions or targets, which can largely reduce the interference of irrelevant scene information.
Tao Deng 0002, Long Qin 0002, Thuyen Ngo, B. S. Manjunath
IEEE Trans. Intell. Transp. Syst.1
2018 Tensor Sensing for Rf Tomographic Imaging
abstract
Radio-frequency (RF) tomographic imaging is a promising technique for inferring multi-dimensional physical space by processing RF signals traversed across a region of interest. However, conventional RF tomography schemes are generally based on vector compressed sensing, which ignores the geometric structures of the target spaces and leads to low recovery precision. The recently proposed transform-based tensor model is more appropriate for sensory data processing, as it helps exploit the geometric structures of the three-dimensional target and improve the recovery precision. In this paper, we propose a novel tensor sensing approach that achieves highly accurate estimation for real-world three-dimensional spaces. First, we use the transform-based tensor model to formulate a tensor sensing problem, and propose a fast alternating minimization algorithm called Alt-Min. Secondly, we drive an algorithm which is optimized to reduce memory and computation requirements. Finally, we present evaluation of our Alt-Min approach using IKEA 3D data and demonstrate significant improvement in recovery error and convergence speed compared to prior tensor-based compressed sensing.
Tao Deng 0002, Feng Qian 0005, Xiao-Yang Liu, Manyuan Zhang, Anwar Elwalid
ICME1
2018 Learning to Boost Bottom-Up Fixation Prediction in Driving Environments via Random Forest
abstract
Saliency detection, an important step in many computer vision applications, can, for example, predict where drivers look in a vehicular traffic environment. While many bottom-up and top-down saliency detection models have been proposed for fixation prediction in outdoor scenes, no specific attempt has been made for traffic images. Here, we propose a learning saliency detection model based on a random forest (RF) to predict drivers' fixation positions in a driving environment. First, we extract low-level (color, intensity, orientation, etc.) and high-level (e.g., the vanishing point and center bias) features and then predict the fixation points via RF-based learning. Finally, we evaluate the performance of our saliency prediction model qualitatively and quantitatively. We use quantitative evaluation metrics that include the revised receiver operating characteristic (ROC), the area under the ROC curve value, and the normalized scan-path saliency score. The experimental results on real traffic images indicate that our model can more accurately predict a driver's fixation area, while driving than the state-of-the-art bottom-up saliency models.
Tao Deng 0002, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.1
2016 Where Does the Driver Look? Top-Down-Based Saliency Detection in a Traffic Driving Environment
abstract
A traffic driving environment is a complex and dynamically changing scene. When driving, drivers always allocate their attention to the most important and salient areas or targets. Traffic saliency detection, which computes the salient and prior areas or targets in a specific driving environment, is an indispensable part of intelligent transportation systems and could be useful in supporting autonomous driving, traffic sign detection, driving training, car collision warning, and other tasks. Recently, advances in visual attention models have provided substantial progress in describing eye movements over simple stimuli and tasks such as free viewing or visual search. However, to date, there exists no computational framework that can accurately mimic a driver's gaze behavior and saliency detection in a complex traffic driving environment. In this paper, we analyzed the eye-tracking data of 40 subjects consisted of nondrivers and experienced drivers when viewing 100 traffic images. We found that a driver's attention was mostly concentrated on the end of the road in front of the vehicle. We proposed that the vanishing point of the road can be regarded as valuable top-down guidance in a traffic saliency detection model. Subsequently, we build a framework of a classic bottom-up and top-down combined traffic saliency detection model. The results show that our proposed vanishing-point-based top-down model can effectively simulate a driver's attention areas in a driving environment.
Tao Deng 0002, Kaifu Yang, Yongjie Li 0001
IEEE Trans. Intell. Transp. Syst.1