EDBT 2026 Demo / reviewers in the wild / expert
Yantao Lu
dblp:131/1381
· DBLP profile ↗
27ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-3103-1067ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STEP-Nav: Spatial-Temporal Efficient Visual Token Pruning for Vision-and-Language Navigation with Large Language ModelsabstractVision-and-Language Navigation (VLN) plays a critical role in tasks of embodied AI, particularly in unseen environments following natural language instructions. Recent advancements leverage large language models (LLMs) to improve the accuracy and generalizability of VLN systems by encoding image sequences as dense token representations. However, this tokenization approach incurs substantial computational overhead due to two key inefficiencies: 1) ego-centric camera views often include navigation-irrelevant re- gions (e.g., sky or distant backgrounds), and 2) high-frame-rate image sequences introduce temporal redundancy. To address these challenges, we propose Spatial-Temporal Efficient Visual Token Pruning (STEP-Nav), a unified frame- work that simultaneously prunes redundant visual tokens and fine-tunes VLN models to preserve navigation performance. In particular, STEP-Nav incorporates a distance- and content-aware token evaluation mechanism to remove irrelevant tokens at the spatial level, along with temporal level similarity-based filtering to reduce redundancy across sequential frames. To ensure pruning does not harm task performance, we introduce a distortion-aware fine-tuning strategy that aligns pruned-token representations with their full-token counterparts while maintaining navigation accuracy. Experiments on the R2R and RxR benchmarks using Navid-CE and NavGPT-2 as base models demonstrate that STEP-Nav preserves over 95% of the performance while reducing 66.7% of tokens, outperforming existing token pruning baselines. Yantao Lu, Ning Liu 0007, Ying Zhang 0060, Jinchao Chen, Chenglie Du |
AAAI | 1 |
| 2026 | CEST: Enhancing Multi-Agent Perception via Communication-Efficient Spatial-Temporal Fusion
Jinchao Chen, Qiuhao Shu, Yantao Lu, Ying Zhang 0060 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | Interactive Vehicle Trajectory Prediction Based on Parameterized Transfer Learning Using Encoder-Decoder NetworkabstractVehicle trajectory prediction is important for automated vehicles to understand driving scenarios. This paper proposes an encoder-decoder network-based parameterized transfer learning (EDN-PTL) model to predict vehicle trajectory. To improve trajectory prediction accuracy, the motion interaction between the target vehicle and the surrounding vehicles is considered, and a multidimensional spatiotemporal input expansion (MSIA) strategy is proposed to extend the feature dimensions. Additionally, global and local scale features, as well as long and short horizon features, are extracted and used for interactive vehicle trajectory prediction by a CNN and LSTM-based encoder-decoder network (CNN-LSTM-EDN). Moreover, the features extracted by CNN-LSTM-EDN are integrated using a stacked convolutional social pooling network (SCSPN). To enhance the environmental adaptability of the trajectory prediction model, a PTL strategy is proposed to enable transfer learning capabilities of EDN-PTL. Based on the PTL strategy, trajectory prediction accuracy is maintained even when applied to untrained environments. The proposed EDN-PTL model is validated on three types of publicly available naturalistic datasets and compared with several baselines and state-of-the-art (SOTA) methods. The validation results demonstrate that the proposed EDN-PTL achieves better prediction accuracy, robustness, and environmental adaptability compared to the baselines and SOTA methods. Ying Zhang 0060, Tingyi Zhao, Chuan Hu 0003, Jinchao Chen, Yantao Lu, Chenglie Du |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2026 | Energy-aware Scheduling of Workflow Applications Towards Schedule Length Optimization in Heterogeneous Distributed Embedded SystemsabstractEnergy optimization constitutes a paramount design consideration in the realm of embedded systems development since these devices are inherently constrained by finite battery resources. Designing and developing an effective energy-aware scheduling approach is a desirable work to provide excellent processing capability while keeping the energy consumption under control. Although previous approaches can obtain reasonable scheduling solutions for tasks with energy consumption constraints, they are computationally expensive and have deficiencies in effectiveness or efficiency due to unfair or inefficient energy pre-assignment strategies. In this article, we study the energy-aware workflow scheduling problem and present a three-stage list-based approach to minimize the schedule length of workflows in heterogeneous distributed embedded systems. First, the workflow applications and energy consumption of processors are modelled, and the energy-aware workflow scheduling problem is formulated as a non-linear mixed integer programming one with various dependency and energy constraints. Then, with an effective task prioritization strategy and a reasonable energy pre-assignment strategy, a three-stage list-based scheduling approach is proposed to schedule the tasks and minimize the schedule length of workflows. Experiments on randomly-generated and real-life workflows demonstrate that our proposed approach constantly outperforms the existing approaches and our algorithm can, respectively, reduce the normalized schedule length and the deviation ratio by 16.7% and 7.6% in average. Jinchao Chen, Qinwei Zhang, Pengcheng Han, Ying Zhang 0060, Yantao Lu, Pengyi Zheng |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | LaTP: LiDAR-aided multimodal token pruning for efficient trajectory prediction of autonomous driving
Yantao Lu, Ning Liu 0007, Yilan Li, Jinchao Chen, Ying Zhang 0060, Yichen Zhu 0001, Senem Velipasalar |
Neural Networks | 1 |
| 2025 | Cross-task and time-aware adversarial attack framework for perception of autonomous driving
Yantao Lu, Ning Liu 0007, Yilan Li, Jinchao Chen, Senem Velipasalar |
Pattern Recognit. | 1 |
| 2025 | Dual-Centralized Q-Network-Based Reinforcement Learning for Cooperative Path Planning of Multiple UAVs
Jinchao Chen, Chongde Ren, Yujiao Hu, Ying Zhang 0060, Yantao Lu, Qing Li 0022, Tao You, Joel J. P. C. Rodrigues |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | QCTF: A Quantized Communication and Transferable Fusion Framework for Multi-Agent Collaborative PerceptionabstractCollaborative perception effectively mitigates issues such as limited field of view and occlusion by enabling multiple agents to share perceptual information. Despite its advantages, challenges persist in complex environments due to factors such as limited communication bandwidth and noisy poses, which may potentially degrade system performance. Meanwhile, a substantial amount of simulation data is widely adopted in collaborative perception to achieve high precision and real-time detection. However, the domain gap between simulated and real-world environments may result in weakened collaborative performance and hindered generalization ability. In this work, we focus on the multi-agent collaborative perception problem and propose a quantized communication and transferable fusion framework, namedQCTF, to efficiently minimize the bandwidth overhead and enhance real-world perception by leveraging unlabeled data for improved adaptability. First, we present a quantized communication method that employs multi-scale residual indices and an optimized codebook to extract robust representations while minimizing bandwidth usage. Then, we design a channel-aware selection strategy that adjusts the bandwidth volume and compensates for the quantized representation by combining the prioritized critical features with the channel dimension. Finally, we adopt a transferable fusion module to effectively bridge the simulation-to-reality domain gaps and improve perceptual capability through multi-scale adaptation discriminators. Experiments on both simulated and real-world datasets are conducted to evaluate the effectiveness of the proposed framework, and the results demonstrate that our approach consistently outperforms the existing methods in limited communication bandwidth and domain adaptation scenarios. Jinchao Chen, Qiuhao Shu, Yantao Lu, Ying Zhang 0060 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Extrinsic-and-Intrinsic Reward-Based Multi-Agent Reinforcement Learning for Multi-UAV Cooperative Target EncirclementabstractDue to their high flexibility and strong maneuverability, unmanned aerial vehicles (UAVs) have attracted lots of attention and are widely employed in many fields. Especially in target encirclement applications, UAVs have shown great advantages in adaptability and reliability, and can efficiently fly to and evenly surround the targets in complex and dynamic environments. In this paper, we concentrate on the cooperative target encirclement problem of heterogeneous UAVs and try to propose a multi-agent reinforcement learning approach to solve the problem. First, with the models of heterogeneous UAVs and obstacles, we analyze the collision avoidance, motion continuity, and energy consumption constraints of UAVs, and formulate the cooperative target encirclement problem as a multi-constraint combinatorial optimization one. Then, inspired by the humans’ learning experience that curiosity provides a powerful motivator for humans to explore, discover, and acquire new knowledge, we propose an extrinsic-and-intrinsic reward-based multi-agent reinforcement learning approach to cooperatively control the behaviors of UAVs and achieve the target encirclement missions. Simulation experiments with randomly generated environments are conducted to evaluate the performance of our approach, and the results show that our approach has a significant advantage in terms of average reward, encirclement success rate, encirclement time, and encirclement energy consumption. Jinchao Chen, Ying Zhang 0060, Yantao Lu, Qiuhao Shu, Yujiao Hu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Vision-Based Geometric Model for Accurate and Fast Lane Recognition in Complex ConditionsabstractLane recognition is an important component of autonomous driving system and advanced driving assistance system (ADAS) for intelligent vehicles. In complex driving conditions, accurate and fast lane recognition is a challenging issue. In this paper, a vision-based geometric model (VBGM) is proposed for accurate and fast lane recognition in complex conditions. The framework of the VBGM includes an image preprocessing stage and a lane recognition stage. In the image preprocessing stage, the region of interest (ROI) is extracted from the original image, and the original image is transformed into an undistorted greyscale image. In the lane recognition stage, the lane contour is first extracted using the Roberts operator. Then, to accurately and quickly recognize the lane marking, a lane recognition coordinate system (LRCS) and a rotational LRCS (R-LRCS) are constructed. The distracting contours in abnormal regions are padded based on the LRCS using a contextual frames correlation (CFC) strategy, and the midpoints of the lane contour are identified based on the R-LRCS. Finally, an adaptive-order polynomial fitting model is built to fit the lane marking according to the midpoints in the LRCS. To evaluate the effectiveness of the proposed method, two state-of-the-art methods are selected for comparison. The comparative results indicate that the proposed method possesses a higher recognition rate and speed for lane recognition in complex conditions. Ying Zhang 0060, Shuaishuai Ge, Tingyi Zhao, Jinchao Chen, Tao You, Yantao Lu, Chenglie Du |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Non-Preemptive Scheduling of Periodic Tasks with Data Dependencies in Heterogeneous Multiprocessor Embedded SystemsabstractHeterogeneous multiprocessor architecture is frequently employed as an economical and efficient means of providing excellent parallel processing capabilities while keeping production cost and power consumption under control. Although this architecture achieves significant performance enhancement and cost reduction, it results in a serious task allocation and scheduling problem, especially for periodic tasks with data dependencies, all of which should be reasonably scheduled and executed in a timely manner such that their deadlines and dependence requirements could be satisfied even if the worst happens. In this article, we concentrate on the non-preemptive scheduling problem of periodic tasks with data dependencies upon heterogeneous multiprocessor platforms. First, with models of data-dependent tasks and heterogeneous processors, we analyze the time, space, precedence, and data dependence constraints of tasks and design an exact formulation based on the mixed integer linear programming to completely explore the solution space and produce the optimal solutions. Then, by constructing a directed acyclic graph to depict the dependence relationship of jobs generated by tasks, we propose an efficient off-line list-based scheduling algorithm to provide a reasonable time and processor allocation for each job, with a view to minimizing the completion time of jobs. Experiments with randomly generated tasks are performed to evaluate the effectiveness and efficiency of the proposed algorithm, and the experimental results show that our algorithm can averagely enhance the scheduling success ratio by 28.5%, and, respectively, reduce the task completion time and the deviation ratio by 23.3% and 17.2%, on average. Jinchao Chen, Ying Zhang 0060, Yantao Lu, Qing Li 0022, Qiuhao Shu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | AlterMOMA: Fusion Redundancy Pruning for Camera-LiDAR Fusion Models with Alternative Modality MaskingabstractCamera-LiDAR fusion models significantly enhance perception performance in autonomous driving. The fusion mechanism leverages the strengths of each modality while minimizing their weaknesses. Moreover, in practice, camera-LiDAR fusion models utilize pre-trained backbones for efficient training. However, we argue that directly loading single-modal pre-trained camera and LiDAR backbones into camera-LiDAR fusion models introduces similar feature redundancy across modalities due to the nature of the fusion mechanism. Unfortunately, existing pruning methods are developed explicitly for single-modal models, and thus, they struggle to effectively identify these specific redundant parameters in camera-LiDAR fusion models. In this paper, to address the issue above on camera-LiDAR fusion models, we propose a novelty pruning framework Alternative Modality Masking Pruning (AlterMOMA), which employs alternative masking on each modality and identifies the redundant parameters. Specifically, when one modality parameters are masked (deactivated), the absence of features from the masked backbone compels the model to reactivate previous redundant features of the other modality backbone. Therefore, these redundant features and relevant redundant parameters can be identified via the reactivation process. The redundant parameters can be pruned by our proposed importance score evaluation function, Alternative Evaluation (AlterEva), which is based on the observation of the loss changes when certain modality parameters are activated and deactivated. Extensive experiments on the nuScene and KITTI datasets encompassing diverse tasks, baseline models, and pruning algorithms showcase that AlterMOMA outperforms existing pruning methods, attaining state-of-the-art performance. Yantao Lu, Ning Liu 0007, Jinchao Chen, Ying Zhang 0060 |
NeurIPS | 2 |
| 2024 | Work-in-Progress: Towards Real-time Collaborative 3D Object Detection Systems with Request-free CommunicationabstractCollaborative 3D object detection by sharing features among agents significantly enhances performance compared to single-agent detection. However, directly sharing full-sized features introduces a large communication bandwidth load. To address this challenge, existing collaborative methods adopt a request-response framework, where the ego agent sends a request, and collaborative agents respond with only the necessary parts of the features after analyzing the request. However, the frequent communication in this request-response cycle impacts real-time system performance in real-world environments by increasing overall processing time and raising the risk of message loss and communication delays. To address this challenge and enable real-time system implementation, we propose a request-free collaborative 3D object detection framework that eliminates the request-response cycle through a novel request-free response generator, named Position and Occlusion Response Generator (PORG). PORG consists of two specialized components, Position-aware Mask Generator (PaMG) and Occlusion-aware Feature Mask Generator (OaMG), which use attention mechanisms to generate the necessary response features without the request from the ego agent. To evaluate the efficiency of our proposed PORG, we conducted evaluations on both public datasets and real-world settings. We provide system implementation for both the request-response and request-free frameworks on Jetson Orin Series embedded devices, and extensive evaluation shows that PORG outperforms the baselines, achieving higher Average Precision (AP) with lower communication bandwidth in public datasets and superior real-time performance on embedded devices. Yantao Lu, Ning Liu 0007, Jinchao Chen, Ying Zhang 0060 |
RTSS | 2 |
| 2024 | CrossPrune: Cooperative pruning for camera-LiDAR fused perception models of autonomous driving
Yantao Lu, Ning Liu 0007, Yilan Li, Jinchao Chen, Ying Zhang 0060, Zifu Wan |
Knowl. Based Syst. | 1 |
| 2024 | Time-aware and task-transferable adversarial attack for perception of autonomous vehicles
Yantao Lu, Haining Ren, Weiheng Chai, Senem Velipasalar, Yilan Li |
Pattern Recognit. Lett. | 1 |
| 2024 | Improving robustness and efficiency of edge computing models
Yilan Li, Yantao Lu, Helei Cui, Senem Velipasalar |
Wirel. Networks | 2 |
| 2023 | Work-in-Progress: Time-Aware Formation Control of Connected and Automated Vehicle Platoon Based on Weighted Graph TheoryabstractThe regulation time is an important index for formation switching control of connected and automated vehicle (CA V) platoon. This paper proposes a time-aware formation control (T AFC) strategy to improve the formation switching performance of CA V platoon. To construct an effective information sharing mechanism among the vehicles in the platoon, a unidirectional weighted graph is designed to construct the relation of the CA V platoon and calculate the impact factor between two different vehicles. Based on the unidirectional weighted graph, the time-aware requirement is converted to the regulation order problem, and the regulation order which corresponding to the minimum time is designed. According to the T AFC, the qualitative regulation strategy of the CA V platoon and the quantitative tune-up strategy of the vehicles are determined. In order to analyze the performance of the TAFC strategy, two state-of-art methods are selected as the benchmarked methods. The validation results demonstrate the proposed method possesses better performance for formation switching control compared with the benchmarked methods. Ying Zhang 0060, Tingyi Zhao, Tao You, Yantao Lu, Jinchao Chen |
RTSS | 5 |
| 2022 | BioKnowPrompt: Incorporating imprecise knowledge into prompt-tuning verbalizer with biomedical text for relation extraction
Qing Li 0022, Tao You, Yantao Lu |
Inf. Sci. | 4 |
| 2021 | Weighted Average Precision: Adversarial Example Detection for Visual Perception Of Autonomous VehiclesabstractRecent works have shown that neural networks are vulnerable to carefully crafted adversarial examples (AE). By adding small perturbations to original images, AEs are able to deceive victim models, and result in incorrect outputs. Research work in adversarial machine learning started to focus on the detection of AEs in autonomous driving applications. However, existing studies either use simplifying assumptions on the outputs of object detectors or ignore the tracking system in the perception pipeline. In this paper, we first propose a novel similarity distance metric for object detection outputs in autonomous driving applications. Then, we bridge the gap between the current AE detection research and the real-world autonomous systems by providing a temporal AE detection algorithm, which takes the impact of tracking system into consideration. We perform evaluations on Berkeley Deep Drive and CityScapes datasets, by using different white-box and black-box attacks, which show that our approach outperforms the mean-average-precision and mean intersection over-union based AE detection baselines by significantly increasing the detection accuracy. Weiheng Chai, Yantao Lu, Senem Velipasalar |
ICIP | 2 |
| 2021 | Fabricate-Vanish: An Effective And Transferable Black-Box Adversarial Attack Incorporating Feature DistortionabstractAdversarial examples have emerged as increasingly severe threats for deep neural networks. Recent works have revealed that these malicious samples can transfer across different neural networks, and effectively attack other models. The state-of-the-art methodologies leverage Fast Gradient Sign Method to generate obstructing textures, which can cause neural networks to make incorrect inferences. However, the over-reliance on task-specific loss functions makes the adversarial examples less transferable across networks. Moreover, recent de-noising based adaptive defences provide promising performance against aforementioned attacks. Therefore, to achieve better transferability and attack effectiveness, we propose a novel attack, referred to as the Fabricate-Vanish (FV) attack, which is able to erase benign representations and generate obstruction textures simultaneously. The proposed FV attack treats the adversarial example transferability as latent contribution for each layer of deep neural networks, and maximizes the attack performance by balancing transferability and task specific loss function. Our experimental results on ImageNet show that the proposed FV attack achieves the best attack performance and better transferability by degrading the accuracy of classifiers 3.8% more on average compared to the state-of-the-art attacks. Yantao Lu, Xueying Du, Bingkun Sun, Haining Ren, Senem Velipasalar |
ICIP | 1 |
| 2021 | Hermes Attack: Steal DNN Models with Lossless Inference Accuracy
Yuankun Zhu, Yueqiang Cheng, Husheng Zhou, Yantao Lu |
USENIX Security Symposium | 4 |
| 2020 | Enhancing Cross-Task Black-Box Transferability of Adversarial Examples With Dispersion ReductionabstractNeural networks are known to be vulnerable to carefully crafted adversarial examples, and these malicious samples often transfer, i.e., they remain adversarial even against other models. Although significant effort has been devoted to the transferability across models, surprisingly little attention has been paid to cross-task transferability, which represents the real-world cybercriminal's situation, where an ensemble of different defense/detection mechanisms need to be evaded all at once. We investigate the transferability of adversarial examples across a wide range of real-world computer vision tasks, including image classification, object detection, semantic segmentation, explicit content detection, and text detection. Our proposed attack minimizes the “dispersion” of the internal feature map, overcoming the limitations of existing attacks, that require task-specific loss functions and/or probing a target model. We conduct evaluation on open-source detection and segmentation models, as well as four different computer vision tasks provided by Google Cloud Vision (GCV) APIs. We demonstrate that our approach outperforms existing attacks by degrading performance of multiple CV tasks by a large margin with only modest perturbations. Yantao Lu, Yunhan Jia, Bai Li 0001, Weiheng Chai, Lawrence Carin, Senem Velipasalar |
CVPR | 1 |
| 2020 | Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object Tracking
Yunhan Jia, Yantao Lu, Junjie Shen 0001, Qi Alfred Chen, Hao Chan, Zhenyu Zhong, Tao Wei 0002 |
ICLR | 2 |
| 2019 | Efficient Human Activity Classification from Egocentric Videos Incorporating Actor-Critic Reinforcement LearningabstractIn this paper, we introduce a novel framework to significantly reduce the computational cost of human temporal activity recognition from egocentric videos while maintaining the accuracy at the same level. We propose to apply the actor-critic model of reinforcement learning to optical flow data to locate a bounding box around region of interest, which is then used for clipping a sub-image from a video frame. We also propose to use one shallow and one deeper 3D convolutional neural network to process the original image and the clipped image region, respectively. We compared our proposed method with another approach using 3D convolutional networks on the recently released Dataset of Multimodal Semantic Egocentric Video. Experimental results show that the proposed method reduces the processing time by 36.4% while providing comparable accuracy at the same time. Yantao Lu, Yilan Li, Senem Velipasalar |
ICIP | 1 |
| 2019 | Autonomous Choice of Deep Neural Network Parameters by a Modified Generative Adversarial NetworkabstractThe choice of parameters, and the design of the network architecture are important factors affecting the performance of deep neural networks. However, this task still heavily depends on trial and error, and empirical results. Considering that there are many design and parameter choices, it is very hard to cover every configuration, and find the optimal structure. In this paper, we propose a novel method that autonomously and simultaneously optimizes multiple parameters of any given deep neural network by using a modified generative adversarial network (GAN). In our approach, two different models compete and improve each other progressively. Without loss of generality, the proposed method has been tested with three different neural network architectures, and three very different datasets and applications. The results show that the presented approach can simultaneously and successfully optimize multiple neural network parameters, and achieve increased accuracy in all three scenarios. Yantao Lu, Senem Velipasalar |
ICIP | 1 |
| 2016 | Robust footstep counting and traveled distance calculation by mobile phones incorporating camera geometryabstractMost available approaches for step counting rely on accelerometer data, and thus are prone to over-counting. In addition, most existing devices calculate the traveled distance based on the counted number of steps and a preset stride length. We present a robust and autonomous method for counting steps and tracking and calculating stride length by using accelerometer, gravity sensor and camera data from smart phones. To provide higher precision, instead of using a preset step and/or stride length, the proposed method calculates the distance traveled with each step by using the camera data. If camera is tilted significantly, the angle data obtained from the gravity sensor is used to account for camera geometry and increase the precision of the calculated step length. Experiments are performed with different subjects and the proposed method is compared with accelerometer-based step counter apps. The results show that incorporating camera geometry increases the accuracy, and the proposed method provides the lowest average error rate in number of steps taken and the calculated traveled distance. Yantao Lu, Senem Velipasalar |
ICIP | 1 |
| 2013 | Cluster analysis based on attractor particle swarm optimization with boundary zoomed for working conditions classification of power plant pulverizing system
Hui Cao 0003, Wenquan Chen, Lixin Jia, Yantao Lu |
Neurocomputing | 5 |