VLDB 2026 Research / reviewers in the wild / expert
Yunsheng Ma
dblp:159/8051
· DBLP profile ↗
16ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMsabstractSpurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machine learning problems. Multimodal Large Language Models (MLLMs), which leverage pretrained vision and language models, have recently demonstrated strong capability in joint vision-language understanding. However, both the presence and severity of spurious biases in MLLMs remain poorly understood. In this work, we address this gap by analyzing the spurious biases in the multimodal setting and uncovering the specific inference-time data patterns that can manifest this problem. To support this analysis, we introduce MM-SpuBench, a comprehensive, human-verified benchmark dataset consisting of image-class pairs annotated with core and spurious attributes, grounded in our taxonomy of nine distinct types of spurious correlations. The benchmark is constructed using human-interpretable attribute information to capture a wide range of spurious patterns reflective of real-world knowledge. Leveraging this benchmark, we conduct a comprehensive evaluation of the state-of-the-art open-source and proprietary MLLMs with both standard accuracy and the proposed Conditional Generation Likelihood Advantage (CGLA). Our findings highlight the persistence of reliance on spurious correlations and the difficulty of mitigation on our benchmark. We hope this work can inspire new technical strides to mitigate these biases. Our benchmark is publicly available at https://huggingface.co/datasets/mmbench/MM-SpuBench. Wenqian Ye, Bohan Liu 0008, Guangtao Zheng, Di Wang 0053, Yunsheng Ma, Bolin Lai, James M. Rehg, Aidong Zhang 0001 |
KDD (1) | 6 |
| 2026 | LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future TrendsabstractWith the broader adoption and highly successful development of large language models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language (NL) understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. This article first introduces the novel concept of designing LLMs for autonomous driving (LLM4AD), followed by a review of existing LLM4AD studies. Then, a comprehensive benchmark is proposed for evaluating the instruction-following and reasoning abilities of LLM4AD systems, which includes LaMPilot-Bench, CARLA Leaderboard 1.0 Benchmark in simulation and NuPlanQA for multiview visual question answering (VQA). Furthermore, extensive real-world experiments are conducted on autonomous vehicle platforms, examining both on-cloud and on-edge LLM deployment for personalized decision-making and motion control. Next, the future trends of integrating language diffusion models into autonomous driving are explored, exemplified by the proposed vision-language diffusion (ViLaD) framework. Finally, the main challenges of LLM4AD are discussed, including latency, deployment, security and privacy, safety, trust and transparency, and personalization. Can Cui 0009, Yunsheng Ma, Sungyeon Park 0001, Zichong Yang, Yupeng Zhou, Peiran Liu 0003, Juanwu Lu, Juntong Peng, Jiaru Zhang, Ruqi Zhang, Lingxi Li 0001, Yaobin Chen, Jitesh H. Panchal, Amr Abdelraouf, Kyungtae Han, Ziran Wang |
Proc. IEEE | 2 |
| 2025 | NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
Sungyeon Park 0001, Can Cui 0009, Yunsheng Ma, Ahmadreza Moradipari, Kyungtae Han, Ziran Wang |
ICCV | 3 |
| 2025 | On-Board Vision-Language Models (VLMs) for Personalized Motion Control of Autonomous VehiclesabstractPersonalized driving refers to an autonomous vehicle’s ability to adapt its driving behavior or control strategies to match individual users’ preferences and driving styles while maintaining safety and comfort standards. However, existing works either fail to capture every individual’s preference precisely or become computationally inefficient as the user base expands. Vision-Language Models (VLMs) offer promising solutions to this front through their natural language understanding and scene reasoning capabilities. In this work, we propose a lightweight yet effective on-board VLM framework that provides low-latency personalized driving performance while maintaining strong reasoning capabilities. Our solution incorporates a Retrieval-Augmented Generation (RAG)-based memory module that enables continuous learning of individual driving preferences through human feedback. Through comprehensive real-world vehicle experiments, our system has demonstrated the ability to provide safe, comfortable, and personalized driving experiences across various scenarios and significantly reduce takeover rates by up to 76.9%. To the best of our knowledge, this work represents the first personalized VLM motion control system in real-world autonomous vehicles. The demo video can be watched at https://tinyurl.com/4xsnz79n. Can Cui 0009, Zichong Yang, Yupeng Zhou, Juntong Peng, Sungyeon Park 0001, Yunsheng Ma, Wenqian Ye, Yiheng Feng, Jitesh H. Panchal, Lingxi Li 0001, Yaobin Chen, Ziran Wang |
IROS | 7 |
| 2025 | Video Token Sparsification for Efficient Multimodal LLMs in Driving Visual Question AnsweringabstractMultimodal large language models (MLLMs) have shown significant potential in enhancing driving scene understanding and visual question answering (VQA) through advanced logical reasoning capabilities. These tasks support driving action generation and explanation, especially in end-to-end autonomous driving applications. However, deploying these models poses a significant challenge due to their substantial parameter sizes and computational demands, which often exceed onboard computational limits. A key limitation stems from the large number of visual tokens needed to capture detailed, long-context visual information, resulting in increased latency and memory use. To address this, we propose Video Token Sparsification (VTS), a novel approach that leverages redundancy in consecutive video frames to reduce visual tokens while preserving critical information. VTS employs a lightweight CNN-based model to identify key frames and prune less informative tokens, mitigating hallucinations and boosting inference throughput without performance loss. Comprehensive experiments on the LingoQA and DRAMA benchmarks show that VTS achieves up to a 33% improvement in inference throughput and a 28% reduction in memory usage compared to baselines, maintaining comparable performance. Yunsheng Ma, Amr Abdelraouf, Ahmadreza Moradipari, Ziran Wang, Kyungtae Han |
IV | 1 |
| 2024 | MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene UnderstandingabstractVision-language generative AI has demonstrated re-markable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However, current benchmark datasets lack multi-modal point cloud, image, and language data pairs. Recent approaches utilize visual instruction learning and cross-modal prompt engineering to expand vision-language models into this domain. In this paper, we pro-pose a new vision-language benchmark that can be used to finetune traffic and HD map domain-specific foundation models. Specifically, we annotate and leverage large-scale, broad-coverage traffic and map data extracted from huge HD map annotations, and use CLIP and LLaMA-2 / Vi-cuna to finetune a baseline model with instruction-following data. Our experimental results across various algorithms reveal that while visual instruction-tuning large language models (LLMs) can effectively learn meaningful represen-tations from MAPLM-QA, there remains significant room for further advancements. To facilitate applying LLMs and multi-modal data into self-driving research, we will release our visual-language QA data, and the baseline models at GitHub.com/LLVM-AD/MAPLM. Yunsheng Ma, Wenqian Ye, Can Cui 0009, Zhipeng Cao 0002, Kaizhao Liang, Ziran Wang, James M. Rehg, Chao Zheng 0004 |
CVPR | 3 |
| 2024 | Quantifying Uncertainty in Motion Prediction with Variational Bayesian MixtureabstractSafety and robustness are crucial factors in developing trustworthy autonomous vehicles. One essential aspect of addressing these factors is to equip vehicles with the capability to predict future trajectories for all moving objects in the surroundings and quantify prediction uncertainties. In this paper, we propose the Sequential Neural Variational Agent (SeNe VA), a generative model that describes the distribution of future trajectories for a single moving object. Our approach can distinguish Out-of-Distribution data while quantifying uncertainty and achieving competitive performance compared to state-of-the-art methods on the Argoverse 2 and INTERACTION datasets. Specifically, a 0.446 meters minimum Final Displacement Error, a 0.203 meters minimum Average Displacement Er-ror, and a 5.35% Miss Rate are achieved on the INTERACTION test set. Extensive qualitative and quantitative analy-sis is also provided to evaluate the proposed model. Our open-source code is available at https://github.com/PurdueDigitalTwin/seneva. Juanwu Lu, Can Cui 0009, Yunsheng Ma, Aniket Bera, Ziran Wang |
CVPR | 3 |
| 2024 | LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model ProgramsabstractAutonomous driving (AD) has made significant strides in recent years. However, existing frameworks struggle to interpret and execute spontaneous user instructions, such as "overtake the car ahead.” Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to bridge this gap. In this paper, we present LaMPilot, a novel framework that integrates LLMs into AD systems, enabling them to follow user instructions by generating code that leverages established functional primitives. We also introduce LaMPilot-Bench, the first bench-mark dataset specifically designed to quantitatively evaluate the efficacy of language model programs in AD. Adopting the LaMPilot framework, we conduct extensive experiments to assess the performance of off-the-shelf LLMs on LaMPilot-Bench. Our results demonstrate the potential of LLMs in handling diverse driving scenarios and following user instructions in driving. To facilitate further research in this area, we release our code and data at GitHub.com/PurdueDigitalTwin/LaMPilot. Yunsheng Ma, Can Cui 0009, Wenqian Ye, Peiran Liu 0003, Juanwu Lu, Amr Abdelraouf, Kyungtae Han, Aniket Bera, James M. Rehg, Ziran Wang |
CVPR | 1 |
| 2024 | ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction DetectionabstractEnsuring traffic safety and mitigating accidents in modern driving is of paramount importance, and computer vision technologies have the potential to significantly contribute to this goal. This paper presents a multi-modal Vision Transformer for Driver Distraction Detection (termed ViT-DD), which incorporates inductive information from training signals related to both distraction detection and driver emotion recognition. Additionally, a self-learning algorithm is developed, allowing for the seamless integration of driver data without emotion labels into the multi-task training process of ViT-DD. Experimental results reveal that the proposed ViT-DD surpasses existing state-of-the-art methods for driver distraction detection by 6.5% and 0.9% on the SFDDD and AUCDD datasets, respectively. Yunsheng Ma, Ziran Wang |
IV | 1 |
| 2024 | MACP: Efficient Model Adaptation for Cooperative PerceptionabstractVehicle-to-vehicle (V2V) communications have greatly enhanced the perception capabilities of connected and automated vehicles (CAVs) by enabling information sharing to "see through the occlusions", resulting in significant performance improvements. However, developing and training complex multi-agent perception models from scratch can be expensive and unnecessary when existing single-agent models show remarkable generalization capabilities. In this paper, we propose a new framework termed MACP, which equips a single-agent pre-trained model with cooperation capabilities. We approach this objective by identifying the key challenges of shifting from single-agent to cooperative settings, adapting the model by freezing most of its parameters and adding a few lightweight modules. We demonstrate in our experiments that the proposed framework can effectively utilize cooperative observations and outperform other state-of-the-art approaches in both simulated and real-world cooperative perception benchmarks while requiring substantially fewer tunable parameters with reduced communication costs. Our ource code is available at https://github.com/PurdueDigitalTwin/MACP. Yunsheng Ma, Juanwu Lu, Can Cui 0009, Sicheng Zhao, Wenqian Ye, Ziran Wang |
WACV | 1 |
| 2023 | Human-Autonomy Teaming on Autonomous Vehicles with Large Language Model-Enabled Human Digital TwinsabstractThe development of autonomous vehicles is dramatically reshaping the transportation landscape, bringing new challenges and opportunities in human-machine interaction. As autonomous vehicles evolve, understanding and responding to human intent becomes significant, and therefore require new ways of human-autonomy teaming. A human digital twin (HDT) is a virtual representation of an individual driver, capturing their preferences, behaviors, and physiological states, enabling machines to better understand and predict human needs and responses. In this paper, we explore how large language models (LLMs), like GPT-4 and LLaMA, together with HDTs are changing the way humans team up with autonomous vehicles. These LLMs help make our conversations with vehicles more natural and intuitive. By pairing them in HDTs, we can get real-time feedback and smarter responses. This combination offers not just easier control but also safer driving experiences. We will break down how this works, why it matters, and what we might expect in the future. Can Cui 0009, Yunsheng Ma, Wenqian Ye, Ziran Wang |
SEC | 2 |
| 2023 | Mitigating Transformer Overconfidence via Lipschitz RegularizationabstractThough Transformers have achieved promising results in many computer vision tasks, they tend to be over-confident in predictions, as the standard Dot Product Self-Attention (DPSA) can barely preserve distance for the unbounded input domain. In this work, we fill this gap by proposing a novel Lipschitz Regularized Transformer (LRFormer). Specifically, we present a new similarity function with the distance within Banach Space to ensure the Lipschitzness and also regularize the term by a contractive Lipschitz Bound. The proposed method is analyzed with a theoretical guarantee, providing a rigorous basis for its effectiveness and reliability. Extensive experiments conducted on standard vision benchmarks demonstrate that our method outperforms the state-of-the-art single forward pass approaches in prediction, calibration, and uncertainty estimation. Wenqian Ye, Yunsheng Ma |
UAI | 2 |
| 2023 | RSAL-iMFS: A framework of randomized stacking with active learning for incremental multi-fidelity surrogate modeling
Zongqi Liu, Xueguan Song, Chao Zhang 0017, Yunsheng Ma, Dacheng Tao |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosabstractEmotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and training classifiers. In this paper, we propose to recognize video emotions in an end-to-end manner based on convolutional neural networks (CNNs). Specifically, we develop a deep Visual-Audio Attention Network (VAANet), a novel architecture that integrates spatial, channel-wise, and temporal attentions into a visual 3D CNN and temporal attentions into an audio 2D CNN. Further, we design a special classification loss, i.e. polarity-consistent cross-entropy loss, based on the polarity-emotion hierarchy constraint to guide the attention generation. Extensive experiments conducted on the challenging VideoEmotion-8 and Ekman-6 datasets demonstrate that the proposed VAANet outperforms the state-of-the-art approaches for video emotion recognition. Our source code is released at: https://github.com/maysonma/VAANet. Sicheng Zhao, Yunsheng Ma, Jufeng Yang, Tengfei Xing, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
AAAI | 2 |
| 2018 | A New Deep Learning-Based Food Recognition System for Dietary Assessment on An Edge Computing Service InfrastructureabstractLiterature has indicated that accurate dietary assessment is very important for assessing the effectiveness of weight loss interventions. However, most of the existing dietary assessment methods rely on memory. With the help of pervasive mobile devices and rich cloud services, it is now possible to develop new computer-aided food recognition system for accurate dietary assessment. However, enabling this future Internet of Things-based dietary assessment imposes several fundamental challenges on algorithm development and system design. In this paper, we set to address these issues from the following two aspects: (1) to develop novel deep learning-based visual food recognition algorithms to achieve the best-in-class recognition accuracy; (2) to design a food recognition system employing edge computing-based service computing paradigm to overcome some inherent problems of traditional mobile cloud computing paradigm, such as unacceptable system latency and low battery life of mobile devices. We have conducted extensive experiments with real-world data. Our results have shown that the proposed system achieved three objectives: (1) outperforming existing work in terms of food recognition accuracy; (2) reducing response time that is equivalent to the minimum of the existing approaches; and (3) lowering energy consumption which is close to the minimum of the state-of-the-art. Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma, Songqing Chen |
IEEE Trans. Serv. Comput. | 6 |
| 2016 | DeepFood: Deep Learning-Based Food Image Recognition for Computer-Aided Dietary Assessment
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma |
ICOST | 6 |