Yunxuan Li

dblp:204/3762 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GIGAS: Adversarial Attacks on Visual Question Answering With Multi-Modal Generative Models
abstract
VQA models, which answer questions about images by combining both visual and textual information, have been proven susceptible to adversarial attacks. These attacks introduce subtle perturbations to the input data to manipulate the model’s predictions. This paper focuses on adversarial attacks targeting VQA models that follow the “pre-training & fine-tuning” paradigm, an area that remains under-explored. We have identified two key issues in the current field. On one hand, existing multi-modal attacks have low ASR due to inter-modal semantic inconsistency from insufficient cross-modal interaction. On the other hand, the dilemma between attack effectiveness and stealthiness limits the practical applicability of adversarial texts. To address these issues, we propose GIGAS, an innovative attack that uses multi-modal generative models to explore multi-modal interaction through three key modules tailored to solve above-mentioned problems. MIGA aligns adversarial visual features with semantics of misleading images generated by multi-modal generative models to mitigate cross-modal inconsistencies. GSA employs MLLMs to generate natural adversarial texts with greater variation and evaluate similarity to filter based on clean images, balancing effectiveness and stealthiness. Iteration Allocation dynamically adjusts attack iterations based on image-text similarity, maximizing the utility of the limited iterations. Experiments conducted on various VL models and VQA datasets demonstrate superior attack performance, with an average ASR of 89.09% on VQAv2.0. Furthermore, our GIGAS exhibits outstanding transferability, around 60% ASR, across diverse models and specific domains. Our code will be available at: https://github.com/Yvonna-cloud/GIGAS.
Yunxuan Li, Jing Yu 0007, Tieyong Zeng, Liyan Ma
IEEE Trans. Circuits Syst. Video Technol.1
2026 Cooperative Control of Traffic Signals and Vehicle Trajectories Using Multi-Agent Actor-Critic Approach With Vehicle-Road-Cloud Integration
abstract
A new mixed traffic flow involving human-driven vehicles (HDVs) and connected and automated vehicles (CAVs) has emerged as a result of developments in autonomous driving and wireless communication. Deep reinforcement learning (DRL) is a promising method for addressing complex traffic flow control issues. However, the majority of DRL-based studies concentrated on optimizing either traffic signals or vehicle trajectories, often overlooking the interaction between these two elements. In this paper, we propose a cooperative control approach for traffic signals and vehicle trajectories using a multi-agent actor-critic (CCTV-MAC) in a mixed traffic flow environment. Our approach incorporates traffic signal control (TSC) agents that optimize signal cycles and green ratios by extracting trajectory features from HDVs and CAVs, as well as vehicle trajectory control (VTC) agents that optimize CAV trajectories based on information from the self-vehicle, the preceding vehicle, and the approaching signal timing provided by the TSC agents. A transformer-based spatio-temporal attention network and a reward-sharing mechanism are incorporated to enhance coordinated decision-making among adjacent TSC agents. Additionally, a vehicle-road-cloud integration structure (VRCIS) is designed for CCTV-MAC, where model training tasks are deployed on the cloud side to accelerate the learning process via parallel computing formulas and algorithms, while decision-making tasks are executed on the vehicle side and roadside to minimize response latency. With the support of VRCIS, cooperative control of TSC and VTC agents is achieved through effective information exchange. Simulations on real-world road networks show that our proposed CCTV-MAC outperforms traditional and other DRL methods.
Yongnan Zhang, Xiaobin Luo, Yunxuan Li, Sirui Peng, Leipeng Zhu, Yonghua Zhou, Hamido Fujita
IEEE Trans. Intell. Transp. Syst.3
2025 Robust Multi-Objective Preference Alignment with Online DPO
abstract
Multi-objective preference alignment of large language models (LLMs) is critical for developing AI systems that are more configurable, personalizable, helpful, and safe. However, optimizing model outputs to satisfy diverse objectives with variable weights at inference time for truly personalized models presents a significant challenge. Existing approaches are either computationally expensive to train or do not sufficiently steer model behaviors. This paper introduces the Multi-Objective Online DPO (MO-ODPO) algorithm, designed to robustly and efficiently align model behaviors with multiple, potentially conflicting human preferences. Our approach incorporates a prompt conditioning mechanism, allowing us to train a single preference-conditional policy, that can adapt to new preference combinations at inference. Experiments on two popular benchmarks show that MO-ODPO Pareto-dominates existing baselines while providing excellent inference-time steerability between diverse objectives.
Ryan Sullivan, Yunxuan Li, Samrat Phatale, Abhinav Rastogi
AAAI3
2025 GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changes
abstract
Visual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are large viewpoint disparities between the query view and the database, such as attempting localization in a corridor previously build from an opposing direction. Intuitively, this issue can be addressed by synthesizing a set of virtual keyframes that cover all viewpoints. However, existing methods for synthesizing novel views to assist localization often fail to ensure geometric accuracy under large viewpoint changes. In this paper, we introduce a confidence-aware geometric prior into 2D Gaussian splatting to ensure the geometric accuracy of the scene. Then we can render novel views through the mesh with clear structures and accurate geometry, even under significant viewpoint changes, enabling the synthesis of a comprehensive set of virtual keyframes. Incorporating this geometry-preserving virtual keyframe database into the localization pipeline significantly enhances the robustness of visual localization.
Yunxuan Li, Lei Fan 0005, Xiaoying Xing, Jianxiong Zhou, Ying Wu 0001
CVPR1
2025 ReFly: A New Reconfigurable Architecture for LLM Training Based on Optical Circuit Switching
abstract
The rapid development of large language model (LLM) has established distributed training as the dominant paradigm, yet communication efficiency remains a critical challenge. Existing training cluster(e.g., Rail-Optimized architecture) based on electrical packet switch (EPS), suffering from high costs, excessive power consumption, and inefficient bandwidth utilization. While optical circuit switch (OCS) offers a promising alternative with its high bandwidth, low latency, and power efficiency, its rigid connectivity struggles to accommodate dynamic multi-task workloads, and its per-unit cost remains prohibitive at scale. To address these limitations, we propose ReFly, a reconfigurable architecture for LLM training based on OCS. By modeling GPU communication requirements, we design a Cycle Decomposition (CD) scheme for cluster construction and an Alternating Decomposition (AD) algorithm to dynamically schedule multiple OCS. Experimental results demonstrate that ReFly reduces deployment costs by 77% and power consumption by 98% compared to state-of-the-art Rail-Optimized architecture while achieving comparable performance, and a 229% higher bus bandwidth than Fat-Tree. These advancements position ReFly as an efficient and cost-effective solution for next-generation LLM training clusters.
Rentao Gu, Yunxuan Li, Mo Guang, Kaiwen Long, Yuefeng Ji
GLOBECOM3
2025 On The Impact of Different Batch Sizes on Byzantine Robustness in Federated Learning
abstract
In recent years, Byzantine attacks within the federal learning (FL) framework have received a great deal of attention. Existing results indicate that the integration of variance-reduced stochastic optimization algorithms with specifically tailored aggregation methods can enhance the resilience of FL systems to Byzantine attacks. However, most existing studies adopt a fixed batch size for all participating nodes when updating their local parameters, which may not be suitable for real-world applications. This paper mainly focuses on the impact of different batch sizes on Byzantine robustness in the FL system. Specifically, we propose a (δ,c)-Weighed-Agnostic Robust Aggregator ((δ,c)-WARAgg) to enhance the Byzantine robustness of the FL system when different clients employ different batch sizes to update their local parameters. By combining (δ,c)-WARAgg with a novel Byzantine-tolerant method that integrates variance reduction and compression, we develop a Byzantine-robust algorithm, termed Byz-VR-MARINA-DB. Theoretically, we prove that Byz-VR-MARINA-DB exhibit superior performance in the presence of Byzantine attacks compared to their original counterparts. Empirical results corroborate the effectiveness of the proposed aggregation rule and the correctness of our theoretical findings.
Xinjian Huang, Yunxuan Li, Yishuo Zhao, Bo Du 0001
IJCNN3
2025 Unveiling Byzantine-robust with Varied Batch Sizes across Different Clients in Federated Learning
abstract
Due to the lack of an effective auditing mechanism for malicious participants, federated learning (FL) framework faces serious threats of Byzantine attacks. Existing Byzantine resilience methods generally use the same batch size by default. However, this same batch size setting may not be applicable in practice, since clients may comprise a diverse array of devices with different data storage and computational capabilities. This paper mainly studies the robustness of Byzantine across different clients with different batch sizes in FL systems. Specifically, we propose a weighted aggregation framework (WAggF) to enhance the Byzantine robustness of the FL systems in case of different clients with different batch sizes. Combining WAggF with Byzantine attack resilient distributed (Byrd-) algorithms, we develop two Byzantine-robustness algorithms: Byrd-SGD with Different Batch sizes (Byrd-DBSGD) and Byrd-SAGA with Different Batch sizes (Byrd-DBSAGA). Theoretically, we prove that these algorithms exhibit superior performance compared to their original counterparts in the presence of Byzantine attacks. Additionally, we introduce a data partition method to solve the over-centralization problem caused by significantly disparate different batch sizes among different clients. Considering the communication efficiency, we utilize an unbiased compressor to ensure synchronous training of the FL system. To the best of our knowledge, this study represents the first investigation into the impact of different batch sizes across different clients on Byzantine robustness in FL. Extensive experimental results confirm that our proposed algorithms exhibit greater resilience to Byzantine attacks. Codes will be released upon publication.
Yunxuan Li, Xinjian Huang, Bo Du 0001
MMAsia1
2025 An Algorithm for Spatio-Temporal Trajectory Conflict Risk Identification in Intersections Considering Lateral Vehicle Movement
abstract
Vehicle trajectories from different directions within intersections intersect and create conflicts on a two-dimensional plane. Compared to highways, the lateral movement of vehicles within intersections is irregular and difficult to predict. This risk from lateral trajectories extends to the surrounding areas, impacting other vehicles within the intersection and subsequently reducing both safety and efficiency. To address this issue, a framework for identifying the conflict risk of vehicle trajectory based on lateral movement at intersections is proposed in this paper. Firstly, the concept of virtual lanes is employed to extract abnormal trajectories of lateral vehicle movement within intersections. Subsequently, a spatio-temporal hexahedral conflict detection algorithm, based on the vehicle border, is developed. Finally, the causes of abnormal lateral trajectories within intersections and their relationship with conflict events are discussed in detail, and hotspot areas of vehicle safety risks within intersections are identified. To validate the proposed model, typical intersection vehicle trajectory data captured by uncrewed aerial vehicles is utilized. The research results indicate that abnormal lateral movements of vehicles within intersections are common and pose significant risks to neighboring vehicles. The data reveal a specific pattern in the abnormal trajectory and trajectory density at intersections, with abnormal lateral trajectories of left-turning vehicles accounting for as much as 95.02% of cases. Furthermore, abnormal lateral movement in risky trajectories accounted for up to 86.6%. The proposed method in this paper provides support for intersection safety assessments, abnormal lateral trajectory warning systems, and real-time risk calculations for connected autonomous vehicles and Car2Car communication.
Zhonghua Wei, Jingxuan Peng, Lishengsa Yue, Yunxuan Li
IEEE Trans. Intell. Transp. Syst.6
2024 Evidential Active Recognition: Intelligent and Prudent Open-World Embodied Perception
abstract
Active recognition enables robots to intelligently explore novel observations, thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data, wherein appropriate actions are more frequently selected when the recognition is accurate. However, most recognition modules are developed under the closed-world assumption, which makes them ill-equipped to handle unexpected inputs, such as the absence of the target object in the current observation. To address this issue, we propose treating active recognition as a sequential evidence-gathering process, providing by-step uncertainty quantification and reliable prediction under the evidence combination theory. Additionally, the reward function developed in this paper effectively characterizes the merit of actions when operating in open-world environments. To evaluate the performance, we collect a dataset from an indoor simulator, encompassing various recognition challenges such as distance, occlusion levels, and visibility. Through a series of experiments on recognition and robustness analysis, we demonstrate the necessity of introducing uncertainties to active recognition and the superior performance of the proposed method.
Lei Fan 0005, Mingfu Liang, Yunxuan Li, Gang Hua 0001, Ying Wu 0001
CVPR3
2024 LALO - A Virtual Data Lake Zone for Composing Tailor-Made Data Products on Demand
Christoph Stach, Yunxuan Li, Laura Schuiki, Bernhard Mitschang
DEXA (2)2
2024 Enabling Lanuguage Models to Implicitly Learn Self-Improvement
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in open-ended text generation tasks. However, the inherent open-ended nature of these tasks implies that there is always room for improvement in the quality of model responses. To address this challenge, various approaches have been proposed to enhance the performance of LLMs. There has been a growing focus on enabling LLMs to self-improve their response quality, thereby reducing the reliance on extensive human annotation efforts for collecting diverse and high-quality training data. Recently, prompting-based methods have been widely explored among self-improvement methods owing to their effectiveness, efficiency, and convenience. However, those methods usually require explicitly and thoroughly written rubrics as inputs to LLMs. It is expensive and challenging to manually derive and provide all necessary rubrics with a real-world complex goal for improvement (e.g., being more helpfulness and less harmful). To this end, we propose an imPlicit self-ImprovemenT (PIT) framework that implicitly learns the improvement goal from human preference data. PIT only requires preference data that are used to train reward models with no extra human efforts. Specifically, we reformulate the training objective of reinforcement learning from human feedback (RLHF) -- instead of maximizing response quality for a given input, we maximize the quality gap of the response conditioned on a reference response. In this way, PIT is implicitly trained with the improvement goal of better aligning with human preferences. Experiments on two real-world datasets and one synthetic dataset show that our method significantly outperforms prompting-based methods.
Ziqi Wang 0003, Le Hou, Tianjian Lu, Yuexin Wu, Yunxuan Li, Hongkun Yu 0001, Heng Ji 0001
ICLR5
2024 Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language Models
abstract
Sparse Mixture-of-Experts (MoE) is a neural architecture design that adds learnable parameters to Large Language Models (LLMs) without increasing computational complexity (FLOPs). Instruction tuning is a technique for training LLMs to follow instructions. We advocate combining these two approaches, as we find that MoE models benefit more from instruction tuning than dense models. In particular, we conduct empirical studies across three experimental setups: (i) Direct finetuning on individual downstream tasks devoid of instruction tuning; (ii) Instruction tuning followed by in-context few-shot or zero-shot generalization on downstream tasks; and (iii) Instruction tuning supplemented by further finetuning on individual downstream tasks. In the first scenario, MoE models overall underperform dense models of identical computational capacity. This narrative, however, dramatically changes with the introduction of instruction tuning (in the second and third scenarios), used independently or in conjunction with task-specific finetuning. Our most powerful model, FLAN-MoE-32B, surpasses the performance of Flan-PaLM-62B on four benchmark tasks, while using only a third of the FLOPs. The advancements embodied by FLAN-MoE inspire a reevaluation of the design principles of large-scale, high-performance language models in the framework of task-agnostic learning.
Sheng Shen 0001, Le Hou, Yanqi Zhou, Nan Du 0002, Shayne Longpre, Jason Wei, Hyung Won Chung, Barret Zoph, William Fedus, Tu Vu, Yuexin Wu, Wuyang Chen 0001, Albert Webson, Yunxuan Li, Vincent Y. Zhao, Hongkun Yu 0001, Kurt Keutzer, Trevor Darrell, Denny Zhou
ICLR15
2024 PaDS: An adaptive and privacy-enabling Data Pipeline for Smart Cars
abstract
The extensive use of onboard sensors in smart cars enables the collection, processing, and dissemination of large amounts of mobile data containing information about the vehicle, its driver, and even bystanders. Despite the undoubted benefits of such smart cars, this leads to significant privacy concerns. Due to their inherent mobility, the situation of smart cars changes frequently, and with it, the appropriate measures to counteract the exposure of private data. However, data management in such vehicles lacks sufficient support for this privacy dynamism. We therefore introduce PaDS, a framework for Privacy adaptive Data Stream. The focus of this paper is to enable adaptive data processing within the vehicle data stream. With PaDS, Privacy-Enhancing Technologies can be deployed dynamically in the data pipeline of a smart car according to the current situation without user intervention. With a comparison of state-of-the-art approaches, we demonstrate that our solution is very efficient as it does not require a complete restart of the data pipeline. Moreover, compared to a static approach, PaDS causes only minimal overhead despite its dynamic adaptation of the data pipeline to react to changing privacy requirements. This renders PaDS an effective privacy solution for smart cars.
Yunxuan Li, Christoph Stach, Bernhard Mitschang
MDM1
2024 Scaling Instruction-Finetuned Language Models
abstract
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation, RealToxicityPrompts). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PaLM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks (at time of release), such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints,1 which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang 0002, Mostafa Dehghani 0001, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu 0001, Slav Petrov, Ed H. Chi, Jeffrey Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei
J. Mach. Learn. Res.7
2023 A Superposition Assessment Framework of Multi-Source Traffic Risks for Mega-Events Using Risk Field Model and Time-Series Generative Adversarial Networks
abstract
In this study, a novel traffic risk assessment framework of mega-events that integrate risk field and deep learning is proposed. Considering the inherent difference of different traffic risks, the risk quantification and standardization is conducted first. Then several risk field models are constructed to quantify the impacts of multi-source traffic risk superposition on mega-events. Then a time-series generative adversarial networks (TimeGAN) is used to predict the evolution of superposition risk. We select 2022 Beijing Winter Olympics as a case to explore the superposition effects of different traffic risks on the convoy entrance of this mega-event. The results illustrate the superposition risks are significantly associated with the strength of each traffic risk, and the distance from the traffic risk location to the convoy entrance. Furthermore, the temporal evolutions for different traffic risks and their superposition are forecasted using TimeGAN. The results show the unexpected traffic congestion risk presents the highest predictive performance (i.e., the average error for RMSE, MAE, and MSE is 0.135%) and the superposition traffic risks present the lowest predictive performance (the average error is 0.536%). Comparison between different methods demonstrates TimeGAN outperforms other methods in predicting both single traffic risks and superposition risks. The research findings could be potentially referenced in multi-source traffic risk management for mega-events.
Zeyang Cheng, Heng Ding, Yunxuan Li, Haijian Bai
IEEE Trans. Intell. Transp. Syst.4
2022 Unsupervised Depth Completion and Denoising for RGB-D Sensors
abstract
Depth information is considered valuable as it describes geometric structures, which benefits various robotic tasks. However, the depth acquired by RGB-D sensors still suffers from two deficiencies, i.e., incompletion and noises. Previous methods complete depth by exploring hand-tuned models or raising surface assumptions, while nowadays, deep approaches intend to solve this problem with rendered image pairs. For depth denoising, as a consequence of different sensor mechanisms, most methods can only work under specific devices. With existing methods, three challenges emerge: the onerous training set collecting process, the mismatch between existing models and present RGB-D sensors, and the non-real-time computation. In this paper, we first state depth completion and denoising are inherently different and without the need to collect or render complete and noiseless ground truths. We address all mentioned challenges with two separate un-supervised learning procedures. The completion network takes color and incomplete depth as input and predicts values to the unobserved area, which combines prior knowledge and color-depth correlations. The denoising step exploits image sequences to construct noise models in a self-supervised manner with the ability to cater to different sensors. Experimental comparisons and ablation studies demonstrate that even without human-labeled ground truths, the proposed method could produce better completion results and also reduce noises in real-time.
Lei Fan 0005, Yunxuan Li, Ying Wu 0001
ICRA2
2017 Quantitative style analysis of Mo Yan and Zhang Wei's novels
abstract
Selecting 24 novels written by Mo Yan and Zhang Wei as corpus, This paper analyzed the stylistic features of Mo Yan and Zhang Wei's novels from the perspective of quantitative style. Features include the pauses in sentences, the relevance of context, the type/token ratio, the frequency of the word string, high-frequency words and text clustering. Through statistic analysis, it is found that Mo Yan and Zhang Wei's works have much in common, which are both very oral, creative and can use all kinds of linguistic materials. However, they are different from one another in the usage of sentence patterns and of words and in the attention of social life. Compared with Mo Yan's language features, Zhang Wei's is more changeable.
Yunxuan Li, Weiyun Ji, Dekuan Xu
WI1