EDBT 2026 Demo / reviewers in the wild / expert
Zexin Li 0001
dblp:221/0808-1
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-8758-2151ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic SystemsabstractAgentic AI aims to create systems that set their own goals, adapt proactively to change, and refine behavior through continuous experience. Recent advances suggest that, when facing multiple and unforeseen tasks, agents could benefit from sharing machine-learned knowledge and reusing policies that have already been fully or partially learned by other agents. However, how to query, select, and retrieve policies from a pool of agents, and how to integrate such policies remains a largely unexplored area. This study explores how an agent decides what knowledge to select, from whom, and when and how to integrate it in its own policy in order to accelerate its own learning. The proposed algorithm, Modular Sharing and Composition in Collective Learning (MOSAIC), improves learning in agentic collectives by combining (1) knowledge selection using performance signals and cosine similarity on Wasserstein task embeddings, (2) modular and transferable neural representations via masks, and (3) policy integration, composition and fine-tuning. MOSAIC outperforms isolated learners and global sharing approaches in both learning speed and overall performance, and in some cases solves tasks that isolated agents cannot. The results also demonstrate that selective, goal-driven reuse leads to less susceptibility to task interference. We also observe the emergence of self-organization, where agents solving simpler tasks accelerate the learning of harder ones through shared knowledge. Saptarshi Nath, Christos Peridis, Eseoghene Benjamin, Soheil Kolouri, Peter Kinnell, Zexin Li 0001, Cong Liu 0005, Shirin Dora, Andrea Soltoggio |
AAAI | 7 |
| 2026 | NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsabstractLLM serving is limited by provider-side resources: longer generations consume more GPU time, increase latency, and reduce throughput in multi-tenant systems.This creates a denial-of-service (DoS) risk, where attackers degrade service by inducing excessive generation.Prior work on LLM DoS primarily relies on adversarial perturbations that delay end-of-sequence termination.We show perturbations are often unnecessary: natural, benignlooking instructions that specify impractical and meaningless tasks can already trigger excessive generation.To study this overlooked vulnerability, we introduce NaturalSloth, an adversarial dataset of natural, instruction-based DoS prompts.Starting from a human-curated seed set spanning diverse attack categories, we design a multi-agent synthesis framework to scale the dataset while preserving malicious intent and increasing semantic diversity.Experiments across a wide range of proprietary and open-source LLMs show that NaturalSloth consistently induces excessive generation, with attack effectiveness further amplified when combined with jailbreak techniques.Our analysis also reveals significant limitations of existing defenses, highlighting the need for dedicated protections against natural DoS attacks. 1 Yiming Chen 0010, Zexin Li 0001, Xianghu Yue, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 2 |
| 2025 | Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationabstractSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang, Zhongwei Wan, Yixin He, Dezhi Ran, Tianle Gu, Haizhou Li, Tao Xie, Baishakhi Ray. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yiming Chen 0010, Zexin Li 0001, Zhongwei Wan, Yixin He 0002, Dezhi Ran, Tianle Gu, Haizhou Li 0001, Tao Xie 0001, Baishakhi Ray |
EMNLP | 3 |
| 2025 | Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsabstractModern deployment of large language models (LLMs) frequently involves both inference serving and continuous retraining to stay aligned with evolving data and user feedback. Common practices separate these workloads onto distinct servers in isolated phases, causing substantial inefficiencies (e.g., GPU idleness) and delayed adaptation to new data in distributed settings. Our empirical analysis reveals that these inefficiencies stem from dynamic request arrivals during serving and workload heterogeneity in pipeline-parallel training. To address these challenges, we propose LeMix, a system for co-locating and managing concurrent LLM serving and training workloads. LeMix integrates offline profiling, execution prediction mechanisms, and runtime scheduling to dynamically adapt resource allocation based on workload characteristics and system conditions. By understanding task-specific behaviors and co-execution interference across shared nodes, LeMix improves utilization and serving quality without compromising serving responsiveness. Our evaluation shows that LeMix improves throughput by up to$3.53 \times$, reduces inference loss by up to$0.61 \times$, and delivers up to$2.12 \times$higher response time SLO attainment over traditional separate setups. To our knowledge, this is the first work to uncover and exploit the opportunities of joint LLM inference and training, paving the way for more resource-efficient deployment of LLMs in production environments. Yufei Li 0001, Zexin Li 0001, Yinglun Zhu, Cong Liu 0005 |
RTSS | 2 |
| 2024 | DuoJoule: Accurate On-Device Deep Reinforcement Learning for Energy and TimelinessabstractDeep Reinforcement Learning (DRL) is critical for autonomous systems to continuously learn and adapt in dynamic environments. However, frequent retraining in DRL leads to high energy consumption, posing significant challenges for mobile and battery-dependent robotic systems. Co-optimizing energy, latency, and algorithm performance is essential for efficient on-device DRL. Current approaches either focus on traditional DNNs like CNNs or target only two out of the three dimensions, rather than addressing all three simultaneously. This paper introduces DuoJoule, a comprehensive framework designed to address the unique challenges of DRL workloads by meeting latency deadlines and adhering to energy budgets while maximizing algorithm performance through both application and system-level configurations. DuoJoule dynamically coordinates adjustments in DRL algorithm parameters and system frequency settings using Dynamic Voltage and Frequency Scaling (DVFS). A key innovation of DuoJoule is its runtime metric tracker, which assesses system status against target budgets and calculates a universal efficiency score. This enables rapid and adaptive tuning at runtime, balancing energy efficiency, latency, and algorithm performance. Extensive evaluation using benchmarks along with a realistic autonomous driving case study demonstrates DuoJoule’s versatile cross-platform efficiency, practicality in real-world scenarios, adaptivity to varying constraints, and low runtime overhead evaluated on two widely used autonomous embedded platforms. Empirical results show that DuoJoule consistently meets latency and energy targets while maintaining near-optimal performance, showcasing its effectiveness in managing the complex trade-off space of on-device DRL. Soheil Shirvani, Aritra Samanta, Zexin Li 0001, Cong Liu 0005 |
RTSS | 3 |
| 2024 | BOXR: Body and head motion Optimization framework for eXtended RealityabstractThe emergence of standalone Extended Reality (XR) systems has enhanced user mobility, accommodating both subtle, frequent head motions and substantial, less frequent body motions. However, the pervasively used Motion-to-Display (M2D) latency metric, which measures the delay between the most recent motion and its corresponding display update, only accounts for head motions. This oversight can leave users prone to motion sickness if significant body motion is involved. Although existing methods optimize M2D latency through asynchronous task scheduling and reprojection methods, they introduce challenges like resource contention between tasks and outdated pose data. These challenges are further complicated by user motion dynamics and scene changes during runtime. To address these issues, we for the first time introduce the Camera-to-Display (C2D) latency metric, which captures the delay caused by body motions, and present BOXR, a framework designed to co-optimize both body and head motion delays within an XR system. BOXR enhances the coordination between M2D and C2D latencies by efficiently scheduling tasks to avoid contentions while maintaining an up-to-date pose in the output frame. Moreover, BOXR incorporates a motion-driven visual inertial odometer to adjust to user motion dynamics and employs scene-dependent foveated rendering to manage changes in the scene effectively. Our evaluations show that BOXR significantly outperforms state-of-the-art solutions in 11 EuRoC MAV datasets across 4 XR applications across 3 hardware platforms. In controlled motion and scene settings, BOXR reduces M2D and C2D latencies by up to $63 \%$ and $27 \%$, respectively and increases frame rate by up to $43 \%$. In practical deployments, BOXR achieves substantial reductions in real-world scenarios-up to $42 \%$ in M2D latency and $31 \%$ in C2D latency-while maintaining remarkably low miss rates of only $1.6 \%$ for M2D requirements and $\mathbf{1. 0 \%}$ for C2D requirements. Zexin Li 0001, Hyoseung Kim 0001, Cong Liu 0005 |
RTSS | 2 |
| 2024 | Transferable Adversarial Attacks Against ASRabstractGiven the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a crucial consideration for maintaining their effectiveness in real-time scenarios. Previous explorations into ASR model robustness have predominantly revolved around evaluating accuracy on white-box settings with full access to ASR models. Nevertheless, full ASR model details are often not available in real-world applications. Therefore, evaluating the robustness of black-box ASR models is essential for a comprehensive understanding of ASR model resilience. In this regard, we thoroughly study the vulnerability of practical black-box attacks in cutting-edge ASR models and propose to employ two advanced time-domain-based transferable attacks alongside our differentiable feature extractor. We also propose a speech-aware gradient optimization approach (SAGO) for ASR, which forces mistranscription with minimal impact on human imperceptibility through voice activity detection rule and a speech-aware gradient-oriented optimizer. Our comprehensive experimental results reveal performance enhancements compared to baseline approaches across five models on two databases. Xiaoxue Gao, Zexin Li 0001, Yiming Chen 0010, Cong Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Dynamic Transformers Provide a False Sense of EfficiencyabstractDespite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1 Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 3 |
| 2023 | White-Box Multi-Objective Adversarial Attack on Dialogue GenerationabstractPre-trained transformers are popular in stateof-the-art dialogue generation (DG) systems.Such language models are, however, vulnerable to various adversarial samples as studied in traditional tasks such as text classification, which inspires our curiosity about their robustness in DG systems.One main challenge of attacking DG models is that perturbations on the current sentence can hardly degrade the response accuracy because the unchanged chat histories are also considered for decision-making.Instead of merely pursuing pitfalls of performance metrics such as BLEU, ROUGE, we observe that crafting adversarial samples to force longer generation outputs benefits attack effectiveness-the generated responses are typically irrelevant, lengthy, and repetitive.To this end, we propose a white-box multi-objective attack method called DGSlow.Specifically, DGSlow balances two objectives-generation accuracy and length, via a gradient-based multiobjective optimizer and applies an adaptive searching mechanism to iteratively craft adversarial samples with only a few modifications.Comprehensive experiments 1 on four benchmark datasets demonstrate that DGSlow could significantly degrade state-of-the-art DG models with a higher success rate than traditional accuracy-based methods.Besides, our crafted sentences also exhibit strong transferability in attacking other models. Yufei Li 0001, Zexin Li 0001, Yingfan Gao, Cong Liu 0005 |
ACL (1) | 2 |
| 2023 | Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face RecognitionabstractA hard challenge in developing practical face recognition (FR) attacks is due to the black-box nature of the target FR model, i.e., inaccessible gradient and parameter information to attackers. While recent research took an important step towards attacking black-box FR models through leveraging transferability, their performance is still limited, especially against online commercial FR systems that can be pessimistic (e.g., a less than 50% ASR-attack success rate on average). Motivated by this, we present Sibling-Attack, a new FR attack technique for the first time explores a novel multi-task perspective (i.e., leveraging extra information from multi-correlated tasks to boost attacking transferability). Intuitively, Sibling-Attack selects a set of tasks correlated with FR and picks the Attribute Recognition (AR) task as the task used in Sibling-Attack based on theoretical and quantitative analysis. Sibling-Attack then develops an optimization framework that fuses adversarial gradient information through (1) constraining the cross-task features to be under the same space, (2) a Jointtask meta optimization framework that enhances the gradient compatibility among tasks, and (3) a cross-task gradient stabilization method which mitigates the oscillation effect during attacking. Extensive experiments demonstrate that SiblingAttack outperforms state-of-the-art FR attack techniques by a non-trivial margin, boosting ASR by 12.61% and 55.77% on average on state-of-the-art pre-trained FR models and two well-known, widely used commercial FR systems. Zexin Li 0001, Bangjie Yin, Taiping Yao, Shouhong Ding, Cong Liu 0005 |
CVPR | 1 |
| 2023 | PIMbot: Policy and Incentive Manipulation for Multi-Robot Reinforcement Learning in Social DilemmasabstractRecent research has demonstrated the potential of reinforcement learning (RL) in enabling effective multi-robot collaboration, particularly in social dilemmas where robots face a trade-off between self-interests and collective benefits. However, environmental factors such as miscommunication and adversarial robots can impact cooperation, making it crucial to explore how multi-robot communication can be manipulated to achieve different outcomes. This paper presents a novel approach, namely PIMbot, to manipulating the reward function in multi-robot collaboration through two distinct forms of manipulation: policy and incentive manipulation. Our work introduces a new angle for manipulation in recent multi-agent RL social dilemmas that utilize a unique reward function for incentivization. By utilizing our proposed PIMbot mechanisms, a robot is able to manipulate the social dilemma environment effectively. PIMbot has the potential for both positive and negative impacts on the task outcome, where positive impacts lead to faster convergence to the global optimum and maximized rewards for any chosen robot. Conversely, negative impacts can have a detrimental effect on the overall task performance. We present comprehensive experimental results that demonstrate the effectiveness of our proposed methods in the Gazebo-simulated multi-robot environment. Our work provides insights into how inter-robot communication can be manipulated and has implications for various robotic applications. Shahab Nikkhoo, Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Cong Liu 0005 |
IROS | 2 |
| 2023 | $\mathrm{R}^{3}$: On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsabstractAutonomous robotic systems, like autonomous vehicles and robotic search and rescue, require efficient on-device training for continuous adaptation of Deep Reinforcement Learning (DRL) models in dynamic environments. This research is fundamentally motivated by the need to understand and address the challenges of on-device real-time DRL, which involves balancing timing and algorithm performance under memory constraints, as exposed through our extensive empirical studies. This intricate balance requires co-optimizing two pivotal parameters of DRL training - batch size and replay buffer size. Configuring these parameters significantly affects timing and algorithm performance, while both (unfortunately) require substantial memory allocation to achieve near-optimal performance. This paper presents$\mathbf{R}^{3}$, a holistic solution for managing timing, memory, and algorithm performance in on-device real-time DRL training.$\mathbf{R}^{3}$employs (i) a deadline-driven feedback loop with dynamic batch sizing for optimizing timing, (ii) efficient memory management to reduce memory footprint and allow larger replay buffer sizes, and (iii) a runtime coordinator guided by heuristic analysis and a runtime profiler for dynamically adjusting memory resource reservations. These components collaboratively tackle the trade-offs in on-device DRL training, improving timing and algorithm performance while minimizing the risk of out-of-memory (OOM) errors. We implemented and evaluated$\mathbf{R}^{3}$extensively across various DRL frameworks and benchmarks on three hardware platforms commonly adopted by autonomous robotic systems. Additionally, we integrate$\mathbf{R}^{3}$with a popular realistic autonomous car simulator to demonstrate its real-world applicability. Evaluation results show that$\mathbf{R}^{3}$achieves efficacy across diverse platforms, ensuring consistent latency performance and timing predictability with minimal overhead. Moreover,$\mathbf{R}^{3}$showcases versatility by handling varied optimization goals and adapting to fluctuating systems scenarios. Zexin Li 0001, Aritra Samanta, Yufei Li 0001, Andrea Soltoggio, Hyoseung Kim 0001, Cong Liu 0005 |
RTSS | 1 |
| 2023 | RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsabstractRecent advancements in language models (LMs) have gained substantial attentions on their capability to generate human-like responses. Though exhibiting a promising future for various applications such as conversation AI, these LMs face deployment challenges on various devices due to their extreme computational cost and unpredictable inference latency. Such varied inference latency, identified as a consequence of uncertainty intrinsic to the nature of language, can lead to computational inefficiency and degrade the overall performance of LMs, especially under high-traffic workloads. Unfortunately, the bandwidth of these uncertainty sources is extensive, complicating the prediction of latency and the effects emanating from such uncertainties. To understand and mitigate the impact of uncertainty on real-time response-demanding systems, we take the first step to comprehend, quantify and optimize these uncertainty-induced latency performance variations in LMs. Specifically, we present RT-LM, an uncertainty-aware resource management ecosystem for real-time inference of LMs. RT-LM innovatively quantifies how specific input uncertainties, recognized within the NLP community, adversely affect latency, often leading to an increased output length. Exploiting these insights, we devise a lightweight yet effective method to dynamically correlate input text uncertainties with output length at runtime. Utilizing this quantification as a latency heuristic, we integrate the uncertainty information into a system-level scheduler which explores several uncertainty-induced optimization opportunities, including uncertainty-aware prioritization, dynamic consolidation, and strategic CPU offloading. Quantitative experiments across five state-of-the-art LMs on two hardware platforms demonstrates that RT-LM can significantly reduce the average response time and improve throughput while incurring a rather small runtime overhead. Yufei Li 0001, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005 |
RTSS | 2 |
| 2023 | RED: A Systematic Real-Time Scheduling Approach for Robotic Environmental DynamicsabstractIntelligent robots are designed to effectively navigate dynamic and unpredictable environments laden with moving mechanical elements and objects. Such environment-induced dynamics, including moving obstacles, can readily alter the computational demand (e.g., the creation of new tasks) and the structure of workloads (e.g., precedence constraints among tasks) during runtime, thereby adversely affecting overall system performance. This challenge is amplified when multi-task inference is expected on robots operating under stringent resource and real-time constraints. To address such a challenge, we introduce RED, a systematic real-time scheduling approach designed to support multi-task deep neural network workloads in resource-limited robotic systems. It is designed to adaptively manage the Robotic Environmental Dynamics (RED) while adhering to real-time constraints. At the core of RED lies a deadline-based scheduler that employs an intermediate deadline assignment policy, effectively managing to change workloads and asynchronous inference prompted by complex, unpredictable environments. This scheduling framework also facilitates the flexible deployment of MIMONet (multi-input multi-output neural networks), which are commonly utilized in multi-tasking robotic systems to circumvent memory bottlenecks. Building on this scheduling framework, RED recognizes and leverages a unique characteristic of MIMONet: its weight-shared architecture. To further accommodate and exploit this feature, RED devises a novel and effective workload refinement and reconstruction process. This process ensures the scheduling framework's compatibility with MIMONet and maximizes efficiency. We have implemented RED on several widely used embedded and mobile platforms, including the NVIDIA Jetson Nano, TX2, Xavier, and Orin platforms. We evaluated its performance using workloads that span a broad range of settings typical in navigation robots. The experimental results demonstrate that RED surpasses existing approaches (often by a significant margin) across critical metrics such as throughput, timing correctness, interference robustness, adaptability, and overhead. Zexin Li 0001, Xiaoxi He, Cong Liu 0005 |
RTSS | 1 |
| 2021 | Efficient algorithms for task mapping on heterogeneous CPU/GPU platforms for fast completion time
Zexin Li 0001, Yuqun Zhang, Husheng Zhou, Cong Liu 0005 |
J. Syst. Archit. | 1 |