Xiangchen Li

dblp:245/7680 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 dLLM-Serve: Bridging the Memory Gap in Diffusion Language Model Serving
abstract
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to Autoregressive Models (ARMs), utilizing parallel decoding to overcome sequential bottlenecks. However, existing research focuses primarily on kernel-level optimizations, lacking a holistic serving framework that addresses the unique memory dynamics of diffusion processes in production. We identify a critical "memory footprint crisis" specific to dLLMs, driven by monolithic logit tensors and the severe resource oscillation between compute-bound "Refresh" phases and bandwidth-bound "Reuse" phases. To bridge this gap, we present dLLM-Serve, an efficient dLLM serving system that co-optimizes memory footprint, computational scheduling, and generation quality. dLLM-Serve introduces Logit-Aware Activation Budgeting to decompose transient tensor peaks, a Phase-Multiplexed Scheduler to interleave heterogeneous request phases, and Head-Centric Sparse Attention to decouple logical sparsity from physical storage. We evaluate dLLM-Serve on diverse workloads (LiveBench, Burst, OSC) and GPUs (RTX 4090, L40S). Relative to the state-of-the-art baseline, dLLM-Serve improves throughput by 1.61 × —1.81 × on the consumer-grade RTX 4090 and 1.60 × —1.74 × on the server-grade NVIDIA L40S, while reducing tail latency by nearly 4 × under heavy contention. dLLM-Serve establishes the first blueprint for scalable dLLM inference, converting theoretical algorithmic sparsity into tangible wall-clock acceleration across heterogeneous hardware.
Jiakun Fan, Yanglin Zhang, Xiangchen Li, Dimitrios S. Nikolopoulos
ICS3
2026 APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
abstract
Deploying large language models (LLMs) for online inference is often constrained by limited GPU memory, particularly due to the growing KV cache during auto-regressive decoding. Hybrid GPU-CPU execution has emerged as a promising solution by offloading KV cache management and parts of attention computation to the CPU. However, a key bottleneck remains: existing schedulers fail to effectively overlap CPU-offloaded tasks with GPU execution during the latency-critical, bandwidth-bound decode phase. This particularly penalizes real-time, decode-heavy applications (e.g., chat, Chain-of-Thought reasoning) which are currently underserved by existing systems, especially under memory pressure typical of edge or low-cost deployments. We present APEX, a novel, profiling-informed scheduling strategy that maximizes CPU-GPU parallelism during hybrid LLM inference. Unlike systems relying on static rules or purely heuristic approaches, APEX dynamically dispatches compute across heterogeneous resources by predicting execution times of CPU and GPU subtasks to maximize overlap while avoiding scheduling overheads. We evaluate APEX on diverse workloads and GPU architectures (NVIDIA T4, A10), using LLaMa-2-7B and LLaMa-3.1-8B models. Compared to GPU-only schedulers like vLLM, APEX improves throughput by 84% - 96% on T4 and 11% - 89% on A10 GPUs, while preserving latency. Against the best existing hybrid schedulers, it delivers up to 72% (T4) and 37% (A10) higher throughput in long-output settings. APEX significantly advances hybrid LLM inference efficiency on such memory-constrained hardware and provides a blueprint for scheduling in heterogeneous AI systems, filling a critical gap for efficient real-time LLM applications.
Jiakun Fan, Yanglin Zhang, Xiangchen Li, Dimitrios S. Nikolopoulos
IPDPS3
2026 Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices
abstract
Real-time multi-label video classification on embedded devices is constrained by limited compute and energy budgets. Yet, video streams exhibit structural properties such as label sparsity, temporal continuity, and label co-occurrence that can be leveraged for more efficient inference. We introduce Polymorph, a context-aware framework that activates a minimal set of lightweight Low Rank Adapters (LoRA) per frame. Each adapter specializes in a subset of classes derived from co-occurrence patterns and is implemented as a LoRA weight over a shared backbone. At runtime, Polymorph dynamically selects and composes only the adapters needed to cover the active labels, avoiding fullmodel switching and weight merging. This modular strategy improves scalability while reducing latency, and energy overhead. Polymorph achieves 40% lower energy consumption and improves mAP by 9 points over strong baselines executing the TAO dataset.
Saeid Ghafouri, Mohsen Fayyaz, Xiangchen Li, Chacko John Deepu, Bo Ji 0001, Dimitrios S. Nikolopoulos, Hans Vandierendonck
WACV3
2025 SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
abstract
The growing gap between the increasing complexity of large language models (LLMs) and the limited computational budgets of edge devices poses a key challenge for efficient on-device inference, despite gradual improvements in hardware capabilities. Existing strategies, such as aggressive quantization, pruning, or remote inference, trade accuracy for efficiency or lead to substantial cost burdens. This position paper introduces a new framework that leverages speculative decoding, previously viewed primarily as a decoding acceleration technique for autoregressive generation of LLMs, as a promising approach specifically adapted for edge computing by orchestrating computation across heterogeneous devices. We propose SLED, a framework that allows lightweight edge devices to draft multiple candidate tokens locally using diverse draft models, while a single, shared edge server verifies the tokens utilizing a more precise target model. To further increase the efficiency of verification, the edge server batches the diverse verification requests from devices. This approach supports heterogeneous devices and reduces server-side memory footprint by sharing a single upstream target model across devices. Our initial experiments with Jetson Orin Nano, Raspberry Pi 4B/5, and an edge server equipped with 4 Nvidia A100 GPUs indicate substantial benefits: ×2.2 higher system throughput, ×2.8 higher system capacity, and better cost efficiency, all without sacrificing model accuracy.
Xiangchen Li, Dimitrios Spatharakis, Saeid Ghafouri, Jiakun Fan, Hans Vandierendonck, Chacko John Deepu, Bo Ji 0001, Dimitrios S. Nikolopoulos
SEC1
2023 TransFlow: a Snakemake workflow for transmission analysis ofMycobacterium tuberculosiswhole-genome sequencing data
abstract
MOTIVATION: Whole-genome sequencing (WGS) is increasingly used to aid the understanding of Mycobacterium tuberculosis (MTB) transmission. The epidemiological analysis of tuberculosis based on the WGS technique requires a diverse collection of bioinformatics tools. Effectively using these analysis tools in a scalable and reproducible way can be challenging, especially for non-experts. RESULTS: Here, we present TransFlow (Transmission Workflow), a user-friendly, fast, efficient and comprehensive WGS-based transmission analysis pipeline. TransFlow combines some state-of-the-art tools to take transmission analysis from raw sequencing data, through quality control, sequence alignment and variant calling, into downstream transmission clustering, transmission network reconstruction and transmission risk factor inference, together with summary statistics and data visualization in a summary report. TransFlow relies on Snakemake and Conda to resolve dependencies among consecutive processing steps and can be easily adapted to any computation environment. AVAILABILITY AND IMPLEMENTATION: TransFlow is free available at https://github.com/cvn001/transflow. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junhang Pan, Xiangchen Li, Mingwu Zhang, Yewei Lu, Yelei Zhu, Kunyang Wu, Zhengwei Liu, Junshun Gao
Bioinform.2
2023 Deep Reinforcement Learning-Based Anti-Jamming Algorithm Using Dual Action Network
abstract
Due to the open nature of wireless communication, malicious electromagnetic jamming has long been a severe threat to the establishment and stability of communication links. To address this anti-jamming problem, a Markov decision process (MDP) with a two-dimensional action space consisting of transmit frequency and power is proposed in this paper, modeling the interaction between a normal communication link and the presence of malicious jammers in a frequency hopping (FH) communication system. Furthermore, we also prove the existence of the deterministic optimal policy of the proposed model theoretically. To obtain a policy for the communication link to avoid being jammed, the Dual Action Network-Based Deep Reinforcement Learning Algorithm, and Action Feedback Mechanism are proposed. The energy consumption and frequency switching overhead are considered and evaluated in both the proposed model and the algorithm. Finally, the proposed model and algorithm are verified not only in a virtual simulation environment but also in the field testing environment. The result suggests that the proposed algorithm is of great practical value for solving anti-jamming problems.
Xiangchen Li, Jienan Chen, Xiang Ling 0002, Tingyong Wu
IEEE Trans. Wirel. Commun.1
2022 PAMS-DP: Building a Unified Open PAMS Human Movement Data Platform
abstract
Research on physical activity can reduce the probability of disease and improve people’s health. Due to the complexity and diversity of human movement, the lack of a unified and standardized human movement data has become a significant bottleneck in human movement research. Building a unified and standardized human movement data platform requires four key technologies: movement modeling, movement coding, movement data organization, and movement data storage. In this paper, we propose a solution to building a large-scale movement data platform by researching these four key technologies and realizing the construction of the PAMS human movement data platform (PAMS-DP). PAMS-DP can integrate massive movement data, storing and automatically analyzing movement relationships and movement properties, publicly providing available movement data files in a unified and standardized format, and enhancing people’s understanding of human movements. Consequently, redundant work associated with collecting and processing massive movement data is reduced.
Mengfei Tang, Jupeng Luo, Xiangchen Li
BIBM4
2022 ByteGraph: A High-Performance Distributed Graph Database in ByteDance
abstract
Most products at ByteDance, e.g., TikTok, Douyin, and Toutiao, naturally generate massive amounts of graph data. To efficiently store, query and update massive graph data is challenging for the broad range of products at ByteDance with various performance requirements. We categorize graph workloads at ByteDance into three types: online analytical, transaction, and serving processing, where each workload has its own characteristics. Existing graph databases have different performance bottlenecks in handling these workloads and none can efficiently handle the scale of graphs at ByteDance. We developed ByteGraph to process these graph workloads with high throughput, low latency and high scalability. There are several key designs in ByteGraph that make it efficient for processing our workloads, including edge-trees to store adjacency lists for high parallelism and low memory usage, adaptive optimizations on thread pools and indexes, and geographic replications to achieve fault tolerance and availability. ByteGraph has been in production use for several years and its performance has shown to be robust for processing a wide range of graph workloads at ByteDance.
Changji Li, Yingqian Hu, Xiangchen Li, Dongqing Han, Huiming Zhu, Xuwei Fu, Tingwei Wu, Hongfei Tan, Hengtian Ding, Mengjin Liu, Kangcheng Wang, Ting Ye, Chenguang Zheng, James Cheng
Proc. VLDB Endow.8
2021 Label Similarity Based Graph Network for Badminton Activity Recognition
Ya Wang 0002, Guowen Pan, Jinwen Ma, Xiangchen Li, Albert Zhong
ICIC (1)4
2019 Automatic Badminton Action Recognition Using CNN with Adaptive Feature Extraction on Sensor Data
Ya Wang 0002, Weichuang Fang, Jinwen Ma, Xiangchen Li, Albert Zhong
ICIC (1)4