EDBT 2026 Demo / reviewers in the wild / expert
Xiang Liu 0017
dblp:31/5736-17
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2026
0009-0006-8550-3767ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lemonshark: Asynchronous DAG-BFT With Early Finality
Michael Yiqing Hu, Alvin Hong Yao Yan, Xiang Liu 0017, Jialin Li 0001 |
NSDI | 4 |
| 2026 | MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge DevicesabstractThe rapid development of the Internet of Things (IoT) applications necessitates resource-efficient computing paradigms that can unify heterogeneous sensing modalities. Spiking Neural Networks (SNNs) meet this need with their event-driven and energy-efficient processing nature. However, deploying SNNs on mobile and embedded platforms is hindered by strict and fluctuating memory budgets. While prior work explores lightweight model design and system-level memory management, these methods either sacrifice accuracy or incur high runtime overhead due to timestep-dependent dynamics. To tackle these challenges, we propose a memory-adaptive framework MASI that enables efficient on-device SNN inference by combining (1) a fine-grained memory-adaptive layer slicing strategy, (2) a timestep-agnostic scheduler that maximizes memory utilization with minimal fragmentation, and (3) a timestep-aware early-exit mechanism that reduces redundant calculations. Evaluated on diverse workloads and edge devices, MASI can dynamically adapt to runtime memory availability, approximately reducing memory usage by 20.67% and inference latency by 58.53% on average with negligible accuracy loss compared to other feasible on-device implementations under memory constraints. Di Yu 0001, Helin Zheng, Changze Lv, Xin Du 0002, Linshan Jiang, Xiang Liu 0017, Gang Pan 0001, Shuiguang Deng |
WWW | 6 |
| 2026 | HyStream: A Hybrid System for Application Streaming via Predictive Delivery and Sequence-Linearized CachingabstractTraditional application delivery requires full local installation, introducing persistent security risks from outdated software and imposing significant download delays. While advances in network bandwidth and latency have made remote content delivery more viable, existing dynamic loading mechanisms, such as network filesystems, often remain constrained by performance bottlenecks. Worse still, these solutions degrade sharply under variable or weak connectivity, where untimely code delivery can stall execution altogether. We propose HyStream, a hybrid application streaming system that combines predictive remote delivery with local sequence-linearized caching to sustain responsive and robust execution without requiring installation. HyStream addresses three key challenges: (1) maintaining microsecond-level latency comparable to local storage; (2) bridging the semantic gap between stateless remote storage and stateful execution; and (3) mitigating the limitations of purely network-based solutions under degraded connectivity. To achieve this, HyStream integrates three core components: a dual-mode transmission mechanism that decouples synchronous demand-driven requests from asynchronous speculative prefetching; a thread-aware Markov-chain model that captures fine-grained, concurrent access patterns for accurate prediction; and a sequence-linearized cache that persists streamed blocks in predicted execution order to support deterministic fallback behavior. Together, these components transform irregular, latency-sensitive I/O into efficient structured access that masks network variability. Evaluation shows HyStream delivers near-native performance across diverse networks. On mobile devices, it achieves 16–29% better per-page access latency than local UFS3.1, even over variable Wi-Fi connectivity. On desktops, it typically sustains startup overheads below 30% relative to local NVMe. Under variable and degraded network conditions, the sequence-linearized cache increasingly serves execution-critical accesses, rendering application performance largely insensitive to network latency and jitter within intra-city and inter-city deployments. Sheng Yue 0001, Xiang Liu 0017, Yongjian Fu 0004, Jialin Li 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | PFL-MD: A Privacy-Preserving Federated Learning Framework for Melanoma Diagnosis with Multiple Party Fully Homomorphic EncryptionabstractFederated learning has emerged as a widely adopted distributed machine learning paradigm. In the field of medical diagnosis, it has become a research hotspot, enabling multiple institutions to collaboratively leverage medical data for accurate analysis. However, the distributed nature of federated learning also introduces new challenges in data security and privacy protection, i.e., a curious server might collude with the client to infer private data of honest clients. In this paper, we implement our own FHE library and integrate it with several widely used federated learning methods, providing a unified framework. Our framework employs multiple-party Fully Homomorphic Encryption (FHE) to remove the requirements of a trusted third party or central key servers and address data security and privacy concerns in federated learning, ensuring that the data of honest participants is never exposed while maintaining the accuracy of the final analysis. We deploy our Privacy-preserving Federated Learning framework in the context of Melanoma Diagnosis, PFL-MD, and across multiple types of widely used benchmark datasets, our method achieves accuracy comparable to that of the original federated learning framework, thereby enabling reliable melanoma diagnosis. Extensive experiments further demonstrate the effectiveness of our framework. Liangxi Liu, Jihe Li, Mengyao Zheng, Zegui Jiang, Yijun Song, Xiang Liu 0017 |
BIBM | 7 |
| 2025 | Efficient Partitioning Deep Learning Models for Medical Image Analysis on Iot DevicesabstractDeep learning models, due to their strong capabilities in learning from data and representing features, have been widely deployed, particularly in the medical field. However, given the limitations of medical devices and deployment scenarios, more and more researchers are focusing on how to better deploy deep learning models on resource-constrained edge devices to enable real-time medical data analysis. Nevertheless, such resourceconstrained IoT devices present significant challenges, making it difficult to achieve both high accuracy and low inference latency. To address this problem, we propose a novel framework, RC-DLM, designed to split and execute complex deep learning models on IoT devices. Specifically, we partition a deep learning model into multiple sub-models according to the computational capacity of each device, with each sub-model responsible for handling a subset of classes. To further reduce computation overhead and inference latency, we integrate a class-wise pruning method to shrink the size of each sub-model. Through large-scale experiments conducted on four popular datasets with three model architectures, we demonstrate that our approach significantly reduces inference latency and model size by up to 5.72 times and 57.5 times, respectively. We further deploy our method on real-world edge devices and compare it with state-of-theart approaches, evaluating the three most important aspects: accuracy, inference time, and model size. The comprehensive experimental results of our RC-DLM confirm the effectiveness of our proposed method. Xiang Liu 0017, Mengyao Zheng, Junyong Cao, Dehui Wei, Kang Lai, Huiying Lan, Yijun Song, Xia Li 0005 |
BIBM | 1 |
| 2025 | High Performance Computing Framework for Secure Variable Selection on Genome-Wide Association Studies with Adaptive Vertical Federated LearningabstractVariable selection for genome-wide association studies (GWAS) has long been a central focus in academic research. However, with the advent of the big data era and the rapid growth of biomedical and healthcare data, scientists are increasingly challenged to extract meaningful information from massive datasets. Worse still, such data are often distributed across multiple parties, making collaborative analysis necessary while also requiring strong privacy preservation. To date, there is still no effective framework that can support high-dimensional data analysis, ensure data privacy across collaborators, and simultaneously capture the relatedness between explanatory and response variables. To address these challenges, we introduce the first high-performance computing framework for variable selection in GWAS with Vertical Federated Learning, termed VS-VFL. Our approach leverages Vertical Federated Learning to enable seamless multi-party collaboration while maintaining data privacy and security. Furthermore, we integrate a wide range of state-of-the-art methods, allowing collaborators to apply their preferred techniques, and we explicitly account for the noni.i.d. nature of biomedical data when analyzing the relatedness between explanatory and response variables. In addition, we employ novel optimization strategies and adaptive algorithms to efficiently handle high-dimensional data with sparse features. This framework empowers researchers to conduct comprehensive analyses and perform accurate linkage mapping of gene associations. Our framework is implemented in Python, can be easily deployed on any platform, and is designed to make advanced GWAS analysis accessible to a broader community of researchers. Mengyao Zheng, Xiang Liu 0017, Junyong Cao, Huiying Lan, Liangxi Liu, Xia Li 0005 |
BIBM | 2 |
| 2025 | SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image GenerationabstractLarge Multimodal Models (LMMs) have demonstrated impressive capabilities in multimodal understanding and generation, pushing forward advancements in text-to-image generation. However, achieving accurate text-image alignment for LMMs, particularly in compositional scenarios, remains challenging. Existing approaches, such as layout planning for multi-step generation and learning3 from human feedback or AI feedback, depend heavily on prompt engineering, costly human annotations, and continual upgrading, limiting flexibility and scalability. In this work, we introduce a model-agnostic iterative self-improvement framework (SILMM) that can enable LMMs to provide helpful and scalable self-feedback and optimize text-image alignment via Direct Preference Optimization (DPO). DPO can readily applied to LMMs that use discrete visual tokens as intermediate image representations; while it is less suitable for LMMs with continuous visual features, as obtaining generation probabilities is challenging. To adapt SILMM to LMMs with continuous features, we propose a diversity mechanism to obtain diverse representations and a kernel-based continuous DPO for alignment. Extensive experiments on three compositional text-to-image generation benchmarks validate the effectiveness and superiority of SILMM, showing improvements exceeding 30% on T2I-CompBench++ and around 20% on DPG-Bench. The code is available at https://silmm.github.io/. Leigang Qu, Haochuan Li, Wenjie Wang 0007, Xiang Liu 0017, Juncheng Li 0006, Liqiang Nie, Tat-Seng Chua |
CVPR | 4 |
| 2025 | Efficient Partitioning Vision Transformer on Edge Devices for Distributed InferenceabstractDeep learning models are increasingly utilized on resource-constrained edge devices for real-time data analytics. Recently, Vision Transformer and their variants have shown exceptional performance in various computer vision tasks. However, their substantial computational requirements and low inference latency create significant challenges for deploying such models on resource-constrained edge devices. To address this issue, we propose a novel framework, ED-ViT, which is designed to efficiently split and execute complex Vision Transformers across multiple edge devices. Our approach involves partitioning Vision Transformer models into several sub-models, while each dedicated to handling a specific subset of data classes. To further reduce computational overhead and inference latency, we introduce a class-wise pruning technique that decreases the size of each sub-model. Through extensive experiments conducted on five datasets using three model architectures and actual implementation on edge devices, we demonstrate that our method significantly cuts down inference latency on edge devices and achieves a reduction in model size by up to 28.9 times and 34.1 times, respectively, while maintaining test accuracy comparable to the original Vision Transformer. Additionally, we compare ED-ViT with two state-of-the-art methods that deploy CNN and SNN models on edge devices, evaluating metrics such as accuracy, inference time, and overall model size. Our comprehensive evaluation underscores the effectiveness of the proposed ED-ViT framework. Xiang Liu 0017, Yijun Song, Xia Li 0005, Huiying Lan, Linshan Jiang, Jialin Li 0001 |
ICDCS | 1 |
| 2025 | One-shot Federated Learning Methods: A Practical GuideabstractOne-shot Federated Learning (OFL) is a distributed machine learning paradigm that constrains client-server communication to a single round, addressing privacy and communication overhead issues associated with multiple rounds of data exchange in traditional Federated Learning (FL). OFL demonstrates the practical potential for integration with future approaches that require collaborative training models, such as large language models (LLMs). However, current OFL methods face two major challenges: data heterogeneity and model heterogeneity, which result in subpar performance compared to conventional FL methods. Worse still, despite numerous studies addressing these limitations, a comprehensive summary is still lacking. To address these gaps, this paper presents a systematic analysis of the challenges faced by OFL and thoroughly reviews the current methods. We also offer an innovative categorization method and analyze the trade-offs of various techniques. Additionally, we discuss the most promising future directions and the technologies that should be integrated into the OFL field. This work aims to provide guidance and insights for future research. Xiang Liu 0017, Zhenheng Tang, Xia Li 0005, Yijun Song, Sijie Ji, Bo Han 0003, Linshan Jiang, Jialin Li 0001 |
IJCAI | 1 |
| 2025 | QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNsabstractVideo content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions. Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | High Performance Computing Framework for Variable Selection on Genome-wide Association StudiesabstractVariable selection for genome-wide association studies (GWAS) has been a major research focus for decades. With the exponential growth of biological and biomedical data in the era of big data, scientists are confronted with the challenge of extracting meaningful information from vast datasets while managing the inherent heterogeneity in bioinformatics. To date, there are no highly effective tools that support high-dimensional datasets and achieve robust variable selection performance, all while accounting for the non-i.i.d. features and structured relatedness among explanatory and response variables.To address these challenges, we introduce the first high-performance computing framework for variable selection in GWAS. Our framework integrates various state-of-the-art methods, allowing researchers to easily combine different techniques and fully explore their potential. Additionally, our approach employs novel optimization strategies to solve the problem efficiently, even for high-dimensional data with sparse characteristics. By processing the data holistically, the framework delivers comprehensive analysis and accurate linkage mapping associations. Designed for ease of use, the framework is implemented in Python and offers seamless deployment, making it accessible to a wide range of researchers. Xiang Liu 0017, Jing Diao, Mengyao Zheng, Jihe Li, Dehui Wei, Qipeng Xie, Xia Li 0005, Linshan Jiang |
BIBM | 1 |
| 2024 | Novel Truncated-rank Graph-structured and Tree-guided Sparse Linear Mixed Models for Variable Selection on Genome-wide Association StudiesabstractVariable selection for genome-wide association studies is a key focus for bioinformatics researchers in high-performance computing. The rapid growth of biological and biomedical data demands has led to high-dimensional, heterogeneous datasets characterized by non-i.i.d. properties and numerous response variables, often resulting in false negatives or positives in recovered results. Traditional methods, when nal̈ively applied, yield suboptimal performance due to confounding factors. To account for the complex interdependencies in heterogeneous data and enhance the practical outcomes of genome-wide association studies, we introduce two methods, TGsLMM and TTsLMM, which balance effects between response and explanatory variables for subpopulation inference. Our unified framework performs sparse variable selection using graph-structured or tree-guided structures in a low-rank linear mixed model. Additionally, we extend our approach to high-dimensional datasets and adaptively select the covariance structure for genomic data. Extensive experiments on synthetic and three real-world datasets emphasize the robustness and effectiveness of our proposed methods, achieving the highest ROC area compared to baselines and superior results for future potential. Xiang Liu 0017, Jing Diao, Mengyao Zheng, Jihe Li, Yongyi Xie, Kang Lai, Xiao Geng, Yijun Song, Linshan Jiang |
BIBM | 1 |
| 2024 | FusionFrame: A Fusion Dataflow Scheduling Framework for DNN Accelerators via Analytical Modeling
Liutao Zheng, Huiying Lan, Xiang Liu 0017, Linshan Jiang, Xuehai Zhou |
ICA3PP (6) | 3 |
| 2024 | FedLPA: One-shot Federated Learning with Layer-Wise Posterior AggregationabstractEfficiently aggregating trained neural networks from local clients into a global model on a server is a widely researched topic in federated learning. Recently, motivated by diminishing privacy concerns, mitigating potential attacks, and reducing communication overhead, one-shot federated learning (i.e., limiting client-server communication into a single round) has gained popularity among researchers. However, the one-shot aggregation performances are sensitively affected by the non-identical training data distribution, which exhibits high statistical heterogeneity in some real-world scenarios. To address this issue, we propose a novel one-shot aggregation method with layer-wise posterior aggregation, named FedLPA. FedLPA aggregates local models to obtain a more accurate global model without requiring extra auxiliary datasets or exposing any private label information, e.g., label distributions. To effectively capture the statistics maintained in the biased local datasets in the practical non-IID scenario, we efficiently infer the posteriors of each layer in each local model using layer-wise Laplace approximation and aggregate them to train the global parameters. Extensive experimental results demonstrate that FedLPA significantly improves learning performance over state-of-the-art methods across several metrics. Xiang Liu 0017, Liangxi Liu, Feiyang Ye 0004, Yunheng Shen, Xia Li 0005, Linshan Jiang, Jialin Li 0001 |
NeurIPS | 1 |
| 2018 | Latency-Optimal Task Offloading for Mobile-Edge Computing System in 5G Heterogeneous NetworksabstractMobile edge computing (MEC) is an emerging technology to improve the quality of computation experience for mobile devices. As a promising paradigm to deal with latency-sensitive and computation-intensive tasks, it provides cloud computing capabilities in close proximity to mobile devices in the fifth-generation (5G) networks. As the radio and computational resources are both limited in 5G networks, reducing system latency by task scheduling and resource allocation has gained renewed interests. To minimize the weighted-sum latency of all users in multi-user MEC system, we formulate an optimization problem based on partial offloading strategy. Since the optimization problem is NP-hard, we transform it into a piece-wise convex problem and get the latency-optimal offloading strategy using the sub- gradient method. We further put forward a simplified algorithm which can achieve close-to- optimal performance in linear time. Our proposed strategies are verified by numerical results, which indicate that our algorithms significantly reduce the weighted-sum latency compared with other baseline strategies. Guoxuan Chi, Yumei Wang, Xiang Liu 0017 |
VTC Spring | 3 |
| 2017 | Multiplex confounding factor correction for genomic association mapping with squared sparse linear mixed modelabstractGenome-wide Association Study has presented a promising way to understand the association between human genomes and complex traits. Many simple polymorphic loci have been shown to explain a significant fraction of phenotypic variability. However, challenges remain in the non-triviality of explaining complex traits associated with multifactorial genetic loci, especially considering the confounding factors caused by population structure, family structure, and cryptic relatedness. In this paper, we propose a Squared-LMM (LMM2) model, aiming to jointly correct population and genetic confounding factors. We offer two strategies of utilizing LMM2for association mapping: 1) It serves as an extension of univariate LMM, which could effectively correct population structure, but consider each SNP in isolation. 2) It is integrated with the multivariate regression model to discover association relationship between complex traits and multifactorial genetic loci. We refer to this second model as sparse Squared-LMM (sLMM2). Further, we extend LMM2/sLMM2by raising the power of our squared model to the LMMn/sLMMnmodel. We demonstrate the practical use of our model with synthetic phenotypic variants generated from genetic loci of Arabidopsis Thaliana. The experiment shows that our method achieves a more accurate and significant prediction on the association relationship between traits and loci. We also evaluate our models on collected phenotypes and genotypes with the number of candidate genes that the models could discover. The results suggest the potential and promising usage of our method in genome-wide association studies. Haohan Wang, Xiang Liu 0017, Yunpeng Xiao 0001, Ming Xu 0008, Eric P. Xing |
BIBM | 2 |