VLDB 2026 Research / reviewers in the wild / expert
Tan Li 0002
dblp:93/9279-2
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-6129-4792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Two-Sided Matching for Batch-Aware LLM Request Scheduling in Edge NetworksabstractEdge Large Language Model (LLM) inference offers reduced latency and enhanced privacy compared to cloud-based approaches. Current inference engines utilize batching mechanisms for computational efficiency. However, static batching creates prompt interdependencies, leading to inefficient memory usage and prolonged processing times due to batch heterogeneity. Existing request scheduling solutions optimize single-server batching but cannot coordinate across multiple edge servers or handle resource constraints at edge. In this paper, we propose a batch-aware request scheduling framework formulated as a two-sided matching game, where batch composition affects individual inference performance through peer effects. We design novel utility functions for users and edge servers based on inference quality, memory and computing capability, construct preference lists, and establish stable assignments via Gale-Shapley algorithm. Based on this, beneficial swaps further optimize batch homogeneity while preserving service quality. Extensive simulations demonstrate substantial improvements in service quality, consistently outperforming baselines across varying request volumes and resource constraints. Tan Li 0002, Yanming Gong |
LCN | 1 |
| 2025 | Layered Randomized Quantization for Communication-Efficient and Privacy-Preserving Distributed LearningabstractIn distributed learning systems, ensuring efficient communication and privacy protection are two significant challenges. Although several existing works have attempted to address these challenges simultaneously, they often overlook essential learning-oriented features such as dynamic gradient and communication characteristics. In this paper, we propose a communication-efficient and privacy-preserving distributed SGD algorithm. Our proposed algorithm employs a layered randomized quantizer (LRQ) to reduce communication overhead, which also ensures that quantization errors follow an exact Gaussian distribution, thus achieving client-level differential privacy. We analyze the trade-off between convergence error, communication, and privacy under non-IID data distributions. Besides, we modify the algorithm to be training-adaptive by adjusting the perround privacy budget allocation in response to i) dynamic gradient features and ii) real-time changing communication rounds. Both closed-form solutions are derived by solving the minimization problem of convergence error subject to the privacy budget constraint. Finally, we evaluate the effectiveness of our approach through extensive experiments on various datasets, including MNIST, CIFAR-10, and CIFAR-100, demonstrating its superiority in terms of communication cost, privacy protection, and model performance compared to state-of-the-art methods. Guangfeng Yan, Tan Li 0002, Kui Wu 0001, Linqi Song |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Unified Multi-Modal Diagnostic Framework With Reconstruction Pre-Training and Heterogeneity-Combat TuningabstractMedical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack high-level semantic information. Furthermore, two significant heterogeneity challenges hinder the transfer of pre-trained knowledge to downstream tasks, i.e., the distribution heterogeneity between pre-training data and downstream data, and the modality heterogeneity within downstream data. To address these challenges, we propose a Unified Medical Multi-modal Diagnostic (UMD) framework with tailored pre-training and downstream tuning strategies. Specifically, to enhance the representation abilities of vision and language encoders, we propose the Multi-level Reconstruction Pre-training (MR-Pretrain) strategy, including a feature-level and data-level reconstruction, which guides models to capture the semantic information from masked inputs of different modalities. Moreover, to tackle two kinds of heterogeneities during the downstream tuning, we present the heterogeneity-combat downstream tuning strategy, which consists of a Task-oriented Distribution Calibration (TD-Calib) and a Gradient-guided Modality Coordination (GM-Coord). In particular, TD-Calib fine-tunes the pre-trained model regarding the distribution of downstream datasets, and GM-Coord adjusts the gradient weights according to the dynamic optimization status of different modalities. Extensive experiments on five public medical datasets demonstrate the effectiveness of our UMD framework, which remarkably outperforms existing approaches on three kinds of downstream tasks. Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma GradingabstractRecently, multimodal deep learning, which integrates histopathology slides and molecular biomarkers, has achieved a promising performance in glioma grading. Despite great progress, due to the intra-modality complexity and intermodality heterogeneity, existing studies suffer from inadequate histopathology representation learning and inefficient molecular-pathology knowledge alignment. These two issues hinder existing methods to precisely interpret diagnostic molecular-pathology features, thereby limiting their grading performance. Moreover, the real-world applicability of existing multimodal approaches is significantly restricted as molecular biomarkers are not always available during clinical deployment. To address these problems, we introduce a novel Focus on Focus (FoF) framework with paired pathology-genomic training and applicable pathology-only inference, enhancing molecular-pathology representation effectively. Specifically, we propose a Focus-oriented Representation Learning (FRL) module to encourage the model to identify regions positively or negatively related to glioma grading and guide it to focus on the diagnostic areas with a consistency constraint. To effectively link the molecular biomarkers to morphological features, we propose a Multi-view Cross-modal Alignment (MCA) module that projects histopathology representations into molecular subspaces, aligning morphological features with corresponding molecular biomarker status by supervised contrastive learning. Experiments on the TCGA GBMLGG dataset demonstrate that our FoF framework significantly improves the glioma grading. Remarkably, our FoF achieves superior performance using only histopathology slides compared to existing multimodal methods. The source code is available at https://github.com/peterlipan/FoF. Li Pan 0004, Qiushi Yang, Tan Li 0002, Xiaohan Xing, Maximus C. F. Yeung, Zhen Chen 0013 |
BIBM | 4 |
| 2024 | Communication-Efficient Multi-Modal Federated Learning via Dynamic Client-Modality MatchingabstractMulti-modal federated learning (MFL) offers the advantage of aggregating models from diverse data modalities to obtain a more powerful fused model while preserving data privacy. However, MFL faces three key challenges: 1) Communication overhead - only a limited number of clients can participate in training due to communication budget constraints; 2) Modality heterogeneity - different modalities contribute unequally to the fused model; 3) Client heterogeneity - clients exhibit variations in data quantity and quality across modalities. To address these challenges, we formulate a joint client-modality selection problem under communication budget constraints. The goal is to determine the participating clients and their uploaded modalities in each communication round, maximizing the performance of the fused model given a limited communication budget. We propose a dynamic many-to-many matching algorithm with two quota budgeting strategies: 1) Round-aware Modality Budgeting (RMB) determines the total number of uploaded modality models per round based on the current training process (i.e., how close the model is to convergence). 2) Modality-aware Client Allocation (MCB) adaptively allocates client quota for each modality by balancing the modality’s contribution to the fusion model against its model size. After quota budgeting, we construct preference lists for clients and modalities to find a stable many-to-many matching of (client, modality) pairs. Experiments demonstrate that our algorithm achieves better model performance than baselines under the same communication budget, validating the benefits of dynamic budget allocation and client scheduling. Tan Li 0002, Yanming Gong, Hai Liu 0001, Zhen Chen 0013, Linqi Song |
IEEE Big Data | 1 |
| 2024 | Truncated Non-Uniform Quantization for Distributed SGDabstractTo address the communication bottleneck challenge in distributed learning, our work introduces a novel two-stage quantization strategy designed to enhance the communication efficiency of distributed Stochastic Gradient Descent (SGD). The proposed method initially employs truncation to mitigate the impact of long-tail noise, followed by a non-uniform quantization of the post-truncation gradients based on their statistical characteristics. We provide a comprehensive convergence analysis of the quantized distributed SGD, establishing theoretical guarantees for its performance. Furthermore, by minimizing the convergence error, we derive optimal closed-form solutions for the truncation threshold and non-uniform quantization levels under given communication constraints. Both theoretical insights and extensive experimental evaluations demonstrate that our proposed algorithm outperforms existing quantization schemes, striking a superior balance between communication efficiency and convergence performance. Guangfeng Yan, Tan Li 0002, Yuanzhang Xiao, Hanxu Hou, Congduan Li, Linqi Song |
ITW | 2 |
| 2024 | Personalized Federated Deep Reinforcement Learning for Heterogeneous Edge Content Caching Networks
Tan Li 0002, Hai Liu 0001, Tse-Tin Chan |
WiOpt | 2 |
| 2023 | Combat Long-Tails in Medical Classification with Relation-Aware Consistency and Virtual Features Compensation
Li Pan 0004, Qiushi Yang, Tan Li 0002, Zhen Chen 0013 |
MICCAI (6) | 4 |
| 2022 | Federated Adaptive Bandits Aided Caching for Heterogeneous Edge Servers with UncertaintyabstractCaching popular content at edge servers is a promising way to achieve high quality of experience for the wireless edge network. However, the design of an effective cache placement scheme still faces two key challenges: 1) content popularity profiles may be unknown in advance. Some online learning techniques can be incorporated to tackle this uncertainty. 2) content popularity profiles may be heterogeneous among different edge servers. Naively utilizing feedback collected by others may substantially hurt the local estimation if there are large disparities between the popularities. Therefore, an adptive information aggregation protocol is needed. In this paper, we first formulate the caching problem as a multi-agent multi-play bandits problem with heterogeneous reward distributions. We then propose a federated adaptive online learning algorithm. Specifically, each MES employs a model mixture technique to aggregate local user feedback and the knowledge captured by the central server. Our theoretical results show that the upper bound on the cache hit loss (e.g., regret) depends on the heterogeneity and information sharing across MESs. The simulation results demonstrate the effectiveness of our method against baseline schemes on both regrets and cache hit rate. Tan Li 0002, Linqi Song |
WCNC | 1 |
| 2022 | Privacy-Preserving Communication-Efficient Federated Multi-Armed BanditsabstractCommunication bottleneck and data privacy are two critical concerns in federated multi-armed bandit (MAB) problems, such as situations in decision-making and recommendations of connected vehicles via wireless. In this paper, we design the privacy-preserving communication-efficient algorithm in such problems and study the interactions among privacy, communication and learning performance in terms of the regret. To be specific, we design privacy-preserving learning algorithms and communication protocols and derive the learning regret when networked private agents are performing online bandit learning in a master-worker, a decentralized and a hybrid structure. Our bandit learning algorithms are based on epoch-wise sub-optimal arm eliminations at each agent and agents exchange learning knowledge with the server/each other at the end of each epoch. Furthermore, we adopt the differential privacy (DP) approach to protect the data privacy at each agent when exchanging information; and we curtail communication costs by making less frequent communications with fewer agents participation. By analyzing the regret of our proposed algorithmic framework in the master-worker, decentralized and hybrid structures, we theoretically show trade-offs between regret and communication costs/privacy. Finally, we empirically show these trade-offs which are consistent with our theoretical analysis. Tan Li 0002, Linqi Song |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | AC-SGD: Adaptively Compressed SGD for Communication-Efficient Distributed LearningabstractGradient compression (e.g., gradient quantization and gradient sparsification) is a core technique in reducing communication costs in distributed learning systems. The recent trend of gradient compression is to use a varying number of bits across iterations, however, relying on empirical observations or engineering heuristics without a systematic treatment and analysis. To the best of our knowledge, a general dynamic gradient compression that leverages both quantization and sparsification techniques is still far from understanding. This paper proposes a novel Adaptively-Compressed Stochastic Gradient Descent (AC-SGD) strategy to adjust the number of quantization bits and the sparsification size with respect to the norm of gradients, the communication budget, and the remaining number of iterations. In particular, we derive an upper bound, tight in some cases, of the convergence error for arbitrary dynamic compression strategy. Then we consider communication budget constraints and propose an optimization formulation - denoted as theAdaptive Compression Problem (ACP)- for minimizing the deep model’s convergence error under such constraints. By solving the ACP, we obtain an enhanced compression algorithm that significantly improves model accuracy under given communication budget constraints. Finally, through extensive experiments on computer vision and natural language processing tasks on MNIST, CIFAR-10, CIFAR-100 and AG-News datasets, respectively, we demonstrate that our compression scheme significantly outperforms the state-of-the-art gradient compression methods in terms of mitigating communication costs. Guangfeng Yan, Tan Li 0002, Shao-Lun Huang, Tian Lan 0001, Linqi Song |
IEEE J. Sel. Areas Commun. | 2 |
| 2020 | Federated Recommendation System via Differential PrivacyabstractIn this paper we are interested in what we term the federated private bandits framework, that combines differential privacy with multi-agent bandit learning. We explore how differential privacy based Upper Confidence Bound (UCB) methods can be applied to multi-agent environments, and in particular to federated learning environments both in ‘master-worker’ and ‘fully decentralized’ settings. We provide theoretical analysis on the privacy and regret performance of the proposed methods and explore the tradeoffs between these two. Tan Li 0002, Linqi Song, Christina Fragouli |
ISIT | 1 |