VLDB 2026 Research / reviewers in the wild / expert
Fan Lai 0001
dblp:179/2330
· DBLP profile ↗
23ranked-venue papers
5as first author
19since 2021 · last 2026
0009-0005-0472-107XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JITServe: SLO-aware LLM Serving with Imprecise Request Information
Wei Zhang 0044, Zhiyu Wu, Yi Mu 0005, Rui Ning, Banruo Liu, Nikhil Sarda, Myungjin Lee, Fan Lai 0001 |
NSDI | 8 |
| 2026 | EMA: Efficient Model Adaptation for Learning-based SystemsabstractMachine learning (ML) is increasingly applied to optimize system performance in tasks such as resource management and network simulation. Unlike traditional ML tasks (e.g., image classification), networked systems often operate in heterogeneous, long-running, and dynamic environment states, where input conditions (e.g., network loads) and operational objectives can shift over time and across settings. Existing learning-based systems offer little support for adaptation, resulting in costly model training, extensive data collection, degraded system performance, and slow responsiveness. Daiyang Yu, Yaqi Qiao, Fan Lai 0001 |
SIGCOMM | 6 |
| 2026 | Compass: SLO-aware Query Planner for Compound AI Serving at Scale
Banruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang, Fan Lai 0001 |
Proc. VLDB Endow. | 5 |
| 2026 | FIITED: Fine-Grained Embedding Dimension Optimization During Training for Recommender SystemsabstractHuge embedding tables in modern deep learning recommender models (DLRM) require prohibitively large memory during training and inference. This paper proposes FIITED, a system to automatically reduce the memory footprint via FIne-grained In-Training Embedding Dimension pruning. By leveraging the key insight that embedding vectors are not equally important, FIITED adaptively adjusts the dimension of each individual embedding vector during model training, assigning larger dimensions to more important embeddings while adapting to dynamic changes in data. We prioritize embedding dimensions with higher frequencies and gradients as more important. To enable efficient pruning of embeddings and their dimensions during model training, we propose an embedding storage system based on virtually-hashed physically-indexed hash tables. Experiments on two industry models and months of realistic datasets show that FIITED can reduce DLRM embedding size by more than 65’ while preserving model quality, outperforming state-of-the-art in-training embedding pruning methods and increasing the reduction ratio by 1.3× to 1.67×. On public datasets, FIITED can reduce the size of embedding tables by 2.2× to 800× with negligible accuracy drop, achieving 1.05× to 12.5× improvement in reduction ratio compared to baselines while improving model throughput. Qinyi Luo, Penghan Wang, Wei Zhang 0044, Fan Lai 0001, Jiachen Mao, Xiaohan Wei, Wei-Yu Tsai, Yuxi Hu 0001, Xuehai Qian |
IEEE Trans. Computers | 4 |
| 2025 | NetZIP: Algorithm/Hardware Co-design of In-network Lossless Compression for Distributed Large Model TrainingabstractIn distributed large model training, the long communication time required to exchange large volumes of gradients and activations among GPUs dominates the training time.To reduce the communication times, lossy or lossless compression of gradients and/or activations can be employed.However, lossy compression of gradients and activations may demand more training iterations to achieve the same model accuracy and cause convergence failure, respectively.Lossless compression, on the other hand, may not reduce the volumes of gradients and activations enough to offset the significant latency associated with compression and decompression on current platforms.To address these challenges, we propose NetZIP, an algorithm/hardware co-design for in-network lossless compression of both gradients and activations.NetZIP consists of two components.(1) NetZIP-algorithm transforms gradients and activations at the bit and value levels to help lightweight standard lossless compression achieve more compression of the gradients and activations.(2) NetZIP-accelerator integrates Net-ZIP-algorithm with a lightweight lossless compression accelerator within a NIC in a bump-in-the-wire fashion to reduce the compression/decompression latency under the resource constraints.NetZIP-algorithm compresses gradients and activations 40-63 and 43-75 percentage points more, respectively, than heavy standard lossless compression for Llama-3 70B, GPT-3 175B, and Llama-3 405B.NetZIP-accelerator, implemented within FPGA-NICs and connected to commodity servers, provides orders of magnitude lower Jinghan Huang 0001, Hyungyo Kim, Nachuan Wang, Jaeyoung Kang 0004, Hrishi Shah, Minjia Zhang, Fan Lai 0001, Nam Sung Kim |
MICRO | 8 |
| 2025 | Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsabstractLarge language models (LLMs) have transformed natural language processing, but their reliable deployment requires effective uncertainty quantification (UQ). Existing UQ methods are often heuristic and lack a fully probabilistic foundation. This paper begins by providing a theoretical justification for the role of perturbations in UQ for LLMs. We then introduce a dual random walk perspective, modeling input–output pairs as two Markov chains with transition probabilities defined by semantic similarity. Building on this, we propose a fully probabilistic framework based on an inverse model, which quantifies uncertainty by evaluating the diversity of the input space conditioned on a given output through systematic perturbations. Within this framework, we define a new uncertainty measure, Inv-Entropy. A key strength of our framework is its flexibility: it supports various definitions of uncertainty measures, embeddings, perturbation strategies, and similarity metrics. We also propose GAAP, a perturbation algorithm based on genetic algorithms, which enhances the diversity of sampled inputs. In addition, we introduce a new evaluation metric, Temperature Sensitivity of Uncertainty (TSU), which directly assesses uncertainty without relying on correctness as a proxy. Extensive experiments demonstrate that Inv-Entropy outperforms existing semantic UQ methods. Haoyi Song, Ruihan Ji, Naichen Shi, Fan Lai 0001, Raed Kontar |
NeurIPS | 4 |
| 2025 | HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-locationabstractLarge language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatbots to throughput-oriented offline workloads like data synthesis. The existing deployment model, which dedicates machines to each workload, simplifies SLO management but often leads to poor resource utilization. This paper introduces HyGen, an interference-aware LLM serving system that enables efficient co-location of online and offline workloads while preserving SLOs. HyGen incorporates two key innovations: (1) performance control mechanisms, including a latency predictor to estimate batch execution time and an SLO-aware profiler to quantify latency interference, and (2) SLO-aware offline scheduling policies that maximize serving throughput and prevent starvation. Our evaluation on production workloads shows that HyGen achieves up to 3.9-5.8× throughput gains over online and hybrid serving baselines, while ensuring latency SLOs. The code of HyGen is publicly available at https://github.com/UIUC-MLSys/HyGen. Ting Sun 0004, Penghan Wang, Fan Lai 0001 |
NeurIPS | 3 |
| 2025 | Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective RolloutsabstractReinforcement learning, such as PPO and GRPO, has powered recent breakthroughs in LLM reasoning. Scaling rollout to sample more prompts enables models to selectively use higher-quality data for training, which can stabilize RL training and improve model performance, but at the cost of significant computational overhead. In this paper, we first show that a substantial portion of this overhead can be avoided by skipping uninformative prompts before rollout. Our analysis of reward dynamics reveals a strong temporal consistency in prompt value: prompts that are uninformative in one epoch of training are likely to remain uninformative in near future epochs. Based on these insights, we propose GRESO (GRPO with Efficient Selective Rollout), an online, lightweight pre-rollout filtering algorithm that predicts and skips uninformative prompts using reward training dynamics. By evaluating GRESO on a broad range of math reasoning benchmarks and models, like Qwen2.5-Math-1.5B, DeepSeek-R1-Distill-Qwen-1.5B, Qwen2.5-Math-7B, Qwen2.5-14B, and Qwen2.5-32B, we show that GRESO achieves up to 2.4x wall-clock time speedup in rollout and up to 2.0x speedup in total training time without accuracy degradation. We make our code publicly available at https://github.com/Infini-AI-Lab/GRESO/. Haizhong Zheng, Brian R. Bartoldson, Bhavya Kailkhura, Fan Lai 0001, Beidi Chen |
NeurIPS | 5 |
| 2025 | IC-Cache: Efficient Large Language Model Serving via In-context CachingabstractLarge language models (LLMs) have excelled in various applications, yet serving them at scale is challenging due to their substantial resource demands and high latency. Our real-world studies reveal that over 70% of user requests to LLMs have semantically similar counterparts, suggesting the potential for knowledge transfer among requests. However, naively caching and reusing past responses leads to a big quality drop. Yu Gan 0002, Nikhil Sarda, Lillian Tsai, Yanqi Zhou, Arvind Krishnamurthy, Fan Lai 0001, Henry M. Levy, David E. Culler |
SOSP | 8 |
| 2025 | Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed DataabstractPrivacy-preserving machine learning (PPML) algorithms use secure computation protocols to allow multiple data parties to collaboratively train machine learning (ML) models while maintaining their data confidentiality. However, current PPML frameworks couple secure protocols with ML models in PPML algorithm implementations, making it challenging for non-experts to develop and optimize PPML applications, limiting their accessibility and performance. We propose Sequoia, a novel PPML framework that decouples ML models and secure protocols to optimize the development and execution of PPML applications across data parties. Sequoia offers JAX-compatible APIs for users to program their ML models, while using a compiler-executor architecture to automatically apply PPML algorithms and system optimizations for model execution over distributed data. The compiler in Sequoia incorporates cross-party PPML processes into user-defined ML models by transparently adding computation, encryption, and communication steps with extensible policies, and the executor efficiently schedules code execution across multiple data parties, considering data dependencies and device heterogeneity. Compared to existing PPML frameworks, Sequoia requires 64%-92% fewer lines of code for users to implement the same PPML algorithms, and achieves 88% speedup of training throughput in horizontal PPML. Kaiqiang Xu, Di Chai, Junxue Zhang 0001, Fan Lai 0001, Kai Chen 0005 |
Proc. ACM Manag. Data | 4 |
| 2024 | Learn To be Efficient: Build Structured Sparsity in Large Language ModelsabstractLarge Language Models (LLMs) have achieved remarkable success with their billion-level parameters, yet they incur high inference overheads. The emergence of activation sparsity in LLMs provides a natural approach to reduce this cost by involving only parts of the parameters for inference. However, existing methods only focus on utilizing this naturally formed activation sparsity in a post-training setting, overlooking the potential for further amplifying this inherent sparsity. In this paper, we hypothesize that LLMs can learn to be efficient by achieving more structured activation sparsity. To achieve this, we introduce a novel training algorithm, Learn-To-be-Efficient (LTE), designed to train efficiency-aware LLMs to learn to activate fewer neurons and achieve a better trade-off between sparsity and performance. Furthermore, unlike SOTA MoEfication methods, which mainly focus on ReLU-based models, LTE can also be applied to LLMs like LLaMA using non-ReLU activations. Extensive evaluation on language understanding, language generation, and instruction tuning tasks show that LTE consistently outperforms SOTA baselines. Along with our hardware-aware custom kernel implementation, LTE reduces LLaMA2-7B inference latency by 25% at 50% sparsity. Haizhong Zheng, Xiaoyan Bai, Xueshen Liu, Z. Morley Mao, Beidi Chen, Fan Lai 0001, Atul Prakash 0001 |
NeurIPS | 6 |
| 2024 | Fed-ensemble: Ensemble Models in Federated Learning for Improved Generalization and Uncertainty QuantificationabstractThe increase in the computational power of edge devices has opened up the possibility of processing some of the data at the edge and distributing model learning. This paradigm is often called federated learning (FL), where edge devices exploit their local computational resources to train models collaboratively. Though FL has seen recent success, it is unclear how to characterize uncertainties in FL predictions. In this paper, we proposeFed-ensemble: a simple approach that brings model ensembling to FL. Instead of aggregating local models to update a single global model,Fed-ensembleuses random permutations to update a group of$K$models and then obtains predictions through model averaging.Fed-ensemblecan be readily utilized within established FL methods and does not impose a computational overhead compared with single-model methods. Empirical results show that our model has superior performance over several FL algorithms on a wide range of data sets and excels in heterogeneous settings often encountered in FL applications. Also, by carefully choosing client-dependent weights in the inference stage,Fed-ensemblebecomes personalized and yields even better performance. Theoretically, we show that predictions on new data from all$K$models belong to the same predictive posterior distribution under a neural tangent kernel regime. This result, in turn, sheds light on the generalization advantages of model averaging and justifies the uncertainty quantification capability. We also illustrate thatFed-ensemblehas an elegant Bayesian interpretation.Note to Practitioners—provides an algorithm that extracts a set of$K$solutions without imposing any additional communication overhead in FL. Given multiple solutions,Fed-ensemblecan be exploited to personalize inference as well as quantify uncertainty. Such capabilities may be beneficial within multiple practical systems that require uncertainty-aware decision-making. Further,Fed-ensemblemay be useful for model validation and hypothesis testing. Naichen Shi, Fan Lai 0001, Raed Kontar, Mosharaf Chowdhury |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | Auxo: Efficient Federated Learning via Scalable Client ClusteringabstractFederated learning (FL) is an emerging machine learning (ML) paradigm that enables heterogeneous edge devices to collaboratively train ML models without revealing their raw data to a logically centralized server. However, beyond the heterogeneous device capacity, FL participants often exhibit differences in their data distributions, which are not independent and identically distributed (Non-IID). Many existing works present point solutions to address issues like slow convergence, low final accuracy, and bias in FL, all stemming from client heterogeneity. Fan Lai 0001, Yinwei Dai, Aditya Akella, Harsha V. Madhyastha, Mosharaf Chowdhury |
SoCC | 2 |
| 2023 | Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingabstractTraining deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledge-guided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy. Decang Sun, Kai Chen 0005, Fan Lai 0001, Mosharaf Chowdhury |
EuroSys | 4 |
| 2023 | Coverage-centric Coreset Selection for High Pruning Rates
Haizhong Zheng, Fan Lai 0001, Atul Prakash 0001 |
ICLR | 3 |
| 2023 | ModelKeeper: Accelerating DNN Training via Automated Training Warmup
Fan Lai 0001, Yinwei Dai, Harsha V. Madhyastha, Mosharaf Chowdhury |
NSDI | 1 |
| 2023 | AdaEmbed: Adaptive Embedding for Large-Scale Recommendation Models
Fan Lai 0001, Wei Zhang 0044, William Tsai, Xiaohan Wei, Yuxi Hu 0001, Sabin Devkota, Jongsoo Park, Zeliang Chen, Ellie Wen, Paul Rivera, Chun-cheng Jason Chen, Mosharaf Chowdhury |
OSDI | 1 |
| 2022 | FedScale: Benchmarking Model and System Performance of Federated Learning at ScaleabstractWe present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging from image classification and object detection to language modeling and speech recognition. Each dataset comes with a unified evaluation protocol using real-world data splits and evaluation metrics. To reproduce realistic FL behavior, FedScale contains a scalable and extensible runtime. It provides high-level APIs to implement FL algorithms, deploy them at scale across diverse hardware and software backends, and evaluate them at scale, all with minimal developer efforts. We combine the two to perform systematic benchmarking experiments and highlight potential opportunities for heterogeneity-aware co-optimizations in FL. FedScale is open-source and actively maintained by contributors from different institutions at http://fedscale.ai. We welcome feedback and contributions from the community. Fan Lai 0001, Yinwei Dai, Sanjay Sri Vallabh Singapuram, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
ICML | 1 |
| 2021 | Oort: Efficient Federated Learning via Guided Participant Selection
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
OSDI | 1 |
| 2020 | Sol: Fast Distributed Computation Over Slow Networks
Fan Lai 0001, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury |
NSDI | 1 |
| 2017 | A Linear Network Code Construction for General Integer Connections Based on the Constraint Satisfaction ProblemabstractThe problem of finding network codes for general connections is inherently difficult in capacity constrained networks. Resource minimization for general connections with network coding is further complicated. Existing methods for identifying solutions mainly rely on highly restricted classes of network codes, and are almost all centralized. In this paper, we introduce linear network mixing coefficients for code constructions of general connections that generalize random linear network coding for multicast connections. For such code constructions, we pose the problem of cost minimization for the subgraph involved in the coding solution and relate this minimization to a path-based constraint satisfaction problem (CSP) and an edge-based CSP. While CSPs are NP-complete in general, we present a path-based probabilistic distributed algorithm and an edge-based probabilistic distributed algorithm with almost sure convergence in finite time by applying communication free learning. Our approach allows fairly general coding across flows, guarantees no greater cost than routing, and shows a possible distributed implementation. Numerical results illustrate the performance improvement of our approach over existing methods. Ying Cui 0001, Muriel Médard, Edmund M. Yeh, Douglas J. Leith, Fan Lai 0001, Ken R. Duffy |
IEEE/ACM Trans. Netw. | 5 |
| 2016 | Optimal Caching and User Association in Cache-Enabled Heterogeneous Wireless NetworksabstractHeterogenous wireless networks (Hetnets) provide a powerful approach to meet the massive growth in traffic demands, but also impose a significant challenge on backhaul. Caching at small base stations (BSs) and wireless small cell backhaul have been proposed as attractive solutions to address this new challenge. In this paper, we consider the optimal caching and user association to minimize the total time to satisfy the average demands in cached-enabled Hetnets with wireless backhaul. We formulate this problem as a mixed discrete- continuous optimization for given bandwidth and cache resources. First, we characterize the structure of the optimal solution. Specifically, we show that the optimal caching is to store the most popular files at each pico BS, and the optimal user association has a threshold form. We also obtain the closed-form optimal solution in the homogenous scenario of pico cells. Then, we analyze the impact of bandwidth and cache resources on the minimum total time to satisfy the average demands. Finally, using numerical simulations, we verify the analytical results. Ying Cui 0001, Fan Lai 0001, Stephen Vaughan Hanly, Phil Whiting |
GLOBECOM | 2 |
| 2016 | Enhanced VIP Algorithms for Forwarding, Caching, and Congestion Control in Named Data NetworksabstractEmerging Information-Centric Networking (ICN) architectures seek to optimally utilize both bandwidth and storage for efficient content distribution over the network. The Virtual Interest Packet (VIP) framework has been proposed to enable joint design of forwarding, caching, and congestion control strategies within the Named Data Networking (NDN) architecture. While the existing VIP algorithms exhibit good performance, they are primarily focused on maximizing network throughput and utility, and do not explicitly consider user delay. In this paper, we develop a new class of enhanced algorithms for joint dynamic forwarding, caching and congestion control within the VIP framework. These enhanced VIP algorithms adaptively stabilize the network and maximize network utility, while improving the delay performance by intelligently making use of VIP information beyond one hop. Generalizing Lyapunov drift techniques, we prove the throughput optimality and characterize the utility-delay tradeoff of the enhanced VIP algorithms. Numerical experiments demonstrate the superior performance of the resulting enhanced algorithms for handling Interest Packets and Data Packets within the actual plane, in terms of low network delay and high network utility. Ying Cui 0001, Fan Lai 0001, Edmund M. Yeh, Ran Liu 0010 |
GLOBECOM | 2 |