Jiamin Li 0002

dblp:81/3437-2 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-8110-2436ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 StitchLLM: Serving LLMs, One Block at a Time
abstract
Bodun Hu, Shuozhe Li, Saurabh Agarwal, Myungjin Lee, Akshay Jajoo, Jiamin Li, Le Xu, Geon-Woo Kim, Donghyun Kim, Hong Xu, Amy Zhang, Aditya Akella. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Bodun Hu, Shuozhe Li, Myungjin Lee, Akshay Jajoo, Jiamin Li 0002, Geon-Woo Kim, Donghyun Kim 0002, Hong Xu 0001, Amy Zhang 0001, Aditya Akella
ACL (1)6
2024 Arlo: Serving Transformer-based Language Models with Dynamic Input Lengths
abstract
A prominent challenge in serving requests for NLP tasks is handling the varying length of input texts. Existing solutions, such as uniform zero-padding and compiler support, suffer from either computational inefficiency or suboptimal latency. To address these practical issues, we propose an approach called polymorphing. Polymorphing involves creating and utilizing multiple runtimes of the model, each statically compiled with a different input length, to serve requests accordingly. This fine-grained use of statically-compiled runtimes reduces the overheads of zero-padding while improving latency performance compared to dynamic compilation. To practically realize polymorphing, we have developed an inference scheduling system, Arlo, which leverages the observed input length distribution to periodically allocate compute resources across multiple runtimes by solving an integer linear program. Upon request arrival, Arlo uses a multi-level queue-based heuristic to dispatch requests to the most suitable runtime instances, efficiently adapting to the dynamics of request length and instance load. Extensive testbed evaluations and large-scale simulations using production traces demonstrate Arlo’s promising potential. It achieves 23.7%–98.1% mean latency reductions compared to existing schemes while significantly reducing tail latency.
Xin Tan 0004, Jiamin Li 0002, Jingzong Li, Hong Xu 0001
ICPP2
2023 Adaptive Gating in Mixture-of-Experts based Language Models
abstract
Large language models, such as OpenAI's Chat-GPT, have demonstrated exceptional language understanding capabilities in various NLP tasks.Sparsely activated mixture-of-experts (MoE) has emerged as a promising solution for scaling models while maintaining a constant number of computational operations.Existing MoE model adopts a fixed gating network where each token is computed by the same number of experts.However, this approach contradicts our intuition that the tokens in each sequence vary in terms of their linguistic complexity and, consequently, require different computational costs.Little is discussed in prior research on the tradeoff between computation per token and model performance.This paper introduces adaptive gating in MoE, a flexible training strategy that allows tokens to be processed by a variable number of experts based on expert probability distribution.The proposed framework preserves sparsity while improving training efficiency.Additionally, curriculum learning is leveraged to further reduce training time.Extensive experiments on diverse NLP tasks show that adaptive gating reduces at most 22.5% training time while maintaining inference quality.Moreover, we conduct a comprehensive analysis of the routing decisions and present our insights when adaptive gating is used.
Jiamin Li 0002, Cong Wang 0001, Hong Xu 0001
EMNLP1
2023 Lyra: Elastic Scheduling for Deep Learning Clusters
abstract
Organizations often build separate training and inference clusters for deep learning, and use separate schedulers to manage them. This leads to problems for both: inference clusters have low utilization when the traffic load is low; training jobs often experience long queuing due to a lack of resources. We introduce Lyra, a new cluster scheduler to address these problems. Lyra introduces capacity loaning to loan idle inference servers for training jobs. It further exploits elastic scaling that scales a training job's resource allocation to better utilize loaned servers. Capacity loaning and elastic scaling create new challenges to cluster management. When the loaned servers need to be returned, we need to minimize job preemptions; when more GPUs become available, we need to allocate them to elastic jobs and minimize the job completion time (JCT). Lyra addresses these combinatorial problems with principled heuristics. It introduces the notion of server preemption cost, which it greedily reduces during server reclaiming. It further relies on the JCT reduction value defined for each additional worker of an elastic job to solve the scheduling problem as a multiple-choice knapsack problem. Prototype implementation on a 64-GPU testbed and large-scale simulation with 15-day traces of over 50,000 production jobs show that Lyra brings 1.53x and 1.48x reductions in average queuing time and JCT, and improves cluster usage by up to 25%.
Jiamin Li 0002, Hong Xu 0001, Yibo Zhu 0001, Zherui Liu, Chuanxiong Guo, Cong Wang 0001
EuroSys1
2023 Accelerating Distributed MoE Training and Inference with Lina
Jiamin Li 0002, Yibo Zhu 0001, Cong Wang 0001, Hong Xu 0001
USENIX ATC1
2023 Bottleneck-Aware Non-Clairvoyant Coflow Scheduling With Fai
abstract
Coflow scheduling is critical to data-parallel applications in data centers. While schemes like Varys can achieve optimal performance, they require a priori information about coflows which is hard to obtain in practice. Existing non-clairvoyant solutions like Aalo generalize least attained service (LAS) scheduling discipline to address this issue. However, they fail to identify the bottleneck flows in a coflow and tend to allocate excessive bandwidth to the non-bottleneck flows, leading to bandwidth wastage and inferior overall performance. To this end, we present Fai that strives to improve the overall coflow performance by accelerating the bottleneck flows without priori knowledge. Fai employs bottleneck-aware scheduling. It adopts loose coordination to update coflow priority and flow rates based on total bytes sent. In addition, Fai detects bottleneck flows based on a flow’s rate and bytes sent, and de-allocates bandwidth for other flows to match the bottleneck rate without affecting the coflow completion time (CCT). The saved bandwidth is then distributed among coflows according to their priority to improve overall performance. Testbed evaluation on a 40-node cluster shows that Fai improves average (P95) CCT by 1.73× (3.43×), compared to Aalo. Large-scale trace-driven simulations also show that Fai outperforms Aalo substantially.
Libin Liu 0001, Chengxi Gao, Peng Wang 0037, Hongming Huang, Jiamin Li 0002, Hong Xu 0001, Wei Zhang 0049
IEEE Trans. Cloud Comput.5
2022 ScaleFlux: Efficient Stateful Scaling in NFV
abstract
Network function virtualization (NFV) enables elastic scaling to middlebox deployment and management. Therefore, efficient stateful scaling is an important task because operators often need to shift traffic and the associated flow states across VNF instances to deal with time-varying loads. Existing NFV scaling methods, however, typically focus on one aspect of the scaling pipeline and does not offer an end-to-end scaling framework. This article presents ScaleFlux, a complete stateful scaling system that efficiently reduces flow-level latency and achieves near-optimal resource usage. ScaleFlux (1) monitors traffic load for each VNF instance and adopts a queue-based mechanism to detect load burstiness timely, (2) deploys a flow bandwidth predictor to predict flow bandwidth time-series with the ABCNN-LSTM model, and (3) schedules the necessary flow and state migration using the simulated annealing algorithm to achieve both flow-level latency guarantee and resource usage minimization. Testbed evaluation with a five-machine cluster shows that ScaleFlux reduces flow completion time by at least 8.7× for all the workloads and achieves near-optimal CPU usage during scaling.
Libin Liu 0001, Hong Xu 0001, Zhixiong Niu, Jingzong Li, Wei Zhang 0049, Peng Wang 0037, Jiamin Li 0002, Chun Jason Xue, Cong Wang 0001
IEEE Trans. Parallel Distributed Syst.7
2021 Two-Dimensional Learning Rate Decay: Towards Accurate Federated Learning with Non-IID Data
abstract
In federated learning a global model is trained with training data geographically distributed over a number of clients. To reduce the communication cost over the expensive wide area network, clients complete multiple local iterations before synchronization. However, since the training data are non-iid, such infrequent synchronization would compromise the accuracy after model convergence. In order to tackle this problem, we propose Two-Dimensional Learning Rate Decay (2D-LRD) in this paper, which aims to improve the model performance by adaptively tuning the learning rate on two dimensions: round-dimension and iteration-dimension during the model training. That is, we gradually decrease the learning rate and decrease the learning rates of local iterations in a synchronization round with different speeds. Based on our experiments and analysis, we find that the sum of the inner product of round updates is a valuable signal for learning rate tuning. We perform evaluation and demonstrate that 2D-LRD can make great progress compared to the baseline scheme.
Kaiwei Mo, Chen Chen 0067, Jiamin Li 0002, Hong Xu 0001, Chun Jason Xue
IJCNN3