Lujia Yin

dblp:226/3707 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MlsDisk: Trusted Block Storage for TEEs Based on Layered Secure Logging
Erci Xu, Lujia Yin, Xinyuan Luo, Shaowei Song, Qingsong Chen, Shoumeng Yan, Jiwu Shu, Hongliang Tian, Yiming Zhang 0003
FAST3
2026 HyperWeave: QoS-aware GPU Overcommitment for Deep Learning Training Jobs
Yong Peng 0006, Lujia Yin, Miao Zhang 0037
IWQoS4
2026 Tree-based Publish-Subscribe Model and Load Partition for Distributed Object System
abstract
Object-oriented systems are commonly built by composing local components, while the composition relation is independent of their physical deployment. Existing message-queue-based middleware usually adopts a flat publish-subscribe model, making objects potentially reachable from one another and causing redundant cross-node communication. This paper proposes a communication-aware partitioning method based on a hierarchical publish-subscribe architecture. The system is modeled as a hierarchical communication tree, in which communication load is quantified through message propagation paths. Deployment mappings, cut-edge variables, and capacity constraints are encoded as an optimization model based on satisfiability modulo theories (SMT). Because directly solving the complete SMT model incurs high overhead in large-scale scenarios, we further design a bottom-up subtree partitioning algorithm as a highly scalable approximate solution. Experimental results show that, compared with the baseline algorithms, the proposed method reduces the weighted cross-node communication cost and provides an efficient near-optimal approximate partitioning scheme for large-scale hierarchical workloads within an acceptable accuracy range. It is suitable for distributed systems with hierarchical characteristics, such as EDA simulation.
Lujia Yin, Zhongxiang Dai, Xiufen Fu, Menglong Lu, Chuan Ai
SIGCOMM1
2025 Multi-agent reinforcement learning for task offloading with hybrid decision space in multi-access edge computing
Miao Zhang 0037, Quanjun Yin, Lujia Yin, Yong Peng 0006
Ad Hoc Networks4
2025 Training large-scale language models with limited GPU memory: a survey
abstract
Large-scale models have gained significant attention in a wide range of fields, such as computer vision and natural language processing, due to their effectiveness across various applications. However, a notable hurdle in training these large-scale models is the limited memory capacity of graphics processing units (GPUs). In this paper, we present a comprehensive survey focused on training large-scale models with limited GPU memory. The exploration commences by scrutinizing the factors that contribute to the consumption of GPU memory during the training process, namely model parameters, model states, and model activations. Following this analysis, we present an in-depth overview of the relevant research work that addresses these aspects individually. Finally, the paper concludes by presenting an outlook on the future of memory optimization in training large-scale language models, emphasizing the necessity for continued research and innovation in this area. This survey serves as a valuable resource for researchers and practitioners keen on comprehending the challenges and advancements in training large-scale language models with limited GPU memory.
Linbo Qiao, Lujia Yin, Peng Liang 0017, Dongsheng Li 0001
Frontiers Inf. Technol. Electron. Eng.3
2025 Koala: Efficient Pipeline Training through Automated Schedule Searching on Domain-Specific Language
abstract
Pipeline parallelism is a crucial technique for large-scale model training, enabling parameter splitting and performance enhancement. However, creating effective pipeline schedules often requires significant manual effort and coding skills, leading to practical inconveniences and complex debugging. Major frameworks such as DeepSpeed and ColossalAI simplify the process by adopting predefined pipeline schedule strategies, such as GPipe and 1F1B. The use of predefined schedules offers limited flexibility and suboptimal training efficiency, as the limited number of manually set candidates cannot provide the optimal strategy for arbitrary model training. To deal with the issue, this article aims to automatically search for the optimal strategy with high efficiency. Since current frameworks only support a limited set of fixed strategies, lacking the technical capability to create a comprehensive strategy search space, we first design a novel domain-specific language (DSL) for pipeline schedule development. The DSL exhibits great understandability, agility, and reusability, supporting the development of all known pipeline schedule strategies and their variants. Second, we are the first to model the complete pipeline schedule strategy space via the DSL, enabling an automated end-to-end globally optimal pipeline schedule searching, while past work may get stuck in a local optimum. Finally, we propose to optimize pipeline performance by modeling and solving the pipeline schedule as a Binary-Tree-Traversing (BTT) optimization problem. Based on the formalization, we further adopt a Dynamic Try-Test Genetic Algorithm to search for the best pipeline schedule strategy, which overwhelms a variety of pre-defined ones. Experimental results show that Koala achieves an enhanced performance by up to \(1.53\times\) over state-of-the-art approaches. Besides, the pipeline schedule strategy searched by Koala outperforms pre-defined pipeline schedule strategies by \(1.10\times \sim 1.55\times\) . Moreover, Koala has superior scalability and effectiveness in combining with data parallelism and tensor parallelism.
Lujia Yin, Qiao Li 0001, Hengjie Li, Xingcheng Zhang, Linbo Qiao, Dongsheng Li 0001
ACM Trans. Archit. Code Optim.2
2025 FlexHMB: A Flexible HMB Design Toward Bufferless Mobile Flash
abstract
Since the last decade, NAND flash has been widely adopted in mobile devices (e.g., smartphones) as the main storage media. Unlike enterprise solid-state drives and hard disk drives, mobile devices organize NAND flash in a DRAMless form, called mobile flash, which lacks an internal DRAM buffer to accommodate write requests due to the space constraints of mobile devices. Instead, it allocates Single-Level-Cell (SLC) flash blocks as a write buffer to accelerate write I/O bursts. However, we observe that this well-known design not only fails to consistently deliver high write performance but also compromises mobile flash capacity and longevity. The former is caused by the severe SLC reclamation interference, while the latter stems from its low density and the occupation of scarce over-provisioning blocks. To address these challenges, we propose FlexHMB, which utilizes the mobile flash controller to manage a high-performance and flexible write buffer allocated from the host-side main memory. Specifically, FlexHMB leverages the host memory buffer (HMB) feature to replace the traditional SLC buffer with a DRAM-based one. By doing so, FlexHMB shifts mobile flash towards bufferless architecture and eliminates penalties from the SLC buffer. To avoid the potential reliability issue arising from the volatility of DRAM, we design a write transaction mechanism to guarantee order consistency. While a large buffer can deliver higher performance, it may lead to competition with mobile applications for memory resources. Taking this into consideration, we design a flexible and dynamic resizing mechanism for the write buffer to make a balance between efficient request handling and memory utilization. The evaluation results show that FlexHMB reduces write latency by 93.54% compared to the SLC write buffer when memory usage is constrained.
Yong Peng 0006, Shaocong Sun, Lujia Yin, Qiao Li 0001, Jie Zhang 0048
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 DELTA: Memory-Efficient Training via Dynamic Fine-Grained Recomputation and Swapping
abstract
To accommodate the increasingly large-scale models within limited-capacity GPU memory, various coarse-grained techniques, such as recomputation and swapping, have been proposed to optimize memory usage. However, these methods have encountered limitations, either in terms of inefficient memory reduction or diminished training performance. In response to this, our article introduces dynamic tensor offloading and recomputation (DELTA), an innovative approach for memory-efficient large-scale model training that combines fine-grained memory optimization and prefetching technology to reduce memory usage while maintaining high training throughput concurrently. Initially, we formulate the problem of memory-throughput joint optimization as an easy-solving 0/1 Knapsack problem. Leveraging this formalization, we use an improving polynomial complexity heuristic algorithm to address the problem effectively. Furthermore, we introduce, to the best of our knowledge, a novel bidirectional prefetching technology into dynamic memory management that significantly accelerates the model training when compared to relying solely on recomputation or swapping. Finally, DELTA offers users an automated training execution library, eliminating the need for manual configuration or specialized expertise. Experimental results demonstrate the effectiveness of DELTA in reducing GPU memory consumption. Compared to state-of-the-art methods, DELTA achieves substantial memory savings ranging from 40% to 72%, while maintaining comparable convergence performance for various models, including ResNet-50, ResNet-101, and BERT-Large. Notably, DELTA enables the training of GPT2-Large and GPT2-XL with batch sizes increased by 5.5× and 6×, respectively, showcasing its versatility and practicality in enabling large-scale model training on GPU hardware.
Qiao Li 0001, Lujia Yin, Dongsheng Li 0001, Yiming Zhang 0003, Xingcheng Zhang, Linbo Qiao, Zhaoning Zhang 0001, Kai Lu 0001
ACM Trans. Archit. Code Optim.3
2024 Low-Cost, High-Reliability Deployment for Cloud Applications With Low-Frequency Periodic Requests
abstract
Low-frequency periodic requests are common in cloud-based enterprise applications. These infrequent requests often leave microservices idle for extended periods, leading to low resource utilization. Furthermore, the randomness of response times may decrease the reliability of the cloud platform. Intuitively, the periodic nature of requests allows for the agile deployment of microservices to promptly free up occupied computing resources. Thus, the key lies in designing low-cost, high-reliability microservice deployment schemes. Traditional approaches relying on specialized expertise are impractical because of intricate interdependencies within microservice frameworks. To address this, the Microservice Deployment Problem for Low-frequency Periodic Requests (MDP-LPR) is formulated, and a Mixed Integer Programming (MIP) model is developed. A deployment framework leveraging statistical analysis and Monte Carlo simulation is proposed to ensure high reliability. Furthermore, a two-stage heuristic algorithm named Relaxation and Precision Mixed Algorithm (RPMA) is introduced to generate low-cost deployment schemes. Finally, experiments are conducted on real-world workflows. The results show that the RPMA outperforms its counterparts in generating low-cost deployment schemes, and the proposed deployment framework enables the automatic acquisition of low-cost, high-reliability deployment schemes.
Zhu Xiang, Lujia Yin, Miao Zhang 0037, Quanjun Yin
IEEE Trans. Serv. Comput.3
2023 AMGmal: Adaptive mask-guided adversarial attack against malware detection with minimal perturbation
Dazhi Zhan, Yexin Duan, Yue Hu 0016, Lujia Yin, Zhisong Pan 0003, Shize Guo
Comput. Secur.4
2022 ParaX : Bandwidth-Efficient Instance Assignment for DL on Multi-NUMA Many-Core CPUs
abstract
Commercial clouds now heavily use CPUs in DL (deep learning) because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on multi-NUMA (non-uniform memory access) many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by$1.73\times \sim 2.93{\times}$.
Yiming Zhang 0003, Lujia Yin, Dongsheng Li 0001, Yuxing Peng 0001, Kai Lu 0001
IEEE Trans. Computers2
2021 MapperX: Adaptive Metadata Maintenance for Fast Crash Recovery of DM-Cache Based Hybrid Storage Devices
Lujia Yin, Yiming Zhang 0003, Yuxing Peng 0001
USENIX ATC1
2021 Elastic scheduler: Heterogeneous and dynamic deep Learning in the cloud
abstract
Abstract GPUs and CPUs have been widely used for model training of deep learning (DL) in the cloud, where both DL workloads and resource usage might heavily change over time. Traditional training methods require beforehand specification on the type (either GPUs or CPUs) and amount of computing devices, and thus cannot elastically schedule the dynamic DL workloads onto available GPUs/CPUs. In this paper, we propose Elastic Scheduler (ES), a novel approach that efficiently supports both heterogeneous training (with different device types) and dynamic training (with varying device numbers). ES (i) accumulates local gradients and simulates multiple virtual workers on one GPU to alleviate the performance gap between GPUs and CPUs for achieving similar accuracy in heterogeneous GPU‐CPU‐hybrid training as in homogeneous training and (ii) uses local gradients stabilizes batch sizes for high accuracy without long compensation. Experiments show that ES achieves significantly higher performance than existing methods for heterogeneous and dynamic training as well as inference.
Lujia Yin, Yiming Zhang 0003, Yuxing Peng 0001, Dongsheng Li 0001
Concurr. Comput. Pract. Exp.1
2021 ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs
abstract
Despite the fact that GPUs and accelerators are more efficient in deep learning (DL), commercial clouds like Facebook and Amazon now heavily use CPUs in DL computation because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by 1.73X ~ 2.93X.
Lujia Yin, Yiming Zhang 0003, Zhaoning Zhang 0001, Yuxing Peng 0001
Proc. VLDB Endow.1
2019 Rise the Momentum: A Method for Reducing the Training Error on Multiple GPUs
Lujia Yin, Zhaoning Zhang 0001, Dongsheng Li 0001
ICA3PP (2)2
2018 A Quick Survey on Large Scale Distributed Deep Learning Systems
abstract
Deep learning have been widely used in various fields and has worked very well as a major role. While the gradual penetration into various fields, data quantity of each applications is increasing tremendously, and so as the computation complexity and model parameters. As an obvious result, the training and inference is time consuming. For example, a classic Resnet50 classification model will be trained in 14 days on a NVIDIA M40 GPU with ImageNet data set. Thus, distributed acceleration is a very useful way to dispatch the computation of training and even inference to scale of nodes in parallel and accelerate the whole process. Facebook's work and UC Berkeley's acceleration can training the Resnet-50 model within hour and minutes by distributed deep learning algorithm and system, representatively. As other distributed accelerations, it gives a possibility to accelerate large models on large data sets from weeks to minutes, which gives researchers and developers more space to explore and search. However, besides acceleration, what other issues will be confronted of the distributed deep learning system? Where is the upper limit of acceleration? What application will acceleration be used for? What is the price and cost of acceleration? In this paper, we will take a simple and quick survey on the distributed deep learning system from algorithm perspective, distributed system perspective and applications perspective. We will present several recent excellent works, and bring analysis on the restricts and prospects of the distributed methods.
Zhaoning Zhang 0001, Lujia Yin, Yuxing Peng 0001, Dongsheng Li 0001
ICPADS2