Jiawen Niu

dblp:390/5998 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0003-0922-1942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 50% Parallel and multicore computing · 35% Cloud and datacenter computing · 10%
Artificial intelligence
3 papers
Efficient and distributed learning · 100%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.732026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data · Proc. VLDB Endow. 2025
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization · Proc. ACM Manag. Data 2025
Machine learning › Efficient and distributed learning › distributed training › parallelization
model partitioning
1.012026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
Distributed systems › fault tolerance
checkpointing
1.012026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
Distributed systems
fault tolerance
1.012026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training
0.912025
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization · Proc. ACM Manag. Data 2025
Machine learning and data management
data management for machine learning
0.912025
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data · Proc. VLDB Endow. 2025
Distributed systems › distributed machine learning
distributed training
0.912025
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment · Proc. ACM Manag. Data 2025
Parallel and multicore computing › parallel computing
heterogeneous parallelism
0.912025
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment · Proc. ACM Manag. Data 2025
Parallel and multicore computing
load balancing
0.912025
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment · Proc. ACM Manag. Data 2025
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
GPU cluster scheduling
0.312026
Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training · PPoPP 2026
High-performance computing › large-scale training
large language model training
0.312025
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment · Proc. ACM Manag. Data 2025
Parallel and multicore computing › parallelization strategies
model and data parallelism
0.312025
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization · Proc. ACM Manag. Data 2025

Methods — techniques the papers use, named apart from their topics

elastic model partitioning · 2.0checkpointing · 2.0workload balancing · 1.7planning algorithm · 1.7model state migration · 1.7low-rank adaptation · 1.7heterogeneous replica scheduling · 1.7two-stage data assignment · 0.9replanning · 0.9re-planning · 0.9dynamic heterogeneous parallel strategies · 0.9
YearPublicationVenuePosition
2026 Elastor: Elastic and Efficient Model Partitioning and Checkpointing for Fault-Tolerant Distributed Training
abstract
Distributed deep learning (DL) training faces instability from GPU/node failures of multi-GPU clusters, necessitating robust fault recovery from model checkpoints. However, we find that existing works only considers node failures but fails to handle partial GPU unavailability, and suffers from inefficient model checkpointing saving and loading, particularly when the GPU availability changes.
Fangcheng Fu, Haoyang Li 0017, Jiawen Niu, Bin Cui 0001
PPoPP6
2025 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
abstract
As the scale of models and training data continues to grow, there is an expanding reliance on more GPUs to train large-scale models, which inevitably increases the likelihood of encountering dynamic stragglers that some devices lag behind in performance occasionally. However, hybrid parallel training, one of the de facto paradigms to train large models, is typically sensitive to the stragglers. This paper presents Malleus , a straggler-resilient hybrid parallel training framework for large-scale models. Malleus quantifies the stragglers at the nuanced, per-GPU granularity during training, and develops a novel planning algorithm to deduce the optimal parallelization of GPU devices, pipeline stages, model layers, and training data, maximizing training efficiency when stragglers exist. In addition, once a shift in the straggler situation is detected, Malleus adaptively adjusts the parallelization via a re-planning process, and seamlessly and efficiently migrates the model states on the fly, without sacrificing the stability of the training tasks. Empirical results on large language models with up to 110B parameters show that Malleus consistently outperforms existing parallel training frameworks under various straggler situations, delivering on average 2.63-5.28x of efficiency improvement.
Haoyang Li 0017, Fangcheng Fu, Jiawen Niu, Hailin Zhang 0004, Xiaonan Nie, Bin Cui 0001
Proc. ACM Manag. Data6
2025 Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
abstract
To optimize large Transformer model training, both efficient parallel computing and advanced data management are indispensable. However, current methods often assume a stable and uniform training workload, neglecting data-induced imbalances-arising from both sampling and packing processes-which can impede training performance. Specifically, data sampling imbalance arises from uneven sequence length distribution of the training data, while data packing imbalance stems from the discrepancy between the linear memory complexity and quadratic time complexity of the attention mechanism. To address these imbalance issues, we develop Hydraulis, which jointly optimizes the parallel strategies and data assignment. For one thing, we introduce large model training with dynamic heterogeneous parallel strategies in response to the sequence length variations within and across training iterations. For another, we devise a two-stage data assignment approach, which strikes a good balance in terms of the training workloads both within and across model replicas. Empirical results demonstrate that Hydraulis outperforms existing systems by 1.32-2.66×. Our source code is available: https://github.com/PKU-DAIR/Hetu.
Haoyang Li 0017, Fangcheng Fu, Jiawen Niu, Jinbao Xue, Yangyu Tao, Di Wang 0052, Jie Jiang 0015, Bin Cui 0001
Proc. ACM Manag. Data6
2025 LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
abstract
With the breakthrough of Transformer-based pre-trained models, the demand for fine-tuning (FT) to adapt the base pre-trained models to downstream applications continues to grow, so it is essential for service providers to reduce the cost of processing FT requests. Low-rank adaption (LoRA) is a widely used FT technique that only trains small-scale adapters and keeps the base model unaltered, conveying the possibility of processing multiple FT tasks by jointly training different LoRA adapters with a shared base model. Nevertheless, through in-depth analysis, we reveal the efficiency of joint FT is dampened by two heterogeneity issues in the training data — the sequence length variation and skewness. To tackle these issues, we develop LobRA, a brand new framework that supports processing multiple FT tasks by jointly training LoRA adapters. Two innovative designs are introduced. Firstly, LobRA deploys the FT replicas (i.e., model replicas for FT) with heterogeneous resource usages and parallel configurations, matching the diverse workloads caused by the sequence length variation. Secondly, for each training step, LobRA takes account of the sequence length skewness and dispatches the training data among the heterogeneous FT replicas to achieve workload balance. We conduct experiments to assess the performance of LobRA, validating that it significantly reduces the GPU seconds required for joint FT by 45.03%-60.67%.
Fangcheng Fu, Haoyang Li 0017, Jiawen Niu, Yaofeng Tu, Bin Cui 0001
Proc. VLDB Endow.6