Tingwen Xie

dblp:354/4212 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0003-4739-9450ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 75% Optimization for machine learning · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 44% Parallel and multicore computing · 44% Hardware accelerators and domain-specific architectures · 13%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
automatic parallelization
1.012026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026
Distributed systems › distributed machine learning
distributed training
1.012026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.712023
SLAMB: Accelerated Large Batch Training with Sparse Communication · ICML 2023
Machine learning › Efficient and distributed learning
distributed training
0.712023
SLAMB: Accelerated Large Batch Training with Sparse Communication · ICML 2023
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.712023
SLAMB: Accelerated Large Batch Training with Sparse Communication · ICML 2023
Machine learning › Optimization for machine learning › large-scale optimization
large batch optimization
0.712023
SLAMB: Accelerated Large Batch Training with Sparse Communication · ICML 2023
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator
0.312026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

random forest · 1.0monte carlo tree search · 1.0cost model · 1.0communication overlap · 1.0sparsification · 0.7momentum masking · 0.7local error compensation · 0.7layer-wise adaptive moments · 0.7
YearPublicationVenuePosition
2026 HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training
abstract
As large neural network models (e.g., LLMs) grow in scale, single-cluster resources become insufficient, making cross-cluster distributed training essential. Cross-cluster training is challenging: hardware heterogeneity complicates load balancing and parallelization strategy and introduces hardware compatibility issues in implementation; cross-cluster communication bottlenecks severely impact training throughput. We present HetAuto, an automatic parallelization system for efficient cross-cluster heterogeneous large model training. HetAuto contributes three key innovations: (1) a principle-guided MCTS algorithm with a random forest-enhanced cost model that efficiently searches parallelization strategies and quickly evaluates their performance under heterogeneous configurations; (2) cross-cluster communication optimizations including Virtual-1F1B scheduling that overlaps communication with computation and an optimized resharding strategy for inter-stage communication; and (3) a unified API enabling seamless integration of diverse accelerators. We evaluate HetAuto across 4 different clusters with up to 736 heterogeneous devices. The evaluation results show that HetAuto achieves up to 1.57× training throughput improvement over representative baselines, and strikes an efficient balance between solution quality and search overhead.
Guicheng Qi, Junwei Su, Liqi Yang, Tingwen Xie, Yerui Sun, Chuan Wu 0001
EuroSys5
2023 SLAMB: Accelerated Large Batch Training with Sparse Communication
abstract
Distributed training of large deep neural networks requires frequent exchange of massive data between machines, thus communication efficiency is a major concern. Existing compressed communication methods are either not compatible with large batch optimization algorithms, or do not provide sufficient speedup in large scale. In this paper, we combine sparsification-based gradient compression with the layer-wise adaptive moments optimizer for large batch training (LAMB). We propose SLAMB, a novel communication-efficient optimizer that supports large batch sizes and scales to thousands of GPUs. SLAMB employs momentum masking, local error compensation, and element-wise adaptive rescaling to achieve accurate layer-wise weight updates, which translates to fast convergence for very large batches. Our empirical results show that, compared to the state-of-the-art, SLAMB transmits half the amount of data in large-batch BERT pre-training, without sacrificing accuracy. Moreover, SLAMB achieves excellent scalability in large computing infrastructures. For instance, SLAMB with 128 GPUs reduces the training time of Swin Transformer pre-training on ImageNet to 5.35 hours, which is 2 hours faster than the state-of-the-art. At the extreme, we trained BERT-XL (2.8B parameters) on 1,024 NVIDIA A100 GPUs, where SLAMB achieved 90% scaling efficiency.
Wenxuan Zhang 0003, Jiawei Fei, Tingwen Xie, Mohamed Elhoseiny 0001, Panos Kalnis
ICML5