Xuanlei Zhao

dblp:367/3224 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0000-4877-3115ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 60% Generative modeling · 20% Deep learning architectures and training · 16%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 47% Distributed systems · 25% High-performance computing · 22%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 23 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.732026
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism · PPoPP 2026
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training · NeurIPS 2025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
Machine learning › Generative modeling
diffusion model
1.722025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Real-Time Video Generation with Pyramid Attention Broadcast · ICLR 2025
Compilers and program optimization
deep learning compiler
1.622025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference · ICLR 2024
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
1.012026
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism · PPoPP 2026
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training
communication optimization
0.912025
Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning · ASPLOS (1) 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
Real-Time Video Generation with Pyramid Attention Broadcast · ICLR 2025
Machine learning › Efficient and distributed learning › efficient training
long-context training
0.912025
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training · NeurIPS 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training › model parallelism
sequence parallelism
0.912025
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training · NeurIPS 2025
Machine learning › Efficient and distributed learning › efficient training
training acceleration
0.912025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers · ICML 2025
Machine learning › Deep learning architectures and training › transformer
transformer training
0.912025
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Real-Time Video Generation with Pyramid Attention Broadcast · ICLR 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
zero-shot domain adaptation
0.912025
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025
Parallel and multicore computing › parallel algorithms › parallel algorithm design
sequence parallelism
0.912025
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers · ICML 2025
Machine learning › Efficient and distributed learning › memory-efficient training
activation memory reduction
0.812024
AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference · ICLR 2024
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
0.812024
AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference · ICLR 2024
Compilers and program optimization
memory optimization
0.812024
AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference · ICLR 2024
Parallel and multicore computing › parallelization strategies
model parallelism
0.812024
FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters · PPoPP 2024
Distributed systems › communication optimization
communication-computation overlap
0.312026
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism · PPoPP 2026
Distributed systems › distributed machine learning
distributed training
0.312026
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism · PPoPP 2026
Distributed systems › distributed machine learning
distributed inference
0.312025
Real-Time Video Generation with Pyramid Attention Broadcast · ICLR 2025

Methods — techniques the papers use, named apart from their topics

recomputation · 2.0pipeline parallelism · 2.0attention parallel partition · 2.0sequence parallelism · 1.7resharding · 1.7dynamic sequence parallelism · 1.7attention broadcasting · 1.7text encoder · 0.9resource-constrained project scheduling · 0.9hyper-convolutional decoder · 0.9auto-decomposition · 0.9LoRA · 0.9dynamic axial parallelism · 0.8code generation · 0.8chunk strategy search · 0.8auto-chunking · 0.8asynchronous operations · 0.8
YearPublicationVenuePosition
2026 HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
abstract
As transformer sequence lengths grow, existing pipeline parallelisms incur suboptimal performance due to the quadratic attention computation and the substantial memory overhead. To relieve these challenges, we propose HelixPipe, a novel pipeline parallelism for long sequence transformer training. First, HelixPipe introduces attention parallel partition, which schedules attention computations of different micro batches across different pipeline stages in parallel, reducing pipeline bubbles. Second, it employs a two-fold first-in-last-out micro batch schedule to balance memory usage and overlap communication with computation. Additionally, HelixPipe utilizes recomputation without attention and chunked MLP to mitigate fragmentation and enable longer sequences. Experiments demonstrate that HelixPipe gains increasing advantages with longer sequence lengths, and outperforms existing methods in throughput and scalability across varying pipeline sizes, model sizes, and cluster configurations. Notably, it achieves a 26% speedup over baseline methods when training a 7B model with 128k sequence length on 64 H20 GPUs. Code is available at https://github.com/zxgx/Megatron-LM.
Geng Zhang 0002, Shenggan Cheng, Xuanlei Zhao, Yang You 0001
PPoPP3
2025 Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning
abstract
With the exponential growth of deep learning (DL), there arises an escalating need for scalability. Despite significant advancements in communication hardware capabilities, the time consumed by communication remains a bottleneck during training. The existing various optimizations are coupled within parallel systems to implement specific computation-communication overlap. These approaches pose challenges in terms of performance, programmability, and generality. In this paper, we introduce Concerto, a compiler framework designed to address these challenges by automatically optimizing and scheduling communication. We formulate the scheduling problem as a resource-constrained project scheduling problem and use off-the-shelf solver to get the near-optimal scheduling. And use auto-decomposition to create overlap opportunity for critical (synchronous) communication. Our evaluation shows Concerto can match or outperform state-of-the-art parallel frameworks, including Megatron-LM, JAX/XLA, DeepSpeed, and Alpa, all of which include extensive hand-crafted optimization. Unlike previous works, Concerto decouples the parallel approach and communication optimization, then can generalize to a wide variety of parallelisms without manual optimization.
Shenggan Cheng, Shengjie Lin, Lansong Diao, Hao Wu 0077, Siyu Wang 0006, Chang Si, Xuanlei Zhao, Jiangsu Du, Wei Lin 0016, Yang You 0001
ASPLOS (1)8
2025 Real-Time Video Generation with Pyramid Attention Broadcast
abstract
We present Pyramid Attention Broadcast (PAB), a real-time, high quality and training-free approach for DiT-based video generation. Our method is founded on the observation that attention difference in the diffusion process exhibits a U-shaped pattern, indicating significant redundancy. We mitigate this by broadcasting attention outputs to subsequent steps in a pyramid style. It applies different broadcast strategies to each attention based on their variance for best efficiency. We further introduce broadcast sequence parallel for more efficient distributed inference. PAB demonstrates up to 10.5x speedup across three models compared to baselines, achieving real-time generation for up to 720p videos. We anticipate that our simple yet effective method will serve as a robust baseline and facilitate future research and application for video generation.
Xuanlei Zhao, Kai Wang 0036, Yang You 0001
ICLR1
2025 DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
abstract
Scaling multi-dimensional transformers to long sequences is indispensable across various domains. However, the challenges of large memory requirements and slow speeds of such sequences necessitate sequence parallelism. All existing approaches fall under the category of embedded sequence parallelism, which are limited to shard along a single sequence dimension, thereby introducing significant communication overhead. However, the nature of multi-dimensional transformers involves independent calculations across multiple sequence dimensions. To this end, we propose Dynamic Sequence Parallelism (DSP) as a novel abstraction of sequence parallelism. DSP dynamically switches the parallel dimension among all sequences according to the computation stage with efficient resharding strategy. DSP offers significant reductions in communication costs, adaptability across modules, and ease of implementation with minimal constraints. Experimental evaluations demonstrate DSP's superiority over state-of-the-art embedded sequence parallelism methods by remarkable throughput improvements ranging from 32.2% to 10x, with less than 25% communication volume.
Xuanlei Zhao, Shenggan Cheng, Zangwei Zheng, Zheming Yang, Yang You 0001
ICML1
2025 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
abstract
Modern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to \textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs. We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research.
Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036
NeurIPS4
2025 StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
abstract
Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements. Specifically, StarTrail introduces an extra parallel dimension and divides the peer-to-peer communication into sub-rings to substantially reduce communication volume and avoid bandwidth bottlenecks. Through comprehensive experiments across diverse hardware environments and on both Natural Language Processing (NLP) and Computer Vision (CV) tasks, we demonstrate that our approach significantly surpasses state-of-the-art methods that support Long sequence lengths, achieving performance improvements of up to 77.12% on GPT-style models and up to 114.33% on DiT (Diffusion Transformer) models without affecting the computations results.
Shenggan Cheng, Kai Wang 0036, Xuanlei Zhao, James Demmel, Yang You 0001
NeurIPS6
2025 REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
abstract
Diffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks.
Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001
NeurIPS7
2024 AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference
abstract
Large deep learning models have achieved impressive performance across a range of applications. However, their large memory requirements, including parameter memory and activation memory, have become a significant challenge for their practical serving. While existing methods mainly address parameter memory, the importance of activation memory has been overlooked. Especially for long input sequences, activation memory is expected to experience a significant exponential growth as the length of sequences increases. In this approach, we propose AutoChunk, an automatic and adaptive compiler system that efficiently reduces activation memory for long sequence inference by chunk strategies. The proposed system generates chunk plans by optimizing through multiple stages. In each stage, the chunk search pass explores all possible chunk candidates and the chunk selection pass identifies the optimal one. At runtime, AutoChunk employs code generation to automatically apply chunk strategies. The experiments demonstrate that AutoChunk can reduce over 80% of activation memory while maintaining speed loss within 10%, extend max sequence length by 3.2x to 11.7x, and outperform state-of-the-art methods by a large margin.
Xuanlei Zhao, Shenggan Cheng, Guangyang Lu, Yang You 0001
ICLR1
2024 FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters
abstract
Protein structure prediction helps to understand gene translation and protein function, which is of growing interest and importance in structural biology. The AlphaFold model, which used transformer architecture to achieve atomic-level accuracy in protein structure prediction, was a significant breakthrough. However, training and inference of AlphaFold model are challenging due to its high computation and memory cost. In this work, we present FastFold, an efficient implementation of AlphaFold for both training and inference. We propose Dynamic Axial Parallelism (DAP) as a novel model parallelism method. Additionally, we have implemented a series of low-level optimizations aimed at reducing communication, computation, and memory costs. These optimizations include Duality Async Operations, highly optimized kernels, and AutoChunk (an automated search algorithm finds the best chunk strategy to reduce memory peaks). Experimental results show that FastFold can efficiently scale to more GPUs using DAP and reduces overall training time from 11 days to 67 hours and achieves 7.5 ~ 9.5× speedup for long-sequence inference. Furthermore, AutoChunk can reduce memory cost by over 80% during inference by automatically partitioning the intermediate tensors during the computation.
Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Ruidong Wu, Jian Peng 0001, Yang You 0001
PPoPP2