Houming Wu

dblp:88/6134 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 75% Parallel and multicore computing · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning › distributed training
communication-efficient distributed training
1.012026
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training · AAAI 2026
Distributed systems › distributed machine learning
distributed training
1.012026
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training · AAAI 2026
Parallel and multicore computing
pipeline parallelism
1.012026
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training · AAAI 2026
Distributed systems › communication optimization
topology-aware communication
1.012026
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training · AAAI 2026

Methods — techniques the papers use, named apart from their topics

topology-aware scheduling · 1.0pipeline parallelism · 1.0communication-computation overlap · 1.0
YearPublicationVenuePosition
2026 TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
abstract
Training large language models (LLMs) is fundamentally constrained by limited device memory and costly inter-device communication. Although pipeline parallelism alleviates memory pressure by partitioning models across devices, it incurs activation communication overhead that scales linearly with sequence length, limiting efficiency in long-context training. Recent weight-passing approaches (e.g., WeiPipe) mitigate this by transmitting model weights instead of activations, but suffer from redundant peer-to-peer (P2P) transfers and underutilized intra-node bandwidth. We propose TawPipe—topology-aware weight pipeline parallelism, which exploits hierarchical bandwidth in distributed clusters for improved communication efficiency. TawPipe: (i) groups devices based on topology to optimize intra-node collective and inter-node P2P communication; (ii) assigns each device a fixed shard of model weights and gradients, avoiding redundant transfers; and (iii) overlaps communication with computation to hide latency. Unlike global collective operations used in fully sharded data parallelism (FSDP), TawPipe confines most communication within node boundaries, significantly reducing cross-node traffic. Extensive experiments on up to 24 GPUs with LLaMA‑style models show that TawPipe achieves superior throughput and scalability compared to state-of-the-art baselines.
Houming Wu, Ling Chen 0001
AAAI1