VLDB 2026 Research / reviewers in the wild / expert
Junwei Cui
dblp:331/1909
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0001-6805-7669ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XTree on EquiMesh: Topology and Algorithm Co-Design for Collective CommunicationabstractMesh topology is widely adopted for both on-chip and chiplet-based interconnects due to its placement-friendly physical layout. However, the low-degree nodes at the edges and corners create bandwidth bottlenecks for common collectives such as AllGather and AllReduce. We address this limitation with EquiMesh, an augmented 2D-Mesh with equivalent-degree nodes without incurring switching complexity. To fully exploit EquiMesh, we propose XTree, a topology-aware algorithm that maximizes utilization of available bandwidth, and MirrorXTree, which constructs ReduceScatter and AllReduce on top of XTree through topology mirroring. Our evaluation shows that EquiMesh with XTree and MirrorXTree achieves 2× and 1.2× higher effective bandwidth than state-of-the-art mesh-based topology-algorithm co-designs for AllGather and AllReduce, respectively. Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
DATE | 1 |
| 2026 | Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference
Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
ISCA | 1 |
| 2025 | Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsabstractExpert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the processing of increasingly large-scale models. However, the All-to-All communication inherent to expert parallelism poses a significant bottleneck, limiting the efficiency of MoE models. Although existing optimization methods partially mitigate this issue, they remain constrained by the sequential dependency between communication and computation operations.
To address this challenge, we propose ScMoE, a novel shortcut-connected MoE architecture integrated with an overlapping parallelization strategy. ScMoE decouples communication from its conventional sequential ordering, enabling up to 100\% overlap with computation.
Compared to the prevalent top-2 MoE baseline, ScMoE achieves speedups of $1.49\times$ in training and $1.82\times$ in inference.
Moreover, our experiments and analyses indicate that ScMoE not only achieves comparable but in some instances surpasses the model quality of existing approaches. Weilin Cai, Juyong Jiang, Junwei Cui, Sunghun Kim 0001, Jiayi Huang 0001 |
ICML | 4 |
| 2025 | Chimera: Communication Fusion for Hybrid Parallelism in Large Language ModelsabstractLarge Language Models (LLMs), exemplified by ChatGPT, have emerged as a predominant workload in current machine learning systems.To achieve efficient training and inference within the constraints of limited single-NPU memory capacity, deploying LLMs on multi-NPU systems typically adopt a hybrid approach that combines various parallelism patterns.This hybrid parallelism within LLMs introduces a significant amount of diverse collective communications.However, these frequent blocking communications impose a substantial burden on the multi-NPU systems.Overcoming the communication bottleneck is crucial to unlocking the potential of multi-NPU systems for efficient and scalable LLM processing.This paper introduces Chimera, a communication fusion mechanism for hybrid parallelism in LLMs.We comprehensively analyze the communication processes of each LLM parallelism pattern, identify the communication redundancy in hybrid parallelism and eliminate redundancy by fusing adjacent communication operators during parallelism transformation.By reordering operations and generating redundancy-free communication operator, Chimera effectively mitigates communication bottleneck in hybrid LLM parallelism.Our results show that Chimera achieves 1.23-7.06×network bandwidth speedup.Additionally, the end-to-end performance of LLM forward pass and backward pass on different typical multi-NPU systems achieves respective 1.32-1.58×and 1.16-1.36×speedups on average compared with those without communication fusion. Junwei Cui, Weilin Cai, Jiayi Huang 0001 |
ISCA | 2 |
| 2025 | Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
Junwei Cui, Weilin Cai, Meng Niu, Jiayi Huang 0001 |
MICRO | 2 |