Chizheng Fang

dblp:430/5698 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0005-4936-167XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 75% Deep learning architectures and training · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 77% GPUs and heterogeneous computing · 23%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
nnScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training · IEEE Trans. Parallel Distributed Syst. 2026
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026
Machine learning › Efficient and distributed learning
model inference
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026

Methods — techniques the papers use, named apart from their topics

search space pruning · 2.0placement-aware cost estimation · 2.0constraint-guided search · 2.0tensor parallelism · 1.0expert parallelism · 1.0dynamic programming · 1.0binary search · 1.0
YearPublicationVenuePosition
2026 SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing
abstract
To accelerate Mixture-of-Experts (MoE) inference, the hybrid parallelism paradigm is first applying pipeline parallelism (PP) to vertically divide the model into stages, with each stage further divided horizontally using tensor or expert parallelism. On the algorithm side, dynamic Top-K routing reduces computation by activating fewer experts per token on average. In this paper, we explore the application of dynamic Top-K routing to PP-enabled MoE inference, aiming to fully unleash their combined potential. We identify key performance bottlenecks arising from Top-K value variation across layers, which conflicts with PP's typically uniform stage partitioning, as well as opportunities to optimize memory usage through their integration. To address these challenges, we present SMIDT, an efficient MoE inference framework tailored for dynamic Top-K routing. SMIDT features: (1) an adaptive, module-level uneven partitioning strategy to balance computation across PP stages, (2) a memory-aware expert replication scheme (DPMoE) that reduces communication overhead, and (3) a lightweight search algorithm combining binary search and dynamic programming to generate efficient parallelism plans. We implement SMIDT on SGLang, a state-of-the-art LLM inference framework, evaluate it on 32 A40 GPUs and 16 A100 GPUs, and compare with manually tuned parallelism strategies. Experimental results show that, when co-locating prefill and decoding phases, SMIDT achieves 1.20–3.13x throughput improvements for prefill-only tasks and 1.05–1.89x for prefill-decoding tasks. When disaggregating prefill and decoding tasks, SMIDT improves average and P99 time-to-first-token (TTFT) by 1.10–1.17x and 1.21–1.26x, respectively.
Zewen Jin, Shen Fu, Chengjie Tang, Youhui Bai, Jiaan Zhu, Chizheng Fang, Ping Gong 0009, Cheng Li 0001
AAAI7
2026 nnScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training
abstract
As deep neural networks grow, training increasingly relies on handcrafted search spaces for efficient parallelization plans. However, our study shows existing spaces exclude optimal plans for models like AlphaFold2 and large language models with large embedding tables. We propose nScaler-M, a framework for generating efficient parallelization plans for deep learning training. Instead of searching within predefined spaces, nScaler-M empowers domain experts to compose custom search spaces using three primitives,op-trans, op-assign, andop-order, which capture model transformation, spatial assignment, and temporal scheduling. Besides, nScaler-M captures device placement and communication patterns viap-meshandc-mesh, which enhances the accuracy of communication cost estimation, ultimately supporting the search for optimal plans on heterogeneous networks. To avoid space explosion, nScaler-M allows constraints to be applied to these primitives, effectively pruning the search space. With the proposed primitives and constraints, nScaler-M can compose existing search spaces as well as new ones. Experiments show that nScaler-M can find new parallelization plans that achieve up to 3.5× speedup for popular DNN models. Additionally, equipped withp-meshandc-mesh, nScaler-M can discover optimized parallelization plans achieving up to 2.79× higher throughput when training LLaMA-3 models, and introduce acceptable searching overhead.
Jiaan Zhu, Chizheng Fang, Zewen Jin, Youhui Bai, Cheng Li 0001
IEEE Trans. Parallel Distributed Syst.4