Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiaan Zhu

dblp:274/3515 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 67% Deep learning architectures and training · 17% Language models and text generation · 15%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 64% GPUs and heterogeneous computing · 19% Distributed systems · 17%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.012026
nnScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training · IEEE Trans. Parallel Distributed Syst. 2026
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026
Machine learning › Efficient and distributed learning
model inference
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
1.012026
SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing · AAAI 2026
Natural language and speech › Language models and text generation › large language model
large language model training and inference
0.912025
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference · AAAI 2025
Distributed systems › distributed machine learning
distributed training
0.312025
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference · AAAI 2025

Methods — techniques the papers use, named apart from their topics

search space pruning · 2.0placement-aware cost estimation · 2.0constraint-guided search · 2.0projection · 1.7all-to-all communication · 1.7tensor parallelism · 1.0expert parallelism · 1.0dynamic programming · 1.0binary search · 1.0mixture-of-experts · 0.9mixture of experts · 0.9
YearPublicationVenuePosition
2026 SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing
abstract
To accelerate Mixture-of-Experts (MoE) inference, the hybrid parallelism paradigm is first applying pipeline parallelism (PP) to vertically divide the model into stages, with each stage further divided horizontally using tensor or expert parallelism. On the algorithm side, dynamic Top-K routing reduces computation by activating fewer experts per token on average. In this paper, we explore the application of dynamic Top-K routing to PP-enabled MoE inference, aiming to fully unleash their combined potential. We identify key performance bottlenecks arising from Top-K value variation across layers, which conflicts with PP's typically uniform stage partitioning, as well as opportunities to optimize memory usage through their integration. To address these challenges, we present SMIDT, an efficient MoE inference framework tailored for dynamic Top-K routing. SMIDT features: (1) an adaptive, module-level uneven partitioning strategy to balance computation across PP stages, (2) a memory-aware expert replication scheme (DPMoE) that reduces communication overhead, and (3) a lightweight search algorithm combining binary search and dynamic programming to generate efficient parallelism plans. We implement SMIDT on SGLang, a state-of-the-art LLM inference framework, evaluate it on 32 A40 GPUs and 16 A100 GPUs, and compare with manually tuned parallelism strategies. Experimental results show that, when co-locating prefill and decoding phases, SMIDT achieves 1.20–3.13x throughput improvements for prefill-only tasks and 1.05–1.89x for prefill-decoding tasks. When disaggregating prefill and decoding tasks, SMIDT improves average and P99 time-to-first-token (TTFT) by 1.10–1.17x and 1.21–1.26x, respectively.
Zewen Jin, Shen Fu, Chengjie Tang, Youhui Bai, Jiaan Zhu, Chizheng Fang, Ping Gong 0009, Cheng Li 0001
AAAI6
2026 nnScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training
abstract
As deep neural networks grow, training increasingly relies on handcrafted search spaces for efficient parallelization plans. However, our study shows existing spaces exclude optimal plans for models like AlphaFold2 and large language models with large embedding tables. We propose nScaler-M, a framework for generating efficient parallelization plans for deep learning training. Instead of searching within predefined spaces, nScaler-M empowers domain experts to compose custom search spaces using three primitives,op-trans, op-assign, andop-order, which capture model transformation, spatial assignment, and temporal scheduling. Besides, nScaler-M captures device placement and communication patterns viap-meshandc-mesh, which enhances the accuracy of communication cost estimation, ultimately supporting the search for optimal plans on heterogeneous networks. To avoid space explosion, nScaler-M allows constraints to be applied to these primitives, effectively pruning the search space. With the proposed primitives and constraints, nScaler-M can compose existing search spaces as well as new ones. Experiments show that nScaler-M can find new parallelization plans that achieve up to 3.5× speedup for popular DNN models. Additionally, equipped withp-meshandc-mesh, nScaler-M can discover optimized parallelization plans achieving up to 2.79× higher throughput when training LLaMA-3 models, and introduce acceptable searching overhead.
Jiaan Zhu, Chizheng Fang, Zewen Jin, Youhui Bai, Cheng Li 0001
IEEE Trans. Parallel Distributed Syst.1
2025 BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
abstract
The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation resources. Recently, a fine-grained DeepSeekMoE structure is proposed, which can further improve the computing efficiency of MoE without performance degradation. However, the All-to-All communication introduced by MoE has become a bottleneck, especially for the fine-grained structure, which typically involves and activates more experts, hence contributing to heavier communication overhead. In this paper, we propose a novel MoE structure named BigMac, which is also fine-grained but with high communication efficiency. The innovation of BigMac is mainly due to that we abandon the Communicate-Descend-Ascend-Communicate (CDAC) manner used by fine-grained MoE, which leads to the All-to-All communication always taking place at the highest dimension. Instead, BigMac designs an efficient Descend-Communicate-Communicate-Ascend (DCCA) manner. Specifically, we add a descending and ascending projection at the entrance and exit of the expert, respectively, which enables the communication to perform at a very low dimension. Furthermore, to adapt to DCCA, we re-design the structure of small experts, ensuring that the expert in BigMac has enough complexity to address tokens. Experimental results show that BigMac achieves comparable or even better model quality than fine-grained MoEs with the same number of experts and a similar number of total parameters. Equally importantly, BigMac reduces the end-to-end latency by up to 3.09 x for training and increases the throughput by up to 3.11 x for inference on state-of-the-art AI computing frameworks including Megatron, Tutel, and DeepSpeed-Inference.
Zewen Jin, Jiaan Zhu, Hongrui Zhan, Youhui Bai, Zhenyu Ming
AAAI3
2024 CEUS-SAM: Cross-Modal Prompt-Based SAM Network for Breast CEUS Image Segmentation
abstract
The precise segmentation of lesions in contrast-enhanced ultrasound (CEUS) videos, especially during the peak enhancement phase, is crucial for early breast cancer diagnosis. However, the dynamic contrast patterns and subtle differences in CEUS images challenge traditional methods. To overcome this, we propose the CEUS-SAM network, a deep learning framework leveraging the Segment Anything Model (SAM) for enhanced lesion segmentation. Our approach first trains on conventional ultrasound (US) data, generating segmentation masks as prompts for CEUS images. A key innovation, the Image Fusion Module (IFM), integrates cross-modal and multi-scale features from US and CEUS, improving tissue differentiation and lesion detection. The CEUS-SAM network significantly reduces manual effort with single-point prompts and minimizes inter-observer variability. Using a breast CEUS dataset with 135 video sequences, our method achieves a Dice score of 78.6% and an IoU score of 66.6%. The code and dataset are available at https://github.com/2284650586/CEUS-SAM.
Min Xu 0003, Ximiao Zhang, Sihua Niu, Jiaan Zhu, Xiuzhuang Zhou
BIBM5