Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shu Pan

dblp:193/7543 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 50% Distributed systems · 38% High-performance computing · 12%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning
distributed training
1.012026
AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026
Parallel and multicore computing
parallelization strategies
1.012026
AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026
Parallel and multicore computing
load balancing
0.312026
AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026
High-performance computing
performance optimization at scale
0.312026
AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026

Methods — techniques the papers use, named apart from their topics

state caching · 1.0memory-aware initialization · 1.0heterogeneity-aware load-balancing estimator · 1.0
YearPublicationVenuePosition
2026 AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training
abstract
Heterogeneous clusters with diverse devices mitigate computational and memory burdens in large language model (LLM) training, yet their inherent resource heterogeneity, characterized by divergent computation, memory, and bandwidth capabilities, renders manual parallelization strategy optimization both challenging and time-intensive. Automatic parallelization is critical for scaling complex workloads across heterogeneous architectures. However, previous methodologies face significant inefficiencies. First, insufficient pruning of the parameter initialization space results in impractically large search spaces. Second, the prevailing automatic parallel search strategies exhibit suboptimal performance in load balancing and resource constraint adaptation. Third, dynamic parallel strategy tuning incurs substantial overhead due to redundant latency calculations for operators with unchanged configurations, leading to unnecessary computational costs. Therefore, insufficient search space pruning, suboptimal load/resource adaptation, and redundant latency computation are identified as the major bottlenecks in our research. To address these challenges, we propose AutoHAAP (Automated Heterogeneity-Aware Asymmetric Partitioning), a novel framework incorporating three core innovations: (1) memory-aware initialization to drastically reduce viable search spaces; (2) a heterogeneity-aware load-balancing estimator that guides resource-efficient configuration search; and (3) state caching mechanisms eliminating redundant latency calculations. Evaluations across GPT3 and Llama3 models of varying scales on both homogeneous and heterogeneous clusters demonstrate that AutoHAAP achieves$\mathbf{0. 6 8}-\mathbf{9 8} \times$search efficiency gains,$\mathbf{6. 5 7 \%} \boldsymbol{-} \mathbf{1 0 6. 9 \%} \boldsymbol{\times}$throughput improvements in homogeneous environments, and$\mathbf{1 0. 1 \%} \boldsymbol{-} 22.28 \% \times$throughput enhancements in heterogeneous setups. These results validate AutoHAAP's effectiveness in distributed LLM training on diverse hardware.
Nana Tang, Shu Pan, Dingding Yu, Zeyue Wang 0003, Mou Sun, Kejie Fu, Fangyu Wang, Yunchuan Chen
HPCA4
2026 HOC: Hierarchical Overlapped Communication Optimization for Parallelism in Distributed Training
Zeyue Wang 0003, Shu Pan, Nana Tang
ISCAS3
2020 Model-driven analysis of mutant fitness experiments improves genome-scale metabolic models of Zymomonas mobilis ZM4
abstract
Genome-scale metabolic models have been utilized extensively in the study and engineering of the organisms they describe. Here we present the analysis of a published dataset from pooled transposon mutant fitness experiments as an approach for improving the accuracy and gene-reaction associations of a metabolic model for Zymomonas mobilis ZM4, an industrially relevant ethanologenic organism with extremely high glycolytic flux and low biomass yield. Gene essentiality predictions made by the draft model were compared to data from individual pooled mutant experiments to identify areas of the model requiring deeper validation. Subsequent experiments showed that some of the discrepancies between the model and dataset were caused by polar effects, mis-mapped barcodes, or mutants carrying both wild-type and transposon disrupted gene copies-highlighting potential limitations inherent to data from individual mutants in these high-throughput datasets. Therefore, we analyzed correlations in fitness scores across all 492 experiments in the dataset in the context of functionally related metabolic reaction modules identified within the model via flux coupling analysis. These correlations were used to identify candidate genes for a reaction in histidine biosynthesis lacking an annotated gene and highlight metabolic modules with poorly correlated gene fitness scores. Additional genes for reactions involved in biotin, ubiquinone, and pyridoxine biosynthesis in Z. mobilis were identified and confirmed using mutant complementation experiments. These discovered genes, were incorporated into the final model, iZM4_478, which contains 747 metabolic and transport reactions (of which 612 have gene-protein-reaction associations), 478 genes, and 616 unique metabolites, making it one of the most complete models of Z. mobilis ZM4 to date. The methods of analysis that we applied here with the Z. mobilis transposon mutant dataset, could easily be utilized to improve future genome-scale metabolic reconstructions for organisms where these, or similar, high-throughput datasets are available.
Wai Kit Ong, Dylan K. Courtney, Shu Pan, Ramon Bonela Andrade, Patricia J. Kiley, Brian F. Pfleger, Jennifer L. Reed
PLoS Comput. Biol.3
2017 Pedestrian tracking for infrared image sequence based on trajectory manifold of spatio-temporal slice
Tao Yang 0010, Dongmei Fu, Shu Pan
Multim. Tools Appl.3