EDBT 2026 Demo / reviewers in the wild / expert
Mou Sun
dblp:365/5987
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 63% Parallel and multicore computing · 30% High-performance computing · 7% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 44% Trustworthy machine learning · 44% Language models and text generation · 13% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems › distributed machine learning
distributed training |
1.9 | 2 | 2026 | AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026 MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization · NeurIPS 2025 |
Parallel and multicore computing
parallelization strategies |
1.0 | 1 | 2026 | AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026 |
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness
fault tolerance |
0.9 | 1 | 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization · NeurIPS 2025 |
Distributed systems › fault tolerance › fault-tolerant distributed systems
fault-tolerant training |
0.9 | 1 | 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization · NeurIPS 2025 |
Parallel and multicore computing
load balancing |
0.3 | 1 | 2026 | AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026 |
High-performance computing
performance optimization at scale |
0.3 | 1 | 2026 | AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training · HPCA 2026 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
recomputation · 1.7low-rank gradient approximation · 1.7state caching · 1.0memory-aware initialization · 1.0heterogeneity-aware load-balancing estimator · 1.0skip-connection · 0.9skip connections · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM TrainingabstractHeterogeneous clusters with diverse devices mitigate computational and memory burdens in large language model (LLM) training, yet their inherent resource heterogeneity, characterized by divergent computation, memory, and bandwidth capabilities, renders manual parallelization strategy optimization both challenging and time-intensive. Automatic parallelization is critical for scaling complex workloads across heterogeneous architectures. However, previous methodologies face significant inefficiencies. First, insufficient pruning of the parameter initialization space results in impractically large search spaces. Second, the prevailing automatic parallel search strategies exhibit suboptimal performance in load balancing and resource constraint adaptation. Third, dynamic parallel strategy tuning incurs substantial overhead due to redundant latency calculations for operators with unchanged configurations, leading to unnecessary computational costs. Therefore, insufficient search space pruning, suboptimal load/resource adaptation, and redundant latency computation are identified as the major bottlenecks in our research. To address these challenges, we propose AutoHAAP (Automated Heterogeneity-Aware Asymmetric Partitioning), a novel framework incorporating three core innovations: (1) memory-aware initialization to drastically reduce viable search spaces; (2) a heterogeneity-aware load-balancing estimator that guides resource-efficient configuration search; and (3) state caching mechanisms eliminating redundant latency calculations. Evaluations across GPT3 and Llama3 models of varying scales on both homogeneous and heterogeneous clusters demonstrate that AutoHAAP achieves$\mathbf{0. 6 8}-\mathbf{9 8} \times$search efficiency gains,$\mathbf{6. 5 7 \%} \boldsymbol{-} \mathbf{1 0 6. 9 \%} \boldsymbol{\times}$throughput improvements in homogeneous environments, and$\mathbf{1 0. 1 \%} \boldsymbol{-} 22.28 \% \times$throughput enhancements in heterogeneous setups. These results validate AutoHAAP's effectiveness in distributed LLM training on diverse hardware. Nana Tang, Shu Pan, Dingding Yu, Zeyue Wang 0003, Mou Sun, Kejie Fu, Fangyu Wang, Yunchuan Chen |
HPCA | 7 |
| 2025 | MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationabstractAs distributed optimization scales to meet the demands of Large Language Model (LLM) training, hardware failures become increasingly non-negligible. Existing fault-tolerant training methods often introduce significant computational or memory overhead, demanding additional resources. To address this challenge, we propose **Me**mory- and **C**omputation- **e**fficient **F**ault-tolerant **O**ptimization (**MeCeFO**), a novel algorithm that ensures robust training with minimal overhead. When a computing node fails, MeCeFO seamlessly transfers its training task to a neighboring node while employing memory- and computation-efficient algorithmic optimizations to minimize the extra workload imposed on the neighboring node handling both tasks. MeCeFO leverages three key algorithmic designs: (i) Skip-connection, which drops the multi-head attention (MHA) module during backpropagation for memory- and computation-efficient approximation; (ii) Recomputation, which reduces activation memory in feedforward networks (FFNs); and (iii) Low-rank gradient approximation, enabling efficient estimation of FFN weight matrix gradients. Theoretically, MeCeFO matches the convergence rate of conventional distributed training, with a rate of $\mathcal{O}(1/\sqrt{nT})$, where $n$ is the data parallelism size and $T$ is the number of iterations. Empirically, MeCeFO maintains robust performance under high failure rates, incurring only a 4.18\% drop in throughput, demonstrating $5.0\times$ to $6.7\times$ greater resilience than previous SOTA approaches. Rizhen Hu, Mou Sun, Binhang Yuan, Kun Yuan 0001 |
NeurIPS | 4 |