EDBT 2026 Demo / reviewers in the wild / expert
Teng Su
dblp:118/8937
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN TrainingabstractPipeline parallelism and activation recomputation are widely adopted optimization techniques, among others, to scale DNN training on large accelerator clusters. However, as DNNs grow in complexity and heterogeneity, it becomes increasingly difficult to determine the optimal combination of pipeline partitioning and recomputation strategies. Existing solutions either propose manual optimization approaches that do not scale or automated approaches that explore only a subset of optimization possibilities due to an explosion of search space. In this paper, we present BMPipe, a bubble-memory co-optimization planner that holistically optimizes computation imbalance, memory under utilization, redundant computation, and schedulinginduced preparation time. At its core, BMPipe uses symbolic representations that unify computation, memory, and bubbles into a single model that is solved by using an ILP-based planner. Using BMPipe, we perform a thorough experimental evaluation where we train several large, state-of-the-art DNN models on a 16K-NPU cluster. We show that BMPipe achieves up to$1.36 \times$speedup compared to the state-of-the-art solution Megatron. Against automatic planners PipeDream, Merak and AdaPipe,, it yields as$1.27 \times$speed-up. In addition, BMPipe boosts peak device-memory utilization by$\mathbf{1. 4 2} \times$compared with Megatron. Ruiwen Wang, Chong Li 0003, Thibaut Tachon, Raja Appuswamy, Teng Su |
CLUSTER | 5 |
| 2025 | Efficient and Automatic 3D Parallelism Strategies Search via Contrastive Reinforcement Learning Pretrained Neural Networks
Jie Ou, Xiaowang Li, Yueming Chen, Jiahong Qian, Teng Su, Wenhong Tian |
DSAA | 5 |
| 2025 | Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and Optimization
Zhibin Wang 0002, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Bingqiang Wang, Yonghong Tian 0001, Yan Zhang 0002, Hui Wang 0030, Fuchun Wei, Boquan Sun, Bin She, Teng Su, Yaoyuan Wang, Guyue Liu |
USENIX ATC | 17 |
| 2025 | EfficientMoE: Optimizing Mixture-of-Experts Model Training With Adaptive Load BalanceabstractMixture-of-Experts (MoE) efficiently trains large models by using sparse activation to lower costs, selecting a few experts based on data characteristics. However, it faces challenges such as All-to-All communication overhead and load imbalance, with most optimizations targeting dynamic graphs rather than the more efficient static graphs. This study identifies two key challenges in training MoE on static graphs: 1) excessive All-to-All communication (up to 75% of iteration time) and load imbalance (70% of tokens handled by two experts) between experts due to the sparse structure of the MoE model and the token distribution; and 2) inefficient zero-padding for static shapes, leading to unnecessary computational overhead(wasting approximately 50% of resources). Thus, EfficientMoE, a scheduling method based on expert load and data characteristics, is introduced. EfficientMoE first designs a sampler to collect real-time information about token distribution, expert load, etc. It constructs a load prediction model to evaluate expert load. Subsequently, EfficientMoE proposes a dynamic schedule strategy for experts with evaluated expert load, reducing All-to-All communication and addressing load-balancing issues. Additionally, an expert capacity model is proposed to set different capacities for replicas of hot experts before static graph compilation, minimizing computation and storage overhead caused by significant padding. This study implements EfficientMoE in MindSpore and uses 32 Ascend AI accelerators to train an MoE model with 21 billion parameters and evaluate its validity. EfficientMoE demonstrated an improvement of 30% in model training time, approximately 12% reduction in communication time, and saved 35% computational resources across different clusters, compared with Switch transformers, and the Fastermoe method for static graphs. Chengchuang Huang, Yipeng Mei, Lifu Zhang 0004, Teng Su |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | Cross-Search With Improved Multi-Dimensional Dichotomy-Based Joint Optimization for Distributed Parallel Training of DNNabstractistributedistributedD parallel training of large-scale deep neural networks (DNN) has attracted the attentions of both artificial intelligence and high-performance distributed computing. One of efficient approaches is the micro-batch-based pipeline parallelism (MBPP), e.g., GPipe and Terapipe. Based on the MBPP, we establish a time-cost model with the basic time function of layers, which considers computing time and communication time simultaneously as well as considers they are nonlinear with the amount of input data. Focusing on the jointly optimal solutions of network division and data partition, we propose a Cross-Search algorithm with Improved Multi-dimensional Dichotomy (CSIMD). Through theoretical derivation, we prove improved multi-dimensional dichotomy (IMD) has appreciable theoretical optimality and linear computational complexity significantly faster than the state-of-the-art methods including dynamic programming and recursive algorithm. Extensive experiments on both CNN-based and transformer-based neural networks demonstrate our proposed CSIMD can obtain optimal network division and data partition schemes under MBPP. On average, the training speeds of CSIMD in CNN- and transformer-based DNNs are respectively (2.0, 2.5)× and (2.66, 5.48)× of (MBPP-R, MBPP-E). Yiqin Fu, Haocheng Lan, Yuanlun Xie, Wenhong Tian, Rajkumar Buyya, Jianhong Qian, Teng Su |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2024 | CSIMD: Cross-Search Algorithm with Improved Multi-dimensional Dichotomy for Micro-Batch-Based Pipeline Parallel Training in DNN
Haocheng Lan, Yuanlun Xie, Wenhong Tian, Jiahong Qian, Teng Su |
Euro-Par (2) | 6 |
| 2023 | CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-XabstractLarge pre-trained code generation models, such as OpenAI Codex, can generate syntax-and function-correct code, making the coding of programmers more productive. In this paper, we introduce CodeGeeX, a multilingual model with 13 billion parameters for code generation. CodeGeeX is pre-trained on 850 billion tokens of 23 programming languages as of June 2022. Our extensive experiments suggest that CodeGeeX outperforms multilingual code models of similar scale for both the tasks of code generation and translation on HumanEval-X. Building upon HumanEval (Python only), we develop the HumanEval-X benchmark for evaluating multilingual models by hand-writing the solutions in C++, Java, JavaScript, and Go. In addition, we build CodeGeeX-based extensions on Visual Studio Code, JetBrains, and Cloud Studio, generating 8 billion tokens for tens of thousands of active users per week. Our user study demonstrates that CodeGeeX can help to increase coding efficiency for 83.4% of its users. Finally, CodeGeeX is publicly accessible since Sep. 2022, we open-sourced its code, model weights, API, extensions, and HumanEval-X at https://github.com/THUDM/CodeGeeX. Qinkai Zheng, Xu Zou 0001, Yuxiao Dong, Shan Wang 0023, Lei Shen 0002, Andi Wang 0003, Yang Li 0074, Teng Su, Zhilin Yang 0001, Jie Tang 0001 |
KDD | 11 |
| 2023 | A Survey on Auto-Parallelism of Large-Scale Deep Learning TrainingabstractDeep learning (DL) has gained great success in recent years, leading to state-of-the-art performance in research community and industrial fields like computer vision and natural language processing. One of the reasons for this success is the huge amount parameters adopted in DL models. However, it is impractical to train a moderately large model with a large number of parameters on a typical single device. Thus, It is necessary to train DL models in clusters with distributed training algorithms. However, traditional distributed training algorithms are usually sub-optimal and highly customized, which owns the drawbacks to train large-scale DL models in varying computing clusters. To handle the above problem, researchers propose auto-parallelism, which is promising to train large-scale DL models efficiently and practically in various computing clusters. In this survey, we perform a broad and thorough investigation on challenges, basis, and strategy searching methods of auto-parallelism in DL training. First, we abstract basic parallelism schemes with their communication cost and memory consumption in DL training. Further, we analyze and compare a series of current auto-parallelism works and investigate strategies and searching methods which are commonly used in practice. At last, we discuss several trends in auto-parallelism which are promising in further research. Peng Liang 0017, Xiaoda Zhang, Youhui Bai, Teng Su, Zhiquan Lai, Linbo Qiao, Dongsheng Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-ParallelismabstractEffective parallelization strategies are crucial for the performance of distributed deep neural network (DNN) training. Recently, several methods have been proposed to search parallelization strategies but they all optimize a single objective (e.g., execution time, memory consumption) and produce only one strategy. We proposeFrontier Tracking(FT), an efficient algorithm that findsa set of Pareto-optimal parallelization strategiesto explore the best trade-off among different objectives. FT can minimize the memory consumption when the number of devices is limited and fully utilize additional resources to reduce the execution time. Based onFT, we develop a user-friendly system, calledTensorOpt, which allows users to run their distributed DNN training jobs without caring the details about searching and coding parallelization strategies. Experimental results show that TensorOpt is more flexible in adapting to resource availability compared with existing frameworks. Zhenkun Cai, Xiao Yan 0002, Kaihao Ma, Yidi Wu 0001, James Cheng, Teng Su, Fan Yu 0004 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2012 | A Family of Fast Hadamard-Fourier Transform AlgorithmsabstractIn this letter, we present a family of fast Hadamard-Fourier transform algorithms which combined Walsh Hadamard and discrete Fourier transforms into one single algorithm. These family algorithms can be computed in butterfly structure, and have similar sparse matrix factorization in each stage, and have less computation stages than the sum of Walsh Hadamard and discrete Fourier transforms. We factorize the algorithms with regular sparse matrices for every stage in radix-R mode, where R is power of 2. Teng Su, Feng Yu 0003 |
IEEE Signal Process. Lett. | 1 |