Chong Li 0003

dblp:50/3011-3 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-4160-7170ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN Training
abstract
Pipeline parallelism and activation recomputation are widely adopted optimization techniques, among others, to scale DNN training on large accelerator clusters. However, as DNNs grow in complexity and heterogeneity, it becomes increasingly difficult to determine the optimal combination of pipeline partitioning and recomputation strategies. Existing solutions either propose manual optimization approaches that do not scale or automated approaches that explore only a subset of optimization possibilities due to an explosion of search space. In this paper, we present BMPipe, a bubble-memory co-optimization planner that holistically optimizes computation imbalance, memory under utilization, redundant computation, and schedulinginduced preparation time. At its core, BMPipe uses symbolic representations that unify computation, memory, and bubbles into a single model that is solved by using an ILP-based planner. Using BMPipe, we perform a thorough experimental evaluation where we train several large, state-of-the-art DNN models on a 16K-NPU cluster. We show that BMPipe achieves up to$1.36 \times$speedup compared to the state-of-the-art solution Megatron. Against automatic planners PipeDream, Merak and AdaPipe,, it yields as$1.27 \times$speed-up. In addition, BMPipe boosts peak device-memory utilization by$\mathbf{1. 4 2} \times$compared with Megatron.
Ruiwen Wang, Chong Li 0003, Thibaut Tachon, Raja Appuswamy, Teng Su
CLUSTER2
2025 ManuMatic: Strategy Injection for Robust Automatic Hybrid Parallelism in Distributed DNN Training
Ruiwen Wang, Chong Li 0003, Raja Appuswamy, Yujie Yuan
NPC (2)2
2024 An Efficient and Scalable Approach to Build Co-occurrence Matrix for DNN's Embedding Layer
abstract
Embedding is a crucial step for deep neural networks. Datasets, from different applications, with different structures, can all be processed through an embedding layer and transformed into a dense matrix. The transformation must minimize both the loss of information and the redundancy of data. Extracting appropriate data features ensures the efficiency of the transformation. The co-occurrence matrix is an excellent way of representing the links between elements in a dataset. However, the dataset size becomes a problem in terms of computation power and memory footprint for using the co-occurrence matrix.
Quentin R. Petit, Chong Li 0003, Nahid Emad
ICS2
2022 Parallelizing Neural Network Models Effectively on GPU by Implementing Reductions Atomically
abstract
Due to the missing of a good orchestration of loop transformations, existing optimizing compilers for deploying neural networks on GPU either parallelize reductions ineffectively or miss the fusion opportunities with other operators. Neural network models thus exhibit sub-optimal performance on GPU. We present a practical approach called Panamera for the effective parallelization of reductions in neural networks on GPU. Panamera first leverages loop coalescing to flatten the loop dimensions of reductions, converting all reduction operators into canonical forms eligible for the polyhedral model. Next, Panamera uses polyhedral transformations to reduce the data movements caused by unfused reductions and perform multi-block hardware binding not considered by many compilers. Finally, Panamera embeds a highly optimized routine implemented using GPU atomic instructions, further improving the performance of neural network models while guaranteeing the correctness of parallel reductions. The experimental results demonstrate the effectiveness of our approach: for single operators our code obtains a mean speedup of 33.7×, 3.5×, 5.4× and 9.6× over cuDNN, CUB, TVM and Ansor, for sub-graphs our approach outperforms cuDNN, TVM and Ansor by 9.5×, 2.6× and 2.7×, and for end-to-end workloads, a tensor compiler integrated with our approach outperforms them by 122.5%, 19.3% and 15.2%.
Jie Zhao 0002, Cédric Bastoul, Yanzhi Yi, Wang Nie, Renwei Zhang, Zhen Geng, Chong Li 0003, Thibaut Tachon, Zhiliang Gan
PACT8
2022 Distributed and Parallel Sparse Computing for Very Large Graph Neural Networks
abstract
Deep learning (DL) requires high-performance processing on big data. Graph Neural Networks, a challenging topic in DL using linear algebra methods, need algorithmic solutions to efficiently assign and process graph data on modern distributed and parallel machines, which are considered with mixed arithmetic and various types of tensor/matrix accelerators. Determining compression techniques for the graph’s sparse data structures is one of the key elements.Our first objective is to design and implement a reusable parallel numerical library to resolve large neural network graphs. Our design strategy is drawn on a component-based approach and targets maximum code reuse in various parallel contexts while allowing for performance optimization. The solution could be later integrated into a DL framework like MindSpore.
Quentin R. Petit, Chong Li 0003, Nahid Emad
IEEE Big Data2
2022 Enhancing Graph Convolutional Networks by Topology Sampling
abstract
Graph Neural Networks (GNNs) play a very important role today. It does analyze not only the graph data itself, but also the data connectivity of the graph. The quality of a GNN is thus altered by the result of extracted graph structure information. The extraction could be enhanced by GNN model design or directly from the training dataset with a GNN-decoupled method. In this paper, we propose RankedDrop, a new sampling method to improve the extraction of graph structure information. This approach is based on droppingout technique, and it adopts a spatial-aware selection of edges to drop. It takes into account structure information of the graph to control the dropping-out, and its random selection of edges to be dropped is under the control of a probability generated with respect to graph’s topological importance. Our experiments point out that RankedDrop provides high-quality and robust training results compared to the leading solutions. Furthermore, RankedDrop could be a framework plugin and combined with GNN model improvements to maximize GNN quality. Furthermore, RankedDrop could be a plugin for AI frameworks like MindSpore and combined with GNN model improvements to maximize GNN quality.
Quentin R. Petit, Chong Li 0003, Serge G. Petiton, Kelun Chai, Nahid Emad
IEEE Big Data2
2021 Efficient and Systematic Partitioning of Large and Deep Neural Networks for Parallelization
Chong Li 0003, Thibaut Tachon, Sébastien Limet, Sophie Robert 0001
Euro-Par2
2012 Implementation of Data-Parallel Skeletons: A Case Study Using a Coarse-Grained Hierarchical Model
abstract
Writing parallel programs is known to be notoriously difficult. Often programmers do not want to reason about message-passing algorithms and only want to combine existing high-level patterns to produce their parallel program. This is the algorithmic skeletons approach to parallel programming. It improves reliability and clarity of source code. But skeletons can be insufficient when complicated communication schemes are needed. Expressing skeletons in a more general and low level language in the form of a library seems to be a good compromise between simplicity and expressive power. In this article, we present a coarsed-grained implementation using a hierarchical model of a set of data-parallel skeletons. Programming experiments and benchmarks complete the article.
Chong Li 0003, Frédéric Gava, Gaétan Hains
ISPDC1