EDBT 2026 Demo / reviewers in the wild / expert
Feiwen Zhu
dblp:19/10106
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ScaleFold: Reducing AlphaFold Initial Training Time to 10 HoursabstractAlphaFold2 has been hailed as a breakthrough in protein folding. It can rapidly predict protein structures with lab-grade accuracy. However, its training procedure is prohibitively time-consuming, and gets diminishing benefits from scaling to more compute resources. In this work, we conducted a comprehensive analysis on the AlphaFold training procedure, identified that inefficient communications and overhead-dominated computations were the key factors that prevented the AlphaFold training from effective scaling. We introduced ScaleFold, a systematic training method that incorporated optimizations specifically for these factors. ScaleFold successfully scaled the AlphaFold training to 2080 NVIDIA H100 GPUs with high resource utilization. In the MLPerf HPC v3.0 benchmark, ScaleFold finished the OpenFold benchmark in 7.51 minutes, shown over 6× speedup than the baseline. For training the AlphaFold model from scratch, ScaleFold completed the pretraining in 10 hours, a significant improvement over the seven days required by the original AlphaFold pretraining baseline. Feiwen Zhu, Arkadiusz Nowaczynski, Jie Xin, Michal Marcinkiewicz, Sukru Burc Eryilmaz, Michael Andersch |
DAC | 1 |
| 2024 | Boosting the Convergence of Reinforcement Learning-Based Auto-Pruning Using Historical DataabstractRecently, neural network compression schemes like channel pruning have been widely used to reduce the model size and computational complexity of deep neural networks (DNNs) for applications in power-constrained scenarios, such as embedded systems. Reinforcement learning (RL)-based auto-pruning has been further proposed to automate the DNN pruning process to avoid expensive hand-crafted work. However, the RL-based pruner involves a time-consuming training process, and pruning and evaluating each network comes at high-computational expense. These problems have greatly restricted the real-world application of RL-based auto-pruning. Thus, we propose an efficient auto-pruning framework that solves this problem by taking advantage of the historical data from the previous auto-pruning process. In our framework, we first boost the convergence of the RL-pruner by transfer learning. Then, an augmented transfer learning scheme is proposed to further speed up the training process by improving the transferability. Finally, an assistant learning process is proposed to improve the sample efficiency of the RL agent. The experiments show that our framework can accelerate the auto-pruning process by$1.5\times $–$2.5\times $for ResNet20, and$1.81\times $–$2.375\times $for other neural networks, such as ResNet56, ResNet18, and MobileNet v1. Jiandong Mu, Mengdi Wang 0001, Feiwen Zhu, Jun Yang 0052, Wei Lin 0016, Wei Zhang 0012 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | FastDimeNet++: Training DimeNet++ in 22 minutesabstractRecently, graph neural network (GNN) has shown significant strength in predicting the quantum mechanical properties of molecules. Based on GNN, the DimeNet++ leverages both distance information of atomic pairs and angle information of atomic triplets via message passing mechanism to predict quantum mechanical properties of molecules and has achieved state-of-the-art results. However, there are more than 10 thousand operators in DimeNet++, which results in low GPU utilization and large CPU launch overhead. The extensive time taken for the training of DimeNet++ is a significant drawback. The training period of DimeNet++ exceeds one month on a single NVIDIA A100 GPU. A common method for reducing training time involves employing data parallelism, which equally distributes the global batch across each GPU. However, data-parallel task partitioning, by default, does not consider load imbalance within the batch. This load imbalance leads to considerable synchronization overhead in a multi-GPU setting, reducing the overall efficiency of the parallelism. For the strong-scaling scenario, it results in 32% of the compute resource being wasted. In light of these observations, we propose a novel approach, FastDimeNet++, which delivers high GPU utilization, low CPU overhead, and extensive scalability, achieved through a series of optimization strategies. These include (i) a communication-free load-balancing sampler, (ii) computation graph reconstruction, and (iii) kernel fusion and redundancy bypass. Our experiments demonstrate that FastDimeNet++ achieves a GPU utilization rate of approximately 88% based on a mini-batch size of 4. Furthermore, we scale FastDimeNet++ to 512 GPUs, reaching 2.8 PetaFLOPS. In the MLPerf HPC V1.0, the winning DimeNet++ submission required a total training time of 111.86 minutes, whereas FastDimeNet++ introduced for the MLPerf HPC V2.0 required just 21.93 minutes, demonstrating a significant performance improvement of over 5 ×. Feiwen Zhu, Michal Futrega, Han Bao 0018, Sukru Burc Eryilmaz, Fei Kong, Kefeng Duan, Xinnian Zheng, Nimrod Angel, Matthias Jouanneaux, Maximilian Stadler, Michal Marcinkiewicz, Fung Xie, June Yang, Michael Andersch |
ICPP | 1 |
| 2022 | AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesabstractThis work reveals that memory-intensive computation is a rising performance-critical factor in recent machine learning models. Due to a unique set of new challenges, existing ML optimizing compilers cannot perform efficient fusion under complex two-level dependencies combined with just-in-time demand. They face the dilemma of either performing costly fusion due to heavy redundant computation, or skipping fusion which results in massive number of kernels. Furthermore, they often suffer from low parallelism due to the lack of support for real-world production workloads with irregular tensor shapes. To address these rising challenges, we propose AStitch, a machine learning optimizing compiler that opens a new multi-dimensional optimization space for memory-intensive ML computations. It systematically abstracts four operator-stitching schemes while considering multi-dimensional optimization objectives, tackles complex computation graph dependencies with novel hierarchical data reuse, and efficiently processes various tensor shapes via adaptive thread mapping. Finally, AStitch provides just-in-time support incorporating our proposed optimizations for both ML training and inference. Although AStitch serves as a stand-alone compiler engine that is portable to any version of TensorFlow, its basic ideas can be generally applied to other ML frameworks and optimization compilers. Experimental results show that AStitch can achieve an average of 1.84x speedup (up to 2.73x) over the state-of-the-art Google's XLA solution across five production workloads. We also deploy AStitch onto a production cluster for ML workloads with thousands of GPUs. The system has been in operation for more than 10 months and saves about 20,000 GPU hours for 70,000 tasks per week. Zhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long, Kai Zhu 0004, Feiwen Zhu, Wenyi Zhao, Jun Yang 0052, Jidong Zhai, Shuaiwen Song, Wei Lin 0016 |
ASPLOS | 6 |
| 2018 | Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip
Feiwen Zhu, Jeff Pool, Michael Andersch, Jeremy Appleyard, Fung Xie |
ICLR (Poster) | 1 |
| 2011 | A Parallel Analysis on Scale Invariant Feature Transform (SIFT) Algorithm
Donglei Yang, Feiwen Zhu |
APPT | 3 |