EDBT 2026 Demo / reviewers in the wild / expert
Jiarui Fang
dblp:147/0860
· DBLP profile ↗
16ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-6724-2763ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers InferenceabstractThis paper presents PipeFusion, an innovative parallel methodology to tackle the high latency issues associated with generating high-resolution images using diffusion transformers (DiTs) models. PipeFusion partitions images into patches and the model layers across multiple GPUs. It employs a patch-level pipeline parallel strategy to orchestrate communication and computation efficiently. By capitalizing on the high similarity between inputs from successive diffusion steps, PipeFusion reuses one-step stale feature maps to provide context for the current pipeline step. This approach notably reduces communication costs compared to existing DiTs inference parallelism, including tensor parallel, sequence parallel and DistriFusion. PipeFusion enhances memory efficiency through parameter distribution across devices, ideal for large DiTs like Flux.1. Experimental results demonstrate that PipeFusion achieves state-of-the-art performance on 8$\times$L40 PCIe GPUs for Pixart, Stable-Diffusion 3, and Flux.1 models. Our Source code is available at \url{https://github.com/xdit-project/xDiT}. Jiarui Fang, Jinzhe Pan, Aoyu Li, Xibo Sun |
NeurIPS | 1 |
| 2024 | Brief Analysis of False Data Injection Attacks Based on Two Data Modalities in IoTs-based Solar Insecticidal LampsabstractOwing to the swift growth in Internet of Things technology, Solar Insecticidal Lamps (SIL) are widely deployed in smart agriculture scenarios. At the same time, various safety and security issues of Solar Insecticidal Lamps - Internet of Things (SIL-IoTs) have attracted considerable attention within the academic community. This paper presents a specific False Data Injection Attack (FDIA) case in SIL-IoTs. Unlike mature urban or commercial IoTs applications, such as smart grids, there is no direct precedent for attack and detection models against FDIA in the SIL-IoTs. Moreover, as outdoor agricultural IoTs are susceptible to numerous uncontrollable environmental factors, data-driven detection methods may fail to detect FDIA in such scenarios. Therefore, we report the FDIA in SIL-IoTs and briefly analyze the detection means from two aspects: insecticidal data and sound signals by conducting experiments. Qin Su, Lei Shu 0001, Xing Yang 0001, Zitian Jiang, Jiarui Fang, Huihsin Chin |
INDIN | 6 |
| 2024 | FastFold: Optimizing AlphaFold Training and Inference on GPU ClustersabstractProtein structure prediction helps to understand gene translation and protein function, which is of growing interest and importance in structural biology. The AlphaFold model, which used transformer architecture to achieve atomic-level accuracy in protein structure prediction, was a significant breakthrough. However, training and inference of AlphaFold model are challenging due to its high computation and memory cost. In this work, we present FastFold, an efficient implementation of AlphaFold for both training and inference. We propose Dynamic Axial Parallelism (DAP) as a novel model parallelism method. Additionally, we have implemented a series of low-level optimizations aimed at reducing communication, computation, and memory costs. These optimizations include Duality Async Operations, highly optimized kernels, and AutoChunk (an automated search algorithm finds the best chunk strategy to reduce memory peaks). Experimental results show that FastFold can efficiently scale to more GPUs using DAP and reduces overall training time from 11 days to 67 hours and achieves 7.5 ~ 9.5× speedup for long-sequence inference. Furthermore, AutoChunk can reduce memory cost by over 80% during inference by automatically partitioning the intermediate tensors during the computation. Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Ruidong Wu, Jian Peng 0001, Yang You 0001 |
PPoPP | 4 |
| 2023 | Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel TrainingabstractThe success of Transformer models has pushed the deep learning model scale to billions of parameters, but the memory limitation of a single GPU has led to an urgent need for training on multi-GPU clusters. However, the best practice for choosing the optimal parallel strategy is still lacking, as it requires domain expertise in both deep learning and parallel computing. The Colossal-AI system addressed the above challenge by introducing a unified interface to scale your sequential code of model training to distributed environments. It supports parallel training methods such as data, pipeline, tensor, and sequence parallelism and is integrated with heterogeneous training and zero redundancy optimizer. Compared to the baseline system, Colossal-AI can achieve up to 2.76 times training speedup on large-scale models. Shenggui Li, Hongxin Liu, Zhengda Bian, Jiarui Fang, Haichen Huang, Yang You 0001 |
ICPP | 4 |
| 2023 | Parallel Training of Pre-Trained Models via Chunk-Based Dynamic Memory ManagementabstractThe pre-trained model (PTM) is revolutionizing Artificial Intelligence (AI) technology. However, the hardware requirement of PTM training is prohibitively high, making it a game for a small proportion of people. Therefore, we proposed PatrickStar system to lower the hardware requirements of PTMs and make them accessible to everyone. PatrickStar uses the CPU-GPU heterogeneous memory space to store the model data. Different from existing works, we organize the model data in memory chunks and dynamically distribute them in the heterogeneous memory. Guided by the runtime memory statistics collected in a warm-up iteration, chunks are orchestrated efficiently in heterogeneous memory and generate lower CPU-GPU data transmission volume and higher bandwidth utilization. Symbiosis with the Zero Redundancy Optimizer, PatrickStar scales to multiple GPUs on multiple nodes. The system can train tasks on bigger models and larger batch sizes, which cannot be accomplished by existing works. Experimental results show that PatrickStar extends model scales 2.27 and 2.5 times of DeepSpeed, and exhibits significantly higher execution speed. PatricStar also successfully runs the 175B GPT3 training task on a 32 GPU cluster. Our code is available athttps://github.com/Tencent/PatrickStar. Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu 0038, Jie Zhou 0016, Yang You 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingabstractLarge-scale pretrained language models have achieved SOTA results on NLP tasks.However, they have been shown vulnerable to adversarial attacks especially for logographic languages like Chinese.In this work, we propose ROCBERT: a pretrained Chinese Bert that is robust to various forms of adversarial attacks like word perturbation, synonyms, typos, etc.It is pretrained with the contrastive learning objective which maximizes the label consistency under different synthesized adversarial examples.The model takes as input multimodal information including the semantic, phonetic and visual features.We show all these features are important to the model robustness since the attack can be performed in all the three forms.Across 5 Chinese NLU tasks, ROCBERT outperforms strong baselines under three blackbox adversarial algorithms without sacrificing the performance on clean testset.It also performs the best in the toxic content detection task under human-made attacks. * Equal contribution. Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Tuo Ji, Jiarui Fang, Jie Zhou 0016 |
ACL (1) | 6 |
| 2021 | TurboTransformers: an efficient GPU serving system for transformer modelsabstractThe transformer is the most critical algorithm innovation of the Nature Language Processing (NLP) field in recent years. Unlike the Recurrent Neural Network (RNN) models, transformers are able to process on dimensions of sequence lengths in parallel, therefore leads to better accuracy on long sequences. However, efficient deployments of them for online services in data centers equipped with GPUs are not easy. First, more computation introduced by transformer structures makes it more challenging to meet the latency and throughput constraints of serving. Second, NLP tasks take in sentences of variable length. The variability of input dimensions brings a severe problem to efficient memory management and serving optimization. Jiarui Fang, Yang Yu 0038, Chengduo Zhao, Jie Zhou 0016 |
PPoPP | 1 |
| 2020 | Efficient AES implementation on Sunway TaihuLight supercomputer: A systematic approach
Liandeng Li, Jiarui Fang, Jinlei Jiang, Lin Gan 0001, Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002 |
J. Parallel Distributed Comput. | 2 |
| 2019 | swATOP: Automatically Optimizing Deep Learning Operators on SW26010 Many-Core ProcessorabstractAchieving an optimized mapping of Deep Learning (DL) operators to new hardware architectures is the key to building a scalable DL system. However, handcrafted optimization involves huge engineering efforts, due to the variety of DL operator implementations and complex programming skills. Targeting the innovative many-core processor SW26010 adopted by the 3rd fastest supercomputer Sunway TaihuLight, an end-to-end automated framework called swATOP is presented as a more practical solution for DL operator optimization. Arithmetic intensive DL operators are expressed into an auto-tuning-friendly form, which is based on tensorized primitives. By describing the algorithm of a DL operator using our domain specific language (DSL), swATOP is able to derive and produce an optimal implementation by separating hardware-dependent optimization and hardware-agnostic optimization. Hardware-dependent optimization is encapsulated in a set of tensorized primitives with sufficient utilization of the underlying hardware features. The hardware-agnostic optimization contains a scheduler, an intermediate representation (IR) optimizer, an auto-tuner, and a code generator. These modules cooperate to perform an automatic design space exploration, to apply a set of programming techniques, to discover a near-optimal solution, and to generate the executable code. Our experiments show that swATOP is able to bring significant performance improvement on DL operators in over 88% of cases, compared with the best-handcrafted optimization. Compared to a black-box autotuner, the tuning and code generation time can be reduced to minutes from days using swATOP. Jiarui Fang, Wenlai Zhao, Jinzhe Yang, Long Wang 0014, Lin Gan 0001, Haohuan Fu, Guangwen Yang 0002 |
ICPP | 2 |
| 2019 | RedSync: Reducing synchronization bandwidth for distributed deep learning training system
Jiarui Fang, Haohuan Fu, Guangwen Yang 0002, Cho-Jui Hsieh |
J. Parallel Distributed Comput. | 1 |
| 2018 | swCaffe: A Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLightabstractThis paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core heterogeneous architecture, with 40,960 SW26010 processors connected through a customized communication network.First, we point out some insightful principles to fully exploit the performance of the innovative many-core architecture.Second, we propose a set of optimization strategies for redesigning a variety of neural network layers based on Caffe.Third, we put forward a topology-aware parameter synchronization scheme to scale the synchronous Stochastic Gradient Descent (SGD) method to multiple processors efficiently.We evaluate our framework by training a variety of widely used neural networks with the ImageNet dataset.On a single node, swCaffe can achieve 23%˜119% overall performance compared with Caffe running on K40m GPU.As compared with the Caffe on CPU, swCaffe runs 3.04˜7.84xfaster on all the networks.Finally, we present the scalability of swCaffe for training of ResNet-50 and AlexNet on the scale of 1024 nodes. Liandeng Li, Jiarui Fang, Haohuan Fu, Jinlei Jiang, Wenlai Zhao, Conghui He, Xin You 0001, Guangwen Yang 0002 |
CLUSTER | 2 |
| 2018 | Optimizing Convolutional Neural Networks on the Sunway TaihuLight SupercomputerabstractThe Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively. Wenlai Zhao, Haohuan Fu, Jiarui Fang, Weijie Zheng 0001, Lin Gan 0001, Guangwen Yang 0002 |
ACM Trans. Archit. Code Optim. | 3 |
| 2017 | swDNN: A Library for Accelerating Deep Learning Applications on Sunway TaihuLightabstractTo explore the potential of training complex deep neural networks (DNNs) on other commercial chips rather than GPUs, we report our work on swDNN, which is a highly-efficient library for accelerating deep learning applications on the newly announced world-leading supercomputer, Sunway TaihuLight. Targeting SW26010 processor, we derive a performance model that guides us in the process of identifying the most suitable approach for mapping the convolutional neural networks (CNNs) onto the 260 cores within the chip. By performing a systematic optimization that explores major factors, such as organization of convolution loops, blocking techniques, register data communication schemes, as well as reordering strategies for the two pipelines of instructions, we manage to achieve a double-precision performance over 1.6 Tflops for the convolution kernel, achieving 54% of the theoretical peak. Compared with Tesla K40m with cuDNNv5, swDNN results in 1.91-9.75x performance speedup in an evaluation with over 100 parameter configurations. Jiarui Fang, Haohuan Fu, Wenlai Zhao, Bingwei Chen, Weijie Zheng 0001, Guangwen Yang 0002 |
IPDPS | 1 |
| 2016 | Cache-Friendly Design for Complex Spatially-Variable Coefficient Stencils on Many-Core ArchitecturesabstractMany-core architectures, such as the NVIDIA graphics processing unit and Intel Xeon Phi, which are characterized by high computation resources but limited on-chip memory capacity, have been used to significantly accelerate various computationally demanding tasks. Stencil operators are naturally suitable for such architectures because of their parallel calculation patterns. However, only simple stencils with points distributed along the axes and with constant coefficients have been fully investigated. This study first provides insights into optimization strategies for stencils with complex shapes, including off-axial points and spatially variable coefficients. Through our proposed stencil-decomposition schemes, we maintain read-only coefficients in on-chip caches to avoid unvectorized memory access. To alleviate the resulting severe cache-starvation situation, a generalized cache-friendly design for many-core architecture is proposed. It can reduce cache miss times and cache space consumption. The proposed methodology significantly improves the performance of stencil operations in a real seismic imaging application and introduces a new option to write highly efficient memory-bound stencil-like loops. Jiarui Fang, Haohuan Fu, Guangwen Yang 0002 |
HiPC | 1 |
| 2016 | Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputerabstractThis paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day. Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002 |
SC | 13 |
| 2015 | Optimizing Complex Spatially-Variant Coefficient Stencils for Seismic Modeling on GPUabstractThe Explicit Time Evolution (ETE) method is an innovative Finite-Difference (FD) type method to simulate the wave propagation in acoustic media with higher spatial and temporal accuracy. However, different from FD, it is difficult to achieve an efficient GPU design because of the poor memory access patterns caused by the off-axis points and spatially-variant coefficients. In this paper, we present a set of new optimization strategies for ETE stencils according to the memory hierarchy of NVIDIA GPU. To handle the problem caused by the complexity of the stencil shapes, we design a one-to-multi updating scheme for shared memory usage. To alleviate the performance damage resulted from the poor memory access pattern of reading spatially-variant coefficients, we propose a stencil decomposition method to reduce un-coalesced global memory access. Based on the state-of-the-art GPU architecture, combining with existing spatial and temporal stencil blocking schemes, we manage to achieve 9.6x and 9.9x speedups compared with a well-tuned 12-core CPUs version for 37-point and 73-point ETE stencils, respectively. Compared with a well-tuned MIC version, the best speedups for the 2 type stencils are 3.7x and 4.7x. Our designs leads to an ETE method that is 31.2x faster than conventional CPU-FD method and make it a practical seismic imaging technology. Jiarui Fang, Haohuan Fu, Nanxun Dai, Lin Gan 0001, Guangwen Yang 0002 |
ICPADS | 1 |