EDBT 2026 Demo / reviewers in the wild / expert
Jiangsu Du
dblp:250/0536
· DBLP profile ↗
39ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0003-4707-9492ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 7 first-author · 29 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication
Zhenyi Zheng, Jiangsu Du |
Euro-Par (1) | 4 |
| 2026 | Learning-Aided Delay-Doppler Filtering for Integrated Sensing and Communication Receivers
Jiangsu Du, Chao Zhang 0009 |
IWCMC | 1 |
| 2026 | MixCache: Mixture-of-Cache for Video Diffusion Transformer AccelerationabstractEfficient video generation models are increasingly vital for multimedia synthetic content generation. Leveraging the Transformer architecture and the diffusion process, video DiT models have emerged as a dominant approach for high-quality video generation. However, their multi-step iterative denoising process incurs high computational cost and inference latency, which limits their practical deployment in large-scale and interactive multimedia applications. Caching, a widely adopted optimization method in DiT models, leverages the redundancy in the diffusion process to skip computations in different granularities (e.g., step, cfg, block). Nevertheless, existing caching methods are limited to single-granularity strategies, struggling to balance generation quality and inference speed in a flexible manner. In this work, we propose MixCache, a training-free caching-based framework for efficient video DiT inference. MixCache first distinguishes the interference and boundary between different caching strategies, and then introduces a context-aware cache triggering strategy to determine when caching should be enabled, along with an adaptive hybrid cache decision strategy for dynamically selecting the optimal caching granularity. Extensive experiments on diverse models demonstrate that MixCache can significantly accelerate video generation (e.g., 1.94× speedup on Wan 14B, 1.97× speedup on HunyuanVideo) while delivering both superior generation quality and inference efficiency compared to baseline methods. Yuanxin Wei, Lansong Diao, Bujiao Chen, Shenggan Cheng, Zhengping Qian, Wenyuan Yu, Nong Xiao 0001, Wei Lin 0016, Jiangsu Du |
ICMR | 9 |
| 2026 | Efficient KV Cache Spillover Management on Memory-Constrained GPU for LLM InferenceabstractThe rapid growth of model parameters presents a significant challenge when deploying large generative models on GPU. Existing LLM runtime memory management solutions tend to maximize batch size to saturate GPU device utilization. Nevertheless, this practice leads to situations where the KV Cache of certain sequences cannot be accommodated on GPUs with limited memory capacity during the model inference, requiring temporary eviction from GPU memory (referred to as KV Cache spillover). However, without careful consideration of the LLM inference's runtime pattern, current LLM inference memory management solutions face issues like one-size-fits-all spillover handling approach for different platforms, under-utilization of GPU in prefill stage, and suboptimal sequence selection due to direct employment of swap or recomputation. In this paper, we introduce FuseSpill, a holistic KV Cache management solution designed to boost LLM inference on memory-constrained GPU by efficiently handling KV Cache spillover. Specifically, FuseSpill consists of a spillover cost model that analyzes the system cost of spillover handling techniques quantitatively, a KV cache swap orchestrator to further refine the basic swap technique to sophisticated disaggregate KV Cache across heterogeneous devices for decoding iterations, a multi-executor scheduler to effectively coordinate task executors across devices, and a response length predictor to exploit the length-aware sequence selection strategy when KV Cache spillover occurs. The experimental results demonstrate that our implementation outperforms existing solutions, delivering a 20% to 40% increase in throughput while simultaneously reducing the inference latency of the spillover sequences. Jiazhi Jiang, Yao Chen 0008, Zining Zhang 0001, Bingsheng He, Pingyi Luo, Mian Lu, Yuqiang Chen, Hongbin Zhang 0006, Jiangsu Du, Dan Huang 0001, Yutong Lu |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2026 | Resource-Efficient Personal Large Language Models Fine-Tuning With Collaborative Edge ComputingabstractLarge language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. Other studies focus on exploiting the potential of edge devices through resource management optimization, yet are ultimately bottlenecked by the resource wall of individual devices. To tackle these challenges, we proposePAC+, a resource efficient collaborative edge AI framework for in-situ personal LLMs fine-tuning.PAC+breaks the resource wall of personal LLMs fine-tuning with a sophisticated algorithm-system co-design. (1) Algorithmically,PAC+implements a personal LLMs fine-tuning technique that is efficient in terms of parameters, time, and memory. It utilizes Parallel Adapters to circumvent the need for a full backward pass through the LLM backbone. Additionally, an activation cache mechanism further streamlining the process by negating the necessity for repeated forward passes across multiple epochs. (2) Systematically,PAC+leverages edge devices in close proximity, pooling them as a collective resource for in-situ personal LLMs fine-tuning, utilizing a hybrid data and pipeline parallelism to orchestrate distributed training. The use of the activation cache eliminates the need for forward pass through the LLM backbone, enabling exclusive fine-tuning of the Parallel Adapters using data parallelism. Extensive evaluation of the prototype implementation demonstrates thatPAC+significantly outperforms existing collaborative edge training systems, achieving up to a$9.7\times$end-to-end speedup. Furthermore, compared to mainstream LLM fine-tuning algorithms,PAC+reduces memory footprint by up to$88.16\%$. Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Jiangsu Du, Xiaowen Chu 0001, Guoliang Xing, Xu Chen 0004 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Doppeladler: Adaptive Tensor Parallelism for Latency-Critical LLM Deployment on CPU-GPU Integrated End-User DeviceabstractLLM deployment on end-user devices has attracted significant interest from tech giants and research institutions due to privacy benefits and the elimination of network roundtrips. Reducing latency is crucial for improving user experience. Enduser devices, such as desktop and mobile processors, often integrate CPU and GPU on a single die, making tensor parallelism a promising approach to distribute workloads and reduce inference latency. However, the predefined tensor parallelism in traditional practices cannot guarantee optimization due to heterogeneity and resource contention in the CPU-GPU integrated end-user devices. In this paper, we propose Doppeladler, a practical framework designed to facilitate parallel inference of LLM on end-user devices. Doppeladler adaptively optimizes tensor parallelism based on real-time conditions and device status, enhancing resource utilization by refining workload partitioning and scheduling with a heuristic-based approach. Additionally, it lowers the costs of CPU-GPU hybrid execution by managing device resource, trigger workload rebalance during runtime to mitigate the penalty of contention, and minimizing communication overhead through zero-copy techniques and alternate access pattern. Experiments demonstrate that Doppeladler outperforms existing methods by 1.2 to 2.5 times. Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Yutong Lu |
PACT | 3 |
| 2025 | Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep LearningabstractWith the exponential growth of deep learning (DL), there arises an escalating need for scalability. Despite significant advancements in communication hardware capabilities, the time consumed by communication remains a bottleneck during training. The existing various optimizations are coupled within parallel systems to implement specific computation-communication overlap. These approaches pose challenges in terms of performance, programmability, and generality. In this paper, we introduce Concerto, a compiler framework designed to address these challenges by automatically optimizing and scheduling communication. We formulate the scheduling problem as a resource-constrained project scheduling problem and use off-the-shelf solver to get the near-optimal scheduling. And use auto-decomposition to create overlap opportunity for critical (synchronous) communication. Our evaluation shows Concerto can match or outperform state-of-the-art parallel frameworks, including Megatron-LM, JAX/XLA, DeepSpeed, and Alpa, all of which include extensive hand-crafted optimization. Unlike previous works, Concerto decouples the parallel approach and communication optimization, then can generalize to a wide variety of parallelisms without manual optimization. Shenggan Cheng, Shengjie Lin, Lansong Diao, Hao Wu 0077, Siyu Wang 0006, Chang Si, Xuanlei Zhao, Jiangsu Du, Wei Lin 0016, Yang You 0001 |
ASPLOS (1) | 9 |
| 2025 | PASK: Cold Start Mitigation for Inference with Proactive and Selective Kernel Loading on GPUsabstractToday, DNN inference is widely adopted, with numerous inference services being spawned from scratch across instances in scenarios such as spot serving, serverless scaling and edge computing, where frequent start-stops are required. In this work, we first delve into the inference workflow and uncover the origins of cold start when invoking a DNN model. Specifically, DNN execution is blocked by the kernel loading process to prepare the code object executing on GPU at the DL primitive library (e.g., cuDNN and MIOpen). To tackle this, we propose PASK, a kernel loading and reusing middleware to mitigate the widespread cold start issue. Unlike the reactive kernel scheduling policy used by existing frameworks, PASK adopts a proactive strategy to interleave code loading, kernel issuing and GPU computation to achieve higher hardware utilization. To further reduce the loading overhead, PASK recycles existing loaded kernels to accomplish the DNN operator, rather than introducing new kernels for every layer. Meanwhile, PASK categorically organizes the cached kernels to efficiently find the applicable kernel for reuse and thus minimize incurred runtime overhead. We implement and evaluate PASK atop of open source DNN inference engine and primitive library on off-the-shelf GPUs. Experiments demonstrate PASK is capable of alleviating the cold start overhead of popular DNN models with $5.62 \times$ speedup on average. Xuanteng Huang, Jiangsu Du, Nong Xiao 0001, Xianwei Zhang 0001 |
DAC | 2 |
| 2025 | AuLoRA: Fine-Grained Loading and Computation Orchestration for Efficient LoRA LLM ServingabstractLoRA is a widely used Parameter-Efficient FineTuning (PEFT) technique for customizing pre-trained backbone models to specific tasks. Serving a backbone model with numerous LoRA adapters, known as multi-tenant LoRA serving, is a common scenario where different users utilize distinct LoRA adapters while sharing the same backbone model. To support more LoRA adapters simultaneously and improve efficiency, existing solutions dynamically load adapters from host memory and separate workloads into batched backbone model computation and adapter computation. However, they introduce complex data dependencies and necessitate careful coordination of loading and computation to enhance efficiency. We introduce AuLoRA, a multi-tenant LoRA serving system that achieves fine-grained orchestration of adapter loading and computation alongside backbone model execution. It optimizes both Time-to-First-Token (TTFT) and throughput by: 1) layer-wise-priority LoRA adapter loading, which reorganizes adapter loading by layer, to perform inference before adapters are fully loaded and overlaps adapter loading with backbone computation. 2) Intra-layer pipelined LoRA adapter execution, which loads LoRA adapters and performs computation in a pipelined manner, further hiding the adapter loading overhead. 3) Dynamic LoRA adapter batching, which explores the optimal LoRA adapter batching plan by comprehensively considering kernel launch overhead and modern hardware parallelism, improving computational efficiency. We compare AuLoRA with S-LoRA, a state-of-the-art multi-tenant LoRA serving system, and the results show that AuLoRA can achieve up to$3.03 \times$TTFT reduction and$1.27 \times$throughput improvement. Jiangsu Du, Zhiguang Chen 0001, Yutong Lu |
ICCD | 2 |
| 2025 | Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core ParallelismabstractIn-situ LLM inference on end-user devices has gained significant interest due to its privacy benefits and reduced dependency on external infrastructure. However, as the decoding process is memory-bandwidth-bound, the diverse processing units in modern end-user devices cannot be fully exploited, resulting in slow LLM inference. This paper presents Ghidorah, an LLM inference system for end-user devices with the unified memory architecture. The key idea of Ghidorah can be summarized in two steps: 1) leveraging speculative decoding approaches to enhance parallelism, and 2) ingeniously distributing workloads across multiple heterogeneous processing units to maximize computing power utilization. Ghidorah includes the hetero-core model parallelism (HCMP) architecture and the architecture-aware profiling (ARCA) approach. The HCMP architecture guides partitioning by leveraging the unified memory design of end-user devices and adapting to the hybrid computational demands of speculative decoding. The ARCA approach is used to determine the optimal speculative strategy and partitioning strategy, balancing acceptance rate with parallel capability to maximize the speedup. Additionally, we optimize sparse computation on ARM CPUs. Experimental results show that Ghidorah can achieve up to$7.6 \times$speedup in the dominant LLM decoding phase compared to the sequential decoding approach on NVIDIA Jetson NX. Jinhui Wei, Yuhui Zhou, Jiazhi Jiang, Jiangsu Du |
ICCD | 5 |
| 2025 | TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM InferenceabstractAs the model size continuously increases, pipeline parallelism shows great promise in throughput-oriented LLM inference due to its low demand on communications. However, imbalanced pipeline workloads and complex data dependencies in the prefill and decode phases result in massive pipeline bubbles and further severe performance reduction. Hongbin Zhang 0006, Taosheng Wei, Zhenyi Zheng, Jiangsu Du, Zhiguang Chen 0001, Yutong Lu |
ICPP | 4 |
| 2025 | gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token ThrottlingabstractPipeline parallelism has emerged as a predominant approach for deploying large language models (LLMs) across distributed nodes, owing to its lower communication overhead compared to tensor parallelism. While demonstrating high throughput in request serving, pipeline parallelism often faces performance limitations caused by pipeline bubbles, which are primarily resulted from imbalanced computation delays across batches. Existing methods like Sarathi-Serve attempt to address this through hybrid scheduling of chunked prefill and decode tokens with a fixed token budget. However, such methods may still experience significant fluctuations, arising either from insufficient prefill tokens or uneven distribution of decode tokens, ultimately leading to computational imbalance. Tianyu Guo 0009, Xianwei Zhang 0001, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001, Yutong Lu |
SC | 3 |
| 2025 | coMtainer: Compilation-assisted HPC Container Images with Enhanced AdaptabilityabstractThe increasing interconnectivity of HPC systems has highlighted the need for efficient application migration across different environments. Containers, widely adopted for this purpose, simplify deployment but often fail to deliver optimal performance due to the separated build and execution container workflow. This leads to generic container images that miss out on system-specific software stack advantages, a challenge we define as the adaptability issue. Yuhao Gu, Haoquan Chen, Xianjie Chen, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001, Xianwei Zhang 0001, Yutong Lu |
SC | 4 |
| 2025 | ORFA: Exploring WebAssembly as a Turing Complete Query Language for Web APIsabstractWeb APIs are the primary communication form for Web services, with RESTful design being the predominant paradigm. However, RESTful APIs are typically fixed once defined, causing data under- or over-fetching as they can't meet clients' varying Web service needs. While semantic enriched API query languages like GraphQL mitigates this problem, they still face expressiveness limitations for logical operations such as indirect queries and loop traversals. To address this, we propose ORFA (One Request For All), the first in literature that employs WebAssembly (Wasm) as a Web API query language to achieve complete expressiveness of client requests. ORFA's key advantage lies in its use of Wasm's Turing completeness to allow clients to compose arbitrary operations within a single request, thus significantly eliminating redundant data transmission and boosting communication efficiency. Technically, ORFA provides a runtime for executing Wasm query programs and incorporates new module splitting strategies and a caching mechanism customized for integrating Wasm into Web API services, which can enable lightweight code transfer and fast request responses. Experimental results on a realistic testbed and popular Web applications show that ORFA effectively reduces latency by 18.4% and network traffic by 24.5% on average, compared to the state-of-the-art GraphQL. Yuhao Gu, Jiangsu Du, Xianwei Zhang 0001 |
WWW | 3 |
| 2025 | Resource-Efficient Collaborative Edge Transformer Inference With Hybrid Model ParallelismabstractTransformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users' privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and proposeGalaxy+, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration.Galaxy+introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity and memory-aware parallelism planning for fully exploiting the resource potential. To mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments,Galaxy+devises a tile-based fine-grained overlapping of communication and computation. Furthermore, a fault-tolerant re-scheduling mechanism is developed to address device-level resource dynamics, ensuring stable and low-latency inference. Extensive evaluation based on prototype implementation demonstrates thatGalaxy+remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving a$1.2\times$to$4.24\times$end-to-end latency reduction. Besides,Galaxy+can adapt to device-level resource dynamics, swiftly rescheduling and restoring inference in the presence of unexpected straggler devices. Shengyuan Ye, Bei Ouyang, Jiangsu Du, Liekang Zeng, Tianyi Qian, Wenzhong Ou, Xiaowen Chu 0001, Deke Guo, Yutong Lu, Xu Chen 0004 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Co-Designing Transformer Architectures for Distributed Inference With Low CommunicationabstractTransformer models have shown significant success in a wide range of tasks. However, the massive resources required for its inference prevent deployment on a single device with relatively constrainted resources, thus leaving a high threshold of integrating their advancements. Observing scenarios such as smart home applications on edge devices and cloud deployment on commodity hardware, it is promising to distribute Transformer inference across multiple devices. Unfortunately, due to the tightly-coupled feature of Transformer model, existing model parallelism approaches necessitate frequent communication to resolve data dependencies, making them unacceptable for distributed inference, especially under relatively weak interconnection. In this paper, we propose DeTransformer, a communication-efficient distributed Transformer inference system. The key idea of DeTransformer involves the co-design of Transformer architecture to reduce the communication during distributed inference. In detail, DeTransformer is based on a novel block parallelism approach, which restructures the original Transformer layer with a single block to the decoupled layer with multiple sub-blocks. Thus, it can exploit model parallelism between sub-blocks. Next, DeTransformer contains an adaptive execution approach that strikes a trade-off among communication capability, computing power and memory budget over multiple devices. It incorporates a two-phase planning for execution, namely static planning and runtime planning. The static planning runs offline, containing a profiling procedure and a weight placement strategy before execution. The runtime planning dynamically determines the optimal parallel computing strategy from an expertly crafted search space based on real-time requests. Notably, this execution approach can adapt to heterogeneous devices by distributing workload based on devices’ computing capabilities. We conduct experiments for both auto-regressive and auto-encoder tasks of Transformer models. Experimental results show that DeTransformer can reduce distributed inference latency by up to 2.81× compared to the SOTA approach on 4 devices, while effectively maintaining task accuracy and a consistent model size. Jiangsu Du, Yuanxin Wei, Shengyuan Ye, Jiazhi Jiang, Xu Chen 0004, Dan Huang 0001, Yutong Lu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Communication-Efficient Model Parallelism for Distributed In-Situ Transformer InferenceabstractTransformer models have shown significant success in a wide range of tasks. Meanwhile, massive resources required by its inference prevent scenarios with resource-constrained devices from in-situ deployment, leaving a high threshold of integrating its advances. Observing that these scenarios, e.g. smart home of edge computing, are usually comprise a rich set of trusted devices with untapped resources, it is promising to distribute Transformer inference onto multiple devices. However, due to the tightly-coupled feature of Transformer model, existing model parallelism approaches necessitate frequent communication to resolve data dependencies, making them unacceptable for distributed inference, especially under weak interconnect of edge scenarios. In this paper, we propose DeTransformer, a communication-efficient distributed in-situ Transformer inference system for edge scenarios. DeTransformer is based on a novel block parallelism approach, with the key idea of restructuring the original Trans-former layer with a single block to the decoupled layer with multi-ple sub-blocks and exploit model parallelism between sub-blocks. Next, DeTransformer contains an adaptive placement approach to automatically select the optimal placement strategy by striking a trade-off among communication capability, computing power and memory budget. Experimental results show that DeTransformer can reduce distributed inference latency by up to 2.81 x compared to the SOTA approach on 4 devices, while effectively maintaining task accuracy and a consistent model size. Yuanxin Wei, Shengyuan Ye, Jiazhi Jiang, Xu Chen 0004, Dan Huang 0001, Jiangsu Du, Yutong Lu |
DATE | 6 |
| 2024 | Efficient Coupling Streaming AI and Ensemble Simulations on HPC Clusters
Jiazhi Jiang, Hongbin Zhang 0006, Deyin Liu, Jiangsu Du, Xiaojiao Yao, Jinhui Wei, Pin Chen, Dan Huang 0001, Yutong Lu |
Euro-Par (1) | 4 |
| 2024 | Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceabstractTransformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users’ privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and propose Galaxy, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration. Galaxy introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity-aware parallelism planning for fully exploiting the resource potential. Furthermore, Galaxy devises a tile-based fine-grained overlapping of communication and computation to mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments. Extensive evaluation based on prototype implementation demonstrates that Galaxy remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 2.5× end-to-end latency reduction. Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu 0001, Yutong Lu, Xu Chen 0004 |
INFOCOM | 2 |
| 2024 | Understanding the Inference Performance of Spatial Temporal Diffusion Transformer
Yuanxin Wei, Jiangsu Du, Dan Huang 0001, Nong Xiao 0001 |
NPC (1) | 3 |
| 2024 | Liger: Interleaving Intra- and Inter-Operator Parallelism for Distributed Large Model InferenceabstractDistributed large model inference is still in a dilemma where balancing cost and effect. The online scenarios demand intraoperator parallelism to achieve low latency and intensive communications makes it costly. Conversely, the inter-operator parallelism can achieve high throughput with much fewer communications, but it fails to enhance the effectiveness. Jiangsu Du, Jinhui Wei, Jiazhi Jiang, Shenggan Cheng, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
PPoPP | 1 |
| 2024 | APTMoE: Affinity-Aware Pipeline Tuning for MoE Models on Bandwidth-Constrained GPU NodesabstractRecently, the sparsely-gated Mixture-Of-Experts (MoE) architecture has garnered significant attention. To benefit a wider audience, fine-tuning MoE models on more affordable clusters, which are typically a limited number of bandwidthconstrained GPU nodes, holds promise. However, it is non-trivial to apply existing cost-effective fine-tuning approaches to MoE models, due to the increased ratio of data to computation. In this paper, we introduce APTMoE, which employs affinityaware pipeline parallelism for fine-tuning MoE models on bandwidth-constrained GPU nodes. We propose an affinity-aware offloading technique that enhances pipeline parallelism for both computational efficiency and model size, and it benefits from a hierarchical loading strategy and a demand-priority scheduling strategy. To improve the computation efficiency and reduce the data movement volume, the hierarchical loading strategy designs three loading phases and efficiently allocates computation across GPUs and CPUs during these phases, leveraging different levels of expert popularity and computation affinity. With the aim of alleviating the mutual interference among the three loading phases and maximizing the bandwidth utilization, the demand-priority scheduling strategy proactively and dynamically coordinates the loading execution order. Experiments demonstrate that APTMoE outperforms existing methods in most cases. Particularly, APTMoE successfully fine-tunes a 61.2B MoE model on 4 Nvidia A800 GPUs(40GB) and achieves up to $33 \%$ throughput improvement compared to the SOTA method. Yuanxin Wei, Jiangsu Du, Jiazhi Jiang, Xianwei Zhang 0001, Dan Huang 0001, Nong Xiao 0001, Yutong Lu |
SC | 2 |
| 2024 | SAIH: A Scalable Evaluation Methodology for Understanding AI Performance Trend on HPC Systems
Jiangsu Du, Yingpeng Wen, Jiazhi Jiang, Dan Huang 0001, Xiangke Liao, Yutong Lu |
J. Comput. Sci. Technol. | 1 |
| 2024 | IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model TrainingabstractTraining large models for modern recommendation systems requires a substantial number of computational devices and extended periods. Since it is essential to store model checkpoints throughout the training progress for accuracy debugging or mitigating potential failures, checkpointing systems are widely used. However, given that recommendation models can scale to hundreds of gigabytes or more, existing solutions often introduce significant overhead in terms of both storage and I/O. In this paper, we present IncrCP, a checkpointing system specifically designed for recommendation models. Given that only a small fraction of model parameters are modified in each iteration, IncrCP creatively leverages the incremental checkpointing strategy and overcomes the inherent slow recovery problem. To support recovering all states throughout the training process, while also ensuring efficient storage utilization and rapid recovery, IncrCP proposes the 2-D chunk approach. It proactively records changed parameters in the training process as well as their indexes, extracts parameters according to duplicated indexes as independent chunk files, and then orchestrates these chunks in the 2-dimensional linked list. In this way, IncrCP achieves fast recovery by loading less unnecessary parameters and performing less deduplication during recovery. Furthermore, IncrCP includes a selective extraction approach to reduce I/O by avoiding worthless extractions and a concatenate approach to reduce random disk access when recovery. Evaluations show that IncrCP achieves up to 6.6× recovery speedup compared to the naive incremental strategy and saves storage space by 60.4% with slight overhead compared to another recovery-friendly strategy. Qingyin Lin, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001 |
Proc. VLDB Endow. | 2 |
| 2024 | Sophisticated Orchestrating Concurrent DLRM Training on CPU/GPU PlatformabstractRecommendation systems are essential to the operation of the majority of internet services, with Deep Learning Recommendation Models (DLRMs) serving as a crucial component. However, due to distinct computation, data access, and memory usage characteristics of recommendation models, the trainning of DLRMs may suffer from low resource utilization on prevalent heterogeneous CPU-GPU hardware platforms. Furthermore, as the majority of high-performance computing systems presently depend on multi-GPU computing nodes, the challenge of addressing low resource utilization becomes even more pronounced. Existing concurrent training solutions cannot be straightforwardly applied to DLRM due to various factors, such as insufficient fine-grained memory management and the lack of collaborative CPU-GPU scheduling. In this paper, we introduce RMixer, a scheduling framework that addresses these challenges by providing an efficient job management and scheduling mechanism for DLRM training jobs on heterogeneous CPU-GPU platforms. To facilitate training co-location, we first estimate the peak memory consumption of each job. Additionally, we track and collect resource utilization for DLRM training jobs. Based on the information of computational patterns, a batched job dispatcher with dynamic resource-complementary scheduling policy is proposed to co-locate DLRM training jobs on CPU-GPU platform. Scheduling strategies for both intra-GPU and inter-GPU scenarios were meticulously devised, with a focus on thoroughly examining individual GPU resource utilization and achieving a balanced state across multiple GPUs. Experimental results demonstrate that our implementation achieved up to 5.3× and 7.5× higher throughput on single GPU and 4 GPU respectively for training jobs involving various recommendation models. Rui Tian 0001, Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Yutong Lu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Enhancing Multi-physics Coupling on ARM Many-Core Cluster
Wencheng Shi, Jiangsu Du, Dan Huang 0001, Yutong Lu |
APPT | 3 |
| 2023 | Accelerating Inference of 3D-CNN on ARMMany-core CPU via Hierarchical Model PartitionabstractMany applications such as biomedical analysis and scientific data analysis involve analyzing volumetric data. This spawns huge demand for 3D CNN. Although accelerators such as GPU may provide higher throughput on deep learning applications, they may not be available in all scenarios. CPU, especially many-core CPU, remains an attractive choice for deep learning in many scenarios. In this paper, we propose a inference solution that targets on the emerging ARM many-core CPU platform. A hierarchical partition approach is claimed to accelerate 3D-CNN inference by exploiting characteristics of memory and cache on ARM many-core CPU. Jiazhi Jiang, Zijiang Huang, Dan Huang 0001, Jiangsu Du, Yutong Lu |
DATE | 4 |
| 2023 | MixRec: Orchestrating Concurrent Recommendation Model Training on CPU-GPU platformabstractThe development of deep learning recommendation models (DLRM) and recommendation systems has significantly improved the precision of information matching. Due to distinct computation, data access, and memory usage characteristics of recommendation models, they may suffer from low resource utilization on prevalent heterogeneous CPU-GPU hardware platforms. Existing concurrent training solutions cannot be directly applied to DLRM due to various factors, such as insufficient fine-grained memory management and the lack of collaborative CPU-GPU scheduling. In this paper, we introduce MixRec, a scheduling framework that addresses these challenges by pro-viding an efficient job management and scheduling mechanism for DLRM training jobs on heterogeneous CPU-GPU platforms. To facilitate training co-location, we first estimate the peak memory consumption of each job. Additionally, we track and collect resource utilization for DLRM training jobs. Based on the information of resource usage, a batched job dispatcher with dynamic resource-complementary scheduling policy is proposed to co-locate DLRM training jobs on CPU-GPU platform. Experimental results demonstrate that our implementation achieved up to 4.42× higher throughput and 3.97× higher resource utilization for training jobs involving various recommendation models. Jiazhi Jiang, Rui Tian 0001, Jiangsu Du, Dan Huang 0001, Yutong Lu |
ICCD | 3 |
| 2023 | Optimizing massively parallel sparse matrix computing on ARM many-core processor
Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Yutong Lu |
Parallel Comput. | 3 |
| 2023 | Improving Computation and Memory Efficiency for Real-world Transformer Inference on GPUsabstractTransformer models have emerged as a leading approach in the field of natural language processing (NLP) and are increasingly being deployed in production environments. Graphic processing units (GPUs) have become a popular choice for the transformer deployment and often rely on the batch processing technique to ensure high hardware performance. Nonetheless, the current practice for transformer inference encounters computational and memory redundancy due to the heavy-tailed distribution of sequence lengths in NLP scenarios, resulting in low practical performance. In this article, we propose a unified solution for improving both computation and memory efficiency of the real-world transformer inference on GPUs. The solution eliminates the redundant computation and memory footprint across a transformer model. At first, a GPU-oriented computation approach is proposed to process the self-attention module in a fine-grained manner, eliminating its redundant computation. Next, the multi-layer perceptron module continues to use the word-accumulation approach to eliminate its redundant computation. Then, to better unify the fine-grained approach and the word-accumulation approach, it organizes the data layout of the self-attention module in block granularity. Since aforementioned approaches make the required memory size largely reduce and constantly fluctuate, we propose the chunk-based approach to enable a better balance between memory footprint and allocation/free efficiency. Our experimental results show that our unified solution achieves a decrease of average latency by 28% on the entire transformer model, 63.8% on the self-attention module, and reduces memory footprint of intermediate results by 7.8×, compared with prevailing frameworks. Jiangsu Du, Jiazhi Jiang, Hongbin Zhang 0006, Dan Huang 0001, Yutong Lu |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | Hierarchical Model Parallelism for Optimizing Inference on Many-core Processor via Decoupled 3D-CNN StructureabstractThe tremendous success of convolutional neural network (CNN) has made it ubiquitous in many fields of human endeavor. Many applications such as biomedical analysis and scientific data analysis involve analyzing volumetric data. This spawns huge demand for 3D-CNN. Although accelerators such as GPU may provide higher throughput on deep learning applications, they may not be available in all scenarios. CPU, especially many-core CPU with non-uniform memory access (NUMA) architecture, remains an attractive choice for deep learning inference in many scenarios. In this article, we propose a distributed inference solution for 3D-CNN that targets on the emerging ARM many-core CPU platform. A hierarchical partition approach is claimed to accelerate 3D-CNN inference by exploiting characteristics of memory and cache on ARM many-core CPU. Based on the hierarchical model partition approach, other optimization techniques such as NUMA-aware thread scheduling and optimization of 3D-img2row convolution are designed to exploit the potential of ARM many-core CPU for 3D-CNN. We evaluate our proposed inference solution with several classic 3D-CNNs: C3D, 3D-resnet34, 3D-resnet50, 3D-vgg11, and P3D. Our experimental results show that our solution can boost the performance of the 3D-CNN inference, and achieve much better scalability, with a negligible fluctuation in accuracy. When employing our 3D-CNN inference solution on ACL libraries, it can outperform naive ACL implementations by 11× to 50× on ARM many-core processor. When employing our 3D-CNN inference solution on NCNN libraries, it can outperform the naive NCNN implementations by 5.2× to 14.2× on ARM many-core processor. Jiazhi Jiang, Zijiang Huang, Dan Huang 0001, Jiangsu Du, Lin Chen 0002, Ziguang Chen, Yutong Lu |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | Full-Stack Optimizing Transformer Inference on ARM Many-Core CPUabstractThe past several years have witnessed tremendous success of transformer models in natural language processing (NLP), and their current landscape is increasingly diverse. Although GPU gradually becomes the dominating workhorse and de facto standard for deep learning, there are still many scenarios where using CPU remains a prevalent choice.Recently, ARM many-core processor starts emigrating to cloud computing and high-performance computing, which is promising to deploy transformer inference. In this paper, we identify several performance bottlenecks of existing inference runtime on many-core CPU including low-core usage, isolated thread configuration, inappropriate implementation of general matrix multiply (GEMM), and redundant computations for variable-length inputs. To tackle these problems, full-stack optimizations are conducted for these challenges from service level to operator level. We explore multi-instance parallelization at the service level to improve CPU core usage. To improve parallel efficiency of the inference runtime, we design NUMA-aware thread scheduling and a look-up table for optimal parallel configurations. The GEMM implementation is tailored for some critical modules to exploit the characteristics of transformer workload. To eliminate redundant computations, a novel storage format is designed and implemented to pack sparse data and a load balancing strategy is proposed for tasks with different sparsity. Experiments show that our implementation can outperform existing solutions by 1.1x to 6x with fixed-length inputs. For variable-length inputs, it achieves 1.9x to 8x speedups on different ARM many-core processors. Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Characterizing and Optimizing Transformer Inference on ARM Many-core ProcessorabstractTransformer has experienced tremendous success and revolutionized the field of natural language processing (NLP). While GPU has become the de facto standard for deep learning computation in many cases, there are still many scenarios where using CPU for deep learning remains a prevalent choice. In particular, ARM many-core processor is emerging as a competitive candidate for HPC systems, which is promising to deploy Transformer inference. Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Yutong Lu |
ICPP | 2 |
| 2022 | Handling heavy-tailed input of transformer inference on GPUsabstractTransformer-based models achieve superior accuracy in the field of natural language processing (NLP) and start to be widely deployed in production. As a popular deployment device, graphic processing units (GPUs) basically adopt the batch processing technique for inferring transformer-based models and achieving high hardware performance. However, as the input sequence lengths of NLP tasks are generally variable and in a heavy-tailed distribution, the batch processing will bring large amounts of redundant computation and hurt the practical efficiency. Jiangsu Du, Jiazhi Jiang, Yang You 0001, Dan Huang 0001, Yutong Lu |
ICS | 1 |
| 2022 | Enhancing Distributed In-Situ CNN Inference in the Internet of ThingsabstractConvolutional neural networks (CNNS) enable machines to view the world as humans and become increasing prevalent for Internet of Things (IoT) applications. Instead of streaming the raw data to the cloud and executing CNN inference remotely, it would be very attractive to use local IoT devices to process as it enables IoT applications with independent decision-making ability. Since a single IoT device can hardly match the requirements of the CNN inference, especially for time-sensitive and high-accuracy tasks, the distributedin-situCNN inference becomes a potential solution. However, because of the inherently tightly coupled structure of existing CNN models, it is difficult to distribute the inference efficiently. In this article, we enhance the distributedin-situCNN inference in the IoT. We fundamentally reduce the communication overhead of distributed CNN inference by designing new loosely coupled structure (LCS). Experimental results demonstrate that LCS achieves the leading performance compared with other popular structures. Next, based on the LCS, we customize the partitioning method to reduce the synchronization points and design the decentralized asynchronous method to optimize communication in each synchronization point. To evaluate the effectiveness, we build a prototype system. When the number of IoT devices increases from 1 to 4, our system accelerates by up to$3.85\times $and reduces the memory footprint in each device by 70% with achieving a competitive accuracy and significantly outperforming other approaches. Jiangsu Du, Yunfei Du 0001, Dan Huang 0001, Yutong Lu, Xiangke Liao |
IEEE Internet Things J. | 1 |
| 2022 | Optimizing small channel 3D convolution on GPU with tensor core
Jiazhi Jiang, Dan Huang 0001, Jiangsu Du, Yutong Lu, Xiangke Liao |
Parallel Comput. | 3 |
| 2021 | Model Parallelism Optimization for Distributed Inference Via Decoupled CNN StructureabstractIt is promising to deploy CNN inference on local end-user devices for high-accuracy and time-sensitive applications. Model parallelism has the potential to provide high throughput and low latency in distributed CNN inference. However, it is non-trivial to use model parallelism as the original CNN model is inherently tightly-coupled structure. In this article, we propose DeCNN, a more effective inference approach that uses decoupled CNN structure to optimize model parallelism for distributed inference on end-user devices. DeCNN is novel consisting of three schemes. Scheme-1 is structure-level optimization. It exploits group convolution and channel shuffle to decouple the original CNN structure for model parallelism. Scheme-2 is partition-level optimization. It is based on channel group to partition the convolutional layers, and then leverages input-based method to partition the fully connected layers, further exposing high degree of parallelism. Scheme-3 is communication-level optimization. It uses inter-sample parallelism to hide communications for better performance and robustness, especially in the weak network connections. We use ImageNet classification task to evaluate the effectiveness of DeCNN on a distributed multi-ARM platform. Notably, when using the number of devices from 1 to 4, DeCNN can accelerate the inference of large-scale ResNet-50 by 3.21×, and reduce 65.3 percent memory footprint, with 1.29 percent accuracy improvement. Jiangsu Du, Xin Zhu 0003, Minghua Shen, Yunfei Du 0001, Yutong Lu, Nong Xiao 0001, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | A Distributed In-Situ CNN Inference System for IoT ApplicationsabstractCNN is a popular deep learning structure able to provide intelligent processing in IoT applications. Instead of deploying the resource-hungry CNN inference workloads on the cloud, it would be promising to utilize local IoT devices for the in-situ processing. Since a single IoT device has only limited resources available, distributing over multiple local devices becomes a potential solution, especially for high-accuracy and time-sensitive tasks. However, it is non-trivial to distribute the inference of existing CNN models efficiently as they are inherently tightly-coupled structure. In this paper, we propose a distributed in-situ CNN inference system with the loosely-coupled CNN structure (LCS), the synchronization-oriented partitioning (SOP), and the decentralized asynchronous communication (DAC) for IoT applications. LCS is based on two novel design ideas, the homogeneous group and the intermittent shuffle. Experiments on ImageNet classification illustrate that LCS has the leading accuracy compared with other structures, under a given computation budget. SOP and DAC target on converting the loosely-coupled feature of LCS into practical performance improvement. SOP tries to partition LCS with fewer synchronization points and DAC reduces the communication overhead by overlapping communications. When the number of IoT devices increases from 1 to 4, our system accelerates by up to 3.85 ×, and reduces the memory footprint in each device by 70%, outperforming other approaches. Jiangsu Du, Minghua Shen, Yunfei Du 0001 |
ICCD | 1 |
| 2019 | Understanding the Resource Demand Differences of Deep Neural Network Training
Jiangsu Du, Xin Zhu 0003, Yunfei Du 0001 |
ICA3PP (2) | 1 |