VLDB 2026 Research / reviewers in the wild / expert
Zhen Zheng
dblp:115/5528
· DBLP profile ↗
30ranked-venue papers
5as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoBSA: RoPE-based Blockwise Sparse Multi-head Latent AttentionabstractLarge Language Models (LLMs) have rapidly advanced in recent years, scaling up in both parameter count and context length.However, as context windows extend from thousands to hundreds of thousands of tokens, attention computation becomes the dominant source of memory usage and runtime in decoding stages, severely limiting the efficiency and scalability of longcontext LLMs.Sparse attention has emerged as a promising solution, reducing complexity by computing attention over only a subset of context tokens.However, the sparse attention for Multi-head Latent Attention (MLA) which is a variant of standard MHA is rarely studied.In this paper, we introduce RoPE-based Blockwise Sparse Attention (RoBSA), a method designed specifically for MLA during the decoding stage of model inference.RoBSA leverages the decoupled nature of RoPE within MLA to implement token selection in a blockwise manner.RoBSA is a lightweight, training-free, and layer-aware algorithm that can be integrated in a plug-and-play fashion.Our method significantly reduces end-to-end inference latency in the decoding stage by up to 2.55× with minimal accuracy loss compared to full attention in long-context scenarios for very large models. Kairong Luo, Zhen Zheng |
ACL (1) | 3 |
| 2026 | A comprehensive taxonomy of prompt engineering techniques for large language modelsabstractAbstract Large Language Models (LLMs) have demonstrated remarkable performance across various downstream tasks, as evidenced by numerous studies. Since 2022, generative AI has shown significant potential in diverse application domains, including gaming, film and television, media, and finance. By 2023, the global AI-generated content (AIGC) industry had attracted over $26 billion in investment. As LLMs become increasingly prevalent, prompt engineering has emerged as a key research area to enhance user-AI interactions and improve LLM performance. The prompt, which serves as the input instruction for the LLM, is closely linked to the model’s responses. Prompt engineering refines the content and structure of prompts, thereby enhancing the performance of LLMs without changing the underlying model parameters. Despite significant advancements in prompt engineering, a comprehensive and systematic summary of existing techniques and their practical applications remains absent. To fill this gap, we investigate existing techniques and applications of prompt engineering. We conduct a thorough review and propose a novel taxonomy that provides a foundational framework for prompt construction. This taxonomy categorizes prompt engineering into four distinct aspects: profile and instruction, knowledge, reasoning and planning, and reliability. By providing a structured framework for understanding its various dimensions, we aim to facilitate the systematic design of prompts. Furthermore, we summarize existing prompt engineering techniques and explore the applications of LLMs across various domains, highlighting their interrelation with prompt engineering strategies. This survey underscores the progress of prompt engineering and its critical role in advancing AI applications, ultimately aiming to provide a systematic reference for future research and applications. Yao-Yang Liu, Zhen Zheng, Feng Zhang 0007, Jin-Cheng Feng, Yi-Yang Fu, Jidong Zhai, Bingsheng He, Xiao Zhang 0001, Xiaoyong Du 0001 |
Frontiers Comput. Sci. | 2 |
| 2026 | Semantic similarity guided contrastive hashing for unsupervised cross-modal retrieval
Limeng Gao, Xinzhong Wang, Zhen Zheng, Mingzhe Yang |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | OVDS-YOLO: An Efficient Detector for Oriented Vehicle Detection in SAR ImagesabstractSynthetic aperture radar (SAR) vehicle detection holds significant value in military reconnaissance and traffic management. Compared to optical images, SAR images exhibit distinctive characteristics. Owing to huge disturbance caused by irrelevant objects and limited expressivity of vehicle features, most mainstream detectors based on deep learning face enormous challenges in SAR vehicle detection. Additionally, the horizontal bounding boxes (HBBs) are replaced by the oriented bounding boxes (OBBs) to suppress background noise and obtain more accurate boundaries of vehicles. However, angle regression remains a challenging problem. To address these issues, we propose an efficient detector for oriented vehicle detection in SAR images, called OVDS-YOLO, combining three major improvements. Initially, in order to distinguish between target vehicles and clutter based on backgrounds, we introduce a background information aggregation module (BIAM), which effectively integrates global information and multi-scale surrounding environment information of vehicles. Furthermore, to guide the network to learn salient features of vehicles, we devise a feature selection module (FSM), which focuses on key features at the pixel and channel levels. Finally, to improve the accuracy of angle prediction, we embed a novel joint optimization paradigm (NJOP) into the detection framework, which performs dual-optimization for angles relying on Gaussian loss and angle loss. The experimental results on the constructed FARAD vehicle dataset confirm that the proposed method achieves excellent performance with low computational overhead. Bingting Zha, Zhen Zheng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2026 | Relative semantic relationship preserving hashing for unsupervised cross-modal retrieval
Limeng Gao, Xinzhong Wang, Zhen Zheng |
Multim. Syst. | 4 |
| 2025 | Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingabstractData centers now allow multiple applications that have lightweight workloads to share a GPU. Existing temporal or spatial sharing systems struggle to provide efficient and accurate quota assignments. We observe that the performance of the multi-user system is often underestimated because of the existence of unused GPU "bubbles" and can be enhanced by squeezing the bubbles. Based on this observation, we design Bless, a bubble-less spatial-temporal sharing GPU system that fine-tunes the GPU resource allocation to improve multi-user performance. Bless leverages precise computing resource management and fine-grained kernel scheduling to ensure stringent quota guarantees and reduce latency fairly for applications with varying GPU quotas. We implement and evaluate Bless with multiple applications and workloads. Our result shows that Bless achieves 21.1% - 37.3% average latency reduction over the state-of-the-art while guaranteeing the promised quota for all applications. Shulai Zhang, Quan Chen 0002, Weihao Cui, Han Zhao 0005, Chunyu Xue, Zhen Zheng, Wei Lin 0016, Minyi Guo |
EuroSys | 6 |
| 2025 | PluS: Highly Efficient and Expandable ML Compiler with Pluggable Graph Schedules
Zhen Zheng, Feng Zhang 0007, Chuanjie Liu, Zaifeng Pan, Jidong Zhai, Xiaoyong Du 0001 |
USENIX ATC | 2 |
| 2024 | WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsabstractGraph Neural Network (GNN) has emerged as an important workload for learning on graphs. With the size of graph data and the complexity of GNN model architectures increasing, developing an efficient GNN system grows more important. As GNN has heavy neural computation workloads on a large graph, it is crucial to partition the entire workload into smaller parts for parallel execution and optimization. However, existing approaches separately partition graph data and GNN operations, resulting in inefficiency and large data movement overhead. Kezhao Huang, Jidong Zhai, Liyan Zheng 0001, Haojie Wang 0004, Yuyang Jin 0001, Qihao Zhang, Runqing Zhang, Zhen Zheng, Youngmin Yi, Xipeng Shen |
EuroSys | 8 |
| 2024 | MonoNN: Enabling a New Monolithic Optimization Space for Neural Network Inference Tasks on Modern GPU-Centric Architectures
Donglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu, Wei Lin 0016, Shuaiwen Song |
OSDI | 2 |
| 2024 | RecFlex: Enabling Feature Heterogeneity-Aware Optimization for Deep Recommendation Models with Flexible SchedulesabstractIndustrial recommendation models typically involve numerous feature fields. The embedding computation workloads are heterogeneous across these fields, thus requiring varied optimal code schedules. While existing solutions apply basic fusion optimization for embedding operations, they inefficiently treat all feature fields with identical schedules, leading to suboptimal performance. In this paper, we introduce RecFlex, which generates fused kernels with distinct schedules for different feature fields. RecFlex employs the interference-aware schedule tuner to tune schedules and the heterogeneous schedule fusion compiler to generate fused kernels, addressing two major challenges. To determine optimal schedules of different feature fields within the fused kernel, RecFlex proposes a two-stage interferencesimulated tuning strategy. To handle dynamic workloads that challenge tuning and fusion, RecFlex combines compile-time schedule tuning with runtime kernel thread mapping. RecFlex surpasses state-of-the-art libraries and compilers, achieving average speedups of $2.64 \times, 20.77 \times$, and $11.31 \times$ over TorchRec, HugeCTR, and RECom, respectively. RecFlex is publicly available at https://github.com/PanZaifeng/RecFlex. Zaifeng Pan, Zhen Zheng, Feng Zhang 0007, Shaden Smith, Chuanjie Liu, Olatunji Ruwase, Xiaoyong Du 0001, Yufei Ding 0001 |
SC | 2 |
| 2024 | Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen 0004, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Song |
USENIX ATC | 2 |
| 2023 | RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsabstractEmbedding columns are important for deep recommendation models to achieve high accuracy, but they can be very time-consuming during inference. Machine learning (ML) compilers are used broadly in real businesses to optimize ML models automatically. Unfortunately, no existing work uses compilers to automatically accelerate the heavy embedding column computations during recommendation model inferences. To fill this gap, we propose RECom, the first ML compiler that aims at optimizing the massive embedding columns in recommendation models on the GPU. RECom addresses three major challenges. First, generating an efficient schedule on the GPU for the massive operators within embedding columns is difficult. Existing solutions usually lead to numerous small kernels and also lack inter-subgraph parallelism. We adopt a novel codegen strategy that fuses massive embedding columns into a single kernel and maps each column into a separate thread block on the GPU. Second, the complex shape computations under dynamic shape scenarios impede further graph optimizations. We develop a symbolic expression-based module to reconstruct all shape computations. Third, ML frameworks inevitably introduce redundant computations due to robustness considerations. We develop a subgraph optimization module that performs graph-level simplifications based on the entire embedding column context. Experiments on both in-house and open-source models show that RECom can achieve 6.61X and 1.91X over state-of-the-art baselines in terms of end-to-end inference latency and throughput, respectively. RECom's source code is publicly available at https://github.com/AlibabaResearch/recom. Zaifeng Pan, Zhen Zheng, Feng Zhang 0007, Hao Liang 0003, Dalin Wang, Xiafei Qiu, Wei Lin 0016, Xiaoyong Du 0001 |
ASPLOS (4) | 2 |
| 2023 | BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachabstractCompiler optimization plays an increasingly important role to boost the performance of machine learning models for data processing and management. With increasingly complex data, the dynamic tensor shape phenomenon emerges for ML models. However, existing ML compilers either can only handle static shape models or expose a series of performance problems for both operator fusion optimization and code generation in dynamic shape scenes. This paper tackles the main challenges of dynamic shape optimization: the fusion optimization without shape value, and code generation supporting arbitrary shapes. To tackle the fundamental challenge of the absence of shape values, it systematically abstracts and excavates the shape information and designs a cross-level symbolic shape representation. With the insight that what fusion optimization relies upon is tensor shape relationships between adjacent operators rather than exact shape values, it proposes the dynamic shape fusion approach based on shape information propagation. To generate code that adapts to arbitrary shapes efficiently, it proposes a compile-time and runtime combined code generation approach. Finally, it presents a complete optimization pipeline for dynamic shape models and implements an industrial-grade ML compiler, named BladeDISC. The extensive evaluation demonstrates that BladeDISC outperforms PyTorch, TorchScript, TVM, ONNX Runtime, XLA, Torch Inductor (dynamic shape), and TensorRT by up to 6.95×, 6.25×, 4.08×, 2.04×, 2.06×, 7.92×, and 4.16× (3.54×, 3.12×, 1.95×, 1.47×, 1.24×, 2.93×, and 1.46× on average) in terms of end-to-end inference speedup on the A10 and T4 GPU, respectively. BladeDISC's source code is publicly available at https://github.com/alibaba/BladeDISC. Zhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu 0004, Wenyi Zhao, Tianyou Guo, Xiafei Qiu, Minmin Sun, Feng Zhang 0007, Xiaoyong Du 0001, Jidong Zhai, Wei Lin 0016 |
Proc. ACM Manag. Data | 1 |
| 2023 | Flash-LLM: Enabling Low-Cost and Highly-Efficient Large Generative Model Inference With Unstructured SparsityabstractWith the fast growth of parameter size, it becomes increasingly challenging to deploy large generative models as they typically require large GPU memory consumption and massive computation. Unstructured model pruning has been a common approach to reduce both GPU memory footprint and the overall computation while retaining good model accuracy. However, the existing solutions do not provide an efficient support for handling unstructured sparsity on modern GPUs, especially on the highly-structured tensor core hardware. Therefore, we propose Flash-LLM for enabling low-cost and highly efficient large generative model inference with the sophisticated support of unstructured sparsity on high-performance but highly restrictive tensor cores. Based on our key observation that the main bottleneck of generative model inference is the several skinny matrix multiplications for which tensor cores would be significantly under-utilized due to low computational intensity, we propose a general Load-as-Sparse and Compute-as-Dense methodology for unstructured sparse matrix multiplication (SpMM). The basic insight is to address the significant memory bandwidth bottleneck while tolerating redundant computations that are not critical for end-to-end performance on tensor cores. Based on this, we design an effective software framework for tensor core based unstructured SpMM, leveraging on-chip resources for efficient sparse data extraction and computation/memory-access overlapping. Extensive evaluations demonstrate that (1) at SpMM kernel level, Flash-LLM significantly outperforms the state-of-the-art library, i.e., Sputnik and SparTA by an average of 2.9X and 1.5X, respectively.(2) At end-to-end framework level on OPT-30B/66B/175B models, for tokens per GPU-second , Flash-LLM achieves up to 3.8X and 3.6X improvement over DeepSpeed and FasterTransformer, respectively, with significantly lower inference cost. Haojun Xia, Zhen Zheng, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li 0045, Wei Lin 0016, Shuaiwen Song |
Proc. VLDB Endow. | 2 |
| 2023 | Expanding the Edge: Enabling Efficient Winograd CNN Inference With Deep Reuse on Edge DeviceabstractDeep learning on edge devices is becoming increasingly important, especially with the explosion of IoT devices. For example, the total number of devices connected to IoT reaches 29 billion in 2022. Convolutional neural networks (CNNs), as common deep learning representatives, are among the most popular neural networks in knowledge and data engineering. However, CNN employs a high degree of computing. In comparison to the training phase, the inference process is more frequently done on low-power computing equipments, such as edge devices. The limited computing resource and high computation pressure limit the effective use of CNN algorithms at the edge. Fortunately, a minimal filtering algorithm called Winograd can reduce convolution calculations by minimizing multiplication operations. We find that Winograd convolution can be accelerated further bydeep reusetechnique, which reuses the similar data and computation processes. In this paper, we propose a new inference method, called DREW, which combines deep reuse with Winograd for further accelerating CNNs. DREW handles three difficulties. First, it can detect the similarities from the complex minimal filtering patterns by clustering. Second, it reduces the online clustering cost in a reasonable range. Third, it provides an adjustable method in clustering granularity balancing the performance and accuracy. We perform evaluation on Raspberry PI and NVIDIA Jetson AGX Xavier edge devices, and experiments show that on five popular networks, 1) DREW further accelerates the Winograd convolution by an average of 8.27× speedup. Even for the highly parallel Winograd implementation, DREW still can provide 2.21× speedup. 2) When DREW is applied to end-to-end Winograd CNN inferences, DREW achieves 5.94× the average performance speedup with no ($< $0.4%) accuracy loss. 3) Energy consumption is an important factor for inference in practice. DREW reduces the number of convolution operations to 10% of the original operations, thus achieving up to 60% energy-efficiency benefits than the original Winograd inference. Feng Zhang 0007, Jiawei Guan, Zhen Zheng, Xiaoguang Guo, Xiao Zhang 0001, Xiaoyong Du 0001, Xipeng Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | AStitch: enabling a new multi-dimensional optimization space for memory-intensive ML training and inference on modern SIMT architecturesabstractThis work reveals that memory-intensive computation is a rising performance-critical factor in recent machine learning models. Due to a unique set of new challenges, existing ML optimizing compilers cannot perform efficient fusion under complex two-level dependencies combined with just-in-time demand. They face the dilemma of either performing costly fusion due to heavy redundant computation, or skipping fusion which results in massive number of kernels. Furthermore, they often suffer from low parallelism due to the lack of support for real-world production workloads with irregular tensor shapes. To address these rising challenges, we propose AStitch, a machine learning optimizing compiler that opens a new multi-dimensional optimization space for memory-intensive ML computations. It systematically abstracts four operator-stitching schemes while considering multi-dimensional optimization objectives, tackles complex computation graph dependencies with novel hierarchical data reuse, and efficiently processes various tensor shapes via adaptive thread mapping. Finally, AStitch provides just-in-time support incorporating our proposed optimizations for both ML training and inference. Although AStitch serves as a stand-alone compiler engine that is portable to any version of TensorFlow, its basic ideas can be generally applied to other ML frameworks and optimization compilers. Experimental results show that AStitch can achieve an average of 1.84x speedup (up to 2.73x) over the state-of-the-art Google's XLA solution across five production workloads. We also deploy AStitch onto a production cluster for ML workloads with thousands of GPUs. The system has been in operation for more than 10 months and saves about 20,000 GPU hours for 70,000 tasks per week. Zhen Zheng, Xuanda Yang, Pengzhan Zhao, Guoping Long, Kai Zhu 0004, Feiwen Zhu, Wenyi Zhao, Jun Yang 0052, Jidong Zhai, Shuaiwen Song, Wei Lin 0016 |
ASPLOS | 1 |
| 2022 | Whale: Efficient Giant Model Training over Heterogeneous GPUs
Xianyan Jia, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang 0135, Langshi Chen, Yong Li 0045, Zhen Zheng, Wei Lin 0016 |
USENIX ATC | 10 |
| 2022 | DREW: Efficient Winograd CNN Inference with Deep ReuseabstractDeep learning has been used in various domains, including Web services. Convolutional neural networks (CNNs), which are deep learning representatives, are among the most popular neural networks in Web systems. However, CNN employs a high degree of computing. In comparison to the training phase, the inference process is more frequently done on low-power computing equipments. The limited computing resource and high computation pressure limit the effective use of CNN algorithms in industry. Fortunately, a minimal filtering algorithm called Winograd can reduce convolution calculations by minimizing multiplication operations. We find that Winograd convolution can be sped up further by deep reuse technique, which reuses the similar data and computation processes. In this paper, we propose a new inference method, called DREW, which combines deep reuse with Winograd for further accelerating CNNs. DREW handles three difficulties. First, it can detect the similarities from the complex minimal filtering patterns by clustering. Second, it reduces the online clustering cost in a reasonable range. Third, it provides an adjustable method in clustering granularity balancing the performance and accuracy. Experiments show that 1) DREW further accelerates the Winograd convolution by an average of 2.06 × speedup; 2) when DREW is applied to end-to-end Winograd CNN inference, it achieves 1.71 × the average performance speedup with no (<0.4%) accuracy loss; 3) DREW reduces the number of convolution operations to 11% of the original operations on average. Feng Zhang 0007, Jiawei Guan, Zhen Zheng, Xiaoyong Du 0001, Xipeng Shen |
WWW | 4 |
| 2022 | Optimizing DNN Compilation for Distributed Training With Joint OP and Tensor FusionabstractThis article proposesDisCo, an automatic deep learning compilation module for data-parallel distributed training. Unlike most deep learning compilers that focus on training or inference on a single device,DisCooptimizes a DNN model for distributed training over multiple GPU machines. Existing single-device compilation strategies do not work well in distributed training, due mainly to communication inefficiency that they incur.DisCogenerates optimized, joint computation operator and communication tensor fusion strategies to enable highly efficient distributed training. A GNN-based simulator is built to effectively estimate per-iteration training time achieved by operator/tensor fusion candidates. A backtracking search algorithm is driven by the simulator, navigating efficiently in the large strategy space to identify good operator/tensor fusion strategies that minimize distributed training time. We compareDisCowith existing DL fusion schemes and show that it achieves good training speed-up close to the ideal, full computation-communication overlap case. Xiaodong Yi 0001, Shiwei Zhang 0002, Lansong Diao, Chuan Wu 0001, Zhen Zheng, Shiqing Fan, Siyu Wang 0006, Jun Yang 0052, Wei Lin 0016 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | DAPPLE: a pipelined data parallel approach for training large modelsabstractIt is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an effective approach for improving device utilization. However, there are still several tricky issues to address: improving computing efficiency while ensuring convergence, and reducing memory usage without incurring additional computing costs. We propose DAPPLE, a synchronous training framework which combines data parallelism and pipeline parallelism for large DNN models. It features a novel parallelization strategy planner to solve the partition and placement problems, and explores the optimal hybrid strategies of data and pipeline parallelism. We also propose a new runtime scheduling algorithm to reduce device memory usage, which is orthogonal to re-computation approach and does not come at the expense of training throughput. Experiments show that DAPPLE planner consistently outperforms strategies generated by PipeDream's planner by up to 3.23× speedup under synchronous training scenarios, and DAPPLE runtime outperforms GPipe by 1.6× speedup of training throughput and saves 12% of memory consumption at the same time. Shiqing Fan, Zongyan Cao, Siyu Wang 0006, Zhen Zheng, Chuan Wu 0001, Guoping Long, Jun Yang 0052, Lixue Xia, Lansong Diao, Wei Lin 0016 |
PPoPP | 6 |
| 2021 | Understanding and bridging the gaps in current GNN performance optimizationsabstractGraph Neural Network (GNN) has recently drawn a rapid increase of interest in many domains for its effectiveness in learning over graphs. Maximizing its performance is essential for many tasks, but remains preliminarily understood. In this work, we provide an in-depth examination of the state-of-the-art GNN frameworks, revealing five major gaps in the current frameworks in optimizing GNN performance, especially in handling the special complexities of GNN over traditional graph or DNN operations. Based on the insights, we put together a set of optimizations to fill the gaps. These optimizations leverage the state-of-the-art GPU optimization techniques and tailor them to the special properties of GNN. Experimental results show that these optimizations achieve 1.37×--15.5× performance improvement over the state-of-the-art frameworks on various GNN models. Kezhao Huang, Jidong Zhai, Zhen Zheng, Youngmin Yi, Xipeng Shen |
PPoPP | 3 |
| 2021 | Exploring deep reuse in winograd CNN inferenceabstractConvolutional neural networks (CNNs), as representatives of deep learning, are one of the most commonly used neural networks in applications such as graphic image analysis. However, CNN has heavy computation patterns; network training processes could take several hours even with modern processors. Different from the training process, the inference process is more often executed on devices with low computing power, such as CPUs. Fortunately, a minimal filtering algorithm, Winograd, can reduce the convolution computations by reducing the number of multiplication operations. We find that the Winograd convolution can be further accelerated by reusing the similar data and computation patterns, which is called deep reuse. Feng Zhang 0007, Zhen Zheng, Xiaoyong Du 0001, Xipeng Shen |
PPoPP | 3 |
| 2020 | GOPipe: A Granularity-Oblivious Programming Framework for Pipelined Stencil Executions on GPUabstractRecent studies have shown promising performance benefits when multiple stages of a pipelined stencil application are mapped to different parts of a GPU to run concurrently. An important factor for the computing efficiency of such pipelines is the granularity of a task. In previous programming frameworks that support true pipelined computations on GPU, the choice has to be made by the programmers during the application development time. Due to many difficulties, programmers' decisions are often far from optimal, causing inferior performance and performance portability. Chanyoung Oh, Zhen Zheng, Xipeng Shen, Jidong Zhai, Youngmin Yi |
PACT | 2 |
| 2020 | Optimizing distributed training deployment in heterogeneous GPU clustersabstractThis paper proposes HeteroG, an automatic module to accelerate deep neural network training in heterogeneous GPU clusters. To train a deep learning model with large amounts of data, distributed training using data or model parallelism has been widely adopted, mostly over homogeneous devices (GPUs, network bandwidth). Heterogeneous training environments may often exist in shared clusters with GPUs of different models purchased in different batches and network connections of different bandwidth availability (e.g., due to contention). Classic data parallelism does not work well in a heterogeneous cluster, while model-parallel training is hard to plan. HeteroG enables highly-efficient distributed training over heterogeneous devices, by automatically converting a single-GPU training model to a distributed one according to the deep learning graph and available resources. HeteroG embraces operation-level hybrid parallelism, communication architecture selection and execution scheduling, based on a carefully designed strategy framework exploiting both GNN-based learning and combinatorial optimization. We compare HeteroG with existing parallelism schemes and show that it achieves up-to 222% training speed-up. HeteroG also enables efficient training of large models over a set of heterogeneous devices where simple parallelism is infeasible. Xiaodong Yi 0001, Shiwei Zhang 0002, Ziyue Luo, Guoping Long, Lansong Diao, Chuan Wu 0001, Zhen Zheng, Jun Yang 0052, Wei Lin 0016 |
CoNEXT | 7 |
| 2019 | HiWayLib: A Software Framework for Enabling High Performance Communications for Heterogeneous Pipeline ComputationsabstractPipeline is a parallel computing model underpinning a class of important applications running on CPU-GPU heterogeneous systems. A critical aspect for the efficiency of such applications is the support of communications among pipeline stages that may reside on CPU and different parts of a GPU. Existing libraries of concurrent data structures do not meet the needs, due to the massive parallelism on GPU and the complexities in CPU-GPU memory and connections. This work gives an in-depth study on the communication problem. It identifies three key issues, namely, slow and error-prone detection of the end of pipeline processing, intensive queue contentions on GPU, and cumbersome inter-device data movements. This work offers solutions to each of the issues, and integrates all together to form a unified library named HiWayLib. Experiments show that HiWayLib significantly boosts the efficiency of pipeline communications in CPU-GPU heterogeneous applications. For real-world applications, HiWayLib produces 1.22~2.13× speedups over the state-of-art implementations with little extra programming effort required. Zhen Zheng, Chanyoung Oh, Jidong Zhai, Xipeng Shen, Youngmin Yi |
ASPLOS | 1 |
| 2019 | GOPipe: a granularity-oblivious programming framework for pipelined stencil executions on GPUabstractRecent studies have shown promising performance benefits of pipelined stencil applications. An important factor for the computing efficiency of such pipelines is the granularity of a task. We presents GOPipe, the first granularity-oblivious programming framework for efficient pipelined stencil executions. With GOPipe, programmers no longer need to specify the appropriate task granularity. GOPipe automatically finds it, and schedules tasks of that granularity while observing all inter-task and inter-stage data dependencies. In our experiments on four real-life applications, GOPipe outperforms the state-of-the-art by up to 4.57× with a much better programming productivity. Chanyoung Oh, Zhen Zheng, Xipeng Shen, Jidong Zhai, Youngmin Yi |
PPoPP | 2 |
| 2017 | Versapipe: a versatile programming framework for pipelined computing on GPUabstractPipeline is an important programming pattern, while GPU, designed mostly for data-level parallel executions, lacks an efficient mechanism to support pipeline programming and executions. This paper provides a systematic examination of various existing pipeline execution models on GPU, and analyzes their strengths and weaknesses. To address their shortcomings, this paper then proposes three new execution models equipped with much improved controllability, including a hybrid model that is capable of getting the strengths of all. These insights ultimately lead to the development of a software programming framework named VersaPipe. With VersaPipe, users only need to write the operations for each pipeline stage. VersaPipe will then automatically assemble the stages into a hybrid execution model and configure it to achieve the best performance. Experiments on a set of pipeline benchmarks and a real-world face detection application show that VersaPipe produces up to 6.90X (2.88X on average) speedups over the original manual implementations. Zhen Zheng, Chanyoung Oh, Jidong Zhai, Xipeng Shen, Youngmin Yi |
MICRO | 1 |
| 2016 | Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputerabstractThis paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day. Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002 |
SC | 17 |
| 2014 | Holistic Modeling and Performance Evaluation for Converged Network-Cloud Service ProvisioningabstractThe crucial role that computer networks play in Cloud computing requires a holistic vision of both networking and computing resources that leads to a convergence of network and Cloud service provisioning. Performance evaluation on converged network-Cloud service provisioning faces some new challenges that are brought in by key features of Cloud computing, such as resource virtualization and service-orientation. In order to address the challenges, we explored application of network calculus theory in analyzing network-Cloud service performance by proposing a novel profile-based modeling approach. In this paper we further extend this approach by particularly considering the impact of data transform between networking and computing systems in network-Cloud service provisioning, thus achieving holistic modeling and analysis on end-to-end service performance across the network and Cloud domains. The developed analytic method is general and applicable to a wide variety of network-Cloud services. The results obtained in this paper also provide useful guidelines for federated management of networking and computing resources for achieving service performance guarantees. Zhen Zheng |
AINA | 2 |
| 2012 | Hybrid Synchronization of Two Delayed Systems with Uncertain Parameters
Zhen Zheng, Qunfang Wang |
ISNN (1) | 1 |