EDBT 2026 Demo / reviewers in the wild / expert
Runxin Zhong
dblp:283/3468
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-8654-0192ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State ReuseabstractRecent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelization techniques accelerate diffusion inference, they face significant scalability challenges due to excessive communication overhead in multi-node environments. We observe that sequence partitions in Context Parallelism (CP) exhibit distinct heterogeneity: spatially proximate partitions contribute more significantly to attention computation results. By mapping this heterogeneous pattern to hierarchical communication topology, we can access high-contribution partitions with reduced communication cost. This insight motivates our novel selective attention state mechanism that strategically balances partial attention computation and historical result reuse across denoising steps. We present DiTango, an efficient parallel framework for DiT generation. DiTango features an anchor-guided state selection planner that optimizes computation-reuse decisions for each partition, complemented by a runtime that orchestrates efficient state-centric operations. This design achieves superior system efficiency while preserving generation quality. Experimental evaluation on popular diffusion models demonstrates that DiTango achieves up to 1.9x end-to-end and 3.2x attention speedup with near-linear scaling in multi-node settings, while maintaining generation quality comparable to state-of-the-art approaches. Runxin Zhong, Zan Zong, Hengjie Li, Yuyang Jin 0001, Jidong Zhai |
HPDC | 2 |
| 2025 | FlashTensor: Optimizing Tensor Programs by Leveraging Fine-grained Tensor PropertyabstractDeep neural networks (DNNs) have shown significant effectiveness in natural language processing and video applications. However, DNN models, especially for long-context tasks, introduce extremely large intermediate tensors, producing substantial memory overhead. Although considerable efforts have been made to optimize DNNs, insufficient awareness of tensor properties has hindered effective memory optimization and can lead to inefficient computations in a long-context scenario. Runxin Zhong, Yuyang Jin 0001, Chen Zhang 0001, Kinman Lei, Shuangyu Li, Jidong Zhai |
PPoPP | 1 |
| 2024 | MAGPY: Compiling Eager Mode DNN Programs by Monitoring Execution States
Chen Zhang 0001, Rongchao Dong, Haojie Wang 0004, Runxin Zhong, Jike Chen, Jidong Zhai |
USENIX ATC | 4 |
| 2024 | Graph-Centric Performance Analysis for Large-Scale Parallel ApplicationsabstractPerformance analysis is essential for understanding the performance behaviors of parallel programs and detecting performance bottlenecks. Whereas, complex interconnections across several types of performance bugs, as well as inter-process communications and data dependence, make efficient performance analysis even more difficult. Despite the fact that many performance tools have been developed, accurately identifying underlying performance bottlenecks for such complex scenarios requires specific in-depth analysis. Significant human efforts and analysis knowledge are often required to implement each specific analytic task. To alleviate the complexity of developing specific performance analytic tasks, we present a programmable performance analysis tool, calledPerFlow. InPerFlow, a step-by-step performance analysis process is represented as an Analysis Flow Diagram, which is constructed with several performance analysis sub-tasks, namely passes, that can be defined by developers or provided byPerFlow's built-in analysis pass library. Furthermore, we define a Performance Abstraction Graph to describe the performance behavior of a parallel program, where the edges indicate the interactions between parallel units, therefore the analytic sub-tasks are converted to graph analysis tasks.PerFlowprovides plentiful Python APIs for developing analytic tasks. Several case studies of real-world applications with up to 700 K lines of code are used to demonstrate the effectiveness ofPerFlow. The results indicate thatPerFlowmakes it much easier to implement specific performance analytic tasks, and these tasks are performed automatically and efficiently to detect underlying performance bottlenecks. Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Xia Liao, Feng Zhang 0007, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Efficient Inference for Pruned CNN Models on Mobile Devices With Holistic Sparsity AlignmentabstractMany artificial intelligence applications based on convolutional neural networks are directly deployed on mobile devices to avoid network unavailability and user privacy leakage. However, the significant increase in model parameter volumes makes it difficult to achieve high-performance convolutional neural network inference on these mobile devices with limited computing power. Weight pruning is one of the main approaches to compress models by reducing model parameters and computational operations, which also introduces irregular sparsity of neural networks, leading to inefficient computation and memory access during inference. This work proposes an end-to-end framework, namely MCPruner, for efficient inference of pruned convolutional neural networks on mobile devices by aligning the sparse patterns with hardware execution features in computation, memory access, and parallelism. It first co-designs pruning methods and code generation optimizations for the alignment of non-zero weight count and vector width, to improve computational efficiency while ensuring accuracy. During the code generation, it applies a sparse pattern-aware format to reduce inefficient memory accesses. Besides, convolution computations are reordered for alignment, and then mapped to parallel threads on accelerated units to achieve high parallelism. Experimental results using several commonly used models and datasets on the ARM-based Hikey970 demonstrate that our work outperforms state-of-the-art methods in inference efficiency, with no accuracy degradation. Yuyang Jin 0001, Runxin Zhong, Saiqin Long, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsabstractPerformance analysis is widely used to identify performance issues of parallel applications. However, complex communications and data dependence, as well as the interactions between different kinds of performance issues make high-efficiency performance analysis even harder. Although a large number of performance tools have been designed, accurately pinpointing root causes for such complex performance issues still needs specific in-depth analysis. To implement each such analysis, significant human efforts and domain knowledge are normally required. Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Jidong Zhai |
PPoPP | 3 |
| 2022 | BaGuaLu: targeting brain scale pretrained models with over 37 million coresabstractLarge-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features. Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai |
PPoPP | 17 |
| 2022 | Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Tsinghua UniversityabstractHidayetoğluet al.propose a novel memory-centric algorithm to reconstruct X-ray CT images in the SC19 article entitled “MemXCT: Memory-Centric X-ray CT Reconstruction with Massive Parallelization”. They formulate the reconstruction with several SpMVs, and propose two memory-centric optimizations to improve cache locality for better memory bandwidth utilization, i.e., a two-level pseudo-Hilbert ordering and a multi-stage input buffering. In this article, we present our results on reproducing that article to show its effectiveness and generality, as part of the SC20 Student Cluster Competition Reproducibility Challenge. We reproduce the execution time and memory bandwidth tests in that article on various architectures, including Intel CPUs, AMD CPUs, and NVIDIA GPUs. We further analyze the bottleneck on different architectures by comparing the achieved memory bandwidth with the peak bandwidth on those architectures. We then reproduce the strong scaling test on CPU and GPU clusters with different scales, and use the proposed algorithm to reconstruct three new X-ray computed tomograms. Runxin Zhong, Chen Zhang 0001, Mingshu Zhai, Lin Gan 0001, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Collaborative Heterogeneity-Aware OS Scheduler for Asymmetric Multicore ProcessorsabstractAsymmetric multicore processors (AMP) offer multiple types of cores under the same programming interface. Extracting the full potential of AMPs requires intelligent scheduling decisions, matching each thread with the right kind of core, the core that will maximize performance or minimize wasted energy for this thread. Existing OS schedulers are not up to this task. While they may handle certain aspects of asymmetry in the system, none can handle all runtime factors affecting AMPs for the general case of multi-threaded multi-programmed workloads. We address this problem by introducing COLAB, a general purpose asymmetry-aware scheduler targeting multi-threaded multi-programmed workloads. It estimates the performance and power of each thread on each type of core and identifies communication patterns and bottleneck threads. With this information, the scheduler makes coordinated core assignment and thread selection decisions that still provide each application its fair share of the processor's time. We evaluate our approach using both the GEM5 simulator on four distinct big.LITTLE configurations and a development board with ARM Cortex-A73/A53 processors and mixed workloads composed of PARSEC and SPLASH2 benchmarks. Compared to the state-of-the art Linux CFS and AMP-aware schedulers, we demonstrate performance gains of up to 25 and 5 to 15 percent on average, together with an average 5 percent energy saving depending on the hardware setup. Runxin Zhong, Vladimir Janjic, Pavlos Petoumenos, Jidong Zhai, Hugh Leather, John Thomson |
IEEE Trans. Parallel Distributed Syst. | 2 |