Yuyang Jin 0001

dblp:259/5264-1 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-2358-3395ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 8 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse
abstract
Recent advances in AI-generated content have driven widespread adoption of Diffusion Transformers (DiTs) for high-resolution, long-duration content generation. While parallelization techniques accelerate diffusion inference, they face significant scalability challenges due to excessive communication overhead in multi-node environments. We observe that sequence partitions in Context Parallelism (CP) exhibit distinct heterogeneity: spatially proximate partitions contribute more significantly to attention computation results. By mapping this heterogeneous pattern to hierarchical communication topology, we can access high-contribution partitions with reduced communication cost. This insight motivates our novel selective attention state mechanism that strategically balances partial attention computation and historical result reuse across denoising steps. We present DiTango, an efficient parallel framework for DiT generation. DiTango features an anchor-guided state selection planner that optimizes computation-reuse decisions for each partition, complemented by a runtime that orchestrates efficient state-centric operations. This design achieves superior system efficiency while preserving generation quality. Experimental evaluation on popular diffusion models demonstrates that DiTango achieves up to 1.9x end-to-end and 3.2x attention speedup with near-linear scaling in multi-node settings, while maintaining generation quality comparable to state-of-the-art approaches.
Runxin Zhong, Zan Zong, Hengjie Li, Yuyang Jin 0001, Jidong Zhai
HPDC5
2026 SYCL++: A Unified Programming Framework for Heterogeneous Supercomputers at Scale
Zitao Shen, Yuyang Jin 0001, Kinman Lei, Zixuan Ma, Zhenchuan Chen, Di Wei, Fei Wang 0096, Ying Liu 0055, Lin Gan 0001, Jidong Zhai
HPDC2
2025 An Efficient 2D Fusion Method for High-Performance Two-Stage Eigensolvers on Modern Heterogeneous Architectures
abstract
Solving a significant portion of the eigensystem is a critical problem in numerical linear algebra and is widely applied in real-world applications.As problem sizes increase, the twostage tridiagonalization method has emerged as the state-ofthe-art approach and has been implemented in well-known libraries such as LAPACK, PLASMA, and MAGMA.Its major performance bottleneck is the tridiagonal-to-band back transformation of eigenvectors (st2sb) due to the dilemma between limited operational intensity and excessive computational cost.This challenge is further exacerbated by the growing imbalance between computational speed and memory bandwidth in modern heterogeneous architectures.To address this challenge, this paper introduces a 2D Fusion method to decouple the operational intensity from the computational cost of st2sb.To reduce the intrinsic overhead of 2D Fusion for large fusion factors, we further propose an effective skipping strategy.Our 2D Fusion enhances the performance of all existing two-stage eigensolvers without loss of accuracy.We evaluated the effectiveness of 2D Fusion in MAGMA and LAPACK across various problem sizes: on the Nvidia A100 GPU, 2D Fusion improves the performance of eigenvalue decomposition in MAGMA by an average speedup of 1.06× for matrices larger than 24k×24k
Yongxiao Zhou, Yi Zong, Yuyang Jin 0001, Wei Xue 0003
ICS3
2025 HSampler : Optimizing Multi-GPU GNN Sampling with Collision-Avoid Selection
Yuyang Jin 0001, Jidong Zhai, Kezhao Huang
NPC (1)1
2025 FlashTensor: Optimizing Tensor Programs by Leveraging Fine-grained Tensor Property
abstract
Deep neural networks (DNNs) have shown significant effectiveness in natural language processing and video applications. However, DNN models, especially for long-context tasks, introduce extremely large intermediate tensors, producing substantial memory overhead. Although considerable efforts have been made to optimize DNNs, insufficient awareness of tensor properties has hindered effective memory optimization and can lead to inefficient computations in a long-context scenario.
Runxin Zhong, Yuyang Jin 0001, Chen Zhang 0001, Kinman Lei, Shuangyu Li, Jidong Zhai
PPoPP2
2025 TraceFlow: Efficient Trace Analysis for Large-Scale Parallel Applications via Interaction Pattern-Aware Trace Distribution
abstract
Trace analysis of large-scale parallel applications is crucial for understanding and optimizing performance. It primarily focuses on the interaction behaviors between different parallel processes, such as synchronization waits and asynchronous overlaps. The trace size explodes as the parallel scale of applications, thus current methods analyze traces in parallel to ensure analysis speed. However, due to the interaction pattern-agnostic trace distribution, they often introduce inter-process communications to fetch non-local event data during interaction analysis, leading to excessively long trace analysis time.
Yuyang Jin 0001, Xirui Shui, Mingshu Zhai, Zan Zong, Feng Zhang 0007, Felix Wolf 0001, Jidong Zhai
SC1
2025 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
abstract
Long-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication.
Zan Zong, Yuyang Jin 0001, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai
SC3
2025 mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai
USENIX ATC7
2025 Leveraging Graph Analysis to Pinpoint Root Causes of Scalability Issues for Parallel Applications
abstract
It is challenging to scale parallel applications to modern supercomputers because of load imbalance, resource contention, and communications between processes. Profiling and tracing are two main performance analysis approaches for detecting these scalability bottlenecks. Profiling is low-cost but lacks detailed dependence for identifying root causes. Tracing records plentiful information but incurs significant overheads. To address these issues, we presentScalAna, which employs static analysis techniques to combine the benefits of profiling and tracing - it enables tracing's analyzability with overhead similar to profiling.ScalAnauses static analysis to capture program structures and data dependence of parallel applications, and leverages lightweight profiling approaches to record performance data during runtime. Then a parallel performance graph is generated with both static and dynamic data. Based on this graph, we design a backtracking detection approach to automatically pinpoint the root causes of scaling issues. We evaluate the efficacy and efficiency ofScalAnausing several real applications with up to 704K lines of code and demonstrate that our approach can effectively pinpoint the root causes of scaling loss with an average overhead of 5.65% for up to 16,384 processes. By fixing the root causes detected by our tool, it achieves up to 33.01% performance improvement.
Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Zhenhua Guo 0003, Yaqian Zhao, Torsten Hoefler, Tao Liu 0029, Xu Liu 0001, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.1
2024 WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and Operations
abstract
Graph Neural Network (GNN) has emerged as an important workload for learning on graphs. With the size of graph data and the complexity of GNN model architectures increasing, developing an efficient GNN system grows more important. As GNN has heavy neural computation workloads on a large graph, it is crucial to partition the entire workload into smaller parts for parallel execution and optimization. However, existing approaches separately partition graph data and GNN operations, resulting in inefficiency and large data movement overhead.
Kezhao Huang, Jidong Zhai, Liyan Zheng 0001, Haojie Wang 0004, Yuyang Jin 0001, Qihao Zhang, Runqing Zhang, Zhen Zheng, Youngmin Yi, Xipeng Shen
EuroSys5
2024 BoostN: Optimizing Imbalanced Neighborhood Communication on Homogeneous Many-Core System
abstract
MPI neighborhood communication with sparse and imbalanced patterns is common in process-level parallel programs. However, these programs often encounter significant performance slowdowns in today’s many-core clusters that feature dozens of cores per node. There are two key causes for this slowdown. First, there is substantial competition for memory and network ports when a large number of processes simultaneously access the MPI library. Second, many neighborhood communications do not align well with the many-core architecture, resulting in performance bottlenecks that could have been mitigated.
Haopeng Huang, Yuyang Jin 0001, Wei Xue 0003
ICPP2
2024 PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
Kinman Lei, Yuyang Jin 0001, Mingshu Zhai, Kezhao Huang, Haoxing Ye, Jidong Zhai
USENIX ATC2
2024 Graph-Centric Performance Analysis for Large-Scale Parallel Applications
abstract
Performance analysis is essential for understanding the performance behaviors of parallel programs and detecting performance bottlenecks. Whereas, complex interconnections across several types of performance bugs, as well as inter-process communications and data dependence, make efficient performance analysis even more difficult. Despite the fact that many performance tools have been developed, accurately identifying underlying performance bottlenecks for such complex scenarios requires specific in-depth analysis. Significant human efforts and analysis knowledge are often required to implement each specific analytic task. To alleviate the complexity of developing specific performance analytic tasks, we present a programmable performance analysis tool, calledPerFlow. InPerFlow, a step-by-step performance analysis process is represented as an Analysis Flow Diagram, which is constructed with several performance analysis sub-tasks, namely passes, that can be defined by developers or provided byPerFlow's built-in analysis pass library. Furthermore, we define a Performance Abstraction Graph to describe the performance behavior of a parallel program, where the edges indicate the interactions between parallel units, therefore the analytic sub-tasks are converted to graph analysis tasks.PerFlowprovides plentiful Python APIs for developing analytic tasks. Several case studies of real-world applications with up to 700 K lines of code are used to demonstrate the effectiveness ofPerFlow. The results indicate thatPerFlowmakes it much easier to implement specific performance analytic tasks, and these tasks are performed automatically and efficiently to detect underlying performance bottlenecks.
Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Xia Liao, Feng Zhang 0007, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.1
2024 Efficient Inference for Pruned CNN Models on Mobile Devices With Holistic Sparsity Alignment
abstract
Many artificial intelligence applications based on convolutional neural networks are directly deployed on mobile devices to avoid network unavailability and user privacy leakage. However, the significant increase in model parameter volumes makes it difficult to achieve high-performance convolutional neural network inference on these mobile devices with limited computing power. Weight pruning is one of the main approaches to compress models by reducing model parameters and computational operations, which also introduces irregular sparsity of neural networks, leading to inefficient computation and memory access during inference. This work proposes an end-to-end framework, namely MCPruner, for efficient inference of pruned convolutional neural networks on mobile devices by aligning the sparse patterns with hardware execution features in computation, memory access, and parallelism. It first co-designs pruning methods and code generation optimizations for the alignment of non-zero weight count and vector width, to improve computational efficiency while ensuring accuracy. During the code generation, it applies a sparse pattern-aware format to reduce inefficient memory accesses. Besides, convolution computations are reordered for alignment, and then mapped to parallel threads on accelerated units to achieve high parallelism. Experimental results using several commonly used models and datasets on the ARM-based Hikey970 demonstrate that our work outperforms state-of-the-art methods in inference efficiency, with no accuracy degradation.
Yuyang Jin 0001, Runxin Zhong, Saiqin Long, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.1
2023 Unified Programming Models for Heterogeneous High-Performance Computers
Zixuan Ma, Yuyang Jin 0001, Shizhi Tang, Haojie Wang 0004, Wei-Cheng Xue, Jidong Zhai
J. Comput. Sci. Technol.2
2022 PerFlow: a domain specific framework for automatic performance analysis of parallel applications
abstract
Performance analysis is widely used to identify performance issues of parallel applications. However, complex communications and data dependence, as well as the interactions between different kinds of performance issues make high-efficiency performance analysis even harder. Although a large number of performance tools have been designed, accurately pinpointing root causes for such complex performance issues still needs specific in-depth analysis. To implement each such analysis, significant human efforts and domain knowledge are normally required.
Yuyang Jin 0001, Haojie Wang 0004, Runxin Zhong, Chen Zhang 0001, Jidong Zhai
PPoPP1
2022 Vapro: performance variance detection and diagnosis for production-run parallel applications
abstract
Performance variance is a serious problem for parallel applications, which can cause performance degradation and make applications' behavior hard to understand. Therefore, detecting and diagnosing performance variance are of crucial importance for users and application developers. However, previous detection approaches either bring too large overhead and hurt applications' performance, or rely on nontrivial source code analysis that is impractical for production-run parallel applications.
Liyan Zheng 0001, Jidong Zhai, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song
PPoPP6
2022 Detecting Performance Variance for Parallel Applications Without Source Code
abstract
For parallel applications, performance variance is a critical issue that can degrade performance and make applications’ behavior difficult to explain. Therefore, users and application developers should be able to detect and diagnose performance variance. Previous detection methods either introduce too much overhead and slow down applications, or rely on nontrivial source code analysis, which is impractical for production-run parallel systems. In this article, we proposeVapro, a framework for detecting and diagnosing performance variance in production-run parallel systems. Our method is based on an observation that most parallel programs contain code snippets that are executed repeatedly with a fixed workload and can be utilized to detect performance variance. We present State Transition Graph (STG) to track program execution and then do light-weight workload analysis on STG to locate performance variance.Vaprois able to successfully identify these snippets at runtime even without program source code. To diagnose the discovered variation,Vaprouses a progressive diagnosis method based on a hybrid model combining variance breakdown and statistical analysis. According to evaluating results,Vapro's performance overhead is only 1.38% on average.Vaprocan identify performance variance in real applications caused by hardware issues, such as memory and IO. The standard deviation of the execution time is decreased by up to 73.5% when the identified variance is fixed.Vaproachieves 30.0% larger detection coverage than the state-of-the-art variance detection approach based on source code analysis.
Jidong Zhai, Liyan Zheng 0001, Feng Zhang 0007, Xiongchao Tang, Haojie Wang 0004, Yuyang Jin 0001, Shuaiwen Song
IEEE Trans. Parallel Distributed Syst.7
2020 Identifying scalability bottlenecks for large-scale parallel programs with graph analysis
abstract
Scaling a parallel program to modern supercomputers is challenging due to inter-process communication, code serialization, and resource contention. Performance analysis tools for finding such scaling bottlenecks either base on profiling or tracing. Profiling incurs lower overheads but does not capture detailed dependencies needed for root-cause analyses. Tracing collects all information at prohibitive overheads. In this work, we develop ScalAna that uses static analysis techniques to achieve the best of both worlds---it enables the analyzability of traces at a cost similar to profiling. We leverage compiler and runtime lightweight techniques to generate performance graph and perform graph analysis algorithm to detect the root cause of scaling issues. We evaluate ScalAna with real applications on the Tianhe-2 supercomputer. Results show that our approach can effectively locate the root cause of scalability bottlenecks for real applications and incur less than 6.38% overhead (1.89% on average) for up to 2,048 processes.
Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Torsten Hoefler, Xu Liu 0001, Jidong Zhai
PPoPP1
2020 ScalAna: automating scaling loss detection with graph analysis
abstract
Scaling a parallel program to modern supercomputers is challenging due to inter-process communication, Amdahl’s law, and resource contention. Performance analysis tools for finding such scaling bottlenecks either base on profiling or tracing. Profiling incurs low overheads but does not capture detailed dependencies needed for root-cause analysis. Tracing collects all information at prohibitive overheads. In this work, we design SCALANA that uses static analysis techniques to achieve the best of both worlds - it enables the analyzability of traces at a cost similar to profiling. SCALANA first leverages static compiler techniques to build a Program Structure Graph, which records the main computation and communication patterns as well as the program’s control structures. At runtime, we adopt lightweight techniques to collect performance data according to the graph structure and generate a Program Performance Graph. With this graph, we propose a novel approach, called backtracking root cause detection, which can automatically and efficiently detect the root cause of scaling loss. We evaluate SCALANA with real applications. Results show that our approach can effectively locate the root cause of scaling loss for real applications and incurs 1.73parcent overhead on average for up to 2,048 processes. We achieve up to 11.11parcent performance improvement by fixing the root causes detected by SCALANA on 2,048 processes.
Yuyang Jin 0001, Haojie Wang 0004, Xiongchao Tang, Torsten Hoefler, Xu Liu 0001, Jidong Zhai
SC1