Xin You 0001

dblp:159/9844-1 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0002-5163-4607ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 6 first-author · 30 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Temporal Graph Network Training via Unified Redundancy Elimination
abstract
Temporal Graph Network (TGN) is increasingly adopted to model evolving relationships in dynamic graphs. However, the training pipeline is plagued by pervasive redundancy in computation, storage, and data loading. These redundancies harm computational efficiency, exacerbate memory pressure, and induce excessive CPU-GPU data transfers. We present PULSE, an end-to-end TGN training framework that systematically eliminates redundancies guided by a unified minimal-unit principle. To realize such principle, PULSE defines three synergetic units: 1) the Minimal Input Unit (MIU) for component-wise deduplication and operator-level reconstruction of redundant computations, 2) the Minimal Storage Unit (MSU) for dependency-guided message reconstruction, only preserving irreproducible entries while enabling on-demand recovery of others, and 3) the Minimal Reuse Unit (MRU) for GPU memory management, combining a BlockPool-based buffer allocator with a bipartite temporal reuse strategy to mitigate fragmentation and exploit inter-batch locality. Experimental results on representative benchmarks demonstrate that PULSE improves training throughput by up to 6.67× over the state-of-the-art baselines.
Hailong Yang 0002, Kejie Ma, Enze Yu, Xin You 0001, Qingxiao Sun, Chenhao Xie 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
ASPLOS (2)6
2026 MatrixFold: Unleashing Manycore CPUs with Outer-Product Units for Mixed-Precision AlphaFold Inference
Shaokang Du, Hailong Yang 0002, Xin You 0001, Baojian Zhou, Depei Qian 0001
Euro-Par (2)4
2026 Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor Cores
abstract
Sparse matrix-vector multiplication (SpMV) is a fundamental operation in scientific computing, machine learning, and graph analytics, demanding efficient execution on modern hardware. Recent advances in hardware accelerators, such as Tensor Cores, have significantly improved the performance of many compute-intensive workloads. However, effectively utilizing Tensor Cores for SpMV remains challenging due to its irregular sparsity patterns and the mismatch between SpMV’s computational characteristics and constrained architecture design, leading to suboptimal performance and underutilization of Tensor Cores. In this paper, we systematically analyze the state-of-the-art SpMV optimizations on Tensor Cores, identify key performance bottlenecks, and propose Drawloom, a Tensor-Core-aware framework for SpMV with efficient Tensor Core mapping and optimized pipeline execution. Drawloom leverages a redesigned Tensor Core mapping strategy with a zig-zag chained sparse storage format, as well as a multi-stage register pipeline to better exploit hardware parallelism. Our evaluation on SuiteSparse dataset demonstrates that Drawloom outperforms cuSPARSE by 2.71×/1.90× (in FP16), 2.95×/2.39× (in FP32), and 2.47×/1.54× (in FP64) on A100 and H100 GPUs, respectively. Compared to the state-of-the-art SpMV implementations, Drawloom achieves a performance speedup of 1.26×/1.18× (in FP16) and 1.49×/1.56× (in FP64) on A100 and H100 GPUs, respectively.
Kaige Zhang 0002, Hailong Yang 0002, Xin You 0001, Tianyu Feng, Yufan Xu 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
PPoPP3
2026 Spatio-Temporal Evolving Anomaly Detection Tool for Large-Scale Heterogeneous Programs Analysis
Zhibo Xuan, Hailong Yang 0002, Xin You 0001, Genshen Chu, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.4
2025 Identifying Potential Anomalous Operations in Graph Neural Network Training
Zhibo Xuan, Hailong Yang 0002, Xin You 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
APPT3
2025 Accelerating the Martian Atmospheric Simulation of GoMars Model with Multi-GPUs
abstract
Mars exploration is at the forefront of space science, which demands robust computational models to decipher its atmospheric dynamics. In this work, we present a significant advancement in computational efficiency for the GoPlanetMars (GoMars), a state-of-the-art Martian atmospheric model. By leveraging the parallel processing capabilities of Graphics Processing Units (GPUs), we accelerate the dynamic core of the GoMars model on multiple NVIDIA A800 GPUs. Through comprehensive performance analysis of GoMars, we optimize both parallel computation and communication patterns to leverage the computational power of multiple GPUs fully, achieving performance comparable to that of a thousand-core CPU cluster. Our evaluation results demonstrate that the GPU-accelerated GoMars model maintains the same level of precision as the native CPU-based implementation, while achieving a substantial speedup, making it a viable solution for high-performance Martian atmospheric simulations.
Guofan Yu, Haoran Kong, Xin You 0001, Hailong Yang 0002, Shaokang Du, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
HPCC3
2025 ESC: Effective Submanifold Convolution using Tensor Cores
abstract
Submanifold convolution is an effective method to process 3D point cloud data, playing a significant role in fields such as robotics, autonomous driving, and AR/VR. However, due to the high sparsity and irregularity of point cloud data, it is challenging to accelerate submanifold convolution on modern GPUs, especially using tensor cores. Previous works have proposed implicit GEMM methods to accelerate submanifold convolution on GPU. However, the performance of such methods is limited by massive redundant computation and suboptimal parameter configurations. In this paper, we propose ESC, a new method to leverage GPU tensor cores for accelerating submanifold convolution with improved performance. Firstly, we propose an online similarity-aware reordering method to increase the point cloud data locality and yield more opportunities for eliminating redundancy. Secondly, we propose TC-aware redundancy elimination to reduce the redundant computation at the fine TC-tile granularity. Moreover, we propose an adaptive configuration selector to select the optimal configuration based on offline profiling results and online input data. Experimental results demonstrate that ESC outperforms the state-of-the-art works on representative datasets.
Hailong Yang 0002, Xin You 0001, Yufan Xu 0001, Kaige Zhang 0002, Mingzhen Li 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
ICPP3
2025 Efficient Locality-aware Instruction Stream Scheduling for Stencil Computation on ARM Processors
abstract
Stencil computation is one of the fundamental computational patterns in scientific computing, commonly adopted in solving partial differential equations (PDEs) and a wide range of application fields.However, due to the memory-bound nature, it is challenging to achieve satisfactory performance on the ARM many-core processors with complex computation and memory hierarchies.In this study, we propose independent instruction stream scheduling with the Serial-FMA to Tree-Based Reduction (SFTBR) technique to decompose the stencil computation into multiple independent instruction streams for improved instruction-level parallelism.Furthermore, we propose a locality-aware block scheduling technique for locality-aware multi-level thread parallelism to address the complexities of cache and memory hierarchies on modern ARM many-core processors.Based on the above techniques, we implement a domain-specific compiler, AOStencil, to automatically generate optimized stencil codes on ARM many-core processors with genetic-algorithm-driven parameter tuning.Our evaluation results demonstrate that AOStencil achieves up to 4.39× speedup over the state-ofthe-art domain-specific compilers on Kunpeng and Phytium platforms.
Shanghao Liu, Hailong Yang 0002, Xin You 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
ICS3
2025 GNNPerf: Towards Effective Performance Profiling and Analysis Across GNN Frameworks
abstract
Graph Neural Networks (GNNs) have been successfully adopted in various application domains and accelerated by parallel processors such as GPUs. Despite the existence of popular frameworks such as Deep Graph Library (DGL) and PyTorch Geometric (PyG), the inconsistent programming paradigms and the lack of a unified analysis toolkit both hinder effective performance comparison among different GNN frameworks. This missing capability not only complicates the selection of the most suitable framework for users, but also impedes developers from optimizing framework implementations. In this paper, we propose GNNPerf, a performance profiling and analysis toolkit for effective performance comparison across GNN frameworks. GNNPerf provides a domain-specific language enabling unified GNN design expression and automatic generation to frameworkspecific implementations. GNNPerf also provides full workflow support for comprehensively evaluating GNN models with easy-to-use profiling, visualization, and analysis. The experimental results demonstrate that the GNNPerf can identify performance bottlenecks and empower users to derive actionable insights, enhancing both GNN model design and framework implementation.
Kejie Ma, Hailong Yang 0002, Zizheng Zhang, Xin You 0001, Zhibo Xuan, Qingxiao Sun, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
IPDPS4
2025 Exploiting Transformer-Based Static Binary Analysis for Identifying Inefficient Locks
Zhibo Xuan, Xin You 0001, Hailong Yang 0002, Jingqi Chen, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
NPC (1)2
2025 Zero-Value Code Specialization via Profile-Guided Control Data Flow Analysis
abstract
Zero-value propagation is a common phenomenon in modern programs, where redundant operations caused by zero-values can severely impact performance. Since zero-values are often generated dynamically at runtime, eliminating such redundancies through static analysis alone is challenging. In this paper, we propose an efficient static control data flow analysis algorithm to identify redundancies resulting from zero-value propagation. Based on this algorithm, we design and implement ZeroSpec, a fully automated profile-guided code optimizer that detects zero-values at runtime and specializes fast paths for them. To maximize performance gains, ZeroSpec also employs a fine-grained cost model that evaluates the optimization potential of individual zero-value instructions to guide the construction of targeted optimization regions. Evaluation on SPEC CPU2017, NPB and real-world applications demonstrates the effectiveness of ZeroSpec, achieving a maximum performance speedup of 1.31 ×.
Shaokang Du, Kelun Lei, Xin You 0001, Hailong Yang 0002, Yufan Xu 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
SC3
2025 Towards Efficient LLM Inference via Collective and Adaptive Speculative Decoding
abstract
Large language models (LLMs) have gained considerable attention for their remarkable performance across a wide range of tasks. However, efficient LLM inference remains challenging because of the autoregressive decoding process, which generates only one token at a time. Speculative decoding has been introduced to address the limitation by using small speculative models (SSMs) to speed up LLM inference. However, the low acceptance rate of SSMs and the high verification cost of LLM prohibit further performance improvement. In this paper, we present Smurfs, an LLM inference system designed to accelerate LLM inference through collective and adaptive speculative decoding. Smurfs adopts a majority-voted mechanism that harnesses multiple SSMs to collaboratively predict LLM outputs in multi-task scenarios, while avoiding high verification cost. It also decouples SSM speculation from LLM verification and uses a pipelined execution to hide the latency of SSM speculation. Additionally, Smurfs proposes a mechanism to dynamically determine the optimal speculation length of SSM at runtime, balancing the performance impact of accepted tokens and verification cost. The experimental results demonstrate the superiority of Smurfs in terms of inference throughput and latency compared to the state-of-the-art LLM inference systems.
Hailong Yang 0002, Tongxuan Liu, Yufan Xu 0001, Xuning Liang, Kejie Ma, Tianyu Feng, Xin You 0001, Ruihao Gong, Rui Wang 0014, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
SC10
2025 Hotspy: identifying performance hotspot with graph neural network based static analysis
Xin You 0001, Zhibo Xuan, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
CCF Trans. High Perform. Comput.2
2025 Exploiting Dynamic Regular Patterns in Irregular Programs for Efficient Vectorization
abstract
Modern optimizing compilers are able to exploit memory access or computation patterns to generate vectorized codes. However, such patterns in irregular programs are unknown until runtime due to the input dependence. Thus, either compiler’s static optimization or profile-guided optimization cannot represent the patterns for any common input, which leads to suboptimal vectorization. To address the above drawback, we propose DynVec , 1 a framework to automatically exploit regular patterns buried deeply inside irregular programs and apply corresponding optimizations for better vectorization. Due to the integration of workload distribution and the ability to represent instruction features and identify regular patterns with effective feature extraction and data re-arranging methods, DynVec can generate highly efficient vectorized codes for both serial and parallel irregular programs by replacing gather / scatter / reduction operations with optimized operation groups. We evaluate DynVec on optimizing irregular programs such as SpMV and graph programs with representative sparse matrix datasets. The experiment results show that DynVec achieves significant speedup compared to the state-of-the-art implementations across a range of X86 and ARM platforms.
Kelun Lei, Shaokang Du, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
ACM Trans. Archit. Code Optim.3
2025 SimTrace: Exploiting Spatial and Temporal Sampling for Large-Scale Performance Analysis
abstract
MPI tracing tools is essential to collect the communication events and performance metrics of large-scale programs for further performance analysis and optimization. However, toward the exascale era, the performance and storage overhead for tracing becomes extremely prohibitive that significantly disturbs the original execution of MPI programs, leading to distorted tracing data and thus mislead analysis results. Although process sampling can effectively reduce the tracing overhead, it can easily miss important execution information that is necessary for subsequent performance analysis. In this article, we propose SimTrace , a scalable MPI tracing tool with novel spatial and temporal sampling strategies that exploits the similarity among MPI processes to achieve both low tracing overhead as well as obtain sufficient tracing information. The experimental results demonstrate that SimTrace can significantly reduce the MPI tracing overhead compared to the state-of-the-art tracing tools, meanwhile enabling effective analysis to guide performance optimization of large-scale programs.
Zhibo Xuan, Xin You 0001, Tianyu Feng, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
ACM Trans. Archit. Code Optim.2
2025 Identifying Performance Inefficiencies of Parallel Program With Spatial and Temporal Trace Analysis
abstract
Performance inefficiencies can lead to performance anomalies in parallel programs. Existing performance analysis tools either have a limited detection scope or require significant domain knowledge to use, which constrains their practical adoption to identify performance inefficiencies. In this paper, we propose STAD, a performance analysis tool for parallel programs that considers both spatial and temporal patterns within trace data. STAD captures the spatial communication patterns between processes using a spatial communication pattern graph. It then adopts a dynamic graph neural network-based unsupervised model to learn the evolving temporal patterns along the timeline. Additionally, STAD diagnoses the root causes of performance anomalies by exploiting the aggregated feature of anomalies along the call tree. Our evaluation results demonstrate that STAD can effectively detect performance anomalies with acceptable overhead and diagnose the root causes attributed to both the program itself and the running environment.
Zhibo Xuan, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.3
2024 Retrospection on the Performance Analysis Tools for Large-Scale HPC Programs
abstract
As the performance gap between hardware and software widens, performance analysis tools are essential for understanding the behavior of large-scale High-Performance Computing (HPC) programs. These tools provide insights into the performance bottlenecks and help in optimizing the performance of the programs. In this paper, we present a comprehensive study of performance analysis tools for large-scale HPC systems including both sampling-based and instrumentation-based tools that are commonly adopted in the HPC community. We investigate the abundance and overheads of data collection as well as the analysis capabilities of HPCToolkit, TAU, and Scalasca with representative programs at scale. Our study shows that different performance analysis tools have distinct strengths and weaknesses, and the choice of a performance analysis tool depends on the specific requirements of the user. We also discuss the challenges and future directions in the field of performance analysis tools for large-scale HPC systems.
Zhibo Xuan, Xin You 0001, Hailong Yang 0002, Mingzhen Li 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
HiPC2
2024 PRoof: A Comprehensive Hierarchical Profiling Framework for Deep Neural Networks with Roofline Analysis
abstract
The increasing diversity of deep neural network (DNN) models and hardware platforms necessitates effective model profiling for high-performance inference deployment. Current DNN profiling tools suffer from either limited optimization insights due to the missing correlation between high-level DNN layer design and low-level hardware performance metrics, or prohibitive profiling overhead due to the large amount of performance measurement through hardware performance counters. Meanwhile, the roofline model has been widely used in the high-performance computing (HPC) domain for identifying performance bottlenecks and guiding optimizations. However, it lacks hierarchical (e.g., kernel/operator/layer), fine-grained, multi-platform support for profiling DNN models.
Siyu Wu 0001, Hailong Yang 0002, Xin You 0001, Ruihao Gong, Yi Liu 0013, Zhongzhi Luan, Depei Qian 0001
ICPP3
2024 GVARP: Detecting Performance Variance on Large-Scale Heterogeneous Systems
abstract
Performance variance is one of the nasty pitfalls of large-scale heterogeneous systems, which can lead to unexpected and unpredictable performance degradation for parallel programs. Such performance issues typically arise from various random hardware and software faults, making it exceedingly difficult to pinpoint the exact causes of performance variance in specific instances. In this paper, we propose GVARP, a performance variance detection tool for large-scale heterogeneous systems. GVARP employs static analysis to identify the performancecritical parameters of kernel functions. Additionally, GVARP segments the program execution with external library calls and asynchronous kernel operations. Then GVARP constructs a state transfer graph and estimates the workload of each program segment to identify and cluster instances of similar workloads, facilitating the detection of performance variance. Our evaluation results demonstrate that GVARP effectively detects performance variance at a large scale with acceptable overhead and provides intuitive insights to locate the sources of performance variance.
Xin You 0001, Zhibo Xuan, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001
SC1
2024 AtRec: Accelerating Recommendation Model Training on CPUs
abstract
The popularity of recommendation models and the enhanced AI processing capability of CPUs have provided massive performance opportunities to deliver satisfactory experiences to a large number of users. Unfortunately, existing recommendation model training methods fail to achieve high efficiency due to unique challenges such as dynamic shape and high parallelism. To address the above limitations, we comprehensively study the distinctive characteristics of recommendation models and discover several unexploited optimization opportunities. To exploit such opportunities, we proposeAtRec, a high-performant recommendation model training engine that significantly accelerates the training process on CPUs. Specifically,AtRecpresents comprehensive approach of training that employs operator-level and graph-level joint optimizations and runtime optimization. At the operator-level,AtRecidentifies and optimizes the time-consuming operators, which enables further efficient graph-level optimizations. At the graph-level,AtRecconducts an in-depth analysis of the inefficiencies in several frequently used subgraphs, enables further performance improvement via eliminating redundant computations and memory accesses. In addition, to achieve better runtime performance,AtRecalso identifies inefficiencies prevalent in the current scheduling and proposes runtime batching. The experiment results demonstrate thatAtReccan significantly outperform state-of-the-art recommendation model training engines. We have open sourced the implementation and corresponding data ofAtRecto boost research in this direction.
Tianyu Feng, Hailong Yang 0002, Xin You 0001, Bangduo Chen, Tongxuan Liu, Zhongzhi Luan, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.4
2023 VClinic: A Portable and Efficient Framework for Fine-Grained Value Profilers
abstract
Fine-grained value profilers reveal a promising way to accurately detect value-related software inefficiencies with binary instrumentation. Due to the architecture-dependent implementation details of binary instrumentation, existing value profilers suffer from poor portability as well as high engineering efforts to achieve efficiency across platforms. In this paper, we propose VClinic, a portable and efficient fine-grained value profiling framework for analyzing highly optimized binaries on both X86 and ARM platforms. VClinic exploits operand-centric two-level designs in its implementation to provide the common building blocks required for value profilers. By constructing four representative value profilers with VClinic, we demonstrate that VClinic can ease the development of value profilers with portability and efficiency across platforms. Guided by the value profilers built upon VClinic, we can achieve up to 89.94% and 74.66% speedup for real-world programs on X86 and ARM platforms, respectively.
Xin You 0001, Hailong Yang 0002, Kelun Lei, Zhongzhi Luan, Depei Qian 0001
ASPLOS (2)1
2023 Efficient Deep Molecular Dynamic Model Training on Heterogeneous System
abstract
Molecular dynamics is a widely adopted simulation method for analyzing the movement of atoms and molecules. Traditional molecular dynamics simulation methods are computationally intensive and difficult to simulate a large number of atoms. In contrast, molecular dynamics based on deep potential models such as DeePMD can leverage deep learning techniques to improve simulation efficiency. Although DeePMD has incorporated mainstream deep learning frameworks, it still suffers from low performance and efficiency during its model training on heterogeneous systems such as CPU and GPU. Particularly, a large number of operators cannot be accelerated by GPU, resulting in low utilization of GPU computational resources. In this paper, we comprehensively analyze the computational bottlenecks and the corresponding root causes of DeePMD. We correspondingly propose several novel optimization strategies. Specifically, for preprocessing, we identify the computation redundancies and the GPU parallelization opportunities for performance optimization. For training, we propose optimization strategies such as operator fusion, redundancy elimination, and concurrent execution of multiple streams and threads in the computation process. Moreover, we apply systematical optimization of computational graphs and operators. The evaluation results show that DeePMD can achieve significant speedups in several cases after applying our proposed optimizations, resulting in a maximum overall speedup of 6.36× with acceptable accuracy.
Shaokang Du, Xin You 0001, Hailong Yang 0002, Jing Shang 0001, Zhiwen Xiao, Zhongzhi Luan, Depei Qian 0001
ICPADS2
2023 Accelerating Big Data Application by Eliminating Redundancy on Hadoop Cluster
abstract
Big data applications are widely adopted to mine valuable information from a tremendous amount of industry data, which is commonly represented as a series of map-reduce operations. Among various map-reduce frameworks, Hadoop is most commonly adopted for data processing at large scale. Although Hadoop eases the development of highly scalable distributed big data applications, inefficient implementation due to poor coding practice and deep software abstractions can cause severe performance issues such as unreasonable slowdown, high response latency, and waste of computing resources, which can lead to unsatisfactory serving delay or significant maintenance cost. In this paper, we first categorize three common types of redundant patterns in big data applications. Then we propose a tool-assisted optimization workflow to detect the redundant patterns automatically, which profiles the application by sampling hardware performance monitoring units. Moreover, we present a profiling visualization method that can help to pinpoint the redundant codes. Based on these approaches, we optimize several big data applications by eliminating redundancies, yielding up to 14.8% performance improvement.
Kelun Lei, Shaokang Du, Xin You 0001, Zhibo Xuan, Haoran Kong, Hailong Yang 0002, Jing Shang 0001, Zhiwen Xiao, Zhongzhi Luan, Depei Qian 0001
ICPADS3
2023 BiRFIA: Selective Binary Rewriting for Function Interception on ARM
abstract
Function interception of fully-optimized binaries is widely used for optimization with its ability to accurately collect runtime information and detect inefficiencies at the function level. However, the implementation of function interception with existing binary rewriting techniques still suffers from limited reliability and performance on ARM platform. In this paper, we propose BiRFIA, an efficient selective binary rewriting framework for function interception targeting highly optimized binaries on ARM platforms. BiRFIA performs static binary rewriting of specific functions and intercepts them through well-formed trampoline sections and external instrumentation libraries. Besides, BiRFIA places complex instrumentation code in the trampoline section and jumps to the trampoline section via an adaptive instruction eviction strategy, which significantly reduces the probability of unexpected errors. For evaluation, we develop two function interception tools based on BiRFIA, including a function performance event counter collector and a function parameter tracer. Guided by these tools, we optimize several benchmarks and real-world programs, yielding up to 8% performance speedup. Our evaluation result demonstrates that BiRFIA incurs negligible runtime overhead of 1.006× on average.
Kelun Lei, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
ICS2
2023 TrivialSpy: Identifying Software Triviality via Fine-grained and Dataflow-based Value Profiling
abstract
Trivial operations cause software inefficiencies that waste functional units and memory bandwidth for executing useless instructions. Although previous works have identified a significant amount of trivial operations in widely used programs, the proposed solutions only provide useful observations, other than actionable guidance to eliminate trivial operations for better performance. In this paper, we propose TrivialSpy - a fine-grained and dataflow-based value profiler to effectively identify software triviality with optimization potential estimation. With the help of dataflow analysis, TrivialSpy can detect software trivialities of heavy operation, trivial chain, and redundant backward slice. In addition, TrivialSpy can identify trivial breakpoints that combine multiple trivial conditions for more optimization opportunities. The evaluation results demonstrate TrivialSpy is capable of identifying software triviality in highly optimized programs. Based on the optimization guidance provided by TrivialSpy, we can achieve 52.09% performance speedup at maximum after eliminating trivial operations.
Xin You 0001, Hailong Yang 0002, Kelun Lei, Zhongzhi Luan, Depei Qian 0001
SC1
2022 Vectorizing SpMV by Exploiting Dynamic Regular Patterns
abstract
Modern optimizing compilers can exploit memory access and computation patterns to generate vectorized codes. However, such patterns in irregular programs such as SpMV are unknown until runtime due to the input dependence. Thus, either compiler’s static optimization or profile-guided optimization cannot represent the patterns for any common input, which leads to suboptimal vectorization. To address the above drawback, we propose DynVec, a framework to automatically exploit regular patterns buried deeply inside SpMV programs and apply corresponding optimizations for better vectorization. Due to the ability to represent instruction features and identify regular patterns with effective feature extraction and data re-arranging methods, DynVec can generate highly efficient vectorized codes by replacing gather/scatter/reduction operations with optimized operation groups. We evaluate DynVec on optimizing SpMV with representative sparse matrix datasets. The experiment results show that DynVec achieves significant speedup compared to the state-of-the-art SpMV implementations across a range of platforms.
Xin You 0001, Changxi Liu, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
ICPP1
2022 PowerSpector: Towards Energy Efficiency with Calling-Context-Aware Profiling
abstract
Energy efficiency has become one of the major concerns in high-performance computing systems towards exascale. On mainstream systems, dynamic voltage and frequency scaling (DVFS) and uncore frequency scaling (UFS) are two popular techniques to trade-off performance and power consumption to achieve better energy efficiency. However, the existing system software is oblivious to application characteristics and thus misses the opportunity for fine-grained power management. Meanwhile, manually instrumenting applications with power management codes are prohibitive due to heavy engineering efforts and thus hardly portable across platforms. In this paper, we propose Powerspector, a fine-grained code profiling and optimization tool with calling context awareness to automatically explore the opportunity for optimizing energy efficiency. The design of Powerspector consists of three phases, including significant region detection, performance profiling and power modeling, and frequency optimization. The first phase automatically identifies the profitable regions for frequency optimization. Then, the second phase guides the core/uncore frequency optimization with power models. The third phase injects frequency optimization codes targeting each significant code region across different calling contexts automatically. The experiment results demonstrate that Powerspector can achieve 1.13×(1.00×), 1.28×(1.09×), and 1.17×(1.06×) improvement on energy efficiency compared to static(region-based) tuning on Haswell, Broadwell, and Skylake platforms, respectively.
Xin You 0001, Hailong Yang 0002, Zhibo Xuan, Zhongzhi Luan, Depei Qian 0001
IPDPS1
2022 Accelerating the cryo-EM structure determination in RELION on GPU cluster
Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
Frontiers Comput. Sci.1
2021 Automatic Code Generation and Optimization of Large-scale Stencil Computation on Many-core Processors
abstract
Stencil computation is an indispensable building block of many scientific applications and is widely used by the numerical solvers of partial differential equations (PDEs). Due to the complex computation patterns of different stencils and the various hardware targets (e.g., many-core processors), many domain-specific languages (DSLs) have been proposed to optimize stencil computation. However, existing stencil DSLs mostly focus on the performance optimizations on homogeneous many-core processors such as CPUs and GPUs, and fail to embrace emerging heterogeneous many-core processors such as Sunway. In addition, few of them can support expressing stencil with multiple time dependencies and optimizations from both spatial and temporal dimensions. Moreover, most stencil DSLs are unable to generate codes that can run efficiently in large scale, which limits their practical applicability. In this paper, we propose MSC, a new stencil DSL designed to express stencil computation in both spatial and temporal dimensions. It can generate high-performance stencil codes for large-scale execution on emerging many-core processors. Specially, we design several optimization primitives for improving parallelism and data locality, and a communication library for efficient halo exchange in large scale execution. The experiment results show that our MSC achieves better performance compared to the state-of-the-art stencil DSLs.
Mingzhen Li 0001, Yi Liu 0013, Hailong Yang 0002, Yongmin Hu, Qingxiao Sun, Bangduo Chen, Xin You 0001, Zhongzhi Luan, Depei Qian 0001
ICPP7
2021 dgQuEST: Accelerating Large Scale Quantum Circuit Simulation through Hybrid CPU-GPU Memory Hierarchies
Tianyu Feng, Xin You 0001, Shuzhang Zhong, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
NPC3
2021 The Deep Learning Compiler: A Comprehensive Survey
abstract
The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers have been proposed from both industry and academia such as Tensorflow XLA and TVM. Similarly, the DL compilers take the DL models described in different DL frameworks as input, and then generate optimized codes for diverse DL hardware as output. However, none of the existing survey has analyzed the unique design architecture of the DL compilers comprehensively. In this paper, we perform a comprehensive survey of existing DL compilers by dissecting the commonly adopted design in details, with emphasis on the DL oriented multi-level IRs, and frontend/backend optimizations. Specifically, we provide a comprehensive comparison among existing DL compilers from various aspects. In addition, we present detailed analysis on the design of multi-level IRs and illustrate the commonly adopted optimization techniques. Finally, several insights are highlighted as the potential research directions of DL compiler. This is the first survey paper focusing on the design architecture of DL compilers, which we hope can pave the road for future research towards DL compiler.
Mingzhen Li 0001, Yi Liu 0013, Qingxiao Sun, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.5
2020 Towards GPU Acceleration of Phonon Computation with ShengBTE
abstract
ShengBTE is one of the software packages that are commonly used in the field of phonon computation (e.g., to determine the lattice thermal conductivity). ShengBTE simulates the phonon diffusion by solving the Boltzmann transport equations, which take long execution time to derive the simulation results due to the high computation complexity. This paper mainly focuses on the performance optimization of ShengBTE on GPU. We identify the performance bottlenecks of ShengBTE and propose corresponding optimizations such as loop-carried dependency elimination, hotspot function acceleration on GPU and performance tuning on thread block. The experiment results show that the proposed optimizations significantly improve the performance of ShengBTE, which achieves an average speedup of 9.06x and 13.74x on discrete temperature simulation and continuous temperature simulation respectively without losing accuracy.
Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001
HPC Asia2
2020 Accelerating De Novo Assembler WTDBG2 on Commodity Servers
Ming Dun, Yunchun Li, Xin You 0001, Qingxiao Sun, Zerong Luan, Hailong Yang 0002
ICA3PP (1)3
2020 ZeroSpy: exploring software inefficiency with redundant zeros
abstract
Redundant zeros cause inefficiencies in which the zero values are loaded and computed repeatedly, resulting in unnecessary memory traffic and identity computation that waste memory bandwidth and CPU resources. optimizing compilers is difficult in eliminating these zero-related inefficiencies due to limitations in static analysis. Hardware approaches, in contrast, optimize inefficiencies without code modification, but are not widely adopted in commodity processors. In this paper, we propose ZeroSpy - a fine-grained profiler to identify redundant zeros caused by both inappropriate use of data structures and useless computation. ZeroSpy also provides intuitive optimization guidance by revealing the locations where the redundant zeros happen in source lines and calling contexts. The experimental results demonstrate ZeroSpy is capable of identifying redundant zeros in programs that have been highly optimized for years. Based on the optimization guidance revealed by ZeroSpy, we can achieve significant speedups after eliminating redundant zeros.
Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001, Xu Liu 0001
SC1
2019 Improving the Parallelism of CESM on GPU
Zehui Jin, Ming Dun, Xin You 0001, Hailong Yang 0002, Yunchun Li, Yingchun Lin, Zhongzhi Luan, Depei Qian 0001
ICA3PP (2)3
2018 swCaffe: A Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
abstract
This paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core heterogeneous architecture, with 40,960 SW26010 processors connected through a customized communication network.First, we point out some insightful principles to fully exploit the performance of the innovative many-core architecture.Second, we propose a set of optimization strategies for redesigning a variety of neural network layers based on Caffe.Third, we put forward a topology-aware parameter synchronization scheme to scale the synchronous Stochastic Gradient Descent (SGD) method to multiple processors efficiently.We evaluate our framework by training a variety of widely used neural networks with the ImageNet dataset.On a single node, swCaffe can achieve 23%˜119% overall performance compared with Caffe running on K40m GPU.As compared with the Caffe on CPU, swCaffe runs 3.04˜7.84xfaster on all the networks.Finally, we present the scalability of swCaffe for training of ResNet-50 and AlexNet on the scale of 1024 nodes.
Liandeng Li, Jiarui Fang, Haohuan Fu, Jinlei Jiang, Wenlai Zhao, Conghui He, Xin You 0001, Guangwen Yang 0002
CLUSTER7