Anxing Xie

dblp:329/1129 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0003-5331-3174ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 GAS: A scheduling primitive dependency analysis-based cost model for tensor program optimization
abstract
Automatically generating high-performance tensor programs has become a promising approach for deploying deep neural networks. A key challenge lies in designing an effective cost model to navigate the vast scheduling search space. Existing approaches typically fall into two categories, each with limitations: offline learning cost models rely on large pre-collected datasets, which may be incomplete or device-specific, and online learning cost models depend on handcrafted features, requiring substantial manual effort and expertise. We propose GAS, a lightweight framework for generating tensor programs for deep learning applications. GAS reformulates feature extraction as a sequence-dependent analysis of scheduling primitives. Our cost model integrates three key factors to uncover performance-critical insights within scheduling sequences: (1) decision factors allocation, quantifying entropy and skewness of scheduling primitive factors to capture their dominance; (2) primitive contribution weights, measuring the relative impact of primitives on overall performance; and (3) structural semantic alignment, capturing correlations between scheduling primitive factors and hardware parallelism mechanisms. This approach reduces the complexity of handcrafted feature engineering and extensive pre-training datasets, significantly improving both efficiency and scalability. Experimental results on NVIDIA GPUs demonstrate that GAS achieves average speedups of 3.79 × over AMOS and 2.22 × over Ansor, while also consistently outperforming other state-of-the-art tensor compilers.
Yonghua Hu, Anxing Xie, Zenghua Cheng, Junyang Tang
J. Syst. Archit.2
2025 GTA: Generating high-performance tensorized program with dual-task scheduling
Anxing Xie, Yonghua Hu, Yuxiang Gao, Zenghua Cheng
J. Syst. Archit.1
2025 Instruction selection optimization for VLIW architecture based on classification node merging
Fangjun Liu, Huifu Zhang, Yonghua Hu, Anxing Xie, Shangfeng Mo
J. Supercomput.4
2024 SORA: Rapid Software Pipelining Optimization after Register Allocation for Vector-DSPs
abstract
Heterogeneous multi-zone processors are key building blocks of high performance computing (HPC), and fully utilizing its hardware resources is essential to unlocking its full power. Software pipelining is a family of techniques that enhance program execution efficiency by accelerating the execution speed of loop programs through parallel execution of instructions from different loop bodies. However, existing software pipelining methods fail to account for the unique hardware architectures of digital signal processors (DSPs), resulting in suboptimal performance when applied directly.In this paper, we present SORA, a rapid Software pipelining Optimization architecture performing efficient instruction scheduling after Register Allocation for the Vector-DSPs. In addition, SORA conducts a comprehensive analysis encompassing instruction type, data flow, functional unit utilization, and register conflict, tailored to accommodate Vector-DSPs. Experimental results demonstrate that SORA can achieve 2.78∼24.1× speedup in the batch normalization, 2D convolution, relu and matrix multiplication, and 1.14× in the ResNet-18 neural network to the original program. Additionally, there is a 1.05∼6.59× improvement compared to loop unrolling optimization.
Anxing Xie, Yonghua Hu
ISPA1