Xinxin Qi

dblp:307/3602 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-8316-2934ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core Processor
abstract
Transformer-based large language models (LLM) are increasingly deployed in high-performance computing environments, where the attention mechanism often becomes a key bottleneck during inference. Although state-of-the-art attention algorithms (e.g., FlashAttention) achieve high efficiency on GPUs, they are ill-suited to emerging heterogeneous many-core processors. In this work, we focus on MT-3000, a representative architecture deployed in the new-generation Tianhe supercomputer, and identify three principal challenges in realizing high-performance attention: complex multi-tier memory requiring manual data movement, excessive reduction overhead caused by sub-tile softmax operations, and static execution pipelines that fail to adapt to inference phases and sequence lengths. To overcome these challenges, we propose DeferAttention , a high-performance attention implementation designed for the MT-3000 many-core processor. DeferAttention introduces a novel deferred-reduction attention strategy to decouple reduction from the fused compute pipeline, enabling more efficient aggregation over large tiles. Moreover, DeferAttention adopts a memory-centric operator design, including data tiling, multi-level software pipelining, and modular micro-kernels, to maximize data reuse and execution throughput. Finally, to support runtime-adaptive execution, DeferAttention integrates a lightweight kernel selection strategy guided by an analytical cost model. Experimental results show that DeferAttention achieves up to 98% of the theoretical peak at the micro-kernel level and 85% at the operator level, outperforming baseline implementations and significantly accelerating end-to-end inference.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che
ACM Trans. Archit. Code Optim.1
2026 mtGEMM: An Efficient GEMM Library for Modern Multi-Core DSPs
abstract
The General Matrix Multiplication (GEMM) is a crucial subprogram in high-performance computing (HPC). With the increasing importance of power and energy consumption, modern Digital Signal Processors (DSPs) are being integrated into general-purpose HPC systems. However, due to architecture disparities, traditional optimizations for CPUs and GPUs are not easily applicable to modern DSPs. This paper shares our experience of optimizing the GEMM operation using a CPU-DSP platform as a case study. Our work employs a set of strategies to improve the performance and scalability of GEMM. These strategies focus on developing micro-kernels based on heterogeneous on-chip memory, addressing the memory access bottleneck in multi-core parallelism, and facilitating efficient transpose-GEMM. These approaches, collectively referred to as an efficient and practical library (a.k.a.mtGEMM), maximize computational capabilities and bandwidth utilization of multi-core DSPs, while achieving high performance for variously-shaped GEMMs. Our experimental results demonstrate thatmtGEMMcan attain between 92% and 96% of the hardware peak, with the multi-core scalability being almost linear.
Jianbin Fang, Kainan Yu, Peng Zhang 0061, Dezun Dong, Xinxin Qi, Xingyu Hou, Ruibo Wang, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.5
2025 Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core Processor
abstract
Optimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Jie Ren 0007
SC1
2024 Optimizing Stencil Computation on Multi-core DSPs
abstract
Stencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart.
Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007
ICPP2
2024 Optimizing General Matrix Multiplications on Modern Multi-core DSPs
abstract
General Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware.
Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001
IPDPS2
2023 PowerDis: Fine-Grained Power Monitoring Through Power Disaggregation Model
Xinxin Qi, Juan Chen 0001, Rongyu Deng, Yuan Yuan 0034, Yonggang Che
ICA3PP (4)1
2023 HighRPM: Combining Integrated Measurement and Sofware Power Modeling for High-Resolution Power Monitoring
abstract
In an era where power and energy are the first-class constraints of computing systems, accurate power information is crucial for energy efficiency optimization in parallel computing systems. Existing power monitoring techniques rely on either software-centric power models that suffer from poor accuracy or integrated hardware measurement schemes that have a low reading update frequency and coarse granularity. These result in a low spatiotemporal resolution for power monitoring. This paper introduces HighRPM, a new method for accurately measuring power consumption on parallel computing systems. HighRPM combines coarse-grained power sensor readings and software power modeling techniques to improve temporal and spatial resolutions. To provide high-frequent power readings in the temporal domain, HighRPM employs statistical modeling and machine learning techniques to predict the long-term power trend and the short-term fluctuations in power consumption. To improve spatial coverage, HighRPM takes low-time resolution node-level power consumption and uses a neural network to distribute the power readings to lower-level computing components like CPUs and memory components. We evaluate HighRPM by applying it to both ARM-based and X86-based platforms. Experimental results show that HighRPM improves time resolution by 10 times, provides accurate readings for CPUs and memory, and reduces error by 7-24% compared to other power modeling methods.
Xinxin Qi, Juan Chen 0001, Yong Dong, Yuan Yuan 0034, Tao Xu 0052, Rongyu Deng, Kexing Zhou, Zheng Wang 0001
ICPP1
2022 Collusion Attack Analysis and Detection of DPoS Consensus Mechanism
Xinxin Qi, Xiaodong Fu, Fei Dai 0002, Li Liu 0032, Jiaman Ding, Wei Peng 0004
BlockSys1
2022 AOA: Adaptive Overclocking Algorithm on CPU-GPU Heterogeneous Platforms
abstract
Abstract Although GPUs have been used to accelerate various convolutional neural network algorithms with good performance, the demand for performance improvement is still continuously increasing. CPU/GPU overclocking technology brings opportunities for further performance improvement in CPU-GPU heterogeneous platforms. However, CPU/GPU overclocking inevitably increases the power of the CPU/GPU, which is not conducive to energy conservation, energy efficiency optimization, or even system stability. How to effectively constrain the total energy to remain roughly unchanged during the CPU/GPU overclocking is a key issue in designing adaptive overclocking algorithms. There are two key factors during solving this key issue. Firstly, the dynamic power upper bound must be set to reflect the real-time behavior characteristics of the program so that algorithm can better meet the total energy unchanging constraints; secondly, instead of independently overclocking at both CPU and GPU sides, coordinately overclocking on CPU-GPU must be considered to adapt to real-time load balance for higher performance improvement and better energy constraints. This paper proposes an Adaptive Overclocking Algorithm (AOA) on CPU-GPU heterogeneous platforms to achieve the goal of performance improvement while the total energy remains roughly unchanged. AOA uses the function $$F_k$$ F k to describe the variable power upper bound and introduces the load imbalance factor W to realize the CPU-GPU coordinated overclocking. Through the verification of several types convolutional neural network algorithms on two CPU-GPU heterogeneous platforms (Intel $$^\circledR $$ ® Xeon E5-2660 & NVIDIA $$^\circledR $$ ® Tesla K80; Intel $$^\circledR $$ ® Core™i9-10920X & NIVIDIA $$^\circledR $$ ® GeForce RTX 2080Ti), AOA achieves an average of 10.7% performance improvement and 4.4% energy savings. To verify the effectiveness of the AOA, we compare AOA with other methods including automatic boost, the highest overclocking and static optimal overclocking.
Zhixin Ou, Juan Chen 0001, Tao Xu 0052, Guodong Jiang, Zhengyuan Tan, Xinxin Qi
ICA3PP7
2022 CP3: Hierarchical Cross-Platform Power/Performance Prediction Using a Transfer Learning Approach
abstract
Abstract Cross-platform power/performance prediction is becoming increasingly important due to the rapid development and variety of software and hardware architectures in an era of heterogeneous multi-core. However, accurate power/performance prediction is faced with an obstacle caused by the large gap between architectures, which is often overcome by laborious and time-consuming fine-grained program profiling on the target platform. To overcome these problems, this paper introduces $$CP^3$$ C P 3 , a hierarchical Cross-platform Power/Performance Prediction framework, which focuses on utilizing architecture differences to migrate built models to target platforms. The core of $$CP^3$$ C P 3 is the three-step hierarchical transfer learning approach, hierarchical division, partial transfer learning, and model fusion, respectively. $$CP^3$$ C P 3 firstly builds a power/performance model on the source platform, then rebuilds it with the reduced training data on the target platform, and finally obtains a cross-platform model. We validate the effectiveness of $$CP^3$$ C P 3 using a group of benchmarks on X86- and ARM-based platforms that use three different types of commonly used processors. Evaluation results show that when applying $$CP^3$$ C P 3 , only 1% of the baseline training data is required to achieve high cross-platform prediction accuracy, with power prediction error being only 0.65%, and performance prediction error being only 4.64%.
Xinxin Qi, Juan Chen 0001
ICA3PP1