Changqing Shi

dblp:380/0876 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0009-5732-2207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CKTI: A Domain-Specific Compiler for Lowering CUDA Kernels to Triton-IR
abstract
CUDA kernels are essential for high-performance computing, yet their deployment has been limited to vendor-specific chips. Unless highly efficient computational kernels are custom-implemented by experts, other chips may face problems such as low utilization and inability to accelerate AI computing. In this paper, we introduce CKTI, a domain-specific compiler that lowers CUDA kernels to Triton-IR, thereby decoupling them from proprietary hardware and fostering diversity across the computing ecosystem. CKTI proposes a scheduling algorithm that transforms the threading model from thread-level to tile-level, along with a mapping scheme from explicit control-structure characteristics to dynamic masks. It also incorporates a custom dialect to express the complete semantics and optimizable properties of the kernel. These collectively ensure performance portability and cross-platform deployment. The results show that CKTI produces correct outputs on different hardware platforms and delivers competitive performance, achieving 1.28X speedup on NVIDIA, 1.17X on AMD, 1.14X on MetaX, and unlocking deployment on Cambricon platforms. Additionally, we validate CKTI’s support for end-to-end workloads across multiple architecture, including NPUs and GPGPUs.
Changqing Shi, Yicheng Sui, Yudong Xie, Shuangpeng Ming
ICS1
2025 Fixing Broken Graphs: LLM-Powered Automatic Code Optimization for DNN Programs
abstract
Deep learning compilers optimize DNN program execution by capturing them as operator-based computation graphs. However, developers’ deep learning programs often contain complex Python language features that prevent compilers from recognizing the entire program as a complete computation graph, resulting in sub-optimal performance. Our analysis reveals that actual capture failures involve only a few lines of code, we believe this problem can be addressed through code repair rather than extensive compiler improvements. To address this challenge, we introduce GraphGlue, a multi-agent system that leverages LLMs to repair and optimize DNN programs for compiler requirements, thereby maximizing the performance benefits of deep learning compilers in inference scenarios. GraphGlue employs (1) graph-break cause mining (GCM) to identify hidden causes of computation graph breaks and facilitate LLM-based repair, and (2) self-correction with reject sampling (SRS) to alternate between code debugging and regeneration, effectively avoiding ineffective feedback attempts caused by incorrect initial optimization strategies. Experimental results demonstrate that programs optimized by GraphGlue achieve up to 2.19x (1.23x on average) speedup compared to using TorchDynamo directly, and deliver up to 15.77x (8.74x on average) memory savings compared to state-of-the-art AI compiler frontends. GraphGlue exhibits strong generalization capabilities across 1,411 real-world user programs, successfully optimizing 92.63% of them. Code is available at https://github.com/Jamesswang/GraphGlue.
Yicheng Sui, Yudong Xie, Changqing Shi
ASE6
2025 TransCL: An Automatic CUDA-to-OpenCL Programs Transformation Framework
abstract
With the rising demand for computational power and the increasing variety of computational scenarios, considerable interest has emerged in transforming existing CUDA programs into more general-purpose OpenCL programs, enabling them to run across diverse hardware platforms. However, manual methods, typically designed for specific applications, lack flexibility. Current automated conversion techniques also face considerable challenges, particularly in handling diverse programming interfaces, memory management, and so on, and are insufficient for converting large-scale, complex CUDA projects. In this article, we propose a novel source-to-source program transformation framework, TransCL, which automates the conversion of CUDA programs in four key aspects: source code, execution model, programming model, and memory model. To achieve this, we abstract a set of conversion rules aligned with the latest CUDA standards, develop a transcoder, implement an OpenCL-compatible programming interface library, and establish a memory mapping mechanism between CUDA and OpenCL. Experiments demonstrate that TransCL provides a high level of automation in converting CUDA-based applications and is effective in handling large, complex projects such as TensorFlow. Moreover, the converted AI framework successfully conducted model training for the first time. The experiment also validates that the converted program can execute correctly across multiple platforms and demonstrate good performance.
Changqing Shi, Chunye Gong, Yicheng Sui, Yutong Jin
ACM Trans. Archit. Code Optim.1
2024 oclCUB: an OpenCL parallel computing library for deep learning operators
Changqing Shi, Yicheng Sui, Yuqiao Chen
CCF Trans. High Perform. Comput.1
2024 Opencl-pytorch: an OpenCL-based extension of PyTorch
Yicheng Sui, Changqing Shi
CCF Trans. High Perform. Comput.3