Chunwei Xia

dblp:189/8618 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-2014-5453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D Vectorization
abstract
CUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators.
Shuaijiang Li, Ying Liu 0055, Shuoming Zhang, Yijin Li, Yangyu Zhang, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng 0002, Huimin Cui
CGO11
2026 Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Yangyu Zhang, Zhaolin Duan, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui
ISCA11
2026 KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
abstract
Language models (LMs) underpin emerging mobile and embedded AI applications like meeting and video summarization and document analysis, which often require processing multiple long-context inputs. Running an LM locally on-device improves privacy, enables offline use, and reduces cost, but long-context inference quickly hits a \emph{memory capacity wall} as the key-value (KV) cache grows linearly with context length and batch size. Existing KV-cache offloading schemes are designed to transfer cache data from GPU memory to CPU memory; however, they are not suitable for embedded and mobile systems, where the CPU and GPU (or NPU) typically share a unified memory and the non-volatile secondary storage (disk) offers limited I/O bandwidth. We present KVSwap, a software framework tailored for local devices that achieves high memory efficiency while effectively leveraging disk storage. KVSwap stores the full cache on disk, uses highly compact in-memory metadata to predict which entries to preload, overlaps computation with hardware-aware disk access, and orchestrates read patterns to match storage device characteristics. Our evaluation shows that across representative LMs and storage types, KVSwap delivers higher throughput under tight memory budgets while maintaining generation quality over existing KV cache offloading schemes.
Chunwei Xia, Zheng Wang 0001
MobiSys2
2026 LEGO-compiler: enhancing neural compilation through translation composability
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Yunji Chen, Xiaobing Feng 0002, Huimin Cui
CCF Trans. High Perform. Comput.4
2026 The new compiler stack: a survey on the synergy of LLMs and compilers
Shuoming Zhang, Qiuchu Yu, Chunwei Xia, Zheng Wang 0001, Xiaobing Feng 0002, Huimin Cui
CCF Trans. High Perform. Comput.4
2025 Accelerating Tensor-Train Decomposition on Graph Neural Networks
abstract
Memory footprint is a major concern when training graph neural networks (GNNs) on large graph data. Tensor-train decomposition (TTD) offers a potential solution by representing high-dimensional tensors with a set of smaller tensors, reducing memory overhead. However, existing TTD-based solutions for GNNs fail to reuse intermediate computation results and minimize memory data transfers to improve GNN performance. We introduce FALCON, a software framework to accelerate TTDbased GNN training. FALCON leverages the observation that a small subset of graph nodes with high edge degrees are frequently accessed, enabling the caching of intermediate results to reduce redundant computation and data transfers. Additionally, it incorporates multi-level graph partitioning and kernel optimization techniques to boost computational efficiency. We evaluated FALCON using three real-world datasets on three GPU platforms-NVIDIA 3090, 4090, and A100. Experimental results show that FALCON outperforms previous TTD-based frameworks, delivering a 1.3 to$8.17 \times$improvement in throughput while maintaining comparable or better efficiencies in memory footprint and model accuracy.
Shenghao Qiu, Chunwei Xia, Zheng Wang 0001
IPDPS2
2025 Leveraging Compilation Statistics for Compiler Phase Ordering
abstract
Choosing the optimal order and combination of compiler optimization passes - known as phase ordering - can enhance the performance of compiled binaries. However, existing approaches struggle to capture the subtle interaction between compiler passes and waste time on low-profitable pass sequences. We introduce CITROEN, a better approach for compiler phase ordering. CITROEN leverages pass-related compilation statistics to reject low-profitable compiler pass sequences to reduce the overhead of phase ordering search. It employs Bayesian optimization to navigate the search space, using compilation statistics instead of traditional tuning parameters to build an online cost model that provides both the performance prediction and the prediction uncertainty of compilation configurations. It dynamically allocates search iterations across source files to optimize search time in multi-file programs. We evaluate CITROEN by integrating it with the LLVM compiler and applying it to benchmarks from cBench and SPEC CPU 2017. CITROEN outperforms existing autotuning methods, discovering high-performing configurations quicker with fewer search iterations.
Chunwei Xia, Zheng Wang 0001
IPDPS2
2024 Optimizing Deep Learning Inference via Global Analysis and Tensor Expressions
abstract
Optimizing deep neural network (DNN) execution is important but becomes increasingly difficult as DNN complexity grows. Existing DNN compilers cannot effectively exploit optimization opportunities across operator boundaries, leaving room for improvement. To address this challenge, we present Souffle, an open-source compiler that optimizes DNN inference across operator boundaries. Souffle creates a global tensor dependency graph using tensor expressions, traces data flow and tensor information, and partitions the computation graph into subprograms based on dataflow analysis and resource constraints. Within a subprogram, Souffle performs local optimization via semantic-preserving transformations, finds an optimized program schedule, and improves instruction-level parallelism and data reuse. We evaluated Souffle using six representative DNN models on an NVIDIA A100 GPU. Experimental results show that Souffle consistently outperforms six state-of-the-art DNN optimizers by delivering a geometric mean speedup of up to 3.7× over TensorRT and 7.8× over Tensorflow XLA.
Chunwei Xia, Qianqi Sun, Zheng Wang 0001, Yuan Wen, Xiaobing Feng 0002, Huimin Cui
ASPLOS (1)1
2024 LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital Accelerator
abstract
Bio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks.
Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang
ICCAD13
2024 Combining Structured Static Code Information and Dynamic Symbolic Traces for Software Vulnerability Prediction
abstract
Deep learning (DL) has emerged as a viable means for identifying software bugs and vulnerabilities. The success of DL relies on having a suitable representation of the problem domain. However, existing DL-based solutions for learning program representations have limitations - they either cannot capture the deep, precise program semantics or suffer from poor scalability. We present Concoction, the first DL system to learn program presentations by combining static source code information and dynamic program execution traces. Concoction employs unsupervised active learning techniques to determine a subset of important paths to collect dynamic symbolic execution traces. By implementing a focused symbolic execution solution, Concoction brings the benefits of static and dynamic code features while reducing the expensive symbolic execution overhead. We integrate Concoction with fuzzing techniques to detect function-level code vulnerabilities in C programs from 20 open-source projects. In 200 hours of automated concurrent test runs, Concoction has successfully uncovered vulnerabilities in all tested projects, identifying 54 unique vulnerabilities and yielding 37 new, unique CVE IDs. Concoction also significantly outperforms 16 prior methods by providing higher accuracy and lower false positive rates.
Huanting Wang, Zhanyong Tang, Shin Hwei Tan, Jie Wang 0110, Hejun Fang, Chunwei Xia, Zheng Wang 0001
ICSE7
2021 ChaoPIM: A PIM-based Protection Framework for DNN Accelerators Using Chaotic Encryption
abstract
Although deep neural networks (DNNs) have been widely used, DNN models running on ASIC- or FPGA-based accelerators still lack effective and efficient protection. Once DNN models are stolen by attackers, it will not only infringe the intellectual property of model providers but also lead to security issues. The existing parameter encryption method brings greater power consumption, which is difficult to apply to resource-constrained edge devices. This paper proposes an effective and efficient framework –ChaoPIM to protect the security of DNN models by utilizing the chaotic encryption and the Processing-In-Memory (PIM) technology. Detailed experimental results show that our framework can effectively prevent attackers from using DNN models normally, as the accuracy of stolen models is quite low. Compared with the powerful Cortex-A53, Kryo-280, Intel-i5-8265U CPUs and TITAN V GPU, ChaoPIM achieves considerable performance improvements on various DNN models.
Ning Lin, Xiaoming Chen 0003, Chunwei Xia, Jing Ye 0001, Xiaowei Li 0001
ATS3
2020 DNNTune: Automatic Benchmarking DNN Models for Mobile-cloud Computing
abstract
Deep Neural Networks (DNNs) are now increasingly adopted in a variety of Artificial Intelligence (AI) applications. Meantime, more and more DNNs are moving from cloud to the mobile devices, as emerging AI chips are integrated into mobiles. Therefore, the DNN models can be deployed in the cloud, on the mobile devices, or even mobile-cloud coordinate processing, making it a big challenge to select an optimal deployment strategy under specific objectives. This article proposes a DNN tuning framework, i.e., DNNTune, that can provide layer-wise behavior analysis across a number of platforms. Using DNNTune, this article further selects 13 representative DNN models, including CNN, LSTM, and MLP, and three mobile devices ranging from low-end to high-end, and two AI accelerator chips to characterize the DNN models on these devices to further assist users finding opportunities for mobile-cloud coordinate computing. Our experimental results demonstrate that DNNTune can find a coordinated deployment achieving up to 1.66× speedup and 15× energy saving comparing with mobile-only and cloud-only deployment.
Chunwei Xia, Huimin Cui, Xiaobing Feng 0002, Jingling Xue
ACM Trans. Archit. Code Optim.1
2018 On Retargeting the AI Programming Framework to New Hardwares
Yisong Chang, Denghui Li, Chunwei Xia, Huimin Cui, Ke Zhang 0017, Xiaobing Feng 0002
NPC4