Changxin Li

dblp:218/6934 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0003-8901-5187ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Low-Latency and Memory-Efficient Deployment of Irregular Sparse Deep Learning Workloads
abstract
Our work introduces a tile-aware scheduling framework for efficient sparse Vision Transformer execution on GPUs. Sparse attention reduces the cost of high-resolution Vision Transformers, but its irregular masks produce blocks with diverse sizes, densities, and locations. Existing FlashAttention-style kernels rely on fixed tile configurations and cannot fully exploit these sparse patterns, leading to wasted computation and underutilized GPU resources. Our framework bridges this gap through four key techniques: sparse attention is represented as an adjacency matrix, structure-aware reordering algorithms improve locality, locally dense blocks are extracted as scheduling units, and offline profiling with integer linear programming (ILP) selects hardware-feasible tile assignments. Results show that our inference scheduler achieves up to 2.13 × end-to-end speedup over fixed-tile FlashAttention and up to 4.6 × speedup in high-resolution images. We further introduce a training-aware extension that reuses the inference tile schedule and augments it with backward computation and activation-memory strategies.
Changxin Li
HPDC1
2026 Achieving Low Latency Inference on High Resolution Images by Exploiting Sparsity in Vision Transformers
Changxin Li, Sanmukh R. Kuppannagari
IPDPS1
2025 Optimizing Deployment of Unstructured Group Convolutions for Low Latency Inference
abstract
Group convolutions are widely adopted in modern CNN architectures such as CondenseNet and ShuffleNet to enable efficient inference on GPUs. However, when the connection pattern between input and output channels does not exhibit regularity (unstructured group convolution), popular deep learning frameworks (e.g., PyTorch) often struggle with load balancing and data reuse issues leading to reduced performance. In this paper, we present a comprehensive optimization framework that combines a Knapsack-based partitioning approach with Integer Linear Programming (ILP) and advanced matrix reordering to optimize the deployment of unstructured group convolutions which are used in popular models such as CondenseNets that learn group connections. Specifically, we use knapsack algorithm to determine partition (group) sizes for the connections to minimize execution time and use an Integer Linear Programming (ILP) to assign connections to the partitions (group) output by the knapsack algorithm. We also employ three matrix reordering strategies-Hierarchical Clustering (HC), Iterative Clustering (IC), and Reverse Cuthill-McKee (RCM) on the matrix representing the input-output connectivity pattern to further improve the performance of our scheduling algorithm. Our experiments on ShuffleNet and CondenseNet demonstrate up to$1.9 \times$speedups over PyTorch. Furthermore, augmenting ILP with reordering achieves an additional$1.3 \times$improvement demonstrating the importance of optimizing for load balancing and data reuse.
Changxin Li, Sanmukh R. Kuppannagari
HiPC1
2025 A Chinese Heart Failure Status Speech Database with Universal and Personalised Classification
Changxin Li, Xingyao Wang 0001, Yili Xia, Hanyue Zhang
INTERSPEECH3
2024 Exploring Algorithmic Design Choices for Low Latency CNN Deployment
abstract
Convolutional Neural Networks (CNN s) have demonstrated significant success in advancing image and video processing technologies, significantly outperforming traditional methods in both accuracy and efficiency. However, deploying CNN s effectively across diverse hardware platforms often faces the challenge of latency, which can critically impact real-time processing applications. In this work, we explore algorithmic design choices aimed at reducing latency in CNN deployments. We implement five convolution algorithms using SYCL and inte-grate them into three popular CNN models: VGG 16, Resnetl0l, and Inception V 4. By replacing the standard PyTorch Conv2d function with our SY CL- based implementations, we evaluate the execution time of each convolution layer and the overall model on G PU s. Our extensive experiments benchmark the performance of these algorithms against the baseline implementations of the PyTorch and Pytorch Extension for Intel. The results demonstrate significant improvements in execution time, underscoring the potential of these algorithmic choices for achieving low latency in CNN deployments.
Changxin Li, Sanmukh R. Kuppannagari
HiPC1
2018 Online customer reviews and consumer evaluation: The role of review font
Yunhui Huang, Changxin Li, Zhijie Lin 0002
Inf. Manag.2