EDBT 2026 Demo / reviewers in the wild / expert
Xueying Wang 0003
dblp:82/1045-3
· DBLP profile ↗
24ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-7835-113XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Low-Precision Approximation of Tensor Operators on GPUs: Enabling Greater Trade-Offs between Performance and AccuracyabstractRecent GPUs integrate specialized hardware for low-precision arithmetic (e.g., FP16, INT8), offering substantial speedups for tensor operations. However, existing methods typically rely on coarse, operator-level trial-and-error tuning, which restricts the performance–accuracy trade-off space and limits achievable gains.We present Platensor, a progressive low-precision approximation framework that expands this trade-off space through ne-grained, tile-level strategies. The key idea is to exploit the tiled computation patterns of GPUs to enable flexible precision control and richer optimization opportunities. Platensor performs a two-phase exploration: a fast rule-based pass that selects promising tile-level configurations, followed by an evolutionary search that refines them. It then automatically generates optimized kernels that combine tiles of different precisions.Experiments on GEMM operators and representative applications—including kNN, LLMs, and HPL-MxP—show that Platensor significantly broadens the attainable performance– accuracy trade-offs and more fully leverages low-precision arithmetic on modern GPUs compared to operator-level tuning. Fan Luo 0003, Guangli Li, Zhaoyang Hao, Xueying Wang 0003, Xiaobing Feng 0002, Huimin Cui, Jingling Xue |
CGO | 4 |
| 2026 | DyPARS: Dynamic-Shape DNN Optimization via Pareto-Aware MCTS for Graph VariantsabstractDynamic-shape DNNs are widely used in applications such as variable-resolution image processing and language modeling with variable-length sequences. Existing DL (Deep-Learning) compilers apply rule-based rewriting to either transform a subgraph into a fixed variant at compile time (leading to suboptimal performance) or generate multiple variants at runtime, incurring significant overhead. The challenge is discovering and applying shape-dependent subgraph variants that maintain high efficiency across diverse inputs with minimal runtime cost.We propose DyPARS, a dynamic-shape DL compiler approach that discovers high-performance subgraph variants at compile time and applies the best ones at runtime. Leveraging Pareto-aware MCTS, DyPARS identifies shape-aware variants, incorporating shape-dependent kernel adaptations. These variants are integrated into a prediction-enhanced computational graph, enabling efficient variant selection based on input shapes with minimal overhead. DyPARS achieves average speedups of 1.31× and 1.80× over TorchInductor (JIT) and BladeDISC (non-JIT), respectively, across five DNN models, demonstrating robust efficiency across diverse inputs. Guangli Li, Qiuchu Yu, Xueying Wang 0003, Jingling Xue |
CGO | 4 |
| 2026 | DACOS: Dependency-Aware Cross-Kernel Overlapping for Optimizing Short-Sequence Workloads in LLM Applications
Zhaoyang Hao, Guangli Li, Fan Luo 0003, Xueying Wang 0003, Huimin Cui, Jingling Xue |
Euro-Par (2) | 5 |
| 2026 | MoonPoly: Bridging Code Generation and Adaptive Execution via Micro-Kernel Polymerization for Optimizing Dynamic-Shape Tensor OperatorsabstractThe prevalence of dynamic tensor shapes, driven by applications like language model serving with varying sequence lengths, is a defining characteristic of modern deep neural networks. This dynamism poses a fundamental challenge: reconciling the need for intensive, offline code generation to achieve peak performance with the demand for low-latency, adaptive execution to handle unpredictable runtime tensor shapes. Consequently, mainstream strategies are ineffective. Vendor-provided libraries, while highly optimized for a subset of common shapes, suffer performance degradation on unconventional ones. Static tensor compilers are hamstrung by prohibitive just-in-time compilation overheads for each new shape. While recent dynamic-shape compilers offer an alternative, they rely on predefined shape ranges, making them brittle when inputs fall outside these bounds. To resolve this tension, we present MoonPoly , a dynamic-shape tensor compiler that introduces micro-kernel polymerization . Our approach decouples these conflicting requirements through a two-stage process. In the offline stage, it performs intensive auto-tuning to generate a set of micro-kernels and corresponding performance models. The online stage then performs adaptive execution, rapidly assembling a near-optimal tensor operator on-the-fly, guided by a lightweight cost model. Evaluated on an NVIDIA A100 GPU, MoonPoly achieves an average operator-level speedup of 1.27× over the cuBLAS library across a diverse set of operators and data types, which in turn yields end-to-end inference acceleration for a variety of models, including BERT, the Vision Transformer, and large language models. Yangyu Zhang, Guangli Li, Feng Yu 0019, Fan Luo 0003, Qianqi Sun, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs
Guangli Li, Feng Yu 0019, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
Euro-Par (2) | 4 |
| 2025 | FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresabstractSparse Matrix-matrix Multiplication (SpMM) and Sampled Dense-dense Matrix Multiplication (SDDMM) are important sparse operators in scientific computing and deep learning. Tensor Core Units (TCUs) enhance modern accelerators with superior computing power, which is promising to boost the performance of matrix operators to a higher level. However, due to the irregularity of unstructured sparse data, it is difficult to deliver practical speedups on TCUs. To this end, we propose FlashSparse, a novel approach to bridge the gap between sparse workloads and the TCU architecture. Specifically, FlashSparse minimizes the sparse granularity for SpMM and SDDMM on TCUs through a novel swap-and-transpose matrix multiplication strategy. Benefiting from the minimum sparse granularity, the computation redundancy is remarkably reduced while the computing power of TCUs is fully utilized. Besides, FlashSparse is equipped with a memory-efficient thread mapping strategy for coalesced data access and a sparse matrix storage format to save memory footprint. Extensive experimental results on H100 and RTX 4090 GPUs show that FlashSparse sets a new state-of-the-art for sparse matrix multiplications (geometric mean 5.5x speedup over DTC-SpMM and 3.22x speedup over RoDe). Jinliang Shi, Shigang Li 0002, Youxuan Xu, Rongtian Fu, Xueying Wang 0003, Tong Wu 0024 |
PPoPP | 5 |
| 2025 | SparkAttention: high-performance multi-head attention for large models on Volta GPU architecture
Youxuan Xu, Tong Wu 0024, Shigang Li 0002, Xueying Wang 0003 |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUsabstractConvolutional Neural Networks (CNNs) are fundamental to advancing computer vision technologies. As CNNs become more complex and larger, optimizing model inference remains a critical challenge in both industry and academia. On modern GPU platforms, CNN operators are typically memory-bound, leading to significant performance degradation due to memory wall effects. While recent advancements have utilized operator fusion–merging multiple operators into one–to enhance inference performance, the fusion of multiple region-based operators like convolution is seldom addressed. This article introduces AFusion , a novel operator fusion technique aimed at improving inference performance, and OptiFX, an automatic optimization framework based on this approach. OptiFX employs a cost-based backtracking search to identify optimal sub-graphs for fusion and utilizes template-based code generation to create efficient kernels for these fused sub-graphs. We evaluate OptiFX across seven prominent CNN architectures–GoogLeNet, ResNet, DenseNet, MobileNet, SqueezeNet, NasNet, and UNet–on Nvidia A6000 Ada, RTX 4090, and Jetson AGX Orin platforms. Our results demonstrate that OptiFX significantly outperforms existing methods, achieving average speedups of \(2.91\times\) , \(3.30\times\) , and \(2.09\times\) in accelerating inference performance on these platforms, respectively. Xueying Wang 0003, Shigang Li 0002, Fan Luo 0003, Zhaoyang Hao, Tong Wu 0024, Ruiyuan Xu, Huimin Cui, Xiaobing Feng 0002, Guangli Li, Jingling Xue |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | Fast Convolution Meets Low Precision: Exploring Efficient Quantized Winograd Convolution on Modern CPUsabstractLow-precision computation has emerged as one of the most effective techniques for accelerating convolutional neural networks and has garnered widespread support on modern hardware. Despite its effectiveness in accelerating convolutional neural networks, low-precision computation has not been commonly applied to fast convolutions, such as the Winograd algorithm, due to numerical issues. In this article, we propose an effective quantized Winograd convolution, named LoWino, which employs an in-side quantization method in the Winograd domain to reduce the precision loss caused by transformations. Meanwhile, we present an efficient implementation that integrates well-designed optimization techniques, allowing us to fully exploit the capabilities of low-precision computation on modern CPUs. We evaluate LoWino on two Intel Xeon Scalable Processor platforms with representative convolutional layers and neural network models. The experimental results demonstrate that our approach can achieve an average of 1.84× and 1.91× operator speedups over state-of-the-art implementations in the vendor library while preserving accuracy loss at a reasonable level. Xueying Wang 0003, Guangli Li, Zhen Jia 0001, Xiaobing Feng 0002, Yida Wang 0003 |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | CoAxNN: Optimizing on-device deep learning with conditional approximate neural networks
Guangli Li, Xiu Ma, Qiuchu Yu, Lei Liu 0040, Huaxiao Liu, Xueying Wang 0003 |
J. Syst. Archit. | 6 |
| 2023 | Facilitating hardware-aware neural architecture search with learning-based predictive models
Xueying Wang 0003, Guangli Li, Xiu Ma, Xiaobing Feng 0002 |
J. Syst. Archit. | 1 |
| 2022 | Accelerating deep neural network filter pruning with mask-aware convolutional computations on modern CPUsabstractFilter pruning, a representative model compression technique , has been widely used to compress and accelerate sophisticated deep neural networks on resource-constrained platforms. Nevertheless, most studies focus on reducing the cost of model inference, whereas the heavy burden of the pruning optimization process is neglected. In this paper, we propose MaskACC, a mask-aware convolutional computation method, which accelerates the prevailing mask-based filter pruning process on modern CPU platforms. MaskACC dynamically reorganizes the tensors used in convolutions with the mask information to avoid unnecessary computations, thereby improving the computational efficiency of the pruning process. Evaluation with state-of-the-art neural network models on CPU cloud platforms demonstrates the effectiveness of our method, which achieves up to 1.61 × speedup under commonly-used pruning rates, compared to conventional computations. Xiu Ma, Guangli Li, Lei Liu 0040, Huaxiao Liu, Xueying Wang 0003 |
Neurocomputing | 5 |
| 2022 | Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning
Guangli Li, Xiu Ma, Xueying Wang 0003, Hengshan Yue, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002, Jingling Xue |
J. Syst. Archit. | 3 |
| 2022 | An Application-oblivious Memory Scheduling System for DNN AcceleratorsabstractDeep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes. Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2021 | Unleashing the Low-Precision Computation Potential of Tensor Cores on GPUsabstractTensor-specialized hardware for supporting low-precision arithmetic has become an inevitable trend due to the ever-increasing demand on computational capability and energy efficiency in intelligent applications. The main challenge faced when accelerating a tensor program on tensor-specialized hardware is how to achieve the best performance possible in reduced precision by fully utilizing its computational resources while keeping the precision loss in a controlled manner. In this paper, we address this challenge by proposing QUANTENSOR, a new approach for accelerating general-purpose tensor programs by replacing its tensor computations with low-precision quantized tensor computations on NVIDIA Tensor Cores. The key novelty is a new residual-based precision refinement technique for controlling the quantization errors, allowing tradeoffs between performance and precision to be made. Evaluation with GEMM, deep neural networks, and linear algebra applications shows that QUANTENSOR can achieve remarkable performance improvements while reducing the precision loss incurred significantly at acceptable overheads. Guangli Li, Jingling Xue, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Jiansong Li, Xiaobing Feng 0002 |
CGO | 4 |
| 2021 | Pinpointing the Memory Behaviors of DNN TrainingabstractThe training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns. Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ISPASS | 5 |
| 2020 | Accelerating Deep Learning Inference with Cross-Layer Data Reuse on GPUs
Xueying Wang 0003, Guangli Li, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002 |
Euro-Par | 1 |
| 2020 | Lance: efficient low-precision quantized winograd convolution for neural networks based on graphics processing unitsabstractAccelerating deep convolutional neural networks has become an active topic and sparked an interest in academia and industry. In this paper, we propose an efficient low-precision quan-tized Winograd convolution algorithm, called LANCE, which combines the advantages of fast convolution and quantization techniques. By embedding linear quantization operations into the Winograd-domain, the fast convolution can be performed efficiently under low-precision computation on graphics processing units. We test neural network models with LANCE on representative image classification datasets, including SVHN, CIFAR, and ImageNet. The experimental results show that our 8-bit quantized Winograd convolution improves the performance by up to 2.40× over the full-precision convolution with trivial accuracy loss. Guangli Li, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Xiaobing Feng 0002 |
ICASSP | 3 |
| 2020 | Compiler-Assisted Operator Template Library for DNN Accelerators
Jiansong Li, Wei Cao 0010, Guangli Li, Xueying Wang 0003, Lei Liu 0030, Xiaobing Feng 0002 |
NPC | 5 |
| 2020 | Fusion-Catalyzed Pruning for Optimizing Deep Learning on Intelligent Edge DevicesabstractThe increasing computational cost of deep neural network models limits the applicability of intelligent applications on resource-constrained edge devices. While a number of neural network pruning methods have been proposed to compress the models, prevailing approaches focus only on parametric operators (e.g., convolution), which may miss optimization opportunities. In this article, we present a novel fusion-catalyzed pruning approach, called FuPruner, which simultaneously optimizes the parametric and nonparametric operators for accelerating neural networks. We introduce an aggressive fusion method to equivalently transform a model, which extends the optimization space of pruning and enables nonparametric operators to be pruned in a similar manner as parametric operators, and a dynamic filter pruning method is applied to decrease the computational cost of models while retaining the accuracy requirement. Moreover, FuPruner provides configurable optimization options for controlling fusion and pruning, allowing much more flexible performance-accuracy tradeoffs to be made. Evaluation with state-of-the-art residual neural networks on five representative intelligent edge platforms, Jetson TX2, Jetson Nano, Edge tensor processing unit, neural compute stick, and neural compute stick 2, demonstrates the effectiveness of our approach, which can accelerate the inference of models on CIFAR-10 and ImageNet datasets. Guangli Li, Xiu Ma, Xueying Wang 0003, Lei Liu 0030, Jingling Xue, Xiaobing Feng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Acorns: A Framework for Accelerating Deep Neural Networks with Input SparsityabstractDeep neural networks have been employed in a broad range of applications, including face detection, natural language processing, and autonomous driving. Yet, the neural networks with the capability to tackle real-world problems are intrinsically expensive in computation, hindering the usage of these models. Sparsity in the input data of neural networks provides an optimizing opportunity. However, harnessing the potential performance improvement on modern CPU faces challenges raised by sparse computations of the neural network, such as cache-unfriendly memory accesses and efficient sparse kernel implementation. In this paper, we propose Acorns, a framework to accelerate deep neural networks with input sparsity. In Acorns, sparse input data is organized into our designed sparse data layout, which allows memory-friendly access for kernels in neural networks and opens the door for many performance-critical optimizations. Upon that, Acorns generates efficient sparse kernels for operators in neural networks from kernel templates, which combine directions that express specific optimizing transformations to be performed, and straightforward code that describes the computation. Comprehensive evaluations demonstrate Acorns can outperform state-of-the-art baselines by significant speedups. On the real-world detection task in autonomous driving, Acorns demonstrates 1.8-22.6× performance improvement over baselines. Specifically, the generated programs achieve 1.8-2.4× speedups over Intel MKL-DNN, 3.0-8.8× speedups over TensorFlow, and 11.1-13.2× speedups over Intel MKL-Sparse. Lei Liu 0030, Peng Zhao 0008, Guangli Li, Jiansong Li, Xueying Wang 0003, Xiaobing Feng 0002 |
PACT | 6 |
| 2019 | Exploiting the input sparsity to accelerate deep neural networks: posterabstractEfficient inference of deep learning models are challenging and of great value in both academic and industrial community. In this paper, we focus on exploiting the sparsity in input data to improve the performance of deep learning models. We propose an end-to-end optimization pipeline to generate programs for the inference with sparse input. The optimization pipeline contains both domain-specific and general optimization techniques and is capable of generating efficient code without relying on the off-the-shelf libraries. Evaluations show that we achieve significant speedups over the state-of-the-art frameworks and libraries on a real-world application, e.g., 9.8× over TensorFlow and 3.6× over Intel MKL on the detection in autonomous driving. Lei Liu 0030, Guangli Li, Jiansong Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Feng 0002 |
PPoPP | 6 |
| 2018 | Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
Guangli Li, Lei Liu 0030, Xueying Wang 0003, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 3 |
| 2018 | Background Subtraction on Depth Videos with Convolutional Neural NetworksabstractBackground subtraction is a significant component of computer vision systems. It is widely used in video surveillance, object tracking, anomaly detection, etc. A new data source for background subtraction appeared as the emergence of low-cost depth sensors like Microsof t Kinect, Asus Xtion PRO, etc. In this paper, we propose a background subtraction approach on depth videos, which is based on convolutional neural networks (CNNs), called BGSNet-D (BackGround Subtraction neural Networks for Depth videos). The method can be used in color unavailable scenarios like poor lighting situations, and can also be applied to combine with existing RGB background subtraction methods. A preprocessing strategy is designed to reduce the influences incurred by noise from depth sensors. The experimental results on the SBM-RGBD dataset show that the proposed method outperforms existing methods on depth data, and even reaches the performance of the methods that use RGB-D data. Xueying Wang 0003, Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
IJCNN | 1 |