VLDB 2026 Research / reviewers in the wild / expert
Guangli Li
dblp:05/7052
· DBLP profile ↗
46ranked-venue papers
13as first author
32since 2021 · last 2026
0000-0002-9738-261XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Low-Precision Approximation of Tensor Operators on GPUs: Enabling Greater Trade-Offs between Performance and AccuracyabstractRecent GPUs integrate specialized hardware for low-precision arithmetic (e.g., FP16, INT8), offering substantial speedups for tensor operations. However, existing methods typically rely on coarse, operator-level trial-and-error tuning, which restricts the performance–accuracy trade-off space and limits achievable gains.We present Platensor, a progressive low-precision approximation framework that expands this trade-off space through ne-grained, tile-level strategies. The key idea is to exploit the tiled computation patterns of GPUs to enable flexible precision control and richer optimization opportunities. Platensor performs a two-phase exploration: a fast rule-based pass that selects promising tile-level configurations, followed by an evolutionary search that refines them. It then automatically generates optimized kernels that combine tiles of different precisions.Experiments on GEMM operators and representative applications—including kNN, LLMs, and HPL-MxP—show that Platensor significantly broadens the attainable performance– accuracy trade-offs and more fully leverages low-precision arithmetic on modern GPUs compared to operator-level tuning. Fan Luo 0003, Guangli Li, Zhaoyang Hao, Xueying Wang 0003, Xiaobing Feng 0002, Huimin Cui, Jingling Xue |
CGO | 2 |
| 2026 | PriTran: Privacy-Preserving Inference for Transformer-Based Language Models under Fully Homomorphic EncryptionabstractTransformer-based language models power many cloud services, but inference on sensitive data raises confidentiality concerns. Fully Homomorphic Encryption (FHE) enables computation on encrypted inputs while preserving privacy, but at high computational cost, making Transformers difficult to deploy. This paper presents PriTran, an efficient CKKS-based library for privacy-preserving Transformer inference on CPUs. Complementing the only prior work, RoLe, which supports only Berttiny(2 encoders), PriTran introduces two novel algorithms with optimized data layouts that accelerate ciphertext–plaintext (CP) and ciphertext–ciphertext (CC) matrix multiplications (MMs) across all Bert models by reducing costly rotations and multiplications. On the MNLI dataset, RoLe fails on inputs longer than 36 tokens within a 5-hour per-token budget, while PriTran achieves average speedups of 29.3% and 22.2% for CP- and CC-MMs, respectively, and 24.1% end-to-end. We further evaluate PriTran on scaled Berttinyvariants with additional encoders and on Bertmini(4 encoders), demonstrating correctness and scalability beyond RoLe’s limits. Within current FHE limits, these gains and RoLe’s failure on longer inputs underscore PriTran’s promise as a practical approach for FHE-based Transformer inference. Yuechen Mu, Guangli Li, Shiping Chen 0001, Jingling Xue |
CGO | 2 |
| 2026 | DyPARS: Dynamic-Shape DNN Optimization via Pareto-Aware MCTS for Graph VariantsabstractDynamic-shape DNNs are widely used in applications such as variable-resolution image processing and language modeling with variable-length sequences. Existing DL (Deep-Learning) compilers apply rule-based rewriting to either transform a subgraph into a fixed variant at compile time (leading to suboptimal performance) or generate multiple variants at runtime, incurring significant overhead. The challenge is discovering and applying shape-dependent subgraph variants that maintain high efficiency across diverse inputs with minimal runtime cost.We propose DyPARS, a dynamic-shape DL compiler approach that discovers high-performance subgraph variants at compile time and applies the best ones at runtime. Leveraging Pareto-aware MCTS, DyPARS identifies shape-aware variants, incorporating shape-dependent kernel adaptations. These variants are integrated into a prediction-enhanced computational graph, enabling efficient variant selection based on input shapes with minimal overhead. DyPARS achieves average speedups of 1.31× and 1.80× over TorchInductor (JIT) and BladeDISC (non-JIT), respectively, across five DNN models, demonstrating robust efficiency across diverse inputs. Guangli Li, Qiuchu Yu, Xueying Wang 0003, Jingling Xue |
CGO | 2 |
| 2026 | DACOS: Dependency-Aware Cross-Kernel Overlapping for Optimizing Short-Sequence Workloads in LLM Applications
Zhaoyang Hao, Guangli Li, Fan Luo 0003, Xueying Wang 0003, Huimin Cui, Jingling Xue |
Euro-Par (2) | 2 |
| 2026 | Topology-aware cross-modal learning with dynamic gradient modulation for survival prediction
Yupeng Yuan, Guangli Li, Nan Jiang 0013, Jingqin Lv, Chengxin Ye, Donghong Ji, Yafeng Ren, Gongning Luo, Hongbin Zhang 0004 |
Expert Syst. Appl. | 2 |
| 2026 | PL-DCP: A pairwise learning framework with domain and class prototypes for EEG emotion recognition under unseen target conditions
Guangli Li, Canbiao Wu, Zhehao Zhou, Tuo Sun, Ping Tan 0003, Li Zhang 0041 |
Neurocomputing | 1 |
| 2026 | Expert-guided bipolar feature disentanglement for multimodal survival prediction
Guangli Li, Fanou Yang, Renzhong Wu, Jingqin Lv, Gongning Luo, Shiying Zeng, Hongbin Zhang 0004 |
Pattern Recognit. | 1 |
| 2026 | MoonPoly: Bridging Code Generation and Adaptive Execution via Micro-Kernel Polymerization for Optimizing Dynamic-Shape Tensor OperatorsabstractThe prevalence of dynamic tensor shapes, driven by applications like language model serving with varying sequence lengths, is a defining characteristic of modern deep neural networks. This dynamism poses a fundamental challenge: reconciling the need for intensive, offline code generation to achieve peak performance with the demand for low-latency, adaptive execution to handle unpredictable runtime tensor shapes. Consequently, mainstream strategies are ineffective. Vendor-provided libraries, while highly optimized for a subset of common shapes, suffer performance degradation on unconventional ones. Static tensor compilers are hamstrung by prohibitive just-in-time compilation overheads for each new shape. While recent dynamic-shape compilers offer an alternative, they rely on predefined shape ranges, making them brittle when inputs fall outside these bounds. To resolve this tension, we present MoonPoly , a dynamic-shape tensor compiler that introduces micro-kernel polymerization . Our approach decouples these conflicting requirements through a two-stage process. In the offline stage, it performs intensive auto-tuning to generate a set of micro-kernels and corresponding performance models. The online stage then performs adaptive execution, rapidly assembling a near-optimal tensor operator on-the-fly, guided by a lightweight cost model. Evaluated on an NVIDIA A100 GPU, MoonPoly achieves an average operator-level speedup of 1.27× over the cuBLAS library across a diverse set of operators and data types, which in turn yields end-to-end inference acceleration for a variety of models, including BERT, the Vision Transformer, and large language models. Yangyu Zhang, Guangli Li, Feng Yu 0019, Fan Luo 0003, Qianqi Sun, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language ModelsabstractActivation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to pursue activation sparsity and acceleration, but few can simultaneously obtain high activation sparsity and comparable model performance. This paper introduces a simple and effective method named “ProSparse” to sparsify LLMs while achieving both targets. Specifically, after introducing ReLU activation, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing for multiple stages. This can enhance activation sparsity and mitigate performance degradation by avoiding radical shifts in activation distributions. With ProSparse, we obtain high sparsity of 89.32% for LLaMA2-7B, 88.80% for LLaMA2-13B, and 87.89% for end-size MiniCPM-1B, respectively, with comparable performance to their original Swish-activated versions. These present the most sparsely activated models among open-source LLaMA versions and competitive end-size models. Inference acceleration experiments further demonstrate the significant practical acceleration potential of LLMs with higher activation sparsity, obtaining up to 4.52x inference speedup. Xu Han 0007, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Zhiyuan Liu 0001, Guangli Li, Maosong Sun 0001 |
COLING | 9 |
| 2025 | TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs
Guangli Li, Feng Yu 0019, Xueying Wang 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
Euro-Par (2) | 2 |
| 2025 | QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code TranslationabstractThe rise of GPU-based high-performance computing (HPC) has driven the widespread adoption of parallel programming models such as CUDA. Yet, the inherent complexity of parallel programming creates a demand for the automated sequential-to-parallel approaches.
However, data scarcity poses a significant challenge for machine learning-based sequential-to-parallel code translation. Although recent back-translation methods show promise, they still fail to ensure functional equivalence in the translated code. In this paper, we propose \textbf{QiMeng-MuPa}, a novel \textbf{Mu}tual-Supervised Learning framework for Sequential-to-\textbf{Pa}rallel code translation, to address the functional equivalence issue. QiMeng-MuPa consists of two models, a Translator and a Tester. Through an iterative loop consisting of Co-verify and Co-evolve steps, the Translator and the Tester mutually generate data for each other and improve collectively. The Tester generates unit tests to verify and filter functionally equivalent translated code, thereby evolving the Translator, while the Translator generates translated code as augmented input to evolve the Tester. Experimental results demonstrate that QiMeng-MuPa significantly enhances the performance of the base models: when applied to Qwen2.5-Coder, it not only improves Pass@1 by up to 28.91\% and boosts Tester performance by 68.90\%, but also outperforms the previous state-of-the-art method CodeRosetta by 1.56 and 6.92 in BLEU and CodeBLEU scores, while achieving performance comparable to DeepSeek-R1 and GPT-4.1. Our code is available at \url{https://github.com/kcxain/mupa}. Changxin Ke, Rui Zhang 0040, Guangli Li, Yuanbo Wen 0001, Shuoming Zhang, Ruiyuan Xu, Jiaming Guo, Chenxi Wang 0005, Ling Li 0001, Qi Guo 0001, Yunji Chen |
NeurIPS | 5 |
| 2025 | Image sentiment analysis based on distillation and sentiment region localization networkabstractAbstract Accurately identifying the emotions in images is crucial for sentiment content analysis. To detect local sentiment regions and acquire discriminative sentiment features, we propose a novel model named Distillation-guided and Contrastive-enhanced Sentiment Region Localization Network (DC-SRLN) to effectively complete image sentiment analysis. Two smart but heterogeneous SRLNs are designed first to pursue local sentiment regions. Then an innovative contrastive learning mode is implemented between global and local features to further enhance the discriminative ability of the sentiment features. Third, the enhanced global and local sentiment features are seamlessly integrated to guide each SRLN accurately capture local sentiment regions. Finally, an adaptive feature fusion module is created to fuse the heterogeneous features from the two SRLNs and generate a new multi-view multi-granularity sentiment semantics with more discriminative ability for image sentiment analysis. Extensive experimental results on three prevailing datasets, namely Twitter I, FI, and ArtPhoto, exhibit that DC-SRLN achieves satisfactory accuracies of 93.2%, 80.6%, and 78.7%, respectively, outperforming recent state-of-the-art baselines. Moreover, DC-SRLN needs less training time, demonstrating its high practicality. The code of DC-SRLN is freely available at https://github.com/Riley6868/DC-SRLN. Hongbin Zhang 0004, Ya Feng, Jingyi Hou, Guangli Li |
Comput. J. | 6 |
| 2025 | Report is a mixture of topics: Topic-guided radiology report generation
Guangli Li, Chentao Huang, Xinjiong Zhou, Donghong Ji, Hongbin Zhang 0004 |
Medical Image Anal. | 1 |
| 2025 | OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUsabstractConvolutional Neural Networks (CNNs) are fundamental to advancing computer vision technologies. As CNNs become more complex and larger, optimizing model inference remains a critical challenge in both industry and academia. On modern GPU platforms, CNN operators are typically memory-bound, leading to significant performance degradation due to memory wall effects. While recent advancements have utilized operator fusion–merging multiple operators into one–to enhance inference performance, the fusion of multiple region-based operators like convolution is seldom addressed. This article introduces AFusion , a novel operator fusion technique aimed at improving inference performance, and OptiFX, an automatic optimization framework based on this approach. OptiFX employs a cost-based backtracking search to identify optimal sub-graphs for fusion and utilizes template-based code generation to create efficient kernels for these fused sub-graphs. We evaluate OptiFX across seven prominent CNN architectures–GoogLeNet, ResNet, DenseNet, MobileNet, SqueezeNet, NasNet, and UNet–on Nvidia A6000 Ada, RTX 4090, and Jetson AGX Orin platforms. Our results demonstrate that OptiFX significantly outperforms existing methods, achieving average speedups of \(2.91\times\) , \(3.30\times\) , and \(2.09\times\) in accelerating inference performance on these platforms, respectively. Xueying Wang 0003, Shigang Li 0002, Fan Luo 0003, Zhaoyang Hao, Tong Wu 0024, Ruiyuan Xu, Huimin Cui, Xiaobing Feng 0002, Guangli Li, Jingling Xue |
ACM Trans. Archit. Code Optim. | 10 |
| 2024 | Optimizing Dynamic-Shape Neural Networks on Accelerators via On-the-Fly Micro-Kernel PolymerizationabstractIn recent times, dynamic-shape neural networks have gained widespread usage in intelligent applications to address complex tasks, introducing challenges in optimizing tensor programs due to their dynamic nature. As the operators' shapes are determined at runtime in dynamic scenarios, the compilation process becomes expensive, limiting the practicality of existing static-shape tensor compilers. To address the need for effective and efficient optimization of dynamic-shape neural networks, this paper introduces MikPoly, a novel dynamic-shape tensor compiler based on micro-kernel polymerization. MikPoly employs a two-stage optimization approach, dynamically combining multiple statically generated micro-kernels using a lightweight cost model based on the shape of a tensor operator known at runtime. We evaluate the effectiveness of MikPoly by employing popular dynamic-shape operators and neural networks on two representative accelerators, namely GPU Tensor Cores and Ascend NPUs. Our experimental results demonstrate that MikPoly effectively optimizes dynamic-shape workloads, yielding an average performance improvement of 1.49× over state-of-the-art vendor libraries. Feng Yu 0019, Guangli Li, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
ASPLOS (2) | 2 |
| 2024 | Time Series Classification Based on Forward Echo State Convolution NetworkabstractAbstract The Echo state network (ESN) is an efficient recurrent neural network that has achieved good results in time series prediction tasks. Still, its application in time series classification tasks has yet to develop fully. In this study, we work on the time series classification problem based on echo state networks. We propose a new framework called forward echo state convolutional network (FESCN). It consists of two parts, the encoder and the decoder, where the encoder part is composed of a forward topology echo state network (FT-ESN), and the decoder part mainly consists of a convolutional layer and a max-pooling layer. We apply the proposed network framework to the univariate time series dataset UCR and compare it with six traditional methods and four neural network models. The experimental findings demonstrate that FESCN outperforms other methods in terms of overall classification accuracy. Additionally, we investigated the impact of reservoir size on network performance and observed that the optimal classification results were obtained when the reservoir size was set to 32. Finally, we investigated the performance of the network under noise interference, and the results show that FESCN has a more stable network performance compared to EMN (echo memory network). Jianfeng Tang, Guangli Li, Shukai Duan 0001, Lidan Wang 0001 |
Neural Process. Lett. | 3 |
| 2024 | Fast Convolution Meets Low Precision: Exploring Efficient Quantized Winograd Convolution on Modern CPUsabstractLow-precision computation has emerged as one of the most effective techniques for accelerating convolutional neural networks and has garnered widespread support on modern hardware. Despite its effectiveness in accelerating convolutional neural networks, low-precision computation has not been commonly applied to fast convolutions, such as the Winograd algorithm, due to numerical issues. In this article, we propose an effective quantized Winograd convolution, named LoWino, which employs an in-side quantization method in the Winograd domain to reduce the precision loss caused by transformations. Meanwhile, we present an efficient implementation that integrates well-designed optimization techniques, allowing us to fully exploit the capabilities of low-precision computation on modern CPUs. We evaluate LoWino on two Intel Xeon Scalable Processor platforms with representative convolutional layers and neural network models. The experimental results demonstrate that our approach can achieve an average of 1.84× and 1.91× operator speedups over state-of-the-art implementations in the vendor library while preserving accuracy loss at a reasonable level. Xueying Wang 0003, Guangli Li, Zhen Jia 0001, Xiaobing Feng 0002, Yida Wang 0003 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | ApproxDup: Developing an Approximate Instruction Duplication Mechanism for Efficient SDC Detection in GPGPUsabstractNowadays, selective instruction duplication (SelDup) is the typical approach to detect silent data corruption (SDC) in GPGPU. However, owing to the up-to-billions fault sites of parallel GPGPU kernel functions, it usually introduces tremendous overhead to perform fault injections (FIs) for obtaining the duplication-candidate instruction set (although can be conducted in parallel). Moreover, current SelDup typically considers all SDCs severe and tends to duplicate more instructions. The nontrivial duplication overhead seriously restricts the deployment of current SelDup on resource-constrained systems (e.g., embedded GPGPUs). To address the above challenges, this article proposes an approximate instruction duplication (ApproxDup) mechanism for efficient SDC detection in GPGPUs. First, to replace the expensive FI-based duplication-candidate instructions identified method, we drive out a machine learning (ML)-based model (SDC-predictor) for instructionwise SDC proneness and severity estimation. Our key insight is that instruction type/functionality and instruction dependency set can efficaciously characterize the instructionwise SDC proneness in GPGPUs. In contrast, the instruction’s original data magnitude, fault propagation range, and error detected features can distinguish its SDC severity. Second, incorporating the concept of approximate computing, we propose ApproxDup that preferentially duplicates severe-SDC-prone instructions while relaxing the detection of minor/detectable SDCs for traditional SelDup overhead reduction. Experimental results exhibit that ApproxDup can cover 92.51% of severe SDCs while merely increasing 38% of dynamic instructions, which achieves a better tradeoff between reliability and performance compared with the state-of-the-art SelDup. Furthermore, we discuss the effectiveness of the proposed method on different ML models/applications/GPGPU architectures. Xiaohui Wei 0002, Nan Jiang 0013, Hengshan Yue, Jianpeng Zhao 0001, Guangli Li, Meikang Qiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Embedded mutual learning: A novel online distillation method integrating diverse knowledge sources
Chuanxiu Li, Guangli Li, Hongbin Zhang 0004, Donghong Ji |
Appl. Intell. | 2 |
| 2023 | Correction to: Embedded mutual learning: a novel online distillation method integrating diverse knowledge sources
Chuanxiu Li, Guangli Li, Hongbin Zhang 0004, Donghong Ji |
Appl. Intell. | 2 |
| 2023 | FASS-pruner: customizing a fine-grained CNN accelerator-aware pruning framework via intra-filter splitting and inter-filter shuffling
Xiaohui Wei 0002, Xinyang Zheng, Guangli Li, Hengshan Yue |
CCF Trans. High Perform. Comput. | 4 |
| 2023 | CoAxNN: Optimizing on-device deep learning with conditional approximate neural networks
Guangli Li, Xiu Ma, Qiuchu Yu, Lei Liu 0040, Huaxiao Liu, Xueying Wang 0003 |
J. Syst. Archit. | 1 |
| 2023 | Facilitating hardware-aware neural architecture search with learning-based predictive models
Xueying Wang 0003, Guangli Li, Xiu Ma, Xiaobing Feng 0002 |
J. Syst. Archit. | 2 |
| 2022 | MF-OMKT: Model fusion based on online mutual knowledge transfer for breast cancer histopathological image classification
Guangli Li, Chuanxiu Li, Guangting Wu, Guangxin Xu, Hongbin Zhang 0004 |
Artif. Intell. Medicine | 1 |
| 2022 | Accelerating deep neural network filter pruning with mask-aware convolutional computations on modern CPUsabstractFilter pruning, a representative model compression technique , has been widely used to compress and accelerate sophisticated deep neural networks on resource-constrained platforms. Nevertheless, most studies focus on reducing the cost of model inference, whereas the heavy burden of the pruning optimization process is neglected. In this paper, we propose MaskACC, a mask-aware convolutional computation method, which accelerates the prevailing mask-based filter pruning process on modern CPU platforms. MaskACC dynamically reorganizes the tensors used in convolutions with the mask information to avoid unnecessary computations, thereby improving the computational efficiency of the pruning process. Evaluation with state-of-the-art neural network models on CPU cloud platforms demonstrates the effectiveness of our method, which achieves up to 1.61 × speedup under commonly-used pruning rates, compared to conventional computations. Xiu Ma, Guangli Li, Lei Liu 0040, Huaxiao Liu, Xueying Wang 0003 |
Neurocomputing | 2 |
| 2022 | Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning
Guangli Li, Xiu Ma, Xueying Wang 0003, Hengshan Yue, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002, Jingling Xue |
J. Syst. Archit. | 1 |
| 2022 | An Application-oblivious Memory Scheduling System for DNN AcceleratorsabstractDeep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes. Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2021 | Unleashing the Low-Precision Computation Potential of Tensor Cores on GPUsabstractTensor-specialized hardware for supporting low-precision arithmetic has become an inevitable trend due to the ever-increasing demand on computational capability and energy efficiency in intelligent applications. The main challenge faced when accelerating a tensor program on tensor-specialized hardware is how to achieve the best performance possible in reduced precision by fully utilizing its computational resources while keeping the precision loss in a controlled manner. In this paper, we address this challenge by proposing QUANTENSOR, a new approach for accelerating general-purpose tensor programs by replacing its tensor computations with low-precision quantized tensor computations on NVIDIA Tensor Cores. The key novelty is a new residual-based precision refinement technique for controlling the quantization errors, allowing tradeoffs between performance and precision to be made. Evaluation with GEMM, deep neural networks, and linear algebra applications shows that QUANTENSOR can achieve remarkable performance improvements while reducing the precision loss incurred significantly at acceptable overheads. Guangli Li, Jingling Xue, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Jiansong Li, Xiaobing Feng 0002 |
CGO | 1 |
| 2021 | LoWino: Towards Efficient Low-Precision Winograd Convolutions on Modern CPUsabstractLow-precision computation, which has been widely supported in contemporary hardware, is considered as one of the most effective methods to accelerate convolutional neural networks. However, low-precision computation is not widely used to speed up Winograd, an algorithm for fast convolution computation, due to the numerical error introduced by combining Winograd transformation and quantization. In this paper, we propose a low-precision Winograd convolution approach, LoWino, based on post-training quantization, which employs a linear quantization method in the Winograd domain to reduce the precision loss caused by transformations. Moreover, we present an efficient implementation that integrates well-designed optimization techniques, thereby adequately exploiting the capability of low-precision computation on modern CPUs. We evaluate our approach on Intel Xeon Scalable Processors by leveraging representative convolutional layers in prevailing deep neural networks. Experimental results show that LoWino achieves up to 2.04 × speedup over state-of-the-art implementations in the vendor library while maintaining the accuracy at a reasonable level. Guangli Li, Zhen Jia 0001, Xiaobing Feng 0002, Yida Wang 0003 |
ICPP | 1 |
| 2021 | Pinpointing the Memory Behaviors of DNN TrainingabstractThe training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns. Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002 |
ISPASS | 3 |
| 2021 | G-SEPM: building an accurate and efficient soft error prediction model for GPGPUsabstractAs GPUs become ubiquitous in large-scale general purpose HPC systems (GPGPUs), ensuring the reliable execution of such systems in the presence of soft errors is increasingly essential. To provide insights into how resilient GPU programs are toward soft errors, researchers typically rely on random Fault Injection (FI) to evaluate the tolerance of programs. However, it is expensive to obtain a statistically significant resilience profile and not suitable to identify all the error-critical fault sites of GPU programs. Hengshan Yue, Xiaohui Wei 0002, Guangli Li, Jianpeng Zhao 0001, Nan Jiang 0013, Jingweijia Tan |
SC | 3 |
| 2021 | Multi-modal visual adversarial Bayesian personalized ranking model for recommendation
Guangli Li, Jianwu Zhuo, Chuanxiu Li, Jin Hua, Zhengyu Niu, Donghong Ji, Renzhong Wu, Hongbin Zhang 0004 |
Inf. Sci. | 1 |
| 2020 | Accelerating Deep Learning Inference with Cross-Layer Data Reuse on GPUs
Xueying Wang 0003, Guangli Li, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002 |
Euro-Par | 2 |
| 2020 | Lance: efficient low-precision quantized winograd convolution for neural networks based on graphics processing unitsabstractAccelerating deep convolutional neural networks has become an active topic and sparked an interest in academia and industry. In this paper, we propose an efficient low-precision quan-tized Winograd convolution algorithm, called LANCE, which combines the advantages of fast convolution and quantization techniques. By embedding linear quantization operations into the Winograd-domain, the fast convolution can be performed efficiently under low-precision computation on graphics processing units. We test neural network models with LANCE on representative image classification datasets, including SVHN, CIFAR, and ImageNet. The experimental results show that our 8-bit quantized Winograd convolution improves the performance by up to 2.40× over the full-precision convolution with trivial accuracy loss. Guangli Li, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Xiaobing Feng 0002 |
ICASSP | 1 |
| 2020 | Compiler-Assisted Operator Template Library for DNN Accelerators
Jiansong Li, Wei Cao 0010, Guangli Li, Xueying Wang 0003, Lei Liu 0030, Xiaobing Feng 0002 |
NPC | 4 |
| 2020 | Novel model to integrate word embeddings and syntactic trees for automatic caption generation from images
Hongbin Zhang 0004, Diedie Qiu, Renzhong Wu, Donghong Ji, Guangli Li, Zhenyu Niu, Tao Li 0022 |
Soft Comput. | 5 |
| 2020 | Fusion-Catalyzed Pruning for Optimizing Deep Learning on Intelligent Edge DevicesabstractThe increasing computational cost of deep neural network models limits the applicability of intelligent applications on resource-constrained edge devices. While a number of neural network pruning methods have been proposed to compress the models, prevailing approaches focus only on parametric operators (e.g., convolution), which may miss optimization opportunities. In this article, we present a novel fusion-catalyzed pruning approach, called FuPruner, which simultaneously optimizes the parametric and nonparametric operators for accelerating neural networks. We introduce an aggressive fusion method to equivalently transform a model, which extends the optimization space of pruning and enables nonparametric operators to be pruned in a similar manner as parametric operators, and a dynamic filter pruning method is applied to decrease the computational cost of models while retaining the accuracy requirement. Moreover, FuPruner provides configurable optimization options for controlling fusion and pruning, allowing much more flexible performance-accuracy tradeoffs to be made. Evaluation with state-of-the-art residual neural networks on five representative intelligent edge platforms, Jetson TX2, Jetson Nano, Edge tensor processing unit, neural compute stick, and neural compute stick 2, demonstrates the effectiveness of our approach, which can accelerate the inference of models on CIFAR-10 and ImageNet datasets. Guangli Li, Xiu Ma, Xueying Wang 0003, Lei Liu 0030, Jingling Xue, Xiaobing Feng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Acorns: A Framework for Accelerating Deep Neural Networks with Input SparsityabstractDeep neural networks have been employed in a broad range of applications, including face detection, natural language processing, and autonomous driving. Yet, the neural networks with the capability to tackle real-world problems are intrinsically expensive in computation, hindering the usage of these models. Sparsity in the input data of neural networks provides an optimizing opportunity. However, harnessing the potential performance improvement on modern CPU faces challenges raised by sparse computations of the neural network, such as cache-unfriendly memory accesses and efficient sparse kernel implementation. In this paper, we propose Acorns, a framework to accelerate deep neural networks with input sparsity. In Acorns, sparse input data is organized into our designed sparse data layout, which allows memory-friendly access for kernels in neural networks and opens the door for many performance-critical optimizations. Upon that, Acorns generates efficient sparse kernels for operators in neural networks from kernel templates, which combine directions that express specific optimizing transformations to be performed, and straightforward code that describes the computation. Comprehensive evaluations demonstrate Acorns can outperform state-of-the-art baselines by significant speedups. On the real-world detection task in autonomous driving, Acorns demonstrates 1.8-22.6× performance improvement over baselines. Specifically, the generated programs achieve 1.8-2.4× speedups over Intel MKL-DNN, 3.0-8.8× speedups over TensorFlow, and 11.1-13.2× speedups over Intel MKL-Sparse. Lei Liu 0030, Peng Zhao 0008, Guangli Li, Jiansong Li, Xueying Wang 0003, Xiaobing Feng 0002 |
PACT | 4 |
| 2019 | Accelerating GPU Computing at Runtime with Binary OptimizationabstractNowadays, many applications use GPUs (Graphics Processing Units) to achieve high performance. When we use GPU servers, the idle CPU resource of the servers is often ignored. In this paper, we explore the idea: using the idle CPU resource to speed up GPU programs. We design a dynamic binary optimization framework for accelerating GPU computing at runtime. A template-based binary optimization method is proposed to optimize kernels, which can avoid the high cost of kernel compilation. This method replaces determined variables with constant values and generates an optimized binary kernel. Based on the analysis results of optimization opportunities, we replace the original kernels with optimized kernels during program execution. The experimental results show that it is feasible to accelerate GPU programs via binary optimization. After applying binary optimization to five convolution layers of deep neural networks, the average performance improvement can reach 20%. Guangli Li, Lei Liu 0030, Xiaobing Feng 0002 |
CGO | 1 |
| 2019 | Exploiting the input sparsity to accelerate deep neural networks: posterabstractEfficient inference of deep learning models are challenging and of great value in both academic and industrial community. In this paper, we focus on exploiting the sparsity in input data to improve the performance of deep learning models. We propose an end-to-end optimization pipeline to generate programs for the inference with sparse input. The optimization pipeline contains both domain-specific and general optimization techniques and is capable of generating efficient code without relying on the off-the-shelf libraries. Evaluations show that we achieve significant speedups over the state-of-the-art frameworks and libraries on a real-world application, e.g., 9.8× over TensorFlow and 3.6× over Intel MKL on the detection in autonomous driving. Lei Liu 0030, Guangli Li, Jiansong Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Feng 0002 |
PPoPP | 3 |
| 2018 | Fast CNN Pruning via Redundancy-Aware Training
Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 3 |
| 2018 | Auto-tuning Neural Network Quantization Framework for Collaborative Inference Between the Cloud and Edge
Guangli Li, Lei Liu 0030, Xueying Wang 0003, Peng Zhao 0008, Xiaobing Feng 0002 |
ICANN (1) | 1 |
| 2018 | Background Subtraction on Depth Videos with Convolutional Neural NetworksabstractBackground subtraction is a significant component of computer vision systems. It is widely used in video surveillance, object tracking, anomaly detection, etc. A new data source for background subtraction appeared as the emergence of low-cost depth sensors like Microsof t Kinect, Asus Xtion PRO, etc. In this paper, we propose a background subtraction approach on depth videos, which is based on convolutional neural networks (CNNs), called BGSNet-D (BackGround Subtraction neural Networks for Depth videos). The method can be used in color unavailable scenarios like poor lighting situations, and can also be applied to combine with existing RGB background subtraction methods. A preprocessing strategy is designed to reduce the influences incurred by noise from depth sensors. The experimental results on the SBM-RGBD dataset show that the proposed method outperforms existing methods on depth data, and even reaches the performance of the methods that use RGB-D data. Xueying Wang 0003, Lei Liu 0030, Guangli Li, Peng Zhao 0008, Xiaobing Feng 0002 |
IJCNN | 3 |
| 2017 | An adaptive scheduling algorithm for heterogeneous Hadoop systemsabstractThe MapReduce framework and its open source implementation Hadoop have established themselves as one of the most polular large data sets analyzers. They are widely used by many cloud service providers such as Amazon EC2 Cloud. However, while latency-sensitive applications becoming more and more important, Hadoop system shows its shortcoming in ensuring jobs completed on time. And currently, user has to provide a metric to evaluate the performance of different clients. Motivated by this, we proposed an algorithm CP-Scheduler (CPS) which uses a optimizer to analyze the best schedule in order to minimize the number of delayed jobs. Otherwise, as Hadoop System is not good at heterogeneous computing, our algorithm can also adapt different remote machines. These two features make it having better efficiency than the scheduler in Hadoop. The proposed algorithm is initially evaluated by a simulator which is designed for Hadoop. Experimental results show that the number of missing deadline jobs decrease by 60 percent on average in different sizes of situations. Jiazhen Han, Zhengheng Yuan, Yiheng Han, Jing Liu 0012, Guangli Li |
ICIS | 6 |
| 2017 | Redundancy checking algorithms based on parallel novel extension ruleabstractRedundancy checking (RC) is a key knowledge reduction technology. Extension rule (ER) is a new reasoning method, first presented in 2003 and well received by experts at home and abroad. Novel extension rule (NER) is an improved ER-based reasoning method, presented in 2009. In this paper, we first analyse the characteristics of the extension rule, and then present a simple algorithm for redundancy checking based on extension rule (RCER). In addition, we introduce MIMF, a type of heuristic strategy. Using the aforementioned rule and strategy, we design and implement RCHER algorithm, which relies on MIMF. Next we design and implement an RCNER (redundancy checking based on NER) algorithm based on NER. Parallel computing greatly accelerates the NER algorithm, which has weak dependence among tasks when executed. Considering this, we present PNER (parallel NER) and apply it to redundancy checking and necessity checking. Furthermore, we design and implement the RCPNER (redundancy checking based on PNER) and NCPPNER (necessary clause partition based on PNER) algorithms as well. The experimental results show that MIMF significantly influences the acceleration of algorithm RCER in formulae on a large scale and high redundancy. Comparing PNER with NER and RCPNER with RCNER, the average speedup can reach up to the number of task decompositions when executed. Comparing NCPNER with the RCNER-based algorithm on separating redundant formulae, speedup increases steadily as the scale of the formulae is incrementing. Finally, we describe the challenges that the extension rule will be faced with and suggest possible solutions. Lei Liu 0030, Guangli Li, Shuai Lü 0001 |
J. Exp. Theor. Artif. Intell. | 3 |
| 2017 | Loss evaluation analysis of illegal attack in SCSKP
Peng Zhang 0053, Lei Liu 0040, Rui Zhang 0040, Guangli Li |
Soft Comput. | 4 |