Jiansong Li

dblp:211/8933 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Few-Shot Learning Based on Multimodal Information Processing
abstract
Few-shot learning aims to develop models with strong generalization capabilities using a small number of training samples. However, most learning methods rely solely on the visual features of a few samples to represent entire categories, leading to poor category representativeness. In contrast, humans can utilize multimodal information to learn category features, thereby making them more representative. Hence, this article emulates the human multimodal learning mechanism by integrating visual features with textual information, thereby facilitating the model's acquisition of more representative and robust category features. Specifically, this article introduces a novel multimodal fusion mechanism-the visual-semantic fusion selection mechanism (VSFSM)-which comprises a fusion selection module (FS-Module) and a category enhancement module (CE-Module). These two modules collaboratively enhance the model's classification performance. The FS-Module aligns and fuses semantic information with visual features across both channel and spatial dimensions, performing feature selection and reconstruction. This process not only generates representative category features but also mitigates the impact of noise. The CE-Module guides the model to emphasize category-specific features in the query images, ultimately yielding representative visual-semantic category features while reducing the interference of noise in the query images. Additionally, to better facilitate few-shot learning, this article introduces a novel objective loss function for optimized training. Extensive comparative and ablation experiments conducted on multiple datasets further validate the effectiveness of the proposed method.
Zhenping Lan, Yanguo Sun, Jiansong Li, Xincheng Yang
IEEE Trans. Neural Networks Learn. Syst.5
2022 MTMG: A Framework for Generating Adversarial Examples Targeting Multiple Learning-Based Malware Detection Systems
Lichen Jia, Jiansong Li
PRICAI (1)3
2022 A systematic study on benchmarking AI inference accelerators
Zihan Jiang 0006, Jiansong Li, Fangxin Liu, Wanling Gao, Lei Wang 0004, Chuanxin Lan, Fei Tang 0003, Lei Liu 0030, Tao Li 0022
CCF Trans. High Perform. Comput.2
2022 Optimal randomized quadrature for weighted Sobolev and Besov classes with the Jacobi weight on the ball
Jiansong Li
J. Complex.1
2022 Optimizing deep neural networks on intelligent edge accelerators via flexible-rate filter pruning
Guangli Li, Xiu Ma, Xueying Wang 0003, Hengshan Yue, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002, Jingling Xue
J. Syst. Archit.5
2022 An Application-oblivious Memory Scheduling System for DNN Accelerators
abstract
Deep Neural Networks (DNNs) tend to go deeper and wider, which poses a significant challenge to the training of DNNs, due to the limited memory capacity of DNN accelerators. Existing solutions for memory-efficient DNN training are densely coupled with the application features of DNN workloads, e.g., layer structures or computational graphs of DNNs are necessary for these solutions. This would result in weak versatility for DNNs with sophisticated layer structures or complicated computation graphs. These schemes usually need to be re-implemented or re-adapted due to the new layer structures or the unusual operators in the computational graphs introduced by these DNNs. In this article, we review the memory pressure issues of DNN training from the perspective of runtime systems and model the memory access behaviors of DNN workloads. We identify the iterative, regularity , and extremalization properties of memory access patterns for DNN workloads. Based on these observations, we propose AppObMem, an application-oblivious memory scheduling system. AppObMem automatically traces the memory behaviors of DNN workloads and schedules the memory swapping to reduce the memory pressure of the device accelerators without the perception of high-level information of layer structures or computation graphs. Evaluations on a variety of DNN models show that, AppObMem obtains 40–60% memory savings with acceptable performance loss. AppObMem is also competitive with other open sourced SOTA schemes.
Jiansong Li, Xueying Wang 0003, Xiaobing Chen, Guangli Li, Peng Zhao 0008, Xianzhi Yu, Yongxin Yang, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002
ACM Trans. Archit. Code Optim.1
2021 Unleashing the Low-Precision Computation Potential of Tensor Cores on GPUs
abstract
Tensor-specialized hardware for supporting low-precision arithmetic has become an inevitable trend due to the ever-increasing demand on computational capability and energy efficiency in intelligent applications. The main challenge faced when accelerating a tensor program on tensor-specialized hardware is how to achieve the best performance possible in reduced precision by fully utilizing its computational resources while keeping the precision loss in a controlled manner. In this paper, we address this challenge by proposing QUANTENSOR, a new approach for accelerating general-purpose tensor programs by replacing its tensor computations with low-precision quantized tensor computations on NVIDIA Tensor Cores. The key novelty is a new residual-based precision refinement technique for controlling the quantization errors, allowing tradeoffs between performance and precision to be made. Evaluation with GEMM, deep neural networks, and linear algebra applications shows that QUANTENSOR can achieve remarkable performance improvements while reducing the precision loss incurred significantly at acceptable overheads.
Guangli Li, Jingling Xue, Lei Liu 0030, Xueying Wang 0003, Xiu Ma, Jiansong Li, Xiaobing Feng 0002
CGO7
2021 Pinpointing the Memory Behaviors of DNN Training
abstract
The training of deep neural networks (DNNs) is usually memory-hungry due to the limited device memory capacity of DNN accelerators. Characterizing the memory behaviors of DNN training is critical to optimize the device memory pressures. In this work, we pinpoint the memory behaviors of each device memory block of GPU during training by instrumenting the memory allocators of the runtime system. Our results show that the memory access patterns of device memory blocks are stable and follow an iterative fashion. These observations are useful for the future optimization of memory-efficient training from the perspective of raw memory access patterns.
Jiansong Li, Guangli Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Chen, Xianzhi Yu, Yongxin Yang, Zihan Jiang 0006, Wei Cao 0010, Lei Liu 0030, Xiaobing Feng 0002
ISPASS1
2020 Accelerating Deep Learning Inference with Cross-Layer Data Reuse on GPUs
Xueying Wang 0003, Guangli Li, Jiansong Li, Lei Liu 0030, Xiaobing Feng 0002
Euro-Par4
2020 Compiler-Assisted Operator Template Library for DNN Accelerators
Jiansong Li, Wei Cao 0010, Guangli Li, Xueying Wang 0003, Lei Liu 0030, Xiaobing Feng 0002
NPC1
2019 Acorns: A Framework for Accelerating Deep Neural Networks with Input Sparsity
abstract
Deep neural networks have been employed in a broad range of applications, including face detection, natural language processing, and autonomous driving. Yet, the neural networks with the capability to tackle real-world problems are intrinsically expensive in computation, hindering the usage of these models. Sparsity in the input data of neural networks provides an optimizing opportunity. However, harnessing the potential performance improvement on modern CPU faces challenges raised by sparse computations of the neural network, such as cache-unfriendly memory accesses and efficient sparse kernel implementation. In this paper, we propose Acorns, a framework to accelerate deep neural networks with input sparsity. In Acorns, sparse input data is organized into our designed sparse data layout, which allows memory-friendly access for kernels in neural networks and opens the door for many performance-critical optimizations. Upon that, Acorns generates efficient sparse kernels for operators in neural networks from kernel templates, which combine directions that express specific optimizing transformations to be performed, and straightforward code that describes the computation. Comprehensive evaluations demonstrate Acorns can outperform state-of-the-art baselines by significant speedups. On the real-world detection task in autonomous driving, Acorns demonstrates 1.8-22.6× performance improvement over baselines. Specifically, the generated programs achieve 1.8-2.4× speedups over Intel MKL-DNN, 3.0-8.8× speedups over TensorFlow, and 11.1-13.2× speedups over Intel MKL-Sparse.
Lei Liu 0030, Peng Zhao 0008, Guangli Li, Jiansong Li, Xueying Wang 0003, Xiaobing Feng 0002
PACT5
2019 Exploiting the input sparsity to accelerate deep neural networks: poster
abstract
Efficient inference of deep learning models are challenging and of great value in both academic and industrial community. In this paper, we focus on exploiting the sparsity in input data to improve the performance of deep learning models. We propose an end-to-end optimization pipeline to generate programs for the inference with sparse input. The optimization pipeline contains both domain-specific and general optimization techniques and is capable of generating efficient code without relying on the off-the-shelf libraries. Evaluations show that we achieve significant speedups over the state-of-the-art frameworks and libraries on a real-world application, e.g., 9.8× over TensorFlow and 3.6× over Intel MKL on the detection in autonomous driving.
Lei Liu 0030, Guangli Li, Jiansong Li, Peng Zhao 0008, Xueying Wang 0003, Xiaobing Feng 0002
PPoPP4