Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shuaiting Li

dblp:380/8076 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0002-7726-4883ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 50% Representation and self-supervised learning · 33% Generative modeling · 17%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 84% Memory systems · 16%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
3.542025
SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting · ICCV 2025
ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba · ICCV 2025
MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization · ASPLOS (1) 2025
Machine learning › Representation and self-supervised learning
vector quantization
3.542025
SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting · ICCV 2025
ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba · ICCV 2025
MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization · ASPLOS (1) 2025
Machine learning › Generative modeling
diffusion model
0.912025
VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers · AAAI 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers · AAAI 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.912025
ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba · ICCV 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers · AAAI 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.912025
MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization · ASPLOS (1) 2025
Memory systems › memory access optimization
memory access reduction
0.312025
SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting · ICCV 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.312025
SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting · ICCV 2025
Hardware accelerators and domain-specific architectures
systolic array
0.312025
MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization · ASPLOS (1) 2025

Methods — techniques the papers use, named apart from their topics

sign-splitting · 1.7progressive freezing · 1.7n:m pruning · 1.7masked k-means · 1.7codebook clustering · 1.7zero-data calibration · 0.9post-training quantization · 0.9incremental vector quantization · 0.9convex combination optimization · 0.9block-wise calibration · 0.9
YearPublicationVenuePosition
2025 VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers
abstract
The Diffusion Transformers Models (DiTs) have transitioned the network architecture from traditional UNets to transformers, demonstrating exceptional capabilities in image generation. Although DiTs have been widely applied to high-definition video generation tasks, their large parameter size hinders inference on edge devices. Vector quantization (VQ) can decompose model weight into a codebook and assignments, allowing extreme weight quantization and significantly reducing memory usage. In this paper, we propose VQ4DiT, a fast post-training vector quantization method for DiTs. We found that traditional VQ methods calibrate only the codebook without calibrating the assignments. This leads to weight sub-vectors being incorrectly assigned to the same assignment, providing inconsistent gradients to the codebook and resulting in a suboptimal result. To address this challenge, VQ4DiT calculates the candidate assignment set for each weight sub-vector based on Euclidean distance and reconstructs the sub-vector based on the weighted average. Then, using the zero-data and block-wise calibration method, the optimal assignment from the set is efficiently selected while calibrating the codebook. VQ4DiT quantizes a DiT XL/2 model on a single NVIDIA A100 GPU within 20 minutes to 5 hours depending on the different quantization settings. Experiments show that VQ4DiT establishes a new state-of-the-art in model size and performance trade-offs, quantizing weights to 2-bit precision while retaining acceptable image generation quality.
Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang
AAAI2
2025 MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
abstract
Vector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional VQ techniques lead to significant accuracy loss because the important weights are not well preserved. To tackle this problem, a novel approach called MVQ is proposed, which aims at better approximating important weights with a limited number of codewords. At the algorithm level, our approach removes the less important weights through N:M pruning and then minimizes the vector clustering error between the remaining weights and codewords by the masked k-means algorithm. Only distances between the unpruned weights and the codewords are computed, which are then used to update the codewords. At the architecture level, our accelerator implements vector quantization on an EWS (Enhanced weight stationary) CNN accelerator and proposes a sparse systolic array design to maximize the benefits brought by masked vector quantization.
Shuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 0010, Zewen Ye, Zongsheng Wang, Haibin Shen, Kejie Huang
ASPLOS (1)1
2025 ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba
abstract
Visual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vector quantization (VQ) decomposes network weights into codebooks and assignments, significantly reducing memory usage and computational latency, thereby enabling the deployment of ViMs on edge devices. Although existing VQ methods have achieved extremely low-bit quantization (e.g., 3-bit, 2-bit, and 1-bit) in convolutional neural networks and Transformer-based networks, directly applying these methods to ViMs results in unsatisfactory accuracy. We identify several key challenges: 1) The weights of Mamba-based blocks in ViMs contain numerous outliers, significantly amplifying quantization errors. 2) When applied to ViMs, the latest VQ methods suffer from excessive memory consumption, lengthy calibration procedures, and suboptimal performance in the search for optimal codewords. In this paper, we propose ViM-VQ, an efficient post-training vector quantization method tailored for ViMs. ViM-VQ consists of two innovative components: 1) a fast convex combination optimization algorithm that efficiently updates both the convex combinations and the convex hulls to search for optimal codewords, and 2) an incremental vector quantization strategy that incrementally confirms optimal codewords to mitigate truncation errors. Experimental results demonstrate that ViM-VQ achieves state-of-the-art performance in low-bit quantization across various visual tasks.
Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang
ICCV2
2025 SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting
abstract
Vector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models, particularly in extreme compression scenarios. However, its efficacy during fine-tuning is limited by the constraint of the compression format, where weight vectors assigned to the same codeword are restricted to updates in the same direction. Consequently, many quantized weights are compelled to move in directions contrary to their local gradient information. To mitigate this issue, we introduce a novel VQ paradigm, Sign-Splitting VQ (SSVQ), which decouples the sign bit of weights from the codebook. Our approach involves extracting the sign bits of uncompressed weights and performing clustering and compression on all-positive weights. We then introduce latent variables for the sign bit and jointly optimize both the signs and the codebook. Additionally, we implement a progressive freezing strategy for the learnable sign to ensure training stability. Extensive experiments on various modern models and tasks demonstrate that SSVQ achieves a significantly superior compression-accuracy trade-off compared to conventional VQ. Furthermore, we validate our algorithm on a hardware accelerator, showing that SSVQ achieves a 3$\times$ speedup over the 8-bit compressed model by reducing memory access. Our code is available at https://github.com/list0830/SSVQ.
Shuaiting Li, Juncan Deng, Chengxuan Wang, Kedong Xu, Rongtao Deng, Haibin Shen, Kejie Huang
ICCV1
2025 A 1FeFET-1T-1C based Compute-in-Memory Macro with Capacitor Reused Pipeline SAR ADC
abstract
Computing-in-memory (CIM) significantly reduces latency and power consumption by combining computation and memory, typically utilizing non-volatile memories (NVM). However, device manufacturing non-uniformity on NVMs can cause output deviations. Additionally, the necessity for bit-shifting circuits and Analog-to-Digital Converters (ADC) increases the area and power overhead. To tackle these challenges, we propose a high-density 1FeFET-1T-1C based CIM macro, integrated with a pipeline Successive-Approximation-Register (SAR) ADC. The design introduces a capacitor structure that counters the non-uniformity issues inherent in FeFET devices. Also, the capacitor array is reused as charge-redistribution and ADCs, substantially minimizing the area and power overhead. Moreover, the pipeline architecture accelerates the conversion process, achieving high speed and high precision. The design is implemented using SMIC 55nm PDK. The energy efficiency (EF) and area efficiency (AF) of the proposed macro are 80.9 TOPS/W and 1.161 TOPS/mm2, respectively. The inference accuracy reaches 91.2% on the CIFAR-10 dataset.
Minghan Jiang, Rui Xiao 0003, Shuaiting Li, Yishu Zhang, Haibin Shen, Kejie Huang
ISCAS4
2025 Cross-Modal Adaptation for Object Detection in Infrared Remote Sensing Imagery
abstract
Modern Thermal InfraRed (TIR) technology has been proven highly significant in Remote Sensing Imagery (RSI). Currently, multimodal RSI object detection based on RGB-TIR image pairs has attracted widespread research. However, capturing features in the TIR domain poses a challenge, as existing object detectors heavily focus on chromatic information in the RGB domain. Furthermore, the quality of RGB images can be influenced by complex environmental conditions, limiting the practicality of multimodal detection. In this paper, we introduce Cross-Modal-YOLO (CM-YOLO), a lightweight yet effective object detector specifically designed for TIR remote sensing images. CM-YOLO employs cross-modal adaptation to enhance the awareness of TIR-RGB modality translation. Specifically, we leverage a Prior Modality Translator (PMT) to learn the InfraRed-Visible (IV) features, which are incorporated into the detection backbone using our IV-Gate modules. Experimental results on the VEDAI dataset demonstrate that CM-YOLO significantly outperforms conventional methods. Moreover, CM-YOLO exhibits a strong generalization ability for TIR-based object detection in urban scenes on the FLIR dataset.
Zeyu Wang 0010, Shuaiting Li, Kejie Huang
IEEE Geosci. Remote. Sens. Lett.2