Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qiang Wu 0012

dblp:87/2533-12 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0009-8981-2876ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 46% Generative modeling · 27% Language models and text generation · 13%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 46% Hardware accelerators and domain-specific architectures · 35% Electronic design automation · 20%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
processing-in-memory
1.722025
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025
Machine learning › Generative modeling
image tokenization
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026
Natural language and speech › Language models and text generation
large language model
1.012026
OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization
1.012026
OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026
Machine learning › Efficient and distributed learning
model quantization
1.012026
OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026
Machine learning › Generative modeling
variational autoencoder
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026
Machine learning › Representation and self-supervised learning
vector quantization
1.012026
VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026
Electronic design automation › power integrity
IR-drop
0.912025
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.912025
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025
Memory systems › processing-in-memory
near-memory processing
0.912025
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.612022
PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization · ECCV (12) 2022
Machine learning › Efficient and distributed learning › model compression
quantization
0.612022
PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization · ECCV (12) 2022
Machine learning › Efficient and distributed learning
on-device inference
0.312026
OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator
0.312025
H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025
Electronic design automation
physical design
0.312025
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025

Methods — techniques the papers use, named apart from their topics

vector quantization · 1.0variational autoencoder · 1.0quantization-aware fine-tuning · 1.0distribution regularization · 1.0bit-width path search · 1.0asynchronous gradient accumulation · 1.0task mapping · 0.9software-hardware co-design · 0.9hybrid bonding · 0.9dataflow co-exploration · 0.9twin uniform quantization · 0.6post-training quantization · 0.6
YearPublicationVenuePosition
2026 OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs
abstract
Large Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real‑world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions.
Shaoyuan Chen, Zhixuan Chen, Zhihang Yuan, Qiang Wu 0012
AAAI5
2026 VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling
abstract
Vector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and after quantization, and poor coherence between the continuous and discrete domains. These issues lead to unstable codeword learning and underutilized codebooks, ultimately degrading the performance of both reconstruction and downstream generation tasks. To this end, we propose VAEVQ, which comprises three key components: (1) Variational Latent Quantization (VLQ), replacing the AE with a VAE for quantization to leverage its structured and smooth latent space, thereby facilitating more effective codeword activation; (2) Representation Coherence Strategy (RCS), adaptively modulating the alignment strength between pre- and post-quantization features to enhance consistency and prevent overfitting to noise; and (3) Distribution Consistency Regularization (DCR), aligning the entire codebook distribution with the continuous latent distribution to improve utilization. Extensive experiments on two benchmark datasets demonstrate that VAEVQ outperforms state-of-the-art methods.
Sicheng Yang 0001, Xing Hu 0010, Qiang Wu 0012
AAAI3
2025 H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference
abstract
Low-batch large language model (LLM) inference has been extensively applied to edge-side generative tasks, such as personal chat helper, virtual assistant, reception bot, private edge server, etc.To efficiently handle both prefill and decoding stages in LLM inference, near-memory processing (NMP) enabled heterogeneous computation paradigm has been proposed.However, existing NMP designs typically embed processing engines into DRAM dies, resulting in limited computation capacity, which in turn restricts their ability to accelerate edge-side low-batch LLM inference.To tackle this problem, we propose H 2 -LLM, a Hybrid-bondingbased Heterogeneous accelerator for edge-side low-batch LLM inference.To balance the trade-off between computation capacity and bandwidth intrinsic to hybrid-bonding technology, we propose * Co-corresponding authors.
Cong Li 0008, Yihan Yin, Xintong Wu, Jingchen Zhu, Zhutianya Gao, Dimin Niu, Qiang Wu 0012, Xin Si, Yuan Xie 0001, Chen Zhang 0001, Guangyu Sun 0003
ISCA7
2025 AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
abstract
SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup.
Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003
ISCA11
2025 RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their exponentially increasing parameters pose significant challenges for deployment on resource-constrained devices. Vector Quantization (VQ) shows great promise for low-bit quantization (e.g., 2 to 4 bits), but existing work faces two key challenges: unconstrained direction error and suboptimal bit allocation. In this paper, we propose RSAVQ, a novel VQ framework to enhance extremely low-bit quantization for LLMs. RSAVQ introduces two geometry-driven innovations that effectively mitigate above limitations: (1) Error Direction Sensitivity Guidance (EDSG), which leverages the Fisher information matrix (FIM)-induced Riemannian metric to project quantization errors onto low-sensitivity directions in the parameter space. Specifically, this projection is performed along the negative natural gradient direction, which effectively suppresses error expansion. (2) Weight Channel Sensitivity Guidance (WCSG) , which constructs a channel-wise sensitivity metric via FIM curvature analysis to dynamically guide bit resource allocation. The approach facilitates a globally optimal quantization solution within prescribed bit constraints. Experiments demonstrate that RSAVQ outperforms existing methods for LLMs. For example, in 2-bit quantization of LLaMA-3 8B, RSAVQ leads baselines like VPTQ and QuIP\# by 0.4 in perplexity (PPL) and 1.5 in zero-shot accuracy. This work offers a practical solution for constrained environments and a theoretical bridge between information geometry and the quantization of neural networks, advancing efficient deep learning.
Zukang Xu, Xing Hu 0010, Qiang Wu 0012
NeurIPS3
2022 PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization
Zhihang Yuan, Chenhao Xue, Qiang Wu 0012, Guangyu Sun 0003
ECCV (12)4