VLDB 2026 Research / reviewers in the wild / expert
Qiang Wu 0012
dblp:87/2533-12
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0009-8981-2876ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 46% Generative modeling · 27% Language models and text generation · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 46% Hardware accelerators and domain-specific architectures · 35% Electronic design automation · 20% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
processing-in-memory |
1.7 | 2 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025 |
Machine learning › Generative modeling
image tokenization |
1.0 | 1 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Efficient and distributed learning
model quantization |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Generative modeling
variational autoencoder |
1.0 | 1 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 |
Machine learning › Representation and self-supervised learning
vector quantization |
1.0 | 1 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 |
Electronic design automation › power integrity
IR-drop |
0.9 | 1 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
0.9 | 1 | 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025 |
Memory systems › processing-in-memory
near-memory processing |
0.9 | 1 | 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.6 | 1 | 2022 | PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization · ECCV (12) 2022 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.6 | 1 | 2022 | PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization · ECCV (12) 2022 |
Machine learning › Efficient and distributed learning
on-device inference |
0.3 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator |
0.3 | 1 | 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM Inference · ISCA 2025 |
Electronic design automation
physical design |
0.3 | 1 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 1.0variational autoencoder · 1.0quantization-aware fine-tuning · 1.0distribution regularization · 1.0bit-width path search · 1.0asynchronous gradient accumulation · 1.0task mapping · 0.9software-hardware co-design · 0.9hybrid bonding · 0.9dataflow co-exploration · 0.9twin uniform quantization · 0.6post-training quantization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMsabstractLarge Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real‑world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions. Shaoyuan Chen, Zhixuan Chen, Zhihang Yuan, Qiang Wu 0012 |
AAAI | 5 |
| 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational ModelingabstractVector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and after quantization, and poor coherence between the continuous and discrete domains. These issues lead to unstable codeword learning and underutilized codebooks, ultimately degrading the performance of both reconstruction and downstream generation tasks. To this end, we propose VAEVQ, which comprises three key components: (1) Variational Latent Quantization (VLQ), replacing the AE with a VAE for quantization to leverage its structured and smooth latent space, thereby facilitating more effective codeword activation; (2) Representation Coherence Strategy (RCS), adaptively modulating the alignment strength between pre- and post-quantization features to enhance consistency and prevent overfitting to noise; and (3) Distribution Consistency Regularization (DCR), aligning the entire codebook distribution with the continuous latent distribution to improve utilization. Extensive experiments on two benchmark datasets demonstrate that VAEVQ outperforms state-of-the-art methods. Sicheng Yang 0001, Xing Hu 0010, Qiang Wu 0012 |
AAAI | 3 |
| 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM InferenceabstractLow-batch large language model (LLM) inference has been extensively applied to edge-side generative tasks, such as personal chat helper, virtual assistant, reception bot, private edge server, etc.To efficiently handle both prefill and decoding stages in LLM inference, near-memory processing (NMP) enabled heterogeneous computation paradigm has been proposed.However, existing NMP designs typically embed processing engines into DRAM dies, resulting in limited computation capacity, which in turn restricts their ability to accelerate edge-side low-batch LLM inference.To tackle this problem, we propose H 2 -LLM, a Hybrid-bondingbased Heterogeneous accelerator for edge-side low-batch LLM inference.To balance the trade-off between computation capacity and bandwidth intrinsic to hybrid-bonding technology, we propose * Co-corresponding authors. Cong Li 0008, Yihan Yin, Xintong Wu, Jingchen Zhu, Zhutianya Gao, Dimin Niu, Qiang Wu 0012, Xin Si, Yuan Xie 0001, Chen Zhang 0001, Guangyu Sun 0003 |
ISCA | 7 |
| 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMabstractSRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup. Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003 |
ISCA | 11 |
| 2025 | RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their exponentially increasing parameters pose significant challenges for deployment on resource-constrained devices. Vector Quantization (VQ) shows great promise for low-bit quantization (e.g., 2 to 4 bits), but existing work faces two key challenges: unconstrained direction error and suboptimal bit allocation. In this paper, we propose RSAVQ, a novel VQ framework to enhance extremely low-bit quantization for LLMs. RSAVQ introduces two geometry-driven innovations that effectively mitigate above limitations: (1) Error Direction Sensitivity Guidance (EDSG), which leverages the Fisher information matrix (FIM)-induced Riemannian metric to project quantization errors onto low-sensitivity directions in the parameter space. Specifically, this projection is performed along the negative natural gradient direction, which effectively suppresses error expansion. (2) Weight Channel Sensitivity Guidance (WCSG) , which constructs a channel-wise sensitivity metric via FIM curvature analysis to dynamically guide bit resource allocation. The approach facilitates a globally optimal quantization solution within prescribed bit constraints. Experiments demonstrate that RSAVQ outperforms existing methods for LLMs. For example, in 2-bit quantization of LLaMA-3 8B, RSAVQ leads baselines like VPTQ and QuIP\# by 0.4 in perplexity (PPL) and 1.5 in zero-shot accuracy. This work offers a practical solution for constrained environments and a theoretical bridge between information geometry and the quantization of neural networks, advancing efficient deep learning. Zukang Xu, Xing Hu 0010, Qiang Wu 0012 |
NeurIPS | 3 |
| 2022 | PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization
Zhihang Yuan, Chenhao Xue, Qiang Wu 0012, Guangyu Sun 0003 |
ECCV (12) | 4 |