VLDB 2026 Research / reviewers in the wild / expert
Xing Hu 0010
dblp:49/10052-10
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2026
0009-0003-4510-898XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Efficient and distributed learning · 54% 3D vision · 13% Generative modeling · 9% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 56% Electronic design automation · 36% Hardware accelerators and domain-specific architectures · 8% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
4.5 | 5 | 2026 | FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection · AAAI 2026 RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization · ICML 2025 MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
3.5 | 4 | 2025 | RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization · ICML 2025 MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance · ICML 2025 MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods · ICLR 2025 |
Computer vision › 3D vision
3d object detection |
1.9 | 2 | 2026 | FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection · AAAI 2026 PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogram · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
1.9 | 2 | 2026 | FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection · AAAI 2026 MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods · ICLR 2025 |
Machine learning › Representation and self-supervised learning
vector quantization |
1.9 | 2 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization · ICML 2025 |
Machine learning › Generative modeling
image tokenization |
1.0 | 1 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 |
Machine learning › Generative modeling
variational autoencoder |
1.0 | 1 | 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › quantization › transformer quantization
large language model quantization |
0.9 | 1 | 2025 | OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting · ICLR 2025 |
Natural language and speech › Language models and text generation › neural language model
mixture-of-experts language model |
0.9 | 1 | 2025 | MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance · ICML 2025 |
Machine learning › Efficient and distributed learning
model quantization |
0.9 | 1 | 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization · ACM Multimedia 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization · ACM Multimedia 2025 |
Computer vision › 3D vision › 3d object detection › point cloud object detection
pillar-based 3d object detection |
0.9 | 1 | 2025 | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogram · CVPR 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.9 | 1 | 2025 | MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods · ICLR 2025 |
Memory systems › processing-in-memory
computing-in-memory |
0.9 | 1 | 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs · Sci. China Inf. Sci. 2025 |
Electronic design automation › power integrity
IR-drop |
0.9 | 1 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 |
Natural language and speech › Language models and text generation › large language model
large language model deployment |
0.3 | 1 | 2025 | MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance · ICML 2025 |
Robotics › Autonomous driving
perception |
0.3 | 1 | 2025 | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogram · CVPR 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.3 | 1 | 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs · Sci. China Inf. Sci. 2025 |
Electronic design automation
physical design |
0.3 | 1 | 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM · ISCA 2025 |
Methods — techniques the papers use, named apart from their topics
vector quantization · 1.0variational autoencoder · 1.0position embedding transformation · 1.0numerical stabilization · 1.0lookup table approximation · 1.0distribution regularization · 1.0task mapping · 0.9software-hardware co-design · 0.9quantization-aware feature encoding · 0.9orthogonal transformation · 0.9height-aware histogram · 0.9KL-Top loss · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VAEVQ: Enhancing Discrete Visual Tokenization Through Variational ModelingabstractVector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and after quantization, and poor coherence between the continuous and discrete domains. These issues lead to unstable codeword learning and underutilized codebooks, ultimately degrading the performance of both reconstruction and downstream generation tasks. To this end, we propose VAEVQ, which comprises three key components: (1) Variational Latent Quantization (VLQ), replacing the AE with a VAE for quantization to leverage its structured and smooth latent space, thereby facilitating more effective codeword activation; (2) Representation Coherence Strategy (RCS), adaptively modulating the alignment strength between pre- and post-quantization features to enhance consistency and prevent overfitting to noise; and (3) Distribution Consistency Regularization (DCR), aligning the entire codebook distribution with the continuous latent distribution to improve utilization. Extensive experiments on two benchmark datasets demonstrate that VAEVQ outperforms state-of-the-art methods. Sicheng Yang 0001, Xing Hu 0010, Qiang Wu 0012 |
AAAI | 2 |
| 2026 | FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object DetectionabstractCamera-based multi-view 3D detection is crucial for autonomous driving. PETR and its variants (PETRs) excel in benchmarks but face deployment challenges due to high computational cost and memory footprint. Quantization is an effective technique for compressing deep neural networks by reducing the bit width of weights and activations. However, directly applying existing quantization methods to PETRs leads to severe accuracy degradation. This issue primarily arises from two key challenges: (1) significant magnitude disparity between multi-modal features—specifically, image features and camera-ray positional embeddings (PE), and (2) the inefficiency and approximation error of quantizing non-linear operators, which commonly rely on hardware-unfriendly computations. In this paper, we propose FQ-PETR, a fully quantized framework for PETRs, featuring three key innovations: (1) Quantization-Friendly LiDAR-ray Position Embedding (QFPE): Replacing multi-point sampling with LiDAR-prior-guided single-point sampling and anchor-based embedding eliminates problematic non-linearities (e.g., inverse-sigmoid) and aligns PE scale with image features, preserving accuracy. (2) Dual-Lookup Table (DULUT): This algorithm approximates complex non-linear functions using two cascaded linear LUTs, achieving high fidelity with minimal entries and no specialized hardware. (3) Quantization After Numerical Stabilization (QANS): Performing quantization after softmax numerical stabilization mitigates attention distortion from large inputs. On PETRs (e.g., PETR, StreamPETR, PETRv2, MV2d), FQ-PETR under W8A8 achieves near-floating-point accuracy ( Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu 0010 |
AAAI | 5 |
| 2025 | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware HistogramabstractReal-time and high-performance 3D object detection plays a critical role in autonomous driving and robotics. Recent pillar-based 3D object detectors have gained significant attention due to their compact representation and low computational overhead, making them suitable for onboard deployment and quantization. However, existing pillar-based detectors still suffer from information loss along height dimension and large numerical distribution difference during pillar feature encoding (PFE), which severely limits their performance and quantization potential. To address above issue, we first unveil the importance of different input information during PFE and identify the height dimension as a key factor in enhancing 3D detection performance. Motivated by this observation, we propose a heightaware pillar feature encoder, called PillarHist. Specifically, PillarHist statistics the discrete distribution of points at different heights within one pillar with the information entropy guidance. This simple yet effective design greatly preserves the information along the height dimension while significantly reducing the computation overhead of the PFE. Meanwhile, PillarHist also constrains the arithmetic distribution of PFE input to a stable range, making it quantization-friendly. Notably, PillarHist operates exclusively within the PFE stage to enhance performance, enabling seamless integration into existing pillar-based methods without introducing complex operations. Extensive experiments show the effectiveness of PillarHist in terms of both efficiency and performance. Sifan Zhou, Zhihang Yuan, Xing Hu 0010, Jian Qian |
CVPR | 4 |
| 2025 | OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution FittingabstractPost-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs).
The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values.
Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space.
In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSTQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ.
OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5\% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32\% on the LLaMA-3-8B model compared to state-of-the-art methods. Code will be available. Xing Hu 0010, Zhixuan Chen, Zukang Xu, Jiangyong Yu, Zhihang Yuan, Zhe Jiang 0004, Sifan Zhou |
ICLR | 1 |
| 2025 | MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation MethodsabstractMamba is an efficient sequence model that rivals Transformers and demonstrates significant potential as a foundational architecture for various tasks. Quantization is commonly used in neural networks to reduce model size and computational latency. However, applying quantization to Mamba remains underexplored, and existing quantization methods, which have been effective for CNN and Transformer models, appear inadequate for Mamba models (e.g., Quarot suffers a 21% accuracy drop on Vim-T$\dagger$ even under W8A8). We have pioneered the exploration of this issue and identified several key challenges. First, significant outliers arepresent in gate projections, output projections, and matrix multiplications. Second, Mamba’s unique parallel scan further amplifies these outliers, leading to uneven and heavy-tailed data distributions. Third, even with the application of the Hadamard transform, the variance across channels in weights and activations still remains inconsistent. To these ends, we propose MambaQuant, a post-training quantization (PTQ) framework consisting of: 1) Karhunen-Lo`eve Transformation (KLT) enhanced rotation, rendering the rotation matrix adaptable to diverse channel distributions. 2) Smooth-Fused rotation, which equalizes channel variances and can merge additional parameters into model weights. Experiments show that MambaQuant can quantize both weights and activations into 8-bit with less than 1% accuracy loss for Mamba-based vision and language tasks. To our knowledge, MambaQuant is the first comprehensive PTQ design for the Mamba family, paving the way for further advancements in its application. Zukang Xu, Yuxuan Yue, Xing Hu 0010, Zhihang Yuan, Zixu Jiang, Zhixuan Chen, Jiangyong Yu, Sifan Zhou |
ICLR | 3 |
| 2025 | MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceabstractMixture-of-Experts (MoE) large language models (LLMs), which leverage dynamic routing and sparse activation to enhance efficiency and scalability, have achieved higher performance while reducing computational costs. However, these models face significant memory overheads, limiting their practical deployment and broader adoption. Post-training quantization (PTQ), a widely used method for compressing LLMs, encounters severe accuracy degradation and diminished generalization performance when applied to MoE models. This paper investigates the impact of MoE’s sparse and dynamic characteristics on quantization and identifies two primary challenges: (1) Inter-expert imbalance, referring to the uneven distribution of samples across experts, which leads to insufficient and biased calibration for less frequently utilized experts; (2) Intra-expert imbalance, arising from MoE’s unique aggregation mechanism, which leads to varying degrees of correlation between different samples and their assigned experts. To address these challenges, we propose MoEQuant, a novel quantization framework tailored for MoE LLMs. MoEQuant includes two novel techniques: 1) Expert-Balanced Self-Sampling (EBSS) is an efficient sampling method that efficiently constructs a calibration set with balanced expert distributions by leveraging the cumulative probabilities of tokens and expert balance metrics as guiding factors. 2) Affinity-Guided Quantization (AGQ), which incorporates affinities between experts and samples into the quantization process, thereby accurately assessing the impact of individual samples on different experts within the MoE layer. Experiments demonstrate that MoEQuant achieves substantial performance gains (more than 10 points accuracy gain in the HumanEval for DeepSeekMoE-16B under 4-bit quantization) and boosts efficiency. Zhixuan Chen, Xing Hu 0010, Zukang Xu, Zhihang Yuan, Sifan Zhou, Jiangyong Yu |
ICML | 2 |
| 2025 | RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector QuantizationabstractRWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14$\times$ speed up. Yuxuan Yue, Zukang Xu, Xing Hu 0010, Jiangyong Yu, Zhixuan Chen, Sifan Zhou, Zhihang Yuan |
ICML | 4 |
| 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMabstractSRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup. Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003 |
ISCA | 2 |
| 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization
Jiangyong Yu, Sifan Zhou, Shuoyu Li, Shuo Wang 0001, Xing Hu 0010, Zukang Xu, Changyong Shu, Zhihang Yuan |
ACM Multimedia | 6 |
| 2025 | RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their exponentially increasing parameters pose significant challenges for deployment on resource-constrained devices. Vector Quantization (VQ) shows great promise for low-bit quantization (e.g., 2 to 4 bits), but existing work faces two key challenges: unconstrained direction error and suboptimal bit allocation. In this paper, we propose RSAVQ, a novel VQ framework to enhance extremely low-bit quantization for LLMs. RSAVQ introduces two geometry-driven innovations that effectively mitigate above limitations: (1) Error Direction Sensitivity Guidance (EDSG), which leverages the Fisher information matrix (FIM)-induced Riemannian metric to project quantization errors onto low-sensitivity directions in the parameter space. Specifically, this projection is performed along the negative natural gradient direction, which effectively suppresses error expansion. (2) Weight Channel Sensitivity Guidance (WCSG) , which constructs a channel-wise sensitivity metric via FIM curvature analysis to dynamically guide bit resource allocation. The approach facilitates a globally optimal quantization solution within prescribed bit constraints. Experiments demonstrate that RSAVQ outperforms existing methods for LLMs. For example, in 2-bit quantization of LLaMA-3 8B, RSAVQ leads baselines like VPTQ and QuIP\# by 0.4 in perplexity (PPL) and 1.5 in zero-shot accuracy. This work offers a practical solution for constrained environments and a theoretical bridge between information geometry and the quantization of neural networks, advancing efficient deep learning. Zukang Xu, Xing Hu 0010, Qiang Wu 0012 |
NeurIPS | 2 |
| 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si |
Sci. China Inf. Sci. | 6 |
| 2024 | Post-training quantization for re-parameterization via coarse & fine weight splitting
Xing Hu 0010, Zhihang Yuan, Jiangyong Yu, Zhe Jiang 0004 |
J. Syst. Archit. | 3 |
| 2022 | 3D Object Detection Based on Multi-scale Feature Fusion and Contrastive Learningabstract3D object detection plays an increasingly important role in the understanding of real natural scenes. In recent years, the method based on Hough voting has attracted more and more attention because of its compact model and high efficiency. However, the current voting strategy only uses poor local proposal information, which is not conducive to model optimization and performance improvement. A 3D object detection model based on multi-scale feature fusion and contrastive learning is proposed in this paper. The proposal stage focuses on voting to generate proposals, including the coordinates of proposals and the corresponding feature vectors. In order to obtain the local structure information, then calculate the multi-scale attention, and trace the multi-scale features of the proposal, an additional proposed coding contrast branch is introduced, which uses the proposed feature coding contrast loss to jointly optimize the feature representation and multi-scale attention modules. We have obtained competitive results on two large datasets SUN RGBD and ScanNet, which shows the effectiveness of our method. Xing Hu 0010, Tulga Khuyag |
SMC | 2 |