VLDB 2026 Research / reviewers in the wild / expert
Zhihang Yuan
dblp:195/4180
· DBLP profile ↗
43ranked-venue papers
7as first author
38since 2021 · last 2026
0000-0001-7846-0240ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 5 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Systems, architecture and hardware · 11 · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMsabstractLarge Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real‑world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions. Shaoyuan Chen, Zhixuan Chen, Zhihang Yuan, Qiang Wu 0012 |
AAAI | 4 |
| 2025 | A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model TrainingabstractTraining diffusion models is always a computation-intensive task. In this paper, we introduce a novel speed-up method for diffusion model training, called SpeeD, which is based on a closer look at time steps. Our key findings are: i) Time steps can be empirically divided into acceleration, deceleration, and convergence areas based on the process increment. ii) These time steps are imbalanced, with many concentrated in the convergence area. iii) The concentrated steps provide limited benefits for diffusion training. To address this, we design an asymmetric sampling strategy that reduces the frequency of steps from the convergence area while increasing the sampling probability in other areas. Additionally, we propose a weighting strategy to emphasize the importance of time steps with rapid-change process increments. As a plug-and-play and architecture-agnostic approach, SpeeD consistently achieves 3 × acceleration across various diffusion architectures, datasets, and tasks. Notably, due to its simple design, our approach significantly reduces the cost of diffusion model training with minimal overhead. Our research enables more researchers to train diffusion models at a lower cost. Kai Wang 0036, Mingjia Shi, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, Yang You 0001 |
CVPR | 5 |
| 2025 | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware HistogramabstractReal-time and high-performance 3D object detection plays a critical role in autonomous driving and robotics. Recent pillar-based 3D object detectors have gained significant attention due to their compact representation and low computational overhead, making them suitable for onboard deployment and quantization. However, existing pillar-based detectors still suffer from information loss along height dimension and large numerical distribution difference during pillar feature encoding (PFE), which severely limits their performance and quantization potential. To address above issue, we first unveil the importance of different input information during PFE and identify the height dimension as a key factor in enhancing 3D detection performance. Motivated by this observation, we propose a heightaware pillar feature encoder, called PillarHist. Specifically, PillarHist statistics the discrete distribution of points at different heights within one pillar with the information entropy guidance. This simple yet effective design greatly preserves the information along the height dimension while significantly reducing the computation overhead of the PFE. Meanwhile, PillarHist also constrains the arithmetic distribution of PFE input to a stable range, making it quantization-friendly. Notably, PillarHist operates exclusively within the PFE stage to enhance performance, enabling seamless integration into existing pillar-based methods without introducing complex operations. Extensive experiments show the effectiveness of PillarHist in terms of both efficiency and performance. Sifan Zhou, Zhihang Yuan, Xing Hu 0010, Jian Qian |
CVPR | 2 |
| 2025 | Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed AcceptanceabstractVision-Language-Action (VLA) models have made substantial progress by leveraging the robust capabilities of Visual Language Models (VLMs).However, VLMs' significant parameter size and autoregressive (AR) decoding nature impose considerable computational demands on VLA models.While Speculative Decoding (SD) has shown efficacy in accelerating Large Language Models (LLMs) by incorporating efficient drafting and parallel verification, allowing multiple tokens to be generated in one forward pass, its application to VLA models remains unexplored.This work introduces Spec-VLA, an SD framework designed to accelerate VLA models.Due to the difficulty of the action prediction task and the greedy decoding mechanism of the VLA models, the direct application of the advanced SD framework to the VLA prediction task yields a minor speed improvement.To boost the generation speed, we propose an effective mechanism to relax acceptance utilizing the relative distances represented by the action tokens of the VLA model.Empirical results across diverse test scenarios affirm the effectiveness of the Spec-VLA framework, and further analysis substantiates the impact of our proposed strategies, which enhance the acceptance length by 44%, achieving 1.42× speedup compared with the OpenVLA baseline, without compromising the success rate.The success of the Spec-VLA framework highlights the potential for broader application of speculative execution in VLA prediction scenarios.We make our code and data publicly available at https: //github.com/PineTreeWss/SpecVLA. Songsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu 0005, Yu Wang 0002, Derek F. Wong |
EMNLP | 3 |
| 2025 | QuEST: Low-Bit Diffusion Model Quantization via Efficient Selective FinetuningabstractThe practical deployment of diffusion models is still hindered by the high memory and computational overhead. Although quantization paves a way for model compression and acceleration, existing methods face challenges in achieving low-bit quantization efficiently. In this paper, we identify imbalanced activation distributions as a primary source of quantization difficulty, and propose to adjust these distributions through weight finetuning to be more quantization-friendly. We provide both theoretical and empirical evidence supporting finetuning as a practical and reliable solution. Building on this approach, we further distinguish two critical types of quantized layers: those responsible for retaining essential temporal information and those particularly sensitive to bit-width reduction. By selectively finetuning these layers under both local and global supervision, we mitigate performance degradation while enhancing quantization efficiency. Our method demonstrates its efficacy across three high-resolution image generation tasks, obtaining state-of-the-art performance across multiple bit-width settings. Haoxuan Wang 0002, Yuzhang Shang, Zhihang Yuan, Junyi Wu 0002, Junchi Yan, Yan Yan 0002 |
ICCV | 3 |
| 2025 | Dlfr-Gen: Diffusion-Based Video Generation With Dynamic Latent Frame Rate
Zhihang Yuan, Yuzhang Shang, Hanling Zhang, Siyuan Wang 0002, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ICCV | 1 |
| 2025 | DiTFastAttnV2: Head-Wise Attention Compression for Multi-Modality Diffusion TransformersabstractText-to-image generation models, especially Multimodal Diffusion Transformers (MMDiT), have shown remarkable progress in generating high-quality images. However, these models often face significant computational bottlenecks, particularly in attention mechanisms, which hinder their scalability and efficiency. In this paper, we introduce DiTFastAttnV2, a post-training compression method designed to accelerate attention in MMDiT. Through an in-depth analysis of MMDiT's attention patterns, we identify key differences from prior DiT-based methods and propose head-wise arrow attention and caching mechanisms to dynamically adjust attention heads, effectively bridging this gap. We also design an Efficient Fused Kernel for further acceleration. By leveraging local metric methods and optimization techniques, our approach significantly reduces the search time for optimal compression schemes to just minutes while maintaining generation quality. Furthermore, with the customized kernel, DiTFastAttnV2 achieves a 68% reduction in attention FLOPs and 1.5x end-to-end speedup on 2K image generation without compromising visual fidelity. Hanling Zhang, Rundong Su, Zhihang Yuan, Pengtao Chen, Mingzhu Shen, Yibo Fan, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ICCV | 3 |
| 2025 | EA-Vit: Efficient Adaptation for Elastic Vision TransformerabstractVision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming and energy-intensive. To address this issue, we propose an efficient ViT adaptation framework that enables a single adaptation process to generate multiple models of varying sizes for deployment on platforms with various resource constraints. Our approach comprises two stages. In the first stage, we enhance a pre-trained ViT with a nested elastic architecture that enables structural flexibility across MLP expansion ratio, number of attention heads, embedding dimension, and network depth. To preserve pre-trained knowledge and ensure stable adaptation, we adopt a curriculum-based training strategy that progressively increases elasticity. In the second stage, we design a lightweight router to select submodels according to computational budgets and downstream task demands. Initialized with Pareto-optimal configurations derived via a customized NSGA-II algorithm, the router is then jointly optimized with the backbone. Extensive experiments on multiple benchmarks demonstrate the effectiveness and versatility of EA-ViT. The code is available at https://github.com/zcxcf/EA-ViT. Wangbo Zhao, Yuhao Zhou 0004, Weidong Tang, Shuo Wang 0001, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Kai Wang 0036 |
ICCV | 7 |
| 2025 | OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution FittingabstractPost-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs).
The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values.
Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space.
In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSTQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ.
OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5\% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32\% on the LLaMA-3-8B model compared to state-of-the-art methods. Code will be available. Xing Hu 0010, Zhixuan Chen, Zukang Xu, Jiangyong Yu, Zhihang Yuan, Zhe Jiang 0004, Sifan Zhou |
ICLR | 8 |
| 2025 | MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation MethodsabstractMamba is an efficient sequence model that rivals Transformers and demonstrates significant potential as a foundational architecture for various tasks. Quantization is commonly used in neural networks to reduce model size and computational latency. However, applying quantization to Mamba remains underexplored, and existing quantization methods, which have been effective for CNN and Transformer models, appear inadequate for Mamba models (e.g., Quarot suffers a 21% accuracy drop on Vim-T$\dagger$ even under W8A8). We have pioneered the exploration of this issue and identified several key challenges. First, significant outliers arepresent in gate projections, output projections, and matrix multiplications. Second, Mamba’s unique parallel scan further amplifies these outliers, leading to uneven and heavy-tailed data distributions. Third, even with the application of the Hadamard transform, the variance across channels in weights and activations still remains inconsistent. To these ends, we propose MambaQuant, a post-training quantization (PTQ) framework consisting of: 1) Karhunen-Lo`eve Transformation (KLT) enhanced rotation, rendering the rotation matrix adaptable to diverse channel distributions. 2) Smooth-Fused rotation, which equalizes channel variances and can merge additional parameters into model weights. Experiments show that MambaQuant can quantize both weights and activations into 8-bit with less than 1% accuracy loss for Mamba-based vision and language tasks. To our knowledge, MambaQuant is the first comprehensive PTQ design for the Mamba family, paving the way for further advancements in its application. Zukang Xu, Yuxuan Yue, Xing Hu 0010, Zhihang Yuan, Zixu Jiang, Zhixuan Chen, Jiangyong Yu, Sifan Zhou |
ICLR | 5 |
| 2025 | MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceabstractMixture-of-Experts (MoE) large language models (LLMs), which leverage dynamic routing and sparse activation to enhance efficiency and scalability, have achieved higher performance while reducing computational costs. However, these models face significant memory overheads, limiting their practical deployment and broader adoption. Post-training quantization (PTQ), a widely used method for compressing LLMs, encounters severe accuracy degradation and diminished generalization performance when applied to MoE models. This paper investigates the impact of MoE’s sparse and dynamic characteristics on quantization and identifies two primary challenges: (1) Inter-expert imbalance, referring to the uneven distribution of samples across experts, which leads to insufficient and biased calibration for less frequently utilized experts; (2) Intra-expert imbalance, arising from MoE’s unique aggregation mechanism, which leads to varying degrees of correlation between different samples and their assigned experts. To address these challenges, we propose MoEQuant, a novel quantization framework tailored for MoE LLMs. MoEQuant includes two novel techniques: 1) Expert-Balanced Self-Sampling (EBSS) is an efficient sampling method that efficiently constructs a calibration set with balanced expert distributions by leveraging the cumulative probabilities of tokens and expert balance metrics as guiding factors. 2) Affinity-Guided Quantization (AGQ), which incorporates affinities between experts and samples into the quantization process, thereby accurately assessing the impact of individual samples on different experts within the MoE layer. Experiments demonstrate that MoEQuant achieves substantial performance gains (more than 10 points accuracy gain in the HumanEval for DeepSeekMoE-16B under 4-bit quantization) and boosts efficiency. Zhixuan Chen, Xing Hu 0010, Zukang Xu, Zhihang Yuan, Sifan Zhou, Jiangyong Yu |
ICML | 6 |
| 2025 | MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-DesignabstractMixture-of-Experts (MoE) models face deployment challenges due to their large parameter counts and computational demands. We explore quantization for MoE models and highlight two key insights: 1) linear blocks exhibit varying quantization sensitivity, and 2) divergent expert activation frequencies create heterogeneous computational characteristics. Based on these observations, we introduce MxMoE, a mixed-precision optimization framework for MoE models that considers both algorithmic and system perspectives. MxMoE navigates the design space defined by parameter sensitivity, expert activation dynamics, and hardware resources to derive efficient mixed-precision configurations. Additionally, MxMoE automatically generates optimized mixed-precision GroupGEMM kernels, enabling parallel execution of GEMMs with different precisions. Evaluations show that MxMoE outperforms existing methods, achieving 2.4 lower Wikitext-2 perplexity than GPTQ at 2.25-bit and delivering up to 3.4x speedup over full precision, as well as up to 29.4% speedup over uniform quantization at equivalent accuracy with 5-bit weight-activation quantization. Our code is available at https://github.com/cat538/MxMoE. Haojie Duanmu, Zhihang Yuan, Size Zheng 0001, Jiangfei Duan, Xingcheng Zhang, Dahua Lin |
ICML | 3 |
| 2025 | RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector QuantizationabstractRWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14$\times$ speed up. Yuxuan Yue, Zukang Xu, Xing Hu 0010, Jiangyong Yu, Zhixuan Chen, Sifan Zhou, Zhihang Yuan |
ICML | 8 |
| 2025 | MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
Zhaofeng Hu, Sifan Zhou, Zhihang Yuan, Shibo Zhao, Ci-Jyun Liang |
ICRA | 3 |
| 2025 | Bidirectional Multitask Learning for Non-Autoregressive Machine TranslationabstractNon-Autoregressive Transformer (NART) models generate tokens independently, resulting in lower translation quality than the Autoregressive Transformer (ART) model. To enhance the generation quality, prior Multitask Learning (MTL) frameworks have incorporated a directional Autoregressive (AR) prediction task in conjunction with the Non-Autoregressive (NAR) task. This work proposes further enhancing the NART model with Bidirectional Autoregressive (Bi-AR) prediction tasks. We propose the Bidirectional Multitask Non-Autoregressive Transformer (BM-NART) framework, which enhances the NART decoder model with a weak twin-decoder block, providing AR prediction supervision signal in both directions. To accommodate the bidirectional decoder, we further enhance the Autoregressive Knowledge Distillation (ARKD) with the introduction of Bidirectional Knowledge Distillation (BiKD), which employs dual directional teacher models to provide Bi-AR knowledge distillation data. The experiment confirms that with BiKD, the BM-NART framework achieves generation quality comparable to ART models in BLEU and BERTScore while retaining the advantage of high parallel generation, achieving a 13.7-20 times acceleration with various parameter scalings. Our LLM-based analysis further reveals that the BM-NART framework surpasses the ART model in handling ambiguous translations, knowledge-dependent translations, and reducing hallucinations, illustrating the substantial potential of future NART models.1 Songsheng Wang, Zhihang Yuan, Dongkun Wang, Derek F. Wong |
IJCNN | 3 |
| 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMabstractSRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup. Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003 |
ISCA | 4 |
| 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization
Jiangyong Yu, Sifan Zhou, Shuoyu Li, Shuo Wang 0001, Xing Hu 0010, Zukang Xu, Changyong Shu, Zhihang Yuan |
ACM Multimedia | 10 |
| 2025 | DLFR-VAE: Dynamic Latent Frame Rate VAE for Video GenerationabstractIn this paper, we propose the Dynamic Latent Frame Rate VAE (DLFR-VAE), a training-free paradigm that can make use of adaptive temporal compression in latent space. While existing video generative models apply fixed compression rates via pretrained VAE, we observe that real-world video content exhibits substantial temporal non-uniformity, with high-motion segments containing more information than static scenes. Based on this insight, DLFR-VAE dynamically adjusts the latent frame rate according to the content complexity. Specifically, DLFR-VAE comprises two core innovations: (1) a Dynamic Latent Frame Rate Scheduler that partitions videos into temporal chunks and adaptively determines optimal frame rates based on information-theoretic content complexity, and (2) a training-free adaptation mechanism that transforms pretrained VAE architectures to dynamic VAE that can process features with variable frame rates. Our simple but effective DLFR-VAE can function as a plug-and-play module, seamlessly integrating with existing video generation models and accelerating the video generation process. Zhihang Yuan, Siyuan Wang 0002, Yuzhang Shang, Hanling Zhang, Tongcheng Fang, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
ACM Multimedia | 1 |
| 2025 | R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token RoutingabstractLarge Language Models (LLMs) achieve impressive reasoning capabilities at the cost of substantial inference overhead, posing substantial deployment challenges. Although distilled Small Language Models (SLMs) significantly enhance efficiency, their performance suffers as they fail to follow LLMs' reasoning paths. Luckily, we reveal that only a small fraction of tokens genuinely diverge reasoning paths between LLMs and SLMs. Most generated tokens are either identical or exhibit neutral differences, such as minor variations in abbreviations or expressions. Leveraging this insight, we introduce **Roads to Rome (R2R)**, a neural token router that selectively utilizes LLMs only for these critical, path-divergent tokens, while leaving the majority of token generation to the SLM. We also develop an automatic data generation pipeline that identifies divergent tokens and generates token-level routing labels to train the lightweight router. We apply R2R to combine R1-1.5B and R1-32B models from the DeepSeek family, and evaluate on challenging math, coding, and QA benchmarks. With an average activated parameter size of 5.6B, R2R surpasses the average accuracy of R1-7B by 1.6×, outperforming even the R1-14B model. Compared to R1-32B, it delivers a 2.8× wall-clock speedup with comparable performance, advancing the Pareto frontier of test-time scaling efficiency. Tianyu Fu 0004, Yi Ge, Yichen You, Enshu Liu, Zhihang Yuan, Guohao Dai 0001, Shengen Yan, Huazhong Yang, Yu Wang 0002 |
NeurIPS | 5 |
| 2025 | SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert IdentificationabstractLarge language models with Mixture-of-Experts (MoE) architectures achieve efficiency and scalability, yet their routing mechanisms introduce safety alignment challenges insufficiently addressed by techniques developed for dense models. In this work, the MoE-specific safety risk of positional vulnerability—that safety-aligned behaviors rely on specific expert modules—is formalized and systematically analyzed. An analytical framework, SAFEx, is presented to robustly identify, characterize, and validate safety-critical experts via a stability-based expert selection procedure, and to decompose them into two functional groups: the Harmful Content Detection Group (HCDG), which specializes in identifying and recognizing harmful content within user inputs, and the Harmful Response Control Group (HRCG), which specializes in controlling and enforcing model behaviors to generate appropriate safety responses. Expert-level interventions are conducted to probe causality and to test mitigation. Targeted masking of SAFEx-selected experts reveals that safety behavior is highly concentrated. On Qwen3-30B-A3B, configured with 48 MoE-FFN layers and 128 experts per layer under top-8 routing (48×128=6,144 experts in total), disabling 12 selected experts reduces the refusal rate by 22%. In addition, lightweight adaptation is performed using LoRA under three configurations—the HRCG, the union of HCDG and HRCG, and all experts—and the resulting updates are composed through negative weight merging targeted at the HRCG, leading to improved refusal under adversarial prompts without full-model retraining. These results establish positional vulnerability as a distinct MoE-specific safety challenge and provide a practical, compute-efficient pathway for expert-level safety interventions within routed architectures. Zhenglin Lai, Mengyao Liao, Bingzhe Wu, Dong Xu 0023, Zebin Zhao 0007, Zhihang Yuan |
NeurIPS | 6 |
| 2025 | A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si |
Sci. China Inf. Sci. | 5 |
| 2024 | Algorithm-Hardware Co-Design for Energy-Efficient A/D Conversion in ReRAM-Based AcceleratorsabstractDeep neural networks are widely deployed in many fields. Due to the in-situ computation (known as processing in memory) capacity of the Resistive Random Access Memory (ReRAM) crossbar, ReRAM-based accelerator shows potential in accelerating DNN with low power and high performance. However, despite power advantage, such kind of accelerators suffer from the high power consumption of peripheral circuits, especially Analog-to-Digital Converter (ADC), which account for over 60 percent of total power consumption. This problem hinders the ReRAM-based accelerator to achieve higher efficiency. Some redundant Analog-to-Digital conversion operations have no contribution to maintaining inference accuracy, and such operations can be eliminated by modifying the ADC searching logic. Based on such observations, we propose an algorithm-hardware co-design method and explore the co-design approach in both hardware design and quantization algorithms. Firstly, we focus on the distribution output along the crossbar's bit-lines and identify the fine-grained redundant ADC sampling bits. To further compress ADC bits, we propose a hardware-friendly quantization method and coding scheme, in which different quantization strategy was applied to the partial results in different intervals. To support the two features above, we propose a lightweight architectural design based on SAR-ADC. It's worth mentioning that our method is not only more energy efficient but also retains the flexibility of the algorithm. Experiments demonstrate that our method can reduce about$1.6\sim 2.3\times$ADC power reduction. Zhihang Yuan, Guangyu Sun 0003 |
DATE | 2 |
| 2024 | PB-LLM: Partially Binarized Large Language ModelsabstractThis paper explores network binarization, a radical form of quantization, compressing model weights to a single bit, specifically for Large Language Models (LLMs) compression.
Due to previous binarization methods collapsing LLMs, we propose a novel approach, Partially-Binarized LLM (PB-LLM), which can achieve extreme low-bit quantization while maintaining the linguistic reasoning capacity of quantized LLMs.
Specifically, our exploration first uncovers the ineffectiveness of naïve applications of existing binarization algorithms and highlights the imperative role of salient weights in achieving low-bit quantization.
Thus, PB-LLM filters a small ratio of salient weights during binarization, allocating them to higher-bit storage, i.e., partially-binarization.
PB-LLM is extended to recover the capacities of quantized LMMs, by analyzing from the perspective of post-training quantization (PTQ) and quantization-aware training (QAT).
Under PTQ, combining the concepts from GPTQ, we reconstruct the binarized weight matrix guided by the Hessian matrix and successfully recover the reasoning capacity of PB-LLM in low-bit.
Under QAT, we freeze the salient weights during training, explore the derivation of optimal scaling factors crucial for minimizing the quantization error, and propose a scaling mechanism based on this derived scaling strategy for residual binarized weights.
Those explorations and the developed methodologies significantly contribute to rejuvenating the performance of low-bit quantized LLMs and present substantial advancements in the field of network binarization for LLMs.
Code is available at https://github.com/hahnyuan/PB-LLM. Zhihang Yuan, Yuzhang Shang, Zhen Dong 0003 |
ICLR | 1 |
| 2024 | Learning High-Frequency Functions Made Easy with Sinusoidal Positional EncodingabstractFourier features based positional encoding (PE) is commonly used in machine learning tasks that involve learning high-frequency features from low-dimensional inputs, such as 3D view synthesis and time series regression with neural tangent kernels. Despite their effectiveness, existing PEs require manual, empirical adjustment of crucial hyperparameters, specifically the Fourier features, tailored to each unique task. Further, PEs face challenges in efficiently learning high-frequency functions, particularly in tasks with limited data. In this paper, we introduce sinusoidal PE (SPE), designed to efficiently learn adaptive frequency features closely aligned with the true underlying function. Our experiments demonstrate that SPE, without hyperparameter tuning, consistently achieves enhanced fidelity and faster training across various tasks, including 3D view synthesis, Text-to-Speech generation, and 1D regression. SPE is implemented as a direct replacement for existing PEs. Its plug-and-play nature lets numerous tasks easily adopt and benefit from SPE. Chuanhao Sun, Zhihang Yuan, Kai Xu 0014, Luo Mai, N. Siddharth 0001, Mahesh K. Marina |
ICML | 2 |
| 2024 | DiTFastAttn: Attention Compression for Diffusion Transformer ModelsabstractDiffusion Transformers (DiT) excel at image and video generation but face computational challenges due to the quadratic complexity of self-attention operators. We propose DiTFastAttn, a post-training compression method to alleviate the computational bottleneck of DiT.
We identify three key redundancies in the attention computation during DiT inference: (1) spatial redundancy, where many attention heads focus on local information; (2) temporal redundancy, with high similarity between the attention outputs of neighboring steps; (3) conditional redundancy, where conditional and unconditional inferences exhibit significant similarity. We propose three techniques to reduce these redundancies: (1) $\textit{Window Attention with Residual Sharing}$ to reduce spatial redundancy; (2) $\textit{Attention Sharing across Timesteps}$ to exploit the similarity between steps; (3) $\textit{Attention Sharing across CFG}$ to skip redundant computations during conditional generation. Zhihang Yuan, Hanling Zhang, Lu Pu, Xuefei Ning, Linfeng Zhang 0001, Tianchen Zhao, Shengen Yan, Guohao Dai 0001, Yu Wang 0002 |
NeurIPS | 1 |
| 2024 | Stabilized activation scale estimation for precise Post-Training Quantization
Zhenyang Hao, Xinggang Wang, Jiawei Liu 0006, Zhihang Yuan, Wenyu Liu 0001 |
Neurocomputing | 4 |
| 2024 | Post-training quantization for re-parameterization via coarse & fine weight splitting
Xing Hu 0010, Zhihang Yuan, Jiangyong Yu, Zhe Jiang 0004 |
J. Syst. Archit. | 4 |
| 2024 | Latency-Aware Unified Dynamic Networks for Efficient Image RecognitionabstractDynamic networks have become a pivotal area of study in deep learning due to their ability to selectively activate computing units (such as layers or channels) or dynamically allocate computation to information-rich regions. This capability significantly curtails unnecessary computations, adapting to varying inputs. Despite these advantages, the practical efficiency of dynamic models often falls short of theoretical computation. This discrepancy arises from three primary challenges: 1) a lack of a unified framework across different dynamic inference paradigms due to the fragmented research landscape; 2) an excessive focus on algorithm design at the expense of scheduling strategies, which are essential for optimizing resource utilization on hardware; and 3) the complexity of latency evaluation, since most current libraries cater to static operators. To tackle these issues, we introduce Latency-Aware Unified Dynamic Networks (LAUDNet), a general framework that integrates three fundamental dynamic paradigms-spatially-adaptive computation, layer skipping, and channel skipping-into a single unified formulation. LAUDNet not only refines algorithmic design but also enhances scheduling optimization with the aid of a latency predictor. This predictor efficiently and accurately predicts the inference latency of dynamic operators on specific hardware setups. Our empirical assessments across multiple vision tasks-image classification, object detection, and instance segmentation-confirm that LAUDNet significantly bridges the gap between theoretical and practical efficiency. For instance, LAUDNet cuts down the practical latency of its static counterpart, ResNet-101, by over 50% on hardware platforms like V100, RTX 3090, and TX2 GPUs. Additionally, LAUDNet excels in the accuracy-efficiency trade-off compared to other methods. Yizeng Han, Zhihang Yuan, Yifan Pu, Chaofei Wang, Shiji Song, Gao Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | PD-Quant: Post-Training Quantization Based on Prediction Difference MetricabstractPost-training quantization (PTQ) is a neural network compression technique that converts a full-precision model into a quantized model using lower-precision data types. Although it can help reduce the size and computational cost of deep neural networks, it can also introduce quantization noise and reduce prediction accuracy, especially in extremely low-bit settings. How to determine the appropriate quantization parameters (e.g., scaling factors and rounding of weights) is the main problem facing now. Existing methods attempt to determine these parameters by minimize the distance between features before and after quantization, but such an approach only considers local information and may not result in the most optimal quantization parameters. We analyze this issue and propose PD-Quant, a method that addresses this limitation by considering global information. It determines the quantization parameters by using the information of differences between network prediction before and after quantization. In addition, PD-Quant can alleviate the overfitting problem in PTQ caused by the small number of calibration sets by adjusting the distribution of activations. Experiments show that PD-Quant leads to better quantization parameters and improves the prediction accuracy of quantized models, especially in low-bit settings. For example, PD-Quant pushes the accuracy of ResNet-18 up to 53.14% and RegNetX-600MF up to 40.67% in weight 2-bit activation 2-bit. The code is released at https://github.com/hustv1/PD-Quant. Jiawei Liu 0006, Lin Niu, Zhihang Yuan, Xinggang Wang, Wenyu Liu 0001 |
CVPR | 3 |
| 2023 | Post-Training Quantization on Diffusion ModelsabstractDenoising diffusion (score-based) generative models have recently achieved significant accomplishments in generating realistic and diverse data. Unfortunately, the generation process of current denoising diffusion models is notoriously slow due to the lengthy iterative noise estimations, which rely on cumbersome neural networks. It prevents the diffusion models from being widely deployed, especially on edge devices. Previous works accelerate the generation process of diffusion model (DM) via finding shorter yet effective sampling trajectories. However, they overlook the cost of noise estimation with a heavy network in every iteration. In this work, we accelerate generation from the perspective of compressing the noise estimation network. Due to the difficulty of retraining DMs, we exclude mainstream training-aware compression paradigms and introduce post-training quantization (PTQ) into DM acceleration. However, the output distributions of noise estimation networks change with time-step, making previous PTQ methods fail in DMs since they are designed for single-time step scenarios. To devise a DM-specific PTQ method, we explore PTQ on DM in three aspects: quantized operations, calibration dataset, and calibration metric. We summarize and use several observations derived from all-inclusive investigations to formulate our method, which especially targets the unique multi-time-step structure of DMs. Experimentally, our method can directly quantize full-precision DMs into 8-bit models while maintaining or even improving their performance in a training-free manner. Importantly, our method can serve as a plug-and-play module on other fast-sampling methods, e.g., DDIM [24]. The code is available at https://https://github.com/42Shawn/PTQ4DM. Yuzhang Shang, Zhihang Yuan, Bingzhe Wu, Yan Yan 0002 |
CVPR | 2 |
| 2023 | MIM4DD: Mutual Information Maximization for Dataset DistillationabstractDataset distillation (DD) aims to synthesize a small dataset whose test performance is comparable to a full dataset using the same model. State-of-the-art (SoTA) methods optimize synthetic datasets primarily by matching heuristic indicators extracted from two networks: one from real data and one from synthetic data (see Fig.1, Left), such as gradients and training trajectories. DD is essentially a compression problem that emphasizes on maximizing the preservation of information contained in the data. We argue that well-defined metrics which measure the amount of shared information between variables in information theory are necessary for success measurement, but are never considered by previous works. Thus, we introduce mutual information (MI) as the metric to quantify the shared information between the synthetic and the real datasets, and devise MIM4DD numerically maximizing the MI via a newly designed optimizable objective within a contrastive learning framework to update the synthetic dataset. Specifically, we designate the samples in different datasets who share the same labels as positive pairs, and vice versa negative pairs. Then we respectively pull and push those samples in positive and negative pairs into contrastive space via minimizing NCE loss. As a result, the targeted MI can be transformed into a lower bound represented by feature maps of samples, which is numerically feasible. Experiment results show that MIM4DD can be implemented as an add-on module to existing SoTA DD methods. Yuzhang Shang, Zhihang Yuan, Yan Yan 0002 |
NeurIPS | 2 |
| 2023 | FD-CNN: A Frequency-Domain FPGA Acceleration Scheme for CNN-Based Image-Processing ApplicationsabstractIn the emerging edge-computing scenarios, FPGAs have been widely adopted to accelerate convolutional neural network (CNN)–based image-processing applications, such as image classification, object detection, and image segmentation, and so on. A standard image-processing pipeline first decodes the collected compressed images from Internet of Things (IoTs) to RGB data, then feeds them into CNN engines to compute the results. Previous works mainly focus on optimizing the CNN inference parts. However, we notice that on the popular ZYNQ FPGA platforms, image decoding can also become the bottleneck due to the poor performance of embedded ARM CPUs. Even with a hardware accelerator, the decoding operations still incur considerable latency. Moreover, conventional RGB-based CNNs have too few input channels at the first layer, which can hardly utilize the high parallelism of CNN engines and greatly slows down the network inference. To overcome these problems, in this article, we propose FD-CNN, a novel CNN accelerator leveraging the partial-decoding technique to accelerate CNNs directly in the frequency domain. Specifically, we omit the most time-consuming IDCT (Inverse Discrete Cosine Transform) operations of image decoding and directly feed the DCT coefficients (i.e., the frequency data) into CNNs. By this means, the image decoder can be greatly simplified. Moreover, compared to the RGB data, frequency data has a narrower input resolution but has 64× more channels. Such an input shape is more hardware friendly than RGB data and can substantially reduce the CNN inference time. We then systematically discuss the algorithm, architecture, and command set design of FD-CNN. To deal with the irregularity of different CNN applications, we propose an image-decoding-aware design-space exploration (DSE) workflow to optimize the pipeline. We further propose an early stopping strategy to tackle the time-consuming progressive JPEG decoding. Comprehensive experiments demonstrate that FD-CNN achieves, on average, 3.24×, 4.29× throughput improvement, 2.55×, 2.54× energy reduction and 2.38×, 2.58× lower latency on ZC-706 and ZCU-102 platforms, respectively, compared to the baseline image-processing pipelines. Xiaoyang Wang 0006, Zhe Zhou 0002, Zhihang Yuan, Jingchen Zhu, Kangrui Sun, Guangyu Sun 0003 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | Tailor: removing redundant operations in memristive analog neural network acceleratorsabstractAnalog in-situ computation based on memristive circuits has been regarded as a promising approach for designing high-performance and low-power neural network accelerators. However, despite the low-cost and highly parallel memristive crossbars, the peripheral circuits especially analog-digital-converters (ADCs) induce significant overhead. Quantitative analysis shows that ADCs can contribute up to 91% energy consumption and 72% chip area, which significantly offset the advantages of memristive NN accelerators. Zhihang Yuan, Guangyu Sun 0003, Zhichao Lu |
DAC | 2 |
| 2022 | PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization
Zhihang Yuan, Chenhao Xue, Qiang Wu 0012, Guangyu Sun 0003 |
ECCV (12) | 1 |
| 2022 | Enabling High-Quality Uncertainty Quantification in a PIM Designed for Bayesian Neural NetworkabstractUncertainty quantification measures the prediction uncertainty of a neural network facing out-of-training-distribution samples. Bayesian Neural Networks (BNNs) can provide high-quality uncertainty quantification by introducing specific noise to the weights during inference. To accelerate BNN inference, ReRAM processing-in-memory (PIM) architecture is a competitive solution to provide both high-efficient computing and in-situ noise generation at the same time. However, there normally exists a huge gap between the generated noise in PIM hardware and that required by a BNN model. We demonstrate that the quality of uncertainty quantification is substantially degraded due to this gap. To solve this problem, we propose a holistic framework called W2W-PIM. We first introduce an efficient method to generate noise in ReRAM PIM design according to the demand of a BNN model. In addition, the PIM architecture is carefully modified to enable the noise generation and evaluate uncertainty quality. Moreover, a calibration unit is further introduced to reduce the noise gap caused by imperfection of the noise model. Comprehensive evaluation results demonstrate that W2W-PIM framework can achieve high-quality uncertainty quantification and high energy-efficiency at the same time. Bingzhe Wu, Guangyu Sun 0003, Zhe Zhang 0006, Zhihang Yuan, Runsheng Wang, Ru Huang 0001, Dimin Niu, Hongzhong Zheng, Zhichao Lu, Meng-Fan Chang, Tianchan Guan, Xin Si |
HPCA | 5 |
| 2022 | Latency-aware Spatial-wise Dynamic NetworksabstractSpatial-wise dynamic convolution has become a promising approach to improving the inference efficiency of deep networks. By allocating more computation to the most informative pixels, such an adaptive inference paradigm reduces the spatial redundancy in image features and saves a considerable amount of unnecessary computation. However, the theoretical efficiency achieved by previous methods can hardly translate into a realistic speedup, especially on the multi-core processors (e.g. GPUs). The key challenge is that the existing literature has only focused on designing algorithms with minimal computation, ignoring the fact that the practical latency can also be influenced by scheduling strategies and hardware properties. To bridge the gap between theoretical computation and practical efficiency, we propose a latency-aware spatial-wise dynamic network (LASNet), which performs coarse-grained spatially adaptive inference under the guidance of a novel latency prediction model. The latency prediction model can efficiently estimate the inference latency of dynamic networks by simultaneously considering algorithms, scheduling strategies, and hardware properties. We use the latency predictor to guide both the algorithm design and the scheduling optimization on various hardware platforms. Experiments on image classification, object detection and instance segmentation demonstrate that the proposed framework significantly improves the practical inference efficiency of deep networks. For example, the average latency of a ResNet-101 on the ImageNet validation set could be reduced by 36% and 46% on a server GPU (Nvidia Tesla-V100) and an edge device (Nvidia Jetson TX2 GPU) respectively without sacrificing the accuracy. Code is available at https://github.com/LeapLabTHU/LASNet. Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun 0003, Gao Huang 0001 |
NeurIPS | 2 |
| 2022 | Flatfish: A Reinforcement Learning Approach for Application-Aware Address MappingabstractThe DRAM performance has become a critical bottleneck of modern computing systems. Prior studies have proposed various optimization techniques on address mapping to bridge the gap between real performance and the peak performance. Nevertheless, these techniques have some common limitations. First, most of them focus on an indirect metric (e.g., bitwise flip ratio) and fail to address the effects of complicated organization hierarchy and timing constraints of DRAM. Second, these approaches do not leverage application-specific information and may not generate the proper address mapping schemes for modern applications. In this article, we propose Flatfish as a comprehensive solution to address these challenges. Flatfish is a self-adaptive memory controller that is able to generate address mapping schemes according to the memory access pattern. Different from prior approaches, Flatfish considers complicated memory hierarchy, including channel, rank, and bank group and addressed critical timing constraints. By mining the characteristics from the memory access traces, Flatfish integrates a reinforcement learning model to generate a binary invertible matrix (BIM) as the address mapping scheme. Flatfish can work in either offline mode or online mode to meet various requirements in different scenarios. The experimental results show that Flatfish can achieve$1.91\times $speedup in the offline mode on GPU, and$1.63\times $speedup in the online mode on CPU, over the commonly used Hynix address mapping scheme. Zhihang Yuan, Yijin Guan, Guangyu Sun 0003, Tao Zhang 0032, Rongshan Wei, Dimin Niu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | NAS4RRAM: neural network architecture search for inference on RRAM-based accelerators
Zhihang Yuan, Jingze Liu, Longhao Yan, Haoxiang Chen 0003, Bingzhe Wu, Yuchao Yang 0001, Guangyu Sun 0003 |
Sci. China Inf. Sci. | 1 |
| 2020 | S2DNAS: Transforming Static CNN Model for Dynamic Inference via Neural Architecture Search
Zhihang Yuan, Bingzhe Wu, Guangyu Sun 0003, Zheng Liang 0003, Shiwan Zhao, Weichen Bi |
ECCV (2) | 1 |
| 2020 | Crane: Mitigating Accelerator Under-utilization Caused by Sparsity Irregularities in CNNsabstractConvolutional neural networks (CNNs) have achieved great success in numerous AI applications. To improve inference efficiency of CNNs, researchers have proposed various pruning techniques to reduce both computation intensity and storage overhead. These pruning techniques result in multi-level sparsity irregularities in CNNs. Together with that in activation matrices, which is induced by employment of ReLU activation function, all these sparsity irregularities cause a serious problem of computation resource under-utilization in sparse CNN accelerators. To mitigate this problem, we propose a method of load-balancing based on a workload stealing technique. We demonstrate that this method can be applied to two major inference data-flows, which cover all state-of-the-art sparse CNN accelerators. Based on this method, we present an accelerator, called Crane, which addresses all kinds of sparsity irregularities in CNNs. We perform a fair comparison between Crane and state-of-the-art prior approaches. Experimental results show that Crane improves performance by 27% ~ 88% and reduces energy consumption by 16% ~ 48%, respectively, compared to the counterparts. Yijin Guan, Guangyu Sun 0003, Zhihang Yuan, Ningyi Xu, Jason Cong, Yuan Xie 0001 |
IEEE Trans. Computers | 3 |
| 2017 | Using Data Compression for Optimizing FPGA-Based Convolutional Neural Network Accelerators
Yijin Guan, Ningyi Xu, Chen Zhang 0001, Zhihang Yuan, Jason Cong |
APPT | 4 |
| 2017 | FPGA-based accelerator for long short-term memory recurrent neural networksabstractLong Short-Term Memory Recurrent neural networks (LSTM-RNNs) have been widely used for speech recognition, machine translation, scene analysis, etc. Unfortunately, general-purpose processors like CPUs and GPGPUs can not implement LSTM-RNNs efficiently due to the recurrent nature of LSTM-RNNs. FPGA-based accelerators have attracted attention of researchers because of good performance, high energy-efficiency and great flexibility. In this work, we present an FPGA-based accelerator for LSTM-RNNs that optimizes both computation performance and communication requirements. The peak performance of our accelerator achieves 7.26 GFLOP/S, which significantly outperforms previous approaches. Yijin Guan, Zhihang Yuan, Guangyu Sun 0003, Jason Cong |
ASP-DAC | 2 |
| 2017 | Reducing Overfitting in Deep Convolutional Neural Networks Using Redundancy Regularizer
Bingzhe Wu, Zhihang Yuan, Guangyu Sun 0003, Charles Wu |
ICANN (2) | 3 |