VLDB 2026 Research / reviewers in the wild / expert
Yushi Huang
dblp:170/8574
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0002-7898-8402ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Efficient and distributed learning · 66% Generative modeling · 29% Language models and text generation · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
3.4 | 4 | 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit · AAAI 2026 Temporal Feature Matters: A Framework for Diffusion Model Quantization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models · ACM Multimedia 2024 |
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Temporal Feature Matters: A Framework for Diffusion Model Quantization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration · ICML 2025 TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models · CVPR 2024 |
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion model quantization |
1.6 | 2 | 2025 | Temporal Feature Matters: A Framework for Diffusion Model Quantization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
1.6 | 2 | 2025 | Temporal Feature Matters: A Framework for Diffusion Model Quantization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models · CVPR 2024 |
Machine learning › Efficient and distributed learning › token reduction
adaptive token pruning |
1.0 | 1 | 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model |
1.0 | 1 | 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
inference acceleration |
1.0 | 1 | 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
KV cache management |
1.0 | 1 | 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization |
1.0 | 1 | 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
1.0 | 1 | 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning
token reduction |
1.0 | 1 | 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
vision-language model compression |
1.0 | 1 | 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion transformer acceleration |
0.9 | 1 | 2025 | HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration · ICML 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
feature caching |
0.9 | 1 | 2025 | HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
post-training sparsity |
0.8 | 1 | 2024 | PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models · ACM Multimedia 2024 |
Performance modeling and evaluation
benchmarking |
0.8 | 1 | 2024 | PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models · ACM Multimedia 2024 |
Machine learning › Efficient and distributed learning
attention computation |
0.3 | 1 | 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.3 | 1 | 2025 | Temporal Feature Matters: A Framework for Diffusion Model Quantization · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
weight quantization |
0.2 | 1 | 2024 | TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
temporal information aware reconstruction · 1.6finite set calibration · 1.6token-level compression · 1.0model-level compression · 1.0dynamic token pruning · 1.0confidence-guided context focusing · 1.0asynchronous KV cache prefetching · 1.0step-wise denoising training · 0.9image error proxy · 0.9caching · 0.9post-training sparsification · 0.8fine-grained pruning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningabstractLong-context inference for Large Language Models (LLMs) is heavily limited by high computational demands. While several existing methods optimize attention computation, they still process the full set of hidden states at each layer, limiting overall efficiency. In this work, we propose SlimInfer, an innovative framework that aims to accelerate inference by directly pruning less critical prompt tokens during the forward pass. Our key insight is an information diffusion phenomenon: As information from critical tokens propagates through layers, it becomes distributed across the entire sequence. This diffusion process suggests that LLMs can maintain their semantic integrity when excessive tokens, even including these critical ones, are pruned in hidden states. Motivated by this, SlimInfer introduces a dynamic fine-grained pruning mechanism that accurately removes redundant tokens of hidden state at intermediate layers. This layer-wise pruning naturally enables an asynchronous KV cache manager that prefetches required token blocks without complex predictors, reducing both memory usage and I/O costs. Extensive experiments show that SlimInfer can achieve up to 2.53× time-to-first-token (TTFT) speedup and 1.88× end-to-end latency reduction for LLaMA3.1-8B-Instruct on a single RTX 4090, without sacrificing performance on LongBench. Lingkun Long, Rubing Yang, Yushi Huang, Desheng Hui, Jianlei Yang 0001 |
AAAI | 3 |
| 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitabstractLarge Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu 0024, Yumeng Shi, Wenya Wang 0001 |
AAAI | 5 |
| 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingabstractLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang, Ao Zhou, Jianlei Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang 0004, Jianlei Yang 0001 |
ACL (1) | 2 |
| 2025 | HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer AccelerationabstractDiffusion Transformers (DiTs) excel in generative tasks but face practical deployment challenges due to high inference costs. Feature caching, which stores and retrieves redundant computations, offers the potential for acceleration. Existing learning-based caching, though adaptive, overlooks the impact of the prior timestep. It also suffers from misaligned objectives-*aligned predicted noise vs. high-quality images*-between training and inference. These two discrepancies compromise both performance and efficiency.
To this end, we *harmonize* training and inference with a novel learning-based *caching* framework dubbed **HarmoniCa**. It first incorporates *Step-Wise Denoising Training* (SDT) to ensure the continuity of the denoising process, where prior steps can be leveraged. In addition, an *Image Error Proxy-Guided Objective* (IEPO) is applied to balance image quality against cache utilization through an efficient proxy to approximate the image error. Extensive experiments across $8$ models, $4$ samplers, and resolutions from $256\times256$ to $2K$ demonstrate superior performance and speedup of our framework. For instance, it achieves over $40\\%$ latency reduction (*i.e.*, $2.07\times$ theoretical speedup) and improved performance on PixArt-$\alpha$. Remarkably, our *image-free* approach reduces training time by $25\\%$ compared with the previous method. Our code is available at https://github.com/ModelTC/HarmoniCa. Yushi Huang, Ruihao Gong, Jing Liu 0048, Jinyang Guo 0002, Xianglong Liu 0001, Jun Zhang 0004 |
ICML | 1 |
| 2025 | Temporal Feature Matters: A Framework for Diffusion Model QuantizationabstractDiffusion models, widely used for image generation, face significant challenges related to their broad applicability due to prolonged inference times and high memory demands. Efficient Post-Training Quantization (PTQ) is crucial to address these issues. However, unlike traditional models, diffusion models critically rely on the time-step for the multi-round denoising. Typically, each time-step is encoded into a hypersensitive temporal feature by several modules. Despite this, existing PTQ methods do not optimize these modules individually. Instead, they employ unsuitable reconstruction objectives and complex calibration methods, leading to significant disturbances in the temporal feature and denoising trajectory, as well as reduced compression efficiency. To address these challenges, we introduce a novel quantization framework that includes three strategies: 1) TIB-based Maintenance: Based on our innovative Temporal Information Block (TIB) definition, Temporal Information-aware Reconstruction (TIAR) and Finite Set Calibration (FSC) are developed to efficiently align original temporal features. 2) Cache-based Maintenance: Instead of indirect and complex optimization for the related modules, pre-computing and caching quantized counterparts of temporal features are developed to minimize errors. 3) Disturbance-aware Selection: Employ temporal feature errors to guide a fine-grained selection between the two maintenance strategies for further disturbance reduction. This framework preserves most of the temporal information and ensures high-quality end-to-end generation. Extensive testing on various datasets, diffusion models and hardware confirms our superior performance and acceleration. Yushi Huang, Ruihao Gong, Xianglong Liu 0001, Jing Liu 0048, Yuhang Li 0001, Jiwen Lu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion ModelsabstractThe Diffusion model, a prevalent framework for image generation, encounters significant challenges in terms of broad applicability due to its extended inference times and substantial memory requirements. Efficient Post-training Quantization (PTQ) is pivotal for addressing these issues in traditional models. Different from traditional models, diffusion models heavily depend on the time-step t to achieve satisfactory multi-round denoising. Usually, t from the finite set {1, …, T} is encoded to a temporal feature by a few modules totally irrespective of the sampling data. However, existing PTQ methods do not optimize these modules separately. They adopt inappropriate reconstruction targets and complex calibration methods, resulting in a severe disturbance of the temporal feature and denoising trajectory, as well as a low compression efficiency. To solve these, we propose a Temporal Feature Maintenance Quantization (TFMQ) framework building upon a Temporal Information Block which is just related to the time-step t and unrelated to the sampling data. Powered by the pioneering block design, we devise temporal information aware reconstruction (TIAR) and finite set calibration (FSC) to align the full-precision temporal features in a limited time. Equipped with the framework, we can maintain the most temporal information and ensure the end-to-end generation quality. Extensive experiments on various datasets and diffusion models prove our state-of-the-art results. Remarkably, our quantization approach, for the first time, achieves model performance nearly on par with the full-precision model under 4-bit weight quantization. Additionally, our method incurs almost no extra computational cost and accelerates quan-tization time by 2.0× on LSUN-Bedrooms 256 × 256 compared to previous works. Our code is publicly available at https://github.com/ModelTC/TFMQ-DM. Yushi Huang, Ruihao Gong, Jing Liu 0048, Tianlong Chen 0001, Xianglong Liu 0001 |
CVPR | 1 |
| 2024 | PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and ModelsabstractWith the increased attention to model efficiency, post-training sparsity (PTS) has become more and more prevalent because of its effectiveness and efficiency. However, there remain questions on better practice of PTS algorithms and the sparsification ability of models, which hinders the further development of this area.Therefore, a benchmark to comprehensively investigate the issues above is urgently needed. In this paper, we propose the first comprehensive post-training sparsity benchmark called PTSBench towards algorithms and models. We benchmark 10+ PTS general-pluggable fine-grained techniques on 3 typical tasks using over 40 off-the-shelf model architectures. Through extensive experiments and analyses, we obtain valuable conclusions and provide several insights from both algorithms and model aspects. Our PTSBench can provide (1) new observations for a better understanding of the PTS algorithms, (2) in-depth and comprehensive evaluations for the sparsification ability of models, and (3) a well-structured and easy-integrate open-source framework. We hope this work will provide illuminating conclusions and advice for future studies of post-training sparsity methods and sparsification-friendly model design. The code for our PTSBench is released at https://github.com/ModelTC/msbench. Jinyang Guo 0002, Ruihao Gong, Yang Yong, Aishan Liu, Yushi Huang, Xianglong Liu 0001 |
ACM Multimedia | 6 |
| 2022 | The Influence of a Virtual Physics Experiment Learning Environment on Grade 9 Students' Motivation towards Physics Learning
Yushi Huang, Alex Wing Cheung Tse |
ICCE | 1 |
| 2011 | Capability as Requirement MetaphorabstractRequirement Engineering (RE) has become an attractive field in both industry and academic. Many RE approaches have been presented in the past years to support eliciting, modeling, analyzing and specifying requirements of system to be built. However, requirement characteristics of kinds of systems like large scale software intensive systems pose several issues to requirements analysis and therefore challenge the extant RE approaches. This paper investigates a number of important metaphors in RE and proposes a novel RE approach that adopts capability as requirement metaphor. We discuss the requirements challenges coming from the changes of system-to-be and argue the necessity to introduce new abstraction and technology into RE to deal with the problems. The notions of capability and the reason to adopt capability as requirement metaphor are analyzed. The meta-model and framework of capability-based requirement engineering is proposed. A case is also studied in order to illustrate our approach. Capability as new abstraction in RE provides a new way to represent, analyze and tradeoff requirements. Jiang Cao, Xinjun Mao, Huining Yan, Yushi Huang, Huaimin Wang 0001, Xicheng Lu |
TrustCom | 4 |