VLDB 2026 Research / reviewers in the wild / expert
Yang Yong
dblp:24/167
· DBLP profile ↗
11ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitabstractLarge Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu 0024, Yumeng Shi, Wenya Wang 0001 |
AAAI | 3 |
| 2026 | Adaptive hybrid speculative decoding for accelerating large language model inferenceabstractProcessing long sequences efficiently is crucial for modern large language models (LLM), yet serving these models incurs high inference costs due to computational bottlenecks in speculative decoding. Existing speculative decoding methods leverage fixed-granularity draft–verification strategies, which introduce a trade-off between inference speed and accuracy: a larger draft step accelerates decoding but increases verification overhead, while a smaller step ensures correctness but reduces efficiency. To overcome this trade-off, we propose Adaptive Hybrid Speculative Decoding (AHSD) , a novel framework built upon two key algorithmic innovations: Adaptive Feedback Thresholds and Confidence-Driven Two-Stage Token Processing, implemented efficiently via a bucketed batching optimization. First, the adaptive feedback threshold dynamically adjusts rollback and validation by monitoring the draft model’s confidence in real time, if confidence exceeds the threshold, verification is skipped; otherwise, a refining check is triggered, maximizing first-pass success while minimizing false positives. Second, the confidence-driven two-stage token processing decouples draft generation from verification, allowing high-confidence tokens to bypass verification entirely. To ensure these algorithmic innovations execute efficiently on modern GPUs, we employ a bucketed batching strategy that groups sequences by dynamically chosen block sizes , pads each homogeneous batch to its maximum length with a unified attention mask, and executes all sequences in parallel, thereby eliminating scheduling overhead and branch divergence caused by length variability.This sharply reduces re-model frequency and enhances parallelism. Comprehensive simulations show that AHSD achieves up to speedup and throughput gain over existing schemes, making it suitable for applications that can tolerate minor quality loss for substantially faster inference. Yang Yong, Ting Ting Yang, Shao Shuai Gao, Ye Hua Yin |
Neurocomputing | 1 |
| 2025 | Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool InvocationabstractThe rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only provide end-to-end scores but lack in-depth analysis and often suffer from issues such as instability. To address this gap, we have meticulously designed the Tool Playgrounds framework, a comprehensive, analyzable, and extensible benchmark. This framework evaluates boundary dimensions such as parameter missing interaction, parameter correction, tool failover, and leveraging internal knowledge. Our findings indicate that even the most advanced commercial models frequently overlook these essential aspects and face challenges in managing complex tool usage. To foster further research and development, we have made our code, dataset, and leaderboard publicly available on https://github.com/zhiwei-dong/ToolPlaygrounds. Zhiwei Dong, Ruihao Gong, Yang Yong, Yongqiang Yao, Song-Lu Chen, Xu-Cheng Yin |
ICASSP | 3 |
| 2025 | A survey of low-bit large language models: Basics, systems, and algorithms
Ruihao Gong, Yifu Ding 0001, Chengtao Lv, Xingyu Zheng, Jinyang Du, Yang Yong, Shiqiao Gu, Haotong Qin, Jinyang Guo 0002, Dahua Lin, Michele Magno, Xianglong Liu 0001 |
Neural Networks | 7 |
| 2024 | Fast and Controllable Post-training Sparsity: Learning Optimal Sparsity Allocation with Global Constraint in MinutesabstractNeural network sparsity has attracted many research interests due to its similarity to biological schemes and high energy efficiency. However, existing methods depend on long-time training or fine-tuning, which prevents large-scale applications. Recently, some works focusing on post-training sparsity (PTS) have emerged. They get rid of the high training cost but usually suffer from distinct accuracy degradation due to neglect of the reasonable sparsity rate at each layer. Previous methods for finding sparsity rates mainly focus on the training-aware scenario, which usually fails to converge stably under the PTS setting with limited data and much less training cost. In this paper, we propose a fast and controllable post-training sparsity (FCPTS) framework. By incorporating a differentiable bridge function and a controllable optimization objective, our method allows for rapid and accurate sparsity allocation learning in minutes, with the added assurance of convergence to a predetermined global sparsity rate. Equipped with these techniques, we can surpass the state-of-the-art methods by a large margin, e.g., over 30\% improvement for ResNet-50 on ImageNet under the sparsity rate of 80\%. Our plug-and-play code and supplementary materials are open-sourced at https://github.com/ModelTC/FCPTS. Ruihao Gong, Yang Yong, Jinyang Guo 0002, Xiuying Wei, Yuqing Ma, Xianglong Liu 0001 |
AAAI | 2 |
| 2024 | PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and ModelsabstractWith the increased attention to model efficiency, post-training sparsity (PTS) has become more and more prevalent because of its effectiveness and efficiency. However, there remain questions on better practice of PTS algorithms and the sparsification ability of models, which hinders the further development of this area.Therefore, a benchmark to comprehensively investigate the issues above is urgently needed. In this paper, we propose the first comprehensive post-training sparsity benchmark called PTSBench towards algorithms and models. We benchmark 10+ PTS general-pluggable fine-grained techniques on 3 typical tasks using over 40 off-the-shelf model architectures. Through extensive experiments and analyses, we obtain valuable conclusions and provide several insights from both algorithms and model aspects. Our PTSBench can provide (1) new observations for a better understanding of the PTS algorithms, (2) in-depth and comprehensive evaluations for the sparsification ability of models, and (3) a well-structured and easy-integrate open-source framework. We hope this work will provide illuminating conclusions and advice for future studies of post-training sparsity methods and sparsification-friendly model design. The code for our PTSBench is released at https://github.com/ModelTC/msbench. Jinyang Guo 0002, Ruihao Gong, Yang Yong, Aishan Liu, Yushi Huang, Xianglong Liu 0001 |
ACM Multimedia | 4 |
| 2022 | Compressing Models with Few Samples: Mimicking then ReplacingabstractFew-sample compression aims to compress a big redundant model into a small compact one with only few samples. If we fine-tune models with these limited few samples directly, models will be vulnerable to overfit and learn almost nothing. Hence, previous methods optimize the compressed model layer-by-layer and try to make every layer have the same outputs as the corresponding layer in the teacher model, which is cumbersome. In this paper, we propose a new framework named Mimicking then Replacing (MiR) for few-sample compression, which firstly urges the pruned model to output the same features as the teacher's in the penultimate layer, and then replaces teacher's layers before penultimate with a well-tuned compact one. Unlike previous layer-wise reconstruction methods, our MiR optimizes the entire network holistically, which is not only simple and effective, but also unsupervised and general. MiR outperforms previous methods with large margins. Codes is available at https://github.com/cjnjuwhy/MiR. Junjie Liu 0003, Xin Ma 0031, Yang Yong, Zhenhua Chai, Jianxin Wu 0001 |
CVPR | 4 |
| 2021 | Hate Speech Detection Based on Sentiment Knowledge SharingabstractXianbing Zhou, Yang Yong, Xiaochao Fan, Ge Ren, Yunfeng Song, Yufeng Diao, Liang Yang, Hongfei Lin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xianbing Zhou, Yang Yong, Xiaochao Fan, Yunfeng Song, Yufeng Diao, Liang Yang 0003, Hongfei Lin |
ACL/IJCNLP (1) | 2 |
| 2019 | An improved general type-2 fuzzy sets type reduction and its application in general type-2 fuzzy controller design
Jianzhong Shi, Liang Shaohua, Yang Yong |
Soft Comput. | 3 |
| 2012 | Position variable structure control for water hydraulic vane actuatorabstractIn the paper a variable structure control (VSC) has been proposed and applied to the position control of a water hydraulic rotary actuator. Based on the model of the water hydraulic vane actuator, the simulation experiments with VSC controller are done and compared with PID controller. Simulation experiment results show that in position control applications, the use of VSC as the controller of the system can improve the dynamical and static state performances of the system. Especially, the VSC system provides higher-precision position as well as stronger robustness to parameters varieties than the PID control system. Also, the chattering has been effectively controlled and reduced in the VSC application by adding a boundary layer. The position control properties of the water hydraulic rotary actuator validate the effectiveness of the proposed VSC method. Yang Yong |
ICARCV | 1 |
| 2006 | Coordination Optimization-based Variable Structure Control for Main Steam Pressure of Power PlantabstractAn optimal variable structure control (VSC) based on a coordination genetic algorithm has been developed. Steady-state error and control switching frequency are used to constitute the system performance indexes in the coordination optimization, while the tuning rate of boundary layer width (BLW) is employed as the optimization parameter. Then based on the mathematical relationship between BLW tuning law and steady state error, an optimized BLW tuning rate is added to the nonlinear control term of VSC. Simulation experiment results applied to the main steam pressure control (MSPC) of power plant show the comprehensive superiority of dynamical and static state performance by using the proposed controller over that of by using an optimized PID control. The proposed VSC system has better robustness against large parameter variations and disturbances. This succeeds in coordinately considering both chattering reduction and high-precision control in VSC Yang Yong, Luo An |
ICARCV | 1 |