Yingtao Shen

dblp:322/3540 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0001-7868-5340ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 River-LLM: Large Language Model Seamless Exit Based on KV Share
abstract
Large Language Models (LLMs) have demonstrated exceptional performance across diverse domains but are increasingly constrained by high inference latency.Early Exit has emerged as a promising solution to accelerate inference by dynamically bypassing redundant layers.However, in decoder-only architectures, the efficiency of Early Exit is severely bottlenecked by the KV Cache Absence problem, where skipped layers fail to provide the necessary historical states for subsequent tokens.Existing solutions, such as recomputation or masking, either introduce significant latency overhead or incur severe precision loss, failing to bridge the gap between theoretical layer reduction and practical wall-clock speedup.In this paper, we propose River-LLM, a training-free framework that enables seamless token-level Early Exit.River-LLM introduces a lightweight KV-Shared Exit River that allows the backbone's missing KV cache to be naturally generated and preserved during the exit process, eliminating the need for costly recovery operations.Furthermore, we utilize state transition similarity within decoder blocks to predict cumulative KV errors and guide precise exit decisions.Extensive experiments on mathematical reasoning and code generation tasks demonstrate that River-LLM achieves 1.71× to 2.16× practical speedup while maintaining high generation quality.
Yingtao Shen, An Zou
ACL (1)1
2026 Compression Space Search: RL-Based Combinational Compression for Neural Networks
abstract
The rising demand for lightweight, high-performance models on mobile and embedded platforms has accelerated the development of model compression techniques. Among these, combinational compression methods—which integrate multiple techniques such as pruning and quantization—offer complementary advantages over using individual methods alone. However, existing research typically focuses on specific combinations designed for a particular model architecture or task. These approaches often overlook the need for a general approach capable of identifying the optimal combination strategy, including the selection, sequence, and degree of applying compression methods. In this paper, we formalize the challenge of combining compression methods—specifically their selection, ordering, and compression degree—as a customized Markov Decision Process defined in a configurable compression space. To solve this, we introduce Compression Space Search (CSS), a practical RL-based framework for automatically and efficiently discovering optimal compression strategies. Experiments across CNN and transformer based vision models demonstrate that the proposed CSS achieves a 30 to 101 times reduction in bit operations while maintaining an accuracy drop of no more than 2%.
Yingtao Shen, Yinchen Ni, Jiace Zhu, An Zou
DATE1
2026 LEAP: Lightweight Neural Network Inference Through Proactive Early-Exiting Prediction
abstract
In recent years, the incorporation of early exit layers into deep neural networks has allowed inference to terminate earlier while maintaining accuracy. However, the passive decision-making involved in the these static exit placement creates a dilemma: fine-grained placement may cause high performance and energy overhead due to frequent exit layer execution, while coarse-grained placement may miss early exit opportunities. Moreover, common energy-saving techniques like adjusting processor configurations are not applicable once inference begins. To overcome these challenges and improve computation and energy efficiency, we propose LEAP, a software-hardware co-design approach. On the software side, LEAP proactively predicts exit points at runtime, reducing computation by enabling early exits without requiring every pre-placed exit layer to be executed. On the hardware side, LEAP adjusts processor settings—such as frequency and voltage—based on single or multiple predicted exits to optimize energy consumption while adhering to latency requirements. Extensive experimental results show that LEAP significantly improves efficiency. Compared to standard inference, LEAP reduces computation by up to 76.4% and saves up to 83.2% in energy. Compared to state-of-the-art early exit methods, LEAP achieves up to 27.9% less computation and 57.1% more energy savings, while maintaining similar accuracy and latency.
Yingtao Shen, Xiangjie Li, Yehan Ma, Weidong Cao 0001, An Zou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 SSMDVFS: Microsecond-Scale DVFS on GPGPUs with Supervised and Self-Calibrated ML
abstract
Over the past decade, as GPUs have evolved to achieve higher computational performance, their power density has also accelerated. Consequently, improving energy efficiency and reducing power consumption has become critically important. Dynamic voltage and frequency scaling (DVFS) is an effective technique for enhancing energy efficiency. With the advent of integrated voltage regulators, DVFS can now operate on microsecond$(\boldsymbol{\mu}\mathbf{s})$timescales. However, developing a practical and effective strategy to guide rapid DVFS remains a significant challenge. This paper proposes a supervised and self-calibrated machine learning framework (SSMDVFS) to guide microsecond-scale GPU voltage and frequency scaling. This framework features an end-to-end design that encompasses data generation, neural network model design, training, compression, and final runtime calibration. Unlike analytical models, which struggle to accurately represent GPU architectures, and reinforcement learning approaches, which can be challenging to converge during runtime, the SSMDVFS offers a practical solution for guiding microsecond-scale voltage and frequency scaling. Experimental results demonstrate that the proposed framework improves energy-delay product (EDP) by 11.09% and outperforms analytical models and reinforcement learning approaches by 13.17% and 36.80 %, respectively.
Minqing Sun, Yingtao Shen, Wei Yan 0005, Qinfen Hao, An Zou
DATE3
2023 Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference
abstract
By adding exiting layers to the deep learning networks, early exit can terminate the inference earlier with accurate results. However, the passive decision-making of whether to exit or continue the next layer has to go through every pre-placed exiting layer until it exits. In addition, it is hard to adjust the configurations of the computing platforms alongside the inference proceeds. By incorporating a low-cost prediction engine, we propose a Predictive Exit framework for computation- and energy-efficient deep learning applications. Predictive Exit can forecast where the network will exit (i.e., establish the number of remaining layers to finish the inference), which effectively reduces the network computation cost by exiting on time without running every pre-placed exiting layer. Moreover, according to the number of remaining layers, proper computing configurations (i.e., frequency and voltage) are selected to execute the network to further save energy. Extensive experimental results demonstrate that Predictive Exit achieves up to 96.2% computation reduction and 72.9% energy-saving compared with classic deep learning networks; and 12.8% computation reduction and 37.6% energy-saving compared with the early exit under state-of-the-art exiting strategies, given the same inference accuracy and latency.
Xiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu, Yingtao Shen, Yehan Ma, An Zou
AAAI5
2023 EENet: Energy Efficient Neural Networks with Run-time Power Management
abstract
Deep learning approaches, such as convolution neural networks (CNNs), have achieved tremendous success in versatile applications. However, one of the challenges to deploy the deep learning models on resource-constrained systems is its huge energy cost. As a dynamic inference approach, early exit adds exiting layers to the networks, which can terminate the inference earlier with accurate results to save energy. The current passive decision-making for energy regulation of early exit cannot adapt to ongoing inference status, varying inference workloads, and timing constraints, let alone guide the reasonable configuration of the computing platforms alongside the inference proceeds for potential energy saving. In this paper, we propose an Energy Efficient Neural Networks (EENet), which introduces a plug-in module to the state-of-the-art networks by incorporating run-time power management. Within each inference, we establish prediction of where the network will exit and adjust computing configurations (i.e., frequency and voltage) accordingly over a small timescale. Considering multiple inferences over a large timescale, we provide frequency and voltage calibration advice, given inference workloads and timing constraints. Finally, the dynamic voltage and frequency scaling (DVFS) governor configures voltage and frequency to execute the network according to the prediction and calibration. Extensive experimental results demonstrate that EENet achieves up to 63.8% energy-saving compared with classic deep learning networks and 21.5% energy-saving compared with the early exit under state-of-the-art exiting strategies, together with improved timing performance.
Xiangjie Li, Yingtao Shen, An Zou, Yehan Ma
DAC2