Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hanshi Sun

dblp:314/7377 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 58% Language models and text generation · 39% Reinforcement learning · 3%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.722025
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models · NeurIPS 2025
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference · ICML 2025
Machine learning › Efficient and distributed learning
model compression
1.722025
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models · NeurIPS 2025
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference · ICML 2025
Machine learning › Efficient and distributed learning
inference efficiency
1.122025
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference · ICML 2025
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference · ICML 2025
Natural language and speech › Language models and text generation
decoding
0.812024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.812024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024
Machine learning › Efficient and distributed learning
memory optimization
0.312025
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models · NeurIPS 2025
Memory systems
memory offloading
0.312025
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.212024
Fast Best-of-N Decoding via Speculative Rejection · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

sparse KV selection · 1.7offloading · 1.7low-rank key cache · 1.7token selection · 0.9redundancy-aware compression · 0.9speculative rejection · 0.8reward model · 0.8
YearPublicationVenuePosition
2025 ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
abstract
With the widespread deployment of long-context large language models (LLMs), there has been a growing demand for efficient support of high-throughput inference. However, as the key-value (KV) cache expands with the sequence length, the increasing memory footprint and the need to access it for decoding both result in low throughput when serving long-context LLMs. While various dynamic sparse attention methods have been proposed to accelerate inference while maintaining generation quality, they either fail to sufficiently reduce GPU memory usage or introduce significant decoding latency by offloading the KV cache to the CPU. We present ShadowKV, a high-throughput long-context LLM inference system that stores the low-rank key cache and offloads the value cache to reduce the memory footprint for larger batch sizes and longer sequences. To minimize decoding latency, ShadowKV employs an accurate KV selection strategy that reconstructs minimal sparse KV pairs on-the-fly. By evaluating ShadowKV on benchmarks like RULER, LongBench, and models such as Llama-3.1-8B and GLM-4-9B-1M, we demonstrate that it achieves up to 6$\times$ larger batch sizes and 3.04$\times$ higher throughput on an A100 GPU without sacrificing accuracy, even surpassing the performance achievable with infinite batch size under the assumption of infinite GPU memory.
Hanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng 0001, Ningxin Zheng, Xin Liu 0086, Harry Dong, Yuejie Chi, Beidi Chen
ICML1
2025 R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
abstract
Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 38% of the KV cache. This KV-cache reduction also leads to a 50% memory saving and a 2x speedup over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.
Zefan Cai, Hanshi Sun, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Anima Anandkumar, Abedelkadir Asi, Junjie Hu 0001
NeurIPS3
2024 Fast Best-of-N Decoding via Speculative Rejection
abstract
The safe and effective deployment of Large Language Models (LLMs) involves a critical step called alignment, which ensures that the model's responses are in accordance with human preferences. Prevalent alignment techniques, such as DPO, PPO and their variants, align LLMs by changing the pre-trained model weights during a phase called post-training. While predominant, these post-training methods add substantial complexity before LLMs can be deployed. Inference-time alignment methods avoid the complex post-training step and instead bias the generation towards responses that are aligned with human preferences. The best-known inference-time alignment method, called Best-of-N, is as effective as the state-of-the-art post-training procedures. Unfortunately, Best-of-N requires vastly more resources at inference time than standard decoding strategies, which makes it computationally not viable. In this work, we introduce Speculative Rejection, a computationally-viable inference-time alignment algorithm. It generates high-scoring responses according to a given reward model, like Best-of-N does, while being between 16 to 32 times more computationally efficient.
Hanshi Sun, Momin Haider, Huitao Yang, Jiahao Qiu, Ming Yin 0003, Mengdi Wang 0001, Peter L. Bartlett, Andrea Zanette
NeurIPS1
2023 Combating medical noisy labels by disentangled distribution learning and consistency regularization
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Hanshi Sun
Future Gener. Comput. Syst.4