Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yusheng Qin

dblp:24/5399 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0004-7452-3696ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 91% Question answering and dialogue systems · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization
1.012026
TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache Transformation · IEEE Trans. Computers 2026
Machine learning › Efficient and distributed learning › inference efficiency
KV cache reuse
1.012026
TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache Transformation · IEEE Trans. Computers 2026
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
0.312026
TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache Transformation · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

QKV computation · 1.0KV cache transformation · 1.0
YearPublicationVenuePosition
2026 TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache Transformation
abstract
Multi-turn conversational systems based on large language models are increasingly being integrated into web platforms and applied across a wide range of domains. However, these systems typically combine the userߣs current query with contextual information from previous interactions, resulting in continuously expanding input prompts. This leads to a significant increase in time-to-first-token (TTFT), causing intolerable delays in web response times. To address this issue, we introduce TransformKV, which maximizes the reuse of the KV cache from previous conversations rather than recomputing, thereby reducing TTFT latency. TransformKV first identifies the specific locations that require transformation to maximize the reuse of the KV cache with minimal operations. It then efficiently transforms the KV cache for a subset of tokens by recomputing only the KV cache that impact semantics. Additionally, TransformKV further reduces TTFT latency by performing only QKV computations in certain layers while skipping other computations that contribute less to overall performance. Experimental results demonstrate that in multi-turn conversation tasks, TransformKV can reduce inference latency by up to 30%, achieving up to a 1.8× improvement in performance compared to similar approaches. Notably, as the context window size increases, the performance gains become even more pronounced.
Jiahang Zhou, Zhiyuan Fang, Yusheng Qin, Wuhui Chen, Tao Zhang 0096, Chuanfu Zhang, Zibin Zheng
IEEE Trans. Computers3