Zhenyu He 0012

dblp:355/4626 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0005-7001-0591ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 30% Deep learning architectures and training · 27% Efficient and distributed learning · 27%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Theoretical computer science
1 paper
Computational complexity · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
positional encoding
1.022025
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation · ICML 2024
Let the Code LLM Edit Itself When You Edit the Code · ICLR 2025
Machine learning › Learning paradigms
lifelong learning
1.012026
Ted-Tok: Maintaining an Evolving Vocabulary for Lifelong Learning · ACL (1) 2026
Natural language and speech › Language models and text generation
tokenization
1.012026
Ted-Tok: Maintaining an Evolving Vocabulary for Lifelong Learning · ACL (1) 2026
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Let the Code LLM Edit Itself When You Edit the Code · ICLR 2025
Machine learning › Efficient and distributed learning › inference efficiency
KV cache reuse
0.912025
Let the Code LLM Edit Itself When You Edit the Code · ICLR 2025
Program synthesis and code generation
code generation with language models
0.912025
Let the Code LLM Edit Itself When You Edit the Code · ICLR 2025
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.812024
Do Efficient Transformers Really Save Computation? · ICML 2024
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation
0.812024
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation · ICML 2024
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.212024
Do Efficient Transformers Really Save Computation? · ICML 2024

Methods — techniques the papers use, named apart from their topics

rotary positional encoding · 1.7KV cache · 1.7dynamic programming modeling · 1.5time-weighted frequency estimation · 1.0relative positional encoding · 0.8absolute positional encoding · 0.8
YearPublicationVenuePosition
2026 Ted-Tok: Maintaining an Evolving Vocabulary for Lifelong Learning
abstract
Lifelong learning investigates how models adapt when exposed to a potentially infinite stream of data.Most conventional approaches focus on updating model parameters (i.e., the neural network weights) as the underlying data distribution evolves over time.However, in natural language processing, model parameters are not the only components that matter.The tokenizer, a foundational part of the system, is usually assumed to remain fixed in lifelong learning scenarios.In this work, we challenge the validity of this assumption: as language evolves, a static tokenizer fragments newly emerging lexical items, reducing compression efficiency and consequently degrading the model performance.We introduce the Temporal Drift Tokenizer (Ted-Tok), which maintains an evolving vocabulary that adapts to emerging linguistic patterns over time.This adaptivity is driven by time-weighted frequency estimators that smooth short-term fluctuations to capture persistent linguistic trends, and a principled addition-deletion strategy targeting sink tokens.Across multiple domains, Ted-Tok consistently improves compression and task performance, with gains increasing under stronger drift, underscoring the role of tokenizer adaptivity in lifelong learning.
Jiameng Huang, Zhi Zhang 0005, Zhenyu He 0012, Di He 0001
ACL (1)3
2025 Let the Code LLM Edit Itself When You Edit the Code
abstract
In this work, we investigate a typical scenario in code generation where a developer edits existing code in real time and requests a code assistant, e.g., a large language model, to re-predict the next token or next line on the fly. Naively, the LLM needs to re-encode the entire KV cache to provide an accurate prediction. However, this process is computationally expensive, especially when the sequence length is long. Simply encoding the edited subsequence and integrating it to the original KV cache meets the temporal confusion problem, leading to significantly worse performance. We address this efficiency and accuracy trade-off by introducing $\underline{\textbf{P}\text{ositional}\ \textbf{I}\text{ntegrity}\ \textbf{E}\text{ncoding}}$ (PIE). Building upon the rotary positional encoding, PIE first removes the rotary matrices in the Key cache that introduce temporal confusion and then reapplies the correct rotary matrices. This process ensures that positional relationships between tokens are correct and requires only a single round of matrix multiplication. We validate the effectiveness of PIE through extensive experiments on the RepoBench-C-8k dataset, utilizing DeepSeek-Coder models with 1.3B, 6.7B, and 33B parameters. Our evaluation includes three real-world coding tasks: code insertion, code deletion, and multi-place code editing. Results demonstrate that PIE reduces computational overhead by over 85% compared to the standard full recomputation approach across all model sizes and tasks while well approximating the model performance.
Zhenyu He 0012, Jun Zhang 0003, Shengjie Luo, Jingjing Xu 0001, Zhi Zhang 0005, Di He 0001
ICLR1
2024 Exploiting Pre-trained Models for Drug Target Affinity Prediction with Nearest Neighbors
abstract
Drug-Target binding Affinity (DTA) prediction is essential for drug discovery. Despite the application of deep learning methods to DTA prediction, the achieved accuracy remain suboptimal. In this work, inspired by the recent success of retrieval methods, we propose kNN-DTA, a non-parametric embedding-based retrieval method adopted on a pre-trained DTA prediction model, which can extend the power of the DTA model with no or negligible cost. Different from existing methods, we introduce two neighbor aggregation ways from both embedding space and label space that are integrated into a unified framework. Specifically, we propose a label aggregation with pair-wise retrieval and a representation aggregation with point-wise retrieval of the nearest neighbors. This method executes in the inference phase and can efficiently boost the DTA prediction performance with no training cost. In addition, we propose an extension, Ada-kNN-DTA, an instance-wise and adaptive aggregation with lightweight learning. Results on four benchmark datasets show that kNN-DTA brings significant improvements, outperforming previous state-of-the-art (SOTA) results, e.g, on BindingDB IC50 and Ki testbeds, kNN-DTA obtains new records of RMSE 0.684 and 0.750 . The extended Ada-kNN-DTA further improves the performance to be 0.675 and 0.735 RMSE. These results strongly prove the effectiveness of our method. Results in other settings and comprehensive studies/analyses also show the great potential of our kNN-DTA approach.
Qizhi Pei, Lijun Wu 0003, Zhenyu He 0012, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Rui Yan 0001
CIKM3
2024 Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
abstract
In this work, we leverage the intrinsic segmentation of language sequences and design a new positional encoding method called Bilevel Positional Encoding (BiPE). For each position, our BiPE blends an intra-segment encoding and an inter-segment encoding. The intra-segment encoding identifies the locations within a segment and helps the model capture the semantic information therein via absolute positional encoding. The inter-segment encoding specifies the segment index, models the relationships between segments, and aims to improve extrapolation capabilities via relative positional encoding. Theoretical analysis shows this disentanglement of positional information makes learning more effective. The empirical results also show that our BiPE has superior length extrapolation capabilities across a wide range of tasks in diverse text modalities.
Zhenyu He 0012, Guhao Feng, Shengjie Luo, Liwei Wang 0001, Jingjing Xu 0001, Zhi Zhang 0005, Hongxia Yang, Di He 0001
ICML1
2024 Do Efficient Transformers Really Save Computation?
abstract
As transformer-based language models are trained on increasingly large datasets and with vast numbers of parameters, finding more efficient alternatives to the standard Transformer has become very valuable. While many efficient Transformers and Transformer alternatives have been proposed, none provide theoretical guarantees that they are a suitable replacement for the standard Transformer. This makes it challenging to identify when to use a specific model and what directions to prioritize for further investigation. In this paper, we aim to understand the capabilities and limitations of efficient Transformers, specifically the Sparse Transformer and the Linear Transformer. We focus on their reasoning capability as exhibited by Chain-of-Thought (CoT) prompts and follow previous works to model them as Dynamic Programming (DP) problems. Our results show that while these models are expressive enough to solve general DP tasks, contrary to expectations, they require a model size that scales with the problem size. Nonetheless, we identify a class of DP problems for which these models can be more efficient than the standard Transformer. We confirm our theoretical results through experiments on representative DP tasks, adding to the understanding of efficient Transformers’ practical strengths and weaknesses.
Jan Ackermann, Zhenyu He 0012, Guhao Feng, Bohang Zhang, Yunzhen Feng, Qiwei Ye, Di He 0001, Liwei Wang 0001
ICML3
2024 REST: Retrieval-Based Speculative Decoding
abstract
Zhenyu He, Zexuan Zhong, Tianle Cai, Jason Lee, Di He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhenyu He 0012, Zexuan Zhong, Tianle Cai, Jason D. Lee, Di He 0001
NAACL-HLT1