Zhengkai Lin

dblp:368/8752 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 62% Efficient and distributed learning · 18% Knowledge representation and reasoning · 16%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Controlling Thinking Speed in Reasoning Models · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model
large reasoning model
0.912025
Controlling Thinking Speed in Reasoning Models · NeurIPS 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Controlling Thinking Speed in Reasoning Models · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge recall
0.812024
Delving into the Reversal Curse: How Far Can Large Language Models Generalize? · NeurIPS 2024
Natural language and speech › Language models and text generation › linguistic generalization
language model generalization
0.812024
Delving into the Reversal Curse: How Far Can Large Language Models Generalize? · NeurIPS 2024
Natural language and speech › Language models and text generation › knowledge editing
representation editing
0.312025
Controlling Thinking Speed in Reasoning Models · NeurIPS 2025
Natural language and speech › Language models and text generation
steering vectors
0.312025
Controlling Thinking Speed in Reasoning Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.212024
Delving into the Reversal Curse: How Far Can Large Language Models Generalize? · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

steering vectors · 0.9representation editing · 0.9difficulty estimation · 0.9probing · 0.8multiple-choice evaluation · 0.8
YearPublicationVenuePosition
2025 Controlling Thinking Speed in Reasoning Models
abstract
Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this work, we enable LRMs to approximate human intelligence through dynamic thinking speed adjustment, optimizing accuracy-efficiency trade-offs. Our approach addresses two key questions: (1) how to control thinking speed in LRMs, and (2) when to adjust it for optimal performance. For the first question, we identify the steering vector that governs slow-fast thinking transitions in LRMs' representation space. Using this vector, we achieve the first representation editing-based test-time scaling effect, outperforming existing prompt-based scaling methods. For the second question, we apply real-time difficulty estimation to signal reasoning segments of varying complexity. Combining these techniques, we propose the first reasoning strategy that enables fast processing of easy steps and deeper analysis for complex reasoning. Without any training or additional cost, our plug-and-play method yields an average +1.3\% accuracy with -8.6\% token usage across leading LRMs and advanced reasoning benchmarks. All of our algorithms are implemented based on vLLM and are expected to support broader applications and inspire future research.
Zhengkai Lin, Zhihang Fu, Ze Chen 0001, Chao Chen 0026, Liang Xie 0003, Wenxiao Wang 0001, Deng Cai 0001, Jieping Ye
NeurIPS1
2024 Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
abstract
While large language models (LLMs) showcase unprecedented capabilities, they also exhibit certain inherent limitations when facing seemingly trivial tasks. A prime example is the recently debated "reversal curse", which surfaces when models, having been trained on the fact "A is B", struggle to generalize this knowledge to infer that "B is A". In this paper, we examine the manifestation of the reversal curse across various tasks and delve into both the generalization abilities and the problem-solving mechanisms of LLMs. This investigation leads to a series of significant insights: (1) LLMs are able to generalize to "B is A" when both A and B are presented in the context as in the case of a multiple-choice question. (2) This generalization ability is highly correlated to the structure of the fact "A is B" in the training documents. For example, this generalization only applies to biographies structured in "[Name] is [Description]" but not to "[Description] is [Name]". (3) We propose and verify the hypothesis that LLMs possess an inherent bias in fact recalling during knowledge application, which explains and underscores the importance of the document structure to successful learning. (4) The negative impact of this bias on the downstream performance of LLMs can hardly be mitigated through training alone. Based on these intriguing findings, our work not only presents a novel perspective for interpreting LLMs' generalization abilities from their intrinsic working mechanism but also provides new insights for the development of more effective learning methods for LLMs.
Zhengkai Lin, Zhihang Fu, Kai Liu 0023, Liang Xie 0003, Binbin Lin 0001, Wenxiao Wang 0001, Deng Cai 0001, Jieping Ye
NeurIPS1