Guangyue Peng

dblp:341/5588 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 47% Language models and text generation · 26% Learning paradigms · 26%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 77% Software testing · 23%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
hallucination
1.012026
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations · ACL (1) 2026
Machine learning › Learning paradigms
continual learning
0.912025
Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory · ACL (1) 2025
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.912025
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation · ACL (1) 2025
Program synthesis and code generation › code generation with language models
repository-level code generation
0.912025
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation · ACL (1) 2025
Natural language and speech › Language models and text generation › text generation
factual text generation
0.312025
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation · ACL (1) 2025
Machine learning › Learning paradigms › incremental learning
scalable continual learning
0.312025
Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

representation analysis · 1.0probing · 1.0transformer · 0.9mixture-of-neighbors induction memory · 0.9large language model · 0.9kNN-LM · 0.9dynamic focus decoding · 0.9
YearPublicationVenuePosition
2026 Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
abstract
Wen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang
ACL (1)2
2025 Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
abstract
Large Language Models (LLMs) are increasingly required to generate text that is both factually accurate and diverse across various openended applications.However, current stochastic decoding methods struggle to balance such objectives.We introduce Dynamic Focus Decoding (DFD), a novel plug-and-play stochastic approach that resolves this trade-off without requiring additional data, knowledge, or models.DFD adaptively adjusts the decoding focus based on distributional differences across layers, leveraging the modular and hierarchical nature of factual knowledge within LLMs.This dynamic adjustment improves factuality in knowledge-intensive decoding steps and promotes diversity in less knowledge-reliant steps.DFD can be easily integrated with existing decoding methods, enhancing both factuality and diversity with minimal computational overhead.Extensive experiments across seven datasets demonstrate that DFD significantly improves performance, providing a scalable and efficient solution for open-ended text generation. 1
Wen Luo 0001, Feifan Song 0001, Wei Li 0101, Guangyue Peng, Shaohang Wei, Houfeng Wang
ACL (1)4
2025 FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
abstract
Wei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0232, Xin Zhang 0099, Zhongxin Guo, Shaoguang Mao, Wen Luo 0001, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li
ACL (1)6
2025 Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory
abstract
Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the internal information flow of the parametric models, limiting scalability and efficiency. Based on recent interpretability theories of LMs, we reconceptualize the non-parametric memory represented by kNN-LM as a learnable Mixture-of-Neighbors Induction Memory (MoNIM), which synergizes the induction capabilities of attention heads with the memorization strength of feed-forward networks (FFN). By integrating into the model’s information flow, MoNIM functions as an FFN-like bypass layer within the Transformer architecture, enabling effective learning of new knowledge. Extensive experiments demonstrate that MoNIM is a retentive and scalable continual learner in both data- and model-wise, enhancing the scalability and continual learning performance of semiparametric LMs.
Guangyue Peng, Tao Ge 0001, Wen Luo 0001, Wei Li 0101, Houfeng Wang
ACL (1)1
2025 Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction
abstract
Wei Li, Wen Luo, Guangyue Peng, Houfeng Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Wei Li 0101, Wen Luo 0001, Guangyue Peng, Houfeng Wang
NAACL (Long Papers)3