VLDB 2026 Research / reviewers in the wild / expert
Guangyue Peng
dblp:341/5588
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 47% Language models and text generation · 26% Learning paradigms · 26% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 77% Software testing · 23% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
hallucination |
1.0 | 1 | 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations · ACL (1) 2026 |
Machine learning › Learning paradigms
continual learning |
0.9 | 1 | 2025 | Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory · ACL (1) 2025 |
Natural language and speech › Language models and text generation › decoding
decoding strategy |
0.9 | 1 | 2025 | Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation · ACL (1) 2025 |
Program synthesis and code generation › code generation with language models
repository-level code generation |
0.9 | 1 | 2025 | FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation · ACL (1) 2025 |
Natural language and speech › Language models and text generation › text generation
factual text generation |
0.3 | 1 | 2025 | Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation · ACL (1) 2025 |
Machine learning › Learning paradigms › incremental learning
scalable continual learning |
0.3 | 1 | 2025 | Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
representation analysis · 1.0probing · 1.0transformer · 0.9mixture-of-neighbors induction memory · 0.9large language model · 0.9kNN-LM · 0.9dynamic focus decoding · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM HallucinationsabstractWen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang |
ACL (1) | 2 |
| 2025 | Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text GenerationabstractLarge Language Models (LLMs) are increasingly required to generate text that is both factually accurate and diverse across various openended applications.However, current stochastic decoding methods struggle to balance such objectives.We introduce Dynamic Focus Decoding (DFD), a novel plug-and-play stochastic approach that resolves this trade-off without requiring additional data, knowledge, or models.DFD adaptively adjusts the decoding focus based on distributional differences across layers, leveraging the modular and hierarchical nature of factual knowledge within LLMs.This dynamic adjustment improves factuality in knowledge-intensive decoding steps and promotes diversity in less knowledge-reliant steps.DFD can be easily integrated with existing decoding methods, enhancing both factuality and diversity with minimal computational overhead.Extensive experiments across seven datasets demonstrate that DFD significantly improves performance, providing a scalable and efficient solution for open-ended text generation. 1 Wen Luo 0001, Feifan Song 0001, Wei Li 0101, Guangyue Peng, Shaohang Wei, Houfeng Wang |
ACL (1) | 4 |
| 2025 | FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature ImplementationabstractWei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wei Li 0232, Xin Zhang 0099, Zhongxin Guo, Shaoguang Mao, Wen Luo 0001, Guangyue Peng, Yangyu Huang, Houfeng Wang, Scarlett Li |
ACL (1) | 6 |
| 2025 | Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction MemoryabstractSemiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the internal information flow of the parametric models, limiting scalability and efficiency. Based on recent interpretability theories of LMs, we reconceptualize the non-parametric memory represented by kNN-LM as a learnable Mixture-of-Neighbors Induction Memory (MoNIM), which synergizes the induction capabilities of attention heads with the memorization strength of feed-forward networks (FFN). By integrating into the model’s information flow, MoNIM functions as an FFN-like bypass layer within the Transformer architecture, enabling effective learning of new knowledge. Extensive experiments demonstrate that MoNIM is a retentive and scalable continual learner in both data- and model-wise, enhancing the scalability and continual learning performance of semiparametric LMs. Guangyue Peng, Tao Ge 0001, Wen Luo 0001, Wei Li 0101, Houfeng Wang |
ACL (1) | 1 |
| 2025 | Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error CorrectionabstractWei Li, Wen Luo, Guangyue Peng, Houfeng Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Wei Li 0101, Wen Luo 0001, Guangyue Peng, Houfeng Wang |
NAACL (Long Papers) | 3 |