VLDB 2026 Research / reviewers in the wild / expert
Yongcheng Zeng
dblp:375/7557
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 53% Reinforcement learning · 39% Multi-agent systems · 8% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model reasoning |
1.0 | 1 | 2026 | Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026 |
Natural language and speech › Language models and text generation › complex reasoning
metacognitive reasoning |
1.0 | 1 | 2026 | Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
1.0 | 1 | 2026 | Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback › learning from human feedback
RLHF |
1.0 | 1 | 2026 | Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Token-level Direct Preference Optimization · ICML 2024 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.8 | 1 | 2024 | Token-level Direct Preference Optimization · ICML 2024 |
Natural language and speech › Language models and text generation
preference optimization |
0.8 | 1 | 2024 | Token-level Direct Preference Optimization · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent environments
real-time strategy games |
0.8 | 1 | 2024 | Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | Token-level Direct Preference Optimization · ICML 2024 |
Machine learning › Reinforcement learning
strategic decision-making |
0.8 | 1 | 2024 | Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024 |
Natural language and speech › Language models and text generation
token-level optimization |
0.8 | 1 | 2024 | Token-level Direct Preference Optimization · ICML 2024 |
Machine learning › Reinforcement learning
agent evaluation |
0.2 | 1 | 2024 | Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.0hierarchical decomposition · 1.0direct preference optimization · 0.8chain-of-thought · 0.8chain of summarization · 0.8bradley-terry model · 0.8KL divergence · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and RefinementabstractZexu Sun, Yongcheng Zeng, Erxue Min, Heyang Gao, Bokai Ji, Dugang Liu, Xing Tang, Xiuqiang He, Xu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zexu Sun, Yongcheng Zeng, Erxue Min, Heyang Gao, Bokai Ji, Dugang Liu, Xing Tang 0007, Xiuqiang He 0001, Xu Chen 0017 |
ACL (1) | 2 |
| 2024 | Token-level Direct Preference OptimizationabstractFine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the generation of these responses occurs in a token level, following a sequential, auto-regressive fashion. In this paper, we introduce Token-level Direct Preference Optimization (TDPO), a novel approach to align LLMs with human preferences by optimizing policy at the token level. Unlike previous methods, which face challenges in divergence efficiency, TDPO integrates forward KL divergence constraints for each token, improving alignment and diversity. Utilizing the Bradley-Terry model for a token-based reward system, our method enhances the regulation of KL divergence, while preserving simplicity without the need for explicit reward modeling. Experimental results across various text tasks demonstrate TDPO’s superior performance in balancing alignment with generation diversity. Notably, fine-tuning with TDPO strikes a better balance than DPO in the controlled sentiment generation and single-turn dialogue datasets, and significantly improves the quality of generated responses compared to both DPO and PPO-based RLHF methods. Yongcheng Zeng, Weiyu Ma, Ning Yang 0005, Haifeng Zhang 0002, Jun Wang 0012 |
ICML | 1 |
| 2024 | Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachabstractWith the continued advancement of Large Language Models (LLMs) Agents in reasoning, planning, and decision-making, benchmarks have become crucial in evaluating these skills. However, there is a notable gap in benchmarks for real-time strategic decision-making. StarCraft II (SC2), with its complex and dynamic nature, serves as an ideal setting for such evaluations. To this end, we have developed TextStarCraft II, a specialized environment for assessing LLMs in real-time strategic scenarios within SC2. Addressing the limitations of traditional Chain of Thought (CoT) methods, we introduce the Chain of Summarization (CoS) method, enhancing LLMs' capabilities in rapid and effective decision-making. Our key experiments included:
1. LLM Evaluation: Tested 10 LLMs in TextStarCraft II, most of them defeating LV5 build-in AI, showcasing effective strategy skills.
2. Commercial Model Knowledge: Evaluated four commercial models on SC2 knowledge; GPT-4 ranked highest by Grandmaster-level experts.
3. Human-AI Matches: Experimental results showed that fine-tuned LLMs performed on par with Gold-level players in real-time matches, demonstrating comparable strategic abilities.
All code and data from this
study have been made pulicly available at https://github.com/histmeisah/Large-Language-Models-play-StarCraftII Weiyu Ma, Qirui Mi, Yongcheng Zeng, Runji Lin, Yuqiao Wu, Jun Wang 0012, Haifeng Zhang 0002 |
NeurIPS | 3 |