Yongcheng Zeng

dblp:375/7557 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 53% Reinforcement learning · 39% Multi-agent systems · 8%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model reasoning
1.012026
Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026
Natural language and speech › Language models and text generation › complex reasoning
metacognitive reasoning
1.012026
Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models
1.012026
Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026
Machine learning › Reinforcement learning › reinforcement learning from human feedback › learning from human feedback
RLHF
1.012026
Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.812024
Token-level Direct Preference Optimization · ICML 2024
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.812024
Token-level Direct Preference Optimization · ICML 2024
Natural language and speech › Language models and text generation
preference optimization
0.812024
Token-level Direct Preference Optimization · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent environments
real-time strategy games
0.812024
Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Token-level Direct Preference Optimization · ICML 2024
Machine learning › Reinforcement learning
strategic decision-making
0.812024
Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024
Natural language and speech › Language models and text generation
token-level optimization
0.812024
Token-level Direct Preference Optimization · ICML 2024
Machine learning › Reinforcement learning
agent evaluation
0.212024
Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.0hierarchical decomposition · 1.0direct preference optimization · 0.8chain-of-thought · 0.8chain of summarization · 0.8bradley-terry model · 0.8KL divergence · 0.8
YearPublicationVenuePosition
2026 Learning from Cognition: Enhancing RL Efficiency for LLM Reasoning via Hierarchical Metacognitive Decomposition and Refinement
abstract
Zexu Sun, Yongcheng Zeng, Erxue Min, Heyang Gao, Bokai Ji, Dugang Liu, Xing Tang, Xiuqiang He, Xu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zexu Sun, Yongcheng Zeng, Erxue Min, Heyang Gao, Bokai Ji, Dugang Liu, Xing Tang 0007, Xiuqiang He 0001, Xu Chen 0017
ACL (1)2
2024 Token-level Direct Preference Optimization
abstract
Fine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the generation of these responses occurs in a token level, following a sequential, auto-regressive fashion. In this paper, we introduce Token-level Direct Preference Optimization (TDPO), a novel approach to align LLMs with human preferences by optimizing policy at the token level. Unlike previous methods, which face challenges in divergence efficiency, TDPO integrates forward KL divergence constraints for each token, improving alignment and diversity. Utilizing the Bradley-Terry model for a token-based reward system, our method enhances the regulation of KL divergence, while preserving simplicity without the need for explicit reward modeling. Experimental results across various text tasks demonstrate TDPO’s superior performance in balancing alignment with generation diversity. Notably, fine-tuning with TDPO strikes a better balance than DPO in the controlled sentiment generation and single-turn dialogue datasets, and significantly improves the quality of generated responses compared to both DPO and PPO-based RLHF methods.
Yongcheng Zeng, Weiyu Ma, Ning Yang 0005, Haifeng Zhang 0002, Jun Wang 0012
ICML1
2024 Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
abstract
With the continued advancement of Large Language Models (LLMs) Agents in reasoning, planning, and decision-making, benchmarks have become crucial in evaluating these skills. However, there is a notable gap in benchmarks for real-time strategic decision-making. StarCraft II (SC2), with its complex and dynamic nature, serves as an ideal setting for such evaluations. To this end, we have developed TextStarCraft II, a specialized environment for assessing LLMs in real-time strategic scenarios within SC2. Addressing the limitations of traditional Chain of Thought (CoT) methods, we introduce the Chain of Summarization (CoS) method, enhancing LLMs' capabilities in rapid and effective decision-making. Our key experiments included: 1. LLM Evaluation: Tested 10 LLMs in TextStarCraft II, most of them defeating LV5 build-in AI, showcasing effective strategy skills. 2. Commercial Model Knowledge: Evaluated four commercial models on SC2 knowledge; GPT-4 ranked highest by Grandmaster-level experts. 3. Human-AI Matches: Experimental results showed that fine-tuned LLMs performed on par with Gold-level players in real-time matches, demonstrating comparable strategic abilities. All code and data from this study have been made pulicly available at https://github.com/histmeisah/Large-Language-Models-play-StarCraftII
Weiyu Ma, Qirui Mi, Yongcheng Zeng, Runji Lin, Yuqiao Wu, Jun Wang 0012, Haifeng Zhang 0002
NeurIPS3