Yipeng Kang

dblp:267/2079 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0000-2743-4747ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 40% Multi-agent systems · 37% Language models and text generation · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.832024
Multi-agent policy transfer via task relationship modeling · Sci. China Inf. Sci. 2024
Non-Linear Coordination Graphs · NeurIPS 2022
Incorporating Pragmatic Reasoning Communication into Emergent Language · NeurIPS 2020
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation
1.012026
Why Are We Moral? An LLM-based Agent Simulation Approach to the Study of Moral Evolution · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice · ACL (1) 2026
Computational social science and digital humanities
legal informatics
1.012026
JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice · ACL (1) 2026
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
0.812024
Multi-agent policy transfer via task relationship modeling · Sci. China Inf. Sci. 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
coordination graph
0.612022
Non-Linear Coordination Graphs · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning › value-based multi-agent reinforcement learning
value decomposition
0.612022
Non-Linear Coordination Graphs · NeurIPS 2022
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.412020
Incorporating Pragmatic Reasoning Communication into Emergent Language · NeurIPS 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › normative reasoning
moral reasoning
0.312026
Why Are We Moral? An LLM-based Agent Simulation Approach to the Study of Moral Evolution · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
0.312025
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia · NeurIPS 2025
Machine learning › Learning paradigms › multi-task learning
task relationship modeling
0.212024
Multi-agent policy transfer via task relationship modeling · Sci. China Inf. Sci. 2024

Methods — techniques the papers use, named apart from their topics

benchmarking · 2.0large language model agents · 1.0zero-shot evaluation · 0.9natural language multi-agent simulation · 0.9Leaky ReLU · 0.6DCOP · 0.6reinforcement learning · 0.4pragmatic reasoning · 0.4
YearPublicationVenuePosition
2026 JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice
abstract
Ziang Chen, Guannan Li, Fanlin Ji, Yipeng Kang, Jiaqi Li, Muhan Zhang, Yangtao Zhang, Li Tianjiao, Jiannan Wang, Xin Guo, Song-Chun Zhu, Bin Ling. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fanlin Ji, Yipeng Kang, Jiaqi Li 0021, Muhan Zhang, Yangtao Zhang, Li Tianjiao, Song-Chun Zhu, Bin Ling
ACL (1)4
2026 Why Are We Moral? An LLM-based Agent Simulation Approach to the Study of Moral Evolution
abstract
Zhou Ziheng, Huacong Tang, Mingjie Bi, Wanying He, Fang Sun, Yizhou Sun, Ying Nian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhou Ziheng, Huacong Tang, Mingjie Bi, Wanying He, Yizhou Sun, Ying Nian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong
ACL (1)9
2025 IBGP: Imperfect Byzantine Generals Problem for Zero-Shot Robustness in Communicative Multi-Agent Systems
Yihuan Mao, Yipeng Kang, Peilun Li, Wei Xu 0005, Chongjie Zhang
AAMAS2
2025 Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
abstract
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement.
Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Beining Zhang, Jim Dilkes, Akash Kundu, Emanuel Tewolde, Jebish Purbey, Ram Mohan Rao Kadiyala, Siddhant Gupta, Aliaksei Korshuk, Buyantuev Alexander, Ilya Makarov, Rolando Fernandez, Zhihan Wang, Caroline Wang, Jiaxun Cui, Lingyun Xiao, Yoonchang Sung, Muhammad Arrasy Rahman, Peter Stone 0001, Yipeng Kang, Hyeonggeun Yun, Ananya, Taehun Cha, Elizaveta Tennant, Olivia Macmillan-Scott, Marta Segura, Diana Riazi, Fuyang Cui, Sriram Ganapathi, Toryn Q. Klassen, Nico Schiavone, Mogtaba Alim, Sheila A. McIlraith, Manuel Ríos, Oswaldo Peña, Manuela Chacon-Chamorro, Rubén Manrique, Luis Felipe Giraldo, Nicanor Quijano, Fangwei Zhong, Wenming Tu, Zhaowei Zhang 0001, Zixia Jia, Zilong Zheng, Chichen Lin, Weijian Fan, Chenao Liu, Sneheel Sarangi, Shuqing Shi, Yali Du 0001, Avinaash Anand Kulandaivel, Yang Liu 0266, Ruiyang Wu 0007, Chetan Talele, Sunjia Lu, Gema Parreno, Shamika Dhuri, Bain McHale, Tim Baarslag, Dylan Hadfield-Menell, Natasha Jaques, José Hernández-Orallo, Joel Z. Leibo
NeurIPS35
2024 Multi-agent policy transfer via task relationship modeling
Rongjun Qin, Feng Chen 0042, Tonghan Wang 0001, Lei Yuan 0005, Xiaoran Wu, Yipeng Kang, Zongzhang Zhang, Chongjie Zhang, Yang Yu 0001
Sci. China Inf. Sci.6
2022 Non-Linear Coordination Graphs
abstract
Value decomposition multi-agent reinforcement learning methods learn the global value function as a mixing of each agent's individual utility functions. Coordination graphs (CGs) represent a higher-order decomposition by incorporating pairwise payoff functions and thus is supposed to have a more powerful representational capacity. However, CGs decompose the global value function linearly over local value functions, severely limiting the complexity of the value function class that can be represented. In this paper, we propose the first non-linear coordination graph by extending CG value decomposition beyond the linear case. One major challenge is to conduct greedy action selections in this new function class to which commonly adopted DCOP algorithms are no longer applicable. We study how to solve this problem when mixing networks with LeakyReLU activation are used. An enumeration method with a global optimality guarantee is proposed and motivates an efficient iterative optimization method with a local optimality guarantee. We find that our method can achieve superior performance on challenging multi-agent coordination tasks like MACO.
Yipeng Kang, Tonghan Wang 0001, Qianlan Yang, Xiaoran Wu, Chongjie Zhang
NeurIPS1
2020 Incorporating Pragmatic Reasoning Communication into Emergent Language
abstract
Emergentism and pragmatics are two research fields that study the dynamics of linguistic communication along quite different timescales and intelligence levels. From the perspective of multi-agent reinforcement learning, they correspond to stochastic games with reinforcement training and stage games with opponent awareness, respectively. Given that their combination has been explored in linguistics, in this work, we combine computational models of short-term mutual reasoning-based pragmatics with long-term language emergentism. We explore this for agent communication in two settings, referential games and Starcraft II, assessing the relative merits of different kinds of mutual reasoning pragmatics models both empirically and theoretically. Our results shed light on their importance for making inroads towards getting more natural, accurate, robust, fine-grained, and succinct utterances.
Yipeng Kang, Tonghan Wang 0001, Gerard de Melo
NeurIPS1