EDBT 2026 Demo / reviewers in the wild / expert
Shao Zhang
dblp:57/1330
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-0111-0776ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 52% Multi-agent systems · 29% Language models and text generation · 9% | |
| Human-computer interaction and pervasive computing
3 papers |
Learning and educational technologies · 43% Human-AI interaction · 43% Collaborative and social computing · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
2.2 | 3 | 2024 | ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination · NeurIPS 2024 Aligning Individual and Collective Objectives in Multi-Agent Cooperation · NeurIPS 2024 Cooperative Open-ended Learning Framework for Zero-Shot Coordination · ICML 2023 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
zero-shot coordination |
1.4 | 2 | 2024 | ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination · NeurIPS 2024 Cooperative Open-ended Learning Framework for Zero-Shot Coordination · ICML 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration |
0.9 | 1 | 2025 | Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration · ACL (1) 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.9 | 1 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 |
Machine learning › Reinforcement learning
reward learning |
0.9 | 1 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration |
0.8 | 1 | 2024 | ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot Coordination · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
narrative question answering |
0.8 | 1 | 2024 | StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning · EMNLP 2024 |
Medical and health informatics
clinical decision support |
0.8 | 1 | 2024 | Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis Diagnosis · CHI 2024 |
Human-AI interaction › explainable AI › uncertainty communication
uncertainty visualization |
0.8 | 1 | 2024 | Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis Diagnosis · CHI 2024 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
cooperative game |
0.7 | 1 | 2023 | Cooperative Open-ended Learning Framework for Zero-Shot Coordination · ICML 2023 |
Machine learning › Reinforcement learning
value function estimation |
0.3 | 1 | 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization · NeurIPS 2025 |
Collaborative and social computing › computer-supported cooperative work › collaborative applications
real-time collaboration |
0.3 | 1 | 2025 | Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration · ACL (1) 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
predictive uncertainty |
0.2 | 1 | 2024 | Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis Diagnosis · CHI 2024 |
Algorithmic game theory and mechanism design
cooperative game theory |
0.2 | 1 | 2023 | Cooperative Open-ended Learning Framework for Zero-Shot Coordination · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
heuristic evaluation · 2.3formative study · 2.3dual process theory · 1.7graph theory · 1.3preference margin regularization · 0.9policy regularization · 0.9gradient adjustment · 0.8differentiable game · 0.8best-response proximity · 0.8best-response diversity · 0.8open-ended learning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI CollaborationabstractShao Zhang, Xihuai Wang, Wenhao Zhang, Chaoran Li, Junru Song, Tingyu Li, Lin Qiu, Xuezhi Cao, Xunliang Cai, Wen Yao, Weinan Zhang, Xinbing Wang, Ying Wen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shao Zhang, Xihuai Wang, Junru Song, Xuezhi Cao, Weinan Zhang 0001, Xinbing Wang, Ying Wen 0001 |
ACL (1) | 1 |
| 2025 | PMAT: Optimizing Action Generation Order in Multi-Agent Reinforcement Learning
Muning Wen, Xihuai Wang, Shao Zhang, Yiwei Shi, Minne Li, Minglong Li, Ying Wen 0001 |
AAMAS | 4 |
| 2025 | STAR: Efficient Preference-based Reinforcement Learning via Dual RegularizationabstractPreference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the reward model overfits to the limited feedback, leading to poor generalization to unseen samples, and (2) the agent exploits the learned reward model, exacerbating overestimation of action values in temporal difference (TD) learning. To address these issues, we propose STAR, an efficient PbRL method that integrates preference margin regularization and policy regularization. Preference margin regularization mitigates overfitting by introducing a bounded margin in reward optimization, preventing excessive bias toward specific feedback. Policy regularization bootstraps a conservative estimate $\widehat{Q}$ from well-supported state-action pairs in the replay memory, reducing overestimation during policy learning. Experimental results show that STAR improves feedback efficiency, achieving 34.8\% higher performance in online settings and 29.7\% in offline settings compared to state-of-the-art methods. Ablation studies confirm that STAR facilitates more robust reward and value function learning. The videos of this project are released at https://sites.google.com/view/pbrl-star. Fengshuo Bai, Rui Zhao 0001, Hongming Zhang 0003, Sijia Cui, Shao Zhang, Bo Xu 0002, Ying Wen 0001, Yaodong Yang 0001 |
NeurIPS | 5 |
| 2024 | Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis DiagnosisabstractToday's AI systems for medical decision support often succeed on benchmark datasets in research papers but fail in real-world deployment. This work focuses on the decision making of sepsis, an acute life-threatening systematic infection that requires an early diagnosis with high uncertainty from the clinician. Our aim is to explore the design requirements for AI systems that can support clinical experts in making better decisions for the early diagnosis of sepsis. The study begins with a formative study investigating why clinical experts abandon an existing AI-powered Sepsis predictive module in their electrical health record (EHR) system. We argue that a human-centered AI system needs to support human experts in the intermediate stages of a medical decision-making process (e.g., generating hypotheses or gathering data), instead of focusing only on the final decision. Therefore, we build SepsisLab based on a state-of-the-art AI algorithm and extend it to predict the future projection of sepsis development, visualize the prediction uncertainty, and propose actionable suggestions (i.e., which additional laboratory tests can be collected) to reduce such uncertainty. Through heuristic evaluation with six clinicians using our prototype system, we demonstrate that SepsisLab enables a promising human-AI collaboration paradigm for the future of AI-assisted sepsis diagnosis and other high-stakes medical decision making. Shao Zhang, Xuhai Xu, Changchang Yin, Yuxuan Lu 0003, Bingsheng Yao, Melanie Tory, Lace M. K. Padilla, Jeffrey M. Caterino, Ping Zhang 0016, Dakuo Wang |
CHI | 1 |
| 2024 | StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based LearningabstractJiaju Chen, Yuxuan Lu, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Ying Xu, Yunyao Li, Qianwen Wang, Dakuo Wang, Yuling Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiaju Chen, Yuxuan Lu 0003, Shao Zhang, Bingsheng Yao, Yuanzhe Dong, Yunyao Li 0001, Dakuo Wang, Yuling Sun |
EMNLP | 3 |
| 2024 | Aligning Individual and Collective Objectives in Multi-Agent CooperationabstractAmong the research topics in multi-agent learning, mixed-motive cooperation is one of the most prominent challenges, primarily due to the mismatch between individual and collective goals. The cutting-edge research is focused on incorporating domain knowledge into rewards and introducing additional mechanisms to incentivize cooperation. However, these approaches often face shortcomings such as the effort on manual design and the absence of theoretical groundings. To close this gap, we model the mixed-motive game as a differentiable game for the ease of illuminating the learning dynamics towards cooperation. More detailed, we introduce a novel optimization method named \textbf{\textit{A}}ltruistic \textbf{\textit{G}}radient \textbf{\textit{A}}djustment (\textbf{\textit{AgA}}) that employs gradient adjustments to progressively align individual and collective objectives. Furthermore, we theoretically prove that AgA effectively attracts gradients to stable fixed points of the collective objective while considering individual interests, and we validate these claims with empirical evidence. We evaluate the effectiveness of our algorithm AgA through benchmark environments for testing mixed-motive collaboration with small-scale agents such as the two-player public good game and the sequential social dilemma games, Cleanup and Harvest, as well as our self-developed large-scale environment in the game StarCraft II. Yang Li 0116, Shao Zhang, Yali Du 0001, Ying Wen 0001, Wei Pan 0004 |
NeurIPS | 4 |
| 2024 | ZSC-Eval: An Evaluation Toolkit and Benchmark for Multi-agent Zero-shot CoordinationabstractZero-shot coordination (ZSC) is a new cooperative multi-agent reinforcement learning (MARL) challenge that aims to train an ego agent to work with diverse, unseen partners during deployment. The significant difference between the deployment-time partners' distribution and the training partners' distribution determined by the training algorithm makes ZSC a unique out-of-distribution (OOD) generalization challenge. The potential distribution gap between evaluation and deployment-time partners leads to inadequate evaluation, which is exacerbated by the lack of appropriate evaluation metrics. In this paper, we present ZSC-Eval, the first evaluation toolkit and benchmark for ZSC algorithms. ZSC-Eval consists of: 1) Generation of evaluation partner candidates through behavior-preferring rewards to approximate deployment-time partners' distribution; 2) Selection of evaluation partners by Best-Response Diversity (BR-Div); 3) Measurement of generalization performance with various evaluation partners via the Best-Response Proximity (BR-Prox) metric. We use ZSC-Eval to benchmark ZSC algorithms in Overcooked and Google Research Football environments and get novel empirical findings. We also conduct a human experiment of current ZSC algorithms to verify the ZSC-Eval's consistency with human evaluation. ZSC-Eval is now available at https://github.com/sjtu-marl/ZSC-Eval. Xihuai Wang, Shao Zhang, Jingxiao Chen, Ying Wen 0001, Weinan Zhang 0001 |
NeurIPS | 2 |
| 2024 | Tackling Cooperative Incompatibility for Zero-Shot Human-AI CoordinationabstractSecuring coordination between AI agent and teammates (human players or AI agents) in contexts involving unfamiliar humans continues to pose a significant challenge in Zero-Shot Coordination. The issue of cooperative incompatibility becomes particularly prominent when an AI agent is unsuccessful in synchronizing with certain previously unknown partners. Traditional algorithms have aimed to collaborate with partners by optimizing fixed objectives within a population, fostering diversity in strategies and behaviors. However, these techniques may lead to learning loss and an inability to cooperate with specific strategies within the population, a phenomenon named cooperative incompatibility in learning. In order to solve cooperative incompatibility in learning and effectively address the problem in the context of ZSC, we introduce the Cooperative Open-ended LEarning (COLE) framework, which formulates open-ended objectives in cooperative games with two players using perspectives of graph theory to evaluate and pinpoint the cooperative capacity of each strategy. We present two practical algorithms, specifically COLESV and COLER, which incorporate insights from game theory and graph theory. We also show that COLE could effectively overcome the cooperative incompatibility from theoretical and empirical analysis. Subsequently, we created an online Overcooked human-AI experiment platform, the COLE platform, which enables easy customization of questionnaires, model weights, and other aspects. Utilizing the COLE platform, we enlist 130 participants for human experiments. Our findings reveal a preference for our approach over state-of-the-art methods using a variety of subjective metrics. Moreover, objective experimental outcomes in the Overcooked game environment indicate that our method surpasses existing ones when coordinating with previously unencountered AI agents and the human proxy model. Our code and demo are publicly available at https://sites.google.com/view/cole-2023. Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004 |
J. Artif. Intell. Res. | 2 |
| 2023 | Cooperative Open-ended Learning Framework for Zero-Shot CoordinationabstractZero-shot coordination in cooperative artificial intelligence (AI) remains a significant challenge, which means effectively coordinating with a wide range of unseen partners. Previous algorithms have attempted to address this challenge by optimizing fixed objectives within a population to improve strategy or behaviour diversity. However, these approaches can result in a loss of learning and an inability to cooperate with certain strategies within the population, known as cooperative incompatibility. To address this issue, we propose the Cooperative Open-ended LEarning (COLE) framework, which constructs open-ended objectives in cooperative games with two players from the perspective of graph theory to assess and identify the cooperative ability of each strategy. We further specify the framework and propose a practical algorithm that leverages knowledge from game theory and graph theory. Furthermore, an analysis of the learning process of the algorithm shows that it can efficiently overcome cooperative incompatibility. The experimental results in the Overcooked game environment demonstrate that our method outperforms current state-of-the-art methods when coordinating with different-level partners. Our demo is available at https://sites.google.com/view/cole-2023. Yang Li 0116, Shao Zhang, Jichen Sun, Yali Du 0001, Ying Wen 0001, Xinbing Wang, Wei Pan 0004 |
ICML | 2 |
| 2022 | Investigating the geometric structure of neural activation spaces with convex hull approximations
Yuting Jia, Shao Zhang, Haiwen Wang, Ying Wen 0001, Luoyi Fu, Huan Long, Xinbing Wang, Chenghu Zhou |
Neurocomputing | 2 |
| 2019 | Keyword Analysis Visualization for Chinese Historical TextsabstractHistorical texts form the basis of the study of antiquities. In the case of Chinese historical texts different genres exist, e.g. chronological and biographical works etc. The contents of these texts normally consist of complex and interrelated information which covers long time period. Traditional history research relies heavily on information extraction and analysis by human researchers. With the recent development of the internet, data science and visualization technologies, digital history gradually attracts more and more attentions and in turn significantly impacts the field of historical study through altering the accessibility of the source materials, the narrative strategy and the analytical methodologies. This paper provides a system that enhances the Chinese historical research using word segmentation, texts analysis and visualization technologies. We can improve the workflow of traditional historical research via automatically detecting important keywords in Chinese historical texts and extracting, analyzing and visualizing the relations between a keyword and other words. This does not only accelerate the text based historical study but also to a great extent increase the scope of the search and analysis of the keywords in Chinese historical texts which used to be limited by the capacity of human researchers. Jihui Zeng, Beibei Zhan, Shao Zhang, Jiajun Bie, Sheng Xiao |
VINCI | 3 |
| 2001 | CVSSearch: Searching through Source Code Using CVS CommentsabstractCVSSearch is a tool that searches for fragments of source code by using CVS comments. CVS is a version control system that is widely used in the open source community. Our search tool takes advantage of the fact that a CVS comment typically describes the lines of code involved in the commit and this description will typically hold for many future versions. In other words, CVSSearch allows one to better search the most recent version of the code by looking at previous versions to better understand the current version. In this paper we describe our algorithm for mapping CVS comments to the corresponding source code, present a search tool based on this technique, and discuss preliminary feedback. Annie Chen, Eric Chou, Joshua Wong, Andrew Y. Yao, Shao Zhang, Amir Michail |
ICSM | 6 |