EDBT 2026 Demo / reviewers in the wild / expert
Ceyao Zhang
dblp:277/1121
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-2544-0718ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Multi-agent systems · 38% Language models and text generation · 31% Robot manipulation · 28% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation › grasping
multifingered grasping |
1.0 | 1 | 2026 | DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026 |
Natural language and speech › Language models and text generation
foundation model agents |
0.9 | 1 | 2025 | Cradle: Empowering Foundation Agents towards General Computer Control · ICML 2025 |
Knowledge, reasoning and agents › Multi-agent systems
cooperative agents |
0.8 | 1 | 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024 |
Natural language and speech › Language models and text generation
instruction following |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent reasoning
intention inference |
0.8 | 1 | 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
LLM-based multi-agent collaboration |
0.8 | 1 | 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024 |
Natural language and speech › Language models and text generation
multimodal language model |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
zero-shot coordination |
0.8 | 1 | 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024 |
Program synthesis and code generation
code generation with language models |
0.8 | 1 | 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024 |
Program synthesis and code generation › code generation with language models
software engineering agents |
0.8 | 1 | 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024 |
Robotics › Robot manipulation
grasping |
0.3 | 1 | 2026 | DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026 |
Robotics › Robot manipulation
non-prehensile grasping |
0.3 | 1 | 2026 | DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration |
0.2 | 1 | 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024 |
Knowledge, reasoning and agents › Multi-agent systems › intelligent agents
open-world agent |
0.2 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Natural language and speech › Language models and text generation › prompting
prompt engineering |
0.2 | 1 | 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.3imitation learning · 1.8self-reflection · 1.7large multimodal model · 1.7action planning · 1.7prompt sequences · 1.5vision-language model · 1.0diffusion policy · 1.0belief updating · 0.8autoregressive transformer · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingabstractDexterous grasping remains a fundamental yet challenging problem in robotics. A general-purpose robot must be capable of grasping diverse objects in arbitrary scenarios. However, existing research typically relies on restrictive assumptions, such as single-object settings or limited environments, showing constrained generalization. We present DexGraspVLA, a hierarchical framework for robust generalization in language-guided general dexterous grasping and beyond. It utilizes a pre-trained Vision-Language model as the high-level planner and learns a diffusion-based low-level Action controller. The key insight to achieve generalization lies in iteratively transforming diverse language and visual inputs into domain-invariant representations via foundation models, where imitation learning can be effectively applied due to the alleviation of domain shift. Notably, our method achieves a 90+% dexterous grasping success rate under thousands of challenging unseen cluttered scenes. Empirical analysis confirms the consistency of internal model behavior across environmental variations, validating our design. DexGraspVLA also, for the first time, simultaneously demonstrates free-form long-horizon prompt execution, robustness to adversarial objects and human disturbance, and failure recovery. Extended application to nonprehensile grasping further proves its generality. Yifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang, Tianrui Guan, Fanlian Zeng, Ka Nam Lui, Yuyao Ye, Yitao Liang, Yaodong Yang 0001, Yuanpei Chen |
AAAI | 4 |
| 2025 | Cradle: Empowering Foundation Agents towards General Computer ControlabstractDespite their success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Information Gathering, Self-Reflection, Task Inference, Skill Curation, Action Planning, and Memory, Cradle is able to understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning and information retrieval, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games (Red Dead Redemption 2, Cities:Skylines, Stardew Valley and Dealer’s Life 2), five software applications (Chrome, Outlook, Feishu, Meitu and CapCut), and a comprehensive benchmark, OSWorld. With a unified interface to interact with any software, Cradle greatly extends the reach of foundation agents thus paving the way for generalist agents. Weihao Tan, Wentao Zhang 0007, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li 0003, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Xiaoqiang Chai, Yifei Bi, Tianbao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Chaojie Wang 0001, Xinrun Wang, Börje Karlsson 0001, Bo An 0001, Shuicheng Yan, Zongqing Lu 0002 |
ICML | 21 |
| 2024 | ProAgent: Building Proactive Cooperative Agents with Large Language ModelsabstractBuilding agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they interact with during the training phase. Such reliance, however, constrains the agents' capacity for strategic adaptation when cooperating with unfamiliar teammates, which becomes a significant challenge in zero-shot coordination scenarios. To address this challenge, we propose ProAgent, a novel framework that harnesses large language models (LLMs) to create proactive agents capable of dynamically adapting their behavior to enhance cooperation with teammates. ProAgent can analyze the present state, and infer the intentions of teammates from observations. It then updates its beliefs in alignment with the teammates' subsequent actual behaviors. Moreover, ProAgent exhibits a high degree of modularity and interpretability, making it easily integrated into various of coordination scenarios. Experimental evaluations conducted within the Overcooked-AI environment unveil the remarkable performance superiority of ProAgent, outperforming five methods based on self-play and population-based training when cooperating with AI agents. Furthermore, in partnered with human proxy models, its performance exhibits an average improvement exceeding 10% compared to the current state-of-the-art method. For more information about our project, please visit https://pku-proagent.github.io. Ceyao Zhang, Kaijie Yang, Siyi Hu 0001, Guanghe Li, Yihang Sun, Zhaowei Zhang 0001, Anji Liu, Song-Chun Zhu, Xiaojun Chang, Junge Zhang, Feng Yin 0001, Yitao Liang, Yaodong Yang 0001 |
AAAI | 1 |
| 2024 | MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkabstractRecently, remarkable progress has been made on automated problem solving through societies of agents based on large language models (LLMs). Previous LLM-based multi-agent systems can already solve simple dialogue tasks. More complex tasks, however, face challenges through logic inconsistencies due to cascading hallucinations caused by naively chaining LLMs. Here we introduce MetaGPT, an innovative meta-programming framework incorporating efficient human workflows into LLM-based multi-agent collaborations. MetaGPT encodes Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together. On collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than previous chat-based multi-agent systems. Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu 0001, Jürgen Schmidhuber |
ICLR | 7 |
| 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following AgentsabstractThis paper presents OmniJARVIS, a novel Vision-Language-Action (VLA) model for open-world instruction-following agents in Minecraft. Compared to prior works that either emit textual goals to separate controllers or produce the control command directly, OmniJARVIS seeks a different path to ensure both strong reasoning and efficient decision-making capabilities via unified tokenization of multimodal interaction data. First, we introduce a self-supervised approach to learn a behavior encoder that produces discretized tokens for behavior trajectories $\tau = \{o_0, a_0, \dots\}$ and an imitation learning policy decoder conditioned on these tokens. These additional behavior tokens will be augmented to the vocabulary of pretrained Multimodal Language Models. With this encoder, we then pack long-term multimodal interactions involving task instructions, memories, thoughts, observations, textual responses, behavior trajectories, etc into unified token sequences and model them with autoregressive transformers. Thanks to the semantically meaningful behavior tokens, the resulting VLA model, OmniJARVIS, can reason (by producing chain-of-thoughts), plan, answer questions, and act (by producing behavior tokens for the imitation learning policy decoder). OmniJARVIS demonstrates excellent performances on a comprehensive collection of atomic, programmatic, and open-ended tasks in open-world Minecraft. Our analysis further unveils the crucial design principles in interaction data formation, unified tokenization, and its scaling potentials. The dataset, models, and code will be released at https://craftjarvis.org/OmniJARVIS. Shaofei Cai, Zhancun Mu, Haowei Lin, Ceyao Zhang, Xuejie Liu, Qing Li 0003, Anji Liu, Xiaojian Ma 0001, Yitao Liang |
NeurIPS | 5 |
| 2023 | Data-adaptive M-estimators for robust regression via bi-level optimization
Ceyao Zhang, Tianjian Zhang, Feng Yin 0001, Abdelhak M. Zoubir |
Signal Process. | 1 |
| 2022 | MetaLoc: Learning to Learn Indoor RSS Fingerprinting Localization over Multiple ScenariosabstractThe existing indoor fingerprinting methods based on received signal strength (RSS) are rather accurate after intensive offline calibration for a specific scenario, but the well-calibrated localization model (can be a pure statistical one or a data-driven one) will present poor generalization ability in a new scenario, which results in big loss in knowledge and human effort. To break the scenario-specific localization bottleneck, we propose a new-fashioned data-driven fingerprinting method for localization based on meta-learning, named by MetaLoc, that can adapt itself rapidly to a new, possibly unseen, scenario with very little calibration work. Specifically, the underlying localization model is taken to be a deep neural network (NN), and we train an optimal set of group-specific meta-parameters by leveraging historical data collected from diverse well-calibrated indoor scenarios and the maximum mean discrepancy criterion. Simulation results confirm that the meta-parameters obtained for MetaLoc achieves very rapid adaptation to new scenarios, competitive localization accuracy, and high resistance to significantly reduced reference points (RPs), saving a lot of calibration effort. Ceyao Zhang, Qinglei Kong, Feng Yin 0001, Lexi Xu, Kai Niu 0001 |
ICC | 2 |