Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ceyao Zhang

dblp:277/1121 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-2544-0718ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Multi-agent systems · 38% Language models and text generation · 31% Robot manipulation · 28%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation › grasping
multifingered grasping
1.012026
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026
Natural language and speech › Language models and text generation
foundation model agents
0.912025
Cradle: Empowering Foundation Agents towards General Computer Control · ICML 2025
Knowledge, reasoning and agents › Multi-agent systems
cooperative agents
0.812024
ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024
Natural language and speech › Language models and text generation
instruction following
0.812024
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent reasoning
intention inference
0.812024
ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent collaboration
LLM-based multi-agent collaboration
0.812024
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024
Natural language and speech › Language models and text generation
multimodal language model
0.812024
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.812024
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
zero-shot coordination
0.812024
ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024
Program synthesis and code generation
code generation with language models
0.812024
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024
Program synthesis and code generation › code generation with language models
software engineering agents
0.812024
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024
Robotics › Robot manipulation
grasping
0.312026
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026
Robotics › Robot manipulation
non-prehensile grasping
0.312026
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping · AAAI 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning
human-AI collaboration
0.212024
ProAgent: Building Proactive Cooperative Agents with Large Language Models · AAAI 2024
Knowledge, reasoning and agents › Multi-agent systems › intelligent agents
open-world agent
0.212024
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024
Natural language and speech › Language models and text generation › prompting
prompt engineering
0.212024
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework · ICLR 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.3imitation learning · 1.8self-reflection · 1.7large multimodal model · 1.7action planning · 1.7prompt sequences · 1.5vision-language model · 1.0diffusion policy · 1.0belief updating · 0.8autoregressive transformer · 0.8
YearPublicationVenuePosition
2026 DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
abstract
Dexterous grasping remains a fundamental yet challenging problem in robotics. A general-purpose robot must be capable of grasping diverse objects in arbitrary scenarios. However, existing research typically relies on restrictive assumptions, such as single-object settings or limited environments, showing constrained generalization. We present DexGraspVLA, a hierarchical framework for robust generalization in language-guided general dexterous grasping and beyond. It utilizes a pre-trained Vision-Language model as the high-level planner and learns a diffusion-based low-level Action controller. The key insight to achieve generalization lies in iteratively transforming diverse language and visual inputs into domain-invariant representations via foundation models, where imitation learning can be effectively applied due to the alleviation of domain shift. Notably, our method achieves a 90+% dexterous grasping success rate under thousands of challenging unseen cluttered scenes. Empirical analysis confirms the consistency of internal model behavior across environmental variations, validating our design. DexGraspVLA also, for the first time, simultaneously demonstrates free-form long-horizon prompt execution, robustness to adversarial objects and human disturbance, and failure recovery. Extended application to nonprehensile grasping further proves its generality.
Yifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang, Tianrui Guan, Fanlian Zeng, Ka Nam Lui, Yuyao Ye, Yitao Liang, Yaodong Yang 0001, Yuanpei Chen
AAAI4
2025 Cradle: Empowering Foundation Agents towards General Computer Control
abstract
Despite their success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Information Gathering, Self-Reflection, Task Inference, Skill Curation, Action Planning, and Memory, Cradle is able to understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning and information retrieval, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games (Red Dead Redemption 2, Cities:Skylines, Stardew Valley and Dealer’s Life 2), five software applications (Chrome, Outlook, Feishu, Meitu and CapCut), and a comprehensive benchmark, OSWorld. With a unified interface to interact with any software, Cradle greatly extends the reach of foundation agents thus paving the way for generalist agents.
Weihao Tan, Wentao Zhang 0007, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li 0003, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Xiaoqiang Chai, Yifei Bi, Tianbao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Chaojie Wang 0001, Xinrun Wang, Börje Karlsson 0001, Bo An 0001, Shuicheng Yan, Zongqing Lu 0002
ICML21
2024 ProAgent: Building Proactive Cooperative Agents with Large Language Models
abstract
Building agents with adaptive behavior in cooperative tasks stands as a paramount goal in the realm of multi-agent systems. Current approaches to developing cooperative agents rely primarily on learning-based methods, whose policy generalization depends heavily on the diversity of teammates they interact with during the training phase. Such reliance, however, constrains the agents' capacity for strategic adaptation when cooperating with unfamiliar teammates, which becomes a significant challenge in zero-shot coordination scenarios. To address this challenge, we propose ProAgent, a novel framework that harnesses large language models (LLMs) to create proactive agents capable of dynamically adapting their behavior to enhance cooperation with teammates. ProAgent can analyze the present state, and infer the intentions of teammates from observations. It then updates its beliefs in alignment with the teammates' subsequent actual behaviors. Moreover, ProAgent exhibits a high degree of modularity and interpretability, making it easily integrated into various of coordination scenarios. Experimental evaluations conducted within the Overcooked-AI environment unveil the remarkable performance superiority of ProAgent, outperforming five methods based on self-play and population-based training when cooperating with AI agents. Furthermore, in partnered with human proxy models, its performance exhibits an average improvement exceeding 10% compared to the current state-of-the-art method. For more information about our project, please visit https://pku-proagent.github.io.
Ceyao Zhang, Kaijie Yang, Siyi Hu 0001, Guanghe Li, Yihang Sun, Zhaowei Zhang 0001, Anji Liu, Song-Chun Zhu, Xiaojun Chang, Junge Zhang, Feng Yin 0001, Yitao Liang, Yaodong Yang 0001
AAAI1
2024 MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
abstract
Recently, remarkable progress has been made on automated problem solving through societies of agents based on large language models (LLMs). Previous LLM-based multi-agent systems can already solve simple dialogue tasks. More complex tasks, however, face challenges through logic inconsistencies due to cascading hallucinations caused by naively chaining LLMs. Here we introduce MetaGPT, an innovative meta-programming framework incorporating efficient human workflows into LLM-based multi-agent collaborations. MetaGPT encodes Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together. On collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than previous chat-based multi-agent systems.
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu 0001, Jürgen Schmidhuber
ICLR7
2024 OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
abstract
This paper presents OmniJARVIS, a novel Vision-Language-Action (VLA) model for open-world instruction-following agents in Minecraft. Compared to prior works that either emit textual goals to separate controllers or produce the control command directly, OmniJARVIS seeks a different path to ensure both strong reasoning and efficient decision-making capabilities via unified tokenization of multimodal interaction data. First, we introduce a self-supervised approach to learn a behavior encoder that produces discretized tokens for behavior trajectories $\tau = \{o_0, a_0, \dots\}$ and an imitation learning policy decoder conditioned on these tokens. These additional behavior tokens will be augmented to the vocabulary of pretrained Multimodal Language Models. With this encoder, we then pack long-term multimodal interactions involving task instructions, memories, thoughts, observations, textual responses, behavior trajectories, etc into unified token sequences and model them with autoregressive transformers. Thanks to the semantically meaningful behavior tokens, the resulting VLA model, OmniJARVIS, can reason (by producing chain-of-thoughts), plan, answer questions, and act (by producing behavior tokens for the imitation learning policy decoder). OmniJARVIS demonstrates excellent performances on a comprehensive collection of atomic, programmatic, and open-ended tasks in open-world Minecraft. Our analysis further unveils the crucial design principles in interaction data formation, unified tokenization, and its scaling potentials. The dataset, models, and code will be released at https://craftjarvis.org/OmniJARVIS.
Shaofei Cai, Zhancun Mu, Haowei Lin, Ceyao Zhang, Xuejie Liu, Qing Li 0003, Anji Liu, Xiaojian Ma 0001, Yitao Liang
NeurIPS5
2023 Data-adaptive M-estimators for robust regression via bi-level optimization
Ceyao Zhang, Tianjian Zhang, Feng Yin 0001, Abdelhak M. Zoubir
Signal Process.1
2022 MetaLoc: Learning to Learn Indoor RSS Fingerprinting Localization over Multiple Scenarios
abstract
The existing indoor fingerprinting methods based on received signal strength (RSS) are rather accurate after intensive offline calibration for a specific scenario, but the well-calibrated localization model (can be a pure statistical one or a data-driven one) will present poor generalization ability in a new scenario, which results in big loss in knowledge and human effort. To break the scenario-specific localization bottleneck, we propose a new-fashioned data-driven fingerprinting method for localization based on meta-learning, named by MetaLoc, that can adapt itself rapidly to a new, possibly unseen, scenario with very little calibration work. Specifically, the underlying localization model is taken to be a deep neural network (NN), and we train an optimal set of group-specific meta-parameters by leveraging historical data collected from diverse well-calibrated indoor scenarios and the maximum mean discrepancy criterion. Simulation results confirm that the meta-parameters obtained for MetaLoc achieves very rapid adaptation to new scenarios, competitive localization accuracy, and high resistance to significantly reduced reference points (RPs), saving a lot of calibration effort.
Ceyao Zhang, Qinglei Kong, Feng Yin 0001, Lexi Xu, Kai Niu 0001
ICC2