EDBT 2026 Demo / reviewers in the wild / expert
Zhancun Mu
dblp:381/4972
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 34% Motion planning and robot control · 14% Multi-agent systems · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 25 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
1.0 | 1 | 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment · AAAI 2026 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › goal reasoning
goal specification |
1.0 | 1 | 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment · AAAI 2026 |
Robotics › Motion planning and robot control
robot learning |
1.0 | 1 | 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment · AAAI 2026 |
Robotics › Motion planning and robot control › robot learning
visuomotor policy |
1.0 | 1 | 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment · AAAI 2026 |
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent |
0.9 | 1 | 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting · CVPR 2025 |
Robotics › Robot navigation and mapping › embodied navigation
embodied decision-making |
0.9 | 1 | 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting · CVPR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting · CVPR 2025 |
Environmental and earth informatics › geophysics
full-waveform inversion |
0.9 | 1 | 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI · NeurIPS 2025 |
Environmental and earth informatics
geophysics |
0.9 | 1 | 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI · NeurIPS 2025 |
Environmental and earth informatics › geophysical imaging
seismic tomography |
0.9 | 1 | 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
automated negotiation |
0.8 | 1 | 2024 | A Contextual Combinatorial Bandit Approach to Negotiation · ICML 2024 |
Machine learning › Reinforcement learning › multi-armed bandit
combinatorial bandits |
0.8 | 1 | 2024 | A Contextual Combinatorial Bandit Approach to Negotiation · ICML 2024 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.8 | 1 | 2024 | A Contextual Combinatorial Bandit Approach to Negotiation · ICML 2024 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
goal-conditioned policy |
0.8 | 1 | 2024 | Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.8 | 1 | 2024 | Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024 |
Natural language and speech › Language models and text generation
instruction following |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Natural language and speech › Language models and text generation
multimodal language model |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.8 | 1 | 2024 | Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
0.8 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Human-robot interaction
intent alignment |
0.3 | 1 | 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal Alignment · AAAI 2026 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.3 | 1 | 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting · CVPR 2025 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting · CVPR 2025 |
Machine learning › Deep learning architectures and training
physics-informed neural network |
0.3 | 1 | 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWI · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems › intelligent agents
open-world agent |
0.2 | 1 | 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents · NeurIPS 2024 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning |
0.2 | 1 | 2024 | Pre-Training Goal-based Models for Sample-Efficient Reinforcement Learning · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
target visibility loss · 2.0cross-view consistency loss · 2.0behavior cloning · 2.0full-waveform inversion · 1.7forward simulation · 1.7visual-temporal context prompting · 0.9SAM2 · 0.9goal prior model · 0.8goal clustering · 0.8behavior regularization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Steering Visuomotor Policy in Open Worlds via Cross-View Goal AlignmentabstractWe aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal alignment framework that allows users to specify target objects using segmentation masks from their camera views rather than the agent’s observations. We highlight that behavior cloning alone fails to align the agent’s behavior with human intent when the human and agent camera views differ significantly. To address this, we introduce two auxiliary objectives: cross-view consistency loss and target visibility loss, which explicitly enhance the agent's spatial reasoning ability. According to this, we develop ROCKET-2, a state-of-the-art agent trained in Minecraft, achieving an improvement in the efficiency of inference 3x to 6x. We demonstrate that ROCKET-2 can directly interpret goals from human camera views, enabling better human-agent interaction. Remarkably, ROCKET-2 demonstrates zero-shot generalization capabilities: despite being trained exclusively on the Minecraft dataset, it can adapt and generalize to other 3D environments like Doom, DMLab, and Unreal through a simple action space mapping. Shaofei Cai, Zhancun Mu, Anji Liu, Yitao Liang |
AAAI | 2 |
| 2025 | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context PromptingabstractVision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges. One critical issue is bridging the gap between discrete entities in low-level observations and the abstract concepts required for effective planning. A common solution is building hierarchical agents, where VLMs serve as high-level reasoners that break down tasks into executable sub-tasks, typically specified using language. However, language suffers from the inability to communicate detailed spatial information. We propose visual-temporal context prompting, a novel communication protocol between VLMs and policy models. This protocol leverages object segmentation from past observations to guide policy-environment interactions. Using this approach, we train ROCKET-1, a low-level policy that predicts actions based on concatenated visual observations and segmentation masks, supported by real-time object tracking from SAM-2. Our method unlocks the potential of VLMs, enabling them to tackle complex tasks that demand spatial reasoning. Experiments in Minecraft show that our approach enables agents to achieve previously unattainable tasks, with a 76% absolute improvement in open-world interaction performance. Codes are available at https://craftjarvis.github.io/ROCKET-1. Shaofei Cai, Kewei Lian, Zhancun Mu, Xiaojian Ma 0001, Anji Liu, Yitao Liang |
CVPR | 4 |
| 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWIabstractGlobal seismic tomography, taking advantage of seismic waves from natural earthquakes, provides essential insights into the earth's internal dynamics. Advanced Full-Waveform Inversion (FWI) techniques, whose aim is to meticulously interpret every detail in seismograms, confront formidable computational demands in forward modeling and adjoint simulations on a global scale. Recent advancements in Machine Learning (ML) offer a transformative potential for accelerating the computational efficiency of FWI and extending its applicability to larger scales. This work presents the first 3D global synthetic dataset tailored for seismic wavefield modeling and full-waveform tomography, referred to as the Global Tomography (GlobalTomo) dataset. This dataset is comprehensive, incorporating explicit wave physics and robust geophysical parameterization at realistic global scales, generated through state-of-the-art forward simulations optimized for 3D global wavefield calculations. Through extensive analysis and the establishment of ML baselines, we illustrate that ML approaches are particularly suitable for global FWI, overcoming its limitations with rapid forward modeling and flexible inversion strategies. This work represents a cross-disciplinary effort to enhance our understanding of the earth's interior through physics-ML modeling. Shiqian Li, Zhancun Mu, Shiji Xin, Zhixiang Dai, Kuangdai Leng, Rita Zhang, Yixin Zhu 0001 |
NeurIPS | 3 |
| 2024 | Pre-Training Goal-based Models for Sample-Efficient Reinforcement LearningabstractPre-training on task-agnostic large datasets is a promising approach for enhancing the sample efficiency of reinforcement learning (RL) in solving complex tasks. We present PTGM, a novel method that pre-trains goal-based models to augment RL by providing temporal abstractions and behavior regularization. PTGM involves pre-training a low-level, goal-conditioned policy and training a high-level policy to generate goals for subsequent RL tasks. To address the challenges posed by the high-dimensional goal space, while simultaneously maintaining the agent's capability to accomplish various skills, we propose clustering goals in the dataset to form a discrete high-level action space. Additionally, we introduce a pre-trained goal prior model to regularize the behavior of the high-level policy in RL, enhancing sample efficiency and learning stability. Experimental results in a robotic simulation environment and the challenging open-world environment of Minecraft demonstrate PTGM’s superiority in sample efficiency and task performance compared to baselines. Moreover, PTGM exemplifies enhanced interpretability and generalization of the acquired low-level skills. Haoqi Yuan, Zhancun Mu, Feiyang Xie, Zongqing Lu 0002 |
ICLR | 2 |
| 2024 | A Contextual Combinatorial Bandit Approach to NegotiationabstractLearning effective negotiation strategies poses two key challenges: the exploration-exploitation dilemma and dealing with large action spaces. However, there is an absence of learning-based approaches that effectively address these challenges in negotiation. This paper introduces a comprehensive formulation to tackle various negotiation problems. Our approach leverages contextual combinatorial multi-armed bandits, with the bandits resolving the exploration-exploitation dilemma, and the combinatorial nature handles large action spaces. Building upon this formulation, we introduce NegUCB, a novel method that also handles common issues such as partial observations and complex reward functions in negotiation. NegUCB is contextual and tailored for full-bandit feedback without constraints on the reward functions. Under mild assumptions, it ensures a sub-linear regret upper bound. Experiments conducted on three negotiation tasks demonstrate the superiority of our approach. Yexin Li, Zhancun Mu, Siyuan Qi |
ICML | 2 |
| 2024 | OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following AgentsabstractThis paper presents OmniJARVIS, a novel Vision-Language-Action (VLA) model for open-world instruction-following agents in Minecraft. Compared to prior works that either emit textual goals to separate controllers or produce the control command directly, OmniJARVIS seeks a different path to ensure both strong reasoning and efficient decision-making capabilities via unified tokenization of multimodal interaction data. First, we introduce a self-supervised approach to learn a behavior encoder that produces discretized tokens for behavior trajectories $\tau = \{o_0, a_0, \dots\}$ and an imitation learning policy decoder conditioned on these tokens. These additional behavior tokens will be augmented to the vocabulary of pretrained Multimodal Language Models. With this encoder, we then pack long-term multimodal interactions involving task instructions, memories, thoughts, observations, textual responses, behavior trajectories, etc into unified token sequences and model them with autoregressive transformers. Thanks to the semantically meaningful behavior tokens, the resulting VLA model, OmniJARVIS, can reason (by producing chain-of-thoughts), plan, answer questions, and act (by producing behavior tokens for the imitation learning policy decoder). OmniJARVIS demonstrates excellent performances on a comprehensive collection of atomic, programmatic, and open-ended tasks in open-world Minecraft. Our analysis further unveils the crucial design principles in interaction data formation, unified tokenization, and its scaling potentials. The dataset, models, and code will be released at https://craftjarvis.org/OmniJARVIS. Shaofei Cai, Zhancun Mu, Haowei Lin, Ceyao Zhang, Xuejie Liu, Qing Li 0003, Anji Liu, Xiaojian Ma 0001, Yitao Liang |
NeurIPS | 3 |