Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Maoqing Yao

dblp:320/3677 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot manipulation · 24% Generative modeling · 22% Reinforcement learning · 14%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › distributed training
asynchronous training
1.012026
Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System · ACL (1) 2026
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
1.012026
Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System · ACL (1) 2026
Machine learning › Generative modeling
diffusion model
0.912025
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation · NeurIPS 2025
Robotics › Autonomous driving
end-to-end driving
0.912025
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) · NeurIPS 2025
Machine learning › Reinforcement learning
model-based reinforcement learning
0.912025
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation · NeurIPS 2025
Computer vision › 3D vision
3d object detection
0.612022
Multi-Class 3D Object Detection with Single-Class Supervision · ICRA 2022
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312025
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation · NeurIPS 2025
Computer vision › 3D vision › 3d scene modeling › scene representation
4d scene representation
0.312025
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.312025
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation · NeurIPS 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.312025
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

coarse-to-fine learning · 1.0asynchronous training · 1.0sparse context memory · 0.9model-based reinforcement learning · 0.9knowledge distillation · 0.9guidance mechanism · 0.9autoregressive video diffusion · 0.94d gaussian splatting · 0.9semi-supervised learning · 0.6pseudo-labeling · 0.6
YearPublicationVenuePosition
2026 Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
abstract
Yifei Wei, Linqing Zhong, Yi Liu, Yuxiang Lu, Xindong He, Maoqing Yao, Guanghui Ren. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yifei Wei, Linqing Zhong, Yi Liu 0081, Xindong He, Maoqing Yao, Guanghui Ren
ACL (1)6
2026 Is Diversity All You Need for Scalable Robotic Manipulation?
abstract
Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood. In this work, we investigate the nuanced role of data diversity in robot learning by examining three critical dimensions-task (what to do), embodiment (which robot to use), and expert (who demonstrates)-challenging the conventional intuition of “more diverse is better”. Throughout extensive experiments on various robot platforms, we reveal that (1) task diversity proves more critical than per-task demonstration quantity, with scene diversity playing a more important role than skill diversity for robustness and generalization under distribution shifts; (2) multi-embodiment pre-training data is non-essential for cross-embodiment transfer-models trained on high-quality single-embodiment data can efficiently transfer to different platforms, showing desirable scaling property during fine-tuning and its potential of replacing large-scale multi-embodiment pre-training; and (3) expert diversity, arising from individual operational preferences and stochastic variations in human demonstrations, can be confounding to policy learning, with action rate multimodality emerging as a key contributing factor. Based on this insight, we propose a distribution debiasing method to mitigate action rate ambiguity, the yielding GO-1-Pro achieves substantial performance gains of 15%, equivalent to using 2.5× pre-training data. Collectively, these findings provide new perspectives and offer practical guidance on how to scale robotic manipulation datasets effectively. The code will be released.
Modi Shi, Li Chen 0008, Chiming Liu, Guanghui Ren, Ping Luo 0002, Di Huang 0001, Maoqing Yao, Hongyang Li 0001
IEEE Trans. Robotics9
2025 EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
abstract
We introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page.
Siyuan Huang 0004, Liliang Chen, Shengcong Chen, Yue Liao, Zhengkai Jiang 0001, Peng Gao 0007, Hongsheng Li 0001, Maoqing Yao, Guanghui Ren
NeurIPS10
2025 Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)
abstract
Reinforcement Learning (RL) can mitigate the causal confusion and distribution shift inherent to imitation learning (IL). However, applying RL to end-to-end autonomous driving (E2E-AD) remains an open problem for its training difficulty, and IL is still the mainstream paradigm in both academia and industry. Recently Model-based Reinforcement Learning (MBRL) have demonstrated promising results in neural planning; however, these methods typically require privileged information as input rather than raw sensor data. We fill this gap by designing Raw2Drive, a dual-stream MBRL approach. Initially, we efficiently train an auxiliary privileged world model paired with a neural planner that uses privileged information as input. Subsequently, we introduce a raw sensor world model trained via our proposed Guidance Mechanism, which ensures consistency between the raw sensor world model and the privileged world model during rollouts. Finally, the raw sensor world model combines the prior knowledge embedded in the heads of the privileged world model to effectively guide the training of the raw sensor policy. Raw2Drive is so far the only RL based end-to-end method on CARLA Leaderboard 2.0, and Bench2Drive and it achieves state-of-the-art performance.
Zhenjie Yang 0001, Xiaosong Jia, Xue Yang 0005, Maoqing Yao, Junchi Yan
NeurIPS5
2022 Multi-Class 3D Object Detection with Single-Class Supervision
abstract
While multi-class 3D detectors are needed in many robotics applications, training them with fully labeled datasets can be expensive in labeling cost. An alternative approach is to have targeted single-class labels on disjoint data samples. In this paper, we are interested in training a multi-class 3D object detection model, while using these single-class labeled data. We begin by detailing the unique stance of our “Single-Class Supervision” (SCS) setting with respect to related concepts such as partial supervision and semi supervision. Then, based on the case study of training the multi-class version of Range Sparse Net (RSN), we adapt a spectrum of algorithms - from supervised learning to pseudo-labeling - to fully exploit the properties of our SCS setting, and perform extensive ablation studies to identify the most effective algorithm and practice. Empirical experiments on the Waymo Open Dataset show that proper training under SCS can approach or match full supervision training while saving labeling costs.
Chenxi Liu 0001, Maoqing Yao, Weiyue Wang 0002, Zhaoqi Leng, Charles R. Qi, Dragomir Anguelov
ICRA3