Tianxing Zhou

dblp:353/3006 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Motion planning and robot control · 34% Robot manipulation · 24% Reinforcement learning · 20%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
learning from demonstration
1.922026
Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic Manipulation · IEEE Trans. Robotics 2026
GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning · CVPR 2025
Robotics › Motion planning and robot control › robot learning
manipulation skill learning
1.822026
Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic Manipulation · IEEE Trans. Robotics 2026
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › cross-embodiment learning
cross-embodiment transfer
0.912025
GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning · CVPR 2025
Robotics › Motion planning and robot control › robot learning › robot policy learning
video-conditioned policy learning
0.912025
GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning · CVPR 2025
Virtual and augmented reality › interactive storytelling
immersive storytelling
0.912025
So Long: Interactive Storytelling, Embodying Collective Historical Memory, and Participatory Archiving in a VR Voyage · ACM Multimedia 2025
Machine learning › Reinforcement learning
imitation learning
0.812024
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions · NeurIPS 2024
Machine learning › Reinforcement learning › imitation learning › learning from observation
visual imitation learning
0.812024
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions · NeurIPS 2024
Machine learning › Graph learning
graph generation
0.622026
Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic Manipulation · IEEE Trans. Robotics 2026
GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning · CVPR 2025
Collaborative and social computing
collective memory
0.312025
So Long: Interactive Storytelling, Embodying Collective Historical Memory, and Participatory Archiving in a VR Voyage · ACM Multimedia 2025
Computer vision › Vision and language
vision-language model
0.212024
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

interactive storytelling · 1.7self-supervised pretraining · 1.0graph-to-graph generative modeling · 1.0policy learning · 0.9graph neural network · 0.9generative modeling · 0.9iterative comparison · 0.8hierarchical constraint representation · 0.8
YearPublicationVenuePosition
2026 Learning From Videos Through Graph-to-Graphs Generative Modeling for Robotic Manipulation
abstract
Learning from demonstration is a powerful method for robotic skill acquisition. Nevertheless, a critical limitation lies in the substantial costs associated with gathering demonstration datasets, typically action-labeled robot data, which creates a fundamental constraint in the field. Video data offer a compelling solution as an alternative rich data source, containing diverse behavioral and physical knowledge. This study introduces G3M, an innovative framework that exploits video data viaGraph-to-GraphsGenerativeModeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. The proposed G3M abstracts video frame into graph representations by identifying object and visual action vertices for capturing state information. It then effectively models internal structures and spatial relationships present in these graph constructions, with the objective of predicting forthcoming graphs. The generated graphs function as conditional inputs that guide the control policy in determining robotic behaviors. This concise method effectively encodes critical spatial relationships while facilitating accurate prediction of subsequent graph sequences, thus allowing the development of resilient control policy despite constraints in action-annotated training samples. Furthermore, these transferable graph representations enable the effective extraction of manipulation knowledge through human videos as well as recordings from robots with different embodiments. The experimental results demonstrate that G3M attains superior performance using merely 20% action-labeled data relative to comparable approaches. Moreover, our method outperforms the state-of-the-art method, showing performance gains exceeding 19% in simulated environments and 23% in real-world experiments, while delivering improvements of over 35% in cross-embodiment transfer experiments and exhibiting strong performance on long-horizon tasks. Our project page is available athttps://g3m-project.github.io/.
Guangyan Chen, Meiling Wang 0002, Te Cui, Chengcai Yang, Mengxiao Hu, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue
IEEE Trans. Robotics8
2025 GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning
abstract
Learning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alternative. In this paper, we present GraphMimic, a novel paradigm that leverages video data via graph-to-graphs generative modeling, which pre-trains models to generate future graphs conditioned on the graph within a video frame. Specifically, GraphMimic abstracts video frames into object and visual action vertices, and constructs graphs for state representations. The graph generative modeling network then effectively models internal structures and spatial relationships within the constructed graphs, aiming to generate future graphs. The generated graphs serve as conditions for the control policy, mapping to robot actions. Our concise approach captures important spatial relations and enhances future graph generation accuracy, enabling the acquisition of robust policies from limited action-labeled data. Furthermore, the transferable graph representations facilitate the effective learning of manipulation skills from cross-embodiment videos. Our experiments exhibit that GraphMimic achieves superior performance using merely 20% action-labeled data. Moreover, our method outperforms the state-of-the-art method by over 17% and 23% in simulation and real-world experiments, and delivers improvements of over 33% in cross-embodiment transfer experiments.
Guangyan Chen, Te Cui, Meiling Wang 0002, Chengcai Yang, Mengxiao Hu, Yao Mu 0001, Zicai Peng, Tianxing Zhou, Xinran Jiang, Yi Yang 0009, Yufeng Yue
CVPR9
2025 Human Demonstrations are Generalizable Knowledge for Robots
abstract
Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object instances. In this paper, we propose a different perspective, considering human demonstration videos not as mere instructions, but as a source of knowledge for robots. Motivated by this perspective and the remarkable comprehension and generalization capabilities exhibited by large language models (LLMs), we propose DigKnow, a method that DIstills Generalizable KNOWledge with a hierarchical structure. Specifically, DigKnow begins by converting human demonstration video frames into observation knowledge. This knowledge is then subjected to analysis to extract human action knowledge and further distilled into pattern knowledge that comprises task and object instances, resulting in the acquisition of generalizable knowledge with a hierarchical structure. In settings with different tasks or object instances, DigKnow retrieves relevant knowledge for the current task and object instances. Subsequently, the LLM-based planner conducts planning based on the retrieved knowledge, and the policy executes actions in line with the plan to achieve the designated task. Utilizing the retrieved knowledge, we validate and rectify planning and execution outcomes, resulting in a substantial enhancement of the success rate. Experimental results across a range of tasks and scenes demonstrate the effectiveness of this approach in facilitating real-world robots to accomplish tasks with the knowledge derived from human demonstrations.
Te Cui, Tianxing Zhou, Mengxiao Hu, Zicai Peng, Haizhou Li 0004, Guangyan Chen, Meiling Wang 0002, Yufeng Yue
IROS2
2025 OpenVox: Real-time Instance-level Open-vocabulary Probabilistic Voxel Representation
abstract
In recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition helps mitigate semantic ambiguity in point-wise feature maps, efficiently obtaining rich semantic understanding and robust incremental reconstruction at the instance-level remains challenging. To address these challenges, we introduce OpenVox, a real-time incremental open-vocabulary probabilistic instance voxel representation. In the front-end, we design an efficient instance segmentation and comprehension pipeline that enhances language reasoning through encoding captions. In the back-end, we implement probabilistic instance voxels and formulate the cross-frame incremental fusion process into two subtasks: instance association and live map evolution, ensuring robustness to sensor and segmentation noise. Extensive evaluations across multiple datasets demonstrate that OpenVox achieves state-of-the-art performance in zero-shot instance segmentation, semantic segmentation, and open-vocabulary retrieval. The project page of OpenVox is available at https://open-vox.github.io/.
Yinan Deng, Bicheng Yao, Yihang Tang, Tianxing Zhou, Yi Yang 0009, Yufeng Yue
IROS4
2025 STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner
abstract
The ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon embodied tasks. In the STEP framework, we construct a subgoal tree through a pair of closed-loop models: a subgoal decomposition model and a leaf node termination model. Within this framework, we develop a hierarchical tree structure that spans from coarse to fine resolutions. The subgoal decomposition model leverages a foundation LLM to break down complex goals into manageable subgoals, thereby spanning the subgoal tree. The leaf node termination model provides real-time feedback based on environmental states, determining when to terminate the tree spanning and ensuring each leaf node can be directly converted into a primitive action. Experiments conducted in both the VirtualHome WAH-NL benchmark and on real robots demonstrate that STEP achieves long-horizon embodied task completion with success rates up to 34% (WAH-NL) and 25% (real robot) outperforming SOTA methods.
Tianxing Zhou, Haojia Ao, Guangyan Chen, Boyang Xing, Cheng Jingwen, Yi Yang 0009, Yufeng Yue
IROS1
2025 So Long: Interactive Storytelling, Embodying Collective Historical Memory, and Participatory Archiving in a VR Voyage
Tianxing Zhou, Chengkai Xu, Xinyue Yao
ACM Multimedia1
2024 VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
abstract
Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, current VIL methods naively employ VLMs to learn high-level plans from human videos, relying on pre-defined motion primitives for executing physical interactions, which remains a major bottleneck. In this work, we present VLMimic, a novel paradigm that harnesses VLMs to directly learn even fine-grained action levels, only given a limited number of human videos. Specifically, VLMimic first grounds object-centric movements from human videos, and learns skills using hierarchical constraint representations, facilitating the derivation of skills with fine-grained action levels from limited human videos. These skills are refined and updated through an iterative comparison strategy, enabling efficient adaptation to unseen environments. Our extensive experiments exhibit that our VLMimic, using only 5 human videos, yields significant improvements of over 27% and 21% in RLBench and real-world manipulation tasks, and surpasses baselines by more than 37% in long-horizon tasks. Code and videos are available on our anonymous homepage.
Guangyan Chen, Meiling Wang 0002, Te Cui, Yao Mu 0001, Tianxing Zhou, Zicai Peng, Mengxiao Hu, Haizhou Li 0004, Li Yuan 0007, Yi Yang 0009, Yufeng Yue
NeurIPS6