Hanzhi Chen

dblp:258/3651 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 46% Robot manipulation · 30% Deep learning architectures and training · 19%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
foundation model
1.012026
An overview of domain-specific foundation model: key technologies, applications and challenges · Sci. China Inf. Sci. 2026
Robotics › Robot manipulation › learning from demonstration
learning from human video
0.912025
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation · CVPR 2025
Robotics › Robot manipulation
grasping
0.812024
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object · ICRA 2024
Computer vision › 3D vision › object pose estimation
6d object pose estimation
0.712023
TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation · CVPR 2023
Computer vision › 3D vision
neural rendering
0.712023
TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation · CVPR 2023
Computer vision › 3D vision › pose estimation › learning-based pose estimation
self-supervised pose estimation
0.712023
TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation · CVPR 2023
Computer vision › 3D vision › 3d scene understanding › affordance grounding
3d affordance learning
0.312025
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation · CVPR 2025
Robotics › Motion planning and robot control
robot learning
0.312025
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation · CVPR 2025
Computer vision › 3D vision › implicit neural representation
neural surface representation
0.212024
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object · ICRA 2024

Methods — techniques the papers use, named apart from their topics

structure from motion · 0.9diffusion model · 0.9depth foundation model · 0.9neural surface grasping fields · 0.8categorical correspondence transfer · 0.8texture regularization · 0.7surfel-based loss · 0.7adversarial training · 0.7
YearPublicationVenuePosition
2026 An overview of domain-specific foundation model: key technologies, applications and challenges
Haolong Chen, Hanzhi Chen, Zijian Zhao 0002, Kaifeng Han, Guangxu Zhu, Yichen Zhao, Wei Xu 0001, Qingjiang Shi
Sci. China Inf. Sci.2
2025 VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
abstract
Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not scale well. We argue that learning from in-the-wild human videos offers a promising solution for robotic manipulation tasks, as vast amounts of relevant data already exist on the internet. In this work, we present VidBot, a framework enabling zero-shot robotic manipulation using learned 3D affordance from in-the-wild monocular RGB-only human videos. VidBot leverages a pipeline to extract explicit representations from them, namely 3D hand trajectories from videos, combining a depth foundation model with structure-from-motion techniques to reconstruct temporally consistent, metric-scale 3D affordance representations agnostic to embodiments. We introduce a coarse-to-fine affordance learning model that first identifies coarse actions from the pixel space and then generates fine-grained interaction trajectories with a diffusion model, conditioned on coarse actions and guided by test-time constraints for context-aware interaction planning, enabling substantial generalization to novel scenes and embodiments. Extensive experiments demonstrate the efficacy of VidBot, which significantly outperforms counterparts across 13 manipulation tasks in zero-shot settings and can be seamlessly deployed across robot systems in real-world environments. VidBot paves the way for leveraging everyday human videos to make robot learning more scalable.
Hanzhi Chen, Marc Pollefeys, Stefan Leutenegger
CVPR1
2025 Scalable Outdoors Autonomous Drone Flight with Visual-Inertial SLAM and Dense Submaps Built without LiDAR
abstract
Autonomous navigation is needed for several robotics applications. In this paper we present an autonomous Micro Aerial Vehicle (MAV) system which purely relies on cost-effective and light-weight passive visual and inertial sensors to perform large-scale autonomous navigation in outdoor, unstructured and cluttered environments. We leverage visual-inertial simultaneous localization and mapping (VI-SLAM) for accurate MAV state estimates and couple it with a volumetric occupancy submapping system to achieve a scalable mapping framework which can be directly used for path planning. To ensure the safety of the MAV during navigation, we also propose a novel reference trajectory anchoring scheme that deforms the reference trajectory the MAV is tracking upon state updates from the VI-SLAM system in a consistent way, even upon large state updates due to loop-closures. We thoroughly validate our system in both real and simulated forest environments and at peak velocities up to 3 m/s – while not encountering a single collision or system failure. To the best of our knowledge, this is the first system which achieves this level of performance in such an unstructured environment using low-cost passive visual sensors and fully on-board computation, including VI-SLAM. Code available at https://github.com/ethz-mrl/mrl_navigation.
Sebastián Barbas Laina, Simon Boche, Sotiris Papatheodorou, Dimos Tzoumanikas, Simon Schaefer, Hanzhi Chen, Stefan Leutenegger
IROS6
2025 Multipath Inflation Factor for Robust GNSS/IMU/VO Fusion-Based Navigation in Urban Areas
abstract
Global navigation satellite systems (GNSS), integrated with an inertial measurement unit (IMU) and visual sensors, are widely used for vehicular navigation. With the advancement of emerging vehicular technologies, the performance requirements for positioning, navigation, and timing (PNT) have become critical, emphasizing not only positioning accuracy but also high reliability. However, GNSS signals are susceptible to reflection and diffraction in urban environments, leading to multipath effects, such as non-line-of-sight (NLOS) reception and multipath interference. The GNSS positioning errors will increase significantly, causing the integrated navigation system to fail to meet the high-performance navigation requirements. To address this issue, we have proposed a robust GNSS/IMU/visual odometry (VO) fusion algorithm with a new GNSS weighting model and an adaptive VO velocity measurement update algorithm for urban navigation. In particular, a multipath inflation factor, based on real-time IMU and VO data, is proposed for the GNSS weighting model to mitigate multipath effects. A VO variance attenuation factor based on zero velocity detection in a robust extended Kalman filter (REKF) is also designed to adaptively adjust the covariance of VO measurements, enhancing the overall robustness of the integrated system. A field test was conducted in urban environments. The results show that the proposed algorithm achieves horizontal and 3-D positioning accuracy of 3.38 m and 5.00 m, respectively, outperforming the conventional GNSS/IMU/VO integration using C/N0-based weighting model. The improvements in horizontal and 3-D positioning accuracy are 63.4% and 56.1%, respectively.
Rui Sun 0005, Hanzhi Chen, Yi Mao 0003, Washington Yotto Ochieng
IEEE Internet Things J.4
2024 FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
abstract
We present FuncGrasp, a framework that can infer dense yet reliable grasp configurations for unseen objects using one annotated object and single-view RGB-D observation via categorical priors. Unlike previous works that only transfer a set of grasp poses, FuncGrasp aims to transfer infinite configurations parameterized by an object-centric continuous grasp function across varying instances. To ease the transfer process, we propose Neural Surface Grasping Fields (NSGF), an effective neural representation defined on the surface to densely encode grasp configurations. Further, we exploit function-to-function transfer using sphere primitives to establish semantically meaningful categorical correspondences, which are learned in an unsupervised fashion without any expert knowledge. We showcase the effectiveness through extensive experiments in both simulators and the real world. Remarkably, our framework significantly outperforms several strong baseline methods in terms of density and reliability for generated grasps.
Hanzhi Chen, Binbin Xu 0001, Stefan Leutenegger
ICRA1
2023 TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation
abstract
In this paper, we introduce neural texture learning for 6D object pose estimation from synthetic data and a few unlabelled real images. Our major contribution is a novel learning scheme which removes the drawbacks of previous works, namely the strong dependency on co-modalities or additional refinement. These have been previously necessary to provide training signals for convergence. We formulate such a scheme as two sub-optimisation problems on texture learning and pose learning. We separately learn to predict realistic texture of objects from real image collections and learn pose estimation from pixel-perfect synthetic data. Combining these two capabilities allows then to synthesise photorealistic novel views to supervise the pose estimator with accurate geometry. To alleviate pose noise and segmentation imperfection present during the texture learning phase, we propose a surfel-based adversarial training loss together with texture regularisation from synthetic data. We demonstrate that the proposed approach significantly outperforms the recent state-of-the-art methods without ground-truth pose annotations and demonstrates substantial generalisation improvements towards unseen scenes. Remarkably, our scheme improves the adopted pose estimators substantially even when initialised with much inferior performance.
Hanzhi Chen, Fabian Manhardt, Nassir Navab, Benjamin Busam
CVPR1
2021 Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth Estimation
abstract
Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer architecture, together with novel regularized loss formulations, can improve depth consistency while preserving accuracy. We propose a spatial attention module that correlates coarse depth predictions to aggregate local geometric information. A novel temporal attention mechanism further processes the local geometric information in a global context across consecutive images. Additionally, we introduce geometric constraints between frames regularized by photometric cycle consistency. By combining our proposed regularization and the novel spatial-temporal-attention module we fully leverage both the geometric and appearance-based consistency across monocular frames. This yields geometrically meaningful attention and improves temporal depth stability and accuracy compared to previous methods.
Patrick Ruhkamp, Daoyi Gao, Hanzhi Chen, Nassir Navab, Benjamin Busam
3DV3