Heng Li 0009

dblp:02/3672-9 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-5143-5061ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
abstract
Scene regression methods, such as VGGT [86], solve the Structure-from-Motion (SfM) problem by directly regressing camera poses and 3D scene structures from input images. They demonstrate impressive performance in handling images under extreme viewpoint changes. However, these methods struggle to handle a large number of input images. To address this problem, we introduce SAIL-Recon, a feed-forward Transformer for large scale SfM, by augmenting the scene regression network with visual localization capabilities. Specifically, our method first computes a neural scene representation tokens from a subset of anchor images. The regression network is then fine-tuned to reconstruct all input images conditioned on this neural scene representation. Comprehensive experiments show that our method not only scales efficiently to large-scale scenes, but also achieves state-of-the-art results on both camera pose estimation and novel view synthesis benchmarks, including TUM-RGBD, CO3Dv2, and Tanks & Temples. Code and models are publicly available here.
Junyuan Deng, Heng Li 0009, Weiqiang Ren, Qian Zhang 0009, Ping Tan 0002
3DV2
2026 SPATIALGEN: Layout-Guided 3D Indoor Scene Generation
Chuan Fang, Heng Li 0009, Yixun Liang, Jia Zheng 0002, Yongsen Mao, Yuan Liu 0025, Rui Tang 0015, Zihan Zhou 0001, Ping Tan 0002
3DV2
2025 Gaussianavatar-Editor: Photorealistic Animatable Gaussian Head Avatar Editor
abstract
We introduce GaussianAvatar-Editor, an innovative framework for text-driven editing of animatable Gaussian head avatars that can be fully controlled in expression, pose, and viewpoint. Unlike static 3D Gaussian editing, editing animatable 4D Gaussian avatars presents challenges related to motion occlusion and spatial-temporal inconsistency. To address these issues, we propose the Weighted Alpha Blending Equation (WABE). This function enhances the blending weight of visible Gaussians while suppressing the influence on non-visible Gaussians, effectively handling motion occlusion during editing. Furthermore, to improve editing quality and ensure 4D consistency, we incorporate conditional adversarial learning into the editing process. This strategy helps to refine the edited results and maintain consistency throughout the animation. By integrating these methods, our GaussianAvatar-Editor
Xiangyue Liu 0001, Kunming Luo, Heng Li 0009, Yuan Liu 0025, Li Yi 0001, Ping Tan 0002
3DV3
2025 Universal Features Guided Zero-Shot Category-Level Object Pose Estimation
abstract
Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to establish semantic similarity-based correspondences and can be extended to unseen categories without additional model fine-tuning. Our method begins with combining efficient 2D universal features to find sparse correspondences between intra-category objects and gets initial coarse pose. To handle the correspondence degradation of 2D universal features if the pose deviates much from the target pose, we use an iterative strategy to optimize the pose. Subsequently, to resolve pose ambiguities due to shape differences between intra-category objects, the coarse pose is refined by optimizing with dense alignment constraint of 3D universal features. Our method outperforms previous methods on the REAL275 and Wild6D benchmarks for unseen categories.
Wentian Qu, Chenyu Meng, Heng Li 0009, Jian Cheng 0006, CuiXia Ma, Hongan Wang, Xiao Zhou 0023, Xiaoming Deng 0001, Ping Tan 0002
AAAI3
2023 Dense RGB Slam with Neural Implicit Maps
Heng Li 0009, Xiaodong Gu 0004, Weihao Yuan 0001, Luwei Yang, Zilong Dong, Ping Tan 0002
ICLR1
2023 Monocular Scene Reconstruction with 3D SDF Transformers
Weihao Yuan 0001, Xiaodong Gu 0004, Heng Li 0009, Zilong Dong, Siyu Zhu 0001
ICLR3
2022 RAGO: Recurrent Graph Optimizer For Multiple Rotation Averaging
abstract
This paper proposes a deep recurrent Rotation Averaging Graph Optimizer (RAGO) for Multiple Rotation Averaging (MRA). Conventional optimization-based methods usually fail to produce accurate results due to corrupted and noisy relative measurements. Recent learning-based approaches regard MRA as a regression problem, while these methods are sensitive to initialization due to the gauge freedom problem. To handle these problems, we propose a learnable iterative graph optimizer minimizing a gauge- invariant cost function with an edge rectification strategy to mitigate the effect of inaccurate measurements. Our graph optimizer iteratively refines the global camera rotations by minimizing each node's single rotation objective function. Besides, our approach iteratively rectifies relative rotations to make them more consistent with the current camera orientations and observed relative rotations. Furthermore,$we$employ a gated recurrent unit to improve the result by tracing the temporal information of the cost graph. Our framework is a real-time learning-to-optimize rotation averaging graph optimizer with a tiny size deployed for real-world applications. RAGO outperforms previous traditional and deep methods on real-world and synthetic datasets. The code is available at github.com/sfu-gruvi-3dv/RAGO.
Heng Li 0009, Zhaopeng Cui, Shuaicheng Liu, Ping Tan 0002
CVPR1
2021 End-to-End Rotation Averaging With Multi-Source Propagation
abstract
This paper presents an end-to-end neural network for multiple rotation averaging in SfM. Due to the manifold constraint of rotations, conventional methods usually take two separate steps involving spanning tree based initialization and iterative nonlinear optimization respectively. These methods can suffer from bad initializations due to the noisy spanning tree or outliers in input relative rotations. To handle these problems, we propose to integrate initialization and optimization together in an unified graph neural network via a novel differentiable multi-source propagation module. Specifically, our network utilizes the image context and geometric cues in feature correspondences to reduce the impact of outliers. Furthermore, unlike the methods that utilize the spanning tree to initialize orientations according to a single reference node in a top-down manner, our net-work initializes orientations according to multiple sources while utilizing information from all neighbors in a differentiable way. More importantly, our end-to-end formulation also enables iterative re-weighting of input relative orientations at test time to improve the accuracy of the final estimation by minimizing the impact of outliers. We demonstrate the effectiveness of our method on two real-world datasets, achieving state-of-the-art performance.
Luwei Yang, Heng Li 0009, Jamal Ahmed Rahim, Zhaopeng Cui, Ping Tan 0002
CVPR2