Jiahe Li 0007

dblp:265/2448-7 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
20since 2021 · last 2026
0009-0007-8669-921XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction
abstract
Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this challenge by employing flattened Gaussian primitives to better fit surface geometry, combined with depth regularization to alleviate geometric ambiguities under limited viewpoints. Nevertheless, the increased anisotropy inherent in flattened Gaussians exacerbates overfitting in sparse-view scenarios, hindering accurate surface fitting and degrading novel view synthesis performance. In this paper, we propose SparseSurf, a method that reconstructs more accurate and detailed surfaces while preserving high-quality novel view rendering. Our key insight is to introduce Stereo Geometry-Texture Alignment, which bridges rendering quality and geometry estimation, thereby jointly enhancing both surface reconstruction and view synthesis. In addition, we present a Pseudo-Feature Enhanced Geometry Consistency that enforces multi-view geometric consistency by incorporating both training and unseen views, effectively mitigating overfitting caused by sparse supervision. Extensive experiments on the DTU, BlendedMVS, and Mip-NeRF360 datasets demonstrate that our method achieves the state-of-the-art performance.
Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Haonan Luo 0002, Xiao Bai 0001
AAAI3
2026 FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
abstract
We present FoundationSLAM, a learning-based monocular dense SLAM system that addresses the absence of geometric consistency in previous flow-based approaches for accurate and robust tracking and mapping. Our core idea is to bridge flow estimation with geometric reasoning by leveraging the guidance from foundation depth models. To this end, we first develop a Hybrid Flow Network that produces geometry-aware correspondences, enabling consistent depth and pose inference across diverse keyframes. To enforce global consistency, we propose a Bi-Consistent Bundle Adjustment Layer that jointly optimizes keyframe pose and depth under multi-view constraints. Furthermore, we introduce a Reliability-Aware Refinement mechanism that dynamically adapts the flow update process by distinguishing between reliable and uncertain regions, forming a closed feedback loop between matching and optimization. Extensive experiments demonstrate that FoundationSLAM achieves superior trajectory accuracy and dense reconstruction quality across multiple challenging datasets, while running in real-time at 18 FPS, demonstrating strong generalization to various scenarios and practical applicability of our method.
Jiahe Li 0007, Fabio Tosi, Matteo Poggi, Xiao Bai 0001
AAAI2
2026 CurvLoc: Surface Curvature Prompted Gaussian Splatting for Visual Localization
Jiahe Li 0007, Botao Jiang, Zihang Wang 0002, Xiaohan Yu 0001, Xiao Bai 0001, Haonan Luo 0002
Int. J. Comput. Vis.3
2026 DNGaussian++: Improving Sparse-View Gaussian Radiance Fields With Depth Normalization
abstract
Synthesizing novel views from sparse views has achieved impressive advances with radiance fields, yet prevailing methods suffer from high consumption or insufficient refinement capability. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian Splatting, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the remarkable advancement of recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Although DNGaussian shows impressive performance, its patch-wise regularization obscures the inconsistency in cross-patch errors. Additionally, primitives can still be irreversibly trapped in local minima under sparse views, even if depth regularization is applied. In this paper, we propose an extended version, DNGaussian++. First, a Geometry Instance Regularizer is developed to enable depth regularization for continuous consistency by exploiting reliable instance-level depth cues. Leveraging the depth gradient guidance, we then propose a Depth-Guided Geometry Reorganization to address the aforementioned local minima problem with high representation efficiency. Extensive experiments show that DNGaussian++ exhibits state-of-the-art performance in multiple datasets and scenarios with high efficiency, and the broad applicability and effectiveness are verified on various backbones and tasks.
Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Lin Gu 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 EIK-Nav: Boosting zero-shot object navigation with explicit and implicit knowledge
Botao Jiang, Haonan Luo 0002, Zihang Wang 0002, Jiahe Li 0007, Xiao Bai 0001
Pattern Recognit.4
2025 Reverse Chain-of-Thought and Causal Path Verification: A Modular Plugin for Aligning LLMs with Knowledge Graphs
abstract
Large language models (LLMs) exhibit strong language understanding capabilities, but encounter challenges when integrating structured knowledge from knowledge graphs (KGs) for complex reasoning tasks such as knowledge graph question answering (KGQA). Existing methods often rely on prompt engineering or fixed templates, which obscure the relational structure and limit generalization. To address these limitations, this paper introduces the Reverse Chain-of-Thought (R-CoT) and Causal Path Verification Plugin, a modular framework that reconstructs retrieved KG triples into reverse chains of sub-questions. Each reasoning step is aligned with a supporting triple, forming interpretable multi-hop paths. In particular, Semantic Causal Scoring (SCS) module is further incorporated to evaluate the causal alignment between each reverse sub-question and the original question through dynamic semantic vector matching. The SCS design avoids frequent interactions with LLMs and effectively filters irrelevant or unsupported reasoning steps. Based on the scoring results, a template-free, model-agnostic R-CoT input format is constructed as a semi-structured sequence. This design preserves the KG structure in natural language form and enables seamless integration with standard LLMs without fine-tuning. Experimental results demonstrate that the R-CoT Plugin consistently improves factual alignment, enhances reasoning stability, and outperforms conventional prompt-based methods in both accuracy and coherence.
Dezhuang Miao, Yibin Du, Xiang Li 0117, Jiahe Li 0007, Bo Zhang 0096, Bingyu Yan, Litian Zhang
CIKM5
2025 InsTaG: Learning Personalized 3D Talking Head from Few-Second Video
abstract
Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a fast learning of realistic personalized 3D talking head from few training data. Built upon a lightweight 3DGS person-specific synthesizer with universal motion priors, InsTaG achieves high-quality and fast adaptation while preserving high-level personalization and efficiency. As preparation, we first propose an Identity-Free Pre-training strategy that enables the pre-training of the person-specific model and encourages the collection of universal motion priors from long-video data corpus. To fully exploit the universal motion priors to learn an unseen new identity, we then present a Motion-Aligned Adaptation strategy to adaptively align the target head to the pre-trained field, and constrain a robust dynamic head structure under few training data. Experiments demonstrate our outstanding performance and efficiency under various data scenarios to render high-quality personalized talking heads. Project page: https://fictionarry.github.io/InsTaG/.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
CVPR1
2025 Towards Feature-Consistent Parameter Collaboration for Personalized Federated Learning
abstract
Personalized federated learning (PFL) aims to improve the performance of the local model on each client with the non-IID data among different clients. This paper introduces FedFPC, a PFL method that allows effective and robust parameter-wise collaboration to achieve outperforming performance. Stem from the idea that similar clients should share more consistent feature representation and benefit more from each other, two strategies are designed to ensure feature consistency during training. First, we present an Attention-Guided Critical Parameter Selection strategy for critical parameter selection, which utilizes the attention prior from current expressive all-purpose features to identify the parameter with the most contribution to feature representation for the local data. Then, a Feature-Consistent Parameter Collaboration strategy is proposed to provide robust parameter collaboration for the local model with help from feature-consistent clients, which obtain similar feature representations to the target client. Experimental results demonstrate that FedFPC stands out by its superior performance in various PFL tasks compared to state-of-the-art methods, meanwhile with better robustness in diverse complex scenarios.
Jiahe Li 0007, Yuchao Zhang 0004, Wendong Wang 0003
ICASSP2
2025 GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
abstract
Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based framework that explores and extends the under-investigated potential of sparse voxels for achieving accurate, detailed, and complete surface reconstruction. As strengths, sparse voxels support preserving the coverage completeness and geometric clarity, while corresponding challenges also arise from absent scene constraints and locality in surface refinement. To ensure correct scene convergence, we first propose a Voxel-Uncertainty Depth Constraint that maximizes the effect of monocular depth cues while presenting a voxel-oriented uncertainty to avoid quality degradation, enabling effective and robust scene constraints yet preserving highly accurate geometries. Subsequently, Sparse Voxel Surface Regularization is designed to enhance geometric consistency for tiny voxels and facilitate the voxel-based formation of sharp and accurate surfaces. Extensive experiments demonstrate our superior performance compared to existing methods across diverse challenging scenarios, excelling in geometric accuracy, detail preservation, and reconstruction completeness while maintaining high efficiency. Code is available at https://github.com/Fictionarry/GeoSVR.
Jiahe Li 0007, Youmin Zhang 0005, Xiao Bai 0001, Xiaohan Yu 0001, Lin Gu 0003
NeurIPS1
2025 Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting
abstract
We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint optimization creates a mutually reinforcing cycle: the priors enhance the quality of 3DGS, which in turn refines the priors, further improving the reconstruction. Additionally, Eve3D introduces a novel optimization step based on bundle adjustment, overcoming the limitations of the highly local supervision in standard 3DGS pipelines. Eve3D achieves state-of-the-art results in surface reconstruction and novel view synthesis on the Tanks & Temples, DTU, and Mip-NeRF360 datasets. while retaining fast convergence, highlighting an unprecedented trade-off between accuracy and speed.
Youmin Zhang 0005, Fabio Tosi, Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Matteo Poggi
NeurIPS5
2025 3D human avatar reconstruction with neural fields: A recent survey
Meiying Gu, Jiahe Li 0007, Haonan Luo 0002, Xiao Bai 0001
Image Vis. Comput.2
2025 Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow Estimation
abstract
With advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning.
Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 ALStereo: Active learning for stereo matching
Jiahe Li 0007, Meiying Gu, Xiaohan Yu 0001, Xiao Bai 0001, Edwin R. Hancock
Pattern Recognit.2
2024 DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization
abstract
Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views, yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian radiance fields, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the highly efficient representation and surprising quality of the recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry reshaping, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Extensive experiments on LLFF, DTU, and Blender datasets demonstrate that DNGaussian outperforms state-of-the-art methods, achieving comparable or better results with significantly reduced memory cost, a 25 × reduction in training time, and over 3000 × faster rendering speed. Code is available at: https://github.com/Fictionarry/DNGaussian.
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
CVPR1
2024 Robust Synthetic-to-Real Transfer for Stereo Matching
abstract
With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have investigated the robustness after fine-tuning them in real-world scenarios, during which the domain generalization ability can be seriously degraded. In this paper, we explore fine-tuning stereo matching networks without compromising their robustness to unseen domains. Our motivation stems from comparing Ground Truth (GT) versus Pseudo Label (PL) for fine-tuning: GT degrades, but PL preserves the domain generalization ability. Empirically, we find the difference between GT and PL implies valuable information that can regularize networks during fine-tuning. We also propose a framework to utilize this difference for fine-tuning, consisting of a frozen Teacher, an exponential moving average (EMA) Teacher, and a Student network. The core idea is to utilize the EMA Teacher to measure what the Student has learned and dynamically improve GT and PL for fine-tuning. We integrate our framework with state-of-the-art networks and evaluate its effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the domain generalization ability during fine-tuning. Code is available at: https://github.com/jiaw-z/DKT-Stereo.
Jiahe Li 0007, Lei Huang 0015, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
CVPR2
2024 TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003
ECCV (10)1
2024 CoR-GS: Sparse-View 3D Gaussian Splatting via Co-regularization
Jiahe Li 0007, Xiaohan Yu 0001, Lei Huang 0015, Lin Gu 0003, Xiao Bai 0001
ECCV (1)2
2024 Efficient Emotional Talking Head Generation via Dynamic 3D Gaussian Rendering
Jiahe Li 0007, Xiao Bai 0001
PRCV (6)2
2024 Expression-aware neural radiance fields for high-fidelity talking portrait synthesis
Xueni Guo, Jiahe Li 0007, Feihu Yan, Guangzhe Zhao, Caiyong Wang
Image Vis. Comput.5
2023 Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis
abstract
This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal contribution of spatial regions to guide talking portrait modeling. Specifically, to improve the accuracy of dynamic head reconstruction, a compact and expressive NeRF-based Tri-Plane Hash Representation is introduced by pruning empty spatial regions with three planar hash encoders. For speech audio, we propose a Region Attention Module to generate region-aware condition feature via an attention mechanism. Different from existing methods that utilize an MLP-based encoder to learn the cross-modal relation implicitly, the attention mechanism builds an explicit connection between audio features and spatial regions to capture the priors of local motions. Moreover, a direct and fast Adaptive Pose Encoding is introduced to optimize the head-torso separation problem by mapping the complex transformation of the head pose into spatial coordinates. Extensive experiments demonstrate that our method renders better high-fidelity and audio-lips synchronized talking portrait videos, with realistic details and high efficiency compared to previous methods. Code is available at https://github.com/Fictionarry/ER-NeRF.
Jiahe Li 0007, Xiao Bai 0001, Jun Zhou 0001, Lin Gu 0003
ICCV1