VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Huo
dblp:126/0919
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-7192-2593ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanoAdapter: Efficient Adaptation of Depth Foundation Models for Immersive Multimedia via Spherical RectificationabstractPanoramic depth estimation is essential for immersive multimedia retrieval and content understanding. However, the geometric distortions of equirectangular projection (ERP) and the scarcity of annotated data remain major obstacles. While monocular depth foundation models excel in perspective domains, their efficacy drops in panoramic scenarios due to the geometric domain discrepancy. We present PanoAdapter, a parameter-efficient framework to adapt perspective priors to panoramas. Our core innovation is a Spherical Rectification (SR) mechanism that injects explicit spherical geometric priors into the latent feature space. By utilizing latitude-aware embeddings, SR facilitates distortion-aware feature modulation, calibrating representations for the non-uniform stretching of ERP. Furthermore, we mitigate data scarcity via a Teacher-Student semi-supervised learning paradigm, leveraging large-scale unlabeled data through sky-masked pseudo-labeling. Evaluations on Matterport3D and Stanford2D3D benchmarks show that PanoAdapter achieves state-of-the-art performance. Its lightweight, modular design offers an efficient and scalable solution for integrating foundation models into immersive multimedia systems. Zhiyong Huo |
ICMR | 2 |
| 2026 | Towards Robust Sparse-View 3D Gaussian Splatting via Hierarchical Depth and Multi-View ConsistencyabstractWhile Three-dimensional Gaussian Splatting (3DGS) has emerged as a premier paradigm for real-time novel view synthesis, its performance suffers from severe geometric degeneracy and overfitting in sparse-view scenarios due to the rank-deficient nature of the optimization. To address these fundamental challenges, we propose SegGeoGS, a framework that reformulates sparse-view reconstruction as a hierarchical consensus problem across geometric and photometric manifolds. First, we introduce a hierarchical depth supervision mechanism that decouples the scene into heterogeneous confidence zones, utilizing semantic-guided manifolds to adaptively calibrate dense monocular priors. Second, to prevent misfitting the underlying geometric structure under sparse observations, we design a stochastic Gaussian depth rendering mode by incorporating probabilistic sampling during optimization. Finally, we develop multi-view consistency constraints that jointly optimize the scene’s geometry and appearance through Gaussian-weighted geometric reprojection and SSIM-driven texture regularization. Extensive evaluations on the NeRF-LLFF benchmark demonstrate that SegGeoGS outperforms state-of-the-art depth-guided paradigms in the vast majority of scenes, achieving superior geometric fidelity and visual realism. Jianhui Zheng, Zhiyong Huo |
ICMR | 2 |
| 2025 | Efficient Monocular Depth Estimation Via Single-Step Latent Diffusion ModelsabstractMonocular depth estimation is a fundamental task in computer vision, aiming to recover the depth information of a scene from a RGB image. Traditional latent diffusion models start from Gaussian noise and progressively restore depth information through iterative denoising, during which errors may accumulate, leading to blurred boundaries and the loss of fine details. In this work, we propose a method that directly predicts depth maps from RGB images by leveraging a pre-trained Stable Diffusion model as the generative backbone. The denoising U-Net of Stable Diffusion is redefined as an encoder-decoder architecture to enable more efficient feature extraction and reconstruction. The method reduces artifacts and cumulative errors while improving inference speed. In addition, the control of depth-specific features is enhanced through a CLIP-based conditional generation mechanism. Extensive evaluations on benchmark datasets demonstrate that the proposed approach achieves competitive performance in both accuracy and efficiency. Zhiyong Huo |
ICMR | 1 |
| 2012 | A cost construction via MSW and linear regression for stereo matching
Tianliang Liu, Xiubin Dai, Zhiyong Huo, Xiuchang Zhu, Limin Luo 0001 |
ICPR | 3 |