VLDB 2026 Research / reviewers in the wild / expert
Shaohui Jiao
dblp:00/7555
· DBLP profile ↗
13ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0002-5350-6809ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang 0001, Zhen Li 0031, Shaohui Jiao, Daquan Zhou, Qibin Hou, Ming-Ming Cheng |
ICCV | 6 |
| 2025 | Geometrically-plausible and Semantically-consistent Generation of Indoor PanoramasabstractWe present PanoGPS, a new approach for generating indoor panoramas, capable of following a room layout sketch and description text to generate panoramas with geometrically-plausible room layouts and semantically-consistent contents. Overall, our idea is to inform the generative model collectively with the two inputs, to inspire it to implicitly generate and refine a unified latent code for panorama generation. Specifically, we first propose using the semantic layout map to encode the room geometry to condition the generative model with plausible geometric information. Second, we establish a pipeline to integrate features from the two inputs and generate a single unified latent code that describes both the room layout and room contents of the target panorama. Third, we progressively refine the latent code to produce a coarse panorama to enforce seamless left-right border connections, while iteratively upsizing the coarse panorama with details. Our method outperforms existing approaches by generating high-quality panoramas with geometrically plausible structures and semantically meaningful content. Code is available at: https://github.com/wumengyangok/PanoGPS Zhiliang Zeng, Mengyang Wu, Xianzhi Li 0001, Wenzhao Gao, Shaohui Jiao, Chi-Wing Fu |
ICME | 5 |
| 2025 | TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsabstractThis paper introduces TempSamp-R1, a new reinforcement fine-tuning framework designed to improve the effectiveness of adapting multimodal large language models (MLLMs) to video temporal grounding tasks. We reveal that existing reinforcement learning methods, such as Group Relative Policy Optimization (GRPO), rely on on-policy sampling for policy updates. However, in tasks with large temporal search spaces, this strategy becomes both inefficient and limited in performance, as it often fails to identify temporally accurate solutions. To address this limitation, TempSamp-R1 leverages ground-truth annotations as off-policy supervision to provide temporally precise guidance, effectively compensating for the sparsity and misalignment in on-policy solutions. To further stabilize training and reduce variance in reward-based updates, TempSamp-R1 provides a non-linear soft advantage computation method that dynamically reshapes the reward feedback via an asymmetric transformation. By employing a hybrid Chain-of-Thought (CoT) training paradigm, TempSamp-R1 optimizes a single unified model to support both CoT and non-CoT inference modes, enabling efficient handling of queries with varying reasoning complexity. Experimental results demonstrate that TempSamp-R1 outperforms GRPO-based baselines, establishing new state-of-the-art performance on benchmark datasets:Charades-STA (R1\@0.7: 52.9\%, +**2.7**\%), ActivityNet Captions (R1\@0.5: 56.0\%, +**5.3**\%), and QVHighlights (mAP: 30.0\%, +**3.0**\%). Moreover, TempSamp-R1 shows robust few-shot generalization capabilities under limited data. Code is available at https://github.com/HVision-NKU/TempSamp-R1. Shaoyong Jia, Hangyi Kuang, Shaohui Jiao, Qibin Hou, Ming-Ming Cheng |
NeurIPS | 5 |
| 2025 | Topology-Aware Optimization of Gaussian Primitives for Human-Centric Volumetric VideosabstractVolumetric video is emerging as a key medium for digitizing the dynamic physical world, creating the virtual environments with six degrees of freedom to deliver immersive user experiences. However, robustly modeling general dynamic scenes, especially those involving topological changes while maintaining long-term tracking remains a fundamental challenge. In this paper, we present TaoGS, a novel topology-aware dynamic Gaussian representation that disentangles motion and appearance to support, both, long-range tracking and topological adaptation. We represent scene motion with a sparse set of motion Gaussians, which are continuously updated by a spatio-temporal tracker and photometric cues that detect structural variations across frames. To capture fine-grained texture, each motion Gaussian anchors and dynamically activates a set of local appearance Gaussians, which are non-rigidly warped to the current frame to provide strong initialization and significantly reduce training time. This activation mechanism enables efficient modeling of detailed textures and maintains temporal coherence, allowing high-fidelity rendering even under challenging scenarios such as changing clothes. To enable seamless integration into codec-based volumetric formats, we introduce a global Gaussian Lookup Table that records the lifespan of each Gaussian and organizes attributes into a lifespan-aware 2D layout. This structure aligns naturally with standard video codecs and supports up to 40× compression. TaoGS provides a unified, adaptive solution for scalable volumetric video under topological variation, capturing moments where “elegance in motion” and “Power in Stillness”— delivering immersive experiences that harmonize with the physical world. Project page: https://guochch.github.io/TaoGS/. Yuheng Jiang, Yize Wu, Shengkun Zhu, Zhehao Shen, Yingliang Zhang, Shaohui Jiao, Zhuo Su 0006, Lan Xu 0003, Marc Habermann, Christian Theobalt |
SIGGRAPH Asia | 8 |
| 2024 | SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth EstimationabstractRecently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details. In this paper, we introduce SQLdepth, a novel approach that can effectively learn fine-grained scene structure priors from ego-motion. In SQLdepth, we propose a novel Self Query Layer (SQL) to build a self-cost volume and infer depth from it, rather than inferring depth from feature maps. We show that, the self-cost volume is an effective inductive bias for geometry learning, which implicitly models the single-frame scene geometry, with each slice of it indicating a relative distance map between points and objects in a latent space. Experimental results on KITTI and Cityscapes show that our method attains remarkable state-of-the-art performance, and showcases computational efficiency, reduced training complexity, and the ability to recover fine-grained scene details. Moreover, the self-matching-oriented relative distance querying in SQL improves the robustness and zero-shot generalization capability of SQLdepth. Code is available at https://github.com/hisfog/SfMNeXt-Impl. Youhong Wang, Yunji Liang, Shaohui Jiao, Hongkai Yu |
AAAI | 4 |
| 2024 | Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene ReconstructionabstractImplicit neural representation has paved the way for new approaches to dynamic scene reconstruction. Nonetheless, cutting-edge dynamic neural rendering methods rely heavily on these implicit representations, which frequently struggle to capture the intricate details of objects in the scene. Furthermore, implicit methods have difficulty achieving real-time rendering in general dynamic scenes, limiting their use in a variety of tasks. To address the issues, we propose a deformable 3D Gaussians splatting method that reconstructs scenes using 3D Gaussians and learns them in canonical space with a deformation field to model monocular dynamic scenes. We also introduce an annealing smoothing training mechanism with no extra overhead, which can mitigate the impact of inaccurate poses on the smoothness of time interpolation tasks in real-world scenes. Through a differential Gaussian rasterizer, the deformable 3D Gaussians not only achieve higher rendering quality but also real-time rendering speed. Experiments show that our method outperforms existing methods significantly in terms of both rendering quality and speed, making it well-suited for tasks such as novel-view synthesis, time interpolation, and real-time rendering. Our code is available at https://github.com/ingra14m/Deformable-3D-Gaussians. Ziyi Yang 0008, Shaohui Jiao, Yuqing Zhang 0005, Xiaogang Jin 0001 |
CVPR | 4 |
| 2024 | Spec-Gaussian: Anisotropic View-Dependent Appearance for 3D Gaussian SplattingabstractThe recent advancements in 3D Gaussian splatting (3D-GS) have not only facilitated real-time rendering through modern GPU rasterization pipelines but have also attained state-of-the-art rendering quality. Nevertheless, despite its exceptional rendering quality and performance on standard datasets, 3D-GS frequently encounters difficulties in accurately modeling specular and anisotropic components. This issue stems from the limited ability of spherical harmonics (SH) to represent high-frequency information. To overcome this challenge, we introduce Spec-Gaussian, an approach that utilizes an anisotropic spherical Gaussian (ASG) appearance field instead of SH for modeling the view-dependent appearance of each 3D Gaussian. Additionally, we have developed a coarse-to-fine training strategy to improve learning efficiency and eliminate floaters caused by overfitting in real-world scenes. Our experimental results demonstrate that our method surpasses existing approaches in terms of rendering quality. Thanks to ASG, we have significantly improved the ability of 3D-GS to model scenes with specular and anisotropic components without increasing the number of 3D Gaussians. This improvement extends the applicability of 3D GS to handle intricate scenarios with specular and anisotropic surfaces. Ziyi Yang 0008, Yang-Tian Sun, Yihua Huang 0002, Xiaoyang Lyu, Shaohui Jiao, Xiaojuan Qi 0001, Xiaogang Jin 0001 |
NeurIPS | 7 |
| 2024 | CamoFormer: Masked Separable Attention for Camouflaged Object DetectionabstractHow to identify and segment camouflaged objects from the background is challenging. Inspired by the multi-head self-attention in Transformers, we present a simple masked separable attention (MSA) for camouflaged object detection. We first separate the multi-head self-attention into three parts, which are responsible for distinguishing the camouflaged objects from the background using different mask strategies. Furthermore, we propose to capture high-resolution semantic representations progressively based on a simple top-down decoder with the proposed MSA to attain precise segmentation results. These structures plus a backbone encoder form a new model, dubbed CamoFormer. Extensive experiments show that CamoFormer achieves new state-of-the-art performance on three widely-used camouflaged object detection benchmarks. To better evaluate the performance of the proposed CamoFormer around the border regions, we propose to use two new metrics, i.e., BR-M and BR-F. There are on average ∼ 5% relative improvements over previous methods in terms of S-measure and weighted F-measure. Bowen Yin, Xuying Zhang, Deng-Ping Fan, Shaohui Jiao, Ming-Ming Cheng, Luc Van Gool, Qibin Hou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Real-time local stereo via edge-aware disparity propagation
Xing Mei, Shaohui Jiao, Mingcai Zhou, Haitao Wang 0006 |
Pattern Recognit. Lett. | 3 |
| 2013 | Principal Observation Ray Calibration for Tiled-Lens-Array Integral Imaging DisplayabstractIntegral imaging display (IID) is a promising technology to provide realistic 3D image without glasses. To achieve a large screen IID with a reasonable fabrication cost, a potential solution is a tiled-lens-array IID (TLA-IID). However, TLA-IIDs are subject to 3D image artifacts when there are even slight misalignments between the lens arrays. This work aims at compensating these artifacts by calibrating the lens array poses with a camera and including them in a ray model used for rendering the 3D image. Since the lens arrays are transparent, this task is challenging for traditional calibration methods. In this paper, we propose a novel calibration method based on defining a set of principle observation rays that pass lens centers of the TLA and the camera's optical center. The method is able to determine the lens array poses with only one camera at an arbitrary unknown position without using any additional markers. The principle observation rays are automatically extracted using a structured light based method from a dense correspondence map between the displayed and captured pixels. Experiments show that lens array misalignments can be estimated with a standard deviation smaller than 0.4 pixels. Based on this, 3D image artifacts are shown to be effectively removed in a test TLA-IID with challenging misalignments. Haitao Wang 0006, Mingcai Zhou, Shandong Wang, Shaohui Jiao, Xing Mei, Hoyoung Lee, Ji Yeun Kim |
CVPR | 5 |
| 2009 | Time-Varying Simulation for Image-Based CarpetsabstractThe paper presents a novel approach for simulating realistic time-varied carpets. By the approach, a 3D carpet is constructed first by image-based techniques from a single photo input, through a texel structure, established to generate realistic carpet with its pattern guided by the captured image. Secondly, a time-varying simulation model is proposed to capture various aspects of the time-dependent appearance of carpets such as dust accumulation, color fading and fiber change. Additionally, a time-varying map is provided to control the specific weathering degrees at different voxels in time through a hierarchal model. Experimental results show that the realistic time-varying simulation is successfully achieved with the proposed techniques. Shaohui Jiao, Youquan Liu, Enhua Wu |
ICIG | 1 |
| 2009 | Realistic grass withering simulation using time-varying texelsabstractGrass is one of the crucial elements in representing real nature scenes. Various methods have been proposed for grass simulation, for example, Boulanger and his colleagues render realistic grass in real-time with dynamic lighting, shadows and animations. Till now, few approaches have investigated the realistic simulation of withering grassland which is still a big challenge in computer graphics, due to the great complexity of both geometry and withering mechanism. In our work, a novel concept called Time-Varying Texels (TVT) is proposed to extend the static texel[Kajiya and Kay 1989] and make it time-serialized. In a grass TVT structure, time-dependent texels are arranged hierarchically in a Time-Space Partitioning (TSP) tree, so that the grass withering process can be efficiently approximated, with acceptable spatial and temporal errors. In this way we facilitate LOD rendering. Traditional texel structures could hardly undertake physical based calculation in each grass blade, so we introduce a point based structure (PBS) to increase the flexibility. Dynamic processes such as geometric deformation and material transformation can be achieved on PBS during the whole withering procedure. Additionally, by clustering the pre-computed TVT samples in low densities using a mingling algorithm, we obtain high-density grass TVT so as to significantly improve the efficiency of TVT generation. Shaohui Jiao, Pheng-Ann Heng, Enhua Wu |
SIGGRAPH ASIA Sketches | 1 |
| 2009 | Weathering fur simulationabstractThe paper presents a novel approach for simulating the weathering fur. Dusty effects on fur is generated by volumetric γ-ton tracing method and the geometry deformation is modeled through a dynamic PBS. The proposed approach can efficiently simulate the weathering effects of fur. Shaohui Jiao, Gang Yang 0007, Enhua Wu |
VRST | 1 |