Shaohui Dai

dblp:365/1139 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0005-4139-7950ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 83% Generative modeling · 14% Segmentation and scene understanding · 4%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 82% Rendering · 18%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.622025
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs · ACM Multimedia 2025
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane · ACM Multimedia 2024
Computer vision › 3D vision › 3d scene understanding
open-vocabulary 3d scene understanding
1.622025
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs · ACM Multimedia 2025
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane · ACM Multimedia 2024
Machine learning › Generative modeling
diffusion model
1.012026
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion · AAAI 2026
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
1.012026
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion · AAAI 2026
Visual content generation and editing › image generation
multi-view image generation
1.012026
DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion · AAAI 2026
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs · ACM Multimedia 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs · ACM Multimedia 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.312025
Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs · ACM Multimedia 2025
Rendering › gaussian splatting
3d gaussian splatting
0.212024
GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 2.0pseudo-ground-truth views · 2.0fine-tuning · 2.0semantic-space hyperplane · 1.53d gaussian splatting · 1.5superpoint graph · 0.9semantic feature lifting · 0.9reprojection · 0.9
YearPublicationVenuePosition
2026 DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
abstract
Reconstructing 3D objects from a single image is a long-standing challenge, particularly under real-world occlusions. While recent diffusion-based view synthesis models can generate consistent novel views from a single RGB image, they generally assume fully visible inputs and struggle when parts of the object are occluded, leading to inconsistent views and degraded 3D reconstruction quality. To address this limitation, we propose DeOcc-1-to-3, an end-to-end framework for occlusion-aware multi-view generation. Our method directly synthesizes six structurally consistent novel views from a single partially occluded image, enabling downstream 3D reconstruction without requiring prior inpainting or manual annotations. We design a self-supervised training pipeline that leverages occluded–unoccluded image pairs and pseudo-ground-truth views to guide structure-aware completion and view consistency. Without modifying the original architecture, we fully fine-tune the diffusion model to jointly learn completion and multi-view generation. Additionally, we introduce the first benchmark for occlusion-aware reconstruction, covering diverse occlusion levels, object categories, and mask patterns, providing a standardized evaluation protocol.
Yansong Qu, Shaohui Dai, Yuze Wang 0006, You Shen, Shengchuan Zhang, Liujuan Cao
AAAI2
2025 Training-Free Hierarchical Scene Understanding for Gaussian Splatting with Superpoint Graphs
abstract
Bridging natural language and 3D geometry is a crucial step toward flexible, language-driven scene understanding. While recent advances in 3D Gaussian Splatting (3DGS) have enabled fast and high-quality scene reconstruction, research has also explored incorporating open-vocabulary understanding into 3DGS. However, most existing methods require iterative optimization over per-view 2D semantic feature maps, which not only results in inefficiencies but also leads to inconsistent 3D semantics across views. To address these limitations, we introduce a training-free framework that constructs a superpoint graph directly from Gaussian primitives. The superpoint graph partitions the scene into spatially compact and semantically coherent regions, forming view-consistent 3D entities and providing a structured foundation for open-vocabulary understanding. Based on the graph structure, we design an efficient reprojection strategy that lifts 2D semantic features onto the superpoints, avoiding costly multi-view iterative training. The resulting representation ensures strong 3D semantic coherence and naturally supports hierarchical understanding, enabling both coarse- and fine-grained open-vocabulary perception within a unified semantic field. Extensive experiments demonstrate that our method achieves state-of-the-art open-vocabulary segmentation performance, with semantic field reconstruction completed over 30× faster.
Shaohui Dai, Yansong Qu, Zheyan Li, Shengchuan Zhang, Liujuan Cao
ACM Multimedia1
2024 GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space Hyperplane
Yansong Qu, Shaohui Dai, Jianghang Lin, Liujuan Cao, Shengchuan Zhang, Rongrong Ji
ACM Multimedia2
2024 CPE COIN++: Towards Optimized Implicit Neural Representation Compression Via Chebyshev Positional Encoding
Haocheng Chu, Shaohui Dai, Wenqi Ding, Tianshuo Xu, Pingyang Dai, Shengchuan Zhang, Yan Zhang 0109, Xiang Chang, Chih-Min Lin, Fei Chao 0001, Changjiang Shang, Qiang Shen 0001
PRCV (9)2
2023 Dynamic Neural Networks for Adaptive Implicit Image Compression
Binru Huang, Yongzhen Hu, Shaohui Dai
PRCV (11)4