VLDB 2026 Research / reviewers in the wild / expert
Nan Qiao 0009
dblp:321/5423
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 73% Vision and language · 14% Information extraction and text analysis · 14% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d scene understanding
3d instance segmentation |
0.9 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › 3D vision › 3d scene understanding › 3d instance segmentation
open-vocabulary 3d instance segmentation |
0.9 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Natural language and speech › Information extraction and text analysis › open vocabulary learning
open-vocabulary recognition |
0.9 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › 3D vision › 3d reconstruction
multimodal 3d reconstruction |
0.7 | 1 | 2023 | Multimodal Neural Radiance Field · ICRA 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | Multimodal Neural Radiance Field · ICRA 2023 |
Computer vision › 3D vision › 3d object detection
3d proposal generation |
0.3 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › 3D vision
3d reconstruction |
0.3 | 1 | 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance Segmentation · ICCV 2025 |
Computer vision › 3D vision
point cloud registration |
0.2 | 1 | 2023 | Multimodal Neural Radiance Field · ICRA 2023 |
Methods — techniques the papers use, named apart from their topics
standardized maximum similarity · 0.9CLIP · 0.9Alpha-CLIP · 0.9point cloud registration · 0.7neural radiance field · 0.7infrared image supervision · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance SegmentationabstractUnlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research, we observe that these individual concepts are not mutually exclusive but complementary. In this paper, we propose a new state-of-the-art solution for OV-3DIS by carefully designing a recipe to combine the concepts together and refining them to address key challenges. Our solution follows the two-stage scheme: 3D proposal generation and instance classification. We employ robust 3D tracking-based proposal aggregation to generate 3D proposals and remove overlapped or partial proposals by iterative merging/removal. For the classification stage, we replace the standard CLIP model with Alpha-CLIP, which incorporates object masks as an alpha channel to reduce background noise and obtain object-centric representation. Additionally, we introduce the standardized maximum similarity (SMS) score to normalize text-to-proposal similarity, effectively filtering out false positives and boosting precision. Our framework achieves state-of-the-art performance on ScanNet200 and S3DIS across all AP and AR metrics, even surpassing an end-to-end closed-vocabulary method. Sanghun Jung, Ke Zhang 0028, Nan Qiao 0009, Albert Chen 0001, Yuyin Sun, Hsiang-Wei Huang, Byron Boots, Min Sun 0001, Cheng-Hao Kuo |
ICCV | 4 |
| 2024 | ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain AdaptationabstractLarge-scale pre-trained vision-language models (VLM) such as CLIP [32] have demonstrated noteworthy zero-shot classification capability, achieving 76.3% top-1 accuracy on ImageNet without seeing any examples. However, while applying CLIP to a downstream target domain, the presence of visual and text domain gaps and cross-modality misalignment can greatly impact the model performance. To address such challenges, we propose ReCLIP, a novel source-free domain adaptation method for VLMs, which does not require any source data or target labeled data. ReCLIP first learns a projection space to mitigate the misaligned visual-text embeddings and learns pseudo labels. Then, it deploys cross-modality self-training with the pseudo labels to update visual and text encoders, refine labels and reduce domain gaps and misalignment iteratively. With extensive experiments, we show that ReCLIP outperforms all the baselines significantly and improves the average accuracy of CLIP from 69.83% to 74.94% on 22 image classification benchmarks. Xuefeng Hu, Ke Zhang 0028, Albert Chen 0001, Jiajia Luo, Yuyin Sun, Ken Wang, Nan Qiao 0009, Min Sun 0001, Cheng-Hao Kuo, Ramakant Nevatia |
WACV | 8 |
| 2023 | Multimodal Neural Radiance FieldabstractThis paper addresses the challenge of reconstructing a scene with a neural radiance field (NeRF) for robot vision and scene understanding using multiple modalities. Researchers have introduced the use of NeRF to represent an object for synthesizing and rendering novel views of complex scenes by optimizing a 3-D radiance field for ray casting and rendering for 2-D RGB images. However, using RGB images alone introduces additional geometry ambiguities with transparent objects or complex scenes and cannot accurately depict the 3-D shapes. We discuss and solve this problem and use multiple modalities as input for the same NeRF model to build a multimodal NeRF by incorporating point clouds and infrared image supervision to prevent such bias. In contrast to RGB images, infrared images and point clouds are typically taken by separate cameras that cannot be aligned with the RGB camera. We further introduce the alignment of different modalities based on point cloud registration to estimate the relative transformation matrices between them before training a NeRF model with multiple modalities. We evaluate our model on chosen scenes from the ScanNet and M2DGR datasets and demonstrate that it outperforms existing state-of-the-art methods. Haidong Zhu, Yuyin Sun, Jiajia Luo, Nan Qiao 0009, Ramakant Nevatia, Cheng-Hao Kuo |
ICRA | 6 |
| 2023 | Human-in-the-Loop Video Semantic Segmentation Auto-AnnotationabstractAccurate per-pixel semantic class annotations of the entire video are crucial for designing and evaluating video semantic segmentation algorithms. However, the annotations are usually limited to a small subset of the video frames due to the high annotation cost and limited budget in practice. In this paper, we propose a novel human-in-the-loop framework called HVSA to generate semantic segmentation annotations for the entire video using only a small annotation budget. Our method alternates between active sample selection and test-time fine-tuning algorithms until annotation quality is satisfied. In particular, the active sample selection algorithm picks the most important samples to get manual annotations, where the sample can be a video frame, a rectangle, or even a super-pixel. Further, the test-time fine-tuning algorithm propagates the manual annotations of selected samples to the entire video. Real-world experiments show that our method generates highly accurate and consistent semantic segmentation annotations while simultaneously enjoys significantly small annotation cost. Nan Qiao 0009, Yuyin Sun, Chong Liu 0007, Jiajia Luo, Ke Zhang 0028, Cheng-Hao Kuo |
WACV | 1 |
| 2023 | CameraPose: Weakly-Supervised Monocular 3D Human Pose Estimation by Leveraging In-the-wild 2D AnnotationsabstractTo improve the generalization of 3D human pose estimators, many existing deep learning based models focus on adding different augmentations to training poses. However, data augmentation techniques are limited to the "seen" pose combinations and hard to infer poses with rare "unseen" joint positions. To address this problem, we present CameraPose, a weakly-supervised framework for 3D human pose estimation from a single image, which can not only be applied on 2D-3D pose pairs but also on 2D alone annotations. By adding a camera parameter branch, any in-the-wild 2D annotations can be fed into our pipeline to boost the training diversity and the 3D poses can be implicitly learned by reprojecting back to 2D. Moreover, CameraPose introduces a refinement network module with confidence-guided loss to further improve the quality of noisy 2D keypoints extracted by 2D pose estimators. Experimental results demonstrate that the CameraPose brings in clear improvements on cross-scenario datasets. Notably, it outperforms the baseline method by 3mm on the most challenging dataset 3DPW. In addition, by combining our proposed refinement network module with existing 3D pose estimators, their performance can be improved in cross-scenario evaluation. Cheng-Yen Yang, Jiajia Luo, Yuyin Sun, Nan Qiao 0009, Ke Zhang 0028, Zhongyu Jiang, Jenq-Neng Hwang, Cheng-Hao Kuo |
WACV | 5 |
| 2022 | A polar-edge context-aware (PECA) network for mirror segmentation
Liqiang He, Jiajia Luo, Ke Zhang 0028, Yuyin Sun, Nan Qiao 0009, Cheng-Hao Kuo, Sinisa Todorovic |
Image Vis. Comput. | 6 |