EDBT 2026 Demo / reviewers in the wild / expert
Cheng-Hao Kuo
dblp:38/3578
· DBLP profile ↗
32ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0001-9464-9625ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 19 since 2021Systems, architecture and hardware · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset GenerationabstractMultimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are trained on images with limited diverse camera-object relations and corresponding textual descriptions. To address this, we propose a synthetic generation pipeline to create large-scale 3D visual instruction datasets. Our framework takes 3D assets as input and uses rendering and diffusion-based image generation models to create photorealistic images preserving precise camera-object relations. Additionally, large language models (LLMs) are used to generate text prompts for guiding visual instruction tuning and controlling image generation. We create Ultimate3D, a dataset of 240K VQAs with precise camera-object annotations, and corresponding benchmark. MLLMs fine-tuned on our proposed dataset outperform commercial models by a large margin, achieving an average accuracy improvement of 33.4% on camera-object relation recognition tasks. Our code, dataset, and benchmark will contribute to broad MLLM applications. Albert Chen 0001, Shashwat Verma, Sankalp Dayal, Min Sun 0001, Cheng-Hao Kuo, Daniel G. Aliaga |
WACV | 9 |
| 2025 | Zero-shot 3D Question Answering via Voxel-based Dynamic Token CompressionabstractRecent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compared to 3D-LMMs on 3D question answering tasks, largely due to the greater scale and diversity of available 2D image data in contrast to the more limited 3D data. Multi-frame VLMs, although achieving superior performance, suffer from the difficulty of retaining all the detailed visual information in the 3D scene while limiting the number of visual tokens. Common methods such as token pooling, reduce visual token usage but often lead to information loss, impairing the model’s ability to preserve visual details essential for 3D question answering tasks. To address this, we propose voxel-based Dynamic Token Compression (DTC), which combines 3D spatial priors and visual semantics to achieve over 90% reduction in visual tokens usage for current multi-frame VLMs. Our method maintains performance comparable to state-of-the-art models on 3D question answering benchmarks including OpenEQA and ScanQA, demonstrating its effectiveness. Hsiang-Wei Huang, Fu-Chen Chen, Wenhao Chai, Che-Chun Su, Sanghun Jung, Cheng-Yen Yang, Jenq-Neng Hwang, Min Sun 0001, Cheng-Hao Kuo |
CVPR | 10 |
| 2025 | UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial Referencesabstract6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous reference images that fully cover the object. Estimating 6D poses from partial references, which capture only fragments of an object’s appearance and geometry, remains challenging. To address this, we propose UA-Pose, an uncertainty-aware approach for 6D object pose estimation and online object completion specifically designed for partial references. We assume access to either (1) a limited set of RGBD images with known poses or (2) a single 2D image. For the first case, we initialize a partial object 3D model based on the provided images and poses, while for the second, we use image-to-3D techniques to generate an initial object 3D model. Our method integrates uncertainty into the incomplete 3D model, distinguishing between seen and unseen regions. This uncertainty enables confidence assessment in pose estimation and guides an uncertainty-aware sampling strategy for online object completion, enhancing robustness in pose estimation accuracy and improving object completeness. We evaluate our method on the YCB-Video, YCBInEOAT, and HO3D datasets, including RGBD sequences of YCB objects manipulated by robots and human hands. Experimental results demonstrate significant performance improvements over existing methods, particularly when object observations are incomplete or partially captured. Project page: https://minfenli.github.io/UA-Pose/ Ming-Feng Li, Fu-En Wang, Hritam Basak, Yuyin Sun, Shreekant Gayaka, Min Sun 0001, Cheng-Hao Kuo |
CVPR | 8 |
| 2025 | POp-GS: Next Best View in 3D-Gaussian Splatting with P-OptimalityabstractIn this paper, we present a novel algorithm for quantifying uncertainty and information gained within 3D Gaussian Splatting (3D-GS) through P-Optimality. While 3D-GS has proven to be a useful world model with high-quality rasterizations, it does not natively quantify uncertainty or information, posing a challenge for real-world applications such as 3D-GS SLAM. We propose to quantify information gain in 3D-GS by reformulating the problem through the lens of optimal experimental design, which is a classical solution widely used in literature. By restructuring information quantification of 3D-GS through optimal experimental design, we arrive at multiple solutions, of which T-Optimality and D-Optimality perform the best quantitatively and qualitatively as measured on two popular datasets. Additionally, we propose a block diagonal covariance approximation which provides a measure of correlation at the expense of a greater computation cost. Joey Wilson, Marcelino M. de Almeida, Sachit Mahajan, Martin Labrie, Maani Ghaffari Jadidi, Omid Ghasemalizadeh, Min Sun 0001, Cheng-Hao Kuo, Arnab Sen |
CVPR | 8 |
| 2025 | OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human Annotations
Peng-Hao Hsu, Ke Zhang 0028, Fu-En Wang, Tao Tu 0002, Ming-Feng Li, Yu-Lun Liu 0001, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo |
ICCV | 9 |
| 2025 | Details Matter for Indoor Open-Vocabulary 3D Instance SegmentationabstractUnlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research, we observe that these individual concepts are not mutually exclusive but complementary. In this paper, we propose a new state-of-the-art solution for OV-3DIS by carefully designing a recipe to combine the concepts together and refining them to address key challenges. Our solution follows the two-stage scheme: 3D proposal generation and instance classification. We employ robust 3D tracking-based proposal aggregation to generate 3D proposals and remove overlapped or partial proposals by iterative merging/removal. For the classification stage, we replace the standard CLIP model with Alpha-CLIP, which incorporates object masks as an alpha channel to reduce background noise and obtain object-centric representation. Additionally, we introduce the standardized maximum similarity (SMS) score to normalize text-to-proposal similarity, effectively filtering out false positives and boosting precision. Our framework achieves state-of-the-art performance on ScanNet200 and S3DIS across all AP and AR metrics, even surpassing an end-to-end closed-vocabulary method. Sanghun Jung, Ke Zhang 0028, Nan Qiao 0009, Albert Chen 0001, Yuyin Sun, Hsiang-Wei Huang, Byron Boots, Min Sun 0001, Cheng-Hao Kuo |
ICCV | 13 |
| 2025 | Modeling Uncertainty in 3D Gaussian Splatting Through Continuous Semantic SplattingabstractIn this paper, we present a novel algorithm for probabilistically updating and rasterizing semantic maps within 3D Gaussian Splatting (3D-GS). Although previous methods have introduced algorithms which learn to rasterize features in 3D-GS for enhanced scene understanding, 3D-GS can fail without warning which presents a challenge for safety-critical robotic applications. To address this gap, we propose a method which advances the literature of continuous semantic mapping from voxels to ellipsoids, combining the precise structure of 3D-GS with the ability to quantify uncertainty of probabilistic robotic maps. Given a set of images, our algorithm performs a probabilistic semantic update directly on the 3D ellipsoids to obtain an expectation and variance through the use of conjugate priors. We also propose a probabilistic rasterization which returns per-pixel segmentation predictions with quantifiable uncertainty. We compare our method with similar probabilistic voxel-based methods to verify our extension to 3D ellipsoids, and perform ablation studies on uncertainty quantification and temporal smoothing. Joey Wilson, Marcelino M. de Almeida, Min Sun 0001, Sachit Mahajan, Maani Ghaffari Jadidi, Parker Ewen, Omid Ghasemalizadeh, Cheng-Hao Kuo, Arnie Sen |
ICRA | 8 |
| 2025 | Enhancing Single Image to 3D Generation using Gaussian Splatting and Hybrid Diffusion Priorsabstract3D object generation from a single unposed RGB image is essential for robotic perception, as reconstructing complete geometry and texture is essential for precise manipulation, grasping, and scene understanding, which is key for autonomous navigation and dexterous interaction. Recent advancements in image-to-3D employ Gaussian Splatting with pre-trained 2D or 3D diffusion models, but a disparity exists: 2D models generate high-fidelity textures yet lack geometric consistency, while 3D models ensure structural coherence but produce overly smooth textures. To address this, we introduce a two-stage frequency-based distillation loss integrated with Gaussian Splatting, leveraging geometric priors from a 3D diffusion model’s low-frequency spectrum for structural consistency and a 2D diffusion model’s high-frequency details for sharper textures. Our approach achieves state-of-the-art 3D reconstruction quality, significantly improving robotic perception pipelines. Additionally, we demonstrate the easy adaptability of our method for highly accurate object pose estimation and tracking, which is critical for precise robotic grasping, manipulation, and scene understanding. Additional results can be found in the supplementary file. Hritam Basak, Hadi Tabatabaee, Shreekant Gayaka, Ming-Feng Li, Cheng-Hao Kuo, Arnie Sen, Min Sun 0001, Zhaozheng Yin |
IROS | 6 |
| 2025 | V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D AnnotationsabstractThe field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D training data, owing to the labor-intensive nature of 3D data collection and annotation processes. In this paper, we present V-MIND (Versatile Monocular INdoor Detector), which enhances the performance of in-door 3D detectors across a diverse set of object classes by harnessing publicly available large-scale 2D datasets. By leveraging well-established monocular depth estimation techniques and camera intrinsic predictors, we can generate 3D training data by converting large-scale 2D images into 3D point clouds and subsequently deriving pseudo 3D bounding boxes. To mitigate distance errors inherent in the converted point clouds, we introduce a novel 3D self-calibration loss for refining the pseudo 3D bounding boxes during training. Additionally, we propose a novel ambiguity loss to address the ambiguity that arises when introducing new classes from 2D datasets. Finally, through joint training with existing 3D datasets and pseudo 3D bounding boxes derived from 2D datasets, V-MIND achieves state-of-the-art object detection performance across a wide range of classes on the Omni3D indoor dataset. Jin-Cheng Jhang, Tao Tu 0002, Fu-En Wang, Ke Zhang 0028, Min Sun 0001, Cheng-Hao Kuo |
WACV | 6 |
| 2024 | GDA: Generalized Diffusion for Robust Test-Time AdaptationabstractMachine learning models face generalization challenges when exposed to out-of-distribution (OOD) samples with unforeseen distribution shifts. Recent research reveals that for vision tasks, test-time adaptation employing diffusion models can achieve state-of-the-art accuracy improvements on OOD samples by generating domain-aligned samples without altering the model's weights. Unfortunately, those studies have primarily focused on pixel-level corruptions, thereby lacking the generalization to adapt to a broader range of OOD types. We introduce Generalized Diffusion Adaptation (GDA), a novel diffusion-based test-time adaptation method robust against diverse OOD types. Specifically, GDA iteratively guides the diffusion by applying a marginal entropy loss derived from the model, in conjunction with style and content preservation losses during the reverse sampling process. In other words, GDA considers the model's output behavior and the samples' semantic information as a whole, reducing ambiguity in downstream tasks. Evaluation across various model architectures and OOD benchmarks indicates that GDA consistently surpasses previous diffusion-based adaptation methods. Notably, it achieves the highest classification accuracy improvements, ranging from 4.4% to 5.02% on ImageNet-C and 2.5% to 7.4% on Rendition, Sketch, and Stylized benchmarks. This performance highlights GDA's generalization to a broader range of OOD benchmarks. Yun-Yun Tsai, Fu-Chen Chen, Albert Chen 0001, Che-Chun Su, Min Sun 0001, Cheng-Hao Kuo |
CVPR | 7 |
| 2024 | No More Ambiguity in 360° Room Layout via Bi-Layout EstimationabstractInherent ambiguity in layout annotations poses significant challenges to developing accurate 360° room layout estimation models. To address this issue, we propose a novel Bi-Layout model capable of predicting two distinct layout types. One stops at ambiguous regions, while the other extends to encompass all visible areas. Our model employs two global context embeddings, where each embedding is designed to capture specific contextual information for each layout type. With our novel feature guidance module, the image feature retrieves relevant context from these embeddings, generating layout-aware features for precise bi-layout predictions. A unique property of our Bi-Layout model is its ability to inherently detect ambiguous regions by comparing the two predictions. To circumvent the need for manual correction of ambiguous annotations during testing, we also introduce a new metric for disambiguating ground truth layouts. Our method demonstrates superior performance on benchmark datasets, notably outperforming leading approaches. Specifically, on the MatterportLayout dataset, it improves 3DIoU from 81.70% to 82.57% across the full test set and notably from 54.80% to 59.97% in subsets with significant ambiguity. Yu-Ju Tsai, Jin-Cheng Jhang, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo, Ming-Hsuan Yang 0001 |
CVPR | 7 |
| 2024 | GenRC: Generative 3D Room Completion from Sparse Image Collections
Ming-Feng Li, Yueh-Feng Ku, Hong-Xuan Yen, Yu-Lun Liu 0001, Albert Chen 0001, Cheng-Hao Kuo, Min Sun 0001 |
ECCV (37) | 7 |
| 2024 | Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding
Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Min Sun 0001, Mohammad Soleymani 0001 |
ECCV (80) | 4 |
| 2024 | Correspondence-Free SE(3) Point Cloud Registration in RKHS via Unsupervised Equivariant Learning
Ray Zhang 0001, Zheming Zhou, Min Sun 0001, Omid Ghasemalizadeh, Cheng-Hao Kuo, Ryan M. Eustice, Maani Ghaffari Jadidi, Arnie Sen |
ECCV (88) | 5 |
| 2024 | PoCo: Point Context Cluster for RGBD Indoor Place RecognitionabstractWe present a novel end-to-end algorithm (PoCo) for the indoor RGB-D place recognition task, aimed at identifying the most likely match for a given query frame within a reference database. The task presents inherent challenges attributed to the constrained field of view and limited range of perception sensors. We propose a new network architecture, which generalizes the recent Context of Clusters (CoCs) to extract global descriptors directly from the noisy point clouds through end-to-end learning. Moreover, we develop the architecture by integrating both color and geometric modalities into the point features to enhance the global descriptor representation. We conducted evaluations on public datasets ScanNet-PR and ARKit with 807 and 5047 scenarios, respectively. PoCo achieves SOTA performance: on ScanNet-PR, we achieve R@1 of 64.63%, a 5.7% improvement from the best-published result CGis (61.12%); on Arkit, we achieve R@1 of 45.12%, a 13.3% improvement from the best-published result CGis (39.82%). In addition, PoCo shows higher efficiency than CGis in inference time (1.75X-faster), and we demonstrate the effectiveness of PoCo in recognizing places within a real-world laboratory environment. Video: https://youtu.be/D8dObAeMiCw; Jing Liang 0006, Zhuo Deng 0002, Zheming Zhou, Omid Ghasemalizadeh, Dinesh Manocha, Min Sun 0001, Cheng-Hao Kuo, Arnie Sen |
IROS | 7 |
| 2024 | ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain AdaptationabstractLarge-scale pre-trained vision-language models (VLM) such as CLIP [32] have demonstrated noteworthy zero-shot classification capability, achieving 76.3% top-1 accuracy on ImageNet without seeing any examples. However, while applying CLIP to a downstream target domain, the presence of visual and text domain gaps and cross-modality misalignment can greatly impact the model performance. To address such challenges, we propose ReCLIP, a novel source-free domain adaptation method for VLMs, which does not require any source data or target labeled data. ReCLIP first learns a projection space to mitigate the misaligned visual-text embeddings and learns pseudo labels. Then, it deploys cross-modality self-training with the pseudo labels to update visual and text encoders, refine labels and reduce domain gaps and misalignment iteratively. With extensive experiments, we show that ReCLIP outperforms all the baselines significantly and improves the average accuracy of CLIP from 69.83% to 74.94% on 22 image classification benchmarks. Xuefeng Hu, Ke Zhang 0028, Albert Chen 0001, Jiajia Luo, Yuyin Sun, Ken Wang, Nan Qiao 0009, Min Sun 0001, Cheng-Hao Kuo, Ramakant Nevatia |
WACV | 11 |
| 2023 | Bidirectional Alignment for Domain Adaptive Detection with TransformersabstractWe propose a Bidirectional Alignment for domain adaptive Detection with Transformers (BiADT) to improve cross domain object detection performance. Existing adversarial learning based methods use gradient reverse layer (GRL) to reduce the domain gap between the source and target domains in feature representations. Since different image parts and objects may exhibit various degrees of domain-specific characteristics, directly applying GRL on a global image or object representation may not be suitable. Our proposed BiADT explicitly estimates token-wise domain-invariant and domain-specific features in the image and object token sequences. BiADT has a novel deformable attention and self-attention, aimed at bi-directional domain alignment and mutual information minimization. These two objectives reduce the domain gap in domain-invariant representations, and simultaneously increase the distinctiveness of domain-specific features. Our experiments show that BiADT achieves very competitive performance to SOTA consistently on Cityscapes-to-FoggyCityscapes, Sim10K-to-Citiscapes and Cityscapes-to-BDD100K, outperforming the strong baseline, AQT, by 2.0, 2.1, and 2.4 in mAP50, respectively. The implementation is available at https://github.com/helq2612/biADT Liqiang He, Albert Chen 0001, Min Sun 0001, Cheng-Hao Kuo, Sinisa Todorovic |
ICCV | 5 |
| 2023 | ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object DetectionabstractWe propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-view images to alleviate the confusion arising from voxels of free space, and during the inference phase, only images from multiple views are required. Besides, a powerful pre-trained 2D feature extractor can be leveraged by our representation, leading to a more robust performance. To evaluate the effectiveness of ImGeoNet, we conduct quantitative and qualitative experiments on three indoor datasets, namely ARKitScenes, ScanNetV2, and ScanNet200. The results demonstrate that ImGeoNet outperforms the current state-of-the-art multiview image-based method, ImVoxelNet, on all three datasets in terms of detection accuracy. In addition, ImGeoNet shows great data efficiency by achieving results comparable to ImVoxelNet with 100 views while utilizing only 40 views. Furthermore, our studies indicate that our proposed image-induced geometry-aware representation can enable image-based methods to attain superior detection accuracy than the seminal point cloud-based method, VoteNet, in two practical scenarios: (1) scenarios where point clouds are sparse and noisy, such as in ARKitScenes, and (2) scenarios involve diverse object classes, particularly classes of small objects, as in the case in ScanNet200. Project page: https://ttaoretw.github.io/imgeonet. Tao Tu 0002, Shun-Po Chuang, Yu-Lun Liu 0001, Cheng Sun 0004, Ke Zhang 0028, Donna Roy, Cheng-Hao Kuo, Min Sun 0001 |
ICCV | 7 |
| 2023 | Multimodal Neural Radiance FieldabstractThis paper addresses the challenge of reconstructing a scene with a neural radiance field (NeRF) for robot vision and scene understanding using multiple modalities. Researchers have introduced the use of NeRF to represent an object for synthesizing and rendering novel views of complex scenes by optimizing a 3-D radiance field for ray casting and rendering for 2-D RGB images. However, using RGB images alone introduces additional geometry ambiguities with transparent objects or complex scenes and cannot accurately depict the 3-D shapes. We discuss and solve this problem and use multiple modalities as input for the same NeRF model to build a multimodal NeRF by incorporating point clouds and infrared image supervision to prevent such bias. In contrast to RGB images, infrared images and point clouds are typically taken by separate cameras that cannot be aligned with the RGB camera. We further introduce the alignment of different modalities based on point cloud registration to estimate the relative transformation matrices between them before training a NeRF model with multiple modalities. We evaluate our model on chosen scenes from the ScanNet and M2DGR datasets and demonstrate that it outperforms existing state-of-the-art methods. Haidong Zhu, Yuyin Sun, Jiajia Luo, Nan Qiao 0009, Ramakant Nevatia, Cheng-Hao Kuo |
ICRA | 8 |
| 2023 | SAAML: A Framework for Semi-supervised Affective Adaptation via Metric LearningabstractSocially intelligent systems such as home robots should be able to perceive emotions and social behaviors. Affect recognition datasets have limited labeled data, and existing large unlabeled datasets, e.g., VoxCeleb2, suitable for pre-training, mostly contain neutral expressions, limiting their application to affective downstream tasks. We introduce a novel Semi-supervised Affective Adaptation framework via Metric Learning (SAAML) to adapt pre-trained audiovisual models (e.g., AV-HuBERT) to expressive behaviors associated with emotions and social communication. The proposed framework automatically retrieves a large number of emotional excerpts (>100 hours) from the VoxCeleb2 dataset via metric learning from two emotion recognition datasets (MSP-IMPROV and CREMA-D), and learns domain-invariant emotion-aware representations. Experimental results show that fine-tuning the proposed affect-aware AV-HuBERT (AW-HuBERT) improves the emotion recognition accuracy by 3-6% compared to fine-tuning the original pre-trained models. We further validate the effectiveness of the AW-HuBERT on human-centered visual understanding tasks, namely, facial expression recognition, video highlight detection, and continuous emotion recognition. The proposed approach consistently outperforms AV-HuBERT and delivers competitive performance compared to the existing methods. With this work, we demonstrate the effectiveness of adaptive pre-training for existing models on domain-specific data to enhance their performance for human-centered tasks. Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Mohammad Soleymani 0001 |
ACM Multimedia | 4 |
| 2023 | Human-in-the-Loop Video Semantic Segmentation Auto-AnnotationabstractAccurate per-pixel semantic class annotations of the entire video are crucial for designing and evaluating video semantic segmentation algorithms. However, the annotations are usually limited to a small subset of the video frames due to the high annotation cost and limited budget in practice. In this paper, we propose a novel human-in-the-loop framework called HVSA to generate semantic segmentation annotations for the entire video using only a small annotation budget. Our method alternates between active sample selection and test-time fine-tuning algorithms until annotation quality is satisfied. In particular, the active sample selection algorithm picks the most important samples to get manual annotations, where the sample can be a video frame, a rectangle, or even a super-pixel. Further, the test-time fine-tuning algorithm propagates the manual annotations of selected samples to the entire video. Real-world experiments show that our method generates highly accurate and consistent semantic segmentation annotations while simultaneously enjoys significantly small annotation cost. Nan Qiao 0009, Yuyin Sun, Chong Liu 0007, Jiajia Luo, Ke Zhang 0028, Cheng-Hao Kuo |
WACV | 7 |
| 2023 | CameraPose: Weakly-Supervised Monocular 3D Human Pose Estimation by Leveraging In-the-wild 2D AnnotationsabstractTo improve the generalization of 3D human pose estimators, many existing deep learning based models focus on adding different augmentations to training poses. However, data augmentation techniques are limited to the "seen" pose combinations and hard to infer poses with rare "unseen" joint positions. To address this problem, we present CameraPose, a weakly-supervised framework for 3D human pose estimation from a single image, which can not only be applied on 2D-3D pose pairs but also on 2D alone annotations. By adding a camera parameter branch, any in-the-wild 2D annotations can be fed into our pipeline to boost the training diversity and the 3D poses can be implicitly learned by reprojecting back to 2D. Moreover, CameraPose introduces a refinement network module with confidence-guided loss to further improve the quality of noisy 2D keypoints extracted by 2D pose estimators. Experimental results demonstrate that the CameraPose brings in clear improvements on cross-scenario datasets. Notably, it outperforms the baseline method by 3mm on the most challenging dataset 3DPW. In addition, by combining our proposed refinement network module with existing 3D pose estimators, their performance can be improved in cross-scenario evaluation. Cheng-Yen Yang, Jiajia Luo, Yuyin Sun, Nan Qiao 0009, Ke Zhang 0028, Zhongyu Jiang, Jenq-Neng Hwang, Cheng-Hao Kuo |
WACV | 9 |
| 2022 | Learning Feature Decomposition for Domain Adaptive Monocular Depth EstimationabstractMonocular depth estimation (MDE) has attracted intense study due to its low cost and critical functions for robotic tasks such as localization, mapping and obstacle detection. Supervised approaches have led to great success with the advance of deep learning, but they rely on large quantities of ground-truth depth annotations that are expensive to acquire. Unsupervised domain adaptation (UDA) transfers knowledge from labeled source data to unlabeled target data, so as to relax the constraint of supervised learning. However, existing UDA approaches may not completely align the domain gap across different datasets because of the domain shift problem. We believe better domain alignment can be achieved via well-designed feature decomposition. In this paper, we propose a novel UDA method for MDE, referred to as Learning Feature Decomposition for Adaptation (LFDA), which learns to decompose the feature space into content and style components. LFDA only attempts to align the content component since it has a smaller domain gap. Meanwhile, it excludes the style component which is specific to the source domain from training the primary task. Furthermore, LFDA uses separate feature distribution estimations to further bridge the domain gap. Extensive experiments on three domain adaptative MDE scenarios show that the proposed method achieves superior accuracy and lower computational cost compared to the state-of-the-art approaches. Shao-Yuan Lo, Jim Thomas 0001, Vishal M. Patel, Cheng-Hao Kuo |
IROS | 6 |
| 2022 | A polar-edge context-aware (PECA) network for mirror segmentation
Liqiang He, Jiajia Luo, Ke Zhang 0028, Yuyin Sun, Nan Qiao 0009, Cheng-Hao Kuo, Sinisa Todorovic |
Image Vis. Comput. | 7 |
| 2021 | Acted vs. Improvised: Domain Adaptation for Elicitation Approaches in Audio-Visual Emotion RecognitionabstractKey challenges in developing generalized automatic emotion recognition systems include scarcity of labeled data and lack of gold-standard references. Even for the cues that are labeled as the same emotion category, the variability of associated expressions can be high depending on the elicitation context e.g., emotion elicited during improvised conversations vs. acted sessions with predefined scripts. In this work, we regard the emotion elicitation approach as domain knowledge, and explore domain transfer learning techniques on emotional utterances collected under different emotion elicitation approaches, particularly with limited labeled target samples. Our emotion recognition model combines the gradient reversal technique with an entropy loss function as well as the softlabel loss, and the experiment results show that domain transfer learning methods can be employed to alleviate the domain mismatch between different elicitation approaches. Our work provides new insights into emotion data collection, particularly the impact of its elicitation strategies, and the importance of domain adaptation in emotion recognition aiming for generalized systems. Yelin Kim, Cheng-Hao Kuo, Shri Narayanan |
Interspeech | 3 |
| 2020 | MEBOW: Monocular Estimation of Body Orientation in the WildabstractBody orientation estimation provides crucial visual cues in many applications, including robotics and autonomous driving. It is particularly desirable when 3-D pose estimation is difficult to infer due to poor image resolution, occlusion or indistinguishable body parts. We present COCO-MEBOW (Monocular Estimation of Body Orientation in the Wild), a new large-scale dataset for orientation estimation from a single in-the-wild image. The body-orientation labels for around 130K human bodies within 55K images from the COCO dataset have been collected using an efficient and high-precision annotation pipeline. We also validated the benefits of the dataset. First, we show that our dataset can substantially improve the performance and the robustness of a human body orientation estimation model, the development of which was previously limited by the scale and diversity of the available training data. Additionally, we present a novel triple-source solution for 3-D human pose estimation, where 3-D pose labels, 2-D pose labels, and our body-orientation labels are all used in joint training. Our model significantly outperforms state-of-the-art dual-source solutions for monocular 3-D human pose estimation, where training only uses 3-D pose labels and 2-D pose labels. This substantiates an important advantage of MEBOW for 3-D human pose estimation, which is particularly appealing because the per-instance labeling cost for body orientations is far less than that for 3-D poses. The work demonstrates high potential of MEBOW in addressing real-world challenges involving understanding human behaviors. Further information of this work is available at https://chenyanwu.github.io/MEBOW/. Chenyan Wu, Jiajia Luo, Che-Chun Su, Anuja Dawane, Bikramjot Hanzra, Bilan Liu, James Z. Wang 0001, Cheng-Hao Kuo |
CVPR | 10 |
| 2013 | Person re-identification using semantic color names and RankBoostabstractWe address the problem of appearance-based person re-identification, which has been drawing an increasing amount of attention in computer vision. It is a very challenging task since the visual appearance of a person can change dramatically due to different backgrounds, camera characteristics, lighting conditions, view-points, and human poses. Among the recent studies on person re-id, color information plays a major role in terms of performance. Traditional color information like color histogram, however, still has much room to improve. We propose to apply semantic color names to describe a person image, and compute probability distribution on those basic color terms as image descriptors. To be better combined with other features, we define our appearance affinity model as linear combination of similarity measurements of corresponding local descriptors, and apply the RankBoost algorithm to find the optimal weights for the similarity measurements. We evaluate our proposed system on the highly challenging VIPeR dataset, and show improvements over the state-of-the-art methods in terms of widely used person re-id evaluation metrics. Cheng-Hao Kuo, Sameh Khamis, Vinay D. Shet |
WACV | 1 |
| 2011 | How does person identity recognition help multi-person tracking?abstractWe address the problem of multi-person tracking in a complex scene from a single camera. Although tracklet-association methods have shown impressive results in several challenging datasets, discriminability of the appearance model remains a limitation. Inspired by the work of person identity recognition, we obtain discriminative appearance-based affinity models by a novel framework to incorporate the merits of person identity recognition, which help multi-person tracking performance. During off-line learning, a small set of local image descriptors is selected to be used in on-line learned appearances-based affinity models effectively and efficiently. Given short but reliable track-lets generated by frame-to-frame association of detection responses, we identify them as query tracklets and gallery tracklets. For each gallery tracklet, a target-specific appearance model is learned from the on-line training samples collected by spatio-temporal constraints. Both gallery tracklets and query tracklets are fed into hierarchical association framework to obtain final tracking results. We evaluate our proposed system on several public datasets and show significant improvements in terms of tracking evaluation metrics. Cheng-Hao Kuo, Ramakant Nevatia |
CVPR | 1 |
| 2010 | Multi-target tracking by on-line learned discriminative appearance modelsabstractWe present an approach for online learning of discriminative appearance models for robust multi-target tracking in a crowded scene from a single camera. Although much progress has been made in developing methods for optimal data association, there has been comparatively less work on the appearance models, which are key elements for good performance. Many previous methods either use simple features such as color histograms, or focus on the discriminability between a target and the background which does not resolve ambiguities between the different targets. We propose an algorithm for learning a discriminative appearance model for different targets. Training samples are collected online from tracklets within a time sliding window based on some spatial-temporal constraints; this allows the models to adapt to target instances. Learning uses an Ad-aBoost algorithm that combines effective image descriptors and their corresponding similarity measurements. We term the learned models as OLDAMs. Our evaluations indicate that OLDAMs have significantly higher discrimination between different targets than conventional holistic color histograms, and when integrated into a hierarchical association framework, they help improve the tracking accuracy, particularly reducing the false alarms and identity switches. Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
CVPR | 1 |
| 2010 | Inter-camera Association of Multi-target Tracks by On-Line Learned Appearance Affinity Models
Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
ECCV (1) | 1 |
| 2009 | Robust multi-view car detection using unsupervised sub-categorizationabstractThis paper presents a novel approach for multi-view car detection using unsupervised sub-categorization instead of manual labeling. Cars have large variability of models and the view-point makes the appearance change dramatically. For object classes with a large intra-class variation like cars, a divide-and-conquer strategy may be applied. Instead of using manually predefined intra-class sub-categorization, we examine several non-linear dimension reduction methods and group samples in the low-dimension embedding in an unsupervised way. The clustered samples have strong view-point similarities internally. A boosting-based cascade tree classifier is trained based on these sub-categorizations. To demonstrate the capability of our multi-view car detector, we create a more challenging test set with annotations. Compared to the UIUC side-view car data set, our test set contains a large range of car models, view points, and complex backgrounds. We compare our approach with previous methods and the result shows that ours outperforms the state-of-the-art methods. Cheng-Hao Kuo, Ramakant Nevatia |
WACV | 1 |
| 2009 | Identifying Noncooperative Subjects at a Distance Using Face Images and Inferred Three-Dimensional Face ModelsabstractWe present an approach to identify noncooperative individuals at a distance from a sequence of images, using 3-D face models. Most biometric features (such as fingerprints, hand shape, iris, or retinal scans) require cooperative subjects in close proximity to the biometric system. We process images acquired with an ultrahigh-resolution video camera, infer the location of the subjects' head, use this information to crop the region of interest, build a 3-D face model, and use this 3-D model to perform biometric identification. To build the 3-D model, we use an image sequence, as natural head and body motion provides enough viewpoint variation to perform stereomotion for 3-D face reconstruction. We have conducted experiments on a 2-D and 3-D databases collected in our laboratory. First, we found that metric 3-D face models can be used for recognition by using simple scaling method even though there is no exact scale in the 3-D reconstruction. Second, experiments using a commercial 3-D matching engine suggest the feasibility of the proposed approach for recognition against 3-D galleries at a distance (3, 6, and 9 m). Moreover, we show initial 3-D face modeling results on various factors including head motion, outdoor lighting conditions, and glasses. The evaluation results suggest that video data alone, at a distance of 3 to 9 meters, can provide a 3-D face shape that supports successful face recognition. The performance of 3-D-3-D recognition with the currently generated models does not quite match that of 2-D-2-D. We attribute this to the quality of the inferred models, and this suggests a clear path for future research. Gérard G. Medioni, Jongmoo Choi, Cheng-Hao Kuo, Douglas Fidaleo |
IEEE Trans. Syst. Man Cybern. Part A | 3 |