EDBT 2026 Demo / reviewers in the wild / expert
Kyungdon Joo
dblp:138/0123
· DBLP profile ↗
42ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-3920-9608ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 16 since 2021Systems, architecture and hardware · 7 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Highlights: Video Retrieval with Salient and Surrounding Contexts
Jae Hun Bang, Moon Ye-Bin, Tae-Hyun Oh, Kyungdon Joo |
WACV | 4 |
| 2026 | AnyBald: Toward Realistic Diffusion-Based Hair Removal In-The-WildabstractWe present AnyBald, a novel framework for realistic hair removal from portrait images captured under diverse in-the-wild conditions. One of the key challenges in this task is the lack of high-quality paired data, as existing datasets are often low-quality, with limited viewpoint variation and diversity, making it difficult to handle real-world cases. To address this, we construct a scalable data augmentation pipeline that synthesizes high-quality hair and non-hair image pairs across diverse real-world scenarios, enabling effective generalization and scalable supervision. With this enriched dataset, we present a new hair removal framework that reformulates pretrained latent diffusion in-painting using learnable text prompts, removing the need for explicit masking at inference. In doing so, our model achieves natural hair removal with semantic preservation via implicit localization. To further enhance spatial precision, we introduce a regularization loss that guides the model to attend specifically to hair regions. Extensive experiments demonstrate that AnyBald outperforms in removing hair while preserving identity and semantics across various in-the-wild domains. Our project page is here: https://vision3d-lab.github.io/anybald/. Yongjun Choi, Seungoh Han, Soomin Kim 0006, Sumin Son, Mohsen Rohani, Edgar Maucourant, Dongbo Min, Kyungdon Joo |
WACV | 8 |
| 2026 | Pose-Diverse Multi-View Virtual Try-on from a Single Frontal Image via Diffusion TransformerabstractWe study multi-view virtual try-on framework from a single reference image of themselves and a single frontal image of a garment. While most existing approaches focus on single-view synthesis, their reliance on a single, fixed viewpoint limits their application in immersive environments that require diverse poses and viewpoints. The ability to generate a multi-view virtual try-on image is crucial for a comprehensive user experience, as if the user to inspect the garment from multiple views, including the back and sides, providing a similar experience to a real fitting room. In this paper, we propose a novel framework for pose-controllable, multi-view virtual try-on from a single image. Our method incorporates the proposed cross attention injection for high-quality synthesis and allows for fine-grained pose control. Our framework consists of two stages: the first stage generates a clean frontal try-on result using an off-the-shelf model, and the second stage creates diverse viewpoints using a diffusion transformer. Additionally, an attention injection mechanism ensures consistent identity and garment preservation across synthesized views. Unlike conventional methods that require multiple images of the user or the garment from various angles, our model eliminates these constraints by synthesizing multi-view results from a single input image pair. Our method not only generates realistic images but also enables users to virtually inspect the fit and arrangement of the garment from multiple angles without the need for additional data. Our extensive experiments demonstrate that our framework outperforms existing 3D-based multi-view virtual try-on methods in terms of image quality and pose diversity. Seonghee Han, Minchang Chung, Gyeongsu Cho, Kyungdon Joo |
WACV | 4 |
| 2026 | LighthouseGS: Indoor Structure-aware 3D Gaussian Splatting for Panorama-Style Mobile CapturesabstractWe introduce LighthouseGS, a practical novel view synthesis framework based on 3D Gaussian Splatting that utilizes simple panorama-style captures from a single mobile device. While convenient, this rotation-dominant motion and narrow baseline make accurate camera pose and 3D point estimation challenging, especially in textureless indoor scenes. To address these challenges, LighthouseGSleverages rough geometric priors, such as mobile device camera poses and monocular depth estimation, and utilizes indoor planar structures. Specifically, we propose a new initialization method called plane scaffold assembly to generate consistent 3D points on these structures, followed by a stable pruning strategy to enhance geometry and optimization stability. Additionally, we present geometric and photometric corrections to resolve inconsistencies from motion drift and auto-exposure in mobile devices. Tested on real and synthetic indoor scenes, LighthouseGSdelivers photorealistic rendering, outperforming state-of-the-art methods and enabling applications like panoramic view synthesis and object placement. Project page: https://vision3d-lab.github.io/lighthousegs/ Seungoh Han, Jaehoon Jang 0001, Hyunsu Kim, Jaeheung Surh, Junhyung Kwak, Hyowon Ha, Kyungdon Joo |
WACV | 7 |
| 2026 | Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion ModelabstractDance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music-synchronized characteristics of dance. In this paper, we propose MambaDance, a new dance generation approach that leverages a Mamba-based diffusion model. Mamba, well-suited to handling long and autoregressive sequences, is integrated into our two-stage diffusion architecture, substituting off-the-shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian-based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at https://vision3d-lab.github.io/mambadance. Sangjune Park, Inhyeok Choi, Donghyeon Soon, Youngwoo Jeon, Kyungdon Joo |
WACV | 5 |
| 2025 | HUSH: Holistic Panoramic 3D Scene Understanding using Spherical HarmonicsabstractMotivated by the efficiency of spherical harmonics (SH) in representing various physical phenomena, we propose a Holistic panoramic 3D scene Understanding framework using Spherical Harmonics, dubbed as HUSH. Our approach focuses on a unified framework adaptable to various 3D scene understanding tasks via SH bases. To achieve this, we first estimate SH coefficients, allowing for the adaptive configuration of the SH bases specific to each scene. HUSH then employs a hierarchical attention module that uses SH bases as queries to generate comprehensive scene features by integrating these scene-adaptive SH bases with image features. Additionally, we introduce an SH basis index module that adaptively emphasizes relevant SH bases to produce task-relevant features, enhancing the versatility of HUSH across different scene understanding tasks. Finally, by combining the scene features with task-relevant features in the task-specific heads, we perform various scene understanding tasks, including depth, surface normal and room layout estimation. Experiments demonstrate that HUSH achieves state-of-the-art performance on depth estimation benchmarks, highlighting the robustness and scalability of using SH in panoramic 3D scene understanding. For more information, you can visit our project page https://vision3d-lab.github.io/hush/. Jongsung Lee 0007, Harin Park, Byeong-Uk Lee, Kyungdon Joo |
CVPR | 4 |
| 2025 | VPOcc: Exploiting Vanishing Point for 3D Semantic Occupancy PredictionabstractUnderstanding 3D scenes semantically and spatially is crucial for the safe navigation of robots and autonomous vehicles, aiding obstacle avoidance and accurate trajectory planning. Camera-based 3D semantic occupancy prediction, which infers complete voxel grids from 2D images, is gaining importance in robot vision for its resource efficiency compared to 3D sensors. However, this task inherently suffers from a 2D–3D discrepancy, where objects of the same size in 3D space appear at different scales in a 2D image depending on their distance from the camera due to perspective projection. To tackle this issue, we propose a novel framework called VPOcc that leverages a vanishing point (VP) to mitigate the 2D-3D discrepancy at both the pixel and feature levels. As a pixel-level solution, we introduce a VPZoomer module, which warps images by counteracting the perspective effect using a VP-based homography transformation. In addition, as a feature-level solution, we propose a VP-guided cross-attention (VPCA) module that performs perspective-aware feature aggregation, utilizing 2D image features that are more suitable for 3D space. Lastly, we integrate two feature volumes extracted from the original and warped images to compensate for each other through a spatial volume fusion (SVF) module. By effectively incorporating VP into the network, our framework achieves improvements in both IoU and mIoU metrics on SemanticKITTI and SSCBench-KITTI360 datasets. Additional details are available at https://vision3d-lab.github.io/vpocc/. Junsu Kim 0001, Ukcheol Shin, Jean Oh, Kyungdon Joo |
IROS | 5 |
| 2025 | Rigidity-Aware 3D Gaussian Deformation from a Single ImageabstractReconstructing object deformation from a single image remains a significant challenge in computer vision and graphics. Existing methods typically rely on multi-view video to recover deformation, limiting their applicability under constrained scenarios. To address this, we propose DeformSplat, a novel framework that effectively guides 3D Gaussian deformation from only a single image. Our method introduces two main technical contributions. First, we present Gaussian-to-Pixel Matching which bridges the domain gap between 3D Gaussian representations and 2D pixel observations. This enables robust deformation guidance from sparse visual cues. Second, we propose Rigid Part Segmentation consisting of initialization and refinement. This segmentation explicitly identifies rigid regions, crucial for maintaining geometric coherence during deformation. By combining these two techniques, our approach can reconstruct consistent deformations from a single image. Extensive experiments demonstrate that our approach significantly outperforms existing methods and naturally extends to various applications, such as frame interpolation and interactive object manipulation. Project page : https://vision3d-lab.github.io/deformsplat Jinhyeok Kim, Jae Hun Bang, Seunghyun Seo, Kyungdon Joo |
SIGGRAPH Asia | 4 |
| 2025 | DogRecon: Canine Prior-Guided Animatable 3D Gaussian Dog Reconstruction From A Single ImageabstractWe tackle animatable 3D dog reconstruction from a single image, noting the overlooked potential of animals. Particularly, we focus on dogs, emphasizing their intrinsic characteristics that complicate 3D observation. First, the considerable variation in shapes across breeds presents a complexity for modeling. Additionally, the nature of quadrupeds leads to frequent joint occlusions compared to humans. These challenges make 3D reconstruction from 2D observations difficult, and it becomes dramatically harder when constrained to a single image. To address these challenges, our insight is to combine the acquisition of appearance from generative models, without additional data, with geometric guidance provided by a parametric representation, aiming to achieve complete geometry. To this end, we present DogRecon, our framework consists of two key components: Canine-centric novel view synthesis with canine prior for multi-view generation of dog and a reliable sampling weight strategy with Gaussian Splatting for animatable 3D dog reconstruction. Extensive experiments on the GART, DFA, and internet-sourced datasets confirm our framework has state-of-the-art performance in image-to-3D generation and comparable performance in animatable 3D reconstruction. Additionally, we demonstrate novel pose animation and text-to-3D dog reconstruction as applications. Project page: https://vision3d-lab.github.io/dogrecon/ Gyeongsu Cho, Changwoo Kang, Donghyeon Soon, Kyungdon Joo |
Int. J. Comput. Vis. | 4 |
| 2024 | ContactGen: Contact-Guided Interactive 3D Human Generation for PartnersabstractAmong various interactions between humans, such as eye contact and gestures, physical interactions by contact can act as an essential moment in understanding human behaviors. Inspired by this fact, given a 3D partner human with the desired interaction label, we introduce a new task of 3D human generation in terms of physical contact. Unlike previous works of interacting with static objects or scenes, a given partner human can have diverse poses and different contact regions according to the type of interaction. To handle this challenge, we propose a novel method of generating interactive 3D humans for a given partner human based on a guided diffusion framework (ContactGen in short). Specifically, we newly present a contact prediction module that adaptively estimates potential contact regions between two input humans according to the interaction label. Using the estimated potential contact regions as complementary guidances, we dynamically enforce ContactGen to generate interactive 3D humans for a given partner human within a guided diffusion model. We demonstrate ContactGen on the CHI3D dataset, where our method generates physically plausible and diverse poses compared to comparison methods. Dongjun Gu, Jaehyeok Shim, Jaehoon Jang 0001, Changwoo Kang, Kyungdon Joo |
AAAI | 5 |
| 2024 | The Devil Is in the Details: Simple Remedies for Image-to-LiDAR Representation Learning
Wonjun Jo, Byung-Ki Kwon, Kim Ji-Yeon, Hawook Jeong, Kyungdon Joo, Tae-Hyun Oh |
ACCV (9) | 5 |
| 2024 | DITTO: Dual and Integrated Latent Topologies for Implicit 3D ReconstructionabstractWe propose a novel concept of dual and integrated latent topologies (D ITto in short) for implicit 3D reconstruction from noisy and sparse point clouds. Most existing methods predominantly focus on single latent type, such as point or grid latents. In contrast, the proposed DITTO leverages both point and grid latents (i.e., dual latent) to enhance their strengths, the stability of grid latents and the detailrich capability of point latents. Concretely, DITTO consists of dual latent encoder and integrated implicit decoder. In the dual latent encoder, a dual latent layer, which is the key module block composing the encoder, refines both latents in parallel, maintaining their distinct shapes and enabling recursive interaction. Notably, a newly proposed dynamic sparse point transformer within the dual latent layer effectively refines point latents. Then, the integrated implicit decoder systematically combines these refined latents, achieving high-fidelity 3D reconstruction and surpassing previous state-of-the-art methods on object- and scene-level datasets, especially in thin and detailed structures. Jaehyeok Shim, Kyungdon Joo |
CVPR | 2 |
| 2024 | Fracture Assembly with Segmentation And Iterative RegistrationabstractReassembling broken fractures back to their original shape remains a complex challenge. While prior research has demonstrated impressive results in domain-specific assembly, these methods largely depend on human-designed structural priors or struggle with assembling diverse shapes. To tackle this issue, we introduce a new fracture assembly framework based on segmentation and iterative registration, so-called FRASIER. By finding broken regions of fractures by segmentation, FRASIERdramatically increases overlap region ratios between fractures, which allows us to align fractures by registration. In addition, we employ point cloud XOR and beam search to make our framework robust. Experiments demonstrate that FRASIERoutperforms state-of-the-art methods. Project page: https://frasier-assembly.github.io Jinhyeok Kim, Inha Lee, Kyungdon Joo |
ICASSP | 3 |
| 2023 | Pose-Guided 3D Human Generation in Indoor SceneabstractIn this work, we address the problem of scene-aware 3D human avatar generation based on human-scene interactions. In particular, we pay attention to the fact that physical contact between a 3D human and a scene (i.e., physical human-scene interactions) requires a geometrical alignment to generate natural 3D human avatar. Motivated by this fact, we present a new 3D human generation framework that considers geometric alignment on potential contact areas between 3D human avatars and their surroundings. In addition, we introduce a compact yet effective human pose classifier that classifies the human pose and provides potential contact areas of the 3D human avatar. It allows us to adaptively use geometric alignment loss according to the classified human pose. Compared to state-of-the-art method, our method can generate physically and semantically plausible 3D humans that interact naturally with 3D scenes without additional post-processing. In our evaluations, we achieve the improvements with more plausible interactions and more variety of poses than prior research in qualitative and quantitative analysis. Project page: https://bupyeonghealer.github.io/phin/. Changwoo Kang, Jeongin Park, Kyungdon Joo |
AAAI | 4 |
| 2023 | Diffusion-Based Signed Distance Fields for 3D Shape GenerationabstractWe propose a 3D shape generation framework (SDF-Diffusion in short) that uses denoising diffusion models with continuous 3D representation via signed distance fields (SDF). Unlike most existing methods that depend on discontinuous forms, such as point clouds, SDF-Diffusion generates high-resolution 3D shapes while alleviating memory issues by separating the generative process into two-stage: generation and super-resolution. In the first stage, a diffusion-based generative model generates a low-resolution SDF of 3D shapes. Using the estimated low-resolution SDF as a condition, the second stage diffusion model performs super-resolution to generate high-resolution SDF. Our framework can generate a high-fidelity 3D shape despite the extreme spatial complexity. On the ShapeNet dataset, our model shows competitive performance to the state-of-the-art methods and shows applicability on the shape completion task without modification. Jaehyeok Shim, Changwoo Kang, Kyungdon Joo |
CVPR | 3 |
| 2023 | SlaBins: Fisheye Depth Estimation using Slanted Bins on Road EnvironmentsabstractAlthough 3D perception for autonomous vehicles has focused on frontal-view information, more than half of fatal accidents occur due to side impacts in practice (e.g., T-bone crash). Motivated by this fact, we investigate the problem of side-view depth estimation, especially for monocular fisheye cameras, which provide wide FoV information. However, since fisheye cameras head road areas, it observes road areas mostly and results in severe distortion on object areas, such as vehicles or pedestrians. To alleviate these issues, we propose a new fisheye depth estimation network, SlaBins, that infers an accurate and dense depth map based on a geometric property of road environments; most objects are standing (i.e., orthogonal) on the road environments. Concretely, we introduce a slanted multi-cylindrical image (MCI) representation, which allows us to describe a distance as a radius to a cylindrical layer orthogonal to the ground regardless of the camera viewing direction. Based on the slanted MCI, we estimate a set of adaptive bins and a per-pixel probability map for depth estimation. Then by combining it with the estimated slanted angle of viewing direction, we directly infer a dense and accurate depth map for fisheye cameras. Experiments demonstrate that SlaBins outperforms the state-of-the-art methods in both qualitative and quantitative evaluation on the SynWoodScape and KITTI-360 depth datasets. For more information, you can visit our project page https://syniez.github.io/SlaBins/. Jongsung Lee 0007, Gyeongsu Cho, Jeongin Park, Kyongjun Kim, Seongoh Lee, Jung Hee Kim 0001, Seong-Gyun Jeong, Kyungdon Joo |
ICCV | 8 |
| 2023 | Hong Kong World: Leveraging Structural Regularity for Line-Based SLAMabstractManhattan and Atlanta worlds hold for the structured scenes with only vertical and horizontal dominant directions (DDs). To describe the scenes with additional sloping DDs, a mixture of independent Manhattan worlds seems plausible, but may lead to unaligned and unrelated DDs. By contrast, we propose a novel structural model called Hong Kong world. It is more general than Manhattan and Atlanta worlds since it can represent the environments with slopes, e.g., a city with hilly terrain, a house with sloping roof, and a loft apartment with staircase. Moreover, it is more compact and accurate than a mixture of independent Manhattan worlds by enforcing the orthogonality constraints between not only vertical and horizontal DDs, but also horizontal and sloping DDs. We further leverage the structural regularity of Hong Kong world for the line-based SLAM. Our SLAM method is reliable thanks to three technical novelties. First, we estimate DDs/vanishing points in Hong Kong world in a semi-searching way. We use a new consensus voting strategy for search, instead of traditional branch and bound. This method is the first one that can simultaneously determine the number of DDs, and achieve quasi-global optimality in terms of the number of inliers. Second, we compute the camera pose by exploiting the spatial relations between DDs in Hong Kong world. This method generates concise polynomials, and thus is more accurate and efficient than existing approaches designed for unstructured scenes. Third, we refine the estimated DDs in Hong Kong world by a novel filter-based method. Then we use these refined DDs to optimize the camera poses and 3D lines, leading to higher accuracy and robustness than existing optimization algorithms. In addition, we establish the first dataset of sequential images in Hong Kong world. Experiments showed that our approach outperforms state-of-the-art methods in terms of accuracy and/or efficiency. Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Pyojin Kim, Kyungdon Joo, Zhenjun Zhao, Yun-Hui Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Real-Time Multi-Car Localization and See-Through System
François Rameau, Oleksandr Bailo, Jinsun Park, Kyungdon Joo, In-So Kweon |
Int. J. Comput. Vis. | 4 |
| 2022 | Linear RGB-D SLAM for Structured EnvironmentsabstractWe propose a new linear RGB-D simultaneous localization and mapping (SLAM) formulation by utilizing planar features of the structured environments. The key idea is to understand a given structured scene and exploit its structural regularities such as the Manhattan world. This understanding allows us to decouple the camera rotation by tracking structural regularities, which makes SLAM problems free from being highly nonlinear. Additionally, it provides a simple yet effective cue for representing planar features, which leads to a linear SLAM formulation. Given an accurate camera rotation, we jointly estimate the camera translation and planar landmarks in the global planar map using a linear Kalman filter. Our linear SLAM method, called L-SLAM, can understand not only the Manhattan world but the more general scenario of the Atlanta world, which consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction. To this end, we introduce a novel tracking-by-detection scheme that infers the underlying scene structure by Atlanta representation. With efficient Atlanta representation, we formulate a unified linear SLAM framework for structured environments. We evaluate L-SLAM on a synthetic dataset and RGB-D benchmarks, demonstrating comparable performance to other state-of-the-art SLAM methods without using expensive nonlinear optimization. We assess the accuracy of L-SLAM on a practical application of augmented reality. Kyungdon Joo, Pyojin Kim, Martial Hebert, In-So Kweon, H. Jin Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Robust and Efficient Estimation of Relative Pose for Cameras on Selfie SticksabstractTaking selfies has become one of the major photographic trends of our time. In this study, we focus on the selfie stick, on which a camera is mounted to take selfies. We observe that a camera on a selfie stick typically travels through a particular type of trajectory around a sphere. Based on this finding, we propose a robust, efficient, and optimal estimation method for relative camera pose between two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of camera motion constrained by a selfie stick and define this motion as spherical joint motion. Utilizing a novel parametrization and calibration scheme, we demonstrate that the pose estimation problem can be reduced to a 3-degrees of freedom (DoF) search problem, instead of a generic 6-DoF problem. This facilitates the derivation of an efficient branch-and-bound optimization method that guarantees a global optimal solution, even in the presence of outliers. Furthermore, as a simplified case of spherical joint motion, we introduce selfie motion, which has a fewer number of DoF than spherical joint motion. We validate the performance and guaranteed optimality of our method on both synthetic and real-world data. Additionally, we demonstrate the applicability of the proposed method for two applications: refocusing and stylization. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Unified 3D Mesh Recovery of Humans and Animals by Learning Animal Exercise
Kim Youwang, Kyungdon Joo, Tae-Hyun Oh |
BMVC | 3 |
| 2021 | Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point EstimationabstractExisting vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts a spherical probability map of VP. Based on this map, we can detect all the VPs. Our method is reliable thanks to four technical novelties. First, we leverage the icosahedral spherical representation to express our probability map. This representation provides uniform pixel distribution, and thus facilitates estimating arbitrary positions of VPs. Second, we design a loss function that enforces the antipodal symmetry and sparsity of our spherical probability map to prevent over-fitting. Third, we generate the ground truth probability map that reasonably expresses the locations and uncertainties of VPs. This map unnecessarily peaks at noisy annotated VPs, and also exhibits various anisotropic dispersions. Fourth, given a predicted probability map, we detect VPs by fitting a Bingham mixture model. This strategy can robustly handle close VPs and provide the confidence level of VP useful for practical applications. Experiments showed that our method achieves the best compromise between generality, accuracy, and efficiency, compared with state-of-the-art approaches. Haoang Li, Kai Chen 0028, Pyojin Kim, Kuk-Jin Yoon, Zhe Liu 0022, Kyungdon Joo, Yun-Hui Liu 0001 |
ICCV | 6 |
| 2021 | Stereo Object Matching NetworkabstractThis paper presents a stereo object matching method that exploits both 2D contextual information from images as well as 3D object-level information. Unlike existing stereo matching methods that exclusively focus on the pixel-level correspondence between stereo images within a volumetric space (i.e., cost volume), we exploit this volumetric structure in a different manner. The cost volume explicitly encompasses 3D information along its disparity axis, therefore it is a privileged structure that can encapsulate the 3D contextual information from objects. However, it is not straightforward since the disparity values map the 3D metric space in a non-linear fashion. Thus, we present two novel strategies to handle 3D objectness in the cost volume space: selective sampling (RoISelect) and 2D-3D fusion (fusion-by-occupancy), which allow us to seamlessly incorporate 3D object-level information and achieve accurate depth performance near the object boundary regions. Our depth estimation achieves competitive performance in the KITTI dataset and the Virtual-KITTI 2.0 dataset. Jaesung Choe, Kyungdon Joo, François Rameau, In-So Kweon |
ICRA | 2 |
| 2020 | Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World
Haoang Li, Pyojin Kim, Ji Zhao 0001, Kyungdon Joo, Zhe Liu 0022, Yun-Hui Liu 0001 |
ECCV (22) | 4 |
| 2020 | Non-local Spatial Propagation Network for Depth Completion
Jinsun Park, Kyungdon Joo, Chi-Kuei Liu, In-So Kweon |
ECCV (13) | 2 |
| 2020 | Globally Optimal Relative Pose Estimation for Camera on a Selfie StickabstractTaking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this scenario, we observe that the camera typically goes through a special trajectory along a sphere surface. Motivated by this observation, in this work, we propose an efficient and globally optimal relative camera pose estimation between a pair of two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of the camera motion constrained by a selfie stick and define its motion as spherical joint motion. By the new parametrization and calibration scheme, we show that the pose estimation problem can be reduced to a 3-DoF (degrees of freedom) search problem, instead of a generic 6-DoF problem. This allows us to derive a fast branch-and-bound global optimization, which guarantees a global optimum. Thereby, we achieve efficient and robust estimation even in the presence of outliers. By experiments on both synthetic and real-world data, we validate the performance as well as the guaranteed optimality of the proposed method. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, Yunsu Bok, In-So Kweon |
ICRA | 1 |
| 2020 | Linear RGB-D SLAM for Atlanta WorldabstractWe present a new linear method for RGB-D based simultaneous localization and mapping (SLAM). Compared to existing techniques relying on the Manhattan world assumption defined by three orthogonal directions, our approach is designed for the more general scenario of the Atlanta world. It consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction and thus can represent a wider range of scenes. Our approach leverages the structural regularity of the Atlanta world to decouple the non-linearity of camera pose estimations. This allows us separately to estimate the camera rotation and then the translation, which bypasses the inherent non-linearity of traditional SLAM techniques. To this end, we introduce a novel tracking-by-detection scheme to estimate the underlying scene structure by Atlanta representation. Thereby, we propose an Atlanta frame-aware linear SLAM framework which jointly estimates the camera motion and a planar map supporting the Atlanta structure through a linear Kalman filter. Evaluations on both synthetic and real datasets demonstrate that our approach provides favorable performance compared to existing state-of-the-art methods while extending their working range to the Atlanta world. Kyungdon Joo, Tae-Hyun Oh, François Rameau, Jean-Charles Bazin, In-So Kweon |
ICRA | 1 |
| 2020 | SideGuide: A Large-scale Sidewalk Dataset for Guiding Impaired PeopleabstractIn this paper, we introduce a new large-scale sidewalk dataset called SideGuide that could potentially help impaired people. Unlike most previous datasets, which are focused on road environments, we paid attention to sidewalks, where understanding the environment could provide the potential for improved walking of humans, especially impaired people. Concretely, we interviewed impaired people and carefully selected target objects from the interviewees' feedback (objects they encounter on sidewalks). We then acquired two different types of data: crowd-sourced data and stereo data. We labeled target objects at instance-level (i.e., bounding box and polygon mask) and generated a ground-truth disparity map for the stereo data. SideGuide consists of 350K images with bounding box annotation, 100K images with a polygon mask, and 180K stereo pairs with the ground-truth disparity. We analyzed our dataset by performing baseline analysis for object detection, instance segmentation, and stereo matching tasks. In addition, we developed a prototype that recognizes the target objects and measures distances, which could potentially assist people with disabilities. The prototype suggests the possibility of practical application of our dataset in real life. Kibaek Park, Youngtaek Oh, Soomin Ham, Kyungdon Joo, Hyokyoung Kim, Hyoyoung Kum, In-So Kweon |
IROS | 4 |
| 2020 | Globally Optimal Inlier Set Maximization for Atlanta World UnderstandingabstractIn this work, we describe man-made structures via an appropriate structure assumption, called the Atlanta world assumption, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world can encompass a wider range of scenes, this makes the search space much larger and the problem more challenging. Our input data is a set of surface normals, for example, acquired from RGB-D cameras or 3D laser scanners, as well as lines from calibrated images. Given this input data, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parametrization, and solve this challenging problem using a branch-and-bound (BnB) framework. To alleviate the computational bottleneck in BnB, i.e., the bound computation, we present two bound computation strategies: rectangular bound and slice bound in an efficient measurement domain, i.e., the extended Gaussian image (EGI). In addition, we propose an efficient two-stage method which automatically estimates the number of horizontal directions of a scene. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach. Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Vehicular Multi-Camera Sensor System for Automated Visual Inspection of Electric Power Distribution EquipmentabstractIn this paper, we present a multi-camera sensor system along with its control algorithm for automated visual inspection from a moving vehicle. To accomplish this task, we propose a unique hardware configuration consisting of a frontal stereo vision system, six lateral cameras motorized to tilt, and a GPS/IMU sensor mounted on the roof of a car. From the frontal stereo system, we detect electric poles and estimate their corresponding 3D positions. Based on this 3D estimation, the tilt angles of the motorized lateral cameras are controlled in real-time to capture high resolution images of the equipment - typically installed a few meters above the road surface. In addition, inertial odometry information from the GPS/IMU module is utilized for pose estimation, object localization, and re-identification among cameras. Experimental results demonstrate the efficiency and robustness of our system for automated electric equipment maintenance, which can reduce human effort significantly. Jinsun Park, Ukcheol Shin, Gyumin Shim, Kyungdon Joo, François Rameau, Junhyeok Kim 0004, Dong-Geol Choi, In-So Kweon |
IROS | 4 |
| 2019 | Accurate 3D Reconstruction from Small Motion Clip for Rolling Shutter CamerasabstractStructure from small motion has become an important topic in 3D computer vision as a method for estimating depth, since capturing the input is so user-friendly. However, major limitations exist with respect to the form of depth uncertainty, due to the narrow baseline and the rolling shutter effect. In this paper, we present a dense 3D reconstruction method from small motion clips using commercial hand-held cameras, which typically cause the undesired rolling shutter artifact. To address these problems, we introduce a novel small motion bundle adjustment that effectively compensates for the rolling shutter effect. Moreover, we propose a pipeline for a fine-scale dense 3D reconstruction that models the rolling shutter effect by utilizing both sparse 3D points and the camera trajectory from narrow-baseline images. In this reconstruction, the sparse 3D points are propagated to obtain an initial depth hypothesis using a geometry guidance term. Then, the depth information on each pixel is obtained by sweeping the plane around each depth search space near the hypothesis. The proposed framework shows accurate dense reconstruction results suitable for various sought-after applications. Both qualitative and quantitative evaluations show that our method consistently generates better depth maps compared to state-of-the-art methods. Sunghoon Im 0001, Hyowon Ha, Gyeongmin Choe, Hae-Gon Jeon, Kyungdon Joo, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Robust and Globally Optimal Manhattan Frame Estimation in Near Real TimeabstractMost man-made environments, such as urban and indoor scenes, consist of a set of parallel and orthogonal planar structures. These structures are approximated by the Manhattan world assumption, in which notion can be represented as a Manhattan frame (MF). Given a set of inputs such as surface normals or vanishing points, we pose an MF estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. Conventionally, this problem can be solved by a branch-and-bound framework, which mathematically guarantees global optimality. However, the computational time of the conventional branch-and-bound algorithms is rather far from real-time. In this paper, we propose a novel bound computation method on an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bound with a constant complexity, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for various synthetic and real-world data. We also show the versatility of our approach through three different applications: extension to multiple MF estimation, 3D rotation based video stabilization, and vanishing point estimation (line clustering). Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Globally Optimal Inlier Set Maximization for Atlanta Frame EstimationabstractIn this work, we describe man-made structures via an appropriate structure assumption, called Atlanta world, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world permits to encompass a wider range of scenes, this makes the solution space larger and the problem more challenging. Given a set of inputs, such as lines in a calibrated image or surface normals, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parameterization, and solve this challenging problem by a branch-and-bound framework. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach. Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin |
CVPR | 1 |
| 2018 | Efficient adaptive non-maximal suppression algorithms for homogeneous spatial keypoint distribution
Oleksandr Bailo, François Rameau, Kyungdon Joo, Jinsun Park, Oleksandr Bogdan, In-So Kweon |
Pattern Recognit. Lett. | 3 |
| 2017 | Personalized Cinemagraphs Using Semantic Understanding and Collaborative LearningabstractCinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph requires isolating objects in a semantically meaningful way and then selecting good start times and looping periods for those objects to minimize visual artifacts (such a tearing). To achieve this, we present a new technique that uses object recognition and semantic segmentation as part of an optimization method to automatically create cinemagraphs from videos that are both visually appealing and semantically meaningful. Given a scene with multiple objects, there are many cinemagraphs one could create. Our method evaluates these multiple candidates and presents the best one, as determined by a model trained to predict human preferences in a collaborative way. We demonstrate the effectiveness of our approach with multiple results and a user study. Tae-Hyun Oh, Kyungdon Joo, Neel Joshi, Baoyuan Wang, In-So Kweon, Sing Bing Kang |
ICCV | 2 |
| 2016 | Globally Optimal Manhattan Frame Estimation in Real-TimeabstractGiven a set of surface normals, we pose a Manhattan Frame (MF) estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. We solve this problem through a branchand-bound framework, which mathematically guarantees a globally optimal solution. However, the computational time of conventional branch-and-bound algorithms are intractable for real-time performance. In this paper, we propose a novel bound computation method within an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bounds in real-time, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for synthetic and real-world data. We also show the versatility of our approach through two applications: extension to multiple MF estimation and video stabilization. Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
CVPR | 1 |
| 2016 | Vision system and depth processing for DRC-HUBO+abstractThis paper presents a vision system and a depth processing algorithm for DRC-HUBO+, the winner of the DRC finals 2015. Our system is designed to reliably capture 3D information of a scene and objects and to be robust to challenging environment conditions. We also propose a depth-map upsampling method that produces an outliers-free depth map by explicitly handling depth outliers. Our system is suitable for robotic applications in which a robot interacts with the real-world, requiring accurate object detection and pose estimation. We evaluate our depth processing algorithm in comparison with state-of-the-art algorithms on several synthetic and real-world datasets. Inwook Shim, Seunghak Shin, Yunsu Bok, Kyungdon Joo, Dong-Geol Choi, Joon-Young Lee, Jaesik Park, Jun-Ho Oh, In-So Kweon |
ICRA | 4 |
| 2016 | A Real-Time Augmented Reality System to See-Through CarsabstractOne of the most hazardous driving scenario is the overtaking of a slower vehicle, indeed, in this case the front vehicle (being overtaken) can occlude an important part of the field of view of the rear vehicle's driver. This lack of visibility is the most probable cause of accidents in this context. Recent research works tend to prove that augmented reality applied to assisted driving can significantly reduce the risk of accidents. In this paper, we present a real-time marker-less system to see through cars. For this purpose, two cars are equipped with cameras and an appropriate wireless communication system. The stereo vision system mounted on the front car allows to create a sparse 3D map of the environment where the rear car can be localized. Using this inter-car pose estimation, a synthetic image is generated to overcome the occlusion and to create a seamless see-through effect which preserves the structure of the scene. François Rameau, Hyowon Ha, Kyungdon Joo, Jinsoo Choi, Kibaek Park, In-So Kweon |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2015 | Accurate Camera Calibration Robust to Defocus Using a SmartphoneabstractWe propose a novel camera calibration method for defocused images using a smartphone under the assumption that the defocus blur is modeled as a convolution of a sharp image with a Gaussian point spread function (PSF). In contrast to existing calibration approaches which require well-focused images, the proposed method achieves accurate camera calibration with severely defocused images. This robustness to defocus is due to the proposed set of unidirectional binary patterns, which simplifies 2D Gaussian deconvolution to a 1D Gaussian deconvolution problem with multiple observations. By capturing the set of patterns consecutively displayed on a smartphone, we formulate the feature extraction as a deconvolution problem to estimate feature point locations in sub-pixel accuracy and the blur kernel in each location. We also compensate the error in camera parameters due to refraction of the glass panel of the display device. We evaluate the performance of the proposed method on synthetic and real data. Even under severe defocus, our method shows accurate camera calibration result. Hyowon Ha, Yunsu Bok, Kyungdon Joo, Jiyoung Jung, In-So Kweon |
ICCV | 3 |
| 2015 | High Quality Structure from Small Motion for Rolling Shutter CamerasabstractWe present a practical 3D reconstruction method to obtain a high-quality dense depth map from narrow-baseline image sequences captured by commercial digital cameras, such as DSLRs or mobile phones. Depth estimation from small motion has gained interest as a means of various photographic editing, but important limitations present themselves in the form of depth uncertainty due to a narrow baseline and rolling shutter. To address these problems, we introduce a novel 3D reconstruction method from narrow-baseline image sequences that effectively handles the effects of a rolling shutter that occur from most of commercial digital cameras. Additionally, we present a depth propagation method to fill in the holes associated with the unknown pixels based on our novel geometric guidance model. Both qualitative and quantitative experimental results show that our new algorithm consistently generates better 3D depth maps than those by the state-of-the-art method. Sunghoon Im 0001, Hyowon Ha, Gyeongmin Choe, Hae-Gon Jeon, Kyungdon Joo, In-So Kweon |
ICCV | 5 |
| 2015 | Line meets as-projective-as-possible image stitching with moving DLTabstractWe propose a spatially varying stitching method with line correspondences. We are motivated by the observation that point features could be spatially biased or not matched in practice, e.g., repeated textures or homogeneous regions of man-made structures. In this scenario, line matches can provide strong correspondences as well as supplement cues, such as the structure preserving property. With these advantages, we adopt a feature fusion method that combines point and line correspondences into a unified framework for spatially varying stitching. We then estimate the balancing parameter between the point and line terms using geometric error. Our experiments show accurate alignment for challenging but common cases. Kyungdon Joo, Namil Kim, Tae-Hyun Oh, In-So Kweon |
ICIP | 1 |
| 2013 | Hierarchical 3D line restoration based on angular proximity in structured environmentsabstractWe present a method based on a hierarchical clustering to restore the 3D lines of structured environments. In previous approaches, the restoration of noisy 3D lines is a challenging problem because it is difficult to define a suitable similarity measure discriminative to other lines. Our motivation to overcome the difficulty is that most structured scenes consist of sets of parallel 3D lines with the same angular proximity, which provides a hierarchical similarity measure for structured 3D lines. Accordingly, our restoration method works in a manner that clustering is hierarchically performed on angular and distance levels. The 3D line restoration is then achieved by finding the center of each cluster. The framework also makes the clustered 3D lines align along the associated angular directions. We compare the proposed algorithm with methods using no knowledge of the angular information, and demonstrate its effectiveness through real-world experiments. Kyungdon Joo, Tae-Hyun Oh, Hyeongwoo Kim, In-So Kweon |
ICIP | 1 |