Pyojin Kim

dblp:155/4572 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-5605-8848ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 SPLiCE: Single-Point LiDAR and Camera Calibration & Estimation Leveraging Manhattan World
abstract
We present a novel calibration method between single-point LiDAR and camera sensors utilizing an easy-to-build customized calibration board satisfying the Manhattan world (MW). Previous methods for LiDAR-camera (LC) calibration focus on line and plane correspondences. However, they require dense 3D point clouds from heavy and expensive LiDAR to simplify alignments; otherwise, these approaches fail for extremely sparse LiDAR. Compact, lightweight, and sparse LiDAR and camera sensors are inevitable for micro drones like Crazyflie with a maximum payload of 15 g, but there are no explicit calibration methods for them. To address these issues, we propose a new extrinsic calibration method with a new calibration board, which rotates like a door to capture geometric features and align them with images. Once we find an initial estimate, we refine the relative rotation by minimizing the angle difference between the grid orientation of the checkerboard and the MW axes. We demonstrate the effectiveness of the proposed method through various LC configurations, achieving its capability and high accuracy compared to other state-of-the-art approaches. We release our calibration toolkit, source codes, and how to make the calibration boards for the robotics community: https://SPLiCE-Calib.github.io/.
Jeahn Han, Jungil Ham, Pyojin Kim
IROS4
2023 Hong Kong World: Leveraging Structural Regularity for Line-Based SLAM
abstract
Manhattan and Atlanta worlds hold for the structured scenes with only vertical and horizontal dominant directions (DDs). To describe the scenes with additional sloping DDs, a mixture of independent Manhattan worlds seems plausible, but may lead to unaligned and unrelated DDs. By contrast, we propose a novel structural model called Hong Kong world. It is more general than Manhattan and Atlanta worlds since it can represent the environments with slopes, e.g., a city with hilly terrain, a house with sloping roof, and a loft apartment with staircase. Moreover, it is more compact and accurate than a mixture of independent Manhattan worlds by enforcing the orthogonality constraints between not only vertical and horizontal DDs, but also horizontal and sloping DDs. We further leverage the structural regularity of Hong Kong world for the line-based SLAM. Our SLAM method is reliable thanks to three technical novelties. First, we estimate DDs/vanishing points in Hong Kong world in a semi-searching way. We use a new consensus voting strategy for search, instead of traditional branch and bound. This method is the first one that can simultaneously determine the number of DDs, and achieve quasi-global optimality in terms of the number of inliers. Second, we compute the camera pose by exploiting the spatial relations between DDs in Hong Kong world. This method generates concise polynomials, and thus is more accurate and efficient than existing approaches designed for unstructured scenes. Third, we refine the estimated DDs in Hong Kong world by a novel filter-based method. Then we use these refined DDs to optimize the camera poses and 3D lines, leading to higher accuracy and robustness than existing optimization algorithms. In addition, we establish the first dataset of sequential images in Hong Kong world. Experiments showed that our approach outperforms state-of-the-art methods in terms of accuracy and/or efficiency.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Pyojin Kim, Kyungdon Joo, Zhenjun Zhao, Yun-Hui Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Scale-Aware Monocular Visual Odometry and Extrinsic Calibration Using Vehicle Kinematics
abstract
This paper proposes a new approach to scale-aware monocular visual odometry (VO) and extrinsic calibration using constraints on camera motion by vehicle kinematics. Main idea is to utilize the Ackermann steering model to observe absolute metric scale in turning motion. To describe motion of the camera attached to the vehicle, we first estimate unknown camera-vehicle relative pose by the proposed extrinsic calibration method. To stably observe scale, we detect turn regions and design an observer to estimate the absolute scale as a function of the camera rotation and direction of translational motion during turning. Using the observed scale, we propose an absolute scale recovery to estimate the unknown scale between turns. Because the proposed scale observer becomes singular near zero rotation, we conduct sensitivity analysis on the scale observer, and investigate appropriate conditions for stable scale estimation. For quantitative evaluation of the extrinsic calibration and the absolute scale recovery, we randomly generate synthetic driving datasets with various noise conditions, and evaluate the performance of each module statistically by Monte-Carlo simulations. To evaluate the overall performance, we implement our method and state-of-the-art monocular and stereo VO methods in the public outdoor driving KITTI dataset, and our method shows competitive scale recovery performance with no external sensor and no assumption on surroundings such as planar ground landmarks. To show promising applicability, we collect real-world driving datasets in two multi-floor underground parking lots, and demonstrate the accurate absolute scale recovery performance of our method in indoor driving situations.
Youngseok Jang, Junha Kim, Pyojin Kim, H. Jin Kim
IEEE Trans. Intell. Transp. Syst.4
2022 Single User WiFi Structure from Motion in the Wild
abstract
This paper proposes a novel motion estimation algorithm using WiFi networks and IMU sensor data in large uncontrolled environments, dubbed “WiFi Structure-from-Motion” (WiFi SfM). Given smartphone sensor data through day-to-day activities from a single user over a month, our WiFi SfM algorithm estimates smartphone motion tra-jectories and the structure of the environment represented as a WiFi radio map. The approach 1) establishes frame-to-frame correspondences based on WiFi fingerprints while exploiting our repetitive behavior patterns; 2) aligns trajectories via bundle adjustment; and 3) trains a self-supervised neural network to extract further motion constraints. We have col-lected 235 hours of smartphone data, spanning 38 days of daily activities in a university campus. Our experiments demonstrate the effectiveness of our approach over the competing methods with qualitative evaluations of the estimated motions and quantitative evaluations of indoor localization accuracy based on the reconstructed WiFi radio map. The WiFi SfM technology will potentially allow digital mapping companies to build better radio maps automatically by asking users to share WiFi/IMU sensor data in their daily activities.
Yiming Qian, Hang Yan 0002, Sachini Herath, Pyojin Kim, Yasutaka Furukawa
ICRA4
2022 Linear RGB-D SLAM for Structured Environments
abstract
We propose a new linear RGB-D simultaneous localization and mapping (SLAM) formulation by utilizing planar features of the structured environments. The key idea is to understand a given structured scene and exploit its structural regularities such as the Manhattan world. This understanding allows us to decouple the camera rotation by tracking structural regularities, which makes SLAM problems free from being highly nonlinear. Additionally, it provides a simple yet effective cue for representing planar features, which leads to a linear SLAM formulation. Given an accurate camera rotation, we jointly estimate the camera translation and planar landmarks in the global planar map using a linear Kalman filter. Our linear SLAM method, called L-SLAM, can understand not only the Manhattan world but the more general scenario of the Atlanta world, which consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction. To this end, we introduce a novel tracking-by-detection scheme that infers the underlying scene structure by Atlanta representation. With efficient Atlanta representation, we formulate a unified linear SLAM framework for structured environments. We evaluate L-SLAM on a synthetic dataset and RGB-D benchmarks, demonstrating comparable performance to other state-of-the-art SLAM methods without using expensive nonlinear optimization. We assess the accuracy of L-SLAM on a practical application of augmented reality.
Kyungdon Joo, Pyojin Kim, Martial Hebert, In-So Kweon, H. Jin Kim
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Learning To Identify Correct 2D-2D Line Correspondences on Sphere
abstract
Given a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both structured and unstructured scenes. Instead of geometric constraint, we leverage the spatial regularity on sphere. Specifically, we propose to map line correspondences into vectors tangent to sphere. We use these vectors to encode both angular and positional variations of image lines, which is more reliable and concise than directly using inclinations, midpoints or endpoints of image lines. Neighboring vectors mapped from correct matches exhibit a spatial regularity called local trend consistency, regardless of the type of scenes. To encode this regularity, we design a neural network and also propose a novel loss function that enforces the smoothness constraint of vector field. In addition, we establish a large real-world dataset for image line matching. Experiments showed that our approach outperforms state-of-the-art ones in terms of accuracy, efficiency and robustness, and also leads to high generalization.
Haoang Li, Kai Chen 0028, Ji Zhao 0001, Jiangliu Wang, Pyojin Kim, Zhe Liu 0022, Yun-Hui Liu 0001
CVPR5
2021 Learning Icosahedral Spherical Probability Map Based on Bingham Mixture Model for Vanishing Point Estimation
abstract
Existing vanishing point (VP) estimation methods rely on pre-extracted image lines and/or prior knowledge of the number of VPs. However, in practice, this information may be insufficient or unavailable. To solve this problem, we propose a network that treats a perspective image as input and predicts a spherical probability map of VP. Based on this map, we can detect all the VPs. Our method is reliable thanks to four technical novelties. First, we leverage the icosahedral spherical representation to express our probability map. This representation provides uniform pixel distribution, and thus facilitates estimating arbitrary positions of VPs. Second, we design a loss function that enforces the antipodal symmetry and sparsity of our spherical probability map to prevent over-fitting. Third, we generate the ground truth probability map that reasonably expresses the locations and uncertainties of VPs. This map unnecessarily peaks at noisy annotated VPs, and also exhibits various anisotropic dispersions. Fourth, given a predicted probability map, we detect VPs by fitting a Bingham mixture model. This strategy can robustly handle close VPs and provide the confidence level of VP useful for practical applications. Experiments showed that our method achieves the best compromise between generality, accuracy, and efficiency, compared with state-of-the-art approaches.
Haoang Li, Kai Chen 0028, Pyojin Kim, Kuk-Jin Yoon, Zhe Liu 0022, Kyungdon Joo, Yun-Hui Liu 0001
ICCV3
2021 Fusion-DHL: WiFi, IMU, and Floorplan Fusion for Dense History of Locations in Indoor Environments
abstract
The paper proposes a multi-modal sensor fusion algorithm that fuses WiFi, IMU, and floorplan information to infer an accurate and dense location history in indoor environments. The algorithm uses 1) an inertial navigation algorithm to estimate a relative motion trajectory from IMU sensor data; 2) a WiFi-based localization API in industry to obtain positional constraints and geo-localize the trajectory; and 3) a convolutional neural network to refine the location history to be consistent with the floorplan. We have developed a data acquisition app to build a new dataset with WiFi, IMU, and floorplan data with ground-truth positions at 4 university buildings and 3 shopping malls. Our qualitative and quantitative evaluations demonstrate that the proposed system is able to produce twice as accurate and a few orders of magnitude denser location history than the current standard, while requiring minimal additional energy consumption. We will publicly share our code and models.
Sachini Herath, Saghar Irandoust, Yiming Qian, Pyojin Kim, Yasutaka Furukawa
ICRA5
2020 Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World
Haoang Li, Pyojin Kim, Ji Zhao 0001, Kyungdon Joo, Zhe Liu 0022, Yun-Hui Liu 0001
ECCV (22)2
2020 Moving object detection for visual odometry in a dynamic environment based on occlusion accumulation
abstract
Detection of moving objects is an essential capability in dealing with dynamic environments. Most moving object detection algorithms have been designed for color images without depth. For robotic navigation where real-time RGBD data is often readily available, utilization of the depth information would be beneficial for obstacle recognition. Here, we propose a simple moving object detection algorithm that uses RGB-D images. The proposed algorithm does not require estimating a background model. Instead, it uses an occlusion model which enables us to estimate the camera pose on a background confused with moving objects that dominate the scene. The proposed algorithm allows to separate the moving object detection and visual odometry (VO) so that an arbitrary robust VO method can be employed in a dynamic situation with a combination of moving object detection, whereas other VO algorithms for a dynamic environment are inseparable. In this paper, we use dense visual odometry (DVO) as a VO method with a bi-square regression weight. Experimental results show the segmentation accuracy and the performance improvement of DVO in the situations. We validate our algorithm in public datasets and our dataset which also publicly accessible.
Haram Kim, Pyojin Kim, H. Jin Kim
ICRA2
2018 Indoor RGB-D Compass From a Single Line and Plane
abstract
We propose a novel approach to estimate the three degrees of freedom (DoF) drift-free rotational motion of an RGB-D camera from only a single line and plane in the Manhattan world (MW). Previous approaches exploit the surface normal vectors and vanishing points to achieve accurate 3-DoF rotation estimation. However, they require multiple orthogonal planes or many consistent lines to be visible throughout the entire rotation estimation process; otherwise, these approaches fail. To overcome these limitations, we present a new method that estimates absolute camera orientation from only a single line and a single plane in RANSAC, which corresponds to the theoretical minimal sampling for 3-DoF rotation estimation. Once we find an initial rotation estimate, we refine the camera orientation by minimizing the average orthogonal distance from the endpoints of the lines parallel to the MW axes. We demonstrate the effectiveness of the proposed algorithm through an extensive evaluation on a variety of RGB-D datasets and compare with other state-of-the-art methods.
Pyojin Kim, Brian Coltin, H. Jin Kim
CVPR1
2018 Linear RGB-D SLAM for Planar Environments
Pyojin Kim, Brian Coltin, H. Jin Kim
ECCV (4)1
2018 Low-Drift Visual Odometry in Structured Environments by Decoupling Rotational and Translational Motion
abstract
We present a low-drift visual odometry algorithm that separately estimates rotational and translational motion from lines, planes, and points found in RGB-D images. Previous methods estimate drift-free rotational motion from structural regularities to reduce drift in the rotation estimate, which is the primary source of positioning inaccuracy in visual odometry. However, multiple orthogonal planes are required to be visible throughout the entire motion estimation process; otherwise, these VO approaches fail. We propose a new approach to estimate drift-free rotational motion jointly from both lines and planes by exploiting environmental regularities. We track the spatial regularities with an efficient SO(3)-manifold constrained mean shift algorithm. Once the drift-free rotation is found, we recover the translational motion from all tracked points with and without depth by minimizing the de-rotated reprojection error. We compare the proposed algorithm to other state-of-the-art visual odometry methods on a variety of RGB-D datasets (including especially challenging pure rotations) and demonstrate improved accuracy and lower drift error.
Pyojin Kim, Brian Coltin, H. Jin Kim
ICRA1
2018 Edge-Based Robust RGB-D Visual Odometry Using 2-D Edge Divergence Minimization
abstract
This paper proposes an edge-based robust RGB-D visual odometry (VO) using 2-D edge divergence minimization. Our approach focuses on enabling the VO to operate in more general environments subject to low texture and changing brightness, by employing image edge regions and their image gradient vectors within the iterative closest points (ICP) framework. For more robust and stable ICP-based optimization, we propose a robust edge matching criterion with image gradient vectors. In addition, to reduce a bad effect of outlier residuals, we propose an improved edge registration problem of 2-D edge divergence minimization in the manner of an iterative re-weight least squares (IRLS) motion estimation. To accelerate the proposed approach, a pixel sub-sampling method is employed. We evaluate estimation performance of our method in changing brightness conditions and low-textured scenes. Our approach shows more robust motion estimation than state-of-the-art methods while maintaining comparable accuracy in challenging image sequences at real-time (25 Hz) operation.
Pyojin Kim, H. Jin Kim
IROS2
2017 Visual Odometry with Drift-Free Rotation Estimation Using Indoor Scene Regularities
Pyojin Kim, Brian Coltin, H. Jin Kim
BMVC1
2017 Robust visual localization in changing lighting conditions
abstract
We present an illumination-robust visual localization algorithm for Astrobee, a free-flying robot designed to autonomously navigate on the International Space Station (ISS). Astrobee localizes with a monocular camera and a pre-built sparse map composed of natural visual features. Astrobee must perform tasks not only during the day, but also at night when the ISS lights are dimmed. However, the localization performance degrades when the observed lighting conditions differ from the conditions when the sparse map was built. We investigate and quantify the effect of lighting variations on visual feature-based localization systems, and discover that maps built in darker conditions can also be effective in bright conditions, but the reverse is not true. We extend Astrobee's localization algorithm to make it more robust to changing-light environments on the ISS by automatically recognizing the current illumination level, and selecting an appropriate map and camera exposure time. We extensively evaluate the proposed algorithm through experiments on Astrobee.
Pyojin Kim, Brian Coltin, Oleg Alexandrov, H. Jin Kim
ICRA1
2015 Robust visual odometry to irregular illumination changes with RGB-D camera
abstract
Sensitivity to illumination conditions poses a challenge when utilizing visual odometry (VO) in various applications. To make VO robust with respect to illumination conditions, they need to be considered explicitly. In this paper, we propose a direct visual odometry method which can handle illumination changes by considering an affine illumination model to compensate abrupt, local light variations during direct motion estimation process. The core of our proposed method is to estimate the relative camera pose and the parameters of the illumination changes by minimizing the sum of squared photometric error with efficient second-order minimization. We evaluate the performance of the proposed algorithm on synthetic and real RGB-D datasets with ground-truth. Our result implies that the proposed method successfully estimates 6-DoF pose under significant illumination changes whereas existing direct visual odometry methods either fail or lose accuracy.
Pyojin Kim, Hyon Lim, H. Jin Kim
IROS1
2014 6-DoF velocity estimation using RGB-D camera based on optical flow
abstract
In this paper, we suggest a new 6-DoF velocity estimation algorithm using RGB and depth images. Autonomous control of mobile robots requires their velocity information. There exist numerous researches on estimating and measuring the velocity. However, more investigations are needed related to vision sensors and depth image. In this work, we propose an algorithm for velocity estimation with an RGB-D sensor based on image jacobian matrix usually used in image-based visual servoing. We validate the performance of the proposed estimation algorithm in various environments with the RGB-D benchmark dataset. The velocity estimation results show the high quality of estimated 6-DoF velocity compared to the ground truth velocity.
Pyojin Kim, Hyon Lim, H. Jin Kim
SMC1