VLDB 2026 Research / reviewers in the wild / expert
Ayoung Kim
dblp:16/7747
· DBLP profile ↗
50ranked-venue papers
4as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 3 first-author · 21 since 2021Systems, architecture and hardware · 34 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoEL: Robust Event-Based 3-D Line ReconstructionabstractEvent cameras in motion tend to detect object boundaries or texture edges, which produce lines of brightness changes, especially in man-made environments. While lines can constitute a robust intermediate representation that is consistently observed, the sparse nature of lines may lead to drastic deterioration with minor estimation errors. Only a few previous works, often accompanied by additional sensors, utilize lines to compensate for the severe domain discrepancies of event sensors along with unpredictable noise characteristics. We propose a method that can stably extract tracks of varying appearances of lines using a clever algorithmic process that observes multiple representations from various time slices of events, compensating for potential adversaries within the event data. We then propose geometric cost functions that can refine the 3D line maps and camera poses, eliminating projective distortions and depth ambiguities. The 3D line maps are highly compact and can be equipped with our proposed cost function, which can be adapted for any observations that can detect and extract line structures or projections of them, including 3D point cloud maps or image observations. We demonstrate that our formulation is powerful enough to exhibit a significant performance boost in event-based mapping and pose refinement across diverse datasets, and can be flexibly applied to multimodal scenarios. Our results confirm that the proposed line-based formulation is a robust and effective approach for the practical deployment of event-based perceptual modules. Project page Gwangtak Bae, Seunggu Kang, Ayoung Kim, Young Min Kim 0001 |
IEEE Trans. Robotics | 5 |
| 2025 | Tracking-Based Adaptive Temporal Resampling for Video Coding for MachinesabstractAs video data is increasingly consumed by machines rather than solely by humans, there is a growing demand for new compression methods that efficiently accommodate this shift. Since object information extracted from videos is crucial for machine consumption, analyzing the similarity between adjacent pictures based on tracked object information can help identify motion-based redundancy which can then be utilized for video compression. In this paper, we perform object tracking on input video and analyze the similarity between adjacent pictures based on the movement of tracked objects. After classifying the pictures based on their similarity, highly redundant pictures within each group are aggressively resampled in the temporal domain to improve compression efficiency while maintaining machine performance. We propose a novel picture grouping method to cluster similar adjacent pictures and describe the process of similarity assessment. We evaluated the compression efficiency of the proposed object tracking-based adaptive temporal resampling through performance evaluation experiments, achieving BD-mAP improvements in object detection of 1.29%, 0.47%, and 2.44% for Random Access (RA), Low-Delay (LD), and All Intra (AI) modes, respectively, and achieving BD-MOTA improvements in object tracking of 0.02%, 0.61%, and 6.86% for RA, LD, and AI modes, respectively. Eun-bin An, Minsuk Kim, Kwang-deok Seo, Sangwoon Kwak, Ayoung Kim, Soon-Heung Jung, Hyon-Gon Choo |
AVSS | 5 |
| 2025 | 2D Gaussian Splatting-Based Sparse-View Transparent Object Depth Reconstruction Via Physics Simulation for Scene Update
Jeongyun Kim, Seunghoon Jeong, Giseop Kim, Myung-Hwan Jeon, Eunji Jun, Ayoung Kim |
ICCV | 6 |
| 2025 | Registration beyond Points: General Affine Subspace Alignment via Geodesic Distance on Grassmann ManifoldabstractAffine Grassmannian has been favored for expressing proximity between lines and planes due to its theoretical exactness in measuring distances among features. Despite this advantage, the existing method can only measure the proximity without yielding the distance as an explicit function of rigid body transformation. Thus, an optimizable distance function on the manifold has remained underdeveloped, stifling its application in registration problems. This paper is the first to explicitly derive an optimizable cost function between two Grassmannian features with respect to rigid body transformation ($\mathbf{R}$ and $\mathbf{t}$). Specifically, we present a rigorous mathematical proof demonstrating that the bases of high-dimensional linear subspaces can serve as an explicit representation of the cost. Finally, we propose an optimizable cost function based on the transformed bases that can be applied to the registration problem of any affine subspace. Compared to vector parameter-based approaches, our method is able to find a globally optimal solution by directly minimizing the geodesic distance which is agnostic to representation ambiguity. The resulting cost function and its extension to the inlier-set maximizing Branch-and-Bound (BnB) solver have been demonstrated to improve the convergence of existing solutions or outperform them in various computer vision tasks. The code is available on https://github.com/joomeok/GrassmannRegistration. Hyeonjae Gil, Junwoo Jang, Maani Ghaffari Jadidi, Ayoung Kim |
ICCV | 5 |
| 2025 | Ephemerality Meets Lidar-Based Lifelong MappingabstractLifelong mapping is crucial for the long-term deployment of robots in dynamic environments. In this paper, we present ELite, an ephemerality-aided LiDAR-based lifelong mapping framework which can seamlessly align multiple session data, remove dynamic objects, and update maps in an end-toend fashion. Map elements are typically classified as static or dynamic, but cases like parked cars indicate the need for more detailed categories than binary. Central to our approach is the probabilistic modeling of the world into two-stage ephemerality, which represent the transiency of points in the map within two different time scales. By leveraging the spatiotemporal context encoded in ephemeralities, ELite can accurately infer transient map elements, maintain a reliable up-to-date static map, and improve robustness in aligning the new data in a more finegrained manner. Extensive real-world experiments on long-term datasets demonstrate the robustness and effectiveness of our system. The source code is publicly available for the robotics community: https://github.com/dongjae0107/ELite. Hyeonjae Gil, Giseop Kim, Ayoung Kim |
ICRA | 4 |
| 2025 | Helios: Heterogeneous Lidar Place Recognition via Overlap-Based Learning and Local Spherical TransformerabstractLiDAR place recognition is a crucial module in localization that matches the current location with previously observed environments. Most existing approaches in LiDAR place recognition dominantly focus on the spinning type LiDAR to exploit its large FOV for matching. However, with the recent emergence of various LiDAR types, the importance of matching data across different LiDAR types has grown significantly-a challenge that has been largely overlooked for many years. To address these challenges, we introduce HeLiOS, a deep network tailored for heterogeneous LiDAR place recognition, which utilizes small local windows with spherical transformers and optimal transport-based cluster assignment for robust global descriptors. Our overlap-based data mining and guided-triplet loss overcome the limitations of traditional distance-based mining and discrete class constraints. HeLiOS is validated on public datasets, demonstrating performance in heterogeneous LiDAR place recognition while including an evaluation for longterm recognition, showcasing its ability to handle unseen LiDAR types. We release the HeLiOS code as an open source for the robotics community at https://github.com/minwoo0611/HeLiOS. Minwoo Jung, Sangwoo Jung 0002, Hyeonjae Gil, Ayoung Kim |
ICRA | 4 |
| 2025 | HeRCULES: Heterogeneous Radar Dataset in Complex Urban Environment for Multi-Session Radar SLAMabstractRecently, radars have been widely featured in robotics for their robustness in challenging weather conditions. Two commonly used radar types are spinning radars and phased-array radars, each offering distinct sensor characteristics. Existing datasets typically feature only a single type of radar, leading to the development of algorithms limited to that specific kind. In this work, we highlight that combining different radar types offers complementary advantages, which can be leveraged through a heterogeneous radar dataset. Moreover, this new dataset fosters research in multi-session and multirobot scenarios where robots are equipped with different types of radars. In this context, we introduce the HeRCULES dataset, a comprehensive, multi-modal dataset with heterogeneous radars, FMCW LiDAR, IMU, GPS, and cameras. This is the first dataset to integrate 4D radar and spinning radar alongside FMCW LiDAR, offering unparalleled localization, mapping, and place recognition capabilities. The dataset covers diverse weather and lighting conditions and a range of urban traffic scenarios, enabling a comprehensive analysis across various environments. The sequence paths with multiple revisits and ground truth pose for each sensor enhance its suitability for place recognition research. We expect the HeRCULES dataset to facilitate odometry, mapping, place recognition, and sensor fusion research. The dataset and development tools are available at https://sites.google.com/view/herculesdataset. Hanjun Kim 0002, Minwoo Jung, Chiyun Noh, Sangwoo Jung 0002, Hyunho Song, Wooseong Yang, Hyesu Jang, Ayoung Kim |
ICRA | 8 |
| 2025 | TranSplat: Surface Embedding-Guided 3D Gaussian Splatting for Transparent Object ManipulationabstractTransparent object manipulation remains a significant challenge in robotics due to the difficulty of acquiring accurate and dense depth measurements. Conventional depth sensors often fail with transparent objects, resulting in incomplete or erroneous depth data. Existing depth completion methods struggle with interframe consistency and incorrectly model transparent objects as Lambertian surfaces, leading to poor depth reconstruction. To address these challenges, we propose TranSplat, a surface embedding-guided 3D Gaussian Splatting method tailored for transparent objects. TranSplat uses a latent diffusion model to generate surface embeddings that provide consistent and continuous representations, making it robust to changes in viewpoint and lighting. By integrating these surface embeddings with input RGB images, TranSplat effectively captures the complexities of transparent surfaces, enhancing the splatting of 3D Gaussians and improving depth completion. Evaluations on synthetic and real-world transparent object benchmarks, as well as robot grasping tasks, show that TranSplat achieves accurate and dense depth completion, demonstrating its effectiveness in practical applications. We open-source synthetic dataset and model: https://github.com/jeongyun0609/TranSplat Jeongyun Kim, Jeongho Noh, Dong-Guw Lee, Ayoung Kim |
ICRA | 4 |
| 2025 | GaRLIO: Gravity Enhanced Radar-LiDAR-Inertial OdometryabstractRecently, gravity has been highlighted as a crucial constraint for state estimation to alleviate potential vertical drift. Existing online gravity estimation methods rely on pose estimation combined with IMU measurements, which is considered best practice when direct velocity measurements are unavailable. However, with radar sensors providing direct velocity data-a measurement not yet utilized for gravity estimation-we found a significant opportunity to improve gravity estimation accuracy substantially. GaRLIO, the proposed gravity-enhanced Radar-LiDAR-Inertial Odometry, can robustly predict gravity to reduce vertical drift while simultaneously enhancing state estimation performance using pointwise velocity measurements. Furthermore, GaRLIO ensures robustness in dynamic environments by utilizing radar to remove dynamic objects from LiDAR point clouds. Our method is validated through experiments in various environments prone to vertical drift, demonstrating superior performance compared to traditional LiDAR-Inertial Odometry methods. We make our source code publicly available to encourage further research and development. https://github.com/ChiyunNoh/GaRLIO Chiyun Noh, Wooseong Yang, Minwoo Jung, Sangwoo Jung 0002, Ayoung Kim |
ICRA | 5 |
| 2025 | Ground-Optimized 4D Radar-Inertial Odometry Via Continuous Velocity Integration Using Gaussian ProcessabstractRadar ensures robust sensing capabilities in adverse weather conditions, yet challenges remain due to its high inherent noise level. Existing radar odometry has overcome these challenges with strategies such as filtering spurious points, exploiting Doppler velocity, or integrating with inertial measurements. This paper presents two novel improvements beyond the existing radar-inertial odometry: ground-optimized noise filtering and continuous velocity preintegration. Despite the widespread use of ground planes in LiDAR odometry, imprecise ground point distributions of radar measurements cause naive plane fitting to fail. Unlike plane fitting in LiDAR, we introduce a zone-based uncertainty-aware ground modeling specifically designed for radar. Secondly, we note that radar velocity measurements can be better combined with IMU for a more accurate preintegration in radar-inertial odometry. Existing methods often ignore temporal discrepancies between radar and IMU by simplifying the complexities of asynchronous data streams with discretized propagation models. Tackling this issue, we leverage GP and formulate a continuous preintegration method for tightly integrating 3-DOF linear velocity with IMU, facilitating full 6-DOF motion directly from the raw measurements. Our approach demonstrates remarkable performance (less than 1 % vertical drift) in public datasets with meticulous conditions, illustrating substantial improvement in elevation accuracy. The code will be released as open source for the community: https://github.com/wooseongY/Go-RIO. Wooseong Yang, Hyesu Jang, Ayoung Kim |
ICRA | 3 |
| 2025 | PlanarMesh: Building Compact 3D Meshes from LiDAR using Incremental Adaptive Resolution ReconstructionabstractBuilding an online 3D LiDAR mapping system that produces a detailed surface reconstruction while remaining computationally efficient is a challenging task. In this paper, we present PlanarMesh, a novel incremental, mesh-based LiDAR reconstruction system that adaptively adjusts mesh resolution to achieve compact, detailed reconstructions in real-time. It introduces a new representation, planar-mesh, which combines plane modeling and meshing to capture both large surfaces and detailed geometry. The planar-mesh can be incrementally updated considering both local surface curvature and free-space information from sensor measurements. We employ a multi-threaded architecture with a Bounding Volume Hierarchy (BVH) for efficient data storage and fast search operations, enabling real-time performance. Experimental results show that our method achieves reconstruction accuracy on par with, or exceeding, state-of-the-art techniques—including truncated signed distance functions, occupancy mapping, and voxel-based meshing—while producing smaller output file sizes (10 times smaller than raw input and more than 5 times smaller than mesh-based methods) and maintaining real-time performance (around 2 Hz for a 64-beam sensor). Nived Chebrolu, Yifu Tao, Lintong Zhang, Ayoung Kim, Maurice Fallon |
IROS | 5 |
| 2024 | Relationship between short-term ozone exposure, cause-specific mortality, and high-risk populations: A nationwide, time-stratified, case-crossover studyabstractGround-level ozone is formed by chemical reactions between nitrogen oxides (NOx) and volatile organic compounds (VOC) under sunlight, primarily from industrial and vehicular emissions. Cinoo Kang, Jieun Oh, Hyewon Yun, Juyeon Yang, Chaerin Park, Seoyeong Ahn, Ayoung Kim, Dohoon Kwon 0001, Jinah Park, Ejin Kim, Ho Kim, Whanhee Lee |
BIBM | 9 |
| 2024 | Unbiased Estimator for Distorted Conics in Camera CalibrationabstractIn the literature, points and conics have been major features for camera geometric calibration. Although conics are more informative features than points, the loss of the conic property under distortion has critically limited the utility of conic features in camera calibration. Many existing approaches addressed conic-based calibration by ig-noring distortion or introducing 3D spherical targets to cir-cumvent this limitation. In this paper, we present a novel formulation for conic-based calibration using moments. Our derivation is based on the mathematical finding that the first moment can be estimated without bias even under dis-tortion. This allows us to track moment changes during pro-jection and distortion, ensuring the preservation of the first moment of the distorted conic. With an unbiased estima-tor, the circular patterns can be accurately detected at the sub-pixel level and can now be fully exploited for an entire calibration pipeline, resulting in significantly improved cal-ibration. The entire code is readily available from https://github.com/ChaehyeonSong/discocal. Chaehyeon Song, Myung-Hwan Jeon, Jongwoo Lim, Ayoung Kim |
CVPR | 5 |
| 2024 | PeLiCal: Targetless Extrinsic Calibration via Penetrating Lines for RGB-D Cameras with Limited Co-visibilityabstractRGB-D cameras are crucial in robotic perception, given their ability to produce images augmented with depth data. However, their limited field of view (FOV) often requires multiple cameras to cover a broader area. In multi-camera RGB-D setups, the goal is typically to reduce camera overlap, optimizing spatial coverage with as few cameras as possible. The extrinsic calibration of these systems introduces additional complexities. Existing methods for extrinsic calibration either necessitate specific tools or highly depend on the accuracy of camera motion estimation. To address these issues, we present PeLiCal, a novel line-based calibration approach for RGB-D camera systems exhibiting limited overlap. Our method leverages long line features from surroundings, and filters out outliers with a novel convergence voting algorithm, achieving targetless, real-time, and outlier-robust performance compared to existing methods. We open source our implementation on https://github.com/joomeok/PeLiCal.git. Seungsang Yun, Ayoung Kim |
ICRA | 3 |
| 2024 | Co-RaL: Complementary Radar-Leg Odometry with 4-DoF Optimization and Rolling ContactabstractRobust and accurate localization in challenging environments is becoming crucial for SLAM. In this paper, we propose a unique sensor configuration for precise and robust odometry by integrating chip radar and a legged robot. Specifically, we introduce a tightly coupled radar-leg odometry algorithm for complementary drift correction. Adopting the 4-DoF optimization and decoupled RANSAC to mmWave chip radar significantly enhances radar odometry beyond the existing method, especially z-directional even when using a single radar. For the leg odometry, we employ rolling contact modeling-aided forward kinematics, accommodating scenarios with the potential possibility of contact drift and radar failure. We evaluate our method by comparing it with other chip radar odometry algorithms using real-world datasets with diverse environments while the datasets will be released for the robotics community. https://github.com/SangwooJung98/Co-RaL-Dataset Sangwoo Jung 0002, Wooseong Yang, Ayoung Kim |
IROS | 3 |
| 2024 | Camera Agnostic Two-Head Network for Ego-Lane InferenceabstractVision-based ego-lane inference using High-Definition (HD) maps is essential in autonomous driving and advanced driver assistance systems. The traditional approach necessitates well-calibrated cameras, which confines variation of camera configuration, as the algorithm relies on intrinsic and extrinsic calibration. In this paper, we propose a learning-based ego-lane inference by directly estimating the ego-lane index from a single image. To enhance robust performance, our model incorporates the two-head structure inferring ego-lane in two perspectives simultaneously. Furthermore, we utilize an attention mechanism guided by vanishing point-and-line to adapt to changes in viewpoint without requiring accurate calibration. The high adaptability of our model was validated in diverse environments, devices, and camera mounting points and orientations. Chaehyeon Song, Sungho Yoon, Minhyeok Heo, Ayoung Kim, Sujung Kim |
IV | 4 |
| 2024 | A New Wave in Robotics: Survey on Recent MmWave Radar Applications in RoboticsabstractWe survey the current state of millimeter-wave (mmWave) radar applications in robotics with a focus on unique capabilities, and discuss future opportunities based on the state of the art. Frequency modulated continuous wave mmWave radars operating in the 76–81 GHz range are an appealing alternative to lidars, cameras, and other sensors operating in the near-visual spectrum. Radar has been made more widely available in new packaging classes, more convenient for robotics and its longer wavelengths have the ability to bypass visual clutter, such as fog, dust, and smoke. We begin by covering radar principles as they relate to robotics. We then review the relevant new research across a broad spectrum of robotics applications beginning with motion estimation, localization, and mapping. We then cover object detection and classification, and then close with an analysis of current datasets and calibration techniques that provide entry points into radar research. Kyle Harlow, Hyesu Jang, Tim D. Barfoot, Ayoung Kim, Christoffer R. Heckman |
IEEE Trans. Robotics | 4 |
| 2023 | Learning Extended Depth of field Hyperspectral ImagingabstractWe propose a learning-based method for snapshot hyper-spectral (HS) imaging of deep 3D scenes. The method combines computational HS imaging and extended depth of field (EDoF) imaging capabilities in a single framework, resulting in novel EDoF-HS camera designs. The camera system incorporates a diffractive optical element at the aperture position, a CFA in front of the sensor and a residual dense network at the post-processing stage. These optical and neural components are jointly optimized through end-to-end learning procedure. We demonstrate high quality HS image reconstructions for scenes as deep as 4 diopters. Erdem Sahin, Ugur Akpinar, Ayoung Kim, Atanas P. Gotchev |
ICIP | 3 |
| 2023 | Edge-guided Multi-domain RGB-to-TIR image Translation for Training Vision Tasks with Challenging LabelsabstractThe insufficient number of annotated thermal infrared (TIR) image datasets not only hinders TIR image-based deep learning networks to have comparable performances to that of RGB but it also limits the supervised learning of TIR image-based tasks with challenging labels. As a remedy, we propose a modified multidomain RGB to TIR image translation model focused on edge preservation to employ annotated RGB images with challenging labels. Our proposed method not only preserves key details in the original image but also leverages the optimal TIR style code to portray accurate TIR characteristics in the translated image, when applied on both synthetic and real world RGB images. Using our translation model, we have enabled the supervised learning of deep TIR image-based optical flow estimation and object detection that ameliorated in deep TIR optical flow estimation by reduction in end point error by 56.5% on average and the best object detection mAP of 23.9% respectively. Our code and supplementary materials are available at https://github.com/rpmsnu/sRGB-TIR. Dong-Guw Lee, Myung-Hwan Jeon, Younggun Cho, Ayoung Kim |
ICRA | 4 |
| 2023 | RaPlace: Place Recognition for Imaging Radar using Radon Transform and Mutable ThresholdabstractDue to the robustness in sensing, radar has been highlighted, overcoming harsh weather conditions such as fog and heavy snow. In this paper, we present a novel radar-only place recognition that measures the similarity score by utilizing Radon-transformed sinogram images and cross-correlation in frequency domain. Doing so achieves rigid transform invariance during place recognition, while ignoring the effects of radar multipath and ring noises. In addition, we compute the radar similarity distance using mutable threshold to mitigate variability of the similarity score, and reduce the time complexity of processing a copious radar data with hierarchical retrieval. We demonstrate the matching performance for both intra-session loop-closure detection and global place recognition using a publicly available imaging radar datasets. We verify reliable performance compared to existing stable radar place recognition method. Furthermore, codes for the proposed imaging radar place recognition is released for community https://github.com/hyesu-jang/RaPlace. Hyesu Jang, Minwoo Jung, Ayoung Kim |
IROS | 3 |
| 2023 | Multitask Learning for Scalable and Dense Multilayer Bayesian Map InferenceabstractIn this article, we present a novel and flexible multitask multilayer Bayesian mapping framework with readily extendable attribute layers. The proposed framework goes beyond modern metric-semantic maps to provide even richer environmental information for robots in a single mapping formalism while exploiting intralayer and interlayer correlations. It removes the need for a robot to access and process information from many separate maps when performing a complex task, advancing the way robots interact with their environments. To this end, we design a multitask deep neural network with attention mechanisms as our front-end to provide heterogeneous observations for multiple map layers simultaneously. Our back-end runs a scalable closed-form Bayesian inference with only logarithmic time complexity. We apply the framework to build a dense robotic map, including metric-semantic occupancy and traversability layers. Traversability ground truth labels are automatically generated from exteroceptive sensory data in a self-supervised manner. We present extensive experimental results on publicly available datasets and data collected by a three-dimensional bipedal robot platform and show reliable mapping performance in different environments. Finally, we also discuss how the current framework can be extended to incorporate more information, such as friction, signal strength, temperature, and physical quantity concentration using Gaussian map layers. The software for reproducing the presented results or running on customized data is made publicly available. Lu Gan 0006, Youngji Kim, Jessy W. Grizzle, Jeffrey M. Walls, Ayoung Kim, Ryan M. Eustice, Maani Ghaffari Jadidi |
IEEE Trans. Robotics | 5 |
| 2022 | LT-mapper: A Modular Framework for LiDAR-based Lifelong MappingabstractLong-term 3D map management is a fundamental capability required by a robot to reliably navigate in the non-stationary real-world. This paper develops open-source, modular, and readily available LiDAR-based lifelong mapping for urban sites. This is achieved by dividing the problem into successive subproblems: multi-session SLAM (MSS), high/low dynamic change detection, and positive/negative change management. The proposed method leverages MSS and minimizes potential trajectory error; thus, a manual or good initial alignment is not required for change detection. Our change management scheme preserves efficacy in both memory and computation costs, providing automatic object segregation from a large-scale point cloud map. We verify the framework's reliability and applicability even under permanent year-level variation, through extensive real-world experiments with multiple temporal gaps (from day to year). Giseop Kim, Ayoung Kim |
ICRA | 2 |
| 2022 | Sequential thermal image-based adult and baby detection robust to thermal residual heat marksabstractThe awareness for preserving privacy in in-home monitoring robots is increasing. Although several studies have proposed privacy-preserved in-home monitoring robot systems for adults, only a limited amount of attention has been paid attention to research on privacy-preserved in-home monitoring of babies. Like previous studies, thermal infrared image-based methods could ensure a privacy-preserved monitoring of babies, yet when existing detection methods were applied to thermal images to detect babies and adults, we discovered a frequent occurrence of misdetection due to the presence of thermal residual heat marks. In this research, we propose a sequential thermal image-based detection that conjugated the characteristics of thermal residual heat marks. The proposed detection reduced misdetection caused by thermal residual heat marks by 98.7% when compared to RetinaNet. In addition, we open-source our collected thermal image-based baby and adult dataset via: https://github.com/donkeymouse/ThermalAdultandBaby. Dong-Guw Lee, Kyu-Seob Song, Young-Hoon Nho, Ayoung Kim, Dong-Soo Kwon |
IROS | 4 |
| 2022 | STheReO: Stereo Thermal Dataset for Research in Odometry and MappingabstractThis paper introduces a stereo thermal camera dataset (STheReO) with multiple navigation sensors to encourage thermal SLAM researches. A thermal camera measures infrared rays beyond the visible spectrum therefore it could provide a simple yet robust solution to visually degraded environments where existing visual sensor-based SLAM would fail. Existing thermal camera datasets mostly focused on monocular configuration using the thermal camera with RGB cameras in a visually challenging environment. A few stereo thermal rig were examined but in computer vision perspective without supporting sequential images for state estimation algorithms. To encourage the academia for the evolving stereo thermal SLAM, we obtain nine sequences in total across three spatial locations and three different times per location (e.g., morning, day, and night) to capture the variety of thermal characteristics. By using the STheReO dataset, we hope diverse types of researches will be made, including but not limited to odometry, mapping, and SLAM (e.g., thermal-LiDAR mapping or long-term thermal localization). Our datasets are available at https://sites.google.com/view/rpmsthereo/. Seungsang Yun, Minwoo Jung, Jeongyun Kim, Sangwoo Jung 0002, Younghun Cho, Myung-Hwan Jeon, Giseop Kim, Ayoung Kim |
IROS | 8 |
| 2022 | Scan Context++: Structural Place Recognition Robust to Rotation and Lateral Variations in Urban EnvironmentsabstractPlace recognition is a key module in robotic navigation. The existing line of studies mostly focuses on visual place recognition to recognize previously visited places solely based on their appearance. In this article, we address structural place recognition by recognizing a place based on structural appearance, namely from range sensors. Extending our previous work on a rotation invariant spatial descriptor, the proposed descriptor completes a generic descriptor robust to both rotation (heading) and translation when roll–pitch motions are not severe. We introduce two subdescriptors and enable topological place retrieval followed by the 1-degree of freedom semimetric localization, thereby bridging the gap between topological place retrieval and metric localization. The proposed method has been evaluated thoroughly in terms of environmental complexity and scale. The source code is available and can easily be integrated into existing light detection and ranging simultaneous localization and mapping. Giseop Kim, Sunwook Choi, Ayoung Kim |
IEEE Trans. Robotics | 3 |
| 2021 | Multi-session Underwater Pose-graph SLAM using Inter-session Opti-acoustic Two-view FactorabstractConcurrent mapping necessitates data association among vehicles to overcome temporal and sensor modality differences. In this work, we focus on an underwater multi-vehicle mapping scenario in which vehicles have various sensor modalities, namely sonar and camera. This inter-session sonar-optical image matching poses two main challenges. First, ensuring covisibility for the opti-acoustic pair is complex due to their projection models and field of view (FOV) difference. Second, even with secured covisible frames, feature matching over various sensor modalities is not trivial. To overcome these challenges, we complete multi-session simultaneous localization and mapping (SLAM) by introducing an opti-acoustic pairwise factor. We alleviate the covisibility requirement by introducing inter-session measurements. We achieved opti-acoustic feature matching by applying a style-transfer and integration with SuperGlue. The proposed method is validated via simulation and real underwater tank tests. Hyesu Jang, Sungho Yoon, Ayoung Kim |
ICRA | 3 |
| 2021 | EventVLAD: Visual Place Recognition with Reconstructed Edges from Event CamerasabstractEvent cameras are neuromorphic vision sensors that are able to capture high dynamic range with low latency in microseconds, without motion blur. Their strength lies in the unique representation of data as asynchronous events, enabling detection of scene structures less invariantly from dynamic luminance changes. However, a single event does not represent spatial information, and events must be integrated to translate into meaningful information. Therefore, state-of-the-art deep learning algorithms have focused on reconstructing the original scene from events. However, as environmental variances are also captured throughout events and restored in reconstructed images, simple reconstruction does not help achieving robust visual place recognition. In this paper, we suggest to use reconstructed event edges denoised for place recognition. While brightness wavers with dynamic environmental variances, edge contours only change with gradient magnitude scale. We utilize the high dynamic range of event cameras to detect these scaled edges from different environments and show that using reconstructed edges shows robust performance in overcoming day-to-night illumination variance without a large training set. Alex Junho Lee, Ayoung Kim |
IROS | 2 |
| 2021 | Auto-detection of acoustic emission signals from cracking of concrete structures using convolutional neural networks: Upscaling from specimen
Gyeol Han, Tae-Min Oh, Ki-Il Song, Ayoung Kim, Youngchul Kim, Youngtae Cho, Tae-Hyuk Kwon |
Expert Syst. Appl. | 6 |
| 2020 | Unsupervised Geometry-Aware Deep LiDAR OdometryabstractLearning-based ego-motion estimation approaches have recently drawn strong interest from researchers, mostly focusing on visual perception. A few learning-based approaches using Light Detection and Ranging (LiDAR) have been re-ported; however, they heavily rely on a supervised learning manner. Despite the meaningful performance of these approaches, supervised training requires ground-truth pose labels, which is the bottleneck for real-world applications. Differing from these approaches, we focus on unsupervised learning for LiDAR odometry (LO) without trainable labels. Achieving trainable LO in an unsupervised manner, we introduce the uncertainty-aware loss with geometric confidence, thereby al-lowing the reliability of the proposed pipeline. Evaluation on the KITTI, Complex Urban, and Oxford RobotCar datasets demonstrate the prominent performance of the proposed method compared to conventional model-based methods. The proposed method shows a comparable result against SuMa (in KITTI), LeGO-LOAM (in Complex Urban), and Stereo-VO (in Oxford RobotCar). The video and extra-information of the paper are described in https://sites.google.com/view/deeplo. Younggun Cho, Giseop Kim, Ayoung Kim |
ICRA | 3 |
| 2020 | MulRan: Multimodal Range Dataset for Urban Place RecognitionabstractThis paper introduces a multimodal range dataset namely for radio detection and ranging (radar) and light detection and ranging (LiDAR) specifically targeting the urban environment. By extending our workshop paper [1] to a larger scale, this dataset focuses on the range sensor-based place recognition and provides 6D baseline trajectories of a vehicle for place recognition ground truth. Provided radar data support both raw-level and image-format data, including a set of time-stamped 1D intensity arrays and 360° polar images, respectively. In doing so, we provide flexibility between raw data and image data depending on the purpose of the research. Unlike existing datasets, our focus is at capturing both temporal and structural diversities for range-based place recognition research. For evaluation, we applied and validated that our previous location descriptor and its search algorithm [2] are highly effective for radar place recognition method. Furthermore, the result shows that radar-based place recognition outperforms LiDAR-based one exploiting its longer-range measurements. The dataset is available from https://sites.google.com/view/mulran-pr. Giseop Kim, Yeong Sang Park, Younghun Cho, Jinyong Jeong, Ayoung Kim |
ICRA | 5 |
| 2020 | PhaRaO: Direct Radar Odometry using Phase CorrelationabstractRecent studies in radar-based navigation present promising navigation performance using scanning radars. These scanning radar-based odometry methods are mostly feature-based; they detect and match salient features within a radar image. Differing from existing feature-based methods, this paper reports on a method using direct radar odometry, PhaRaO, which infers relative motion from a pair of radar scans via phase correlation. Specifically, we apply the Fourier Mellin transform (FMT) for Cartesian and log-polar radar images to sequentially estimate rotation and translation. In doing so, we decouple rotation and translation estimations in a coarse-to-fine manner to achieve real-time performance. The proposed method is evaluated using large-scale radar data obtained from various environments. The inferred trajectory yields a 2.34% (translation) and 2.93° (rotation) Relative Error (RE) over a 4km path length on average for the odometry estimation. Yeong Sang Park, Young-Sik Shin, Ayoung Kim |
ICRA | 3 |
| 2020 | Remove, then Revert: Static Point cloud Map Construction using Multiresolution Range ImagesabstractWe present a novel static point cloud map construction algorithm, called Removert, for use within dynamic urban environments. Leaving only static points and excluding dynamic objects is a critical problem in various robust robot missions in changing outdoors, and the procedure commonly contains comparing a query to the noisy map that has dynamic points. In doing so, however, the estimated discrepancies between a query scan and the noisy map tend to possess errors due to imperfect pose estimation, which degrades the static map quality. To tackle the problem, we propose a multiresolution range image-based false prediction reverting algorithm. We first conservatively retain definite static points and iteratively recover more uncertain static points by enlarging the query-to- map association window size, which implicitly compensates the LiDAR motion or registration errors. We validate our method on the KITTI dataset using SemanticKITTI as ground truth, and show our method qualitatively competes or outperforms the human-labeled data (SemanticKITTI) in ambiguous regions. Giseop Kim, Ayoung Kim |
IROS | 2 |
| 2020 | Balanced Depth Completion between Dense Depth Inference and Sparse Range Measurements via KISS-GPabstractEstimating a dense and accurate depth map is the key requirement for autonomous driving and robotics. Recent advances in deep learning have allowed depth estimation in full resolution from a single image. Despite this impressive result, many deep-learning-based monocular depth estimation (MDE) algorithms have failed to keep their accuracy yielding a meter-level estimation error. In many robotics applications, accurate but sparse measurements are readily available from Light Detection and Ranging (LiDAR). Although they are highly accurate, the sparsity limits full resolution depth map reconstruction. Targeting the problem of dense and accurate depth map recovery, this paper introduces the fusion of these two modalities as a depth completion (DC) problem by dividing the role of depth inference and depth regression. Utilizing the state-of-the-art MDE and our Gaussian process (GP) based depth-regression method, we propose a general solution that can flexibly work with various MDE modules by enhancing its depth with sparse range measurements. To overcome the major limitation of GP, we adopt Kernel Interpolation for Scalable Structured (KISS)-GP and mitigate the computational complexity from O(N3) to O(N). Our experiments demonstrate that the accuracy and robustness of our method outperform state-of-the-art unsupervised methods for sparse and biased measurements. Sungho Yoon, Ayoung Kim |
IROS | 2 |
| 2020 | Proactive Camera Attribute Control Using Bayesian Optimization for Illumination-Resilient Visual NavigationabstractIllumination variance is a major challenge for vision-based robotics. Most approaches focus on alleviating illumination changes in already captured images. Despite the large utility, camera attributes have been empirically determined to function in a highly passive manner, yielding vision algorithm failure under radical illumination variance. Recent studies have proposed exposure and gain control schemes that could maximize image information and eschew saturation. In this article, we propose a proactive control scheme for the camera's two dominant attributes-exposure time and gain control. Unlike existing approaches, we formulate this camera attribute control as an optimization problem in which the underlying function is not known a priori. We first define a new metric of the image regarding these two major attributes to include both image gradients and signal-to-noise ratio simultaneously. Based on this metric, we introduce a new formulation for this attribute control via Bayesian optimization (BO) and learn the environmental change from the captured image. During the control, to mitigate the burden of image acquisition and Bayesian optimization, images are synthesized using a camera response function and avoided the actual frame grab from the camera. The proposed method was validated in light-flickering indoor, outdoor near sunset, and indoor-outdoor transient environments where light changes rapidly, supporting 20-40 Hz frame rates. Joowan Kim, Younggun Cho, Ayoung Kim |
IEEE Trans. Robotics | 3 |
| 2019 | Radar Localization and Mapping for Indoor Disaster Environments via Multi-modal Registration to Prior LiDAR MapabstractThis paper presents a localization and mapping algorithm that leverages a radar system in low-visibility environments. We aim to address disaster situations in which prior knowledge of a place is available from CAD or light detection and ranging (LiDAR) maps, but incoming visibility is severely limited. In smoky environments, typical sensors (e.g., cameras and LiDARs) fail to perform reliably due to the large particles in the air. Radars recently attracted attention for their robust perception in low-visibility environments; however, radar measurements' angular ambiguity and low resolution prevented the direct application to the simultaneous localization and mapping (SLAM) framework. In this paper, we propose registering radar measurements against a previously built dense LiDAR map for localization and applying radar-map refinement for mapping. Our proposed method overcomes the significant density discrepancy between LiDAR and radar with a density-independent point registration algorithm. We validate the proposed method in an environment containing dense fog. Yeong Sang Park, Joowan Kim, Ayoung Kim |
IROS | 3 |
| 2018 | DHSGAN: An End to End Dehazing Network for Fog and Smoke
Ramavtar Malav, Ayoung Kim, Soumya Ranjan Sahoo, Gaurav Pandey 0004 |
ACCV (5) | 2 |
| 2018 | Complex Urban LiDAR Data SetabstractThis paper presents a Light Detection and Ranging (LiDAR) data set that targets complex urban environments. Urban environments with high-rise buildings and congested traffic pose a significant challenge for many robotics applications. The presented data set is unique in the sense it is able to capture the genuine features of an urban environment (e.g. metropolitan areas, large building complexes and underground parking lots). Data of two-dimensional (2D) and three-dimensional (3D) LiDAR, which are typical types of LiDAR sensors, are provided in the data set. The two 16-ray 3D LiDARs are tilted on both sides for maximal coverage. One 2D LiDAR faces backward while the other faces forwards to collect data of roads and buildings, respectively. Raw sensor data from Fiber Optic Gyro (FOG), Inertial Measurement Unit (IMU), and the Global Positioning System (GPS) are presented in a file format for vehicle pose estimation. The pose information of the vehicle estimated at 100 Hz is also presented after applying the graph simultaneous localization and mapping (SLAM) algorithm. For the convenience of development, the file player and data viewer in Robot Operating System (ROS) environment were also released via the web page. The full data sets are available at: http://irap.kaist.ac.kr/dataset. In this website, 3D preview of each data set is provided using WebGL. Jinyong Jeong, Younggun Cho, Young-Sik Shin, Hyun Chul Roh, Ayoung Kim |
ICRA | 5 |
| 2018 | Exposure Control Using Bayesian Optimization Based on Entropy Weighted Image GradientabstractUnder- and oversaturation can cause severe image degradation in many vision-based robotic applications. To control camera exposure in dynamic lighting conditions, we introduce a novel metric for image information measure. Measuring an image gradient is typical when evaluating its level of image detail. However, emphasizing more informative pixels substantially improves the measure within an image. By using this entropy weighted image gradient, we introduce an optimal exposure value for vision-based approaches. Using this newly invented metric, we also propose an effective exposure control scheme that covers a wide range of light conditions. When evaluating the function (e.g., image frame grab) is expensive, the next best estimation needs to be carefully considered. Through Bayesian optimization, the algorithm can estimate the optimal exposure value with minimal cost. We validated the proposed image information measure and exposure control scheme via a series of thorough experiments using various exposure conditions. Joowan Kim, Younggun Cho, Ayoung Kim |
ICRA | 3 |
| 2018 | Direct Visual SLAM Using Sparse Depth for Camera-LiDAR SystemabstractThis paper describes a framework for direct visual simultaneous localization and mapping (SLAM) combining a monocular camera with sparse depth information from Light Detection and Ranging (LiDAR). To ensure realtime performance while maintaining high accuracy in motion estimation, we present (i) a sliding window-based tracking method, (ii) strict pose marginalization for accurate pose-graph SLAM and (iii) depth-integrated frame matching for large-scale mapping. Unlike conventional feature-based visual and LiDAR mapping, the proposed approach is direct, eliminating the visual feature in the objective function. We evaluated results using our portable camera-LiDAR system as well as KITTI odometry benchmark datasets. The experimental results prove that the characteristics of two complementary sensors are very effective in improving real-time performance and accuracy. Via validation, we achieved low drift error of 0.98 % in the KITTI benchmark including various environments such as a highway and residential areas. Young-Sik Shin, Yeong Sang Park, Ayoung Kim |
ICRA | 3 |
| 2018 | Stereo Camera Localization in 3D LiDAR MapsabstractAs simultaneous localization and mapping (SLAM) techniques have flourished with the advent of 3D Light Detection and Ranging (LiDAR) sensors, accurate 3D maps are readily available. Many researchers turn their attention to localization in a previously acquired 3D map. In this paper, we propose a novel and lightweight camera-only visual positioning algorithm that involves localization within prior 3D LiDAR maps. We aim to achieve the consumer level global positioning system (GPS) accuracy using vision within the urban environment, where GPS signal is unreliable. Via exploiting a stereo camera, depth from the stereo disparity map is matched with 3D LiDAR maps. A full six degree of freedom (DOF) camera pose is estimated via minimizing depth residual. Powered by visual tracking that provides a good initial guess for the localization, the proposed depth residual is successfully applied for camera pose estimation. Our method runs online, as the average localization error is comparable to ones resulting from state-of-the-art approaches. We validate the proposed method as a stand-alone localizer using KITTI dataset and as a module in the SLAM framework using our own dataset. Youngji Kim, Jinyong Jeong, Ayoung Kim |
IROS | 3 |
| 2018 | Scan Context: Egocentric Spatial Descriptor for Place Recognition Within 3D Point Cloud MapabstractCompared to diverse feature detectors and descriptors used for visual scenes, describing a place using structural information is relatively less reported. Recent advances in simultaneous localization and mapping (SLAM) provides dense 3D maps of the environment and the localization is proposed by diverse sensors. Toward the global localization based on the structural information, we propose Scan Context, a non-histogram-based global descriptor from 3D Light Detection and Ranging (LiDAR) scans. Unlike previously reported methods, the proposed approach directly records a 3D structure of a visible space from a sensor and does not rely on a histogram or on prior training. In addition, this approach proposes the use of a similarity score to calculate the distance between two scan contexts and also a two-phase search algorithm to efficiently detect a loop. Scan context and its search algorithm make loop-detection invariant to LiDAR viewpoint changes so that loops can be detected in places such as reverse revisit and corner. Scan context performance has been evaluated via various benchmark datasets of 3D LiDAR scans, and the proposed method shows a sufficiently improved performance. Giseop Kim, Ayoung Kim |
IROS | 2 |
| 2017 | Visibility enhancement for underwater visual SLAM based on underwater light scattering modelabstractThis paper presents a real-time visibility enhancement algorithm for effective underwater visual simultaneous localization and mapping (SLAM). Unlike an aerial environment, an underwater environment contains larger particles and is dominated by a different image degradation model. Our method starts with a thorough understanding of underwater particle physics (e.g., forward, back, multiple scattering, blur and noise). Targeting underwater image enhancement in a real-world application, we include an artificial light model in the derivation. The proposed method is effective for both color and gray images with substantial improvement in the process time compared to conventional methods. The proposed method is validated by using simulated synthetic images (color) and real-world underwater images (color and grayscale). Using two underwater image sets acquired from the same area but with different water turbidity, we evaluate the proposed visibility enhancement and camera registration improvement in SLAM. Younggun Cho, Ayoung Kim |
ICRA | 2 |
| 2017 | On the uncertainty propagation: Why uncertainty on lie groups preserves monotonicity?abstractResearchers in the simultaneous localization and mapping (SLAM) community have taken for granted that uncertainty associated with the robot pose increases until the loop is closed. However, recently identified by [1], the monotonicity of uncertainty during exploration breaks when the robot returns to the initial position. In this paper, we propose a hypothesis that the monotonicity of pose uncertainty is preserved when the uncertainty is propagated on Lie groups rather than on Euclidean vector space. After deriving covariance propagated over Lie groups and Euclidean vector space, respectively, the monotonicity of uncertainty in each case is thoroughly investigated. Experiments with simulated and real-world scenarios on dead-reckoning validate our hypothesis on the monotonicity of uncertainty. Youngji Kim, Ayoung Kim |
IROS | 2 |
| 2017 | Road-SLAM : Road marking based SLAM with lane-level accuracyabstractIn this paper, we propose the Road-SLAM algorithm, which robustly exploits road markings obtained from camera images. Road markings are well categorized and informative but susceptible to visual aliasing for global localization. To enable loop-closures using road marking matching, our method defines a feature consisting of road markings and surrounding lanes as a sub-map. The proposed method uses random forest method to improve the accuracy of matching using a sub-map containing road information. The random forest classifies road markings into six classes and only incorporates informative classes to avoid ambiguity. The proposed method is validated by comparing the SLAM result with RTK-Global Positioning System (GPS) data. Accurate loop detection improves global accuracy by compensating for cumulative errors in odometry sensors. This method achieved an average global accuracy of 1.098 m over 4.7 km of path length, while running at real-time performance. Jinyong Jeong, Younggun Cho, Ayoung Kim |
Intelligent Vehicles Symposium | 3 |
| 2016 | Simultaneous segmentation, estimation and analysis of articulated motion from dense point cloud sequenceabstractIn this paper, we present a unified approach for Expectation Maximization (EM) based motion segmentation, estimation and analysis from dense point cloud data. When identifying an underlying motion, literature mainly focuses on three related topics: motion segmentation, estimation and analysis. These topics are, however, mostly considered separately while integrated approaches are rare. Our approach specifically focuses on analyzing articulated motion from dense point cloud data by simultaneously solving for three topics using an integrated approach. No prior knowledge, such as background regions, number of segments and correspondence, is required since two iterations in this algorithm allow us to seamlessly accomplish integration of the three tasks. The first iteration of the algorithm is performed between segmentation and estimation, followed by the second iteration between motion estimation and analysis. For the first iteration, we propose EM based subspace clustering algorithm. For the second iteration, we simply fuse the motion analysis method from [1] into an iterative motion estimation algorithm. As a result, we can extract label, correspondence and motion of moving objects simultaneously from dense point cloud sequence. In experiment, we validate the performance of the proposed method on both synthetic and real world data. Youngji Kim, Hwasup Lim, Sang Chul Ahn, Ayoung Kim |
IROS | 4 |
| 2014 | Opportunistic sampling-based planning for active visual SLAMabstractThis paper reports on an active visual SLAM path planning algorithm that plans loop-closure paths in order to decrease visual navigation uncertainty. Loop-closing revisit actions bound the robot's uncertainty but also contribute to redundant area coverage and increased path length.We propose an opportunistic path planner that leverages sampling-based techniques and information filtering for planning revisit paths that are coverage efficient. Our algorithm employs Gaussian Process regression for modeling the prediction of camera registrations and uses a two-step optimization for selecting revisit actions. We show that the proposed method outperforms existing solutions for bounding navigation uncertainty with a hybrid simulation experiment using a real-world dataset collected by a ship hull inspection robot. Stephen M. Chaves, Ayoung Kim, Ryan M. Eustice |
IROS | 2 |
| 2013 | Perception-driven navigation: Active visual SLAM for robotic area coverageabstractThis paper reports on an integrated navigation algorithm for the visual simultaneous localization and mapping (SLAM) robotic area coverage problem. In the robotic area coverage problem, the goal is to explore and map a given target area in a reasonable amount of time. This goal necessitates the use of minimally redundant overlap trajectories for coverage efficiency; however, visual SLAM's navigation estimate will inevitably drift over time in the absence of loop-closures. Therefore, efficient area coverage and good SLAM navigation performance represent competing objectives. To solve this decision-making problem, we introduce perception-driven navigation (PDN), an integrated navigation algorithm that automatically balances between exploration and revisitation using a reward framework. This framework accounts for vehicle localization uncertainty, area coverage performance, and the identification of good candidate regions in the environment for loop-closure. Results are shown for a hybrid simulation using synthetic and real imagery from an autonomous underwater ship hull inspection application. Ayoung Kim, Ryan M. Eustice |
ICRA | 1 |
| 2013 | Real-Time Visual SLAM for Autonomous Underwater Hull Inspection Using Visual SaliencyabstractThis paper reports a real-time monocular visual simultaneous localization and mapping (SLAM) algorithm and results for its application in the area of autonomous underwater ship hull inspection. The proposed algorithm overcomes some of the specific challenges associated with underwater visual SLAM, namely, limited field of view imagery and feature-poor regions. It does so by exploiting our SLAM navigation prior within the image registration pipeline and by being selective about which imagery is considered informative in terms of our visual SLAM map. A novel online bag-of-words measure for intra and interimage saliency are introduced and are shown to be useful for image key-frame selection, information-gain-based link hypothesis, and novelty detection. Results from three real-world hull inspection experiments evaluate the overall approach, including one survey comprising a 3.4-h/2.7-km-long trajectory. Ayoung Kim, Ryan M. Eustice |
IEEE Trans. Robotics | 1 |
| 2011 | Combined visually and geometrically informative link hypothesis for pose-graph visual SLAM using bag-of-wordsabstractThis paper reports on a method to combine expected information gain with visual saliency scores in order to choose geometrically and visually informative loop-closure candidates for pose-graph visual simultaneous localization and mapping (SLAM). Two different bag-of-words saliency metrics are introduced—global saliency and local saliency. Global saliency measures the rarity of an image throughout the entire data set, while local saliency describes the amount of texture richness in an image. The former is important in measuring an overall global saliency map for a given area, and is motivated from inverse document frequency (a measure of rarity) in information retrieval. Local saliency is defined by computing the entropy of the bag-of-words histogram, and is useful to avoid adding visually benign key frames to the map. The two different metrics are presented and experimentally evaluated with indoor and underwater imagery to verify their utility. Ayoung Kim, Ryan M. Eustice |
IROS | 1 |
| 2009 | Pose-graph visual SLAM with geometric model selection for autonomous underwater ship hull inspectionabstractThis paper reports the application of vision based simultaneous localization and mapping (SLAM) to the problem of autonomous ship hull inspection by an underwater vehicle. The goal of this work is to automatically map and navigate the underwater surface area of a ship hull for foreign object detection and maintenance inspection tasks. For this purpose we employ a pose-graph SLAM algorithm using an extended information filter for inference. For perception, we use a calibrated monocular camera system mounted on a tilt actuator so that the camera approximately maintains a nadir view to the hull. A combination of SIFT and Harris features detectors are used within a pairwise image registration framework to provide camera-derived relative-pose constraints (modulo scale). Because the ship hull surface can vary from being locally planar to highly three-dimensional (e.g., screws, rudder), we employ a geometric model selection framework to appropriately choose either an essential matrix or homography registration model during image registration. This allows the image registration engine to exploit geometry information at the early stages of estimation, which results in better navigation and structure reconstruction via more accurate and robust camera-constraints. Preliminary results are reported for mapping a 1,300 image data set covering a 30 m by 5 m section of the hull of a USS aircraft carrier. The post-processed result validates the algorithm's potential to provide in-situ navigation in the underwater environment for trajectory control, while generating a texture-mapped 3D model of the ship hull as a byproduct for inspection. Ayoung Kim, Ryan M. Eustice |
IROS | 1 |