EDBT 2026 Demo / reviewers in the wild / expert
Bhoram Lee
dblp:94/11301
· DBLP profile ↗
13ranked-venue papers
7as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 4 since 2021Systems, architecture and hardware · 11 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SayNav: Grounding Large Language Models for Dynamic Planning to Navigation in New EnvironmentsabstractSemantic reasoning and dynamic planning capabilities are crucial for an autonomous agent to perform complex navigation tasks in unknown environments. It requires a large amount of common-sense knowledge, that humans possess, to succeed in these tasks. We present SayNav, a new approach that leverages human knowledge from Large Language Models (LLMs) for efficient generalization to complex navigation tasks in unknown large-scale environments. SayNav uses a novel grounding mechanism, that incrementally builds a 3D scene graph of the explored environment as inputs to LLMs, for generating feasible and contextually appropriate high-level plans for navigation. The LLM-generated plan is then executed by a pre-trained low-level planner, that treats each planned step as a short-distance point-goal navigation sub-task. SayNav dynamically generates step-by-step instructions during navigation and continuously refines future steps based on newly perceived information. We evaluate SayNav on multi-object navigation (MultiON) task, that requires the agent to utilize a massive amount of human knowledge to efficiently search multiple different objects in an unknown environment. We also introduce a benchmark dataset for MultiON task employing ProcTHOR framework that provides large photo-realistic indoor environments with variety of objects. SayNav achieves state-of-the-art results and even outperforms an oracle based baseline with strong ground-truth assumptions by more than 8% in terms of success rate, highlighting its ability to generate dynamic plans for successfully locating objects in large-scale new environments. The code, benchmark dataset and demonstration videos are accessible at https://www.sri.com/ics/computer-vision/saynav. Abhinav Rajvanshi, Karan Sikka, Bhoram Lee, Han-Pang Chiu, Alvaro Velasquez |
ICAPS | 4 |
| 2024 | VioLA: Aligning Videos to 2D LiDAR ScansabstractWe study the problem of aligning a video that captures a local portion of an environment to the 2D LiDAR scan of the entire environment. We introduce a method (VioLA) that starts with building a semantic map of the local scene from the image sequence, then extracts points at a fixed height for registering to the LiDAR map. Due to reconstruction errors or partial coverage of the camera scan, the reconstructed semantic map may not contain sufficient information for registration. To address this problem, VioLA makes use of a pre-trained text-to-image inpainting model paired with a depth completion model for filling in the missing scene content in a geometrically consistent fashion to support pose registration. We evaluate VioLA on two real-world RGB-D benchmarks, as well as a self-captured dataset of a large office scene. Notably, our proposed scene completion module improves the pose registration performance by up to 20%. Jun-Jee Chao, Kazim Selim Engin, Nikhil Chavan Dafle, Bhoram Lee, Volkan Isler |
ICRA | 4 |
| 2024 | HIO-SDF: Hierarchical Incremental Online Signed Distance FieldsabstractA good representation of a large, complex mobile robot workspace must be space-efficient yet capable of encoding relevant geometric details. When exploring unknown environments, it needs to be updatable incrementally in an online fashion. We introduce HIO-SDF, a new method that represents the environment as a Signed Distance Field (SDF). State of the art representations of SDFs are based on either neural networks or voxel grids. Neural networks are capable of representing the SDF continuously. However, they are hard to update incrementally as neural networks tend to forget previously observed parts of the environment unless an extensive sensor history is stored for training. Voxel-based representations do not have this problem but they are not space-efficient especially in large environments with fine details. HIO-SDF combines the advantages of these representations using a hierarchical approach which employs a coarse voxel grid that captures the observed parts of the environment together with high-resolution local information to train a neural network. HIO-SDF achieves a 46% lower mean global SDF error across all test scenes than a state of the art continuous representation, and a 30% lower error than a discrete representation at the same resolution as our coarse global SDF grid. Videos and code are available at: https://samsunglabs.github.io/HIO-SDF-project-page/ Vasileios Vasilopoulos, Suveer Garg, Jinwook Huh, Bhoram Lee, Volkan Isler |
ICRA | 4 |
| 2022 | Towards Safe, Realistic Testbed for Robotic Systems with Human InteractionabstractSimulation has been a necessary, safe testbed for robotics systems (RS). However, testing in simulation alone is not enough for robotic systems operating in close proximity, or interacting directly with, humans, because simulated humans are very limited. Furthermore, testing with real humans can be unsafe and costly. As recent advances in machine learning are being brought to physical robotic systems, how to collect data as well as evaluate them with human interactions safely yet realistically is a critical question. This paper presents a Mixed-Reality (MR) system toward human-centered development of robotic systems emphasizing benefits as a data collection and testbed tool. MR testbeds allow humans to interact with various levels of virtuality to maintain both realism and safety. We detail the advantages and limitations of these different levels of realism or virtualization, and report our MR-based RS testbed implemented using off-the-shelf MR devices with the Unity game engine and ROS. We demonstrate our testbed in a multi-robot, multi-person tracking and monitoring application. We share our vision and insights earned during the development and data collection. Bhoram Lee, Jonathan Brookshire, Rhys Yahata, Supun Samarasekera |
ICRA | 1 |
| 2019 | Online Continuous Mapping using Gaussian Process Implicit SurfacesabstractThe representation of the environment strongly affects how robots can move and interact with it. This paper presents an online approach for continuous mapping using Gaussian Process Implicit Surfaces (GPISs). Compared with grid-based methods, GPIS better utilizes sparse measurements to represent the world seamlessly. It provides direct access to the signed-distance function (SDF) and its derivatives which are invaluable for other robotic tasks and it incorporates uncertainty in the sensor measurements. Our approach incrementally and efficiently updates GPIS by employing a regressor on observations and a spatial tree structure. The effectiveness of the suggested approach is demonstrated using simulations and real world 2D/3D data. Bhoram Lee, Clark Zhang, Zonghao Huang, Daniel D. Lee |
ICRA | 1 |
| 2019 | Pixels to Plans: Learning Non-Prehensile Manipulation by Imitating a PlannerabstractWe present a novel method enabling robots to quickly learn to manipulate objects by leveraging a motion planner to generate “expert” training trajectories from a small amount of human-labeled data. In contrast to the traditional sense-plan-act cycle, we propose a deep learning architecture and training regimen called PtPNet that can estimate effective end-effector trajectories for manipulation directly from a single RGB-D image of an object. Additionally, we present a data collection and augmentation pipeline that enables the automatic generation of large numbers (millions) of training image and trajectory examples with almost no human labeling effort.We demonstrate our approach in a non-prehensile tool-based manipulation task, specifically picking up shoes with a hook. In hardware experiments, PtPNet generates motion plans (open-loop trajectories) that reliably (89% success over 189 trials) pick up four very different shoes from a range of positions and orientations, and reliably picks up a shoe it has never seen before. Compared with a traditional sense-plan-act paradigm, our system has the advantages of operating on sparse information (single RGB-D frame), producing high-quality trajectories much faster than the expert planner (300ms versus several seconds), and generalizing effectively to previously unseen shoes. Video available at https://youtu.be/voIkyiBtwn4. Tarik Tosun, Eric Mitchell, Ben Eisner, Jinwook Huh, Bhoram Lee, Dae-Won Lee, Volkan Isler, H. Sebastian Seung, Daniel D. Lee |
IROS | 5 |
| 2018 | Constrained Sampling-Based Planning for Grasping and ManipulationabstractThis paper presents a novel constrained, sampling-based motion planning method for grasp and transport tasks with a redundant robotic manipulator. We utilize a planning margin for grasping with constraints that allow the best grasp configuration and approach direction to be determined automatically. For manipulators with many degrees of freedom, our method efficiently chooses the optimal grasp pose when there are many redundant solutions. The method also introduces a parameterized intermediate pose that is optimized to determine the approach direction, increasing robustness under sensor uncertainty and execution errors. Our method also considers transporting the grasped object to the desired target position using a Rapidly-exploring Random Tree (RRT) algorithm that incorporates soft constraints via appropriate cost penalties. We demonstrate the effectiveness and efficiency of our algorithms on a number of simulated and experimental applications. Our experimental results show a marked improvement in computational efficiency in comparison to previously studied approaches. Jinwook Huh, Bhoram Lee, Daniel D. Lee |
ICRA | 2 |
| 2017 | Adaptive motion planning with high-dimensional mixture modelsabstractThis paper presents a novel adaptive approach to fast sampling-based motion planning by learning models of collision and collision-free regions in configuration spaces in an online manner. The proposed approach incrementally learns Gaussian Mixture Models (GMMs) for collision detection in high dimensional configuration spaces. In practical applications for robotic manipulation, the representation of collision and collision-free regions in configuration space can change due to relative motion between the robot base and workspace. We show how to rapidly adapt to such changes by using inverse kinematics to transform the parameters of the Gaussian mixture model to new configurations. The transformed model is initially used as a prior and then continually updated and refined as the RRT planning algorithm proceeds in real-time. This approach is extremely computationally efficient, and our proposed method is compared with traditional sampling-based planning methods on a number of experimental robot arm planning scenarios. Jinwook Huh, Bhoram Lee, Daniel D. Lee |
ICRA | 2 |
| 2017 | Self-supervised online learning of appearance for 3D trackingabstractThis paper presents a self-supervised online learning approach for 3D object tracking that requires no pretraining of appearance. Our method focuses on selecting the most relevant parts of the RGBD input by continuously updating appearance classifiers in conjunction with the spatial occupancy of the target. Fine-grained regions selected via the learned bottom-up saliency, together with spatial cues of the 3D shape model, are used to identify and localize the target via shape registration. The subsequent 3-D pose estimate along with positive and negative labels from the registration are used for online learning appearance. The proposed method outperforms competing model-based tracking algorithms on public datasets as well as on a new motion scene dataset that we have collected. Bhoram Lee, Daniel D. Lee |
IROS | 1 |
| 2016 | Learning anisotropic ICP (LA-ICP) for robust and efficient 3D registrationabstractThis paper presents an online learning approach to 3D object registration that vastly improves the performance of Iterative Closest Point (ICP) methods. Our approach achieves better robustness and stable convergence by learning generalized distance functions directly from a stream of object depth data. The proposed algorithm, Learning Anisotropic ICP (LA-ICP), parameterizes the point uncertainty of the underlying object surface as an anisotropic Gaussian, and estimates the covariance parameters of the likelihood function for ICP from data. Our learning scheme does not require manual tuning and the parameters of the algorithm are continually updated from observed data. Experiments on various RGB-D object datasets demonstrate the effectiveness of our approach in terms of convergence and pose accuracy as well as robustness to initial conditions. Bhoram Lee, Daniel D. Lee |
ICRA | 1 |
| 2016 | Online learning of visibility and appearance for object pose estimationabstractThis paper presents an online self-supervised approach to improve the quality and relevance of input point cloud to a 3D registration algorithm. The suggested method considers the visibility of the model points and learns discriminative appearance of the object under gradual changes. It selectively reduces the amount of information to process by excluding non-visible points of the model and removing outliers from data stream, which results in better alignment between the input data and the model. Thus, by providing a good initial pose, it speeds up the iterative procedure of EM-like optimization for pose estimation (i.e., ICP) to achieve better efficiency and robustness. We compiled a new object dataset of RGBD images under camera motion with ground truth poses of the camera and the objects. We have performed experiments on this dataset and obtained promising results. Bhoram Lee, Daniel D. Lee |
IROS | 1 |
| 2015 | Online self-supervised monocular visual odometry for ground vehiclesabstractThis paper presents an online self-supervised approach to monocular visual odometry and ground classification applied to ground vehicles. We solve the motion and structure problem based on a constrained kinematic model. The true scale of the monocular scene is recovered by estimating the ground surface. We consider a general parametric ground surface model and use the Random Sample Consensus (RANSAC) algorithm for robust fitting of the parameters. The estimated ground surface provides training samples to learn a probabilistic appearance-based ground classifier in an online and self-supervised manner. The appearance-based classifier is then used to bias the RANSAC sampling to generate better hypotheses for parameter estimation of the ground surface model. Thus, without relying on any prior information, we combine geometric estimates with appearance-based classification to achieve an online self-learning scheme from monocular vision. Experimental results demonstrate that online learning improves the computational efficiency and accuracy compared to standard sampling in RANSAC. Evaluations on the KITTI benchmark dataset demonstrate the stability and accuracy of our overall methods in comparison to previous approaches. Bhoram Lee, Kostas Daniilidis, Daniel D. Lee |
ICRA | 1 |
| 2012 | Evaluation of human tangential force input performanceabstractWhile interacting with mobile devices, users may press against touch screens and also exert tangential force to the display in a sliding manner. We seek to guide UI design based on the tangential force applied by a user to the surface of a hand-held device. A prototype of an interface using tangential force input was implemented utilizing a force sensitive layer and an elastic layer and used for the user experiment. We investigated user controllability to reach and maintain target force levels and considered the effects of hand pose and direction of force input. Our results imply no significant difference in performance when applying force holding the device in one hand and in two hands. We also observed that users have more physical and perceived loads when applying tangential force in the left-right direction compared to the up-down direction. Based on the experimental results, we discuss considerations for user interface applications of tangential-force-based interface. Bhoram Lee, Hyunjeong Lee, Soo-Chul Lim, Hyungkew Lee, Seungju Han 0001, Joonah Park |
CHI | 1 |