EDBT 2026 Demo / reviewers in the wild / expert
Changhyun Choi
dblp:30/1218
· DBLP profile ↗
42ranked-venue papers
15as first author
21since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 10 first-author · 15 since 2021Systems, architecture and hardware · 28 · 10 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Routing Manipulation of Deformable Linear Object Using Reinforcement Learning and Diffusion PolicyabstractTasks involving deformable linear objects (DLOs) are prevalent in daily life but pose significant challenges due to their infinite degrees of freedom and underactuated nature. Frequent contact between DLOs and surrounding objects with unknown physical parameters, such as friction, further complicates their manipulation. Performing tasks like routing ropes through a hole requires gentle yet robust manipulation, making it particularly challenging. Previous research has not adequately addressed general DLO manipulation tasks that involve intensive contact, especially in environments with rough surfaces. This paper presents a robust and delicate manipulation learning approach for the DLO routing task, leveraging reinforcement learning (RL) and diffusion policy. First, reinforcement learning agents are trained separately for rope insertion and pulling. During training, the agents are encouraged to minimize rope tension throughout task execution in environments with randomized friction to achieve delicate motion. Next, the rollouts from these agents are collected as expert demonstrations to train a diffusion policy. Our approach generates delicate motions to prevent the rope from being damaged or getting stuck on rough surfaces while remaining robust against environmental disturbances. Please refer to our project page: https://lmeee.github.io/DLOPull/ Mingen Li, Houjian Yu, Changhyun Choi |
ICRA | 3 |
| 2025 | A Parameter-Efficient Tuning Framework for Language-Guided Object Grounding and Robot GraspingabstractThe language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their extensive computation and data demands limit the feasibility of local deployment and customization. To address this, we propose a novel CLIP-based [1] multimodal parameter-efficient tuning (PET) framework designed for three language-guided object grounding and grasping tasks: (1) Referring Expression Segmentation (RES), (2) Referring Grasp Synthesis (RGS), and (3) Referring Grasp Affordance (RGA). Our approach introduces two key innovations: a bi-directional vision-language adapter that aligns multimodal inputs for pixel-level language understanding and a depth fusion branch that incorporates geometric cues to facilitate robot grasping predictions. Experiment results demonstrate superior performance in the RES object grounding task compared with existing CLIP-based full-model tuning or PET approaches. In the RGS and RGA tasks, our model not only effectively interprets object attributes based on simple language descriptions but also shows strong potential for comprehending complex spatial reasoning scenarios, such as multiple identical objects present in the workspace. Project page: https://z.umn.edu/etog-etrg Houjian Yu, Mingen Li, Alireza Rezazadeh, Yang Yang 0083, Changhyun Choi |
ICRA | 5 |
| 2025 | InvSlotGNN: Unsupervised Discovery of Viewpoint Invariant Multiobject Representations and Visual DynamicsabstractLearning multiobject dynamics purely from visual data is challenging due to the need for robust object representations that can be learned through robot interactions. In previous work (Rezazadeh et al., 2023), we introduced two novel architectures: SlotTransport for discovering object-centric representations from singleview RGB images, referred to as slots, and SlotGNN for predicting scene dynamics from singleview RGB images and robot interactions using the discovered slots. This article introduces InvSlotGNN, a novel framework for learning multiview slot discovery and dynamics that are invariant to the camera viewpoint. First, we demonstrate that SlotTransport can be trained on multiview data such that a single model discovers temporally aligned, object-centric representations from a wide range of different camera angles. These slots bind to objects from various viewpoints, even under occlusion or absence. Next, we introduce InvSlotGNN, an extension of SlotGNN, that learns multiobject dynamics invariant to the camera angle and predicts the future state from observations taken by uncalibrated cameras. InvSlotGNN learns a graph representation of the scene using the slots from SlotTransport and performs relational and spatial reasoning to predict the future state of the scene for arbitrary viewpoints, conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning multiview object-centric features that accurately encode visual and positional information. Furthermore, we highlight the accuracy of InvSlotGNN in downstream robotic tasks, including long-horizon prediction and multiobject rearrangement. Finally, with minimal real data, our framework robustly predicts slots and their dynamics in real-world multiview scenarios. Alireza Rezazadeh, Houjian Yu, Karthik Desingh, Changhyun Choi |
IEEE Trans. Robotics | 4 |
| 2024 | Learning for Deformable Linear Object Insertion Leveraging Flexibility Estimation from Visual CuesabstractManipulation of deformable Linear objects (DLOs), including iron wire, rubber, silk, and nylon rope, is ubiquitous in daily life. These objects exhibit diverse physical properties, such as Young’s modulus and bending stiffness. Such diversity poses challenges for developing generalized manipulation policies. However, previous research limited their scope to single-material DLOs and engaged in time-consuming data collection for the state estimation. In this paper, we propose a two-stage manipulation approach consisting of a material property (e.g., flexibility) estimation and policy learning for DLO insertion with reinforcement learning. Firstly, we design a flexibility estimation scheme that characterizes the properties of different types of DLOs. The ground truth flexibility data is collected in simulation to train our flexibility estimation module. During the manipulation, the robot interacts with the DLOs to estimate flexibility by analyzing their visual configurations. Secondly, we train a policy conditioned on the estimated flexibility to perform challenging DLO insertion tasks. Our pipeline trained with diverse insertion scenarios achieves an 85.6% success rate in simulation and 66.67% in real robot experiments. Please refer to our project page: https://lmeee.github.io/DLOInsert/ Mingen Li, Changhyun Choi |
ICRA | 2 |
| 2024 | SlotGNN: Unsupervised Discovery of Multi-Object Representations and Visual DynamicsabstractLearning multi-object dynamics from visual data using unsupervised techniques is challenging due to the need for robust, object representations that can be learned through robot interactions. This paper presents a novel framework with two new architectures: SlotTransport for discovering object representations from RGB images and SlotGNN for predicting their collective dynamics from RGB images and robot interactions. Our SlotTransport architecture is based on slot attention for unsupervised object discovery and uses a feature transport mechanism to maintain temporal alignment in object-centric representations. This enables the discovery of slots that consistently reflect the composition of multi-object scenes. These slots robustly bind to distinct objects, even under heavy occlusion or absence. Our SlotGNN, a novel unsupervised graph-based dynamics model, predicts the future state of multi-object scenes. SlotGNN learns a graph representation of the scene using the discovered slots from SlotTransport and performs relational and spatial reasoning to predict the future appearance of each slot conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning object-centric features that accurately encode both visual and positional information. Further, we highlight the accuracy of SlotGNN in downstream robotic tasks, including challenging multi-object rearrangement and long-horizon prediction. Finally, our unsupervised approach proves effective in the real world. With only minimal additional data, our framework robustly predicts slots and their corresponding dynamics in real-world control tasks. Our project webpage: bit.ly/slotgnn. Alireza Rezazadeh, Athreyi Badithela, Karthik Desingh, Changhyun Choi |
ICRA | 4 |
| 2024 | Attribute-Based Robotic Grasping With Data-Efficient AdaptationabstractRobotic grasping is one of the most fundamental robotic manipulation tasks and has been the subject of extensive research. However, swiftly teaching a robot to grasp a novel target object in clutter remains challenging. This paper attempts to address the challenge by leveraging object attributes that facilitate recognition, grasping, and rapid adaptation to new domains. In this work, we present an end-to-end encoder-decoder network to learn attribute-based robotic grasping with data-efficient adaptation capability. We first pre-train the end-to-end model with a variety of basic objects to learn generic attribute representation for recognition and grasping. Our approach fuses the embeddings of a workspace image and a query text using a gated-attention mechanism and learns to predict instance grasping affordances. To train the joint embedding space of visual and textual attributes, the robot utilizes object persistence before and after grasping. Our model is self-supervised in a simulation that only uses basic objects of various colors and shapes but generalizes to novel objects in new environments. To further facilitate generalization, we propose two adaptation methods, adversarial adaption and one-grasp adaptation. Adversarial adaptation regulates the image encoder using augmented data of unlabeled images, whereas one-grasp adaptation updates the overall end-to-end model using augmented data from one grasp trial. Both adaptation methods are data-efficient and considerably improve instance grasping performance. Experimental results in both simulation and the real world demonstrate that our approach achieves over 81% instance grasping success rate on unknown objects, which outperforms several baselines by large margins. Supplementary material is available athttps://z.umn.edu/attr-grasp. Yang Yang 0083, Houjian Yu, Xibai Lou, Yuanhao Liu 0003, Changhyun Choi |
IEEE Trans. Robotics | 5 |
| 2023 | FOGL: Federated Object Grasping LearningabstractFederated learning is a promising technique for training global models in a data-decentralized environment. In this paper, we propose a federated learning approach for robotic object grasping. The main challenge is that the data collected by multiple robots deployed in different environments tends to form heterogeneous data distributions (i.e., non-IID) and that the existing federated learning methods on such data distributions show serious performance degradation. To tackle this problem, we propose federated object grasping learning (FOGL) that uses cross-evaluation in a general federated learning process to assess the training performance of robots. We cluster robots with similar training patterns and perform independent federated learning on each cluster. Finally, we integrate the global models for each cluster through an ensemble inference. We apply FOGL to various federated learning scenarios in robotic object grasping and show state-of-the-art performance on the Cornell grasping dataset. Seok-Kyu Kang, Changhyun Choi |
ICRA | 2 |
| 2023 | Adversarial Object Rearrangement in Constrained Environments with Heterogeneous Graph Neural NetworksabstractAdversarial object rearrangement in the real world (e.g., previously unseen or oversized items in kitchens and stores) could benefit from understanding task scenes, which inherently entail heterogeneous components such as current objects, goal objects, and environmental constraints. The semantic relationships among these components are distinct from each other and crucial for multi-skilled robots to perform efficiently in everyday scenarios. We propose a hierarchical robotic manipulation system that learns the underlying relationships and maximizes the collaborative power of its diverse skills (e.g., PICK-PLACE, PUSH) for rearranging adversarial objects in constrained environments. The high-level coordinator employs a heterogeneous graph neural network (HetGNN), which reasons about the current objects, goal objects, and environmental constraints; the low-level 3D Convolutional Neural Network-based actors execute the action primitives. Our approach is trained entirely in simulation, and achieved an average success rate of 87.88% and a planning cost of 12.82 in real-world experiments, surpassing all baseline methods. Supplementary material is available at https://sites.google.com/umn.edu/versatile-rearrangement. Xibai Lou, Houjian Yu, Ross Worobel, Yang Yang 0083, Changhyun Choi |
IROS | 5 |
| 2023 | IOSG: Image-Driven Object Searching and GraspingabstractWhen robots retrieve specific objects from cluttered scenes, such as home and warehouse environments, the target objects are often partially occluded or completely hidden. Robots are thus required to search, identify a target object, and successfully grasp it. Preceding works have relied on pre-trained object recognition or segmentation models to find the target object. However, such methods require laborious manual annotations to train the models and even fail to find novel target objects. In this paper, we propose an Image-driven Object Searching and Grasping (IOSG) approach where a robot is provided with the reference image of a novel target object and tasked to find and retrieve it. We design a Target Similarity Network that generates a probability map to infer the location of the novel target. IOSG learns a hierarchical policy; the high-level policy predicts the subtask type, whereas the low-level policies, explorer and coordinator, generate effective push and grasp actions. The explorer is responsible for searching the target object when it is hidden or occluded by other objects. Once the target object is found, the coordinator conducts target-oriented pushing and grasping to retrieve the target from the clutter. The proposed pipeline is trained with full self-supervision in simulation and applied to a real environment. Our model achieves a 96.0% and 94.5% task success rate on coordination and exploration tasks in simulation respectively, and 85.0% success rate on a real robot for the search-and-grasp task. Please refer to our project page for more information: https://z.umn.edu/iosg. Houjian Yu, Xibai Lou, Yang Yang 0083, Changhyun Choi |
IROS | 4 |
| 2023 | Active Planar Mass Distribution Estimation with Robotic ManipulationabstractIn this work, we present a method to estimate the planar mass distribution of a rigid object through robotic interactions and force/torque feedback. This is a challenging problem because of the complexity of modeling physical dynamics and the action dependencies across the model parameters. We propose a sequential estimation strategy combined with a set of robot action selection rules based on the analytical formulation of a discrete-time dynamics model. To evaluate the performance of our approach, we also manufactured re-configurable block objects that allow us to modify the object mass distribution while having access to the ground truth values. We compare our approach against multiple baselines and show that it can estimate the mass distribution with around 10% error, while the baselines have errors ranging from 18% to 68%. Jiacheng Yuan, Changhyun Choi, Ellad B. Tadmor, Volkan Isler |
IROS | 2 |
| 2022 | Self-supervised Interactive Object Segmentation Through a Singulation-and-Grasping Approach
Houjian Yu, Changhyun Choi |
ECCV (39) | 2 |
| 2022 | Interactive Robotic Grasping with Attribute-Guided DisambiguationabstractInteractive robotic grasping using natural language is one of the most fundamental tasks in human-robot interaction. However, language can be a source of ambiguity, particularly when there are ambiguous visual or linguistic contents. This paper investigates the use of object attributes in disambiguation and develops an interactive grasping system capable of effectively resolving ambiguities via dialogues. Our approach first predicts target scores and attribute scores through vision-and-language grounding. To handle ambiguous objects and commands, we propose an attribute-guided formulation of the partially observable Markov decision process (Attr-POMDP) for disambiguation. The Attr-POMDP utilizes target and attribute scores as the observation model to calculate the expected return of an attribute-based (e.g., “what is the color of the target, red or green?”) or a pointing-based (e.g., “do you mean this one?”) question. Our disambiguation module runs in real time on a real robot, and the interactive grasping system achieves a 91.43 % selection accuracy in the real-robot experiments, outperforming several baselines by large margins. Supplementary material is available at https://sites.google.com/umn.eduJattr-disam. Yang Yang 0083, Xibai Lou, Changhyun Choi |
ICRA | 3 |
| 2022 | Learning Object Relations with Graph Neural Networks for Target-Driven Grasping in Dense ClutterabstractRobots in the real world frequently come across identical objects in dense clutter. When evaluating grasp poses in these scenarios, a target-driven grasping system requires knowledge of spatial relations between scene objects (e.g., proximity, adjacency, and occlusions). To efficiently complete this task, we propose a target-driven grasping system that simultaneously considers object relations and predicts 6-DoF grasp poses. A densely cluttered scene is first formulated as a grasp graph with nodes representing object geometries in the grasp coordinate frame and edges indicating spatial relations between the objects. We design a Grasp Graph Neural Network (G2N2) that evaluates the grasp graph and finds the most feasible 6-DoF grasp pose for a target object. Additionally, we develop a shape completion-assisted grasp pose sampling method that improves sample quality and consequently grasping efficiency. We compare our method against several baselines in both simulated and real settings. In real-world experiments with novel objects, our approach achieves a 77.78% grasping accuracy in densely cluttered scenarios, surpassing the best-performing baseline by more than 15%. Supplementary material is available at https://sites.google.com/umn.edu/graph-grasping. Xibai Lou, Yang Yang 0083, Changhyun Choi |
ICRA | 3 |
| 2022 | Fusion of Tandem-X and Gedi Data for Mapping Forest Height in the Brazilian AmazonabstractThe combination of TanDEM-X interferometric measurements with GEDI lidar full waveform measurements can provide continuous high-resolution forest height maps at global scale with sufficient accuracy without using external information about the underlyingtopography. In previous studies, the GEDI lidar full waveforms have been used to provide an approximation of the TanDEM-X X-band (radar) vertical reflectivity function in the height inversion of an entire TanDEM-X scene. This framework has been applied to the whole the whole Brazilian Amazon, and the obtained results are presented and analyzed in this paper. More than 12,000 TanDEM-X scenes and 250 millions GEDI lidar measurements have been processed.have. Changhyun Choi, Matteo Pardini, Roman Guliaev, Konstantinos Papathanassiou |
IGARSS | 1 |
| 2022 | Forest Parameter Estimation by Means of Multi-Baseline Pol-Insar Techniques: State-of-the-Art and Future ChallengesabstractPolarimetric SAR Interferometry (Pol-InSAR) is a SAR remote sensing discipline with unique and powerful applications related to the vertical structure of natural and man-made volume scatterers. The coherent combination of single- or multi -baseline interferograms acquired at different polarisations provides sensitivity to the vertical distribution of scattering processes and allows their characterisation by using the associated (volume) interferometric coherences. Konstantinos Papathanassiou, Roman Guliaev, Changhyun Choi, Lea Albrecht, Noelia Romero-Puig, Alberto Alonso-González, Jun Su Kim, Matteo Pardini |
IGARSS | 3 |
| 2022 | Fixture-Aware DDQN for Generalized Environment-Enabled GraspingabstractThis paper expands on the problem of grasping an object that can only be grasped by a single parallel gripper when a fixture (e.g., wall, heavy object) is harnessed. Preceding work that tackle this problem are limited in that the employed networks implicitly learn specific targets and fixtures to leverage. However, the notion of a usable fixture can vary in different environments, at times without any outwardly noticeable differences. In this paper, we propose a method to relax this limitation and further handle environments where the fixture location is unknown. The problem is formulated as visual affordance learning in a partially observable setting. We present a self-supervised reinforcement learning algorithm, Fixture-Aware Double Deep Q-Network (FA-DDQN), that processes the scene observation to 1) identify the target object based on a reference image, 2) distinguish possible fixtures based on interaction with the environment, and finally 3) fuse the information to generate a visual affordance map to guide the robot to successful Slide-to-Wall grasps. We demonstrate our proposed solution in simulation and in real robot experiments to show that in addition to achieving higher success than baselines, it also performs zero-shot generalization to novel scenes with unseen object configurations. Eddie Sasagawa, Changhyun Choi |
IROS | 2 |
| 2021 | Attribute-Based Robotic Grasping with One-Grasp AdaptationabstractRobotic grasping is one of the most fundamental robotic manipulation tasks and has been actively studied. However, how to quickly teach a robot to grasp a novel target object in clutter remains challenging. This paper attempts to tackle the challenge by leveraging object attributes that facilitate recognition, grasping, and quick adaptation. In this work, we introduce an end-to-end learning method of attribute-based robotic grasping with one-grasp adaptation capability. Our approach fuses the embeddings of a workspace image and a query text using a gated-attention mechanism and learns to predict instance grasping affordances. Besides, we utilize object persistence before and after grasping to learn a joint metric space of visual and textual attributes. Our model is self-supervised in a simulation that only uses basic objects of various colors and shapes but generalizes to novel objects and real-world scenes. We further demonstrate that our model is capable of adapting to novel objects with only one grasp data and improving instance grasping performance significantly. Experimental results in both simulation and the real world demonstrate that our approach achieves over 80% instance grasping success rate on unknown objects, which outperforms several baselines by large margins. Supplementary material is available at https://sites.google.com/umn.edu/attributes-grasping. Yang Yang 0083, Yuanhao Liu 0003, Hengyue Liang, Xibai Lou, Changhyun Choi |
ICRA | 5 |
| 2021 | Learning Visual Affordances with Target-Orientated Deep Q-Network to Grasp Objects by Harnessing Environmental FixturesabstractThis paper introduces a challenging object grasping task and proposes a self-supervised learning approach. The goal of the task is to grasp an object which is not feasible with a single parallel gripper, but only with harnessing environment fixtures (e.g., walls, furniture, heavy objects). This Slide-to-Wall grasping task assumes no prior knowledge except the partial observation of a target object. Hence the robot should learn an effective policy given a scene observation that may include the target object, environmental fixtures, and any other disturbing objects. We formulate the problem as visual affordances learning for which Target-Oriented Deep Q-Network (TO-DQN) is proposed to efficiently learn visual affordance maps (i.e., Q-maps) to guide robot actions. Since the training necessitates robot's exploration and collision with the fixtures, TO-DQN is first trained safely with a simulated robot manipulator and then applied to a real robot. We empirically show that TO-DQN can learn to solve the task in different environment settings in simulation and outperforms a standard and a variant of Deep Q-Network (DQN) in terms of training efficiency and robustness. The testing performance in both simulation and real-robot experiments shows that the policy trained by TO-DQN achieves comparable performance to humans. Hengyue Liang, Xibai Lou, Yang Yang 0083, Changhyun Choi |
ICRA | 4 |
| 2021 | Collision-Aware Target-Driven Object Grasping in Constrained EnvironmentsabstractGrasping a novel target object in constrained environments (e.g., walls, bins, and shelves) requires intensive reasoning about grasp pose reachability to avoid collisions with the surrounding structures. Typical 6-DoF robotic grasping systems rely on the prior knowledge about the environment and intensive planning computation, which is ungeneralizable and inefficient. In contrast, we propose a novel Collision-Aware Reachability Predictor (CARP) for 6-DoF grasping systems. The CARP learns to estimate the collision-free probabilities for grasp poses and significantly improves grasping in challenging environments. The deep neural networks in our approach are trained fully by self-supervision in simulation. The experiments in both simulation and the real world show that our approach achieves more than 75% grasping rate on novel objects in various surrounding structures. The ablation study demonstrates the effectiveness of the CARP, which improves the 6-DoF grasping rate by 95.7%. Xibai Lou, Yang Yang 0083, Changhyun Choi |
ICRA | 3 |
| 2021 | Tandem-X and Gedi Data Fusion for a Continuous Forest Height Mapping at Large ScalesabstractThe TerraSAR-X add on for Digital Elevation Measurement (TanDEM-X) mission provides Interferometric Synthetic Aperture Radar (InSAR) wall-to-wall data (not sparse) at high resolution and at global scale. In addition, the NASA Global Ecosystem Dynamics Investigation (GEDI) is a new spaceborne system that provides (from 51.6°N and 51.6°S) sparse measurements (not images) through LiDAR waveforms. Both systems are sensitivity to the canopy structure such as the forest height but with their own limitations. The TanDEM-X single polarization (HH) interferometric coherence magnitude at X-band provides a continuous mapping of the forest while GEDI provides accurate (but sparse) measurements of the forest. In this paper a methodology of how to combine both systems to estimated forest height is presented and applied to more than 900 TanDEM-X scenes over Gabon in Africa. The forest height results over an area of 1° by 1° are shown and compared respect to GEDI. Finally, a wall-to-wall forest map over the entire country of Gabon is presented as an example of large scale mapping towards a potential global (entire earth) forest height map. Victor Cazcarra-Bes, Matteo Pardini, Changhyun Choi, Roman Guliaev, Konstantinos Papathanassiou |
IGARSS | 3 |
| 2021 | Optimal InSAR Conditions for Monitoring Creek Changes in Tidal FlatsabstractMonitoring of tidal creeks is important for safety in tidal flats and for understanding coastal changes due to sedimentation and erosion. The tidal creek is a small stream developed on tidal flats and requires very precise Digital Elevation Model (DEM) to detect it. Interferometric SAR (InSAR) technique was usually applied for generating DEM over inland regions. In tidal flat regions, low backscattering due to flatness and remnant water of the tidal flat prevents from generating high precision DEM. In this study, optimum InSAR baseline condition was investigated to generate such a precise DEM in the tidal flat. The long-baseline TanDEM-X and airborne InSAR system were found to be useful, and the tidal creek information (width and depth) was successfully extracted from the DEM that was generated from optimal baseline InSAR systems. Duk-jin Kim, Changhyun Choi, Jungkyo Jung, Ji-Hwan Hwang |
IGARSS | 2 |
| 2020 | Helping Robots Learn: A Human-Robot Master-Apprentice Model Using Demonstrations via Virtual Reality TeleoperationabstractAs artificial intelligence becomes an increasingly prevalent method of enhancing robotic capabilities, it is important to consider effective ways to train these learning pipelines and to leverage human expertise. Working towards these goals, a master-apprentice model is presented and is evaluated during a grasping task for effectiveness and human perception. The apprenticeship model augments self-supervised learning with learning by demonstration, efficiently using the human's time and expertise while facilitating future scalability to supervision of multiple robots; the human provides demonstrations via virtual reality when the robot cannot complete the task autonomously. Experimental results indicate that the robot learns a grasping task with the apprenticeship model faster than with a solely self-supervised approach and with fewer human interventions than a solely demonstration-based approach; 100% grasping success is obtained after 150 grasps with 19 demonstrations. Preliminary user studies evaluating workload, usability, and effectiveness of the system yield promising results for system scalability and deployability. They also suggest a tendency for users to overestimate the robot's skill and to generalize its capabilities, especially as learning improves. Joseph DelPreto, Jeffrey Lipton, Lindsay Sanneman, Aidan J. Fay, Christopher K. Fourie, Changhyun Choi, Daniela Rus |
ICRA | 6 |
| 2020 | Learning to Generate 6-DoF Grasp Poses with Reachability AwarenessabstractMotivated by the stringent requirements of unstructured real-world where a plethora of unknown objects reside in arbitrary locations of the surface, we propose a voxel-based deep 3D Convolutional Neural Network (3D CNN) that generates feasible 6-DoF grasp poses in unrestricted workspace with reachability awareness. Unlike the majority of works that predict if a proposed grasp pose within the restricted workspace will be successful solely based on grasp pose stability, our approach further learns a reachability predictor that evaluates if the grasp pose is reachable or not from robot's own experience. To avoid the laborious real training data collection, we exploit the power of simulation to train our networks on a large-scale synthetic dataset. This work is an early attempt that simultaneously learns grasping reachability while proposing feasible grasp poses with 3D CNN. Experimental results in both simulation and real-world demonstrate that our approach outperforms several other methods and achieves 82.5% grasping success rate on unknown objects. Xibai Lou, Yang Yang 0083, Changhyun Choi |
ICRA | 3 |
| 2020 | Forest Height Estimation from Tandem-X InSAR Coherence Magnitude Towards Large Scale ApplicationsabstractTanDEM-X experiments have shown that forest height can be estimated with single polarization X-band interferometric coherences. An external digital terrain model (DTM) not only allows to use both coherence magnitude and phase information, but also to overcome X-band penetration limitations. However, DTM information is not available for large areas. Using coherence magnitudes makes height inversion feasible, but it requires a model relating coherence to height. Here we report an experiment using the X-band local phase center variations. Results over a tropical forest site show that in those stands in which the low X-band penetration is not a limitation, the there is a good correlation between the obtained TanDEM-X heights and the heights from Lidar measurements. Changhyun Choi, Roman Guliaev, Victor Cazcarra-Bes, Matteo Pardini, Konstantinos Papathanassiou |
IGARSS | 1 |
| 2019 | A Structure-Based Framework for the Combination of GEDI and Tandem-X Measurements Over Forest ScenariosabstractNASA's Global Ecosystem Dynamics Investigation (GEDI) waveform lidar is expected to provide unprecedented measurements of forest structure and biomass in tropical and temperate environments. In order to bridge the limitations induced by the ground sampling of the GEDI waveforms, and to obtain enhanced forest structure estimates, the potential of combining TanDEM-X (high resolution) singlepass interferometric coherences and lidar waveforms is currently investigated. In this work, a combination framework based on the ability of lidar and TanDEM-X measurements to express physical forest structure by means of appropriate indices is discussed. In particular, commonalities and complementarities between the different measurements are addressed by means of experimental results obtained in temperate and tropical forest sites in which comparisons among structure indices from lidar, T anDEM-X and field inventories can be established. Changhyun Choi, Matteo Pardini, Konstantinos Papathanassiou |
IGARSS | 1 |
| 2018 | Task-Specific Sensor Planning for Robotic Assembly TasksabstractWhen performing multi-robot tasks, sensory feedback is crucial in reducing uncertainty for correct execution. Yet the utilization of sensors should be planned as an integral part of the task planning, taken into account several factors such as the tolerance of different inferred properties of the scene and interaction with different agents. In this paper we handle this complex problem in a principled, yet efficient way. We use surrogate predictors based on open-loop simulation to estimate and bound the probability of success for specific tasks. We reason about such task-specific uncertainty approximants and their effectiveness. We show how they can be incorporated into a multi-robot planner, and demonstrate results with a team of robots performing assembly tasks. Guy Rosman, Changhyun Choi, Mehmet Remzi Dogar, John W. Fisher III, Daniela Rus |
ICRA | 2 |
| 2018 | Quantification of Horizontal Forest Structure from High Resolution Tandem-X Interferometric CoherencesabstractRecent TanDEM-X experiments have shown that the limited penetration capability at X-band in forest volumes allow the estimation of the height variability of the top canopy layer, which can be used as a proxy to the horizontal structure (i.e. heterogeneity), by using high resolution digital elevation models (DEMs). However, the use of an external digital terrain model (DTM) is necessary to separate the (high resolution) canopy height variations from the topographic ones. In this work, the possibility of compensating terrain topographic variation by using a low resolution TanDEM-X DEM instead of an external DTM is investigated. The results show that the use of a reference DEM with a resolution on the order of 100 m allows to compensate the terrain-induced topographic variations and to preserve the information on forest horizontal heterogeneity at a large extent. Changhyun Choi, Matteo Pardini, Konstantinos Papathanassiou |
IGARSS | 1 |
| 2017 | Duckietown: An open, inexpensive and flexible platform for autonomy education and researchabstractDuckietown is an open, inexpensive and flexible platform for autonomy education and research. The platform comprises small autonomous vehicles (“Duckiebots”) built from off-the-shelf components, and cities (“Duckietowns”) complete with roads, signage, traffic lights, obstacles, and citizens (duckies) in need of transportation. The Duckietown platform offers a wide range of functionalities at a low cost. Duckiebots sense the world with only one monocular camera and perform all processing onboard with a Raspberry Pi 2, yet are able to: follow lanes while avoiding obstacles, pedestrians (duckies) and other Duckiebots, localize within a global map, navigate a city, and coordinate with other Duckiebots to avoid collisions. Duckietown is a useful tool since educators and researchers can save money and time by not having to develop all of the necessary supporting infrastructure and capabilities. All materials are available as open source, and the hope is that others in the community will adopt the platform for education and research. Liam Paull, Jacopo Tani, Heejin Ahn, Javier Alonso-Mora, Luca Carlone, Michal Cáp, Yu Fan Chen, Changhyun Choi, Jeff Dusek, Yajun Fang, Daniel Hoehener, Shih-Yuan Liu, Michael Novitzky, Igor Franzoni Okuyama, Jason Pazis, Guy Rosman, Valerio Varricchio, Hsueh-Cheng Wang, Dmitry S. Yershov, Hang Zhao 0021, Michael Benjamin, Christopher Carr, Maria T. Zuber, Sertac Karaman, Emilio Frazzoli, Domitilla Del Vecchio, Daniela Rus, Jonathan P. How, John J. Leonard, Andrea Censi |
ICRA | 8 |
| 2017 | Intertidal flat topographies measured by long-baseline airborne sar and tandem-xabstractIntertidal flats which are located between land and ocean are productive and rapidly changing places. But these places are now facing many environmental challenges related to climate change and human-induced impacts. Topographic change related to sedimentation or erosion in intertidal flats is the most evident sign of the environmental changes. The intertidal flats usually have small topographic variations (less than 5m) and experience ebb and flood tides every day. Thus, the conventional SAR interferometric techniques (repeat-pass InSAR) cannot be applied for generating DEMs in these intertidal flats. In this study, we developed a long-baseline single-pass airborne interferometric SAR (InSAR) system that has small ambiguity height and collected airborne InSAR data in several intertidal flats, the west coast of Korean peninsula. The constructed topographies using the long-baseline airborne InSAR system were compared with TanDEM-X DEM obtained during long-baseline mission phase. Duk-jin Kim, Changhyun Choi, Jungkyo Jung, Ki-mook Kang, Seung Hee Kim, Ji-Hwan Hwang |
IGARSS | 2 |
| 2016 | Probabilistic visual verification for robotic assembly manipulationabstractIn this paper we present a visual verification approach for robotic assembly manipulation which enables robots to verify their assembly state. Given shape models of objects and their expected placement configurations, our approach estimates the probability of the success of the assembled state using a depth sensor. The proposed approach takes into account uncertainties in object pose. Probability distributions of depth and surface normal depending on the uncertainties are estimated to classify the assembly state in a Bayesian formulation. The effectiveness of our approach is validated in comparative experiments with other approaches. Changhyun Choi, Daniela Rus |
ICRA | 1 |
| 2016 | Evaluation of long-baseline TanDEM-X DEM in tidal flatabstractTanDEM-X (TerraSAR-X add-on for Digital Elevation Measurement), a German SAR satellite, was launched in June 2010. Since 2012, TanDEM-X mission has provided the most detailed global DEM with unprecedented accuracy. Also, they carried out various experimental mode including long baseline XTI (cross track interferometry). Despite their progressive instrument calibration and high accurate orbit determination, the error of topographic height was still remained in the level of global DEM and was not entirely estimated. Particularly, the DEM accuracy in coastal area including tidal flats has never been reported. In this study, we evaluated the accuracy of TanDEM-X DEM in coastal area by comparing with RTK GPS measurements, and we investigated the optimal baseline condition to generate significant DEM in tidal flat. Changhyun Choi, Duk-jin Kim |
IGARSS | 1 |
| 2016 | Measurements of intertidal flat topography using a long-baseline airborne interferometric SARabstractIntertidal flats are productive and rapidly changing places. However, these places are facing many environmental challenges related to climate change and human-induced impacts. Topographic change due to sedimentation or erosion in intertidal flats can be the key indicator for recognizing these environmental changes. The intertidal flats usually have small topographic variations (less than 5m) and high-moistured soil surfaces (greater than 50%). The conventional SAR interferometric techniques cannot be used for generating DEMs in these intertidal flats. In this study, we developed a long-baseline airborne interferometric SAR (InSAR) system that has small ambiguity height and collected airborne InSAR data in Jebu intertidal flat, west coast of Korean peninsula. The constructed topographies using the long-baseline airborne InSAR were compared with TanDEM-X DEM (acquired during long-baseline mission phase) and GPS-RTK measurements. Duk-jin Kim, Changhyun Choi, Jungkyo Jung, Ki-mook Kang, Seung Hee Kim, Ji-Hwan Hwang |
IGARSS | 2 |
| 2013 | RGB-D object tracking: A particle filter approach on GPUabstractThis paper presents a particle filtering approach for 6-DOF object pose tracking using an RGB-D camera. Our particle filter is massively parallelized in a modern GPU so that it exhibits real-time performance even with several thousand particles. Given an a priori 3D mesh model, the proposed approach renders the object model onto texture buffers in the GPU, and the rendered results are directly used by our parallelized likelihood evaluation. Both photometric (colors) and geometric (3D points and surface normals) features are employed to determine the likelihood of each particle with respect to a given RGB-D scene. Our approach is compared with a tracker in the PCL both quantitatively and qualitatively in synthetic and real RGB-D sequences, respectively. Changhyun Choi, Henrik I. Christensen |
IROS | 1 |
| 2013 | RGB-D edge detection and edge-based registrationabstractWe present a 3D edge detection approach for RGB-D point clouds and its application in point cloud registration. Our approach detects several types of edges, and makes use of both 3D shape information and photometric texture information. Edges are categorized as occluding edges, occluded edges, boundary edges, high-curvature edges, and RGB edges. We exploit the organized structure of the RGB-D image to efficiently detect edges, enabling near real-time performance. We present two applications of these edge features: edge-based pair-wise registration and a pose-graph SLAM approach based on this registration, which we compare to state-of-the-art methods. Experimental results demonstrate the performance of edge detection and edge-based registration both quantitatively and qualitatively. Changhyun Choi, Alexander J. B. Trevor, Henrik I. Christensen |
IROS | 1 |
| 2012 | Voting-based pose estimation for robotic assembly using a 3D sensorabstractWe propose a voting-based pose estimation algorithm applicable to 3D sensors, which are fast replacing their 2D counterparts in many robotics, computer vision, and gaming applications. It was recently shown that a pair of oriented 3D points, which are points on the object surface with normals, in a voting framework enables fast and robust pose estimation. Although oriented surface points are discriminative for objects with sufficient curvature changes, they are not compact and discriminative enough for many industrial and real-world objects that are mostly planar. As edges play the key role in 2D registration, depth discontinuities are crucial in 3D. In this paper, we investigate and develop a family of pose estimation algorithms that better exploit this boundary information. In addition to oriented surface points, we use two other primitives: boundary points with directions and boundary line segments. Our experiments show that these carefully chosen primitives encode more information compactly and thereby provide higher accuracy for a wide class of industrial parts and enable faster computation. We demonstrate a practical robotic bin-picking system using the proposed algorithm and a 3D sensor. Changhyun Choi, Yuichi Taguchi, Oncel Tuzel, Ming-Yu Liu 0001, Srikumar Ramalingam |
ICRA | 1 |
| 2012 | 3D pose estimation of daily objects using an RGB-D cameraabstractIn this paper, we present an object pose estimation algorithm exploiting both depth and color information. While many approaches assume that a target region is cleanly segmented from background, our approach does not rely on that assumption, and thus it can estimate pose of a target object in heavy clutter. Recently, an oriented point pair feature was introduced as a low dimensional description of object surfaces. The feature has been employed in a voting scheme to find a set of possible 3D rigid transformations between object model and test scene features. While several approaches using the pair features require an accurate 3D CAD model as training data, our approach only relies on several scanned views of a target object, and hence it is straightforward to learn new objects. In addition, we argue that exploiting color information significantly enhances the performance of the voting process in terms of both time and accuracy. To exploit the color information, we define a color point pair feature, which is employed in a voting scheme for more effective pose estimation. We show extensive quantitative results of comparative experiments between our approach and a state-of-the-art. Changhyun Choi, Henrik I. Christensen |
IROS | 1 |
| 2012 | 3D textureless object detection and tracking: An edge-based approachabstractThis paper presents an approach to textureless object detection and tracking of the 3D pose. Our detection and tracking schemes are coherently integrated in a particle filtering framework on the special Euclidean group, SE(3), in which the visual tracking problem is tackled by maintaining multiple hypotheses of the object pose. For textureless object detection, an efficient chamfer matching is employed so that a set of coarse pose hypotheses is estimated from the matching between 2D edge templates of an object and a query image. Particles are then initialized from the coarse pose hypotheses by randomly drawing based on costs of the matching. To ensure the initialized particles are at or close to the global optimum, an annealing process is performed after the initialization. While a standard edge-based tracking is employed after the annealed initialization, we employ a refinement process to establish improved correspondences between projected edge points from the object model and edge points from an input image. Comparative results for several image sequences with clutter are shown to validate the effectiveness of our approach. Changhyun Choi, Henrik I. Christensen |
IROS | 1 |
| 2011 | Please smileabstractNowadays, with reductions in manufacturing costs and a transition toward lifestyles of convenience, robots are becoming pervasive in our homes, museums, and hospitals. In addition to increased demands for robots in these domains, recently more artistic robots that interact with audiences on a personal instead of a practical level are now being exhibited in art exhibition. This paper explains how people interpret artistic robots as more than mere machines in the theory of intentionality and introduces the implementation of the artistic robot, Please Smile, which consists of five robotic skeleton arms that gesture in response to a viewer's facial expressions. The paper also explores how individuals can use experimental designs to create artistic robots that can express various ideas that traditional, practical robots can often not convey. Hye Yeon Nam, Changhyun Choi, Sam Mendenhall |
Creativity & Cognition | 2 |
| 2011 | Robust 3D visual tracking using particle filtering on the SE(3) groupabstractIn this paper, we present a 3D model-based object tracking approach using edge and keypoint features in a particle filtering framework. Edge points provide 1D information for pose estimation and it is natural to consider multiple hypotheses. Recently, particle filtering based approaches have been proposed to integrate multiple hypotheses and have shown good performance, but most of the work has made an assumption that an initial pose is given. To remove this assumption, we employ keypoint features for initialization of the filter. Given 2D-3D keypoint correspondences, we choose a set of minimum correspondences to calculate a set of possible pose hypotheses. Based on the inlier ratio of correspondences, the set of poses are drawn to initialize particles. For better performance, we employ an autoregressive state dynamics and apply it to a coordinate-invariant particle filter on the SE(3) group. Based on the number of effective particles calculated during tracking, the proposed system re-initializes particles when the tracked object goes out of sight or is occluded. The robustness and accuracy of our approach is demonstrated via comparative experiments. Changhyun Choi, Henrik I. Christensen |
ICRA | 1 |
| 2010 | Real-time 3D model-based tracking using edge and keypoint features for robotic manipulationabstractWe propose a combined approach for 3D real-time object recognition and tracking, which is directly applicable to robotic manipulation. We use keypoints features for the initial pose estimation. This pose estimate serves as an initial estimate for edge-based tracking. The combination of these two complementary methods provides an efficient and robust tracking solution. The main contributions of this paper includes: 1) While most of the RAPiD style tracking methods have used simplified CAD models or at least manually well designed models, our system can handle any form of polygon mesh model. To achieve the generality of object shapes, salient edges are automatically identified during an offline stage. Dull edges usually invisible in images are maintained as well for the cases when they constitute the object boundaries. 2) Our system provides a fully automatic recognition and tracking solution, unlike most of the previous edge-based tracking that require a manual pose initialization scheme. Since the edge-based tracking sometimes drift because of edge ambiguity, the proposed system monitors the tracking results and occasionally re-initialize when the tracking results are inconsistent. Experimental results demonstrate our system's efficiency as well as robustness. Changhyun Choi, Henrik I. Christensen |
ICRA | 1 |
| 2009 | Cognitive vision for efficient scene processing and object categorization in highly cluttered environmentsabstractOne of the key competencies required in modern robots is finding objects in complex environments. For the last decade, significant progress in computer vision and machine learning literatures has increased the recognition performance of well localized objects. However, the performance of these techniques is still far from human performance, especially in cluttered environments. We believe that the performance gap between robots and humans is due in part to humans' use of an attention system. According to cognitive psychology, the human visual system uses two stages of visual processing to interpret visual input. The first stage is a pre-attentive process perceiving scenes fast and coarsely to select potentially interesting regions. The second stage is a more complex process analyzing the regions hypothesized in the previous stage. These two stages play an important role in enabling efficient use of the limited cognitive resources available. Inspired by this biological fact, we propose a visual attentional object categorization approach for robots that enables object recognition in real environments under a critical time limitation. We quantitatively evaluate the performance for recognition of objects in highly cluttered scenes without significant loss of detection rates across several experimental settings. Changhyun Choi, Henrik I. Christensen |
IROS | 1 |
| 2008 | Real-time 3D object pose estimation and tracking for natural landmark based visual servoabstractA real-time solution for estimating and tracking the 3D pose of a rigid object is presented for image-based visual servo with natural landmarks. The many state-of-the-art technologies that are available for recognizing the 3D pose of an object in a natural setting are not suitable for real-time servo due to their time lags. This paper demonstrates that a real-time solution of 3D pose estimation become feasible by combining a fast tracker such as KLT [7] [8] with a method of determining the 3D coordinates of tracking points on an object at the time of SIFT based tracking point initiation, assuming that a 3D geometric model with SIFT description of an object is known a-priori. Keeping track of tracking points with KLT, removing the tracking point outliers automatically, and reinitiating the tracking points using SIFT once deteriorated, the 3D pose of an object can be estimated and tracked in real-time. This method can be applied to both mono and stereo camera based 3D pose estimation and tracking. The former guarantees higher frame rates with about 1 ms of local pose estimation, while the latter assures of more precise pose results but with about 16 ms of local pose estimation. The experimental investigations have shown the effectiveness of the proposed approach with real-time performance. Changhyun Choi, Seungmin Baek, Sukhan Lee 0001 |
IROS | 1 |