Daehyung Park

dblp:153/7654 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-1287-9433ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 8 since 2021Systems, architecture and hardware · 12 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 LINGO-Space: Language-Conditioned Incremental Grounding for Space
abstract
We aim to solve the problem of spatially localizing composite instructions referring to space: space grounding. Compared to current instance grounding, space grounding is challenging due to the ill-posedness of identifying locations referred to by discrete expressions and the compositional ambiguity of referring expressions. Therefore, we propose a novel probabilistic space-grounding methodology (LINGO-Space) that accurately identifies a probabilistic distribution of space being referred to and incrementally updates it, given subsequent referring expressions leveraging configurable polar distributions. Our evaluations show that the estimation using polar distributions enables a robot to ground locations successfully through 20 table-top manipulation benchmark tests. We also show that updating the distribution helps the grounding method accurately narrow the referring space. We finally demonstrate the robustness of the space grounding with simulated manipulation and real quadruped robot navigation tasks. Code and videos are available at https://lingo-space.github.io.
Nayoung Oh, Deokmin Hwang, Daehyung Park
AAAI4
2024 Constrained Nonlinear Disturbance Observer for Robotic Systems
abstract
Disturbance observer (DOB) is a well-known two-loop control structure that imparts robustness to a controller with a simple implementation. As a nonlinear DOB for the robotic systems, we proposed so-called nonlinear robust internal-loop compensator (NRIC) framework in our previous work. In this paper, we further extend the NRIC in such a way that an optimization scheme can be embedded in the control structure. The proposed method is called constrained NRIC (C-NRIC), because the optimization allows us to impose constraints, by which a controller acquires additional properties. As a particular use case of the C-NRIC framework, we design contact-responsive motion controllers that enables a robot to react to unknown interactions while accurately tracking the desired trajectory in free motion. The effectiveness of such designs is validated through the real-world experiments.
Ji Wan Han, Daehyung Park, Minjun Kim 0003
ICRA2
2024 Graph-based 3D Collision-distance Estimation Network with Probabilistic Graph Rewiring
abstract
We aim to solve the problem of data-driven collision-distance estimation given 3-dimensional (3D) geometries. Conventional algorithms suffer from low accuracy due to their reliance on limited representations, such as point clouds. In contrast, our previous graph-based model, GraphDistNet, achieves high accuracy using edge information but incurs higher message-passing costs with growing graph size, limiting its applicability to 3D geometries. To overcome these challenges, we propose GDN-R, a novel 3D graph-based estimation network. GDN-R employs a layer-wise probabilistic graph-rewiring algorithm leveraging the differentiable Gumbel-top-K relaxation. Our method accurately infers minimum distances through iterative graph rewiring and updating relevant embeddings. The probabilistic rewiring enables fast and robust embedding with respect to unforeseen categories of geometries. Through 41, 412 random benchmark tasks with 150 pairs of 3D objects, we show GDN-R outperforms state-of-the-art baseline methods in terms of accuracy and generalizability. We also show that the proposed rewiring improves the update performance reducing the size of the estimation model. We finally show its batch prediction and auto-differentiation capabilities for trajectory optimization in both simulated and real-world scenarios.
Minjae Song, Yeseung Kim, Minjun Kim 0003, Daehyung Park
ICRA4
2023 A Reachability Tree-Based Algorithm for Robot Task and Motion Planning
abstract
This paper presents a novel algorithm for robot task and motion planning (TAMP) problems by utilizing a reachability tree. While tree-based algorithms are known for their speed and simplicity in motion planning (MP), they are not well-suited for TAMP problems that involve both abstracted and geometrical state variables. To address this challenge, we propose a hierarchical sampling strategy, which first generates an abstracted task plan using Monte Carlo tree search (MCTS) and then fills in the details with a geometrically feasible motion trajectory. Moreover, we show that the performance of the proposed method can be significantly enhanced by selecting an appropriate reward for MCTS and by using a pre-generated goal state that is guaranteed to be geometrically feasible. A comparative study using TAMP benchmark problems demonstrates the effectiveness of the proposed approach.
Kanghyun Kim, Daehyung Park, Minjun Kim 0003
ICRA2
2023 Learning-based Initialization of Trajectory Optimization for Path-following Problems of Redundant Manipulators
abstract
Trajectory optimization (TO) is an efficient tool to generate a redundant manipulator's joint trajectory following a 6-dimensional Cartesian path. The optimization performance largely depends on the quality of initial trajectories. However, the selection of a high-quality initial trajectory is non-trivial and requires a considerable time budget due to the extremely large space of the solution trajectories and the lack of prior knowledge about task constraints in configuration space. To alleviate the issue, we present a learning-based initial trajectory generation method that generates high-quality initial trajectories in a short time budget by adopting example-guided reinforcement learning. In addition, we suggest a null-space projected imitation reward to consider null-space constraints by efficiently learning kinematically feasible motion captured in expert demonstrations. Our statistical evaluation in simulation shows the improved optimality, efficiency, and applicability of TO when we plug in our method's output, compared with three other baselines. We also show the performance improvement and feasibility via real-world experiments with a seven-degree-of-freedom manipulator.
Minsung Yoon, Mincheul Kang, Daehyung Park, Sung-Eui Yoon
ICRA3
2023 SGGNet2: Speech-Scene Graph Grounding Network for Speech-guided Navigation
abstract
The spoken language serves as an accessible and efficient interface, enabling non-experts and disabled users to interact with complex assistant robots. However, accurately grounding language utterances gives a significant challenge due to the acoustic variability in speakers’ voices and environmental noise. In this work, we propose a novel speech-scene graph grounding network (SGGNet2) that robustly grounds spoken utterances by leveraging the acoustic similarity between correctly recognized and misrecognized words obtained from automatic speech recognition (ASR) systems. To incorporate the acoustic similarity, we extend our previous grounding model, the scene-graph-based grounding network (SGGNet), with the ASR model from NVIDIA NeMo. We accomplish this by feeding the latent vector of speech pronunciations into the BERT-based grounding network within SGGNet. We evaluate the effectiveness of using latent vectors of speech commands in grounding through qualitative and quantitative studies. We also demonstrate the capability of SGGNet2 in a speech-based navigation task using a real quadruped robot, RBQ-3, from Rainbow Robotics.
Yeseung Kim, Jaehwi Jang, Minjae Song, Woojin Choi, Daehyung Park
RO-MAN6
2022 Confidence-Based Robot Navigation Under Sensor Occlusion with Deep Reinforcement Learning
abstract
This paper considers the problem of prolonged occlusions on navigation sensors due to dust, smudges, soils, etc. Such uncontrollable occlusions often cause lower visibility as well as higher uncertainty that require considerably sophisticated behavior. To secure visibility (i.e., confidence about the world), we propose a confidence-based navigation method that encourages the robot to explore the uncertain region around the robot maximizing its local confidence. To effectively extract features from the variable size of sensor occlusions, we adopt a point-cloud based representation network. Our method returns a resilient navigation policy via deep reinforcement learning, autonomously avoiding collisions under sensor occlusions while reaching a goal. We evaluate our method in simulated and real-world environments with either static or dynamic obstacles under various sensor-occlusion scenarios. The experimental result shows that our method outperforms baseline methods under the highly occurring sensor occlusion, and achieves maximum 90% and 80% success rates in the tested static and dynamic environments, respectively.
Hyeongyeol Ryu, Minsung Yoon, Daehyung Park, Sung-Eui Yoon
ICRA3
2021 Reactive Task and Motion Planning under Temporal Logic Specifications
abstract
We present a task-and-motion planning (TAMP) algorithm robust against a human operator's cooperative or adversarial interventions. Interventions often invalidate the current plan and require replanning on the fly. Replanning can be computationally expensive and often interrupts seamless task execution. We introduce a dynamically reconfigurable planning methodology with behavior tree-based control strategies toward reactive TAMP, which takes the advantage of previous plans and incremental graph search during temporal logic-based reactive synthesis. Our algorithm also shows efficient recovery functionalities that minimize the number of replanning steps. Finally, our algorithm produces a robust, efficient, and complete TAMP solution. Our experimental results show the algorithm results in superior manipulation performance in both simulated and real-world tasks.
Shen Li 0003, Daehyung Park, Yoonchang Sung, Julie A. Shah, Nicholas Roy
ICRA2
2019 Leveraging Past References for Robust Language Grounding
abstract
Grounding referring expressions to objects in an environment has traditionally been considered a one-off, ahistorical task.However, in realistic applications of grounding, multiple users will repeatedly refer to the same set of objects.As a result, past referring expressions for objects can provide strong signals for grounding subsequent referring expressions.We therefore reframe the grounding problem from the perspective of coreference detection and propose a neural network that detects when two expressions are referring to the same object.The network combines information from vision and past referring expressions to resolve which object is being referred to.Our experiments show that detecting referring expression coreference is an effective way to ground objects described by subtle visual properties, which standard visual grounding models have difficulty capturing.We also show the ability to detect object coreference allows the grounding model to perform well even when it encounters object categories not seen in the training data.
Subhro Roy, Michael Noseworthy, Rohan Paul, Daehyung Park, Nicholas Roy
CoNLL4
2018 3D Human Pose Estimation on a Configurable Bed from a Pressure Image
abstract
Robots have the potential to assist people in bed, such as in healthcare settings, yet bedding materials like sheets and blankets can make observation of the human body difficult for robots. A pressure-sensing mat on a bed can provide pressure images that are relatively insensitive to bedding materials. However, prior work on estimating human pose from pressure images has been restricted to 2D pose estimates and flat beds. In this work, we present two convolutional neural networks to estimate the 3D joint positions of a person in a configurable bed from a single pressure image. The first network directly outputs 3D joint positions, while the second outputs a kinematic model that includes estimated joint angles and limb lengths. We evaluated our networks on data from 17 human participants with two bed configurations: supine and seated. Our networks achieved a mean joint position error of 77 mm when tested with data from people outside the training set, outperforming several baselines. We also present a simple mechanical model that provides insight into ambiguity associated with limbs raised off of the pressure mat, and demonstrate that Monte Carlo dropout can be used to estimate pose confidence in these situations. Finally, we provide a demonstration in which a mobile manipulator uses our network's estimated kinematic model to reach a location on a person's body in spite of the person being seated in a bed and covered by a blanket.
Henry M. Clever, Ariel Kapusta, Daehyung Park, Zackory Erickson, Yash Chitalia, Charles C. Kemp
IROS3
2017 A multimodal execution monitor with anomaly classification for robot-assisted feeding
abstract
Activities of daily living (ADLs) are important for quality of life. Robotic assistance offers the opportunity for people with disabilities to perform ADLs on their own. However, when a complex semi-autonomous system provides real-world assistance, occasional anomalies are likely to occur. Robots that can detect, classify and respond appropriately to common anomalies have the potential to provide more effective and safer assistance. We introduce a multimodal execution monitor to detect and classify anomalous executions when robots operate near humans. Our system builds on our past work on multimodal anomaly detection. Our new monitor classifies the type and cause of common anomalies using an artificial neural network. We implemented and evaluated our execution monitor in the context of robot-assisted feeding with a general-purpose mobile manipulator. In our evaluations, our monitor outperformed baseline methods from the literature. It succeeded in detecting 12 common anomalies from 8 able-bodied participants with 83% accuracy and classifying the types and causes of the detected anomalies with 90% and 81% accuracies, respectively. We then performed an in-home evaluation with Henry Evans, a person with severe quadriplegia. With our system, Henry successfully fed himself while the monitor detected, classified the types, and classified the causes of anomalies with 86%, 90%, and 54% accuracy, respectively.
Daehyung Park, Hokeun Kim, Yuuna Hoshi, Zackory Erickson, Ariel Kapusta, Charles C. Kemp
IROS1
2016 Multimodal execution monitoring for anomaly detection during robot manipulation
abstract
Online detection of anomalous execution can be valuable for robot manipulation, enabling robots to operate more safely, determine when a behavior is inappropriate, and otherwise exhibit more common sense. By using multiple complementary sensory modalities, robots could potentially detect a wider variety of anomalies, such as anomalous contact or a loud utterance by a human. However, task variability and the potential for false positives make online anomaly detection challenging, especially for long-duration manipulation behaviors. In this paper, we provide evidence for the value of multimodal execution monitoring and the use of a detection threshold that varies based on the progress of execution. Using a data-driven approach, we train an execution monitor that runs in parallel to a manipulation behavior. Like previous methods for anomaly detection, our method trains a hidden Markov model (HMM) using multimodal observations from non-anomalous executions. In contrast to prior work, our system also uses a detection threshold that changes based on the execution progress. We evaluated our approach with haptic, visual, auditory, and kinematic sensing during a variety of manipulation tasks performed by a PR2 robot. The tasks included pushing doors closed, operating switches, and assisting able-bodied participants with eating yogurt. In our evaluations, our anomaly detection method performed substantially better with multimodal monitoring than single modality monitoring. It also resulted in more desirable ROC curves when compared with other detection threshold methods from the literature, obtaining higher true positive rates for comparable false positive rates.
Daehyung Park, Zackory Erickson, Tapomayukh Bhattacharjee, Charles C. Kemp
ICRA1
2015 Combining tactile sensing and vision for rapid haptic mapping
abstract
We consider the problem of enabling a robot to efficiently obtain a dense haptic map of its visible surroundings using the complementary properties of vision and tactile sensing. Our approach assumes that visible surfaces that look similar to one another are likely to have similar haptic properties. We present an iterative algorithm that enables a robot to infer dense haptic labels across visible surfaces when given a color-plus-depth (RGB-D) image along with a sequence of sparse haptic labels representative of what could be obtained via tactile sensing. Our method uses a color-based similarity measure and connected components on color and depth data. We evaluated our method using several publicly available RGBD image datasets with indoor cluttered scenes pertinent to robot manipulation. We analyzed the effects of algorithm parameters and environment variation, specifically the level of clutter and the type of setting, like a shelf, table top, or sink area. In these trials, the visible surface for each object consisted of an average of 8602 pixels, and we provided the algorithm with a sequence of haptically-labeled pixels up to a maximum of 40 times the number of objects in the image. On average, our algorithm correctly assigned haptic labels to 76.02% of all of the object pixels in the image given this full sequence of labels. We also performed experiments with the humanoid robot DARCI reaching in a cluttered foliage environment while using our algorithm to create a haptic map. Doing so enabled the robot to reach goal locations using a single plan after a single greedy reach, while our previous tactile-only mapping method required 5 or more plans to reach each goal.
Tapomayukh Bhattacharjee, Ashwin A. Shenoi, Daehyung Park, James M. Rehg, Charles C. Kemp
IROS3
2015 Task-centric selection of robot and environment initial configurations for assistive tasks
abstract
When a mobile manipulator functions as an assistive device, the robot's initial configuration and the configuration of the environment can impact the robot's ability to provide effective assistance. Selecting initial configurations for assistive tasks can be challenging due to the high number of degrees of freedom of the robot, the environment, and the person, as well as the complexity of the task. In addition, rapid selection of initial conditions can be important, so that the system will be responsive to the user and will not require the user to wait a long time while the robot makes a decision. To address these challenges, we present Task-centric initial Configuration Selection (TCS), which unlike previous work uses a measure of task-centric manipulability to accommodate state estimation error, considers various environmental degrees of freedom, and can find a set of configurations from which a robot can perform a task. TCS performs substantial offline computation, so that it can rapidly provide solutions at run time. At run time, the system performs an optimization over candidate initial configurations using a utility function that can include factors such as movement costs for the robot's mobile base. To evaluate TCS, we created models of 11 activities of daily living (ADLs) and evaluated TCS's performance with these 11 assistive tasks in a computer simulation of a PR2, a robotic bed, and a model of a human body. TCS performed as well or better than a baseline algorithm in all of our tests against state estimation error.
Ariel Kapusta, Daehyung Park, Charles C. Kemp
IROS2
2014 Learning to reach into the unknown: Selecting initial conditions when reaching in clutter
abstract
Often in highly-cluttered environments, a robot can observe the exterior of the environment with ease, but cannot directly view nor easily infer its detailed internal structure (e.g., dense foliage or a full refrigerator shelf). We present a data-driven approach that greatly improves a robot's success at reaching to a goal location in the unknown interior of an environment based on observable external properties, such as the category of the clutter and the locations of openings into the clutter (i.e., apertures). We focus on the problem of selecting a good initial configuration for a manipulator when reaching with a greedy controller. We use density estimation to model the probability of a successful reach given an initial condition and then perform constrained optimization to find an initial condition with the highest estimated probability of success. We evaluate our approach with two simulated robots reaching in clutter, and provide a demonstration with a real PR2 robot reaching to locations through random apertures. In our evaluations, our approach significantly outperformed two alternative approaches when making two consecutive reach attempts to goals in distinct categories of unknown clutter. Our approach only uses sparse readily-apparent features.
Daehyung Park, Ariel Kapusta, You Keun Kim, James M. Rehg, Charles C. Kemp
IROS1