VLDB 2026 Research / reviewers in the wild / expert
Tucker Hermans
dblp:67/4241
· DBLP profile ↗
32ranked-venue papers
2as first author
17since 2021 · last 2025
0000-0003-2496-2768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 2 first-author · 13 since 2021Systems, architecture and hardware · 23 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Points2Plans: From Point Clouds to Long-Horizon Plans with Composable Relational DynamicsabstractWe present Points2Plans, a framework for composable planning with a relational dynamics model that enables robots to solve long-horizon manipulation tasks from partial-view point clouds. Given a language instruction and a point cloud of the scene, our framework initiates a hierarchical planning procedure, whereby a language model generates a high-level plan and a sampling-based planner produces constraint-satisfying continuous parameters for manipulation primitives sequenced according to the high-level plan. Key to our approach is the use of a relational dynamics model as a unifying interface between the continuous and symbolic representations of states and actions, thus facilitating language-driven planning from high-dimensional perceptual input such as point clouds. Whereas previous relational dynamics models require training on datasets of multi-step manipulation scenarios that align with the intended test scenarios, Points2Plans uses only single-step simulated training data while generalizing zero-shot to a variable number of steps during real-world evaluations. We evaluate our approach on tasks involving geometric reasoning, multi-object interactions, and occluded object reasoning in both simulated and real-world settings. Results demonstrate that Points2Plans offers strong generalization to unseen long-horizon tasks in the real world, where it solves over 85% of evaluated tasks while the next best baseline solves only 50%. Christopher Agia, Jimmy Wu, Tucker Hermans, Jeannette Bohg |
ICRA | 4 |
| 2024 | Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory ModelsabstractRobots need to have a memory of previously observed, but currently occluded objects to work reliably in realistic environments. We investigate the problem of encoding object-oriented memory into a multi-object manipulation reasoning and planning framework. We propose DOOM and LOOM, which leverage transformer relational dynamics to encode the history of trajectories given partial-view point clouds and an object discovery and tracking engine. Our approaches can perform multiple challenging tasks including reasoning with occluded objects, novel objects appearance, and object reappearance. Throughout our extensive simulation and real-world experiments, we find that our approaches perform well in terms of different numbers of objects and different numbers of distractor actions. Furthermore, we show our approaches outperform an implicit memory baseline. Jialin Yuan, Chanho Kim, Pupul Pradhan, Bryan Chen, Fuxin Li, Tucker Hermans |
ICRA | 7 |
| 2024 | Point Cloud Models Improve Visual Robustness in Robotic LearnersabstractVisual control policies can encounter significant performance degradation when visual conditions like lighting or camera position differ from those seen during training – often exhibiting sharp declines in capability even for minor differences. In this work, we examine robustness to a suite of these types of visual changes for RGB-D and point cloud based visual control policies. To perform these experiments on both model-free and model-based reinforcement learners, we introduce a novel Point Cloud World Model (PCWM) and point cloud based control policies. Our experiments show that policies that explicitly encode point clouds are significantly more robust than their RGB-D counterparts. Further, we find our proposed PCWM significantly outperforms prior works in terms of sample efficiency during training. Taken together, these results suggest reasoning about the 3D scene through point clouds can improve performance, reduce learning time, and increase robustness for robotic learners. Code: https://github.com/pvskand/pcwm Skand Peri, Iain Lee, Chanho Kim, Fuxin Li, Tucker Hermans, Stefan Lee |
ICRA | 5 |
| 2024 | DefGoalNet: Contextual Goal Learning from Demonstrations for Deformable Object ManipulationabstractShape servoing, a robotic task dedicated to controlling objects to desired goal shapes, is a promising approach to deformable object manipulation. An issue arises, however, with the reliance on the specification of a goal shape. This goal has been obtained either by a laborious domain knowledge engineering process or by manually manipulating the object into the desired shape and capturing the goal shape at that specific moment, both of which are impractical in various robotic applications. In this paper, we solve this problem by developing a novel neural network DefGoalNet, which learns deformable object goal shapes directly from a small number of human demonstrations. We demonstrate our method’s effectiveness on various robotic tasks, both in simulation and on a physical robot. Notably, in the surgical retraction task, even when trained with as few as 10 demonstrations, our method achieves a median success percentage of nearly 90%. These results mark a substantial advancement in enabling shape servoing methods to bring deformable object manipulation closer to practical real-world applications. Bao Thach, Tanner Watts, Shing-Hei Ho, Tucker Hermans, Alan Kuntz |
ICRA | 4 |
| 2024 | Neural Kinodynamic Planning: Learning for KinoDynamic Tree ExpansionabstractWe integrate neural networks into kinodynamic motion planning and present the Learning for KinoDynamic Tree Expansion (L4KDE) method. Tree-based planning approaches, such as rapidly exploring random tree (RRT), are the dominant approach to finding globally optimal plans in continuous state-space motion planning. Central to these approaches is tree expansion, the procedure in which new nodes are added to an ever-expanding tree. We study the kinodynamic variants of tree-based planning, where we have known system dynamics and kinematic constraints. In the interest of quickly selecting nodes to connect newly sampled coordinates, existing methods typically cannot optimise the finding of nodes that have a low cost to transition to sampled coordinates. Instead, they use metrics like Euclidean distance between coordinates as a heuristic for selecting candidate nodes to connect to the search tree. We propose L4KDE to address this issue. L4KDE uses a neural network to predict transition costs between queried states, which can be efficiently computed in batch, providing much higher quality estimates of transition cost compared to commonly used heuristics while maintaining almost-surely asymptotic optimality guarantee. We empirically demonstrate the significant performance improvement provided by L4KDE on a variety of challenging system dynamics,with the ability to generalise across different instances of the same model class and in conjunction with a suite of modern tree-based motion planners. Tin Lai, Weiming Zhi, Tucker Hermans, Fabio Ramos 0001 |
IROS | 3 |
| 2024 | V-PRISM: Probabilistic Mapping of Unknown Tabletop ScenesabstractThe ability to construct concise scene representations from sensor input is central to the field of robotics. This paper addresses the problem of robustly creating a 3D representation of a tabletop scene from a segmented RGBD image. These representations are then critical for a range of downstream manipulation tasks. Many previous attempts to tackle this problem do not capture accurate uncertainty, which is required to subsequently produce safe motion plans. In this paper, we cast the representation of 3D tabletop scenes as a multi-class classification problem. To tackle this, we introduce V-PRISM, a framework and method for robustly creating probabilistic 3D segmentation maps of tabletop scenes. Our maps contain both occupancy estimates, segmentation information, and principled uncertainty measures. We evaluate the robustness of our method in (1) procedurally generated scenes using open-source object datasets, and (2) real-world tabletop data collected from a depth camera. Our experiments show that our approach outperforms alternative continuous reconstruction approaches that do not explicitly reason about objects in a multi-class formulation. Herbert Wright, Weiming Zhi, Matthew Johnson-Roberson, Tucker Hermans |
IROS | 4 |
| 2024 | Position Regulation of a Conductive Nonmagnetic Object With Two Stationary Rotating-Magnetic-Dipole Field SourcesabstractEddy currents induced by rotating magnetic dipole fields can produce forces and torques that enable dexterous manipulation of conductive nonmagnetic objects. This paradigm shows promise for application in the remediation of space debris. The induced force from each rotating-magnetic-dipole field source always includes a repulsive component, suggesting that the object should be surrounded by field sources to some degree to ensure the object does not leave the dexterous workspace during manipulation. In this article, we show that it is possible to fully control the position of an object in a workspace near the midpoint between just two stationary field sources. A given position controller requires a low-level force controller. We propose two new force controllers, and compare them with the state-of-the-art method from the literature. One of the new force controllers is particularly good at not inducing parasitic torques, which is hypothesized to be beneficial for future tasks manipulating and detumbling rotating resident space objects. We perform experimental verification using numerical and physical simulators of microgravity. Devin K. Dalton, Griffin F. Tabor, Tucker Hermans, Jake J. Abbott |
IEEE Trans. Robotics | 3 |
| 2024 | Latent Space Planning for Multiobject Manipulation With Environment-Aware Relational ClassifiersabstractObjects rarely sit in isolation in everyday human environments. If we want robots to operate and perform tasks in our human environments, they must understand how the objects they manipulate will interact with structural elements of the environment for all but the simplest of tasks. As such, we would like our robots to reason about how multiple objects and environmental elements relate to one another and how those relations may change as the robot interacts with the world. We examine the problem of predicting interobject and object–environment relations between previously unseen objects and novel environments purely from partial-view point clouds. Our approach enables robots to plan and execute sequences to complete multiobject manipulation tasks defined from logical relations. This removes the burden of providing explicit, continuous object states as goals to the robot. We explore several different neural network architectures for this task. We find the best performing model to be a novel transformer-based neural network that both predicts object–environment relations and learns a latent-space dynamics function. We achieve reliable sim-to-real transfer without any fine-tuning. Our experiments show that our model understands how changes in observed environmental geometry relate to semantic relations between objects. Nichols Crawford Taylor, Adam Conkey, Tucker Hermans |
IEEE Trans. Robotics | 5 |
| 2023 | Planning for Multi-Object Manipulation with Graph Neural Network Relational ClassifiersabstractObjects rarely sit in isolation in human environments. As such, we'd like our robots to reason about how multiple objects relate to one another and how those relations may change as the robot interacts with the world. To this end, we propose a novel graph neural network framework for multi-object manipulation to predict how inter-object relations change given robot actions. Our model operates on partial-view point clouds and can reason about multiple objects dynamically interacting during the manipulation, By learning a dynamics model in a learned latent graph embedding space, our model enables multi-step planning to reach target goal relations. We show our model trained purely in simulation transfers well to the real world. Our planner enables the robot to rearrange a variable number of objects with a range of shapes and sizes using both push and pick-and-place skills. Adam Conkey, Tucker Hermans |
ICRA | 3 |
| 2023 | DefGraspNets: Grasp Planning on 3D Fields with Graph Neural NetsabstractRobotic grasping of 3D deformable objects is critical for real-world applications such as food handling and robotic surgery. Unlike rigid and articulated objects, 3D deformable objects have infinite degrees of freedom. Fully defining their state requires 3D deformation and stress fields, which are exceptionally difficult to analytically compute or experimentally measure. Thus, evaluating grasp candidates for grasp planning typically requires accurate, but slow 3D finite element method (FEM) simulation. Sampling-based grasp planning is often impractical, as it requires evaluation of a large number of grasp candidates. Gradient-based grasp planning can be more efficient, but requires a differentiable model to synthesize optimal grasps from initial candidates. Differentiable FEM simulators may fill this role, but are typically no faster than standard FEM. In this work, we propose learning a predictive graph neural network (GNN), DefGraspNets, to act as our differentiable model. We train DefGraspNets to predict 3D stress and deformation fields based on FEM-based grasp simulations. DefGraspNets not only runs up to 1500x faster than the FEM simulator, but also enables fast gradient-based grasp optimization over 3D stress and deformation metrics. We design DefGraspNets to align with real-world grasp planning practices and demonstrate generalization across multiple test sets, including real-world experiments. Isabella Huang, Yashraj Narang, Ruzena Bajcsy, Fabio Ramos 0001, Tucker Hermans, Dieter Fox |
ICRA | 5 |
| 2022 | StructFormer: Learning Spatial Structure for Language-Guided Semantic Rearrangement of Novel ObjectsabstractGeometric organization of objects into semantically meaningful arrangements pervades the built world. As such, assistive robots operating in warehouses, offices, and homes would greatly benefit from the ability to recognize and rearrange objects into these semantically meaningful structures. To be useful, these robots must contend with previously unseen objects and receive instructions without significant programming. While previous works have examined recognizing pairwise semantic relations and sequential manipulation to change these simple relations none have shown the ability to arrange objects into complex structures such as circles or table settings. To address this problem we propose a novel transformer-based neural network, StructFormer, which takes as input a partial-view point cloud of the current object arrangement and a structured language command encoding the desired object configuration. We show through rigorous experiments that StructFormer enables a physical robot to rearrange novel objects into semantically meaningful structures with multi-object relational constraints inferred from the language command. Chris Paxton 0001, Tucker Hermans, Dieter Fox |
ICRA | 3 |
| 2022 | Learning Visual Shape Control of Novel 3D Deformable Objects from Partial-View Point CloudsabstractIf robots could reliably manipulate the shape of 3D deformable objects, they could find applications in fields ranging from home care to warehouse fulfillment to surgical assistance. Analytic models of elastic, 3D deformable objects require numerous parameters to describe the potentially infinite degrees of freedom present in determining the object's shape. Previous attempts at performing 3D shape control rely on hand-crafted features to represent the object shape and require training of object-specific control models. We overcome these issues through the use of our novel DeformerNet neural network architecture, which operates on a partial-view point cloud of the object being manipulated and a point cloud of the goal shape to learn a low-dimensional representation of the object shape. This shape embedding enables the robot to learn to define a visual servo controller that provides Cartesian pose changes to the robot end-effector causing the object to deform towards its target shape. Crucially, we demonstrate both in simulation and on a physical robot that DeformerNet reliably generalizes to object shapes and material stiffness not seen during training and outperforms comparison methods for both the generic shape control and the surgical task of retraction. Bao Thach, Brian Y. Cho, Alan Kuntz, Tucker Hermans |
ICRA | 4 |
| 2022 | DULA and DEBA: Differentiable Ergonomic Risk Models for Postural Assessment and Optimization in Ergonomically Intelligent pHRIabstractErgonomics and human comfort are essential concerns in physical human-robot interaction applications. Defining an accurate and easy-to-use ergonomic assessment model stands as an important step in providing feedback for postural correction to improve operator health and comfort. Common practical methods in the area suffer from inaccurate ergonomics models in performing postural optimization. In order to retain assessment quality, while improving computational considerations, we propose a novel framework for postural assessment and optimization for ergonomically intelligent physical human-robot interaction. We introduce DULA and DEBA, differentiable and continuous ergonomics models learned to replicate the popular and scientifically validated RULA and REBA assessments with more than 99% accuracy. We show that DULA and DEBA provide assessment comparable to RULA and REBA while providing computational benefits when being used in postural optimization. We evaluate our framework through human and simulation experiments. We highlight DULA and DEBA's strength in a demonstration of postural optimization for a simulated pHRI task. Amir Yazdani, Roya Sabbagh Novin, Andrew Merryweather, Tucker Hermans |
IROS | 4 |
| 2021 | Optimizing Hospital Room Layout to Reduce the Risk of Patient FallsabstractDespite years of research into patient falls in hospital rooms, falls and related injuries remain a serious concern to patient safety. In this work, we formulate a gradient-free constrained optimization problem to generate and reconfigure the hospital room interior layout to minimize the risk of falls. We define a cost function built on a hospital room fall model that takes into account the supportive or hazardous effect of the patient's surrounding objects, as well as simulated patient trajectories inside the room. We define a constraint set that ensures the functionality of the generated room layouts in addition to conforming to architectural guidelines. We solve this problem efficiently using a variant of simulated annealing. We present results for two real-world hospital room types and demonstrate a significant improvement of 18% on average in patient fall risk when compared with a traditional hospital room layout and 41% when compared with randomly generated layouts. Sarvenaz Chaeibakhsh, Roya Sabbagh Novin, Tucker Hermans, Andrew Merryweather, Alan Kuntz |
ICORES | 3 |
| 2021 | Risk-Aware Decision Making for Service Robots to Minimize Risk of Patient Falls in HospitalsabstractPlanning under uncertainty is a crucial capability for autonomous systems to operate reliably in uncertain and dynamic environments. The concern of safety becomes even more critical in healthcare settings where robots interact with human patients. In this paper, we propose a novel risk-aware planning framework to minimize the risk of falls by providing a patient with an assistive device. Our approach combines learning-based prediction with model-based control to plan for the fall prevention task. This provides advantages compared to end-to-end learning methods in which the robot's performance is limited to specific scenarios, or purely model-based approaches that use relatively simple function approximators and are prone to high modeling errors. We compare various risk metrics and the results from simulated scenarios show that using the proposed cost function, the robot can plan interventions to avoid high fall score events. Roya Sabbagh Novin, Amir Yazdani, Andrew Merryweather, Tucker Hermans |
ICRA | 4 |
| 2021 | Near-Optimal Area-Coverage Path Planning of Energy-Constrained Aerial Robots With Application in Autonomous Environmental MonitoringabstractThis article describes a Voronoi-based path generation (VPG) algorithm for an energy-constrained mobile robot, such as an unmanned aerial vehicle (UAV). The algorithm solves a variation of the coverage path-planning problem where complete coverage of an area is not possible due to path-length limits caused by energy constraints on the robot. The algorithm works by modeling the path as a connected network of mass-spring-damper systems. The approach further leverages the properties of Voronoi diagrams to generate a potential field to move path waypoints to near-optimal configurations while maintaining path-length constraints. Simulation and physical experiments on an aerial vehicle are described. Simulated runtimes show linear-time complexity with respect to the number of path waypoints. Tests in variously shaped areas demonstrate that the method can generate paths in both convex and nonconvex areas. Comparison tests with other path generation methods demonstrate that the VPG algorithm strikes a good balance between runtime and optimality, with significantly better runtime than direct optimization, lower cost coverage paths than a lawnmower-style coverage path, and moderately better performance in both metrics than the most conceptually similar method. Physical experiments demonstrate the applicability of the VPG method to a physical UAV, and comparisons between real-world results and simulations show that the costs of the generated paths are within a few percent of each other, implying that analysis performed in simulation will hold for real-world application, assuming that the robot is capable of closely following the path and a good energy model is available.Note to Practitioners—For autonomous mobile-robotics-based applications where a robot equipped with a tool or sensor is required to survey an area for inspection, monitoring, cleaning, and so on, effectively covering the area is desirable. However, for energy-constrained systems such as aerial vehicles with limited flight time, complete coverage is not possible. Presented here is a new Voronoi-based path generation algorithm that takes energy constraints into account to generate waypoints for the robot to follow in a near-optimal configuration while maintaining path-length constraints. The approach is applied in simulation and experiments for an application in environmental monitoring using unmanned aerial vehicles. Katharin R. Jensen-Nau, Tucker Hermans, Kam K. Leang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2021 | In-Hand Object-Dynamics Inference Using Tactile FingertipsabstractHaving the ability to estimate an object's properties through interaction will enable robots to manipulate novel objects. Object's dynamics, specifically the friction and inertial parameters have only been estimated in a lab environment with precise and often external sensing. Could we infer an object's dynamics in the wild with only the robot's sensors? In this article, we explore the estimation of dynamics of a grasped object in motion, with tactile force sensing at multiple fingertips. Our estimation approach does not rely on torque sensing to estimate the dynamics. To estimate friction, we develop a control scheme to actively interact with the object until slip is detected. To robustly perform the inertial estimation, we setup a factor graph that fuses all our sensor measurements on physically consistent manifolds and perform inference. We show that tactile fingertips enable in-hand dynamics estimation of low mass objects. Balakumar Sundaralingam, Tucker Hermans |
IEEE Trans. Robotics | 2 |
| 2020 | Learning Continuous 3D Reconstructions for Geometrically Aware GraspingabstractDeep learning has enabled remarkable improvements in grasp synthesis for previously unseen objects from partial object views. However, existing approaches lack the ability to explicitly reason about the full 3D geometry of the object when selecting a grasp, relying on indirect geometric reasoning derived when learning grasp success networks. This abandons explicit geometric reasoning, such as avoiding undesired robot object collisions. We propose to utilize a novel, learned 3D reconstruction to enable geometric awareness in a grasping system. We leverage the structure of the reconstruction network to learn a grasp success classifier which serves as the objective function for a continuous grasp optimization. We additionally explicitly constrain the optimization to avoid undesired contact, directly using the reconstruction. We examine the role of geometry in grasping both in the training of grasp metrics and through 96 robot grasping trials. Our results can be found on https://sites.google.com/view/reconstruction-grasp/. Mark Van der Merwe, Qingkai Lu, Balakumar Sundaralingam, Martin Matak, Tucker Hermans |
ICRA | 5 |
| 2020 | Multi-Fingered Active Grasp LearningabstractLearning-based approaches to grasp planning are preferred over analytical methods due to their ability to better generalize to new, partially observed objects. However, data collection remains one of the biggest bottlenecks for grasp learning methods, particularly for multi-fingered hands. The relatively high dimensional configuration space of the hands coupled with the diversity of objects common in daily life requires a significant number of samples to produce robust and confident grasp success classifiers. In this paper, we present the first active deep learning approach to grasping that searches over the grasp configuration space and classifier confidence in a unified manner. We base our approach on recent success in planning multi-fingered grasps as probabilistic inference with a learned neural network likelihood function. We embed this within a multi-armed bandit formulation of sample selection. We show that our active grasp learning approach uses fewer training samples to produce grasp success rates comparable with the passive supervised learning method trained with grasping data generated by an analytical planner. We additionally show that grasps generated by the active learner have greater qualitative and quantitative diversity in shape. Qingkai Lu, Mark Van der Merwe, Tucker Hermans |
IROS | 3 |
| 2019 | Robust Learning of Tactile Force Estimation through Robot InteractionabstractCurrent methods for estimating force from tactile sensor signals are either inaccurate analytic models or task-specific learned models. In this paper, we explore learning a robust model that maps tactile sensor signals to force. We specifically explore learning a mapping for the SynTouch BioTac sensor via neural networks. We propose a voxelized input feature layer for spatial signals and leverage information about the sensor surface to regularize the loss function. To learn a robust tactile force model that transfers across tasks, we generate ground truth data from three different sources: (1) the BioTac rigidly mounted to a force torque (FT) sensor, (2) a robot interacting with a ball rigidly attached to the same FT sensor, and (3) through force inference on a planar pushing task by formalizing the mechanics as a system of particles and optimizing over the object motion. A total of 140k samples were collected from the three sources. We achieve a median angular accuracy of 3.5 degrees in predicting force direction (66% improvement over the current state of the art) and a median magnitude accuracy of 0.06 N (93% improvement) on a test dataset. Additionally, we evaluate the learned force model in a force feedback grasp controller performing object lifting and gentle placement. Our results can be found on https: //sites.google.com/view/tactile-force. Balakumar Sundaralingam, Alexander Lambert, Ankur Handa, Byron Boots, Tucker Hermans, Stanley T. Birchfield, Nathan D. Ratliff, Dieter Fox |
ICRA | 5 |
| 2018 | Geometric In-Hand Regrasp Planning: Alternating Optimization of Finger Gaits and In-Grasp ManipulationabstractThis paper explores the problem of autonomous, in-hand regrasping-the problem of moving from an initial grasp on an object to a desired grasp using the dexterity of a robot's fingers. We propose a planner for this problem which alternates between finger gaiting, and in-grasp manipulation. Finger gaiting enables the robot to move a single finger to a new contact location on the object, while the remaining fingers stably hold the object. In-grasp manipulation moves the object to a new pose relative to the robot's palm, while maintaining the contact locations between the hand and object. Given the object's geometry (as a mesh), the hand's kinematic structure, and the initial and desired grasps, we plan a sequence of finger gaits and object reposing actions to reach the desired grasp without dropping the object. We propose an optimization based approach and report in-hand regrasping plans for 5 objects over 5 in-hand regrasp goals each. The plans generated by our planner are collision free and guarantee kinematic feasibility. Balakumar Sundaralingam, Tucker Hermans |
ICRA | 2 |
| 2018 | Dynamic Model Learning and Manipulation Planning for Objects in Hospitals Using a Patient Assistant Mobile (PAM)RobotabstractOne of the most concerning and costly problems in hospitals is patients falls. We address this problem by introducing PAM, a patient assistant mobile robot, that maneuvers mobility aids to assist with fall prevention. Common objects found inside hospitals include objects with legs (i.e. walkers, tables, chairs, equipment stands). For a mobile robot operating in such environments, safely maneuvering these objects without collision is essential. Since providing the robot with dynamic models of all possible legged objects that may exist in such environments is not feasible, autonomous learning of an approximate dynamic model for these objects would significantly improve manipulation planning. We describe a probabilistic method to do this by fitting pre-categorized object models learned from minimal force and motion interactions with an object. In addition, we account for multiple manipulation strategies, which requires a hybrid control system comprised of discrete grasps on legs and continuous applied forces. To do this, we use a simple one-wheel point-mass model. A hybrid MPC-based manipulation planning algorithm was developed to compensate for modeling errors. While the proposed algorithm applies to a broad range of legged objects, we only show results for the case of a 2-wheel, 4-legged walker in this paper. Simulation and experimental tests show that the obtained dynamic model is sufficiently accurate for safe and collision-free manipulation. When combined with the proposed manipulation planning algorithm, the robot can successfully move the object to a desired position without collision. Roya Sabbagh Novin, Amir Yazdani, Tucker Hermans, Andrew Merryweather |
IROS | 3 |
| 2017 | First demonstration of simultaneous localization and propulsion of a magnetic capsule in a lumen using a single rotating magnetabstractThis paper presents a method for closed-loop propulsion of a screw-type magnetic capsule with embedded Hall-effect sensors using a single rotating actuator magnet. The method estimates the six-degree-of-freedom (6-DOF) pose of the capsule while it is synchronously rotating with the applied field. It is intended for application in active capsule endoscopy of the intestines. An extended Kalman filter, which uses a simplified 2-DOF process model restricting the capsule to forward or backward movement and rotation about its principle axis, is used to provide a full 6-DOF estimate of the capsule's pose as the capsule travels through a lumen. The capsule's movement in the applied field is constantly monitored to determine if the capsule is synchronously rotating with the applied field. Based on this information, the rotation speed of the external source is adjusted to prevent a loss in the desired magnetic coupling. We experimentally demonstrate, for the first time, simultaneous localization and closed-loop propulsion of a capsule through a lumen using a single rotating magnet. Prior work assumed the capsule had no net motion during the localization phase, requiring decoupled localization and propulsion. This closed-loop performance results in a three times speed up in completion time, compared to the previous decoupled approach. Katie M. Popek, Tucker Hermans, Jake J. Abbott |
ICRA | 2 |
| 2017 | Planning Multi-fingered Grasps as Probabilistic Inference in a Learned Deep Network
Qingkai Lu, Kautilya Chenna, Balakumar Sundaralingam, Tucker Hermans |
ISRR | 4 |
| 2016 | Active tactile object exploration with Gaussian processesabstractAccurate object shape knowledge provides important information for performing stable grasping and dexterous manipulation. When modeling an object using tactile sensors, touching the object surface at a fixed grid of points can be sample inefficient. In this paper, we present an active touch strategy to efficiently reduce the surface geometry uncertainty by leveraging a probabilistic representation of object surface. In particular, we model the object surface using a Gaussian process and use the associated uncertainty information to efficiently determine the next point to explore. We validate the resulting method for tactile object surface modeling using a real robot to reconstruct multiple, complex object surfaces. Zhengkun Yi, Roberto Calandra, Filipe Veiga, Herke van Hoof, Tucker Hermans, Jan Peters 0001 |
IROS | 5 |
| 2015 | Stabilizing novel objects by learning to predict tactile slipabstractDuring grasping and other in-hand manipulation tasks maintaining a stable grip on the object is crucial for the task's outcome. Inherently connected to grip stability is the concept of slip. Slip occurs when the contact between the fingertip and the object is partially lost, resulting in sudden undesired changes to the objects state. While several approaches for slip detection have been proposed in the literature, they frequently rely on previous knowledge of the manipulated object. This previous knowledge may be unavailable, seeing that robots operating in real-world scenarios often must interact with previously unseen objects. In our work we explore the generalization capabilities of well known supervised learning methods, using random forest classifiers to create generalizable slip predictors. We utilize these classifiers in the feedback loop of an object stabilization controller. We show that the controller can successfully stabilize previously unknown objects by predicting and counteracting slip events. Filipe Veiga, Herke van Hoof, Jan Peters 0001, Tucker Hermans |
IROS | 4 |
| 2013 | An In Depth View of SaliencyabstractVisual saliency is a computational process that identifies important locations and structure in the visual field.Most current methods for saliency rely on cues such as color and texture while ignoring depth information, which is known to be an important saliency cue in the human cognitive system.We propose a novel computational model of visual saliency which incorporates depth information.We compare our approach to several state of the art visual saliency methods and we introduce a method for saliency based segmentation of generic objects.We demonstrate that by explicitly constructing 3D layout and shape features from depth measurements, we can obtain better performance than methods which treat the depth map as just another image channel.Our method requires no learning and can operate on scenes for which the system has no previous knowledge.We conduct object segmentation experiments on a new dataset of registered RGB-D images captured on a mobile-manipulator robot. Arridhana Ciptadi, Tucker Hermans, James M. Rehg |
BMVC | 2 |
| 2013 | Decoupling behavior, perception, and control for autonomous learning of affordancesabstractA novel behavior representation is introduced that permits a robot to systematically explore the best methods by which to successfully execute an affordance-based behavior for a particular object. The approach decomposes affordance-based behaviors into three components. We first define controllers that specify how to achieve a desired change in object state through changes in the agent's state. For each controller we develop at least one behavior primitive that determines how the controller outputs translate to specific movements of the agent. Additionally we provide multiple perceptual proxies that define the representation of the object that is to be computed as input to the controller during execution. A variety of proxies may be selected for a given controller and a given proxy may provide input for more than one controller. When developing an appropriate affordance-based behavior strategy for a given object, the robot can systematically vary these elements as well as note the impact of additional task variables such as location in the workspace. We demonstrate the approach using a PR2 robot that explores different combinations of controller, behavior primitive, and proxy to perform a push or pull positioning behavior on a selection of household objects, learning which methods best work for each object. Tucker Hermans, James M. Rehg, Aaron F. Bobick |
ICRA | 1 |
| 2012 | Guided pushing for object singulationabstractWe propose a novel method for a robot to separate and segment objects in a cluttered tabletop environment. The method leverages the fact that external object boundaries produce visible edges within an object cluster. We achieve this singulation of objects by using the robot arm to perform pushing actions specifically selected to test whether particular visible edges correspond to object boundaries. We verify the separation of objects after a push by examining the clusters formed by geometric segmentation of regions residing on the table surface. To avoid explicitly representing and tracking edges across push behaviors we aggregate over all edges in a given orientation by representing the push-history as an orientation histogram. By tracking the history of directions pushed for each object cluster we can build evidence that a cluster cannot be further separated. We present quantitative and qualitative experimental results performed in a real home environment by a mobile manipulator using input from an RGB-D camera mounted on the robot's head. We show that our pushing strategy can more reliably obtain singulation in fewer pushes than an approach, that does not explicitly reason about boundary information. Tucker Hermans, James M. Rehg, Aaron F. Bobick |
IROS | 1 |
| 2011 | Push planning for object placement on cluttered table surfacesabstractWe present a novel planning algorithm for the problem of placing objects on a cluttered surface such as a table, counter or floor. The planner (1) selects a placement for the target object and (2) constructs a sequence of manipulation actions that create space for the object. When no continuous space is large enough for direct placement, the planner leverages means-end analysis and dynamic simulation to find a sequence of linear pushes that clears the necessary space. Our heuristic for determining candidate placement poses for the target object is used to guide the manipulation search. We show successful results for our algorithm in simulation. Akansel Cosgun, Tucker Hermans, Victor Emeli, Mike Stilman |
IROS | 2 |
| 2010 | Movie genre classification via scene categorizationabstractThis paper presents a method for movie genre categorization of movie trailers, based on scene categorization. We view our approach as a step forward from using only low-level visual feature cues, towards the eventual goal of high-level seman- tic understanding of feature films. Our approach decom- poses each trailer into a collection of keyframes through shot boundary analysis. From these keyframes, we use state-of- the-art scene detectors and descriptors to extract features, which are then used for shot categorization via unsuper- vised learning. This allows us to represent trailers using a bag-of-visual-words (bovw) model with shot classes as vo- cabularies. We approach the genre classification task by mapping bovw temporally structured trailer features to four high-level movie genres: action, comedy, drama or horror films. We have conducted experiments on 1239 annotated trailers. Our experimental results demonstrate that exploit- ing scene structures improves film genre classification com- pared to using only low-level visual features. Howard Zhou, Tucker Hermans, Asmita V. Karandikar, James M. Rehg |
ACM Multimedia | 2 |
| 2008 | Player Positioning in the Four-Legged League
Henry Work, Eric Chown, Tucker Hermans, Jesse Butterfield, Mark McGranaghan 0001 |
RoboCup | 3 |