EDBT 2026 Demo / reviewers in the wild / expert
Gregory D. Hager
dblp:12/5814 · also Greg Hager
· DBLP profile ↗
223ranked-venue papers
20as first author
17since 2021 · last 2026
0000-0002-6662-9763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 144 · 18 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 89 · 5 first-author · 8 since 2021Systems, architecture and hardware · 78 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 64 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3Human-computer interaction and ubiquitous computing · 3Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAST: Evaluating Multi-Object Trackers with Context-Aware Switch and Transfer ScoresabstractMulti-object tracking (MOT) has been a subject of intensive research for decades. Multiple standard datasets and benchmarks have been set up, and several evaluation metrics, such as MOTA, IDF1 and HOTA. These metrics have become the de facto standard for comparing and ranking trackers on standardized datasets to measure progress. In this paper, we focus on MOTA and HOTA, and present a study of cases where these metrics’ behaviors may not be desirable. In addition, we demonstrate how they might not be ideal when used as a tool to inspect a tracker’s failure cases. We point out that these issues are related to the sizes of the context windows in which they measure association quality, where MOTA is too nearsighted while HOTA can be too holistic depending on the task settings.In this paper, we rethink the familiar notion of identity switches (IDSw) proposed in MOTA, and propose a generalized version of it by introducing a context window when evaluating the ID assignment choice for each detection. We show that the proposed metric, CAST, mitigates the limitations of MOTA and HOTA, and demonstrate its usefulness when diagnosing model failures through examples. Our code and toolkit will be made available at https://github.com/bkkm78/cast. Jin Bai 0001, Gregory D. Hager |
WACV | 2 |
| 2024 | Domain Adaptation of Visual Policies with a Single DemonstrationabstractDeploying machine learning algorithms for robot tasks in real-world applications presents a core challenge: overcoming the domain gap between the training and the deployment environment. This is particularly difficult for visuomotor policies that utilize high-dimensional images as input, particularly when those images are generated via simulation. A common method to tackle this issue is through domain randomization, which aims to broaden the span of the training distribution to cover the test-time distribution. However, this approach is only effective when the domain randomization encompasses the actual shifts in the test-time distribution. We take a different approach, where we make use of a single demonstration (a prompt) to learn policy that adapts to the testing target environment. Our proposed framework, PromptAdapt, leverages the Transformer architecture’s capacity to model sequential data to learn demonstration-conditioned visual policies, allowing for in-context adaptation to a target domain that is distinct from training. Our experiments in both simulation and real-world settings show that PromptAdapt is a strong domain-adapting policy that outperforms baseline methods by a large margin under a range of domain shifts, including variations in lighting, color, texture, and camera pose. Videos and more information can be viewed at project webpage: https://sites.google.com/view/promptadapt. Weiyao Wang 0002, Gregory D. Hager |
ICRA | 2 |
| 2024 | VIHE: Virtual In-Hand Eye Transformer for 3D Robotic ManipulationabstractIn this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong inductive bias for effectively recognizing the correct pose for the hand, especially for challenging high-precision tasks such as peg insertion. On 18 manipulation tasks in RLBench simulated environments, VIHE achieves a new state-of-the-art, with a 12% absolute improvement, increasing from 65% to 77% over the existing state-of-the-art model using 100 demonstrations per task. In real-world scenarios, VIHE can learn manipulation tasks with just a handful of demonstrations, highlighting its practical utility. Videos and code implementation can be found at our project site: https://vihe-3d.github.io. Weiyao Wang 0002, Shiyu Jin, Gregory D. Hager, Liangjun Zhang |
IROS | 4 |
| 2024 | Embedding Task Structure for Action DetectionabstractWe present a straightforward, flexible method to enhance the accuracy and quality of action detection by expressing temporal and structural relationships of actions in the loss function of a deep network. We describe ways to represent otherwise implicit structure in video data and demonstrate how these structures reflect natural biases that improve network training. Our experiments show that our approach improves both accuracy and edit-distance of action recognition and detection models over a baseline. Our framework leads to improvements over prior work and obtains state-of-the-art results on multiple benchmarks. The code is available here. Michael Peven, Gregory D. Hager |
WACV | 2 |
| 2023 | Naive human judges can accurately predict expertise in children's block building. Can embedded motion sensors do just as well?
E. Emory Davis, Anand Malpani, Gregory D. Hager, Amy Lynne Shelton, Barbara Landau |
CogSci | 4 |
| 2023 | Mapping DNN Embedding Manifolds for Network Generalization PredictionabstractDeep Neural Networks(DNN) often fail in surprising ways, and predicting how well a trained DNN will generalize in a new, external operating domain is essential for deploying DNNs in safety critical applications, e.g., perception for self-driving vehicles or medical image analysis. Recently, the task of Network Generalization Prediction (NGP) has been proposed to predict how a DNN will generalize in an external operating domain. Previous NGP approaches have leveraged multiple labeled test sets or labeled metadata. In this study, we propose an embedding map, the first NGP approach that predicts DNN performance based on how unlabeled images from an external operating domain map in the DNN embedding space. We evaluate our proposed Embedding Map and other recently proposed NGP approaches for pedestrian, melanoma, and animal classification tasks. We find that our embedding map has the best average NGP performance, and that our embedding map is effective at modeling complex, non-linear embedding space structures. Molly O'Brien, Brett Wolfinger, Julia V. Bukowski, Mathias Unberath, Aria Pezeshk, Gregory D. Hager |
WACV | 6 |
| 2022 | Coarse-To-Fine Incremental Few-Shot Learning
Xiang Xiang 0001, Yuwen Tan, Alan L. Yuille, Gregory D. Hager |
ECCV (31) | 6 |
| 2022 | SAGE: SLAM with Appearance and Geometry Prior for Endoscopyabstract., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry of the observed anatomy from a monocular endoscopic video. To this end, we develop a Simultaneous Localization and Mapping system by combining the learning-based appearance and optimizable geometry priors and factor graph optimization. The appearance and geometry priors are explicitly learned in an end-to-end differentiable training pipeline to master the task of pair-wise image alignment, one of the core components of the SLAM system. In our experiments, the proposed SLAM system is shown to robustly handle the challenges of texture scarceness and illumination variation that are commonly seen in endoscopy. The system generalizes well to unseen endoscopes and subjects and performs favorably compared with a state-of-the-art feature-based SLAM system. The code repository is available at https://github.com/lppllppl920/SAGE-SLAM.git. Xingtong Liu, Zhaoshuo Li, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
ICRA | 4 |
| 2022 | Network Generalization Prediction for Safety Critical Tasks in Novel Operating DomainsabstractIt is well known that Neural Network (network) performance often degrades when a network is used in novel operating domains that differ from its training and testing domains. This is a major limitation, as networks are being integrated into safety critical, cyber-physical systems that must work in unconstrained environments, e.g., perception for autonomous vehicles. Training networks that generalize to novel operating domains and that extract robust features is an active area of research, but previous work fails to predict what the network performance will be in novel operating domains. We propose the task Network Generalization Prediction: predicting the expected network performance in novel operating domains. We describe the network performance in terms of an interpretable Context Subspace, and we propose a methodology for selecting the features of the Context Subspace that provide the most information about the network performance. We identify the Context Subspace for a pretrained Faster RCNN network performing pedestrian detection on the Berkeley Deep Drive (BDD) Dataset, and demonstrate Network Generalization Prediction accuracy within 5% of observed performance. We also demonstrate that the Context Subspace from the BDD Dataset is informative for completely unseen datasets, JAAD and Cityscapes, where predictions have a bias of 10% or less. Molly O'Brien, Mike Medoff, Julia V. Bukowski, Gregory D. Hager |
WACV | 4 |
| 2022 | Surgical data science - from concepts toward clinical translationabstractRecent developments in data science in general and machine learning in particular have transformed the way experts envision the future of surgery. Surgical Data Science (SDS) is a new research field that aims to improve the quality of interventional healthcare through the capture, organization, analysis and modeling of data. While an increasing number of data-driven approaches and clinical applications have been studied in the fields of radiological and clinical data science, translational success stories are still lacking in surgery. In this publication, we shed light on the underlying reasons and provide a roadmap for future advances in the field. Based on an international workshop involving leading researchers in the field of SDS, we review current practice, key achievements and initiatives as well as available standards and tools for a number of topics relevant to the field, namely (1) infrastructure for data acquisition, storage and access in the presence of regulatory constraints, (2) data annotation and sharing and (3) data analytics. We further complement this technical perspective with (4) a review of currently available SDS products and the translational progress from academia and (5) a roadmap for faster clinical translation and exploitation of the full potential of SDS, based on an international multi-round Delphi process. Lena Maier-Hein, Matthias Eisenmann, Duygu Sarikaya, Keno März, Toby Collins, Anand Malpani, Johannes Fallert, Hubertus Feußner, Stamatia Giannarou, Pietro Mascagni, Hirenkumar Nakawala, Adrian Park 0001, Carla M. Pugh, Danail Stoyanov, S. Swaroop Vedula, Kevin Cleary, Gabor Fichtinger, Germain Forestier, Bernard Gibaud, Teodor P. Grantcharov, Makoto Hashizume, Doreen Heckmann-Nötzel, Hannes Kenngott, Ron Kikinis, Lars Mündermann, Nassir Navab, Sinan Onogur, Tobias Roß, Raphael Sznitman, Russell H. Taylor, Minu Tizabi, Martin Wagner 0001, Gregory D. Hager, Thomas Neumuth, Nicolas Padoy, Justin Collins, Ines Gockel, Jan Goedeke, Daniel A. Hashimoto, Luc Joyeux, Kyle Lam, Daniel Richard Leff, Amin Madani, Hani J. Marcus, Ozanan R. Meireles, Alexander Seitel, Dogu Teber, Frank Ückert, Beat P. Müller-Stich, Pierre Jannin, Stefanie Speidel |
Medical Image Anal. | 33 |
| 2021 | DASZL: Dynamic Action Signatures for Zero-shot LearningabstractThere are many realistic applications of activity recognition where the set of potential activity descriptions is combinatorially large. This makes end-to-end supervised training of a recognition system impractical as no training set is practically able to encompass the entire label set. In this paper, we present an approach to fine-grained recognition that models activities as compositions of dynamic action signatures. This compositional approach allows us to reframe fine-grained recognition as zero-shot activity recognition, where a detector is composed "on the fly" from simple first-principles state machines supported by deep-learned components. We evaluate our method on the Olympic Sports and UCF101 datasets, where our model establishes a new state of the art under multiple experimental paradigms. We also extend this method to form a unique framework for zero-shot joint segmentation and classification of activities in video and demonstrate the first results in zero- shot decoding of complex action sequences on a widely-used surgical dataset. Lastly, we show that we can use off-the-shelf object detectors to recognize activities in completely de-novo settings with no additional training. Tae Soo Kim 0001, Jonathan D. Jones, Michael Peven, Zihao Xiao 0001, Jin Bai 0001, Yi Zhang 0099, Weichao Qiu, Alan L. Yuille, Gregory D. Hager |
AAAI | 9 |
| 2021 | Neighborhood Normalization for Robust Geometric Feature LearningabstractExtracting geometric features from 3D models is a common first step in applications such as 3D registration, tracking, and scene flow estimation. Many hand-crafted and learning-based methods aim to produce consistent and distinguishable geometric features for 3D models with partial overlap. These methods work well in cases where the point density and scale of the overlapping 3D objects are similar, but struggle in applications where 3D data are obtained independently with unknown global scale and scene overlap. Unfortunately, instances of this resolution mismatch are common in practice, e.g., when aligning data from multiple sensors. In this work, we introduce a new normalization technique, Batch-Neighborhood Normalization, aiming to improve robustness to mean-std variation of local feature distributions that presumably can happen in samples with varying point density. We empirically demonstrate that the presented normalization method’s performance compares favorably to comparison methods in indoor and outdoor environments, and on a clinical dataset, on common point registration benchmarks in both standard and, particularly, resolution-mismatch settings. The source code and clinical dataset are available at https://github.com/lppllppl920/NeighborhoodNormalization-Pytorch. Xingtong Liu, Benjamin Killeen, Ayushi Sinha, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
CVPR | 5 |
| 2021 | Motion Guided Attention Fusion to Recognize Interactions from VideosabstractWe present a dual-pathway approach for recognizing fine-grained interactions from videos. We build on the success of prior dual-stream approaches, but make a distinction between the static and dynamic representations of objects and their interactions explicit by introducing separate motion and object detection pathways. Then, using our new Motion-Guided Attention Fusion module, we fuse the bottom-up features in the motion pathway with features captured from object detections to learn the temporal aspects of an action. We show that our approach can generalize across appearance effectively and recognize actions where an actor interacts with previously unseen objects. We validate our approach using the compositional action recognition task from the Something-Something-v2 dataset where we outperform existing state-of-the-art methods. We also show that our method can generalize well to real world tasks by showing state-of-the-art performance on recognizing humans assembling various IKEA furniture on the IKEA-ASM dataset. Tae Soo Kim 0001, Jonathan D. Jones, Gregory D. Hager |
ICCV | 3 |
| 2021 | Out-of-Distribution Robustness with Deep Recursive FiltersabstractAccurate state and uncertainty estimation is imperative for mobile robots and self driving vehicles to achieve safe navigation in pedestrian rich environments. A critical component of state and uncertainty estimation for robot navigation is to perform robustly under out-of-distribution noise. Traditional methods of state estimation decouple perception and state estimation making it difficult to operate on noisy, high dimensional data. Here, we describe an approach that combines the expressiveness of deep neural networks with principled approaches to uncertainty estimation found in recursive filters. We particularly focus on techniques that provide better robustness to out-of-distribution noise and demonstrate applicability of our approach on two scenarios: a simple noisy pendulum state estimation problem and real world pedestrian localization using the nuScenes dataset [1]. We show that our approach improves state and uncertainty estimation compared to baselines while achieving approximately 3× improvement in computational efficiency. Kapil D. Katyal, I-Jeng Wang, Gregory D. Hager |
ICRA | 3 |
| 2021 | Cumulative Assessment for Urban 3D ModelingabstractUrban 3D modeling from satellite images requires accurate semantic segmentation to delineate urban features, multiple view stereo for 3D reconstruction of surface heights, and 3D model fitting to produce compact models with accurate surface slopes. In this work, we present a cumulative assessment metric that succinctly captures error contributions from each of these components. We demonstrate our approach by providing challenging public datasets and extending two open source projects to provide an end-to-end 3D modeling baseline solution to stimulate further research and evaluation with a public leaderboard. Shea Hagstrom, Hee Won Pak, Stephanie Ku, Sean Wang 0001, Gregory D. Hager, Myron Z. Brown |
IGARSS | 5 |
| 2021 | Localization and Control of Magnetic Suture Needles in Cluttered Surgical Site with Blood and TissueabstractReal-time visual localization of needles is necessary for various surgical applications, including surgical automation and visual feedback. In this study we investigate localization and autonomous robotic control of needles in the context of our magneto-suturing system. Our system holds the potential for surgical manipulation with the benefit of minimal invasiveness and reduced patient side effects. However, the nonlinear magnetic fields produce unintuitive forces and demand delicate position-based control that exceeds the capabilities of direct human manipulation. This makes automatic needle localization a necessity. Our localization method combines neural network-based segmentation and classical techniques, and we are able to consistently locate our needle with 0.73 mm RMS error in clean environments and 2.72 mm RMS error in challenging environments with blood and occlusion. The average localization RMS error is 2.16 mm for all environments we used in the experiments. We combine this localization method with our closed-loop feedback control system to demonstrate the further applicability of localization to autonomous control. Our needle is able to follow a running suture path in (1) no blood, no tissue; (2) heavy blood, no tissue; (3) no blood, with tissue; and (4) heavy blood, with tissue environments. The tip position tracking error ranges from 2.6 mm to 3.7 mm RMS, opening the door towards autonomous suturing tasks. Will Pryor, Yotam Barnoy, Suraj Raval, Xiaolong Liu 0002, Lamar O. Mair, Daniel Lerner, Onder Erin, Gregory D. Hager, Yancy Diaz-Mercado, Axel Krieger |
IROS | 8 |
| 2021 | Robust Policy Search for an Agile Ground Vehicle Under Perception UncertaintyabstractLearning robust policies for robotic systems operating in presence of uncertainty is a challenging task. For safe navigation, in addition to the natural stochasticity of the environment and vehicle dynamics, the perception uncertainty associated with dynamic entities, e.g. pedestrians, must be accounted for during motion planning. To this end, we construct an algorithm with built-in robustness to uncertainty by directly minimizing an upper confidence bound on the expected cost of trajectories instead of employing a standard approach based on minimizing the expected cost itself. Perception uncertainty is incorporated into the policy search framework by predicting each pedestrian’s intent belief and propagating their state distribution in time using closed-loop goal-directed dynamics. We train the policy in simulation and show that it could be transferred to an agile ground vehicle for successful autonomous robot navigation in presence of pedestrians with perception uncertainty. We further show the superior performance of this policy over a policy that does not consider pedestrian intent and perception uncertainty. Shahriar Sefati, Subhransu Mishra, Matthew Sheckells, Kapil D. Katyal, Jin Bai 0001, Gregory D. Hager, Marin Kobilarov |
IROS | 6 |
| 2020 | Learning Geocentric Object Pose in Oblique Monocular ImagesabstractAn object's geocentric pose, defined as the height above ground and orientation with respect to gravity, is a powerful representation of real-world structure for object detection, segmentation, and localization tasks using RGBD images. For close-range vision tasks, height and orientation have been derived directly from stereo-computed depth and more recently from monocular depth predicted by deep networks. For long-range vision tasks such as Earth observation, depth cannot be reliably estimated with monocular images. Inspired by recent work in monocular height above ground prediction and optical flow prediction from static images, we develop an encoding of geocentric pose to address this challenge and train a deep network to compute the representation densely, supervised by publicly available airborne lidar. We exploit these attributes to rectify oblique images and remove observed object parallax to dramatically improve the accuracy of localization and to enable accurate alignment of multiple images taken from very different oblique viewpoints. We demonstrate the value of our approach by extending two large-scale public datasets for semantic segmentation in oblique satellite images. All of our data and code are publicly available. Gordon A. Christie, Rodrigo Rene Rai Munoz Abujder, Kevin Foster, Shea Hagstrom, Gregory D. Hager, Myron Z. Brown |
CVPR | 5 |
| 2020 | Semantic Image Manipulation Using Scene GraphsabstractImage manipulation can be considered a special case of image generation where the image to be produced is a modification of an existing image. Image generation and manipulation have been, for the most part, tasks that operate on raw pixels. However, the remarkable progress in learning rich image and object representations has opened the way for tasks such as text-to-image or layout-to-image generation that are mainly driven by semantics. In our work, we address the novel problem of image manipulation from scene graphs, in which a user can edit images by merely applying changes in the nodes or edges of a semantic graph that is generated from the image. Our goal is to encode image information in a given constellation and from there on generate new constellations, such as replacing objects or even changing relationships between objects, while respecting the semantics and style from the original image. We introduce a spatio-semantic scene graph network that does not require direct supervision for constellation changes or image edits. This makes it possible to train the system from existing real-world datasets with no additional annotation effort. Helisa Dhamo, Azade Farshad, Iro Laina, Nassir Navab, Gregory D. Hager, Federico Tombari, Christian Rupprecht 0001 |
CVPR | 5 |
| 2020 | Extremely Dense Point Correspondences Using a Learned Feature DescriptorabstractHigh-quality 3D reconstructions from endoscopy video play an important role in many clinical applications, including surgical navigation where they enable direct video-CT registration. While many methods exist for general multi-view 3D reconstruction, these methods often fail to deliver satisfactory performance on endoscopic video. Part of the reason is that local descriptors that establish pair-wise point correspondences, and thus drive reconstruction, struggle when confronted with the texture-scarce surface of anatomy. Learning-based dense descriptors usually have larger receptive fields enabling the encoding of global information, which can be used to disambiguate matches. In this work, we present an effective self-supervised training scheme and novel loss design for dense descriptor learning. In direct comparison to recent local and dense descriptors on an in-house sinus endoscopy dataset, we demonstrate that our proposed dense descriptor can generalize to unseen patients and scopes, thereby largely improving the performance of Structure from Motion (SfM) in terms of model density and completeness. We also evaluate our method on a public dense optical flow dataset and a small-scale SfM public dataset to further demonstrate the effectiveness and generality of our method. The source code is available at https://github.com/lppllppl920/DenseDescriptorLearning-Pytorch. Xingtong Liu, Yiping Zheng, Benjamin Killeen, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
CVPR | 5 |
| 2020 | Learning From Synthetic AnimalsabstractDespite great success in human parsing, progress for parsing other deformable articulated objects, like animals, is still limited by the lack of labeled data. In this paper, we use synthetic images and ground truth generated from CAD animal models to address this challenge. To bridge the domain gap between real and synthetic images, we propose a novel consistency-constrained semi-supervised learning method (CC-SSL). Our method leverages both spatial and temporal consistencies, to bootstrap weak models trained on synthetic data with unlabeled real images. We demonstrate the effectiveness of our method on highly deformable animals, such as horses and tigers. Without using any real image label, our method allows for accurate keypoint prediction on real images. Moreover, we quantitatively show that models using synthetic data achieve better generalization performance than models trained on real images across different domains in the Visual Domain Adaptation Challenge dataset. Our synthetic dataset contains 10+ animals with diverse poses and rich ground truth, which enables us to use the multi-task learning strategy to further boost models' performance. Jiteng Mu, Weichao Qiu, Gregory D. Hager, Alan L. Yuille |
CVPR | 3 |
| 2020 | Anatomy-Aware Siamese Network: Exploiting Semantic Asymmetry for Accurate Pelvic Fracture Detection in X-Ray Images
Haomin Chen, Yirui Wang 0002, Weijian Li 0001, Chi-Tung Chang, Adam P. Harrison, Jing Xiao 0006, Gregory D. Hager, Le Lu 0001, Chien-Hung Liao, Shun Miao |
ECCV (23) | 8 |
| 2020 | Intent-Aware Pedestrian Prediction for Adaptive Crowd NavigationabstractMobile robots capable of navigating seamlessly and safely in pedestrian rich environments promise to bring robotic assistance closer to our daily lives. In this paper we draw on insights of how humans move in crowded spaces to explore how to recognize pedestrian navigation intent, how to predict pedestrian motion and how a robot may adapt its navigation policy dynamically when facing unexpected human movements. Our approach is to develop algorithms that replicate this behavior. We experimentally demonstrate the effectiveness of our prediction algorithm using real-world pedestrian datasets and achieve comparable or better prediction accuracy compared to several state-of-the-art approaches. Moreover, we show that confidence of pedestrian prediction can be used to adjust the risk of a navigation policy adaptively to afford the most comfortable level as measured by the frequency of personal space violation in comparison with baselines. Furthermore, our adaptive navigation policy is able to reduce the number of collisions by 43% in the presence of novel pedestrian motion not seen during training. Kapil D. Katyal, Gregory D. Hager, Chien-Ming Huang 0001 |
ICRA | 2 |
| 2020 | Autonomously Navigating a Surgical Tool Inside the Eye by Learning from DemonstrationabstractA fundamental challenge in retinal surgery is safely navigating a surgical tool to a desired goal position on the retinal surface while avoiding damage to surrounding tissues, a procedure that typically requires tens-of-microns accuracy. In practice, the surgeon relies on depth-estimation skills to localize the tool-tip with respect to the retina in order to perform the tool-navigation task, which can be prone to human error. To alleviate such uncertainty, prior work has introduced ways to assist the surgeon by estimating the tooltip distance to the retina and providing haptic or auditory feedback. However, automating the tool-navigation task itself remains unsolved and largely unexplored. Such a capability, if reliably automated, could serve as a building block to streamline complex procedures and reduce the chance for tissue damage. Towards this end, we propose to automate the tool-navigation task by learning to mimic expert demonstrations of the task. Specifically, a deep network is trained to imitate expert trajectories toward various locations on the retina based on recorded visual servoing to a given goal specified by the user. The proposed autonomous navigation system is evaluated in simulation and in physical experiments using a silicone eye phantom. We show that the network can reliably navigate a needle surgical tool to various desired locations within 137 μm accuracy in physical experiments and 94 μm in simulation on average, and generalizes well to unseen situations such as in the presence of auxiliary surgical tools, variable eye backgrounds, and brightness conditions. Ji Woong Kim, Changyan He, Müller G. Urias, Peter Gehlbach, Gregory D. Hager, Iulian Iordachita, Marin Kobilarov |
ICRA | 5 |
| 2020 | Reconstructing Sinus Anatomy from Endoscopic Video - Towards a Radiation-Free Approach for Quantitative Longitudinal Assessment
Xingtong Liu, Maia Stiber, Jindan Huang, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
MICCAI (3) | 5 |
| 2020 | Deep hiearchical multi-label classification applied to chest X-ray abnormality taxonomies
Haomin Chen, Shun Miao, Daguang Xu, Gregory D. Hager, Adam P. Harrison |
Medical Image Anal. | 4 |
| 2020 | Dense Depth Estimation in Monocular Endoscopy With Self-Supervised Learning MethodsabstractWe present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or shading. Our method only requires monocular endoscopic videos and a multi-view stereo method, e.g., structure from motion, to supervise learning in a sparse manner. Consequently, our method requires neither manual labeling nor patient computed tomography (CT) scan in the training and application phases. In a cross-patient experiment using CT scans as groundtruth, the proposed method achieved submillimeter mean residual error. In a comparison study to recent self-supervised depth estimation methods designed for natural video on in vivo sinus endoscopy data, we demonstrate that the proposed approach outperforms the previous methods by a large margin. The source code for this work is publicly available online at https://github.com/lppllppl920/EndoscopyDepthEstimation-Pytorch. Xingtong Liu, Ayushi Sinha, Masaru Ishii, Gregory D. Hager, Austin Reiter, Russell H. Taylor, Mathias Unberath |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Uncertainty-Aware Occupancy Map Prediction Using Generative Networks for Robot NavigationabstractEfficient exploration through unknown environments remains a challenging problem for robotic systems. In these situations, the robot's ability to reason about its future motion is often severely limited by sensor field of view (FOV). By contrast, biological systems routinely make decisions by taking into consideration what might exist beyond their FOV based on prior experience. We present an approach for predicting occupancy map representations of sensor data for future robot motions using deep neural networks. We develop a custom loss function used to make accurate prediction while emphasizing physical boundaries. We further study extensions to our neural network architecture to account for uncertainty and ambiguity inherent in mapping and exploration. Finally, we demonstrate a combined map prediction and information-theoretic exploration strategy using the variance of the generated hypotheses as the heuristic for efficient exploration of unknown environments. Kapil D. Katyal, Katie M. Popek, Chris Paxton 0001, Philippe Burlina, Gregory D. Hager |
ICRA | 5 |
| 2019 | Visual Robot Task PlanningabstractProspection is key to solving challenging problems in new environments, but it has not been deeply explored as applied to task planning for perception-driven robotics. We propose visual robot task planning, where we take in an input image and must generate a sequence of high-level actions and associated observations that achieve some task. In this paper, we describe a neural network architecture and associated planning algorithm that (1) learns a representation of the world that can generate prospective futures, (2) uses this generative model to simulate the result of sequences of high-level actions in a variety of environments, and (3) evaluates these actions via a variant of Monte Carlo Tree Search to find a viable solution to a particular problem. Our approach allows us to visualize intermediate motion goals and learn to plan complex activity from visual information, and used this to generate and visualize task plans on held-out examples of a block-stacking simulation. Chris Paxton 0001, Yotam Barnoy, Kapil D. Katyal, Raman Arora, Gregory D. Hager |
ICRA | 5 |
| 2019 | The CoSTAR Block Stacking Dataset: Learning with Workspace ConstraintsabstractA robot can now grasp an object more effectively than ever before, but once it has the object what happens next? We show that a mild relaxation of the task and workspace constraints implicit in existing object grasping datasets can cause neural network based grasping algorithms to fail on even a simple block stacking task when executed under more realistic circumstances. To address this, we introduce the JHU CoSTAR Block Stacking Dataset (BSD), where a robot interacts with 5.1 cm colored blocks to complete an order-fulfillment style block stacking task. It contains dynamic scenes and real time-series data in a less constrained environment than comparable datasets. There are nearly 12,000 stacking attempts and over 2 million frames of real data. We discuss the ways in which this dataset provides a valuable resource for a broad range of other topics of investigation. We find that hand-designed neural networks that work on prior datasets do not generalize to this task. Thus, to establish a baseline for this dataset, we demonstrate an automated search of neural network based models using a novel multiple-input HyperTree MetaModel, and find a final model which makes reasonable 3D pose predictions for grasping and stacking on our dataset. The CoSTAR BSD, code, and instructions are available at sites.google.com/site/costardataset. Andrew Hundt, Varun Jain, Chris Paxton 0001, Gregory D. Hager |
IROS | 5 |
| 2019 | Automated Surgical Activity Recognition with One Labeled Sequence
Robert S. DiPietro, Gregory D. Hager |
MICCAI (5) | 2 |
| 2019 | Semantic Stereo for Incidental Satellite ImagesabstractThe increasingly common use of incidental satellite images for stereo reconstruction versus rigidly tasked binocular or trinocular coincident collection is helping to enable timely global-scale 3D mapping; however, reliable stereo correspondence from multi-date image pairs remains very challenging due to seasonal appearance differences and scene change. Promising recent work suggests that semantic scene segmentation can provide a robust regularizing prior for resolving ambiguities in stereo correspondence and reconstruction problems. To enable research for pairwise semantic stereo and multi-view semantic 3D reconstruction with incidental satellite images, we have established a large-scale public dataset including multi-view, multi-band satellite images and ground truth geometric and semantic labels for two large cities. To demonstrate the complementary nature of the stereo and segmentation tasks, we present lightweight public baselines adapted from recent state of the art convolutional neural network models and assess their performance. Marc Bosch, Kevin Foster, Gordon A. Christie, Sean Wang 0001, Gregory D. Hager, Myron Z. Brown |
WACV | 5 |
| 2019 | Toward Computer Vision Systems That Understand Real-World Assembly ProcessesabstractMany applications of computer vision require robust systems that can parse complex structures as they evolve in time. Using a block construction task as a case study, we illustrate the main components involved in building such systems. We evaluate performance at three increasingly-detailed levels of spatial granularity on two multimodal (RGBD + IMU) datasets. On the first, designed to match the assumptions of the model, we report better than 90% accuracy at the finest level of granularity. On the second, designed to test the robustness of our model under adverse, real-world conditions, we report 67% accuracy and 91% precision at the mid-level of granularity. We show that this seemingly simple process presents many opportunities to expand the frontiers of computer vision and action recognition. Jonathan D. Jones, Gregory D. Hager, Sanjeev Khudanpur |
WACV | 2 |
| 2019 | The deformable most-likely-point paradigm
Ayushi Sinha, Seth Billings, Austin Reiter, Xingtong Liu, Masaru Ishii, Gregory D. Hager, Russell H. Taylor |
Medical Image Anal. | 6 |
| 2019 | Deep Supervision with Intermediate ConceptsabstractRecent data-driven approaches to scene interpretation predominantly pose inference as an end-to-end black-box mapping, commonly performed by a Convolutional Neural Network (CNN). However, decades of work on perceptual organization in both human and machine vision suggest that there are often intermediate representations that are intrinsic to an inference task, and which provide essential structure to improve generalization. In this work, we explore an approach for injecting prior domain structure into neural network training by supervising hidden layers of a CNN with intermediate concepts that normally are not observed in practice. We formulate a probabilistic framework which formalizes these notions and predicts improved generalization via this deep supervision method. One advantage of this approach is that we are able to train only from synthetic CAD renderings of cluttered scenes, where concept values can be extracted, but apply the results to real images. Our implementation achieves the state-of-the-art performance of 2D/3D keypoint localization and image classification on real image benchmarks including KITTI, PASCAL VOC, PASCAL3D+, IKEA, and CIFAR100. We provide additional evidence that our approach outperforms alternative forms of supervision, such as multi-task networks. M. Zeeshan Zia, Quoc-Huy Tran, Xiang Yu 0002, Gregory D. Hager, Manmohan Krishna Chandraker |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Constraints and Development in Children's Block Construction
Cathryn S. Cortesa, Jonathan D. Jones, Gregory D. Hager, Sanjeev Khudanpur, Barbara Landau, Amy Lynne Shelton |
CogSci | 3 |
| 2018 | Guide Me: Interacting With Deep NetworksabstractInteraction and collaboration between humans and intelligent machines has become increasingly important as machine learning methods move into real-world applications that involve end users. While much prior work lies at the intersection of natural language and vision, such as image captioning or image generation from text descriptions, less focus has been placed on the use of language to guide or improve the performance of a learned visual processing algorithm. In this paper, we explore methods to flexibly guide a trained convolutional neural network through user input to improve its performance during inference. We do so by inserting a layer that acts as a spatio-semantic guide into the network. This guide is trained to modify the network's activations, either directly via an energy minimization scheme or indirectly through a recurrent model that translates human language queries to interaction weights. Learning the verbal interaction is fully automatic and does not require manual text annotations. We evaluate the method on two datasets, showing that guiding a pre-trained network can improve performance, and provide extensive insights into the interaction between the guide and the CNN. Christian Rupprecht 0001, Iro Laina, Nassir Navab, Gregory D. Hager, Federico Tombari |
CVPR | 4 |
| 2018 | A Unified Framework for Multi-view Multi-class Object Pose Estimation
Jin Bai 0001, Gregory D. Hager |
ECCV (16) | 3 |
| 2018 | S3D: Stacking Segmental P3D for Action Quality AssessmentabstractAction quality assessment is crucial in areas of sports, surgery and assembly line where action skills can be evaluated. In this paper, we propose the Segment-based P3D-fused network S3D built-upon ED-TCN and push the performance on the UNLV-Dive dataset by a significant margin. We verify that segment-aware training performs better than full-video training which turns out to focus on the water spray. We show that temporal segmentation can be embedded with few efforts. Xiang Xiang 0001, Austin Reiter, Gregory D. Hager, Trac D. Tran |
ICIP | 4 |
| 2018 | Evaluating Methods for End-User Creation of Robot Task PlansabstractHow can we enable users to create effective, perception-driven task plans for collaborative robots? We conducted a 35-person user study with the Behavior Tree-based CoSTAR system to determine which strategies for end user creation of generalizable robot task plans are most usable and effctive. CoSTAR allows domain experts to author complex, perceptually grounded task plans for collaborative robots. As a part of CoSTAR's wide range of capabilities, it allows users to specify SmartMoves: abstract goals such as “pick up component A from the right side of the table.” Users were asked to perform pick-and-place assembly tasks with either SmartMoves or one of three simpler baseline versions of CoSTAR. Overall, participants found CoSTAR to be highly usable, with an average System Usability Scale score of 73.4 out of 100. SmartMove also helped users perform tasks faster and more effectively; all SmartMove users completed the first two tasks, while not all users completed the tasks using the other strategies. SmartMove users showed better performance for incorporating perception across all three tasks. Chris Paxton 0001, Felix Jonathan, Andrew Hundt, Bilge Mutlu, Gregory D. Hager |
IROS | 5 |
| 2018 | Unsupervised Learning for Surgical Motion by Learning to Predict the Future
Robert S. DiPietro, Gregory D. Hager |
MICCAI (4) | 2 |
| 2018 | Endoscopic Navigation in the Absence of CT Imaging
Ayushi Sinha, Xingtong Liu, Austin Reiter, Masaru Ishii, Gregory D. Hager, Russell H. Taylor |
MICCAI (4) | 5 |
| 2018 | Evaluation and Stability Analysis of Video-Based Navigation System for Functional Endoscopic Sinus Surgery on In Vivo Clinical DataabstractFunctional endoscopic sinus surgery (FESS) is one of the most common outpatient surgical procedures performed in the head and neck region. It is used to treat chronic sinusitis, a disease characterized by inflammation in the nose and surrounding paranasal sinuses, affecting about 15% of the adult population. During FESS, the nasal cavity is visualized using an endoscope, and instruments are used to remove tissues that are often within a millimeter of critical anatomical structures, such as the optic nerve, carotid arteries, and nasolacrimal ducts. To maintain orientation and to minimize the risk of damage to these structures, surgeons use surgical navigation systems to visualize the 3-D position of their tools on patients' preoperative Computed Tomographies (CTs). This paper presents an image-based method for enhanced endoscopic navigation. The main contributions are: (1) a system that enables a surgeon to asynchronously register a sequence of endoscopic images to a CT scan with higher accuracy than other reported solutions using no additional hardware; (2) the ability to report the robustness of the registration; and (3) evaluation on in vivo human data. The system also enables the overlay of anatomical structures, visible, or occluded, on top of video images. The methods are validated on four different data sets using multiple evaluation metrics. First, for experiments on synthetic data, we observe a mean absolute position error of 0.21mm and a mean absolute orientation error of 2.8° compared with ground truth. Second, for phantom data, we observe a mean absolute position error of 0.97mm and a mean absolute orientation error of 3.6° compared with the same motion tracked by an electromagnetic tracker. Third, for cadaver data, we use fiducial landmarks and observe an average reprojection distance error of 0.82mm. Finally, for in vivo clinical data, we report an average ICP residual error of 0.88mm in areas that are not composed of erectile tissue and an average ICP residual error of 1.09mm in areas that are composed of erectile tissue. Simon Léonard, Ayushi Sinha, Austin Reiter, Masaru Ishii, Gary L. Gallia, Russell H. Taylor, Gregory D. Hager |
IEEE Trans. Medical Imaging | 7 |
| 2017 | Characterizing spatial construction processes: Toward computational tools to understand cognition
Cathryn S. Cortesa, Jonathan D. Jones, Gregory D. Hager, Sanjeev Khudanpur, Amy Lynne Shelton, Barbara Landau |
CogSci | 3 |
| 2017 | Temporal Convolutional Networks for Action Segmentation and DetectionabstractThe ability to identify and temporally segment fine-grained human actions throughout a video is crucial for robotics, surveillance, education, and beyond. Typical approaches decouple this problem by first extracting local spatiotemporal features from video frames and then feeding them into a temporal classifier that captures high-level temporal patterns. We describe a class of temporal models, which we call Temporal Convolutional Networks (TCNs), that use a hierarchy of temporal convolutions to perform fine-grained action segmentation or detection. Our Encoder-Decoder TCN uses pooling and upsampling to efficiently capture long-range temporal patterns whereas our Dilated TCN uses dilated convolutions. We show that TCNs are capable of capturing action compositions, segment durations, and long-range dependencies, and are over a magnitude faster to train than competing LSTM-based Recurrent Neural Networks. We apply these models to three challenging fine-grained datasets and show large improvements over the state of the art. Colin Lea, Michael D. Flynn, René Vidal, Austin Reiter, Gregory D. Hager |
CVPR | 5 |
| 2017 | Deep Supervision with Shape Concepts for Occlusion-Aware 3D Object ParsingabstractMonocular 3D object parsing is highly desirable in various scenarios including occlusion reasoning and holistic scene interpretation. We present a deep convolutional neural network (CNN) architecture to localize semantic parts in 2D image and 3D space while inferring their visibility states, given a single RGB image. Our key insight is to exploit domain knowledge to regularize the network by deeply supervising its hidden layers, in order to sequentially infer intermediate concepts associated with the final task. To acquire training data in desired quantities with ground truth 3D shape and relevant concepts, we render 3D object CAD models to generate large-scale synthetic data and simulate challenging occlusion configurations between objects. We train the network only on synthetic data and demonstrate state-of-the-art performances on real image benchmarks including an extended version of KITTI, PASCAL VOC, PASCAL3D+ and IKEA for 2D and 3D keypoint localization and instance segmentation. The empirical results substantiate the utility of our deep supervision scheme by demonstrating effective transfer of knowledge from synthetic data to real images, resulting in less overfitting compared to standard end-to-end training. M. Zeeshan Zia, Quoc-Huy Tran, Xiang Yu 0002, Gregory D. Hager, Manmohan Krishna Chandraker |
CVPR | 5 |
| 2017 | Regularizing face verification nets for pain intensity regressionabstractLimited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille |
ICIP | 6 |
| 2017 | CoSTAR: Instructing collaborative robots with behavior trees and visionabstractFor collaborative robots to become useful, end users who are not robotics experts must be able to instruct them to perform a variety of tasks. With this goal in mind, we developed a system for end-user creation of robust task plans with a broad range of capabilities. CoSTAR: the Collaborative System for Task Automation and Recognition is our winning entry in the 2016 KUKA Innovation Award competition at the Hannover Messe trade show, which this year focused on Flexible Manufacturing. CoSTAR is unique in how it creates natural abstractions that use perception to represent the world in a way users can both understand and utilize to author capable and robust task plans. Our Behavior Tree-based task editor integrates high-level information from known object segmentation and pose estimation with spatial reasoning and robot actions to create robust task plans. We describe the cross-platform design and implementation of this system on multiple industrial robots and evaluate its suitability for a wide variety of use cases. Chris Paxton 0001, Andrew Hundt, Felix Jonathan, Kelleher Guerin, Gregory D. Hager |
ICRA | 5 |
| 2017 | Combining neural networks and tree search for task and motion planning in challenging environmentsabstractTask and motion planning subject to Linear Temporal Logic (LTL) specifications in complex, dynamic environments requires efficient exploration of many possible future worlds. Model-free reinforcement learning has proven successful in a number of challenging tasks, but shows poor performance on tasks that require long-term planning. In this work, we integrate Monte Carlo Tree Search with hierarchical neural net policies trained on expressive LTL specifications. We use reinforcement learning to find deep neural networks representing both low-level control policies and task-level “option policies” that achieve high-level goals. Our combined architecture generates safe and responsive motion plans that respect the LTL constraints. We demonstrate our approach in a simulated autonomous driving setting, where a vehicle must drive down a road in traffic, avoid collisions, and navigate an intersection, all while obeying rules of the road. Chris Paxton 0001, Vasumathi Raman, Gregory D. Hager, Marin Kobilarov |
IROS | 3 |
| 2016 | Segmental Spatiotemporal CNNs for Fine-Grained Action Segmentation
Colin Lea, Austin Reiter, René Vidal, Gregory D. Hager |
ECCV (3) | 4 |
| 2016 | Semi-Autonomous Telerobotic Assembly over High-Latency NetworksabstractWe report the development and a preliminary multi-user evaluation of an assisted teleoperation architecture for assembly tasks over high-latency networks. While such tasks still require human insight, there are often elements of these tasks which can be executed autonomously. Our architecture contextualizes a user's intended actions in a remote scene. It assists the user by performing the intended action more precisely based on the latest local scene model. We report the results of a multi-user study of nine participants to evaluate performance of this approach using a full-3D dynamic robotic simulation in a laboratory environment. The study required users to assemble part of a structure with 200 or 4000 milliseconds of time delay, with or without assistance from the reported system. For each condition, we evaluated performance based both on partial and final task completion times, and we estimated each user's workload with the NASA Task Load Index. The study showed that our architecture has the potential to increase feasibility of the task both with and without large telemetry delay and to enable users to complete elements of the task more rapidly. We also found that the assistance dramatically reduced the users' workload. We provide open-source implementations for all components of our system, which can be adapted to other robotic platforms. Jonathan Bohren, Chris Paxton 0001, Ryan Howarth, Gregory D. Hager, Louis L. Whitcomb |
HRI | 4 |
| 2016 | Unsupervised surgical data alignment with application to automatic activity annotationabstractRobotic surgery and other minimally-invasive surgical techniques are an integral part of patient care, and readily yield large amounts of data. Surgical tool motion (kinematic data) contains information that is useful for assessment and education. Typically, assessment and education tools that rely upon the kinematic data require substantial manual processing such as activity annotations. The goal of this paper was to develop an automated method to align surgical recordings and assign activity annotations. We developed an approach based on unsupervised alignment to efficient annotate kinematic data for its constituent activity segments. Our method includes extracting non-linear features from the kinematic data using a stacked de-noising autoencoder, and using modified dynamic time warping to align the kinematic data from different trials of the study task. We combined alignment between a test and one or a small set of template trials (with prior manual annotations) with voting based on kernel density estimation to transfer labels from the template to the test trial. Our experiments on performance of this method using two datasets captured in the training laboratory demonstrate an accuracy of 72% to 94% for annotating activity segments within a surgical training task. Our findings are robust to data captured from several surgeons, and to deviations in activity from a canonical activity sequence. S. Swaroop Vedula, Gyusung I. Lee, Mija R. Lee, Sanjeev Khudanpur, Gregory D. Hager |
ICRA | 6 |
| 2016 | Learning convolutional action primitives for fine-grained action recognitionabstractFine-grained action recognition is important for many applications of human-robot interaction, automated skill assessment, and surveillance. The goal is to segment and classify all actions occurring in a time series sequence. While recent recognition methods have shown strong performance in robotics applications, they often require hand-crafted features, use large amounts of domain knowledge, or employ overly simplistic representations of how objects change throughout an action. In this paper we present the Latent Convolutional Skip Chain Conditional Random Field (LC-SC-CRF). This time series model learns a set of interpretable and composable action primitives from sensor data. We apply our model to cooking tasks using accelerometer data from the University of Dundee 50 Salads dataset and to robotic surgery training tasks using robot kinematic data from the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS). Our performance on 50 Salads and JIGSAWS are 18.0% and 5.3% higher than the state of the art, respectively. This model performs well without requiring hand-crafted features or intricate domain knowledge. The code and features have been made public. Colin Lea, René Vidal, Gregory D. Hager |
ICRA | 3 |
| 2016 | Hierarchical semantic parsing for object pose estimation in densely cluttered scenesabstractDensely cluttered scenes are composed of multiple objects which are in close contact and heavily occlude each other. Few existing 3D object recognition systems are capable of accurately predicting object poses in such scenarios. This is mainly due to the presence of objects with textureless surfaces, similar appearances and the difficulty of object instance segmentation. In this paper, we present a hierarchical semantic segmentation algorithm which partitions a densely cluttered scene into different object regions. A RANSAC-based registration method is subsequently applied to estimate 6-DoF object poses within each object class. Part of this algorithm includes a generalized pooling scheme used to construct robust and discriminative object representations from a convolutional architecture with multiple pooling domains. We also provide a new RGB-D dataset which serves as a benchmark for object pose estimation in densely cluttered scenes. This dataset contains five thousand scene frames and over twenty thousand labeled poses of ten common hand tools. We show that our method demonstrates improved performance of pose estimation on this new dataset compared with other state-of-the-art methods. Jonathan Bohren, Eric Carlson, Gregory D. Hager |
ICRA | 4 |
| 2016 | Incremental scene understanding on dense SLAMabstractWe present an architecture for online, incremental scene modeling which combines a SLAM-based scene understanding framework with semantic segmentation and object pose estimation. The core of this approach comprises a probabilistic inference scheme that predicts semantic labels for object hypotheses at each new frame. From these hypotheses, recognized scene structures are incrementally constructed and tracked. Semantic labels are inferred using a multi-domain convolutional architecture which operates on the image time series and which enables efficient propagation of features as well as robust model registration. To evaluate this architecture, we introduce a large-scale RGB-D dataset JHUSEQ-25 as a new benchmark for the sequence-based scene understanding in complex and densely cluttered scenes. This dataset contains 25 RGB-D video sequences with 100,000 labeled frames in total. We validate our method on this dataset and demonstrate improved performance of semantic segmentation and 6-DoF object pose estimation compared with methods based on the single view. Keisuke Tateno, Federico Tombari, Nassir Navab, Gregory D. Hager |
IROS | 6 |
| 2016 | Do what i want, not what i did: Imitation of skills by planning sequences of actionsabstractWe propose a learning-from-demonstration approach for grounding actions from expert data and an algorithm for using these actions to perform a task in new environments. Our approach is based on an application of sampling-based motion planning to search through the tree of discrete, high-level actions constructed from a symbolic representation of a task. Recursive sampling-based planning is used to explore the space of possible continuous-space instantiations of these actions. We demonstrate the utility of our approach with a magnetic structure assembly task, showing that the robot can intelligently select a sequence of actions in different parts of the workspace and in the presence of obstacles. This approach can better adapt to new environments by selecting the correct high-level actions for the particular environment while taking human preferences into account. Chris Paxton 0001, Felix Jonathan, Marin Kobilarov, Gregory D. Hager |
IROS | 4 |
| 2016 | Sensor substitution for video-based action recognitionabstractThere are many applications where domain-specific sensing, such as accelerometers, kinematics, or force sensing, provide unique and important information for control or for analysis of motion. However, it is not always the case that these sensors can be deployed or accessed beyond laboratory environments. For example, it is possible to instrument humans or robots to measure motion in the laboratory in ways that it is not possible to replicate in the wild. An alternative, which we explore in this paper, is to address situations where accurate sensing is available while training an algorithm, but for which only video is available for deployment. We present two examples of this sensory substitution methodology. The first variation trains a convolutional neural network to regress real-valued signals, including robot end-effector pose, from video. The second example regresses binary signals derived from accelerometer data which signifies when specific objects are in motion. We evaluate these on the JIGSAWS dataset for robotic surgery training assessment and the 50 Salads dataset for modeling complex structured cooking tasks. We evaluate the trained models for video-based action recognition and show that the trained models provide information that is comparable to the sensory signals they replace. Christian Rupprecht 0001, Colin Lea, Federico Tombari, Nassir Navab, Gregory D. Hager |
IROS | 5 |
| 2016 | Anatomically Constrained Video-CT Registration via the V-IMLOP AlgorithmabstractFunctional endoscopic sinus surgery (FESS) is a surgical procedure used to treat acute cases of sinusitis and other sinus diseases. FESS is fast becoming the preferred choice of treatment due to its minimally invasive nature. However, due to the limited field of view of the endoscope, surgeons rely on navigation systems to guide them within the nasal cavity. State of the art navigation systems report registration accuracy of over 1mm, which is large compared to the size of the nasal airways. We present an anatomically constrained video-CT registration algorithm that incorporates multiple video features. Our algorithm is robust in the presence of outliers. We also test our algorithm on simulated and in-vivo data, and test its accuracy against degrading initializations. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Seth Billings, Ayushi Sinha, Austin Reiter, Simon Léonard, Masaru Ishii, Gregory D. Hager, Russell H. Taylor |
MICCAI (3) | 6 |
| 2016 | Recognizing Surgical Activities with Recurrent Neural Networks
Robert S. DiPietro, Colin Lea, Anand Malpani, Narges Ahmidi, S. Swaroop Vedula, Gyusung I. Lee, Mija R. Lee, Gregory D. Hager |
MICCAI (1) | 8 |
| 2016 | Guest Editorial: Special Section on CVPR 2013abstractThis special section contains selected papers from the IEEE Computer Vision and Pattern Recognition (CVPR), June, 2013, jointly sponsored by the IEEE and the Computer Vision Foundation. William T. Freeman, Richard Szeliski, Gregory D. Hager |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | VESICLE: Volumetric Evaluation of Synaptic Inferfaces using Computer Vision at Large Scale
William R. Gray Roncal, Michael J. Pekala, Verena Kaynig, Dean Kleissas, Joshua T. Vogelstein, Hanspeter Pfister, Randal C. Burns, R. Jacob Vogelstein, Mark A. Chevillet, Gregory D. Hager |
BMVC | 10 |
| 2015 | Beyond spatial pooling: Fine-grained representation learning in multiple domainsabstractObject recognition systems have shown great progress over recent years. However, creating object representations that are robust to changes in viewpoint while capturing local visual details continues to be a challenge. In particular, recent convolutional architectures employ spatial pooling to achieve scale and shift invariances, but they are still sensitive to out-of-plane rotations. In this paper, we formulate a probabilistic framework for analyzing the performance of pooling. This framework suggests two directions for improvement. First, we apply multiple scales of filters coupled with different pooling granularities, and second we make use of color as an additional pooling domain, thereby reducing the sensitivity to spatial deformations. We evaluate our algorithm on the object instance recognition task using two independent publicly available RGB-D datasets, and demonstrate significant improvements over the current state-of-the-art. In addition, we present a new dataset for industrial objects to further validate the effectiveness of our approach versus other state-of-the-art approaches for object recognition using RGB-D data. Austin Reiter, Gregory D. Hager |
CVPR | 3 |
| 2015 | Hierarchical Sparse and Collaborative Low-Rank representation for emotion recognitionabstractIn this paper, we design a Collaborative-Hierarchical Sparse and Low-Rank (C-HiSLR) model that is natural for recognizing human emotion in visual data. Previous attempts require explicit expression components, which are often unavailable and difficult to recover. Instead, our model exploits the low-rank property to subtract neutral faces from expressive facial frames as well as performs sparse representation on the expression components with group sparsity enforced. For the CK+ dataset, C-HiSLR on raw expressive faces performs as competitive as the Sparse Representation based Classification (SRC) applied on manually prepared emotions. Our C-HiSLR performs even better than SRC in terms of true positive rate. Xiang Xiang 0001, Minh Dao, Gregory D. Hager, Trac D. Tran |
ICASSP | 3 |
| 2015 | An incremental approach to learning generalizable robot tasks from human demonstrationabstractDynamic Movement Primitives (DMPs) are a common method for learning a control policy for a task from demonstration. This control policy consists of differential equations that can create a smooth trajectory to a new goal point. However, DMPs only have a limited ability to generalize the demonstration to new environments and solve problems such as obstacle avoidance. Moreover, standard DMP learning does not cope with the noise inherent to human demonstrations. Here, we propose an approach for robot learning from demonstration that can generalize noisy task demonstrations to a new goal point and to an environment with obstacles. This strategy for robot learning from demonstration results in a control policy that incorporates different types of learning from demonstration, which correspond to different types of observational learning as outlined in developmental psychology. Amir M. Ghalamzan E., Chris Paxton 0001, Gregory D. Hager, Luca Bascetta |
ICRA | 3 |
| 2015 | A framework for end-user instruction of a robot assistant for manufacturingabstractSmall Manufacturing Entities (SMEs) have not incorporated robotic automation as readily as large companies due to rapidly changing product lines, complex and dexterous tasks, and the high cost of start-up. While recent low-cost robots such as the Universal Robots UR5 and Rethink Robotics Baxter are more economical and feature improved programming interfaces, based on our discussions with manufacturers further incorporation of robots into the manufacturing work flow is limited by the ability of these systems to generalize across tasks and handle environmental variation. Our goal is to create a system designed for small manufacturers that contains a set of capabilities useful for a wide range of tasks, is both powerful and easy to use, allows for perceptually grounded actions, and is able to accumulate, abstract, and reuse plans that have been taught. We present an extension to Behavior Trees that allows for representing the system capabilities of a robot as a set of generalizable operations that are exposed to an end-user for creating task plans. We implement this framework in CoSTAR, the Collaborative System for Task Automation and Recognition, and demonstrate its effectiveness with two case studies. We first perform a complex tool-based object manipulation task in a laboratory setting. We then show the deployment of our system in an SME where we automate a machine tending task that was not possible with current off the shelf robots. Kelleher Guerin, Colin Lea, Chris Paxton 0001, Gregory D. Hager |
ICRA | 4 |
| 2015 | Transition State Clustering: Unsupervised Surgical Trajectory Segmentation for Robot Learning
Sanjay Krishnan, Animesh Garg, Sachin Patil, Colin Lea, Gregory D. Hager, Pieter Abbeel, Kenneth Y. Goldberg |
ISRR (2) | 5 |
| 2015 | Bridging the Robot Perception Gap with Mid-Level Vision
Jonathan Bohren, Gregory D. Hager |
ISRR (2) | 3 |
| 2015 | An Improved Model for Segmentation and Recognition of Fine-Grained Activities with Application to Surgical Training TasksabstractAutomated segmentation and recognition of fine-grained activities is important for enabling new applications in industrial automation, human-robot collaboration, and surgical training. Many existing approaches to activity recognition assume that a video has already been segmented and perform classification using an abstract representation based on spatio-temporal features. While some approaches perform joint activity segmentation and recognition, they typically suffer from a poor modeling of the transitions between actions and a representation that does not incorporate contextual information about the scene. In this paper, we propose a model for action segmentation and recognition that improves upon existing work in two directions. First, we develop a variation of the Skip-Chain Conditional Random Field that captures long-range state transitions between actions by using higher-order temporal relationships. Second, we argue that in constrained environments, where the relevant set of objects is known, it is better to develop features using high-level object relationships that have semantic meaning instead of relying on abstract features. We apply our approach to a set of tasks common for training in robotic surgery: suturing, knot tying, and needle passing, and show that our method increases micro and macro accuracy by 18.46% and 44.13% relative to the state of the art on a widely used robotic surgery dataset. Colin Lea, Gregory D. Hager, René Vidal |
WACV | 2 |
| 2014 | Adjutant: A framework for flexible human-machine collaborative systemsabstractFlexible interaction and instruction is a key enabling technology for expanding robotics into small to medium scale manufacturing, in-home assistance for physically disabled individuals, and robotic surgery. In these cases, performing a task manually is neither practical nor scalable, yet complete automation is cost-prohibitive or impossible. Thus, our interest is in collaborative systems that can be easily trained to work with an operator. This collaborative robotic system should be instructable in a generalizable way for a wide range of tasks, and should generalize to new tasks gracefully with minimal retraining. At the same time, for a given task, the system should take advantage of user interaction modalities needed to accomplish the task, subject to the constraints of the available interfaces. These ideas motivate the Adjutant framework. Adjutant supports human-robot collaborative operations for ranges of user roles and robot capability. Adjutant models human-robot systems via sets of robot capabilities, composable high-level functions that can be specialized to specific tasks, and collaborative behaviors which relate these capabilities to specific user interfaces or interaction paradigms. Adjutant also incorporates several methods encapsulating reusable task information into capabilities, thus specializing them, including tool affordances, perceptual grounding templates, and tool movement primitives. We have implemented Adjutant as a software framework in ROS and, in this paper, explore the utility of Adjutant for performing several real-world collaborative manufacturing tasks on an industrial robot test-bed. Kelleher Guerin, Sebastian Riedel 0002, Jonathan Bohren, Gregory D. Hager |
IROS | 4 |
| 2014 | Needle Guidance Using Handheld Stereo Vision and Projection for Ultrasound-Based Interventions
Philipp J. Stolka, Pezhman Foroughi, Matthew Rendina, Clifford R. Weiss, Gregory D. Hager, Emad Boctor |
MICCAI (2) | 5 |
| 2014 | Ultrasound elastography using multiple images
Hassan Rivaz, Emad Boctor, Michael A. Choti, Gregory D. Hager |
Medical Image Anal. | 4 |
| 2014 | Fundus Image Mosaicking for Information Augmentation in Computer-Assisted Slit-Lamp ImagingabstractLaser photocoagulation is currently the standard treatment for sight-threatening diseases worldwide, namely diabetic retinopathy and retinal vein occlusions. The slit lamp biomicroscope is the most commonly used device for this procedure, specially for the treatment of the eye periphery. However, only a small portion of the retina can be visualized through the biomicroscope, complicating the task of localizing and identifying surgical targets, increasing treatment duration and patient discomfort. In order to assist surgeons, we propose a method for creating intraoperative retina maps for view expansion using a slit-lamp device. Based on the mosaicking method described by Richa et al, 2012, the proposed method is a combination of direct and feature-based methods, suitable for the textured nature of the human retina. In this paper, we describe three major enhancements to the original formulation. The first is a visual tracking method using local illumination compensation to cope with the challenging visualization conditions. The second is an efficient pixel selection scheme for increased computational efficiency. The third is an entropy-based mosaic update method to dynamically improve the retina map during exploration. To evaluate the performance of the proposed method, we conducted several experiments on human subjects with a computer-assisted slit-lamp prototype. We also demonstrate the practical value of the system for photo documentation, diagnosis and intraoperative navigation. Rogério Richa, Rodrigo Linhares, Eros Comunello, Aldo von Wangenheim, Jean-Yves Schnitzler, Benjamin Wassmer, Claire Guillemot, Gilles Thuret, Philippe Gain, Gregory D. Hager, Russell H. Taylor |
IEEE Trans. Medical Imaging | 10 |
| 2013 | A pilot study in vision-based augmented telemanipulation for remote assembly over high-latency networksabstractIn this paper we present an approach to extending the capabilities of telemanipulation systems by intelligently augmenting a human operator's motion commands based on quantitative three-dimensional scene perception at the remote telemanipulation site. This framework is the first prototype of the Augmented Shared-Control for Efficient, Natural Telemanipulation (ASCENT) System. ASCENT aims to enable new robotic applications in environments where task complexity precludes autonomous execution or where low-bandwidth and/or high-latency communication channels exist between the nearest human operator and the application site. These constraints can constrain the domain of telemanipulation to simple or static environments, reduce the effectiveness of telemanipulation, and even preclude remote intervention entirely. ASCENT is a semi-autonomous framework that increases the speed and accuracy of a human operator's actions via seamless transitions between one-to-one teleoperation and autonomous interventions. We report the promising results of a pilot study validating ASCENT in a transatlantic telemanipulation experiment between The Johns Hopkins University in Baltimore, MD, USA and the German Aerospace Center (DLR) in Oberpfaffenhofen, Germany. In these experiments, we observed average telemetry delays of 200ms, and average video delays of 2s with peaks of up to 6s for all data. We also observed 75% frame loss for video streams due to bandwidth limits, giving 4fps video. Jonathan Bohren, Chavdar Papazov, Darius Burschka, Kai Krieger, Sven Parusel, Sami Haddadin, William L. Shepherdson, Gregory D. Hager, Louis L. Whitcomb |
ICRA | 8 |
| 2013 | String Motif-Based Description of Tool Motion for Detecting Skill and Gestures in Robotic Surgery
Narges Ahmidi, Benjamín Béjar Haro, S. Swaroop Vedula, Sanjeev Khudanpur, René Vidal, Gregory D. Hager |
MICCAI (1) | 7 |
| 2013 | Surgical Gesture Segmentation and Recognition
Lingling Tao, Luca Zappella, Gregory D. Hager, René Vidal |
MICCAI (3) | 3 |
| 2013 | Dynamic Template Tracking and Recognition
Rizwan Chaudhry, Gregory D. Hager, René Vidal |
Int. J. Comput. Vis. | 2 |
| 2013 | Surgical gesture classification from video and kinematic data
Luca Zappella, Benjamín Béjar Haro, Gregory D. Hager, René Vidal |
Medical Image Anal. | 3 |
| 2013 | Unified Detection and Tracking of Instruments during Retinal MicrosurgeryabstractMethods for tracking an object have generally fallen into two groups: tracking by detection and tracking through local optimization. The advantage of detection-based tracking is its ability to deal with target appearance and disappearance, but it does not naturally take advantage of target motion continuity during detection. The advantage of local optimization is efficiency and accuracy, but it requires additional algorithms to initialize tracking when the target is lost. To bridge these two approaches, we propose a framework for unified detection and tracking as a time-series Bayesian estimation problem. The basis of our approach is to treat both detection and tracking as a sequential entropy minimization problem, where the goal is to determine the parameters describing a target in each frame. To do this we integrate the Active Testing (AT) paradigm with Bayesian filtering, and this results in a framework capable of both detecting and tracking robustly in situations where the target object enters and leaves the field of view regularly. We demonstrate our approach on a retinal tool tracking problem and show through extensive experiments that our method provides an efficient and robust tracking solution. Raphael Sznitman, Rogério Richa, Russell H. Taylor, Bruno Jedynak, Gregory D. Hager |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2013 | Evaluation of a System for High-Accuracy 3D Image-Based Registration of Endoscopic Video to C-Arm Cone-Beam CT for Image-Guided Skull Base SurgeryabstractThe safety of endoscopic skull base surgery can be enhanced by accurate navigation in preoperative computed tomography (CT) or, more recently, intraoperative cone-beam CT (CBCT). The ability to register real-time endoscopic video with CBCT offers an additional advantage by rendering information directly within the visual scene to account for intraoperative anatomical change. However, tracker localization error ( ∼ 1-2 mm ) limits the accuracy with which video and tomographic images can be registered. This paper reports the first implementation of image-based video-CBCT registration, conducts a detailed quantitation of the dependence of registration accuracy on system parameters, and demonstrates improvement in registration accuracy achieved by the image-based approach. Performance was evaluated as a function of parameters intrinsic to the image-based approach, including system geometry, CBCT image quality, and computational runtime. Overall system performance was evaluated in a cadaver study simulating transsphenoidal skull base tumor excision. Results demonstrated significant improvement in registration accuracy with a mean reprojection distance error of 1.28 mm for the image-based approach versus 1.82 mm for the conventional tracker-based method. Image-based registration was highly robust against CBCT image quality factors of noise and resolution, permitting integration with low-dose intraoperative CBCT. Daniel Mirota, Ali Uneri, Sebastian Schafer, Sajendra Nithiananthan, Douglas D. Reh, Masaru Ishii, Gary L. Gallia, Russell H. Taylor, Gregory D. Hager, Jeffrey H. Siewerdsen |
IEEE Trans. Medical Imaging | 9 |
| 2012 | Robust Object Tracking in Crowd Dynamic Scenes Using Explicit Stereo Depth
Le Lu 0001, Gregory D. Hager, Jianyu Tang, Hanzi Wang |
ACCV (3) | 3 |
| 2012 | Deformable Tracking of Textured Curvilinear ObjectsabstractThreads and wires are deformable 3-dimensional (3D) and curvilinear objects which are commonly manipulated by humans in various medical and manufacturing tasks. Several applications, including computerassisted evaluation, augmented reality guidance, and autonomous robotic manipulation [2, 3] would benefit from the real-time estimation of the 3D shapes of these deformable objects from images. This estimation is however challenging due to multiple factors: 1) little information is available within an image to visually detect and distinguish a curvilinear object due to its thin and usually uniform appearance; 2) different 3D shapes may lead to the same visual perception, even in a stereo setting in case portions of the objects lie in an epipolar plane; and 3) the motions and deformations can be large, depending on the stiffness of the object. Additionally, a tracking approach that can consistently track specific points along the object defined by their arclength, such as the extremities or midpoint, would be particularly useful in the aforementioned applications. To deal with visual ambiguities such as drift along the curve, we propose to texture the object with a coarse pattern of alternating colors and formulate the shape estimation as a deformable 1D template tracking problem. Tracking is expressed as an energy minimization over a set of control pointsQ parameterizing a 3D NURBS C3D modeling the object: Nicolas Padoy, Gregory D. Hager |
BMVC | 2 |
| 2012 | Sequential scene parsing using range and intensity informationabstractThis paper describes an extension of the sequential scene analysis system presented by Hager and Wegbreit [12]. In contrast to the original system, which was limited to scenes consisting of geometric primitives, such as spheres, cuboids, and cylinders computed from range data, the extended system is capable of dealing with arbitrarily shaped objects computed from range and intensity images. An object model composed of a triangulated geometry and intensity-based SURF features is introduced. The integration of prior object models into the sequential scene parsing framework is described. The extended system is evaluated with respect to pose estimation and its ability to handle complex scene sequences. It is shown that the new object models enable accurate pose estimation and reliable recognition even in highly cluttered scenes. Manuel Brucker, Simon Léonard, Tim Bodenmüller, Gregory D. Hager |
ICRA | 4 |
| 2012 | Robotic Path Planning for Surgeon Skill Evaluation in Minimally-Invasive Sinus Surgery
Narges Ahmidi, Gregory D. Hager, Lisa Ishii, Gary L. Gallia, Masaru Ishii |
MICCAI (1) | 2 |
| 2012 | Hybrid Tracking and Mosaicking for Information Augmentation in Retinal Surgery
Rogério Richa, Balázs Vágvölgyi, Marcin Balicki, Gregory D. Hager, Russell H. Taylor |
MICCAI (1) | 4 |
| 2012 | Data-Driven Visual Tracking in Retinal Microsurgery
Raphael Sznitman, Karim Ali 0002, Rogério Richa, Russell H. Taylor, Gregory D. Hager, Pascal Fua |
MICCAI (2) | 5 |
| 2012 | A System for Video-Based Navigation for Endoscopic Endonasal Skull Base SurgeryabstractSurgeries of the skull base require accuracy to safely navigate the critical anatomy. This is particularly the case for endoscopic endonasal skull base surgery (ESBS) where the surgeons work within millimeters of neurovascular structures at the skull base. Today's navigation systems provide approximately 2 mm accuracy. Accuracy is limited by the indirect relationship of the navigation system, the image and the patient. We propose a method to directly track the position of the endoscope using video data acquired from the endoscope camera. Our method first tracks image feature points in the video and reconstructs the image feature points to produce 3D points, and then registers the reconstructed point cloud to a surface segmented from preoperative computed tomography (CT) data. After the initial registration, the system tracks image features and maintains the 2D-3D correspondence of image features and 3D locations. These data are then used to update the current camera pose. We present a method for validation of our system, which achieves submillimeter (0.70 mm mean) target registration error (TRE) results. Daniel Mirota, Hanzi Wang, Russell H. Taylor, Masaru Ishii, Gary L. Gallia, Gregory D. Hager |
IEEE Trans. Medical Imaging | 6 |
| 2011 | Handheld micromanipulation with vision-based virtual fixturesabstractPrecise movement during micromanipulation becomes difficult in submillimeter workspaces, largely due to the destabilizing influence of tremor. Robotic aid combined with filtering techniques that suppress tremor frequency bands increases performance; however, if knowledge of the operator's goals is available, virtual fixtures have been shown to greatly improve micromanipulator precision. In this paper, we derive a control law for position-based virtual fixtures within the framework of an active handheld micromanipulator, where the fixtures are generated in real-time from microscope video. Additionally, we develop motion scaling behavior centered on virtual fixtures as a simple and direct extension to our formulation. We demonstrate that hard and soft (motion-scaled) virtual fixtures outperform state-of-the-art tremor cancellation performance on a set of artificial but medically relevant tasks: holding, move-and-hold, curve tracing, and volume restriction. Brian C. Becker, Robert A. MacLachlan, Gregory D. Hager, Cameron N. Riviere |
ICRA | 3 |
| 2011 | Towards integrating task information in skills assessment for dexterous tasks in surgery and simulationabstractWith the increasing popularity of robotic surgery, several studies in the literature have investigated automatically assessing skill measures based on motion and video data captured from these systems. A range of simulation environments for robotic surgery are now in development. Skill assessment in these environments has so far only focused on evaluating the utility and validity of statistics such as task completion time, and instrument distance measured during a simulated task. We present the first work using motion data from a robotic surgery simulation environment in development for classifying users of varying skills and detecting completion of trainee. Given the standardized environment of the simulator, and the availability of the ground truth, skill measurements and feedback based on task motion hold the promise of effective automated objective assessment. Based on motion data of a simulated manipulation task from 17 users of varying skills, we demonstrate binary classification (proficient vs. trainee) of user skill with 87.5% accuracy. Alternate measures based on instrument pose more relevant in the simulated environment including a new measure of motion efficiency are also presented and evaluated. Amod Jog, Brandon Itkowitz, May Liu, Simon P. DiMaio, Gregory D. Hager, Myriam Curet, Rajesh Kumar 0001 |
ICRA | 5 |
| 2011 | Human-Machine Collaborative surgery using learned modelsabstractIn the future of surgery, tele-operated robotic assistants will offer the possibility of performing certain commonly occurring tasks autonomously. Using a natural division of tasks into subtasks, we propose a novel surgical Human-Machine Collaborative (HMC) system in which portions of a surgical task are performed autonomously under complete surgeon's control, and other portions manually. Our system automatically identifies the completion of a manual subtask, seamlessly executes the next automated task, and then returns control back to the surgeon. Our approach is based on learning from demonstration. It uses Hidden Markov Models for the recognition of task completion and temporal curve averaging for learning the executed motions. We demonstrate our approach using a da Vinci tele-surgical robot. We show on two illustrative tasks where such human-machine collaboration is intuitive that automated control improves the usage of the master manipulator workspace. Because such a system does not limit the traditional use of the robot, but merely enhances its capabilities while leaving full control to the surgeon, it provides a safe and acceptable solution for surgical performance enhancement. Nicolas Padoy, Gregory D. Hager |
ICRA | 2 |
| 2011 | Object mapping, recognition, and localization from tactile geometryabstractWe present a method for performing object recognition using multiple images acquired from a tactile sensor. The method relies on using the tactile sensor as an imaging device, and builds an object representation based on mosaics of tactile measurements. We then describe an algorithm that is able to recognize an object using a small number of tactile sensor readings. Our approach makes extensive use of sequential state estimation techniques from the mobile robotics literature, whereby we view the object recognition problem as one of estimating a consistent location within a set of object maps. We examine and test approaches based on both traditional particle filtering and histogram filtering. We demonstrate both the mapping and recognition / localization techniques on a set of raised letter shapes using real tactile sensor data. Zachary A. Pezzementi, Caitlin Reyda, Gregory D. Hager |
ICRA | 3 |
| 2011 | Towards validation of robotic surgery training assessment across training platformsabstractRobotic surgery is increasingly popular for a wide range of complex minimally invasive surgery procedures. To improve robotic surgery training, a skills trainer simulator called dV-Trainer has recently been introduced, and a da Vinci Skills Simulator is in advanced evaluation. These platforms report a range of time and motion based task metrics and literature has investigated the validity of these metrics in training studies. However, the lack of a cross-platform data collection system has so far prevented a cross-platform investigation. Using a new architecture for collecting cross-platform motion data, we present the first study investigating whether metrics previously validated in simulation environments also hold in training exercises with a real robotic system. Preliminary experiments for an anastomosis needle throwing task in both simulated and real robotic environments are presented, and corresponding performance metrics for both proficient and trainee users are reported. Mert Sedef, Amod Jog, Peter Peng, Michael A. Choti, Gregory D. Hager, Jeff Berkley, Rajesh Kumar 0001 |
IROS | 6 |
| 2011 | 3D thread tracking for robotic assistance in tele-surgeryabstractRemote tele-manipulation tasks can be both long and exhausting. The operative workload can however be reduced through contextual systems, in which routine or dexterous actions are performed automatically. In this paper, we investigate this idea in tele-surgery by proposing automatic scissors, namely the possibility for a surgeon to invoke a third robotic arm to come and automatically cut the thread that he/she is holding. In particular, we address the problem of tracking deformable 3-dimensional (3D) curvilinear objects from stereo images. We propose an approach based on discrete Markov random field (MRF) optimization to track, in 3D, a thread modeled by a non-uniform rational B-spline (NURBS). We evaluate its accuracy off-line on synthetic and real data and illustrate its use for an automatic scissors command within an assistance system based on the da Vinci tele-surgical robot. Nicolas Padoy, Gregory D. Hager |
IROS | 2 |
| 2011 | Visual tracking using the sum of conditional varianceabstractThe goal of this paper is to introduce a direct visual tracking method based on an image similarity measure called the sum of conditional variance (SCV). The SCV was originally proposed in the medical imaging domain for registering multi-modal images. In the context of visual tracking, the SCV is invariant to non-linear illumination variations, multi-modal and computationally inexpensive. Compared to information theoretic tracking methods, it requires less iterations to converge and has a significantly larger convergence radius. The novelty in this paper is a generalization of the efficient second-order minimization formulation for tracking using the SCV, allowing us to combine the efficient second-order approximation of the Hessian with a similarity metric invariant to non-linear illumination variations. The result is a visual tracking method that copes with non-linear illumination variations without requiring the estimation of photometric correction parameters at every iteration. We demonstrate the superior performance of the proposed method through comparative studies and tracking experiments under challenging illumination conditions and rapid motions. Rogério Richa, Raphael Sznitman, Russell H. Taylor, Gregory D. Hager |
IROS | 4 |
| 2011 | Tactile Object Recognition and Localization Using Spatially-Varying Appearance
Zachary A. Pezzementi, Gregory D. Hager |
ISRR | 2 |
| 2011 | Spatio-Temporal Registration of Multiple Trajectories
Nicolas Padoy, Gregory D. Hager |
MICCAI (1) | 2 |
| 2011 | Ultrasound Elastography Using Three Images
Hassan Rivaz, Emad Boctor, Michael A. Choti, Gregory D. Hager |
MICCAI (1) | 4 |
| 2011 | Unified Detection and Tracking in Retinal Microsurgery
Raphael Sznitman, Anasuya Basu, Rogério Richa, Jim Handa, Peter Gehlbach, Russell H. Taylor, Bruno Jedynak, Gregory D. Hager |
MICCAI (1) | 8 |
| 2011 | Real-Time Regularized Ultrasound ElastographyabstractThis paper introduces two real-time elastography techniques based on analytic minimization (AM) of regularized cost functions. The first method (1D AM) produces axial strain and integer lateral displacement, while the second method (2D AM) produces both axial and lateral strains. The cost functions incorporate similarity of radio-frequency (RF) data intensity and displacement continuity, making both AM methods robust to small decorrelations present throughout the image. We also exploit techniques from robust statistics to make the methods resistant to large local decorrelations. We further introduce Kalman filtering for calculating the strain field from the displacement field given by the AM methods. Simulation and phantom experiments show that both methods generate strain images with high SNR, CNR and resolution. Both methods work for strains as high as 10% and run in real-time. We also present in vivo patient trials of ablation monitoring. An implementation of the 2D AM method as well as phantom and clinical RF-data can be downloaded. Hassan Rivaz, Emad Boctor, Michael A. Choti, Gregory D. Hager |
IEEE Trans. Medical Imaging | 4 |
| 2011 | A Meta Method for Image MatchingabstractThis paper presents a novel system for image matching in optical endoscopy. The proposed metamatching system approaches the challenge of matching images in a complex scene by incorporating multiple matchers and a decision function. Experiments are presented for Crohn's disease lesion matching in capsule endoscopy with a metamatcher consisting of five independent matchers. We compare the performance of six different types of decision functions. Results show that the F-measure of the metamatching system containing all five matchers is 4%-7% greater than the performance of using the best matcher only, with a maximum F-measure of 0.811. The robustness of the method is validated using simulated data generated by controlled deformations of the image. We also demonstrate how the addition of simulated data to the training set can be used to augment the performance of the metamatcher by up to 10%. Sharmishtaa Seshamani, Rajesh Kumar 0001, Gerard Mullin, Themistocles Dassopoulos, Gregory D. Hager |
IEEE Trans. Medical Imaging | 5 |
| 2011 | Tactile-Object Recognition From Appearance InformationabstractThis paper explores the connection between sensor-based perception and exploration in the context of haptic object identification. The proposed approach combines 1) object recognition from tactile appearance with 2) purposeful haptic exploration of unknown objects to extract appearance information. The recognition component brings to bear computer-vision techniques by viewing tactile-sensor readings as images. We present a bag-of-features framework that uses several tactile-image descriptors, some that are adapted from the vision domain and others that are novel, to estimate a probability distribution over object identity as an unknown object is explored. Haptic exploration is treated as a search problem in a continuous space to take advantage of sampling-based motion planning to explore the unknown object and construct its tactile appearance. Simulation experiments of a robot arm equipped with a haptic sensor at the end-effector provide promising validation, thereby indicating high accuracy in identifying complex shapes from tactile information gathered during exploration. The proposed approach is also validated by using readings from actual tactile sensors to recognize real objects. Zachary A. Pezzementi, Erion Plaku, Caitlin Reyda, Gregory D. Hager |
IEEE Trans. Robotics | 4 |
| 2010 | Adaptive and Generic Corner Detection Based on the Accelerated Segment Test
Elmar Mair, Gregory D. Hager, Darius Burschka, Michael Suppa, Gerd Hirzinger |
ECCV (2) | 2 |
| 2010 | Epipolar-Based Stereo Tracking Without Explicit 3D ReconstructionabstractWe present a general framework for tracking image regions in two views simultaneously based on sum-of-squared differences (SSD) minimization. Our method allows for motion models up to affine transformations. Contrary to earlier approaches, we incorporate the well-known epipolar constraints directly into the SSD optimization process. Since the epipolar geometry can be computed from the image directly, no prior calibration is necessary. Our algorithm has been tested in different applications including camera localization, wide-baseline stereo, object tracking and medical imaging. We show experimental results on robustness and accuracy compared to the known ground truth given by a conventional tracking device. Andre Gaschler, Darius Burschka, Gregory D. Hager |
ICPR | 3 |
| 2010 | Sampling-Based Motion and Symbolic Action Planning with geometric and differential constraintsabstractTo compute collision-free and dynamically-feasibile trajectories that satisfy high-level specifications given in a planning-domain definition language, this paper proposes to combine sampling-based motion planning with symbolic action planning. The proposed approach, Sampling-based Motion and Symbolic Action Planner (SMAP), leverages from sampling-based motion planning the underlying idea of searching for a solution trajectory by selectively sampling and exploring the continuous space of collision-free and dynamically-feasible motions. Drawing from AI, SMAP uses symbolic action planning to identify actions and regions of the continuous space that sampling-based motion planning can further explore to significantly advance the search. The planning layers interact with each-other through estimates on the utility of each action, which are computed based on information gathered during the search. Simulation experiments with dynamical models of vehicles carrying out tasks given by high-level STRIPS specifications provide promising initial validation, showing that SMAP efficiently solves challenging problems. Erion Plaku, Gregory D. Hager |
ICRA | 2 |
| 2010 | Surgical Task and Skill Classification from Eye Tracking and Tool Motion in Minimally Invasive Surgery
Narges Ahmidi, Gregory D. Hager, Lisa Ishii, Gabor Fichtinger, Gary L. Gallia, Masaru Ishii |
MICCAI (3) | 2 |
| 2010 | Tracked Ultrasound Elastography (TrUE)
Pezhman Foroughi, Hassan Rivaz, Ioana Fleming, Gregory D. Hager, Emad Boctor |
MICCAI (2) | 4 |
| 2010 | Augmenting Capsule Endoscopy Diagnosis: A Similarity Learning Approach
Sharmishtaa Seshamani, Rajesh Kumar 0001, Themistocles Dassopoulos, Gerard Mullin, Gregory D. Hager |
MICCAI (2) | 5 |
| 2010 | Adaptive Multispectral Illumination for Retinal Microsurgery
Raphael Sznitman, Diego Rother, James Handa, Peter Gehlbach, Gregory D. Hager, Russell H. Taylor |
MICCAI (3) | 5 |
| 2010 | A Generalized Kernel Consensus-Based Robust EstimatorabstractIn this paper, we present a new Adaptive-Scale Kernel Consensus (ASKC) robust estimator as a generalization of the popular and state-of-the-art robust estimators such as RANdom SAmple Consensus (RANSAC), Adaptive Scale Sample Consensus (ASSC), and Maximum Kernel Density Estimator (MKDE). The ASKC framework is grounded on and unifies these robust estimators using nonparametric kernel density estimation theory. In particular, we show that each of these methods is a special case of ASKC using a specific kernel. Like these methods, ASKC can tolerate more than 50 percent outliers, but it can also automatically estimate the scale of inliers. We apply ASKC to two important areas in computer vision, robust motion estimation and pose estimation, and show comparative results on both synthetic and real data. Hanzi Wang, Daniel Mirota, Gregory D. Hager |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Histograms of oriented optical flow and Binet-Cauchy kernels on nonlinear dynamical systems for the recognition of human actionsabstractSystem theoretic approaches to action recognition model the dynamics of a scene with linear dynamical systems (LDSs) and perform classification using metrics on the space of LDSs, e.g. Binet-Cauchy kernels. However, such approaches are only applicable to time series data living in a Euclidean space, e.g. joint trajectories extracted from motion capture data or feature point trajectories extracted from video. Much of the success of recent object recognition techniques relies on the use of more complex feature descriptors, such as SIFT descriptors or HOG descriptors, which are essentially histograms. Since histograms live in a non-Euclidean space, we can no longer model their temporal evolution with LDSs, nor can we classify them using a metric for LDSs. In this paper, we propose to represent each frame of a video using a histogram of oriented optical flow (HOOF) and to recognize human actions by classifying HOOF time-series. For this purpose, we propose a generalization of the Binet-Cauchy kernels to nonlinear dynamical systems (NLDS) whose output lives in a non-Euclidean space, e.g. the space of histograms. This can be achieved by using kernels defined on the original non-Euclidean space, leading to a well-defined metric for NLDSs. We use these kernels for the classification of actions in video sequences using (HOOF) as the output of the NLDS. We evaluate our approach to recognition of human actions in several scenarios and achieve encouraging results. Rizwan Chaudhry, Avinash Ravichandran, Gregory D. Hager, René Vidal |
CVPR | 3 |
| 2009 | Active guidance of a handheld micromanipulator using visual servoingabstractIn microsurgery, a surgeon often deals with anatomical structures of sizes that are close to the limit of the human hand accuracy. Robotic assistants can help to push beyond the current state of practice by integrating imaging and robot-assisted tools. This paper demonstrates control of a handheld tremor reduction micromanipulator with visual servo techniques, aiding the operator by providing three behaviors: snap-to, motion-scaling, and standoff-regulation. A stereo camera setup viewing the workspace under high magnification tracks the tip of the micromanipulator and the desired target object being manipulated. Individual behaviors activate in task-specific situations when the micromanipulator tip is in the vicinity of the target. We show that the snap-to behavior can reach and maintain a position at a target with an accuracy of 17.5 ± 0.4μm Root Mean Squared Error (RMSE) distance between the tip and target. Scaling the operator's motions and preventing unwanted contact with non-target objects also provides a larger margin of safety. Brian C. Becker, Sandrine Voros, Robert A. MacLachlan, Gregory D. Hager, Cameron N. Riviere |
ICRA | 4 |
| 2009 | Analysis of Crohn's disease lesions in capsule endoscopy imagesabstractCapsule endoscopy (CE) is aimed at diagnosing disease in areas of the gastrointestinal (GI) tract beyond the reach of conventional endoscopy. Recent work has addressed various methods for reducing the complexity of CE diagnosis and the time needed for analyzing the data. This includes detection of lumen and its contractions, fluids such as blood and intestinal juices, as well as extraneous matter such as food and bubbles. This paper outlines our ongoing work to segment lesions (in particular Crohn's disease) and other abnormalities in CE images. In particular, here we describe the data collection and clinical analysis for our project and preliminary results for segmenting abnormal and extraneous images from a set of 10 CE studies. Srdan Bejakovic, Rajesh Kumar 0001, Themistocles Dassopoulos, Gerard Mullin, Gregory D. Hager |
ICRA | 5 |
| 2009 | Articulated object tracking by rendering consistent appearance partsabstractWe describe a general methodology for tracking 3-dimensional objects in monocular and stereo video that makes use of GPU-accelerated filtering and rendering in combination with machine learning techniques. The method operates on targets consisting of kinematic chains with known geometry. The tracked target is divided into one or more areas of consistent appearance. The appearance of each area is represented by a classifier trained to assign a class-conditional probability to image feature vectors. A search is then performed on the configuration space of the target to find the maximum likelihood configuration. In the search, candidate hypotheses are evaluated by rendering a 3D model of the target object and measuring its consistency with the class probability map. The method is demonstrated for tool tracking on videos from two surgical domains, as well as in a human hand-tracking task. Zachary A. Pezzementi, Sandrine Voros, Gregory D. Hager |
ICRA | 3 |
| 2009 | Toward Video-Based Navigation for Endoscopic Endonasal Skull Base Surgery
Daniel Mirota, Hanzi Wang, Russell H. Taylor, Masaru Ishii, Gregory D. Hager |
MICCAI (1) | 5 |
| 2009 | Task versus Subtask Surgical Skill Evaluation of Robotic Minimally Invasive Surgery
Carol E. Reiley, Gregory D. Hager |
MICCAI (1) | 2 |
| 2009 | Tracked Regularized Ultrasound Elastography for Targeting Breast Radiotherapy
Hassan Rivaz, Pezhman Foroughi, Ioana Fleming, Richard Zellars, Emad Boctor, Gregory D. Hager |
MICCAI (1) | 6 |
| 2009 | A Meta Registration Framework for Lesion Matching
Sharmishtaa Seshamani, Purnima Rajan, Rajesh Kumar 0001, Hani Z. Girgis, Themistocles Dassopoulos, Gerard Mullin, Gregory D. Hager |
MICCAI (1) | 7 |
| 2009 | Data-Derived Models for Segmentation with Application to Surgical Assessment and Training
Balakrishnan Varadarajan, Carol E. Reiley, Henry C. Lin 0001, Sanjeev Khudanpur, Gregory D. Hager |
MICCAI (1) | 5 |
| 2009 | Intelligent frame selection for anatomic reconstruction from endoscopic videoabstractUsing endoscopic video, it is possible to perform 3D reconstruction of the anatomy using the well known epipolar constraint between matched feature points. Through this constraint, it is possible to recover the translation and rotation between camera positions and thus reconstruct the 3D anatomy by triangulation. However, these motion estimates are not stable for small camera motions. In this work, we propose a covariance estimation scheme to select pairs of frames which give rise to stable motion estimates, i.e. minimal variance with respect to pixel match error. We parameterize the essential matrix using a minimal 5 parameter representation and estimate motion covariance based upon the estimated feature match variance. The proposed algorithm is applied to endoscopic video sequences recorded in porcine sinus passages in order to extract stable motion estimates. Daniel Abretske, Daniel Mirota, Gregory D. Hager, Masaru Ishii |
WACV | 3 |
| 2009 | Image description with features that summarize
Jason J. Corso, Gregory D. Hager |
Comput. Vis. Image Underst. | 2 |
| 2008 | Robust motion estimation and structure recovery from endoscopic image sequences with an Adaptive Scale Kernel Consensus estimatorabstractTo correctly estimate the camera motion parameters and reconstruct the structure of the surrounding tissues from endoscopic image sequences, we need not only to deal with outliers (e.g., mismatches), which may involve more than 50% of the data, but also to accurately distinguish inliers (correct matches) from outliers. In this paper, we propose a new robust estimator, Adaptive Scale Kernel Consensus (ASKC), which can tolerate more than 50 percent outliers while automatically estimating the scale of inliers. With ASKC, we develop a reliable feature tracking algorithm. This, in turn, allows us to develop a complete system for estimating endoscopic camera motion and reconstructing anatomical structures from endoscopic image sequences. Preliminary experiments on endoscopic sinus imagery have achieved promising results. Hanzi Wang, Daniel Mirota, Masaru Ishii, Gregory D. Hager |
CVPR | 4 |
| 2008 | Control methods for guidance virtual fixtures in compliant human-machine interfacesabstractThis work focuses on the implementation of a vision-based motion guidance method, called virtual fixtures, on admittance-controlled human-machine cooperative robots with compliance. The robot compliance here refers to the structural elastic deformation of the device. The high mechanical stiffness and non-backdrivability of a typical admittance-controlled robot allow for slow and precise motions, making it highly suitable for tasks that require accuracy near human physical limits, such as microsurgery. However, previous experiments have shown that even small robot compliance degraded virtual fixture performance, especially at the micro scale. In this work, control methods to minimize the effect of robot compliance on virtual fixture performance were developed for admittance-controlled cooperative systems. Based on a linear model of the robot dynamics, we applied a Kalman filter to integrate the measurements obtained from the camera and encoders to estimate the robot end-effector position. A partitioned control law was used to achieve end-effector trajectory following on the desired velocity commanded by the admittance and virtual fixture control laws. The effectiveness of the Kalman filter and the controller was validated on a one degree-of-freedom admittance-controlled cooperative testbed. Panadda Marayong, Gregory D. Hager, Allison M. Okamura |
IROS | 2 |
| 2008 | Cooperative Robot Assistant for Retinal Microsurgery
Ioana Fleming, Marcin Balicki, John Koo, Iulian Iordachita, Ben Mitchell, James Handa, Gregory D. Hager, Russell H. Taylor |
MICCAI (2) | 7 |
| 2008 | Ablation Monitoring with Elastography: 2D In-vivoand 3D Ex-vivoStudies
Hassan Rivaz, Ioana Fleming, Lia Assumpcao, Gabor Fichtinger, Ulrike M. Hamper, Michael A. Choti, Gregory D. Hager, Emad Boctor |
MICCAI (2) | 7 |
| 2008 | Intraoperative Visualization of Anatomical Targets in Retinal SurgeryabstractCertain surgical procedures require a high degree of precise manual control within a very restricted area. Retinal surgeries are part of this group of procedures. During vitreoretinal surgery, the surgeon must visualize, using a microscope, an area spanning a few hundreds of microns in diameter and manually correct the potential pathology using direct contact, free hand techniques. In addition, the surgeon must find an effective compromise between magnification, depth perception, field of view, and clarity of view. Pre-operative images are used to locate interventional targets, and also to assess and plan the surgical procedure. This paper proposes a method of fusing information contained in pre-operative imagery, such as fundus and OCT images, with intra-operative video to increase accuracy in finding the target areas. We describe methods for maintaining, in real-time, registration with anatomical features and target areas using image processing. This registration allows us to produce information enhanced displays that ensure that the retinal surgeon is always in visual contact with his/her area of interest. Ioana Fleming, Sandrine Voros, Balázs Vágvölgyi, Zachary A. Pezzementi, James Handa, Russell H. Taylor, Gregory D. Hager |
WACV | 7 |
| 2008 | Ultrasound Elastography: A Dynamic Programming ApproachabstractThis paper introduces a 2-D strain imaging technique based on minimizing a cost function using dynamic programming (DP). The cost function incorporates similarity of echo amplitudes and displacement continuity. Since tissue deformations are smooth, the incorporation of the smoothness into the cost function results in reduced decorrelation noise. As a result, the method generates high-quality strain images of freehand palpation elastography with up to 10% compression, showing that the method is more robust to signal decorrelation (caused by scatterer motion in high axial compression and nonaxial motions of the probe) in comparison to the standard correlation techniques. The method operates in less than 1 s and is thus also potentially suitable for real time elastography. Hassan Rivaz, Emad Boctor, Pezhman Foroughi, Richard Zellars, Gabor Fichtinger, Gregory D. Hager |
IEEE Trans. Medical Imaging | 6 |
| 2007 | Deformable Motion Tracking of Cardiac Structures (DEMOTRACS) for Improved MR ImagingabstractThe speed and quality of imaging cardiac structures (coronary arteries, cardiac valves etc) in MR can be improved by tracking and predicting their motion in MR images. The problem is challenging not only due to the complex motion of these structures that significantly changes the appearance of the region of interest, but also the ability to track at different spatial and temporal resolutions depending on the application. We have developed a multiple-template based tracking approach to track the cardiac structures in MR images. The algorithm has two novel features. First a bidirectional coordinate-descent algorithm is derived to improve accuracy and performance of tracking. Second we propose a method for choosing an optimal set of templates for tracking. The efficacy of the algorithm has been validated by tracking the coronary artery and cardiac valves reliably and accurately in thousands of high resolution cine and low-re solution real-time MR images. Maneesh Dewan, Christine H. Lorenz, Gregory D. Hager |
CVPR | 3 |
| 2007 | A Nonparametric Treatment for Location/Segmentation Based Visual TrackingabstractIn this paper, we address two closely related visual tracking problems: 1) localizing a target's position in low or moderate resolution videos and 2) segmenting a target's image support in moderate to high resolution videos. Both tasks are treated as an online binary classification problem using dynamic foreground/background appearance models. Our major contribution is a novel nonparametric approach that successfully maintains a temporally changing appearance model for both foreground and background. The appearance models are formulated as "bags of image patches" that approximate the true two-class appearance distributions. They are maintained using a temporal-adaptive importance resampling procedure that is based on simple nonparametric statistics of the appearance patch bags. The overall framework is independent of an specific foreground/background classification process and thus offers the freedom to use different classifiers. We demonstrate the effectiveness of our approach with extensive comparative experimental results on sequences from previous visual tracking [1, 12] and video matting [4] work as well as our own data. Le Lu 0001, Gregory D. Hager |
CVPR | 2 |
| 2007 | Full Motion Tracking in Ultrasound Using Image Speckle Information and Visual ServoingabstractThis paper presents a new visual servoing method that is able to stabilize a moving area of soft tissue within an ultrasound B-mode imaging plane. The approach consists of moving the probe in order to minimize the relative position between a target imaging plane and the ultrasound plane observed by the probe of the moving tissue target. The problem is decoupled into motion out-of-plane and motion within plane. For the former, a new original method based on the speckle information contained in the images is developed. For the latter, an image region tracker is used to provide the in-plane motion. A visual servoing control scheme is then developed to perform the tracking robotic task. The method is validated on simulated motions of a probe on a static ultrasound volume acquired from a phantom. Alexandre Krupa, Gabor Fichtinger, Gregory D. Hager |
ICRA | 3 |
| 2007 | Development and Application of a New Steady-Hand Manipulator for Retinal SurgeryabstractThis paper describes the development and initial testing of a new and optimized version of a steady-hand manipulator for retinal microsurgery. In the steady-hand paradigm, the surgeon and the robot share control of a tool attached to the robot through a force sensor. The robot controller senses forces exerted by the operator on the tool and uses this information in various control modes to provide smooth, tremor-free, precise positional control and force scaling. The steady-hand manipulator reported here has been specifically designed with the unique constraints of retinal microsurgery in mind. In particular, the system makes use of a compact wrist design that places the bulk of the robot away from the operating field. The resulting system has high efficacy, flexibility and ergonomics while meeting the accuracy and safety requirements of microsurgery. We have now tested this robot on a biological model system and we report a protocol for reliably cannulating ~80 mum OD veins (the size of veins in the human retina) using the system Ben Mitchell, John Koo, Iulian Iordachita, Peter Kazanzides, Ankur Kapoor, James Handa, Gregory D. Hager, Russell H. Taylor |
ICRA | 7 |
| 2007 | Dynamic Guidance with Pseudoadmittance Virtual FixturesabstractHuman machine collaborative systems (HMCS) have been developed to enhance sensation and suppress extraneous motions or forces during surgical tasks requiring precise motion. However, to date such systems have enforced constraints on the position or path of a tool, but have not considered the dynamics of motion. Also, the focus has been on the effect of guidance of motion during a task, rather than on the learning of motion skills through repetition. We present a pseudo-admittance framework for HMCS design to guide the user's velocity in such tasks. Two different fixture design approaches are analyzed, implemented and compared. Three tests are then conducted, showing the fixtures' promise for both guiding and learning motions with dynamics Zachary A. Pezzementi, Allison M. Okamura, Gregory D. Hager |
ICRA | 3 |
| 2007 | Kernel-based visual servoingabstractTraditionally, visual servoing is separated into tracking and control subsystems. This separation, though convenient, is not necessarily well justified. When tracking and control strategies are designed independently, it is not clear how to optimize them to achieve a certain task. In this work, we propose a framework in which spatial sampling kernels - borrowed from the tracking and registration literature - are used to design feedback controllers for visual servoing. The use of spatial sampling kernels provides natural hooks for Lyapunov theory, thus unifying tracking and control and providing a framework for optimizing a particular servoing task. As a first step, we develop kernel-based visual servos for a subset of relative motions between camera and target scene. The subset of motions we consider are 2D translation, scale, and roll of the target relative to the camera. Our approach provides formal guarantees on the convergence/stability of visual servoing algorithms under putatively generic conditions. Vinutha Kallem, Maneesh Dewan, John P. Swensen, Gregory D. Hager, Noah J. Cowan |
IROS | 4 |
| 2007 | Real-Time Tissue Tracking with B-Mode Ultrasound Using Speckle and Visual Servoing
Alexandre Krupa, Gabor Fichtinger, Gregory D. Hager |
MICCAI (2) | 3 |
| 2007 | Editorial: Special Issue on Vision and Robotics, Parts I and II
Gregory D. Hager, Martial Hebert, Seth Hutchinson 0001 |
Int. J. Comput. Vis. | 1 |
| 2006 | Portability and Applicability of Virtual Fixtures across Medical and Manufacturing TasksabstractVirtual fixtures are virtual constraints that enhance human performance in motion tasks. They can either confine and/or guide a user's motion. In this paper, we use a commercially available motion platform to explore the portability and applicability of virtual fixtures and document how people interact with them. Two micromanipulation tasks are analyzed and the effects of similarly designed virtual fixtures are discussed. One task simulates a medical task, retinal vein cannulation, and the other simulates a manufacturing task, fine leads soldering. Preliminary experimental results show that the virtual fixtures increase the accuracy of both medical and manufacturing tasks, lending support to its portability and applicability across unrelated tasks Henry C. Lin 0001, Keith Mills, Peter Kazanzides, Gregory D. Hager, Panadda Marayong, Allison M. Okamura, Ray Karam |
ICRA | 4 |
| 2006 | Ultrasound Monitoring of Tissue Ablation Via Deformation Model and Shape Priors
Emad Boctor, Michelle de Oliveira, Michael A. Choti, Roger G. Ghanem, Russell H. Taylor, Gregory D. Hager, Gabor Fichtinger |
MICCAI (2) | 6 |
| 2006 | Real-Time Endoscopic Mosaicking
Sharmishtaa Seshamani, William W. Lau, Gregory D. Hager |
MICCAI (1) | 3 |
| 2006 | Dynamic Foreground/Background Extraction from Images and Videos using Random PatchesabstractIn this paper, we propose a novel exemplar-based approach to extract dynamic foreground regions from a changing background within a collection of images or a video sequence. By using image segmentation as a pre-processing step, we convert this traditional pixel-wise labeling problem into a lower-dimensional supervised, binary labeling procedure on image segments. Our approach consists of three steps. First, a set of random image patches are spatially and adaptively sampled within each segment. Second, these sets of extracted samples are formed into two "bags of patches" to model the foreground/background appearance, respectively. We perform a novel bidirectional consistency check between new patches from incoming frames and current "bags of patches" to reject outliers, control model rigidity and make the model adaptive to new observations. Within each bag, image patches are further partitioned and resampled to create an evolving appearance model. Finally, the foreground/background decision over segments in an image is formulated using an aggregation function defined on the similarity measurements of sampled patches relative to the foreground and background models. The essence of the algorithm is conceptually simple and can be easily implemented within a few hundred lines of Matlab code. We evaluate and validate the proposed approach by extensive real examples of the object-level image mapping and tracking within a variety of challenging environments. We also show that it is straightforward to apply our problem formulation on non-rigid object tracking with difficult surveillance videos. Le Lu 0001, Gregory D. Hager |
NIPS | 2 |
| 2006 | Efficient particle filtering using RANSAC with application to 3D face tracking
Le Lu 0001, Xiangtian Dai, Gregory D. Hager |
Image Vis. Comput. | 3 |
| 2005 | Coherent Regions for Concise and Stable Image DescriptionabstractWe present a new method for summarizing images for the purposes of matching and registration. We take the point of view that large, coherent regions in the image provide a concise and stable basis for image description. We develop a new algorithm for image segmentation that operates on several projections (feature spaces) of the image, using kernel-based optimization techniques to locate local extrema of a continuous scale-space of image regions. Descriptors of these image regions and their relative geometry then form the basis of an image description. We present experimental results of these methods applied to the problem of image retrieval. On a moderate sized database, we find that our method performs comparably to two published techniques: Blobworld and SIFT features. However, compared to these techniques two significant advantages of our method are its 1) stability under large changes in the images and 2) its representational efficiency. As a result we argue our proposed method will scale well with larger image sets. Jason J. Corso, Gregory D. Hager |
CVPR (2) | 2 |
| 2005 | A Two Level Approach for Scene RecognitionabstractClassifying pictures into one of several semantic categories is a classical image understanding problem. In this paper, we present a stratified approach to both binary (outdoor-indoor) and multiple category of scene classification. We first learn mixture models for 20 basic classes of local image content based on color and texture information. Once trained, these models are applied to a test image, and produce 20 probability density response maps (PDRM) indicating the likelihood that each image region was produced by each class. We then extract some very simple features from those PDRMs, and use them to train a bagged LDA classifier for 10 scene categories. For this process, no explicit region segmentation or spatial context model are computed. To test this classification system, we created a labeled database of 1500 photos taken under very different environment and lighting conditions, using different cameras, and from 43 persons over 5 years. The classification rate of outdoor-indoor classification is 93.8%, and the classification rate for 10 scene categories is 90.1%. As a byproduct, local image patches can be contextually labeled into the 20 basic material classes by using loopy belief propagation (Yedidia et al., 2001) as an anisotropic filter on PDRMs, producing an image-level segmentation if desired. Le Lu 0001, Kentaro Toyama, Gregory D. Hager |
CVPR (1) | 3 |
| 2005 | Vision-Based 3D Scene Analysis for Driver AssistanceabstractWe present a vision-based system for traffic sign detection and ego-motion estimation in road scenarios. The system is capable of autonomous scene reconstruction and classification. It is used to pre-select candidate surfaces in the vicinity of the road that should be inspected more closely by a sign recognition system. We compare two approaches based on a binocular and a monocular camera system, respectively. We discuss their advantages and disadvantages for applications in driver assistance systems. Darius Burschka, Gregory D. Hager |
ICRA | 2 |
| 2005 | Real-Time Quality Control of Tracked Ultrasound
Emad Boctor, Iulian Iordachita, Gabor Fichtinger, Gregory D. Hager |
MICCAI | 4 |
| 2005 | DaVinci Canvas: A Telerobotic Surgical System with Integrated, Robot-Assisted, Laparoscopic Ultrasound Capability
Joshua Leven, Darius Burschka, Rajesh Kumar 0001, Gary Zhang, Steve Blumenkranz, Xiangtian Dai, Michael Awad, Gregory D. Hager, Mike Marohn, Michael A. Choti, Christopher J. Hasser, Russell H. Taylor |
MICCAI | 8 |
| 2005 | Automatic Detection and Segmentation of Robot-Assisted Surgical Motions
Henry C. Lin 0001, Izhak Shafran, Todd E. Murphy, Allison M. Okamura, David D. Yuh, Gregory D. Hager |
MICCAI | 6 |
| 2005 | Scale-invariant registration of monocular endoscopic images to CT-scans for sinus surgery
Darius Burschka, Ming Li 0052, Masaru Ishii, Russell H. Taylor, Gregory D. Hager |
Medical Image Anal. | 5 |
| 2004 | Probabilistic Data Association Methods in Visual Tracking of Group
Giambattista Gennari, Gregory D. Hager |
CVPR (2) | 2 |
| 2004 | Multiple Kernel Tracking with SSD
Gregory D. Hager, Maneesh Dewan, Charles V. Stewart |
CVPR (1) | 1 |
| 2004 | V-GPS(SLAM): Vision-based Inertial System for Mobile RobotsabstractWe present a novel vision-based approach to simultaneous localization and mapping (SLAM). We discuss it in the context of estimating the 6 DoF pose of a mobile robot from the perception of a monocular camera using a minimum set of three natural landmarks. In contrast to our previously presented V-GPS system, which navigates based on a set of known landmarks, the current approach allows to estimate the required information about the landmarks on-the-fly during the exploration of an unknown environment The method is applicable to indoor and outdoor environments. The calculation is done from the image position of a set of natural landmarks that are tracked in a continuous video stream at frame-rate. An automatic hand-off process allows an update of the set to compensate for occlusions and decreasing reconstruction accuracies with the distance to an imaged landmark. A generic sensor model allows a system configuration with a variety of physical sensors including: monocular perspective cameras, omni-directional cameras and laser range finders. Darius Burschka, Gregory D. Hager |
ICRA | 2 |
| 2004 | Scale-invariant registration of monocular stereo images to 3D surface modelsabstractWe present an approach for scale recovery from monocular stereo images of an endoscopic camera with simultaneous registration to dense 3D surface models. We assume the camera motion to be unknown or at least uncertain. An example application is the registration of endoscope images to pre-operative CT scans that allows instrument navigation during surgical procedures. The application field is not restricted to the medical field. It can be extended to registration of monocular video images to laser-based surface reconstructions in, e.g., mobile navigation area or to autonomous aircraft navigation from topological surveys. A novel way for depth estimation from arbitrary camera motion is presented. In this paper, we focus on the robust initialization of the system and on the scale recovery for the reconstructed 3D point clouds with accurate registration to the candidate surfaces extracted from the CT data. We provide experimental validation of the algorithm with data obtained from our experiments with a phantom skull. Darius Burschka, Ming Li 0052, Russell H. Taylor, Gregory D. Hager |
IROS | 4 |
| 2004 | Scale-Invariant Registration of Monocular Endoscopic Images to CT-Scans for Sinus Surgery
Darius Burschka, Ming Li 0052, Russell H. Taylor, Gregory D. Hager |
MICCAI (2) | 4 |
| 2004 | Vision-Based Assistance for Ophthalmic Micro-Surgery
Maneesh Dewan, Panadda Marayong, Allison M. Okamura, Gregory D. Hager |
MICCAI (2) | 4 |
| 2004 | Stereo-Based Endoscopic Tracking of Cardiac Surface Deformation
William W. Lau, Nicholas A. Ramey, Jason J. Corso, Nitish V. Thakor, Gregory D. Hager |
MICCAI (2) | 5 |
| 2004 | Immediate Ultrasound Calibration with Three Poses and Minimal Image Processing
Anand Viswanathan, Emad Boctor, Russell H. Taylor, Gregory D. Hager, Gabor Fichtinger |
MICCAI (2) | 4 |
| 2004 | A Three Tiered Approach for Articulated Object Action Modeling and RecognitionabstractVisual action recognition is an important problem in computer vision. In this paper, we propose a new method to probabilistically model and recognize actions of articulated objects, such as hand or body gestures, in image sequences. Our method consists of three levels of representa- tion. At the low level, we first extract a feature vector invariant to scale and in-plane rotation by using the Fourier transform of a circular spatial histogram. Then, spectral partitioning [20] is utilized to obtain an initial clustering; this clustering is then refined using a temporal smoothness constraint. Gaussian mixture model (GMM) based clustering and density estimation in the subspace of linear discriminant analysis (LDA) are then applied to thousands of image feature vectors to obtain an intermediate level representation. Finally, at the high level we build a temporal multi- resolution histogram model for each action by aggregating the clustering weights of sampled images belonging to that action. We discuss how this high level representation can be extended to achieve temporal scaling in- variance and to include Bi-gram or Multi-gram transition information. Both image clustering and action recognition/segmentation results are given to show the validity of our three tiered representation. Le Lu 0001, Gregory D. Hager, Laurent Younes |
NIPS | 2 |
| 2004 | VICs: A modular HCI framework using spatiotemporal dynamics
Guangqi Ye, Jason J. Corso, Darius Burschka, Gregory D. Hager |
Mach. Vis. Appl. | 4 |
| 2004 | Vision-assisted control for manipulation using virtual fixturesabstractWe present the design and implementation of a vision-based system for cooperative manipulation at millimeter to micrometer scales. The system is based on an admittance control algorithm that implements a broad class of guidance modes called virtual fixtures. A virtual fixture, like a real fixture, limits the motion of a tool to a prescribed class or range of motions. We describe how both hard (unyielding) and soft (yielding) virtual fixtures can be implemented in this control framework. We then detail the construction of virtual fixtures for point positioning and curve following as well as extensions of these to tubes, cones, and sequences thereof. We also describe an implemented system using the JHU Steady Hand Robot. The system uses computer vision as a sensor for providing a reference trajectory, and the virtual fixture control algorithm then provides haptic feedback to implemented direct, shared manipulation. We provide extensive experimental results detailing both system performance and the effects of virtual fixtures on human speed and accuracy. Alessandro Bettini, Panadda Marayong, Samuel Lang, Allison M. Okamura, Gregory D. Hager |
IEEE Trans. Robotics | 5 |
| 2003 | Optimal landmark configuration for vision-based control of mobile robotsabstractWe analyze the problem of finding the optimal placement of tracked primitives for robust vision-based control of a mobile robot. The analysis evaluates the properties of the Image Jacobian matrix, used for direct generation of the control signals from the error signal in the image, and the accuracy of the underlying sensor system. The analysis is then used to select optimal tracking primitives that ensure good observability and controllability of the mobile system for a variety of sensor system configurations. The theoretical results are validated with our mobile robot for system configurations that use standard video cameras mounted on a pan-tilt head and catadioptric systems. Darius Burschka, Jeremy Geiman, Gregory D. Hager |
ICRA | 3 |
| 2003 | Direct plane tracking in stereo images for mobile navigationabstractWe present a novel plane tracking algorithm based on the direct update of surface parameters from two stereo images. The plane tracking algorithm is posed as an optimization problem, and maintains an iteratively re-weighted least squares approximation of the plane's orientation using direct pixel measurements. To facilitate autonomous operation, we include an algorithm for robust detection of significant planes in the environment. The algorithms have been implemented in a robot navigation system. Jason J. Corso, Darius Burschka, Gregory D. Hager |
ICRA | 3 |
| 2003 | Spatial motion constraints: theory and demonstrations for robot guidance using virtual fixturesabstractIn this article, we describe and demonstrate control algorithms for general motion constraints. These constraints are designed to enhance the accuracy and speed of a user manipulating in an environment with the assistance of a cooperative or telerobotic system. Our method uses a basis of preferred directions, created off-line or in real-time using sensor data, to generate virtual fixtures that may constrain the user to a curve, surface, orientation, etc. in space. Open loop virtual fixtures seek only to maintain user motion along preferred directions, whereas closed loop fixtures additionally guide the user toward a point, line, or surface. This article demonstrates and compares the effects of open and closed loop fixtures in both autonomous and human-machine cases. Panadda Marayong, Ming Li 0052, Allison M. Okamura, Gregory D. Hager |
ICRA | 4 |
| 2003 | Functional reactive programming as a hybrid system frameworkabstractIn previous work we presented functional reactive programming (FRP), a general framework for designing hybrid systems and developing domain-specific languages for related domains. FRP's synchronous dataflow features, like event driven switching, supported by higher-order lazy functional abstractions of Haskell allows rapid development of modular and reusable specifications. In this paper, we look at more closely to the relation of arrowized FRP (AFRP), the FRP implementation, and formal specification of hybrid systems. We show how a formally specified hybrid system can be expressed in FRP and present a constructive proof showing that, for a subset of AFRP programs, there is a corresponding formal hybrid system specification. Izzet Pembeci, Gregory D. Hager |
ICRA | 2 |
| 2003 | VICs: A Modular Vision-Based HCI Framework
Guangqi Ye, Jason J. Corso, Darius Burschka, Gregory D. Hager |
ICVS | 4 |
| 2003 | V-GPS - image-based control for 3D guidance systemsabstractWe present our approach for pose verification with monocular cameras in 3-dimensional space based on the image-based control paradigm. We describe the extensions to our previous control system for mobile navigation that allow us to estimate the complete set of those parameters in space. The major contribution of this approach is a sensor-independent formulation that allows a flexible configuration with a variety of sensor systems including standard cameras, omnidirectional cameras and laser systems. Our second contribution is a way to re-initialize the tracked landmarks during a multi-segment navigation in applications with significant derivations from the pre-taught trajectory as it is the case for handheld systems and flying robots. The presented system can be used as a guidance system for visitors. The localization is based on known landmarks that are in our case natural landmarks in the environment. These landmarks correspond to the satellites of a GPS system. We call it V-GPS (vision-based GPS) because of this similarity in the concept. A camera carried by a person allows to navigate along pre-specified paths through environments, like galleries, hospitals, parks, and other public places. Darius Burschka, Gregory D. Hager |
IROS | 2 |
| 2003 | Handling discontinuities in stereovisual alignment tasksabstractFor more than a decade the capabilities of incompletely calibrated stereo systems have been formally characterized. These characterizations provide theoretical limits on the tasks that visually guided robotic systems can achieve. Within these limits there are several feature-alignment tasks whose image-space representations are inherently discontinuous, including collinearity and coplanarity. This paper presents smooth, locally convergent image-space approximations for such tasks. In addition, we propose the general notion of a "safe" image-based approximation to spatial tasks, in which complete image-space characterization is sacrificed in return for partial characterization along with local controllability. As a result, this article bridges the gap between theoretically specifiable tasks and locally controllable tasks for common stereovisual servoing systems. Zachary Dodds, Gregory D. Hager |
IROS | 2 |
| 2003 | Task modeling and specification for modular sensory based human-machine cooperative systemsabstractThis paper is directed towards developing human-machine cooperative systems (HCMS) for augmented surgical manipulation tasks. These tasks are commonly repetitive, sequential, and consist of simple steps. The transitions between these steps can be driven either by the surgeon's input or sensory information. Consequently, complex tasks can be effectively modeled using a set of basic primitives, where each primitive defines some basic type of motion (e.g. translational motion along a line, rotation about an axis, etc.). These steps can be "open-loop" (simply complying to user's demands) or "closed-loop, in which case external sensing is used to define a nominal reference trajectory. The particular research problem considered here is the development of a system that supports simple design of complex surgical procedures from a set of basic control primitives. The three system levels considered are: i) task graph generation which allows the user to easily design or model a task, ii) task graph execution which executes the task graph, and iii) at the lowest level, the specification of primitives which allows the user to easily specify new types of primitive motions. The system has been developed and validated using the JHU Steady Hand Robot as an experimental platform. Danica Kragic, Gregory D. Hager |
IROS | 2 |
| 2003 | Erratum: Human-Machine Collaborative Systems for Microsurgical Applications
Danica Kragic, Panadda Marayong, Ming Li 0052, Allison M. Okamura, Gregory D. Hager |
ISRR | 5 |
| 2003 | VisHap: augmented reality combining haptics and visionabstractHaptic devices have been successfully incorporated into the human-computer interaction model. However, a drawback common to almost all haptic systems is that the user must be attached to the haptic devices at all times even though force feedback is not always being rendered. This constant contact hinders perception of the virtual environment, primarily because it prevents the user from feeling new tactile sensations upon contact with virtual objects. We present the design and implementation of an augmented reality system called VisHap that uses visual tracking to seamlessly integrate force feedback with tactile feedback to generate a "complete" haptic experience. The VisHap framework allows the user to interact with combinations of virtual and real objects naturally, thereby combining active and passive haptics. An example application of this framework is also presented. The flexibility and extensibility of our framework is promising in that it supports many interaction modes and allows further integration with other augmented reality system. Guangqi Ye, Jason J. Corso, Gregory D. Hager, Allison M. Okamura |
SMC | 3 |
| 2003 | Recent Methods for Image-Based Modeling and RenderingabstractA long-standing goal in image-based modeling and rendering is to capture a scene from camera images and construct a sufficient model to allow photo-realistic rendering of new views. With the confluence of computer graphics and vision, the combination of research on recovering geometric structure from un-calibrated cameras with modeling and rendering has yielded numerous new methods. Yet, many challenging issues remain to be addressed before a sufficiently general and robust system could be built to (for instance) allow an average user to model their home and garden from camcorder video. This tutorial aims to give researchers and students in computer graphics a working knowledge of relevant theory and techniques covering the steps from real-time vision for tracking and the capture of scene geometry and appearance, to the efficient representation and real-time rendering of image-based models. It also includes hands-on demos of real-time visual tracking, modeling and rendering systems. Darius Burschka, Gregory D. Hager, Zachary Dodds, Martin Jägersand, Dana Cobzas, Keith Yerex |
VR | 2 |
| 2003 | Advances in Computational StereoabstractExtraction of three-dimensional structure of a scene from stereo images is a problem that has been studied by the computer vision community for decades. Early work focused on the fundamentals of image correspondence and stereo geometry. Stereo research has matured significantly throughout the years and many advances in computational stereo continue to be made, allowing stereo to be applied to new and more demanding problems. We review recent advances in computational stereo, focusing primarily on three important topics: correspondence methods, methods for occlusion, and real-time implementations. Throughout, we present tables that summarize and draw distinctions among key ideas and approaches. Where available, we provide comparative analyses and we make suggestions for analyses yet to be done. Myron Z. Brown, Darius Burschka, Gregory D. Hager |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Vision Assisted Control for Manipulation using Virtual Fixtures: Experiments at Macro and Micro ScalesabstractWe present the design and implementation of a vision-based system for micron-scale, cooperative manipulation of a surgical tool. The system is based on a control algorithm that implements a broad class of guidance modes called virtual fixtures. A virtual fixture, like a real fixture, limits the motion of a tool to a prescribed class or range. The implemented system uses vision as a sensor for providing a reference trajectory, and the control algorithm then provides haptic feedback involving direct, shared manipulation of a surgical tool. We have tested this system on the JHU Steady Hand robot and provide experimental results for path following and positioning on structures at both macroscopic and microscopic scales. Alessandro Bettini, Samuel Lang, Allison M. Okamura, Gregory D. Hager |
ICRA | 4 |
| 2002 | Stereo-Based Obstacle Avoidance in Indoor Environments with Active Sensor Re-CalibrationabstractWe present a stereo-based obstacle avoidance system for mobile vehicles. The system operates in three steps. First, it models the surface geometry of the supporting surface and removes the supporting surface from the scene. Next, it segments the remaining stereo disparities into connected components in image and disparity space. Finally, it projects the resulting connected components onto the supporting surface and plans a path around them. One interesting aspect of this system is that it can detect both positive and "negative" obstacles (e.g. stairways) in its path. The algorithms we have developed have been implemented on a mobile robot equipped with a real-time stereo system. We present experimental results on indoor environments with planar supporting surfaces that show the algorithms to be both fast and robust. Darius Burschka, Stephen Lee, Gregory D. Hager |
ICRA | 3 |
| 2002 | Specifying Behavior in C++abstractMost robot programming takes place in the "time domain", that is, the goal is to specify the behavior of a system that is acquiring a continual temporal stream of inputs, and is required to provide a continual, temporal stream of outputs. We present a reactive programming language, based on the functional reactive programming paradigm, for specifying such behavior. The major attributes of this language are: 1) it provides for both synchronous and asynchronous definitions of behavior; 2) specification is equational in nature; 3) it is type safe; and 4) it is embedded in C++. In particular the latter makes it simple to "lift" existing C++ libraries into the language. Xiangtian Dai, Gregory D. Hager, John Peterson |
ICRA | 2 |
| 2002 | Functional reactive robotics: an exercise in principled integration of domain-specific languagesabstractSoftware for (semi-) autonomous robots tends to be a complex combination of components from many different application domains such as control theory, vision, and artificial intelligence. Components are often developed using their own domain-specific tools and abstractions. System integration can thus be a significant challenge, in particular when the application calls for a dynamic, adaptable system structure in which rigid boundaries between the subsystems are a performance impediment. We believe that, by identifying suitably abstract notions common to the different domains in question, it is possible to create a broader framework for software integration and to recast existing domain-specific frameworks in these terms. This approach simplifies integration and leads to improved reliability. In this paper, we show how Functional Reactive Programming (FRP) can serve as such a unifying framework for programming vision-guided, semi-autonomous robots and illustrate the benefits this approach entails. The key abstractions in FRP, reactive components describing continuous or discrete behavior in a declarative style, are first class entities, allowing the resulting systems to exhibit a dynamic, adaptable structure which we regard as especially important in the area of autonomous robots. Izzet Pembeci, Henrik Nilsson, Gregory D. Hager |
PPDP | 3 |
| 2001 | Vision Based Control of Mobile RobotsabstractThis paper presents an approach for direct control of a mobile robot to keep it on a pre-taught path based solely on the perception from a monocular CCD camera. In particular, we present a novel vision-based control algorithm for mobile systems equipped with a conventional camera and a pan-tilt head or with an omnidirection camera. This algorithm avoids numerical instabilities of previously reported approaches. The experimental performance of the method as well as its practical limitations are discussed. Darius Burschka, Gregory D. Hager |
ICRA | 2 |
| 2001 | Vision assisted control for manipulation using virtual fixturesabstractThe "steady hand" concept is a way of providing assistance for direct manipulation by applying constraints on the motion of a tool shared by a user and a robot. We explore in detail one family of constraints: virtual fixtures for use in path following tasks. Vision is used to sense the desired path, and then the robot encourages motion toward and along the path through a direction-based control law. This "soft" virtual fixture allows the user to move in other, non-preferred directions, maintaining the user's sense of autonomy and control. Experimental results show that user performance in assisted path following improves with virtual fixture augmentation, and differs with varying fixture compliance. Alessandro Bettini, Samuel Lang, Allison M. Okamura, Gregory D. Hager |
IROS | 4 |
| 2001 | Performance Evaluation of a Cooperative Manipulation Microsurgical Assistant Robot Applied to Stapedotomy
Peter J. Berkelman, Daniel L. Rothbaum, Jaydeep Roy, Samuel Lang, Louis L. Whitcomb, Gregory D. Hager, Patrick S. Jensen, Eugene de Juan, Russell H. Taylor, John K. Niparko |
MICCAI | 6 |
| 2001 | Applications of Task-Level Augmentation for Cooperative Fine Manipulation Tasks in Surgery
Rajesh Kumar 0001, Aaron C. Barnes, Gregory D. Hager, Patrick S. Jensen, Russell H. Taylor |
MICCAI | 3 |
| 2001 | FVision: A Declarative Language for Visual Tracking
John Peterson, Paul Hudak, Alastair Reid 0001, Gregory D. Hager |
PADL | 4 |
| 2001 | Probabilistic Data Association Methods for Tracking Complex Visual ObjectsabstractWe describe a framework that explicitly reasons about data association to improve tracking performance in many difficult visual environments. A hierarchy of tracking strategies results from ascribing ambiguous or missing data to: 1) noise-like visual occurrences, 2) persistent, known scene elements (i.e., other tracked objects), or 3) persistent, unknown scene elements. First, we introduce a randomized tracking algorithm adapted from an existing probabilistic data association filter (PDAF) that is resistant to clutter and follows agile motion. The algorithm is applied to three different tracking modalities-homogeneous regions, textured regions, and snakes-and extensibly defined for straightforward inclusion of other methods. Second, we add the capacity to track multiple objects by adapting to vision a joint PDAF which oversees correspondence choices between same-modality trackers and image features. We then derive a related technique that allows mixed tracker modalities and handles object overlaps robustly. Finally, we represent complex objects as conjunctions of cues that are diverse both geometrically (e.g., parts) and qualitatively (e.g., attributes). Rigid and hinge constraints between part trackers and multiple descriptive attributes for individual parts render the whole object more distinctive, reducing susceptibility to mistracking. Results are given for diverse objects such as people, microscopic cells, and chess pieces. Christopher Rasmussen, Gregory D. Hager |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | On Specifying and Performing Visual Tasks with Qualitative Object ModelsabstractVision-based control has aimed to develop general-purpose, high accuracy systems for manipulating objects. While much of the scientific and technological infrastructure needed to accomplish this aim is now in place, several stumbling blocks still remain. One continuing issue is accuracy, and its relationship to system calibration. We describe a generative task structure for vision-based control of motion that admits a simple, geometric approach to task specification. At the same time, this approach allows one to state precisely what types of miscalibration lead to errors in task performance. A second hurdle has been the programmability of hand-eye systems. However, we argue that a structured object representation sufficient for flexible hand-eye coordination is a possibility. The result is a high-level, object-centered language for expressing hand-eye tasks. Gregory D. Hager, Zachary Dodds |
ICRA | 1 |
| 2000 | An Augmentation System for Fine Manipulation
Rajesh Kumar 0001, Gregory D. Hager, Aaron C. Barnes, Patrick S. Jensen, Russell H. Taylor |
MICCAI | 2 |
| 2000 | Fast and Globally Convergent Pose Estimation from Video ImagesabstractDetermining the rigid transformation relating 2D images to known 3D geometry is a classical problem in photogrammetry and computer vision. Heretofore, the best methods for solving the problem have relied on iterative optimization methods which cannot be proven to converge and/or which do not effectively account for the orthonormal structure of rotation matrices. We show that the pose estimation problem can be formulated as that of minimizing an error metric based on collinearity in object (as opposed to image) space. Using object space collinearity error, we derive an iterative algorithm which directly computes orthogonal rotation matrices and which is globally convergent. Experimentally, we show that the method is computationally efficient, that it is no less accurate than the best currently employed optimization methods, and that it outperforms all tested methods in robustness to outliers. Chien-Ping Lu, Gregory D. Hager, Eric Mjolsness |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Fast 3D Boundary Computation from Occluding Contour MotionabstractPresents a fast method for computing a bounding volume within which an observed object must lie from the observed motion of the occluding contour during a straight-line motion of the camera. The bounding volume is represented as a set of planar cross sections each consisting of multiple convex polygons within which the object lies. The algorithm's worst-case runtime performance is O(nmk) operations, where n is the number of viewpoints used, m is the number of polygons created during the execution of the algorithm, and k is a parameter dependent on the geometric complexity of the object being viewed. Experimental examples are demonstrated; bounding polygon computation from sampled contours in 20 images required less than one second on a 50 MHz i486 CPU. Aage Bendiksen, Gregory D. Hager |
ICRA | 2 |
| 1999 | Task Specification and Monitoring for Uncalibrated Hand/Eye CoordinationabstractMost of the work in robotic manipulation and visual servoing has emphasized how to specify and perform particular tasks. Recent results have formally shown what tasks are possible with uncalibrated imaging systems. This paper extends those results by characterizing in a constructive manner the set of tasks which can be performed with different types of uncalibrated camera models. The tasks resulting structure provides a principle foundation both for a specification language and for automatic execution monitoring in uncalibrated environments. Zachary Dodds, Gregory D. Hager, A. Stephen Morse, João Pedro Hespanha |
ICRA | 2 |
| 1999 | Model-Based 3D Object Tracking Using Projective InvarianceabstractThis paper describes a method of 3D object tracking by projective invariance. Usually, projective invariance has been used for object recognition and its property discards the need for cumbersome camera calibration to obtain metric information. We focus on the fact the projective invariance is not only regarded as recognition cues but also able to verify, the status of a 3D object for visual tracking. With projective invariance relevant to points, we develop a fast and reliable visual tracking strategy by combining a recognition module using projective invariance, window-based corner trackers for object's center tracking and a visibility checking module. The proposed method is experimented in single IBM-PC platform successfully. Sung-Woo Lee, Bum-Jae You, Gregory D. Hager |
ICRA | 3 |
| 1999 | A Language for Declarative Robotic ProgrammingabstractWe have applied methodologies developed for domain-specific embedded languages to create a high-level robot control language called Frob, for functional robotics. Frob supports a programming style that cleanly separates the what from the how of a robotic control program. That is, the what is a simple, easily understood definition of the control strategy using groups of equations and primitives which combine sets of these control system equations into a complex system. The how aspect of the program addresses the unpleasant details, such as the method used to realize these equations, the connection between the control equations and the sensors and effectors in the robot, and communication with other elements of the system. Frob is a system that supports rapid prototyping of new control strategies, enables software reuse through composition, and defines a system in a way that can be formally reasoned about and transformed. John Peterson, Gregory D. Hager, Paul Hudak |
ICRA | 2 |
| 1999 | Prototyping Real-Time Vision Systems: An Experiment in DSL DesignabstractDescribes the enhancement of XVision, a large library of C++ code for real-time vision processing, into FVision (pronounced fission), a fully-featured domain-specific language (DSL) embedded in Haskell. The resulting prototype system substantiates the claims of increased modularity, effective code reuse and rapid prototyping that characterize the DSL approach to systems design. It also illustrates the need for judicious interface design: relegating computationally expensive tasks to XVision (pre-existing C++ components) and leaving modular compositional tasks to FVision (Haskell). At the same time, our experience demonstrates how Haskell's advanced language features (specifically, parametric polymorphism, lazy evaluation, higher-order functions and automatic storage reclamation) permit a rapid DSL design that is itself highly modular and easily modified. Overall, the resulting hybrid system exceeded our expectations: visual tracking programs continue to spend most of their time executing low-level image processing code, while Haskell's advanced features allow us to quickly develop and test small prototype systems within a matter of a few days, and to develop realistic applications within a few weeks. Alastair Reid 0001, John Peterson, Gregory D. Hager, Paul Hudak |
ICSE | 3 |
| 1999 | A Hierarchical Vision Architecture for Robotic Manipulation Tasks
Zachary Dodds, Martin Jägersand, Gregory D. Hager, Kentaro Toyama |
ICVS | 3 |
| 1999 | Three-dimensional pose determination for a humanoid robot using binocular head systemabstractThere has been much interests on three-dimensional pose estimation of an object since there are a lot of robotic applications such as intelligent robotic assembly and/or vision-guided control of a humanoid or human-friendly robot. In this paper, there are proposed a simple and fast stereo matching algorithm for real-time robotic applications and an improved scheme for feature matching using three-dimensional information by adopting backprojection as a feedback element measuring the degree of feature matching. The degree of feature matching is determined by checking existence ratio of each feature point in real images after back-projecting an object model into image plane. The proposed pose determination algorithm is applied for a humanoid robot in KIST, whose name is CENTAUR, successfully and determine the pose of several polyhedral objects using 3D information of vertexes on the outline of an object in image plane. Hong-Jae Kim, Bum-Jae You, Gregory D. Hager, Sang-Rok Oh, Chong-Won Lee |
IROS | 3 |
| 1999 | Computational Vision at Yale
Peter N. Belhumeur, James S. Duncan, Gregory D. Hager, Drew McDermott, A. Stephen Morse, Steven W. Zucker |
Int. J. Comput. Vis. | 3 |
| 1999 | What Tasks can be Performed with an Uncalibrated Stereo Vision System?
João Pedro Hespanha, Zachary Dodds, Gregory D. Hager, A. Stephen Morse |
Int. J. Comput. Vis. | 3 |
| 1999 | Incremental Focus of Attention for Robust Vision-Based Tracking
Kentaro Toyama, Gregory D. Hager |
Int. J. Comput. Vis. | 2 |
| 1999 | Tracking in 3D: Image Variability Decomposition for Recovering Object Pose and Illumination
Peter N. Belhumeur, Gregory D. Hager |
Pattern Anal. Appl. | 2 |
| 1998 | Joint Probabilistic Techniques for Tracking Multi-Part ObjectsabstractCommon objects such as people and cars comprise many visual parts and attributes, yet image-based tracking algorithms are often keyed to only one of a target's identifying characteristics. In this paper, we present a framework for combining and sharing information among several state estimation processes operating on the same underlying visual object. Well-known techniques for joint probabilistic data association are adapted to yield increased robustness when multiple trackers attuned to disparate visual cues are deployed simultaneously. We also formulate a measure of tracker confidence, based on distinctiveness and occlusion probability, which permits the deactivation of trackers before erroneous state estimates adversely affect the ensemble. We discuss experiments focusing on color-region- and snake-based tracking that demonstrate the efficacy of this approach. Christopher Rasmussen, Gregory D. Hager |
CVPR | 2 |
| 1998 | What Can Be Done with an Uncalibrated Stereo System ?abstractOver the last several years, there has been an increasing appreciation of the impact of control architecture on the accuracy of visual servoing systems. In particular, it is generally acknowledged that so-called image-based methods provide the highest guarantees of accuracy on inaccurately calibrated hand-eye systems. Less clear is the impact of the control architecture on the set of tasks which the system can perform. In this article, we present a formal analysis of control architectures for hand-eye coordination. Specifically, we first state a formal characterization of what makes a task performable under three possible encoding methods. Then, for the specific case of cameras modeled using projective geometry, we relate this characterization to notions of projective invariance and demonstrate the limits of achievable performance in this regard. João Pedro Hespanha, Zachary Dodds, Gregory D. Hager, A. Stephen Morse |
ICRA | 3 |
| 1998 | Dynamic Sensor Planning in Visual ServoingabstractWe present an approach to dynamic sensor planning problems in visual servoing. Specifically, one of the main problems an image-based visual servoing is to plan the camera trajectory in order to avoid undesired configurations (e.g., features out of view, collision with obstacles, etc.). Our approach uses the robot redundancy and employs a control scheme based on the task function approach. It combines the regulation of the selected vision-based task with the minimization of a secondary cost function, which reflects given constraints on the manipulator trajectory. We describe how this methodology is applied to common problems in robotic vision: occlusion avoidance, field of view constraint and obstacle avoidance. We demonstrate the validity of this approach with various experiments. Éric Marchand, Gregory D. Hager |
ICRA | 2 |
| 1998 | Joint probabilistic techniques for tracking objects using multiple visual cuesabstractRobots relying on vision as a primary sensor frequently need to track common objects such as people, cars, and tools in order to successfully perform autonomous navigation or grasping tasks. These objects may comprise many visual parts and attributes, yet image-based tracking algorithms are often keyed to only one of a target's identifying characteristics. In this paper, we present a framework for sharing information among disparate state estimation processes operating on the same underlying visual object. Well-known techniques for joint probabilistic data association are adapted to yield increased robustness when multiple trackers attuned to different visual cues are deployed simultaneously. We also formulate a measure of tracker confidence, based on distinctiveness and occlusion probability, which permits the deactivation of trackers before erroneous state estimates adversely affect the ensemble. We will discuss experiments using color-region- and snake-based tracking in tandem that demonstrate the efficacy of this approach. Christopher Rasmussen, Gregory D. Hager |
IROS | 2 |
| 1998 | X Vision: A Portable Substrate for Real-Time Vision ApplicationsabstractIn the past several years, the speed of standard processors has reached the point where interesting problems requiring visual tracking can be carried out on standard workstations. However, relatively little attention has been devoted to developing visual tracking technology in its own right. In this article, we describe X Vision, a modular, portable framework for visual tracking. X Vision is designed to be a programming environment for real-time vision which provides high performance on standard workstations outfitted with a simple digitizer. X Vision consists of a small set of image-level tracking primitives, and a framework for combining tracking primitives to form complex tracking systems. Efficiency and robustness are achieved by propagating geometric and temporal constraints to the feature detection level, where image warping and specialized image processing are combined to perform feature detection quickly and robustly. Over the past several years, we have used X Vision to construct several vision-based systems. We present some of these applications as an illustration of how useful, robust tracking systems can be constructed by simple combinations of a few basic primitives combined with the appropriate task-specific constraints. Gregory D. Hager, Kentaro Toyama |
Comput. Vis. Image Underst. | 1 |
| 1998 | Efficient Region Tracking With Parametric Models of Geometry and IlluminationabstractAs an object moves through the field of view of a camera, the images of the object may change dramatically. This is not simply due to the translation of the object across the image plane; complications arise due to the fact that the object undergoes changes in pose relative to the viewing camera, in illumination relative to light sources, and may even become partially or fully occluded. We develop an efficient general framework for object tracking, which addresses each of these complications. We first develop a computationally efficient method for handling the geometric distortions produced by changes in pose. We then combine geometry and illumination into an algorithm that tracks large image regions using no more computation than would be required to track with no accommodation for illumination changes. Finally, we augment these methods with techniques from robust statistics and treat occluded regions on the object as statistical outliers. Experimental results are given to demonstrate the effectiveness of our methods. Gregory D. Hager, Peter N. Belhumeur |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Image-based prediction of landmark features for mobile robot navigationabstractWe have been developing an architecture for vision-based navigation which relies on continuous feedback from visual "landmarks" to control robot motion, In this approach, landmarks are consistently located and acquired as they come into view. To make this process efficient and robust, it is important that the image locations of these features can be predicted from available image information. In this article, we discuss methods for direct image-based prediction of point and line features for a mobile system operating on a planar surface. Preliminary experimental results suggest that image-based prediction con be performed efficiently and with sufficient accuracy to ensure robust acquisition of navigational landmarks. Gregory D. Hager, David J. Kriegman, Erliang Yeh, Christopher Rasmussen |
ICRA | 1 |
| 1997 | A modular system for robust positioning using feedback from stereo visionabstractThis paper introduces a modular framework for robot motion control using stereo vision. The approach is based on a small number of generic motion control operations referred to as primitive skills. Each primitive skill uses visual feedback to enforce a specific task-space kinematic constraint between a robot end-effector and a set of target features. By observing both the end-effector and target features, primitive skills are able to position with an accuracy that is independent of errors in hand-eye calibration. Furthermore, primitive skills are easily combined to form more complex kinematic constraints as required by different applications. These control laws have been integrated into a system that performs tracking and control on a single processor at real-time rates. Experiments with this system have shown that it is extremely accurate, and that it is insensitive to camera calibration error. The system has been applied to a number of example problems, showing that modular, high precision, vision-based motion control is easily achieved with off-the-shelf hardware. Gregory D. Hager |
IEEE Trans. Robotics Autom. | 1 |
| 1996 | Real-time tracking of image regions with changes in geometry and illuminationabstractHistorically, SSD or correlation-based visual tracking algorithms have been sensitive to changes in illumination and shading across the target region. This paper describes methods for implementing SSD tracking that is both insensitive to illumination variations and computationally efficient. We first describe a vector-space formulation of the tracking problem, showing how to recover geometric deformations. We then show that the same vector space formulation can be used to account for changes in illumination. We combine geometry and illumination into an algorithm that tracks large image regions on live video sequences using no more computation than would be required to trade with no accommodation for illumination changes. We present experimental results which compare the performance of SSD tracking with and without illumination compensation. Gregory D. Hager, Peter N. Belhumeur |
CVPR | 1 |
| 1996 | Incremental Focus of Attention for Robust Visual TrackingabstractWe present the Incremental Focus of Attention (IFA) architecture for adding robustness to software-based, real-time, motion trackers. The framework provides a structure which, when given the entire camera image to search, efficiently focuses the attention of the system into a narrow set of possible slates that includes the target state. IFA offers a means for automatic tracking initialization and reinitialization when environmental conditions momentarily deteriorate and cause the system to lose track of its target. Systems based on the framework degrade gracefully as various assumptions about the environment are violated. In particular, multiple tracking algorithms are layered so that the failure of a single algorithm causes another algorithm of less precision to take over, thereby allowing the system to return approximate feature state information. Kentaro Toyama, Gregory D. Hager |
CVPR | 2 |
| 1996 | X Vision: Combining Image Warping and Geometric Constraints for Fast Visual Tracking
Gregory D. Hager, Kentaro Toyama |
ECCV (2) | 1 |
| 1996 | Servomatic: a modular system for robust positioning using stereo visual servoingabstractWe introduce Servomatic, a modular system for robot motion control based on calibration-insensitive visual servoing. A small number of generic motion control operations referred to as primitive skills use stereo visual feedback to enforce a specific task-space kinematic constraint between a robot end-effector and a set of target features. Primitive skills are able to position with an accuracy that is independent of errors in hand-eye calibration and are easily combined to form more complex kinematic constraints as required by different applications. The system has been applied to a number of example problems, showing that modular, high precision, vision-based motion control is easily achieved with off-the-shelf hardware. Our continuing goal is to develop a system where low-level robot control ceases to be a concern to higher-level robotics researchers. Kentaro Toyama, Gregory D. Hager, Jonathan G. Wang |
ICRA | 2 |
| 1996 | Preliminary results on grasping with vision and touchabstractThis paper presents initial results in integrating touch with vision for delicate manipulation tasks. A generalizable framework of behavioral primitives for tactile and visual feedback control is proposed. Since vision provides position and shape information at a distance, while tactile provides small-scale geometric and force information, we focus on the complimentary roles of vision and touch. We demonstrate that visual feedback can perform the rough positioning needed for tactile sensor feedback, and that grasp force and object orientation can be sensed and controlled with tactile sensing. A force sensor based approach provides a comparison measure, and we observe that the use of tactile sensing results in a more gentle grasp. Jae S. Son, Robert D. Howe, Jonathan G. Wang, Gregory D. Hager |
IROS | 4 |
| 1996 | A tutorial on visual servo controlabstractThis article provides a tutorial introduction to visual servo control of robotic manipulators. Since the topic spans many disciplines our goal is limited to providing a basic conceptual framework. We begin by reviewing the prerequisite topics from robotics and computer vision, including a brief review of coordinate transformations, velocity representation, and a description of the geometric aspects of the image formation process. We then present a taxonomy of visual servo control systems. The two major classes of systems, position-based and image-based systems, are then discussed in detail. Since any visual servo system must be capable of tracking image features in a sequence of images, we also include an overview of feature-based and correlation-based methods for tracking. We conclude the tutorial with a number of observations on the current directions of the research field of visual servo control. Seth Hutchinson 0001, Gregory D. Hager, Peter I. Corke |
IEEE Trans. Robotics Autom. | 2 |
| 1995 | Calibration-Free Visual Control Using Projective InvarianceabstractMuch of the previous work on hand-eye coordination has emphasized the reconstructive aspects of vision. Recently, techniques that avoid explicit reconstruction by placing visual feedback into a control loop have been developed. When properly defined, these methods lead to calibration insensitive hand-eye coordination. Recent work on projective geometry as applied to vision is used to extend this paradigm in two ways. First, it is shown how results from projective geometry can be used to perform online calibration. Second, results on projective invariance are used to define setpoints for visual control that are independent of viewing location. These ideas are illustrated through a number of examples and have been tested on an implemented system.> Gregory D. Hager |
ICCV | 1 |
| 1995 | A "robust" convergent visual servoing systemabstractThis paper describes a simple visual servoing control algorithm capable of robustly positioning a three degree of freedom end effector based only on information from a stereo vision system. The proposed control algorithm does not require estimates of the gripper's spatial position, a significant source of calibration sensitivity. The controller is completely immune to positional camera calibration errors, and we demonstrate robustness to orientation miscalibration through a series of simulations and experiments. Alfred A. Rizzi, Gregory D. Hager, Daniel E. Koditschek |
IROS (1) | 3 |
| 1995 | Keeping your eye on the ball: tracking occluding contours of unfamiliar objects without distractionabstractVisual tracking is prone to distractions, where features similar to the target features guide the track away from its intended object. Global shape models and dynamic models are necessary for completely distraction-free contour tracking, but there are cases when component feature trackers alone can be expected to avoid distraction. We define the tracking problem in general and devise a method for local, window-based, feature trackers to track accurately in spite of background distractions. The algorithm is applied to a generic line tracker and a snake-like contour tracker which are then analyzed with respect to previous contour-trackers. We discuss the advantages and disadvantages of our approach and suggest that existing model-based trackers can be improved by incorporating similar techniques at the local level. Kentaro Toyama, Gregory D. Hager |
IROS (1) | 2 |
| 1994 | Real-time feature tracking and projective invariance as a basis for hand-eye coordinationabstractThis article presents a methodology for using stereo visual feedback to perform manipulation tasks. The two major innovations in the approach are: 1) the use of feature-based tracking methods that perform in real-time on standard workstations without specialized hardware; and 2) the use of closed-loop feedback control based on projective invariants to make positioning accuracy independent of hand-eye calibration error. Particular attention is given to the feature tracking component of the system. The feature tracker is a programming environment that supports a variety of low-level detection methods (basic features), and features defined in terms of other features (composite features). Basic and composite features are combined into feature networks. Experimental results from two feature networks are presented. One computes corresponding epipolar lines using eight corresponding features in two images. The second computes a visual trajectory for a visual servoing system.> Gregory D. Hager |
CVPR | 1 |
| 1994 | A Vision-Based Grasping System for Unfamiliar Planar ObjectsabstractThis paper describes a vision based robotic system for grasping unfamiliar planar objects with a parallel jaw gripper. The input to the grasp synthesis system consists of projected edge data from a single camera video system calibrated to a planar working surface. A simple search procedure is used to find an acceptable grasp. Our grasp analysis metric is the required squeezing force, computed using an approximation of the rigid body equilibrium conditions. The resulting analysis conditions are formulated as a linear program and solved using the simplex method. A Zebra-ZERO robot arm with a parallel-jaw gripper is used to execute the chosen grasp. In a simple test of the system, it successfully grasped and lifted 10 of 14 test objects.> Aage Bendiksen, Gregory D. Hager |
ICRA | 2 |
| 1994 | Robot Feedback Control Based on Stereo Vision: Towards Calibration-Free Hand-Eye CoordinationabstractThis article describes the theory and implementation of a system that positions a robot manipulator using visual information from two cameras. The system simultaneously tracks the robot end-effector and visual features used to define goal positions. An error signal based on the visual distance between the end-effector and the target as defined and a control law that moves the robot to drive this error to zero is derived. The control law has been integrated into a system that performs tracking and stereo control on a single processor with no special purpose hardware at real-time rates. Experiments with the system have shown that the controller is so robust to calibration error that the cameras can be moved several centimeters and rotated several degrees while the system is running with no adverse effects.> Gregory D. Hager, Wen-Chung Chang, A. Stephen Morse |
ICRA | 1 |
| 1994 | Feature-based visual servoing and its application to teleroboticsabstractAdvances in visual servoing theory and practice now make it possible to accurately and robustly position a robot manipulator relative to a target. Both the vision and control algorithms are extremely simple, however they must be initialized on task-relevant features in order to be applied. Consequently, they are particularly well-suited to telerobotics systems where an operator can initialize the system but round-trip delay prohibits direct operator feedback during motion. This paper describes the basic theory behind feature-based visual servoing, and discusses the issues involved in integrating visual servoing into the ROTEX space teleoperation system.> Gregory D. Hager, Gerhard Grunwald, Gerd Hirzinger |
IROS | 1 |
| 1994 | Task-directed computation of qualitative decisions from sensor dataabstractDescribes a novel approach to sensor-based decision making based on formulating and solving large systems of parametric constraints. The constraints describe both a model for sensor data and the criteria for correct decisions about the data. An incremental constraint solving technique that performs decision-directed model recovery is developed. This method is straightforward to apply, is easily parallelized, and convergence can be demonstrated under very reasonable structural and statistical assumptions. This approach is demonstrated on several different decision-making problems involving manipulation and categorization of objects observed with a range scanner. The experiments indicate that simultaneous solution of both model constraints and decision criteria can lead to efficient and effective decision making, even when the observed data does not strongly determine a data model.> Gregory D. Hager |
IEEE Trans. Robotics Autom. | 1 |
| 1993 | Real-time vision-based robot localizationabstractThis paper describes an algorithm for determining robot location from visual landmarks. This algorithm determines both the correspondence between observed landmarks (in this case vertical edges in the environment) and a stored map, and computes the location of the robot using those correspondences. The primary advantages of this algorithm are its use of a single geometric tolerance to describe observation error, its ability to recognize ambiguous sets of correspondences, its ability to compute bounds on the error in localization, and fast execution. The algorithm has been implemented and tested on a mobile robot system. In several hundred trials it has never failed, and computes location accurate to within a centimeter in less than 0.5 s.> Sami Atiya, Gregory D. Hager |
IEEE Trans. Robotics Autom. | 2 |
| 1992 | Constraint solving methods and sensor-based decision-makingabstractThe author describes a novel approach to sensor-based decision-making that involves formulating and solving large systems of parametric constraints. The constraints describe a model for sensor data and the criteria for correct decisions about the data. An incremental constraint solving technique performs the minimal model recovery required to reach a decision. The approach was demonstrated on two different problems, graspability and categorization, using range data and a superellipsoid data model. The experiments indicated that simultaneous solution of both data constraints and decision criteria can lead to be efficient and effective decision-making. even when the observed data was imprecise and incomplete.> Gregory D. Hager |
ICRA | 1 |
| 1991 | Real-time vision-based robot localizationabstractAn algorithm for robot localization using visual landmarks is described. This algorithm determines both the correspondence between observed landmarks (in this case vertical edges in the environment) and a preloaded map, and the location of the robot from those correspondences. The primary advantages of this algorithm are its use of a single geometric tolerance to describe observation error, its ability to recognize ambiguous sets of correspondences, its ability to compute bounds on the error in localization, and its fast execution. The current version of the algorithm has been implemented and tested on a mobile robot system. In several hundred trials the algorithm has not failed, and computes location accurate to within a centimeter in less than half a second.> Sami Aitya, Gregory D. Hager |
ICRA | 2 |
| 1991 | Towards geometric decision making in unstructured environmentsabstractPresents an approach to sensor-based decision making in unstructured environments that relies on describing geometric structures by parameterized volumes. This approach leads to large systems of nonlinear stochastic inequalities. The author describes how these inequalities can be solved using interval bisection, discusses the structural and statistical convergence of the technique, and presents some preliminary experimental results.> Gregory D. Hager |
IROS | 1 |
| 1989 | Task-directed multisensor fusionabstractThe authors consider the problem of task-directed information gathering. They first develop a decision-theoretic model of task-directed sensing. In this framework, sensors are modeled as noise-contaminated, uncertain measurement systems. A sensor task is modelled as consisting of a function describing the type of information required by the task, a utility function describing sensitivity to error, and a cost function describing time or resource constraints on the system. From this description, the authors develop a computational method approximating a standard Bayesian decision-making model. This algorithm, which relies on a finite-element computation, is applicable to a wide variety of sensor fusion problems. The authors describe its derivation, analyze its error properties, and indicate how it can be made robust to errors in the description of sensors and discrepancies between geometric models and sensed objects. They also present the result of applying this fusion technique to several different information gathering tasks in simulated situations and in a distributed sensing system.> Gregory D. Hager, Max Mintz |
ICRA | 1 |
| 1988 | Egomotion And The Stabilized WorldabstractThis paper formulates the tasks of moving-object detection and motion-based depth recovery as a problem of sensor fusion in the presence of uncertainty. We utilize two sensor systems, one providing information about the local image velocity (spsecifying a point in image-velocity space), and the othter providing information about the camera motion (specifying a line segment in image-velocity space). 1Ne utilize Mahalanobis distance as a threshold rule for determining consistency between the measurements from the two sensor systems. We also suggest a framework for using the resulting segmented flow field to update estimates of the egomotion parameters. David J. Heeger, Gregory D. Hager |
ICCV | 2 |
| 1988 | Estimation procedures for robust sensor control
Gregory D. Hager, Max Mintz |
Int. J. Approx. Reason. | 1 |
| 1987 | Estimation Procedures for Robust Sensor Control
Gregory D. Hager, Max Mintz |
UAI | 1 |
| 1986 | Information and multi-sensor coordination
Gregory D. Hager, Hugh F. Durrant-Whyte |
UAI | 1 |