Martial Hebert

dblp:h/MartialHebert · DBLP profile ↗
← Back
262ranked-venue papers
10as first author
19since 2021 · last 2025
0000-0003-4566-5930ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 240 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 149 · 2 first-author · 9 since 2021Systems, architecture and hardware · 65 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 ReferevErything: Towards Segmenting Everything we can Speak of in Videos
abstract
We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data by fine-tuning them on small-scale Referring Object Segmentation datasets. Our key insight is to preserve the entirety of the generative model's architecture by shifting its objective from predicting noise to predicting mask latents. The resulting model can accurately segment rare and unseen objects, despite only being trained on a limited set of categories. Additionally, it can effortlessly generalize to non-object dynamic concepts, such as smoke or raindrops, as demonstrated in our new benchmark for Referring Video Process Segmentation (Ref-VPS). REM performs on par with the state-of-the-art on in-domain datasets, like Ref-DAVIS, while outperforming them by up to 12 IoU points out-of-domain, leveraging the power of generative pre-training. We also show that advancements in video generation directly improve segmentation.
Anurag Bagchi, Zhipeng Bao, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert
ICCV5
2025 Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
abstract
Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing them either solely for off-the-shelf data augmentation or as mere feature extractors. In contrast to these isolated and thus sub-optimal efforts, we introduce an integrated, versatile, diffusion-based framework, Diff-2-in-1, that can simultaneously handle both multi-modal data generation and dense visual perception, through a unique exploitation of the diffusion-denoising process. Within this framework, we further enhance discriminative visual perception via multi-modal generation, by utilizing the denoising network to create multi-modal data that mirror the distribution of the original training set. Importantly, Diff-2-in-1 optimizes the utilization of the created diverse and faithful data by leveraging a novel self-improving learning mechanism. Comprehensive experimental evaluations validate the effectiveness of our framework, showcasing consistent performance improvements across various discriminative backbones and high-quality multi-modal data generation characterized by both realism and usefulness. Our project website is available at https://zsh2000.github.io/diff-2-in-1.github.io/.
Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang
ICLR4
2024 Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
abstract
Complex 3D scene understanding has gained increasing attention, with scene encoding strategies built on top of visual foundation models playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present the first comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image, video, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluation yields key intriguing findings: Unsupervised image foundation models demonstrate superior overall performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, language-pretrained models show unexpected limitations in language-related tasks, and the mixture-of-vision-expert (MoVE) strategy leads to consistent performance improvement. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene understanding tasks.
Yunze Man, Shuhong Zheng, Zhipeng Bao, Martial Hebert, Liangyan Gui, Yu-Xiong Wang
NeurIPS4
2023 Object Discovery from Motion-Guided Tokens
abstract
Object discovery – separating objects from the background without manual labels – is a fundamental open challenge in computer vision. Previous methods struggle to go beyond clustering of low-level cues, whether handcrafted (e.g., color, texture) or learned (e.g., from auto-encoders). In this work, we augment the auto-encoder representation learning framework with two key components: motion-guidance and mid-level feature tokenization. Although both have been separately investigated, we introduce a new transformer decoder showing that their benefits can compound thanks to motion-guided vector quantization. We show that our architecture effectively leverages the synergy between motion and tokenization, improving upon the state of the art on both synthetic and real datasets. Our approach enables the emergence of interpretable object-specific mid-level features, demonstrating the benefits of motion-guidance (no labeling) and quantization (interpretability, memory efficiency).
Zhipeng Bao, Pavel Tokmakov, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert
CVPR5
2023 Multi-task View Synthesis with Neural Radiance Fields
abstract
Multi-task visual learning is a critical aspect of computer vision. Current research, however, predominantly concentrates on the multi-task dense prediction setting, which overlooks the intrinsic 3D world and its multi-view consistent structures, and lacks the capability for versatile imagination. In response to these limitations, we present a novel problem setting – multi-task view synthesis (MTVS), which reinterprets multi-task prediction as a set of novel-view synthesis tasks for multiple scene properties, including RGB. To tackle the MTVS problem, we propose MuvieNeRF, a framework that incorporates both multi-task and cross-view knowledge to simultaneously synthesize multiple scene properties. MuvieNeRF integrates two key modules, the Cross-Task Attention (CTA) and Cross-View Attention (CVA) modules, enabling the efficient use of information across multiple views and tasks. Extensive evaluation on both synthetic and realistic benchmarks demonstrates that MuvieNeRF is capable of simultaneously synthesizing different scene properties with promising visual quality, even outperforming conventional discriminative models in various settings. Notably, we show that MuvieNeRF exhibits universal applicability across a range of NeRF backbones. Our code is available at https://github.com/zsh2000/MuvieNeRF.
Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang
ICCV3
2023 Discovering Multiple Algorithm Configurations
abstract
Many practitioners in robotics regularly depend on classic, hand-designed algorithms. Often the performance of these algorithms is tuned across a dataset of annotated examples which represent typical deployment conditions. Automatic tuning of these settings is traditionally known as algorithm configuration. In this work, we extend algorithm configuration to automatically discover multiple modes in the tuning dataset. Unlike prior work, these configuration modes represent multiple dataset instances and are detected automatically during the course of optimization. We propose three methods for mode discovery: a post hoc method, a multistage method, and an online algorithm using a multi-armed bandit. Our results characterize these methods on synthetic test functions and in multiple robotics application domains: stereoscopic depth estimation, differentiable rendering, motion planning, and visual odometry. We show the clear benefits of detecting multiple modes in algorithm configuration space.
Leonid Keselman, Martial Hebert
ICRA2
2023 Optimizing Algorithms from Pairwise User Preferences
abstract
Typical black-box optimization approaches in robotics focus on learning from metric scores. However, that is not always possible, as not all developers have ground truth available. Learning appropriate robot behavior in human-centric contexts often requires querying users, who typically cannot provide precise metric scores. Existing approaches leverage human feedback in an attempt to model an implicit reward function; however, this reward may be difficult or impossible to effectively capture. In this work, we introduce SortCMA to optimize algorithm parameter configurations in high dimensions based on pairwise user preferences. SortCMA efficiently and robustly leverages user input to find parameter sets without directly modeling a reward. We apply this method to tuning a commercial depth sensor without ground truth, and to robot social navigation, which involves highly complex preferences over robot behavior. We show that our method succeeds in optimizing for the user's goals and perform a user study to evaluate social navigation results.
Leonid Keselman, Katherine Shih, Martial Hebert, Aaron Steinfeld
IROS3
2023 Beyond RGB: Scene-Property Synthesis with Neural Radiance Fields
abstract
Comprehensive 3D scene understanding, both geometrically and semantically, is important for real-world applications such as robot perception. Most of the existing work has focused on developing data-driven discriminative models for scene understanding. This paper provides a new approach to scene understanding, from a synthesis model perspective, by leveraging the recent progress on implicit scene representation and neural rendering. Building upon the great success of Neural Radiance Fields (NeRFs), we introduce Scene-Property Synthesis with NeRF (SS-NeRF) that is able to not only render photo-realistic RGB images from novel viewpoints, but also render various accurate scene properties (e.g., appearance, geometry, and semantics). By doing so, we facilitate addressing a variety of scene understanding tasks under a unified framework, including semantic segmentation, surface normal estimation, reshading, keypoint detection, and edge detection. Our SS-NeRF framework can be a powerful tool for bridging generative learning and discriminative learning, and thus be beneficial to the investigation of a wide range of interesting problems, such as studying task relationships within a synthesis paradigm, transferring knowledge to novel tasks, facilitating downstream discriminative tasks as ways of data augmentation, and serving as auto-labeller for data creation. Our code is available at https://github.com/zsh2000/SS-NeRF.
Shuhong Zheng, Zhipeng Bao, Martial Hebert, Yu-Xiong Wang
WACV4
2023 A Large-Scale Virtual Dataset and Egocentric Localization for Disaster Responses
abstract
With the increasing social demands of disaster response, methods of visual observation for rescue and safety have become increasingly important. However, because of the shortage of datasets for disaster scenarios, there has been little progress in computer vision and robotics in this field. With this in mind, we present the first large-scale synthetic dataset of egocentric viewpoints for disaster scenarios. We simulate pre- and post-disaster cases with drastic changes in appearance, such as buildings on fire and earthquakes. The dataset consists of more than 300K high-resolution stereo image pairs, all annotated with ground-truth data for the semantic label, depth in metric scale, optical flow with sub-pixel precision, and surface normal as well as their corresponding camera poses. To create realistic disaster scenes, we manually augment the effects with 3D models using physically-based graphics tools. We train various state-of-the-art methods to perform computer vision tasks using our dataset, evaluate how well these methods recognize the disaster situations, and produce reliable results of virtual scenes as well as real-world images. We also present a convolutional neural network-based egocentric localization method that is robust to drastic appearance changes, such as the texture changes in a fire, and layout changes from a collapse. To address these key challenges, we propose a new model that learns a shape-based representation by training on stylized images, and incorporate the dominant planes of query images as approximate scene coordinates. We evaluate the proposed method using various scenes including a simulated disaster dataset to demonstrate the effectiveness of our method when confronted with significant changes in scene layout. Experimental results show that our method provides reliable camera pose predictions despite vastly changed conditions.
Hae-Gon Jeon, Sunghoon Im 0001, Byeong-Uk Lee, François Rameau, Dong-Geol Choi, Jean Oh, In-So Kweon, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.8
2022 Discovering Objects that Can Move
abstract
This paper studies the problem of object discovery - separating objects from the background without manual labels. Existing approaches utilize appearance cues, such as color, texture, and location, to group pixels into object-like regions. However, by relying on appearance alone, these methods fail to separate objects from the background in cluttered scenes. This is a fundamental limitation since the definition of an object is inherently ambiguous and context-dependent. To resolve this ambiguity, we choose to focus on dynamic objects - entities that can move independently in the world. We then scale the recent auto-encoder based frameworks for unsuper-vised object discovery from toy synthetic images to complex real-world scenes. To this end, we simplify their architecture, and augment the resulting model with a weak learning signal from general motion segmentation algorithms. Our experiments demonstrate that, despite only capturing a small subset of the objects that move, this signal is enough to generalize to segment both moving and static instances of dynamic objects. We show that our model scales to a newly collected, photo- realistic synthetic dataset with street driving scenarios. Additionally, we leverage ground truth segmentation and flow annotations in this dataset for thorough ablation and evaluation. Finally, our experiments on the real-world KITTI benchmark demonstrate that the proposed approach outperforms both heuristic- and learning-based methods by capitalizing on motion cues.
Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, Martial Hebert
CVPR6
2022 Learning Continuous Implicit Representation for Near-Periodic Patterns
Bowei Chen 0004, Tiancheng Zhi, Martial Hebert, Srinivasa G. Narasimhan
ECCV (15)3
2022 Approximate Differentiable Rendering with Algebraic Surfaces
Leonid Keselman, Martial Hebert
ECCV (32)2
2022 Generative Modeling for Multi-task Visual Learning
abstract
Generative modeling has recently shown great promise in computer vision, but it has mostly focused on synthesizing visually realistic images. In this paper, motivated by multi-task learning of shareable feature representations, we consider a novel problem of learning a shared generative model that is useful across various visual perception tasks. Correspondingly, we propose a general multi-task oriented generative modeling (MGM) framework, by coupling a discriminative multi-task network with a generative network. While it is challenging to synthesize both RGB images and pixel-level annotations in multi-task scenarios, our framework enables us to use synthesized images paired with only weak annotations (i.e., image-level scene labels) to facilitate multiple visual tasks. Experimental evaluation on challenging multi-task benchmarks, including NYUv2 and Taskonomy, demonstrates that our MGM framework improves the performance of all the tasks by large margins, consistently outperforming state-of-the-art multi-task approaches in different sample-size regimes.
Zhipeng Bao, Martial Hebert, Yu-Xiong Wang
ICML2
2022 CMSNet: Deep Color and Monochrome Stereo
Hae-Gon Jeon, Sunghoon Im 0001, Jaesung Choe, Minjun Kang, Joon-Young Lee, Martial Hebert
Int. J. Comput. Vis.6
2022 Linear RGB-D SLAM for Structured Environments
abstract
We propose a new linear RGB-D simultaneous localization and mapping (SLAM) formulation by utilizing planar features of the structured environments. The key idea is to understand a given structured scene and exploit its structural regularities such as the Manhattan world. This understanding allows us to decouple the camera rotation by tracking structural regularities, which makes SLAM problems free from being highly nonlinear. Additionally, it provides a simple yet effective cue for representing planar features, which leads to a linear SLAM formulation. Given an accurate camera rotation, we jointly estimate the camera translation and planar landmarks in the global planar map using a linear Kalman filter. Our linear SLAM method, called L-SLAM, can understand not only the Manhattan world but the more general scenario of the Atlanta world, which consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction. To this end, we introduce a novel tracking-by-detection scheme that infers the underlying scene structure by Atlanta representation. With efficient Atlanta representation, we formulate a unified linear SLAM framework for structured environments. We evaluate L-SLAM on a synthetic dataset and RGB-D benchmarks, demonstrating comparable performance to other state-of-the-art SLAM methods without using expensive nonlinear optimization. We assess the accuracy of L-SLAM on a practical application of augmented reality.
Kyungdon Joo, Pyojin Kim, Martial Hebert, In-So Kweon, H. Jin Kim
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Semantically supervised appearance decomposition for virtual staging from a single panorama
abstract
We describe a novel approach to decompose a single panorama of an empty indoor environment into four appearance components: specular, direct sunlight, diffuse and diffuse ambient without direct sunlight. Our system is weakly supervised by automatically generated semantic maps (with floor, wall, ceiling, lamp, window and door labels) that have shown success on perspective views and are trained for panoramas using transfer learning without any further annotations. A GAN-based approach supervised by coarse information obtained from the semantic map extracts specular reflection and direct sunlight regions on the floor and walls. These lighting effects are removed via a similar GAN-based approach and a semantic-aware inpainting step. The appearance decomposition enables multiple applications including sun direction estimation, virtual furniture insertion, floor material replacement, and sun direction change, providing an effective tool for virtual home staging. We demonstrate the effectiveness of our approach on a large and recently released dataset of panoramas of empty homes.
Tiancheng Zhi, Bowei Chen 0004, Ivaylo Boyadzhiev, Sing Bing Kang, Martial Hebert, Srinivasa G. Narasimhan
ACM Trans. Graph.5
2021 Learning to Hallucinate Examples from Extrinsic and Intrinsic Supervision
abstract
Learning to hallucinate additional examples has recently been shown as a promising direction to address few-shot learning tasks. This work investigates two important yet overlooked natural supervision signals for guiding the hallucination process – (i) extrinsic: classifiers trained on hallucinated examples should be close to strong classifiers that would be learned from a large amount of real examples; and (ii) intrinsic: clusters of hallucinated and real examples belonging to the same class should be pulled together, while simultaneously pushing apart clusters of hallucinated and real examples from different classes. We achieve (i) by introducing an additional mentor model on data-abundant base classes for directing the hallucinator, and achieve (ii) by performing contrastive learning between hallucinated and real examples. As a general, model-agnostic framework, our dual mentor-and self-directed (DMAS) hallucinator significantly improves few-shot learning performance on widely-used benchmarks in various scenarios.
Liangke Gui, Adrien Bardes, Ruslan Salakhutdinov, Alex Hauptmann 0001, Martial Hebert, Yu-Xiong Wang
ICCV5
2021 Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis
Zhipeng Bao, Yu-Xiong Wang, Martial Hebert
ICLR3
2021 ZePHyR: Zero-shot Pose Hypothesis Rating
abstract
Pose estimation is a basic module in many robot manipulation pipelines. Estimating the pose of objects in the environment can be useful for grasping, motion planning, or manipulation. However, current state-of-the-art methods for pose estimation either rely on large annotated training sets or simulated data. Further, the long training times for these methods prohibit quick interaction with novel objects. To address these issues, we introduce a novel method for zero-shot object pose estimation in clutter. Our approach uses a hypothesis generation and scoring framework, with a focus on learning a scoring function that generalizes to objects not used for training. We achieve zero-shot generalization by rating hypotheses as a function of unordered point differences. We evaluate our method on challenging datasets with both textured and untextured objects in cluttered scenes and demonstrate that our method significantly outperforms previous methods on this task. We also demonstrate how our system can be used by quickly scanning and building a model of a novel object, which can immediately be used by our method for pose estimation. Our work allows users to estimate the pose of novel objects without requiring any retraining. Additional information can be found on our website https://bokorn.github.io/zephyr/
Brian Okorn, Qiao Gu, Martial Hebert, David Held
ICRA3
2020 PanoNet3D: Combining Semantic and Geometric Understanding for LiDAR Point Cloud Detection
abstract
Visual data in autonomous driving perception, such as camera image and LiDAR point cloud, can be interpreted as a mixture of two aspects: semantic feature and geometric structure. Semantics come from the appearance and context of objects to the sensor, while geometric structure is the actual 3D shape of point clouds. Most detectors on LiDAR point clouds focus only on analyzing the geometric structure of objects in real 3D space. Unlike previous works, we propose to learn both semantic feature and geometric structure via a unified multi-view framework. Our method exploits the nature of LiDAR scans - 2D range images, and applies well-studied 2D convolutions to extract semantic features. By fusing semantic and geometric features, our method outperforms state-of-the-art approaches in all categories by a large margin. The methodology of combining semantic and geometric features provides a unique perspective of looking at the problems in real-world 3D point cloud detection.
Jianren Wang, David Held, Martial Hebert
3DV4
2020 Learning Shape-based Representation for Visual Localization in Extremely Changing Conditions
abstract
Visual localization is an important task for applications such as navigation and augmented reality, but is a challenging problem when there are changes in scene appearances through day, seasons, or environments. In this paper, we present a convolutional neural network (CNN)-based approach for visual localization across normal to drastic appearance variations such as pre- and post-disaster cases. Our approach aims to address two key challenges: (1) to reduce the biases based on scene textures as in traditional CNNs, our model learns a shape-based representation by training on stylized images; (2) to make the model robust against layout changes, our approach uses the estimated dominant planes of query images as approximate scene coordinates. Our method is evaluated on various scenes including a simulated disaster dataset to demonstrate the effectiveness of our method in significant changes of scene layout. Experimental results show that our method provides reliable camera pose predictions in various changing conditions.
Hae-Gon Jeon, Sunghoon Im 0001, Jean Oh, Martial Hebert
ICRA4
2020 MAPPER: Multi-Agent Path Planning with Evolutionary Reinforcement Learning in Mixed Dynamic Environments
abstract
Multi-agent navigation in dynamic environments is of great industrial value when deploying a large scale fleet of robot to real-world applications. This paper proposes a decentralized partially observable multi-agent path planning with evolutionary reinforcement learning (MAPPER) method to learn an effective local planning policy in mixed dynamic environments. Reinforcement learning-based methods usually suffer performance degradation on long-horizon tasks with goal-conditioned sparse rewards, so we decompose the long-range navigation task into many easier sub-tasks under the guidance of a global planner, which increases agents' performance in large environments. Moreover, most existing multi-agent planning approaches assume either perfect information of the surrounding environment or homogeneity of nearby dynamic agents, which may not hold in practice. Our approach models dynamic obstacles' behavior with an image-based representation and trains a policy in mixed dynamic environments without homogeneity assumption. To ensure multi-agent training stability and performance, we propose an evolutionary training approach that can be easily scaled to large and complex environments. Experiments show that MAPPER is able to achieve higher success rates and more stable performance when exposed to a large number of non-cooperative dynamic obstacles compared with traditional reaction-based planner LRA* and the state-of-the-art learning-based method.
Zuxin Liu, Baiming Chen, Hongyi Zhou, Guru Koushik Senthil Kumar, Martial Hebert, Ding Zhao
IROS5
2020 Learning Orientation Distributions for Object Pose Estimation
abstract
For robots to operate robustly in the real world, they should be aware of their uncertainty. However, most methods for object pose estimation return a single point estimate of the object's pose. In this work, we propose two learned methods for estimating a distribution over an object's orientation. Our methods take into account both the inaccuracies in the pose estimation as well as the object symmetries. Our first method, which regresses from deep learned features to an isotropic Bingham distribution, gives the best performance for orientation distribution estimation for non-symmetric objects. Our second method learns to compare deep features and generates a non-parameteric histogram distribution. This method gives the best performance on objects with unknown symmetries, accurately modeling both symmetric and non-symmetric objects, without any requirement of symmetry annotation. We show that both of these methods can be used to augment an existing pose estimator. Our evaluation compares our methods to a large number of baseline approaches for uncertainty estimation across a variety of different types of objects. Code available at https://bokorn.github.io/orientation-distributions/.
Brian Okorn, Mengyun Xu, Martial Hebert, David Held
IROS3
2020 Quadtree Generating Networks: Efficient Hierarchical Scene Parsing with Sparse Convolutions
abstract
Semantic segmentation with Convolutional Neural Networks is a memory-intensive task due to the high spatial resolution of feature maps and output predictions. In this paper, we present Quadtree Generating Networks (QGNs), a novel approach able to drastically reduce the memory footprint of modern semantic segmentation networks. The key idea is to use quadtrees to represent the predictions and target segmentation masks instead of dense pixel grids. Our quadtree representation enables hierarchical processing of an input image, with the most computationally demanding layers only being used at regions in the image containing boundaries between classes. In addition, given a trained model, our representation enables flexible inference schemes to trade-off accuracy and computational cost, allowing the network to adapt in constrained situations such as embedded devices. We demonstrate the benefits of our approach on the Cityscapes, SUN-RGBD and ADE20k datasets. On Cityscapes, we obtain an relative 3% mIoU improvement compared to a dilated network with similar memory consumption; and only receive a 3% relative mIoU drop compared to a large dilated network, while reducing memory consumption by over 4×. Our code is available at https://github.com/kashyap7x/QGN.
Kashyap Chitta, José M. Álvarez 0004, Martial Hebert
WACV3
2020 Special Issue: Advances in Architectures and Theories for Computer Vision
Yair Weiss, Vittorio Ferrari, Cristian Sminchisescu, Martial Hebert
Int. J. Comput. Vis.4
2019 Learning Anytime Predictions in Neural Networks via Adaptive Loss Balancing
abstract
This work considers the trade-off between accuracy and testtime computational cost of deep neural networks (DNNs) via anytime predictions from auxiliary predictions. Specifically, we optimize auxiliary losses jointly in an adaptive weighted sum, where the weights are inversely proportional to average of each loss. Intuitively, this balances the losses to have the same scale. We demonstrate theoretical considerations that motivate this approach from multiple viewpoints, including connecting it to optimizing the geometric mean of the expectation of each loss, an objective that ignores the scale of losses. Experimentally, the adaptive weights induce more competitive anytime predictions on multiple recognition data-sets and models than non-adaptive approaches including weighing all losses equally. In particular, anytime neural networks (ANNs) can achieve the same accuracy faster using adaptive weights on a small network than using static constant weights on a large one. For problems with high performance saturation, we also show a sequence of exponentially deepening ANNs can achieve near-optimal anytime results at any budget, at the cost of a const fraction of extra computation.
Hanzhang Hu, Debadeepta Dey, Martial Hebert, J. Andrew Bagnell
AAAI3
2019 Image Deformation Meta-Networks for One-Shot Learning
abstract
Humans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts from one or few examples. Our key insight is that, while the deformed images may not be visually realistic, they still maintain critical semantic information and contribute significantly to formulating classifier decision boundaries. Inspired by the recent progress of meta-learning, we combine a meta-learner with an image deformation sub-network that produces additional training examples, and optimize both models in an end-to-end manner. The deformation sub-network learns to deform images by fusing a pair of images --- a probe image that keeps the visual content and a gallery image that diversifies the deformations. We demonstrate results on the widely used one-shot learning benchmarks (miniImageNet and ImageNet 1K Challenge datasets), which significantly outperform state-of-the-art approaches.
Zitian Chen, Yanwei Fu 0001, Yu-Xiong Wang, Lin Ma 0002, Wei Liu 0005, Martial Hebert
CVPR6
2019 Coordinate-Free Carlsson-Weinshall Duality and Relative Multi-View Geometry
abstract
We present a coordinate-free description of Carlsson-Weinshall duality between scene points and camera pinholes and use it to derive a new characterization of primal/dual multi-view geometry. In the case of three views, a particular set of reduced trilinearities provide a novel parameterization of camera geometry that, unlike existing ones, is subject only to very simple internal constraints. These trilinearities lead to new "quasi-linear" algorithms for primal and dual structure from motion. We include some preliminary experiments with real and synthetic data.
Matthew Trager, Martial Hebert, Jean Ponce
CVPR2
2019 A Structured Model for Action Detection
abstract
A dominant paradigm for learning-based approaches in computer vision is training generic models, such as ResNet for image recognition, or I3D for video understanding, on large datasets and allowing them to discover the optimal representation for the problem at hand. While this is an obviously attractive approach, it is not applicable in all scenarios. We claim that action detection is one such challenging problem - the models that need to be trained are large, and labeled data is expensive to obtain. To address this limitation, we propose to incorporate domain knowledge into the structure of the model, simplifying optimization. In particular, we augment a standard I3D network with a tracking module to aggregate long-term motion patterns, and use a graph convolutional network to reason about interactions between actors and objects. Evaluated on the challenging AVA dataset, the proposed approach improves over the I3D baseline by 5.5% mAP and over the state-of-the-art by 4.8% mAP.
Pavel Tokmakov, Martial Hebert, Cordelia Schmid
CVPR3
2019 Multispectral Imaging for Fine-Grained Recognition of Powders on Complex Backgrounds
abstract
Hundreds of materials, such as drugs, explosives, makeup, food additives, are in the form of powder. Recognizing such powders is important for security checks, criminal identification, drug control, and quality assessment. However, powder recognition has drawn little attention in the computer vision community. Powders are hard to distinguish: they are amorphous, appear matte, have little color or texture variation and blend with surfaces they are deposited on in complex ways. To address these challenges, we present the first comprehensive dataset and approach for powder recognition using multi-spectral imaging. By using Shortwave Infrared (SWIR) multi-spectral imaging together with visible light (RGB) and Near Infrared (NIR), powders can be discriminated with reasonable accuracy. We present a method to select discriminative spectral bands to significantly reduce acquisition time while improving recognition accuracy. We propose a blending model to synthesize images of powders of various thickness deposited on a wide range of surfaces. Incorporating band selection and image synthesis, we conduct fine-grained recognition of 100 powders on complex backgrounds, and achieve 60%~70% accuracy on recognition with known powder location, and over 40% mean IoU without known location.
Tiancheng Zhi, Bernardo Rodrigues Pires, Martial Hebert, Srinivasa G. Narasimhan
CVPR3
2019 Learning Compositional Representations for Few-Shot Recognition
abstract
One of the key limitations of modern deep learning approaches lies in the amount of data required to train them. Humans, by contrast, can learn to recognize novel categories from just a few examples. Instrumental to this rapid learning ability is the compositional structure of concept representations in the human brain --- something that deep learning models are lacking. In this work, we make a step towards bridging this gap between human and machine learning by introducing a simple regularization technique that allows the learned representation to be decomposable into parts. Our method uses category-level attribute annotations to disentangle the feature space of a network into subspaces corresponding to the attributes. These attributes can be either purely visual, like object parts, or more abstract, like openness and symmetry. We demonstrate the value of compositional representations on three datasets: CUB-200-2011, SUN397, and ImageNet, and show that they require fewer examples to learn classifiers for novel categories.
Pavel Tokmakov, Yu-Xiong Wang, Martial Hebert
ICCV3
2019 Meta-Learning to Detect Rare Objects
abstract
Few-shot learning, i.e., learning novel concepts from few examples, is fundamental to practical visual recognition systems. While most of existing work has focused on few-shot classification, we make a step towards few-shot object detection, a more challenging yet under-explored task. We develop a conceptually simple but powerful meta-learning based framework that simultaneously tackles few-shot classification and few-shot localization in a unified, coherent way. This framework leverages meta-level knowledge about "model parameter generation" from base classes with abundant data to facilitate the generation of a detector for novel classes. Our key insight is to disentangle the learning of category-agnostic and category-specific components in a CNN based detection model. In particular, we introduce a weight prediction meta-model that enables predicting the parameters of category-specific components from few examples. We systematically benchmark the performance of modern detectors in the small-sample size regime. Experiments in a variety of realistic scenarios, including within-domain, cross-domain, and long-tailed settings, demonstrate the effectiveness and generality of our approach under different notions of novel classes.
Yu-Xiong Wang, Deva Ramanan, Martial Hebert
ICCV3
2019 DISC: A Large-scale Virtual Dataset for Simulating Disaster Scenarios
abstract
In this paper, we present the first large-scale synthetic dataset for visual perception in disaster scenarios, and analyze state-of-the-art methods for multiple computer vision tasks with reference baselines. We simulated before and after disaster scenarios such as fire and building collapse for fifteen different locations in realistic virtual worlds. The dataset consists of more than 300K high-resolution stereo image pairs, all annotated with ground-truth data for semantic segmentation, depth, optical flow, surface normal estimation and camera pose estimation. To create realistic disaster scenes, we manually augmented the effects with 3D models using physical-based graphics tools. We use our dataset to train state-of-the-art methods and evaluate how well these methods can recognize the disaster situations and produce reliable results on virtual scenes as well as real-world images. The results obtained from each task are then used as inputs to the proposed visual odometry network for generating 3D maps of buildings on fire. Finally, we discuss challenges for future research.
Hae-Gon Jeon, Sunghoon Im 0001, Byeong-Uk Lee, Dong-Geol Choi, Martial Hebert, In-So Kweon
IROS5
2018 PCN: Point Completion Network
abstract
Shape completion, the problem of estimating the complete geometry of objects from partial observations, lies at the core of many vision and robotics applications. In this work, we propose Point Completion Network (PCN), a novel learning-based approach for shape completion. Unlike existing shape completion methods, PCN directly operates on raw point clouds without any structural assumption (e.g. symmetry) or annotation (e.g. semantic class) about the underlying shape. It features a decoder design that enables the generation of fine-grained completions while maintaining a small number of parameters. Our experiments show that PCN produces dense, complete point clouds with realistic structures in the missing regions on inputs with various levels of incompleteness and noise, including cars from LiDAR scans in the KITTI dataset.
Tejas Khot, David Held, Christoph Mertz, Martial Hebert
3DV5
2018 Learning by Asking Questions
abstract
We introduce an interactive learning framework for the development and testing of intelligent visual systems, called learning-by-asking (LBA). We explore LBA in context of the Visual Question Answering (VQA) task. LBA differs from standard VQA training in that most questions are not observed during training time, and the learner must ask questions it wants answers to. Thus, LBA more closely mimics natural learning and has the potential to be more data-efficient than the traditional VQA setting. We present a model that performs LBA on the CLEVR dataset, and show that it automatically discovers an easy-to-hard curriculum when learning interactively from an oracle. Our LBA generated data consistently matches or outperforms the CLEVR train data and is more sample efficient. We also show that our model asks questions that generalize to state-of-the-art VQA models and to novel test time distributions.
Ishan Misra, Ross B. Girshick, Rob Fergus, Martial Hebert, Abhinav Gupta 0001, Laurens van der Maaten
CVPR4
2018 Low-Shot Learning From Imaginary Data
abstract
Humans can quickly learn new visual concepts, perhaps because they can easily visualize or imagine what novel objects look like from different views. Incorporating this ability to hallucinate novel instances of new concepts might help machine vision systems perform better low-shot learning, i.e., learning concepts from few examples. We present a novel approach to low-shot learning that uses this idea. Our approach builds on recent progress in meta-learning ("learning to learn") by combining a meta-learner with a "hallucinator" that produces additional training examples, and optimizing both models jointly. Our hallucinator can be incorporated into a variety of meta-learners and provides significant gains: up to a 6 point boost in classification accuracy when only a single training example is available, yielding state-of-the-art performance on the challenging ImageNet low-shot classification benchmark.
Yu-Xiong Wang, Ross B. Girshick, Martial Hebert, Bharath Hariharan
CVPR3
2018 Deep Material-Aware Cross-Spectral Stereo Matching
abstract
Cross-spectral imaging provides strong benefits for recognition and detection tasks. Often, multiple cameras are used for cross-spectral imaging, thus requiring image alignment, or disparity estimation in a stereo setting. Increasingly, multi-camera cross-spectral systems are embedded in active RGBD devices (e.g. RGB-NIR cameras in Kinect and iPhone X). Hence, stereo matching also provides an opportunity to obtain depth without an active projector source. However, matching images from different spectral bands is challenging because of large appearance variations. We develop a novel deep learning framework to simultaneously transform images across spectral bands and estimate disparity. A material-aware loss function is incorporated within the disparity prediction network to handle regions with unreliable matching such as light sources, glass windshields and glossy surfaces. No depth supervision is required by our method. To evaluate our method, we used a vehicle-mounted RGB-NIR stereo system to collect 13.7 hours of video data across a range of areas in and around a city. Experiments show that our method achieves strong performance and reaches real-time speed.
Tiancheng Zhi, Bernardo Rodrigues Pires, Martial Hebert, Srinivasa G. Narasimhan
CVPR3
2017 Gradient Boosting on Stochastic Data Streams
abstract
Boosting is a popular ensemble algorithm that generates more powerful learners by linearly combining base models from a simpler hypothesis class. In this work, we investigate the problem of adapting batch gradient boosting for minimizing convex loss functions to online setting where the loss at each iteration is i.i.d sampled from an unknown distribution. To generalize from batch to online, we first introduce the definition of online weak learning edge with which for strongly convex and smooth loss functions, we present an algorithm, Streaming Gradient Boosting (SGB) with exponential shrinkage guarantees in the number of weak learners. We further present an adaptation of SGB to optimize non-smooth loss functions, for which we derive a $O(\ln(N)/N)$ convergence rate. We also show that our analysis can extend to adversarial online learning setting under a stronger assumption that the online weak learning edge will hold in adversarial setting. We finally demonstrate experimental results showing that in practice our algorithms can achieve competitive results as classic gradient boosting while using less computation.
Hanzhang Hu, Wen Sun 0002, Arun Venkatraman, Martial Hebert, J. Andrew Bagnell
AISTATS4
2017 From Red Wine to Red Tomato: Composition with Context
abstract
Compositionality and contextuality are key building blocks of intelligence. They allow us to compose known concepts to generate new and complex ones. However, traditional learning methods do not model both these properties and require copious amounts of labeled data to learn new concepts. A large fraction of existing techniques, e.g., using late fusion, compose concepts but fail to model contextuality. For example, red in red wine is different from red in red tomatoes. In this paper, we present a simple method that respects contextuality in order to compose classifiers of known visual concepts. Our method builds upon the intuition that classifiers lie in a smooth space where compositional transforms can be modeled. We show how it can generalize to unseen combinations of concepts. Our results on composing attributes, objects as well as composing subject, predicate, and objects demonstrate its strong generalization performance compared to baselines. Finally, we present detailed analysis of our method and highlight its properties.
Ishan Misra, Abhinav Gupta 0001, Martial Hebert
CVPR3
2017 General Models for Rational Cameras and the Case of Two-Slit Projections
abstract
The rational camera model recently introduced in [18] provides a general methodology for studying abstract nonlinear imaging systems and their multi-view geometry. This paper builds on this framework to study physical realizations of rational cameras. More precisely, we give an explicit account of the mapping between between physical visual rays and image points (missing in the original description), which allows us to give simple analytical expressions for direct and inverse projections. We also consider primitive camera models, that are orbits under the action of various projective transformations, and lead to a general notion of intrinsic parameters. The methodology is general, but it is illustrated concretely by an in-depth study of two-slit cameras, that we model using pairs of linear projections. This simple analytical form allows us to describe models for the corresponding primitive cameras, to introduce intrinsic parameters with a clear geometric meaning, and to define an epipolar tensor characterizing two-view correspondences. In turn, this leads to new algorithms for structure from motion and self-calibration.
Matthew Trager, Bernd Sturmfels, John F. Canny, Martial Hebert, Jean Ponce
CVPR4
2017 Growing a Brain: Fine-Tuning by Increasing Model Capacity
abstract
CNNs have made an undeniable impact on computer vision through the ability to learn high-capacity models with large annotated training sets. One of their remarkable properties is the ability to transfer knowledge from a large source dataset to a (typically smaller) target dataset. This is usually accomplished through fine-tuning a fixed-size network on new target data. Indeed, virtually every contemporary visual recognition system makes use of fine-tuning to transfer knowledge from ImageNet. In this work, we analyze what components and parameters change during fine-tuning, and discover that increasing model capacity allows for more natural model adaptation through fine-tuning. By making an analogy to developmental learning, we demonstrate that growing a CNN with additional units, either by widening existing layers or deepening the overall network, significantly outperforms classic fine-tuning approaches. But in order to properly grow a network, we show that newly-added units must be appropriately normalized to allow for a pace of learning that is consistent with existing units. We empirically validate our approach on several benchmark datasets, producing state-of-the-art results.
Yu-Xiong Wang, Deva Ramanan, Martial Hebert
CVPR3
2017 Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection
abstract
A major impediment in rapidly deploying object detection models for instance detection is the lack of large annotated datasets. For example, finding a large labeled dataset containing instances in a particular kitchen is unlikely. Each new environment with new instances requires expensive data collection and annotation. In this paper, we propose a simple approach to generate large annotated instance datasets with minimal effort. Our key insight is that ensuring only patch-level realism provides enough training signal for current object detector models. We automatically `cut' object instances and `paste' them on random backgrounds. A naive way to do this results in pixel artifacts which result in poor performance for trained models. We show how to make detectors ignore these artifacts during training and generate data that gives competitive performance on real data. Our method outperforms existing synthesis approaches and when combined with real images improves relative performance by more than 21% on benchmark datasets. In a cross-domain setting, our synthetic data combined with just 10% real data outperforms models trained on all real data.
Debidatta Dwibedi, Ishan Misra, Martial Hebert
ICCV3
2017 The Pose Knows: Video Forecasting by Generating Pose Futures
abstract
Current approaches to video forecasting attempt to generate videos directly in pixel space using Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs). However, since these approaches try to model all the structure and scene dynamics at once, in unconstrained settings they often generate uninterpretable results. Our insight is that forecasting needs to be done first at a higher level of abstraction. Specifically, we exploit human pose detectors as a free source of supervision and break the video forecasting problem into two discrete steps. First we explicitly model the high level structure of active objects in the scene (humans) and use a VAE to model the possible future movements of humans in the pose space. We then use the future poses generated as conditional information to a GAN to predict the future frames of the video in pixel space. By using the structured space of pose as an intermediate representation, we sidestep the problems that GANs have in generating video pixels directly. We show through quantitative and qualitative evaluation that our method outperforms state-of-the-art methods for video prediction.
Jacob Walker, Kenneth Marino, Abhinav Gupta 0001, Martial Hebert
ICCV4
2017 Learning robust failure response for autonomous vision based flight
abstract
The ability of autonomous mobile robots to react to and recover from potential failures of on-board systems is an important area of ongoing robotics research. With increasing emphasis on robust systems and long-term autonomy, mobile robots must be able to respond safely and intelligently to dangerous situations. Recent developments in computer vision have made autonomous vision based navigation possible. However, vision systems are known to be imperfect and prone to failure due to variable lighting, terrain changes, and other environmental variables. We describe a system for learning simple failure recovery maneuvers based on experience. This involves both recognizing when the vision system is prone to failure, and associating failures with appropriate responses that will most likely help the robot recover. We implement this system on an autonomous quadrotor and demonstrate that behaviors learned with our system are effective in recovering from situational perception failure, thereby improving reliability in cluttered and uncertain environments.
Dhruv Mauria Saxena, Vincent Kurtz, Martial Hebert
ICRA3
2017 Predictive-State Decoders: Encoding the Future into Recurrent Networks
abstract
Recurrent neural networks (RNNs) are a vital modeling technique that rely on internal states learned indirectly by optimization of a supervised, unsupervised, or reinforcement training loss. RNNs are used to model dynamic processes that are characterized by underlying latent states whose form is often unknown, precluding its analytic representation inside an RNN. In the Predictive-State Representation (PSR) literature, latent state processes are modeled by an internal state representation that directly models the distribution of future observations, and most recent work in this area has relied on explicitly representing and targeting sufficient statistics of this probability distribution. We seek to combine the advantages of RNNs and PSRs by augmenting existing state-of-the-art recurrent neural networks with Predictive-State Decoders (PSDs), which add supervision to the network's internal state representation to target predicting future observations. PSDs are simple to implement and easily incorporated into existing training pipelines via additional loss regularization. We demonstrate the effectiveness of PSDs with experimental results in three different domains: probabilistic filtering, Imitation Learning, and Reinforcement Learning. In each, our method improves statistical performance of state-of-the-art recurrent baselines and does so with fewer iterations and less data.
Arun Venkatraman, Nicholas Rhinehart, Wen Sun 0002, Lerrel Pinto, Martial Hebert, Byron Boots, Kris Makoto Kitani, J. Andrew Bagnell
NIPS5
2017 Learning to Model the Tail
abstract
We describe an approach to learning from long-tailed, imbalanced datasets that are prevalent in real-world settings. Here, the challenge is to learn accurate "few-shot'' models for classes in the tail of the class distribution, for which little data is available. We cast this problem as transfer learning, where knowledge from the data-rich classes in the head of the distribution is transferred to the data-poor classes in the tail. Our key insights are as follows. First, we propose to transfer meta-knowledge about learning-to-learn from the head classes. This knowledge is encoded with a meta-network that operates on the space of model parameters, that is trained to predict many-shot model parameters from few-shot model parameters. Second, we transfer this meta-knowledge in a progressive manner, from classes in the head to the "body'', and from the "body'' to the tail. That is, we transfer knowledge in a gradual fashion, regularizing meta-networks for few-shot regression with those trained with more training data. This allows our final network to capture a notion of model dynamics, that predicts how model parameters are likely to change as more training data is gradually added. We demonstrate results on image classification datasets (SUN, Places, and ImageNet) tuned for the long-tailed setting, that significantly outperform common heuristics, such as data resampling or reweighting.
Yu-Xiong Wang, Deva Ramanan, Martial Hebert
NIPS3
2016 Online Instrumental Variable Regression with Applications to Online Linear System Identification
abstract
Instrumental variable regression (IVR) is a statistical technique utilized to recover unbiased estimators when there are errors in the independent variables. Estimator bias in learned time series models can yield poor performance in applications such as long-term prediction and filtering where the recursive use of the model results in the accumulation of propagated error. However, prior work addressed the IVR objective in the batch setting, where it is necessary to store the entire dataset in memory — an infeasible requirementin large dataset scenarios. In this work, we develop Online Instrumental Variable Regression (OIVR), an algorithm that is capable of updating the learned estimator with streaming data. We show that the online adaptation of IVR enjoys a no-regret performance guarantee with respect the original batchsetting by taking advantage of any no-regret online learning algorithm inside OIVR for the underlying update steps. We experimentally demonstrate the efficacy of our algorithm in combination with popular no-regret onlinealgorithms for the task of learning predictive dynamical system models and on a prototypical econometrics instrumental variable regression problem.
Arun Venkatraman, Wen Sun 0002, Martial Hebert, J. Andrew Bagnell, Byron Boots
AAAI3
2016 Learning by Transferring from Unsupervised Universal Sources
abstract
Category classifiers trained from a large corpus of annotated data are widely accepted as the sources for (hypothesis) transfer learning. Sources generated in this way are tied to a particular set of categories, limiting their transferability across a wide spectrum of target categories. In this paper, we address this largely-overlooked yet fundamental source problem by both introducing a systematic scheme for generating universal source hypotheses and proposing a principled, scalable approach to automatically tuning the transfer process. Our approach is based on the insights that expressive source hypotheses could be generated without any supervision and that a sparse combination of such hypotheses facilitates recognition of novel categories from few samples. We demonstrate improvements over the state-of-the-art on object and scene classification in the small sample size regime.
Yu-Xiong Wang, Martial Hebert
AAAI2
2016 Learning to Extract Motion from Videos in Convolutional Neural Networks
Damien Teney, Martial Hebert
ACCV (5)2
2016 Cross-Stitch Networks for Multi-task Learning
abstract
Multi-task learning in Convolutional Networks has displayed remarkable success in the field of recognition. This success can be largely attributed to learning shared representations from multiple supervisory tasks. However, existing multi-task approaches rely on enumerating multiple network architectures specific to the tasks at hand, that do not generalize. In this paper, we propose a principled approach to learn shared representations in ConvNets using multitask learning. Specifically, we propose a new sharing unit: "cross-stitch" unit. These units combine the activations from multiple networks and can be trained end-to-end. A network with cross-stitch units can learn an optimal combination of shared and task-specific representations. Our proposed method generalizes across multiple tasks and shows dramatically improved performance over baseline methods for categories with few training examples.
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta 0001, Martial Hebert
CVPR4
2016 Consistency of Silhouettes and Their Duals
abstract
Silhouettes provide rich information on three-dimensional shape, since the intersection of the associated visual cones generates the "visual hull", which encloses and approximates the original shape. However, not all silhouettes can actually be projections of the same object in space: this simple observation has implications in object recognition and multi-view segmentation, and has been (often implicitly) used as a basis for camera calibration. In this paper, we investigate the conditions for multiple silhouettes, or more generally arbitrary closed image sets, to be geometrically "consistent". We present this notion as a natural generalization of traditional multi-view geometry, which deals with consistency for points. After discussing some general results, we present a "dual" formulation for consistency, that gives conditions for a family of planar sets to be sections of the same object. Finally, we introduce a more general notion of silhouette "compatibility" under partial knowledge of the camera projections, and point out some possible directions for future research.
Matthew Trager, Martial Hebert, Jean Ponce
CVPR2
2016 A Discriminative Framework for Anomaly Detection in Large Videos
Allison Del Giorno, J. Andrew Bagnell, Martial Hebert
ECCV (5)3
2016 Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification
Ishan Misra, C. Lawrence Zitnick, Martial Hebert
ECCV (1)3
2016 An Uncertain Future: Forecasting from Static Images Using Variational Autoencoders
Jacob Walker, Carl Doersch, Abhinav Gupta 0001, Martial Hebert
ECCV (7)4
2016 Learning to Learn: Model Regression Networks for Easy Small Sample Learning
Yu-Xiong Wang, Martial Hebert
ECCV (6)2
2016 Inference Machines for Nonparametric Filter Learning
Arun Venkatraman, Wen Sun 0002, Martial Hebert, Byron Boots, J. Andrew Bagnell
IJCAI3
2016 Introspective perception: Learning to predict failures in vision systems
abstract
As robots aspire for long-term autonomous operations in complex dynamic environments, the ability to reliably take mission-critical decisions in ambiguous situations becomes critical. This motivates the need to build systems that have situational awareness to assess how qualified they are at that moment to make a decision. We call this self-evaluating capability as introspection. In this paper, we take a small step in this direction and propose a generic framework for introspective behavior in perception systems. Our goal is to learn a model to reliably predict failures in a given system, with respect to a task, directly from input sensor data. We present this in the context of vision-based autonomous MAV flight in outdoor natural environments, and show that it effectively handles uncertain situations.
Shreyansh Daftry, Sam Zeng, J. Andrew Bagnell, Martial Hebert
IROS4
2016 Map-supervised road detection
abstract
We propose an approach to detect drivable road area in monocular images. It is a self-supervised approach which doesn't require any human road annotations on images to train the road detection algorithm. Our approach reduces human labeling effort and makes training scalable. We combine the best of both supervised and unsupervised methods in our approach. First, we automatically generate training road annotations for images using OpenStreetMap1, vehicle pose estimation sensors, and camera parameters. Next, we train a Convolutional Neural Network (CNN) for road detection using these annotations. We show that we are able to generate reasonably accurate training annotations in KITTI data-set [1]. We achieve state-of-the-art performance among the methods which do not require human annotation effort.
Ankit Laddha, Mehmet Kemal Kocamaz, Luis E. Navarro-Serment, Martial Hebert
Intelligent Vehicles Symposium4
2016 Learning from Small Sample Sets by Combining Unsupervised Meta-Training with CNNs
abstract
This work explores CNNs for the recognition of novel categories from few examples. Inspired by the transferability properties of CNNs, we introduce an additional unsupervised meta-training stage that exposes multiple top layer units to a large amount of unlabeled real-world images. By encouraging these units to learn diverse sets of low-density separators across the unlabeled data, we capture a more generic, richer description of the visual world, which decouples these units from ties to a specific set of categories. We propose an unsupervised margin maximization that jointly estimates compact high-density regions and infers low-density separators. The low-density separator (LDS) modules can be plugged into any or all of the top layers of a standard CNN architecture. The resulting CNNs significantly improve the performance in scene classification, fine-grained recognition, and action recognition with small training samples.
Yu-Xiong Wang, Martial Hebert
NIPS2
2016 Efficient Feature Group Sequencing for Anytime Linear Prediction
Hanzhang Hu, Alexander Grubb, J. Andrew Bagnell, Martial Hebert
UAI4
2016 Cutting through the clutter: Task-relevant features for image matching
abstract
Where do we focus our attention in an image? Humans have an amazing ability to cut through the clutter to the parts of an image most relevant to the task at hand. Consider the task of geo-localizing tourist photos by retrieving other images taken at that location. Such photos naturally contain friends and family, and perhaps might even be nearly filled by a person's face if it is a selfie. Humans have no trouble ignoring these `distractions' and recognizing the parts that are indicative of location (e.g., the towers of Neuschwanstein Castle instead of their friend's face, a tree, or a car). In this paper, we investigate learning this ability automatically. At training-time, we learn how informative a region is for localization. At test-time, we use this learned model to determine what parts of a query image to use for retrieval. We introduce a new dataset, People at Landmarks, that contains large amounts of clutter in query images. Our system is able to outperform the existing state of the art approach to retrieval by more than 10% mAP, as well as improve results on a standard dataset without heavy occluders (Oxford5K).
Rohit Girdhar, David F. Fouhey, Kris Makoto Kitani, Abhinav Gupta 0001, Martial Hebert
WACV5
2016 Dealing with small data and training blind spots in the Manhattan world
abstract
Leveraging Manhattan assumption we generate metrically rectified novel views from a single image, even for non-box scenarios. Our novel views enable the already trained classifiers to handle training data missing views (blind spots) without additional training. We demonstrate this on end-to-end scene text spotting under perspective. Additionally, utilizing our fronto-parallel views, we discover unsuspended invariant mid-level patches given a few widely separated training examples (small data domain). These invariant patches outperform various baselines on small data image retrieval challenge.
Muhammad Wajahat Hussain, Javier Civera 0001, Luis Montano, Martial Hebert
WACV4
2016 Trinocular Geometry Revisited
Matthew Trager, Jean Ponce, Martial Hebert
Int. J. Comput. Vis.3
2015 Toward Mobile Robots Reasoning Like Humans
abstract
Robots are increasingly becoming key players in human-robot teams. To become effective teammates, robots must possess profound understanding of an environment, be able to reason about the desired commands and goals within a specific context, and be able to communicate with human teammates in a clear and natural way. To address these challenges, we have developed an intelligence architecture that combines cognitive components to carry out high-level cognitive tasks, semantic perception to label regions in the world, and a natural language component to reason about the command and its relationship to the objects in the world. This paper describes recent developments using this architecture on a fielded mobile robot platform operating in unknown urban environments. We report a summary of extensive outdoor experiments; the results suggest that a multidisciplinary approach to robotics has the potential to create competent human-robot teams.
Jean Oh, Arne Suppé, Felix Duvallet, Abdeslam Boularias, Luis E. Navarro-Serment, Martial Hebert, Anthony Stentz, Jerry Vinokurov, Oscar J. Romero, Christian Lebiere, Robert M. Dean
AAAI6
2015 Improving Multi-Step Prediction of Learned Time Series Models
abstract
Most typical statistical and machine learning approaches to time series modeling optimize a single-step prediction error. In multiple-step simulation, the learned model is iteratively applied, feeding through the previous output as its new input. Any such predictor however, inevitably introduces errors, and these compounding errors change the input distribution for future prediction steps, breaking the train-test i.i.d assumption common in supervised learning. We present an approach that reuses training data to make a no-regret learner robust to errors made during multi-step prediction. Our insight is to formulate the problem as imitation learning; the training data serves as a "demonstrator" by providing corrections for the errors made during multi-step prediction. By this reduction of multi-step time series prediction to imitation learning, we establish theoretically a strong performance guarantee on the relation between training error and the multi-step prediction error. We present experimental results of our method, DaD, and show significant improvement over the traditional approach in two notably different domains, dynamic system modeling and video texture prediction.
Arun Venkatraman, Martial Hebert, J. Andrew Bagnell
AAAI2
2015 Watch and learn: Semi-supervised learning of object detectors from videos
abstract
We present a semi-supervised approach that localizes multiple unknown object instances in long videos. We start with a handful of labeled boxes and iteratively learn and label hundreds of thousands of object instances. We propose criteria for reliable object detection and tracking for constraining the semi-supervised learning process and minimizing semantic drift. Our approach does not assume exhaustive labeling of each object instance in any single frame, or any explicit annotation of negative data. Working in such a generic setting allow us to tackle multiple object instances in video, many of which are static. In contrast, existing approaches either do not consider multiple object instances per video, or rely heavily on the motion of the objects present. The experiments demonstrate the effectiveness of our approach by evaluating the automatically labeled data on a variety of metrics like quality, coverage (recall), diversity, and relevance to training an object detector.
Ishan Misra, Abhinav Shrivastava, Martial Hebert
CVPR3
2015 Inferring 3D layout of building facades from a single image
abstract
In this paper, we propose a novel algorithm that infers the 3D layout of building facades from a single 2D image of an urban scene. Different from existing methods that only yield coarse orientation labels or qualitative block approximations, our algorithm quantitatively reconstructs building facades in 3D space using a set of planes mutually related by 3D geometric constraints. Each plane is characterized by a continuous orientation vector and a depth distribution. An optimal solution is reached through inter-planar interactions. Due to the quantitative and plane-based nature of our geometric reasoning, our model is more expressive and informative than existing approaches. Experiments show that our method compares competitively with the state of the art on both 2D and 3D measures, while yielding a richer interpretation of the 3D scene behind the image.
Jiyan Pan, Martial Hebert, Takeo Kanade
CVPR2
2015 Model recommendation: Generating object detectors from few samples
abstract
In this paper, we explore an approach to generating detectors that is radically different from the conventional way of learning a detector from a large corpus of annotated positive and negative data samples. Instead, we assume that we have evaluated “off-line” a large library of detectors against a large set of detection tasks. Given a new target task, we evaluate a subset of the models on few samples from the new task and we use the matrix of models-tasks ratings to predict the performance of all the models in the library on the new task, enabling us to select a good set of detectors for the new task. This approach has three key advantages of great interest in practice: 1) generating a large collection of expressive models in an unsupervised manner is possible; 2) a far smaller set of annotated samples is needed compared to that required for training from scratch; and 3) recommending models is a very fast operation compared to the notoriously expensive training procedures of modern detectors. (1) will make the models informative across different categories; (2) will dramatically reduce the need for manually annotating vast datasets for training detectors; and (3) will enable rapid generation of new detectors.
Yu-Xiong Wang, Martial Hebert
CVPR2
2015 Predicting Multiple Structured Visual Interpretations
abstract
We present a simple approach for producing a small number of structured visual outputs which have high recall, for a variety of tasks including monocular pose estimation and semantic scene segmentation. Current state-of-the-art approaches learn a single model and modify inference procedures to produce a small number of diverse predictions. We take the alternate route of modifying the learning procedure to directly optimize for good, high recall sequences of structured-output predictors. Our approach introduces no new parameters, naturally learns diverse predictions and is not tied to any specific structured learning or inference procedure. We leverage recent advances in the contextual submodular maximization literature to learn a sequence of predictors and empirically demonstrate the simplicity and performance of our approach on multiple challenging vision tasks including achieving state-of-the-art results on multiple predictions for monocular pose-estimation and image foreground/background segmentation.
Debadeepta Dey, Varun Ramakrishna, Martial Hebert, J. Andrew Bagnell
ICCV3
2015 Single Image 3D without a Single 3D Image
abstract
Do we really need 3D labels in order to learn how to predict 3D? In this paper, we show that one can learn a mapping from appearance to 3D properties without ever seeing a single explicit 3D label. Rather than use explicit supervision, we use the regularity of indoor scenes to learn the mapping in a completely unsupervised manner. We demonstrate this on both a standard 3D scene understanding dataset as well as Internet images for which 3D is unavailable, precluding supervised learning. Despite never seeing a 3D label, our method produces competitive results.
David F. Fouhey, Muhammad Wajahat Hussain, Abhinav Gupta 0001, Martial Hebert
ICCV4
2015 The Joint Image Handbook
abstract
Given multiple perspective photographs, point correspondences form the "joint image", effectively a replica of three dimensional space distributed across its two-dimensional projections. This set can be characterized by multilinear equations over image coordinates, such as epipolar and trifocal constraints. We revisit in this paper the geometric and algebraic properties of the joint image, and address fundamental questions such as how many and which multilinearities are necessary and/or sufficient to determine camera geometry and/or image correspondences. The new theoretical results in this paper answer these questions in a very general setting and, in turn, are intended to serve as a "handbook" reference about multilinearities for practitioners.
Matthew Trager, Martial Hebert, Jean Ponce
ICCV2
2015 Dense Optical Flow Prediction from a Static Image
abstract
Given a scene, what is going to move, and in what direction will it move? Such a question could be considered a non-semantic form of action prediction. In this work, we present a convolutional neural network (CNN) based approach for motion prediction. Given a static image, this CNN predicts the future motion of each and every pixel in the image in terms of optical flow. Our CNN model leverages the data in tens of thousands of realistic videos to train our model. Our method relies on absolutely no human labeling and is able to predict motion based on the context of the scene. Because our CNN model makes no assumptions about the underlying scene, it can predict future optical flow on a diverse set of scenarios. We outperform all previous approaches by large margins.
Jacob Walker, Abhinav Gupta 0001, Martial Hebert
ICCV3
2015 Mitigating memory requirements for random trees/ferns
abstract
Randomized sets of binary tests have appeared to be quite effective in solving a variety of image processing and vision problems. The exponential growth of their memory usage with the size of the sets however hampers their implementation on the memory-constrained hardware generally available on low-power embedded systems. Our paper addresses this limitation by formulating the conventional semi-naive Bayesian ensemble decision rule in terms of posterior class probabilities, instead of class conditional distributions of binary tests realizations. Subsequent clustering of the posterior class distributions computed at training allows for sharp reduction of large binary tests sets memory footprint, while preserving their high accuracy. Our validation considers a smart metering applicative scenario, and demonstrates that up to 80% of the memory usage can be saved, at constant accuracy.
Christophe De Vleeschouwer, Antoine Legrand, Laurent Jacques, Martial Hebert
ICIP4
2015 Visual chunking: A list prediction framework for region-based object detection
abstract
We consider detecting objects in an image by iteratively selecting from a set of arbitrarily shaped candidate regions. Our generic approach, which we term visual chunking, reasons about the locations of multiple object instances in an image while expressively describing object boundaries. We design an optimization criterion for measuring the performance of a list of such detections as a natural extension to a common per-instance metric. We present an efficient algorithm with provable performance for building a high-quality list of detections from any candidate set of region-based proposals. We also develop a simple class-specific algorithm to generate a candidate region instance in near-linear time in the number of low-level superpixels that outperforms other region generating methods. In order to make predictions on novel images at testing time without access to ground truth, we develop learning approaches to emulate these algorithms' behaviors. We demonstrate that our new approach outperforms sophisticated baselines on benchmark datasets.
Nicholas Rhinehart, Jiaji Zhou, Martial Hebert, J. Andrew Bagnell
ICRA3
2015 Inferring door locations from a teammate's trajectory in stealth human-robot team operations
abstract
Robot perception is generally viewed as the interpretation of data from various types of sensors such as cameras. In this paper, we study indirect perception where a robot can perceive new information by making inferences from non-visual observations of human teammates. As a proof-of-concept study, we specifically focus on a door detection problem in a stealth mission setting where a team operation must not be exposed to the visibility of the team's opponents. We use a special type of the Noisy-OR model known as BN2O model of Bayesian inference network to represent the inter-visibility and to infer the locations of the doors, i.e., potential locations of the opponents. Experimental results on both synthetic data and real person tracking data achieve an F-measure of over .9 on average, suggesting further investigation on the use of non-visual perception in human-robot team operations.
Jean Oh, Luis E. Navarro-Serment, Arne Suppé, Anthony Stentz, Martial Hebert
IROS5
2015 Efficient Model Evaluation with Bilinear Separation Model
abstract
In this paper, we investigate the issue of evaluating efficiently a large set of models on an input image in detection and classification tasks. We show that by formulating the visual task as a large matrix multiplication problem, something that is possible for a broad set of modern detectors and classifiers, we are able to dramatically reduce the rate of growth of computation as the number of models increases. The approach, based on a bilinear separation model, combines standard matrix factorization with a task dependent term which ensures that the resulting smaller size problem maintains performance on the original task. Experiments show that we are able to maintain, or even exceed, the level of performance compared to the default approach of using all the models directly, in both detection and classification tasks. This approach is complementary to other efforts in the literature on speeding up computation through GPU implementation, fast matrix operations, or quantization, in that any of these optimizations can be incorporated.
Fanyi Xiao, Martial Hebert
WACV2
2015 3DNN: 3D Nearest Neighbor
Scott Satkin, Maheen Rashid, Martial Hebert
Int. J. Comput. Vis.4
2015 Data-Driven Objectness
abstract
We propose a data-driven approach to estimate the likelihood that an image segment corresponds to a scene object (its "objectness") by comparing it to a large collection of example object regions. We demonstrate that when the application domain is known, for example, in our case activity of daily living (ADL), we can capture the regularity of the domain specific objects using millions of exemplar object regions. Our approach to estimating the objectness of an image region proceeds in two steps: 1) finding the exemplar regions that are the most similar to the input image segment; 2) calculating the objectness of the image segment by combining segment properties, mutual consistency across the nearest exemplar regions, and the prior probability of each exemplar region. In previous work, parametric objectness models were built from a small number of manually annotated objects regions, instead, our data-driven approach uses 5 million object regions along with their metadata information. Results on multiple data sets demonstrates our data-driven approach compared to the existing model based techniques. We also show the application of our approach in improving the performance of object discovery algorithms.
Hongwen Kang, Martial Hebert, Alexei A. Efros, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Detailed 3D Model Driven Single View Scene Understanding
abstract
We present a data driven approach to holistic scene understanding. From a single image of an indoor scene, our approach estimates its detailed 3D geometry, i.e. The location of its walls and floor, and the 3D appearance of its containing objects, as well as its semantic meaning, i.e. A prediction of what objects it contains. This is made possible by using large datasets of detailed 3D models alongside appearance based detectors. We first estimate the 3D layout of a room, and extrapolate 2D object detection hypotheses to three dimensions to form bounding cuboids. Cuboids are converted to detailed 3D models of the predicted semantic category. Combinations of 3D models are used to create a large list of layout hypotheses for each image -- where each layout hypothesis is semantically meaningful and geometrically plausible. The likelihood of each layout hypothesis is ranked using a learned linear model -- and the hypothesis with the highest predicted likelihood is the final predicted 3D layout. Our approach is able to recover the detailed geometry of scenes, provide precise segmentation of objects in the image plane, and estimate objects' pose in 3D.
Maheen Rashid, Martial Hebert
3DV2
2014 Usability of a Novel Wearable Camera System to Inform Tailored Intervention with Dementia Family Caregivers
Jennifer H. Lingler, Julie Klinger, Laurel Person Mecca, Grace Campbell, Amanda Hunsaker, Sally Hostein, Bernardo Rodrigues Pires, Martial Hebert, Richard Schulz, Judith T. Matthews
AMIA9
2014 Trinocular Geometry Revisited
abstract
When do the visual rays associated with triplets of point correspondences converge, that is, intersect in a common point? Classical models of trinocular geometry based on the fundamental matrices and trifocal tensor associated with the corresponding cameras only provide partial answers to this fundamental question, in large part because of underlying, but seldom explicit, general configuration assumptions. This paper uses elementary tools from projective line geometry to provide necessary and sufficient geometric and analytical conditions for convergence in terms of transversals to triplets of visual rays, without any such assumptions. In turn, this yields a novel and simple minimal parameterization of trinocular geometry for cameras with non-collinear or collinear pinholes.
Jean Ponce, Martial Hebert
CVPR2
2014 Patch to the Future: Unsupervised Visual Prediction
abstract
In this paper we present a conceptually simple but surprisingly powerful method for visual prediction which combines the effectiveness of mid-level visual elements with temporal modeling. Our framework can be learned in a completely unsupervised manner from a large collection of videos. However, more importantly, because our approach models the prediction framework on these mid-level elements, we can not only predict the possible motion in the scene but also predict visual appearances - how are appearances going to change with time. This yields a visual "hallucination" of probable events on top of the scene. We show that our method is able to accurately predict and visualize simple future events, we also show that our approach is comparable to supervised methods for event prediction.
Jacob Walker, Abhinav Gupta 0001, Martial Hebert
CVPR3
2014 Predicting Failures of Vision Systems
abstract
Computer vision systems today fail frequently. They also fail abruptly without warning or explanation. Alleviating the former has been the primary focus of the community. In this work, we hope to draw the community's attention to the latter, which is arguably equally problematic for real applications. We promote two metrics to evaluate failure prediction. We show that a surprisingly straightforward and general approach, that we call ALERT, can predict the likely accuracy (or failure) of a variety of computer vision systems - semantic segmentation, vanishing point and camera parameter estimation, and image memorability prediction - on individual input images. We also explore attribute prediction, where classifiers are typically meant to generalize to new unseen categories. We show that ALERT can be useful in predicting failures of this transfer. Finally, we leverage ALERT to improve the performance of a downstream application of attribute prediction: zero-shot learning. We show that ALERT can outperform several strong baselines for zero-shot learning on four datasets.
Peng Zhang 0023, Jiuling Wang, Ali Farhadi, Martial Hebert, Devi Parikh
CVPR4
2014 Unfolding an Indoor Origami World
David F. Fouhey, Abhinav Gupta 0001, Martial Hebert
ECCV (6)3
2014 Self-explanatory Sparse Representation for Image Classification
Baodi Liu, Yu-Xiong Wang, Bin Shen 0002, Yu-Jin Zhang, Martial Hebert
ECCV (2)5
2014 On Image Contours of Projective Shapes
Jean Ponce, Martial Hebert
ECCV (4)2
2014 Pose Machines: Articulated Pose Estimation via Inference Machines
Varun Ramakrishna, Daniel Munoz, Martial Hebert, J. Andrew Bagnell, Yaser Sheikh
ECCV (2)3
2014 Motion Words for Videos
Ekaterina H. Taralova, Fernando De la Torre, Martial Hebert
ECCV (1)3
2014 Visual sensing for developing autonomous behavior in snake robots
abstract
Snake robots are uniquely qualified to investigate a large variety of settings including archaeological sites, natural disaster zones, and nuclear power plants. For these applications, modular snake robots have been tele-operated to perform specific tasks using images returned to it from an onboard camera in the robots head. In order to give the operator an even richer view of the environment and to enable the robot to perform autonomous tasks we developed a structured light sensor that can make three-dimensional maps of the environment. This paper presents a sensor that is uniquely qualified to meet the severe constraints in size, power and computational footprint of snake robots. Using range data, in the form of 3D pointclouds, we show that it is possible to pair high-level planning with mid-level control to accomplish complex tasks without operator intervention.
Hugo Ponte, Max Queenan, Chaohui Gong, Christoph Mertz, Matthew J. Travers, Florian Enner, Martial Hebert, Howie Choset
ICRA7
2014 Physical querying with multi-modal sensing
abstract
We present Marvin, a system that can search physical objects using a mobile or wearable device. It integrates HOG-based object recognition, SURF-based localization information, automatic speech recognition, and user feedback information with a probabilistic model to recognize the “object of interest” at high accuracy and at interactive speeds. Once the object of interest is recognized, the information that the user is querying, e.g. reviews, options, etc., is displayed on the user's mobile or wearable device. We tested this prototype in a real-world retail store during business hours, with varied degree of background noise and clutter. We show that this multi-modal approach achieves superior recognition accuracy compared to using a vision system alone, especially in cluttered scenes where a vision system would be unable to distinguish which object is of interest to the user without additional input. It is computationally able to scale to large numbers of objects by focusing compute-intensive resources on the objects most likely to be of interest, inferred from user speech and implicit localization information. We present the system architecture, the probabilistic model that integrates the multi-modal information, and empirical results showing the benefits of multi-modal integration.
Iljoo Baek, Taylor Stine, Denver Dash, Fanyi Xiao, Yaser Sheikh, Yair Movshovitz-Attias, Martial Hebert, Takeo Kanade
WACV8
2014 Data-driven exemplar model selection
abstract
We consider the problem of discovering discriminative exemplars suitable for object detection. Due to the diversity in appearance in real world objects, an object detector must capture variations in scale, viewpoint, illumination etc. The current approaches do this by using mixtures of models, where each mixture is designed to capture one (or a few) axis of variation. Current methods usually rely on heuristics to capture these variations; however, it is unclear which axes of variation exist and are relevant to a particular task. Another issue is the requirement of a large set of training images to capture such variations. Current methods do not scale to large training sets either because of training time complexity [31] or test time complexity [26]. In this work, we explore the idea of compactly capturing task-appropriate variation from the data itself. We propose a two stage data-driven process, which selects and learns a compact set of exemplar models for object detection. These selected models have an inherent ranking, which can be used for anytime/budgeted detection scenarios. Another benefit of our approach (beyond the computational speedup) is that the selected set of exemplar models performs better than the entire set.
Ishan Misra, Abhinav Shrivastava, Martial Hebert
WACV3
2014 Occlusion Reasoning for Object Detectionunder Arbitrary Viewpoint
abstract
We present a unified occlusion model for object instance detection under arbitrary viewpoint. Whereas previous approaches primarily modeled local coherency of occlusions or attempted to learn the structure of occlusions from data, we propose to explicitly model occlusions by reasoning about 3D interactions of objects. Our approach accurately represents occlusions under arbitrary viewpoint without requiring additional training data, which can often be difficult to obtain. We validate our model by incorporating occlusion reasoning with the state-of-the-art LINE2D and Gradient Network methods for object instance detection and demonstrate significant improvement in recognizing texture-less objects under severe occlusions.
Edward Hsiao, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Fast and Scalable Approximate Spectral Matching for Higher Order Graph Matching
abstract
This paper presents a fast and efficient computational approach to higher order spectral graph matching. Exploiting the redundancy in a tensor representing the affinity between feature points, we approximate the affinity tensor with the linear combination of Kronecker products between bases and index tensors. The bases and index tensors are highly compressed representations of the approximated affinity tensor, requiring much smaller memory than in previous methods, which store the full affinity tensor. We compute the principal eigenvector of the approximated affinity tensor using the small bases and index tensors without explicitly storing the approximated tensor. To compensate for the loss of matching accuracy by the approximation, we also adopt and incorporate a marginalization scheme that maps a higher order tensor to matrix as well as a one-to-one mapping constraint into the eigenvector computation process. The experimental results show that the proposed method is faster and requires smaller memory than the existing methods with little or no loss of accuracy.
Soonyong Park, Sung-Kee Park, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Gradient Networks: Explicit Shape Matching Without Extracting Edges
abstract
We present a novel framework for shape-based template matching in images. While previous approaches required brittle contour extraction, considered only local information, or used coarse statistics, we propose to match the shape explicitly on low-level gradients by formulating the problem as traversing paths in a gradient network. We evaluate our algorithm on a challenging dataset of objects in cluttered environments and demonstrate significant improvement over state-of-the-art methods for shape matching and object detection.
Edward Hsiao, Martial Hebert
AAAI2
2013 Data-Driven 3D Primitives for Single Image Understanding
abstract
What primitives should we use to infer the rich 3D world behind an image? We argue that these primitives should be both visually discriminative and geometrically informative and we present a technique for discovering such primitives. We demonstrate the utility of our primitives by using them to infer 3D surface normals given a single image. Our technique substantially outperforms the state-of-the-art and shows improved cross-dataset performance.
David F. Fouhey, Abhinav Gupta 0001, Martial Hebert
ICCV3
2013 Style-Aware Mid-level Representation for Discovering Visual Connections in Space and Time
abstract
We present a weakly-supervised visual data mining approach that discovers connections between recurring mid-level visual elements in historic (temporal) and geographic (spatial) image collections, and attempts to capture the underlying visual style. In contrast to existing discovery methods that mine for patterns that remain visually consistent throughout the dataset, our goal is to discover visual elements whose appearance changes due to change in time or location; i.e., exhibit consistent stylistic variations across the label space (date or geo-location). To discover these elements, we first identify groups of patches that are style-sensitive. We then incrementally build correspondences to find the same element across the entire dataset. Finally, we train style-aware regressors that model each element's range of stylistic differences. We apply our approach to date and geo-location prediction and show substantial improvement over several baselines that do not model visual style. We also demonstrate the method's effectiveness on the related task of fine-grained classification.
Yong Jae Lee, Alexei A. Efros, Martial Hebert
ICCV3
2013 3DNN: Viewpoint Invariant 3D Geometry Matching for Scene Understanding
abstract
We present a new algorithm 3DNN (3D Nearest-Neighbor), which is capable of matching an image with 3D data, independently of the viewpoint from which the image was captured. By leveraging rich annotations associated with each image, our algorithm can automatically produce precise and detailed 3D models of a scene from a single image. Moreover, we can transfer information across images to accurately label and segment objects in a scene. The true benefit of 3DNN compared to a traditional 2D nearest-neighbor approach is that by generalizing across viewpoints, we free ourselves from the need to have training examples captured from all possible viewpoints. Thus, we are able to achieve comparable results using orders of magnitude less data, and recognize objects from never-before-seen viewpoints. In this work, we describe the 3DNN algorithm and rigorously evaluate its performance for the tasks of geometry estimation and object detection/segmentation. By decoupling the viewpoint and the geometry of an image, we develop a scene matching approach which is truly 100% viewpoint invariant, yielding state-of-the-art performance on challenging data.
Scott Satkin, Martial Hebert
ICCV2
2013 Exploiting domain knowledge for Object Discovery
abstract
In this paper, we consider the problem of Lifelong Robotic Object Discovery (LROD) as the long-term goal of discovering novel objects in the environment while the robot operates, for as long as the robot operates. As a first step towards LROD, we automatically process the raw video stream of an entire workday of a robotic agent to discover objects. We claim that the key to achieve this goal is to incorporate domain knowledge whenever available, in order to detect and adapt to changes in the environment. We propose a general graph-based formulation for LROD in which generic domain knowledge is encoded as constraints. Our formulation enables new sources of domain knowledge—metadata—to be added dynamically to the system, as they become available or as conditions change. By adding domain knowledge, we discover 2.7· more objects and decrease processing time 190 times. Our optimized implementation, HerbDisc, processes 6 h 20 min of RGBD video of real human environments in 18 min 30 s, and discovers 121 correct novel objects with their 3D models.
Alvaro Collet, Corina Gurau, Martial Hebert, Siddhartha S. Srinivasa
ICRA4
2013 Efficient 3-D scene analysis from streaming data
abstract
Rich scene understanding from 3-D point clouds is a challenging task that requires contextual reasoning, which is typically computationally expensive. The task is further complicated when we expect the scene analysis algorithm to also efficiently handle data that is continuously streamed from a sensor on a mobile robot. Hence, we are typically forced to make a choice between 1) using a precise representation of the scene at the cost of speed, or 2) making fast, though inaccurate, approximations at the cost of increased misclassifications. In this work, we demonstrate that we can achieve the best of both worlds by using an efficient and simple representation of the scene in conjunction with recent developments in structured prediction in order to obtain both efficient and state-of-the-art classifications. Furthermore, this efficient scene representation naturally handles streaming data and provides a 300% to 500% speedup over more precise representations.
Hanzhang Hu, Daniel Munoz, J. Andrew Bagnell, Martial Hebert
ICRA4
2013 Multi-armed recommendation bandits for selecting state machine policies for robotic systems
abstract
We investigate the problem of selecting a state-machine from a library to control a robot. We are particularly interested in this problem when evaluating such state machines on a particular robotics task is expensive. As a motivating example, we consider a problem where a simulated vacuuming robot must select a driving state machine well-suited for a particular (unknown) room layout. By borrowing concepts from collaborative filtering (recommender systems such as Netflix and Amazon.com), we present a multi-armed bandit formulation that incorporates recommendation techniques to efficiently select state machines for individual room layouts. We show that this formulation outperforms the individual approaches (recommendation, multi-armed bandits) as well as the baseline of selecting the `average best' state machine across all rooms.
Pyry Matikainen, P. Michael Furlong, Rahul Sukthankar, Martial Hebert
ICRA4
2013 Efficient temporal consistency for streaming video scene analysis
abstract
We address the problem of image-based scene analysis from streaming video, as would be seen from a moving platform, in order to efficiently generate spatially and temporally consistent predictions of semantic categories over time. In contrast to previous techniques which typically address this problem in batch and/or through graphical models, we demonstrate that by learning visual similarities between pixels across frames, a simple filtering algorithfiltering algorithmm is able to achieve high performance predictions in an efficient and online/causal manner. Our technique is a meta-algorithm that can be efficiently wrapped around any scene analysis technique that produces a per-pixel semantic category distribution. We validate our approach over three different scene analysis techniques on three different datasets that contain different semantic object categories. Our experiments demonstrate that our approach is very efficient in practice and substantially improves the consistency of the predictions over time.
Ondrej Miksik, Daniel Munoz, J. Andrew Bagnell, Martial Hebert
ICRA4
2013 Learning monocular reactive UAV control in cluttered natural environments
abstract
Autonomous navigation for large Unmanned Aerial Vehicles (UAVs) is fairly straight-forward, as expensive sensors and monitoring devices can be employed. In contrast, obstacle avoidance remains a challenging task for Micro Aerial Vehicles (MAVs) which operate at low altitude in cluttered environments. Unlike large vehicles, MAVs can only carry very light sensors, such as cameras, making autonomous navigation through obstacles much more challenging. In this paper, we describe a system that navigates a small quadrotor helicopter autonomously at low altitude through natural forest environments. Using only a single cheap camera to perceive the environment, we are able to maintain a constant velocity of up to 1.5m/s. Given a small set of human pilot demonstrations, we use recent state-of-the-art imitation learning techniques to train a controller that can avoid trees by adapting the MAVs heading. We demonstrate the performance of our system in a more controlled environment indoors, and in real natural forest environments outdoors.
Stéphane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar, Andreas Wendel, Debadeepta Dey, J. Andrew Bagnell, Martial Hebert
ICRA7
2013 Fast and scalable approximate spectral graph matching for correspondence problems
U Kang, Martial Hebert, Soonyong Park
Inf. Sci.2
2012 Using Expectations to Drive Cognitive Behavior
abstract
Generating future states of the world is an essential component of high-level cognitive tasks such as planning. We explore the notion that such future-state generation is more widespread and forms an integral part of cognition. We call these generated states expectations, and propose that cognitive systems constantly generate expectations, match them to observed behavior and react when a difference exists between the two. We describe an ACT-R model that performs expectation-driven cognition on two tasks – pedestrian tracking and behavior classification. The model generates expectations of pedestrian movements to track them. The model also uses differences in expectations to identify distinctive features that differentiate these tracks. During learning, the model learns the association between these features and the various behaviors. During testing, it classifies pedestrian tracks by recalling the behavior associated with the features of each track. We tested the model on both single and multiple behavior datasets and compared the results against a k-NN classifier. The k-NN classifier outperformed the model in correct classifications, but the model had fewer incorrect classifications in the multiple behavior case, and both systems had about equal incorrect classifications in the single behavior case.
Unmesh Kurup, Christian Lebiere, Anthony Stentz, Martial Hebert
AAAI4
2012 Object Instance Sharing by Enhanced Bounding Box Correspondence
abstract
Most contemporary object detection approaches assume each object instance in the training data to be uniquely represented by a single bounding box. In this paper, we go beyond this conventional view by allowing an object instance to be described by multiple bounding boxes. The new bounding box annotations are determined based on the alignment of an object instance with the other training instances in the dataset. Our proposal enables the training data to be reused multiple times for training richer multi-component category models. We operationalize this idea by two complementary operations: bounding box shrinking, which finds subregions of an object instance that could be shared; and bounding box enlarging, which enlarges object instances to include local contextual cues. We empirically validate our approach on the PASCAL VOC detection dataset.
Santosh Kumar Divvala, Alexei A. Efros, Martial Hebert
BMVC3
2012 Data-Driven Scene Understanding from 3D Models
abstract
In this paper, we propose a data-driven approach to leverage repositories of 3D models for scene understanding. Our ability to relate what we see in an image to a large collection of 3D models allows us to transfer information from these models, creating a rich understanding of the scene. We develop a framework for auto-calibrating a camera, rendering 3D models from the viewpoint an image was taken, and computing a similarity measure between each 3D model and an input image. We demonstrate this data-driven approach in the context of geometry estimation and show the ability to find the identities and poses of object in a scene. Additionally, we present a new dataset with annotated scene geometry. This data allows us to measure the performance of our algorithm in 3D, rather than in the image plane.
Scott Satkin, Martial Hebert
BMVC3
2012 Occlusion reasoning for object detection under arbitrary viewpoint
abstract
We present a unified occlusion model for object instance detection under arbitrary viewpoint. Whereas previous approaches primarily modeled local coherency of occlusions or attempted to learn the structure of occlusions from data, we propose to explicitly model occlusions by reasoning about 3D interactions of objects. Our approach accurately represents occlusions under arbitrary viewpoint without requiring additional training data, which can often be difficult to obtain. We validate our model by extending the state-of-the-art LINE2D method for object instance detection and demonstrate significant improvement in recognizing textureless objects under severe occlusions.
Edward Hsiao, Martial Hebert
CVPR2
2012 Model recommendation for action recognition
abstract
Simply choosing one model out of a large set of possibilities for a given vision task is a surprisingly difficult problem, especially if there is limited evaluation data with which to distinguish among models, such as when choosing the best “walk” action classifier from a large pool of classifiers tuned for different viewing angles, lighting conditions, and background clutter. In this paper we suggest that this problem of selecting a good model can be recast as a recommendation problem, where the goal is to recommend a good model for a particular task based on how well a limited probe set of models appears to perform. Through this conceptual remapping, we can bring to bear all the collaborative filtering techniques developed for consumer recommender systems (e.g., Netflix, Amazon.com). We test this hypothesis on action recognition, and find that even when every model has been directly rated on a training set, recommendation finds better selections for the corresponding test set than the best performers on the training set.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert
CVPR3
2012 Connecting Missing Links: Object Discovery from Sparse Observations Using 5 Million Product Images
Hongwen Kang, Martial Hebert, Alexei A. Efros, Takeo Kanade
ECCV (6)2
2012 Activity Forecasting
Kris Makoto Kitani, Brian D. Ziebart, J. Andrew Bagnell, Martial Hebert
ECCV (4)4
2012 Co-inference for Multi-modal Scene Analysis
Daniel Munoz, J. Andrew Bagnell, Martial Hebert
ECCV (6)3
2012 An integrated system for autonomous robotics manipulation
abstract
We describe the software components of a robotics system designed to autonomously grasp objects and perform dexterous manipulation tasks with only high-level supervision. The system is centered on the tight integration of several core functionalities, including perception, planning and control, with the logical structuring of tasks driven by a Behavior Tree architecture. The advantage of the implementation is to reduce the execution time while integrating advanced algorithms for autonomous manipulation. We describe our approach to 3-D perception, real-time planning, force compliant motions, and audio processing. Performance results for object grasping and complex manipulation tasks of in-house tests and of an independent evaluation team are presented.
J. Andrew Bagnell, Felipe Cavalcanti, Lei Cui 0005, Thomas Galluzzo, Martial Hebert, Moslem Kazemi, Matthew Klingensmith, Jacqueline Libby, Tian Yu Liu, Nancy S. Pollard, Mihail Pivtoraiko, Jean-Sebastien Valois, Ranqi Zhu
IROS5
2012 Unsupervised Learning for Graph Matching
Marius Leordeanu, Rahul Sukthankar, Martial Hebert
Int. J. Comput. Vis.3
2012 First-Person Vision
abstract
For understanding the behavior, intent, and environment of a person, the surveillance metaphor is traditional; that is, install cameras and observe the subject, and his/her interaction with other people and the environment. Instead, we argue that the first-person vision (FPV), which senses the environment and the subject's activities from a wearable sensor, is more advantageous with images about the subject's environment as taken from his/her view points, and with readily available information about head motion and gaze through eye tracking. In this paper, we review key research challenges that need to be addressed to develop such FPV systems, and describe our ongoing work to address them using examples from our prototype systems.
Takeo Kanade, Martial Hebert
Proc. IEEE2
2011 From 3D scene geometry to human workspace
abstract
We present a human-centric paradigm for scene understanding. Our approach goes beyond estimating 3D scene geometry and predicts the "workspace" of a human which is represented by a data-driven vocabulary of human interactions. Our method builds upon the recent work in indoor scene understanding and the availability of motion capture data to create a joint space of human poses and scene geometry by modeling the physical interactions between the two. This joint space can then be used to predict potential human poses and joint locations from a single image. In a way, this work revisits the principle of Gibsonian affordances, reinterpreting it for the modern, data-driven era.
Abhinav Gupta 0001, Scott Satkin, Alexei A. Efros, Martial Hebert
CVPR4
2011 Learning message-passing inference machines for structured prediction
abstract
Nearly every structured prediction problem in computer vision requires approximate inference due to large and complex dependencies among output labels. While graphical models provide a clean separation between modeling and inference, learning these models with approximate inference is not well understood. Furthermore, even if a good model is learned, predictions are often inaccurate due to approximations. In this work, instead of performing inference over a graphical model, we instead consider the inference procedure as a composition of predictors. Specifically, we focus on message-passing algorithms, such as Belief Propagation, and show how they can be viewed as procedures that sequentially predict label distributions at each node over a graph. Given labeled graphs, we can then train the sequence of predictors to output the correct labeling s. The result no longer corresponds to a graphical model but simply defines an inference procedure, with strong theoretical properties, that can be used to classify new graphs. We demonstrate the scalability and efficacy of our approach on 3D point cloud classification and 3D surface estimation from single images.
Stéphane Ross, Daniel Munoz, Martial Hebert, J. Andrew Bagnell
CVPR3
2011 Prop-free pointing detection in dynamic cluttered environments
abstract
Vision-based prop-free pointing detection is challenging both from an algorithmic and a systems standpoint. From a computer vision perspective, accurately determining where multiple users are pointing is difficult in cluttered environments with dynamic scene content. Standard approaches relying on appearance models or background subtraction to segment users operate poorly in this domain. We propose a method that focuses on motion analysis to detect pointing gestures and robustly estimate the pointing direction. Our algorithm is self-initializing; as the user points, we analyze the observed motion from two cameras and infer rotation centers that best explain the observed motion. From these, we group pixel-level flow into dominant pointing vectors that each originate from a rotation center and merge across views to obtain 3D pointing vectors. However, our proposed algorithm is computationally expensive, posing systems challenges even with current computing infrastructure. We achieve interactive speeds by exploiting coarse-grained parallelization over a cluster of computers. In unconstrained environments, we obtain an average angular precision of 2.7°.
Pyry Matikainen, Padmanabhan Pillai, Lily B. Mummert, Rahul Sukthankar, Martial Hebert
FG5
2011 Discovering object instances from scenes of Daily Living
abstract
We propose an approach to identify and segment objects from scenes that a person (or robot) encounters in Activities of Daily Living (ADL). Images collected in those cluttered scenes contain multiple objects. Each image provides only a partial, possibly very different view of each object. An object instance discovery program must be able to link pieces of visual information from multiple images and extract the consistent patterns.
Hongwen Kang, Martial Hebert, Takeo Kanade
ICCV2
2011 Feature seeding for action recognition
abstract
Progress in action recognition has been in large part due to advances in the features that drive learning-based methods. However, the relative sparsity of training data and the risk of overfitting have made it difficult to directly search for good features. In this paper we suggest using synthetic data to search for robust features that can more easily take advantage of limited data, rather than using the synthetic data directly as a substitute for real data. We demonstrate that the features discovered by our selection method, which we call seeding, improve performance on an action classification task on real data, even though the synthetic data from which the features are seeded differs significantly from the real data, both in terms of appearance and the set of action classes.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert
ICCV3
2011 Source constrained clustering
abstract
We consider the problem of quantizing data generated from disparate sources, e.g. subjects performing actions with different styles, movies with particular genre bias, various conditions in which images of objects are taken, etc. These are scenarios where unsupervised clustering produces inadequate codebooks because algorithms like K-means tend to cluster samples based on data biases (e.g. cluster subjects), rather than cluster similar samples across sources (e.g. cluster actions). We propose a new quantization technique, Source Constrained Clustering (SCC), which extends the K-means algorithm by enforcing clusters to group samples from multiple sources. We evaluate the method in the context of activity recognition from videos in an unconstrained environment. Experiments on several tasks and features show that using source information improves classification performance.
Ekaterina H. Taralova, Fernando De la Torre, Martial Hebert
ICCV3
2011 Structure discovery in multi-modal data: A region-based approach
abstract
The ability of a perception system to discern what is important in a scene and what is not is an invaluable asset, with multiple applications in object recognition, people detection and SLAM, among others. In this paper, we aim to analyze all sensory data available to separate a scene into a few physically meaningful parts, which we term structure, while discarding background clutter. In particular, we consider the combination of image and range data, and base our decision in both appearance and 3D shape. Our main contribution is the development of a framework to perform scene segmentation that preserves physical objects using multi-modal data. We combine image and range data using a novel mid-level fusion technique based on the concept of regions that avoids any pixel-level correspondences between data sources. We associate groups of pixels with 3D points into multi-modal regions that we term regionlets, and measure the structure-ness of each regionlet using simple, bottom-up cues from image and range features. We show that the highest-ranked regionlets correspond to the most prominent objects in the scene. We verify the validity of our approach on 105 scenes of household environments.
Alvaro Collet, Siddhartha S. Srinivasa, Martial Hebert
ICRA3
2011 3-D scene analysis via sequenced predictions over points and regions
abstract
We address the problem of understanding scenes from 3-D laser scans via per-point assignment of semantic labels. In order to mitigate the difficulties of using a graphical model for modeling the contextual relationships among the 3-D points, we instead propose a multi-stage inference procedure to capture these relationships. More specifically, we train this procedure to use point cloud statistics and learn relational information (e.g., tree-trunks are below vegetation) over fine (point-wise) and coarse (region-wise) scales. We evaluate our approach on three different datasets, that were obtained from different sensors, and demonstrate improved performance.
Xuehan Xiong, Daniel Munoz, J. Andrew Bagnell, Martial Hebert
ICRA4
2011 Image matching with distinctive visual vocabulary
abstract
In this paper we propose an image indexing and matching algorithm that relies on selecting distinctive high dimensional features. In contrast with conventional techniques that treated all features equally, we claim that one can benefit significantly from focusing on distinctive features. We propose a bag-of-words algorithm that combines the feature distinctiveness in visual vocabulary generation. Our approach compares favorably with the state of the art in image matching tasks on the University of Kentucky Recognition Benchmark dataset and on an indoor localization dataset. We also show that our approach scales up more gracefully on a large scale Flickr dataset.
Hongwen Kang, Martial Hebert, Takeo Kanade
WACV2
2011 Recovering Occlusion Boundaries from an Image
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.3
2010 Representations for object and action recognition
Martial Hebert
BMVC1
2010 Making specific features less discriminative to improve point-based 3D object recognition
abstract
We present a framework that retains ambiguity in feature matching to increase the performance of 3D object recognition systems. Whereas previous systems removed ambiguous correspondences during matching, we show that ambiguity should be resolved during hypothesis testing and not at the matching phase. To preserve ambiguity during matching, we vector quantize and match model features in a hierarchical manner. This matching technique allows our system to be more robust to the distribution of model descriptors in feature space. We also show that we can address recognition under arbitrary viewpoint by using our framework to facilitate matching of additional features extracted from affine transformed model images. The evaluation of our algorithms in 3D object recognition is demonstrated on a difficult dataset of 620 images.
Edward Hsiao, Alvaro Collet, Martial Hebert
CVPR3
2010 Blocks World Revisited: Image Understanding Using Qualitative Geometry and Mechanics
Abhinav Gupta 0001, Alexei A. Efros, Martial Hebert
ECCV (4)3
2010 Representing Pairwise Spatial and Temporal Relations for Action Recognition
Pyry Matikainen, Martial Hebert, Rahul Sukthankar
ECCV (1)2
2010 Stacked Hierarchical Labeling
Daniel Munoz, J. Andrew Bagnell, Martial Hebert
ECCV (6)3
2010 Modeling the Temporal Extent of Actions
Scott Satkin, Martial Hebert
ECCV (1)2
2010 People helping robots helping people: Crowdsourcing for grasping novel objects
abstract
For successful deployment, personal robots must adapt to ever-changing indoor environments. While dealing with novel objects is a largely unsolved challenge in AI, it is easy for people. In this paper we present a framework for robot supervision through Amazon Mechanical Turk. Unlike traditional models of teleoperation, people provide semantic information about the world and subjective judgements. The robot then autonomously utilizes the additional information to enhance its capabilities. The information can be collected on demand in large volumes and at low cost. We demonstrate our approach on the task of grasping unknown objects.
Alexander Sorokin, Dmitry Berenson, Siddhartha S. Srinivasa, Martial Hebert
IROS4
2010 Estimating Spatial Layout of Rooms using Volumetric Reasoning about Objects and Surfaces
abstract
There has been a recent push in extraction of 3D spatial layout of scenes. However, none of these approaches model the 3D interaction between objects and the spatial layout. In this paper, we argue for a parametric representation of objects in 3D, which allows us to incorporate volumetric constraints of the physical world. We show that augmenting current structured prediction techniques with volumetric reasoning significantly improves the performance of the state-of-the-art.
David C. Lee, Abhinav Gupta 0001, Martial Hebert, Takeo Kanade
NIPS3
2010 Volumetric Features for Video Event Detection
Yan Ke, Rahul Sukthankar, Martial Hebert
Int. J. Comput. Vis.3
2009 An empirical study of context in object detection
abstract
This paper presents an empirical evaluation of the role of context in a contemporary, challenging object detection task - the PASCAL VOC 2008. Previous experiments with context have mostly been done on home-grown datasets, often with non-standard baselines, making it difficult to isolate the contribution of contextual information. In this work, we present our analysis on a standard dataset, using top-performing local appearance detectors as baseline. We evaluate several different sources of context and ways to utilize it. While we employ many contextual cues that have been used before, we also propose a few novel ones including the use of geographic context and a new approach for using object spatial support.
Santosh Kumar Divvala, Derek Hoiem, James Hays, Alexei A. Efros, Martial Hebert
CVPR5
2009 Geometric reasoning for single image structure recovery
abstract
We study the problem of generating plausible interpretations of a scene from a collection of line segments automatically extracted from a single indoor image. We show that we can recognize the three dimensional structure of the interior of a building, even in the presence of occluding objects. Several physically valid structure hypotheses are proposed by geometric reasoning and verified to find the best fitting model to line segments, which is then converted to a full 3D model. Our experiments demonstrate that our structure recovery from line segments is comparable with methods using full image appearance. Our approach shows how a set of rules describing geometric constraints between groups of segments can be used to prune scene interpretation hypotheses and to generate the most plausible interpretation.
David C. Lee, Martial Hebert, Takeo Kanade
CVPR2
2009 Unsupervised learning for graph matching
abstract
Graph matching is an important problem in computer vision. It is used in 2D and 3D object matching and recognition. Despite its importance, there is little literature on learning the parameters that control the graph matching problem, even though learning is important for improving the matching rate, as shown by this and other work. In this paper we show for the first time how to perform parameter learning in an unsupervised fashion, that is when no correct correspondences between graphs are given during training. We show empirically that unsupervised learning is comparable in efficiency and quality with the supervised one, while avoiding the tedious manual labeling of ground truth correspondences. We also verify experimentally that this learning method can improve the performance of several state-of-the art graph matching algorithms.
Marius Leordeanu, Martial Hebert
CVPR2
2009 Contextual classification with functional Max-Margin Markov Networks
abstract
We address the problem of label assignment in computer vision: given a novel 3D or 2D scene, we wish to assign a unique label to every site (voxel, pixel, superpixel, etc.). To this end, the Markov Random Field framework has proven to be a model of choice as it uses contextual information to yield improved classification results over locally independent classifiers. In this work we adapt a functional gradient approach for learning high-dimensional parameters of random fields in order to perform discrete, multi-label classification. With this approach we can learn robust models involving high-order interactions better than the previously used learning method. We validate the approach in the context of point cloud classification and improve the state of the art. In addition, we successfully demonstrate the generality of the approach on the challenging vision problem of recovering 3-D geometric surfaces from images.
Daniel Munoz, J. Andrew Bagnell, Nicolas Vandapel, Martial Hebert
CVPR4
2009 Onboard contextual classification of 3-D point clouds with learned high-order Markov Random Fields
abstract
Contextual reasoning through graphical models such as Markov random fields often show superior performance against local classifiers in many domains. Unfortunately, this performance increase is often at the cost of time consuming, memory intensive learning and slow inference at testing time. Structured prediction for 3-D point cloud classification is one example of such an application. In this paper we present two contributions. First we show how efficient learning of a random field with higher-order cliques can be achieved using subgradient optimization. Second, we present a context approximation using random fields with high-order cliques designed to make this model usable online, onboard a mobile vehicle for environment modeling. We obtained results with the mobile vehicle on a variety of terrains, at 1/3 Hz for a map 25 times 50 meters and a vehicle speed of 1-2 m/s.
Daniel Munoz, Nicolas Vandapel, Martial Hebert
ICRA3
2009 Planning-based prediction for pedestrians
abstract
We present a novel approach for determining robot movements that efficiently accomplish the robot's tasks while not hindering the movements of people within the environment. Our approach models the goal-directed trajectories of pedestrians using maximum entropy inverse optimal control. The advantage of this modeling approach is the generality of its learned cost function to changes in the environment and to entirely different environments. We employ the predictions of this model of pedestrian trajectories in a novel incremental planner and quantitatively show the improvement in hindrance-sensitive robot trajectory planning provided by our approach.
Brian D. Ziebart, Nathan D. Ratliff, Garratt Gallagher, Christoph Mertz, Kevin M. Peterson, J. Andrew Bagnell, Martial Hebert, Anind K. Dey, Siddhartha S. Srinivasa
IROS7
2009 An Integer Projected Fixed Point Method for Graph Matching and MAP Inference
abstract
Graph matching and MAP inference are essential problems in computer vision and machine learning. We introduce a novel algorithm that can accommodate both problems and solve them efficiently. Recent graph matching algorithms are based on a general quadratic programming formulation, that takes in consideration both unary and second-order terms reflecting the similarities in local appearance as well as in the pairwise geometric relationships between the matched features. In this case the problem is NP-hard and a lot of effort has been spent in finding efficiently approximate solutions by relaxing the constraints of the original problem. Most algorithms find optimal continuous solutions of the modified problem, ignoring during the optimization the original discrete constraints. The continuous solution is quickly binarized at the end, but very little attention is put into this final discretization step. In this paper we argue that the stage in which a discrete solution is found is crucial for good performance. We propose an efficient algorithm, with climbing and convergence properties, that optimizes in the discrete domain the quadratic score, and it gives excellent results either by itself or by starting from the solution returned by any graph matching algorithm. In practice it outperforms state-or-the art algorithms and it also significantly improves their performance if used in combination. When applied to MAP inference, the algorithm is a parallel extension of Iterated Conditional Modes (ICM) with climbing and convergence properties that make it a compelling alternative to the sequential ICM. In our experiments on MAP inference our algorithm proved its effectiveness by outperforming ICM and Max-Product Belief Propagation.
Marius Leordeanu, Martial Hebert, Rahul Sukthankar
NIPS2
2009 Occlusion Boundaries from Motion: Low-Level Detection and Mid-Level Reasoning
Andrew N. Stein, Martial Hebert
Int. J. Comput. Vis.2
2009 Local detection of occlusion boundaries in video
Andrew N. Stein, Martial Hebert
Image Vis. Comput.2
2008 Fast Motion Consistency through Matrix Quantization
abstract
Determining the motion consistency between two video clips is a key component for many applications such as video event detection and human pose estimation. Shechtman and Irani recently proposed a method for measuring the motion consistency between two videos by representing the motion about each point with a space-time Harris matrix of spatial and temporal derivatives. A motion-consistency measure can be accurately estimated without explicitly calculating the optical flow from the videos, which could be noisy. However, the motion consistency calculation is computationally expensive and it must be evaluated between all possible pairs of points between the two videos. We propose a novel quantization method for the space-time Harris matrices that reduces the consistency calculation to a fast table lookup for any arbitrary consistency measure. We demonstrate that for the continuous rank drop consistency measure used by Shechtman and Irani, our quantization method is much faster and achieves the same accuracy as the existing approximation.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert, Yan Ke
BMVC3
2008 Closing the loop in scene interpretation
abstract
Image understanding involves analyzing many different aspects of the scene. In this paper, we are concerned with how these tasks can be combined in a way that improves the performance of each of them. Inspired by Barrow and Tenenbaum, we present a flexible framework for interfacing scene analysis processes using intrinsic images. Each intrinsic image is a registered map describing one characteristic of the scene. We apply this framework to develop an integrated 3D scene understanding system with estimates of surface orientations, occlusion boundaries, objects, camera viewpoint, and relative depth. Our experiments on a set of 300 outdoor images demonstrate that these tasks reinforce each other, and we illustrate a coherent scene understanding with automatically reconstructed 3D models.
Derek Hoiem, Alexei A. Efros, Martial Hebert
CVPR3
2008 Unsupervised modeling of object categories using link analysis techniques
abstract
We propose an approach for learning visual models of object categories in an unsupervised manner in which we first build a large-scale complex network which captures the interactions of all unit visual features across the entire training set and we infer information, such as which features are in which categories, directly from the graph by using link analysis techniques. The link analysis techniques are based on well-established graph mining techniques used in diverse applications such as WWW, bioinformatics, and social networks. The techniques operate directly on the patterns of connections between features in the graph rather than on statistical properties, e.g., from clustering in feature space. We argue that the resulting techniques are simpler, and we show that they perform similarly or better compared to state of the art techniques on common data sets. We also show results on more challenging data sets than those that have been used in prior work on unsupervised modeling.
Gunhee Kim, Christos Faloutsos, Martial Hebert
CVPR3
2008 Smoothing-based Optimization
abstract
We propose an efficient method for complex optimization problems that often arise in computer vision. While our method is general and could be applied to various tasks, it was mainly inspired from problems in computer vision, and it borrows ideas from scale space theory. One of the main motivations for our approach is that searching for the global maximum through the scale space of a function is equivalent to looking for the maximum of the original function, with the advantage of having to handle fewer local optima. Our method works with any non-negative, possibly non-smooth function, and requires only the ability of evaluating the function at any specific point. The algorithm is based on a growth transformation, which is guaranteed to increase the value of the scale space function at every step, unlike gradient methods. To demonstrate its effectiveness we present its performance on a few computer vision applications, and show that in our experiments it is more effective than some well established methods such as MCMC, Simulated Annealing and the more local Nelder-Mead optimization method.
Marius Leordeanu, Martial Hebert
CVPR2
2008 Towards unsupervised whole-object segmentation: Combining automated matting with boundary detection
abstract
We propose a novel step toward the unsupervised segmentation of whole objects by combining ldquohintsrdquo of partial scene segmentation offered by multiple soft, binary mattes. These mattes are implied by a set of hypothesized object boundary fragments in the scene. Rather than trying to find or define a single ldquobestrdquo segmentation, we generate multiple segmentations of an image. This reflects contemporary methods for unsupervised object discovery from groups of images, and it allows us to define intuitive evaluation metrics for our sets of segmentations based on the accurate and parsimonious delineation of scene objects. Our proposed approach builds on recent advances in spectral clustering, image matting, and boundary detection. It is demonstrated qualitatively and quantitatively on a dataset of scenes and is suitable for current work in unsupervised object discovery without top-down knowledge.
Andrew N. Stein, Thomas S. Stepleton, Martial Hebert
CVPR3
2008 Discriminative Sparse Image Models for Class-Specific Edge Detection and Image Interpretation
Julien Mairal, Marius Leordeanu, Francis R. Bach, Martial Hebert, Jean Ponce
ECCV (3)4
2008 Object Recognition by Integrating Multiple Image Segmentations
Caroline Pantofaru, Cordelia Schmid, Martial Hebert
ECCV (3)3
2008 Segmentation of Salient Regions in Outdoor Scenes Using Imagery and 3-D Data
abstract
This paper describes a segmentation method for extracting salient regions in outdoor scenes using both 3-D laser scans and imagery information. Our approach is a bottom- up attentive process without any high-level priors, models, or learning. As a mid-level vision task, it is not only robust against noise and outliers but it also provides valuable information for other high-level tasks in the form of optimal segments and their ranked saliency. In this paper, we propose a new saliency definition for 3-D point clouds and we incorporate it with saliency features from color information.
Gunhee Kim, Daniel F. Huber, Martial Hebert
WACV3
2008 Putting Objects in Perspective
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.3
2007 A Framework for Learning to Recognize and Segment Object Classes using Weakly Supervised Training Data
abstract
The continual improvement of object recognition systems has resulted in an increased demand for their application to problems which require an exact pixel-level object segmentation. In this paper, we illustrate an example of an object class recognition and segmentation system which is trained using weakly supervised training data, with the goal of examining the influence that different model choices can have on its performance. In order to achieve pixel-level labeling for rigid and deformable objects, we employ regions generated by unsupervised segmentation as the spatial support for our image features, and explore model selection issues related to their representation. Numerical results for pixel-level accuracy are presented on two challenging and varied datasets.
Caroline Pantofaru, Martial Hebert
BMVC2
2007 Combining Local Appearance and Motion Cues for Occlusion Boundary Detection
abstract
Building on recent advances in the detection of appearance edges from multiple local cues, we present an approach for detecting occlusion boundaries which also incorporates local motion information. We argue that these boundaries have physical significance which makes them important for many high-level vision tasks and that motion offers a unique, often critical source of additional information for detecting them. We provide a new dataset of natural image sequences with labeled occlusion boundaries, on which we learn a classifier that leverages appearance cues along with motion estimates from either side of an edge. We demonstrate improved performance for pixelwise differentiation of occlusion boundaries from non-occluding edges by combining these weak local cues, as compared to using them separately. The results are suitable as improved input to subsequent mid- or high-level reasoning methods.
Andrew N. Stein, Martial Hebert
BMVC2
2007 Denoising Manifold and Non-Manifold Point Clouds
abstract
The faithful reconstruction of 3-D models from irregular and noisy point samples is a task central to many applications of computer vision and graphics. We present an approach to denoising that naturally handles intersections of manifolds, thus preserving high-frequency details without oversmoothing. This is accomplished through the use of a modified locally weighted regression algorithm that models a neighborhood of points as an implicit product of linear subspaces. By posing the problem as one of energy minimization subject to constraints on the coefficients of a higher order polynomial, we can also incorporate anisotropic error models appropriate for data acquired with a range sensor. We demonstrate the effectiveness of our approach through some preliminary results in denoising synthetic data in 2-D and 3-D domains.
Ranjith Unnikrishnan, Martial Hebert
BMVC2
2007 Spatio-temporal Shape and Flow Correlation for Action Recognition
abstract
This paper explores the use of volumetric features for action recognition. First, we propose a novel method to correlate spatio-temporal shapes to video clips that have been automatically segmented. Our method works on over-segmented videos, which means that we do not require background subtraction for reliable object segmentation. Next, we discuss and demonstrate the complementary nature of shape- and flow-based features for action recognition. Our method, when combined with a recent flow-based correlation technique, can detect a wide range of actions in video, as demonstrated by results on a long tennis video. Although not specifically designed for whole-video classification, we also show that our method's performance is competitive with current action classification techniques on a standard video classification dataset.
Yan Ke, Rahul Sukthankar, Martial Hebert
CVPR3
2007 Beyond Local Appearance: Category Recognition from Pairwise Interactions of Simple Features
abstract
We present a discriminative shape-based algorithm for object category localization and recognition. Our method learns object models in a weakly-supervised fashion, without requiring the specification of object locations nor pixel masks in the training data. We represent object models as cliques of fully-interconnected parts, exploiting only the pairwise geometric relationships between them. The use of pairwise relationships enables our algorithm to successfully overcome several problems that are common to previously-published methods. Even though our algorithm can easily incorporate local appearance information from richer features, we purposefully do not use them in order to demonstrate that simple geometric relationships can match (or exceed) the performance of state-of-the-art object recognition algorithms.
Marius Leordeanu, Martial Hebert, Rahul Sukthankar
CVPR2
2007 Recovering Occlusion Boundaries from a Single Image
abstract
Occlusion reasoning, necessary for tasks such as navigation and object search, is an important aspect of everyday life and a fundamental problem in computer vision. We believe that the amazing ability of humans to reason about occlusions from one image is based on an intrinsically 3D interpretation. In this paper, our goal is to recover the occlusion boundaries and depth ordering of free-standing structures in the scene. Our approach is to learn to identify and label occlusion boundaries using the traditional edge and region cues together with 3D surface and depth cues. Since some of these cues require good spatial support (i.e., a segmentation), we gradually create larger regions and use them to improve inference over the boundaries. Our experiments demonstrate the power of a scene-based approach to occlusion reasoning.
Derek Hoiem, Andrew N. Stein, Alexei A. Efros, Martial Hebert
ICCV4
2007 Event Detection in Crowded Videos
abstract
Real-world actions occur often in crowded, dynamic environments. This poses a difficult challenge for current approaches to video event detection because it is difficult to segment the actor from the background due to distracting motion from other objects in the scene. We propose a technique for event recognition in crowded videos that reliably identifies actions in the presence of partial occlusion and background clutter. Our approach is based on three key ideas: (1) we efficiently match the volumetric representation of an event against oversegmented spatio-temporal video volumes; (2) we augment our shape-based features using flow; (3) rather than treating an event template as an atomic entity, we separately match by parts (both in space and time), enabling robustness against occlusions and actor variability. Our experiments on human actions, such as picking up a dropped object or waving in a crowd show reliable detection with few false positives.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV3
2007 Learning to Find Object Boundaries Using Motion Cues
abstract
While great strides have been made in detecting and localizing specific objects in natural images, the bottom-up segmentation of unknown, generic objects remains a difficult challenge. We believe that occlusion can provide a strong cue for object segmentation and "pop-out", but detecting an object's occlusion boundaries using appearance alone is a difficult problem in itself. If the camera or the scene is moving, however, that motion provides an additional powerful indicator of occlusion. Thus, we use standard appearance cues (e.g. brightness/color gradient) in addition to motion cues that capture subtle differences in the relative surface motion (i.e. parallax) on either side of an occlusion boundary. We describe a learned local classifier and global inference approach which provide a frame-work for combining and reasoning about these appearance and motion cues to estimate which region boundaries of an initial over-segmentation correspond to object/occlusion boundaries in the scene. Through results on a dataset which contains short videos with labeled boundaries, we demonstrate the effectiveness of motion cues for this task.
Andrew N. Stein, Derek Hoiem, Martial Hebert
ICCV3
2007 Potential negative obstacle detection by occlusion labeling
abstract
In this paper, we present an approach for potential negative obstacle detection, based on missing data interpretation that extends traditional techniques driven by data only, which capture the occupancy of the scene. The approach is decomposed into three steps: three-dimensional (3D) data accumulation and low level classification, 3D occluder propagation, and context-based occlusion labeling. The approach is validated using logged laser data collected in various outdoor natural terrains and also demonstrated live on-board the Demo-III experimental unmanned vehicle (XUV).
Nicholas Heckman, Jean-François Lalonde, Nicolas Vandapel, Martial Hebert
IROS4
2007 Background estimation under rapid gain change in thermal imagery
Hulya Yalcin, Robert T. Collins, Martial Hebert
Comput. Vis. Image Underst.3
2007 Editorial: Special Issue on Vision and Robotics, Parts I and II
Gregory D. Hager, Martial Hebert, Seth Hutchinson 0001
Int. J. Comput. Vis.2
2007 Recovering Surface Layout from an Image
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.3
2007 Toward Objective Evaluation of Image Segmentation Algorithms
abstract
Unsupervised image segmentation is an important component in many image understanding algorithms and practical vision systems. However, evaluation of segmentation algorithms thus far has been largely subjective, leaving a system designer to judge the effectiveness of a technique based only on intuition and results in the form of a few example segmented images. This is largely due to image segmentation being an ill-defined problem-there is no unique ground-truth segmentation of an image against which the output of an algorithm may be compared. This paper demonstrates how a recently proposed measure of similarity, the Normalized Probabilistic Rand (NPR) index, can be used to perform a quantitative comparison between image segmentation algorithms using a hand-labeled set of ground-truth segmentations. We show that the measure allows principled comparisons between segmentations created by different algorithms, as well as segmentations on different images. We outline a procedure for algorithm evaluation through an example evaluation of some familiar algorithms-the mean-shift-based algorithm, an efficient graph-based segmentation algorithm, a hybrid algorithm that combines the strengths of both methods, and expectation maximization. Results are presented on the 300 images in the publicly available Berkeley Segmentation Data Set.
Ranjith Unnikrishnan, Caroline Pantofaru, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Local Detection of Occlusion Boundaries in Video
abstract
Occlusion boundaries are notoriously difficult for many patch-based computer vision algorithms, but they also provide potentially useful information about scene structure and shape. Using short video clips, we present a novel method for scoring the degree to which occlusion is visible at detected edges. We first utilise a spatio-temporal edge detector which estimates edge strength, orientation, and normal motion. By then extracting patches from either side of each detected (possibly moving) edge pixel, we can estimate and compare motion to determine if occlusion is present. In experiments on synthetic and natural images, we demonstrate our ability to differentiate occlusion boundary pixels from simple edge pixels by using motion information. In terms of precision versus recall, our occlusion scoring metric outperforms a rank-based motion inconsistency measure from the literature. The completely local, bottom-up approach described here is intended to provide powerful low-level information for use by higher-level reasoning methods.
Andrew N. Stein, Martial Hebert
BMVC2
2006 Extracting Scale and Illuminant Invariant Regions through Color
abstract
Despite the fact that color is a powerful cue in object recognition, the extraction of scale-invariant interest regions from color images frequently begins with a conversion of the image to grayscale. The isolation of interest points is then completely determined by luminance, and the use of color is deferred to the stage of descriptor formation. This seemingly innocuous conversion to grayscale is known to suppress saliency and can lead to representative regions being undetected by procedures based only on luminance. Furthermore, grayscaled images of the same scene under even slightly different illuminants can appear sufficiently different as to affect the repeatability of detections across images. We propose a method that combines information from the color channels to drive the detection of scale-invariant keypoints. By factoring out the local effect of the illuminant using an expressive linear model, we demonstrate robustness to a change in the illuminant without having to estimate its properties from the image. Results are shown on challenging images from two commonly used color constancy datasets.
Ranjith Unnikrishnan, Martial Hebert
BMVC2
2006 Putting Objects in Perspective
abstract
Image understanding requires not only individually estimating elements of the visual world but also capturing the interplay among them. In this paper, we provide a framework for placing local object detection in the context of the overall 3D scene by modeling the interdependence of objects, surface orientations, and camera viewpoint. Most object detection methods consider all scales and locations in the image as equally likely. We show that with probabilistic estimates of 3D geometry, both in terms of surfaces and world coordinates, we can put objects into perspective and model the scale and location variance in the image. Our approach reflects the cyclical nature of the problem by allowing probabilistic object hypotheses to refine geometry and vice-versa. Our framework allows painless substitution of almost any object detector and is easily extended to include other aspects of image understanding. Our results confirm the benefits of our integrated approach.
Derek Hoiem, Alexei A. Efros, Martial Hebert
CVPR (2)3
2006 Efficient MAP approximation for dense energy functions
abstract
We present an efficient method for maximizing energy functions with first and second order potentials, suitable for MAP labeling estimation problems that arise in undirected graphical models. Our approach is to relax the integer constraints on the solution in two steps. First we efficiently obtain the relaxed global optimum following a procedure similar to the iterative power method for finding the largest eigenvector of a matrix. Next, we map the relaxed optimum on a simplex and show that the new energy obtained has a certain optimal bound. Starting from this energy we follow an efficient coordinate ascent procedure that is guaranteed to increase the energy at every step and converge to a solution that obeys the initial integral constraints. We also present a sufficient condition for ascent procedures that guarantees the increase in energy at every step.
Marius Leordeanu, Martial Hebert
ICML2
2006 Opportunistic Use of Vision to Push Back the Path-Planning Horizon
abstract
Mobile robots need maps or other forms of geometric information about the environment to navigate. The mobility sensors (LADAR, stereo, etc.) on these robotic vehicles can however populate these maps only up to a distance of a few tens of meters. A navigation system has no knowledge about the world beyond this sensing horizon. As a result, path planners that rely only on this knowledge are unable to anticipate obstacles sufficiently early and have no choice but to resort to an inefficient local obstacle avoidance behavior. However, recent developments in the computer vision community allows us to collect geometric information about the environment far beyond this sensing horizon. The coarse 3D geometric estimation that can be recovered is derived from an appearance-based model. That uses a multiple-hypothesis framework to robustly estimate scene structure from a single image and estimating confidences for each geometric label. This 3D geometric estimation is used with a previously presented navigation strategy that reasons about sensor constraints and plans for measurements while navigating towards the goal. The validity of the sensing method and navigation strategy is supported by results from simulations as well as field experiments with a real robotic platform. These results also show that significant reduction in path length can be achieved by using this framework
Bart C. Nabbe, Derek Hoiem, Alexei A. Efros, Martial Hebert
IROS4
2006 Discriminative Random Fields
Sanjiv Kumar, Martial Hebert
Int. J. Comput. Vis.2
2006 Rapid Object Indexing Using Locality Sensitive Hashing and Joint 3D-Signature Space Estimation
abstract
We propose a new method for rapid 3D object indexing that combines feature-based methods with coarse alignment-based matching techniques. Our approach achieves a sublinear complexity on the number of models, maintaining at the same time a high degree of performance for real 3D sensed data that is acquired in largely uncontrolled settings. The key component of our method is to first index surface descriptors computed at salient locations from the scene into the whole model database using the Locality Sensitive Hashing (LSH), a probabilistic approximate nearest neighbor method. Progressively complex geometric constraints are subsequently enforced to further prune the initial candidates and eliminate false correspondences due to inaccuracies in the surface descriptors and the errors of the LSH algorithm. The indexed models are selected based on the MAP rule using posterior probability of the models estimated in the joint 3D-signature space. Experiments with real 3D data employing a large database of vehicles, most of them very similar in shape, containing 1,000,000 features from more than 365 models demonstrate a high degree of performance in the presence of occlusion and obscuration, unmodeled vehicle interiors and part articulations, with an average processing time between 50 and 100 seconds per query.
Bogdan Matei, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.7
2005 A Flow-Based Approach to Vehicle Detection and Background Mosaicking in Airborne Video
abstract
In this work, we address the detection of vehicles in a video stream obtained from a moving airborne platform. We propose a Bayesian framework for estimating dense optical flow over time that explicitly estimates a persistent model of background appearance. The approach assumes that the scene can be described by background and occlusion layers, estimated within an expectation-maximization framework. The mathematical formulation of the paper is an extension of the work in (H. Yalcin et al., 2005) where motion and appearance models for foreground and background layers are estimated simultaneously in a Bayesian framework.
Hulya Yalcin, Martial Hebert, Robert T. Collins, Michael J. Black
CVPR (2)2
2005 Geometric Context from a Single Image
abstract
Many computer vision algorithms limit their performance by ignoring the underlying 3D geometric structure in the image. We show that we can estimate the coarse geometric properties of a scene by learning appearance-based models of geometric classes, even in cluttered natural scenes. Geometric classes describe the 3D orientation of an image region with respect to the camera. We provide a multiple-hypothesis framework for robustly estimating scene structure from a single image and obtaining confidences for each geometric label. These confidences can then be used to improve the performance of many other applications. We provide a thorough quantitative evaluation of our algorithm on a set of outdoor images and demonstrate its usefulness in two applications: object detection and automatic single-view reconstruction.
Derek Hoiem, Alexei A. Efros, Martial Hebert
ICCV3
2005 Efficient Visual Event Detection Using Volumetric Features
abstract
This paper studies the use of volumetric features as an alternative to popular local descriptor approaches for event detection in video sequences. Motivated by the recent success of similar ideas in object detection on static images, we generalize the notion of 2D box features to 3D spatio-temporal volumetric features. This general framework enables us to do real-time video analysis. We construct a realtime event detector for each action of interest by learning a cascade of filters based on volumetric features that efficiently scans video sequences in space and time. This event detector recognizes actions that are traditionally problematic for interest point methods - such as smooth motions where insufficient space-time interest points are available. Our experiments demonstrate that the technique accurately detects actions on real-world sequences and is robust to changes in viewpoint, scale and action speed. We also adapt our technique to the related task of human action classification and confirm that it achieves performance comparable to a current interest point based human activity recognizer on a standard database of human activities.
Yan Ke, Rahul Sukthankar, Martial Hebert
ICCV3
2005 A Hierarchical Field Framework for Unified Context-Based Classification
abstract
We present a two-layer hierarchical formulation to exploit different levels of contextual information in images for robust classification. Each layer is modeled as a conditional field that allows one to capture arbitrary observation-dependent label interactions. The proposed framework has two main advantages. First, it encodes both the short-range interactions (e.g., pixelwise label smoothing) as well as the long-range interactions (e.g., relative configurations of objects or regions) in a tractable manner. Second, the formulation is general enough to be applied to different domains ranging from pixelwise image labeling to contextual object detection. The parameters of the model are learned using a sequential maximum-likelihood approximation. The benefits of the proposed framework are demonstrated on four different datasets and comparison results are presented
Sanjiv Kumar, Martial Hebert
ICCV2
2005 A Spectral Technique for Correspondence Problems Using Pairwise Constraints
abstract
We present an efficient spectral method for finding consistent correspondences between two sets of features. We build the adjacency matrix M of a graph whose nodes represent the potential correspondences and the weights on the links represent pairwise agreements between potential correspondences. Correct assignments are likely to establish links among each other and thus form a strongly connected cluster. Incorrect correspondences establish links with the other correspondences only accidentally, so they are unlikely to belong to strongly connected clusters. We recover the correct assignments based on how strongly they belong to the main cluster of M, by using the principal eigenvector of M and imposing the mapping constraints required by the overall correspondence mapping (one-to-one or one-to-many). The experimental evaluation shows that our method is robust to outliers, accurate in terms of matching rate, while being much faster than existing methods
Marius Leordeanu, Martial Hebert
ICCV2
2005 Analysis and Removal of Artifacts in 3-D LADAR Data
abstract
Errors in laser based range measurements can be divided into two categories: intrinsic sensor errors (range drift with temperature, systematic and random errors), and errors due to the interaction of the laser beam with the environment. The former have traditionally received attention and can be modeled. The latter in contrast have long been observed but not well characterized. We propose to do so in this paper. In addition, we present a sensor independent method to remove such artifacts. The objective is to improve the overall quality of 3-D scene reconstruction to perform terrain classification of scenes with vegetation.
John Tuley, Nicolas Vandapel, Martial Hebert
ICRA3
2005 Automatic photo pop-up
abstract
This paper presents a fully automatic method for creating a 3D model from a single photograph. The model is made up of several texture-mapped planar billboards and has the complexity of a typical children's pop-up book illustration. Our main insight is that instead of attempting to recover precise geometry, we statistically modelgeometric classesdefined by their orientations in the scene. Our algorithm labels regions of the input image into coarse categories: "ground", "sky", and "vertical". These labels are then used to "cut and fold" the image into a pop-up model using a set of simple assumptions. Because of the inherent ambiguity of the problem and the statistical nature of the approach, the algorithm is not expected to work on every image. However. it performs surprisingly well for a wide range of scenes taken from a typical person's photo album.
Derek Hoiem, Alexei A. Efros, Martial Hebert
ACM Trans. Graph.3
2004 A Hybrid Object-Level/Pixel-Level Framework For Shape-based Recognition
abstract
This paper presents a technique for shape-based recognition that fuses pixellevel and object-level approaches into a unified framework. A pixel-level algorithm classifies individual pixels as belonging to a target object or clutter based on automatically-selected shape features computed in a spatial arrangement around them; an object-level algorithm classifies object-sized rectangular image regions as objects or clutter by aggregating pixel classifier scores in the regions. We train a cascade of interleaved pixel-level and objectlevel modules to quickly localize complex-shaped objects in highly cluttered scenes under arbitrary out-of-image-plane rotation. Experimental results on a large set of real, highly-cluttered images of a common object under arbitrary out of image plane rotation demonstrate improvements over cascades of strictly pixel-level modules. 1
Owen T. Carmichael, Martial Hebert
BMVC2
2004 Parts-Based 3D Object Classification
Daniel F. Huber, Anuj Kapuria, Raghavendra Donamukkala, Martial Hebert
CVPR (2)4
2004 Linear Model Hashing and Batch RANSAC for Rapid and Accurate Object Recognition
Ying Shan, Bogdan Matei, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert
CVPR (2)6
2004 Enabling Learning from Large Datasets: Applying Active Learning to Mobile Robotics
abstract
Autonomous navigation in outdoor, off-road environments requires solving complex classification problems. Obstacle detection, road following and terrain classification are examples of tasks which have been successfully approached using supervised machine learning techniques for classification. Large amounts of training data are usually necessary in order to achieve satisfactory generalization. In such cases, manually labeling data becomes an expensive and tedious process. This work describes a method for reducing the amount of data that needs to be presented to a human trainer. The algorithm relies on kernel density estimation in order to identify "interesting" scenes in a dataset. Our method does not require any interaction with a human expert for selecting the images, and only minimal amounts of tuning are necessary. We demonstrate its effectiveness in several experiments using data collected with two different vehicles. We first show that our method automatically selects those scenes from a large dataset that a person would consider "important" for classification tasks. Secondly, we show that by labeling only few of the images selected by our method, we obtain classification performance that is comparable to the one reached after labeling hundreds of images from the same dataset.
Cristian Dima, Martial Hebert, Anthony Stentz
ICRA2
2004 Classifier Fusion for Outdoor Obstacle Detection
abstract
This work describes an approach for using several levels of data fusion in the domain of autonomous off-road navigation. We are focusing on outdoor obstacle detection, and we present techniques that leverage on data fusion and machine learning for increasing the reliability of obstacle detection systems. We are combining color and infrared (IR) imagery with range information from a laser range finder. We show that in addition to fusing data at the pixel level, performing high level classifier fusion is beneficial in our domain. Our general approach is to use machine learning techniques for automatically deriving effective models of the classes of interest (obstacle and non-obstacle for example). We train classifiers on different subsets of the features we extract from our sensor suite and show how different classifier fusion schemes can be applied for obtaining a multiple classifier system that is more robust than any of the classifiers presented as input. We present experimental results we obtained on data collected with both the experimental unmanned vehicle (XUV) and a CMU developed robotic tractor.
Cristian Dima, Nicolas Vandapel, Martial Hebert
ICRA3
2004 Natural Terrain Classification using 3-D Ladar Data
abstract
Because of the difficulty of interpreting laser data in a meaningful way, safe navigation in vegetated terrain is still a daunting challenge. In this paper, we focus on the segmentation of ladar data using local 3-D point statistics into three classes: clutter to capture grass and tree canopy, linear to capture thin objects like wires or tree branches, and finally surface to capture solid objects like ground terrain surface, rocks or tree trunks. We present the details of the method proposed, the modifications we made to implement it on-board an autonomous ground vehicle. Finally, we present results from field tests using this rover and results produced from different stationary laser sensors.
Nicolas Vandapel, Daniel F. Huber, Anuj Kapuria, Martial Hebert
ICRA4
2004 Path planning with hallucinated worlds
abstract
We describe an approach that integrates midrange sensing into a dynamic path planning algorithm. The algorithm is based on measuring the reduction in path cost that would be caused by taking a sensor reading from candidate locations. The planner uses this measure in order to decide where to take the next sensor reading. Ideally, one would like to evaluate a path based on a map that is as close as possible to the true underlying world. In practice, however, the map is only sparsely populated by data derived from sensor readings. A key component of the approach described in this paper is a mechanism to infer (or "hallucinate") more complete maps from sparse sensor readings. We show how this hallucination mechanism is integrated with the planner to produce better estimates of the gain in path cost occurred when taking sensor readings. We show results on a real robot as well as a statistical analysis on a large set of randomly generated path planning problems on elevation maps from real terrain.
Bart C. Nabbe, Sanjiv Kumar, Martial Hebert
IROS3
2004 Finding organized structures in 3-D ladar data
abstract
In this paper, we address the problem of finding organized thin structures in three-dimensional (3-D) data. Linear and planar structures segmentation received much attention but thin structures organized in complex patterns remain a challenge for segmentation algorithms. We are interested especially in the problems posed by repetitive and symmetric structures acquired with a laser range finder. The method relies on 3-D data projections along specific directions and 2-D histograms comparison. The sensitivity of the classification algorithm to the parameter settings is evaluated and a segmentation method proposed. We illustrate our approach with data from a concertina wire in terrain with vegetation.
Nicolas Vandapel, Martial Hebert
IROS2
2004 Shape-Based Recognition of Wiry Objects
abstract
We present an approach to the recognition of complex-shaped objects in cluttered environments based on edge information. We first use example images of a target object in typical environments to train a classifier cascade that determines whether edge pixels in an image belong to an instance of the desired object or the clutter. Presented with a novel image, we use the cascade to discard clutter edge pixels and group the object edge pixels into overall detections of the object. The features used for the edge pixel classification are localized, sparse edge density operations. Experiments validate the effectiveness of the technique for recognition of a set of complex objects in a variety of cluttered indoor scenes under arbitrary out-of-image-plane rotation. Furthermore, our experiments suggest that the technique is robust to variations between training and testing environments and is efficient at runtime.
Owen T. Carmichael, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Shape-based Recognition Of Wiry Objects
abstract
We present an approach to the recognition of complex-shaped objects in cluttered environments based on edge cues. We first use example images of the desired object in typical backgrounds to train a classifier cascade which determines whether edge pixels in an image belong to an instance of the object or the clutter. Presented with a novel image, we use the cascade to discard clutter edge pixels. The features used for this classification are localized, sparse edge density operations. Experiments validate the effectiveness of the technique for recognition of complex objects in cluttered indoor scenes under arbitrary out-of-image-plane rotation.
Owen T. Carmichael, Martial Hebert
CVPR (2)2
2003 3D Modeling Using a Statistical Sensor Model and Stochastic Search
abstract
Accurate and robust registration of multiple three-dimensional (3D) views is crucial for creating digital 3D models of real-world scenes. In this paper, we present a framework for evaluating the quality of model hypotheses during the registration phase. We use maximum likelihood estimation to learn a probabilistic model of registration success. This method provides a principled way to combine multiple measures of registration accuracy. Also, we describe a stochastic algorithm for robustly searching the large space of possible models for the best model hypothesis. This new approach can detect situations in which no solution exists, outputting a set of model parts if a single model using all the views cannot be found. We show results for a large collection of automatically modeled scenes and demonstrate that our algorithm works independently of scene size and the type of range sensor. This work is part of a system we have developed to automate the 3D modeling process for a set of 3D views obtained from unknown sensor viewpoints.
Daniel F. Huber, Martial Hebert
CVPR (1)2
2003 Man-Made Structure Detection in Natural Images using a Causal Multiscale Random Field
abstract
This paper presents a generative model based approach to man-made structure detection in 2D (two-dimensional) natural images. The proposed approach uses a causal multiscale random field suggested by Bouman and Shapiro (1994) as a prior model on the class labels on the image sites. However, instead of assuming the conditional independence of the observed data, we propose to capture the local dependencies in the data using a multiscale feature vector. The distribution of the multiscale feature vectors is modeled as mixture of Gaussians. A set of robust multi-scale features is presented that captures the general statistical properties of man-made structures at multiple scales without relying on explicit edge detection. The proposed approach was validated on real-world images from the Corel data set, and a performance comparison with other techniques is presented.
Sanjiv Kumar, Martial Hebert
CVPR (1)2
2003 The Optimal Distance Measure for Object Detection
abstract
We develop a multi-class object detection framework whose core component is a nearest neighbor search over object part classes. The performance of the overall system is critically dependent on the distance measure used in the nearest neighbor search. A distance measure that minimizes the misclassification risk for the 1-nearest neighbor search can be shown to be the probability that a pair of input image measurements belong to different classes. In practice, we model the optimal distance measure using a linear logistic model that combines the discriminative powers of more elementary distance measures associated with a collection of simple to construct feature spaces like color, texture and local shape properties. Furthermore, in order to perform search over large training sets efficiently, the same framework was extended to find hamming distance measures associated with simple discriminators. By combining this discrete distance model with the continuous model, we obtain a hierarchical distance model that is both fast and accurate. Finally, the nearest neighbor search over object part classes was integrated into a whole object detection system and evaluated against an indoor detection task yielding good results.
Shyjan Mahamud, Martial Hebert
CVPR (1)2
2003 Discriminative Random Fields: A Discriminative Framework for Contextual Interaction in Classification
abstract
In this work we present discriminative random fields (DRFs), a discriminative framework for the classification of image regions by incorporating neighborhood interactions in the labels as well as the observed data. The discriminative random fields offer several advantages over the conventional Markov random field (MRF) framework. First, the DRFs allow to relax the strong assumption of conditional independence of the observed data generally used in the MRF framework for tractability. This assumption is too restrictive for a large number of applications in vision. Second, the DRFs derive their classification power by exploiting the probabilistic discriminative models instead of the generative models used in the MRF framework. Finally, all the parameters in the DRF model are estimated simultaneously from the training data unlike the MRF framework where likelihood parameters are usually learned separately from the field parameters. We illustrate the advantages of the DRFs over the MRF framework in an application of man-made structure detection in natural images taken from the Corel database.
Sanjiv Kumar, Martial Hebert
ICCV2
2003 Minimum Risk Distance Measure for Object Recognition
abstract
The optimal distance measure for a given discrimination task under the nearest neighbor framework has been shown to be the likelihood that a pair of measurements have different class labels [S. Mahamud et al., (2002)]. For implementation and efficiency considerations, the optimal distance measure was approximated by combining more elementary distance measures defined on simple feature spaces. We address two important issues that arise in practice for such an approach: (a) What form should the elementary distance measure in each feature space take? We motivate the need to use the optimal distance measure in simple feature spaces as the elementary distance measures; such distance measures have the desirable property that they are invariant to distance-respecting transformations, (b) How do we combine the elementary distance measures ? We present the precise statistical assumptions under which a linear logistic model holds exactly. We benchmark our model with three other methods on a challenging face discrimination task and show that our approach is competitive with the state of the art.
Shyjan Mahamud, Martial Hebert
ICCV2
2003 Fast 3D tracking of non-rigid objects
abstract
A fast 3D tracking method of non-rigid objects is proposed. The method consists of fast ICP and modified RPM techniques for range images. By using the techniques, the authors aim real time 3D non-rigid object tracking. The authors developed an actual tracking system and show the effectiveness of the method by some experimental results.
Nobuhiro Okada, Martial Hebert
ICRA2
2003 Where and when to look: how to extend the myopic planning horizon
abstract
In this paper we describe an approach towards integrating mid-range sensing data into a dynamic path planning algorithm. The key problem, sensing for planning is addressed in the context of outdoor navigation. An algorithmic approach is described towards solving these problems and both simulation results and initial experimental results for outdoor navigation using wide baseline stereo data are presented.
Bart C. Nabbe, Martial Hebert
IROS2
2003 Toward generating labeled maps from color and range data for robot navigation
abstract
This paper addresses the problem of extracting information from range and color data acquired by a mobile robot in urban environments. Our approach extracts geometric structures from clouds of 3-D points and regions from the corresponding color images, labels them based on prior models of the objects expected in the environment - buildings in the current experiments - and combines the two sources of information into a composite labeled map. Ultimately, our goal is to generate maps that are segmented into objects of interest, each of which is labeled by its type, e.g., building, vegetation, etc. Such a map provides a higher-level representation of the environment than the geometric maps normally used for mobile robot navigation. The techniques presented here are a step toward the automatic construction of such labeled maps.
Caroline Pantofaru, Ranjith Unnikrishnan, Martial Hebert
IROS3
2003 Robust extraction of multiple structures from non-uniformly sampled data
abstract
The extraction of multiple coherent structures from point clouds is crucial to the problem of scene modeling. While many statistical methods exist for robust estimation from noisy data, they are inadequate for addressing issues of scale, semi-structured clutter, and large point density variation together with the computational restriction of autonomous navigation. This paper extends an approach of nonparametric projection-pursuit based regression to compensate for the non-uniform and directional nature of data sampled in outdoor environments. The proposed algorithm is employed for extraction of planar structures and clutter grouping. Results are shown for scene abstraction of 3D range data in large urban scenes.
Ranjith Unnikrishnan, Martial Hebert
IROS2
2003 Quality assessment of traversability maps from aerial LIDAR data for an unmanned ground vehicle
abstract
In this paper we address the problem of assessing quantitatively the quality of traversability maps computed from data collected by an airborne laser range finder. Such data is used to plan paths for an unmanned ground vehicle (UGV) prior to the execution of long range traverses. Little attention has been devoted to the problem we address in this paper. We use a unique data set of geodetic control points, real robot navigation data, ground LIDAR (light detection and ranging) data and aerial imagery, collected during a week long demonstration to support our work.
Nicolas Vandapel, Raghavendra Donamukkala, Martial Hebert
IROS3
2003 Discriminative Fields for Modeling Spatial Dependencies in Natural Images
abstract
In this paper we present Discriminative Random Fields (DRF), a discrim- inative framework for the classification of natural image regions by incor- porating neighborhood spatial dependencies in the labels as well as the observed data. The proposed model exploits local discriminative models and allows to relax the assumption of conditional independence of the observed data given the labels, commonly used in the Markov Random Field (MRF) framework. The parameters of the DRF model are learned using penalized maximum pseudo-likelihood method. Furthermore, the form of the DRF model allows the MAP inference for binary classifica- tion problems using the graph min-cut algorithms. The performance of the model was verified on the synthetic as well as the real-world images. The DRF model outperforms the MRF model in the experiments.
Sanjiv Kumar, Martial Hebert
NIPS2
2003 Fully automatic registration of multiple 3D data sets
Daniel F. Huber, Martial Hebert
Image Vis. Comput.2
2003 An observation-constrained generative approach for probabilistic classification of image regions
Sanjiv Kumar, Alexander C. Loui, Martial Hebert
Image Vis. Comput.3
2002 Object Recognition by a Cascade of Edge Probes
abstract
We frame the problem of object recognition from edge cues in terms of determining whether individual edge pixels belong to the target object or to clutter, based on the configuration of edges in their vicinity. A classifier solves this problem by computing sparse, localized edge features at image locations determined at training time. In order to save computation and solve the aperture problem, we apply a cascade of these classifiers to the image, each of which computes edge features over larger image regions than its predecessors. Experiments apply this approach to the recognition of real objects with holes and wiry components in cluttered scenes under arbitrary out-of-image-plane rotation.
Owen T. Carmichael, Martial Hebert
BMVC2
2002 Training Object Detection Models with Weakly Labeled Data
abstract
Appearance based object detection systems utilizing statistical models to capture real world variations in appearance have been shown to exhibit good detection performance. The parameters of these statistical models are typically learned automatically from labeled training images. This process can be difficult in that a large number of labeled training examples may be needed to accurately model appearance variation. In this work we describe a method whereby a training set consisting of a small number of fully labeled training examples augmented with a set of weakly labeled examples can be used to train a detector which exhibits performance better than that which can be obtained with a reduced set of fully labeled training examples alone.
Charles R. Rosenberg, Martial Hebert
BMVC2
2002 On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert
ECCV (3)6
2002 Combining Simple Discriminators for Object Discrimination
Shyjan Mahamud, Martial Hebert, John D. Lafferty
ECCV (3)2
2002 Robust Tracking and Structure from Motion with Sample Based Uncertainty Representation
abstract
Geometric reconstruction of the environment from images is critical in autonomous mapping and robot navigation. Geometric reconstruction involves feature tracking, i.e., locating corresponding image features in consecutive images, and structure from motion (SFM), i.e., recovering the 3D structure of the environment from a set of correspondences between images. Although algorithms for feature tracking and structure from motion are well-established, their use in practical mobile robot applications is still difficult because of occluded features, non-smooth motion between frames, and ambiguous patterns in images. We show how a sampling-based representation can be used in place of the traditional Gaussian representation of uncertainty. We show how sampling can be used for both feature tracking and SFM and we show how they are combined in this framework. The approach is exercised in the context of a mobile robot navigating through an outdoor environment with an omnidirectional camera.
Martial Hebert
ICRA2
2002 Toward Practical Cooperative Stereo for Robotic Colonies
abstract
In this paper we describe an approach towards cooperative stereo. The key problems, significant different views, different scale and occlusion, will be addressed in the context of a distributed robotic system. An algorithmic approach is described towards solving these problems and results on real scenes are presented.
Bart C. Nabbe, Martial Hebert
ICRA2
2001 Provably-Convergent Iterative Methods for Projective Structure from Motion
abstract
The estimation of the projective structure of a scene from image correspondences can be formulated as the minimization of the mean-squared distance between predicted and observed image points with respect to the projection matrices, the scene point positions, and their depths. Since these unknowns are not independent, constraints must be chosen to ensure that the optimization process. is well posed. This paper examines three plausible choices, and shows that the first one leads to the Sturm-Triggs projective factorization algorithm, while the other two lead to new provably-convergent approaches. Experiments with synthetic and real data are used to compare the proposed techniques to the Sturm-Triggs algorithm and bundle adjustment.
Shyjan Mahamud, Martial Hebert, Yasuhiro Omori, Jean Ponce
CVPR (1)2
2001 Object Recognition using Boosted Discriminants
abstract
We approach the task of object discrimination as that of learning efficient "codes" for each object class in terms of responses to a set of chosen discriminants. We formulate this approach in an energy minimization framework. The "code" is built incrementally by successively constructing discriminants that focus on pairs of training images of objects that are currently hard to classify. The particular discriminants that we use partition the set of objects of interest into two well-separated groups. We find the optimal discriminant as well as partition by formulating an objective criteria that measures the well-separateness of the partition. We derive an iterative solution that alternates between the solutions for two generalized eigenproblems, one for the discriminant parameters and the other for the indicator variables denoting the partition. We show how the optimization can easily be biased to focus on hard to classify pairs, which enables us to choose new discriminants one by one in a sequential manner We validate our approach on a challenging face discrimination task using parts as features and show that it compares favorably with the performance of an eigenspace method.
Shyjan Mahamud, Martial Hebert, Jianbo Shi
CVPR (1)2
2001 Color Constancy Using KL-Divergence
Charles R. Rosenberg, Martial Hebert, Sebastian Thrun
ICCV2
2000 Iterative Projective Reconstruction from Multiple Views
abstract
We propose an iterative method for the recovery of the projective structure and motion from multiple images. It has been recently noted that by scaling the measurement matrix by the true projective depths, recovery of the structure and motion is possible by factorization. The reliable determination of the projective depths is crucial to the success of this approach. The previous approach recovers these projective depths using pairwise constraints among images. We first discuss a few important drawbacks with this approach. We then propose an iterative method where we simultaneously recover both the projective depths as well as the structure and motion that avoids some of these drawbacks by utilizing all of the available data uniformly. The new approach makes use of a subspace constraint on the projections of a 3D point onto an arbitrary number of images. The projective depths are readily determined by solving a generalized eigenvalue problem derived from the subspace constraint. We also formulate a dual subspace constraint on all the points in a given image, which can be used for verifying the projective geometry of a scene or object that was modeled. We prove the monotonic convergence of the iterative scheme to a local maximum. We show the robustness of the approach on both synthetic and real data despite large perspective distortions and varying initializations.
Shyjan Mahamud, Martial Hebert
CVPR2
2000 Invariant Filtering for Simultaneous Localization and Mapping
abstract
This paper presents an algorithm for simultaneous localization and map building for a mobile robot moving in an unknown environment. The robot can measure only the bearings to identifiable targets and its own relative motion. The approach is to recursively estimate features of the environment which are invariant to the robot pose in order to decouple the pose error from the map error. The highly nonlinear nature of this problem requires more explicit reasoning about the spatial relationships between landmarks and between the robot and landmarks than those used in previous methods.
Matthew C. Deans, Martial Hebert
ICRA2
2000 Active and Passive Range Sensing for Robotics
abstract
In this paper, we present a brief survey of the technologies currently available for range sensing, of their use in robotics applications, and of emerging technologies for future systems. The paper is organized by type of sensing: laser range finders, triangulation range finders, and passive stereo. A separate section focuses on current development in the area of nonscanning sensors, a critical area to achieve range sensing performance comparable to that of conventional cameras. The presentation of the different technologies is based on many recent examples from robotics research.
Martial Hebert
ICRA1
2000 3-D Map Reconstruction from Range Data
abstract
We present techniques for building models of complex environments from range data gathered at multiple viewpoints. The challenges in this problem are: the matching of unregistered views without prior knowledge of pose, the use of very large data sets, and the manipulation of data sets of different resolutions and from different sensors. Our approach is unique in that no prior knowledge of the relative viewpoints is needed in order to register the data. We show results in building maps of interior environment from range finding data, building large terrain maps from ground-based and from aerial data, and from an operational for mapping from stereo data for hazardous environment characterization. The paper summarizes the major results obtained so far in this area.
Daniel F. Huber, Owen T. Carmichael, Martial Hebert
ICRA3
1999 Harmonic Maps and Their Applications in Surface Matching
abstract
The surface-matching problem is investigated in this paper using a mathematical tool called harmonic maps. The theory of harmonic maps studies the mapping between different metric manifolds from the energy-minimization point of view. With the application of harmonic maps, a surface representation called harmonic shape images is generated to represent and match 3D freeform surfaces. The basic idea of harmonic shape images is to map a 3D surface patch with disc topology to a 2D domain and encode the shape information of the surface patch into the 2D image. This simplifies the surface-matching problem to a 2D image-matching problem. Due to the application of harmonic maps in generating harmonic shape images, harmonic shape images have the following advantages: they have sound mathematical background; they preserve both the shape and continuity of the underlying surfaces; and they are robust to occlusion and independent of any specific surface sampling scheme. The performance of surface matching using harmonic maps is evaluated using real data. Preliminary results are presented in the paper.
Dongmei Zhang 0001, Martial Hebert
CVPR2
1999 Efficient Recovery of Low-Dimensional Structure from High-Dimensional Data
abstract
Many modeling tasks in computer vision, e.g. structure from motion, shape/reflectance from shading, filter synthesis have a low-dimensional intrinsic structure even though the dimension of the input data can be relatively large. We propose a simple but surprisingly effective iterative randomized algorithm that drastically cuts down the time required for recovering the intrinsic structure. The computational cost depends only on the intrinsic dimension of the structure of the task. It is based on the recently proposed Cascade Basis Reduction (CBR) algorithm that was developed in the context of steerable filters. A key feature of our algorithm compared with CBR is that an arbitrary a priori basis for the task is not required. This allows us to extend the applicability of the algorithm to tasks beyond steerable filters such as structure from motion. We prove the convergence for the new algorithm. In practice the new algorithm is much faster than CBR for the same modeling error. We demonstrate this speed-up for the construction of a steerable basis for Gabor filters. We also demonstrate the generality of the new algorithm by applying it to to an example from structure from motion without missing data.
Shyjan Mahamud, Martial Hebert
ICCV2
1999 3-D Cueing: A Data Filter for Object Recognition
abstract
Presents a method for quickly filtering range data points to make object recognition in large 3D data sets feasible. The general approach, called "3D cueing", uses shape signatures from object models as the basis for a fast, probabilistic classification system which rates scene points in terms of their likelihood of belonging to a model. This algorithm which could be used as a front-end for any traditional 3D matching technique, is demonstrated using several models and cluttered scenes in which the model occupies between 1% and 50% of the data points.
Owen T. Carmichael, Martial Hebert
ICRA2
1999 A new approach to 3-D terrain mapping
abstract
We discuss the problem of building larger high-resolution three-dimensional representations of unstructured terrain using terrestrial range sensors, which operate at the scale of meters to hundreds of meters. Issues specific to this sensing modality include widely varying resolution, absence of reliably detectable features, and very large data sets. We have developed a map building algorithm that registers and integrates sequences of range images, and we demonstrate its capabilities by building large terrain maps (260/spl times/166 meters) using ground-based and low-altitude terrestrial range sensors.
Daniel F. Huber, Martial Hebert
IROS2
1999 Using Spin Images for Efficient Object Recognition in Cluttered 3D Scenes
abstract
We present a 3D shape-based object recognition system for simultaneous recognition of multiple objects in scenes containing clutter and occlusion. Recognition is based on matching surfaces by matching points using the spin image representation. The spin image is a data level shape descriptor that is used to match surfaces represented as surface meshes. We present a compression scheme for spin images that results in efficient multiple object recognition which we verify with results showing the simultaneous recognition of multiple objects from a library of 20 models. Furthermore, we demonstrate the robust performance of recognition in the presence of clutter and occlusion through analysis of recognition trials on 100 scenes.
Andrew E. Johnson 0002, Martial Hebert
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Efficient Multiple Model Recognition in Cluttered 3-D Scenes
abstract
We present a 3-D shape-based object recognition system for simultaneous recognition of multiple objects in scenes containing clutter and occlusion. Recognition is based on matching surfaces by matching points using the spin-image representation. The spin-image is a data level shape descriptor that is used to match surfaces represented as surface meshes. We present a compression scheme for spin-images that results in efficient multiple object recognition which we verify with results showing the simultaneous recognition of multiple objects from a library of 20 models. Furthermore, we demonstrate the robust performance of recognition in the presence of clutter and occlusion through analysis of recognition trials on 100 scenes.
Andrew E. Johnson 0002, Martial Hebert
CVPR2
1998 Experiments in Autonomous Driving with Concurrent Goals and Multiple Vehicles
abstract
In this paper we report on experiments with a system for autonomously driving two vehicles based on complex mission specifications. We show that the system is able to plan local paths in obstacle fields based on sensor data, to plan and update global paths to goals based on frequent obstacle map updates, and to modify mission execution, e.g., the ordering of the goals, based on the updated paths to the goals. Two recently developed sensors are used for obstacle detection: a high-speed laser rangefinder and a video-rate stereo system. An updated version of a dynamic path planner D* is used for online computation of routes. A new mission planning and execution monitoring tool, GRAMMPS, is used for managing the allocation and ordering of goals between vehicles. We report on experiments conducted in an outdoor test site with two HMMWVs. Implementation details and performance analysis, including failure modes, are described based on a series of twelve experiments, each over 1/2 km distance with up to nine goals. This system is the first multivehicle and multigoal system to be demonstrated in real, natural environments with this degree of generality. The work reported here includes a number of results not previously published, including the use of a real-time stereo machine, a high performance laser rangefinder and the GRAMMPS planning system.
Barry Brumitt, Martial Hebert
ICRA2
1998 Active Laser Radar for High Performance Measurements
abstract
Laser scanners, or laser radars (ladar), have been used for a number of years for mobile robot navigation and inspection tasks. Although previous scanners were sufficient for low speed applications, they often did not have the range or angular resolution necessary for mapping at the long distances. Many also did not provide an ample field of view with high accuracy and high precision. In this paper we will present the development of state-of-the-art, high speed, high accuracy, 3D laser radar technology. This work has been a joint effort between CMU and K2T and Z+F. The scanner mechanism provides an unobstructed 360/spl deg/ horizontal field of view, and a 70/spl deg/ vertical field of view. Resolution of the scanner is variable with a maximum resolution of approximately 0.06 degrees per pixel in both azimuth and elevation. The laser is amplitude-modulated, continuous-wave with an ambiguity interval of 52 m, a range resolution of 1.6 mm, and a maximum pixel rate of 625 kHz. This paper will focus on the design and performance of the laser radar and will discuss several potential applications for the technology. It reports on performance data of the system including noise, drift over time, precision, and accuracy with measurements. Influences of ambient light, surface material of the target and ambient temperature for range accuracy are discussed. Example data of applications will be shown and improvements will also be discussed.
John A. Hancock, Dirk Langer, Martial Hebert, Ryan Sullivan, Darin Ingimarson, Eric Hoffmann, Markus Mettenleiter, Christoph Fröhlich
ICRA3
1998 Unconstrained registration of large 3D point sets for complex model building
abstract
We present a method for building models of complex environments from range data gathered at multiple viewpoints. Our approach is unique in that no prior knowledge of the relative positions of the viewpoints is needed in order to register data from them. Furthermore, we present a technique for specification and utilization of so-called "common-sense" constraints on the transformations between views to improve the accuracy and speed of the registration process. Results are shown from our effort to map a 60 m by 20 m multiple-room storage area containing a cluttered array of objects.
Owen T. Carmichael, Martial Hebert
IROS2
1998 Omni-directional visual servoing for human-robot interaction
abstract
We describe a visual servoing system developed as a human-robot interface to drive a mobile robot toward any chosen target. An omni-directional camera is used to get the 360 degree of field of view, and an efficient tracking technique is developed to track the target. The use of the omni-directional geometry eliminates, many of the problems common in visual tracking and makes the use of visual servoing a practical alternative for robot-human interaction. The experiments demonstrate that it is an effective and robust way to guide a robot. In particular the experiments show robustness of the tracker to loss of template, vehicle motion, and change in scale and orientation.
Martial Hebert
IROS2
1998 Laser intensity-based obstacle detection
abstract
We present a novel method for obstacle detection for automated highway environments. Laser range scanners have frequently been used for obstacle detection for mobile robots. Although most laser scanners provide intensity information in addition to range, laser intensity has been ignored by most researchers. We show that laser intensity, on its own, is sufficient (and better) for detecting obstacles at long ranges in mild terrain such as an automated highway.
John A. Hancock, Martial Hebert, Charles E. Thorpe
IROS2
1998 Control of Polygonal Mesh Resolution for 3-D Computer Vision
Andrew E. Johnson 0002, Martial Hebert
Graph. Model. Image Process.2
1998 Surface matching for object recognition in complex three-dimensional scenes
Andrew E. Johnson 0002, Martial Hebert
Image Vis. Comput.2
1997 Recognizing Objects by Matching Oriented Points
abstract
We present an approach to recognition of complex objects in cluttered 3-D scenes that does not require feature extraction or segmentation. Our object representation comprises descriptive images associated with each oriented point on the surface of an object. Using a single point basis constructed from an oriented point, the position of other points on the surface of the object can be described by two parameters. The accumulation of these parameters for many points on the surface of the object results in an image at each oriented point. These images, localized descriptions of the global shape of the object, are invariant to rigid transformations. Through correlation of images, point correspondences between a model and scene data are established and then grouped using geometric consistency. The effectiveness of our algorithm is demonstrated with results showing recognition of complex objects in cluttered scenes with occlusion.
Andrew E. Johnson 0002, Martial Hebert
CVPR2
1997 Multi-Scale Classification of 3-D Objects
abstract
We describe an approach to the classification of 3-D objects using a multi-scale representation. This approach starts with a smoothing algorithm for representing objects at different scales. Smoothing is applied in curvature space directly, thus avoiding the usual shrinkage problems and allowing for efficient implementations. A 3-D similarity measure that integrates the representations of the objects at multiple scales is introduced. Given a library of models, objects that are similar based on this multi-scale measure are grouped together into classes. The objects that are in the same class are combined into a single prototype object. Finally, the prototypes are used for hierarchical recognition by first comparing the scene representation to the prototypes and then matching it only to the objects in the most likely class rather than to the entire library of models. Beyond its application to object recognition, this approach provides an attractive implementation of the intuitive notions of scale and approximate similarity for 3-D shapes.
Dongmei Zhang 0001, Martial Hebert
CVPR2
1997 An Integral Approach to Free-Form Object Modeling
abstract
Presents an approach to free-form object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views, so that correspondence can be established among different views. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the techniques of principal component analysis with missing data (PCAMD). Experiments show that our approach is robust against noise and mismatch.
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 On 3D Shape Similarity
abstract
This paper addresses the problem of 3D shape similarity between closed surfaces. A curved or polyhedral 3D object of genus zero is represented by a mesh that has nearly uniform distribution with known connectivity among mesh nodes. A shape similarity metric is defined based on the L/sub 2/ distance between the local curvature distributions over the mesh representations of the two objects. For both convex and concave objects, the shape metric can be computed in time O(n/sup 2/), where n is the number of tessellations of the sphere or the number of meshes which approximate the surface. Experiments show that our method produces good shape similarity measurements.
Harry Shum, Martial Hebert, Katsushi Ikeuchi
CVPR2
1995 Weakly-Calibrated Stereo Perception for Rover Navigation
abstract
Presents a vision system for autonomous navigation based on stereo perception without 3D reconstruction. This approach uses weakly calibrated stereo images, i.e. images for which only the epipolar geometry is known. The vision system first rectifies the images, matches selected points between the two images, and then computes the relative elevation of the points relative to a reference plane as well as the images of their projections on this plane. We have integrated this vision module into a complete navigation system. In this system, the relative elevation is used as a shape indicator in order to compute appropriate steering directions everytime a new stereo pair is processed. We have conducted initial experiments in unstructured, outdoor environments with an wheeled rover.>
Luc Robert, Michel Buffa, Martial Hebert
ICCV3
1995 An Integral Approach to Free-Formed Object Modeling
abstract
Presents a new approach to free-formed object modeling from multiple range images. In most conventional approaches, successive views are registered sequentially. In contrast to the sequential approaches, we propose an integral approach which reconstructs statistically optimal object models by simultaneously aggregating all data from multiple views into a weighted least-squares (WLS) formulation. The integral approach has two components. First, a global resampling algorithm constructs partial representations of the object from individual views so that correspondences can be established among different views. The global resampling algorithm is based on the spherical attribute image (SAI) previously introduced in the context of object representation and recognition. Second, a weighted least-squares algorithm integrates resampled partial representations of multiple views, using the technique of principal component analysis with missing data (PCAMD). Experiments using real range images show that our approach is robust against noise and mismatches, and generates accurate object models.>
Harry Shum, Martial Hebert, Katsushi Ikeuchi, Raj Reddy
ICCV2
1995 Mapping and Positioning for a Prototype Lunar Rover
abstract
In this paper, we describe practical, effective approaches to outdoor mapping and positioning, and present results from systems implemented for a prototype lunar rover. For mapping, we have developed a binocular head and mounted it on a motion-averaging mast. This head provides images to a normalized correlation matcher, that intelligently selects what part of the image to process (saving time), and subsamples the images (again saving time) without subsampling disparities (which would reduce accuracy). The mapping system has operated successfully during long-duration field exercises, processing streams of thousands of images. The positioning system employs encoders, inclinometers, a compass, and a turn-rate sensor to maintain the position and orientation of the rover as it traverses. The system succeeds in the face of significant sensor noise by virtue of sensor modelling, plus extensive filtering and data screening.
Eric Krotkov, Martial Hebert
ICRA2
1995 3-D object modeling and recognition for telerobotic manipulation
abstract
This paper describes a system that semi-automatically builds a virtual world for remote operations by constructing 3-D models of a robot's work environment. With a minimum of human interaction, planar and quadric surface representations of objects typically found in man-made facilities are generated from laser rangefinder data. The surface representations are used to recognize complex models of objects in the scene. These object models are incorporated into a larger world model that can be viewed and analyzed by the operator, accessed by motion planning and robot safeguarding algorithms, and ultimately used by the operator to command the robot through graphical programming and other high level constructs. Limited operator interaction, combined with assumptions about the robots task environment, make the problem of modeling and recognizing objects tractable and yields a solution that can be readily incorporated into many telerobotic control schemes.
Andrew E. Johnson 0002, Patrick Leger, Regis Hoffman, Martial Hebert, James Osborn
IROS (1)4
1995 Experience with rover navigation for lunar-like terrains
abstract
Reliable navigation is critical for a lunar rover, both for autonomous traverses and safeguarded remote teleoperation. This paper describes an implemented system that has autonomously driven a prototype wheeled lunar rover over a kilometer in natural, outdoor terrain. The navigation system uses stereo terrain maps to perform local obstacle avoidance, and arbitrates steering recommendations from both the user and the rover. The paper describes the system architecture, each of the major components, and the experimental results to date.
Reid G. Simmons, Eric Krotkov, Lonnie Chrisman, Fábio G. Cozman, Richard Goodwin, Martial Hebert, Lalitesh Katragadda, Sven Koenig, Gita Krishnaswamy, Yoshikazu Shinoda, William Whittaker, Paul R. Klarer
IROS (1)6
1995 A complete navigation system for goal acquisition in unknown environments
abstract
Most autonomous outdoor navigation systems tested on actual robots have centered on local navigation tasks such as avoiding obstacles or following roads. Global navigation has been limited to simple wandering, path tracking, straight-line goal seeking behaviors, or executing a sequence of scripted local behaviors. These capabilities are insufficient for unstructured and unknown environments, where replanning may be needed to account for new information discovered in every sensor image. To address these problems, the authors developed a complete system that integrates local and global navigation. The local system uses a scanning laser rangefinder to detect and avoid obstacles. The global system uses an incremental path planning algorithm to optimally replan the global path for each detected obstacle. A control arbiter steers the robot to achieve the proper balance between safety and goal acquisition. This system was tested on a real robot and successfully drove it 1.4 kilometers to find a goal given no a priori map of the environment.
Anthony Stentz, Martial Hebert
IROS (1)2
1995 Building 3-D Models from Unregistered Range Images
Ken Higuchi, Martial Hebert, Katsushi Ikeuchi
CVGIP Graph. Model. Image Process.2
1995 A Spherical Representation for Recognition of Free-Form Surfaces
abstract
Introduces a new surface representation for recognizing curved objects. The authors approach begins by representing an object by a discrete mesh of points built from range data or from a geometric model of the object. The mesh is computed from the data by deforming a standard shaped mesh, for example, an ellipsoid, until it fits the surface of the object. The authors define local regularity constraints that the mesh must satisfy. The authors then define a canonical mapping between the mesh describing the object and a standard spherical mesh. A surface curvature index that is pose-invariant is stored at every node of the mesh. The authors use this object representation for recognition by comparing the spherical model of a reference object with the model extracted from a new observed scene. The authors show how the similarity between reference model and observed data can be evaluated and they show how the pose of the reference object in the observed scene can be easily computed using this representation. The authors present results on real range images which show that this approach to modelling and recognizing 3D objects has three main advantages: (1) it is applicable to complex curved surfaces that cannot be handled by conventional techniques; (2) it reduces the recognition problem to the computation of similarity between spherical distributions; in particular, the recognition algorithm does not require any combinatorial search; and (3) even though it is based on a spherical mapping, the approach can handle occlusions and partial views.>
Martial Hebert, Katsushi Ikeuchi, Hervé Delingette
IEEE Trans. Pattern Anal. Mach. Intell.1
1994 Object representation for object recognition
abstract
This paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.>
Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman
CVPR6
1994 Deriving Orientation Cues from Stereo Images
Luc Robert, Martial Hebert
ECCV (1)2
1994 Pixel-Based Range Processing for Autonomous Driving
abstract
We describe a pixel-based approach to range processing for obstacle detection and autonomous driving as an alternative to the traditional image- or map-based approaches. The pixel-based approach eliminates the delays due to image acquisition and map building and permits the integration of traversability evaluation and path generation into a single module without the latency involved in distributed systems. We describe the algorithm used for updating a local map using individual range pixels, for detecting obstacles on the fly, and for generating steering commands. We illustrate the performance of the algorithm using an implementation of a cross-country driving system with a scanning laser range finder.>
Martial Hebert
ICRA1
1994 Building 3-D Models from Unregistered Range Images
abstract
The authors describe an approach to building a three-dimensional model from a set of range images. The authors' goal is to build models of free-form surfaces obtained from arbitrary viewing directions, with no initial estimate of the relative viewing directions. The approach is based on building discrete meshes representing the surfaces observed in each of the range images, to map each of the meshes to a spherical image, and to compute the transformations between the views by matching the spherical images. The meshes are built using an iterative fitting algorithm previously developed; the spherical images are built by matching the nodes of the surface meshes to the nodes of a reference mesh on the unit sphere and by storing a measure of curvature at every node. The authors describe the algorithms used for building such models from range images and for matching them. The authors give results obtained using range images of complex objects.>
Ken Higuchi, Martial Hebert, Katsushi Ikeuchi
ICRA2
1994 Real-Thme 3-D Pose Estimation Using a High-Speed Range Sensor
abstract
This paper describes a system which can perform full 3-D pose estimation of a single arbitrarily shaped, rigid object at rates up to 10 Hz. A triangular mesh model of the object to be tracked is generated offline using conventional range sensors. Real-time range data of the object is sensed by the CMU high speed VLSI range sensor. Pose estimation is performed by registering the real-time range data to the triangular mesh model using an enhanced implementation of the Iterative Closest Point (ICP) Algorithm introduced by Besl and McKay (1992). The method does not require explicit feature extraction or specification of correspondence. Pose estimation accuracies of the order of 1% of the object size in translation, and 1 degree in rotation have been measured.>
David A. Simon, Martial Hebert, Takeo Kanade
ICRA2
1994 An Integrated System for Autonomous Off-Road Navigation
abstract
In this paper, we report on experiments with a core system for autonomous navigation in outdoor natural terrain. The system consists of three parts: a perception module which processes range images to identify untraversable regions of the terrain, a local map management module which maintains a representation of the environment in the vicinity of the vehicle, and a planning module which issues commands to the vehicle controller. Our approach uses reactive planning for generating commands to drive the vehicle along with "early traversability evaluation," in which the perception module decides which parts of the terrain are traversable as soon as a new image is taken. We argue that our approach leads to a robust and efficient navigation system. We illustrate our approach by an experiment in which a vehicle travelled autonomously for one kilometer through unmapped cross-country terrain.>
Dirk Langer, Julio Rosenblatt, Martial Hebert
ICRA3
1994 A behavior-based system for off-road navigation
abstract
We describe a core system for autonomous navigation in outdoor natural terrain. The system consists of three parts: a perception module that processes range images to identify untraversable regions of the terrain, a local map management module that maintains a representation of the environment in the vicinity of the vehicle, and a planning module that issues commands to the vehicle controller. We illustrate our approach by an experiment in which a vehicle travelled autonomously for one kilometer through unmapped cross-country terrain.>
Dirk Langer, Julio Rosenblatt, Martial Hebert
IEEE Trans. Robotics Autom.3
1993 A spherical representation for the recognition of curved objects
abstract
The authors introduce a surface representation for recognizing curved objects. The approach begins by representing an object by a discrete mesh of points built from range data or from a geometric model of the object. The mesh is computed from the data by deforming a standard shaped mesh, for example, an ellipsoid, until it fits the surface of the object. Local regularity constraints that the mesh must satisfy are defined. A canonical mapping is then defined between the mesh describing the object and a standard spherical mesh. A surface curvature index which is pose-invariant is stored at every node of the mesh. This object representation is used for recognition by comparing the spherical model of a reference object with the model extracted from a new observed scene. It is shown that the similarity between reference model and observed data can be evaluated, and it is also demonstrated that the pose of the reference object in the observed scene can be easily computed using this representation.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
ICCV2
1992 3-D landmark recognition from range images
abstract
Progress in building and recognizing models of objects for an autonomous vehicle for on-road and cross-country navigation is reported. The object models are stored in a map and are used as landmarks for estimating vehicle position. The landmarks can be used as intermediate control points at which the vehicle must take some prescribed action in the case of a complex mission. Robust object tracking using sequences of range images and building and updating 3-D object representations is presented. Tracking uses object prediction from one image to the next to accurately compute object locations. Object representations are built by merging sets of points from individual images into a single set in an object-centered coordinate frame. The sparse set of points is then segmented into shapes yielding compact and general object representations. An algorithm for landmark identification in range images is introduced in the context of map-based navigation.>
Martial Hebert
CVPR1
1992 Task Oriented Vision
Katsushi Ikeuchi, Martial Hebert
IROS2
1992 Shape representation and image segmentation using deformable surfaces
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
Image Vis. Comput.2
1992 3D measurements from imaging laser radars: how good are they?
Martial Hebert, Eric Krotkov
Image Vis. Comput.1
1991 Shape representation and image segmentation using deformable surfaces
abstract
A technique for constructing shape representation from images using free-form deformable surfaces is presented. The authors model an object as a closed surface that is deformed subject to attractive fields generated by input data points and features. Features affect the global shape of the surface, while data points control its local shape. This approach is used to segment objects even in cluttered or unstructured environments. The algorithm is general in that it makes few assumptions on the type of features, the nature of the data, and the type of objects. Results for a wide range of applications are presented: reconstruction of smooth isolated objects such as human faces, reconstruction of structured objects such as polyhedra, and segmentation of complex scenes with mutually occluding objects. The algorithm has been successfully tested using data from different sensors including grey-coding range finders and video cameras, using one or several images.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
CVPR2
1991 A three-finger gripper for manipulation in unstructured environments
abstract
A gripper is described for manipulation in natural, unstructured environments. The specific manipulation task is to pick up surface material such as pebbles or small rocks in a natural terrain. The application is to give autonomous sampling capabilities to an autonomous vehicle for planetary exploration. The authors describe the task analysis process that led to the selection of a configuration with three soft fingers. They carry out a complete analysis of the stability of a grasp for this gripper including an analysis of the deformation of the fingers at the points of contact. The implementation of a grasp selection algorithm is described, and results on three-dimensional representations of objects computed from range data are presented.>
C. Francois, Katsushi Ikeuchi, Martial Hebert
ICRA3
1991 Building qualitative elevation maps from side scan sonar data for autonomous underwater navigation
abstract
Deriving a terrain model from sensor data is an important task for the autonomous navigation of a mobile robot. An approach is presented for autonomous underwater vehicles using a side scan sonar system. Some general aspects of the type of data and filtering techniques to improve it are discussed. An estimated bottom contour is derived using a geometric reflection model and information about shadows and highlights. Several techniques of surface reconstruction and their limitations are presented. A method is presented for feature extraction which is important for future data matching/fusion procedures.>
Dirk Langer, Martial Hebert
ICRA2
1991 Trajectory generation with curvature constraint based on energy minimization
abstract
The trajectory generation problem for mobile robots consists in providing a set of trajectories that are 'smooth' and meet certain boundary conditions. The authors present a method to generate curvature continuous trajectories for which the curvature profile is a polynomial function of arc length. An algorithm based on the deformation of a curve by energy minimization allows one to solve general geometric constraints which was not possible by previous methods. Furthermore, it is able to take into account the limitation of radius of curvature of the robot by controlling the extrema of curvature along the path.>
Hervé Delingette, Martial Hebert, Katsushi Ikeuchi
IROS2
1991 3-D measurements from imaging laser radars: how good are they?
abstract
The authors analyze a class of imaging range finders-amplitude-modulated continuous-wave laser radars-in the context of computer vision and robotics. The analysis develops measurement models from the fundamental principles of laser radar operation, and identifies the nature and cause of key problems that plague measurements from this class of sensors. They classify the problems as fundamental (e.g. related to the signal-to-noise ratio), as architectural (e.g. limited by encoding distance by angles (0.2 pi )), and as artifacts of particular hardware implementations (e.g. insufficient temperature compensation). Experimental results from two different scanning laser range finders designed for autonomous navigation illustrate and support the analysis.>
Martial Hebert, Eric Krotkov
IROS1
1989 Building and navigating maps of road scenes using an active sensor
abstract
The author presents algorithms for building maps of road scenes using an active range and reflectance sensor and for using the maps to traverse a portion of the world already explored. He describes some advantages of an active sensor, namely, that it is independent of the illumination conditions, does not require complex calibration in order to transform observed features to the vehicle's reflectance frame, and provides 3-D terrain models as well as road models. Using this map built from sensor data facilitates navigation in two respects: the vehicle may navigate faster, since less perception processing is necessary, and the vehicle may follow a more accurate path, since the navigation system does not rely entirely on inaccurate visual data. The author presents a complete system that includes road-following, map-building, and map-based navigation using the ERIM laser rangefinder. Experimental results are presented.>
Martial Hebert
ICRA1
1989 Terrain mapping for a roving planetary explorer
abstract
The authors are prototyping a legged vehicle, the Ambler, for an exploratory mission on another planet, conceivably Mars, where it is to traverse uncharted areas and collect material samples. They describe how the rover can construct from range imagery a geometric terrain representation, i.e., elevation map that includes uncertainty, unknown areas, and local features. First, they present an algorithm for constructing an elevation map from a single range image. By virtue of working in spherical-polar space, the algorithm is independent of the desired map resolution and the orientation of the sensor, unlike algorithms that work in Cartesian space. Secondly, the authors present a two-stage matching technique (feature matching followed by iconic matching) that identifies the transformation T corresponding to the vehicle displacement between two viewing positions. Thirdly, to support legged locomotion over rough terrain, they describe methods for evaluating regions of the constructed elevation maps as footholds.>
Martial Hebert, Claude Caillas, Eric Krotkov, In-So Kweon, Takeo Kanade
ICRA1
1987 An architecture and two cases in range-based modeling and planning
abstract
This paper presents a framework for autonomous robots that reason from range data. We argue for spatial reasoning as a basic cognition mode for robots operating in unpredictable work environments and present a three-level architecture for modeling and planning from range data. Two implementations, robotic excavation with sonar ranging and mine navigation with laser ranging, illustrate the techniques and provide two experiences to evaluate the architecture.
William Whittaker, George M. Turkiyyah, Martial Hebert
ICRA3
1986 Outdoor scene analysis using range data
abstract
This paper describes techniques for outdoor scene analysis using range data. The purpose of these techniques is to build a 3-D representation of the environment of an mobile robot equipped with a range sensor. Algorithms are presented for scene segmentation, object detection, map building, and object recognition. We present results obtained in an outdoor navigation environment in which a laser range finder is mounted on a vehicle. These results have been successfully applied to the problem of path planning through obstacles.
Martial Hebert
ICRA1
1984 Polyhedral approximation of 3-D objects without holes
Olivier D. Faugeras, Martial Hebert, Philippe Mussi, Jean-Daniel Boissonnat
Comput. Vis. Graph. Image Process.2
1983 A 3-D Recognition and Positioning Algorithm Using Geometrical Matching Between Primitive Surfaces
Olivier D. Faugeras, Martial Hebert
IJCAI2