VLDB 2026 Research / reviewers in the wild / expert
Stanley T. Birchfield
dblp:b/StanBirchfield · also Stan Birchfield
· DBLP profile ↗
88ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0001-7366-2441ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 9 first-author · 28 since 2021Systems, architecture and hardware · 36 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-GENERALIST: Vision-Language-Action Models for Crafting 3D WorldsabstractCreating 3D graphics content for immersive and interactive worlds remains labor-intensive, limiting our ability to create large-scale synthetic data for training foundation models. Recent methods aim to alleviate this, but they often focus on a single aspect (e.g., layout) and do not improve generation quality by simply scaling computational resources. We recast 3D environment generation as a sequential decision-making problem, using Vision-Language Models (VLMs) as policies that output actions to jointly craft a 3D environment's layout, materials, lighting, and assets. Our framework, 3D-Generalist, trains VLMs to generate more prompt-aligned 3D environments via self-improvement fine-tuning. We demonstrate the effectiveness of 3D-Generalist and our training strategy in generating simulation-ready 3D environments. We also demonstrate its quality and scalability for synthetic data generation by pretraining a vision foundation model on the generated data. After fine-tuning on downstream tasks, we show that it surpasses models pre-trained on meticulously human-crafted synthetic data and approaches results achieved when training with orders of magnitude larger real data. Fan-Yun Sun, Shengguang Wu, Christian Jacobsen, Thomas Yim, Haoming Zou, Alexander Zook, Shangru Li, Yu-Hsin Chou, Ethem Can, Xunlei Wu, Clemens Eppner, Valts Blukis, Jonathan Tremblay, Jiajun Wu 0001, Stanley T. Birchfield, Nick Haber |
3DV | 15 |
| 2025 | RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsabstractSpatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant challenges in spatial reasoning tasks, as their training data are based on general-purpose image datasets that often lack sophisticated spatial understanding. For example, datasets frequently do not capture reference frame comprehension, yet effective spatial reasoning requires understanding whether to reason from ego-, world-, or object-centric perspectives. To address this issue, we introduce RoboSpatial, a large-scale dataset for spatial understanding in robotics. It consists of real indoor and tabletop scenes, captured as 3D scans and egocentric images, and annotated with rich spatial information relevant to robotics. The dataset includes 1M images, 5k 3D scans, and 3M annotated spatial relationships, and the pairing of 2D egocentric images with 3D scans makes it both 2D- and 3D- ready. Our experiments show that models trained with RoboSpatial outperform baselines on downstream tasks such as spatial affordance prediction, spatial relationship prediction, and robotics manipulation. Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su 0001, Stanley T. Birchfield |
CVPR | 6 |
| 2025 | FoundationStereo: Zero-Shot Stereo MatchingabstractTremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization — a hallmark of foundation models in other computer vision tasks — remains challenging for stereo matching. We introduce FoundationStereo, a foundation model for stereo depth estimation designed to achieve strong zero-shot generalization. To this end, we first construct a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism, followed by an automatic self-curation pipeline to remove ambiguous samples. We then design a number of network architecture components to enhance scalability, including a side-tuning feature backbone that adapts rich monocular priors from vision foundation models to mitigate the sim-to-real gap, and long-range context reasoning for effective cost volume filtering. Together, these components lead to strong robustness and accuracy across domains, establishing a new standard in zero-shot stereo depth estimation. Project page: https://nvlabs.github.io/FoundationStereo/ Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz, Orazio Gallo, Stanley T. Birchfield |
CVPR | 6 |
| 2025 | SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric ManipulationabstractWe introduce SPOT, an object-centric imitation learning framework. The key idea is to capture each task by an object-centric representation, specifically the SE(3) object pose trajectory relative to the target. This approach decouples embodiment actions from sensory inputs, facilitating learning from various demonstration types, including both action-based and action-less human hand demonstrations, as well as crossembodiment generalization. Additionally, object pose trajectories inherently capture planning constraints from demonstrations without the need for manually-crafted rules. To guide the robot in executing the task, the object trajectory is used to condition a diffusion policy. We systematically evaluate our method on simulation and real-world tasks. In real-world evaluation, using only eight demonstrations shot on an iPhone, our approach completed all tasks while fully complying with task constraints. Project page: https://nvlabs.github.io/object_centric_diffusion Cheng-Chun Hsu, Bowen Wen, Jie Xu 0028, Yashraj S. Narang, Xiaolong Wang 0004, Yuke Zhu, Joydeep Biswas, Stanley T. Birchfield |
ICRA | 8 |
| 2025 | RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completionabstract3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D consistency, are computationally expensive, and struggle to capture sharp object boundaries.
Our work (RaySt3R) addresses these limitations by recasting 3D shape completion as a novel view synthesis problem.
Specifically, given a single RGB-D image,
and a novel viewpoint (encoded as a collection of query rays),
we train a feedforward transformer to predict depth maps, object masks, and per-pixel confidence scores for those query rays.
RaySt3R fuses these predictions across multiple query views
to reconstruct complete 3D shapes.
We evaluate RaySt3R on synthetic and real-world datasets, and observe it achieves state-of-the-art performance,
outperforming the baselines on all datasets by up to 44% in 3D chamfer distance. Bardienus Pieter Duisterhof, Jan Oberst, Bowen Wen, Stanley T. Birchfield, Deva Ramanan, Jeffrey Ichnowski |
NeurIPS | 4 |
| 2025 | Audio-Visual Segmentation with Semantics
Jinxing Zhou, Xuyang Shen, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong |
Int. J. Comput. Vis. | 7 |
| 2024 | Partial-View Object View Synthesis via Filtering InversionabstractWe propose Filtering Inversion (FINV), a learning framework and optimization process that predicts a renderable 3D object representation from one or few partial views. FINV addresses the challenge of synthesizing novel views of objects from partial observations, spanning cases where the object is not entirely in view, is partially occluded, or is only observed from similar views. To achieve this, FINV learns shape priors by training a 3D generative model. At inference, given one or more views of a novel real-world object, FINV first finds a set of latent codes for the object by inverting the generative model from multiple initial seeds. Maintaining the set of latent codes, FINV filters and resamples them after receiving each new observation, akin to particle filtering. The generator is then finetuned for each latent code on the available views in order to adapt to novel objects. We show that FINV successfully synthesizes novel views of real-world objects (e.g., chairs, tables, and cars), even if the generative prior is trained only on synthetic objects. The ability to address the sim-to-real problem allows FINV to be used for object categories without real-world datasets. FINV achieves state-of-the-art performance on multiple real-world datasets, recovers object shape and texture from partial and sparse views, is robust to occlusion, and is able to incrementally improves its representation with more observations. Fan-Yun Sun, Jonathan Tremblay, Valts Blukis, Danfei Xu, Boris Ivanovic, Péter Karkus, Stanley T. Birchfield, Dieter Fox, Yunzhu Li, Jiajun Wu 0001, Marco Pavone 0001, Nick Haber |
3DV | 8 |
| 2024 | NeRFDeformer: NeRF Transformation from a Single View via 3D Scene FlowsabstractWe present a method for automatically modifying a NeRF representation based on a single observation of a non-rigid transformed version of the original scene. Our method defines the transformation as a 3D flow, specifically as a weighted linear blending of rigid transformations of 3D anchor points that are defined on the surface of the scene. In order to identify anchor points, we introduce a novel correspondence algorithm that first matches RGB-based pairs, then leverages multi-view information and 3D reprojection to robustly filter false positives in two steps. We also introduce a new dataset for exploring the problem of modifying a NeRF scene through a single observation. Our dataset11https://nerfdeformer.github.io/ contains 113 synthetic scenes leveraging 47 3D assets. We show that our proposed method outperforms NeRF editing methods as well as diffusion-based methods, and we also explore different methods for filtering correspondences. Zhenggang Tang, Zhongzheng Ren, Xiaoming Zhao 0001, Bowen Wen, Jonathan Tremblay, Stanley T. Birchfield, Alexander G. Schwing |
CVPR | 6 |
| 2024 | FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsabstractWe present FoundationPose, a unified foundation model for 6D object pose estimation and tracking, supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without finetuning, as long as its CAD model is given, or a small number of reference images are captured. Thanks to the unified framework, the downstream pose estimation modules are the same in both setups, with a neural implicit representation used for efficient novel view synthesis when no CAD model is available. Strong generalizability is achieved via large-scale synthetic training, aided by a large language model (LLM), a novel transformer-based architecture, and contrastive learning formulation. Extensive evaluation on multiple public datasets involving challenging scenarios and objects indicate our unified approach outperforms existing methods specialized for each task by a large margin. In addition, it even achieves comparable results to instance-level methods despite the reduced assumptions. Project page: https://nvlabs.github.io/FoundationPose/ Bowen Wen, Wei Yang 0019, Jan Kautz, Stanley T. Birchfield |
CVPR | 4 |
| 2024 | Neural Implicit Representation for Building Digital Twins of Unknown Articulated ObjectsabstractWe address the problem of building digital twins of unknown articulated objects from two RGBD scans of the object at different articulation states. We decompose the problem into two stages, each addressing distinct aspects. Our method first reconstructs object-level shape at each state, then recovers the underlying articulation model in-cluding part segmentation and joint articulations that as-sociate the two states. By explicitly modeling point-level correspondences and exploiting cues from images, 3D reconstructions, and kinematics, our method yields more accurate and stable results compared to prior work. It also handles more than one movable part and does not rely on any object shape or structure priors. Project page: https://github.com/NVlabs/DigitalTwinArt Yijia Weng, Bowen Wen, Jonathan Tremblay, Valts Blukis, Dieter Fox, Leonidas J. Guibas, Stanley T. Birchfield |
CVPR | 7 |
| 2023 | TTA-COPE: Test-Time Adaptation for Category-Level Object Pose EstimationabstractTest-time adaptation methods have been gaining attention recently as a practical solution for addressing source-to-target domain gaps by gradually updating the model without requiring labels on the target data. In this paper, we propose a method of test-time adaptation for category-level object pose estimation called TTA-COPE. We design a pose ensemble approach with a self-training loss using pose-aware confidence. Unlike previous unsupervised domain adaptation methods for category-level object pose estimation, our approach processes the test data in a sequential, online manner, and it does not require access to the source domain at runtime. Extensive experimental results demonstrate that the proposed pose ensemble and the self-training loss improve category-level object pose performance during test time under both semi-supervised and unsupervised settings. Taeyeop Lee, Jonathan Tremblay, Valts Blukis, Bowen Wen, Byeong-Uk Lee, Inkyu Shin, Stanley T. Birchfield, In-So Kweon, Kuk-Jin Yoon |
CVPR | 7 |
| 2023 | BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown ObjectsabstractWe present a near real-time (10Hz) method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the object. Our method works for arbi-trary rigid objects, even when visual texture is largely ab-sent. The object is assumed to be segmented in the first frame only. No additional information is required, and no assumption is made about the interaction agent. Key to our method is a Neural Object Field that is learned concurrently with a pose graph optimization process in order to robustly accumulate information into a consistent 3D representation capturing both geometry and appearance. A dynamic pool of posed memory frames is automatically main-tained to facilitate communication between these threads. Our approach handles challenging sequences with large pose changes, partial and full occlusion, untextured surfaces, and specular highlights. We show results on HO3D, YCBInEOAT, and BEHAVE datasets, demonstrating that our method significantly outperforms existing approaches. Project page: https://bundlesdf.github.io/ Bowen Wen, Jonathan Tremblay, Valts Blukis, Stephen Tyree, Thomas Müller 0013, Alex Evans, Dieter Fox, Jan Kautz, Stanley T. Birchfield |
CVPR | 9 |
| 2023 | Affordance Diffusion: Synthesizing Hand-Object InteractionsabstractRecent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting objects into a user-specified region. In contrast, in this work we focus on synthesizing complex interactions (i.e., an articulated hand) with a given object. Given an RGB image of an object, we aim to hallucinate plausible images of a human hand interacting with it. We propose a two-step generative approach: a LayoutNet that samples an articulation-agnostic hand-object-interaction layout, and a ContentNet that synthesizes images of a hand grasping the object given the predicted layout. Both are built on top of a large-scale pretrained diffusion model to make use of its latent representation. Compared to baselines, the proposed method is shown to generalize better to novel objects and perform surprisingly well on out-of-distribution in-the-wild scenes of portable-sized objects. The resulting system allows us to predict descriptive affordance information, such as hand articulation and approaching orientation. Yufei Ye 0001, Abhinav Gupta 0001, Shalini De Mello, Stanley T. Birchfield, Jiaming Song, Shubham Tulsiani, Sifei Liu |
CVPR | 5 |
| 2023 | Parallel Inversion of Neural Radiance Fields for Robust Pose EstimationabstractWe present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual between pixels rendered from a fast NeRF model and pixels in the observed image. We integrate a momentum-based camera extrinsic optimization procedure into Instant Neural Graphics Primitives, a recent exceptionally fast NeRF implementation. By introducing parallel Monte Carlo sampling into the pose estimation task, our method overcomes local minima and improves efficiency in a more extensive search space. We also show the importance of adopting a more robust pixel-based loss function to reduce error. Experiments demonstrate that our method can achieve improved generalization and robustness on both synthetic and real-world benchmarks. Yunzhi Lin, Thomas Müller 0013, Jonathan Tremblay, Bowen Wen, Stephen Tyree, Alex Evans, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 8 |
| 2023 | RGB-Only Reconstruction of Tabletop Scenes for Collision-Free Manipulator ControlabstractWe present a system for collision-free control of a robot manipulator that uses only RGB views of the world. Perceptual input of a tabletop scene is provided by multiple images of an RGB camera (without depth) that is either handheld or mounted on the robot end effector. A NeRF-like process is used to reconstruct the 3D geometry of the scene, from which the Euclidean full signed distance function (ESDF) is computed. A model predictive control algorithm is then used to control the manipulator to reach a desired pose while avoiding obstacles in the ESDF. We show results on a real dataset collected and annotated in our lab. Our results are also available at https://ngp-mpc.github.io/. Zhenggang Tang, Balakumar Sundaralingam, Jonathan Tremblay, Bowen Wen, Stephen Tyree, Charles T. Loop, Alexander G. Schwing, Stanley T. Birchfield |
ICRA | 9 |
| 2023 | HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and ReconstructionsabstractWe present the HANDAL dataset for category-level object pose estimation and affordance prediction. Unlike previous datasets, ours is focused on robotics-ready manipulable objects that are of the proper size and shape for functional grasping by robot manipulators, such as pliers, utensils, and screwdrivers. Our annotation process is streamlined, requiring only a single off-the-shelf camera and semi-automated processing, allowing us to produce high-quality 3D annotations without crowd-sourcing. The dataset consists of 308k annotated image frames from 2.2k videos of 212 real-world objects in 17 categories. We focus on hardware and kitchen tool objects to facilitate research in practical scenarios in which a robot manipulator needs to interact with the environment beyond simple pushing or indiscriminate grasping. We outline the usefulness of our dataset for 6-DoF category-level pose+scale estimation and related tasks. We also provide 3D reconstructed meshes of all objects, and we outline some of the bottlenecks to be addressed for democratizing the collection of datasets like this one. Project website: https://nvlabs.github.io/HANDAL/ Andrew Guo, Bowen Wen, Jianhe Yuan, Jonathan Tremblay, Stephen Tyree, Jeffrey Smith 0002, Stanley T. Birchfield |
IROS | 7 |
| 2023 | Vicinity Vision TransformerabstractVision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches. Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Audio-Visual Segmentation
Jinxing Zhou, Weixuan Sun, Jing Zhang 0052, Stanley T. Birchfield, Dan Guo 0001, Lingpeng Kong, Meng Wang 0001, Yiran Zhong |
ECCV (37) | 6 |
| 2022 | PredictionNet: Real-Time Joint Probabilistic Traffic Prediction for Planning, Control, and SimulationabstractPredicting the future motion of traffic agents is crucial for safe and efficient autonomous driving. To this end, we present PredictionNet, a deep neural network (DNN) that predicts the motion of all surrounding traffic agents together with the ego-vehicle's motion. All predictions are probabilistic and are represented in a simple top-down rasterization that allows an arbitrary number of agents. Conditioned on a multi-layer map with lane information, the network outputs future positions, velocities, and backtrace vectors jointly for all agents including the ego-vehicle in a single pass. Trajectories are then extracted from the output. The network can be used to simulate realistic traffic, and it produces competitive results on popular benchmarks. More importantly, it has been used to successfully control a real-world vehicle for hundreds of kilometers, by combining it with a motion planning/control subsystem. The network runs faster than real-time on an embedded GPU, and the system shows good generalization (across sensory modalities and locations) due to the choice of input representation. Furthermore, we demonstrate that by extending the DNN with reinforcement learning (RL), it can better handle rare or unsafe events like aggressive maneuvers and crashes. Alexey Kamenev, Lirui Wang, Ollin Boer Bohan, Ishwar Kulkarni, Bilal Kartal, Artem Molchanov, Stanley T. Birchfield, David Nistér, Nikolai Smolyanskiy |
ICRA | 7 |
| 2022 | Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty EstimationabstractWe propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predict the bounding cuboid and 6- DoF pose (up to scale). Internally, a deep network predicts distributions over object keypoints (vertices of the bounding cuboid) in image coordinates, after which a novel probabilistic filtering process integrates across estimates before computing the final pose using PnP. Our framework allows the system to take previous uncertainties into consideration when predicting the current frame, resulting in predictions that are more accurate and stable than single frame methods. Extensive experiments show that our method outperforms existing approaches on the challenging Objectron benchmark of annotated object videos. We also demonstrate the usability of our work in an augmented reality setting. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 5 |
| 2022 | Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB ImageabstractPrior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstructured, real-world scenarios. In this work, we propose a single-stage, keypoint-based approach for category-level object pose estimation that operates on unknown object instances within a known category using a single RGB image as input. The proposed network performs 2D object detection, detects 2D keypoints, estimates 6- DoF pose, and regresses relative bounding cuboid dimensions. These quantities are estimated in a sequential fashion, leveraging the recent idea of convGRU for propagating information from easier tasks to those that are more difficult. We favor simplicity in our design choices: generic cuboid vertex coordinates, single-stage network, and monocular RGB input. We conduct extensive experiments on the challenging Objectron benchmark, outperforming state-of-the-art methods on the 3D IoU metric (27.6% higher than the MobilePose single-stage approach and 7.1 % higher than the related two-stage approach). Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
ICRA | 5 |
| 2022 | 6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and BenchmarkabstractWe present a new dataset for 6-DoF pose estimation of known objects, with a focus on robotic manipulation research. We propose a set of toy grocery objects, whose physical instantiations are readily available for purchase and are appropriately sized for robotic grasping and manipulation. We provide 3D scanned textured models of these objects, suitable for generating synthetic training data, as well as RGBD images of the objects in challenging, cluttered scenes exhibiting partial occlusion, extreme lighting variations, multiple instances per image, and a large variety of poses. Using semi-automated RGBD-to-model texture correspondences, the images are annotated with ground truth poses accurate within a few millimeters. We also propose a new pose evaluation metric called ADD-H based on the Hungarian assignment algorithm that is robust to symmetries in object geometry without requiring their explicit enumeration. We share pre-trained pose estimators for all the toy grocery objects, along with their baseline performance on both validation and test sets. We offer this dataset to the community to help connect the efforts of computer vision researchers with the needs of roboticists.11https://github.com/swtyree/hope-dataset Stephen Tyree, Jonathan Tremblay, Thang To, Terry Mosier, Jeffrey Smith 0002, Stanley T. Birchfield |
IROS | 7 |
| 2022 | Displacement-Invariant Cost Computation for Stereo MatchingabstractAbstract Although deep learning-based methods have dominated stereo matching leaderboards by yielding unprecedented disparity accuracy, their inference time is typically slow, i.e., less than 4 FPS for a pair of 540p images. The main reason is that the leading methods employ time-consuming 3D convolutions applied to a 4D feature volume. A common way to speed up the computation is to downsample the feature volume, but this loses high-frequency details. To overcome these challenges, we propose a displacement-invariant cost computation module to compute the matching costs without needing a 4D feature volume. Rather, costs are computed by applying the same 2D convolution network on each disparity-shifted feature map pair independently. Unlike previous 2D convolution-based methods that simply perform context mapping between inputs and disparity maps, our proposed approach learns to match features between the two images. We also propose an entropy-based refinement strategy to refine the computed disparity map, which further improves the speed by avoiding the need to compute a second disparity map on the right image. Extensive experiments on standard datasets (SceneFlow, KITTI, ETH3D, and Middlebury) demonstrate that our method achieves competitive accuracy with much less inference time. On typical image sizes (e.g., $$540\times 960$$ 540 × 960 ), our method processes over 100 FPS on a desktop GPU, making our method suitable for time-critical applications such as autonomous driving. We also show that our approach generalizes well to unseen datasets, outperforming 4D-volumetric methods. We will release the source code to ensure the reproducibility. Yiran Zhong, Charles T. Loop, Wonmin Byeon, Stanley T. Birchfield, Yuchao Dai, Kaihao Zhang, Alexey Kamenev, Thomas M. Breuel, Hongdong Li, Jan Kautz |
Int. J. Comput. Vis. | 4 |
| 2021 | DexYCB: A Benchmark for Capturing Hand Grasping of ObjectsabstractWe introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estimation, and 3D hand pose estimation. Finally, we evaluate a new robotics-relevant task: generating safe robot grasps in human-to-robot object handover.1 Yu-Wei Chao, Wei Yang 0019, Yu Xiang 0001, Pavlo Molchanov 0001, Ankur Handa, Jonathan Tremblay, Yashraj Narang, Karl Van Wyk, Umar Iqbal 0001, Stanley T. Birchfield, Jan Kautz, Dieter Fox |
CVPR | 10 |
| 2021 | Deep Two-View Structure-From-Motion RevisitedabstractTwo-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting a depth map from a single image, both of which are ill-posed problems. In contrast, we propose to revisit the problem of deep two-view SfM by leveraging the well-posedness of the classic pipeline. Our method consists of 1) an optical flow estimation network that predicts dense correspondences between two frames; 2) a normalized pose estimation module that computes relative camera poses from the 2D optical flow correspondences, and 3) a scale-invariant depth estimation network that leverages epipolar geometry to reduce the search space, refine the dense correspondences, and estimate relative depth maps. Extensive experiments show that our method outperforms all state-of-the-art two-view SfM methods by a clear margin on KITTI depth, KITTI VO, MVS, Scenes11, and SUN3D datasets in both relative pose and depth estimation. Yiran Zhong, Yuchao Dai, Stanley T. Birchfield, Kaihao Zhang, Nikolai Smolyanskiy, Hongdong Li |
CVPR | 4 |
| 2021 | Self-Supervised Real-to-Sim Scene GenerationabstractSynthetic data is emerging as a promising solution to the scalability issue of supervised deep learning, especially when real data are difficult to acquire or hard to annotate. Synthetic data generation, however, can itself be prohibitively expensive when domain experts have to manually and painstakingly oversee the process. More-over, neural networks trained on synthetic data often do not perform well on real data because of the domain gap. To solve these challenges, we propose Sim2SG, a self-supervised automatic scene generation technique for matching the distribution of real data. Importantly, Sim2SG does not require supervision from the real-world dataset, thus making it applicable in situations for which such annotations are difficult to obtain. Sim2SG is designed to bridge both the content and appearance gaps, by matching the content of real data, and by matching the features in the source and target domains. We select scene graph (SG) generation as the downstream task, due to the limited availability of labeled datasets. Experiments demonstrate significant improvements over leading baselines in reducing the domain gap both qualitatively and quantitatively, on several synthetic datasets as well as the real-world KITTI dataset. Aayush Prakash, Shoubhik Debnath, Jean-Francois Lafleche, Eric Cameracci, Gavriel State, Stanley T. Birchfield, Marc T. Law |
ICCV | 6 |
| 2021 | Fast Uncertainty Quantification for Deep Object Pose EstimationabstractDeep learning-based object pose estimators are often unreliable and overconfident especially when the input image is outside the training domain, for instance, with sim2real transfer. Efficient and robust uncertainty quantification (UQ) in pose estimators is critically needed in many robotic tasks. In this work, we propose a simple, efficient, and plug-and-play UQ method for 6-DoF object pose estimation. We ensemble 2–3 pre-trained models with different neural network architectures and/or training data sources, and compute their average pair-wise disagreement against one another to obtain the uncertainty quantification. We propose four disagreement metrics, including a learned metric, and show that the average distance (ADD) is the best learning-free metric and it is only slightly worse than the learned metric, which requires labeled target data. Our method has several advantages compared to the prior art: 1) our method does not require any modification of the training process or the model inputs; and 2) it needs only one forward pass for each model. We evaluate the proposed UQ method on three tasks where our uncertainty quantification yields much stronger correlations with pose estimation errors than the baselines. Moreover, in a real robot grasping task, our method increases the grasping success rate from 35% to 90%. Video and code are available at https://sites.google.com/view/fastuq. Guanya Shi, Jonathan Tremblay, Stanley T. Birchfield, Fabio Ramos 0001, Anima Anandkumar, Yuke Zhu |
ICRA | 4 |
| 2021 | Hierarchical Planning for Long-Horizon Manipulation with Geometric and Symbolic Scene GraphsabstractWe present a visually grounded hierarchical planning algorithm for long-horizon manipulation tasks. Our algorithm offers a joint framework of neuro-symbolic task planning and low-level motion generation conditioned on the specified goal. At the core of our approach is a two-level scene graph representation, namely geometric scene graph and symbolic scene graph. This hierarchical representation serves as a structured, object-centric abstraction of manipulation scenes. Our model uses graph neural networks to process these scene graphs for predicting high-level task plans and low-level motions. We demonstrate that our method scales to long-horizon tasks and generalizes well to novel task goals. We validate our method in a kitchen storage task in both physical simulation and the real world. Experiments show that our method achieves over 70% success rate and nearly 90% of subgoal completion rate on the real robot while being four orders of magnitude faster in computation time compared to standard search-based task-and-motion planner.1 Jonathan Tremblay, Stanley T. Birchfield, Yuke Zhu |
ICRA | 3 |
| 2021 | Joint Space Control via Deep Reinforcement LearningabstractThe dominant way to control a robot manipulator uses hand-crafted differential equations leveraging some form of inverse kinematics / dynamics. We propose a simple, versatile joint-level controller that dispenses with differential equations entirely. A deep neural network, trained via model-free reinforcement learning, is used to map from task space to joint space. Experiments show the method capable of achieving similar error to traditional methods, while greatly simplifying the process by automatically handling redundancy, joint limits, and acceleration / deceleration profiles. The basic technique is extended to avoid obstacles by augmenting the input to the network with information about the nearest obstacles. Results are shown both in simulation and on a real robot via sim-to-real transfer of the learned policy. We show that it is possible to achieve sub-centimeter accuracy, both in simulation and the real world, with a moderate amount of training. Visak Kumar, David Hoeller, Balakumar Sundaralingam, Jonathan Tremblay, Stanley T. Birchfield |
IROS | 5 |
| 2021 | Multi-view Fusion for Multi-level Robotic Scene UnderstandingabstractWe present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of unknown objects from categories corresponding to primitive shapes (e.g., cuboids and cylinders), and 3) full 6-DoF pose of known objects. By developing and fusing recent techniques in these domains, we provide a rich scene representation for robot awareness. We demonstrate the importance of each of these modules, their complementary nature, and the potential benefits of the system in the context of robotic manipulation. Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, Stanley T. Birchfield |
IROS | 5 |
| 2021 | RMPflow: A Geometric Framework for Generation of Multitask Motion PoliciesabstractGenerating robot motion for multiple tasks in dynamic environments is challenging, requiring an algorithm to respond reactively while accounting for complex nonlinear relationships between tasks. In this article, we develop a novel policy synthesis algorithm, Riemannian motion policy (RMP)flow, based on geometrically consistent transformations of RMPs. RMPs are a class of reactive motion policies that parameterize non-Euclidean behaviors as dynamical systems in intrinsically nonlinear task spaces. Given a set of RMPs designed for individual tasks, RMPflow can combine these policies to generate an expressive global policy, while simultaneously exploiting sparse structure for computational efficiency. We study the geometric properties of RMPflow and provide sufficient conditions for stability. Finally, we experimentally demonstrate that accounting for the natural Riemannian geometry of task policies can simplify classically difficult problems, such as planning through the clutter on high-degree-of-freedom manipulation systems.Note to Practitioners—Requirements on safety and responsiveness for collaborative robots have driven a need for new ideas in control design that bridge between standard objectives in low-level control (such as trajectory tracking) and high-level behavioral objectives (such as collision avoidance) often relegated to planning systems. Modern results from geometric control, which promise stable controllers that can smoothly and safely transition between many behavioral tasks, therefore, become highly relevant. However, for years, this field has remained inaccessible due to its mathematical complexity. This article aims to: 1) make those ideas accessible to robotics and control experts by recasting them in a concrete algorithmic framework amenable to controller design and 2) to additionally generalize them to better satisfy the specific needs of robotic behavior generation. Our experiments demonstrate that the resulting controllers can engender natural behavior that adapts instantaneously to changing surroundings with zero planning while performing manipulation tasks. The framework is gaining traction within the robotics community, finding increasing application in areas, such as autonomous navigation, tactile servoing, and multi-agent systems. Future research will address learning these controllers from data to simplify that process of design and tuning, which at present can require experience. Ching-An Cheng, Mustafa Mukadam, Jan Issac, Stanley T. Birchfield, Dieter Fox, Byron Boots, Nathan D. Ratliff |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2020 | DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm SystemabstractTeleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usually offer reduced degrees of control. Herein, a low-cost, depth-based teleoperation system, DexPilot, was developed that allows for complete control over the full 23 DoA robotic system by merely observing the bare human hand. DexPilot enabled operators to solve a variety of complex manipulation tasks that go beyond simple pick-and-place operations and performance was measured through speed and reliability metrics. DexPilot cost-effectively enables the production of high dimensional, multi-modality, state-action data that can be leveraged in the future to learn sensorimotor policies for challenging manipulation tasks. The videos of the experiments can be found at https://sites.google.com/view/dex-pilot. Ankur Handa, Karl Van Wyk, Wei Yang 0019, Jacky Liang, Yu-Wei Chao, Stanley T. Birchfield, Nathan D. Ratliff, Dieter Fox |
ICRA | 7 |
| 2020 | Toward Sim-to-Real Directional Semantic GraspingabstractWe address the problem of directional semantic grasping, that is, grasping a specific object from a specific direction. We approach the problem using deep reinforcement learning via a double deep Q-network (DDQN) that learns to map downsampled RGB input images from a wrist-mounted camera to Q-values, which are then translated into Cartesian robot control commands via the cross-entropy method (CEM). The network is learned entirely on simulated data generated by a custom robot simulator that models both physical reality (contacts) and perceptual quality (high-quality rendering). The reality gap is bridged using domain randomization. The system is an example of end-to-end (mapping input monocular RGB images to output Cartesian motor commands) grasping of objects from multiple pre-defined object-centric orientations, such as from the side or top. We show promising results in both simulation and the real world, along with some challenges faced and the need for future research in this area. Shariq Iqbal, Jonathan Tremblay, Andy Campbell, Kirby Leung, Thang To, Erik Leitch, Duncan McKay, Stanley T. Birchfield |
ICRA | 9 |
| 2020 | Camera-to-Robot Pose Estimation from a Single ImageabstractWe present an approach for estimating the pose of an external camera with respect to a robot using a single RGB image of the robot. The image is processed by a deep neural network to detect 2D projections of keypoints (such as joints) associated with the robot. The network is trained entirely on simulated data using domain randomization to bridge the reality gap. Perspective-n-point (PnP) is then used to recover the camera extrinsics, assuming that the camera intrinsics and joint configuration of the robot manipulator are known. Unlike classic hand-eye calibration systems, our method does not require an off-line calibration step. Rather, it is capable of computing the camera extrinsics from a single frame, thus opening the possibility of on-line calibration. We show experimental results for three different robots and camera sensors, demonstrating that our approach is able to achieve accuracy with a single frame that is comparable to that of classic off-line hand-eye calibration using multiple frames. With additional frames from a static pose, accuracy improves even further. Code, datasets, and pretrained models for three widely-used robot manipulators are made available. Timothy E. Lee, Jonathan Tremblay, Thang To, Terry Mosier, Oliver Kroemer, Dieter Fox, Stanley T. Birchfield |
ICRA | 8 |
| 2020 | MVLidarNet: Real-Time Multi-Class Scene Understanding for Autonomous Driving Using Multiple ViewsabstractAutonomous driving requires the inference of actionable information such as detecting and classifying objects, and determining the drivable space. To this end, we present Multi-View LidarNet (MVLidarNet), a two-stage deep neural network for multi-class object detection and drivable space segmentation using multiple views of a single LiDAR point cloud. The first stage processes the point cloud projected onto a perspective view in order to semantically segment the scene. The second stage then processes the point cloud (along with semantic labels from the first stage) projected onto a bird's eye view, to detect and classify objects. Both stages use an encoder-decoder architecture. We show that our multi-view, multi-stage, multi-class approach is able to detect and classify objects while simultaneously determining the drivable space using a single LiDAR scan as input, in challenging scenes with more than one hundred vehicles and pedestrians at a time. The system operates efficiently at 150 fps on an embedded GPU designed for a self-driving car, including a postprocessing step to maintain identities over time. We show results on both KITTI and a much larger internal dataset, thus demonstrating the method's ability to scale by an order of magnitude. Ryan Oldja, Nikolai Smolyanskiy, Stanley T. Birchfield, Alexander Popov, David Wehr, Ibrahim Eden, Joachim Pehserl |
IROS | 4 |
| 2020 | Indirect Object-to-Robot Pose Estimation from an External Monocular RGB CameraabstractWe present a robotic grasping system that uses a single external monocular RGB camera as input. The object-to-robot pose is computed indirectly by combining the output of two neural networks: one that estimates the object-to-camera pose, and another that estimates the robot-to-camera pose. Both networks are trained entirely on synthetic data, relying on domain randomization to bridge the sim-to-real gap. Because the latter network performs online camera calibration, the camera can be moved freely during execution without affecting the quality of the grasp. Experimental results analyze the effect of camera placement, image resolution, and pose refinement in the context of grasping several household objects. We also present results on a new set of 28 textured household toy grocery objects, which have been selected to be accessible to other researchers. To aid reproducibility of the research, we offer 3D scanned textured models, along with pre-trained weights for pose estimation. Jonathan Tremblay, Stephen Tyree, Terry Mosier, Stanley T. Birchfield |
IROS | 4 |
| 2019 | Few-Shot Viewpoint Estimation
Hung-Yu Tseng, Shalini De Mello, Jonathan Tremblay, Sifei Liu, Stanley T. Birchfield, Ming-Hsuan Yang 0001, Jan Kautz |
BMVC | 5 |
| 2019 | CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-IdentificationabstractUrban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km. To the best of our knowledge, CityFlow is the largest-scale dataset in terms of spatial coverage and the number of cameras/videos in an urban environment. The dataset contains more than 200K annotated bounding boxes covering a wide range of scenes, viewing angles, vehicle models, and urban traffic flow conditions. Camera geometry and calibration information are provided to aid spatio-temporal analysis. In addition, a subset of the benchmark is made available for the task of image-based vehicle re-identification (ReID). We conducted an extensive experimental evaluation of baselines/state-of-the-art approaches in MTMC tracking, multi-target single-camera (MTSC) tracking, object detection, and image-based ReID on this dataset, analyzing the impact of different network architectures, loss functions, spatio-temporal models and their combinations on task effectiveness. An evaluation server is launched with the release of our benchmark at the 2019 AI City Challenge (https://www.aicitychallenge.org/) that allows researchers to compare the performance of their newest techniques. We expect this dataset to catalyze research in this field, propel the state-of-the-art forward, and lead to deployed traffic optimization(s) in the real world. Milind Naphade, Ming-Yu Liu 0001, Xiaodong Yang 0001, Stanley T. Birchfield, Ratnesh Kumar 0004, David C. Anastasiu, Jenq-Neng Hwang |
CVPR | 5 |
| 2019 | PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic DataabstractIn comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (caused by the dependency of shape and appearance on viewpoint), and 2) small inter-class variability (caused by the similarity in shape and appearance between vehicles produced by different manufacturers). To address these challenges, we propose a Pose-Aware Multi-Task Re-Identification (PAMTRI) framework. This approach includes two innovations compared with previous methods. First, it overcomes viewpoint-dependency by explicitly reasoning about vehicle pose and shape via keypoints, heatmaps and segments from pose estimation. Second, it jointly classifies semantic vehicle attributes (colors and types) while performing ReID, through multi-task learning with the embedded pose representations. Since manually labeling images with detailed pose and attribute information is prohibitive, we create a large-scale highly randomized synthetic dataset with automatically annotated vehicle attributes for training. Extensive experiments validate the effectiveness of each proposed component, showing that PAMTRI achieves significant improvement over state-of-the-art on two mainstream vehicle ReID benchmarks: VeRi and CityFlow-ReID. Milind Naphade, Stanley T. Birchfield, Jonathan Tremblay, William Hodge, Ratnesh Kumar 0004, Xiaodong Yang 0001 |
ICCV | 3 |
| 2019 | Structured Domain Randomization: Bridging the Reality Gap by Context-Aware Synthetic DataabstractWe present structured domain randomization (SDR), a variant of domain randomization (DR) that takes into account the structure of the scene in order to add context to the generated data. In contrast to DR, which places objects and distractors randomly according to a uniform probability distribution, SDR places objects and distractors randomly according to probability distributions that arise from the specific problem at hand. In this manner, SDR-generated imagery enables the neural network to take the context around an object into consideration during detection. We demonstrate the power of SDR for the problem of 2D bounding box car detection, achieving competitive results on real data after training only on synthetic data. On the KITTI easy, moderate, and hard tasks, we show that SDR outperforms other approaches to generating synthetic data (VKITTI, Sim 200k, or DR), as well as real data collected in a different domain (BDD100K). Moreover, synthetic SDR data combined with real KITTI data outperforms real KITTI data alone. Aayush Prakash, Shaad Boochoon, Mark Brophy, David Acuna, Eric Cameracci, Gavriel State, Omer Shapira, Stanley T. Birchfield |
ICRA | 8 |
| 2019 | Robust Learning of Tactile Force Estimation through Robot InteractionabstractCurrent methods for estimating force from tactile sensor signals are either inaccurate analytic models or task-specific learned models. In this paper, we explore learning a robust model that maps tactile sensor signals to force. We specifically explore learning a mapping for the SynTouch BioTac sensor via neural networks. We propose a voxelized input feature layer for spatial signals and leverage information about the sensor surface to regularize the loss function. To learn a robust tactile force model that transfers across tasks, we generate ground truth data from three different sources: (1) the BioTac rigidly mounted to a force torque (FT) sensor, (2) a robot interacting with a ball rigidly attached to the same FT sensor, and (3) through force inference on a planar pushing task by formalizing the mechanics as a system of particles and optimizing over the object motion. A total of 140k samples were collected from the three sources. We achieve a median angular accuracy of 3.5 degrees in predicting force direction (66% improvement over the current state of the art) and a median magnitude accuracy of 0.06 N (93% improvement) on a test dataset. Additionally, we evaluate the learned force model in a force feedback grasp controller performing object lifting and gentle placement. Our results can be found on https: //sites.google.com/view/tactile-force. Balakumar Sundaralingam, Alexander Lambert, Ankur Handa, Byron Boots, Tucker Hermans, Stanley T. Birchfield, Nathan D. Ratliff, Dieter Fox |
ICRA | 6 |
| 2018 | Synthetically Trained Neural Networks for Learning Human-Readable Plans from Real-World DemonstrationsabstractWe present a system to infer and execute a human-readable program from a real-world demonstration. The system consists of a series of neural networks to perform perception, program generation, and program execution. Leveraging convolutional pose machines, the perception network reliably detects the bounding cuboids of objects in real images even when severely occluded, after training only on synthetic images using domain randomization. To increase the applicability of the perception network to new scenarios, the network is formulated to predict in image space rather than in world space. Additional networks detect relationships between objects, generate plans, and determine actions to reproduce a real-world demonstration. The networks are trained entirely in simulation, and the system is tested in the real world on the pick-and-place problem of stacking colored cubes using a Baxter robot. Jonathan Tremblay, Thang To, Artem Molchanov, Stephen Tyree, Jan Kautz, Stanley T. Birchfield |
ICRA | 6 |
| 2018 | RMPflow: A Computational Graph for Automatic Motion Policy Generation
Ching-An Cheng, Mustafa Mukadam, Jan Issac, Stanley T. Birchfield, Dieter Fox, Byron Boots, Nathan D. Ratliff |
WAFR | 4 |
| 2017 | Toward low-flying autonomous MAV trail navigation using deep neural networks for environmental awarenessabstractWe present a micro aerial vehicle (MAV) system, built with inexpensive off-the-shelf hardware, for autonomously following trails in unstructured, outdoor environments such as forests. The system introduces a deep neural network (DNN) called TrailNet for estimating the view orientation and lateral offset of the MAV with respect to the trail center. The DNN-based controller achieves stable flight without oscillations by avoiding overconfident behavior through a loss function that includes both label smoothing and entropy reward. In addition to the TrailNet DNN, the system also utilizes vision modules for environmental awareness, including another DNN for object detection and a visual odometry component for estimating depth for the purpose of low-level obstacle detection. All vision systems run in real time on board the MAV via a Jetson TX1. We provide details on the hardware and software used, as well as implementation details. We present experiments showing the ability of our system to navigate forest trails more robustly than previous techniques, including autonomous flights of 1 km. Nikolai Smolyanskiy, Alexey Kamenev, Jeffrey Smith 0002, Stanley T. Birchfield |
IROS | 4 |
| 2015 | RGBD Point Cloud Alignment Using Lucas-Kanade Data Association and Automatic Error Metric SelectionabstractWe propose to overcome a significant limitation of the iterative closest point (ICP) algorithm used by KinectFusion, namely, its sole reliance upon geometric information. Our approach uses both geometric and color information in a direct manner that uses all the data in order to accurately estimate camera pose. Data association is performed by Lucas-Kanade to compute an affine warp between the color images associated with two RGBD point clouds. A subsequent step then estimates the Euclidean transformation between the point clouds using either a point-to-point or point-to-plane error metric, with a novel method based on a normal covariance test for automatically selecting between them. Together, Lucas-Kanade data association with covariance testing enables robust camera tracking through areas of low geometric features, without sacrificing accuracy in environments in which the existing ICP technique succeeds. Experimental results on several publicly available datasets demonstrate the improved performance both qualitatively and quantitatively. Brian Peasley, Stanley T. Birchfield |
IEEE Trans. Robotics | 2 |
| 2014 | Efficient Hierarchical Graph-Based Segmentation of RGBD VideosabstractWe present an efficient and scalable algorithm for segmenting 3D RGBD point clouds by combining depth, color, and temporal information using a multistage, hierarchical graph-based approach. Our algorithm processes a moving window over several point clouds to group similar regions over a graph, resulting in an initial over-segmentation. These regions are then merged to yield a dendrogram using agglomerative clustering via a minimum spanning tree algorithm. Bipartite graph matching at a given level of the hierarchical tree yields the final segmentation of the point clouds by maintaining region identities over arbitrarily long periods of time. We show that a multistage segmentation with depth then color yields better results than a linear combination of depth and color. Due to its incremental processing, our algorithm can process videos of any length and in a streaming pipeline. The algorithm's ability to produce robust, efficient segmentation is demonstrated with numerous experimental results on challenging sequences from our own as well as public RGBD data sets. Steven Hickson, Stanley T. Birchfield, Irfan A. Essa, Henrik I. Christensen |
CVPR | 2 |
| 2014 | An inexpensive method for evaluating the localization performance of a mobile robot navigation systemabstractWe propose a method for evaluating the localization accuracy of an indoor navigation system in arbitrarily large environments. Instead of using externally mounted sensors, as required by most ground-truth systems, our approach involves mounting only landmarks consisting of distinct patterns printed on inexpensive foam boards. A pose estimation algorithm computes the pose of the robot with respect to the landmark using the image obtained by an on-board camera. We demonstrate that such an approach is capable of providing accurate estimates of a mobile robot's position and orientation with respect to the landmarks in arbitrarily-sized environments over arbitrarily-long trials. Furthermore, because the approach involves minimal outfitting of the environment, we show that only a small amount of setup time is needed to apply the method to a new environment. Experiments involving a state-of-the-art navigation system demonstrate the ability of the method to facilitate accurate localization measurements over arbitrarily long periods of time. Harsha Kikkeri, Gershon Parent, Mihai Jalobeanu, Stanley T. Birchfield |
ICRA | 4 |
| 2014 | Fast and accurate PoseSLAM by combining relative and global state spacesabstractWe revisit the question of state space in the context of performing loop closure. Although a relative state space has been previously discounted, we show that such a state space is actually extremely powerful, able to achieve recognizable results after just one iteration. The power behind the technique (called POReSS) is the coupling between parameters that causes the orientation of one node to affect the position and orientation of other nodes. At the same time, the approach is fast because, like the more popular incremental state space, the Jacobian never needs to be explicitly computed. Furthermore, we show that while POReSS is able to quickly compute a solution near the global optimum, it is not precise enough to perform the fine adjustments necessary to reach the global minimum. As a result, we augment POReSS with a fast variant of Gauss-Seidel (called Graph-Seidel) on a global state space to allow the solution to settle closer to the global minimum. We show that this combination of POReSS and Graph-Seidel converges more quickly and scales to very large graphs better than other techniques while at the same time computing a competitive residual. Brian Peasley, Stanley T. Birchfield |
ICRA | 2 |
| 2014 | Program synthesis by examples for object repositioning tasksabstractWe address the problem of synthesizing human-readable computer programs for robotic object repositioning tasks based on human demonstrations. A stack-based domain specific language (DSL) is introduced for object repositioning tasks, and a learning algorithm is proposed to synthesize a program in this DSL based on human demonstrations. Once the synthesized program has been learned, it can be rapidly verified and refined in the simulator via further demonstrations if necessary, then finally executed on an actual robot to accomplish the corresponding learned tasks in the physical world. By performing demonstrations on a novel tablet interface, the time required for teaching is greatly reduced compared with using a real robot. Experiments show a variety of object repositioning tasks such as sorting, kitting, and packaging can be programmed using this approach. Ashley Feniello, Hao Dang, Stanley T. Birchfield |
IROS | 3 |
| 2013 | Replacing Projective Data Association with Lucas-Kanade for KinectFusionabstractWe propose to overcome a significant limitation of the KinectFusion algorithm, namely, its sole reliance upon geometric information to estimate camera pose. Our approach uses both geometric and color information in a direct manner that uses all the data in order to perform the association of data between two RGBD point clouds. Data association is performed by aligning the two color images associated with the two point clouds by estimating a projective warp using the Lucas-Kanade algorithm. This projective warp is then used to create a correspondence map between the two point clouds, which is then used as the data association for a point-to-plane error minimization. This approach to correspondence allows camera tracking to be maintained through areas of low geometric features. We show that our proposed LKDA data association technique enables accurate scene reconstruction in environments in which low geometric texture causes the existing approach to fail, while at the same time demonstrating that the new technique does not adversely affect results in environments in which the existing technique succeeds. Brian Peasley, Stanley T. Birchfield |
ICRA | 2 |
| 2013 | 3D non-rigid deformable surface estimation without feature correspondenceabstractWe propose an algorithm, that extends our previous work, to estimate the current configuration of a non-rigid object using energy minimization and graph cuts. Our approach removes the need for feature correspondence or texture information and extends the boundary energy term. The object segmentation process is improved by using graph cuts along with a skin detector. We introduce an automatic mesh generator that provides a triangular mesh encapsulating the entire non-rigid object without predefined values. Our approach also handles in-plane rotation by reinitializing the mesh after data has been lost in the image sequence. Results display the proposed algorithm over a dataset consisting of seven shirts, two pairs of shorts, two posters, and a pair of pants. Bryan Willimon, Ian D. Walker, Stanley T. Birchfield |
ICRA | 3 |
| 2013 | A new approach to clothing classification using mid-level layersabstractWe present a novel approach for classifying items from a pile of laundry. The classification procedure exploits color, texture, shape, and edge information from 2D and 3D local and global information for each article of clothing using a Kinect sensor. The key contribution of this paper is a novel method of classifying clothing which we term L-M-H, more specifically L-C-S-H using characteristics and selection masks. Essentially, the method decomposes the problem into high (H), low (L) and multiple mid-level (characteristics(C), selection masks(S)) layers and produces “local” solutions to solve the global classification problem. Experiments demonstrate the ability of the system to efficiently classify and label into one of three categories (shirts, socks, or dresses). These results show that, on average, the classification rates, using this new approach with mid-level layers, achieve a true positive rate of 90%. Bryan Willimon, Ian D. Walker, Stanley T. Birchfield |
ICRA | 3 |
| 2013 | An Energy Minimization Approach to Automatic Traffic Camera CalibrationabstractWe present a method for automatic calibration of traffic cameras. The problem is formulated as one of energy minimization in reduced road-parameter space, from which internal and external camera parameters are determined. Our approach combines bottom-up processing of a video to find a vanishing point, lines in the background, and a directed activity map, along with top-down processing to fit a road model to these detected features using Markov chain Monte Carlo (MCMC). Enhanced autocorrelation along the dashed lines is used in conjunction with a best-fit road model to find road-to-image parameters. To maximize both robustness to noise and flexibility (e.g., to handle cases in which the camera is looking straight down the road), a single-vanishing-point length-based approach (VWL, according to the taxonomy in the work of Kanhere and Birchfield) is used. On a large number of data sets exhibiting a wide variety of conditions (including distractions such as bridges and on/off-ramps), our approach performs well, achieving less than 10% error in measuring test lengths in all cases. Douglas N. Dawson, Stanley T. Birchfield |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2012 | Occlusion-aware reconstruction and manipulation of 3D articulated objectsabstractWe present a method to recover complete 3D models of articulated objects. Structure-from-motion techniques are used to capture 3D point cloud models of the object in two different configurations. A novel combination of Procrustes analysis and RANSAC facilitates a straightforward geometric approach to recovering the joint axes, as well as classifying them automatically as either revolute or prismatic. With the resulting articulated model, a robotic system is able to manipulate the object along its joint axes at a specified grasp point in order to exercise its degrees of freedom. Because the models capture all sides of the object, they are occluded-aware, enabling the robotic system to plan paths to parts of the object that are not visible in the current view. Our algorithm does not require prior knowledge of the object, nor does it make any assumptions about the planarity of the object or scene. Experiments with a PUMA 500 robotic arm demonstrate the effectiveness of the approach on a variety of objects with both revolute and prismatic joints. Ian D. Walker, Stanley T. Birchfield |
ICRA | 3 |
| 2012 | Accurate on-line 3D occupancy grids using Manhattan world constraintsabstractIn this paper we present an algorithm for constructing nearly drift-free 3D occupancy grids of large indoor environments in an online manner. Our approach combines data from an odometry sensor with output from a visual registration algorithm, and it enforces a Manhattan world constraint by utilizing factor graphs to produce an accurate online estimate of the trajectory of a mobile robotic platform. We also examine the advantages and limitations of the octree data structure representation of a 3D environment. Through several experiments in environments with varying sizes and construction we show that our method reduces rotational and translational drift significantly without performing any loop closing techniques. Brian Peasley, Stanley T. Birchfield, Alexander Cunningham, Frank Dellaert |
IROS | 2 |
| 2012 | An energy minimization approach to 3D non-rigid deformable surface estimation using RGBD dataabstractWe propose an algorithm that uses energy minimization to estimate the current configuration of a non-rigid object. Our approach utilizes an RGBD image to calculate corresponding SURF features, depth, and boundary information. We do not use predetermined features, thus enabling our system to operate on unmodified objects. Our approach relies on a 3D nonlinear energy minimization framework to solve for the configuration using a semi-implicit scheme. Results show various scenarios of dynamic posters and shirts in different configurations to illustrate the performance of the method. In particular, we show that our method is able to estimate the configuration of a textureless nonrigid object with no correspondences available. Bryan Willimon, Steven Hickson, Ian D. Walker, Stanley T. Birchfield |
IROS | 4 |
| 2011 | Classification of clothing using interactive perceptionabstractWe present a system for automatically extracting and classifying items in a pile of laundry. Using only visual sensors, the robot identifies and extracts items sequentially from the pile. When an item has been removed and isolated, a model is captured of the shape and appearance of the object, which is then compared against a database of known items. The classification procedure relies upon silhouettes, edges, and other low-level image measurements of the articles of clothing. The contributions of this paper are a novel method for extracting articles of clothing from a pile of laundry and a novel method of classifying clothing using interactive perception. Experiments demonstrate the ability of the system to efficiently classify and label into one of six categories (pants, shorts, short-sleeve shirt, long-sleeve shirt, socks, or underwear). These results show that, on average, classification rates using robot interaction are 59% higher than those that do not use interaction. Bryan Willimon, Stanley T. Birchfield, Ian D. Walker |
ICRA | 2 |
| 2011 | Model for unfolding laundry using interactive perceptionabstractWe present an algorithm for automatically unfolding a piece of clothing. A piece of laundry is pulled in different directions at various points of the cloth in order to flatten the laundry. The features of the cloth are extracted and calculated to determine a valid location and orientation in which to interact with it. The features include the peak region, corner locations, and continuity / discontinuity of the cloth. In this paper we present a two-stage algorithm, introducing a novel solution to the unfolding / flattening problem using interactive perception. Simulations using 3D simulation software, and experiments with robot hardware demonstrate the ability of the algorithm to flatten pieces of laundry using different starting configurations. These results show that, at most, the algorithm flattens out a piece of cloth from 11.1% to 95.6% of the canonical configuration. Bryan Willimon, Stanley T. Birchfield, Ian D. Walker |
IROS | 2 |
| 2010 | Image-based segmentation of indoor corridor floors for a mobile robotabstractWe present a novel method for image-based floor detection from a single image. In contrast with previous approaches that rely upon homographies, our approach does not require multiple images (either stereo or optical flow). It also does not require the camera to be calibrated, even for lens distortion. The technique combines three visual cues for evaluating the likelihood of horizontal intensity edge line segments belonging to the wall-floor boundary. The combination of these cues yields a robust system that works even in the presence of severe specular reflections, which are common in indoor environments. The nearly real-time algorithm is tested on a large database of images collected in a wide variety of conditions, on which it achieves nearly 90% detection accuracy. Yinxiao Li, Stanley T. Birchfield |
IROS | 2 |
| 2010 | Rigid and non-rigid classification using interactive perceptionabstractRobotics research tends to focus upon either non-contact sensing or machine manipulation, but not both. This paper explores the benefits of combining the two by addressing the problem of classifying unknown objects, such as found in service robot applications. In the proposed approach, an object lies on a flat background, and the goal of the robot is to interact with and classify each object so that it can be studied further. The algorithm considers each object to be classified using color, shape, and flexibility. Experiments on a number of different objects demonstrate the ability of efficiently classifying and labeling each item through interaction. Bryan Willimon, Stanley T. Birchfield, Ian D. Walker |
IROS | 2 |
| 2010 | Iris segmentation in non-ideal images using graph cuts
Shrinivas J. Pundlik, Damon L. Woodard, Stanley T. Birchfield |
Image Vis. Comput. | 3 |
| 2010 | Rapid automated detection of roots in minirhizotron images
Stanley T. Birchfield, Christina E. Wells |
Mach. Vis. Appl. | 2 |
| 2010 | A Taxonomy and Analysis of Camera Calibration Methods for Traffic Monitoring ApplicationsabstractMany vision-based automatic traffic-monitoring systems require a calibrated camera to compute the speeds and length-based classifications of tracked vehicles. A number of techniques, both manual and automatic, have been proposed for performing such calibration, but no study has yet focused on evaluating the relative strengths of these different alternatives. We present a taxonomy for roadside camera calibration that not only encompasses the existing methods (VVW, VWH, and VWL) but also includes several novel methods (VVH, VVL, VLH, VVD, VWD, and VHD). We also introduce an overconstrained (OC) approach that takes into account all the available measurements, resulting in reduced error and overcoming the inherent ambiguity in single-vanishing-point solutions. This important but oft-neglected ambiguity has not received the attention that it deserves; we analyze it and propose several ways of overcoming it. Our analysis includes the relative tradeoffs between two-vanishing-point solutions, single-vanishing-point solutions, and solutions that require the distance to the road to be known. The various methods are compared using simulations and experiments with real images, showing that methods that use a known length generally outperform the others in terms of error and that the OC method reduces errors even further. Neeraj K. Kanhere, Stanley T. Birchfield |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2009 | Adaptive fragments-based tracking of non-rigid objects using level setsabstractWe present an approach to visual tracking based on dividing a target into multiple regions, or fragments. The target is represented by a Gaussian mixture model in a joint feature-spatial space, with each ellipsoid corresponding to a different fragment. The fragments are automatically adapted to the image data, being selected by an efficient region-growing procedure and updated according to a weighted average of the past and present image statistics. Modeling of target and background are performed in a Chan-Vese manner, using the framework of level sets to preserve accurate boundaries of the target. The extracted target boundaries are used to learn the dynamic shape of the target over time, enabling tracking to continue under total occlusion. Experimental results on a number of challenging sequences demonstrate the effectiveness of the technique. Prakash Chockalingam, S. Nalin Pradeep, Stanley T. Birchfield |
ICCV | 3 |
| 2009 | Qualitative Vision-Based Path FollowingabstractWe present a simple approach for vision-based path following for a mobile robot. Based upon a novel concept called the funnel lane, the coordinates of feature points during the replay phase are compared with those obtained during the teaching phase in order to determine the turning direction. Increased robustness is achieved by coupling the feature coordinates with odometry information. The system requires a single off-the-shelf, forward-looking camera with no calibration (either external or internal, including lens distortion). Implicit calibration of the system is needed only in the form of a single controller gain. The algorithm is qualitative in nature, requiring no map of the environment, no image Jacobian, no homography, no fundamental matrix, and no assumption about a flat ground plane. Experimental results demonstrate the capability of real-time autonomous navigation in both indoor and outdoor environments and on flat, slanted, and rough terrain with dynamic occluding objects for distances of hundreds of meters. We also demonstrate that the same approach works with wide-angle and omnidirectional cameras with only slight modification. Stanley T. Birchfield |
IEEE Trans. Robotics | 2 |
| 2008 | Joint tracking of features and edgesabstractSparse features have traditionally been tracked from frame to frame independently of one another. We propose a framework in which features are tracked jointly. Combining ideas from Lucas-Kanade and Horn-Schunck, the estimated motion of a feature is influenced by the estimated motion of neighboring features. The approach also handles the problem of tracking edges in a unified way by estimating motion perpendicular to the edge, using the motion of neighboring features to resolve the aperture problem. Results are shown on several image sequences to demonstrate the improved results obtained by the approach. Stanley T. Birchfield, Shrinivas J. Pundlik |
CVPR | 1 |
| 2008 | Limbus/pupil switching for wearable eye tracking under variable lighting conditionsabstractWe present a low-cost wearable eye tracker built from off-the-shelf components. Based on the open source openEyes project (the only other similar effort that we are aware of), our eye tracker operates in the visible spectrum and variable lighting conditions. The novelty of our approach rests in automatically switching between tracking the pupil/iris boundary in bright light to tracking the iris/sclera boundary (limbus) in dim light. Additional improvements include a semi-automatic procedure for calibrating the eye and scene cameras, as well as an automatic procedure for initializing the location of the pupil in the first image frame. The system is accurate to two degrees visual angle in both indoor and outdoor environments. Wayne J. Ryan, Andrew T. Duchowski, Stanley T. Birchfield |
ETRA | 3 |
| 2008 | Real-Time Incremental Segmentation and Tracking of Vehicles at Low Camera Angles Using Stable FeaturesabstractWe present a method for segmenting and tracking vehicles on highways using a camera that is relatively low to the ground. At such low angles, 3-D perspective effects cause significant changes in appearance over time, as well as severe occlusions by vehicles in neighboring lanes. Traditional approaches to occlusion reasoning assume that the vehicles initially appear well separated in the image; however, in our sequences, it is not uncommon for vehicles to enter the scene partially occluded and remain so throughout. By utilizing a 3-D perspective mapping from the scene to the image, along with a plumb line projection, we are able to distinguish a subset of features whose 3-D coordinates can be accurately estimated. These features are then grouped to yield the number and locations of the vehicles, and standard feature tracking is used to maintain the locations of the vehicles over time. Additional features are then assigned to these groups and used to classify vehicles as cars or trucks. Our technique uses a single grayscale camera beside the road, incrementally processes image frames, works in real time, and produces vehicle counts with over 90% accuracy on challenging sequences. Neeraj K. Kanhere, Stanley T. Birchfield |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2008 | Real-Time Motion Segmentation of Sparse Feature Points at Any SpeedabstractWe present a real-time incremental approach to motion segmentation operating on sparse feature points. In contrast to previous work, the algorithm allows for a variable number of image frames to affect the segmentation process, thus enabling an arbitrary number of objects traveling at different relative speeds to be detected. Feature points are detected and tracked throughout an image sequence, and the features are grouped using a spatially constrained expectation-maximization (EM) algorithm that models the interactions between neighboring features using the Markov assumption. The primary parameter used by the algorithm is the amount of evidence that must accumulate before features are grouped. A statistical goodness-of-fit test monitors the change in the motion parameters of a group over time in order to automatically update the reference frame. Experimental results on a number of challenging image sequences demonstrate the effectiveness and computational efficiency of the technique. Shrinivas J. Pundlik, Stanley T. Birchfield |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Isomap Tracking with Particle FilteringabstractThe problem of tracking involves challenges like in-plane and out-of-plane rotations, scaling, variations in ambient light and occlusions. In this paper we look at the problem of tracking a person's head and also estimating its pose in each frame. Robust tracking can be achieved by reducing the dimensionality of high-dimensional training data and using the recovered low-dimensional structure to estimate the state of an object at every time-step with recursive Bayesian filtering. Isometric feature mapping, also known as Isomap, provides an unsupervised framework to find the true degrees of freedom in high-dimensional input data like a person's head with varying poses. After the data has been reduced to lower dimensions a particle filter can be used to track and at the same time approximate the pose of a person's head in any image sequence. Isomap tracking with particle filtering is capable of handling rapid translation and out-of-plane rotation of a person's head with a relatively small amount of training data. The performance of the tracker is demonstrated on an image sequence with a person's head undergoing translation and out-of-plane rotation. Nikhil Rane, Stanley T. Birchfield |
ICIP (2) | 2 |
| 2007 | Person following with a mobile robot using binocular feature-based trackingabstractWe present the Binocular Sparse Feature Segmentation (BSFS) algorithm for vision-based person following with a mobile robot. BSFS uses Lucas-Kanade feature detection and matching in order to determine the location of the person in the image and thereby control the robot. Matching is performed between two images of a stereo pair, as well as between successive video frames. We use the Random Sample Consensus (RANSAC) scheme for segmenting the sparse disparity map and estimating the motion models of the person and background. By fusing motion and stereo information, BSFS handles difficult situations such as dynamic backgrounds, out-of-plane rotation, and similar disparity and/or motion between the person and background. Unlike color-based approaches, the person is not required to wear clothing with a different color from the environment. Our system is able to reliably follow a person in complex dynamic, cluttered environments in real time. Stanley T. Birchfield |
IROS | 2 |
| 2007 | Correspondence as energy-based segmentation
Stanley T. Birchfield, Braga Natarajan, Carlo Tomasi |
Image Vis. Comput. | 1 |
| 2006 | Automated Patient Facial Image Capture to Reduce Medical Error
Michael Gillam, Craig Feied, Stanley T. Birchfield, Jonathan A. Handler, Mark S. Smith |
AMIA | 3 |
| 2006 | Motion Segmentation at Any SpeedabstractWe present an incremental approach to motion segmentation. Feature points are detected and tracked throughout an image sequence, and the features are grouped using a region-growing algorithm with an affine motion model. The primary parameter used by the algorithm is the amount of evidence that must accumulate before features are grouped. Contrasted with previous work, the algorithm allows for a variable number of image frames to affect the decision process, thus enabling objects to be detected independently of their velocity in the image. Procedures are presented for grouping features, measuring the consistency of the resulting groups, assimilating new features into existing groups, and splitting groups over time. Experimental results on a number of challenging image sequences demonstrate the effectiveness of the technique. 1 Shrinivas J. Pundlik, Stanley T. Birchfield |
BMVC | 2 |
| 2006 | Qualitative Vision-based Mobile Robot NavigationabstractWe present a novel, simple algorithm for mobile robot navigation. Using a teach-replay approach, the robot is manually led along a desired path in a teaching phase, then the robot autonomously follows that path in a replay phase. The technique requires a single off-the-shelf, forward-looking camera with no calibration (including no calibration for lens distortion). Feature points are automatically detected and tracked throughout the image sequence, and the feature coordinates in the replay phase are compared with those computed previously in the teaching phase to determine the turning commands for the robot. The algorithm is entirely qualitative in nature, requiring no map of the environment, no image Jacobian, no homography, no fundamental matrix, and no assumption about a flat ground plane. Experimental results demonstrate the capability of autonomous navigation in both indoor and outdoor environments, on both flat and slanted surfaces, with dynamic occluding objects, for distances over 100 m Stanley T. Birchfield |
ICRA | 2 |
| 2006 | Detecting and Measuring Fine Roots in Minirhizotron Images Using Matched Filtering and Local Entropy Thresholding
Stanley T. Birchfield, Christina E. Wells |
Mach. Vis. Appl. | 2 |
| 2005 | Spatiograms versus Histograms for Region-Based TrackingabstractWe introduce the concept of a spatiogram, which is a generalization of a histogram that includes potentially higher order moments. A histogram is a zeroth-order spatiogram, while second-order spatiograms contain spatial means and covariances for each histogram bin. This spatial information still allows quite general transformations, as in a histogram, but captures a richer description of the target to increase robustness in tracking. We show how to use spatiograms in kernel-based trackers, deriving a mean shift procedure in which individual pixels vote not only for the amount of shift but also for its direction. Experiments show improved tracking results compared with histograms, using both mean shift and exhaustive local search. Stanley T. Birchfield, Sriram Rangarajan |
CVPR (2) | 1 |
| 2005 | Vehicle Segmentation and Tracking from a Low-Angle Off-Axis CameraabstractWe present a novel method for visually monitoring a highway when the camera is relatively low to the ground and on the side of the road. In such a case, occlusion and the perspective effects due to the heights of the vehicles cannot be ignored. Features are detected and tracked throughout the image sequence, and then grouped together using a multilevel homography, which is an extension of the standard homography to the low-angle situation. We derive a concept called the relative height constraint that makes it possible to estimate the 3D height of feature points on the vehicles from a single camera, a key part of the technique. Experimental results on several different highways demonstrate the system's ability to successfully segment and track vehicles at low angles, even in the presence of severe occlusion and significant perspective changes. Neeraj K. Kanhere, Shrinivas J. Pundlik, Stanley T. Birchfield |
CVPR (2) | 3 |
| 2005 | Acoustic localization by interaural level differenceabstractInteraural level difference (ILD) is an important cue for acoustic localization. Although its behavior has been studied extensively in natural systems, it remains an untapped resource for computer-based systems. We investigate the possibility of using ILD for acoustic localization, deriving constraints on the location of a sound source given the relative energy level of the signals received by two microphones. We then present an algorithm for computing the sound source location by combining likelihood functions, one for each microphone pair. Experimental results show that accurate acoustic localization can be achieved using ILD alone. Stanley T. Birchfield, Rajitha Gangishetty |
ICASSP (4) | 1 |
| 2005 | Microphone Array Position Calibration by Basis-Point Classical Multidimensional ScalingabstractClassical multidimensional scaling (MDS) is a global, noniterative technique for finding coordinates of points given their interpoint distances. We describe the algorithm and show how it yields a simple, inexpensive method for calibrating an array of microphones with a tape measure (or similar measuring device). We present an extension to the basic algorithm, called basis-point classical MDS (BCMDS), which handles the case when many of the distances are unavailable, thus yielding a technique that is practical for microphone arrays with a large number of microphones. We also show that BCMDS, when combined with a calibration target consisting of four synchronized sound sources, can be used for automatic calibration via time-delay estimation. We evaluate the accuracy of both classical MDS and BCMDS, investigating the sensitivity of the algorithms to noise and to the design parameters to yield insight as to the choice of those parameters. Our results validate the practical applicability of the algorithms, showing that errors on the order of 10-20 mm can be achieved in real scenarios. Stanley T. Birchfield, Amarnag Subramanya |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Geometric microphone array calibration by multidimensional scalingabstractClassical multidimensional scaling is a simple, linear technique for finding coordinates of points given their interpoint distances. In this paper we describe the algorithm and show how it can be used to solve the geometric microphone array calibration problem. The method requires no complicated hardware or calibration targets, just a tape measure (or similar measuring device). We also extend the basic algorithm to handle the case when some distances are unavailable, which makes the technique practical for microphone arrays with relatively large numbers of microphones. Stanley T. Birchfield |
ICASSP (5) | 1 |
| 2002 | Fast Bayesian acoustic localizationabstractWe derive a probabilistic formulation, based upon Bayes' rule, for the acoustic localization problem. The resulting formula is shown to be closely related to the energy of a conventionally beamformed signal. We then present a close approximation to both which is much faster to compute — by two orders of magnitude with our experimental setup. The fast algorithm is essentially a generalization of approaches based upon time delay estimates (TDE's), by applying the principle of least commitment. Experiments on real signals demonstrate accurate localization in noisy, reverberant environments (less than 3 dB SNR) several times faster than real time. Stanley T. Birchfield, Daniel Kahn Gillmor |
ICASSP | 1 |
| 2001 | Acoustic source direction by hemisphere samplingabstractA method for estimating the direction to a sound source, using a compact array of microphones, is presented. For each pair of microphones, the signals are prefiltered and correlated. Rather than taking the peak of the correlation vectors as estimates for the time delay between the microphones, all the correlation vectors are accumulated in a common coordinate system, namely a unit hemisphere centered on the microphone array. The maximum cell in the hemisphere then indicates the azimuthal and elevation angles to the source. Unlike previous techniques, this algorithm is applicable to arbitrary microphone configurations, handles more than two microphone pairs, and has no blind spots. Experiments demonstrate significantly increased robustness to noise, compared with previous techniques. Stanley T. Birchfield, Daniel Kahn Gillmor |
ICASSP | 1 |
| 1999 | Multiway Cut for Stereo and Motion with Slanted SurfacesabstractSlanted surfaces pose a problem for correspondence algorithms utilizing search because of the greatly increased number of possibilities, when compared with fronto-parallel surfaces. In this paper we propose an algorithm to compute correspondence between stereo images or between frames of a motion sequence by minimizing an energy functional that accounts for slanted surfaces. The energy is minimized in a greedy strategy that alternates between segmenting the image into a number of non-overlapping regions (using the multiway-cut algorithm of Boykov, Veksler, and Zabih) and finding the affine parameters describing the displacement function of each region. A follow-up step enables the algorithm to escape local minima due to oversegmentation. Experiments on real images show the algorithm's ability to find an accurate segmentation and displacement map, as well as discontinuities and creases, from a wide variety of stereo and motion imagery. Stanley T. Birchfield, Carlo Tomasi |
ICCV | 1 |
| 1999 | Depth Discontinuities by Pixel-to-Pixel Stereo
Stanley T. Birchfield, Carlo Tomasi |
Int. J. Comput. Vis. | 1 |
| 1998 | Elliptical Head Tracking Using Intensity Gradients and Color HistogramsabstractAn algorithm for tracking a person's head is presented. The head's projection onto the image plane is modeled as an ellipse whose position and size are continually updated by a local search combining the output of a module concentrating on the intensity gradient around the ellipse's perimeter with that of another module focusing on the color histogram of the ellipse's interior. Since these two modules have roughly orthogonal failure modes, they serve to complement one another. The result is a robust, real-time system that is able to track a person's head with enough accuracy to automatically control the camera's pan, tilt, and zoom in order to keep the person centered in the field of view at a desired size. Extensive experimentation shows the algorithm's robustness with respect to full 360-degree out-of-plane rotation, up to 90-degree tilting, severe but brief occlusion, arbitrary camera movement, and multiple moving people in the background. Stanley T. Birchfield |
CVPR | 1 |
| 1998 | Depth Discontinuities by Pixel-to-Pixel StereoabstractAn algorithm to detect depth discontinuities from a stereo pair of images is presented. The algorithm matches individual pixels in corresponding scanline pairs while allowing occluded pixels to remain unmatched, then propagates the information between scanlines by means of a fast postprocessor. The algorithm handles large untextured regions, uses a measure of pixel dissimilarity that is insensitive to image sampling, and prunes bad search nodes to increase the speed of dynamic programming. The computation is relatively fast, taking about 1.5 microseconds per pixel per disparity on a workstation. Approximate disparity maps and precise depth discontinuities (along both horizontal and vertical boundaries) are shown for five stereo images containing textured, untextured, fronto-parallel, and slanted objects. Stanley T. Birchfield, Carlo Tomasi |
ICCV | 1 |
| 1998 | A Pixel Dissimilarity Measure That Is Insensitive to Image SamplingabstractBecause of image sampling, traditional measures of pixel dissimilarity can assign a large value to two corresponding pixels in a stereo pair, even in the absence of noise and other degrading effects. We propose a measure of dissimilarity that is provably insensitive to sampling because it uses the linearly interpolated intensity functions surrounding the pixels. Experiments on real images show that our measure alleviates the problem of sampling with little additional computational overhead. Stanley T. Birchfield, Carlo Tomasi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |