Derek Hoiem

dblp:08/6948 · DBLP profile ↗
← Back
95ranked-venue papers
15as first author
21since 2021 · last 2025
0000-0001-6260-5708ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 78 · 12 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 9 first-author · 13 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
abstract
The goal of this paper is to encode a 3D scene into an extremely compact representation from$2 D$images and to enable its transmittance, decoding and rendering in real-time across various platforms. Despite the progress in NeRFs and Gaussian Splats, their large model size and specialized renderers make it challenging to distribute free-viewpoint 3D content as easily as images. To address this, we have designed a novel 3D representation that encodes the plenoptic function into sinusoidal function indexed dense volumes. This approach facilitates feature sharing across different locations, improving compactness over traditional spatial voxels. The memory footprint of the dense 3D feature grid can be further reduced using spatial decomposition techniques. This design combines the strengths of spatial hashing functions and voxel decomposition, resulting in a model size as small as 150 KB for each 3D scene. Moreover, PPNG features a lightweight rendering pipeline with only 300 lines of code that decodes its representation into standard GL textures and fragment shaders. This enables realtime rendering using the traditional GL pipeline, ensuring universal compatibility and efficiency across various platforms without additional dependencies. Our results are available at: https://jyl.kr/ppng
Jae Yong Lee 0006, Yuqun Wu, Chuhang Zou, Derek Hoiem, Shenlong Wang
3DV4
2025 MonoPatchNeRF: Improving Neural Radiance Fields with Patch-Based Monocular Guidance
abstract
The latest regularized Neural Radiance Field (NeRF) approaches produce poor geometry and view extrapolation for large scale sparse view scenes, such as ETH3D. Density-based approaches tend to be under-constrained, while surface-based approaches tend to miss details. In this paper, we take a density-based approach, sampling patches instead of individual rays to better incorporate monocular depth and normal estimates and patch-based photometric consistency constraints between training views and sampled virtual views. Loosely constraining densities based on estimated depth aligned to sparse points further improves geometric accuracy. While maintaining similar view synthesis quality, our approach significantly improves geometric accuracy on the ETH3D benchmark, e.g. increasing the F1@2cm score by 4x-8x compared to other regularized density-based approaches, with much lower training and inference time than other approaches.
Yuqun Wu, Jae Yong Lee 0006, Chuhang Zou, Shenlong Wang, Derek Hoiem
3DV5
2025 RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
abstract
We present Relocate, a simple training-free baseline designed to perform the challenging task of visual query localization in long videos. To eliminate the need for task-specific training and efficiently handle long videos, Relocate leverages a region-based representation derived from pretrained vision models. At a high level, it follows the classic object localization approach: (1) identify all objects in each video frame, (2) compare the objects with the given query and select the most similar ones, and (3) perform bidirectional tracking to get a spatio-temporal response. However, we propose some key enhancements to handle small objects, cluttered scenes, partial visibility, and varying appearances. Notably, we refine the selected objects for accurate localization and generate additional visual queries to capture visual variations. We evaluate Relocate on the challenging Ego4D Visual Query 2D Localization dataset, establishing a new baseline that outperforms prior task-specific methods by 49% (relative improvement) in spatio-temporal average precision.
Savya Khosla, Sethuraman TV, Alexander G. Schwing, Derek Hoiem
CVPR4
2025 PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
abstract
Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning—yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward task. In this work, we introduce PARTONOMY, an LMM benchmark designed for pixel-level part grounding. We construct PARTONOMY from existing part datasets and our own rigorously annotated set of images, encompassing 862 parts and 5346 objects for evaluation. Unlike existing datasets that simply ask models to identify generic parts, PARTONOMY utilizes highly technical concepts and challenges models to compare objects’ parts, consider part-whole relationships, and justify textual predictions with visual segmentations. Our experiments demonstrate significant limitations in state-of-the-art LMMs (e.g., LISA-13B achieves only 5.9% gIoU), highlighting a critical gap in their part grounding abilities. We note that existing segmentation-enabled LMMs (segmenting LMMs) have two key architectural shortcomings: they use special [SEG] tokens not seen during pretraining which induce distribution shift, and they discard predicted segmentations instead of using past predictions to guide future ones. To address these deficiencies, we train several part-centric LMMs and propose PLUM, a novel segmenting LMM that utilizes span tagging instead of segmentation tokens and that conditions on prior predictions in a feedback loop. We find that pretrained PLUM dominates existing segmenting LMMs on reasoning segmentation, VQA, and visual hallucination benchmarks. In addition, PLUM finetuned on our proposed Explanatory Part Segmentation task is competitive with segmenting LMMs trained on significantly more segmentation data. Our work opens up new avenues towards enabling fine-grained, grounded visual understanding in LMMs.
Ansel Blume, Hyeonjeong Ha, Elen Chatikyan, Xiaomeng Jin, Khanh Duy Nguyen, Nanyun Peng 0001, Kai-Wei Chang 0001, Derek Hoiem, Heng Ji 0001
NeurIPS9
2025 REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
abstract
We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnostic segmenters (e.g., SAM) with patch-based image encoders (e.g., DINO) to produce compact and effective region representations, but they suffer from high computational cost due to the segmentation step. REN bypasses this bottleneck using a lightweight module that directly generates region tokens, enabling 60x faster token generation with 35x less memory, while also improving token quality. It uses a few cross-attention blocks that take point prompts as queries and features from a patch-based image encoder as keys and values to produce region tokens that correspond to the prompted objects. We train REN with three popular encoders—DINO, DINOv2, and OpenCLIP—and show that it can be extended to other encoders without dedicated training. We evaluate REN on semantic segmentation and retrieval tasks, where it consistently outperforms the original encoders in both performance and compactness, and matches or exceeds SAM-based region methods while being significantly faster. Notably, REN achieves state-of-the-art results on the challenging Ego4D VQ2D benchmark and outperforms proprietary LMMs on Visual Haystacks' single-needle challenge. The code and pretrained models are available at https://github.com/savya08/ren.
Savya Khosla, Sethuraman TV, Barnett Lee, Alexander G. Schwing, Derek Hoiem
NeurIPS5
2024 Addressing Low-Shot MVS by Detecting and Completing Planar Surfaces
abstract
Multiview stereo (MVS) systems typically require at least three views to reconstruct each scene point. This requirement increases the burden of image captures and leads to incomplete reconstructions. Our main idea to address this low-shot MVS problem is to detect planar surfaces in depth maps generated by any MVS system and complete these surfaces by reformulating the MVS depth prediction task to a simpler planar surface assignment problem. We use single and multi-view cues (when available) and employ the DeepLabv3 architecture to infer the extent of planar regions and accurately complete missing surfaces. We show that our approach reconstructs portions of surfaces viewed by only one image, yielding denser models than existing MVS systems.
Rajbir Kataria, Zhizhong Li 0001, Joseph DeGol, Derek Hoiem
3DV4
2024 Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
abstract
We present Unified-IO 2,the. first autoregressive multi-modal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs - images, text, audio, action, bounding boxes etc., into a shared semantic space and then process them with a single encoder-decoder transformer model. Since training with such diverse modalities is challenging, we propose various architectural improvements to stabilize model training. We train our model from scratch on a large multimodal pre-training corpus from diverse sources with a multimodal mixture of denoisers objective. To learn an expansive set of skills, such as following multimodal instructions, we construct and. finetune on an ensemble of 120 datasets with prompts and augmentations. With a single unified model, Unified-io 2 achieves state-of-the-art performance on the GRIT benchmark and strong results in more than 35 benchmarks, including image generation and understanding, natural language understanding, video and audio understanding, and robotic manipulation. We release all our models to the research community.
Jiasen Lu, Sangho Lee 0008, Zichen Zhang 0016, Savya Khosla, Ryan Marten, Derek Hoiem, Aniruddha Kembhavi
CVPR7
2024 Region-Based Representations Revisited
abstract
We investigate whether region-based representations are effective for recognition. Regions were once a mainstay in recognition approaches, but pixel and patch-based features are now used almost exclusively. We show that recent class-agnostic segmenters like SAM can be effectively combined with strong self-supervised representations, like those from DINOv2, and used for a wide variety of tasks, including semantic segmentation, object-based image re-trieval, and multi-image analysis. Once the masks and features are extracted, these representations, even with linear decoders, enable competitive performance, making them well suited to applications that require custom queries. The representations' compactness also makes them well-suited to video analysis and other problems requiring inference across many images.
Michal Shlapentokh-Rothman, Ansel Blume, Yuqun Wu, Sethuraman TV, Heyi Tao, Jae Yong Lee 0006, Wilfredo Torres, Yu-Xiong Wang, Derek Hoiem
CVPR10
2024 Anytime Continual Learning for Open Vocabulary Classification
Zhen Zhu 0006, Yiming Gong, Derek Hoiem
ECCV (6)3
2024 MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
abstract
We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.
Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001
ACM Multimedia29
2024 Consistent Multimodal Generation via A Unified GAN Framework
abstract
We investigate how to generate multimodal image outputs, such as RGB, depth, and surface normals, with a single generative model. The challenge is to produce outputs that are realistic, and also consistent with each other. Our solution builds on the StyleGAN3 architecture, with a shared backbone and modality-specific branches in the last layers of the synthesis network, and we propose per-modality fidelity discriminators and a cross-modality consistency discriminator. In experiments on the Stanford2D3D dataset, we demonstrate realistic and consistent generation of RGB, depth, and normal images. We also show a training recipe to easily extend our pretrained model on a new domain, even with a few pairwise data. We further evaluate the use of synthetically generated RGB and depth pairs for training or fine-tuning depth estimators. Code will be available at here.
Zhen Zhu 0006, Yijun Li 0001, Weijie Lyu, Krishna Kumar Singh, Zhixin Shu, Sören Pirk, Derek Hoiem
WACV7
2023 ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation
abstract
State-of-the-art vision-language models (VLMs) still have limited performance in structural knowledge extraction, such as relations between objects.In this work, we present ViStruct, a training framework to learn VLMs for effective visual structural knowledge extraction.Two novel designs are incorporated.First, we propose to leverage the inherent structure of programming language to depict visual structural information.This approach enables explicit and consistent representation of visual structural information of multiple granularities, such as concepts, relations, and events, in a well-organized structured format.Second, we introduce curriculum-based learning for VLMs to progressively comprehend visual structures, from fundamental visual concepts to intricate event structures.Our intuition is that lower-level knowledge may contribute to complex visual structure understanding.Furthermore, we compile and release a collection of datasets tailored for visual structural knowledge extraction.We adopt a weakly-supervised approach to directly generate visual event structures from captions for ViStruct training, capitalizing on abundant image-caption pairs from the web.In experiments, we evaluate ViStruct on visual structure prediction tasks, demonstrating its effectiveness in improving the understanding of visual structures.The code is public at https://github.com/ Yangyi-Chen/vi-struct.
Yangyi Chen, Xingyao Wang 0002, Manling Li, Derek Hoiem, Heng Ji 0001
EMNLP4
2023 StyleGAN knows Normal, Depth, Albedo, and More
abstract
Intrinsic images, in the original sense, are image-like maps of scene properties like depth, normal, albedo, or shading. This paper demonstrates that StyleGAN can easily be induced to produce intrinsic images. The procedure is straightforward. We show that if StyleGAN produces $G({\bf w})$ from latent ${\bf w}$, then for each type of intrinsic image, there is a fixed offset ${\bf d}_c$ so that $G({\bf w}+{\bf d}_c)$ is that type of intrinsic image for $G({\bf w})$. Here ${\bf d}_c$ is {\em independent of ${\bf w}$}. The StyleGAN we used was pretrained by others, so this property is not some accident of our training regime. We show that there are image transformations StyleGAN will {\em not} produce in this fashion, so StyleGAN is not a generic image regression engine. It is conceptually exciting that an image generator should ``know'' and represent intrinsic images. There may also be practical advantages to using a generative model to produce intrinsic images. The intrinsic images obtained from StyleGAN compare well both qualitatively and quantitatively with those obtained by using SOTA image regression techniques; but StyleGAN's intrinsic images are robust to relighting effects, unlike SOTA methods.
Anand Bhattad, Daniel McKee, Derek Hoiem, David A. Forsyth
NeurIPS3
2022 Towards General Purpose Vision Systems: An End-to-End Task-Agnostic Vision-Language Architecture
abstract
Computer vision systems today are primarily N-purpose systems, designed and trained for a predefined set of tasks. Adapting such systems to new tasks is challenging and often requires nontrivial modifications to the network architecture (e.g. adding new output heads) or training process (e.g. adding new losses). To reduce the time and expertise required to develop new applications, we would like to create general purpose vision systems that can learn and perform a range of tasks without any modification to the architecture or learning process. In this paper, we propose GPV-1, a task-agnostic vision-language architecture that can learn and perform tasks that involve receiving an image and producing text and/or bounding boxes, including classification, localization, visual question answering, captioning, and more. We also propose evaluations of generality of architecture, skill-concept11For this work, we define concepts, skills and tasks as follows: Concepts - nouns (e.g. car, person, dog), Skills - operations that we wish to perform on the given inputs (e.g. classification, object detection, image captioning), Tasks - predefined combinations of a set of skills performed on a set of concepts (e.g. ImageNet classification task involves the skill of image classification across 1000 concepts). transfer, and learning efficiency that may informfuture work on general purpose vision. Our experiments indicate GPV-1 is effective at multiple tasks, reuses some concept knowledge across tasks, can perform the Referring Expressions task zero-shot, and further improves upon the zero-shot performance using a few training samples.
Tanmay Gupta, Amita Kamath, Aniruddha Kembhavi, Derek Hoiem
CVPR4
2022 Webly Supervised Concept Expansion for General Purpose Vision Models
Amita Kamath, Tanmay Gupta, Eric Kolve, Derek Hoiem, Aniruddha Kembhavi
ECCV (36)5
2022 Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
abstract
The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets.Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei 0003, Xudong Lin 0003, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001
NeurIPS10
2021 PatchMatch-RL: Deep MVS with Pixelwise Depth, Normal, and Visibility
abstract
Recent learning-based multi-view stereo (MVS) methods show excellent performance with dense cameras and small depth ranges. However, non-learning based approaches still outperform for scenes with large depth ranges and sparser wide-baseline views, in part due to their PatchMatch optimization over pixelwise estimates of depth, normals, and visibility. In this paper, we propose an end-to-end trainable PatchMatch-based MVS approach that combines advantages of trainable costs and regularizations with pixelwise estimates. To overcome the challenge of the non-differentiable PatchMatch optimization that involves iterative sampling and hard decisions, we use reinforcement learning to minimize expected photometric cost and maximize likelihood of ground truth depth and normals. We incorporate normal estimation by using dilated patch kernels and propose a recurrent cost regularization that applies beyond frontal plane-sweep algorithms to our pixelwise depth/normal estimates. We evaluate our method on widely used MVS benchmarks, ETH3D and Tanks and Temples (TnT). On ETH3D, our method outperforms other recent learning-based approaches and performs comparably on advanced TnT.
Jae Yong Lee 0006, Joseph DeGol, Chuhang Zou, Derek Hoiem
ICCV4
2021 Learning Curves for Analysis of Deep Networks
abstract
Learning curves model a classifier’s test error as a function of the number of training samples. Prior works show that learning curves can be used to select model parameters and extrapolate performance. We investigate how to use learning curves to evaluate design choices, such as pretraining, architecture, and data augmentation. We propose a method to robustly estimate learning curves, abstract their parameters into error and data-reliance, and evaluate the effectiveness of different parameterizations. Our experiments exemplify use of learning curves for analysis and yield several interesting observations.
Derek Hoiem, Tanmay Gupta, Zhizhong Li 0001, Michal Shlapentokh-Rothman
ICML1
2021 Task-Assisted Domain Adaptation with Anchor Tasks
Zhizhong Li 0001, Linjie Luo, Sergey Tulyakov, Qieyun Dai, Derek Hoiem
WACV5
2021 Manhattan Room Layout Reconstruction from a Single $360^{\circ }$ Image: A Comparative Study of State-of-the-Art Methods
Chuhang Zou, Jheng-Wei Su, Chihan Peng, Alex Colburn, Qi Shan, Peter Wonka, Hung-Kuo Chu, Derek Hoiem
Int. J. Comput. Vis.8
2021 Editorial: Introduction to the Special Section on CVPR2019 Best Papers
Gang Hua 0001, Derek Hoiem, Abhinav Gupta 0001, Zhuowen Tu
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Improving Structure from Motion with Reliable Resectioning
abstract
A common cause of failure in structure-from-motion (SfM) is misregistration of images due to visual patterns that occur in more than one scene location. Most work to solve this problem ignores image matches that are inconsistent according to the statistics of the tracks graph, but these methods often need to be tuned for each dataset and can lead to reduced completeness of normally good reconstructions when valid matches are removed. Our key idea is to address ambiguity directly in the reconstruction process by using only a subset of reliable matches to determine resectioning order and the initial pose. We also introduce a new measure of similarity that adjusts the influence of feature matches based on their track length. We show this improves reconstruction robustness for two state-of-the-art SfM algorithms on many diverse datasets.
Rajbir Kataria, Joseph DeGol, Derek Hoiem
3DV3
2020 Improving Confidence Estimates for Unfamiliar Examples
abstract
Intuitively, unfamiliarity should lead to lack of confidence. In reality, current algorithms often make highly confident yet wrong predictions when faced with relevant but unfamiliar examples. A classifier we trained to recognize gender is 12 times more likely to be wrong with a 99% confident prediction if presented with a subject from a different age group than those seen during training. In this paper, we compare and evaluate several methods to improve confidence estimates for unfamiliar and familiar samples. We propose a testing methodology of splitting unfamiliar and familiar samples by attribute (age, breed, subcategory) or sampling (similar datasets collected by different people at different times). We evaluate methods including confidence calibration, ensembles, distillation, and a Bayesian model and use several metrics to analyze label, likelihood, and calibration error. While all methods reduce over-confident errors, the ensemble of calibrated models performs best overall, and T-scaling performs best among the approaches with fastest inference.
Zhizhong Li 0001, Derek Hoiem
CVPR2
2020 Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion
abstract
We introduce DeepInversion, a new method for synthesizing images from the image distribution used to train a deep neural network. We ``invert'' a trained network (teacher) to synthesize class-conditional input images starting from random noise, without using any additional information about the training dataset. Keeping the teacher fixed, our method optimizes the input while regularizing the distribution of intermediate feature maps using information stored in the batch normalization layers of the teacher. Further, we improve the diversity of synthesized images using Adaptive DeepInversion, which maximizes the Jensen-Shannon divergence between the teacher and student network logits. The resulting synthesized images from networks trained on the CIFAR-10 and ImageNet datasets demonstrate high fidelity and degree of realism, and help enable a new breed of data-free applications - ones that do not require any real images or labeled data. We demonstrate the applicability of our proposed method to three tasks of immense practical importance - (i) data-free network pruning, (ii) data-free knowledge transfer, and (iii) data-free continual learning.
Hongxu Yin, Pavlo Molchanov 0001, José M. Álvarez 0004, Zhizhong Li 0001, Arun Mallya, Derek Hoiem, Niraj K. Jha, Jan Kautz
CVPR6
2020 Contrastive Learning for Weakly Supervised Phrase Grounding
Tanmay Gupta, Arash Vahdat, Gal Chechik, Xiaodong Yang 0001, Jan Kautz, Derek Hoiem
ECCV (3)6
2020 Silhouette Guided Point Cloud Reconstruction beyond Occlusion
abstract
One major challenge in 3D reconstruction is to infer the complete shape geometry from partial foreground occlusions. In this paper, we propose a method to reconstruct the complete 3D shape of an object from a single RGB image, with robustness to occlusion. Given the image and a silhouette of the visible region, our approach completes the silhouette of the occluded region and then generates a point cloud. We show improvements for reconstruction of non-occluded and partially occluded objects by providing the predicted complete silhouette as guidance. We also improve state-of-the-art for 3D shape prediction with a 2D reprojection loss from multiple synthetic views and a surface-based smoothing and refinement step. Experiments demonstrate the efficacy of our approach both quantitatively and qualitatively on synthetic and real scene datasets.
Chuhang Zou, Derek Hoiem
WACV2
2019 ViCo: Word Embeddings From Visual Co-Occurrences
abstract
We propose to learn word embeddings from visual co-occurrences. Two words co-occur visually if both words apply to the same image or image region. Specifically, we extract four types of visual co-occurrences between object and attribute words from large-scale, textually-annotated visual databases like VisualGenome and ImageNet. We then train a multi-task log-bilinear model that compactly encodes word “meanings” represented by each co-occurrence type into a single visual word-vector. Through unsupervised clustering, supervised partitioning, and a zero-shot-like generalization analysis we show that our word embeddings complement text-only embeddings like Glove by better representing similarities and differences between visual concepts that are difficult to obtain from text corpora alone. We further evaluate our embeddings on five downstream applications, four of which are vision-language tasks. Augmenting Glove with our embeddings yields gains on all tasks. We also find that random embeddings perform comparably to learned embeddings on all supervised vision-language tasks, contrary to conventional wisdom.
Tanmay Gupta, Alexander G. Schwing, Derek Hoiem
ICCV3
2019 No-Frills Human-Object Interaction Detection: Factorization, Layout Encodings, and Training Techniques
abstract
We show that for human-object interaction detection a relatively simple factorized model with appearance and layout encodings constructed from pre-trained object detectors outperforms more sophisticated approaches. Our model includes factors for detection scores, human and object appearance, and coarse (box-pair configuration) and optionally fine-grained layout (human pose). We also develop training techniques that improve learning efficiency by: (1) eliminating a train-inference mismatch; (2) rejecting easy negatives during mini-batch training; and (3) using a ratio of negatives to positives that is two orders of magnitude larger than existing approaches. We conduct a thorough ablation study to understand the importance of different factors and training techniques using the challenging HICO-Det dataset.
Tanmay Gupta, Alexander G. Schwing, Derek Hoiem
ICCV3
2019 Complete 3D Scene Parsing from an RGBD Image
Chuhang Zou, Zhizhong Li 0001, Derek Hoiem
Int. J. Comput. Vis.4
2018 FEATS: Synthetic Feature Tracks for Structure from Motion Evaluation
abstract
We present FEATS (Feature Extraction and Tracking Simulator), that synthesizes feature tracks using a camera trajectory and scene geometry (e.g. CAD, multi-view stereo). We introduce 2D feature and matching noise models that can be controlled using a few parameters. We also provide a new dataset of images and ground truth camera pose. We process this data (and a synthetic version) with several current SfM algorithms and show that the synthetic tracks are representative of the real tracks. We then show two practical uses of FEATS: (1) we generate hundreds of trajectories with varying noise and show that COLMAP is more robust to noise than OpenSfM and VisualSfM; and (2) we calculate 3D point error and show that accurate camera pose estimates do not guarantee accurate 3D maps.
Joseph DeGol, Jae Yong Lee 0006, Rajbir Kataria, Daniel Yuan, Timothy Bretl, Derek Hoiem
3DV6
2018 Pixels, Voxels, and Views: A Study of Shape Representations for Single View 3D Object Shape Prediction
abstract
The goal of this paper is to compare surface-based and volumetric 3D object shape representations, as well as viewer-centered and object-centered reference frames for single-view 3D shape prediction. We propose a new algorithm for predicting depth maps from multiple viewpoints, with a single depth or RGB image as input. By modifying the network and the way models are evaluated, we can directly compare the merits of voxels vs. surfaces and viewer-centered vs. object-centered for familiar vs. unfamiliar objects, as predicted from RGB or depth images. Among our findings, we show that surface-based methods outperform voxel representations for objects from novel classes and produce higher resolution outputs. We also find that using viewer-centered coordinates is advantageous for novel objects, while object-centered representations are better for more familiar objects. Interestingly, the coordinate frame significantly affects the shape representation learned, with object-centered placing more importance on implicitly recognizing the object category and viewer-centered producing shape representations with less dependence on category recognition.
Daeyun Shin, Charless C. Fowlkes, Derek Hoiem
CVPR3
2018 LayoutNet: Reconstructing the 3D Room Layout From a Single RGB Image
abstract
We propose an algorithm to predict room layout from a single image that generalizes across panoramas and perspective images, cuboid layouts and more general layouts (e.g. "L"-shape room). Our method operates directly on the panoramic image, rather than decomposing into perspective images as do recent works. Our network architecture is similar to that of RoomNet [15], but we show improvements due to aligning the image based on vanishing points, predicting multiple layout elements (corners, boundaries, size and translation), and fitting a constrained Manhattan layout to the resulting predictions. Our method compares well in speed and accuracy to other existing work on panoramas, achieves among the best accuracy for perspective images, and can handle both cuboid-shaped and more general Manhattan layouts.
Chuhang Zou, Alex Colburn, Qi Shan, Derek Hoiem
CVPR4
2018 Improved Structure from Motion Using Fiducial Marker Matching
Joseph DeGol, Timothy Bretl, Derek Hoiem
ECCV (3)3
2018 Imagine This! Scripts to Compositions to Videos
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, Aniruddha Kembhavi
ECCV (8)4
2018 Learning without Forgetting
abstract
When building a unified vision system or gradually adding new apabilities to a system, the usual assumption is that training data for all tasks is always available. However, as the number of tasks grows, storing and retraining on such data becomes infeasible. A new problem arises where we add new capabilities to a Convolutional Neural Network (CNN), but the training data for its existing capabilities are unavailable. We propose our Learning without Forgetting method, which uses only new task data to train the network while preserving the original capabilities. Our method performs favorably compared to commonly used feature extraction and fine-tuning adaption techniques and performs similarly to multitask learning that uses original task data we assume unavailable. A more surprising observation is that Learning without Forgetting may be able to replace fine-tuning with similar old and new task datasets for improved new task performance.
Zhizhong Li 0001, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 ChromaTag: A Colored Marker and Fast Detection Algorithm
abstract
Current fiducial marker detection algorithms rely on marker IDs for false positive rejection. Time is wasted on potential detections that will eventually be rejected as false positives. We introduce ChromaTag, a fiducial marker and detection algorithm designed to use opponent colors to limit and quickly reject initial false detections and grayscale for precise localization. Through experiments, we show that ChromaTag is significantly faster than current fiducial markers while achieving similar or better detection accuracy. We also show how tag size and viewing direction effect detection accuracy. Our contribution is significant because fiducial markers are often used in real-time applications (e.g. marker assisted robot navigation) where heavy computation is required by other parts of the system.
Joseph DeGol, Timothy Bretl, Derek Hoiem
ICCV3
2017 Aligned Image-Word Representations Improve Inductive Transfer Across Vision-Language Tasks
Tanmay Gupta, Kevin J. Shih, Saurabh Singh 0005, Derek Hoiem
ICCV4
2017 3D-PRNN: Generating Shape Primitives with Recurrent Neural Networks
abstract
The success of various applications including robotics, digital content creation, and visualization demand a structured and abstract representation of the 3D world from limited sensor data. Inspired by the nature of human perception of 3D shapes as a collection of simple parts, we explore such an abstract shape representation based on primitives. Given a single depth image of an object, we present 3DPRNN, a generative recurrent neural network that synthesizes multiple plausible shapes composed of a set of primitives. Our generative model encodes symmetry characteristics of common man-made objects, preserves long-range structural coherence, and describes objects of varying complexity with a compact representation. We also propose a method based on Gaussian Fields to generate a large scale dataset of primitive-based shape representations to train our network. We evaluate our approach on a wide range of examples and show that it outperforms nearest-neighbor based shape retrieval methods and is on-par with voxelbased generative models while using a significantly reduced parameter space.
Chuhang Zou, Ersin Yumer, Jimei Yang, Duygu Ceylan, Derek Hoiem
ICCV5
2016 Geometry-Informed Material Recognition
abstract
Our goal is to recognize material categories using images and geometry information. In many applications, such as construction management, coarse geometry information is available. We investigate how 3D geometry (surface normals, camera intrinsic and extrinsic parameters) can be used with 2D features (texture and color) to improve material classification. We introduce a new dataset, GeoMat, which is the first to provide both image and geometry data in the form of: (i) training and testing patches that were extracted at different scales and perspectives from real world examples of each material category, and (ii) a large scale construction site scene that includes 160 images and over 800,000 hand labeled 3D points. Our results show that using 2D and 3D features both jointly and independently to model materials improves classification accuracy across multiple scales and viewing directions for both material patches and images of a large scale construction site scene.
Joseph DeGol, Mani Golparvar Fard, Derek Hoiem
CVPR3
2016 Where to Look: Focus Regions for Visual Question Answering
abstract
We present a method that learns to answer visual questions by selecting image regions relevant to the text-based query. Our method maps textual queries and visual features from various regions into a shared space where they are compared for relevance with an inner product. Our method exhibits significant improvements in answering questions such as "what color," where it is necessary to evaluate a specific location, and "what room," where it selectively identifies informative image regions. Our model is tested on the recently released VQA [1] dataset, which features free-form human-annotated questions and answers.
Kevin J. Shih, Saurabh Singh 0005, Derek Hoiem
CVPR3
2016 Learning to Localize Little Landmarks
abstract
We interact everyday with tiny objects such as the door handle of a car or the light switch in a room. These little landmarks are barely visible and hard to localize in images. We describe a method to find such landmarks by finding a sequence of latent landmarks, each with a prediction model. Each latent landmark predicts the next in sequence, and the last localizes the target landmark. For example, to find the door handle of a car, our method learns to start with a latent landmark near the wheel, as it is globally distinctive, subsequent latent landmarks use the context from the earlier ones to get closer to the target. Our method is supervised solely by the location of the little landmark and displays strong performance on more difficult variants of established tasks and on two new tasks.
Saurabh Singh 0005, Derek Hoiem, David A. Forsyth
CVPR2
2016 Learning Without Forgetting
Zhizhong Li 0001, Derek Hoiem
ECCV (4)2
2016 Swapout: Learning an ensemble of deep architectures
abstract
We describe Swapout, a new stochastic training method, that outperforms ResNets of identical network structure yielding impressive results on CIFAR-10 and CIFAR-100. Swapout samples from a rich set of architectures including dropout, stochastic depth and residual architectures as special cases. When viewed as a regularization method swapout not only inhibits co-adaptation of units in a layer, similar to dropout, but also across network layers. We conjecture that swapout achieves strong regularization by implicitly tying the parameters across layers. When viewed as an ensemble training method, it samples a much richer set of architectures than existing methods such as dropout or stochastic depth. We propose a parameterization that reveals connections to exiting architectures and suggests a much richer set of architectures to be explored. We show that our formulation suggests an efficient training method and validate our conclusions on CIFAR-10 and CIFAR-100 matching state of the art accuracy. Remarkably, our 32 layer wider model performs similar to a 1001 layer ResNet model.
Saurabh Singh 0005, Derek Hoiem, David A. Forsyth
NIPS2
2015 Part Localization using Multi-Proposal Consensus for Fine-Grained Categorization
abstract
We present a simple deep learning framework to simultaneously predict keypoint locations and their respective visibilities and use those to achieve state-of-the-art performance for fine-grained classification. We show that by conditioning the predictions on object proposals with sufficient image support, our method can do well without complicated spatial reasoning. Instead, inference methods with robustness to outliers, yield state-of-the-art for keypoint localization. We demonstrate the effectiveness of our accurate keypoint localization and visibility prediction on the fine-grained bird recognition task with and without ground truth bird bounding boxes, and outperform existing state-of-the-art methods by over 2%.
Kevin J. Shih, Arun Mallya, Saurabh Singh 0005, Derek Hoiem
BMVC4
2015 Completing 3D object shape from one depth image
abstract
Our goal is to recover a complete 3D model from a depth image of an object. Existing approaches rely on user interaction or apply to a limited class of objects, such as chairs. We aim to fully automatically reconstruct a 3D model from any category. We take an exemplar-based approach: retrieve similar objects in a database of 3D models using view-based matching and transfer the symmetries and surfaces from retrieved models. We investigate completion of 3D models in three cases: novel view (model in database); novel model (models for other objects of the same category in database); and novel category (no models from the category in database).
Jason Rock, Tanmay Gupta, Justin Thorsen, JunYoung Gwak, Daeyun Shin, Derek Hoiem
CVPR6
2015 Learning a sequential search for landmarks
abstract
We propose a general method to find landmarks in images of objects using both appearance and spatial context. This method is applied without changes to two problems: parsing human body layouts, and finding landmarks in images of birds. Our method learns a sequential search for localizing landmarks, iteratively detecting new landmarks given the appearance and contextual information from the already detected ones. The choice of landmark to be added is opportunistic and depends on the image; for example, in one image a head-shoulder group might be expanded to a head-shoulder-hip group but in a different image to a head-shoulder-elbow group. The choice of initial landmark is similarly image dependent. Groups are scored using a learned function, which is used to expand them greedily. Our scoring function is learned from data labelled with landmarks but without any labeling of a detection order. Our method represents a novel spatial model for the kinematics of groups of landmarks, and displays strong performance on two different model problems.
Saurabh Singh 0005, Derek Hoiem, David A. Forsyth
CVPR2
2015 Family Member Identification from Photo Collections
abstract
Family photo collections often contain richer semantics than arbitrary images of people because families contain a handful of specific individuals who can be associated with certain social roles (e.g. father, mother, or child). As a result, family photo collections have unique challenges and opportunities for face recognition compared to random groups of photos containing people. We address the problem of unsupervised family member discovery: given a collection of family photos, we infer the size of the family, as well as the visual appearance and social role of each family member. As a result, we are able to recognize the same individual across many different photos. We propose an unsupervised EM-style joint inference algorithm with a probabilistic CRF that models identity and role assignments for all detected faces, along with associated pair wise relationships between them. Our experiments illustrate how joint inference of both identity and role (across all photos simultaneously) outperforms independent estimates of each. Joint inference also improves the ability to recognize the same individual across many different photos.
Qieyun Dai, Peter Carr 0001, Leonid Sigal, Derek Hoiem
WACV4
2015 Labeling Complete Surfaces in Scene Understanding
Derek Hoiem
Int. J. Comput. Vis.2
2015 Guest Editorial: Scene Understanding
Derek Hoiem, James Hays, Jianxiong Xiao, Aditya Khosla
Int. J. Comput. Vis.1
2015 Learning Discriminative Collections of Part Detectors for Object Recognition
abstract
We propose a method to learn a diverse collection of discriminative parts from object bounding box annotations. Part detectors can be trained and applied individually, which simplifies learning and extension to new features or categories. We apply the parts to object category detection, pooling part detections within bottom-up proposed regions and using a boosted classifier with proposed sigmoid weak learners for scoring. On PASCAL VOC2010, we evaluate the part detectors' ability to discriminate and localize annotated keypoints and their effectiveness in detecting object categories.
Kevin J. Shih, Ian Endres, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Category-Independent Object Proposals with Diverse Ranking
abstract
We propose a category-independent method to produce a bag of regions and rank them, such that top-ranked regions are likely to be good segmentations of different objects. Our key objectives are completeness and diversity: Every object should have at least one good proposed region, and a diverse set should be top-ranked. Our approach is to generate a set of segmentations by performing graph cuts based on a seed region and a learned affinity function. Then, the regions are ranked using structured learning based on various cues. Our experiments on the Berkeley Segmentation Data Set and Pascal VOC 2011 demonstrate our ability to find most objects within a small bag of proposed regions.
Ian Endres, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Learning Collections of Part Models for Object Recognition
abstract
We propose a method to learn a diverse collection of discriminative parts from object bounding box annotations. Part detectors can be trained and applied individually, which simplifies learning and extension to new features or categories. We apply the parts to object category detection, pooling part detections within bottom-up proposed regions and using a boosted classifier with proposed sigmoid weak learners for scoring. On PASCAL VOC 2010, we evaluate the part detectors' ability to discriminate and localize annotated key points. Our detection system is competitive with the best-existing systems, outperforming other HOG-based detectors on the more deformable categories.
Ian Endres, Kevin J. Shih, Johnston Jiaa, Derek Hoiem
CVPR4
2013 Boundary Cues for 3D Object Shape Recovery
abstract
Early work in computer vision considered a host of geometric cues for both shape reconstruction and recognition. However, since then, the vision community has focused heavily on shading cues for reconstruction, and moved towards data-driven approaches for recognition. In this paper, we reconsider these perhaps overlooked "boundary" cues (such as self occlusions and folds in a surface), as well as many other established constraints for shape reconstruction. In a variety of user studies and quantitative tasks, we evaluate how well these cues inform shape reconstruction (relative to each other) in terms of both shape quality and shape recognition. Our findings suggest many new directions for future research in shape reconstruction, such as automatic boundary cue detection and relaxing assumptions in shape from shading (e.g. orthographic projection, Lambertian surfaces).
Kevin Karsch, Zicheng Liao, Jason Rock, Jonathan T. Barron, Derek Hoiem
CVPR5
2013 Support Surface Prediction in Indoor Scenes
abstract
In this paper, we present an approach to predict the extent and height of supporting surfaces such as tables, chairs, and cabinet tops from a single RGBD image. We define support surfaces to be horizontal, planar surfaces that can physically support objects and humans. Given a RGBD image, our goal is to localize the height and full extent of such surfaces in 3D space. To achieve this, we created a labeling tool and annotated 1449 images with rich, complete 3D scene models in NYU dataset. We extract ground truth from the annotated dataset and developed a pipeline for predicting floor space, walls, the height and full extent of support surfaces. Finally we match the predicted extent with annotated scenes in training scenes and transfer the the support surface configuration from training scenes. We evaluate the proposed approach in our dataset and demonstrate its effectiveness in understanding scenes in 3D space.
Derek Hoiem
ICCV2
2013 Paired Regions for Shadow Detection and Removal
abstract
In this paper, we address the problem of shadow detection and removal from single images of natural scenes. Differently from traditional methods that explore pixel or edge information, we employ a region-based approach. In addition to considering individual regions separately, we predict relative illumination conditions between segmented regions from their appearances and perform pairwise classification based on such information. Classification results are used to build a graph of segments, and graph-cut is used to solve the labeling of shadow and nonshadow regions. Detection results are later refined by image matting, and the shadow-free image is recovered by relighting each pixel based on our lighting model. We evaluate our method on the shadow detection dataset in Zhu et al. In addition, we created a new dataset with shadow-free ground truth images, which provides a quantitative basis for evaluating shadow removal. We study the effectiveness of features for both unary and pairwise classification.
Qieyun Dai, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Improved Object Categorization and Detection Using Comparative Object Similarity
abstract
Due to the intrinsic long-tailed distribution of objects in the real world, we are unlikely to be able to train an object recognizer/detector with many visual examples for each category. We have to share visual knowledge between object categories to enable learning with few or no training examples. In this paper, we show that local object similarity information--statements that pairs of categories are similar or dissimilar--is a very useful cue to tie different categories to each other for effective knowledge transfer. The key insight: Given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. To exploit this category-dependent similarity regularization, we develop a regularized kernel machine algorithm to train kernel classifiers for categories with few or no training examples. We also adapt the state-of-the-art object detector to encode object similarity constraints. Our experiments on hundreds of categories from the Labelme dataset show that our regularized kernel classifiers can make significant improvement on object categorization. We also evaluate the improved object detector on the PASCAL VOC 2007 benchmark dataset.
Gang Wang 0012, David A. Forsyth, Derek Hoiem
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Learning to localize detected objects
abstract
In this paper, we propose an approach to accurately localize detected objects. The goal is to predict which features pertain to the object and define the object extent with segmentation or bounding box. Our initial detector is a slight modification of the DPM detector by Felzenszwalb et al., which often reduces confusion with background and other objects but does not cover the full object. We then describe and evaluate several color models and edge cues for local predictions, and we propose two approaches for localization: learned graph cut segmentation and structural bounding box prediction. Our experiments on the PASCAL VOC 2010 dataset show that our approach leads to accurate pixel assignment and large improvement in bounding box overlap, sometimes leading to large overall improvement in detection accuracy.
Qieyun Dai, Derek Hoiem
CVPR2
2012 A data driven method for feature transformation
abstract
Most image understanding algorithms begin with the extraction of information thought to be relevant to the particular task. This is commonly known as feature extraction and has, up to this date, been a largely manual process, where a reasonable method is chosen through validation on the experimented dataset. In this work we propose a data driven, local histogram based feature extraction method that reduces the manual intervention during the feature computation process and improves on the performance of widely used gradient histogram based features (e.g., HOG). We demonstrate favorable object detection results against HOG on the Inria Pedestrian[7], Pascal 2007[10] data.
Mert Dikmen, Derek Hoiem, Thomas S. Huang
CVPR2
2012 Learning shared body plans
abstract
We cast the problem of recognizing related categories as a unified learning and structured prediction problem with shared body plans. When provided with detailed annotations of objects and their parts, these body plans model objects in terms of shared parts and layouts, simultaneously capturing a variety of categories in varied poses. We can use these body plans to jointly train many detectors in a shared framework with structured learning, leading to significant gains for each supervised task. Using our model, we can provide detailed predictions of objects and their parts for both familiar and unfamiliar categories.
Ian Endres, Vivek Srikumar, Ming-Wei Chang, Derek Hoiem
CVPR4
2012 Recovering free space of indoor scenes from a single image
abstract
In this paper we consider the problem of recovering the free space of an indoor scene from its single image. We show that exploiting the box like geometric structure of furniture and constraints provided by the scene, allows us to recover the extent of major furniture objects in 3D. Our “boxy” detector localizes box shaped objects oriented parallel to the scene across different scales and object types, and thus blocks out the occupied space in the scene. To localize the objects more accurately in 3D we introduce a set of specially designed features that capture the floor contact points of the objects. Image based metrics are not very indicative of performance in 3D. We make the first attempt to evaluate single view based occupancy estimates for 3D errors and propose several task driven performance measures towards it. On our dataset of 592 indoor images marked with full 3D geometry of the scene, we show that: (a) our detector works well using image based metrics; (b) our refinement method produces significant improvements in localization in 3D; and (c) if one evaluates using 3D metrics, our method offers major improvements over other single view based scene geometry estimation methods.
Varsha Hedau, Derek Hoiem, David A. Forsyth
CVPR2
2012 Beyond the Line of Sight: Labeling the Underlying Surfaces
Derek Hoiem
ECCV (5)2
2012 Diagnosing Error in Object Detectors
Derek Hoiem, Yodsawalai Chodpathumwan, Qieyun Dai
ECCV (3)1
2012 Indoor Segmentation and Support Inference from RGBD Images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, Rob Fergus
ECCV (5)2
2012 Learning Image Similarity from Flickr Groups Using Fast Kernel Machines
abstract
Measuring image similarity is a central topic in computer vision. In this paper, we propose to measure image similarity by learning from the online Flickr image groups. We do so by: Choosing 103 Flickr groups, building a one-versus-all multiclass classifier to classify test images into a group, taking the set of responses of the classifiers as features, calculating the distance between feature vectors to measure image similarity. Experimental results on the Corel dataset and the PASCAL VOC 2007 dataset show that our approach performs better on image matching, retrieval, and classification than using conventional visual features. To build our similarity measure, we need one-versus-all classifiers that are accurate and can be trained quickly on very large quantities of data. We adopt an SVM classifier with a histogram intersection kernel. We describe a novel fast training algorithm for this classifier: the Stochastic Intersection Kernel MAchine (SIKMA) training algorithm. This method can produce a kernel classifier that is more accurate than a linear classifier on tens of thousands of examples in minutes.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Single-image shadow detection and removal using paired regions
abstract
In this paper, we address the problem of shadow detection and removal from single images of natural scenes. Different from traditional methods that explore pixel or edge information, we employ a region based approach. In addition to considering individual regions separately, we predict relative illumination conditions between segmented regions from their appearances and perform pairwise classification based on such information. Classification results are used to build a graph of segments, and graph-cut is used to solve the labeling of shadow and non-shadow regions. Detection results are later refined by image matting, and the shadow free image is recovered by relighting each pixel based on our lighting model. We evaluate our method on the shadow detection dataset. In addition, we created a new dataset with shadow-free ground truth images, which provides a quantitative basis for evaluating shadow removal.
Qieyun Dai, Derek Hoiem
CVPR3
2011 Recovering Occlusion Boundaries from an Image
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.1
2011 Rendering synthetic objects into legacy photographs
Kevin Karsch, Varsha Hedau, David A. Forsyth, Derek Hoiem
ACM Trans. Graph.4
2010 Attribute-centric recognition for cross-category generalization
abstract
We propose an approach to find and describe objects within broad domains. We introduce a new dataset that provides annotation for sharing models of appearance and correlation across categories. We use it to learn part and category detectors. These serve as the visual basis for an integrated model of objects. We describe objects by the spatial arrangement of their attributes and the interactions between them. Using this model, our system can find animals and vehicles that it has not seen and infer attributes, such as function and pose. Our experiments demonstrate that we can more reliably locate and describe both familiar and unfamiliar objects, compared to a baseline that relies purely on basic category detectors.
Ali Farhadi, Ian Endres, Derek Hoiem
CVPR3
2010 Comparative object similarity for improved recognition with few or no examples
abstract
Learning models for recognizing objects with few or no training examples is important, due to the intrinsic long-tailed distribution of objects in the real world. In this paper, we propose an approach to use comparative object similarity. The key insight is that: given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. We develop a regularized kernel machine algorithm to use this category dependent similarity regularization. Our experiments on hundreds of categories show that our method can make significant improvement, especially for categories with no examples.
Gang Wang 0012, David A. Forsyth, Derek Hoiem
CVPR3
2010 Category Independent Object Proposals
Ian Endres, Derek Hoiem
ECCV (5)2
2010 Thinking Inside the Box: Using Appearance Models and Context Based on Room Geometry
Varsha Hedau, Derek Hoiem, David A. Forsyth
ECCV (6)2
2010 It's All About the Data
abstract
Modern computer vision research consumes labelled data in quantity, and building datasets has become an important activity. The Internet has become a tremendous resource for computer vision researchers. By seeing the Internet as a vast, slightly disorganized collection of visual data, we can build datasets. The key point is that visual data are surrounded by contextual information like text and HTML tags, which is a strong, if noisy, cue to what the visual data means. In a series of case studies, we illustrate how useful this contextual information is. It can be used to build a large and challenging labelled face dataset with no manual intervention. With very small amounts of manual labor, contextual data can be used together with image data to identify pictures of animals. In fact, these contextual data are sufficiently reliable that a very large pool of noisily tagged images can be used as a resource to build image features, which reliably improve on conventional visual features. By seeing the Internet as a marketplace that can connect sellers of annotation services to researchers, we can obtain accurately annotated datasets quickly and cheaply. We describe methods to prepare data, check quality, and set prices for work for this annotation process. The problems posed by attempting to collect very big research datasets are fertile for researchers because collecting datasets requires us to focus on two important questions: What makes a good picture? What is the meaning of a picture?
Tamara L. Berg, Alexander Sorokin, Gang Wang 0012, David A. Forsyth, Derek Hoiem, Ian Endres, Ali Farhadi
Proc. IEEE5
2009 An empirical study of context in object detection
abstract
This paper presents an empirical evaluation of the role of context in a contemporary, challenging object detection task - the PASCAL VOC 2008. Previous experiments with context have mostly been done on home-grown datasets, often with non-standard baselines, making it difficult to isolate the contribution of contextual information. In this work, we present our analysis on a standard dataset, using top-performing local appearance detectors as baseline. We evaluate several different sources of context and ways to utilize it. While we employ many contextual cues that have been used before, we also propose a few novel ones including the use of geographic context and a new approach for using object spatial support.
Santosh Kumar Divvala, Derek Hoiem, James Hays, Alexei A. Efros, Martial Hebert
CVPR2
2009 Describing objects by their attributes
abstract
We propose to shift the goal of recognition from naming to describing. Doing so allows us not only to name familiar objects, but also: to report unusual aspects of a familiar object (“spotty dog”, not just “dog”); to say something about unfamiliar objects (“hairy and four-legged”, not just “unknown”); and to learn how to recognize new objects with few or no visual examples. Rather than focusing on identity assignment, we make inferring attributes the core problem of recognition. These attributes can be semantic (“spotty”) or discriminative (“dogs have it but sheep do not”). Learning attributes presents a major new challenge: generalization across object categories, not just across instances within a category. In this paper, we also introduce a novel feature selection method for learning attributes that generalize well across categories. We support our claims by thorough evaluation that provides insights into the limitations of the standard recognition paradigm of naming and demonstrates the new abilities provided by our attribute-based framework.
Ali Farhadi, Ian Endres, Derek Hoiem, David A. Forsyth
CVPR3
2009 Building text features for object image classification
abstract
We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text feature of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well, consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
CVPR2
2009 Recovering the spatial layout of cluttered rooms
abstract
In this paper, we consider the problem of recovering the spatial layout of indoor scenes from monocular images. The presence of clutter is a major problem for existing single-view 3D reconstruction algorithms, most of which rely on finding the ground-wall boundary. In most rooms, this boundary is partially or entirely occluded. We gain robustness to clutter by modeling the global room space with a parameteric 3D “box” and by iteratively localizing clutter and refitting the box. To fit the box, we introduce a structured learning algorithm that chooses the set of parameters to minimize error, based on global perspective cues. On a dataset of 308 images, we demonstrate the ability of our algorithm to recover spatial layout in cluttered rooms and show several examples of estimated free space.
Varsha Hedau, Derek Hoiem, David A. Forsyth
ICCV2
2009 Learning image similarity from Flickr groups using Stochastic Intersection Kernel MAchines
abstract
Measuring image similarity is a central topic in computer vision. In this paper, we learn similarity from Flickr groups and use it to organize photos. Two images are similar if they are likely to belong to the same Flickr groups. Our approach is enabled by a fast Stochastic Intersection Kernel MAchine (SIKMA) training algorithm, which we propose. This proposed training method will be useful for many vision problems, as it can produce a classifier that is more accurate than a linear classifier, trained on tens of thousands of examples in two minutes. The experimental results show our approach performs better on image matching, retrieval, and classification than using conventional visual features.
Gang Wang 0012, Derek Hoiem, David A. Forsyth
ICCV2
2008 Closing the loop in scene interpretation
abstract
Image understanding involves analyzing many different aspects of the scene. In this paper, we are concerned with how these tasks can be combined in a way that improves the performance of each of them. Inspired by Barrow and Tenenbaum, we present a flexible framework for interfacing scene analysis processes using intrinsic images. Each intrinsic image is a registered map describing one characteristic of the scene. We apply this framework to develop an integrated 3D scene understanding system with estimates of surface orientations, occlusion boundaries, objects, camera viewpoint, and relative depth. Our experiments on a set of 300 outdoor images demonstrate that these tasks reinforce each other, and we illustrate a coherent scene understanding with automatically reconstructed 3D models.
Derek Hoiem, Alexei A. Efros, Martial Hebert
CVPR1
2008 Learning CRFs Using Graph Cuts
Martin Szummer, Pushmeet Kohli, Derek Hoiem
ECCV (2)3
2008 Putting Objects in Perspective
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.1
2007 3D LayoutCRF for Multi-View Object Class Recognition and Segmentation
abstract
We introduce an approach to accurately detect and segment partially occluded objects in various viewpoints and scales. Our main contribution is a novel framework for combining object-level descriptions (such as position, shape, and color) with pixel-level appearance, boundary, and occlusion reasoning. In training, we exploit a rough 3D object model to learn physically localized part appearances. To find and segment objects in an image, we generate proposals based on the appearance and layout of local parts. The proposals are then refined after incorporating object-level information, and overlapping objects compete for pixels to produce a final description and segmentation of objects in the scene. A further contribution is a novel instance penalty, which is handled very efficiently during inference. We experimentally validate our approach on the challenging PASCAL'06 car database.
Derek Hoiem, Carsten Rother, John M. Winn
CVPR1
2007 Recovering Occlusion Boundaries from a Single Image
abstract
Occlusion reasoning, necessary for tasks such as navigation and object search, is an important aspect of everyday life and a fundamental problem in computer vision. We believe that the amazing ability of humans to reason about occlusions from one image is based on an intrinsically 3D interpretation. In this paper, our goal is to recover the occlusion boundaries and depth ordering of free-standing structures in the scene. Our approach is to learn to identify and label occlusion boundaries using the traditional edge and region cues together with 3D surface and depth cues. Since some of these cues require good spatial support (i.e., a segmentation), we gradually create larger regions and use them to improve inference over the boundaries. Our experiments demonstrate the power of a scene-based approach to occlusion reasoning.
Derek Hoiem, Andrew N. Stein, Alexei A. Efros, Martial Hebert
ICCV1
2007 Learning to Find Object Boundaries Using Motion Cues
abstract
While great strides have been made in detecting and localizing specific objects in natural images, the bottom-up segmentation of unknown, generic objects remains a difficult challenge. We believe that occlusion can provide a strong cue for object segmentation and "pop-out", but detecting an object's occlusion boundaries using appearance alone is a difficult problem in itself. If the camera or the scene is moving, however, that motion provides an additional powerful indicator of occlusion. Thus, we use standard appearance cues (e.g. brightness/color gradient) in addition to motion cues that capture subtle differences in the relative surface motion (i.e. parallax) on either side of an occlusion boundary. We describe a learned local classifier and global inference approach which provide a frame-work for combining and reasoning about these appearance and motion cues to estimate which region boundaries of an initial over-segmentation correspond to object/occlusion boundaries in the scene. Through results on a dataset which contains short videos with labeled boundaries, we demonstrate the effectiveness of motion cues for this task.
Andrew N. Stein, Derek Hoiem, Martial Hebert
ICCV2
2007 Recovering Surface Layout from an Image
Derek Hoiem, Alexei A. Efros, Martial Hebert
Int. J. Comput. Vis.1
2007 Photo clip art
abstract
We present a system for inserting new objects into existing photographs by querying a vast image-based object library, pre-computed using a publicly available Internet object database. The central goal is to shield the user from all of the arduous tasks typically involved in image compositing. The user is only asked to do two simple things: 1) pick a 3D location in the scene to place a new object; 2) select an object to insert using a hierarchical menu. We pose the problem of object insertion as a data-driven, 3D-based, context-sensitive object retrieval task. Instead of trying to manipulate the object to change its orientation, color distribution, etc. to fit the new image, we simply retrieve an object of a specified class that has all the required properties (camera pose, lighting, resolution, etc) from our large object library. We present new automatic algorithms for improving object segmentation and blending, estimating true 3D object size and orientation, and estimating scene lighting conditions. We also present an intuitive user interface that makes object insertion fast and simple even for the artistically challenged.
Jean-François Lalonde, Derek Hoiem, Alexei A. Efros, Carsten Rother, John M. Winn, Antonio Criminisi
ACM Trans. Graph.2
2006 Putting Objects in Perspective
abstract
Image understanding requires not only individually estimating elements of the visual world but also capturing the interplay among them. In this paper, we provide a framework for placing local object detection in the context of the overall 3D scene by modeling the interdependence of objects, surface orientations, and camera viewpoint. Most object detection methods consider all scales and locations in the image as equally likely. We show that with probabilistic estimates of 3D geometry, both in terms of surfaces and world coordinates, we can put objects into perspective and model the scale and location variance in the image. Our approach reflects the cyclical nature of the problem by allowing probabilistic object hypotheses to refine geometry and vice-versa. Our framework allows painless substitution of almost any object detector and is easily extended to include other aspects of image understanding. Our results confirm the benefits of our integrated approach.
Derek Hoiem, Alexei A. Efros, Martial Hebert
CVPR (2)1
2006 Opportunistic Use of Vision to Push Back the Path-Planning Horizon
abstract
Mobile robots need maps or other forms of geometric information about the environment to navigate. The mobility sensors (LADAR, stereo, etc.) on these robotic vehicles can however populate these maps only up to a distance of a few tens of meters. A navigation system has no knowledge about the world beyond this sensing horizon. As a result, path planners that rely only on this knowledge are unable to anticipate obstacles sufficiently early and have no choice but to resort to an inefficient local obstacle avoidance behavior. However, recent developments in the computer vision community allows us to collect geometric information about the environment far beyond this sensing horizon. The coarse 3D geometric estimation that can be recovered is derived from an appearance-based model. That uses a multiple-hypothesis framework to robustly estimate scene structure from a single image and estimating confidences for each geometric label. This 3D geometric estimation is used with a previously presented navigation strategy that reasons about sensor constraints and plans for measurements while navigating towards the goal. The validity of the sensing method and navigation strategy is supported by results from simulations as well as field experiments with a real robotic platform. These results also show that significant reduction in path length can be achieved by using this framework
Bart C. Nabbe, Derek Hoiem, Alexei A. Efros, Martial Hebert
IROS2
2005 Computer Vision for Music Identification
abstract
We describe how certain tasks in the audio domain can be effectively addressed using computer vision approaches. This paper focuses on the problem of music identification, where the goal is to reliably identify a song given a few seconds of noisy audio. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem. By employing pairwise boosting on a large set of Viola-Jones features, our system learns compact, discriminative, local descriptors that are amenable to efficient indexing. During the query phase, we retrieve the set of song snippets that locally match the noisy sample and employ geometric verification in conjunction with an EM-based "occlusion" model to identify the song that is most consistent with the observed signal. We have implemented our algorithm in a practical system that can quickly and accurately recognize music from short audio samples in the presence of distortions such as poor recording quality and significant ambient noise. Our experiments demonstrate that this approach significantly outperforms the current state-of-the-art in content-based music identification.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (1)2
2005 Computer Vision for Music Identification: Video Demonstration
abstract
This paper describes a demonstration video for our music identification system. The goal of music identification is to reliably recognize a song from a small sample of noisy audio. This problem is challenging because the recording is often corrupted by noise and because the audio sample will only match a small portion of the target song. Additionally, a practical music identification system should scale (in both accuracy and speed) to databases containing hundreds of thousands of songs. Recently, the music identification problem has attracted considerable attention. However, the task remains unsolved, particularly for noisy real-world queries. We cast music identification into an equivalent sub-image retrieval framework: identify the portion of a spectrogram image from the database that best matches a given query snippet. Our approach treats the spectrogram of each music clip as a 2D image and transforms music identification into a corrupted sub-image retrieval problem.
Yan Ke, Derek Hoiem, Rahul Sukthankar
CVPR (2)2
2005 SOLAR: sound object localization and retrieval in complex audio environments
abstract
The ability to identify sounds in complex audio environments is highly useful for multimedia retrieval, security, and many mobile robotic applications, but very little work has been done in this area. We present the SOLAR system, a system capable of finding sound objects, such as dog barks or car horns, in complex audio data extracted from movies. SOLAR avoids the need for segmentation by scanning over the audio data in fixed increments and classifying each short audio window separately. SOLAR employs boosted decision tree classifiers to select suitable features for modeling each sound object and to discriminate between the object of interest and all other sounds. We demonstrate the effectiveness of our approach with experiments on thirteen sound object classes trained using only tens of positive examples and tested on hours of audio data extracted from popular movies.
Derek Hoiem, Yan Ke, Rahul Sukthankar
ICASSP (5)1
2005 Geometric Context from a Single Image
abstract
Many computer vision algorithms limit their performance by ignoring the underlying 3D geometric structure in the image. We show that we can estimate the coarse geometric properties of a scene by learning appearance-based models of geometric classes, even in cluttered natural scenes. Geometric classes describe the 3D orientation of an image region with respect to the camera. We provide a multiple-hypothesis framework for robustly estimating scene structure from a single image and obtaining confidences for each geometric label. These confidences can then be used to improve the performance of many other applications. We provide a thorough quantitative evaluation of our algorithm on a set of outdoor images and demonstrate its usefulness in two applications: object detection and automatic single-view reconstruction.
Derek Hoiem, Alexei A. Efros, Martial Hebert
ICCV1
2005 Automatic photo pop-up
abstract
This paper presents a fully automatic method for creating a 3D model from a single photograph. The model is made up of several texture-mapped planar billboards and has the complexity of a typical children's pop-up book illustration. Our main insight is that instead of attempting to recover precise geometry, we statistically modelgeometric classesdefined by their orientations in the scene. Our algorithm labels regions of the input image into coarse categories: "ground", "sky", and "vertical". These labels are then used to "cut and fold" the image into a pop-up model using a set of simple assumptions. Because of the inherent ambiguity of the problem and the statistical nature of the approach, the algorithm is not expected to work on every image. However. it performs surprisingly well for a wide range of scenes taken from a typical person's photo album.
Derek Hoiem, Alexei A. Efros, Martial Hebert
ACM Trans. Graph.1
2004 Object-Based Image Retrieval Using the Statistical Structure of Images
Derek Hoiem, Rahul Sukthankar, Henry Schneiderman, Larry Huston
CVPR (2)1
2004 SnapFind: brute force interactive image retrieval
abstract
SnapFind is an image retrieval system that enables efficient interactive search of large data sets by exploiting active disk technology. In contrast to earlier approaches, where data is typically pre-indexed for efficient retrieval according to a fixed scheme, SnapFind provides users with the flexibility to search non-indexed data in a brute force manner. The query is translated into a customized searchlet that is executed in parallel by processors near the storage devices. This enables the majority of irrelevant images to be discarded where they are stored. Partial results are displayed during search execution allowing users to interactively refine the query without waiting for search termination. This paper argues that algorithms with user-adjustable parameters are preferable to black-box image retrieval techniques.
Larry Huston, Rahul Sukthankar, Derek Hoiem
ICIG3
1994 Designing and using integrated data collection and analysis tools: challenges and considerations
abstract
This paper describes the design and evolution of mi integrated set of computer-aided usability engineering (CAUSE) tools for data collection and analysis. The tools were designed to collect and analyse observational, video, and system event data in both the usability laboratory and in the field. Three generations of tools are described and the problems with each generation are discussed. Solutions to the problems arc presented, where available. Conclusions about the strengths and weaknesses of particular types of data, CAUSE tool design, and the importance of multiple data sources are drawn. An agenda for future work is also outlined.
Derek Hoiem, Kent D. Sullivan
Behav. Inf. Technol.1