VLDB 2026 Research / reviewers in the wild / expert
Gregory Shakhnarovich
dblp:17/1926 · also Greg Shakhnarovich
· DBLP profile ↗
73ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-4700-9398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastMap: Revisiting Structure from Motion Through First-Order OptimizationabstractWe propose FastMap, a new global structure from motion method focused on speed and simplicity. Previous methods like COLMAP and GLOMAP are able to estimate high-precision camera poses, but suffer from poor scalability when the number of matched keypoint pairs becomes large, mainly due to the time-consuming process of secondorder Gauss-Newton optimization. Instead, we design our method solely based on first-order optimizers. To obtain maximal speedup, we identify and eliminate two key performance bottlenecks: computational complexity and the kernel implementation of each optimization step. Through extensive experiments, we show that FastMap is up to 10 × faster than COLMAP and GLOMAP with GPU acceleration and achieves comparable pose accuracy. Project webpage: https://jiahao.ai/fastmap. Muhammad Zubair Irshad, Igor Vasiljevic, Matthew R. Walter, Vitor Campagnolo Guizilini, Gregory Shakhnarovich |
3DV | 7 |
| 2026 | Cross-Modal Taxonomic Generalization in (Vision-) Language ModelsabstractTianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyang Xu 0002, Marcelo Sandoval-Castañeda, Karen Livescu, Gregory Shakhnarovich, Kanishka Misra |
ACL (1) | 4 |
| 2025 | SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster PredictionabstractSign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks.Pretraining methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent (frame or video segment) representations, which ignore the effects of relationships across time in sign language.We introduce SHuBERT (Sign Hidden-Unit BERT), a self-supervised contextual representation model learned from approximately 1,000 hours of American Sign Language video.SHu-BERT adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams.SHuBERT achieves state-of-theart performance across multiple tasks including sign language translation, isolated sign language recognition, and fingerspelling detection. Shester Gueuwou, Xiaodan Du 0001, Gregory Shakhnarovich, Karen Livescu, Alexander H. Liu |
ACL (1) | 3 |
| 2025 | Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric DiffusionabstractCurrent methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation. Vitor Campagnolo Guizilini, Muhammad Zubair Irshad, Dian Chen 0005, Gregory Shakhnarovich, Rares Ambrus |
CVPR | 4 |
| 2025 | SPLART: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian SplattingabstractReconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability limitations (requiring 3D supervision or costly annotations), robustness issues (being susceptible to local optima), and rendering shortcomings (lacking speed or photorealism). We introduce SplArt, a self-supervised, category-agnostic framework that leverages 3D Gaussian Splatting (3DGS) to reconstruct articulated objects and infer kinematics from two sets of posed RGB images captured at different articulation states, enabling real-time photorealistic rendering for novel viewpoints and articulations. SplArt augments 3DGS with a differentiable mobility parameter per Gaussian, achieving refined part segmentation. A multi-stage optimization strategy is employed to progressively handle reconstruction, part segmentation, and articulation estimation, significantly enhancing robustness and accuracy. SplArt exploits geometric self-supervision, effectively addressing challenging scenarios without requiring 3D annotations or category-specific priors. Evaluations on established and newly proposed benchmarks, along with applications to real-world scenarios using a handheld RGB camera, demonstrate SplArt's state-of-the-art performance and real-world practicality. Code is publicly available at https://github.com/ripl/splart. Shengjie Lin, Jiading Fang, Muhammad Zubair Irshad, Vitor Campagnolo Guizilini, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter |
ICCV | 6 |
| 2025 | OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real WorldabstractWe would like to estimate the pose and full shape of an object from a single observation, without assuming known 3D model or category. In this work, we propose OmniShape, the first method of its kind to enable probabilistic pose and shape estimation. OmniShape is based on the key insight that shape completion can be decoupled into two multi-modal distributions: one capturing how measurements project into a normalized object reference frame defined by the dataset and the other modelling a prior over object geometries represented as triplanar neural fields. By training separate conditional diffusion models for these two distributions, we enable sampling multiple hypotheses from the joint pose and shape distri-bution. OmniShape demonstrates compelling performance on challenging real world datasets. Project website: https://tri-ml.glthub.io/omnishape. Katherine Liu, Sergey Zakharov, Dian Chen 0005, Takuya Ikeda, Gregory Shakhnarovich, Adrien Gaidon, Rares Ambrus |
ICRA | 5 |
| 2024 | Alpha Invariance: On Inverse Scaling Between Distance and Volume Density in Neural Radiance FieldsabstractScale-ambiguity in 3D scene dimensions leads to magnitude-ambiguity of volumetric densities in neural radiance fields, i.e., the densities double when scene size is halved, and vice versa. We call this property alpha invariance. For NeRFs to better maintain alpha invariance, we recommend 1) parameterizing both distance and volume densities in log space, and 2) a discretization-agnostic initialization strategy to guarantee high ray transmittance. We revisit a few popular radiance field models and find that these systems use various heuristics to deal with issues arising from scene scaling. We test their behaviors and show our recipe to be more robust. Visit our project page at https://pals.ttic.edu/p/alpha-invariance. Joshua Ahn, Raymond A. Yeh, Gregory Shakhnarovich |
CVPR | 4 |
| 2024 | Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelabstractText-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low-quality results due to the scarcity of 3D training data. In this paper, we propose Instant3D, a novel method that generates high-quality and diverse 3D assets from text prompts in a feed-forward manner. We adopt a two-stage paradigm, which first generates a sparse set of four structured and consistent views from text in one shot with a fine-tuned 2D text-to-image diffusion model, and then directly regresses the NeRF from the generated images with a novel transformer-based sparse-view reconstructor. Through extensive experiments, we demonstrate that our method can generate diverse 3D assets of high visual quality within 20 seconds, which is two orders of magnitude faster than previous optimization-based methods that can take 1 to 10 hours. Our project webpage is: https://jiahao.ai/instant3d/. Hao Tan 0002, Kai Zhang 0045, Zexiang Xu, Fujun Luan, Yinghao Xu 0001, Yicong Hong, Kalyan Sunkavalli, Gregory Shakhnarovich, Sai Bi |
ICLR | 9 |
| 2024 | HyperFields: Towards Zero-Shot Generation of NeRFs from TextabstractWe introduce HyperFields, a method for generating text-conditioned Neural Radiance Fields (NeRFs) with a single forward pass and (optionally) some fine-tuning. Key to our approach are: (i) a dynamic hypernetwork, which learns a smooth mapping from text token embeddings to the space of NeRFs; (ii) NeRF distillation training, which distills scenes encoded in individual NeRFs into one dynamic hypernetwork. These techniques enable a single network to fit over a hundred unique scenes. We further demonstrate that HyperFields learns a more general map between text and NeRFs, and consequently is capable of predicting novel in-distribution and out-of-distribution scenes — either zero-shot or with a few finetuning steps. Finetuning HyperFields benefits from accelerated convergence thanks to the learned general map, and is capable of synthesizing novel scenes 5 to 10 times faster than existing neural optimization-based methods. Our ablation experiments show that both the dynamic architecture and NeRF distillation are critical to the expressivity of HyperFields. Sudarshan Babu, Richard Liu, Avery Zhou, Michael Maire, Gregory Shakhnarovich, Rana Hanocka |
ICML | 5 |
| 2024 | Transcrib3D: 3D Referring Expression Resolution through Large Language ModelsabstractIf robots are to work effectively alongside people, they must be able to interpret natural language references to objects in their 3D environment. Understanding 3D referring expressions is challenging—it requires the ability to both parse the 3D structure of the scene and correctly ground free-form language in the presence of distraction and clutter. We introduce Transcrib3D, an approach that brings together 3D detection methods and the emergent reasoning capabilities of large language models (LLMs). Transcrib3D uses text as the unifying medium, which allows us to sidestep the need to learn shared representations connecting multi-modal inputs, which would require massive amounts of annotated 3D data. As a demonstration of its effectiveness, Transcrib3D achieves state-of-the-art results on 3D reference resolution benchmarks, with a great leap in performance from previous multi-modality baselines. To improve upon zero-shot performance and facilitate local deployment on edge computers and robots, we propose self-correction for fine-tuning that trains smaller models, resulting in performance close to that of large models. We show that our method enables a real robot to perform pick-and-place tasks given queries that contain challenging referring expressions. Code will be available at https://ripl.github.io/Transcrib3D. Jiading Fang, Xiangshan Tan, Shengjie Lin, Igor Vasiljevic, Vitor Campagnolo Guizilini, Hongyuan Mei, Rares Ambrus, Gregory Shakhnarovich, Matthew R. Walter |
IROS | 8 |
| 2023 | Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D GenerationabstractA diffusion model learns to predict a vector field of gradients. We propose to apply chain rule on the learned gradients, and back-propagate the score of a diffusion model through the Jacobian of a differentiable renderer, which we instantiate to be a voxel radiance field. This setup aggregates 2D scores at multiple camera viewpoints into a 3D score, and re-purposes a pretrained 2D model for 3D data generation. We identify a technical challenge of distribution mismatch that arises in this application, and propose a novel estimation mechanism to resolve it. We run our algorithm on several off-the-shelf diffusion image generative models, including the recently released Stable Diffusion trained on the large-scale LAION 5B dataset. Xiaodan Du 0001, Raymond A. Yeh, Gregory Shakhnarovich |
CVPR | 5 |
| 2022 | Searching for fingerspelled content in American Sign LanguageabstractNatural language processing for sign language video-including tasks like recognition, translation, and search-is crucial for making artificial intelligence technologies accessible to deaf individuals, and is gaining research interest in recent years.In this paper, we address the problem of searching for fingerspelled keywords or key phrases in raw sign language videos.This is an important task since significant content in sign language is often conveyed via fingerspelling, and to our knowledge the task has not been studied before.We propose an end-to-end model for this task, FSS-Net, that jointly detects fingerspelling and matches it to a text sequence.Our experiments, done on a large public dataset of ASL fingerspelling in the wild, show the importance of fingerspelling detection as a component of a search and retrieval model.Our model significantly outperforms baseline methods adapted from prior work on related tasks. Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
ACL (1) | 3 |
| 2022 | Depth Field Networks For Generalizable Multi-view Scene Representation
Vitor Campagnolo Guizilini, Igor Vasiljevic, Jiading Fang, Rare Ambru, Gregory Shakhnarovich, Matthew R. Walter, Adrien Gaidon |
ECCV (32) | 5 |
| 2022 | Open-Domain Sign Language Translation Learned from Online VideoabstractExisting work on sign language translationthat is, translation from sign language videos into sentences in a written language-has focused mainly on (1) data collected in a controlled environment or (2) data in a specific domain, which limits the applicability to realworld settings.In this paper, we introduce Ope-nASL, a large-scale American Sign Language (ASL) -English dataset collected from online video sites (e.g., YouTube).OpenASL contains 288 hours of ASL videos in multiple domains from over 200 signers and is the largest publicly available ASL translation dataset to date.To tackle the challenges of sign language translation in realistic settings and without glosses, we propose a set of techniques including sign search as a pretext task for pre-training and fusion of mouthing and handshape features.The proposed techniques produce consistent and large improvements in translation quality, over baseline models based on prior work. 1 Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
EMNLP | 3 |
| 2022 | Self-Supervised Camera Self-Calibration from VideoabstractCamera calibration is integral to robotics and computer vision algorithms that seek to infer geometric properties of the scene from visual input streams. In practice, calibration is a laborious procedure requiring specialized data collection and careful tuning. This process must be repeated whenever the parameters of the camera change, which can be a frequent occurrence for mobile robots and autonomous vehicles. In contrast, self-supervised depth and ego-motion estimation approaches can bypass explicit calibration by in-ferring per-frame projection models that optimize a view-synthesis objective. In this paper, we extend this approach to explicitly calibrate a wide range of cameras from raw videos in the wild. We propose a learning algorithm to regress per-sequence calibration parameters using an efficient family of general camera models. Our procedure achieves self-calibration results with sub-pixel reprojection error, outperforming other learning-based methods. We validate our approach on a wide variety of camera geometries, including perspective, fisheye, and catadioptric. Finally, we show that our approach leads to improvements in the downstream task of depth estimation, achieving state-of-the-art results on the EuRoC dataset with greater computational efficiency than contemporary methods. The project page: https://sites.google.com/ttic.edu/self-sup-self-calib Jiading Fang, Igor Vasiljevic, Vitor Campagnolo Guizilini, Rares Ambrus, Gregory Shakhnarovich, Adrien Gaidon, Matthew R. Walter |
ICRA | 5 |
| 2022 | Boosting Barely Robust Learners: A New Perspective on Adversarial RobustnessabstractWe present an oracle-efficient algorithm for boosting the adversarial robustness of barely robust learners. Barely robust learning algorithms learn predictors that are adversarially robust only on a small fraction $\beta \ll 1$ of the data distribution. Our proposed notion of barely robust learning requires robustness with respect to a ``larger'' perturbation set; which we show is necessary for strongly robust learning, and that weaker relaxations are not sufficient for strongly robust learning. Our results reveal a qualitative and quantitative equivalence between two seemingly unrelated problems: strongly robust learning and barely robust learning. Avrim Blum, Omar Montasser, Gregory Shakhnarovich, Hongyang Zhang 0001 |
NeurIPS | 3 |
| 2022 | Erratum to "Deep Back-Projection Networks for Single Image Super-Resolution"abstractIn the above article [1], the article title was incorrect. The correct article title is "Deep Back-Projection Networks for Single Image Super-Resolution." Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Information-Theoretic Segmentation by Inpainting Error MaximizationabstractWe study image segmentation from an information-theoretic perspective, proposing a novel adversarial method that performs unsupervised segmentation by partitioning images into maximally independent sets. More specifically, we group image pixels into foreground and background, with the goal of minimizing predictability of one set from the other. An easily computed loss drives a greedy search process to maximize inpainting error over these partitions. Our method does not involve training deep networks, is computationally cheap, class-agnostic, and even applicable in isolation to a single unlabeled image. Experiments demonstrate that it achieves a new state-of-the-art in unsupervised segmentation quality, while being substantially faster and more general than competing approaches.1 Pedro Savarese, Sunnie S. Y. Kim, Michael Maire, Gregory Shakhnarovich, David McAllester |
CVPR | 4 |
| 2021 | Fingerspelling Detection in American Sign LanguageabstractFingerspelling, in which words are signed letter by letter, is an important component of American Sign Language. Most previous work on automatic fingerspelling recognition has assumed that the boundaries of fingerspelling regions in signing videos are known beforehand. In this paper, we consider the task of fingerspelling detection in raw, untrimmed sign language videos. This is an important step towards building real-world fingerspelling recognition systems. We propose a benchmark and a suite of evaluation metrics, some of which reflect the effect of detection on the downstream fingerspelling recognition task. In addition, we propose a new model that learns to detect fingerspelling via multi-task training, incorporating pose estimation and fingerspelling recognition (transcription) along with detection, and compare this model to several alternatives. The model outperforms all alternative approaches across all metrics, establishing a state of the art on the benchmark. Bowen Shi 0002, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
CVPR | 3 |
| 2021 | Task-Driven Super Resolution: Object Detection in Low-Resolution Images
Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
ICONIP (5) | 2 |
| 2021 | Deep Back-ProjectiNetworks for Single Image Super-ResolutionabstractPrevious feed-forward architectures of recently proposed deep super-resolution networks learn the features of low-resolution inputs and the non-linear mapping from those to a high-resolution output. However, this approach does not fully address the mutual dependencies of low- and high-resolution images. We propose Deep Back-Projection Networks (DBPN), the winner of two image super-resolution challenges (NTIRE2018 and PIRM2018), that exploit iterative up- and down-sampling layers. These layers are formed as a unit providing an error feedback mechanism for projection errors. We construct mutually-connected up- and down-sampling units each of which represents different types of low- and high-resolution components. We also show that extending this idea to demonstrate a new insight towards more efficient network design substantially, such as parameter sharing on the projection module and transition layer on projection step. The experimental results yield superior results and in particular establishing new state-of-the-art results across multiple data sets, especially for large scaling factors such as 8×. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motionabstractSelf-supervised learning has emerged as a powerful tool for depth and ego-motion estimation, leading to state-of-the-art results on benchmark datasets. However, one significant limitation shared by current methods is the assumption of a known parametric camera model - usually the standard pinhole geometry - leading to failure when applied to imaging systems that deviate significantly from this assumption (e.g., catadioptric cameras or underwater imaging). In this work, we show that self-supervision can be used to learn accurate depth and ego-motion estimation without prior knowledge of the camera model. Inspired by the geometric model of Grossberg and Nayar, we introduce Neural Ray Surfaces (NRS), convolutional networks that represent pixel-wise projection rays, approximating a wide range of cameras. NRS are fully differentiable and can be learned end-to-end from unlabeled raw videos. We demonstrate the use of NRS for self-supervised learning of visual odometry and depth estimation from raw videos obtained using a wide variety of camera systems, including pinhole, fisheye, and catadioptric. Igor Vasiljevic, Vitor Campagnolo Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Gregory Shakhnarovich, Adrien Gaidon |
3DV | 6 |
| 2020 | Space-Time-Aware Multi-Resolution Video EnhancementabstractWe consider the problem of space-time super-resolution (ST-SR): increasing spatial resolution of video frames and simultaneously interpolating frames to increase the frame rate. Modern approaches handle these axes one at a time. In contrast, our proposed model called STARnet super-resolves jointly in space and time. This allows us to leverage mutually informative relationships between time and space: higher resolution can provide more detailed information about motion, and higher frame-rate can provide better pixel alignment. The components of our model that generate latent low- and high-resolution representations during ST-SR can be used to finetune a specialized mechanism for just spatial or just temporal super-resolution. Experimental results demonstrate that STARnet improves the performances of space-time, spatial, and temporal video super-resolution by substantial margins on publicly available datasets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 2 |
| 2020 | Pixel Consensus Voting for Panoptic SegmentationabstractThe core of our approach, Pixel Consensus Voting, is a framework for instance segmentation based on the generalized Hough transform. Pixels cast discretized, probabilistic votes for the likely regions that contain instance centroids. At the detected peaks that emerge in the voting heatmap, backprojection is applied to collect pixels and produce instance masks. Unlike a sliding window detector that densely enumerates object proposals, our method detects instances as a result of the consensus among pixel-wise votes. We implement vote aggregation and backprojection using native operators of a convolutional neural network. The discretization of centroid voting reduces the training of instance segmentation to pixel labeling, analogous and complementary to FCN-style semantic segmentation, leading to an efficient and unified architecture that jointly models things and stuff. We demonstrate the effectiveness of our pipeline on COCO and Cityscapes Panoptic Segmentation and obtain competitive results. Code will be open-sourced. Ruotian Luo, Michael Maire, Gregory Shakhnarovich |
CVPR | 4 |
| 2020 | Deformable Style Transfer
Sunnie S. Y. Kim, Nicholas I. Kolkin, Jason Salavon, Gregory Shakhnarovich |
ECCV (26) | 4 |
| 2019 | Recurrent Back-Projection Network for Video Super-ResolutionabstractWe proposed a novel architecture for the problem of video super-resolution. We integrate spatial and temporal contexts from continuous video frames using a recurrent encoder-decoder module, that fuses multi-frame information with the more traditional, single frame super-resolution path for the target frame. In contrast to most prior work where frames are pooled together by stacking or warping, our model, the Recurrent Back-Projection Network (RBPN) treats each context frame as a separate source of information. These sources are combined in an iterative refinement framework inspired by the idea of back-projection in multiple-image super-resolution. This is aided by explicitly representing estimated inter-frame motion with respect to the target, rather than explicitly aligning frames. We propose a new video super-resolution benchmark, allowing evaluation at a larger scale and considering videos in different motion regimes. Experimental results demonstrate that our RBPN is superior to existing methods on several datasets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 2 |
| 2019 | Style Transfer by Relaxed Optimal Transport and Self-SimilarityabstractThe goal of style transfer algorithms is to render the content of one image using the style of another. We propose Style Transfer by Relaxed Optimal Transport and Self-Similarity (STROTSS), a new optimization-based style transfer algorithm. We extend our method to allow user specified point-to-point or region-to-region control over visual similarity between the style image and the output. Such guidance can be used to either achieve a particular visual effect or correct errors made by unconstrained style transfer. In order to quantitatively compare our method to prior work, we conduct a large-scale user study designed to assess the style-content tradeoff across settings in style transfer algorithms. Our results indicate that for any desired level of content preservation, our method provides higher quality stylization than prior work. Nicholas I. Kolkin, Jason Salavon, Gregory Shakhnarovich |
CVPR | 3 |
| 2019 | Fingerspelling Recognition in the Wild With Iterative Visual AttentionabstractSign language recognition is a challenging gesture sequence recognition problem, characterized by quick and highly coarticulated motion. In this paper we focus on recognition of fingerspelling sequences in American Sign Language (ASL) videos collected in the wild, mainly from YouTube and Deaf social media. Most previous work on sign language recognition has focused on controlled settings where the data is recorded in a studio environment and the number of signers is limited. Our work aims to address the challenges of real-life data, reducing the need for detection or segmentation modules commonly used in this domain. We propose an end-to-end model based on an iterative attention mechanism, without explicit hand detection or segmentation. Our approach dynamically focuses on increasingly high-resolution regions of interest. It out-performs prior work by a large margin. We also introduce a newly collected data set of crowdsourced annotations of fingerspelling in the wild, and show that performance can be further improved with this additional data set. Bowen Shi 0002, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
ICCV | 5 |
| 2019 | Semantic Speech Retrieval With a Visually Grounded Model of Untranscribed SpeechabstractThere is growing interest in models that can learn from unlabelled speech paired with visual context. This setting is relevant for low-resource speech processing, robotics, and human language acquisition research. Here we study how a visually grounded speech model, trained on images of scenes paired with spoken captions, captures aspects of semantics. We use an external image tagger to generate soft text labels from images, which serve as targets for a neural model that maps untranscribed speech to (semantic) keyword labels. We introduce a newly collected data set of human semantic relevance judgements and an associated task, semantic speech retrieval, where the goal is to search for spoken utterances that are semantically relevant to a given text query. Without seeing any text, the model trained on parallel speech and images achieves a precision of almost 60% on its top ten semantic retrievals. Compared to a supervised model trained on transcriptions, our model matches human judgements better by some measures, especially in retrieving non-verbatim semantic matches. We perform an extensive analysis of the model and its resulting representations. Herman Kamper, Gregory Shakhnarovich, Karen Livescu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Deep Back-Projection Networks for Super-ResolutionabstractThe feed-forward architectures of recently proposed deep super-resolution networks learn representations of low-resolution inputs, and the non-linear mapping from those to high-resolution output. However, this approach does not fully address the mutual dependencies of low- and high-resolution images. We propose Deep Back-Projection Networks (DBPN), that exploit iterative up- and downsampling layers, providing an error feedback mechanism for projection errors at each stage. We construct mutually-connected up- and down-sampling stages each of which represents different types of image degradation and high-resolution components. We show that extending this idea to allow concatenation of features across up- and downsampling stages (Dense DBPN) allows us to reconstruct further improve super-resolution, yielding superior results and in particular establishing new state of the art results for large scaling factors such as 8× across multiple data sets. Muhammad Haris 0002, Gregory Shakhnarovich, Norimichi Ukita |
CVPR | 2 |
| 2018 | Discriminability Objective for Training Descriptive CaptionsabstractOne property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning. Ruotian Luo, Brian L. Price, Scott Cohen, Gregory Shakhnarovich |
CVPR | 4 |
| 2018 | Regularizing Deep Networks by Modeling and Predicting Label Structure
Mohammadreza Mostajabi, Michael Maire, Gregory Shakhnarovich |
CVPR | 3 |
| 2018 | Self-Supervised Relative Depth Learning for Urban Scene Understanding
Huaizu Jiang, Gustav Larsson, Michael Maire, Gregory Shakhnarovich, Erik G. Learned-Miller |
ECCV (11) | 4 |
| 2018 | American Sign Language Fingerspelling Recognition in the WildabstractWe address the problem of American Sign Language fingerspelling recognition “in the wild”, using videos collected from websites. We introduce the largest data set available so far for the problem of fingerspelling recognition, and the first using naturally occurring video data. Using this data set, we present the first attempt to recognize fingerspelling sequences in this challenging setting. Unlike prior work, our video data is extremely challenging due to low frame rates and visual variability. To tackle the visual challenges, we train a special-purpose signing hand detector using a small subset of our data. Given the hand detector output, a sequence model decodes the hypothesized fingerspelled letter sequence. For the sequence model, we explore attention-based recurrent encoder-decoders and CTC-based approaches. As the first attempt at fingerspelling recognition in the wild, this work is intended to serve as a baseline for future work on sign language recognition in realistic conditions. We find that, as expected, letter error rates are much higher than in previous work on more controlled data, and we analyze the sources of error and effects of model variants. Bowen Shi 0002, Aurora Martinez Del Rio, Jonathan Keane, Jonathan Michaux, Diane Brentari, Gregory Shakhnarovich, Karen Livescu |
SLT | 6 |
| 2017 | Colorization as a Proxy Task for Visual UnderstandingabstractWe investigate and improve self-supervision as a drop-in replacement for ImageNet pretraining, focusing on automatic colorization as the proxy task. Self-supervised training has been shown to be more promising for utilizing unlabeled data than other, traditional unsupervised learning methods. We build on this success and evaluate the ability of our self-supervised network in several contexts. On VOC segmentation and classification tasks, we present results that are state-of-the-art among methods not using ImageNet labels for pretraining representations. Moreover, we present the first in-depth analysis of self-supervision via colorization, concluding that formulation of the loss, training details and network architecture play important roles in its effectiveness. This investigation is further expanded by revisiting the ImageNet pretraining paradigm, asking questions such as: How much training data is needed? How many labels are needed? How much do features change when fine-tuned? We relate these questions back to self-supervision by showing that colorization provides a similarly powerful supervisory signal as various flavors of ImageNet pretraining. Gustav Larsson, Michael Maire, Gregory Shakhnarovich |
CVPR | 3 |
| 2017 | Comprehension-Guided Referring ExpressionsabstractWe consider generation and comprehension of natural language referring expression for objects in an image. Unlike generic image captioning which lacks natural standard evaluation criteria, quality of a referring expression may be measured by the receivers ability to correctly infer which object is being described. Following this intuition, we propose two approaches to utilize models trained for comprehension task to generate better expressions. First, we use a comprehension module trained on human-generated expressions, as a critic of referring expression generator. The comprehension module serves as a differentiable proxy of human evaluation, providing training signal to the generation module. Second, we use the comprehension model in a generate-and-rerank pipeline, which chooses from candidate expressions generated by a model according to their performance on the comprehension task. We show that both approaches lead to improved referring expression generation on multiple benchmark datasets. Ruotian Luo, Gregory Shakhnarovich |
CVPR | 2 |
| 2017 | Training Deep Networks to be Spatially SensitiveabstractIn many computer vision tasks, for example saliency prediction or semantic segmentation, the desired output is a foreground map that predicts pixels where some criteria is satisfied. Despite the inherently spatial nature of this task commonly used learning objectives do not incorporate the spatial relationships between misclassified pixels and the underlying ground truth. The Weighted F-measure, a recently proposed evaluation metric, does reweight errors spatially, and has been shown to closely correlate with human evaluation of quality, and stably rank predictions with respect to noisy ground truths (such as a sloppy human annotator might generate). However it suffers from computational complexity which makes it intractable as an optimization objective for gradient descent, which must be evaluated thousands or millions of times while learning a model's parameters. We propose a differentiable and efficient approximation of this metric. By incorporating spatial information into the objective we can use a simpler model than competing methods without sacrificing accuracy, resulting in faster inference speeds and alleviating the need for pre/post-processing. We match (or improve) performance on several tasks compared to prior state of the art by traditional metrics, and in many cases significantly improve performance by the weighted F-measure. Nicholas I. Kolkin, Gregory Shakhnarovich, Eli Shechtman |
ICCV | 2 |
| 2017 | FractalNet: Ultra-Deep Neural Networks without Residuals
Gustav Larsson, Michael Maire, Gregory Shakhnarovich |
ICLR (Poster) | 3 |
| 2017 | Visually Grounded Learning of Keyword Prediction from Untranscribed SpeechabstractDuring language acquisition, infants have the benefit of visual cues to ground spoken language. Robots similarly have access to audio and visual sensors. Recent work has shown that images and spoken captions can be mapped into a meaningful common space, allowing images to be retrieved using speech and vice versa. In this setting of images paired with untranscribed spoken captions, we consider whether computer vision systems can be used to obtain textual labels for the speech. Concretely, we use an image-to-words multi-label visual classifier to tag images with soft textual labels, and then train a neural network to map from the speech to these soft targets. We show that the resulting speech system is able to predict which words occur in an utterance---acting as a spoken bag-of-words classifier---without seeing any parallel speech and text. We find that the model often confuses semantically related words, e.g. "man" and "person", making it even more effective as a semantic keyword spotter. Herman Kamper, Shane Settle, Gregory Shakhnarovich, Karen Livescu |
INTERSPEECH | 3 |
| 2017 | Lexicon-free fingerspelling recognition from video: Data, models, and signer adaptation
Taehwan Kim 0003, Jonathan Keane, Hao Tang 0002, Jason Riggle, Gregory Shakhnarovich, Diane Brentari, Karen Livescu |
Comput. Speech Lang. | 6 |
| 2016 | Learning Representations for Automatic Colorization
Gustav Larsson, Michael Maire, Gregory Shakhnarovich |
ECCV (4) | 3 |
| 2016 | Depth from a Single Image by Harmonizing Overcomplete Local Network PredictionsabstractA single color image can contain many cues informative towards different aspects of local geometric structure. We approach the problem of monocular depth estimation by using a neural network to produce a mid-level representation that summarizes these cues. This network is trained to characterize local scene geometry by predicting, at every image location, depth derivatives of different orders, orientations and scales. However, instead of a single estimate for each derivative, the network outputs probability distributions that allow it to express confidence about some coefficients, and ambiguity about others. Scene depth is then estimated by harmonizing this overcomplete set of network predictions, using a globalization procedure that finds a single consistent depth map that best matches all the local derivative distributions. We demonstrate the efficacy of this approach through evaluation on the NYU v2 depth data set. Ayan Chakrabarti, Jingyu Shao, Gregory Shakhnarovich |
NIPS | 3 |
| 2015 | Feedforward semantic segmentation with zoom-out featuresabstractWe introduce a purely feed-forward architecture for semantic segmentation. We map small image elements (superpixels) to rich feature representations extracted from a sequence of nested regions of increasing extent. These regions are obtained by “zooming out” from the superpixel all the way to scene-level resolution. This approach exploits statistical structure in the image and in the label space without setting up explicit structured prediction mechanisms, and thus avoids complex and expensive inference. Instead superpixels are classified by a feedforward multilayer network. Our architecture achieves 69.6% average accuracy on the PASCAL VOC 2012 test set. Mohammadreza Mostajabi, Payman Yadollahpour, Gregory Shakhnarovich |
CVPR | 3 |
| 2014 | Knowing a Good HOG Filter When You See It: Efficient Selection of Filters for Detection
Ejaz Ahmed 0002, Gregory Shakhnarovich, Subhransu Maji |
ECCV (1) | 2 |
| 2014 | Discriminative Metric Learning by Neighborhood Gerrymandering
Shubhendu Trivedi, David A. McAllester, Gregory Shakhnarovich |
NIPS | 3 |
| 2014 | A Consistent Estimator of the Expected Gradient Outerproduct
Shubhendu Trivedi, Samory Kpotufe, Gregory Shakhnarovich |
UAI | 4 |
| 2014 | Part and Attribute Discovery from Relative Annotations
Subhransu Maji, Gregory Shakhnarovich |
Int. J. Comput. Vis. | 2 |
| 2013 | Part Discovery from Partial CorrespondenceabstractWe study the problem of part discovery when partial correspondence between instances of a category are available. For visual categories that exhibit high diversity in structure such as buildings, our approach can be used to discover parts that are hard to name, but can be easily expressed as a correspondence between pairs of images. Parts naturally emerge from point-wise landmark matches across many instances within a category. We propose a learning framework for automatic discovery of parts in such weakly supervised settings, and show the utility of the rich part library learned in this way for three tasks: object detection, category-specific saliency estimation, and fine-grained image parsing. 1. Subhransu Maji, Gregory Shakhnarovich |
CVPR | 2 |
| 2013 | Image Segmentation by Cascaded Region AgglomerationabstractWe propose a hierarchical segmentation algorithm that starts with a very fine over segmentation and gradually merges regions using a cascade of boundary classifiers. This approach allows the weights of region and boundary features to adapt to the segmentation scale at which they are applied. The stages of the cascade are trained sequentially, with asymetric loss to maximize boundary recall. On six segmentation data sets, our algorithm achieves best performance under most region-quality measures, and does it with fewer segments than the prior work. Our algorithm is also highly competitive in a dense over segmentation (super pixel) regime under boundary-based measures. Zhile Ren, Gregory Shakhnarovich |
CVPR | 2 |
| 2013 | Discriminative Re-ranking of Diverse SegmentationsabstractThis paper introduces a hybrid, two-stage approach to semantic image segmentation. In the first stage a probabilistic model generates a set of diverse plausible segmentations. In the second stage, a discriminatively trained re-ranking model selects the best segmentation from this set. The re-ranking stage can use much more complex features than what could be tractably used in the probabilistic model, allowing a better exploration of the solution space than possible by simply producing the most probable solution from the probabilistic model. While our proposed approach already achieves state-of-the-art results (48%) on the challenging VOC 2012 dataset, our machine and human analyses suggest that even larger gains are possible with such an approach. Payman Yadollahpour, Dhruv Batra, Gregory Shakhnarovich |
CVPR | 3 |
| 2013 | A Systematic Exploration of Diversity in Machine TranslationabstractThis paper addresses the problem of producing a diverse set of plausible translations.We present a simple procedure that can be used with any statistical machine translation (MT) system.We explore three ways of using diverse translations: (1) system combination, (2) discriminative reranking with rich features, and (3) a novel post-editing scenario in which multiple translations are presented to users.We find that diversity can improve performance on these tasks, especially for sentences that are difficult for MT. Kevin Gimpel, Dhruv Batra, Chris Dyer, Gregory Shakhnarovich |
EMNLP | 4 |
| 2013 | Fingerspelling Recognition with Semi-Markov Conditional Random FieldsabstractRecognition of gesture sequences is in general a very difficult problem, but in certain domains the difficulty may be mitigated by exploiting the domain's ``grammar''. One such grammatically constrained gesture sequence domain is sign language. In this paper we investigate the case of finger spelling recognition, which can be very challenging due to the quick, small motions of the fingers. Most prior work on this task has assumed a closed vocabulary of finger spelled words, here we study the more natural open-vocabulary case, where the only domain knowledge is the possible finger spelled letters and statistics of their sequences. We develop a semi-Markov conditional model approach, where feature functions are defined over segments of video and their corresponding letter labels. We use classifiers of letters and linguistic hand shape features, along with expected motion profiles, to define segmental feature functions. This approach improves letter error rate (Levenshtein distance between hypothesized and correct letter sequences) from 16.3% using a hidden Markov model baseline to 11.6% using the proposed semi-Markov model. Taehwan Kim 0003, Gregory Shakhnarovich, Karen Livescu |
ICCV | 2 |
| 2012 | Diverse M-Best Solutions in Markov Random Fields
Dhruv Batra, Payman Yadollahpour, Abner Guzmán-Rivera, Gregory Shakhnarovich |
ECCV (5) | 4 |
| 2012 | American sign language fingerspelling recognition with phonological feature-based tandem modelsabstractWe study the recognition of fingerspelling sequences in American Sign Language from video using tandem-style models, in which the outputs of multilayer perceptron (MLP) classifiers are used as observations in a hidden Markov model (HMM)-based recognizer. We compare a baseline HMM-based recognizer, a tandem recognizer using MLP letter classifiers, and a tandem recognizer using MLP classifiers of phonological features. We present experiments on a database of fingerspelling videos. We find that the tandem approaches outperform an HMM-based baseline, and that phonological feature-based tandem models outperform letter-based tandem models. Taehwan Kim 0003, Karen Livescu, Gregory Shakhnarovich |
SLT | 3 |
| 2012 | Viewpoint-aware object detection and continuous pose estimation
Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
Image Vis. Comput. | 5 |
| 2011 | Viewpoint-aware object detection and pose estimationabstractWe describe an approach to category-level detection and viewpoint estimation for rigid 3D objects from single 2D images. In contrast to many existing methods, we directly integrate 3D reasoning with an appearance-based voting architecture. Our method relies on a nonparametric representation of a joint distribution of shape and appearance of the object class. Our voting method employs a novel parametrization of joint detection and viewpoint hypothesis space, allowing efficient accumulation of evidence. We combine this with a re-scoring and refinement mechanism, using an ensemble of view-specific Support Vector Machines. We evaluate the performance of our approach in detection and pose estimation of cars on a number of benchmark datasets. Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
ICCV | 5 |
| 2010 | Sparse Coding for Learning Interpretable Spatio-Temporal PrimitivesabstractSparse coding has recently become a popular approach in computer vision to learn dictionaries of natural images. In this paper we extend sparse coding to learn interpretable spatio-temporal primitives of human motion. We cast the problem of learning spatio-temporal primitives as a tensor factorization problem and introduce constraints to learn interpretable primitives. In particular, we use group norms over those tensors, diagonal constraints on the activations as well as smoothness constraints that are inherent to human motion. We demonstrate the effectiveness of our approach to learn interpretable representations of human motion from motion capture data, and show that our approach outperforms recently developed matching pursuit and sparse coding algorithms. Taehwan Kim 0003, Gregory Shakhnarovich, Raquel Urtasun |
NIPS | 2 |
| 2009 | Discovery of phosphorylation motif mixtures in phosphoproteomics dataabstractMOTIVATION: Modification of proteins via phosphorylation is a primary mechanism for signal transduction in cells. Phosphorylation sites on proteins are determined in part through particular patterns, or motifs, present in the amino acid sequence. RESULTS: We describe an algorithm that simultaneously discovers multiple motifs in a set of peptides that were phosphorylated by several different kinases. Such sets of peptides are routinely produced in proteomics experiments.Our motif-finding algorithm uses the principle of minimum description length to determine a mixture of sequence motifs that distinguish a foreground set of phosphopeptides from a background set of unphosphorylated peptides. We show that our algorithm outperforms existing motif-finding algorithms on synthetic datasets consisting of mixtures of known phosphorylation sites. We also derive a motif specificity score that quantifies whether or not the phosphoproteins containing an instance of a motif have a significant number of known interactions. Application of our motif-finding algorithm to recently published human and mouse proteomic studies recovers several known phosphorylation motifs and reveals a number of novel motifs that are enriched for interactions with a particular kinase or phosphatase. Our tools provide a new approach for uncovering the sequence specificities of uncharacterized kinases or phosphatases. Anna M. Ritz, Gregory Shakhnarovich, Arthur R. Salomon, Benjamin J. Raphael |
Bioinform. | 2 |
| 2008 | A Slicing-Based Coherence Measure for Clusters of DTI Integral Curves
Çagatay Demiralp, Gregory Shakhnarovich, Song Zhang 0004, David H. Laidlaw |
MICCAI (1) | 2 |
| 2008 | Nearest-Neighbor Methods in Learning and VisionabstractIn this excellent book, the editors deal with the state-of-the-art, current best practices, and some innovative applications of nearest-neighbor methods in learning and vision. This volume brings together contributions of top-level researchers in theory of computation, machine learning, and computer vision with the goal of closing up the gaps between disciplines and current state-of-the-art methods for emerging applications. All the content is well-written, highly relevant, original, and timely. The audience for this book consists of researchers, scientists, engineers, professionals, and academics, working not only in this field, but also in any field that could benefit from these powerful methods. This book can be particularly useful to researchers working on the basis set expansion networks. Gregory Shakhnarovich, Trevor Darrell, Piotr Indyk |
IEEE Trans. Neural Networks | 1 |
| 2006 | Conditional Random People: Tracking Humans with CRFs and Grid FiltersabstractWe describe a state-space tracking approach based on a Conditional Random Field (CRF) model, where the observation potentials are learned from data. We find functions that embed both state and observation into a space where similarity corresponds to L1 distance, and define an observation potential based on distance in this space. This potential is extremely fast to compute and in conjunction with a grid-filtering framework can be used to reduce a continuous state estimation problem to a discrete one. We show how a state temporal prior in the grid-filter can be computed in a manner similar to a sparse HMM, resulting in real-time system performance. The resulting system is used for human pose tracking in video sequences. Leonid Taycher, David Demirdjian, Trevor Darrell, Gregory Shakhnarovich |
CVPR (1) | 4 |
| 2006 | An investigation of computational and informational limits in Gaussian mixture clusteringabstractWe investigate under what conditions clus-tering by learning a mixture of spherical Gaussians is (a) computationally tractable; and (b) statistically possible. We show that using principal component projection greatly aids in recovering the clustering using EM; present empirical evidence that even using such a projection, there is still a large gap between the number of samples needed to re-cover the clustering using EM, and the num-ber of samples needed without computational restrictions; and characterize the regime in which such a gap exists. 1. Nathan Srebro, Gregory Shakhnarovich, Sam T. Roweis |
ICML | 2 |
| 2006 | Nonlinear physically-based models for decoding motor-cortical population activityabstractNeural motor prostheses (NMPs) require the accurate decoding of motor cortical population activity for the control of an artificial motor system. Previous work on cortical decoding for NMPs has focused on the recovery of hand kinematics. Human NMPs however may require the control of computer cursors or robotic devices with very different physical and dynamical properties. Here we show that the firing rates of cells in the primary motor cortex of non-human primates can be used to control the parameters of an artificial physical system exhibiting realistic dynamics. The model represents 2D hand motion in terms of a point mass connected to a system of idealized springs. The nonlinear spring coefficients are estimated from the firing rates of neurons in the motor cortex. We evaluate linear and a nonlinear decoding algorithms using neural recordings from two monkeys performing two different tasks. We found that the decoded spring coefficients produced accurate hand trajectories compared with state-of-the-art methods for direct decoding of hand kinematics. Furthermore, using a physically-based system produced decoded movements that were more “natural” in that their frequency spectrum more closely matched that of natural hand movements. Gregory Shakhnarovich, Sung-Phil Kim, Michael J. Black |
NIPS | 1 |
| 2005 | Face Recognition with Image Sets Using Manifold Density DivergenceabstractIn many automatic face recognition applications, a set of a person's face images is available rather than a single image. In this paper, we describe a novel method for face recognition using image sets. We propose a flexible, semi-parametric model for learning probability densities confined to highly non-linear but intrinsically low-dimensional manifolds. The model leads to a statistical formulation of the recognition problem in terms of minimizing the divergence between densities estimated on these manifolds. The proposed method is evaluated on a large data set, acquired in realistic imaging conditions with severe illumination variation. Our algorithm is shown to match the best and outperform other state-of-the-art algorithms in the literature, achieving 94% recognition rate on average. Ognjen Arandjelovic, Gregory Shakhnarovich, John Fisher, Roberto Cipolla, Trevor Darrell |
CVPR (1) | 2 |
| 2005 | Avoiding the "Streetlight Effect": Tracking by Exploring Likelihood ModesabstractClassic methods for Bayesian inference effectively constrain search to lie within regions of significant probability of the temporal prior. This is efficient with an accurate dynamics model, but otherwise is prone to ignore significant peaks in the true posterior. A more accurate posterior estimate can be obtained by explicitly finding modes of the likelihood function and combining them with a weak temporal prior. In our approach, modes are found using efficient example-based matching followed by local refinement to find peaks and estimate peak bandwidth. By reweighting these peaks according to the temporal prior we obtain an estimate of the full posterior model. We show comparative results on real and synthetic images in a high degree of freedom articulated tracking task. David Demirdjian, Leonid Taycher, Gregory Shakhnarovich, Kristen Grauman, Trevor Darrell |
ICCV | 3 |
| 2005 | Learning silhouette features for control of human motionabstractWe present a vision-based performance interface for controlling animated human characters. The system interactively combines information about the user's motion contained in silhouettes from three viewpoints with domain knowledge contained in a motion capture database to produce an animation of high quality. Such an interactive system might be useful for authoring, for teleconferencing, or as a control interface for a character in a game. In our implementation, the user performs in front of three video cameras; the resulting silhouettes are used to estimate his orientation and body configuration based on a set of discriminative local features. Those features are selected by a machine-learning algorithm during a preprocessing step. Sequences of motions that approximate the user's actions are extracted from the motion database and scaled in time to match the speed of the user's motion. We use swing dancing, a complex human motion, to demonstrate the effectiveness of our approach. We compare our results to those obtained with a set of global features, Hu moments, and ground truth measurements from a motion capture system. Liu Ren 0001, Gregory Shakhnarovich, Jessica K. Hodgins, Hanspeter Pfister, Paul A. Viola |
ACM Trans. Graph. | 2 |
| 2003 | A Bayesian Approach to Image-Based Visual Hull ReconstructionabstractWe present a Bayesian approach to image-based visual hull reconstruction. The 3D (three-dimensional) shape of an object of a known class is represented by sets of silhouette views simultaneously observed from multiple cameras. We show how the use of a class-specific prior in a visual hull reconstruction can reduce the effect of segmentation errors from the silhouette extraction process. In our representation, 3D information is implicit in the joint observations of multiple contours from known viewpoints. We model the prior density using a probabilistic principal components analysis-based technique and estimate a maximum a posteriori reconstruction of multi-view contours. The proposed method is applied to a dataset of pedestrian images, and improvements in the approximate 3D models under various noise conditions are shown. Kristen Grauman, Gregory Shakhnarovich, Trevor Darrell |
CVPR (1) | 2 |
| 2003 | Inferring 3D Structure with a Statistical Image-Based Shape ModelabstractWe present an image-based approach to infer 3D structure parameters using a probabilistic "shape+structure" model. The 3D shape of an object class is represented by sets of contours from silhouette views simultaneously observed from multiple calibrated cameras, while structural features of interest on the object are denoted by a number of 3D locations. A prior density over the multiview shape and corresponding structure is constructed with a mixture of probabilistic principal components analyzers. Given a novel set of contours, we infer the unknown structure parameters from the new shape's Bayesian reconstruction. Model matching and parameter inference are done entirely in the image domain and require no explicit 3D construction. Our shape model enables accurate estimation of structure despite segmentation errors or missing views in the input silhouettes, and it works even with only a single input view. Using a training set of thousands of pedestrian images generated from a synthetic model, we can accurately infer the 3D locations of 19 joints on the body based on observed silhouette contours from real images. Kristen Grauman, Gregory Shakhnarovich, Trevor Darrell |
ICCV | 2 |
| 2003 | Fast Pose Estimation with Parameter-Sensitive HashingabstractExample-based methods are effective for parameter estimation problems when the underlying system is simple or the dimensionality of the input is low. For complex and high-dimensional problems such as pose estimation, the number of required examples and the computational complexity rapidly become prohibitively high. We introduce a new algorithm that learns a set of hashing functions that efficiently index examples relevant to a particular estimation task. Our algorithm extends locality-sensitive hashing, a recently developed method to find approximate neighbors in time sublinear in the number of examples. This method depends critically on the choice of hash functions that are optimally relevant to a particular estimation problem. Experiments demonstrate that the resulting algorithm, which we call parameter-sensitive hashing, can rapidly and accurately estimate the articulated pose of human figures from a large database of example images. Gregory Shakhnarovich, Paul A. Viola, Trevor Darrell |
ICCV | 1 |
| 2002 | Face Recognition from Long-Term Observations
Gregory Shakhnarovich, John W. Fisher III, Trevor Darrell |
ECCV (3) | 1 |
| 2002 | Boosted Dyadic Kernel DiscriminantsabstractWe introduce a novel learning algorithm for binary classi(cid:12)cation with hyperplane discriminants based on pairs of training points from opposite classes (dyadic hypercuts). This algorithm is further extended to nonlinear discriminants using kernel functions satisfy- ing Mercer’s conditions. An ensemble of simple dyadic hypercuts is learned incrementally by means of a con(cid:12)dence-rated version of Ad- aBoost, which provides a sound strategy for searching through the (cid:12)nite set of hypercut hypotheses. In experiments with real-world datasets from the UCI repository, the generalization performance of the hypercut classi(cid:12)ers was found to be comparable to that of SVMs and k-NN classi(cid:12)ers. Furthermore, the computational cost of classi(cid:12)cation (at run time) was found to be similar to, or bet- ter than, that of SVM. Similarly to SVMs, boosted dyadic kernel discriminants tend to maximize the margin (via AdaBoost). In contrast to SVMs, however, we o(cid:11)er an on-line and incremental learning machine for building kernel discriminants whose complex- ity (number of kernel evaluations) can be directly controlled (traded o(cid:11) for accuracy). Baback Moghaddam, Gregory Shakhnarovich |
NIPS | 2 |
| 2001 | Integrated Face and Gait Recognition From Multiple ViewsabstractWe develop a view-normalization approach to multi-view face and gait recognition. An image-based visual hull (IBVH) is computed from a set of monocular views and used to render virtual views for tracking and recognition. We determine canonical viewpoints by examining the 3D structure, appearance (texture), and motion of the moving person. For optimal face recognition, we place virtual cameras to capture frontal face appearance; for gait recognition we place virtual cameras to capture a side-view of the person. Multiple cameras can be rendered simultaneously, and camera position is dynamically updated as the person moves through the workspace. Image sequences from each canonical view are passed to an unmodified face or gait recognition algorithm. We show that our approach provides greater recognition accuracy than is obtained using the unnormalized input sequences, and that integrated face and gait recognition provides improved performance over either modality alone. Canonical view estimation, rendering, and recognition have been efficiently implemented and can run at near real-time speeds. Gregory Shakhnarovich, Lily Lee, Trevor Darrell |
CVPR (1) | 1 |
| 2001 | Smoothed Bootstrap and Statistical Data Cloning for Classifier Evaluation
Gregory Shakhnarovich, Ran El-Yaniv, Yoram Baram |
ICML | 1 |