VLDB 2026 Research / reviewers in the wild / expert
David J. Fleet
dblp:07/2099
· DBLP profile ↗
125ranked-venue papers
9as first author
24since 2021 · last 2025
0000-0003-0734-7114ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 114 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Resolution Frame Interpolation with Patch-based Cascaded DiffusionabstractDespite the recent progress, existing frame interpolation methods still struggle with processing extremely high resolution input and handling challenging cases such as repetitive textures, thin objects, and large motion. To address these issues, we introduce a patch-based cascaded pixel diffusion model for high resolution frame interpolation, HiFI, that excels in these scenarios while achieving competitive performance on standard benchmarks. Cascades, which generate a series of images from low to high resolution, can help significantly with large or complex motion that require both global context for a coarse solution and detailed context for high resolution output. However, contrary to prior work on cascaded diffusion models which perform diffusion on increasingly large resolutions, we use a single model that always performs diffusion at the same resolution and upsamples by processing patches of the inputs and the prior solution. At inference time, this drastically reduces memory usage and allows a single model, solving both frame interpolation (base model’s task) and spatial up-sampling, saving training cost as well. HiFI excels at high-resolution images and complex repeated textures that require global context, achieving comparable or state-of-the-art performance on various benchmarks (Vimeo, Xiph, X-Test, and SEPE-8K). We further introduce a new dataset, LaMoR, that focuses on particularly challenging cases, and HiFI significantly outperforms other baselines. Junhwa Hur, Charles Herrmann, Saurabh Saxena, Janne Kontkanen, Wei-Sheng Lai, Michael Rubinstein, David J. Fleet, Deqing Sun |
AAAI | 8 |
| 2025 | RoMo: Robust Motion Segmentation Improves Structure from MotionabstractThere has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin. Lily Goli, Sara Sabour, Mark J. Matthews, Marcus A. Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Saxena, Andrea Tagliasacchi |
ICCV | 7 |
| 2025 | Controlling Space and Time with Diffusion ModelsabstractWe present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works which generally operate in limited domains (e.g., object centric).
4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of
image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes.
See https://4d-diffusion.github.io for video samples. Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet |
ICLR | 5 |
| 2025 | Reconstructing Heterogeneous Biomolecules via Hierarchical Gaussian Mixtures and Part DiscoveryabstractCryo-EM is a transformational paradigm in molecular biology where computational methods are used to infer 3D molecular structure at atomic resolution from extremely noisy 2D electron microscope images.
At the forefront of research is how to model the structure when the imaged particles exhibit non-rigid conformational flexibility and compositional variation where parts are sometimes missing.
We introduce a novel 3D reconstruction framework with a hierarchical Gaussian mixture model, inspired in part by Gaussian Splatting for 4D scene reconstruction.
In particular, the structure of the model is grounded in an initial process that infers a part-based segmentation of the particle, providing essential inductive bias in order to handle both conformational and compositional variability. The framework, called \methodName, is shown to reveal biologically meaningful structures on complex experimental datasets, and establishes a new state-of-the-art on CryoBench, a benchmark for cryo-EM heterogeneity methods. Shayan Shekarforoush, David B. Lindell, Marcus A. Brubaker, David J. Fleet |
NeurIPS | 4 |
| 2025 | Personalized Video-Based Hand Taxonomy Using Egocentric Video in the WildabstractOBJECTIVE: Hand function is central to inter- actions with our environment. Developing a comprehen- sive model of hand grasps in naturalistic environments is crucial across various disciplines, including robotics, ergonomics, and rehabilitation. Creating such a taxonomy poses challenges due to the significant variation in grasping strategies that individuals may employ. For instance, individuals with impaired hands, such as those with spinal cord injuries (SCI), may develop unique grasps not used by unimpaired individuals. These grasping techniques may differ from person to person, influenced by variable senso- rimotor impairment, creating a need for personalized meth- ods of analysis. METHOD: This study aimed to automatically identify the dominant distinct hand grasps for each indi- vidual without reliance on a priori taxonomies, by applying semantic clustering to egocentric video. Egocentric video recordings collected in the homes of 19 individual with cervical SCI were used to cluster grasping actions with semantic significance. A deep learning model integrating posture and appearance data was employed to create a personalized hand taxonomy. RESULTS: Quantitative analysis reveals a cluster purity of 67.6% ± 24.2% with 18.0% ± 21.8% redundancy. Qualitative assessment revealed meaningful clusters in video content. DISCUSSION: This methodology provides a flexible and effective strategy to analyze hand function in the wild, with applications in clinical assess- ment and in-depth characterization of human-environment interactions in a variety of contexts. Mehdy Dousty, David J. Fleet, José Zariffa |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | SpotLessSplats: Ignoring Distractors in 3D Gaussian SplattingabstractThree-dimensional Gaussian Splatting (3DGS) is a promising technique for 3D reconstruction, offering efficient training and rendering speeds, making it suitable for real-time applications. However, current methods require highly controlled environments–no moving people or wind-blown elements, and consistent lighting–to meet the interview consistency assumption of 3DGS. This makes reconstruction of real-world captures problematic. We present SpotLessSplats, an approach that leverages pre-trained and general-purpose features coupled with robust optimization to effectively ignore transient distractors. Our method achieves state-of-the-art reconstruction quality both visually and quantitatively, on casual captures. Sara Sabour, Lily Goli, Georgios Kopanas, Mark J. Matthews, Dmitry Lagun, Leonidas J. Guibas, Alec Jacobson, David J. Fleet, Andrea Tagliasacchi |
ACM Trans. Graph. | 8 |
| 2024 | Directly Fine-Tuning Diffusion Models on Differentiable RewardsabstractWe present Direct Reward Fine-Tuning (DRaFT), a simple and effective method for fine-tuning diffusion models to maximize differentiable reward functions, such as scores from human preference models. We first show that it is possible to backpropagate the reward function gradient through the full sampling procedure, and that doing so achieves strong performance on a variety of rewards, outperforming reinforcement learning-based approaches. We then propose more efficient variants of DRaFT: DRaFT-K, which truncates backpropagation to only the last K steps of sampling, and DRaFT-LV, which obtains lower-variance gradient estimates for the case when K=1. We show that our methods work well for a variety of reward functions and can be used to substantially improve the aesthetic quality of images generated by Stable Diffusion 1.4. Finally, we draw connections between our approach and prior work, providing a unifying perspective on the design space of gradient-based fine-tuning algorithms. Kevin Clark, Paul Vicol, Kevin Swersky, David J. Fleet |
ICLR | 4 |
| 2024 | CryoSPIN: Improving Ab-Initio Cryo-EM Reconstruction with Semi-Amortized Pose InferenceabstractCryo-EM is an increasingly popular method for determining the atomic resolution 3D structure of macromolecular complexes (eg, proteins) from noisy 2D images captured by an electron microscope. The computational task is to reconstruct the 3D density of the particle, along with 3D pose of the particle in each 2D image, for which the posterior pose distribution is highly multi-modal. Recent developments in cryo-EM have focused on deep learning for which amortized inference has been used to predict pose. Here, we address key problems with this approach, and propose a new semi-amortized method, cryoSPIN, in which reconstruction begins with amortized inference and then switches to a form of auto-decoding to refine poses locally using stochastic gradient descent. Through evaluation on synthetic datasets, we demonstrate that cryoSPIN is able to handle multi-modal pose distributions during the amortized inference stage, while the later, more flexible stage of direct pose optimization yields faster and more accurate convergence of poses compared to baselines. On experimental data, we show that cryoSPIN outperforms the state-of-the-art cryoAI in speed and reconstruction quality. Shayan Shekarforoush, David B. Lindell, Marcus A. Brubaker, David J. Fleet |
NeurIPS | 4 |
| 2024 | Hand Grasp Classification in Egocentric Video After Cervical Spinal Cord InjuryabstractOBJECTIVE: The hand function of individuals with spinal cord injury (SCI) plays a crucial role in their independence and quality of life. Wearable cameras provide an opportunity to analyze hand function in non-clinical environments. Summarizing the video data and documenting dominant hand grasps and their usage frequency would allow clinicians to quickly and precisely analyze hand function. METHOD: We introduce a new hierarchical model to summarize the grasping strategies of individuals with SCI at home. The first level classifies hand-object interaction using hand-object contact estimation. We developed a new deep model in the second level by incorporating hand postures and hand-object contact points using contextual information. RESULTS: In the first hierarchical level, a mean of 86% ±1.0% was achieved among 17 participants. At the grasp classification level, the mean average accuracy was 66.2 ±12.9%. The grasp classifier's performance was highly dependent on the participants, with accuracy varying from 41% to 78%. The highest grasp classification accuracy was obtained for the model with smoothed grasp classification, using a ResNet50 backbone architecture for the contextual head and a temporal pose head. DISCUSSION: We introduce a novel algorithm that, for the first time, enables clinicians to analyze the quantity and type of hand movements in individuals with spinal cord injury at home. The algorithm can find applications in other research fields, including robotics, and most neurological diseases that affect hand function, notably, stroke and Parkinson's. Mehdy Dousty, David J. Fleet, José Zariffa |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingabstractText-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen [36] on text-guided image inpainting. Imagen Editor's edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment – such that Imagen Editor is preferred over DALL-E 2 [31] and Stable Diffusion [33] – and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes. Su Wang 0001, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mohammad Norouzi 0002 |
CVPR | 9 |
| 2023 | RobustNeRF: Ignoring Distractors with Robust LossesabstractNeural radiance fields (NeRF) excel at synthesizing new views given multi-view, calibrated images of a static scene. When scenes include distractors, which are not persistent during image capture (moving objects, lighting variations, shadows), artifacts appear as view-dependent effects or ‘floaters’. To cope with distractors, we advocate a form of robust estimation for NeRF training, modeling distractors in training data as outliers of an optimization problem. Our method successfully removes outliers from a scene and improves upon our baselines, on synthetic and real-world scenes. Our technique is simple to incorporate in modern NeRF frameworks, with few hyper-parameters. It does not assume a priori knowledge of the types of distractors, and is instead focused on the optimization problem rather than pre-processing or modeling transient objects. More results at https://robustnerf.github.io/public. Sara Sabour, Suhani Vora, Daniel Duckworth, Ivan Krasin, David J. Fleet, Andrea Tagliasacchi |
CVPR | 5 |
| 2023 | A Generalist Framework for Panoptic Segmentation of Images and VideosabstractPanoptic segmentation assigns semantic and instance ID labels to every pixel of an image. As permutations of instance IDs are also valid solutions, the task requires learning of high-dimensional one-to-many mapping. As a result, state-of-the-art approaches use customized architectures and task-specific loss functions. We formulate panoptic segmentation as a discrete data generation problem, without relying on inductive bias of the task. A diffusion model is proposed to model panoptic masks, with a simple architecture and generic loss function. By simply adding past predictions as a conditioning signal, our method is capable of modeling video (in a streaming setting) and thereby learns to track object instances automatically. With extensive experiments, we demonstrate that our simple approach can perform competitively to state-of-the-art specialist methods in similar settings.1 Ting Chen 0007, Lala Li, Saurabh Saxena, Geoffrey E. Hinton, David J. Fleet |
ICCV | 5 |
| 2023 | Scalable Adaptive Computation for Iterative GenerationabstractNatural data is redundant yet predominant architectures tile computation uniformly across their input and output space. We propose the Recurrent Interface Network (RIN), an attention-based architecture that decouples its core computation from the dimensionality of the data, enabling adaptive computation for more scalable generation of high-dimensional data. RINs focus the bulk of computation (i.e. global self-attention) on a set of latent tokens, using cross-attention to read and write (i.e. route) information between latent and data tokens. Stacking RIN blocks allows bottom-up (data to latent) and top-down (latent to data) feedback, leading to deeper and more expressive routing. While this routing introduces challenges, this is less problematic in recurrent computation settings where the task (and routing problem) changes gradually, such as iterative generation with diffusion models. We show how to leverage recurrence by conditioning the latent tokens at each forward pass of the reverse diffusion process with those from prior computation, i.e. latent self-conditioning. RINs yield state-of-the-art pixel diffusion models for image and video generation, scaling to1024×1024 images without cascades or guidance, while being domain-agnostic and up to 10× more efficient than 2D and 3D U-Nets. Allan Jabri, David J. Fleet, Ting Chen 0007 |
ICML | 2 |
| 2023 | The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationabstractDenoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity.
We show that they also excel in estimating optical flow and monocular depth, surprisingly without task-specific architectures and loss functions that are predominant for these tasks.
Compared to the point estimates of conventional regression-based methods, diffusion models also enable Monte Carlo inference, e.g., capturing uncertainty and ambiguity in flow and depth.
With self-supervised pre-training, the combined use of synthetic and real data for supervised training, and technical innovations (infilling and step-unrolled denoising diffusion training) to handle noisy-incomplete training data, one can train state-of-the-art diffusion models for depth and optical flow estimation, with additional zero-shot coarse-to-fine refinement for high resolution estimates.
Extensive experiments focus on quantitative performance against benchmarks, ablations, and the model's ability to capture uncertainty and multimodality, and impute missing values. Our model obtains a state-of-the-art relative depth error of 0.074 on the indoor NYU benchmark and an Fl-all score of 3.26\% on the KITTI optical flow benchmark, about 25\% better than the best published method. Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi 0002, Deqing Sun, David J. Fleet |
NeurIPS | 7 |
| 2023 | Image Super-Resolution via Iterative RefinementabstractWe present SR3, an approach to image Super-Resolution via Repeated Refinement. SR3 adapts denoising diffusion probabilistic models (Ho et al. 2020), (Sohl-Dickstein et al. 2015) to image-to-image translation, and performs super-resolution through a stochastic iterative denoising process. Output images are initialized with pure Gaussian noise and iteratively refined using a U-Net architecture that is trained on denoising at various noise levels, conditioned on a low-resolution input image. SR3 exhibits strong performance on super-resolution tasks at different magnification factors, on faces and natural images. We conduct human evaluation on a standard 8× face super-resolution task on CelebA-HQ for which SR3 achieves a fool rate close to 50%, suggesting photo-realistic outputs, while GAN baselines do not exceed a fool rate of 34%. We evaluate SR3 on a 4× super-resolution task on ImageNet, where SR3 outperforms baselines in human evaluation and classification accuracy of a ResNet-50 classifier trained on high-resolution images. We further show the effectiveness of SR3 in cascaded image generation, where a generative model is chained with super-resolution models to synthesize high-resolution images with competitive FID scores on the class-conditional 256×256 ImageNet generation challenge. Chitwan Saharia, Jonathan Ho, Tim Salimans, David J. Fleet, Mohammad Norouzi 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Kubric: A scalable dataset generatorabstractData is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification. Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J. Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam H. Laradji, Hsueh-Ti Derek Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, A. Cengiz Öztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Tianhao Wu 0003, Kwang Moo Yi, Fangcheng Zhong, Andrea Tagliasacchi |
CVPR | 7 |
| 2022 | Disentangling Architecture and Training for Optical Flow
Deqing Sun, Charles Herrmann, Fitsum A. Reda, Michael Rubinstein, David J. Fleet, William T. Freeman |
ECCV (22) | 5 |
| 2022 | Pix2seq: A Language Modeling Framework for Object Detection
Ting Chen 0007, Saurabh Saxena, Lala Li, David J. Fleet, Geoffrey E. Hinton |
ICLR | 4 |
| 2022 | A Unified Sequence Interface for Vision TasksabstractWhile language tasks are naturally expressed in a single, unified, modeling framework, i.e., generating sequences of tokens, this has not been the case in computer vision. As a result, there is a proliferation of distinct architectures and loss functions for different vision tasks. In this work we show that a diverse set of "core" computer vision tasks can also be unified if formulated in terms of a shared pixel-to-sequence interface. We focus on four tasks, namely, object detection, instance segmentation, keypoint detection, and image captioning, all with diverse types of outputs, e.g., bounding boxes or dense masks. Despite that, by formulating the output of each task as a sequence of discrete tokens with a unified interface, we show that one can train a neural network with a single model architecture and loss function on all these tasks, with no task-specific customization. To solve a specific task, we use a short prompt as task description, and the sequence output adapts to the prompt so it can produce task-specific output. We show that such a model can achieve competitive performance compared to well-established task-specific models. Ting Chen 0007, Saurabh Saxena, Lala Li, Tsung-Yi Lin, David J. Fleet, Geoffrey E. Hinton |
NeurIPS | 5 |
| 2022 | Video Diffusion ModelsabstractGenerating temporally coherent high fidelity video is an important milestone in generative modeling research. We make progress towards this milestone by proposing a diffusion model for video generation that shows very promising initial results. Our model is a natural extension of the standard image diffusion architecture, and it enables jointly training from image and video data, which we find to reduce the variance of minibatch gradients and speed up optimization. To generate long and higher resolution videos we introduce a new conditional sampling technique for spatial and temporal video extension that performs better than previously proposed methods. We present the first results on a large text-conditioned video generation task, as well as state-of-the-art results on established benchmarks for video prediction and unconditional video generation. Supplementary material is available at https://video-diffusion.github.io/. Jonathan Ho, Tim Salimans, Alexey A. Gritsenko, Mohammad Norouzi 0002, David J. Fleet |
NeurIPS | 6 |
| 2022 | Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingabstractWe present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g., T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment. Chitwan Saharia, Saurabh Saxena, Lala Li, Jay Whang, Remi Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi 0002 |
NeurIPS | 12 |
| 2022 | Residual Multiplicative Filter Networks for Multiscale ReconstructionabstractCoordinate networks like Multiplicative Filter Networks (MFNs) and BACON offer some control over the frequency spectrum used to represent continuous signals such as images or 3D volumes. Yet, they are not readily applicable to problems for which coarse-to-fine estimation is required, including various inverse problems in which coarse-to-fine optimization plays a key role in avoiding poor local minima. We introduce a new coordinate network architecture and training scheme that enables coarse-to-fine optimization with fine-grained control over the frequency support of learned reconstructions. This is achieved with two key innovations. First, we incorporate skip connections so that structure at one scale is preserved when fitting finer-scale structure. Second, we propose a novel initialization scheme to provide control over the model frequency spectrum at each stage of optimization. We demonstrate how these modifications enable multiscale optimization for coarse-to-fine fitting to natural images. We then evaluate our model on synthetically generated datasets for the the problem of single-particle cryo-EM reconstruction. We learn high resolution multiscale structures, on par with the state-of-the art. Project webpage: https://shekshaa.github.io/ResidualMFN/. Shayan Shekarforoush, David B. Lindell, David J. Fleet, Marcus A. Brubaker |
NeurIPS | 3 |
| 2022 | Cascaded Diffusion Models for High Fidelity Image GenerationabstractWe show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample quality. A cascaded diffusion model comprises a pipeline of multiple diffusion models that generate images of increasing resolution, beginning with a standard diffusion model at the lowest resolution, followed by one or more super-resolution diffusion models that successively upsample the image and add higher resolution details. We find that the sample quality of a cascading pipeline relies crucially on conditioning augmentation, our proposed method of data augmentation of the lower resolution conditioning inputs to the super-resolution models. Our experiments show that conditioning augmentation prevents compounding error during sampling in a cascaded model, helping us to train cascading pipelines achieving FID scores of 1.48 at 64x64, 3.52 at 128x128 and 4.88 at 256x256 resolutions, outperforming BigGAN-deep, and classification accuracy scores of 63.02% (top-1) and 84.06% (top-5) at 256x256, outperforming VQ-VAE-2. Jonathan Ho, Chitwan Saharia, David J. Fleet, Mohammad Norouzi 0002, Tim Salimans |
J. Mach. Learn. Res. | 4 |
| 2021 | Unsupervised Part Representation by Flow CapsulesabstractCapsule networks aim to parse images into a hierarchy of objects, parts and relations. While promising, they remain limited by an inability to learn effective low level part descriptions. To address this issue we propose a way to learn primary capsule encoders that detect atomic parts from a single image. During training we exploit motion as a powerful perceptual cue for part definition, with an expressive decoder for part generation within a layered image model with occlusion. Experiments demonstrate robust part discovery in the presence of multiple objects, cluttered backgrounds, and occlusion. The learned part decoder is shown to infer the underlying shape masks, effectively filling in occluded regions of the detected shapes. We evaluate FlowCapsules on unsupervised part segmentation and unsupervised image classification. Sara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E. Hinton, David J. Fleet |
ICML | 5 |
| 2020 | Exemplar VAE: Linking Generative Models, Nearest Neighbor Retrieval, and Data AugmentationabstractWe introduce Exemplar VAEs, a family of generative models that bridge the gap between parametric and non-parametric, exemplar based generative models. Exemplar VAE is a variant of VAE with a non-parametric latent prior based on a Parzen window estimator. To sample from it, one first draws a random exemplar from a training set, then stochastically transforms that exemplar into a latent code and a new observation. We propose retrieval augmented training (RAT) as a way to speed up Exemplar VAE training by using approximate nearest neighbor search in the latent space to define a lower bound on log marginal likelihood. To enhance generalization, model parameters are learned using exemplar leave-one-out and subsampling. Experiments demonstrate the effectiveness of Exemplar VAEs on density estimation and representation learning. Importantly, generative data augmentation using Exemplar VAEs on permutation invariant MNIST and Fashion MNIST reduces classification error from 1.17% to 0.69% and from 8.56% to 8.16%. Sajad Norouzi, David J. Fleet, Mohammad Norouzi 0002 |
NeurIPS | 2 |
| 2019 | Differentiable Probabilistic Models of Scientific Imaging with the Fourier Slice Theorem
Karen Ullrich, Rianne van den Berg, Marcus A. Brubaker, David J. Fleet, Max Welling |
UAI | 4 |
| 2018 | VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Fartash Faghri, David J. Fleet, Jamie Kiros, Sanja Fidler |
BMVC | 2 |
| 2017 | Subspace selection to suppress confounding source domain information in AAM transfer learningabstractActive appearance models (AAMs) have seen tremendous success in face analysis. However, model learning depends on the availability of detailed annotation of canonical landmark points. As a result, when accurate AAM fitting is required on a different set of variations (expression, pose, identity), a new dataset is collected and annotated. To overcome the need for time consuming data collection and annotation, transfer learning approaches have received recent attention. The goal is to transfer knowledge from previously available datasets (source) to a new dataset (target). We propose a subspace transfer learning method, in which we select a subspace from the source that best describes the target space. We propose a metric to compute the directional similarity between the source eigenvectors and the target subspace. We show an equivalence between this metric and the variance of target data when projected onto source eigenvectors. Using this equivalence, we select a subset of source principal directions that capture the variance in target data. To define our model, we augment the selected source subspace with the target subspace learned from a handful of target examples. In experiments done on six public datasets, we show that our approach outperforms the state of the art in terms of the RMS fitting error as well as the percentage of test examples for which AAM fitting converges to the ground truth. Azin Asgarian, Ahmed Ashraf 0001, David J. Fleet, Babak Taati |
IJCB | 3 |
| 2017 | Building Proteins in a Day: Efficient 3D Molecular Structure Estimation with Electron CryomicroscopyabstractDiscovering the 3D atomic-resolution structure of molecules such as proteins and viruses is one of the foremost research problems in biology and medicine. Electron Cryomicroscopy (cryo-EM) is a promising vision-based technique for structure estimation which attempts to reconstruct 3D atomic structures from a large set of 2D transmission electron microscope images. This paper presents a new Bayesian framework for cryo-EM structure estimation that builds on modern stochastic optimization techniques to allow one to scale to very large datasets. We also introduce a novel Monte-Carlo technique that reduces the cost of evaluating the objective function during optimization by over five orders of magnitude. The net result is an approach capable of estimating 3D molecular structure from large-scale datasets in about a day on a single CPU workstation. Ali Punjani, Marcus A. Brubaker, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Guest Editorial: Human Activity Understanding from 2D and 3D Data
Junsong Yuan 0001, Wanqing Li 0001, Zhengyou Zhang, David J. Fleet, Jamie Shotton |
Int. J. Comput. Vis. | 4 |
| 2015 | Building proteins in a day: Efficient 3D molecular reconstructionabstractDiscovering the 3D atomic structure of molecules such as proteins and viruses is a fundamental research problem in biology and medicine. Electron Cryomicroscopy (Cryo-EM) is a promising vision-based technique for structure estimation which attempts to reconstruct 3D structures from 2D images. This paper addresses the challenging problem of 3D reconstruction from 2D Cryo-EM images. A new framework for estimation is introduced which relies on modern stochastic optimization techniques to scale to large datasets. We also introduce a novel technique which reduces the cost of evaluating the objective function during optimization by over five orders or magnitude. The net result is an approach capable of estimating 3D molecular structure from large scale datasets in about a day on a single workstation. Marcus A. Brubaker, Ali Punjani, David J. Fleet |
CVPR | 3 |
| 2015 | Efficient Non-greedy Optimization of Decision TreesabstractDecision trees and randomized forests are widely used in computer vision and machine learning. Standard algorithms for decision tree induction optimize the split functions one node at a time according to some splitting criteria. This greedy procedure often leads to suboptimal trees. In this paper, we present an algorithm for optimizing the split functions at all levels of the tree jointly with the leaf parameters, based on a global objective. We show that the problem of finding optimal linear-combination (oblique) splits for decision trees is related to structured prediction with latent variables, and we formulate a convex-concave upper bound on the tree's empirical loss. Computing the gradient of the proposed surrogate objective with respect to each training exemplar is O(d^2), where d is the tree depth, and thus training deep trees is feasible. The use of stochastic gradient descent for optimization enables effective training with large datasets. Experiments on several classification benchmarks demonstrate that the resulting non-greedy decision trees outperform greedy decision tree baselines. Mohammad Norouzi 0002, Maxwell D. Collins, Matthew Johnson 0003, David J. Fleet, Pushmeet Kohli |
NIPS | 4 |
| 2015 | Editorial
Björn Stenger, Norimichi Ukita, Yoichi Sato 0001, Pascal Fua, David J. Fleet |
Comput. Vis. Image Underst. | 5 |
| 2015 | Efficient Optimization for Sparse Gaussian Process RegressionabstractWe propose an efficient optimization algorithm to select a subset of training data as the inducing set for sparse Gaussian process regression. Previous methods either use different objective functions for inducing set and hyperparameter selection, or else optimize the inducing set by gradient-based continuous optimization. The former approaches are harder to interpret and suboptimal, whereas the latter cannot be applied to discrete input domains or to kernel functions that are not differentiable with respect to the input. The algorithm proposed in this work estimates an inducing set and the hyperparameters using a single objective. It can be used to optimize either the marginal likelihood or a variational free energy. Space and time complexity are linear in training set size, and the algorithm can be applied to large regression problems on discrete or continuous domains. Empirical evaluation shows state-of-art performance in discrete cases, competitive prediction results as well as a favorable trade-off between training and test time in continuous cases. Yanshuai Cao, Marcus A. Brubaker, David J. Fleet, Aaron Hertzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Posebits for Monocular Human Pose EstimationabstractWe advocate the inference of qualitative information about 3D human pose, called posebits, from images. Posebits represent Boolean geometric relationships between body parts (e.g., left-leg in front of right-leg or hands close to each other). The advantages of posebits as a mid-level representation are 1) for many tasks of interest, such qualitative pose information may be sufficient (e.g., semantic image retrieval), 2) it is relatively easy to annotate large image corpora with posebits, as it simply requires answers to yes/no questions, and 3) they help resolve challenging pose ambiguities and therefore facilitate the difficult talk of image-based 3D pose estimation. We introduce posebits, a posebit database, a method for selecting useful posebits for pose estimation and a structural SVM model for posebit inference. Experiments show the use of posebits for semantic image retrieval and for improving 3D pose estimation. Gerard Pons-Moll, David J. Fleet, Bodo Rosenhahn |
CVPR | 2 |
| 2014 | Fast Exact Search in Hamming Space With Multi-Index HashingabstractThere is growing interest in representing image data and feature descriptors using compact binary codes for fast near neighbor search. Although binary codes are motivated by their use as direct indices (addresses) into a hash table, codes longer than 32 bits are not being used as such, as it was thought to be ineffective. We introduce a rigorous way to build multiple hash tables on binary code substrings that enables exact k-nearest neighbor search in Hamming space. The approach is storage efficient and straight-forward to implement. Theoretical analysis shows that the algorithm exhibits sub-linear run-time behavior for uniformly distributed codes. Empirical results show dramatic speedups over a linear scan baseline for datasets of up to one billion codes of 64, 128, or 256 bits. Mohammad Norouzi 0002, Ali Punjani, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Cartesian K-MeansabstractA fundamental limitation of quantization techniques like the k-means clustering algorithm is the storage and run-time cost associated with the large numbers of clusters required to keep quantization errors small and model fidelity high. We develop new models with a compositional parameterization of cluster centers, so representational capacity increases super-linearly in the number of parameters. This allows one to effectively quantize data using billions or trillions of centers. We formulate two such models, Orthogonal k-means and Cartesian k-means. They are closely related to one another, to k-means, to methods for binary hash function optimization like ITQ (Gong and Lazebnik, 2011), and to Product Quantization for vector quantization (Jegou et al., 2011). The models are tested on large-scale ANN retrieval tasks (1M GIST, 1B SIFT features), and on codebook learning for object recognition (CIFAR-10). Mohammad Norouzi 0002, David J. Fleet |
CVPR | 2 |
| 2013 | Efficient Optimization for Sparse Gaussian Process RegressionabstractWe propose an efficient discrete optimization algorithm for selecting a subset of training data to induce sparsity for Gaussian process regression. The algorithm estimates this inducing set and the hyperparameters using a single objective, either the marginal likelihood or a variational free energy. The space and time complexity are linear in the training set size, and the algorithm can be applied to large regression problems on discrete or continuous domains. Empirical evaluation shows state-of-art performance in the discrete case and competitive results in the continuous case. Yanshuai Cao, Marcus A. Brubaker, David J. Fleet, Aaron Hertzmann |
NIPS | 3 |
| 2012 | Fast search in Hamming space with multi-index hashingabstractThere has been growing interest in mapping image data onto compact binary codes for fast near neighbor search in vision applications. Although binary codes are motivated by their use as direct indices (addresses) into a hash table, codes longer than 32 bits are not being used in this way, as it was thought to be ineffective. We introduce a rigorous way to build multiple hash tables on binary code substrings that enables exact K-nearest neighbor search in Hamming space. The algorithm is straightforward to implement, storage efficient, and it has sub-linear run-time behavior for uniformly distributed codes. Empirical results show dramatic speed-ups over a linear scan baseline and for datasets with up to one billion items, 64- or 128-bit codes, and search radii up to 25 bits. Mohammad Norouzi 0002, Ali Punjani, David J. Fleet |
CVPR | 3 |
| 2012 | Hamming Distance Metric LearningabstractMotivated by large-scale multimedia applications we propose to learn mappings from high-dimensional data to binary codes that preserve semantic similarity. Binary codes are well suited to large-scale applications as they are storage efficient and permit exact sub-linear kNN search. The framework is applicable to broad families of mappings, and uses a flexible form of triplet ranking loss. We overcome discontinuous optimization of the discrete mappings by minimizing a piecewise-smooth upper bound on empirical loss, inspired by latent structural SVMs. We develop a new loss-augmented inference algorithm that is quadratic in the code length. We show strong retrieval performance on CIFAR-10 and MNIST, with promising classification results using no more than kNN on the binary codes. Mohammad Norouzi 0002, David J. Fleet, Ruslan Salakhutdinov |
NIPS | 2 |
| 2012 | Human attributes from 3D pose tracking
Micha Livne, Leonid Sigal, Nikolaus F. Troje, David J. Fleet |
Comput. Vis. Image Underst. | 4 |
| 2012 | Shared Kernel Information Embedding for Discriminative InferenceabstractLatent variable models, such as the GPLVM and related methods, help mitigate overfitting when learning from small or moderately sized training sets. Nevertheless, existing methods suffer from several problems: 1) complexity, 2) the lack of explicit mappings to and from the latent space, 3) an inability to cope with multimodality, and 4) the lack of a well-defined density over the latent space. We propose an LVM called the Kernel Information Embedding (KIE) that defines a coherent joint density over the input and a learned latent space. Learning is quadratic, and it works well on small data sets. We also introduce a generalization, the shared KIE (sKIE), that allows us to model multiple input spaces (e.g., image features and poses) using a single, shared latent representation. KIE and sKIE permit missing data during inference and partially labeled data during learning. We show that with data sets too large to learn a coherent global model, one can use the sKIE to learn local online models. We use sKIE for human pose inference. Roland Memisevic, Leonid Sigal, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Minimal Loss Hashing for Compact Binary Codes
Mohammad Norouzi 0002, David J. Fleet |
ICML | 2 |
| 2011 | Simultaneous Tracking and Activity RecognitionabstractMany tracking problems involve several distinct objects interacting with each other. We develop a framework that takes into account interactions between objects allowing the recognition of complex activities. In contrast to classic approaches that consider distinct phases of tracking and activity recognition, our framework performs these two tasks simultaneously. In particular, we adopt a Bayesian standpoint where the system maintains a joint distribution of the positions, the interactions and the possible activities. This turns out to be advantegeous, as information about the ongoing activities can be used to improve the prediction step of the tracking, while, at the same time, tracking information can be used for online activity recognition. Experimental results in two different settings show that our approach 1) decreases the error rate and improves the identity maintenance of the positional tracking and 2) identifies the correct activity with higher accuracy than standard approaches. Cristina E. Manfredotti, David J. Fleet, Howard J. Hamilton, Sandra Zilles |
ICTAI | 2 |
| 2011 | Bone graphs: Medial shape parsing and abstraction
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi |
Comput. Vis. Image Underst. | 3 |
| 2011 | Object categorization using bone graphs
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi |
Comput. Vis. Image Underst. | 3 |
| 2011 | Model-Based 3D Hand Pose Estimation from Monocular VideoabstractA novel model-based approach to 3D hand tracking from monocular video is presented. The 3D hand pose, the hand texture, and the illuminant are dynamically estimated through minimization of an objective function. Derived from an inverse problem formulation, the objective function enables explicit use of temporal texture continuity and shading information while handling important self-occlusions and time-varying illumination. The minimization is done efficiently using a quasi-Newton method, for which we provide a rigorous derivation of the objective function gradient. Particular attention is given to terms related to the change of visibility near self-occlusion boundaries that are neglected in existing formulations. To this end, we introduce new occlusion forces and show that using all gradient terms greatly improves the performance of the method. Qualitative and quantitative experimental results demonstrate the potential of the approach. Martin de La Gorce, David J. Fleet, Nikos Paragios |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Dynamical binary latent variable models for 3D human pose trackingabstractWe introduce a new class of probabilistic latent variable model called the Implicit Mixture of Conditional Restricted Boltzmann Machines (imCRBM) for use in human pose tracking. Key properties of the imCRBM are as follows: (1) learning is linear in the number of training exemplars so it can be learned from large datasets; (2) it learns coherent models of multiple activities; (3) it automatically discovers atomic “movemes” and (4) it can infer transitions between activities, even when such transitions are not present in the training set. We describe the model and how it is learned and we demonstrate its use in the context of Bayesian filtering for multi-view and monocular pose tracking. The model handles difficult scenarios including multiple activities and transitions among activities. We report state-of-the-art results on the HumanEva dataset. Graham W. Taylor, Leonid Sigal, David J. Fleet, Geoffrey E. Hinton |
CVPR | 3 |
| 2010 | Human Attributes from 3D Pose Tracking
Leonid Sigal, David J. Fleet, Nikolaus F. Troje, Micha Livne |
ECCV (3) | 2 |
| 2010 | Physics-Based Person Tracking Using the Anthropomorphic Walker
Marcus A. Brubaker, David J. Fleet, Aaron Hertzmann |
Int. J. Comput. Vis. | 2 |
| 2010 | Optimizing walking controllers for uncertain inputs and environmentsabstractWe introduce methods for optimizing physics-based walking controllers for robustness to uncertainty. Many unknown factors, such as external forces, control torques, and user control inputs, cannot be known in advance and must be treated as uncertain. These variables are represented with probability distributions, and a return function scores the desirability of a single motion. Controller optimization entails maximizing the expected value of the return, which is computed by Monte Carlo methods. We demonstrate examples with different sources of uncertainty and task constraints. Optimizing control strategies under uncertainty increases robustness and produces natural variations in style. Jack M. Wang, David J. Fleet, Aaron Hertzmann |
ACM Trans. Graph. | 2 |
| 2009 | Backing Off: Hierarchical Decomposition of Activity for 3D Novel Pose RecoveryabstractFor model-based 3D human pose estimation, even simple models of the human body lead to high-dimensional state spaces. Where the class of activity is known a priori, low-dimensional activity models learned from training data make possible a thorough and efficient search for the best pose. Conversely, searching for solutions in the full state space places no restriction on the class of motion to be recovered, but is both difficult and expensive. This paper explores a potential middle ground between these approaches, using the hierarchical Gaussian process latent variable model to learn activity at different hierarchical scales within the human skeleton. We show that by training on full-body activity data then descending through the hierarchy in stages and exploring subtrees independently of one another, novel poses may be recovered. Experimental results on motion capture data and monocular video sequences demonstrate the utility of the approach, and comparisons are drawn with existing low-dimensional activity models. © 2009. The copyright of this document resides with its authors. John Darby, Baihua Li, Nicholas Costen, David J. Fleet, Neil D. Lawrence |
BMVC | 4 |
| 2009 | Stochastic Image DenoisingabstractWe present a novel algorithm for image denoising. Our algorithm is based on random walks over arbitrary neighbourhoods surrounding a given pixel. The size and shape of each neighbourhood are determined by the configuration and similarity of nearby pixels. Assuming that pixels within the neighbourhood of x0 are likely to have been generated by the same random process, we want the weights used to mix these pixels during denoising to depend on the similarity between them and x0. At the same time, we require the random walk to follow a smooth path from x0 to any other pixel in the neighbourhood, so the transition probabilities should also depend on the similarity between pairs of neighbouring pixels along any given path. With this in mind, we define a random walk originating at pixel x0 as an ordered sequence of pixels T0,k = {x0,x1, . . . ,xk} visited along the path from x0 to xk. Within this sequence, the probability of a transition between two consecutive pixels x j and x j+1 is defined to be Francisco J. Estrada, David J. Fleet, Allan Douglas Jepson |
BMVC | 2 |
| 2009 | Shared Kernel Information Embedding for discriminative inferenceabstractLatent variable models (LVM), like the shared-GPLVM and the spectral latent variable model, help mitigate over-fitting when learning discriminative methods from small or moderately sized training sets. Nevertheless, existing methods suffer from several problems: (1) complexity; (2) the lack of explicit mappings to and from the latent space; (3) an inability to cope with multi-modality; and (4) the lack of a well-defined density over the latent space. We propose a LVM called the shared kernel information embedding (sKIE). It defines a coherent density over a latent space and multiple input/output spaces (e.g., image features and poses), and it is easy to condition on a latent state, or on combinations of the input/output states. Learning is quadratic, and it works well on small datasets. With datasets too large to learn a coherent global model, one can use sKIE to learn local online models. sKIE permits missing data during inference, and partially labelled data during learning. We use sKIE for human pose inference. Leonid Sigal, Roland Memisevic, David J. Fleet |
CVPR | 3 |
| 2009 | Estimating contact dynamicsabstractMotion and interaction with the environment are fundamentally intertwined. Few people-tracking algorithms exploit such interactions, and those that do assume that surface geometry and dynamics are given. This paper concerns the converse problem, i.e., the inference of contact and environment properties from motion. For 3D human motion, with a 12-segment articulated body model, we show how one can estimate the forces acting on the body in terms of internal forces (joint torques), gravity, and the parameters of a contact model (e.g., the geometry and dynamics of a spring-based model). This is tested on motion capture data and video-based tracking data, with walking, jogging, cartwheels, and jumping. Marcus A. Brubaker, Leonid Sigal, David J. Fleet |
ICCV | 3 |
| 2009 | TurboPixels: Fast Superpixels Using Geometric FlowsabstractWe describe a geometric-flow-based algorithm for computing a dense oversegmentation of an image, often referred to as superpixels. It produces segments that, on one hand, respect local image boundaries, while, on the other hand, limiting undersegmentation through a compactness constraint. It is very fast, with complexity that is approximately linear in image size, and can be applied to megapixel sized images with high superpixel densities in a matter of minutes. We show qualitative demonstrations of high-quality results on several complex images. The Berkeley database is used to quantitatively compare its performance to a number of oversegmentation algorithms, showing that it yields less undersegmentation than algorithms that lack a compactness constraint while offering a significant speedup over N-cuts, which does enforce compactness. Alex Levinshtein, Adrian Stere, Kiriakos N. Kutulakos, David J. Fleet, Sven J. Dickinson, Kaleem Siddiqi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2009 | Optimizing walking controllersabstractThis paper describes a method for optimizing the parameters of a physics-based controller for full-body, 3D walking. A modified version of the SIMBICON controller [Yin et al. 2007] is optimized for characters of varying body shape, walking speed and step length. The objective function includes terms for power minimization, angular momentum minimization, and minimal head motion, among others. Together these terms produce a number of important features of natural walking, including active toe-off, near-passive knee swing, and leg extension during swing. We explain the specific form of our objective criteria, and show the importance of each term to walking style. We demonstrate optimized controllers for walking with different speeds, variation in body shape, and in ground slope. Jack M. Wang, David J. Fleet, Aaron Hertzmann |
ACM Trans. Graph. | 2 |
| 2008 | The Kneed Walker for human pose trackingabstractThe Kneed Walker is a physics-based model derived from a planar biomechanical characterization of human locomotion. By controlling torques at the knees, hips and torso, the model captures a full range of walking motions with foot contact and balance. Constraints are used to properly handle ground collisions and joint limits. A prior density over walking motions is based on dynamics that are optimized for efficient cyclic gaits over a wide range of natural human walking speeds and step lengths, on different slopes. The generative model used for monocular tracking comprises the Kneed Walker prior, a 3D kinematic model constrained to be consistent with the underlying dynamics, and a simple measurement model in terms of appearance and optical flow. The tracker is applied to people walking with varying speeds, on hills, and with occlusion. Marcus A. Brubaker, David J. Fleet |
CVPR | 2 |
| 2008 | Model-based hand tracking with texture, shading and self-occlusionsabstractA novel model-based approach to 3D hand tracking from monocular video is presented. The 3D hand pose, the hand texture and the illuminant are dynamically estimated through minimization of an objective function. Derived from an inverse problem formulation, the objective function enables explicit use of texture temporal continuity and shading information, while handling important self-occlusions and time-varying illumination. The minimization is done efficiently using a quasi-Newton method, for which we propose a rigorous derivation of the objective function gradient. Particular attention is given to terms related to the change of visibility near self-occlusion boundaries that are neglected in existing formulations. In doing so we introduce new occlusion forces and show that using all gradient terms greatly improves the performance of the method. Experimental results demonstrate the potential of the formulation. Martin de La Gorce, Nikos Paragios, David J. Fleet |
CVPR | 3 |
| 2008 | Topologically-constrained latent variable modelsabstractIn dimensionality reduction approaches, the data are typically embedded in a Euclidean latent space. However for some data sets this is inappropriate. For example, in human motion data we expect latent spaces that are cylindrical or a toroidal, that are poorly captured with a Euclidean space. In this paper, we present a range of approaches for embedding data in a non-Euclidean latent space. Our focus is the Gaussian Process latent variable model. In the context of human motion modeling this allows us to (a) learn models with interpretable latent directions enabling, for example, style/content separation, and (b) generalise beyond the data set enabling us to learn transitions between motion styles even though such transitions are not present in the data. Raquel Urtasun, David J. Fleet, Andreas Geiger 0001, Jovan Popovic, Trevor Darrell, Neil D. Lawrence |
ICML | 2 |
| 2008 | Introduction of New EditorsabstractO support the continuing increase in submission arising from the popularity of TPAMI with authors, we are pleased to announce that Professor Zoubin Ghahramani will be joining David Fleet as an Associate Editor-in-Chief (AEIC) of TPAMI. He will help to maintain the high quality that readers expect and the timeliness and thoroughness of reviewing that authors’ demand. The AEIC works hand-in-hand with the Editor-in-Chief in all aspects of TPAMI’s editorial review process, including establishing policies, selecting special issues, selecting editors, helping the fl ow of papers through the review process, handling appeals, etc. Dr. Ghahramani is a leader in the fi eld of machine learning, an area of increasing importance to TPAMI, for both the quality of his research and his service to the community. We are also happy to announce that the TPAMI editorial board is expanding with the addition of four new Associate Editors, Dr. Sing Bing Kang, Professor Kevin Murphy. Dr. Salil Prabhakar, and Professor Dale Schuurmans. Dr. Kang will handle papers on image-based rendering, vision for graphics, image and video processing/enhancement, and stereopsis and structure from motion. Professor Murphy will oversee papers about graphical models, structure learning, causal inference, unsupervised learning, and Bayesian methods. Dr. Prabhakar will oversee the review process of papers in all aspects of biometric systems and theoretical and empirical evaluation of computer vision algorithms. Professor Schuurmans’ expertise includes machine learning techniques (support vector machines, kernel methods, clustering, dimensionality reduction), graphical models, optimization, and Monte Carlo methods, and he has begun to handle papers in these areas. Their brief biographies are below. Welcome to TPAMI’s editorial board! David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Editorial-State of the Transactions
David J. Kriegman, David J. Fleet, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Gaussian Process Dynamical Models for Human MotionabstractWe introduce Gaussian process dynamical models (GPDM) for nonlinear time series analysis, with applications to learning models of human pose and motion from high-dimensionalmotion capture data. A GPDM is a latent variable model. It comprises a low-dimensional latent space with associated dynamics, and a map from the latent space to an observation space. We marginalize out the model parameters in closed-form, using Gaussian process priors for both the dynamics and the observation mappings. This results in a non-parametric model for dynamical systems that accounts for uncertainty in the model. We demonstrate the approach, and compare four learning algorithms on human motion capture data in which each pose is 50-dimensional. Despite the use of small data sets, the GPDM learns an effective representation of the nonlinear dynamics in these spaces. Jack M. Wang, David J. Fleet, Aaron Hertzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Correction to "Gaussian Process Dynamical Models for Human Motion"abstractIn the above titled paper (ibid., vol. 30, no. 2, pp. 283-298, Feb 08), two figures were misprinted. The correct figures are presented here. Jack M. Wang, David J. Fleet, Aaron Hertzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Higher-order Autoregressive Models for Dynamic TexturesabstractDynamic textured sequences are characterized by the interactions between many particles or objects in the scene. Based on earlier work the images of the sequence are interpreted as the output of a linear autoregressive process driven by white Gaussian noise. We extend earlier work by increasing the amount temporal information included when learning the motion in the scene, allowing the models to capture complex motion patterns which extend over multiple frames, thereby increasing the perceptual accuracy of the synthesized results. To overcome problems of dynamic model stability, we apply Burg’s Maximum Entropy Spectral Analysis technique for parameter estimation, which is found to be reliably stable on smaller samples of training data, even with higher-order dynamics. 1 Midori Hyndman, Allan Douglas Jepson, David J. Fleet |
BMVC | 3 |
| 2007 | Physics-Based Person Tracking Using Simplified Lower-Body DynamicsabstractWe introduce a physics-based model for 3D person tracking. Based on a biomechanical characterization of lower-body dynamics, the model captures important physical properties of bipedal locomotion such as balance and ground contact, generalizes naturally to variations in style due to changes in speed, step-length, and mass, and avoids common problems such as footskate that arise with existing trackers. The model dynamics comprises a two degree-of-freedom representation of human locomotion with inelastic ground contact. A stochastic controller generates impulsive forces during the toe-off stage of walking and spring-like forces between the legs. A higher-dimensional kinematic observation model is then conditioned on the underlying dynamics. We use the model for tracking walking people from video, including examples with turning, occlusion, and varying gait. Marcus A. Brubaker, David J. Fleet, Aaron Hertzmann |
CVPR | 2 |
| 2007 | Multifactor Gaussian process models for style-content separationabstractWe introduce models for density estimation with multiple, hidden, continuous factors. In particular, we propose a generalization of multilinear models using nonlinear basis functions. By marginalizing over the weights, we obtain a multifactor form of the Gaussian process latent variable model. In this model, each factor is kernelized independently, allowing nonlinear mappings from any particular factor to the data. We learn models for human locomotion data, in which each pose is generated by factors representing the person's identity, gait, and the current state of motion. We demonstrate our approach using time-series prediction, and by synthesizing novel animation from the model. Jack M. Wang, David J. Fleet, Aaron Hertzmann |
ICML | 2 |
| 2007 | Special issue on spatial coherence for visual motion analysis
W. James MacLean, Nikos Paragios, David J. Fleet |
Comput. Vis. Image Underst. | 3 |
| 2007 | Editorial-State of the Transactions
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | 3D People Tracking with Gaussian Process Dynamical ModelsabstractWe advocate the use of Gaussian Process Dynamical Models (GPDMs) for learning human pose and motion priors for 3D people tracking. A GPDM provides a lowdimensional embedding of human motion data, with a density function that gives higher probability to poses and motions close to the training data. With Bayesian model averaging a GPDM can be learned from relatively small amounts of data, and it generalizes gracefully to motions outside the training set. Here we modify the GPDM to permit learning from motions with significant stylistic variation. The resulting priors are effective for tracking a range of human walking styles, despite weak and noisy image measurements and significant occlusions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (1) | 2 |
| 2006 | Temporal motion models for monocular and multiview 3D human body tracking
Raquel Urtasun, David J. Fleet, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2006 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Editorial-State of the TransactionsabstractIT was another good year for the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)—a year with more and better papers and a more competitive publication environment. It was also a year of some challenges. There are two major sets of metrics that are commonly used to evaluate the performance of a journal: citation statistics, which assess the significance of the papers published, and editorial statistics, which characterize the flow of papers through the editorial process. Certainly, the former is connected to the latter. A vibrant and healthy journal starts with authors submitting their most important papers. Between conferences and journals, authors have a great deal of choice about where to publish. They want a prestigious journal where their papers will receive a fair review in a timely manner. Starting with citation statistics, the most widely used measure is the impact factor which is the average number of times papers published in the two previous years referenced by journals in the given year. Based on the publications tracked by Thomson ISI, TPAMI papers were cited 14,708 times in 2006 and the impact factor in 2006 was 4.3, up from 3.8 in 2005. TPAMI is the IEEE’s most cited publication, the second most cited journal in all of electrical engineering, and the fifth most cited journal in all of computer science. In 2006, TPAMI received 912 submissions, up from 749 the year before. The acceptance rate is about 25 percent. The peer review cycle has been improving and the average time from submission to first decision is three months, and the average time from submission to acceptance is nine months. It is worth noting how favorably the submission to first decision time compares to that of conferences. The reason for the length of time between submission and acceptance is that many papers undergo a major and a minor revision and, consequently, require additional reviewing. There is an asymmetry between accepted and rejected papers since those that are rejected (the majority) leave the review process more quickly. TPAMI posts accepted papers digitally in the CS Digital Library and IEEE Xplore in advance of the printed version, providing early access to readers. Once papers are posted online, they are considered published and with this new posting online upon acceptance, submission to publication has been dramatically reduced. In 2007, TPAMI published a very successful special issue on Progress and Directions in Biometrics, edited by Josef Kittler, Davide Maltoni, Lawrence O’Gorman, Salil Prabhakar, and Tieniu Tan. The issue received 85 submissions, of which 19 were accepted. Handling this many papers was a daunting task for the guest editors. Looking toward 2008, there are two special issues in the works: First, we are pleased to announce that TPAMI will be having a special section devoted to the award-winning papers from the 2007 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). There have always been very strong ties between the major conferences in our field and TPAMI, and it was high time for the best papers at the conferences to be recognized by TPAMI. Please note that these papers are still fully reviewed by TPAMI. Second, papers are under review for a special issue on Real-World Image Annotation and Retrieval. This issue points to the large scale applications that are becoming possible as the methods of computer vision and machine learning become more effective and robust. The year (2007) started with a conversion of the Webbased software for tracking manuscripts through peer review to a new system (Manuscript Central v3.0). The new version has a backend which is much more secure and a user interface which has been updated to provide more information to authors, reviewers, and editors and a more contemporary look and feel. While security, unless violated, is invisible to the user, the new interface is inescapable—the response has been mixed. While many bugs were fixed and most of the more serious issues have been addressed, there is still much to be done. We apologize to all who were inconvenienced. We would like to thank everyone who helps to make TPAMI a great journal, starting with the authors who submit their best works to TPAMI. The largest group is composed of the reviewers who collectively provide about 2,500 each year. Reviewing can be very rewarding, but it is also a heavy commitment and we appreciate the efforts and time these reviewers put into TPAMI and the community for which they are volunteering. For seasoned researchers, reviewing is a chance to see research results before they hit the press—a sneak preview—and to provide a valuable service to the community. For a novice reviewer, the first few dozen papers will highlight to the reviewer what makes a good versus a poor submission. Papers that are published are among 25 percent of submissions passing through the sieve and can be viewed as “positive training examples.” The other 75 percent of submissions are “negative training examples,” and these are only accessible to reviewers. The implication from a pattern classification perspective should be apparent to readers of this editorial. We would also like to thank the Associate Editors (AEs) who make a long term (four year) and ongoing commitment to the journal. The AEs are among the top researchers in the field, and their time is very valuable. They must make personal decisions about how much time they devote to IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 30, NO. 2, FEBRUARY 2008 193 David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Introduction of New Associate EditorsabstractThe EiC and Associate EiC express our gratitude to David Forsyth, Brendan Frey, Venu Govindaraju, and Cordelia Schmid who are retiring as associate editors of TPAMI. While we will miss their dedication to the transactions, we hope that they will be enjoying a bit more free time. We are also pleased to announce that Professor Daniel Lopresti, Professor B.S. Manjunath, Professor Marc Pollefeys, and Professor Ramin Zabih have joined the editorial board. Professor Lopresti will oversee papers in document and handwriting analysis, biometrics, approximate string matching algorithms, and performance evaluation. Professor Manjunath will be considering papers in feature extraction, segmentation, image/video retrieval, and image registration. Professor Pollefeys will be responsible for submissions in structure from motion, stereo, multiple view geometry and camera calibration, 3D and appearance modeling, shape-from-X techniques, and novel sensors. Professor Zabih will handle papers on stereo and medical imaging as well as energy minimization and graph algorithms. We look forward to working with them. Their brief biographies appear herein. David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Monocular 3-D Tracking of the Golf SwingabstractWe propose an approach to incorporating dynamic models into the human body tracking process that yields full 3D reconstructions from monocular sequences. We formulate the tracking problem in terms of minimizing a differentiable criterion whose differential structure is rich enough for successful optimization using a simple hill-climbing approach as opposed to a multihypotheses probabilistic one. In other words, we avoid the computational complexity of multihypotheses algorithms while obtaining excellent results under challenging conditions. To demonstrate this, we focus on monocular tracking of a golf swing from ordinary video. It involves both dealing with potentially very different swing styles, recovering arm motions that are perpendicular to the camera plane and handling strong self-occlusions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (2) | 2 |
| 2005 | Monocular 3D Tracking of the Golf SwingabstractThis paper showcases our tracking algorithm described in (R. Urtasun et al., 2005). The proposed approach incorporates dynamic motion models into the human body tracking process yielding full 3D reconstruction from monocular sequences. The tracking is formulated in terms of minimizing a differentiable criterion whose differential structure is rich enough for successful optimization using a single hypothesis. In other words we avoid the computational complexity of multihypotheses algorithms while obtaining excellent results under challenging conditions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (2) | 2 |
| 2005 | Priors for People Tracking from Small Training SetsabstractWe advocate the use of scaled Gaussian process latent variable models (SGPLVM) to learn prior models of 3D human pose for 3D people tracking. The SGPLVM simultaneously optimizes a low-dimensional embedding of the high-dimensional pose data and a density function that both gives higher probability to points close to training data and provides a nonlinear probabilistic mapping from the low-dimensional latent space to the full-dimensional pose space. The SGPLVM is a natural choice when only small amounts of training data are available. We demonstrate our approach with two distinct motions, golfing and walking. We show that the SGPLVM sufficiently constrains the problem such that tracking can be accomplished with straightforward deterministic optimization. Raquel Urtasun, David J. Fleet, Aaron Hertzmann, Pascal Fua |
ICCV | 2 |
| 2005 | Learning Sensor Network Topology through Monte Carlo Expectation MaximizationabstractWe consider the problem of inferring sensor positions and a topological (i.e. qualitative) map of an environment given a set of cameras with non-overlapping fields of view. In this way, without prior knowledge of the environment nor the exact position of sensors within the environment, one can infer the topology of the environment, and common traffic patterns within it. In particular, we consider sensors stationed at the junctions of the hallways of a large building. We infer the sensor connectivity graph and the travel times between sensors (and hence the hallway topology) from the sequence of events caused by unlabeled agents (i.e. people) passing within view of the different sensors. We do this based on a first-order semi-Markov model of the agent's behavior. The paper describes a problem formulation and proposes a stochastic algorithm for its solution. The result of the algorithm is a probabilistic model of the sensor network connectivity graph and the underlying traffic patterns. We conclude with results from numerical simulations Dimitri Marinakis, Gregory Dudek, David J. Fleet |
ICRA | 3 |
| 2005 | Gaussian Process Dynamical ModelsabstractThis paper introduces Gaussian Process Dynamical Models (GPDM) for nonlinear time series analysis. A GPDM comprises a low-dimensional latent space with associated dynamics, and a map from the latent space to an observation space. We marginalize out the model parameters in closed-form, using Gaussian Process (GP) priors for both the dynamics and the observation mappings. This results in a nonparametric model for dynamical systems that accounts for uncertainty in the model. We demonstrate the approach on human motion capture data in which each pose is 62-dimensional. Despite the use of small data sets, the GPDM learns an effective representation of the nonlinear dynamics in these spaces. Webpage: http://www.dgp.toronto.edu/ Jack M. Wang, David J. Fleet, Aaron Hertzmann |
NIPS | 2 |
| 2005 | Introduction of New Associate EditorsabstractThe EiC and Associate EiC express our gratitude to David Forsyth, Brendan Frey, Venu Govindaraju, and Cordelia Schmid who are retiring as associate editors of TPAMI. While we will miss their dedication to the transactions, we hope that they will be enjoying a bit more free time. We are also pleased to announce that Professor Daniel Lopresti, Professor B.S. Manjunath, Professor Marc Pollefeys, and Professor Ramin Zabih have joined the editorial board. Professor Lopresti will oversee papers in document and handwriting analysis, biometrics, approximate string matching algorithms, and performance evaluation. Professor Manjunath will be considering papers in feature extraction, segmentation, image/video retrieval, and image registration. Professor Pollefeys will be responsible for submissions in structure from motion, stereo, multiple view geometry and camera calibration, 3D and appearance modeling, shape-from-X techniques, and novel sensors. Professor Zabih will handle papers on stereo and medical imaging as well as energy minimization and graph algorithms. We look forward to working with them. Their brief biographies appear herein. David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Introduction of New Associate Editors
David J. Kriegman, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Perceptually-supported image editing of text and graphicsabstractThis extended abstract reprises our UIST '03 paper on "Perceptually-Supported Image Editing of Text and Graphics." We introduce a novel image editing program, called ScanScribe , that emphasizes easy selection and manipulation of material found in informal, casual documents such as sketches, handwritten notes, whiteboard images, screen snapshots, and scanned documents. Eric Saund, David J. Fleet, Daniel Larner, James Mahoney |
ACM Trans. Graph. | 2 |
| 2003 | Error-in-variables likelihood functions for motion estimationabstractOver-determined linear systems with noise in all measurements are common in computer vision, and particularly in motion estimation. Maximum likelihood estimators have been proposed to solve such problems, but except for simple cases, the corresponding likelihood functions are extremely complex, and accurate confidence measures do not exist. This paper derives the form of simple likelihood functions for such linear systems in the general case of heteroscedastic noise. We also derive a new algorithm for computing maximum likelihood solutions based on a modified Newton method. The new algorithm is more accurate, and exhibits more reliable convergence behavior than existing methods. We present an application to affine motion estimation, a simple heteroscedastic estimation problem. Oscar Nestares, David J. Fleet |
ICIP (3) | 2 |
| 2003 | Perceptually-supported image editing of text and graphicsabstractThis paper presents a novel image editing program emphasizing easy selection and manipulation of material found in informal, casual documents such as sketches, handwritten notes, whiteboard images, screen snapshots, and scanned documents. The program, called ScanScribe, offers four significant advances. First, it presents a new, intuitive model for maintaining image objects and groups, along with underlying logic for updating these in the course of an editing session. Second, ScanScribe takes advantage of newly developed image processing algorithms to separate foreground markings from a white or light background, and thus can automatically render the background transparent so that image material can be rearranged without occlusion by background pixels. Third, ScanScribe introduces new interface techniques for selecting image objects with a pointing device without resorting to a palette of tool modes. Fourth, ScanScribe presents a platform for exploiting image analysis and recognition methods to make perceptually significant structure readily available to the user. As a research prototype, ScanScribe has proven useful in the work of members of our laboratory, and has been released on a limited basis for user testing and evaluation. Eric Saund, David J. Fleet, Daniel Larner, James Mahoney |
UIST | 2 |
| 2003 | Robust Online Appearance Models for Visual TrackingabstractWe propose a framework for learning robust, adaptive, appearance models to be used for motion-based tracking of natural objects. The model adapts to slowly changing appearance, and it maintains a natural measure of the stability of the observed image structure during tracking. By identifying stable properties of appearance, we can weight them more heavily for motion estimation, while less stable properties can be proportionately downweighted. The appearance model involves a mixture of stable image structure, learned over long time courses, along with two-frame motion information and an outlier process. An online EM-algorithm is used to adapt the appearance model parameters over time. An implementation of this approach is developed for an appearance model based on the filter responses from a steerable pyramid. This model is used in a motion-based tracking algorithm to provide robustness in the face of image outliers, such as those caused by occlusions, while adapting to natural changes in appearance such as those due to facial expressions or variations in 3D pose. Allan Douglas Jepson, David J. Fleet, Thomas F. El-Maraghi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A Probabilistic Theory of Occupancy and Emptiness
Rahul Bhotika, David J. Fleet, Kiriakos N. Kutulakos |
ECCV (3) | 2 |
| 2002 | A Layered Motion Representation with Occlusion and Compact Spatial Support
Allan Douglas Jepson, David J. Fleet, Michael J. Black |
ECCV (1) | 2 |
| 2001 | Robust Online Appearance Models for Visual TrackingabstractAbstract—We propose a framework for learning robust, adaptive, appearance models to be used for motion-based tracking of natural objects. The model adapts to slowly changing appearance, and it maintains a natural measure of the stability of the observed image structure during tracking. By identifying stable properties of appearance, we can weight them more heavily for motion estimation, while less stable properties can be proportionately downweighted. The appearance model involves a mixture of stable image structure, learned over long time courses, along with two-frame motion information and an outlier process. An online EM-algorithm is used to adapt the appearance model parameters over time. An implementation of this approach is developed for an appearance model based on the filter responses from a steerable pyramid. This model is used in a motion-based tracking algorithm to provide robustness in the face of image outliers, such as those caused by occlusions, while adapting to natural changes in appearance such as those due to facial expressions or variations in 3D pose. Allan Douglas Jepson, David J. Fleet, Thomas F. El-Maraghi |
CVPR (1) | 2 |
| 2001 | Probabilistic Tracking of Motion Boundaries with Spatiotemporal PredictionsabstractWe describe a probabilistic framework for detecting and tracking motion boundaries. It builds on previous work (M.J. Black and D.J. Fleet, 2000) that used a particle filter to compute a posterior distribution over multiple, local motion models, one of which was specific for motion boundaries. We extend that framework in two ways: 1) with an enhanced likelihood that combines motion and edge support, 2) with a spatiotemporal model that propagates beliefs between adjoining image neighborhoods to encourage boundary continuity and provide better temporal predictions for motion boundaries. Approximate inference is achieved with a combination of tools: sampled representations allow us to represent multimodal non-Gaussian distributions and to apply nonlinear dynamics, while mixture models are used to simplify the computation of joint prediction distributions. Oscar Nestares, David J. Fleet |
CVPR (2) | 2 |
| 2001 | People Tracking Using Hybrid Monte Carlo FilteringabstractParticle filters are used for hidden state estimation with nonlinear dynamical systems. The inference of 3-D human motion is a natural application, given the nonlinear dynamics of the body and the nonlinear relation between states and image observations. However, the application of particle filters has been limited to cases where the number of state variables is relatively small, because the number of samples needed with high dimensional problems can be prohibitive. We describe a filter that uses hybrid Monte Carlo (HMC) to obtain samples in high dimensional spaces. It uses multiple Markov chains that use posterior gradients to rapidly explore the state space, yielding fair samples from the posterior. We find that the HMC filter is several thousand times faster than a conventional particle filter on a 28 D people tracking problem. Kiam Choo, David J. Fleet |
ICCV | 2 |
| 2001 | Lattice Particle Filters
Dirk Ormoneit, Christiane Lemieux, David J. Fleet |
UAI | 3 |
| 2001 | Computing Optical Flow with Physical Models of Brightness VariationabstractAlthough most optical flow techniques presume brightness constancy, it is well-known that this constraint is often violated, producing poor estimates of image motion. This paper describes a generalized formulation of optical flow estimation based on models of brightness variations that are caused by time-dependent physical processes. These include changing surface orientation with respect to a directional illuminant, motion of the illuminant, and physical models of heat transport in infrared images. With these models, we simultaneously estimate the 2D image motion and the relevant physical parameters of the brightness change model. The estimation problem is formulated using total least squares, with confidence bounds on the parameters. Experiments in four domains, with both synthetic and natural inputs, show how this formulation produces superior estimates of the 2D image motion. Horst W. Haussecker, David J. Fleet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | Computing Optical Flow with Physical Models of Brightness VariationabstractThis paper exploits physical models of time-varying brightness in image sequences to estimate optical flow and physical parameters of the scene. Previous approaches handled violations of brightness constancy with the use of robust statistics or with generalized brightness constancy constraints that allow generic types of contrast and illumination changes. We consider models of brightness variation that have time-dependent physical causes, namely, changing surface orientation with respect to a directional illuminant, motion of the illuminant, and physical models of heat transport in infrared images. We simultaneously estimate the optical flow and the relevant physical parameters. The estimation problem is formulated using total least squares (TLS), with confidence bounds on the parameters. Horst W. Haussecker, David J. Fleet |
CVPR | 2 |
| 2000 | Likelihood Functions and Confidence Bounds for Total-Least-Squares ProblemsabstractThis paper addresses the derivation of likelihood functions and confidence bounds for problems involving over-determined linear systems with noise in all measurements, often referred to as total-least-squares (TLS). It has been shown previously that TLS provides maximum likelihood estimates. But rather than being a function solely of the variables of interest, the associated likelihood functions increase in dimensionality with the number of equations. This has made it difficult to derive suitable confidence bounds, and impractical to use these probability functions with Bayesian belief propagation or Bayesian tracking. This paper derives likelihood functions that are defined only on the parameters of interest. This has two main advantages. First, the likelihood functions are much easier to use within a Bayesian framework; and second it is straightforward to obtain a reliable confidence bound on the estimates. We demonstrate the accuracy of our confidence bound in relation to others that have been proposed. Also, we use our theoretical results to obtain likelihood functions for estimating the direction of 3D camera translation. Oscar Nestares, David J. Fleet, David J. Heeger |
CVPR | 2 |
| 2000 | Stochastic Tracking of 3D Human Figures Using 2D Image Motion
Hedvig Kjellström, Michael J. Black, David J. Fleet |
ECCV (2) | 3 |
| 2000 | Robustly Estimating Changes in Image Appearance
Michael J. Black, David J. Fleet, Yaser Yacoob |
Comput. Vis. Image Underst. | 2 |
| 2000 | Probabilistic Detection and Tracking of Motion Boundaries
Michael J. Black, David J. Fleet |
Int. J. Comput. Vis. | 2 |
| 2000 | Design and Use of Linear Models for Image Motion Analysis
David J. Fleet, Michael J. Black, Yaser Yacoob, Allan Douglas Jepson |
Int. J. Comput. Vis. | 1 |
| 1999 | Probabilistic Detection and Tracking of Motion DiscontinuitiesabstractWe propose a Bayesian framework for representing and recognizing local image motion in terms of two primitive models: translation and motion discontinuity. Motion discontinuities are represented using a nonlinear generative model that explicitly encodes the orientation of the boundary, the velocities on either side, the motion of the occluding edge over time, and the appearance/disappearance of pixels at the boundary. We represent the posterior distribution over the model parameters given the image data using discrete samples. This distribution is propagated over time using the Condensation algorithm. To efficiently represent such a high-dimensional space we initialize samples using the responses of a low-level motion discontinuity detector. Michael J. Black, David J. Fleet |
ICCV | 2 |
| 1999 | Spotlights: A Robust Method for Surface-Based Registration in Orthopedic Surgery
Burton Ma, Randy E. Ellis, David J. Fleet |
MICCAI | 3 |
| 1998 | Motion Feature Detection Using Steerable Flow FieldsabstractThe estimation and detection of occlusion boundaries and moving bars are important and challenging problems in image sequence analysis. Here, we model such motion features as linear combinations of steerable basis flow fields. These models constrain the interpretation of image motion, and are used in the same way as translational or affine motion models. We estimate the subspace coefficients of the motion feature models directly from spatiotemporal image derivatives using a robust regression method. From the subspace coefficients we detect the presence of a motion feature and solve for the orientation of the feature and the relative velocities of the surfaces. Our method does not require the prior computation of optical flow and recovers accurate estimates of orientation and velocity. David J. Fleet, Michael J. Black, Allan Douglas Jepson |
CVPR | 1 |
| 1998 | A Framework for Modeling Appearance Change in Image SequencesabstractImage "appearance" may change over time due to a variety of causes such as: 1) object or camera motion; 2) generic photometric events including variations in illumination (e.g. shadows) and specular reflections; and 3) "iconic changes" which are specific to the objects being viewed and include complex occlusion events and changes in the material properties of the objects. We propose a general framework for representing and recovering these "appearance changes" in an image sequence as a "mixture" of different causes. The approach generalizes previous work on optical flow to provide a richer description of image events and more reliable estimates of image motion. Michael J. Black, David J. Fleet, Yaser Yacoob |
ICCV | 2 |
| 1997 | Learning Parameterized Models of Image MotionabstractA framework for learning parameterized models of optical flow from image sequences is presented. A class of motions is represented by a set of orthogonal basis flow fields that are computed from a training set using principal component analysis. Many complex image motions can be represented by a linear combination of a small number of these basis flows. The learned motion models may be used for optical flow estimation and for model-based recognition. For optical flow estimation we describe a robust, multi-resolution scheme for directly computing the parameters of the learned flow models from image derivatives. As examples we consider learning motion discontinuities, non-rigid motion of human mouths, and articulated human motion. Michael J. Black, Yaser Yacoob, Allan Douglas Jepson, David J. Fleet |
CVPR | 4 |
| 1997 | Embedding Invisible Information in Color ImagesabstractWe describe a method for embedding information in color images. A model of human color vision is used to ensure that the embedded signal is invisible. Sinusoidal signals are embedded so that they can be detected (decoded) without use of the original image. The sinusoids act as a grid, providing a coordinate frame on the image. We use the grid to automatically scale and align (deskew) images that have been printed and then scanned. David J. Fleet, David J. Heeger |
ICIP (1) | 1 |
| 1995 | Recursive Filters for Optical FlowabstractWorking toward efficient (real-time) implementations of optical flow methods, the authors have applied simple recursive filters to achieve temporal smoothing and differentiation of image intensity, and to compute 2d flow from component velocity constraints using spatiotemporal least-squares minimization. Accuracy in simulation is similar to that obtained in the study by Barren et al. (1994), while requiring much less storage, less computation, and shorter delays.> David J. Fleet, Keith Langley |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | Stable Estimation of Image OrientationabstractWe examine the performance of phase-based and energy-based techniques for estimating image orientation, emphasizing measurement stability under deformations that commonly occur in natural images. We also present a new energy-based technique for estimating multiple orientations in a neighbourhood and for reliably extracting multiple orientations from textured image regions.> Leif Haglund, David J. Fleet |
ICIP (3) | 2 |
| 1994 | Performance of optical flow techniques
John L. Barron, David J. Fleet, Steven S. Beauchemin |
Int. J. Comput. Vis. | 2 |
| 1993 | Stability of Phase InformationabstractThis paper concerns the robustness of local phase information for measuring image velocity and binocular disparity. It addresses the dependence of phase behavior on the initial filters as well as the image variations that exist between different views of a 3D scene. We are particularly interested in the stability of phase with respect to geometric deformations, and its linearity as a function of spatial position. These properties are important to the use of phase information, and are shown to depend on the form of the filters as well as their frequency bandwidths. Phase instabilities are also discussed using the model of phase singularities described by Jepson and Fleet. In addition to phase-based methods, these results are directly relevant to differential optical flow methods and zero-crossing tracking.> David J. Fleet, Allan Douglas Jepson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | On Transparent Motion Computation
Keith Langley, David J. Fleet, Tim J. Atherton |
BMVC | 2 |
| 1992 | Performance of optical flow techniquesabstractThe performance of six optical flow techniques is compared, emphasizing measurement accuracy. The most accurate methods are found to be the local differential approaches, where nu is computed explicitly in terms of a locally constant or linear model. Techniques using global smoothness constraints appear to produce visually attractive flow fields, but in general seem to be accurate enough for qualitative use only and insufficient as precursors to the computations of egomotion and 3D structures. It is found that some form of confidence measure/threshold is crucial for all techniques in order to separate the inaccurate from the accurate. Drawbacks of the six techniques are discussed.> John L. Barron, David J. Fleet, Steven S. Beauchemin, T. A. Burkitt |
CVPR | 2 |
| 1992 | Multiple motions from instantaneous frequencyabstractThe measurement of multiple velocities using phase-based methods is discussed. In particular, phase gradients (instantaneous frequency) from different bandpass channels (quadrature filter outputs) are used to estimate multiple image velocities in a single neighborhood. The approach is similar to that of M. Shizawa and K. Mase (1990) in which nth-order differential operators are required to compute n simultaneous velocity estimates. However, to use instantaneous frequency, the output of each channel must be differentiated only once.> Keith Langley, David J. Fleet, Tim J. Atherton |
CVPR | 2 |
| 1991 | Phase-based disparity measurement
David J. Fleet, Allan Douglas Jepson, Michael R. M. Jenkin |
CVGIP Image Underst. | 1 |
| 1991 | Phase singularities in scale-space
Allan Douglas Jepson, David J. Fleet |
Image Vis. Comput. | 2 |
| 1990 | Scale-Space Singularities
Allan Douglas Jepson, David J. Fleet |
ECCV | 2 |
| 1990 | Computation of component image velocity from local phase information
David J. Fleet, Allan Douglas Jepson |
Int. J. Comput. Vis. | 1 |
| 1989 | Computation of normal velocity from local phase informationabstractA technique for the estimation of 2-D normal velocity is presented. The image sequence is first represented by a family of velocity-tuned linear filters. Normal velocity, in the individual filter outputs, is expressed as the local first-order behavior of surfaces of constant phase. Justification for this is discussed, and it is shown to provide an effective basis for the local computation of normal velocity. The resultant approach is local in space-time. It permits multiple velocity estimates within a single neighborhood, and it yields accurate velocity estimates that are robust with respect to noise and perspective deformation.> David J. Fleet, Allan Douglas Jepson |
CVPR | 1 |
| 1989 | Hierarchical Construction of Orientation and Velocity Selective FiltersabstractAs a step towards the early measurement of visual primitives, the authors outline design criteria for the extraction of orientation and velocity information, and present a variety of tools useful in the construction of simple linear filters. A hierarchical parallel-processing scheme is used in which nodes compute a weighted sum of inputs from within a small spatio-temporal neighborhood. The resulting scheme is easily analyzed and provides mechanisms sensitive to narrow ranges of both image velocity and orientation. The hierarchical approach in combination with separability in the first levels yields an efficient implementation.> David J. Fleet, Allan Douglas Jepson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |