VLDB 2026 Research / reviewers in the wild / expert
David A. Forsyth
dblp:f/DavidAForsyth · also David Alexander Forsyth
· DBLP profile ↗
187ranked-venue papers
34as first author
31since 2021 · last 2026
0000-0002-2278-0752ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 149 · 30 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 125 · 16 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved Convex Decomposition with Ensembling and Negative PrimitivesabstractDescribing a scene in terms of primitives - geometrically simple shapes that offer a parsimonious but accurate abstraction of structure - is an established and difficult fitting problem. Different scenes require different numbers of primitives, and these primitives interact strongly. Existing methods are evaluated by comparing predicted depth, normals, and segmentation against ground truth. The state of the art method involves a learned regression procedure to predict a start point consisting of a fixed number of primitives, followed by a descent method to refine the geometry and remove redundant primitives. CSG (Constructive Solid Geometry) representations are significantly enhanced by a set-differencing operation. Our representation incorporates negative primitives, which are differenced from the positive primitives. These notably enrich the geometry that the model can encode, while complicating the fitting problem. This paper presents a method that can (a) incorporate these negative primitives and (b) choose the overall number of positive and negative primitives by ensembling. Extensive experiments on the standard NYUv2 dataset confirm that (a) this approach results in substantial improvements in depth representation and segmentation over SOTA and (b) negative primitives improve fitting accuracy. Our method is robustly applicable across datasets: in a first, we evaluate primitive prediction for LAION images. Vaibhav Vavilala, Florian Kluger, Seemandhar Jain, Bodo Rosenhahn, Anand Bhattad, David A. Forsyth |
3DV | 6 |
| 2025 | UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single VideoabstractWe present UrbanIR (Urban Scene Inverse Rendering), a new inverse graphics model that enables realistic, free-viewpoint renderings of scenes under various lighting conditions with a single video. It accurately infers shape, albedo, visibility, and sun and sky illumination from wide-baseline videos, such as those from car-mounted cameras, differing from NeRF's dense view settings. In this context, standard methods often yield subpar geometry and material estimates, such as inaccurate roof representations and numerous ‘floaters’. UrbanIR addresses these issues with novel losses that reduce errors in inverse graphics inference and rendering artifacts. Its techniques allow for precise shadow volume estimation in the original scene. The model's outputs support controllable editing, enabling photorealistic free-viewpoint renderings of night simulations, relit scenes, and inserted objects, marking a significant improvement over existing state-of-the-art methods. Our code and data will be made publicly available upon acceptance. Chih-Hao Lin, Kuan-Sheng Chen, David A. Forsyth, Jia-Bin Huang 0001, Anand Bhattad, Shenlong Wang |
3DV | 5 |
| 2025 | Denoising Monte Carlo Renders with Diffusion ModelsabstractPhysically-based renderings contain Monte Carlo noise, with variance that increases as the number of rays per pixel decreases. This noise, while zero-mean for good modern renderers, can have heavy tails (most notably, for scenes containing specular or refractive objects). Learned methods for restoring low fidelity renders are highly developed, because suppressing render noise means one can save compute and use fast renders with few rays per pixel. We demonstrate that a diffusion model can denoise low fidelity renders successfully. Furthermore, our method can be conditioned on a variety of natural render information, and this conditioning helps performance. Quantitative experiments show that our method is competitive with SOTA across a range of sampling rates. Qualitative examination of the reconstructions suggests that the image prior applied by a diffusion method strongly favors reconstructions that are “like” real images - so have straight shadow boundaries, curved specularities and no “fireflies.” Vaibhav Vavilala, Rahul Vasanth, David A. Forsyth |
3DV | 3 |
| 2025 | InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance ModelingabstractWe present InvRGB+L, a novel inverse rendering model that reconstructs large, relightable, and dynamic scenes from a single RGB+LiDAR sequence. Conventional inverse graphics methods rely primarily on RGB observations and use LiDAR mainly for geometric information, often resulting in suboptimal material estimates due to visible light interference. We find that LiDAR's intensity values-captured with active illumination in a different spectral range-offer complementary cues for robust material estimation under variable lighting. Inspired by this, InvRGB+L leverages LiDAR intensity cues to overcome challenges inherent in RGB-centric inverse graphics through two key innovations: (1) a novel physics-based LiDAR shading model and (2) RGB-LiDAR material consistency losses. The model produces novel-view RGB and LiDAR renderings of urban and indoor scenes and supports relighting, night simulations, and dynamic object insertions, achieving results that surpass current state-of-the-art methods in both scene-level urban inverse rendering and LiDAR simulation. Xiaoxue Chen, Bhargav Chandaka, Chih-Hao Lin, Ya-Qin Zhang, David A. Forsyth, Shenlong Wang |
ICCV | 5 |
| 2025 | Empirical Privacy VarianceabstractWe propose the notion of empirical privacy variance and study it in the context of differentially private fine-tuning of language models. Specifically, we show that models calibrated to the same $(\varepsilon, \delta)$-DP guarantee using DP-SGD with different hyperparameter configurations can exhibit significant variations in empirical privacy, which we quantify through the lens of memorization. We investigate the generality of this phenomenon across multiple dimensions and discuss why it is surprising and relevant. Through regression analysis, we examine how individual and composite hyperparameters influence empirical privacy. The results reveal a no-free-lunch trade-off: existing practices of hyperparameter tuning in DP-SGD, which focus on optimizing utility under a fixed privacy budget, often come at the expense of empirical privacy. To address this, we propose refined heuristics for hyperparameter selection that explicitly account for empirical privacy, showing that they are both precise and practically useful. Finally, we take preliminary steps to understand empirical privacy variance. We propose two hypotheses, identify limitations in existing techniques like privacy auditing, and outline open questions for future research. Yuzheng Hu, Fan Wu 0011, Ruicheng Xian, Lydia Zakynthinou, Pritish Kamath, Chiyuan Zhang, David A. Forsyth |
ICML | 8 |
| 2025 | A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningabstractOnline reinforcement learning (RL) excels in complex, safety-critical domains but suffers from sample inefficiency, training instability, and limited interpretability. Data attribution provides a principled way to trace model behavior back to training samples, yet existing methods assume fixed datasets, which is violated in online RL where each experience both updates the policy and shapes future data collection.
In this paper, we initiate the study of data attribution for online RL, focusing on the widely used Proximal Policy Optimization (PPO) algorithm. We start by establishing a *local* attribution framework, interpreting model checkpoints with respect to the records in the recent training buffer. We design two target functions, capturing agent action and cumulative return respectively, and measure each record's contribution through gradient similarity between its training loss and these targets. We demonstrate the power of this framework through three concrete applications: diagnosis of learning, temporal analysis of behavior formation, and targeted intervention during training. Leveraging this framework, we further propose an algorithm, iterative influence-based filtering (IIF), for online RL training that iteratively performs experience filtering to refine policy updates. Across standard RL benchmarks (classic control, navigation, locomotion) to RLHF for large language models, IIF reduces sample complexity, speeds up training, and achieves higher returns. Together, these results open a new direction for making online RL more interpretable, efficient, and effective. Yuzheng Hu, Fan Wu 0011, Haotian Ye, David A. Forsyth, James Zou 0001, Nan Jiang 0008, Jiaqi W. Ma, Han Zhao 0002 |
NeurIPS | 4 |
| 2025 | Copy or Not? Reference-Based Face Image Restoration with Fine Details
Min Jin Chong, Dejia Xu, Zhangyang Wang, David A. Forsyth, Gurunandan Krishnan |
WACV | 5 |
| 2025 | Dequantization and Color Transfer with Diffusion ModelsabstractWe demonstrate an image dequantizing diffusion model that enables novel edits on natural images. We propose operating on quantized images because they offer easy abstraction for patch-based edits and palette transfer. In particular, we show that color palettes can make the output of the diffusion model easier to control and interpret. We first establish that existing image restoration methods are not sufficient, such as JPEG noise reduction models. We then demonstrate that our model can generate natural images that respect the color palette the user asked for. For palette transfer, we propose a method based on weighted bipartite matching. We then show that our model generates plausible images even after extreme palette transfers, respecting user query. Our method can optionally condition on the source texture in part or all of the image. In doing so, we overcome a common problem in existing image colorization methods that are unable to produce colors with a different luminance than the input. We evaluate several possibilities for texture conditioning and their tradeoffs, including luminance, image gradients, and thresholded gradients, the latter of which performed best in maintaining texture and color control simultaneously. Our method can be usefully extended to another practical edit: recoloring patches of an image while respecting the source texture. Our procedure is supported by several qualitative and quantitative evaluations. Vaibhav Vavilala, Faaris Shaik, David A. Forsyth |
WACV | 3 |
| 2024 | StyLitGAN: Image-Based Relighting via Latent ControlabstractWe describe a novel method, StyLitGAN, for relighting and resurfacing images in the absence of labeled data. StyL-itGAN generates images with realistic lighting effects, including cast shadows, soft shadows, inter-reflections, and glossy effects, without the need for paired or CGI data. StyLit-GAN uses an intrinsic image method to decompose an image, followed by a search of the latent space of a pretrained Style-GAN to identify a set of directions. By prompting the model to fix one component (e.g., albedo) and vary another (e.g., shading), we generate relighted images by adding the identi-fied directions to the latent style codes. Quantitative metrics of change in albedo and lighting diversity allow us to choose effective directions using a forward selection process. Qual-itative evaluation confirms the effectiveness of our method. Anand Bhattad, James Soole, David A. Forsyth |
CVPR | 3 |
| 2024 | Shadows Don't Lie and Lines Can't Bend! Generative Models Don't know Projective Geometry...for NowabstractGenerative models can produce impressively realistic images. This paper demonstrates that generated images have geometric features different from those of real images. We build a set of collections of generated images, prequalified to fool simple, signal-based classifiers into believing they are real. We then show that prequalified generated images can be identified reliably by classifiers that only look at geometric properties. We use three such classifiers. All three classifiers are denied access to image pixels, and look only at derived geometric features. The first classifier looks at the perspective field of the image, the second looks at lines detected in the image, and the third looks at relations between detected objects and shadows. Our procedure detects generated images more reliably than SOTA local signal based detectors, for images from a number of distinct generators. Saliency maps suggest that the classifiers can identify geometric problems reliably. We conclude that current generators cannot reliably reproduce geometric properties of real images. Ayush Sarkar, Hanlin Mai, Amitabh Mahapatra, Svetlana Lazebnik, David A. Forsyth, Anand Bhattad |
CVPR | 5 |
| 2024 | HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalabstractAutomated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigorously assess new methods. To address this issue, we introduce HarmBench, a standardized evaluation framework for automated red teaming. We identify several desirable properties previously unaccounted for in red teaming evaluations and systematically design HarmBench to meet these criteria. Using HarmBench, we conduct a large-scale comparison of 18 red teaming methods and 33 target LLMs and defenses, yielding novel insights. We also introduce a highly efficient adversarial training method that greatly enhances LLM robustness across a wide range of attacks, demonstrating how HarmBench enables codevelopment of attacks and defenses. We open source HarmBench at https://github.com/centerforaisafety/HarmBench. Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang 0001, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li 0026, David A. Forsyth, Dan Hendrycks |
ICML | 11 |
| 2024 | Latent Intrinsics Emerge from Training to RelightabstractImage relighting is the task of showing what a scene from a source image would look like if illuminated differently. Inverse graphic schemes recover an explicit representation of geometry and a set of chosen intrinsics, then relight with some form of renderer. But error control for inverse graphics is difficult, and inverse graphics methods can represent only the effects of the chosen intrinsics. This paper describes a relighting method that is entirely data-driven, where intrinsics and lighting are each represented as latent variables. Our approach produces SOTA relightings of real scenes, as measured by standard metrics. We show that albedo can be recovered from our latent intrinsics without using any example albedos, and that the albedos recovered are competitive with SOTA methods. William Gao, Seemandhar Jain, Michael Maire, David A. Forsyth, Anand Bhattad |
NeurIPS | 5 |
| 2024 | SoK: Privacy-Preserving Data SynthesisabstractAs the prevalence of data analysis grows, safeguarding data privacy has become a paramount concern. Consequently, there has been an upsurge in the development of mechanisms aimed at privacy-preserving data analyses. However, these approaches are task-specific; designing algorithms for new tasks is a cumbersome process. As an alternative, one can create synthetic data that is (ideally) devoid of private information. This paper focuses on privacy-preserving data synthesis (PPDS) by providing a comprehensive overview, analysis, and discussion of the field. Specifically, we put forth a master recipe that unifies two prominent strands of research in PPDS: statistical methods and deep learning (DL)-based methods. Under the master recipe, we further dissect the statistical methods into choices of modeling and representation, and investigate the DL-based methods by different generative modeling principles. To consolidate our findings, we provide comprehensive reference tables, distill key takeaways, and identify open problems in the existing literature. In doing so, we aim to answer the following questions: What are the design principles behind different PPDS methods? How can we categorize these methods, and what are the advantages and disadvantages associated with each category? Can we provide guidelines for method selection in different real-world scenarios? We proceed to benchmark several prominent DL-based methods on the task of private image synthesis and conclude that DP-MERF is an all-purpose approach. Finally, upon systematizing the work over the past decade, we identify future directions and call for actions from researchers. Yuzheng Hu, Fan Wu 0011, Qinbin Li, Yunhui Long, Gonzalo Munilla Garrido, Chang Ge 0002, Bolin Ding, David A. Forsyth, Bo Li 0026, Dawn Song |
SP | 8 |
| 2024 | Preserving Image Properties Through Initializations in Diffusion ModelsabstractRetail photography imposes specific requirements on images. For instance, images may need uniform background colors, consistent model poses, centered products, and consistent lighting. Minor deviations from these standards impact a site’s aesthetic appeal, making the images unsuitable for use. We show that Stable Diffusion methods, as currently applied, do not respect these requirements. The usual practice of training the denoiser with a very noisy image and starting inference with a sample of pure noise leads to inconsistent generated images during inference. This inconsistency occurs because it is easy to tell the difference between samples of the training and inference distributions. As a result, a network trained with centered retail product images with uniform backgrounds generates images with erratic backgrounds. The problem is easily fixed by initializing inference with samples from an approximation of noisy images. However, in using such an approximation, the joint distribution of text and noisy image at inference time still slightly differs from that at training time. This discrepancy is corrected by training the network with samples from the approximate noisy image distribution. Extensive experiments on real application data show significant qualitative and quantitative improvements in performance from adopting these procedures. Finally, our procedure can interact well with other control-based methods to further enhance the controllability of diffusion-based methods. Jeffrey Zhang 0004, Shao-Yu Chang, Kedan Li, David A. Forsyth |
WACV | 4 |
| 2024 | P2D: Plug and Play Discriminator for accelerating GAN frameworksabstractMost image classification tasks benefit from using pre-trained feature stacks. In contrast, the discriminator for adversarial losses is trained at the same time as the model because using a pretrained feature stack yields a very poor model. Recent work has shown that an implicit regularization scheme allows using pretrained feature stacks to construct a discriminator, which improves both speed of training and quality of results. However, we observe that changes in hyperparameters can result in substantial changes in generator behavior.We show that using a modified version of the R1 regularization scheme that regularizes in the feature space instead of the image space results in a plug-and-play discriminator– P2D. Our scheme results in a method that is highly stable across changes in architecture and framework; that significantly speeds up training; and that produces models that reliably beat SOTA in quality. The huge reduction in training resources required means that P2D could make training powerful generative models over specific datasets accessible to most researchers. Min Jin Chong, Krishna Kumar Singh, Yijun Li 0001, Jingwan Lu, David A. Forsyth |
WACV | 5 |
| 2024 | Controlling Virtual Try-on Pipeline Through Rendering PoliciesabstractThis paper shows how to impose rendering policies on a virtual try-on (VTON) pipeline. Our rendering policies are lightweight procedural descriptions of how the pipeline should render outfits or render particular types of garments. Our policies are procedural expressions describing offsets to the control points for each set of garment types. The policies are easily authored and are generalizable to any outfit composed of garments of similar types. We describe a VTON pipeline that accepts our policies to modify garment drapes and produce high-quality try-on images with garment attributes preserved.Layered outfits are a particular challenge to VTON systems because learning to coordinate warps between multiple garments so that nothing sticks out is difficult. Our rendering policies offer a lightweight and effective procedure to achieve this coordination, while also allowing precise manipulation of drape. Drape describes the way in which a garment is worn (for example, a shirt could be tucked or untucked).Quantitative and qualitative evaluations demonstrate that our method allows effective manipulation of drape and produces significant measurable improvements in rendering quality for complicated layering interactions. Kedan Li, Jeffrey Zhang 0004, Shao-Yu Chang, David A. Forsyth |
WACV | 4 |
| 2023 | ClimateNeRF: Extreme Weather Synthesis in Neural Radiance FieldabstractPhysical simulations produce excellent predictions of weather effects. Neural radiance fields produce SOTA scene models. We describe a novel NeRF-editing procedure that can fuse physical simulations with NeRF models of scenes, producing realistic movies of physical phenomena in those scenes. Our application – Climate NeRF – allows people to visualize what climate change outcomes will do to them.ClimateNeRF allows us to render realistic weather effects, including smog, snow, and flood. Results can be controlled with physically meaningful variables like water level. Qualitative and quantitative studies show that our simulated results are significantly more realistic than those from SOTA 2D image editing and SOTA 3D NeRF stylization. Zhi-Hao Lin, David A. Forsyth, Jia-Bin Huang 0001, Shenlong Wang |
ICCV | 3 |
| 2023 | Convex Decomposition of Indoor ScenesabstractWe describe a method to parse a complex, cluttered indoor scene into primitives which offer a parsimonious abstraction of scene structure. Our primitives are simple convexes. Our method uses a learned regression procedure to parse a scene into a fixed number of convexes from RGBD input, and can optionally accept segmentations to improve the decomposition. The result is then polished with a descent method which adjusts the convexes to produce a very good fit, and greedily removes superfluous primitives. Because the entire scene is parsed, we can evaluate using traditional depth, normal, and segmentation error metrics. Our evaluation procedure demonstrates that the error from our primitive representation is comparable to that of predicting depth from a single image. Vaibhav Vavilala, David A. Forsyth |
ICCV | 2 |
| 2023 | Improving Equivariance in State-of-the-Art Supervised Depth and Normal PredictorsabstractDense depth and surface normal predictors should possess the equivariant property to cropping-and-resizing – cropping the input image should result in cropping the same output image. However, we find that state-of-the-art depth and normal predictors, despite having strong performances, surprisingly do not respect equivariance. The problem exists even when crop-and-resize data augmentation is employed during training. To remedy this, we propose an equivariant regularization technique, consisting of an averaging procedure and a self-consistency loss, to explicitly promote cropping-and-resizing equivariance in depth and normal networks. Our approach can be applied to both CNN and Transformer architectures, does not incur extra cost during testing, and notably improves the supervised and semi-supervised learning performance of dense predictors on Taskonomy tasks. Finally, finetuning with our loss on unlabeled images improves not only equivariance but also accuracy of state-of-the-art depth and normal predictors when evaluated on NYU-v2. Yuanyi Zhong, Anand Bhattad, Yu-Xiong Wang, David A. Forsyth |
ICCV | 4 |
| 2023 | StyleGAN knows Normal, Depth, Albedo, and MoreabstractIntrinsic images, in the original sense, are image-like maps of scene properties like depth, normal, albedo, or shading. This paper demonstrates that StyleGAN can easily be induced to produce intrinsic images. The procedure is straightforward. We show that if StyleGAN produces $G({\bf w})$ from latent ${\bf w}$, then for each type of intrinsic image, there is a fixed offset ${\bf d}_c$ so that $G({\bf w}+{\bf d}_c)$ is that type of intrinsic image for $G({\bf w})$. Here ${\bf d}_c$ is {\em independent of ${\bf w}$}. The StyleGAN we used was pretrained by others, so this property is not some accident of our training regime. We show that there are image transformations StyleGAN will {\em not} produce in this fashion, so StyleGAN is not a generic image regression engine.
It is conceptually exciting that an image generator should ``know'' and represent intrinsic images. There may also be practical advantages to using a generative model to produce intrinsic images. The intrinsic images obtained from StyleGAN compare well both qualitatively and quantitatively with those obtained by using SOTA image regression techniques; but StyleGAN's intrinsic images are robust to relighting effects, unlike SOTA methods. Anand Bhattad, Daniel McKee, Derek Hoiem, David A. Forsyth |
NeurIPS | 4 |
| 2023 | POVNet: Image-Based Virtual Try-On Through Accurate Warping and ResidualabstractVirtual dressing room applications help online shoppers visualize outfits. Such a system, to be commercially viable, must satisfy a set of performance criteria. The system must produce high quality images that faithfully preserve garment properties, allow users to mix and match garments of various types and support human models varying in skin tone, hair color, body shape, and so on. This paper describes POVNet, a framework that meets all these requirements (except body shapes variations). Our system uses warping methods together with residual data to preserve garment texture at fine scales and high resolution. Our warping procedure adapts to a wide range of garments and allows swapping in and out of individual garments. A learned rendering procedure using an adversarial loss ensures that fine shading, etc. is accurately reflected. A distance transform representation ensures that hems, cuffs, stripes, and so on are correctly placed. We demonstrate improvements in garment rendering over state of the art resulting from these procedures. We demonstrate that the framework is scalable, responds in real-time, and works robustly with a variety of garment categories. Finally, we demonstrate that using this system as a virtual dressing room interface for fashion e-commerce websites has significantly boosted user-engagement rates. Kedan Li, Jeffrey Zhang 0004, David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Cut-and-Paste Object Insertion by Enabling Deep Image Prior for ReshadingabstractWe show how to insert an object from one image to another and get realistic results in the hard case, where the shading of the inserted object clashes with the shading of the scene. Rendering objects using an illumination model of the scene doesn't work, because doing so requires a geometric and material model of the object, which is hard to recover from a single image. In this paper, we introduce a method that corrects shading inconsistencies of the inserted object without requiring a geometric and physical model or an environment map. Our method uses a deep image prior (DIP), trained to produce reshaded renderings of inserted objects via consistent image decomposition inferential losses. The resulting image from DIP aims to have (a) an albedo similar to the cut-and-paste albedo, (b) a similar shading field to that of the target scene, and (c) a shading that is consistent with the cut-and-paste surface normals. The result is a simple procedure that produces convincing shading of the inserted object. We show the efficacy of our method both qualitatively and quantitatively for several objects with complex surface properties and also on a dataset of spherical lampshades for quantitative evaluation. Our method significantly outperforms an Image Harmonization (IH) baseline for all these objects. They also outperform the cut-and-paste and IH baselines in a user study with over 100 users. Anand Bhattad, David A. Forsyth |
3DV | 2 |
| 2022 | DIVeR: Real-time and Accurate Neural Radiance Fields with Deterministic Integration for Volume RenderingabstractDIVeR builds on the key ideas of NeRF and its variants-density models and volume rendering – to learn 3D object models that can be rendered realistically from small numbers of images. In contrast to all previous NeRF methods, DIVeR uses deterministic rather than stochastic estimates of the volume rendering integral. DIVeR's representation is a voxel based field of features. To compute the volume rendering integral, a ray is broken into intervals, one per voxel; components of the volume rendering integral are estimated from the features for each interval using an MLP, and the components are aggregated. As a result, DIVeR can render thin translucent structures that are missed by other integrators. Furthermore, DIVeR's representation has semantics that is relatively exposed compared to other such methods – moving feature vectors around in the voxel space results in natural edits. Extensive qualitative and quantitative comparisons to current state-of-the-art methods show that DIVeR produces models that (1) render at or above state-of-the-art quality, (2) are very small without being baked, (3) render very fast without being baked, and (4) can be edited in natural ways. Our real-time code is available at: https://github.com/lwwu2/diver-rt Liwen Wu, Jae Yong Lee 0006, Anand Bhattad, Yu-Xiong Wang, David A. Forsyth |
CVPR | 5 |
| 2022 | JoJoGAN: One Shot Face Stylization
Min Jin Chong, David A. Forsyth |
ECCV (16) | 2 |
| 2022 | On the Importance of Firth Bias Reduction in Few-Shot Classification
Saba Ghaffari, Ehsan Saleh, David A. Forsyth, Yu-Xiong Wang |
ICLR | 3 |
| 2022 | How to Steer Your Adversary: Targeted and Efficient Model Stealing Defenses with Gradient RedirectionabstractModel stealing attacks present a dilemma for public machine learning APIs. To protect financial investments, companies may be forced to withhold important information about their models that could facilitate theft, including uncertainty estimates and prediction explanations. This compromise is harmful not only to users but also to external transparency. Model stealing defenses seek to resolve this dilemma by making models harder to steal while preserving utility for benign users. However, existing defenses have poor performance in practice, either requiring enormous computational overheads or severe utility trade-offs. To meet these challenges, we present a new approach to model stealing defenses called gradient redirection. At the core of our approach is a provably optimal, efficient algorithm for steering an adversary’s training updates in a targeted manner. Combined with improvements to surrogate networks and a novel coordinated defense strategy, our gradient redirection defense, called GRAD^2, achieves small utility trade-offs and low computational overhead, outperforming the best prior defenses. Moreover, we demonstrate how gradient redirection enables reprogramming the adversary with arbitrary behavior, which we hope will foster work on new avenues of defense. Mantas Mazeika, Bo Li 0026, David A. Forsyth |
ICML | 3 |
| 2022 | How Would The Viewer Feel? Estimating Wellbeing From Video ScenariosabstractIn recent years, deep neural networks have demonstrated increasingly strong abilities to recognize objects and activities in videos. However, as video understanding becomes widely used in real-world applications, a key consideration is developing human-centric systems that understand not only the content of the video but also how it would affect the wellbeing and emotional state of viewers. To facilitate research in this setting, we introduce two large-scale datasets with over 60,000 videos manually annotated for emotional response and subjective wellbeing. The Video Cognitive Empathy (VCE) dataset contains annotations for distributions of fine-grained emotional responses, allowing models to gain a detailed understanding of affective states. The Video to Valence (V2V) dataset contains annotations of relative pleasantness between videos, which enables predicting a continuous spectrum of wellbeing. In experiments, we show how video models that are primarily trained to recognize actions and find contours of objects can be repurposed to understand human preferences and the emotional content of videos. Although there is room for improvement, predicting wellbeing and emotional response is on the horizon for state-of-the-art models. We hope our datasets can help foster further advances at the intersection of commonsense video understanding and human preference learning. Mantas Mazeika, Eric Tang, Andy Zou, Steven Basart, Jun Shern Chan, Dawn Song, David A. Forsyth, Jacob Steinhardt, Dan Hendrycks |
NeurIPS | 7 |
| 2022 | Controlled GAN-Based Creature Synthesis via a Challenging Game Art Dataset - Addressing the Noise-Latent Trade-OffabstractThe state-of-the-art StyleGAN2 network supports powerful methods to create and edit art, including generating random images, finding images "like" some query, and modifying content or style. Further, recent advancements enable training with small datasets. We apply these methods to synthesize card art, by training on a novel Yu-Gi-Oh dataset. While noise inputs to StyleGAN2 are essential for good synthesis, we find that coarse-scale noise interferes with latent variables on this dataset because both control long-scale image effects. We observe over-aggressive variation in art with changes in noise and weak content control via latent variable edits. Here, we demonstrate that training a modified StyleGAN2, where coarse-scale noise is suppressed, removes these unwanted effects. We obtain a superior FID; changes in noise result in local exploration of style; and identity control is markedly improved. These results and analysis lead towards a GAN-assisted art synthesis tool for digital artists of all skill levels, which can be used in film, games, or any creative industry for artistic ideation. Vaibhav Vavilala, David A. Forsyth |
WACV | 2 |
| 2022 | Intrinsic Image Decomposition Using ParadigmsabstractIntrinsic image decomposition is the task of mapping image to albedo and shading. Classical approaches derive methods from spatial models. The modern literature stresses evaluation, by comparing predictions to human judgements ("lighter", "same as", "darker"). The best modern intrinsic image methods train a map from image to albedo using images rendered from computer graphics models and example human judgements. This approach yields practical methods, but obtaining rendered images can be inconvenient. Furthermore, the approach cannot explain how a one could learn to recover intrinsic images without geometric, surface and illumination models, as people and animals appear to do. This paper describes a method that learns intrinsic image decomposition without seeing human annotations, rendered data, or ground truth data. Instead, the method relies on paradigms - spatial models of albedo and of shading. Rather than finding the "best" albedo and shading for an image via optimization, our approach trains a neural network on synthetic images. The synthetic images are constructed by multiplying albedos and shading fields sampled from our models. The network is subject to a novel smoothing procedure that ensures good behavior at short scales on real images. An averaging procedure ensures that reported albedo and shading are largely equivariant - different crops and scalings of an image will report the same albedo and shading at shared points. This averaging procedure controls long scale error. The standard evaluation for an intrinsic image method is a WHDR score. Our method achieves WHDR scores competitive with those of strong recent methods allowed to see training WHDR annotations, rendered data, and ground truth data. Our method produces albedo and shading maps with attractive qualitative properties - for example, albedo fields do not suppress wood grain and represent narrow grooves in surfaces well. Because our method is unsupervised, we can compute estimates of the test/train variance of WHDR scores; these are quite large, and suggest is unsafe to rely small differences in reported WHDR. David A. Forsyth, Jason Rock |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Retrieve in Style: Unsupervised Facial Feature Transfer and RetrievalabstractWe present Retrieve in Style (RIS), an unsupervised framework for facial feature transfer and retrieval on real images. Recent work shows capabilities of transferring local facial features by capitalizing on the disentanglement property of the StyleGAN latent space. RIS improves existing art on the following: 1) Introducing more effective feature disentanglement to allow for challenging transfers (i.e., hair, pose) that were not shown possible in SoTA methods. 2) Eliminating the need for per-image hyperparameter tuning, and for computing a catalog over a large batch of images. 3) Enabling fine-grained face retrieval using disentangled facial features (e.g., eyes). To our best knowledge, this is the first work to retrieve face images at this fine level. 4) Demonstrating robust, natural editing on real images. Our qualitative and quantitative analyses show RIS achieves both high-fidelity feature transfers and accurate fine-grained retrievals on real images. We also discuss the responsible applications of RIS. Our code is available at https://github.com/mchong6/RetrieveInStyle. Min Jin Chong, Wen-Sheng Chu, David A. Forsyth |
ICCV | 4 |
| 2021 | LSD-StructureNet: Modeling Levels of Structural Detail in 3D Part HierarchiesabstractGenerative models for 3D shapes represented by hierarchies of parts can generate realistic and diverse sets of outputs. However, existing models suffer from the key practical limitation of modelling shapes holistically and thus cannot perform conditional sampling, i.e. they are not able to generate variants on individual parts of generated shapes without modifying the rest of the shape. This is limiting for applications such as 3D CAD design that involve adjusting created shapes at multiple levels of detail. To address this, we introduce LSD-StructureNet, an augmentation to the StructureNet architecture that enables re-generation of parts situated at arbitrary positions in the hierarchies of its outputs. We achieve this by learning individual, probabilistic conditional decoders for each hierarchy depth. We evaluate LSD-StructureNet on the PartNet dataset, the largest dataset of 3D shapes represented by hierarchies of parts. Our results show that contrarily to existing methods, LSD-StructureNet can perform conditional sampling without impacting inference speed or the realism and diversity of its outputs. Dominic Roberts, Ara Danielyan, Hang Chu, Mani Golparvar Fard, David A. Forsyth |
ICCV | 5 |
| 2020 | Effectively Unbiased FID and Inception Score and Where to Find ThemabstractThis paper shows that two commonly used evaluation metrics for generative models, the Fréchet Inception Distance (FID) and the Inception Score (IS), are biased -- the expected value of the score computed for a finite sample set is not the true value of the score. Worse, the paper shows that the bias term depends on the particular model being evaluated, so model A may get a better score than model B simply because model A's bias term is smaller. This effect cannot be fixed by evaluating at a fixed number of samples. This means all comparisons using FID or IS as currently computed are unreliable. We then show how to extrapolate the score to obtain an effectively bias-free estimate of scores computed with an infinite number of samples, which we term FID Infinity and IS Infinity. In turn, this effectively bias-free estimate requires good estimates of scores with a finite number of samples. We show that using Quasi-Monte Carlo integration notably improves estimates of FID and IS for finite sample sets. Our extrapolated scores are simple, drop-in replacements for the finite sample scores. Additionally, we show that using low discrepancy sequence in GAN training offers small improvements in the resulting generator. Min Jin Chong, David A. Forsyth |
CVPR | 2 |
| 2020 | Why Do These Match? Explaining the Behavior of Image Similarity Models
Bryan A. Plummer, Mariya I. Vasileva, Vitali Petsiuk, Kate Saenko, David A. Forsyth |
ECCV (11) | 5 |
| 2020 | Unrestricted Adversarial Examples via Semantic Manipulation
Anand Bhattad, Min Jin Chong, Kaizhao Liang, Bo Li 0026, David A. Forsyth |
ICLR | 5 |
| 2020 | Improving Style Transfer with Calibrated MetricsabstractStyle transfer produces a transferred image which is a rendering of a content image in the manner of a style image. We seek to understand how to improve style transfer.To do so requires quantitative evaluation procedures, but current evaluation is qualitative, mostly involving user studies. We describe a novel quantitative evaluation procedure. Our procedure relies on two statistics: the Effectiveness (E) statistic measures the extent that a given style has been transferred to the target, and the Coherence (C) statistic measures the extent to which the original image's content is preserved. Our statistics are calibrated to human preference: targets with larger values of E and C will reliably be preferred by human subjects in comparisons of style and content, respectively.We use these statistics to investigate relative performance of a number of Neural Style Transfer (NST) methods, revealing a number of intriguing properties. Admissible methods lie on a Pareto frontier (i.e. improving E reduces C, or vice versa). Three methods are admissible: Universal style transfer produces very good C but weak E; modifying the optimization used for Gatys' loss produces a method with strong E and strong C; and a modified cross-layer method has slightly better E at strong cost in C. While the histogram loss improves the E statistics of Gatys' method, it does not make the method admissible. Surprisingly, style weights have relatively little effect in improving EC scores, and most variability in transfer is explained by the style itself (meaning experimenters can be misguided by selecting styles). Our GitHub Link is available1. Mao-Chuang Yeh, Anand Bhattad, Chuhang Zou, David A. Forsyth |
WACV | 5 |
| 2020 | Guest Editors' Introduction to the Special Section on Computational Photography
Ayan Chakrabarti, Kalyan Sunkavalli, David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Fast, Diverse and Accurate Image Captioning Guided by Part-Of-SpeechabstractImage captioning is an ambiguous problem, with many suitable captions for an image. To address ambiguity, beam search is the de facto method for sampling multiple captions. However, beam search is computationally expensive and known to produce generic captions. To address this concern, some variational auto-encoder (VAE) and generative adversarial net (GAN) based methods have been proposed. Though diverse, GAN and VAE are less accurate. In this paper, we first predict a meaningful summary of the image, then generate the caption based on that summary. We use part-of-speech as summaries, since our summary should drive caption generation. We achieve the trifecta: (1) High accuracy for the diverse captions as evaluated by standard captioning metrics and user studies; (2) Faster computation of diverse captions compared to beam search and diverse beam search; and (3) High diversity as evaluated by counting novel sentences, distinct n-grams and mutual overlap (i.e., mBleu-4) scores. Aditya Deshpande, Jyoti Aneja, Liwei Wang 0009, Alexander G. Schwing, David A. Forsyth |
CVPR | 5 |
| 2019 | Max-Sliced Wasserstein Distance and Its Use for GANsabstractGenerative adversarial nets (GANs) and variational auto-encoders have significantly improved our distribution modeling capabilities, showing promise for dataset augmentation, image-to-image translation and feature learning. However, to model high-dimensional distributions, sequential training and stacked architectures are common, increasing the number of tunable hyper-parameters as well as the training time. Nonetheless, the sample complexity of the distance metrics remains one of the factors affecting GAN training. We first show that the recently proposed sliced Wasserstein distance has compelling sample complexity properties when compared to the Wasserstein distance. To further improve the sliced Wasserstein distance we then analyze its `projection complexity' and develop the max-sliced Wasserstein distance which enjoys compelling sample complexity while reducing projection complexity, albeit necessitating a max estimation. We finally illustrate that the proposed distance trains GANs on high-dimensional images up to a resolution of 256x256 easily. Ishan Deshpande, Yuan-Ting Hu, Ruoyu Sun 0001, Ayis Pyrros, Nasir Siddiqui, Oluwasanmi Koyejo, Zhizhen Zhao 0001, David A. Forsyth, Alexander G. Schwing |
CVPR | 8 |
| 2019 | An Approximate Shading Model with Detail Decomposition for Object Relighting
Zicheng Liao, Kevin Karsch, David A. Forsyth |
Int. J. Comput. Vis. | 4 |
| 2018 | Structural Consistency and Controllability for Diverse Colorization
Safa Messaoud, David A. Forsyth, Alexander G. Schwing |
ECCV (6) | 2 |
| 2018 | Learning Type-Aware Embeddings for Fashion Compatibility
Mariya I. Vasileva, Bryan A. Plummer, Krishna Dusad, Shreya Rajpal, Ranjitha Kumar, David A. Forsyth |
ECCV (16) | 6 |
| 2017 | Learning Diverse Image ColorizationabstractColorization is an ambiguous problem, with multiple viable colorizations for a single grey-level image. However, previous methods only produce the single most probable colorization. Our goal is to model the diversity intrinsic to the problem of colorization and produce multiple colorizations that display long-scale spatial co-ordination. We learn a low dimensional embedding of color fields using a variational autoencoder (VAE). We construct loss terms for the VAE decoder that avoid blurry outputs and take into account the uneven distribution of pixel colors. Finally, we build a conditional model for the multi-modal distribution between grey-level image and the color field embeddings. Samples from this conditional model result in diverse colorization. We demonstrate that our method obtains better diverse colorizations than a standard conditional variational autoencoder (CVAE) model, as well as a recently proposed conditional generative adversarial network (cGAN). Aditya Deshpande, Jiajun Lu, Mao-Chuang Yeh, Min Jin Chong, David A. Forsyth |
CVPR | 5 |
| 2017 | SafetyNet: Detecting and Rejecting Adversarial Examples RobustlyabstractWe describe a method to produce a network where current methods such as DeepFool have great difficulty producing adversarial samples. Our construction suggests some insights into how deep networks work. We provide a reasonable analyses that our construction is difficult to defeat, and show experimentally that our method is hard to defeat with both Type I and Type II attacks using several standard networks and datasets. This SafetyNet architecture is used to an important and novel application SceneProof, which can reliably detect whether an image is a picture of a real scene or not. SceneProof applies to images captured with depth maps (RGBD images) and checks if a pair of image and depth map is consistent. It relies on the relative difficulty of producing naturalistic depth maps for images in post processing. We demonstrate that our SafetyNet is robust to adversarial examples built from currently known attacking approaches. Jiajun Lu, Theerasit Issaranon, David A. Forsyth |
ICCV | 3 |
| 2017 | State of the JournalabstractPresents the editor's view of the current state of this journal publication. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Learning to Localize Little LandmarksabstractWe interact everyday with tiny objects such as the door handle of a car or the light switch in a room. These little landmarks are barely visible and hard to localize in images. We describe a method to find such landmarks by finding a sequence of latent landmarks, each with a prediction model. Each latent landmark predicts the next in sequence, and the last localizes the target landmark. For example, to find the door handle of a car, our method learns to start with a latent landmark near the wheel, as it is globally distinctive, subsequent latent landmarks use the context from the earlier ones to get closer to the target. Our method is supervised solely by the location of the little landmark and displays strong performance on more difficult variants of established tasks and on two new tasks. Saurabh Singh 0005, Derek Hoiem, David A. Forsyth |
CVPR | 3 |
| 2016 | Keynote speakers: Computational imaging: How much imaging - How much computation?abstractThis talk will discuss (from the viewpoint of a physicist with background in optical engineering and information) how much imaging and how much computing is (or should be) in computational imaging, aiming for high information efficiency. For an opticist the order of the keywords “computational imaging” is inverted: imaging is the first operation in the sequence, as optics does the encoding for redundancy reduction, which has necessarily to be done before electronic noise is added in the channel. Before computers emerged, decoding was performed solely by optical "hardware" and the observer. As for today, it appears as if digital image processing evolved to a new quality: computational imaging, with much more imaging involved than during the times of digital imaging processing. Optics and computational imaging are not anymore in different faculties. Computational imaging will definitely have a great future — and it has a history: opticists do computational imaging for many years, without even knowing the term, as the author of this abstract had to notice to its own astonishment. As paradigms, we will discuss a few 3D-"cameras" which exploit different complexity of "imaging = encoding" and “computation = decoding”. Among others deflectometry, with a dynamic range of up to 106, SEM like images by pure optics, with large depth of field and very low noise and the 3D motion picture camera, with full 3D information in each camera frame. Gerd Häusler, David A. Forsyth, Marc Walton, Francis Halzen, Aydogan Ozcan |
ICCP | 2 |
| 2016 | Swapout: Learning an ensemble of deep architecturesabstractWe describe Swapout, a new stochastic training method, that outperforms ResNets of identical network structure yielding impressive results on CIFAR-10 and CIFAR-100. Swapout samples from a rich set of architectures including dropout, stochastic depth and residual architectures as special cases. When viewed as a regularization method swapout not only inhibits co-adaptation of units in a layer, similar to dropout, but also across network layers. We conjecture that swapout achieves strong regularization by implicitly tying the parameters across layers. When viewed as an ensemble training method, it samples a much richer set of architectures than existing methods such as dropout or stochastic depth. We propose a parameterization that reveals connections to exiting architectures and suggests a much richer set of architectures to be explored. We show that our formulation suggests an efficient training method and validate our conclusions on CIFAR-10 and CIFAR-100 matching state of the art accuracy. Remarkably, our 32 layer wider model performs similar to a 1001 layer ResNet model. Saurabh Singh 0005, Derek Hoiem, David A. Forsyth |
NIPS | 3 |
| 2016 | State of the JournalabstractPresents information on the state of the journal for this issue of the publication. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | An approximate shading model for object relightingabstractWe propose an approximate shading model for image-based object modeling and insertion. Our approach is a hybrid of 3D rendering and image-based composition. It avoids the difficulties of physically accurate shape estimation from a single image, and allows for more flexible image composition than pure image-based methods. The model decomposes the shading field into (a) a rough shape term that can be reshaded, (b) a parametric shading detail that encodes missing features from the first term, and (c) a geometric detail term that captures fine-scale material properties. With this object model, we build an object relighting system that allows an artist to select an object from an image and insert it into a 3D scene. Through simple interactions, the system can adjust illumination on the inserted object so that it appears more naturally in the scene. Our quantitative evaluation and extensive user study suggest our method is a promising alternative to existing methods of object insertion. Zicheng Liao, Kevin Karsch, David A. Forsyth |
CVPR | 3 |
| 2015 | Sparse depth super resolutionabstractWe describe a method to produce detailed high resolution depth maps from aggressively subsampled depth measurements. Our method fully uses the relationship between image segmentation boundaries and depth boundaries It uses an image combined with a low resolution depth map. 1) The image is segmented with the guidance of sparse depth samples 2) Each segment has its depth field reconstructed independently using a novel smoothing method. 3) For videos, time-stamped samples from near frames are incorporated. The paper shows reconstruction results of super resolution from x4 to x100, while previous methods mainly work on x2 to xl6. The method is tested on four different datasets and six video sequences, covering quite different regimes, and it outperforms recent state of the art methods quantitatively and qualitatively We also demonstrate that depth maps produced by our method can be used by applications such as hand trackers, while depth maps from other methods have problems. Jiajun Lu, David A. Forsyth |
CVPR | 2 |
| 2015 | Learning a sequential search for landmarksabstractWe propose a general method to find landmarks in images of objects using both appearance and spatial context. This method is applied without changes to two problems: parsing human body layouts, and finding landmarks in images of birds. Our method learns a sequential search for localizing landmarks, iteratively detecting new landmarks given the appearance and contextual information from the already detected ones. The choice of landmark to be added is opportunistic and depends on the image; for example, in one image a head-shoulder group might be expanded to a head-shoulder-hip group but in a different image to a head-shoulder-elbow group. The choice of initial landmark is similarly image dependent. Groups are scored using a learned function, which is used to expand them greedily. Our scoring function is learned from data labelled with landmarks but without any labeling of a detection order. Our method represents a novel spatial model for the kinematics of groups of landmarks, and displays strong performance on two different model problems. Saurabh Singh 0005, Derek Hoiem, David A. Forsyth |
CVPR | 3 |
| 2015 | Learning Large-Scale Automatic Image ColorizationabstractWe describe an automated method for image colorization that learns to colorize from examples. Our method exploits a LEARCH framework to train a quadratic objective function in the chromaticity maps, comparable to a Gaussian random field. The coefficients of the objective function are conditioned on image features, using a random forest. The objective function admits correlations on long spatial scales, and can control spatial error in the colorization of the image. Images are then colorized by minimizing this objective function. We demonstrate that our method strongly outperforms a natural baseline on large-scale experiments with images of real scenes using a demanding loss function. We demonstrate that learning a model that is conditioned on scene produces improved results. We show how to incorporate a desired color histogram into the objective function, and that doing so can lead to further improvements in results. Aditya Deshpande, Jason Rock, David A. Forsyth |
ICCV | 3 |
| 2015 | Projectibles: Optimizing Surface Color For ProjectionabstractTypically video projectors display images onto white screens, which can result in a washed out image. Projectibles algorithmically control the display surface color to increase the contrast and resolution. By combining a printed image with projected light, we can create animated, high resolution, high dynamic range visual experiences for video sequences. We present two algorithms for separating an input video sequence into a printed component and projected component, maximizing the combined contrast and resolution while minimizing any visual artifacts introduced from the decomposition. We present empirical measurements of real-world results of six example video sequences, subjective viewer feedback ratings, and we discuss the benefits and limitations of Projectibles. This is the first approach to combine a static display with a dynamic display for the display of video, and the first to optimize surface color for projection of video. Brett R. Jones, Rajinder Sodhi, Pulkit Budhiraja, Kevin Karsch, Brian P. Bailey, David A. Forsyth |
UIST | 6 |
| 2015 | State of the JournalabstractReports on the state of the journal. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | 30Hz Object Detection with DPM V5
Mohammad Amin Sadeghi, David A. Forsyth |
ECCV (1) | 2 |
| 2014 | Human parsing with a cascade of hierarchical poselet based prunersabstractWe address the problem of human parsing using part-based models. In particular, we consider part-based models that exploit rich pairwise relationship between parts, e.g. the color symmetry between left/right limbs. This poses a computational challenge since the state space of each part is very large, and algorithmic tricks (e.g. the distance transform) cannot be applied to handle these types of pairwise relationships. We propose to prune the state space of each part using a cascade of pruners. These pruners can filter out 99.6% of the states per part to about 500 states per part, while keeping the ground-truth states in the pruned state most of the time. In the pruned space, we can afford to apply human parsing models with more complex pairwise relationships between parts, such as the color symmetry. We demonstrate our method on a challenging human parsing dataset. Duan Tran, Yang Wang 0003, David A. Forsyth |
ICME | 3 |
| 2014 | Recognizing activities in multiple views with fusion of frame judgments
Selen Pehlivan, David A. Forsyth |
Image Vis. Comput. | 2 |
| 2014 | Editorial: State of the JournalabstractThe two factors that make TPAMI a wonderful journal are largely immune to disruption by a change of editors. Our community is a fertile source of exciting intellectual creations and scientific discoveries, and this factor ensures there are fine papers for the journal to publish. The other factor is the large community of volunteers who find and promote strong papers. The journal owes a great deal to the tremendous efforts, skill, and professionalism of the Associate Editors in Chief (AEICs). In 2011, TPAMI received 944 submissions, of which 171 were accepted. On average, from submission to first decision took 4.8 months, to accept took 10 months, to online publication 11.1 months, and to paper 17.5 months. Note that these numbers are not cumulative. In 2012, TPAMI received 1,033 submissions, of which 166 have thus far been accepted. On average, from submission to first decision took 3.6 months, to accept took 7.8 months, to online publication 8.4 months, and to paper 14.6 months. You can see the effect of the extra pages that Ramin organized. Figures for 2013 are not yet in, but as of September there were 703 submissions, of which 14 had already been accepted. These statistics suggest that the journal is generally efficient at handling papers. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Video Event Detection: From Subvolume Localization to Spatiotemporal Path SearchabstractAlthough sliding window-based approaches have been quite successful in detecting objects in images, it is not a trivial problem to extend them to detecting events in videos. We propose to search for spatiotemporal paths for video event detection. This new formulation can accurately detect and locate video events in cluttered and crowded scenes, and is robust to camera motions. It can also well handle the scale, shape, and intraclass variations of the event. Compared to event detection using spatiotemporal sliding windows, the spatiotemporal paths correspond to the event trajectories in the video space, thus can better handle events composed by moving objects. We prove that the proposed search algorithm can achieve the global optimal solution with the lowest complexity. Experiments are conducted on realistic video data sets with different event detection tasks, such as anomaly event detection, walking person detection, and running detection. Our proposed method is compatible with different types of video features or object detectors and robust to false and missed local detections. It significantly improves the overall detection and localization accuracy over the state-of-the-art methods. Du Tran, Junsong Yuan 0001, David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | ConstructAide: analyzing and visualizing construction sites through photographs and building modelsabstractWe describe a set of tools for analyzing, visualizing, and assessing architectural/construction progress with unordered photo collections and 3D building models. With our interface, a user guides the registration of the model in one of the images, and our system automatically computes the alignment for the rest of the photos using a novel Structure-from-Motion (SfM) technique; images with nearby viewpoints are also brought into alignment with each other. After aligning the photo(s) and model(s), our system allows a user, such as a project manager or facility owner, to explore the construction site seamlessly in time, monitor the progress of construction, assess errors and deviations, and create photorealistic architectural visualizations. These interactions are facilitated by automatic reasoning performed by our system: static and dynamic occlusions are removed automatically, rendering information is collected, and semantic selection tools help guide user input. We also demonstrate that our user-assisted SfM method outperforms existing techniques on both real-world construction data and established multi-view datasets. David A. Forsyth |
ACM Trans. Graph. | 1 |
| 2014 | Automatic Scene Inference for 3D Object CompositingabstractWe present a user-friendly image editing system that supports a drag-and-drop object insertion (where the user merely drags objects into the image, and the system automatically places them in 3D and relights them appropriately), postprocess illumination editing, and depth-of-field manipulation. Underlying our system is a fully automatic technique for recovering a comprehensive 3D scene model (geometry, illumination, diffuse albedo, and camera parameters) from a single, low dynamic range photograph. This is made possible by two novel contributions: an illumination inference algorithm that recovers a full lighting model of the scene (including light sources that are not directly visible in the photograph), and a depth estimation algorithm that combines data-driven depth transfer with geometric reasoning about the scene layout. A user study shows that our system produces perceptually convincing results, and achieves the same level of realism as techniques that require significant user interaction. Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr 0001, Hailin Jin, Rafael Fonte, Michael Sittig, David A. Forsyth |
ACM Trans. Graph. | 8 |
| 2013 | BeThere: 3D mobile collaboration with spatial inputabstractWe present BeThere, a proof-of-concept system designed to explore 3D input for mobile collaborative interactions. With BeThere, we explore 3D gestures and spatial input which allow remote users to perform a variety of virtual interactions in a local user's physical environment. Our system is completely self-contained and uses depth sensors to track the location of a user's fingers as well as to capture the 3D shape of objects in front of the sensor. We illustrate the unique capabilities of our system through a series of interactions that allow users to control and manipulate 3D virtual content. We also provide qualitative feedback from a preliminary user study which confirmed that users can complete a shared collaborative task using our system. Rajinder Sodhi, Brett R. Jones, David A. Forsyth, Brian P. Bailey, Giuliano Maciocci |
CHI | 3 |
| 2013 | Non-parametric Filtering for Geometric Detail Extraction and Material RepresentationabstractGeometric detail is a universal phenomenon in real world objects. It is an important component in object modeling, but not accounted for in current intrinsic image works. In this work, we explore using a non-parametric method to separate geometric detail from intrinsic image components. We further decompose an image as albedo * (coarse-scale shading + shading detail). Our decomposition offers quantitative improvement in albedo recovery and material classification. Our method also enables interesting image editing activities, including bump removal, geometric detail smoothing/enhancement and material transfer. Zicheng Liao, Jason Rock, Yang Wang 0003, David A. Forsyth |
CVPR | 4 |
| 2013 | Large multi-class image categorization with ensembles of label treesabstractWe consider sublinear test-time algorithms for image categorization when the number of classes is very large. Our method builds upon the label tree approach proposed in [1], which decomposes the label set into a tree structure and classify a test example by traversing the tree. Even though this method achieves logarithmic run-time, its performance is limited by the fact that any errors made in an internal node of the tree cannot be recovered. In this paper, we propose label forests - ensembles of label trees. Each tree in a label forest will decompose the label set in a slightly different way. The final classification decision is made by aggregating information across all trees in the label forest. The test running time of label forest is still logarithmic in the number of categories. But using an ensemble of label trees achieves much better performance in terms of accuracies. We demonstrate our approach on an image classification task that involves 1000 categories. Yang Wang 0003, David A. Forsyth |
ICME | 2 |
| 2013 | Fast Template Evaluation with Vector QuantizationabstractApplying linear templates is an integral part of many object detection systems and accounts for a significant portion of computation time. We describe a method that achieves a substantial end-to-end speedup over the best current methods, without loss of accuracy. Our method is a combination of approximating scores by vector quantizing feature windows and a number of speedup techniques including cascade. Our procedure allows speed and accuracy to be traded off in two ways: by choosing the number of Vector Quantization levels, and by choosing to rescore windows or not. Our method can be directly plugged into any recognition system that relies on linear templates. We demonstrate our method to speed up the original Exemplar SVM detector [1] by an order of magnitude and Deformable Part models [2] by two orders of magnitude with no loss of accuracy. Mohammad Amin Sadeghi, David A. Forsyth |
NIPS | 2 |
| 2013 | TPAMI CVPR Special SectionabstractThe articles in this special issue include papers from the CVPR'11 conference which was held in Colorado Spring, CO, June 2011. Pedro F. Felzenszwalb, David A. Forsyth, Pascal Fua, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Editor's noteabstractJiri Matas has served two terms as Associate Editor in Chief, and the rules require he leave the role. His contribution to the smooth running of the journal has been spectacular. I know our community understands and appreciates the work that Jiri has done. As of writing, I cannot announce the name of a new AEIC, but I expect to do so shortly. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Editor's note [new Editorial Board members announced]abstractThe Editor-in-Chief announces that Kalle Astrom, Karsten Borgwardt, Francois Fleuret, Gert Lankriet, Deva Ramanan, Peter Sturm, Rene Vidal, and Ruigang Yang have agreed to join the Editorial Board. They will be handling a broad range of papers, primarily focused on computer vision. Brief biographies of these distinguished additions to our masthead appear here. David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Improved Object Categorization and Detection Using Comparative Object SimilarityabstractDue to the intrinsic long-tailed distribution of objects in the real world, we are unlikely to be able to train an object recognizer/detector with many visual examples for each category. We have to share visual knowledge between object categories to enable learning with few or no training examples. In this paper, we show that local object similarity information--statements that pairs of categories are similar or dissimilar--is a very useful cue to tie different categories to each other for effective knowledge transfer. The key insight: Given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. To exploit this category-dependent similarity regularization, we develop a regularized kernel machine algorithm to train kernel classifiers for categories with few or no training examples. We also adapt the state-of-the-art object detector to encode object similarity constraints. Our experiments on hundreds of categories from the Labelme dataset show that our regularized kernel classifiers can make significant improvement on object categorization. We also evaluate the improved object detector on the PASCAL VOC 2007 benchmark dataset. Gang Wang 0012, David A. Forsyth, Derek Hoiem |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Recovering free space of indoor scenes from a single imageabstractIn this paper we consider the problem of recovering the free space of an indoor scene from its single image. We show that exploiting the box like geometric structure of furniture and constraints provided by the scene, allows us to recover the extent of major furniture objects in 3D. Our “boxy” detector localizes box shaped objects oriented parallel to the scene across different scales and object types, and thus blocks out the occupied space in the scene. To localize the objects more accurately in 3D we introduce a set of specially designed features that capture the floor contact points of the objects. Image based metrics are not very indicative of performance in 3D. We make the first attempt to evaluate single view based occupancy estimates for 3D errors and propose several task driven performance measures towards it. On our dataset of 592 indoor images marked with full 3D geometry of the scene, we show that: (a) our detector works well using image based metrics; (b) our refinement method produces significant improvements in localization in 3D; and (c) if one evaluates using 3D metrics, our method offers major improvements over other single view based scene geometry estimation methods. Varsha Hedau, Derek Hoiem, David A. Forsyth |
CVPR | 3 |
| 2012 | Building a dictionary of image fragmentsabstractWe show how to build large dictionaries of meaningful image fragments. These fragments could represent objects, objects in a local context, or parts of scenes. Our fragments operate as region-based exemplars, and we show how they can be used for image classification, to localize objects, and to compose new images. While each of these activities has been demonstrated before, each has required manually extracted fragments. Because our method for fragment extraction is automatic it can operate at a large scale. Our method uses recent advances in generic object detection techniques, together with discriminative tests to obtain good, clean fragment sets with extensive diversity. Our fragments are organized by the tags of the source images to build a semantically organized fragment table. A good set of fragment exemplars describes only the object, rather than object+context. Context could help identify an object; but it could also contribute noise, because other objects might appear in the same context. We show a slight improvement in classification performance by two standard exemplar matching methods using our fragment dictionary over such methods using image exemplars. This suggests that knowing the support of an exemplar is valuable. Furthermore, we demonstrate our automatically built fragment dictionary is capable of good localization. Finally, our fragment dictionary supports a keyword based fragment search system, which allows artists to get the fragments they need to make image collages. Zicheng Liao, Ali Farhadi, Yang Wang 0003, Ian Endres, David A. Forsyth |
CVPR | 5 |
| 2012 | Attribute Discovery via Predictable Discriminative Binary Codes
Mohammad Rastegari, Ali Farhadi, David A. Forsyth |
ECCV (6) | 3 |
| 2012 | Around device interaction for multiscale navigationabstractIn this paper we study the design space of free-space interactions for multiscale navigation afforded by mobile depth sensors. Such interactions will have a greater working volume, more fluid control and avoid screen occlusion effects intrinsic to touch screens. This work contributes the first study to show that mobile free-space interactions can be as good as touch. We also analyze sensor orientation and interaction volume usage, resulting in strong implications for how sensors should be placed on mobile devices. We describe a user study evaluating mobile free-space navigation techniques and the impacts of sensor orientation on user experience. Finally, we discuss guidelines for future mobile free-space interaction techniques and sensor design. Brett R. Jones, Rajinder Sodhi, David A. Forsyth, Brian P. Bailey, Giuliano Maciocci |
Mobile HCI | 3 |
| 2012 | Discriminative hierarchical part-based models for human parsing and action recognition
Yang Wang 0003, Duan Tran, Zicheng Liao, David A. Forsyth |
J. Mach. Learn. Res. | 4 |
| 2012 | Learning Image Similarity from Flickr Groups Using Fast Kernel MachinesabstractMeasuring image similarity is a central topic in computer vision. In this paper, we propose to measure image similarity by learning from the online Flickr image groups. We do so by: Choosing 103 Flickr groups, building a one-versus-all multiclass classifier to classify test images into a group, taking the set of responses of the classifiers as features, calculating the distance between feature vectors to measure image similarity. Experimental results on the Corel dataset and the PASCAL VOC 2007 dataset show that our approach performs better on image matching, retrieval, and classification than using conventional visual features. To build our similarity measure, we need one-versus-all classifiers that are accurate and can be trained quickly on very large quantities of data. We adopt an SVM classifier with a histogram intersection kernel. We describe a novel fast training algorithm for this classifier: the Stochastic Intersection Kernel MAchine (SIKMA) training algorithm. This method can produce a kernel classifier that is more accurate than a linear classifier on tens of thousands of examples in minutes. Gang Wang 0012, Derek Hoiem, David A. Forsyth |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | A Subdivision-Based Representation for Vector Image EditingabstractVector graphics has been employed in a wide variety of applications due to its scalability and editability. Editability is a high priority for artists and designers who wish to produce vector-based graphical content with user interaction. In this paper, we introduce a new vector image representation based on piecewise smooth subdivision surfaces, which is a simple, unified and flexible framework that supports a variety of operations, including shape editing, color editing, image stylization, and vector image processing. These operations effectively create novel vector graphics by reusing and altering existing image vectorization results. Because image vectorization yields an abstraction of the original raster image, controlling the level of detail of this abstraction is highly desirable. To this end, we design a feature-oriented vector image pyramid that offers multiple levels of abstraction simultaneously. Our new vector image representation can be rasterized efficiently using GPU-accelerated subdivision. Experiments indicate that our vector image representation achieves high visual quality and better supports editing operations than existing representations. Zicheng Liao, Hugues Hoppe, David A. Forsyth, Yizhou Yu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Still looking at peopleabstractThere is a great need for programs that can describe what people are doing from video. Among other applications, such programs could be used to search for scenes in consumer video; in surveillance applications; to support the design of buildings and of public places; to screen humans for diseases; and to build enhanced human computer interfaces. David A. Forsyth |
ICMI | 1 |
| 2011 | Variable-Source Shading Analysis
David A. Forsyth |
Int. J. Comput. Vis. | 1 |
| 2011 | Lightness Recovery for Pictorial Surfaces
Anna Paviotti, David A. Forsyth, Guido M. Cortelazzo |
Int. J. Comput. Vis. | 2 |
| 2011 | Rendering synthetic objects into legacy photographs
Kevin Karsch, Varsha Hedau, David A. Forsyth, Derek Hoiem |
ACM Trans. Graph. | 3 |
| 2010 | Comparative object similarity for improved recognition with few or no examplesabstractLearning models for recognizing objects with few or no training examples is important, due to the intrinsic long-tailed distribution of objects in the real world. In this paper, we propose an approach to use comparative object similarity. The key insight is that: given a set of object categories which are similar and a set of categories which are dissimilar, a good object model should respond more strongly to examples from similar categories than to examples from dissimilar categories. We develop a regularized kernel machine algorithm to use this category dependent similarity regularization. Our experiments on hundreds of categories show that our method can make significant improvement, especially for categories with no examples. Gang Wang 0012, David A. Forsyth, Derek Hoiem |
CVPR | 2 |
| 2010 | Every Picture Tells a Story: Generating Sentences from Images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young 0001, Cyrus Rashtchian, Julia Hockenmaier, David A. Forsyth |
ECCV (4) | 7 |
| 2010 | Thinking Inside the Box: Using Appearance Models and Context Based on Room Geometry
Varsha Hedau, Derek Hoiem, David A. Forsyth |
ECCV (6) | 3 |
| 2010 | Improved Human Parsing with a Full Relational Model
Duan Tran, David A. Forsyth |
ECCV (4) | 2 |
| 2010 | Seeing People in Social Context: Recognizing People and Social Relationships
Gang Wang 0012, Andrew C. Gallagher, Jiebo Luo 0001, David A. Forsyth |
ECCV (5) | 4 |
| 2010 | It's All About the DataabstractModern computer vision research consumes labelled data in quantity, and building datasets has become an important activity. The Internet has become a tremendous resource for computer vision researchers. By seeing the Internet as a vast, slightly disorganized collection of visual data, we can build datasets. The key point is that visual data are surrounded by contextual information like text and HTML tags, which is a strong, if noisy, cue to what the visual data means. In a series of case studies, we illustrate how useful this contextual information is. It can be used to build a large and challenging labelled face dataset with no manual intervention. With very small amounts of manual labor, contextual data can be used together with image data to identify pictures of animals. In fact, these contextual data are sufficiently reliable that a very large pool of noisily tagged images can be used as a resource to build image features, which reliably improve on conventional visual features. By seeing the Internet as a marketplace that can connect sellers of annotation services to researchers, we can obtain accurately annotated datasets quickly and cheaply. We describe methods to prepare data, check quality, and set prices for work for this annotation process. The problems posed by attempting to collect very big research datasets are fertile for researchers because collecting datasets requires us to focus on two important questions: What makes a good picture? What is the meaning of a picture? Tamara L. Berg, Alexander Sorokin, Gang Wang 0012, David A. Forsyth, Derek Hoiem, Ian Endres, Ali Farhadi |
Proc. IEEE | 4 |
| 2009 | Describing objects by their attributesabstractWe propose to shift the goal of recognition from naming to describing. Doing so allows us not only to name familiar objects, but also: to report unusual aspects of a familiar object (“spotty dog”, not just “dog”); to say something about unfamiliar objects (“hairy and four-legged”, not just “unknown”); and to learn how to recognize new objects with few or no visual examples. Rather than focusing on identity assignment, we make inferring attributes the core problem of recognition. These attributes can be semantic (“spotty”) or discriminative (“dogs have it but sheep do not”). Learning attributes presents a major new challenge: generalization across object categories, not just across instances within a category. In this paper, we also introduce a novel feature selection method for learning attributes that generalize well across categories. We support our claims by thorough evaluation that provides insights into the limitations of the standard recognition paradigm of naming and demonstrates the new abilities provided by our attribute-based framework. Ali Farhadi, Ian Endres, Derek Hoiem, David A. Forsyth |
CVPR | 4 |
| 2009 | Building text features for object image classificationabstractWe introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text feature of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well, consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small. Gang Wang 0012, Derek Hoiem, David A. Forsyth |
CVPR | 3 |
| 2009 | A latent model of discriminative aspectabstractRecognition using appearance features is confounded by phenomena that cause images of the same object to look different, or images of different objects to look the same. This may occur because the same object looks different from different viewing directions, or because two generally different objects have views from which they look similar. In this paper, we introduce the idea of discriminative aspect, a set of latent variables that encode these phenomena. Changes in view direction are one cause of changes in discriminative aspect, but others include changes in texture or lighting. However, images are not labelled with relevant discriminative aspect parameters. We describe a method to improve discrimination by inferring and then using latent discriminative aspect parameters. We apply our method to two parallel problems: object category recognition and human activity recognition. In each case, appearance features are powerful given appropriate training data, but traditionally fail badly under large changes in view. Our method can recognize an object quite reliably in a view for which it possesses no training example. Our method also reweights features to discount accidental similarities in appearance. We demonstrate that our method produces a significant improvement on the state of the art for both object and activity recognition. Ali Farhadi, Mostafa Kamali Tabrizi, Ian Endres, David A. Forsyth |
ICCV | 4 |
| 2009 | Recovering the spatial layout of cluttered roomsabstractIn this paper, we consider the problem of recovering the spatial layout of indoor scenes from monocular images. The presence of clutter is a major problem for existing single-view 3D reconstruction algorithms, most of which rely on finding the ground-wall boundary. In most rooms, this boundary is partially or entirely occluded. We gain robustness to clutter by modeling the global room space with a parameteric 3D “box” and by iteratively localizing clutter and refitting the box. To fit the box, we introduce a structured learning algorithm that chooses the set of parameters to minimize error, based on global perspective cues. On a dataset of 308 images, we demonstrate the ability of our algorithm to recover spatial layout in cluttered rooms and show several examples of estimated free space. Varsha Hedau, Derek Hoiem, David A. Forsyth |
ICCV | 3 |
| 2009 | Unlabeled data improvesword predictionabstractLabeling image collections is a tedious task, especially when multiple labels have to be chosen for each image. In this paper we introduce a new framework that extends state of the art models in word prediction to incorporate information from unlabeled examples, using manifold regularization. To the best of our knowledge this is the first semi-supervised multi-task model used in vision problems. The new model can be solved using gradient descent and is fast and efficient. We show remarkable improvements for cases with few labeled examples for challenging multi-task learning problems in vision (predicting words for images and attributes for objects). Nicolas Loeff, Ali Farhadi, Ian Endres, David A. Forsyth |
ICCV | 4 |
| 2009 | Joint learning of visual attributes, object classes and visual saliencyabstractWe present a method to learn visual attributes (eg.“red”, “metal”, “spotted”) and object classes (eg. “car”, “dress”, “umbrella”) together. We assume images are labeled with category, but not location, of an instance. We estimate models with an iterative procedure: the current model is used to produce a saliency score, which, together with a homogeneity cue, identifies likely locations for the object (resp. attribute); then those locations are used to produce better models with multiple instance learning. Crucially, the object and attribute models must agree on the potential locations of an object. This means that the more accurate of the two models can guide the improvement of the less accurate model. Our method is evaluated on two data sets of images of real scenes, one in which the attribute is color and the other in which it is material. We show that our joint learning produces improved detectors. We demonstrate generalization by detecting attribute-object pairs which do not appear in our training data. The iteration gives significant improvement in performance. Gang Wang 0012, David A. Forsyth |
ICCV | 2 |
| 2009 | Learning image similarity from Flickr groups using Stochastic Intersection Kernel MAchinesabstractMeasuring image similarity is a central topic in computer vision. In this paper, we learn similarity from Flickr groups and use it to organize photos. Two images are similar if they are likely to belong to the same Flickr groups. Our approach is enabled by a fast Stochastic Intersection Kernel MAchine (SIKMA) training algorithm, which we propose. This proposed training method will be useful for many vision problems, as it can produce a classifier that is more accurate than a linear classifier, trained on tens of thousands of examples in two minutes. The experimental results show our approach performs better on image matching, retrieval, and classification than using conventional visual features. Gang Wang 0012, Derek Hoiem, David A. Forsyth |
ICCV | 3 |
| 2009 | Generalizing motion edits with Gaussian processesabstractOne way that artists create compelling character animations is by manipulating details of a character's motion. This process is expensive and repetitive. We show that we can make such motion editing more efficient by generalizing the edits an animator makes on short sequences of motion to other sequences. Our method predicts frames for the motion using Gaussian process models of kinematics and dynamics. These estimates are combined with probabilistic inference. Our method can be used to propagate edits from examples to an entire sequence for an existing character, and it can also be used to map a motion from a control character to a very different target character. The technique shows good generalization. For example, we show that an estimator, learned from a few seconds of edited example animation using our methods, generalizes well enough to edit minutes of character animation in a high-quality fashion. Learning is interactive: An animator who wants to improve the output can provide small, correcting examples and the system will produce improved estimates of motion. We make this interactive learning process efficient and natural with a fast, full-body IK system with novel features. Finally, we present data from interviews with professional character animators that indicate that generalizing and propagating animator edits can save artists significant time and work. Leslie Ikemoto, Okan Arikan, David A. Forsyth |
ACM Trans. Graph. | 3 |
| 2008 | Object image retrieval by exploiting online knowledge resourcesabstractWe describe a method to retrieve images found on web pages with specified object class labels, using an analysis of text around the image and of image appearance. Our method determines whether an object is both described in text and appears in a image using a discriminative image model and a generative text model. Our models are learnt by exploiting established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for image). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of ten animal categories; on this data set, we outperform previous approaches [7, 33]. We have also collected five more categories. Experimental results show the effectiveness of our approach on this new data set. Gang Wang 0012, David A. Forsyth |
CVPR | 2 |
| 2008 | ManifoldBoost: stagewise function approximation for fully-, semi- and un-supervised learningabstractWe describe a manifold learning framewor that naturally accommodates supervised learning, partially supervised learning and unsupervised clustering as particular cases. Our method chooses a function by minimizing loss subject to a manifold regularization penalty. This augmented cost is minimized using a greedy, stagewise, functional minimization procedure, as in Gradientboost. Each stage of boosting is fast and efficient. We demonstrate our approach using both radial basis function approximations and trees. The performance of our method is at the state of the art on many standard semi-supervised learning benchmarks, and we produce results for large scale datasets. Nicolas Loeff, David A. Forsyth, Deepak Ramachandran |
ICML | 2 |
| 2008 | Searching for Complex Human Activities with No Visual Examples
Nazli Ikizler, David A. Forsyth |
Int. J. Comput. Vis. | 2 |
| 2007 | Unsupervised Segmentation of Objects using Efficient LearningabstractWe describe an unsupervised method to segment objects detected in images using a novel variant of an interest point template, which is very efficient to train and evaluate. Once an object has been detected, our method segments an image using a conditional random field (CRF) model. This model integrates image gradients, the location and scale of the object, the presence of object parts, and the tendency of these parts to have characteristic patterns of edges nearby. We enhance our method using multiple unsegmented images of objects to learn the parameters of the CRF, in an iterative conditional maximization framework. We show quantitative results on images of real scenes that demonstrate the accuracy of segmentation. Himanshu Arora, Nicolas Loeff, David A. Forsyth, Narendra Ahuja |
CVPR | 3 |
| 2007 | Transfer Learning in Sign languageabstractWe build word models for American Sign Language (ASL) that transfer between different signers and different aspects. This is advantageous because one could use large amounts of labelled avatar data in combination with a smaller amount of labelled human data to spot a large number of words in human data. Transfer learning is possible because we represent blocks of video with novel intermediate discriminative features based on splits of the data. By constructing the same splits in avatar and human data and clustering appropriately, our features are both discriminative and semantically similar: across signers similar features imply similar words. We demonstrate transfer learning in two scenarios: from avatar to a frontally viewed human signer and from an avatar to human signer in a 3/4 view. Ali Farhadi, David A. Forsyth, Ryan White |
CVPR | 2 |
| 2007 | Searching Video for Complex Activities with Finite State ModelsabstractWe describe a method of representing human activities that allows a collection of motions to be queried without examples, using a simple and effective query language. Our approach is based on units of activity at segments of the body, that can be composed across space and across the body to produce complex queries. The presence of search units is inferred automatically by tracking the body, lifting the tracks to 3D and comparing to models trained using motion capture data. We show results for a large range of queries applied to a collection of complex motion and activity. Our models of short time scale limb behaviour are built using labelled motion capture set. We compare with discriminative methods applied to tracker data; our method offers significantly improved performance. We show experimental evidence that our method is robust to view direction and is unaffected by the changes of clothing. Nazli Ikizler, David A. Forsyth |
CVPR | 2 |
| 2007 | Configuration Estimates Improve Pedestrian FindingabstractFair discriminative pedestrian finders are now available. In fact, these pedestrian finders make most errors on pedestrians in configurations that are uncommon in the training data, for example, mounting a bicycle. This is undesirable. However, the human configuration can itself be estimated discriminatively using structure learning. We demonstrate a pedestrian finder which first finds the most likely human pose in the window using a discriminative procedure trained with structure learning on a small dataset. We then present features (local histogram of oriented gradient and local PCA of gradient) based on that configuration to an SVM classifier. We show, using the INRIA Person dataset, that estimates of configuration significantly improve the accuracy of a discriminative pedestrian finder. Duan Tran, David A. Forsyth |
NIPS | 2 |
| 2007 | Quick transitions with cached multi-way blendsabstractWe describe a discriminative method for distinguishing natural-looking from unnatural-looking motion. Our method is based on physical and data-driven features of motion to which humans seem sensitive. We demonstrate that our technique is significantly more accurate than current alternatives. Leslie Ikemoto, Okan Arikan, David A. Forsyth |
SI3D | 3 |
| 2007 | Tracking People by Learning Their AppearanceabstractAn open vision problem is to automatically track the articulations of people from a video sequence. This problem is difficult because one needs to determine both the number of people in each frame and estimate their configurations. But, finding people and localizing their limbs is hard because people can move fast and unpredictably, can appear in a variety of poses and clothes, and are often surrounded by limb-like clutter. We develop a completely automatic system that works in two stages; it first builds a model of appearance of each person in a video and then it tracks by detecting those models in each frame ("tracking by model-building and detection"). We develop two algorithms that build models; one bottom-up approach groups together candidate body parts found throughout a sequence. We also describe a top-down approach that automatically builds people-models by detecting convenient key poses within a sequence. We finally show that building a discriminative model of appearance is quite helpful since it exploits structure in a background (without background-subtraction). We demonstrate the resulting tracker on hundreds of thousands of frames of unscripted indoor and outdoor activity, a feature-length film ("Run Lola Run"), and legacy sports footage (from the 2002 World Series and 1998 Winter Olympics). Experiments suggest that our system 1) can count distinct individuals, 2) can identify and track them, 3) can recover when it loses track, for example, if individuals are occluded or briefly leave the view, 4) can identify body configuration accurately, and 5) is not dependent on particular models of human motion. Deva Ramanan, David A. Forsyth, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Capturing and animating occluded clothabstractWe capture the shape of moving cloth using a custom set of color markers printed on the surface of the cloth. The output is a sequence of triangle meshes with static connectivity and with detail at the scale of individual markers in both smooth and folded regions. We compute markers' coordinates in space using correspondence across multiple synchronized video cameras. Correspondence is determined from color information in small neighborhoods and refined using a novel strain pruning process. Final correspondence does not require neighborhood information. We use a novel data driven hole-filling technique to fill occluded regions. Our results include several challenging examples: a wrinkled shirt sleeve, a dancing pair of pants, and a rag tossed onto a cup. Finally, we demonstrate that cloth capture is reusable by animating a pair of pants using human motion capture data. Ryan White, Keenan Crane, David A. Forsyth |
ACM Trans. Graph. | 3 |
| 2006 | Discriminating Image Senses by Clustering with Multimodal Features
Nicolas Loeff, Cecilia O. Alm, David A. Forsyth |
ACL | 3 |
| 2006 | Animals on the WebabstractWe demonstrate a method for identifying images containing categories of animals. The images we classify depict animals in a wide range of aspects, configurations and appearances. In addition, the images typically portray multiple species that differ in appearance (e.g. ukari’s, vervet monkeys, spider monkeys, rhesus monkeys, etc.). Our method is accurate despite this variation and relies on four simple cues: text, color, shape and texture. Visual cues are evaluated by a voting method that compares local image phenomena with a number of visual exemplars for the category. The visual exemplars are obtained using a clustering method applied to text on web pages. The only supervision required involves identifying which clusters of exemplars refer to which sense of a term (for example, "monkey" can refer to an animal or a bandmember). Because our method is applied to web pages with free text, the word cue is extremely noisy. We show unequivocal evidence that visual information improves performance for our task. Our method allows us to produce large, accurate and challenging visual datasets mostly automatically. Tamara L. Berg, David A. Forsyth |
CVPR (2) | 2 |
| 2006 | Searching Off-line Arabic DocumentsabstractCurrently an abundance of historical manuscripts, journals, and scientific notes remain largely unaccessible in library archives. Manual transcription and publication of such documents is unlikely, and automatic transcription with high enough accuracy to support a traditional text search is difficult. In this work we describe a lexicon-free system for performing text queries on off-line printed and handwritten Arabic documents. Our segmentation-based approach utilizes gHMMs with a bigram letter transition model, and KPCA/LDA for letter discrimination. The segmentation stage is integrated with inference. We show that our method is robust to varying letter forms, ligatures, and overlaps. Additionally, we find that ignoring letters beyond the adjoining neighbors has little effect on inference and localization, which leads to a significant performance increase over standard dynamic programming. Finally, we discuss an extension to perform batch searches of large word lists for indexing purposes. Jim Chan, Celal Ziftci, David A. Forsyth |
CVPR (2) | 3 |
| 2006 | Aligning ASL for Statistical Translation Using a Discriminative Word ModelabstractWe describe a method to align ASL video subtitles with a closed-caption transcript. Our alignments are partial, based on spotting words within the video sequence, which consists of joined (rather than isolated) signs with unknown word boundaries. We start with windows known to contain an example of a word, but not limited to it. We estimate the start and end of the word in these examples using a voting method. This provides a small number of training examples (typically three per word). Since there is no shared structure, we use a discriminative rather than a generative word model. While our word spotters are not perfect, they are sufficient to establish an alignment. We demonstrate that quite small numbers of good word spotters results in an alignment good enough to produce simple English-ASL translations, both by phrase matching and using word substitution. Ali Farhadi, David A. Forsyth |
CVPR (2) | 2 |
| 2006 | Combining Cues: Shape from Shading and TextureabstractWe demonstrate a method for reconstructing the shape of a deformed surface from a single view. After decomposing an image into irradiance and albedo components, we combine normal cues from shading and texture to produce a field of unambiguous normals. Using these normals, we reconstruct the 3D geometry. Our method works in two regimes: either requiring the frontal appearance of the texture or building it automatically from a series of images of the deforming texture. We can recover geometry with errors below four percent of object size on arbitrary textures, and estimate specific geometric parameters using a custom texture even more accurately. Ryan White, David A. Forsyth |
CVPR (2) | 2 |
| 2006 | Retexturing Single Views Using Texture and Shading
Ryan White, David A. Forsyth |
ECCV (4) | 2 |
| 2006 | Knowing when to put your foot downabstractFootskate, where a character's foot slides on the ground when it should be planted firmly, is a common artifact resulting from almost any attempt to modify motion capture data. We describe an online method for fixing footskate that requires no manual clean-up. An important part of fixing footskate is determining when the feet should be planted. We introduce an oracle that can automatically detect when foot plants should occur. Our method is more accurate than baseline methods that check the height or speed of the feet. These baseline methods perform especially poorly on noisy or imperfect data, requiring manual fixing. Once trained, our oracle is robust and can be used without manual clean-up, making it suitable for large databases of motion. After the foot plants are detected we use an off-the-shelf inverse kinematics based method to maintain ground contact during each foot plant. Our foot plant detection mechanism coupled with an IK based fixer can be treated as a black box that produces natural-looking motion of the feet, making it suitable for interactive systems. We demonstrate several applications which would produce unrealistic motion without our method. Leslie Ikemoto, Okan Arikan, David A. Forsyth |
SI3D | 3 |
| 2006 | Shape from Texture without Boundaries
Anthony Lobay, David A. Forsyth |
Int. J. Comput. Vis. | 2 |
| 2006 | Building Models of Animals from VideoabstractThis paper argues that tracking, object detection, and model building are all similar activities. We describe a fully automatic system that builds 2D articulated models known as pictorial structures from videos of animals. The learned model can be used to detect the animal in the original video--in this sense, the system can be viewed as a generalized tracker (one that is capable of modeling objects while tracking them). The learned model can be matched to a visual library; here, the system can be viewed as a video recognition algorithm. The learned model can also be used to detect the animal in novel images--in this case, the system can be seen as a method for learning models for object recognition. We find that we can significantly improve the pictorial structures by augmenting them with a discriminative texture model learned from a texture library. We develop a novel texture descriptor that outperforms the state-of-the-art for animal textures. We demonstrate the entire system on real video sequences of three different animals. We show that we can automatically track and identify the given animal. We use the learned models to recognize animals from two data sets; images taken by professional photographers from the Corel collection, and assorted images from the Web returned by Google. We demonstrate quite good performance on both data sets. Comparing our results with simple baselines, we show that, for the Google set, we can detect, localize, and recover part articulations from a collection demonstrably hard for object recognition. Deva Ramanan, David A. Forsyth, Kobus Barnard |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Skeletal Parameter Estimation from Optical Motion Capture DataabstractIn this paper we present an algorithm for automatically estimating a subject's skeletal structure from optical motion capture data. Our algorithm consists of a series of steps that cluster markers into segment groups, determine the topological connectivity between these groups, and locate the positions of their connecting joints. Our problem formulation makes use of fundamental distance constraints that must hold for markers attached to an articulated structure, and we solve the resulting systems using a combination of spectral clustering and nonlinear optimization. We have tested our algorithms using data from both passive and active optical motion capture devices. Our results show that the system works reliably even with as few as one or two markers on each segment. For data recorded from human subjects, the system determines the correct topology and qualitatively accurate structure. Tests with a mechanical calibration linkage demonstrate errors for inferred segment lengths on average of only two percent. We discuss applications of our methods for commercial human figure animation, and for identifying human or animal subjects based on their motion independent of marker placement or feature selection. Adam G. Kirk, James F. O'Brien, David A. Forsyth |
CVPR (2) | 3 |
| 2005 | Skeletal Parameter Estimation from Optical Motion Capture DataabstractIn this paper, we present results of our algorithm for automatically estimating a subject's skeletal structure from optical motion capture data. Our algorithm consists of a series of steps that cluster markers into groups representing body segments, determine their topological connectivity, and locate the positions of the connecting joints. Our results show that the system works reliably even when only one or two markers are attached to each segment. We tested an implementation of this algorithm with both passive and active motion capture data and found it to work well. Its computed skeletal estimates closely match measured values, and the algorithm behaves robustly in the presence of noise, marker occlusion, and other errors typical of motion capture data. Adam G. Kirk, James F. O'Brien, David A. Forsyth |
CVPR (2) | 3 |
| 2005 | Finding GlassabstractThis paper addresses the problem of finding glass objects in images. Visual cues obtained by combining the systematic distortions in background texture occurring at the boundaries of transparent objects with the strong highlights typical of glass surfaces are used to train a hierarchy of classifiers, identify glass edges, and find consistent support regions for these edges. Qualitative and quantitative experiments involving a number of different classifiers and real images are presented. Kenton McHenry, Jean Ponce, David A. Forsyth |
CVPR (2) | 3 |
| 2005 | Detecting, Localizing and Recovering Kinematics of Textured AnimalsabstractWe develop and demonstrate an object recognition system capable of accurately detecting, localizing, and recovering the kinematic configuration of textured animals in real images. We build a deformation model of shape automatically from videos of animals and an appearance model of texture from a labeled collection of animal images, and combine the two models automatically. We develop a simple texture descriptor that outperforms the state of the art. We test our animal models on two datasets; images taken by professional photographers from the Corel collection, and assorted images from the Web returned by Google. We demonstrate quite good performance on both datasets. Comparing our results with simple baselines, we show that for the Google set, we can recognize objects from a collection demonstrably hard for object recognition. Deva Ramanan, David A. Forsyth, Kobus Barnard |
CVPR (2) | 2 |
| 2005 | Strike a Pose: Tracking People by Finding Stylized PosesabstractWe develop an algorithm for finding and kinematically tracking multiple people in long sequences. Our basic assumption is that people tend to take on certain canonical poses, even when performing unusual activities like throwing a baseball or figure skating. We build a person detector that quite accurately detects and localizes limbs of people in lateral walking poses. We use the estimated limbs from a detection to build a discriminative appearance model; we assume the features that discriminate a figure in one frame will discriminate the figure in other frames. We then use the models as limb detectors in a pictorial structure framework, detecting figures in unrestricted poses in both previous and successive frames. We have run our tracker on hundreds of thousands of frames, and present and apply a methodology for evaluating tracking on such a large scale. We test our tracker on real sequences including a feature-length film, an hour of footage from a public park, and various sports sequences. We find that we can quite accurately automatically find and track multiple people interacting with each other while performing fast and unusual motions. Deva Ramanan, David A. Forsyth, Andrew Zisserman |
CVPR (1) | 2 |
| 2005 | Tracking People and Recognizing Their ActivitiesabstractWe present a system for automatic people tracking and activity recognition. Our basic approach to people-tracking is to build an appearance model for the person in the video. The video illustrates our method of using a stylized-pose detector. Our system builds a model of limb appearance from those sparse stylized detections. Our algorithm then reprocesses the video, using the learned appearance models to find people in unrestricted configuration. We can use our tracker to recover 3D configurations and activity labels. We assume we have a motion capture library where the 3D poses have been labeled offline with activity descriptions. Deva Ramanan, David A. Forsyth, Andrew Zisserman |
CVPR (2) | 2 |
| 2005 | Searching for Character ModelsabstractWe introduce a method to automatically improve character models for a handwritten script without the use of transcriptions and using a minimum of document specific training data. We show that we can use searches for the words in a dictionary to identify portions of the document whose transcriptions are unambiguous. Using templates extracted from those regions, we retrain our character prediction model to drastically improve our search retrieval performance for words in the document. Jaety Edwards, David A. Forsyth |
NIPS | 2 |
| 2005 | Efficient Unsupervised Learning for Localization and Detection in Object CategoriesabstractWe describe a novel method for learning templates for recognition and localization of objects drawn from categories. A generative model repre- sents the configuration of multiple object parts with respect to an object coordinate system; these parts in turn generate image features. The com- plexity of the model in the number of features is low, meaning our model is much more efficient to train than comparative methods. Moreover, a variational approximation is introduced that allows learning to be or- ders of magnitude faster than previous approaches while incorporating many more features. This results in both accuracy and localization im- provements. Our model has been carefully tested on standard datasets; we compare with a number of recent template models. In particular, we demonstrate state-of-the-art results for detection and localization. Nicolas Loeff, Himanshu Arora, Alexander Sorokin, David A. Forsyth |
NIPS | 4 |
| 2005 | Fast and detailed approximate global illumination by irradiance decompositionabstractIn this paper we present an approximate method for accelerated computation of the final gathering step in a global illumination algorithm. Our method operates by decomposing the radiance field close to surfaces into separate far- and near-field components that can be approximated individually. By computing surface shading using these approximations, instead of directly querying the global illumination solution, we have been able to obtain rendering time speed ups on the order of 10x compared to previous acceleration methods. Our approximation schemes rely mainly on the assumptions that radiance due to distant objects will exhibit low spatial and angular variation, and that the visibility between a surface and nearby surfaces can be reasonably predicted by simple location and orientation-based heuristics. Motivated by these assumptions, our far-field scheme uses scattered-data interpolation with spherical harmonics to represent spatial and angular variation, and our near-field scheme employs an aggressively simple visibility heuristic. For our test scenes, the errors introduced when our assumptions fail do not result in visually objectionable artifacts or easily noticeable deviation from a ground-truth solution. We also discuss how our near-field approximation can be used with standard local illumination algorithms to produce significantly improved images at only negligible additional cost. Okan Arikan, David A. Forsyth, James F. O'Brien |
ACM Trans. Graph. | 2 |
| 2004 | Names and Faces in the News
Tamara L. Berg, Alexander C. Berg, Jaety Edwards, Michael Maire, Ryan White, Yee Whye Teh, Erik G. Learned-Miller, David A. Forsyth |
CVPR (2) | 8 |
| 2004 | Recovering Shape and Irradiance Maps from Rich Dense Texton Fields
Anthony Lobay, David A. Forsyth |
CVPR (1) | 2 |
| 2004 | Towards auto-documentary: tracking the evolution of news storiesabstractNews videos constitute an important source of information for tracking and documenting important events. In these videos, news stories are often accompanied by short video shots that tend to be repeated during the course of the event. Automatic detection of such repetitions is essential for creating auto-documentaries, for alleviating the limitation of traditional textual topic detection methods. In this paper, we propose novel methods for detecting and tracking the evolution of news over time. The proposed method exploits both visual cues and textual information to summarize evolving news stories. Experiments are carried on the TREC-VID data set consisting of 120 hours of news videos from two different channels. Pinar Duygulu, Jia-Yu Pan, David A. Forsyth |
ACM Multimedia | 3 |
| 2004 | Whos In the PictureabstractThe context in which a name appears in a caption provides powerful cues as to who is depicted in the associated image. We obtain 44,773 face im- ages, using a face detector, from approximately half a million captioned news images and automatically link names, obtained using a named en- tity recognizer, with these faces. A simple clustering method can pro- duce fair results. We improve these results significantly by combining the clustering process with a model of the probability that an individual is depicted given its context. Once the labeling procedure is over, we have an accurately labeled set of faces, an appearance model for each individual depicted, and a natural language model that can produce ac- curate results on captions in isolation. Tamara L. Berg, Alexander C. Berg, Jaety Edwards, David A. Forsyth |
NIPS | 4 |
| 2004 | Making Latin Manuscripts Searchable using gHMMsabstractWe describe a method that can make a scanned, handwritten mediaeval latin manuscript accessible to full text search. A generalized HMM is fitted, using transcribed latin to obtain a transition model and one exam- ple each of 22 letters to obtain an emission model. We show results for unigram, bigram and trigram models. Our method transcribes 25 pages of a manuscript of Terence with fair accuracy (75% of letters correctly transcribed). Search results are very strong; we use examples of vari- ant spellings to demonstrate that the search respects the ink of the doc- ument. Furthermore, our model produces fair searches on a document from which we obtained no training data. Jaety Edwards, Yee Whye Teh, David A. Forsyth, Roger Bock, Michael Maire, Grace Vesom |
NIPS | 3 |
| 2003 | The Effects of Segmentation and Feature Choice in a Translation Model of Object RecognitionabstractWe work with a model of object recognition where words must be placed on image regions. This approach means that large scale experiments are relatively easy, so we can evaluate the effects of various early and midlevel vision algorithms on recognition performance. We evaluate various image segmentation algorithms by determining word prediction accuracy for images segmented in various ways and represented by various features. We take the view that good segmentations respect object boundaries, and so word prediction should be better for a better segmentation. However, it is usually very difficult in practice to obtain segmentations that do not break up objects, so most practitioners attempt to merge segments to get better putative object representations. We demonstrate that our paradigm of word prediction easily allows us to predict potentially useful segment merges, even for segments that do not look similar (for example, merging the black and white halves of a penguin is not possible with feature-based segmentation; the main cue must be "familiar configuration"). These studies focus on unsupervised learning of recognition. However, we show that word prediction can be markedly improved by providing supervised information for a relatively small number of regions together with large quantities of unsupervised information. This supervisory information allows a better and more discriminative choice of features and breaks possible symmetries. Kobus Barnard, Pinar Duygulu, Raghavendra Guru, Prasad Gabbur, David A. Forsyth |
CVPR (2) | 5 |
| 2003 | Finding and Tracking People from the Bottom UpabstractWe describe a tracker that can track moving people in long sequences without manual initialization. Moving people are modeled with the assumption that, while configuration can vary quite substantially from frame to frame, appearance does not. This leads to an algorithm that firstly builds a model of the appearance of the body of each individual by clustering candidate body segments, and then uses this model to find all individuals in each frame. Unusually, the tracker does not rely on a model of human dynamics to identify possible instances of people; such models are unreliable, because human motion is fast and large accelerations are common. We show our tracking algorithm can be interpreted as a loopy inference procedure on an underlying Bayes net. Experiments on video of real scenes demonstrate that this tracker can (a) count distinct individuals; (b) identify and track them; (c) recover when it loses track, for example, if individuals are occluded or briefly leave the view; (d) identify the configuration of the body largely correctly; and (e) is not dependent on particular models of human motion. Deva Ramanan, David A. Forsyth |
CVPR (2) | 2 |
| 2003 | Using Temporal Coherence to Build Models of AnimalsabstractWe describe a system that can build appearance models of animals automatically from a video sequence of the relevant animal with no explicit supervisory information. The video sequence need not have any form of special background. Animals are modeled as a 2D kinematic chain of rectangular segments, where the number of segments and the topology of the chain are unknown. The system detects possible segments, clusters segments whose appearance is coherent over time, and then builds a spatial model of such segment clusters. The resulting representation of the spatial configuration of the animal in each frame can be seen either as a track - in which case the system described should be viewed as a generalized tracker, that is capable of modeling objects while tracking them - or as the source of an appearance model which can be used to build detectors for the particular animal. This is because knowing a video sequence is temporally coherent - i.e. that a particular animal is present through the sequence - is a strong supervisory signal. The method is shown to be successful as a tracker on video sequences of real scenes showing three different animals. For the same reason it is successful as a tracker, the method results in detectors that can be used to find each animal fairly reliably within the Corel collection of images. Deva Ramanan, David A. Forsyth |
ICCV | 2 |
| 2003 | Automatic Annotation of Everyday MovementsabstractThis paper describes a system that can annotate a video sequence with: a description of the appearance of each actor; when the actor is in view; and a representation of the actor’s activity while in view. The system does not require a fixed background, and is automatic. The system works by (1) tracking people in 2D and then, using an annotated motion capture dataset, (2) synthesizing an annotated 3D motion sequence matching the 2D tracks. The 3D motion capture data is manually annotated off-line using a class structure that describes everyday motions and allows mo- tion annotations to be composed — one may jump while running, for example. Descriptions computed from video of real motions show that the method is accurate. Deva Ramanan, David A. Forsyth |
NIPS | 2 |
| 2003 | Matching Words and Pictures
Kobus Barnard, Pinar Duygulu, David A. Forsyth, Nando de Freitas, David M. Blei, Michael I. Jordan |
J. Mach. Learn. Res. | 3 |
| 2003 | Motion synthesis from annotationsabstractThis paper describes a framework that allows a user to synthesize human motion while retaining control of its qualitative properties. The user paints a timeline with annotations --- like walk, run or jump --- from a vocabulary which is freely chosen by the user. The system then assembles frames from a motion database so that the final motion performs the specified actions at specified times. The motion can also be forced to pass through particular configurations at particular times, and to go to a particular position and orientation. Annotations can be painted positively (for example, must run), negatively (for example, may not run backwards) or as a don't-care . The system uses a novel search method, based around dynamic programming at several scales, to obtain a solution efficiently so that authoring is interactive. Our results demonstrate that the method can generate smooth, natural-looking motion.The annotation vocabulary can be chosen to fit the application, and allows specification of composite motions (run and jump simultaneously, for example). The process requires a collection of motion data that has been annotated with the chosen vocabulary. This paper also describes an effective tool, based around repeated use of support vector machines, that allows a user to annotate a large collection of motions quickly and easily so that they may be used with the synthesis algorithm. Okan Arikan, David A. Forsyth, James F. O'Brien |
ACM Trans. Graph. | 2 |
| 2002 | Object Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary
Pinar Duygulu, Kobus Barnard, João F. G. de Freitas, David A. Forsyth |
ECCV (4) | 4 |
| 2002 | Shape from Texture without Boundaries
David A. Forsyth |
ECCV (3) | 1 |
| 2002 | Interactive motion generation from examplesabstractThere are many applications that demand large quantities of natural looking motion. It is difficult to synthesize motion that looks natural, particularly when it is people who must move. In this paper, we present a framework that generates human motions by cutting and pasting motion capture data. Selecting a collection of clips that yields an acceptable motion is a combinatorial problem that we manage as a randomized search of a hierarchy of graphs. This approach can generate motion sequences that satisfy a variety of constraints automatically. The motions are smooth and human-looking. They are generated in real time so that we can author complex motions interactively. The algorithm generates multiple motions that satisfy a given set of constraints, allowing a variety of choices for the animator. It can easily synthesize multiple motions that interact with each other using constraints. This framework allows the extensive re-use of motion capture data for new purposes. Okan Arikan, David A. Forsyth |
ACM Trans. Graph. | 2 |
| 2001 | Clustering ArtabstractWe extend a recently developed method (K. Barnard and D. Forsyth, 2001) for learning the semantics of image databases using text and pictures. We incorporate statistical natural language processing in order to deal with free text. We demonstrate the current system on a difficult dataset, namely 10000 images of work from the Fine Arts Museum of San Francisco. The images include line drawings, paintings, and pictures of sculpture and ceramics. Many of the images have associated free text which varies greatly from physical description to interpretation and mood. We use WordNet to provide semantic grouping information and to help disambiguate word senses, as well as emphasize the hierarchical nature of semantic relationships. This allows us to impose a natural structure on the image collection that reflects semantics to a considerable degree. Our method produces a joint probability distribution for words and picture elements. We demonstrate that this distribution can be used: (a) to provide illustrations for given captions, and (b) to generate words for images outside the training set. Results from this annotation process yield a quantitative study of our method. Finally, the annotation process can be seen as a form of object recognizer that has been learned through a partially supervised process. Kobus Barnard, Pinar Duygulu, David A. Forsyth |
CVPR (2) | 3 |
| 2001 | Mixtures of Trees for Object RecognitionabstractEfficient detection of objects in images is complicated by variations of object appearance due to intra-class object differences, articulation, lighting, occlusions, and aspect variations. To reduce the search required for detection, we employ the bottom-up approach where we find candidate image features and associate some of them with parts of the object model. We represent objects as collections of local features, and would like to allow any of them to be absent, with only a small subset sufficient for detection;furthermore, our model should allow efficient correspondence search. We propose a model, Mixture of Trees, that achieves these goals. With a mixture of trees, we can model the individual appearances of the features, relationships among them, and the aspect, and handle occlusions. Independences captured in the model make efficient inference possible. In our earlier work, we have shown that mixtures of trees can be used to model objects with a natural tree structure, in the context of human tracking. Now we show that a natural tree structure is not required, and use a mixture of trees for both frontal and view-invariant face detection. We also show that by modeling faces as collections of features we can establish an intrinsic coordinate frame for a face, and estimate the out-of-plane rotation of a face. Sergey Ioffe, David A. Forsyth |
CVPR (2) | 2 |
| 2001 | Learning the Semantics of Words and PicturesabstractWe present a statistical model for organizing image collections which integrates semantic information provided by associate text and visual information provided by image features. The model is very promising for information retrieval tasks such as database browsing and searching for images based on text and/or image features. Furthermore, since the model learns relationships between text and image features, it can be used for novel applications such as associating words with pictures, and unsupervised learning for object recognition. Kobus Barnard, David A. Forsyth |
ICCV | 2 |
| 2001 | Shape from Texture and IntegrabilityabstractWe describe a shape from texture method that constructs a maximum a posteriori estimate of surface coefficients using both the deformation of individual texture elements-as in local methods-and the overall distribution of elements-as in global methods. The method described applies to a much larger family of textures than any previous method, local or global. We demonstrate an analogy with shape from shading, and use this to produce a numerical method. Examples of reconstructions for synthetic images of surfaces are provided, and compared with ground truth. The method is defined for orthographic views, but can be generalised to perspective views simply. David A. Forsyth |
ICCV | 1 |
| 2001 | Noise in Bilinear ProblemsabstractDespite the wide application of bilinear problems to problems both in computer vision and in other fields, their behaviour under the effects of noise is still poorly understood. In this paper, we show analytically that marginal distributions on the solution components of a bilinear problem can be bimodal, even with Gaussian measurement error. We demonstrate and compare three different methods of estimating the covariance of a solution. We show that the Hessian at the mode substantially underestimates covariance. Many problems in computer vision can be posed as bilinear problems: i.e. one must find a solution to a set of equations of the form. John A. Haddon, David A. Forsyth |
ICCV | 2 |
| 2001 | Human Tracking with Mixtures of Trees
Sergey Ioffe, David A. Forsyth |
ICCV | 2 |
| 2001 | The Joy of Sampling
David A. Forsyth, John A. Haddon, Sergey Ioffe |
Int. J. Comput. Vis. | 1 |
| 2001 | Probabilistic Methods for Finding People
Sergey Ioffe, David A. Forsyth |
Int. J. Comput. Vis. | 2 |
| 2000 | How Does CONDENSATION Behave with a Finite Number of Samples?
Oliver D. King, David A. Forsyth |
ECCV (1) | 2 |
| 2000 | Sampling plausible solutions to multi-body constraint problemsabstractTraditional collision intensive multi-body simulations are difficult to control due to extreme sensitivity to initial conditions or model parameters. Furthermore, there may be multiple ways to achieve any one goal, and it may be difficult to codify a user's preferences before they have seen the available solutions. In this paper we extend simulation models to include plausible sources of uncertainty, and then use a Markov chain Monte Carlo algorithm to sample multiple animations that satisfy constraints. A user can choose the animation they prefer, or applications can take direct advantage of the multiple solutions. Our technique is applicable when a probability can be attached to each animation, with "good" animations having high probability, and for such cases we provide a definition of physical plausibility for animations. We demonstrate our approach with examples of multi-body rigid-body simulations that satisfy constraints of various kinds, for each case presenting animations that are true to a physical model, are significantly different from each other, and yet still satisfy the constraints. CR Descriptors: I.3.7 [Computer Graphics]: Three-Dimensional Graphics and Realism - Animation; I.3.5 [Computer Graphics]: Computational Geometry and Object Modeling - Physically based modeling; I.6.5 [Simulation and Modeling]: Model Development - Modeling methodologies G.3 [Probability and Statistics]: Probabilistic algorithms; Keywords: plausible motion, Markov chain Monte Carlo, motion synthesis, spacetime constraints 1 Stephen Chenney, David A. Forsyth |
SIGGRAPH | 2 |
| 1999 | Sampling, Resampling and Colour ConstancyabstractWe formulate colour constancy as a problem of Bayesian inference, where one is trying to represent the posterior on possible interpretations given image data. We represent the posterior as a set of samples, drawn from that distribution using a Markov chain Monte Carlo method. We show how to build an efficient sampler. This approach has the advantage that it unifies the constraints on the problem, and represents possible ambiguities. In turn, a good description of possible ambiguities means that new information, instead of producing contradictions, is easily incorporated by resampling existing samples. The method is demonstrated on the case where surfaces seen in two distinct images are later discovered to be the same. We show examples using images of real scenes. David A. Forsyth |
CVPR | 1 |
| 1999 | Bayesian Structure from MotionabstractFormulates structure from motion as a Bayesian inference problem and uses a Markov-chain Monte Carlo sampler to sample the posterior on this problem. This results in a method that can identify both small and large tracker errors and yields reconstructions that are stable in the presence of these errors. Furthermore, the method gives detailed information on the range of ambiguities in structure given a particular data set and requires no special geometric formulation to cope with degenerate situations. Motion segmentation is obtained by a layer of discrete variables associating a point with an object. We demonstrate a sampler that successfully samples an approximation to the marginal on this domain, producing a relatively unambiguous segmentation. David A. Forsyth, Sergey Ioffe, John A. Haddon |
ICCV | 1 |
| 1999 | Finding People by SamplingabstractWe show how to use a sampling method to find sparsely clad people in static images. People are modeled as an assembly of nine cylindrical segments. Segments are found using an EM algorithm and then assembled into hypotheses incrementally, using a learned likelihood model. Each assembly step passes on a set of samples of its likelihood to the next; this yields effective pruning of the space of hypotheses. The collection of available nine-segment hypotheses is then represented by a set of equivalence classes, which yield an efficient pruning process. The posterior for the number of people is obtained from the class representatives. People are counted quite accurately in images of real scenes using a MAP estimate. We show the method allows top-down as well as bottom up reasoning. While the method can be overwhelmed by very large numbers of segments, we show that this problem can be avoided by quite simple pruning steps. Sergey Ioffe, David A. Forsyth |
ICCV | 2 |
| 1999 | Interactive ray tracing with the visibility complex
Franklin S. Cho, David A. Forsyth |
Comput. Graph. | 2 |
| 1999 | Automatic Detection of Human Nudes
David A. Forsyth, Margaret M. Fleck |
Int. J. Comput. Vis. | 1 |
| 1998 | Shape Representations from Shading Primitives
John A. Haddon, David A. Forsyth |
ECCV (2) | 2 |
| 1998 | Shading Primitives: Finding Folds and Shallow GroovesabstractDiffuse interreflections cause effects that make current theories of shape from shading unsatisfactory. We show that distant radiating surfaces produce radiosity effects at low spatial frequencies. This means that, if a shading pattern has a small region of support, unseen surfaces in the environment can only produce effects that vary slowly over the support region. It is therefore relatively easy to construct matching processes for such patterns that are robust to interreflections. We call regions with these patterns "shading primitives". Folds and grooves on surfaces provide two examples of shading primitives; the shading pattern is relatively independent of surface shape at a fold or a groove, and the pattern is localised. We show that the pattern of shading can be predicted accurately by a simple model, and derive a matching process from this model. Both groove and fold matchers are shown to work well on images of real scenes. John A. Haddon, David A. Forsyth |
ICCV | 2 |
| 1998 | Learning to Find Pictures of People
Sergey Ioffe, David A. Forsyth |
NIPS | 2 |
| 1997 | Body plansabstractThis paper describes a representation for people and animals, called a body plan, which is adapted to segmentation and to recognition in complex environments. The representation is an organized collection of grouping hints obtained from a combination of constraints on color and texture and constraints on geometric properties such as the structure of individual parts and the relationships between parts. Body plans can be learned from image data, using established statistical learning techniques. The approach is illustrated with two examples of programs that successfully use body plans for recognition: one example involves determining whether a picture contains a scantily clad human, using a body plan built by hand; the other involves determining whether a picture contains a horse, using a body plan learned from image data. In both cases, the system demonstrates excellent performance on large, uncontrolled test sets and very large and diverse control sets. David A. Forsyth, Margaret M. Fleck |
CVPR | 1 |
| 1997 | Finding People and Animals by Guided AssemblyabstractThis paper describes a new representation for people and animals, called a body plan. The representation is an organized collection of grouping hints obtained from constraints on color, texture, shape, and geometrical relations. Body plans can be learned from image data, using established statistical learning techniques. Body plans are well adapted to segmentation and recognition in complex environments, such as the huge libraries of digitized images now becoming widely available. Two specific applications of body plans are presented: an algorithm that determines whether an image depicts a scantily clad human and an algorithm that learns and uses a body plan to find pictures of horses. Both algorithms demonstrate excellent performance on large, poorly controlled input data. David A. Forsyth, Margaret M. Fleck |
ICIP (3) | 1 |
| 1997 | View-Dependent Culling of Dynamic Systems in Virtual EnvironmentsabstractScalable rendering of virtual environments requires culling objects that have no etTect on the view.This paper explores culling moving objects by not solving the equations of motion of objects that don't tiect the view.While this approach could be scalable for many kinds of environments, it raises two problems: consistency -ensuring that objects that come back into view do so in the right state -and completeness -ensuring that objects that would have entered the view volume as a result of their motions, do so.Solutions to these problems lie in studying the statistics of the motion of objects.We show strategies for addressing the problem of consistency, with a number of examples that illustrate both the difficulties involved in, and the potential gains to be obtained by, formulating a comprehensive approach. Stephen Chenney, David A. Forsyth |
SI3D | 2 |
| 1996 | Finding Naked People
Margaret M. Fleck, David A. Forsyth, Christoph Bregler |
ECCV (2) | 2 |
| 1996 | Finding objects in image databases by groupingabstractRetrieving images from very large collections, using image content as a key, is becoming an important problem. Finding objects in image databases is a big challenge in the field. The paper describes our approach to object recognition, which is distinguished by: a rich involvement of early visual primitives, including color and texture; hierarchical grouping and learning strategies in the classification process; the ability to deal with rather general objects in uncontrolled configurations and contexts. We illustrate these properties with three case studies: one demonstrating the use of color and texture descriptors; one learning scenery concepts using grouped features; and one demonstrating a possible application domain in detecting naked people in a scene. Jitendra Malik, David A. Forsyth, Margaret M. Fleck, Hayit Greenspan, Thomas K. Leung, Chad Carson, Serge J. Belongie, Christoph Bregler |
ICIP (2) | 2 |
| 1996 | Identifying nude picturesabstractThis paper demonstrates an automatic system for telling whether there are naked people present in an image. The approach combines color and texture properties to obtain a mask for skin regions, which is shown to be effective for a wide range of shades and colors of skin. These skin regions are then fed to a specialized grouper, which attempts to group a human figure using geometric constraints on human structure. This approach introduces a new view of object recognition, where an object model is an organized collection of grouping hints obtained from a combination of constraints on color and texture and constraints on geometric properties such as the structure of individual parts and the relationships between parts. The system demonstrates excellent performance on a test set of 565 uncontrolled images of naked people, mostly obtained from the internet, and 4289 assorted control images, drawn from a wide collection of sources. David A. Forsyth, Margaret M. Fleck |
WACV | 1 |
| 1996 | Recognizing algebraic surfaces from their outlines
David A. Forsyth |
Int. J. Comput. Vis. | 1 |
| 1995 | MORSE: An Architecture for 3D Object Recognition Based on Invariants
Joseph L. Mundy, Rupert W. Curwen, Jane Liu, Charlie Rothwell, Andrew Zisserman, David A. Forsyth |
ACCV | 6 |
| 1995 | Class-Based Grouping in Perspective ImagesabstractIn any object recognition system a major and primary task is to associate those image features, within an image of a complex scene, that arise from an individual object. The key idea here is that a geometric class defined in 3D induces relationships in the image which must hold between points on the image outline (the perspective projection of the object). The resulting image constraints enable both identification and grouping of image features belonging to objects of that class. The classes include surfaces of revolution, canal surfaces (pipes) and polyhedra. Recognition proceeds by first recognising an object as belonging to one of the classes (for example a surface of revolution) and subsequently identifying the object (for example as a particular vase). This differs from conventional object recognition systems where recognition is generally targetted at particular objects. These classes also support the computation of 3D invariant descriptions including symmetry axes, canonical coordinate frames and projective signatures. The constraints and grouping methods are viewpoint invariant, and proceed with no information on object pose. We demonstrate the effectiveness of this class-based grouping on real, cluttered scenes using grouping algorithms developed for rotationally symmetric surfaces, canal-surfaces and polyhedra.> Andrew Zisserman, Joseph L. Mundy, David A. Forsyth, Jane Liu, Nic Pillow, Charlie Rothwell, Sven Utcke |
ICCV | 3 |
| 1995 | 3D Object Recognition Using Invariance
Andrew Zisserman, David A. Forsyth, Joseph L. Mundy, Charlie Rothwell, Jane Liu, Nic Pillow |
Artif. Intell. | 2 |
| 1995 | Planar object recognition using projective shape representation
Charlie Rothwell, Andrew Zisserman, David A. Forsyth, Joseph L. Mundy |
Int. J. Comput. Vis. | 3 |
| 1994 | Using global consistency to recognise Euclidean objects with an uncalibrated cameraabstractA recognition strategy consisting of a mixture of indexing on invariants and search, allows objects to be recognised up to a Euclidean ambiguity with an uncalibrated camera. The approach works by using projective invariants to determine all the possible projectively equivalent models for a particular imaged object; then a system of global consistency constraints is used to determine which of these projectively equivalent, but Euclidean distinct, models corresponds to the objects viewed. These constraints follow from properties of the imaging geometry. In particular, a recognition hypothesis is equivalent to an assertion about, among other things, viewing conditions and geometric relationships between objects, and these assertions must be consistent for hypotheses to be correct. The approach is demonstrated to work on images of real scenes consisting of polygonal objects and polyhedra.> David A. Forsyth, Joseph L. Mundy, Andrew Zisserman, Charlie Rothwell |
CVPR | 1 |
| 1994 | Object representation for object recognitionabstractThis paper discusses some representation issues and challenges involved in object recognition. It is intended as a step toward assessing current object representation schemes and proposing design and evaluation criteria for future ones.> Jean Ponce, Ruzena Bajcsy, Dimitris N. Metaxas, Thomas O. Binford, David A. Forsyth, Martial Hebert, Katsushi Ikeuchi, Avinash C. Kak, Linda G. Shapiro, Stan Sclaroff, Alex Pentland, George C. Stockman |
CVPR | 5 |
| 1993 | Efficient recognition of rotationally symmetric surfaces and straight homogeneous generalized cylindersabstractIt is known that rotationally symmetric surfaces can be recognized from their outlines alone, using cross-ratios of bitangent intersections. A successful implementation of this technique is demonstrated using a novel bitangent finder which works on images of real scenes. The stability of the cross-ratios is reported and compared to affine invariants. The recognition technique is shown to extend to the case of straight homogeneous generalized cylinders.> Jane Liu, Joseph L. Mundy, David A. Forsyth, Andrew Zisserman, Charlie Rothwell |
CVPR | 3 |
| 1993 | Recognizing algebraic surfaces from their outlinesabstractThe author shows that the projective invariants of an algebraic surface can be computed from the outline of that surface in a perspective view using an uncalibrated camera, by showing that an outline completely determines the projective geometry of an algebraic surface. A single perspective view of a generic algebraic surface of degree three or greater uniquely determines the projective geometry of the surface. The result holds for an unknown focal point, and an uncalibrated camera. The projective ambiguity is not improved by using a calibrated camera.> David A. Forsyth |
ICCV | 1 |
| 1993 | Extracting projective structure from single perspective views of 3D point setsabstractA number of recent papers have argued that invariants do not exist for three-dimensional point sets in general position, which has often been misinterpreted to mean that invariants cannot be computed for any three-dimensional structure. It is proved by example that although the general statement is true, invariants do exist for structured three-dimensional point sets. Projective invariants are derived for two object classes: the first is for points that lie on the vertices of polyhedra, and the second for objects that are projectively equivalent to ones possessing a bilateral symmetry. The motivations for computing such invariants are twofold: they can be used for recognition, and they can be used to compute projective structure. Examples of invariants computed from real images are given.> Charlie Rothwell, David A. Forsyth, Andrew Zisserman, Joseph L. Mundy |
ICCV | 2 |
| 1992 | Efficient model library access by projectively invariant indexing functionsabstractProjectively invariant shape descriptors allow fast indexing into model libraries without the need for pose computation or camera calibration. Progress in building a model-based vision system for plane objects that uses algebraic projective invariants is described. A brief account of these descriptors is given, and the recognition system is described, giving examples of the invariant techniques working on real images.> Charlie Rothwell, Andrew Zisserman, Joseph L. Mundy, David A. Forsyth |
CVPR | 4 |
| 1992 | Recognising rotationally symmetric surfaces from their outlines
David A. Forsyth, Joseph L. Mundy, Andrew Zisserman, Charlie Rothwell |
ECCV | 1 |
| 1992 | Canonical Frames for Planar Object Recognition
Charlie Rothwell, Andrew Zisserman, David A. Forsyth, Joseph L. Mundy |
ECCV | 3 |
| 1992 | Transformational invariance - a primer
David A. Forsyth, Joseph L. Mundy, Andrew Zisserman |
Image Vis. Comput. | 1 |
| 1992 | Relative motion and pose from arbitrary plane curves
Charlie Rothwell, Andrew Zisserman, Constantinos Marinos, David A. Forsyth, Joseph L. Mundy |
Image Vis. Comput. | 4 |
| 1991 | Using Projective Invariants for Constant Time Library Indexing in Model Based Vision
Charlie Rothwell, Andrew Zisserman, David A. Forsyth, Joseph L. Mundy |
BMVC | 3 |
| 1991 | Projectively invariant representations using implicit algebraic curves
David A. Forsyth, Joseph L. Mundy, Andrew Zisserman |
Image Vis. Comput. | 1 |
| 1991 | Invariant Descriptors for 3D Object Recognition and PoseabstractInvariant descriptors are shape descriptors that are unaffected by object pose, by perspective projection, or by the intrinsic parameters of the camera. These descriptors can be constructed using the methods of invariant theory, which are briefly surveyed. A range of applications of invariant descriptors in 3D model-based vision is demonstrated. First, a model-based vision system that recognizes curved plane objects irrespective of their pose is demonstrated. Curves are not reduced to polyhedral approximations but are handled as objects in their own right. Models are generated directly from image data. Once objects have been recognized, their pose can be computed. Invariant descriptors for 3D objects with plane faces are described. All these ideas are demonstrated using images of real scenes. The stability of a range of invariant descriptors to measurement error is treated in detail.> David A. Forsyth, Joseph L. Mundy, Andrew Zisserman, Chris Coelho, Aaron Heller, Charlie Rothwell |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1991 | Reflections on ShadingabstractIt is demonstrated that mutual illumination can produce significant effects in real scenes. An example is presented to illustrate the difficulties that mutual illumination presents to shape recovery schemes. These effects are qualitatively modeled by the radiosity equation. Using the radiosity equation, the authors predict the occurrence of spectral events in the radiance, namely, discontinuities in the radiance and its derivatives. Experimental evidence establishes the validity of this approach. Mutual illumination can generate discontinuities in the derivatives of radiance unrelated to local geometry. It is argued that it is not possible to obtain veridical dense depth or normal maps from a shading analysis. However, discontinuities in radiance are tractably related to scene geometry and, moreover, can be detected.> David A. Forsyth, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Transformational invariance - a primerabstractAbstract The shape of objects seen in images depends on the viewpoint. This effect confounds recognition. We demonstrate a theoretical framework within which it is possible to construct descriptors for curves which do not vary with viewpoint. These descriptors are known as invariants. We use this framework to construct invariant shape descriptors for plane curves. These invariant shape descriptors make it possible to recognise plane curves, without explicitly determining the relationship between the curve reference frame and the camera coordinate system, and can be used to index quickly and efficiently into a large model base of curves. Many of these ideas are demonstrated by experiments on real image data. David A. Forsyth, Joseph L. Mundy, Andrew Zisserman |
BMVC | 1 |
| 1990 | Relative motion and pose from invariantsabstractProjectively invariant shape descriptors efficiently identify instances of object models in images without reference to object pose. These descriptions rely on frame independent representations of planar curves, using plane conies. We show that object pose can be determined from coplanar curves, given such a frame independent representation. This result is demonstrated for real image data. The shape of objects in images changes as the camera is moved around. This extremely simple observation represents the dominant problem in model based vision. Nielsen [4, 5] first suggested using projectively invariant labels as landmarks for navigation. Recent papers [1, 2] have shown that it is possible to compute shape descriptors of arbitrary plane objects that are unaffected by Andrew Zisserman, Constantinos Marinos, David A. Forsyth, Joseph L. Mundy, Charlie Rothwell |
BMVC | 3 |
| 1990 | Projectively Invariant Representations Using Implicit Algebraic Curves
David A. Forsyth, Joseph L. Mundy, Andrew Zisserman |
ECCV | 1 |
| 1990 | Invariance-a new framework for visionabstractIt is shown that curved planar objects have shape descriptors that are unaffected by the position, orientation and intrinsic parameters of the camera. These shape descriptors can be used to index quickly and efficiently into a large model base of curved planar objects, because their value is independent of pose and unaffected by perspective. Thus, recognition can proceed independent of calculating pose. Object curves are represented using conics, attached with a fitting technique that commutes with projection. This means that the pose of an object can be determined by backprojecting known conics. The authors show examples of recognition and pose determination using real image data.> David A. Forsyth, Joseph L. Mundy, Andrew Zisserman |
ICCV | 1 |
| 1990 | A novel algorithm for color constancy
David A. Forsyth |
Int. J. Comput. Vis. | 1 |
| 1990 | Shape from shading in the light of mutual illumination
David A. Forsyth, Andrew Zisserman |
Image Vis. Comput. | 1 |
| 1989 | Mutual illuminationabstractThe authors report theoretical and experimental results which underline the importance of mutual illumination to visual modules dealing with shape and with surface lightness. The experiments are in good agreement with results obtained with a simple theoretical model. These results show the effects of mutual illumination in pictures of simple objects, and indicate that these effects must be accounted for in modeling image intensities. The data imply that shape from shading based on the image irradiance equation make real errors on images of concave objects, and that edge detectors that respond to only step edges perform badly on polyhedral scenes and waste information.> David A. Forsyth, Andrew Zisserman |
CVPR | 1 |
| 1988 | A Novel Approach To Colour ConstancyabstractBy approaching colour constancy as a problem of predicting colour appearance, we derive the colour constancy equation, which we use to enumerate those properties of illuminant and surface reflectance required for colour constancy. We then use a physical realisability constraint on surface reflectances to construct the set of illuminants under which the image observed can have arisen. Two distinct algorithms arise from employing this constraint in conjunction with the colour constancy equation: the first corresponds to normalisation according to a coefficient rule, the second is considerably more complex, and allows a large number of parameters in the illuminant to be recovered. The simpler algorithm has been tested extensively on images of real Mondriaan’s, taken under different coloured lights and displays good constancy. The results also indicate that good constancy requires that receptoral gain be controlled. David A. Forsyth |
ICCV | 1 |