Allan Douglas Jepson

dblp:93/4241 · also Allan D. Jepson · DBLP profile ↗
← Back
83ranked-venue papers
9as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 A Truncated Newton Method for Optimal Transport
abstract
Developing a contemporary optimal transport (OT) solver requires navigating trade-offs among several critical requirements: GPU parallelization, scalability to high-dimensional problems, theoretical convergence guarantees, empirical performance in terms of precision versus runtime, and numerical stability in practice. With these challenges in mind, we introduce a specialized truncated Newton algorithm for entropic-regularized OT. In addition to proving that locally quadratic convergence is possible without assuming a Lipschitz Hessian, we provide strategies to maximally exploit the high rate of local convergence in practice. Our GPU-parallel algorithm exhibits exceptionally favorable runtime performance, achieving high precision orders of magnitude faster than many existing alternatives. This is evidenced by wall-clock time experiments on 24 problem sets (12 datasets $\times$ 2 cost functions). The scalability of the algorithm is showcased on an extremely large OT problem with $n \approx 10^6$, solved approximately under weak entropic regularization.
Mete Kemertas, Amir-massoud Farahmand, Allan Douglas Jepson
ICLR3
2025 Probabilistic Directed Distance Fields for Ray-Based Shape Representations
abstract
In modern computer vision, the optimal representation of 3D shape remains task-dependent. One fundamental operation applied to such representations is differentiable rendering, which enables learning-based inverse graphics approaches. Standard explicit representations are often easily rendered, but can suffer from limited geometric fidelity, among other issues. On the other hand, implicit representations generally preserve greater fidelity, but suffer from difficulties with rendering, limiting scalability. In this work, we devise Directed Distance Fields (DDFs), which map a ray or oriented point (position and direction) to surface visibility and depth. This enables efficient differentiable rendering, obtaining depth with a single forward pass per pixel, as well as higher-order geometry with only additional backward passes. Using probabilistic DDFs (PDDFs), we can model the inherent discontinuities in the underlying field. We then apply DDFs to single-shape fitting, generative modelling, and 3D reconstruction, showcasing strong performance with simple architectural components via the versatility of our representation. Finally, since the dimensionality of DDFs permits view-dependent geometric artifacts, we conduct a theoretical investigation of the constraints necessary for view consistency. We find a small set of field properties that are sufficient to guarantee a DDF is consistent, without knowing which shape the field is expressing.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Shape-Based Measures Improve Scene Categorization
abstract
Converging evidence indicates that deep neural network models that are trained on large datasets are biased toward color and texture information. Humans, on the other hand, can easily recognize objects and scenes from images as well as from bounding contours. Mid-level vision is characterized by the recombination and organization of simple primary features into more complex ones by a set of so-called Gestalt grouping rules. While described qualitatively in the human literature, a computational implementation of these perceptual grouping rules is so far missing. In this article, we contribute a novel set of algorithms for the detection of contour-based cues in complex scenes. We use the medial axis transform (MAT) to locally score contours according to these grouping rules. We demonstrate the benefit of these cues for scene categorization in two ways: (i) Both human observers and CNN models categorize scenes most accurately when perceptual grouping information is emphasized. (ii) Weighting the contours with these measures boosts performance of a CNN model significantly compared to the use of unweighted contours. Our work suggests that, even though these measures are computed directly from contours in the image, current CNN models do not appear to extract or utilize these grouping cues.
Morteza Rezanejad, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos
abstract
Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction steps in such videos, i.e. the task called key-step localization. Traditional methods for key-step localization require video-level human annotations and thus do not scale to large datasets. In this work, we tackle the problem with no human supervision and introduce StepFormer, a self-supervised model that discovers and localizes instruction steps in a video. StepFormer is a transformer decoder that attends to the video with learnable queries, and produces a sequence of slots capturing the key-steps in the video. We train our system on a large dataset of instructional videos, using their automatically-generated subtitles as the only source of supervision. In particular, we supervise our system with a sequence of text narrations using an order-aware loss function that filters out irrelevant phrases. We show that our model outperforms all previous unsupervised and weakly-supervised approaches on step detection and localization by a large margin on three challenging benchmarks. Moreover, our model demonstrates an emergent property to solve zero-shot multi-step localization and outperforms all relevant baselines at this task.
Nikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Richard P. Wildes, Allan Douglas Jepson
CVPR6
2023 Efficient Flow-Guided Multi-frame De-fencing
abstract
Taking photographs "in-the-wild" is often hindered by fence obstructions that stand between the camera user and the scene of interest, and which are hard or impossible to avoid. De-fencing is the algorithmic process of automatically removing such obstructions from images, revealing the invisible parts of the scene. While this problem can be formulated as a combination of fence segmentation and image inpainting, this often leads to implausible hallucinations of the occluded regions. Existing multi-frame approaches rely on propagating information to a selected keyframe from its temporal neighbors, but they are often inefficient and struggle with alignment of severely obstructed images. In this work we draw inspiration from the video completion literature, and develop a simplified framework for multi-frame de-fencing that computes high quality flow maps directly from obstructed frames, and uses them to accurately align frames. Our primary focus is efficiency and practicality in a real world setting: the input to our algorithm is a short image burst (5 frames) – a data modality commonly available in modern smartphones– and the output is a single reconstructed keyframe, with the fence removed. Our approach leverages simple yet effective CNN modules, trained on carefully generated synthetic data, and outperforms more complicated alternatives real bursts, both quantitatively and qualitatively, while running real-time.
Stavros Tsogkas, Fengjia Zhang, Allan Douglas Jepson, Alex Levinshtein
WACV3
2023 Disentangling Geometric Deformation Spaces in Generative Latent Shape Models
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
Int. J. Comput. Vis.4
2022 P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak Supervision
abstract
In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work leverages intermediate visual observations as supervision, which requires expensive annotation efforts to localize precisely all the instructional steps in training videos. In contrast, we remove the need for expensive temporal video annotations and propose a weakly supervised approach by learning from natural language instructions. Our model is based on a transformer equipped with a memory module, which maps the start and goal observations to a sequence of plausible actions. Furthermore, we augment our model with a probabilistic generative module to capture the uncertainty inherent to procedure planning, an aspect largely overlooked by previous work. We evaluate our model on three datasets and show our weakly-supervised approach outperforms previous fully supervised state-of-the-art models on multiple metrics.
He Zhao 0004, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis, Richard P. Wildes, Allan Douglas Jepson
CVPR6
2022 Representing 3D Shapes with Probabilistic Directed Distance Fields
abstract
Differentiable rendering is an essential operation in modern vision, allowing inverse graphics approaches to 3D understanding to be utilized in modern machine learning frameworks. Explicit shape representations (voxels, point clouds, or meshes), while relatively easily rendered, often suffer from limited geometric fidelity or topological con-straints. On the other hand, implicit representations (occu-pancy, distance, or radiance fields) preserve greater fidelity, but suffer from complex or inefficient rendering processes, limiting scalability. In this work, we endeavour to address both shortcomings with a novel shape representation that allows fast differentiable rendering within an implicit ar-chitecture. Building on implicit distance representations, we define Directed Distance Fields (DDFs), which map an oriented point (position and direction) to surface visibility and depth. Such a field can render a depth map with a single forward pass per pixel, enable differential surface geometry extraction (e.g., surface normals and curvatures) via network derivatives, be easily composed, and permit extraction of classical unsigned distance fields. Using probabilistic DDFs (PDDFs), we show how to model inherent discontinuities in the underlying field. Finally, we apply our method to fitting single shapes, unpaired 3D-aware generative image modelling, and single-image 3D reconstruction tasks, showcasing strong performance with simple architectural components via the versatility of our representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
CVPR4
2022 Flow Graph to Video Grounding for Weakly-Supervised Multi-step Localization
Nikita Dvornik, Isma Hadji, Hai X. Pham, Dhaivat Bhatt, Brais Martínez, Afsaneh Fazly, Allan Douglas Jepson
ECCV (35)7
2022 GraN-GAN: Piecewise Gradient Normalization for Generative Adversarial Networks
abstract
Modern generative adversarial networks (GANs) predominantly use piecewise linear activation functions in discriminators (or critics), including ReLU and LeakyReLU. Such models learn piecewise linear mappings, where each piece handles a subset of the input space, and the gradients per subset are piecewise constant. Under such a class of discriminator (or critic) functions, we present Gradient Normalization (GraN), a novel input-dependent normalization method, which guarantees a piecewise K-Lipschitz constraint in the input space. In contrast to spectral normalization, GraN does not constrain processing at the individual network layers, and, unlike gradient penalties, strictly enforces a piecewise Lipschitz constraint almost everywhere. Empirically, we demonstrate improved image generation performance across multiple datasets (incl. CIFAR-10/100, STL-10, LSUN bedrooms, and CelebA), GAN loss functions, and metrics. Further, we analyze altering the often untuned Lipschitz constant in several standard GANs, not only attaining significant performance gains, but also finding connections between and training dynamics, particularly in low-gradient loss plateaus, with the common Adam optimizer.
Vineeth S. Bhaskara, Tristan Aumentado-Armstrong, Allan Douglas Jepson, Alex Levinshtein
WACV3
2021 Representation Learning via Global Temporal Alignment and Cycle-Consistency
abstract
We introduce a weakly supervised method for representation learning based on aligning temporal sequences (e.g., videos) of the same process (e.g., human action). The main idea is to use the global temporal ordering of latent correspondences across sequence pairs as a supervisory signal. In particular, we propose a loss based on scoring the optimal sequence alignment to train an embedding network. Our loss is based on a novel probabilistic path finding view of dynamic time warping (DTW) that contains the following three key features: (i) the local path routing decisions are contrastive and differentiable, (ii) pairwise distances are cast as probabilities that are contrastive as well, and (iii) our formulation naturally admits a global cycle-consistency loss that verifies correspondences. For evaluation, we consider the tasks of fine-grained action classification, few shot learning, and video synchronization. We report significant performance increases over previous methods. In addition, we report two applications of our temporal alignment framework, namely 3D pose reconstruction and fine-grained audio/visual retrieval.
Isma Hadji, Konstantinos G. Derpanis, Allan Douglas Jepson
CVPR3
2021 Dependency parsing with structure preserving embeddings
abstract
Ákos Kádár, Lan Xiao, Mete Kemertas, Federico Fancellu, Allan Jepson, Afsaneh Fazly. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ákos Kádár, Mete Kemertas, Federico Fancellu, Allan Douglas Jepson, Afsaneh Fazly
EACL5
2021 Personalized Multi-modal Video Retrieval on Mobile Devices
abstract
Current video retrieval systems on mobile devices cannot process complex natural language queries, especially if they contain personalized concepts, such as proper names. To address these shortcomings, we propose an efficient and privacy-preserving video retrieval system that works well with personalized queries containing proper names, without re-training using personalized labelled data from users. Our system first computes an initial ranking of a video collection by using a generic attention-based video-text matching model (i.e., a model designed for non-personalized queries), and then uses a face detector to conduct personalized adjustments to these initial rankings. These adjustments are done by reasoning over the face information from the detector and the attention information provided by the generic model. We show that our system significantly outperforms existing keyword-based retrieval systems, and achieves comparable performance to the generic matching model fine-tuned on plenty of labelled data. Our results suggest that the proposed system can effectively capture both semantic context and personalized information in queries.
Allan Douglas Jepson, Iqbal Mohomed, Konstantinos G. Derpanis, Afsaneh Fazly
ACM Multimedia2
2021 Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers
abstract
In this work, we consider the problem of sequence-to-sequence alignment for signals containing outliers. Assuming the absence of outliers, the standard Dynamic Time Warping (DTW) algorithm efficiently computes the optimal alignment between two (generally) variable-length sequences. While DTW is robust to temporal shifts and dilations of the signal, it fails to align sequences in a meaningful way in the presence of outliers that can be arbitrarily interspersed in the sequences. To address this problem, we introduce Drop-DTW, a novel algorithm that aligns the common signal between the sequences while automatically dropping the outlier elements from the matching. The entire procedure is implemented as a single dynamic program that is efficient and fully differentiable. In our experiments, we show that Drop-DTW is a robust similarity measure for sequence retrieval and demonstrate its effectiveness as a training loss on diverse applications. With Drop-DTW, we address temporal step localization on instructional videos, representation learning from noisy videos, and cross-modal representation learning for audio-visual retrieval and localization. In all applications, we take a weakly- or unsupervised approach and demonstrate state-of-the-art results under these settings.
Nikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg, Allan Douglas Jepson
NeurIPS5
2020 Cycle-Consistent Generative Rendering for 2D-3D Modality Translation
abstract
For humans, visual understanding is inherently generative: given a 3D shape, we can postulate how it would look in the world; given a 2D image, we can infer the 3D structure that likely gave rise to it. We can thus translate between the 2D visual and 3D structural modalities of a given object. In the context of computer vision, this corresponds to a learnable module that serves two purposes: (i) generate a realistic rendering of a 3D object (shape-to-image translation) and (ii) infer a realistic 3D shape from an image (image-to-shape translation). In this paper, we learn such a module while being conscious of the difficulties in obtaining large paired 2D-3D datasets. By leveraging generative domain translation methods, we are able to define a learning algorithm that requires only weak supervision, with unpaired data. The resulting model is not only able to perform 3D shape, pose, and texture inference from 2D images, but can also generate novel textured 3D shapes and renders, similar to a graphics pipeline. More specifically, our method (i) infers an explicit 3D mesh representation, (ii) utilizes example shapes to regularize inference, (iii) requires only an image mask (no keypoints or camera extrinsics), and (iv) has generative capabilities. While prior work explores subsets of these properties, their combination is novel. We demonstrate the utility of our learned representation, as well as its performance on image generation and unpaired 3D shape inference tasks.
Tristan Aumentado-Armstrong, Alex Levinshtein, Stavros Tsogkas, Konstantinos G. Derpanis, Allan Douglas Jepson
3DV5
2019 Scene Categorization From Contours: Medial Axis Based Salience Measures
Morteza Rezanejad, Gabriel Downs, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi
CVPR5
2019 Geometric Disentanglement for Generative Latent Shape Models
abstract
Representing 3D shapes is a fundamental problem in artificial intelligence, which has numerous applications within computer vision and graphics. One avenue that has recently begun to be explored is the use of latent representations of generative models. However, it remains an open problem to learn a generative model of shapes that is interpretable and easily manipulated, particularly in the absence of supervised labels. In this paper, we propose an unsupervised approach to partitioning the latent space of a variational autoencoder for 3D point clouds in a natural way, using only geometric information, that builds upon prior work utilizing generative adversarial models of point sets. Our method makes use of tools from spectral geometry to separate intrinsic and extrinsic shape information, and then considers several hierarchical disentanglement penalties for dividing the latent space in this manner. We also propose a novel disentanglement penalty that penalizes the predicted change in the latent representation of the output,with respect to the latent variables of the initial shape. We show that the resulting latent representation exhibits intuitive and interpretable behaviour, enabling tasks such as pose transfer that cannot easily be performed by models with an entangled representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Allan Douglas Jepson, Sven J. Dickinson
ICCV3
2013 Fast Rigid Motion Segmentation via Incrementally-Complex Local Models
abstract
The problem of rigid motion segmentation of trajectory data under orthography has been long solved for non-degenerate motions in the absence of noise. But because real trajectory data often incorporates noise, outliers, motion degeneracies and motion dependencies, recently proposed motion segmentation methods resort to non-trivial representations to achieve state of the art segmentation accuracies, at the expense of a large computational cost. This paper proposes a method that dramatically reduces this cost (by two or three orders of magnitude) with minimal accuracy loss (from 98.8% achieved by the state of the art, to 96.2% achieved by our method on the standard Hopkins 155 dataset). Computational efficiency comes from the use of a simple but powerful representation of motion that explicitly incorporates mechanisms to deal with noise, outliers and motion degeneracies. Subsets of motion models with the best balance between prediction accuracy and model complexity are chosen from a pool of candidates, which are then used for segmentation.
Fernando Flores-Mangas, Allan Douglas Jepson
CVPR2
2010 Polynomial shape from shading
abstract
We examine the shape from shading problem without boundary conditions as a polynomial system. This view allows, in generic cases, a complete solution for ideal polyhedral objects. For the general case we propose a semidefinite programming relaxation procedure, and an exact line search iterative procedure with a new smoothness term that favors folds at edges. We use this numerical technique to inspect shading ambiguities.
Ady Ecker, Allan Douglas Jepson
CVPR2
2010 Non-rigid structure from locally-rigid motion
abstract
We introduce locally-rigid motion, a general framework for solving the M-point, N-view structure-from-motion problem for unknown bodies deforming under orthography. The key idea is to first solve many local 3-point, N-view rigid problems independently, providing a “soup” of specific, plausibly rigid, 3D triangles. The main advantage here is that the extraction of 3D triangles requires only very weak assumptions: (1) deformations can be locally approximated by near-rigid motion of three points (i.e., stretching not dominant) and (2) local motions involve some generic rotation in depth. Triangles from this soup are then grouped into bodies, and their depth flips and instantaneous relative depths are determined. Results on several sequences, both our own and from related work, suggest these conditions apply in diverse settings - including very challenging ones (e.g., multiple deforming bodies). Our starting point is a novel linear solution to 3-point structure from motion, a problem for which no general algorithms currently exist.
Allan Douglas Jepson, Kiriakos N. Kutulakos
CVPR2
2009 Stochastic Image Denoising
abstract
We present a novel algorithm for image denoising. Our algorithm is based on random walks over arbitrary neighbourhoods surrounding a given pixel. The size and shape of each neighbourhood are determined by the configuration and similarity of nearby pixels. Assuming that pixels within the neighbourhood of x0 are likely to have been generated by the same random process, we want the weights used to mix these pixels during denoising to depend on the similarity between them and x0. At the same time, we require the random walk to follow a smooth path from x0 to any other pixel in the neighbourhood, so the transition probabilities should also depend on the similarity between pairs of neighbouring pixels along any given path. With this in mind, we define a random walk originating at pixel x0 as an ordered sequence of pixels T0,k = {x0,x1, . . . ,xk} visited along the path from x0 to xk. Within this sequence, the probability of a transition between two consecutive pixels x j and x j+1 is defined to be
Francisco J. Estrada, David J. Fleet, Allan Douglas Jepson
BMVC3
2009 Benchmarking Image Segmentation Algorithms
Francisco J. Estrada, Allan Douglas Jepson
Int. J. Comput. Vis.2
2009 The quantitative characterization of the distinctiveness and robustness of local image descriptors
Gustavo Carneiro 0001, Allan Douglas Jepson
Image Vis. Comput.2
2008 Semidefinite Programming Heuristics for Surface Reconstruction Ambiguities
Ady Ecker, Allan Douglas Jepson, Kiriakos N. Kutulakos
ECCV (1)2
2007 Higher-order Autoregressive Models for Dynamic Textures
abstract
Dynamic textured sequences are characterized by the interactions between many particles or objects in the scene. Based on earlier work the images of the sequence are interpreted as the output of a linear autoregressive process driven by white Gaussian noise. We extend earlier work by increasing the amount temporal information included when learning the motion in the scene, allowing the models to capture complex motion patterns which extend over multiple frames, thereby increasing the perceptual accuracy of the synthesized results. To overcome problems of dynamic model stability, we apply Burg’s Maximum Entropy Spectral Analysis technique for parameter estimation, which is found to be reliably stable on smaller samples of training data, even with higher-order dynamics. 1
Midori Hyndman, Allan Douglas Jepson, David J. Fleet
BMVC2
2007 Shape from Planar Curves: A Linear Escape from Flatland
abstract
We revisit the problem of recovering 3D shape from the projection of planar curves on a surface. This problem is strongly motivated by perception studies. Applications include single-view modeling and fully uncalibrated structured light. When the curves intersect, the problem leads to a linear system for which a direct least-squares method is sensitive to noise. We derive a more stable solution and show examples where the same method produces plausible surfaces from the projection of parallel (non-intersecting) planar cross sections.
Ady Ecker, Kiriakos N. Kutulakos, Allan Douglas Jepson
CVPR3
2007 Flexible Spatial Configuration of Local Image Features
abstract
Local image features have been designed to be informative and repeatable under rigid transformations and illumination deformations. Even though current state-of-the-art local image features present a high degree of repeatability, their local appearance alone usually does not bring enough discriminative power to support a reliable matching, resulting in a relatively high number of mismatches in the correspondence set formed during the data association procedure. As a result, geometric filters, commonly based on global spatial configuration, have been used to reduce this number of mismatches. However, this approach presents a trade off between the effectiveness to reject mismatches and the robustness to non-rigid deformations. In this paper, we propose two geometric filters, based on semilocal spatial configuration of local features, that are designed to be robust to non-rigid deformations and to rigid transformations, without compromising its efficacy to reject mismatches. We compare our methods to the Hough transform, which is an efficient and effective mismatch rejection step based on global spatial configuration of features. In these comparisons, our methods are shown to be more effective in the task of rejecting mismatches for rigid transformations and non-rigid deformations at comparable time complexity figures. Finally, we demonstrate how to integrate these methods in a probabilistic recognition system such that the final verification step uses not only the similarity between features, but also their semi-local configuration.
Gustavo Carneiro 0001, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 The Distinctiveness, Detectability, and Robustness of Local Image Features
abstract
We introduce a new method that characterizes typical local image features (e.g., SIFT, phase feature) in terms of their distinctiveness, detectability, and robustness to image deformations. This is useful for the task of classifying local image features in terms of those three properties. The importance of this classification process for a recognition system using local features is as follows: a) reduce the recognition time due to a smaller number of features present in the test image and in the database of model features; b) improve the recognition accuracy since only the most useful features for the recognition task are kept in the model database; and c) increase the scalability of the recognition system given the smaller number of features per model. A discriminant classifier is trained to select well behaved feature points. A regression network is then trained to provide quantitative models of the detection distributions for each selected feature point. It is important to note that both the classifier and the regression network use image data alone as their input. Experimental results show that the use of these trained networks not only improves the performance of our recognition system, but it also significantly reduces the computation time for the recognition process.
Gustavo Carneiro 0001, Allan Douglas Jepson
CVPR (2)2
2005 Quantitative Evaluation of a Novel Image Segmentation Algorithm
abstract
We present a quantitative evaluation of SE-MinCut, a novel segmentation algorithm based on spectral embedding and minimum cut. We use human segmentations from the Berkeley segmentation database as ground truth and propose suitable measures to evaluate segmentation quality. With these measures we generate precision/recall curves for SE-MinCut and three of the leading segmentation algorithms: mean-shift, normalized Cuts, and the local variation algorithm. These curves characterize the performance of each algorithm over a range of input parameters. We compare the precision/recall curves for the four algorithms and show segmented images that support the conclusions obtained from the quantitative evaluation.
Francisco J. Estrada, Allan Douglas Jepson
CVPR (2)2
2004 Spectral Embedding and Min Cut for Image Segmentation
abstract
Recently it has been shown that min-cut algorithms can provide perceptually salient image segments when they are given appropriate proposals for source and sink regions. Here we explore the use of random walks and associated spectral embedding techniques for the automatic generation of suitable proposal regions. To do this, we first derive a mathematical connection between spectral embedding and anisotropic image smoothing kernels. We then use properties of the spectral embedding and the associated smoothing kernels to select multiple pairs of source and sink regions for min-cut. This typically provides an over-segmentation, and therefore region merging is used to form the final image segmentation. We demonstrate this process on several sample images. 1
Francisco J. Estrada, Allan Douglas Jepson, S. Chakra Chennubhotla
BMVC2
2004 Flexible Spatial Models for Grouping Local Image Features
Gustavo Carneiro 0001, Allan Douglas Jepson
CVPR (2)2
2004 Variational Mixture Smoothing for Non-Linear Dynamical Systems
Cristian Sminchisescu, Allan Douglas Jepson
CVPR (2)2
2004 Generative modeling for continuous non-linearly embedded visual inference
abstract
Many difficult visual perception problems, like 3D human motion estimation, can be formulated in terms of inference using complex generative models, defined over high-dimensional state spaces. Despite progress, optimizing such models is difficult because prior knowledge cannot be flexibly integrated in order to reshape an initially designed representation space. Nonlinearities, inherent sparsity of high-dimensional training sets, and lack of global continuity makes dimensionality reduction challenging and low-dimensional search inefficient. To address these problems, we present a learning and inference algorithm that restricts visual tracking to automatically extracted, non-linearly embedded, low-dimensional spaces. This formulation produces a layered generative model with reduced state representation, that can be estimated using efficient continuous optimization methods. Our prior flattening method allows a simple analytic treatment of low-dimensional intrinsic curvature constraints, and allows consistent interpolation operations. We analyze reduced manifolds for human interaction activities, and demonstrate that the algorithm learns continuous generative models that are useful for tracking and for the reconstruction of 3D human motion in monocular video.
Cristian Sminchisescu, Allan Douglas Jepson
ICML2
2004 Hierarchical Eigensolver for Transition Matrices in Spectral Methods
abstract
We show how to build hierarchical, reduced-rank representation for large stochastic matrices and use this representation to design an efficient al- gorithm for computing the largest eigenvalues, and the corresponding eigenvectors. In particular, the eigen problem is first solved at the coars- est level of the representation. The approximate eigen solution is then interpolated over successive levels of the hierarchy. A small number of power iterations are employed at each stage to correct the eigen solution. The typical speedups obtained by a Matlab implementation of our fast eigensolver over a standard sparse matrix eigensolver [13] are at least a factor of ten for large image sizes. The hierarchical representation has proven to be effective in a min-cut based segmentation algorithm that we proposed recently [8]. 1 Spectral Methods Graph-theoretic spectral methods have gained popularity in a variety of application do- mains: segmenting images [22]; embedding in low-dimensional spaces [4, 5, 8]; and clus- tering parallel scientific computation tasks [19]. Spectral methods enable the study of prop- erties global to a dataset, using only local (pairwise) similarity or affinity measurements be- tween the data points. The global properties that emerge are best understood in terms of a random walk formulation on the graph. For example, the graph can be partitioned into clus- ters by analyzing the perturbations to the stationary distribution of a Markovian relaxation process defined in terms of the affinity weights [17, 18, 24, 7]. The Markovian relaxation process need never be explicitly carried out; instead, it can be analytically expressed using the leading order eigenvectors, and eigenvalues, of the Markov transition matrix. In this paper we consider the practical application of spectral methods to large datasets. In particular, the eigen decomposition can be very expensive, on the order of O(n3), where n is the number of nodes in the graph. While it is possible to compute analytically the first eigenvector (see x3 below), the remaining subspace of vectors (necessary for say clustering) has to be explicitly computed. A typical approach to dealing with this difficulty is to first sparsify the links in the graph [22] and then apply an efficient eigensolver [13, 23, 3]. In comparison, we propose in this paper a specialized eigensolver suitable for large stochas- tic matrices with known stationary distributions. In particular, we exploit the spectral prop- erties of the Markov transition matrix to generate hierarchical, successively lower-ranked approximations to the full transition matrix. The eigen problem is solved directly at the coarsest level of representation. The approximate eigen solution is then interpolated over successive levels of the hierarchy, using a small number of power iterations to correct the solution at each stage. 2 Previous Work One approach to speeding up the eigen decomposition is to use the fact that the columns of the affinity matrix are typically correlated. The idea then is to pick a small number of representative columns to perform eigen decomposition via SVD. For example, in the Nystrom approximation procedure, originally proposed for integral eigenvalue problems, the idea is to randomly pick a small set of m columns; generate the corresponding affinity matrix; solve the eigenproblem and finally extend the solution to the complete graph [9, 10]. The Nystrom method has also been recently applied in the kernel learning methods for fast Gaussian process classification and regression [25]. Other sampling-based approaches include the work reported in [1, 2, 11]. Our starting point is the transition matrix generated from affinity weights and we show how building a representational hierarchy follows naturally from considering the stochas- tic matrix. A closely related work is the paper by Lin on reduced rank approximations of transition matrices [14]. We differ in how we approximate the transition matrices, in par- ticular our objective function is computationally less expensive to solve. In particular, one of our goals in reducing transition matrices is to develop a fast, specialized eigen solver for spectral clustering. Fast eigensolving is also the goal in ACE [12], where successive levels in the hierarchy can potentially have negative affinities. A graph coarsening process for clustering was also pursued in [21, 3]. 3 Markov Chain Terminology We first provide a brief overview of the Markov chain terminology here (for more details see [17, 15, 6]). We consider an undirected graph G = (V; E) with vertices vi, for i = f1; : : : ; ng, and edges ei;j with non-negative weights ai;j. Here the weight ai;j represents the affinity between vertices vi and vj. The affinities are represented by a non-negative, symmetric n (cid:2) n matrix A having weights ai;j as elements. The degree of a node j is j=1 aj;i, where we define D = diag(d1; : : : ; dn). A Markov chain is defined using these affinities by setting a transition probability matrix M = AD(cid:0)1, where the columns of M each sum to 1. The transition probability matrix defines the random walk of a particle on the graph G. The random walk need never be explicitly carried out; instead, it can be analytically ex- pressed using the leading order eigenvectors, and eigenvalues, of the Markov transition matrix. Because the stochastic matrices need not be symmetric in general, a direct eigen decomposition step is not preferred for reasons of instability. This problem is easily circum- vented by considering a normalized affinity matrix: L = D(cid:0)1=2AD(cid:0)1=2, which is related to the stochastic matrix by a similarity transformation: L = D(cid:0)1=2M D1=2. Because L is symmetric, it can be diagonalized: L = U (cid:3)U T , where U = [~u1; ~u2; (cid:1) (cid:1) (cid:1) ; ~un] is an orthogonal set of eigenvectors and (cid:3) is a diagonal matrix of eigenvalues [(cid:21)1; (cid:21)2; (cid:1) (cid:1) (cid:1) ; (cid:21)n] sorted in decreasing order. The eigenvectors have unit length k~ukk = 1 and from the form of A and D it can be shown that the eigenvalues (cid:21)i 2 ((cid:0)1; 1], with at least one eigenvalue equal to one. Without loss of generality, we take (cid:21)1 = 1. Because L and M are similar we can perform an eigen decomposition of the Markov transition matrix as: M = D1=2LD(cid:0)1=2 = D1=2U (cid:3) U T D(cid:0)1=2. Thus an eigenvector ~u of L corresponds to an eigenvector D1=2~u of M with the same eigenvalue (cid:21). The Markovian relaxation process after (cid:12) iterations, namely M (cid:12), can be represented as: M (cid:12) = D1=2U (cid:3)(cid:12)U T D(cid:0)1=2. Therefore, a particle undertaking a random walk with an initial distribution ~p 0 acquires after (cid:12) steps a distribution ~p (cid:12) given by: ~p (cid:12) = M (cid:12)~p 0. Assuming the graph is connected, as (cid:12) ! 1, the Markov chain approaches a unique i=1 di, and thus, M 1 = ~(cid:25)1T , where 1 is a n-dim column vector of all ones. Observe that ~(cid:25) is an eigenvector of M as it is easy to show that M~(cid:25) = ~(cid:25) and the corresponding eigenvalue is 1. Next, we show how to generate hierarchical, successively low-ranked approximations for the transition matrix M. stationary distribution given by ~(cid:25) = diag(D)=Pn defined to be: dj = Pn
S. Chakra Chennubhotla, Allan Douglas Jepson
NIPS2
2003 Multi-scale Phase-based Local Features
abstract
Local feature methods suitable for image feature based object recognition and for the estimation of motion and structure are composed of two steps, namely the 'where' and 'what' steps. The 'where' step (e.g., interest point detector) must select image points that are robustly localizable under common image deformations and whose neighborhoods are relatively informative. The 'what' step (e.g., local feature extractor) then provides a representation of the image neighborhood that is semi-invariant to image deformations, but distinctive enough to provide model identification. We present a quantitative evaluation of both the 'where' and the 'what' steps for three recent local feature methods: a) phase-based local features (Carneiro and Jepson, 2002), b) differential invariants (Schmid and Mohr, 1997), and c) the scale invariant feature transform (SIFT) (Lowe, 1999). Moreover, in order to make the phase-based approach more comparable to the other two approaches, we also introduce a new form of multi-scale interest point detector to be used for its 'where' step. The results show that the phase-based local features lead to better performance than the other two approaches when dealing with common illumination changes, 2D rotation, and sub-pixel translation. On the other hand, the phase-based local features are somewhat more sensitive to scale and large shear changes than the other two methods. Finally, we demonstrate the viability of the phase-based local feature in a simple object recognition system.
Gustavo Carneiro 0001, Allan Douglas Jepson
CVPR (1)2
2003 Video Input Driven Animation (VIDA)
abstract
There are many challenges associated with the integration of synthetic and real imagery. One particularly difficult problem is the automatic extraction of salient parameters of natural phenomena in real video footage for subsequent application to synthetic objects. We can ensure that the hair and clothing of a synthetic actor placed in a meadow of swaying grass will move consistently with the wind that moved that grass. The video footage can be seen as a controller for the motion of synthetic features, a concept we call video input driven animation (VIDA). We propose a schema that analyzes an input video sequence, extracts parameters from the motion of objects in the video, and uses this information to drive the motion of synthetic objects. To validate the principles of VIDA, we approximate the inverse problem to harmonic oscillation, which we use to extract parameters of wind and of regular water waves. We observe the effect of wind on a tree in a video, estimate wind speed parameters from its motion, and then use this to make synthetic objects move. We also extract water elevation parameters from the observed motion of boats and apply the resulting water waves to synthetic boats.
Allan Douglas Jepson, Eugene Fiume
ICCV2
2003 Robust Online Appearance Models for Visual Tracking
abstract
We propose a framework for learning robust, adaptive, appearance models to be used for motion-based tracking of natural objects. The model adapts to slowly changing appearance, and it maintains a natural measure of the stability of the observed image structure during tracking. By identifying stable properties of appearance, we can weight them more heavily for motion estimation, while less stable properties can be proportionately downweighted. The appearance model involves a mixture of stable image structure, learned over long time courses, along with two-frame motion information and an outlier process. An online EM-algorithm is used to adapt the appearance model parameters over time. An implementation of this approach is developed for an appearance model based on the filter responses from a steerable pyramid. This model is used in a motion-based tracking algorithm to provide robustness in the face of image outliers, such as those caused by occlusions, while adapting to natural changes in appearance such as those due to facial expressions or variations in 3D pose.
Allan Douglas Jepson, David J. Fleet, Thomas F. El-Maraghi
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Phase-Based Local Features
Gustavo Carneiro 0001, Allan Douglas Jepson
ECCV (1)2
2002 A Layered Motion Representation with Occlusion and Compact Spatial Support
Allan Douglas Jepson, David J. Fleet, Michael J. Black
ECCV (1)1
2002 Half-Lives of EigenFlows for Spectral Clustering
abstract
Using a Markov chain perspective of spectral clustering we present an algorithm to automatically find the number of stable clusters in a dataset. The Markov chain’s behaviour is characterized by the spectral properties of the matrix of transition probabilities, from which we derive eigenflows along with their halflives. An eigenflow describes the flow of probabil- ity mass due to the Markov chain, and it is characterized by its eigen- value, or equivalently, by the halflife of its decay as the Markov chain is iterated. A ideal stable cluster is one with zero eigenflow and infi- nite half-life. The key insight in this paper is that bottlenecks between weakly coupled clusters can be identified by computing the sensitivity of the eigenflow’s halflife to variations in the edge weights. We propose a novel EIGENCUTS algorithm to perform clustering that removes these identified bottlenecks in an iterative fashion.
S. Chakra Chennubhotla, Allan Douglas Jepson
NIPS2
2001 Robust Online Appearance Models for Visual Tracking
abstract
Abstract—We propose a framework for learning robust, adaptive, appearance models to be used for motion-based tracking of natural objects. The model adapts to slowly changing appearance, and it maintains a natural measure of the stability of the observed image structure during tracking. By identifying stable properties of appearance, we can weight them more heavily for motion estimation, while less stable properties can be proportionately downweighted. The appearance model involves a mixture of stable image structure, learned over long time courses, along with two-frame motion information and an outlier process. An online EM-algorithm is used to adapt the appearance model parameters over time. An implementation of this approach is developed for an appearance model based on the filter responses from a steerable pyramid. This model is used in a motion-based tracking algorithm to provide robustness in the face of image outliers, such as those caused by occlusions, while adapting to natural changes in appearance such as those due to facial expressions or variations in 3D pose.
Allan Douglas Jepson, David J. Fleet, Thomas F. El-Maraghi
CVPR (1)1
2001 Sparse PCA: Extracting Multi-scale Structure from Data
S. Chakra Chennubhotla, Allan Douglas Jepson
ICCV2
2001 SAVI: an actively controlled teleconferencing system
Rainer Herpers, Konstantinos G. Derpanis, W. James MacLean, Gilbert Verghese, Michael R. M. Jenkin, Evangelos E. Milios, Allan Douglas Jepson, John K. Tsotsos
Image Vis. Comput.7
2000 Design and Use of Linear Models for Image Motion Analysis
David J. Fleet, Michael J. Black, Yaser Yacoob, Allan Douglas Jepson
Int. J. Comput. Vis.4
1999 Qualitative Probabilities for Image Interpretation
abstract
Two basic problems in image interpretation are: a) determining which interpretations are the most plausible amongst many possibilities; and b) controlling the search for plausible interpretations. We address these issues using a Bayesian approach, with the plausibility ordering and search pruning based on the posterior probabilities of interpretations. However, due to the need for detailed quantitative prior probabilities and the need to evaluate complex integrals over various conditional distributions, a full Bayesian approach is currently impractical except in tightly constrained domains. To circumvent these difficulties we introduce the notion of qualitative probabilistic analysis. In particular, given spatial and contrast resolution parameters, we consider only the asymptotic order of the posterior probability for any interpretation as these resolutions are made finer. We introduce this approach for a simple card-world domain, and present computational results for blocks-world images.
Allan Douglas Jepson, Richard Mann
ICCV1
1998 Motion Feature Detection Using Steerable Flow Fields
abstract
The estimation and detection of occlusion boundaries and moving bars are important and challenging problems in image sequence analysis. Here, we model such motion features as linear combinations of steerable basis flow fields. These models constrain the interpretation of image motion, and are used in the same way as translational or affine motion models. We estimate the subspace coefficients of the motion feature models directly from spatiotemporal image derivatives using a robust regression method. From the subspace coefficients we detect the presence of a motion feature and solve for the orientation of the feature and the relative velocities of the surfaces. Our method does not require the prior computation of optical flow and recovers accurate estimates of orientation and velocity.
David J. Fleet, Michael J. Black, Allan Douglas Jepson
CVPR3
1998 Towards the Computational Perception of Action
abstract
Understanding observations of interacting objects requires one to reason about qualitative scene dynamics. For example, on observing a hand lifting a can, we may infer that an 'active' hand is applying an upwards force (by grasping) to lift a 'passive' can. Previously we presented a system that infers qualitative scene dynamics from the instantaneous motion of objects. However; since that analysis only considered single frames in isolation, there were often multiple interpretations for each frame. In this work we show how the dynamic information inferred at each frame can be integrated over time to reduce ambiguity. Our approach to integrating information is to extend our representation to describe objects by a set of properties or capabilities that are assumed to persist over time. Given this extended representation we find interpretations that require the smallest set(s) of properties over the whole image sequence.
Richard Mann, Allan Douglas Jepson
CVPR2
1998 A Probabilistic Framework for Matching Temporal Trajectories: CONDENSATION-Based Recognition of Gestures and Expressions
Michael J. Black, Allan Douglas Jepson
ECCV (1)2
1998 Recognizing Temporal Trajectories Using the Condensation Algorithm
Michael J. Black, Allan Douglas Jepson
FG2
1998 EigenTracking: Robust Matching and Tracking of Articulated Objects Using a View-Based Representation
Michael J. Black, Allan Douglas Jepson
Int. J. Comput. Vis.2
1998 PLAYBOT A visually-guided robot for physically disabled children
John K. Tsotsos, Gilbert Verghese, Sven J. Dickinson, Michael R. M. Jenkin, Allan Douglas Jepson, Evangelos E. Milios, Fernando Nuflo, Suzanne Stevenson, Michael J. Black, Dimitris N. Metaxas
Image Vis. Comput.5
1997 Learning Parameterized Models of Image Motion
abstract
A framework for learning parameterized models of optical flow from image sequences is presented. A class of motions is represented by a set of orthogonal basis flow fields that are computed from a training set using principal component analysis. Many complex image motions can be represented by a linear combination of a small number of these basis flows. The learned motion models may be used for optical flow estimation and for model-based recognition. For optical flow estimation we describe a robust, multi-resolution scheme for directly computing the parameters of the learned flow models from image derivatives. As examples we consider learning motion discontinuities, non-rigid motion of human mouths, and articulated human motion.
Michael J. Black, Yaser Yacoob, Allan Douglas Jepson, David J. Fleet
CVPR3
1997 The Computational Perception of Scene Dynamics
Richard Mann, Allan Douglas Jepson, Jeffrey Mark Siskind
Comput. Vis. Image Underst.2
1996 Skin and Bones: Multi-layer, Locally Affine, Optical Flow and Regularization with Transparency
abstract
This paper describes a new method for estimating optical flow that strikes a balance between the flexibility of local dense computations and the robustness and accuracy of global parameterized flow models. An affine model of image motion is used within local image patches while a spatial smoothness constraint on the affine flow parameters of neighboring patches enforces continuity of the motion. We refer to this as a "Skin and Bones" model in which the affine patches can be thought of as rigid "bones" connected by a flexible "skin". Since local image patches may contain multiple motions we use a layered representation for the affine bones. To regularize this layered motion representation we develop a new framework for regularization with transparency.
Shanon X. Ju, Michael J. Black, Allan Douglas Jepson
CVPR3
1996 EigenTracking: Robust Matching and Tracking of Articulated Objects Using a View-Based Representation
Michael J. Black, Allan Douglas Jepson
ECCV (1)2
1996 Computational Perception of Scene Dynamics
Richard Mann, Allan Douglas Jepson, Jeffrey Mark Siskind
ECCV (2)2
1996 Estimating Optical Flow in Segmented Images Using Variable-Order Parametric Models With Local Deformations
abstract
This paper presents a new model for estimating optical flow based on the motion of planar regions plus local deformations. The approach exploits brightness information to organize and constrain the interpretation of the motion by using segmented regions of piecewise smooth brightness to hypothesize planar regions in the scene. Parametric flow models are estimated in these regions in a two step process which first computes a coarse fit and then estimates the appropriate parametrization of the motion of the region. The initial fit is refined using a generalization of the standard area-based regression approaches. Since the assumption of planarity is likely to be violated, we allow local deformations from the planar assumption in the same spirit as physically-based approaches which model shape using coarse parametric models plus local deformations. This parametric plus deformation model exploits the strong constraints of parametric approaches while retaining the adaptive nature of regularization approaches. Experimental results on a variety of images model produces accurate flow estimates while the incorporation of brightness segmentation boundaries.
Michael J. Black, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Detecting Floor Anomalies
abstract
When a robot moves about a 2D world such as a planar surface, it is important that obstacles to the robot's motions be detected. This classical problem of has proven to be difficult. Many researchers have formulated this problem as being the process of determining where a robot cannot move due to the presence of obstacles. An alternative approach presented here is to determine where an robot can go by identifying floor regions for which the planar floor assumption can be verified. A stereo vision system is developed for Floor Anomaly Detection (FAD), and its relationship to existing stereo obstacle detection algorithms is described.
Michael R. M. Jenkin, Allan Douglas Jepson
BMVC2
1994 Recovery of Egomotion and Segmentation of Independent Object Motion Using the EM Algorithm
abstract
This paper examines the use of the EM algorithm to perform motion segmentation on image sequences that contain independent object motion. The input data are linear constraints on 3-D translational motion and bilinear constraints on 3-D translation and rotation, derived from computed optical flow using subspace methods. The problems of outlier detection, deciding how many processes, and the initial guesses for the EM algorithm are considered. Results obtained from an image sequence are presented. 1 Introduction In order for an observer to navigate in its environment, it is important that the observer can detect other independently moving objects and avoid collisions. The motion of the observer complicates this task. For the purpose of this paper we divide image motion into two categories: egomotion and motion due to independent moving objects. Egomotion is defined as the image motion induced by an observer moving through a static environment. Motion due to independently moving objects ...
W. James MacLean, Allan Douglas Jepson, Richard C. Frecker
BMVC2
1994 A new closed-form solution for absolute orientation
abstract
Presents a closed-form solution for the determination of 3D displacement and rotation parameters, given a set of 3D point correspondences. The method applies to the general case of arbitrary displacement and rotation with respect to an arbitrary stationary scene. The approach is based on the linear subspace method, which allows separate linear equations to be extracted for the displacement and rotation. This leads to a simple algorithm, suitable for real-time implementations, for the determination of both the 3D transformation and error estimates for the estimated 3D displacement. Preliminary experimental results with range data are presented.>
Zhengyan Wang, Allan Douglas Jepson
CVPR2
1994 ARK: autonomous mobile robot for an industrial environment
abstract
This paper describes research on the ARK (Autonomous Mobile Robot in a Known Environment) project. The technical objective of the project is to build a robot that can navigate and carry out survey/inspection tasks in a complex but known industrial environment. Rather than altering the robots environment by adding easily identifiable beacons the robot relies on naturally occurring objects to use as visual landmarks for navigation. The robot is equipped with various sensors that are used to detect unmapped obstacles, landmarks and objects. This paper describes the robot's industrial environment, it's control architecture, and some results in processing the robot's range and vision sensor data for navigation.>
Michael R. M. Jenkin, N. Bains, J. Bruce, T. Campbell, Brian Down, Piotr Jasiobedzki, Allan Douglas Jepson, B. Majarais, Evangelos E. Milios, S. B. Nickerson, James R. R. Service, Demetri Terzopoulos, John K. Tsotsos, David Wilkes
IROS7
1993 Mixture models for optical flow computation
abstract
The computation of optical flow relies on merging information available over an image patch to form an estimate of 2-D image velocity at a point. This merging process raises many issues. These include the treatment of outliers in component velocity measurements and the modeling of multiple motions within a patch which arise from occlusion boundaries or transparency. A new approach for dealing with these issues is presented. It is based on the use of a probabilistic mixture model to explicitly represent multiple motions within a patch. A simple extension of the EM-algorithm is used to compute a maximum likelihood estimate for the various motion parameters. Preliminary experiments indicate that this approach is computationally efficient, and that it can provide robust estimates of the optical flow values in the presence of outliers and multiple motions.>
Allan Douglas Jepson, Michael J. Black
CVPR1
1993 Stability of Phase Information
abstract
This paper concerns the robustness of local phase information for measuring image velocity and binocular disparity. It addresses the dependence of phase behavior on the initial filters as well as the image variations that exist between different views of a 3D scene. We are particularly interested in the stability of phase with respect to geometric deformations, and its linearity as a function of spatial position. These properties are important to the use of phase information, and are shown to depend on the form of the filters as well as their frequency bandwidths. Phase instabilities are also discussed using the model of phase singularities described by Jepson and Fleet. In addition to phase-based methods, these results are directly relevant to differential optical flow methods and zero-crossing tracking.>
David J. Fleet, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 From Features to Perceptual Categories
Whitman Richards, Jacob Feldman, Allan Douglas Jepson
BMVC3
1992 Subspace methods for recovering rigid motion I: Algorithm and implementation
David J. Heeger, Allan Douglas Jepson
Int. J. Comput. Vis.2
1992 A lattice framework for integrating vision modules
abstract
How information processing by individual modules such as stereopsis, motion, texture, and color modules may be integrated, or assimilated, is addressed. A framework for assimilation based on a partial ordering of constraints implicit in all active modules is proposed. Such constraints, for example the rigidity constraint for motion, although often robust are also fallible, and hence are more properly regarded as premises. Such premises are used to construct a preference ordering for (classes of) interpretations of the image. Interpretations associated with maximal states in the ordering are taken as the assimilated interpretation of the modules. This approach stresses the need to use world knowledge to reason about the plausibility and consistency of interpretations of the image data.>
Allan Douglas Jepson, Whitman Richards
IEEE Trans. Syst. Man Cybern.1
1991 Phase-based disparity measurement
David J. Fleet, Allan Douglas Jepson, Michael R. M. Jenkin
CVGIP Image Underst.2
1991 Techniques for disparity measurement
Michael R. M. Jenkin, Allan Douglas Jepson, John K. Tsotsos
CVGIP Image Underst.2
1991 Phase singularities in scale-space
Allan Douglas Jepson, David J. Fleet
Image Vis. Comput.1
1990 Scale-Space Singularities
Allan Douglas Jepson, David J. Fleet
ECCV1
1990 Simple method for computing 3D motion and depth
abstract
The nonlinear equation that relates the optical-flow field to 3-D motion and depth can be split by an exact algebraic manipulation to form three sets of equations. The first set relates the image velocities to the translational component of the 3-D motion alone. Thus, the depth and the rotational velocity need not be known or estimated prior to solving for the translational velocity. Once the translation has been recovered, the second set of equations can be used to solve for the rotational velocity. Finally, depth can be estimated with the third set of equations, given the recovered transaction and rotation. The algorithm applies to the general case of arbitrary motion with respect to an arbitrary scene. The authors show that the performance of the algorithm compares quite favorably with other proposed approaches.>
David J. Heeger, Allan Douglas Jepson
ICCV2
1990 The feasibility of motion and structure from noisy time-varying image velocity information
John L. Barron, Allan Douglas Jepson, John K. Tsotsos
Int. J. Comput. Vis.2
1990 Computation of component image velocity from local phase information
David J. Fleet, Allan Douglas Jepson
Int. J. Comput. Vis.2
1990 Visual Perception of Three-Dimensional Motion
abstract
As an observer moves and explores the environment, the visual stimulation in his eye is constantly changing. Somehow he is able to perceive the spatial layout of the scene, and to discern his movement through space. Computational vision researchers have been trying to solve this problem for a number of years with only limited success. It is a difficult problem to solve because the relationship between the optical-flow field, the 3D motion parameters, and depth is nonlinear. We have come to understand that this nonlinear equation describing the optical-flow field can be split by an exact algebraic manipulation to yield an equation that relates the image velocities to the translational component of the 3D motion alone. Thus, the depth and the rotational velocity need not be known or estimated prior to solving for the translational velocity. The algorithm applies to the general case of arbitrary motion with respect to an arbitrary scene. It is simple to compute and it is plausible biologically.
David J. Heeger, Allan Douglas Jepson
Neural Comput.2
1989 Computation of normal velocity from local phase information
abstract
A technique for the estimation of 2-D normal velocity is presented. The image sequence is first represented by a family of velocity-tuned linear filters. Normal velocity, in the individual filter outputs, is expressed as the local first-order behavior of surfaces of constant phase. Justification for this is discussed, and it is shown to provide an effective basis for the local computation of normal velocity. The resultant approach is local in space-time. It permits multiple velocity estimates within a single neighborhood, and it yields accurate velocity estimates that are robust with respect to noise and perspective deformation.>
David J. Fleet, Allan Douglas Jepson
CVPR2
1989 The fast computation of disparity from phase differences
abstract
Previous work has demonstrated that the task of recovering local disparity measurements can be reduced to the task of measuring the local phase between bandpass signals extracted from the left and right cameras. In computing this local phase difference, earlier algorithms expressed the computational task as a nonlinear differential equation to be solved at each image point. Although this approach has great appeal as a model for biological disparity measurement, the solving of a differential equation at a large number of image points and disparities makes the algorithm unsuitable for serial digital computer applications. Here, the authors demonstrate how the approach of recovering disparity from the measurement of local phase differences can be accomplished without the computational expense exhibited by previous algorithms. This disparity measurement technique is embedded within a simple coarse-to-fine stereopsis similar to the algorithm proposed by H.K. Nishihara (1984) and the resulting algorithm is applied to a number of stereo pairs.>
Allan Douglas Jepson, Michael R. M. Jenkin
CVPR1
1989 Hierarchical Construction of Orientation and Velocity Selective Filters
abstract
As a step towards the early measurement of visual primitives, the authors outline design criteria for the extraction of orientation and velocity information, and present a variety of tools useful in the construction of simple linear filters. A hierarchical parallel-processing scheme is used in which nodes compute a weighted sum of inputs from within a small spatio-temporal neighborhood. The resulting scheme is easily analyzed and provides mechanisms sensitive to narrow ranges of both image velocity and orientation. The hierarchical approach in combination with separability in the first levels yields an efficient implementation.>
David J. Fleet, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.2
1989 Response profiles of trajectory detectors
abstract
It has previously been demonstrated that detectors can be constructed that are sensitive to the 3D trajectory of a structure. The general approach is based upon a biological model for looming detectors proposed by K. Beverley and D. Regan (J. Physiol., vol.193, p.17-29, 1973). The computational model is based upon recent results in motion analysis and static stereopsis concerning the measurement of trajectory without prior form recognition. Some of the characteristics of these detectors are examined, and a number of experiments that show the responses of particular detectors to structure with different trajectories, disparities, and velocities are described. Some of the similarities and differences between the detectors presented and models for looming detectors present in biological vision systems are indicated.>
Michael R. M. Jenkin, Allan Douglas Jepson
IEEE Trans. Syst. Man Cybern.2
1988 The Feasibility Of Motion And Structure Computations
abstract
We address the problem of interpreting image velocity fields generated by a moving monocular observer viewing a stationary environment under perspective projection to obtain 3-D information about the relative motion of the observer (egomotion) and the relative depth of environmental surface points (environmental layout). The algorithm presented in this paper involves computing motion and structure from a spatio-temporal distribution of image velocities that are hypothesized to belong to the same 3-D planar surface. However, the main result of this paper is not just another motion and structure algorithm that exhibits some novel features but rather an extensive error analysis of the algorithm’s preformance for various types of noise in the image velocities. Waxman and Ullman [83] have devised an algorithm for computing motion and structure using image velocity and its 1st and 2d order spatial derivatives at one image point. We generalize this result to include derivative information in time as well. Further, we show the approximate equivalence of reconstruction algorithms that use only image velocities and those that use one image velocity and its 1st and/or 2”d spatio-temporal derivatives at one image point. The main question addressed in this paper is: “How accurate do the input image velocities have to be?’ or equivalently, “How accurate does the input image velocity and its I~ and 2& order derivatives have to be?“. The answer to this question involves worst case error analysis. We end the paper by drawing some conclusions about the feasibility of motion and structure calculations in general. I.1 Introduction In this paper, we present a algorithm for computing the motion and strncture parameters that describe egomotion and environmental layout from image velocity fields generated by a moving monocular observer viewing a stationary environment. Egomotion is defined as the motion of the observer relative to his environment and can be described by 6 parameters; 3 dvth-scaled translational parameters, Z and 3 rotation parameters, o. Environmental layout refers to the 3-D shape and location of objects in the environment. For monocular image sequences, en$ronmental layout is described by the normalized surface gradient, a, at each image point. To determine these motion and structure parameters we derive nonlinear equations relating image velocity at some image int ?(?*,t ‘) to the underlying motion and structure parameters at (P,c). The computaP tion of egomotion and environmental layout from image velocity is sometimes called the reconstruction problem; we reconstruct the observer’s motion, and the layout of his environment, from (timevarying) image velocity. A lot of research has been devoted to
John L. Barron, Allan Douglas Jepson, John K. Tsotsos
ICCV2
1987 The Sensitivity of Motion and Structure Computations
John L. Barron, Allan Douglas Jepson, John K. Tsotsos
AAAI2
1987 Determination of Egomotion and Environmental Layout from Noisy Time-Varying Image Velocity in Binocular Image Sequences
John L. Barron, Allan Douglas Jepson, John K. Tsotsos
IJCAI2
1987 The Use of Color in Highlight Identification
Ron Gershon, Allan Douglas Jepson, John K. Tsotsos
IJCAI2
1987 From [R, G, B] to Surface Reflectance: Computing Color Constant Descriptors in Images
Ron Gershon, Allan Douglas Jepson, John K. Tsotsos
IJCAI2