P. Anandan 0001

dblp:34/659-1 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
0since 2021 · last 2006
0000-0003-3457-7281ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
3D vision · 86% Video understanding and tracking · 10% Representation and self-supervised learning · 4%
Computer graphics and multimedia
10 papers
Image and video processing · 59% Multimedia analysis and retrieval · 21% Geometric modeling and processing · 17%
Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 77% Data models and query languages · 23%

Topics — the 30 heaviest of 58, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
structure from motion
0.162002
Direct Recovery of Planar-Parallax from Multiple Frames · IEEE Trans. Pattern Anal. Mach. Intell. 2002
Factorization with Uncertainty · ECCV (1) 2000
Layer Extraction from Multiple Images Containing Reflections and Transparency · CVPR 2000
Computer vision › 3D vision › 3d scene modeling › scene representation
layered scene representation
0.132001
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2001
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · ICCV 1999
A Layered Approach to Stereo Reconstruction · CVPR 1998
Computer vision › 3D vision
motion estimation
0.132005
Motion Recovery by Integrating over the Joint Image Manifold · Int. J. Comput. Vis. 2005
A framework for the robust estimation of optical flow · ICCV 1993
A model for the detection of motion over time · ICCV 1990
Computer vision › 3D vision › motion estimation
camera motion estimation
0.122003
Recovery of Epipolar Geometry as a Manifold Fitting Problem · ICCV 2003
A Unified Approach to Moving Object Detection in 2D and 3D Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Computer vision › 3D vision
3d reconstruction
0.122002
Direct Recovery of Planar-Parallax from Multiple Frames · IEEE Trans. Pattern Anal. Mach. Intell. 2002
Implicit Representation and Scene Reconstruction from Probability Density Functions · CVPR 1999
Computer vision › 3D vision
multi-view geometry
0.142002
From Reference Frames to Reference Planes: Multi-View Parallax Geometry and Applications · ECCV (2) 1998
Parallax Geometry of Pairs of Points for 3D Scene Analysis · ECCV (1) 1996
What Does the Scene Look Like from a Scene Point? · ECCV (2) 2002
Computer vision › 3D vision
3d scene reconstruction
0.021999
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · ICCV 1999
Implicit Representation and Scene Reconstruction from Probability Density Functions · CVPR 1999
Image and video processing
image registration
0.031998
Robust Multi-Sensor Image Alignment · ICCV 1998
Mosaic Based Representations of Video Sequences and Their Applications · ICCV 1995
Adaptive-complexity registration of images · CVPR 1994
Computer vision › 3D vision › multi-view geometry
epipolar geometry
0.012003
Recovery of Epipolar Geometry as a Manifold Fitting Problem · ICCV 2003
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
manifold fitting
0.012003
Recovery of Epipolar Geometry as a Manifold Fitting Problem · ICCV 2003
Computer vision › 3D vision
novel view synthesis
0.012002
What Does the Scene Look Like from a Scene Point? · ECCV (2) 2002
Computer vision › 3D vision › multi-view geometry › two-view geometry
planar parallax
0.012002
Direct Recovery of Planar-Parallax from Multiple Frames · IEEE Trans. Pattern Anal. Mach. Intell. 2002
Computer vision › 3D vision › motion estimation
optical flow
0.042000
A framework for the robust estimation of optical flow · ICCV 1993
Segmenting Visual Actions Based on Spatio-Temporal Motion Patterns · CVPR 2000
Robust dynamic motion estimation over time · CVPR 1991
Computer vision › 3D vision
3d scene modeling
0.012001
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Computer vision › 3D vision › mid-level vision
layer extraction
0.012001
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Computer vision › Video understanding and tracking
motion segmentation
0.012001
An Integrated Bayesian Approach to Layer Extraction from Image Sequences · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Computer vision › Video understanding and tracking
action segmentation
0.012000
Segmenting Visual Actions Based on Spatio-Temporal Motion Patterns · CVPR 2000
Computer vision › 3D vision › multi-view geometry
camera geometry
0.012000
Integrating Local Affine into Global Projective Images in the Joint Image Space · ECCV (1) 2000
Image and video processing › image decomposition › image separation
layer separation
0.012000
Layer Extraction from Multiple Images Containing Reflections and Transparency · CVPR 2000
Computational geometry
projective geometry
0.012000
Integrating Local Affine into Global Projective Images in the Joint Image Space · ECCV (1) 2000
Multimedia analysis and retrieval
video indexing
0.021998
Video indexing based on mosaic representations · Proc. IEEE 1998
Mosaic Based Representations of Video Sequences and Their Applications · ICCV 1995
Computer vision › 3D vision › 3d reconstruction › projective reconstruction
affine reconstruction
0.011999
Implicit Representation and Scene Reconstruction from Probability Density Functions · CVPR 1999
Computer vision › 3D vision › camera calibration
camera model
0.011999
What Can Be Determined from a Full and a Weak Perspective Image? · ICCV 1999
Computer vision › 3D vision › camera calibration › camera model
weak perspective camera
0.011999
What Can Be Determined from a Full and a Weak Perspective Image? · ICCV 1999
Geometric modeling and processing › shape representation
implicit representation
0.011999
Implicit Representation and Scene Reconstruction from Probability Density Functions · CVPR 1999
Geometric modeling and processing
shape representation
0.011999
Implicit Representation and Scene Reconstruction from Probability Density Functions · CVPR 1999
Computer vision › 3D vision
depth estimation
0.011998
A Layered Approach to Stereo Reconstruction · CVPR 1998
Computer vision › Video understanding and tracking › motion detection
moving object detection
0.011998
A Unified Approach to Moving Object Detection in 2D and 3D Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.011998
A Layered Approach to Stereo Reconstruction · CVPR 1998
Computer vision › 3D vision › 3d reconstruction › multi-view stereo
stereo reconstruction
0.011998
A Layered Approach to Stereo Reconstruction · CVPR 1998

Methods — techniques the papers use, named apart from their topics

uncertainty propagation · 0.1epipolar geometry · 0.1expectation-maximization · 0.1bayesian inference · 0.1manifold fitting · 0.0view synthesis · 0.0plane+parallax · 0.0direct brightness-based estimation · 0.0markov chain monte carlo · 0.0RANSAC · 0.0minimum- and maximum-composites · 0.0dominant motion estimation · 0.0constrained least squares · 0.0maximum likelihood estimation · 0.0robust estimation · 0.0normalized correlation · 0.0geometric indexing · 0.0dynamic indexing · 0.0
YearPublicationVenuePosition
2006 Globalization: Challenges to Database Community
Sang Kyun Cha, P. Anandan 0001, Meichun Hsu, C. Mohan 0001, Rajeev Rastogi, Vishal Sikka, Honesty C. Young
VLDB2
2005 Extracting layers and analyzing their specular properties using epipolar-plane-image analysis
Antonio Criminisi, Sing Bing Kang, Rahul Swaminathan, Richard Szeliski, P. Anandan 0001
Comput. Vis. Image Underst.5
2005 Motion Recovery by Integrating over the Joint Image Manifold
Liran Goshen, Ilan Shimshoni, P. Anandan 0001, Daniel Keren
Int. J. Comput. Vis.3
2004 Guest Editorial: Computer Vision Research at Microsoft Corporation
P. Anandan 0001, Andrew Blake 0001
Int. J. Comput. Vis.1
2003 Recovery of Epipolar Geometry as a Manifold Fitting Problem
abstract
The introduction of the joint image manifold allows to treat the problem of recovering camera motion and epipolar geometry as the problem of fitting a manifold to the data measured in a stereo pair. The manifold has a singularity and boundary, therefore care must be taken when fitting it. This paper reviews the notion of joint image manifold, and how previous motion recovery methods can be viewed in its context, and then offers a new fitting method, which improves upon previous results, especially when the extent of the data and/or the motion are small.
Liran Goshen, Ilan Shimshoni, P. Anandan 0001, Daniel Keren
ICCV3
2002 What Does the Scene Look Like from a Scene Point?
Michal Irani, Tal Hassner, P. Anandan 0001
ECCV (2)3
2002 Factorization with Uncertainty
P. Anandan 0001, Michal Irani
Int. J. Comput. Vis.1
2002 Direct Recovery of Planar-Parallax from Multiple Frames
abstract
We present an algorithm that estimates dense planar-parallax motion from multiple uncalibrated views of a 3D scene. This generalizes the "plane+parallax" recovery methods to more than two frames. The parallax motion of pixels across multiple frames (relative to a planar surface) is related to the 3D scene structure and the camera epipoles. The parallax field, the epipoles, and the 3D scene structure are estimated directly from image brightness variations across multiple frames, without precomputing correspondences.
Michal Irani, P. Anandan 0001, Meir Cohen
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 An Integrated Bayesian Approach to Layer Extraction from Image Sequences
abstract
This paper describes a Bayesian approach for modeling 3D scenes as collection of approximately planar layers that are arbitrarily positioned and oriented in the scene. In contrast to much of the previous work on layer-based motion modeling, which computes layered descriptions of 2D image motion, our work leads to a 3D description of the scene. There are two contributions within the paper. The first is to formulate the prior assumptions about the layers and scene within a Bayesian decision making framework which is used to automatically determine the number of layers and the assignment of individual pixels to layers. The second is algorithmic. In order to achieve the optimization, a Bayesian version of RANSAC is developed with which to initialize the segmentation. Then, a generalized expectation maximization method is used to find the MAP solution.
Philip Torr 0001, Richard Szeliski, P. Anandan 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2000 Segmenting Visual Actions Based on Spatio-Temporal Motion Patterns
abstract
The analysis of human action captured in video sequences has been a topic of considerable interest in computer vision. Much of the previous work has focused on the problem of action or activity recognition, but ignored the problem of detecting action boundaries in a video sequence containing unfamiliar and arbitrary visual actions. This paper presents an approach to this problem based on detecting temporal discontinuities of the spatial pattern of image motion that captures the action. We represent frame to frame optical-flow in terms of the coefficients of the most significant principal components computed from all the flow-fields within a given video sequence. We then detect the discontinuities in the temporal trajectories of these coefficients based on three different measures. We compare our segment boundaries against those detected by human observers on the same sequences in a recent independent psychological study of human perception of visual events. We show experimental results on the two sequences that were used in this study. Our experimental results are promising both from visual evaluation and when compared against the results of the psychological study.
Yong Rui, P. Anandan 0001
CVPR2
2000 Layer Extraction from Multiple Images Containing Reflections and Transparency
abstract
Many natural images contain reflections and transparency, i.e., they contain mixtures of reflected and transmitted light. When viewed from a moving camera, these appear as the superposition of component layer images moving relative to each other. The problem of multiple motion recovery has been previously studied by a number of researchers. However no one has yet demonstrated how to accurately recover the component images themselves. In this paper we develop an optimal approach to recovering layer images and their associated motions from an arbitrary number of composite images. We develop two different techniques for estimating the component layer images given known motion estimates. The first approach uses constrained least squares to recover the layer images. The second approach iteratively refines lower and upper bounds on the layer images using two novel compositing operations, namely minimum- and maximum-composites of aligned images. We combine these layer extraction techniques with a dominant motion estimator and a subsequent motion refinement stage. This results in a completely automated system that recovers transparent images and motions from a collection of input images.
Richard Szeliski, Shai Avidan, P. Anandan 0001
CVPR3
2000 Integrating Local Affine into Global Projective Images in the Joint Image Space
P. Anandan 0001, Shai Avidan
ECCV (1)1
2000 Factorization with Uncertainty
Michal Irani, P. Anandan 0001
ECCV (1)2
2000 The Geometry-Image Representation Tradeoff for Rendering
abstract
It is generally recognized that 3-D models are compact representations for rendering. While pure image-based rendering techniques are capable of producing highly photorealistic outputs, the size of the input "model" is usually very large. The important issues in trading off geometry versus images include compactness of representation, photorealism of reconstructed views, and speed of rendering. We describe our past work in modeling and rendering, and articulate lessons learnt. We then delineate our vision of an ideal rendering system.
Sing Bing Kang, Richard Szeliski, P. Anandan 0001
ICIP3
1999 Implicit Representation and Scene Reconstruction from Probability Density Functions
abstract
A technique is presented for representing linear features as probability density functions in two or three dimensions. Three chief advantages of this approach are (1) a unified representation and algebra for manipulating points, lines, and planes, (2) seamless incorporation of uncertainty information, and (3) a very simple recursive solution for maximum likelihood shape estimation. Applications to uncalibrated affine scene reconstruction are presented, with results on images of an outdoor environment.
Steven M. Seitz, P. Anandan 0001
CVPR2
1999 An Integrated Bayesian Approach to Layer Extraction from Image Sequences
abstract
This paper describes a Bayesian approach for modeling 3D scenes as a collection of approximately planar layers that are arbitrarily positioned and oriented in the scene. In contrast to much of the previous work on layer based motion modeling, which compute layered descriptions of 2D image motion, our work leads to a 3D description of the scene. We focus on the key problem of automatically segmenting the scene into layers as a precursor to recovery of stereo disparity data. The prior assumptions about the scene are formulated within a Bayesian decision making framework, and are then used to automatically determine the number of layers and the assignment of individual pixels to layers. Although using a collection of 3D layers has been previously proposed as an efficient and effective representation for multimedia applications, results to date have relied on hand segmentation. In contrast, the work described aims at fully automatic segmentation.
Philip Torr 0001, Richard Szeliski, P. Anandan 0001
ICCV3
1999 What Can Be Determined from a Full and a Weak Perspective Image?
abstract
This paper presents a first investigation on the structure from motion problem from the combination of full and weak perspective images. This problem arises in multiresolution object modeling where multiple zoomed-in or close-up views are combined with wider or distant reference views. The narrow field-of-view (FOV) images from the zoomed-in or closeup views can be approximated as weak perspective projection. Using a full perspective projection model for the narrow FOV images, although more accurate, actually leads to instabilities during the estimation process due to the non-linearities in the imaging model. The weak perspective approximation leads to more stable estimation algorithms, although at the cost of a small amount of modeling inaccuracy. Previous work in structure from motion focused either on two (or more) perspective images or on a set of weak perspective (more generally, affine) images. The main contribution of this paper is the study of the SFM problem for the much neglected case of one perspective and one (or more) weak perspective image. We show that in contrast to the case of a pair of weak perspective images, there is adequate information to recover Euclidean structure from a single perspective and a single weak perspective image. The epipolar geometry is simpler than with two perspective images leading to simpler and more stable estimation algorithms. Computer simulation shows that more stable results can be obtained with the technique presented in this paper than if two images are both considered to be full perspective.
Zhengyou Zhang, P. Anandan 0001, Harry Shum
ICCV2
1998 A Layered Approach to Stereo Reconstruction
abstract
We propose a framework for extracting structure from stereo which represents the scene as a collection of approximately planar layers. Each layer consists of an explicit 3D plane equation, a colored image with per-pixel opacity (a sprite), and a per-pixel depth offset relative to the plane. Initial estimates of the layers are recovered using techniques taken from parametric motion estimation. These initial estimates are then refined using a re-synthesis algorithm which takes into account both occlusions and mixed pixels. Reasoning about such effects allows the recovery of depth and color information with high accuracy even in partially occluded regions. Another important benefit of our framework is that the output consists of a collection of approximately planar regions, a representation which is far more appropriate than a dense depth map for many applications such as rendering and video parsing.
Simon Baker, Richard Szeliski, P. Anandan 0001
CVPR3
1998 From Reference Frames to Reference Planes: Multi-View Parallax Geometry and Applications
Michal Irani, P. Anandan 0001, Daphna Weinshall
ECCV (2)2
1998 Robust Multi-Sensor Image Alignment
abstract
This paper presents a method for alignment of images acquired by sensors of different modalities (e.g., EO and IR). The paper has two main contributions: (i) It identifies an appropriate image representation, for multi-sensor alignment, i.e., a representation which emphasizes the common information between the two multi-sensor images, suppresses the non-common information, and is adequate for coarse-to-fine processing. (ii) It presents a new alignment technique which applies global estimation to any choice of a local similarity measure. In particular, it is shown that when this registration technique is applied to the chosen image representation with a local normalized-correlation similarity measure, it provides a new multi-sensor alignment algorithm which is robust to outliers, and applies to a wide variety of globally complex brightness transformations between the two images. Our proposed image representation does not rely on sparse image features (e.g., edge, contour, or point features). It is continuous and does not eliminate the detailed variations within local image regions. Our method naturally extends to coarse-to-fine processing, and applies even in situations when the multi-sensor signals are globally characterized by low statistical correlation.
Michal Irani, P. Anandan 0001
ICCV2
1998 Interactive 3D modeling from multiple images using scene regularities
abstract
Due to the complexity of real scenes and the fragility of fully automated vision techniques, results from many automated modeling systems are disappointing. Automated techniques often require manual clean-up and postprocessing to segment the scene into coherent objects and surfaces, or to triangulate sparse point matches. They may also be required to enforce geometric constraints such as known orientations of surfaces. For instance, building interiors and exteriors provide vertical and horizontal lines and parallel and perpendicular planes. In this paper, we attack the 3D modeling problem from the other side: we specify some geometric knowledge ahead of time (e.g., known orientations of lines, co-planarity of points, initial scene segmentations), and use these constraints to guide our matching and reconstruction algorithms. We present two interactive (semi-automated) systems for recovering 3D models of large-scale environments from multiple images.
Harry Shum, Richard Szeliski, Simon Baker, P. Anandan 0001
WACV5
1998 A Unified Approach to Moving Object Detection in 2D and 3D Scenes
abstract
The detection of moving objects is important in many tasks. Previous approaches to this problem can be broadly divided into two classes: 2D algorithms which apply when the scene can be approximated by a flat surface and/or when the camera is only undergoing rotations and zooms, and 3D algorithms which work well only when significant depth variations are present in the scene and the camera is translating. We describe a unified approach to handling moving object detection in both 2D and 3D scenes, with a strategy to gracefully bridge the gap between those two extremes. Our approach is based on a stratification of the moving object detection problem into scenarios which gradually increase in their complexity. We present a set of techniques that match the above stratification. These techniques progressively increase in their complexity, ranging from 2D techniques to more complex 3D techniques. Moreover, the computations required for the solution to the problem at one complexity level become the initial processing step for the solution at the next complexity level. We illustrate these techniques using examples from real-image sequences.
Michal Irani, P. Anandan 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Video indexing based on mosaic representations
abstract
Video is a rich source of information. It provides visual information about scenes. This information is implicitly buried inside the raw video data, however, and is provided with the cost of very high temporal redundancy. While the standard sequential form of video storage is adequate for viewing in a movie mode, it fails to support rapid access to information of interest that is required in many of the emerging applications of video. This paper presents an approach for efficient access, use and manipulation of video data. The video data are first transformed from their sequential and redundant frame-based representation, in which the information about the scene is distributed over many frames, to an explicit and compact scene-based representation, to which each frame can be directly related. This compact reorganization of the video data supports nonlinear browsing and efficient indexing to provide rapid access directly to information of interest. This paper describes a new set of methods for indexing into the video sequence based on the scene-based representation. These indexing methods are based on geometric and dynamic information contained in the video. These methods complement the more traditional content-based indexing methods, which utilize image appearance information (namely, color and texture properties) but are considerably simpler to achieve and are highly computationally efficient.
Michal Irani, P. Anandan 0001
Proc. IEEE2
1997 Interactive content-based video indexing and browsing
abstract
In this paper we present a framework for efficient representation, access, and manipulation of video data. Our approach is based on decomposing video information into its spatial (appearance), temporal (dynamics), geometric components. This derived information is organized into data representations that support non-linear browsing and efficient indexing to provide rapid access directly to the information of interest.
Michal Irani, Harpreet Sawhney, Rakesh Kumar 0001, P. Anandan 0001
MMSP4
1996 Parallax Geometry of Pairs of Points for 3D Scene Analysis
Michal Irani, P. Anandan 0001
ECCV (1)2
1996 A unified approach to moving object detection in 2D and 3D scenes
abstract
The detection of moving objects is important in many tasks. Previous approaches to this problem can be broadly divided into two classes: 2D algorithms which apply when the scene can be approximated by a flat surface and/or when the camera is only undergoing rotations and zooms; and 3D algorithms which work well only when significant depth variations are present in the scene and the camera is translating. In this paper, we describe a unified approach to handling moving object detection in both 2D and 3D scenes, with a strategy to gracefully bridge the gap between those two extremes. Our approach is based on a stratification of the moving object detection problem into scenarios and corresponding techniques which gradually increase in their complexity. Moreover, the computations required for the solution to the problem at one complexity level become the initial processing step for the solution at the next complexity level.
Michal Irani, P. Anandan 0001
ICPR2
1996 The Robust Estimation of Multiple Motions: Parametric and Piecewise-Smooth Flow Fields
Michael J. Black, P. Anandan 0001
Comput. Vis. Image Underst.2
1996 Efficient representations of video sequences and their applications
Michal Irani, P. Anandan 0001, James R. Bergen, Rakesh Kumar 0001, Steven C. Hsu
Signal Process. Image Commun.2
1995 Mosaic Based Representations of Video Sequences and Their Applications
abstract
Recently, there has been a growing interest in the use of mosaic images to represent the information contained in video sequences. The paper systematically investigates how to go beyond thinking of the mosaic simply as a visualization device, but rather as a basis for efficient representation of video sequences. We describe two different types of mosaics called the static and the dynamic mosaic that are suitable for different needs and scenarios. We discuss a series of extensions to these basic mosaics to provide representations at multiple spatial and temporal resolutions and to handle 3D scene information. We describe techniques for the basic elements of the mosaic construction process, namely alignment, integration, and residual analysis. We describe several applications of mosaic representations including video compression, enhancement, enhanced visualization, and other applications in video indexing, search, and manipulation.>
Michal Irani, P. Anandan 0001, Steven C. Hsu
ICCV2
1995 Video as an image data source: efficient representations and applications
abstract
The two fundamental advantages of video over still imagery are: (i) the ability capture temporal information, and (ii) the ability to acquire a continuously varying set of views of a scene. These advantages are obtained, however, at the cost of vastly increased amount of data. This paper describes an approach to video representation that is based on frame-to-frame alignment, mosaic construction, and 3D parallax recovery. The basic motivation behind our approach is to enable rapid access to the contents, while maintaining the data in a form as close to the source as possible. This representation supports a wide variety of applications that involve transmission, storage, visualization, retrieval, analysis, and manipulation of video sequences.
P. Anandan 0001, Michal Irani, Rakesh Kumar 0001, James R. Bergen
ICIP1
1995 Video compression using mosaic representations
Michal Irani, Steven C. Hsu, P. Anandan 0001
Signal Process. Image Commun.3
1994 Adaptive-complexity registration of images
abstract
We present a framework for image registration algorithms that finds a lowest-order model of the flow between two images. Low-order models are useful in image registration, because they leave scene structure intact. But in real images complexity varies, and cannot be determined ahead of time. Algorithms in our framework adapt model complexity to image data during a coarse-fine parameter estimation process. Complexity increases keep residual flow small enough that motion can be correctly estimated at each subsequent resolution level. We present one algorithm within this framework which increases complexity by replacing global estimates with estimates over successively smaller patches. We show results of applying this algorithm to the task of mosaicing panoramic aerial images with unknown lens distortion and unknown camera position.>
J. R. Muller, P. Anandan 0001, James R. Bergen
CVPR2
1994 Direct recovery of shape from multiple views: a parallax based approach
abstract
Given two arbitrary views of a scene under central projection, if the motion of points on a parametric surface is compensated, the residual parallax displacement field on the reference image is an epipolar field. If the surface aligned is a plane, the parallax magnitude at an image point is directly proportional to the height of the point from the plane and inversely proportional to its depth from the camera. The authors exploit the above theorem to infer 3D height information from oblique aerial 2D images. The authors use direct methods to register the aerial images and develop methods to infer height information under the following three conditions: (i) focal length and image center are both known, (ii) only the focal length is known, and (iii) both are unknown.
Rakesh Kumar 0001, P. Anandan 0001, Keith J. Hanna
ICPR (1)2
1994 Accurate computation of optical flow by using layered motion representations
abstract
This paper presents a framework combining two prevailing approaches to motion analysis: optical flows which describes motion at each point, and methods that define global motions for larger regions. Image motion is represented by layers-image regions whose coherent motion can be approximated by some parametric motion model. The motion at every point is obtained by the parametric motion estimate of the entire layer, corrected by a residual flow which captures the difference between the real image motion and the layer's motion model. The new approach is able to construct accurate flow fields in the presence of multiple motions, motion boundaries, and transparent motions.
Steven C. Hsu, P. Anandan 0001, Shmuel Peleg
ICPR (1)2
1994 Real-time scene stabilization and mosaic construction
abstract
We describe a real-time system designed to construct a stable view of a scene through aligning images of an incoming video stream and dynamically constructing an image mosaic. This system uses a video processing unit developed by the David Sarnoff Research Center called the Vision Front End (VFE-100) for the pyramid-based image processing tasks required to implement this process. This paper includes a description of the multiresolution coarse-to-fine image registration strategy, the techniques used for mosaic construction, the implementation of this process on the VFE-100 system, and experimental results showing image mosaics constructed with the VFE-100.>
Michael W. Hansen, P. Anandan 0001, Kristin J. Dana, Gooitzen S. van der Wal, Peter J. Burt
WACV2
1994 Frameless registration of MR and CT 3D volumetric data sets
abstract
In this paper we present techniques for frameless registration of 3D Magnetic Resonance (MR) and Computed Tomography (CT) volumetric data of the head and spine. We present techniques for estimating a 3D affine or rigid transform which can be used to resample the CT (or MR) data to align with the MR (or CT) data. Our technique transforms the MR and CT data sets with spatial filters so they can be directly matched. The matching is done by a direct optimization technique using a gradient based descent approach and a coarse-to-fine control strategy over a 4D pyramid. We present results on registering the head and spine data by matching 3D edges and results on registering cranial ventricle data by matching images filtered by a Laplacian of a Gaussian.>
Rakesh Kumar 0001, Kristin J. Dana, P. Anandan 0001, Neil E. Okamoto, James R. Bergen, Paul F. Hemler, Thilaka S. Sumanaweera, Petra A. van den Elsen, John R. Adler Jr.
WACV3
1993 A framework for the robust estimation of optical flow
abstract
The authors consider the problem of robustly estimating optical flow from a pair of images using a new framework based on robust estimation which addresses violations of the brightness constancy and spatial smoothness assumptions. They also show the relationship between the robust estimation framework and line-process approaches for coping with spatial discontinuities. In doing so, the notion of a line process is generalized to that of an outlier process that can account for violations in both the brightness and smoothness assumptions. A graduated non-convexity algorithm is presented for recovering optical flow and motion discontinuities. The performance of the robust formulation is demonstrated on both synthetic data and natural images.>
Michael J. Black, P. Anandan 0001
ICCV2
1992 Hierarchical Model-Based Motion Estimation
James R. Bergen, P. Anandan 0001, Keith J. Hanna, Rajesh Hingorani
ECCV2
1991 Robust dynamic motion estimation over time
abstract
A novel approach to incrementally estimating visual motion over a sequence of images is presented. The authors start by formulating constraints on image motion to account for the possibility of multiple motions. This is achieved by exploiting the notions of weak continuity and robust statistics in the formulation of a minimization problem. The resulting objective function is non-convex. Traditional stochastic relaxation techniques for minimizing such functions prove inappropriate for the task. A highly parallel incremental stochastic minimization algorithm is presented which has a number of advantages over previous approaches. The incremental nature of the scheme makes it dynamic and permits the detection of occlusion and disocclusion boundaries.>
Michael J. Black, P. Anandan 0001
CVPR2
1991 Measurement of non-rigid motion using contour shape descriptors
abstract
The problem of measuring the motion of deformable objects from image sequences is addressed. The approach is based upon modeling the overall boundary of the object as a deformable contour and then tracking local segments of the contour through the temporal sequence. Motion computation involves first matching the local segments between pairs of contours by minimizing the deformation between the segments using a measure of bending energy. Results from the match process are incorporated into an optimization functional, along with a general smoothness term, whose local minimum results in a smooth flow field that is consistent with the match data. The computation is performed for all pairs of frames in the temporal sequence, resulting in a composite flow field over the entire sequence. The technique is applied to synthetic contour sequences and the problem of tracking left ventricular (LV) endocardial motion from medical image sequences.>
James S. Duncan, R. L. Owen, Lawrence H. Staib, P. Anandan 0001
CVPR4
1990 Constraints for the Early Detection of Discontinuity from Motion
Michael J. Black, P. Anandan 0001
AAAI2
1990 A model for the detection of motion over time
abstract
A model is proposed for the incremental estimation of visual motion fields from image sequences. The authors' model exploits three standard constraints on image motion within an optimization framework: (1) data conservation-the intensity structure of a surface patch changes gradually over time; (2) spatial coherence-neighboring points have similar motions; and (3) temporal coherence-the image velocity of a surface patch changes gradually. The authors' formulation takes into account the possibility of multiple motions at a particular location. They present an incremental scheme for the minimization of the objective function, based on simulated annealing. All computations are parallel, local, and incremental, and occlusion and disocclusion boundaries are estimated.>
Michael J. Black, P. Anandan 0001
ICCV2
1989 A computational framework and an algorithm for the measurement of visual motion
P. Anandan 0001
Int. J. Comput. Vis.1
1989 Optimization in Model Matching and Perceptual Organization
abstract
We introduce an optimization approach for solving problems in computer vision that involve multiple levels of abstraction. Our objective functions include compositional and specialization hierarchies. We cast vision problems as inexact graph matching problems, formulate graph matching in terms of constrained optimization, and use analog neural networks to perform the optimization. The method is applicable to perceptual grouping and model matching. Preliminary experimental results are shown.
Eric Mjolsness, Gene Gindi, P. Anandan 0001
Neural Comput.3
1988 Neural Networks for Model Matching and Perceptual Organization
Eric Mjolsness, Gene Gindi, P. Anandan 0001
NIPS3
1988 Neural network for model based recognition: Simulation results
abstract
The minimum (symmetric) rank of a simple graph G over a field F is the smallest possible rank among all symmetric matrices over F whose ijth entry (for i≠j) is nonzero whenever {i,j} is an edge in G and is zero otherwise. The problem of determining minimum (symmetric) rank has been studied extensively. We define the minimum skew rank of a simple graph G to be the smallest possible rank among all skew-symmetric matrices over F whose ijth entry (for i≠j) is nonzero whenever {i,j} is an edge in G and is zero otherwise. We apply techniques from the minimum (symmetric) rank problem and from skew-symmetric matrices to obtain results about the minimum skew rank problem.
Tony Zador, Gene Gindi, Eric Mjolsness, P. Anandan 0001
Neural Networks4
1985 Pattern-recognizing stochastic learning automata
abstract
A class of learning tasks is described that combines aspects of learning automation tasks and supervised learning pattern-classification tasks. These tasks are called associative reinforcement learning tasks. An algorithm is presented, called the associative reward-penalty, or AR-Palgorithm for which a form of optimal performance is proved. This algorithm simultaneously generalizes a class of stochastic learning automata and a class of supervised learning pattern-classification methods related to the Robbins-Monro stochastic approximation procedure. The relevance of this hybrid algorithm is discussed with respect to the collective behaviour of learning automata and the behaviour of networks of pattern-classifying adaptive elements. Simulation results are presented that illustrate the associative reinforcement learning task and the performance of the AR-Palgorithm as compared with that of several existing algorithms.
Andrew G. Barto, P. Anandan 0001
IEEE Trans. Syst. Man Cybern.2
1980 An Application of File-Comparison Algorithms to the Study of Program Editors
P. Anandan 0001, David W. Embley, George Nagy
Int. J. Man Mach. Stud.1