VLDB 2026 Research / reviewers in the wild / expert
Amit K. Agrawal
dblp:a/AmitKAgrawal
· DBLP profile ↗
44ranked-venue papers
19as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 19 first-authorArtificial intelligence and machine learning · 29 · 12 first-authorSystems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
27 papers |
Computational photography and imaging · 56% Image and video processing · 25% Rendering · 11% | |
| Artificial intelligence
13 papers |
3D vision · 88% Image recognition and object detection · 4% Face, body and person analysis · 4% |
Topics — the 30 heaviest of 61, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
camera calibration |
0.6 | 4 | 2013 | Extrinsic Camera Calibration without a Direct View Using Spherical Mirror · ICCV 2013 Single Image Calibration of Multi-axial Imaging Systems · CVPR 2013 A theory of multi-layer flat refractive geometry · CVPR 2012 |
Computational photography and imaging › image acquisition › imaging system design › camera design
coded exposure |
0.4 | 5 | 2009 | Invertible motion blur in video · ACM Trans. Graph. 2009 Coded exposure deblurring: Optimized codes for PSF estimation and invertibility · CVPR 2009 Optimal single image capture for motion deblurring · CVPR 2009 |
Image and video processing › image restoration › image deblurring
motion deblurring |
0.4 | 5 | 2009 | Invertible motion blur in video · ACM Trans. Graph. 2009 Coded exposure deblurring: Optimized codes for PSF estimation and invertibility · CVPR 2009 Optimal single image capture for motion deblurring · CVPR 2009 |
Computational photography and imaging
time-of-flight imaging |
0.3 | 2 | 2014 | Decomposing Global Light Transport Using Time of Flight Imaging · Int. J. Comput. Vis. 2014 Decomposing global light transport using time of flight imaging · CVPR 2012 |
Computational photography and imaging
3d scanning |
0.3 | 2 | 2013 | A Practical Approach to 3D Scanning in the Presence of Interreflections, Subsurface Scattering and Defocus · Int. J. Comput. Vis. 2013 Structured light 3D scanning in the presence of global illumination · CVPR 2011 |
Computational photography and imaging
light field imaging |
0.3 | 2 | 2013 | Towards Motion Aware Light Field Video for Dynamic Scenes · ICCV 2013 Axial light field for curved mirrors: Reflect your perspective, widen your view · CVPR 2010 |
Computational photography and imaging › light field imaging
light field capture |
0.2 | 3 | 2008 | Shield fields: modeling and capturing 3D occluders · ACM Trans. Graph. 2008 Non-refractive modulators for encoding and capturing scene appearance and depth · CVPR 2008 Dappled photography: mask enhanced cameras for heterodyned light fields and coded aperture refocusing · ACM Trans. Graph. 2007 |
Image and video processing › image restoration › image deblurring
blur kernel estimation |
0.2 | 2 | 2009 | Invertible motion blur in video · ACM Trans. Graph. 2009 Coded exposure deblurring: Optimized codes for PSF estimation and invertibility · CVPR 2009 |
Geometric modeling and processing › 3d reconstruction › shape from silhouette
visual hull reconstruction |
0.2 | 2 | 2012 | Convex bricks: A new primitive for visual hull modeling and reconstruction · ICRA 2012 Shield fields: modeling and capturing 3D occluders · ACM Trans. Graph. 2008 |
Computer vision › 3D vision › camera calibration
extrinsic calibration |
0.2 | 1 | 2013 | Extrinsic Camera Calibration without a Direct View Using Spherical Mirror · ICCV 2013 |
Computer vision › 3D vision › camera calibration
mirror-based calibration |
0.2 | 1 | 2013 | Extrinsic Camera Calibration without a Direct View Using Spherical Mirror · ICCV 2013 |
Image and video processing
image reconstruction |
0.2 | 1 | 2013 | Towards Motion Aware Light Field Video for Dynamic Scenes · ICCV 2013 |
Computational photography and imaging › light field imaging
light field video |
0.2 | 1 | 2013 | Towards Motion Aware Light Field Video for Dynamic Scenes · ICCV 2013 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.2 | 2 | 2010 | Specular surface reconstruction from sparse reflection correspondences · CVPR 2010 An Algebraic Approach to Surface Reconstruction from Gradient Fields · ICCV 2005 |
Geometric modeling and processing › surface reconstruction
gradient field integration |
0.2 | 2 | 2009 | Enforcing integrability by error correction using l1-minimization · CVPR 2009 What Is the Range of Surface Reconstructions from a Gradient Field? · ECCV (1) 2006 |
Image and video processing › image restoration
image deblurring |
0.2 | 2 | 2009 | Invertible motion blur in video · ACM Trans. Graph. 2009 Coded exposure photography: motion deblurring using fluttered shutter · ACM Trans. Graph. 2006 |
Geometric modeling and processing
surface reconstruction |
0.2 | 2 | 2009 | Enforcing integrability by error correction using l1-minimization · CVPR 2009 What Is the Range of Surface Reconstructions from a Gradient Field? · ECCV (1) 2006 |
Computer vision › 3D vision › range sensing
structured light |
0.1 | 1 | 2012 | Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns · ECCV (5) 2012 |
Rendering
global illumination |
0.1 | 1 | 2012 | Decomposing global light transport using time of flight imaging · CVPR 2012 |
Rendering
subsurface scattering |
0.1 | 1 | 2012 | Decomposing global light transport using time of flight imaging · CVPR 2012 |
Computer vision › 3D vision
depth estimation |
0.1 | 4 | 2012 | Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns · ECCV (5) 2012 Decomposing global light transport using time of flight imaging · CVPR 2012 Axial-cones: modeling spherical catadioptric cameras for wide-angle light field rendering · ACM Trans. Graph. 2010 |
Computer vision › 3D vision
3d reconstruction |
0.1 | 1 | 2011 | Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors · CVPR 2011 |
Computer vision › 3D vision › structure from motion
bundle adjustment |
0.1 | 1 | 2011 | Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrors · CVPR 2011 |
Computational photography and imaging › physics-based vision › light transport analysis
global illumination effects |
0.1 | 1 | 2011 | Structured light 3D scanning in the presence of global illumination · CVPR 2011 |
Computational photography and imaging › 3d scanning
structured light |
0.1 | 1 | 2011 | Structured light 3D scanning in the presence of global illumination · CVPR 2011 |
Computer vision › 3D vision › pose estimation
model-based pose estimation |
0.1 | 1 | 2010 | Pose estimation in heavy clutter using a multi-flash camera · ICRA 2010 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2010 | Pose estimation in heavy clutter using a multi-flash camera · ICRA 2010 |
Computer vision › 3D vision
object pose estimation |
0.1 | 1 | 2010 | Pose estimation in heavy clutter using a multi-flash camera · ICRA 2010 |
Computer vision › 3D vision › 3d shape reconstruction
specular surface reconstruction |
0.1 | 1 | 2010 | Specular surface reconstruction from sparse reflection correspondences · CVPR 2010 |
Computational photography and imaging
camera model |
0.1 | 1 | 2010 | Axial-cones: modeling spherical catadioptric cameras for wide-angle light field rendering · ACM Trans. Graph. 2010 |
Methods — techniques the papers use, named apart from their topics
gaussian modeling · 0.3exponential decay modeling · 0.3eigenvalue problem · 0.2time-of-flight imaging · 0.2spherical mirror reflection · 0.2sparse representation · 0.2multiplexing matrices · 0.2light transport modeling · 0.2dictionary learning · 0.2defocus modeling · 0.2axial camera model · 0.2analytical solution · 0.2temporal encoding · 0.1structured light · 0.1linear programming · 0.1five-point algorithm · 0.1essential matrix computation · 0.1convex decomposition · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | PoseNet3D: Learning Temporally Consistent 3D Human Pose via Knowledge DistillationabstractRecovering 3D human pose from 2D joints is a highly unconstrained problem. We propose a novel neural network framework, PoseNet3D, that takes 2D joints as input and outputs 3D skeletons and SMPL body model parameters. By casting our learning approach in a student-teacher framework, we avoid using any 3D data such as paired/unpaired 3D data, motion capture sequences, depth images or multi-view images during training. We first train a teacher network that outputs 3D skeletons, using only 2D poses for training. The teacher network distills its knowledge to a student network that predicts 3D pose in SMPL representation. Finally, both the teacher and the student networks are jointly fine-tuned in an end-to-end manner using temporal, self-consistency and adversarial losses, improving the accuracy of each individual network. Results on Human3.6M dataset for 3D human pose estimation demonstrate that our approach reduces the 3D joint prediction error by 18% compared to previous unsupervised methods. Qualitative results on in-the-wild datasets show that the recovered 3D poses and meshes are natural, realistic, and flow smoothly over consecutive frames. Shashank Tripathi, Siddhant Ranade, Ambrish Tyagi, Amit K. Agrawal |
3DV | 4 |
| 2014 | Decomposing Global Light Transport Using Time of Flight Imaging
Di Wu 0006, Andreas Velten, Matthew O'Toole, Belén Masiá, Amit K. Agrawal, Qionghai Dai, Ramesh Raskar |
Int. J. Comput. Vis. | 5 |
| 2013 | Single Image Calibration of Multi-axial Imaging SystemsabstractImaging systems consisting of a camera looking at multiple spherical mirrors (reflection) or multiple refractive spheres (refraction) have been used for wide-angle imaging applications. We describe such setups as multi-axial imaging systems, since a single sphere results in an axial system. Assuming an internally calibrated camera, calibration of such multi-axial systems involves estimating the sphere radii and locations in the camera coordinate system. However, previous calibration approaches require manual intervention or constrained setups. We present a fully automatic approach using a single photo of a 2D calibration grid. The pose of the calibration grid is assumed to be unknown and is also recovered. Our approach can handle unconstrained setups, where the mirrors/refractive balls can be arranged in any fashion, not necessarily on a grid. The axial nature of rays allows us to compute the axis of each sphere separately. We then show that by choosing rays from two or more spheres, the unknown pose of the calibration grid can be obtained linearly and independently of sphere radii and locations. Knowing the pose, we derive analytical solutions for obtaining the sphere radius and location. This leads to an interesting result that 6-DOF pose estimation of a multi-axial camera can be done without the knowledge of full calibration. Simulations and real experiments demonstrate the applicability of our algorithm. Amit K. Agrawal, Srikumar Ramalingam |
CVPR | 1 |
| 2013 | Extrinsic Camera Calibration without a Direct View Using Spherical MirrorabstractWe consider the problem of estimating the extrinsic parameters (pose) of a camera with respect to a reference 3D object without a direct view. Since the camera does not view the object directly, previous approaches have utilized reflections in a planar mirror to solve this problem. However, a planar mirror based approach requires a minimum of three reflections and has degenerate configurations where estimation fails. In this paper, we show that the pose can be obtained using a single reflection in a spherical mirror of known radius. This makes our approach simpler and easier in practice. In addition, unlike planar mirrors, the spherical mirror based approach does not have any degenerate configurations, leading to a robust algorithm. While a planar mirror reflection results in a virtual perspective camera, a spherical mirror reflection results in a non-perspective axial camera. The axial nature of rays allows us to compute the axis (direction of sphere center) and few pose parameters in a linear fashion. We then derive an analytical solution to obtain the distance to the sphere center and remaining pose parameters and show that it corresponds to solving a 16th degree equation. We present comparisons with a recent method that use planar mirrors and show that our approach recovers more accurate pose in the presence of noise. Extensive simulations and results on real data validate our algorithm. Amit K. Agrawal |
ICCV | 1 |
| 2013 | Towards Motion Aware Light Field Video for Dynamic ScenesabstractCurrent Light Field (LF) cameras offer fixed resolution in space, time and angle which is decided a-priori and is independent of the scene. These cameras either trade-off spatial resolution to capture single-shot LF or tradeoff temporal resolution by assuming a static scene to capture high spatial resolution LF. Thus, capturing high spatial resolution LF video for dynamic scenes remains an open and challenging problem. We present the concept, design and implementation of a LF video camera that allows capturing high resolution LF video. The spatial, angular and temporal resolution are not fixed a-priori and we exploit the scene-specific redundancy in space, time and angle. Our reconstruction is motion-aware and offers a continuum of resolution tradeoff with increasing motion in the scene. The key idea is (a) to design efficient multiplexing matrices that allow resolution tradeoffs, (b) use dictionary learning and sparse representations for robust reconstruction, and (c) perform local motion-aware adaptive reconstruction. We perform extensive analysis and characterize the performance of our motion-aware reconstruction algorithm. We show realistic simulations using a graphics simulator as well as real results using a LCoS based programmable camera. We demonstrate novel results such as high resolution digital refocusing for dynamic moving objects. Salil Tambe, Ashok Veeraraghavan, Amit K. Agrawal |
ICCV | 3 |
| 2013 | A Practical Approach to 3D Scanning in the Presence of Interreflections, Subsurface Scattering and Defocus
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
Int. J. Comput. Vis. | 2 |
| 2012 | A theory of multi-layer flat refractive geometryabstractFlat refractive geometry corresponds to a perspective camera looking through single/multiple parallel flat refractive mediums. We show that the underlying geometry of rays corresponds to an axial camera. This realization, while missing from previous works, leads us to develop a general theory of calibrating such systems using 2D-3D correspondences. The pose of 3D points is assumed to be unknown and is also recovered. Calibration can be done even using a single image of a plane. We show that the unknown orientation of the refracting layers corresponds to the underlying axis, and can be obtained independently of the number of layers, their distances from the camera and their refractive indices. Interestingly, the axis estimation can be mapped to the classical essential matrix computation and 5-point algorithm [15] can be used. After computing the axis, the thicknesses of layers can be obtained linearly when refractive indices are known, and we derive analytical solutions when they are unknown. We also derive the analytical forward projection (AFP) equations to compute the projection of a 3D point via multiple flat refractions, which allows non-linear refinement by minimizing the reprojection error. For two refractions, AFP is either 4th or 12th degree equation depending on the refractive indices. We analyze ambiguities due to small field of view, stability under noise, and show how a two layer system can be well approximated as a single layer system. Real experiments using a water tank validate our theory. Amit K. Agrawal, Srikumar Ramalingam, Yuichi Taguchi, Visesh Chari |
CVPR | 1 |
| 2012 | Decomposing global light transport using time of flight imagingabstractGlobal light transport is composed of direct and indirect components. In this paper, we take the first steps toward analyzing light transport using high temporal resolution information via time of flight (ToF) images. The time profile at each pixel encodes complex interactions between the incident light and the scene geometry with spatially-varying material properties. We exploit the time profile to decompose light transport into its constituent direct, subsurface scattering, and interreflection components. We show that the time profile is well modelled using a Gaussian function for the direct and interreflection components, and a decaying exponential function for the subsurface scattering component. We use our direct, subsurface scattering, and interreflection separation algorithm for four computer vision applications: recovering projective depth maps, identifying subsurface scattering objects, measuring parameters of analytical subsurface scattering models, and performing edge detection using ToF images. Di Wu 0006, Matthew O'Toole, Andreas Velten, Amit K. Agrawal, Ramesh Raskar |
CVPR | 4 |
| 2012 | Motion-Aware Structured Light Using Spatio-Temporal Decodable Patterns
Yuichi Taguchi, Amit K. Agrawal, Oncel Tuzel |
ECCV (5) | 2 |
| 2012 | Variable focus video: Reconstructing depth and video for dynamic scenesabstractTraditional depth from defocus (DFD) algorithms assume that the camera and the scene are static during acquisition time. In this paper, we examine the effects of camera and scene motion on DFD algorithms. We show that, given accurate estimates of optical flow (OF), one can robustly warp the focal stack (FS) images to obtain a virtual static FS and apply traditional DFD algorithms on the static FS. Acquiring accurate OF in the presence of varying focal blur is a challenging task. We show how defocus blur variations cause inherent biases in the estimates of optical flow. We then show how to robustly handle these biases and compute accurate OF estimates in the presence of varying focal blur. This leads to an architecture and an algorithm that converts a traditional 30 fps video camera into a co-located 30 fps image and a range sensor. Further, the ability to extract image and range information allows us to render images with artistic depth-of field effects, both extending and reducing the depth of field of the captured images. We demonstrate experimental results on challenging scenes captured using a camera prototype. Nitesh Shroff, Ashok Veeraraghavan, Yuichi Taguchi, Oncel Tuzel, Amit K. Agrawal, Rama Chellappa |
ICCP | 5 |
| 2012 | Convex bricks: A new primitive for visual hull modeling and reconstructionabstractIndustrial automation tasks typically require a 3D model of the object for robotic manipulation. The ability to reconstruct the 3D model using a sample object is useful when CAD models are not available. For textureless objects, visual hull of the object obtained using silhouette-based reconstruction can avoid expensive 3D scanners for 3D modeling. We propose convex brick (CB), a new 3D primitive for modeling and reconstructing a visual hull from silhouettes. CB's are powerful in modeling arbitrary non-convex 3D shapes. Using CB, we describe an algorithm to generate a polyhedral visual hull from polygonal silhouettes; the visual hull is reconstructed as a combination of 3D convex bricks. Our approach uses well-studied geometric operations such as 2D convex decomposition and intersection of 3D convex cones using linear programming. The shape of CB can adapt to the given silhouettes, thereby significantly reducing the number of primitives required for a volumetric representation. Our framework allows easy control of reconstruction parameters such as accuracy and the number of required primitives. We present an extensive analysis of our algorithm and show visual hull reconstruction on challenging real datasets consisting of highly non-convex shapes. We also show real results on pose estimation of an industrial part in a bin-picking system using the reconstructed visual hull. Visesh Chari, Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
ICRA | 2 |
| 2011 | Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrorsabstractCatadioptric cameras are widely used to increase the field of view using mirrors. Central catadioptric systems having an effective single viewpoint are easy to model and use, but severely constraint the camera positioning with respect to the mirror. On the other hand, non-central catadioptric systems allow greater flexibility in camera placement, but are often approximated using central or linear models due to the lack of an exact model. We bridge this gap and describe an exact projection model for non-central catadioptric systems. We derive an analytical `forward projection' equation for the projection of a 3D point reflected by a quadric mirror on the imaging plane of a perspective camera, with no restrictions on the camera placement, and show that it is an 8thdegree equation in a single unknown. While previous non-central catadioptric cameras primarily use an axial configuration where the camera is placed on the axis of a rotationally symmetric mirror, we allow off-axis (any) camera placement. Using this analytical model, a non-central catadioptric camera can be used for sparse as well as dense 3D reconstruction similar to perspective cameras, using well-known algorithms such as bundle adjustment and plane sweeping. Our paper is the first to show such results for off-axis placement of camera with multiple quadric mirrors. Simulation and real results using parabolic mirrors and an off-axis perspective camera are demonstrated. Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
CVPR | 1 |
| 2011 | Structured light 3D scanning in the presence of global illuminationabstractGlobal illumination effects such as inter-reflections, diffusion and sub-surface scattering severely degrade the performance of structured light-based 3D scanning. In this paper, we analyze the errors caused by global illumination in structured light-based shape recovery. Based on this analysis, we design structured light patterns that are resilient to individual global illumination effects using simple logical operations and tools from combinatorial mathematics. Scenes exhibiting multiple phenomena are handled by combining results from a small ensemble of such patterns. This combination also allows us to detect any residual errors that are corrected by acquiring a few additional images. Our techniques do not require explicit separation of the direct and global components of scene radiance and hence work even in scenarios where the separation fails or the direct component is too low. Our methods can be readily incorporated into existing scanning systems without significant overhead in terms of capture time or hardware. We show results on a variety of scenes with complex shape and material properties and challenging global illumination effects. Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 2 |
| 2010 | Optimal coded sampling for temporal super-resolutionabstractConventional low frame rate cameras result in blur and/or aliasing in images while capturing fast dynamic events. Multiple low speed cameras have been used previously with staggered sampling to increase the temporal resolution. However, previous approaches are inefficient: they either use small integration time for each camera which does not provide light benefit, or use large integration time in a way that requires solving a big ill-posed linear system. We propose coded sampling that address these issues: using N cameras it allows N times temporal superresolution while allowing ~N/2 times more light compared to an equivalent high speed camera. In addition, it results in a well-posed linear system which can be solved independently for each frame, avoiding reconstruction artifacts and significantly reducing the computational time and memory. Our proposed sampling uses optimal multiplexing code considering additive Gaussian noise to achieve the maximum possible SNR in the recovered video. We show how to implement coded sampling on off-the-shelf machine vision cameras. We also propose a new class of invertible codes that allow continuous blur in captured frames, leading to an easier hardware implementation. Amit K. Agrawal, Mohit Gupta 0001, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
CVPR | 1 |
| 2010 | Specular surface reconstruction from sparse reflection correspondencesabstractWe present a practical approach for surface reconstruction of smooth mirror-like objects using sparse reflection correspondences (RCs). Assuming finite object motion with a fixed camera and un-calibrated environment, we derive the relationship between RC and the surface shape. We show that by locally modeling the surface as a quadric, the relationship between the RCs and unknown surface parameters becomes linear. We develop a simple surface reconstruction algorithm that amounts to solving either an eigenvalue problem or a second order cone program (SOCP). Ours is the first method that allows for reconstruction of mirror surfaces from sparse RCs, obtained from standard algorithms such as SIFT. Our approach overcomes the practical issues in shape from specular flow (SFSF) such as the requirement of dense optical flow and undefined/infinite flow at parabolic points. We also show how to incorporate auxiliary information such as sparse surface normals into our framework. Experiments, both real and synthetic are shown that validate the theory presented. Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
CVPR | 4 |
| 2010 | Axial light field for curved mirrors: Reflect your perspective, widen your viewabstractMirrors have been used to enable wide field-of-view (FOV) catadioptric imaging. The mapping between the incoming and reflected light rays depends non-linearly on the mirror shape and has been well-studied using caustics. We analyze this mapping using two-plane light field parameterization, which provides valuable insight into the geometric structure of reflected rays. Using this analysis, we study the problem of generating a single-viewpoint virtual perspective image for catadioptric systems, which is unachievable for several common configurations. Instead of minimizing distortions appearing in a single image, we propose to capture all the rays required to generate a virtual perspective by capturing a light field. We consider rotationally symmetric mirrors and show that a traditional planar light field results in significant aliasing artifacts. We propose axial light field, captured by moving the camera along the mirror rotation axis, for efficient sampling and to remove aliasing artifacts. This allows us to computationally generate wide FOV virtual perspectives using a wider class of mirrors than before, without using scene priors or depth estimation. We analyze the relationship between the axial light field parameters and the FOV/resolution of the resulting virtual perspective. Real results using a spherical mirror demonstrate generating 140° FOV virtual perspective using multiple 30° FOV images. Yuichi Taguchi, Amit K. Agrawal, Srikumar Ramalingam, Ashok Veeraraghavan |
CVPR | 2 |
| 2010 | Analytical Forward Projection for Axial Non-central Dioptric and Catadioptric Cameras
Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
ECCV (3) | 1 |
| 2010 | Flexible Voxels for Motion-Aware Videography
Mohit Gupta 0001, Amit K. Agrawal, Ashok Veeraraghavan, Srinivasa G. Narasimhan |
ECCV (1) | 2 |
| 2010 | Image Invariants for Smooth Reflective Surfaces
Aswin C. Sankaranarayanan, Ashok Veeraraghavan, Oncel Tuzel, Amit K. Agrawal |
ECCV (2) | 4 |
| 2010 | Pose estimation in heavy clutter using a multi-flash cameraabstractWe propose a novel solution to object detection, localization and pose estimation with applications in robot vision. The proposed method is especially applicable when the objects of interest may not be richly textured and are immersed in heavy clutter. We show that a multi-flash camera (MFC) provides accurate separation of depth edges and texture edges in such scenes. Then, we reformulate the problem, as one of finding matches between the depth edges obtained in one or more MFC images to the rendered depth edges that are computed offline using 3D CAD model of the objects. In order to facilitate accurate matching of these binary depth edge maps, we introduce a novel cost function that respects both the position and the local orientation of each edge pixel. This cost function is significantly superior to traditional Chamfer cost and leads to accurate matching even in heavily cluttered scenes where traditional methods are unreliable. We present a sub-linear time algorithm to compute the cost function using techniques from 3D distance transforms and integral images. Finally, we also propose a multi-view based pose-refinement algorithm to improve the estimated pose. We implemented the algorithm on an industrial robot arm and obtained location and angular estimation accuracy of the order of 1 mm and 2° respectively for a variety of parts with minimal texture. Ming-Yu Liu 0001, Oncel Tuzel, Ashok Veeraraghavan, Rama Chellappa, Amit K. Agrawal, Haruhisa Okuda |
ICRA | 5 |
| 2010 | Reinterpretable Imager: Towards Variable Post-Capture Space, Angle and Time Resolution in PhotographyabstractAbstract We describe a novel multiplexing approach to achieve tradeoffs in space, angle and time resolution in photography. We explore the problem of mapping useful subsets of time‐varying 4D lightfields in a single snapshot. Our design is based on using a dynamic mask in the aperture and a static mask close to the sensor. The key idea is to exploit scene‐specific redundancy along spatial, angular and temporal dimensions and to provide a programmable or variable resolution tradeoff among these dimensions. This allows a user to reinterpret the single captured photo as either a high spatial resolution image, a refocusable image stack or a video for different parts of the scene in post‐processing. A lightfield camera or a video camera forces a‐priori choice in space‐angle‐time resolution. We demonstrate a single prototype which provides flexible post‐capture abilities not possible using either a single‐shot lightfield camera or a multi‐frame video camera. We show several novel results including digital refocusing on objects moving in depth and capturing multiple facial expressions in a single photo. Amit K. Agrawal, Ashok Veeraraghavan, Ramesh Raskar |
Comput. Graph. Forum | 1 |
| 2010 | Axial-cones: modeling spherical catadioptric cameras for wide-angle light field renderingabstractCatadioptric imaging systems are commonly used for wide-angle imaging, but lead to multi-perspective images which do not allow algorithms designed for perspective cameras to be used. Efficient use of such systems requires accurate geometric ray modeling as well as fast algorithms. We present accurate geometric modeling of the multi-perspective photo captured with a spherical catadioptric imaging system usingaxial-cone cameras:multiple perspective cameras lying on an axis each with a different viewpoint and a different cone of rays. This modeling avoids geometric approximations and allows several algorithms developed for perspective cameras to be applied to multi-perspective catadioptric cameras. We demonstrate axial-cone modeling in the context of rendering wide-angle light fields, captured using a spherical mirror array. We present several applications such as spherical distortion correction, digital refocusing for artistic depth of field effects in wide-angle scenes, and wide-angle dense depth estimation. Our GPU implementation using axial-cone modeling achieves up to three orders of magnitude speed up over ray tracing for these applications. Yuichi Taguchi, Amit K. Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam, Ramesh Raskar |
ACM Trans. Graph. | 2 |
| 2009 | Optimal single image capture for motion deblurringabstractDeblurring images of moving objects captured from a traditional camera is an ill-posed problem due to the loss of high spatial frequencies in the captured images. Techniques have attempted to engineer the motion point spread function (PSF) by either making it invertible using coded exposure, or invariant to motion by moving the camera in a specific fashion. We address the problem of optimal single image capture strategy for best deblurring performance. We formulate the problem of optimal capture as maximizing the signal to noise ratio (SNR) of the deconvolved image given a scene light level. As the exposure time increases, the sensor integrates more light, thereby increasing the SNR of the captured signal. However, for moving objects, larger exposure time also results in more blur and hence more deconvolution noise. We compare the following three single image capture strategies: (a) traditional camera, (b) coded exposure camera, and (c) motion invariant photography, as well as the best exposure time for capture by analyzing the rate of increase of deconvolution noise with exposure time. We analyze which strategy is optimal for known/unknown motion direction and speed and investigate how the performance degrades for other cases. We present real experimental results by simulating the above capture strategies using a high speed video camera. Amit K. Agrawal, Ramesh Raskar |
CVPR | 1 |
| 2009 | Coded exposure deblurring: Optimized codes for PSF estimation and invertibilityabstractWe consider the problem of single image object motion deblurring from a static camera. It is well-known that deblurring of moving objects using a traditional camera is ill-posed, due to the loss of high spatial frequencies in the captured blurred image. A coded exposure camera modulates the integration pattern of light by opening and closing the shutter within the exposure time using a binary code. The code is chosen to make the resulting point spread function (PSF) invertible, for best deconvolution performance. However, for a successful deconvolution algorithm, PSF estimation is as important as PSF invertibility. We show that PSF estimation is easier if the resulting motion blur is smooth and the optimal code for PSF invertibility could worsen PSF estimation, since it leads to non-smooth blur. We show that both criterions of PSF invertibility and PSF estimation can be simultaneously met, albeit with a slight increase in the deconvolution noise. We propose design rules for a code to have good PSF estimation capability and outline two search criteria for finding the optimal code for a given length. We present theoretical analysis comparing the performance of the proposed code with the code optimized solely for PSF invertibility. We also show how to easily implement coded exposure on a consumer grade machine vision camera with no additional hardware. Real experimental results demonstrate the effectiveness of the proposed codes for motion deblurring. Amit K. Agrawal |
CVPR | 1 |
| 2009 | 3D pose estimation and segmentation using specular cuesabstractWe present a system for fast model-based segmentation and 3D pose estimation of specular objects using appearance based specular features. We use observed (a) specular reflection and (b) specular flow as cues, which are matched against similar cues generated from a CAD model of the object in various poses. We avoid estimating 3D geometry or depths, which is difficult and unreliable for specular scenes. In the first method, the environment map of the scene is utilized to generate a database containing synthesized specular reflections of the object for densely sampled 3D poses. This database is compared with captured images of the scene at run time to locate and estimate the 3D pose of the object. In the second method, specular flows are generated for dense 3D poses as illumination invariant features and are matched to the specular flow of the scene. We incorporate several practical heuristics such as use of saturated/highlight pixels for fast matching and normal selection to minimize the effects of inter-reflections and cluttered backgrounds. Despite its simplicity, our approach is effective in scenes with multiple specular objects, partial occlusions, inter-reflections, cluttered backgrounds and changes in ambient illumination. Experimental results demonstrate the effectiveness of our method for various synthetic and real objects. Ju Yong Chang, Ramesh Raskar, Amit K. Agrawal |
CVPR | 3 |
| 2009 | Enforcing integrability by error correction using l1-minimizationabstractSurface reconstruction from gradient fields is an important final step in several applications involving gradient manipulations and estimation. Typically, the resulting gradient field is non-integrable due to linear/non-linear gradient manipulations, or due to presence of noise/outliers in gradient estimation. In this paper, we analyze integrability as error correction, inspired from recent work in compressed sensing, particulary ℓ0- ℓ1equivalence. We propose to obtain the surface by finding the gradient field which best fits the corrupted gradient field in ℓ1sense. We present an exhaustive analysis of the properties of ℓ1solution for gradient field integration using linear algebra and graph analogy. We consider three cases: (a) noise, but no outliers (b) no-noise but outliers and (c) presence of both noise and outliers in the given gradient field. We show that ℓ1solution performs as well as least squares in the absence of outliers. While previous ℓ0- ℓ1equivalence work has focused on the number of errors (outliers), we show that the location of errors is equally important for gradient field integration. We characterize the ℓ1solution both in terms of location and number of outliers, and outline scenarios where ℓ1solution is equivalent to ℓ0solution. We also show that when ℓ1solution is not able to remove outliers, the property of local error confinement holds: i.e., the errors do not propagate to the entire surface as in least squares. We compare with previous techniques and show that ℓ1solution performs well across all scenarios without the need for any tunable parameter adjustments. Dikpal Reddy, Amit K. Agrawal, Rama Chellappa |
CVPR | 2 |
| 2009 | Invertible motion blur in videoabstractWe show that motion blur in successive video frames is invertible even if the point-spread function (PSF) due to motion smear in a single photo is non-invertible. Blurred photos exhibit nulls (zeros) in the frequency transform of the PSF, leading to an ill-posed deconvolution. Hardware solutions to avoid this require specialized devices such as the coded exposure camera or accelerating sensor motion. We employ ordinary video cameras and introduce the notion of null-filling along with joint-invertibility of multiple blur-functions. The key idea is to record the same object with varying PSFs, so that the nulls in the frequency component of one frame can be filled by other frames. The combined frequency transform becomes null-free, making deblurring well-posed. We achieve jointly-invertible blur simply by changing the exposure time of successive frames. We address the problem of automatic deblurring of objects moving with constant velocity by solving the four critical components: preservation of all spatial frequencies, segmentation of moving parts, motion estimation of moving parts, and non-degradation of the static parts of the scene. We demonstrate several challenging cases of object motion blur including textured backgrounds and partial occluders. Amit K. Agrawal, Ramesh Raskar |
ACM Trans. Graph. | 1 |
| 2008 | Non-refractive modulators for encoding and capturing scene appearance and depthabstractWe analyze the modulation of a light field via non-refracting attenuators. In the most general case, any desired modulation can be achieved with attenuators having four degrees of freedom in ray-space. We motivate the discussion with a universal 4D ray modulator (ray-filter) which can attenuate the intensity of each ray independently. We describe operation of such a fantasy ray-filter in the context of altering the 4D light field incident on a 2D camera sensor. Ray-filters are difficult to realize in practice but we can achieve reversible encoding for light field capture using patterned attenuating mask. Two mask-based designs are analyzed in this framework. The first design closely mimics the angle-dependent ray-sorting possible with the ray filter. The second design exploits frequency-domain modulation to achieve a more efficient encoding. We extend these designs for optimal sampling of light field by matching the modulation function to the specific shape of the band-limit frequency transform of light field. We also show how a hand-held version of an attenuator based light field camera can be built using a medium-format digital camera and an inexpensive mask. Ashok Veeraraghavan, Amit K. Agrawal, Ramesh Raskar, Ankit Mohan, Jack Tumblin |
CVPR | 2 |
| 2008 | Shield fields: modeling and capturing 3D occludersabstractWe describe a unified representation of occluders in light transport and photography using shield fields: the 4D attenuation function which acts on any light field incident on an occluder. Our key theoretical result is that shield fields can be used to decouple the effects of occluders and incident illumination. We first describe the properties of shield fields in the frequency-domain and briefly analyze the "forward" problem of efficiently computing cast shadows. Afterwards, we apply the shield field signal-processing framework to make several new observations regarding the "inverse" problem of reconstructing 3D occluders from cast shadows -- extending previous work on shape-from-silhouette and visual hull methods. From this analysis we develop the first single-camera, single-shot approach to capture visual hulls without requiring moving or programmable illumination. We analyze several competing camera designs, ultimately leading to the development of a new large-format, mask-based light field camera that exploits optimal tiled-broadband codes for light-efficient shield field capture. We conclude by presenting a detailed experimental analysis of shield field capture and 3D occluder reconstruction. Douglas Lanman, Ramesh Raskar, Amit K. Agrawal, Gabriel Taubin |
ACM Trans. Graph. | 3 |
| 2008 | Glare aware photography: 4D ray sampling for reducing glare effects of camera lensesabstractGlare arises due to multiple scattering of light inside the camera's body and lens optics and reduces image contrast. While previous approaches have analyzed glare in 2D image space, we show that glare is inherently a 4D ray-space phenomenon. By statistically analyzing the ray-space inside a camera, we can classify and remove glare artifacts. In ray-space, glare behaves as high frequency noise and can be reduced by outlier rejection. While such analysis can be performed by capturing the light field inside the camera, it results in the loss of spatial resolution. Unlike light field cameras, we do not need to reversibly encode the spatial structure of the ray-space, leading to simpler designs. We explore masks for uniform and non-uniform ray sampling and show a practical solution to analyze the 4D statistics without significantly compromising image resolution. Although diffuse scattering of the lens introduces 4D low-frequency glare, we can produce useful solutions in a variety of common scenarios. Our approach handles photography looking into the sun and photos taken without a hood, removes the effect of lens smudges and reduces loss of contrast due to camera body reflections. We show various applications in contrast enhancement and glare manipulation. Ramesh Raskar, Amit K. Agrawal, Cyrus A. Wilson, Ashok Veeraraghavan |
ACM Trans. Graph. | 2 |
| 2007 | Detecting and Segmenting Un-occluded Items by Actively Casting Shadows
Tze Ki Koh, Amit K. Agrawal, Ramesh Raskar, Steve Morgan, Nicholas Miles, Barrie Hayes-Gill |
ACCV (1) | 2 |
| 2007 | Resolving Objects at Higher Resolution from a Single Motion-blurred ImageabstractMotion blur can degrade the quality of images and is considered a nuisance for computer vision problems. In this paper, we show that motion blur can in-fact be used for increasing the resolution of a moving object. Our approach utilizes the information in a single motion-blurred image without any image priors or training images. As the blur size increases, the resolution of the moving object can be enhanced by a larger factor, albeit with a corresponding increase in reconstruction noise. Traditionally, motion deblurring and super-resolution have been ill-posed problems. Using a coded-exposure camera that preserves high spatial frequencies in the blurred image, we present a linear algorithm for the combined problem of deblurring and resolution enhancement and analyze the invertibility of the resulting linear system. We also show a method to selectively enhance the resolution of a narrow region of high-frequency features, when the resolution of the entire moving object cannot be increased due to small motion blur. Results on real images showing up to four times resolution enhancement are presented. Amit K. Agrawal, Ramesh Raskar |
CVPR | 1 |
| 2007 | Dappled photography: mask enhanced cameras for heterodyned light fields and coded aperture refocusingabstractWe describe a theoretical framework for reversibly modulating 4D light fields using an attenuating mask in the optical path of a lens based camera. Based on this framework, we present a novel design to reconstruct the 4D light field from a 2D camera image without any additional refractive elements as required by previous light field cameras. The patterned mask attenuates light rays inside the camera instead of bending them, and the attenuation recoverably encodes the rays on the 2D sensor. Our mask-equipped camera focuses just as a traditional camera to capture conventional 2D photos at full sensor resolution, but the raw pixel values also hold a modulated 4D light field. The light field can be recovered by rearranging the tiles of the 2D Fourier transform of sensor values into 4D planes, and computing the inverse Fourier transform. In addition, one can also recover the full resolution image information for the in-focus parts of the scene. We also show how a broadband mask placed at the lens enables us to compute refocused images at full sensor resolution for layered Lambertian scenes. This partial encoding of 4D ray-space data enables editing of image contents by depth, yet does not require computational recovery of the complete 4D light field. Ashok Veeraraghavan, Ramesh Raskar, Amit K. Agrawal, Ankit Mohan, Jack Tumblin |
ACM Trans. Graph. | 3 |
| 2006 | Edge Suppression by Gradient Field Transformation Using Cross-Projection TensorsabstractWe propose a new technique for edge-suppressing operations on images. We introduce cross projection tensors to achieve affine transformations of gradient fields. We use these tensors, for example, to remove edges in one image based on the edge-information in a second image. Traditionally, edge suppression is achieved by setting image gradients to zero based on thresholds. A common application is in the Retinex problem, where the illumination map is recovered by suppressing the reflectance edges, assuming it is slowly varying. We present a class of problems where edge-suppression can be a useful tool. These problems involve analyzing images of the same scene under variable illumination. Instead of resetting gradients, the key idea in our approach is to derive local tensors using one image and to transform the gradient field of another image using them. Reconstructed image from the modified gradient field shows suppressed edges or textures at the corresponding locations. All operations are local and our approach does not require any global analysis. We demonstrate the algorithm in the context of several applications such as (a) recovering the foreground layer under varying illumination, (b) estimating intrinsic images in non-Lambertian scenes, (c) removing shadows from color images and obtaining the illumination map, and (d) removing glass reflections. Amit K. Agrawal, Ramesh Raskar, Rama Chellappa |
CVPR (2) | 1 |
| 2006 | What Is the Range of Surface Reconstructions from a Gradient Field?
Amit K. Agrawal, Ramesh Raskar, Rama Chellappa |
ECCV (1) | 1 |
| 2006 | Robust ego-motion estimation and 3-D model refinement using surface parallaxabstractWe present an iterative algorithm for robustly estimating the ego-motion and refining and updating a coarse depth map using parametric surface parallax models and brightness derivatives extracted from an image pair. Given a coarse depth map acquired by a range-finder or extracted from a digital elevation map (DEM), ego-motion is estimated by combining a global ego-motion constraint and a local brightness constancy constraint. Using the estimated camera motion and the available depth estimate, motion of the three-dimensional (3-D) points is compensated. We utilize the fact that the resulting surface parallax field is an epipolar field, and knowing its direction from the previous motion estimates, estimate its magnitude and use it to refine the depth map estimate. The parallax magnitude is estimated using a constant parallax model (CPM) which assumes a smooth parallax field and a depth based parallax model (DBPM), which models the parallax magnitude using the given depth map. We obtain confidence measures for determining the accuracy of the estimated depth values which are used to remove regions with potentially incorrect depth estimates for robustly estimating ego-motion in subsequent iterations. Experimental results using both synthetic and real data (both indoor and outdoor sequences) illustrate the effectiveness of the proposed algorithm. Amit K. Agrawal, Rama Chellappa |
IEEE Trans. Image Process. | 1 |
| 2006 | Coded exposure photography: motion deblurring using fluttered shutterabstractIn a conventional single-exposure photograph, moving objects or moving cameras cause motion blur. The exposure time defines a temporal box filter that smears the moving object across the image by convolution. This box filter destroys important high-frequency spatial details so that deblurring via deconvolution becomes an ill-posed problem.Rather than leaving the shutter open for the entire exposure duration, we "flutter" the camera's shutter open and closed during the chosen exposure time with a binary pseudo-random sequence. The flutter changes the box filter to a broad-band filter that preserves high-frequency spatial details in the blurred image and the corresponding deconvolution becomes a well-posed problem. We demonstrate that manually-specified point spread functions are sufficient for several challenging cases of motion-blur removal including extremely large motions, textured backgrounds and partial occluders. Ramesh Raskar, Amit K. Agrawal, Jack Tumblin |
ACM Trans. Graph. | 2 |
| 2005 | Why I Want a Gradient CameraabstractWe propose a camera that measures static gradients instead of static intensities. Quantizing sensed intensity differences between adjacent pixel values permits an ordinary A/D converter to measure detailed high contrast (HDR) scenes. We measure alternating 'cliques' of sensors (small groups) that locally determine their own best exposure, and reconstruct the image using a Poisson solver. This intrinsically differential design suppresses common-mode noise, hides and smoothes quantization, and can correct for its own saturated sensors. Simulations demonstrate these capabilities in side-by-side comparisons. Jack Tumblin, Amit K. Agrawal, Ramesh Raskar |
CVPR (1) | 2 |
| 2005 | Moving Object Segmentation and Dynamic Scene Reconstruction Using Two FramesabstractIn this paper, a two-frame approach is presented for segmentation of independent moving objects in video along with estimation of ego-motion, independent object motion and reconstruction of the dynamic scene using intensity images. The proposed method utilizes the least median of squares in estimating ego-motion and parallax constraints for segmenting independently moving objects. A 3D structure for the static scene is also estimated using surface parallax. The motion of moving objects is estimated by first fitting a parametric flow model followed by subspace analysis. The algorithm works well for unconstrained translational motion of moving objects. Amit K. Agrawal, Rama Chellappa |
ICASSP (2) | 1 |
| 2005 | An Algebraic Approach to Surface Reconstruction from Gradient FieldsabstractSeveral important problems in computer vision such as shape from shading (SFS) and photometric stereo (PS) require reconstructing a surface from an estimated gradient field, which is usually non-integrable, i.e. have non-zero curl. We propose a purely algebraic approach to enforce integrability in discrete domain. We first show that enforcing integrability can be formulated as solving a single linear system Ax =b over the image. In general, this system is under-determined. We show conditions under which the system can be solved and a method to get to those conditions based on graph theory. The proposed approach is non-iterative, has the important property of local error confinement and can be applied to several problems. Results on SFS and PS demonstrate the applicability of our method. Amit K. Agrawal, Rama Chellappa, Ramesh Raskar |
ICCV | 1 |
| 2005 | Removing photography artifacts using gradient projection and flash-exposure samplingabstractFlash images are known to suffer from several problems: saturation of nearby objects, poor illumination of distant objects, reflections of objects strongly lit by the flash and strong highlights due to the reflection of flash itself by glossy surfaces. We propose to use a flash and no-flash (ambient) image pair to produce better flash images. We present a novel gradient projection scheme based on a gradient coherence model that allows removal of reflections and highlights from flash images. We also present a brightness-ratio based algorithm that allows us to compensate for the falloff in the flash image brightness due to depth. In several practical scenarios, the quality of flash/no-flash images may be limited in terms of dynamic range. In such cases, we advocate using several images taken under different flash intensities and exposures. We analyze the flash intensity-exposure space and propose a method for adaptively sampling this space so as to minimize the number of captured images for any given scene. We present several experimental results that demonstrate the ability of our algorithms to produce improved flash images. Amit K. Agrawal, Ramesh Raskar, Shree K. Nayar, Yuanzhen Li |
ACM Trans. Graph. | 1 |
| 2004 | 3D model refinement using surface-parallaxabstractWe present an approach to update and refine coarse 3D models of urban environments from a sequence of intensity images using surface parallax. This generalizes the plane-parallax recovery methods to surface-parallax using arbitrary surfaces. A coarse and potentially incomplete depth map of the scene obtained from a digital elevation map (DEM) is used as a reference surface which is refined and updated using this approach. The reference depth map is used to estimate the camera motion and the motion of the 3D points on the reference surface is compensated. The resulting parallax, which is an epipolar field, is estimated using an adaptive windowing technique and used to obtain the refined depth map. Amit K. Agrawal, Rama Chellappa |
ICASSP (3) | 1 |
| 2004 | Robust ego-motion estimation and 3d model refinement using depth based parallax modelabstractWe present an iterative algorithm for robustly estimating the ego-motion and refining and updating a coarse, noisy and partial depth map using a depth based parallax model and brightness derivatives extracted from an image pair. Given a coarse, noisy and partial depth map acquired by a range-finder or obtained from a Digital Elevation Map (DFM), we first estimate the ego-motion by combining a global ego-motion constraint and a local brightness constancy constraint. Using the estimated camera motion and the available depth map estimate, motion of the 3D points is compensated. We utilize the fact that the resulting surface parallax field is an epipolar field and knowing its direction from the previous motion estimates, estimate its magnitude and use it to refine the depth map estimate. Instead of assuming a smooth parallax field or locally smooth depth models, we locally model the parallax magnitude using the depth map, formulate the problem as a generalized eigen-value analysis and obtain better results. In addition, confidence measures for depth estimates are provided which can be used to remove regions with potentially incorrect (and outliers in) depth estimates for robustly estimating ego-motion in the next iteration. Results on both synthetic and real examples are presented. Amit K. Agrawal, Rama Chellappa |
ICIP | 1 |
| 2002 | An Experimental Evaluation of Linear and Kernel-Based Methods for Face RecognitionabstractIn this paper we present the results of a comparative study of linear and kernel-based methods for face recognition. The methods used for dimensionality reduction are Principal Component Analysis (PCA), Kernel Principal Component Analysis (KPCA), Linear Discriminant Analysis (LDA) and Kernel Discriminant Analysis (KDA). The methods used for classification are Nearest Neighbor (NN) and Support Vector Machine (SVM). In addition, these classification methods are applied on raw images to gauge the performance of these dimensionality reduction techniques. All experiments have been performed on images from UMIST Face Database. Himaanshu Gupta, Amit K. Agrawal, Tarun Pruthi, Chandra Shekhar 0002, Rama Chellappa |
WACV | 2 |