VLDB 2026 Research / reviewers in the wild / expert
Sudipta N. Sinha
dblp:05/257 · also Sudipta Narayan Sinha
· DBLP profile ↗
53ranked-venue papers
12as first author
5since 2021 · last 2024
0000-0002-4186-3289ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 40 · 10 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
29 papers |
3D vision · 68% Robot navigation and mapping · 14% Segmentation and scene understanding · 4% | |
| Computer graphics and multimedia
13 papers |
Geometric modeling and processing · 32% Computational photography and imaging · 27% Image and video processing · 25% | |
| Network and information security
4 papers |
Privacy and data protection · 81% Security and privacy of machine learning · 19% | |
| Computer networks
2 papers |
Edge and fog computing · 57% Internet of things and sensor networks · 43% |
Topics — the 30 heaviest of 87, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
structure from motion |
1.1 | 6 | 2019 | Revealing Scenes by Inverting Structure From Motion Reconstructions · CVPR 2019 Flight Dynamics-Based Recovery of a UAV Trajectory Using Ground Cameras · CVPR 2017 Real-time image-based 6-DOF localization in large-scale environments · CVPR 2012 |
Computer vision › 3D vision
3d reconstruction |
1.0 | 7 | 2018 | Multiview Rectification of Folded Documents · IEEE Trans. Pattern Anal. Mach. Intell. 2018 Detecting and Reconstructing 3D Mirror Symmetric Objects · ECCV (2) 2012 Real-time image-based 6-DOF localization in large-scale environments · CVPR 2012 |
Computer vision › 3D vision
camera pose estimation |
0.8 | 2 | 2019 | Privacy Preserving Image Queries for Camera Localization · ICCV 2019 Privacy Preserving Image-Based Localization · CVPR 2019 |
Computer vision › 3D vision › correspondence estimation
dense correspondence |
0.8 | 2 | 2021 | PatchMatch-Based Neighborhood Consensus for Semantic Correspondence · CVPR 2021 Joint Recovery of Dense Correspondence and Cosegmentation in Two Images · CVPR 2016 |
Computer vision › 3D vision
visual localization |
0.7 | 2 | 2022 | Learning to Detect Scene Landmarks for Camera Localization · CVPR 2022 Real-time image-based 6-DOF localization in large-scale environments · CVPR 2012 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.6 | 3 | 2018 | Learning to Fuse Proposals from Multiple Scanline Optimizations in Semi-Global Matching · ECCV (13) 2018 Efficient High-Resolution Stereo Matching Using Local Plane Sweeps · CVPR 2014 Object stereo - Joint stereo matching and object segmentation · CVPR 2011 |
Geometric modeling and processing
3d reconstruction |
0.6 | 5 | 2017 | Robust Multiview Photometric Stereo Using Planar Mesh Parameterization · IEEE Trans. Pattern Anal. Mach. Intell. 2017 Piecewise planar stereo for image-based rendering · ICCV 2009 Interactive 3D architectural modeling from unordered photo collections · ACM Trans. Graph. 2008 |
Computer vision › 3D vision
pose estimation |
0.6 | 1 | 2022 | Learning to Detect Scene Landmarks for Camera Localization · CVPR 2022 |
Computer vision › 3D vision
object pose estimation |
0.5 | 2 | 2019 | Privacy Preserving Image Queries for Camera Localization · ICCV 2019 Real-time image-based 6-DOF localization in large-scale environments · CVPR 2012 |
Robotics › Robot navigation and mapping
visual odometry |
0.5 | 2 | 2017 | Fast Multi-frame Stereo Scene Flow with Motion Segmentation · CVPR 2017 Monocular Localization of a moving person onboard a Quadrotor MAV · ICRA 2015 |
Computer vision › 3D vision
neighborhood consensus |
0.5 | 1 | 2021 | PatchMatch-Based Neighborhood Consensus for Semantic Correspondence · CVPR 2021 |
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
0.5 | 1 | 2021 | PatchMatch-Based Neighborhood Consensus for Semantic Correspondence · CVPR 2021 |
Computer vision › 3D vision
3d human pose estimation |
0.4 | 1 | 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion Capture · CVPR 2020 |
Computer vision › 3D vision › 3d human pose estimation
monocular 3d pose estimation |
0.4 | 1 | 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion Capture · CVPR 2020 |
Robotics › Robot navigation and mapping
view planning |
0.4 | 1 | 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion Capture · CVPR 2020 |
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation |
0.4 | 2 | 2016 | Joint Recovery of Dense Correspondence and Cosegmentation in Two Images · CVPR 2016 Multiple View Object Cosegmentation Using Appearance and Stereo Cues · ECCV (5) 2012 |
Robotics › Robot navigation and mapping › localization
vision-based localization |
0.4 | 1 | 2019 | Privacy Preserving Image-Based Localization · CVPR 2019 |
Security and privacy of machine learning
privacy attack |
0.4 | 1 | 2019 | Revealing Scenes by Inverting Structure From Motion Reconstructions · CVPR 2019 |
Privacy and data protection
privacy-preserving computation |
0.4 | 1 | 2019 | Privacy Preserving Image-Based Localization · CVPR 2019 |
Privacy and data protection
privacy-preserving data analysis |
0.4 | 1 | 2019 | Privacy Preserving Image Queries for Camera Localization · ICCV 2019 |
Privacy and data protection › location privacy
privacy-preserving localization |
0.4 | 1 | 2019 | Privacy Preserving Image-Based Localization · CVPR 2019 |
Computer vision › 3D vision › structure from motion
bundle adjustment |
0.3 | 2 | 2017 | Flight Dynamics-Based Recovery of a UAV Trajectory Using Ground Cameras · CVPR 2017 Discovering and exploiting 3D symmetries in structure from motion · CVPR 2012 |
Computer vision › 3D vision › object pose estimation
6d object pose estimation |
0.3 | 1 | 2018 | Real-Time Seamless Single Shot 6D Object Pose Prediction · CVPR 2018 |
Machine learning › Reinforcement learning
exploration |
0.3 | 1 | 2018 | Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility · ECCV (15) 2018 |
Robotics › Robot navigation and mapping › view planning
next-best-view planning |
0.3 | 1 | 2018 | Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility · ECCV (15) 2018 |
Machine learning › Reinforcement learning › exploration
scene exploration |
0.3 | 1 | 2018 | Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility · ECCV (15) 2018 |
Computer vision › 3D vision › stereo vision › stereo matching
semi-global matching |
0.3 | 1 | 2018 | Learning to Fuse Proposals from Multiple Scanline Optimizations in Semi-Global Matching · ECCV (13) 2018 |
Image and video processing
document image analysis |
0.3 | 1 | 2018 | Multiview Rectification of Folded Documents · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Image and video processing › image restoration › document image restoration
document image rectification |
0.3 | 1 | 2018 | Multiview Rectification of Folded Documents · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Robotics › Legged, aerial and field robots
aerial robots |
0.3 | 1 | 2017 | Submodular Trajectory Optimization for Aerial 3D Scanning · ICCV 2017 |
Methods — techniques the papers use, named apart from their topics
subspace-to-subspace distance · 1.0localized 3d reconstruction · 1.0graph-based subgraph identification · 1.0affine subspace embedding · 1.0line-based feature representation · 0.8feature point concealment · 0.8scene landmark detection · 0.6convolutional neural network · 0.6patchmatch · 0.5end-to-end training · 0.54d scoring function · 0.5uncertainty estimation · 0.4motion planning · 0.4deep learning · 0.4cascaded u-net · 0.4SIFT descriptors · 0.43d line cloud representation · 0.4ridge-aware 3d reconstruction · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improved Scene Landmark Detection for Camera LocalizationabstractCamera localization methods based on retrieval, local feature matching, and 3D structure-based pose estimation are accurate but require high storage, are slow, and are not privacy-preserving. A method based on scene landmark detection (SLD) was recently proposed to address these limitations. It involves training a convolutional neural network (CNN) to detect a few predetermined, salient, scene-specific 3D points or landmarks and computing camera pose from the associated 2D–3D correspondences. Although SLD outperformed existing learning-based approaches, it was notably less accurate than 3D structure-based methods. In this paper, we show that the accuracy gap was due to insufficient model capacity and noisy labels during training. To mitigate the capacity issue, we propose to split the landmarks into subgroups and train a separate network for each subgroup. To generate better training labels, we propose using dense reconstructions to estimate visibility of scene landmarks. Finally, we present a compact architecture to improve memory efficiency. Accuracy wise, our approach is on par with state of the art structurebased methods on the INDOOR-6 dataset but runs significantly faster and uses less storage. Code and models can be found at https://github.com/microsoft/SceneLandmarkLocalization. Tien Do, Sudipta N. Sinha |
3DV | 2 |
| 2022 | Learning to Detect Scene Landmarks for Camera LocalizationabstractModern camera localization methods that use image retrieval, feature matching, and 3D structure-based pose estimation require long-term storage of numerous scene images or a vast amount of image features. This can make them unsuitable for resource constrained VR/AR devices and also raises serious privacy concerns. We present a new learned camera localization technique that eliminates the need to store features or a detailed 3D point cloud. Our key idea is to implicitly encode the appearance of a sparse yet salient set of 3D scene points into a convolutional neural network (CNN) that can detect these scene points in query images whenever they are visible. We refer to these points as scene landmarks. We also show that a CNN can be trained to regress bearing vectors for such landmarks even when they are not within the camera's field-of-view. We demonstrate that the predicted landmarks yield accurate pose estimates and that our method outperforms DSAC*, the state-of-the-art in learned localization. Furthermore, extending HLoc (an accurate method) by combining its correspondences with our predictions boosts its accuracy even further. Tien Do, Ondrej Miksik, Joseph DeGol, Hyun Soo Park, Sudipta N. Sinha |
CVPR | 5 |
| 2021 | Privacy-Preserving Image Features via Adversarial Affine Subspace EmbeddingsabstractMany computer vision systems require users to upload image features to the cloud for processing and storage. These features can be exploited to recover sensitive information about the scene or subjects, e.g., by reconstructing the appearance of the original image. To address this privacy concern, we propose a new privacy-preserving feature representation. The core idea of our work is to drop constraints from each feature descriptor by embedding it within an affine subspace containing the original feature as well as adversarial feature samples. Feature matching on the privacy-preserving representation is enabled based on the notion of subspace-to-subspace distance. We experimentally demonstrate the effectiveness of our method and its high practical relevance for the applications of visual localization and mapping as well as face authentication. Compared to the original features, our approach makes it significantly more difficult for an adversary to recover private information. Mihai Dusmanu, Johannes L. Schönberger, Sudipta N. Sinha, Marc Pollefeys |
CVPR | 3 |
| 2021 | PatchMatch-Based Neighborhood Consensus for Semantic CorrespondenceabstractWe address estimating dense correspondences between two images depicting different but semantically related scenes. End-to-end trainable deep neural networks incorporating neighborhood consensus cues are currently the best methods for this task. However, these architectures require exhaustive matching and 4D convolutions over matching costs for all pairs of feature map pixels. This makes them computationally expensive. We present a more efficient neighborhood consensus approach based on PatchMatch. For higher accuracy, we propose to use a learned local 4D scoring function for evaluating candidates during the PatchMatch iterations. We have devised an approach to jointly train the scoring function and the feature extraction modules by embedding them into a proxy model which is end-to-end differentiable. The modules are trained in a supervised setting using a cross-entropy loss to directly incorporate sparse keypoint supervision. Our evaluation on PF-Pascal and SPair-71K shows that our method significantly outperforms the state-of-the-art on both datasets while also being faster and using less memory. Jae Yong Lee 0006, Joseph DeGol, Victor Fragoso, Sudipta N. Sinha |
CVPR | 4 |
| 2021 | Visage: enabling timely analytics for drone imageryabstractAnalytics with three-dimensional imagery from drones are driving the next generation of remote monitoring applications. Today, there is an unmet need in providing such analytics in an interactive manner, especially over weak Internet connections, to quickly diagnose and solve problems in the commercial industry space of monitoring assets using drones in remote parts of the world. Existing mechanisms either compromise on the quality of insights by not building 3D images and analyze individual 2D images in isolation, or spend tens of minutes building a 3D image before obtaining and uploading insights. We present Visage, a system that accelerates 3D image analytics by identifying smaller parts of the data that can actually benefit from 3D analytics and prioritizing building, and uploading the localized 3D images for those parts. To achieve this, Visage uses a graph to represent raw 2D images and their relative content overlap, and then identifies the various subgraphs using application knowledge that are good candidates for localized 3D image based insights. We evaluate Visage using data from multiple real deployments and show that it can reduce analytics-latency by up to four orders of magnitude. Sagar Jha, Youjie Li, Shadi A. Noghabi, Vaishnavi Nattar Ranganathan, Peeyush Kumar, Michael Toelle, Sudipta N. Sinha, Ranveer Chandra, Anirudh Badam |
MobiCom | 8 |
| 2020 | Generalized Pose-and-Scale Estimation using 4-Point Congruence ConstraintsabstractWe present gP4Pc, a new method for computing the absolute pose of a generalized camera with unknown internal scale from four corresponding 3D point-and-ray pairs. Unlike most pose-and-scale methods, gP4Pc is based on constraints arising from the congruence of shapes defined by two sets of four points related by an unknown similarity transformation. By choosing a novel parametrization for the problem, we derive a system of four quadratic equations in four scalar variables. The variables represent the distances of 3D points along the rays from the camera centers. After solving this system via Gröbner basis-based automatic polynomial solvers, we compute the similarity transformation using an efficient 3D point-point alignment method. We also propose a specialized variant of our solver for the case of coplanar points, which is computationally very efficient and about 3× faster than the fastest existing solver. Our experiments on real and synthetic datasets, demonstrate that gP4Pc is among the fastest methods in terms of total running time when used within a RANSAC framework, while achieving competitive numerical stability, accuracy, and robustness to noise. Victor Fragoso, Sudipta N. Sinha |
3DV | 2 |
| 2020 | Depth Completion Using a View-constrained Deep PriorabstractRecent work has shown that the structure of convolutional neural networks (CNNs) induces a strong prior that favors natural images. This prior, known as a deep image prior (DIP), is an effective regularizer in inverse problems such as image denoising and inpainting. We extend the concept of the DIP to depth images. Given color images and noisy and incomplete target depth maps, we optimize a randomly-initialized CNN model to reconstruct a depth map restored by virtue of using the CNN network structure as a prior combined with a view-constrained photo-consistency loss. This loss is computed using images from a geometrically calibrated camera from nearby viewpoints. We apply this deep depth prior for inpainting and refining incomplete and noisy depth maps within both binocular and multi-view stereo pipelines. Our quantitative and qualitative evaluation shows that our refined depth maps are more accurate and complete, and after fusion, produces dense 3D models of higher quality. Pallabi Ghosh, Vibhav Vineet, Larry Davis 0001, Abhinav Shrivastava, Sudipta N. Sinha, Neel Joshi |
3DV | 5 |
| 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion CaptureabstractThe accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the location which will yield the highest accuracy remains an open problem. This is the problem that we address in this paper. Specifically, given a short video sequence, we introduce an algorithm that predicts which viewpoints should be chosen to capture future frames so as to maximize 3D human pose estimation accuracy. The key idea underlying our approach is a method to estimate the uncertainty of the 3D body pose estimates. We integrate several sources of uncertainty, originating from deep learning based regressors and temporal smoothness. Our motion planner yields improved 3D body pose estimates and outperforms or matches existing ones that are based on person following and orbiting. Sena Kiciroglu, Helge Rhodin, Sudipta N. Sinha, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2019 | Revealing Scenes by Inverting Structure From Motion ReconstructionsabstractMany 3D vision systems localize cameras within a scene using 3D point clouds. Such point clouds are often obtained using structure from motion (SfM), after which the images are discarded to preserve privacy. In this paper, we show, for the first time, that such point clouds retain enough information to reveal scene appearance and compromise privacy. We present a privacy attack that reconstructs color images of the scene from the point cloud. Our method is based on a cascaded U-Net that takes as input, a 2D multichannel image of the points rendered from a specific viewpoint containing point depth and optionally color and SIFT descriptors and outputs a color image of the scene from that viewpoint. Unlike previous feature inversion methods, we deal with highly sparse and irregular 2D point distributions and inputs where many point attributes are missing, namely keypoint orientation and scale, the descriptor image source and the 3D point visibility. We evaluate our attack algorithm on public datasets and analyze the significance of the point cloud attributes. Finally, we show that novel views can also be generated thereby enabling compelling virtual tours of the underlying scene. Francesco Pittaluga, Sanjeev J. Koppal, Sing Bing Kang, Sudipta N. Sinha |
CVPR | 4 |
| 2019 | Privacy Preserving Image-Based LocalizationabstractImage-based localization is a core component of many augmented/mixed reality (AR/MR) and autonomous robotic systems. Current localization systems rely on the persistent storage of 3D point clouds of the scene to enable camera pose estimation, but such data reveals potentially sensitive scene information. This gives rise to significant privacy risks, especially as for many applications 3D mapping is a background process that the user might not be fully aware of. We pose the following question: How can we avoid disclosing confidential information about the captured 3D scene, and yet allow reliable camera pose estimation? This paper proposes the first solution to what we call privacy preserving image-based localization. The key idea of our approach is to lift the map representation from a 3D point cloud to a 3D line cloud. This novel representation obfuscates the underlying scene geometry while providing sufficient geometric constraints to enable robust and accurate 6-DOF camera pose estimation. Extensive experiments on several datasets and localization scenarios underline the high practical relevance of our proposed approach. Pablo Speciale, Johannes L. Schönberger, Sing Bing Kang, Sudipta N. Sinha, Marc Pollefeys |
CVPR | 4 |
| 2019 | Low-cost aerial imaging for small holder farmersabstractRecent work in networked systems has shown that using aerial imagery for farm monitoring can enable precision agriculture by lowering the cost and reducing the overhead of large scale sensor deployment. However, acquiring aerial imagery requires a drone, which has high capital and operational costs, often beyond the reach of farmers in the developing world. In this paper, we present TYE (Tethered eYE), an inexpensive platform for aerial imagery. It consists of a tethered helium balloon with a custom mount that can hold a smartphone (or a camera) with a battery pack. The balloon can be carried using a tether by a person or a vehicle. We incorporate various techniques to increase the operational time of the system, and to provide actionable insights even with unstable imagery. We develop path-planning algorithms and use that to develop an interactive mobile phone application that provides the user instant feedback to guide users to efficiently traverse large areas of land. We use computer vision algorithms to stitch orthomosaics by effectively countering wind-induced motion of the camera. We have used TYE for aerial imaging of agricultural land for over a year, and envision it as a low-cost aerial imaging platform for similar applications. Zerina Kapetanovic, Akshit Kumar, Vasuki Narasimha Swamy, Rohit Patil, Deepak Vasisht, Rahul Sharma 0001, S. Manohar 0001, Ranveer Chandra, Anirudh Badam, Gireeja Ranade, Sudipta N. Sinha, Akshay Uttama Nambi |
COMPASS | 12 |
| 2019 | Privacy Preserving Image Queries for Camera LocalizationabstractAugmented/mixed reality and robotic applications are increasingly relying on cloud-based localization services, which require users to upload query images to perform camera pose estimation on a server. This raises significant privacy concerns when consumers use such services in their homes or in confidential industrial settings. Even if only image features are uploaded, the privacy concerns remain as the images can be reconstructed fairly well from feature locations and descriptors. We propose to conceal the content of the query images from an adversary on the server or a man-in-the-middle intruder. The key insight is to replace the 2D image feature points in the query image with randomly oriented 2D lines passing through their original 2D positions. It will be shown that this feature representation hides the image contents, and thereby protects user privacy, yet still provides sufficient geometric constraints to enable robust and accurate 6-DOF camera pose estimation from feature correspondences. Our proposed method can handle single- and multi-image queries as well as exploit additional information about known structure, gravity, and scale. Numerous experiments demonstrate the high practical relevance of our approach. Pablo Speciale, Johannes L. Schönberger, Sudipta N. Sinha, Marc Pollefeys |
ICCV | 3 |
| 2019 | Photorealistic Image Synthesis for Object Instance DetectionabstractWe present an approach to synthesize highly photorealistic images of 3D object models, which we use to train a convolutional neural network for detecting the objects in real images. The proposed approach has three key ingredients: (1) 3D object models are rendered in 3D models of complete scenes with realistic materials and lighting, (2) plausible geometric configuration of objects and cameras in a scene is generated using physics simulation, and (3) high photorealism of the synthesized images is achieved by physically based rendering. When trained on images synthesized by the proposed approach, the Faster R-CNN object detector [1] achieves a 24% absolute improvement of [email protected] on Rutgers APC [2] and 11% on LineMod-Occluded [3] datasets, compared to a baseline where the training images are synthesized by rendering object models on top of random photographs. This work is a step towards being able to effectively train object detectors without capturing or annotating any real images. A dataset of 400K synthetic images with ground truth annotations for various computer vision tasks will be released on the project website: thodan.github.io/objectsynth. Tomas Hodan, Vibhav Vineet, Ran Gal, Emanuel Shalev, Jon Hanzelka, Treb Connell, Pedro Urbina, Sudipta N. Sinha, Brian Guenter |
ICIP | 8 |
| 2018 | Learning to Align Images Using Weak Geometric SupervisionabstractImage alignment tasks require accurate pixel correspondences, which are usually recovered by matching local feature descriptors. Such descriptors are often derived using supervised learning on existing datasets with ground truth correspondences. However, the cost of creating such datasets is usually prohibitive. In this paper, we propose a new approach to align two images related by an unknown 2D homography where the local descriptor is learned from scratch from the images and the homography is estimated simultaneously. Our key insight is that a siamese convolutional neural network can be trained jointly while iteratively updating the homography parameters by optimizing a single loss function. Our method is currently weakly supervised because the input images need to be roughly aligned. We have used this method to align images of different modalities such as RGB and near-infra-red (NIR) without using any prior labeled data. Images automatically aligned by our method were then used to train descriptors that generalize to new images. We also evaluated our method on RGB images. On the HPatches benchmark, our method achieves comparable accuracy to deep local descriptors that were trained offline in a supervised setting. Jing Dong 0002, Byron Boots, Frank Dellaert, Ranveer Chandra, Sudipta N. Sinha |
3DV | 5 |
| 2018 | Real-Time Seamless Single Shot 6D Object Pose PredictionabstractWe propose a single-shot approach for simultaneously detecting an object in an RGB image and predicting its 6D pose without requiring multiple stages or having to examine multiple hypotheses. Unlike a recently proposed single-shot technique for this task [10] that only predicts an approximate 6D pose that must then be refined, ours is accurate enough not to require additional post-processing. As a result, it is much faster - 50 fps on a Titan X (Pascal) GPU - and more suitable for real-time processing. The key component of our method is a new CNN architecture inspired by [27, 28] that directly predicts the 2D image locations of the projected vertices of the object's 3D bounding box. The object's 6D pose is then estimated using a PnP algorithm. For single object and multiple object pose estimation on the LINEMOD and OCCLUSION datasets, our approach substantially outperforms other recent CNN-based approaches [10, 25] when they are all used without postprocessing. During post-processing, a pose refinement step can be used to boost the accuracy of these two methods, but at 10 fps or less, they are much slower than our method. Bugra Tekin, Sudipta N. Sinha, Pascal Fua |
CVPR | 2 |
| 2018 | Learn-to-Score: Efficient 3D Scene Exploration by Predicting View Utility
Benjamin Hepp, Debadeepta Dey, Sudipta N. Sinha, Ashish Kapoor, Neel Joshi, Otmar Hilliges |
ECCV (15) | 3 |
| 2018 | Learning to Fuse Proposals from Multiple Scanline Optimizations in Semi-Global Matching
Johannes L. Schönberger, Sudipta N. Sinha, Marc Pollefeys |
ECCV (13) | 2 |
| 2018 | Multiview Rectification of Folded DocumentsabstractDigitally unwrapping images of paper sheets is crucial for accurate document scanning and text recognition. This paper presents a method for automatically rectifying curved or folded paper sheets from a few images captured from multiple viewpoints. Prior methods either need expensive 3D scanners or model deformable surfaces using over-simplified parametric representations. In contrast, our method uses regular images and is based on general developable surface models that can represent a wide variety of paper deformations. Our main contribution is a new robust rectification method based on ridge-aware 3D reconstruction of a paper sheet and unwrapping the reconstructed surface using properties of developable surfaces via conformal mapping. We present results on several examples including book pages, folded letters and shopping receipts. Shaodi You, Yasuyuki Matsushita, Sudipta N. Sinha, Yusuke Bou, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Semi-global Stereo Matching with Surface Orientation PriorsabstractSemi-Global Matching (SGM) is a widely-used efficient stereo matching technique. It works well for textured scenes, but fails on untextured slanted surfaces due to its fronto-parallel smoothness assumption. To remedy this problem, we propose a simple extension, termed SGM-P, to utilize precomputed surface orientation priors. Such priors favor different surface slants in different 2D image regions or 3D scene regions and can be derived in various ways. In this paper we evaluate plane orientation priors derived from stereo matching at a coarser resolution and show that such priors can yield significant performance gains for difficult weakly-textured scenes. We also explore surface normal priors derived from Manhattan-world assumptions, and we analyze the potential performance gains using oracle priors derived from ground-truth data. SGM-P only adds a minor computational overhead to SGM and is an attractive alternative to more complex methods employing higher-order smoothness terms. Daniel Scharstein, Tatsunori Taniai, Sudipta N. Sinha |
3DV | 3 |
| 2017 | Flight Dynamics-Based Recovery of a UAV Trajectory Using Ground CamerasabstractWe propose a new method to estimate the 6-dof trajectory of a flying object such as a quadrotor UAV within a 3D airspace monitored using multiple fixed ground cameras. It is based on a new structure from motion formulation for the 3D reconstruction of a single moving point with known motion dynamics. Our main contribution is a new bundle adjustment procedure, which in addition to optimizing the camera poses, regularizes the point trajectory using a prior based on motion dynamics (or specifically flight dynamics). Furthermore, we can infer the underlying control input sent to the UAVs autopilot that determined its flight trajectory. Our method requires neither perfect single-view tracking nor appearance matching across views. For robustness, we allow the tracker to generate multiple detections per frame in each video. The true detections and the data association across videos is estimated using robust multi-view triangulation and subsequently refined in our bundle adjustment formulation. Quantitative evaluation on simulated data and experiments on real videos from indoor and outdoor scenes shows that our technique is superior to existing methods. Artem Rozantsev, Sudipta N. Sinha, Debadeepta Dey, Pascal Fua |
CVPR | 2 |
| 2017 | Fast Multi-frame Stereo Scene Flow with Motion SegmentationabstractWe propose a new multi-frame method for efficiently computing scene flow (dense depth and optical flow) and camera ego-motion for a dynamic scene observed from a moving stereo camera rig. Our technique also segments out moving objects from the rigid scene. In our method, we first estimate the disparity map and the 6-DOF camera motion using stereo matching and visual odometry. We then identify regions inconsistent with the estimated camera motion and compute per-pixel optical flow only at these regions. This flow proposal is fused with the camera motion-based flow proposal using fusion moves to obtain the final optical flow and motion segmentation. This unified framework benefits all four tasks - stereo, optical flow, visual odometry and motion segmentation leading to overall higher accuracy and efficiency. Our method is currently ranked third on the KITTI 2015 scene flow benchmark. Furthermore, our CPU implementation runs in 2-3 seconds per frame which is 1-3 orders of magnitude faster than the top six methods. We also report a thorough evaluation on challenging Sintel sequences with fast camera and object motion, where our method consistently outperforms OSF [30], which is currently ranked second on the KITTI benchmark. Tatsunori Taniai, Sudipta N. Sinha, Yoichi Sato 0001 |
CVPR | 2 |
| 2017 | Submodular Trajectory Optimization for Aerial 3D ScanningabstractDrones equipped with cameras are emerging as a powerful tool for large-scale aerial 3D scanning, but existing automatic flight planners do not exploit all available information about the scene, and can therefore produce inaccurate and incomplete 3D models. We present an automatic method to generate drone trajectories, such that the imagery acquired during the flight will later produce a high-fidelity 3D model. Our method uses a coarse estimate of the scene geometry to plan camera trajectories that: (1) cover the scene as thoroughly as possible; (2) encourage observations of scene geometry from a diverse set of viewing angles; (3) avoid obstacles; and (4) respect a user-specified flight time budget. Our method relies on a mathematical model of scene coverage that exhibits an intuitive diminishing returns property known as submodularity. We leverage this property extensively to design a trajectory planning algorithm that reasons globally about the non-additive coverage reward obtained across a trajectory, jointly with the cost of traveling between views. We evaluate our method by using it to scan three large outdoor scenes, and we perform a quantitative evaluation using a photorealistic video game simulator. Mike Roberts 0001, Shital Shah, Debadeepta Dey, Anh Truong, Sudipta N. Sinha, Ashish Kapoor, Pat Hanrahan, Neel Joshi |
ICCV | 5 |
| 2017 | FarmBeats: An IoT Platform for Data-Driven Agriculture
Deepak Vasisht, Zerina Kapetanovic, Jongho Won, Xinxin Jin, Ranveer Chandra, Sudipta N. Sinha, Ashish Kapoor, Madhusudhan Sudarshan, Sean Stratman |
NSDI | 6 |
| 2017 | Robust Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a robust uncalibrated multiview photometric stereo method for high quality 3D shape reconstruction. In our method, a coarse initial 3D mesh obtained using a multiview stereo method is projected onto a 2D planar domain using a planar mesh parameterization technique. We describe methods for surface normal estimation that work in the parameterized 2D space that jointly incorporates all geometric and photometric cues from multiple viewpoints. Using an estimated surface normal map, a refined 3D mesh is then recovered by computing an optimal displacement map in the same 2D planar domain. Our method avoids the need of merging view-dependent surface normal maps that is often required in conventional methods. We conduct evaluation on various real-world objects containing surfaces with specular reflections, multiple albedos, and complex topologies in both controlled and uncontrolled settings and demonstrate that accurate 3D meshes with fine geometric details can be recovered by our method. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Efficient and Robust Color Consistency for Community Photo CollectionsabstractWe present an efficient technique to optimize color consistency of a collection of images depicting a common scene. Our method first recovers sparse pixel correspondences in the input images and stacks them into a matrix with many missing entries. We show that this matrix satisfies a rank two constraint under a simple color correction model. These parameters can be viewed as pseudo white balance and gamma correction parameters for each input image. We present a robust low-rank matrix factorization method to estimate the unknown parameters of this model. Using them, we improve color consistency of the input images or perform color transfer with any input image as the source. Our approach is insensitive to outliers in the pixel correspondences thereby precluding the need for complex pre-processing steps. We demonstrate high quality color consistency results on large photo collections of popular tourist landmarks and personal photo collections containing images of people. Jaesik Park, Yu-Wing Tai, Sudipta N. Sinha, In-So Kweon |
CVPR | 3 |
| 2016 | Joint Recovery of Dense Correspondence and Cosegmentation in Two ImagesabstractWe propose a new technique to jointly recover cosegmentation and dense per-pixel correspondence in two images. Our method parameterizes the correspondence field using piecewise similarity transformations and recovers a mapping between the estimated common "foreground" regions in the two images allowing them to be precisely aligned. Our formulation is based on a hierarchical Markov random field model with segmentation and transformation labels. The hierarchical structure uses nested image regions to constrain inference across multiple scales. Unlike prior hierarchical methods which assume that the structure is given, our proposed iterative technique dynamically recovers the structure along with the labeling. This joint inference is performed in an energy minimization framework using iterated graph cuts. We evaluate our method on a new dataset of 400 image pairs with manually obtained ground truth, where it outperforms state-of-the-art methods designed specifically for either cosegmentation or correspondence estimation. Tatsunori Taniai, Sudipta N. Sinha, Yoichi Sato 0001 |
CVPR | 2 |
| 2015 | Monocular Localization of a moving person onboard a Quadrotor MAVabstractIn this paper, we propose a novel method to recover the 3D trajectory of a moving person from a monocular camera mounted on a quadrotor micro aerial vehicle (MAV). The key contribution is an integrated approach that simultaneously performs visual odometry (VO) and persistent tracking of a person automatically detected in the scene. All computation pertaining to VO, detection and tracking runs onboard the MAV from a front-facing monocular RGB camera. Given the gravity direction from an inertial sensor and the knowledge of the individual's height, a complete 3D trajectory of the person within the reconstructed scene can be estimated. When the ground plane is detected from the triangulated 3D points, the absolute metric scale of the trajectory and the 3D map is also recovered. Our extensive indoor and outdoor experiments show that the system can localize a person moving naturally within a large area. The system runs at 17 frames per second on the onboard computer. A walking person was successfully tracked for two minutes and an accurate trajectory was recovered over a distance of 140 meters with our system running onboard. Hyon Lim, Sudipta N. Sinha |
ICRA | 2 |
| 2014 | Calibrating a Non-isotropic Near Point Light Source Using a PlaneabstractWe show that a non-isotropic near point light source rigidly attached to a camera can be calibrated using multiple images of a weakly textured planar scene. We prove that if the radiant intensity distribution (RID) of a light source is radially symmetric with respect to its dominant direction, then the shading observed on a Lambertian scene plane is bilaterally symmetric with respect to a 2D line on the plane. The symmetry axis detected in an image provides a linear constraint for estimating the dominant light axis. The light position and RID parameters can then be estimated using a linear method. Specular highlights if available can also be used for light position estimation. We also extend our method to handle non-Lambertian reflectances which we model using a biquadratic BRDF. We have evaluated our method on synthetic data quantitavely. Our experiments on real scenes show that our method works well in practice and enables light calibration without the need of a specialized hardware. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 2 |
| 2014 | Efficient High-Resolution Stereo Matching Using Local Plane SweepsabstractWe present a stereo algorithm designed for speed and efficiency that uses local slanted plane sweeps to propose disparity hypotheses for a semi-global matching algorithm. Our local plane hypotheses are derived from initial sparse feature correspondences followed by an iterative clustering step. Local plane sweeps are then performed around each slanted plane to produce out-of-plane parallax and matching-cost estimates. A final global optimization stage, implemented using semi-global matching, assigns each pixel to one of the local plane hypotheses. By only exploring a small fraction of the whole disparity space volume, our technique achieves significant speedups over previous algorithms and achieves state-of-the-art accuracy on high-resolution stereo pairs of up to 19 megapixels. Sudipta N. Sinha, Daniel Scharstein, Richard Szeliski |
CVPR | 1 |
| 2014 | Car make and model recognition using 3D curve alignmentabstractWe present a new approach for recognizing the make and model of a car from a single image. While most previous methods are restricted to fixed or limited viewpoints, our system is able to verify a car's make and model from an arbitrary view. Our model consists of 3D space curves obtained by backprojecting image curves onto silhouette-based visual hulls and then refining them using three-view curve matching. We also build an appearance model of taillights which is used as an additional cue. Our approach is able to verify the exact make and model of a car over a wide range of viewpoints and background clutter. Edward Hsiao, Sudipta N. Sinha, Krishnan Ramnath, Simon Baker, C. Lawrence Zitnick, Richard Szeliski |
WACV | 2 |
| 2014 | AutoCaption: Automatic caption generation for personal photosabstractAutoCaption is a system that helps a smartphone user generate a caption for their photos. It operates by uploading the photo to a cloud service where a number of parallel modules are applied to recognize a variety of entities and relations. The outputs of the modules are combined to generate a large set of candidate captions, which are returned to the phone. The phone client includes a convenient user interface that allows users to select their favorite caption, reorder, add, or delete words to obtain the grammatical style they prefer. The user can also select from multiple candidates returned by the recognition modules. Krishnan Ramnath, Simon Baker, Lucy Vanderwende, Motaz Ahmad El-Saban, Sudipta N. Sinha, Anitha Kannan, Noran Hassan, Michel Galley, Yi Yang 0007, Deva Ramanan, Alessandro Bergamo, Lorenzo Torresani |
WACV | 5 |
| 2014 | Car make and model recognition using 3D curve alignmentabstractWe present a new approach for recognizing the make and model of a car from a single image. While most previous methods are restricted to fixed or limited viewpoints, our system is able to verify a car's make and model from an arbitrary view. Our model consists of 3D space curves obtained by backprojecting image curves onto silhouette-based visual hulls and then refining them using three-view curve matching. These 3D curves are then matched to 2D image curves using a 3D view-based alignment technique. We present two different methods for estimating the pose of a car, which we then use to initialize the 3D curve matching. Our approach is able to verify the exact make and model of a car over a wide range of viewpoints in cluttered scenes. Krishnan Ramnath, Sudipta N. Sinha, Richard Szeliski, Edward Hsiao |
WACV | 2 |
| 2013 | Leveraging Structure from Motion to Learn Discriminative Codebooks for Scalable Landmark ClassificationabstractIn this paper we propose a new technique for learning a discriminative codebook for local feature descriptors, specifically designed for scalable landmark classification. The key contribution lies in exploiting the knowledge of correspondences within sets of feature descriptors during code-book learning. Feature correspondences are obtained using structure from motion (SfM) computation on Internet photo collections which serve as the training data. Our codebook is defined by a random forest that is trained to map corresponding feature descriptors into identical codes. Unlike prior forest-based codebook learning methods, we utilize fine-grained descriptor labels and address the challenge of training a forest with an extremely large number of labels. Our codebook is used with various existing feature encoding schemes and also a variant we propose for importance-weighted aggregation of local features. We evaluate our approach on a public dataset of 25 landmarks and our new dataset of 620 landmarks (614K images). Our approach significantly outperforms the state of the art in landmark classification. Furthermore, our method is memory efficient and scalable. Alessandro Bergamo, Sudipta N. Sinha, Lorenzo Torresani |
CVPR | 2 |
| 2013 | Multiview Photometric Stereo Using Planar Mesh ParameterizationabstractWe propose a method for accurate 3D shape reconstruction using uncalibrated multiview photometric stereo. A coarse mesh reconstructed using multiview stereo is first parameterized using a planar mesh parameterization technique. Subsequently, multiview photometric stereo is performed in the 2D parameter domain of the mesh, where all geometric and photometric cues from multiple images can be treated uniformly. Unlike traditional methods, there is no need for merging view-dependent surface normal maps. Our key contribution is a new photometric stereo based mesh refinement technique that can efficiently reconstruct meshes with extremely fine geometric details by directly estimating a displacement texture map in the 2D parameter domain. We demonstrate that intricate surface geometry can be reconstructed using several challenging datasets containing surfaces with specular reflections, multiple albedos and complex topologies. Jaesik Park, Sudipta N. Sinha, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
ICCV | 2 |
| 2012 | Discovering and exploiting 3D symmetries in structure from motionabstractMany architectural scenes contain symmetric or repeated structures, which can generate erroneous image correspondences during structure from motion (Sfm) computation. Prior work has shown that the detection and removal of these incorrect matches is crucial for accurate and robust recovery of scene structure. In this paper, we point out that these incorrect matches, in fact, provide strong cues to the existence of symmetries and structural regularities in the unknown 3D structure. We make two key contributions. First, we propose a method to recover various symmetry relations in the structure using geometric and appearance cues. A set of structural constraints derived from the symmetries are imposed within a new constrained bundle adjustment formulation, where symmetry priors are also incorporated. Second, we show that the recovered symmetries enable us to choose a natural coordinate system for the 3D structure where gauge freedom in rotation is held fixed. Furthermore, based on the symmetries, 3D structure completion is also performed. Our approach significantly reduces drift through ”structural” loop closures and improves the accuracy of reconstructions in urban scenes. Andrea Cohen, Christopher Zach, Sudipta N. Sinha, Marc Pollefeys |
CVPR | 3 |
| 2012 | Real-time image-based 6-DOF localization in large-scale environmentsabstractWe present a real-time approach for image-based localization within large scenes that have been reconstructed offline using structure from motion (Sfm). From monocular video, our method continuously computes a precise 6-DOF camera pose, by efficiently tracking natural features and matching them to 3D points in the Sfm point cloud. Our main contribution lies in efficiently interleaving a fast keypoint tracker that uses inexpensive binary feature descriptors with a new approach for direct 2D-to-3D matching. The 2D-to-3D matching avoids the need for online extraction of scale-invariant features. Instead, offline we construct an indexed database containing multiple DAISY descriptors per 3D point extracted at multiple scales. The key to the efficiency of our method lies in invoking DAISY descriptor extraction and matching sparingly during localization, and in distributing this computation over a window of successive frames. This enables the algorithm to run in real-time, without fluctuations in the latency over long durations. We evaluate the method in large indoor and outdoor scenes. Our algorithm runs at over 30 Hz on a laptop and at 12 Hz on a low-power, mobile computer suitable for onboard computation on a quadrotor micro aerial vehicle. Hyon Lim, Sudipta N. Sinha, Michael F. Cohen, Matthew Uyttendaele |
CVPR | 2 |
| 2012 | Multiple View Object Cosegmentation Using Appearance and Stereo Cues
Adarsh Kowdle, Sudipta N. Sinha, Richard Szeliski |
ECCV (5) | 2 |
| 2012 | Detecting and Reconstructing 3D Mirror Symmetric Objects
Sudipta N. Sinha, Krishnan Ramnath, Richard Szeliski |
ECCV (2) | 1 |
| 2012 | Image-based rendering for scenes with reflectionsabstractWe present a system for image-based modeling and rendering of real-world scenes containing reflective and glossy surfaces. Previous approaches to image-based rendering assume that the scene can be approximated by 3D proxies that enable view interpolation using traditional back-to-front or z-buffer compositing. In this work, we show how these can be generalized to multiple layers that are combined in an additive fashion to model the reflection and transmission of light that occurs at specular surfaces such as glass and glossy materials. To simplify the analysis and rendering stages, we model the world using piecewise-planar layers combined using both additive and opaque mixing of light. We also introduce novel techniques for estimating multiple depths in the scene and separating the reflection and transmission components into different layers. We then use our system to model and render a variety of real-world scenes with reflections. Sudipta N. Sinha, Johannes Kopf 0001, Michael Goesele, Daniel Scharstein, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2011 | Object stereo - Joint stereo matching and object segmentationabstractThis paper presents a method for joint stereo matching and object segmentation. In our approach a 3D scene is represented as a collection of visually distinct and spatially coherent objects. Each object is characterized by three different aspects: a color model, a 3D plane that approximates the object's disparity distribution, and a novel 3D connectivity property. Inspired by Markov Random Field models of image segmentation, we employ object-level color models as a soft constraint, which can aid depth estimation in powerful ways. In particular, our method is able to recover the depth of regions that are fully occluded in one input view, which to our knowledge is new for stereo matching. Our model is formulated as an energy function that is optimized via fusion moves. We show high-quality disparity and object segmentation results on challenging image pairs as well as standard benchmarks. We believe our work not only demonstrates a novel synergy between the areas of image segmentation and stereo matching, but may also inspire new work in the domain of automatic and interactive object-level scene manipulation. Michael Bleyer, Carsten Rother, Pushmeet Kohli, Daniel Scharstein, Sudipta N. Sinha |
CVPR | 5 |
| 2011 | Structure from motion for scenes with large duplicate structuresabstractMost existing structure from motion (SFM) approaches for unordered images cannot handle multiple instances of the same structure in the scene. When image pairs containing different instances are matched based on visual similarity, the pairwise geometric relations as well as the correspondences inferred from such pairs are erroneous, which can lead to catastrophic failures in the reconstruction. In this paper, we investigate the geometric ambiguities caused by the presence of repeated or duplicate structures and show that to disambiguate between multiple hypotheses requires more than pure geometric reasoning. We couple an expectation maximization (EM)-based algorithm that estimates camera poses and identifies the false match-pairs with an efficient sampling method to discover plausible data association hypotheses. The sampling method is informed by geometric and image-based cues. Our algorithm usually recovers the correct data association, even in the presence of large numbers of false pairwise matches. Richard Roberts 0001, Sudipta N. Sinha, Richard Szeliski, Drew Steedly |
CVPR | 2 |
| 2011 | Feature tracking and matching in video using programmable graphics hardware
Sudipta N. Sinha, Jan-Michael Frahm, Marc Pollefeys, Yakup Genc |
Mach. Vis. Appl. | 1 |
| 2010 | Camera Network Calibration and Synchronization from Silhouettes in Archived Video
Sudipta N. Sinha, Marc Pollefeys |
Int. J. Comput. Vis. | 1 |
| 2009 | Piecewise planar stereo for image-based renderingabstractWe present a novel multi-view stereo method designed for image-based rendering that generates piecewise planar depth maps from an unordered collection of photographs. Sudipta N. Sinha, Drew Steedly, Richard Szeliski |
ICCV | 1 |
| 2008 | Detailed Real-Time Urban 3D Reconstruction from Video
Marc Pollefeys, David Nistér, Jan-Michael Frahm, Amir Akbarzadeh, Philippos Mordohai, Brian Clipp, Chris Engels, David Gallup, Seon Joo Kim, Paul Merrell, C. Salmi, Sudipta N. Sinha, B. Talton, Liang Wang 0002, Qingxiong Yang, Henrik Stewénius, Ruigang Yang, Greg Welch, Herman Towles |
Int. J. Comput. Vis. | 12 |
| 2008 | Interactive 3D architectural modeling from unordered photo collectionsabstractWe present an interactive system for generating photorealistic, textured, piecewise-planar 3D models of architectural structures and urban scenes from unordered sets of photographs. To reconstruct 3D geometry in our system, the user draws outlines overlaid on 2D photographs. The 3D structure is then automatically computed by combining the 2D interaction with the multi-view geometric information recovered by performing structure from motion analysis on the input photographs. We utilize vanishing point constraints at multiple stages during the reconstruction, which is particularly useful for architectural scenes where parallel lines are abundant. Our approach enables us to accurately model polygonal faces from 2D interactions in a single image. Our system also supports useful operations such as edge snapping and extrusions. Seamless texture maps are automatically generated by combining multiple input photographs using graph cut optimization and Poisson blending. The user can add brush strokes as hints during the texture generation stage to remove artifacts caused by unmodeled geometric structures. We build models for a variety of architectural scenes from collections of up to about a hundred photographs. Sudipta N. Sinha, Drew Steedly, Richard Szeliski, Maneesh Agrawala, Marc Pollefeys |
ACM Trans. Graph. | 1 |
| 2007 | Multi-View Stereo via Graph Cuts on the Dual of an Adaptive Tetrahedral MeshabstractWe formulate multi-view 3D shape reconstruction as the computation of a minimum cut on the dual graph of a semi- regular, multi-resolution, tetrahedral mesh. Our method does not assume that the surface lies within a finite band around the visual hull or any other base surface. Instead, it uses photo-consistency to guide the adaptive subdivision of a coarse mesh of the bounding volume. This generates a multi-resolution volumetric mesh that is densely tesselated in the parts likely to contain the unknown surface. The graph-cut on the dual graph of this tetrahedral mesh produces a minimum cut corresponding to a triangulated surface that minimizes a global surface cost functional. Our method makes no assumptions about topology and can recover deep concavities when enough cameras observe them. Our formulation also allows silhouette constraints to be enforced during the graph-cut step to counter its inherent bias for producing minimal surfaces. Local shape refinement via surface deformation is used to recover details in the reconstructed surface. Reconstructions of the Multi- View Stereo Evaluation benchmark datasets and other real datasets show the effectiveness of our method. Sudipta N. Sinha, Philippos Mordohai, Marc Pollefeys |
ICCV | 1 |
| 2006 | Pan-tilt-zoom camera calibration and high-resolution mosaic generation
Sudipta N. Sinha, Marc Pollefeys |
Comput. Vis. Image Underst. | 1 |
| 2005 | Synchronization and Calibration of a Camera Network for 3D Event Reconstruction from Live VideoabstractWe present an approach for automatic reconstruction of a dynamic event using multiple video cameras recording from different viewpoints. Our approach recovers all the necessary information by analyzing the motion of the silhouettes in the multiple video streams. The first step consists of computing the calibration and synchronization for pairs of cameras. We compute the temporal offset and epipolar geometry using an efficient RANSAC-based algorithm to search for the epipoles as well as for robustness. In the next stage the calibration and synchronization for the complete camera network is recovered and then refined through maximum likelihood estimation. Finally, a visual hull algorithm is used to the recover the dynamic shape of the observed object. Sudipta N. Sinha, Marc Pollefeys |
CVPR (2) | 1 |
| 2005 | Calibration of Pan-Tilt-Zoom (PTZ) Cameras and Omni-Directional CamerasabstractIn the first part we discuss the problem of recovering the calibration of a network of pan-tilt-zoom cameras. The intrinsic parameters of each camera over its full range of zoom settings are estimated through a two step procedure. We first determine the intrinsic parameters at the camera's lowest zoom setting very accurately by capturing an extended panorama. The camera intrinsics are then determined at discrete steps in a monotonically increasing zoom sequence that spans the full zoom range of the cameras. Both steps are fully automatic and do not assume any knowledge of the scene structure. We validate our approach by calibrating two different types of pan tilt zoom cameras placed in an outdoor environment. We also show the high-resolution panoramic mosaics built from the images captured during this process. The second section deals with the calibration of omnidirectional cameras. A broad class of both central and non-central cameras, such as fish-eye and catadioptric cameras, can be reduced to 1D radial cameras under the assumption of known center of radial distortion. We study the multi-view geometry of 1D radial cameras. For cameras in general configuration, we introduce a quadrifocal tensor. From this tensor a metric reconstruction of the 1D cameras as well as the observed features can be obtained. In a second phase this reconstruction can then be used as a calibration object to estimate a non-parametric non-central model for the cameras. We study some degenerate cases, including pure rotation. In the case of a purely rotating camera we obtain a trifocal tensor. This allows us to obtain a metric reconstruction of the plane at infinity. Next, we use the plane at infinity as a calibration device to non-parametrically estimate the radial distortion. SriRam Thirthala, Sudipta N. Sinha, Marc Pollefeys |
CVPR (2) | 2 |
| 2005 | Multi-View Reconstruction Using Photo-consistency and Exact Silhouette Constraints: A Maximum-Flow FormulationabstractThis paper describes a novel approach for reconstructing a closed continuous surface of an object from multiple calibrated color images and silhouettes. Any accurate reconstruction must satisfy (1) photo-consistency and (2) silhouette consistency constraints. Most existing techniques treat these cues identically in optimization frameworks where silhouette constraints are traded off against photo-consistency and smoothness priors. Our approach strictly enforces silhouette constraints, while optimizing photo-consistency and smoothness in a global graph-cut framework. We transform the reconstruction problem into computing max-flow/min-cut in a geometric graph, where any cut corresponds to a surface satisfying exact silhouette constraints (its silhouettes should exactly coincide with those of the visual hull); a minimum cut is the most photo-consistent surface amongst them. Our graph-cut formulation is based on the rim mesh, (the combinatorial arrangement of rims or contour generators from many views) which can be computed directly from the silhouettes. Unlike other methods, our approach enforces silhouette constraints without introducing a bias near the visual hull boundary and also recovers the rim curves. Results are presented for synthetic and real datasets. Sudipta N. Sinha, Marc Pollefeys |
ICCV | 1 |
| 2004 | Camera Network Calibration from Dynamic Silhouettes
Sudipta N. Sinha, Marc Pollefeys, Leonard McMillan |
CVPR (1) | 1 |
| 2004 | Iso-disparity Surfaces for General Stereo Configurations
Marc Pollefeys, Sudipta N. Sinha |
ECCV (3) | 2 |