VLDB 2026 Research / reviewers in the wild / expert
Srikumar Ramalingam
dblp:17/4216
· DBLP profile ↗
71ranked-venue papers
16as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 15 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 9 first-author · 4 since 2021Systems, architecture and hardware · 13 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SurfR: Surface Reconstruction with Multi-Scale AttentionabstractWe propose a fast and accurate surface reconstruction algorithm for unorganized point clouds using an implicit representation. Recent learning methods are either singleobject representations with small neural models that allow for high surface details but require per-object training or generalized representations that require larger models and generalize to newer shapes but lack details, and inference is slow. We propose a new implicit representation for general 3D shapes that is faster than all the baselines at their optimum resolution, with only a marginal loss in performance compared to the state-of-the-art. We achieve the best accuracy-speed trade-off using three key contributions. Many implicit methods extract features from the point cloud to classify whether a query point is inside or outside the object. First, to speed up the reconstruction, we show that this feature extraction does not need to use the query point at an early stage (lazy query). Second, we use a parallel multi-scale grid representation to develop robust features for different noise levels and input resolutions. Finally, we show that attention across scales can provide improved reconstruction results. The code will be made available. Siddhant Ranade, Gonçalo Dias Pais, Ross Tyler Whitaker, Jacinto C. Nascimento, Pedro Miraldo, Srikumar Ramalingam |
3DV | 6 |
| 2025 | GIST: Greedy Independent Set Thresholding for Max-Min Diversification with Submodular UtilityabstractThis work studies a novel subset selection problem called *max-min diversification with monotone submodular utility* (MDMS), which has a wide range of applications in machine learning, e.g., data sampling and feature selection.
Given a set of points in a metric space,
the goal of MDMS is to maximize $f(S) = g(S) + \lambda \cdot \text{div}(S)$
subject to a cardinality constraint $|S| \le k$,
where
$g(S)$ is a monotone submodular function
and
$\text{div}(S) = \min_{u,v \in S : u \ne v} \text{dist}(u,v)$ is the *max-min diversity* objective.
We propose the `GIST` algorithm, which gives a $\frac{1}{2}$-approximation guarantee for MDMS
by approximating a series of maximum independent set problems with a bicriteria greedy algorithm.
We also prove that it is NP-hard to approximate within a factor of $0.5584$.
Finally, we show in our empirical study that `GIST` outperforms state-of-the-art benchmarks
for a single-shot data sampling task on ImageNet. Matthew Fahrbach, Srikumar Ramalingam, Morteza Zadimoghaddam, Sara Ahmadian, Gui Citovsky, Giulia DeSalvo |
NeurIPS | 2 |
| 2025 | Analyzing Similarity Metrics for Data Selection for Language Model PretrainingabstractMeasuring similarity between training examples is critical for curating high-quality and diverse pretraining datasets for language models.
However, similarity is typically computed with a generic off-the-shelf embedding model that has been trained for tasks such as retrieval.
Whether these embedding-based similarity metrics are well-suited for pretraining data selection remains largely unexplored.
In this paper, we propose a new framework to assess the suitability of a similarity metric specifically for data curation in language model pretraining applications. Our framework's first evaluation criterion captures how well distances reflect generalization in pretraining loss between different training examples. Next, we use each embedding model to guide a standard diversity-based data curation algorithm and measure its utility by pretraining a language model on the selected data and evaluating downstream task performance. Finally, we evaluate the capabilities of embeddings to distinguish between examples from different data sources. With these evaluations, we demonstrate that standard off-the-shelf embedding models are not well-suited for the pretraining data curation setting, underperforming even remarkably simple embeddings that are extracted from models trained on the same pretraining corpus. Our experiments are performed on the Pile, for pretraining a 1.7B parameter language model on 200B tokens. We believe our analysis and evaluation framework serves as a foundation for the future design of embeddings that specifically reason about similarity in pretraining datasets. Dylan Sam, Ayan Chakrabarti, Afshin Rostamizadeh, Srikumar Ramalingam, Gui Citovsky, Sanjiv Kumar |
NeurIPS | 4 |
| 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image GenerationabstractModern text-to-image generation models produce high-quality images that are both photorealistic and faithful to the text prompts. However, this quality comes at significant computational cost: nearly all of these models are iterative and require running sampling multiple times with large models. This iterative process is needed to ensure that different regions of the image are not only aligned with the text prompt, but also compatible with each other. In this work, we propose a light-weight approach to achieving this compatibility between different regions of an image, using a Markov Random Field (MRF) model. We demonstrate the effectiveness of this method on top of the latent token-based Muse text-to-image model. The MRF richly encodes the compatibility among image tokens at different spatial locations to improve quality and significantly reduce the required number of Muse sampling steps. Inference with the MRF is significantly cheaper, and its parameters can be quickly learned through back-propagation by modeling MRF inference as a differentiable neural-network layer. Our full model, MarkovGen, uses this proposed MRF model to both speed up Muse by 1.5 x and produce higher quality images by decreasing undesirable image artifacts. Sadeep Jayasumana, Daniel Glasner, Srikumar Ramalingam, Andreas Veit, Ayan Chakrabarti, Sanjiv Kumar |
CVPR | 3 |
| 2024 | Rethinking FID: Towards a Better Evaluation Metric for Image GenerationabstractAs with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of Inception-v3 features of real images, and those of images generated by the algorithm. We highlight important drawbacks of FID: Inception's poor representation of the rich and varied content generated by modern text-to-image models, incorrect normality assumptions, and poor sample complexity. We call for a reevaluation of FID's use as the primary quality metric for generated images. We empirically demonstrate that FID contradicts human raters, it does not reflect gradual improvement of iterative text-to-image models, it does not capture distortion levels, and that it produces inconsistent results when varying the sample size. We also propose an alternative new metric, CMMD, based on richer CLIP embeddings and the maximum mean discrepancy distance with the Gaussian RBF kernel. It is an unbiased estimator that does not make any assumptions on the probability distribution of the embeddings and is sample efficient. Through extensive experiments and analysis, we demonstrate that FID-based evaluations of text-to-image models may be unreliable, and that CMMD of-fers a more robust and reliable assessment of image quality. A reference implementation of CMMD is available at: https://github.com/google-research/google-research/tree/master/cmmd. Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, Sanjiv Kumar |
CVPR | 2 |
| 2023 | Leveraging Importance Weights in Subset Selection
Gui Citovsky, Giulia DeSalvo, Sanjiv Kumar, Srikumar Ramalingam, Afshin Rostamizadeh, Yunjuan Wang |
ICLR | 4 |
| 2023 | Fast and Accurate 3D Registration from Line Intersection Constraints
André Mateus 0001, Siddhant Ranade, Srikumar Ramalingam, Pedro Miraldo |
Int. J. Comput. Vis. | 3 |
| 2022 | Learning ABCs: Approximate Bijective Correspondence for isolating factors of variation with weak supervisionabstractRepresentational learning forms the backbone of most deep learning applications, and the value of a learned representation is intimately tied to its information content regarding different factors of variation. Finding good representations depends on the nature of supervision and the learning algorithm. We propose a novel algorithm that utilizes a weak form of supervision where the data is partitioned into sets according to certain inactive (common) factors of variation which are invariant across elements of each set. Our key insight is that by seeking correspondence between elements of different sets, we learn strong representations that exclude the inactive factors of variation and isolate the active factors that vary within all sets. As a consequence of focusing on the active factors, our method can leverage a mix of setsupervised and wholly unsupervised data, which can even belong to a different domain. We tackle the challenging problem of synthetic-to-real object pose transfer, without pose annotations on anything, by isolating pose information which generalizes to the category level and across the synthetic/real domain gap. The method can also boost performance in supervised settings, by strengthening intermediate representations, as well as operate in practically attainable scenarios with set-supervised natural images, where quantity is limited and nuisance factors of variation are more plentiful. Accompanying code may be found on github. Kieran A. Murphy, Varun Jampani, Srikumar Ramalingam, Ameesh Makadia |
CVPR | 3 |
| 2022 | The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural NetworksabstractNeural networks tend to achieve better accuracy with training if they are larger {—} even if the resulting models are overparameterized. Nevertheless, carefully removing such excess of parameters before, during, or after training may also produce models with similar or even improved accuracy. In many cases, that can be curiously achieved by heuristics as simple as removing a percentage of the weights with the smallest absolute value {—} even though absolute value is not a perfect proxy for weight relevance. With the premise that obtaining significantly better performance from pruning depends on accounting for the combined effect of removing multiple weights, we revisit one of the classic approaches for impact-based pruning: the Optimal Brain Surgeon (OBS). We propose a tractable heuristic for solving the combinatorial extension of OBS, in which we select weights for simultaneous removal, and we combine it with a single-pass systematic update of unpruned weights. Our selection method outperforms other methods for high sparsity, and the single-pass weight update is also advantageous if applied after those methods. Xin Yu 0003, Thiago Serra, Srikumar Ramalingam, Shandian Zhe |
ICML | 3 |
| 2021 | Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation ManifoldabstractIn the deep learning era, the vast majority of methods to predict pose from a single image are trained to classify or regress to a single given ground truth pose per image. Such methods have two main shortcomings, i) they cannot represent uncertainty about the predictions, and ii) they cannot handle symmetric objects, where multiple (potentially infinite) poses may be correct. Only recently these shortcomings have been addressed, but current approaches as limited in that they cannot express the full rich space of distributions on the rotation manifold. To this end, we introduce a method to estimate arbitrary, non-parametric distributions on SO(3). Our key idea is to represent the distributions implicitly, with a neural network that estimates the probability density, given the input image and a candidate pose. At inference time, grid sampling or gradient ascent can be used to find the most likely pose, but it is also possible to evaluate the density at any pose, enabling reasoning about symmetries and uncertainty. This is the most general way of representing distributions on manifolds, and to demonstrate its expressive power we introduce a new dataset containing symmetric and nearly-symmetric objects. Our method also shows advantages on the popular object pose estimation benchmarks ModelNet10-SO(3) and T-LESS. Code, data, and visualizations may be found at implicit-pdf.github.io. Kieran A. Murphy, Carlos Esteves, Varun Jampani, Srikumar Ramalingam, Ameesh Makadia |
ICML | 4 |
| 2021 | Scaling Up Exact Neural Network Compression by ReLU StabilityabstractWe can compress a rectifier network while exactly preserving its underlying functionality with respect to a given input domain if some of its neurons are stable. However, current approaches to determine the stability of neurons with Rectified Linear Unit (ReLU) activations require solving or finding a good approximation to multiple discrete optimization problems. In this work, we introduce an algorithm based on solving a single optimization problem to identify all stable neurons. Our approach is on median 183 times faster than the state-of-art method on CIFAR-10, which allows us to explore exact compression on deeper (5 x 100) and wider (2 x 800) networks within minutes. For classifiers trained under an amount of L1 regularization that does not worsen accuracy, we can remove up to 56% of the connections on the CIFAR-10 dataset. The code is available at the following link, https://github.com/yuxwind/ExactCompression . Thiago Serra, Xin Yu 0003, Abhinav Kumar 0004, Srikumar Ramalingam |
NeurIPS | 4 |
| 2020 | Empirical Bounds on Linear Regions of Deep Rectifier NetworksabstractWe can compare the expressiveness of neural networks that use rectified linear units (ReLUs) by the number of linear regions, which reflect the number of pieces of the piecewise linear functions modeled by such networks. However, enumerating these regions is prohibitive and the known analytical bounds are identical for networks with same dimensions. In this work, we approximate the number of linear regions through empirical bounds based on features of the trained network and probabilistic inference. Our first contribution is a method to sample the activation patterns defined by ReLUs using universal hash functions. This method is based on a Mixed-Integer Linear Programming (MILP) formulation of the network and an algorithm for probabilistic lower bounds of MILP solution sets that we call MIPBound, which is considerably faster than exact counting and reaches values in similar orders of magnitude. Our second contribution is a tighter activation-based bound for the maximum number of linear regions, which is particularly stronger in networks with narrow layers. Combined, these bounds yield a fast proxy for the number of linear regions of a deep neural network. Thiago Serra, Srikumar Ramalingam |
AAAI | 2 |
| 2020 | Mapping of Sparse 3D Data Using Alternating Projection
Siddhant Ranade, Xin Yu 0003, Shantnu Kakkar, Pedro Miraldo, Srikumar Ramalingam |
ACCV (1) | 5 |
| 2020 | Lossless Compression of Deep Neural Networks
Thiago Serra, Abhinav Kumar 0004, Srikumar Ramalingam |
CPAIOR | 3 |
| 2020 | Minimal Solvers for 3D Scan Alignment With Pairs of Intersecting LinesabstractWe explore the possibility of using line intersection constraints for 3D scan registration. Typical 3D registration algorithms exploit point and plane correspondences, while line intersection constraints have not been used in the context of 3D scan registration before. Constraints from a match of pairs of intersecting lines in two 3D scans can be seen as two 3D line intersections, a plane correspondence, and a point correspondence. In this paper, we present minimal solvers that combine these different type of constraints: 1) three line intersections and one point match; 2) one line intersection and two point matches; 3) three line intersections and one plane match; 4) one line intersection and two plane matches; and 5) one line intersection, one point match, and one plane match. To use all the available solvers, we present a hybrid RANSAC loop. We propose a non-linear refinement technique using all the inliers obtained from the RANSAC. Vast experiments with simulated data and two real-data data-sets show that the use of these features and the combined solvers improve the accuracy. The code is available. André Mateus 0001, Srikumar Ramalingam, Pedro Miraldo |
CVPR | 2 |
| 2020 | 3DRegNet: A Deep Neural Network for 3D Point RegistrationabstractWe present 3DRegNet, a novel deep learning architecture for the registration of 3D scans. Given a set of 3D point correspondences, we build a deep neural network to address the following two challenges: (i) classification of the point correspondences into inliers/outliers, and (ii) regression of the motion parameters that align the scans into a common reference frame. With regard to regression, we present two alternative approaches: (i) a Deep Neural Network (DNN) registration and (ii) a Procrustes approach using SVD to estimate the transformation. Our correspondence-based approach achieves a higher speedup compared to competing baselines. We further propose the use of a refinement network, which consists of a smaller 3DRegNet as a refinement to improve the accuracy of the registration. Extensive experiments on two challenging datasets demonstrate that we outperform other methods and achieve state-of-the-art results. Gonçalo Dias Pais, Srikumar Ramalingam, Venu Madhav Govindu, Jacinto C. Nascimento, Rama Chellappa, Pedro Miraldo |
CVPR | 2 |
| 2019 | Minimal Solvers for Mini-Loop Closures in 3D Multi-Scan Alignmentabstract3D scan registration is a classical, yet a highly useful problem in the context of 3D sensors such as Kinect and Velodyne. While there are several existing methods, the techniques are usually incremental where adjacent scans are registered first to obtain the initial poses, followed by motion averaging and bundle-adjustment refinement. In this paper, we take a different approach and develop minimal solvers for jointly computing the initial poses of cameras in small loops such as 3-, 4-, and 5-cycles. Note that the classical registration of 2 scans can be done using a minimum of 3 point matches to compute 6 degrees of relative motion. On the other hand, to jointly compute the 3D registrations in n-cycles, we take 2 point matches between the first n-1 consecutive pairs (i.e., Scan 1 & Scan 2, ... , and Scan n-1 & Scan n) and 1 or 2 point matches between Scan 1 and Scan n. Overall, we use 5, 7, and 10 point matches for 3-, 4-, and 5-cycles, and recover 12, 18, and 24 degrees of transformation variables, respectively. Using simulations and real-data we show that the 3D registration using mini n-cycles are computationally efficient, and can provide alternate and better initial poses compared to standard pairwise methods. Pedro Miraldo, Surojit Saha, Srikumar Ramalingam |
CVPR | 3 |
| 2019 | Building anatomically realistic jaw kinematics model from data
Wenwu Yang, Nathan Marshak, Daniel Sýkora, Srikumar Ramalingam, Ladislav Kavan |
Vis. Comput. | 4 |
| 2018 | Novel Single View Constraints for Manhattan 3D Line ReconstructionabstractThis paper proposes a novel and exact method to reconstruct line-based 3D structure from a single image using Manhattan world assumption. This problem is a distinctly unsolved problem because there can be multiple 3D reconstructions from a single image. Thus, we are often forced to look for priors on Manhattan world and common scene structures. In addition to the standard orthogonality, linear perspective, and parallelism constraints, we investigate a few novel constraints based on the physical realizability of the scene structure. We treat the line segments in the image to be part of a graph similar to \emph{\bf straws and connectors game}, where the goal is to back-project the line segments in 3D space and while ensuring that some of these 3D line segments connect with each other (i.e., truly intersect in 3D space) to form the 3D structure. We consider three sets of novel constraints while solving the reconstruction: (1) constraints on a series of Manhattan line intersections that form cycles, but are not all physically realizable, (2) constraints on true and false intersections in the case of nearby lines lying on the same Manhattan plane, and (3) constraints on the intersections of boundary and non-boundary line segments. The reconstruction is achieved using mixed integer linear programming (MILP), and we show compelling results on real images. Along with this paper, we release a challenging Single View Line Reconstruction dataset with ground truth 3D line models for research purposes. Siddhant Ranade, Srikumar Ramalingam |
3DV | 2 |
| 2018 | LS3D: Single-View Gestalt 3D Surface Reconstruction from Manhattan Line Segments
Yiming Qian, Srikumar Ramalingam, James H. Elder |
ACCV (4) | 2 |
| 2018 | Analytical Modeling of Vanishing Points and Curves in Catadioptric CamerasabstractVanishing points and vanishing lines are classical geometrical concepts in perspective cameras that have a lineage dating back to 3 centuries. A vanishing point is a point on the image plane where parallel lines in 3D space appear to converge, whereas a vanishing line passes through 2 or more vanishing points. While such concepts are simple and intuitive in perspective cameras, their counterparts in catadioptric cameras (obtained using mirrors and lenses) are more involved. For example, lines in the 3D space map to higher degree curves in catadioptric cameras. The projection of a set of 3D parallel lines converges on a single point in perspective images, whereas they converge to more than one point in catadioptric cameras. To the best of our knowledge, we are not aware of any systematic development of analytical models for vanishing points and vanishing curves in different types of catadioptric cameras. In this paper, we derive parametric equations for vanishing points and vanishing curves using the calibration parameters, mirror shape coefficients, and direction vectors of parallel lines in 3D space. We show compelling experimental results on vanishing point estimation and absolute pose estimation for a wide range of catadioptric cameras in both simulations and real experiments. Pedro Miraldo, Francisco Girbal Eiras, Srikumar Ramalingam |
CVPR | 3 |
| 2018 | Learning Strict Identity Mappings in Deep Residual NetworksabstractA family of super deep networks, referred to as residual networks or ResNet [14], achieved record-beating performance in various visual tasks such as image recognition, object detection, and semantic segmentation. The ability to train very deep networks naturally pushed the researchers to use enormous resources to achieve the best performance. Consequently, in many applications super deep residual networks were employed for just a marginal improvement in performance. In this paper, we propose ε-ResNet that allows us to automatically discard redundant layers, which produces responses that are smaller than a threshold ε, without any loss in performance. The ε-ResNet architecture can be achieved using a few additional rectified linear units in the original ResNet. Our method does not use any additional variables nor numerous trials like other hyperparameter optimization techniques. The layer selection is achieved using a single training process and the evaluation is performed on CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets. In some instances, we achieve about 80% reduction in the number of parameters. Xin Yu 0003, Zhiding Yu, Srikumar Ramalingam |
CVPR | 3 |
| 2018 | A Minimal Closed-Form Solution for Multi-perspective Pose Estimation using Points and Lines
Pedro Miraldo, Tiago J. Dias, Srikumar Ramalingam |
ECCV (16) | 3 |
| 2018 | Simultaneous Edge Alignment and Learning
Zhiding Yu, Weiyang Liu, Yang Zou 0003, Chen Feng 0002, Srikumar Ramalingam, B. V. K. Vijaya Kumar, Jan Kautz |
ECCV (3) | 5 |
| 2018 | Bounding and Counting Linear Regions of Deep Neural NetworksabstractWe investigate the complexity of deep neural networks (DNN) that represent piecewise linear (PWL) functions. In particular, we study the number of linear regions, i.e. pieces, that a PWL function represented by a DNN can attain, both theoretically and empirically. We present (i) tighter upper and lower bounds for the maximum number of linear regions on rectifier networks, which are exact for inputs of dimension one; (ii) a first upper bound for multi-layer maxout networks; and (iii) a first method to perform exact enumeration or counting of the number of regions by modeling the DNN with a mixed-integer linear formulation. These bounds come from leveraging the dimension of the space defining each linear region. The results also indicate that a deep rectifier network can only have more linear regions than every shallow counterpart with same number of neurons if that number exceeds the dimension of the input. Thiago Serra, Christian Tjandraatmadja, Srikumar Ramalingam |
ICML | 3 |
| 2018 | VLASE: Vehicle Localization by Aggregating Semantic EdgesabstractWe propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pairs of distinct objects such as building-sky, road-sidewalk, and building-ground. While prior work has shown promising results by utilizing the boundary between prominent classes such as sky and building using skylines, we generalize this to consider 19 semantic classes. We extract semantic edge features using CASENet architecture and utilize VLAD framework to perform image retrieval. We achieve improvement over state-of-the-art localization algorithms such as SIFT-VLAD and its deep variant NetVLAD. Ablation study shows the importance of different semantic classes, and our unified approach achieves better performance compared to individual prominent features such as skylines. We also introduce SLC Marathon dataset, a challenging dataset covering most of Salt Lake City with sufficient lighting variations. Xin Yu 0003, Sagar Chaturvedi, Chen Feng 0002, Yuichi Taguchi, Teng-Yok Lee, Clinton Fernandes, Srikumar Ramalingam |
IROS | 7 |
| 2017 | Barcode: Global Binary Patterns for Fast Visual InferenceabstractWe present Barcode, a global binary descriptor for images captured from a vehicle-mounted camera with two applications: localization and turn classification. Barcode characterizes an image by encoding the distribution of vertical lines into a binary descriptor: in each vertical stripe of an image, if any vertical line exists the corresponding bit is set to 1, otherwise 0. For localization, our approach uses a database of geolocated images, each having its Barcode precomputed during a preprocessing stage. In the run time, we first generate the binary descriptor for each image and then use the descriptor to find the location in the database via Hamming distance metric. For turn classification, we train a deep neural network that uses a set of Barcodes from consecutive images to classify turns (left, right, straight, and stationary). We show that Barcode extraction can be done at 100-1000~Hz, localization at 10~kHz, and turn classification at 1~kHz. We show compelling experimental results on KITTI dataset and other sequences captured near Harvard and Purdue campuses. Teng-Yok Lee, Sonali Patil, Srikumar Ramalingam, Yuichi Taguchi, Bedrich Benes |
3DV | 3 |
| 2017 | CASENet: Deep Category-Aware Semantic Edge DetectionabstractBoundary and edge cues are highly beneficial in improving a wide variety of vision tasks such as semantic segmentation, object recognition, stereo, and object proposal generation. Recently, the problem of edge detection has been revisited and significant progress has been made with deep learning. While classical edge detection is a challenging binary problem in itself, the category-aware semantic edge detection by nature is an even more challenging multi-label problem. We model the problem such that each edge pixel can be associated with more than one class as they appear in contours or junctions belonging to two or more semantic classes. To this end, we propose a novel end-to-end deep semantic edge learning architecture based on ResNet and a new skip-layer architecture where category-wise edge activations at the top convolution layer share and are fused with the same set of bottom layer features. We then propose a multi-label loss function to supervise the fused activations. We show that our proposed architecture benefits this problem with better performance, and we outperform the current state-of-the-art semantic edge detection methods by a large margin on standard data sets such as SBD and Cityscapes. Zhiding Yu, Chen Feng 0002, Ming-Yu Liu 0001, Srikumar Ramalingam |
CVPR | 4 |
| 2017 | MonoRGBD-SLAM: Simultaneous localization and mapping using both monocular and RGBD camerasabstractRGBD SLAM systems have shown impressive results, but the limited field of view (FOV) and depth range of typical RGBD cameras still cause problems for registering distant frames. Monocular SLAM systems, in contrast, can exploit wide-angle cameras and do not have the depth range limitation, but are unstable for textureless scenes. We present a SLAM system that uses both an RGBD camera and a wide-angle monocular camera for combining the advantages of the two sensors. Our system extracts 3D point features from RGBD frames and 2D point features from monocular frames, which are used to perform both RGBD-to-RGBD and RGBD-to-monocular registration. To compensate for different FOV and resolution of the cameras, we generate multiple virtual images for each wide-angle monocular image and use the feature descriptors computed on the virtual images to perform the RGBD-to-monocular matches. To compute the poses of the frames, we construct a graph where nodes represent RGBD and monocular frames and edges denote the pairwise registration results between the nodes. We compute the global poses of the nodes by first finding the minimum spanning trees (MSTs) of the graph and then pruning edges that have inconsistent poses due to possible mismatches using the MST result. We finally run bundle adjustment on the graph using all the consistent edges. Experimental results show that our system registers a larger number of frames than using only an RGBD camera, leading to larger-scale 3D reconstruction. Khalid Yousif, Yuichi Taguchi, Srikumar Ramalingam |
ICRA | 3 |
| 2017 | ROS2D: Image feature detector using rank order statisticsabstractWe present a new image feature detection method. Our method selects features based on segmenting points with high local intensity variations across different scales using a robust rank order statistics approach. Our method produces a large number of repeatable features that are invariant to several image transformations such as rotation, scaling, viewpoint, and lighting variations. We show the advantages of our feature in comparison to other existing features using the Oxford dataset. We also show that, when used in monocular and stereo SLAM systems, our feature outperforms SIFT in terms of the pose estimation accuracy using several public datasets including the KITTI dataset. Khalid Yousif, Yuichi Taguchi, Srikumar Ramalingam, Alireza Bab-Hadiashar |
ICRA | 3 |
| 2017 | Efficient minimization of higher order submodular functions using monotonic Boolean functions
Srikumar Ramalingam, Chris Russell 0001, Lubor Ladicky, Philip Torr 0001 |
Discret. Appl. Math. | 1 |
| 2017 | A Unifying Model for Camera CalibrationabstractThis paper proposes a unified theory for calibrating a wide variety of camera models such as pinhole, fisheye, cata-dioptric, and multi-camera networks. We model any camera as a set of image pixels and their associated camera rays in space. Every pixel measures the light traveling along a (half-) ray in 3-space, associated with that pixel. By this definition, calibration simply refers to the computation of the mapping between pixels and the associated 3D rays. Such a mapping can be computed using images of calibration grids, which are objects with known 3D geometry, taken from unknown positions. This general camera model allows to represent non-central cameras; we also consider two special subclasses, namely central and axial cameras. In a central camera, all rays intersect in a single point, whereas the rays are completely arbitrary in a non-central one. Axial cameras are an intermediate case: the camera rays intersect a single line. In this work, we show the theory for calibrating central, axial and non-central models using calibration grids, which can be either three-dimensional or planar. Srikumar Ramalingam, Peter F. Sturm |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Pinpoint SLAM: A hybrid of 2D and 3D simultaneous localization and mapping for RGB-D sensorsabstractConventional SLAM systems with an RGB-D sensor use depth measurements only in a limited depth range due to hardware limitation and noise of the sensor, ignoring regions that are too far or too close from the sensor. Such systems introduce registration errors especially in scenes with large depth variations. In this paper, we present a novel RGB-D SLAM system that makes use of both 2D and 3D measurements. Our system first extracts keypoints from RGB images and generates 2D and 3D point features from the keypoints with invalid and valid depth values, respectively. It then establishes 3D-to-3D, 2D-to-3D, and 2D-to-2D point correspondences among frames. For the 2D-to-3D point correspondences, we use the rays defined by the 2D point features to “pinpoint” the corresponding 3D point features, generating longer-range constraints than using only 3D-to-3D correspondences. For the 2D-to-2D point correspondences, we triangulate the rays to generate 3D points that are used as 3D point features in the subsequent process. We use the hybrid correspondences in both online SLAM and offline postprocessing: the online SLAM focuses more on the speed by computing correspondences among consecutive frames for real-time operations, while the offline postprocessing generates more correspondences among all the frames for higher accuracy. The results on RGB-D SLAM benchmarks show that the online SLAM provides higher accuracy than conventional SLAM systems, while the postprocessing further improves the accuracy. Esra Ataer Cansizoglu, Yuichi Taguchi, Srikumar Ramalingam |
ICRA | 3 |
| 2016 | High-performance and tunable stereo reconstructionabstractTraditional stereo algorithms have focused their efforts on reconstruction quality and have largely avoided prioritizing for run time performance. Robots, on the other hand, require quick maneuverability and effective computation to observe its immediate environment and perform tasks within it. In this work, we propose a high-performance and tunable stereo disparity estimation method, with a peak frame-rate of 120Hz (VGA resolution, on a single CPU-thread), that can potentially enable robots to quickly reconstruct their immediate surroundings and maneuver at high-speeds. Our key contribution is a disparity estimation algorithm that iteratively approximates the scene depth via a piece-wise planar mesh from stereo imagery, with a fast depth validation step for semi-dense reconstruction. The mesh is initially seeded with sparsely matched keypoints, and is recursively tessellated and refined as needed (via a resampling stage), to provide the desired stereo disparity accuracy. The inherent simplicity and speed of our approach, with the ability to tune it to a desired reconstruction quality and runtime performance makes it a compelling solution for applications in high-speed vehicles. Sudeep Pillai, Srikumar Ramalingam, John J. Leonard |
ICRA | 2 |
| 2016 | Fast localization and tracking using event sensorsabstractThe success of many robotics applications hinges on the speed at which the underlying sensing and inference tasks are carried out. Many high-speed applications such as autonomous driving and evasive maneuvering of quadrotors require high run time performance, which traditional cameras can seldom provide. In this paper we develop a fast localization and tracking algorithm using an event sensor, which produces on the order of million asynchronous events per second at pixels where luminance changes. The events are usually fired at the high gradient pixels (edges), where luminance changes occur as the sensor moves. We develop a fast spatio-temporal binning scheme to detect lines from these events at the edges. We represent the 3D model of the world using vertical lines, and the sensor pose can be estimated using the correspondences from 2D event lines to 3D world lines. The inherent simplicity of our method enables us to achieve a run time performance of 1000 Hertz. Wenzhen Yuan 0001, Srikumar Ramalingam |
ICRA | 2 |
| 2016 | Parameter learning for improving binary descriptor matchingabstractBinary descriptors allow fast detection and matching algorithms in computer vision problems. Though binary descriptors can be computed at almost two orders of magnitude faster than traditional gradient based descriptors, they suffer from poor matching accuracy in challenging conditions. In this paper we propose three improvements for binary descriptors in their computation and matching that enhance their performance in comparison to traditional binary and non-binary descriptors without compromising their speed. This is achieved by learning some weights and threshold parameters that allow customized matching under some variations such as lighting and viewpoint. Our suggested improvements can be easily applied to any binary descriptor. We demonstrate our approach on the ORB (Oriented FAST and Rotated BRIEF) descriptor and compare its performance with the traditional ORB and SIFT descriptors on a wide variety of datasets. In all instances, our enhancements outperform standard ORB and are comparable to SIFT. Bharath Sankaran, Srikumar Ramalingam, Yuichi Taguchi |
IROS | 2 |
| 2015 | Semantic Classification of Boundaries of an RGBD ImageabstractThe problem of labeling the edges present in a single color image as convex, concave, and occluding entities is one of the fundamental problems in computer vision. It has been shown that this information can contribute to segmentation, reconstruction and recognition problems. Recently, it has been shown that this classification is not straightforward even using RGBD data. This makes us wonder whether this apparent simple cue has more information than a depth map? In this paper, we propose a novel algorithm using random forest for classifying edges into convex, concave and occluding entities. We release a data set with more than 500 RGBD images with pixel-wise ground labels. Our method produces promising results and achieves an F-score of 0.84 on the data set. Nishit Soni, Anoop M. Namboodiri, C. V. Jawahar, Srikumar Ramalingam |
BMVC | 4 |
| 2015 | Line-sweep: Cross-ratio for wide-baseline matching and 3D reconstructionabstractWe propose a simple and useful idea based on cross-ratio constraint for wide-baseline matching and 3D reconstruction. Most existing methods exploit feature points and planes from images. Lines have always been considered notorious for both matching and reconstruction due to the lack of good line descriptors. We propose a method to generate and match new points using virtual lines constructed using pairs of keypoints, which are obtained using standard feature point detectors. We use cross-ratio constraints to obtain an initial set of new point matches, which are subsequently used to obtain line correspondences. We develop a method that works for both calibrated and uncalibrated camera configurations. We show compelling line-matching and large-scale 3D reconstruction. Srikumar Ramalingam, Michel Antunes, Daniel Snow, Gim Hee Lee, Sudeep Pillai |
CVPR | 1 |
| 2015 | Estimating Drivable Collision-Free Space from Monocular VideoabstractIn this paper we propose a novel algorithm for estimating the drivable collision-free space for autonomous navigation of on-road and on-water vehicles. In contrast to previous approaches that use stereo cameras or LIDAR, we show a method to solve this problem using a single camera. Inspired by the success of many vision algorithms that employ dynamic programming for efficient inference, we reduce the free space estimation task to an inference problem on a 1D graph, where each node represents a column in the image and its label denotes a position that separates the free space from the obstacles. Our algorithm exploits several image and geometric features based on edges, color, and homography to define potential functions on the 1D graph, whose parameters are learned through structured SVM. We show promising results on the challenging KITTI dataset as well as video collected from boats. Srikumar Ramalingam, Yuichi Taguchi, Yohei Miki 0003, Raquel Urtasun |
WACV | 2 |
| 2015 | Guest Editors' Introduction: Special Section on Higher Order Graphical Models in Computer VisionabstractThe papers in this special section address the programs and services supported by graphical models in computer vision. This section explores the main challenges in this framework—modeling novel priors, learning, inference—and presents innovative solutions. The papers cover the aspects of modeling novel priors, inference algorithms and parameter learning methods in the context of higher order graphical models. Karteek Alahari, Dhruv Batra, Srikumar Ramalingam, Nikos Paragios, Richard S. Zemel |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Calibration of Non-overlapping Cameras Using an External SLAM SystemabstractWe present a simple method for calibrating a set of cameras that may not have overlapping field of views. We reduce the problem of calibrating the non-overlapping cameras to the problem of localizing the cameras with respect to a global 3D model reconstructed with a simultaneous localization and mapping (SLAM) system. Specifically, we first reconstruct such a global 3D model by using a SLAM system using an RGB-D sensor. We then perform localization and intrinsic parameter estimation for each camera by using 2D-3D correspondences between the camera and the 3D model. Our method locates the cameras within the 3D model, which is useful for visually inspecting camera poses and provides a model-guided browsing interface of the images. We demonstrate the advantages of our method using several indoor scenes. Esra Ataer Cansizoglu, Yuichi Taguchi, Srikumar Ramalingam, Yohei Miki 0003 |
3DV | 3 |
| 2014 | A vanishing point-based global descriptor for Manhattan scenesabstractViewpoint-invariant object matching is challenging due to image distortions caused by several factors such as rotation, translation, illumination, cropping and occlusion. We propose a compact, global image descriptor for Manhattan scenes that captures relative locations and strengths of edges along vanishing directions. To construct the descriptor, an edge map is determined per vanishing point, capturing the edge strengths over a range of angles measured at the vanishing point. For matching, descriptors from two scenes are compared across multiple candidate scales and displacements. The matching performance is refined by comparing edge shapes at the local maxima of the scale-displacement plots. The proposed descriptor matching algorithm achieves an equal error rate of 7% for the Zurich Buildings Database, indicating significant gains in discriminative ability over other global descriptors that rely on aggregate image statistics but do not exploit the underlying scene geometry. Rohit Naini, Shantanu Rane, Srikumar Ramalingam |
ICASSP | 3 |
| 2014 | Entropy-Rate Clustering: Cluster Analysis via Maximizing a Submodular Function Subject to a Matroid ConstraintabstractWe propose a new objective function for clustering. This objective function consists of two components: the entropy rate of a random walk on a graph and a balancing term. The entropy rate favors formation of compact and homogeneous clusters, while the balancing function encourages clusters with similar sizes and penalizes larger clusters that aggressively group samples. We present a novel graph construction for the graph associated with the data and show that this construction induces a matroid--a combinatorial structure that generalizes the concept of linear independence in vector spaces. The clustering result is given by the graph topology that maximizes the objective function under the matroid constraint. By exploiting the submodular and monotonic properties of the objective function, we develop an efficient greedy algorithm. Furthermore, we prove an approximation bound of (1/2) for the optimality of the greedy solution. We validate the proposed algorithm on various benchmarks and show its competitive performances with respect to popular clustering algorithms. We further apply it for the task of superpixel segmentation. Experiments on the Berkeley segmentation data set reveal its superior performances over the state-of-the-art superpixel segmentation algorithms in all the standard evaluation metrics. Ming-Yu Liu 0001, Oncel Tuzel, Srikumar Ramalingam, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Single Image Calibration of Multi-axial Imaging SystemsabstractImaging systems consisting of a camera looking at multiple spherical mirrors (reflection) or multiple refractive spheres (refraction) have been used for wide-angle imaging applications. We describe such setups as multi-axial imaging systems, since a single sphere results in an axial system. Assuming an internally calibrated camera, calibration of such multi-axial systems involves estimating the sphere radii and locations in the camera coordinate system. However, previous calibration approaches require manual intervention or constrained setups. We present a fully automatic approach using a single photo of a 2D calibration grid. The pose of the calibration grid is assumed to be unknown and is also recovered. Our approach can handle unconstrained setups, where the mirrors/refractive balls can be arranged in any fashion, not necessarily on a grid. The axial nature of rays allows us to compute the axis of each sphere separately. We then show that by choosing rays from two or more spheres, the unknown pose of the calibration grid can be obtained linearly and independently of sphere radii and locations. Knowing the pose, we derive analytical solutions for obtaining the sphere radius and location. This leads to an interesting result that 6-DOF pose estimation of a multi-axial camera can be done without the knowledge of full calibration. Simulations and real experiments demonstrate the applicability of our algorithm. Amit K. Agrawal, Srikumar Ramalingam |
CVPR | 2 |
| 2013 | Manhattan Junction Catalogue for Spatial Reasoning of Indoor ScenesabstractJunctions are strong cues for understanding the geometry of a scene. In this paper, we consider the problem of detecting junctions and using them for recovering the spatial layout of an indoor scene. Junction detection has always been challenging due to missing and spurious lines. We work in a constrained Manhattan world setting where the junctions are formed by only line segments along the three principal orthogonal directions. Junctions can be classified into several categories based on the number and orientations of the incident line segments. We provide a simple and efficient voting scheme to detect and classify these junctions in real images. Indoor scenes are typically modeled as cuboids and we formulate the problem of the cuboid layout estimation as an inference problem in a conditional random field. Our formulation allows the incorporation of junction features and the training is done using structured prediction techniques. We outperform other single view geometry estimation methods on standard datasets. Srikumar Ramalingam, Jaishanker K. Pillai, Yuichi Taguchi |
CVPR | 1 |
| 2013 | Lifting 3D Manhattan Lines from a Single ImageabstractWe propose a novel and an efficient method for reconstructing the 3D arrangement of lines extracted from a single image, using vanishing points, orthogonal structure, and an optimization procedure that considers all plausible connectivity constraints between lines. Line detection identifies a large number of salient lines that intersect or nearly intersect in an image, but relatively a few of these apparent junctions correspond to real intersections in the 3D scene. We use linear programming (LP) to identify a minimal set of least-violated connectivity constraints that are sufficient to unambiguously reconstruct the 3D lines. In contrast to prior solutions that primarily focused on well-behaved synthetic line drawings with severely restricting assumptions, we develop an algorithm that can work on real images. The algorithm produces line reconstruction by identifying 95% correct connectivity constraints in York Urban database, with a total computation time of 1 second per image. Srikumar Ramalingam, Matthew Brand |
ICCV | 1 |
| 2013 | Point-plane SLAM for hand-held 3D sensorsabstractWe present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. We show that it is possible to register 3D data in two different coordinate systems using any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points). Our algorithm uses the minimal set of primitives in a RANSAC framework to robustly compute correspondences and estimate the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D data, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of indoor spaces using a hand-held Kinect sensor. Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002 |
ICRA | 3 |
| 2013 | A Theory of Minimal 3D Point to 3D Plane Registration and Its Generalization
Srikumar Ramalingam, Yuichi Taguchi |
Int. J. Comput. Vis. | 1 |
| 2012 | A theory of multi-layer flat refractive geometryabstractFlat refractive geometry corresponds to a perspective camera looking through single/multiple parallel flat refractive mediums. We show that the underlying geometry of rays corresponds to an axial camera. This realization, while missing from previous works, leads us to develop a general theory of calibrating such systems using 2D-3D correspondences. The pose of 3D points is assumed to be unknown and is also recovered. Calibration can be done even using a single image of a plane. We show that the unknown orientation of the refracting layers corresponds to the underlying axis, and can be obtained independently of the number of layers, their distances from the camera and their refractive indices. Interestingly, the axis estimation can be mapped to the classical essential matrix computation and 5-point algorithm [15] can be used. After computing the axis, the thicknesses of layers can be obtained linearly when refractive indices are known, and we derive analytical solutions when they are unknown. We also derive the analytical forward projection (AFP) equations to compute the projection of a 3D point via multiple flat refractions, which allows non-linear refinement by minimizing the reprojection error. For two refractions, AFP is either 4th or 12th degree equation depending on the refractive indices. We analyze ambiguities due to small field of view, stability under noise, and show how a two layer system can be well approximated as a single layer system. Real experiments using a water tank validate our theory. Amit K. Agrawal, Srikumar Ramalingam, Yuichi Taguchi, Visesh Chari |
CVPR | 2 |
| 2012 | Convex bricks: A new primitive for visual hull modeling and reconstructionabstractIndustrial automation tasks typically require a 3D model of the object for robotic manipulation. The ability to reconstruct the 3D model using a sample object is useful when CAD models are not available. For textureless objects, visual hull of the object obtained using silhouette-based reconstruction can avoid expensive 3D scanners for 3D modeling. We propose convex brick (CB), a new 3D primitive for modeling and reconstructing a visual hull from silhouettes. CB's are powerful in modeling arbitrary non-convex 3D shapes. Using CB, we describe an algorithm to generate a polyhedral visual hull from polygonal silhouettes; the visual hull is reconstructed as a combination of 3D convex bricks. Our approach uses well-studied geometric operations such as 2D convex decomposition and intersection of 3D convex cones using linear programming. The shape of CB can adapt to the given silhouettes, thereby significantly reducing the number of primitives required for a volumetric representation. Our framework allows easy control of reconstruction parameters such as accuracy and the number of required primitives. We present an extensive analysis of our algorithm and show visual hull reconstruction on challenging real datasets consisting of highly non-convex shapes. We also show real results on pose estimation of an industrial part in a bin-picking system using the reconstructed visual hull. Visesh Chari, Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
ICRA | 4 |
| 2012 | Voting-based pose estimation for robotic assembly using a 3D sensorabstractWe propose a voting-based pose estimation algorithm applicable to 3D sensors, which are fast replacing their 2D counterparts in many robotics, computer vision, and gaming applications. It was recently shown that a pair of oriented 3D points, which are points on the object surface with normals, in a voting framework enables fast and robust pose estimation. Although oriented surface points are discriminative for objects with sufficient curvature changes, they are not compact and discriminative enough for many industrial and real-world objects that are mostly planar. As edges play the key role in 2D registration, depth discontinuities are crucial in 3D. In this paper, we investigate and develop a family of pose estimation algorithms that better exploit this boundary information. In addition to oriented surface points, we use two other primitives: boundary points with directions and boundary line segments. Our experiments show that these carefully chosen primitives encode more information compactly and thereby provide higher accuracy for a wide class of industrial parts and enable faster computation. We demonstrate a practical robotic bin-picking system using the proposed algorithm and a 3D sensor. Changhyun Choi, Yuichi Taguchi, Oncel Tuzel, Ming-Yu Liu 0001, Srikumar Ramalingam |
ICRA | 5 |
| 2012 | SLAM using both points and planes for hand-held 3D sensorsabstractWe present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. Our algorithm uses any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points) in a RANSAC framework to efficiently compute the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D scenes, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of office spaces using a hand-held Kinect sensor. Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002 |
ISMAR | 3 |
| 2011 | Beyond Alhazen's problem: Analytical projection model for non-central catadioptric cameras with quadric mirrorsabstractCatadioptric cameras are widely used to increase the field of view using mirrors. Central catadioptric systems having an effective single viewpoint are easy to model and use, but severely constraint the camera positioning with respect to the mirror. On the other hand, non-central catadioptric systems allow greater flexibility in camera placement, but are often approximated using central or linear models due to the lack of an exact model. We bridge this gap and describe an exact projection model for non-central catadioptric systems. We derive an analytical `forward projection' equation for the projection of a 3D point reflected by a quadric mirror on the imaging plane of a perspective camera, with no restrictions on the camera placement, and show that it is an 8thdegree equation in a single unknown. While previous non-central catadioptric cameras primarily use an axial configuration where the camera is placed on the axis of a rotationally symmetric mirror, we allow off-axis (any) camera placement. Using this analytical model, a non-central catadioptric camera can be used for sparse as well as dense 3D reconstruction similar to perspective cameras, using well-known algorithms such as bundle adjustment and plane sweeping. Our paper is the first to show such results for off-axis placement of camera with multiple quadric mirrors. Simulation and real results using parabolic mirrors and an off-axis perspective camera are demonstrated. Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
CVPR | 3 |
| 2011 | Entropy rate superpixel segmentationabstractWe propose a new objective function for superpixel segmentation. This objective function consists of two components: entropy rate of a random walk on a graph and a balancing term. The entropy rate favors formation of compact and homogeneous clusters, while the balancing function encourages clusters with similar sizes. We present a novel graph construction for images and show that this construction induces a matroid - a combinatorial structure that generalizes the concept of linear independence in vector spaces. The segmentation is then given by the graph topology that maximizes the objective function under the matroid constraint. By exploiting submodular and mono-tonic properties of the objective function, we develop an efficient greedy algorithm. Furthermore, we prove an approximation bound of ½ for the optimality of the solution. Extensive experiments on the Berkeley segmentation benchmark show that the proposed algorithm outperforms the state of the art in all the standard evaluation metrics. Ming-Yu Liu 0001, Oncel Tuzel, Srikumar Ramalingam, Rama Chellappa |
CVPR | 3 |
| 2011 | The light-path less traveledabstractThis paper extends classical object pose and relative camera motion estimation algorithms for imaging sensors sampling the scene through light-paths. Many algorithms in multi-view geometry assume that every pixel observes light traveling in a single line in space. We wish to relax this assumption and address various theoretical and practical issues in modeling camera rays as piece-wise linear-paths. Such paths consisting of finitely many linear segments are typical of any simple camera configuration with reflective and refractive elements. Our main contribution is to propose efficient algorithms that can work with the complete light-path without knowing the correspondence between their individual segments and the scene points. Second, we investigate light-paths containing infinitely many and small piece-wise linear segments that can be modeled using simple parametric curves such as conics. We show compelling simulations and real experiments, involving catadioptric configurations and mirages, to validate our study. Srikumar Ramalingam, Sofien Bouaziz, Peter F. Sturm, Philip Torr 0001 |
CVPR | 1 |
| 2011 | Pose estimation using both points and lines for geo-localizationabstractThis paper identifies and fills the probably last two missing items in minimal pose estimation algorithms using points and lines. Pose estimation refers to the problem of recovering the pose of a calibrated camera given known features (points or lines) in the world and their projections on the image. There are four minimal configurations using point and line features: 3 points, 2 points and 1 line, 1 point and 2 lines, 3 lines. The first and the last scenarios that depend solely on either points or lines have been studied a few decades earlier. However the mixed scenarios, which are more common in practice, have not been solved yet. In this paper we show that it is indeed possible to develop a general technique that can solve all four scenarios using the same approach and that the solutions involve computing the roots of either a 4th degree or an 8th degree equation. The centerpiece of our method is a simple and generic method that uses collinearity and coplanarity constraints for solving the pose. In addition to validating the performance of these algorithms in simulations, we also show a compelling application for geo-localization using image sequences and coarse (plane-based) 3D models of GPS challenged urban canyons. Srikumar Ramalingam, Sofien Bouaziz, Peter F. Sturm |
ICRA | 1 |
| 2011 | Finding a needle in a specular haystackabstractProgress in machine vision algorithms has led to widespread adoption of these techniques to automate several industrial assembly tasks. Nevertheless, shiny or specular objects which are common in industrial environments still present a great challenge for vision systems. In this paper, we take a step towards this problem under the context of vision-aided robotic assembly. We show that when the illumination source moves, the specular highlights remain in a region whose radius is inversely proportional to the surface curvature. This allows us to extract regions of the object that have high surface curvature. These points of high curvature can be used as features for specular objects. Further, an inexpensive multi-flash camera (MFC) design can be used to reliably extract these features. We show that one can use multiple views of the object using the MFC in order to triangulate and obtain the 3D location and pose of the shiny objects. Finally, we show a system consisting of a robot arm with an MFC that can perform automated detection and pose estimation of shiny screws within a cluttered bin, achieving position and orientation errors less than 0.5 mm and 0.8° respectively. Nitesh Shroff, Yuichi Taguchi, Oncel Tuzel, Ashok Veeraraghavan, Srikumar Ramalingam, Haruhisa Okuda |
ICRA | 5 |
| 2010 | Axial light field for curved mirrors: Reflect your perspective, widen your viewabstractMirrors have been used to enable wide field-of-view (FOV) catadioptric imaging. The mapping between the incoming and reflected light rays depends non-linearly on the mirror shape and has been well-studied using caustics. We analyze this mapping using two-plane light field parameterization, which provides valuable insight into the geometric structure of reflected rays. Using this analysis, we study the problem of generating a single-viewpoint virtual perspective image for catadioptric systems, which is unachievable for several common configurations. Instead of minimizing distortions appearing in a single image, we propose to capture all the rays required to generate a virtual perspective by capturing a light field. We consider rotationally symmetric mirrors and show that a traditional planar light field results in significant aliasing artifacts. We propose axial light field, captured by moving the camera along the mirror rotation axis, for efficient sampling and to remove aliasing artifacts. This allows us to computationally generate wide FOV virtual perspectives using a wider class of mirrors than before, without using scene priors or depth estimation. We analyze the relationship between the axial light field parameters and the FOV/resolution of the resulting virtual perspective. Real results using a spherical mirror demonstrate generating 140° FOV virtual perspective using multiple 30° FOV images. Yuichi Taguchi, Amit K. Agrawal, Srikumar Ramalingam, Ashok Veeraraghavan |
CVPR | 3 |
| 2010 | Analytical Forward Projection for Axial Non-central Dioptric and Catadioptric Cameras
Amit K. Agrawal, Yuichi Taguchi, Srikumar Ramalingam |
ECCV (3) | 3 |
| 2010 | P2Pi: A Minimal Solution for Registration of 3D Points to 3D Planes
Srikumar Ramalingam, Yuichi Taguchi, Tim K. Marks, Oncel Tuzel |
ECCV (5) | 1 |
| 2010 | SKYLINE2GPS: Localization in urban canyons using omni-skylinesabstractThis paper investigates the problem of geo-localization in GPS challenged urban canyons using only skylines. Our proposed solution takes a sequence of upward facing omnidirectional images and coarse 3D models of cities to compute the geo-trajectory. The camera is oriented upwards to capture images of the immediate skyline, which is generally unique and serves as a fingerprint for a specific location in a city. Our goal is to estimate global position by matching skylines extracted from omni-directional images to skyline segments from coarse 3D city models. Under day-time and clear sky conditions, we propose a sky-segmentation algorithm using graph cuts for estimating the geo-location. In cases where the skyline gets affected by partial fog, night-time and occlusions from trees, we propose a shortest path algorithm that computes the location without prior sky detection. We show compelling experimental results for hundreds of images taken in New York, Boston and Tokyo under various weather and lighting conditions (daytime, foggy dawn and night-time). Srikumar Ramalingam, Sofien Bouaziz, Peter F. Sturm, Matthew Brand |
IROS | 1 |
| 2010 | Generic self-calibration of central cameras
Srikumar Ramalingam, Peter F. Sturm, Suresh K. Lodha |
Comput. Vis. Image Underst. | 1 |
| 2010 | Axial-cones: modeling spherical catadioptric cameras for wide-angle light field renderingabstractCatadioptric imaging systems are commonly used for wide-angle imaging, but lead to multi-perspective images which do not allow algorithms designed for perspective cameras to be used. Efficient use of such systems requires accurate geometric ray modeling as well as fast algorithms. We present accurate geometric modeling of the multi-perspective photo captured with a spherical catadioptric imaging system usingaxial-cone cameras:multiple perspective cameras lying on an axis each with a different viewpoint and a different cone of rays. This modeling avoids geometric approximations and allows several algorithms developed for perspective cameras to be applied to multi-perspective catadioptric cameras. We demonstrate axial-cone modeling in the context of rendering wide-angle light fields, captured using a spherical mirror array. We present several applications such as spherical distortion correction, digital refocusing for artistic depth of field effects in wide-angle scenes, and wide-angle dense depth estimation. Our GPU implementation using axial-cone modeling achieves up to three orders of magnitude speed up over ray tracing for these applications. Yuichi Taguchi, Amit K. Agrawal, Ashok Veeraraghavan, Srikumar Ramalingam, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2008 | Exact inference in multi-label CRFs with higher order cliquesabstractThis paper addresses the problem of exactly inferring the maximum a posteriori solutions of discrete multi-label MRFs or CRFs with higher order cliques. We present a framework to transform special classes of multi-label higher order functions to submodular second order Boolean functions (referred to as Fs2), which can be minimized exactly using graph cuts and we characterize those classes. The basic idea is to use two or more Boolean variables to encode the states of a single multi-label variable. There are many ways in which this can be done and much interesting research lies in finding ways which are optimal or minimal in some sense. We study the space of possible encodings and find the ones that can transform the most general class of functions to Fs2. Our main contributions are two-fold. First, we extend the subclass of submodular energy functions that can be minimized exactly using graph cuts. Second, we show how higher order potentials can be used to improve single view 3D reconstruction results. We believe that our work on exact minimization of higher order energy functions will lead to similar improvements in solutions of other labelling problems. Srikumar Ramalingam, Pushmeet Kohli, Karteek Alahari, Philip Torr 0001 |
CVPR | 1 |
| 2008 | Minimal solutions for generic imaging modelsabstractA generic imaging model refers to a non-parametric camera model where every camera is treated as a set of unconstrained projection rays. Calibration would simply be a method to map the projection rays to image pixels; such a mapping can be computed using plane based calibration grids. However, existing algorithms for generic calibration use more point correspondences than the theoretical minimum. It has been well-established that non-minimal solutions for calibration and structure-from-motion algorithms are generally noise-prone compared to minimal solutions. In this work we derive minimal solutions for generic calibration algorithms. Our algorithms for generally central cameras use 4 point correspondences in three calibration grids to compute the motion between the grids. Using simulations we show that our minimal solutions are more robust to noise compared to non-minimal solutions. We also show very accurate distortion correction results on fisheye images. Srikumar Ramalingam, Peter F. Sturm |
CVPR | 1 |
| 2008 | Randomized trees for human pose detectionabstractThis paper addresses human pose recognition from video sequences by formulating it as a classification problem. Unlike much previous work we do not make any assumptions on the availability of clean segmentation. The first step of this work consists in a novel method of aligning the training images using 3D Mocap data. Next we define classes by discretizing a 2D manifold whose two dimensions are camera viewpoint and actions. Our main contribution is a pose detection algorithm based on random forests. A bottom-up approach is followed to build a decision tree by recursively clustering and merging the classes at each level. For each node of the decision tree we build a list of potentially discriminative features using the alignment of training images; in this paper we consider Histograms of Orientated Gradient (HOG). We finally grow an ensemble of trees by randomly sampling one of the selected HOG blocks at each node. Our proposed approach gives promising results with both fixed and moving cameras. Grégory Rogez, Jonathan Rihan, Srikumar Ramalingam, Carlos Orrite-Uruñuela, Philip Torr 0001 |
CVPR | 3 |
| 2007 | Learning priors for calibrating families of stereo camerasabstractOnline camera recalibration is necessary for long-term deployment of computer vision systems. Existing algorithms assume that the source of recalibration information is a set of features in a general 3D scene; and that enough features are observed that the calibration problem is well-constrained. However; these assumptions are frequently invalid outside the laboratory. Real-world scenes often lack texture, contain repeated texture, or are mostly planar, making calibration difficult or impossible. In this paper we consider the calibration of families of stereo cameras, where each camera is assumed to have parameters drawn from a common but unknown prior distribution. We show how estimation of this prior using a small-number of offline-calibrated cameras (e.g. from the same production line) allows online calibration of additional cameras using a small number of point correspondences; and that using the estimated prior significantly increases the accuracy and robustness of stereo camera calibration. Andrew W. Fitzgibbon, Duncan P. Robertson, Antonio Criminisi, Srikumar Ramalingam, Andrew Blake 0001 |
ICCV | 4 |
| 2006 | Theory and Calibration for Axial Cameras
Srikumar Ramalingam, Peter F. Sturm, Suresh K. Lodha |
ACCV (1) | 1 |
| 2006 | A generic structure-from-motion framework
Srikumar Ramalingam, Suresh K. Lodha, Peter F. Sturm |
Comput. Vis. Image Underst. | 1 |
| 2005 | Towards Complete Generic Camera CalibrationabstractWe consider the problem of calibrating a highly generic imaging model, that consists of a non-parametric association of a projection ray in 3D to every pixel in an image. Previous calibration approaches for this model do not seem to be directly applicable for cameras with large fields of view and non-central cameras. In this paper, we describe a complete calibration approach that should in principle be able to handle any camera that can be described by the generic imaging model. Initial calibration is performed using multiple-images of overlapping calibration grids simultaneously. This is then improved using pose estimation and bundle adjustment-type algorithms. The approach has been applied on a wide variety of central and non-central cameras including fisheye lens, catadioptric cameras with spherical and hyperbolic mirrors, and multi-camera setups. We also consider the question if non-central models are more appropriate for certain cameras than central models. Srikumar Ramalingam, Peter F. Sturm, Suresh K. Lodha |
CVPR (1) | 1 |
| 2004 | A Generic Concept for Camera Calibration
Peter F. Sturm, Srikumar Ramalingam |
ECCV (2) | 2 |