VLDB 2026 Research / reviewers in the wild / expert
Suryansh Kumar 0001
dblp:124/2783
· DBLP profile ↗
26ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-2755-8744ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time-archival camera virtualization for sports and visual performances
William Stone, Suryansh Kumar 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Stereo Risk: A Continuous Modeling Approach to Stereo MatchingabstractWe introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing the scene disparity values, yet via discretization of scene disparity values. Such discretization often fails to capture the nuanced, continuous nature of scene depth. Stereo Risk departs from the conventional discretization approach by formulating the scene disparity as an optimal solution to a continuous risk minimization problem, hence the name "stereo risk". We demonstrate that $L^1$ minimization of the proposed continuous risk function enhances stereo-matching performance for deep networks, particularly for disparities with multi-modal probability distributions. Furthermore, to enable the end-to-end network training of the non-differentiable $L^1$ risk optimization, we exploited the implicit function theorem, ensuring a fully differentiable network. A comprehensive analysis demonstrates our method's theoretical soundness and superior performance over the state-of-the-art methods across various benchmark datasets, including KITTI 2012, KITTI 2015, ETH3D, SceneFlow, and Middlebury 2014. Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Yao Yao 0008, Luc Van Gool |
ICML | 2 |
| 2024 | ICGNet: A Unified Approach for Instance-Centric GraspingabstractAccurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible grasps. These grasps need to be compliant with the local object geometry. Second, for each proposed grasp, the robot needs to reason about the interactions with other objects in the scene. Finally, the robot must compute a collision-free grasp trajectory while taking into account the geometry of the target object. Most grasp detection algorithms directly predict grasp poses in a monolithic fashion, which does not capture the composability of the environment. In this paper, we introduce an end-to-end architecture for object-centric grasping. The method uses pointcloud data from a single arbitrary viewing direction as an input and generates an instance-centric representation for each partially observed object in the scene. This representation is further used for object reconstruction and grasp detection in cluttered table-top scenes. We show the effectiveness of the proposed method by extensively evaluating it against state-of-the-art methods on synthetic datasets, indicating superior performance for grasping and reconstruction. Additionally, we demonstrate real-world applicability by decluttering scenes with varying numbers of objects. Videos and Code icgraspnet.github.io. René Zurbrügg, Yifan Liu 0001, Francis Engelmann, Suryansh Kumar 0001, Marco Hutter 0001, Vaishakh Patil, Fisher Yu 0001 |
ICRA | 4 |
| 2024 | Learning Robust Multi-scale Representation for Neural Radiance Fields from Unposed ImagesabstractWe introduce an improved solution to the neural image-based rendering problem in computer vision. Given a set of images taken from a freely moving camera at train time, the proposed approach could synthesize a realistic image of the scene from a novel viewpoint at test time. The key ideas presented in this paper are (i) Recovering accurate camera parameters via a robust pipeline from unposed day-to-day images is equally crucial in neural novel view synthesis problem; (ii) It is rather more practical to model object’s content at different resolutions since dramatic camera motion is highly likely in day-to-day unposed images. To incorporate the key ideas, we leverage the fundamentals of scene rigidity, multi-scale neural scene representation, and single-image depth prediction. Concretely, the proposed approach makes the camera parameters as learnable in a neural fields-based modeling framework. By assuming per view depth prediction is given up to scale, we constrain the relative pose between successive frames. From the relative poses, absolute camera pose estimation is modeled via a graph-neural network-based multiple motion averaging within the multi-scale neural-fields network, leading to a single loss function. Optimizing the introduced loss function provides camera intrinsic, extrinsic, and image rendering from unposed images. We demonstrate, with examples, that for a unified framework to accurately model multiscale neural scene representation from day-to-day acquired unposed multi-view images, it is equally essential to have precise camera-pose estimates within the scene representation framework. Without considering robustness measures in the camera pose estimation pipeline, modeling for multi-scale aliasing artifacts can be counterproductive. We present extensive experiments on several benchmark datasets to demonstrate the suitability of our approach. Nishant Jain, Suryansh Kumar 0001, Luc Van Gool |
Int. J. Comput. Vis. | 2 |
| 2023 | Enhanced Stable View SynthesisabstractWe introduce an approach to enhance the novel view synthesis from images taken from a freely moving camera. The introduced approach focuses on outdoor scenes where recovering accurate geometric scaffold and camera pose is challenging, leading to inferior results using the state-of-the-art stable view synthesis (SVS) method. SVS and related methods fail for outdoor scenes primarily due to (i) overrelying on the multiview stereo (MVS) for geometric scaffold recovery and (ii) assuming COLMAP computed camera poses as the best possible estimates, despite it being well-studied that MVS 3D reconstruction accuracy is limited to scene disparity and camera-pose accuracy is sensitive to key-point correspondence selection. This work proposes a principled way to enhance novel view synthesis solutions drawing inspiration from the basics of multiple view geometry. By leveraging the complementary behavior of MVS and monocular depth, we arrive at a better scene depth per view for nearby and far points, respectively. Moreover, our approach jointly refines camera poses with image-based rendering via multiple rotation averaging graph optimization. The recovered scene depth and the camera-pose help better view-dependent on-surface feature aggregation of the entire scene. Extensive evaluation of our approach on the popular benchmark dataset, such as Tanks and Temples, shows substantial improvement in view synthesis results compared to the prior art. For instance, our method shows 1.5 dB of PSNR improvement on the Tank and Temples. Similar statistics are observed when tested on other benchmark datasets such as FVS, Mip-NeRF 360, and DTU. Nishant Jain, Suryansh Kumar 0001, Luc Van Gool |
CVPR | 2 |
| 2023 | Single Image Depth Prediction Made Better: A Multivariate Gaussian TakeabstractNeural-network-based single image depth prediction (SIDP) is a challenging task where the goal is to predict the scene's per-pixel depth at test time. Since the problem, by definition, is illposed, the fundamental goal is to come up with an approach that can reliably model the scene depth from a set of training examples. In the pursuit of perfect depth estimation, most existing state-of-the-art learning techniques predict a single scalar depth value per-pixel. Yet, it is well-known that the trained model has accuracy limits and can predict imprecise depth. Therefore, an SIDP approach must be mindful of the expected depth variations in the model's prediction at test time. Accordingly, we introduce an approach that performs continuous modeling of per-pixel depth, where we can predict and reason about the per-pixel depth and its distribution. To this end, we model per-pixel scene depth using a multivariate Gaussian distribution. Moreover, contrary to the existing uncertainty modeling methods—in the same spirit, where per-pixel depth is assumed to be independent, we introduce per-pixel covariance modeling that encodes its depth dependency w.r.t. all the scene points. Unfortunately, per-pixel depth covariance modeling leads to a computationally expensive continuous loss function, which we solve efficiently using the learned low-rank approximation of the overall covariance matrix. Notably, when tested on benchmark datasets such as KITTI, NYU, and SUN-RGB-D, the SIDP model obtained by optimizing our loss function shows state-of-the-art results. Our method's accuracy (named MG) is among the top on the KITTI depth-prediction benchmark leaderboard11http://www.cvlibs.net/datasets/kitti/eval–depth.php?benchmark=depth–prediction. Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool |
CVPR | 2 |
| 2023 | VA-DepthNet: A Variational Approach to Single Image Depth Prediction
Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Luc Van Gool |
ICLR | 2 |
| 2023 | Multi-View Photometric Stereo RevisitedabstractMulti-view photometric stereo (MVPS) is a preferred method for detailed and precise 3D acquisition of an object from images. Although popular methods for MVPS can provide outstanding results, they are often complex to execute and limited to isotropic material objects. To address such limitations, we present a simple, practical approach to MVPS, which works well for isotropic as well as other object material types such as anisotropic and glossy. The proposed approach in this paper exploits the benefit of uncertainty modeling in a deep neural network for a reliable fusion of photometric stereo (PS) and multi-view stereo (MVS) network predictions. Yet, contrary to the recently proposed state-of-the-art, we introduce neural volume rendering methodology for a trustworthy fusion of MVS and PS measurements. The advantage of introducing neural volume rendering is that it helps in the reliable modeling of objects with diverse material types, where existing MVS methods, PS methods, or both may fail. Furthermore, it allows us to work on neural 3D shape representation, which has recently shown outstanding results for many geometric processing tasks. Our suggested new loss function aims to fit the zero level set of the implicit neural function using the most certain MVS and PS network predictions coupled with weighted neural volume rendering cost. The proposed approach shows state-of-the-art results when tested extensively on several benchmark datasets. Berk Kaya, Suryansh Kumar 0001, Carlos Eduardo Porto de Oliveira, Vittorio Ferrari, Luc Van Gool |
WACV | 2 |
| 2022 | Robustifying the Multi-Scale Representation of Neural Radiance Fields
Nishant Jain, Suryansh Kumar 0001, Luc Van Gool |
BMVC | 2 |
| 2022 | Uncertainty-Aware Deep Multi-View Photometric StereoabstractThis paper presents a simple and effective solution to the longstanding classical multi-view photometric stereo (MVPS) problem. It is well-known that photometric stereo (PS) is excellent at recovering high-frequency surface details, whereas multi-view stereo (MVS) can help remove the low-frequency distortion due to PS and retain the global geometry of the shape. This paper proposes an approach that can effectively utilize such complementary strengths of PS and MVS. Our key idea is to combine them suitably while considering the per-pixel uncertainty of their estimates. To this end, we estimate per-pixel surface normals and depth using an uncertainty-aware deep-PS network and deep-MVS network, respectively. Uncertainty modeling helps select reliable surface normal and depth estimates at each pixel which then act as a true representative of the dense surface geometry. At each pixel, our approach either selects or discards deep-PS and deep-MVS network prediction depending on the prediction uncertainty measure. For dense, detailed, and precise inference of the object's surface profile, we propose to learn the implicit neural shape representation via a multilayer perceptron (MLP). Our approach encourages the MLP to converge to a natural zero-level set surface using the confident prediction from deep-PS and deep-MVS networks, providing superior dense surface reconstruction. Extensive experiments on the DiLiGenT-MV benchmark dataset show that our method provides high-quality shape recovery with a much lower memory footprint while outperforming almost all of the existing approaches. Berk Kaya, Suryansh Kumar 0001, Carlos Eduardo Porto de Oliveira, Vittorio Ferrari, Luc Van Gool |
CVPR | 2 |
| 2022 | Generative Flows with Invertible AttentionsabstractFlow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains understudied, while it has made breakthroughs in other domains. To fill the gap, this paper introduces two types of invertible attention mechanisms, i.e., map-based and transformer-based attentions, for both unconditional and conditional generative flows. The key idea is to exploit a masked scheme of these two attentions to learn long-range data dependencies in the context of generative flows. The masked scheme allows for invertible attention modules with tractable Jacobian determinants, enabling its seamless integration at any positions of the flow-based models. The proposed attention mechanisms lead to more efficient generative flows, due to their capability of modeling the long-term data dependencies. Evaluation on multiple image synthesis tasks shows that the proposed attention flows result in efficient models and compare favorably against the state-of-the-art unconditional and conditional generative flows. Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar 0001, Radu Timofte, Luc Van Gool |
CVPR | 3 |
| 2022 | Organic Priors in Non-rigid Structure from Motion
Suryansh Kumar 0001, Luc Van Gool |
ECCV (2) | 1 |
| 2022 | Learning Online Multi-sensor Depth Fusion
Erik Sandström, Martin R. Oswald, Suryansh Kumar 0001, Silvan Weder, Fisher Yu 0001, Cristian Sminchisescu, Luc Van Gool |
ECCV (32) | 3 |
| 2022 | Neural Radiance Fields Approach to Deep Multi-View Photometric StereoabstractWe present a modern solution to the multi-view photometric stereo problem (MVPS). Our work suitably exploits the image formation model in a MVPS experimental setup to recover the dense 3D reconstruction of an object from images. We procure the surface orientation using a photometric stereo (PS) image formation model and blend it with a multi-view neural radiance field representation to recover the object’s surface geometry. Contrary to the previous multi-staged framework to MVPS, where the position, iso-depth contours, or orientation measurements are estimated independently and then fused later, our method is simple to implement and realize. Our method performs neural rendering of multi-view images while utilizing surface normals estimated by a deep photometric stereo network. We render the MVPS images by considering the object’s surface normals for each 3D sample point along the viewing direction rather than explicitly using the density gradient in the volume space via 3D occupancy information. We optimize the proposed neural radiance field representation for the MVPS setup efficiently using a fully connected deep network to recover the 3D geometry of an object. Extensive evaluation on the DiLiGenT-MV benchmark dataset shows that our method performs better than the approaches that perform only PS or only multi-view stereo (MVS) and provides comparable results against the state-of-the-art multistage fusion methods. Berk Kaya, Suryansh Kumar 0001, Francesco Sarno, Vittorio Ferrari, Luc Van Gool |
WACV | 2 |
| 2022 | Neural Architecture Search for Efficient Uncalibrated Deep Photometric StereoabstractWe present an automated machine learning approach for uncalibrated photometric stereo (PS). Our work aims at discovering lightweight and computationally efficient PS neural networks with excellent surface normal accuracy. Unlike previous uncalibrated deep PS networks, which are handcrafted and carefully tuned, we leverage differentiable neural architecture search (NAS) strategy to find uncalibrated PS architecture automatically. We begin by defining a discrete search space for a light calibration network and a normal estimation network, respectively. We then perform a continuous relaxation of this search space and present a gradient-based optimization strategy to find an efficient light calibration and normal estimation network. Directly applying the NAS methodology to uncalibrated PS is not straightforward as certain task-specific constraints must be satisfied, which we impose explicitly. Moreover, we search for and train the two networks separately to account for the Generalized Bas-Relief (GBR) ambiguity. Extensive experiments on the DiLiGenT dataset show that the automatically searched neural architectures performance compares favorably with the state-of-the-art uncalibrated PS methods while having a lower memory footprint. Francesco Sarno, Suryansh Kumar 0001, Berk Kaya, Zhiwu Huang, Vittorio Ferrari, Luc Van Gool |
WACV | 2 |
| 2021 | Uncalibrated Neural Inverse Rendering for Photometric Stereo of General SurfacesabstractThis paper presents an uncalibrated deep neural network framework for the photometric stereo problem. For training models to solve the problem, existing neural network-based methods either require exact light directions or ground-truth surface normals of the object or both. However, in practice, it is challenging to procure both of this information precisely, which restricts the broader adoption of photometric stereo algorithms for vision application. To bypass this difficulty, we propose an uncalibrated neural inverse rendering approach to this problem. Our method first estimates the light directions from the input images and then optimizes an image reconstruction loss to calculate the surface normals, bidirectional reflectance distribution function value, and depth. Additionally, our formulation explicitly models the concave and convex parts of a complex surface to consider the effects of interreflections in the image formation process. Extensive evaluation of the proposed method on the challenging subjects generally shows comparable or better results than the supervised and classical approaches. Berk Kaya, Suryansh Kumar 0001, Carlos Eduardo Porto de Oliveira, Vittorio Ferrari, Luc Van Gool |
CVPR | 2 |
| 2021 | Neural Architecture Search of SPD Manifold NetworksabstractIn this paper, we propose a new neural architecture search (NAS) problem of Symmetric Positive Definite (SPD) manifold networks, aiming to automate the design of SPD neural architectures. To address this problem, we first introduce a geometrically rich and diverse SPD neural architecture search space for an efficient SPD cell design. Further, we model our new NAS problem with a one-shot training process of a single supernet. Based on the supernet modeling, we exploit a differentiable NAS algorithm on our relaxed continuous search space for SPD neural architecture search. Statistical evaluation of our method on drone, action, and emotion recognition tasks mostly provides better results than the state-of-the-art SPD networks and traditional NAS algorithms. Empirical results show that our algorithm excels in discovering better performing SPD network design and provides models that are more than three times lighter than searched by the state-of-the-art NAS algorithms. Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar 0001, Erik Goron Endsjo, Yan Wu 0019, Luc Van Gool |
IJCAI | 3 |
| 2021 | Superpixel Soup: Monocular Dense 3D Reconstruction of a Complex Dynamic SceneabstractThis work addresses the task of dense 3D reconstruction of a complex dynamic scene from images. The prevailing idea to solve this task is composed of a sequence of steps and is dependent on the success of several pipelines in its execution. To overcome such limitations with the existing algorithm, we propose a unified approach to solve this problem. We assume that a dynamic scene can be approximated by numerous piecewise planar surfaces, where each planar surface enjoys its own rigid motion, and the global change in the scene between two frames is as-rigid-as-possible (ARAP). Consequently, our model of a dynamic scene reduces to a soup of planar structures and rigid motion of these local planar structures. Using planar over-segmentation of the scene, we reduce this task to solving a "3D jigsaw puzzle" problem. Hence, the task boils down to correctly assemble each rigid piece to construct a 3D shape that complies with the geometry of the scene under the ARAP assumption. Further, we show that our approach provides an effective solution to the inherent scale-ambiguity in structure-from-motion under perspective projection. We provide extensive experimental results and evaluation on several benchmark datasets. Quantitative comparison with competing approaches shows state-of-the-art performance. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Non-Rigid Structure from Motion: Prior-Free Factorization Method RevisitedabstractA simple prior free factorization algorithm [9] is quite often cited work in the field of Non-Rigid Structure from Motion (NRSfM). The benefit of this work lies in its simplicity of implementation, strong theoretical justification to the motion and structure estimation, and its invincible originality. Despite this, the prevailing view is, that it performs exceedingly inferior to other methods on several benchmark datasets [14], [1]. However, our subtle investigation provides some empirical statistics which made us think against such views. The statistical results we obtained supersedes Dai et al.[9] originally reported results on the benchmark datasets by a significant margin under some elementary changes in their core algorithmic idea [9]. Now, these results not only exposes some unrevealed areas for research in NRSfM but also give rise to new mathematical challenges for NRSfM researchers. We argue that by properly utilizing the well-established assumptions about a non-rigidly deforming shape i.e, it deforms smoothly over frames [27] and it spans a low-rank space, the simple prior-free idea can provide results which is comparable to the best available algorithms. In this paper, we explore some of the hidden intricacies missed by Dai et. al. work [9] and how some elementary measures and modifications can enhance its performance, as high as approx. 18% on the benchmark dataset. The improved performance is justified and empirically verified by extensive experiments on several datasets. We believe our work has both practical and theoretical importance for the development of better NRSfM algorithms. Suryansh Kumar 0001 |
WACV | 1 |
| 2019 | Jumping Manifolds: Geometry Aware Dense Non-Rigid Structure From MotionabstractGiven dense image feature correspondences of a non-rigidly moving object across multiple frames, this paper proposes an algorithm to estimate its 3D shape for each frame. To solve this problem accurately, the recent state-of-the-art algorithm reduces this task to set of local linear subspace reconstruction and clustering problem using Grassmann manifold representation [34]. Unfortunately, their method missed on some of the critical issues associated with the modeling of surface deformations, for e.g., the dependence of a local surface deformation on its neighbors. Furthermore, their representation to group high dimensional data points inevitably introduce the drawbacks of categorizing samples on the high-dimensional Grassmann manifold [32, 31]. Hence, to deal with such limitations with [34], we propose an algorithm that jointly exploits the benefit of high-dimensional Grassmann manifold to perform reconstruction, and its equivalent lower-dimensional representation to infer suitable clusters. To accomplish this, we project each Grassmannians onto a lower-dimensional Grassmann manifold which preserves and respects the deformation of the structure w.r.t its neighbors. These Grassmann points in the lower-dimension then act as a representative for the selection of high-dimensional Grassmann samples to perform each local reconstruction. In practice, our algorithm provides a geometrically efficient way to solve dense NRSfM by switching between manifolds based on its benefit and usage. Experimental results show that the proposed algorithm is very effective in handling noise with reconstruction accuracy as good as or better than the competing methods. Suryansh Kumar 0001 |
CVPR | 1 |
| 2018 | Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian PerspectiveabstractThis paper addresses the task of dense non-rigid structure-front-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sparse feature points or ignore local non-linearities of shape deformations, and thus cannot reliably model complex non-rigid deformations. To address these issues, in this paper, we propose a new approach for dense NRSfM by modeling the problem on a Grassmann manifold. Specifically, we assume the complex non-rigid deformations lie on a union of local linear subspaces both spatially and temporally. This naturally allows for a compact representation of the complex non-rigid deformation over frames. We provide experimental results on several synthetic and real benchmark datasets. The procured results clearly demonstrate that our method, apart from being scalable and more accurate than state-of-the-art methods, is also more robust to noise and generalizes to highly nonlinear deformations. Suryansh Kumar 0001, Anoop Cherian, Yuchao Dai, Hongdong Li |
CVPR | 1 |
| 2017 | Monocular Dense 3D Reconstruction of a Complex Dynamic Scene from Two Perspective FramesabstractThis paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel over-segmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way, we reduce the dynamic reconstruction problem to a “3D jigsaw puzzle ” problem which takes pieces from an unorganized “soup of superpixels". We show that our method provides an effective solution to the inherent relative scale ambiguity in structure-from-motion. Since our method does not assume a template prior, or per-object segmentation, or knowledge about the rigidity of the dynamic scene, it is applicable to a wide range of scenarios. Extensive experiments on both synthetic and real monocular sequences demonstrate the superiority of our method compared with the state-of-the-art methods. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
ICCV | 1 |
| 2017 | Spatio-temporal union of subspaces for multi-body non-rigid structure-from-motion
Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
Pattern Recognit. | 1 |
| 2016 | Multi-Body Non-Rigid Structure-from-MotionabstractIn this paper, we present the first multi-body non-rigid structure-from-motion (SFM) method, which simultaneously reconstructs and segments multiple objects that are undergoing non-rigid deformation over time. Under our formulation, 3D trajectories for each non-rigid object can be well approximated with a sparse affine combination of other 3D trajectories from the same object. The resultant optimization is solved by the alternating direction method of multipliers (ADMM). We demonstrate the efficacy of the proposed method through extensive experiments on both synthetic and real data sequences. Our method outperforms other alternative methods, such as first clustering the 2D feature tracks to groups and then doing non-rigid reconstruction in each group or first conducting 3D reconstruction by using single subspace assumption and then clustering the 3D trajectories into groups. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
3DV | 1 |
| 2014 | Small Object Discovery and Recognition Using Actively Guided RobotabstractIn the field of active perception, object search is a widely studied problem. To search for an object in large rooms, it would be expensive to explore and check each object's similarity with the object of interest. The expense could uncontrollably bloat as the number of objects to be searched increases. If the objects are of the order of a 2-5cm, they appear very small, making it difficult for the present algorithms to recognize them. A general human strategy in such cases is to sparsely identify, from far away (4-6m), if the object of interest is present in the scene. Subsequently, each of the possible objects is analysed from closer proximity to recognize, for further manipulation. In this work, we present a similar framework. We reduce search-space, by identifying existential probability of a small object from a distance followed by a closer 3-D analysis of its point cloud to accurately recognize it. This is achieved by 2-D modelling of the objects using Gaussian Mixture Models followed by recognizing objects using efficient RGB-Depth based algorithm. Sudhanshu Mittal, M. Siva Karthik, Suryansh Kumar 0001, K. Madhava Krishna |
ICPR | 3 |
| 2014 | Markov Random Field based small obstacle discovery over imagesabstractSmall obstacles of the order of 0.5–3cms and homogeneous scenes often pose a problem for indoor mobile robots. These obstacles cannot be clearly distinguished even with the state of the art depth sensors or laser range finders using existing vision based algorithms. With the advent of sophisticated image processing algorithms like SLIC [1] and LSD [9], it is possible to extract rich information from an image which led us to develop a novel architecture to detect very small obstacles on the floor using a monocular camera. This information is further processed using a Markov Random Field based graph cut formalism that precisely segments the floor and detects obstacles which are extremely low. We show robust and accurate obstacle detection and floor segmentation in diverse environments over a large variety of objects found indoors. In our case, low lying obstacles, changing floor patterns and extremely homogeneous environments are properly classified which leads to a drastic decrease in the number of obstacles that may not be classified by existing robotic vision algorithms. Suryansh Kumar 0001, M. Siva Karthik, K. Madhava Krishna |
ICRA | 1 |