K. Madhava Krishna

dblp:90/4844 · also Krishnan Madhava Krishna, Madhava Krishna 0001 · DBLP profile ↗
← Back
124ranked-venue papers
2as first author
43since 2021 · last 2025
0000-0001-7846-7901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 121 · 2 first-author · 42 since 2021Systems, architecture and hardware · 93 · 2 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
abstract
Dynamic scene rendering and reconstruction play a crucial role in computer vision and augmented reality. Recent methods based on 3D Gaussian Splatting (3DGS), have enabled accurate modeling of dynamic urban scenes, but for urban scenes they require both camera and LiDAR data, ground-truth 3D segmentations and motion data in the form of tracklets or pre-defined object templates such as SMPL. In this work, we explore whether a combination of 2D object agnostic priors in the form of depth and point tracking coupled with a signed distance function (SDF) representation for dynamic objects can be used to relax some of these requirements. We present a novel approach that integrates Signed Distance Functions (SDFs) with 3D Gaussian Splatting (3DGS) to create a more robust object representation by harnessing the strengths of both methods. Our unified optimization framework enhances the geometric accuracy of 3D Gaussian splatting and improves deformation modeling within the SDF, resulting in a more adaptable and precise representation. We demonstrate that our method achieves state-of-the-art performance in rendering metrics even without LiDAR data on urban scenes. When incorporating LiDAR, our approach improved further in reconstructing and generating novel views across diverse object categories, without ground-truth 3D motion annotation. Additionally, our method enables various scene editing tasks, including scene decomposition, and scene composition.
Siddharth Tourani, Jayaram Reddy, Akash Kumbar, Satyajit Tourani, Nishant Goyal, K. Madhava Krishna, N. Dinesh Reddy, Muhammad Haris Khan
ICCV6
2025 Da-Vil: Adaptive Dual-Arm Manipulation with Reinforcement Learning and Variable Impedance Control
abstract
Dual-arm manipulation is an area of growing interest in the robotics community. Enabling robots to perform tasks that require the coordinated use of two arms, is essential for complex manipulation tasks such as handling large objects, assembling components, and performing human-like interactions. However, achieving effective dual-arm manipulation is challenging due to the need for precise coordination, dynamic adaptability, and the ability to manage interaction forces between the arms and the objects being manipulated. We propose a novel pipeline that combines the advantages of policy learning based on environment feedback and gradient-based optimization to learn controller gains required for the control outputs. This allows the robotic system to dynamically modulate its impedance in response to task demands, ensuring stability and dexterity in dual-arm operations. We evaluate our pipeline on a trajectory-tracking task involving a variety of large, complex objects with different masses and geometries. The performance is then compared to three other established methods for controlling dual-arm robots, demonstrating superior results. Project page: https://dualarmvil.github.io/Dual-Arm-VIL/
Md Faizal Karim, Shreya Bollimuntha, Mohammed Saad Hashmi, Autrio Das, Gaurav Singh 0012, Srinath Sridhar 0002, Arun Kumar Singh 0001, Nagamanikandan Govindan, K. Madhava Krishna
ICRA9
2025 CrowdSurfer: Sampling Optimization Augmented with Vector-Quantized Variational AutoEncoder for Dense Crowd Navigation
abstract
Navigation amongst densely packed crowds remains a challenge for mobile robots. The complexity increases further if the environment layout changes, making the prior computed global plan infeasible. In this paper, we show that it is possible to dramatically enhance crowd navigation by just improving the local planner. Our approach combines generative modelling with inference-time optimization to generate sophisticated long-horizon local plans at interactive rates. More specifically, we train a Vector Quantized Variational AutoEncoder to learn a prior over the expert trajectory distribution conditioned on the perception input. At run-time, this is used as an initialization for a sampling-based optimizer for further refinement. Our approach does not require any sophisticated prediction of dynamic obstacles and yet provides state-of-theart performance. In particular, we compare against the recent DRL-VO approach [2] and show a 40% improvement in success rate and a 6% improvement in travel time.
Naman Kumar 0003, Antareep Singha, Laksh Nanwani, Dhruv Potdar, Tarun R, Fatemeh Rastgar, Simon Idoko, Arun Kumar Singh 0001, K. Madhava Krishna
ICRA9
2025 AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement
abstract
An embodied agent assisting humans is often asked to complete new tasks, and there may not be sufficient time or labeled examples to train the agent to perform these new tasks. Large Language Models (LLMs) trained on considerable knowledge across many domains can be used to predict a sequence of abstract actions for completing such tasks, although the agent may not be able to execute this sequence due to task-, agent-, or domain-specific constraints. Our framework addresses these challenges by leveraging the generic predictions provided by LLM and the prior domain knowledge encoded in a Knowledge Graph (KG), enabling an agent to quickly adapt to new tasks. The robot also solicits and uses human input as needed to refine its existing knowledge. Based on experimental evaluation in the context of cooking and cleaning tasks in simulation domains, we demonstrate that the interplay between LLM, KG, and human input leads to substantial performance gains compared with just using the LLM. Project website1§Project supported in part by TCS Research India: https://sssshivvvv.github.io/adaptbot/
Shivam Singh, Karthik Swaminathan, Nabanita Dash, Snehasis Banerjee, Mohan Sridharan, K. Madhava Krishna
ICRA7
2025 Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies
abstract
World Model-based Reinforcement Learning (WMRL) enables sample efficient policy learning by reducing the need for online interactions which can potentially be costly and unsafe, especially for autonomous driving. However, existing world models often suffer from low prediction fidelity and compounding one-step errors, leading to policy degradation over long horizons. Additionally, traditional RL policies, often deterministic or single Gaussian-based, fail to capture the multi-modal nature of decision-making in complex driving scenarios. To address these challenges, we propose Imagine-2-Drive, a novel WMRL framework that integrates a high-fidelity world model with a multi-modal diffusion-based policy actor. It consists of two key components: DiffDreamer, a diffusion-based world model that generates future observations simultaneously, mitigating error accumulation, and DPA (Diffusion Policy Actor), a diffusion-based policy that models diverse and multi-modal trajectory distributions. By training DPA within DiffDreamer, our method enables robust policy learning with minimal online interactions. We evaluate our method in CARLA using standard driving benchmarks and demonstrate that it outperforms prior world model baselines, improving Route Completion and Success Rate by 15% and 20% respectively.Project page: https://imagine-2-drive.github.io/
Anant Garg, K. Madhava Krishna
IROS2
2025 Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
abstract
Drivable Free-space prediction is a fundamental and crucial problem in autonomous driving. Recent works have addressed the problem by representing the entire non-obstacle road regions as the free-space. In contrast our aim is to estimate the driving corridors that are a navigable subset of the entire road region. Unfortunately, existing corridor estimation methods directly assume a BEV-centric representation, which is hard to obtain. In contrast, we frame drivable free-space corridor prediction as a pure image perception task, using only monocular camera input. However such a formulation poses several challenges as one doesn’t have the corresponding data for such free-space corridor segments in the image. Consequently, we develop a novel self-supervised approach for free-space sample generation by leveraging future ego trajectories and front-view camera images, making the process of visual corridor estimation dependent on the ego trajectory. We then employ a diffusion process to model the distribution of such segments in the image. However, the existing binary mask-based representation for a segment poses many limitations. Therefore, we introduce ContourDiff, a specialized diffusion-based architecture that denoises over contour points rather than relying on binary mask representations, enabling structured and interpretable free-space predictions. We evaluate our approach qualitatively and quantitatively on both nuScenes and CARLA, demonstrating its effectiveness in accurately predicting safe multimodal navigable corridors in the image.
Tejas S. Stanley, Pranjal Paul, Arun Kumar Singh 0001, K. Madhava Krishna
IROS5
2025 DG16M: A Large-Scale Dataset for Dual-Arm Grasping with Force-Optimized Grasps
abstract
Dual-arm robotic grasping is crucial for handling large objects that require stable and coordinated manipulation. While single-arm grasping has been extensively studied, datasets tailored for dual-arm settings remain scarce. We introduce a large-scale dataset of 16 million dual-arm grasps, evaluated under improved force-closure constraints. Additionally, we develop a benchmark dataset containing 300 objects with approximately 30,000 grasps, evaluated in a physics simulation environment, providing a better grasp quality assessment for dual-arm grasp synthesis methods. Finally, we demonstrate the effectiveness of our dataset by training a Dual-Arm Grasp Classifier network that outperforms the state-of-the-art methods by 15%, achieving higher grasp success rates and improved generalization across objects. Project page: https://dg16m.github.io/DG-16M/
Md Faizal Karim, Mohammed Saad Hashmi, Shreya Bollimuntha, Mahesh Reddy Tapeti, Gaurav Singh 0012, Nagamanikandan Govindan, K. Madhava Krishna
IROS7
2025 SparseLoc: Sparse Open-Set Landmark-based Global Localization for Autonomous Navigation
abstract
Global localization is a critical problem in autonomous navigation, enabling precise positioning without reliance on GPS. Modern techniques often depend on dense LiDAR maps, which, while precise, require extensive storage and computational resources. Alternative approaches have explored sparse maps and learned features, but suffer from poor robustness and generalization. We propose SparseLoc1, a global localization framework that leverages vision-language foundation models to generate sparse, semantic-topometric maps in a zero-shot manner. Our approach combines this representation with Monte Carlo localization enhanced by a novel late optimization strategy for improved pose estimation. By constructing compact yet discriminative maps and refining poses through retrospective optimization, SparseLoc overcomes limitations of existing sparse methods, offering a more efficient and robust solution. Our system achieves over 5× improvement in localization accuracy compared to existing sparse mapping techniques. Despite utilizing only 1/500thof the points used by dense methods, it achieves comparable performance, maintaining average global localization error below 5m and 2° on KITTI. We further demonstrate the practical applicability of our method through cross-sequence localization experiments and downstream navigation tasks.
Pranjal Paul, Vineeth Bhat, Tejas Salian, Mohd. Omama, Krishna Murthy Jatavallabhula, Naveen Arulselvan, K. Madhava Krishna
IROS7
2025 SegMASt3R: Geometry Grounded Segment Matching
abstract
Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features, segment matching captures structured regions, offering greater robustness to occlusions, lighting variations, and viewpoint changes. In this paper, we leverage the spatial understanding of 3D foundation models to tackle wide-baseline segment matching, a challenging setting involving extreme viewpoint shifts. We propose an architecture that uses the inductive bias of these 3D foundation models to match segments across image pairs with up to $180^\circ$ rotation. Extensive experiments show that our approach outperforms state-of-the-art methods, including the SAM2 video propagator and local feature matching methods, by up to 30\% on the AUPRC metric, on ScanNet++ and Replica datasets. We further demonstrate benefits of the proposed model on relevant downstream tasks, including 3D instance mapping and object-relative navigation.
Rohit Jayanti, Swayam Agrawal, Vansh Garg, Siddharth Tourani, Muhammad Haris Khan, Sourav Garg, K. Madhava Krishna
NeurIPS7
2024 Revisit Anything: Visual Place Recognition via Image Segment Retrieval
Kartik Garg, Sai Shubodh Puligilla, Shishir Kolathaya, K. Madhava Krishna, Sourav Garg
ECCV (68)4
2024 Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments†
abstract
Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-of-the-art methods for task anticipation use data-driven deep networks and Large Language Models (LLMs), but they do so at the level of high-level tasks and/or require many training examples. Our framework leverages the generic knowledge of LLMs through a small number of prompts to perform high-level task anticipation, using the anticipated tasks as goals in a classical planning system to compute a sequence of finer-granularity actions that jointly achieve these goals. We ground and evaluate our framework’s abilities in realistic scenarios in the VirtualHome environment and demonstrate a 31% reduction in execution time compared with a system that does not consider upcoming tasks.
Raghav Arora, Shivam Singh, Karthik Swaminathan, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
ICRA9
2024 Talk2BEV: Language-enhanced Bird's-eye View Maps for Autonomous Driving
abstract
This work introduces Talk2BEV, a large vision-language model (LVLM)1interface for bird’s-eye view (BEV) maps commonly used in autonomous driving. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set of object categories and driving scenarios, Talk2BEV eliminates the need for BEV-specific training, relying instead on well-performing pre-trained LVLMs. This enables a single system to cater to a variety of autonomous driving tasks encompassing visual and spatial reasoning, predicting the intents of traffic actors, and decision-making based on visual cues. We extensively evaluate Talk2BEV on a large number of scene understanding tasks that rely on both the ability to interpret freeform natural language queries, and in grounding these queries to the visual context embedded into the language-enhanced BEV map. To enable further research in LVLMs for autonomous driving scenarios, we develop and release Talk2BEV-Bench, a benchmark encompassing 1000 human-annotated BEV scenarios, with more than 20,000 questions and ground-truth responses from the NuScenes dataset. We encourage the reader to view the demos on our project page: https://llmbev.github.io/talk2bev/
Tushar Choudhary, Vikrant Dewangan, Shivam Chandhok, Shubham Priyadarshan, Anushka Jain, Arun Kumar Singh 0001, Siddharth Srivastava 0004, Krishna Murthy Jatavallabhula, K. Madhava Krishna
ICRA9
2024 ATPPNet: Attention based Temporal Point cloud Prediction Network
abstract
Point cloud prediction is an important yet challenging task in the field of autonomous driving. The goal is to predict future point cloud sequences that maintain object structures while accurately representing their temporal motion. These predicted point clouds help in other subsequent tasks like object trajectory estimation for collision avoidance or estimating locations with the least odometry drift. In this work, we present ATPPNet, a novel architecture that predicts future point cloud sequences given a sequence of previous time step point clouds obtained with LiDAR sensor. ATPPNet leverages Conv-LSTM along with channel-wise and spatial attention dually complemented by a 3D-CNN branch for extracting an enhanced spatio-temporal context to recover high quality fidel predictions of future point clouds. We conduct extensive experiments on publicly available datasets and report impressive performance outperforming the existing methods. We also conduct a thorough ablative study of the proposed architecture and provide an application study that highlights the potential of our model for tasks like odometry estimation.
Kaustab Pal, Aditya Sharma 0001, Avinash Sharma 0001, K. Madhava Krishna
ICRA4
2024 EDMP: Ensemble-of-costs-guided Diffusion for Motion Planning
abstract
Classical motion planning for robotic manipulation includes a set of general algorithms that aim to minimize a scene-specific cost of executing a given plan. This approach offers remarkable adaptability, as they can be directly used off-the-shelf for any new scene without needing specific training datasets. However, without a prior understanding of what diverse valid trajectories are and without specially designed cost functions for a given scene, the overall solutions tend to have low success rates within a certain time limit. While deep-learning-based algorithms tremendously improve success rates, they are much harder to adopt without specialized training datasets. We propose EDMP, an Ensemble-of-costs-guided Diffusion for Motion Planning that aims to combine the strengths of classical and deep-learning-based motion planning. Our diffusion-based network is trained on a set of diverse kinematically valid trajectories. Like classical planning, for any new scene at the time of inference, we compute scene-specific costs such as "collision cost" and guide the diffusion to generate valid trajectories that satisfy the scene-specific constraints. Further, instead of a single cost function that may be insufficient in capturing diversity across scenes, we use an ensemble of costs to guide the diffusion process, significantly improving the success rate compared to classical planners. EDMP performs comparably with SOTA deep-learning-based methods while retaining the generalization capabilities primarily associated with classical planners.
Kallol Saha, Vishal Reddy Mandadi, Jayaram Reddy, Ajit Srikanth, Aditya Agarwal, Bipasha Sen, Arun Kumar Singh 0001, K. Madhava Krishna
ICRA8
2024 Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
abstract
With the rise in consumer depth cameras, a wealth of unlabeled RGB-D data has become available. This prompts the question of how to utilize this data for geometric reasoning of scenes. While many RGB-D registration methods rely on geometric and feature-based similarity, we take a different approach. We use cycle-consistent keypoints as salient points to enforce spatial coherence constraints during matching, improving correspondence accuracy. Additionally, we introduce a novel pose block that combines a GRU recurrent unit with transformation synchronization, blending historical and multi-view data. Our approach surpasses previous self-supervised registration methods on ScanNet and 3DMatch, even outperforming some older supervised methods. We also integrate our components into existing methods, showing their effectiveness.
Siddharth Tourani, Jayaram Reddy, Sarvesh Thakur, K. Madhava Krishna, Muhammad Haris Khan, N. Dinesh Reddy
ICRA4
2024 DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions
abstract
Semantic segmentation in adverse weather scenarios is a critical task for autonomous driving systems. While foundation models have shown promise, the need for specialized adaptors becomes evident for handling more challenging scenarios. We introduce DiffPrompter, a novel differentiable visual and latent prompting mechanism aimed at expanding the learning capabilities of existing adaptors in foundation models. Our proposed ∇HFC (High Frequency Components) based image processing block excels particularly in adverse weather conditions, where conventional methods often fall short. Furthermore, we investigate the advantages of jointly training visual and latent prompts, demonstrating that this combined approach significantly enhances performance in out-of-distribution scenarios. Our differentiable visual prompts leverage parallel and series architectures to generate prompts, effectively improving object segmentation tasks in adverse conditions. Through a comprehensive series of experiments and evaluations, we provide empirical evidence to support the efficacy of our approach. Project page: diffprompter.github.io
Sanket Kalwar, Mihir Ungarala, Aaron Monis, Krishna Reddy Konda, Sourav Garg, K. Madhava Krishna
IROS7
2024 Bi-level Trajectory Optimization on Uneven Terrains with Differentiable Wheel-Terrain Interaction Model
abstract
Navigation of wheeled vehicles on uneven terrain necessitates going beyond the 2D approaches for trajectory planning. Specifically, it is essential to incorporate the full 6dof variation of vehicle pose and its associated stability cost in the planning process. To this end, most recent works aim to learn a neural network model to predict vehicle evolution. However, such approaches are data-intensive and fraught with generalization issues.In this paper, we present a purely model-based approach that just requires the digital elevation information of the terrain. Specifically, we express the wheel-terrain interaction and 6dof pose prediction as a non-linear least squares (NLS) problem. As a result, trajectory planning can be viewed as a bi-level optimization. The inner optimization layer predicts the pose on the terrain along a given trajectory, while the outer layer deforms the trajectory itself to reduce the stability and kinematic costs of the pose.We improve the state-of-the-art in the following respects. First, we show that our NLS-based pose prediction closely matches the output of a high-fidelity physics engine. This result, coupled with the fact that we can query gradients of the NLS solver, makes our pose predictor a differentiable wheel-terrain interaction model. We further leverage this differentiability to efficiently solve the proposed bi-level trajectory optimization problem. Finally, we perform extensive experiments and comparisons with a baseline to showcase the effectiveness of our approach in obtaining smooth, stable trajectories.
Amith Manoharan, Aditya Sharma 0001, Himani Belsare, Kaustab Pal, K. Madhava Krishna, Arun Kumar Singh 0001
IROS5
2024 QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
abstract
Robotic tasks such as planning and navigation require a hierarchical semantic understanding of a scene, which could include multiple floors and rooms. Current methods primarily focus on object segmentation for 3D scene understanding. However, such methods struggle to segment out topological regions like "kitchen" in the scene. In this work, we introduce a two-step pipeline to solve this problem. First, we extract a topological map, i.e., floorplan of the indoor scene using a novel multi-channel occupancy representation. Then, we generate CLIP-aligned features and semantic labels for every room instance based on the objects it contains using a self-attention transformer. Our language-topology alignment supports natural language querying, e.g., a "place to cook" locates the "kitchen". We outperform the current state-of-the-art on room segmentation by ~20% and room classification by ~12%. Our detailed qualitative analysis and ablation studies provide insights into the problem of joint structural and semantic 3D scene understanding. Project Page: quest-maps.github.io
Yash Mehan, Kumaraditya Gupta, Rohit Jayanti, Anirudh Govil, Sourav Garg, K. Madhava Krishna
IROS6
2024 Imagine2Servo: Intelligent Visual Servoing with Diffusion-Driven Goal Generation for Robotic Tasks
abstract
Visual servoing, the method of controlling robot motion through feedback from visual sensors, has seen significant advancements with the integration of optical flow-based methods. However, its application remains limited by inherent challenges, such as the necessity for a target image at test time, the requirement of substantial overlap between initial and target images, and the reliance on feedback from a single camera. This paper introduces Imagine2Servo†, an innovative approach leveraging diffusion-based image editing techniques to enhance visual servoing algorithms by generating intermediate goal images. This methodology allows for the extension of visual servoing applications beyond traditional constraints, enabling tasks like long-range navigation and manipulation without predefined goal images. We propose a pipeline that synthesizes subgoal images grounded in the task at hand, facilitating servoing in scenarios with minimal initial and target image overlap and integrating multi-camera feedback for comprehensive task execution. Our contributions demonstrate a novel application of image generation to robotic control, significantly broadening the capabilities of visual servoing systems. Real-world experiments validate the effectiveness and versatility of the Imagine2Servo framework in accomplishing a variety of tasks, marking a notable advancement in the field of visual servoing.
Pranjali Pathre, Gunjan Gupta, M. Nomaan Qureshi, Mandyam Brunda, Samarth Brahmbhatt, K. Madhava Krishna
IROS6
2024 LeGo-Drive: Language-enhanced Goal-oriented Closed-Loop End-to-End Autonomous Driving
abstract
Existing Vision-Language Models (VLMs) produce long-term trajectory waypoints or directly control actions based on their perception input and language prompt. However, these VLMs are not explicitly aware of the constraints imposed by the scene or kinematics of the vehicle. As a result, the generated trajectories or control inputs are likely to be unsafe and/or infeasible. In this paper, we introduce LeGo-Drive†, which aims to address these issues. Our key idea is to use the VLM to just predict a goal location based on the given language command and perception input, which is then fed to a downstream differentiable trajectory optimizer with learnable components. We train the VLM and the trajectory optimizer in an end-to-end fashion using a loss function that captures the ego-vehicle’s ability to reach the predicted goal while satisfying safety and kinematic constraints. The gradients during the back-propagation flow through the optimization layer and make the VLM aware of the planner’s capabilities, making more feasible goal predictions. We compare our end-to-end approach with a decoupled framework where the planner is just used at the inference time to drive to the VLM-predicted goal location and report a goal reaching Success Rate of 81%. We demonstrate the versatility of LeGo-Drive†across various driving scenarios and navigation commands, highlighting its potential for practical deployment in autonomous vehicles.
Pranjal Paul, Anant Garg, Tushar Choudhary, Arun Kumar Singh 0001, K. Madhava Krishna
IROS5
2024 Constrained 6-DoF Grasp Generation on Complex Shapes for Improved Dual-Arm Manipulation
abstract
Efficiently generating grasp poses tailored to specific regions of an object is vital for various robotic manipulation tasks, especially in a dual-arm setup. This scenario presents a significant challenge due to the complex geometries involved, requiring a deep understanding of the local geometry to generate grasps efficiently on the specified constrained regions. Existing methods only explore settings involving table-top/small objects and require augmented datasets to train, limiting their performance on complex objects. We propose CGDF: Constrained Grasp Diffusion Fields, a diffusion-based grasp generative model that generalizes to objects with arbitrary geometries, as well as generates dense grasps on the target regions. CGDF uses a part-guided diffusion approach that enables it to get high sample efficiency in constrained grasping without explicitly training on massive constraint-augmented datasets. We provide qualitative and quantitative comparisons using analytical metrics and in simulation, in both unconstrained and constrained settings to show that our method can generalize to generate stable grasps on complex objects, especially useful for dual-arm manipulation settings, while existing methods struggle to do so. More results, code and an extended version of the paper can be found on the project page: https://constrained-grasp-diffusion.github.io/
Gaurav Singh 0012, Sanket Kalwar, Md Faizal Karim, Bipasha Sen, Nagamanikandan Govindan, Srinath Sridhar 0002, K. Madhava Krishna
IROS7
2024 FinderNet: A Data Augmentation Free Canonicalization aided Loop Detection and Closure technique for Point clouds in 6-DOF separation
abstract
We focus on the problem of LiDAR point cloud based loop detection (or Finding) and closure (LDC) for mobile robots. State-of-the-art (SOTA) methods directly generate learned embeddings from a given point cloud, require large data augmentation, and are not robust to wide viewpoint variations in 6 Degrees-of-Freedom (DOF). Moreover, the absence of strong priors in an unstructured point cloud leads to highly inaccurate LDC. In this original approach, we propose independent roll and pitch canonicalization of point clouds using a common dominant ground plane. We discretize the canonicalized point clouds along the axis perpendicular to the ground plane leads to images similar to digital elevation maps (DEMs), which expose strong spatial priors in the scene. Our experiments show that LDC based on learnt embeddings from such DEMs is not only data efficient but also significantly more robust, and generalizable than the current SOTA. We report an (average precision for loop detection, mean absolute translation/rotation error) improvement of (8.4, 16.7/5.43)% on the KITTI08 sequence, and (11.0, 34.0/25.4)% on GPR10 sequence, over the current SOTA. To further test the robustness of our technique on point clouds in 6-DOF motion we create and opensource a custom dataset called Lidar-UrbanFly Dataset (LUF) which consists of point clouds obtained from a LiDAR mounted on a quadrotor. More details on our website https://gsc2001.github.io/FinderNet/
Sudarshan S. Harithas, Gurkirat Singh, Aneesh Chavan, Sarthak Sharma, Suraj Patni, Chetan Arora 0001, K. Madhava Krishna
WACV7
2023 Canonical Fields: Self-Supervised Learning of Pose-Canonicalized Neural Fields
abstract
Coordinate-based implicit neural networks, or neural fields, have emerged as useful representations of shape and appearance in 3D computer vision. Despite advances, however, it remains challenging to build neural fields for categories of objects without datasets like ShapeNet that provide “canonicalized” object instances that are consistently aligned for their 3D position and orientation (pose). We present Canonical Field Network (CaFi-Net), a self-supervised method to canonicalize the 3D pose of instances from an object category represented as neural fields, specifically neural radiance fields (NeRFs). CaFi-Net directly learns from continuous and noisy radiance fields using a Siamese network architecture that is designed to extract equivariant field features for category-level canonicalization. During inference, our method takes pre-trained neural radiance fields of novel object instances at arbitrary 3D pose and estimates a canonical field with consistent 3D pose across the entire category. Extensive experiments on a new dataset of 1300 NeRF models across 13 object categories show that our method matches or exceeds the performance of 3D point cloud-based methods.
Rohith Agaram, Shaurya Dewan, Rahul Sajnani, Adrien Poulenard, K. Madhava Krishna, Srinath Sridhar 0002
CVPR5
2023 Sequence-Agnostic Multi-Object Navigation
abstract
The Multi-Object Navigation (MultiON) task requires a robot to localize an instance (each) of multiple object classes. It is a fundamental task for an assistive robot in a home or a factory. Existing methods for MultiON have viewed this as a direct extension of Object Navigation (ON), the task of localising an instance of one object class, and are pre-sequenced, i.e., the sequence in which the object classes are to be explored is provided in advance. This is a strong limitation in practical applications characterized by dynamic changes. This paper describes a deep reinforcement learning framework for sequence-agnostic MultiON based on an actor-critic architecture and a suitable reward specification. Our framework leverages past experiences and seeks to reward progress toward individual as well as multiple target object classes. We use photo-realistic scenes from the Gibson benchmark dataset in the AI Habitat 3D simulation environment to experimentally show that our method performs better than a pre-sequenced approach and a state of the art ON method extended to MultiON.
Nandiraju Gireesh, Ahana Datta, Snehasis Banerjee, Mohan Sridharan, Brojeshwar Bhowmick, K. Madhava Krishna
ICRA7
2023 Ground then Navigate: Language-guided Navigation in Dynamic Scenes
abstract
We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable regions corresponding to the textual command. At each timestamp, the model predicts a segmentation mask corresponding to the intermediate or the final navigable region. Our work contrasts with existing efforts in VLN, which pose this task as a node selection problem, given a discrete connected graph corresponding to the environment. We do not assume the availability of such a discretised map. Our work moves towards continuity in action space, provides interpretability through visual feedback and allows VLN on commands requiring finer manoeuvres like “park between the two cars”. Furthermore, we propose a novel meta-dataset CARLA-NAV to allow efficient training and validation. The dataset comprises pre-recorded training sequences and a live environment for validation and testing. We provide extensive qualitative and quantitative em-pirical results to validate the efficacy of the proposed approach. Code is available at https://github.com/kanji95/carla_nav.
Kanishk Jain, Varun Chhangani, Amogh Tiwari, K. Madhava Krishna, Vineet Gandhi
ICRA4
2023 GDIP: Gated Differentiable Image Processing for Object Detection in Adverse Conditions
abstract
Detecting objects under adverse weather and lighting conditions is crucial for the safe and continuous operation of an autonomous vehicle, and remains an unsolved problem. We present a Gated Differentiable Image Processing (GDIP) block, a domain-agnostic network architecture, which can be plugged into existing object detection networks (e.g., Yolo) and trained end-to-end with adverse condition images such as those captured under fog and low lighting. Our pro-posed GDIP block learns to enhance images directly through the downstream object detection loss. This is achieved by learning parameters of multiple image pre-processing (IP) techniques that operate concurrently, with their outputs combined using weights learned through a novel gating mechanism. We further improve GDIP through a multi-stage guidance procedure for progressive image enhancement. Finally, trading off accuracy for speed, we propose a variant of GDIP that can be used as a regularizer for training Yolo, which eliminates the need for GDIP-based image enhancement during inference, resulting in higher throughput and plausible real-world deployment. We demonstrate significant improvement in detection performance over several state-of-the-art methods through quantitative and qualitative studies on synthetic datasets such as PascalVOC, and real-world foggy (RTTS) and low-lighting (ExDark) datasets.
Sanket Kalwar, Aakash Aanegola, Krishna Reddy Konda, Sourav Garg, K. Madhava Krishna
ICRA6
2023 SCARP: 3D Shape Completion in ARbitrary Poses for Improved Grasping
abstract
Recovering full 3D shapes from partial observations is a challenging task that has been extensively addressed in the computer vision community. Many deep learning methods tackle this problem by training 3D shape generation networks to learn a prior over the full 3D shapes. In this training regime, the methods expect the inputs to be in a fixed canonical form, without which they fail to learn a valid prior over the 3D shapes. We propose SCARP, a model that performs Shape C ompletion in ARbitrary Poses. Given a partial pointcloud of an object, SCARP learns a disentangled feature representation of pose and shape by relying on rotationally equivariant pose features and geometric shape features trained using a multi-tasking objective. Unlike existing methods that depend on an external canonicalization method, SCARP performs canonicalization, pose estimation, and shape completion in a single network, improving the performance by 45% over the existing baselines. In this work, we use SCARP for improving grasp proposals on tabletop objects. By completing partial tabletop objects directly in their observed poses, SCARP enables a SOTA grasp proposal network improve their proposals by 71.2% on partial shapes. Project page: https://bipashasen.github.io/scarp
Bipasha Sen, Aditya Agarwal, Gaurav Singh 0012, Brojeshwar Bhowmick, Srinath Sridhar 0002, K. Madhava Krishna
ICRA6
2023 Learning Arc-Length Value Function for Fast Time-Optimal Pick and Place Sequence Planning and Execution
abstract
This paper presents a real-time algorithm for computing the optimal sequence and motion plans for a fixed-base manipulator to pick and place a set of given objects. The optimality is defined in terms of the total execution time of the sequence or its proxy, the arc-length in the joint-space. The fundamental complexity stems from the fact that the optimality metric depends on the joint motion, but the task specification is in the end-effector space. Moreover, mapping between a pair of end-effector positions to the shortest arc-length joint trajectory is not analytic; instead, it entails solving a complex trajectory optimization problem. Existing works ignore this complex mapping and use the Euclidean distance in the end-effector space to compute the sequence. In this paper, we overcome the reliance on the Euclidean distance heuristic by introducing a novel data-driven technique to estimate the optimal arc-length cost in joint space (a.k.a the value function) between two given end-effector positions. We parametrize the value function as a Neural Network and motivate a niche choice for its architecture, inspired by the works on metric learning. The learned value function is then used as an edge cost in a capacitated vehicle routing problem (CVRP) set-up to compute the optimal visitation sequence. Finally, we optimize over the input space of the learnt value function network to propose a novel Inverse Kinematics (IK) algorithm that produces substantially shorter joint arc-length trajectories than existing approaches while executing the computed optimal sequence. We show that our sequence planner, in combination with our proposed IK, offers a substantial improvement in joint arc-length over existing state-of-the-art while maintaining scalability to a large number of objects.
Prajwal Thakur, M. Nomaan Qureshi, Arun Kumar Singh 0001, Y. V. S. Harish, Pushkal Katara, Houman Masnavi, K. Madhava Krishna, Brojeshwar Bhowmick
IJCNN7
2023 Hilbert Space Embedding-Based Trajectory Optimization for Multi-Modal Uncertain Obstacle Trajectory Prediction
abstract
Safe autonomous driving critically depends on how well the ego-vehicle can predict the trajectories of neighboring vehicles. To this end, several trajectory prediction algorithms have been presented in the existing literature. Many of these approaches output a multimodal distribution of obstacle trajectories instead of a single deterministic prediction to account for the underlying uncertainty. However, existing planners cannot handle the multimodality based on just sample-level information of the predictions. With this motivation, this paper proposes a trajectory optimizer that can leverage the distributional aspects of the prediction in a computationally tractable and sample-efficient manner. Our optimizer can work with arbitrarily complex distributions and thus can be used with output distribution represented as a deep neural network. The core of our approach is built on embedding distribution in Reproducing Kernel Hilbert Space (RKHS), which we leverage in two ways. First, we propose an RKHS embedding approach to select probable samples from the obstacle trajectory distribution. Second, we rephrase chance-constrained optimization as distribution matching in RKHS and propose a novel sampling-based optimizer for its solution. We validate our approach with handcrafted and neural network-based predictors trained on real-world datasets and show improvement over the existing stochastic optimization approaches in safety metrics.
Basant Sharma, Aditya Sharma 0001, K. Madhava Krishna, Arun Kumar Singh 0001
IROS3
2023 HyP-NeRF: Learning Improved NeRF Priors using a HyperNetwork
abstract
Neural Radiance Fields (NeRF) have become an increasingly popular representation to capture high-quality appearance and shape of scenes and objects. However, learning generalizable NeRF priors over categories of scenes or objects has been challenging due to the high dimensionality of network weight space. To address the limitations of existing work on generalization, multi-view consistency and to improve quality, we propose HyP-NeRF, a latent conditioning method for learning generalizable category-level NeRF priors using hypernetworks. Rather than using hypernetworks to estimate only the weights of a NeRF, we estimate both the weights and the multi-resolution hash encodings resulting in significant quality gains. To improve quality even further, we incorporate a denoise and finetune strategy that denoises images rendered from NeRFs estimated by the hypernetwork and finetunes it while retaining multiview consistency. These improvements enable us to use HyP-NeRF as a generalizable prior for multiple downstream tasks including NeRF reconstruction from single-view or cluttered scenes and text-to-NeRF. We provide qualitative comparisons and evaluate HyP-NeRF on three tasks: generalization, compression, and retrieval, demonstrating our state-of-the-art results.
Bipasha Sen, Gaurav Singh 0012, Aditya Agarwal, Rohith Agaram, K. Madhava Krishna, Srinath Sridhar 0002
NeurIPS5
2023 CLIPGraphs: Multimodal Graph Networks to Infer Object-Room Affinities
abstract
This paper introduces a novel method for determining the best room to place an object in, for embodied scene rearrangement. While state-of-the-art approaches rely on large language models (LLMs) or reinforcement learned (RL) policies for this task, our approach, CLIPGraphs, efficiently combines commonsense domain knowledge, data-driven methods, and recent advances in multimodal learning. Specifically, it (a) encodes a knowledge graph of prior human preferences about the room location of different objects in home environments, (b) incorporates vision-language features to support multimodal queries based on images or text, and (c) uses a graph network to learn object-room affinities based on embeddings of the prior knowledge and the vision-language features. We demonstrate that our approach provides better estimates of the most appropriate location of objects from a benchmark set of object categories in comparison with state-of-the-art baselines.11Supplementary material and code: https://clipgraphs.github.io
Raghav Arora, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
RO-MAN8
2023 Instance-Level Semantic Maps for Vision Language Navigation
abstract
Humans have a natural ability to perform semantic associations with the surrounding objects in the environment. This allows them to create a mental map of the environment, allowing them to navigate on-demand when given linguistic instructions. A natural goal in Vision Language Navigation (VLN) research is to impart autonomous agents with similar capabilities. Recent works take a step towards this goal by creating a semantic spatial map representation of the environment without any labeled data. However, their representations are limited for practical applicability as they do not distinguish between different instances of the same object. In this work, we address this limitation by integrating instance-level information into spatial map representation using a community detection algorithm and utilizing word ontology learned by large language models (LLMs) to perform open-set semantic associations in the mapping representation. The resulting map representation improves the navigation performance by two-fold (233%) on realistic language commands with instance-specific descriptions compared to the baseline. We validate the practicality and effectiveness of our approach through extensive qualitative and quantitative experiments.
Laksh Nanwani, Anmol Agarwal, Kanishk Jain, Raghav Prabhakar, Aaron Monis, Aditya P. Mathur, Krishna Murthy Jatavallabhula, A. H. Abdul Hafez, Vineet Gandhi, K. Madhava Krishna
RO-MAN10
2022 CCO-VOXEL: Chance Constrained Optimization over Uncertain Voxel-Grid Representation for Safe Trajectory Planning
abstract
We present CCO-VOXEL: the very first chance-constrained optimization (CCO) algorithm that can compute trajectory plans with probabilistic safety guarantees in real-time directly on the voxel-grid representation of the world. CCO-VOXEL maps the distribution over the distance to the closest obstacle to a distribution over collision-constraint violation and computes an optimal trajectory that minimizes the violation probability. Importantly, unlike existing works, we never assume the nature of the sensor uncertainty or the probability distribution of the resulting collision-constraint violations. We leverage the notion of Hilbert Space embedding of distributions and Maximum Mean Discrepancy (MMD) to compute a tractable surrogate for the original chance-constrained optimization problem and employ a combination of A* based graph-search and Cross-Entropy Method for obtaining its minimum. We show tangible performance gain in terms of collision avoidance and trajectory smoothness as a consequence of our probabilistic formulation vis a vis state-of-the-art planning methods that do not account for such non-parametric noise. Finally, we also show how a combination of low-dimensional feature embedding and pre-caching of Kernel Matrices of MMD allow us to achieve real-time performance in simulations as well as in implementations on on-board commodity hardware that controls the quadrotor flight.
Sudarshan S. Harithas, Rishabh Dev Yadav, Arun Kumar Singh 0001, K. Madhava Krishna
ICRA5
2022 Drift Reduced Navigation with Deep Explainable Features
abstract
Modern autonomous vehicles (AVs) often rely on vision, LIDAR, and even radar-based simultaneous localization and mapping (SLAM) frameworks for precise localization and navigation. However, modern SLAM frameworks often lead to unacceptably high levels of drift (i.e., localization error) when AVs observe few visually distinct features or encounter occlusions due to dynamic obstacles. This paper argues that minimizing drift must be a key desiderata in AV motion planning, which requires an AV to take active control decisions to move towards feature-rich regions while also minimizing conventional control cost. To do so, we first introduce a novel data-driven perception module that observes LIDAR point clouds and estimates which features/regions an AV must navigate towards for drift minimization. Then, we introduce an interpretable model predictive controller (MPC) that moves an AV toward such feature-rich regions while avoiding visual occlusions and gracefully trading off drift and control cost. Our experiments on challenging, dynamic scenarios in the state-of-the-art CARLA simulator indicate our method reduces drift up to 76.76% compared to benchmark approaches.
Mohd. Omama, Sundar Sripada V. S., Sandeep Chinchali, Arun Kumar Singh 0001, K. Madhava Krishna
IROS5
2022 IndoLayout: Leveraging Attention for Extended Indoor Layout Estimation from an RGB Image
abstract
In this work, we propose IndoLayout, a novel real-time approach for generating high-quality occupancy maps from an RGB image for indoor scenes. Such occupancy maps are often crucial for path-planning and mapping in indoor environments but are often built using only information contained in the ego view. In contrast, our approach also predicts occupancy values beyond immediately visible regions from just a monocular image, leveraging learnt priors from indoor scenes. Hence, our proposed network can produce a hallucinated, amodal scene layout that includes areas occluded in the RGB image, such as a navigable floor behind a desk. Specifically, we propose a novel architecture that uses self-attention and adversarial learning to vastly improve the quality of the predicted layout. We evaluate our model on several photorealistic indoor datasets and outperform previous relevant work on all metrics that measure layout quality, including newly adopted ones. Finally, we demonstrate the effectiveness of our method by showing significant improvements on the PointNav task over similar approaches using IndoLayout. For more details, please refer to the project page: https://indolayout.github.io/.
Shantanu Singh, Jaidev Shriram, Shaantanu Kulkarni, Brojeshwar Bhowmick, K. Madhava Krishna
IROS5
2021 DRACO: Weakly Supervised Dense Reconstruction And Canonicalization of Objects
abstract
We present DRACO, a method for Dense Reconstruction And Canonicalization of Object shape from one or more RGB images. Canonical shape reconstruction— estimating 3D object shape in a coordinate space canonicalized for scale, rotation, and translation parameters—is an emerging paradigm that holds promise for a multitude of robotic applications. Prior approaches either rely on painstakingly gathered dense 3D supervision, or produce only sparse canonical representations, limiting real-world applicability. DRACO performs dense canonicalization using only weak supervision in the form of camera poses and semantic keypoints at train time. During inference, DRACO predicts dense object-centric depth maps in a canonical coordinate-space, solely using one or more RGB images of an object. Extensive experiments on canonical shape reconstruction and pose estimation show that DRACO is competitive or superior to fully-supervised methods.
Rahul Sajnani, AadilMehdi J. Sanchawala, Krishna Murthy Jatavallabhula, Srinath Sridhar 0002, K. Madhava Krishna
ICRA5
2021 RoRD: Rotation-Robust Descriptors and Orthographic Views for Local Feature Matching
abstract
The use of local detectors and descriptors in typical computer vision pipelines works well until variations in viewpoint and appearance change become extreme. Past research in this area has typically focused on one of two approaches to this challenge: the use of projections into spaces more suitable for feature matching under extreme viewpoint changes, and attempting to learn features that are inherently more robust to viewpoint change. In this paper, we present a novel framework that combines the learning of invariant descriptors through data augmentation and orthographic viewpoint projection. We propose rotation-robust local descriptors, learnt through training data augmentation based on rotation homographies, and a correspondence ensemble technique that combines vanilla feature correspondences with those obtained through rotation-robust features. Using a range of benchmark datasets as well as contributing a new bespoke dataset for this research domain, we evaluate the effectiveness of the proposed approach on key tasks including pose estimation and visual place recognition. Our system outperforms a range of baseline and state-of-the-art techniques, including enabling higher levels of place recognition precision across opposing place viewpoints, and achieves practically useful performance levels even under extreme viewpoint changes. We reduce pose estimation error by 86.72% relative to state of the art.
Udit Singh Parihar, Aniket Gujarathi, Kinal Mehta, Satyajit Tourani, Sourav Garg, Michael Milford, K. Madhava Krishna
IROS7
2021 RTVS: A Lightweight Differentiable MPC Framework for Real-Time Visual Servoing
abstract
Recent data-driven approaches to visual servoing have shown improved performances over classical methods due to precise feature matching and depth estimation. Some recent servoing approaches use a model predictive control (MPC) framework which generalise well to novel environments and are capable of incorporating dynamic constraints, but are computationally intractable in real-time, making it difficult to deploy in real-world scenarios. On the contrary, single-step methods optimise greedily and achieve high servoing rates, but lack the benefits of the MPC multi-step ahead formulation. In this paper, we make the best of both worlds and propose a lightweight visual servoing MPC framework which generates optimal control near real-time at a frequency of 10.52 Hz. This work utilises the differential cross-entropy sampling method for quick and effective control generation along with a lightweight neural network, significantly improving the servoing frequency. We also propose a flow depth normalisation layer which ameliorates the issue of inferior predictions of two view depth from the flow network. We conduct extensive experimentation on the Habitat simulator and show a notable decrease in servoing time in comparison with other approaches that optimise over a time horizon. We achieve the right balance between time and performance for visual servoing in six degrees of freedom (6DoF), while retaining the advantageous MPC formulation. Our code and dataset are publicly available†.
M. Nomaan Qureshi, Pushkal Katara, Harit Pandya, Y. V. S. Harish, AadilMehdi J. Sanchawala, Gourav Kumar, Brojeshwar Bhowmick, K. Madhava Krishna
IROS9
2021 RP-VIO: Robust Plane-based Visual-Inertial Odometry for Dynamic Environments
abstract
Modern visual-inertial navigation systems (VINS) are faced with a critical challenge in real-world deployment: they need to operate reliably and robustly in highly dynamic environments. Current best solutions merely filter dynamic objects as outliers based on the semantics of the object category. Such an approach does not scale as it requires semantic classifiers to encompass all possibly-moving object classes; this is hard to define, let alone deploy. On the other hand, many realworld environments exhibit strong structural regularities in the form of planes such as walls and ground surfaces, which are also crucially static. We present RP-VIO, a monocular visual-inertial odometry system that leverages the simple geometry of these planes for improved robustness and accuracy in challenging dynamic environments. Since existing datasets have a limited number of dynamic elements, we also present a highly-dynamic, photorealistic synthetic dataset for a more effective evaluation of the capabilities of modern VINS systems. We evaluate our approach on this dataset, and three diverse sequences from standard datasets including two real-world dynamic sequences and show a significant improvement in robustness and accuracy over a state-of-the-art monocular visual-inertial odometry system. We also show in simulation an improvement over a simple dynamic-features masking approach. Our code and dataset are publicly available†.
Karnik Ram, Chaitanya Kharyal, Sudarshan S. Harithas, K. Madhava Krishna
IROS4
2021 Grounding Linguistic Commands to Navigable Regions
abstract
Humans have a natural ability to effortlessly comprehend linguistic commands such as “park next to the yellow sedan” and instinctively know which region of the road the vehicle should navigate. Extending this ability to autonomous vehicles is the next step towards creating fully autonomous agents that respond and act according to human commands. To this end, we propose the novel task of Referring Navigable Regions (RNR), i.e., grounding regions of interest for navigation based on the linguistic command. RNR is different from Referring Image Segmentation (RIS), which focuses on grounding an object referred to by the natural language expression instead of grounding a navigable region. For example, for a command “park next to the yellow sedan,” RIS will aim to segment the referred sedan, and RNR aims to segment the suggested parking region on the road. We introduce a new dataset, Talk2Car-RegSeg, which extends the existing Talk2car [1] dataset with segmentation masks for the regions described by the linguistic commands. A separate test split with concise manoeuvre-oriented commands is provided to assess the practicality of our dataset. We benchmark the proposed dataset using a novel transformer-based architecture. We present extensive ablations and show superior performance over baselines on multiple evaluation metrics. A downstream path planner generating trajectories based on RNR outputs confirms the efficacy of the proposed framework.
Nivedita Rufus, Kanishk Jain, Unni Krishnan R. Nair, Vineet Gandhi, K. Madhava Krishna
IROS5
2021 Modular Pipe Climber III with Three-Output Open Differential
abstract
The paper introduces the novel Modular Pipe Climber III with a Three-Output Open Differential (3-OOD) mechanism to eliminate slipping of the tracks due to the changing cross-sections of the pipe. This will be achieved in any orientation of the robot. Previous pipe climbers use three-wheel/track modules, each with an individual driving mechanism to achieve stable traversing. Slipping of tracks is prevalent in such robots when it encounters the pipe turns. Thus, active control of each module’s speed is employed to mitigate the slip, thereby requiring substantial control effort. The proposed pipe climber implements the 3-OOD to address this issue by allowing the robot to mechanically modulate the track speeds as it encounters a turn. The proposed 3-OOD is the first three-output differential to realize the functional abilities of a traditional two-output differential.
Rama Vadapalli, Saharsh Agarwal, Vishnu Kumar, Kartik Suryavanshi, Nagamanikandan Govindan, K. Madhava Krishna
IROS6
2021 Probabilistic Collision Avoidance For Multiple Robots: A Closed Form PDF Approach
abstract
This paper proposes a novel method for reactive multiagent collision avoidance by characterizing the longitudinal and lateral intent uncertainty along a trajectory as a closed-form probability density function. Intent uncertainty is considered as the set of reachable velocities in a planning interval and distributed as a Gaussian distribution over the robot's instantaneous velocity. We utilize the Time Scaled Collision Cone(TSCC) approach, which characterizes the space of instantaneous collision avoidance velocities available to the ego-agent. We introduce intent uncertainty into the characteristic equation of the TSCC to derive the closed-form probability density function, which allows the collision avoidance problem to be rewritten as a deterministic optimization procedure. The formulation also allows the flexibility for the inclusion of confidence intervals for collision avoidance. We thus demonstrate the results and ablation studies of this derived collision avoidance formulation on various confidence intervals.
Josyula Gopala Krishna, Anirudha Ramesh, K. Madhava Krishna
IV3
2021 GCExp: Goal-Conditioned Exploration for Object Goal Navigation
abstract
In this paper, we address the highly challenging problem of object goal navigation. The agent, in an unseen environment, has to perceive its surroundings to identify and navigate towards potential regions where the specified goal category can occur. Rather than developing goal driven exploration policies, we aim to adapt the existing exploration policies that maximize scene coverage to be goal-conditioned. Thus, we propose a standalone scene understanding module to identify potential regions where the goal occurs. We also propose Goal-Conditioned Exploration (GCExp), an algorithm that entails the integration of our novel scene understanding module with any existing exploration policy. We test our solution in photo-realistic simulation environments using state-of-the-art exploration policy, Active Neural Slam [1], and show improved performance over the same on every evaluation metric.
Gulshan Kumar, Narasimhan Sai Shankar, Himansu Didwania, Ruddra Dev Roychoudhury, Brojeshwar Bhowmick, K. Madhava Krishna
RO-MAN6
2020 DFVS: Deep Flow Guided Scene Agnostic Image Based Visual Servoing
abstract
Existing deep learning based visual servoing approaches regress the relative camera pose between a pair of images. Therefore, they require a huge amount of training data and sometimes fine-tuning for adaptation to a novel scene. Furthermore, current approaches do not consider underlying geometry of the scene and rely on direct estimation of camera pose. Thus, inaccuracies in prediction of the camera pose, especially for distant goals, lead to a degradation in the servoing performance. In this paper, we propose a two-fold solution: (i) We consider optical flow as our visual features, which are predicted using a deep neural network. (ii) These flow features are then systematically integrated with depth estimates provided by another neural network using interaction matrix. We further present an extensive benchmark in a photo-realistic 3D simulation across diverse scenes to study the convergence and generalisation of visual servoing approaches. We show convergence for over 3m and 40 degrees while maintaining precise positioning of under 2cm and 1 degree on our challenging benchmark where the existing approaches that are unable to converge for majority of scenarios for over 1.5m and 20 degrees. Furthermore, we also evaluate our approach for a real scenario on an aerial robot. Our approach generalizes to novel scenarios producing precise and robust servoing performance for 6 degrees of freedom positioning tasks with even large camera transformations without any retraining or fine-tuning.
Y. V. S. Harish, Harit Pandya, Ayush Gaud, Shreya Terupally, Narasimhan Sai Shankar, K. Madhava Krishna
ICRA6
2020 Topological Mapping for Manhattan-like Repetitive Environments
abstract
We showcase a topological mapping framework for a challenging indoor warehouse setting. At the most abstract level, the warehouse is represented as a Topological Graph where the nodes of the graph represent a particular warehouse topological construct (e.g. rackspace, corridor) and the edges denote the existence of a path between two neighbouring nodes or topologies. At the intermediate level, the map is represented as a Manhattan Graph where the nodes and edges are characterized by Manhattan properties and as a Pose Graph at the lower-most level of detail. The topological constructs are learned via a Deep Convolutional Network while the relational properties between topological instances are learnt via a Siamese-style Neural Network. In the paper, we show that maintaining abstractions such as Topological Graph and Manhattan Graph help in recovering an accurate Pose Graph starting from a highly erroneous and unoptimized Pose Graph. We show how this is achieved by embedding topological and Manhattan relations as well as Manhattan Graph aided loop closure relations as constraints in the backend Pose Graph optimization framework. The recovery of near ground-truth Pose Graph on real-world indoor warehouse scenes vindicate the efficacy of the proposed framework.
Sai Shubodh Puligilla, Satyajit Tourani, Tushar Vaidya, Udit Singh Parihar, Ravi Kiran Sarvadevabhatla, K. Madhava Krishna
ICRA6
2020 Bi-Convex Approximation of Non-Holonomic Trajectory Optimization
abstract
Autonomous cars and fixed-wing aerial vehicles have the so-called non-holonomic kinematics which non-linearly maps control input to states. As a result, trajectory optimization with such a motion model becomes highly non-linear and non-convex. In this paper, we improve the computational tractability of non-holonomic trajectory optimization by reformulating it in terms of a set of bi-convex cost and constraint functions along with a non-linear penalty. The bi-convex part acts as a relaxation for the non-holonomic trajectory optimization while the residual of the penalty dictates how well its output obeys the non-holonomic behavior. We adopt an alternating minimization approach for solving the reformulated problem and show that it naturally leads to the replacement of the challenging non-linear penalty with a globally valid convex surrogate. Along with the common cost functions modeling goal-reaching, trajectory smoothness, etc., the proposed optimizer can also accommodate a class of non-linear costs for modeling goal-sets, while retaining the bi-convex structure. We benchmark the proposed optimizer against off-the-shelf solvers implementing sequential quadratic programming and interior-point methods and show that it produces solutions with similar or better cost as the former while significantly outperforming the latter. Furthermore, as compared to both off-the-shelf solvers, the proposed optimizer achieves more than 20x reduction in computation time.
Arun Kumar Singh 0001, Raghu Ram Theerthala, Mithun Babu, Unni Krishnan R. Nair, K. Madhava Krishna
ICRA5
2020 Omnidirectional Tractable Three Module Robot
abstract
This paper introduces the Omnidirectional Tractable Three Module Robot for traversing inside complex pipe networks. The robot consists of three omnidirectional modules fixed 120° apart circumferentially which can rotate about their own axis allowing holonomic motion of the robot. The holonomic motion enables the robot to overcome motion singularity when negotiating T-junctions and further allows the robot to arrive in a preferred orientation while taking turns inside a pipe. We have developed a closed-form kinematic model for the robot in the paper and propose the `Motion Singularity Region' that the robot needs to avoid while negotiating T-junction. The design and motion capabilities of the robot are demonstrated both by conducting simulations in MSC ADAMS on a simplified lumped-model of the robot and with experiments on its physical embodiment.
Kartik Suryavanshi, Rama Vadapalli, Ruchita Vucha, K. Madhava Krishna
ICRA5
2020 AutoLay: Benchmarking amodal layout estimation for autonomous driving
abstract
Given an image or a video captured from a monocular camera, amodal layout estimation is the task of predicting semantics and occupancy in bird's eye view. The term amodal implies we also reason about entities in the scene that are occluded or truncated in image space. While several recent efforts have tackled this problem, there is a lack of standardization in task specification, datasets, and evaluation protocols. We address these gaps with AutoLay, a dataset and benchmark for amodal layout estimation from monocular images. AutoLay encompasses driving imagery from two popular datasets: KITTI [1] and Argoverse [2]. In addition to fine-grained attributes such as lanes, sidewalks, and vehicles, we also provide semantically annotated 3D point clouds. We implement several baselines and bleeding edge approaches, and release our data and code.1.
Kaustubh Mani, Narasimhan Sai Shankar, Krishna Murthy Jatavallabhula, K. Madhava Krishna
IROS4
2020 Understanding Dynamic Scenes using Graph Convolution Networks
abstract
We present a novel Multi-Relational Graph Convolutional Network (MRGCN) based framework to model on-road vehicle behaviors from a sequence of temporally ordered frames as grabbed by a moving monocular camera. The input to MRGCN is a multi-relational graph where the graph's nodes represent the active and passive agents/objects in the scene, and the bidirectional edges that connect every pair of nodes are encodings of their Spatio-temporal relations.We show that this proposed explicit encoding and usage of an intermediate spatio-temporal interaction graph to be well suited for our tasks over learning end-end directly on a set of temporally ordered spatial relations. We also propose an attention mechanism for MRGCNs that conditioned on the scene dynamically scores the importance of information from different interaction types.The proposed framework achieves significant performance gain over prior methods on vehicle-behavior classification tasks on four datasets. We also show a seamless transfer of learning to multiple datasets without resorting to fine-tuning. Such behavior prediction methods find immediate relevance in a variety of navigation tasks such as behavior planning, state estimation, and applications relating to the detection of traffic violations over videos.
Sravan Mylavarapu, Mahtab Sandhu, Priyesh Vijayan, K. Madhava Krishna, Balaraman Ravindran, Anoop M. Namboodiri
IROS4
2020 LiDAR guided Small obstacle Segmentation
abstract
Detecting small obstacles on the road is critical for autonomous driving. In this paper, we present a method to reliably detect such obstacles through a multi-modal framework of sparse LiDAR(VLP-16) and Monocular vision. LiDAR is employed to provide additional context in the form of confidence maps to monocular segmentation networks. We show significant performance gains when the context is fed as an additional input to monocular semantic segmentation frameworks. We further present a new semantic segmentation dataset to the community, comprising of over 3000 image frames with corresponding LiDAR observations. The images come with pixel-wise annotations of three classes off-road, road, and small obstacle. We stress that precise calibration between LiDAR and camera is crucial for this task and thus propose a novel Hausdorff distance based calibration refinement method over extrinsic parameters. As a first benchmark over this dataset, we report our results with 73 % instance detection up to a distance of 50 meters on challenging scenarios. Qualitatively by showcasing accurate segmentation of obstacles less than 15 cms at 50m depth and quantitatively through favourable comparisons vis a vis prior art, we vindicate the method's efficacy. Our project and dataset is hosted at https://small-obstacle-dataset.github.io/.
Aasheesh Singh, Aditya Kamireddypalli, Vineet Gandhi, K. Madhava Krishna
IROS4
2020 Towards Accurate Vehicle Behaviour Classification With Multi-Relational Graph Convolutional Networks
abstract
Understanding on-road vehicle behaviour from a temporal sequence of sensor data is gaining in popularity. In this paper, we propose a pipeline for understanding vehicle behaviour from a monocular image sequence or video. A monocular sequence along with scene semantics, optical flow and object labels are used to get spatial information about the object (vehicle) of interest and other objects (semantically contiguous set of locations) in the scene. This spatial information is encoded by a Multi-Relational Graph Convolutional Network (MR-GCN), and a temporal sequence of such encodings is fed to a recurrent network to label vehicle behaviours. The proposed framework can classify a variety of vehicle behaviours to high fidelity on datasets that are diverse and include European, Chinese and Indian on-road scenes. The framework also provides for seamless transfer of models across datasets without entailing re-annotation, retraining and even fine-tuning. We show comparative performance gain over baseline Spatio-temporal classifiers and detail a variety of ablations to showcase the efficacy of the framework.
Sravan Mylavarapu, Mahtab Sandhu, Priyesh Vijayan, K. Madhava Krishna, Balaraman Ravindran, Anoop M. Namboodiri
IV4
2020 Multi-object Monocular SLAM for Dynamic Environments
abstract
In this paper, we tackle the problem of multibody SLAM from a monocular camera. The term multibody, implies that we track the motion of the camera, as well as that of other dynamic participants in the scene. The quintessential challenge in dynamic scenes is unobservability: it is not possible to unambiguously triangulate a moving object from a moving monocular camera. Existing approaches solve restricted variants of the problem, but the solutions suffer relative scale ambiguity (i.e., a family of infinitely many solutions exist for each pair of motions in the scene). We solve this rather intractable problem by leveraging single-view metrology, advances in deep learning, and category-level shape estimation. We propose a multi pose-graph optimization formulation, to resolve the relative and absolute scale factor ambiguities involved. This optimization helps us reduce the average error in trajectories of multiple bodies over real-world datasets, such as KITTI [1]. To the best of our knowledge, our method is the first practical monocular multi-body SLAM system to perform dynamic multi-object and ego localization in a unified framework in metric scale.
Gokul B. Nair, Swapnil Daga, Rahul Sajnani, Anirudha Ramesh, Junaid Ahmed Ansari, Krishna Murthy Jatavallabhula, K. Madhava Krishna
IV7
2020 SROM: Simple Real-time Odometry and Mapping using LiDAR data for Autonomous Vehicles
abstract
In this paper, we present SROM, a novel realtime Simultaneous Localization and Mapping (SLAM) system for autonomous vehicles. The keynote of the paper showcases SROM's ability to maintain localization at low sampling rates or at high linear or angular velocities where most popular LiDAR based localization approaches get degraded fast. We also demonstrate SROM to be computationally efficient and capable of handling high-speed maneuvers. It also achieves low drifts without the need for any other sensors like IMU and/or GPS. Our method has a two-layer structure wherein first, an approximate estimate of the rotation angle and translation parameters are calculated using a Phase Only Correlation (POC) method. Next, we use this estimate as an initialization for a point-to-plane ICP algorithm to obtain fine matching and registration. Another key feature of the proposed algorithm is the removal of dynamic objects before matching the scans. This improves the performance of our system as the dynamic objects can corrupt the matching scheme and derail localization. Our SLAM system can build reliable maps at the same time generating high-quality odometry. We exhaustively evaluated the proposed method in many challenging highways/country/urban sequences from the KITTI dataset and the results demonstrate better accuracy in comparisons to other state-of-the-art methods with reduced computational expense aiding in real-time realizations. We have also integrated our SROM system with our in-house autonomous vehicle and compared it with the state-of-the-art methods like LOAM and LeGO-LOAM.
Nivedita Rufus, Unni Krishnan R. Nair, A. V. S. Sai Bhargav Kumar, Vashist Madiraju, K. Madhava Krishna
IV5
2020 Mono Lay out: Amodal scene layout from a single image
abstract
In this paper, we address the novel, highly challenging problem of estimating the layout of a complex urban driving scenario. Given a single color image captured from a driving platform, we aim to predict the bird's eye view layout of the road and other traffic participants. The estimated layout should reason beyond what is visible in the image, and compensate for the loss of 3D information due to projection. We dub this problem amodal scene layout estimation, which involves hallucinating scene layout for even parts of the world that are occluded in the image. To this end, we present MonoLayout, a deep neural network for realtime amodal scene layout estimation from a single image. We represent scene layout as a multi-channel semantic occupancy grid, and leverage adversarial feature learning to hallucinate " plausible completions for occluded image parts. We extend several state-of-the-art approaches for road-layout estimation and vehicle occupancy estimation in bird's eye view to the amodal setup and thoroughly evaluate against them. By leveraging temporal sensor fusion to generate training labels, we significantly outperform current art over a number of datasets.
Kaustubh Mani, Swapnil Daga, Shubhika Garg, Narasimhan Sai Shankar, Krishna Murthy Jatavallabhula, K. Madhava Krishna
WACV6
2019 Talk to the Vehicle: Language Conditioned Autonomous Navigation of Self Driving Cars
abstract
We propose a novel pipeline that blends encodings from natural language and 3D semantic maps obtained from visual imagery to generate local trajectories that are executed by a low-level controller. The pipeline precludes the need for a prior registered map through a local waypoint generator neural network. The waypoint generator network (WGN) maps semantics and natural language encodings (NLE) to local waypoints. A local planner then generates a trajectory from the ego location of the vehicle (an outdoor car in this case) to these locally generated waypoints while a low-level controller executes these plans faithfully. The efficacy of the pipeline is verified in the CARLA simulator environment as well as on local semantic maps built from real-world KITTI dataset. In both these environments (simulated and real-world) we show the ability of the WGN to generate waypoints accurately by mapping NLE of varying sequence lengths and levels of complexity. We compare with baseline approaches and show significant performance gain over them. And finally, we show real implementations on our electric car verifying that the pipeline lends itself to practical and tangible realizations in uncontrolled outdoor settings. In loop execution of the proposed pipeline that involves repetitive invocations of the network is critical for any such language-based navigation framework. This effort successfully accomplishes this thereby bypassing the need for prior metric maps or strategies for metric level localization during traversal.
Sriram N. N., Tirth Maniar, Jayaganesh Kalyanasundaram, Vineet Gandhi, Brojeshwar Bhowmick, K. Madhava Krishna
IROS6
2019 INFER: INtermediate representations for FuturE pRediction
abstract
In urban driving scenarios, forecasting future trajectories of surrounding vehicles is of paramount importance. While several approaches for the problem have been proposed, the best-performing ones tend to require extremely detailed input representations (e.g. image sequences). As a result, such methods do not generalize to datasets they have not been trained on. In this paper, we propose intermediate representations that are particularly well-suited for future prediction. As opposed to using texture (color) information from images, we condition on semantics and train an autoregressive model to accurately predict future trajectories of traffic participants (vehicles) (see Fig. above). We demonstrate that semantics provide a significant boost over techniques that operate over raw pixel intensities/disparities. Uncharacteristic of state-of-the-art approaches, our representations and models generalize across different sensing modalities (stereo imagery, LiDAR, a combination of both), and also across completely different datasets, collected across several cities, and also across countries where people drive on opposite sides of the road (left-handed vs right-handed driving). Additionally, we demonstrate an application of our approach in multi-object tracking (data association). To foster further research in transferable representations and ensure reproducibility, we release all our code and data.33More qualitative and quantitative results, along with code and data can be found at https://rebrand.ly/INFER-results.
Shashank Srikanth, Junaid Ahmed Ansari, Karnik Ram R., Sarthak Sharma, Krishna Murthy Jatavallabhula, K. Madhava Krishna
IROS6
2019 A Hierarchical Network for Diverse Trajectory Proposals
abstract
Autonomous explorative robots frequently encounter scenarios where multiple future trajectories can be pursued. Often these are cases with multiple paths around an obstacle or trajectory options towards various frontiers. Humans in such situations can inherently perceive and reason about the surrounding environment to identify several possibilities of either manoeuvring around the obstacles or moving towards various frontiers. In this work, we propose a 2 stage Convolutional Neural Network architecture which mimics such an ability to map the perceived surroundings to multiple trajectories that a robot can choose to traverse. The first stage is a Trajectory Proposal Network which suggests diverse regions in the environment which can be occupied in the future. The second stage is a Trajectory Sampling network which provides a finegrained trajectory over the regions proposed by Trajectory Proposal Network. We evaluate our framework in diverse and complicated real life settings. For the outdoor case, we use the KITTI dataset and our own outdoor driving dataset. In the indoor setting, we use an autonomous drone to navigate various scenarios and also a ground robot which can explore the environment using the trajectories proposed by our framework. Our experiments suggest that the framework is able to develop a semantic understanding of the obstacles, open regions and identify diverse trajectories that a robot can traverse. Our comparisons portray the performance gain of the proposed architecture over a diverse set of methods against which it is compared.
Sriram N. N., Gourav Kumar, Abhay Singh, M. Siva Karthik, Saket Saurav, Brojeshwar Bhowmick, K. Madhava Krishna
IV7
2019 Motion Planning Framework for Autonomous Vehicles: A Time Scaled Collision Cone Interleaved Model Predictive Control Approach
abstract
Planning frameworks for autonomous vehicles must be robust and computationally efficient for real time realization. At the same time, they should accommodate the unpredictable behavior of the other participants and produce safe trajectories. In this paper, we present a computationally efficient hierarchical planning framework for autonomous vehicles that can generate safe trajectories in complex driving scenarios, which are commonly encountered in urban traffic settings. The first level of the proposed framework constructs a Model Predictive Control(MPC)routine using an efficient difference of convex programmingapproach, that generates smooth and collision-free trajectories. The constraints on curvature and road boundaries are seamlessly integrated into this optimization routine. The second layer is mainly responsible to handle the unpredictable behaviors that are typically exhibited by the other participants of traffic. It is built along the lines of time scaled collision cone(TSCC)which optimize for the velocities along the trajectory to handle such disturbances. We additionally show that our framework maintains optimal balance between temporal and path deviations while executing safe trajectories. To demonstrate the efficacy of the presented framework we validated it in extensive simulations in different driving scenarios like over taking, lane merging and jaywalking among many dynamic and static obstacles.
Raghu Ram Theerthala, A. V. S. Sai Bhargav Kumar, Mithun Babu, Phani-Teja Singamaneni, K. Madhava Krishna
IV5
2019 Probabilistic obstacle avoidance and object following: An overlap of Gaussians approach
abstract
Autonomous navigation and obstacle avoidance are core capabilities that enable robots to execute tasks in the real world. We propose a new approach to collision avoidance that accounts for uncertainty in the states of the agent and the obstacles. We first demonstrate that measures of entropy- used in current approaches for uncertainty-aware obstacle avoidance-are an inappropriate design choice. We then propose an algorithm that solves an optimal control sequence with a guaranteed risk bound, using a measure of overlap between the two distributions that represent the state of the robot and the obstacle, respectively. Furthermore, we provide closed form expressions that can characterize the overlap as a function of the control input. The proposed approach enables model-predictive control framework to generate bounded-confidence control commands. An extensive set of simulations have been conducted in various constrained environments in order to demonstrate the efficacy of the proposed approach over the prior art. We demonstrate the usefulness of the proposed scheme under tight spaces where computing risk-sensitive control maneuvers is vital. We also show how this framework generalizes to other problems, such as object-following.
Dhaivat Bhatt, Akash Garg, Bharath Gopalakrishnan, K. Madhava Krishna
RO-MAN4
2019 PIVO: Probabilistic Inverse Velocity Obstacle for Navigation under Uncertainty
abstract
In this paper, we present an algorithmic framework which computes the collision-free velocities for the robot in a human shared dynamic and uncertain environment. We extend the concept of Inverse Velocity Obstacle (IVO) to a probabilistic variant to handle the state estimation and motion uncertainties that arise due to the other participants of the environment. These uncertainties are modeled as non-parametric probability distributions. In our PIVO: Probabilistic Inverse Velocity Obstacle, we propose the collision-free navigation as an optimization problem by reformulating the velocity conditions of IVO as chance constraints that takes the uncertainty into account. The space of collision-free velocities that result from the presented optimization scheme are associated to a confidence measure as a specified probability. We demonstrate the efficacy of our PIVO through numerical simulations and demonstrating its ability to generate safe trajectories under highly uncertain environments.
P. S. Naga Jyotish, Yash Goel, A. V. S. Sai Bhargav Kumar, K. Madhava Krishna
RO-MAN4
2018 MergeNet: A Deep Net Architecture for Small Obstacle Discovery
abstract
We present here, a novel network architecture called MergeNet for discovering small obstacles for on-road scenes in the context of autonomous driving. The basis of the architecture rests on the central consideration of training with less amount of data since the physical setup and the annotation process for small obstacles is hard to scale. For making effective use of the limited data, we propose a multi-stage training procedure involving weight-sharing, separate learning of low and high level features from the RGBD input and a refining stage which learns to fuse the obtained complementary features. The model is trained and evaluated on the Lost and Found dataset and is able to achieve state-of-art results with just 135 images in comparison to the 1000 images used by the previous benchmark. Additionally, we also compare our results with recent methods trained on 6000 images and show that our method achieves comparable performance with only 1000 training samples.
Krishnam Gupta, Syed Ashar Javed, Vineet Gandhi, K. Madhava Krishna
ICRA4
2018 Constructing Category-Specific Models for Monocular Object-SLAM
abstract
We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available. To alleviate the need for huge amounts of labeled data, we develop a rendering pipeline that enables synthesis of large datasets from a limited amount of manually labeled data. Using data thus synthesized, we learn category-level models for object deformations in 3D, as well as discriminative object features in 2D. These category models are instance-independent and aid in the design of object landmark observations that can be incorporated into a generic monocular SLAM framework. Where typical object-SLAM approaches usually solve only for object and camera poses, we also estimate object shape on-the-fty, allowing for a wide range of objects from the category to be present in the scene. Moreover, since our 2D object features are learned discriminatively, the proposed object-SLAM system succeeds in several scenarios where sparse feature-based monocular SLAM fails due to insufficient features or parallax. Also, the proposed category-models help in object instance retrieval, useful for Augmented Reality (AR) applications. We evaluate the proposed framework on multiple challenging real-world scenes and show - to the best of our knowledge - first results of an instance-independent monocular object-SLAM system and the benefits it enjoys over feature-based SLAM methods.
Parv Parkhiya, Rishabh Khawad, Krishna Murthy Jatavallabhula, Brojeshwar Bhowmick, K. Madhava Krishna
ICRA5
2018 Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking
abstract
This paper introduces geometry and object shape and pose costs for multi-object tracking in urban driving scenarios. Using images from a monocular camera alone, we devise pairwise costs for object tracks, based on several 3D cues such as object pose, shape, and motion. The proposed costs are agnostic to the data association method and can be incorporated into any optimization framework to output the pairwise data associations. These costs are easy to implement, can be computed in real-time, and complement each other to account for possible errors in a tracking-by-detection framework. We perform an extensive analysis of the designed costs and empirically demonstrate consistent improvement over the state-of-the-art under varying conditions that employ a range of object detectors, exhibit a variety in camera and object motions, and, more importantly, are not reliant on the choice of the association framework. We also show that, by using the simplest of associations frameworks (two-frame Hungarian assignment), we surpass the state-of-the-art in multi-object-tracking on road scenes. More qualitative and quantitative results can be found at https://junaidcs032.github.io/Geometry_ObjectShape_MOT/.
Sarthak Sharma, Junaid Ahmed Ansari, Krishna Murthy Jatavallabhula, K. Madhava Krishna
ICRA4
2018 The Earth Ain't Flat: Monocular Reconstruction of Vehicles on Steep and Graded Roads from a Moving Camera
abstract
Accurate localization of other traffic participants is a vital task in autonomous driving systems. State-of-the-art systems employ a combination of sensing modalities such as RGB cameras and LiDARs for localizing traffic participants, but monocular localization demonstrations have been confined to plain roads. We demonstrate - to the best of our knowledge - the first results for monocular object localization and shape estimation on surfaces that are non-coplanar with the moving ego vehicle mounted with a monocular camera. We approximate road surfaces by local planar patches and use semantic cues from vehicles in the scene to initialize a local bundle-adjustment like procedure that simultaneously estimates the 3D pose and shape of the vehicles, and the orientation of the local ground plane on which the vehicle stands. We also demonstrate that our approach transfers from synthetic to real data, without any hyperparameter-/fine-tuning. We evaluate the proposed approach on the KITTI and SYNTHIA-SF benchmarks, for a variety of road plane configurations. The proposed approach significantly improves the state-of-the-art for monocular object localization on arbitrarily-shaped roads.
Junaid Ahmed Ansari, Sarthak Sharma, Anshuman Majumdar, Krishna Murthy Jatavallabhula, K. Madhava Krishna
IROS5
2018 CalibNet: Geometrically Supervised Extrinsic Calibration using 3D Spatial Transformer Networks
abstract
3D LiDARs and 2D cameras are increasingly being used alongside each other in sensor rigs for perception tasks. Before these sensors can be used to gather meaningful data, however, their extrinsics (and intrinsics) need to be accurately calibrated, as the performance of the sensor rig is extremely sensitive to these calibration parameters. A vast majority of existing calibration techniques require significant amounts of data and/or calibration targets and human effort, severely impacting their applicability in large-scale production systems. We address this gap with CalibNet: a geometrically supervised deep network capable of automatically estimating the 6-DoF rigid body transformation between a 3D LiDAR and a 2D camera in real-time. CalibNet alleviates the need for calibration targets, thereby resulting in significant savings in calibration efforts. During training, the network only takes as input a LiDAR point cloud, the corresponding monocular image, and the camera calibration matrix K. At train time, we do not impose direct supervision (i.e., we do not directly regress to the calibration parameters, for example). Instead, we train the network to predict calibration parameters that maximize the geometric and photometric consistency of the input images and point clouds. CalibNet learns to iteratively solve the underlying geometric problem and accurately predicts extrinsic calibration parameters for a wide range of mis-calibrations, without requiring retraining or domain adaptation. The project page is hosted at https://epiception.github.io/CalibNet.
Ganesh Iyer, Karnik Ram R., Krishna Murthy Jatavallabhula, K. Madhava Krishna
IROS4
2018 Towards View-Invariant Intersection Recognition from Videos using Deep Network Ensembles
abstract
This paper strives to answer the following question: Is it possible to recognize an intersection when seen from different road segments that constitute the intersection? An intersection or a junction typically is a meeting point of three or four road segments. Its recognition from a road segment that is transverse to or 180 degrees apart from its previous sighting is an extremely challenging and yet a very relevant problem to be addressed from the point of view of both autonomous driving as well as loop detection. This paper formulates this as a problem of video recognition and proposes a novel LSTM based Siamese style deep network for video recognition. For what is indeed a challenging problem and the limited annotated dataset available we show competitive results of recognizing intersections when approached from diverse viewpoints or road segments. Specifically, we tabulate effective recognition accuracy even as the approaches to the intersection being compared are disparate both in terms of viewpoints and weather/illumination conditions. We show competitive results on both synthetic yet highly realistic data mined from the gaming platform GTA as well as on real world data made available through Mapillary.
Gunshi Gupta, Avinash Sharma 0001, K. Madhava Krishna
IROS4
2018 Image Based Visual Servoing for Tumbling Objects
abstract
Objects in space often exhibit a tumbling motion around the major inertial axis. In this paper, we address the image based visual servoing of a robotic system towards an uncooperative tumbling object. In contrast to previous approaches that require explicit reconstruction of the object and an estimation of its velocity, we propose a novel controller that is able to minimize the feature error directly in image space. This is achieved by observing that the feature points on the tumbling object follow a circular path around the axis of rotation and their projection creates an elliptical track in the image plane. Our controller minimizes the error between this elliptical track and the desired features, such that at the desired pose the features lie on the circumference of the ellipse. The effectiveness of our framework is exhibited by implementing the algorithm in simulation as well on a mobile robot.
P. Mithun, Harit Pandya, Ayush Gaud, Suril Vijaykumar Shah, K. Madhava Krishna
IROS5
2018 Overtaking Maneuvers in Simulated Highway Driving using Deep Reinforcement Learning
abstract
Most methods that attempt to tackle the problem of Autonomous Driving and overtaking usually try to either directly minimize an objective function or iteratively in a Reinforcement Learning like framework to generate motor actions given a set of inputs. We follow a similar trend but train the agent in a way similar to a curriculum learning approach where the agent is first given an easier problem to solve, followed by a harder problem. We use Deep Deterministic Policy Gradients to learn overtaking maneuvers for a car, in presence of multiple other cars, in a simulated highway scenario. The novelty of our approach lies in the training strategy used where we teach the agent to drive in a manner similar to the way humans learn to drive and the fact that our reward function uses only the raw sensor data at the current time step. This method, which resembles a curriculum learning approach is able to learn smooth maneuvers, largely collision free, wherein the agent overtakes all other cars, independent of the track and number of cars in the scene.
Meha Kaushik, Vignesh Prasad, K. Madhava Krishna, Balaraman Ravindran
Intelligent Vehicles Symposium3
2018 A Novel Lane Merging Framework with Probabilistic Risk based Lane Selection using Time Scaled Collision Cone
abstract
Conventionally, planning frameworks for autonomous vehicles consider large safety margins and pre- defined paths for performing the merge maneuvers. These considerations often increase the wait time at the intersec- tions leading to traffic disruption. In this paper, we present a motion planning framework for autonomous vehicles to perform merge maneuver in dense traffic. Our framework is divided into a two-layer structure, Lane Selection layer and Scale optimization layer. The Lane Selection layer computes the likelihood of collision along the lanes. This likelihood represents the collision risk associated with each lane and is used for lane selection. Subsequently, the Scale optimization layer solves the time scaled collision cone (TSCC) constraint re- actively for collision-free velocities. Our framework guarantees a collision-free merging even in dense traffic with minimum disruption. Furthermore, we show the simulation results in different merging scenarios to demonstrate the efficacy of our framework.
A. V. S. Sai Bhargav Kumar, Adarsh Modh, Mithun Babu, Bharath Gopalakrishnan, K. Madhava Krishna
Intelligent Vehicles Symposium5
2018 Fast Multi Model Motion Segmentation on Road Scenes
abstract
We propose a novel motion clustering formulation over spatio-temporal depth images obtained from stereo sequences that segments multiple motion models in the scene in an unsupervised manner. The motion models are obtained at frame rates that compete with the speed of the stereo depth computation. This is possible due to a decoupling framework that first delineates spatial clusters and subsequently assigns motion labels to each of these cluster with analysis of a novel motion graph model. A principled computation of the weights of the motion graph that signifies the relative shear and stretch between possible clusters lends itself to a high fidelity segmentation of the motion models in the scene. The fidelity is vindicated through accuracies reaching 89.61% on KITTI and complex native sequences.
Mahtab Sandhu, Nazrul Haque, Avinash Sharma 0001, K. Madhava Krishna, Shanti Medasani
Intelligent Vehicles Symposium4
2017 Reconstructing vehicles from a single image: Shape priors for road scene understanding
abstract
We present an approach for reconstructing vehicles from a single (RGB) image, in the context of autonomous driving. Though the problem appears to be ill-posed, we demonstrate that prior knowledge about how 3D shapes of vehicles project to an image can be used to reason about the reverse process, i.e., how shapes (back-)project from 2D to 3D. We encode this knowledge in shape priors, which are learnt over a small keypoint-annotated dataset. We then formulate a shape-aware adjustment problem that uses the learnt shape priors to recover the 3D pose and shape of a query object from an image. For shape representation and inference, we leverage recent successes of Convolutional Neural Networks (CNNs) for the task of object and keypoint localization, and train a novel cascaded fully-convolutional architecture to localize vehicle keypoints in images. The shape-aware adjustment then robustly recovers shape (3D locations of the detected keypoints) while simultaneously filling in occluded keypoints. To tackle estimation errors incurred due to erroneously detected keypoints, we use an Iteratively Re-weighted Least Squares (IRLS) scheme for robust optimization, and as a by-product characterize noise models for each predicted keypoint. We evaluate our approach on autonomous driving benchmarks, and present superior results to existing monocular, as well as stereo approaches.
Krishna Murthy Jatavallabhula, G. V. Sai Krishna, Falak Chhaya, K. Madhava Krishna
ICRA4
2017 Exploring convolutional networks for end-to-end visual servoing
abstract
Present image based visual servoing approaches rely on extracting hand crafted visual features from an image. Choosing the right set of features is important as it directly affects the performance of any approach. Motivated by recent breakthroughs in performance of data driven methods on recognition and localization tasks, we aim to learn visual feature representations suitable for servoing tasks in unstructured and unknown environments. In this paper, we present an end-to-end learning based approach for visual servoing in diverse scenes where the knowledge of camera parameters and scene geometry is not available a priori. This is achieved by training a convolutional neural network over color images with synchronised camera poses. Through experiments performed in simulation and on a quadrotor, we demonstrate the efficacy and robustness of our approach for a wide range of camera poses in both indoor as well as outdoor environments.
Aseem Saxena, Harit Pandya, Gourav Kumar, Ayush Gaud, K. Madhava Krishna
ICRA5
2017 Detecting, localizing, and recognizing trees with a monocular MAV: Towards preventing deforestation
abstract
We propose a novel pipeline for detecting, localizing, and recognizing trees with a quadcoptor equipped with monocular camera. The quadcoptor flies in an area of semidense plantation filled with many trees of more than 5 meter in height. Trees are detected on a per frame basis using state of the art Convolutional Neural Networks inspired by recent rapid advancements showcased in Deep Learning literature. Once detected, the trees are tagged with a GPS coordinate through our global localizing and positioning framework. Further the localized trees are segmented, characterized by feature descriptors, and stored in a database by their GPS coordinates. In a subsequent run in the same area, the trees that get detected are queried to the database and get associated with the trees in the database. The association problem is posed as a dynamic programming problem and the optimal association is inferred. The algorithm has been verified in various zones in our campus infested with trees with varying density on the Bebop 2 drone equipped with omnidirectional vision. High percentage of successful recognition and association of the trees between two or more runs is the cornerstone of this effort. The proposed method is also able to identify if trees are missing from their expected GPS tagged locations thereby making it possible to immediately alert concerned authorities about possible unlawful felling of trees. We also propose a novel way of obtaining dense disparity map for quadcopter with monocular camera.
Utsav Shah, Rishabh Khawad, K. Madhava Krishna
ICRA3
2017 Detachable modular robot capable of cooperative climbing and multi agent exploration
abstract
At the cross section of the fields of Uneven Terrain Navigation and Multi Agent Systems (MAS), in this work, a Detachable Compliant Modular Robot (DCMR) which can perform concurrent scene exploration by detaching into numerous parts, while preserving its ability to climb stairs is proposed and built. A spring is designed and used in the modular robot taking the worst-case-scenario of stairs encountered in an urban setting. In addition to the actuators at the wheels, an additional set of actuators per module are introduced to enable the detachment and re-attachment. The design additions and their trade-offs are discussed. Potential applications are presented with special focus on improving coverage of a map with obstacles/slabs large enough to merit exploration by climbing them. The problem of turning in crammed spaces is solved using the ability to detach of DCMR. The detaching & re-attaching capability, and stair climbing of the composite modular robot are demonstrated through experimentation using the prototype.
Sri Harsha Turlapati, K. Madhava Krishna, Suril Vijaykumar Shah
ICRA3
2017 Have i reached the intersection: A deep learning-based approach for intersection detection from monocular cameras
abstract
Long-short term memory networks(LSTM) models have shown considerable performance on variety of problems dealing with sequential data. In this paper, we propose a variant of Long-Term Recurrent Convolutional Network(LRCN) to detect road intersection. We call this network as IntersectNet. We pose road intersection detection as binary classification task over sequence of frames. The model combines deep hierarchical visual feature extractor with recurrent sequence model. The model is end to end trainable with capability of capturing the temporal dynamics of the system. We exploit this capability to identify road intersection in a sequence of temporally consistent images. The model has been rigorously trained and tested on various different datasets. We think that our findings could be useful to model behavior of autonomous agent in the real-world.
Dhaivat Bhatt, Danish Sodhi, Arghya Pal, Vineeth N. Balasubramanian, K. Madhava Krishna
IROS5
2017 Multi-trajectory pose correspondences using scale-dependent topological analysis of pose-graphs
abstract
This paper considers the problem of finding pose matches between trajectories of multiple robots in their respective coordinate frames or equivalent matches between trajectories obtained during different sessions. Pose correspondences between trajectories are mediated by common landmarks represented in a topological map lacking distinct metric coordinates. Despite such lack of explicit metric level associations, we mine preliminary pose level correspondences between trajectories through a novel multi-scale heat-kernel descriptor and correspondence graph framework. These serve as an improved initialization for ICP (Iterative Closest Point) to yield dense pose correspondences. We perform extensive analysis of the proposed method under varying levels of pose and landmark noise and showcase its superiority in obtaining pose matches in comparison with standard ICP like methods. To the best of our knowledge, this is the first work of the kind that brings in elements from spectral graph theory to solve the problem of pose correspondences in a multi-robotic setting and differentiates itself from other works.
Sayantan Datta, Avinash Sharma 0001, K. Madhava Krishna
IROS3
2017 PRVO: Probabilistic Reciprocal Velocity Obstacle for multi robot navigation under uncertainty
abstract
We present PRVO, a probabilistic variant of Reciprocal Velocity Obstacle (RVO) for decentralized multi-robot navigation under uncertainty. PRVO characterizes the space of velocities that would allow each robot to fulfill its share in collision avoidance with a specified probability. PRVO is modeled as chance constraints over the velocity level constraints defined by RVO and takes into account the uncertainty associated with both state estimation as well as the actuation of each robot. Since chance constraints are in general computationally intractable, we propose a series of reformulations which when combined with time scaling based concepts leads to a closed form characterization of solution space of PRVO for a given probability of collision avoidance. We validate our formulation through numerical simulations in which we highlight the advantages of PRVO over the related existing formulations.
Bharath Gopalakrishnan, Arun Kumar Singh 0001, Meha Kaushik, K. Madhava Krishna, Dinesh Manocha
IROS4
2017 Pose induction for visual servoing to a novel object instance
abstract
Present visual servoing approaches are instance specific i.e. they control camera motion between two views of the same object. However, in practical scenarios where a robot is required to handle various instances of a category, classical visual servoing techniques are less suitable. We formulate across instance visual servoing as a pose induction and pose alignment problem. Initially, the desired pose given for any known instance is transferred to the novel instance through pose induction. Then the pose alignment problem is solved by estimating the current pose using the part aware keypoints reconstruction followed by a pose based visual servoing (PBVS) iteration. To tackle large variation in appearance across object instances in a category, we employ visual features that uniquely correspond to locations of object's parts in images. These part-aware keypoints are learned from annotated images using a convolutional neural network (CNN). Advantages of using such part-aware semantics are two-fold. Firstly, it conceals the illumination and textural variations from the visual servoing algorithm. Secondly, semantic keypoints enables us to match descriptors across instances accurately. We validate the efficacy of our approach through experiments in simulation as well as on a quadcopter. Our approach results in acceptable desired camera pose and smooth velocity profile. We also show results for large camera transformations with no overlap between current and desired pose for 3D objects, which is desirable in servoing context.
Gourav Kumar, Harit Pandya, Ayush Gaud, K. Madhava Krishna
IROS4
2017 Shape priors for real-time monocular object localization in dynamic environments
abstract
Reconstruction of dynamic objects in a scene is a highly challenging problem in the context of SLAM. In this paper, we present a real-time monocular object localization system that estimates the shape and pose of dynamic objects in real-time, using video frames captured from a moving monocular camera. Although the problem seems to be ill-posed, we demonstrate that, by incorporating prior knowledge of the object category, we can obtain more detailed instance-level reconstructions. As opposed to earlier object model specifications, the proposed shape-prior model leads to the formulation of a Bundle Adjustment-like optimization problem for simultaneous shape and pose estimation. Leveraging recent successes of Convolutional Neural Networks (CNNs) for object keypoint localization, we present a CNN architecture that performs precise keypoint localization. We then demonstrate how these keypoints can be used to recover 3D object properties, while accounting for any 2D localization errors and self-occlusion. We show significant performance improvements compared to state-of-the-art monocular competitors for 2D keypoint detection, as well as 3D localization and reconstruction of dynamic objects.
Krishna Murthy Jatavallabhula, Sarthak Sharma, K. Madhava Krishna
IROS3
2017 COCrIP: Compliant OmniCrawler in-pipeline robot
abstract
This paper presents a modular in-pipeline climbing robot with a novel compliant foldable OmniCrawler mechanism. The robot has a series of 3 compliant foldable OmniCrawler modules interconnected by links via passive joints. The circular cross-section of the module enables a holonomic motion to facilitate the alignment of the robot in the direction of bends. Additionally, the crawler mechanism provides a fair amount of traction, even on slippery pipe surfaces. These advantages of crawler modules have been further augmented by incorporating active compliance in the module, which helps to negotiate sharp bends in small diameter pipes. Introducing compliance in the crawler module with a single chain-lugs assembly is the the key novelty of this design. For the desirable pipe diameter and curvature of the bends, the spring stiffness value for each passive joint is determined by formulating a constrained optimization problem using the quasi-static model of the robot. Moreover, a minimum friction coefficient value between the module-pipe surface which can be vertically climbed by the robot without slipping is estimated. The numerical simulation results have further been validated by experiments on real robot prototype.
Enna Sachdeva, K. Madhava Krishna
IROS4
2016 Monocular reconstruction of vehicles: Combining SLAM with shape priors
abstract
Reasoning about objects in images and videos using 3D representations is re-emerging as a popular paradigm in computer vision. Specifically, in the context of scene understanding for roads, 3D vehicle detection and tracking from monocular videos still needs a lot of attention to enable practical applications. Current approaches leverage two kinds of information to deal with the vehicle detection and tracking problem: (1) 3D representations (eg. wireframe models or voxel based or CAD models) for diverse vehicle skeletal structures learnt from data, and (2) classifiers trained to detect vehicles or vehicle parts in single images built on top of a basic feature extraction step. In this paper, we propose to extend current approaches in two ways. First, we extend detection to a multiple view setting. We show that leveraging information given by feature or part detectors in multiple images can lead to more accurate detection results than single image detection. Secondly, we show that given multiple images of a vehicle, we can also leverage 3D information from the scene generated using a unique structure from motion algorithm. This helps us localize the vehicle in 3D, and constrain the parameters of optimization for fitting the 3D model to image data. We show results on the KITTI dataset, and demonstrate superior results compared with recent state-of-the-art methods, with upto 14.64 % improvement in localization error.
Falak Chhaya, N. Dinesh Reddy, Sarthak Upadhyay, Visesh Chari, M. Zeeshan Zia, K. Madhava Krishna
ICRA6
2016 Plantation monitoring and yield estimation using autonomous quadcopter for precision agriculture
abstract
Recently, quadcopters with their advance sensors and imaging capabilities have become an imperative part of the precision agriculture. In this work, we have described a framework which performs plantation monitoring and yield estimation using the supervised learning approach, while autonomously navigating through an inter-row path of the plantation. The proposed navigation framework assists the quadcopter to follow a sequence of collision-free GPS way points and has been integrated with ROS (Robot Operating System). The trajectory planning and control module of the navigation framework employ convex programming techniques to generate minimum time trajectory between way-points and produces appropriate control inputs for the quadcopter. A new ‘pomegranate dataset’ comprising of plantation surveillance video and annotated frames capturing the varied stages of pomegranate growth along with the navigation framework are being delivered as a part of this work.
Vishakh Duggal, Mohak Sukhwani, Kumar Bipin, G. Syamasundar Reddy, K. Madhava Krishna
ICRA5
2016 Discriminative learning based visual servoing across object instances
abstract
Classical visual servoing approaches use visual features based on geometry of the object such as points, lines, region, etc. to attain the desired camera pose. However, geometrical features are not suited for visual servoing across different object instances due to large variations in appearance and shape. In this paper, we present a new framework for visual servoing across object instances. Our approach is based on a discriminative learning framework where the desired pose is estimated using previously seen examples. Specifically, we learn a binary classifier that separates the desired pose from all other poses for that object category. The classification error is then used to control the end-effector so that the desired pose is attained. We present controllers for linear, kernel and exemplar Support Vector Machine (SVM) and empirically discuss their performance in the visual servoing context. To address large intra-category variation in appearance, we propose a modified version of Histogram of Oriented Gradients (HOG) features for visual servoing. We show effective servoing across diverse instances over 3 object categories with zero terminal velocity and acceptable camera pose error at termination.
Harit Pandya, K. Madhava Krishna, C. V. Jawahar
ICRA2
2016 Rolling shutter and motion blur removal for depth cameras
abstract
Structured light range sensors (SLRS) like the Microsoft Kinect have electronic rolling shutters (ERS). The output of such a sensor while in motion is subject to significant motion blur (MB) and rolling shutter (RS) distortion. Most robotic literature still does not explicitly model this distortion, resulting in inaccurate camera motion estimation. In RGBD cameras, we show via experimentation that the distortion undergone by depth images is different from that of color images and provide a mathematical model for it. We propose an algorithm that rectifies for these RS and MB distortions. To assess the performance of the algorithm we conduct an extensive set of experiments for each step of the pipeline. We assess the performance of our algorithm by comparing the performance of the rectified images on scene-flow and camera pose estimation, and show that with our proposed rectification, the performance improvement is significant.
Siddharth Tourani, Sudhanshu Mittal, Akhil Nagariya, Visesh Chari, K. Madhava Krishna
ICRA5
2015 Autonomous navigation of generic monocular quadcopter in natural environment
abstract
Autonomous navigation of generic monocular quadcopter in the natural environment requires sophisticated mechanism for perception, planning and control. In this work, we have described a framework which performs perception using monocular camera and generates minimum time collision free trajectory and control for any commercial quadcopter flying through cluttered unknown environment. The proposed framework first utilizes supervised learning approach to estimate the dense depth map for video stream obtained from frontal monocular camera. This depth map is initially transformed into Ego Dynamic Space and subsequently, is used for computing locally traversable way-points utilizing binary integer programming methodology. Finally, trajectory planning and control module employs a convex programming technique to generate collision-free trajectory which follows these way-points and produces appropriate control inputs for the quadcopter. These control inputs are computed from the generated trajectory in each update. Hence, they are applicable to achieve closed-loop control similar to model predictive controller. We have demonstrated the applicability of our system in controlled indoors and in unstructured natural outdoors environment.
Kumar Bipin, Vishakh Duggal, K. Madhava Krishna
ICRA3
2015 Servoing across object instances: Visual servoing for object category
abstract
Traditional visual servoing is able to navigate a robotic system between two views of the same object. However, it is not designed to servo between views of different objects. In this paper, we consider a novel problem of servoing any instance (exemplar) of an object category to a desired pose (view) and propose a strategy to accomplish the task. We use features that semantically encode the locations of object parts and define the servoing error as the difference between positions of corresponding parts in the image space. Our controller is based on the linear combination of 3D models, such that the resulting model interpolates between the given and desired instances. We conducted our experiments on five different object categories in simulation framework and show that our approach achieves the desired pose with smooth trajectory. Furthermore, we show the performance gain achieved by using a linear combination of models (instances) vis a vis a controller that switches across models during servoing in terms of trajectory's length, smoothness and error in camera pose and image features.
Harit Pandya, K. Madhava Krishna, C. V. Jawahar
ICRA2
2015 Closed form characterization of collision free velocities and confidence bounds for non-holonomic robots in uncertain dynamic environments
abstract
Navigating non-holonomic mobile robots in dynamic environments is challenging because it requires computing at each instant, the space of collision free velocities, characterized by a set of highly non-linear and non-convex inequalities. Moreover, uncertainty in obstacle trajectories further increases the complexity of the problem, as it now becomes imperative to relate the space of collision free velocities to a confidence measure. In this paper, we present a novel perspective towards analyzing and solving probabilistic collision avoidance constraints based on our previous works on non-linear time scaling. In particular, we have shown earlier that a time scaled version of collision cone constraints can be solved in closed form and thus can be used to efficiently characterize the space of collision free velocities. In the current proposed work, we present a probabilistic version of time scaled collision cone constraints obtained by representing obstacle states through generic probability distributions. We present a novel reformulation of the probabilistic constraints into a family of deterministic algebraic constraints. The solution space of each member of the family can be derived in closed form and at the same time, can also be related to the lower bound on confidence measure through Cantelli's inequality. Thus, the proposed work represents a significant improvement over the current state of the art frameworks where probabilistic collision avoidance constraints are solved through exhaustive sampling in the state-control space. We also present a cost metric which serves as the basis for the construction of the various collision avoidance maneuvers based on factors like deviation from the current path, acceleration/de-acceleration capability of the robot, confidence of collision avoidance etc. We very briefly explain how the current robot state can be connected to the solution space of safe velocities in smooth time optimal fashion. Finally, the validity of the proposed formulation is exhibited through extensive numerical simulation results.
Bharath Gopalakrishnan, Arun Kumar Singh 0001, K. Madhava Krishna
IROS3
2015 Dynamic body VSLAM with semantic constraints
abstract
Image based reconstruction of urban environments is a challenging problem that deals with optimization of large number of variables, and has several sources of errors like the presence of dynamic objects. Since most large scale approaches make the assumption of observing static scenes, dynamic objects are relegated to the noise modelling section of such systems. This is an approach of convenience since the RANSAC based framework used to compute most multiview geometric quantities for static scenes naturally confine dynamic objects to the class of outlier measurements. However, reconstructing dynamic objects along with the static environment helps us get a complete picture of an urban environment. Such understanding can then be used for important robotic tasks like path planning for autonomous navigation, obstacle tracking and avoidance, and other areas. In this paper, we propose a system for robust SLAM that works in both static and dynamic environments. To overcome the challenge of dynamic objects in the scene, we propose a new model to incorporate semantic constraints into the reconstruction algorithm. While some of these constraints are based on multi-layered dense CRFs trained over appearance as well as motion cues, other proposed constraints can be expressed as additional terms in the bundle adjustment optimization process that does iterative refinement of 3D structure and camera / object motion trajectories. We show results on the challenging KITTI urban dataset for accuracy of motion segmentation and reconstruction of the trajectory and shape of moving objects relative to ground truth. We are able to show average relative error reduction by 41 % for moving object trajectory reconstruction relative to state-of-the-art methods like TriTrack[16], as well as on standard bundle adjustment algorithms with motion segmentation.
N. Dinesh Reddy, Prateek Singhal, Visesh Chari, K. Madhava Krishna
IROS4
2015 A class of non-linear time scaling functions for smooth time optimal control along specified paths
abstract
Computing time optimal motions along specified paths forms an integral part of the solution methodology for many motion planning problems. Conventionally, this optimal control problem is solved considering piece-wise constant parametrization for the control input which leads to convexity and sparsity in the optimization structure. However, it also results in discontinuous control trajectory which is difficult to track. Thus, in this paper we revisit this time optimal control problem with the primary motivation of ensuring a high degree of smoothness in the resulting motion profile. In particular, we solve it with continuity constraints in control and higher order motion derivatives like jerk, snap etc. It is clear that such constraints would necessitate the use of time varying control inputs over the commonly used piece-wise constant form. The primary contribution of the current work lies in the introduction of a C∞class of time scaling functions represented as parametric exponentials. This in turn allows us to represent time varying control inputs as products of parametric exponential and a polynomial functions. We present the motivation behind adopting such representation of time scaling function over more common polynomial forms, both from mathematical as well as implementation standpoint. We also show that the proposed representation of time scaling function and control input leads to a very simple optimization structure where most of the constraints are linear. The non-linearity has a quasi-convex structure which can be reformulated into a simple difference of convex form. Thus, the resulting optimization can be efficiently solved through sequential convex programming where, at each iteration, the constraints in difference of convex form are further simplified to more conservative linear constraints.
Arun Kumar Singh 0001, K. Madhava Krishna
IROS2
2015 Stair Climbing using a compliant modular robot
abstract
Stair Climbing is a key functionality desired for robots deployed in Urban Search and Rescue (USAR) scenarios. A novel compliant modular robot was proposed earlier to climb steep and big obstacles. This work extends the functionality of this robot to ascend and descend stairs of dimensions that are also typical of an urban setting. Stair Climbing is realized by equipping the robot's link joints with optimally designed passive spring pairs that resist clockwise and counter clockwise moments generated by the ground during the climbing motion. This 3-module robot is only propelled by wheel actuators. Desirable stair climbing configurations are estimated a-priori and used to obtain the optimal stiffness for springs. Extensive numerical simulation results over different stair configurations are shown. The numerical simulations are corroborated by experimentation using the prototype and its performance is tabulated for different types of surfaces.
Sri Harsha Turlapati, Mihir Shah, Phani-Teja Singamaneni, Avinash Siravuru, Suril Vijaykumar Shah, K. Madhava Krishna
IROS6
2015 Overtaking maneuvers by non linear time scaling over reduced set of learned motion primitives
abstract
Overtaking of a vehicle moving on structured roads is one of the most frequent driving behavior. In this work, we have described a Real Time Control System based framework for overtaking maneuver of autonomous vehicles. Proposed framework incorporates Intelligent Planning and Modular control modules. Intelligent Planning module of the framework enables the vehicle to intelligently select the most appropriate behavioral characteristics given the perceived operating environment. Subsequently, Modular control module reduces the search space of overtaking trajectories through an SVM based learning approach. These trajectories are then examined for possible future time collision using Velocity Obstacle. It employs non linear time scaling that provides for continuous trajectories in the space of linear and angular velocities to achieve continuous curvature overtaking maneuvers respecting velocity and acceleration bounds. Further time scaling also can scale velocities to avoid collisions and can compute a time optimal trajectory for the learned behavior. The preliminary results show the appropriateness of our proposed framework in virtual urban environment.
Vishakh Duggal, Kumar Bipin, Arun Kumar Singh 0001, Bharath Gopalakrishnan, Brijendra K. Bharti, Abdelaziz Khiat, K. Madhava Krishna
Intelligent Vehicles Symposium7
2014 Small Object Discovery and Recognition Using Actively Guided Robot
abstract
In the field of active perception, object search is a widely studied problem. To search for an object in large rooms, it would be expensive to explore and check each object's similarity with the object of interest. The expense could uncontrollably bloat as the number of objects to be searched increases. If the objects are of the order of a 2-5cm, they appear very small, making it difficult for the present algorithms to recognize them. A general human strategy in such cases is to sparsely identify, from far away (4-6m), if the object of interest is present in the scene. Subsequently, each of the possible objects is analysed from closer proximity to recognize, for further manipulation. In this work, we present a similar framework. We reduce search-space, by identifying existential probability of a small object from a distance followed by a closer 3-D analysis of its point cloud to accurately recognize it. This is achieved by 2-D modelling of the objects using Gaussian Mixture Models followed by recognizing objects using efficient RGB-Depth based algorithm.
Sudhanshu Mittal, M. Siva Karthik, Suryansh Kumar 0001, K. Madhava Krishna
ICPR4
2014 A compliant multi-module robot for climbing big step-like obstacles
abstract
A novel compliant robot is proposed for traversing on unstructured terrains. The robot consists of modules, each containing a link and an active wheel-pair, and neighboring modules are connected using a passive joint. This type of robots are lighter and provide high durability due to the absence of link-actuators. However, they have limited climbing ability due to tendency of tipping over while climbing big obstacles. To overcome this disadvantage, the use of compliant joints is proposed in this work. Stiffness of each compliant joint is estimated by formulating an optimization problem with an objective to minimize link joint moments while maintaining static-equilibrium. This is one of the key novelties of the proposed work. A design methodology is also proposed for developing an n-module compliant robot for climbing a given height on a known surface. The efficacy of the proposed formulation is illustrated using numerical simulations of the three and five module robots. The robot is successfully able to climb maximum heights upto three times and six times the wheel diameter using three and five modules, respectively. A working prototype was developed and the simulation results were successfully validated on it.
Avinash Siravuru, Akshaya Purohit, Suril Vijaykumar Shah, K. Madhava Krishna
ICRA5
2014 Reactionless visual servoing of a dual-arm space robot
abstract
This paper presents a novel visual servoing controller for a satellite mounted dual-arm space robot. The controller is designed to complete the task of servoing the robot's endeffectors to the desired pose, while regulating orientation of the base-satellite. Task redundancy approach is utilized to coordinate the servoing process and attitude of the base satellite. The visual task is defined as a primary task, while regulating attitude of the base satellite to zero is defined as a secondary task. The secondary task is formulated as an optimization problem in such a way that it does not affect the primary task, and simultaneously minimizes its cost function. A set of numerical experiments are carried out on a dual-arm space robot showing efficacy of the proposed control methodology.
A. H. Abdul Hafez, V. V. Anurag, Suril Vijaykumar Shah, K. Madhava Krishna, C. V. Jawahar
ICRA4
2014 Markov Random Field based small obstacle discovery over images
abstract
Small obstacles of the order of 0.5–3cms and homogeneous scenes often pose a problem for indoor mobile robots. These obstacles cannot be clearly distinguished even with the state of the art depth sensors or laser range finders using existing vision based algorithms. With the advent of sophisticated image processing algorithms like SLIC [1] and LSD [9], it is possible to extract rich information from an image which led us to develop a novel architecture to detect very small obstacles on the floor using a monocular camera. This information is further processed using a Markov Random Field based graph cut formalism that precisely segments the floor and detects obstacles which are extremely low. We show robust and accurate obstacle detection and floor segmentation in diverse environments over a large variety of objects found indoors. In our case, low lying obstacles, changing floor patterns and extremely homogeneous environments are properly classified which leads to a drastic decrease in the number of obstacles that may not be classified by existing robotic vision algorithms.
Suryansh Kumar 0001, M. Siva Karthik, K. Madhava Krishna
ICRA3
2014 Posture control of a three-segmented tracked robot with torque minimization during step climbing
abstract
In this paper, we present a posture control scheme for step climbing by an in-house developed three-segmented tracked robot, miniUGV. The posture control scheme results in minimum torque at the actuated joints of the segments. Non-linear optimization is carried out offline for progressively decreasing distance of the robot from the step with torque minimization as objective function and force balance, motor torque limits, slippage avoidance and interference avoidance constraints. The resulting angles of the joints are fitted to a third degree polynomial as a function of the robot distance from the step and the step height. It is shown that a single set of polynomial functions is sufficient for climbing steps of all permissible heights and angles of attack of the front segment. The methodology has been verified through simulation followed by implementation on the real robot. As a consequence of this optimization we find that the average current reduced by more than thirty percent, reducing power consumption and confirming the efficacy of the optimization framework.
Sartaj Singh, Babu D. Jadhav, K. Madhava Krishna
ICRA3
2014 Time scaled collision cone based trajectory optimization approach for reactive planning in dynamic environments
abstract
The current paper proposes a trajectory optimization approach for navigating a non-holonomic wheeled mobile robot in dynamic environments. The dynamic obstacle's motion is not known and hence is represented by a band of predicted trajectories. The trajectory optimization can account for large number of predicted obstacle trajectories and seeks to avoid each predicted trajectory of every obstacle in the sensing range of the robot. The two primary contributions of the proposed trajectory optimization are (1): A computationally efficient method for computing the intersection space of collision avoidance constraints of large number of predicted obstacle trajectories. (2): A optimization framework to connect the current state to the solution space in time optimal fashion. The intersection/solution space computation is build on our earlier proposed concept of time scaled collision cone, which can be solved in closed form to obtain a set of formulae. These formulae describe how much and in what manner the temporal specification of a trajectory needs to be changed to avoid a given set of dynamic obstacles. This allows us to quickly evaluate solution space of time scaled collision cone over various candidate trajectories, thus reducing the problem of computing the intersection space to that of generating multiple homotopic trajectories. The optimization framework used to connect the current state to the solution space in time optimal fashion is based on the concept of non-linear time scaling, which induces a difference of convex form structure. Thus, on the theoretical side, we show that the various components of the proposed framework are computationally simple and involves solving sets of linear equations and using state of the art convex programming techniques. On the practical side we show that the proposed planner performs better than sampling based planners which treat dynamic obstacles as static over a short duration of time.
Bharath Gopalakrishnan, Arun Kumar Singh 0001, K. Madhava Krishna
IROS3
2013 Depth really Matters: Improving Visual Salient Region Detection with Depth
abstract
Depth information has been shown to affect identification of visually salient regions in images. In this paper, we investigate the role of depth in saliency detection in the presence of (i) competing saliencies due to appearance, (ii) depth-induced blur and (iii) centre-bias. Having established through experiments that depth continues to be a significant contributor to saliency in the presence of these cues, we propose a 3D-saliency formulation that takes into account structural features of objects in an indoor setting to identify regions at salient depth levels. Computed 3D-saliency is used in conjunction with 2D-saliency models through non-linear regression using SVM to improve saliency maps. Experiments on benchmark datasets containing depth information show that the proposed fusion of 3D-saliency with 2D-saliency models results in an average improvement in ROC scores of about 9% over state-of-the-art 2D saliency models. The main contributions of this paper are: (i) The development of a 3D-saliency model that integrates depth and geometric features of object surfaces in indoor scenes. (ii) Fusion of appearance (RGB) saliency with depth saliency through non-linear regression using SVM. (iii) Experiments to support the hypothesis that depth improves saliency detection in the presence of blur and centre-bias. The effectiveness of the 3D-saliency model and its fusion with RGB-saliency is illustrated through experiments on two benchmark datasets that contain depth information. Current stateof-the-art saliency detection algorithms perform poorly on these datasets that depict indoor scenes due to the presence of competing saliencies in the form of color contrast. For example in Fig. 1, saliency maps of [1] is shown for different scenes, along with its human eye fixations and our proposed saliency map after fusion. It is seen from the first scene of Fig. 1, that illumination plays spoiler role in RGB-saliency map. In second scene of Fig. 1, the RGB-saliency is focused on the cap though multiple salient objects are present in the scene. Last scene at the bottom of Fig. 1, shows the limitation of the RGB-saliency when the object is similar in appearance with the background. Effect of depth on Saliency: In [4], it is shown that depth is an important cue for saliency. In this paper we go further and verify if the depth alone influences the saliency. Different scenes were captured for experimentation using Kinect sensor. Observations resulted out of these experiments are (i) Humans fixate on the objects at closer depth, in the presence of visually competing salient objects in the background, (ii) Early attention happens on the objects at closer depth, (iii) Effective fixations are high at the low contrast foreground compared to the high contrast objects in the background which are blurred, (iv) Low contrast object placed at the center of the field of view, gets more attention compared to other locations. As a result of all these observations, we develop a 3D-saliency that captures the depth information of the regions in the scene. 3D-Saliency: We adapt the region based contrast method from Cheng et al. [1] in computing contrast strengths for the segmented 3D surfaces or regions. Each segmented region is assigned a contrast score using surface normals as the feature. Structure of the surface can be described based on the distribution of normals in the region. We compute a histogram of angular distances formed by every pair of normals in the region. Every region Rk is associated with a histogram Hk. Contrast score Ck of a region Rk is computed as the sum of the dot products of its histogram with histograms of other regions in the scene. Since the depth of the region is influencing the visual attention, the contrast score is scaled by a value Zk, which is the depth of the region Rk from the sensor. In order to define the saliency, sizes of the regions i.e. the number of the points in the region, have to be considered. We find the ratio of the region dimension to the half of the scene dimension. Considering nk as the number of 3D points in the region Rk, the constrast score becomes Figure 1: Four different scenes and their saliency maps; For each scene from top left (i) Original Image, (ii) RGB-Saliency map using RC [1], (iii) Human fixations from eye-tracker and (iv) Fused RGBD-saliency map
Karthik Desingh, K. Madhava Krishna, Deepu Rajan, C. V. Jawahar
BMVC2
2013 Multibody VSLAM with relative scale solution for curvilinear motion reconstruction
abstract
A solution to the relative scale problem where reconstructed moving objects and the stationary world are represented in a unified common scale has proven equivalent to a conjecture. Motion reconstruction from a moving monocular camera is considered ill posed due to known problems of observability. We show for the first time several significant motion reconstruction of outdoor vehicles moving along non-holonomic curves and straight lines. The reconstructed motion is represented in the unified frame which also depicts the estimated camera trajectory and the reconstructed stationary world. This is possible due to our Multibody VSLAM framework with a novel solution for relative scale proposed in the current paper. Two solutions that compute the relative scale are proposed. The solutions provide for a unified representation within four views of reconstruction of the moving object and are thus immediate. In one, the solution for the scale is that which satisfies the planarity constraint of the object motion. The assumption of planar object motion while being generic enough is subject to stringent degenerate situations that are more widespread. To circumvent such degeneracies we assume that the object motion to be locally circular or linear and find the relative scale solution for such object motions. Precise reconstruction is achieved in synthetic data. The fidelity of reconstruction is further vindicated with reconstructions of moving cars and vehicles in uncontrolled outdoor scenes.
Rahul Kumar Namdev, K. Madhava Krishna, C. V. Jawahar
ICRA2
2013 Heterogeneous UGV-MAV exploration using integer programming
abstract
This paper presents a novel exploration strategy for coordinated exploration between unmanned ground vehicles (UGV) and micro-air vehicles (MAV). The exploration is modeled as an Integer Programming (IP) optimization problem and the allocation of the vehicles(agents) to frontier locations is modeled using binary variables. The formulation is also studied for distributed system, where agents are divided into multiple teams using graph partitioning. Optimization seamlessly integrates several practical constraints that arise in exploration between such heterogeneous agents and provides an elegant solution for assigning task to agents. We have also presented comparison with previous methods based on distance traversed and computational time to signify advantages of presented method. We also show practical realization of such an exploration where an UGV-MAV team efficiently builds a map of an indoor environment.
Ayush Dewan, Aravindh Mahendran, Nikhil Soni, K. Madhava Krishna
IROS4
2013 Visual localization in highly crowded urban environments
abstract
Visual localization in crowded dynamic environments requires information about static and dynamic objects. This paper presents a robust method that learns the useful features from multiple runs in highly crowded urban environments. Useful features are identified as distinctive ones that are also reliable to extract in diverse imaging conditions. Relative importance of features is used to derive the weight for each feature. The popular Bag-of-words model is used for image retrieval and localization, where query image is the current view of the environment and database contains the visual experience from previous runs. Based on the reliability, features are augmented and eliminated over runs. This reduces the size of representation, and makes it more reliable in crowded scenes. We tested the proposed method on data sets collected from highly crowded Indian urban outdoor settings. Experiments have shown that with the help of a small subset (10%) of the detected features, we can reliably localize the camera. We achieve superior results in terms of localization accuracy even when more than 90% of the pixels are occluded or dynamic.
A. H. Abdul Hafez, K. Madhava Krishna, C. V. Jawahar
IROS3
2013 Coordinating mobile manipulator's motion to produce stable trajectories on uneven terrain based on feasible acceleration count
abstract
In this paper we consider the problem of coordinating the motion of the manipulator and the vehicle to produce stable trajectories for the combined mobile manipulator system on uneven terrain. These kinds of situations often arise in planetary exploration, where rovers equipped with a manipulator are required to navigate over general uneven terrain. Moreover the framework can also be used in situations where the mobile manipulator is required to transport objects on uneven terrain. We generate feasible trajectories for the vehicle between a given start and a goal point considering the dynamics of the manipulator. The framework proposed in the paper plans such motion profile of the manipulator that maximizes vehicle stability which is measured by a novel concept called Feasible Acceleration Count (FAC). We show that, from the point of view of motion planning of mobile manipulator on uneven terrains, FAC gives a better estimate of vehicle stability than more popular metrics like Tip-Over Stability. The trajectory planner closely resembles motion primitive based graph based planning and is combined with a novel cost function derived from FAC. The efficacy of the approach is shown through simulations of a mobile manipulator system on a 2.5D uneven terrain.
Arun Kumar Singh 0001, K. Madhava Krishna
IROS2
2012 Motion segmentation of multiple objects from a freely moving monocular camera
abstract
Motion segmentation is an inevitable component for mobile robotic systems such as the case with robots performing SLAM and collision avoidance in dynamic worlds. This paper proposes an incremental motion segmentation system that efficiently segments multiple moving objects and simultaneously build the map of the environment using visual SLAM modules. Multiple cues based on optical flow and two view geometry are integrated to achieve this segmentation. A dense optical flow algorithm is used for dense tracking of features. Motion potentials based on geometry are computed for each of these dense tracks. These geometric potentials along with the optical flow potentials are used to form a graph like structure. A graph based segmentation algorithm then clusters together nodes of similar potentials to form the eventual motion segments. Experimental results of high quality segmentation on different publicly available datasets demonstrate the effectiveness of our method.
Rahul Kumar Namdev, Abhijit Kundu, K. Madhava Krishna, C. V. Jawahar
ICRA3
2012 Planning trajectories on uneven terrain using optimization and non-linear time scaling techniques
abstract
In this paper we introduce a novel framework of generating trajectories which explicitly satisfies the stability constraints such as no-slip and permanent ground contact on uneven terrain. The main contributions of this paper are: (1) It derives analytical functions depicting the evolution of the vehicle on uneven terrain. These functional descriptions enable us to have a fast evaluation of possible vehicle stability along various directions on the terrain and this information is used to control the shape of the trajectory. (2) It introduces a novel paradigm wherein non-linear time scaling brought about by parametrized exponential functions are used to modify the velocity and acceleration profile of the vehicle so that these satisfy the no-slip and contact constraints. We show that nonlinear time scaling manipulates velocity and acceleration profile in a versatile manner and consequently has exceptional utility not only in uneven terrain navigation but also in general in any problem where it is required to change the velocity of the robot while keeping the path unchanged like collision avoidance.
Arun Kumar Singh 0001, K. Madhava Krishna, Srikanth Saripalli
IROS2
2011 Realtime multibody visual SLAM with a smoothly moving monocular camera
abstract
This paper presents a realtime, incremental multibody visual SLAM system that allows choosing between full 3D reconstruction or simply tracking of the moving objects. Motion reconstruction of dynamic points or objects from a monocular camera is considered very hard due to well known problems of observability. We attempt to solve the problem with a Bearing only Tracking (BOT) and by integrating multiple cues to avoid observability issues. The BOT is accomplished through a particle filter, and by integrating multiple cues from the reconstruction pipeline. With the help of these cues, many real world scenarios which are considered unobservable with a monocular camera is solved to reasonable accuracy. This enables building of a unified dynamic 3D map of scenes involving multiple moving objects. Tracking and reconstruction is preceded by motion segmentation and detection which makes use of efficient geometric constraints to avoid difficult degenerate motions, where objects move in the epipolar plane. Results reported on multiple challenging real world image sequences verify the efficacy of the proposed framework.
Abhijit Kundu, K. Madhava Krishna, C. V. Jawahar
ICCV2
2011 Large scale visual localization in urban environments
abstract
This paper introduces a vision based localization method for large scale urban environments. The method is based upon Bag-of-Words image retrieval techniques and handles problems that arise in urban environments due to repetitive scene structure and the presence of dynamic objects like vehicles. The localization system was experimentally verified it localization experiments along a 5km long path in an urban environment.
Supreeth Achar, C. V. Jawahar, K. Madhava Krishna
ICRA3
2011 Quasi-static motion planning on uneven terrain for a wheeled mobile robot
abstract
In this paper we present a motion planning algorithm connecting a starting and ending goal positions of a wheeled mobile robot (WMR) with a passive variable camber (PVC) on a fully 3D uneven terrain without slipping. The overall planning framework is along the lines of the RRT (Rapidly Exploring Random Tree). The curve connecting the adjacent nodes of the RRT is a quasi-static path which is generated using the forward motion problem based on the Peshkin's minimum energy principle which combines the force and kinematic relationships of the WMR into a nonlinear optimization problem. The output of this optimization routine is a set of ordinary differential equations (ODEs) representing the non-holonomic constraints and wheel ground contact conditions of the robot along with a set of differential algebraic equations (DAEs) representing the geometric/holonomic constraints of the robot. In general a complete simulation of a WMR on a fully 3D terrain has been a difficult problem to solve. Previous methods for continuous evolution of the WMR have only incorporated the wheel ground contact constraints within the DAE framework. This work goes beyond the previous methods by incorporating the quasi-static and friction cone constraints within the DAE framework. This evolution is now extended to a motion planning algorithm which guarantees that the vehicle traverses along quasi-static stable paths.
Vijay Eathakota, Gattupalli Aditya, K. Madhava Krishna
IROS3
2010 Mapping large scale environments by combining Particle Filter and Information Filter
abstract
This paper presents two approaches to combine two popular mapping strategies, namely Particle Filters and Information Filters. The first method describes how the Particle Filter can be incorporated into the Information Filter framework by building local submaps using the Particle Filter and combining them using a Information Filter to obtain a global map. Using the Particle Filter locally reduces the linearization errors and is useful in handling ambiguous data associations, while the Information Filter keeps track of the uncertainty over long periods of time, thereby avoiding FastSLAM's tendency to become overconfident. The second method shows how the Information Filter can be used in the Particle Filter framework as a simple means of remembering the filter's uncertainty. This can then be used to repopulate particles while closing loops. This not only handles non linearities but is also robust for loop closing because, unlike the Particle Filter, the Information Filter does not exhibit forgetfulness of the trajectory's past.
Mahesh Mohan, K. Madhava Krishna
ICARCV2
2010 Fast and Spatially-Smooth Terrain Classification Using Monocular Camera
abstract
In this paper, we present a monocular camera based terrain classification scheme. The uniqueness of the proposed scheme is that it inherently incorporates spatial smoothness while segmenting a image, without requirement of post-processing smoothing methods. The algorithm is extremely fast because it is build on top of a Random Forest classifier. We present comparison across features and classifiers. The baseline algorithm uses color, texture and their combination with classifiers such as SVM and Random Forests. We further enhance the algorithm through a label transfer method. The efficacy of the proposed solution can be seen as we reach a low error rates on both our dataset and other publicly available datasets.
Chetan Jakkoju, K. Madhava Krishna, C. V. Jawahar
ICPR2
2010 An adaptive outdoor terrain classification methodology using monocular camera
abstract
An adaptive partition based Random Forests classifier for outdoor terrain classification is presented in this paper. The classifier is a combination of two underlying classifiers. One of which is a random forest learnt over bootstrapped or offline dataset, the second is another random forest that adapts to changes on the fly. Posterior probabilities of both the static and changing/online classifiers are fused to assign the eventual label for the online image data. The online classifier learns at frequent intervals of time through a sparse and stable set of tracked patches, which makes it lightweight and real-time friendly. The learning which is actuated at frequent intervals during the sojourn significantly improves the performance of the classifier vis-a-vis a scheme that only uses the classifier learnt offline or at bootstrap. The method is well suited and finds immediate applications for outdoor autonomous driving where the classifier needs to be updated frequently based on what shows up recently on the terrain and without largely deviating from those learnt at bootstrapping. The role of the partition based classifier to enhance the performance of a regular multi class classifier such as random forests and multi class SVMs is also summarized in this paper.
Chetan Jakkoju, K. Madhava Krishna, C. V. Jawahar
IROS2
2010 A visual exploration algorithm using semantic cues that constructs image based hybrid maps
abstract
A vision based exploration algorithm that invokes semantic cues for constructing a hybrid map of images - a combination of semantic and topological maps is presented in this paper. At the top level the map is a graph of semantic constructs. Each node in the graph is a semantic construct or label such as a room or a corridor, the edge represented by a transition region such as a doorway that links the two semantic constructs. Each semantic node embeds within it a topological graph that constitutes the map at the middle level. The topological graph is a set of nodes, each node representing an image of the higher semantic construct. At the low level the topological graph embeds metric values and relations, where each node embeds the pose of the robot from which the image was taken and any two nodes in the graph are related by a transformation consisting of a rotation and translation. The exploration algorithm explores a semantic construct completely before moving or branching onto a new construct. Within each semantic construct it uses a local feature based exploration algorithm that uses a combination of local and global decisions to decide the next best place to move. During the process of exploring a semantic construct it identifies transition regions that serve as gateways to move from that construct to another. The exploration is deemed complete when all transition regions are marked visited. Loop detection happens at transition regions and graph relaxation techniques are used to close loops when detected to obtain a consistent metric embedding of the robot poses. Semantic constructs are labeled using a visual bag of words(VBOW) representation with a probabilistic SVM classifier.
Aravindhan K. Krishnan, K. Madhava Krishna
IROS2
2010 A two phase recursive tree propagation based multi-robotic exploration framework with fixed base station constraint
abstract
A multi-robotic exploration with the requirement of communication link to a fixed base station is presented in this paper. The robots organize themselves into roles of maintainers of communication (hinged robots or robot nodes) or explorers of the environment ensuring that every robot is in contact with the base station directly or through the hinged robots. A two phased strategy for the same is presented. The first phase is characterized by a recursive growth of trees that starts from the root node or the base station and then repeated from other nodes of the hitherto grown tree in a depth first fashion. The second phase constitutes the recursive tree growth invoked repeatedly from the frontier nodes. While the first phase rapidly explores areas around the base station in a concentric fashion, the second phase extends the depth of the explored area to increase the limits of coverage. The strategy is consistent in that none of the robots loose contact with the base station. Extensive simulations confirm the efficacy of the method and comparisons portray performance gain in terms of exploration time and absence of deadlocks vis-a-vis the few methods previously reported in the literature.
Piyoosh Mukhija, K. Madhava Krishna, Vamshi Krishna
IROS2
2010 A novel compliant rover for rough terrain mobility
abstract
In this paper a novel suspension mechanism for rough terrain mobility is proposed. The proposed mechanism is simpler than the existing suspension mechanism in the sense that the number of links and joints has been significantly reduced without compromising the climbing ability of the rover. We explore the use of compliant elements like springs for passively controlling the degree of freedom of the proposed mechanism and a framework for optimizing the spring parameters has been proposed. A performance evaluation of the proposed mechanism has been shown in terms of extensive simulations.
Arun Kumar Singh 0001, Rahul Kumar Namdev, Vijay Eathakota, K. Madhava Krishna
IROS4
2009 Secured Multi-robotic Active Localization without Exchange of Maps: A Case of Secure Cooperation Amongst Non-trusting Robots
abstract
Secure multiparty protocols have found applications in numerous domains, where multiple nontrusting parties wish to evaluate a function of their private inputs. In this paper, we consider the case of multiple robots wishing to localize themselves, with maps as their private inputs. Though localization of robots has been a well studied problem, only recent studies have shown how to actively localize multiple robots through coordination. In all such studies, localization has typically been achieved through constructing a publicly known global map. Here, we show how a similar solution can be given in the case of nontrusting robots, which do not wish to disclose their local maps.
Sarat C. Addepalli, Piyush Bansal, K. Srinathan 0001, K. Madhava Krishna
ARES4
2009 Moving object detection by multi-view geometric techniques from a single camera mounted robot
abstract
The ability to detect, and track multiple moving objects like person and other robots, is an important prerequisite for mobile robots working in dynamic indoor environments. We approach this problem by detecting independently moving objects in image sequence from a monocular camera mounted on a robot. We use multi-view geometric constraints to classify a pixel as moving or static. The first constraint, we use, is the epipolar constraint which requires images of static points to lie on the corresponding epipolar lines in subsequent images. In the second constraint, we use the knowledge of the robot motion to estimate a bound in the position of image pixel along the epipolar line. This is capable of detecting moving objects followed by a moving camera in the same direction, a so-called degenerate configuration where the epipolar constraint fails. To classify the moving pixels robustly, a Bayesian framework is used to assign a probability that the pixel is stationary or dynamic based on the above geometric properties and the probabilities are updated when the pixels are tracked in subsequent images. The same framework also accounts for the error in estimation of camera motion. Successful and repeatable detection and pursuit of people and other moving objects in realtime with a monocular camera mounted on the Pioneer 3DX, in a cluttered environment confirms the efficacy of the method.
Abhijit Kundu, K. Madhava Krishna, Jayanthi Sivaswamy
IROS2
2008 Active global localization for multiple robots by disambiguating multiple hypotheses
abstract
In environments which possess relatively few features that enable a robot to unambiguously determine its location, global localization algorithms can result in multiple hypotheses locations of a robot. In such a scenario the robot, for effective localization, has to be actively guided to those locations where there is a maximum chance of eliminating most of the ambiguous states - which is often referred to as dasiaactive localizationpsila. When extended to multi robotic scenarios where all robots possess more than one hypothesis of their position, there is the opportunity to do better by using robots apart from obstacles as dasiahypotheses resolving agentspsila. The paper presents a unified framework accounting for the map structure as well as measurement amongst robots while guiding a set of robots to locations where they can singularize to a unique state. The appropriateness of our approach is demonstrated empirically in both simulation & real-time (on Amigobots) and its efficacy verified. Extensive comparative analysis portrays the advantage of the current method over others that do not perform active localization in a multi-robotic sense.
Shivudu Bhuvanagiri, K. Madhava Krishna
IROS2
2008 On-line convex optimization based solution for mapping in VSLAM
abstract
This paper presents a novel real-time algorithm to sequentially solve the triangulation problem. The problem addressed is estimation of 3D point coordinates given its images and the matrices of the respective cameras used in the imaging process. The algorithm has direct application to real time systems like visual SLAM. This article demonstrates the application of the proposed algorithm to the mapping problem in visual SLAM. Experiments have been carried out for the general triangulation problem as well as the application to visual SLAM. Results show that the application of the proposed method to mapping in visual SLAM outperforms the state of the art mapping methods.
A. H. Abdul Hafez, Shivudu Bhuvanagiri, K. Madhava Krishna, C. V. Jawahar
IROS3
2008 Covering hostile terrains with partial and complete visibilities: On minimum distance paths
abstract
We present a method for finding paths for multiple Unmanned Air Vehicles (UAVs) such that the sum over their lengths is minimum as they cover a 3D terrain (represented as height fields). The paths are constrained to lie beneath an exposure surface to ensure stealth from enemy outposts. The exposure surface is also computed as a height field. The algorithm greedily clusters the terrain such that gain in visibility per distance would be higher for intra-cluster points than points across clusters. Paths generated on clusters formed by such a per distance visibility metric are reduced by more than 25% over other related decoupled methods. The method is extended to cover terrains with partial visibilities. The advantage of the coupled metric extends under constrained visibility also. We again show performance gain by comparing with an existing decoupled algorithm that solves a similar problem of minimum distance terrain coverage with constrained visibility. The paper reveals that decomposing the terrain based on visibility first and then distance is always better than the other way round to cover the terrain in shorter distances.
Mahesh Mohan, Rahul Sawhney, K. Madhava Krishna, K. Srinathan 0001, Manohar B. Srikanth
IROS3
2007 Optimal Multi-Sensor Based Multi Target Detection by Moving Sensors to the Maximal Clique in a Covering Graph
Ganesh P. Kumar 0003, K. Madhava Krishna
IJCAI2
2007 Feature Based Occupancy Grid Maps for Sonar Based Safe-Mapping
Amit Kumar Pandey, K. Madhava Krishna, Mainak Nath
IJCAI2
2006 Extension of Reeds and Shepp Paths to a Robot with Front and Rear Wheel Steer
abstract
This paper presents an algorithm for extending RS paths for a robot with both front and rear wheel steer. We call such robots as FR steer. The occurrence of such paths is due to the additional maneuver possible in such a robot which we call parallel steer, in addition to the ones already present in a vehicle with only front wheel steering. Hence we extend the optimal path set / , containing only a single element to a set : , containing n elements, thereby extending its configuration set along the optimal path from the initial to the final configuration. This extension of the set / to set : is made possible by introducing a special set, which we call the Parallel Steer (PS) Set. Such an extension of the configuration set would increase the size of the final configuration set achievable by a path that is optimal in free space. In the following discussion, we shall term all paths whose length is equal to an RS path as optimal.
Siddharth Sanan, Darshan Santani, K. Madhava Krishna, Henry Hexmoor
ICRA3
2005 A t-step ahead constrained optimal target detection algorithm for a multi sensor surveillance system
abstract
We present a methodology for optimal target detection in a multi sensor surveillance system. The system consists of mobile sensors that guard a rectangular surveillance zone crisscrossed by moving targets. Targets penetrate the surveillance zone with poisson rates at uniform velocities. Under these conditions we present a motion strategy computation for each sensor such that it maximizes target detection for the next T time-steps. A coordination mechanism among sensors ensures that overlapping and overlooked regions of observation among sensors are minimized. This coordination mechanism is interleaved with the motion strategy computation to reduce detections of the same target by more than one sensor for the same time-step. To avoid an exhaustive search in the joint space of all the sensors the coordination mechanism constrains the search by assigning priorities to the sensors and thereby arbitrating among sensory tasks. A comparison of this methodology with other multi target tracking schemes verifies its efficacy in maximizing detections. "Sample" and "time-step" are used equivalently and interchangeably in this paper.
K. Madhava Krishna, Henry Hexmoor, Shravan Kumar Sogani
IROS1
2004 Reactive Collision Avoidance of Multiple Moving Agents by Cooperation and Conflict Propagation
abstract
A strategy for collision avoidance between several moving robots that are not in possession of each other's plans is presented here. A robot's awareness of other robots is limited to the knowledge of their current states represented by their present and impending velocities and their motion direction. A robot is aware of the presence of other robots when they fall within its field of vision. Collision avoidance is attempted at three levels namely at individual, cooperative and propagation levels through velocity control. At individual level it suffices that one of the robots involved in a forthcoming collision modifies its velocity. The cooperative level is characterized by the requirement that all the robots involved in collision modify their velocities in a synchronized fashion. In the third level robots not involved in a collision are entailed to participate by altering their velocities in a manner that resolves collision conflicts between the robots involved. The third level is termed as the propagation level since the collision conflict is propagated to robots not a part of the conflict and their assistance sought in avoiding conflicts. The strategy is implemented in a distributed fashion across all robots in the system. Simulation results are presented to authenticate the efficacy of the proposed method.
K. Madhava Krishna, Henry Hexmoor
ICRA1
2002 On the influence of sensor capacities and environment dynamics onto collision-free motion plans
abstract
A methodology for computing the maximum velocity profile for a planned trajectory of the robot is described in this paper. The profile is computed considering the robot and environment dynamics as well as the constraints of the sensing apparatus. The mobile objects can be arbitrary in number and their direction and velocity of motion is not known. The only known information about the moving objects is the maximum velocity they can possess. The robot that moves with the computed velocity profile can assure from its side that it would not collide onto any of the numerous moving objects that could intercept its future trajectory. The methodology has been incorporated onto a motion planner for a nonholonomous robot and the results presented. The motivation here is to facilitate the process of having safe and understanding robots. Hence the planned velocity profiles are in general conservative though the robot could perhaps do better on-line. However at planning time the robot's immobility before collision is guaranteed.
Rachid Alami 0001, Thierry Siméon, K. Madhava Krishna
IROS3