EDBT 2026 Demo / reviewers in the wild / expert
Harit Pandya
dblp:155/3207
· DBLP profile ↗
14ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-3842-1949ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 5 since 2021Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ReCoRe: Regularized Contrastive Representation Learning of World ModelabstractWhile recent model-free Reinforcement Learning (RL) methods have demonstrated human-level effectiveness in gaming environments, their success in everyday tasks like visual navigation has been limited, particularly under significant appearance variations. This limitation arises from (i) poor sample efficiency and (ii) over-fitting to training scenarios. To address these challenges, we present a world model that learns invariant features using (i) contrastive unsupervised learning and (ii) an intervention-invariant regularizer. Learning an explicit representation of the world dynamics i.e. a world model, improves sample efficiency while contrastive learning implicitly enforces learning of invariant features, which improves generalization. However, the naïve integration of contrastive loss to world models is not good enough, as world-model-based RL methods independently optimize representation learning and agent policy. To overcome this issue, we propose an intervention-invariant regularizer in the form of an auxiliary task such as depth prediction, image denoising, image segmentation, etc., that explicitly enforces invariance to style interventions. Our method outperforms current state-of-the-art model-based and model-free RL methods and significantly improves on out-of-distribution point navigation tasks evaluated on the iGibson benchmark. With only visual observations, we further demonstrate that our approach outperforms recent language-guided foundation models for point navigation, which is essential for deployment on robots with limited computation capabilities. Finally, we demonstrate that our proposed model excels at the sim-to-real transfer of its perception module on the Gibson benchmark. Rudra P. K. Poudel, Harit Pandya, Stephan Liwicki, Roberto Cipolla |
CVPR | 2 |
| 2022 | Push-to-See: Learning Non-Prehensile Manipulation to Enhance Instance Segmentation via Deep Q-LearningabstractEfficient robotic manipulation of objects for sorting and searching often rely upon how well the objects are perceived and the available grasp poses. The challenge arises when the objects are irregular, have similar visual features (e.g., textureless objects) and the scene is densely cluttered. In such cases, non-prehensile manipulation (e.g., pushing) can facilitate grasping or searching by improving object perception and singulating the objects from the clutter via physical interaction. The current robotics literature in interactive segmentation focuses solely on isolated cases, where the central aim is on searching or singulating a single target object, or segmenting sparsely cluttered scenes, mainly through matching visual futures in successive scenes before and after the robotic interaction. On the other hand, in this paper, we introduce the first interactive segmentation model in the literature that can autonomously enhance the instance segmentation of such challenging scenes as a whole via optimising a Q-value function that predicts appropriate pushing actions for singulation. We achieved this by training a deep reinforcement learning model with reward signals generated by a Mask-RCNN trained solely on depth images. We evaluated our model in experiments by comparing its success on segmentation quality with a heuristic baseline, as well as the state-of-the-art Visual Pushing and Grasping (VPG) model [1]. Our model significantly outperformed both baselines in all benchmark scenarios. Furthermore, decreasing the segmentation error inherently enabled the autonomous singulation of the scene as a whole. Our evaluation experiments also serve as a benchmark for interactive segmentation research. Baris Serhan, Harit Pandya, Ayse Küçükyilmaz, Gerhard Neumann |
ICRA | 2 |
| 2022 | CoMBiNED: Multi-Constrained Model Based Planning for Navigation in Dynamic EnvironmentsabstractRecent deep reinforcement learning (DRL) approaches have achieved high success rate in map-less dynamic obstacle avoidance tasks. However, navigation in unseen dynamic scenarios without a pre-built map in the presence of dynamic obstacles still remains an open challenge. Since, learning accurate models for complex robotic scenarios such as navigation directly from high dimensional sensory measurements requires a large amount of data and training. Furthermore, even a small change on robot configuration such as kino-dynamics or sensor in the inference time requires re-training of the policy. In this paper, we address these issues in a principled fashion through a multi-constraint model based online planning (CoMBiNED) framework that does not require any retraining or modifications on the existing policy. We disentangle the given task into sub-tasks and learn dynamical models for them. Treating these dynamical models as soft-constraints, we employ stochastic optimisation to jointly optimize these sub-tasks on-the-fly at the inference time. We consider navigation as central application in this work and evaluate our approach on publicly available benchmark with complex dynamic scenarios and achieved significant improvement over recent approaches both in the cases of with-and-without given map of the environment. Harit Pandya, Rudra P. K. Poudel, Stephan Liwicki |
IROS | 1 |
| 2021 | RTVS: A Lightweight Differentiable MPC Framework for Real-Time Visual ServoingabstractRecent data-driven approaches to visual servoing have shown improved performances over classical methods due to precise feature matching and depth estimation. Some recent servoing approaches use a model predictive control (MPC) framework which generalise well to novel environments and are capable of incorporating dynamic constraints, but are computationally intractable in real-time, making it difficult to deploy in real-world scenarios. On the contrary, single-step methods optimise greedily and achieve high servoing rates, but lack the benefits of the MPC multi-step ahead formulation. In this paper, we make the best of both worlds and propose a lightweight visual servoing MPC framework which generates optimal control near real-time at a frequency of 10.52 Hz. This work utilises the differential cross-entropy sampling method for quick and effective control generation along with a lightweight neural network, significantly improving the servoing frequency. We also propose a flow depth normalisation layer which ameliorates the issue of inferior predictions of two view depth from the flow network. We conduct extensive experimentation on the Habitat simulator and show a notable decrease in servoing time in comparison with other approaches that optimise over a time horizon. We achieve the right balance between time and performance for visual servoing in six degrees of freedom (6DoF), while retaining the advantageous MPC formulation. Our code and dataset are publicly available†. M. Nomaan Qureshi, Pushkal Katara, Harit Pandya, Y. V. S. Harish, AadilMehdi J. Sanchawala, Gourav Kumar, Brojeshwar Bhowmick, K. Madhava Krishna |
IROS | 4 |
| 2021 | Intent-Aware Predictive Haptic Guidance and its Application to Shared Control TeleoperationabstractThis paper presents a haptic shared control paradigm that modulates the level of robotic guidance, based on predictions of human motion intentions. The proposed method incorporates robot trajectories learned from human demonstrations and dynamically adjusts the level of robotic assistance based on how closely the detected intentions match these trajectories. An experimental study is conducted to demonstrate the paradigm on a teleoperated pick-and-place task using a Franka Emika Panda robot arm, controlled via a 3D Systems Touch X haptic interface. In the experiment, the human operator teleoperates a remote robot arm while observing the environment on a 2D screen. While the human teleoperates the robot arm, the objects are tracked, and the human’s motion intentions (e.g., which object will be picked or which bin will be approached) are predicted using a Deep Q-Network (DQN). The predictions are made considering the current robot state and baseline robot trajectories that are learned from human demonstrations using Probabilistic Movement Primitives (ProMPs). The detected intentions are then used to condition the ProMP trajectories to modulate the movement and accommodate changing object configurations. Consequently, the system generates adaptive force guidance as weighted virtual fixtures that are rendered on the haptic device. The outcomes of the user study, conducted with 12 participants, indicate that the proposed paradigm can successfully guide users to robust grasping configurations and brings better performance by reducing the number of grasp attempts and improving trajectory smoothness and length. Kim Tien Ly, P. Mithun, Harit Pandya, Gerhard Neumann, Ayse Küçükyilmaz |
RO-MAN | 3 |
| 2020 | DFVS: Deep Flow Guided Scene Agnostic Image Based Visual ServoingabstractExisting deep learning based visual servoing approaches regress the relative camera pose between a pair of images. Therefore, they require a huge amount of training data and sometimes fine-tuning for adaptation to a novel scene. Furthermore, current approaches do not consider underlying geometry of the scene and rely on direct estimation of camera pose. Thus, inaccuracies in prediction of the camera pose, especially for distant goals, lead to a degradation in the servoing performance. In this paper, we propose a two-fold solution: (i) We consider optical flow as our visual features, which are predicted using a deep neural network. (ii) These flow features are then systematically integrated with depth estimates provided by another neural network using interaction matrix. We further present an extensive benchmark in a photo-realistic 3D simulation across diverse scenes to study the convergence and generalisation of visual servoing approaches. We show convergence for over 3m and 40 degrees while maintaining precise positioning of under 2cm and 1 degree on our challenging benchmark where the existing approaches that are unable to converge for majority of scenarios for over 1.5m and 20 degrees. Furthermore, we also evaluate our approach for a real scenario on an aerial robot. Our approach generalizes to novel scenarios producing precise and robust servoing performance for 6 degrees of freedom positioning tasks with even large camera transformations without any retraining or fine-tuning. Y. V. S. Harish, Harit Pandya, Ayush Gaud, Shreya Terupally, Narasimhan Sai Shankar, K. Madhava Krishna |
ICRA | 2 |
| 2019 | Recurrent Kalman Networks: Factorized Inference in High-Dimensional Deep Feature SpacesabstractIn order to integrate uncertainty estimates into deep time-series modelling, Kalman Filters (KFs) (Kalman et al., 1960) have been integrated with deep learning models, however, such approaches typically rely on approximate inference tech- niques such as variational inference which makes learning more complex and often less scalable due to approximation errors. We propose a new deep approach to Kalman filtering which can be learned directly in an end-to-end manner using backpropagation without additional approximations. Our approach uses a high-dimensional factorized latent state representation for which the Kalman updates simplify to scalar operations and thus avoids hard to backpropagate, computationally heavy and potentially unstable matrix inversions. Moreover, we use locally linear dynamic models to efficiently propagate the latent state to the next time step. The resulting network architecture, which we call Recurrent Kalman Network (RKN), can be used for any time-series data, similar to a LSTM (Hochreiter & Schmidhuber, 1997) but uses an explicit representation of uncertainty. As shown by our experiments, the RKN obtains much more accurate uncertainty estimates than an LSTM or Gated Recurrent Units (GRUs) (Cho et al., 2014) while also showing a slightly improved prediction performance and outperforms various recent generative models on an image imputation task. Philipp Becker, Harit Pandya, Gregor H. W. Gebhardt, Cheng Zhao 0002, C. James Taylor, Gerhard Neumann |
ICML | 2 |
| 2018 | Image Based Visual Servoing for Tumbling ObjectsabstractObjects in space often exhibit a tumbling motion around the major inertial axis. In this paper, we address the image based visual servoing of a robotic system towards an uncooperative tumbling object. In contrast to previous approaches that require explicit reconstruction of the object and an estimation of its velocity, we propose a novel controller that is able to minimize the feature error directly in image space. This is achieved by observing that the feature points on the tumbling object follow a circular path around the axis of rotation and their projection creates an elliptical track in the image plane. Our controller minimizes the error between this elliptical track and the desired features, such that at the desired pose the features lie on the circumference of the ellipse. The effectiveness of our framework is exhibited by implementing the algorithm in simulation as well on a mobile robot. P. Mithun, Harit Pandya, Ayush Gaud, Suril Vijaykumar Shah, K. Madhava Krishna |
IROS | 2 |
| 2018 | Combining Method of Alternating Projections and Augmented Lagrangian for Task Constrained Trajectory OptimizationabstractMotion planning for manipulators under task space constraints is difficult as it constrains the joint configurations to always lie on an implicitly defined manifold. It is possible to view task constrained motion planning as an optimization problem with non-linear equality constraints, which can be solved by general non-linear optimization techniques. In this paper, we present a novel custom optimizer which exploits the underlying structure present in many task constraints. At the core of our approach are some simple reformulations, which when coupled with the method of alternating projection, leads to an efficient convex optimization based routine for computing a feasible solution to the task constraints. We subsequently build on this result and use the concept of Augmented Lagrangian to guide the feasible solutions towards those that also minimize the user defined cost function. We show that the proposed optimizer is fully distributive and thus, can be easily parallelized. We validate our formulation on some common robotic benchmark problems. In particular, we show that the proposed optimizer achieves cyclic motion in the joint space corresponding to a similar nature trajectory in the task space. Furthermore, as a baseline, we compare the proposed optimizer with an off-the-shelf non-linear solver provide in open source package SciPy. We show that for similar task constraint residuals and smoothness cost, it can be upto more than three times faster than the SciPy alternative. Reza Ghabcheloo, Andreas Müller 0002, Harit Pandya |
IROS | 4 |
| 2017 | Exploring convolutional networks for end-to-end visual servoingabstractPresent image based visual servoing approaches rely on extracting hand crafted visual features from an image. Choosing the right set of features is important as it directly affects the performance of any approach. Motivated by recent breakthroughs in performance of data driven methods on recognition and localization tasks, we aim to learn visual feature representations suitable for servoing tasks in unstructured and unknown environments. In this paper, we present an end-to-end learning based approach for visual servoing in diverse scenes where the knowledge of camera parameters and scene geometry is not available a priori. This is achieved by training a convolutional neural network over color images with synchronised camera poses. Through experiments performed in simulation and on a quadrotor, we demonstrate the efficacy and robustness of our approach for a wide range of camera poses in both indoor as well as outdoor environments. Aseem Saxena, Harit Pandya, Gourav Kumar, Ayush Gaud, K. Madhava Krishna |
ICRA | 2 |
| 2017 | Pose induction for visual servoing to a novel object instanceabstractPresent visual servoing approaches are instance specific i.e. they control camera motion between two views of the same object. However, in practical scenarios where a robot is required to handle various instances of a category, classical visual servoing techniques are less suitable. We formulate across instance visual servoing as a pose induction and pose alignment problem. Initially, the desired pose given for any known instance is transferred to the novel instance through pose induction. Then the pose alignment problem is solved by estimating the current pose using the part aware keypoints reconstruction followed by a pose based visual servoing (PBVS) iteration. To tackle large variation in appearance across object instances in a category, we employ visual features that uniquely correspond to locations of object's parts in images. These part-aware keypoints are learned from annotated images using a convolutional neural network (CNN). Advantages of using such part-aware semantics are two-fold. Firstly, it conceals the illumination and textural variations from the visual servoing algorithm. Secondly, semantic keypoints enables us to match descriptors across instances accurately. We validate the efficacy of our approach through experiments in simulation as well as on a quadcopter. Our approach results in acceptable desired camera pose and smooth velocity profile. We also show results for large camera transformations with no overlap between current and desired pose for 3D objects, which is desirable in servoing context. Gourav Kumar, Harit Pandya, Ayush Gaud, K. Madhava Krishna |
IROS | 2 |
| 2016 | Discriminative learning based visual servoing across object instancesabstractClassical visual servoing approaches use visual features based on geometry of the object such as points, lines, region, etc. to attain the desired camera pose. However, geometrical features are not suited for visual servoing across different object instances due to large variations in appearance and shape. In this paper, we present a new framework for visual servoing across object instances. Our approach is based on a discriminative learning framework where the desired pose is estimated using previously seen examples. Specifically, we learn a binary classifier that separates the desired pose from all other poses for that object category. The classification error is then used to control the end-effector so that the desired pose is attained. We present controllers for linear, kernel and exemplar Support Vector Machine (SVM) and empirically discuss their performance in the visual servoing context. To address large intra-category variation in appearance, we propose a modified version of Histogram of Oriented Gradients (HOG) features for visual servoing. We show effective servoing across diverse instances over 3 object categories with zero terminal velocity and acceptable camera pose error at termination. Harit Pandya, K. Madhava Krishna, C. V. Jawahar |
ICRA | 1 |
| 2015 | Servoing across object instances: Visual servoing for object categoryabstractTraditional visual servoing is able to navigate a robotic system between two views of the same object. However, it is not designed to servo between views of different objects. In this paper, we consider a novel problem of servoing any instance (exemplar) of an object category to a desired pose (view) and propose a strategy to accomplish the task. We use features that semantically encode the locations of object parts and define the servoing error as the difference between positions of corresponding parts in the image space. Our controller is based on the linear combination of 3D models, such that the resulting model interpolates between the given and desired instances. We conducted our experiments on five different object categories in simulation framework and show that our approach achieves the desired pose with smooth trajectory. Furthermore, we show the performance gain achieved by using a linear combination of models (instances) vis a vis a controller that switches across models during servoing in terms of trajectory's length, smoothness and error in camera pose and image features. Harit Pandya, K. Madhava Krishna, C. V. Jawahar |
ICRA | 1 |
| 2014 | Identifying Ragas in Indian MusicabstractIn this work, we propose a method to identify the ragas of an Indian Carnatic music signal. This has several interesting applications in digital music indexing, recommendation and retrieval. However, this problem is hard due to (i) the absence of a fixed frequency for a note (ii) relative scale of notes (iii) oscillations around a note, and (iv) improvisations. In this work, we attempt the raga classification problem in a non-linear SVM framework using a combination of two kernels that represent the similarities of a music signal using two different features-pitch-class profile and n-gram distribution of notes. This differs from the previous pitch-class profile based approaches where the temporal information of notes is ignored. We evaluated the proposed approach on our own raga dataset and Comp Music dataset and show an improvement of 10.19% by combining the information from two features relevant to Indian Carnatic music. Vijay Kumar 0004, Harit Pandya, C. V. Jawahar |
ICPR | 2 |