EDBT 2026 Demo / reviewers in the wild / expert
Saurabh Gupta 0001
dblp:06/5843-1
· DBLP profile ↗
56ranked-venue papers
7as first author
28since 2021 · last 2025
0000-0002-1195-9028ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 14 since 2021Systems, architecture and hardware · 8 · 5 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PhysGen3D: Crafting a Miniature Interactive World from a Single ImageabstractEnvisioning physically plausible outcomes from a single image requires a deep understanding of the world’s dynamics. To address this, we introduce PhysGen3D, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geometric and semantic understanding with physics-based simulation, PhysGen3D creates an interactive 3D world from a static image, enabling us to "imagine" and simulate future scenarios based on user input. At its core, PhysGen3D estimates 3D shapes, poses, physical and lighting properties of objects, thereby capturing essential physical attributes that drive realistic object interactions. This framework allows users to specify precise initial conditions, such as object speed or material properties, for enhanced control over generated video outcomes. We evaluate PhysGen3D’s performance against closed-source state-of-the-art (SOTA) image-to-video models, including Pika, Kling, and Gen-3, showing PhysGen3D’s capacity to generate videos with realistic physics while offering greater flexibility and fine-grained control. Our results show that PhysGen3D achieves a unique balance of photorealism, physical plausibility, and user-driven interactivity, opening new possibilities for generating dynamic, physics-grounded video from an image. Project page: https://by-luckk.github.io/PhysGen3D. Hanxiao Jiang 0001, Saurabh Gupta 0001, Yunzhu Li, Shenlong Wang |
CVPR | 4 |
| 2025 | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday InteractionsabstractWe tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) Interaction Codebook: a VQVAE model to learn a latent codebook of hand poses and contact points, effectively tokenizing interaction trajectories, (2) Interaction Predictor: a transformer-decoder module to predict the interaction trajectory from test time inputs by using an indexer module to retrieve a latent affordance from the learned codebook. To train our model, we develop a data engine that extracts 3D hand poses and contact trajectories from the diverse HoloAssist dataset. We evaluate our model on a benchmark that is 2.5-10× larger than existing works, in terms of diversity of objects and interactions observed, and test for generalization of the model across object categories, action categories, tasks, and scenes. Experimental results show the effectiveness of our approach over transformer & diffusion baselines across all settings. Ben Lundell, Dmitry Andreychuk, David Forsyth, Saurabh Gupta 0001, Harpreet Sawhney |
CVPR | 5 |
| 2025 | AlphaOne: Reasoning Models Thinking Slow and Fast at Test TimeabstractJunyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, Huan Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Runpei Dong, Han Wang 0019, Xuying Ning, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta 0001, Huan Zhang 0001 |
EMNLP | 10 |
| 2025 | Visual Sync: Multi-Camera Synchronization via Cross-View Object MotionabstractToday, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross‑camera streams remains challenging. Existing methods assume controlled settings, specific targets, manual correction, or costly hardware.
We present VisualSync, an optimization framework based on multi‑view dynamics that aligns unposed, unsynchronized videos at millisecond accuracy. Our key insight is that any moving 3D point, when co‑visible in two cameras, obeys epipolar constraints once properly synchronized. To exploit this, VisualSync leverages off‑the‑shelf 3D reconstruction, feature matching, and dense tracking to extract tracklets, relative poses, and cross‑view correspondences. It then jointly minimizes the epipolar error to estimate each camera’s time offset. Experiments on four diverse, challenging datasets show that VisualSync outperforms baseline methods, achieving an average synchronization error below 130 ms. David Yifan Yao, Saurabh Gupta 0001, Shenlong Wang |
NeurIPS | 3 |
| 2024 | Bootstrapping Autonomous Driving Radars with Self-Supervised LearningabstractThe perception of autonomous vehicles using radars has attracted increased research interest due its ability to operate in fog and bad weather. However, training radar models is hindered by the cost and difficulty of annotating largescale radar data. To overcome this bottleneck, we propose a self-supervised learning framework to leverage the large amount of unlabeled radar data to pre-train radar only embeddings for self-driving perception tasks. The proposed method combines radar-to-radar and radar-to-vision contrastive losses to learn a general representation from unlabeled radar heatmaps paired with their corresponding camera images. When used for downstream object detection, we demonstrate that the proposed self-supervision framework can improve the accuracy of state-of-the-art supervised baselines by 5.8% in mAP. Code is available at https://github.com/yiduohao/Radical. Yiduo Hao, Sohrab Madani, Junfeng Guan, Mohammed Alloulah, Saurabh Gupta 0001, Haitham Hassanieh |
CVPR | 5 |
| 2024 | Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao |
ECCV (25) | 19 |
| 2024 | PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
Zhongzheng Ren, Saurabh Gupta 0001, Shenlong Wang |
ECCV (82) | 3 |
| 2024 | 3D Reconstruction of Objects in Hands Without Real World 3D Supervision
Matthew Chang, Matthew Jin, Ruisen Tu, Saurabh Gupta 0001 |
ECCV (78) | 5 |
| 2024 | Mitigating Perspective Distortion-Induced Shape Ambiguity in Image Crops
Saurabh Gupta 0001 |
ECCV (78) | 3 |
| 2024 | 3D Hand Pose Estimation in Everyday Egocentric Images
Ruisen Tu, Matthew Chang, Saurabh Gupta 0001 |
ECCV (78) | 4 |
| 2024 | Toward Control of Wheeled Humanoid Robots with Unknown Payloads: Equilibrium Point Estimation via Real-to-Sim AdaptationabstractModel-based controllers using a linearized model around the system’s equilibrium point is a common approach in the control of a wheeled humanoid due to their less computational load and ease of stability analysis. However, controlling a wheeled humanoid robot while it lifts an unknown object presents significant challenges, primarily due to the lack of knowledge in object dynamics. This paper presents a framework designed for predicting the new equilibrium point explicitly to control a wheeled-legged robot with unknown dynamics. We estimated the total mass and center of mass of the system from its response to initially unknown dynamics, then calculated the new equilibrium point accordingly. To avoid using additional sensors (e.g., force torque sensor) and reduce the effort of obtaining expensive real data, a data-driven approach is utilized with a novel real-to-sim adaptation. A more accurate nonlinear dynamics model, offering a closer representation of real-world physics, is injected into a rigid-body simulation for real-to-sim adaptation. The nonlinear dynamics model parameters were optimized using Particle Swarm Optimization. The efficacy of this framework was validated on a physical wheeled inverted pendulum, a simplified model of a wheeled-legged robot. The experimental results indicate that employing a more precise analytical model with optimized parameters significantly reduces the gap between simulation and reality, thus improving the efficiency of a model-based controller in controlling a wheeled robot with unknown dynamics. Donghoon Baek, Youngwoo Sim 0001, Amartya Purushottam, Saurabh Gupta 0001, João Ramos 0004 |
IROS | 4 |
| 2024 | Estimating Perceptual Uncertainty to Predict Robust Motion PlansabstractA typical sense-plan-act robotics pipeline is brittle due to the inherent inaccuracies in the output of the sensing module and the lack of awareness of the planning module to those inaccuracies. This paper develops a framework to predict uncertainty estimates for neural network-based vision models used for state estimation in robotics pipelines. Our uncertainty estimates are based directly on the image observation data and are explicitly trained to model the error distribution on a held-out calibration set. We also demonstrate how predicted uncertainties can be used to select robust control strategies. We conduct experiments on the mobile manipulation problem of articulating everyday objects (e.g. opening a cupboard) and demonstrate the quality of estimated uncertainty and its downstream impact on robustness of inferred control strategies. Michelle Zhang, Saurabh Gupta 0001 |
IROS | 3 |
| 2024 | Around the Corner mmWave Imaging in Practical EnvironmentsabstractWe present the design, implementation, and evaluation of RFlect, a mmWave imaging system capable of producing around-the-corner high-resolution images in practical environments. RFlect leverages signals reflected off complex surfaces (e.g., poles, concave surfaces, or composition of multiple surfaces) to image objects that are not in the RF line-of-sight. RFlect models the reflections and introduces reconstruction algorithms for different types of surfaces. It also leverages a novel method for precisely mapping the location and geometry of the reflecting surface. We also derive the theoretical resolution and coverage for different reflecting surface geometries. We built a prototype of RFlect and performed extensive evaluations to demonstrate its ability to reconstruct the shape of objects around the corner, with an average Chamfer Distance of 2cm and 3D F-Score of 88.6%. Laura Dodds, Hailan Shanbhag, Junfeng Guan, Saurabh Gupta 0001, Haitham Hassanieh |
MobiCom | 4 |
| 2023 | Building Rearticulable Models for Arbitrary 3D Objects from 4D Point CloudsabstractWe build rearticulable models for arbitrary everyday man-made objects containing an arbitrary number of parts that are connected together in arbitrary ways via 1 degree-of-freedom joints. Given point cloud videos of such everyday objects, our method identifies the distinct object parts, what parts are connected to what other parts, and the properties of the joints connecting each part pair. We do this by jointly optimizing the part segmentation, transformation, and kinematics using a novel energy minimization frame-work. Our inferred animatable models, enables retargeting to novel poses with sparse point correspondences guidance. We test our method on a new articulating robot dataset, and the Sapiens dataset with common daily objects. Experiments show that our method outperforms two leading prior works on various metrics. Saurabh Gupta 0001, Shenlong Wang |
CVPR | 2 |
| 2023 | Exploiting Virtual Array Diversity for Accurate Radar DetectionabstractUsing millimeter-wave radars as a perception sensor provides self-driving cars with robust sensing capability in adverse weather. However, mmWave radars currently lack sufficient spatial resolution for semantic scene understanding. This paper introduces Radatron++, a system leverages cascaded MIMO (Multiple-Input Multiple-Output) radar to achieve accurate vehicle detection for self-driving cars. We develop a novel hybrid radar processing and deep learning approach to leverage the 10× finer angular resolution while combating unique challenges of cascaded MIMO radars. We train and evaluate Radatron++ with a novel cascaded radar dataset. Radatron++ achieves 93.9% and 58.5% Average Precisions with 0.5 and 0.75 Intersection over Union thresholds respectively in 2D bounding box detection, outperforming prior work using low-resolution radars by 9.3% and 18.1% respectively. Junfeng Guan, Sohrab Madani, Waleed Ahmed, Samah Hussein, Saurabh Gupta 0001, Haitham Hassanieh |
ICASSP | 5 |
| 2023 | ContactGen: Generative Contact Modeling for Grasp GenerationabstractThis paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part. Given an input object, we propose a conditional generative model to predict ContactGen and adopt model-based optimization to predict diverse and geometrically feasible grasps. Experimental results demonstrate our method can generate high-fidelity and diverse human grasps for various objects. Jimei Yang, Saurabh Gupta 0001, Shenlong Wang |
ICCV | 4 |
| 2023 | One-shot Visual Imitation via Attributed Waypoints and Demonstration AugmentationabstractIn this paper, we analyze the behavior of existing techniques and design new solutions for the problem of one-shot visual imitation. In this setting, an agent must solve a novel instance of a novel task given just a single visual demonstration. Our analysis reveals that current methods fall short because of three errors: the DAgger problem arising from purely offline training, last centimeter errors in interacting with objects, and mis-fitting to the task context rather than to the actual task. This motivates the design of our modular approach where we a) separate out task inference (what to do) from task execution (how to do it), and b) develop data augmentation and generation techniques to mitigate mis-fitting. The former allows us to leverage hand-crafted motor primitives for task execution which side-steps the DAgger problem and last centimeter errors, while the latter gets the model to focus on the task rather than the task context. Our model gets 100% and 48% success rates on two recent benchmarks, improving upon the current state-of-the-art by absolute 90% and 20% respectively. Matthew Chang, Saurabh Gupta 0001 |
ICRA | 2 |
| 2023 | Predicting Motion Plans for Articulating Everyday ObjectsabstractMobile manipulation tasks such as opening a door, pulling open a drawer, or lifting a toilet seat require constrained motion of the end-effector under environmental and task constraints. This, coupled with partial information in novel environments, makes it challenging to employ classical motion planning approaches at test time. Our key insight is to cast it as a learning problem to leverage past experience of solving similar planning problems to directly predict motion plans for mobile manipulation tasks in novel situations at test time. To enable this, we develop a simulator, ArtObjSim, that simulates articulated objects placed in real scenes. We then introduce$\mathbf{SeqIK}+\theta_{0}$, a fast and flexible representation for motion plans. Finally, we learn models that use$\mathbf{SeqIK}+\theta_{0}$to quickly predict motion plans for articulating novel objects at test time. Experimental evaluation shows improved speed and accuracy at generating motion plans than pure search-based methods and pure learning methods. Max E. Shepherd, Saurabh Gupta 0001 |
ICRA | 3 |
| 2023 | Contactless Material Identification with Millimeter Wave VibrometryabstractThis paper introduces RFVibe, a system that enables contactless material and object identification through the fusion of millimeter wave wireless signals with acoustic signals. In particular, RFVibe plays an audio sound next to the object that generates micro-vibrations in the object. These micro-vibrations can be captured by shining a millimeter wave radar signal on the object and analyzing the phase of the reflected wireless signal. RFVibe can then extract several features including resonance frequencies and vibration modes, damping time of vibrations, and wireless reflection coefficients. These features are then used to enable more accurate identification, with a step towards generalizing towards different setups and locations. We implement RFVibe using an off-the-shelf millimeter-wave radar and an acoustic speaker. We evaluate it on 23 objects of 7 material types (Metal, Wood, Ceramic, Glass, Plastic, Cardboard, and Foam), obtaining 81.3% accuracy for material classification, a 30% improvement over prior work. RFVibe is able to classify with reasonable accuracy in scenarios that it has not encountered before, including different locations, angles, boundary conditions, and objects. Hailan Shanbhag, Sohrab Madani, Akhil Isanaka, Deepak Nair, Saurabh Gupta 0001, Haitham Hassanieh |
MobiSys | 5 |
| 2023 | Poster: Contactless Material Identification with Millimeter Wave VibrometryabstractThis paper introduces RFVibe, a system that enables contactless material and object identification through the fusion of millimeter wave wireless signals with acoustic signals. In particular, RFVibe plays an audio sound next to the object that generates micro-vibrations in the object. These micro-vibrations can be captured by shining a millimeter wave radar signal on the object and analyzing the phase of the reflected wireless signal. RFVibe can then extract several features including resonance frequencies and vibration modes, damping time of vibrations, and wireless reflection coefficients. These features are then used to enable more accurate identification, with a step towards generalizing towards different setups and locations. We implement RFVibe using an off-the-shelf millimeter-wave radar and an acoustic speaker. We evaluate it on 23 objects of 7 material types (Metal, Wood, Ceramic, Glass, Plastic, Cardboard, and Foam), obtaining 81.3% accuracy for material classification, a 30% improvement over prior work. RFVibe is able to classify with reasonable accuracy in scenarios that it has not encountered before, including different locations, angles, boundary conditions, and objects. Hailan Shanbhag, Sohrab Madani, Akhil Isanaka, Deepak Nair, Saurabh Gupta 0001, Haitham Hassanieh |
MobiSys | 5 |
| 2023 | Look Ma, No Hands! Agent-Environment Factorization of Egocentric VideosabstractThe analysis and use of egocentric videos for robotics tasks is made challenging by occlusion and the visual mismatch between the human hand and a robot end-effector. Past work views the human hand as a nuisance and removes it from the scene. However, the hand also provides a valuable signal for learning. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving the in-painting quality in egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos. Matthew Chang, Saurabh Gupta 0001 |
NeurIPS | 3 |
| 2022 | Human Hands as Probes for Interactive Object UnderstandingabstractInteractive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos. We demonstrate that observation of what human hands interact with and how can provide both the relevant data and the necessary supervision. Attending to hands, readily localizes and stabilizes active objects for learning and reveals places where interactions with objects occur. Analyzing the hands shows what we can do to objects and how. We apply these basic principles on the EPIC-KITCHENS dataset, and successfully learn state-sensitive features, and object affordances (regions of interaction and afforded grasps), purely by observing hands in egocentric videos. Mohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh Gupta 0001 |
CVPR | 4 |
| 2022 | Radatron: Accurate Detection Using Multi-resolution Cascaded MIMO Radar
Sohrab Madani, Jayden Guan, Waleed Ahmed, Saurabh Gupta 0001, Haitham Hassanieh |
ECCV (39) | 4 |
| 2022 | TIDEE: Tidying Up Novel Rooms Using Visuo-Semantic Commonsense Priors
Gabriel Sarch, Zhaoyuan Fang, Adam W. Harley, Paul Schydlo, Michael J. Tarr, Saurabh Gupta 0001, Katerina Fragkiadaki |
ECCV (39) | 6 |
| 2022 | Learning Value Functions from Undirected State-only Experience
Matthew Chang, Saurabh Gupta 0001 |
ICLR | 3 |
| 2022 | On-Device CPU Scheduling for Robot SystemsabstractRobots have to take highly responsive real-time actions, driven by complex decisions involving a pipeline of sensing, perception, planning, and reaction tasks. These tasks must be scheduled on resource-constrained devices such that the performance goals and the requirements of the application are met. This is a difficult problem that requires handling multiple scheduling dimensions, and variations in computational resource usage and availability. In practice, system designers manually tune parameters for their specific hardware and application, which results in poor generalization and increases the development burden. In this work, we highlight the emerging need for scheduling CPU resources at runtime in robot systems. We use robot navigation as a case-study to understand the key scheduling requirements for such systems. Armed with this understanding, we develop a CPU scheduling framework, Catan, that dynamically schedules compute resources across different components of an app so as to meet the specified application requirements. Through experiments with a prototype implemented on ROS, we show the impact of system scheduling on meeting the application's performance goals, and how Catan dynamically adapts to runtime variations. Aditi Partap, Samuel Grayson, Muhammad Huzaifa, Sarita V. Adve, Brighten Godfrey, Saurabh Gupta 0001, Kris Hauser, Radhika Mittal |
IROS | 6 |
| 2021 | SEAL: Self-supervised Embodied Active Learning using Exploration and 3D ConsistencyabstractIn this paper, we explore how we can build upon the data and models of Internet images and use them to adapt to robot vision without requiring any extra labels. We present a framework called Self-supervised Embodied Active Learning (SEAL). It utilizes perception models trained on internet images to learn an active exploration policy. The observations gathered by this exploration policy are labelled using 3D consistency and used to improve the perception model. We build and utilize 3D semantic maps to learn both action and perception in a completely self-supervised manner. The semantic map is used to compute an intrinsic motivation reward for training the exploration policy and for labelling the agent observations using spatio-temporal 3D consistency and label propagation. We demonstrate that the SEAL framework can be used to close the action-perception loop: it improves object detection and instance segmentation performance of a pretrained perception model by just moving around in training environments and the improved perception model can be used to improve Object Goal Navigation. Devendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta 0001, Jitendra Malik, Ruslan Salakhutdinov |
NeurIPS | 3 |
| 2021 | On the Use of ML for Blackbox System Performance Prediction
Silvery D. Fu, Saurabh Gupta 0001, Radhika Mittal, Sylvia Ratnasamy |
NSDI | 2 |
| 2020 | Neural Topological SLAM for Visual NavigationabstractThis paper studies the problem of image-goal navigation which involves navigating to the location indicated by a goal image in a novel previously unseen environment. To tackle this problem, we design topological representations for space that effectively leverage semantics and afford approximate geometric reasoning. At the heart of our representations are nodes with associated semantic features, that are interconnected using coarse geometric information. We describe supervised learning-based algorithms that can build, maintain and use such representations under noisy actuation. Experimental study in visually and physically realistic simulation suggests that our method builds effective representations that capture structural regularities and efficiently solve long-horizon navigation problems. We observe a relative improvement of more than 50% over existing methods that study this task. Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta 0001, Saurabh Gupta 0001 |
CVPR | 4 |
| 2020 | Use the Force, Luke! Learning to Predict Physical Forces by Simulating EffectsabstractWhen we humans look at a video of human-object interaction, we can not only infer what is happening but we can even extract actionable information and imitate those interactions. On the other hand, current recognition or geometric approaches lack the physicality of action representation. In this paper, we take a step towards more physical understanding of actions. We address the problem of inferring contact points and the physical forces from videos of humans interacting with objects. One of the main challenges in tackling this problem is obtaining ground-truth labels for forces. We sidestep this problem by instead using a physics simulator for supervision. Specifically, we use a simulator to predict effects, and enforce that estimated forces must lead to same effect as depicted in the video. Our quantitative and qualitative results show that (a) we can predict meaningful forces from videos whose effects lead to accurate imitation of the motions observed, (b) by jointly optimizing for contact point and force prediction, we can improve the performance on both tasks in comparison to independent training, and (c) we can learn a representation from this model that generalizes to novel objects using few shot examples. Kiana Ehsani, Shubham Tulsiani, Saurabh Gupta 0001, Ali Farhadi, Abhinav Gupta 0001 |
CVPR | 3 |
| 2020 | Through Fog High-Resolution Imaging Using Millimeter Wave RadarabstractThis paper demonstrates high-resolution imaging using millimeter Wave (mmWave) radars that can function even in dense fog. We leverage the fact that mmWave signals have favorable propagation characteristics in low visibility conditions, unlike optical sensors like cameras and LiDARs which cannot penetrate through dense fog. Millimeter-wave radars, however, suffer from very low resolution, specularity, and noise artifacts. We introduce HawkEye, a system that leverages a cGAN architecture to recover high-frequency shapes from raw low-resolution mmWave heat-maps. We propose a novel design that addresses challenges specific to the structure and nature of the radar signals involved. We also develop a data synthesizer to aid with large-scale dataset generation for training. We implement our system on a custom-built mmWave radar platform and demonstrate performance improvement over both standard mmWave radars and other competitive baselines. Junfeng Guan, Sohrab Madani, Suraj Jog, Saurabh Gupta 0001, Haitham Hassanieh |
CVPR | 4 |
| 2020 | Semantic Curiosity for Active Visual Learning
Devendra Singh Chaplot, Helen Jiang, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ECCV (6) | 3 |
| 2020 | Aligning Videos in Space and TimeabstractIn this paper, we focus on the task of extracting visual correspondences across videos. Given a query video clip from an action class, we aim to align it with training videos in space and time. Obtaining training data for such a fine-grained alignment task is challenging and often ambiguous. Hence, we propose a novel alignment procedure that learns such correspondence in space and time via cross video cycle-consistency. During training, given a pair of videos, we compute cycles that connect patches in a given frame in the first video by matching through frames in the second video. Cycles that connect overlapping patches together are encouraged to score higher than cycles that connect non-overlapping patches. Our experiments on the Penn Action and Pouring datasets demonstrate that the proposed method can successfully learn to correspond semantically similar patches across videos, and learns representations that are sensitive to object and action states. Senthil Purushwalkam, Tian Ye 0006, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ECCV (26) | 3 |
| 2020 | Learning To Explore Using Active Neural SLAM
Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta 0001, Abhinav Gupta 0001, Ruslan Salakhutdinov |
ICLR | 3 |
| 2020 | Intrinsic Motivation for Encouraging Synergistic Behavior
Rohan Chitnis, Shubham Tulsiani, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ICLR | 3 |
| 2020 | Learning to Move with Affordance Maps
William Qi, Ravi Teja Mullapudi, Saurabh Gupta 0001, Deva Ramanan |
ICLR | 3 |
| 2020 | Efficient Bimanual Manipulation Using Learned Task SchemasabstractWe address the problem of effectively composing skills to solve sparse-reward tasks in the real world. Given a set of parameterized skills (such as exerting a force or doing a top grasp at a location), our goal is to learn policies that invoke these skills to efficiently solve such tasks. Our insight is that for many tasks, the learning process can be decomposed into learning a state-independent task schema (a sequence of skills to execute) and a policy to choose the parameterizations of the skills in a state-dependent manner. For such tasks, we show that explicitly modeling the schema's state-independence can yield significant improvements in sample efficiency for model-free reinforcement learning algorithms. Furthermore, these schemas can be transferred to solve related tasks, by simply re-learning the parameterizations with which the skills are invoked. We find that doing so enables learning to solve sparse-reward tasks on real-world robotic systems very efficiently. We validate our approach experimentally over a suite of robotic bimanual manipulation tasks, both in simulation and on real hardware. See videos at http://tinyurl.com/chitnis-schema. Rohan Chitnis, Shubham Tulsiani, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ICRA | 3 |
| 2020 | Semantic Visual Navigation by Watching YouTube VideosabstractSemantic cues and statistical regularities in real-world environment layouts can improve efficiency for navigation in novel environments. This paper learns and leverages such semantic cues for navigating to objects of interest in novel environments, by simply watching YouTube videos. This is challenging because YouTube videos don't come with labels for actions or goals, and may not even showcase optimal behavior. Our method tackles these challenges through the use of Q-learning on pseudo-labeled transition quadruples (image, action, next image, reward). We show that such off-policy Q-learning from passive data is able to learn meaningful semantic cues for navigation. These cues, when used in a hierarchical navigation policy, lead to improved efficiency at the ObjectGoal task in visually realistic simulations. We observe a relative improvement of 15-83% over end-to-end RL, behavior cloning, and classical methods, while using minimal direct interaction. Matthew Chang, Saurabh Gupta 0001 |
NeurIPS | 3 |
| 2020 | Cognitive Mapping and Planning for Visual Navigation
Saurabh Gupta 0001, Varun Tolani, James Davidson, Sergey Levine, Rahul Sukthankar, Jitendra Malik |
Int. J. Comput. Vis. | 1 |
| 2019 | Learning Exploration Policies for Navigation
Tao Chen 0046, Saurabh Gupta 0001, Abhinav Gupta 0001 |
ICLR (Poster) | 2 |
| 2019 | Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic DataabstractThe ability to segment unknown objects in depth images has potential to enhance robot skills in grasping and object tracking. Recent computer vision research has demonstrated that Mask R-CNN can be trained to segment specific categories of objects in RGB images when massive hand-labeled datasets are available. As generating these datasets is time-consuming, we instead train with synthetic depth images. Many robots now use depth sensors, and recent results suggest training on synthetic depth data can transfer successfully to the real world. We present a method for automated dataset generation and rapidly generate a synthetic training dataset of 50,000 depth images and 320,000 object masks using simulated heaps of 3D CAD models. We train a variant of Mask R-CNN with domain randomization on the generated dataset to perform category-agnostic instance segmentation without any hand-labeled data and we evaluate the trained network, which we refer to as Synthetic Depth (SD) Mask R-CNN, on a set of real, high-resolution depth images of challenging, densely-cluttered bins containing objects with highly-varied geometry. SD Mask R-CNN outperforms point cloud clustering baselines by an absolute 15% in Average Precision and 20% in Average Recall on COCO benchmarks, and achieves performance levels similar to a Mask R-CNN trained on a massive, hand-labeled RGB dataset and fine-tuned on real images from the experimental setup. We deploy the model in an instance-specific grasping pipeline to demonstrate its usefulness in a robotics application. Code, the synthetic training dataset, and supplementary material are available at https://bit.ly/2letCuE. Michael Danielczuk, Matthew Matl, Saurabh Gupta 0001, Andrew Li, Andrew Lee 0002, Jeffrey Mahler, Kenneth Y. Goldberg |
ICRA | 3 |
| 2018 | Factoring Shape, Pose, and Layout From the 2D Image of a 3D SceneabstractThe goal of this paper is to take a single 2D image of a scene and recover the 3D structure in terms of a small set of factors: a layout representing the enclosing surfaces as well as a set of objects represented in terms of shape and pose. We propose a convolutional neural network-based approach to predict this representation and benchmark it on a large dataset of indoor scenes. Our experiments evaluate a number of practical design questions, demonstrate that we can infer this representation, and quantitatively and qualitatively demonstrate its merits compared to alternate representations. Shubham Tulsiani, Saurabh Gupta 0001, David F. Fouhey, Alexei A. Efros, Jitendra Malik |
CVPR | 2 |
| 2018 | Visual Memory for Robust Path FollowingabstractHumans routinely retrace a path in a novel environment both forwards and backwards despite uncertainty in their motion. In this paper, we present an approach for doing so. Given a demonstration of a path, a first network generates an abstraction of the path. Equipped with this abstraction, a second network then observes the world and decides how to act in order to retrace the path under noisy actuation and a changing environment. The two networks are optimized end-to-end at training time. We evaluate the method in two realistic simulators, performing path following both forwards and backwards. Our experiments show that our approach outperforms both a classical approach to solving this task as well as a number of other baselines. Ashish Kumar 0007, Saurabh Gupta 0001, David F. Fouhey, Sergey Levine, Jitendra Malik |
NeurIPS | 2 |
| 2017 | Cognitive Mapping and Planning for Visual NavigationabstractWe introduce a neural architecture for navigation in novel environments. Our proposed architecture learns to map from first-person views and plans a sequence of actions towards goals in the environment. The Cognitive Mapper and Planner (CMP) is based on two key ideas: a) a unified joint architecture for mapping and planning, such that the mapping is driven by the needs of the planner, and b) a spatial memory with the ability to plan given an incomplete set of observations about the world. CMP constructs a top-down belief map of the world and applies a differentiable neural net planner to produce the next action at each time step. The accumulated belief of the world enables the agent to track visited regions of the environment. Our experiments demonstrate that CMP outperforms both reactive strategies and standard memory-based architectures and performs well in novel environments. Furthermore, we show that CMP can also achieve semantically specified goals, such as "go to a chair". Saurabh Gupta 0001, James Davidson, Sergey Levine, Rahul Sukthankar, Jitendra Malik |
CVPR | 1 |
| 2016 | Cross Modal Distillation for Supervision TransferabstractIn this work we propose a technique that transfers supervision between images from different modalities. We use learned representations from a large labeled modality as supervisory signal for training representations for a new unlabeled paired modality. Our method enables learning of rich representations for unlabeled modalities and can be used as a pre-training procedure for new modalities with limited labeled data. We transfer supervision from labeled RGB images to unlabeled depth and optical flow images and demonstrate large improvements for both these cross modal supervision transfers. Saurabh Gupta 0001, Judy Hoffman, Jitendra Malik |
CVPR | 1 |
| 2016 | Learning with Side Information through Modality HallucinationabstractWe present a modality hallucination architecture for training an RGB object detection model which incorporates depth side information at training time. Our convolutional hallucination network learns a new and complementary RGB image representation which is taught to mimic convolutional mid-level features from a depth network. At test time images are processed jointly through the RGB and hallucination networks to produce improved detection performance. Thus, our method transfers information commonly extracted from depth training data to a network which can extract that information from the RGB counterpart. We present results on the standard NYUDv2 dataset and report improvement on the RGB detection task. Judy Hoffman, Saurabh Gupta 0001, Trevor Darrell |
CVPR | 2 |
| 2016 | Cross-modal adaptation for RGB-D detectionabstractIn this paper we propose a technique to adapt convolutional neural network (CNN) based object detectors trained on RGB images to effectively leverage depth images at test time to boost detection performance. Given labeled depth images for a handful of categories we adapt an RGB object detector for a new category such that it can now use depth images in addition to RGB images at test time to produce more accurate detections. Our approach is built upon the observation that lower layers of a CNN are largely task and category agnostic and domain specific while higher layers are largely task and category specific while being domain agnostic. We operationalize this observation by proposing a mid-level fusion of RGB and depth CNNs. Experimental evaluation on the challenging NYUD2 dataset shows that our proposed adaptation technique results in an average 21% relative improvement in detection performance over an RGB-only baseline even when no depth training data is available for the particular category evaluated. We believe our proposed technique will extend advances made in computer vision to RGB-D data leading to improvements in performance at little additional annotation effort. Judy Hoffman, Saurabh Gupta 0001, Jian Leong, Sergio Guadarrama, Trevor Darrell |
ICRA | 2 |
| 2016 | The three R's of computer vision: Recognition, reconstruction and reorganization
Jitendra Malik, Pablo Andrés Arbeláez, João Carreira 0001, Katerina Fragkiadaki, Ross B. Girshick, Georgia Gkioxari, Saurabh Gupta 0001, Bharath Hariharan, Abhishek Kar, Shubham Tulsiani |
Pattern Recognit. Lett. | 7 |
| 2015 | From captions to visual concepts and backabstractThis paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time. Hao Fang 0002, Saurabh Gupta 0001, Forrest N. Iandola, Rupesh Kumar Srivastava, Li Deng 0001, Piotr Dollár, Jianfeng Gao 0001, Xiaodong He 0001, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, Geoffrey Zweig |
CVPR | 2 |
| 2015 | Aligning 3D models to RGB-D images of cluttered scenesabstractThe goal of this work is to represent objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instances in the scene and then using a convolutional neural network (CNN) to predict the pose of the object. This CNN is trained using pixel surface normals in images containing renderings of synthetic objects. When tested on real data, our method outperforms alternative algorithms trained on real data. We then use this coarse pose estimate along with the inferred pixel support to align a small number of prototypical models to the data, and place into the scene the model that fits best. We observe a 48% relative improvement in performance at the task of 3D detection over the current state-of-the-art [34], while being an order of magnitude faster. Saurabh Gupta 0001, Pablo Andrés Arbeláez, Ross B. Girshick, Jitendra Malik |
CVPR | 1 |
| 2015 | Indoor Scene Understanding with RGB-D Images: Bottom-up Segmentation, Object Detection and Semantic Segmentation
Saurabh Gupta 0001, Pablo Andrés Arbeláez, Ross B. Girshick, Jitendra Malik |
Int. J. Comput. Vis. | 1 |
| 2014 | Learning Rich Features from RGB-D Images for Object Detection and Segmentation
Saurabh Gupta 0001, Ross B. Girshick, Pablo Andrés Arbeláez, Jitendra Malik |
ECCV (7) | 1 |
| 2013 | Perceptual Organization and Recognition of Indoor Scenes from RGB-D ImagesabstractWe address the problems of contour detection, bottom-up grouping and semantic segmentation using RGB-D data. We focus on the challenging setting of cluttered indoor scenes, and evaluate our approach on the recently introduced NYU-Depth V2 (NYUD2) dataset [27]. We propose algorithms for object boundary detection and hierarchical segmentation that generalize the gPb-ucm approach of [2] by making effective use of depth information. We show that our system can label each contour with its type (depth, normal or albedo). We also propose a generic method for long-range amodal completion of surfaces and show its effectiveness in grouping. We then turn to the problem of semantic segmentation and propose a simple approach that classifies super pixels into the 40 dominant object categories in NYUD2. We use both generic and class-specific features to encode the appearance and geometry of objects. We also show how our approach can be used for scene classification, and how this contextual information in turn improves object recognition. In all of these tasks, we report significant improvements over the state-of-the-art. Saurabh Gupta 0001, Pablo Andrés Arbeláez, Jitendra Malik |
CVPR | 1 |
| 2013 | A Data Driven Approach for Algebraic Loop Invariants
Rahul Sharma 0001, Saurabh Gupta 0001, Bharath Hariharan, Alex Aiken, Percy Liang, Aditya V. Nori |
ESOP | 2 |
| 2013 | Verification as Learning Geometric Concepts
Rahul Sharma 0001, Saurabh Gupta 0001, Bharath Hariharan, Alex Aiken, Aditya V. Nori |
SAS | 2 |
| 2012 | Semantic segmentation using regions and partsabstractWe address the problem of segmenting and recognizing objects in real world images, focusing on challenging articulated categories such as humans and other animals. For this purpose, we propose a novel design for region-based object detectors that integrates efficiently top-down information from scanning-windows part models and global appearance cues. Our detectors produce class-specific scores for bottom-up regions, and then aggregate the votes of multiple overlapping candidates through pixel classification. We evaluate our approach on the PASCAL segmentation challenge, and report competitive performance with respect to current leading techniques. On VOC2010, our method obtains the best results in 6/20 categories and the highest performance on articulated objects. Pablo Andrés Arbeláez, Bharath Hariharan, Chunhui Gu, Saurabh Gupta 0001, Lubomir D. Bourdev, Jitendra Malik |
CVPR | 4 |