Nikhil Chavan Dafle

dblp:151/9654 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 6 since 2021Systems, architecture and hardware · 12 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FineControlNet: Fine-level Text Control for Image Generation with Spatially Aligned Text Control Injection
abstract
Recently introduced ControlNet has the ability to steer the text-driven image generation process with geometric input such as 2D human pose, or edge representations. While ControlNet provides control over the geometric form of the instances in the generated image, it lacks the capability to dictate the visual appearance of each instance. We present FineControlNet to provide fine control over each instance's appearance while maintaining the pose control capability. Specifically, we develop and demonstrate FineControlNet with geometric control via human pose images and appearance control via instance-level text prompts. The spatial alignment of 2D poses and instance-specific text prompts in latent space enables the fine control of multiple instances. We evaluate the performance of FineControlNet with rigorous comparison against state-of-the-art pose-conditioned text-to-image diffusion models. FineControlNet achieves superior performance in generating high quality images that follow instance-specific controls. We will release the code and the dataset.
Hongsuk Choi, Isaac Kasahara, Kazim Selim Engin, Moritz A. Graule, Nikhil Chavan Dafle, Volkan Isler
WACV5
2024 VioLA: Aligning Videos to 2D LiDAR Scans
abstract
We study the problem of aligning a video that captures a local portion of an environment to the 2D LiDAR scan of the entire environment. We introduce a method (VioLA) that starts with building a semantic map of the local scene from the image sequence, then extracts points at a fixed height for registering to the LiDAR map. Due to reconstruction errors or partial coverage of the camera scan, the reconstructed semantic map may not contain sufficient information for registration. To address this problem, VioLA makes use of a pre-trained text-to-image inpainting model paired with a depth completion model for filling in the missing scene content in a geometrically consistent fashion to support pose registration. We evaluate VioLA on two real-world RGB-D benchmarks, as well as a self-captured dataset of a large office scene. Notably, our proposed scene completion module improves the pose registration performance by up to 20%.
Jun-Jee Chao, Kazim Selim Engin, Nikhil Chavan Dafle, Bhoram Lee, Volkan Isler
ICRA3
2024 HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image
abstract
This paper presents a method to learn hand-object interaction prior for reconstructing a 3D hand-object scene from a single RGB image. The inference as well as training-data generation for 3D hand-object scene reconstruction is challenging due to the depth ambiguity of a single image and occlusions by the hand and object. We turn this challenge into an opportunity by utilizing the hand shape to constrain the possible relative configuration of the hand and object geometry. We design a generalizable implicit function, HandNeRF, that explicitly encodes the correlation of the 3D hand shape features and 2D object features to predict the hand and object scene geometry. With experiments on real-world datasets, we show that HandNeRF can reconstruct hand-object scenes of novel grasp configurations more accurately than comparable methods. Moreover, we demonstrate that object reconstruction from HandNeRF ensures more accurate execution of downstream tasks, such as grasping for robotic hand-over.
Hongsuk Choi, Nikhil Chavan Dafle, Jiacheng Yuan, Volkan Isler, Hyunsoo Park
ICRA2
2024 RIC: Rotate-Inpaint-Complete for Generalizable Scene Reconstruction
abstract
General scene reconstruction refers to the task of estimating the full 3D geometry and texture of a scene containing previously unseen objects. In many practical applications such as AR/VR, autonomous navigation, and robotics, only a single view of the scene may be available, making the scene reconstruction task challenging. In this paper, we present a method for scene reconstruction by structurally breaking the problem into two steps: rendering novel views via inpainting and 2D to 3D scene lifting. Specifically, we leverage the generalization capability of large visual language models (DALL•E 2) to inpaint the missing areas of scene color images rendered from different views. Next, we lift these inpainted images to 3D by predicting normals of the inpainted image and solving for the missing depth values. By predicting for normals instead of depth directly, our method allows for robustness to changes in depth distributions and scale. With rigorous quantitative evaluation, we show that our method outperforms multiple baselines while providing generalization to novel objects and scenes. Code and data is available at https://samsunglabs.github.io/RIC-project-page/.
Isaac Kasahara, Kazim Selim Engin, Nikhil Chavan Dafle, Shuran Song, Volkan Isler
ICRA4
2023 Pick2Place: Task-aware 6DoF Grasp Estimation via Object-Centric Perspective Affordance
abstract
The choice of a grasp plays a critical role in the success of downstream manipulation tasks. Consider a task of placing an object in a cluttered scene; the majority of possible grasps may not be suitable for the desired placement. In this paper, we study the synergy between the picking and placing of an object in a cluttered scene to develop an algorithm for task-aware grasp estimation. We present an object-centric action space that encodes the relationship between the geometry of the placement scene and the object to be placed in order to provide placement affordance maps directly from perspective views of the placement scene. This action space enables the computation of a one-to-one mapping between the placement and picking actions allowing the robot to generate a diverse set of pick-and-place proposals and to optimize for a grasp under other task constraints such as robot kinematics and collision avoidance. With experiments both in simulation and on a real robot we demonstrate that with our method, the robot is able to successfully complete the task of placement-aware grasping with over 89 % accuracy in such a way that generalizes to novel objects and scenes.
Zhanpeng He, Nikhil Chavan Dafle, Jinwook Huh, Shuran Song, Volkan Isler
ICRA2
2023 Real-Time Simultaneous Multi-Object 3D Shape Reconstruction, 6DoF Pose Estimation and Dense Grasp Prediction
abstract
In this paper, we present a realtime method for simultaneous object-level scene understanding and grasp prediction. Specifically, given a single RGBD image of a scene, our method localizes all the objects in the scene and for each object, it generates the following: full 3D shape, scale, pose with respect to the camera frame, and a dense set of feasible grasps. The main advantage of our method is its computation speed as it avoids sequential perception and grasp planning. With detailed quantitative analysis of reconstruction quality and grasp accuracy, we show that our method delivers competitive performance compared to the state-of-the-art methods, while providing fast inference at 30 frames per second speed.
Nikhil Chavan Dafle, Isaac Kasahara, Kazim Selim Engin, Jinwook Huh, Volkan Isler
IROS2
2022 Simultaneous Object Reconstruction and Grasp Prediction using a Camera-centric Object Shell Representation
abstract
Being able to grasp objects is a fundamental component of most robotic manipulation systems. In this paper, we present a new approach to simultaneously reconstruct a mesh and a dense grasp quality map of an object from a depth image. At the core of our approach is a novel camera-centric object representation called the “object shell” which is composed of an observed “entry image” and a predicted “exit image”. We present an image-to-image residual ConvNet architecture in which the object shell and a grasp-quality map are predicted as separate output channels. The main advantage of the shell representation and the corresponding neural network architecture, ShellGrasp-Net, is that the input-output pixel correspondences in the shell representation are explicitly represented in the architecture. We show that this coupling yields superior generalization capabilities for object reconstruction and accurate grasp quality estimation implicitly considering the object geometry. Our approach yields an efficient dense grasp quality map and an object geometry estimate in a single forward pass. Both of these outputs can be used in a wide range of robotic manipulation applications. With rigorous experimental validation, both in simulation and on a real setup, we show that our shell-based method can be used to generate precise grasps and the associated grasp quality with over 90% accuracy. Diverse grasps computed on shell reconstructions allow the robot to select and execute grasps in cluttered scenes with more than 93% success rate.
Nikhil Chavan Dafle, Sergiy Popovych, Daniel D. Lee, Volkan Isler
IROS1
2020 PnuGrip: An Active Two-Phase Gripper for Dexterous Manipulation
abstract
We present the design of an active two-phase finger for mechanically mediated dexterous manipulation. The finger enables re-orientation of a grasped object by using a pneumatic braking mechanism to transition between free-rotating and fixed (i.e., braked) phases. Our design allows controlled high-bandwidth (5 Hz) phase transitions independent of the grasping force for manipulation of a variety of objects. Moreover, its thin profile (1 cm) facilitates picking and placing in clutter. Finally, the design features a sensor for measuring fingertip rotation to support feedback control. We experimentally characterize the finger's load handling capacity in the brake phase and rotational resistance in the free phase. We also demonstrate several pick-and-place manipulations common to industrial and laboratory automation settings that are simplified by our design.
Ian H. Taylor, Nikhil Chavan Dafle, Godric Li, Neel Doshi, Alberto Rodriguez 0003
IROS2
2018 Stable Prehensile Pushing: In-Hand Manipulation with Alternating Sticking Contacts
abstract
This paper presents an approach to in-hand manipulation planning that exploits the mechanics of alternating sticking contact. Particularly, we consider the problem of manipulating a grasped object using external pushes for which the pusher sticks to the object. Given the physical properties of the object, frictional coefficients at contacts and a desired regrasp on the object, we propose a sampling-based planning framework that builds a pushing strategy concatenating different feasible stable pushes to achieve the desired regrasp. An efficient dynamics formulation allows us to plan in-hand manipulations 100-1000 times faster than our previous work which builds upon a complementarity formulation. Experimental observations for the generated plans show that the object precisely moves in the grasp as expected by the planner.
Nikhil Chavan Dafle, Alberto Rodriguez 0003
ICRA1
2018 Robotic Pick-and-Place of Novel Objects in Clutter with Multi-Affordance Grasping and Cross-Domain Image Matching
abstract
This paper presents a robotic pick-and-place system that is capable of grasping and recognizing both known and novel objects in cluttered environments. The key new feature of the system is that it handles a wide range of object categories without needing any task-specific training data for novel objects. To achieve this, it first uses a category-agnostic affordance prediction algorithm to select and execute among four different grasping primitive behaviors. It then recognizes picked objects with a cross-domain image classification framework that matches observed images to product images. Since product images are readily available for a wide range of objects (e.g., from the web), the system works out-of-the-box for novel objects without requiring any additional training data. Exhaustive experimental results demonstrate that our multi-affordance grasping achieves high success rates for a wide variety of objects in clutter, and our recognition algorithm achieves high accuracy for both known and novel grasped objects. The approach was part of the MIT-Princeton Team system that took 1st place in the stowing task at the 2017 Amazon Robotics Challenge. All code, datasets, and pre-trained models are available online at http://arc.cs.princeton.edu.
Andy Zeng 0001, Shuran Song, Kuan-Ting Yu, Elliott Donlon, Francois Robert Hogan, Maria Bauzá 0001, Daolin Ma, Orion Taylor, Melody Liu, Eudald Romo Grau, Nima Fazeli, Ferran Alet, Nikhil Chavan Dafle, Rachel M. Holladay, Isabella Morona, Prem Qu Nair, Druck Green, Ian H. Taylor, Weber Liu, Thomas A. Funkhouser, Alberto Rodriguez 0003
ICRA13
2017 Sampling-Based Planning of In-Hand Manipulation with External Pushes
Nikhil Chavan Dafle, Alberto Rodriguez 0003
ISRR1
2015 Prehensile pushing: In-hand manipulation with push-primitives
abstract
This paper explores the manipulation of a grasped object by pushing it against its environment. Relying on precise arm motions and detailed models of frictional contact, prehensile pushing enables dexterous manipulation with simple manipulators, such as those currently available in industrial settings, and those likely affordable by service and field robots.
Nikhil Chavan Dafle, Alberto Rodriguez 0003
IROS1
2014 Extrinsic dexterity: In-hand manipulation with external forces
abstract
“In-hand manipulation” is the ability to reposition an object in the hand, for example when adjusting the grasp of a hammer before hammering a nail. The common approach to in-hand manipulation with robotic hands, known as dexterous manipulation [1], is to hold an object within the fingertips of the hand and wiggle the fingers, or walk them along the object's surface. Dexterous manipulation, however, is just one of the many techniques available to the robot. The robot can also roll the object in the hand by using gravity, or adjust the object's pose by pressing it against a surface, or if fast enough, it can even toss the object in the air and catch it in a different pose. All these techniques have one thing in common: they rely on resources extrinsic to the hand, either gravity, external contacts or dynamic arm motions. We refer to them as “extrinsic dexterity”. In this paper we study extrinsic dexterity in the context of regrasp operations, for example when switching from a power to a precision grasp, and we demonstrate that even simple grippers are capable of ample in-hand manipulation. We develop twelve regrasp actions, all open-loop and hand-scripted, and evaluate their effectiveness with over 1200 trials of regrasps and sequences of regrasps, for three different objects (see video [2]). The long-term goal of this work is to develop a general repertoire of these behaviors, and to understand how such a repertoire might eventually constitute a general-purpose in-hand manipulation capability.
Nikhil Chavan Dafle, Alberto Rodriguez 0003, Robert Paolini, Bowei Tang, Siddhartha S. Srinivasa, Michael A. Erdmann, Matthew T. Mason, Ivan Lundberg, Harald Staab, Thomas A. Fuhlbrigge
ICRA1
2014 Regrasping objects using extrinsic dexterity
abstract
This video presents the application of Extrinsic Dexterity to change the pose of an object in the hand, i.e., to regrasp the object.
Nikhil Chavan Dafle, Alberto Rodriguez 0003, Robert Paolini, Bowei Tang, Siddhartha S. Srinivasa, Michael A. Erdmann, Matthew T. Mason, Ivan Lundberg, Harald Staab, Thomas A. Fuhlbrigge
ICRA1