EDBT 2026 Demo / reviewers in the wild / expert
Nilesh Kulkarni
dblp:194/7595
· DBLP profile ↗
12ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion ModelabstractThe computer vision community has developed numerous techniques for digitally restoring true scene information from single-view degraded photographs, an important yet extremely ill-posed task. In this work, we tackle image restoration from a different perspective by jointly denoising multiple photographs of the same scene. Our core hypothesis is that degraded images capturing a shared scene contain complementary information that, when combined, better constrains the restoration problem. To this end, we implement a powerful multi-view diffusion model that jointly generates uncorrupted views by extracting rich information from multi-view relationships. Our experiments show that our multi-view approach outperforms existing single-view image and even video-based methods on image deblurring and super-resolution tasks. Critically, our model is trained to output 3D consistent images, making it a promising tool for applications requiring robust multi-view integration, such as 3D reconstruction or pose estimation. Project website: https://myc634.github.io/sirdiff/ Yucheng Mao, Nilesh Kulkarni, Jeong Joon Park |
CVPR | 3 |
| 2024 | FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationabstractEstimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose directly using neural networks are more robust to limited overlap and can infer absolute translation scale, but at the expense of reduced precision. We show how to combine the best of both methods; our approach yields results that are both precise and robust, while also accurately inferring translation scales. At the heart of our model lies a Transformer that (1) learns to balance between solved and learned pose estimations, and (2) provides a prior to guide a solver. A comprehensive analy-sis supports our design choices and demonstrates that our method adapts flexibly to various feature extractors and correspondence estimators, showing state-of-the-art performance in 6DoF pose estimation on Matterport3D, Inte-rio rNet, StreetLearn, and Map-free Relocalization. Project page: https://crockwell.github.io/farl Chris Rockwell 0001, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park, Justin Johnson 0001, David F. Fouhey |
CVPR | 2 |
| 2024 | 3DFIRES: Few Image 3D REconstruction for Scenes with Hidden SurfacesabstractThis paper introduces 3DFIRES, a novel system for scene-level 3D reconstruction from posed images. Designed to work with as few as one view, 3DFIRES reconstructs the complete geometry of unseen scenes, including hidden surfaces. With multiple view inputs, our method pro-duces full reconstruction within all camera frustums. A key feature of our approach is the fusion of multi-view information at the feature level, enabling the production of coherent and comprehensive 3D reconstruction. We train our system on non-watertight scans from large-scale real scene dataset. We show it matches the efficacy of single-view reconstruction methods with only one input and surpasses existing techniques in both quantitative and qualitative measures for sparse-view 3D reconstruction. Project page: https://jinlinyi.github.io/3DFIRES/ Linyi Jin, Nilesh Kulkarni, David F. Fouhey |
CVPR | 2 |
| 2024 | NIFTY: Neural Object Interaction Fields for Guided Human Motion SynthesisabstractWe address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object, which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field guides the sampling of an object-conditioned human motion diffusion model, so as to encourage plausible contacts and affordance semantics. To support interactions with scarcely available data, we propose an automated synthetic data pipeline. For this, we seed a pre-trained motion model, which has priors for the basics of human movement, with interaction-specific anchor poses extracted from limited motion capture data. Using our guided diffusion model trained on generated synthetic data, we synthesize realistic motions for sitting and lifting with several objects, outperforming alternative approaches in terms of motion quality and successful action completion. We call our framework NIFTY: Neural Interaction Fields for Trajectory sYnthesis. NIFTY results are available on https://nileshkulkarni.github.io/nifty. Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu, Justin Johnson 0001, David F. Fouhey, Leonidas J. Guibas |
CVPR | 1 |
| 2023 | Learning to Predict Scene-Level Implicit 3D from Posed RGBD DataabstractWe introduce a method that can learn to predict scenelevel implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often been tied to meshes, we show that we can train one using only a set of posed RGBD images. This setting may help 3D reconstruction unlock the sea of accelerometer+RGBD data that is coming with new phones. Our system, D2-DRDF, can match and sometimes outperform current methods that use mesh supervision and shows better robustness to sparse data. Nilesh Kulkarni, Linyi Jin, Justin Johnson 0001, David F. Fouhey |
CVPR | 1 |
| 2022 | Directed Ray Distance Functions for 3D Scene Reconstruction
Nilesh Kulkarni, Justin Johnson 0001, David F. Fouhey |
ECCV (2) | 1 |
| 2021 | Collision Replay: What Does Bumping Into Things Tell You About Scene Geometry?
Alexander Raistrick, Nilesh Kulkarni, David F. Fouhey |
BMVC | 2 |
| 2020 | Articulation-Aware Canonical Surface MappingabstractWe tackle the tasks of: 1) predicting a Canonical Surface Mapping (CSM) that indicates the mapping from 2D pixels to corresponding points on a canonical template shape , and 2) inferring the articulation and pose of the template corresponding to the input image. While previous approaches rely on keypoint supervision for learning, we present an approach that can learn without such annotations. Our key insight is that these tasks are geometrically related, and we can obtain supervisory signal via enforcing consistency among the predictions. We present results across a diverse set of animal object categories, showing that our method can learn articulation and CSM prediction from image collections using only foreground mask labels for training. We empirically show that allowing articulation helps learn more accurate CSM prediction, and that enforcing the consistency with predicted CSM is similarly critical for learning meaningful articulation. Nilesh Kulkarni, Abhinav Gupta 0001, David F. Fouhey, Shubham Tulsiani |
CVPR | 1 |
| 2019 | 3D-RelNet: Joint Object and Relational Network for 3D PredictionabstractWe propose an approach to predict the 3D shape and pose for the objects present in a scene. Existing learning based methods that pursue this goal make independent predictions per object, and do not leverage the relationships amongst them. We argue that reasoning about these relationships is crucial, and present an approach to incorporate these in a 3D prediction framework. In addition to independent per-object predictions, we predict pairwise relations in the form of relative 3D pose, and demonstrate that these can be easily incorporated to improve object level estimates. We report performance across different datasets (SUNCG, NYUv2), and show that our approach significantly improves over independent prediction approaches while also outperforming alternate implicit reasoning methods. Nilesh Kulkarni, Ishan Misra, Shubham Tulsiani, Abhinav Gupta 0001 |
ICCV | 1 |
| 2019 | Canonical Surface Mapping via Geometric Cycle ConsistencyabstractWe explore the task of Canonical Surface Mapping (CSM). Specifically, given an image, we learn to map pixels on the object to their corresponding locations on an abstract 3D model of the category. But how do we learn such a mapping? A supervised approach would require extensive manual labeling which is not scalable beyond a few hand-picked categories. Our key insight is that the CSM task (pixel to 3D), when combined with 3D projection (3D to pixel), completes a cycle. Hence, we can exploit a geometric cycle consistency loss, thereby allowing us to forgo the dense manual supervision. Our approach allows us to train a CSM model for a diverse set of classes, without sparse or dense keypoint annotation, by leveraging only foreground mask labels for training. We show that our predictions also allow us to infer dense correspondence between two images, and compare the performance of our approach against several methods that predict correspondence by leveraging varying amount of supervision. Nilesh Kulkarni, Shubham Tulsiani, Abhinav Gupta 0001 |
ICCV | 1 |
| 2016 | Content-based recommendation for podcast audio-items using natural language processing techniquesabstractA podcast combines the liveliness of a FM radio channel with the economy of Internet blog posting. They are especially convenient for scenarios when there is limited internet ability and connectivity for example in the car, the gym, etc. While both the volume and heterogeneity of content is huge it becomes operationally difficult to manually categorize or tag these audio items, thus manage them in a system for users to discover. Furthermore, due to the incompleteness of audio associated meta data there are not enough features for a typical recommender system to learn the item similarities thus make recommendations. In this paper we propose and examine a novel approach to generate latent embeddings for podcast items utilizing the aggregated information from all the text-based features associated with the audioitems. These embeddings that are generated using well established Natural Language Processing (NLP) techniques for the podcast items can be used to measure or indicate the content similarity among the various podcast items. Both GPU (CUDA) and CPU computing architectures are experimented and bench marked for the model training, cross-validation of the content predictions on large scale datasets. Zhou Xing, Marzieh Parandehgheibi, Nilesh Kulkarni, Chris Pouliot |
IEEE BigData | 4 |
| 2016 | Robust kernel principal nested spheresabstractKernel principal component analysis (kPCA) learns nonlinear modes of variation in the data by nonlinearly mapping the data to kernel feature space and performing (linear) PCA in the associated reproducing kernel Hilbert space (RKHS). However, several widely-used Mercer kernels map data to a Hilbert sphere in RKHS. For such directional data in RKHS, linear analyses can be unnatural or suboptimal. Hence, we propose an alternative to kPCA by extending principal nested spheres (PNS) to RKHS without needing the explicit lifting map underlying the kernel, but solely relying on the kernel trick. It generalizes the model for the residual errors by penalizing the Lpnorm / quasi-norm to enable robust learning from corrupted training data. Our method, termed robust kernel PNS (rkPNS), relies on the Riemannian geometry of the Hilbert sphere in RKHS. Relying on rkPNS, we propose novel algorithms for dimensionality reduction and classification (with and without outliers in the training data). Evaluation on real-world datasets shows that rkPNS compares favorably to the state of the art. Suyash P. Awate, Manik Dhar, Nilesh Kulkarni |
ICPR | 3 |