EDBT 2026 Demo / reviewers in the wild / expert
Nikhil Varma Keetha
dblp:261/3637
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-2770-0835ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 74% Robot navigation and mapping · 19% Robot manipulation · 5% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
2.5 | 3 | 2025 | UFM: A Simple Path towards Unified Dense Correspondence with Flow · NeurIPS 2025 FlowR: Flowing from Sparse to Dense 3D Reconstructions · ICCV 2025 SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Computer vision › 3D vision › correspondence estimation
dense correspondence |
0.9 | 1 | 2025 | UFM: A Simple Path towards Unified Dense Correspondence with Flow · NeurIPS 2025 |
Computer vision › 3D vision › motion estimation
optical flow |
0.9 | 1 | 2025 | UFM: A Simple Path towards Unified Dense Correspondence with Flow · NeurIPS 2025 |
Computer vision › 3D vision › feature matching › local feature matching
wide-baseline matching |
0.9 | 1 | 2025 | UFM: A Simple Path towards Unified Dense Correspondence with Flow · NeurIPS 2025 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.8 | 1 | 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Robotics › Robot navigation and mapping › SLAM › dense SLAM
dense RGB-D SLAM |
0.8 | 1 | 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Robotics › Robot navigation and mapping
map prediction |
0.8 | 1 | 2024 | Map It Anywhere: Empowering BEV Map Prediction using Large-scale Public Datasets · NeurIPS 2024 |
Computer vision › 3D vision
novel view synthesis |
0.8 | 1 | 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Computer vision › 3D vision › 3d scene modeling
scene representation |
0.8 | 1 | 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Robotics › Robot navigation and mapping
SLAM |
0.8 | 1 | 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM · CVPR 2024 |
Robotics › Robot manipulation › object perception
object identification |
0.6 | 1 | 2022 | AirObject: A Temporally Evolving Graph Embedding for Object Identification · CVPR 2022 |
Computer vision › 3D vision
object representation |
0.6 | 1 | 2022 | AirObject: A Temporally Evolving Graph Embedding for Object Identification · CVPR 2022 |
Robotics › Autonomous driving
perception |
0.2 | 1 | 2024 | Map It Anywhere: Empowering BEV Map Prediction using Large-scale Public Datasets · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9flow regression · 0.9pre-training · 0.8online tracking and mapping · 0.8camera model-agnostic learning · 0.83d gaussian representation · 0.8temporal convolutional network · 0.6graph attention network · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MapAnything: Universal Feed-Forward Metric 3D Reconstruction; map-anything.github.ioabstractWe introduce MapAnything, a unified transformer-based feed-forward model that ingests one or more images along with optional geometric inputs such as camera intrinsics, poses, depth, or partial reconstructions, and directly regresses the metric 3D scene geometry and cameras. MapAnything lever-ages a factored representation of multi-view scene geome-try, i.e., a collection of depth maps, local raymaps, camera poses, and a metric scale factor that effectively upgrades local reconstructions into a globally consistent metric frame. Standardizing the supervision and training across diverse datasets, along with flexible input augmentation, enables MapAnything to address a broad range of 3D vision tasks in a single feed-forward pass, including uncalibrated structure-from-motion, calibrated multi-view stereo, monocular depth estimation, camera localization, depth completion, and more. We provide extensive experimental analyses and model ab-lations demonstrating that MapAnything We provide extensive experimental analyses and model ab-lations demonstrating that MapAnything outperforms or matches specialist feed-forward models while offering more efficient joint training behavior, thus paving the way toward a universal 3D reconstruction backbone. Nikhil Varma Keetha, Norman Müller, Johannes Schönberger, Lorenzo Porzi, Tobias Fischer 0004, Arno Knapitsch, Duncan Zauss, Ethan Weber, Nelson Antunes, Jonathon Luiten, Manuel Lopez-Antequera, Samuel Rota Bulò, Christian Richardt, Deva Ramanan, Sebastian A. Scherer, Peter Kontschieder |
3DV | 1 |
| 2025 | FlowR: Flowing from Sparse to Dense 3D Reconstructions
Tobias Fischer 0004, Samuel Rota Bulò, Yung-Hsu Yang, Nikhil Varma Keetha, Lorenzo Porzi, Norman Müller, Katja Schwarz, Jonathon Luiten, Marc Pollefeys, Peter Kontschieder |
ICCV | 4 |
| 2025 | RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and ExplorationabstractOpen-set semantic mapping is crucial for openworld robots. Current mapping approaches either are limited by the depth range or only map beyond-range entities in constrained settings, where overall they fail to combine within-range and beyond-range observations. Furthermore, these methods make a trade-off between fine-grained semantics and efficiency. We introduce RayFronts, a unified representation that enables both dense and beyond-range efficient semantic mapping. RayFronts encodes task-agnostic openset semantics to both in-range voxels and beyond-range rays encoded at map boundaries, empowering the robot to reduce search volumes significantly and make informed decisions both within & beyond sensory range, while running at 8.84 Hz on an Orin AGX. Benchmarking the within-range semantics shows that RayFronts’s fine-grained image encoding provides 1.34× zero-shot 3D semantic segmentation performance while improving throughput by 16.5×. Traditionally, online mapping performance is entangled with other system components, complicating evaluation. We propose a planner-agnostic evaluation framework that captures the utility for online beyond-range search and exploration, and show RayFronts reduces search volume 2.2× more efficiently than the closest online baselines. Omar Alama, Avigyan Bhattacharya, Haoyang He, Seungchan Kim, Yuheng Qiu, Cherie Ho, Nikhil Varma Keetha, Sebastian A. Scherer |
IROS | 8 |
| 2025 | UFM: A Simple Path towards Unified Dense Correspondence with FlowabstractDense image correspondence is central to many applications, such as visual odometry, 3D reconstruction, object association, and re-identification. Historically, dense correspondence has been tackled separately for wide-baseline scenarios and optical flow estimation, despite the common goal of matching content between two images. In this paper, we develop a Unified Flow \& Matching model (UFM), which is trained on unified data for pixels that are co-visible in both source and target images. UFM uses a simple, generic transformer architecture that directly regresses the $(u,v)$ flow. It is easier to train and more accurate for large flows compared to the typical coarse-to-fine cost volumes in prior work. UFM is 28\% more accurate than state-of-the-art flow methods (Unimatch), while also having 62\% less error and 6.7x faster than dense wide-baseline matchers (RoMa). UFM is the first to demonstrate that unified training can outperform specialized approaches across both domains. This result enables fast, general-purpose correspondence and opens new directions for multi-modal, long-range, and real-time correspondence tasks. Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb, Yuheng Qiu, Jay Karhade, Shreyas Jha, Yaoyu Hu, Deva Ramanan, Sebastian A. Scherer |
NeurIPS | 2 |
| 2024 | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAMabstractDense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications. However, current methods are often hampered by the non-volumetric or implicit way they represent a scene. This work introduces SplaTAM, an approach that, for the first time, leverages explicit volumetric representations, i.e., 3D Gaussians, to enable high-fidelity reconstruction from a single unposed RGB-D camera, surpassing the capabilities of existing methods. SplaTAM employs a simple online tracking and mapping system tailored to the underlying Gaussian representation. It utilizes a silhouette mask to elegantly capture the presence of scene density. This combination enables several benefits over prior representations, including fast rendering and dense optimization, quickly determining if areas have been previously mapped, and structured map expansion by adding more Gaussians. Extensive experiments show that SplaTAM achieves up to 2 x superior performance in camera pose estimation, map construction, and novel-view synthesis over existing methods, paving the way for more immersive high-fidelity SLAM applications. Nikhil Varma Keetha, Jay Karhade, Krishna Murthy Jatavallabhula, Gengshan Yang, Sebastian A. Scherer, Deva Ramanan, Jonathon Luiten |
CVPR | 1 |
| 2024 | Targeted Image Transformation for Improving Robustness in Long Range Aircraft DetectionabstractIn the field of aviation, the Detect and Avoid (DAA) problem deals with incorporating collision avoidance capabilities into current autopilot navigation systems. As an application of the Small Object Detection (SOD) problem, DAA presents the difficulties of a low signal-to-noise ratio and far range detection. Visual DAA is also susceptible to changing weather and lighting conditions at deployment. While current literature has presented many solutions for this, prior work has yet to study the robustness of the learning-based models for DAA. In this work, we show that standard techniques for improving robustness for object detection do not produce the desired results for DAA given the SOD constraints. We present targeted transformations, a zero-shot technique that can significantly improve robustness with minimal impact on accuracy. We demonstrate how to construct these transformations and evaluate our method on the current SOTA model for DAA, showing a 53.6% increase in recall. This makes our pipeline more robust to changes in lighting and environmental factors, and better able to detect potential threats. In the future, we hope to automate the transformation selection process, making it easier to adopt in different use cases. Rebecca Martin, Clement Fung, Nikhil Varma Keetha, Lujo Bauer, Sebastian A. Scherer |
IROS | 3 |
| 2024 | Map It Anywhere: Empowering BEV Map Prediction using Large-scale Public DatasetsabstractTop-down Bird's Eye View (BEV) maps are a popular perception representation for ground robot navigation due to their richness and flexibility for downstream tasks. While recent methods have shown promise for predicting BEV maps from First-Person View (FPV) images, their generalizability is limited to small regions captured by current autonomous vehicle-based datasets. In this context, we show that a more scalable approach towards generalizable map prediction can be enabled by using two large-scale crowd-sourced mapping platforms, Mapillary for FPV images and OpenStreetMap for BEV semantic maps.We introduce Map It Anywhere (MIA), a data engine that enables seamless curation and modeling of labeled map prediction data from existing open-source map platforms. Using our MIA data engine, we display the ease of automatically collecting a 1.2 million FPV & BEV pair dataset encompassing diverse geographies, landscapes, environmental factors, camera models & capture scenarios. We further train a simple camera model-agnostic model on this data for BEV map prediction.Extensive evaluations using established benchmarks and our dataset show that the data curated by MIA enables effective pretraining for generalizable BEV map prediction, with zero-shot performance far exceeding baselines trained on existing datasets by 35%. Our analysis highlights the promise of using large-scale public maps for developing & testing generalizable BEV perception, paving the way for more robust autonomous navigation.Website: mapitanywhere.github.io Cherie Ho, Jiaye Zou, Omar Alama, Sai Mitheran Jagadesh Kumar, Cheng-Yu Chiang, Taneesh Gupta, Chen Wang 0033, Nikhil Varma Keetha, Katia P. Sycara, Sebastian A. Scherer |
NeurIPS | 8 |
| 2022 | AirObject: A Temporally Evolving Graph Embedding for Object IdentificationabstractObject encoding and identification are vital for robotic tasks such as autonomous exploration, semantic scene understanding, and relocalization. Previous approaches have attempted to either track objects or generate descriptors for object identification. However, such systems are limited to a “fixed” partial object representation from a single viewpoint. In a robot exploration setup, there is a requirement for a temporally “evolving” global object representation built as the robot observes the object from multiple viewpoints. Furthermore, given the vast distribution of unknown novel objects in the real world, the object identification process must be class-agnostic. In this context, we propose a novel temporal 3D object encoding approach, dubbed AirObject, to obtain global keypoint graph-based embeddings of objects. Specifically, the global 3D object embeddings are generated using a temporal convolutional network across structural information of multiple frames obtained from a graph attention-based encoding method. We demonstrate that AirObject achieves the state-of-the-art performance for video object identification and is robust to severe occlusion, perceptual aliasing, viewpoint shift, deformation, and scale transform, outperforming the state-of-the-art single-frame and sequential descriptors. To the best of our knowledge, AirObject is one of the first temporal object encoding methods. Source code is available at https://github.com/Nik-v9/AirObject. Nikhil Varma Keetha, Chen Wang 0033, Yuheng Qiu, Sebastian A. Scherer |
CVPR | 1 |