Nikhil Kakodkar

dblp:192/5952 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot navigation and mapping · 26% Vision and language · 21% Language models and text generation · 16%
Computer graphics and multimedia
1 paper
Computational photography and imaging · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › 3d vision and language
3d scene grounding
0.812024
CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots · ICRA 2024
Natural language and speech › Language models and text generation
instruction following
0.812024
CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots · ICRA 2024
Robotics › Robot navigation and mapping › embodied navigation
natural language navigation instructions
0.812024
CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots · ICRA 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
spatial language understanding
0.812024
CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots · ICRA 2024
Robotics › Motion planning and robot control
robot planning
0.712023
ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence · ICRA 2023
Knowledge, reasoning and agents › Multi-agent systems
pursuit-evasion
0.312018
Model-Based Probabilistic Pursuit via Inverse Reinforcement Learning · ICRA 2018
Computer vision › Vision and language
vision-language model
0.212023
ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence · ICRA 2023
Robotics › Robot navigation and mapping
mobile robot navigation
0.112018
Model-Based Probabilistic Pursuit via Inverse Reinforcement Learning · ICRA 2018
Robotics › Robot navigation and mapping
target search
0.112018
Model-Based Probabilistic Pursuit via Inverse Reinforcement Learning · ICRA 2018

Methods — techniques the papers use, named apart from their topics

large language model · 2.1vision-language model · 1.3inverse reinforcement learning · 0.3combinatorial search · 0.3POMDP · 0.3MDP · 0.3
YearPublicationVenuePosition
2024 CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots
abstract
This work explores the capacity of large language models (LLMs) to address problems at the intersection of spatial planning and natural language interfaces for navigation. We focus on following complex instructions that are more akin to natural conversation than traditional explicit procedural directives typically seen in robotics. Unlike most prior work where navigation directives are provided as simple imperative commands (e.g., "go to the fridge"), we examine implicit directives obtained through conversational interactions.We leverage the 3D simulator AI2Thor to create household query scenarios at scale, and augment it by adding complex language queries for 40 object types. We demonstrate that a robot using our method CARTIER (Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots) can parse descriptive language queries up to 42% more reliably than existing LLM-enabled methods by exploiting the ability of LLMs to interpret the user interaction in the context of the objects in the scenario.
Dmitriy Rivkin, Nikhil Kakodkar, Francois Robert Hogan, Bobak H. Baghi, Gregory Dudek
ICRA2
2023 ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence
abstract
Our work examines the way in which large language models can be used for robotic planning and sampling in the context of automated photographic documentation. Specifically, we illustrate how to produce a photo-taking robot with an exceptional level of semantic awareness by leveraging recent advances in general purpose language (LM) and vision-language (VLM) models. Given a high-level description of an event we use an LM to generate a natural-language list of photo descriptions that one would expect a photographer to capture at the event. We then use a VLM to identify the best matches to these descriptions in the robot's video stream. The photo portfolios generated by our method are consistently rated as more appropriate to the event by human evaluators than those generated by existing methods.
Dmitriy Rivkin, Gregory Dudek, Nikhil Kakodkar, David Meger, Oliver Limoyo, Michael R. M. Jenkin, Xue Liu 0004, Francois Robert Hogan
ICRA3
2018 Model-Based Probabilistic Pursuit via Inverse Reinforcement Learning
abstract
We address the integrated prediction, planning, and control problem that enables a single follower robot (the photographer) to quickly re-establish visual contact with a moving target (the subject) that has escaped the follower's field of view. We deal with this scenario, which reactive controllers are typically ill-equipped to handle, by making plausible predictions about the long- and short-term behavior of the target, and planning pursuit paths that will maximize the chance of seeing the target again. At the core of our pursuit method is the use of predictive models of target behavior, which help narrow down the set of possible future locations of the target to a few discrete hypotheses, as well as the use of combinatorial search in physical space to check those hypotheses efficiently. We model target behavior in terms of a learned navigation reward function, using Inverse Reinforcement Learning, based on semantic terrain features of satellite maps. Our pursuit algorithm continuously predicts the latent destination of the target and its position in the future, and relies on efficient graph representation and search methods in order to navigate to locations at which the target is most likely to be seen at an anticipated time. We perform extensive evaluation of our predictive pursuit algorithm over multiple satellite maps, thousands of simulation scenarios, against state-of-the art MDP and POMDP solvers. We show that our method significantly outperforms them by exploiting domain-specific knowledge, while being able to run in real-time.
Florian Shkurti, Nikhil Kakodkar, Gregory Dudek
ICRA2