Snehasis Banerjee

dblp:53/11206 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6497-2085ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Planning, search and constraint satisfaction · 48% Knowledge representation and reasoning · 26% Robot navigation and mapping · 20%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
1.622025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base refinement
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical problem solving
task decomposition
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
classical planning
0.812024
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024
Robotics › Robot navigation and mapping › object goal navigation
multi-object navigation
0.712023
Sequence-Agnostic Multi-Object Navigation · ICRA 2023
Robotics › Robot navigation and mapping
object goal navigation
0.712023
Sequence-Agnostic Multi-Object Navigation · ICRA 2023
Natural language and speech › Language models and text generation
prompting
0.212024
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.212023
Sequence-Agnostic Multi-Object Navigation · ICRA 2023

Methods — techniques the papers use, named apart from their topics

large language model · 2.5knowledge graph · 1.7classical planning · 0.8reward shaping · 0.7actor-critic · 0.7
YearPublicationVenuePosition
2025 AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement
abstract
An embodied agent assisting humans is often asked to complete new tasks, and there may not be sufficient time or labeled examples to train the agent to perform these new tasks. Large Language Models (LLMs) trained on considerable knowledge across many domains can be used to predict a sequence of abstract actions for completing such tasks, although the agent may not be able to execute this sequence due to task-, agent-, or domain-specific constraints. Our framework addresses these challenges by leveraging the generic predictions provided by LLM and the prior domain knowledge encoded in a Knowledge Graph (KG), enabling an agent to quickly adapt to new tasks. The robot also solicits and uses human input as needed to refine its existing knowledge. Based on experimental evaluation in the context of cooking and cleaning tasks in simulation domains, we demonstrate that the interplay between LLM, KG, and human input leads to substantial performance gains compared with just using the LLM. Project website1§Project supported in part by TCS Research India: https://sssshivvvv.github.io/adaptbot/
Shivam Singh, Karthik Swaminathan, Nabanita Dash, Snehasis Banerjee, Mohan Sridharan, K. Madhava Krishna
ICRA5
2024 Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments†
abstract
Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-of-the-art methods for task anticipation use data-driven deep networks and Large Language Models (LLMs), but they do so at the level of high-level tasks and/or require many training examples. Our framework leverages the generic knowledge of LLMs through a small number of prompts to perform high-level task anticipation, using the anticipated tasks as goals in a classical planning system to compute a sequence of finer-granularity actions that jointly achieve these goals. We ground and evaluate our framework’s abilities in realistic scenarios in the VirtualHome environment and demonstrate a 31% reduction in execution time compared with a system that does not consider upcoming tasks.
Raghav Arora, Shivam Singh, Karthik Swaminathan, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
ICRA5
2024 Indoor Surveillance Robot with Person Following and Re-identification
abstract
Can a surveillance robot autonomously raise alerts for anomalous activity from its ego view camera perception in indoor spaces? Can it also follow an un-authorized person to investigate any intrusion? However, there exists a lack of datasets, trained models and methodology to handle surveillance use cases in indoor scenarios. In this work, we have created a hand annotated dataset involving indoor objects, specific to office spaces to understand context of perceived scenes. We demonstrate an end-to-end pipeline to find real time security hazards based on camera perception of a mobile robot to raise appropriate alerts to stakeholders. Additionally, following a person by a mobile robot is an essential feature in the surveillance domain, to check the whereabouts in case the person is a guest or an unauthorized person. While existing work has focused on learning the person model from frontal view, our work has focused on building an online model of a person to follow from any combination of back, side and frontal views. We have specially focused on person re-identification and trajectory, in case the person goes out of view or gets occluded – which is a challenging problem. To address this, we have presented a system and method that learns person’s distinct features on the fly and builds a strategy to navigate while keeping a safe and optimal distance from the person being followed. We also discuss deployment scenarios of the system in a real robot in an office environment.
Snehasis Banerjee, Abhijit Kumar, Apoorv Shekhar
IJCNN1
2023 Sequence-Agnostic Multi-Object Navigation
abstract
The Multi-Object Navigation (MultiON) task requires a robot to localize an instance (each) of multiple object classes. It is a fundamental task for an assistive robot in a home or a factory. Existing methods for MultiON have viewed this as a direct extension of Object Navigation (ON), the task of localising an instance of one object class, and are pre-sequenced, i.e., the sequence in which the object classes are to be explored is provided in advance. This is a strong limitation in practical applications characterized by dynamic changes. This paper describes a deep reinforcement learning framework for sequence-agnostic MultiON based on an actor-critic architecture and a suitable reward specification. Our framework leverages past experiences and seeks to reward progress toward individual as well as multiple target object classes. We use photo-realistic scenes from the Gibson benchmark dataset in the AI Habitat 3D simulation environment to experimentally show that our method performs better than a pre-sequenced approach and a state of the art ON method extended to MultiON.
Nandiraju Gireesh, Ahana Datta, Snehasis Banerjee, Mohan Sridharan, Brojeshwar Bhowmick, K. Madhava Krishna
ICRA4
2023 Anomalous Activity Detection from Ego View Camera of Surveillance Robots
abstract
Can a surveillance robot autonomously detect anomalous activity from its ego view camera perception? This is a challenging task as it requires identifying what is normal and what is an abnormal pattern - given the variations of possible anomalies and abnormalities. This paper presents an architecture and method based on a spatio-temporal convolution neural network to detect and classify anomalies. This work is inspired by the ‘Konio-Magno-Parvocellular’ cells of the human brain, which is claimed to aid humans in organizing changes in perceived scenes. The model is trained and tested on a benchmark video dataset [1] of human activity. We have obtained 91% testing accuracy on this dataset. Experiments in simulation as well as deployment on a real robot shows that the proposed methodology can identify anomalous activities effectively. We have also listed down the observations from practical deployment of the model.
Mritunjoy Halder, Snehasis Banerjee, P. Balamuralidhar
IJCNN2
2023 CLIPGraphs: Multimodal Graph Networks to Infer Object-Room Affinities
abstract
This paper introduces a novel method for determining the best room to place an object in, for embodied scene rearrangement. While state-of-the-art approaches rely on large language models (LLMs) or reinforcement learned (RL) policies for this task, our approach, CLIPGraphs, efficiently combines commonsense domain knowledge, data-driven methods, and recent advances in multimodal learning. Specifically, it (a) encodes a knowledge graph of prior human preferences about the room location of different objects in home environments, (b) incorporates vision-language features to support multimodal queries based on images or text, and (c) uses a graph network to learn object-room affinities based on embeddings of the prior knowledge and the vision-language features. We demonstrate that our approach provides better estimates of the most appropriate location of objects from a benchmark set of object categories in comparison with state-of-the-art baselines.11Supplementary material and code: https://clipgraphs.github.io
Raghav Arora, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
RO-MAN4