Shivam Singh

dblp:214/0769 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Planning, search and constraint satisfaction · 45% Generative modeling · 24% Knowledge representation and reasoning · 24%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
1.622025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024
Machine learning › Generative modeling
diffusion model
0.912025
RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions · ICCV 2025
Machine learning › Generative modeling › diffusion model › image editing
instruction-based image editing
0.912025
RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions · ICCV 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base refinement
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › hierarchical problem solving
task decomposition
0.912025
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement · ICRA 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
classical planning
0.812024
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024
Computer vision › Vision and language › visual grounding
referring expression
0.312025
RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions · ICCV 2025
Natural language and speech › Language models and text generation
prompting
0.212024
Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments† · ICRA 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.5knowledge graph · 1.7synthetic data generation · 0.9inversion · 0.9classical planning · 0.8
YearPublicationVenuePosition
2025 RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions
abstract
Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.
Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang, Chitta Baral
ICCV3
2025 AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement
abstract
An embodied agent assisting humans is often asked to complete new tasks, and there may not be sufficient time or labeled examples to train the agent to perform these new tasks. Large Language Models (LLMs) trained on considerable knowledge across many domains can be used to predict a sequence of abstract actions for completing such tasks, although the agent may not be able to execute this sequence due to task-, agent-, or domain-specific constraints. Our framework addresses these challenges by leveraging the generic predictions provided by LLM and the prior domain knowledge encoded in a Knowledge Graph (KG), enabling an agent to quickly adapt to new tasks. The robot also solicits and uses human input as needed to refine its existing knowledge. Based on experimental evaluation in the context of cooking and cleaning tasks in simulation domains, we demonstrate that the interplay between LLM, KG, and human input leads to substantial performance gains compared with just using the LLM. Project website1§Project supported in part by TCS Research India: https://sssshivvvv.github.io/adaptbot/
Shivam Singh, Karthik Swaminathan, Nabanita Dash, Snehasis Banerjee, Mohan Sridharan, K. Madhava Krishna
ICRA1
2024 Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments†
abstract
Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-of-the-art methods for task anticipation use data-driven deep networks and Large Language Models (LLMs), but they do so at the level of high-level tasks and/or require many training examples. Our framework leverages the generic knowledge of LLMs through a small number of prompts to perform high-level task anticipation, using the anticipated tasks as goals in a classical planning system to compute a sequence of finer-granularity actions that jointly achieve these goals. We ground and evaluate our framework’s abilities in realistic scenarios in the VirtualHome environment and demonstrate a 31% reduction in execution time compared with a system that does not consider upcoming tasks.
Raghav Arora, Shivam Singh, Karthik Swaminathan, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
ICRA2