Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Siddhant Agarwal

dblp:197/5870 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 55% Multi-agent systems · 13% Language models and text generation · 7%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 85% Medical and health informatics · 15%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 25 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
goal-conditioned reinforcement learning
1.522025
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning · ICLR 2025
f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-Divergences · NeurIPS 2023
Natural language and speech › Language models and text generation › large language model
large language model applications
1.012026
MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes · AAAI 2026
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
1.012026
MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes · AAAI 2026
Machine learning › Reinforcement learning
hindsight relabeling
0.912025
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning
0.912025
RLZero: Direct Policy Inference from Language Without In-Domain Supervision · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning
latent representation learning
0.912025
Hybrid Latent Representations for PDE Emulation · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › multi-robot systems
robot soccer
0.912025
Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer · ICRA 2025
Machine learning › Reinforcement learning
sample efficiency
0.912025
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
successor representation
0.912025
Proto Successor Measure: Representing the Behavior Space of an RL Agent · ICML 2025
Machine learning › Generative modeling
video generation
0.912025
RLZero: Direct Policy Inference from Language Without In-Domain Supervision · NeurIPS 2025
Machine learning › Reinforcement learning › generalization in reinforcement learning
zero-shot reinforcement learning
0.912025
Proto Successor Measure: Representing the Behavior Space of an RL Agent · ICML 2025
Computational science and engineering › scientific machine learning
neural PDE emulators
0.912025
Hybrid Latent Representations for PDE Emulation · NeurIPS 2025
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.912025
Hybrid Latent Representations for PDE Emulation · NeurIPS 2025
Multimedia analysis and retrieval
audio-visual learning
0.912025
Clink! Chop! Thud! - Learning Object Sounds From Real-World Interactions · ICCV 2025
Multimedia analysis and retrieval
multimodal action understanding
0.912025
Clink! Chop! Thud! - Learning Object Sounds From Real-World Interactions · ICCV 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation
0.712023
What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes · AAAI 2023
Machine learning › Reinforcement learning
exploration
0.712023
f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-Divergences · NeurIPS 2023
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.712023
f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-Divergences · NeurIPS 2023
Computer vision › Vision and language › visual grounding
visual semantic role labeling
0.712023
What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes · AAAI 2023
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.512021
Learning to Deceive Knowledge Graph Augmented Models via Targeted Perturbation · ICLR 2021
Medical and health informatics
mental health informatics
0.312026
MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes · AAAI 2026
Machine learning › Transfer learning and domain adaptation › cross-embodiment learning
cross-embodiment transfer
0.312025
RLZero: Direct Policy Inference from Language Without In-Domain Supervision · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.312025
Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer · ICRA 2025
Robotics › Robot manipulation › object manipulation
object-centric manipulation
0.312025
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning · ICLR 2025
Natural language and speech › Language models and text generation
text generation
0.212023
What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes · AAAI 2023

Methods — techniques the papers use, named apart from their topics

large language model · 2.0cognitive analytic therapy · 2.0autoencoder · 1.7visitation distributions · 0.9slot attention · 0.9sim2real · 0.9segmentation mask prediction · 0.9reward-free learning · 0.9null counterfactual interaction inference · 0.9model-free reinforcement learning · 0.9learned dynamics model · 0.9fourier neural operator · 0.9convolutional neural network · 0.9behavior decomposition · 0.9targeted perturbation · 0.5
YearPublicationVenuePosition
2026 MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes
abstract
Over the past years, memes have evolved from being exclusively a medium of humorous exchanges to one that allows users to express a range of emotions freely and easily. With the ever-growing utilization of memes in expressing depressive sentiments, we conduct a study on identifying depressive symptoms exhibited by memes shared by users of online social media platforms. We introduce RESTOREx as a vital resource for detecting depressive symptoms in memes on social media through the Large Language Model (LLM) generated and human-annotated explanations. We introduce MAMA-Memeia, a collaborative multi-agent multi-aspect discussion framework grounded in the clinical psychology method of Cognitive Analytic Therapy (CAT) Competencies. MAMA-Memeia improves upon the current state-of-the-art by 7.55% in macro-F1 and is established as the new benchmark compared to over 30 methods.
Siddhant Agarwal, Adya Dhuler, Polly Ruhnke, Melvin Speisman, Md. Shad Akhtar, Shweta Yadav 0001
AAAI1
2025 Clink! Chop! Thud! - Learning Object Sounds From Real-World Interactions
abstract
Can a model distinguish between the sound of a spoon hitting a hardwood floor versus a carpeted one? Everyday object interactions produce sounds unique to the objects involved. We introduce the sounding object detection task to evaluate a model's ability to link these sounds to the objects directly involved. Inspired by human perception, our multimodal object-aware framework learns from in-the-wild egocentric videos. To encourage an object-centric approach, we first develop an automatic pipeline to compute segmentation masks of the objects involved to guide the model's focus during training towards the most informative regions of the interaction. A slot attention visual encoder is used to further enforce an object prior. We demonstrate state of the art performance on our new task along with existing multimodal action understanding tasks.
Mengyu Yang, Haozheng Pei, Siddhant Agarwal, Arun Balajee Vasudevan, James Hays
ICCV4
2025 Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
abstract
Hindsight relabeling is a powerful tool for overcoming sparsity in goal-conditioned reinforcement learning (GCRL), especially in certain domains such as navigation and locomotion. However, hindsight relabeling can struggle in object-centric domains. For example, suppose that the goal space consists of a robotic arm pushing a particular target block to a goal location. In this case, hindsight relabeling will give high rewards to any trajectory that does not interact with the block. However, these behaviors are only useful when the object is already at the goal---an extremely rare case in practice. A dataset dominated by these kinds of trajectories can complicate learning and lead to failures. In object-centric domains, one key intuition is that meaningful trajectories are often characterized by object-object interactions such as pushing the block with the gripper. To leverage this intuition, we introduce Hindsight Relabeling using Interactions (HInt), which combines interactions with hindsight relabeling to improve the sample efficiency of downstream RL. However, interactions do not have a consensus statistical definition that is tractable for downstream GCRL. Therefore, we propose a definition of interactions based on the concept of _null counterfactual_: a cause object is interacting with a target object if, in a world where the cause object did not exist, the target object would have different transition dynamics. We leverage this definition to infer interactions in Null Counterfactual Interaction Inference (NCII), which uses a ``nulling'' operation with a learned model to simulate absences and infer interactions. We demonstrate that NCII is able to achieve significantly improved interaction inference accuracy in both simple linear dynamics domains and dynamic robotic domains in Robosuite, Robot Air Hockey, and Franka Kitchen. Furthermore, we demonstrate that HInt improves sample efficiency by up to $4\times$ in these domains as goal-conditioned tasks.
Caleb Chuck, Carl Qi, Chang Shi, Siddhant Agarwal, Amy Zhang 0001, Scott Niekum
ICLR5
2025 Proto Successor Measure: Representing the Behavior Space of an RL Agent
abstract
Having explored an environment, intelligent agents should be able to transfer their knowledge to most downstream tasks within that environment without additional interactions. Referred to as "zero-shot learning", this ability remains elusive for general-purpose reinforcement learning algorithms. While recent works have attempted to produce zero-shot RL agents, they make assumptions about the nature of the tasks or the structure of the MDP. We present Proto Successor Measure: the basis set for all possible behaviors of a Reinforcement Learning Agent in a dynamical system. We prove that any possible behavior (represented using visitation distributions) can be represented using an affine combination of these policy-independent basis functions. Given a reward function at test time, we simply need to find the right set of linear weights to combine these bases corresponding to the optimal policy. We derive a practical algorithm to learn these basis functions using reward-free interaction data from the environment and show that our approach can produce the near-optimal policy at test time for any given reward function without additional environmental interactions. Project page: agarwalsiddhant10.github.io/projects/psm.html.
Siddhant Agarwal, Harshit Sikchi, Peter Stone 0001, Amy Zhang 0001
ICML1
2025 Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer
abstract
Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) is a promising approach to learning decisionmaking in such domains, however, end-to-end RL in complex environments is often intractable. To address this challenge in the RoboCup Standard Platform League (SPL) domain, we developed a novel architecture integrating RL within a classical robotics stack, while employing a multi-fidelity sim2real approach and decomposing behavior into learned sub-behaviors with heuristic selection. Our architecture led to victory in the 2024 RoboCup SPL Challenge Shield Division. In this work, we fully describe our system's architecture and empirically analyze key design decisions that contributed to its success. Our approach demonstrates how RL-based behaviors can be integrated into complete robot behavior architectures.
Adam Labiosa, Zhihan Wang, Siddhant Agarwal, William Cong, Geethika Hemkumar, Abhinav Narayan Harish, Benjamin Hong, Josh Kelle, Zisen Shao, Peter Stone 0001, Josiah Hanna
ICRA3
2025 Hybrid Latent Representations for PDE Emulation
abstract
For classical PDE solvers, adjusting the spatial resolution and time step offers a trade-off between speed and accuracy. Neural emulators often achieve better speed-accuracy trade-offs by operating on a compact representation of the PDE system. Coarsened PDE fields are a simple and effective representation, but cannot exploit fine spatial scales in the high-fidelity numerical solutions. Alternatively, unstructured latent representations provide efficient autoregressive rollouts, but cannot enforce local interactions or physical laws as inductive biases. To overcome these limitations, we introduce hybrid representations that augment coarsened PDE fields with spatially structured latent variables extracted from high-resolution inputs. Hybrid representations provide efficient rollouts, can be trained on a simple loss defined on coarsened PDE fields, and support hard physical constraints. When predicting fine- and coarse-scale features across multiple PDE emulation tasks, they outperform or match the speed-accuracy trade-offs of the best convolutional, attentional, Fourier operator-based and autoencoding baselines.
Ali Can Bekar, Siddhant Agarwal, Christian Hüttig, Nicola Tosi, David S. Greenberg
NeurIPS2
2025 RLZero: Direct Policy Inference from Language Without In-Domain Supervision
abstract
The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward signal is notoriously difficult, as humans are often unable to predict the optimal behavior corresponding to a reward function. Natural language offers an intuitive alternative for instructing reinforcement learning (RL) agents, yet previous language-conditioned approaches either require costly supervision or test-time training given a language instruction. In this work, we present a new approach that uses a pretrained RL agent trained using only unlabeled, offline interactions—without task-specific supervision or labeled trajectories—to get zero-shot test-time policy inference from arbitrary natural language instructions. We introduce a framework comprising three steps: *imagine*, *project*, and *imitate*. First, the agent imagines a sequence of observations corresponding to the provided language description using video generative models. Next, these imagined observations are projected into the target environment domain. Finally, an agent pretrained in the target environment with unsupervised RL instantly imitates the projected observation sequence through a closed-form solution. To the best of our knowledge, our method, RLZero, is the first approach to show direct language-to-behavior generation abilities on a variety of tasks and environments without any in-domain supervision. We further show that components of RLZero can be used to generate policies zero-shot from cross-embodied videos, such as those available on YouTube, even for complex embodiments like humanoids.
Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli, Caleb Chuck, Max Rudolph, Peter Stone 0001, Amy Zhang 0001, Scott Niekum
NeurIPS2
2023 What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes
abstract
Memes are powerful means for effective communication on social media. Their effortless amalgamation of viral visuals and compelling messages can have far-reaching implications with proper marketing. Previous research on memes has primarily focused on characterizing their affective spectrum and detecting whether the meme's message insinuates any intended harm, such as hate, offense, racism, etc. However, memes often use abstraction, which can be elusive. Here, we introduce a novel task - EXCLAIM, generating explanations for visual semantic role labeling in memes. To this end, we curate ExHVV, a novel dataset that offers natural language explanations of connotative roles for three types of entities - heroes, villains, and victims, encompassing 4,680 entities present in 3K memes. We also benchmark ExHVV with several strong unimodal and multimodal baselines. Moreover, we posit LUMEN, a novel multimodal, multi-task learning framework that endeavors to address EXCLAIM optimally by jointly learning to predict the correct semantic roles and correspondingly to generate suitable natural language explanations. LUMEN distinctly outperforms the best baseline across 18 standard natural language generation evaluation metrics. Our systematic evaluation and analyses demonstrate that characteristic multimodal cues required for adjudicating semantic roles are also helpful for generating suitable explanations.
Siddhant Agarwal, Tharun Suresh, Preslav Nakov, Md. Shad Akhtar, Tanmoy Chakraborty 0002
AAAI2
2023 f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-Divergences
abstract
Goal-Conditioned Reinforcement Learning (RL) problems often have access to sparse rewards where the agent receives a reward signal only when it has achieved the goal, making policy optimization a difficult problem. Several works augment this sparse reward with a learned dense reward function, but this can lead to sub-optimal policies if the reward is misaligned. Moreover, recent works have demonstrated that effective shaping rewards for a particular problem can depend on the underlying learning algorithm. This paper introduces a novel way to encourage exploration called $f$-Policy Gradients, or $f$-PG. $f$-PG minimizes the f-divergence between the agent's state visitation distribution and the goal, which we show can lead to an optimal policy. We derive gradients for various f-divergences to optimize this objective. Our learning paradigm provides dense learning signals for exploration in sparse reward settings. We further introduce an entropy-regularized policy optimization objective, that we call $state$-MaxEnt RL (or $s$-MaxEnt RL) as a special case of our objective. We show that several metric-based shaping rewards like L2 can be used with $s$-MaxEnt RL, providing a common ground to study such metric-based shaping rewards with efficient exploration. We find that $f$-PG has better performance compared to standard policy gradient methods on a challenging gridworld as well as the Point Maze and FetchReach environments. More information on our website https://agarwalsiddhant10.github.io/projects/fpg.html.
Siddhant Agarwal, Ishan Durugkar, Peter Stone 0001, Amy Zhang 0001
NeurIPS1
2021 Learning to Deceive Knowledge Graph Augmented Models via Targeted Perturbation
Mrigank Raman, Aaron Chan, Siddhant Agarwal, Peifeng Wang, Hansen Wang, Sungchul Kim, Ryan Rossi, Handong Zhao, Nedim Lipka, Xiang Ren 0001
ICLR3
2019 Traffic Sign Classification using Hybrid HOG-SURF Features and Convolutional Neural Networks
Rishabh Madan, Deepank Agrawal, Shreyas Kowshik, Harsh Maheshwari, Siddhant Agarwal, Debashish Chakravarty
ICPRAM5