Karmesh Yadav

dblp:264/3702 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Robot navigation and mapping · 22% Reinforcement learning · 15% 3D vision · 12%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
object goal navigation
1.322023
Navigating to Objects Specified by Images · ICCV 2023
Habitat-Matterport 3D Semantics Dataset · CVPR 2023
Machine learning › Reinforcement learning
long-horizon tasks
0.912025
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › partially observable reinforcement learning
memory-based reinforcement learning
0.912025
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control › robot learning
embodied manipulation
0.812024
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control · NeurIPS 2024
Robotics › Robot navigation and mapping
embodied navigation
0.812024
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024
Computer vision › 3D vision
environmental understanding
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024
Robotics › Robot manipulation
learning from demonstration
0.812024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Computer vision › 3D vision
3d scene understanding
0.712023
Habitat-Matterport 3D Semantics Dataset · CVPR 2023
Computer vision › Segmentation and scene understanding
3d semantic segmentation
0.712023
Habitat-Matterport 3D Semantics Dataset · CVPR 2023
Machine learning › Representation and self-supervised learning › pre-training
pre-trained visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › model adaptation
task adaptation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Representation and self-supervised learning
visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Learning paradigms
continual learning
0.412020
Look-ahead Meta Learning for Continual Learning · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
meta-learning
0.412020
Look-ahead Meta Learning for Continual Learning · NeurIPS 2020
Machine learning › Learning theory
online learning
0.412020
Look-ahead Meta Learning for Continual Learning · NeurIPS 2020
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
zero-shot sim-to-real transfer
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024

Methods — techniques the papers use, named apart from their topics

transformer · 0.9summarization tokens · 0.9representation learning · 0.8pre-trained visual representation · 0.8large language model evaluation · 0.8foundation model · 0.8fine-tuning · 0.8diffusion model · 0.8data augmentation · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2025 Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
abstract
To enable embodied agents to operate effectively over extended timeframes, it is crucial to develop models that form and access memories to stay contextualized in their environment. In the current paradigm of training transformer-based policies for embodied sequential decision-making tasks, visual inputs often overwhelm the context limits of transformers, while humans can maintain and utilize a lifetime of experience compressed as memories. Significant compression is possible in principle, as much of the input is irrelevant and can be abstracted. However, existing approaches predominantly focus on either recurrent models with fixed-size memory or transformers with full-context reliance. In this work, we propose Memo, a transformer-based architecture and training recipe for reinforcement learning (RL) on memory-intensive, long-horizon tasks. Memo incorporates the creation and retrieval of memory by interleaving periodic summarization tokens with the inputs of a model during training. We demonstrate Memo’s effectiveness on a grid-world meta-RL benchmark and a multi-object navigation task in photo-realistic indoor settings. Memo outperforms naive long-context transformer baselines while being more compute and storage efficient. Additionally, Memo generalizes better to longer contexts at inference time and remains robust in streaming settings, where historical context must be truncated to fit inference constraints.
Gunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal, Rahaf Aljundi
NeurIPS2
2024 OpenEQA: Embodied Question Answering in the Era of Foundation Models
abstract
We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models.
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran
CVPR11
2024 What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?
abstract
We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluation using three different robots and two different policy learning paradigms. From this e ort, we can arrive at three insights: 1) the performance trends of PVRs in the simulation are generally indicative of their trends in the real world, 2) the use of PVRs enables a first-of-its-kind result with indoor ImageNav (zero-shot transfer to a held-out scene in the real world), and 3) the benefits from variations in PVRs, primarily data-augmentation and fine-tuning, also transfer to the real-world performance. See project website1for additional details and visuals.
Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, Oleksandr Maksymets
ICRA2
2024 Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
abstract
Embodied AI agents require a fine-grained understanding of the physical world mediated through visual and language inputs. Such capabilities are difficult to learn solely from task-specific data. This has led to the emergence of pre-trained vision-language models as a tool for transferring representations learned from internet-scale data to downstream tasks and new domains. However, commonly used contrastively trained representations such as in CLIP have been shown to fail at enabling embodied agents to gain a sufficiently fine-grained scene understanding—a capability vital for control. To address this shortcoming, we consider representations from pre-trained text-to-image diffusion models, which are explicitly optimized to generate images from text prompts and as such, contain text-conditioned representations that reflect highly fine-grained visuo-spatial information. Using pre-trained text-to-image diffusion models, we construct Stable Control Representations which allow learning downstream control policies that generalize to complex, open-ended environments. We show that policies learned using Stable Control Representations are competitive with state-of-the-art representation learning approaches across a broad range of simulated control settings, encompassing challenging manipulation and navigation tasks. Most notably, we show that Stable Control Representations enable learning policies that exhibit state-of-the-art performance on OVMM, a difficult open-vocabulary navigation benchmark.
Gunshi Gupta, Karmesh Yadav, Yarin Gal, Dhruv Batra, Zsolt Kira, Cong Lu, Tim G. J. Rudner
NeurIPS2
2023 Habitat-Matterport 3D Semantics Dataset
abstract
We present the Habitat-Matterport 3D Semantics (HM3DSem) dataset. HM3DSem is the largest dataset of 3D real-world spaces with densely annotated semantics that is currently available to the academic community. It consists of 142,646 object instance annotations across 216 3D spaces and 3,100 rooms within those spaces. The scale, quality, and diversity of object annotations far exceed those of prior datasets. A key difference setting apart HM3DSem from other datasets is the use of texture information to annotate pixel-accurate object boundaries. We demonstrate the effectiveness of HM3DSem dataset for the Object Goal Navigation task using different methods. Policies trained using HM3DSem perform outperform those trained on prior datasets. Introduction of HM3DSem in the Habitat ObjectNav Challenge lead to an increase in participation from 400 submissions in 2021 to 1022 submissions in 2022. Project page: https://aihabitat.org/datasets/hm3d-semantics/
Karmesh Yadav, Ram Ramrakhya, Santhosh K. Ramakrishnan, Théophile Gervet, John M. Turner, Aaron Gokaslan, Noah Maestre, Angel X. Chang, Dhruv Batra, Manolis Savva, Alexander Clegg, Devendra Singh Chaplot
CVPR1
2023 Navigating to Objects Specified by Images
abstract
Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modular method solves sub-tasks of exploration, goal instance re-identification, goal localization, and local navigation. We re-identify the goal instance in egocentric vision using feature-matching and localize the goal instance by projecting matched features to a map. Each sub-task is solved using off-the-shelf components requiring zero fine-tuning. On the HM3D InstanceImageNav benchmark, this system outperforms a baseline end-to-end RL policy 7x and a state-of-the-art ImageNav model 2.3x (56% vs . 25% success). We deploy this system to a mobile robot platform and demonstrate effective real-world performance, achieving an 88% success rate across a home and an office environment.
Jacob Krantz, Théophile Gervet, Karmesh Yadav, Austin S. Wang, Chris Paxton 0001, Roozbeh Mottaghi, Dhruv Batra, Jitendra Malik, Stefan Lee, Devendra Singh Chaplot
ICCV3
2023 Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
abstract
We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we systematically evaluate existing PVRs and find that none are universally dominant. To study the effect of pre-training data size and diversity, we combine over 4,000 hours of egocentric videos from 7 different sources (over 4.3M images) and ImageNet to train different-sized vision transformers using Masked Auto-Encoding (MAE) on slices of this data. Contrary to inferences from prior work, we find that scaling dataset size and diversity does not improve performance universally (but does so on average). Our largest model, named VC-1, outperforms all prior PVRs on average but does not universally dominate either. Next, we show that task- or domain-specific adaptation of VC-1 leads to substantial gains, with VC-1 (adapted) achieving competitive or superior performance than the best known results on all of the benchmarks in CortexBench. Finally, we present real-world hardware experiments, in which VC-1 and VC-1 (adapted) outperform the strongest pre-existing PVR. Overall, this paper presents no new techniques but a rigorous systematic evaluation, a broad set of findings about PVRs (that in some cases, refute those made in narrow domains in prior work), and open-sourced code and models (that required over 10,000 GPU-hours to train) for the benefit of the research community.
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 0001, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Oleksandr Maksymets, Aravind Rajeswaran, Franziska Meier
NeurIPS2
2020 Look-ahead Meta Learning for Continual Learning
abstract
The continual learning problem involves training models with limited capacity to perform well on a set of an unknown number of sequentially arriving tasks. While meta-learning shows great potential for reducing interference between old and new tasks, the current training procedures tend to be either slow or offline, and sensitive to many hyper-parameters. In this work, we propose Look-ahead MAML (La-MAML), a fast optimisation-based meta-learning algorithm for online-continual learning, aided by a small episodic memory. By incorporating the modulation of per-parameter learning rates in our meta-learning update, our approach also allows us to draw connections to and exploit prior work on hypergradients and meta-descent. This provides a more flexible and efficient way to mitigate catastrophic forgetting compared to conventional prior-based methods. La-MAML achieves performance superior to other replay-based, prior-based and meta-learning based approaches for continual learning on real-world visual classification benchmarks.
Gunshi Gupta, Karmesh Yadav, Liam Paull
NeurIPS2