VLDB 2026 Research / reviewers in the wild / expert
Sneha Silwal
dblp:317/0840
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Robot manipulation · 26% Representation and self-supervised learning · 21% Reinforcement learning · 14% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Computer vision › 3D vision
environmental understanding |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Robotics › Robot manipulation
learning from demonstration |
0.8 | 1 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 |
Robotics › Robot manipulation
dexterous manipulation |
0.7 | 1 | 2023 | Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation · ICRA 2023 |
Machine learning › Reinforcement learning
imitation learning |
0.7 | 1 | 2023 | Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation · ICRA 2023 |
Machine learning › Representation and self-supervised learning › pre-training
pre-trained visual representation |
0.7 | 1 | 2023 | Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › model adaptation
task adaptation |
0.7 | 1 | 2023 | Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
visual representation |
0.7 | 1 | 2023 | Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023 |
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation |
0.2 | 1 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 |
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
zero-shot sim-to-real transfer |
0.2 | 1 | 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024 |
Robotics › Robot manipulation › dexterous manipulation
in-hand manipulation |
0.2 | 1 | 2023 | Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous Manipulation · ICRA 2023 |
Methods — techniques the papers use, named apart from their topics
pre-trained visual representation · 0.8large language model evaluation · 0.8foundation model · 0.8fine-tuning · 0.8data augmentation · 0.8vision transformer · 0.7masked autoencoding · 0.7imitation learning · 0.7RGB camera teleoperation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation ModelsabstractWe present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models. Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran |
CVPR | 7 |
| 2024 | What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?abstractWe present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluation using three different robots and two different policy learning paradigms. From this e ort, we can arrive at three insights: 1) the performance trends of PVRs in the simulation are generally indicative of their trends in the real world, 2) the use of PVRs enables a first-of-its-kind result with indoor ImageNav (zero-shot transfer to a held-out scene in the real world), and 3) the benefits from variations in PVRs, primarily data-augmentation and fine-tuning, also transfer to the real-world performance. See project website1for additional details and visuals. Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, Oleksandr Maksymets |
ICRA | 1 |
| 2023 | Dexterous Imitation Made Easy: A Learning-Based Framework for Efficient Dexterous ManipulationabstractOptimizing behaviors for dexterous manipulation has been a longstanding challenge in robotics, with a variety of methods from model-based control to model-free reinforcement learning having been previously explored in literature. Such prior work often require extensive trial-and-error training along with task-specific tuning of reward functions, which makes applying dexterous manipulation for general purpose problems quite impractical. A sample-efficient and practical alternate to trial-and-error learning is imitation learning. However, collecting and learning from demonstrations in dexterous manipulation is quite challenging due to the high-dimensional action-space involved with multi-finger control. In this work, we propose ‘Dexterous Imitation Made Easy’ (DIME) a new imitation learning framework for dexterous manipulation. DIME only requires a single RGB camera that observes a human operator to teleoperate a robotic hand. Once demonstrations are collected, DIME employs state-of-the-art imitation learning methods to train dexterous manipulation policies. On real robot benchmarks we demonstrate that DIME can be used to solve complex, in-hand manipulation tasks such as ‘flipping’, ‘spinning’, and ‘rotating’ objects with just 30 demonstrations and no additional robot training. Our code, pre-collected demonstrations, and robot videos are publicly available at: https://nyu-robot-learning.github.io/dime. Sridhar Pandian Arunachalam, Sneha Silwal, Ben Evans, Lerrel Pinto |
ICRA | 2 |
| 2023 | Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?abstractWe present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we systematically evaluate existing PVRs and find that none are universally dominant. To study the effect of pre-training data size and diversity, we combine over 4,000 hours of egocentric videos from 7 different sources (over 4.3M images) and ImageNet to train different-sized vision transformers using Masked Auto-Encoding (MAE) on slices of this data. Contrary to inferences from prior work, we find that scaling dataset size and diversity does not improve performance universally (but does so on average). Our largest model, named VC-1, outperforms all prior PVRs on average but does not universally dominate either. Next, we show that task- or domain-specific adaptation of VC-1 leads to substantial gains, with VC-1 (adapted) achieving competitive or superior performance than the best known results on all of the benchmarks in CortexBench. Finally, we present real-world hardware experiments, in which VC-1 and VC-1 (adapted) outperform the strongest pre-existing PVR. Overall, this paper presents no new techniques but a rigorous systematic evaluation, a broad set of findings about PVRs (that in some cases, refute those made in narrow domains in prior work), and open-sourced code and models (that required over 10,000 GPU-hours to train) for the benefit of the research community. Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 0001, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Oleksandr Maksymets, Aravind Rajeswaran, Franziska Meier |
NeurIPS | 6 |