Jay Vakil

dblp:345/8174 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0002-1166-1897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 26% Motion planning and robot control · 23% Representation and self-supervised learning · 20%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › hierarchical reinforcement learning
action chunking
0.812024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Machine learning › Reinforcement learning
imitation learning
0.812024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Robotics › Robot manipulation
learning from demonstration
0.812024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Robotics › Motion planning and robot control › robot learning › manipulation task learning
multi-task manipulation
0.812024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Robotics › Motion planning and robot control
robot learning
0.812024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Machine learning › Representation and self-supervised learning › pre-training
pre-trained visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › model adaptation
task adaptation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Representation and self-supervised learning
visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Deep learning architectures and training
data augmentation
0.212024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Machine learning › Deep learning architectures and training › data augmentation
semantic augmentation
0.212024
RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking · ICRA 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
zero-shot sim-to-real transfer
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024

Methods — techniques the papers use, named apart from their topics

task conditioning · 0.8semantic augmentation · 0.8pre-trained visual representation · 0.8fine-tuning · 0.8data augmentation · 0.8action chunking · 0.8vision transformer · 0.7masked autoencoding · 0.7
YearPublicationVenuePosition
2024 RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking
abstract
The grand aim of having a single robot that can manipulate arbitrary objects in diverse settings is at odds with the paucity of robotics datasets. Acquiring and growing such datasets is strenuous due to manual efforts, operational costs, and safety challenges. A path toward such a universal agent requires an efficient framework capable of generalization but within a reasonable data budget. In this paper, we develop an efficient framework (MT-ACT) for training universal agents capable of multi-task manipulation skills using (a) semantic augmentations that can rapidly multiply existing datasets and (b) action representations that can extract performant policies with small yet diverse multi-modal datasets without overfitting. In addition, reliable task conditioning and an expressive policy architecture enables our agent to exhibit a diverse repertoire of skills in novel situations specified using task commands. Using merely 7500 demonstrations, we are able to train a single policy RoboAgent capable of 12 unique skills, and demonstrate its generalization over 38 tasks spread across common daily activities in diverse kitchen scenes. On average, RoboAgent outperforms prior methods by over 40% in unseen situations while being more sample efficient. See https://robopen.github.io/for video results and appendix.
Homanga Bharadhwaj, Jay Vakil, Mohit Sharma 0001, Abhinav Gupta 0001, Shubham Tulsiani
ICRA2
2024 What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?
abstract
We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluation using three different robots and two different policy learning paradigms. From this e ort, we can arrive at three insights: 1) the performance trends of PVRs in the simulation are generally indicative of their trends in the real world, 2) the use of PVRs enables a first-of-its-kind result with indoor ImageNav (zero-shot transfer to a held-out scene in the real world), and 3) the benefits from variations in PVRs, primarily data-augmentation and fine-tuning, also transfer to the real-world performance. See project website1for additional details and visuals.
Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, Oleksandr Maksymets
ICRA4
2023 Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
abstract
We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we systematically evaluate existing PVRs and find that none are universally dominant. To study the effect of pre-training data size and diversity, we combine over 4,000 hours of egocentric videos from 7 different sources (over 4.3M images) and ImageNet to train different-sized vision transformers using Masked Auto-Encoding (MAE) on slices of this data. Contrary to inferences from prior work, we find that scaling dataset size and diversity does not improve performance universally (but does so on average). Our largest model, named VC-1, outperforms all prior PVRs on average but does not universally dominate either. Next, we show that task- or domain-specific adaptation of VC-1 leads to substantial gains, with VC-1 (adapted) achieving competitive or superior performance than the best known results on all of the benchmarks in CortexBench. Finally, we present real-world hardware experiments, in which VC-1 and VC-1 (adapted) outperform the strongest pre-existing PVR. Overall, this paper presents no new techniques but a rigorous systematic evaluation, a broad set of findings about PVRs (that in some cases, refute those made in narrow domains in prior work), and open-sourced code and models (that required over 10,000 GPU-hours to train) for the benefit of the research community.
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 0001, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Oleksandr Maksymets, Aravind Rajeswaran, Franziska Meier
NeurIPS10