Sashwat Mahalingam

dblp:351/0792 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Motion planning and robot control · 44% Transfer learning and domain adaptation · 44% Representation and self-supervised learning · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot learning › visuomotor learning
visuomotor skill learning
0.812024
SpawnNet: Learning Generalizable Visuomotor Skills from Pre-trained Network · ICRA 2024
Machine learning › Representation and self-supervised learning › pre-training
pre-trained visual representation
0.212024
SpawnNet: Learning Generalizable Visuomotor Skills from Pre-trained Network · ICRA 2024

Methods — techniques the papers use, named apart from their topics

multi-layer representation fusion · 0.8imitation learning · 0.8
YearPublicationVenuePosition
2024 SpawnNet: Learning Generalizable Visuomotor Skills from Pre-trained Network
abstract
The existing internet-scale image and video datasets cover a wide range of everyday objects and tasks, bringing the potential of learning policies that generalize in diverse scenarios. Prior works have explored visual pre-training with different self-supervised objectives. Still, the generalization capabilities of the learned policies and the advantages over well-tuned baselines remain unclear from prior studies. In this work, we present a focused study of the generalization capabilities of the pre-trained visual representations at the categorical level. We identify the key bottleneck in using a frozen pre-trained visual backbone for policy learning and then propose SpawnNet, a novel two-stream architecture that learns to fuse pre-trained multi-layer representations into a separate network to learn a robust policy. Through extensive simulated and real experiments, we show significantly better categorical generalization compared to prior approaches in imitation learning settings. Open-sourced code and videos can be found on our website: https://xingyu-lin.github.io/spawnnet/.
John So, Sashwat Mahalingam, Fangchen Liu, Pieter Abbeel
ICRA3