Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sergio Arnaud

dblp:344/3641 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 41% 3D vision · 27% Representation and self-supervised learning · 10%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › 3d vision and language
3d vision-language grounding
1.722025
Unifying 2D and 3D Vision-Language Understanding · ICML 2025
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025
Computer vision › Vision and language › 3d vision and language
3d language grounding
0.912025
Unifying 2D and 3D Vision-Language Understanding · ICML 2025
Computer vision › 3D vision › 3d object detection
3d object localization
0.912025
LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025
Computer vision › 3D vision
3d scene understanding
0.912025
Unifying 2D and 3D Vision-Language Understanding · ICML 2025
Computer vision › Vision and language › 3d vision and language
3d vision-language understanding
0.912025
Unifying 2D and 3D Vision-Language Understanding · ICML 2025
Computer vision › 3D vision › 3d scene understanding › 3d instance segmentation
open-vocabulary 3d instance segmentation
0.912025
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025
Computer vision › Vision and language › visual grounding
referential grounding
0.912025
LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024
Computer vision › 3D vision
environmental understanding
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024
Robotics › Robot manipulation
learning from demonstration
0.812024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Machine learning › Representation and self-supervised learning › pre-training
pre-trained visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › model adaptation
task adaptation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Machine learning › Representation and self-supervised learning
visual representation
0.712023
Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? · NeurIPS 2023
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312025
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025
Robotics › Robot navigation and mapping › mobile robot navigation
indoor navigation
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
zero-shot sim-to-real transfer
0.212024
What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments? · ICRA 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.9transfer learning · 0.9self-supervised learning · 0.9masked prediction · 0.9language-conditioned mask decoder · 0.9knowledge distillation · 0.9foundation model features · 0.9differentiable rendering · 0.92d-to-3d lifting · 0.9foundation model · 0.8
YearPublicationVenuePosition
2025 From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
abstract
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes–a six-order-of-magnitude gap that severely limits performance. We introduce LIFT-GS, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This render-supervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7% mAP on open-vocabulary instance segmentation (vs. 20.2% prior SOTA) and consistent 10-30% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies fine-tuning datasets by 2$\times$, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project page: https://liftgs.github.io.
Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson 0001, Jeong Joon Park, Alexander Sax
ICML2
2025 Unifying 2D and 3D Vision-Language Understanding
abstract
Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language understanding that bridges the gap between existing 2D-centric models and the rich 3D sensory data available in embodied systems. Our approach initializes most model weights from pre-trained 2D models and trains on both 2D and 3D vision-language data. We propose a novel language-conditioned mask decoder shared across 2D and 3D modalities to ground objects effectively in both RGB and RGB-D images, outperforming box-based approaches. To further reduce the domain gap between 2D and 3D, we incorporate 2D-to-3D lifting strategies, enabling UniVLG to utilize 2D data to enhance 3D performance. With these innovations, our model achieves state-of-the-art performance across multiple 3D vision-language grounding tasks, demonstrating the potential of transferring advances from 2D vision-language learning to the data-constrained 3D domain. Furthermore, co-training on both 2D and 3D data enhances performance across modalities without sacrificing 2D capabilities. By removing the reliance on 3D mesh reconstruction and ground-truth object proposals, UniVLG sets a new standard for realistic, embodied-aligned evaluation. Code and additional visualizations are available at https://univlg.github.io.
Alexander Swerdlow, Sergio Arnaud, Ada Martin, Alexander Sax, Franziska Meier, Katerina Fragkiadaki
ICML4
2025 LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D
abstract
We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model. Code, models and dataset can be found at the project website: locate3d.atmeta.com
Paul McVay, Sergio Arnaud, Ada Martin, Arjun Majumdar, Krishna Murthy Jatavallabhula, Phillip Thomas, Ruslan Partsey, Daniel Dugas, Abha Gejji, Alexander Sax, Vincent-Pierre Berges, Mikael Henaff, Ang Cao, Ishita Prasad, Mrinal Kalakrishnan, Michael G. Rabbat, Nicolas Ballas, Mido Assran, Oleksandr Maksymets, Aravind Rajeswaran
ICML2
2024 OpenEQA: Embodied Question Answering in the Era of Foundation Models
abstract
We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models.
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran
CVPR10
2024 What Do We Learn from a Large-Scale Study of Pre-Trained Visual Representations in Sim and Real Environments?
abstract
We present a large empirical investigation on the use of pre-trained visual representations (PVRs) for training downstream policies that execute real-world tasks. Our study involves five different PVRs, each trained for five distinct manipulation or indoor navigation tasks. We performed this evaluation using three different robots and two different policy learning paradigms. From this e ort, we can arrive at three insights: 1) the performance trends of PVRs in the simulation are generally indicative of their trends in the real world, 2) the use of PVRs enables a first-of-its-kind result with indoor ImageNav (zero-shot transfer to a held-out scene in the real world), and 3) the benefits from variations in PVRs, primarily data-augmentation and fine-tuning, also transfer to the real-world performance. See project website1for additional details and visuals.
Sneha Silwal, Karmesh Yadav, Tingfan Wu, Jay Vakil, Arjun Majumdar, Sergio Arnaud, Vincent-Pierre Berges, Dhruv Batra, Aravind Rajeswaran, Mrinal Kalakrishnan, Franziska Meier, Oleksandr Maksymets
ICRA6
2023 Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
abstract
We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we systematically evaluate existing PVRs and find that none are universally dominant. To study the effect of pre-training data size and diversity, we combine over 4,000 hours of egocentric videos from 7 different sources (over 4.3M images) and ImageNet to train different-sized vision transformers using Masked Auto-Encoding (MAE) on slices of this data. Contrary to inferences from prior work, we find that scaling dataset size and diversity does not improve performance universally (but does so on average). Our largest model, named VC-1, outperforms all prior PVRs on average but does not universally dominate either. Next, we show that task- or domain-specific adaptation of VC-1 leads to substantial gains, with VC-1 (adapted) achieving competitive or superior performance than the best known results on all of the benchmarks in CortexBench. Finally, we present real-world hardware experiments, in which VC-1 and VC-1 (adapted) outperform the strongest pre-existing PVR. Overall, this paper presents no new techniques but a rigorous systematic evaluation, a broad set of findings about PVRs (that in some cases, refute those made in narrow domains in prior work), and open-sourced code and models (that required over 10,000 GPU-hours to train) for the benefit of the research community.
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 0001, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, Pieter Abbeel, Jitendra Malik, Dhruv Batra, Oleksandr Maksymets, Aravind Rajeswaran, Franziska Meier
NeurIPS3