EDBT 2026 Demo / reviewers in the wild / expert
Paul McVay
dblp:313/2226
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 45% 3D vision · 43% Question answering and dialogue systems · 12% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d object detection
3d object localization |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Computer vision › Vision and language › 3d vision and language
3d vision-language grounding |
0.9 | 1 | 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Computer vision › 3D vision › 3d scene understanding › 3d instance segmentation
open-vocabulary 3d instance segmentation |
0.9 | 1 | 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Computer vision › Vision and language › visual grounding
referential grounding |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Computer vision › 3D vision
environmental understanding |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.3 | 1 | 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9self-supervised learning · 0.9masked prediction · 0.9knowledge distillation · 0.9foundation model features · 0.9differentiable rendering · 0.9large language model evaluation · 0.8foundation model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMsabstract3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes–a six-order-of-magnitude gap that severely limits performance. We introduce LIFT-GS, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This render-supervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7% mAP on open-vocabulary instance segmentation (vs. 20.2% prior SOTA) and consistent 10-30% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies fine-tuning datasets by 2$\times$, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project page: https://liftgs.github.io. Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson 0001, Jeong Joon Park, Alexander Sax |
ICML | 8 |
| 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3DabstractWe present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model. Code, models and dataset can be found at the project website: locate3d.atmeta.com Paul McVay, Sergio Arnaud, Ada Martin, Arjun Majumdar, Krishna Murthy Jatavallabhula, Phillip Thomas, Ruslan Partsey, Daniel Dugas, Abha Gejji, Alexander Sax, Vincent-Pierre Berges, Mikael Henaff, Ang Cao, Ishita Prasad, Mrinal Kalakrishnan, Michael G. Rabbat, Nicolas Ballas, Mido Assran, Oleksandr Maksymets, Aravind Rajeswaran |
ICML | 1 |
| 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation ModelsabstractWe present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models. Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran |
CVPR | 8 |