EDBT 2026 Demo / reviewers in the wild / expert
Alexander Sax
dblp:218/6484 · also Sasha Sax
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0009-1613-076XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
3D vision · 43% Vision and language · 29% Transfer learning and domain adaptation · 7% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% |
Topics — the 23 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › 3d vision and language
3d vision-language grounding |
1.7 | 2 | 2025 | Unifying 2D and 3D Vision-Language Understanding · ICML 2025 From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Computer vision › Vision and language › 3d vision and language
3d language grounding |
0.9 | 1 | 2025 | Unifying 2D and 3D Vision-Language Understanding · ICML 2025 |
Computer vision › 3D vision › 3d object detection
3d object localization |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass · CVPR 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | Unifying 2D and 3D Vision-Language Understanding · ICML 2025 |
Computer vision › Vision and language › 3d vision and language
3d vision-language understanding |
0.9 | 1 | 2025 | Unifying 2D and 3D Vision-Language Understanding · ICML 2025 |
Computer vision › 3D vision
camera pose estimation |
0.9 | 1 | 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass · CVPR 2025 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.9 | 1 | 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass · CVPR 2025 |
Computer vision › 3D vision › 3d scene understanding › 3d instance segmentation
open-vocabulary 3d instance segmentation |
0.9 | 1 | 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Computer vision › Vision and language › visual grounding
referential grounding |
0.9 | 1 | 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3D · ICML 2025 |
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Computer vision › 3D vision
environmental understanding |
0.8 | 1 | 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024 |
Machine learning › Transfer learning and domain adaptation
cross-task transfer |
0.7 | 2 | 2019 | Taskonomy: Disentangling Task Transfer Learning · IJCAI 2019 Taskonomy: Disentangling Task Transfer Learning · CVPR 2018 |
Computer vision › 3D vision
depth estimation |
0.5 | 1 | 2021 | Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans · ICCV 2021 |
Computer vision › 3D vision
surface normal estimation |
0.5 | 1 | 2021 | Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scans · ICCV 2021 |
Machine learning › Transfer learning and domain adaptation › model adaptation
network adaptation |
0.4 | 1 | 2020 | Side-Tuning: A Baseline for Network Adaptation via Additive Side Networks · ECCV (3) 2020 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.4 | 1 | 2020 | Robust Learning Through Cross-Task Consistency · CVPR 2020 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2020 | Robust Learning Through Cross-Task Consistency · CVPR 2020 |
Robotics › Robot navigation and mapping
embodied perception |
0.3 | 1 | 2018 | Gibson Env: Real-World Perception for Embodied Agents · CVPR 2018 |
Machine learning › Learning paradigms
multi-task learning |
0.3 | 1 | 2018 | Taskonomy: Disentangling Task Transfer Learning · CVPR 2018 |
Computer vision › 3D vision › neural rendering
3d gaussian splatting |
0.3 | 1 | 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs · ICML 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.1 | 1 | 2018 | Gibson Env: Real-World Perception for Embodied Agents · CVPR 2018 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9transformer · 0.9transfer learning · 0.9masked prediction · 0.9language-conditioned mask decoder · 0.9knowledge distillation · 0.9foundation model features · 0.9feed-forward inference · 0.9differentiable rendering · 0.92d-to-3d lifting · 0.9virtualizing real spaces · 0.3domain adaptation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassabstractMulti-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R’s Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy. Alexander Sax, Kevin J. Liang, Mikael Henaff, Ang Cao, Joyce Y. Chai, Franziska Meier, Matt Feiszli |
CVPR | 2 |
| 2025 | From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMsabstract3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes–a six-order-of-magnitude gap that severely limits performance. We introduce LIFT-GS, a practical distillation technique that overcomes this limitation by using differentiable rendering to bridge 3D and 2D supervision. LIFT-GS predicts 3D Gaussian representations from point clouds and uses them to render predicted language-conditioned 3D masks into 2D views, enabling supervision from 2D foundation models (SAM, CLIP, LLaMA) without requiring any 3D annotations. This render-supervised formulation enables end-to-end training of complete encoder-decoder architectures and is inherently model-agnostic. LIFT-GS achieves state-of-the-art results with 25.7% mAP on open-vocabulary instance segmentation (vs. 20.2% prior SOTA) and consistent 10-30% improvements on referential grounding tasks. Remarkably, pretraining effectively multiplies fine-tuning datasets by 2$\times$, demonstrating strong scaling properties that suggest 3D VLG currently operates in a severely data-scarce regime. Project page: https://liftgs.github.io. Ang Cao, Sergio Arnaud, Oleksandr Maksymets, Ada Martin, Vincent-Pierre Berges, Paul McVay, Ruslan Partsey, Aravind Rajeswaran, Franziska Meier, Justin Johnson 0001, Jeong Joon Park, Alexander Sax |
ICML | 14 |
| 2025 | Unifying 2D and 3D Vision-Language UnderstandingabstractProgress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language understanding that bridges the gap between existing 2D-centric models and the rich 3D sensory data available in embodied systems. Our approach initializes most model weights from pre-trained 2D models and trains on both 2D and 3D vision-language data. We propose a novel language-conditioned mask decoder shared across 2D and 3D modalities to ground objects effectively in both RGB and RGB-D images, outperforming box-based approaches. To further reduce the domain gap between 2D and 3D, we incorporate 2D-to-3D lifting strategies, enabling UniVLG to utilize 2D data to enhance 3D performance. With these innovations, our model achieves state-of-the-art performance across multiple 3D vision-language grounding tasks, demonstrating the potential of transferring advances from 2D vision-language learning to the data-constrained 3D domain. Furthermore, co-training on both 2D and 3D data enhances performance across modalities without sacrificing 2D capabilities. By removing the reliance on 3D mesh reconstruction and ground-truth object proposals, UniVLG sets a new standard for realistic, embodied-aligned evaluation. Code and additional visualizations are available at https://univlg.github.io. Alexander Swerdlow, Sergio Arnaud, Ada Martin, Alexander Sax, Franziska Meier, Katerina Fragkiadaki |
ICML | 6 |
| 2025 | LOCATE 3D: Real-World Object Localization via Self-Supervised Learning in 3DabstractWe present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model. Code, models and dataset can be found at the project website: locate3d.atmeta.com Paul McVay, Sergio Arnaud, Ada Martin, Arjun Majumdar, Krishna Murthy Jatavallabhula, Phillip Thomas, Ruslan Partsey, Daniel Dugas, Abha Gejji, Alexander Sax, Vincent-Pierre Berges, Mikael Henaff, Ang Cao, Ishita Prasad, Mrinal Kalakrishnan, Michael G. Rabbat, Nicolas Ballas, Mido Assran, Oleksandr Maksymets, Aravind Rajeswaran |
ICML | 10 |
| 2024 | OpenEQA: Embodied Question Answering in the Era of Foundation ModelsabstractWe present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models. Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran |
CVPR | 23 |
| 2021 | Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D ScansabstractThis paper introduces a pipeline to parametrically sample and render static multi-task vision datasets from comprehensive 3D scans from the real-world. In addition to enabling interesting lines of research, we show the tooling and generated data suffice to train robust vision models. Familiar architectures trained on a generated starter dataset reached state-of-the-art performance on multiple common vision tasks and benchmarks, despite having seen no benchmark or non-pipeline data. The depth estimation network outperforms MiDaS and the surface normal estimation network is the first to achieve human-level performance for in-the-wild surface normal estimation—at least according to one metric on the OASIS benchmark. The Dockerized pipeline with CLI, the (mostly python) code, PyTorch dataloaders for the generated data, the generated starter dataset, download scripts and other utilities are all available ${\color{Magenta}through}\;{\color{Magenta}our}\;{\color{Magenta}project}\;{\color{Magenta}website}$. Ainaz Eftekhar, Alexander Sax, Jitendra Malik, Amir Zamir |
ICCV | 2 |
| 2020 | Robust Learning Through Cross-Task ConsistencyabstractVisual perception entails solving a wide set of tasks (e.g., object detection, depth estimation, etc). The predictions made for different tasks out of one image are not independent, and therefore, are expected to be 'consistent'. We propose a flexible and fully computational framework for learning while enforcing Cross-Task Consistency (X-TAC). The proposed formulation is based on 'inference path invariance' over an arbitrary graph of prediction domains. We observe that learning with cross-task consistency leads to more accurate predictions, better generalization to out-of-distribution samples, and improved sample efficiency. This framework also leads to a powerful unsupervised quantity, called 'Consistency Energy, based on measuring the intrinsic consistency of the system. Consistency Energy well correlates with the supervised error (r=0.67), thus it can be employed as an unsupervised robustness metric as well as for detection of out-of-distribution inputs (AUC=0.99). The evaluations were performed on multiple datasets, including Taskonomy, Replica, CocoDoom, and ApolloScape. Amir Zamir, Alexander Sax, Nikhil Cheerla, Rohan Suri, Zhangjie Cao, Jitendra Malik, Leonidas J. Guibas |
CVPR | 2 |
| 2020 | Side-Tuning: A Baseline for Network Adaptation via Additive Side Networks
Jeffrey O. Zhang, Alexander Sax, Amir Zamir, Leonidas J. Guibas, Jitendra Malik |
ECCV (3) | 2 |
| 2019 | Taskonomy: Disentangling Task Transfer Learning
Amir Zamir, Alexander Sax, Bokui Shen, Leonidas J. Guibas, Jitendra Malik, Silvio Savarese |
IJCAI | 2 |
| 2018 | Gibson Env: Real-World Perception for Embodied AgentsabstractDeveloping visual perception models for active agents and sensorimotor control in the physical world are cumbersome as existing algorithms are too slow to efficiently learn in real-time and robots are fragile and costly. This has given rise to learning-in-simulation which consequently casts a question on whether the results transfer to real-world. In this paper, we investigate developing real-world perception for active agents, propose Gibson Environment for this purpose, and showcase a set of perceptual tasks learned therein. Gibson is based upon virtualizing real spaces, rather than artificially designed ones, and currently includes over 1400 floor spaces from 572 full buildings. The main characteristics of Gibson are: I. being from the real-world and reflecting its semantic complexity, II. having an internal synthesis mechanism "Goggles" enabling deploying the trained models in real-world without needing domain adaptation, III. embodiment of agents and making them subject to constraints of physics and space. Fei Xia 0002, Amir Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, Silvio Savarese |
CVPR | 4 |
| 2018 | Taskonomy: Disentangling Task Transfer LearningabstractDo visual tasks have relationships, or are they unrelated? For instance, could having surface normals simplify estimating the depth of an image? Intuition answers these questions positively, implying existence of a certain structure among visual tasks. Knowing this structure has notable values; it provides a principled way for identifying relationships across tasks, for instance, in order to reuse supervision among tasks with redundancies or solve many tasks in one system without piling up the complexity. We propose a fully computational approach for modeling the transfer learning structure of the space of visual tasks. This is done via finding transfer learning dependencies across tasks in a dictionary of twenty-six 2D, 2.5D, 3D, and semantic tasks. The product is a computational taxonomic map among tasks for transfer learning, and we exploit it to reduce the demand for labeled data. For example, we show that the total number of labeled datapoints needed for solving a set of 10 tasks can be reduced by roughly 2/3 (compared to training independently) while keeping the performance nearly the same. We provide a set of tools for computing and visualizing this taxonomical structure at http://taskonomy.vision. Amir Zamir, Alexander Sax, Bokui Shen, Leonidas J. Guibas, Jitendra Malik, Silvio Savarese |
CVPR | 2 |