EDBT 2026 Demo / reviewers in the wild / expert
Hanoona Rasheed
dblp:405/4533
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 19% Vision and language · 19% Language models and text generation · 19% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Image recognition and object detection
spatial alignment |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
vision encoder |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Video understanding and tracking
video classification |
0.3 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 0.9alignment method · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Perception Encoder: The best visual embeddings are not at the output of the networkabstractWe introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models Daniel Bolya, Po-Yao Huang 0001, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei 0005, Tengyu Ma 0005, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Marco Monteiro, Hu Xu 0001, Shiyu Dong, Nikhila Ravi, Shang-Wen Li 0001, Piotr Dollár, Christoph Feichtenhofer |
NeurIPS | 10 |