EDBT 2026 Demo / reviewers in the wild / expert
Tengyu Ma 0005
dblp:384/7631
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Segmentation and scene understanding · 31% Video understanding and tracking · 14% Deep learning architectures and training · 14% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.9 | 1 | 2025 | SAM 2: Segment Anything in Images and Videos · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Video understanding and tracking › video object segmentation
promptable video segmentation |
0.9 | 1 | 2025 | SAM 2: Segment Anything in Images and Videos · ICLR 2025 |
Computer vision › Segmentation and scene understanding
prompt-based segmentation |
0.9 | 1 | 2025 | SAM 2: Segment Anything in Images and Videos · ICLR 2025 |
Computer vision › Image recognition and object detection
spatial alignment |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding
video segmentation |
0.9 | 1 | 2025 | SAM 2: Segment Anything in Images and Videos · ICLR 2025 |
Machine learning › Deep learning architectures and training
vision encoder |
0.9 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.3 | 1 | 2025 | SAM 2: Segment Anything in Images and Videos · ICLR 2025 |
Computer vision › Video understanding and tracking
video classification |
0.3 | 1 | 2025 | Perception Encoder: The best visual embeddings are not at the output of the network · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9streaming memory · 0.9contrastive learning · 0.9alignment method · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAM 2: Segment Anything in Images and VideosabstractWe present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. Our model is a simple transformer architecture with streaming memory for real-time video processing. SAM 2 trained on our data provides strong performance across a wide range of tasks. In video segmentation, we observe better accuracy, using 3x fewer interactions than prior approaches. In image segmentation, our model is more accurate and 6x faster than the Segment Anything Model (SAM). We believe that our data, model, and insights will serve as a significant milestone for video segmentation and related perception tasks. We are releasing our main model, the dataset, an interactive demo and code. Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma 0005, Haitham Khedr, Roman Rädle, Chloé Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross B. Girshick, Piotr Dollár, Christoph Feichtenhofer |
ICLR | 6 |
| 2025 | Perception Encoder: The best visual embeddings are not at the output of the networkabstractWe introduce Perception Encoder (PE), a family of state-of-the-art vision encoders for image and video understanding. Traditionally, vision encoders have relied on a variety of pretraining objectives, each excelling at different downstream tasks. Surprisingly, after scaling a carefully tuned image pretraining recipe and refining with a robust video data engine, we find that contrastive vision-language training alone can produce strong, general embeddings for all of these downstream tasks. There is only one caveat: these embeddings are hidden within the intermediate layers of the network. To draw them out, we introduce two alignment methods: language alignment for multimodal language modeling, and spatial alignment for dense prediction. Together, our PE family of models achieves state-of-the-art results on a wide variety of tasks, including zero-shot image and video classification and retrieval; document, image, and video Q&A; and spatial tasks such as detection, tracking, and depth estimation. We release our models, code, and novel dataset of synthetically and human-annotated videos: https://github.com/facebookresearch/perception_models Daniel Bolya, Po-Yao Huang 0001, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei 0005, Tengyu Ma 0005, Jiale Zhi, Jathushan Rajasegaran, Hanoona Rasheed, Marco Monteiro, Hu Xu 0001, Shiyu Dong, Nikhila Ravi, Shang-Wen Li 0001, Piotr Dollár, Christoph Feichtenhofer |
NeurIPS | 7 |
| 2025 | PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingabstractVision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM–VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about ''what'', ''where'', ''when'', and ''how'' of a video. We make our work fully reproducible by providing data, training recipes, code & models. Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras, Tushar Nagarajan, Muhammad Maaz 0001, Yale Song, Tengyu Ma 0005, Shuming Hu, Suyog Dutt Jain, Hanoona Abdul Rasheed, Peize Sun, Po-Yao Huang 0001, Daniel Bolya, Nikhila Ravi, Shashank Jain, Tammy Stark, Seungwhan Moon, Babak Damavandi, Vivian Lee, Andrew Westbury, Salman Khan 0001, Philipp Krähenbühl, Piotr Dollár, Lorenzo Torresani, Kristen Grauman, Christoph Feichtenhofer |
NeurIPS | 8 |