EDBT 2026 Demo / reviewers in the wild / expert
Orr Zohar
dblp:335/1624
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-9091-7140ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 27% Language models and text generation · 20% Vision and language · 15% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Machine learning › Generative modeling
image generation |
0.9 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Machine learning › Generative modeling
image tokenization |
0.9 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | Apollo: An Exploration of Video Understanding in Large Multimodal Models · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.9 | 1 | 2025 | Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Computer vision › Vision and language
video-language model |
0.9 | 1 | 2025 | Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
video multimodal large language model |
0.9 | 1 | 2025 | Apollo: An Exploration of Video Understanding in Large Multimodal Models · CVPR 2025 |
Computer vision › Video understanding and tracking
video question answering |
0.9 | 1 | 2025 | Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024 |
Natural language and speech › Language models and text generation
LLM agents |
0.8 | 1 | 2024 | VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024 |
Computer vision › Video understanding and tracking
long video understanding |
0.8 | 1 | 2024 | VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024 |
Natural language and speech › Language models and text generation › LLM agents
video agent |
0.8 | 1 | 2024 | VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024 |
Machine learning › Learning theory
model selection |
0.7 | 1 | 2023 | LOVM: Language-Only Vision Model Selection · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection
objectness estimation |
0.7 | 1 | 2023 | PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023 |
Computer vision › Image recognition and object detection › object detection
open-world object detection |
0.7 | 1 | 2023 | PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
performance prediction |
0.7 | 1 | 2023 | LOVM: Language-Only Vision Model Selection · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection › open-world object detection
unknown object detection |
0.7 | 1 | 2023 | PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023 |
Computer vision › Vision and language
vision-language model |
0.7 | 1 | 2023 | LOVM: Language-Only Vision Model Selection · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
0.3 | 1 | 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
weak supervision · 0.9vision transformer · 0.9video sampling · 0.9self-training · 0.9scaling consistency analysis · 0.9large multimodal model · 0.9diffusion transformer · 0.9autoencoding · 0.9architecture search · 0.9large language model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Apollo: An Exploration of Video Understanding in Large Multimodal ModelsabstractDespite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), what drives their video perception remains poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models and limited open research hinder the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs. We begin by critically examining the primary contributors to the high computational requirements associated with video-LMM research and discover Scaling Consistency, wherein design and training decisions made on smaller models and datasets (up to a critical size) effectively transfer to larger models. Leveraging these insights, we explored many video-specific aspects of video-LMMs, including video sampling, architectures, data composition, training schedules, and more. Guided by these findings, we introduce Apollo, a state-of-the-art family of LMMs that achieve superior performance across different model sizes. Our models process over 1-hour videos efficiently, with the 3B parameter variant outperforming most existing 7B models. Apollo-7B is state-of-the-art compared to 7B LMMs with a 70.9 on MLVU, and 63.3 on Video-MME. Orr Zohar, Yann Dubois, Nikhil Mehta 0002, Tong Xiao 0003, Philippe Hansen-Estruch, Licheng Yu, Felix Juefei-Xu, Serena Yeung-Levy, Xide Xia |
CVPR | 1 |
| 2025 | Video-STaR: Self-Training Enables Video Instruction Tuning with Any SupervisionabstractThe performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused.
Meanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial.
Herein, we present $\underline{\text{Video}}$ $\underline{\text{S}}\text{elf}$-$\underline{\text{T}}\text{raining}$ $\text{with}$ $\underline{\text{a}}\text{ugmented}$ $\underline{\text{R}}\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning.
Video-STaR allows the utilization of *any* labeled video dataset for video instruction tuning.
In Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision.
During instruction generation, an LMM is prompted to propose an answer. The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset.
By training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning.
Our results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability. Orr Zohar, Yonatan Bitton, Idan Szpektor, Serena Yeung-Levy |
ICLR | 1 |
| 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and GenerationabstractVisual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions remain about how auto-encoder design impacts reconstruction and downstream generative performance. This work explores scaling in auto-encoders for reconstruction and generation by replacing the convolutional backbone with an enhanced Vision Transformer for Tokenization (ViTok). We find scaling the auto-encoder bottleneck correlates with reconstruction but exhibits a nuanced relationship with generation. Separately, encoder scaling yields no gains, while decoder scaling improves reconstruction with minimal impact on generation. As a result, we determine that scaling the current paradigm of auto-encoders is not effective for improving generation performance. Coupled with Diffusion Transformers, ViTok achieves competitive image reconstruction and generation performance on 256p and 512p ImageNet-1K. In videos, ViTok achieves SOTA reconstruction and generation performance on 16-frame 128p UCF-101. Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang 0001, Tingbo Hou, Sriram Vishwanath, Peter Vajda, Xinlei Chen |
ICML | 4 |
| 2024 | VideoAgent: Long-Form Video Understanding with Large Language Model as Agent
Orr Zohar, Serena Yeung-Levy |
ECCV (80) | 3 |
| 2023 | PROB: Probabilistic Objectness for Open World Object DetectionabstractOpen World Object Detection (OWOD) is a new and challenging computer vision task that bridges the gap between classic object detection (OD) benchmarks and object detection in the real world. In addition to detecting and classifying seen/labeled objects, OWOD algorithms are expected to detect novel/unknown objects - which can be classified and incrementally learned. In standard OD, object proposals not overlapping with a labeled object are automatically classified as background. Therefore, simply applying OD methods to OWOD fails as unknown objects would be predicted as background. The challenge of detecting unknown objects stems from the lack of supervision in distinguishing unknown objects and background object proposals. Previous OWOD methods have attempted to overcome this issue by generating supervision using pseudo-labeling - however, unknown object detection has remained low. Probabilistic/generative models may provide a solution for this challenge. Herein, we introduce a novel probabilistic framework for objectness estimation, where we alternate between probability distribution estimation and objectness likelihood maximization of known objects in the embedded feature space - ultimately allowing us to estimate the objectness probability of different proposals. The resulting Probabilistic Objectness transformer-based open-world detector, PROB, integrates our framework into traditional object detection models, adapting them for the open-world setting. Comprehensive experiments on OWOD benchmarks show that PROB outperforms all existing OWOD methods in both unknown object detection (~ 2 × unknown recall) and known object detection (~ 10% mAP). Our code is available at https://github.com/orrzohar/PROB. Orr Zohar, Kuan-Chieh Wang, Serena Yeung-Levy |
CVPR | 1 |
| 2023 | LOVM: Language-Only Vision Model SelectionabstractPre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot settings. However, selecting the best-performing VLM for some downstream applications is non-trivial, as it is dataset and task-dependent. Meanwhile, the exhaustive evaluation of all available VLMs on a novel application is not only time and computationally demanding but also necessitates the collection of a labeled dataset for evaluation. As the number of open-source VLM variants increases, there is a need for an efficient model selection strategy that does not require access to a curated evaluation dataset. This paper proposes a novel task and benchmark for efficiently evaluating VLMs' zero-shot performance on downstream applications without access to the downstream task dataset. Specifically, we introduce a new task LOVM: Language-Only Vision Model Selection , where methods are expected to perform both model selection and performance prediction based solely on a text description of the desired downstream application. We then introduced an extensive LOVM benchmark consisting of ground-truth evaluations of 35 pre-trained VLMs and 23 datasets, where methods are expected to rank the pre-trained VLMs and predict their zero-shot performance. Orr Zohar, Shih-Cheng Huang, Kuan-Chieh Wang, Serena Yeung-Levy |
NeurIPS | 1 |