Orr Zohar

dblp:335/1624 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-9091-7140ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 27% Language models and text generation · 20% Vision and language · 15%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025
Machine learning › Generative modeling
image generation
0.912025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025
Machine learning › Generative modeling
image tokenization
0.912025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025
Natural language and speech › Language models and text generation
multimodal language model
0.912025
Apollo: An Exploration of Video Understanding in Large Multimodal Models · CVPR 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.912025
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025
Machine learning › Generative modeling
video generation
0.912025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025
Computer vision › Vision and language
video-language model
0.912025
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
video multimodal large language model
0.912025
Apollo: An Exploration of Video Understanding in Large Multimodal Models · CVPR 2025
Computer vision › Video understanding and tracking
video question answering
0.912025
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision · ICLR 2025
Natural language and speech › Language models and text generation
large language model
0.812024
VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024
Computer vision › Video understanding and tracking
long video understanding
0.812024
VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024
Natural language and speech › Language models and text generation › LLM agents
video agent
0.812024
VideoAgent: Long-Form Video Understanding with Large Language Model as Agent · ECCV (80) 2024
Machine learning › Learning theory
model selection
0.712023
LOVM: Language-Only Vision Model Selection · NeurIPS 2023
Computer vision › Image recognition and object detection › object detection
objectness estimation
0.712023
PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023
Computer vision › Image recognition and object detection › object detection
open-world object detection
0.712023
PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
performance prediction
0.712023
LOVM: Language-Only Vision Model Selection · NeurIPS 2023
Computer vision › Image recognition and object detection › object detection › open-world object detection
unknown object detection
0.712023
PROB: Probabilistic Objectness for Open World Object Detection · CVPR 2023
Computer vision › Vision and language
vision-language model
0.712023
LOVM: Language-Only Vision Model Selection · NeurIPS 2023
Machine learning › Efficient and distributed learning › large-scale learning
model scaling
0.312025
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation · ICML 2025

Methods — techniques the papers use, named apart from their topics

weak supervision · 0.9vision transformer · 0.9video sampling · 0.9self-training · 0.9scaling consistency analysis · 0.9large multimodal model · 0.9diffusion transformer · 0.9autoencoding · 0.9architecture search · 0.9large language model · 0.8
YearPublicationVenuePosition
2025 Apollo: An Exploration of Video Understanding in Large Multimodal Models
abstract
Despite the rapid integration of video perception capabilities into Large Multimodal Models (LMMs), what drives their video perception remains poorly understood. Consequently, many design decisions in this domain are made without proper justification or analysis. The high computational cost of training and evaluating such models and limited open research hinder the development of video-LMMs. To address this, we present a comprehensive study that helps uncover what effectively drives video understanding in LMMs. We begin by critically examining the primary contributors to the high computational requirements associated with video-LMM research and discover Scaling Consistency, wherein design and training decisions made on smaller models and datasets (up to a critical size) effectively transfer to larger models. Leveraging these insights, we explored many video-specific aspects of video-LMMs, including video sampling, architectures, data composition, training schedules, and more. Guided by these findings, we introduce Apollo, a state-of-the-art family of LMMs that achieve superior performance across different model sizes. Our models process over 1-hour videos efficiently, with the 3B parameter variant outperforming most existing 7B models. Apollo-7B is state-of-the-art compared to 7B LMMs with a 70.9 on MLVU, and 63.3 on Video-MME.
Orr Zohar, Yann Dubois, Nikhil Mehta 0002, Tong Xiao 0003, Philippe Hansen-Estruch, Licheng Yu, Felix Juefei-Xu, Serena Yeung-Levy, Xide Xia
CVPR1
2025 Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
abstract
The performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused. Meanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial. Herein, we present $\underline{\text{Video}}$ $\underline{\text{S}}\text{elf}$-$\underline{\text{T}}\text{raining}$ $\text{with}$ $\underline{\text{a}}\text{ugmented}$ $\underline{\text{R}}\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning. Video-STaR allows the utilization of *any* labeled video dataset for video instruction tuning. In Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision. During instruction generation, an LMM is prompted to propose an answer. The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset. By training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning. Our results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability.
Orr Zohar, Yonatan Bitton, Idan Szpektor, Serena Yeung-Levy
ICLR1
2025 Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
abstract
Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions remain about how auto-encoder design impacts reconstruction and downstream generative performance. This work explores scaling in auto-encoders for reconstruction and generation by replacing the convolutional backbone with an enhanced Vision Transformer for Tokenization (ViTok). We find scaling the auto-encoder bottleneck correlates with reconstruction but exhibits a nuanced relationship with generation. Separately, encoder scaling yields no gains, while decoder scaling improves reconstruction with minimal impact on generation. As a result, we determine that scaling the current paradigm of auto-encoders is not effective for improving generation performance. Coupled with Diffusion Transformers, ViTok achieves competitive image reconstruction and generation performance on 256p and 512p ImageNet-1K. In videos, ViTok achieves SOTA reconstruction and generation performance on 16-frame 128p UCF-101.
Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang 0001, Tingbo Hou, Sriram Vishwanath, Peter Vajda, Xinlei Chen
ICML4
2024 VideoAgent: Long-Form Video Understanding with Large Language Model as Agent
Orr Zohar, Serena Yeung-Levy
ECCV (80)3
2023 PROB: Probabilistic Objectness for Open World Object Detection
abstract
Open World Object Detection (OWOD) is a new and challenging computer vision task that bridges the gap between classic object detection (OD) benchmarks and object detection in the real world. In addition to detecting and classifying seen/labeled objects, OWOD algorithms are expected to detect novel/unknown objects - which can be classified and incrementally learned. In standard OD, object proposals not overlapping with a labeled object are automatically classified as background. Therefore, simply applying OD methods to OWOD fails as unknown objects would be predicted as background. The challenge of detecting unknown objects stems from the lack of supervision in distinguishing unknown objects and background object proposals. Previous OWOD methods have attempted to overcome this issue by generating supervision using pseudo-labeling - however, unknown object detection has remained low. Probabilistic/generative models may provide a solution for this challenge. Herein, we introduce a novel probabilistic framework for objectness estimation, where we alternate between probability distribution estimation and objectness likelihood maximization of known objects in the embedded feature space - ultimately allowing us to estimate the objectness probability of different proposals. The resulting Probabilistic Objectness transformer-based open-world detector, PROB, integrates our framework into traditional object detection models, adapting them for the open-world setting. Comprehensive experiments on OWOD benchmarks show that PROB outperforms all existing OWOD methods in both unknown object detection (~ 2 × unknown recall) and known object detection (~ 10% mAP). Our code is available at https://github.com/orrzohar/PROB.
Orr Zohar, Kuan-Chieh Wang, Serena Yeung-Levy
CVPR1
2023 LOVM: Language-Only Vision Model Selection
abstract
Pre-trained multi-modal vision-language models (VLMs) are becoming increasingly popular due to their exceptional performance on downstream vision applications, particularly in the few- and zero-shot settings. However, selecting the best-performing VLM for some downstream applications is non-trivial, as it is dataset and task-dependent. Meanwhile, the exhaustive evaluation of all available VLMs on a novel application is not only time and computationally demanding but also necessitates the collection of a labeled dataset for evaluation. As the number of open-source VLM variants increases, there is a need for an efficient model selection strategy that does not require access to a curated evaluation dataset. This paper proposes a novel task and benchmark for efficiently evaluating VLMs' zero-shot performance on downstream applications without access to the downstream task dataset. Specifically, we introduce a new task LOVM: Language-Only Vision Model Selection , where methods are expected to perform both model selection and performance prediction based solely on a text description of the desired downstream application. We then introduced an extensive LOVM benchmark consisting of ground-truth evaluations of 35 pre-trained VLMs and 23 datasets, where methods are expected to rank the pre-trained VLMs and predict their zero-shot performance.
Orr Zohar, Shih-Cheng Huang, Kuan-Chieh Wang, Serena Yeung-Levy
NeurIPS1