Ariel Shaulov

dblp:356/3859 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 60% Video understanding and tracking · 20% 3D vision · 20%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation · NeurIPS 2025
Computer vision › 3D vision › motion perception
motion coherence
0.912025
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation · NeurIPS 2025
Computer vision › Video understanding and tracking › temporal localization
temporal event localization
0.912025
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds · CVPR 2025
Machine learning › Generative modeling › video generation
text-to-video generation
0.912025
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance
0.912025
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation · NeurIPS 2025
Multimedia analysis and retrieval › audio-visual learning
audio-visual event perception
0.912025
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds · CVPR 2025

Methods — techniques the papers use, named apart from their topics

score-level fusion · 1.7dynamic thresholding · 1.7variance-based guidance · 0.9temporal representation · 0.9
YearPublicationVenuePosition
2025 Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
abstract
In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their training data. This limitation significantly impedes their capacity to generalize to novel, unseen event categories. Furthermore, the annotation process for this task is labor-intensive, requiring extensive manual labeling across modalities and temporal segments, limiting the scalability of current methods. Current state-of-the-art models ignore the shifts in event distributions over time, reducing their ability to adjust to changing video dynamics. Additionally, previous methods rely on late fusion to combine audio and visual information. While straightforward, this approach results in a significant loss of multimodal interactions. To address these challenges, we propose Audio-Visual Adaptive Video Analysis (AV2A), a model-agnostic approach that requires no further training and integrates a score-level fusion technique to retain richer multimodal interactions. AV2A also includes a within-video label shift algorithm, leveraging input video data and predictions from prior frames to dynamically adjust event distributions for subsequent frames. Moreover, we present the first training-free, open-vocabulary baseline for audio-visual event perception, demonstrating that AV2A achieves substantial improvements over naive training-free baselines. We demonstrate the effectiveness of AV2A on both zero-shot and weakly-supervised state-of-the-art methods, achieving notable improvements in performance metrics over existing approaches. Our code is available on Github.
Eitan Shaar, Ariel Shaulov, Gal Chechik, Lior Wolf
CVPR2
2025 Classifier-Guided Captioning Across Modalities
abstract
Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This limitation hinders performance in tasks like audio or video captioning, where different semantic cues are needed. Addressing this challenge is crucial for creating more adaptable and versatile captioning frameworks applicable across diverse real-world contexts. In this work, we introduce a method to adapt captioning networks to the semantics of alternative settings, such as capturing audibility in audio captioning, where it is crucial to describe sounds and their sources. Our framework consists of two main components: (i) a frozen captioning system incorporating a language model (LM), and (ii) a text classifier that guides the captioning system. The classifier is trained on a dataset automatically generated by GPT-4, using tailored prompts specifically designed to enhance key aspects of the generated captions. Importantly, the framework operates solely during inference, eliminating the need for further training of the underlying captioning model. We evaluated the framework on various models and modalities, with a focus on audio captioning, and report promising results. Notably, when combined with an existing zero-shot audio captioning system, our framework improves its quality and sets state-of-the-art performance in zero-shot audio captioning.
Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik, Lior Wolf
ICASSP1
2025 FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
abstract
Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external conditioning signals to enforce temporal consistency. In this work, we explore whether a meaningful temporal representation can be extracted directly from the predictions of a pre-trained model without any additional training or auxiliary inputs. We introduce __FlowMo__, a novel training-free guidance method that enhances motion coherence using only the model's own predictions in each diffusion step. FlowMo first derives an appearance-debiased temporal representation by measuring the distance between latents corresponding to consecutive frames. This highlights the implicit temporal structure predicted by the model. It then estimates motion coherence by measuring the patch-wise variance across the temporal dimension, and guides the model to reduce this variance dynamically during sampling. Extensive experiments across multiple text-to-video models demonstrate that FlowMo significantly improves motion coherence without sacrificing visual quality or prompt alignment, offering an effective plug-and-play solution for enhancing the temporal fidelity of pre-trained video diffusion models.
Ariel Shaulov, Itay Hazan 0001, Lior Wolf, Hila Chefer
NeurIPS1