Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Michael Finkelson

dblp:405/3115 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › audio generation
text-to-audio generation
0.912025
CAFA: A Controllable Automatic Foley Artist · ICCV 2025
Audio and music processing › sound synthesis › video-to-audio generation
foley sound generation
0.912025
CAFA: A Controllable Automatic Foley Artist · ICCV 2025
Audio and music processing › sound synthesis
video-to-audio generation
0.912025
CAFA: A Controllable Automatic Foley Artist · ICCV 2025

Methods — techniques the papers use, named apart from their topics

text-to-audio model · 1.7modality adapter · 1.7diffusion model · 1.7
YearPublicationVenuePosition
2025 CAFA: A Controllable Automatic Foley Artist
abstract
Foley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in automatic video-to-audio models have increased the demand for greater user control in the process. One possible approach is to incorporate text to guide audio generation. While supported by existing methods, challenges remain in ensuring compatibility between modalities, particularly when the text introduces additional information or contradicts the sounds naturally inferred from the visuals. In this work, we introduce CAFA (Controllable Automatic Foley Artist) a video-and-text-to-audio model that generates semantically and temporally aligned audio for a given video, guided by text input. CAFA is built upon a text-to-audio model and integrates video information through a modality adapter mechanism. By incorporating text, users can refine semantic details and introduce creative variations, guiding the audio synthesis beyond the expected video contextual cues. Experiments show that besides its superior quality in terms of semantic alignment and audio-visual synchronization the proposed method enable high textual controllability as demonstrated in subjective and objective evaluations.
Roi Benita, Michael Finkelson, Tavi Halperin, Gleb Sterkin, Yossi Adi
ICCV2