EDBT 2026 Demo / reviewers in the wild / expert
Yogesh Virkar
dblp:277/3561
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SWAN: SubWord Alignment Network for HMM-free word timing estimation in end-to-end automatic speech recognition
Woo Hyun Kang, Srikanth Vishnubhotla, Rudolf Braun, Yogesh Virkar, Raghuveer Peri, Kyu J. Han |
INTERSPEECH | 4 |
| 2023 | Improving Isochronous Machine Translation with Target Factors and Auxiliary Counters
Proyag Pal, Brian Thompson 0001, Yogesh Virkar, Prashant Mathur, Alexandra Chronopoulou, Marcello Federico |
INTERSPEECH | 3 |
| 2023 | Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic DubbingabstractAbstract We investigate how humans perform the task of dubbing video content from one language into another, leveraging a novel corpus of 319.57 hours of video from 54 professionally produced titles. This is the first such large-scale study we are aware of. The results challenge a number of assumptions commonly made in both qualitative literature on human dubbing and machine-learning literature on automatic dubbing, arguing for the importance of vocal naturalness and translation quality over commonly emphasized isometric (character length) and lip-sync constraints, and for a more qualified view of the importance of isochronic (timing) constraints. We also find substantial influence of the source-side audio on human dubs through channels other than the words of the translation, pointing to the need for research on ways to preserve speech characteristics, as well as transfer of semantic properties such as emphasis and emotion, in automatic dubbing systems. William Brannon, Yogesh Virkar, Brian Thompson 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | Duration Modeling of Neural TTS for Automatic DubbingabstractAutomatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our automatic dubbing system, isochrony is modeled by controlling the verbosity of machine translation; inserting pauses in the translations, a.k.a. prosodic alignment; and controlling the duration of text-to-speech (TTS) utterances. The latter two steps heavily rely on speech duration information, either to predict or control TTS duration. So far, duration prediction was based on a proxy method while duration control on linear warping of the TTS speech spectrogram. In this study, we propose novel duration models for neural TTS that can be leveraged both to predict and control TTS duration. Experimental results show that compared to previous work, the new models improve or match the performance of prosodic alignment and significantly enhance neural TTS speech quality for both slow and fast speaking rates. Johanes Effendi, Yogesh Virkar, Roberto Barra-Chicote, Marcello Federico |
ICASSP | 2 |
| 2022 | ISOMETRIC MT: Neural Machine Translation for Automatic DubbingabstractAutomatic dubbing (AD) is among the machine translation (MT) use cases where translations should match a given length to allow for synchronicity between source and target speech. For neural MT, generating translations of length close to the source length (e.g. within ±10% in character count), while preserving quality is a challenging task. Controlling MT output length comes at a cost to translation quality, which is usually mitigated with a two step approach of generating N-best hypotheses and then re-ranking based on length and quality. This work introduces a self-learning approach that allows a transformer model to directly learn to generate outputs that closely match the source length, in short Isometric MT. In particular, our approach does not require to generate multiple hypotheses nor any auxiliary ranking function. We report results on four language pairs (English → French, Italian, German, Spanish) with a publicly available benchmark. Automatic and manual evaluations show that our method for Isometric MT outperforms more complex approaches proposed in the literature. Surafel Melaku Lakew, Yogesh Virkar, Prashant Mathur, Marcello Federico |
ICASSP | 2 |
| 2022 | Isochrony-Aware Neural Machine Translation for Automatic DubbingabstractWe introduce the task of isochrony-aware machine translation which aims at generating translations suitable for dubbing.Dubbing of a spoken sentence requires transferring the content as well as the speech-pause structure of the source into the target language to achieve audiovisual coherence.Practically, this implies correctly projecting pauses from the source to the target and ensuring that target speech segments have roughly the same duration of the corresponding source speech segments.In this work, we propose implicit and explicit modeling approaches to integrate isochrony information into neural machine translation.Experiments on English-German/French language pairs with automatic metrics show that the simplest of the considered approaches works best.Results are confirmed by human evaluations of translations and dubbed videos. Derek Tam, Surafel Melaku Lakew, Yogesh Virkar, Prashant Mathur, Marcello Federico |
INTERSPEECH | 3 |
| 2022 | Prosodic alignment for off-screen automatic dubbingabstractThe goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence.This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses, especially when the speaker's mouth is visible.In previous work, we introduced a prosodic alignment model to address isochrone or on-screen dubbing.In this work, we extend the prosodic alignment model to also address off-screen dubbing that requires less stringent synchronization constraints.We conduct experiments on four dubbing directions -English to French, Italian, German and Spanish -on a publicly available collection of TED Talks and on publicly available YouTube videos.Empirical results show that compared to our previous work the extended prosodic alignment model provides significantly better subjective viewing experience on videos in which on-screen and off-screen automatic dubbing is applied for sentences with speakers mouth visible and not visible, respectively. Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote |
INTERSPEECH | 1 |
| 2021 | Machine Translation Verbosity Control for Automatic DubbingabstractAutomatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterances. In this paper, we focus on the problem of controlling the verbosity of machine translation out-put, so that subsequent steps of our automatic dubbing pipeline can generate dubs of better quality. We propose new methods to control the verbosity of MT output and compare them against the state of the art with both intrinsic and extrinsic evaluations. For our experiments we use a public data set to dub English speeches into French, Italian, German and Spanish. Finally, we report extensive subjective tests that measure the impact of MT verbosity control on the final quality of dubbed video clips. Surafel Melaku Lakew, Marcello Federico, Yue Wang 0034, Cuong Hoang, Yogesh Virkar, Roberto Barra-Chicote, Robert Enyedi |
ICASSP | 5 |
| 2021 | Improvements to Prosodic Alignment for Automatic DubbingabstractAutomatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.e., arranging the translation of the original speech to optimally match its sequence of phrases and pauses. To this end, we present improvements to the prosodic alignment component of our recently introduced dubbing architecture. We present empirical results for four dubbing directions – English to French, Italian, German and Spanish – on a publicly available collection of TED Talks. Compared to previous work, our enhanced prosodic alignment model significantly improves prosodic alignment accuracy and provides segmentation perceptibly better or on par with manually annotated reference segmentation. Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote |
ICASSP | 1 |
| 2021 | Intra-Sentential Speaking Rate Control in Neural Text-To-Speech for Automatic Dubbing
Yogesh Virkar, Marcello Federico, Roberto Barra-Chicote, Robert Enyedi |
Interspeech | 2 |
| 2020 | Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing
Marcello Federico, Yogesh Virkar, Robert Enyedi, Roberto Barra-Chicote |
INTERSPEECH | 2 |