VLDB 2026 Research / reviewers in the wild / expert
Fadi Biadsy
dblp:38/6162
· DBLP profile ↗
29ranked-venue papers
13as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 12 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 10 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Streaming Parrotron for on-device speech-to-speech conversion
Oleg Rybakov, Fadi Biadsy, Liyang Jiang, Phoenix Meadowlark, Shivani Agrawal |
INTERSPEECH | 2 |
| 2023 | Real time spectrogram inversion on mobile phone
Oleg Rybakov, Marco Tagliasacchi, Yunpeng Li 0008, Liyang Jiang, Fadi Biadsy |
INTERSPEECH | 6 |
| 2022 | A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model PersonalizationabstractModel fine-tuning and adaptation have become a common approach for model specialization for downstream tasks or domains. Fine-tuning the entire model or a subset of the parameters using light-weight adaptation has shown considerable success across different specialization tasks. Fine-tuning a model for a large number of domains typically requires starting a new training job for every domain posing scaling limitations. Once these models are trained, deploying them also poses significant scalability challenges for inference for real-time applications. In this paper, building upon prior light-weight adaptation techniques, we propose a modular framework that enables us to substantially improve scalability for model training and inference. We introduce Submodels that can be quickly and dynamically loaded for on-the-fly inference. We also propose multiple approaches for training those Submodels in parallel using an embedding space in the same training job. We test our framework on an extreme use-case which is speech model personalization for atypical speech, requiring a Submodel for each user. We obtain 128x Submodel throughput with a fixed computation budget without a loss of accuracy. We also show that learning a speaker-embedding space can scale further and reduce the amount of personalization training data required per speaker. Fadi Biadsy, Youzheng Chen, Oleg Rybakov, Andrew Rosenberg, Pedro J. Moreno 0001 |
INTERSPEECH | 1 |
| 2022 | Non-Parallel Voice Conversion for ASR AugmentationabstractAutomatic speech recognition (ASR) needs to be robust to speaker differences.Voice Conversion (VC) modifies speaker characteristics of input speech.This is an attractive feature for ASR data augmentation.In this paper, we demonstrate that voice conversion can be used as a data augmentation technique to improve ASR performance, even on LibriSpeech, which contains 2,456 speakers.For ASR augmentation, it is necessary that the VC model be robust to a wide range of input speech.This motivates the use of a non-autoregressive, non-parallel VC model, and the use of a pretrained ASR encoder within the VC model.This work suggests that despite including many speakers, speaker diversity may remain a limitation to ASR quality.Finally, interrogation of our VC performance has provided useful metrics for objective evaluation of VC quality. Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy, Jesse Emond, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2021 | Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented SpeechabstractAutomatic Speech Recognition (ASR) systems are often optimized to work best for speakers with canonical speech patterns.Unfortunately, these systems perform poorly when tested on atypical speech and heavily accented speech.It has previously been shown that personalization through model fine-tuning substantially improves performance.However, maintaining such large models per speaker is costly and difficult to scale.We show that by adding a relatively small number of extra parameters to the encoder layers via socalled residual adapter, we can achieve similar adaptation gains compared to model finetuning, while only updating a tiny fraction (less than 0.5%) of the model parameters.We demonstrate this on two speech adaptation tasks (atypical and accented speech) and for two state-of-the-art ASR architectures. Katrin Tomanek, Victoria Zayats, Dirk Padfield, Kara Vaillancourt, Fadi Biadsy |
EMNLP (1) | 5 |
| 2021 | Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical SpeechabstractWe present an extended Parrotron model: a single, end-to-end network that enables voice conversion and recognition simultaneously. Input spectrograms are transformed to output spectrograms in the voice of a predetermined target speaker while also generating hypotheses in a target vocabulary. We study the performance of this novel architecture, which jointly predicts speech and text, on atypical (e.g. dysarthric) speech. We show that with as little as an hour of atypical speech, speaker adaptation can yield a 77% relative reduction in Word Error Rate (WER), measured by ASR performance on the converted speech. We also show that data augmentation using a customized synthesizer built on atypical speech can provide an additional 10% relative improvement over the best speaker-adapted model. Finally, we show how these methods generalize across 8 types of atypical speech for a range of speech impairment severities. Rohan Doshi, Youzheng Chen, Liyang Jiang, Fadi Biadsy, Bhuvana Ramabhadran, Fang Chu, Andrew Rosenberg, Pedro J. Moreno 0001 |
ICASSP | 5 |
| 2021 | Conformer Parrotron: A Faster and Stronger End-to-End Speech Conversion and Recognition Model for Atypical Speech
Zhehuai Chen, Bhuvana Ramabhadran, Fadi Biadsy, Youzheng Chen, Liyang Jiang, Fang Chu, Rohan Doshi, Pedro J. Moreno 0001 |
Interspeech | 3 |
| 2019 | Comparison of Data Augmentation and Adaptation Strategies for Code-switched Automatic Speech RecognitionabstractCode-switching occurs when the speaker alternates between two or more languages or dialects. It is a pervasive phenomenon in most Indic spoken languages. Code-switching poses a challenge in language modeling as it complicates the orthographic realization of text, and generally, there is a shortage of code-switched data. In this paper, we investigate data augmentation and adaptation strategies for language modeling. Using Bengali and English as an example, we study augmenting the code-switched transcripts with separate transliterated Bengali and English corpora. We present results on two speech recognition tasks, namely, voice search and dictation. We show improvements on both tasks with Maximum Entropy (MaxEnt) and Long Short-Term Memory (LSTM) language models (LMs). We also explore different adaptation strategies for MaxEnt LM and LSTM LM, demonstrating that the transliteration-based data-augmented LSTM LM matches the adapted MaxEnt LM which is trained on more Bengali-English data. Bhuvana Ramabhadran, Jesse Emond, Andrew Rosenberg, Fadi Biadsy |
ICASSP | 5 |
| 2019 | Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech SeparationabstractWe describe Parrotron, an end-to-end-trained speech-to-speech conversion model that maps an input spectrogram directly to another spectrogram, without utilizing any intermediate discrete representation.The network is composed of an encoder, spectrogram and phoneme decoders, followed by a vocoder to synthesize a time-domain waveform.We demonstrate that this model can be trained to normalize speech from any speaker regardless of accent, prosody, and background noise, into the voice of a single canonical target speaker with a fixed accent and consistent articulation and prosody.We further show that this normalization model can be adapted to normalize highly atypical speech from a deaf speaker, resulting in significant improvements in intelligibility and naturalness, measured via a speech recognizer and listening tests.Finally, demonstrating the utility of this model on other speech tasks, we show that the same model architecture can be trained to perform a speech separation task. Fadi Biadsy, Ron J. Weiss, Pedro J. Moreno 0001, Dimitri Kanvesky, Ye Jia |
INTERSPEECH | 1 |
| 2019 | Direct Speech-to-Speech Translation with a Sequence-to-Sequence ModelabstractWe present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation.The network is trained end-to-end, learning to map speech spectrograms into target spectrograms in another language, corresponding to the translated content (in a different canonical voice).We further demonstrate the ability to synthesize translated speech using the voice of the source speaker.We conduct experiments on two Spanish-to-English speech translation datasets, and find that the proposed model slightly underperforms a baseline cascade of a direct speech-to-text translation model and a text-to-speech synthesis model, demonstrating the feasibility of the approach on this very challenging task. Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson |
INTERSPEECH | 3 |
| 2018 | Modeling Non-Linguistic Contextual Signals in LSTM Language Models Via Domain AdaptationabstractLanguage Models (LMs) for Automatic Speech Recognition (ASR) can benefit from utilizing non-linguistic contextual signals in modeling. Examples of these signals include the geographical location of the user speaking to the system and/or the identity of the application (app) being spoken to. In practice, the vast majority of input speech queries typically lack annotations of such signals, which poses a challenge to directly train domain-specific LMs. To obtain robust domain LMs, generally an LM which has been pre-trained on general data will be adapted to specific domains. We propose four domain adaptation schemes to improve the domain performance of Long Short-Term Memory (LSTM) LMs, by incorporating app based contextual signals of voice search queries. We show that most of our adaptation strategies are effective, reducing word perplexity up to 21 % relative to a fine-tuned baseline on a held-out domain-specific development set. Initial experiments using a state-of-the-art Italian ASR system show a 3 % relative reduction in WER on top of an unadapted 5-gram LM. In addition, human evaluations show significant improvements on sub-domains from using app signals. Shankar Kumar, Fadi Biadsy, Michael Nirschl, Tomas Vykruta, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2017 | Effectively Building Tera Scale MaxEnt Language Models Incorporating Non-Linguistic Signals
Fadi Biadsy, Mohammadreza Ghodsi, Diamantino Caseiro |
INTERSPEECH | 1 |
| 2017 | Sparse Non-Negative Matrix Language Modeling: Maximum Entropy Flexibility on the Cheap
Ciprian Chelba, Diamantino Caseiro, Fadi Biadsy |
INTERSPEECH | 3 |
| 2017 | Approaches for Neural-Network Language Model Adaptation
Michael Nirschl, Fadi Biadsy, Shankar Kumar |
INTERSPEECH | 3 |
| 2014 | Backoff inspired features for maximum entropy language modelsabstractMaximum Entropy (MaxEnt) language models [1, 2] are linear models that are typically regularized via well-known L1 or L2 terms in the likelihood objective, hence avoiding the need for the kinds of backoff or mixture weights used in smoothed n-gram language models using Katz backoff [3] and similar tech-niques. Even though backoff cost is not required to regularize the model, we investigate the use of backoff features in Max-Ent models, as well as some backoff-inspired variants. These features are shown to improve model quality substantially, as shown in perplexity and word-error rate reductions, even in very large scale training scenarios of tens or hundreds of billions of words and hundreds of millions of features. Index Terms: maximum entropy modeling, language model-ing, n-gram models, linear models Fadi Biadsy, Keith B. Hall, Pedro J. Moreno 0001, Brian Roark |
INTERSPEECH | 1 |
| 2013 | Automatic detection of speaker state: Lexical, prosodic, and phonetic approaches to level-of-interest and intoxication classification
William Yang Wang, Fadi Biadsy, Andrew Rosenberg, Julia Hirschberg |
Comput. Speech Lang. | 2 |
| 2012 | Google's cross-dialect Arabic voice searchabstractWe present a large scale effort to build a commercial Automatic Speech Recognition (ASR) product for Arabic. Our goal is to support voice search, dictation, and voice control for the general Arabic-speaking public, including support for multiple Arabic dialects. We describe our ASR system design and compare recognizers for five Arabic dialects, with the potential to reach more than 125 million people in Egypt, Jordan, Lebanon, Saudi Arabia, and the United Arab Emirates (UAE). We compare systems built on diacritized vs. non-diacritized text. We also conduct cross-dialect experiments, where we train on one dialect and test on the others. Our average word error rate (WER) is 24.8% for voice search. Fadi Biadsy, Pedro J. Moreno 0001, Martin Jansche |
ICASSP | 1 |
| 2012 | On-the-fly Topic Adaptation for YouTube Video TranscriptionabstractAutomatic closed-captioning of video is a useful application of speech recognition technology but poses numerous challenges when applied to open-domain user-uploaded videos such as those on YouTube.In this work, we explore a strategy to improve decoding accuracy for video transcription by decoding each video with a language model (LM) adapted specifically to the topics that the video covers.Taxonomic topic classifiers are used to determine the topic content of videos and to build a large set of topic-specific LMs from web documents.We consider strategies for selecting and interpolating LMs in both supervised and unsupervised scenarios in a two-pass lattice rescoring framework.Experiments on a YouTube video corpus show a 3.6 absolute reduction in WER over generic single-pass transcriptions as well as a statistically significant 0.8 absolute improvement over rescoring with a very large non-adapted LM built from all the documents. Kapil Thadani, Fadi Biadsy, Dan Bikel |
INTERSPEECH | 2 |
| 2011 | The IBM 2011 GALE Arabic speech transcription systemabstractWe describe the Arabic broadcast transcription system fielded by IBM in the GALE Phase 5 machine translation evaluation. Key advances over our Phase 4 system include a new Bayesian Sensing HMM acoustic model; multistream neural network features; a MADA vowelized acoustic model; and the use of a variety of language model techniques with significant additive gains. These advances were instrumental in achieving a word error rate of 7.4% on the Phase 5 evaluation set, and an absolute improvement of 0.9% word error rate over our 2009 system on the unsequestered Phase 4 evaluation data. Lidia Mangu, Hong-Kwang Jeff Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy |
ASRU | 7 |
| 2011 | From Modern Standard Arabic to Levantine ASR: Leveraging GALE for dialectsabstractWe report a series of experiments about how we can progress from Modern Standard Arabic (MSA) to Levantine ASR, in the context of the GALE DARPA program. While our GALE models achieved very low error rates, we still see error rates twice as high when decoding dialectal data. In this paper, we make use of a state-of-the-art Arabic dialect recognition system to automatically identify Levantine and MSA subsets in mixed speech of a variety of dialects including MSA. Training separate models on these subsets, we show a significant reduction in word error rate over using the entire data set to train one system for both dialects. During decoding, we use a tree array structure to mix Levantine and MSA models automatically using the posterior probabilities of the dialect classifier as soft weights. This technique allows us to mix these models without sacrificing performance for either varieties. Furthermore, using the initial acoustic-based dialect recognition system's output, we show that we can bootstrap a text-based dialect classifier and use it to identify relevant text data for building Levantine language models. Moreover, we compare different vowelization approaches when transitioning from MSA to Levantine models. Hagen Soltau, Lidia Mangu, Fadi Biadsy |
ASRU | 3 |
| 2011 | Dialect and Accent Recognition Using Phonetic-Segmentation SupervectorsabstractWe describe a new approach to automatic dialect and accent recognition which exceeds state-of-the-art performance in three recognition tasks.This approach improves the accuracy and substantially lower the time complexity of our earlier phoneticbased kernel approach for dialect recognition.In contrast to state-of-the-art acoustic-based systems, our approach employs phone labels and segmentation to constrain the acoustic models.Given a speaker's utterance, we first obtain phone hypotheses using a phone recognizer and then extract GMM-supervectors for each phone type, effectively summarizing the speaker's phonetic characteristics in a single vector of phone-type supervectors.Using these vectors, we design a kernel function that computes the phonetic similarities between pairs of utterances to train SVM classifiers to identify dialects.Comparing this approach to the state-of-the-art, we obtain a 12.9% relative improvement in EER on Arabic dialects, and a 17.9% relative improvement for American vs. Indian English dialects.We also see a 53.5% relative improvement over a GMM-UBM on American Southern vs. Non-Southern English. Fadi Biadsy, Julia Hirschberg, Daniel P. W. Ellis |
INTERSPEECH | 1 |
| 2011 | Intoxication Detection Using Phonetic, Phonotactic and Prosodic Cues
Fadi Biadsy, William Yang Wang, Andrew Rosenberg, Julia Hirschberg |
INTERSPEECH | 1 |
| 2011 | Segmentation-Free Online Arabic Handwriting RecognitionabstractArabic script is naturally cursive and unconstrained and, as a result, an automatic recognition of its handwriting is a challenging problem. The analysis of Arabic script is further complicated in comparison to Latin script due to obligatory dots/stokes that are placed above or below most letters. In this paper, we introduce a new approach that performs online Arabic word recognition on a continuous word-part level, while performing training on the letter level. In addition, we appropriately handle delayed strokes by first detecting them and then integrating them into the word-part body. Our current implementation is based on Hidden Markov Models (HMM) and correctly handles most of the Arabic script recognition difficulties. We have tested our implementation using various dictionaries and multiple writers and have achieved encouraging results for both writer-dependent and writer-independent recognition. Fadi Biadsy, Raid Saabni, Jihad El-Sana |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2010 | Dialect recognition using a phone-GMM-supervector-based SVM kernelabstractIn this paper, we introduce a new approach to dialect recognition which relies on the hypothesis that certain phones are realized differently across dialects. Given a speaker’s utterance, we first obtain the most likely phone sequence using a phone recognizer. We then extract GMM Supervectors for each phone instance. Using these vectors, we design a kernel function that computes the similarities of phones between pairs of utterances. We employ this kernel to train SVM classifiers that estimate posterior probabilities, used during recognition. Testing our approach on four Arabic dialects from 30s cuts, we compare our performance to five approaches: PRLM; GMM-UBM; our own improved version of GMM-UBM which employs fMLLR adaptation; our recent discriminative phonotactic approach; and a state-of-the-art system: SDC-based GMM-UBM discriminatively trained. Our kernel-based technique outperforms all these previous approaches; the overall EER of our system is 4.9%. Fadi Biadsy, Julia Hirschberg, Michael Collins 0001 |
INTERSPEECH | 1 |
| 2009 | Contextual Phrase-Level Polarity Analysis Using Lexical Affect Scoring and Syntactic N-Grams
Apoorv Agarwal, Fadi Biadsy, Kathy McKeown |
EACL | 2 |
| 2009 | Using prosody and phonotactics in Arabic dialect identificationabstractWhile Modern Standard Arabic is the formal spoken and written language of the Arab world, dialects are the major communication mode for everyday life; identifying a speaker’s dialect is thus critical to speech processing tasks such as automatic speech recognition, as well as speaker identification. We examine the role of prosodic features (intonation and rhythm) across four Arabic dialects: Gulf, Iraqi, Levantine, and Egyptian, for the purpose of automatic dialect identification. We show that prosodic features can significantly improve identification, over a purely phonotactic-based approach, with an identification accuracy of 86.33 % for 2m utterances. 1. Fadi Biadsy, Julia Hirschberg |
INTERSPEECH | 1 |
| 2009 | Improving the Arabic Pronunciation Dictionary for Phone and Word Recognition with Linguistically-Based Pronunciation Rules
Fadi Biadsy, Nizar Habash, Julia Hirschberg |
HLT-NAACL | 1 |
| 2008 | An Unsupervised Approach to Biography Production Using Wikipedia
Fadi Biadsy, Julia Hirschberg, Elena Filatova |
ACL | 1 |
| 2007 | Comparing american and palestinian perceptions of charisma using acoustic-prosodic and lexical analysisabstractCharisma, the ability to lead by virtue of personality alone, is difficult to define but relatively easy to identify. However, cultural factors clearly affect perceptions of charisma. In this paper we compare results from parallel perception studies investigating charismatic speech in Palestinian Arabic and American English. We examine acoustic/prosodic and lexical correlates of charisma ratings to determine how the two cultures differ with respect to their views of charismatic speech. Fadi Biadsy, Julia Hirschberg, Andrew Rosenberg, Wisam Dakka |
INTERSPEECH | 1 |