VLDB 2026 Research / reviewers in the wild / expert
Sefik Emre Eskimez
dblp:170/0089
· DBLP profile ↗
34ranked-venue papers
14as first author
25since 2021 · last 2025
0000-0001-6259-5925ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 13 first-author · 23 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Speech Extraction with Human Feedback
Malek Itani, Ashton Graves, Sefik Emre Eskimez, Shyamnath Gollakota |
INTERSPEECH | 3 |
| 2025 | TS3-Codec: Transformer-Based Simple Streaming Single Codec
Naoyuki Kanda, Sefik Emre Eskimez, Jinyu Li 0001 |
INTERSPEECH | 3 |
| 2024 | An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
Xiaofei Wang 0009, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Yufei Xia, Jinzhu Li, Sheng Zhao 0002, Jinyu Li 0001, Naoyuki Kanda |
INTERSPEECH | 2 |
| 2024 | Target conversation extraction: Source separation using turn-taking dynamicsabstractExtracting the speech of participants in a conversation amidst interfering speakers and noise presents a challenging problem. In this paper, we introduce the novel task of target conversation extraction, where the goal is to extract the audio of a target conversation based on the speaker embedding of one of its participants. To accomplish this, we propose leveraging temporal patterns inherent in human conversations, particularly turn-taking dynamics, which uniquely characterize speakers engaged in conversation and distinguish them from interfering speakers and noise. Using neural networks, we show the feasibility of our approach on English and Mandarin conversation datasets. In the presence of interfering speakers, our results show an 8.19 dB improvement in signal-to-noise ratio for 2-speaker conversations and a 7.92 dB improvement for 2-4-speaker conversations. Code, dataset available at https://github.com/chentuochao/Target-Conversation-Extraction. Tuochao Chen, Bohan Wu, Malek Itani, Sefik Emre Eskimez, Takuya Yoshioka, Shyamnath Gollakota |
INTERSPEECH | 5 |
| 2024 | Total-Duration-Aware Duration Modeling for Text-to-Speech SystemsabstractAccurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications.However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored.In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration.We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations.Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models.We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts. Sefik Emre Eskimez, Xiaofei Wang 0009, Manthan Thakker, Chung-Hsien Tsai, Canrun Li, Hemin Yang, Zirun Zhu, Jinyu Li 0001, Sheng Zhao 0002, Naoyuki Kanda |
INTERSPEECH | 1 |
| 2024 | Knowledge boosting during low-latency inference
Vidya Srinivas, Malek Itani, Tuochao Chen, Sefik Emre Eskimez, Takuya Yoshioka, Shyamnath Gollakota |
INTERSPEECH | 4 |
| 2024 | E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTSabstractThis paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility. In the E2 TTS framework, the text input is converted into a character sequence with filler tokens. The flow-matching-based mel spectrogram generator is then trained based on the audio infilling task. Unlike many previous works, it does not require additional components (e.g., duration model, grapheme-to-phoneme) or complex techniques (e.g., monotonic alignment search). Despite its simplicity, E2 TTS achieves state-of-the-art zero-shot TTS capabilities that are comparable to or surpass previous works, including Voicebox and NaturalSpeech 3. The simplicity of E2 TTS also allows for flexibility in the input representation. We propose several variants of E2 TTS to improve usability during inference. See https://aka.ms/e2tts/for demo samples. Sefik Emre Eskimez, Xiaofei Wang 0009, Manthan Thakker, Canrun Li, Chung-Hsien Tsai, Hemin Yang, Zirun Zhu, Xu Tan 0003, Sheng Zhao 0002, Naoyuki Kanda |
SLT | 1 |
| 2024 | Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-To-SpeechabstractPeople change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich emotions, including NVs. This paper introduces EmoCtrl-TTS, an emotion-controllable zeroshot TTS that can generate highly emotional speech with NVs for any speaker. EmoCtrl-TTS leverages arousal and valence values, as well as laughter embeddings, to condition the flow-matchingbased zero-shot TTS. To achieve high-quality emotional speech generation, EmoCtrl-TTS is trained using more than 27,000 hours of expressive data curated based on pseudo-labeling. Comprehensive evaluations demonstrate that EmoCtrl-TTS excels in mimicking the emotions of audio prompts in speech-to-speech translation scenarios. We also show that EmoCtrl-TTS can capture emotion changes, express strong emotions, and generate various NVs in zero-shot TTS. See https://aka.ms/emoctrl-tts for demo samples. Xiaofei Wang 0009, Sefik Emre Eskimez, Manthan Thakker, Daniel Tompkins, Chung-Hsien Tsai, Canrun Li, Sheng Zhao 0002, Jinyu Li 0001, Naoyuki Kanda |
SLT | 3 |
| 2024 | SpeechX: Neural Codec Language Model as a Versatile Speech TransformerabstractRecent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. Xiaofei Wang 0007, Manthan Thakker, Zhuo Chen 0006, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Shujie Liu 0001, Jinyu Li 0001, Takuya Yoshioka |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Speech Separation with Large-Scale Self-Supervised LearningabstractSelf-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massively scaling up both the pre-training data (more than 300K hours) and fine-tuning data (10K hours). We also investigate various techniques to efficiently integrate the pre-trained model with the SS network under a limited computation budget, including a low frame rate SSL model training setup and a fine-tuning scheme using only the part of the pre-trained model. Compared with a supervised baseline and the WavLM-based SS model using feature embeddings obtained with the previously released 94K hours trained WavLM, our proposed model obtains 15.9% and 11.2% of relative word error rate (WER) reductions, respectively, for a simulated far-field speech mixture test set. For conversation transcription on real meeting recordings using continuous speech separation, the proposed model achieves 6.8% and 10.6% of relative WER reductions over the purely supervised baseline on AMI and ICSI evaluation sets, respectively, while reducing the computational cost by 38%. Zhuo Chen 0006, Naoyuki Kanda, Jian Wu 0027, Yu Wu 0012, Xiaofei Wang 0009, Takuya Yoshioka, Jinyu Li 0001, Sunit Sivasankaran, Sefik Emre Eskimez |
ICASSP | 9 |
| 2023 | Breaking the Trade-Off in Personalized Speech Enhancement With Cross-Task Knowledge DistillationabstractPersonalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in addition to background noise. Unlike unconditional speech enhancement, causal PSE models may occasionally remove the target speech by mistake. The PSE models also tend to leak interfering speech when the target speaker is silent for an extended period. We show that existing PSE methods suffer from a trade-off between speech over-suppression and interference leakage by addressing one problem at the expense of the other. We propose a new PSE model training framework using cross-task knowledge distillation to mitigate this trade-off. Specifically, we utilize a personalized voice activity detector (pVAD) during training to exclude the non-target speech frames that are wrongly identified as containing the target speaker with hard or soft classification. This prevents the PSE model from being too aggressive while still allowing the model to learn to suppress the input speech when it is likely to be spoken by interfering speakers. Comprehensive evaluation results are presented, covering various PSE usage scenarios. Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka |
ICASSP | 2 |
| 2023 | Real-Time Audio-Visual End-To-End Speech EnhancementabstractAudio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers’ voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model’s performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase. Zirun Zhu, Hemin Yang, Sefik Emre Eskimez, Huaming Wang |
ICASSP | 5 |
| 2023 | Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation
Sefik Emre Eskimez, Takuya Yoshioka, Alex Ju, Tanel Pärnamaa, Huaming Wang |
INTERSPEECH | 1 |
| 2023 | Improving Readability for Automatic Speech Recognition TranscriptionabstractModern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to grammatical errors, disfluency, and other noises common in spoken communication. These readable issues introduced by speakers and ASR systems will impair the performance of downstream tasks and the understanding of human readers. In this work, we present a task called ASR post-processing for readability (APR) and formulate it as a sequence-to-sequence text generation problem. The APR task aims to transform the noisy ASR output into a readable text for humans and downstream tasks while maintaining the semantic meaning of speakers. We further study the APR task from the benchmark dataset, evaluation metrics, and baseline models: First, to address the lack of task-specific data, we propose a method to construct a dataset for the APR task by using the data collected for grammatical error correction. Second, we utilize metrics adapted or borrowed from similar tasks to evaluate model performance on the APR task. Lastly, we use several typical or adapted pre-trained models as the baseline models for the APR task. Furthermore, we fine-tune the baseline models on the constructed dataset and compare their performance with a traditional pipeline method in terms of proposed evaluation metrics. Experimental results show that all the fine-tuned baseline models perform better than the traditional pipeline method, and our adapted RoBERTa model outperforms the pipeline method by 4.95 and 6.63 BLEU points on two test sets, respectively. The human evaluation and case study further reveal the ability of the proposed model to improve the readability of ASR transcripts. Junwei Liao, Sefik Emre Eskimez, Liyang Lu, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Hong Qu 0002, Michael Zeng 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Icassp 2022 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios. Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, Robert Aichner |
ICASSP | 7 |
| 2022 | Personalized speech enhancement: new models and Comprehensive evaluationabstractPersonalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing systems for various acoustic scenarios. In this work, we propose two neural networks for PSE that achieve superior performance to the previously proposed VoiceFilter. In addition, we create test sets that capture a variety of scenarios that users can encounter during video conferencing. Furthermore, we propose a new metric to measure the target speaker over-suppression (TSOS) problem, which was not sufficiently investigated before despite its critical importance in deployment. Besides, we propose multi-task training with a speech recognition back-end. Our results show that the proposed models can yield better speech recognition accuracy, speech intelligibility, and perceptual quality than the baseline models, and the multi-task training can alleviate the TSOS issue in addition to improving the speech recognition accuracy. Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Xiaofei Wang 0009, Zhuo Chen 0006, Xuedong Huang 0001 |
ICASSP | 1 |
| 2022 | One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech EnhancementabstractWith the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared with the unconditional speech enhancement (SE) methods in these scenarios due to their ability to remove interfering speech in addition to the environmental noise. In this work, we leverage spatial information afforded by microphone arrays to improve such systems’ performance further. We investigate the relative importance of speaker embeddings and spatial features. Moreover, we propose a new causal array-geometry-agnostic multi-channel PSE model, which can generate a high-quality enhanced signal from arbitrary microphone geometry. Experimental results show that the proposed geometry agnostic model outperforms the model trained on a specific microphone array geometry in both speech quality and automatic speech recognition accuracy. We also demonstrate the effectiveness of the proposed approach for unseen array geometries. Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Zhuo Chen 0006, Xuedong Huang 0001 |
ICASSP | 2 |
| 2022 | All-Neural Beamformer for Continuous Speech SeparationabstractContinuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is transcribing a meeting conversation recorded by a microphone array. Prior studies explored various deep learning models for time-frequency mask estimation, followed by a minimum variance distortionless response (MVDR) filter to improve the automatic speech recognition (ASR) accuracy. The performance of these methods is fundamentally upper-bounded by MVDR’s spatial selectivity. Recently, the all deep learning MVDR (ADL-MVDR) model was proposed for neural beamforming and demonstrated superior performance in a target speech extraction task using pre-segmented input. In this paper, we further adapt ADL-MVDR to the CSS task with several enhancements to enable end-to-end neural beamforming. The proposed system achieves significant word error rate reduction over a baseline spectral masking system on the LibriCSS dataset. Moreover, the proposed neural beamformer is shown to be comparable to a state-of-the-art MVDR-based system in real meeting transcription tasks, including AMI, while showing potentials to further simplify the run-time implementation and reduce the system latency with frame-wise processing. Zhuohuang Zhang, Takuya Yoshioka, Naoyuki Kanda, Zhuo Chen 0006, Xiaofei Wang 0009, Dongmei Wang, Sefik Emre Eskimez |
ICASSP | 7 |
| 2022 | Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge DistillationabstractThis paper investigates how to improve the runtime speed of personalized speech enhancement (PSE) networks while maintaining the model quality.Our approach includes two aspects: architecture and knowledge distillation (KD).We propose an end-to-end enhancement (E3Net) model architecture, which is 3× faster than a baseline STFT-based model.Besides, we use KD techniques to develop compressed student models without significantly degrading quality.In addition, we investigate using noisy data without reference clean signals for training the student models, where we combine KD with multi-task learning (MTL) using an automatic speech recognition (ASR) loss.Our results show that E3Net provides better speech and transcription quality with a lower target speaker over-suppression (TSOS) rate than the baseline model.Furthermore, we show that the KD methods can yield student models that are 2 -4× faster than the teacher and provides reasonable quality.Combining KD and MTL improves the ASR and TSOS metrics without degrading the speech quality. Manthan Thakker, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang |
INTERSPEECH | 2 |
| 2022 | Leveraging Real Conversational Data for Multi-Channel Continuous Speech SeparationabstractExisting multi-channel continuous speech separation (CSS) models are heavily dependent on supervised data - either simulated data which causes data mismatch between the training and real-data testing, or the real transcribed overlapping data, which is difficult to be acquired, hindering further improvements in the conversational/meeting transcription tasks. In this paper, we propose a three-stage training scheme for the CSS model that can leverage both supervised data and extra large-scale unsupervised real-world conversational data. The scheme consists of two conventional training approaches -- pre-training using simulated data and ASR-loss-based training using transcribed data -- and a novel continuous semi-supervised training between the two, in which the CSS model is further trained by using real data based on the teacher-student learning framework. We apply this scheme to an array-geometry-agnostic CSS model, which can use the multi-channel data collected from any microphone array. Large-scale meeting transcription experiments are carried out on both Microsoft internal meeting data and the AMI meeting corpus. The steady improvement by each training stage has been observed, showing the effect of the proposed method that enables leveraging real conversational data for CSS model training. Xiaofei Wang 0009, Dongmei Wang, Naoyuki Kanda, Sefik Emre Eskimez, Takuya Yoshioka |
INTERSPEECH | 4 |
| 2022 | Separating Long-Form Speech with Group-wise Permutation Invariant TrainingabstractMulti-talker conversational speech processing has drawn many interests for various applications such as meeting transcription.Speech separation is often required to handle overlapped speech that is commonly observed in conversation.Although the original utterancelevel permutation invariant training-based continuous speech separation approach has proven to be effective in various conditions, it lacks the ability to leverage the long-span relationship of utterances and is computationally inefficient due to the highly overlapped sliding windows.To overcome these drawbacks, we propose a novel training scheme named Group-PIT, which allows direct training of the speech separation models on the long-form speech with a low computational cost for label assignment.Two different speech separation approaches with Group-PIT are explored, including direct long-span speech separation and short-span speech separation with long-span tracking.The experiments on the simulated meeting-style data demonstrate the effectiveness of our proposed approaches, especially in dealing with a very long speech input. Wangyou Zhang, Zhuo Chen 0006, Naoyuki Kanda, Shujie Liu 0001, Jinyu Li 0001, Sefik Emre Eskimez, Takuya Yoshioka, Zhong Meng, Yanmin Qian, Furu Wei |
INTERSPEECH | 6 |
| 2022 | Speech Driven Talking Face Generation From a Single Image and an Emotion ConditionabstractVisual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an end-to-end talking face generation system that takes a speech utterance, a single face image, and a categorical emotion label as input to render a talking face video synchronized with the speech and expressing the conditioned emotion. Objective evaluation on image quality, audiovisual synchronization, and visual emotion expression shows that the proposed system outperforms a state-of-the-art baseline system. Subjective evaluation of visual emotion expression and video realness also demonstrates the superiority of the proposed system. Furthermore, we conduct a human emotion recognition pilot study using generated videos with mismatched emotions among the audio and visual modalities. Results show that humans respond to the visual modality more significantly than the audio modality on this task. Sefik Emre Eskimez, You Zhang 0001, Zhiyao Duan |
IEEE Trans. Multim. | 1 |
| 2021 | Generating Human Readable Transcript for Automatic Speech Recognition with Pre-Trained Language ModelabstractModern Automatic Speech Recognition (ASR) systems can achieve high performance in terms of recognition accuracy. However, a perfectly accurate transcript still can be challenging to read due to disfluency, filter words, and other errata common in spoken communication. Many downstream tasks and human readers rely on the output of the ASR system; therefore, errors introduced by the speaker and ASR system alike will be propagated to the next task in the pipeline. In this work, we propose an ASR post-processing model that aims to transform the incorrect and noisy ASR output into a readable text for humans and downstream tasks. We leverage the Metadata Extraction (MDE) corpus to construct a task-specific dataset for our study. Since the dataset is small, we propose a novel data augmentation method and use a two-stage training strategy to fine-tune the RoBERTa pre-trained model. On the constructed test set, our model outperforms a production two-step pipeline-based post-processing method by a large margin of 13.26 on readability-aware WER (RA-WER) and 17.53 on BLEU metrics. Human evaluation also demonstrates that our method can generate more human-readable transcripts than the baseline method. Junwei Liao, Yu Shi 0001, Ming Gong 0001, Linjun Shou, Sefik Emre Eskimez, Liyang Lu, Hong Qu 0002, Michael Zeng 0001 |
ICASSP | 5 |
| 2021 | One-Shot Voice Conversion with Speaker-Agnostic StarGAN
Sefik Emre Eskimez, Dimitrios Dimitriadis, Ken'ichi Kumatani, Robert Gmyr |
Interspeech | 1 |
| 2021 | Human Listening and Live Captioning: Multi-Task Training for Speech EnhancementabstractWith the surge of online meetings, it has become more critical than ever to provide high-quality speech audio and live captioning under various noise conditions.However, most monaural speech enhancement (SE) models introduce processing artifacts and thus degrade the performance of downstream tasks, including automatic speech recognition (ASR).This paper proposes a multi-task training framework to make the SE models unharmful to ASR.Because most ASR training samples do not have corresponding clean signal references, we alternately perform two model update steps called SE-step and ASR-step.The SEstep uses clean and noisy signal pairs and a signal-based loss function.The ASR-step applies a pre-trained ASR model to training signals enhanced with the SE model.A cross-entropy loss between the ASR output and reference transcriptions is calculated to update the SE model parameters.Experimental results with realistic large-scale settings using ASR models trained on 75,000-hour data show that the proposed framework improves the word error rate for the SE output by 11.82% with little compromise in the SE quality.Performance analysis is also carried out by changing the ASR model, the data used for the ASR-step, and the schedule of the two update steps. Sefik Emre Eskimez, Xiaofei Wang 0009, Hemin Yang, Zirun Zhu, Zhuo Chen 0006, Huaming Wang, Takuya Yoshioka |
Interspeech | 1 |
| 2020 | End-To-End Generation of Talking Faces from Noisy SpeechabstractAcoustic cues are not the only component in speech communication; if the visual counterpart is present, it is shown to benefit speech comprehension. In this work, we propose an end-to-end (no pre- or post-processing) system that can generate talking faces from arbitrarily long noisy speech. We propose a mouth region mask to encourage the network to focus on mouth movements rather than speech irrelevant movements. In addition, we use generative adversarial network (GAN) training to improve the image quality and mouth-speech synchronization. Furthermore, we employ noise-resilient training to make our network robust to unseen non-stationary noise. We evaluate our system with image quality and mouth shape (landmark) measures on noisy speech utterances with five types of unseen non-stationary noise between -10 dB and 30 dB signal-to-noise ratio (SNR) with increments of 1 dB SNR. Results show that our system outperforms a state-of-the-art baseline system significantly, and our noise-resilient training improves performance for noisy speech in a wide range of SNR. Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan |
ICASSP | 1 |
| 2020 | A Federated Approach in Training Acoustic ModelsabstractIn this paper, a novel platform for Acoustic Model training based on Federated Learning (FL) is described. This is the first attempt to introduce Federated Learning techniques in Speech Recognition (SR) tasks. Besides the novelty of the task, the paper describes an easily generalizable FL platform and presents the design decisions used for this task. Amongst the novel algorithms introduced is a hierarchical optimization scheme employing pairs of optimizers and an algorithm for gradient selection, leading to improvements in training time and SR performance. The experimental validation of the proposed system is based on the LibriSpeech task, presenting a speed-up of x1.5 and 6% WERR. The proposed Federated Learning system appears to outperform the golden standard of distributed training in both convergence speed and overall model performance. Further improvements have been experienced in internal tasks. Dimitrios Dimitriadis, Ken'ichi Kumatani, Robert Gmyr, Yashesh Gaur, Sefik Emre Eskimez |
INTERSPEECH | 5 |
| 2020 | GAN-Based Data Generation for Speech Emotion RecognitionabstractIn this work, we propose a GAN-based method to generate synthetic data for speech emotion recognition. Specifically, we investigate the usage of GANs for capturing the data manifold when the data is eyes-off, i.e., where we can train networks using the data but cannot copy it from the clients. We propose a CNN-based GAN with spectral normalization on both the generator and discriminator, both of which are pre-trained on large unlabeled speech corpora. We show that our method provides better speech emotion recognition performance than a strong baseline. Furthermore, we show that even after the data on the client is lost, our model can generate similar data that can be used for model bootstrapping in the future. Although we evaluated our method for speech emotion recognition, it can be applied to other tasks. Sefik Emre Eskimez, Dimitrios Dimitriadis, Robert Gmyr, Kenichi Kumanati |
INTERSPEECH | 1 |
| 2020 | Sequence-Level Self-Learning with Multiple HypothesesabstractIn this work, we develop new self-learning techniques with an attention-based sequence-to-sequence (seq2seq) model for automatic speech recognition (ASR). For untranscribed speech data, the hypothesis from an ASR system must be used as a label. However, the imperfect ASR result makes unsupervised learning difficult to consistently improve recognition performance especially in the case that multiple powerful teacher models are unavailable. In contrast to conventional unsupervised learning approaches, we adopt the \emph{multi-task learning} (MTL) framework where the $n$-th best ASR hypothesis is used as the label of each task. The seq2seq network is updated through the MTL framework so as to find the common representation that can cover multiple hypotheses. By doing so, the effect of the \emph{hard-decision} errors can be alleviated. We first demonstrate the effectiveness of our self-learning methods through ASR experiments in an accent adaptation task between the US and British English speech. Our experiment results show that our method can reduce the WER on the British speech data from 14.55\% to 10.36\% compared to the baseline model trained with the US English data only. Moreover, we investigate the effect of our proposed methods in a federated learning scenario. Ken'ichi Kumatani, Dimitrios Dimitriadis, Yashesh Gaur, Robert Gmyr, Sefik Emre Eskimez, Jinyu Li 0001, Michael Zeng 0001 |
INTERSPEECH | 5 |
| 2020 | Noise-Resilient Training Method for Face Landmark Generation From SpeechabstractVisual cues such as lip movements, when available, play an important role in speech communication. They are especially helpful for the hearing impaired population or in noisy environments. When not available, having a system to automatically generate talking faces in sync with input speech would enhance speech communication and enable many novel applications. In this article, we present a new system that can generate 3D talking face landmarks from speech in an online fashion. We employ a neural network that accepts the raw waveform as an input. The network contains convolutional layers with 1D kernels and outputs the active shape model (ASM) coefficients of face landmarks. To promote smoother transitions between video frames, we present a variant of the model that has the same architecture but also accepts the previous frame's ASM coefficients as an additional input. To cope with background noise, we propose a new training method to incorporate speech enhancement ideas at the feature level. Objective evaluations on landmark prediction show that the proposed system yields statistically significantly smaller errors than two state-of-the-art baseline methods on both a single-speaker dataset and a multi-speaker dataset. Experiments on noisy speech input with five types of non-stationary unseen noise show statistically significant improvements of the system performance thanks to the noise-resilient training method. Finally, subjective evaluations show that the generated talking faces have a significantly more convincing match with the input audio, achieving a similarly convincing level of realism as the ground-truth landmarks. Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Speech Super Resolution Generative Adversarial NetworkabstractThe goal of speech super-resolution (SSR) or speech bandwidth expansion is to generate the missing high-frequency components for a given low-resolution speech signal. It has the potential to improve the quality of telecommunications. We propose a new method for SSR that leverages the generative adversarial networks (GANs) and a regularization method for stabilizing the GAN training. The generator network is a convolutional autoencoder with 1D convolution kernels, operating along time-axis and generating the high-frequency log-power spectra from the low-frequency log-power spectra input. We employ two recent deep neural network (DNN) based approaches to compare them with our proposed method, including both objective speech quality metrics and subjective perceptual tests. We show that our proposed method outperforms the baseline methods in terms of both objective and subjective evaluations. Sefik Emre Eskimez, Kazuhito Koishida |
ICASSP | 1 |
| 2018 | Unsupervised Learning Approach to Feature Analysis for Automatic Speech Emotion RecognitionabstractThe scarcity of emotional speech data is a bottleneck of developing automatic speech emotion recognition (ASER) systems. One way to alleviate this issue is to use unsupervised feature learning techniques to learn features from the widely available general speech and use these features to train emotion classifiers. These unsupervised methods, such as denoising autoencoder (DAE), variational autoencoder (VAE), adversarial autoencoder (AAE) and adversarial variational Bayes (AVB), can capture the intrinsic structure of the data distribution in the learned feature representation. In this work, we systematically investigate four kinds of unsupervised feature learning methods for improving speaker-independent ASER. We show that all methods improve the performance regarding unweighted accuracy rating (UAR) and Fl-score over methods that use hand-crafted features or that do not perform feature learning on external datasets. We also show that VAE, AAE and AVB methods, which control the distribution of the latent representation, outperform DAE that does not control such distribution. This suggests the benefits of using variational inference methods to learn features from general speech for the speech tasks such as ASER that has very limited labeled data. Sefik Emre Eskimez, Zhiyao Duan, Wendi B. Heinzelman |
ICASSP | 1 |
| 2018 | Front-end speech enhancement for commercial speaker verification systems
Sefik Emre Eskimez, Peter Soufleris, Zhiyao Duan, Wendi B. Heinzelman |
Speech Commun. | 1 |
| 2016 | Emotion classification: How does an automated system compare to Naive human coders?abstractThe fact that emotions play a vital role in social interactions, along with the demand for novel human-computer interaction applications, have led to the development of a number of automatic emotion classification systems. However, it is still debatable whether the performance of such systems can compare with human coders. To address this issue, in this study, we present a comprehensive comparison in a speech-based emotion classification task between 138 Amazon Mechanical Turk workers (Turkers) and a state-of-the-art automatic computer system. The comparison includes classifying speech utterances into six emotions (happy, neutral, sad, anger, disgust and fear), into three arousal classes (active, passive, and neutral), and into three valence classes (positive, negative, and neutral). The results show that the computer system outperforms the naive Turkers in almost all cases. Furthermore, the computer system can increase the classification accuracy by rejecting to classify utterances for which it is not confident, while the Turkers do not show a significantly higher classification accuracy on their confident utterances versus unconfi-dent ones. Sefik Emre Eskimez, Kenneth Imade, Melissa Sturge-Apple, Zhiyao Duan, Wendi B. Heinzelman |
ICASSP | 1 |