VLDB 2026 Research / reviewers in the wild / expert
Berrak Sisman
dblp:184/9511
· DBLP profile ↗
53ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0001-8078-3305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language ModelsabstractEmotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) encode it internally.We present the first neuron-level interpretability study of emotion-sensitive neurons (ESNs) in LALMs and provide causal evidence supporting the existence of such units in Qwen2.5-Omni,Kimi-Audio, and Audio Flamingo 3. Across these three widely used open-source models, we compare frequency-, entropy-, mean-deviation-, and contrast-based neuron selectors on multiple emotion recognition benchmarks.Using inference-time interventions, we reveal a consistent emotion-specific signature: deactivating neurons selected for a given emotion disproportionately degrades recognition of that emotion while largely preserving other classes, whereas targeted steering amplifies these units to bias predictions toward the target emotion.These effects arise with modest amounts of identification data and scale systematically with intervention strength.We further observe that ESNs exhibit non-uniform layerwise clustering with partial cross-dataset transfer.Taken together, our results offer a causal, neuron-level account of emotion decisions in LALMs and highlight targeted neuron interventions as an actionable handle for controllable affective behaviors. Xiutian Zhao, Björn W. Schuller, Berrak Sisman |
ACL (1) | 3 |
| 2025 | Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH).The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the target utterance.However, MDH contains fine-grained semantic and prosody knowledge at the word level.Existing methods overlook the fine-grained semantic and prosodic interaction modeling.To address this gap, we propose MFCIG-CSS, a novel Multimodal Fine-grained Context Interaction Graph-based CSS system.Our approach constructs two specialized multimodal fine-grained dialogue interaction graphs: a semantic interaction graph and a prosody interaction graph.These two interaction graphs effectively encode interactions between word-level semantics, prosody, and their influence on subsequent utterances in MDH.The encoded interaction features are then leveraged to enhance synthesized speech with natural conversational prosody.Experiments on the DailyTalk dataset demonstrate that MFCIG-CSS outperforms all baseline models in terms of prosodic expressiveness.Code and speech samples are available at https://github.com/AI-S2-Lab/MFCIG-CSS. Zhenqi Jia, Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
EMNLP | 3 |
| 2025 | Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
Rui Liu 0008, Pu Gao, Jiatian Xi, Berrak Sisman, Carlos Busso, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2025 | EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast
Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso, Berrak Sisman |
INTERSPEECH | 5 |
| 2025 | Can Emotion Fool Anti-spoofing?
Aurosweta Mahapatra, Ismail Rasim Ülgen, Abinay Reddy Naini, Carlos Busso, Berrak Sisman |
INTERSPEECH | 5 |
| 2025 | The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso |
INTERSPEECH | 10 |
| 2025 | Advancing Pediatric ASR: The Role of Voice Generation in Disordered Speech
Karen Rosero, Ali N. Salman, Shreeram Suresh Chandra, Berrak Sisman, Cortney Van't Slot, Alex A. Kane, Rami R. Hallac, Carlos Busso |
INTERSPEECH | 4 |
| 2025 | PRESENT: Zero-Shot Text-to-Prosody ControlabstractCurrent strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-to-speech (TTS) models, we present PRESENT (PRosody Editing without Style Embeddings or New Training), which exploits explicit prosody prediction in FastSpeech2-based models by modifying the inference process directly. We apply our text-to-prosody framework to zero-shot language transfer using a JETS model exclusively trained on English LJSpeech data. We obtain character error rates (CER) of 12.8%, 18.7% and 5.9% for German, Hungarian and Spanish respectively, beating the previous state-of-the-art CER by over 2× for all three languages. Furthermore, we allow subphoneme-level control, a first in this field. To evaluate its effectiveness, we show that PRESENT can improve the prosody of questions, and use it to generate Mandarin, a tonal language where vowel pitch varies at subphoneme level. We attain 25.3% hanzi CER and 13.0% pinyin CER with the JETS model. All our code and audio samples11https://github.com/iamanigeeit/presentandhttps://present2024.web.app/are available online. Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans |
IEEE Signal Process. Lett. | 4 |
| 2025 | Versatile Audio-Visual Learning for Emotion RecognitionabstractMost current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented interchangeably for either predicting emotional attributes or recognizing categorical emotions. Achieving such flexibility in a multimodal emotion recognition system is difficult due to the inherent challenges in accurately interpreting and integrating varied data sources. It is also a challenge to robustly handle missing or partial information while allowing direct switch between regression or classification tasks. This study proposes a versatile audio-visual learning (VAVL) framework for handling unimodal and multimodal systems for emotion regression or emotion classification tasks. We implement an audio-visual framework that can be trained even when audio and visual paired data is not available for part of the training set (i.e., audio only or only video is present). We achieve this effective representation learning with audio-visual shared layers, residual connections over shared layers, and a unimodal reconstruction task. Our experimental results reveal that our architecture significantly outperforms strong baselines on the CREMA-D, MSP-IMPROV, and CMU-MOSEI corpora. Notably, VAVL attains a new state-of-the-art performance in the emotional attribute prediction task on the MSP-IMPROV corpus. Lucas Goncalves, Seong-Gyun Leem, Berrak Sisman, Carlos Busso |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Enhanced Facial Landmarks Detection for Patients with Repaired Cleft Lip and PalateabstractCleft lip and palate (CLP) is a congenital condition causing deformities in the oral and labial tissues. Post-surgery, patients often experience residual issues like facial asymmetry, and speech disorders. Tracking points in the orofacial area using a facial landmark detector (FLD) contributes to the assessment of speech development and movement impairments. However, off-the-shelf FLDs fail at delineating the lips of patients with repaired CLP. To address this need, our study introduces the CLP-Trans strategy, a domain transfer solution that employs tailor-made affine transformations to modify facial images sourced from publicly available datasets, which constitute our source domain, whereas images of patients with repaired CLP form our target domain. We aim to reduce distribution disparities between the source and target domains for FLD by simulating common outcomes of CLP repair surgery. The system utilizes a deep convolutional neural network (CNN) to learn from transformed images, therefore, preserving the privacy and facilitating the reproducibility of the findings. The strategy achieves statistically significant improvements in the normalized mean square error (NMSE), reducing it from 2.417 to 2.086 (i.e., 13.7% error reduction) by using the proposed strategy when evaluating images of patients with CLP. Karen Rosero, Ali N. Salman, Berrak Sisman, Rami R. Hallac, Carlos Busso |
FG | 3 |
| 2024 | Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion RecognitionabstractSpeaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, leading to their under-utilization. Our study reveals a direct and useful link between emotion and state-of-the-art speaker embeddings in the form of intra-speaker clusters. By conducting a thorough clustering analysis, we demonstrate that emotion information can be readily extracted from speaker embeddings. In order to leverage this information, we introduce a novel contrastive pretraining approach applied to emotion-unlabeled data for speech emotion recognition. The proposed approach involves the sampling of positive and the negative examples based on the intra-speaker clusters of speaker embeddings. The proposed strategy, which leverages extensive emotion-unlabeled data, leads to a significant improvement in SER performance, whether employed as a standalone pretraining task or integrated into a multi-task pretraining setting. Ismail Rasim Ülgen, Zongyang Du, Carlos Busso, Berrak Sisman |
ICASSP | 4 |
| 2024 | Unsupervised Domain Adaptation for Speech Emotion Recognition using K-Nearest Neighbors Voice Conversion
Pravin Mote, Berrak Sisman, Carlos Busso |
INTERSPEECH | 2 |
| 2024 | Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ülgen, Carlos Busso, Berrak Sisman |
INTERSPEECH | 6 |
| 2024 | Discrete Unit Based Masking For Improving Disentanglement in Voice ConversionabstractVoice conversion (VC) aims to modify the speaker’s identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker’s identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility. Philip H. Lee, Ismail Rasim Ülgen, Berrak Sisman |
SLT | 3 |
| 2024 | SNIPER Training: Single-Shot Sparse Training for Text-to-SpeechabstractText-to-speech (TTS) models have achieved remarkable naturalness in recent years, yet like most deep neural models, they have more parameters than necessary. Sparse TTS models can improve on dense models via pruning and extra retraining, or converge faster than dense models with some performance loss. Thus, we propose training TTS models using decaying sparsity, i.e. a high initial sparsity to accelerate training first, followed by a progressive rate reduction to obtain better eventual performance. This decremental approach differs from current methods of incrementing sparsity to a desired target, which costs significantly more time than dense training. We call our method SNIPER training: Single-shot Initialization Pruning Evolving-Rate training. Our experiments on FastSpeech2 show that we were able to obtain better losses in the first few training epochs with SNIPER, and that the final SNIPER-trained models outperformed constant-sparsity models and edged out dense models, with negligible difference in training time. Our code is available on Github11https://github.com/iamanigeeit/sniper. Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans |
TENCON | 4 |
| 2024 | Accented Text-to-Speech Synthesis with a Conditional Variational AutoencoderabstractAccent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis based on a Conditional Variational Autoencoder. It has the ability to synthesize a selected speaker's voice, and convert this to any desired target accent. Our thorough experiments validate the effectiveness of the proposed framework using both objective and subjective evaluations. The results also show remarkable performance in terms of the model's ability to manipulate accents in the synthesized speech. Overall, our proposed framework presents a promising avenue for future accented TTS research. Jan Melechovský, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
TENCON | 3 |
| 2024 | Accent Conversion in Text-to-Speech Using Multi-Level VAE and Adversarial TrainingabstractWith rapid globalization, the need to build inclu-sive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers. Inclusive speech technology aims to erase any biases towards specific groups, such as people of certain accent. We note that state-of-the-art Text-to-Speech (TTS) systems may currently not be suitable for all people, regardless of their background, as they are designed to generate high-quality voices without focusing on accent. In this paper, we propose a TTS model that utilizes a Multi-Level Variational Autoencoder with adversarial learning to address accented speech synthesis and conversion in TTS, with a vision for more inclusive systems in the future. We evaluate the performance through both objective metrics and subjective listening tests. The results show an improvement in accent conversion ability compared to the baseline. Jan Melechovský, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
TENCON | 3 |
| 2024 | Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity RenderingabstractAccented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1), which is challenging as L2 is different from L1 in terms of phonetic rendering and prosody pattern (pitch, energy, and duration variance, etc.). Accented TTS has several significant real-world applications, such as language learning, preserving and documenting endangered languages and dialects, etc. that make it an important area of research and development. Moreover, changing the accent intensity of any conversational AI system has the potential to allow specific users to understand its produced speech better. However, there is no intuitive solution for the control of the accent intensity for an utterance at both fine and coarse-grained levels, that are phoneme and utterance levels respectively. In this work, we propose a neural TTS architecture that allows us to control the accent style and its intensity. This is achieved through two novel mechanisms: 1) the front-end and back-end accent knowledge injection mechanism to enhance the accent interpretability of TTS modeling; and 2) an automatic speech recognition (ASR) based accent intensity modeling strategy to quantify the accent intensity in both L2 phoneme and utterance levels. In the front-end, a newaccent variation adaptorseeks to project the accent-aware pitch, energy and duration features at a phoneme level, with the help of the fine-grained accent intensity information; In the back-end, a consistency constraint module that ensures the synthesized L2 speech manifests the expected accent intensity, is injected in the front-end, precisely. Experiments show that the proposed system attains superior performance to the baseline models in terms of accent rendering and intensity control. To our knowledge, this is the first study of accented TTS with explicit intensity control at both fine and coarse-grained levels. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
Mirazul Haque, Rutvij Shah, Berrak Sisman, Cong Liu 0005, Wei Yang 0013 |
INTERSPEECH | 4 |
| 2023 | High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units
Junchen Lu, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2023 | Emotion Intensity and its Control for Emotional Voice ConversionabstractEmotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech also conveys emotions with various intensity levels that the listener can perceive. In this paper, we aim to explicitly characterize and control the intensity of emotion. We propose to disentangle the speaker style from linguistic content and encode the speaker style into a style embedding in a continuous space that forms the prototype of emotion embedding. We further learn the actual emotion encoder from an emotion-labelled database and study the use of relative attributes to represent fine-grained emotion intensity. To ensure emotional intelligibility, we incorporateemotion classification lossandemotion embedding similarity lossinto the training of the EVC network. As desired, the proposed network controls the fine-grained emotion intensity in the output speech. Through both objective and subjective evaluations, we validate the effectiveness of the proposed network for emotional expressiveness and emotion intensity control. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Speech Synthesis With Mixed EmotionsabstractEmotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech with a mixture of emotions at run-time. We propose a novel formulation that measures the relative difference between the speech samples of different emotions. We then incorporate our formulation into a sequence-to-sequence emotional text-to-speech framework. During the training, the framework does not only explicitly characterize emotion styles but also explores the ordinal nature of emotions by quantifying the differences with other emotions. At run-time, we control the model to produce the desired emotion mixture by manually defining an emotion attribute vector. The objective and subjective evaluations have validated the effectiveness of the proposed framework. To our best knowledge, this research is the first study on modelling, synthesizing, and evaluating mixed emotions in speech. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice OverabstractIn this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but also perfect lip-speech synchronization. A natural solution to AVO is to condition the speech rendering on the temporal progression of lip sequence in the video. We propose a novel text-to-speech model that is conditioned on visual input, named VisualTTS, for accurate lip-speech synchronization. The proposed VisualTTS adopts two novel mechanisms that are 1) textual-visual attention, and 2) visual fusion strategy during acoustic decoding, which both contribute to forming accurate alignment between the input text content and lip motion in input lip sequence. Experimental results show that VisualTTS achieves accurate lip-speech synchronization and outperforms all baseline systems. Junchen Lu, Berrak Sisman, Rui Liu 0008, Mingyang Zhang 0003, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep LearningabstractEmotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet. Rui Liu 0008, Berrak Sisman, Björn W. Schuller, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2022 | Disentanglement of Emotional Style and Speaker Identity for Expressive Voice ConversionabstractExpressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional style for different speakers. Inspired by the recent success of speaker disentanglement with variational autoencoder (VAE), we propose an any-to-any expressive voice conversion framework, that is called StyleVC. StyleVC is designed to disentangle linguistic content, speaker identity, pitch, and emotional style information. We study the use of style encoder to model emotional style explicitly. At run-time, StyleVC converts both speaker identity and emotional style for arbitrary speakers. Experiments validate the effectiveness of our proposed framework in both objective and subjective evaluations. Zongyang Du, Berrak Sisman, Kun Zhou 0003, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2022 | EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech ModelsabstractNeural models are known to be over-parameterized, and recent work has shown that sparse text-to-speech (TTS) models can outperform dense models. Although a plethora of sparse methods has been proposed for other domains, such methods have rarely been applied in TTS. In this work, we seek to answer the question: what are the characteristics of selected sparse techniques on the performance and model complexity? We compare a Tacotron2 baseline and the results of applying five techniques. We then evaluate the performance via the factors of naturalness, intelligibility and prosody, while reporting model size and training time. Complementary to prior research, we find that pruning before or during training can achieve similar performance to pruning after training and can be trained much faster, while removing entire neurons degrades performance much more than removing parameters. To our best knowledge, this is the first work that compares sparsity paradigms in text-to-speech synthesis. Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman |
INTERSPEECH | 4 |
| 2022 | Learning Accent Representation with Multi-Level VAE Towards Controllable Speech SynthesisabstractAccent is a crucial aspect of speech that helps define one's identity. We note that the state-of-the-art Text-to-Speech (TTS) systems can achieve high-quality generated voice, but still lack in terms of versatility and customizability. Moreover, they generally do not take into account accent, which is an important feature of speaking style. In this work, we utilize the concept of Multi-level VAE (ML-VAE) to build a control mechanism that aims to disentangle accent from a reference accented speaker; and to synthesize voices in different accents such as English, American, Irish, and Scottish. The proposed framework can also achieve high-quality accented voice generation for multi-speaker setup, which we believe is remarkable. We investigate the performance through objective metrics and conduct listening experiments for a subjective performance assessment. We showed that the proposed method achieves good performance for naturalness, speaker similarity, and accent similarity. Jan Melechovský, Ambuj Mehrish, Dorien Herremans, Berrak Sisman |
SLT | 4 |
| 2022 | Emotional voice conversion: Theory, databases and ESDabstractIn this paper, we first provide a review of the state-of-the-art emotional voice conversion research, and the existing emotional speech databases. We then motivate the development of a novel emotional speech database (ESD) that addresses the increasing research need. With this paper, the ESD database1 is now made available to the research community. The ESD database consists of 350 parallel utterances spoken by 10 native English and 10 native Chinese speakers and covers 5 emotion categories (neutral, happy, angry, sad and surprise). More than 29 h of speech data were recorded in a controlled acoustic environment. The database is suitable for multi-speaker and cross-lingual emotional voice conversion studies. As case studies, we implement several state-of-the-art emotional voice conversion systems on the ESD database. This paper provides a reference study on ESD in conjunction with its release. Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2022 | Decoding Knowledge Transfer for Neural Text-to-Speech TrainingabstractNeural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways. However, the exposure bias problem, that arises from the mismatch between the training and inference process in autoregressive models, remains an issue. It often leads to performance degradation in face of out-of-domain test data. To address this problem, we study a novel decoding knowledge transfer strategy, and propose a multi-teacher knowledge distillation (MT-KD) network for Tacotron2 TTS model. The idea is to pre-train two Tacotron2 TTS teacher models in teacher forcing and scheduled sampling modes, and transfer the pre-trained knowledge to a student model that performs free running decoding. We show that the MT-KD network provides an adequate platform for neural TTS training, where the student model learns to emulate the behaviors of the two teachers, at the same time, minimizing the mismatch between training and run-time inference. Experiments on both Chinese and English data show that MT-KD system consistently outperforms the competitive baselines in terms of naturalness, robustness and expressiveness for in-domain and out-of-domain test data. Furthermore, we show that knowledge distillation outperforms adversarial learning and data augmentation in addressing the exposure bias problem. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style TransferabstractTraditional voice conversion (VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in daily communication, and the emotional style of speech can be speaker-dependent. In this paper, we study a technique to jointly convert the speaker identity and speaker-dependent emotional style, that is called expressive voice conversion. We propose a StarGAN-based framework to learn a many-to-many mapping across different speakers, that takes into account speaker-dependent emotional style without the need for parallel data. To this end, we condition the generator on emotional style encoding derived from a pre-trained speech emotion recognition (SER) model. The experiments validate the effectiveness of our proposed framework in both objective and subjective evaluations. To our best knowledge, this is the first study on expressive voice conversion. Zongyang Du, Berrak Sisman, Kun Zhou 0003, Haizhou Li 0001 |
ASRU | 2 |
| 2021 | DEEPA: A Deep Neural Analyzer for Speech and Singing VocodingabstractConventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under certain assumptions about the signals following signal processing principle, therefore, not easily generalizable to different audio, for example, from speech to singing. In this paper, we propose a deep neural analyzer, denoted as DeepA – a neural vocoder that extracts F0 and timbre/aperiodicity encoding from the input speech that emulate those defined in conventional vocoders. Therefore, the resulting parameters are more interpretable than other latent neural representations. At the same time, as the deep neural analyzer is learnable, it is expected to be more accurate for signal reconstruction and manipulation, and generalizable from speech to singing. The proposed neural analyzer is built based on a variational autoencoder (VAE) architecture. We show that DeepA improves F0 estimation over the conventional vocoder (WORLD). To our best knowledge, this is the first study dedicated to the development of a neural framework for extracting learnable vocoder-like parameters. Sergey Nikonorov, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
ASRU | 2 |
| 2021 | Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech SynthesisabstractAttention-based end-to-end text-to-speech synthesis (TTS) is superior to conventional statistical methods in many ways. Transformer-based TTS is one of such successful implementations. While Transformer TTS models the speech frame sequence well with a self-attention mechanism, it does not associate input text with output utterances from a syntactic point of view at sentence level. We propose a novel neural TTS model, denoted as GraphSpeech, that is formulated under graph neural network framework. GraphSpeech encodes explicitly the syntactic relation of input lexical tokens in a sentence, and incorporates such information to derive syntactically motivated character embeddings for TTS attention mechanism. Experiments show that GraphSpeech consistently outperforms the Transformer TTS baseline in terms of spectrum and prosody rendering of utterances. Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
ICASSP | 2 |
| 2021 | Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech DatasetabstractEmotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network conditioned on discrete representation, such as one-hot emotion labels. Such networks learn to remember a fixed set of emotional styles. In this paper, we propose a novel framework based on variational auto-encoding Wasserstein generative adversarial network (VAW-GAN), which makes use of a pre-trained speech emotion recognition (SER) model to transfer emotional style during training and at run-time inference. In this way, the network is able to transfer both seen and unseen emotional style to a new utterance. We show that the proposed framework achieves remarkable performance by consistently outperforming the baseline framework. This paper also marks the release of an emotional speech dataset (ESD) for voice conversion, which has multiple speakers and languages. Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001 |
ICASSP | 2 |
| 2021 | Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion DiscriminabilityabstractEmotional text-to-speech synthesis (ETTS) has seen much progress in recent years.However, the generated voice is often not perceptually identifiable by its intended emotion category.To address this problem, we propose a new interactive training paradigm for ETTS, denoted as i-ETTS, which seeks to directly improve the emotion discriminability by interacting with a speech emotion recognition (SER) model.Moreover, we formulate an iterative training strategy with reinforcement learning to ensure the quality of i-ETTS optimization.Experimental results demonstrate that the proposed i-ETTS outperforms the state-of-the-art baselines by rendering speech with more accurate emotion style.To our best knowledge, this is the first study of reinforcement learning in emotional text-to-speech synthesis. Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
Interspeech | 2 |
| 2021 | Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence TrainingabstractEmotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional voice conversion with a limited amount of emotional speech data. We note that the proposed EVC framework leverages text-to-speech (TTS) as they share a common goal that is to generate high-quality expressive voice. In stage 1, we perform style initialization with a multi-speaker TTS corpus, to disentangle speaking style and linguistic content. In stage 2, we perform emotion training with a limited amount of emotional speech data, to learn how to disentangle emotional style and linguistic information from the speech. The proposed framework can perform both spectrum and prosody conversion and achieves significant improvement over the state-of-the-art baselines in both objective and subjective evaluation. Kun Zhou 0003, Berrak Sisman, Haizhou Li 0001 |
Interspeech | 2 |
| 2021 | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li |
SIGDIAL | 5 |
| 2021 | Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In SpeechabstractEmotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional elements in speech through variational autoencoding Wasserstein generative adversarial network (VAW-GAN). We propose a speaker-dependent EVC framework based on VAW-GAN, that includes two VAW-GAN pipelines, one for spectrum conversion, and another for prosody conversion. We train a spectral encoder that disentangles emotion and prosody (F0) information from spectral features; we also train a prosodic encoder that disentangles emotion modulation of prosody (affective prosody) from linguistic prosody. At run-time, the decoder of spectral VAW-GAN is conditioned on the output of prosodic VAW-GAN. The vocoder takes the converted spectral and prosodic features to generate the target emotional speech. Experiments validate the effectiveness of our proposed method in both objective and subjective evaluations. Kun Zhou 0003, Berrak Sisman, Haizhou Li 0001 |
SLT | 2 |
| 2021 | FastTalker: A neural text-to-speech architecture with shallow and group autoregression
Rui Liu 0008, Berrak Sisman, Yixing Lin, Haizhou Li 0001 |
Neural Networks | 2 |
| 2021 | Exploiting Morphological and Phonological Features to Improve Prosodic Phrasing for Mongolian Speech SynthesisabstractProsodic phrasing is an important factor that affects naturalness and intelligibility in text-to-speech synthesis. Studies show that deep learning techniques improve prosodic phrasing when large text and speech corpus are available. However, for low-resource languages, such as Mongolian, prosodic phrasing remains a challenge for various reasons. First, the database suitable for system training is limited. Second, word composition knowledge that is prosody-informing has not been used in prosodic phrase modeling. To address these problems, in this article, we propose a feature augmentation method in conjunction with a self-attention neural classifier. We augment input text with morphological and phonological decompositions of words to enhance the text encoder. We study the use of self-attention classifier, that makes use of global context of a sentence, as a decoder for phrase break prediction. Both objective and subjective evaluations validate the effectiveness of the proposed phrase break prediction framework, that consistently improves voice quality in a Mongolian text-to-speech synthesis system. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Expressive TTS Training With Frame and Style Reconstruction LossabstractWe propose a novel training strategy for Tacotron-based text-to-speech (TTS) system that improves the speech styling at utterance level. One of the key challenges in prosody modeling is the lack of reference that makes explicit modeling difficult. The proposed technique doesn’t require prosody annotations from training data. It doesn’t attempt to model prosody explicitly either, but rather encodes the association between input text and its prosody styles using a Tacotron-based TTS framework. This study marks a departure from the style token paradigm where prosody is explicitly modeled by a bank of prosody embeddings. It adopts a combination of two objective functions: 1) frame level reconstruction loss, that is calculated between the synthesized and target spectral features; 2) utterance level style reconstruction loss, that is calculated between the deep style features of synthesized and target speech. The style reconstruction loss is formulated as a perceptual loss to ensure that utterance level speech style is taken into consideration during training. Experiments show that the proposed training strategy achieves remarkable performance and outperforms the state-of-the-art baseline in both naturalness and expressiveness. To our best knowledge, this is the first study to incorporate utterance level perceptual quality as a loss function into Tacotron training for improved expressiveness. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep LearningabstractSpeaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this article, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research. Berrak Sisman, Junichi Yamagishi, Simon King 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Teacher-Student Training For Robust Tacotron-Based TTSabstractWhile neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that results in unpredictable performance for out-of-domain test data at run-time. To overcome this, we propose a teacher-student training scheme for Tacotron-based TTS by introducing a distillation loss function in addition to the feature loss function. We first train a Tacotron2-based TTS model by always providing natural speech frames to the decoder, that serves as a teacher model. We then train another Tacotron2-based model as a student model, of which the decoder takes the predicted speech frames as input, similar to how the decoder works during run-time inference. With the distillation loss, the student model learns the output probabilities from the teacher model, that is called knowledge distillation. Experiments show that our proposed training scheme consistently improves the voice quality for out-of-domain test data both in Chinese and English systems. Rui Liu 0008, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 2 |
| 2020 | Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice ConversionabstractEmotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity.The prior studies on emotional voice conversion are mostly carried out under the assumption that emotion is speaker-dependent.We consider that there is a common code between speakers for emotional expression in a spoken language, therefore, a speaker-independent mapping between emotional states is possible.In this paper, we propose a speaker-independent emotional voice conversion framework, that can convert anyone's emotion without the need for parallel data.We propose a VAW-GAN based encoderdecoder structure to learn the spectrum and prosody mapping.We perform prosody conversion by using continuous wavelet transform (CWT) to model the temporal dependencies.We also investigate the use of F0 as an additional input to the decoder to improve emotion conversion performance.Experiments show that the proposed speaker-independent framework achieves competitive results for both seen and unseen speakers. Kun Zhou 0003, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2020 | DeepConversion: Voice conversion with limited parallel training dataabstractA deep neural network approach to voice conversion usually depends on a large amount of parallel training data from source and target speakers. In this paper, we propose a novel conversion pipeline, DeepConversion, that leverages a large amount of non-parallel, multi-speaker data, but requires only a small amount of parallel training data. It is believed that we can represent the shared characteristics of speakers by training a speaker independent general model on a large amount of publicly available, non-parallel, multi-speaker speech data. Such general model can then be used to learn the mapping between source and target speaker more effectively from a limited amount of parallel training data. We also propose a strategy to make full use of the parallel data in all models along the pipeline. In particular, the parallel data is used to adapt the general model towards the source-target speaker pair to achieve a coarse grained conversion, and to develop a compact Error Reduction Network (ERN) for a fine-grained conversion. The parallel data is also used to adapt the WaveNet vocoder towards the source-target pair. The experiments show that DeepConversion that only uses a limited amount of parallel training data, consistently outperforms the traditional approaches that use a large amount of parallel training data, in both objective and subjective evaluations. Mingyang Zhang 0003, Berrak Sisman, Li Zhao 0003, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2020 | Modeling Prosodic Phrasing With Multi-Task Learning in Tacotron-Based TTSabstractTacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur frequently. In this letter, we extend the Tacotron-based speech synthesis framework to explicitly model the prosodic phrase breaks. We propose a multi-task learning scheme for Tacotron training, that optimizes the system to predict both Mel spectrum and phrase breaks. To our best knowledge, this is the first implementation of multi-task learning for Tacotron based TTS with a prosodic phrasing model. Experiments show that our proposed training scheme consistently improves the voice quality for both Chinese and Mongolian systems. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | On the Study of Generative Adversarial Networks for Cross-Lingual Voice ConversionabstractCross-lingual voice conversion (VC) aims to convert the source speaker's voice to sound like that of the target speaker, when the source and target speakers speak different languages. In this paper, we propose to use Generative Adversarial Networks (GANs) for cross-lingual voice-conversion. We further the studies on Variational Autoencoding Wasserstein GAN (VAW-GAN) and cycle-consistent adversarial network (CycleGAN), that are known to be effective for mono-lingual voice conversion. As cross-lingual voice conversion needs to converts the voice across different phonetic system, it is more challenging than mono-lingual voice conversion. By using VAW-GAN and CycleGAN, we successfully convert the speaker identity while carrying over the source speaker's linguistic content. The proposed idea is unique in the sense that it neither relies on bilingual data and their alignment, nor any external process, such as ASR. Moreover, it works with limited amount of training data of any two languages. To our best knowledge, this is the first comprehensive study of Generative Adversarial Networks in cross-lingual voice conversion. In the experiments, we achieve high-quality converted voice, that performs equally well or better than mono-lingual voice conversion. Berrak Sisman, Mingyang Zhang 0003, Minghui Dong, Haizhou Li 0001 |
ASRU | 1 |
| 2019 | VQVAE Unsupervised Unit Discovery and Multi-Scale Code2Spec Inverter for Zerospeech Challenge 2019abstractWe describe our submitted system for the ZeroSpeech Challenge 2019.The current challenge theme addresses the difficulty of constructing a speech synthesizer without any text or phonetic labels and requires a system that can (1) discover subword units in an unsupervised way, and (2) synthesize the speech with a target speaker's voice.Moreover, the system should also balance the discrimination score ABX, the bit-rate compression rate, and the naturalness and the intelligibility of the constructed voice.To tackle these problems and achieve the best tradeoff, we utilize a vector quantized variational autoencoder (VQ-VAE) and a multi-scale codebook-tospectrogram (Code2Spec) inverter trained by mean square error and adversarial loss.The VQ-VAE extracts the speech to a latent space, forces itself to map it into the nearest codebook and produces compressed representation.Next, the inverter generates a magnitude spectrogram to the target voice, given the codebook vectors from VQ-VAE.In our experiments, we also investigated several other clustering algorithms, including K-Means and GMM, and compared them with the VQ-VAE result on ABX scores and bit rates.Our proposed approach significantly improved the intelligibility (in CER), the MOS, and discrimination ABX scores compared to the official ZeroSpeech 2019 baseline or even the topline. Andros Tjandra, Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
INTERSPEECH | 2 |
| 2019 | Group Sparse Representation With WaveNet Vocoder Adaptation for Spectrum and Prosody ConversionabstractThe statistical approach to voice conversion typically consists of a feature conversion module followed by a vocoder. So far, the feature conversion studies are mainly focused on the conversion of spectrum. However, speaker identity is also characterized by prosodic features, such as fundamental frequency F0 and energy contour among others. In this paper, we study the transformation of speaker characteristics both in terms of spectrum and prosody. We propose two novel techniques that effectively use a limited amount of source-target training data and leverage a large general speech corpus to improve the voice conversion quality. First, we study the phonetic sparse representation under the group sparsity mathematical formulation. We use phonetic posteriorgrams PPGs together with spectral and prosody features to form tandem feature in the phonetic dictionary. The tandem feature allow us to estimate an activation matrix that is less dependent on source speakers, thus providing a better voice conversion quality. Second, we study the use of WaveNet vocoder that can be trained on general speech corpus from multiple speakers and adapted on target speaker data to improve the vocoding quality. We benefit from the large general speech databases that are used to train the PPG generator, and the WaveNet vocoder. The experiments show that the proposed conversion framework outperforms the traditional spectrum and prosody conversion techniques in both objective and subjective evaluations. Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Wavelet Analysis of Speaker Dependent and Independent Prosody for Voice Conversion
Berrak Sisman, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2018 | A Voice Conversion Framework with Tandem Feature Sparse Representation and Speaker-Adapted WaveNet Vocoder
Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2018 | Adaptive Wavenet Vocoder for Residual Compensation in GAN-Based Voice ConversionabstractIn this paper, we propose to use generative adversarial networks (GAN) together with a WaveNet vocoder to address the over-smoothing problem arising from the deep learning approaches to voice conversion, and to improve the vocoding quality over the traditional vocoders. As GAN aims to minimize the divergence between the natural and converted speech parameters, it effectively alleviates the over-smoothing problem in the converted speech. On the other hand, WaveNet vocoder allows us to leverage from the human speech of a large speaker population, thus improving the naturalness of the synthetic voice. Furthermore, for the first time, we study how to use WaveNet vocoder for residual compensation to improve the voice conversion performance. The experiments show that the proposed voice conversion framework consistently outperforms the baselines. Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
SLT | 1 |
| 2017 | Sparse representation of phonetic features for voice conversion with and without parallel dataabstractThis paper presents a voice conversion framework that uses phonetic information in an exemplar-based voice conversion approach. The proposed idea is motivated by the fact that phone-dependent exemplars lead to better estimation of activation matrix, therefore, possibly better conversion. We propose to use the phone segmentation results from automatic speech recognition (ASR) to construct a sub-dictionary for each phone. The proposed framework can work with or without parallel training data. With parallel training data, we found that phonetic sub-dictionary outperforms the state-of-the-art baseline in objective and subjective evaluations. Without parallel training data, we use Phonetic PosteriorGrams (PPGs) as the speaker-independent exemplars in the phonetic sub-dictionary to serve as a bridge between speakers. We report that such technique achieves a competitive performance without the need of parallel training data. Berrak Sisman, Haizhou Li 0001, Kay Chen Tan |
ASRU | 1 |
| 2016 | Energy and data cooperation in energy harvesting multiple access channelabstractWe consider the energy harvesting two user Gaussian multiple access channel (MAC), where both users harvest energy from nature. The users cooperate at the physical layer (data cooperation) by establishing common messages through overheard signals and then cooperatively sending them. In addition, the users cooperate at the battery level (energy cooperation) by wirelessly transferring energy to each other. We find the jointly optimal offline transmit power and rate allocation policy together with the energy transfer policy that maximizes the departure region. We provide necessary conditions for energy transfer, and prove some properties of the optimal transmit policy, thereby shedding some light on the interplay between energy and data cooperation. Berk Gurakan, Berrak Sisman, Onur Kaya, Sennur Ulukus |
WCNC | 2 |