EDBT 2026 Demo / reviewers in the wild / expert
Edresson Casanova
dblp:258/8636
· DBLP profile ↗
16ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0003-0160-7173ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open Full-duplex Voice Agent with Speech-to-Speech Language ModelabstractWe present the system demonstration and opensource code release of a novel, data-efficient framework that converts any standard text Large Language Model (LLM) into a full-duplex end-to-end (E2E) speech-to-speech (S2S) model, for building conversational voice agents. Our new modeling method enables any LLMs to simultaneously listen and speak without requiring extensive speech-text pretraining. Moreover, we demonstrate how to put together a low-latency and full-duplex voice agent with open-source modeling, inference optimization, and serving solutions. This work significantly lowers the barrier to entry for developing low-latency, human-like voice agents by providing a generalizable, end-to-end solution built on open-source technologies. Edresson Casanova, Chen Chen 0075, Kevin Hu, Ankita Pasad, Elena Rastorgueva, Seelan Lakshmi Narasimhan, Slyne Deng, Ehsan Hosseini-Asl, Piotr Zelasko, Valentin Mendelev, Subhankar Ghosh, Yifan Peng 0003, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Vitaly Lavrukhin, Boris Ginsburg |
ASRU | 1 |
| 2025 | MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic ModelingabstractRecently, several public datasets for automatic speech recognition (ASR) in Brazilian Portuguese (BP) have been released, improving ASR systems performance. However, these datasets lack diversity in terms of age groups, regional accents, and education levels. In this paper, we present a new publicly available dataset consisting of 289 life story interviews (365 hours), featuring a broad range of speakers varying in age, education, and regional accents. First, we demonstrated the presence of bias in current BP ASR models concerning education levels and age groups. Second, we showed that our dataset helps mitigate these biases. Additionally, an ASR model trained on our dataset performed better during evaluation on a diverse test set. Finally, the ASR model trained with our dataset was extrinsically evaluated through a topic modeling task that utilized the automatically transcribed output. Sidney Evaldo Leal, Arnaldo Cândido Jr., Ricardo M. Marcacini, Edresson Casanova, Odilon Gonçalves, Anderson da Silva Soares, Rodrigo Lima 0004, Lucas Gris, Sandra M. Aluísio |
COLING | 4 |
| 2025 | Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free GuidanceabstractShehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shehzeen Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li 0007 |
EMNLP | 4 |
| 2025 | Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and InferenceabstractLarge language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models. Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain, Jason Li 0007, Subhankar Ghosh, Ante Jukic, Sang-gil Lee |
ICASSP | 1 |
| 2025 | FreeSVC: Towards Zero-shot Multilingual Singing Voice ConversionabstractThis work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available1. Alef Iury Siqueira Ferreira, Lucas Gris, Augusto Seben da Rosa, Frederico Santos de Oliveira, Edresson Casanova, Rafael Teixeira Sousa, Arnaldo Cândido Jr., Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho |
ICASSP | 5 |
| 2025 | NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukic, Jason Li 0007, Boris Ginsburg |
INTERSPEECH | 1 |
| 2025 | Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Ehsan Hosseini-Asl, Chen Chen 0075, Edresson Casanova, Subhankar Ghosh, Piotr Zelasko, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 4 |
| 2025 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain, Edresson Casanova, Evelina Bakhturina, Jason Li 0007 |
INTERSPEECH | 5 |
| 2024 | MLAAD: The Multi-Language Audio Anti-Spoofing DatasetabstractText-to-Speech (TTS) technology brings significant advantages, such as giving a voice to those with speech impairments, but also enables audio deepfakes and spoofs. The former mislead individuals and may propagate misinformation, while the latter undermine voice biometric security systems. AI-based detection can help to address these challenges by automatically differentiating between genuine and fabricated voice recordings. However, these models are only as good as their training data, which currently is severely limited due to an overwhelming concentration on English and Chinese audio in anti-spoofing databases, thus restricting its worldwide effectiveness.In response, this paper presents the Multi-Language Audio Anti-Spoof Dataset (MLAAD), created using 52 TTS models, comprising 22 different architectures, to generate 160.2 hours of synthetic voice in 23 different languages. We train and evaluate three state-of-the-art deepfake detection models with MLAAD, and observe that MLAAD demonstrates superior performance over comparable datasets like InTheWild or FakeOrReal when used as a training resource. Furthermore, in comparison with the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, both excelling on four datasets.By publishing1MLAAD and making trained models accessible via an interactive webserver2, we aim to democratize antispoofing technology, making it accessible beyond the realm of specialists, thus contributing to global efforts against audio spoofing and deepfakes. Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Piotr Syga, Philip Sperl, Konstantin Böttinger |
IJCNN | 4 |
| 2024 | XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Logan Hart, Aya Aljafari, Reuben Morais, Samuel Olayemi, Julian Weber |
INTERSPEECH | 1 |
| 2023 | ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti |
INTERSPEECH | 1 |
| 2022 | YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneabstractYourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training. Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Cândido Jr., Eren Gölge, Moacir Ponti |
ICML | 1 |
| 2022 | BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpusabstractBibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa.The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models.The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba.This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica.We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language.We present results for text-to-speech models with Coqui TTS. David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack, Julian Weber, Salomon Kabongo, Elizabeth Salesky, Iroro Orife, Colin Leong, Perez Ogayo, Chris C. Emezue, Jonathan Mukiibi, Salomey Osei, Apelete Agbolo, Victor Akinode, Bernard Opoku, Samuel Olanrewaju, Jesujoba O. Alabi, Shamsuddeen Hassan Muhammad |
INTERSPEECH | 3 |
| 2021 | Transfer Learning and Data Augmentation Techniques to the COVID-19 Identification Tasks in ComParE 2021abstractIn this work, we propose several techniques to address data scarceness in ComParE 2021 COVID-19 identification tasks for the application of deep models such as Convolutional Neural Networks.Data is initially preprocessed into spectrogram or MFCC-gram formats.After preprocessing, we combine three different data augmentation techniques to be applied in model training.Then we employ transfer learning techniques from pretrained audio neural networks.Those techniques are applied to several distinct neural architectures.For COVID-19 identification in speech segments, we obtained competitive results.On the other hand, in the identification task based on cough data, we succeeded in producing a noticeable improvement on existing baselines, reaching 75.9% unweighted average recall (UAR). Edresson Casanova, Arnaldo Cândido Jr., Ricardo Corso Fernandes Junior, Marcelo Finger, Lucas Gris, Moacir Ponti, Daniel Peixoto Pinto da Silva |
Interspeech | 1 |
| 2021 | SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech ModelabstractIn this paper, we propose SC-GlowTTS: an efficient zero-shot multi-speaker text-to-speech model that improves similarity for speakers unseen during training. We propose a speaker-conditional architecture that explores a flow-based decoder that works in a zero-shot scenario. As text encoders, we explore a dilated residual convolutional-based encoder, gated convolutional-based encoder, and transformer-based encoder. Additionally, we have shown that adjusting a GAN-based vocoder for the spectrograms predicted by the TTS model on the training dataset can significantly improve the similarity and speech quality for new speakers. Our model converges using only 11 speakers, reaching state-of-the-art results for similarity with new speakers, as well as high speech quality. Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas M. Müller, Frederico Santos de Oliveira, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti |
Interspeech | 1 |
| 2020 | Evaluating Sentence Segmentation in Different Datasets of Neuropsychological Language Tests in Brazilian PortugueseabstractAutomatic analysis of connected speech by natural language processing techniques is a promising direction for diagnosing cognitive impairments. However, some difficulties still remain: the time required for manual narrative transcription and the decision on how transcripts should be divided into sentences for successful application of parsers used in metrics, such as Idea Density, to analyze the transcripts. The main goal of this paper was to develop a generic segmentation system for narratives of neuropsychological language tests. We explored the performance of our previous single-dataset-trained sentence segmentation architecture in a richer scenario involving three new datasets used to diagnose cognitive impairments, comprising different stories and two types of stimulus presentation for eliciting narratives — visual and oral — via illustrated story-book and sequence of scenes, and by retelling. Also, we proposed and evaluated three modifications to our previous RCNN architecture: (i) the inclusion of a Linear Chain CRF; (ii) the inclusion of a self-attention mechanism; and (iii) the replacement of the LSTM recurrent layer by a Quasi-Recurrent Neural Network layer. Our study allowed us to develop two new models for segmenting impaired speech transcriptions, along with an ideal combination of datasets and specific groups of narratives to be used as the training set. Edresson Casanova, Marcos V. Treviso, Lilian Hübner, Sandra M. Aluísio |
LREC | 1 |