Arnaldo Cândido Jr.

dblp:10/7438 · also Arnaldo Cândido Júnior · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-5647-0891ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic Modeling
abstract
Recently, several public datasets for automatic speech recognition (ASR) in Brazilian Portuguese (BP) have been released, improving ASR systems performance. However, these datasets lack diversity in terms of age groups, regional accents, and education levels. In this paper, we present a new publicly available dataset consisting of 289 life story interviews (365 hours), featuring a broad range of speakers varying in age, education, and regional accents. First, we demonstrated the presence of bias in current BP ASR models concerning education levels and age groups. Second, we showed that our dataset helps mitigate these biases. Additionally, an ASR model trained on our dataset performed better during evaluation on a diverse test set. Finally, the ASR model trained with our dataset was extrinsically evaluated through a topic modeling task that utilized the automatically transcribed output.
Sidney Evaldo Leal, Arnaldo Cândido Jr., Ricardo M. Marcacini, Edresson Casanova, Odilon Gonçalves, Anderson da Silva Soares, Rodrigo Lima 0004, Lucas Gris, Sandra M. Aluísio
COLING2
2025 FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
abstract
This work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available1.
Alef Iury Siqueira Ferreira, Lucas Gris, Augusto Seben da Rosa, Frederico Santos de Oliveira, Edresson Casanova, Rafael Teixeira Sousa, Arnaldo Cândido Jr., Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho
ICASSP7
2023 Discriminant Audio Properties in Deep Learning Based Respiratory Insufficiency Detection in Brazilian Portuguese
Marcelo M. Gauy, Larissa Cristina Berti, Arnaldo Cândido Jr., Augusto Camargo Neto, Alfredo Goldman, Anna Sara Shafferman Levin, Marcus Martins, Beatriz Raposo de Medeiros, Marcelo Queiroz, Ester C. Sabino, Flaviane Romani Fernandes Svartman, Marcelo Finger
AIME3
2023 ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti
INTERSPEECH4
2022 YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone
abstract
YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training.
Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Cândido Jr., Eren Gölge, Moacir Ponti
ICML4
2021 Using Open Information Extraction to Extract Relations: An Extended Systematic Mapping
abstract
Context: For thousands of years humans have been using natural language to register their knowledge on important information to enable its access to future generations. With internet, a large amount of textual data is produced and shared on a daily basis. So, scientists started to research techniques for efficiently process knowledge stored in textual format. In this context, Natural Language Processing (NLP) became a popular area studying linguistic phenomena and using computational methods to process texts in natural language. In particular, Open Information Extraction (Open IE) was proposed to gather information from plain text. Despite the advances in this area, it is still necessary to map details about how these approaches were proposed to support the community while creating more efficient Open IE systems. Objective: In this paper, we identify, in the literature, the main characteristics of proposed Open IE approaches. Method: First, we extended the search performed in a systematic mapping previously published by using backward snowballing and a manual search. Next, we updated the electronic database search including ACL Anthology. Finally, 159 studies proposing Open IE approaches were considered for data extraction. Results: Data analysis showed a significant increase in the number of studies published about Open IE in the last years. In addition, we provide important details about how these techniques were proposed (e.g., data sets used and output evaluation techniques). Results indicate that researchers started to adopt neural networks to perform Open IE instead of using conventional supervised learning techniques. Conclusion: Recent advances in Artificial Intelligence and neural networks techniques allowed scientists to have a new perspective on how to perform efficient textual data management. Therefore, Open IE approaches gained much attention as they can help in many contexts, especially in knowledge management tasks.
Vinícius G. dos Santos, Patrick Rodrigo da Silva, Erica Ferreira 0001, Kátia Romero Felizardo, Willian Massami Watanabe, Arnaldo Cândido Jr., Giovani Volnei Meinerz, Sandra M. Aluísio, Nandamudi Lankalapalli Vijaykumar
CLEI6
2021 Using Natural Language Processing to Build Graphical Abstracts to be used in Studies Selection Activity in Secondary Studies
abstract
Context: Secondary studies, as Systematic Literature Reviews (SLRs) and Systematic Mappings (SMs), have been providing methodological and structured processes to identify and select research evidence in Computer Science, especially in Software Engineering (SE). One of the main activities of a secondary study process is to read the abstracts to decide on including or excluding studies. This activity is considered costly and time-consuming. In order to speed up the selection activity, some alternatives such as, structured abstracts and graphical abstracts (e.g. Concept Maps – CMs), have been proposed. Objective: This study presents an approach to automatically build CMs using Natural Language Processing (NLP) to support the selection activity of secondary studies. Method: First, we proposed an approach composed by two pipelines: (1) perform the triple extraction of concept-relation-concept based on NLP; and (2) attach the extracted triples in a structure used as a template to scientific studies. Second, we evaluated both pipelines conducting experiments. Results: The preliminary evaluation revealed that CMs extracted are coherent when compared with their source text. Conclusions: NLP can assist the automatic construction of CMs. In addition, the experiment results show that the approach can be useful to support researchers in the selection of studies in the selection activity of secondary studies.
Vinícius G. dos Santos, Erica Ferreira 0001, Kátia Romero Felizardo, Willian Massami Watanabe, Arnaldo Cândido Jr., Sandra M. Aluísio, Nandamudi Lankalapalli Vijaykumar
SEAA5
2021 Transfer Learning and Data Augmentation Techniques to the COVID-19 Identification Tasks in ComParE 2021
abstract
In this work, we propose several techniques to address data scarceness in ComParE 2021 COVID-19 identification tasks for the application of deep models such as Convolutional Neural Networks.Data is initially preprocessed into spectrogram or MFCC-gram formats.After preprocessing, we combine three different data augmentation techniques to be applied in model training.Then we employ transfer learning techniques from pretrained audio neural networks.Those techniques are applied to several distinct neural architectures.For COVID-19 identification in speech segments, we obtained competitive results.On the other hand, in the identification task based on cough data, we succeeded in producing a noticeable improvement on existing baselines, reaching 75.9% unweighted average recall (UAR).
Edresson Casanova, Arnaldo Cândido Jr., Ricardo Corso Fernandes Junior, Marcelo Finger, Lucas Gris, Moacir Ponti, Daniel Peixoto Pinto da Silva
Interspeech2
2021 SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model
abstract
In this paper, we propose SC-GlowTTS: an efficient zero-shot multi-speaker text-to-speech model that improves similarity for speakers unseen during training. We propose a speaker-conditional architecture that explores a flow-based decoder that works in a zero-shot scenario. As text encoders, we explore a dilated residual convolutional-based encoder, gated convolutional-based encoder, and transformer-based encoder. Additionally, we have shown that adjusting a GAN-based vocoder for the spectrograms predicted by the TTS model on the training dataset can significantly improve the similarity and speech quality for new speakers. Our model converges using only 11 speakers, reaching state-of-the-art results for similarity with new speakers, as well as high speech quality.
Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas M. Müller, Frederico Santos de Oliveira, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti
Interspeech6
2020 Reducing efforts of software engineering systematic literature reviews updates using text classification
Willian Massami Watanabe, Kátia Romero Felizardo, Arnaldo Cândido Jr., Erica Ferreira 0001, José Ede de Campos Neto, Nandamudi Lankalapalli Vijaykumar
Inf. Softw. Technol.3
2012 Rhetorical Move Detection in English Abstracts: Multi-label Sentence Classifiers and their Annotated Corpora
Carmen Dayrell, Arnaldo Cândido Jr., Gabriel Lima, Danilo Machado Jr., Ann A. Copestake, Valéria Delisandra Feltrim, Stella E. O. Tagnin, Sandra M. Aluísio
LREC2