VLDB 2026 Research / reviewers in the wild / expert
Arlindo Rodrigues Galvão Filho
dblp:134/0945 · also Arlindo R. G. Filho, Arlindo R. Galvão Filho
· DBLP profile ↗
14ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-2151-8039ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ToxSyn-PT: A Synthetic Fine-Grained Dataset of Minority-Targeted Toxic Language in Portuguese
Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Diogo Fernandes, Arlindo Rodrigues Galvão Filho |
LREC | 5 |
| 2026 | MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese SpeakersabstractWhile large language models (LLMs) show transformative potential in healthcare, their development remains focused on high-resource languages. This creates a critical barrier for other languages, as simple translation fails to capture unique clinical and cultural nuances, such as endemic diseases. To address this, we introduce MedPT, the first large-scale, real-world corpus of patient-doctor interactions for the Brazilian Portuguese medical domain. Comprising 384,095 authentic question-answer pairs and covering over 3,200 distinct health-related conditions, the dataset was refined through a rigorous multi-stage curation protocol that employed a hybrid quantitative-qualitative analysis to filter noise and contextually enrich thousands of ambiguous queries, resulting in a corpus of approximately 57 million tokens. We further utilize of LLM-driven annotation to classify queries into seven semantic types to capture user intent. To validate MedPT's utility, we benchmark it in a medical specialty classification task: fine-tuning a 1.7B parameter model achieves an outstanding 94\% F1-score on a 20-class setup. Furthermore, our qualitative error analysis shows misclassifications are not random but reflect genuine clinical ambiguities (e.g., between comorbid conditions), proving the dataset's deep semantic richness. We publicly release MedPT on Hugging Face to support the development of more equitable, accurate, and culturally-aware medical technologies for the Portuguese-speaking world. Fernanda Bufon Färber, Iago Alves Brito, Julia Soares Dollis, Pedro Schindler Freire Brasil Ribeiro, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho |
LREC | 6 |
| 2025 | BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTSabstractThe detection of audio deepfakes (ADD) has become increasingly important due to the rapid evolution of generative speech models. However, progress in this field remains uneven across languages, particularly for low-resource languages like Portuguese, which lack high-quality datasets. In this paper, we introduce BRSpeech-DF, the first publicly available ADD dataset for Portuguese, encompassing both Brazilian and European variants. The dataset contains over 458,000 utterances, including a smaller portion of real speech from 62 speakers and a large collection of synthetic samples generated using multiple zero-shot text-to-speech (TTS) models, each conditioned on the original speaker’s voice. By providing this resource, our objective is to support the development of robust, multilingual detection systems, thereby advancing equity in speech forensics and security research. BRSpeech-DF addresses a significant gap in annotated data for underrepresented languages, facilitating more inclusive and generalizable advancements in synthetic speech detection. Alexandre Costa Ferro Filho, Rafaello Virgilli, Lucas Alcântara Souza, Frederico Santos de Oliveira, Marcelo Henrique Lopes Ferreira, Daniel Tunnermann, Gustavo dos Reis Oliveira, Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho |
EMNLP | 9 |
| 2025 | "Is It Responsible?"Emerging Results on Comparing Guardrails for Harm Mitigation in LLM-Enhanced Software ApplicationsabstractBackground: The rapid adoption of Large Language Models (LLMs) in software engineering, such as in customer service chatbots, has brought substantial benefits but also significant risks. Biased, inappropriate, or harmful responses may occur when integrating LLMs as commercial off-the-shelf (COTS) components into user-facing applications. Aims: This paper aims to present emerging results from an exploratory study that compares commercial guardrail frameworks, assessing their effectiveness in filtering inappropriate content during chat interactions. Method: We empirically evaluated three guardrail frameworks LLM Guard, Llama Guard, and OpenAI Moderation - using two datasets containing toxic and offensive content. Results: The frameworks achieved high accuracy for one dataset (over 90%) but underperformed in other metrics, indicating that toxic or dangerous content could still reach users in certain scenarios, such as deployment in chatbots. Conclusions: These findings$h$ighlight$t$he$n$eed$f$or further improvements in guardrail frameworks and provide insights for researchers and practitioners to support the selection of appropriate solutions for enhancing harm mitigation in LLMbased applications. Manoel Veríssimo dos Santos Neto, Valdemar Vicente Graciano Neto, Arlindo Rodrigues Galvão Filho, Mohamad Kassab, Edson OliveiraJr |
ESEM | 3 |
| 2025 | FreeSVC: Towards Zero-shot Multilingual Singing Voice ConversionabstractThis work presents FreeSVC, a promising multilingual singing voice conversion approach that leverages an enhanced VITS model with Speaker-invariant Clustering (SPIN) for better content representation and the State-of-the-Art (SOTA) speaker encoder ECAPA2. FreeSVC incorporates trainable language embeddings to handle multiple languages and employs an advanced speaker encoder to disentangle speaker characteristics from linguistic content. Designed for zero-shot learning, FreeSVC enables cross-lingual singing voice conversion without extensive language-specific training. We demonstrate that a multilingual content extractor is crucial for optimal cross-language conversion. Our source code and models are publicly available1. Alef Iury Siqueira Ferreira, Lucas Gris, Augusto Seben da Rosa, Frederico Santos de Oliveira, Edresson Casanova, Rafael Teixeira Sousa, Arnaldo Cândido Jr., Anderson da Silva Soares, Arlindo Rodrigues Galvão Filho |
ICASSP | 9 |
| 2025 | Evaluating Deep Speaker Embedding Robustness to Domain, Sampling Rate, and Codec Variations
Alexandre Ferro Filho, Diogo Fernandes Costa Silva, Pedro Elias Engelberg Silva Borges, Arlindo Rodrigues Galvão Filho |
INTERSPEECH | 4 |
| 2025 | An LLM-Enhanced Framework for Bridging Simulators and Game Engines towards Realistic 3D SimulationsabstractModeling and Simulation (M&S) is a fundamental approach for engineering disruptive solutions, widely applied in critical domains such as aerospace, military, healthcare, traffic, and energy.These domains often exhibit high complexity, uncertainty, and nonlinearity, making simulation and visualization particularly challenging.While M&S allows for the evaluation and analysis of innovative systems, enhancing the experience for analysts and engineers through complementary tools is crucial for improving realism, accuracy, and trustworthiness.In this paper, we present preliminary results on the development of a framework that integrates M&S with advanced visualization capabilities.Our approach bridges the Discrete-Event System Specification (DEVS) formalism with the Unity engine to provide a realistic graphical representation of simulations.By leveraging a Large Language Model (LLM) trained for this purpose, we automatically generate both simulation and animation codes, ensuring synchronized execution.The generated models exchange data, with the simulator serving as the computational engine and Unity as the visualization platform.Additionally, we provide a library of pre-defined simulation models for reuse in various domains.This integration enhances the interpretability of simulations, offering a more intuitive and immersive experience for system analysis. Luiza M. F. Cintra, Elisa Ayumi Masasi de Oliveira, Rafael Teixeira Sousa, Valdemar Vicente Graciano Neto, Arlindo Rodrigues Galvão Filho, Gustavo Higino Webster Barbosa, Sofia Larissa da Costa Paiva |
IMX | 5 |
| 2025 | Immersive Virtual Museums with Spatially-Aware Retrieval-Augmented GenerationabstractVirtual Reality has significantly expanded possibilities for immersive museum experiences, overcoming traditional constraints such as space, preservation, and geographic limitations.However, existing virtual museum platforms typically lack dynamic, personalized, and contextually accurate interactions.To address this, we propose Spatially-Aware Retrieval-Augmented Generation (SA-RAG), an innovative framework integrating visual attention tracking with Retrieval-Augmented Generation systems and advanced Large Language Models.By capturing users' visual attention in real time, SA-RAG dynamically retrieves contextually relevant data, enhancing the accuracy, personalization, and depth of user interactions within immersive virtual environments.The system's effectiveness is initially demonstrated through our preliminary tests within a realistic VR museum implemented using Unreal Engine.Although promising, comprehensive human evaluations involving broader user groups are planned for future studies to rigorously validate SA-RAG's effectiveness, educational enrichment potential, and accessibility improvements in virtual museums.The framework also presents opportunities for broader applications in immersive educational and storytelling domains. Elisa Ayumi Masasi de Oliveira, Rafael Teixeira Sousa, Andressa Araújo Bastos, Luiza M. F. Cintra, Arlindo Rodrigues Galvão Filho |
IMX | 5 |
| 2024 | Multi-Strain Dynamics: Modeling Dengue Transmission in Brazilian Regions with SIR ModelsabstractDengue is a viral disease that represents a significant public health challenge in Brazil, with different virus strains complicating control efforts. This study aims to analyze the transmission dynamics of dengue fever in Brazilian regions, considering the influence of different virus strains. Multi-strain virus spread exhibit distinct transmission dynamics, leading to strain-specific impacts on the timing, intensity, and geographic spread of outbreaks. This study highlights the importance of understanding multi-strain dynamics for effective dengue control in Brazil. The study case for Goiás and Federal District in 2023 achieved a MAPE of 0.1901 and 0.3053, respectively. In 2024, although relatively lower MAPE values of 0.1884 and 0.2045 were achieved, more observations are necessary to enhance the predictive accuracy of the model. In that context, this approach has proven to be a promising tool for analyzing the dynamics of dengue transmission in Goiás and Federal District. Arthur Ricardo de Sousa Vitória, Adriel Lenner Vinhal Mori, Clarimar José Coelho, Arlindo Rodrigues Galvão Filho |
CBMS | 4 |
| 2023 | Live Births Prediction using Legendre Memory Unit: A Case Study for the Health Regions of GoiásabstractThe use of forecasting models is becoming even more common in healthcare and administration applications because it can be a reliable decision support tool. Live birth rate is a health index that is directly linked with maternal and newborn health and its prediction can assist health managers to anticipate resources destined for obstetric and pediatric services. Thus, the objective of this work is to forecast the number of live births in the state of Golás (Brazil) for a 24-month horizon, providing useful information to support the planning and implementation of public policies. The model suggested is the Legendre Memory Unit (LMU) which is applied to data provided by the information system on live births of the information department of the single health system (SINASC-DATASUS). The dataset is composed of 252 monthly records of the number of live births for the 18 health regions of Golás. The results were measured in prediction ability by Mean Absolute Percentual Error (MAPE) and Mean Absolute Error (MAE). The average MAPE and MAE were 6.4614 and 19.9136, respectively. Gabriela Kaori Diógenes, Arthur Ricardo de Sousa Vitória, Diogo Fernandes Costa Silva, Daniel do Prado Pagotto, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho |
CBMS | 6 |
| 2023 | Live Birth Forecasting in Brazillian Health Regions with Tree-based Machine Learning ModelsabstractThis paper aims to do time series forecasting of live births in Brazil with modern tree-based machine learning models. These models are popular choices for time series forecasting due to their ability to model non-linear relationships, so they were applied to live birth forecasting with multiple covariates. The study uses data from the Brazilian Ministry of Health to train and evaluate forecasting models, following guidelines of the Ministry's expectations and needs for using forecasts for public policy planning. The study uses data from all 450 micro-regions in Brazil with records between the years 2000 and 2020. The objective is to train a tree-based model with all months between 2000 and 2018 years to assess the performance of forecasting the number of births over the years 2019 and 2020. LightGBM, XGBoost, and Catboost were evaluated and compared to AutoARIMA and simple linear regression. LightGBM performed slightly better than other models evaluated achieving a MAPE of 0.0797, with more consistent performance over the 24 months of the forecasting horizon. The results show that the tree-based models are reliable for dealing with multiple covariates and can be a useful tool for public policy planning. Douglas Vieira Do Nascimento, Rafael Teixeira Sousa, Diogo Fernandes Costa Silva, Daniel do Prado Pagotto, Clarimar José Coelho, Arlindo Rodrigues Galvão Filho |
CBMS | 6 |
| 2023 | Pancreatic Cancer Detection Using Hyperspectral Imaging and Machine LearningabstractPancreatic cancer is a highly lethal disease, for which mortality is similar to incidence. Most patients with pancreatic cancer do not show symptoms until the disease has reached an advanced stage. The high mortality of pancreatic cancer is mainly due to fact that more than 50% of patients already discover it with metastasis, which reduces treatment options and chances of cure. The success of treatment depends on discovering disease as early as possible. Diagnosis in pancreatic cancer is traditionally confirmed by tissue biopsy of organ. This work presents a methodology to aid the diagnosis based on hyperspectral image for carcinogenic tissue classification using partial least squares and discriminant analysis to optimize process of diagnosing pancreatic adenocarcinoma. The results showed overlapping of areas classified by proposed model and by images used for diagnosis, proving to be a potential tool to aid in the diagnosis of pancreatic cancer. Arlindo Rodrigues Galvão Filho, Isabela Jubé Wastowski, Marise A. R. Moreira, Maria A. de P. C. Cysneiros, Clarimar José Coelho |
ICIP | 1 |
| 2017 | Integer-based genetic algorithm for feature selection in multivariate calibrationabstractFeature selection is a important tast to reduce dimensionality in large datasets. Datasets from multivariate calibration problems are a good example lf datasets with a large number of features. In literature, there are several types of techniques to reduce the number of features for this problem, among them, evolutionary algorithms such as genetic algorithms (GAs). They have been successfully used with binary encoding to select features in multivariate calibration. However, as far as we know, there is no work in literature which provides an integer encoding GA in such context. Thus, this paper presents an integer-based GA implementation for feature selection in multivariate calibration models. The results demonstrated that our proposal is able to outperform the outcomes of participants from 2014 IDRC regarding model prediction error as well as number of selected features. In this dataset, the samples correspond to oils from petroleum reservoirs around the world and gas mixtures in the gas phase measured in transmittance. The gain of our proposed implementation in relation to the winner was from 20.9% up to 88.8%. Rhelcris S. Sousa, Telma Woerle de Lima Soares, Lauro Cássio Martins de Paula, Roney Lopes Lima, Arlindo Rodrigues Galvão Filho, Anderson da Silva Soares |
CEC | 5 |
| 2013 | Multi-objective evolutionary algorithm for variable selection in calibration problems: A case study for protein concentration predictionabstractThis paper presents a multi-objective formulation for variable selection in calibration problems. The prediction of protein concentration on wheat is obtained by a linear regression model using variables obtained by a spectrophotometer device. This device measure hundreds of correlated variables related with physicochemical properties and that can be used to estimate the protein concentration. The problem is the selection of a subset informative and uncorrelated variables that help the minimization of prediction error. In this work we propose the use of two objectives in this problem: the prediction error and the number of variables in the model, both related to linear equations system stability. We proposed a multi-objective formulation using two multi-objective algorithms: the NSGA-II and the SPEA-II. Additionally we propose a final decision maker method to choice the final subset of variables from the Pareto front. For the case study is used wheat data obtained by NIR spectrometry where the objective is the determination of a variable subgroup with information about protein concentration. The results of traditional techniques of multivariate calibration as the Successive Projections Algorithm (SPA), Partial Least Square (PLS) and mono-objective genetic algorithm are presents for comparisons. For NIR spectral analysis of protein concentration on wheat, the number of variables selected from 775 spectral variables was reduced for just 10 in the SPEA-II algorithm. The prediction error decreased from 0.2 in the classical methods to 0.09 in proposed approach, a reduction of 45%. The model using variables selected by SPEA-II had better prediction performance than classical algorithms and full-spectrum partial least-squares (PLS). Daniel Vitor de Lucena, Telma Woerle de Lima Soares, Anderson da Silva Soares, Alexandre C. B. Delbem, Arlindo Rodrigues Galvão Filho, Clarimar José Coelho, Gustavo Teodoro Laureano |
IEEE Congress on Evolutionary Computation | 5 |