VLDB 2026 Research / reviewers in the wild / expert
Anderson R. Avila
dblp:158/4090 · also Anderson Raymundo Avila
· DBLP profile ↗
22ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-3088-5116ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Infox-QC: A Quebec-Focused French Corpus for Misinformation Detection and AI Robustness Assessment
Moetaz Doghmane, Hazem Amamou, Thiziri Sefsaf, Alan Davoust, Anderson R. Avila |
LREC | 5 |
| 2026 | PhonemeDF: A Synthetic Speech Dataset for Audio Deepfake Detection and Naturalness Evaluation
Vamshi Nallaguntla, Aishwarya Fursule, Shruti Rajendra Kshirsagar, Anderson R. Avila |
LREC | 4 |
| 2025 | Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative AnalysisabstractRetrieval-Augmented Generation (RAG) was first introduced to enhance the capabilities of Large Language Models (LLMs) beyond their encoded-prior knowledge. This is achieved by providing LLMs with an external source of knowledge, which helps to reduce factual hallucinations and enables the access to new information, typically not available during their pretraining phase. Despite its benefits, there is an increasing concern with the impact of inconsistent retrieved information towards LLMs’ responses. Hence, the Retrieval-Augmented Generation Benchmark (RGB) was introduced as a new testbed for RAG evaluation, meant to assess the robustness of LLMs towards inconsistency in the retrieved information. In this work, we use the RGB corpus to evaluate LLMs in four scenarios: (1) noise robustness; (2) information integration; (3) negative rejection; and (4) counterfactual robustness. We perform a comparative analysis between the RAG baseline defined by the RGB and variations of GraphRAG, which is a RAG system based on a Knowledge Graph (KG) and developed to retrieve relevant information from large documents. We tested GraphRAG under three customization to improve its robustness. Our approach demonstrates improvements compared to the RGB baseline, providing insights on how to design more reliable RAG systems, tailored for real-world scenarios. Hazem Amamou, Stéphane Gagnon, Alan Davoust, Anderson R. Avila |
SMC | 4 |
| 2025 | Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial AttacksabstractAdversarial examples can represent a serious threat to machine learning (ML) algorithms. If used to manipulate the behaviour of ML-based Network Intrusion Detection Systems (NIDS), they can jeopardize network security. In this work, we aim to mitigate such risks by increasing the robustness of NIDS towards adversarial attacks. To that end, we explore two adversarial methods for generating malicious network traffic. The first method is based on Generative Adversarial Networks (GAN) and the second one is the Fast Gradient Sign Method (FGSM). The adversarial examples generated by these methods are then used to evaluate a novel multilayer defense mechanism, specifically designed to mitigate the vulnerability of ML-based NIDS. Our solution consists of one layer of stacking classifiers and a second layer based on an autoencoder. If the incoming network data are classified as benign by the first layer, the second layer is activated to ensure that the decision made by the stacking classifier is correct. We also incorporated adversarial training to further improve the robustness of our solution. Experiments on two datasets, namely UNSW-NB15 and NSL-KDD, demonstrate that the proposed approach increases resilience to adversarial attacks. Nasim Soltani, Shayan Nejadshamsi, Zakaria Abou El Houda, Raphaël Khoury, Kelton A. P. Costa, Tiago H. Falk, Anderson R. Avila |
SMC | 7 |
| 2025 | Phonetic Analysis of Real and Synthetic Speech Using HuBERT Embeddings: Perspectives for Deepfake DetectionabstractThe growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can now produce convincing synthetic speech with high quality and intelligibility. This poses a serious threat to voice biometric security systems, such as automatic speaker recognition. It also increases the risks associated to the spread of spoken disinformation, where synthetic voices can be used to disseminate malicious content. In this study, we conduct an analysis of real and synthetic speech at phonetic and word levels. For that, a parallel dataset comprising real and synthetic speech signals were developed based on a subset of the LibriSpeech ASR corpus. Synthetic speech samples were generated using two TTS and one VC systems: Coqui TTS, VITS TTS, and StarGANv2 VC. We adopted HuBERT, a self-supervised speech model, to extract speech embeddings. The motivation for using this model stems from its ability to recognize sound units corresponding to the so-called pseudo phonemes. Our analysis is based on the KL divergence (KLD) between the distributions of synthetic and real phonemes, which allowed us to rank synthetic phonemes based on their alignment with their real counterpart. We also trained several classifiers per phoneme to distinguish between real and synthetic samples. We then compute the correlations between KLD and accuracies per phoneme. Besides showing a list of phonemes that are more discriminative, our findings suggest that vowels correlate better with the classifiers’ performance, suggesting that the KLD can be an indicator of the most distinguishable phonemes for deepfake detection. Dia Elhak Temmar, Assia Hamadene, Vamshi Nallaguntla, Aishwarya Fursule, Mohand Saïd Allili, Shruti Rajendra Kshirsagar, Anderson R. Avila |
SMC | 7 |
| 2024 | VIC-KD: Variance-Invariance-Covariance Knowledge Distillation to Make Keyword Spotting More Robust Against Adversarial AttacksabstractKeyword spotting (KWS) refers to the task of identifying a set of predefined words in audio streams. With the advances seen recently with deep neural networks, it has become a popular technology to activate and control small devices, such as voice assistants. Relying on such models for edge devices, however, can be challenging due to hardware constraints. Moreover, as adversarial attacks have increased against voice-based technologies, developing solutions robust to such attacks has become crucial. In this work, we propose VIC-KD, a robust distillation recipe for model compression and adversarial robustness. Using self-supervised speech representations, we show that imposing geometric priors to the latent representations of both Teacher and Student models leads to more robust target models. Experiments on the Google Speech Commands datasets show that the proposed methodology improves upon current state-of-the-art robust distillation methods, such as ARD and RSLAD, by 12% and 8% in robust accuracy, respectively. Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila, Tiago H. Falk |
ICASSP | 3 |
| 2023 | A Max-Min Security Game for Coordinated Backdoor Attacks on Federated LearningabstractWe address in this paper the challenge of data poisoning attacks on Federated Learning. We consider a particularly challenging attack scenario in which a single poisoning attack is coordinated over a set of clients to complicate its detection. In response, the federated learning server assigns a weight to each client’s model update with the aim of mitigating the effects of the poisoning on the global model. To address this challenge, we first design a trust mechanism that enables the federated learning server to assess the trustworthiness of each client on the basis of the client’s adherence to the federated learning protocol and the quality of data contributed by the client. Capitalizing on the trust mechanism, we model the interactions between the attacker and federated learning server as a security max-min game. The outcome of the game guides the server on the optimal weight assignment strategy over the set of clients’ model updates, so as to minimize the effects of the data poisoning on the global model. Simulations conducted on the MNIST and CIFAR-10 datasets suggest that our proposed solution decreases the coordinated attack success rate, as well as the false positive and false negative percentages compared to two baseline solutions. Omar Abdel Wahab 0001, Anderson R. Avila |
IEEE Big Data | 2 |
| 2023 | Robustdistiller: Compressing Universal Speech Representations for Enhanced Environment RobustnessabstractSelf-supervised speech pre-training enables deep neural network models to capture meaningful and disentangled factors from raw waveform signals. The learned universal speech representations can then be used across numerous down-stream tasks. These representations, however, are sensitive to distribution shifts caused by environmental factors, such as noise and/or room reverberation. Their large sizes, in turn, make them unfeasible for edge applications. In this work, we propose a knowledge distillation methodology termed RobustDistiller which compresses universal representations while making them more robust against environmental artifacts via a multi-task learning objective. The proposed layer-wise distillation recipe is evaluated on top of three well-established universal representations, as well as with three downstream tasks. Experimental results show the proposed methodology applied on top of the WavLM Base+ teacher model outperforming all other benchmarks across noise types and levels, as well as reverberation times. Oftentimes, the obtained results with the student model (24M parameters) achieved results inline with those of the teacher model (95M). Heitor R. Guimarães, Arthur Pimentel, Anderson R. Avila, Mehdi Rezagholizadeh, Boxing Chen, Tiago H. Falk |
ICASSP | 3 |
| 2023 | Assessing the Vulnerability of Self-Supervised Speech Representations for Keyword Spotting Under White-Box Adversarial AttacksabstractSelf-supervised speech pre-training has emerged as a useful tool to extract representations from speech that can be used across different tasks. While these models are starting to appear in commercial systems, their robustness to so-called adversarial attacks have yet to be fully characterized. This paper evaluates the vulnerability of three self-supervised speech representations (wav2vec 2.0, HuBERT and WavLM) to three white-box adversarial attacks under different signal-to-noise ratios (SNR). The study uses keyword spotting as a downstream task and shows that the models are very vulnerable to attacks, even at high SNRs. The paper also investigates the transferability of attacks between models and analyses the generated noise patterns in order to develop more effective defence mechanisms. The modulation spectrum shows to be a potential tool for detection of adversarial attacks to speech systems. Heitor R. Guimarães, Yi Zhu 0011, Orson Mengara, Anderson R. Avila, Tiago H. Falk |
SMC | 4 |
| 2023 | How Secure is Code Generated by ChatGPT?abstractIn recent years, large language models have been responsible for great advances in the field of artificial intelligence (AI). ChatGPT in particular, an AI chatbot developed and recently released by OpenAI, has taken the field to the next level. The conversational model is able not only to process human-like text, but also to translate natural language into code. However, the safety of programs generated by ChatGPT should not be overlooked. In this paper, we perform an experiment to address this issue. Specifically, we ask ChatGPT to generate a number of computer programs in order to evaluate the security of the resulting source code. We further investigate whether ChatGPT can be prodded to improve code security by appropriate prompts, and discuss the ethical aspects of using AI to generate code. Results suggest that ChatGPT is aware of potential vulnerabilities, but nonetheless often generates source code that are not robust to certain attacks. Raphaël Khoury, Anderson R. Avila, Jacob Brunelle, Baba Mamadou Camara |
SMC | 2 |
| 2022 | Low-bit Shift Network for End-to-End Spoken Language UnderstandingabstractDeep neural networks (DNN) have achieved impressive success in multiple domains.Over the years, the accuracy of these models has increased with the proliferation of deeper and more complex architectures.Thus, state-of-the-art solutions are often computationally expensive, which makes them unfit to be deployed on edge computing platforms.In order to mitigate the high computation, memory, and power requirements of inferring convolutional neural networks (CNNs), we propose the use of power-of-two quantization, which quantizes continuous parameters into low-bit power-of-two values.This reduces computational complexity by removing expensive multiplication operations and with the use of low-bit weights.ResNet is adopted as the building block of our solution and the proposed model is evaluated on a spoken language understanding (SLU) task.Experimental results show improved performance for shift neural network architectures, with our low-bit quantization achieving 98.76 % on the test set which is comparable performance to its full-precision counterpart and state-of-the-art solutions. Anderson R. Avila, Khalil Bibi, Rui Heng Yang, Xinlin Li 0001 |
INTERSPEECH | 1 |
| 2021 | A Streaming End-to-End Framework For Spoken Language UnderstandingabstractEnd-to-end spoken language understanding (SLU) recently attracted increasing interest. Compared to the conventional tandem-based approach that combines speech recognition and language understanding as separate modules, the new approach extracts users' intentions directly from the speech signals, resulting in joint optimization and low latency. Such an approach, however, is typically designed to process one intent at a time, which leads users to have to take multiple rounds to fulfill their requirements while interacting with a dialogue system. In this paper, we propose a streaming end-to-end framework that can process multiple intentions in an online and incremental way. The backbone of our framework is a unidirectional RNN trained with the connectionist temporal classification (CTC) criterion. By this design, an intention can be identified when sufficient evidence has been accumulated, and multiple intentions will be identified sequentially. We evaluate our solution on the Fluent Speech Commands (FSC) dataset and the detection accuracy is about 97 % on all multi-intent settings. This result is comparable to the performance of the state-of-the-art non-streaming models, but is achieved in an online and incremental way. We also employ our model to an keyword spotting task using the Google Speech Commands dataset, and the results are also highly promising. Nihal Potdar, Anderson R. Avila, Yiran Cao |
IJCAI | 2 |
| 2021 | Sequential End-to-End Intent and Slot Label Classification and LocalizationabstractHuman-computer interaction (HCI) is significantly impacted by delayed responses from a spoken dialogue system.Hence, endto-end (e2e) spoken language understanding (SLU) solutions have recently been proposed to decrease latency.Such approaches allow for the extraction of semantic information directly from the speech signal, thus bypassing the need for a transcript from an automatic speech recognition (ASR) system.In this paper, we propose a compact e2e SLU architecture for streaming scenarios, where chunks of the speech signal are processed continuously to predict intent and slot values.Our model is based on a 3D convolutional neural network (3D-CNN) and a unidirectional long short-term memory (LSTM).We compare the performance of two alignment-free losses: the connectionist temporal classification (CTC) method and its adapted version, namely connectionist temporal localization (CTL).The latter performs not only the classification but also localization of sequential audio events.The proposed solution is evaluated on the Fluent Speech Command dataset and results show our model ability to process incoming speech signal, reaching accuracy as high as 98.97 % for CTC and 98.78 % for CTL on single-label classification, and as high as 95.69 % for CTC and 95.28 % for CTL on two-label prediction. Yiran Cao, Nihal Potdar, Anderson R. Avila |
Interspeech | 3 |
| 2021 | On the use of blind channel response estimation and a residual neural network to detect physical access attacks to speaker verification systems
Anderson R. Avila, Jahangir Alam 0001, Fabiano O. Costa Prado, Douglas D. O'Shaughnessy, Tiago H. Falk |
Comput. Speech Lang. | 1 |
| 2021 | Automatic speaker verification from affective speech using Gaussian mixture model based estimation of neutral speech characteristics
Anderson R. Avila, Douglas D. O'Shaughnessy, Tiago H. Falk |
Speech Commun. | 1 |
| 2021 | Feature Pooling of Modulation Spectrum Features for Improved Speech Emotion Recognition in the WildabstractInterest in affective computing is burgeoning, in great part due to its role in emerging affective human-computer interfaces (HCI). To date, the majority of existing research on automated emotion analysis has relied on data collected in controlled environments. With the rise of HCI applications on mobile devices, however, so-called “in-the-wild” settings have posed a serious threat for emotion recognition systems, particularly those based on voice. In this case, environmental factors such as ambient noise and reverberation severely hamper system performance. In this paper, we quantify the detrimental effects that the environment has on emotion recognition and explore the benefits achievable with speech enhancement. Moreover, we propose a modulation spectral feature pooling scheme that is shown to outperform a state-of-the-art benchmark system for environment-robust prediction of spontaneous arousal and valence emotional primitives. Experiments on an environment-corrupted version of the RECOLA dataset of spontaneous interactions show the proposed feature pooling scheme, combined with speech enhancement, outperforming the benchmark across different noise-only, reverberation-only and noise-plus-reverberation conditions. Additional tests with the SEWA database show the benefits of the proposed method for in-the-wild applications. Anderson R. Avila, Zahid Akhtar, João Felipe Santos, Douglas D. O'Shaughnessy, Tiago H. Falk |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Non-intrusive Speech Quality Assessment Using Neural NetworksabstractEstimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neural networks for non-intrusive audio quality assessment. We propose three neural network-based approaches for mean opinion score (MOS) estimation. We compare our results to three instrumental measures: the perceptual evaluation of speech quality (PESQ), the ITU-T Recommendation P.563, and the speech-to-reverberation energy ratio. Our evaluation uses a speech dataset contaminated with convolutive and additive noise, labeled using a crowd-based QoE evaluation, evaluated with Pearson correlation with MOS labels, and mean-squared-error of the estimated MOS. Our proposed approaches outperform the aforementioned instrumental measures, with a fully connected deep neural network using Mel-frequency features providing the best correlation (0.87) and the lowest mean squared error (0.15). Anderson R. Avila, Hannes Gamper, Chandan K. A. Reddy, Ross Cutler, Ivan Tashev, Johannes Gehrke |
ICASSP | 1 |
| 2019 | Blind Channel Response Estimation for Replay Attack Detection
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 1 |
| 2019 | Intrusive Quality Measurement of Noisy and Enhanced Speech based on i-Vector SimilarityabstractIn this paper, the i-vector framework is investigated as an intrusive quality measure for noisy and enhanced speech. While widely used across numerous speech applications, the potential of using i-vectors to summarize the quality of a speech recording has been overlooked. This paper aims to fill this gap. We show that the i-vector framework is well-suited for assessing speech signal quality, surpassing well-established instrumental measures such as the Perceptual Evaluation of Speech Quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA). Three datasets are used in our experiments. First, the TIMIT database is used to train the i-vector extractor on clean speech. To evaluate the proposed method, the noisy speech corpus (NOIZEUS) and the evaluation set of the 2014 IEEE REVERB challenge are used with both containing subjective ratings of perceived quality. Correlations with mean opinion scores (MOS) as high as 0.90 are achieved. Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
QoMEX | 1 |
| 2018 | Investigating Speech Enhancement and Perceptual Quality for Speech Emotion Recognition
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 1 |
| 2018 | Towards a Neuro-Inspired No-Reference Instrumental Quality Measure for Text-to-Speech SystemsabstractSubjective evaluation of synthesized speech is not an easy task as various quality dimensions can be affected, including naturalness, prosody, pronunciation, and continuity, to name a few. Evaluations typically rely on naive listeners, thus more closely representing the consumers of commercial products. As such, while the results of these costly and time consuming tests may provide text-to-speech (TTS) system developers with feedback on the perceived quality and acceptability of their devices, it provides little information on what the source of the problems are and what can be done about it. In this paper, we propose the use of neuroimaging to probe the unconscious cognitive processing of naive listeners as they listen to synthesized speech generated by different systems of varying quality. The obtained neural insights have allowed us to extract a small subset of very relevant features from the speech signals and to use these features to build a simple, no-reference instrumental quality metric specifically tailored to TTS speech. The metric is tested on an unseen dataset and shown to significantly outperform a benchmark algorithm. Anderson R. Avila, Tiago H. Falk |
QoMEX | 2 |
| 2014 | Improving the performance of far-field speaker verification using multi-condition training: the case of GMM-UBM and i-vector systems
Anderson R. Avila, Milton Orlando Sarria-Paja, Francisco J. Fraga, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 1 |