EDBT 2026 Demo / reviewers in the wild / expert
Adria Mallol-Ragolta
dblp:227/5520
· DBLP profile ↗
18ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0001-6855-485XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 9 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Early Detection of ALS in Absence of Speech Impairments with Computer Audition
Adria Mallol-Ragolta, Monica González Machorro, Ricarda von Heynitz, Katia Scherzer, Isabell Cordts, Björn W. Schuller |
AIME (1) | 1 |
| 2025 | ProtoCLAP - Prototypical Contrastive Language-Audio PretrainingabstractWe propose ProtoCLAP, a framework that integrates prototypical representations of the targeted classes in the languageaudio contrastive learning paradigm. Projecting the audio and the language representations in a shared embeddings space – where the prototypical representations are computed –, ProtoCLAP aims to maximise the similarity of the audio embeddings and their corresponding audio and language prototypes, while enforcing the similarity between both prototypical representations. We conduct our experiments on the MASCFLICHT Corpus and the Second DiCOVA Challenge Dataset. ProtoCLAP achieves the best results in three out of the six scenarios investigated. For face mask type and face mask coverage area recognition, ProtoCLAP scores the best Unweighted Average Recall on the test set, 62.8% and 56.7%, respectively. For COVID-19 detection, ProtoCLAP obtains the highest Area Under the Curve on the test set when exploiting the breathing sounds, 84.77%. Adria Mallol-Ragolta, Björn W. Schuller |
ASRU | 1 |
| 2023 | COVID-19 Detection from Speech in Noisy ConditionsabstractWe explore the integration of audio enhancement into a speech-based COVID-19 detection system in an attempt to make speech captured in noisy environments from everyday life useful for the detection of the virus. For this purpose, two multi-task learning approaches are exploited to jointly optimise a front-end speech enhancement model and a subsequent COVID-19 detection model. In comparison to several baseline methods, such as noisy data augmentation, cold cascade of speech enhancement, and COVID-19 models, our proposed solutions are able to recover a substantial percentage of the performance reduction caused by real-world noises. Our best-performing model, which is trained using the synthetic data of the DiCOVA speech corpus and AudioSet environmental backgrounds, can achieve an average AUC of 76.87 % on the test data covering a wide range of noise intensities, which is over 10 % better than a COVID-19 model trained with clean audio. Shuo Liu 0012, Adria Mallol-Ragolta, Björn W. Schuller |
ICASSP | 2 |
| 2023 | The MASCFLICHT Corpus: Face Mask Type and Coverage Area Recognition from Speech
Adria Mallol-Ragolta, Nils Urbach, Shuo Liu 0012, Anton Batliner, Björn W. Schuller |
INTERSPEECH | 1 |
| 2022 | Multi-Type Outer Product-Based Fusion of Respiratory Sounds for Detecting COVID-19abstractComunicació presentada a Interspeech 2022, celebrat del 18 al 22 de setembre de 2022 a Inchon, Corea del Sud. Adria Mallol-Ragolta, Helena Cuesta, Emilia Gómez, Björn W. Schuller |
INTERSPEECH | 1 |
| 2022 | The ACM Multimedia 2022 Computational Paralinguistics Challenge: Vocalisations, Stuttering, Activity, & MosquitoesabstractThe ACM Multimedia 2022 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Vocalisations and Stuttering Sub-Challenges, a classification on human non-verbal vocalisations and speech has to be made; the Activity Sub-Challenge aims at beyond-audio human activity recognition from smartwatch sensor data; and in the Mosquitoes Sub-Challenge, mosquitoes need to be detected. We describe the Sub-Challenges, baseline feature extraction, and classifiers based on the 'usual' ComParE and BoAW features, the auDeep toolkit, and deep feature extraction from pre-trained CNNs using the DeepSpectrum toolkit; in addition, we add end-to-end sequential modelling, and a log-mel-128-BNN. Björn W. Schuller, Anton Batliner, Shahin Amiriparian, Christian Bergler, Maurice Gerczuk, Natalie Holz, Pauline Larrouy-Maestri, Sebastian P. Bayerl, Korbinian Riedhammer, Adria Mallol-Ragolta, Maria Pateraki, Harry Coppock, Ivan Kiskin, Marianne Sinka, Stephen J. Roberts |
ACM Multimedia | 10 |
| 2022 | Face mask recognition from audio: The MASC database and an overview on the mask challenge
Mostafa M. Mohamed, Mina A. Nessiem, Anton Batliner, Christian Bergler, Simone Hantke, Maximilian Schmitt, Alice Baird, Adria Mallol-Ragolta, Vincent Karas, Shahin Amiriparian, Björn W. Schuller |
Pattern Recognit. | 8 |
| 2022 | Capturing Time Dynamics From Speech Using Neural Networks for Surgical Mask DetectionabstractThe importance of detecting whether a person wears a face mask while speaking has tremendously increased since the outbreak of SARS-CoV-2 (COVID-19), as wearing a mask can help to reduce the spread of the virus and mitigate the public health crisis. Besides affecting human speech characteristics related to frequency, face masks cause temporal interferences in speech, altering the pace, rhythm, and pronunciation speed. In this regard, this paper presents two effective neural network models to detect surgical masks from audio. The proposed architectures are both based on Convolutional Neural Networks (CNNs), chosen as an optimal approach for the spatial processing of the audio signals. One architecture applies a Long Short-Term Memory (LSTM) network to model the time-dependencies. Through an additional attention mechanism, the LSTM-based architecture enables the extraction of more salient temporal information. The other architecture (named ConvTx) retrieves the relative position of a sequence through the positional encoder of a transformer module. In order to assess to which extent both architectures can complement each other when modelling temporal dynamics, we also explore the combination of LSTM and Transformers in three hybrid models. Finally, we also investigate whether data augmentation techniques, such as, using transitions between audio frames and considering gender-dependent frameworks might impact the performance of the proposed architectures. Our experimental results show that one of the hybrid models achieves the best performance, surpassing existing state-of-the-art results for the task at hand. Shuo Liu 0012, Adria Mallol-Ragolta, Tianhao Yan, Kun Qian 0003, Emilia Parada-Cabaleiro, Bin Hu 0001, Björn W. Schuller |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | harAGE: A Novel Multimodal Smartwatch-based Dataset for Human Activity RecognitionabstractThis work introduces the harAGEdataset: a novel multimodal smartwatch-based dataset for Human Activity Recognition (HAR) with more than 17 hours of data collected from 19 participants using a Garmin Vivoactive 3 device. The dataset contains samples from resting, lying, sitting, standing, washing hands, walking, running, stairs climbing, strength workout, flexibility workout, and cycling activities. The resting activity, excluded from the set of activities to recognise, was explicitly conducted while avoiding stressors and external stimuli, so the data collected can be used to compute the personal, baseline heart rate at rest. We also present the HAR-based models trained using the accelerometer data to recognise different sets of activities. Specifically, we focus on different strategies to combine, fuse, and enrich the accelerometer measurements, so they can be used end-to-end. Model performances are assessed following a Leave-One-Subject-Out Cross-Validation (LOSO-CV) approach, and we use the Unweighted Average Recall (UAR) as the evaluation metric to compare the ground truth and the inferred information. The best UAR score of 98.1 % is obtained when recognising the static and the dynamic activities, excluding the samples corresponding to the washing hands, strength workout, and flexibility workout activities. When recognising the specific activities included in these two sets, the model with the best performance scores a UAR of 70.1 %. Finally, when recognising all the activities considered in the harAGEdataset, the highest UAR achieved is 64.3 %. Adria Mallol-Ragolta, Anastasia Semertzidou, Maria Pateraki, Björn W. Schuller |
FG | 1 |
| 2021 | A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion RecognitionabstractNotwithstanding the significant advancements in the field of deep learning, the basic long short-term memory (LSTM) or Gated Recurrent Unit (GRU) units have largely remained unchanged and unexplored. There are several possibilities in advancing the state-of-art by rightly adapting and enhancing the various elements of these units. Activation functions are one such key element. In this work, we explore using diverse activation functions within GRU and bi-directional GRU (BiGRU) cells in the context of speech emotion recognition (SER). We also propose a novel Attention ReLU GRU (AR-GRU) that employs attention-based Rectified Linear Unit (AReLU) activation within GRU and BiGRU cells. We demonstrate the effectiveness of AR-GRU on one exemplary application using the recently proposed network for SER namely Interaction-Aware Attention Network (IAAN). Our proposed method utilising AR-GRU within this network yields significant performance gain and achieves an unweighted accuracy of 68.3% (2% over the baseline) and weighted accuracy of 66.9 % (2.2 % absolute over the baseline) in four class emotion recognition on the IEMOCAP database. Srividya Tirunellai Rajamani, Kumar T. Rajamani, Adria Mallol-Ragolta, Shuo Liu 0012, Björn W. Schuller |
ICASSP | 3 |
| 2021 | Cough-Based COVID-19 Detection with Contextual Attention Convolutional Neural Networks and Gender InformationabstractThe aim of this contribution is to automatically detect COVID-19 patients by analysing the acoustic information embedded in coughs.COVID-19 affects the respiratory system, and, consequently, respiratory-related signals have the potential to contain salient information for the task at hand.We focus on analysing the spectrogram representations of cough samples with the aim to investigate whether COVID-19 alters the frequency content of these signals.Furthermore, this work also assesses the impact of gender in the automatic detection of COVID-19.To extract deep-learnt representations of the spectrograms, we compare the performance of a cough-specific, and a Resnet18 pre-trained Convolutional Neural Network (CNN).Additionally, our approach explores the use of contextual attention, so the model can learn to highlight the most relevant deep-learnt features extracted by the CNN.We conduct our experiments on the dataset released for the Cough Sound Track of the DICOVA 2021 Challenge.The best performance on the test set is obtained using the Resnet18 pre-trained CNN with contextual attention, which scored an Area Under the Curve (AUC) of 70.91 % at 80 % sensitivity. Adria Mallol-Ragolta, Helena Cuesta, Emilia Gómez, Björn W. Schuller |
Interspeech | 1 |
| 2021 | Frustration recognition from speech during game interaction using wide residual networksabstractAlthough frustration is a common emotional reaction while playing games, an excessive level of frustration can negatively impact a user's experience, discouraging them from further game interactions. The automatic detection of frustration can enable the development of adaptive systems that can adapt a game to a user's specific needs through real-time difficulty adjustment, thereby optimizing the player's experience and guaranteeing game success. To this end, we present a speech-based approach for the automatic detection of frustration during game interactions, a specific task that remains underexplored in research. The experiments were performed on the Multimodal Game Frustration Database (MGFD), an audiovisual dataset—collected within the Wizard-of-Oz framework—that is specially tailored to investigate verbal and facial expressions of frustration during game interactions. We explored the performance of a variety of acoustic feature sets, including Mel-Spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), and the low-dimensional knowledge-based acoustic feature set eGeMAPS. Because of the continual improvements in speech recognition tasks achieved by the use of convolutional neural networks (CNNs), unlike the MGFD baseline, which is based on the Long Short-Term Memory (LSTM) architecture and Support Vector Machine (SVM) classifier—in the present work, we consider typical CNNs, including ResNet, VGG, and AlexNet. Furthermore, given the unresolved debate on the suitability of shallow and deep networks, we also examine the performance of two of the latest deep CNNs: WideResNet and EfficientNet. Our best result, achieved with WideResNet and Mel-Spectrogram features, increases the system performance from 58.8% unweighted average recall (UAR) to 93.1% UAR for speech-based automatic frustration recognition. Meishu Song, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Zijiang Yang 0007, Shuo Liu 0012, Zhao Ren, Ziping Zhao 0001, Björn W. Schuller |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | Latent-Based Adversarial Neural Networks for Facial Affect EstimationsabstractThere is a growing interest in affective computing research nowadays given its crucial role in bridging humans with computers. This progress has recently been accelerated due to the emergence of bigger dataset. One recent advance in this field is the use of adversarial learning to improve model learning through augmented samples. However, the use of latent features, which is feasible through adversarial learning, is not largely explored, yet. This technique may also improve the performance of affective models, as analogously demonstrated in related fields, such as computer vision. To expand this analysis, in this work, we explore the use of latent features through our proposed adversarial-based networks for valence and arousal recognition in the wild. Specifically, our models operate by aggregating several modalities to our discriminator, which is further conditioned to the extracted latent features by the generator. Our experiments on the recently released SEWA dataset suggest the progressive improvements of our results. Finally, we show our competitive results on the Affective Behavior Analysis in-the-Wild (ABAW) challenge dataset. Decky Aspandi, Adria Mallol-Ragolta, Björn W. Schuller, Xavier Binefa |
FG | 2 |
| 2020 | A Curriculum Learning Approach for Pain Intensity Recognition from Facial ExpressionsabstractThe high prevalence of chronic pain in society raises the need to develop new digital tools that can automatically and objectively assess pain intensity in individuals. These tools can contribute to an optimisation of clinical resources, as they offer cost-effective solutions for early detection, continuous monitoring, and treatment personalisation by utilising Artificial Intelligence techniques. In this work, we present our contribution to the Pain Intensity Estimation from Facial Expressions task of the EMOPAIN 2020 Challenge. Specifically, we compare the performance of Recurrent Neural Networks trained with standard or Curriculum Learning (CL) approaches to predict the pain intensity level of individuals reported in an 11-point scale from facial expressions. The results obtained using the test partition support the use of CL-based approaches in the automatic prediction of pain from facial features. The best model trained using a CL approach achieved a Concordance Correlation Coefficient (CCC) of 0.196 in the test partition, while the model trained using a standard approach, without CL, achieved a CCC of 0.174. In terms of CCC, these results respectively represent an improvement of 0.136 and 0.114 on the best results of the baseline system reported by the Challenge organisers using the test partition. Adria Mallol-Ragolta, Shuo Liu 0012, Nicholas Cummins, Björn W. Schuller |
FG | 1 |
| 2020 | An Investigation of Cross-Cultural Semi-Supervised Learning for Continuous Affect RecognitionabstractOne of the keys for supervised learning techniques to succeed resides in the access to vast amounts of labelled training data. The process of data collection, however, is expensive, time- consuming, and application dependent. In the current digital era, data can be collected continuously. This continuity renders data annotation into an endless task, which potentially, in problems such as emotion recognition, requires annotators with different cultural backgrounds. Herein, we study the impact of utilising data from different cultures in a semi-supervised learning ap- proach to label training material for the automatic recognition of arousal and valence. Specifically, we compare the performance of culture-specific affect recognition models trained with man- ual or cross-cultural automatic annotations. The experiments performed in this work use the dataset released for the Cross- cultural Emotion Sub-challenge of the Audio/Visual Emotion Challenge (AVEC) 2019. The results obtained convey that the cultures used for training impact on the system performance. Furthermore, in most of the scenarios assessed, affect recogni- tion models trained with hybrid solutions, combining manual and automatic annotations, surpass the baseline model, which was exclusively trained with manual annotations. Adria Mallol-Ragolta, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 1 |
| 2019 | Performance Analysis of Unimodal and Multimodal Models in Valence-Based Empathy RecognitionabstractThe human ability to empathise is a core aspect of successful interpersonal relationships. In this regard, human-robot interaction can be improved through the automatic perception of empathy, among other human attributes, allowing robots to affectively adapt their actions to interactants' feelings in any given situation. This paper presents our contribution to the generalised track of the One-Minute Gradual (OMG) Empathy Prediction Challenge by describing our approach to predict a listener's valence during semi-scripted actor-listener interactions. We extract visual and acoustic features from the interactions and feed them into a bidirectional long short-term memory network to capture the time-dependencies of the valence-based empathy during the interactions. Generalised and personalised unimodal and multimodal valence-based empathy models are then trained to assess the impact of each modality on the system performance. Furthermore, we analyse if intra-subject dependencies on empathy perception affect the system performance. We assess the models by computing the concordance correlation coefficient (CCC) between the predicted and self-annotated valence scores. The results support the suitability of employing multimodal data to recognise participants' valence-based empathy during the interactions, and highlight the subject-dependency of empathy. In particular, we obtained our best result with a personalised multimodal model, which achieved a CCC of 0.11 on the test set. Adria Mallol-Ragolta, Maximilian Schmitt, Alice Baird, Nicholas Cummins, Björn W. Schuller |
FG | 1 |
| 2019 | A Hierarchical Attention Network-Based Approach for Depression Detection from Transcribed Clinical InterviewsabstractThe high prevalence of depression in society has given rise to a need for new digital tools that can aid its early detection. Among other effects, depression impacts the use of language. Seeking to exploit this, this work focuses on the detection of depressed and non-depressed individuals through the analysis of linguistic information extracted from transcripts of clinical interviews with a virtual agent. Specifically, we investigated the advantages of employing hierarchical attention-based networks for this task. Using Global Vectors (GloVe) pretrained word embedding models to extract low-level representations of the words, we compared hierarchical local-global attention networks and hierarchical contextual attention networks. We performed our experiments on the Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WoZ) dataset, which contains audio, visual, and linguistic information acquired from participants during a clinical session. Our results using the DAIC-WoZ test set indicate that hierarchical contextual attention networks are the most suitable configuration to detect depression from transcripts. The configuration achieves an Unweighted Average Recall (UAR) of .66 using the test set, surpassing our baseline, a Recurrent Neural Network that does not use attention. Adria Mallol-Ragolta, Ziping Zhao 0001, Lukas Stappen, Nicholas Cummins, Björn W. Schuller |
INTERSPEECH | 1 |
| 2018 | A Multimodal Approach for Predicting Changes in PTSD Symptom SeverityabstractThe rising prevalence of mental illnesses is increasing the demand for new digital tools to support mental wellbeing. Numerous collaborations spanning the fields of psychology, machine learning and health are building such tools. Machine-learning models that estimate effects of mental health interventions currently rely on either user self-reports or measurements of user physiology. In this paper, we present a multimodal approach that combines self-reports from questionnaires and skin conductance physiology in a web-based trauma-recovery regime. We evaluate our models on the EASE multimodal dataset and create PTSD symptom severity change estimators at both total and cluster-level. We demonstrate that modeling the PTSD symptom severity change at the total-level with self-reports can be statistically significantly improved by the combination of physiology and self-reports or just skin conductance measurements. Our experiments show that PTSD symptom cluster severity changes using our novel multimodal approach are significantly better modeled than using self-reports and skin conductance alone when extracting skin conductance features from triggers modules for avoidance, negative alterations in cognition & mood and alterations in arousal & reactivity symptoms, while it performs statistically similar for intrusion symptom. Adria Mallol-Ragolta, Svati Dhamija, Terrance E. Boult |
ICMI | 1 |