VLDB 2026 Research / reviewers in the wild / expert
Dimitra Emmanouilidou
dblp:55/3103
· DBLP profile ↗
12ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Functional Near-Infrared Spectroscopy Feature Extraction with Application in Workload EstimationabstractFunctional near-infrared spectroscopy (fNIRS) is a brain imaging technique used to estimate neuronal activity by measuring blood oxygenation. In this paper, we develop and evaluate an extensive set of fNIRS features for workload estimation, combining them with respiration and heartbeat signals. Our subject- and session-independent workload estimator is validated in a virtual flight simulator, where workload is objectively assessed based on task performance. We experiment with various regression models and feature ablations, identifying the most effective fNIRS features. The best fNIRS-based model achieves a correlation of 0.3188 with objective workload labels, improving to 0.3268 when incorporating breathing signals. This study demonstrates the value of our novel fNIRS feature set for workload estimation. Elisabeth R. M. Heremans, David Johnston, Dimitra Emmanouilidou, Andre Golard, Ivan Tashev, Ryen W. White |
ICASSP | 3 |
| 2025 | Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio DistanceabstractThe complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations of using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD derived from multiple encoders to provide a more objective measure of musical emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variability and prominence of generated musical emotion, thereby enhancing its realism. Additionally, we investigate the differences in realism between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the issue of emotion bias in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate musical emotion more objectively and effectively. Yuanchao Li, Azalea Gui, Dimitra Emmanouilidou, Hannes Gamper |
ICME | 3 |
| 2024 | Training Audio Captioning Models without AudioabstractAutomated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in general data scarcity for the task. In this work, we address this major limitation and propose an approach to train AAC systems using only text. Our approach leverages the multimodal space of contrastively trained audio-text models, such as CLAP. During training, a decoder generates captions conditioned on the pretrained CLAP text encoder. During inference, the text encoder is replaced with the pretrained CLAP audio encoder. To bridge the modality gap between text and audio embeddings, we propose the use of noise injection or a learnable adapter, during training. We find that the proposed text-only framework performs competitively with stateof-the-art models trained with paired audio, showing that efficient text-to-audio transfer is possible. Finally, we showcase both stylized audio captioning and caption enrichment while training without audio or human-created text captions. Soham Deshmukh, Benjamin Elizalde, Dimitra Emmanouilidou, Bhiksha Raj, Rita Singh, Huaming Wang |
ICASSP | 3 |
| 2024 | Adapting Frechet Audio Distance for Generative Music EvaluationabstractThe growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may be hampered by sample size bias, poor choice of audio embeddings, or the use of biased or low-quality reference sets. We propose reducing sample size bias by extrapolating scores towards an infinite sample size. Through comparisons with MusicCaps labels and a listening test we identify audio embeddings and music reference sets that yield FAD scores well-correlated with acoustic and musical quality. Our results suggest that per-song FAD can be useful to identify outlier samples and predict perceptual quality for a range of music sets and generative models. Finally, we release a toolkit that allows adapting FAD for generative music evaluation. Azalea Gui, Hannes Gamper, Sebastian Braun, Dimitra Emmanouilidou |
ICASSP | 4 |
| 2023 | Multi-View Learning for Speech Emotion Recognition with Categorical Emotion, Categorical Sentiment, and Dimensional ScoresabstractPsychological research has postulated that emotions and sentiment are correlated to dimensional scores of valence, arousal, and dominance. However, the literature of Speech Emotion Recognition focuses on independently predicting the three of them for a given speech audio. In this paper, we evaluate and quantify the predictive power of the dimensional scores towards categorical emotions and sentiment for two publicly available speech emotion datasets. We utilize the three emotional views in a joined multi-view training framework. The views comprise the dimensional scores, emotions categories, and sentiment categories. We present a comparison for each emotional view or combination of, utilizing two general-purpose models for speech-related applications: CNN14 and Wav2Vec2. To our knowledge this is the first time such a joint framework is explored. We found that a joined multi-view training framework can produce results as strong or stronger than models trained independently for each view. Daniel Tompkins, Dimitra Emmanouilidou, Soham Deshmukh, Benjamin Elizalde |
ICASSP | 2 |
| 2021 | Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer InterfaceabstractPeople enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device integrated into a pair of headphones. In a user study of the proposed system, participants were asked to pay attention to one of three musical instruments playing simultaneously from separate spatial directions. We used a stimulus reconstruction method to decode attention from EEG signals. Results show that the proposed system can achieve good decoding accuracy (>70%) while providing superior user-friendliness compared to a traditional EEG setup. Winko W. An, Barbara G. Shinn-Cunningham, Hannes Gamper, Dimitra Emmanouilidou, David Johnston, Mihai Jalobeanu, Edward Cutrell, Andrew D. Wilson, Kuan-Jung Chiang, Ivan Tashev |
ICASSP | 4 |
| 2021 | Design and Comparative Performance of a Robust Lung Auscultation System for Noisy Clinical SettingsabstractChest auscultation is a widely used clinical tool for respiratory disease detection. The stethoscope has undergone a number of transformative enhancements since its invention, including the introduction of electronic systems in the last two decades. Nevertheless, stethoscopes remain riddled with a number of issues that limit their signal quality and diagnostic capability, rendering both traditional and electronic stethoscopes unusable in noisy or non-traditional environments (e.g., emergency rooms, rural clinics, ambulatory vehicles). This work outlines the design and validation of an advanced electronic stethoscope that dramatically reduces external noise contamination through hardware redesign and real-time, dynamic signal processing. The proposed system takes advantage of an acoustic sensor array, an external facing microphone, and on-board processing to perform adaptive noise suppression. The proposed system is objectively compared to six commercially-available acoustic and electronic devices in varying levels of simulated noisy clinical settings and quantified using two metrics that reflect perceptual audibility and statistical similarity, normalized covariance measure (NCM) and magnitude squared coherence (MSC). The analyses highlight the major limitations of current stethoscopes and the significant improvements the proposed system makes in challenging settings by minimizing both distortion of lung sounds and contamination by ambient noise. Ian McLane, Dimitra Emmanouilidou, James E. West, Mounya Elhilali |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Electronic Stethoscope Filtering Mimics the Perceived Sound Characteristics of Acoustic StethoscopeabstractElectronic stethoscopes offer several advantages over conventional acoustic stethoscopes, including noise reduction, increased amplification, and ability to store and transmit sounds. However, the acoustical characteristics of electronic and acoustic stethoscopes can differ significantly, introducing a barrier for clinicians to transition to electronic stethoscopes. This work proposes a method to process lung sounds recorded by an electronic stethoscope, such that the sounds are perceived to have been captured by an acoustic stethoscope. The proposed method calculates an electronic-to-acoustic stethoscope filter by measuring the difference between the average frequency responses of an acoustic and an electronic stethoscope to multiple lung sounds. To validate the method, a change detection experiment was conducted with 51 medical professionals to compare filtered electronic, unfiltered electronic, and acoustic stethoscope lung sounds. Participants were asked to detect when transitions occurred in sounds comprising several sections of the three types of recordings. Transitions between the filtered electronic and acoustic stethoscope sections were detected, on average, by chance (sensitivity index equal to zero) and also detected significantly less than transitions between the unfiltered electronic and acoustic stethoscope sections ( ), demonstrating the effectiveness of the method to filter electronic stethoscopes to mimic an acoustic stethoscope. This processing could incentivize clinicians to adopt electronic stethoscopes by providing a means to shift between the sound characteristics of acoustic and electronic stethoscopes in a single device, allowing for a faster transition to new technology and greater appreciation for the electronic sound quality. Valerie E. Rennoll, Ian McLane, Dimitra Emmanouilidou, James E. West, Mounya Elhilali |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Predicting Word Error Rate for Reverberant SpeechabstractReverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propose predicting ASR performance in terms of the word error rate (WER) directly from acoustic parameters via a polynomial, sigmoidal, or neural network fit, as well as blindly from reverberant speech samples using a convolutional neural network (CNN). We carry out experiments on two state-of-the-art ASR models and a large set of acoustic impulse responses (AIRs). The results confirm C50 and C80 to be highly correlated with WER, allowing WER to be predicted with the proposed fitting approaches. The proposed non-intrusive CNN model outperforms C50-based WER prediction, indicating that WER can be estimated blindly, i.e., directly from the reverberant speech samples without knowledge of the acoustic parameters. Hannes Gamper, Dimitra Emmanouilidou, Sebastian Braun, Ivan Tashev |
ICASSP | 2 |
| 2020 | Supervised Deep Hashing for Efficient Audio Event RetrievalabstractEfficient retrieval of audio events can facilitate real-time implementation of numerous query and search-based systems. This work investigates the potency of different hashing techniques for efficient audio event retrieval. Multiple state-of-the-art weak audio embeddings are employed for this purpose. The performance of four classical unsupervised hashing algorithms is explored as part of off-the-shelf analysis. Then, we propose a partially supervised deep hashing framework that transforms the weak embeddings into a low-dimensional space while optimizing for efficient hash codes. The model uses only a fraction of the available labels and is shown here to significantly improve the retrieval accuracy on two widely employed audio event datasets. The extensive analysis and comparison between supervised and unsupervised hashing methods presented here, give insights on the quantizability of audio embeddings. This work provides a first look in efficient audio event retrieval systems and hopes to set baselines for future research. Arindam Jati, Dimitra Emmanouilidou |
ICASSP | 2 |
| 2010 | Decision support in heart failure through processing of electro- and echocardiograms
Franco Chiarugi, Sara Colantonio, Dimitra Emmanouilidou, Massimo Martinelli, Davide Moroni, Ovidio Salvetti |
Artif. Intell. Medicine | 3 |
| 2009 | A Decision Support System for Aiding Heart Failure ManagementabstractThe purpose of this paper is to present an effective way to achieve a high-level integration of a clinical decision support system in the general process of heart failure care and to discuss the advantages of such an approach. In particular, the relevant and significant medical knowledge and experts' know-how have been modelled according to an ontological formalism extended with a base of rules for inferential reasoning. These have been also combined with advanced analytical tools for data processing. In particular, methods for the segmentation of echocardiographic image sequences and algorithms for ECG processing have been implemented and integrated into the system. Sara Colantonio, Massimo Martinelli, Davide Moroni, Ovidio Salvetti, Franco Chiarugi, Dimitra Emmanouilidou |
ISDA | 6 |