VLDB 2026 Research / reviewers in the wild / expert
Dimitris Sgouropoulos
dblp:151/7128 · also Dimitrios Sgouropoulos
· DBLP profile ↗
7ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RobuSER: A Robustness Benchmark for Speech Emotion RecognitionabstractThe recent surge in deep learning has improved Speech Emotion Recognition (SER) model performance; however, ensuring robustness across diverse scenarios beyond the training dataset remains a problem. This challenge becomes pronounced in real-world situations characterized by noisy conditions, where model adaptability to unclean data is crucial. Despite ongoing efforts to develop noise-robust models, the lack of standardized evaluation protocols hampers fair comparisons among different models. This paper tackles this issue by introducing Robuser, a benchmarking procedure designed specifically for evaluating the robustness of SER models under noise. Robuser is a comprehensive open-source benchmark that can be applied to any speech dataset, focusing on diverse corruption types in two pivotal dimensions: additive background noise and various signal distortion corruptions, each in varying levels of severity. Furthermore, through the evaluation of a state-of-the-art SER model against this benchmark, we offer quantitative insights into the impact of the different corruption types and severity levels on performance. The baseline model reveals a notable performance degradation of up to 22.77% in Unweighted Accuracy (UA) and 20.32% in Weighted Accuracy (WA) on corrupted IEMOCAP, underscoring the substantial room for improvement in this domain. Our code is openly available at the following URL: https://github.com/BehavioralSignalTechnologies/ser_robustness.git Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Sofia Eleftheriou, Petros Mitseas, Dimitris Sgouropoulos, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan |
ACII | 6 |
| 2024 | Emotion-Aware Speech Popularity Prediction: A Use-Case on TED TalksabstractIn the context of the ever-growing influence of social media, understanding and predicting the popularity of content has become crucial for creators and marketers alike. Our research addresses this need by introducing a method to forecast the success of oral presentations, focusing on the nuanced use of paralinguistic features and insights derived from speech emotion recognition models. This innovative approach is designed to enhance verbal communication skills by providing public speakers with targeted feedback. We leverage a dataset of 2,462 TED talk videos, complete with metadata such as user comments, tags, and views, to establish a set of four objective metrics for determining presentation popularity. These metrics form the foundation of our analysis, enabling us to evaluate the efficacy of our predictive methodology. By integrating audio-based emotional cues with text-based content analysis we showcase the capability of the proposed speech analytics system to capture user assessments of presentation quality. This research highlights the role of emotional expression in speech as a component of content's appeal, advocating for a broader analytical perspective beyond just text-only analysis. It suggests new directions for improving the impact of public speaking and calls for further investigation into multimodal content analysis, aiming to deepen our understanding of audience engagement on social media and content delivery platforms. Dimitris Sgouropoulos, Petros Mitseas, Sofia Eleftheriou, Theodoros Giannakopoulos, Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Athanasios Katsamanis, Shri Narayanan |
ACII | 1 |
| 2023 | Cross-Lingual Features for Alzheimer's Dementia Detection from Speech
Thomas Melistas, Lefteris Kapelonis, Nikolaos Antoniou, Petros Mitseas, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 5 |
| 2022 | Audio and ASR-based Filled Pause DetectionabstractFilled pauses (or fillers) are the most common form of speech disfluencies and they can be recognized as hesitation markers (“um”, “uh” and “er”) made by speakers, usually to gain extra time while thinking their next words. Filled pauses are very frequent in spontaneous speech. Their detection is therefore rather important for two basic reasons: (a) their existence influences the performance of individual components, like Automatic Speech Recognition system (ASR), in human-machine interaction and (b) their frequency can characterize the overall speech quality of a particular speaker, as it can be strongly associated with the speaker's confidence. Despite that, only limited work has been published for the detection of filled pauses in speech, especially through audio. In this work, we propose a framework for filled pause detection using both audio and textual information. For the audio modality, we transfer knowledge from a plethora of supervised tasks, such as emotion or speaking rate, using Convolutional Neural Networks (CNNs). For the text modality, we develop a temporal Recurrent Neural Network (RNN) method that takes into account textual information derived from an ASR system. In addition, the proposed transfer learning approach for the audio classifier leads to better results when benchmarked on our internal dataset for which the text is not transcribed but estimated by an ASR system. In this case, a simple late fusion approach boosts the performance even further. This proves that the audio approach is suitable for real-world applications where the transcribed text is not available and has to leverage imperfect ASR results, or even the absence of textual information (to reduce computational cost). Aggelina Chatziagapi, Dimitris Sgouropoulos, Constantinos Karouzos, Thomas Melistas, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
ACII | 2 |
| 2019 | Using Oliver API for emotion-aware movie content characterizationabstractThis paper demonstrates the utilization of Oliver11https://behavioralsignals.com/oliver/, the speech emotion recognition (SER) API created by Behavioral Signals, in the context of a movie content visualization application. Oliver API provides an emotion recognition as-a-service solution that can be accessed via a Web API. In this work, we demonstrate how one can send sound recordings from famous movies, retrieve respective emotional descriptors and use simple aggregations on these descriptors to visualize movie content. We have compiled a dataset of 60 movies, categorized over 8 directors. The classification examples included in this paper indicate the ability of simple emotion aggregations to discriminate between movie directors. In order for others to also experiment with the output of both the API's Emotional and Automatic Speech Recognition, the responses are provided as JSON files in this link: https://tinyurl.com/yxeqvvy2. Theodoros Giannakopoulos, Spiros Dimopoulos, Georgios Pantazopoulos, Aggelina Chatziagapi, Dimitris Sgouropoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
CBMI | 5 |
| 2019 | Data Augmentation Using GANs for Speech Emotion Recognition
Aggelina Chatziagapi, Georgios Paraskevopoulos, Dimitris Sgouropoulos, Georgios Pantazopoulos, Malvina Nikandrou, Theodoros Giannakopoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 3 |
| 2015 | Fusing multiple audio sensors for acoustic event detectionabstractThis paper presents an Internet-of-Things approach to fusing audio sensors towards the detection of audio events, in a meeting room scenario. The different types of audio sensors (microphones) and respective individual audio analysis modules are incorporated within the context of an IoT framework that follows a message-oriented architecture. Each individual audio analysis module is composed by a feature extraction stage and a Support-Vector-Machine (SVM) classifier. A fusion module is also adopted to combine the individual sensor-level decisions, in order to extract the final classification decision. A detailed experimental evaluation on a publicly available real-world dataset proves a rather significant performance boosting in terms of overall classification accuracy. In addition, the proposed architecture enables an easy-to-use training procedure that can easily handle any number of audio sensors and respective classifiers without any prior knowledge of the room's geometry or any other constraints regarding the topological condition of the sensors. Giorgos Siantikos, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Evaggelos Spyrou |
ISPA | 2 |