VLDB 2026 Research / reviewers in the wild / expert
Theodoros Giannakopoulos
dblp:64/1130 · also Theodore Giannakopoulos
· DBLP profile ↗
33ranked-venue papers
11as first author
10since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RobuSER: A Robustness Benchmark for Speech Emotion RecognitionabstractThe recent surge in deep learning has improved Speech Emotion Recognition (SER) model performance; however, ensuring robustness across diverse scenarios beyond the training dataset remains a problem. This challenge becomes pronounced in real-world situations characterized by noisy conditions, where model adaptability to unclean data is crucial. Despite ongoing efforts to develop noise-robust models, the lack of standardized evaluation protocols hampers fair comparisons among different models. This paper tackles this issue by introducing Robuser, a benchmarking procedure designed specifically for evaluating the robustness of SER models under noise. Robuser is a comprehensive open-source benchmark that can be applied to any speech dataset, focusing on diverse corruption types in two pivotal dimensions: additive background noise and various signal distortion corruptions, each in varying levels of severity. Furthermore, through the evaluation of a state-of-the-art SER model against this benchmark, we offer quantitative insights into the impact of the different corruption types and severity levels on performance. The baseline model reveals a notable performance degradation of up to 22.77% in Unweighted Accuracy (UA) and 20.32% in Weighted Accuracy (WA) on corrupted IEMOCAP, underscoring the substantial room for improvement in this domain. Our code is openly available at the following URL: https://github.com/BehavioralSignalTechnologies/ser_robustness.git Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Sofia Eleftheriou, Petros Mitseas, Dimitris Sgouropoulos, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan |
ACII | 8 |
| 2024 | Emotion-Aware Speech Popularity Prediction: A Use-Case on TED TalksabstractIn the context of the ever-growing influence of social media, understanding and predicting the popularity of content has become crucial for creators and marketers alike. Our research addresses this need by introducing a method to forecast the success of oral presentations, focusing on the nuanced use of paralinguistic features and insights derived from speech emotion recognition models. This innovative approach is designed to enhance verbal communication skills by providing public speakers with targeted feedback. We leverage a dataset of 2,462 TED talk videos, complete with metadata such as user comments, tags, and views, to establish a set of four objective metrics for determining presentation popularity. These metrics form the foundation of our analysis, enabling us to evaluate the efficacy of our predictive methodology. By integrating audio-based emotional cues with text-based content analysis we showcase the capability of the proposed speech analytics system to capture user assessments of presentation quality. This research highlights the role of emotional expression in speech as a component of content's appeal, advocating for a broader analytical perspective beyond just text-only analysis. It suggests new directions for improving the impact of public speaking and calls for further investigation into multimodal content analysis, aiming to deepen our understanding of audience engagement on social media and content delivery platforms. Dimitris Sgouropoulos, Petros Mitseas, Sofia Eleftheriou, Theodoros Giannakopoulos, Antonia Petrogianni, Lefteris Kapelonis, Nikolaos Antoniou, Athanasios Katsamanis, Shri Narayanan |
ACII | 4 |
| 2024 | Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based LossesabstractWhat do different contrastive learning (CL) losses actually optimize for? Although multiple CL methods have demonstrated remarkable representation learning capabilities, the differences in their inner workings remain largely opaque. In this work, we analyse several CL families and prove that, under certain conditions, they admit the same minimisers when optimizing either their batch-level objectives or their expectations asymptotically. In both cases, an intimate connection with the hyperspherical energy minimisation (HEM) problem resurfaces. Drawing inspiration from this, we introduce a novel CL objective, coined Decoupled Hyperspherical Energy Loss (DHEL). DHEL simplifies the problem by decoupling the target hyperspherical energy from the alignment of positive examples while preserving the same theoretical guarantees. Going one step further, we show the same results hold for another relevant CL family, namely kernel contrastive learning (KCL), with the additional advantage of the expected loss being independent of batch size, thus identifying the minimisers in the non-asymptotic regime. Empirical results demonstrate improved downstream performance and robustness across combinations of different batch sizes and hyperparameters and reduced dimensionality collapse, on several computer vision datasets. Panagiotis Koromilas, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis |
ICML | 3 |
| 2023 | Unsupervised Temporal Analysis of Mouse VocalizationsabstractMice communicate using ultrasonic vocalizations (USVs) that vary according to parameters such as sex, genetic background, and environmental stimuli. The study of USVs production provides useful models of the underlying neurobiology mechanisms of human speech and many methods exist to detect USVs in mice recordings. In order to achieve temporal analysis of these vocalizations, one must first group them into categories. This grouping of USVs is a rather demanding task considering the high volume of USVs even in small recordings. Most existing tools can recognize a predefined number of categories and offer no temporal analysis capabilities. In this work, we used the open-source software Analysis of Mouse VOcal Communication (AMVOC) for USVs detection and propose an unsupervised learning approach based on features extracted from a Convolutional Autoencoder (CAE). For the evaluation of the CAE approach we built a benchmark dataset. Using USVs transition matrices we propose three metrics that quantity differences in the temporal structure between different recordings. We evaluate these metrics using a dataset from mice with a FoxP2 mutation, a gene involved in speech function. In this way, a researcher can perform batch comparisons of the temporal structure of recordings, extract insights and identify differences in syntax composition prior to more thorough analysis. Christodoulos Bochalis, César D. M. Vargas, Erich D. Jarvis, Theodoros Giannakopoulos |
CIBCB | 4 |
| 2023 | Designing and Evaluating Speech Emotion Recognition Systems: A Reality Check Case Study with IEMOCAPabstractThere is an imminent need for guidelines and standard test sets to allow direct and fair comparisons of speech emotion recognition (SER). While resources, such as the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database, have emerged as widely-adopted reference corpora for researchers to develop and test models for SER, published work reveals a wide range of assumptions and variety in its use that challenge reproducibility and generalization. Based on a critical review of the latest advances in SER using IEMOCAP as the use case, our work aims at two contributions: First, using an analysis of the recent literature, including assumptions made and metrics used therein, we provide a set of SER evaluation guidelines. Second, using recent publications with open-sourced implementations, we focus on reproducibility assessment in SER. Nikolaos Antoniou, Athanasios Katsamanis, Theodoros Giannakopoulos, Shri Narayanan |
ICASSP | 3 |
| 2023 | MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor MethodsabstractDespite the considerable research output on Multimodal Learning for Affect-related tasks, most of the current methods are very complex in terms of the number of trainable parameters, and thus do not constitute effective solutions for real-life applications. In this work we try to alleviate this gap in the literature by introducing the Multimodal Attention Tensor Regression (MMATR) network, a lightweight model that is based on: (i) a static input representation (2D matrix of dimensions time × features) for each modality, which helps to avoid high-parameterized sequential models by incorporating a CNN, (ii) the replacement of the usual pooling and flattening operations as well as the linear layers by tensor contraction and tensor regression layers that are able to reduce the number of parameters, while keeping the high-order structure of the multimodal data, and (iii) a bimodal attention layer that learns multimodal co-occurrences. By a set of experiments comparing with a variety of state-of-the-art techniques, we show that the proposed MMATR can achieve results competitive to the state-of-the-art in the task of Multimodal Sentiment Analysis, albeit having four orders of magnitude fewer parameters. Panagiotis Koromilas, Mihalis A. Nicolaou, Theodoros Giannakopoulos, Yannis Panagakis |
ICASSP | 3 |
| 2023 | Cross-Lingual Features for Alzheimer's Dementia Detection from Speech
Thomas Melistas, Lefteris Kapelonis, Nikolaos Antoniou, Petros Mitseas, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
INTERSPEECH | 6 |
| 2022 | Audio and ASR-based Filled Pause DetectionabstractFilled pauses (or fillers) are the most common form of speech disfluencies and they can be recognized as hesitation markers (“um”, “uh” and “er”) made by speakers, usually to gain extra time while thinking their next words. Filled pauses are very frequent in spontaneous speech. Their detection is therefore rather important for two basic reasons: (a) their existence influences the performance of individual components, like Automatic Speech Recognition system (ASR), in human-machine interaction and (b) their frequency can characterize the overall speech quality of a particular speaker, as it can be strongly associated with the speaker's confidence. Despite that, only limited work has been published for the detection of filled pauses in speech, especially through audio. In this work, we propose a framework for filled pause detection using both audio and textual information. For the audio modality, we transfer knowledge from a plethora of supervised tasks, such as emotion or speaking rate, using Convolutional Neural Networks (CNNs). For the text modality, we develop a temporal Recurrent Neural Network (RNN) method that takes into account textual information derived from an ASR system. In addition, the proposed transfer learning approach for the audio classifier leads to better results when benchmarked on our internal dataset for which the text is not transcribed but estimated by an ASR system. In this case, a simple late fusion approach boosts the performance even further. This proves that the audio approach is suitable for real-world applications where the transcribed text is not available and has to leverage imperfect ASR results, or even the absence of textual information (to reduce computational cost). Aggelina Chatziagapi, Dimitris Sgouropoulos, Constantinos Karouzos, Thomas Melistas, Theodoros Giannakopoulos, Athanasios Katsamanis, Shri Narayanan |
ACII | 5 |
| 2022 | Real-time Feasibility of a Human Intention Method Evaluated Through a Competitive Human-Robot Reaching GameabstractPredicting human behavior is a necessary robot ability for safe and fluent human-robot collaboration in shared workspaces. Robots should recognize human intended actions by combining information from ongoing movements and other environmental cues. In many cases, visual sensors might be required for obtaining information regarding the human move-ment. While using cameras, especially a single one, offers several benefits, the information captured is noisy and of relative low frequency considering the requirements for intention prediction. The purpose of this study was to evaluate the feasibility of obtaining real-time intention prediction and using it timely for robot action, when human behavior is observed by a single RGB-D camera. Visual information is used to obtain human joint data using Openpose. Based on this we then construct appropriate features and train several Machine Learning models. We evaluate the feasibility of timely robot action using a competitive human-robot game. The results show that a prediction available at about 288ms is early enough to enable timely robot action provided that the robot has to act on objects that are no farther than 10 cm away. Athanasios C. Tsitos, Maria Dagioglou, Theodoros Giannakopoulos |
HRI | 3 |
| 2022 | A Dataset for Speech Emotion Recognition in Greek Theatrical PlaysabstractMachine learning methodologies can be adopted in cultural applications and propose new ways to distribute or even present the cultural content to the public. For instance, speech analytics can be adopted to automatically generate subtitles in theatrical plays, in order to (among other purposes) help people with hearing loss. Apart from a typical speech-to-text transcription with Automatic Speech Recognition (ASR), Speech Emotion Recognition (SER) can be used to automatically predict the underlying emotional content of speech dialogues in theatrical plays, and thus to provide a deeper understanding how the actors utter their lines. However, real-world datasets from theatrical plays are not available in the literature. In this work we present GreThE, the Greek Theatrical Emotion dataset, a new publicly available data collection for speech emotion recognition in Greek theatrical plays. The dataset contains utterances from various actors and plays, along with respective valence and arousal annotations. Towards this end, multiple annotators have been asked to provide their input for each speech recording and inter-annotator agreement is taken into account in the final ground truth generation. In addition, we discuss the results of some indicative experiments that have been conducted with machine and deep learning frameworks, using the dataset, along with some widely used databases in the field of speech emotion recognition. Maria Moutti, Sofia Eleftheriou, Panagiotis Koromilas, Theodoros Giannakopoulos |
LREC | 4 |
| 2019 | Using Oliver API for emotion-aware movie content characterizationabstractThis paper demonstrates the utilization of Oliver11https://behavioralsignals.com/oliver/, the speech emotion recognition (SER) API created by Behavioral Signals, in the context of a movie content visualization application. Oliver API provides an emotion recognition as-a-service solution that can be accessed via a Web API. In this work, we demonstrate how one can send sound recordings from famous movies, retrieve respective emotional descriptors and use simple aggregations on these descriptors to visualize movie content. We have compiled a dataset of 60 movies, categorized over 8 directors. The classification examples included in this paper indicate the ability of simple emotion aggregations to discriminate between movie directors. In order for others to also experiment with the output of both the API's Emotional and Automatic Speech Recognition, the responses are provided as JSON files in this link: https://tinyurl.com/yxeqvvy2. Theodoros Giannakopoulos, Spiros Dimopoulos, Georgios Pantazopoulos, Aggelina Chatziagapi, Dimitris Sgouropoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
CBMI | 1 |
| 2019 | Data Augmentation Using GANs for Speech Emotion Recognition
Aggelina Chatziagapi, Georgios Paraskevopoulos, Dimitris Sgouropoulos, Georgios Pantazopoulos, Malvina Nikandrou, Theodoros Giannakopoulos, Athanasios Katsamanis, Alexandros Potamianos, Shri Narayanan |
INTERSPEECH | 6 |
| 2019 | Unsupervised Low-Rank Representations for Speech Emotion RecognitionabstractWe examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptors and their aggregations (IS10) and one modeling recurrence dynamics of speech (RQA), as well as their fusion. We report speech emotion recognition (SER) results for learned representations on two databases using different classification methods. Classification with low-dimensional representations yields performance improvement in a variety of settings. This indicates that dimensionality reduction is an effective way to combat the curse of dimensionality for SER. Visualization of features in two dimensions provides insight into discriminatory abilities of reduced feature sets. Georgios Paraskevopoulos, Efthymios Tzinis, Nikolaos Ellinas, Theodoros Giannakopoulos, Alexandros Potamianos |
INTERSPEECH | 4 |
| 2019 | Athens Urban Soundscape (ATHUS): A Dataset for Urban Soundscape Quality Recognition
Theodoros Giannakopoulos, Margarita Orfanidi, Stavros J. Perantonis |
MMM (1) | 1 |
| 2018 | Enhanced movie content similarity based on textual, auditory and visual information
Konstantinos Bougiatiotis, Theodoros Giannakopoulos |
Expert Syst. Appl. | 2 |
| 2018 | Speech-music discrimination using deep visual feature extractors
Michalis Papakostas, Theodoros Giannakopoulos |
Expert Syst. Appl. | 2 |
| 2018 | Curriculum learning of visual attribute clusters for multi-task classification
Nikolaos Sarafianos, Theodoros Giannakopoulos, Christophoros Nikou, Ioannis A. Kakadiaris |
Pattern Recognit. | 2 |
| 2017 | Towards predicting task performance from EEG signalsabstractSmart wearable devices have lead to an increased need for processing and sharing large streams of physiological data in real-time. Modern Human-Machine Interaction (HMI) systems, especially applications designed for user training and assessment (e.g., educational or smart-rehabilitation systems), should be able to track and monitor those signals and adapt their parameters accordingly in order to optimally facilitate the special needs of each individual. Towards this end, we propose a passive Brain-Computer Interface (BCI), using a wireless non-intrusive EEG sensor under a robot assisted training task designed for cognitive assessment. As part of this ongoing work, we demonstrate our initial results on predicting user's task performance, from the EEG signals, before task completion. Our findings highlight the potentials of our hypotheses as we achieve a maximum accuracy rate equal to 74% when evaluated on 69 real subjects. Michalis Papakostas, Konstantinos Tsiakas, Theodoros Giannakopoulos, Fillia Makedon |
IEEE BigData | 3 |
| 2016 | Computation and communication challenges to deploy robots in assisted living environments
Georgios Keramidas, Christos P. Antonopoulos, Nikos S. Voros, Fynn Schwiegelshohn, Philipp Wehner, Jens Rettkowski, Diana Göhringer, Michael Hübner 0001, Stasinos Konstantopoulos, Theodoros Giannakopoulos, Vangelis Karkaletsis, Evaggelinos P. Mariatos |
DATE | 10 |
| 2016 | Audio-visual speaker diarization using fisher linear semi-discriminant analysis
Nikolaos Sarafianos, Theodoros Giannakopoulos, Sergios Petridis |
Multim. Tools Appl. | 2 |
| 2015 | Visual sentiment analysis for brand monitoring enhancementabstractBrand monitoring and reputation management are vital tasks in all modern business intelligence frameworks. However, recent related technologies rely mostly on the textual aspect of online content, in order to extract the underlying sentiment with respect to particular brands. In this work, we demonstrate the sentiment analysis method in the context of a brand monitoring framework, breaking the text-only barrier in the field. Towards this end, a wide range of visual features is extracted, some of which focus on the underlying semiotics and aesthetics of the images. In addition, we employ textual information embedded in the images under study, by adopting text mining techniques that focus on extracting sentiment. We evaluate the classification task for the particular binary task (negative vs positive sentiment) and propose a fusion approach that combines the two different modalities. Finally, the evaluation procedure has been carried out in the context of two different use cases, namely: (a) a general image sentiment classifier for brand and advertising images and (b) a brand-specific classification procedure, according to which the brand of the input images is known a-priori. Results have proven that the visual-based sentiment classification of brand and advertising information can outperform the respective text-based classification. In addition, fusing the two modalities leads to significant performance boosting. Theodoros Giannakopoulos, Michalis Papakostas, Stavros J. Perantonis, Vangelis Karkaletsis |
ISPA | 1 |
| 2015 | Fusing multiple audio sensors for acoustic event detectionabstractThis paper presents an Internet-of-Things approach to fusing audio sensors towards the detection of audio events, in a meeting room scenario. The different types of audio sensors (microphones) and respective individual audio analysis modules are incorporated within the context of an IoT framework that follows a message-oriented architecture. Each individual audio analysis module is composed by a feature extraction stage and a Support-Vector-Machine (SVM) classifier. A fusion module is also adopted to combine the individual sensor-level decisions, in order to extract the final classification decision. A detailed experimental evaluation on a publicly available real-world dataset proves a rather significant performance boosting in terms of overall classification accuracy. In addition, the proposed architecture enables an easy-to-use training procedure that can easily handle any number of audio sensors and respective classifiers without any prior knowledge of the room's geometry or any other constraints regarding the topological condition of the sensors. Giorgos Siantikos, Dimitris Sgouropoulos, Theodoros Giannakopoulos, Evaggelos Spyrou |
ISPA | 3 |
| 2012 | Fisher Linear Semi-Discriminant Analysis for Speaker DiarizationabstractGiven an audio signal with an unknown number of people speaking, speaker diarization aims to automatically answer the question “who spoke when.” Crucial to the success of diarization is the distance metric between speech segments, a factor depending on the choice of the feature space: distances should be low for segments of the same speaker and high for segments of different speakers. Starting from an Mel-frequency cepstrum coefficient (MFCC)-based feature space, an algorithm is proposed that finds a Fisher near-optimal linear discriminant subspace, adapted to the particular speakers which exist in the audio signal. The proposed approach relies on a semi-supervised version of Fisher linear discriminant analysis (FLD), leveraging information from the sequential structure of the audio signal as a substitute for unknown speaker labels. The resulting algorithm is completely unsupervised; therefore, the need for speaker labels in the provided or an independent set is dismissed. The eigenvalue perturbation theory is applied in order to provide optimality bounds with respect to FLD, showing the effectiveness of the approach under the assumption that speakers do not significantly modify the characteristics of their voice. A complete diarization system is then proposed, using fuzzy clustering, a non-parametric K-nearest neighbors classifier and a hidden Markov model. The experimental results show a major improvement of speaker diarization accuracy when using the optimal subspace found by the proposed approach with respect to using the initial MFCC feature space or subspaces found by competitive approaches. Theodoros Giannakopoulos, Sergios Petridis |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Multimodal and ontology-based fusion approaches of audio and visual processing for violence detection in movies
Thanassis Perperis, Theodoros Giannakopoulos, Alexandros Makris, Dimitrios I. Kosmopoulos, Sofia Tsekeridou, Stavros J. Perantonis, Sergios Theodoridis |
Expert Syst. Appl. | 2 |
| 2010 | Unsupervised Speaker Clustering in a Linear Discriminant SubspaceabstractWe present an approach for grouping single-speaker speech segments into speaker-specific clusters. Our approach is based on applying the K-means clustering algorithm to a suitable discriminant subspace, where the euclidean distance reflect speaker differences. A core feature of our approach is approximating speaker-conditional statistics, that are not available, with single-speaker segments statistics, which can be evaluated, thus making possible to apply the LDA algorithm for finding the optimal discriminative subspace, using unlabeled data. To illustrate our method, we present examples of clusters generated by our approach when applied to the ICMLA 2010 Speaker Clustering Challenge datasets. Theodoros Giannakopoulos, Sergios Petridis |
ICMLA | 1 |
| 2010 | A Multimodal Approach to Violence Detection in Video Sharing SitesabstractThis paper presents a method for detecting violent content in video sharing sites. The proposed approach operates on a fusion of three modalities: audio, moving image and text data, the latter being collected from the accompanying user comments. The problem is treated as a binary classification task (violent vs non-violent content) on a 9-dimensional feature space, where 7 out of 9 features are extracted from the audio stream. The proposed method has been evaluated on 210 YouTube videos and the overall accuracy has reached 82%. Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
ICPR | 1 |
| 2009 | A dimensional approach to emotion recognition of speech from moviesabstractIn this paper we present a novel method for extracting affective information from movies, based on speech data. The method is based on a 2D representation of speech emotions (Emotion Wheel). The goal is twofold. First, to investigate whether the Emotion Wheel offers a good representation for emotions associated with speech signals. To this end, several humans have manually annotated speech data from movies using the Emotion Wheel and the level of disagreement has been computed as a measure of representation quality. The results indicate that the emotion wheel is a good representation of emotions in speech data. Second, a regression approach is adopted, in order to predict the location of an unknown speech segment in the Emotion Wheel. Each speech segment is represented by a vector of ten audio features. The results indicate that the resulting architecture can estimate emotion states of speech from movies, with sufficient accuracy. Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
ICASSP | 1 |
| 2008 | Gunshot detection in audio streams from movies by means of dynamic programming and Bayesian networksabstractThis paper treats gunshot detection in audio streams from movies as a maximization task, where the solution is obtained by means of dynamic programming. The proposed method seeks the sequence of segments and respective class labels, i.e., gunshots vs. all other audio types, that maximize the product of posterior class label probabilities, given the segments' data. The required posterior probabilities are estimated by combining soft classification decisions from a set of Bayesian Network combiners. Tests that have been performed on a large set of audio streams indicate that the proposed method yields high performance in terms of both precision and recall of detected gunshot events. Aggelos Pikrakis, Theodoros Giannakopoulos, Sergios Theodoridis |
ICASSP | 2 |
| 2008 | A novel efficient approach for audio segmentationabstractIn this paper, a novel approach to audio segmentation is presented. The problem of detecting audio segmentspsila limits is treated as a binary classification task. Frames are classified as ldquosegment limitsrdquo vs ldquononsegment limitsrdquo. For each audio frame a spectrogram is computed and eight feature values are extracted from respective frequency bands. Final decisions are taken based on a classifier combination scheme. The algorithm has very low complexity with almost real time performance. It achieves 86% accuracy rate on real audio streams extracted from movies. Moreover, it introduces a general framework to audio segmentation, which does not depend explicitly on the number of audio classes. Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
ICPR | 1 |
| 2008 | Music tracking in audio streams from moviesabstractThis paper presents a robust and computationally efficient method for tracking music in audio streams from movies. The audio stream is first mid-term processed with a fixed length moving window and four features are extracted per window. Each feature is fed as input to a simple classifier which produces a soft output for the binary problem of music vs. all other types of audio. The soft outputs are then combined to yield a measure of confidence quantifying whether the segment corresponds to music or not. At a final step, thresholding is applied to filter out segments where the confidence measure is low. The proposed approach has been tested with audio streams from various movies and its performance was measured both on a mid-term segment basis as well as on an event detection basis. Reported results demonstrate that the method exhibits high performance even when music is mixed with other types of audio in the stream. Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
MMSP | 1 |
| 2008 | A Speech/Music Discriminator of Radio Recordings Based on Dynamic Programming and Bayesian NetworksabstractThis paper presents a multistage system for speech/music discrimination which is based on a three-step procedure. The first step is a computationally efficient scheme consisting of a region growing technique and operates on a 1-D feature sequence, which is extracted from the raw audio stream. This scheme is used as a preprocessing stage and yields segments with high music and speech precision at the expense of leaving certain parts of the audio recording unclassified. The unclassified parts of the audio stream are then fed as input to a more computationally demanding scheme. The latter treats speech/music discrimination of radio recordings as a probabilistic segmentation task, where the solution is obtained by means of dynamic programming. The proposed scheme seeks the sequence of segments and respective class labels (i.e., speech/music) that maximize the product of posterior class probabilities, given the data that form the segments. To this end, a Bayesian Network combiner is embedded as a posterior probability estimator. At a final stage, an algorithm that performs boundary correction is applied to remove possible errors at the boundaries of the segments (speech or music) that have been previously generated. The proposed system has been tested on radio recordings from various sources. The overall system accuracy is approximately 96%. Performance results are also reported on a musical genre basis and a comparison with existing methods is given. Aggelos Pikrakis, Theodoros Giannakopoulos, Sergios Theodoridis |
IEEE Trans. Multim. | 2 |
| 2007 | A Multi-Class Audio Classification Method With Respect To Violent Content In Movies Using Bayesian NetworksabstractIn this work, we present a multi-class classification algorithm for audio segments recorded from movies, focusing on the detection of violent content, for protecting sensitive social groups (e.g. children). Towards this end, we have used twelve audio features stemming from the nature of the signals under study. In order to classify the audio segments into six classes (three of them violent), Bayesian networks have been used in combination with the one versus all classification architecture. The overall system has been trained and tested on a large data set (5000 audio segments), recorded from more than 30 movies of several genres. Experiments showed, that the proposed method can be used as an accurate multi-class classification scheme, but also, as a binary classifier for the problem of violent -non violent audio content. Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
MMSP | 1 |
| 2006 | A Speech/Music Discriminator for Radio Recordings Using Bayesian NetworksabstractThis paper presents a speech/music discriminator for radio recordings. The segmentation stage is based on the detection of changes in the energy distribution of the audio signal. For the classification stage, Bayesian networks have been adopted in order to combine the results of nine k-nearest neighbor classifiers trained on individual features. To this end, a comparison of the performance of three popular Bayesian network architectures is presented. Furthermore, in order to reduce the number of features used for classification, a new feature selection scheme is introduced, that is also based on the properties of Bayesian networks. The proposed system has been tested on real Internet broadcasts of BBC radio stations Theodoros Giannakopoulos, Aggelos Pikrakis, Sergios Theodoridis |
ICASSP (5) | 1 |