Stavros Ntalampiras

dblp:13/928 · DBLP profile ↗
← Back
43ranked-venue papers
26as first author
16since 2021 · last 2025
0000-0003-3482-9215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 17 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 7 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Speech-Based Depression Assessment: A Comprehensive Survey
abstract
Depression (major depressive disorder) is one of the most common mental illnesses worldwide, causing feelings of sadness and loss of interest, and is a leading cause of suicidal ideation. Limited access to mental health services, stigma, patient privacy and delay in seeking help are the most significant barriers to assessment and effective treatment. In order to enhance the accuracy of depression prediction, automated strategies employing computational models have been widely explored in literature. To this end, automatic Speech Depression Recognition (SDR) methods stand out, as speech comprises a valuable marker of mental health. Interestingly, recording speech comprises a less intrusive and more portable approach than capturing video, thus more easily accepted, especially by the younger generations, who are at a considerable risk of social isolation due to addiction to social networks and excessive use of mobile devices. In this context, this paper presents an up-to-date survey on SDR. More specifically, we a) detail the major challenges and key issues on SDR, b) summarise the most recent approaches existing in the related literature, and c) highlight the open problems. At the same time, we illustrate a framework encompassing the latest tendencies for SDR, along with a suitable comparison of the achieved performances. Finally, we highlight future trends and present the overall findings, providing researchers with best practices and techniques to address the major challenges of SDR, as well as stimulating discussion and improvement in the field.
Samara Soares Leal, Stavros Ntalampiras, Roberto Sassi
IEEE Trans. Affect. Comput.2
2023 Ensemble Learning for Cough-Based Subject-Independent COVID-19 Detection
abstract
This paper belongs to the medical acoustics field and presents a solution for COVID-19 detection based on the cough sound events.Unfortunately, the use of RT-PCR Molecular Swab tests for the diagnosis of COVID-19 is associated with considerable cost, is based on availability of suitable equipment, requires a specific time period to produce the result, let alone the potential errors in the execution of the tests.Interestingly, in addition to Swab tests, cough sound events could facilitate the detection of COVID-19.Currently, there is a great deal of research in this direction, which has led to the development of publicly available datasets which have been processed, segmented, and labeled by medical experts.This work proposes an ensemble composed of a variety of classifiers suitably adapted to the present problem.Such classifiers are based on a standardized feature extraction front-end representing the involved audio signals limiting the necessity to design handcrafted features.In addition, we elaborate on a prearranged publicly available dataset and introduce an experimental protocol taking into account model bias originating from subject dependency.After thorough experiments, the proposed model was able to outperform the state of the art both in patient-dependent and -independent settings. 798
Vincenzo Conversano, Stavros Ntalampiras
ICPRAM2
2023 A Hierarchical Approach for Multilingual Speech Emotion Recognition
Marco Nicolini, Stavros Ntalampiras
ICPRAM2
2023 Lightweight Audio-Based Human Activity Classification Using Transfer Learning
abstract
This paper employs the acoustic modality to address the human activity recognition (HAR) problem.The cornerstone of the proposed solution is the YAMNet deep neural network, the embeddings of which comprise the input to a fully-connected linear layer trained for HAR.Importantly, the dataset is publicly available and includes the following human activities: preparing coffee, frying egg, no activity, showering, using microwave, washing dishes, washing hands, and washing teeth.The specific set of activities is representative of a standard home environment facilitating a wide range of applications.The performance offered by the proposed transfer learning-based framework surpasses the state of the art, while being able to be executed on mobile devices, such as smartphones, tablets, etc.In fact, the obtained model has been exported and thoroughly tested for real-time HAR on a smartphone device with the input being the audio captured from its microphone.
Marco Nicolini, Federico Simonetta, Stavros Ntalampiras
ICPRAM3
2023 Adversarial Attacks Against Acoustic Monitoring of Industrial Machines
abstract
The recent rise of adversarial machine learning exposed the serious vulnerabilities existing in current frameworks depending on the smooth operation of such automated solutions. This article focuses on the critical field of monitoring the health of industrial machines based on the respective acoustic emissions. After building an audio-based monitoring solution using log-Mel spectrograms and convolutional neural networks, we systematically evaluate the applicability of four types of adversarial attacks: 1) fast gradient sign; 2) projected gradient descent; 3) Jacobian saliency map; and 4) Carlini and Wagner$\ell _{\infty }$. Seeing the problem from the attacker perspective, we designed two different attack types, aiming at inducing either false positives or false negatives. We define three figures of merit specifically designed to assess the performance of each attack type from diverse points of view. The experimental setup relies on a publicly available data set including acoustic emissions representing four industrial machines, i.e., fan, pump, slider rail, and valve.
Stavros Ntalampiras
IEEE Internet Things J.1
2023 Few-shot learning for modeling cyber physical systems in non-stationary environments
abstract
Abstract This paper proposes a modeling scheme for cyber physical systems operating in non-stationary, small data environments. Unlike the traditional modeling logic, we introduce the few-shot learning paradigm, the operation of which is based on quantifying both similarities and dissimilarities. As such, we designed a suitable change detection mechanism able to reveal previously unknown operational states, which are incorporated in the dictionary online. We elaborate on spectrograms extracted from high-resolution ultrasound depth sensor timeseries, while the backbone of the proposed method is a Siamese Neural Network. The experimental scenario considers data representing liquid containers for fuel/water when the following five operational states are present: normal, accident, breakdown, sabotage, and cyber-attack. Thorough experiments were carried out assessing every aspect of the present framework and demonstrating its efficacy even when very few samples per class are available. In addition, we propose a probabilistic data selection scheme facilitating one-shot learning. Last but not least, responding to the wide requirement for interpretable AI, we explain the obtained predictions by examining the layer-wise activation maps.
Stavros Ntalampiras, Ilyas Potamitis
Neural Comput. Appl.1
2023 Explainable Siamese Neural Network for Classifying Pediatric Respiratory Sounds
abstract
The field of medical acoustics is gaining constantly-increasing attention by the scientific community, with the the general goal being the automatic understanding of medical related signals to assist medical personnel in the decision making process. In this direction, this work introduces a framework able to differentiate between normal and abnormal respiratory sounds by learning relationships characterizing pairs of sounds. More specifically, considering the nature of respiratory sounds, we designed a feature set able to capture the coarse and fine structure exhibited such signals by means of multiresolution analysis. Similar/dissimilar relationships are modeled via a suitably-learned Siamese Neural Network encompassing a series of convolutional layers. Interestingly, such a relationship learning framework conveniently solves the existing class imbalance problem as it is trained on pairs of similar/dissimilar audio signals. Importantly, we employed the dataset designed for the IEEE BioCAS 2022 Grand challenge on Respiratory Sound Classification along with a standardized experimental protocol allowing reproducibility and reliable comparison between different approaches. After extensive experiments assessing the proposed framework from diverse points of view, including an ablation study, it is shown that it outperforms existing approaches, while providing explainable predictions via a Q&A scheme allowing interaction with the medical experts.
Stavros Ntalampiras
IEEE J. Biomed. Health Informatics1
2022 Variational Autoencoders for Anomaly Detection in Respiratory Sounds
Michele Cozzatti, Federico Simonetta, Stavros Ntalampiras
ICANN (4)3
2022 Deep Feature Learning for Medical Acoustics
Alessandro Maria Poirè, Federico Simonetta, Stavros Ntalampiras
ICANN (4)3
2022 Learning relationships between audio signals based on reservoir networks
abstract
Inspired by the respective human capacity, this paper investigates learning through analogies via the reservoir computing paradigm. Even though the vast majority of the literature considers modeling of patterns/objects serving classification-based applications, we move to learning of relationships characterizing input pairs. To this end, we describe the design process of a suitable echo state network capturing relationships existing between audio signals, i.e. similar, dissimilar, reverberation, and noise. Such a neural network operates on frequency-domain representations of sound belonging to the following four classes: male and female speech, applause and laughter. After extensive experiments, we demonstrate that echo state networks are able to learn through analogies and achieve superior performance when compared with Siamese neural networks.
Stavros Ntalampiras
IJCNN1
2022 Acoustics-specific Piano Velocity Estimation
abstract
Motivated by the state-of-art psychological research, we note that a piano performance transcribed with existing Automatic Music Transcription (AMT) methods cannot be successfully resynthesized without affecting the artistic content of the performance. This is due to 1) the different mappings between MIDI parameters used by different instruments, and 2) the fact that musicians adapt their way of playing to the surrounding acoustic environment. To face this issue, we propose a methodology to build acoustics-specific AMT systems that are able to model the adaptations that musicians apply to convey their interpretation. Specifically, we train models tailored for virtual instruments in a modular architecture that takes as input an audio recording and the relative aligned music score and outputs the acoustics-specific velocities of each note. We test different model shapes and show that the proposed methodology generally outperforms the usual AMT pipeline which does not consider the specificities of the instrument and of the acoustic environment. Interestingly, such a methodology is extensible in a straightforward way since only slight efforts are required to train models for the inference of other piano parameters, such as pedaling.
Federico Simonetta, Stavros Ntalampiras, Federico Avanzini
MMSP2
2022 A perceptual measure for evaluating the resynthesis of automatic music transcriptions
abstract
Abstract This study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change. We propose to distinguish the concept of “performance” from the one of “interpretation”, which expresses the “artistic intention”. Towards assessing this distinction, we carried out an experimental evaluation where 91 subjects were invited to listen to various audio recordings created by resynthesizing MIDI data obtained through Automatic Music Transcription (AMT) systems and a sensorized acoustic piano. During the resynthesis, we simulated different contexts and asked listeners to evaluate how much the interpretation changes when the context changes. Results show that: (1) MIDI format alone is not able to completely grasp the artistic intention of a music performance; (2) usual objective evaluation measures based on MIDI data present low correlations with the average subjective evaluation. To bridge this gap, we propose a novel measure which is meaningfully correlated with the outcome of the tests. In addition, we investigate multimodal machine learning by providing a new score-informed AMT method and propose an approximation algorithm for thep-dispersion problem.
Federico Simonetta, Federico Avanzini, Stavros Ntalampiras
Multim. Tools Appl.3
2021 CatMeows: A Publicly-Available Dataset of Cat Vocalizations
abstract
This work presents a dataset of cat vocalizations focusing on the meows emitted in three different contexts: brushing, isolation in an unfamiliar environment, and waiting for food. The dataset contains vocalizations produced by 21 cats belonging to two breeds, namely Maine Coon and European Shorthair. Sounds have been recorded using low-cost devices easily available on the marketplace, and the data acquired are representative of real-world cases both in terms of audio quality and acoustic conditions. The dataset is open-access, released under Creative Commons Attribution 4.0 International licence, and it can be retrieved from the Zenodo web repository.
Luca A. Ludovico, Stavros Ntalampiras, Giorgio Presti, Simona Cannas, Monica Battini, Silvana Mattiello
MMM (2)2
2021 Audio-to-Score Alignment Using Deep Automatic Music Transcription
abstract
Audio-to-score alignment (A2SA) is a multimodal task consisting in the alignment of audio signals to music scores. Recent literature confirms the benefits of Automatic Music Transcription (AMT) for A2SA at the frame-level. In this work, we aim to elaborate on the exploitation of AMT Deep Learning (DL) models for achieving alignment at the note-level. We propose a method which benefits from HMM-based score-to-score alignment and AMT, showing a remarkable advancement beyond the state-of-the-art. We design a systematic procedure to take advantage of large datasets which do not offer an aligned score. Finally, we perform a thorough comparison and extensive tests on multiple datasets.
Federico Simonetta, Stavros Ntalampiras, Federico Avanzini
MMSP2
2021 One-shot learning for acoustic diagnosis of industrial machines
Stavros Ntalampiras
Expert Syst. Appl.1
2021 Speech emotion recognition via learning analogies
Stavros Ntalampiras
Pattern Recognit. Lett.1
2020 One-shot learning for acoustic identification of bird species in non-stationary environments
abstract
This work introduces the one-shot learning paradigm in the computational bioacoustics domain. Even though, most of the related literature assumes availability of data characterizing the entire class dictionary of the problem at hand, that is rarely true as a habitat's species composition is only known up to a certain extent. Thus, the problem needs to be addressed by methodologies able to cope with non-stationarity. To this end, we propose a framework able to detect changes in the class dictionary and incorporate new classes on the fly. We design an one-shot learning architecture composed of a Siamese Neural Network operating in the loglvlel spectrogram space. We extensively examine the proposed approach on two datasets of various bird species using suitable figures of merit. Interestingly, such a learning scheme exhibits state of the art performance, while taking into account extreme non-stationarity cases.
Michelangelo Acconcjaioco, Stavros Ntalampiras
ICPR2
2020 Collaborative framework for automatic classification of respiratory sounds
abstract
There are several diseases (e.g. asthma, pneumonia etc.) affecting the human respiratory apparatus altering its airway path substantially, thus characterising its acoustic properties. This work unfolds an automatic audio signal processing framework achieving classification between normal and abnormal respiratory sounds. Thanks to a recent challenge, a real‐world dataset specifically designed to address the needs of the specific problem is available to the scientific community. Unlike previous works in the literature, the authors take advantage of information provided by several stethoscopes simultaneously, i.e. elaborating at the acoustic sensor network level. To this end, they employ two features sets extracted from different domains, i.e. spectral and wavelet. These are modelled by convolutional neural networks, hidden Markov models and Gaussian mixture models. Subsequently, a synergistic scheme is designed operating at the decision level of the best‐performing classifier with respect to each stethoscope. Interestingly, such a scheme was able to boost the classification accuracy surpassing the current state of the art as it is able to identify respiratory sound patterns with a 66.7% accuracy.
Stavros Ntalampiras
IET Signal Process.1
2020 Emotional quantification of soundscapes by learning between samples
abstract
Abstract Predicting the emotional responses of humans to soundscapes is a relatively recent field of research coming with a wide range of promising applications. This work presents the design of two convolutional neural networks, namely ArNet and ValNet, each one responsible for quantifying arousal and valence evoked by soundscapes. We build on the knowledge acquired from the application of traditional machine learning techniques on the specific domain, and design a suitable deep learning framework. Moreover, we propose the usage of artificially created mixed soundscapes, the distributions of which are located between the ones of the available samples, a process that increases the variance of the dataset leading to significantly better performance. The reported results outperform the state of the art on a soundscape dataset following Schafer’s standardized categorization considering both sound’s identity and the respective listening context.
Stavros Ntalampiras
Multim. Tools Appl.1
2019 Classification of Sounds Indicative of Respiratory Diseases
Stavros Ntalampiras, Ilyas Potamitis
EANN1
2018 Forecasting Mobile Service Demands for Anticipatory MEC
abstract
The accurate estimation of future traffic loads is a key enabler for anticipatory mobile networking. In this paper, we investigate the prediction of the traffic generated by different mobile service classes over base station clusters, at an order-of-minute granularity and using relatively short historical data. This scenario is relevant to mobile edge computing (MEC), where resources need to be orchestrated for individual services separately across multiple base stations, at fairly long timescales. To address the prediction problem, we propose a novel forecasting model based on an autoregressive multiple-input single-output (MISO) approach, where the inputs are collected from regions exhibiting strong correlations in the offered load of a specific mobile service. Experiments on real-world data collected in an operational 3G/4G network demonstrate the effectiveness of our model, which attains average relative errors between 0.4 % and 5 % when forecasting 5-minute-aggregate traffic of individual mobile service classes.
Stavros Ntalampiras
WOWMOM1
2017 Emotion Prediction of Sound Events Based on Transfer Learning
Stavros Ntalampiras, Ilyas Potamitis
EANN1
2017 Hybrid framework for categorising sounds of Mysticete whales
abstract
This study addresses a problem belonging to the domain of whale audio processing, more specifically the automatic classification of sounds produced by the Mysticete species. The specific task is quite challenging given the vast repertoire of the involved species, the adverse acoustic conditions and the nearly inexistent prior scientific work. Two feature sets coming from different domains (frequency and wavelet) were designed to tackle the problem. These are modelled by means of a hybrid technique taking advantage of the merits of a generative and a discriminative classifier. The dataset includes five species ( Blue , Fin , Bowhead , Southern Right , and Humpback ) and it is publicly available at http://www.mobysound.org/ . The authors followed a thorough experimental procedure and achieved quite encouraging recognition rates.
Stavros Ntalampiras
IET Signal Process.1
2016 Online model-free sensor fault identification and dictionary learning in Cyber-Physical Systems
abstract
This paper presents a model-free method for the online identification of sensor faults and learning of their fault dictionary. The method, designed having in mind Cyber-Physical Systems (CPSs), takes advantage of functional relationships among the datastreams acquired by CPS sensing units. Existing model-free change detection mechanisms are proposed to detect faults and identify the fault type thanks to a fault dictionary which is built over time. The main features of the proposed algorithm are its ability to operate without requiring any a priori information about the system under inspection or the nature of the possibly occurring faults. As such, the method follows the model-free approach, characterized by the fact the fault dictionary is constructed online once faults are detected. Whenever available, humans can be considered in the loop to label a fault or a fault class in the dictionary as well as introduce fault instances generated thanks to a priori information. Experimental results on both synthetic and real datasets corroborate the effectiveness of the proposed fault diagnosis system.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN2
2016 Automatic identification of integrity attacks in cyber-physical systems
Stavros Ntalampiras
Expert Syst. Appl.1
2015 Detection of Integrity Attacks in Cyber-Physical Critical Infrastructures Using Ensemble Modeling
abstract
This paper presents an anomaly-based methodology for reliable detection of integrity attacks in cyber-physical critical infrastructures. Such malicious events compromise the smooth operation of the infrastructure while the attacker is able to exploit the respective resources according to his/her purposes. Even though the operator may not understand the attack, since the overall system appears to remain in a steady state, the consequences may be of catastrophic nature with a huge negative impact. Here, we apply a computational intelligent technique which incorporates the merits of two of the heterogeneous modeling approaches (linear time-invariant and neural networks), while considering both temporal and functional dependencies existing among the elements of an infrastructure. The experimental platform includes a power grid simulator of the IEEE 30 bus model and a cyber network emulator. Subsequently, we implemented a wide range of integrity attacks (replay, ramp, pulse, scaling, and random) with different intensity levels. A thorough evaluation procedure is carried out while the results demonstrate the ability of the proposed method to produce a desired result in terms of false positive rate, false negative rate, and detection delay.
Stavros Ntalampiras
IEEE Trans. Ind. Informatics1
2015 Fault Identification in Distributed Sensor Networks Based on Universal Probabilistic Modeling
abstract
This paper proposes a holistic modeling scheme for fault identification in distributed sensor networks. The proposed scheme is based on modeling the relationship between two datastreams by means of a hidden Markov model (HMM) trained on the parameters of linear time-invariant dynamic systems, which estimate the specific relationship over consecutive time windows. Every system state, including the nominal one, is represented by an HMM and the novel data are categorized according to the model producing the highest likelihood. The system is able to understand whether the novel data belong to the fault dictionary, are fault-free, or represent a new fault type. We extensively evaluated the discrimination capabilities of the proposed approach and contrasted it with a multilayer perceptron using data coming from the Barcelona water distribution network. Nine system states are present in the dataset and the recognition rates are provided in the confusion matrix form.
Stavros Ntalampiras
IEEE Trans. Neural Networks Learn. Syst.1
2014 Automatic Fault Identification in Sensor Networks Based on Probabilistic Modeling
Stavros Ntalampiras, Georgios Giannopoulos
CRITIS1
2014 Faults and Cyber Attacks Detection in Critical Infrastructures
Yannis Soupionis, Stavros Ntalampiras, Georgios Giannopoulos
CRITIS2
2014 On predicting the unpleasantness level of a sound event
Stavros Ntalampiras, Ilyas Potamitis
INTERSPEECH1
2013 Model ensemble for an effective on-line reconstruction of missing data in sensor networks
abstract
The literature has shown that model ensemble techniques are particularly effective to solve regression/classification applications by providing, given a suitable aggregation mechanism, a better generalization ability than the generic model of the ensemble. However, only few recent results consider the use of ensembles for a time-dependent framework, with focus on time-series forecasting. Here, we propose the use of ensemble of models to an on-line reconstruction of missing data coming from a sensor network. Reconstructing missing data is of paramount importance for any further data processing and must be carried out on-line not to introduce unnecessary latency when data lead to a decision or control action. The ensemble is designed by both exploiting temporal and spatial dependencies existing among the sensors composing the network. An effective aggregation mechanism is proposed for the considered models to improve the generalization ability of the ensemble. Results demonstrate the effectiveness of the proposed approach in reconstructing missing data.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN2
2013 A Novel Holistic Modeling Approach for Generalized Sound Recognition
abstract
Nowadays, generalized sound recognition technology is constantly gaining attention within the generic context of scene analysis and understanding (smart-home, surveillance, bioacoustics, etc.). It is typically achieved using a set of relevant to the task at hand descriptors modelled by means of a statistical tool, e.g., hidden Markov model. This work exhaustively applies the Universal Modeling (UM) (or class-independent) approach on the particular task. The feature extraction engine extracts descriptors belonging to time, frequency and wavelet domains. We describe a novel data selection scheme based on Gaussian mixture model clustering for the creation of the UM. The scheme takes into account the dataset characteristics, adapts itself to them and leads to higher recognition rates than the standard UM approach.
Stavros Ntalampiras
IEEE Signal Process. Lett.1
2013 A Cognitive Fault Diagnosis System for Distributed Sensor Networks
abstract
This paper introduces a novel cognitive fault diagnosis system (FDS) for distributed sensor networks that takes advantage of spatial and temporal relationships among sensors. The proposed FDS relies on a suitable functional graph representation of the network and a two-layer hierarchical architecture designed to promptly detect and isolate faults. The lower processing layer exploits a novel change detection test (CDT) based on hidden Markov models (HMMs) configured to detect variations in the relationships between couples of sensors. HMMs work in the parameter space of linear time-invariant dynamic systems, approximating, over time, the relationship between two sensors; changes in the approximating model are detected by inspecting the HMM likelihood. Information provided by the CDT layer is then passed to the cognitive one, which, by exploiting the graph representation of the network, aggregates information to discriminate among faults, changes in the environment, and false positives induced by the model bias of the HMMs.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IEEE Trans. Neural Networks Learn. Syst.2
2012 An HMM-based change detection method for intelligent embedded sensors
abstract
In this work we address the problem of automatically detecting changes either induced by faults or concept drifts in data streams coming from multi-sensor units. The proposed methodology is based on the fact that the relationships among different sensor measurements follow a probabilistic pattern sequence when normal data, i.e. data which do not present a change, are observed. Differently, when a change in the process generating the data occurs the probabilistic pattern sequence is modified. The relationship between two generic data streams is modelled through a sequence of linear dynamic time-invariant models whose trained coefficients are used as features feeding a Hidden Markov Model (HMM) which, in turn, extracts the pattern structure. Change detection is achieved by thresholding the log-likelihood value associated with incoming new patterns, hence comparing the affinity between the structure of new acquisitions with that learned through the HMM. Experiments on both artificial and real data demonstrate the appreciable performance of the method both in terms of detection delay, false positive and false negative rates.
Cesare Alippi, Stavros Ntalampiras, Manuel Roveri
IJCNN2
2012 Modeling the Temporal Evolution of Acoustic Parameters for Speech Emotion Recognition
abstract
During recent years, the field of emotional content analysis of speech signals has been gaining a lot of attention and several frameworks have been constructed by different researchers for recognition of human emotions in spoken utterances. This paper describes a series of exhaustive experiments which demonstrate the feasibility of recognizing human emotional states via integrating low level descriptors. Our aim is to investigate three different methodologies for integrating subsequent feature values. More specifically, we used the following methods: 1) short-term statistics, 2) spectral moments, and 3) autoregressive models. Additionally, we employed a newly introduced group of parameters which is based on the wavelet decomposition. These are compared with a baseline set comprised of descriptors which are usually used for the specific task. Subsequently, we experimented on fusing these sets on the feature and log-likelihood levels. The classification step is based on hidden Markov models, while several algorithms which can handle redundant information were used during fusion. We report results on the well-known and freely available database BERLIN using data of six emotional states. Our experiments show the importance of including information which is captured by the set based on multiresolution analysis and the efficacy of merging subsequent feature values.
Stavros Ntalampiras, Nikos Fakotakis
IEEE Trans. Affect. Comput.1
2011 Probabilistic Novelty Detection for Acoustic Surveillance Under Real-World Conditions
abstract
Novelty detection in the machine learning context refers to identifying unknown/novel data, i.e., data which vary greatly from the ones that the system was trained with. This paper explores this technique as applied to acoustic surveillance of abnormal situations. The ultimate goal of the system is to help an authorized person towards taking the appropriate actions for preventing life/property loss. A wide variety of acoustic parameters is employed towards forming a multidomain feature vector, which captures diverse characteristics of the audio signals. Subsequently the feature coefficients are fed to three probabilistic novelty detection methodologies. Their performance is computed using two measures which take into account misdetections and false alarms. Out dataset was recorded under real-world conditions including three different locations where various types of normal and abnormal sound events were captured. A smart-home environment, an open public space, and an office corridor were used. The results indicate that probabilistic novelty detection can provide an accurate analysis of the audio scene to identify abnormal events.
Stavros Ntalampiras, Ilyas Potamitis, Nikos Fakotakis
IEEE Trans. Multim.1
2010 Fusion of acoustic and optical sensor data for automatic fight detection in urban environments
Maria Andersson, Stavros Ntalampiras, Todor Ganchev, Joakim Rydell, Jörgen Ahlberg, Nikos Fakotakis
FUSION2
2010 A multidomain approach for automatic home environmental sound classification
Stavros Ntalampiras, Ilyas Potamitis, Nikos Fakotakis
INTERSPEECH1
2010 Identification of abnormal audio events based on probabilistic novelty detection
Stavros Ntalampiras, Ilyas Potamitis, Nikos Fakotakis
INTERSPEECH1
2010 Heterogeneous Sensor Database in Support of Human Behaviour Analysis in Unrestricted Environments: The Audio Part
Stavros Ntalampiras, Todor Ganchev, Ilyas Potamitis, Nikos Fakotakis
LREC1
2009 On acoustic surveillance of hazardous situations
abstract
The present study presents a practical methodology for automatic space monitoring based solely on the perceived acoustic information. We consider the case where atypical situations such as screams, explosions and gunshots take place in a metro station environment. Our approach is based on a two stage recognition schema, each one exploiting HMMs for approximating the density function of the corresponding sound class. The main objective is to detect abnormal events that take place in a noisy environment. A thorough evaluation procedure is carried out under different SNR conditions and we report high detection rates with respect to false alarm and miss probabilities rates.
Stavros Ntalampiras, Ilyas Potamitis, Nikos Fakotakis
ICASSP1
2008 A comparative study in automatic recognition of broadcast audio
abstract
This paper provides a thorough description of a methodology which leads to high accuracy as regards automatic analysis of broadcast audio. The main objective is to find a feature set for efficient speech/music discrimination while keeping the number of its dimensions as small as possible. Three groups of parameters based on Mel-scale filterbank, MPEG-7 standard and wavelet decomposition are examined in detail. We annotated on-line radio recordings characterized by great diversity, for building probabilistic models and testing four frameworks. The proposed approach utilizes wavelets and MPEG-7 ASP descriptor for modeling speech and music respectively, and results to 98.5 % average recognition rate.
Stavros Ntalampiras, Nikos Fakotakis
INTERSPEECH1
2008 Audio Database in Support of Potentiel Threat and Crisis Situation Management
Stavros Ntalampiras, Ilyas Potamitis, Todor Ganchev, Nikos Fakotakis
LREC1