Irene Martín-Morató

dblp:189/8281 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
2since 2021 · last 2023
0000-0002-2115-0193ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Audio and music processing · 68% Visualization and visual analytics · 32%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › sound event detection
sound event recognition
0.822020
Adaptive Distance-Based Pooling in Convolutional Neural Networks for Audio Event Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2020
Adaptive Mid-Term Representations for Robust Audio Event Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Visualization and visual analytics
crowdsourced annotation
0.712023
Strong Labeling of Sound Events Using Crowdsourced Weak Labels and Annotator Competence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Audio and music processing
sound event detection
0.712023
Strong Labeling of Sound Events Using Crowdsourced Weak Labels and Annotator Competence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2023

Methods — techniques the papers use, named apart from their topics

weak label aggregation · 0.7multi-annotator competence estimation · 0.7data augmentation · 0.4convolutional neural network · 0.4adaptive distance-based pooling · 0.4texture windows · 0.3nonlinear time normalization · 0.3distance-based subsampling · 0.3
YearPublicationVenuePosition
2023 Training Sound Event Detection with Soft Labels from Crowdsourced Annotations
abstract
In this paper, we study the use of soft labels to train a system for sound event detection (SED). Soft labels can result from annotations which account for human uncertainty about categories, or emerge as a natural representation of multiple opinions in annotation. Converting annotations to hard labels results in unambiguous categories for training, at the cost of losing the details about the labels distribution. This work investigates how soft labels can be used, and what benefits they bring in training a SED system. The results show that the system is capable of learning information about the activity of the sounds which is reflected in the soft labels and is able to detect sounds that are missed in the typical binary target training setup. We also release a new dataset produced through crowdsourcing, containing temporally strong labels for sound events in real-life recordings, with both soft and hard labels.
Irene Martín-Morató, Manu Harju, Paul Ahokas, Annamaria Mesaros
ICASSP1
2023 Strong Labeling of Sound Events Using Crowdsourced Weak Labels and Annotator Competence Estimation
abstract
Crowdsourcing is a popular tool for collecting large amounts of annotated data, but the specific format of the strong labels necessary for sound event detection is not easily obtainable through crowdsourcing. In this work, we propose a novel annotation workflow that leverages the efficiency of crowdsourcing weak labels, and uses a high number of annotators to produce reliable and objective strong labels. The weak labels are collected in a highly redundant setup, to allow reconstruction of the temporal information. To obtain reliable labels, the annotators' competence is estimated using MACE (Multi-Annotator Competence Estimation) and incorporated into the strong labels estimation through weighing of individual opinions. We show that the proposed method produces consistently reliable strong annotations not only for synthetic audio mixtures, but also for audio recordings of real everyday environments. While only a maximum 80% coincidence with the complete and correct reference annotations was obtained for synthetic data, these results are explained by an extended study of how polyphony and SNR levels affect the identification rate of the sound events by the annotators. On real data, even though the estimated annotators' competence is significantly lower and the coincidence with reference labels is under 69%, the proposed majority opinion approach produces reliable aggregated strong labels in comparison with the more difficult task of crowdsourcing directly strong labels.
Irene Martín-Morató, Annamaria Mesaros
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Adaptive Distance-Based Pooling in Convolutional Neural Networks for Audio Event Classification
abstract
In the last years, deep convolutional neural networks have become a standard for the development of state-of-the-art audio classification systems, taking the lead over traditional approaches based on feature engineering. While they are capable of achieving human performance under certain scenarios, it has been shown that their accuracy is severely degraded when the systems are tested over noisy or weakly segmented events. Although better generalization could be obtained by increasing the size of the training dataset, e.g. by applying data augmentation techniques, this also leads to longer and more complex training procedures. In this article, we propose a new type of pooling layer aimed at compensating non-relevant information of audio events by applying an adaptive transformation of the convolutional feature maps in the temporal axis. The proposed layer performs a non-linear temporal transformation that follows a uniform distance subsampling criterion on the learned feature space. The experiments conducted over different datasets show significant performance improvements when the proposed layer is added to baseline models, resulting in systems that generalize better to mismatching test conditions and learn more robustly from weakly labeled data.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Sound Event Envelope Estimation in Polyphonic Mixtures
abstract
Sound event detection is the task of identifying automatically the presence and temporal boundaries of sound events within an input audio stream. In the last years, deep learning methods have established themselves as the state-of-the-art approach for the task, using binary indicators during training to denote whether an event is active or inactive. However, such binary activity indicators do not fully describe the events, and estimating the envelope of the sounds could provide more precise modeling of their activity. This paper proposes to estimate the amplitude envelopes of target sound event classes in polyphonic mixtures. For training, we use the amplitude envelopes of the target sounds, calculated from mixture signals and, for comparison, from their isolated counterparts. The model is then used to perform envelope estimation and sound event detection. Results show that the envelope estimation allows good modeling of the sounds activity, with detection results comparable to current state-of-the art.
Irene Martín-Morató, Annamaria Mesaros, Toni Heittola, Tuomas Virtanen, Maximo Cobos, Francesc J. Ferri
ICASSP1
2018 Adaptive Mid-Term Representations for Robust Audio Event Classification
abstract
Low-level audio features are commonly used in many audio analysis tasks, such as audio scene classification or acoustic event detection. Due to the variable length of audio signals, it is a common approach to create fixed-length feature vectors consisting of a set of statistics that summarize the temporal variability of such short-term features. To avoid the loss of temporal information, the audio event can be divided into a set of mid-term segments or texture windows. However, such an approach requires to estimate accurately the onset and offset times of the audio events in order to obtain a robust mid-term statistical description of their temporal evolution. This paper proposes the use of an alternative event representation based on nonlinear time normalization prior to the extraction of mid-term statistics. The short-term features are transformed into a new fixed-length representation that considers uniform distance subsampling over a defined feature space in contrast to the classical short-term temporal framing. The results show that the use of distance-based texture windows provides an improved statistical description of the event robust to errors in the event segmentation stage under noisy conditions.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Analysis of data fusion techniques for multi-microphone audio event detection in adverse environments
abstract
Acoustic event detection (AED) is currently a very active research area with multiple applications in the development of smart acoustic spaces. In this context, the advances brought by Internet of Things (IoT) platforms where multiple distributed microphones are available have also contributed to this interest. In such scenarios, the use of data fusion techniques merging information from several sensors becomes an important aspect in the design of multi-microphone AED systems. In this paper, we present a preliminary analysis of several data-fusion techniques aimed at improving the recognition accuracy of an AED system by taking advantage of the diversity provided by multiple microphones in adverse acoustic conditions. The results confirm that, under appropriate processing schemes, the recognition rate can be increased as well as the corresponding independence on event location.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
MMSP1