Maximo Cobos

dblp:95/7644 · also Máximo Cobos-Serrano · DBLP profile ↗
← Back
43ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-7318-3192ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 1 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Real-time object tracking with on-device deep learning for adaptive beamforming in dynamic acoustic environments
abstract
Abstract Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning–based tracking with beamforming to achieve precise sound source localization and directional audio capture in dynamic environments. The approach combines single-camera depth estimation and stereo vision to enable accurate 3D localization of moving objects. A planar concentric circular microphone array constructed with MEMS microphones provides a compact, energy-efficient platform supporting 2D beam steering across azimuth and elevation. Real-time tracking outputs continuously adapt the array’s focus, synchronizing the acoustic response with the target’s position. By uniting learned spatial awareness with dynamic steering, the system maintains robust performance in the presence of multiple or moving sources. Experimental evaluation demonstrates significant gains in signal-to-interference ratio, making the design well-suited for teleconferencing, smart home devices, and assistive technologies.
Jorge Ortigoso-Narro, Jose A. Belloch, Adrian Amor-Martin, Sandra Roger 0002, Maximo Cobos
J. Supercomput.5
2025 An Incremental Selection Method for Semi-Supervised Speaker Adaptation in Speech Emotion Recognition
abstract
Adapting Speech Emotion Recognition (SER) to new, previously unseen speakers, remains a significant challenge due to the variability in emotional expression across speakers and the scarcity of labeled data for adaptation. This study introduces a novel incremental adaptation framework designed to address these challenges by leveraging a modified k-means algorithm to iteratively select representative samples in a latent space. By progressively refining the model in this way, the method facilitates the alignment of a given source domain to a new speaker enhancing generalization. Experiments conducted using data from diverse datasets, under both balanced and unbalanced conditions, demonstrate that the proposed approach outperforms random selection and labeling, achieving comparable or superior results to state-of-the-art, non-incremental methods. These findings underscore the potential of the proposed incremental strategy for improving speaker adaptation in SER tasks, particularly in data-limited scenarios.
Carlos Castorena, Maximo Cobos, Francesc J. Ferri
IEEE Signal Process. Lett.2
2025 Edge computing for driving safety: evaluating deep learning models for cost-effective sound event detection
Carlos Castorena, Jesús López Ballester, Juan Antonio De Rus, Maximo Cobos, Francesc J. Ferri
J. Supercomput.4
2024 Evaluation of Real-Time Acoustic Event Detection Models in Driving Scenarios
abstract
This paper delves into addressing road safety concerns by exploring cost-effective solutions for detecting sound events (SED) specifically designed for driving scenarios. While advanced technologies such as deep learning show promise in enhancing road safety, their practical implementation often involves expensive sensors and hardware. Given that distractions are a significant contributor to accidents, it is crucial to have effective detection and mitigation measures in place. This study focuses on auditory distractions and employs SED with affordable edge devices to identify and timestamp relevant audio events, offering valuable insights into the driving environment. We assess the performance of state-of-the-art deep learning models on various edge devices, including the 2023 DCASE baseline with convolutional recurrent neural networks (CRNN) and a customized YOLO vision model for audio spectrograms. Our analysis spans different hardware options, ranging from single board computers (SBCs) to desktop equipment, providing guidance on selecting cost-effective hardware for in-vehicle SED applications. The research aims to contribute to the development of affordable SED solutions in the context of driving safety, with the ultimate goal of advancing road safety initiatives globally.
Carlos Castorena, Jesús López Ballester, Juan Antonio De Rus, Francesc J. Ferri, Maximo Cobos
EATIS5
2024 Deep Learning Based AoA and AoD Estimation for Millimeter Wave MIMO Systems
abstract
In this work, we propose using a deep learning method for parametric millimeter-wave (mmWave) channel estimation, specifically focusing on angle-of-arrival (AoA) and angle-of-departure (AoD) parameters in the frequency domain. Channel estimation is fundamental in mmWave to efficiently implement beamforming techniques. Our approach involves adapting a residual convolutional neural network (ResNet) to the task, incorporating a technique from topological data analysis to accurately estimate angular frequencies. Additionally, we enhance the model’s performance by including a posterior model fitting to improve the probability of detection. Through simulations, we compare our ResNet-based approach with existing signal processing methods and the Crámer-Rao lower bound. Our findings demonstrate significant enhancements in system robustness, increasing the probability of detection while minimizing estimation errors.
Diego Lloria, Sandra Roger 0002, Carmen Botella-Mascarell, Maximo Cobos
EATIS4
2024 A ResNet Approach for AoA and AoD Estimation in Analog Millimeter Wave MIMO Systems
abstract
Parametric millimeter-wave (mmWave) channel estimation involves modeling the channel matrix by combining direction-dependent signal paths, exploiting the sparse nature of mmWave channels. In our study, we propose a deep learning-based approach to estimate angle-of-arrival (AoA) and angle-of-departure (AoD) parameters from input observations in the frequency domain. To address this challenge, we have adapted a residual convolutional neural network (ResNet) to this specific problem and incorporated a technique from topological data analysis, enabling us to accurately retrieve the angular frequencies. Furthermore, we have extended this basic architecture by incorporating a posterior model fitting to enhance the system performance in terms of probability of detection. In our research, we compare the ResNet and extended ResNet approaches with state-of-the-art signal processing techniques and the Crámer-Rao lower bound through simulation. Our results indicate significant improvements in system robustness by increasing the probability of detection while maintaining a reduced estimation error.
Diego Lloria, Sandra Roger 0002, Carmen Botella-Mascarell, Maximo Cobos, Tommy Svensson
PIMRC4
2023 Acoustic Source Localization in the Spherical Harmonics Domain Exploiting Low-Rank Approximations
abstract
Acoustic signal processing in the spherical harmonics domain (SHD) is an active research area that exploits the signals acquired by higher order microphone arrays. A very important task is that concerning the localization of active sound sources. In this paper, we propose a simple yet effective method to localize prominent acoustic sources in adverse acoustic scenarios. By using a proper normalization and arrangement of the estimated spherical harmonic coefficients, we exploit low-rank approximations to estimate the far field modal directional pattern of the dominant source at each time-frame. The experiments confirm the validity of the proposed approach, with superior performance compared to other recent SHD-based approaches.
Maximo Cobos, Mirco Pezzoli, Fabio Antonacci, Augusto Sarti
ICASSP1
2023 Zero-Shot Anomalous Sound Detection in Domestic Environments Using Large-Scale Pretrained Audio Pattern Recognition Models
abstract
Anomalous sound detection is central to audio-based surveillance and monitoring. In a domestic environment, however, the classes of sounds to be considered anomalous are situation-dependent and cannot be determined in advance. At the same time, it is not feasible to expect a demanding labeling effort from the end user. To address these problems, we present a novel zero-shot method relying on an auxiliary large-scale pretrained audio neural network in support of an unsupervised anomaly detector. The auxiliary module is tasked to generate a fingerprint for each sound occasionally registered by the user. These fingerprints are then compared with those extracted from the input audio stream, and the resulting similarity score is used to increase or reduce the sensitivity of the base detector. Experimental results on synthetic data show that the proposed method substantially improves upon the unsupervised base detector and is capable of outperforming existing few-shot learning systems developed for machine condition monitoring without involving additional training.
Alessandro Ilic Mezza, Giulio Zanetti, Maximo Cobos, Fabio Antonacci
ICASSP3
2023 Towards the Creation of Scalable Tools for automatic Quality of Experience Evaluation and a Multi-Purpose Dataset for Affective Computing
abstract
Traditional tools used to evaluate the Quality of Experience (QoE) of users after browsing an ad, using a product, or performing any kind of task typically involves surveys, user testing, and analytics. However, these methods provide limited insights and have limitations due to the need of users’ active cooperation and sincerity, the long testing time, the high cost, and the limited scalability. On this work we present the tools we are developing to automatically evaluate QoE in different use cases such as dashboards that show on real time reactions to different events in the form of emotions and affections predicted by different models based on physiological data. To develop these tools, we require datasets on affective computing. We highlight some limitations of the available ones, the difficulties during the creation of such data, and our current work in the confection of a new one with automatic annotation of ground truth.
Juan Antonio De Rus, Mario Montagud, Maximo Cobos
IMX3
2023 Towards the Creation of Tools for Automatic Quality of Experience Evaluation with Focus on Interactive Virtual Environments
abstract
This paper contains the research proposal of Juan Antonio De Rus presented at the IMX 23 Doctoral Symposium. Virtual Reality (VR) applications are already used to support diverse tasks such as online meetings, education, or training, and the usages grow every year. To enrich the experience VR scenarios, include multimodal content (video, audio, text, synthetic content) and multi-sensory stimuli are typically included. Tools to evaluate the Quality of Experience (QoE) of such scenarios are needed. Traditional tools used to evaluate the QoE of users performing any kind of task typically involves surveys, user testing or analytics. However, these methods provide limited insights for our tasks with VR and have shortcomings and a limited scalability. In this doctoral study we have formulated a set of open research questions and objectives on which we plan to generate contributions and knowledge in the field of Affective Computing (AC) and Multimodal Interactive Virtual Environments. Hence, in this paper we present a set of tools we are developing to automatically evaluate QoE in different use cases. They include dashboards to monitor in real time reactions to different events in the form of emotions and affections predicted by different models based on physiological data, as well as the creation of a dataset for AC and its associated methodology.
Juan Antonio De Rus, Mario Montagud, Maximo Cobos
IMX3
2023 AI-IoT Platform for Blind Estimation of Room Acoustic Parameters Based on Deep Neural Networks
abstract
Room acoustical parameters have been widely used to describe sound perception in indoor environments, such as concert halls, conference rooms, etc. Many of them have been standardized and often have a high computational demand. With the increasing presence of deep learning approaches in automatic monitoring systems, wireless acoustic sensor networks (WASNs) offer great potential to facilitate the estimation of such parameters. In this scenario, convolutional neural networks (CNNs) offer significant reductions in the computational requirements for in-node parameter predictions, enabling the so-called Artificial Intelligence-Internet of Things (AI-IoT). In this article, we describe the design and analysis of a CNN trained to predict simultaneously a set of common room acoustical parameters directly from speech signals, without the need for specific impulse response measurements. The results show that the proposed CNN-based prediction of room acoustical parameters and speech intelligibility achieves a relative error rate of less than a 5.5%, accompanied by a computational speedup factor close to 250 with respect to the conventional signal processing approach.
Jesús López Ballester, Santiago Felici-Castell, Jaume Segura-Garcia, Maximo Cobos
IEEE Internet Things J.4
2022 Acceleration of the TSDCE MIMO Channel Estimation Algorithm on a Multi-core Platform
abstract
The use of Multi-Processor System-on-Chip (MPSoC) is becoming widespread in a huge number of signal processing systems, including wireless communications and vehicular technology applications. In those scenarios, when Multiple-Input Multiple-Output (MIMO) communication schemes are considered, the system usually has to deal with a high number of communication links that involve sensors and antennas from different vehicles and users. The use of MIMO systems with a high number of antennas increases the complexity of many signal processing algorithms which could benefit from computationally efficient implementations. The Xilinx Zynq UltraScale+ EG Heterogeneous MPSoC is a well-positioned platform to manage computationally-demanding communication systems. This platform holds a dual-core ARM Cortex-R5, a quad-core ARM Cortex-A53, a graphics processing unit (GPU) and a high-end Field Programmable Gate Array (FPGA). In particular, this work aims to evaluate the computational performance of the Transformed Spatial Domain Channel Estimation (TSDCE), a novel millimeter-wave MIMO channel estimation algorithm, on the proposed embedded platform. This work focuses firstly on developing an efficient sequential implementation that runs on the ARM Cortex-A53, so that we can afterwards leverage the use of the multi-core system to accelerate the sequential performance.
Pablo M. Aviles, Diego Lloria, Jose A. Belloch, Sandra Roger 0002, Almudena Lindoso, Maximo Cobos
EATIS6
2022 Sparsity-Based Sound Field Separation in the Spherical Harmonics Domain
abstract
Sound field analysis and reconstruction has been a topic of intense research in the last decades for its multiple applications in spatial audio processing tasks. In this context, the identification of the direct and reverberant sound field components is a problem of great interest, where several solutions exploiting spherical harmonics representations have already been proposed. However, the available techniques demand a large number of high-order microphones (HOMs) and high computational power in order to fulfill the necessary spatial sampling requirements, which can only be reduced by prior information obtained through acoustic measurements. Inspired by compressed sensing approaches, this paper proposes an alternative sparse formulation for estimating the exterior and interior sound field components in the spherical harmonics domain that allows to reduce hardware requirements without the need for additional acoustic measurements. The results show that a considerable reduction in the number of HOMs can be achieved while improving the estimation of the sound field components.
Mirco Pezzoli, Maximo Cobos, Fabio Antonacci, Augusto Sarti
ICASSP2
2022 AI-assisted affective computing and spatial audio for interactive multimodal virtual environments: research proposal
Juan Antonio De Rus, Mario Montagud, Maximo Cobos
MMSys3
2022 An Open-Set Recognition and Few-Shot Learning Dataset for Audio Event Classification in Domestic Environments
abstract
The problem of training with a small set of positive samples is known as few-shot learning (FSL). It is widely known that traditional deep learning algorithms usually show very good performance when trained with large datasets. However, in many applications, it is not possible to obtain such a high number of samples. This paper deals with the application of FSL to the detection of specific and intentional acoustic events given by different types of sound alarms, such as door bells or fire alarms, using a limited number of samples. These sounds typically occur in domestic environments where many events corresponding to a wide variety of sound classes take place. Therefore, the detection of such alarms in a practical scenario can be considered an open-set recognition (OSR) problem. To address the lack of a dedicated public dataset for audio FSL, researchers usually make modifications on other available datasets. This paper is aimed at providing the audio recognition community with a carefully annotated dataset1 for FSL in an OSR context comprised of 1360 clips from 34 classes divided into pattern sounds and unwanted sounds. To facilitate and promote research on this area, results with state-of-the-art baseline systems based on transfer learning are also presented.
Javier Naranjo-Alcazar, Sergi Perez-Castanos, Pedro Zuccarello, Ana M. Torres, José J. López, Francesc J. Ferri, Maximo Cobos
Pattern Recognit. Lett.7
2022 Performance analysis of a millimeter wave MIMO channel estimation method in an embedded multi-core processor
abstract
Abstract The emerging Multi-Processor System-on-Chip (MPSoC) technology, which combines heterogeneous computing with the high performance of field programmable gate arrays (FPGA), is a promising platform for a large number of applications, including wireless communications and vehicular technology. In this specific application context, when multiple-input multiple-output (MIMO) scenarios are considered, the system usually has to manage a large number of communication links among sensors and antennas involving different vehicles and users. Millimeter wave (mmWave) communications are one of the key technology enablers toward achieving high data rates in beyond 5G systems (B5G). Communication at these frequency bands usually involves the use of large antenna arrays, often requiring high computational resources. One of the candidate platforms able to manage a huge number of communications is the Xilinx Zynq UltraScale+ EG Heterogeneous MPSoC, which is composed of a dual-core Cortex-R5, a quad-core ARM Cortex-A53, a graphics processing unit (GPU) and a high-end FPGA. This work analyzes the computational performance that requires a recent mmWave MIMO channel estimation algorithm in a platform of this kind. As a first approach, we will focus our work on the performance that can be achieved via the quad-core ARM Cortex-A53. To this end, we will use the libraries for numerical algebra (BLAS and LAPACK). The results show that our reference implementation is able to manage a large MIMO communication system with 256 antennas without exhausting platform resources.
Pablo M. Aviles, Diego Lloria, Jose A. Belloch, Sandra Roger 0002, Almudena Lindoso, Maximo Cobos
J. Supercomput.6
2021 Low-complexity AoA and AoD Estimation in the Transformed Spatial Domain for Millimeter Wave MIMO Channels
abstract
High-accuracy angle of arrival (AoA) and angle of departure (AoD) estimation is critical for cell search, stable communications and positioning in millimeter wave (mmWave) cellular systems. Moreover, the design of low-complexity AoA/AoD estimation algorithms is also of major importance in the deployment of practical systems to enable a fast and resource-efficient computation of beamforming weights. Parametric mmWave channel estimation allows to describe the channel matrix as a combination of direction-dependent signal paths, exploiting the sparse characteristics of mmWave channels. In this context, a fast Transformed Spatial Domain Channel Estimation (TSDCE) algorithm was recently proposed to perform parametric channel estimation with low complexity, which in turn results in a full characterization of the transmitting and receiving angles for dominant signal paths. In this paper, we analyze the AoA/AoD estimation capability and accuracy of the TSDCE algorithm in detail. We find that the TSDCE algorithm has a significant performance advantage with respect to the traditional approach, which is based on frequency domain processing, in complexity-constrained environments, especially at high signal-to-noise ratios.
Sandra Roger 0002, Carmen Botella-Mascarell, Diego Lloria, Maximo Cobos, Gábor Fodor 0001
PIMRC4
2021 Ray-Space-Based Multichannel Nonnegative Matrix Factorization for Audio Source Separation
abstract
Nonnegative matrix factorization (NMF) has been traditionally considered a promising approach for audio source separation. While standard NMF is only suited for single-channel mixtures, extensions to consider multi-channel data have been also proposed. Among the most popular alternatives, multichannel NMF (MNMF) and further derivations based on constrained spatial covariance models have been successfully employed to separate multi-microphone convolutive mixtures. This letter proposes a MNMF extension by considering a mixture model with Ray-Space-transformed signals, where magnitude data successfully encodes source locations as frequency-independent linear patterns. We show that the MNMF algorithm can be seamlessly adapted to consider Ray-Space-transformed data, providing competitive results with recent state-of-the-art MNMF algorithms in a number of configurations using real recordings.
Mirco Pezzoli, Julio J. Carabias-Orti, Maximo Cobos, Fabio Antonacci, Augusto Sarti
IEEE Signal Process. Lett.3
2021 Fast Channel Estimation in the Transformed Spatial Domain for Analog Millimeter Wave Systems
abstract
Fast channel estimation in millimeter-wave (mmWave) systems is a fundamental enabler of high-gain beamforming, which boosts coverage and capacity. The channel estimation stage typically involves an initial beam training process where a subset of the possible beam directions at the transmitter and receiver is scanned along a predefined codebook. Unfortunately, the high number of transmit and receive antennas deployed in mmWave systems increase the complexity of the beam selection and channel estimation tasks. In this work, we tackle the channel estimation problem in analog systems from a different perspective than used by previous works. In particular, we propose to move the channel estimation problem from the angular domain into the transformed spatial domain, in which estimating the angles of arrivals and departures corresponds to estimating the angular frequencies of paths constituting the mmWave channel. The proposed approach, referred to as transformed spatial domain channel estimation (TSDCE) algorithm, exhibits robustness to additive white Gaussian noise by combining low-rank approximations and sample autocorrelation functions for each path in the transformed spatial domain. Numerical results evaluate the mean square error of the channel estimation and the direction of arrival estimation capability. TSDCE significantly reduces the first, while exhibiting a remarkably low computational complexity compared with well-known benchmarking schemes.
Sandra Roger 0002, Maximo Cobos, Carmen Botella-Mascarell, Gábor Fodor 0001
IEEE Trans. Wirel. Commun.2
2020 Time Difference of Arrival Estimation from Frequency-Sliding Generalized Cross-Correlations Using Convolutional Neural Networks
abstract
The interest in deep learning methods for solving traditional signal processing tasks has been steadily growing in the last years. Time delay estimation (TDE) in adverse scenarios is a challenging problem, where classical approaches based on generalized cross-correlations (GCCs) have been widely used for decades. Recently, the frequency-sliding GCC (FS-GCC) was proposed as a novel technique for TDE based on a sub-band analysis of the cross-power spectrum phase, providing a structured two-dimensional representation of the time delay information contained across different frequency bands. Inspired by deep-learning-based image denoising solutions, we propose in this paper the use of convolutional neural networks (CNNs) to learn the time-delay patterns contained in FS-GCCs extracted in adverse acoustic conditions. Our experiments confirm that the proposed approach provides excellent TDE performance while being able to generalize to different room and sensor setups.
Luca Comanducci, Maximo Cobos, Fabio Antonacci, Augusto Sarti
ICASSP2
2020 Performance comparison of container orchestration platforms with low cost devices in the fog, assisting Internet of Things applications
Rafael Fayos-Jordan, Santiago Felici-Castell, Jaume Segura-Garcia, Jesús López Ballester, Maximo Cobos
J. Netw. Comput. Appl.5
2020 Frequency-Sliding Generalized Cross-Correlation: A Sub-Band Time Delay Estimation Approach
abstract
The generalized cross-correlation(GCC) is regarded as the most popular approach for estimating the time difference of arrival (TDOA) between the signals received at two sensors. Time delay estimates are obtained by maximizing the GCC output, where the direct-path delay is usually observed as a prominent peak. Moreover, GCCs play also an important role in steered response power (SRP) localization algorithms, where the SRP functional can be written as an accumulation of the GCCs computed from multiple sensor pairs. Unfortunately, the accuracy of TDOA estimates is affected by multiple factors, including noise, reverberation and signal bandwidth. In this paper, a sub-band approach for time delay estimation aimed at improving the performance of the conventional GCC is presented. The proposed method is based on the extraction of multiple GCCs corresponding to different frequency bands of the cross-power spectrum phase in a sliding-window fashion. The major contributions of this paper include: 1) a sub-band GCC representation of the cross-power spectrum phase that, despite having a reduced temporal resolution, provides a more suitable representation for estimating the true TDOA; 2) such matrix representation is shown to be rank one in the ideal noiseless case, a property that is exploited in more adverse scenarios to obtain a more robust and accurate GCC; 3) we propose a set of low-rank approximation alternatives for processing the sub-band GCC matrix, leading to better TDOA estimates and source localization performance. An extensive set of experiments is presented to demonstrate the validity of the proposed approach.
Maximo Cobos, Fabio Antonacci, Luca Comanducci, Augusto Sarti
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Adaptive Distance-Based Pooling in Convolutional Neural Networks for Audio Event Classification
abstract
In the last years, deep convolutional neural networks have become a standard for the development of state-of-the-art audio classification systems, taking the lead over traditional approaches based on feature engineering. While they are capable of achieving human performance under certain scenarios, it has been shown that their accuracy is severely degraded when the systems are tested over noisy or weakly segmented events. Although better generalization could be obtained by increasing the size of the training dataset, e.g. by applying data augmentation techniques, this also leads to longer and more complex training procedures. In this article, we propose a new type of pooling layer aimed at compensating non-relevant information of audio events by applying an adaptive transformation of the convolutional feature maps in the temporal axis. The proposed layer performs a non-linear temporal transformation that follows a uniform distance subsampling criterion on the learned feature space. The experiments conducted over different datasets show significant performance improvements when the proposed layer is added to baseline models, resulting in systems that generalize better to mismatching test conditions and learn more robustly from weakly labeled data.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Sound Event Envelope Estimation in Polyphonic Mixtures
abstract
Sound event detection is the task of identifying automatically the presence and temporal boundaries of sound events within an input audio stream. In the last years, deep learning methods have established themselves as the state-of-the-art approach for the task, using binary indicators during training to denote whether an event is active or inactive. However, such binary activity indicators do not fully describe the events, and estimating the envelope of the sounds could provide more precise modeling of their activity. This paper proposes to estimate the amplitude envelopes of target sound event classes in polyphonic mixtures. For training, we use the amplitude envelopes of the target sounds, calculated from mixture signals and, for comparison, from their isolated counterparts. The model is then used to perform envelope estimation and sound event detection. Results show that the envelope estimation allows good modeling of the sounds activity, with detection results comparable to current state-of-the art.
Irene Martín-Morató, Annamaria Mesaros, Toni Heittola, Tuomas Virtanen, Maximo Cobos, Francesc J. Ferri
ICASSP5
2019 Practical Considerations for Acoustic Source Localization in the IoT Era: Platforms, Energy Efficiency, and Performance
abstract
The rapid development of the Internet of Things (IoT) has posed important changes in the way emerging acoustic signal processing applications are conceived. While traditional acoustic processing applications have been developed taking into account high-throughput computing platforms equipped with expensive multichannel audio interfaces, the IoT paradigm is demanding the use of more flexible and energy-efficient systems. In this context, algorithms for source localization and ranging in wireless acoustic sensor networks can be considered an enabling technology for many IoT-based environments, including security, industrial, and health-care applications. This paper is aimed at evaluating important aspects dealing with the practical deployment of IoT systems for acoustic source localization. Recent systems-on-chip composed of low-power multicore processors, combined with a small graphics accelerator (or GPU), yield a notable increment of the computational capacity needed in intensive signal processing algorithms while partially retaining the appealing low power consumption of embedded systems. Different algorithms and implementations over several state-of-the-art platforms are discussed, analyzing important aspects, such as the tradeoffs between performance, energy efficiency, and exploitation of parallelism by taking into account real-time constraints.
Jose A. Belloch, José M. Badía, Francisco D. Igual, Maximo Cobos
IEEE Internet Things J.4
2019 Accelerating the SRP-PHAT algorithm on multi- and many-core platforms using OpenCL
José M. Badía, Jose A. Belloch, Maximo Cobos, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.3
2018 On the Design of Probe Signals in Wireless Acoustic Sensor Networks Self-Positioning Algorithms
abstract
A wireless acoustic sensor network comprises a distributed group of devices equipped with audio transducers. Typically, these devices can interoperate with each other using wireless links and perform collaborative audio signal processing. Ranging and self-positioning of the network nodes are examples of tasks that can be carried out collaboratively using acoustic signals. However, the environmental conditions can distort the emitted signals and complicate the ranging process. In this context, the selection of proper acoustic signals can facilitate the attainment of this goal and improve the localization accuracy. This letter deals with the design and evaluation of acoustic probe signals allowing the implementation of ranging applications using unsynchronized audio signals acquired by different network nodes. The proposed signals have been evaluated in diverse simulated environments and implemented in a real deployment. The results confirm the suitability of these signals for implementing localization applications with an accuracy that is in the order of 1 cm.
Juan José Pérez Solano, Maximo Cobos, Jaume Segura-Garcia, Santiago Felici-Castell
IEEE Signal Process. Lett.2
2018 Adaptive Mid-Term Representations for Robust Audio Event Classification
abstract
Low-level audio features are commonly used in many audio analysis tasks, such as audio scene classification or acoustic event detection. Due to the variable length of audio signals, it is a common approach to create fixed-length feature vectors consisting of a set of statistics that summarize the temporal variability of such short-term features. To avoid the loss of temporal information, the audio event can be divided into a set of mid-term segments or texture windows. However, such an approach requires to estimate accurately the onset and offset times of the audio events in order to obtain a robust mid-term statistical description of their temporal evolution. This paper proposes the use of an alternative event representation based on nonlinear time normalization prior to the extraction of mid-term statistics. The short-term features are transformed into a new fixed-length representation that considers uniform distance subsampling over a defined feature space in contrast to the classical short-term temporal framing. The results show that the use of distance-based texture windows provides an improved statistical description of the event robust to errors in the event segmentation stage under noisy conditions.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Analysis of data fusion techniques for multi-microphone audio event detection in adverse environments
abstract
Acoustic event detection (AED) is currently a very active research area with multiple applications in the development of smart acoustic spaces. In this context, the advances brought by Internet of Things (IoT) platforms where multiple distributed microphones are available have also contributed to this interest. In such scenarios, the use of data fusion techniques merging information from several sensors becomes an important aspect in the design of multi-microphone AED systems. In this paper, we present a preliminary analysis of several data-fusion techniques aimed at improving the recognition accuracy of an AED system by taking advantage of the diversity provided by multiple microphones in adverse acoustic conditions. The results confirm that, under appropriate processing schemes, the recognition rate can be increased as well as the corresponding independence on event location.
Irene Martín-Morató, Maximo Cobos, Francesc J. Ferri
MMSP2
2017 Steered Response Power Localization of Acoustic Passband Signals
abstract
The vast majority of localization approaches using phase transform (PHAT) consider that the sources of interest are wideband low-pass sources. While this may be the usual case for common audio signals such as speech, PHAT methods are affected negatively by modulation artifacts when the sources to be localized are passband signals. In these cases, steered response power PHAT localization becomes less robust. This letter analyzes the form of generalized cross-correlation functions with PHAT when passband acoustic signals are considered, proposing approaches for increasing the localization performance through the mitigation of these negative effects.
Maximo Cobos, Miguel Garcia 0001, Miguel Arevalillo-Herráez
IEEE Signal Process. Lett.1
2017 A Parallel Approach to HRTF Approximation and Interpolation Based on a Parametric Filter Model
abstract
Spatial audio-rendering techniques using head-related transfer functions (HRTFs) are currently used in many different contexts such as immersive teleconferencing systems, gaming, or 3-D audio reproduction. Since all these applications usually involve real-time constraints, efficient processing structures for HRTF modeling and interpolation are necessary for providing real-time binaural audio solutions. This letter presents a parametric parallel model that allows us to perform HRTF filtering and interpolation efficiently from an input HRTF dataset. The resulting model, which is an adaptation from a recently proposed modeling technique, not only reduces the size of HRTF datasets significantly, but also allows for simplified interpolation and real-time computation over parallel processors. In order to discuss the suitability of this new model, an implementation over a graphic processing unit is presented.
Germán Ramos, Maximo Cobos, Balázs Bank, Jose A. Belloch
IEEE Signal Process. Lett.2
2017 A Robust Wrap Reduction Algorithm for Fringe Projection Profilometry and Applications in Magnetic Resonance Imaging
abstract
In this paper, we present an effective algorithm to reduce the number of wraps in a 2D phase signal provided as input. The technique is based on an accurate estimate of the fundamental frequency of a 2D complex signal with the phase given by the input, and the removal of a dependent additive term from the phase map. Unlike existing methods based on the discrete Fourier transform (DFT), the frequency is computed by using noise-robust estimates that are not restricted to integer values. Then, to deal with the problem of a non-integer shift in the frequency domain, an equivalent operation is carried out on the original phase signal. This consists of the subtraction of a tilted plane whose slope is computed from the frequency, followed by a re-wrapping operation. The technique has been exhaustively tested on fringe projection profilometry (FPP) and magnetic resonance imaging (MRI) signals. In addition, the performance of several frequency estimation methods has been compared. The proposed methodology is particularly effective on FPP signals, showing a higher performance than the state-of-the-art wrap reduction approaches. In this context, it contributes to canceling the carrier effect at the same time as it eliminates any potential slope that affects the entire signal. Its effectiveness on other carrier-free phase signals, e.g., MRI, is limited to the case that inherent slopes are present in the phase data.
Miguel Arevalillo-Herráez, Maximo Cobos, Miguel Garcia 0001
IEEE Trans. Image Process.2
2017 A Survey of Sound Source Localization Methods in Wireless Acoustic Sensor Networks
abstract
Wireless acoustic sensor networks (WASNs) are formed by a distributed group of acoustic-sensing devices featuring audio playing and recording capabilities. Current mobile computing platforms offer great possibilities for the design of audio-related applications involving acoustic-sensing nodes. In this context, acoustic source localization is one of the application domains that have attracted the most attention of the research community along the last decades. In general terms, the localization of acoustic sources can be achieved by studying energy and temporal and/or directional features from the incoming sound at different microphones and using a suitable model that relates those features with the spatial location of the source (or sources) of interest. This paper reviews common approaches for source localization in WASNs that are focused on different types of acoustic features, namely, the energy of the incoming signals, their time of arrival (TOA) or time difference of arrival (TDOA), the direction of arrival (DOA), and the steered response power (SRP) resulting from combining multiple microphone signals. Additionally, we discuss methods not only aimed at localizing acoustic sources but also designed to locate the nodes themselves in the network. Finally, we discuss current challenges and frontiers in this field.
Maximo Cobos, Fabio Antonacci, Anastasios Alexandridis, Athanasios Mouchtaris, Bowon Lee
Wirel. Commun. Mob. Comput.1
2017 Wireless Acoustic Sensor Networks and Applications
Maximo Cobos, Fabio Antonacci, Athanasios Mouchtaris, Bowon Lee
Wirel. Commun. Mob. Comput.1
2015 On the performance of multi-GPU-based expert systems for acoustic localization involving massive microphone arrays
Jose A. Belloch, Alberto González 0001, Antonio M. Vidal, Maximo Cobos
Expert Syst. Appl.4
2015 Subjective quality assessment of multichannel audio accompanied with video in representative broadcasting genres
Maximo Cobos, José J. López, Juan M. Navarro, Germán Ramos
Multim. Syst.1
2014 Cumulative-sum-based localization of sound events in low-cost wireless acoustic sensor networks
abstract
Wireless acoustic sensor networks (WASNs) are known for their potential applications in multiple areas, such as audio-based surveillance, binaural hearing aids or advanced acoustic monitoring. The knowledge of the spatial position of a source of interest is usually a requirement for many of these applications. Therefore, source localization is an important problem to be addressed in WASNs. Unfortunately, most localization algorithms need costly signal processing stages that prevent them from being implemented in low-cost sensor networks, requiring additional modules for signal acquisition and processing. This paper presents a low-complexity method for acoustic event detection and localization considering a change detection statistical framework. Two possible implementation approaches based on the efficient cumulative sum (CUSUM) algorithm are presented and discussed. Results from simulations and a real deployment show that the proposed techniques can be easily implemented in low-cost sensor networks, providing good localization accuracy and making good use of the available node resources.
Maximo Cobos, Juan José Pérez Solano, Santiago Felici-Castell, Jaume Segura-Garcia, Juan M. Navarro
IEEE ACM Trans. Audio Speech Lang. Process.1
2012 Maximum a Posteriori Binary Mask Estimation for Underdetermined Source Separation Using Smoothed Posteriors
abstract
Sound source separation has become a topic of intensive research in the last years. The research effort has been specially relevant for the underdetermined case, where a considerable number of sparse methods working in the time-frequency (T-F) domain have appeared. In this context, although binary masking seems to be a preferred choice for source demixing, the estimated masks differ substantially from the ideal ones. This paper proposes a maximum a posteriori (MAP) framework for binary mask estimation. To this end, class-conditional source probabilities according to the observed mixing parameters are modeled via ratios of dependent Cauchy distributions while source priors are iteratively calculated from the observed histograms. Moreover, spatially smoothed posteriors in the T-F domain are proposed to avoid noisy estimates, showing that the estimated masks are closer to the ideal ones in terms of objective performance measures.
Maximo Cobos, José J. López
IEEE Trans. Speech Audio Process.1
2011 Real time speaker localization and detection system for camera steering in multiparticipant videoconferencing environments
abstract
A real time speaker localization and detection system for videoconferencing environments is presented. In this system, a recently proposed modified Steered Response Power - Phase Transform (SRP-PHAT) algorithm has been used as the core processing scheme. The new SRP-PHAT functional has been shown to provide robust localization performance in indoor environments without the need for having a very fine spatial grid, thus reducing the computational cost required in a practical implementation. Moreover, it has been demonstrated that the statistical distribution of location estimates when a speaker is active can be successfully used to discriminate between speech and non-speech frames by using a criterion of peakedness. As a result, talking participants can be detected and located with significant accuracy following a common processing framework.
Amparo Marti, Maximo Cobos, José J. López
ICASSP2
2011 Computer-based detection and classification of flaws in citrus fruits
José J. López, Maximo Cobos, Emanuel Aguilera
Neural Comput. Appl.2
2011 A Modified SRP-PHAT Functional for Robust Real-Time Sound Source Localization With Scalable Spatial Sampling
abstract
The Steered Response Power - Phase Transform (SRP-PHAT) algorithm has been shown to be one of the most robust sound source localization approaches operating in noisy and reverberant environments. However, its practical implementation is usually based on a costly fine grid-search procedure, making the computational cost of the method a real issue. In this letter, we introduce an effective strategy that extends the conventional SRP-PHAT functional with the aim of considering the volume surrounding the discrete locations of the spatial grid. As a result, the modified functional performs a full exploration of the sampled space rather than computing the SRP at discrete spatial positions, increasing its robustness and allowing for a coarser spatial grid. To this end, the Generalized Cross-Correlation (GCC) function corresponding to each microphone pair must be properly accumulated according to the defined microphone setup. Experiments carried out under different acoustic conditions confirm the validity of the proposed approach.
Maximo Cobos, Amparo Marti, José J. López
IEEE Signal Process. Lett.1
2009 Defect Detection and Classification in Citrus Using Computer Vision
José J. López, Emanuel Aguilera, Maximo Cobos
ICONIP (2)3
2008 Improving Isolation of Blindly Separated Sources Using Time-Frequency Masking
abstract
A refinement technique based on time-frequency masking is proposed to improve source isolation in blind audio source separation algorithms. The refinement technique uses an energy-normalized source-to-interference ratio in order to identify and eliminate interfering energy from the extracted sources. Some examples using this refinement method with different separation algorithms are discussed. The results show that source isolation can be significantly enhanced with negligible degradation of the separated sources.
Maximo Cobos, José J. López
IEEE Signal Process. Lett.1