Juan Azcarreta

dblp:204/3478 · also Juan Azcarreta Ortiz · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Sound Event Detection With Boundary-Aware Optimization and Inference
abstract
Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers—Recurrent Event Detection (RED) and Event Proposal Network (EPN)—which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.
Florian Schmid, Chi Ian Tang, Sanjeel Parekh, Vamsi K. Ithapu, Juan Azcarreta, Giacomo Ferroni, Yijun Qian, Arnoldas Jasonas, Cosmin Frateanu, Camilla Clark, Gerhard Widmer, Cagdas Bilen
IEEE Signal Process. Lett.5
2025 Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
abstract
Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications.
Heitor R. Guimarães, Ke Tan 0001, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey 0004, Buye Xu
ASRU3
2025 Advancing Active Speaker Detection for Egocentric Videos
abstract
This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing the lip region in the visual input, and (ii) applying motion blur augmentation. These methods significantly enhance the model’s performance in handling the challenges typical of egocentric videos. We showcase the effectiveness of these techniques on a simple but efficient causal audio-visual model. The proposed model, named EgoASD, demonstrates state-of-the-art performance on the EasyCom dataset, beating the previous SOTA by 1.7% mean Average Precision (mAP) with a model 2.5 times smaller. Our ablations highlight the importance of visual input, motion blur augmentation, the pretraining method and the importance of temporal context. To demonstrate its applicability in the real world, we apply our model to audio-visual speaker diarization, outperforming other baselines on EasyCom.
Jaesung Huh, Juan Azcarreta, Anurag Kumar 0003, Ashutosh Pandey 0004, Ali Aroudi, Daniel D. E. Wong, Francesco Nesta, Buye Xu, Jacob Donley
ICASSP2
2025 Ultra low-compute complex spectral masking for multichannel speech enhancement
abstract
We present a streamlined framework for complex spectral masking that processes multichannel speech with minimal computational demands, enhancing both spectral magnitude and phase by integrating low-compute models with the Multi-Channel Wiener Filter (MCWF). Our methodology employs a two-stage, end-to-end training approach where a deep neural network (DNN) first estimates MCWF weights, followed by another DNN that refines the MCWF output, enhancing spectral masking quality. This architecture not only outperforms the traditional oracle Minimum Variance Distortionless Response (MVDR) beamformer but also maintains high efficiency, requiring less than 50MMACs for processing one second of 8-channel audio. Empirical results demonstrate that our framework exceeds the performance of existing low-compute models, offering significant enhancements with minimal computational demands, making it ideal for deployment on edge devices with limited computational resources.
Ashutosh Pandey 0004, Juan Azcarreta
ICASSP2
2025 Efficient Neural and Numerical Methods for High-QualityOnline Speech Spectrogram Inversion via Gradient Theorem
Andres Fernandez, Juan Azcarreta, Cagdas Bilen, Jesus Monge-Alvarez
INTERSPEECH2
2025 A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
Yuan-Kuei Wu, Juan Azcarreta, Kashyap Patel, Buye Xu, Jung-Suk Lee, Sanha Lee, Ashutosh Pandey 0004
INTERSPEECH2
2024 All Neural Low-latency Directional Speech Extraction
Ashutosh Pandey 0004, Sanha Lee, Juan Azcarreta, Buye Xu
INTERSPEECH3
2021 Improving Sound Event Detection Metrics: Insights from DCASE 2020
abstract
The ranking of sound event detection (SED) systems may be biased by assumptions inherent to evaluation criteria and to the choice of an operating point. This paper compares conventional event-based and segment-based criteria against the Polyphonic Sound Detection Score (PSDS)'s intersection-based criterion, over a selection of systems from DCASE 2020 Challenge Task 4. It shows that, by relying on collars, the conventional event-based criterion introduces different strictness levels depending on the length of the sound events, and that the segment-based criterion may lack precision and be application dependent. Alternatively, PSDS's intersection-based criterion overcomes the dependency of the evaluation on sound event duration and provides robustness to labelling subjectivity, by allowing valid detections of interrupted events. Furthermore, PSDS enhances the comparison of SED systems by measuring sound event modelling performance independently from the systems' operating points.
Giacomo Ferroni, Nicolas Turpault, Juan Azcarreta, Francesco Tuveri, Romain Serizel, Cagdas Bilen, Sacha Krstulovic
ICASSP3
2020 A Framework for the Robust Evaluation of Sound Event Detection
abstract
This work defines a new framework for performance evaluation of polyphonic sound event detection (SED) systems, which overcomes the limitations of the conventional collar-based event decisions, event F-scores and event error rates. The proposed framework introduces a definition of event detection that is more robust against labelling subjectivity. It also resorts to polyphonic receiver operating characteristic (ROC) curves to deliver more global insight into system performance than F1-scores, and proposes a reduction of these curves into a single polyphonic sound detection score (PSDS), which allows system comparison independently from operating points (OPs). The presented method also delivers better insight into data biases and classification stability across sound classes. Furthermore, it can be tuned to varying applications in order to match a variety of user experience requirements. The benefits of the proposed approach are demonstrated by re-evaluating the baseline and two of the top-performing systems from DCASE 2019 Task 4.
Cagdas Bilen, Giacomo Ferroni, Francesco Tuveri, Juan Azcarreta, Sacha Krstulovic
ICASSP4
2018 Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source Counting
abstract
Here we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effective for frequency bin-wise clustering when the number of sources is known. However, it cannot resolve the permutation ambiguity, and is inapplicable when the number of sources is unknown. The proposed PF-cGMM is an extension of the cGMM, which resolves these issues. The resolution of the permutation ambiguity can be realized by a spatial prior called a complex inverse Wishart mixture model (cIWMM). The absence of the permutation ambiguity facilitates source counting, which is performed by hierarchical clustering in this paper. Experiments showed that the PF-cGMM was able to (1) resolve the permutation ambiguity and (2) realize source separation even when the number of sources was unknown with little performance degradation compared to when it was known.
Juan Azcarreta, Nobutaka Ito, Shoko Araki, Tomohiro Nakatani
ICASSP1
2017 Hardware and software for reproducible research in audio array signal processing
abstract
In our demo, we present two hardware platforms for prototyping audio array signal processing. Pyramic is a 48-channel microphone array fitted on an FPGA and Compact Six is a portable microphone array with six microphones, closer to the technical constraints of consumer electronics. A browser based interface was developed that allows the user to interact with the audio stream from the arrays in real time. The software component of this demo is a Python module with implementations of basic audio signal processing blocks and popular techniques like STFT, beamforming, and DoA. Both the hardware design files and the software are open source and freely shared. As part of a collaboration with IBM Research, their beamforming and imaging technologies will also be portrayed. The hardware will be demonstrated through an installation processing the microphone signals into light patterns on a circular LED array. The demo will be interactive and let visitors play with different algorithms for DoA (SRP, FRIDA [1], Bluebild) and beamforming (MVDR, Flexibeam [2]). The availability of an open platform with reference implementations encourages reproducible research and minimizes setup-time when testing and benchmarking new audio array signal processing algorithms. It can also serve as a useful educational tool, providing a means to work with real-life signals.
Eric Bezzam, Robin Scheibler, Juan Azcarreta, Hanjie Pan, Matthieu Simeoni, Rene Beuchat, Paul Hurley, Basile Bruneau, Corentin Ferry, Sepand Kashani
ICASSP3