EDBT 2026 Demo / reviewers in the wild / expert
Juan Azcarreta
dblp:204/3478 · also Juan Azcarreta Ortiz
· DBLP profile ↗
11ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sound Event Detection With Boundary-Aware Optimization and InferenceabstractTemporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers—Recurrent Event Detection (RED) and Event Proposal Network (EPN)—which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes. Florian Schmid, Chi Ian Tang, Sanjeel Parekh, Vamsi K. Ithapu, Juan Azcarreta, Giacomo Ferroni, Yijun Qian, Arnoldas Jasonas, Cosmin Frateanu, Camilla Clark, Gerhard Widmer, Cagdas Bilen |
IEEE Signal Process. Lett. | 5 |
| 2025 | Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder RefinementabstractDeploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their computational cost limits their feasibility on embedded platforms. This work presents an efficient end-to-end SE framework that leverages a Differentiable Digital Signal Processing (DDSP) vocoder for high-quality speech synthesis. First, a compact neural network predicts enhanced acoustic features from noisy speech: spectral envelope, fundamental frequency (F0), and periodicity. These features are fed into the DDSP vocoder to synthesize the enhanced waveform. The system is trained end-to-end with STFT and adversarial losses, enabling direct optimization at the feature and waveform levels. Experimental results show that our method improves intelligibility and quality by 4% (STOI) and 19% (DNSMOS) over strong baselines without significantly increasing computation, making it well-suited for real-time applications. Heitor R. Guimarães, Ke Tan 0001, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey 0004, Buye Xu |
ASRU | 3 |
| 2025 | Advancing Active Speaker Detection for Egocentric VideosabstractThis paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing the lip region in the visual input, and (ii) applying motion blur augmentation. These methods significantly enhance the model’s performance in handling the challenges typical of egocentric videos. We showcase the effectiveness of these techniques on a simple but efficient causal audio-visual model. The proposed model, named EgoASD, demonstrates state-of-the-art performance on the EasyCom dataset, beating the previous SOTA by 1.7% mean Average Precision (mAP) with a model 2.5 times smaller. Our ablations highlight the importance of visual input, motion blur augmentation, the pretraining method and the importance of temporal context. To demonstrate its applicability in the real world, we apply our model to audio-visual speaker diarization, outperforming other baselines on EasyCom. Jaesung Huh, Juan Azcarreta, Anurag Kumar 0003, Ashutosh Pandey 0004, Ali Aroudi, Daniel D. E. Wong, Francesco Nesta, Buye Xu, Jacob Donley |
ICASSP | 2 |
| 2025 | Ultra low-compute complex spectral masking for multichannel speech enhancementabstractWe present a streamlined framework for complex spectral masking that processes multichannel speech with minimal computational demands, enhancing both spectral magnitude and phase by integrating low-compute models with the Multi-Channel Wiener Filter (MCWF). Our methodology employs a two-stage, end-to-end training approach where a deep neural network (DNN) first estimates MCWF weights, followed by another DNN that refines the MCWF output, enhancing spectral masking quality. This architecture not only outperforms the traditional oracle Minimum Variance Distortionless Response (MVDR) beamformer but also maintains high efficiency, requiring less than 50MMACs for processing one second of 8-channel audio. Empirical results demonstrate that our framework exceeds the performance of existing low-compute models, offering significant enhancements with minimal computational demands, making it ideal for deployment on edge devices with limited computational resources. Ashutosh Pandey 0004, Juan Azcarreta |
ICASSP | 2 |
| 2025 | Efficient Neural and Numerical Methods for High-QualityOnline Speech Spectrogram Inversion via Gradient Theorem
Andres Fernandez, Juan Azcarreta, Cagdas Bilen, Jesus Monge-Alvarez |
INTERSPEECH | 2 |
| 2025 | A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
Yuan-Kuei Wu, Juan Azcarreta, Kashyap Patel, Buye Xu, Jung-Suk Lee, Sanha Lee, Ashutosh Pandey 0004 |
INTERSPEECH | 2 |
| 2024 | All Neural Low-latency Directional Speech Extraction
Ashutosh Pandey 0004, Sanha Lee, Juan Azcarreta, Buye Xu |
INTERSPEECH | 3 |
| 2021 | Improving Sound Event Detection Metrics: Insights from DCASE 2020abstractThe ranking of sound event detection (SED) systems may be biased by assumptions inherent to evaluation criteria and to the choice of an operating point. This paper compares conventional event-based and segment-based criteria against the Polyphonic Sound Detection Score (PSDS)'s intersection-based criterion, over a selection of systems from DCASE 2020 Challenge Task 4. It shows that, by relying on collars, the conventional event-based criterion introduces different strictness levels depending on the length of the sound events, and that the segment-based criterion may lack precision and be application dependent. Alternatively, PSDS's intersection-based criterion overcomes the dependency of the evaluation on sound event duration and provides robustness to labelling subjectivity, by allowing valid detections of interrupted events. Furthermore, PSDS enhances the comparison of SED systems by measuring sound event modelling performance independently from the systems' operating points. Giacomo Ferroni, Nicolas Turpault, Juan Azcarreta, Francesco Tuveri, Romain Serizel, Cagdas Bilen, Sacha Krstulovic |
ICASSP | 3 |
| 2020 | A Framework for the Robust Evaluation of Sound Event DetectionabstractThis work defines a new framework for performance evaluation of polyphonic sound event detection (SED) systems, which overcomes the limitations of the conventional collar-based event decisions, event F-scores and event error rates. The proposed framework introduces a definition of event detection that is more robust against labelling subjectivity. It also resorts to polyphonic receiver operating characteristic (ROC) curves to deliver more global insight into system performance than F1-scores, and proposes a reduction of these curves into a single polyphonic sound detection score (PSDS), which allows system comparison independently from operating points (OPs). The presented method also delivers better insight into data biases and classification stability across sound classes. Furthermore, it can be tuned to varying applications in order to match a variety of user experience requirements. The benefits of the proposed approach are demonstrated by re-evaluating the baseline and two of the top-performing systems from DCASE 2019 Task 4. Cagdas Bilen, Giacomo Ferroni, Francesco Tuveri, Juan Azcarreta, Sacha Krstulovic |
ICASSP | 4 |
| 2018 | Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source CountingabstractHere we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effective for frequency bin-wise clustering when the number of sources is known. However, it cannot resolve the permutation ambiguity, and is inapplicable when the number of sources is unknown. The proposed PF-cGMM is an extension of the cGMM, which resolves these issues. The resolution of the permutation ambiguity can be realized by a spatial prior called a complex inverse Wishart mixture model (cIWMM). The absence of the permutation ambiguity facilitates source counting, which is performed by hierarchical clustering in this paper. Experiments showed that the PF-cGMM was able to (1) resolve the permutation ambiguity and (2) realize source separation even when the number of sources was unknown with little performance degradation compared to when it was known. Juan Azcarreta, Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 1 |
| 2017 | Hardware and software for reproducible research in audio array signal processingabstractIn our demo, we present two hardware platforms for prototyping audio array signal processing. Pyramic is a 48-channel microphone array fitted on an FPGA and Compact Six is a portable microphone array with six microphones, closer to the technical constraints of consumer electronics. A browser based interface was developed that allows the user to interact with the audio stream from the arrays in real time. The software component of this demo is a Python module with implementations of basic audio signal processing blocks and popular techniques like STFT, beamforming, and DoA. Both the hardware design files and the software are open source and freely shared. As part of a collaboration with IBM Research, their beamforming and imaging technologies will also be portrayed. The hardware will be demonstrated through an installation processing the microphone signals into light patterns on a circular LED array. The demo will be interactive and let visitors play with different algorithms for DoA (SRP, FRIDA [1], Bluebild) and beamforming (MVDR, Flexibeam [2]). The availability of an open platform with reference implementations encourages reproducible research and minimizes setup-time when testing and benchmarking new audio array signal processing algorithms. It can also serve as a useful educational tool, providing a means to work with real-life signals. Eric Bezzam, Robin Scheibler, Juan Azcarreta, Hanjie Pan, Matthieu Simeoni, Rene Beuchat, Paul Hurley, Basile Bruneau, Corentin Ferry, Sepand Kashani |
ICASSP | 3 |