VLDB 2026 Research / reviewers in the wild / expert
Alberto N. Escalante
dblp:48/1647 · also Alberto N. Escalante-B., Alberto Nicolás Escalante Bañuelos
· DBLP profile ↗
14ranked-venue papers
4as first author
6since 2021 · last 2023
0000-0002-8704-1432ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | DeepFilterNet: Perceptually Motivated Real-Time Speech Enhancement
Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
INTERSPEECH | 2 |
| 2023 | Deep Multi-Frame Filtering for Hearing Aids
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
INTERSPEECH | 3 |
| 2022 | Deepfilternet: A Low Complexity Speech Enhancement Framework for Full-Band Audio Based On Deep FilteringabstractComplex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks (CM) are usually preferred over real-valued masks due to their ability to modify the phase. Recent work proposed to use a complex filter instead of a point-wise multiplication with a mask. This allows to incorporate information from previous and future time steps exploiting local correlations within each frequency band.In this work, we propose DeepFilterNet, a two stage speech enhancement framework utilizing deep filtering. First, we enhance the spectral envelope using ERB-scaled gains modeling the human frequency perception. The second stage employs deep filtering to enhance the periodic components of speech. Additionally to taking advantage of perceptual properties of speech, we enforce network sparsity via separable convolutions and extensive grouping in linear and recurrent layers to design a low complexity architecture.We further show that our two stage deep filtering approach outperforms complex masks over a variety of frequency resolutions and latencies and demonstrate convincing performance compared to other state-of-the-art models. Hendrik Schröter, Alberto N. Escalante, Tobias Rosenkranz, Andreas K. Maier |
ICASSP | 2 |
| 2022 | Low Latency Speech Enhancement for Hearing Aids Using Deep FilteringabstractNoise reduction is an important feature supporting hearing aid (HA) users in their daily routines and is thus included in most commercially available devices. Latency requirements of HAs require short processing windows resulting in a poor frequency resolution in the whole processing chain including noise reduction. Previous studies have shown that deep neural network (DNN) based algorithms outperform conventional noise reduction algorithms especially for non-stationary noises. This study explores a DNN based noise reduction method using deep filtering targeted for wideband spectrograms given the employed HA filter bank. That is, we predict complex filter coefficients that are linearly applied to the noisy spectrum. We assess different filter sizes over time and frequency axis, and provide evidence for a superior performance over a complex ratio mask. Furthermore, we introduce a frequency response loss that operates on a per-frequency-band basis to fully utilize the deep filtering concept. We objectively demonstrate on-par performance with related state-of-the-art deep learning methods and show in a subjective user study that our method is perceptually preferred to existing HA noise reduction algorithms. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | LACOPE: Latency-Constrained Pitch Estimation for Speech Enhancement
Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Andreas K. Maier |
Interspeech | 3 |
| 2021 | Fricative Phoneme Detection Using Deep Neural Networks and its Comparison to Traditional MethodsabstractS.3171-3175 Metehan Yurt, Pavan Kantharaju, Sascha Disch, Andreas Niedermeier, Alberto N. Escalante, Veniamin I. Morgenshtern |
Interspeech | 5 |
| 2020 | CLCNET: Deep Learning-Based Noise Reduction for Hearing aids using Complex Linear CodingabstractNoise reduction is an important part of modern hearing aids and is included in most commercially available devices. Deep learning-based state-of-the-art algorithms, however, either do not consider real-time and frequency resolution constrains or result in poor quality under very noisy conditions.To improve monaural speech enhancement in noisy environments, we propose CLCNet, a framework based on complex valued linear coding. First, we define complex linear coding (CLC) motivated by linear predictive coding (LPC) that is applied in the complex frequency domain. Second, we propose a framework that incorporates complex spectrogram input and coefficient output. Third, we define a parametric normalization for complex valued spectrograms that complies with low-latency and on-line processing.Our CLCNet was evaluated on a mixture of the EUROM database and a real-world noise dataset recorded with hearing aids and compared to traditional real-valued Wiener-Filter gains. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Marc Aubreville, Andreas K. Maier |
ICASSP | 3 |
| 2020 | Lightweight Online Noise Reduction on Embedded Devices Using Hierarchical Recurrent Neural NetworksabstractDeep-learning based noise reduction algorithms have proven their success especially for non-stationary noises, which makes it desirable to also use them for embedded devices like hearing aids (HAs). This, however, is currently not possible with state-of-the-art methods. They either require a lot of parameters and computational power and thus are only feasible using modern CPUs. Or they are not suitable for online processing, which requires constraints like low-latency by the filter bank and the algorithm itself. In this work, we propose a mask-based noise reduction approach. Using hierarchical recurrent neural networks, we are able to drastically reduce the number of neurons per layer while including temporal context via hierarchical connections. This allows us to optimize our model towards a minimum number of parameters and floating-point operations (FLOPs), while preserving noise reduction quality compared to previous work. Our smallest network contains only 5k parameters, which makes this algorithm applicable on embedded devices. We evaluate our model on a mixture of EUROM and a real-world noise database and report objective metrics on unseen noise. Hendrik Schröter, Tobias Rosenkranz, Alberto N. Escalante, Pascal Zobel, Andreas K. Maier |
INTERSPEECH | 3 |
| 2020 | Improved graph-based SFA: information preservation complements the slowness principleabstractAbstract Slow feature analysis (SFA) is an unsupervised learning algorithm that extracts slowly varying features from a multi-dimensional time series. SFA has been extended to supervised learning (classification and regression) by an algorithm called graph-based SFA (GSFA). GSFA relies on a particular graph structure to extract features that preserve label similarities. Processing of high dimensional input data (e.g., images) is feasible via hierarchical GSFA (HGSFA), resulting in a multi-layer neural network. Although HGSFA has useful properties, in this work we identify a shortcoming, namely, that HGSFA networks prematurely discard quickly varying but useful features before they reach higher layers, resulting in suboptimal global slowness and an under-exploited feature space. To counteract this shortcoming, which we call unnecessary information loss, we propose an extension called hierarchical information-preserving GSFA (HiGSFA), where some features fulfill a slowness objective and other features fulfill an information preservation objective. The efficacy of the extension is verified in three experiments: (1) an unsupervised setup where the input data is the visual stimuli of a simulated rat, (2) the localization of faces in image patches, and (3) the estimation of human age from facial photographs of the MORPH-II database. Both HiGSFA and HGSFA can learn multiple labels and offer a rich feature space, feed-forward training, and linear complexity in the number of samples and dimensions. However, the proposed algorithm, HiGSFA, outperforms HGSFA in terms of feature slowness, estimation accuracy, and input reconstruction, giving rise to a promising hierarchical supervised-learning approach. Moreover, for age estimation, HiGSFA achieves a mean absolute error of 3.41 years, which is a competitive performance for this challenging problem. Alberto N. Escalante, Laurenz Wiskott |
Mach. Learn. | 1 |
| 2019 | Measuring the Data Efficiency of Deep Learning MethodsabstractIn this paper, we propose a new experimental protocol and use it to benchmark the data efficiency --- performance as a function of training set size --- of two deep learning algorithms, convolutional neural networks (CNNs) and hierarchical information-preserving graph-based slow feature analysis (HiGSFA), for tasks in classification and transfer learning scenarios. The algorithms are trained on different-sized subsets of the MNIST and Omniglot data sets. HiGSFA outperforms standard CNN networks when the models are trained on 50 and 200 samples per class for MNIST classification. In other cases, the CNNs perform better. The results suggest that there are cases where greedy, locally optimal bottom-up learning is equally or more powerful than global gradient-based learning. Hlynur Davíð Hlynsson, Alberto N. Escalante, Laurenz Wiskott |
ICPRAM | 2 |
| 2016 | Theoretical Analysis of the Optimal Free Responses of Graph-Based SFA for the Design of Training GraphsabstractSlow feature analysis (SFA) is an unsupervised learning algorithm that extracts slowly varying features from a multi- dimensional time series. Graph-based SFA (GSFA) is an extension to SFA for supervised learning that can be used to successfully solve regression problems if combined with a simple supervised post-processing step on a small number of slow features. The objective function of GSFA minimizes the squared output differences between pairs of samples specified by the edges of a structure called training graph. The edges of current training graphs, however, are derived only from the relative order of the labels. Exploiting the exact numerical value of the labels enables further improvements in label estimation accuracy. In this article, we propose the exact label learning (ELL) method to create a more precise training graph that encodes the desired labels explicitly and allows GSFA to extract a normalized version of them directly (i.e., without supervised post- processing). The ELL method is used for three tasks: (1) We estimate gender from artificial images of human faces (regression) and show the advantage of coding additional labels, particularly skin color. (2) We analyze two existing graphs for regression. (3) We extract compact discriminative features to classify traffic sign images. When the number of output features is limited, such compact features provide a higher classification rate compared to a graph that generates features equivalent to those of nonlinear Fisher discriminant analysis. The method is versatile, directly supports multiple labels, and provides higher accuracy compared to current graphs for the problems considered. Alberto N. Escalante, Laurenz Wiskott |
J. Mach. Learn. Res. | 1 |
| 2013 | How to solve classification and regression problems on high-dimensional data with a supervised extension of slow feature analysis
Alberto N. Escalante, Laurenz Wiskott |
J. Mach. Learn. Res. | 1 |
| 2010 | Gender and Age Estimation from Synthetic Face Images
Alberto N. Escalante, Laurenz Wiskott |
IPMU | 1 |
| 2008 | Secure Multi-Coupons for Federated Environments: Privacy-Preserving and Customer-Friendly
Frederik Armknecht, Alberto N. Escalante, Hans Löhr, Mark Manulis, Ahmad-Reza Sadeghi |
ISPEC | 2 |