VLDB 2026 Research / reviewers in the wild / expert
Leonardo Pepino
dblp:267/2291
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-5037-3700ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmarking Time-localized Explanations for Audio Classification ModelsabstractMost modern approaches for audio processing are opaque, in the sense that they do not provide an explanation for their decisions. For this reason, various methods have been proposed to explain the outputs generated by these models. Good explanations can result in interesting insights about the data or the model, as well as increase trust in the system. Unfortunately, evaluating the quality of explanations is far from trivial since, for most tasks, there is no clear ground truth explanation to use as reference. In this work, we propose a benchmark for time-localized explanations for audio classification models that uses time annotations of target events as a proxy for ground truth explanations. We use this benchmark to systematically optimize and compare various approaches for model-agnostic post-hoc explanation, obtaining, in some cases, close to perfect explanations. Finally, we illustrate the utility of the explanations for uncovering spurious correlations. Cecilia Bolaños, Leonardo Pepino, Martín Meza, Luciana Ferrer |
INTERSPEECH | 2 |
| 2025 | EnCodecMAE: leveraging neural codecs for universal audio representation learning
Leonardo Pepino, Pablo Riera, Luciana Ferrer |
INTERSPEECH | 1 |
| 2025 | A Dataset for Automatic Assessment of TTS Quality in Spanish
Alejandro Sosa Welford, Leonardo Pepino |
INTERSPEECH | 2 |
| 2023 | Towards detecting the level of trust in the skills of a virtual assistant from the user's speech
Lara Gauder, Leonardo Pepino, Pablo Riera, Silvina Brussino, Jazmín Vidal, Agustín Gravano, Luciana Ferrer |
Comput. Speech Lang. | 2 |
| 2022 | Study of Positional Encoding Approaches for Audio Spectrogram TransformersabstractTransformers have revolutionized the world of deep learning, specially in the field of natural language processing. Recently, the Audio Spectrogram Transformer (AST) was proposed for audio classification, leading to state of the art results in several datasets. However, in order for ASTs to outperform CNNs, pretraining with ImageNet is needed. In this paper, we study one component of the AST, the positional encoding, and propose several variants to improve the performance of ASTs trained from scratch, without ImageNet pretraining. Our best model, which incorporates conditional positional encodings, significantly improves performance on Audioset and ESC-50 compared to the original AST. Leonardo Pepino, Pablo Riera, Luciana Ferrer |
ICASSP | 1 |
| 2021 | Alzheimer Disease Recognition Using Speech-Based Embeddings From Pre-Trained Models
Lara Gauder, Leonardo Pepino, Luciana Ferrer, Pablo Riera |
Interspeech | 2 |
| 2021 | Emotion Recognition from Speech Using wav2vec 2.0 EmbeddingsabstractEmotion recognition datasets are relatively small, making the use of the more sophisticated deep learning approaches challenging. In this work, we propose a transfer learning method for speech emotion recognition where features extracted from pre-trained wav2vec 2.0 models are modeled using simple neural networks. We propose to combine the output of several layers from the pre-trained model using trainable weights which are learned jointly with the downstream model. Further, we compare performance using two different wav2vec 2.0 models, with and without finetuning for speech recognition. We evaluate our proposed approaches on two standard emotion databases IEMOCAP and RAVDESS, showing superior performance compared to results in the literature. Leonardo Pepino, Pablo Riera, Luciana Ferrer |
Interspeech | 1 |
| 2020 | Fusion Approaches for Emotion Recognition from Speech Using Acoustic and Text-Based FeaturesabstractIn this paper, we study different approaches for classifying emotions from speech using acoustic and text-based features. We propose to obtain contextualized word embeddings with BERT to represent the information contained in speech transcriptions and show that this results in better performance than using Glove embeddings. We also propose and compare different strategies to combine the audio and text modalities, evaluating them on IEMOCAP and MSPPODCAST datasets. We find that fusing acoustic and text-based systems is beneficial on both datasets, though only subtle differences are observed across the evaluated fusion approaches. Finally, for IEMOCAP, we show the large effect that the criteria used to define the cross-validation folds have on results. In particular, the standard way of creating folds for this dataset results in a highly optimistic estimation of performance for the text-based system, suggesting that some previous works may overestimate the advantage of incorporating transcriptions. Leonardo Pepino, Pablo Riera, Luciana Ferrer, Agustín Gravano |
ICASSP | 1 |