VLDB 2026 Research / reviewers in the wild / expert
Alex Peiró Lilja
dblp:173/6405
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LaFresCat: A studio-quality Catalan multi-accent speech dataset for text-to-speech synthesisabstractCurrent text-to-speech (TTS) systems are capable of learning the phonetics of a language accurately given that the speech data used to train such models covers all phonetic phenomena. For languages with different varieties, this includes all their richness and accents. This is the case of Catalan, a mid-resourced language with several dialects or accents. Although there are various publicly available corpora, there is a lack of high-quality open-access data for speech technologies covering its variety of accents. Common Voice includes recordings of Catalan speakers from different regions; however, accent labeling has been shown to be inaccurate, and artificially enhanced samples may be unsuitable for TTS. To address these limitations, we present LaFresCat, the first studio-quality Catalan multi-accent dataset. LaFresCat comprises 3.5 h of professionally recording speech covering four of the most prominent Catalan accents: Balearic, Central, North-Western, and Valencian. In this work, we provide a detailed description of the dataset design: utterances were selected to be phonetically balanced, detailed speaker instructions were provided, native speakers from the regions corresponding to the Catalan accents were hired, and the recordings were formatted and post-processed. The resulting dataset, LaFresCat, is publicly available. To preliminarily evaluate the dataset, we trained and assessed a lightweight flow-based TTS system, which is also provided as a by-product. We also analyzed LaFresCat samples and the corresponding TTS-generated samples at the phonetic level, employing expert annotations and Pillai scores to quantify acoustic vowel overlap. Preliminary results suggest a significant improvement in predicted mean opinion score (UTMOS), with an increase of 0.42 points when the TTS system is fine-tuned on LaFresCat rather than trained from scratch, starting from a pre-trained version based on Central Catalan data from Common Voice. Subsequent human expert annotations achieved nearly 90% accuracy in accent classification for LaFresCat recordings. However, although the TTS tends to homogenize pronunciation, it still learns distinct dialectal patterns. This assessment offers key insights for establishing a baseline to guide future evaluations of Catalan multi-accent TTS systems and further studies of LaFresCat. Alex Peiró Lilja, Carme Armentano-Oller, José Giraldo, Wendy Elvira-García, Ignasi Esquerra, Rodolfo Zevallos, Cristina España-Bonet, Martí Llopart-Font, Baybars Külebi, Mireia Farrús |
Comput. Speech Lang. | 1 |
| 2025 | Evaluating Speech Enhancement Performance Across Demographics and Language
José Giraldo, Alex Peiró Lilja, Carme Armentano-Oller, Rodolfo Zevallos, Cristina España-Bonet |
INTERSPEECH | 2 |
| 2025 | Towards Domain-Specific Spoken Language Understanding for a Catalan Voice-Controlled Video Game
Alex Peiró Lilja, Rodolfo Zevallos, Carme Armentano-Oller, José Giraldo, Cristina España-Bonet, Mireia Farrús |
INTERSPEECH | 1 |
| 2025 | Assessing the Performance and Efficiency of Mamba ASR in Low-Resource ScenariosabstractMamba, a state space model-based architecture, is emerging as a strong alternative to Transformer models, showing equal or superior performance in sequence generation, including speech. However, analyses have focused mainly on highresource scenarios. This paper explores Mamba’s potential in ASR for low-resource scenarios. We compare the Transformerbased Conformer and its state-space counterpart, ConMamba, across nine languages with varying training data. Our results show that ConMamba achieves similar WER to Conformer for short-context inputs but significantly improves performance on long-context inputs, reducing WER by up to 50% on average. Additionally, ConMamba enhances efficiency, requiring 40–45% less training time, using 50% less memory, and accelerating inference by 63–70%, making it a more effective ASR solution across different data availability scenarios. Rodolfo Zevallos, Martí Cortada Garcia, Sarah Solito, Carlos Mena, Alex Peiró Lilja, Javier Hernando |
INTERSPEECH | 5 |
| 2024 | Multi-speaker and multi-dialectal Catalan TTS models for video gaming
Alex Peiró Lilja, José Giraldo, Martí Llopart-Font, Carme Armentano-Oller, Baybars Külebi, Mireia Farrús |
INTERSPEECH | 1 |
| 2021 | The INGENIOUS Multilingual Operations App
Joan Codina, Guillermo Cámbara, Alex Peiró Lilja, Jens Grivolla, Roberto Carlini, Mireia Farrús |
Interspeech | 3 |
| 2020 | CATOTRON - A Neural Text-to-Speech System in Catalan
Baybars Külebi, Alp Öktem, Alex Peiró Lilja, Santiago Pascual, Mireia Farrús |
INTERSPEECH | 3 |
| 2020 | Naturalness Enhancement with Linguistic Information in End-to-End TTS Using Unsupervised Parallel EncodingabstractComunicació presentada a Interspeech 2020 celebrat del 25 al 29 d'octubre de 2020 a Shanghai, Xina. Alex Peiró Lilja, Mireia Farrús |
INTERSPEECH | 1 |
| 2015 | Automatic detection of equipment alarms in a neonatal intensive care unit environment: a knowledge-based approachabstractAlarm sounds triggered by biomedical equipment play a key role in providing healthcare in a neonatal intensive care unit (NICU). This paper presents our work on automatic detection of acoustic alarms in a noisy NICU environment, where knowledge about the particular characteristics of each alarm class is integrated at different stages of the detection system. The feature extraction is based on applying, around alarm-specific frequencies, a method for detection of sinusoidal signals, which employs the normalised short-term magnitude and phase spectrum. Also, the ratios of magnitudes at those frequencies are\ntaken as features. The system consists of a set of GMM-based detectors, each designed to deal with a specific alarm. Temporal structure of alarms, in terms of duration of signal and silence intervals in every alarm period, is incorporated by aggregating the frame-level posterior probabilities. The experimental evaluations are performed with a database recorded in a real-world hospital environment. The performance of the detection system is assessed both at the frame level and at the alarm period level. Ganna Raboshchuk, Peter Jancovic, Climent Nadeu, Alex Peiró Lilja, Münevver Köküer, Blanca Muñoz Mahamud, Ana Riverola de Veciana |
INTERSPEECH | 4 |