EDBT 2026 Demo / reviewers in the wild / expert
Thomas Rolland
dblp:161/4165
· DBLP profile ↗
13ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-6468-8391ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
Francisco Teixeira, Carlos Carvalho 0003, Mariana Julião, Catarina Botelho, Rubén Solera-Ureña, Sérgio Paulo, Thomas Rolland, Ben Peters, Isabel Trancoso, Alberto Abad |
LREC | 7 |
| 2025 | CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European PortugueseabstractExisting resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties underexplored. To bridge this gap, we introduce CAMÕES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46 h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of $\mathbf{4 2 5 h}$ of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties. Carlos Carvalho 0003, Francisco Teixeira, Catarina Botelho, Anna Pompili, Rubén Solera-Ureña, Sérgio Paulo, Mariana Julião, Thomas Rolland, John Mendonça, Diogo A. P. Nunes, Isabel Trancoso, Alberto Abad |
ASRU | 8 |
| 2025 | Acoustic and Linguistic Biomarkers for Cognitive Impairment Detection from Speech
Catarina Botelho, David Gimeno-Gómez, Francisco Teixeira, John Mendonça, Patrícia Pereira, Diogo A. P. Nunes, Thomas Rolland, Anna Pompili, Rubén Solera-Ureña, Maria Ponte, David Martins de Matos, Carlos D. Martínez-Hinarejos, Isabel Trancoso, Alberto Abad |
INTERSPEECH | 7 |
| 2025 | Children's Voice Privacy: First Steps and Emerging ChallengesabstractInternational audience Ajinkya Kulkarni, Francisco Teixeira, Enno Hermann, Thomas Rolland, Isabel Trancoso, Mathew Magimai-Doss |
INTERSPEECH | 4 |
| 2025 | Exploring Shared-Weight Mechanisms in Transformer and Conformer Architectures for Automatic Speech Recognition
Thomas Rolland, Alberto Abad |
INTERSPEECH | 1 |
| 2024 | Exploring Adapters with Conformers for Children's Automatic Speech RecognitionabstractThe high variability in acoustic, pronunciation, and linguistic characteristics of children’s speech makes of children’s automatic speech recognition (ASR) a complex task. Training a dedicated ASR model from scratch for children remains challenging, mainly due to the limited availability of children’s data. To tackle this limitation, a common strategy involves fine-tuning a pre-trained ASR model. However, this approach faces challenges due to the diversity of speakers and data scarcity, especially when dealing with large ASR models like the Conformer. In this study, we explore an alternative approach known as Adapter transfer. Adapter transfer requires training fewer parameters and can be more effective in adapting large ASR models for children’s speech. In this paper, we assess various Adapter configurations in the literature and introduce a novel configuration called Two Serial Adapter (TSA). The experimental results indicate that Adapter transfer consistently outperforms traditional fine-tuning across various configurations for the Conformer model. Thomas Rolland, Alberto Abad |
ICASSP | 1 |
| 2024 | Improved Children's Automatic Speech Recognition Combining Adapters and Synthetic Data AugmentationabstractChildren’s automatic speech recognition (ASR) poses a significant challenge due to the high variability nature of children’s speech. The limited availability of training datasets hampers the effective modelling of this variability, which can be partially addressed using a text-to-speech (TTS) system for data augmentation. However, generated data may contain imperfections, potentially impacting performance. In this work, we use Adapters to handle the domain mismatch when fine-tuning with TTS data. This involves a two-step training process: training adapter layers with a frozen pre-trained model using synthetic data, then fine-tuning both adapters and the entire model with a mix of synthetic and real data, where only synthetic data passes through the adapters. Experimental results demonstrate up to 6% relative reduction in WER compared to the straightforward use of synthetic data, indicating the effectiveness of adapter-based architectures in learning from imperfect synthetic data. Thomas Rolland, Alberto Abad |
ICASSP | 1 |
| 2024 | Shared-Adapters: A Novel Transformer-based Parameter Efficient Transfer Learning Approach For Children's Automatic Speech Recognition
Thomas Rolland, Alberto Abad |
INTERSPEECH | 1 |
| 2024 | Introduction To Partial Fine-tuning: A Comprehensive Evaluation Of End-to-end Children's Automatic Speech Recognition AdaptationabstractAutomatic Speech Recognition (ASR) encounters unique challenges when dealing with children's speech, mainly due to the scarcity of available data.Training large ASR models with constrained data presents a significant challenge.To address this, fine-tuning strategy is frequently employed.However, fine-tuning an entire large pre-trained model with limited children's speech data may overfit leading to decreased performance.This study offers a granular evaluation of children's ASR fine-tuning, departing from conventional whole-network tunning.We present a partial fine-tuning approach spotlighting the importance of the Encoder and Feedforward Neural Network modules in Transformer-based models.Remarkably, this method surpasses the efficacy of whole-model fine-tuning, with a relative word error rate improvement of 9% when dealing with limited data.Our findings highlight the critical role of partial fine-tuning in advancing children's ASR model development. Thomas Rolland, Alberto Abad |
INTERSPEECH | 1 |
| 2022 | Multilingual Transfer Learning for Children Automatic Speech RecognitionabstractDespite recent advances in automatic speech recognition (ASR), the recognition of children’s speech still remains a significant challenge. This is mainly due to the high acoustic variability and the limited amount of available training data. The latter problem is particularly evident in languages other than English, which are usually less-resourced. In the current paper, we address children ASR in a number of less-resourced languages by combining several small-sized children speech corpora from these languages. In particular, we address the following research question: Does a novel two-step training strategy in which multilingual learning is followed by language-specific transfer learning outperform conventional single language/task training for children speech, as well as multilingual and transfer learning alone? Based on previous experimental results with English, we hypothesize that multilingual learning provides a better generalization of the underlying characteristics of children’s speech. Our results provide a positive answer to our research question, by showing that using transfer learning on top of a multilingual model for an unseen language outperforms conventional single language-specific learning. Thomas Rolland, Alberto Abad, Catia Cucchiarini, Helmer Strik |
LREC | 1 |
| 2021 | Transfer Learning-Based Cough Representations for Automatic Detection of COVID-19abstractIn the last months, there has been an increasing interest in developing reliable, cost-effective, immediate and easy to use machine learning based tools that can help health care operators, institutions, companies, etc. to optimize their screening campaigns.In this line, several initiatives emerged aimed at the automatic detection of COVID-19 from speech, breathing and coughs, with inconclusive preliminary results.The ComParE 2021 COVID-19 Cough Sub-challenge provides researchers from all over the world a suitable test-bed for the evaluation and comparison of their work.In this paper, we present the INESC-ID contribution to the ComParE 2021 COVID-19 Cough Sub-challenge.We leverage transfer learning to develop a set of three expert classifiers based on deep cough representation extractors.A calibrated decision-level fusion system provides the final classification of coughs recordings as either COVID-19 positive or negative.Results show unweighted average recalls of 72.3% and 69.3% in the development and test sets, respectively.Overall, the experimental assessment shows the potential of this approach although much more research on extended respiratory sounds datasets is needed. Rubén Solera-Ureña, Catarina Botelho, Francisco Teixeira, Thomas Rolland, Alberto Abad, Isabel Trancoso |
Interspeech | 4 |
| 2020 | The INESC-ID Multi-Modal System for the ADReSS 2020 ChallengeabstractThis paper describes a multi-modal approach for the automatic detection of Alzheimer's disease proposed in the context of the INESC-ID Human Language Technology Laboratory participation in the ADReSS 2020 challenge.Our classification framework takes advantage of both acoustic and textual feature embeddings, which are extracted independently and later combined.Speech signals are encoded into acoustic features using DNN speaker embeddings extracted from pre-trained models.For textual input, contextual embedding vectors are first extracted using an English Bert model and then used either to directly compute sentence embeddings or to feed a bidirectional LSTM-RNNs with attention.Finally, an SVM classifier with linear kernel is used for the individual evaluation of the three systems.Our best system, based on the combination of linguistic and acoustic information, attained a classification accuracy of 81.25%.Results have shown the importance of linguistic features in the classification of Alzheimer's Disease, which outperforms the acoustic ones in terms of accuracy.Early stage features fusion did not provide additional improvements, confirming that the discriminant ability conveyed by speech in this case is smooth out by linguistic data. Anna Pompili, Thomas Rolland, Alberto Abad |
INTERSPEECH | 2 |
| 2015 | Protein Domain-Level Landscape of Cancer-Type-Specific Somatic MutationsabstractIdentifying driver mutations and their functional consequences is critical to our understanding of cancer. Towards this goal, and because domains are the functional units of a protein, we explored the protein domain-level landscape of cancer-type-specific somatic mutations. Specifically, we systematically examined tumor genomes from 21 cancer types to identify domains with high mutational density in specific tissues, the positions of mutational hotspots within these domains, and the functional and structural context where possible. While hotspots corresponding to specific gain-of-function mutations are expected for oncoproteins, we found that tumor suppressor proteins also exhibit strong biases toward being mutated in particular domains. Within domains, however, we observed the expected patterns of mutation, with recurrently mutated positions for oncogenes and evenly distributed mutations for tumor suppressors. For example, we identified both known and new endometrial cancer hotspots in the tyrosine kinase domain of the FGFR2 protein, one of which is also a hotspot in breast cancer, and found new two hotspots in the Immunoglobulin I-set domain in colon cancer. Thus, to prioritize cancer mutations for further functional studies aimed at more precise cancer treatments, we have systematically correlated mutations and cancer types at the protein domain level. Evangelia Petsalaki, Thomas Rolland, David E. Hill, Marc Vidal, Frederick P. Roth |
PLoS Comput. Biol. | 3 |