VLDB 2026 Research / reviewers in the wild / expert
Adriana Fernandez-Lopez
dblp:199/1754
· DBLP profile ↗
9ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0001-8839-6537ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition ModelsabstractThis paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viability of this training paradigm for such models, yielding several notable findings. Firstly, we discover that applying a low-rank structure exclusively to the attention modules can unexpectedly enhance performance, even with a significant rank reduction of 12%. In contrast, feed-forward layers present greater challenges, as they begin to exhibit performance degradation with a moderate 50% rank reduction. Furthermore, we find that both initialization and layer-wise rank assignment play critical roles in successful low-rank training. Specifically, employing SVD initialization and linear layer-wise rank mapping significantly boosts the efficacy of low-rank weight training. Building on these insights, we introduce the Low-Rank Speech Model from Scratch (LR-SMS), an approach that achieves performance parity with full-rank training while delivering substantial reductions in parameters count (by at least 2×), and training time speedups (by 1.3× for ASR and 1.15× for AVSR). Adriana Fernandez-Lopez, Shiwei Liu 0003, Lu Yin 0006, Stavros Petridis, Maja Pantic |
ICASSP | 1 |
| 2024 | MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma 0001, Lu Yin 0006, Qiao Xiao, Stavros Petridis, Shiwei Liu 0003, Maja Pantic |
INTERSPEECH | 1 |
| 2024 | Dynamic Data Pruning for Automatic Speech RecognitionabstractThe recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss. Qiao Xiao, Pingchuan Ma 0001, Adriana Fernandez-Lopez, Boqian Wu, Lu Yin 0006, Stavros Petridis, Mykola Pechenizkiy, Maja Pantic, Decebal Constantin Mocanu, Shiwei Liu 0003 |
INTERSPEECH | 3 |
| 2023 | Auto-AVSR: Audio-Visual Speech Recognition with Automatic LabelsabstractAudio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger models and training sets. However, accurate labelling of datasets is time-consuming and expensive. Hence, in this work, we investigate the use of automatically-generated transcriptions of unlabelled datasets to increase the training set size. For this purpose, we use publicly-available pre-trained ASR models to automatically transcribe unlabelled datasets such as AVSpeech and Vox-Celeb2. Then, we train ASR, VSR and AV-ASR models on the augmented training set, which consists of the LRS2 and LRS3 datasets as well as the additional automatically-transcribed data. We demonstrate that increasing the size of the training set, a recent trend in the literature, leads to reduced WER despite using noisy transcriptions. The proposed model achieves new state-of-the-art performance on AV-ASR on LRS2 and LRS3. In particular, it achieves a WER of 0.9 % on LRS3, a relative improvement of 30 % over the current state-of-the–art approach, and outperforms methods that have been trained on non-publicly available datasets with 26 times more training data. Pingchuan Ma 0001, Alexandros Haliassos, Adriana Fernandez-Lopez, Honglie Chen, Stavros Petridis, Maja Pantic |
ICASSP | 3 |
| 2023 | SparseVSR: Lightweight and Noise Robust Visual Speech Recognition
Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma 0001, Alexandros Haliassos, Stavros Petridis, Maja Pantic |
INTERSPEECH | 1 |
| 2022 | End-to-End Lip-Reading Without Large-Scale DataabstractThe development of Automatic Lip-Reading (ALR) systems for continuous speech recognition has so far limited their applicability to English since this is the only language with large-scale datasets sufficient to train end-to-end ALR systems. In this work, we show that it is possible to train competitive end-to-end ALR systems in alternative languages with challenging small-scale data as long as the appropriate restrictions are made to the learning process of the visual front-end objective. To this end, we hypothesize that the visual front-end should be trained in a self-supervised setting, allowing it to target its ownvisual units. We specifically definevisual unitsas a collection of visually similar images constrained by linguistics and provide an algorithmic implementation to automatically generate them. We show thatvisual unitscan be used to add an intermediate classification task between the visual and temporal modules that facilitates meaningful learning of visual features and, as a consequence, reduces the amount of data required to train an end-to-end ALR system. Additionally, we present a data augmentation strategy for enriching the temporal context. We synthesize realistic video sequences by appropriately combining characters-like sub-sequences from existing videos. We test the proposed ALR system on i) the VLRF dataset, a small-scale database that is one of the largest in Spanish, and achieve 44.77$\%$CER and 72.90% WER, which are competitive with the state-of-the-art and significant for this volume of training material; ii) the TCD-TIMIT dataset, a comparable medium-scale database in English, where we achieve 36.58% CER and 56.29% WER, which are also state-of-the-art results on speaker-dependent experiments. Adriana Fernandez-Lopez, Federico Sukno |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Cogans For Unsupervised Visual Speech Adaptation To New SpeakersabstractAudio-Visual Speech Recognition (AVSR) faces the difficult task of exploiting acoustic and visual cues simultaneously. Augmenting speech with the visual channel creates its own challenges, e.g. every person has unique mouth movements, making the generalization of visual models very difficult. This factor motivates our focus on the generalization of speaker-independent (SI) AVSR systems especially in noisy environments by exploiting the visual domain. Specifically, we are the first to explore the visual adaptation of an SI-AVSR system to an unknown and unlabelled speaker. We adapt an AVSR system trained in a source domain to decode samples in a target domain without the need for labels in the target domain. For the domain adaptation of the unknown speaker, we use Coupled Generative Adversarial Networks to automatically learn a joint distribution of multi-domain images. We evaluate our character-based AVSR system on the TCD-TIMIT dataset and obtain up to a 10% average improvement with respect to its AVSR system equivalent. Adriana Fernandez-Lopez, Ali Karaali, Naomi Harte, Federico Sukno |
ICASSP | 1 |
| 2018 | Survey on automatic lip-reading in the era of deep learning
Adriana Fernandez-Lopez, Federico Sukno |
Image Vis. Comput. | 1 |
| 2017 | Towards Estimating the Upper Bound of Visual-Speech Recognition: The Visual Lip-Reading Feasibility DatabaseabstractSpeech is the most used communication method between humans and it involves the perception of auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, although the video can provide information that is complementary to the audio. Exploiting the visual information, however, has proven challenging. On one hand, researchers have reported that the mapping between phonemes and visemes (visual units) is one-to-many because there are phonemes which are visually similar and indistinguishable between them. On the other hand, it is known that some people are very good lip-readers (e.g: deaf people). We study the limit of visual only speech recognition in controlled conditions. With this goal, we designed a new database in which the speakers are aware of being read and aim to facilitate lip-reading. In the literature, there are discrepancies on whether hearing-impaired people are better lip-readers than normal-hearing people. Then, we analyze if there are differences between the lip-reading abilities of 9 hearing-impaired and 15 normal-hearing people. Finally, human abilities are compared with the performance of a visual automatic speech recognition system. In our tests, hearing-impaired participants outperformed the normal-hearing participants but without reaching statistical significance. Human observers were able to decode 44% of the spoken message. In contrast, the visual only automatic system achieved 20% of word recognition rate. However, if we repeat the comparison in terms of phonemes both obtained very similar recognition rates, just above 50%. This suggests that the gap between human lip-reading and automatic speech-reading might be more related to the use of context than to the ability to interpret mouth appearance. Adriana Fernandez-Lopez, Oriol Martínez, Federico Sukno |
FG | 1 |