EDBT 2026 Demo / reviewers in the wild / expert
Mireia Farrús
dblp:70/1200 · also Mireia Farrús Cabeceran
· DBLP profile ↗
31ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-7160-9513ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 7 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
Pol Pastells, Mireia Farrús, Mariona Taulé |
LREC | 3 |
| 2026 | LaFresCat: A studio-quality Catalan multi-accent speech dataset for text-to-speech synthesisabstractCurrent text-to-speech (TTS) systems are capable of learning the phonetics of a language accurately given that the speech data used to train such models covers all phonetic phenomena. For languages with different varieties, this includes all their richness and accents. This is the case of Catalan, a mid-resourced language with several dialects or accents. Although there are various publicly available corpora, there is a lack of high-quality open-access data for speech technologies covering its variety of accents. Common Voice includes recordings of Catalan speakers from different regions; however, accent labeling has been shown to be inaccurate, and artificially enhanced samples may be unsuitable for TTS. To address these limitations, we present LaFresCat, the first studio-quality Catalan multi-accent dataset. LaFresCat comprises 3.5 h of professionally recording speech covering four of the most prominent Catalan accents: Balearic, Central, North-Western, and Valencian. In this work, we provide a detailed description of the dataset design: utterances were selected to be phonetically balanced, detailed speaker instructions were provided, native speakers from the regions corresponding to the Catalan accents were hired, and the recordings were formatted and post-processed. The resulting dataset, LaFresCat, is publicly available. To preliminarily evaluate the dataset, we trained and assessed a lightweight flow-based TTS system, which is also provided as a by-product. We also analyzed LaFresCat samples and the corresponding TTS-generated samples at the phonetic level, employing expert annotations and Pillai scores to quantify acoustic vowel overlap. Preliminary results suggest a significant improvement in predicted mean opinion score (UTMOS), with an increase of 0.42 points when the TTS system is fine-tuned on LaFresCat rather than trained from scratch, starting from a pre-trained version based on Central Catalan data from Common Voice. Subsequent human expert annotations achieved nearly 90% accuracy in accent classification for LaFresCat recordings. However, although the TTS tends to homogenize pronunciation, it still learns distinct dialectal patterns. This assessment offers key insights for establishing a baseline to guide future evaluations of Catalan multi-accent TTS systems and further studies of LaFresCat. Alex Peiró Lilja, Carme Armentano-Oller, José Giraldo, Wendy Elvira-García, Ignasi Esquerra, Rodolfo Zevallos, Cristina España-Bonet, Martí Llopart-Font, Baybars Külebi, Mireia Farrús |
Comput. Speech Lang. | 10 |
| 2025 | The Potential of Speech Features to Discriminate between Original and Machine-Translated TextsabstractDiscriminating between original texts and machine translations involves identifying whether a text was originally authored in the target language or generated through machine translation. To our knowledge, all methods to date depend exclusively on text-based features. In this study, we move beyond this unimodal approach by incorporating speech features. Machine-translated texts display linguistic deviations from original texts, such as those in lexicon and syntax, which can also manifest in speech characteristics. We evaluate the effectiveness of using text features, speech features, and their bimodal fusion to train classifiers capable of discerning original from machine-translated texts. Additionally, we explore various classification algorithms and fusion techniques. Our results show that speech features alone surpass chance accuracy, while combining text and speech features enhances performance beyond text-only methods. Furthermore, although no single classification or fusion method proves consistently superior, advanced fusion techniques outperform simple feature concatenation. Yongjian Chen, Mireia Farrús, Antonio Toral |
ICASSP | 2 |
| 2025 | Towards Domain-Specific Spoken Language Understanding for a Catalan Voice-Controlled Video Game
Alex Peiró Lilja, Rodolfo Zevallos, Carme Armentano-Oller, José Giraldo, Cristina España-Bonet, Mireia Farrús |
INTERSPEECH | 6 |
| 2025 | SCRIBAL: A Digital Transcription Tool in Higher Education
Javier Román, Pol Pastells, Mauro Vázquez Chas, Clara Puigventós, Montserrat Nofre, Mariona Taulé, Mireia Farrús |
INTERSPEECH | 7 |
| 2024 | Improving NMT from a Low-Resource Source Language: A Use Case from Catalan to Chinese via SpanishabstractThe effectiveness of neural machine translation is markedly constrained in low-resource scenarios, where the scarcity of parallel data hampers the development of robust models. This paper focuses on the scenario where the source language is low-resourceand there exists a related high-resource language, for which we introduce a novel approach that combines pivot translation and multilingual training. As a use case we tackle the automatic translation from Catalan to Chinese, using Spanish as an additional language. Our evaluation, conducted on the FLORES-200 benchmark, compares our new approach against a vanilla baseline alongside other models representing various low-resource techniques in the Catalan-to-Chinese context. Experimental results highlight the efficacy of our proposed method, which outperforms existing models, notably demonstrating significant improvements both in translation quality and in lexical diversity. Yongjian Chen, Antonio Toral, Mireia Farrús |
EAMT (1) | 4 |
| 2024 | TEMA: Token Embeddings Mapping for Enriching Low-Resource Language ModelsabstractThe objective of the research we present is to remedy the problem of the low quality of language models for low-resource languages.We introduce an algorithm, the Token Embedding Mapping Algorithm (TEMA), that maps the token embeddings of a richly pre-trained model L1 to a poorly trained model L2, thus creating a richer L2' model.Our experiments show that the L2' model reduces perplexity with respect to the original monolingual model L2, and that for downstream tasks, including SuperGLUE, the results are state-of-the-art or better for the most semantic tasks.The models obtained with TEMA are also competitive or better than multilingual or extended models proposed as solutions for mitigating the low-resource language problems. Rodolfo Zevallos, Núria Bel, Mireia Farrús |
EMNLP | 3 |
| 2024 | Multi-speaker and multi-dialectal Catalan TTS models for video gaming
Alex Peiró Lilja, José Giraldo, Martí Llopart-Font, Carme Armentano-Oller, Baybars Külebi, Mireia Farrús |
INTERSPEECH | 6 |
| 2022 | Recycle Your Wav2Vec2 Codebook: A Speech Perceiver for Keyword SpottingabstractSpeech information in a pretrained wav2vec2.0 model is usually leveraged through its encoder, which has at least 95M parameters, being not so suitable for small footprint Keyword Spotting. In this work, we show an efficient way of profiting from wav2vec2.0’s linguistic knowledge, by recycling the phonetic information encoded in its latent codebook, which has been typically thrown away after pretraining. We do so by transferring the codebook as weights for the latent bottleneck of a Keyword Spotting Perceiver, thus initializing such model with phonetic embeddings already. The Perceiver design relies on cross-attention between these embeddings and input data to generate better representations. Our method delivers accuracy gains compared to random initialization, at no latency costs. Plus, we show that the phonetic embeddings can easily be downsampled with k-means clustering, speeding up inference in 3.5 times at only slight accuracy penalties. Guillermo Cámbara, Jordi Luque, Mireia Farrús |
COLING | 3 |
| 2022 | Data Augmentation for Low-Resource Quechua ASR ImprovementabstractComunicació presentada a INTERSPEECH 2022, celebrat del 18 al 22 de setembre de 2022 a Inchon, Corea del Sud. Rodolfo Zevallos, Núria Bel, Guillermo Cámbara, Mireia Farrús, Jordi Luque |
INTERSPEECH | 4 |
| 2021 | The INGENIOUS Multilingual Operations App
Joan Codina, Guillermo Cámbara, Alex Peiró Lilja, Jens Grivolla, Roberto Carlini, Mireia Farrús |
Interspeech | 6 |
| 2021 | Mobile eHealth Platform for Home Monitoring of Bipolar Disorder
Joan Codina, Sergio Escalera, Joan Escudero, Coen Antens, Pau Buch-Cardona, Mireia Farrús |
MMM (2) | 6 |
| 2020 | Detection of Speech Events and Speaker Characteristics through Photo-Plethysmographic Signal Neural ProcessingabstractThe use of photoplethysmogram signal (PPG) for heart and sleep monitoring is commonly found nowadays in smart-phones and wrist wearables. Besides common usages, it has been proposed and reported that person information can be extracted from PPG for other uses, like biometry tasks. In this work, we explore several end-to-end convolutional neural network architectures for detection of human's characteristics such as gender or person identity. In addition, we evaluate whether speech/non-speech events may be inferred from PPG signal, where speech might translate in fluctuations into the pulse signal. The obtained results are promising and clearly show the potential of fully end-to-end topologies for automatic extraction of meaningful biomarkers, even from a noisy signal sampled by a low-cost PPG sensor. The AUCs for best architectures put forward PPG wave as biological discriminant, reaching 79% and 89.0%, respectively for gender and person verification tasks. Furthermore, speech detection experiments reporting AUCs around 69% encourage us for further exploration about the feasibility of PPG for speech processing tasks. Guillermo Cámbara, Jordi Luque, Mireia Farrús |
ICASSP | 3 |
| 2020 | CATOTRON - A Neural Text-to-Speech System in Catalan
Baybars Külebi, Alp Öktem, Alex Peiró Lilja, Santiago Pascual, Mireia Farrús |
INTERSPEECH | 5 |
| 2020 | Naturalness Enhancement with Linguistic Information in End-to-End TTS Using Unsupervised Parallel EncodingabstractComunicació presentada a Interspeech 2020 celebrat del 25 al 29 d'octubre de 2020 a Shanghai, Xina. Alex Peiró Lilja, Mireia Farrús |
INTERSPEECH | 2 |
| 2020 | Integrating lexical and prosodic features for automatic paragraph segmentation
Catherine Lai, Mireia Farrús, Johanna D. Moore |
Speech Commun. | 2 |
| 2019 | Prosodic Phrase Alignment for Machine DubbingabstractDubbing is a type of audiovisual translation where dialogues are translated and enacted so that they give the impression that the media is in the target language. It requires a careful alignment of dubbed recordings with the lip movements of performers in order to achieve visual coherence. In this paper, we deal with the specific problem of prosodic phrase synchronization within the framework of machine dubbing. Our methodology exploits the attention mechanism output in neural machine translation to find plausible phrasing for the translated dialogue lines and then uses them to condition their synthesis. Our initial work in this field records comparable speech rate ratio to professional dubbing translation, and improvement in terms of lip-syncing of long dialogue lines. Alp Öktem, Mireia Farrús, Antonio Bonafonte |
INTERSPEECH | 2 |
| 2018 | Visualizing Punctuation Restoration in Speech Transcripts with Prosograph
Alp Öktem, Mireia Farrús, Antonio Bonafonte |
INTERSPEECH | 2 |
| 2018 | Compilation of Corpora for the Study of the Information Structure-Prosody Interface
Alicia Burga, Mónica Domínguez, Mireia Farrús, Leo Wanner |
LREC | 3 |
| 2017 | A Thematicity-Based Prosody Enrichment Tool for CTS
Mónica Domínguez, Mireia Farrús, Leo Wanner |
INTERSPEECH | 2 |
| 2017 | Using Prosody to Classify Discourse RelationsabstractComunicació presentada a: The 18th Annual Conference of the International Speech Communication Association (INTERSPEECH 2017), celebrada a Estocolm, Suència, del 20 al 24 d'agost de 2017. Janine Kleinhans, Mireia Farrús, Agustín Gravano, Juan Manuel Pérez, Catherine Lai, Leo Wanner |
INTERSPEECH | 2 |
| 2017 | Prosograph: A Tool for Prosody Visualisation of Large Speech Corpora
Alp Öktem, Mireia Farrús, Leo Wanner |
INTERSPEECH | 2 |
| 2016 | An Automatic Prosody Tagger for Spontaneous SpeechabstractSpeech prosody is known to be central in advanced communication technologies. However, despite the advances of theoretical studies in speech prosody, so far, no large scale prosody annotated resources that would facilitate empirical research and the development of empirical computational approaches are available. This is to a large extent due to the fact that current common prosody annotation conventions offer a descriptive framework of intonation contours and phrasing based on labels. This makes it difficult to reach a satisfactory inter-annotator agreement during the annotation of gold standard annotations and, subsequently, to create consistent large scale annotations. To address this problem, we present an annotation schema for prominence and boundary labeling of prosodic phrases based upon acoustic parameters and a tagger for prosody annotation at the prosodic phrase level. Evaluation proves that inter-annotator agreement reaches satisfactory values, from 0.60 to 0.80 Cohen’s kappa, while the prosody tagger achieves acceptable recall and f-measure figures for five spontaneous samples used in the evaluation of monologue and dialogue formats in English and Spanish. The work presented in this paper is a first step towards a semi-automatic acquisition of large corpora for empirical prosodic analysis. Mónica Domínguez, Mireia Farrús, Leo Wanner |
COLING | 2 |
| 2016 | Automatic Paragraph Segmentation with Lexical and Prosodic FeaturesabstractAs long-form spoken documents become more ubiquitous in everyday life, so does the need for automatic discourse segmentation in spoken language processing tasks. Although previous work has focused on broad topic segmentation, detection of finer-grained discourse units, such as paragraphs, is highly desirable for presenting and analyzing spoken content. To better understand how different aspects of speech cue these subtle discourse transitions, we investigate automatic paragraph segmentation of TED talks. We build lexical and prosodic paragraph segmenters using Support Vector Machines, AdaBoost, and Long Short Term Memory (LSTM) recurrent neural networks. In general, we find that induced cue words and supra-sentential prosodic features outperform features based on topical coherence, syntactic form and complexity. However, our best performance is achieved by combining a wide range of individually weak lexical and prosodic features, with the sequence modelling LSTM generally outperforming the other classifiers by a large margin. Moreover, we find that models that allow lower level interactions between different feature types produce better results than treating lexical and prosodic contributions as separate, independent information sources. Catherine Lai, Mireia Farrús, Johanna D. Moore |
INTERSPEECH | 2 |
| 2012 | Study and correlation analysis of linguistic, perceptual, and automatic machine translation evaluationsabstractAbstract Evaluation of machine translation output is an important task. Various human evaluation techniques as well as automatic metrics have been proposed and investigated in the last decade. However, very few evaluation methods take the linguistic aspect into account. In this article, we use an objective evaluation method for machine translation output that classifies all translation errors into one of the five following linguistic levels: orthographic, morphological, lexical, semantic, and syntactic. Linguistic guidelines for the target language are required, and human evaluators use them in to classify the output errors. The experiments are performed on English‐to‐Catalan and Spanish‐to‐Catalan translation outputs generated by four different systems: 2 rule‐based and 2 statistical. All translations are evaluated using the 3 following methods: a standard human perceptual evaluation method, several widely used automatic metrics, and the human linguistic evaluation. Pearson and Spearman correlation coefficients between the linguistic, perceptual, and automatic results are then calculated, showing that the semantic level correlates significantly with both perceptual evaluation and automatic metrics. Mireia Farrús, Marta R. Costa-jussà, Maja Popovic |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Linguistic-based Evaluation Criteria to identify Statistical Machine Translation Errors
Mireia Farrús, Marta R. Costa-jussà, José B. Mariño, José A. R. Fonollosa |
EAMT | 1 |
| 2010 | Automatic and Human Evaluation Study of a Rule-based and a Statistical Catalan-Spanish Machine Translation Systems
Marta R. Costa-jussà, Mireia Farrús, José B. Mariño, José A. R. Fonollosa |
LREC | 2 |
| 2009 | Improving a Catalan-Spanish Statistical Translation System using Morphosyntactic Knowledge
Mireia Farrús, Marta R. Costa-jussà, Marc Poch, Adolfo Hernández, José B. Mariño |
EAMT | 1 |
| 2008 | Robustness of prosodic features to voice imitationabstractComunicació presentada a 9th Annual Conference of the International Speech Communication Association celebrada a Brisbane (Australia) del 22 al 26 de setembre de 2008. Mireia Farrús, Michael Wagner 0004, Jan Anguita, Javier Hernando |
INTERSPEECH | 1 |
| 2007 | Jitter and shimmer measurements for speaker recognitionabstractComunicació presentada a: 8th Annual Conference of the International Speech Communication Association a Antwerp (Belgium) celebrada del 27 al 31 d'agost de 2007. Mireia Farrús, Javier Hernando, Pascual Ejarque |
INTERSPEECH | 1 |
| 2006 | Person Verification by Fusion of Prosodic, Voice Spectral and Facial Parameters
Javier Hernando, Mireia Farrús, Pascual Ejarque, Ainara Garde, Jordi Luque |
SECRYPT | 2 |