EDBT 2026 Demo / reviewers in the wild / expert
Helena Moniz
dblp:65/6363
· DBLP profile ↗
42ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-0900-6938ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Information Extraction and Template Filling from Client Style GuidesabstractStyle guides are a centrepiece of professional translation workflows. Yet, their integration into automatic pipelines remains underexplored. This paper presents exploratory work on information extraction from client style guides and application to a templated style guide, developed to be a system prompt. This template is then applied during an LLM-based translation to automatically produce outputs that are compliant to client’s requirements. The study focused on seven language pairs~(LP), evaluating the automatic extraction, and translation quality and compliance with the style guide. The extraction demonstrated reliable performance across languages and file formats. Translation quality and adherence were evaluated using human preference annotation, comparing two Tower models (Tower Zen 9B and Tower+ 72B). The results indicate a modest advantage for Tower+, but with mutual acceptability in certain instances. These findings establish a viable semi-automatic framework for style guide integration in translation workflows, and motivate further investigation across broader domains, clients, and LPs. Leonor Graça, Vera Cabarrão, Helena Moniz |
EAMT (2) | 3 |
| 2026 | A Multilingual Red Teaming-Driven Safety Analysis of LLMsabstractThis work benchmarks safety across several large language models (LLMs) and compares their performances through red teaming, which simulates adversarial attacks and identifies vulnerabilities in the systems. Using two public datasets and a proprietary dataset, the models were tested with three purposes. First, a red teaming test was conducted to establish a safety comparison between five models in English and Portuguese. The results revealed that, in general, Sugarloaf 3.1 is the safest model, but that Vesuvius 4.0 slightly outperforms it in Portuguese, also revealing that both outperform GPT-4o. Afterwards, three models were tested with one guardrailing prompt, that encourages safe interactions, and two content moderation prompts, in both languages, to understand the strengths of the current guardrails, as well as the effectiveness of the content moderation task. The results show that current guardrails are sufficient, notwithstanding room for improvement (particularly for Portuguese), but that the performance of the content moderation task was substandard, even for the best performing model – GPT-4o. Finally, the 3.0 TowerLLM models were tested in English to evaluate the effect that tokens and temperature have on the output, revealing that an intermediate token limit leads to safer responses while a higher temperature causes performance degradation. Patrícia Pandeiro, Vera Cabarrão, Helena Moniz |
EAMT (1) | 3 |
| 2025 | Diving into Gender Translation Bias for the Portuguese LanguageabstractBias in Machine Translation models has become a significant concern. Despite extensive research in several language pairs, Portuguese remains under-explored. This study investigates gender bias in English-to-Portuguese translation. We extend an established dataset to include inter-sentence examples and conduct a series of experiments to compare commercial Machine Translation systems, general-purpose Large Language Models, and non-commercial translation-specific models across various dimensions of gender bias. Additionally, we compare gender bias in Portuguese translation with that in other Romance languages (French, Spanish, and Italian). Finally, we explore whether sentiment influences gender bias in English-to-Portuguese translation. Our contributions include (1) a detailed analysis of gender bias in English-to-Portuguese Machine Translation, (2) an extended dataset incorporating inter-sentence evaluations, and (3) a multi-faceted comparative analysis across models and Romance languages. Sofia Bonifácio, Helena Moniz, Luísa Coheur |
ECAI | 2 |
| 2025 | Exploratory Study of Filled Pauses in Ukrainian Language: Phonetic Properties of Filled Pauses
Anna Havras, Carlos Mendes, Helena Moniz, Gueorgui Hristovsky, João Miranda |
INTERSPEECH | 3 |
| 2025 | The BridgeAI ProjectabstractThis paper presents an updated overview of the ‘BridgeAI’ project, a science-for-policy initiative funded by the Portuguese Foundation for Science and Technology (FCT) and the Recovery and Resilience Programme. In its second stage of implementation, BridgeAI continues to build upon its original goals, working towards a strategy to align AI research, policy, regulatory frameworks, and practical application. The project provides Portugal with an evidence-based framework to implement the EU Artificial Intelligence (AI) Act (AIA), ensuring responsible AI innovation through multidisciplinary collaboration. BridgeAI connects academia, industry, public administration, and civil society to create actionable insights and regulatory recommendations. This paper details the project’s latest advancements, key recommendations, and future directions. Helena Moniz, António Novais, Joana Lamego, Nuno André |
MTSummit (2) | 1 |
| 2025 | Cultural Transcreation in Asian Languages with Prompt-Based LLMsabstractThis research explores Cultural Transcreation (CT) for East Asian languages, focusing primarily on Mandarin Chinese (ZH) and the customer service (CS) market. We combined Large Language Models (LLMs) with prompt engineering to develop a CT product that, aligned with the Augmented Translation concept, enhances multilingual CS communication, enables professionals to engage with their target audience effortlessly, and improves overall service quality. Through a series of preparatory steps, including guideline establishment, benchmark validation, iterative prompt refinement, and LLM testing, we integrated the CT product into the CS platform, assessed its performance, and refined prompts based on a pilot feedback. The results highlight its success in empowering agents, regardless of linguistic or cultural expertise, to bridge effective communication gaps through AI-assisted cultural rephrasing, thus achieving its market launch. Beyond CS, the study extends the concept of transcreation and prompt-based LLM applications to other fields, discussing its performance in the language conversion of website content and advertising. Helena Wu, Beatriz Silva, Vera Cabarrão, Helena Moniz |
MTSummit (2) | 4 |
| 2024 | The Center for Responsible AI ProjectabstractThis paper describes the project “NextGenAI: Center for Responsible AI”, a 39-month Mobilizing and Green Agenda for Business Innovation funded by the Portuguese Recovery and Resilience Plan, under the Recovery and Resilience Facility (RRF). The project aims to create a new Center for Responsible AI in Portugal, capable of delivering more than 20 AI products in crucial areas like “Life Sciences”, many of which use generative AI, particularly NLP models such as those for Machine Translation, contributing to translating into legislation the European Law included in the EU AI Act, and creating a critical mass in the development of responsible AI technologies. To accomplish this mission, the Center for Responsible AI is formed by an ecosystem of startups and research institutions driving research in a virtuous way by addressing real market needs and opportunities in Responsible AI. Maria Ana Henriques, Ana C. Farinha, Nuno André, António Novais, Sara Guerreiro de Sousa, Bruno Prezado Silva, Helena Moniz, André F. T. Martins, Paulo Dimas |
EAMT (2) | 8 |
| 2024 | The BridgeAI ProjectabstractThis paper describes the project “BridgeAI: Boosting Regulatory Implementation with Data-driven insights, Global expertise, and Ethics for AI”, a one-year science-for-policy research project funded by the Portuguese Foundation for Science and Technology (FCT). The project aims to provide decision-makers in Portugal with the best context to implement the EU Artificial Intelligence (AI) Act and bridge the gap between AI research and policy. Although not exclusively on machine translation, the project pertains to natural language processing in general and ultimately to each of us as citizens. Helena Moniz, Joana Lamego, Nuno André, António Novais, Maria Henriques, Mariana Dalblon, Paulo Dimas |
EAMT (2) | 1 |
| 2024 | Cultural Transcreation with LLMs as a new productabstractWe present how at Unbabel we have been using Large Language Models to apply a Cultural Transcreation (CT) product on customer support (CS) emails and how we have been testing the quality and potential of this product. We discuss our preliminary evaluation of the performance of different MT models in the task of translating rephrased content and the quality of the translation outputs. Furthermore, we introduce the live pilot programme and the corresponding relevant findings, showing that transcreated content is not only culturally adequate but it is also of high rephrasing and translation quality. Beatriz Silva, Helena Wu, Yan Jingxuan, Vera Cabarrão, Helena Moniz, Sara Guerreiro de Sousa, Malene Sjørslev Søholm, Ana C. Farinha, Paulo Dimas |
EAMT (2) | 5 |
| 2024 | Generating subject-matter expertise assessment questions with GPT-4: a medical translation use-caseabstractThis paper examines the suitability of a large language model (LLM), GPT-4, for generating multiple choice questions (MCQs) aimed at assessing subject matter expertise (SME) in the domain of medical translation. The main objective of these questions is to model the skills of potential subject matter experts in a human-in-the-loop machine translation (MT) flow, to ensure that tasks are matched to the individuals with the right skill profile. The investigation was conducted at Unbabel, an artificial intelligence-powered human translation platform. Two medical translation experts evaluated the GPT-4-generated questions and answers, one focusing on English–European Portuguese, and the other on English–German. We present a methodology for creating prompts to elicit high-quality GPT-4 outputs for this use case, as well as for designing evaluation scorecards for human review of such output. Our findings suggest that GPT-4 has the potential to generate suitable items for subject matter expertise tests, providing a more efficient approach compared to relying solely on humans. Furthermore, we propose recommendations for future research to build on our approach and refine the quality of the outputs generated by LLMs. Diana Silveira, Marina Torrón, Helena Moniz |
EAMT (1) | 3 |
| 2023 | Quality Fit for Purpose: Building Business Critical Errors Test SuitesabstractThis paper illustrates a new methodology based on Test Suites (Avramidis et al., 2018) with focus on Business Critical Errors (BCEs) (Stewart et al., 2022) to evaluate the output of Machine Translation (MT) and Quality Estimation (QE) systems. We demonstrate the value of relying on semi-automatic evaluation done through scalable BCE-focused Test Suites to monitor both MT and QE systems’ performance for 8 language pairs (LPs) and a total of 4 error categories. This approach allows us to not only track the impact of new features and implementations in a real business environment, but also to identify strengths and weaknesses in models regarding different error types, and subsequently know what to improve henceforth. Mariana Cabeça, Marianna Buchicchio, Madalena Gonçalves, Christine Maroti, João Godinho, Pedro Coelho, Helena Moniz, Alon Lavie |
EAMT | 7 |
| 2023 | Context-aware and gender-neutral Translation MemoriesabstractThis work proposes an approach to use Part-Of-Speech (POS) information to automatically detect context-dependent Translation Units (TUs) from a Translation Memory database pertaining to the customer support domain. In line with our goal to minimize context-dependency in TUs, we show how this mechanism can be deployed to create new gender-neutral and context-independent TUs. Our experiments, conducted across Portuguese (PT), Brazilian Portuguese (PT-BR), Spanish (ES), and Spanish-Latam (ES-LATAM), show that the occurrence of certain POS with specific words is accurate in identifying context dependency. In a cross-client analysis, we found that ~10% of the most frequent 13,200 TUs were context-dependent, with gender determining context-dependency in 98% of all confirmed cases. We used these findings to suggest gender-neutral equivalents for the most frequent TUs with gender constraints. Our approach is in use in the Unbabel translation pipeline, and can be integrated into any other Neural Machine Translation (NMT) pipeline. Marjolene Paulo, Vera Cabarrão, Helena Moniz, Miguel Menezes, Rachel Grewcock, Eduardo Farah |
EAMT | 3 |
| 2023 | A Context-Aware Annotation Framework for Customer Support Live Chat Machine TranslationabstractTo measure context-aware machine translation (MT) systems quality, existing solutions have recommended human annotators to consider the full context of a document. In our work, we revised a well known Machine Translation quality assessment framework, Multidimensional Quality Metrics (MQM), (Lommel et al., 2014) by introducing a set of nine annotation categories that allows to map MT errors to source document contextual phenomenon, for simplicity sake we named such phenomena as contextual triggers. Our analysis shows that the adapted categories set enhanced MQM’s potential for MT error identification, being able to cover up to 61% more errors, when compared to traditional non-context core MQM’s application. Subsequently, we analyzed the severity of these MT “contextual errors”, showing that the majority fall under the critical and major levels, further indicating the impact of such errors. Finally, we measured the ability of existing evaluation metrics in detecting the proposed MT “contextual errors”. The results have shown that current state-of-the-art metrics fall short in detecting MT errors that are caused by contextual triggers on the source document side. With the work developed, we hope to understand how impactful context is for enhancing quality within a MT workflow and draw attention to future integration of the proposed contextual annotation framework into current MQM’s core typology. Miguel Menezes, M. Amin Farajian, Helena Moniz, João Varelas Graça |
MTSummit (1) | 3 |
| 2022 | Multi3Generation: Multitask, Multilingual, Multimodal Language GenerationabstractThis paper presents the Multitask, Multilingual, Multimodal Language Generation COST Action – Multi3Generation (CA18231), an interdisciplinary network of research groups working on different aspects of language generation. This “meta-paper” will serve as reference for citations of the Action in future publications. It presents the objectives, challenges and a the links for the achieved outcomes. Anabela Barreiro, José Guilherme Camargo de Souza, Albert Gatt, Mehul Bhatt, Elena Lloret, Aykut Erdem, Dimitra Gkatzia, Helena Moniz, Irene Russo, Fábio N. Kepler, Iacer Calixto, Marcin Paprzycki, François Portet, Isabelle Augenstein, Mirela Alhasani |
EAMT | 8 |
| 2022 | Agent and User-Generated Content and its Impact on Customer Support MTabstractThis paper illustrates a new evaluation framework developed at Unbabel for measuring the quality of source language text and its effect on both Machine Translation (MT) and Human Post-Edition (PE) performed by non-professional post-editors. We examine both agent and user-generated content from the Customer Support domain and propose that differentiating the two is crucial to obtaining high quality translation output. Furthermore, we present results of initial experimentation with a new evaluation typology based on the Multidimensional Quality Metrics (MQM) Framework Lommel et al., 2014), specifically tailored toward the evaluation of source language text. We show how the MQM Framework Lommel et al., 2014) can be adapted to assess errors of monolingual source texts and demonstrate how very specific source errors propagate to the MT and PE targets. Finally, we illustrate how MT systems are not robust enough to handle very specific source noise in the context of Customer Support data. Madalena Gonçalves, Marianna Buchicchio, Craig Stewart, Helena Moniz, Alon Lavie |
EAMT | 4 |
| 2022 | A Case Study on the Importance of Named Entities in a Machine Translation Pipeline for Customer Support ContentabstractThis paper describes the research developed at Unbabel, a Portuguese Machine-translation start-up, that combines MT with human post-edition and focuses strictly on customer service content. We aim to contribute to furthering MT quality and good-practices by exposing the importance of having a continuously-in-development robust Named Entity Recognition system compliant with General Data Protection Regulation (GDPR). Moreover, we have tested semiautomatic strategies that support and enhance the creation of Named Entities gold standards to allow a more seamless implementation of Multilingual Named Entities Recognition Systems. The project described in this paper is the result of a shared work between Unbabel ́s linguists and Unbabel ́s AI engineering team, matured over a year. The project should, also, be taken as a statement of multidisciplinary, proving and validating the much-needed articulation between the different scientific fields that compose and characterize the area of Natural Language Processing (NLP). Miguel Menezes, Vera Cabarrão, Pedro Mota, Helena Moniz, Alon Lavie |
EAMT | 4 |
| 2022 | QUARTZ: Quality-Aware Machine TranslationabstractThis paper presents QUARTZ, QUality-AwaRe machine Translation, a project led by Unbabel which aims at developing machine translation systems that are more robust and produce fewer critical errors. With QUARTZ we want to enable machine translation for user-generated conversational content types that do not tolerate critical errors in automatic translations. José Guilherme Camargo de Souza, Ricardo Rei, Ana C. Farinha, Helena Moniz, André F. T. Martins |
EAMT | 4 |
| 2021 | Retrieval Augmentation for Deep Neural NetworksabstractDeep neural networks have achieved state-of-the-art results in various vision and/or language tasks. Despite the use of large training datasets, most models are trained by iterating over single input-output pairs, discarding the remaining examples for the current prediction. In this work, we actively exploit the training data, using the information from nearest training examples to aid the prediction both during training and testing. Specifically, our approach uses the target of the most similar training example to initialize the memory state of an LSTM model, or to guide attention mechanisms. We apply this approach to image captioning and sentiment analysis, respectively through image and text retrieval. Results confirm the effectiveness of the proposed approach for the two tasks, on the widely used Flickr8 and IMDB datasets. Our code is publicly available11htttp://github.com/RitaRamo/retrieval-augmentation-nn. Rita Ramos, Patrícia Pereira, Helena Moniz, João Paulo Carvalho 0001, Bruno Martins 0001 |
IJCNN | 3 |
| 2020 | Project MAIA: Multilingual AI Agent AssistantabstractThis paper presents the Multilingual Artificial Intelligence Agent Assistant (MAIA), a project led by Unbabel with the collaboration of CMU, INESC-ID and IT Lisbon. MAIA will employ cutting-edge machine learning and natural language processing technologies to build multilingual AI agent assistants, eliminating language barriers. MAIA’s translation layer will empower human agents to provide customer support in real-time, in any language, with human quality. André F. T. Martins, João Graça, Paulo Dimas, Helena Moniz, Graham Neubig |
EAMT | 4 |
| 2020 | Exploring Text and Audio Embeddings for Multi-Dimension Elderly Emotion Recognition
Mariana Julião, Alberto Abad, Helena Moniz |
INTERSPEECH | 3 |
| 2019 | Unbabel Talk - Human Verified Translations for Voice Instant Messaging
Luís Bernardo, Mathieu Giquel, Sebastião Quintas, Paulo Dimas, Helena Moniz, Isabel Trancoso |
INTERSPEECH | 5 |
| 2018 | Acoustic-prosodic Entrainment in Structural Metadata EventsabstractThis paper presents an acoustic-prosodic analysis of entrain- ment in a Portuguese map-task corpus. Our aim is to ana- lyze how turn-by-turn entrainment varies with distinct structural metadata events: types of sentence-like units (SU) in consecu- tive turns (e.g. interrogatives followed by declaratives, or both declaratives), and with the presence of discourse markers, affir- mative cue words, and disfluencies in the beginning of turns. Entrainment at turn-exchanges may be observed in terms of pitch, energy, duration, and voice quality. Regarding SU types, question-answer turns are the ones with stronger similarity, and declarative-interrogative pairs are the ones where less entrain- ment occurs, as expected. Moreover, in question-answer pairs, there is also stronger evidence of entrainment with Yes/No and Tag questions than with Wh- questions. In fact, these subtypes are coded in distinctive prosodic ways (moreover, the first sub- type has no associated lexical-syntactic cues in Portuguese, only prosodic). As for turn-initial structures, entrainment is stronger when the second turn begins with an affirmative cue word; less strong with ambiguous structures (such as ‘OK’), emphatic af- firmative answers, and negative answers; and scarce with dis- fluencies and discourse markers. The different degrees of local entrainment may be related with the informative structure of distinct structural metadata events. Vera Cabarrão, Fernando Batista, Helena Moniz, Isabel Trancoso, Ana Isabel Mata |
INTERSPEECH | 3 |
| 2017 | A Semi-Supervised Learning Approach for Acoustic-Prosodic Personality Perception in Under-Resourced DomainsabstractAutomatic personality analysis has gained attention in the last years as a fundamental dimension in human-To-human and human-To-machine interaction. However, it still suffers from limited number and size of speech corpora for specific domains, such as the assessment of children's personality. This paper investigates a semi-supervised training approach to tackle this scenario. We devise an experimental setup with age and language mismatch and two training sets: A small labeled training set from the Interspeech 2012 Personality Sub-challenge, containing French adult speech labeled with personality OCEAN traits, and a large unlabeled training set of Portuguese children's speech. As test set, a corpus of Portuguese children's speech labeled with OCEAN traits is used. Based on this setting, we investigate a weak supervision approach that iteratively refines an initial model trained with the labeled data-set using the unlabeled data-set. We also investigate knowledge-based features, which leverage expert knowledge in acoustic-prosodic cues and thus need no extra data. Results show that, despite the large mismatch imposed by language and age differences, it is possible to attain improvements with these techniques, pointing both to the benefits of using a weak supervision and expert-based acoustic-prosodic features across age and language. Rubén Solera-Ureña, Helena Moniz, Fernando Batista, Vera Cabarrão, Anna Pompili, Ramón Fernandez Astudillo, Joana Campos 0001, Ana Paiva 0001, Isabel Trancoso |
INTERSPEECH | 2 |
| 2017 | The INTERACT Project and Crisis MT
Sharon O'Brien, Chao-Hong Liu, Andy Way, João Graça, Helena Moniz, Ellie Kemp, Rebecca Petras |
MTSummit (2) | 6 |
| 2016 | SPA: Web-based Platform for easy Access to Speech Processing Modules
Fernando Batista, Pedro Curto, Isabel Trancoso, Alberto Abad, Jaime Ferreira, Eugénio Ribeiro, Helena Moniz, David Martins de Matos, Ricardo Ribeiro 0001 |
LREC | 7 |
| 2016 | The SpeDial datasets: datasets for Spoken Dialogue Systems analytics
José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Helena Moniz, Alberto Abad, Katerina Louka, Elias Iosif, Alexandros Potamianos |
LREC | 4 |
| 2015 | Combining multiple approaches to predict the degree of nativenessabstractAutomatic speaker nativeness assessment has multiple applications, such as second language learning and IVR systems. In this paper we view this as a regression problem, since the available labels are on a continuous scale. Multiple approaches were applied, such as phonotactic models, i-vectors, and goodness of pronunciation, covering both segmental and suprasegmental features. Different phonotactic models were adopted, either trained with the challenge data, or using additional multilingual data from other domains. The obtained values were later combined in multiple ways and fed to a support vector machine regressor. Results on the test set surpass the provided baseline and are in line with the results obtained on the remaining sets. This suggests that our models generalize well to other datasets Eugénio Ribeiro, Jaime Ferreira, Julia Olcoz, Alberto Abad, Helena Moniz, Fernando Batista, Isabel Trancoso |
INTERSPEECH | 5 |
| 2014 | OpenLogos Semantico-Syntactic Knowledge-Rich Bilingual Dictionaries
Anabela Barreiro, Fernando Batista, Ricardo Ribeiro 0001, Helena Moniz, Isabel Trancoso |
LREC | 4 |
| 2014 | Revising the annotation of a Broadcast News corpus: a linguistic approach
Vera Cabarrão, Helena Moniz, Fernando Batista, Ricardo Ribeiro 0001, Nuno J. Mamede, Hugo Meinedo, Isabel Trancoso, Ana Isabel Mata, David Martins de Matos |
LREC | 2 |
| 2014 | Teenage and adult speech in school context: building and processing a corpus of European Portuguese
Ana Isabel Mata, Helena Moniz, Fernando Batista, Julia Hirschberg |
LREC | 2 |
| 2014 | Prosodic, syntactic, semantic guidelines for topic structures across domains and corpora
Ana Isabel Mata, Helena Moniz, Telmo Móia, Anabela Gonçalves, Fátima Silva, Fernando Batista, Inês Duarte, Fátima De Cassia E. Oliveira, Isabel Falé |
LREC | 2 |
| 2014 | Speaking style effects in the production of disfluencies
Helena Moniz, Fernando Batista, Ana Isabel Mata, Isabel Trancoso |
Speech Commun. | 1 |
| 2013 | Disfluency detection based on prosodic features for university lecturesabstractThis paper focuses on the identification of disfluent sequences and their distinct structural regions, based on acoustic and prosodic features. Reported experiments are based on a corpus of university lectures in European Portuguese, with roughly 32h, and a relatively high percentage of disfluencies (7.6%). The set of features automatically extracted from the corpus proved to be discriminant of the regions contained in the production of a disfluency. Several machine learning methods have been applied, but the best results were achieved using Classification and Regression Trees (CART). The set of features which was most informative for cross-region identification encompasses word duration ratios, word confidence score, silent ratios, and pitch and energy slopes. Features such as the number of phones and syllables per word proved to be more useful for the identification of the interregnum, whereas energy slopes were most suited for identifying the interruption point. Henrique Medeiros, Helena Moniz, Fernando Batista, Isabel Trancoso, Luís Nunes |
INTERSPEECH | 2 |
| 2012 | Prosodic contex-based analysis of disfluenciesabstractThis work explores prosodic cues of disfluencies in a corpus of university lectures. Results show three significant (p < 0.001) trends: pitch and energy slopes are significantly different between the disfluency and the onset of fluency; those features are also relevant to disfluency type differentiation; and they do not seem to be a speakereffect. The best combination of linguistic features one can use to better predict the onset of fluency are pitch and energy resets as well as the presence of a silent pause immediately before a repair. Our results, thus, point out to a strategy of prosodic contrast rather than of parallelism. With this work we hope to contribute to the analysis of the prosodic behaviors in the production of the so called disfluencies and in the fluency repair in European Portuguese. Helena Moniz, Fernando Batista, Isabel Trancoso, Ana Isabel Mata |
INTERSPEECH | 1 |
| 2012 | Bilingual Experiments on Automatic Recovery of Capitalization and Punctuation of Automatic Speech TranscriptsabstractThis paper focuses on the tasks of recovering capitalization and punctuation marks from texts without that information, such as spoken transcripts, produced by automatic speech recognition systems. These two practical rich transcription tasks were performed using the same discriminative approach, based on maximum entropy, suitable for on-the-fly usage. Reported experiments were conducted both over Portuguese and English broadcast news data. Both force aligned and automatic transcripts were used, allowing to measure the impact of the speech recognition errors. Capitalized words and named entities are intrinsically related, and are influenced by time variation effects. For that reason, the so-called language dynamics have been addressed for the capitalization task. Language adaptation results indicate, for both languages, that the capitalization performance is affected by the temporal distance between the training and testing data. In what regards the punctuation task, this paper covers the three most frequent punctuation marks: full stop, comma, and question marks. Different methods were explored for improving the baseline results for full stop and comma. The first uses punctuation information extracted from large written corpora. The second applies different levels of linguistic structure, including lexical, prosodic, and speaker related features. The comma detection improved significantly in the first method, thus indicating that it depends more on lexical features. The second method provided even better results, for both languages and both punctuation marks, best results being achieved mainly for full stop. As for question marks, there is a small gain, but differences are not very significant, due to the relatively small number of question marks in the corpora. Fernando Batista, Helena Moniz, Isabel Trancoso, Nuno J. Mamede |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Processing 'yup!' and other short utterances in interactive speechabstractThe detection of short utterances in conversational or interactive speech is essential to the proper processing of meaning in spoken interaction. Short, simple utterances are extremely common, and because of their highly variable prosody, carry many different forms of subtle interpersonal information. This paper reports on our approach to this problem and describes some corpora we are working with as well as the results of an analysis showing overlapping segments to be significantly different in their prosodic characteristics. Nick Campbell 0001, John Kane 0002, Helena Moniz |
ICASSP | 3 |
| 2010 | Extending the punctuation module for european portugueseabstractThis paper describes our recent work on extending the punctuation module of automatic subtitles for Portuguese Broadcast News. The main improvement was achieved by the use of prosodic information. This enabled the extension of the previous module which covered only full stops and commas, to cover question marks as well. The approach uses lexical, acoustic and prosodic information. Our results show that the latter is relevant for all types of punctuation. An analysis of the results also shows what type of interrogative is better dealt with by our method, taking into account the specificities of Portuguese. This may lead to different results for different types of corpora, depending on the types of interrogatives that are more frequent. Fernando Batista, Helena Moniz, Isabel Trancoso, Hugo Meinedo, Ana Isabel Mata, Nuno J. Mamede |
INTERSPEECH | 2 |
| 2009 | Classification of disfluent phenomena as fluent communicative devices in specific prosodic contextsabstractThis work explores prosodic cues of disfluent phenomena. In our previous work, we conducted a perceptual experiment regarding (dis)fluency ratings. Results suggested that some disfluencies may be considered felicitous by listeners, namely filled pauses and prolongations. In an attempt to discriminate which linguistic features are more salient in the classification of disfluencies as either fluent or disfluent phenomena, we used CART techniques on a corpus of 3.5 hours of spontaneous and prepared non-scripted speech. CART results pointed out 2 splits: break indices and contour shape. The first split indicates that events uttered at breaks 3 and 4 are considered felicitous. The second shows that these events must have flat or ascending contours to be considered as such; otherwise they are strongly penalized. Our preliminary results suggest that there are regular trends in the production of these events, namely, prosodic phrasing and contour shape. Index Terms: prosody, disfluency, fluency rating Helena Moniz, Isabel Trancoso, Ana Isabel Mata |
INTERSPEECH | 1 |
| 2008 | How can you use disfluencies and still sound as a good speaker?
Helena Moniz, Ana Isabel Mata, Isabel Trancoso, Céu Viana |
INTERSPEECH | 1 |
| 2008 | The LECTRA Corpus - Classroom Lecture Transcriptions in European Portuguese
Isabel Trancoso, Rui Martins, Helena Moniz, Ana Isabel Mata, Céu Viana |
LREC | 3 |
| 2007 | On filled-pauses and prolongations in european portugueseabstractThis paper reports preliminary results from a study of disfluencies in European Portuguese, based on a corpus of prepared (non-scripted) and spontaneous oral presentations in high school context. We will focus on the contextual distribution and temporal patterns of filled pauses and segmental prolongations, as well as on the way those are rated by listeners. Results suggest that filled pauses and segmental prolongations behave alike, have similar functions and may be considered in complementary distribution, obeying general syntactic and prosodic constraints. Index Terms spontaneous speech, disfluencies, prosody. 1. Helena Moniz, Ana Isabel Mata, Céu Viana |
INTERSPEECH | 1 |
| 2006 | Recognition of classroom lectures in european portugueseabstractClassroom lectures may be very challenging for automatic speech recognizers, because the vocabulary may be very specific and the speaking style very spontaneous. Our first experiments using a recognizer trained for Broadcast News resulted in word error rates near 60%, clearly confirming the need for adaptation to the specific topic of the lectures, on one hand, and for better strategies for handling spontaneous speech. This paper describes our efforts in these two directions: the different domain adaptation steps that lowered the error rate to 45%, with very little transcribed adaptation material, and the exploratory study of spontaneous speech phenomena in European Portuguese, namely concerning filled pauses. Index Terms: spontaneous speech recognition, Portuguese. 1. Isabel Trancoso, Ricardo Nunes, Luís Neves, Céu Viana, Helena Moniz, Diamantino Caseiro, Ana Isabel Mata |
INTERSPEECH | 5 |