Luísa Coheur

dblp:27/6767 · also Maria Luísa Torres Ribeiro Marques da Silva Coheur · DBLP profile ↗
← Back
50ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-2456-5028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 2 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Genetic algorithms as an optimization strategy to reduce the complexity of fuzzy fingerprints from large language models
abstract
Large Language Models currently represent the forefront of Natural Language Processing for classification tasks. Recent studies have integrated Fuzzy Fingerprints as a novel classification layer within these models to enhance result interpretability and reduce overall model complexity, without substantially compromising performance. However, in more challenging settings, this framework requires a larger fingerprint size to achieve competitive results when compared to using the full classification model. In this work, we employ Genetic Algorithms to the Fuzzy Fingerprint framework to further optimize these fingerprints to surpass the performance of conventional large pre-trained classifiers, while also offering improvements in interpretability and reduced complexity. Our empirical analysis demonstrates that optimizing to a smaller fingerprint size not only improves interpretability but also, when compared to a baseline with a larger fingerprint size, delivers higher performance relative to leading-edge methodologies.
João Paulo Carvalho 0001, Luísa Coheur
Fuzzy Sets Syst.3
2025 Diving into Gender Translation Bias for the Portuguese Language
abstract
Bias in Machine Translation models has become a significant concern. Despite extensive research in several language pairs, Portuguese remains under-explored. This study investigates gender bias in English-to-Portuguese translation. We extend an established dataset to include inter-sentence examples and conduct a series of experiments to compare commercial Machine Translation systems, general-purpose Large Language Models, and non-commercial translation-specific models across various dimensions of gender bias. Additionally, we compare gender bias in Portuguese translation with that in other Romance languages (French, Spanish, and Italian). Finally, we explore whether sentiment influences gender bias in English-to-Portuguese translation. Our contributions include (1) a detailed analysis of gender bias in English-to-Portuguese Machine Translation, (2) an extended dataset incorporating inter-sentence evaluations, and (3) a multi-faceted comparative analysis across models and Romance languages.
Sofia Bonifácio, Helena Moniz, Luísa Coheur
ECAI3
2025 Towards Cyberbullying Detection: Building, Benchmarking and Longitudinal Analysis of Aggressiveness and Conflicts/Attacks Datasets From Twitter
abstract
Offense and hate speech are a source of online conflicts which have become common in social media and, as such, their study is a growing topic of research in machine learning and natural language processing. This article presents two Portuguese language offense-related datasets that deepen the study of the subject: an Aggressiveness dataset and a Conflicts/Attacks dataset. While the former is similar to other offense detection related datasets, the latter constitutes a novelty due to the use of the history of the interaction between users. Several studies were carried out to construct and analyze the data in the datasets. The first study included gathering expressions of verbal aggression witnessed by adolescents to guide data extraction for the datasets. The second study included extracting data from Twitter (in Portuguese) that matched the most frequent expressions/words/sentences that were identified in the previous study. The third study consisted in the development of the Aggressiveness dataset, the Conflicts/Attacks dataset, and classification models. In our fourth study, we proposed to examine whether online aggression and conflicts/attacks revealed any trend changes over time with a sample of 86 adolescents. With this study, we also proposed to investigate whether the amount of tweets sent over a period of 273 days was related to online aggression and conflicts/attacks. Finally, we analyzed the percentage of participants who participated in the aggressions and/or attacks/conflicts.
Paula Ferreira 0004, Nádia Salgado Pereira, Hugo Rosa, Sofia Oliveira, Luísa Coheur, Sofia Mateus Francisco, Sidclay Bezerra de Souza, Ricardo Ribeiro 0001, João Paulo Carvalho 0001, Paula Paulino, Isabel Trancoso, Ana Margarida Veiga Simão
IEEE Trans. Affect. Comput.5
2024 Information Retrieval Using Fuzzy Fingerprints
Gonçalo Raposo, João Paulo Carvalho 0001, Luísa Coheur, Bruno Martins 0001
IPMU (1)3
2024 xcomet : Transparent Machine Translation Evaluation through Fine-grained Error Detection
abstract
Abstract Widely used learned metrics for machine translation evaluation, such as Comet and Bleurt, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xcomet, an open-source learned metric designed to bridge the gap between these approaches. xcomet integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xcomet is largely capable of identifying localized critical errors and hallucinations.
Nuno Miguel Guerreiro, Ricardo Rei, Daan van Stigt, Luísa Coheur, Pierre Colombo, André F. T. Martins
Trans. Assoc. Comput. Linguistics4
2023 Towards Realistic Sign Language Animations
abstract
Current signing avatars are often described as unnatural as they cannot accurately reproduce all the subtleties of synchronized body behaviors of a human signer. In this paper, we investigate a new dynamic approach for transitions between signs and the effect of mouthing behaviors. Although native signers preferred animations with dynamic transitions, we did not find significant differences in comprehension and perceived naturalness scores. On the other hand, we show that including mouthing behaviors improved comprehension and perceived naturalness for novice Portuguese sign language learners.
Inês Lacerda, Hugo Nicolau, Luísa Coheur
IVA3
2023 Who Said That?: Selecting the Correct Persona from Conversational Text
abstract
In this paper, we explore the ability to detect personas from conversational data. First, we adapt the Persona-Chat dataset, a well-known dialogue dataset, to support the task of selecting the correct persona out of various candidates. Then, we introduce persona perturbations to create additional identical personas that act as more challenging distractors. We train three different BERT-based models in a multiple-choice fashion to select the correct persona from a group of distractor personas. We show that this approach is able to discern between the group of original persona candidates, however, these models struggle to maintain high performance when we employ very identical distractors obtained from the proposed perturbations.
João Paulo Carvalho 0001, Luísa Coheur
IVA3
2023 Prompting, Retrieval, Training: An exploration of different approaches for task-oriented dialogue generation
abstract
Task-oriented dialogue systems need to generate appropriate responses to help fulfill users' requests.This paper explores different strategies, namely prompting, retrieval, and finetuning, for task-oriented dialogue generation.Through a systematic evaluation, we aim to provide valuable insights and guidelines for researchers and practitioners working on developing efficient and effective dialogue systems for real-world applications.Evaluation is performed on the MultiWOZ and Taskmaster-2 datasets, and we test various versions of FLAN-T5, GPT-3.5, and GPT-4 models.Costs associated with running these models are analyzed, and dialogue evaluation is briefly discussed.Our findings suggest that when testing data differs from the training data, fine-tuning may decrease performance, favoring a combination of a more general language model and a prompting mechanism based on retrieved examples.
Gonçalo Raposo, Luísa Coheur, Bruno Martins 0001
SIGDIAL2
2023 PGTask: Introducing the Task of Profile Generation from Dialogues
abstract
Recent approaches have attempted to personalize dialogue systems by leveraging profile information into models.However, this knowledge is scarce and difficult to obtain, which makes the extraction/generation of profile information from dialogues a fundamental asset.To surpass this limitation, we introduce the Profile Generation Task (PGTask).We contribute with a new dataset for this problem, comprising profile sentences aligned with related utterances, extracted from a corpus of dialogues.Furthermore, using state-of-the-art methods, we provide a benchmark for profile generation on this novel dataset.Our experiments disclose the challenges of profile generation, and we hope that this introduces a new research direction.
João Paulo Carvalho 0001, Luísa Coheur
SIGDIAL3
2023 Onception: Active Learning with Expert Advice for Real World Machine Translation
abstract
Active learning can play an important role in low-resource settings (i.e., where annotated data is scarce), by selecting which instances may be more worthy to annotate. Most active learning approaches for Machine Translation assume the existence of a pool of sentences in a source language, and rely on human annotators to provide translations or post-edits, which can still be costly. In this article, we apply active learning to a real-world human-in-the-loop scenario in which we assume that: (1) the source sentences may not be readily available, but instead arrive in a stream; (2) the automatic translations receive feedback in the form of a rating, instead of a correct/edited translation, since the human-in-the-loop might be a user looking for a translation, but not be able to provide one. To tackle the challenge of deciding whether each incoming pair source–translations is worthy to query for human feedback, we resort to a number of stream-based active learning query strategies. Moreover, because we do not know in advance which query strategy will be the most adequate for a certain language pair and set of Machine Translation models, we propose to dynamically combine multiple strategies using prediction with expert advice. Our experiments on different language pairs and feedback settings show that using active learning allows us to converge on the best Machine Translation systems with fewer human interactions. Furthermore, combining multiple strategies using prediction with expert advice outperforms several individual active learning strategies with even fewer interactions, particularly in partial feedback settings.
Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha
Comput. Linguistics3
2022 Searching for COMETINHO: The Little Metric That Could
abstract
In recent years, several neural fine-tuned machine translation evaluation metrics such as COMET and BLEURT have been proposed. These metrics achieve much higher correlations with human judgments than lexical overlap metrics at the cost of computational efficiency and simplicity, limiting their applications to scenarios in which one has to score thousands of translation hypothesis (e.g. scoring multiple systems or Minimum Bayes Risk decoding). In this paper, we explore optimization techniques, pruning, and knowledge distillation to create more compact and faster COMET versions. Our results show that just by optimizing the code through the use of caching and length batching we can reduce inference time between 39% and 65% when scoring multiple systems. Also, we show that pruning COMET can lead to a 21% model reduction without affecting the model’s accuracy beyond 0.01 Kendall tau correlation. Furthermore, we present DISTIL-COMET a lightweight distilled version that is 80% smaller and 2.128x faster while attaining a performance close to the original model and above strong baselines such as BERTSCORE and PRISM.
Ricardo Rei, Ana C. Farinha, José Guilherme Camargo de Souza, Pedro G. Ramos, André F. T. Martins, Luísa Coheur, Alon Lavie
EAMT6
2022 Question Rewriting? Assessing Its Importance for Conversational Question Answering
Gonçalo Raposo, Bruno Martins 0001, Luísa Coheur
ECIR (2)4
2022 A Study on the Best Way to Compress Natural Language Processing Models
abstract
Current research in Natural Language Processing shows a growing number of models extensively trained with large computational budgets. However, these models present computationally demanding requirements, preventing them from being deployed in devices with strict resource and response latency limitations. In this paper, we apply state-of-the-art model compression techniques to create compact versions of several of these models. In order to evaluate whether the trade-off between model performance and budget is worthwhile, we evaluate them in terms of efficiency, model simplicity and environmental foot-print. We also present a brief comparison between uncompressed and compressed models when running in low-end hardware.
João Antunes, Miguel L. Pardal, Luísa Coheur
FUZZ-IEEE3
2022 Towards a sentiment-aware conversational agent
abstract
We propose an end-to-end sentiment-aware conversational agent based on two models: a reply sentiment prediction model and a text generation model, conditioned on the predicted sentiment and the context of the dialogue. Additionally, we propose to use a sentiment classification model to evaluate the sentiment expressed by the agent during the development of the model. Results show that explicitly guiding the text generation model with a pre-defined set of sentiment sentences leads to clear improvements, regarding the expressed sentiment and the quality of the generated text.
Isabel Dias, Ricardo Rei, Patrícia Pereira, Luísa Coheur
IVA4
2021 Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort
abstract
Vânia Mendonça, Ricardo Rei, Luisa Coheur, Alberto Sardinha, Ana Lúcia Santos. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Vânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha, Ana Lúcia Santos
ACL/IJCNLP (1)3
2020 Query Strategies, Assemble! Active Learning with Expert Advice for Low-resource Natural Language Processing
abstract
Active learning plays an important role in low-resource scenarios, i.e., when only a small amount of annotated instances is available. However, one does not know what is the best active learning strategy before actually testing a handful of strategies on a labeled set, which might not be viable in a real world low-resource scenario. Instead, it would be desirable to dynamically obtain the results from the best strategy on a given scenario, while using as little annotated resources as possible.In this paper, we present a novel application of prediction with expert advice to combine different query strategies as experts, giving a greater weight to those which select the most useful instances. We evaluated our approach in two Natural Language Processing (NLP) tasks: Part-of-Speech tagging (for English) and Named Entity Recognition (for Portuguese). Results show that our solution keeps up with the results of the best strategy in each scenario, nearly reaching fully supervised performance with only half of the annotated data.
Vânia Mendonça, Alberto Sardinha, Luísa Coheur, Ana Lúcia Santos
FUZZ-IEEE3
2020 From Eliza to Siri and Beyond
Luísa Coheur
IPMU (1)1
2020 To BERT or Not to BERT Dealing with Possible BERT Failures in an Entailment Task
Pedro Fialho, Luísa Coheur, Paulo Quaresma
IPMU (1)2
2020 HamNoSyS2SiGML: Translating HamNoSys Into SiGML
abstract
Sign Languages are visual languages and the main means of communication used by Deaf people. However, the majority of the information available online is presented through written form. Hence, it is not of easy access to the Deaf community. Avatars that can animate sign languages have gained an increase of interest in this area due to their flexibility in the process of generation and edition. Synthetic animation of conversational agents can be achieved through the use of notation systems. HamNoSys is one of these systems, which describes movements of the body through symbols. Its XML-compliant, SiGML, is a machine-readable input of HamNoSys able to animate avatars. Nevertheless, current tools have no freely available open source libraries that allow the conversion from HamNoSys to SiGML. Our goal is to develop a tool of open access, which can perform this conversion independently from other platforms. This system represents a crucial intermediate step in the bigger pipeline of animating signing avatars. Two cases studies are described in order to illustrate different applications of our tool.
Carolina C. Neves, Luísa Coheur, Hugo Nicolau
LREC2
2020 AIA-BDE: A Corpus of FAQs in Portuguese and their Variations
abstract
We present AIA-BDE, a corpus of 380 domain-oriented FAQs in Portuguese and their variations, i.e., paraphrases or entailed questions, created manually, by humans, or automatically, with Google Translate. Its aims to be used as a benchmark for FAQ retrieval and automatic question-answering, but may be useful in other contexts, such as the development of task-oriented dialogue systems, or models for natural language inference in an interrogative context. We also report on two experiments. Matching variations with their original questions was not trivial with a set of unsupervised baselines, especially for manually created variations. Besides high performances obtained with ELMo and BERT embeddings, an Information Retrieval system was surprisingly competitive when considering only the first hit. In the second experiment, text classifiers were trained with the original questions, and tested when assigning each variation to one of three possible sources, or assigning them as out-of-domain. Here, the difference between manual and automatic variations was not so significant.
Hugo Gonçalo Oliveira, João Ferreira 0004, Pedro Fialho, Ricardo Rodrigues 0001, Luísa Coheur, Ana Alves 0001
LREC6
2019 BeamSeg: A Joint Model for Multi-Document Segmentation and Topic Identification
abstract
We propose BeamSeg, a joint model for segmentation and topic identification of documents from the same domain.The model assumes that lexical cohesion can be observed across documents, meaning that segments describing the same topic use a similar lexical distribution over the vocabulary.The model implements lexical cohesion in an unsupervised Bayesian setting by drawing from the same language model segments with the same topic.Contrary to previous approaches, we assume that language models are not independent, since the vocabulary changes in consecutive segments are expected to be smooth and not abrupt.We achieve this by using a dynamic Dirichlet prior that takes into account data contributions from other topics.BeamSeg also models segment length properties of documents based on modality (textbooks, slides, etc.).The evaluation is carried out in three datasets.In two of them, improvements of up to 4.8% and 7.3% are obtained in the segmentation and topic identifications tasks, indicating that both tasks should be jointly modeled.
Pedro Mota, Maxine Eskénazi, Luísa Coheur
CoNLL3
2018 Improving Question Generation with the Teacher's Implicit Feedback
Hugo Rodrigues 0001, Luísa Coheur, Eric Nyberg
AIED (2)2
2018 Efficient Navigation in Learning Materials: An Empirical Study on the Linking Process
Pedro Mota, Luísa Coheur, Maxine Eskénazi
AIED (2)2
2018 Using Fuzzy Fingerprints for Cyberbullying Detection in Social Networks
abstract
As cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, we study how a recent technique with proven success in similar tasks, Fuzzy Fingerprints, performs when detecting textual cyberbullying in social networks. Despite being commonly treated as binary classification task, we argue that this is in fact a retrieval problem where the only relevant performance is that of retrieving cyberbullying interactions. Experiments show that the Fuzzy Fingerprints slightly outperforms baseline classifiers when tested in a close to real life scenario, where cyberbullying instances are rarer than those without cyberbullying.
Hugo Rosa, João Paulo Carvalho 0001, Pável Calado, Bruno Martins 0001, Ricardo Ribeiro 0001, Luísa Coheur
FUZZ-IEEE6
2018 A "Deeper" Look at Detecting Cyberbullying in Social Networks
abstract
As cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, a detailed look at the current state-of-the-art in cyberbullying detection reveals that deep learning techniques have seldom been used to tackle this problem, despite growing reputation in other text-based classification tasks. Motivated by neural networks' documented success, three architectures are implemented from similar works: a simple CNN, a hybrid CNN-LSTM and a mixed CNN-LSTM-DNN. In addition, three text representations are trained from three different sources, via the word2vec model: Google-News, Twitter and Formspring. The experiment shows that these models with one of the above embeddings beat other benchmark classifiers (Support Vector Machines and Logistic Regression) both in an unbalanced and balanced version of the same dataset.
Hugo Rosa, David Martins de Matos, Ricardo Ribeiro 0001, Luísa Coheur, João Paulo Carvalho 0001
IJCNN4
2018 MUSED: A multimedia multi-document dataset for topic segmentation
abstract
Abstract Research on topic segmentation has recently focused on segmenting documents by taking advantage of documents covering the same topics. In order to properly evaluate such approaches, a dataset of related documents is needed. However, existing datasets are limited in the number of related documents per domain. In addition, most of the available datasets do not consider documents from different media sources (PowerPoints, videos, etc.), which pose specific challenges to segmentation. We fill this gap with the MUltimedia SEgmentation Dataset (MUSED), a collection of documents manually segmented, from different media sources, in seven different domains, with an average of twenty related documents per domain. In this paper, we describe the process of building MUSED. A multi-annotator study is carried out to determine if it is possible to observe agreement among human judges and characterize their disagreement patterns. In addition, we use MUSED to compare the state-of-the-art topic segmentation techniques, including the ones that take advantage of related documents. Moreover, we study the impact of having documents from different media sources in the dataset. To the best of our knowledge, MUSED is the first dataset that allows a straightforward evaluation of both single- and multiple-documents topic segmentation techniques, as well as to study how these behave in the presence of documents from different media sources. Results show that some techniques are, indeed, sensitive to different media sources, and also that current multi-document segmentation models do not outperform previous models, pointing to a research line that needs to be boosted.
Pedro Mota, Maxine Eskénazi, Luísa Coheur
Nat. Lang. Eng.3
2016 QGASP: a Framework for Question Generation Based on Different Levels of Linguistic Information
Hugo Rodrigues 0001, Luísa Coheur, Eric Nyberg
INLG2
2016 Building a Corpus of Errors and Quality in Machine Translation: Experiments on Error Impact
Ângela Costa, Rui Correia, Luísa Coheur
LREC3
2015 ChatWoz: Chatting through a Wizard of Oz
abstract
Several cases of autistic children successfully interacting with virtual assistants such as Siri or Cortana have been recently reported. In this demo we describe ChatWoz, an application that can be used as a Wizard of Oz, to collect real data for dialogue systems, but also to allow children to interact with their caregivers through it, as it is based on a virtual agent. ChatWoz is composed of an interface controlled by the caregiver, which establishes what the agent will utter, in a synthesised voice. Several elements of the interface can be controlled, such as the agent's face emotions. In this paper we focus on the scenario of child-caregiver interaction and detail the features implemented in order to couple with it.
Pedro Fialho, Luísa Coheur
ASSETS2
2015 VITHEA-Kids: a Platform for Improving Language Skills of Children with Autism Spectrum Disorder
abstract
In this work, we present a platform designed for children with Autism Spectrum Disorder to develop language and generalization skills, in response to the lack of applications tailored for the unique abilities, symptoms, and challenges of the autistic children. This platform allows caregivers to build customized multiple choice exercises while taking into account specific needs/characteristics of each child. We also propose a module for the automatic generation of exercises, aiming to ease the task of exercise creation for caregivers.
Vânia Mendonça, Luísa Coheur, Alberto Sardinha
ASSETS2
2015 A linguistically motivated taxonomy for Machine Translation error analysis
Ângela Costa, Wang Ling, Tiago Luís, Rui Correia, Luísa Coheur
Mach. Transl.5
2014 Luke, I am Your Father: Dealing with Out-of-Domain Requests by Using Movies Subtitles
David Ameixa, Luísa Coheur, Pedro Fialho, Paulo Quaresma
IVA2
2014 Translation errors from English to Portuguese: an annotated corpus
Ângela Costa, Tiago Luís, Luísa Coheur
LREC3
2014 JUST.ASK, a QA system that learns to answer new questions from previous interactions
Sérgio Curto, Ana Cristina Mendes, Pedro Curto, Luísa Coheur, Ângela Costa
LREC4
2013 Introducing UWS - A fuzzy based word similarity function with good discrimination capability: Preliminary results
abstract
This paper introduces a novel word similarity function, the Uke Similarity Function (UWS), that fuses the most interesting characteristics of the two main philosophies in word and string matching: the edit distance and the n-gram similarity approach. It also uses fuzzy sets to integrate expert knowledge about typographical errors and to easily include phonetic and token related errors. The UWS was developed with the goal of automatic detection and correction of typographical and other word errors in unedited corpus data when creating word lists.
João Paulo Carvalho 0001, Luísa Coheur
FUZZ-IEEE2
2013 When the answer comes into question in question-answering: survey and open issues
abstract
Abstract The answer determines the success of a Question-Answering (QA) system. In redundancy-based QA systems, a common approach is to extract the candidate answers from the information sources and select the most frequent answers as the final answers. However, this strategy has some pitfalls. For instance, if a system is not able to detect equivalences between the candidate answers, their frequencies might be erroneously calculated. Moreover, the user who posed the question should also be taken into account when answering: different persons require different (correct) answers. This can involve the use of suitable vocabulary and/or information details. In these situations, the generation of a response can be a more suitable strategy, instead of the extraction and direct retrieval of the answer from the information sources. The present survey targets the state of the art in the answering task in QA under three different lines of research. First, we present several works that focus on relating candidate answers. Then, we recover the concept of cooperative answer – a correct, useful, and non-misleading answer – and we bring up attempts to address cooperative answering. Finally, we investigate the research community endeavors on response generation. We will also present our perspective on each of these three topics throughout this paper.
Ana Cristina Mendes, Luísa Coheur
Nat. Lang. Eng.2
2012 A critical survey on the use of Fuzzy Sets in Speech and Natural Language Processing
abstract
This paper shows how the use and applications of Fuzzy Sets (FS) in Speech and Natural Language Processing (SNLP) have seen a steady decline to a point where FS are virtually unknown or unappealing for most of the researchers currently working in the SNLP field, tries to find the reasons behind this decline, and proposes some guidelines on what could be done to reverse it and make FS assume a relevant role in SNLP.
João Paulo Carvalho 0001, Fernando Batista, Luísa Coheur
FUZZ-IEEE3
2012 An English-Portuguese parallel corpus of questions: translation guidelines and application in SMT
Ângela Costa, Tiago Luís, Joana Ribeiro, Ana Cristina Mendes, Luísa Coheur
LREC5
2012 Extending a wordnet framework for simplicity and scalability
Pedro Fialho, Sérgio Curto, Ana Cristina Mendes, Luísa Coheur
LREC4
2012 Dealing with unknown words in statistical machine translation
João Pedro Carlos Gomes da Silva, Luísa Coheur, Ângela Costa, Isabel Trancoso
LREC2
2011 Bootstrapping Multiple-Choice Tests with The-Mentor
Ana Cristina Mendes, Sérgio Curto, Luísa Coheur
CICLing (1)3
2011 BP2EP - Adaptation of Brazilian Portuguese texts to European Portuguese
Luís Marujo, Nuno Grazina, Tiago Luís, Wang Ling, Luísa Coheur, Isabel Trancoso
EAMT5
2011 It's Answer Time - Taking the Next Step in Question-Answering
Ana Cristina Mendes, Luísa Coheur
ICAART (1)2
2011 An Approach to Answer Selection in Question-Answering Based on Semantic Relations
abstract
A usual strategy to select the final answer in factoid Question-Answering (QA) relies on redundancy. A score is given to each candidate answer as a function of its frequency of occurrence, and the final answer is selected from the set of candidates sorted in decreasing order of score. For that purpose, systems often try to group together semantically equivalent answers. However, they hold several other semantic relations, such as inclusion, which are not considered, and candidates are mostly seen independently, as competitors. Our hypothesis is that not just equivalence, but other relations between candidate answers have impact on the performance of a redundancy-based QA system. In this paper, we describe experimental studies to back up this hypothesis. Our findings show that, with relatively simple techniques to recognize relations, systems’ accuracy can be improved for answers of categories NUMBER, DATE and ENTITY.
Ana Cristina Mendes, Luísa Coheur
IJCAI2
2011 Towards the Rapid Development of a Natural Language Understanding Module
Catarina Moreira, Ana Cristina Mendes, Luísa Coheur, Bruno Martins 0001
IVA3
2011 Controlling Complexity in Part-of-Speech Induction
abstract
We consider the problem of fully unsupervised learning of grammatical (part-of-speech) categories from unlabeled text. The standard maximum-likelihood hidden Markov model for this task performs poorly, because of its weak inductive bias and large model capacity. We address this problem by refining the model and modifying the learning objective to control its capacity via para- metric and non-parametric constraints. Our approach enforces word-category association sparsity, adds morphological and orthographic features, and eliminates hard-to-estimate parameters for rare words. We develop an efficient learning algorithm that is not much more computationally intensive than standard training. We also provide an open-source implementation of the algorithm. Our experiments on five diverse languages (Bulgarian, Danish, English, Portuguese, Spanish) achieve significant improvements compared with previous methods for the same task.
João Graça, Kuzman Ganchev, Luísa Coheur, Fernando Pereira 0003, Ben Taskar
J. Artif. Intell. Res.3
2010 Named Entity Recognition in Questions: Towards a Golden Collection
Ana Cristina Mendes, Luísa Coheur, Paula Vaz Lobo
LREC2
2009 Adapting a Virtual Agent to Users' Vocabulary and Needs
Ana Cristina Mendes, Rui Prada, Luísa Coheur
IVA3
2008 Building a Golden Collection of Parallel Multi-Language Word Alignment
João Graça, Joana Paulo Pardal, Luísa Coheur, Diamantino Caseiro
LREC3
2008 Supporting Named Entity Recognition and Syntactic Analysis with Full-Text Queries
Luísa Coheur, Ana Guimarães, Nuno J. Mamede
NLDB1