Roser Morante

dblp:00/4850 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
8since 2021 · last 2025
0000-0002-8356-6469ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 11 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
abstract
In this article we present UNED-ACCESS 2024, a bilingual dataset that consists of 1003 multiple-choice questions of university entrance level exams in Spanish and English. Questions are originally formulated in Spanish and manually translated into English, and have not ever been publicly released, ensuring minimal contamination when evaluating Large Language Models with this dataset. A selection of current open-source and proprietary models are evaluated in a uniform zero-shot experimental setting both on the UNED-ACCESS 2024 dataset and on an equivalent subset of MMLU questions. Results show that (i) Smaller models not only perform worse than the largest models, but also degrade faster in Spanish than in English. The performance gap between both languages is negligible for the best models, but grows up to 37% for smaller models; (ii) Model ranking on UNED-ACCESS 2024 is almost identical (0.98 Pearson correlation) to the one obtained with MMLU (a similar, but publicly available benchmark), suggesting that contamination affects similarly to all models, and (iii) As in publicly available datasets, reasoning questions in UNED-ACCESS are more challenging for models of all sizes.
Eva Sánchez-Salido, Roser Morante, Julio Gonzalo 0001, Guillermo Marco, Jorge Carrillo de Albornoz, Laura Plaza, Enrique Amigó, Andrés Fernández García, Alejandro Benito-Santos, Adrián Ghajari Espinosa, Víctor Fresno-Fernández
COLING2
2025 EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos
Laura Plaza, Jorge Carrillo de Albornoz, Iván Árcos, Paolo Rosso, Damiano Spina, Enrique Amigó, Julio Gonzalo 0001, Roser Morante
ECIR (5)8
2024 A Web Portal about the State of the Art of NLP Tasks in Spanish
abstract
This paper presents a new web portal with information about the state of the art of natural language processing tasks in Spanish. It provides information about forums, competitions, tasks and datasets in Spanish, that would otherwise be spread in multiple articles and web sites. The portal consists of overview pages where information can be searched for and filtered by several criteria and individual pages with detailed information and hyperlinks to facilitate navigation. Information has been manually curated from publications that describe competitions and NLP tasks from 2013 until 2023 and will be updated as new tasks appear. A total of 185 tasks and 128 datasets from 94 competitions have been introduced.
Enrique Amigó, Jorge Carrillo de Albornoz, Andrés Fernández, Julio Gonzalo 0001, Guillermo Marco, Roser Morante, Laura Plaza, Jacobo Pedrosa-Marín
LREC/COLING6
2024 EXIST 2024: sEXism Identification in Social neTworks and Memes
Laura Plaza, Jorge Carrillo de Albornoz, Enrique Amigó, Julio Gonzalo 0001, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Víctor Ruiz
ECIR (5)5
2023 Overview of EXIST 2023: sEXism Identification in Social NeTworks
Laura Plaza, Jorge Carrillo de Albornoz, Roser Morante, Enrique Amigó, Julio Gonzalo 0001, Damiano Spina, Paolo Rosso
ECIR (3)3
2022 Identifying Copied Fragments in a 18th Century Dutch Chronicle
abstract
We apply computational stylometric techniques to an 18th century Dutch chronicle to determine which fragments of the manuscript represent the author’s own original work and which show signs of external source use through either direct copying or paraphrasing. Through stylometric methods the majority of text fragments in the chronicle can be correctly labelled as either the author’s own words, direct copies from sources or paraphrasing. Our results show that clustering text fragments based on stylometric measures is an effective methodology for authorship verification of this document; however, this approach is less effective when personal writing style is masked by author independent styles or when applied to paraphrased text.
Roser Morante, Eleanor L. T. Smith, Lianne Wilhelmus, Alie Lassche, Erika Kuijpers
LREC1
2021 Processing negation: An introduction to the special issue
Eduardo Blanco 0002, Roser Morante
Nat. Lang. Eng.2
2021 Recent advances in processing negation
abstract
Abstract Negation is a complex linguistic phenomenon present in all human languages. It can be seen as an operator that transforms an expression into another expression whose meaning is in some way opposed to the original expression. In this article, we survey previous work on negation with an emphasis on computational approaches. We start defining negation and two important concepts: scope and focus of negation. Then, we survey work in natural language processing that considers negation primarily as a means to improve the results in some task. We also provide information about corpora containing negation annotations in English and other languages, which usually include a combination of annotations of negation cues, scopes, foci, and negated events. We continue the survey with a description of automated approaches to process negation, ranging from early rule-based systems to systems built with traditional machine learning and neural networks. Finally, we conclude with some reflections on current progress and future directions.
Roser Morante, Eduardo Blanco 0002
Nat. Lang. Eng.1
2020 Must Children be Vaccinated or not? Annotating Modal Verbs in the Vaccination Debate
abstract
In this paper we analyze the use of modal verbs in a corpus of texts related to the vaccination debate. Broadly speaking, the vaccination debate centers around whether vaccination is safe, and whether it is morally acceptable to enforce mandatory vaccination. In order to successfully intervene and curb the spread of preventable diseases due to low vaccination rates, health practitioners need to be adequately informed on public perception of the safety and necessity of vaccines. Public perception can relate to the strength of conviction that an individual may have towards a proposition (e.g. ‘one must vaccinate’ versus ‘one should vaccinate’), as well as qualify the type of proposition, be it related to morality (‘government should not interfere in my personal choice’) or related to possibility (‘too many vaccines at once could hurt my child’). Text mining and analysis of modal auxiliaries are economically viable means of gaining insights into these perspectives, particularly on a large scale due to the widespread use of social media and blogs as vehicles of communication.
Liza King, Roser Morante
LREC2
2020 Annotating Perspectives on Vaccination
abstract
In this paper we present the Vaccination Corpus, a corpus of texts related to the online vaccination debate that has been annotated with three layers of information about perspectives: attribution, claims and opinions. Additionally, events related to the vaccination debate are also annotated. The corpus contains 294 documents from the Internet which reflect different views on vaccinations. It has been compiled to study the language of online debates, with the final goal of experimenting with methodologies to extract and contrast perspectives in the framework of the vaccination debate.
Roser Morante, Chantal van Son, Isa Maks, Piek Vossen
LREC1
2020 Detecting Negation Cues and Scopes in Spanish
abstract
In this work we address the processing of negation in Spanish. We first present a machine learning system that processes negation in Spanish. Specifically, we focus on two tasks: i) negation cue detection and ii) scope identification. The corpus used in the experimental framework is the SFU Corpus. The results for cue detection outperform state-of-the-art results, whereas for scope detection this is the first system that performs the task for Spanish. Moreover, we provide a qualitative error analysis aimed at understanding the limitations of the system and showing which negation cues and scopes are straightforward to predict automatically, and which ones are challenging.
Salud M. Jiménez-Zafra, Roser Morante, Eduardo Blanco 0002, María Teresa Martín Valdivia, Luis Alfonso Ureña López
LREC2
2020 Corpora Annotated with Negation: An Overview
abstract
Negation is a universal linguistic phenomenon with a great qualitative impact on natural language processing applications. The availability of corpora annotated with negation is essential to training negation processing systems. Currently, most corpora have been annotated for English, but the presence of languages other than English on the Internet, such as Chinese or Spanish, is greater every day. In this study, we present a review of the corpora annotated with negation information in several languages with the goal of evaluating what aspects of negation have been annotated and how compatible the corpora are. We conclude that it is very difficult to merge the existing corpora because we found differences in the annotation schemes used, and most importantly, in the annotation guidelines: the way in which each corpus was tokenized and the negation elements that have been annotated. Differently than for other well established tasks like semantic role labeling or parsing, for negation there is no standard annotation scheme nor guidelines, which hampers progress in its treatment.
Salud M. Jiménez-Zafra, Roser Morante, María Teresa Martín Valdivia, Luis Alfonso Ureña López
Comput. Linguistics2
2018 Scoring and Classifying Implicit Positive Interpretations: A Challenge of Class Imbalance
abstract
This paper reports on a reimplementation of a system on detecting implicit positive meaning from negated statements. In the original regression experiment, different positive interpretations per negation are scored according to their likelihood. We convert the scores to classes and report our results on both the regression and classification tasks. We show that a baseline taking the mean score or most frequent class is hard to beat because of class imbalance in the dataset. Our error analysis indicates that an approach that takes the information structure into account (i.e. which information is new or contrastive) may be promising, which requires looking beyond the syntactic and semantic characteristics of negated statements.
Chantal van Son, Roser Morante, Lora Aroyo, Piek Vossen
COLING2
2018 A review of Spanish corpora annotated with negation
abstract
The availability of corpora annotated with negation information is essential to develop negation processing systems in any language. However, there is a lack of these corpora even for languages like English, and when there are corpora available they are small and the annotations are not always compatible across corpora. In this paper we review the existing corpora annotated with negation in Spanish with the purpose of first, gathering the information to make it available for other researchers and, second, analyzing how compatible are the corpora and how has the linguistic phenomenon been addressed. Our final aim is to develop a supervised negation processing system for Spanish, for which we need training and test data. Our analysis shows that it will not be possible to merge the small corpora existing for Spanish due to lack of compatibility in the annotations.
Salud M. Jiménez-Zafra, Roser Morante, María Teresa Martín Valdivia, Luis Alfonso Ureña López
COLING2
2018 Systems' Agreements and Disagreements in Temporal Processing: An Extensive Error Analysis of the TempEval-3 Task
Tommaso Caselli, Roser Morante
LREC2
2018 Resource Interoperability for Sustainable Benchmarking: The Case of Events
Chantal van Son, Oana Inel, Roser Morante, Lora Aroyo, Piek Vossen
LREC3
2016 GRaSP: A Multilayered Annotation Scheme for Perspectives
Chantal van Son, Tommaso Caselli, Antske Fokkens, Isa Maks, Roser Morante, Lora Aroyo, Piek Vossen
LREC5
2012 A Statistical Relational Learning Approach to Identifying Evidence Based Medicine Categories
Mathias Verbeke, Vincent Van Asch, Roser Morante, Paolo Frasconi, Walter Daelemans, Luc De Raedt
EMNLP-CoNLL3
2012 The Netlog Corpus. A Resource for the Study of Flemish Dutch Internet Language
Mike Kestemont, Claudia Peersman, Benny De Decker, Guy De Pauw, Kim Luyckx, Roser Morante, Frederik Vaassen, Janneke van de Loo, Walter Daelemans
LREC6
2012 ConanDoyle-neg: Annotation of negation cues and their scope in Conan Doyle stories
Roser Morante, Walter Daelemans
LREC1
2012 Processing modality and negation
Roser Morante
HLT-NAACL1
2012 Modality and Negation: An Introduction to the Special Issue
abstract
Traditionally, most research in NLP has focused on propositional aspects of meaning. To truly understand language, however, extra-propositional aspects are equally important. Modality and negation typically contribute significantly to these extra-propositional meaning aspects. Although modality and negation have often been neglected by mainstream computational linguistics, interest has grown in recent years, as evidenced by several annotation projects dedicated to these phenomena. Researchers have started to work on modeling factuality, belief and certainty, detecting speculative sentences and hedging, identifying contradictions, and determining the scope of expressions of modality and negation. In this article, we will provide an overview of how modality and negation have been modeled in computational linguistics.
Roser Morante, Caroline Sporleder
Comput. Linguistics1
2011 Kernel-Based Logical and Relational Learning with kLog for Hedge Cue Detection
Mathias Verbeke, Paolo Frasconi, Vincent Van Asch, Roser Morante, Walter Daelemans, Luc De Raedt
ILP4
2010 Descriptive Analysis of Negation Cues in Biomedical Texts
Roser Morante
LREC1
2010 Highlights of the BioTM 2010 workshop on advances in bio text mining
abstract
Recently, the application of text mining (TM) and natural language processing (NLP) techniques to the biological and medical sciences has received increasing interest. In addition to many new workshops and conferences arising in this domain, recently also a number of community-wide tasks were conducted to benchmark text mining techniques on specific challenges (e.g. BioCreative, BioNLP Shared Task, ...)
Thomas Abeel, Sofie Van Landeghem, Roser Morante, Vincent Van Asch, Yves Van de Peer, Walter Daelemans, Yvan Saeys
BMC Bioinform.3
2009 A Metalearning Approach to Processing the Scope of Negation
Roser Morante, Walter Daelemans
CoNLL1
2008 Analysis of Joint Inference Strategies for the Semantic Role Labeling of Spanish and Catalan
Mihai Surdeanu, Roser Morante, Lluís Màrquez
CICLing2
2008 A Combined Memory-Based Semantic Role Labeler of English
Roser Morante, Walter Daelemans, Vincent Van Asch
CoNLL1
2008 Learning the Scope of Negation in Biomedical Texts
Roser Morante, Anthony M. L. Liekens, Walter Daelemans
EMNLP1
2008 CNTS: Memory-Based Learning of Generating Repeated References
Iris Hendrickx, Walter Daelemans, Kim Luyckx, Roser Morante, Vincent Van Asch
INLG4
2008 Semantic Role Labeling Tools Trained on the Cast3LB-CoNNL-SemRol Corpus
Roser Morante
LREC1