Patricia Martín-Rodilla

dblp:117/0754 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-1540-883XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
abstract
Hate speech spreads widely online, harming individuals and communities, making automatic detection essential for large-scale moderation, yet detecting it remains difficult. Part of the challenge lies in subjectivity: what one person flags as hate speech, another may see as benign. Traditional annotation agreement metrics, such as Cohen's $κ$, oversimplify this disagreement, treating it as an error rather than meaningful diversity. Meanwhile, Large Language Models (LLMs) promise scalable annotation, but prior studies demonstrate that they cannot fully replace human judgement, especially in subjective tasks. In this work, we reexamine LLM reliability using a subjectivity-aware framework, cross-Rater Reliability (xRR), revealing that even under fairer lens, LLMs still diverge from humans. Yet this limitation opens an opportunity: we find that LLM-generated annotations can reliably reflect performance trends across classification models, correlating with human evaluations. We test this by examining whether LLM-generated annotations preserve the relative ordering of model performance derived from human evaluation (i.e. whether models ranked as more reliable by human annotators preserve the same order when evaluated with LLM-generated labels). Our results show that, although LLMs differ from humans at the instance level, they reproduce similar ranking and classification patterns, suggesting their potential as proxy evaluators. While not a substitute for human annotators, they might serve as a scalable proxy for evaluation in subjective NLP tasks.
Paloma Piot-Perez-Abadin, David Otero 0001, Patricia Martín-Rodilla, Javier Parapar
LREC3
2025 Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated Data
abstract
We explore the use of discourse parsers for extracting a particular discourse structure in a real-world social media scenario. Specifically, we focus on enhancing parser performance through the integration of synthetic data generated by large language models (LLMs). We conduct experiments using a newly developed dataset of 1,170 local RST discourse structures, including 900 synthetic and 270 gold examples, covering three social media platforms: online news comments sections, a discussion forum (Reddit), and a social media messaging platform (Twitter). Our primary goal is to assess the impact of LLM-generated synthetic training data on parser performance in a raw text setting without pre-identified discourse units. While both top-down and bottom-up RST architectures greatly benefit from synthetic data, challenges remain in classifying evaluative discourse structures.
Martial Pastor, Nelleke Oostdijk, Patricia Martín-Rodilla, Javier Parapar
COLING3
2024 eRisk 2024: Depression, Anorexia, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani
ECIR (5)2
2024 MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
abstract
Hate speech represents a pervasive and detrimental form of online discourse, often manifested through an array of slurs, from hateful tweets to defamatory posts. As such speech proliferates, it connects people globally and poses significant social, psychological, and occasionally physical threats to targeted individuals and communities. Current computational linguistic approaches for tackling this phenomenon rely on labelled social media datasets for training. For unifying efforts, our study advances in the critical need for a comprehensive meta-collection, advocating for an extensive dataset to help counteract this problem effectively. We scrutinized over 60 datasets, selectively integrating those pertinent into MetaHate. This paper offers a detailed examination of existing collections, highlighting their strengths and limitations. Our findings contribute to a deeper understanding of the existing datasets, paving the way for training more robust and adaptable models. These enhanced models are essential for effectively combating the dynamic and complex nature of hate speech in the digital realm.
Paloma Piot-Perez-Abadin, Patricia Martín-Rodilla, Javier Parapar
ICWSM2
2024 IAT/ML: a metamodel and modelling approach for discourse analysis
abstract
Abstract Language technologies are gaining momentum as textual information saturates social networks and media outlets, compounded by the growing role of fake news and disinformation. In this context, approaches to represent and analyse public speeches, news releases, social media posts and other types of discourses are becoming crucial. Although there is a large body of literature on text-based machine learning, it tends to focus on lexical and syntactical issues rather than semantic or pragmatic. Being useful, these advances cannot tackle the nuanced and highly context-dependent problems of discourse evaluation that society demands. In this paper, we present IAT/ML, a metamodel and modelling approach to represent and analyse discourses. IAT/ML focuses on semantic and pragmatic issues, thus tackling a little researched area in language technologies. It does so by combining three different modelling approaches: ontological, which focuses on what the discourse is about; argumentation, which deals with how the text justifies what it says; and agency, which provides insights into the speakers’ beliefs, desires and intentions. Together, these three modelling approaches make IAT/ML a comprehensive solution to represent and analyse complex discourses towards their understanding, evaluation and fact checking.
Cesar Gonzalez-Perez, Martin Pereira-Fariña, Beatriz Calderón-Cerrato, Patricia Martín-Rodilla
Softw. Syst. Model.4
2023 eRisk 2023: Depression, Pathological Gambling, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani
ECIR (3)2
2022 eRisk 2022: Pathological Gambling, Depression, and Eating Disorder Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani
ECIR (2)2
2021 eRisk 2021: Pathological Gambling, Self-harm and Depression Challenges
Javier Parapar, Patricia Martín-Rodilla, David E. Losada, Fabio Crestani
ECIR (2)2
2021 Experimental Analysis of the Relevance of Features and Effects on Gender Classification Models for Social Media Author Profiling
abstract
[Abstract] Automatic user profiling from social networks has become a popular task due to its commercial applications (targeted advertising, market studies...). Automatic profiling models infer demographic characteristics of social network users from their generated content or interactions. Users’ demographic information is also precious for more social worrying tasks such as automatic early detection of mental disorders. For this type of users’ analysis tasks, it has been shown that the way how they use language is an important indicator which contributes to the effectiveness of the models. Therefore, we also consider that for identifying aspects such as gender, age or user’s origin, it is interesting to consider the use of the language both from psycho-linguistic and semantic features. A good selection of features will be vital for the performance of retrieval, classification, and decision-making software systems. In this paper, we will address gender classification as a part of the automatic profiling task. We show an experimental analysis of the performance of existing gender classification models based on external corpus and baselines for automatic profiling. We analyse in-depth the influence of the linguistic features in the classification accuracy of the model. After that analysis, we have put together a feature set for gender classification models in social networks with an accuracy performance above existing baselines.
Paloma Piot-Perez-Abadin, Patricia Martín-Rodilla, Javier Parapar
ENASE2
2021 Enriching linguistic descriptions of data: A framework for composite protoforms
abstract
One of the current limitations of fuzzy linguistic descriptions of data is the lack of diversity of protoforms that can be used to linguistically summarize data. Despite an important effort in providing protoforms with improved semantics that are applicable to time series data or specific application domains, type-I and type-II fuzzy quantified sentences are still predominant in the literature. In this context, we propose a different approach for defining new types of protoforms. Instead of understanding protoforms as individual primitives, our proposal draws inspiration from Rhetorical Structure Theory to provide a framework that allows to define new types of complex protoforms based on semantic relations among simpler protoforms. Based on this framework, we propose an initial taxonomy of relations among protoforms and provide an illustrative use case based on real data and evaluated by human users.
Alejandro Ramos-Soto, Patricia Martín-Rodilla
Fuzzy Sets Syst.2
2020 Adding Temporal Dimension to Ontology Learning Models for Depression Signs Detection from Social Media Texts
abstract
Approaches to early detection of depression based on individual's language are receiving increasing attention, with detection software systems based on lexical, grammatical or discursive components applied to medical corpus or social media texts. However, these first detection systems are defragmented, each attending to a specific feature or linguistic level, and not addressing a more conceptual level. Existing ontology learning (OL) methods extract the ontology referred in the text. In addition, existing systems perform language analysis for the detection of depression as a snapshot of each individual, regardless of their temporal dimension. Is it possible that suitable linguistic features to detect early signs of depression vary over time? And the underlying ontology? This paper presents a model that adds the temporal component to current ontology learning models to perform evolutionary analysis of both linguistic and ontological features to texts from social networks. The model has been applied to an external corpus of depression from social media texts, with a two-fold goal: 1) validating the model by contrasting it with OL models without temporal component 2) producing a corpus of evolutionary OL results applied to the depression detection from social media texts.
Patricia Martín-Rodilla
ENASE1
2019 Metainformation scenarios in Digital Humanities: Characterization and conceptual modelling strategies
abstract
Requirements for the analysis, interpretation and reuse of information are becoming more and more ambitious as we generate larger and more complex datasets. This is leading to the development and widespread use of information about information, often called metainformation (or metadata) in most disciplines. The Digital Humanities are not an exception. We often assume that metainformation helps us in documenting information for future reference by recording who has created it, when and how, among other aspects. We also assume that recording metainformation will facilitate the tasks of interpreting information at later stages. However, some works have identified some issues with existing metadata approaches, related to 1) the proliferation of too many “standards” and difficulties to choose between them; 2) the generalized assumption that metadata and data (or metainformation and information) are essentially different, and the subsequent development of separate sets of languages and tools for each (introducing redundant models); and 3) the combination of conceptual and implementation concerns within most approaches, violating basic engineering principles of modularity and separation of concerns. Some of these problems are especially relevant in Digital Humanities. In addition, we argue here that the lack of characterization of the scenarios in which metainformation plays a relevant role in humanistic projects often results in metainformation being recorded and managed without a specific purpose in mind. In turn, this hinders the process of decision making on issues such as what metainformation must be recorded in a specific project, and how it must be conceptualized, stored and managed. This paper presents a review of the most used metadata approaches in Digital Humanities and, taking a conceptual modelling perspective, analyses their major issues as outlined above. It also describes what the most common scenarios for the use of metainformation in Digital Humanities are, presenting a characterization that can assist in the setting of goals for metainformation recording and management in each case. Based on these two aspects, a new approach is proposed for the conceptualization, recording and management of metainformation in the Digital Humanities, using the ConML conceptual modelling language, and adopting the overall view that metainformation is not essentially different to information. The proposal is validated in Digital Humanities scenarios through case studies employing real-world datasets.
Patricia Martín-Rodilla, Cesar Gonzalez-Perez
Inf. Syst.1
2018 Assessing data analysis performance in research contexts: An experiment on accuracy, efficiency, productivity and researchers' satisfaction
Patricia Martín-Rodilla, José Ignacio Panach, Cesar Gonzalez-Perez, Oscar Pastor 0001
Data Knowl. Eng.1
2017 An Alternative Approach to Metainformation Conceptualisation and Use
Cesar Gonzalez-Perez, Patricia Martín-Rodilla
ER2
2017 A metamodel and code generation approach for symmetric unary associations
abstract
The concept of association appears in almost every modelling language, and plays a crucial role in defining how classes (or other kinds of types) can be related to each other, both in conceptual models and code. Often, associations are assumed to be binary (i.e. linking two types), and sometimes higher-arity associations are also considered, such as in UML. However, little attention has been paid to unary associations, which link a type back to itself. Some unary associations establish a symmetric relation on the instances of the type they are attached to, and in this paper we argue that mainstream modelling languages, especially UML, provide no support whatsoever to model this kind of associations, despite being extremely common in real life. To address this need, we propose a simple and powerful metamodel that describes symmetric unary associations, explain how this metamodel has been implemented as part of the ConML conceptual modelling language, and describe how this kind of associations can be implemented in code generation scenarios.
Cesar Gonzalez-Perez, Patricia Martín-Rodilla
RCIS2
2016 Using model views to assist with model conformance and extension
abstract
Most literature in conceptual modelling focuses on the development of models. However, models, once created, must be used, which requires that the model is as usable as possible. Often, usage scenarios for a model are only vaguely clear at model creation time, so model usability should be considered as a relevant problem, especially in relation to conformance (creating instance models that conform to a base type model) and extension (creating extended models that build on a base one). In this paper, we propose a particular mechanism, model views, that allow model users to customise, to certain extent, what a model looks like, and thus adapt it to their usage scenario. Model views are fully described, and two case studies regarding very different situations are reported as a form of validation. The results obtained show that model views significantly add flexibility and customisation control during model conformance and extension efforts.
Cesar Gonzalez-Perez, Patricia Martín-Rodilla
RCIS2
2015 Automatic process model discovery from textual methodologies
abstract
Process mining has been successfully used in automatic knowledge discovery and in providing guidance or support. The known process mining approaches rely on processes being executed with the help of information systems thus enabling the automatic capture of process traces as event logs. However, there are many other fields such as Humanities, Social Sciences and Medicine where workers follow processes and log their execution manually in textual forms instead. The problem we tackle in this paper is mining process instance models from unstructured, text-based process traces. Using natural language processing with a focus on the verb semantics, we created a novel unsupervised technique TextProcessMiner that discovers process instance models in two steps: 1.ActivityMiner mines the process activities; 2.ActivityRelationshipMiner mines the sequence, parallelism and mutual exclusion relationships between activities. We employed technical action research through which we validated and preliminarily evaluated our proposed technique in an Archaeology case. The results are very satisfactory with 88% correctly discovered activities in the log and a process instance model that adequately reflected the original process. Moreover, the technique we created emerged as domain independent.
Elena V. Epure, Patricia Martín-Rodilla, Charlotte Hug, Rébecca Deneckère, Camille Salinesi
RCIS2
2014 An ISO/IEC 24744-derived modelling language for discourse analysis
abstract
Most of the information that is used as input for the development of information systems is originally produced in non-structured forms, such as verbal communications or free-style text documents. Some examples are requirements specifications or documents associated to translation of contents. In cases like these, there is a need to support the structuring and extraction of the underlying semantic relations embedded in the text, in order to fully understand and post-process it. There are models for the analysis of textual information, such as information retrieval solutions or topic maps. However, these models do not offer an integral, flexible and reusable approach that can assist in giving structure and semantics to the information within a standard framework. Discourse analysis techniques, in contrast, identify semantic relations (such as authors' intentionality or possible dependencies between text elements) in the textual information according to well-established linguistic patterns. Based on these techniques, this paper presents a modelling language based on the ISO/IEC 24744 metamodel that is capable of representing pieces of textual information in a highly structured manner, describing the semantic relations in the associated discourse. In addition, the paper shows an application of the proposed language to the domain of requirements engineering, illustrating the benefits of the application of the suggested approach as well as its possibilities in other textual domains.
Patricia Martín-Rodilla, Cesar Gonzalez-Perez
RCIS1
2014 User interface design guidelines for rich applications in the context of cultural heritage data
abstract
Advanced interaction techniques are necessary to explore the potential of large data-volume systems. In this context, rich internet application patterns were defined, but usually reduced to the development of social web applications. However, other types of applications, such as data-analysis applications, require also advanced interaction solutions to assist users in making decisions and data-analysis. This paper identifies a set of problems emerged in the interaction between humans and data-analysis applications. We propose a set of guidelines for rich applications as a solution for these problems. As illustrative example of a real data-analysis environment, the paper focuses on a case study in the cultural heritage domain, highlighting the existing interaction problems and how they can be solved through the design guidelines proposed. The set of design guidelines allows to specify interfaces abstractly, creating a repository to solve interaction problems. These guidelines aim to serve as a basis for a future identification of new rich applications design patterns.
Patricia Martín-Rodilla, José Ignacio Panach, Oscar Pastor 0001
RCIS1
2012 The role of software in cultural heritage issues: Types, user needs and design guidelines based on principles of interaction
abstract
In most cases, the studied software in the cultural heritage domain has been designed from the perspective of other disciplines, such as forestry engineering, geography or documentation. In the Institute of Heritage Sciences, the cultural heritage is studied as a research topic, with methodologies to study the cultural heritage activities and considering the processing of data derived from these processes like a way to add value and knowledge in these contexts [4]. From this perspective, this paper shows a process of requirement elicitation with cultural heritage professionals and the needs identified by them. It mainly focuses on the identification of interaction human-computer (IHC) needs.
Patricia Martín-Rodilla
RCIS1