Miriam Fernández

dblp:54/2749 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-5939-4321ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 22 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Prescriptive analytics motivating distance learning students to take remedial action: A case study of a student-facing dashboard
Christothea Herodotou, Jessica Carr, Sagun Shrestha, Catherine Comfort, Vaclav Bayer, Claire Maguire, Paul Mulholland, Miriam Fernández
LAK9
2025 Co-creating an Ontology of Online Gender-Based Harms: An Interdisciplinary Perspective
Miriam Fernández, Alba Catalina Morales Tirado, Ángel Pavón Pérez, Keely Duddin, Min Zhang 0027, Ksenia Bakina, Arosha K. Bandara, Rose Capdevila, Lisa Lazard, Olga Jurasz
ISWC (2)1
2024 Enhancing Hate Speech Annotations with Background Semantics
abstract
Most automated hate speech detection models rely on human annotations for training and evaluation. Logic and research indicate that people who belong to groups targeted by hate speech are better at identifying it, often due to their increased familiarity with the topic and associated hate speech terminology. However, most hate speech annotation practices overlook this issue, and hence the labels produced tend to have a reduced accuracy. In this paper, we describe an approach where the text to be annotated is supplemented with background semantics, to expose the meaning of hate speech terminology that is less likely to be known to general annotators. We test the impact of this approach by measuring change in inter-annotator agreement, before and after introducing semantics, between two groups of annotators; those who belong to the target group of hate speech, and those who are not. Our experiments show that infusing text with semantic background increases inter-annotator agreement by up to 11.3% on average, aligning the annotations from annotators who do not belong to the target groups with those from the target groups.
Paula Reyero Lobo, Enrico Daga, Harith Alani, Miriam Fernández
ECAI4
2023 Annotators' Perspectives: Exploring the Influence of Identity on Interpreting Misogynoir
abstract
Social Networking Sites are home to different forms of hate, including "Misogynoir", which specifically targets Black women through a combination of racism and sexism. Detecting misogynoir presents challenges due to its subjective nature and the varied interpretations of hate speech. Using annotator justifications from four distinct demographic groups; including Black women, Black men, White women and White men, we seek to gain a deeper understanding of the factors that influence annotators' reasoning process and labelling decisions for potential cases of Misogynoir and Allyship. Given the unique experiences of Black women who face both racism and sexism, the study sought to understand how their intersectional identities shape their perspectives compared to other groups. The research employed a qualitative analysis of responses from participants to identify key themes and patterns. Three significant themes emerged from our in-depth qualitative analysis of these annotator justifications: prior knowledge and experience, the language of the social media post, and its context. Our results revealed that annotators historically at risk of abuse demonstrated a nuanced understanding of how their intersecting identities inform their interpretations and judgement of tweets, drawing on their personal encounters with misogyny and racism compared to their non-target counterparts of this type of hate. This study underscores the significance of diverse annotator perspectives and content comprehension in understanding and addressing hate speech, particularly when it intersects with multiple forms of discrimination. Our study contributes to the methodological advancements in social network analysis and mining, highlighting the importance of considering annotator characteristics in the development of tools and approaches for detecting and addressing intersectional hate.
Joseph Kwarteng, Tracie Farrell, Aisling Third, Miriam Fernández
ASONAM4
2023 A Multidisciplinary Lens of Bias in Hate Speech
abstract
Hate speech detection systems may exhibit discriminatory behaviours. Research in this field has focused primarily on issues of discrimination toward the language use of minoritised communities and non-White aligned English. The interrelated issues of bias, model robustness, and disproportionate harms are weakly addressed by recent evaluation approaches, which capture them only implicitly. In this paper, we recruit a multidisciplinary group of experts to bring closer this divide between fairness and trustworthy model evaluation. Specifically, we encourage the experts to discuss not only the technical, but the social, ethical, and legal aspects of this timely issue. The discussion sheds light on critical bias facets that require careful considerations when deploying hate speech detection systems in society. Crucially, they bring clarity to different approaches for assessing, becoming aware of bias from a broader perspective, and offer valuable recommendations for future research in this field.
Paula Reyero Lobo, Joseph Kwarteng, Mayra Russo, Miriam Fahimi, Kristen M. Scott, Antonio Ferrara 0003, Indira Sen, Miriam Fernández
ASONAM8
2023 CERSEI: Cognitive Effort Based Recommender System for Enhancing Inclusiveness
Geoffray Bonnin, Vaclav Bayer, Miriam Fernández, Christothea Herodotou, Martin Hlosta, Paul Mulholland
EC-TEL3
2021 Learning Analytics and Fairness: Do Existing Algorithms Serve Everyone Equally?
Vaclav Bayer, Martin Hlosta, Miriam Fernández
AIED (2)3
2021 Impact of Predictive Learning Analytics on Course Awarding Gap of Disadvantaged Students in STEM
Martin Hlosta, Christothea Herodotou, Vaclav Bayer, Miriam Fernández
AIED (2)4
2021 Misogynoir: public online response towards self-reported misogynoir
abstract
"Misogynoir" refers to the specific forms of misogyny that Black women experience, which couple racism and sexism together. To better understand the online manifestations of this type of hate, and to propose methods that can automatically identify it, in this paper, we conduct a study on 4 cases of Black women in Tech reporting experiences of misogynoir on the Twitter platform. We follow the reactions to these cases (both supportive and non-supportive responses), and categorise them within a model of misogynoir that highlights experiences of Tone Policing, White Centring, Racial Gaslighting and Defensiveness. As an intersectional form of abusive or hateful speech, we investigate the possibilities and challenges to detect online instances of misogynoir in an automated way. We then conduct a closer qualitative analysis on messages of support and non-support to look at some of these categories in more detail. The purpose of this investigation is to understand responses to misogynoir online, including doubling down on misogynoir, engaging in performative allyship, and showing solidarity with Black women in tech.
Joseph Kwarteng, Serena Coppolino Perfumi, Tracie Farrell, Miriam Fernández
ASONAM4
2020 Exploiting Citation Knowledge in Personalised Recommendation of Recent Scientific Publications
abstract
In this paper we address the problem of providing personalised recommendations of recent scientific publications to a particular user, and explore the use of citation knowledge to do so. For this purpose, we have generated a novel dataset that captures authors’ publication history and is enriched with different forms of paper citation knowledge, namely citation graphs, citation positions, citation contexts, and citation types. Through a number of empirical experiments on such dataset, we show that the exploitation of the extracted knowledge, particularly the type of citation, is a promising approach for recommending recently published papers that may not be cited yet. The dataset, which we make publicly available, also represents a valuable resource for further investigation on academic information retrieval and filtering.
Anita Khadka, Iván Cantador, Miriam Fernández
LREC3
2020 Capturing and Exploiting Citation Knowledge for Recommending Recently Published Papers
abstract
With the continuous growth of scientific literature, discovering relevant academic papers for a researcher has become a challenging task, especially when looking for the latest, most recent papers. In this case, traditional collaborative filtering systems are ineffective, since they are unable to recommend items not previously seen, rated or cited. This is known as the item cold-start problem. In this paper, we explore the potential of exploiting citation knowledge to provide a given user with relevant suggestions about recent scientific publications. A novel hybrid recommendation method that encapsulates such citation knowledge is proposed. Experimental results show improvements over baseline methods, evidencing benefits of using citation knowledge to recommend recently published papers in a personalised way. Moreover, as a result of our work, we also provide a unique dataset that, differently to previous corpora, contains detailed paper citation information.
Anita Khadka, Iván Cantador, Miriam Fernández
WETICE3
2020 Exploiting Open Data to analyze discussion and controversy in online citizen participation
Iván Cantador, María E. Cortés-Cediel, Miriam Fernández
Inf. Process. Manag.3
2018 What's going on in my city?: recommender systems and electronic participatory budgeting
abstract
In this paper, we present electronic participatory budgeting (ePB) as a novel application domain for recommender systems. On public data from the ePB platforms of three major US cities - Cambridge, Miami and New York City-, we evaluate various methods that exploit heterogeneous sources and models of user preferences to provide personalized recommendations of citizen proposals. We show that depending on characteristics of the cities and their participatory processes, particular methods are more effective than others for each city. This result, together with open issues identified in the paper, call for further research in the area.
Iván Cantador, María E. Cortés-Cediel, Miriam Fernández, Harith Alani
RecSys3
2017 A Semantic Graph-Based Approach for Radicalisation Detection on Social Media
Hassan Saif, Thomas Dickinson, Leon Kastler, Miriam Fernández, Harith Alani
ESWC (1)4
2017 HESML: A scalable ontology-based semantic similarity measures library with a set of reproducible experiments and a replication dataset
abstract
• This work is a detailed companion reproducibility paper of the methods and experiments proposed in three previous works by Lastra-Díaz and García-Serrano, which introduce a set of reproducible experiments on word similarity based on HESML and ReproZip with the aim of exactly reproducing the experimental surveys in the aforementioned works. • This work introduces a new representation model for taxonomies called PosetHERep, and a Java software library called Half-Edge Semantic Measures Library (HESML) based on it, which implements most ontology-based semantic similarity measures and Information Content (IC) models based on WordNet reported in the literature. • PosetHERep proposes a memory-efficient representation for taxonomies which linearly scales with the size of the taxonomy and provides an efficient implementation of a large set of topological queries and graph-based algorithms, which is an adaptation of the half-edge data structure commonly used to represent discrete manifolds and planar graphs in computational geometry. • This work also introduces a replication framework and dataset, called WNSimRep v1, which is provided as supplementary material and whose aim is to assist the exact replication of most similarity measures and IC models reported in the literature. • Finally, this work introduces an experimental survey on the performance and scalability of the most recent state-of-the-art semantic measures libraries. This latter experimental survey confirms the statistically significant outperformance of HESML on the state-of-the-art libraries in terms of performance and scalability, as well as the possibility to improve significantly the performance and scalability of the semantic measures libraries without caching using PosetHERep. This work is a detailed companion reproducibility paper of the methods and experiments proposed by Lastra-Díaz and García-Serrano in (2015, 2016) [56–58], which introduces the following contributions: (1) a new and efficient representation model for taxonomies, called PosetHERep , which is an adaptation of the half-edge data structure commonly used to represent discrete manifolds and planar graphs; (2) a new Java software library called the Half-Edge Semantic Measures Library ( HESML) based on PosetHERep , which implements most ontology-based semantic similarity measures and Information Content (IC) models reported in the literature; (3) a set of reproducible experiments on word similarity based on HESML and ReproZip with the aim of exactly reproducing the experimental surveys in the three aforementioned works; (4) a replication framework and dataset, called WNSimRep v1 , whose aim is to assist the exact replication of most methods reported in the literature; and finally, (5) a set of scalability and performance benchmarks for semantic measures libraries. PosetHERep and HESML are motivated by several drawbacks in the current semantic measures libraries, especially the performance and scalability, as well as the evaluation of new methods and the replication of most previous methods. The reproducible experiments introduced herein are encouraged by the lack of a set of large, self-contained and easily reproducible experiments with the aim of replicating and confirming previously reported results. Likewise, the WNSimRep v1 dataset is motivated by the discovery of several contradictory results and difficulties in reproducing previously reported methods and experiments. PosetHERep proposes a memory-efficient representation for taxonomies which linearly scales with the size of the taxonomy and provides an efficient implementation of most taxonomy-based algorithms used by the semantic measures and IC models, whilst HESML provides an open framework to aid research into the area by providing a simpler and more efficient software architecture than the current software libraries. Finally, we prove the outperformance of HESML on the state-of-the-art libraries, as well as the possibility of significantly improving their performance and scalability without caching using PosetHERep .
Juan J. Lastra-Díaz, Ana García-Serrano, Montserrat Batet, Miriam Fernández, Fernando Seabra Chirigati
Inf. Syst.4
2016 Contextual semantics for sentiment analysis of Twitter
Hassan Saif, Yulan He 0001, Miriam Fernández, Harith Alani
Inf. Process. Manag.3
2015 Identifying Prominent Life Events on Twitter
abstract
Social media is a common place for people to post and share digital reflections of their life events, including major events such as getting married, having children, graduating, etc. Although the creation of such posts is straightforward, the identification of events on online media remains a challenge. Much research in recent years focused on extracting major events from Twitter, such as earthquakes, storms, and floods. This paper however, targets the automatic detection of personal life events, focusing on five events that psychologists found to be the most prominent in people lives. We define a variety of features (user, content, semantic and interaction) to capture the characteristics of those life events and present the results of several classification methods to automatically identify these events in Twitter. Our proposed classification methods obtain results between 0.84 and 0.92 F1-measure for the different types of life events. A novel contribution of this work also lies in a new corpus of tweets, which has been annotated by using crowdsourcing and that constitutes, to the best of our knowledge, the first publicly available dataset for the automatic identification of personal life events from Twitter.
Thomas Dickinson, Miriam Fernández, Lisa Thomas 0001, Paul Mulholland, Pamela Briggs, Harith Alani
K-CAP2
2014 SentiCircles for Contextual and Conceptual Semantic Sentiment Analysis of Twitter
Hassan Saif, Miriam Fernández, Yulan He 0001, Harith Alani
ESWC2
2014 On Stopwords, Filtering and Data Sparsity for Sentiment Analysis of Twitter
Hassan Saif, Miriam Fernández, Yulan He 0001, Harith Alani
LREC2
2014 Semantic Patterns for Sentiment Analysis of Twitter
Hassan Saif, Yulan He 0001, Miriam Fernández, Harith Alani
ISWC (2)3
2013 Community analysis through semantic rules and role composition derivation
Matthew Rowe 0001, Miriam Fernández, Sofia Angeletou, Harith Alani
J. Web Semant.2
2012 Workshop on multimodal crowd sensing (CrowdSens 2012)
abstract
This paper provides an overview of the 1st International Workshop on Multimodal Crowd Sensing (CrowdSens 2012), held at the 21st ACM International Conference on Information and Knowledge Management (CIKM 2012). This workshop aimed to provide an open forum for researchers from various fields such as fields such as Natural Language Processing, Information Extraction, Data Mining, Information Retrieval, User Modeling and Personalization, Stream Processing, and Sensor Networks, for addressing the challenges of effectively mining, analyzing, fusing, and exploiting information sourced from multimodal physical and social sensor data sources.
Haggai Roitman, Iván Cantador, Miriam Fernández
CIKM3
2011 Ontology augmentation: combining semantic web and text resources
abstract
This work investigates the process of selecting, extracting and reorganizing content from Semantic Web information sources, to produce an ontology meeting the specifications of a particular domain and/or task. The process is combined with traditional text-based ontology learning methods to achieve tolerance to knowledge incompleteness. The paper describes the approach and presents experiments in which an ontology was built for a diet evaluation task. Although the example presented concerns the specific case of building a nutritional ontology, the methods employed are domain independent and transferrable to other use cases.
Miriam Fernández, Ziqi Zhang 0001, Vanessa López, Victoria S. Uren, Enrico Motta
K-CAP1
2011 Linking Data across Universities: An Integrated Video Lectures Dataset
Miriam Fernández, Mathieu d'Aquin, Enrico Motta
ISWC (2)1
2011 Semantically enhanced Information Retrieval: An ontology-based approach
Miriam Fernández, Iván Cantador, Vanessa López, David Vallet, Pablo Castells, Enrico Motta
J. Web Semant.1
2009 Evaluating Semantic Relations by Exploring Ontologies on the Semantic Web
Marta Sabou, Miriam Fernández, Enrico Motta
NLDB2
2007 Personalized Content Retrieval in Context Using Ontological Knowledge
abstract
Personalized content retrieval aims at improving the retrieval process by taking into account the particular interests of individual users. However, not all user preferences are relevant in all situations. It is well known that human preferences are complex, multiple, heterogeneous, changing, even contradictory, and should be understood in context with the user goals and tasks at hand. In this paper, we propose a method to build a dynamic representation of the semantic context of ongoing retrieval tasks, which is used to activate different subsets of user interests at runtime, in a way that out-of-context preferences are discarded. Our approach is based on an ontology-driven representation of the domain of discourse, providing enriched descriptions of the semantics involved in retrieval actions and preferences, and enabling the definition of effective means to relate preferences and context
David Vallet, Pablo Castells, Miriam Fernández, Phivos Mylonas, Yannis Avrithis
IEEE Trans. Circuits Syst. Video Technol.3
2007 An Adaptation of the Vector-Space Model for Ontology-Based Information Retrieval
abstract
Semantic search has been one of the motivations of the semantic Web since it was envisioned. We propose a model for the exploitation of ontology-based knowledge bases to improve search over large document repositories. In our view of information retrieval on the semantic Web, a search engine returns documents rather than, or in addition to, exact values in response to user queries. For this purpose, our approach includes an ontology-based scheme for the semiautomatic annotation of documents and a retrieval system. The retrieval model is based on an adaptation of the classic vector-space model, including an annotation weighting algorithm, and a ranking algorithm. Semantic search is combined with conventional keyword-based retrieval to achieve tolerance to knowledge base incompleteness. Experiments are shown where our approach is tested on corpora of significant scale, showing clear improvements with respect to keyword-based search
Pablo Castells, Miriam Fernández, David Vallet
IEEE Trans. Knowl. Data Eng.2
2006 Probabilistic Score Normalization for Rank Aggregation
Miriam Fernández, David Vallet, Pablo Castells
ECIR1
2006 Adaptive multimedia access: from user needs to semantic personalisation
abstract
We discuss the use of a reliable user requirements methodology for gathering essential data relating to user needs in advanced, personalised multimedia content applications. We revise the implications of these requirements in the development of advanced personalised searching and browsing tools aimed at assisting the end-consumer by complementing explicit user requests with implicit user preferences, to better meet individual user needs. We examine the important technical challenges such as representing conceptual-level user interests, and coping with the subtleties of user preferences, such as variability and heterogeneity, which we see as critical for the success of personalisation techniques. We consider the fact that personalisation is not always appropriate, and it is sometimes preferable to disable it to avoid obtrusiveness.
A. Evans, Miriam Fernández, David Vallet, Pablo Castells
ISCAS2
2006 Using historical data to enhance rank aggregation
abstract
Rank aggregation is a pervading operation in IR technology. We hypothesize that the performance of score-based aggregation may be affected by artificial, usually meaningless deviations consis-tently occurring in the input score distributions, which distort the combined result when the individual biases differ from each other. We propose a score-based rank aggregation model where the source scores are normalized to a common distribution before being combined. Early experiments on available data from several TREC collections are shown to support our proposal.
Miriam Fernández, David Vallet, Pablo Castells
SIGIR1
2005 An Ontology-Based Information Retrieval Model
David Vallet, Miriam Fernández, Pablo Castells
ESWC2