VLDB 2026 Research / reviewers in the wild / expert
Kathleen C. Fraser
dblp:140/2850
· DBLP profile ↗
17ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0002-0752-6705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncovering Bias in Large Vision-Language Models at Scale with CounterfactualsabstractPhillip Howard, Kathleen C. Fraser, Anahita Bhiwandiwalla, Svetlana Kiritchenko. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Phillip Howard, Kathleen C. Fraser, Anahita Bhiwandiwalla, Svetlana Kiritchenko |
NAACL (Long Papers) | 2 |
| 2025 | Detecting AI-Generated Text: Factors Influencing Detectability with Current MethodsabstractLarge language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize stateof-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how “detectable” AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge. Kathleen C. Fraser, Hillary Dawkins, Svetlana Kiritchenko |
J. Artif. Intell. Res. | 1 |
| 2024 | Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-StereotypesabstractGender stereotypes are pervasive beliefs about individuals based on their gender that play a significant role in shaping societal attitudes, behaviours, and even opportunities. Recognizing the negative implications of gender stereotypes, particularly in online communications, this study investigates eleven strategies to automatically counteract and challenge these views. We present AI-generated gender-based counter-stereotypes to (self-identified) male and female study participants and ask them to assess their offensiveness, plausibility, and potential effectiveness. The strategies of counter-facts and broadening universals (i.e., stating that anyone can have a trait regardless of group membership) emerged as the most robust approaches, while humour, perspective-taking, counter-examples, and empathy for the speaker were perceived as less effective. Also, the differences in ratings were more pronounced for stereotypes about the different targets than between the genders of the raters. Alarmingly, many AI-generated counter-stereotypes were perceived as offensive and/or implausible. Our analysis and the collected dataset offer foundational insight into counter-stereotype generation, guiding future efforts to develop strategies that effectively challenge gender stereotypes in online interactions. Isar Nejadgholi, Kathleen C. Fraser, Anna Kerkhof, Svetlana Kiritchenko |
LREC/COLING | 2 |
| 2024 | Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel ImagesabstractFollowing on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged.Such models can incorporate images as input in addition to text, and perform tasks such as visual question answering, image captioning, story generation, etc.Here, we examine potential gender and racial biases in such systems, based on the perceived characteristics of the people in the input images.To accomplish this, we present a new dataset PAIRS (PArallel Images for eveRyday Scenarios).The PAIRS dataset contains sets of AI-generated images of people, such that the images are highly similar in terms of background and visual content, but differ along the dimensions of gender (man, woman) and race (Black, white).By querying the LVLMs with such images, we observe significant differences in the responses according to the perceived gender or race of the person depicted. Kathleen C. Fraser, Svetlana Kiritchenko |
EACL (1) | 1 |
| 2024 | Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender DiscourseabstractThis work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language.We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsing views that reflect sexist assumptions.With both human and automatic evaluation, we show that all eight models produce comprehensible and contextually relevant text, which is helpful in understanding diverse views on how sexism is perceived.Also, through analysis of moral foundations cited by LLMs in their arguments, we uncover the diverse ideological perspectives in models' outputs, with some models aligning more with progressive or conservative views on gender roles and sexism.Based on our observations, we caution against the potential misuse of LLMs to justify sexist language.We also highlight that LLMs can serve as tools for understanding the roots of sexist beliefs and designing well-informed interventions.Given this dual capacity, it is crucial to monitor LLMs and design safety mechanisms for their use in applications that involve sensitive societal topics, such as sexism.Warning: This paper includes examples that might be offensive and upsetting. Rongchen Guo, Isar Nejadgholi, Hillary Dawkins, Kathleen C. Fraser, Svetlana Kiritchenko |
EMNLP | 4 |
| 2023 | Diversity is Not a One-Way Street: Pilot Study on Ethical Interventions for Racial Bias in Text-to-Image Systems
Kathleen C. Fraser, Svetlana Kiritchenko, Isar Nejadgholi |
ICCC | 1 |
| 2022 | Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation VectorsabstractRobustness of machine learning models on ever-changing real-world data is critical, especially for applications affecting human wellbeing such as content moderation.New kinds of abusive language continually emerge in online discussions in response to current events (e.g., COVID-19), and the deployed abuse detection systems should be updated regularly to remain accurate.In this paper, we show that general abusive language classifiers tend to be fairly reliable in detecting out-of-domain explicitly abusive utterances but fail to detect new types of more subtle, implicit abuse.Next, we propose an interpretability technique, based on the Testing Concept Activation Vector (TCAV) method from computer vision, to quantify the sensitivity of a trained model to the humandefined concepts of explicit and implicit abusive language, and use that to explain the generalizability of the model on new data, in this case, COVID-related anti-Asian hate speech.Extending this technique, we introduce a novel metric, Degree of Explicitness, for a single instance and show that the new metric is beneficial in suggesting out-of-domain unlabeled examples to effectively enrich the training data with informative, implicitly abusive texts. Isar Nejadgholi, Kathleen C. Fraser, Svetlana Kiritchenko |
ACL (1) | 2 |
| 2022 | Extracting Age-Related Stereotypes from Social Media TextsabstractAge-related stereotypes are pervasive in our society, and yet have been under-studied in the NLP community. Here, we present a method for extracting age-related stereotypes from Twitter data, generating a corpus of 300,000 over-generalizations about four contemporary generations (baby boomers, generation X, millennials, and generation Z), as well as “old” and “young” people more generally. By employing word-association metrics, semi-supervised topic modelling, and density-based clustering, we uncover many common stereotypes as reported in the media and in the psychological literature, as well as some more novel findings. We also observe trends consistent with the existing literature, namely that definitions of “young” and “old” age appear to be context-dependent, stereotypes for different generations vary across different topics (e.g., work versus family life), and some age-based stereotypes are distinct from generational stereotypes. The method easily extends to other social group labels, and therefore can be used in future work to study stereotypes of different social categories. By better understanding how stereotypes are formed and spread, and by tracking emerging stereotypes, we hope to eventually develop mitigating measures against such biased statements. Kathleen C. Fraser, Svetlana Kiritchenko, Isar Nejadgholi |
LREC | 1 |
| 2022 | Necessity and Sufficiency for Explaining Text Classifiers: A Case Study in Hate Speech DetectionabstractEsma Balkir, Isar Nejadgholi, Kathleen Fraser, Svetlana Kiritchenko. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Esma Balkir, Isar Nejadgholi, Kathleen C. Fraser, Svetlana Kiritchenko |
NAACL-HLT | 3 |
| 2021 | Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content ModelabstractKathleen C. Fraser, Isar Nejadgholi, Svetlana Kiritchenko. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kathleen C. Fraser, Isar Nejadgholi, Svetlana Kiritchenko |
ACL/IJCNLP (1) | 1 |
| 2021 | Confronting Abusive Language Online: A Survey from the Ethical and Human Rights PerspectiveabstractThe pervasiveness of abusive content on the internet can lead to severe psychological and physical harm. Significant effort in Natural Language Processing (NLP) research has been devoted to addressing this problem through abusive content detection and related sub-areas, such as the detection of hate speech, toxicity, cyberbullying, etc. Although current technologies achieve high classification performance in research studies, it has been observed that the real-life application of this technology can cause unintended harms, such as the silencing of under-represented groups. We review a large body of NLP research on automatic abuse detection with a new focus on ethical challenges, organized around eight established ethical principles: privacy, accountability, safety and security, transparency and explainability, fairness and non-discrimination, human control of technology, professional responsibility, and promotion of human values. In many cases, these principles relate not only to situational ethical codes, which may be context-dependent, but are in fact connected to universal human rights, such as the right to privacy, freedom from discrimination, and freedom of expression. We highlight the need to examine the broad social impacts of this technology, and to bring ethical and human rights considerations to every stage of the application life-cycle, from task formulation and dataset design, to model training and evaluation, to application deployment. Guided by these principles, we identify several opportunities for rights-respecting, socio-technical solutions to detect and confront online abuse, including ‘nudging’, ‘quarantining’, value sensitive design, counter-narratives, style transfer, and AI-driven public education applications.evaluation, to application deployment. Guided by these principles, we identify several opportunities for rights-respecting, socio-technical solutions to detect and confront online abuse, including 'nudging', 'quarantining', value sensitive design, counter-narratives, style transfer, and AI-driven public education applications. Svetlana Kiritchenko, Isar Nejadgholi, Kathleen C. Fraser |
J. Artif. Intell. Res. | 3 |
| 2019 | Multilingual word embeddings for the assessment of narrative speech in mild cognitive impairmentabstractWe analyze the information content of narrative speech samples from individuals with mild cognitive impairment (MCI), in both English and Swedish, using a combination of supervised and unsupervised learning techniques. We extract information units using topic models trained on word embeddings in monolingual and multilingual spaces, and find that the multilingual approach leads to significantly better classification accuracies than training on the target language alone. In many cases, we find that augmenting the topic model training corpus with additional clinical data from a different language is more effective than training on additional monolingual data from healthy controls. Ultimately we are able to distinguish MCI speakers from healthy older adults with accuracies of up to 63% (English) and 72% (Swedish) on the basis of information content alone. We also compare our method against previous results measuring information content in Alzheimer’s disease, and report an improvement over other topic-modeling approaches. Furthermore, our results support the hypothesis that subtle differences in language can be detected in narrative speech, even at the very early stages of cognitive decline, when scores on screening tools such as the Mini-Mental State Exam are still in the “normal” range. Kathleen C. Fraser, Kristina Lundholm Fors, Dimitrios Kokkinakis |
Comput. Speech Lang. | 1 |
| 2018 | A Swedish Cookie-Theft Corpus
Dimitrios Kokkinakis, Kristina Lundholm Fors, Kathleen C. Fraser, Arto Nordlund |
LREC | 3 |
| 2017 | An analysis of eye-movements during reading for the detection of mild cognitive impairmentabstractWe present a machine learning analysis of eye-tracking data for the detection of mild cognitive impairment, a decline in cognitive abilities that is associated with an increased risk of developing dementia.We compare two experimental configurations (reading aloud versus reading silently), as well as two methods of combining information from the two trials (concatenation and merging).Additionally, we annotate the words being read with information about their frequency and syntactic category, and use these annotations to generate new features.Ultimately, we are able to distinguish between participants with and without cognitive impairment with up to 86% accuracy. Kathleen C. Fraser, Kristina Lundholm Fors, Dimitrios Kokkinakis, Arto Nordlund |
EMNLP | 1 |
| 2016 | Speech Recognition in Alzheimer's Disease and in its Assessment
Luke Zhou, Kathleen C. Fraser, Frank Rudzicz |
INTERSPEECH | 2 |
| 2015 | Sentence segmentation of aphasic speechabstractKathleen C. Fraser, Naama Ben-David, Graeme Hirst, Naida Graham, Elizabeth Rochon. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Kathleen C. Fraser, Naama Ben-David, Graeme Hirst, Naida L. Graham, Elizabeth Rochon |
HLT-NAACL | 1 |
| 2013 | Using text and acoustic features to diagnose progressive aphasia and its subtypesabstractThis paper presents experiments in automatically diagnosing primary progressive aphasia (PPA) and two of its subtypes, semantic dementia (SD) and progressive nonfluent aphasia (PNFA), from the acoustics of recorded narratives and textual analysis of the resultant transcripts. In order to train each of three types of classifier (naive Bayes, support vector machine, random forest), a large set of 81 available features must be reduced in size. Two methods of feature selection are therefore compared – one based on statistical significance and the other based on minimum-redundancy-maximum-relevance. After classifier optimization, PPA (or absence thereof) is correctly diagnosed across 87.4% of conditions, and the two subtypes of PPA are correctly classified 75.6% of the time. Kathleen C. Fraser, Frank Rudzicz, Elizabeth Rochon |
INTERSPEECH | 1 |