VLDB 2026 Research / reviewers in the wild / expert
Zeerak Talat
dblp:305/7414 · also Zeerak Waseem
· DBLP profile ↗
23ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0001-5503-867XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media DataabstractSocial media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors.However, sharing this sensitive data poses privacy risks and limits the growth of benchmark datasets.We comprehensively evaluate whether privacy-preserving ML techniques can enable safer data sharing while preserving performance.Specifically, we apply federated learning (FL) and Differentially Private FL for two widely-studied mental health prediction tasks: depression detection on X (Twitter) and suicide crisis detection on Reddit.We simulate realistic data-sharing scenarios by treating each user as a client in a non-IID setting, evaluating across different client fractions, aggregation strategies, and privacy budgets.While FL achieves comparable performance to centralized training (centralized 𝐹 1 = 85.63; best FL model 𝐹 1 = 83.16) on depression identification, we find that Differentially Private FL has a large performance-privacy trade-off (up to 𝐹 1 = 27.01 drop) even with low levels of noise (𝜖 = 50).This is due to the distortion of highly informative yet sparse mental health linguistic markers related to mental health, like health topics and emotion words.This research empirically demonstrates the potential and limitations of current privacy preservation techniques for mental health inference tasks. Nuredin Ali Abdelkadir, Anjali Ratnam, Zeerak Talat, Stevie Chancellor |
ACL (1) | 3 |
| 2026 | Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature ReviewabstractMotivated by a growing research interest into automatic speech recognition (ASR), and the growing body of work for languages in which code-switching (CS) often occurs, we present a systematic literature review of code-switching in end-to-end ASR models. We collect and manually annotate papers published in peer reviewed venues. We document the languages considered, datasets, metrics, model choices, and performance, and present a discussion of challenges in end-to-end ASR for code-switching. Our analysis thus provides insights on current research efforts and available resources as well as opportunities and gaps to guide future research. Maha Tufail Agro, Atharva Kulkarni, Karima Kadaoui, Zeerak Talat, Hanan Aldarmaki |
LREC | 4 |
| 2025 | The Role of Expertise in Effectively Moderating Harmful Social Media Content
Nuredin Ali Abdelkadir, Tianling Yang, Shivani Kapania, Meron Estefanos, Fasica Berhane Gebrekidan, Zecharias Zelalem, Messai Ali, Rishan Berhe, Dylan K. Baker, Zeerak Talat, Milagros Miceli, Alex Hanna, Timnit Gebru |
CHI | 10 |
| 2025 | Exploring the Limitations of Detecting Machine-Generated TextabstractRecent improvements in the quality of the generations by large language models have spurred research into identifying machine-generated text. Such work often presents high-performing detectors. However, humans and machines can produce text in different styles and domains, yet the the performance impact of such on machine generated text detection systems remains unclear. In this paper, we audit the classification performance for detecting machine-generated text by evaluating on texts with varying writing styles. We find that classifiers are highly sensitive to stylistic changes and differences in text complexity, and in some cases degrade entirely to random classifiers. We further find that detection systems are particularly susceptible to misclassify easy-to-read texts while they have high performance for complex texts, leading to concerns about the reliability of detection systems. We recommend that future work attends to stylistic factors and reading difficulty levels of human-written and machine-generated text. Jad Doughman, Osama Mohammed Afzal, Hawau Olamide Toyin, Shady Shehata, Preslav Nakov, Zeerak Talat |
COLING | 6 |
| 2025 | The Only Way is Ethics: A Guide to Ethical Research with Large Language ModelsabstractThere is a significant body of work looking at the ethical considerations of large language models (LLMs): critiquing tools to measure performance and harms; proposing toolkits to aid in ideation; discussing the risks to workers; considering legislation around privacy and security etc. As yet there is no work that integrates these resources into a single practical guide that focuses on LLMs; we attempt this ambitious goal. We introduce LLM Ethics Whitepaper, which we provide as an open and living resource for NLP practitioners, and those tasked with evaluating the ethical implications of others’ work. Our goal is to translate ethics literature into concrete recommendations for computer scientists. LLM Ethics Whitepaper distils a thorough literature review into clear Do’s and Don’ts, which we present also in this paper. We likewise identify useful toolkits to support ethical work. We refer the interested reader to the full LLM Ethics Whitepaper, which provides a succinct discussion of ethical considerations at each stage in a project lifecycle, as well as citations for the hundreds of papers from which we drew our recommendations. The present paper can be thought of as a pocket guide to conducting ethical research with LLMs. Eddie L. Ungless, Nikolas Vitsakis, Zeerak Talat, James Garforth, Björn Ross, Arno Onken, Atoosa Kasirzadeh, Alexandra Birch |
COLING | 3 |
| 2025 | SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language ModelsabstractMargaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat |
NAACL (Long Papers) | 54 |
| 2024 | Classist Tools: Social Class Correlates with Performance in NLPabstractThe field of sociolinguistics has studied factors affecting language use for the last century.Labov (1964) and Bernstein (1960) showed that socioeconomic class strongly influences our accents, syntax and lexicon.However, despite growing concerns surrounding fairness and bias in Natural Language Processing (NLP), there is a dearth of studies delving into the effects it may have on NLP systems.We show empirically that NLP systems' performance is affected by speakers' SES, potentially disadvantaging less-privileged socioeconomic groups.We annotate a corpus of 95K utterances from movies with social class, ethnicity and geographical language variety and measure the performance of NLP systems on three tasks: language modelling, automatic speech recognition, and grammar error correction.We find significant performance disparities that can be attributed to socioeconomic status as well as ethnicity and geographical differences.1 With NLP technologies becoming ever more ubiquitous and quotidian, they must accommodate all language varieties to avoid disadvantaging already marginalised groups.We argue for the inclusion of socioeconomic class in future language technologies. Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Dirk Hovy |
ACL (1) | 3 |
| 2024 | Impoverished Language Technology: The Lack of (Social) Class in NLPabstractSince Labov’s foundational 1964 work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov’s initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies. Amanda Cercas Curry, Zeerak Talat, Dirk Hovy |
LREC/COLING | 2 |
| 2024 | Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment LexiconabstractFajri Koto, Tilman Beck, Zeerak Talat, Iryna Gurevych, Timothy Baldwin. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Fajri Koto, Tilman Beck, Zeerak Talat, Iryna Gurevych, Timothy Baldwin |
EACL (1) | 3 |
| 2024 | Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPabstractThis paper introduces the concept of actionability in the context of bias measures in natural language processing (NLP).We define actionability as the degree to which a measurement's results enable informed action and propose a set of desiderata for assessing it.Building on existing frameworks such as measurement modeling, we argue that actionability is a crucial aspect of bias measures that has been largely overlooked in the literature.We conduct a comprehensive review of 146 papers proposing bias measures in NLP, examining whether and how they provide the information required for actionable results.Our findings reveal that many key elements of actionability, including a measure's intended use and reliability assessment, are often unclear or absent.This study highlights a significant gap in the current approach to developing and reporting bias measures in NLP.We argue that this lack of clarity may impede the effective implementation and utilization of these measures.To address this issue, we offer recommendations for more comprehensive and actionable metric development and reporting practices in NLP bias research. Pieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett, Zeerak Talat |
EMNLP | 5 |
| 2024 | Understanding "Democratization" in NLP and ML ResearchabstractRecent improvements in natural language processing (NLP) and machine learning (ML) and increased mainstream adoption have led to researchers frequently discussing the "democratization" of artificial intelligence.In this paper, we seek to clarify how democratization is understood in NLP and ML publications, through large-scale mixed-methods analyses of papers using the keyword "democra*" published in NLP and adjacent venues.We find that democratization is most frequently used to convey (ease of) access to or use of technologies, without meaningfully engaging with theories of democratization, while research using other invocations of "democra*" tends to be grounded in theories of deliberation and debate.Based on our findings, we call for researchers to enrich their use of the term democratization with appropriate theory, towards democratic technologies beyond superficial access. 1 Arjun Subramonian, Vagrant Gautam, Dietrich Klakow, Zeerak Talat |
EMNLP | 4 |
| 2024 | AraOffence: Detecting Offensive Speech Across Dialects in Arabic Media
Youssef Nafea, Shady Shehata, Zeerak Talat, Ahmed Abo Eitta, Ahmed Sharshar, Preslav Nakov |
INTERSPEECH | 3 |
| 2024 | The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human LabelsabstractEve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Eve Fleisig, Su Lin Blodgett, Daniel Klein 0001, Zeerak Talat |
NAACL-HLT | 4 |
| 2023 | Bound by the Bounty: Collaboratively Shaping Evaluation Processes for Queer AI HarmsabstractBias evaluation benchmarks and dataset and model documentation have emerged as central processes for assessing the biases and harms of artificial intelligence (AI) systems. However, these auditing processes have been criticized for their failure to integrate the knowledge of marginalized communities and consider the power dynamics between auditors and the communities. Consequently, modes of bias evaluation have been proposed that engage impacted communities in identifying and assessing the harms of AI systems (e.g., bias bounties). Even so, asking what marginalized communities want from such auditing processes has been neglected. In this paper, we ask queer communities for their positions on, and desires from, auditing processes. To this end, we organized a participatory workshop to critique and redesign bias bounties from queer perspectives. We found that when given space, the scope of feedback from workshop participants goes far beyond what bias bounties afford, with participants questioning the ownership, incentives, and efficacy of bounties. We conclude by advocating for community ownership of bounties and complementing bounties with participatory processes (e.g., co-creation). Nathaniel Dennler, Anaelia Ovalle, Ashwin Singh, Luca Soldaini, Arjun Subramonian, Huy Tu, William Agnew, Avijit Ghosh, Kyra Yee, Irene Font Peradejordi, Zeerak Talat, Mayra Russo, Jessica de Jesus de Pinho Pinhal |
AIES | 11 |
| 2023 | A Federated Approach for Hate Speech DetectionabstractHate speech detection has been the subject of high research attention, due to the scale of content created on social media.In spite of the attention and the sensitive nature of the task, privacy preservation in hate speech detection has remained under-studied.The majority of research has focused on centralised machine learning infrastructures which risk leaking data.In this paper, we show that using federated machine learning can help address privacy the concerns that are inherent to hate speech detection while obtaining up to 6.81% improvement in terms of F1-score. Jay Gala, Deep Gandhi, Jash Mehta, Zeerak Talat |
EACL | 4 |
| 2023 | Mirages. On Anthropomorphism in Dialogue SystemsabstractAutomated dialogue or conversational systems are anthropomorphised by developers and personified by users.While a degree of anthropomorphism may be inevitable due to the choice of medium, conscious and unconscious design choices can guide users to personify such systems to varying degrees.Encouraging users to relate to automated systems as if they were human can lead to high risk scenarios caused by over-reliance on their outputs.As a result, natural language processing researchers have investigated the factors that induce personification and develop resources to mitigate such effects.However, these efforts are fragmented, and many aspects of anthropomorphism have yet to be explored.In this paper, we discuss the linguistic factors that contribute to the anthropomorphism of dialogue systems and the harms that can arise, including reinforcing gender stereotypes and notions of acceptable language.We recommend that future efforts towards developing dialogue systems take particular care in their design, development, release, and description; and attend to the many linguistic cues that can elicit personification by users. Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, Zeerak Talat |
EMNLP | 5 |
| 2022 | Directions for NLP Practices Applied to Online Hate Speech DetectionabstractAddressing hate speech in online spaces has been conceptualized as a classification task that uses Natural Language Processing (NLP) techniques.Through this conceptualization, the hate speech detection task has relied on common conventions and practices from NLP.For instance, inter-annotator agreement is conceptualized as a way to measure dataset quality and certain metrics and benchmarks are used to assure model generalization.However, hate speech is a deeply complex and situated concept that eludes such static and disembodied practices.In this position paper, we critically reflect on these methodologies for hate speech detection, we argue that many conventions in NLP are poorly suited for the problem and encourage researchers to develop methods that are more appropriate for the task. Paula Fortuna, Mónica Domínguez, Leo Wanner, Zeerak Talat |
EMNLP | 4 |
| 2022 | A Federated Approach to Predicting Emojis in Hindi TweetsabstractThe use of emojis affords a visual modality to, often private, textual communication.The task of predicting emojis however provides a challenge for machine learning as emoji use tends to cluster into the frequently used and the rarely used emojis.Much of the machine learning research on emoji use has focused on high resource languages and has conceptualised the task of predicting emojis around traditional server-side machine learning approaches.However, traditional machine learning approaches for private communication can introduce privacy concerns, as these approaches require all data to be transmitted to a central storage.In this paper, we seek to address the dual concerns of emphasising high resource languages for emoji prediction and risking the privacy of people's data.We introduce a new dataset of 118k tweets (augmented from 25k unique tweets) for emoji prediction in Hindi, 1 and propose a modification to the federated learning algorithm, CausalFedGSD, which aims to strike a balance between model performance and user privacy.We show that our approach obtains comparative scores with more complex centralised models while reducing the amount of data required to optimise the models and minimising risks to user privacy. Deep Gandhi, Jash Mehta, Nirali Parekh, Karan Waghela, Lynette D'Mello, Zeerak Talat |
EMNLP | 6 |
| 2022 | On the Machine Learning of Ethical Judgments from Natural LanguageabstractZeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, Adina Williams. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, Adina Williams |
NAACL-HLT | 1 |
| 2021 | A Survey of Race, Racism, and Anti-Racism in NLPabstractAnjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia Tsvetkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Anjalie Field, Su Lin Blodgett, Zeerak Talat, Yulia Tsvetkov |
ACL/IJCNLP (1) | 3 |
| 2021 | HateCheck: Functional Tests for Hate Speech Detection ModelsabstractDetecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses. Paul Röttger, Bertie Vidgen, Dong Nguyen 0002, Zeerak Talat, Helen Z. Margetts, Janet B. Pierrehumbert |
ACL/IJCNLP (1) | 4 |
| 2021 | Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionabstractBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe Kiela. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bertie Vidgen, Tristan Thrush, Zeerak Talat, Douwe Kiela |
ACL/IJCNLP (1) | 3 |
| 2021 | Dynabench: Rethinking Benchmarking in NLPabstractDouwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel 0001, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams |
NAACL-HLT | 14 |