VLDB 2026 Research / reviewers in the wild / expert
Indira Sen
dblp:219/5568
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-3475-0371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language ModelsabstractPrompt-based language models like GPT4 and LLaMa have been used for a wide variety of use cases such as simulating agents, searching for information, or for content analysis.For all of these applications and others, political biases in these models can affect their performance.Several researchers have attempted to study political bias in language models using evaluation suites based on surveys, such as the Political Compass Test (PCT), often finding a particular leaning favored by these models.However, there is some variation in the exact prompting techniques, leading to diverging findings, and most research relies on constrained-answer settings to extract model responses.Moreover, the Political Compass Test is not a scientifically valid survey instrument.In this work, we contribute a political bias measured informed by political science theory, building on survey design principles to test a wide variety of input prompts, while taking into account prompt sensitivity.We then prompt 11 different open and commercial models, differentiating between instruction-tuned and non-instructiontuned models, and automatically classify their political stances from 88,110 responses.Leveraging this dataset, we compute political bias profiles across different prompt variations and find that while PCT exaggerates bias in certain models like GPT3.5, measures of political bias are often unstable, but generally more leftleaning for instruction-tuned models.Code and data are available on GitHub 1 . Mats Faulborn, Indira Sen, Max Pellert, Andreas Spitz, David García 0001 |
ACL (1) | 2 |
| 2024 | An Open Multilingual System for Scoring Readability of WikipediaabstractWith over 60M articles, Wikipedia has become the largest platform for open and freely accessible knowledge.While it has more than 15B monthly visits, its content is believed to be inaccessible to many readers due to the lack of readability of its text.However, previous investigations of the readability of Wikipedia have been restricted to English only, and there are currently no systems supporting the automatic readability assessment of the 300+ languages in Wikipedia.To bridge this gap, we develop a multilingual model to score the readability of Wikipedia articles.To train and evaluate this model, we create a novel multilingual dataset spanning 14 languages, by matching articles from Wikipedia to simplified Wikipedia and online children encyclopedias.We show that our model performs well in a zero-shot scenario, yielding a ranking accuracy of more than 80% across 14 languages and improving upon previous benchmarks.These results demonstrate the applicability of the model at scale for languages in which there is no ground-truth data available for model fine-tuning.Furthermore, we provide the first overview on the state of readability in Wikipedia beyond English. Mykola Trokhymovych, Indira Sen, Martin Gerlach |
ACL (1) | 2 |
| 2023 | A Multidisciplinary Lens of Bias in Hate SpeechabstractHate speech detection systems may exhibit discriminatory behaviours. Research in this field has focused primarily on issues of discrimination toward the language use of minoritised communities and non-White aligned English. The interrelated issues of bias, model robustness, and disproportionate harms are weakly addressed by recent evaluation approaches, which capture them only implicitly. In this paper, we recruit a multidisciplinary group of experts to bring closer this divide between fairness and trustworthy model evaluation. Specifically, we encourage the experts to discuss not only the technical, but the social, ethical, and legal aspects of this timely issue. The discussion sheds light on critical bias facets that require careful considerations when deploying hate speech detection systems in society. Crucially, they bring clarity to different approaches for assessing, becoming aware of bias from a broader perspective, and offer valuable recommendations for future research in this field. Paula Reyero Lobo, Joseph Kwarteng, Mayra Russo, Miriam Fahimi, Kristen M. Scott, Antonio Ferrara 0003, Indira Sen, Miriam Fernández |
ASONAM | 7 |
| 2023 | People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language DetectionabstractNLP models are used in a variety of critical social computing tasks, such as detecting sexist, racist, or otherwise hateful content.Therefore, it is imperative that these models are robust to spurious features.Past work has attempted to tackle such spurious features using training data augmentation, including Counterfactually Augmented Data (CADs).CADs introduce minimal changes to existing training data points and flip their labels; training on them may reduce model dependency on spurious features.However, manually generating CADs can be time-consuming and expensive.Hence in this work, we assess if this task can be automated using generative NLP models.We automatically generate CADs using Polyjuice, Chat-GPT, and Flan-T5, and evaluate their usefulness in improving model robustness compared to manually-generated CADs.By testing both model performance on multiple out-of-domain test sets and individual data point efficacy, our results show that while manual CADs are still the most effective, CADs generated by Chat-GPT come a close second.One key reason for the lower performance of automated methods is that the changes they introduce are often insufficient to flip the original label. 1Warning: This paper has instances of hateful and sexist language to serve as examples. Indira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein, Wil M. P. van der Aalst, Claudia Wagner 0001 |
EMNLP | 1 |
| 2022 | Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech DetectionabstractIndira Sen, Mattia Samory, Claudia Wagner, Isabelle Augenstein. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Indira Sen, Mattia Samory, Claudia Wagner 0001, Isabelle Augenstein |
NAACL-HLT | 1 |
| 2022 | Depression at Work: Exploring Depression in Major US Companies from Online ReviewsabstractStudies on depression in the workplace have mostly investigated its impact on individual employees. Little is known about its association with the company as a whole, or the state where the company is based. This is due to the lack of scalable methodologies operationalizing depression in the specific context of the workplace, and of data documenting potential distress. In this work, we adapted a work-related depression scale called Occupational Depression Inventory (ODI), gathered more than 350K employee reviews of 104 major companies across the whole US for the (2008-2020) years, and developed a deep-learning framework (called AutoODI) scoring these reviews on a composite ODI score. Presence of ODI mentions manifested itself not only at micro-level (companies scoring high in ODI suffered from low stock growth) but also at macro-level (states hosting these companies were associated with high depression rates, talent shortage, and economic deprivation). This new way of applying AutoODI onto company reviews offers both theoretical implications for the literature in computational social science, occupational health and economic geography, and practical implications for companies and policy makers. Indira Sen, Daniele Quercia, Marios Constantinides, Matteo Montecchi, Licia Capra, Sanja Scepanovic, Renzo Bianchi |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?abstractAs NLP models are increasingly deployed in socially situated settings such as online abusive content detection, it is crucial to ensure that these models are robust.One way of improving model robustness is to generate counterfactually augmented data (CAD) for training models that can better learn to distinguish between core features and data artifacts.While models trained on this type of data have shown promising out-of-domain generalizability, it is still unclear what the sources of such improvements are.We investigate the benefits of CAD for social NLP models by focusing on three social computing constructs -sentiment, sexism, and hate speech.Assessing the performance of models trained with and without CAD across different types of datasets, we find that while models trained on CAD show lower in-domain performance, they generalize better out-of-domain.We unpack this apparent discrepancy using machine explanations and find that CAD reduces model reliance on spurious features.Leveraging a novel typology of CAD to analyze their relationship with model performance, we find that CAD which acts on the construct directly or a diverse set of CAD leads to higher performance. Indira Sen, Mattia Samory, Fabian Flöck, Claudia Wagner 0001, Isabelle Augenstein |
EMNLP (1) | 1 |
| 2021 | "Call me sexist, but..." : Revisiting Sexism Detection Using Psychological Scales and Adversarial Samples
Mattia Samory, Indira Sen, Julian Kohne, Fabian Flöck, Claudia Wagner 0001 |
ICWSM | 2 |
| 2020 | On the Reliability and Validity of Detecting Approval of Political Actors in TweetsabstractSocial media sites like Twitter possess the potential to complement surveys that measure political opinions and, more specifically, political actors' approval.However, new challenges related to the reliability and validity of social-media-based estimates arise.Various sentiment analysis and stance detection methods have been developed and used in previous research to measure users' political opinions based on their content on social media.In this work, we attempt to gauge the efficacy of untargeted sentiment, targeted sentiment, and stance detection methods in labeling various political actors' approval by benchmarking them across several datasets.We also contrast the performance of these pretrained methods that can be used in an off-the-shelf (OTS) manner against a set of models trained on minimal custom data.We find that OTS methods have low generalizability on unseen and familiar targets, while low-resource custom models are more robust.Our work sheds light on the strengths and limitations of existing methods proposed for understanding politicians' approval from tweets. Indira Sen, Fabian Flöck, Claudia Wagner 0001 |
EMNLP (1) | 1 |