EDBT 2026 Demo / reviewers in the wild / expert
Rezvaneh Rezapour
dblp:160/1666 · also Rezvaneh (Shadi) Rezapour
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8185-4785ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit DiscussionsabstractAbstract The rise of generative AI (GenAI) has impacted many aspects of life. As these systems become embedded in everyday practices, understanding public trust in them also becomes essential for responsible adoption and governance. Prior work on trust in AI has largely drawn from psychology and human–computer interaction, but there is a lack of computational, large-scale, and longitudinal approaches to measuring trust and distrust in GenAI and large language models. We present the first computational study of Trust and Distrust in GenAI, using a multi-year Reddit dataset (2022–2025) spanning 39 subreddits and 230,576 posts. Crowd-sourced annotations of a representative sample were combined with classification models for scale analysis. Our results show that Trust and Distrust are nearly balanced over time, although Trust modestly outweighs Distrust. Technical performance and usability dominate as dimensions for both categories, while personal experience is the most frequent reason shaping attitudes. Distinct patterns also emerge across trustor groups: while industry professionals and tech leaders predominantly express Trust in GenAI, Distrust remains more prevalent among AI ethicists, journalists, and the general public. Our results provide a methodological framework for large-scale Trust analysis and insights into evolving public perceptions towards GenAI. Aria Pessianzadeh, Naima Sultana, Hilde Van den Bulck, David Gefen, Shahin Jabbari, Rezvaneh Rezapour |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | Detecting Impact Relevant Sections in Scientific ResearchabstractImpact assessment is an evolving area of research that aims at measuring and predicting the potential effects of projects or programs. Measuring the impact of scientific research is a vibrant subdomain, closely intertwined with impact assessment. A recurring obstacle pertains to the absence of an efficient framework which can facilitate the analysis of lengthy reports and text labeling. To address this issue, we propose a framework for automatically assessing the impact of scientific research projects by identifying pertinent sections in project reports that indicate the potential impacts. We leverage a mixed-method approach, combining manual annotations with supervised machine learning, to extract these passages from project reports. We experiment with different machine learning algorithms, including traditional statistical models as well as pre-trained transformer language models. Our experiments show that our proposed method achieves accuracy scores up to 0.81, and that our method is generalizable to scientific research from different domains and different languages. Maria Becker, Kanyao Han, Antonina Werthmann, Rezvaneh Rezapour, Haejin Lee, Jana Diesner |
LREC/COLING | 4 |
| 2024 | Words Matter: Reducing Stigma in Online Conversations about Substance Use with Large Language ModelsabstractStigma is a barrier to treatment for individuals struggling with substance use disorders (SUD), which leads to significantly lower treatment engagement rates.With only 7% of those affected receiving any form of help, societal stigma not only discourages individuals with SUD from seeking help but isolates them, hindering their recovery journey and perpetuating a cycle of shame and self-doubt.This study investigates how stigma manifests on social media, particularly Reddit, where anonymity can exacerbate discriminatory behaviors.We analyzed over 1.2 million posts, identifying 3,207 that exhibited stigmatizing language related to people who use substances (PWUS).Of these, 1,649 posts were classified as containing directed stigma towards PWUS, which became the focus of our de-stigmatization efforts.Using Informed and Stylized LLMs, we developed a model to transform these instances into more empathetic language.Our paper contributes to the field by proposing a computational framework for analyzing stigma and destigmatizing online content, and delving into the linguistic features that propagate stigma towards PWUS.Our work not only enhances understanding of stigma's manifestations online but also provides practical tools for fostering a more supportive environment for those affected by SUD. Layla Bouzoubaa, Elham Aghakhani, Rezvaneh Rezapour |
EMNLP | 3 |
| 2023 | Exploring the Landscape of Drug Communities on Reddit: A Network StudyabstractIn a world where the use of social media is almost universal, understanding its influence on drug information consumption has become a crucial topic of research. This study delves into the networks of drug-centric subreddits to gain insights into community structures and the nature of user interactions within Reddit. After conducting a thematic analysis of 131 drug-related subreddits, we selected four dominant ones for a comprehensive network analysis, examining the broader interactions of their active users. Our results show that drug-related subreddits are highly connected to each other, resembling small-world networks, and tangential subreddits, like r/AskReddit and r/gaming, are influential in connecting different parts of networks. While each community within the sub-networks has its unique focus, there is a clear convergence around shared themes and interests, notably harm reduction and mutual support. Our research not only underscores the pivotal role of online platforms like Reddit in bridging the active members of drug-related subreddits to a broader spectrum of content but also highlights the underlying commonalities that weave these communities together, emphasizing their mutual interests and shared values. Layla Bouzoubaa, Jordyn Young, Rezvaneh Rezapour |
ASONAM | 3 |
| 2023 | Exploring Moral Principles Exhibited in OSS: A Case Study on GitHub Heated IssuesabstractTo foster collaboration and inclusivity in Open Source Software (OSS) projects, it is crucial to understand and detect patterns of toxic language that may drive contributors away, especially those from underrepresented communities. Although machine learning-based toxicity detection tools trained on domain-specific data have shown promise, their design lacks an understanding of the unique nature and triggers of toxicity in OSS discussions, highlighting the need for further investigation. In this study, we employ Moral Foundations Theory to examine the relationship between moral principles and toxicity in OSS. Specifically, we analyze toxic communications in GitHub issue threads to identify and understand five types of moral principles exhibited in text, and explore their potential association with toxic behavior. Our preliminary findings suggest a possible link between moral principles and toxic comments in OSS communications, with each moral principle associated with at least one type of toxicity. The potential of MFT in toxicity detection warrants further investigation. Ramtin Ehsani, Rezvaneh Rezapour, Preetha Chatterjee |
ESEC/SIGSOFT FSE | 2 |
| 2023 | An expert-in-the-loop method for domain-specific document categorization based on small training dataabstractAbstract Automated text categorization methods are of broad relevance for domain experts since they free researchers and practitioners from manual labeling, save their resources (e.g., time, labor), and enrich the data with information helpful to study substantive questions. Despite a variety of newly developed categorization methods that require substantial amounts of annotated data, little is known about how to build models when (a) labeling texts with categories requires substantial domain expertise and/or in‐depth reading, (b) only a few annotated documents are available for model training, and (c) no relevant computational resources, such as pretrained models, are available. In a collaboration with environmental scientists who study the socio‐ecological impact of funded biodiversity conservation projects, we develop a method that integrates deep domain expertise with computational models to automatically categorize project reports based on a small sample of 93 annotated documents. Our results suggest that domain expertise can improve automated categorization and that the magnitude of these improvements is influenced by the experts' understanding of categories and their confidence in their annotation, as well as data sparsity and additional category characteristics such as the portion of exclusive keywords that can identify a category. Kanyao Han, Rezvaneh Rezapour, Katia Nakamura, Dikshya Devkota, Daniel C. Miller, Jana Diesner |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2022 | Information Extraction from Social Media: A Hands-on Tutorial on Tasks, Data, and Open Source ToolsabstractInformation extraction (IE) is a common sub-area of natural language processing that focuses on identifying structured data from unstructured data. One application domain of IE is Information Retrieval (IR), which relies on accurate and high-performance IE to retrieve high quality results from massive datasets. Another example of IE is to identify named entities in a text. For example, in the the sentence "Katy Perry lives in the USA", Katy Perry and USA are named entities of types of PERSON and LOCATION, respectively. Also, identify the sentiment expressed in a text is another instance of IE: in the sentence, "This movie was awesome", the expressed sentiment is positive. Finally, IE is concerned with identifying various linguistic aspects of text data, e.g., part of speech of words, noun phrases, dependency parses, etc., which can serve as features for additional IE tasks. This tutorial introduces participants to a) the usage of Python based, open-source tools that support IE from social media data (mainly Twitter), and b) best practices for ensuring the responsible use of IE and research data. Participants will learn and practice various lexical, semantic, and syntactic IE techniques that are commonly used for analyzing tweets. Participants will also be familiarized with the landscape of publicly available social media data (including popular NLP and IE benchmarks) and methods for collecting and preparing them for analysis. Furthermore, participants will be trained to use a suite of open source tools (SAIL for active learning, TwitterNER for named entity recognition, TweetNLP for transformer based NLP, and SocialMediaIE for multi task learning), which utilize advanced machine learning techniques (e.g., deep learning, active learning with human-in-the-loop, multi-lingual, and multi-task learning) to perform IE on their own or existing datasets. Participants will also learn how social contexts of text production and usage of results can be integrated into IE systems to improve these systems and to consider the role of time in improving social media IE quality. Finally, participants will learn about the governance of social media data for research purposes. The tools introduced in the tutorial will focus on the three main stages of IE, namely, collection of data (including annotation), data processing and analytics, and visualization of the extracted information. More details can be found at: https://socialmediaie.github.io/tutorials/ Shubhanshu Mishra, Rezvaneh Rezapour, Jana Diesner |
CIKM | 2 |
| 2022 | Information Extraction from Social Media: A Hands-On Tutorial on Tasks, Data, and Open Source Tools
Shubhanshu Mishra, Rezvaneh Rezapour, Jana Diesner |
ECIR (2) | 2 |
| 2022 | What Makes a Good Podcast Summary?abstractAbstractive summarization of podcasts is motivated by the growing popularity of podcasts and the needs of their listeners. Podcasting is a markedly different domain from news and other media that are commonly studied in the context of automatic summarization. As such, the qualities of a good podcast summary are yet unknown. Using a collection of podcast summaries produced by different algorithms alongside human judgments of summary quality obtained from the TREC 2020 Podcasts Track, we study the correlations between various automatic evaluation metrics and human judgments, as well as the linguistic aspects of summaries that result in strong evaluations. Rezvaneh Rezapour, Sravana Reddy, Rosie Jones, Ian Soboroff |
SIGIR | 1 |
| 2021 | Detecting Extraneous Content in PodcastsabstractSravana Reddy, Yongze Yu, Aasish Pappu, Aswin Sivaraman, Rezvaneh Rezapour, Rosie Jones. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Sravana Reddy, Aasish Pappu, Aswin Sivaraman, Rezvaneh Rezapour, Rosie Jones |
EACL | 5 |
| 2020 | 100, 000 Podcasts: A Spoken English Document CorpusabstractAnn Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Ann Clifton, Sravana Reddy, Aasish Pappu, Rezvaneh Rezapour, Hamed R. Bonab, Maria Eskevich, Gareth J. F. Jones, Jussi Karlgren, Ben Carterette, Rosie Jones |
COLING | 5 |
| 2020 | Beyond Citations: Corpus-based Methods for Detecting the Impact of Research Outcomes on SocietyabstractThis paper proposes, implements and evaluates a novel, corpus-based approach for identifying categories indicative of the impact of research via a deductive (top-down, from theory to data) and an inductive (bottom-up, from data to theory) approach. The resulting categorization schemes differ in substance. Research outcomes are typically assessed by using bibliometric methods, such as citation counts and patterns, or alternative metrics, such as references to research in the media. Shortcomings with these methods are their inability to identify impact of research beyond academia (bibliometrics) and considering text-based impact indicators beyond those that capture attention (altmetrics). We address these limitations by leveraging a mixed-methods approach for eliciting impact categories from experts, project personnel (deductive) and texts (inductive). Using these categories, we label a corpus of project reports per category schema, and apply supervised machine learning to infer these categories from project reports. The classification results show that we can predict deductively and inductively derived impact categories with 76.39% and 78.81% accuracy (F1-score), respectively. Our approach can complement solutions from bibliometrics and scientometrics for assessing the impact of research and studying the scope and types of advancements transferred from academia to society. Rezvaneh Rezapour, Jutta Bopp, Norman Fiedler, Diana Steffen, Andreas Witt, Jana Diesner |
LREC | 1 |
| 2017 | Classification and Detection of Micro-Level Impact of Issue-Focused Documentary Films based on ReviewsabstractWe present novel research at the intersection of review mining and impact assessment of issue-focused information products, namely documentary films. We develop and evaluate a theoretically grounded classification schema, related codebook, corpus annotation, and prediction model for detecting multiple types of impact that documentaries can have on individuals, such as change versus reaffirmation of behavior, cognition, and emotions, based on user-generated content, i.e., reviews. This work broadens the scope of review mining tasks, which typically comprise the prediction of ratings, helpfulness, and opinions. Our results suggest that documentaries can change or reinforce peoples' conception of an issue. We perform supervised learning to predict impact on the sentence level by using data driven as well as predefined linguistic, lexical, and psychological features; achieving an accuracy rate of 81% (F1) when using a Random Forest classifier, and 73% with a Support Vector Machine. Rezvaneh Rezapour, Jana Diesner |
CSCW | 1 |