EDBT 2026 Demo / reviewers in the wild / expert
Scott A. Hale
dblp:32/10840 · also Scott Hale
· DBLP profile ↗
8ranked-venue papers in the field
1as first author
6since 2021 · last 2024
0000-0002-6894-4951ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 8 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | "I Am 30F and Need Advice!": A Mixed-Method Analysis of the Effects of Advice-Seekers' Self-Disclosure on Received RepliesabstractIn community question answering sites, users can easily make a post to ask questions or seek advice. Others volunteer replies to these posts to provide answers of varying quality, detail, and helpfulness. In the advice-seeking process, self-disclosure enables posters to provide a relatable context for their requests but comes at a cost of greater identifiability. We focus on the "r/Advice" Reddit community and present a mixed-method study on how self-disclosure of advice-seekers shapes the prevalence and detail of the feedback received. We focus particularly on age and gender disclosure as both are reliably detected and normatively considered in the context of giving advice. We use both hurdle negative binomial regression models and discourse analysis to examine the relationship between self-disclosure and the replies received and explore themes related to disclosure. The results show that advice-seekers' age or gender disclosure correlates with more replies and more helpful replies, but the effects of age and gender disclosure are not additive. We also find both reciprocity and homophily effects in disclosure as reply-givers are more likely to self-disclose when the advice-seeker does so. The lack of additive effects alongside the thematic analysis suggests disclosure practices are used to elicit sufficient credibility or basis for empathy, whereas too much or too little disclosure creates uncertainty or inhibits the applicability of the received advice. Scott A. Hale, Bernie Hogan |
ICWSM | 2 |
| 2024 | A Multilingual Similarity Dataset for News Article FrameabstractUnderstanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frame remains a challenge due to the absence of a concrete and unified standard dataset that considers the comprehensive nuances within news content. To address this gap, we introduce an extended version of a large labeled news article dataset with 16,687 new labeled pairs. Leveraging the pairwise comparison of news articles, our method frees the work of manual identification of frame classes in traditional news frame analysis studies. Overall we introduce the most extensive cross-lingual news article similarity dataset available to date with 26,555 labeled news article pairs across 10 languages. Each data point has been meticulously annotated according to a codebook detailing eight critical aspects of news content, under a human-in-the-loop framework. Application examples demonstrate its potential in unearthing country communities within global news coverage, exposing media bias among news outlets, and quantifying the factors related to news creation. We envision that this news similarity dataset will broaden our understanding of the media ecosystem in terms of news coverage of events and perspectives across countries, locations, languages, and other social constructs. By doing so, it can catalyze advancements in social science research and applied methodologies, thereby exerting a profound impact on our society. Xi Chen 0125, Mattia Samory, Scott A. Hale, David Jurgens, Przemyslaw A. Grabowicz |
ICWSM | 3 |
| 2024 | Misinformation Mitigation Praxis: Lessons Learned and Future Directions from Co·InsightsabstractMisinformation is a global challenge, but successful mitigations must come from the communities affected and not be imposed by external entities. Co·Insights is a multi-year NSF-funded initiative building capacity to respond to misinformation and harmful narratives in Asian American and Pacific Islander (AAPI) communities, based on a deep involvement with grassroots organizations and a co-construction of tools grounded on community needs. Co·Insights' unique cross-sectoral, cross-disciplinary collaboration is a convergence of information retrieval, computational social science, and ethnographic inquiry with a unique platform that enables community organizations, fact-checkers, and academics to work together to respond effectively to harmful content targeting communities. In this SIRIP talk designed for a technical audience, we will share lessons learned from the first 2.5 years of Co·Insights, and how we are bridging the academic--praxis divide by integrating state-of-the-art research developments into large-scale deployed systems. Topics covered will include: the challenges and joys of collaborating across disciplinary and sectoral boundaries; community-driven approaches that utilize information retrieval techniques such as claim matching in concert with emerging best practices in misinformation mitigation; open problems we have encountered in the space; and future directions we find promising. Co·Insights is led by Meedan, a global non-profit providing award-winning software solutions to mitigate misinformation; and in partnership with AAPI community organizations and several academic institutions. Scott A. Hale, Venkata Rama Kiran Garimella, Shiri Dori-Hacohen |
SIGIR | 1 |
| 2024 | SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation TasksabstractDiaspora communities are disproportionately impacted by off-the-radar misinformation and often neglected by mainstream fact-checking efforts, creating a critical need to scale-up efforts of nascent fact-checking initiatives. In this paper we present SynDy, a framework for Synthetic Dynamic Dataset Generation to leverage the capabilities of the largest frontier Large Language Models (LLMs) to train local, specialized language models. To the best of our knowledge, SynDy is the first paper utilizing LLMs to create fine-grained synthetic labels for tasks of direct relevance to misinformation mitigation, namely Claim Matching, Topical Clustering, and Claim Relationship Classification. SynDy utilizes LLMs and social media queries to automatically generate distantly-supervised, topically-focused datasets with synthetic labels on these three tasks, providing essential tools to scale up human-led fact-checking at a fraction of the cost of human-annotated data. Training on SynDy's generated labels shows improvement over a standard baseline and is not significantly worse compared to training on human labels (which may be infeasible to acquire). SynDy is being integrated into Meedan's chatbot tiplines that are used by over 50 organizations, serve over 230K users annually, and automatically distribute human-written fact-checks via messaging apps such as WhatsApp. SynDy will also be integrated into our deployed Co-Insights toolkit, enabling low-resource organizations to launch tiplines for their communities. Finally, we envision SynDy enabling additional fact-checking tools such as matching new misinformation claims to high-quality explainers on common misinformation topics. Michael Shliselberg, Ashkan Kazemi, Scott A. Hale, Shiri Dori-Hacohen |
SIGIR | 3 |
| 2024 | Global News Synchrony and Diversity During the Start of the COVID-19 PandemicabstractNews coverage profoundly affects how countries and individuals behave in international relations. Yet, we have little empirical evidence of how news coverage varies across countries. To enable studies of global news coverage, we develop an efficient computational methodology that comprises three components: (i) a transformer model to estimate multilingual news similarity; (ii) a global event identification system that clusters news based on a similarity network of news articles; and (iii) measures of news synchrony across countries and news diversity within a country, based on country-specific distributions of news coverage of the global events. Each component achieves state-of-the art performance, scaling seamlessly to massive datasets of millions of news articles. Xi Chen 0125, Scott A. Hale, David Jurgens, Mattia Samory, Ethan Zuckerman, Przemyslaw A. Grabowicz |
WWW | 2 |
| 2022 | Information Ecosystem Threats in Minoritized Communities: Challenges, Open Problems and Research DirectionsabstractJournalists, fact-checkers, academics, and community media are overwhelmed in their attempts to support communities suffering from gender-, race- and ethnicity-targeted information ecosystem threats, including but not limited to misinformation, hate speech, weaponized controversy and online-to-offline harassment. Yet, for a plethora of reasons, minoritized groups are underserved by current approaches to combat such threats. In this panel, we will present and discuss the challenges and open problems facing such communities and the researchers hoping to serve them. We will also discuss the current state-of-the-art as well as the most promising future directions, both within IR specifically, across Computer Science more broadly, as well as that requiring transdisciplinary and cross-sectoral collaborations. The panel will attract both IR practitioners and researchers and include at least one panelist outside of IR, with unique expertise in this space. Shiri Dori-Hacohen, Scott A. Hale |
SIGIR | 2 |
| 2019 | Demographic Inference and Representative Population Estimates from Multilingual Social Media DataabstractSocial media provide access to behavioural data at an unprecedented scale and granularity. However, using these data to understand phenomena in a broader population is difficult due to their non-representativeness and the bias of statistical inference tools towards dominant languages and groups. While demographic attribute inference could be used to mitigate such bias, current techniques are almost entirely monolingual and fail to work in a global environment. We address these challenges by combining multilingual demographic inference with post-stratification to create a more representative population sample. To learn demographic attributes, we create a new multimodal deep neural architecture for joint classification of age, gender, and organization-status of social media users that operates in 32 languages. This method substantially outperforms current state of the art while also reducing algorithmic bias. To correct for sampling biases, we propose fully interpretable multilevel regression methods that estimate inclusion probabilities from inferred joint population counts and ground-truth population counts. Zijian Wang 0002, Scott A. Hale, David Ifeoluwa Adelani, Przemyslaw A. Grabowicz, Timo Hartmann, Fabian Flöck, David Jurgens |
WWW | 2 |
| 2018 | Online Petitioning Through Data Exploration and What We Found There: A Dataset of Petitions from Avaaz.org
Pablo Aragón, Diego Sáez-Trumper, Miriam Redi, Scott A. Hale, Vicenç Gómez, Andreas Kaltenbrunner |
ICWSM | 4 |