VLDB 2026 Research / reviewers in the wild / expert
Libby Hemphill
dblp:62/2987
· DBLP profile ↗
14ranked-venue papers in the field
4as first author
11since 2021 · last 2026
0000-0002-3793-7281ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 13 (4 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building a Test Collection for Social Science Dataset RetrievalabstractData reuse expedites scientific progress and conserves resources, yet connecting researchers with datasets available for reuse remains a challenge. The scientific community has proposed several recommendation systems to help identify relevant data and maximize reuse. However, test collections or benchmarks for evaluating the performance of dataset recommendation systems are rare, particularly for social science dataset retrieval. To address this gap, we created a novel test collection for evaluating social science dataset recommendation systems. Our collection includes 249,102 query–dataset pairs with relevance judgments, featuring 262 unique search queries and 10,749 datasets. We describe how we created this collection using datasets archived at the Inter-university Consortium for Political and Social Research (ICPSR) and the ICPSR Bibliography of Data-related Literature. Additionally, we demonstrate a potential use case by evaluating the performance of embedding-based recommendation models on our test collection. The test collection is available through ICPSR at https://doi.org/10.3886/E238682V1. Ji Eun Kim, Sara Lafia, Libby Hemphill |
CHIIR | 3 |
| 2025 | What's in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text AnnotationsabstractManually annotating data for computational social science tasks can be costly, time-consuming, and emotionally draining. While recent work suggests that LLMs can perform such annotation tasks in zero-shot settings, little is known about how prompt design impacts LLMs' compliance and accuracy. We conduct a large-scale multi-prompt experiment to test how model selection (GPT-4o, GPT-3.5, PaLM2, and Falcon7b) and prompt design features (definition inclusion, output type, explanation, and prompt length) impact the compliance and accuracy of LLM-generated annotations on four highly relevant and diverse CSS tasks (toxicity, sentiment, rumor stance, and news frames). Our results show that LLM compliance and accuracy are prompt-dependent. For instance, prompting for numerical scores instead of labels reduces all LLMs' compliance and accuracy. Concise prompts can significantly reduce prompting costs but also lead to lower accuracy on tasks like toxicity. Furthermore, minor prompt changes like asking for an explanation can cause large changes in the distribution of LLM-generated labels. By assessing the impact of prompt design on the quality and distribution of LLM-generated annotations, this work serves as both a practical guide and a warning for using LLMs in CSS research. Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, Libby Hemphill |
ICWSM | 5 |
| 2025 | Data, not documents: Moving beyond theories of information-seeking behavior to advance data discoveryabstractAbstract Many theories of human information behavior (HIB) assume that information objects are in text document format. This paper argues four important HIB theories are insufficient for describing users' search strategies for data because of assumptions about the attributes of objects that users seek. We first review and compare four HIB theories: Bates' berrypicking , Marchionni's electronic information search , Dervin's sense‐making , and Meho and Tibbo's social scientist information‐seeking . All four theories assume that information‐seekers search for text documents. Next, we compare these theories to search behavior by analyzing Google Analytics data from the Inter‐university Consortium for Political and Social Research (ICPSR). Users took direct, scenic, and orienting paths when searching for data. We also interviewed ICPSR users ( n = 20), and they said they needed dataset documentation and contextual information to find data. However, Dervin's sense‐making alone cannot explain the information‐seeking behaviors that we observed. Instead, what mattered most were object attributes determined by the type of information that users sought (i.e., data, not documents). We conclude by suggesting an alternative frame for building user‐centered data discovery tools. Anthony J. Million, Jeremy York, Sara Lafia, Libby Hemphill |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2025 | Valuing curation infrastructuresabstractAbstract This study uses a theoretical lens of infrastructural dimensions to examine stakeholders' perceptions of the value of curation, focusing on the social science data repository, the Inter‐university Consortium for Political and Social Research (ICPSR). Drawing on 67 interviews with both internal (ICPSR staff) and external (funders, data producers, and reusers) stakeholders, we analyze how value is ascribed to curation across technical, organizational, and social components of infrastructure. We identify five key ways interviewees conceptualized the value of curation infrastructures: supporting sustainability and durability, enabling research efficiency, fostering trust, building community, and advancing data equity. Our findings highlight the role of curation in knowledge generation by reframing curation as infrastructure rather than a set of discrete practices. We clarify how transparency operates in dual—and sometimes conflicting—ways: as both understandability and invisibility, shaping trust in and access to data repositories. Second, we demonstrate how data equity is increasingly perceived by stakeholders as a core infrastructural value, enacted through practices that lower barriers to access. Finally, we surface the persistent challenges in evaluating and funding curation infrastructures due to their long time horizons and often‐invisible nature. This work advocates recognizing and funding curation infrastructures as essential for long‐term scientific and societal progress. Morgan F. Wofford, Andrea K. Thomer, Libby Hemphill, Katherine Polasek, Elizabeth Yakel |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2024 | Landscape of Large Language Models in Global English News: Topics, Sentiments, and Spatiotemporal AnalysisabstractGenerative AI has exhibited considerable potential to transform various industries and public life. The role of news media coverage of generative AI is pivotal in shaping public perceptions and judgments about this significant technological innovation. This paper provides in-depth analysis and rich insights into the temporal and spatial distribution of topics, sentiment, and substantive themes within global news coverage focusing on the latest emerging technology—generative AI. We collected a comprehensive dataset of English news articles (January 2018 to November 2023, N = 24,827) through ProQuest databases. For topic modeling, we employed the BERTopic technique and combined it with qualitative coding to identify semantic themes. Subsequently, sentiment analysis was conducted using the RoBERTa-base model. Analysis of temporal patterns in the data reveals notable variability in coverage across key topics—business, corporate technological development, regulation and security, and education—with spikes in articles coinciding with major AI developments and policy discussions. Sentiment analysis shows a predominantly neutral to positive media stance, with the business-related articles exhibiting more positive sentiment, while regulation and security articles receive a reserved, neutral to negative sentiment. Our study offers a valuable framework to investigate global news discourse and evaluate news attitudes and themes related to emerging technologies. Lu Xian, Lingyao Li, Ben Zefeng Zhang, Libby Hemphill |
ICWSM | 5 |
| 2024 | A Bibliometric Review of Large Language Models Research from 2017 to 2023abstractLarge language models (LLMs), such as OpenAI's Generative Pre-trained Transformer (GPT), are a class of language models that have demonstrated outstanding performance across a range of natural language processing (NLP) tasks. LLMs have become a highly sought-after research area because of their ability to generate human-like language and their potential to revolutionize science and technology. In this study, we conduct bibliometric and discourse analyses of scholarly literature on LLMs. Synthesizing over 5,000 publications, this article serves as a roadmap for researchers, practitioners, and policymakers to navigate the current landscape of LLMs research. We present the research trends from 2017 to early 2023, identifying patterns in research paradigms and collaborations. We start with analyzing the core algorithm developments and NLP tasks that are fundamental in LLMs research. We then investigate the applications of LLMs in various fields and domains, including medicine, engineering, social science, and humanities. Our review also reveals the dynamic, fast-paced evolution of LLMs research. Overall, this article offers valuable insights into the current state, impact, and potential of LLMs research and its applications. Lizhou Fan, Lingyao Li, Zihui Ma, Sanggyu Lee, Huizi Yu, Libby Hemphill |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2024 | "HOT" ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social MediaabstractHarmful textual content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to this issue is developing detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful textual content. We used ChatGPT to investigate this potential and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful textual content on social media: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with the provided HOT definitions. However, ChatGPT classifies “hateful” and “offensive” as subsets of “toxic.” Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these insights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understanding and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance on the potential of using generative AI models for moderating large volumes of user-generated textual content on social media. Lingyao Li, Lizhou Fan, Shubham Atreja, Libby Hemphill |
ACM Trans. Web | 4 |
| 2023 | Direct, Orienting, and Scenic Paths: How Users Navigate Search in a Research Data ArchiveabstractSocial scientists increasingly share data so others can evaluate, replicate, and extend their research. To understand the process of data discovery as a precursor to data use, we study prospective users’ interactions with archived data. We gathered data for 98,000 user sessions initiated at a large social science data archive, the Inter-university Consortium for Political and Social Research (ICPSR). Our data reflect four years (2012-16) of users’ interactions with archival resources, including a data catalog, study-level metadata, variables, and publications that cite nearly 10,000 datasets. We constructed a network of user interactions linking website landing (e.g., site entrances) to exit pages, from which we identified three types of paths that users take through the research data archive: direct, orienting, and scenic. We also interpreted points of failure (e.g., drop-offs) and recurring behaviors (e.g., sensemaking) that support or impede data discovery along search paths. We articulate strategies that users adopt as they navigate data search and suggest ways to enhance the accessibility of data, metadata, and the systems that organize each. Sara Lafia, Anthony J. Million, Libby Hemphill |
CHIIR | 3 |
| 2022 | Leaders or Followers? A Temporal Analysis of Tweets from IRA Trolls
Siva K. Balasubramanian, Mustafa Bilgic 0001, Aron Culotta, Libby Hemphill, Anita Nikolich, Matthew A. Shapiro |
ICWSM | 4 |
| 2022 | How do properties of data, their curation, and their funding relate to reuse?abstractDespite large public investments in facilitating the secondary use of data, there is little information about the specific factors that predict data's reuse. Using data download logs from the Inter-university Consortium for Political and Social Research (ICPSR), this study examines how data properties, curation decisions, and repository funding models relate to data reuse. We find that datasets deposited by institutions, subject to many curatorial tasks, and whose access and preservation is funded externally, are used more often. Our findings confirm that investments in data collection, curation, and preservation are associated with more data reuse. Libby Hemphill, Amy M. Pienta, Sara Lafia, Dharma Akmon, David A. Bleckley |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2021 | Saving social media data: Understanding data management practices among social media researchers and their implications for archivesabstractAbstract Social media data (SMD) offer researchers new opportunities to leverage those data for their work in broad areas such as public opinion, digital culture, labor trends, and public health. The success of efforts to save SMD for reuse by researchers will depend on aligning data management and archiving practices with evolving norms around the capture, use, sharing, and security of datasets. This paper presents an initial foray into understanding how established practices for managing and preserving data should adapt to demands from researchers who use and reuse SMD, and from people who are subjects in SMD. We examine the data management practices of researchers who use SMD through a survey, and we analyze published articles that used data from Twitter. We discuss how researchers describe their data management practices and how these practices may differ from the management of conventional data types. We explore conceptual, technical, and ethical challenges for data archives based on the similarities and differences between SMD and other types of research data, focusing on the social sciences. Finally, we suggest areas where archives may need to revise policies, practices, and services in order to create secure, persistent, and usable collections of SMD. Libby Hemphill, Margaret L. Hedstrom, Susan H. Leonard |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2020 | Two Computational Models for Analyzing Political Attention in Social Media
Libby Hemphill, Angela M. Schöpke-Gonzalez |
ICWSM | 1 |
| 2018 | Forecasting the Presence and Intensity of Hostility on Instagram Using Linguistic and Social Features
Ping Liu 0002, Joshua Guberman, Libby Hemphill, Aron Culotta |
ICWSM | 3 |
| 2007 | Human-machine reconfigurations: Plans and situated actions, 2nd edabstractChapter 6 ("Georeferencing Elements in Metadata Standards") reviews the main metadata standards used for describing information objects that include geospatial references.Employing very helpful diagrams, the author explores the structure of the current main metadata standards and examines in detail their common points and differences.This comparison is then used as a departing point to open the discussion about overall interoperability across standards and the excessive complexity of the Geography Markup Language (GML) standard proposed by the Open Geospatial Consortium's.As a result, this chapter proposes a new generic structure for georeferencing and discusses the basic requirements of a geometry language that tries to meet the needs of general interoperability.The final objective of all the concepts, structures, and techniques introduced in the previous chapters is the development of systems that enable an effective retrieval of information objects related to specific geographic areas of interest.Chapter 7 ("Geographic Information Retrieval" [GIR]) presents the foundations of information retrieval based on geographic relationships.Unless GIR and text-based information retrieval can be used in combination, this chapter only explores the use of spatial characteristics for retrieval purposes.After presenting topological and geometric relations, the author introduces the main spatial operations that enable the matching of georeferenced information objects against users' spatial queries and analyses several currently running GIR systems.The final part of the chapter reviews current Libby Hemphill |
J. Assoc. Inf. Sci. Technol. | 1 |