EDBT 2026 Demo / reviewers in the wild / expert
Sara Lafia
dblp:182/9740
· DBLP profile ↗
8ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-5896-7295ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building a Test Collection for Social Science Dataset RetrievalabstractData reuse expedites scientific progress and conserves resources, yet connecting researchers with datasets available for reuse remains a challenge. The scientific community has proposed several recommendation systems to help identify relevant data and maximize reuse. However, test collections or benchmarks for evaluating the performance of dataset recommendation systems are rare, particularly for social science dataset retrieval. To address this gap, we created a novel test collection for evaluating social science dataset recommendation systems. Our collection includes 249,102 query–dataset pairs with relevance judgments, featuring 262 unique search queries and 10,749 datasets. We describe how we created this collection using datasets archived at the Inter-university Consortium for Political and Social Research (ICPSR) and the ICPSR Bibliography of Data-related Literature. Additionally, we demonstrate a potential use case by evaluating the performance of embedding-based recommendation models on our test collection. The test collection is available through ICPSR at https://doi.org/10.3886/E238682V1. Ji Eun Kim, Sara Lafia, Libby Hemphill |
CHIIR | 2 |
| 2025 | Data, not documents: Moving beyond theories of information-seeking behavior to advance data discoveryabstractAbstract Many theories of human information behavior (HIB) assume that information objects are in text document format. This paper argues four important HIB theories are insufficient for describing users' search strategies for data because of assumptions about the attributes of objects that users seek. We first review and compare four HIB theories: Bates' berrypicking , Marchionni's electronic information search , Dervin's sense‐making , and Meho and Tibbo's social scientist information‐seeking . All four theories assume that information‐seekers search for text documents. Next, we compare these theories to search behavior by analyzing Google Analytics data from the Inter‐university Consortium for Political and Social Research (ICPSR). Users took direct, scenic, and orienting paths when searching for data. We also interviewed ICPSR users ( n = 20), and they said they needed dataset documentation and contextual information to find data. However, Dervin's sense‐making alone cannot explain the information‐seeking behaviors that we observed. Instead, what mattered most were object attributes determined by the type of information that users sought (i.e., data, not documents). We conclude by suggesting an alternative frame for building user‐centered data discovery tools. Anthony J. Million, Jeremy York, Sara Lafia, Libby Hemphill |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2023 | Direct, Orienting, and Scenic Paths: How Users Navigate Search in a Research Data ArchiveabstractSocial scientists increasingly share data so others can evaluate, replicate, and extend their research. To understand the process of data discovery as a precursor to data use, we study prospective users’ interactions with archived data. We gathered data for 98,000 user sessions initiated at a large social science data archive, the Inter-university Consortium for Political and Social Research (ICPSR). Our data reflect four years (2012-16) of users’ interactions with archival resources, including a data catalog, study-level metadata, variables, and publications that cite nearly 10,000 datasets. We constructed a network of user interactions linking website landing (e.g., site entrances) to exit pages, from which we identified three types of paths that users take through the research data archive: direct, orienting, and scenic. We also interpreted points of failure (e.g., drop-offs) and recurring behaviors (e.g., sensemaking) that support or impede data discovery along search paths. We articulate strategies that users adopt as they navigate data search and suggest ways to enhance the accessibility of data, metadata, and the systems that organize each. Sara Lafia, Anthony J. Million, Libby Hemphill |
CHIIR | 1 |
| 2022 | How do properties of data, their curation, and their funding relate to reuse?abstractDespite large public investments in facilitating the secondary use of data, there is little information about the specific factors that predict data's reuse. Using data download logs from the Inter-university Consortium for Political and Social Research (ICPSR), this study examines how data properties, curation decisions, and repository funding models relate to data reuse. We find that datasets deposited by institutions, subject to many curatorial tasks, and whose access and preservation is funded externally, are used more often. Our findings confirm that investments in data collection, curation, and preservation are associated with more data reuse. Libby Hemphill, Amy M. Pienta, Sara Lafia, Dharma Akmon, David A. Bleckley |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2022 | The Craft and Coordination of Data Curation: Complicating Workflow Views of Data ScienceabstractData curation is the process of making a dataset fit-for-use and archivable. It is critical to data-intensive science because it makes complex data pipelines possible, studies reproducible, and data reusable. Yet the complexities of the hands-on, technical, and intellectual work of data curation is frequently overlooked or downplayed. Obscuring the work of data curation not only renders the labor and contributions of data curators invisible but also hides the impact that curators' work has on the later usability, reliability, and reproducibility of data. To better understand the work and impact of data curation, we conducted a close examination of data curation at a large social science data repository, the Inter-university Consortium for Political and Social Research (ICPSR). We asked: What does curatorial work entail at ICPSR, and what work is more or less visible to different stakeholders and in different contexts? And, how is that curatorial work coordinated across the organization? We triangulated accounts of data curation from interviews and records of curation in Jira tickets to develop a rich and detailed account of curatorial work. While we identified numerous curatorial actions performed by ICPSR curators, we also found that curators rely on a number of craft practices to perform their jobs. The reality of their work practices defies the rote sequence of events implied by many life cycle or workflow models. Further, we show that craft practices are needed to enact data curation best practices and standards. The craft that goes into data curation is often invisible to end users, but it is well recognized by ICPSR curators and their supervisors. Explicitly acknowledging and supporting data curators as craftspeople is important in creating sustainable and successful curatorial infrastructures. Andrea K. Thomer, Dharma Akmon, Jeremy York, Allison R. B. Tyler, Faye Polasek, Sara Lafia, Libby Hemphill, Elizabeth Yakel |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2021 | Leveraging Machine Learning to Detect Data Curation ActivitiesabstractThis paper describes a machine learning approach for annotating and analyzing data curation work logs at ICPSR, a large social sciences data archive. The systems we studied track curation work and coordinate team decision-making at ICPSR. Archive staff use these systems to organize, prioritize, and document curation work done on datasets, making them promising resources for studying curation work and its impact on data reuse, especially in combination with data usage analytics. A key challenge, however, is classifying similar activities so that they can be measured and associated with impact metrics. This paper contributes: 1) a set of data curation activities; 2) a computational model for identifying curation actions in work log descriptions; and 3) an analysis of frequent data curation activities at ICPSR over time. We first propose a set of data curation actions to help us analyze the impact of curation work. We then use this set to annotate a set of data curation logs, which contain records of data transformations and project management decisions completed by archive staff. Finally, we train a text classifier to detect the frequency of curation actions in a large set of work logs. Our approach supports the analysis of curation work documented in work log systems as an important step toward studying the relationship between research data curation and data reuse. Sara Lafia, Andrea K. Thomer, David A. Bleckley, Dharma Akmon, Libby Hemphill |
e-Science | 1 |
| 2019 | Enabling the Discovery of Thematically Related Research Objects with Systematic SpatializationsabstractIt is challenging for scholars to discover thematically related research in a multidisciplinary setting, such as that of a university library. In this work, we use spatialization techniques to convey the relatedness of research themes without requiring scholars to have specific knowledge of disciplinary search terminology. We approach this task conceptually by revisiting existing spatialization techniques and reframing them in terms of core concepts of spatial information, highlighting their different capacities. To apply our design, we spatialize masters and doctoral theses (two kinds of research objects available through a university library repository) using topic modeling to assign a relatively small number of research topics to the objects. We discuss and implement two distinct spaces for exploration: a field view of research topics and a network view of research objects. We find that each space enables distinct visual perceptions and questions about the relatedness of research themes. A field view enables questions about the distribution of research objects in the topic space, while a network view enables questions about connections between research objects or about their centrality. Our work contributes to spatialization theory a systematic choice of spaces informed by core concepts of spatial information. Its application to the design of library discovery tools offers two distinct and intuitive ways to gain insights into the thematic relatedness of research objects, regardless of the disciplinary terms used to describe them. Sara Lafia, Christina Last, Werner Kuhn |
COSIT | 1 |
| 2019 | Talk of the Town: Discovering Open Public Data via Voice Assistants (Short Paper)abstractAccess to public data in the United States and elsewhere has steadily increased as governments have launched geospatially-enabled web portals like Socrata, CKAN, and Esri Hub. However, data discovery in these portals remains a challenge for the average user. Differences between users' colloquial search terms and authoritative metadata impede data discovery. For example, a motivated user with expertise can leverage valuable public data about transportation, real estate values, and crime, yet it remains difficult for the average user to discover and leverage data. To close this gap, community dashboards that use public data are being developed to track initiatives for public consumption; however, dashboards still require users to discover and interpret data. Alternatively, local governments are now developing data discovery systems that use voice assistants like Amazon Alexa and Google Home as conversational interfaces to public data portals. We explore these emerging technologies, examining the application areas they are designed to address and the degree to which they currently leverage existing open public geospatial data. In the context of ongoing technological advances, we envision using core concepts of spatial information to organize the geospatial themes of data exposed through voice assistant applications. This will allow us to curate them for improved discovery, ultimately supporting more meaningful user questions and their translation into spatial computations. Sara Lafia, Jingyi Xiao, Thomas Hervey, Werner Kuhn |
COSIT | 1 |