VLDB 2026 Research / reviewers in the wild / expert
Amy Rechkemmer
dblp:266/0873
· DBLP profile ↗
7ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0001-7572-751XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource LanguagesabstractIn automated fact-checking (AFC), checkworthiness detection identifies claims requiring verification based on domain-specific criteria.On Wikipedia, this task instantiates as Citation Needed Detection (CND), which flags claims lacking supporting citations.However, existing research has largely overlooked lowerresource languages, and recent AFC pipelines rely on large language models (LLMs), which are inaccessible to low-resource organizations.We introduce MCN, a multilingual CND corpus spanning 18 languages across three resource levels, on which we conduct an extensive study of small decoder-based language models (SLMs).Our experiments show that SLMs fine-tuned with an encoder-style objective substantially outperform prompted LLMs across languages.We further present one of the first studies on cross-lingual CND, demonstrating that SLMs fine-tuned solely on English claims surpass LLMs, even with little to no target-language adaptation.Our findings have important implications for lower-resource Wikipedia communities and suggest that compact, task-specific models are preferable to LLMs for CND.We release all data and code at https://github.com/gerritq/mcnDataset Task Domain Languages (resource-level) Size NLP4IF 2021 (Shaar et al., 2021) CD/CWD Covid-19/Politics 3 (2 high, 1 medium) 1.3K-4K CitationNeeded (Redi et al., 2019) CWD Wikipedia 3 (3 high) 20K Kazemi et al. (2021) CD Covid-19, Politics 5 (2 high, 2 medium, 1 low) 5K Dutta et al. (2023) CD Politics 3 (2 high, 1 medium) 600-1.4KHalitaj and Zubiaga (2024) CWD Wikipedia 3 (2 high, 1 medium) 31K-1.1mCheckThat 2024 (Hasanain et al., 2024) CD/CWD Politics, Covid-19 3 (3 high) 1K-23K Baigutanova et al. (2026) CWD Wikipedia 5 (5 high) 20K MCN (ours) CWD Wikipedia 18 (8 high, 5 medium, 5 low) 2K-1.1m Gerrit Quaremba, Amy Rechkemmer, Elizabeth Black, Denny Vrandecic, Elena Simperl |
ACL (1) | 2 |
| 2024 | Snapper: Accelerating Bounding Box Annotation in Object Detection Tasks with Find-and-Snap ToolingabstractObject detection tasks are central to the development of datasets and algorithms in computer vision and machine learning. Despite its centrality, object detection remains tedious and time-consuming due to the inherent interactions that are often associated with drawing precise annotations. In this paper, we introduce Snapper, an interactive and intelligent annotation tool that intercepts bounding box annotations as they’re drawn and “snaps” them to the nearby object edges in real-time. Through a mixed-design user study with 18 full-time annotators, we compare Snapper’s annotation mode to alternative modes of annotation and find that Snapper enables participants to complete object detection tasks 39% more quickly without diminishing annotation quality. Further, we find that participants perceive Snapper as a tool that is interactively intuitive, trustworthy, and helpful. We conclude by discussing the implications of our findings as they relate to augmenting annotators’ conventions for drawing annotations in practice. Alex C. Williams, Min Bai, Jonathan Buck, Tristan McKinney, Amy Rechkemmer, Koushik Kalyanaraman, Matthew Lease, Patrick Haffner, Li Erran Li |
IUI | 5 |
| 2022 | When Confidence Meets Accuracy: Exploring the Effects of Multiple Performance Indicators on Trust in Machine Learning ModelsabstractPrevious research shows that laypeople’s trust in a machine learning model can be affected by both performance measurements of the model on the aggregate level and performance estimates on individual predictions. However, it is unclear how people would trust the model when multiple performance indicators are presented at the same time. We conduct an exploratory human-subject experiment to answer this question. We find that while the level of model confidence significantly affects people’s belief in model accuracy, both the model’s stated and observed accuracy generally have a larger impact on people’s willingness to follow the model’s predictions as well as their self-reported levels of trust in the model, especially after observing the model’s performance in practice. We hope the empirical evidence reported in this work could open doors to further studies to advance understanding of how people perceive, process, and react to performance-related information of machine learning. Amy Rechkemmer, Ming Yin 0001 |
CHI | 1 |
| 2022 | Understanding the Microtask Crowdsourcing Experience for Workers with Disabilities: A Comparative ViewabstractMicrotask crowdsourcing holds great potential as an employment opportunity with the flexibility and anonymity that individuals with disability may require. Though prior research has explored the accessibility of crowd work, the lived crowd work experiences of the broader community of workers with disability are still largely under-explored, especially when it comes to how their experiences are similar to or different from the experiences of workers without disability. In this work, we aim to obtain a deeper understanding of the microtask crowdsourcing experience for people with disabilities, especially regarding their financial and social experiences of participating in crowd work, along with the benefits and challenges that they encounter through this work. Specifically, we first surveyed 1,200 crowd workers both with and without disability about their experiences using the Amazon Mechanical Turk platform, and the differences we found inspired the design of a follow-up survey to gain greater understanding of the crowd work experience for workers with disability. Our findings reveal that workers with disability receive unique benefits from performing crowd work, such as a greater sense of purpose, but also encounter many challenges, such as completing tasks on time and earning a livable wage, causing them to turn to online communities for assistance. Although many of the challenges they face are not unique to crowd workers with disability, workers with disability may be disproportionately impacted by these challenges. From our findings, we provide implications for crowd platforms, as well as the gig economy as a whole, that seek to promote greater consideration of workers with a diverse range of conditions to create a more valuable work experience for them. Amy Rechkemmer, Ming Yin 0001 |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2021 | Exploring the Effects of Goal Setting When Training for Complex Crowdsourcing Tasks (Extended Abstract)abstractTraining is one way of enabling novice workers to work on complex crowdsourcing tasks. Based on goal setting theory in psychology, we conduct a randomized experiment to study whether and how setting different goals---including performance goal, learning goal, and behavioral goal---when training workers for a complex crowdsourcing task affects workers' learning perception, learning gain, and post-training performance. We find that setting different goals during training significantly affects workers' learning perception, but does not have an effect on learning gain or post-training performance. Further, exploratory analysis helps shed light on when and why various goals may or may not work in the crowdsourcing context. Amy Rechkemmer, Ming Yin 0001 |
IJCAI | 1 |
| 2020 | Motivating Novice Crowd Workers through Goal Setting: An Investigation into the Effects on Complex Crowdsourcing Task TrainingabstractTraining workers within a task is one way of enabling novice workers, who may lack domain knowledge or experience, to work on complex crowdsourcing tasks. Based on goal setting theory in psychology, we conduct a randomized experiment to study whether and how setting different goals—including performance goal, learning goal, and behavioral goal—when training workers for a complex crowdsourcing task affects workers’ learning perception, learning gain, and post-training performance. We find that setting different goals during training significantly affects workers’ learning perception, but overall does not have an effect on learning gain or post-training performance. However, higher levels of learning gain can be obtained when setting learning goals for workers who are highly learning-oriented. Additionally, giving workers a challenging behavioral goal can nudge them to adopt desirable behavior meant to improve learning and performance, though the adoption of such behavior does not lead to as much improvement as when the worker decides to take part in the behavior themselves. We conclude by discussing the lessons we’ve learned on how to effectively utilize goals in complex crowdsourcing task training. Amy Rechkemmer, Ming Yin 0001 |
HCOMP | 1 |
| 2020 | Small Town or Metropolis? Analyzing the Relationship between Population Size and LanguageabstractThe variance in language used by different cultures has been a topic of study for researchers in linguistics and psychology, but often times, language is compared across multiple countries in order to show a difference in culture. As a geographically large country that is diverse in population in terms of the background and experiences of its citizens, the U.S. also contains cultural differences within its own borders. Using a set of over 2 million posts from distinct Twitter users around the country dating back as far as 2014, we ask the following question: is there a difference in how Americans express themselves online depending on whether they reside in an urban or rural area? We categorize Twitter users as either urban or rural and identify ideas and language that are more commonly expressed in tweets written by one population over the other. We take this further by analyzing how the language from specific cities of the U.S. compares to the language of other cities and by training predictive models to predict whether a user is from an urban or rural area. We publicly release the tweet and user IDs that can be used to reconstruct the dataset for future studies in this direction. Amy Rechkemmer, Steven R. Wilson 0001, Rada Mihalcea |
LREC | 1 |