Aaron Halfaker

dblp:26/2369 · also Aaron Lee Halfaker · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-8907-6367ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 21 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 Summaries, Highlights, and Action Items: Design, Implementation and Evaluation of an LLM-powered Meeting Recap System
abstract
Meetings play a critical infrastructural role in coordinating work. The recent surge of hybrid and remote meetings in computer-mediated spaces has led to new problems (e.g., more time spent in less engaging meetings) and new opportunities (e.g., automated transcription/captioning and recap support). Advances in dialogue summarization offer the potential for improving post-meeting experiences, but fixed-length summaries often fail to meet diverse needs, such as quick overviews or detailed insights. To address these gaps, we use cognitive science and discourse theories to conceptualize two recap designs: important highlights and a structured, hierarchical minutes view, targeting complementary recap needs. We operationalize these representations into high-fidelity prototypes using dialogue summarization. Finally, we evaluate the representations' effectiveness with seven users in the context of their work meetings at Microsoft. Our results show both recap types are valuable in different contexts, enabling collaboration through discussions and consensus-building. Exploring the meaning of users adding, editing, and deleting from recaps suggests varying alignment for using these actions to improve AI-recap. Our design implications, such as incorporating organizational artifacts (e.g., linking presentations) in recaps and personalizing context, advance the discourse of effective recap designs for organizational work and support past results from cognition studies.
Sumit Asthana, Sagih Hilleli, Aaron Halfaker
Proc. ACM Hum. Comput. Interact.4
2024 Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
abstract
AI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation.
Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Meng-Hsin Wu, Sherry Tongshuang Wu, Kenneth Holstein, Haiyi Zhu
CHI2
2023 On Improving Summarization Factual Consistency from Natural Language Feedback
abstract
Yixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir Radev, Ahmed Hassan Awadallah. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yixin Liu 0003, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir R. Radev, Ahmed Awadallah 0001
ACL (1)4
2021 Wikipedia ORES Explorer: Visualizing Trade-offs For Designing Applications With Machine Learning API
abstract
With the growing industry applications of Artificial Intelligence (AI) systems, pre-trained models and APIs have emerged and greatly lowered the barrier of building AI-powered products. However, novice AI application designers often struggle to recognize the inherent algorithmic trade-offs and evaluate model fairness before making informed design decisions. In this study, we examined the Objective Revision Evaluation System (ORES), a machine learning (ML) API in Wikipedia used by the community to build anti-vandalism tools. We designed an interactive visualization system to communicate model threshold trade-offs and fairness in ORES. We evaluated our system by conducting 10 in-depth interviews with potential ORES application designers. We found that our system helped application designers who have limited ML backgrounds learn about in-context ML knowledge, recognize inherent value trade-offs, and make design decisions that aligned with their goals. By demonstrating our system in a real-world domain, this paper presents a novel visualization approach to facilitate greater accessibility and human agency in AI application design.
Zining Ye, Xinran Yuan, Shaurya Gaur, Aaron Halfaker, Jodi Forlizzi, Haiyi Zhu
Conference on Designing Interactive Systems4
2021 Automatically Labeling Low Quality Content on Wikipedia By Leveraging Patterns in Editing Behaviors
abstract
Wikipedia articles aim to be definitive sources of encyclopedic content. Yet, only 0.6% of Wikipedia articles have high quality according to its quality scale due to insufficient number of Wikipedia editors and enormous number of articles. Supervised Machine Learning (ML) quality improvement approaches that can automatically identify and fix content issues rely on manual labels of individual Wikipedia sentence quality. However, current labeling approaches are tedious and produce noisy labels. Here, we propose an automated labeling approach that identifies the semantic category (e.g., adding citations, clarifications) of historic Wikipedia edits and uses the modified sentences prior to the edit as examples that require that semantic improvement. Highest-rated article sentences are examples that no longer need semantic improvements. We show that training existing sentence quality classification algorithms on our labels improves their performance compared to training them on existing labels. Our work shows that editing behaviors of Wikipedia editors provide better labels than labels generated by crowdworkers who lack the context to make judgments that the editors would agree with.
Sumit Asthana, Sabrina Tobar Thommel, Aaron Halfaker, Nikola Banovic 0001
Proc. ACM Hum. Comput. Interact.3
2021 Effects of Algorithmic Flagging on Fairness: Quasi-experimental Evidence from Wikipedia
abstract
Online community moderators often rely on social signals such as whether or not a user has an account or a profile page as clues that users may cause problems. Reliance on these clues can lead to "overprofiling'' bias when moderators focus on these signals but overlook the misbehavior of others. We propose that algorithmic flagging systems deployed to improve the efficiency of moderation work can also make moderation actions more fair to these users by reducing reliance on social signals and making norm violations by everyone else more visible. We analyze moderator behavior in Wikipedia as mediated by RCFilters, a system which displays social signals and algorithmic flags, and estimate the causal effect of being flagged on moderator actions. We show that algorithmically flagged edits are reverted more often, especially those by established editors with positive social signals, and that flagging decreases the likelihood that moderation actions will be undone. Our results suggest that algorithmic flagging systems can lead to increased fairness in some contexts but that the relationship is complex and contingent.
Nathan TeBlunthuis, Benjamin Mako Hill, Aaron Halfaker
Proc. ACM Hum. Comput. Interact.3
2020 Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based Systems
abstract
On Wikipedia, sophisticated algorithmic tools are used to assess the quality of edits and take corrective actions. However, algorithms can fail to solve the problems they were designed for if they conflict with the values of communities who use them. In this study, we take a Value-Sensitive Algorithm Design approach to understanding a community-created and -maintained machine learning-based algorithm called the Objective Revision Evaluation System (ORES)---a quality prediction system used in numerous Wikipedia applications and contexts. Five major values converged across stakeholder groups that ORES (and its dependent applications) should: (1) reduce the effort of community maintenance, (2) maintain human judgement as the final authority, (3) support differing peoples' differing workflows, (4) encourage positive engagement with diverse editor groups, and (5) establish trustworthiness of people and algorithms within the community. We reveal tensions between these values and discuss implications for future research to improve algorithms like ORES.
C. Estelle Smith, Bowen Yu 0001, Anjali Srivastava, Aaron Halfaker, Loren G. Terveen, Haiyi Zhu
CHI4
2020 ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia
abstract
Algorithmic systems---from rule-based bots to machine learning classifiers---have a long history of supporting the essential work of content moderation and other curation work in peer production projects. From counter-vandalism to task routing, basic machine prediction has allowed open knowledge projects like Wikipedia to scale to the largest encyclopedia in the world, while maintaining quality and consistency. However, conversations about how quality control should work and what role algorithms should play have generally been led by the expert engineers who have the skills and resources to develop and modify these complex algorithmic systems. In this paper, we describe ORES: an algorithmic scoring service that supports real-time scoring of wiki edits using multiple independent classifiers trained on different datasets. ORES decouples several activities that have typically all been performed by engineers: choosing or curating training data, building models to serve predictions, auditing predictions, and developing interfaces or automated agents that act on those predictions. This meta-algorithmic system was designed to open up socio-technical conversations about algorithms in Wikipedia to a broader set of participants. In this paper, we discuss the theoretical mechanisms of social change ORES enables and detail case studies in participatory machine learning around ORES from the 5 years since its deployment.
Aaron Halfaker, R. Stuart Geiger
Proc. ACM Hum. Comput. Interact.1
2018 Distance and Attraction: Gravity Models for Geographic Content Production
abstract
Volunteered Geographic Information (VGI), such as contributions to OpenStreetMap and geotagged Wikipedia articles, is often assumed to be produced locally. However, recent work has found that peer-produced VGI is frequently contributed by non-locals. We evaluate this approach across hundreds of content types from Wikipedia, OpenStreetMap, and eBird, and show that these models can describe more than 90% of "VGI flows" for some content types. Our findings advance geographic HCI theory, suggesting some spatial mechanisms underpinning VGI production. We also discuss design implications that can help (a) human and algorithmic consumers of VGI evaluate the perspectives it contains and (b) address geographic coverage variations in these platforms (e.g. via more effective volunteer recruitment strategies).
Jacob Thebault-Spieker, Aaron Halfaker, Loren G. Terveen, Brent J. Hecht
CHI2
2018 Information Fortification: An Online Citation Behavior
abstract
In this multi-method study, we examine citation activity on English-language Wikipedia to understand how information claims are supported in a non-scientific open collaboration context. We draw on three data sources-edit logs, interview data, and document analysis-to present an integrated interpretation of citation activity and found pervasive themes related to controversy and conflict. Based on this analysis, we present and discuss information fortification as a concept that explains online citation activity that arises from both naturally occurring and manufactured forms of controversy. This analysis challenges a workshop position paper from Group 2005 by Forte and Bruckman, which draws on Latour's sociology of science and citation to explain citation in Wikipedia with a focus on credibility seeking. We discuss how information fortification differs from theories of citation that have arisen from bibliometrics scholarship and are based on scientific citation practices.
Andrea Forte, Nazanin Andalibi, Tim Gorichanaz, Meen Chul Kim, Thomas H. Park, Aaron Halfaker
GROUP6
2018 Evaluating the impact of the Wikipedia Teahouse on newcomer socialization and retention
abstract
Effective socialization of new contributors is vital for the long-term sustainability of open collaboration projects. Previous research has identified many common barriers to participation. However, few interventions employed to increase newcomer retention over the long term by improving aspects of the onboarding experience have demonstrated success. This study presents an evaluation of the impact of one such intervention, the Wikipedia Teahouse, on new editor survival. In a controlled experiment, we find that new editors invited to the Teahouse are retained at a higher rate than editors who do not receive an invite. The effect is observed for both low-and high-activity newcomers, and for both short- and long-term survival.
Jonathan T. Morgan, Aaron Halfaker
OpenSym2
2018 With Few Eyes, All Hoaxes are Deep
abstract
Quality control is critical to open production communities like Wikipedia. Wikipedia editors enact border quality control with edits (counter-vandalism) and new article creations (new page patrolling) shortly after they are saved. In this paper, we describe a long-standing set of inefficiencies that have plagued new page patrolling by drawing a contrast to the more efficient, distributed processes for counter-vandalism. Further, to address this issue, we demonstrate an effective automated topic model based on a labeling strategy that leverages a folksonomy developed by subject specific working groups in Wikipedia (WikiProject tags) and a flexible ontology (WikiProjects Directory) to arrive at a hierarchical and uniform label set. We are able to attain very high fitness measures (macro ROC-AUC: 95.2%, macro PR-AUC: 74.5%) and real-time performance using word2vec-based features. Finally, we present a proposal for how incorporating this model into current tools will shift the dynamics of new article review positively.
Sumit Asthana, Aaron Halfaker
Proc. ACM Hum. Comput. Interact.2
2018 Bot Detection in Wikidata Using Behavioral and Other Informal Cues
abstract
Bots have been important to peer production's success. Wikipedia, OpenStreetMap, and Wikidata all have taken advantage of automation to perform work at a rate and scale exceeding that of human contributors. Understanding the ways in which humans and bots behave in these communities is an important topic, and one that relies on accurate bot recognition. Yet, in many cases, bot activities are not explicitly flagged and could be mistaken for human contributions. We develop a machine classifier to detect previously unidentified bots using implicit behavioral and other informal editing characteristics. We show that this method yields a high level of fitness under both formal evaluation (PR-AUC: 0.845, ROC-AUC: 0.985) and a qualitative analysis of "anonymous" contributor edit sessions. We also show that, in some cases, unflagged bot activities can significantly misrepresent human behavior in analyses. Our model has the potential to support future research and community patrolling activities.
Andrew Hall, Loren G. Terveen, Aaron Halfaker
Proc. ACM Hum. Comput. Interact.3
2018 Value-Sensitive Algorithm Design: Method, Case Study, and Lessons
abstract
Most commonly used approaches to developing automated or artificially intelligent algorithmic systems are Big Data-driven and machine learning-based. However, these approaches can fail, for two notable reasons: (1) they may lack critical engagement with users and other stakeholders; (2) they rely largely on historical human judgments, which do not capture and incorporate human insights into how the world can be improved in the future. We propose and describe a novel method for the design of such algorithms, which we call Value Sensitive Algorithm Design. Value Sensitive Algorithm Design incorporates stakeholders' tacit knowledge and explicit feedback in the early stages of algorithm creation. This increases the chance to avoid biases in design choices or to compromise key stakeholder values. Generally, we believe that algorithms should be designed to balance multiple stakeholders' needs, motivations, and interests, and to help achieve important collective goals. We also describe a specific project "Designing Intelligent Socialization Algorithms for WikiProjects in Wikipedia" to illustrate our method. We intend this paper to contribute to the rich ongoing conversation concerning the use of algorithms in supporting critical decision-making in society.
Haiyi Zhu, Bowen Yu 0001, Aaron Halfaker, Loren G. Terveen
Proc. ACM Hum. Comput. Interact.3
2017 Identifying Semantic Edit Intentions from Revisions in Wikipedia
abstract
Most studies on human editing focus merely on syntactic revision operations, failing to capture the intentions behind revision changes, which are essential for facilitating the single and collaborative writing process.In this work, we develop in collaboration with Wikipedia editors a 13-category taxonomy of the semantic intention behind edits in Wikipedia articles.Using labeled article edits, we build a computational classifier of intentions that achieved a micro-averaged F1 score of 0.621.We use this model to investigate edit intention effectiveness: how different types of edits predict the retention of newcomers and changes in the quality of articles, two key concerns for Wikipedia today.Our analysis shows that the types of edits that users make in their first session predict their subsequent survival as Wikipedia editors, and articles in different stages need different types of edits.
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy
EMNLP2
2017 Interpolating Quality Dynamics in Wikipedia and Demonstrating the Keilana Effect
abstract
For open, volunteer generated content like Wikipedia, quality is a prominent concern. To measure Wikipedia's quality, researchers have historically relied on expert evaluation or assessments of article quality by Wikipedians themselves. While both of these methods have proven effective for answering many questions about Wikipedia's quality and processes, they are both problematic: expert evaluation is expensive and Wikipedian quality assessments are sporadic and unpredictable. Studies that explore Wikipedia's quality level or the processes that result in quality improvements have only examined small snapshots of Wikipedia and often rely on complex propensity models to deal with the unpredictable nature of Wikipedians' own assessments. In this paper, I describe a method for measuring article quality in Wikipedia historically and at a finer granularity than was previously possible. I use this method to demonstrate an important coverage dynamic in Wikipedia (specifically, articles about women scientists) and offer this method, dataset, and open API to the research community studying Wikipedia quality dynamics.
Aaron Halfaker
OpenSym1
2017 Operationalizing Conflict and Cooperation between Automated Software Agents in Wikipedia: A Replication and Expansion of 'Even Good Bots Fight'
abstract
This paper replicates, extends, and refutes conclusions made in a study published in PLoS ONE ("Even Good Bots Fight"), which claimed to identify substantial levels of conflict between automated software agents (or bots) in Wikipedia using purely quantitative methods. By applying an integrative mixed-methods approach drawing on trace ethnography, we place these alleged cases of bot-bot conflict into context and arrive at a better understanding of these interactions. We found that overwhelmingly, the interactions previously characterized as problematic instances of conflict are typically better characterized as routine, productive, even collaborative work. These results challenge past work and show the importance of qualitative/quantitative collaboration. In our paper, we present quantitative metrics and qualitative heuristics for operationalizing bot-bot conflict. We give thick descriptions of kinds of events that present as bot-bot reverts, helping distinguish conflict from non-conflict. We computationally classify these kinds of events through patterns in edit summaries. By interpreting found/trace data in the socio-technical contexts in which people give that data meaning, we gain more from quantitative measurements, drawing deeper understandings about the governance of algorithmic systems in Wikipedia. We have also released our data collection, processing, and analysis pipeline, to facilitate computational reproducibility of our findings and to help other researchers interested in conducting similar mixed-method scholarship in other platforms and contexts.
R. Stuart Geiger, Aaron Halfaker
Proc. ACM Hum. Comput. Interact.2
2017 Simulation Experiments on (the Absence of) Ratings Bias in Reputation Systems
abstract
As the gig economy continues to grow and freelance work moves online, five-star reputation systems are becoming more and more common. At the same time, there are increasing accounts of race and gender bias in evaluations of gig workers, with negative impacts for those workers. We report on a series of four Mechanical Turk-based studies in which participants who rated simulated gig work did not show race- or gender bias, while manipulation checks showed they reliably distinguished between low- and high-quality work. Given prior research, this was a striking result. To explore further, we used a Bayesian approach to verify absence of ratings bias (as opposed to merely not detecting bias). This Bayesian test let us identify an upper- bound: if any bias did exist in our studies, it was below an average of 0.2 stars on a five-star scale. We discuss possible interpretations of our results and outline future work to better understand the results.
Jacob Thebault-Spieker, Daniel Kluver, Maximilian A. Klein, Aaron Halfaker, Brent J. Hecht, Loren G. Terveen, Joseph A. Konstan
Proc. ACM Hum. Comput. Interact.4
2016 Not at Home on the Range: Peer Production and the Urban/Rural Divide
abstract
Wikipedia articles about places, OpenStreetMap features, and other forms of peer-produced content have become critical sources of geographic knowledge for humans and intelligent technologies. In this paper, we explore the effectiveness of the peer production model across the rural/urban divide, a divide that has been shown to be an important factor in many online social systems. We find that in both Wikipedia and OpenStreetMap, peer-produced content about rural areas is of systematically lower quality, is less likely to have been produced by contributors who focus on the local area, and is more likely to have been generated by automated software agents (i.e. "bots"). We then codify the systemic challenges inherent to characterizing rural phenomena through peer production and discuss potential solutions.
Isaac L. Johnson, Allen Yilun Lin, Toby Jia-Jun Li, Andrew Hall, Aaron Halfaker, Johannes Schöning, Brent J. Hecht
CHI5
2016 Who Did What: Editor Role Identification in Wikipedia
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy
ICWSM2
2016 Edit Categories and Editor Role Identification in Wikipedia
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy
LREC2
2015 User Session Identification Based on Strong Regularities in Inter-activity Time
abstract
Session identification is a common strategy used to develop metrics for web analytics and perform behavioral analyses of user-facing systems. Past work has argued that session identification strategies based on an inactivity threshold is inherently arbitrary or has advocated that thresholds be set at about 30 minutes. In this work, we demonstrate a strong regularity in the temporal rhythms of user initiated events across several different domains of online activity (incl. video gaming, search, page views and volunteer contributions). We describe a methodology for identifying clusters of user activity and argue that the regularity with which these activity clusters appear implies a good rule-of-thumb inactivity threshold of about 1 hour. We conclude with implications that these temporal rhythms may have for system design based on our observations and theories of goal-directed human activity.
Aaron Halfaker, Oliver Keyes, Daniel Kluver, Jacob Thebault-Spieker, Tien T. Nguyen, Kenneth Shores, Anuradha Uduwage, Morten Warncke-Wang
WWW1
2014 Snuggle: designing for efficient socialization and ideological critique
abstract
Wikipedia, the encyclopedia "anyone can edit", has become increasingly less so. Recent academic research and popular discourse illustrates the often aggressive ways newcomers are treated by veteran Wikipedians. These are complex sociotechnical issues, bound up in infrastructures based on problematic ideologies. In response, we worked with a coalition of Wikipedians to design, develop, and deploy Snuggle, a new user interface that served two critical functions: making the work of newcomer socialization more effective, and bringing visibility to instances in which Wikipedians? current practice of gatekeeping socialization breaks down. Snuggle supports positive socialization by helping mentors quickly find newcomers whose good-faith mistakes were reverted as damage. Snuggle also supports ideological critique and reflection by bringing visibility to the consequences of viewing newcomers through a lens of suspiciousness.
Aaron Halfaker, R. Stuart Geiger, Loren G. Terveen
CHI1
2014 Accept, decline, postpone: How newcomer productivity is reduced in English Wikipedia by pre-publication review
abstract
Wikipedia needs to attract and retain newcomers while also increasing the quality of its content. Yet new Wikipedia users are disproportionately affected by the quality assurance mechanisms designed to thwart spammers and promoters. English Wikipedia's Articles for Creation provides a protected space for drafting new articles, which are reviewed against minimum quality guidelines before they are published. In this study we explore how this drafting process has affected the productivity of newcomers in Wikipedia. Using a mixed qualitative and quantitative approach, we show how the process's pre-publication review, which is intended to improve the success of newcomers, in fact decreases newcomer productivity in English Wikipedia and offer recommendations for system designers.
Jodi Schneider, Bluma S. Gelley, Aaron Halfaker
OpenSym3
2013 Using edit sessions to measure participation in wikipedia
abstract
Many quantitative, log-based studies of participation and contribution in CSCW and CMC systems measure the activity of users in terms of output, based on metrics like posts to forums, edits to Wikipedia articles, or commits to code repositories. In this paper, we instead seek to estimate the amount of time users have spent contributing. Through an analysis of Wikipedia log data, we identify a pattern of punctuated bursts in editors' activity that we refer to as edit sessions. Based on these edit sessions, we build a metric that approximates the labor hours of editors in the encyclopedia. Using this metric, we first compare labor-based analyses with output-based analyses, finding that the activity of many editors can appear quite differently based on the kind of metric used. Second, we use edit session data to examine phenomena that cannot be adequately studied with purely output-based metrics, such as the total number of labor hours for the entire project.
R. Stuart Geiger, Aaron Halfaker
CSCW2
2013 Making peripheral participation legitimate: reader engagement experiments in wikipedia
abstract
Open collaboration communities thrive when participation is plentiful. Recent research has shown that the English Wikipedia community has constructed a vast and accurate information resource primarily through the monumental effort of a relatively small number of active, volunteer editors. Beyond Wikipedia's active editor community is a substantially larger pool of potential participants: readers. In this paper we describe a set of field experiments using the Article Feedback Tool, a system designed to elicit lightweight contributions from Wikipedia's readers. Through the lens of social learning theory and comparisons to related work in open bug tracking software, we evaluate the costs and benefits of the expanded participation model and show both qualitatively and quantitatively that peripheral contributors add value to an open collaboration community as long as the cost of identifying low quality contributions remains low.
Aaron Halfaker, Oliver Keyes, Dario Taraborelli
CSCW1
2013 When the levee breaks: without bots, what happens to Wikipedia's quality control processes?
abstract
In the first half of 2011, ClueBot NG -- one of the most prolific counter-vandalism bots in the English-language Wikipedia -- went down for four distinct periods, each period of downtime lasting from days to weeks. In this paper, we use these periods of breakdown as naturalistic experiments to study Wikipedia's heterogeneous quality control network, which we analyze as a multi-tiered system in which distinct classes of reviewers use various reviewing technologies to patrol for different kinds of damage at staggered time periods. Our analysis showed that the overall time-to-revert edits was almost doubled when this software agent was down. Yet while a significantly fewer proportion of edits made during the bot's downtime were reverted, we found that those edits were later eventually reverted. This suggests that other agents in Wikipedia took over this quality control work, but performed it at a far slower rate.
R. Stuart Geiger, Aaron Halfaker
OpenSym2
2012 Defense Mechanism or Socialization Tactic? Improving Wikipedia's Notifications to Rejected Contributors
R. Stuart Geiger, Aaron Halfaker, Maryana Pinchuk, Steven Walling
ICWSM2
2009 Wikipedians are born, not made: a study of power editors on Wikipedia
abstract
Open content web sites depend on users to produce information of value. Wikipedia is the largest and most well-known such site. Previous work has shown that a small fraction of editors --Wikipedians -- do most of the work and produce most of the value. Other work has offered conjectures about how Wikipedians differ from other editors and how Wikipedians change over time. We quantify and test these conjectures. Our key findings include: Wikipedians' edits last longer; Wikipedians invoke community norms more often to justify their edits; on many dimensions of activity, Wikipedians start intensely, tail off a little, then maintain a relatively high level of activity over the course of their career. Finally, we show that the amount of work done by Wikipedians and non-Wikipedians differs significantly from their very first day. Our results suggest a design opportunity: customizing the initial user experience to improve retention and channel new users' intense energy.
Katherine A. Panciera, Aaron Halfaker, Loren G. Terveen
GROUP2