Milagros Miceli

dblp:257/8177 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-0585-3072ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 12 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Beyond Content Exposure: Systemic Factors Driving Moderators' Mental Health Crisis in Africa
abstract
Content moderators review disturbing content to protect social media users, often at significant cost to their mental health. Recent reports document the mental health conditions of African moderators as notably problematic. Beyond the content itself, what factors contribute to the deteriorating mental health of these workers? We surveyed 134 moderators across Africa to understand their mental health and interviewed 15 moderators to contextualize their experiences. We found that African moderators suffer from high psychological distress and lower well-being compared to moderators in other areas. Former moderators showed significantly higher distress levels, demonstrating long-term impact that extends beyond their moderation work. Our interviews showed that systemic and structural labor conditions contribute to moderators’ severe psychological distress and diminished mental well-being. Corporate wellness programs promoted by platforms were found ineffective and inadequate. We discuss how this requires holistic attention and structural solutions by all involved parties to improve moderators’ mental health.
Nuredin Ali Abdelkadir, Tianling Yang, Shivani Kapania, Kauna Ibrahim Malgwi, Fasica Berhane Gebrekidan, Adio-Adet Dinika, Elaine O. Nsoesie, Milagros Miceli, Stevie Chancellor
CHI8
2026 'The plan is just survival': Data Work in Kenya and the Regime of Entrapment
abstract
The rapid expansion of the AI industry relies heavily on the production, verification, and maintenance of data, otherwise known as "data work". Companies outsource and offshore this work through global AI supply chains that operate under exploitative conditions. Drawing on semi-structured interviews with Kenyan data workers across platforms and BPOs, this paper examines how such conditions take shape and persist. We argue that workers are caught within a regime of entrapment, a system of interconnected mechanisms that make it difficult for workers to leave or improve their positions. These mechanisms include the push to invest in the promise of ‘AI’ jobs, the use of precarious contracts to govern workers, the capture of regulatory institutions, and the exploitation of global labor arbitrage. Using complementary lenses of neoliberal governmentality, precarity, and supply chain capitalism, we analyze why labor mobilization in this sector remains uniquely constrained. We conclude by outlining an orientation for research and scholarly practice that can support workers’ organizing efforts and contest the structural conditions sustaining this regime.
Shivani Kapania, Tianling Yang, Nuredin Ali Abdelkadir, Morgan Klaus Scheuerman, Milagros Miceli, Alex S. Taylor, Sarah E. Fox
CHI5
2025 The Role of Expertise in Effectively Moderating Harmful Social Media Content
Nuredin Ali Abdelkadir, Tianling Yang, Shivani Kapania, Meron Estefanos, Fasica Berhane Gebrekidan, Zecharias Zelalem, Messai Ali, Rishan Berhe, Dylan K. Baker, Zeerak Talat, Milagros Miceli, Alex Hanna, Timnit Gebru
CHI11
2025 The Making of Performative Accuracy in AI Training: Precision Labor and Its Consequences
abstract
Peer Reviewed
Ben Zefeng Zhang, Tianling Yang, Milagros Miceli, Oliver L. Haimson, Michaelanne Thomas
CHI3
2025 What Knowledge Do We Produce from Social Media Data and How?
abstract
HCI and CSCW research that uses social media data to make inferences about individuals and communities has proliferated in the last decade. Previous studies have elaborated on methodological concerns and challenges and examined the assumptions and values underlying knowledge production through quantification and data. We expand this line of research by making visible and explicit the conventions and practices that establish, sustain, and reinforce current discourses in social media research. We conducted a Critical Discourse Analysis on 84 research papers published between 2010 and 2023 in CHI, CSCW, and GROUP that combine social media data and computational methods. Our findings show that plenty of this work legitimizes social media data as a valid source of information by centering its public availability, unobtrusiveness, and volume. Furthermore, to justify computational techniques, these papers prioritize computational expediency over data and method appropriateness. We argue that these embedded strategies may result in a methodological and epistemological distance between researchers and the studied communities, impacting problem framing, data collection, and finding application. With this work, we join the voices that have advocated for increased reflexivity in HCI and CSCW communities to scrutinize knowledge production and the role of researchers as knowledge producers.
Adriana Alvarado Garcia, Tianling Yang, Milagros Miceli
Proc. ACM Hum. Comput. Interact.3
2024 "Guilds" as Worker Empowerment and Control in a Chinese Data Work Platform
abstract
Data work plays a fundamental role in the development of algorithmic systems and the AI industry. It is often performed in business process outsourcing (BPO) companies and crowdsourcing platforms, involving a global and distributed workforce as well as networks of collaborative actors. Previous work on community building among data workers centers organization and mutual support or focuses on the structuring and instrumentalization of crowdworker groups for complicated projects. We add to these lines of research by focusing on a specific form of community building encouraged and facilitated by platforms in China: guilds. Based on ethnographic work on a Chinese crowdsourcing platform and 14 semi-structured interviews with data workers, our findings show that guilds are a form of both worker empowerment and control. With this work, we add a nuanced empirical case to the interconnection of BPOs, online communities and crowdsourcing platforms in the current data production sector in China, thus expanding previous investigations on global perspectives of data production. We discuss guilds in relation to individual workers and highlight their effects on data work, including efficient coordination, enhanced standardization, and flattened power structure.
Tianling Yang, Milagros Miceli
Proc. ACM Hum. Comput. Interact.2
2023 Mobilizing Social Media Data: Reflections of a Researcher Mediating between Data and Organization
abstract
This paper examines the practices involved in mobilizing social media data from their site of production to the institutional context of non-profit organizations. We report on nine months of fieldwork with a transnational and intergovernmental organization using social media data to understand the role of grassroots initiatives in Mexico, in the unique context of the COVID-19 pandemic. We show how different stakeholders negotiate the definition of problems to be addressed with social media data, the collective creation of ground-truth, and the limitations involved in the process of extracting value from data. The meanings of social media data are not defined in advance; instead, they are contingent on the practices and needs of the organization that seeks to extract insights from the analysis. We conclude with a list of reflections and questions for researchers who mediate in the mobilization of social media data into non-profit organizations to inform humanitarian action.
Adriana Alvarado Garcia, Marisol Wong-Villacres, Milagros Miceli, Benjamín Hernández, Christopher A. Le Dantec
CHI3
2022 The Data-Production Dispositif
abstract
Machine learning (ML) depends on data to train and verify models. Very often, organizations outsource processes related to data work (i.e., generating and annotating data and evaluating outputs) through business process outsourcing (BPO) companies and crowdsourcing platforms. This paper investigates outsourced ML data work in Latin America by studying three platforms in Venezuela and a BPO in Argentina. We lean on the Foucauldian notion of dispositif to define the data-production dispositif as an ensemble of discourses, actions, and objects strategically disposed to (re)produce power/knowledge relations in data and labor. Our dispositif analysis comprises the examination of 210 data work instruction documents, 55 interviews with data workers, managers, and requesters, and participant observation. Our findings show that discourses encoded in instructions reproduce and normalize the worldviews of requesters. Precarious working conditions and economic dependency alienate workers, making them obedient to instructions. Furthermore, discourses and social contexts materialize in artifacts, such as interfaces and performance metrics, limiting workers' agency and normalizing specific ways of interpreting data. We conclude by stressing the importance of counteracting the data-production dispositif by fighting alienation and precarization, and empowering data workers to become assets in the quest for high-quality data.
Milagros Miceli, Julian Posada
Proc. ACM Hum. Comput. Interact.1
2022 Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?
abstract
Research in machine learning (ML) has argued that models trained on incomplete or biased datasets can lead to discriminatory outputs. In this commentary, we propose moving the research focus beyond bias-oriented framings by adopting a power-aware perspective to "study up" ML datasets. This means accounting for historical inequities, labor conditions, and epistemological standpoints inscribed in data. We draw on HCI and CSCW work to support our argument, critically analyze previous research, and point at two co-existing lines of work within our research community \,---\,one bias-centered, the other power-aware. We highlight the need for dialogue and cooperation in three areas: data quality, data work, and data documentation. In the first area, we argue that reducing societal problems to "bias" misses the context-based nature of data. In the second one, we highlight the corporate forces and market imperatives involved in the labor of data workers that subsequently shape ML datasets. Finally, we propose expanding current transparency-oriented efforts in dataset documentation to reflect the social contexts of data design and production.
Milagros Miceli, Julian Posada, Tianling Yang
Proc. ACM Hum. Comput. Interact.1
2022 Documenting Data Production Processes: A Participatory Approach for Data Work
abstract
The opacity of machine learning data is a significant threat to ethical data work and intelligible systems. Previous research has addressed this issue by proposing standardized checklists to document datasets. This paper expands that field of inquiry by proposing a shift of perspective: from documenting datasets towards documenting data production. We draw on participatory design and collaborate with data workers at two companies located in Bulgaria and Argentina, where the collection and annotation of data for machine learning are outsourced. Our investigation comprises 2.5 years of research, including 33 semi-structured interviews, five co-design workshops, the development of prototypes, and several feedback instances with participants. We identify key challenges and requirements related to the integration of documentation practices in real-world data production scenarios. Our findings comprise important design considerations and highlight the value of designing data documentation based on the needs of data workers. We argue that a view of documentation as a boundary object, i.e., an object that can be used differently across organizations and teams but holds enough immutable content to maintain integrity, can be useful when designing documentation to retrieve heterogeneous, often distributed, contexts of data production.
Milagros Miceli, Tianling Yang, Adriana Alvarado Garcia, Julian Posada, Sonja Mei Wang, Marc Pohl, Alex Hanna
Proc. ACM Hum. Comput. Interact.1
2020 Biased Priorities, Biased Outcomes: Three Recommendations for Ethics-oriented Data Annotation Practices
abstract
In this paper, we analyze the relation between data-related biases and practices of data annotation, by placing them in the context of market economy. We understand annotation as a praxis related to the sensemaking of data and investigate annotation practices for vision models by focusing on the values that are prioritized by industrial decision-makers and practitioners. The quality of data is critical for machine learning models as it holds the power to (mis-)represent the population it is intended to analyze. For autonomous systems to be able to make sense of the world, humans first need to make sense of the data these systems will be trained on. This paper addresses this issue, guided by the following research questions: Which goals are prioritized by decision-makers at the data annotation stage? How do these priorities correlate with data-related bias issues? Focusing on work practices and their context, our research goal aims at understanding the logics driving companies and their impact on the performed annotations. The study follows a qualitative design and is based on 24 interviews with relevant actors and extensive participatory observations, including several weeks of fieldwork at two companies dedicated to data annotation for vision models in Buenos Aires, Argentina and Sofia, Bulgaria. The prevalence of market-oriented values over socially responsible approaches is argued based on three corporate priorities that inform work practices in this field and directly shape the annotations performed: profit (short deadlines connected to the strive for profit are prioritized over alternative approaches that could prevent biased outcomes), standardization (the strive for standardized and, in many cases, reductive or biased annotations to make data fit the products and revenue plans of clients), and opacity (related to client's power to impose their criteria on the annotations that are performed. Criteria that most of the times remain opaque due to corporate confidentiality). Finally, we introduce three elements, aiming at developing ethics-oriented practices of data annotation, that could help prevent biased outcomes: transparency (regarding the documentation of data transformations, including information on responsibilities and criteria for decision-making.), education (training on the potential harms caused by AI and its ethical implications, that could help data annotators and related roles adopt a more critical approach towards the interpretation and labeling of data), and regulations (clear guidelines for ethical AI developed at the governmental level and applied both in private and public organizations).
Gunay Kazimzade, Milagros Miceli
AIES2
2020 Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer Vision
abstract
The interpretation of data is fundamental to machine learning. This paper investigates practices of image data annotation as performed in industrial contexts. We define data annotation as a sense-making practice, where annotators assign meaning to data through the use of labels. Previous human-centered investigations have largely focused on annotators? subjectivity as a major cause of biased labels. We propose a wider view on this issue: guided by constructivist grounded theory, we conducted several weeks of fieldwork at two annotation companies. We analyzed which structures, power relations, and naturalized impositions shape the interpretation of data. Our results show that the work of annotators is profoundly informed by the interests, values, and priorities of other actors above their station. Arbitrary classifications are vertically imposed on annotators, and through them, on data. This imposition is largely naturalized. Assigning meaning to data is often presented as a technical matter. This paper shows it is, in fact, an exercise of power with multiple implications for individuals and society.
Milagros Miceli, Martin Schuessler, Tianling Yang
Proc. ACM Hum. Comput. Interact.1