Ding Wang 0006

dblp:99/4292-6 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-5037-5017ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Designing Around Stigma: Human-Centered LLMs for Menstrual Health
abstract
Menstrual health education (MHE) in Pakistan is constrained by cultural taboos and inadequate formal curricula, leaving women with few trusted resources to lean on. In response to these challenges, we introduce a WhatsApp-based chatbot powered by a large language model (LLM) and Retrieval-Augmented Generation (RAG), co-designed with Pakistani college women. Workshops (N=30) revealed key design requirements—support for Roman Urdu, use of subsidized platforms, and an expert-curated knowledge base. We then deployed the chatbot with 13 participants for two weeks (403 messages + interviews). Women used it to challenge cultural taboos, legitimize health concerns often dismissed as “normal”, and build reproductive health knowledge through iterative questioning. Yet, interactions also exposed tensions: reliance on cultural explanatory models, questions of trust and validation, and gendered persona of the chatbot itself. We contribute empirical insights, a stigma-aware design framework for culturally sensitive conversational AI, and a methodological lens foregrounding expert validation in intimate health domains.
Amna Shahnawaz, Ayesha Shafique, Ding Wang 0006, Maryam Mustafa
CHI3
2026 Clue and Context Fusion for Sarcasm Detection with Large Multimodal Models
abstract
Detecting sarcasm in social media is fundamentally different from general VLM benchmarks: it is a pragmatic contradiction problem in which the literal signal in one modality is intentionally misaligned with the intended meaning, while dominant pre-training (e.g., CLIP-style contrastive agreement) biases models toward modality alignment rather than incongruity detection. We present SCARF, a contradiction-aware framework that equips large multimodal models with explicit sarcasm cues and context-sensitive retrieval. SCARF constructs coarse scene cues and fine localized evidence via tag-constrained QA, then distills them with visual tokens into a [FUSION] control vector for the LLM; a label-contrastive retriever supplies type- and context-matched exemplars, and a local multi-view encoder surfaces micro-cues. With the same backbone and training data, SCARF attains 87.92% Acc/86.67% F1 on MMSD2.0 and 77.14% Acc/76.44% F1 zero-shot on XDMSD, outperforming a comparably fine-tuned LLaVA-1.5. Ablations show sarcasm clue fusion is the main driver of gains, and tag-constrained QA improves rationale grounding and reduces hallucinations.
Yushan Pan, Ding Wang 0006, Wei Wang 0042, Xiaowei Huang 0001, Zhijie Xu
ACM Trans. Intell. Syst. Technol.3
2025 Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
abstract
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful.
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo
NeurIPS5
2024 The Problems with Proxies: Making Data Work Visible through Requester Practices
abstract
Fairness in AI and ML systems is increasingly linked to the proper treatment and recognition of data workers involved in training dataset development. Yet, those who collect and annotate the data, and thus have the most intimate knowledge of its development, are often excluded from critical discussions. This exclusion prevents data annotators, who are domain experts, from contributing effectively to dataset contextualization. Our investigation into the hiring and engagement practices of 52 data work requesters on platforms like Amazon Mechanical Turk reveals a gap: requesters frequently hold naive or unchallenged notions of worker identities and capabilities and rely on ad-hoc qualification tasks that fail to respect the workers’ expertise. These practices not only undermine the quality of data but also the ethical standards of AI development. To rectify these issues, we advocate for policy changes to enhance how data annotation tasks are designed and managed and to ensure data workers are treated with the respect they deserve.
Annabel Rothschild, Ding Wang 0006, Niveditha Jayakumar Vilvanathan, Lauren Wilcox, Carl F. DiSalvo, Betsy James DiSalvo
AIES (1)2
2024 GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives
abstract
Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex Taylor, Mark Diaz, Ding Wang, Gregory Serapio-García. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Vinodkumar Prabhakaran, Christopher Homan, Lora Aroyo, Aida Mostafazadeh Davani, Alicia Parrish, Alex S. Taylor, Mark Diaz, Ding Wang 0006, Gregory Serapio-García
NAACL-HLT8
2024 Making Data Work Count
abstract
In this paper, we examine the work of data annotation. Specifically, we focus on the role of counting or quantification in organising annotation work. Based on an ethnographic study of data annotation in two outsourcing centres in India, we observe that counting practices and its associated logics are an integral part of day-to-day annotation activities. In particular, we call attention to the presumption of total countability observed in annotation - the notion that everything, from tasks, datasets and deliverables, to workers, work time, quality and performance, can be managed by applying the logics of counting. To examine this, we draw on sociological and socio-technical scholarship on quantification and develop the lens of a 'regime of counting' that makes explicit the specific counts, practices, actors and structures that underpin the pervasive counting in annotation. We find that within the AI supply chain and data work, counting regimes aid the assertion of authority by the AI clients (also called requesters) over annotation processes, constituting them as reductive, standardised, and homogenous. We illustrate how this has implications for i) how annotation work and workers get valued, ii) the role human discretion plays in annotation, and iii) broader efforts to introduce accountable and more just practices in AI. Through these implications, we illustrate the limits of operating within the logic of total countability. Instead, we argue for a view of counting as partial - located in distinct geographies, shaped by specific interests and accountable in only limited ways. This, we propose, sets the stage for a fundamentally different orientation to counting and what counts in data annotation.
Srravya Chandhiramowuli, Alex S. Taylor, Sara Heitlinger, Ding Wang 0006
Proc. ACM Hum. Comput. Interact.4
2023 A hunt for the Snark: Annotator Diversity in Data Practices
abstract
Diversity in datasets is a key component to building responsible AI/ML. Despite this recognition, we know little about the diversity among the annotators involved in data production. We investigated the approaches to annotator diversity through 16 semi-structured interviews and a survey with 44 AI/ML practitioners. While practitioners described nuanced understandings of annotator diversity, they rarely designed dataset production to account for diversity in the annotation process. The lack of action was explained through operational barriers: from the lack of visibility in the annotator hiring process, to the conceptual difficulty in incorporating worker diversity. We argue that such operational barriers and the widespread resistance to accommodating annotator diversity surface a prevailing logic in data practices—where neutrality, objectivity and ‘representationalist thinking’ dominate. By understanding this logic to be part of a regime of existence, we explore alternative ways of accounting for annotator subjectivity and diversity in data practices.
Shivani Kapania, Alex S. Taylor, Ding Wang 0006
CHI3
2023 DICES Dataset: Diversity in Conversational AI Evaluation for Safety
abstract
Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This requirement overly simplifies the natural subjectivity present in many tasks, and obscures the inherent diversity in human perceptions and opinions about many content items. Preserving the variance in content and diversity in human perceptions in datasets is often quite expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is socio-culturally situated in this context. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographics information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. The DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of safety for conversational AI. We further describe a set of metrics that show how rater diversity influences safety perception across different geographic regions, ethnicity groups, age groups, and genders. The goal of the DICES dataset is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.
Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher Homan, Alicia Parrish, Gregory Serapio-García, Vinodkumar Prabhakaran, Ding Wang 0006
NeurIPS8
2022 Whose AI Dream? In search of the aspiration in data annotation
abstract
Data is fundamental to AI/ML models. This paper investigates the work practices concerning data annotation as performed in the industry, in India. Previous human-centred investigations have largely focused on annotators’ subjectivity, bias and efficiency. We present a wider perspective of the data annotation: following a grounded approach, we conducted three sets of interviews with 25 annotators, 10 industry experts and 12 ML/AI practitioners. Our results show that the work of annotators is dictated by the interests, priorities and values of others above their station. More than technical, we contend that data annotation is a systematic exercise of power through organizational structure and practice. We propose a set of implications for how we can cultivate and encourage better practice to balance the tension between the need for high quality data at low cost and the annotators’ aspiration for well-being, career perspective, and active participation in building the AI dream.
Ding Wang 0006, Shantanu Prabhat, Nithya Sambasivan
CHI1
2020 Making Chat at Home in the Hospital: Exploring Chat Use by Nurses
abstract
In this paper, we examine WhatsApp use by nurses in India. Globally, personal chat apps have taken the workplace by storm, and healthcare is no exception. In the hospital setting, this raises questions around how chat apps are integrated into hospital work and the consequences of using such personal tools for work. To address these questions, we conducted an ethnographic study of chat use in nurses' work in a large multi-specialty hospital. By examining how chat is embedded in the hospital, rather than focusing on individual use of personal tools, we throw new light on the adoption of personal tools at work — specifically what happens when such tools are adopted and used as though they were organisational tools. In doing so, we explicate their impact on invisible work [77] and the creep of work into personal time, as well as how hierarchy and power play out in technology use. Thus, we point to the importance of looking beyond individual adoption by knowledge workers when studying the impact of personal tools at work.
Naveena Karusala, Ding Wang 0006, Jacki O'Neill
CHI2
2020 Please Call the Specialism: Using WeChat to Support Patient Care in China
abstract
We examine how WeChat has been adopted to support nurse-patient communication in an IVF clinic in China. In this setting, the biggest challenge to delivering high-quality patient-centred care is the large number of patients. Nurses typically spend less than five minutes with each patient during clinical visits. To compensate for such minimal in-person consultation, nurse-facilitated patient groups were created on WeChat, to extend medical care and facilitate peer support. Through an ethnographic study, we examined how these groups fit into the clinic's communication ecosystem, and the challenges they raise for nurse-facilitators who receive thousands of messages daily. We propose a set of design suggestions aiming to make the work of the nurse-facilitator easier and more effective. In highlighting the opportunities and challenges of using chat to extend care beyond the clinic, we contribute to a burgeoning discussion of how chat can support patient care in the Global South.
Ding Wang 0006, Santosh D. Kale, Jacki O'Neill
CHI1