VLDB 2026 Research / reviewers in the wild / expert
Xiangyu Zhou 0001
dblp:187/3858-1
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0000-0776-8045ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trauma-Informed Data Donation: Integrating Expert and Donor Perspectives on Designing Against Re-Traumatization During Collection of Sexual Violence DataabstractData donation has received attention as a more consensual means of collecting personal data for scientific inquiry and AI technology. Yet the nature of data often donated–such as harmful online messages and menstrual tracking logs–carries risk of retraumatization (the forced reliving of traumatic experience). While the well-being of data donors is considered in prior work, approaches to retraumatization remain ad hoc. We present Trauma-Informed Data Donation (TIDD): a context-specific, exploratory design framework for adapting the Trauma-Informed Approach (TIA) from the Public Health domain to data donation. TIDD was the product of a 2-year research through design process with experts on sexual violence and trauma, and observational interviews of data donors. We use a case study applying TIDD to our custom data donation platform, Ube, as an invitation for designers to consider how TIDD could be used as a malleable foundation for donation of data associated with other forms of trauma. Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko |
DIS | 4 |
| 2026 | Not All Tokens Are Meant to Be ForgottenabstractLarge Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they tend to memorize unwanted information, such as private or copyrighted content, raising significant privacy and legal concerns. Unlearning has emerged as a promising solution, but existing methods face a significant challenge of over-forgetting. This issue arises because they indiscriminately suppress the generation of all the tokens in forget samples, leading to a substantial loss of model utility. To overcome this challenge, we introduce the Targeted Information Forgetting (TIF) framework, which consists of (1) a flexible targeted information identifier designed to differentiate between unwanted words (UW) and general words (GW) in the forget samples, and (2) a novel Targeted Preference Optimization approach that leverages Logit Preference Loss to unlearn unwanted information associated with UW and Preservation Loss to retain general information in GW, effectively improving the unlearning process while mitigating utility degradation. Extensive experiments on the TOFU and MUSE benchmarks demonstrate that the proposed TIF framework enhances unlearning effectiveness while preserving model utility and achieving state-of-the-art results. Xiangyu Zhou 0001, Yao Qiang, Saleh Zare Zade, Douglas Zytko, Prashant Khanduri, Dongxiao Zhu |
AAAI | 1 |
| 2025 | Automatic Calibration for Membership Inference Attack on Large Language ModelsabstractMembership Inference Attacks (MIAs) have recently been employed to determine whether a specific text was part of the pre-training data of Large Language Models (LLMs). However, existing methods often misinfer non-members as members, leading to a high false positive rate, or depend on additional reference models for probability calibration, which limits their practicality. To overcome these challenges, we introduce a novel framework called Automatic Calibration Membership Inference Attack (ACMIA), which utilizes a tunable temperature to calibrate output probabilities effectively. This approach is inspired by our theoretical insights into maximum likelihood estimation during the pre-training of LLMs. We introduce ACMIA in three configurations designed to accommodate different levels of model access and increase the probability gap between members and non-members, improving the reliability and robustness of membership inference. Extensive experiments on various open-source LLMs demonstrate that our proposed attack is highly effective, robust, and generalizable, surpassing state-of-the-art baselines across three widely used benchmarks. The source code is publicly available. Saleh Zare Zade, Yao Qiang, Xiangyu Zhou 0001, Hui Zhu 0016, Mohammad Amin Roshani, Prashant Khanduri, Dongxiao Zhu |
ECAI | 3 |
| 2025 | Collective Consent: Who Needs to Consent to the Donation of Data Representing Multiple People?abstractData donation is a growing form of personal data collection that foregrounds consent and conscious participation of the data donor. There remains little guidance on who must consent to data donation, particularly when the data represents multiple people. We provide empirical perspectives on this question through in-situ observation and interviews (N=18) with online daters who chose to donate messaging interactions with potential sexual partners for sexual violence research. Findings elucidate two diverging perspectives. Participants advocating for ''unilateral consent'' argued that consent of their messaging partner is not necessary, in part, because the anticipated benefit of data donation superseded consent. Participants advocating for ''collective consent'' wanted both messaging partners to consent to its donation, citing concerns for privacy of, and personal relationships with, the other person. Findings suggest that collective consent interfaces should be incorporated in data donation platforms, even if not strictly required by legal regulation, to improve donation of multi-person data. Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2024 | "It's Not What We Were Trying to Get At, but I Think Maybe It Should Be": Learning How to Do Trauma-Informed Design with a Data Donation Platform for Online Dating Sexual ViolenceabstractA majority of people experience trauma, spurring calls to incorporate trauma-informed approaches (TIA) from public health and social work into technology design. While technologies touted as trauma-informed are starting to propagate the literature, there persists a gap in knowledge around how design teams apply TIA and qualify their technology as adhering to trauma-informed principles. We address this through a 12-month development project with trauma and sexual violence experts to produce Ube, a data donation platform for collecting online dating sexual consent data to improve sexual risk detection AI. Through analysis of design documentation we retrospectively articulate a trauma-informed design process that evolved through the course of Ube’s development, comprising three elements for integrating trauma-informed principles: design goals that adapt the definition of TIA to the application domain, design activities that map to trauma-informed principles, and consequent design choices. We conclude with methodological recommendations to improve trauma-informed design processes. Emma Walquist, Isha Datey, Xiangyu Zhou 0001, Kelly Berishaj, Melissa Mcdonald, Michele Parkhill, Dongxiao Zhu, Douglas Zytko |
CHI | 4 |