Ruyuan Wan

dblp:276/9656 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-0357-5139ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 "Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews
abstract
Coded language is an important part of human communication.It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood.Current language models handle coded language poorly.Progress has been limited by the lack of real-world datasets and clear taxonomies.This paper introduces CODEDLANG, a dataset of 7,744 Chinese Google Maps reviews, including 900 reviews with span-level annotations of coded language.We developed a seven-class taxonomy that captures common encoding strategies, including phonetic, orthographic, and cross-lingual substitutions.We benchmarked language models on coded language detection, classification, and review rating prediction.Results show that even strong models can fail to identify or understand coded language.Because many coded expressions rely on pronunciation-based strategies, we further conducted a phonetic analysis of coded and decoded forms.Our code and dataset are publicly available 1 .Together, our results highlight coded language as an important and underexplored challenge for real-world NLP systems.
Ruyuan Wan, Changye Li 0001, Ting-Hao 'Kenneth' Huang
ACL (1)1
2026 Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
abstract
Recently, red teaming, with roots in security, has become a key evaluative approach to ensure the safety and reliability of Generative Artificial Intelligence. However, most existing work emphasizes technical benchmarks and attack success rates, leaving the socio-technical practices of how red teaming datasets are defined, created, and evaluated under-examined. Drawing on 22 interviews with practitioners who design and evaluate red teaming datasets, we examine the data practices and standards that underpin this work. Because adversarial datasets determine the scope and accuracy of model evaluations, they are critical artifacts for assessing potential harms from large language models. Our contributions are first, empirical evidence of practitioners conceptualizing red teaming and developing and evaluating red teaming datasets. Second, we reflect on how practitioners’ conceptualization of risk leads to overlooking the context, interaction type, and user specificity. We conclude with three opportunities for HCI researchers to expand the conceptualization and data practices for red-teaming.
Adriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla A. Badillo-Urquiola
CHI2
2025 Hashtag Re-Appropriation for Audience Control on Recommendation-Driven Social Media Xiaohongshu (rednote)
abstract
Algorithms have played a central role in personalized recommendations on social media. However, they also present significant obstacles for content creators trying to predict and manage their audience reach. This issue is particularly challenging for marginalized groups seeking to maintain safe spaces. Our study explores how women on Xiaohongshu (rednote), a recommendation-driven social platform, proactively re-appropriate hashtags (e.g., #Baby Supplemental Food) by using them in posts unrelated to their literal meaning. The hashtags were strategically chosen from topics that would be uninteresting to the male audience they wanted to block. Through a mixed-methods approach, we analyzed the practice of hashtag re-appropriation based on 5,800 collected posts and interviewed 24 active users from diverse backgrounds to uncover users' motivations and reactions towards the re-appropriation. This practice highlights how users can reclaim agency over content distribution on recommendation-driven platforms, offering insights into self-governance within algorithmic-centered power structures.
Ruyuan Wan, Lingbo Tong, Tiffany Knearem, Toby Jia-Jun Li, Ting-Hao 'Kenneth' Huang, Qunfang Wu
CHI1
2024 CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent Agents
abstract
In recent years, there has been a growing interest in employing intelligent agents in writing. Previous work emphasizes the evaluation of the quality of end product—whether it was coherent and polished, overlooking the journey that led to the product, which is an invaluable dimension of the creative process. To understand how to recognize human efforts in co-writing with intelligent writing systems, we adapt Flower and Hayes’ cognitive process theory of writing and propose CoCo Matrix, a two-dimensional taxonomy of entropy and information gain, to depict the new human-agent co-writing model. We define four quadrants and situate thirty-four published systems within the taxonomy. Our research found that low entropy and high information gain systems are under-explored, yet offer promising future directions in writing tasks that benefit from the agent’s divergent planning and the human’s focused translation. CoCo Matrix, not only categorizes different writing systems but also deepens our understanding of the cognitive processes in human-agent co-writing. By analyzing minimal changes in the writing process, CoCo Matrix serves as a proxy for the writer’s mental model, allowing writers to reflect on their contributions. This reflection is facilitated through the measured metrics of information gain and entropy, which provide insights irrespective of the writing system used.
Ruyuan Wan, Simret Araya Gebreegziabher, Toby Jia-Jun Li, Karla A. Badillo-Urquiola
Creativity & Cognition1
2024 Tricky vs. Transparent: Towards an Ecologically Valid and Safe Approach for Evaluating Online Safety Nudges for Teens
abstract
HCI research has been at the forefront of designing interventions for protecting teens online; yet, how can we test and evaluate these solutions without endangering the youth we aim to protect? Towards this goal, we conducted focus groups with 20 teens to inform the design of a social media simulation platform and study for evaluating online safety nudges co-designed with teens. Participants evaluated risk scenarios, personas, platform features, and our research design to provide insight regarding the ecological validity of these artifacts. Teens expected risk scenarios to be subtle and tricky, while also higher in risk to be believable. The teens iterated on the nudges to prioritize risk prevention without reducing autonomy, risk coping, and community accountability. For the simulation, teens recommended using transparency with some deceit to balance realism and respect for participants. Our meta-level research provides a teen-centered action plan to evaluate online safety interventions safely and effectively.
Zainab Agha, Jinkyung Park, Ruyuan Wan, Naima Samreen Ali, Dominic DiFranzo, Karla A. Badillo-Urquiola, Pamela J. Wisniewski
CHI3
2024 CoCoLoFa: A Dataset of News Comments with Common Logical Fallacies Written by LLM-Assisted Crowds
abstract
Detecting logical fallacies in texts can help users spot argument flaws, but automating this detection is not easy.Manually annotating fallacies in large-scale, real-world text data to create datasets for developing and validating detection models is costly.This paper introduces COCOLOFA, the largest known English logical fallacy dataset, containing 7,706 comments for 648 news articles, with each comment labeled for fallacy presence and type.We recruited 143 crowd workers to write comments embodying specific fallacy types (e.g., slippery slope) in response to news articles.Recognizing the complexity of this writing task, we built an LLM-powered assistant into the workers' interface to aid in drafting and refining their comments.Experts rated the writing quality and labeling validity of COCOLOFA as high and reliable.BERT-based models fine-tuned using COCOLOFA achieved the highest fallacy detection (F1=0.86)and classification (F1=0.87)performance on its test set, outperforming the stateof-the-art LLMs.Our work shows that combining crowdsourcing and LLMs enables us to more effectively construct datasets for complex linguistic phenomena that crowd workers find challenging to produce on their own.COCOLOFA is public at CoCoLoFa.org/.
Min-Hsuan Yeh, Ruyuan Wan, Ting-Hao Huang
EMNLP2
2023 Everyone's Voice Matters: Quantifying Annotation Disagreement Using Demographic Information
abstract
In NLP annotation, it is common to have multiple annotators label the text and then obtain the ground truth labels based on major annotators’ agreement. However, annotators are individuals with different backgrounds and various voices. When annotation tasks become subjective, such as detecting politeness, offense, and social norms, annotators’ voices differ and vary. Their diverse voices may represent the true distribution of people’s opinions on subjective matters. Therefore, it is crucial to study the disagreement from annotation to understand which content is controversial from the annotators. In our research, we extract disagreement labels from five subjective datasets, then fine-tune language models to predict annotators’ disagreement. Our results show that knowing annotators’ demographic information (e.g., gender, ethnicity, education level), in addition to the task text, helps predict the disagreement. To investigate the effect of annotators’ demographics on their disagreement level, we simulate different combinations of their artificial demographics and explore the variance of the prediction to distinguish the disagreement from the inherent controversy from text content and the disagreement in the annotators’ perspective. Overall, we propose an innovative disagreement prediction mechanism for better design of the annotation process that will achieve more accurate and inclusive results for NLP systems. Our code and dataset are publicly available.
Ruyuan Wan, Jaehyung Kim 0001, Dongyeop Kang
AAAI1