Emily Chen

dblp:136/8702 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Instruction Composition for Automated LLM Red-Teaming
abstract
Various approaches to LLM red-teaming leverage an attacker LLM to discover jailbreaks against a target.Several of them task attacker with formulating its own strategies to transform harmful queries into jailbreaks through trial and error, resulting in a semantically limited range of successful attacks.Another recent approach discovers diverse attacks by combining crowdsourced queries and tactics within the attacker's instructions, but does so at random, limiting effectiveness.This article introduces a novel framework, ADAPTIVE INSTRUCTION COM-POSITION, that combines crowdsourced texts according to an adaptive mechanism trained to jointly optimize attack effectiveness with diversity.We use reinforcement learning to balance exploration with exploitation in a combinatorial space of instructions to guide the attacker toward diverse generations tailored to target vulnerabilities.We show that our strategy substantially outperforms random combination on a set of effectiveness and diversity metrics, even under model transfer.Further, we show that it surpasses a host of recent adaptive approaches on the public benchmark Harmbench.We employ a lightweight neural contextual bandit that adapts to contrastively pretrained embeddings, and provide ablations to suggest that the contrastive property enables the network to generalize and scale to the massive space.Warning: this article discusses malicious content and methods for generating it using LLMs.
Jesse Zymet, Andy Luo, Swapnil Shinde, Sahil Wadhwa, Emily Chen
ACL (1)5
2026 Change is Hard: Consistent Player Behavior Across Games with Conflicting Incentives
abstract
This paper examines how player flexibility – a player’s willingness to engage in a breadth of options or specialize – manifests across two gaming environments: League of Legends (League) and Teamfight Tactics (TFT). We analyze the gameplay decisions of 4,830 players who have played at least 50 competitive games in both titles and explore cross-game dynamics of behavior retention and consistency. Our work introduces a novel cross-game analysis that tracks the same players’ behavior across two different environments, reducing self-selection bias. Our findings reveal that while games incentivize different behaviors (specialization in League versus flexibility in TFT) for performance-based success, players exhibit consistent behavior across platforms. This study contributes to long-standing debate about agency versus structure, showing individual agency may be more predictive of cross-platform behavior than game-imposed structure in competitive settings. These insights offer implications for game developers, designers and researchers interested in building systems to promote behavior change.
Emily Chen, Alexander J. Bisberg, Dmitri Williams, Magy Seif El-Nasr, Emilio Ferrara
CHI1
2026 Extending STRIVE to World of Tanks: A Cross-Game Validation of a Socio-behavioral Player Taxonomy
abstract
STRIVE is a taxonomy for multiplayer games organized around player behavior features such as sociality, communication, and experience. We extend STRIVE from Sky: Children of the Light, a relationship-focused social game, to World of Tanks (WoT), a game centered on team combat, player clans, and competitive performance. Using three player data snapshots from 2020 and the original STRIVE framework, we recover four stable player types in WoT: Veterans, Socialites, Squad Players, and Newbies. These segments persist across time and align with the original STRIVE dimensions, while adding a WoT-specific performance dimension. Prediction experiments illustrate that future battle participation varies in predictability across player types: Veterans and Squad Players are consistently more predictable than Socialites and Newbies. Together, these findings support STRIVE as a generalizable framework for comparing social play across multiple games.
Alexander J. Bisberg, Emily Chen, Dmitri Williams, Emilio Ferrara
FDG2
2025 GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
abstract
We address the problem of data scarcity in harmful text classification for guardrailing applications and introduce GRAID (Geometric and Reflective AI-Driven Data Augmentation), a novel pipeline that leverages Large Language Models (LLMs) for dataset augmentation.GRAID consists of two stages: (i) generation of geometrically controlled examples using a constrained LLM, and (ii) augmentation through a multi-agentic reflective process that promotes stylistic diversity and uncovers edge cases.This combination enables both reliable coverage of the input space and nuanced exploration of harmful content.Using two benchmark data sets, we demonstrate that augmenting a harmful text classification dataset with GRAID leads to significant improvements in downstream guardrail model performance.Warning: This paper contains techniques to synthetically generate offensive and malicious content using LLMs.
Melissa Kazemi Rad, Alberto Purpura, Emily Chen, Mohammad Shahed Sorower
EMNLP4
2025 STRIVE: Socio-behavioral Taxonomy Representation for Interactive Virtual Environments
abstract
We introduce STRIVE, a framework to build socio-behavioral taxonomies in multiplayer online games using unsupervised learning on common features across many social games.This work demonstrates this framework on "Sky: Children of the Light," a social adventure game by Thatgamecompany, using features such as cooperative play, social bonding, and in-game communication.After performing descriptive statistics, clustering and dimensionality reduction, we assign semantic categories to these behavior clusters.Next we perform a behavior prediction experiment where the most social cluster's behavior has a higher correlation with future play time and chats sent than predicting on the full dataset.These results suggest the importance of customized, or personalized, player behavior prediction models.Moreover, this framework could be easily extended to other games and further augment our understanding of human behavior in virtual worlds, ultimately aiding social scientists and game designers to better match players together for healthier online interactions.
Alexander J. Bisberg, Emily Chen, Marlon Twyman, Dmitri Williams, Emilio Ferrara
FDG2
2025 Treebeard: A Scalable and Fault Tolerant ORAM Datastore
Amin Setayesh, Cheran Mahalingam, Emily Chen, Sujaya Maiyya
USENIX Security Symposium3
2025 Communication Patterns Predict Team Skill in Multiplayer Online Games
abstract
The present research on team collaboration is typically performed through qualitative interview based studies or social network measurements of connectedness through co-play. In this study, we take the unique approach to build networks from direct messages between players in the massive online game World of Tanks where players self-organize into clans with specific roles assigned from military rankings (from Private to Commander). We explore the relationship between team communication volume and skill level, the impact of communication features on clan rating, and the differences in communication hierarchy between high and low-rated clans. Our findings reveal that higher-rated clans send more pre-battle chat messages, suggesting that effective communication and strategic planning are key to team performance. Evidence shows teams who use voice chat during battle are significantly higher ranked. Finally, we reveal that the highest rated clans have more connected lower-ranked members emphasizing that these teams are ''only as strong as their weakest link.'' This research is guided by the Transactive Memory Systems and Collective Intelligence theories which serve to expand the contribution of this research outside of games to other forms of virtual collaboration.
Alexander J. Bisberg, Sonia Jawaid Shaikh, Yilei Zeng, Fred Morstatter, Emily Chen, Emilio Ferrara, Dmitri Williams
Proc. ACM Hum. Comput. Interact.5
2024 "Can You Play Anything Else?" Understanding Play Style Flexibility in League of Legends
abstract
This study investigates the concept of flexibility within League of Legends, a popular online multiplayer game, focusing on the relationship between user adaptability and team success. Utilizing a dataset encompassing players of varying skill levels and play styles, we calculate two measures of flexibility for each player: overall flexibility and temporal flexibility. Our findings suggest that the flexibility of a user is dependent upon a user’s preferred play style, and flexibility does impact match outcome. This work also shows that skill level not only indicates how willing a player is to adapt their play style but also how their adaptability changes over time. This paper highlights the duality and balance of specialization versus flexibility, providing insights that can inform strategic planning, collaboration and resource allocation in competitive environments.
Emily Chen, Alexander J. Bisberg, Emilio Ferrara
CoG1
2024 What's in your PIE? Understanding the contents of personalized information environments with PIEGraph
abstract
Abstract Social media have long been studied from platform‐centric perspectives, which entail sampling messages based on criteria such as keywords and specific accounts. In contrast, user‐centric approaches attempt to reconstruct the personalized information environments users create for themselves. Most user‐centric studies analyze what users have accessed directly through browsers (e.g., through clicks) rather than what they may have seen in their social media feeds. This study introduces a data collection system of our own design called PIEGraph that links survey data with posts collected from participants' personalized X (formerly known as Twitter) timelines. Thus, in contrast with previous research, our data include much more than what users decide to click on. We measure the total amount of data in our participants' respective feeds and conduct descriptive and inferential analyses of three other quantities of interest: political content, ideological skew, and fact quality ratings. Our results are relevant to ongoing debates about digital echo chambers, misinformation, and conspiracy theories; and our general methodological approach could be applied to social media beyond X/Twitter contingent on data availability.
Deen Freelon, Meredith Pruden, Daniel Malmer, Qunfang Wu, Yiping Xia, Emily Chen, Andrew Crist
J. Assoc. Inf. Sci. Technol.7
2023 Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War between Ukraine and Russia
abstract
On February 24, 2022, Russia invaded Ukraine. In the days that followed, reports kept flooding in from laymen to news anchors of a conflict quickly escalating into war. Russia faced immediate backlash and condemnation from the world at large. While the war continues to contribute to an ongoing humanitarian and refugee crisis in Ukraine, a second battlefield has emerged in the online space, both in the use of social media to garner support for both sides of the conflict and also in the context of information warfare. In this paper, we present a collection of nearly half a billion tweets, from February 22, 2022, through January 8, 2023, that we are publishing for the wider research community to use. This dataset can be found at https://github.com/echen102/ukraine-russia. Our preliminary analysis on a subset of our dataset already shows evidence of public engagement with Russian state-sponsored media and other domains that are known to push unreliable information towards the beginning of the war; the former saw a spike in activity on the day of the Russian invasion, while the other saw spikes in engagement within the first month of the war. Our hope is that this public dataset can help the research community to further understand the ever-evolving role that social media plays in information dissemination, influence campaigns, grassroots mobilization, and much more, during a time of conflict.
Emily Chen, Emilio Ferrara
ICWSM1
2022 Validating child-friendly neuroimaging language localizer in adults
Halie Olson, Emily Chen, Hana Ro, Somaia Saba, Kirsten Lydic, Rebecca Saxe
CogSci2
2022 The Gift that Keeps on Giving: Generosity is Contagious in Multiplayer Online Games
abstract
Understanding social interactions and generous behaviors have long been of considerable interest in the social sciences community. While the contagion of generosity is documented in the real world, less is known about such phenomenon in virtual worlds and whether it has an actionable impact on user behavior and retention. In this work, we analyze social dynamics in the virtual world of the popular massively multiplayer online role-playing game (MMORPG) Sky: Children of Light. We develop a framework to reveal the patterns of generosity in such social environments and provide empirical evidence of social contagion and contagious generosity. Players become more engaged in the game after playing with others and especially with friends. We also find that players who experience generosity first-hand or even observe other players conduct generous acts become more generous themselves in the future. Additionally, we show that both receiving and observing generosity lead to higher future engagement in the game. Since Sky resembles the real world from a social play aspect, the implications of our findings also go beyond this virtual world.
Alexander J. Bisberg, Julie Jiang, Yilei Zeng, Emily Chen, Emilio Ferrara
Proc. ACM Hum. Comput. Interact.4
2021 Pictures as a Form of Protest: A Survey and Analysis of Images Posted During the Stop Asian Hate Movement on Twitter
abstract
Modern protests are not limited to on-the-ground operations, and the ease and speed at which users can upload images to social media platforms has enabled protests to manifest online. Previous analysis of protest imagery from social media sites categorized these images into groups including texts, screenshots, memes, and artwork. However, large-scale manual annotation to identify different types of images is not feasible. By applying machine learning to a large Twitter dataset focused on the Stop Asian Hate movement, we found the type of image an account posted during protests on Twitter is tied to the credibility and political leaning of posted content, type of witnessing (remote or connective), and community formation.
Oliver Melbourne Allen, Emily Chen, Emilio Ferrara
MASS2
2021 EIT-kit: An Electrical Impedance Tomography Toolkit for Health and Motion Sensing
abstract
In this paper, we propose EIT-kit, an electrical impedance tomography toolkit for designing and fabricating health and motion sensing devices. EIT-kit contains (1) an extension to a 3D editor for personalizing the form factor of electrode arrays and electrode distribution, (2) a customized EIT sensing motherboard for performing the measurements, (3) a microcontroller library that automates signal calibration and facilitates data collection, and (4) an image reconstruction library for mobile devices for interpolating and visualizing the measured data. Together, these EIT-kit components allow for applications that require 2- or 4-terminal setups, up to 64 electrodes, and single or multiple (up to four) electrode arrays simultaneously.
Junyi Zhu 0001, Jackson C. Snowden, Joshua Verdejo, Emily Chen, Hamid Ghaednia, Joseph H. Schwab, Stefanie Mueller 0001
UIST4
2020 Improved Finite-State Morphological Analysis for St. Lawrence Island Yupik Using Paradigm Function Morphology
abstract
St. Lawrence Island Yupik is an endangered polysynthetic language of the Bering Strait region. While conducting linguistic fieldwork between 2016 and 2019, we observed substantial support within the Yupik community for language revitalization and for resource development to support Yupik education. To that end, Chen & Schwartz (2018) implemented a finite-state morphological analyzer as a critical enabling technology for use in Yupik language education and technology. Chen & Schwartz (2018) reported a morphological analysis coverage rate of approximately 75% on a dataset of 60K Yupik tokens, leaving considerable room for improvement. In this work, we present a re-implementation of the Chen & Schwartz (2018) finite-state morphological analyzer for St. Lawrence Island Yupik that incorporates new linguistic insights; in particular, in this implementation we make use of the Paradigm Function Morphology (PFM) theory of morphology. We evaluate this new PFM-based morphological analyzer, and demonstrate that it consistently outperforms the existing analyzer of Chen & Schwartz (2018) with respect to accuracy and coverage rate across multiple datasets.
Emily Chen, Hyunji Hayley Park, Lane Schwartz
LREC1
2019 Human Evaluation of Models Built for Interpretability
abstract
Recent years have seen a boom in interest in interpretable machine learning systems built on models that can be understood, at least to some degree, by domain experts. However, exactly what kinds of models are truly human-interpretable remains poorly understood. This work advances our understanding of precisely which factors make models interpretable in the context of decision sets, a specific class of logic-based model. We conduct carefully controlled human-subject experiments in two domains across three tasks based on human-simulatability through which we identify specific types of complexity that affect performance more heavily than others-trends that are consistent across tasks and domains. These results can inform the choice of regularizers during optimization to learn more interpretable models, and their consistency suggests that there may exist common design principles for interpretable machine learning systems.
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel Gershman, Finale Doshi-Velez
HCOMP2
2018 A Morphological Analyzer for St. Lawrence Island / Central Siberian Yupik
Emily Chen, Lane Schwartz
LREC1