VLDB 2026 Research / reviewers in the wild / expert
Rachel Greenstadt
dblp:93/655
· DBLP profile ↗
55ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0002-1831-1785ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 36 · 5 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What's in a Label? Propaganda Labels and User Sharing Behavior on Social Media PlatformsabstractAuthentic information is vital for a society's ability to make rational decisions. Fabricated and manipulative information can be harmful to society as seen in cases of threatening events that were consequences of foreign propaganda and radical ideologies. While past research has studied dis- and misinformation on social media platforms, the study of propaganda has received much less attention. This study explores the sharing intentions of propaganda on social media platforms and develops an intervention to help detect it. In a randomized controlled trial setting, we added indicators to social media posts that used propaganda techniques to advance an agenda, including techniques that rely on fallacious reasoning, emotional rather than logical reasoning, etc. We then asked our participants (n=1,187) about their intention to engage with these posts. We found that participants were significantly (2.4 times) less likely to share these posts with indicators. We also found that participants’ political affiliation moderated their sharing intentions. We believe our findings provide valuable insights for the study of propaganda on social media platforms. Julia Jose, Chris Geeng, Kediel O. Morales, Damon McCoy, Rachel Greenstadt |
ICWSM | 5 |
| 2025 | An Analysis of Chinese Censorship Bias in LLMsabstractWhen a large language model (LLM) has been trained on text featuring social biases, those biases implicitly impact the outputs of the model. Training an LLM on sanitized content, i.e., those pieces of content which remain after being subjected to state censorship (including alterations, deletions, and self-imposed censorship), results in what we term censorship bias. A model impacted by censorship bias may be less likely to reflect views that are routinely prohibited and more likely to reflect views that are not. This may particularly be an issue when interfacing with a model in a language that is predominantly used in a region with strong censorship laws. In this work, we outline what censorship bias is, introduce a novel methodology for identifying and measuring it, and apply that methodology to evaluate the most popular current LLMs. As part of the contributions of this work we designed and evaluated CensorshipDetector, a Chinese language text classification model which we use as part of our experimental design. Our evaluation of CensorshipDetector found it to be 91% accurate at differentiating between sanitized content and non-sanitized content. Our testing revealed evidence of censorship bias across all of the models we evaluated. Finally, we outline the potential harms of censorship bias, namely the exportation of information manipulation that would have primarily harmed a domestic audience to diaspora, as well as recommendations to various stakeholders to limit the harms of censorship bias and prevent it in the future. Jeffrey Knockel, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 3 |
| 2025 | More and Scammier Ads: The Perils of YouTube's Ad Privacy SettingsabstractWhen users disable online ad personalization, they might be anticipating to see fewer ads that are "relevant" to them as a trade-off for more privacy. In this paper, we show that the tradeoff can go much further than this intuition. We conducted controlled experiments on YouTube in Australia, Canada, Ireland, the United Kingdom, and the United States to investigate the impact of disabling ad personalization on the quantity and quality of ads that users receive. Through experiments where emulated users with different ad privacy settings watched sequences of 400 videos, we show that disabling ad personalization can lead to the user being shown as much as 1.30 times more pre-roll ads than the default (least private) setting. More concerning is that in our experiments, the proportion of predatory ads increased 2.69 times compared to the default setting, from 2.5% to 8.7% of ads. This result highlights that certain user demographics (in this case, privacy-conscious users) can be exposed to significantly higher rates of predatory ads, and suggests that the platform's efforts to curb such ads are still falling short. Cat Mai, Bruno Coelho, Julia B. Kieserman, Lexie Matsumoto, Kyle Spinelli, Eric Yang, Athanasios Andreou, Rachel Greenstadt, Tobias Lauinger, Damon McCoy |
Proc. Priv. Enhancing Technol. | 8 |
| 2025 | Sheep's clothing, wolfish intent: Automated detection and evaluation of problematic 'allowed' advertisementsabstractThe digital advertising ecosystem sustains the free web and drives global innovation, but often at the cost of user privacy through intrusive tracking and non-compliant ads, especially harmful to under-age users. This has led to widespread adoption of privacy tools like adblockers and anti-trackers, which, while disrupting ad revenues, expose users to alternate forms of tracking and fingerprinting. To address this, many adblockers now allow 'non-intrusive' ads by default. In this study, we evaluate Adblock Plus's Acceptable Ads feature and find a 13.6% increase in problematic ads compared to no adblocker use—challenging claims of improved user experience. We also find that ad exchanges on allowlists are more likely to serve problematic content, underscoring the hidden cost privacy-aware users pay when relying on such technologies. While prior work in the domain has been limited by their practical viability, we further propose a methodology to automate the detection of problematic ads using LLMs with zero-shot prompting, achieving substantial agreement with human annotators (IAA score: 0.79). This establishes the efficacy of LLMs in problematic content detection under well-defined environments. As In-browser LLMs emerge, adversaries may exploit problematic ad content to fingerprint privacy-conscious ABP users. At the same time, these advances present new opportunities for adblockers to develop robust defenses, detect malicious exchanges, and uphold both user privacy and the sustainability of the ad-supported web. Ritik Roongta, Julia Jose, Hussam Habib, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 4 |
| 2024 | From User Insights to Actionable Metrics: A User-Focused Evaluation of Privacy-Preserving Browser ExtensionsabstractThe rapid growth of web tracking via advertisements has led to an increased adoption of privacy-preserving browser extensions. These extensions are crucial for blocking trackers and enhancing the overall web browsing experience. The advertising industry is constantly changing, leading to ongoing development and improvements in both new and existing ad-blocking and anti-tracking extensions. Despite this, there is a lack of comprehensive studies exploring the set of user concerns associated with these extensions. Our research addresses this gap by identifying five user concerns and establishing a privacy and usability topics framework, specific to privacy-preserving extensions. Ritik Roongta, Rachel Greenstadt |
AsiaCCS | 2 |
| 2024 | Stoking the Flames: Understanding Escalation in an Online Harassment CommunityabstractOnline harassment remains a prevalent problem for internet users. Its impact is made orders of magnitude worse when multiple harassers coordinate to conduct networked attacks. This paper presents an analysis of 231 threads in Kiwi Farms, a notorious online harassment community. We find that networked online harassment campaigns consists of three phases: target introduction, network decision, and network response. The first stage consists of the initial narrative elements, that are approved or not in stage two and expanded in stage three. Narrative building is a common element of all three stages. The network plays a key role in narrative building, adding elements to the narrative in at least 80 % of the threads, resulting in sustained harassment. This finding is central to our model of Continuous Narrative Escalation (CNE), that has two parts: (1) narrative continuation, the action of repeatedly adding new information to the existing narrative and (2) escalation, the aggravation of harassment that occurs as a consequence. In addition, we present insights from our analysis of 100 takedown requests threads, discussing received abuse reports. We find that these takedown requests are misused by the community and are used as elements to further fuel the narrative. We use our findings and framework to come up with a set of recommendations, that can inform harassment interventions and make online spaces safer. Kejsi Take, Victoria Zhong, Chris Geeng, Emmi Bevensee, Damon McCoy, Rachel Greenstadt |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2024 | Challenges in Restructuring Community-based ModerationabstractContent moderation practices and technologies need to change over time as requirements and community expectations shift. However, attempts to restructure the existing moderation practices can be difficult, especially for platforms that rely on their communities to moderate, because changes can transform the workflow and workload participants' reward systems. By examining the extensive archival discussions around a prepublication moderation technology on Wikipedia named Flagged Revisions, complemented by seven semi-structured interviews, we identify various challenges in restructuring community-based moderation practices. Thus, we find that while a new system might sound good in theory and perform well in terms of quantitative metrics, it may conflict with existing social norms. Furthermore, our findings underscore how the relationship between platforms and self-governed communities can hinder the ability to assess the performance of any new system and introduce considerable costs related to maintaining, overhauling, or scrapping any piece of infrastructure. Chau Tran, Kejsi Take, Kaylea Champion, Benjamin Mako Hill, Rachel Greenstadt |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | What to Expect When You're Accessing: An Exploration of User Privacy Rights in People Search WebsitesabstractPeople Search Websites, a category of data brokers, collect, catalog, monetize and often publicly display individuals' personally identifiable information (PII). We present a study of user privacy rights in 20 such websites assessing the usability of data access and data removal mechanisms. We combine insights from these two processes to determine connections between sites, such as shared access mechanisms or removal effects. We find that data access requests are mostly unsuccessful. Instead, sites cite a variety of legal exceptions or misinterpret the nature of the requests. By purchasing reports, we find that only one set of connected sites provided access to the same report they sell to customers. We leverage a multiple step removal process to investigate removal effects between suspected connected sites. In general, data removal is more streamlined than data access, but not very transparent; questions about the scope of removal and reappearance of information remain. Confirming and expanding the connections observed in prior phases, we find that four main groups are behind 14 of the sites studied, indicating the need to further catalog these connections to simplify removal. Kejsi Take, Jordyn Young, Rasika Bhalerao, Kevin Gallagher 0001, Andrea Forte, Damon McCoy, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 7 |
| 2023 | How Library IT Staff Navigate Privacy and Security Challenges and Responsibilities
Alan F. Luo, Noel Warford, Samuel Dooley, Rachel Greenstadt, Michelle L. Mazurek, Nora McDonald |
USENIX Security Symposium | 4 |
| 2023 | "I'm going to trust this until it burns me" Parents' Privacy Concerns and Delegation of Trust in K-8 Educational Technology
Victoria Zhong, Susan E. McGregor, Rachel Greenstadt |
USENIX Security Symposium | 3 |
| 2023 | Intersectional Thinking about PETs: A Study of Library PrivacyabstractThis qualitative study examines the privacy challenges perceived by librarians who afford access to physical and electronic spaces and are in a unique position of safeguarding the privacy of their patrons. As internet “service providers,” librarians represent a bridge between the physical and internet world, and thus offer a unique sight line to the convergence of privacy, identity, and social disadvantage. Drawing on interviews with 16 librarians, we describe how they often interpret or define their own rules when it comes to privacy to protect patrons who face challenges that stem from structures of inequality outside their walls. We adopt the term “intersectional thinking” to describe how librarians reported thinking about privacy solutions, which is focused on identity and threats of structural discrimination (the rules, norms, and other determinants of discrimination embedded in institutions and other societal structures that present barriers to certain groups or individuals), and we examine the role that low/no-tech strategies play in ameliorating these threats. We then discuss how librarians act as "privacy intermediaries" for patrons, the potential analogue for this role for developers of systems, the power of low/no-tech strategies, and implications for design and research of privacy-enhancing technologies (PETs). Nora McDonald, Rachel Greenstadt, Andrea Forte |
Proc. Priv. Enhancing Technol. | 2 |
| 2022 | Using Authorship Verification to Mitigate Abuse in Online Communities
Janith Weerasinghe, Rhia Singh, Rachel Greenstadt |
ICWSM | 3 |
| 2022 | Conspiracy Brokers: Understanding the Monetization of YouTube Conspiracy TheoriesabstractConspiracy theories are increasingly a subject of research interest as society grapples with their rapid growth in areas such as politics or public health. Previous work has established YouTube as one of the most popular sites for people to host and discuss different theories. In this paper, we present an analysis of monetization methods of conspiracy theorist YouTube creators and the types of advertisers potentially targeting this content. We collect 184,218 ad impressions from 6,347 unique advertisers found on conspiracy-focused channels and mainstream YouTube content. We classify the ads into business categories and compare their prevalence between conspiracy and mainstream content. We also identify common offsite monetization methods. In comparison with mainstream content, conspiracy videos had similar levels of ads from well-known brands, but an almost eleven times higher prevalence of likely predatory or deceptive ads. Additionally, we found that conspiracy channels were more than twice as likely as mainstream channels to use offsite monetization methods, and 53% of the demonetized channels we observed were linking to third-party sites for alternative monetization opportunities. Our results indicate that conspiracy theorists on YouTube had many potential avenues to generate revenue, and that predatory ads were more frequently served for conspiracy videos. Cameron Ballard, Ian Goldstein, Pulak Mehta, Genesis Smothers, Kejsi Take, Victoria Zhong, Rachel Greenstadt, Tobias Lauinger, Damon McCoy |
WWW | 7 |
| 2022 | The Risks, Benefits, and Consequences of Prepublication Moderation: Evidence from 17 Wikipedia Language EditionsabstractMany online communities rely on postpublication moderation where contributors-even those that are perceived as being risky-are allowed to publish material immediately and where moderation takes place after the fact. An alternative arrangement involves moderating content before publication. A range of communities have argued against prepublication moderation by suggesting that it makes contributing less enjoyable for new members and that it will distract established community members with extra moderation work. We present an empirical analysis of the effects of a prepublication moderation system called FlaggedRevs that was deployed by several Wikipedia language editions. We used panel data from 17 large Wikipedia editions to test a series of hypotheses related to the effect of the system on activity levels and contribution quality. We found that the system was very effective at keeping low-quality contributions from ever becoming visible. Although there is some evidence that the system discouraged participation among users without accounts, our analysis suggests that the system's effects on contribution volume and quality were moderate at most. Our findings imply that concerns regarding the major negative effects of prepublication moderation systems on contribution quality and project productivity may be overstated. Chau Tran, Kaylea Champion, Benjamin Mako Hill, Rachel Greenstadt |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | "It Feels Like Whack-a-mole": User Experiences of Data Removal from People Search WebsitesabstractPeople Search Websites aggregate and publicize users’ Personal Identifiable Information (PII), previously sourced from data brokers. This paper presents a qualitative study of the perceptions and experiences of 18 participants who sought information removal by hiring a removal service or requesting removal from the sites. The users we interviewed were highly motivated and had sophisticated risk perceptions. We found that they encountered obstacles during the removal process, resulting in a high cost of removal, whether they requested it themselves or hired a service. Participants perceived that the successful monetization of users PII motivates data aggregators to make the removal more difficult. Overall, self management of privacy by attempting to keep information off the internet is difficult and its’ success is hard to evaluate. We provide recommendations to users, third parties, removal services and researchers aiming to improve the removal process. Kejsi Take, Kevin Gallagher 0001, Andrea Forte, Damon McCoy, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 5 |
| 2021 | A large-scale characterization of online incitements to harassment across platformsabstractAttack strategies used by online harassers have evolved over time to inflict increasing harm to their targets. In addition to scaling harassment through incitement and coordination, online communities that commonly engage in harassment are likely a source of "innovation" for harassment attack strategies. We use the incitements or calls to harassment posted by members of these communities as a lens through which to holistically measure and understand this ecosystem. We create a filtering pipeline to discover 14,679 incitements to harassment within four large-scale data sets of messages and posts that span multiple platforms. Max Aliapoulios, Kejsi Take, Prashanth Ramakrishna, Daniel Borkan, Beth Goldberg, Jeffrey S. Sorensen, Anna Turner, Rachel Greenstadt, Tobias Lauinger, Damon McCoy |
Internet Measurement Conference | 8 |
| 2021 | Supervised Authorship Segmentation of Open Source Code ProjectsabstractAbstract Source code authorship attribution can be used for many types of intelligence on binaries and executables, including forensics, but introduces a threat to the privacy of anonymous programmers. Previous work has shown how to attribute individually authored code files and code segments. In this work, we examine authorship segmentation, in which we determine authorship of arbitrary parts of a program. While previous work has performed segmentation at the textual level, we attempt to attribute subtrees of the abstract syntax tree (AST). We focus on two primary problems: identifying the primary author of an arbitrary AST subtree and identifying on which edges of the AST primary authorship changes. We demonstrate that the former is a difficult problem but the later is much easier. We also demonstrate methods by which we can leverage the easier problem to improve accuracy for the harder problem. We show that while identifying the author of subtrees is difficult overall, this is primarily due to the abundance of small subtrees: in the validation set we can attribute subtrees of at least 25 nodes with accuracy over 80% and at least 33 nodes with accuracy over 90%, while in the test set we can attribute subtrees of at least 33 nodes with accuracy of 70%. While our baseline accuracy for single AST nodes is 20.21% for the validation set and 35.66% for the test set, we present techniques by which we can increase this accuracy to 42.01% and 49.21% respectively. We further present observations about collaborative code found on GitHub that may drive further research. Edwin Dauber, Robert F. Erbacher, Gregory Shearer, Michael J. Weisman, Frederica Free-Nelson, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 6 |
| 2020 | Are anonymity-seekers just like everybody else? An analysis of contributions to Wikipedia from TorabstractUser-generated content sites routinely block contributions from users of privacy-enhancing proxies like Tor because of a perception that proxies are a source of vandalism, spam, and abuse. Although these blocks might be effective, collateral damage in the form of unrealized valuable contributions from anonymity seekers is invisible. One of the largest and most important user-generated content sites, Wikipedia, has attempted to block contributions from Tor users since as early as 2005. We demonstrate that these blocks have been imperfect and that thousands of attempts to edit on Wikipedia through Tor have been successful. We draw upon several data sources and analytical techniques to measure and describe the history of Tor editing on Wikipedia over time and to compare contributions from Tor users to those from other groups of Wikipedia users. Our analysis suggests that although Tor users who slip through Wikipedia's ban contribute content that is more likely to be reverted and to revert others, their contributions are otherwise similar in quality to those from other unregistered participants and to the initial contributions of registered users. Chau Tran, Kaylea Champion, Andrea Forte, Benjamin Mako Hill, Rachel Greenstadt |
SP | 5 |
| 2020 | The Tools and Tactics Used in Intimate Partner Surveillance: An Analysis of Online Infidelity Forums
Emily Tseng, Rosanna Bellini, Nora McDonald, Matan Danos, Rachel Greenstadt, Damon McCoy, Nicola Dell, Thomas Ristenpart |
USENIX Security Symposium | 5 |
| 2020 | The Pod People: Understanding Manipulation of Social Media Popularity via Reciprocity AbuseabstractOnline Social Network (OSN) Users’ demand to increase their account popularity has driven the creation of an underground ecosystem that provides services or techniques to help users manipulate content curation algorithms. One method of subversion that has recently emerged occurs when users form groups, called pods, to facilitate reciprocity abuse, where each member reciprocally interacts with content posted by other members of the group. We collect 1.8 million Instagram posts that were posted in pods hosted on Telegram. We first summarize the properties of these pods and how they are used, uncovering that they are easily discoverable by Google search and have a low barrier to entry. We then create two machine learning models for detecting Instagram posts that have gained interaction through two different kinds of pods, achieving 0.91 and 0.94 AUC, respectively. Finally, we find that pods are effective tools for increasing users’ Instagram popularity, we estimate that pod utilization leads to a significantly increased level of likely organic comment interaction on users’ subsequent posts. Janith Weerasinghe, Bailey Flanigan, Aviel J. Stein, Damon McCoy, Rachel Greenstadt |
WWW | 5 |
| 2020 | "So-called privacy breeds evil": Narrative Justifications for Intimate Partner Surveillance in Online ForumsabstractA growing body of research suggests that intimate partner abusers use digital technologies to surveil their partners, including by installing spyware apps, compromising devices and online accounts, and employing social engineering tactics. However, to date, this form of privacy violation, called intimate partner surveillance (IPS), has primarily been studied from the perspective of victim-survivors. We present a qualitative study of how potential perpetrators of IPS harness the emotive power of sharing personal narratives to validate and legitimise their abusive behaviours. We analysed 556 stories of IPS posted on publicly accessible online forums dedicated to the discussion of sexual infidelity. We found that many users share narrative posts describing IPS as they boast about their actions, advise others on how to perform IPS without detection, and seek suggestions for next steps to take. We identify a set of common thematic story structures, justifications for abuse, and outcomes within the stories that provide a window into how these individuals believe their behaviour to be justified. Using these stories, we develop a four-stage framework that captures the change in a potential perpetrator's approach to IPS. We use our findings and framework to guide a discussion of efforts to combat abuse, including how we can identify crucial moments where interventions might be safely applied to prevent or deescalate IPS. Rosanna Bellini, Emily Tseng, Nora McDonald, Rachel Greenstadt, Damon McCoy, Thomas Ristenpart, Nicola Dell |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | Privacy, Anonymity, and Perceived Risk in Open Collaboration: A Study of Service ProvidersabstractAnonymity can enable both healthy online interactions like support-seeking and toxic behaviors like hate speech. How do online service providers balance these threats and opportunities? This two-part qualitative study examines the challenges perceived by open collaboration service providers in allowing anonymous contributions to their projects. We interviewed eleven people familiar with organizational decisions related to privacy and security at five open collaboration projects and followed up with an analysis of public discussions about anonymous contribution to Wikipedia. We contrast our findings with prior work on threats perceived by project volunteers and explore misalignment between policies aiming to serve contributors and the privacy practices of contributors themselves. Nora McDonald, Benjamin Mako Hill, Rachel Greenstadt, Andrea Forte |
CHI | 3 |
| 2019 | A Forensic Qualitative Analysis of Contributions to Wikipedia from Anonymity Seeking UsersabstractBy choice or by necessity, some contributors to commons-based peer production sites use privacy-protecting services to remain anonymous. As anonymity seekers, users of the Tor network have been cast both as ill-intentioned vandals and as vulnerable populations concerned with their privacy. In this study, we use a dataset drawn from a corpus of Tor edits to Wikipedia to uncover the character of Tor users' contributions. We build in-depth narrative descriptions of Tor users' actions and conduct a thematic analysis that places their editing activity into seven broad groups. We find that although their use of a privacy-protecting service marks them as unusual within Wikipedia, the character of many Tor users' contributions is in line with the expectations and norms of Wikipedia. However, our themes point to several important places where lack of trust promotes disorder, and to contributions where risks to contributors, service providers, and communities are unaligned. Kaylea Champion, Nora McDonald, Stephanie Bankes, Joseph Zhang, Rachel Greenstadt, Andrea Forte, Benjamin Mako Hill |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2019 | Git Blame Who?: Stylistic Authorship Attribution of Small, Incomplete Source Code FragmentsabstractAbstract Program authorship attribution has implications for the privacy of programmers who wish to contribute code anonymously. While previous work has shown that individually authored complete files can be attributed, these efforts have focused on such ideal data sets as contest submissions and student assignments. We explore the problem of authorship attribution “in the wild,” examining source code obtained from open-source version control systems, and investigate how contributions can be attributed to their authors, either on an individual or a per-account basis. In this work, we present a study of attribution of code collected from collaborative environments and identify factors which make attribution of code fragments more or less successful. For individual contributions, we show that previous methods (adapted to be applied to short code fragments) yield an accuracy of approximately 50% or 60%, depending on whether we average by sample or by author, at identifying the correct author out of a set of 104 programmers. By ensembling the classification probabilities of a sufficiently large set of samples belonging to the same author we achieve much higher accuracy for assigning the set of samples to the correct author from a known suspect set. Additionally, we propose the use of calibration curves to identify which samples are by unknown and previously unencountered authors. Edwin Dauber, Aylin Caliskan, Richard E. Harang, Gregory Shearer, Michael J. Weisman, Frederica Free-Nelson, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 7 |
| 2019 | "Because... I was told... so much": Linguistic Indicators of Mental Health Status on TwitterabstractAbstract Recent studies have shown that machine learning can identify individuals with mental illnesses by analyzing their social media posts. Topics and words related to mental health are some of the top predictors. These findings have implications for early detection of mental illnesses. However, they also raise numerous privacy concerns. To fully evaluate the implications for privacy, we analyze the performance of different machine learning models in the absence of tweets that talk about mental illnesses. Our results show that machine learning can be used to make predictions even if the users do not actively talk about their mental illness. To fully understand the implications of these findings, we analyze the features that make these predictions possible. We analyze bag-of-words, word clusters, part of speech n-gram features, and topic models to understand the machine learning model and to discover language patterns that differentiate individuals with mental illnesses from a control group. This analysis confirmed some of the known language patterns and uncovered several new patterns. We then discuss the possible applications of machine learning to identify mental illnesses, the feasibility of such applications, associated privacy implications, and analyze the feasibility of potential mitigations. Janith Weerasinghe, Kediel O. Morales, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 3 |
| 2018 | When Coding Style Survives Compilation: De-anonymizing Programmers from Executable Binaries
Aylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard E. Harang, Konrad Rieck, Rachel Greenstadt, Arvind Narayanan |
NDSS | 6 |
| 2018 | Editors' Introduction
Rachel Greenstadt, Damon McCoy, Carmela Troncoso |
Proc. Priv. Enhancing Technol. | 1 |
| 2018 | Editors' Introduction
Rachel Greenstadt, Damon McCoy, Carmela Troncoso |
Proc. Priv. Enhancing Technol. | 1 |
| 2018 | Editors' Introduction
Rachel Greenstadt, Damon McCoy, Carmela Troncoso |
Proc. Priv. Enhancing Technol. | 1 |
| 2018 | Editors' Introduction
Rachel Greenstadt, Damon McCoy, Carmela Troncoso |
Proc. Priv. Enhancing Technol. | 1 |
| 2017 | How Unique is Your .onion?: An Analysis of the Fingerprintability of Tor Onion ServicesabstractRecent studies have shown that Tor onion (hidden) service websites are particularly vulnerable to website fingerprinting attacks due to their limited number and sensitive nature. In this work we present a multi-level feature analysis of onion site fingerprintability, considering three state-of-the-art website fingerprinting methods and 482 Tor onion services, making this the largest analysis of this kind completed on onion services to date. Rebekah Overdorf, Marc Juarez, Gunes Acar, Rachel Greenstadt, Claudia Díaz |
CCS | 4 |
| 2017 | Privacy, Anonymity, and Perceived Risk in Open Collaboration: A Study of Tor Users and WikipediansabstractThis qualitative study examines privacy practices and concerns among contributors to open collaboration projects. We collected interview data from people who use the anonymity network Tor who also contribute to online projects and from Wikipedia editors who are concerned about their privacy to better understand how privacy concerns impact participation in open collaboration projects. We found that risks perceived by contributors to open collaboration projects include threats of surveillance, violence, harassment, opportunity loss, reputation loss, and fear for loved ones. We explain participants' operational and technical strategies for mitigating these risks and how these strategies affect their contributions. Finally, we discuss chilling effects associated with privacy loss, the need for open collaboration projects to go beyond attracting and educating participants to consider their privacy, and some of the social and technical approaches that could be explored to mitigate risk at a project or community level. Andrea Forte, Nazanin Andalibi, Rachel Greenstadt |
CSCW | 3 |
| 2017 | Source Code Authorship Attribution Using Long Short-Term Memory Based Networks
Bander Alsulami, Edwin Dauber, Richard E. Harang, Spiros Mancoridis, Rachel Greenstadt |
ESORICS (1) | 5 |
| 2017 | Using Stylometry to Attribute Programmers and WritersabstractIn this talk, I will discuss my lab's work in the emerging field of adversarial stylometry and machine learning. Machine learning algorithms are increasingly being used in security and privacy domains, in areas that go beyond intrusion or spam detection. For example, in digital forensics, questions often arise about the authors of documents: their identity, demographic background, and whether they can be linked to other documents. The field of stylometry uses linguistic features and machine learning techniques to answer these questions. We have applied stylometry to difficult domains such as underground hacker forums, open source projects (code), and tweets. I will discuss our Doppelgnger Finder algorithm, which enables us to group Sybil accounts on underground forums and detect blogs from Twitter feeds and reddit comments. In addition, I will discuss our work attributing unknown source code and binaries. Rachel Greenstadt |
IH&MMSec | 1 |
| 2017 | Editors' Introduction
Claudia Díaz, Rachel Greenstadt, Damon McCoy |
Proc. Priv. Enhancing Technol. | 2 |
| 2017 | Editors' Introduction
Claudia Díaz, Rachel Greenstadt, Damon McCoy |
Proc. Priv. Enhancing Technol. | 2 |
| 2017 | Editors' Introduction
Claudia Díaz, Rachel Greenstadt, Damon McCoy |
Proc. Priv. Enhancing Technol. | 2 |
| 2017 | Editors' Introduction
Claudia Díaz, Rachel Greenstadt, Damon McCoy |
Proc. Priv. Enhancing Technol. | 2 |
| 2016 | Blogs, Twitter Feeds, and Reddit Comments: Cross-domain Authorship AttributionabstractAbstract Stylometry is a form of authorship attribution that relies on the linguistic information to attribute documents of unknown authorship based on the writing styles of a suspect set of authors. This paper focuses on the cross-domain subproblem where the known and suspect documents differ in the setting in which they were created. Three distinct domains, Twitter feeds, blog entries, and Reddit comments, are explored in this work. We determine that state-of-the-art methods in stylometry do not perform as well in cross-domain situations (34.3% accuracy) as they do in in-domain situations (83.5% accuracy) and propose methods that improve performance in the cross-domain setting with both feature and classification level techniques which can increase accuracy to up to 70%. In addition to testing these approaches on a large real world dataset, we also examine real world adversarial cases where an author is actively attempting to hide their identity. Being able to identify authors across domains facilitates linking identities across the Internet making this a key security and privacy concern; users can take other measures to ensure their anonymity, but due to their unique writing style, they may not be as anonymous as they believe. Rebekah Overdorf, Rachel Greenstadt |
Proc. Priv. Enhancing Technol. | 2 |
| 2015 | De-anonymizing Programmers via Code Stylometry
Aylin Caliskan, Richard E. Harang, Arvind Narayanan, Clare R. Voss, Fabian Yamaguchi, Rachel Greenstadt |
USENIX Security Symposium | 7 |
| 2014 | A Critical Evaluation of Website Fingerprinting AttacksabstractRecent studies on Website Fingerprinting (WF) claim to have found highly effective attacks on Tor. However, these studies make assumptions about user settings, adversary capabilities, and the nature of the Web that do not necessarily hold in practical scenarios. The following study critically evaluates these assumptions by conducting the attack where the assumptions do not hold. We show that certain variables, for example, user's browsing habits, differences in location and version of Tor Browser Bundle, that are usually omitted from the current WF model have a significant impact on the efficacy of the attack. We also empirically show how prior work succumbs to the base rate fallacy in the open-world scenario. We address this problem by augmenting our classification method with a verification step. We conclude that even though this approach reduces the number of false positives over 63\%, it does not completely solve the problem, which remains an open issue for WF attacks. Marc Juarez, Sadia Afroz 0001, Gunes Acar, Claudia Díaz, Rachel Greenstadt |
CCS | 5 |
| 2014 | Active Linguistic Authentication Using Real-Time Stylometric Evaluation for Multi-Modal Decision Fusion
Ariel Stolerman, Lex Fridman 0001, Rachel Greenstadt, Patrick Brennan, Patrick Juola |
IFIP Int. Conf. Digital Forensics | 3 |
| 2014 | Breaking the Closed-World Assumption in Stylometric Authorship Attribution
Ariel Stolerman, Rebekah Overdorf, Sadia Afroz 0001, Rachel Greenstadt |
IFIP Int. Conf. Digital Forensics | 4 |
| 2014 | Doppelgänger Finder: Taking Stylometry to the UndergroundabstractStylometry is a method for identifying anonymous authors of anonymous texts by analyzing their writing style. While stylometric methods have produced impressive results in previous experiments, we wanted to explore their performance on a challenging dataset of particular interest to the security research community. Analysis of underground forums can provide key information about who controls a given bot network or sells a service, and the size and scope of the cybercrime underworld. Previous analyses have been accomplished primarily through analysis of limited structured metadata and painstaking manual analysis. However, the key challenge is to automate this process, since this labor intensive manual approach clearly does not scale. We consider two scenarios. The first involves text written by an unknown cybercriminal and a set of potential suspects. This is standard, supervised stylometry problem made more difficult by multilingual forums that mix l33t-speak conversations with data dumps. In the second scenario, you want to feed a forum into an analysis engine and have it output possible doppelgangers, or users with multiple accounts. While other researchers have explored this problem, we propose a method that produces good results on actual separate accounts, as opposed to data sets created by artificially splitting authors into multiple identities. For scenario 1, we achieve 77% to 84% accuracy on private messages. For scenario 2, we achieve 94% recall with 90% precision on blogs and 85.18% precision with 82.14% recall for underground forum users. We demonstrate the utility of our approach with a case study that includes applying our technique to the Carders forum and manual analysis to validate the results, enabling the discovery of previously undetected doppelganger accounts. Sadia Afroz 0001, Aylin Caliskan, Ariel Stolerman, Rachel Greenstadt, Damon McCoy |
IEEE Symposium on Security and Privacy | 4 |
| 2013 | Towards Active Linguistic Authentication
Patrick Juola, John Noecker Jr., Ariel Stolerman, Michael Ryan, Patrick Brennan, Rachel Greenstadt |
IFIP Int. Conf. Digital Forensics | 6 |
| 2013 | From Language to Family and Back: Native Language and Language Family Identification from English Text
Ariel Stolerman, Aylin Caliskan, Rachel Greenstadt |
HLT-NAACL | 3 |
| 2013 | GMM Based Semi-Supervised Learning for Channel-Based Authentication SchemeabstractAuthentication schemes based on wireless physical layer channel information have gained significant attention in recent years. It has been shown in recent studies, that the channel based authentication can either cooperate with existing higher layer security protocols or provide some degree of security to networks without central authority such as sensor networks. We propose a Gaussian Mixture Model based semi-supervised learning technique to identify intruders in the network by building a probabilistic model of the wireless channel of the network users. We show that even without having a complete apriori knowledge of the statistics of intruders and users in the network, our technique can learn and update the model in an online fashion while maintaining high detection rate. We experimentally demonstrate our proposed technique leveraging pattern diversity and show using measured channels that miss detection rates as low as 0.1% for false alarm rate of 0.3% can be achieved. Nikhil Gulati, Rachel Greenstadt, Kapil R. Dandekar, John MacLaren Walsh |
VTC Fall | 2 |
| 2013 | The illiterate editor: metadata-driven revert detection in WikipediaabstractAs the community depends more heavily on Wikipedia as a source of reliable information, the ability to quickly detect and remove detrimental information becomes increasingly important. The longer incorrect or malicious information lingers in a source perceived as reputable, the more likely that information will be accepted as correct and the greater the loss to source reputation. We present The Illiterate Editor (IllEdit), a content-agnostic, metadata-driven classification approach to Wikipedia revert detection. Our primary contribution is in building a metadata-based feature set for detecting edit quality, which is then fed into a Support Vector Machine for edit classification. By analyzing edit histories, the IllEdit system builds a profile of user behavior, estimates expertise and spheres of knowledge, and determines whether or not a given edit is likely to be eventually reverted. The success of the system in revert detection (0.844 F-measure) as well as its disjoint feature set as compared to existing, content-analyzing vandalism detection systems, shows promise in the synergistic usage of IllEdit for increasing the reliability of community information. Jeffrey Segall, Rachel Greenstadt |
OpenSym | 2 |
| 2012 | Use Fewer Instances of the Letter "i": Toward Writing Style Anonymization
Andrew W. E. McDonald, Sadia Afroz 0001, Aylin Caliskan, Ariel Stolerman, Rachel Greenstadt |
Privacy Enhancing Technologies | 5 |
| 2012 | Detecting Hoaxes, Frauds, and Deception in Writing Style OnlineabstractIn digital forensics, questions often arise about the authors of documents: their identity, demographic background, and whether they can be linked to other documents. The field of stylometry uses linguistic features and machine learning techniques to answer these questions. While stylometry techniques can identify authors with high accuracy in non-adversarial scenarios, their accuracy is reduced to random guessing when faced with authors who intentionally obfuscate their writing style or attempt to imitate that of another author. While these results are good for privacy, they raise concerns about fraud. We argue that some linguistic features change when people hide their writing style and by identifying those features, stylistic deception can be recognized. The major contribution of this work is a method for detecting stylistic deception in written documents. We show that using a large feature set, it is possible to distinguish regular documents from deceptive documents with 96.6% accuracy (F-measure). We also present an analysis of linguistic features that can be modified to hide writing style. Sadia Afroz 0001, Michael Brennan, Rachel Greenstadt |
IEEE Symposium on Security and Privacy | 3 |
| 2012 | Adversarial stylometry: Circumventing authorship recognition to preserve privacy and anonymityabstractThe use of stylometry, authorship recognition through purely linguistic means, has contributed to literary, historical, and criminal investigation breakthroughs. Existing stylometry research assumes that authors have not attempted to disguise their linguistic writing style. We challenge this basic assumption of existing stylometry methodologies and present a new area of research: adversarial stylometry. Adversaries have a devastating effect on the robustness of existing classification methods. Our work presents a framework for creating adversarial passages including obfuscation , where a subject attempts to hide her identity, and imitation , where a subject attempts to frame another subject by imitating his writing style, and translation where original passages are obfuscated with machine translation services. This research demonstrates that manual circumvention methods work very well while automated translation methods are not effective. The obfuscation method reduces the techniques' effectiveness to the level of random guessing and the imitation attempts succeed up to 67% of the time depending on the stylometry technique used. These results are more significant given the fact that experimental subjects were unfamiliar with stylometry, were not professional writers, and spent little time on the attacks. This article also contributes to the field by using human subjects to empirically validate the claim of high accuracy for four current techniques (without adversaries). We have also compiled and released two corpora of adversarial stylometry texts to promote research in this field with a total of 57 unique authors. We argue that this field is important to a multidisciplinary approach to privacy, security, and anonymity. Michael Brennan, Sadia Afroz 0001, Rachel Greenstadt |
ACM Trans. Inf. Syst. Secur. | 3 |
| 2009 | Practical Attacks Against Authorship Recognition Techniques
Michael Brennan, Rachel Greenstadt |
IAAI | 2 |
| 2006 | Privatizing Constraint Optimization
Rachel Greenstadt |
AAAI | 1 |
| 2006 | Analysis of Privacy Loss in Distributed Constraint Optimization
Rachel Greenstadt, Jonathan P. Pearce, Milind Tambe |
AAAI | 1 |
| 2003 | Why we can't be bothered to read privacy policies models of privacy economics as a lemons marketabstractConsumers want to interact with web sites, but they also want to keep control of their private information. Asymmetric information about whether web sites will sell private information or not leads to a lemons market for privacy. We discuss privacy policies as signals in a lemons market and ways in which current realizations of privacy policies may fail to be effective signals. As a result of these shortcomings, we consider a "lemons market with testing," where consumers have a cost of determining whether a site meets their privacy requirement. Our model explains empirical data concerning privacy policies and privacy seals. We end by discussing cyclic instability in the number of web sites that sell consumer information. Tony Vila, Rachel Greenstadt, David Molnar |
ICEC | 2 |