Sarah Michele Rajtmajer

dblp:42/8717 · also Sarah Rajtmajer · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-1464-0848ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 A Tale of Two Identities: An Ethical Audit of AI-Crafted Synthetic Personas
abstract
As LLMs (large language models) are increasingly used to generate synthetic personas, particularly in data-limited domains such as health, privacy, and HCI, it becomes necessary to understand how these narratives represent identity, especially that of minority communities. In this paper, we audit synthetic personas generated by 3 LLMs (GPT4o, Gemini 1.5 Pro, Deepseek v2.5) through the lens of representational harm, focusing specifically on racial identity. Using a mixed-methods approach combining close reading, lexical analysis, and a parameterized creativity framework, we compare 1,512 LLM-generated persona to human-authored responses. Our findings reveal that LLMs disproportionately foreground racial markers, overproduce culturally coded language, and construct personas that are syntactically elaborate yet narratively reductive. These patterns result in a range of sociotechnical harms, including stereotyping, exoticism, erasure, and benevolent bias, that are often obfuscated by superficially positive narrations. We formalize this phenomenon as algorithmic othering, where minoritized identities are rendered hypervisible but less authentic.
Pranav Venkit, Yingfan Zhou, Sarah Michele Rajtmajer, Shomir Wilson
AAAI4
2025 Can Third Parties Read Our Emotions?
abstract
Natural Language Processing tasks that aim to infer an author’s private states, e.g., emotions and opinions, from their written text, typically rely on datasets annotated by third-party annotators. However, the assumption that third-party annotators can accurately capture authors’ private states remains largely unexamined. In this study, we present human subjects experiments on emotion recognition tasks that directly compare third-party annotations with first-party (author-provided) emotion labels. Our findings reveal significant limitations in third-party annotations—whether provided by human annotators or large language models (LLMs)—in faithfully representing authors’ private states. However, LLMs outperform human annotators nearly across the board. We further explore methods to improve third-party annotation quality. We find that demographic similarity between first-party authors and third-party human annotators enhances annotation performance. While incorporating first-party demographic information into prompts leads to a marginal but statistically significant improvement in LLMs’ performance. We introduce a framework for evaluating the limitations of third-party annotations and call for refined annotation practices to accurately represent and model authors’ private states.
Yingfan Zhou, Pranav Venkit, Halima Binte Islam, Sneha Arya, Shomir Wilson, Sarah Michele Rajtmajer
ACL (1)7
2025 Have LLMs Reopened the Pandora's Box of AI-Generated Fake News?
abstract
Xinyu Wang, Wenbo Zhang, Sai Koneru, Hangzhi Guo, Bonam Mingole, S. Shyam Sundar, Sarah Rajtmajer, Amulya Yadav. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xinyu Wang 0029, Wenbo Zhang 0008, Sai Dileep Koneru, Hangzhi Guo, Bonam Mingole, S. Shyam Sundar, Sarah Michele Rajtmajer, Amulya Yadav
NAACL (Long Papers)7
2025 Effect of AI Performance, Risk Perception, and Trust on Human Dependence in Deepfake Detection AI System
abstract
Synthetic images, audio, and video can now be generated and edited by Artificial Intelligence (AI). In particular, the malicious use of synthetic data has raised concerns about potential harms to cybersecurity, personal privacy, and public trust. Although AI-based detection tools exist to help identify synthetic content, their limitations often lead to user mistrust and confusion between real and fake content. This study examines the role of AI performance in influencing human trust and decision making in synthetic data identification. Through an online human subject experiment involving 400 participants, we examined how varying AI performance impacts human trust and dependence on AI in deepfake detection. Our findings indicate how participants calibrate their dependence on AI based on their perceived risk and the prediction results provided by AI. These insights contribute to the development of transparent and explainable AI systems that better support everyday users in mitigating the harms of synthetic media.
Yingfan Zhou, Ester Chen, Manasa Pisipati, Aiping Xiong, Sarah Michele Rajtmajer
Proc. ACM Hum. Comput. Interact.5
2024 Integrating measures of replicability into scholarly search: Challenges and opportunities
abstract
Challenges to reproducibility and replicability have gained widespread attention, driven by large replication projects with lukewarm success rates. A nascent work has emerged developing algorithms to estimate the replicability of published findings. The current study explores ways in which AI-enabled signals of confidence in research might be integrated into the literature search. We interview 17 PhD researchers about their current processes for literature search and ask them to provide feedback on a replicability estimation tool. Our findings suggest that participants tend to confuse replicability with generalizability and related concepts. Information about replicability can support researchers throughout the research design processes. However, the use of AI estimation is debatable due to the lack of explainability and transparency. The ethical implications of AI-enabled confidence assessment must be further studied before such tools could be widely accepted. We discuss implications for the design of technological tools to support scholarly activities and advance replicability.
Chuhao Wu, Tatiana Chakravorti, John M. Carroll 0001, Sarah Michele Rajtmajer
CHI4
2024 Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
abstract
Hypothesis formulation and testing are central to empirical research. A strong hypothesis is a best guess based on existing evidence and informed by a comprehensive view of relevant literature. However, with exponential increase in the number of scientific articles published annually, manual aggregation and synthesis of evidence related to a given hypothesis is a challenge. Our work explores the ability of current large language models (LLMs) to discern evidence in support or refute of specific hypotheses based on the text of scientific abstracts. We share a novel dataset for the task of scientific hypothesis evidencing using community-driven annotations of studies in the social sciences. We compare the performance of LLMs to several state of the art methods and highlight opportunities for future research in this area. Our dataset is shared with the research community: https://github.com/Sai90000/ScientificHypothesisEvidencing.git
Sai Dileep Koneru, Jian Wu 0006, Sarah Michele Rajtmajer
LREC/COLING3
2024 An Audit on the Perspectives and Challenges of Hallucinations in NLP
abstract
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, Shomir Wilson. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Pranav Venkit, Tatiana Chakravorti, Heidi R. Biggs, Mukund Srinath, Koustava Goswami, Sarah Michele Rajtmajer, Shomir Wilson
EMNLP7
2023 A Study on Reproducibility and Replicability of Table Structure Recognition Methods
Kehinde Ajayi, Muntabir Hasan Choudhury, Sarah Michele Rajtmajer, Jian Wu 0006
ICDAR (2)3
2023 Information Operations in Turkey: Manufacturing Resilience with Free Twitter Accounts
abstract
Following the 2016 US elections Twitter launched their Information Operations (IO) hub where they archive account activity connected to state linked information operations. In June 2020, Twitter took down and released a set of accounts linked to Turkey's ruling political party (AKP). We investigate these accounts in the aftermath of the takedown to explore whether AKP-linked operations are ongoing and to understand the strategies they use to remain resilient to disruption. We collect live accounts that appear to be part of the same network, ~30% of which have been suspended by Twitter since our collection. We create a BERT-based classifier that shows similarity between these two networks, develop a taxonomy to categorize these accounts, find direct sequel accounts between the Turkish takedown and the live accounts, and find evidence that Turkish IO actors deliberately construct their network to withstand large-scale shutdown by utilizing explicit and implicit signals of coordination. We compare our findings from the Turkish operation to Russian and Chinese IO on Twitter and find that Turkey's IO utilizes a unique group structure to remain resilient. Our work highlights the fundamental imbalance between IO actors quickly and easily creating free accounts and the social media platforms spending significant resources on detection and removal, and contributes novel findings about Turkish IO on Twitter.
Maya Merhi, Sarah Michele Rajtmajer, Dongwon Lee 0001
ICWSM2
2023 Differential Privacy enabled Dementia Classification: An Exploration of the Privacy-Accuracy Trade-off in Speech Signal Data
Suhas BN, Sarah Michele Rajtmajer, Saeed Abdullah
INTERSPEECH2
2023 Content Sharing Design for Social Welfare in Networked Disclosure Game
abstract
This work models the costs and benefits of personal information sharing, or self-disclosure, in online social networks as a networked disclosure game. In a networked population where edges represent visibility amongst users, we assume a leader can influence network structure through content promotion, and we seek to optimize social welfare through network design. Our approach considers user interaction non-homogeneously, where pairwise engagement amongst users can involve or not involve sharing personal information. We prove that this problem is NP-hard. As a solution, we develop a Mixed-integer Linear Programming algorithm, which can achieve an exact solution, and also develop a time-efficient heuristic algorithm that can be used at scale. We conduct numerical experiments to demonstrate the properties of the algorithms and map theoretical results to a dataset of posts and comments in 2020 and 2021 in a COVID-related Subreddit community where privacy risks and sharing tradeoffs were particularly pronounced.
Feiran Jia, Chenxi Qiu, Sarah Michele Rajtmajer, Anna Cinzia Squicciarini
UAI3
2022 A Synthetic Prediction Market for Estimating Confidence in Published Work
abstract
Explainably estimating confidence in published scholarly work offers opportunity for faster and more robust scientific progress. We develop a synthetic prediction market to assess the credibility of published claims in the social and behavioral sciences literature. We demonstrate our system and detail our findings using a collection of known replication projects. We suggest that this work lays the foundation for a research agenda that creatively uses AI for peer review.
Sarah Michele Rajtmajer, Christopher Griffin 0001, Jian Wu 0006, Robert Fraleigh, Laxmaan Balaji, Anna Cinzia Squicciarini, Anthony Kwasnica, David M. Pennock, Michael McLaughlin, Timothy Fritton, Nishanth Nakshatri, Arjun Manoj Menon, Sai Ajay Modukuri, Rajal Nivargi, C. Lee Giles
AAAI1
2022 The Contribution of Verified Accounts to Self-Disclosure in COVID-Related Twitter Conversations
Tingting Du, Prasanna Umar, Sarah Michele Rajtmajer, Anna Cinzia Squicciarini
ICWSM3
2022 An Extended Ultimatum Game for Multi-Party Access Control in Social Networks
abstract
In this article, we aim to answer an important set of questions about the potential longitudinal effects of repeated sharing and privacy settings decisions over jointly managed content among users in a social network. We model user interactions through a repeated game in a network graph. We present a variation of the one-shot Ultimatum Game, wherein individuals interact with peers to make a decision on a piece of shared content. The outcome of this game is either success or failure, wherein success implies that a satisfactory decision for all parties is made and failure instead implies that the parties could not reach an agreement. Our proposed game is grounded in empirical data about individual decisions in repeated pairwise negotiations about jointly managed content in a social network. We consider both a “continuous” privacy model as well the “discrete” case of a model wherein privacy values are to be chosen among a fixed set of options. We formally demonstrate that over time, the system converges toward a “fair” state, wherein each individual’s preferences are accounted for. Our discrete model is validated by way of a user study, where participants are asked to propose privacy settings for own shared content from a small, discrete set of options.
Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, Justin Semonsen, Andrew Belmonte, Pratik Agarwal
ACM Trans. Web2
2021 Comparative assessment of cyber-physical threats to megacities
abstract
By 2030, forecasts suggest that urban areas will house 60 percent of the world’s population and one in every three people will live in cities with at least half a million inhabitants. Within the same time frame, the number of global megacities is expected to jump from 33 today to 43 in 2030 [1]. Underpinning these large urban areas will be an interconnected network of critical physical infrastructures reliant on Internet-connected Industrial Control Systems and susceptible to increasingly sophisticated, e.g., AI-enabled, cyber threats. In hand, the cyber threat landscape is shifting rapidly. We are seeing a sharp rise in the number of cyberattacks on critical infrastructure [2] with significant impacts cascading across multiple sectors and causing disruption to the provisioning of essential goods and services. Security scholars suggest that these impacts are not always equitable and that disruption to critical infrastructure can affect vulnerable groups differently [3], which further emphasizes the need to improve cybersecurity between critical infrastructure sectors [4]. Through structured analysis of city statistics, demographic information, cyber incidents, and current cyber policy, our presentation will articulate potential social implications of megacity growth through the lens of cyber-physical infrastructure disruption. We investigate the largest 15 megacities in the world and find that megacities continue to grow in population but not in cyber policy. We highlight recent examples of cyber-physical disruption in Mumbai and New York City with focus on implications for vulnerable populations. Our work suggests the need for future research on social responsibility regarding security of these critical infrastructure sectors and on the need for technology-focused law, policy, and regulation guidelines.
Jordyn Dennis, Caitlin A. Grady, Sarah Michele Rajtmajer
ISTAS3
2021 Self-disclosure on Twitter During the COVID-19 Pandemic: A Network Perspective
Prasanna Umar, Chandan Akiti, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer
ECML/PKDD (4)4
2021 Digital Inequality Through the Lens of Self-Disclosure
abstract
Abstract Recent work has brought to light disparities in privacy-related concerns based on socioeconomic status, race and ethnicity. This paper examines relationships between U.S. based Twitter users’ socio-demographic characteristics and their privacy behaviors. Income, gender, age, race/ethnicity, education level and occupation are correlated with stated and observed privacy preferences of 110 active Twitter users. Contrary to our expectations, analyses suggest that neither socioeconomic status ( SES ) nor demographics is a significant predictor of the use of account security features. We do find that gender and education predict rate of self-disclosure, or voluntary sharing of personal information. We explore variability in the types of information disclosed amongst socio-demographic groups. Exploratory findings indicate that: 1) participants shared less personal information than they recall having shared in exit surveys; 2) there is no strong correlation between people’s stated attitudes and their observed behaviors.
Sarah Michele Rajtmajer, Eesha Srivatsavaya, Shomir Wilson
Proc. Priv. Enhancing Technol.2
2020 A Study of Self-Privacy Violations in Online Public Discourse
abstract
User engagement in online public discourse often includes self-disclosure - the revelation of personal information. Such disclosures on online public platforms (e.g., news forums) become a shared history, vulnerable to detrimental use by advertisers and malicious parties. Yet, users indulge in self-disclosing behavior to attain strategic goals like relational development, social connectedness, identity clarification, and social control. In this work, we develop supervised models to detect instances of self-disclosure in users' comments in the context of public discourse. Using three different datasets, we validate the performance of our models. Our detection models achieve an accuracy of 75.8 percent in a news discourse dataset. The performances on evaluation against existing methods on two secondary datasets are on par if not better. We examine the rate at which users self-disclose to understand when and to what extent users abide by group norms of such behavior. Our results show that self-disclosing users are often similar in their alignment or divergence with the group norm. As such, these similarly divergent users in a conversation use similar language in their disclosures. Finally, we reflect o n the implications o f alignment with or divergence from group norms in light of online privacy.
Prasanna Umar, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer
IEEE BigData3
2019 Toward Image Privacy Classification and Spatial Attribution of Private Content
abstract
Machine labeling of image content as private or public is a notoriously difficult problem, with the usual image processing challenges compounded by the highly personal, subjective, and contextual nature of access control decision making. In general, a user's privacy expectation for a given image is consequential to specific contents therein and the presence of sensitive content somewhere in the image is sufficient to warrant a private label. In this work, we extend the problem of determining a single privacy label for a given image to jointly inferring a privacy label and detecting the specific areas of sensitive content within a privately labeled image. We propose a stochastic spatial attribution model which exploits sophisticated (deep neural net derived) image features over randomly selected image patches, as well as image saliency quantification. We validate our detected private regions through extensive user study experiments. This effort to achieve spatial attribution of private image content helps to lay a foundation for warning mechanisms which may serve to aid both social media sites and their users.
Haoti Zhong, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, David J. Miller 0001
IEEE BigData4
2019 Rating Mechanisms for Sustainability of Crowdsourcing Platforms
abstract
Crowdsourcing leverages the diverse skill sets of large collections of individual contributors to solve problems and execute projects, where contributors may vary significantly in experience, expertise, and interest in completing tasks. Hence, to ensure the satisfaction of its task requesters, most existing crowdsourcing platforms focus primarily on supervising contributors' behavior. This lopsided approach to supervision negatively impacts contributor engagement and platform sustainability.
Chenxi Qiu, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer
CIKM3
2019 Detection and Analysis of Self-Disclosure in Online News Commentaries
abstract
Online users engage in self-disclosure - revealing personal information to others - in pursuit of social rewards. However, there are associated costs of disclosure to users' privacy. User profiling techniques support the use of contributed content for a number of purposes, e.g., micro-targeting advertisements. In this paper, we study self-disclosure as it occurs in newspaper comment forums. We explore a longitudinal dataset of about 60,000 comments on 2202 news articles from four major English news websites. We start with detection of language indicative of various types of self-disclosure, leveraging both syntactic and semantic information present in texts. Specifically, we use dependency parsing for subject, verb, and object extraction from sentences, in conjunction with named entity recognition to extract linguistic indicators of self-disclosure. We then use these indicators to examine the effects of anonymity and topic of discussion on self-disclosure. We find that anonymous users are more likely to self-disclose than identifiable users, and that self-disclosure varies across topics of discussion. Finally, we discuss the implications of our findings for user privacy.
Prasanna Umar, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer
WWW3
2019 Opinion Dynamics in the Presence of Increasing Agreement Pressure
abstract
In this paper, we study a model of agent consensus in a social network in the presence increasing interagent influence, i.e., increasing peer pressure. Each agent in the social network has a distinct social stress function given by a weighted sum of internal and external behavioral pressures. We assume a weighted average update rule consistent with the classic DeGroot model and prove conditions under which a connected group of agents converge to a fixed opinion distribution, and under which conditions the group reaches consensus. We show that the update rule converges to gradient descent and explain its transient and asymptotic convergence properties. Through simulation, we study the rate of convergence on a scale-free network.
Justin Semonsen, Christopher Griffin 0001, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer
IEEE Trans. Cybern.4
2018 Multi-Party Access Control: Requirements, State of the Art and Open Challenges
abstract
Multi-party access control is gaining attention and prominence within the community, as access control models and systems are faced with complex, jointly-owned and jointly-managed content. Traditional single-user approaches lack the richness and flexibility to accommodate these scenarios, resulting in undesired disclosure of sensitive data and resources. Moving forward fundamental work in this area is critical. In particular, as personal data amasses and algorithms for data mining improve, personally identifiable information is more readily inferred and the practical implications of privacy decisions are relatively opaque. This is true even at the individual level, but the parallel problem for jointly managed content involves the cross product of these complex outcomes. In this presentation, we discuss fundamental requirements of successful multi-party access control mechanisms and contextualize these concepts with respect to the state of the art. Based on this analysis, we identify open challenges and draw a roadmap for future work.
Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, Nicola Zannone
SACMAT2
2017 Dynamic Contract Design for Heterogenous Workers in Crowdsourcing for Quality Control
abstract
Crowdsourcing sites heavily rely on paid workers to ensure completion of tasks. Yet, designing a pricing strategies able to incentivize users' quality and retention is non trivial. Existing payment strategies either simply set a fixed payment per task without considering changes in workers' behaviors, or rule out poor quality responses and workers based on coarse criteria. Hence, task requesters may be investing significantly in work that is inaccurate or even misleading. In this paper, we design a dynamic contract to incentivize high-quality work. Our proposed approach offers a theoretically proven algorithm to calculate the contract for each worker in a cost-efficient manner. In contrast to existing work, our contract design is not only adaptive to changes in workers' behavior, but also adjusts pricing policy in the presence of malicious behavior. Both theoretical and experimental analysis over real Amazon review traces show that our contract design can achieve a near optimal solution. Furthermore, experimental results demonstrate that our contract design 1) can promote high-quality work and prevent malicious behavior, and 2) outperforms the intuitive strategy of excluding all malicious workers in terms of the requester's utility.
Chenxi Qiu, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, James Caverlee
ICDCS3
2016 Content-Driven Detection of Cyberbullying on the Instagram Social Network
Haoti Zhong, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, Christopher Griffin 0001, David J. Miller 0001, Cornelia Caragea
IJCAI4
2015 A Hybrid Epidemic Model for Antinormative Behavior in Online Social Networks
abstract
In this paper, we describe a novel approach to investigate negative behavior dynamics in online social networks as epidemic phenomena. We present a finite-state machine model for time-varying epidemic dynamics, and validate this model with experiments over a large dataset of Youtube commentaries, indicating how different epidemic patterns of behavior can be tied to specific interaction patterns among users. A full version of this paper is available on arXiv.org.
Cong Liao, Anna Cinzia Squicciarini, Christopher Griffin 0001, Sarah Michele Rajtmajer
ASONAM4
2015 Identification and characterization of cyberbullying dynamics in an online social network
abstract
Cyberbullying is an increasingly prevalent phenomenon impacting young adults. In this paper, we present a study on both detecting cyberbullies in online social networks and identifying the pairwise interactions between users through which the influence of bullies seems to spread. In particular, we investigate the role of user demographics and social network features in predicting how users will respond to a cyberbullying comment. We characterize the influencer/influenced relationship by which a user who has no history of abuse observes a peer engaging in bullying and follows suit. To our knowledge, this is the first effort modeling peer pressure and social dynamics with analytical models. We validate our models on two distinct social network datasets, totalling over 16,000 posts. Our results offer insight into the dynamics of bullying and confirm social theories on the power of peer groups in the cyberworld. A full version of this paper is available on arXiv.org.
Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, Christopher Griffin 0001
ASONAM2