Sai Teja Peddinti

dblp:91/8283 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
12since 2021 · last 2026
0009-0007-3242-9353ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 14 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
abstract
Large Language Models (LLMs) such as ChatGPT can infer personal attributes from seemingly innocuous text, raising privacy risks beyond memorized data leakage. While prior work has demonstrated these risks, little is known about how users estimate and respond. We conducted a survey with 240 U.S. participants who judged text snippets for inference risks, reported concern levels, and attempted rewrites to block inference. We compared their rewrites with those generated by ChatGPT and Rescriber, a state-of-the-art sanitization tool. Results show that participants struggled to anticipate inference, performing a little better than chance. User rewrites were effective in just 28% of cases - better than Rescriber but worse than ChatGPT. We examined our participants’ rewriting strategies, and observed that while paraphrasing was the most common strategy it is also the least effective; instead abstraction and adding ambiguity were more successful. Our work highlights the importance of inference-aware design in LLM interactions.
Qia Wang 0001, Sai Teja Peddinti, Nina Taft, Nick Feamster
CHI2
2026 Understanding U.S. Users' Security and Privacy Transparency Needs for Consumer-Facing Generative AI
Jiaxun Cao, Chunxi Zhan, Rithvik Neti, Sai Teja Peddinti, Pardis Emami Naeini
SOUPS5
2026 "You Have Been Selected as the Winner": Characterizing User-Reported Scams on TikTok
Smirity Kaushik, Kyle Beadle, Gauri Nayak, Madelyn Sanfilippo, Mainack Mondal, Yang Wang 0005, Sai Teja Peddinti, Yixin Zou
SOUPS7
2026 Nudging Developers Toward Privacy: Evaluating the Impact of Personalized App Review Reports
Sai Teja Peddinti, Omer Akgul, Michelle L. Mazurek, Nina Taft
SOUPS1
2025 "We are not Future-ready": Understanding AI Privacy Risks and Existing Mitigation Strategies from the Perspective of AI Developers in Europe
Alexandra Klymenko, Stephen Meisenbacher, Patrick Gage Kelley, Sai Teja Peddinti, Kurt Thomas, Florian Matthes
SOUPS4
2024 A Decade of Privacy-Relevant Android App Reviews: Large Scale Trends
Omer Akgul, Sai Teja Peddinti, Nina Taft, Michelle L. Mazurek, Hamza Harkous, Animesh Srivastava, Benoit Seguin
USENIX Security Symposium2
2023 Towards Fine-Grained Localization of Privacy Behaviors
abstract
Privacy labels help developers communicate their application’s privacy behaviors (i.e., how and why an application uses personal information) to users. But, studies show that developers face several challenges in creating them and the resultant labels are often inconsistent with their application’s privacy behaviors. In this paper, we create a novel methodology called fine-grained localization of privacy behaviors to locate individual statements in source code which encode privacy behaviors and predict their privacy labels. We design and develop an attention-based multi-head encoder model which creates individual representations of multiple methods and uses attention to identify relevant statements that implement privacy behaviors. These statements are then used to predict privacy labels for the application’s source code and can help developers write privacy statements that can be used as notices. Our quantitative analysis shows that our approach can achieve high accuracy in identifying privacy labels, with the lowest accuracy of 91.41% and the highest of 98.45%. We also evaluate the efficacy of our approach with six software professionals from our university. The results demonstrate that our approach reduces the time and mental effort required by developers to create high-quality privacy statements and can finely localize statements in methods that implement privacy behaviors.
Vijayanta Jain, Sepideh Ghanavati, Sai Teja Peddinti, Collin McMillan
EuroS&P3
2022 Analyzing User Perspectives on Mobile App Privacy at Scale
abstract
In this paper we present a methodology to analyze users' concerns and perspectives about privacy at scale. We leverage NLP techniques to process millions of mobile app reviews and extract privacy concerns. Our methodology is composed of a binary classifier that distinguishes between privacy and non-privacy related reviews. We use clustering to gather reviews that discuss similar privacy concerns, and employ summarization metrics to extract representative reviews to summarize each cluster. We apply our methods on 287M reviews for about 2M apps across the 29 categories in Google Play to identify top privacy pain points in mobile apps. We identified approximately 440K privacy related reviews. We find that privacy related reviews occur in all 29 categories, with some issues arising across numerous app categories and other issues only surfacing in a small set of app categories. We show empirical evidence that confirms dominant privacy themes - concerns about apps requesting unnecessary permissions, collection of personal information, frustration with privacy controls, tracking and the selling of personal data. As far as we know, this is the first large scale analysis to confirm these findings based on hundreds of thousands of user inputs. We also observe some unexpected findings such as users warning each other not to install an app due to privacy issues, users uninstalling apps due to privacy reasons, as well as positive reviews that reward developers for privacy friendly apps. Finally we discuss the implications of our method and findings for developers and app stores.
Preksha Nema, Pauline Anthonysamy, Nina Taft, Sai Teja Peddinti
ICSE4
2022 Hark: A Deep Learning System for Navigating Privacy Feedback at Scale
abstract
Integrating user feedback is one of the pillars for building successful products. However, this feedback is generally collected in an unstructured free-text form, which is challenging to understand at scale. This is particularly demanding in the privacy domain due to the nuances associated with the concept and the limited existing solutions. In this work, we present Hark1, a system for discovering and summarizing privacy-related feedback at scale. Hark automates the entire process of summarizing privacy feedback, starting from unstructured text and resulting in a hierarchy of high-level privacy themes and fine-grained issues within each theme, along with representative reviews for each issue. At the core of Hark is a set of new deep learning models trained on different tasks, such as privacy feedback classification, privacy issues generation, and high-level theme creation. We illustrate Hark’s efficacy on a corpus of 626 M Google Play reviews. Out of this corpus, our privacy feedback classifier extracts $6 M$ privacy-related reviews (with an AUC-ROC of 0.92). With three annotation studies, we show that Hark’s generated issues are of high accuracy and coverage and that the theme titles are of high quality. We illustrate Hark’s capabilities by presenting high-level insights from $1.3 M$ Android apps.1an English verb meaning to “pay close attention”
Hamza Harkous, Sai Teja Peddinti, Rishabh Khandelwal, Animesh Srivastava, Nina Taft
SP2
2022 PAcT: Detecting and Classifying Privacy Behavior of Android Applications
abstract
Interpreting and describing mobile applications' privacy behaviors to ensure creating consistent and accurate privacy notices is a challenging task for developers. Traditional approaches to creating privacy notices are based on predefined templates or questionnaires and do not rely on any traceable behaviors in code which may result in inconsistent and inaccurate notices. In this paper, we present an automated approach to detect privacy behaviors in code of Android applications. We develop Privacy Action Taxonomy (PAcT), which includes labels for Practice (i.e. how applications use personal information) and Purpose (i.e. why). We annotate ~5,200 code segments based on the labels and create a multi-label multi-class dataset with ~14,000 labels. We develop and train deep learning models to classify code segments. We achieve the highest F-1 scores across all label types of 79.62% and 79.02% for Practice and Purpose.
Vijayanta Jain, Sanonda Datta Gupta, Sepideh Ghanavati, Sai Teja Peddinti, Collin McMillan
WISEC4
2021 PriGen: Towards Automated Translation of Android Applications' Code to Privacy Captions
Vijayanta Jain, Sanonda Datta Gupta, Sepideh Ghanavati, Sai Teja Peddinti
RCIS4
2021 A Large Scale Study of User Behavior, Expectations and Engagement with Android Permissions
Weicheng Cao, Chunqiu Xia, Sai Teja Peddinti, David Lie, Nina Taft, Lisa M. Austin
USENIX Security Symposium3
2019 Reducing Permission Requests in Mobile Apps
abstract
Users of mobile apps sometimes express discomfort or concerns with what they see as unnecessary or intrusive permission requests by certain apps. However encouraging mobile app developers to request fewer permissions is challenging because there are many reasons why permissions are requested; furthermore, prior work [25] has shown it is hard to disambiguate the purpose of a particular permission with high certainty.
Sai Teja Peddinti, Igor Bilogrevic, Nina Taft, Martin Pelikan, Úlfar Erlingsson, Pauline Anthonysamy, Giles Hogben
Internet Measurement Conference1
2017 Exploring decision making with Android's runtime permission dialogs using in-context surveys
Bram Bonné, Sai Teja Peddinti, Igor Bilogrevic, Nina Taft
SOUPS2
2016 Finding Sensitive Accounts on Twitter: An Automated Approach Based on Follower Anonymity
Sai Teja Peddinti, Keith W. Ross, Justin Cappos
ICWSM1
2014 Cloak and Swagger: Understanding Data Sensitivity through the Lens of User Anonymity
abstract
Most of what we understand about data sensitivity is through user self-report (e.g., surveys), this paper is the first to use behavioral data to determine content sensitivity, via the clues that users give as to what information they consider private or sensitive through their use of privacy enhancing product features. We perform a large-scale analysis of user anonymity choices during their activity on Quora, a popular question-and-answer site. We identify categories of questions for which users are more likely to exercise anonymity and explore several machine learning approaches towards predicting whether a particular answer will be written anonymously. Our findings validate the viability of the proposed approach towards an automatic assessment of data sensitivity, show that data sensitivity is a nuanced measure that should be viewed on a continuum rather than as a binary concept, and advance the idea that machine learning over behavioral data can be effectively used in order to develop product features that can help keep users safe.
Sai Teja Peddinti, Aleksandra Korolova, Elie Bursztein, Geetanjali Sampemane
IEEE Symposium on Security and Privacy1
2014 Web search query privacy: Evaluating query obfuscation and anonymizing networks
abstract
Web Search is one of the most rapidly growing applications on the internet today. However, the current practice followed by most search engines – of logging and analyzing users' queries – raises serious privacy concerns. In this paper, we concentrate on two existing solutions which are relatively easy to deploy – namely Query Obfuscation and Anonymizing Networks. In query obfuscation, a client-side software attempts to mask real user queries via injection of certain noisy queries. Anonymizing networks route the user queries through a series of relay servers, hiding the actual query source from the search engine. A fundamental problem with these solutions, however, is that user queries are still obviously revealed to the search engine, although they are “mixed” among queries generated either by a machine or by other users. We focus on TrackMeNot (TMN), a popular query obfuscation tool, and the Tor anonymizing network, and try to analyse whether these solutions can actually preserve users' privacy in practice against an adversarial search engine. We demonstrate that a search engine, equipped with only a short-term history of a user's search queries, can break the privacy guarantees of TMN and Tor by only utilizing off-the-shelf machine learning techniques.
Sai Teja Peddinti, Nitesh Saxena
J. Comput. Secur.1
2011 On the effectiveness of anonymizing networks for web search privacy
abstract
Web search has emerged as one of the most important applications on the internet, with several search engines available to the users. There is a common practice among these search engines to log and analyse the user queries, which leads to serious privacy implications. One well known solution to search privacy involves issuing the queries via an anonymizing network, such as Tor, thereby hiding one's identity from the search engine. A fundamental problem with this solution, however, is that user queries are still obviously revealed to the search engine, although they are "mixed" among the queries issued by other users of the same anonymization service.
Sai Teja Peddinti, Nitesh Saxena
AsiaCCS1
2011 On the limitations of query obfuscation techniques for location privacy
abstract
A promising approach to location privacy is query obfuscation, which involves reporting k -- 1 false locations along with the real location. In this paper, we examine the level of privacy protection provided by the current query obfuscation techniques against adversarial location service providers. As a representative and realistic implementation of query obfuscation, we focus on SybilQuery. We present two types of attacks depending upon whether or not a short-term query history is available. When history is available, using machine learning, we were able to identify 93.67% of user trips, with only 2.02% of fake trips misclassified, for the security parameter k = 5. In the absence of history, we used trip correlations to form a smaller set of trips effectively increasing the user query identification probability from 20% to about 40%. Our work demonstrates that the use of aggregate statistical information alone is not sufficient to generate simulated trips. We identify areas for improvement in the existing query obfuscation techniques.
Sai Teja Peddinti, Nitesh Saxena
UbiComp1
2010 On the Privacy of Web Search Based on Query Obfuscation: A Case Study of TrackMeNot
Sai Teja Peddinti, Nitesh Saxena
Privacy Enhancing Technologies1