EDBT 2026 Demo / reviewers in the wild / expert
Mahmood Sharif
dblp:136/8393
· DBLP profile ↗
24ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-7661-2220ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 18 · 4 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bot Among Us: Exploring User Awareness and Privacy Concerns About Chatbots in Group ChatsabstractAs chatbots become increasingly integrated into group conversations on instant messaging platforms, concerns arise about their impact on user privacy. While prior research has examined chatbot risks in one-on-one interactions, little is known about how users perceive and respond to privacy threats in group settings, where chatbots may silently access messages and metadata. To address this gap, we conducted an online survey (N=374) across five popular messaging platforms—WhatsApp, Discord, Telegram, Viber, and LINE—to evaluate user awareness, understanding of chatbot access, privacy concerns, and behavioral responses. We found that many users were unaware of bots in their group chats and significantly underestimated their data access: only 41.7% correctly identified what messages chatbots could access. Privacy concerns also rose sharply after users learned about actual bot permissions. Based on our findings, we propose a five-stage model that captures how users detect, interpret, and respond to chatbot-related privacy risks. We further analyzed the designs of platforms with official chatbot support through this model and found mismatches between design choices and user expectations. Finally, we offer design recommendations to improve transparency and user control in group chatbot-interactions. Kai-Hsiang Chou, Yi-An Wang, Chong Kai Lau, Mahmood Sharif, Hsu-Chun Hsiao |
Proc. Priv. Enhancing Technol. | 4 |
| 2026 | Universal Jailbreak Suffixes Are Strong Attention HijackersabstractAbstract We study suffix-based jailbreaks—a powerful family of attacks against large language models (LLMs) that optimize adversarial suffixes to circumvent safety alignment. Focusing on the widely used foun-dational GCG attack (Zou et al., 2023b), we observe that suffixes vary in efficacy: some are markedly more universal—generalizing to many unseen harmful instructions—than others. We first show that a shallow, critical mechanism drives GCG’s effectiveness. This mechanism builds on the information flow from the adversarial suffix to the final chat template tokens before generation. Quantifying the dominance of this mechanism during generation, we find GCG irregularly and aggressively hijacks the contex-tualization process. Crucially, we tie hijacking to the universality phenomenon, with more universal suffixes being stronger hijackers. Subsequently, we show that these insights have practical implications: GCG’s universality can be efficiently enhanced (up to ×5 in some cases) at no additional computational cost, and can also be surgically mitigated, at least halving the attack’s success with minimal utility loss.1 Matan Ben-Tov, Mor Geva, Mahmood Sharif |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | GASLITEing the Retrieval: Exploring Vulnerabilities in Dense Embedding-based SearchabstractDense embedding-based text retrieval—retrieval of relevant passages from corpora via deep learning encodings—has emerged as a powerful method attaining state-of-the-art search results and popularizing Retrieval Augmented Generation (RAG). Still, like other search methods, embedding-based retrieval may be susceptible to search-engine optimization (SEO) attacks, where adversaries promote malicious content by introducing adversarial passages to corpora. Prior work has shown such SEO is feasible, mostly demonstrating attacks against retrieval-integrated systems (e.g., RAG). Yet, these consider relaxed SEO threat models (e.g., targeting single queries), use baseline attack methods, and provide small-scale retrieval evaluation, thus obscuring our comprehensive understanding of retrievers' worst-case behavior. Matan Ben-Tov, Mahmood Sharif |
CCS | 2 |
| 2025 | Safety Perceptions of Generative AI Conversational Agents: Uncovering Perceptual Differences in Trust, Risk, and Fairness
Jan Tolsdorf, Alan F. Luo, Monica Kodwani, Junho Eum, Mahmood Sharif, Michelle L. Mazurek, Adam J. Aviv |
SOUPS | 5 |
| 2024 | Training Robust ML-based Raw-Binary Malware Detectors in Hours, not MonthsabstractMachine-learning (ML) classifiers are increasingly used to distinguish malware from benign binaries. Recent work has shown that ML-based detectors can be evaded by adversarial examples, but also that one may defend against such attacks via adversarial training. However, adversarial training, and subsequent robustness evaluation, is computationally expensive in the raw-binary malware-detection domain because it requires producing many adversarial examples for both training and evaluation. Prior work found that Greedy-training, a faster robust training technique that forgoes using adversarial examples, showed some promise in producing robust malware detectors. However, Greedy-training was far less effective in inducing robustness than the more expensive adversarial training, and it also severely hurt natural accuracy (i.e., accuracy on the original data). To faster train models, this work presents GreedyBlock-training, an enhanced version of Greedy-training that we empirically show achieves not only state-of-the-art robustness in malware detectors, exceeding even adversarial training, but also retains natural accuracy better than adversarial training. Furthermore, as it does not require creating adversarial (or functional) examples, GreedyBlock-training is significantly faster than adversarial training. Specifically, we show that GreedyBlock-training can produce more robust (+54% on average), more naturally accurate (+7% on average), and more efficiently trained (-91% average computation) malware detectors than prior work. To faster evaluate models, we also develop methods to faster gauge the robustness of ML-based raw-binary malware detectors by introducing robustness proxies, which can be used either to predict which models are likely to be the most robust, thus helping prioritize which detectors to evaluate with expensive attacks, or aiding in deciding which detectors are worthwhile to continue training. Experimentally, we show these proxy measures can find the most robust detector in a pool of detectors while using only ~20-50% of the computation that would otherwise be required. Keane Lucas, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif |
CCS | 5 |
| 2024 | Group-based Robustness: A General Framework for Customized Robustness in the Real World
Weiran Lin, Keane Lucas, Neo Eyal, Lujo Bauer, Michael K. Reiter, Mahmood Sharif |
NDSS | 6 |
| 2024 | CaFA: Cost-aware, Feasible Attacks With Database Constraints Against Neural Tabular ClassifiersabstractThis work presents CaFA, a system for Cost-aware Feasible Attacks for assessing the robustness of neural tabular classifiers against adversarial examples realizable in the problem space, while minimizing adversaries’ effort. To this end, CaFA leverages TabPGD—an algorithm we set forth to generate adversarial perturbations suitable for tabular data— and incorporates integrity constraints automatically mined by state-of-the-art database methods. After producing adversarial examples in the feature space via TabPGD, CaFA projects them on the mined constraints, leading, in turn, to better attack realizability. We tested CaFA with three datasets and two architectures and found, among others, that the constraints we use are of higher quality (measured via soundness and completeness) than ones employed in prior work. Moreover, CaFA achieves higher feasible success rates—i.e., it generates adversarial examples that are often misclassified while satisfying constraints—than prior attacks while simultaneously perturbing few features with lower magnitudes, thus saving effort and improving inconspicuousness. We open-source CaFA,1hoping it will serve as a generic system enabling machine-learning engineers to assess their models’ robustness against realizable attacks, thus advancing deployed models’ trustworthiness. Matan Ben-Tov, Daniel Deutch, Nave Frost, Mahmood Sharif |
SP | 4 |
| 2024 | DrSec: Flexible Distributed Representations for Efficient Endpoint SecurityabstractThe increasing complexity of attacks has given rise to varied security applications tackling profound tasks, ranging from alert triage to attack reconstruction. Yet, security products, such as Endpoint Detection and Response, bring together applications that are developed in isolation, trigger many false positives, miss actual attacks, and produce limited labels useful in supervised learning schemes. To address these challenges, we propose DrSec—a system employing self-supervised learning to pre-train foundation language models (LMs) that ingest event-sequence data and emit distributed representations for processes. Once pre-trained, the LMs can be adapted to solve different downstream tasks with limited to no supervision, helping unify the currently fractured application ecosystem. We trained DrSec with two LM types on a real-world dataset containing ∼91M processes and ∼2.55B events, and tested it in three application domains. We found that DrSec enables accurate, unsupervised process identification; outperforms leading methods on alert triage to reduce alert fatigue (e.g., 75.11% vs. ≤64.31% precision-recall area under curve); and accurately learns expert-developed rules, allowing tuning incident detectors to control false positives and negatives. Mahmood Sharif, Pubali Datta, Andy Riddle, Kim Westfall, Adam Bates 0001, Vijay Ganti, Matthew Lentz, David Ott |
SP | 1 |
| 2024 | A High Coverage Cybersecurity Scale Predictive of User Behavior
Yukiko Sawaya, Sarah Lu, Takamasa Isohara, Mahmood Sharif |
USENIX Security Symposium | 4 |
| 2023 | Accessorize in the Dark: A Security Analysis of Near-Infrared Face Recognition
Amit Cohen, Mahmood Sharif |
ESORICS (3) | 2 |
| 2023 | Adversarial Training for Raw-Binary Malware Classifiers
Keane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif |
USENIX Security Symposium | 6 |
| 2022 | "I Have No Idea What a Social Bot Is": On Users' Perceptions of Social Bots and Ability to Detect ThemabstractSocial bots—software agents controlling accounts on online social networks (OSNs)—have been employed for various malicious purposes, including spreading disinformation and scams. Understanding user perceptions of bots and ability to distinguish them from other accounts can inform mitigations. To this end, we conducted an online study with 297 users of seven OSNs to explore their mental models of bots and evaluate their ability to classify bots and non-bots correctly. We found that while some participants were aware of bots’ primary characteristics, others provided abstract descriptions or confused bots with other phenomena. Participants also struggled to classify accounts correctly (e.g., misclassifying > 50% of accounts) and were more likely to misclassify bots than non-bots. Furthermore, we observed that perceptions of bots had a significant effect on participants’ classification accuracy. For example, participants with abstract perceptions of bots were more likely to misclassify. Informed by our findings, we discuss directions for developing user-centered interventions against bots. Daniel Kats, Mahmood Sharif |
HAI | 2 |
| 2022 | Constrained Gradient Descent: A Powerful and Principled Evasion Attack Against Neural NetworksabstractWe propose new, more efficient targeted white-box attacks against deep neural networks. Our attacks better align with the attacker’s goal: (1) tricking a model to assign higher probability to the target class than to any other class, while (2) staying within an $\epsilon$-distance of the attacked input. First, we demonstrate a loss function that explicitly encodes (1) and show that Auto-PGD finds more attacks with it. Second, we propose a new attack method, Constrained Gradient Descent (CGD), using a refinement of our loss function that captures both (1) and (2). CGD seeks to satisfy both attacker objectives—misclassification and bounded $\ell_{p}$-norm—in a principled manner, as part of the optimization, instead of via ad hoc post-processing techniques (e.g., projection or clipping). We show that CGD is more successful on CIFAR10 (0.9–4.2%) and ImageNet (8.6–13.6%) than state-of-the-art attacks while consuming less time (11.4–18.8%). Statistical tests confirm that our attack outperforms others against leading defenses on different datasets and values of $\epsilon$. Weiran Lin, Keane Lucas, Lujo Bauer, Michael K. Reiter, Mahmood Sharif |
ICML | 5 |
| 2022 | Scalable verification of GNN-based job schedulersabstractRecently, Graph Neural Networks (GNNs) have been applied for scheduling jobs over clusters, achieving better performance than hand-crafted heuristics. Despite their impressive performance, concerns remain over whether these GNN-based job schedulers meet users’ expectations about other important properties, such as strategy-proofness, sharing incentive, and stability. In this work, we consider formal verification of GNN-based job schedulers. We address several domain-specific challenges such as networks that are deeper and specifications that are richer than those encountered when verifying image and NLP classifiers. We develop vegas, the first general framework for verifying both single-step and multi-step properties of these schedulers based on carefully designed algorithms that combine abstractions, refinements, solvers, and proof transfer. Our experimental results show that vegas achieves significant speed-up when verifying important properties of a state-of-the-art GNN-based scheduler compared to previous methods. Haoze Wu 0001, Clark W. Barrett, Mahmood Sharif, Nina Narodytska, Gagandeep Singh 0001 |
Proc. ACM Program. Lang. | 3 |
| 2021 | Malware Makeover: Breaking ML-based Static Analysis by Modifying Executable BytesabstractMotivated by the transformative impact of deep neural networks (DNNs) in various domains, researchers and anti-virus vendors have proposed DNNs for malware detection from raw bytes that do not require manual feature engineering. In this work, we propose an attack that interweaves binary-diversification techniques and optimization frameworks to mislead such DNNs while preserving the functionality of binaries. Unlike prior attacks, ours manipulates instructions that are a functional part of the binary, which makes it particularly challenging to defend against. We evaluated our attack against three DNNs in white- and black-box settings, and found that it often achieved success rates near 100%. Moreover, we found that our attack can fool some commercial anti-viruses, in certain cases with a success rate of 85%. We explored several defenses, both new and old, and identified some that can foil over 80% of our evasion attempts. However, these defenses may still be susceptible to evasion by attacks, and so we advocate for augmenting malware-detection systems with methods that do not rely on machine learning. Keane Lucas, Mahmood Sharif, Lujo Bauer, Michael K. Reiter, Saurabh Shintre |
AsiaCCS | 2 |
| 2019 | A Field Study of Computer-Security Perceptions Using Anti-Virus Customer-Support ChatsabstractUnderstanding users' perceptions of suspected computer-security problems can help us tailor technology to better protect users. To this end, we conducted a field study of users' perceptions using 189,272 problem descriptions sent to the customer-support desk of a large anti-virus vendor from 2015 to 2018. Using qualitative methods, we analyzed 650 problem descriptions to study the security issues users faced and the symptoms that led users to their own diagnoses. Subsequently, we investigated to what extent and for what types of issues user diagnoses matched those of experts. We found, for example, that users and experts were likely to agree for most issues, but not for attacks (e.g., malware infections), for which they agreed only in 44% of the cases. Our findings inform several user-security improvements, including how to automate interactions with users to resolve issues and to better communicate issues to users. Mahmood Sharif, Kevin A. Roundy, Matteo Dell'Amico, Christopher Gates 0002, Daniel Kats, Lujo Bauer, Nicolas Christin |
CHI | 1 |
| 2019 | A General Framework for Adversarial Examples with ObjectivesabstractImages perturbed subtly to be misclassified by neural networks, calledadversarial examples, have emerged as a technically deep challenge and an important concern for several application domains. Most research on adversarial examples takes as its only constraint that the perturbed images are similar to the originals. However, real-world application of these ideas often requires the examples to satisfy additional objectives, which are typically enforced through custom modifications of the perturbation process. In this article, we proposeadversarial generative nets(AGNs), a general methodology to train ageneratorneural network to emit adversarial examples satisfying desired objectives. We demonstrate the ability of AGNs to accommodate a wide range of objectives, including imprecise ones difficult to model, in two application domains. In particular, we demonstratephysicaladversarial examples—eyeglass frames designed to fool face recognition—with better robustness, inconspicuousness, and scalability than previous approaches, as well as a new attack to fool a handwritten-digit classifier. Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. Reiter |
ACM Trans. Priv. Secur. | 1 |
| 2018 | Predicting Impending Exposure to Malicious Content from User BehaviorabstractMany computer-security defenses are reactive---they operate only when security incidents take place, or immediately thereafter. Recent efforts have attempted to predict security incidents before they occur, to enable defenders to proactively protect their devices and networks. These efforts have primarily focused on long-term predictions. We propose a system that enables proactive defenses at the level of a single browsing session. By observing user behavior, it can predict whether they will be exposed to malicious content on the web seconds before the moment of exposure, thus opening a window of opportunity for proactive defenses. We evaluate our system using three months' worth of HTTP traffic generated by 20,645 users of a large cellular provider in 2017 and show that it can be helpful, even when only very low false positive rates are acceptable, and despite the difficulty of making "on-the-fly'' predictions. We also engage directly with the users through surveys asking them demographic and security-related questions, to evaluate the utility of self-reported data for predicting exposure to malicious content. We find that self-reported data can help forecast exposure risk over long periods of time. However, even on the long-term, self-reported data is not as crucial as behavioral measurements to accurately predict exposure. Mahmood Sharif, Junpei Urakawa, Nicolas Christin, Ayumu Kubota, Akira Yamada 0001 |
CCS | 1 |
| 2018 | Riding out DOMsday: Towards Detecting and Preventing DOM Cross-Site Scripting
William Melicher, Anupam Das 0001, Mahmood Sharif, Lujo Bauer, Limin Jia 0001 |
NDSS | 3 |
| 2017 | Self-Confidence Trumps Knowledge: A Cross-Cultural Study of Security BehaviorabstractComputer security tools usually provide universal solutions without taking user characteristics (origin, income level, ...) into account. In this paper, we test the validity of using such universal security defenses, with a particular focus on culture. We apply the previously proposed Security Behavior Intentions Scale (SeBIS) to 3,500 participants from seven countries. We first translate the scale into seven languages while preserving its reliability and structure validity. We then build a regression model to study which factors affect participants' security behavior. We find that participants from different countries exhibit different behavior. For instance, participants from Asian countries, and especially Japan, tend to exhibit less secure behavior. Surprisingly to us, we also find that actual knowledge influences user behavior much less than user self-confidence in their computer security knowledge. Stated differently, what people think they know affects their security behavior more than what they do know. Yukiko Sawaya, Mahmood Sharif, Nicolas Christin, Ayumu Kubota, Akihiro Nakarai, Akira Yamada 0001 |
CHI | 2 |
| 2017 | Topics of Controversy: An Empirical Analysis of Web Censorship ListsabstractAbstract Studies of Internet censorship rely on an experimental technique called probing. From a client within each country under investigation, the experimenter attempts to access network resources that are suspected to be censored, and records what happens. The set of resources to be probed is a crucial, but often neglected, element of the experimental design. We analyze the content and longevity of 758,191 webpages drawn from 22 different probe lists, of which 15 are alleged to be actual blacklists of censored webpages in particular countries, three were compiled using a priori criteria for selecting pages with an elevated chance of being censored, and four are controls. We find that the lists have very little overlap in terms of specific pages. Mechanically assigning a topic to each page, however, reveals common themes, and suggests that handcurated probe lists may be neglecting certain frequently censored topics. We also find that pages on controversial topics tend to have much shorter lifetimes than pages on uncontroversial topics. Hence, probe lists need to be continuously updated to be useful. To carry out this analysis, we have developed automated infrastructure for collecting snapshots of webpages, weeding out irrelevant material (e.g. site “boilerplate” and parked domains), translating text, assigning topics, and detecting topic changes. The system scales to hundreds of thousands of pages collected. Zachary Weinberg, Mahmood Sharif, Janos Szurdi, Nicolas Christin |
Proc. Priv. Enhancing Technol. | 2 |
| 2016 | Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionabstractMachine learning is enabling a myriad innovations, including new algorithms for cancer diagnosis and self-driving cars. The broad use of machine learning makes it important to understand the extent to which machine-learning algorithms are subject to attack, particularly when used in applications where physical security or safety is at risk. Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. Reiter |
CCS | 1 |
| 2016 | (Do Not) Track Me Sometimes: Users' Contextual Preferences for Web TrackingabstractAbstract Online trackers compile profiles on users for targeting ads, customizing websites, and selling users’ information. In this paper, we report on the first detailed study of the perceived benefits and risks of tracking-and the reasons behind them-conducted in the context of users’ own browsing histories. Prior work has studied this in the abstract; in contrast, we collected browsing histories from and interviewed 35 people about the perceived benefits and risks of online tracking in the context of their own browsing behavior. We find that many users want more control over tracking and think that controlled tracking has benefits, but are unwilling to put in the effort to control tracking or distrust current tools. We confirm previous findings that users’ general attitudes about tracking are often at odds with their comfort in specific situations. We also identify specific situational factors that contribute to users’ preferences about online tracking and explore how and why. Finally, we examine a sample of popular tools for controlling tracking and show that they only partially address the situational factors driving users’ preferences.We suggest opportunities to improve such tools, and explore the use of a classifier to automatically determine whether a user would be comfortable with tracking on a particular page visit; our results suggest this is a promising direction for future work. William Melicher, Mahmood Sharif, Lujo Bauer, Mihai Christodorescu, Pedro Giovanni Leon |
Proc. Priv. Enhancing Technol. | 2 |
| 2013 | Secure authentication from facial attributeswith no privacy lossabstractBiometric authentication is more secure than using regular passwords, as biometrics cannot be "forgotten" and contain high entropy. Thus, many constructions rely on biometric features for authentication, and use them as a source for "good" cryptographic keys. At the same time, biometric systems carry with them many privacy concerns. Orr Dunkelman, Margarita Osadchy, Mahmood Sharif |
CCS | 3 |