EDBT 2026 Demo / reviewers in the wild / expert
Sadia Afroz 0001
dblp:29/7562-1
· DBLP profile ↗
19ranked-venue papers
5as first author
2since 2021 · last 2026
0000-0002-8427-0635ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 14 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
8 papers |
Network security · 50% Digital forensics and information hiding · 29% Privacy and data protection · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Web and social media mining · 56% Data mining · 44% | |
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 100% | |
| Computer networks
3 papers |
Network measurement and analytics · 100% |
Topics — the 17 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Network security
anonymity networks |
0.5 | 2 | 2017 | Characterizing the Nature and Dynamics of Tor Exit Blocking · USENIX Security Symposium 2017 A Critical Evaluation of Website Fingerprinting Attacks · CCS 2014 |
Digital forensics and information hiding
authorship attribution |
0.3 | 2 | 2014 | Doppelgänger Finder: Taking Stylometry to the Underground · IEEE Symposium on Security and Privacy 2014 Detecting Hoaxes, Frauds, and Deception in Writing Style Online · IEEE Symposium on Security and Privacy 2012 |
Digital forensics and information hiding › authorship attribution
stylometry |
0.3 | 2 | 2014 | Doppelgänger Finder: Taking Stylometry to the Underground · IEEE Symposium on Security and Privacy 2014 Detecting Hoaxes, Frauds, and Deception in Writing Style Online · IEEE Symposium on Security and Privacy 2012 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.3 | 1 | 2017 | Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain Adaptation · EMNLP 2017 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › domain adaptive classification
fine-grained domain adaptation |
0.3 | 1 | 2017 | Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain Adaptation · EMNLP 2017 |
Data mining › text mining
information extraction and text analysis |
0.3 | 1 | 2017 | Tools for Automated Analysis of Cybercriminal Markets · WWW 2017 |
Digital forensics and information hiding › cryptocurrency forensics
blockchain transaction analysis |
0.3 | 1 | 2017 | Backpage and Bitcoin: Uncovering Human Traffickers · KDD 2017 |
Privacy and data protection
anonymity |
0.2 | 1 | 2016 | Do You See What I See? Differential Treatment of Anonymous Users · NDSS 2016 |
Network security › anonymity networks
censorship circumvention |
0.2 | 1 | 2016 | SoK: Towards Grounding Censorship Circumvention in Empiricism · IEEE Symposium on Security and Privacy 2016 |
Network security › censorship
censorship measurement |
0.2 | 1 | 2016 | SoK: Towards Grounding Censorship Circumvention in Empiricism · IEEE Symposium on Security and Privacy 2016 |
Network security › anonymity networks
tor |
0.2 | 1 | 2014 | A Critical Evaluation of Website Fingerprinting Attacks · CCS 2014 |
Network security
traffic analysis |
0.2 | 1 | 2014 | A Critical Evaluation of Website Fingerprinting Attacks · CCS 2014 |
Digital forensics and information hiding
underground forum analysis |
0.2 | 1 | 2014 | Doppelgänger Finder: Taking Stylometry to the Underground · IEEE Symposium on Security and Privacy 2014 |
Network security › traffic analysis
website fingerprinting |
0.2 | 1 | 2014 | A Critical Evaluation of Website Fingerprinting Attacks · CCS 2014 |
Biometric security
deception detection |
0.1 | 1 | 2012 | Detecting Hoaxes, Frauds, and Deception in Writing Style Online · IEEE Symposium on Security and Privacy 2012 |
Network measurement and analytics
traffic analysis |
0.1 | 1 | 2017 | Characterizing the Nature and Dynamics of Tor Exit Blocking · USENIX Security Symposium 2017 |
Network measurement and analytics
internet measurement |
0.1 | 1 | 2016 | SoK: Towards Grounding Censorship Circumvention in Empiricism · IEEE Symposium on Security and Privacy 2016 |
Methods — techniques the papers use, named apart from their topics
blocklist analysis · 0.9RIPE Atlas measurement · 0.9BitTorrent DHT crawling · 0.9stylometry · 0.6machine learning classification · 0.6dataset construction · 0.6bitcoin mempool analysis · 0.6survey · 0.5measurement study · 0.5natural language processing · 0.3machine learning · 0.3classification with verification · 0.2base rate analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring and Evaluating the Performance of Generative Ai Models for Scam DetectionabstractOnline scams continue to cause substantial financial and personal harm. As a result, detection systems based on Large Language Models (LLMs) have been integrated into security products ranging from email gateways and browser extensions to fraud-monitoring dashboards. As this adoption accelerates, a common belief has taken hold: that these models are broadly suitable for scam detection. In this work, we investigate whether LLMs, with their strong capabilities in understanding intent, context, and reasoning, can effectively detect scams across diverse scenarios without task-specific fine-tuning. We curate and release a unique benchmark dataset of real-world scams spanning multiple formats and topics. We evaluate nine LLMs of varying sizes and architectures, examining their performance under different prompting strategies and comparing them to a fine-tuned BERT-based classifier. Our results show that while larger LLMs generally outperform smaller ones, effective prompting substantially boosts the performance of smaller models. Moreover, LLMs are better at generalizing to unseen scams compared to fine-tuned models, suggesting that pre-trained knowledge contributes meaningfully to scam detection. We release our dataset and evaluation framework to facilitate future research in robust scam detection using language models. Cem Topcuoglu, Seyed Ali Akhavani, Harel Berger, Sadia Afroz 0001, Michalis Pachilakis, Vibhor Sehgal, Leyla Bilge, Engin Kirda |
COMPSAC | 4 |
| 2022 | MAB-Malware: A Reinforcement Learning Framework for Blackbox Generation of Adversarial MalwareabstractModern commercial antivirus systems increasingly rely on machine learning (ML) to keep up with the rampant inflation of new malware. However, it is well-known that machine learning models are vulnerable to adversarial examples (AEs). Previous works have shown that ML malware classifiers are fragile to the white-box adversarial attacks. However, ML models used in commercial antivirus (AV) products are usually not available to attackers and only return hard classification labels. Therefore, it is more practical to evaluate the robustness of ML models and real-world AVs in a pure black-box manner. We propose a black-box Reinforcement Learning (RL) based framework to generate AEs for PE malware classifiers and AV engines. It regards the adversarial attack problem as a multi-armed bandit problem, which finds an optimal balance between exploiting the successful patterns and exploring more varieties. Compared to other frameworks, our improvements lie in three points: 1) limiting the exploration space by modeling the generation process as a stateless process to avoid combination explosions, 2) reusing the successful payload in modeling; and 3) minimizing the changes on AE samples to correctly assign the rewards in RL learning (which also helps identify the root cause of evasions). As a result, our framework has much higher evasion rates than other off-the-shelf frameworks. Results show it has over 74%--97% evasion rate for two state-of-the-art ML detectors and over 32%--48% evasion rate for commercial AVs in a pure black-box setting. We also demonstrate that the transferability of adversarial attacks among ML-based classifiers is higher than that between ML-based classifiers and commercial AVs. Xuezixiang Li, Sadia Afroz 0001, Deepali Garg, Dmitry Kuznetsov, Heng Yin 0001 |
AsiaCCS | 3 |
| 2020 | AISec'20: 13th Workshop on Artificial Intelligence and SecurityabstractRecent years have seen a dramatic increase in applications of Artificial Intelligence (AI), Machine Learning (ML), and data mining to security and privacy problems. The analytic tools and intelligent behavior provided by these techniques make AI and ML increasingly important for autonomous real-time analysis and decision making in domains with a wealth of data or that require quick reactions to constantly changing situations. The use of learning methods in security-sensitive domains, in which adversaries may attempt to mislead or evade intelligent machines, creates new frontiers for security research. The recent widespread adoption of "deep learning'' techniques, whose security properties are difficult to reason about directly, has only added to the importance of this research. In addition, data mining and machine learning techniques create a wealth of privacy issues, due to the abundance and accessibility of data. The AISec workshop provides a venue for presenting and discussing new developments in the intersection of security and privacy with AI and machine learning. Sadia Afroz 0001, Nicholas Carlini, Ambra Demontis |
CCS | 1 |
| 2020 | Quantifying the Impact of Blocklisting in the Age of Address ReuseabstractBlocklists, consisting of known malicious IP addresses, can be used as a simple method to block malicious traffic. However, blocklists can potentially lead to unjust blocking of legitimate users due to IP address reuse, where more users could be blocked than intended. IP addresses can be reused either at the same time (Network Address Translation) or over time (dynamic addressing). We propose two new techniques to identify reused addresses. We built a crawler using the BitTorrent Distributed Hash Table to detect NATed addresses and use the RIPE Atlas measurement logs to detect dynamically allocated address spaces. We then analyze 151 publicly available IPv4 blocklists to show the implications of reused addresses and find that 53-60% of blocklists contain reused addresses having about 30.6K-45.1K listings of reused addresses. We also find that reused addresses can potentially affect as many as 78 legitimate users for as many as 44 days. Sivaramakrishnan Ramanathan, Anushah Hossain, Jelena Mirkovic, Minlan Yu, Sadia Afroz 0001 |
Internet Measurement Conference | 5 |
| 2019 | AISec'19: 12th ACM Workshop on Artificial Intelligence and SecurityabstractRecent years have seen a dramatic increase in applications of Artificial Intelligence (AI) and Machine Learning (ML) to security and privacy problems. The analytic tools and intelligent behavior provided by these techniques make AI and ML increasingly important for autonomous real-time analysis and decision making in domains with a wealth of data or that require quick reactions to constantly changing situations. The use of learning methods in security-sensitive domains, in which adversaries may attempt to mislead or evade intelligent machines, creates new frontiers for security research. The recent widespread adoption of deep-learning techniques, whose security properties are difficult to reason about directly, has only added to the importance of this research. In addition, data mining and machine learning techniques create a wealth of privacy issues, due to the abundance and accessibility of data. The 12th ACM Workshop on Artificial Intelligence and Security (AISec) is one of the historical, leading venues for presenting and discussing new developments in the intersection of security and privacy with AI and ML. Sadia Afroz 0001, Battista Biggio, Nicholas Carlini, Yuval Elovici, Asaf Shabtai |
CCS | 1 |
| 2018 | 11th International Workshop on Artificial Intelligence and Security (AISec 2018)
Sadia Afroz 0001, Battista Biggio, Yuval Elovici, David Mandell Freeman, Asaf Shabtai |
CCS | 1 |
| 2017 | Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain AdaptationabstractGreg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca Portnoff, Sadia Afroz, Damon McCoy, Kirill Levchenko, Vern Paxson. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca S. Portnoff, Sadia Afroz 0001, Damon McCoy, Kirill Levchenko, Vern Paxson |
EMNLP | 5 |
| 2017 | Backpage and Bitcoin: Uncovering Human TraffickersabstractSites for online classified ads selling sex are widely used by human traffickers to support their pernicious business. The sheer quantity of ads makes manual exploration and analysis unscalable. In addition, discerning whether an ad is advertising a trafficked victim or an independent sex worker is a very difficult task. Very little concrete ground truth (i.e., ads definitively known to be posted by a trafficker) exists in this space. In this work, we develop tools and techniques that can be used separately and in conjunction to group sex ads by their true owner (and not the claimed author in the ad). Specifically, we develop a machine learning classifier that uses stylometry to distinguish between ads posted by the same vs. different authors with 90% TPR and 1% FPR. We also design a linking technique that takes advantage of leakages from the Bitcoin mempool, blockchain and sex ad site, to link a subset of sex ads to Bitcoin public wallets and transactions. Finally, we demonstrate via a 4-week proof of concept using Backpage as the sex ad site, how an analyst can use these automated approaches to potentially find human traffickers. Rebecca S. Portnoff, Danny Yuxing Huang, Periwinkle Doerfler, Sadia Afroz 0001, Damon McCoy |
KDD | 4 |
| 2017 | Characterizing the Nature and Dynamics of Tor Exit Blocking
Rachee Singh, Rishab Nithyanand, Sadia Afroz 0001, Paul Pearce, Michael Carl Tschantz, Phillipa Gill, Vern Paxson |
USENIX Security Symposium | 3 |
| 2017 | Tools for Automated Analysis of Cybercriminal MarketsabstractUnderground forums are widely used by criminals to buy and sell a host of stolen items, datasets, resources, and criminal services. These forums contain important resources for understanding cybercrime. However, the number of forums, their size, and the domain expertise required to understand the markets makes manual exploration of these forums unscalable. In this work, we propose an automated, top-down approach for analyzing underground forums. Our approach uses natural language processing and machine learning to automatically generate high-level information about underground forums, first identifying posts related to transactions, and then extracting products and prices. We also demonstrate, via a pair of case studies, how an analyst can use these automated approaches to investigate other categories of products and transactions. We use eight distinct forums to assess our tools: Antichat, Blackhat World, Carders, Darkode, Hack Forums, Hell, L33tCrew and Nulled. Our automated approach is fast and accurate, achieving over 80% accuracy in detecting post category, product, and prices. Rebecca S. Portnoff, Sadia Afroz 0001, Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Damon McCoy, Kirill Levchenko, Vern Paxson |
WWW | 2 |
| 2016 | Reviewer Integration and Performance Measurement for Malware Detection
Brad Miller 0002, Alex Kantchelian, Michael Carl Tschantz, Sadia Afroz 0001, Rekha Bachwani, Riyaz Faizullabhoy, Ling Huang 0001, Vaishaal Shankar, Tony Wu 0002, George Yiu, Anthony D. Joseph, J. D. Tygar |
DIMVA | 4 |
| 2016 | Do You See What I See? Differential Treatment of Anonymous Users
Sheharbano Khattak, David Fifield, Sadia Afroz 0001, Mobin Javed, Srikanth Sundaresan, Damon McCoy, Vern Paxson, Steven J. Murdoch |
NDSS | 3 |
| 2016 | SoK: Towards Grounding Censorship Circumvention in EmpiricismabstractEffective evaluations of approaches to circumventing government Internet censorship require incorporating perspectives of how censors operate in practice. We undertake an extensive examination of real censors by surveying prior measurement studies and analyzing field reports and bug tickets from practitioners. We assess both deployed circumvention approaches and research proposals to consider the criteria employed in their evaluations and compare these to the observed behaviors of real censors, identifying areas where evaluations could more faithfully and effectively incorporate the practices of modern censors. These observations lead to an agenda realigning research with the predominant problems of today. Michael Carl Tschantz, Sadia Afroz 0001, Vern Paxson |
IEEE Symposium on Security and Privacy | 2 |
| 2014 | A Critical Evaluation of Website Fingerprinting AttacksabstractRecent studies on Website Fingerprinting (WF) claim to have found highly effective attacks on Tor. However, these studies make assumptions about user settings, adversary capabilities, and the nature of the Web that do not necessarily hold in practical scenarios. The following study critically evaluates these assumptions by conducting the attack where the assumptions do not hold. We show that certain variables, for example, user's browsing habits, differences in location and version of Tor Browser Bundle, that are usually omitted from the current WF model have a significant impact on the efficacy of the attack. We also empirically show how prior work succumbs to the base rate fallacy in the open-world scenario. We address this problem by augmenting our classification method with a verification step. We conclude that even though this approach reduces the number of false positives over 63\%, it does not completely solve the problem, which remains an open issue for WF attacks. Marc Juarez, Sadia Afroz 0001, Gunes Acar, Claudia Díaz, Rachel Greenstadt |
CCS | 2 |
| 2014 | Breaking the Closed-World Assumption in Stylometric Authorship Attribution
Ariel Stolerman, Rebekah Overdorf, Sadia Afroz 0001, Rachel Greenstadt |
IFIP Int. Conf. Digital Forensics | 3 |
| 2014 | Doppelgänger Finder: Taking Stylometry to the UndergroundabstractStylometry is a method for identifying anonymous authors of anonymous texts by analyzing their writing style. While stylometric methods have produced impressive results in previous experiments, we wanted to explore their performance on a challenging dataset of particular interest to the security research community. Analysis of underground forums can provide key information about who controls a given bot network or sells a service, and the size and scope of the cybercrime underworld. Previous analyses have been accomplished primarily through analysis of limited structured metadata and painstaking manual analysis. However, the key challenge is to automate this process, since this labor intensive manual approach clearly does not scale. We consider two scenarios. The first involves text written by an unknown cybercriminal and a set of potential suspects. This is standard, supervised stylometry problem made more difficult by multilingual forums that mix l33t-speak conversations with data dumps. In the second scenario, you want to feed a forum into an analysis engine and have it output possible doppelgangers, or users with multiple accounts. While other researchers have explored this problem, we propose a method that produces good results on actual separate accounts, as opposed to data sets created by artificially splitting authors into multiple identities. For scenario 1, we achieve 77% to 84% accuracy on private messages. For scenario 2, we achieve 94% recall with 90% precision on blogs and 85.18% precision with 82.14% recall for underground forum users. We demonstrate the utility of our approach with a case study that includes applying our technique to the Carders forum and manual analysis to validate the results, enabling the discovery of previously undetected doppelganger accounts. Sadia Afroz 0001, Aylin Caliskan, Ariel Stolerman, Rachel Greenstadt, Damon McCoy |
IEEE Symposium on Security and Privacy | 1 |
| 2012 | Use Fewer Instances of the Letter "i": Toward Writing Style Anonymization
Andrew W. E. McDonald, Sadia Afroz 0001, Aylin Caliskan, Ariel Stolerman, Rachel Greenstadt |
Privacy Enhancing Technologies | 2 |
| 2012 | Detecting Hoaxes, Frauds, and Deception in Writing Style OnlineabstractIn digital forensics, questions often arise about the authors of documents: their identity, demographic background, and whether they can be linked to other documents. The field of stylometry uses linguistic features and machine learning techniques to answer these questions. While stylometry techniques can identify authors with high accuracy in non-adversarial scenarios, their accuracy is reduced to random guessing when faced with authors who intentionally obfuscate their writing style or attempt to imitate that of another author. While these results are good for privacy, they raise concerns about fraud. We argue that some linguistic features change when people hide their writing style and by identifying those features, stylistic deception can be recognized. The major contribution of this work is a method for detecting stylistic deception in written documents. We show that using a large feature set, it is possible to distinguish regular documents from deceptive documents with 96.6% accuracy (F-measure). We also present an analysis of linguistic features that can be modified to hide writing style. Sadia Afroz 0001, Michael Brennan, Rachel Greenstadt |
IEEE Symposium on Security and Privacy | 1 |
| 2012 | Adversarial stylometry: Circumventing authorship recognition to preserve privacy and anonymityabstractThe use of stylometry, authorship recognition through purely linguistic means, has contributed to literary, historical, and criminal investigation breakthroughs. Existing stylometry research assumes that authors have not attempted to disguise their linguistic writing style. We challenge this basic assumption of existing stylometry methodologies and present a new area of research: adversarial stylometry. Adversaries have a devastating effect on the robustness of existing classification methods. Our work presents a framework for creating adversarial passages including obfuscation , where a subject attempts to hide her identity, and imitation , where a subject attempts to frame another subject by imitating his writing style, and translation where original passages are obfuscated with machine translation services. This research demonstrates that manual circumvention methods work very well while automated translation methods are not effective. The obfuscation method reduces the techniques' effectiveness to the level of random guessing and the imitation attempts succeed up to 67% of the time depending on the stylometry technique used. These results are more significant given the fact that experimental subjects were unfamiliar with stylometry, were not professional writers, and spent little time on the attacks. This article also contributes to the field by using human subjects to empirically validate the claim of high accuracy for four current techniques (without adversaries). We have also compiled and released two corpora of adversarial stylometry texts to promote research in this field with a total of 57 unique authors. We argue that this field is important to a multidisciplinary approach to privacy, security, and anonymity. Michael Brennan, Sadia Afroz 0001, Rachel Greenstadt |
ACM Trans. Inf. Syst. Secur. | 2 |