Roberto Perdisci

dblp:60/6768 · DBLP profile ↗
← Back
73ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-7339-0041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 60 · 7 first-author · 14 since 2021Computer networks · 8 · 3 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-authorArtificial intelligence and machine learning · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 DNS Trap: Unveiling Reactive DNS Monitoring via Stimulated Network Emissions
Aaron Faulkenberry, Athanasios Avgetidis, Omar Alrawi, Zane Ma, Roberto Perdisci, Manos Antonakakis
EuroS&P5
2025 Revealing the True Indicators: Understanding and Improving IoC Extraction From Threat Reports
abstract
Indicators of Compromise (IoCs) are critical for threat detection and response, marking malicious activity across networks and systems. Yet, the effectiveness of automated IoC extraction systems is fundamentally limited by one key issue: the lack of high-quality ground truth. Current extraction tools rely either on manually extracted ground truth, which is labor-intensive and costly, or on automated ground truth creation methods that include non-malicious artifacts, leading to inflated false positive (FP) rates and unreliable threat intelligence. In this work, we analyze the shortcomings of existing ground truth creation strategies and address them by introducing the first hybrid human-in-the-Ioop pipeline for IoC extraction, which combines a large language model-based classifier (LANCE) with expert analyst validation. Our system improves precision through explainable, context-aware labeling and reduces analysts' work factor by 43% compared to manual annotation, as demonstrated in our evaluation with six analysts. Using this approach, we produce PRISM, a high-quality, publicly available benchmark of 1,791 labeled IoCs from 50 real-world threat reports. PRISM supports both fair evaluation and training of IoC extraction methods and enables reproducible research grounded in expert-validated indicators.
Evangelos Froudakis, Athanasios Avgetidis, Sean Tyler Frankum, Roberto Perdisci, Manos Antonakakis, Angelos D. Keromytis
ACSAC4
2025 PP3D: An In-Browser Vision-Based Defense Against Web Behavior Manipulation Attacks
abstract
Web-based behavior-manipulation attacks (BMAs)—such as scareware, fake software downloads, tech support scams, etc.—are a class of social engineering (SE) attacks that exploit human decision-making vulnerabilities. These attacks remain under-studied compared to other attacks such as information harvesting attacks (e.g., phishing) or malware infections. Prior technical work has primarily focused on measuring BMAs, offering little in the way of generic defenses. To address this gap, we introduce Pixel Patrol 3D (PP3D), the first end-to-end browser framework for discovering, detecting, and defending against behavior-manipulating SE attacks in real time. PP3D consists of a visual detection model implemented within a browser extension, which deploys the model client-side to protect users across desktop and mobile devices while preserving privacy. Our evaluation shows that PP3D can achieve above 99% detection rate at 1% false positives, while maintaining good latency and overhead performance across devices. Even when faced with new BMA samples collected months after training the detection model, our defense system can still achieve above 97% detection rate at 1% false positives. These results demonstrate that our framework offers a practical, effective, and generalizable defense against a broad and evolving class of web behavior-manipulation attacks.
Spencer King, Irfan Ozen, Karthika Subramani, Saranyan Senthivel, Phani Vadrevu, Roberto Perdisci
ACSAC6
2025 From Concealment to Exposure: Understanding the Lifecycle and Infrastructure of APT Domains
abstract
Advanced Persistent Threats (APTs) are sophisticated and long-lived attacks that are often backed by nationstates. Despite the security community’s efforts to design and deploy specialized systems to combat them, APTs have remained prevalent while persisting undetected for significantly more time than commodity cyber threats. In this paper, we measure this difference by conducting the first longitudinal analysis of APT infrastructure by shedding light on the lifecycle of their domain names. To enable this study, we build Atropos, a novel measurement methodology that automatically and accurately labels DNS records of APT domain names, enabling us to understand their lifecycle and gain a more comprehensive and contextualized infrastructure picture than the one that is shared in public reports. Using the comprehensive infrastructure view that Atropos provides, we study 405 APT actors over a period spanning a decade and unveil several novel findings regarding their utilization of network infrastructure that have practical implications. We find that APT actors provision their IPs to their domain names 317 days on average before an attack is publicly reported. Furthermore, $73.6 \%$ of the APT IPs that are part of the attack infrastructure no longer point to their domains at the time of first public disclosure, highlighting that researchers and security practitioners need to consider historic DNS data in order to get a more comprehensive and accurate picture when training network detection, investigation, or attribution systems. Organizations that are more sensitive to APT attacks will need to retain network logs for at least 19 to 25 months in order to have higher probabilities of discovering whether they have been a target of an APT attack. Finally, we provide evidence that APT actors re-use hosting providers, deploy APT network infrastructure close to their intended attack targets, and increasingly utilize more cloud-fronting. These findings are important because they can guide future threat detection and attribution works.
Athanasios Avgetidis, Aaron Faulkenberry, Boladji Vinny Adjibi, Tillson Galloway, Panagiotis Kintis, Omar Alrawi, Zane Ma, Fabian Monrose, Angelos D. Keromytis, Roberto Perdisci, Manos Antonakakis
RAID10
2025 Uncontained Danger: Quantifying Remote Dependencies in Containerized Applications
abstract
Containers benefit software developers, aiding them with increased portability, scalability, and consistency across different environments. From a security point of view, containers incentivize the conversion of monolithic software into microservices which then can be isolated from each other, better lending themselves to least-privilege deployments. In this paper, we shed light to the unexplored issue of remote dependencies in Docker images and containers. Unless a Docker image is fully self-contained, every dependence to the outside world is an opportunity for attackers to hijack these dependencies and conduct supply-chain attacks against these images. To do so, we curate a dataset of 200 K Docker images and design DockerGym, a dynamic analysis system which automatically installs, executes, and stimulates running containers, while monitoring their network communications. We discuss multiple approaches for activating the images in our dataset and the types of remote dependencies that we were able to discover. Among others, we observe that 13% of evaluated Docker images have some form of remote dependencies, with approximately 10 K images resolving public domain names. We observe the use of unencrypted protocols (such as HTTP) and a range of other issues that could be straightforwardly exploited by attackers in the context of supply-chain attacks.
Chris Tsoukaladelis, Roberto Perdisci, Nick Nikiforakis
RAID2
2025 Online Adaptive Anomaly Detection in Networked Electrical Machines by Adaptive Enveloped Singular Spectrum Transformation
abstract
The emergence of networked electrical machines has increased susceptibility to anomalies, including cyber-attack and physical faults, potentially leading to significant operational disruptions. In this article, we propose an online adaptive anomaly detection algorithm, adaptive enveloped singular spectrum transformation (AdaESST), which aims to identify hard-to-detect anomalies effectively. AdaESST first extracts informative components of signals by embedding the waveform data into subspaces using singular value decomposition, and then calculates anomalous score based on the subspace distance between two subsequence time series. AdaESST outperforms traditional detection methods by its capacity to adjust to new operational scenarios, thereby offering persistent protection in dynamic industrial environments. Throughout all numerical experiments simulating real-world industrial conditions, AdaESST exhibits high detection accuracy in monitoring motor and point of common coupling (PCC) currents, demonstrating its capability to safeguard against sophisticated anomalies. The detection accuracy for PCC currents is on par with that for motor currents. In essence, AdaESST has the potential to reduce the requirements for sensors, thereby lowering maintenance costs while maintaining high data integrity and security. The work contributes to enhancing the security of networked electrical machines, presenting a resilient and cost-efficient strategy in the face of emerging anomalies.
Shushan Wu, Stephen James Coshatt, Xilin Gong, Ramviyas Parasuraman, Justin Conrad, Roberto Perdisci, Wenxuan Zhong, Jin Ye 0001, Ping Ma 0001, Wen-Zhan Song 0001
IEEE Internet Things J.8
2024 Practical Attacks Against DNS Reputation Systems
abstract
DNS reputation systems are a critical layer of network defense that use ML to identify potentially malicious domains based on DNS-related behaviors. Despite their importance in protecting against spam, malware, and social engineering, little is known about the adversarial robustness of real-world DNS reputation systems. This work takes a first look at general attacks against DNS reputation systems. To overcome the black-box setting of deployed DNS reputation systems, we begin by creating an open-source reference DNS reputation system that 1) overcomes common pitfalls in data collection, preprocessing, training, and evaluation found in prior work, 2) approximates DNS reputation systems from prior research, and 3) enables future reproducible research. We find that general adversarial ML techniques are impractical due to a highly constrained input space, complex feature interdependencies, and difficult inversion from feature vectors to raw input samples. We then implement two classes of practical attacks, mimicry and popularity manipulation, that achieve high success rates against both our reference model and a popular commercial DNS reputation system, highlighting the transferability of the attacks to the real world. Finally, we develop constraint models that assess the time and financial cost required to execute our attacks. Using these models, we demonstrate that an adversary with US$10 can evade a leading security vendor with a 100% success rate in two weeks.
Tillson Galloway, Kleanthis Karakolios, Zane Ma, Roberto Perdisci, Angelos D. Keromytis, Manos Antonakakis
SP4
2024 C-Frame: Characterizing and measuring in-the-wild CAPTCHA attacks
abstract
In this paper, we design and implement C-Frame, the first measurement system to collect real-time, in-the-wild data on modern CAPTCHA attacks. For this, we study the recent evolution in the protocols of CAPTCHAs as well as human-driven farms that facilitate attacks against CAPTCHAs. This study leads us directly to the discovery of a unique vantage point to conduct a global-scale CAPTCHA attack measurement study. Harnessing this, we design and build C-Frame to be CAPTCHA-agnostic and ethically considerate. We then deploy our system for a 92-day period resulting in capturing of 425,257 CAPTCHA attacks on 1417 sites.In order to characterize these attacks, we leverage a carefully designed qualitative analysis approach using 3 analysts. Our study results in delineation of 34 different CAPTCHA-attack categories with several interesting real world attack examples. Twitter received the largest number of CAPTCHA attacks overall (about 255,480 attack requests) most of which attempt to create bot accounts. We also categorized and captured attacks such as ticket scalping attempts (e.g. a Taylor Swift concert event in Brazil), fraudulent lawsuit claims, and abusive appointment booking attempts (e.g. a Spain visa site in China). We also found CAPTCHA-assisted attempts to download data from government website (e.g. websites of 20 US states). We ascribe our attacks to 58 different countries across 5 continents. We present a detailed measurement analysis to give insights on this attack data and also suggest some future potential remediation measures that can be inspired by our system.
Hoang Dai Nguyen, Karthika Subramani, Bhupendra Acharya, Roberto Perdisci, Phani Vadrevu
SP4
2024 WEBRR: A Forensic System for Replaying and Investigating Web-Based Attacks in The Modern Web
Joey Allen, Matthew Landen, Roberto Perdisci, Wenke Lee
USENIX Security Symposium5
2024 Discovering and Measuring CDNs Prone to Domain Fronting
abstract
Domain fronting is a network communication technique that involves leveraging (or abusing) content delivery networks (CDNs) to disguise the final destination of network packets by presenting them as if they were intended for a different domain than their actual endpoint. This technique can be used for both benign and malicious purposes, such as circumventing censorship or hiding malware-related communications from network security systems. Since domain fronting has been known for a few years, some popular CDN providers have implemented traffic filtering approaches to curb its use at their CDN infrastructure. However, it remains unclear to what extent domain fronting has been mitigated.
Karthika Subramani, Roberto Perdisci, Pierros Skafidas, Manos Antonakakis
WWW2
2023 Understanding, Measuring, and Detecting Modern Technical Support Scams
abstract
Technical support scams (TSS) are social engineering attacks that aim to exploit users that have limited knowledge about technology, such as the elderly, causing significant financial loss to vulnerable citizens. The security community has attempted to respond to these web-based scams with different countermeasures. However, to the best of our knowledge, no robust countermeasures have been proposed thus far to defend against modern TSS campaigns that abuse web search engines to inflate their rankings in search results and lure many potential victims.To defend against these TSS attacks, in this paper we first study the TSS ecosystem, with particular focus on how modern TSS campaigns are operated and promoted on the web. Then, we capitalize on our findings by proposing a novel detection system named TASR that can be used to differentiate TSS websites from legitimate technical support websites in a topic-agnostic way, by leveraging features that capture key traits of how TSS web pages are promoted. Our cross-validation tests show that TASR can detect 94.5% of the TSS links in web search results at a false positive rate of less than 1%, significantly outperforming previous work.
Jienan Liu, Pooja Pun, Phani Vadrevu, Roberto Perdisci
EuroS&P4
2023 Combating Robocalls with Phone Virtual Assistant Mediated Interaction
Sharbani Pandit, Krishanu Sarker, Roberto Perdisci, Mustaque Ahamad, Diyi Yang
USENIX Security Symposium3
2023 TRIDENT: Towards Detecting and Mitigating Web-based Social Engineering Attacks
Joey Allen, Matthew Landen, Roberto Perdisci, Wenke Lee
USENIX Security Symposium4
2022 SoK: Workerounds - Categorizing Service Worker Attacks and Mitigations
abstract
Service Workers (SWs) are a powerful feature at the core of Progressive Web Apps, namely web applications that can continue to function when the user's device is offline and that have access to device sensors and capabilities previously accessible only by native applications. During the past few years, researchers have found a number of ways in which SWs may be abused to achieve different malicious purposes. For instance, SWs may be abused to build a web-based botnet, launch DDoS attacks, or perform cryptomining; they may be hijacked to create persistent cross-site scripting (XSS) attacks; they may be leveraged in the context of side-channel attacks to compromise users' privacy; or they may be abused for phishing or social engineering attacks using web push notifications-based malvertising. In this paper, we reproduce and analyze known attack vectors related to SWs and explore new abuse paths that have not previously been considered. We systematize the attacks into different categories, and then analyze whether, how, and estimate when these attacks have been published and mitigated by different browser vendors. Then, we discuss a number of open SW security problems that are currently unmitigated, and propose SW behavior monitoring approaches and new browser policies that we believe should be implemented by browsers to further improve SW security. Furthermore, we implement a proof-of-concept version of several policies in the Chromium code base, and also measure the behavior of SWs used by highly popular web applications with respect to these new policies. Our measurements show that it should be feasible to implement and enforce stricter SW security policies without a significant impact on most legitimate production SWs.
Karthika Subramani, Jordan Jueckstock, Alexandros Kapravelos, Roberto Perdisci
EuroS&P4
2022 PhishInPatterns: measuring elicited user interactions at scale on phishing websites
abstract
Despite phishing attacks and detection systems being extensively studied, phishing is still on the rise and has recently reached an all-time high. Attacks are becoming increasingly sophisticated, leveraging new web design patterns to add perceived legitimacy and, at the same time, evade state-of-the-art detectors and web security crawlers.
Karthika Subramani, William Melicher, Oleksii Starov, Phani Vadrevu, Roberto Perdisci
IMC5
2021 Detecting and Measuring In-The-Wild DRDoS Attacks at IXPs
Karthika Subramani, Roberto Perdisci, Maria Konte
DIMVA2
2021 C^2SR: Cybercrime Scene Reconstruction for Post-mortem Forensic Analysis
Yonghwi Kwon 0001, Weihang Wang 0001, Jinho Jung 0001, Kyu Hyung Lee, Roberto Perdisci
NDSS5
2020 Towards a Practical Differentially Private Collaborative Phone Blacklisting System
abstract
Spam phone calls have been rapidly growing from nuisance to an increasingly effective scam delivery tool. To counter this increasingly successful attack vector, a number of commercial smartphone apps that promise to block spam phone calls have appeared on app stores, and are now used by hundreds of thousands or even millions of users. However, following a business model similar to some online social network services, these apps often collect call records or other potentially sensitive information from users’ phones with little or no formal privacy guarantees.
Daniele Ucci, Roberto Perdisci, Mustaque Ahamad
ACSAC2
2020 Mnemosyne: An Effective and Efficient Postmortem Watering Hole Attack Investigation System
abstract
Compromising a website that is routinely visited by employees of a targeted organization has become a popular technique for nation-state level adversaries to penetrate an enterprise's network. This technique, dubbed a "watering hole" attack, leverages a compromised website to serve as a stepping stone into the true victims' network. Despite watering hole attacks being one of the main techniques used by attackers to achieve the initial compromise stage of the cyber kill chain, there has been relatively little research related to detecting or investigating complex watering hole attacks. While there is existing work that seeks to detect malicious modifications made to an otherwise benign website, we argue that simply detecting that the website is compromised is only the first stage of the investigation. In this paper, we propose Mnemosyne, a postmortem forensic analysis engine that relies on browser-based attack provenance to accurately reconstruct, investigate, and assess the ramifications of watering hole attacks. Mnemosyne relies on a lightweight browser-modification-free auditing daemon to passively collect causality logs related to the browser's execution. Next, Mnemosyne applies a set of versioning techniques on top of these causality logs to precisely pinpoint when the website was compromised and what modifications were made by the adversary. Following this step, Mnemosyne relies on a novel user-level analysis to assess how the malicious modifications affected the targeted enterprise and seeks to identify exactly which employees fell victim to the attack. Throughout our extensive evaluation, we found that Mnemosyne's forensic analysis engine was able to identify the true victims in all seven real-world watering hole scenarios, while also reducing the amount of manual analysis required by the forensic analyst by 98.17% on average.
Joey Allen, Matthew Landen, Raghav Bhat, Harsh Grover, Yang Ji 0002, Roberto Perdisci, Wenke Lee
CCS8
2020 IoTFinder: Efficient Large-Scale Identification of IoT Devices via Passive DNS Traffic Analysis
abstract
Being able to enumerate potentially vulnerable IoT devices across the Internet is important, because it allows for assessing global Internet risks and enables network operators to check the hygiene of their own networks. To this end, in this paper we propose IoTFinder, a system for efficient, large-scalepassiveidentification of IoT devices. Specifically, we leverage distributed passive DNS data collection, and develop a machine learning-based system that aims to accurately identify a large variety of IoT devices based solely on theirDNS fingerprints. Our system is independent of whether the devices reside behind a NAT or other middleboxes, or whether they are assigned an IPv4 or IPv6 address. We design IoTFinder as a multi-label classifier, and evaluate its accuracy in several different settings, including computing detection results over a third-party IoT traffic dataset and DNS traffic collected at a US-based ISP hosting more than 40 million clients. The experimental results show that our approach allows for accurately detecting many diverse IoT devices, even when they are hosted behind a NAT and their traffic is “mixed” with traffic generated by other IoT and non-IoT devices hosted in the same local network.
Roberto Perdisci, Thomas Papastergiou, Omar Alrawi, Manos Antonakakis
EuroS&P1
2020 When Push Comes to Ads: Measuring the Rise of (Malicious) Push Advertising
abstract
The rapid growth of online advertising has fueled the growth of ad-blocking software, such as new ad-blocking and privacy-oriented browsers or browser extensions. In response, both ad publishers and ad networks are constantly trying to pursue new strategies to keep up their revenues. To this end, ad networks have started to leverage the Web Push technology enabled by modern web browsers. As web push notifications (WPNs) are relatively new, their role in ad delivery has not been yet studied in depth. Furthermore, it is unclear to what extent WPN ads are being abused for malvertising (i.e., to deliver malicious ads). In this paper, we aim to fill this gap. Specifically, we propose a system called PushAdMiner that is dedicated to (1) automatically registering for and collecting a large number of web-based push notifications from publisher websites, (2) finding WPN-based ads among these notifications, and (3) discovering malicious WPN-based ad campaigns. Using PushAdMiner, we collected and analyzed 21,541 WPN messages by visiting thousands of different websites. Among these, our system identified 572 WPN ad campaigns, for a total of 5,143 WPN-based ads that were pushed by a variety of ad networks. Furthermore, we found that 51% of all WPN ads we collected are malicious, and that traditional ad-blockers and malicious URL filters are remarkably ineffective against WPN-based malicious ads, leaving a significant abuse vector unchecked.
Karthika Subramani, Xingzi Yuan, Omid Setayeshfar, Phani Vadrevu, Kyu Hyung Lee, Roberto Perdisci
Internet Measurement Conference6
2019 What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns
abstract
Malicious ads often use social engineering (SE) tactics to coax users into downloading unwanted software, purchasing fake products or services, or giving up valuable personal information. These ads are often served by low-tier ad networks that may not have the technical means (or simply the will) to patrol the ad content they serve to curtail abuse.
Phani Vadrevu, Roberto Perdisci
Internet Measurement Conference2
2018 Towards Measuring the Role of Phone Numbers in Twitter-Advertised Spam
abstract
The telephony channel has become an attractive target for cyber criminals, who are using it to craft a variety of attacks. In addition to delivering voice and messaging spam, this channel is also being used to lure victims into calling phone numbers that are controlled by the attackers. One way this is done is by aggressively advertising phone numbers on social media (e.g., Twitter). This form of spam is then monetized over the telephony channel, via messages/calls made by victims. We refer to this type of attacks as outgoing phone communication (OPC) attacks.
Payas Gupta, Roberto Perdisci, Mustaque Ahamad
AsiaCCS2
2018 Augmenting Telephone Spam Blacklists by Mining Large CDR Datasets
abstract
Telephone spam has become an increasingly prevalent problem in many countries all over the world. For example, the US Federal Trade Commission's (FTC) National Do Not Call Registry's number of cumulative complaints of spam/scam calls reached 30.9 million submissions in 2016. Naturally, telephone carriers can play an important role in the fight against spam. However, due to the extremely large volume of calls that transit across large carrier networks, it is challenging to mine their vast amounts of call detail records (CDRs) to accurately detect and block spam phone calls. This is because CDRs only contain high-level metadata (e.g., source and destination numbers, call start time, call duration, etc.) related to each phone calls. In addition, ground truth about both benign and spam-related phone numbers is often very scarce (only a tiny fraction of all phone numbers can be labeled). More importantly, telephone carriers are extremely sensitive to false positives, as they need to avoid blocking any non-spam calls, making the detection of spam-related numbers even more challenging. In this paper, we present a novel detection system that aims to discover telephone numbers involved in spam campaigns. Given a small seed of known spam phone numbers, our system uses a combination of unsupervised and supervised machine learning methods to mine new, previously unknown spam numbers from large datasets of call detail records (CDRs). Our objective is not to detect all possible spam phone calls crossing a carrier's network, but rather to expand the list of known spam numbers while aiming for zero false positives, so that the newly discovered numbers may be added to a phone blacklist, for example. To evaluate our system, we have conducted experiments over a large dataset of real-world CDRs provided by a leading telephony provider in China, while tuning the system to produce no false positives. The experimental results show that our system is able to greatly expand on the initial seed of known spam numbers by up to about 250%.
Jienan Liu, Babak Rahbarinia, Roberto Perdisci, Haitao Du
AsiaCCS3
2018 JSgraph: Enabling Reconstruction of Web Attacks via Efficient Tracking of Live In-Browser JavaScript Executions
Bo Li 0058, Phani Vadrevu, Kyu Hyung Lee, Roberto Perdisci
NDSS4
2018 Towards Measuring the Effectiveness of Telephony Blacklists
Sharbani Pandit, Roberto Perdisci, Mustaque Ahamad, Payas Gupta
NDSS2
2017 Practical Attacks Against Graph-based Clustering
abstract
Graph modeling allows numerous security problems to be tackled in a general way, however, little work has been done to understand their ability to withstand adversarial attacks. We design and evaluate two novel graph attacks against a state-of-the-art network-level, graph-based detection system. Our work highlights areas in adversarial machine learning that have not yet been addressed, specifically: graph-based clustering techniques, and a global feature space where realistic attackers without perfect knowledge must be accounted for (by the defenders) in order to be practical. Even though less informed attackers can evade graph clustering with low cost, we show that some practical defenses are possible.
Yizheng Chen 0001, Yacin Nadji, Athanasios Kountouras, Fabian Monrose, Roberto Perdisci, Manos Antonakakis, Nikolaos Vasiloglou
CCS5
2017 Exploring the Long Tail of (Malicious) Software Downloads
abstract
In this paper, we present a large-scale study of global trends in software download events, with an analysis of both benign and malicious downloads, and a categorization of events for which no ground truth is currently available. Our measurement study is based on a unique, real-world dataset collected at Trend Micro containing more than 3 million in-the-wild web-based software download events involving hundreds of thousands of Internet machines, collected over a period of seven months. Somewhat surprisingly, we found that despite our best efforts and the use of multiple sources of ground truth, more than 83% of all downloaded software files remain unknown, i.e. cannot be classified as benign or malicious, even two years after they were first observed. If we consider the number of machines that have downloaded at least one unknown file, we find that more than 69% of the entire machine/user population downloaded one or more unknown software file. Because the accuracy of malware detection systems reported in the academic literature is typically assessed only over software files that can be labeled, our findings raise concerns on their actual effectiveness in large-scale real-world deployments, and on their ability to defend the majority of Internet machines from infection. To better understand what these unknown software files may be, we perform a detailed analysis of their properties. We then explore whether it is possible to extend the labeling of software downloads by building a rule-based system that automatically learns from the available ground truth and can be used to identify many more benign and malicious files with very high confidence. This allows us to greatly expand the number of software files that can be labeled with high confidence, thus providing results that can benefit the evaluation of future malware detection systems.
Babak Rahbarinia, Marco Balduzzi, Roberto Perdisci
DSN3
2017 Enabling Reconstruction of Attacks on Users via Efficient Browsing Snapshots
Phani Vadrevu, Jienan Liu, Bo Li 0058, Babak Rahbarinia, Kyu Hyung Lee, Roberto Perdisci
NDSS6
2017 Still Beheading Hydras: Botnet Takedowns Then and Now
abstract
Devices infected with malicious software typically form botnet armies under the influence of one or more command and control (C&C) servers. The botnet problem reached such levels where federal law enforcement agencies have to step in and take actions against botnets by disrupting (or “taking down”) their C&Cs, and thus their illicit operations. Lately, more and more private companies have started to independently take action against botnet armies, primarily focusing on their DNS-based C&Cs. While well-intentioned, their C&C takedown methodology is in most cases ad-hoc, and limited by the breadth of knowledge available around the malware that facilitates the botnet. With this paper, we aim to bring order, measure, and reason to the botnet takedown problem. We improve an existing takedown analysis system called rza. Specifically, we examine additional botnet takedowns, enhance the risk calculation to use botnet population counts, and include a detailed discussion of policy improvements that can be made to improve takedowns. As part of our system evaluation, we perform a postmortem analysis of the recent 3322.org, Citadel, and No-IP takedowns.
Yacin Nadji, Roberto Perdisci, Manos Antonakakis
IEEE Trans. Dependable Secur. Comput.2
2016 Real-Time Detection of Malware Downloads via Large-Scale URL->File->Machine Graph Mining
abstract
In this paper we propose Mastino, a novel defense system to detect malware download events. A download event is a 3-tuple that identifies the action of downloading a file from a URL that was triggered by a client (machine). Mastino utilizes global situation awareness and continuously monitors various network- and system-level events of the clients' machines across the Internet and provides real time classification of both files and URLs to the clients upon submission of a new, unknown file or URL to the system. To enable detection of the download events, Mastino builds a large download graph that captures the subtle relationships among the entities of download events, i.e. files, URLs, and machines. We implemented a prototype version of Mastino and evaluated it in a large-scale real-world deployment. Our experimental evaluation shows that Mastino can accurately classify malware download events with an average of 95.5% true positive (TP), while incurring less than 0.5% false positives (FP). In addition, we show the Mastino can classify a new download event as either benign or malware in just a fraction of a second, and is therefore suitable as a real time defense system.
Babak Rahbarinia, Marco Balduzzi, Roberto Perdisci
AsiaCCS3
2016 MAXS: Scaling Malware Execution with Sequential Multi-Hypothesis Testing
abstract
In an attempt to coerce useful information about the behavior of new malware families, threat analysts commonly force newly collected malicious software samples to run within a sandboxed environment. The main goal is to gather intelligence that can later be leveraged to detect and enumerate new malware infections within a network. Currently, most analysis environments "blindly" execute each newly collected malware sample for a predetermined amount of time (e.g., four to five minutes). However, a large majority of malware samples that are forced through sandbox execution are simply repackaged versions of previously seen (and already analyzed) malware. Consequently, a significant amount of time may be wasted in analyzing samples that do not generate new intelligence.
Phani Vadrevu, Roberto Perdisci
AsiaCCS2
2016 Towards Measuring and Mitigating Social Engineering Software Download Attacks
Terry Nelms, Roberto Perdisci, Manos Antonakakis, Mustaque Ahamad
USENIX Security Symposium2
2016 Efficient and Accurate Behavior-Based Tracking of Malware-Control Domains in Large ISP Networks
abstract
In this article, we propose Segugio , a novel defense system that allows for efficiently tracking the occurrence of new malware-control domain names in very large ISP networks. Segugio passively monitors the DNS traffic to build a machine-domain bipartite graph representing who is querying what . After labeling nodes in this query behavior graph that are known to be either benign or malware-related, we propose a novel approach to accurately detect previously unknown malware-control domains. We implemented a proof-of-concept version of Segugio and deployed it in large ISP networks that serve millions of users. Our experimental results show that Segugio can track the occurrence of new malware-control domains with up to 94% true positives (TPs) at less than 0.1% false positives (FPs). In addition, we provide the following results: (1) we show that Segugio can also detect control domains related to new, previously unseen malware families, with 85% TPs at 0.1% FPs; (2) Segugio’s detection models learned on traffic from a given ISP network can be deployed into a different ISP network and still achieve very high detection accuracy; (3) new malware-control domains can be detected days or even weeks before they appear in a large commercial domain-name blacklist; (4) Segugio can be used to detect previously unknown malware-infected machines in ISP networks; and (5) we show that Segugio clearly outperforms domain-reputation systems based on Belief Propagation.
Babak Rahbarinia, Roberto Perdisci, Manos Antonakakis
ACM Trans. Priv. Secur.2
2015 WebCapsule: Towards a Lightweight Forensic Engine for Web Browsers
abstract
Performing detailed forensic analysis of real-world web security incidents targeting users, such as social engineering and phishing attacks, is a notoriously challenging and time-consuming task. To reconstruct web-based attacks, forensic analysts typically rely on browser cache files and system logs. However, cache files and logs provide only sparse information often lacking adequate detail to reconstruct a precise view of the incident. To address this problem, we need an always-on and lightweight (i.e., low overhead) forensic data collection system that can be easily integrated with a variety of popular browsers, and that allows for recording enough detailed information to enable a full reconstruction of web security incidents, including phishing attacks.
Christopher Neasbitt, Bo Li 0058, Roberto Perdisci, Long Lu, Kapil Singh, Kang Li 0001
CCS3
2015 Segugio: Efficient Behavior-Based Tracking of Malware-Control Domains in Large ISP Networks
abstract
In this paper, we propose Segugio, a novel defense system that allows for efficiently tracking the occurrence of new malware-control domain names in very large ISP networks. Segugio passively monitors the DNS traffic to build a machine-domain bipartite graph representing who is querying what. After labelling nodes in this query behavior graph that are known to be either benign or malware-related, we propose a novel approach to accurately detect previously unknown malware-control domains. We implemented a proof-of-concept version of Segugio and deployed it in large ISP networks that serve millions of users. Our experimental results show that Segugio can track the occurrence of new malware-control domains with up to 94% true positives (TPs) at less than 0.1% false positives (FPs). In addition, we provide the following results: (1) we show that Segugio can also detect control domains related to new, previously unseen malware families, with 85% TPs at 0.1% FPs, (2) Segugio's detection models learned on traffic from a given ISP network can be deployed into a different ISP network and still achieve very high detection accuracy, (3) new malware-control domains can be detected days or even weeks before they appear in a large commercial domain name blacklist, and (4) we show that Segugio clearly outperforms Notos, a previously proposed domain name reputation system.
Babak Rahbarinia, Roberto Perdisci, Manos Antonakakis
DSN2
2015 ASwatch: An AS Reputation System to Expose Bulletproof Hosting ASes
abstract
Bulletproof hosting Autonomous Systems (ASes)-malicious ASes fully dedicated to supporting cybercrime-provide freedom and resources for a cyber-criminal to operate.Their services include hosting a wide range of illegal content, botnet C&C servers, and other malicious resources.Thousands of new ASes are registered every year, many of which are often used exclusively to facilitate cybercrime.A natural approach to squelching bulletproof hosting ASes is to develop a reputation system that can identify them for takedown by law enforcement and as input to other attack detection systems (e.g., spam filters, botnet detection systems).Unfortunately, current AS reputation systems rely primarily on data-plane monitoring of malicious activity from IP addresses (and thus can only detect malicious ASes after attacks are underway), and are not able to distinguish between malicious and legitimate but abused ASes.As a complement to these systems, in this paper, we explore a fundamentally different approach to establishing AS reputation.We present ASwatch, a system that identifies malicious ASes using exclusively the control-plane (i.e., routing) behavior of ASes.ASwatch's design is based on the intuition that, in an attempt to evade possible detection and remediation efforts, malicious ASes exhibit "agile" control plane behavior (e.g., short-lived routes, aggressive re-wiring).We evaluate our system on known malicious ASes; our results show that ASwatch detects up to 93% of malicious ASes with a 5% false positive rate, which is reasonable to effectively complement existing defense systems. CCS Concepts• Security and privacy →
Maria Konte, Roberto Perdisci, Nick Feamster
SIGCOMM2
2015 WebWitness: Investigating, Categorizing, and Mitigating Malware Download Paths
Terry Nelms, Roberto Perdisci, Manos Antonakakis, Mustaque Ahamad
USENIX Security Symposium2
2015 Understanding Malvertising Through Ad-Injecting Browser Extensions
abstract
Malvertising is a malicious activity that leverages advertising to distribute various forms of malware. Because advertising is the key revenue generator for numerous Internet companies, large ad networks, such as Google, Yahoo and Microsoft, invest a lot of effort to mitigate malicious ads from their ad networks. This drives adversaries to look for alternative methods to deploy malvertising. In this paper, we show that browser extensions that use ads as their monetization strategy often facilitate the deployment of malvertising. Moreover, while some extensions simply serve ads from ad networks that support malvertising, other extensions maliciously alter the content of visited webpages to force users into installing malware. To measure the extent of these behaviors we developed Expector, a system that automatically inspects and identifies browser extensions that inject ads, and then classifies these ads as malicious or benign based on their landing pages. Using Expector, we automatically inspected over 18,000 Chrome browser extensions. We found 292 extensions that inject ads, and detected 56 extensions that participate in malvertising using 16 different ad networks and with a total user base of 602,417.
Xinyu Xing 0001, Wei Meng 0001, Byoungyoung Lee, Udi Weinsberg, Anmol Sheth, Roberto Perdisci, Wenke Lee
WWW6
2014 ClickMiner: Towards Forensic Reconstruction of User-Browser Interactions from Network Traces
abstract
Recent advances in network traffic capturing techniques have made it feasible to record full traffic traces, often for extended periods of time. Among the applications enabled by full traffic captures, being able to automatically reconstruct user-browser interactions from archived web traffic traces would be helpful in a number of scenarios, such as aiding the forensic analysis of network security incidents. Unfortunately, the modern web is becoming increasingly complex, serving highly dynamic pages that make heavy use of scripting languages, a variety of browser plugins, and asynchronous content requests. Consequently, the semantic gap between user-browser interactions and the network traces has grown significantly, making it challenging to analyze the web traffic produced by even a single user.
Christopher Neasbitt, Roberto Perdisci, Kang Li 0001, Terry Nelms
CCS2
2014 DNS Noise: Measuring the Pervasiveness of Disposable Domains in Modern DNS Traffic
abstract
In this paper, we present an analysis of a new class of domain names: disposable domains. We observe that popular web applications, along with other Internet services, systematically use this new class of domain names. Disposable domains are likely generated automatically, characterized by a "one-time use" pattern, and appear to be used as a way of "signaling" via DNS queries. To shed light on the pervasiveness of disposable domains, we study 24 days of live DNS traffic spanning a year observed at a large Internet Service Provider. We find that disposable domains increased from 23.1% to 27.6% of all queried domains, and from 27.6% to 37.2% of all resolved domains observed daily. While this creative use of DNS may enable new applications, it may also have unanticipated negative consequences on the DNS caching infrastructure, DNSSEC validating resolvers, and passive DNS data collection systems.
Yizheng Chen 0001, Manos Antonakakis, Roberto Perdisci, Yacin Nadji, David Dagon, Wenke Lee
DSN3
2014 PeerRush: Mining for unwanted P2P traffic
Babak Rahbarinia, Roberto Perdisci, Andrea Lanzi, Kang Li 0001
J. Inf. Secur. Appl.2
2014 Building a Scalable System for Stealthy P2P-Botnet Detection
abstract
Peer-to-peer (P2P) botnets have recently been adopted by botmasters for their resiliency against take-down efforts. Besides being harder to take down, modern botnets tend to be stealthier in the way they perform malicious activities, making current detection approaches ineffective. In addition, the rapidly growing volume of network traffic calls for high scalability of detection systems. In this paper, we propose a novel scalable botnet detection system capable of detecting stealthy P2P botnets. Our system first identifies all hosts that are likely engaged in P2P communications. It then derives statistical fingerprints to profile P2P traffic and further distinguish between P2P botnet traffic and legitimate P2P traffic. The parallelized computation with bounded complexity makes scalability a built-in feature of our system. Extensive evaluation has demonstrated both high detection accuracy and great scalability of the proposed system.
Junjie Zhang 0004, Roberto Perdisci, Wenke Lee, Xiapu Luo, Unum Sarfraz
IEEE Trans. Inf. Forensics Secur.2
2013 Beheading hydras: performing effective botnet takedowns
abstract
Devices infected with malicious software typically form botnet armies under the influence of one or more command and control (C&C) servers. The botnet problem reached such levels where federal law enforcement agencies have to step in and take actions against botnets by disrupting (or "taking down") their C&Cs, and thus their illicit operations. Lately, more and more private companies have started to independently take action against botnet armies, primarily focusing on their DNS-based C&Cs. While well-intentioned, their C&C takedown methodology is in most cases ad-hoc, and limited by the breadth of knowledge available around the malware that facilitates the botnet.
Yacin Nadji, Manos Antonakakis, Roberto Perdisci, David Dagon, Wenke Lee
CCS3
2013 PeerRush: Mining for Unwanted P2P Traffic
Babak Rahbarinia, Roberto Perdisci, Andrea Lanzi, Kang Li 0001
DIMVA2
2013 Measuring and Detecting Malware Downloads in Live Network Traffic
Phani Vadrevu, Babak Rahbarinia, Roberto Perdisci, Kang Li 0001, Manos Antonakakis
ESORICS3
2013 Connected Colors: Unveiling the Structure of Criminal Networks
Yacin Nadji, Manos Antonakakis, Roberto Perdisci, Wenke Lee
RAID3
2013 ExecScent: Mining for New C&C Domains in Live Networks with Adaptive Control Protocol Templates
Terry Nelms, Roberto Perdisci, Mustaque Ahamad
USENIX Security Symposium2
2013 Scalable fine-grained behavioral clustering of HTTP-based malware
Roberto Perdisci, Davide Ariu, Giorgio Giacinto
Comput. Networks1
2012 VAMO: towards a fully automated malware clustering validity analysis
abstract
Malware clustering is commonly applied by malware analysts to cope with the increasingly growing number of distinct malware variants collected every day from the Internet. While malware clustering systems can be useful for a variety of applications, assessing the quality of their results is intrinsically hard. In fact, clustering can be viewed as an unsupervised learning process over a dataset for which the complete ground truth is usually not available. Previous studies propose to evaluate malware clustering results by leveraging the labels assigned to the malware samples by multiple anti-virus scanners (AVs). However, the methods proposed thus far require a (semi-)manual adjustment and mapping between labels generated by different AVs, and are limited to selecting a reference sub-set of samples for which an agreement regarding their labels can be reached across a majority of AVs. This approach may bias the reference set towards "easy to cluster" malware samples, thus potentially resulting in an overoptimistic estimate of the accuracy of the malware clustering results.
Roberto Perdisci, Man Chon U
ACSAC1
2012 From Throw-Away Traffic to Bots: Detecting the Rise of DGA-Based Malware
Manos Antonakakis, Roberto Perdisci, Yacin Nadji, Nikolaos Vasiloglou, Saeed Abu-Nimeh, Wenke Lee, David Dagon
USENIX Security Symposium2
2012 Early Detection of Malicious Flux Networks via Large-Scale Passive DNS Traffic Analysis
abstract
In this paper, we present FluxBuster, a novel passive DNS traffic analysis system for detecting and tracking malicious flux networks. FluxBuster applies large-scale monitoring of DNS traffic traces generated by recursive DNS (RDNS) servers located in hundreds of different networks scattered across several different geographical locations. Unlike most previous work, our detection approach is not limited to the analysis of suspicious domain names extracted from spam emails or precompiled domain blacklists. Instead, FluxBuster is able to detect malicious flux service networks in-the-wild, i.e., as they are "accessed” by users who fall victim of malicious content, independently of how this malicious content was advertised. We performed a long-term evaluation of our system spanning a period of about five months. The experimental results show that FluxBuster is able to accurately detect malicious flux networks with a low false positive rate. Furthermore, we show that in many cases FluxBuster is able to detect malicious flux domains several days or even weeks before they appear in public domain blacklists.
Roberto Perdisci, Igino Corona, Giorgio Giacinto
IEEE Trans. Dependable Secur. Comput.1
2011 Exposing invisible timing-based traffic watermarks with BACKLIT
abstract
Traffic watermarking is an important element in many network security and privacy applications, such as tracing botnet C&C communications and deanonymizing peer-to-peer VoIP calls. The state-of-the-art traffic watermarking schemes are usually based on packet timing information and they are notoriously difficult to detect. In this paper, we show for the first time that even the most sophisticated timing-based watermarking schemes (e.g., RAINBOW and SWIRL) are not invisible by proposing a new detection system called BACKLIT. BACKLIT is designed according to the observation that any practical timing-based traffic watermark will cause noticeable alterations in the intrinsic timing features typical of TCP flows. We propose five metrics that are sufficient for detecting four state-of-the-art traffic watermarks for bulk transfer and interactive traffic. BACKLIT can be easily deployed in stepping stones and anonymity networks (e.g., Tor), because it does not rely on strong assumptions and can be realized in an active or passive mode. We have conducted extensive experiments to evaluate BACKLIT's detection performance using the PlanetLab platform. The results show that BACKLIT can detect watermarked network flows with high accuracy and few false positives.
Xiapu Luo, Peng Zhou 0002, Junjie Zhang 0004, Roberto Perdisci, Wenke Lee, Rocky K. C. Chang
ACSAC4
2011 Understanding the prevalence and use of alternative plans in malware with network games
abstract
In this paper we describe and evaluate a technique to improve the amount of information gained from dynamic malware analysis systems. By playing network games during analysis, we explore the behavior of malware when it believes its network resources are malfunctioning. This forces the malware to reveal its alternative plan to the analysis system resulting in a more complete understanding of malware behavior. Network games are similar to multipath exploration techniques, but are resistant to conditional code obfuscation. Our experimental results show that network games discover highly useful network information from malware. Of the 161,000 domain names and over three million IP addresses coerced from malware during three weeks, over 95% never appeared on public blacklists. We show that this information is both likely to be malicious and can be used to improve existing domain name and IP address reputation systems, blacklists, and network-based malware clustering systems.
Yacin Nadji, Manos Antonakakis, Roberto Perdisci, Wenke Lee
ACSAC3
2011 SURF: detecting and measuring search poisoning
abstract
Search engine optimization (SEO) techniques are often abused to promote websites among search results. This is a practice known as blackhat SEO. In this paper we tackle a newly emerging and especially aggressive class of blackhat SEO, namely search poisoning. Unlike other blackhat SEO techniques, which typically attempt to promote a website's ranking only under a limited set of search keywords relevant to the website's content, search poisoning techniques disregard any term relevance constraint and are employed to poison popular search keywords with the sole purpose of diverting large numbers of users to short-lived traffic-hungry websites for malicious purposes. To accurately detect search poisoning cases, we designed a novel detection system called SURF. SURF runs as a browser component to extract a number of robust (i.e., difficult to evade) detection features from search-then-visit browsing sessions, and is able to accurately classify malicious search user redirections resulted from user clicking on poisoned search results. Our evaluation on real-world search poisoning instances shows that SURF can achieve a detection rate of 99.1% at a false positive rate of 0.9%. Furthermore, we applied SURF to analyze a large dataset of search-related browsing sessions collected over a period of seven months starting in September 2010. Through this long-term measurement study we were able to reveal new trends and interesting patterns related to a great variety of poisoning cases, thus contributing to a better understanding of the prevalence and gravity of the search poisoning problem.
Long Lu, Roberto Perdisci, Wenke Lee
CCS2
2011 Boosting the scalability of botnet detection using adaptive traffic sampling
abstract
Botnets pose a serious threat to the health of the Internet. Most current network-based botnet detection systems require deep packet inspection (DPI) to detect bots. Because DPI is a computational costly process, such detection systems cannot handle large volumes of traffic typical of large enterprise and ISP networks. In this paper we propose a system that aims to efficiently and effectively identify a small number of suspicious hosts that are likely bots. Their traffic can then be forwarded to DPI-based botnet detection systems for fine-grained inspection and accurate botnet detection. By using a novel adaptive packet sampling algorithm and a scalable spatial-temporal flow correlation approach, our system is able to substantially reduce the volume of network traffic that goes through DPI, thereby boosting the scalability of existing botnet detection systems. We implemented a proof-of-concept version of our system, and evaluated it using real-world legitimate and botnet-related network traces. Our experimental results are very promising and suggest that our approach can enable the deployment of botnet-detection systems in large, high-speed networks.
Junjie Zhang 0004, Xiapu Luo, Roberto Perdisci, Guofei Gu, Wenke Lee, Nick Feamster
AsiaCCS3
2011 Detecting stealthy P2P botnets using statistical traffic fingerprints
abstract
Peer-to-peer (P2P) botnets have recently been adopted by botmasters for their resiliency to take-down efforts. Besides being harder to take down, modern botnets tend to be stealthier in the way they perform malicious activities, making current detection approaches, including, ineffective. In this paper, we propose a novel botnet detection system that is able to identify stealthy P2P botnets, even when malicious activities may not be observable. First, our system identifies all hosts that are likely engaged in P2P communications. Then, we derive statistical fingerprints to profile different types of P2P traffic, and we leverage these fingerprints to distinguish between P2P botnet traffic and other legitimate P2P traffic. Unlike previous work, our system is able to detect stealthy P2P botnets even when the underlying compromised hosts are running legitimate P2P applications (e.g., Skype) and the P2P bot software at the same time. Our experimental evaluation based on real-world data shows that the proposed system can achieve high detection accuracy with a low false positive rate.
Junjie Zhang 0004, Roberto Perdisci, Wenke Lee, Unum Sarfraz, Xiapu Luo
DSN2
2011 HTTPOS: Sealing Information Leaks with Browser-side Obfuscation of Encrypted Flows
Xiapu Luo, Peng Zhou 0002, Edmond W. W. Chan, Wenke Lee, Rocky K. C. Chang, Roberto Perdisci
NDSS6
2011 Detecting Malware Domains at the Upper DNS Hierarchy
Manos Antonakakis, Roberto Perdisci, Wenke Lee, Nikolaos Vasiloglou, David Dagon
USENIX Security Symposium2
2010 On the Secrecy of Spread-Spectrum Flow Watermarks
Xiapu Luo, Junjie Zhang 0004, Roberto Perdisci, Wenke Lee
ESORICS3
2010 Behavioral Clustering of HTTP-Based Malware and Signature Generation Using Malicious Network Traces
Roberto Perdisci, Wenke Lee, Nick Feamster
NSDI1
2010 A Centralized Monitoring Infrastructure for Improving DNS Security
Manos Antonakakis, David Dagon, Xiapu Luo, Roberto Perdisci, Wenke Lee, Justin Bellmor
RAID4
2010 Building a Dynamic Reputation System for DNS
Manos Antonakakis, Roberto Perdisci, David Dagon, Wenke Lee, Nick Feamster
USENIX Security Symposium2
2009 Detecting Malicious Flux Service Networks through Passive Analysis of Recursive DNS Traces
abstract
In this paper we propose a novel, passive approach for detecting and tracking malicious flux service networks. Our detection system is based on passive analysis of recursive DNS (RDNS) traffic traces collected from multiple large networks. Contrary to previous work, our approach is not limited to the analysis of suspicious domain names extracted from spam emails or precompiled domain blacklists. Instead, our approach is able to detect malicious flux service networks in-the-wild, i.e., as they are accessed by users who fall victims of malicious content advertised through blog spam, instant messaging spam, social Website spam, etc., beside email spam. We experiment with the RDNS traffic passively collected at two large ISP networks. Overall, our sensors monitored more than 2.5 billion DNS queries per day from millions of distinct source IPs for a period of 45 days. Our experimental results show that the proposed approach is able to accurately detect malicious flux service networks. Furthermore, we show how our passive detection and tracking of malicious flux service networks may benefit spam filtering applications.
Roberto Perdisci, Igino Corona, David Dagon, Wenke Lee
ACSAC1
2009 WSEC DNS: Protecting recursive DNS resolvers from poisoning attacks
abstract
Recently, a new attack for poisoning the cache of Recursive DNS (RDNS) resolvers was discovered and revealed to the public. In response, major DNS vendors released a patch to their software. However, the released patch does not completely protect DNS servers from cache poisoning attacks in a number of practical scenarios. DNSSEC seems to offer a definitive solution to the vulnerabilities of the DNS protocol, but unfortunately DNSSEC has not yet been widely deployed. In this paper, we proposeWild-card SECure DNS (WSEC DNS), a novel solution to DNS cache poisoning attacks. WSEC DNS relies on existing properties of the DNS protocol and is based on wild-card domain names. We show that WSEC DNS is able to decrease the probability of success of cache poisoning attacks by several orders of magnitude. That is, with WSEC DNS in place, an attacker has to persistently run a cache poisoning attack for years, before having a non-negligible chance of success. Furthermore, WSEC DNS offers complete backward compatibility to DNS servers that may for any reason decide not to implement it, therefore allowing an incremental large-scale deployment. Contrary to DNSSEC, WSEC DNS is deployable immediately because it does not have the technical and political problems that have so far hampered a large-scale deployment of DNSSEC.
Roberto Perdisci, Manos Antonakakis, Xiapu Luo, Wenke Lee
DSN1
2009 McPAD: A multiple classifier system for accurate payload-based anomaly detection
Roberto Perdisci, Davide Ariu, Prahlad Fogla, Giorgio Giacinto, Wenke Lee
Comput. Networks1
2008 McBoost: Boosting Scalability in Malware Collection and Analysis Using Statistical Classification of Executables
abstract
In this work, we propose Malware Collection Booster (McBoost), a fast statistical malware detection tool that is intended to improve the scalability of existing malware collection and analysis approaches. Given a large collection of binaries that may contain both hitherto unknown malware and benign executables, McBoost reduces the overall time of analysis by classifying and filtering out the least suspicious binaries and passing only the most suspicious ones to a detailed binary analysis process for signature extraction.The McBoost framework consists of a classifier specialized in detecting whether an executable is packed or not, a universal unpacker based on dynamic binary analysis, and a classifier specialized in distinguishing between malicious or benign code. We developed a proof-of-concept version of McBoost and evaluated it on 5,586 malware and 2,258 benign programs. McBoost has an accuracy of 87.3%, and an Area Under the ROC curve (AUC) equal to 0.977. Our evaluation also shows that McBoost reduces the overall time of analysis to only a fraction (e.g., 13.4%) of the computation time that would otherwise be required to analyze large sets of mixed malicious and benign executables.
Roberto Perdisci, Andrea Lanzi, Wenke Lee
ACSAC1
2008 BotMiner: Clustering Analysis of Network Traffic for Protocol- and Structure-Independent Botnet Detection
Guofei Gu, Roberto Perdisci, Junjie Zhang 0004, Wenke Lee
USENIX Security Symposium2
2008 Classification of packed executables for accurate computer virus detection
Roberto Perdisci, Andrea Lanzi, Wenke Lee
Pattern Recognit. Lett.1
2006 Using an Ensemble of One-Class SVM Classifiers to Harden Payload-based Anomaly Detection Systems
abstract
Unsupervised or unlabeled learning approaches for network anomaly detection have been recently proposed. In particular, recent work on unlabeled anomaly detection focused on high speed classification based on simple payload statistics. For example, PAYL, an anomaly IDS, measures the occurrence frequency in the payload of n-grams. A simple model of normal traffic is then constructed according to this description of the packets' content. It has been demonstrated that anomaly detectors based on payload statistics can be "evaded" by mimicry attacks using byte substitution and padding techniques. In this paper we propose a new approach to construct high speed payload-based anomaly IDS intended to be accurate and hard to evade. We propose a new technique to extract the features from the payload. We use a feature clustering algorithm originally proposed for text classification problems to reduce the dimensionality of the feature space. Accuracy and hardness of evasion are obtained by constructing our anomaly-based IDS using an ensemble of one-class SVM classifiers that work on different feature spaces.
Roberto Perdisci, Guofei Gu, Wenke Lee
ICDM1
2006 MisleadingWorm Signature Generators Using Deliberate Noise Injection
abstract
Several syntactic-based automatic worm signature generators, e.g., Polygraph, have recently been proposed. These systems typically assume that a set of suspicious flows are provided by a flow classifier, e.g., a honeynet or an intrusion detection system, that often introduces "noise" due to difficulties and imprecision inflow classification. The algorithms for extracting the worm signatures from the flow data are designed to cope with the noise. It has been reported that these systems can handle a fairly high noise level, e.g., 80% for Polygraph. In this paper, we show that if noise is introduced deliberately to mislead a worm signature generator, a much lower noise level, e.g., 50%, can already prevent the system from reliably generating useful worm signatures. Using Polygraph as a case study, we describe a new and general class of attacks whereby a worm can combine polymorphism and misleading behavior to intentionally pollute the dataset of suspicious flows during its propagation and successfully mislead the automatic signature generation process. This study suggests that unless an accurate and robust flow classification process is in place, automatic syntactic-based signature generators are vulnerable to such noise injection attacks.
Roberto Perdisci, David Dagon, Wenke Lee, Prahlad Fogla, Monirul Islam Sharif
S&P1
2006 Polymorphic Blending Attacks
Prahlad Fogla, Monirul Islam Sharif, Roberto Perdisci, Oleg M. Kolesnikov, Wenke Lee
USENIX Security Symposium3
2006 Alarm clustering for intrusion detection systems in computer networks
Roberto Perdisci, Giorgio Giacinto, Fabio Roli
Eng. Appl. Artif. Intell.1