EDBT 2026 Demo / reviewers in the wild / expert
Vern Paxson
dblp:p/VernPaxson
· DBLP profile ↗
143ranked-venue papers
17as first author
3since 2021 · last 2023
0009-0005-2673-543XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 79 · 5 first-author · 2 since 2021Computer networks · 54 · 10 first-author · 1 since 2021Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GGFAST: Automating Generation of Flexible Network Traffic ClassifiersabstractWhen employing supervised machine learning to analyze network traffic, the heart of the task often lies in developing effective features for the ML to leverage. We develop GGFAST, a unified, automated framework that can build powerful classifiers for specific network traffic analysis tasks, built on interpretable features. The framework uses only packet sizes, directionality, and sequencing, facilitating analysis in a payload-agnostic fashion that remains applicable in the presence of encryption. Julien Piet, Dubem Nwoji, Vern Paxson |
SIGCOMM | 3 |
| 2023 | Network Detection of Interactive SSH Impostors Using Deep Learning
Julien Piet, Aashish Sharma, Vern Paxson, David A. Wagner 0001 |
USENIX Security Symposium | 3 |
| 2021 | Hopper: Modeling and Detecting Lateral Movement
Grant Ho, Mayank Dhiman, Devdatta Akhawe, Vern Paxson, Stefan Savage, Geoffrey M. Voelker, David A. Wagner 0001 |
USENIX Security Symposium | 4 |
| 2020 | Composition Kills: A Case Study of Email Sender Authentication
Jianjun Chen 0005, Vern Paxson, Jian Jiang 0002 |
USENIX Security Symposium | 2 |
| 2019 | Detecting and Characterizing Lateral Phishing at Scale
Grant Ho, Asaf Cidon, Lior Gavish, Marco Schweighauser, Vern Paxson, Stefan Savage, Geoffrey M. Voelker, David A. Wagner 0001 |
USENIX Security Symposium | 5 |
| 2018 | We Still Don't Have Secure Cross-Domain Requests: an Empirical Study of CORS
Jianjun Chen 0005, Jian Jiang 0002, Hai-Xin Duan, Tao Wan 0004, Shuo Chen 0001, Vern Paxson, Min Yang 0002 |
USENIX Security Symposium | 6 |
| 2017 | A Large-Scale Empirical Study of Security PatchesabstractGiven how the "patching treadmill" plays a central role for enabling sites to counter emergent security concerns, it behooves the security community to understand the patch development process and characteristics of the resulting fixes. Illumination of the nature of security patch development can inform us of shortcomings in existing remediation processes and provide insights for improving current practices. In this work we conduct a large-scale empirical study of security patches, investigating more than 4,000 bug fixes for over 3,000 vulnerabilities that affected a diverse set of 682 open-source software projects. For our analysis we draw upon the National Vulnerability Database, information scraped from relevant external references, affected software repositories, and their associated security fixes. Leveraging this diverse set of information, we conduct an analysis of various aspects of the patch development life cycle, including investigation into the duration of impact a vulnerability has on a code base, the timeliness of patch development, and the degree to which developers produce safe and reliable fixes. We then characterize the nature of security fixes in comparison to other non-security bug fixes, exploring the complexity of different types of patches and their impact on code bases. Frank Li 0001, Vern Paxson |
CCS | 2 |
| 2017 | Data Breaches, Phishing, or Malware?: Understanding the Risks of Stolen CredentialsabstractIn this paper, we present the first longitudinal measurement study of the underground ecosystem fueling credential theft and assess the risk it poses to millions of users. Over the course of March, 2016--March, 2017, we identify 788,000 potential victims of off-the-shelf keyloggers; 12.4 million potential victims of phishing kits; and 1.9 billion usernames and passwords exposed via data breaches and traded on blackmarket forums. Using this dataset, we explore to what degree the stolen passwords---which originate from thousands of online services---enable an attacker to obtain a victim's valid email credentials---and thus complete control of their online identity due to transitive trust. Drawing upon Google as a case study, we find 7--25% of exposed passwords match a victim's Google account. For these accounts, we show how hardening authentication mechanisms to include additional risk signals such as a user's historical geolocations and device profiles helps to mitigate the risk of hijacking. Beyond these risk metrics, we delve into the global reach of the miscreants involved in credential theft and the blackhat tools they rely on. We observe a remarkable lack of external pressure on bad actors, with phishing kit playbooks and keylogger capabilities remaining largely unchanged since the mid-2000s. Kurt Thomas, Frank Li 0001, Ali Zand, Jacob Barrett, Juri Ranieri, Luca Invernizzi, Yarik Markov, Oxana Comanescu, Vijay Eranti, Angelique Moscicki, Dan Margolis, Vern Paxson, Elie Bursztein |
CCS | 12 |
| 2017 | Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain AdaptationabstractGreg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca Portnoff, Sadia Afroz, Damon McCoy, Kirill Levchenko, Vern Paxson. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017. Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca S. Portnoff, Sadia Afroz 0001, Damon McCoy, Kirill Levchenko, Vern Paxson |
EMNLP | 8 |
| 2017 | Packetlab: a universal measurement endpoint interfaceabstractThe right vantage point is critical to the success of any active measurement. However, most research groups cannot afford to design, deploy, and maintain their own network of measurement endpoints, and thus rely measurement infrastructure shared by others. Unfortunately, the mechanism by which we share access to measurement endpoints today is not frictionless; indeed, issues of compatibility, trust, and a lack of incentives get in the way of efficiently sharing measurement infrastructure. Kirill Levchenko, Amogh Dhamdhere, Bradley Huffaker, K. C. Claffy, Mark Allman, Vern Paxson |
Internet Measurement Conference | 6 |
| 2017 | Target generation for internet-wide IPv6 scanningabstractFast IPv4 scanning has enabled researchers to answer a wealth of new security and measurement questions. However, while increased network speeds and computational power have enabled comprehensive scans of the IPv4 address space, a brute-force approach does not scale to IPv6. Systems are limited to scanning a small fraction of the IPv6 address space and require an algorithmic approach to determine a small set of candidate addresses to probe. In this paper, we first explore the considerations that guide designing such algorithms. We introduce a new approach that identifies dense address space regions from a set of known "seed" addresses and generates a set of candidates to scan. We compare our algorithm 6Gen against Entropy/IP---the current state of the art---finding that we can recover between 1--8 times as many addresses for the five candidate datasets considered in the prior work. However, during our analysis, we uncover widespread IP aliasing in IPv6 networks. We discuss its effect on target generation and explore preliminary approaches for detecting aliased regions. Austin Murdock, Frank Li 0001, Paul Bramsen, Zakir Durumeric, Vern Paxson |
Internet Measurement Conference | 5 |
| 2017 | The Security Impact of HTTPS Interception
Zakir Durumeric, Zane Ma, Drew Springall, Richard Barnes 0001, Nick Sullivan, Elie Bursztein, Michael D. Bailey, J. Alex Halderman, Vern Paxson |
NDSS | 9 |
| 2017 | Augur: Internet-Wide Detection of Connectivity DisruptionsabstractAnecdotes, news reports, and policy briefings collectively suggest that Internet censorship practices are pervasive. The scale and diversity of Internet censorship practices makes it difficult to precisely monitor where, when, and how censorship occurs, as well as what is censored. The potential risks in performing the measurements make this problem even more challenging. As a result, many accounts of censorship begin-and end-with anecdotes or short-term studies from only a handful of vantage points. We seek to instead continuously monitor information about Internet reachability, to capture the onset or termination of censorship across regions and ISPs. To achieve this goal, we introduce Augur, a method and accompanying system that utilizes TCP/IP side channels to measure reachability between two Internet locations without directly controlling a measurement vantage point at either location. Using these side channels, coupled with techniques to ensure safety by not implicating individual users, we develop scalable, statistically robust methods to infer network-layer filtering, and implement a corresponding system capable of performing continuous monitoring of global censorship. We validate our measurements of Internet-wide disruption in nearly 180 countries over 17 days against sites known to be frequently blocked, we also identify the countries where connectivity disruption is most prevalent. Paul Pearce, Roya Ensafi, Frank Li 0001, Nick Feamster, Vern Paxson |
IEEE Symposium on Security and Privacy | 5 |
| 2017 | Detecting Credential Spearphishing in Enterprise Settings
Grant Ho, Aashish Sharma, Mobin Javed, Vern Paxson, David A. Wagner 0001 |
USENIX Security Symposium | 4 |
| 2017 | Global Measurement of DNS Manipulation
Paul Pearce, Ben Jones, Frank Li 0001, Roya Ensafi, Nick Feamster, Nicholas Weaver, Vern Paxson |
USENIX Security Symposium | 7 |
| 2017 | Characterizing the Nature and Dynamics of Tor Exit Blocking
Rachee Singh, Rishab Nithyanand, Sadia Afroz 0001, Paul Pearce, Michael Carl Tschantz, Phillipa Gill, Vern Paxson |
USENIX Security Symposium | 7 |
| 2017 | Tools for Automated Analysis of Cybercriminal MarketsabstractUnderground forums are widely used by criminals to buy and sell a host of stolen items, datasets, resources, and criminal services. These forums contain important resources for understanding cybercrime. However, the number of forums, their size, and the domain expertise required to understand the markets makes manual exploration of these forums unscalable. In this work, we propose an automated, top-down approach for analyzing underground forums. Our approach uses natural language processing and machine learning to automatically generate high-level information about underground forums, first identifying posts related to transactions, and then extracting products and prices. We also demonstrate, via a pair of case studies, how an analyst can use these automated approaches to investigate other categories of products and transactions. We use eight distinct forums to assess our tools: Antichat, Blackhat World, Carders, Darkode, Hack Forums, Hell, L33tCrew and Nulled. Our automated approach is fast and accurate, achieving over 80% accuracy in detecting post category, product, and prices. Rebecca S. Portnoff, Sadia Afroz 0001, Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Damon McCoy, Kirill Levchenko, Vern Paxson |
WWW | 8 |
| 2017 | Social Engineering Attacks on Government Opponents: Target PerspectivesabstractAbstract New methods of dissident surveillance employed by repressive nation-states increasingly involve socially engineering targets into unwitting cooperation (e.g., by convincing them to open a malicious attachment or link). While a fair amount is understood about the nature of these threat actors and the types of tools they use, there is comparatively little understood about targets’ perceptions of the risks associated with their online activity, and their security posture. We conducted in-depth interviews of 30 potential targets of Middle Eastern and Horn of Africa-based governments, also examining settings and software on their computers and phones. Our engagement illuminates the ways that likely targets are vulnerable to the types of social engineering employed by nation-states. William R. Marczak, Vern Paxson |
Proc. Priv. Enhancing Technol. | 2 |
| 2016 | Host of Troubles: Multiple Host Ambiguities in HTTP ImplementationsabstractThe Host header is a security-critical component in an HTTP request, as it is used as the basis for enforcing security and caching policies. While the current specification is generally clear on how host-related protocol fields should be parsed and interpreted, we find that the implementations are problematic. We tested a variety of widely deployed HTTP implementations and discover a wide range of non-compliant and inconsistent host processing behaviours. The particular problem is that when facing a carefully crafted HTTP request with ambiguous host fields (e.g., with multiple Host headers), two different HTTP implementations often accept and understand it differently when operating on the same request in sequence. We show a number of techniques to induce inconsistent interpretations of host between HTTP implementations and how the inconsistency leads to severe attacks such as HTTP cache poisoning and security policy bypass. The prevalence of the problem highlights the potential negative impact of gaps between the specifications and implementations of Internet protocols. Jianjun Chen 0005, Jian Jiang 0002, Hai-Xin Duan, Nicholas Weaver, Tao Wan 0004, Vern Paxson |
CCS | 6 |
| 2016 | PREDATOR: Proactive Recognition and Elimination of Domain Abuse at Time-Of-RegistrationabstractMiscreants register thousands of new domains every day to launch Internet-scale attacks, such as spam, phishing, and drive-by downloads. Quickly and accurately determining a domain's reputation (association with malicious activity) provides a powerful tool for mitigating threats and protecting users. Yet, existing domain reputation systems work by observing domain use (e.g., lookup patterns, content hosted) often too late to prevent miscreants from reaping benefits of the attacks that they launch. As a complement to these systems, we explore the extent to which features evident at domain registration indicate a domain's subsequent use for malicious activity. We develop PREDATOR, an approach that uses only time-of-registration features to establish domain reputation. We base its design on the intuition that miscreants need to obtain many domains to ensure profitability and attack agility, leading to abnormal registration behaviors (e.g., burst registrations, textually similar names). We evaluate PREDATOR using registration logs of second-level .com and .net domains over five months. PREDATOR achieves a 70% detection rate with a false positive rate of 0.35%, thus making it an effective and early first line of defense against the misuse of DNS domains. It predicts malicious domains when they are registered, which is typically days or weeks earlier than existing DNS blacklists. Shuang Hao 0001, Alex Kantchelian, Brad Miller 0002, Vern Paxson, Nick Feamster |
CCS | 4 |
| 2016 | An Analysis of the Privacy and Security Risks of Android VPN Permission-enabled Apps
Muhammad Ikram 0001, Narseo Vallina-Rodriguez, Suranga Seneviratne, Mohamed Ali Kâafar, Vern Paxson |
Internet Measurement Conference | 5 |
| 2016 | A Multi-perspective Analysis of Carrier-Grade NAT Deployment
Philipp Richter, Florian Wohlfart, Narseo Vallina-Rodriguez, Mark Allman, Randy Bush, Anja Feldmann, Christian Kreibich, Nicholas Weaver, Vern Paxson |
Internet Measurement Conference | 9 |
| 2016 | Forwarding-Loop Attacks in Content Delivery Networks
Jianjun Chen 0005, Hai-Xin Duan, Jinjin Liang, Jian Jiang 0002, Kang Li 0001, Tao Wan 0004, Vern Paxson |
NDSS | 8 |
| 2016 | Do You See What I See? Differential Treatment of Anonymous Users
Sheharbano Khattak, David Fifield, Sadia Afroz 0001, Mobin Javed, Srikanth Sundaresan, Damon McCoy, Vern Paxson, Steven J. Murdoch |
NDSS | 7 |
| 2016 | VAST: A Unified Platform for Interactive Network Forensics
Matthias Vallentin, Vern Paxson, Robin Sommer |
NSDI | 2 |
| 2016 | Detecting DNS Root Manipulation
Ben Jones, Nick Feamster, Vern Paxson, Nicholas Weaver, Mark Allman |
PAM | 3 |
| 2016 | SoK: Towards Grounding Censorship Circumvention in EmpiricismabstractEffective evaluations of approaches to circumventing government Internet censorship require incorporating perspectives of how censors operate in practice. We undertake an extensive examination of real censors by surveying prior measurement studies and analyzing field reports and bug tickets from practitioners. We assess both deployed circumvention approaches and research proposals to consider the criteria employed in their evaluations and compare these to the observed behaviors of real censors, identifying areas where evaluations could more faithfully and effectively incorporate the practices of modern censors. These observations lead to an agenda realigning research with the predominant problems of today. Michael Carl Tschantz, Sadia Afroz 0001, Vern Paxson |
IEEE Symposium on Security and Privacy | 4 |
| 2016 | You've Got Vulnerability: Exploring Effective Vulnerability Notifications
Frank Li 0001, Zakir Durumeric, Jakub Czyz, Mohammad Karami, Michael D. Bailey, Damon McCoy, Stefan Savage, Vern Paxson |
USENIX Security Symposium | 8 |
| 2016 | Remedying Web Hijacking: Notification Effectiveness and Webmaster ComprehensionabstractAs miscreants routinely hijack thousands of vulnerable web servers weekly for cheap hosting and traffic acquisition, security services have turned to notifications both to alert webmasters of ongoing incidents as well as to expedite recovery. In this work we present the first large-scale measurement study on the effectiveness of combinations of browser, search, and direct webmaster notifications at reducing the duration a site remains compromised. Our study captures the life cycle of 760,935 hijacking incidents from July, 2014--June, 2015, as identified by Google Safe Browsing and Search Quality. We observe that direct communication with webmasters increases the likelihood of cleanup by over 50% and reduces infection lengths by at least 62%. Absent this open channel for communication, we find browser interstitials---while intended to alert visitors to potentially harmful content---correlate with faster remediation. As part of our study, we also explore whether webmasters exhibit the necessary technical expertise to address hijacking incidents. Based on appeal logs where webmasters alert Google that their site is no longer compromised, we find 80% of operators successfully clean up symptoms on their first appeal. However, a sizeable fraction of site owners do not address the root cause of compromise, with over 12% of sites falling victim to a new attack within 30 days. We distill these findings into a set of recommendations for improving web security and best practices for webmasters. Frank Li 0001, Grant Ho, Eric Kuan, Yuan Niu, Lucas Ballard, Kurt Thomas, Elie Bursztein, Vern Paxson |
WWW | 8 |
| 2016 | Towards Mining Latent Client Identifiers from Network TrafficabstractAbstract Websites extensively track users via identifiers that uniquely map to client machines or user accounts. Although such tracking has desirable properties like enabling personalization and website analytics, it also raises serious concerns about online user privacy, and can potentially enable illicit surveillance by adversaries who broadly monitor network traffic. In this work we seek to understand the possibilities of latent identifiers appearing in user traffic in forms beyond those already well-known and studied, such as browser and Flash cookies. We develop a methodology for processing large network traces to semi-automatically discover identifiers sent by clients that distinguish users/devices/browsers, such as usernames, cookies, custom user agents, and IMEI numbers. We address the challenges of scaling such discovery up to enterprise-sized data by devising multistage filtering and streaming algorithms. The resulting methodology reflects trade-offs between reducing the ultimate analysis burden and the risk of missing potential identifier strings. We analyze 15 days of data from a site with several hundred users and capture dozens of latent identifiers, primarily in HTTP request components, but also in non-HTTP protocols. Sakshi Jain, Mobin Javed, Vern Paxson |
Proc. Priv. Enhancing Technol. | 3 |
| 2015 | Examining How the Great Firewall Discovers Hidden Circumvention ServersabstractRecently, the operators of the national censorship infrastructure of China began to employ "active probing" to detect and block the use of privacy tools. This probing works by passively monitoring the network for suspicious traffic, then actively probing the corresponding servers, and blocking any that are determined to run circumvention servers such as Tor. Roya Ensafi, David Fifield, Philipp Winter, Nick Feamster, Nicholas Weaver, Vern Paxson |
Internet Measurement Conference | 6 |
| 2015 | Measurement and Analysis of Traffic Exchange ServicesabstractTraffic exchange services enable members to bring traffic to their websites from a diverse pool of IP addresses, in return for visiting sites of other members. We examine the world of traffic exchanges to characterize their makeup, usage, and monetization. We find that the ecosystem includes a range of services, from manual exchanges where participants must solve CAPTCHAs between successive page views, to exchanges that provide tools that automatically surf without requiring any user action. By "milking" a sample of these exchanges, we analyze month-long datasets to examine the nature of URLs that members submit to them. We find a wide prevalence of URLs for services that pay users in return for views to their content, and at least 30% of the requested impressions are for pages that clearly participate in a class of impression fraud called referrer spoofing. We also analyze the size and composition of a sample of these exchange networks by making purchases, finding that the exchanges delivered visits from roughly 200K unique IP~addresses, and that in some exchange networks, the majority of visits came from cloud hosting services. Mobin Javed, Cormac Herley, Marcus Peinado, Vern Paxson |
Internet Measurement Conference | 4 |
| 2015 | Beyond the Radio: Illuminating the Higher Layers of Mobile NetworksabstractCellular network performance is often viewed as primarily dominated by the radio technology. However, reality proves more complex: mobile operators deploy and configure their networks in different ways, and sometimes establish network sharing agreements with other mobile carriers. Moreover, regulators have encouraged newer operational models such as Mobile Virtual Network Operators (MVNOs) to promote competition. In this paper we draw upon data collected by the ICSI Netalyzr app for Android to characterize how operational decisions, such as network configurations, business models, and relationships between operators introduce diversity in service quality and affect user security and privacy. We delve in detail beyond the radio link and into network configuration and business relationships in six countries. We identify the widespread use of transparent middleboxes such as HTTP and DNS proxies, analyzing how they actively modify user traffic, compromise user privacy, and potentially undermine user security. In addition, we identify network sharing agreements between operators, highlighting the implications of roaming and characterizing the properties of MVNOs, including that a majority are simply rebranded versions of major operators. More broadly, our findings highlight the importance of considering higher-layer relationships when seeking to analyze mobile traffic in a sound fashion. Narseo Vallina-Rodriguez, Srikanth Sundaresan, Christian Kreibich, Nicholas Weaver, Vern Paxson |
MobiSys | 5 |
| 2015 | Temporal Lensing and Its Application in Pulsing Denial-of-Service AttacksabstractWe introduce "temporal lensing": a technique that concentrates a relatively low-bandwidth flood into a short, high-bandwidth pulse. By leveraging existing DNS infrastructure, we experimentally explore lensing and the properties of the pulses it creates. We also empirically show how attackers can use lensing alone to achieve peak bandwidths more than an order of magnitude greater than their upload bandwidth. While formidable by itself in a pulsing DoS attack, attackers can also combine lensing with amplification to potentially produce pulses with peak bandwidths orders of magnitude larger than their own. Ryan Rasti, Mukul Murthy, Nicholas Weaver, Vern Paxson |
IEEE Symposium on Security and Privacy | 4 |
| 2015 | Ad Injection at Scale: Assessing Deceptive Advertisement ModificationsabstractToday, web injection manifests in many forms, but fundamentally occurs when malicious and unwanted actors tamper directly with browser sessions for their own profit. In this work we illuminate the scope and negative impact of one of these forms, ad injection, in which users have ads imposed on them in addition to, or different from, those that websites originally sent them. We develop a multi-staged pipeline that identifies ad injection in the wild and captures its distribution and revenue chains. We find that ad injection has entrenched itself as a cross-browser monetization platform impacting more than 5% of unique daily IP addresses accessing Google -- tens of millions of users around the globe. Injected ads arrive on a client's machine through multiple vectors: our measurements identify 50,870 Chrome extensions and 34,407 Windows binaries, 38% and 17% of which are explicitly malicious. A small number of software developers support the vast majority of these injectors who in turn syndicate from the larger ad ecosystem. We have contacted the Chrome Web Store and the advertisers targeted by ad injectors to alert each of the deceptive practices involved. Kurt Thomas, Elie Bursztein, Chris Grier, Grant Ho, Nav Jagpal, Alexandros Kapravelos, Damon McCoy, Antonio Nappa, Vern Paxson, Paul Pearce, Niels Provos, Moheeb Abu Rajab |
IEEE Symposium on Security and Privacy | 9 |
| 2015 | Blocking-resistant communication through domain frontingabstractAbstract We describe “domain fronting,” a versatile censorship circumvention technique that hides the remote endpoint of a communication. Domain fronting works at the application layer, using HTTPS, to communicate with a forbidden host while appearing to communicate with some other host, permitted by the censor. The key idea is the use of different domain names at different layers of communication. One domain appears on the “outside” of an HTTPS request—in the DNS request and TLS Server Name Indication—while another domain appears on the “inside”—in the HTTP Host header, invisible to the censor under HTTPS encryption. A censor, unable to distinguish fronted and nonfronted traffic to a domain, must choose between allowing circumvention traffic and blocking the domain entirely, which results in expensive collateral damage. Domain fronting is easy to deploy and use and does not require special cooperation by network intermediaries. We identify a number of hard-to-block web services, such as content delivery networks, that support domain-fronted connections and are useful for censorship circumvention. Domain fronting, in various forms, is now a circumvention workhorse. We describe several months of deployment experience in the Tor, Lantern, and Psiphon circumvention systems, whose domain-fronting transports now connect thousands of users daily and transfer many terabytes per month. David Fifield, Chang Lan, Rod Hynes, Percy Wegmann, Vern Paxson |
Proc. Priv. Enhancing Technol. | 5 |
| 2014 | Characterizing Large-Scale Click Fraud in ZeroAccessabstractClick fraud is a scam that hits a criminal sweet spot by both tapping into the vast wealth of online advertising and exploiting that ecosystem's complex structure to obfuscate the flow of money to its perpetrators. In this work, we illuminate the intricate nature of this activity through the lens of ZeroAccess--one of the largest click fraud botnets in operation. Using a broad range of data sources, including peer-to-peer measurements, command-and-control telemetry, and contemporaneous click data from one of the top ad networks, we construct a view into the scale and complexity of modern click fraud operations. By leveraging the dynamics associated with Microsoft's attempted takedown of ZeroAccess in December 2013, we employ this coordinated view to identify "ad units" whose traffic (and hence revenue) primarily derived from ZeroAccess. While it proves highly challenging to extrapolate from our direct observations to a truly global view, by anchoring our analysis in the data for these ad units we estimate that the botnet's fraudulent activities plausibly induced advertising losses on the order of $100,000 per day. Paul Pearce, Vacha Dave, Chris Grier, Kirill Levchenko, Saikat Guha 0002, Damon McCoy, Vern Paxson, Stefan Savage, Geoffrey M. Voelker |
CCS | 7 |
| 2014 | Consequences of Connectivity: Characterizing Account Hijacking on TwitterabstractIn this study we expose the serious large-scale threat of criminal account hijacking and the resulting damage incurred by users and web services. We develop a system for detecting large-scale attacks on Twitter that identifies 14 million victims of compromise. We examine these accounts to track how attacks spread within social networks and to determine how criminals ultimately realize a profit from hijacked credentials. We find that compromise is a systemic threat, with victims spanning nascent, casual, and core users. Even brief compromises correlate with 21% of victims never returning to Twitter after the service wrests control of a victim's account from criminals. Infections are dominated by social contagions---phishing and malware campaigns that spread along the social graph. These contagions mirror information diffusion and biological diseases, growing in virulence with the number of neighboring infections. Based on the severity of our findings, we argue that early outbreak detection that stems the spread of compromise in 24 hours can spare 70% of victims. Kurt Thomas, Frank Li 0001, Chris Grier, Vern Paxson |
CCS | 4 |
| 2014 | A Tangled Mass: The Android Root Certificate StoresabstractThe security of today's Web rests in part on the set of X.509 certificate authorities trusted by each user's browser. Users generally do not themselves configure their browser's root store but instead rely upon decisions made by the suppliers of either the browsers or the devices upon which they run. In this work we explore the nature and implications of these trust decisions for Android users. Drawing upon datasets collected by Netalyzr for Android and ICSI's Certificate Notary, we characterize the certificate root store population present in mobile devices in the wild. Motivated by concerns that bloated root stores increase the attack surface of mobile users, we report on the interplay of certificate sets deployed by the device manufacturers, mobile operators, and the Android OS. We identify certificates installed exclusively by apps on rooted devices, thus breaking the audited and supervised root store model, and also discover use of TLS interception via HTTPS proxies employed by a market research company. Narseo Vallina-Rodriguez, Johanna Amann, Christian Kreibich, Nicholas Weaver, Vern Paxson |
CoNEXT | 5 |
| 2014 | The Matter of HeartbleedabstractThe Heartbleed vulnerability took the Internet by surprise in April 2014. The vulnerability, one of the most consequential since the advent of the commercial Internet, allowed attackers to remotely read protected memory from an estimated 24--55% of popular HTTPS sites. In this work, we perform a comprehensive, measurement-based analysis of the vulnerability's impact, including (1) tracking the vulnerable population, (2) monitoring patching behavior over time, (3) assessing the impact on the HTTPS certificate ecosystem, and (4) exposing real attacks that attempted to exploit the bug. Furthermore, we conduct a large-scale vulnerability notification experiment involving 150,000 hosts and observe a nearly 50% increase in patching by notified hosts. Drawing upon these analyses, we discuss what went well and what went poorly, in an effort to understand how the technical community can respond more effectively to such events in the future. Zakir Durumeric, James Kasten, David Adrian, J. Alex Halderman, Michael D. Bailey, Frank Li 0001, Nicholas Weaver, Johanna Amann, Jethro G. Beekman, Mathias Payer, Vern Paxson |
Internet Measurement Conference | 11 |
| 2014 | A Look at the Consequences of Internet Censorship Through an ISP LensabstractInternet censorship artificially changes the dynamics of resource production and consumption, affecting a range of stakeholders that include end users, service providers, and content providers. We analyze two large-scale censorship events in Pakistan: blocking of pornographic content in 2011 and of YouTube in 2012. Using traffic datasets collected at home and SOHO networks before and after the censorship events, we: a) quantify the demand for blocked content, b) illuminate challenges encountered by service providers in implementing the censorship policies, c) investigate changes in user behavior (e.g., with respect to circumvention) after censorship, and d) assess benefits extracted by competing content providers of blocked content. Sheharbano Khattak, Mobin Javed, Syed Ali Khayam, Zartash Afzal Uzmi, Vern Paxson |
Internet Measurement Conference | 5 |
| 2014 | HILTI: an Abstract Execution Environment for Deep, Stateful Network Traffic AnalysisabstractWhen developing networking systems such as firewalls, routers, and intrusion detection systems, one faces a striking gap between the ease with which one can often describe a desired analysis in high-level terms, and the tremendous amount of low-level implementation details that one must still grapple with to come to a robust solution. We present HILTI, a platform that bridges this divide by providing to application developers much of the low-level functionality, without tying it to a specific analysis structure. HILTI consists of two parts: (1) an abstract machine model that we tailor specifically to the networking domain, directly supporting the field's common abstractions and idioms in its instruction set; and (2) a compilation strategy for turning programs written for the abstract machine into optimized, natively executable code. We have developed a prototype of the HILTI compiler toolchain that fully implements the design's functionality, and ported exemplars of networking applications to the HILTI model to demonstrate the aptness of its abstractions. Our evaluation of HILTI's functionality and performance confirms its potential to become a powerful platform for future application development. Robin Sommer, Matthias Vallentin, Lorenzo De Carli, Vern Paxson |
Internet Measurement Conference | 4 |
| 2014 | Here Be Web Proxies
Nicholas Weaver, Christian Kreibich, Martin Dam, Vern Paxson |
PAM | 4 |
| 2014 | Native actors: how to scale network forensicsabstractWhen an organization detects a security breach, it undertakes a forensic analysis to figure out what happened. This investigation involves inspecting a wide range of heterogeneous data sources spanning over a long period of time. The iterative nature of the analysis procedure requires an interactive experience with the data. However, the distributed processing paradigms we find in practice today fail to provide this requirement: the batch-oriented nature of MapReduce cannot deliver sub-second round-trip times, and distributed in-memory processing cannot store the terabytes of activity logs needed to inspect during an incident. Matthias Vallentin, Dominik Charousset, Thomas C. Schmidt, Vern Paxson, Matthias Wählisch |
SIGCOMM | 4 |
| 2014 | Hulk: Eliciting Malicious Behavior in Browser Extensions
Alexandros Kapravelos, Chris Grier, Neha Chachra, Christopher Krügel, Giovanni Vigna, Vern Paxson |
USENIX Security Symposium | 6 |
| 2014 | When Governments Hack Opponents: A Look at Actors and Technology
William R. Marczak, John Scott-Railton, Morgan Marquis-Boire, Vern Paxson |
USENIX Security Symposium | 4 |
| 2013 | Detecting stealthy, distributed SSH brute-forcingabstractIn this work we propose a general approach for detecting distributed malicious activity in which individual attack sources each operate in a stealthy, low-profile manner. We base our approach on observing statistically significant changes in a parameter that summarizes aggregate activity, bracketing a distributed attack in time, and then determining which sources present during that interval appear to have coordinated their activity. We apply this approach to the problem of detecting stealthy distributed SSH bruteforcing activity, showing that we can model the process of legitimate users failing to authenticate using a beta-binomial distribution, which enables us to tune a detector that trades off an expected level of false positives versus time-to-detection. Using the detector we study the prevalence of distributed bruteforcing, finding dozens of instances in an extensive 8-year dataset collected from a site with several thousand SSH users. Many of the attacks---some of which last months---would be quite difficult to detect individually. While a number of the attacks reflect indiscriminant global probing, we also find attacks that targeted only the local site, as well as occasional attacks that succeeded. Mobin Javed, Vern Paxson |
CCS | 2 |
| 2013 | Understanding the domain registration behavior of spammersabstractSpammers register a tremendous number of domains to evade blacklisting and takedown efforts. Current techniques to detect such domains rely on crawling spam URLs or monitoring lookup traffic. Such detection techniques are only effective after the spammers have already launched their campaigns, and thus these countermeasures may only come into play after the spammer has already reaped significant benefits from the dissemination of large volumes of spam. In this paper we examine the registration process of such domains, with a particular eye towards features that might indicate that a given domain likely has a malicious purpose at registration time, before it is ever used for an attack. Our assessment includes exploring the characteristics of registrars, domain life cycles, registration bursts, and naming patterns. By investigating zone changes from the .com TLD over a 5-month period, we discover that spammers employ bulk registration, that they often re-use domains previously registered by others, and that they tend to register and host their domains over a small set of registrars. Our findings suggest steps that registries or registrars could use to frustrate the efforts of miscreants to acquire domains in bulk, ultimately reducing their agility for mounting large-scale attacks. Shuang Hao 0001, Matthew Thomas, Vern Paxson, Nick Feamster, Christian Kreibich, Chris Grier, Scott Hollenbeck |
Internet Measurement Conference | 3 |
| 2013 | Practical Comprehensive Bounds on Surreptitious Communication over DNS
Vern Paxson, Mihai Christodorescu, Mobin Javed, Josyula R. Rao, Reiner Sailer, Douglas Lee Schales, Marc Ph. Stoecklin, Kurt Thomas, Wietse Z. Venema, Nicholas Weaver |
USENIX Security Symposium | 1 |
| 2013 | Trafficking Fraudulent Accounts: The Role of the Underground Market in Twitter Spam and Abuse
Kurt Thomas, Damon McCoy, Chris Grier, Alek Kolcz, Vern Paxson |
USENIX Security Symposium | 5 |
| 2012 | Manufacturing compromise: the emergence of exploit-as-a-serviceabstractWe investigate the emergence of the exploit-as-a-service model for driveby browser compromise. In this regime, attackers pay for an exploit kit or service to do the "dirty work" of exploiting a victim's browser, decoupling the complexities of browser and plugin vulnerabilities from the challenges of generating traffic to a website under the attacker's control. Upon a successful exploit, these kits load and execute a binary provided by the attacker, effectively transferring control of a victim's machine to the attacker. Chris Grier, Lucas Ballard, Juan Caballero, Neha Chachra, Christian Dietrich 0005, Kirill Levchenko, Panayiotis Mavrommatis, Damon McCoy, Antonio Nappa, Andreas Pitsillidis, Niels Provos, M. Zubair Rafique, Moheeb Abu Rajab, Christian Rossow, Kurt Thomas, Vern Paxson, Stefan Savage, Geoffrey M. Voelker |
CCS | 16 |
| 2012 | Fathom: a browser-based network measurement platformabstractFor analyzing network performance issues, there can be great utility in having the capability to measure directly from the perspective of end systems. Because end systems do not provide any external programming interface to measurement functionality, obtaining this capability today generally requires installing a custom executable on the system, which can prove prohibitively expensive. In this work we leverage the ubiquity of web browsers to demonstrate the possibilities of browsers themselves offering such a programmable environment. We present Fathom, a Firefox extension that implements a number of measurement primitives that enable websites or other parties to program network measurements using JavaScript. Fathom is lightweight, imposing < 3.2% overhead in page load times for popular web pages, and often provides 1 ms timestamp accuracy. We demonstrate Fathom's utility with three case studies: providing a JavaScript version of the Netalyzr network characterization tool, debugging web access failures, and enabling web sites to diagnose performance problems of their clients. Mohan Dhawan, Justin Samuel, Renata Teixeira, Christian Kreibich, Mark Allman, Nicholas Weaver, Vern Paxson |
Internet Measurement Conference | 7 |
| 2012 | The BIZ Top-Level Domain: Ten Years Later
Tristan Halvorson, Janos Szurdi, Gregor Maier, Márk Félegyházi, Christian Kreibich, Nicholas Weaver, Kirill Levchenko, Vern Paxson |
PAM | 8 |
| 2012 | Prudent Practices for Designing Malware Experiments: Status Quo and OutlookabstractMalware researchers rely on the observation of malicious code in execution to collect datasets for a wide array of experiments, including generation of detection models, study of longitudinal behavior, and validation of prior research. For such research to reflect prudent science, the work needs to address a number of concerns relating to the correct and representative use of the datasets, presentation of methodology in a fashion sufficiently transparent to enable reproducibility, and due consideration of the need not to harm others. In this paper we study the methodological rigor and prudence in 36 academic publications from 2006 -- 2011 that rely on malware execution. 40% of these papers appeared in the 6 highest-ranked academic security conferences. We find frequent shortcomings, including problematic assumptions regarding the use of execution-driven datasets (25% of the papers), absence of description of security precautions taken during experiments (71% of the articles), and oftentimes insufficient description of the experimental setup. Deficiencies occur in top-tier venues and elsewhere alike, highlighting a need for the community to improve its handling of malware datasets. In the hope of aiding authors, reviewers, and readers, we frame guidelines regarding transparency, realism, correctness, and safety for collecting and using malware datasets. Christian Rossow, Christian Dietrich 0005, Chris Grier, Christian Kreibich, Vern Paxson, Norbert Pohlmann, Herbert Bos, Maarten van Steen |
IEEE Symposium on Security and Privacy | 5 |
| 2012 | Cloud Terminal: Secure Access to Sensitive Applications from Untrusted Systems
Lorenzo Martignoni, Pongsin Poosankam, Matei Zaharia, Jun Han 0001, Stephen McCamant, Dawn Song, Vern Paxson, Adrian Perrig, Scott Shenker, Ion Stoica |
USENIX ATC | 7 |
| 2011 | An Assessment of Overt Malicious Activity Manifest in Residential Networks
Gregor Maier, Anja Feldmann, Vern Paxson, Robin Sommer, Matthias Vallentin |
DIMVA | 3 |
| 2011 | What's Clicking What? Techniques and Innovations of Today's Clickbots
Brad Miller 0002, Paul Pearce, Chris Grier, Christian Kreibich, Vern Paxson |
DIMVA | 5 |
| 2011 | GQ: practical containment for measuring modern malware systemsabstractMeasurement and analysis of modern malware systems such as botnets relies crucially on execution of specimens in a setting that enables them to communicate with other systems across the Internet. Ethical, legal, and technical constraints however demand containment of resulting network activity in order to prevent the malware from harming others while still ensuring that it exhibits its inherent behavior. Current best practices in this space are sorely lacking: measurement researchers often treat containment superficially, sometimes ignoring it altogether. In this paper we present GQ, a malware execution "farm" that uses explicit containment primitives to enable analysts to develop containment policies naturally, iteratively, and safely. We discuss GQ's architecture and implementation, our methodology for developing containment policies, and our experiences gathered from six years of development and operation of the system. Christian Kreibich, Nicholas Weaver, Chris Kanich, Weidong Cui, Vern Paxson |
Internet Measurement Conference | 5 |
| 2011 | Suspended accounts in retrospect: an analysis of twitter spamabstractIn this study, we examine the abuse of online social networks at the hands of spammers through the lens of the tools, techniques, and support infrastructure they rely upon. To perform our analysis, we identify over 1.1 million accounts suspended by Twitter for disruptive activities over the course of seven months. In the process, we collect a dataset of 1.8 billion tweets, 80 million of which belong to spam accounts. We use our dataset to characterize the behavior and lifetime of spam accounts, the campaigns they execute, and the wide-spread abuse of legitimate web services such as URL shorteners and free web hosting. We also identify an emerging marketplace of illegitimate programs operated by spammers that include Twitter account sellers, ad-based URL shorteners, and spam affiliate programs that help enable underground market diversification. Kurt Thomas, Chris Grier, Dawn Song, Vern Paxson |
Internet Measurement Conference | 4 |
| 2011 | Detecting and Analyzing Automated Activity on Twitter
Chao Michael Zhang, Vern Paxson |
PAM | 2 |
| 2011 | Click Trajectories: End-to-End Analysis of the Spam Value ChainabstractSpam-based advertising is a business. While it has engendered both widespread antipathy and a multi-billion dollar anti-spam industry, it continues to exist because it fuels a profitable enterprise. We lack, however, a solid understanding of this enterprise's full structure, and thus most anti-Spam interventions focus on only one facet of the overall spam value chain (e.g., spam filtering, URL blacklisting, site takedown).In this paper we present a holistic analysis that quantifies the full set of resources employed to monetize spam email -- including naming, hosting, payment and fulfillment -- using extensive measurements of three months of diverse spam data, broad crawling of naming and hosting infrastructures, and over 100 purchases from spam-advertised sites. We relate these resources to the organizations who administer them and then use this data to characterize the relative prospects for defensive interventions at each link in the spam value chain. In particular, we provide the first strong evidence of payment bottlenecks in the spam value chain, 95% of spam-advertised pharmaceutical, replica and software products are monetized using merchant services from just a handful of banks. Kirill Levchenko, Andreas Pitsillidis, Neha Chachra, Brandon Enright, Márk Félegyházi, Chris Grier, Tristan Halvorson, Chris Kanich, Christian Kreibich, Damon McCoy, Nicholas Weaver, Vern Paxson, Geoffrey M. Voelker, Stefan Savage |
IEEE Symposium on Security and Privacy | 13 |
| 2011 | Design and Evaluation of a Real-Time URL Spam Filtering ServiceabstractOn the heels of the widespread adoption of web services such as social networks and URL shorteners, scams, phishing, and malware have become regular threats. Despite extensive research, email-based spam filtering techniques generally fall short for protecting other web services. To better address this need, we present Monarch, a real-time system that crawls URLs as they are submitted to web services and determines whether the URLs direct to spam. We evaluate the viability of Monarch and the fundamental challenges that arise due to the diversity of web service spam. We show that Monarch can provide accurate, real-time protection, but that the underlying characteristics of spam do not generalize across web services. In particular, we find that spam targeting email qualitatively differs in significant ways from spam campaigns targeting Twitter. We explore the distinctions between email and Twitter spam, including the abuse of public web hosting and redirector services. Finally, we demonstrate Monarch's scalability, showing our system could protect a service such as Twitter -- which needs to process 15 million URLs/day -- for a bit under $800/day. Kurt Thomas, Chris Grier, Justin Ma, Vern Paxson, Dawn Song |
IEEE Symposium on Security and Privacy | 4 |
| 2011 | Measuring Pay-per-Install: The Commoditization of Malware Distribution
Juan Caballero, Chris Grier, Christian Kreibich, Vern Paxson |
USENIX Security Symposium | 4 |
| 2011 | Show Me the Money: Characterizing Spam-advertised Revenue
Chris Kanich, Nicholas Weaver, Damon McCoy, Tristan Halvorson, Christian Kreibich, Kirill Levchenko, Vern Paxson, Geoffrey M. Voelker, Stefan Savage |
USENIX Security Symposium | 7 |
| 2011 | Towards Situational Awareness of Large-Scale Botnet Probing EventsabstractBotnets dominate today's attack landscape. In this work, we investigate ways to analyze collections of malicious probing traffic in order to understand the significance of large-scale “botnet probes.” In such events, an entire collection of remote hosts together probes the address space monitored by a sensor in some sort of coordinated fashion. Our goal is to develop methodologies by which sites receiving such probes can infer-using purely local observation-information about the probing activity: What scanning strategies does the probing employ? Is this an attack that specifically targets the site, or is the site only incidentally probed as part of a larger, indiscriminant attack? Our analysis draws upon extensive honeynet data to explore the prevalence of different types of scanning, including properties, such as trend, uniformity, coordination, and darknet avoidance. In addition, we design schemes to extrapolate the global properties of scanning events (e.g., total population and target scope) as inferred from the limited local view of a honeynet. Cross-validating with data from DShield shows that our inferences exhibit promising accuracy. Zhichun Li, Anup Goyal, Yan Chen 0004, Vern Paxson |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2010 | @spam: the underground on 140 characters or lessabstractIn this work we present a characterization of spam on Twitter. We find that 8% of 25 million URLs posted to the site point to phishing, malware, and scams listed on popular blacklists. We analyze the accounts that send spam and find evidence that it originates from previously legitimate accounts that have been compromised and are now being puppeteered by spammers. Using clickthrough data, we analyze spammers' use of features unique to Twitter and the degree that they affect the success of spam. We find that Twitter is a highly successful platform for coercing users to visit spam pages, with a clickthrough rate of 0.13%, compared to much lower rates previously reported for email spam. We group spam URLs into campaigns and identify trends that uniquely distinguish phishing, malware, and spam, to gain an insight into the underlying techniques used to attract users. Chris Grier, Kurt Thomas, Vern Paxson, Chao Michael Zhang |
CCS | 3 |
| 2010 | Using strongly typed networking to architect for tussleabstractToday's networks discriminate towards or against traffic for a wide range of reasons, and in response end users and their applications increasingly attempt to evade monitoring and control, resulting in an ongoing tussle whose roots run deep. In this work we explore an architectural paradigm that can accommodate such tussles in a systematic and transparent fashion. The key idea at the core of our design is strongly typed networking: the notion that application messages contain type information that fully describes the content being transferred. Our framework allows for transparency between parties which then leads to dialog and choice for both users and service providers. While in the early stages, we provide a possible framework for directly addressing the tussle between end users and "the network" without resorting to an ever-increasing degree of obfuscation and inference. Chitra Muthukrishnan, Vern Paxson, Mark Allman, Aditya Akella |
HotNets | 2 |
| 2010 | Netalyzr: illuminating the edge networkabstractIn this paper we present Netalyzr, a network measurement and debugging service that evaluates the functionality provided by people's Internet connectivity. The design aims to prove both comprehensive in terms of the properties we measure and easy to employ and understand for users with little technical background. We structure Netalyzr as a signed Java applet (which users access via their Web browser) that communicates with a suite of measurement-specific servers. Traffic between the two then probes for a diverse set of network properties, including outbound port filtering, hidden in-network HTTP caches, DNS manipulations, NAT behavior, path MTU issues, IPv6 support, and access-modem buffer capacity. In addition to reporting results to the user, Netalyzr also forms the foundation for an extensive measurement of edge-network properties. To this end, along with describing Netalyzr 's architecture and system implementation, we present a detailed study of 130,000 measurement sessions that the service has recorded since we made it publicly available in June 2009. Christian Kreibich, Nicholas Weaver, Boris Nechaev, Vern Paxson |
Internet Measurement Conference | 4 |
| 2010 | Botnet Judo: Fighting Spam with Itself
Andreas Pitsillidis, Kirill Levchenko, Christian Kreibich, Chris Kanich, Geoffrey M. Voelker, Vern Paxson, Nicholas Weaver, Stefan Savage |
NDSS | 6 |
| 2010 | A Longitudinal View of HTTP Traffic
Tom Callahan, Mark Allman, Vern Paxson |
PAM | 3 |
| 2010 | Outside the Closed World: On Using Machine Learning for Network Intrusion DetectionabstractIn network intrusion detection research, one popular strategy for finding attacks is monitoring a network's activity for anomalies: deviations from profiles of normality previously learned from benign traffic, typically identified using tools borrowed from the machine learning community. However, despite extensive academic research one finds a striking gap in terms of actual deployments of such systems: compared with other intrusion detection approaches, machine learning is rarely employed in operational "real world" settings. We examine the differences between the network intrusion detection problem and other areas where machine learning regularly finds much more success. Our main claim is that the task of finding attacks is fundamentally different from these other applications, making it significantly harder for the intrusion detection community to employ machine learning effectively. We support this claim by identifying challenges particular to network intrusion detection, and provide a set of guidelines meant to strengthen future research on anomaly detection. Robin Sommer, Vern Paxson |
IEEE Symposium on Security and Privacy | 2 |
| 2009 | Automating analysis of large-scale botnet probing eventsabstractBotnets dominate today's attack landscape. In this work we investigate ways to analyze collections of malicious probing traffic in order to understand the significance of large-scale "botnet probes". In such events, an entire collection of remote hosts together probes the address space monitored by a sensor in some sort of coordinated fashion. Our goal is to develop methodologies by which sites receiving such probes can infer---using purely local observation---information about the probing activity: What scanning strategies does the probing employ? Is this an attack that specifically targets the site, or is the site only incidentally probed as part of a larger, indiscriminant attack? Zhichun Li, Anup Goyal, Yan Chen 0004, Vern Paxson |
AsiaCCS | 4 |
| 2009 | Securing Mediated Trace Access Using Black-box Permutation Analysis
Prateek Mittal, Vern Paxson, Robin Sommer, Mark Winterrowd |
HotNets | 2 |
| 2009 | On dominant characteristics of residential broadband internet trafficabstractWhile residential broadband Internet access is popular in many parts of the world, only a few studies have examined the characteristics of such traffic. In this paper we describe observations from monitoring the network activity for more than 20,000 residential DSL customers in an urban area. To ensure privacy, all data is immediately anonymized. We augment the anonymized packet traces with information about DSL-level sessions, IP (re-)assignments, and DSL link bandwidth. Gregor Maier, Anja Feldmann, Vern Paxson, Mark Allman |
Internet Measurement Conference | 3 |
| 2009 | On calibrating enterprise switch measurementsabstractThe complexity of modern enterprise networks is ever-increasing, and our understanding of these important networks is not keeping pace. Our insight into intra-subnet traffic (staying within a single LAN) is particularly limited, due to the widespread use of Ethernet switches that preclude ready LAN-wide monitoring. We have recently undertaken an approach to obtaining extensive intra-subnet visibility based on tapping sets of Ethernet switch ports simultaneously. However, doing so leads to a number of measurement calibration issues that require careful consideration to address. First, one must correctly account for redundant copies of packets that appear due to switch flooding, which if not accurately identified can greatly skew subsequent analysis results. We show that a simple, natural rule one might use for doing so in fact introduces systematic errors, but an altered version of the rule performs significantly better. We then employ this revised rule to aid with calibration issues concerning the fidelity of packet timestamps and the amount of measurement loss that our collection apparatus incurred. Additionally, we develop techniques to "map" the monitored network in terms of identifying key topological components, such as subnet boundaries, which hosts were directly monitored, and the presence of "hidden" switches and hubs. Finally, we present initial analyses demonstrating that the magnitude and diversity of traffic at the subnet level is in fact striking, highlighting the importance of obtaining and correctly calibrating switch-level enterprise traces. Boris Nechaev, Vern Paxson, Mark Allman, Andrei V. Gurtov |
Internet Measurement Conference | 2 |
| 2009 | Detecting Forged TCP Reset Packets
Nicholas Weaver, Robin Sommer, Vern Paxson |
NDSS | 3 |
| 2009 | An architecture for exploiting multi-core processors to parallelize network intrusion preventionabstractAbstract It is becoming increasingly difficult to implement effective systems for preventing network attacks, due to the combination of the rising sophistication of attacks requiring more complex analyses to detect; the relentless growth in the volume of network traffic that we must analyze; and, critically, the failure in recent years for uniprocessor performance to sustain the exponential gains that for so many years CPUs have enjoyed. For commodity hardware, tomorrow's performance gains will instead come frommulti‐corearchitectures in which a whole set of CPUs executes concurrently. Taking advantage of the full power of multi‐core processors for network intrusion prevention requires an in‐depth approach. In this work we frame an architecture customized for parallel execution of network attack analysis. At the lowest layer of the architecture is an ‘Active Network Interface’, a custom device based on an inexpensive FPGA platform. The analysis itself is structured as an event‐based system, which allows us to find many opportunities for concurrent execution, since events introduce a natural asynchrony into the analysis while still maintaining good cache locality. A preliminary evaluation demonstrates the potential of this architecture. Copyright © 2009 John Wiley & Sons, Ltd. Robin Sommer, Vern Paxson, Nicholas Weaver |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Spamalytics: an empirical analysis of spam marketing conversionabstractThe of spam--the probability that an unsolicited e-mail will ultimately elicit a sale--underlies the entire spam value proposition. However, our understanding of this critical behavior is quite limited, and the literature lacks any quantitative study concerning its true value. In this paper we present a methodology for measuring the conversion rate of spam. Using a parasitic infiltration of an existing botnet's infrastructure, we analyze two spam campaigns: one designed to propagate a malware Trojan, the other marketing on-line pharmaceuticals. For nearly a half billion spam e-mails we identify the number that are successfully delivered, the number that pass through popular anti-spam filters, the number that elicit user visits to the advertised sites, and the number of sales and infections produced. Chris Kanich, Christian Kreibich, Kirill Levchenko, Brandon Enright, Geoffrey M. Voelker, Vern Paxson, Stefan Savage |
CCS | 6 |
| 2008 | A Tool for Offline and Live Testing of Evasion Resilience in Network Intrusion Detection Systems
Leo Juan, Christian Kreibich, Chih-Hung Lin, Vern Paxson |
DIMVA | 4 |
| 2008 | A Reactive Measurement Framework
Mark Allman, Vern Paxson |
PAM | 2 |
| 2008 | Predicting the Resource Consumption of Network Intrusion Detection Systems
Holger Dreger, Anja Feldmann, Vern Paxson, Robin Sommer |
RAID | 3 |
| 2008 | Enriching network security analysis with time travelabstractIn many situations it can be enormously helpful to archive the raw contents of a network traffic stream to disk, to enable later inspection of activity that becomes interesting only in retrospect. We present a Time Machine (TM) for network traffic that provides such a capability. The TM leverages the heavy-tailed nature of network flows to capture nearly all of the likely-interesting traffic while storing only a small fraction of the total volume. An initial proof-of-principle prototype established the forensic value of such an approach, contributing to the investigation of numerous attacks at a site with thousands of users. Based on these experiences, a rearchitected implementation of the system provides flexible, highperformance traffic stream capture, indexing and retrieval, including an interface between the TM and a real-time network intrusion detection system (NIDS). The NIDS controls the TM by dynamically adjusting recording parameters, instructing it to permanently store suspicious activity for offline forensics, and fetching traffic from the past for retrospective analysis. We present a detailed performance evaluation of both stand-alone and joint setups, and report on experiences with running the system live in high-volume environments. Gregor Maier, Robin Sommer, Holger Dreger, Anja Feldmann, Vern Paxson, Fabian Schneider 0001 |
SIGCOMM | 5 |
| 2008 | Predicting the resource consumption of network intrusion detection systemsabstractWhen installing network intrusion detection systems (NIDSs), operators are faced with a large number of parameters and analysis options for tuning trade-offs between detection accuracy versus resource requirements. In this work we set out to assist this process by understanding and predicting the CPU and memory consumption of such systems. Holger Dreger, Anja Feldmann, Vern Paxson, Robin Sommer |
SIGMETRICS | 3 |
| 2008 | Efficient and Robust TCP Stream NormalizationabstractNetwork intrusion detection and prevention systems are vulnerable to evasion by attackers who craft ambiguous traffic to breach the defense of such systems. A normalizer is an inline network element that thwarts evasion attempts by removing ambiguities in network traffic. A particularly challenging step in normalization is the sound detection of inconsistent TCP retransmissions, wherein an attacker sends TCP segments with different payloads for the same sequence number space to present a network monitor with ambiguous analysis. Normalizers that buffer all unacknowledged data to verify the consistency of subsequent retransmissions consume inordinate amounts of memory on highspeed links. On the other hand, normalizers that buffer only the hashes of unacknowledged segments cannot verify the consistency of 20-30% of retransmissions that, according to our traces, do not align with the original transmissions. This paper presents the design of RoboNorm, a normalizer that buffers only the hashes of unacknowledged segments, and yet can detect all inconsistent retransmissions in any TCP byte stream. RoboNorm consumes 1-2 orders of magnitude less memory than normalizers that buffers all unacknowledged data, and is amenable to a high-speed implementation. RoboNorm is also robust to attacks that attempt to compromise its operation or exhaust its resources. Mythili Vutukuru, Hari Balakrishnan, Vern Paxson |
SP | 3 |
| 2008 | Principles for Developing Comprehensive Network Visibility
Mark Allman, Christian Kreibich, Vern Paxson, Robin Sommer, Nicholas Weaver |
HotSec | 3 |
| 2007 | An inquiry into the nature and causes of the wealth of internet miscreantsabstractArticle An inquiry into the nature and causes of the wealth of internet miscreants Share on CCS '07: Proceedings of the 14th ACM conference on Computer and communications securityOctober 2007 Pages 375–388https://doi.org/10.1145/1315245.1315292Online:28 October 2007Publication History 66citation1,671DownloadsMetricsTotal Citations66Total Downloads1,671Last 12 Months59Last 6 weeks14 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jason Franklin, Adrian Perrig, Vern Paxson, Stefan Savage |
CCS | 3 |
| 2007 | Shunting: a hardware/software architecture for flexible, high-performance network intrusion preventionabstractStateful, in-depth, inline traffic analysis for intrusion detection and prevention is growing increasingly more difficult as the data rates of modern networks rise. Yet it remains the case that in many environments, much of the traffic comprising a high-volume stream can, after some initial analysis, be qualified as of "likely uninteresting." We present a combined hardware/software architecture, Shunting, that provides a lightweight mechanism for an intrusion prevention system (IPS) to take advantage of the "heavy-tailed" nature of network traffic to offload work from software to hardware. José M. González, Vern Paxson, Nicholas Weaver |
CCS | 2 |
| 2007 | On the Adaptive Real-Time Detection of Fast-Propagating Network Worms
Jaeyeon Jung, Rodolfo A. Milito, Vern Paxson |
DIMVA | 3 |
| 2007 | The shunt: an FPGA-based accelerator for network intrusion preventionabstractThe sophistication and complexity of analysis performed by today's network intrusion prevention systems (IPSs) benefits greatly from implementation using general-purpose CPUs. Yet the performance of such CPUs increasingly lags behind that necessary to process today's high-rate traffic streams. A key observation, however, is that much of the traffic comprising a high-volume stream can, after some initial analysis, be qualified as "likely uninteresting." To this end, we have developed an in-line, FPGA-based IPS ac-celerator, the Shunt, using the NetFPGA2 platform. The Shunt functions as the forwarding device used by the IPS; it alone processes the bulk of the traffic, offloading the memory bus and leaving the CPU free to inspect the subset of the traffic deemed germane for security analysis. To do so, the Shunt maintains several large state tables indexed by packet header fields, including IP/TCP flags, source and destination IP addresses, and connection tuples. The tables yield decision values the element makes on a packet-by-packet basis: forward the packet, drop it, or divert it through the IPS. By manipulating table entries, the IPS can specify the traffic it wishes to examine, directly block malicious traffic, and "cut through" traffic streams once it has had an opportunity to "vet" them, all on a fine-grained basis. We base our design on a novel series of caches, with a "fail safe" miss policy, coupled to a host PC to handle both cache management and higher level IPS analysis. The design requires only 2 MB of SRAM for its extensive caches, and can sup-port four Gbps Ethernets on a single Virtex 2 Pro 30. Nicholas Weaver, Vern Paxson, José M. González |
FPGA | 2 |
| 2007 | Enabling an Energy-Efficient Future Internet Through Selectively Connected End Systems
Mark Allman, Kenneth J. Christensen, Bruce Nordman, Vern Paxson |
HotNets | 4 |
| 2007 | Issues and etiquette concerning use of shared measurement dataabstractIn this note we discuss issues surrounding how to provide and use network measurement data made available for sharing among researchers. While previous work has focused on the technical details of enabling sharing via traffic anonymization, we focus on higher-level aspects of the process such as potential harm to the provider (e.g., by de-anonymizing a shared dataset) or interactions to strengthen subsequent research (e.g., helping to establish ground truth). We believe the community would benefit from a dialog regarding expectations and responsibilities of data providers, and the etiquette involved with using others' measurement data. To this end, we provide a set of guidelines that aim to aid the process of sharing measurement data. We present these not as specific rules, but rather a framework under which providers and users can better attain a mutual understanding about how to treat particular datasets. Mark Allman, Vern Paxson |
Internet Measurement Conference | 2 |
| 2007 | A brief history of scanningabstractIncessant scanning of hosts by attackers looking for vulnerable servers has become a fact of Internet life. In this paper we present an initial study of the scanning activity observed at one site over the past 12.5 years. We study the onset of scanning in the late 1990s and its evolution in terms of characteristics such as the number of scanners, targets and probing patterns. While our study is preliminary in many ways, it provides the first longitudinal examination of a now ubiquitous Internet phenomenon. Mark Allman, Vern Paxson, Jeff Terrell |
Internet Measurement Conference | 2 |
| 2007 | The NIDS Cluster: Scalable, Stateful Network Intrusion Detection on Commodity Hardware
Matthias Vallentin, Robin Sommer, Jason Lee 0001, Craig Leres, Vern Paxson, Brian Tierney |
RAID | 5 |
| 2007 | The Strengths of Weaker Identities: Opportunistic Personas
Mark Allman, Christian Kreibich, Vern Paxson, Robin Sommer, Nicholas Weaver |
HotSec | 3 |
| 2006 | Fighting Coordinated Attackers with Cross-Organizational Information Sharing
Mark Allman, Ethan Blanton, Vern Paxson, Scott Shenker |
HotNets | 3 |
| 2006 | Network System Challenges in Selective Sharing and Verification for Personal, Social, and Urban-Scale Sensing Applications
Andrew Parker 0001, Sasank Reddy, Thomas Schmid 0002, Kevin K. Chang, Saurabh Ganeriwal, Mani Srivastava 0001, Mark H. Hansen, Jeff Burke, Deborah Estrin, Mark Allman, Vern Paxson |
HotNets | 11 |
| 2006 | Semi-automated discovery of application session structureabstractWhile the problem of analyzing network traffic at the granularity of individual connections has seen considerable previous work and tool development, understanding traffic at a higher level - the structure of user-initiated sessions comprised of groups of related connections - remains much less explored. Some types of session structure, such as the coupling between an FTP control connection and the data connections it spawns, have prespecified forms, though the specifications do not guarantee how the forms appear in practice. Other types of sessions, such as a user reading email with a browser, only manifest empirically. Still other sessions might exist without us even knowing of their presence, such as a botnet zombie receiving instructions from its master and proceeding in turn to carry them out. We present algorithms rooted in the statistics of Poisson processes that can mine a large corpus of network connection logs to extract the apparent structure of application sessions embedded in the connections. Our methods are semi-automated in that we aim to present an analyst with high-quality information (expressed as regular expressions) reflecting different possible abstractions of an application's session structure. We develop and test our methods using traces from a large Internet site, finding diversity in the number of applications that manifest, their different session structures, and the presence of abnormal behavior. Our work has applications to traffic characterization and monitoring, source models for synthesizing network traffic, and anomaly detection. Jayanthkumar Kannan, Jaeyeon Jung, Vern Paxson, Can Emre Koksal |
Internet Measurement Conference | 3 |
| 2006 | binpac: a yacc for writing application protocol parsersabstractA key step in the semantic analysis of network traffic is to parse the traffic stream according to the high-level protocols it contains. This process transforms raw bytes into structured, typed, and semantically meaningful data fields that provide a high-level representation of the traffic. However, constructing protocol parsers by hand is a tedious and error-prone affair due to the complexity and sheer number of application protocols.This paper presents binpac, a declarative language and compiler designed to simplify the task of constructing robust and efficient semantic analyzers for complex network protocols. We discuss the design of the binpac language and a range of issues in generating efficient parsers from high-level specifications. We have used binpac to build several protocol parsers for the "Bro" network intrusion detection system, replacing some of its existing analyzers (handcrafted in C++), and supplementing its operation with analyzers for new protocols. We can then use Bro's powerful scripting language to express application-level analysis of network traffic in high-level terms that are both concise and expressive. binpac is now part of the open-source Bro distribution. Ruoming Pang, Vern Paxson, Robin Sommer, Larry L. Peterson |
Internet Measurement Conference | 2 |
| 2006 | Protocol-Independent Adaptive Replay of Application Dialog
Weidong Cui, Vern Paxson, Nicholas Weaver, Randy H. Katz |
NDSS | 2 |
| 2006 | Enhancing Network Intrusion Detection with Integrated Sampling and Filtering
José M. González, Vern Paxson |
RAID | 2 |
| 2006 | Rethinking Hardware Support for Network Analysis and Intrusion Prevention
Vern Paxson, Krste Asanovic, Sarang Dharmapurikar, John W. Lockwood, Ruoming Pang, Robin Sommer, Nicholas Weaver |
HotSec | 1 |
| 2006 | Network loss tomography using striped unicast probes
Nick G. Duffield, Francesco Lo Presti, Vern Paxson, Don Towsley |
IEEE/ACM Trans. Netw. | 3 |
| 2006 | Observed structure of addresses in IP traffic
Eddie Kohler, Jinyang Li 0001, Vern Paxson, Scott Shenker |
IEEE/ACM Trans. Netw. | 3 |
| 2005 | Exploiting Independent State For Network Intrusion DetectionabstractNetwork intrusion detection systems (NIDSs) critically rely on processing a great deal of state. Often much of this state resides solely in the volatile processor memory accessible to a single user-level process on a single machine. In this work, we highlight the power of independent state, i.e., internal fine-grained state that can be propagated from one instance of a NIDS to others running either concurrently or subsequently. Independent state provides us with a wealth of possible applications that hold promise for enhancing the capabilities of NIDSs. We discuss an implementation of independent state for the Bro NIDS and examine how we can then leverage independent state for distributed processing, load parallelization, selective preservation of state across restarts and crashes, dynamic reconfiguration, high level policy maintenance, and support for profiling and debugging. We have experimented with each of these applications in several large environments and are now working to integrate them into the sites' operational monitoring. A performance evaluation shows that our implementation is suitable for use even in large scale environments Robin Sommer, Vern Paxson |
ACSAC | 2 |
| 2005 | Enhancing the Accuracy of Network-Based Intrusion Detection with Host-Based Context
Holger Dreger, Christian Kreibich, Vern Paxson, Robin Sommer |
DIMVA | 3 |
| 2005 | Building a Time Machine for Efficient Recording and Retrieval of High-Volume Network Traffic
Stefan Kornexl, Vern Paxson, Holger Dreger, Anja Feldmann, Robin Sommer |
Internet Measurement Conference | 2 |
| 2005 | Exploiting Underlying Structure for Detailed Reconstruction of an Internet-scale Event
Abhishek Kumar 0003, Vern Paxson, Nicholas Weaver |
Internet Measurement Conference | 2 |
| 2005 | A First Look at Modern Enterprise Traffic
Ruoming Pang, Mark Allman, Mike Bennett, Jason Lee 0001, Vern Paxson, Brian Tierney |
Internet Measurement Conference | 5 |
| 2005 | Robust TCP Stream Reassembly in the Presence of Adversaries
Sarang Dharmapurikar, Vern Paxson |
USENIX Security Symposium | 2 |
| 2005 | Guest Editor's Introduction: 2005 IEEE Symposium on Security and PrivacyabstractSINCE 1980, the IEEE Symposium on Security and Privacy has been the premier annual forum for the presentation of scientific developments in information security and privacy technology, and for bringing together researchers and practitioners in the field. It is sponsored by the IEEE Computer Society Technical Committee on Security and Privacy, in co-operation with The International Association for Cryptologic Research (IACR). The program committee of the 2005 conference received 192 submissions, and selected 17 papers to be presented, on the basis of excellence of scientific contribution. Out of these 17 high quality papers, the program committee selected three as the most highly rated papers for this special issue. In no particular order, they are: “Hardware-Assisted Circumvention of Self-Hashing Software Tamper Resistance” by P.C. van Oorschot, Anil Somayaji, and Glenn Wurster; “Remote Physical Device Fingerprinting” by Tadayoshi Kohno, Andre Broido, and K.C. Claffy; “Relating Symbolic and Cryptographic Secrecy” by Michael Backes and Birgit Pfitzmann. Like all scientific conferences, the IEEE Symposium on Security and Privacy lives from the voluntary and hard work of many people. We wish to thank all of them-authors, reviewers, participants and organizers-but in particular the members of the program committee: William Arbaugh, Michael Backes, Josh Benaloh, Marc Dacier, Herve Debar, George Dinolt, Riccardo Focardi, Virgil Gligor, Peter Gutmann, Dogan Kesdogan, Helmut Kurth, Wenke Lee, Roy Maxion, John McHugh, Catherine Meadows, Radia Perlman, Birgit Pfitzmann, Joachim Posegga, Niels Provos, Josyula R. Rao, Michael Reiter Eric Rescorla, Rei SafaviNaini, Pierangela Samarati, Andrei Serjantov, Giovanni Vigna, Dan S. Wallach, Andreas Wespi, and Marianne Winslett. We also thank the anonymous journal reviewers of the three papers published in this special issue for their work. Vern Paxson received the MS and PhD degrees from the University of California, Berkeley, and has been (and continues to be) a staff scientist with the Lawrence Berkeley National Laboratory’s Network Research Group for many years. He began at the ICIR group of the International Computer Science Institute (ICSI) in 1999. His main active research projects are Bro, worms (including the network telescope project), DETER, and PREDICT. He has been the vice chair of ACM SIGCOMM; program cochair for IEEE Security and Privacy 2005 (Program); and program committee member for SRUTI 2005, RAID 2005, ACSAC 2005, and USENIX/ACM NSDI ’05. He was on the editorial board of IEEE/ACM Transactions on Networking from 2000-2004. Vern Paxson, Michael Waidner |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2004 | Operational experiences with high-volume network intrusion detectionabstract... (NIDSs) face extreme challenges with respect to traffic volume, traffic diversity, and resource management. While crucial for acceptance and operational deployment, the research literature mainly omits such practical difficulties. In this paper, we offer an evaluation based on extensive operational experience. More specifically, we identify and explore key factors with respect to resource management and efficient packet processing and highlight their impact using a set of real-world traces. On the one hand, these insights help us gauge the trade-offs of tuning a NIDS. On the other hand, they motivate us to explore several novel ways of reducing resource requirements. These enable us to improve the state management considerably as well as balance the processing load dynamically. Overall this enables us to operate a NIDS successfully in our highvolume network environments. Holger Dreger, Anja Feldmann, Vern Paxson, Robin Sommer |
CCS | 3 |
| 2004 | Characteristics of internet background radiationabstractMonitoring any portion of the Internet address space reveals incessant activity. This holds even when monitoring traffic sent to unused addresses, which we term "background radiation. " Background radiation reflects fundamentally nonproductive traffic, either malicious (flooding backscatter, scans for vulnerabilities, worms) or benign (misconfigurations). While the general presence of background radiation is well known to the network operator community, its nature has yet to be broadly characterized. We develop such a characterization based on data collected from four unused networks in the Internet. Two key elements of our methodology are (i) the use of filtering to reduce load on the measurement system, and (ii) the use of active responders to elicit further activity from scanners in order to differentiate different types of background radiation. We break down the components of background radiation by protocol, application, and often specific exploit; analyze temporal patterns and correlated activity; and assess variations across different networks and over time. While we find a menagerie of activity, probes from worms and autorooters heavily dominate. We conclude with considerations of how to incorporate our characterizations into monitoring and detection activities. Ruoming Pang, Vinod Yegneswaran, Paul Barford, Vern Paxson, Larry L. Peterson |
Internet Measurement Conference | 4 |
| 2004 | Strategies for sound internet measurementabstractConducting an Internet measurement study in a sound fashion can be much more difficult than it might first appear. We present a number of strategies drawn from experiences for avoiding or overcoming some of the pitfalls. In particular, we discuss dealing with errors and inaccuracies; the importance of associating meta-data with measurements; the technique of calibrating measurements by examining outliers and testing for consistencies; difficulties that arise with large-scale measurements; the utility of developing a discipline for reliably reproducing analysis results; and issues with making datasets publicly available. We conclude with thoughts on the sorts of tools and community practices that can assist researchers with conducting sound measurement studies. Vern Paxson |
Internet Measurement Conference | 1 |
| 2004 | Measuring adversariesabstractMany concepts and techniques developed for general Internet measurement have counterparts in the domain of detecting and analyzing network attacks. The task is greatly complicated, however, by the fact that the object of study is adversarial: attackers do not wish to be "measured" and will take steps to thwart observation. We look at the far-ranging consequences of this different measurement environment: the analysis difficulties-some fundamental-that arise due to subtle ambiguities in the true semantics of observed traffic; new notions of "active measurement"; the highly challenging task of rapidly characterizing Internet-scale pheonmena such as global worm pandemics; the need for detailed application-level analysis and related policy and legal difficulties; attacks that target passive analysis tools; and the inherent "arms race" nature of the undertaking. Vern Paxson |
SIGMETRICS | 1 |
| 2004 | Fast Portscan Detection Using Sequential Hypothesis TestingabstractAttackers routinely perform random portscans of IP addresses to find vulnerable servers to compromise. Network intrusion detection systems (NIDS) attempt to detect such behavior and flag these portscanners as malicious. An important need in such systems is prompt response: the sooner a NIDS detects malice, the lower the resulting damage. At the same time, a NIDS should not falsely implicate benign remote hosts as malicious. Balancing the goals of promptness and accuracy in detecting malicious scanners is a delicate and difficult task. We develop a connection between this problem and the theory of sequential hypothesis testing and show that one can model accesses to local IP addresses as a random walk on one of two stochastic processes, corresponding respectively to the access patterns of benign remote hosts and malicious ones. The detection problem then becomes one of observing a particular trajectory and inferring from it the most likely classification for the remote host. We use this insight to develop TRW (Threshold Random Walk), an online detection algorithm that identifies malicious remote hosts. Using an analysis of traces from two qualitatively different sites, we show that TRW requires a much smaller number of connection attempts (4 or 5 in practice) to detect malicious activity compared to previous schemes, while also providing theoretical bounds on the low (and configurable) probabilities of missed detection and false alarms. In summary, TRW performs significantly faster and also more accurately than other current solutions. Jaeyeon Jung, Vern Paxson, Arthur W. Berger, Hari Balakrishnan |
S&P | 2 |
| 2004 | Very Fast Containment of Scanning Worms
Nicholas Weaver, Stuart Staniford-Chen, Vern Paxson |
USENIX Security Symposium | 3 |
| 2003 | Enhancing byte-level network intrusion detection signatures with contextabstractMany network intrusion detection systems (NIDS) use byte sequences as signatures to detect malicious activity. While being highly efficient, they tend to suffer from a high false-positive rate. We develop the concept of contextual signatures as an improvement of string-based signature-matching. Rather than matching fixed strings in isolation, we augment the matching process with additional context. When designing an efficient signature engine for the NIDS bro, we provide low-level context by using regular expressions for matching, and high-level context by taking advantage of the semantic information made available by bro's protocol analysis and scripting language. Therewith, we greatly enhance the signature's expressiveness and hence the ability to reduce false positives. We present several examples such as matching requests with replies, using knowledge of the environment, defining dependencies between signatures to model step-wise attacks, and recognizing exploit scans.To leverage existing efforts, we convert the comprehensive signature set of the popular freeware NIDS snort into bro's language. While this does not provide us with improved signatures by itself, we reap an established base to build upon. Consequently, we evaluate our work by comparing to snort, discussing in the process several general problems of comparing different NIDSs. Robin Sommer, Vern Paxson |
CCS | 2 |
| 2003 | A high-level programming environment for packet trace anonymization and transformationabstractPacket traces of operational Internet traffic are invaluable to network research, but public sharing of such traces is severely limited by the need to first remove all sensitive information. Current trace anonymization technology leaves only the packet headers intact, completely stripping the contents; to our knowledge, there are no publicly available traces of any significant size that contain packet payloads. We describe a new approach to transform and anonymize packet traces. Our tool provides high-level language support for packet transformation, allowing the user to write short policy scripts to express sophisticated trace transformations. The resulting scripts can anonymize both packet headers and payloads, and can perform application-level transformations such as editing HTTP or SMTP headers, replacing the content of Web items with MD5 hashes, or altering filenames or reply codes that match given patterns. We discuss the critical issue of verifying that anonymizations are both correctly applied and correctly specified, and experiences with anonymizing FTP traces from the Lawrence Berkeley National Laboratory for public release. Ruoming Pang, Vern Paxson |
SIGCOMM | 2 |
| 2003 | Active Mapping: Resisting NIDS Evasion without Altering TrafficabstractA critical problem faced by a network intrusion detection system (NIDS) is that of ambiguity. The NIDS cannot always determine what traffic reaches a given host nor how that host will interpret the traffic, and attackers may exploit this ambiguity to avoid detection or cause misleading alarms. We present a lightweight solution, active mapping, which eliminates TCP/IP-based ambiguity in a NIDS analysis with minimal runtime cost. Active mapping efficiently builds profiles of the network topology and the TCP/IP policies of hosts on the network; a NIDS may then use the host profiles to disambiguate the interpretation of the network traffic on a per-host basis. Active mapping avoids the semantic and performance problems of traffic normalization, in which traffic streams are modified to remove ambiguities. We have developed a prototype implementation of active mapping and modified a NIDS to use the active mapping-generated profile database in our tests. We found wide variation across operating systems' TCP/IP stack policies in real-world tests (about 6700 hosts), underscoring the need for this sort of disambiguation. Umesh Shankar, Vern Paxson |
S&P | 2 |
| 2002 | Observed structure of addresses in IP trafficabstractThis paper investigates the structure of addresses contained in IP traffic. Specifically, we analyze the structural characteristics of destination IP addresses seen on Interuet links, considered as a subset of the address space. These characteristics may have implications for algorithms that deal with IP address aggregates, such as routing lookups and aggregatebased congestion control. We find that address structures are well modeled by a multifractal Cantor dust with two parameters. The model may be useful for simulations where realistic IP addresses are preferred. We also develop concise characterizations of address structures, including active aggregate counts and discriminating prefixes. Our structural characterizations are stable over short time scales at a given site, and different sites have visibly different characterizations, so that the characterizations make useful fingerprints of the traffic seen at a site. Also, changing traffic conditions, such as worm propagation, significantly alter these fingerprints. Eddie Kohler, Jinyang Li 0001, Vern Paxson, Scott Shenker |
Internet Measurement Workshop | 3 |
| 2002 | Multiscale Stepping-Stone Detection: Detecting Pairs of Jittered Interactive Streams by Exploiting Maximum Tolerable Delay
David L. Donoho, Ana Georgina Flesia, Umesh Shankar, Vern Paxson, Jason Coit, Stuart Staniford-Chen |
RAID | 4 |
| 2002 | On the characteristics and origins of internet flow ratesabstractThis paper considers the distribution of the rates at which flows transmit data, and the causes of these rates. First, using packet level traces from several Internet links, and summary flow statistics from an ISP backbone, we examine Internet flow rates and the relationship between the rate and other flow characteristics such as size and duration. We find, as have others, that while the distribution of flow rates is skewed, it is not as highly skewed as the distribution of flow sizes. We also find that for large flows the size and rate are highly correlated. Second, we attempt to determine the cause of the rates at which flows transmit data by developing a tool, T-RAT, to analyze packet-level TCP dynamics. In our traces, the most frequent causes appear to be network congestion and receiver window limits. Yin Zhang 0001, Lee Breslau, Vern Paxson, Scott Shenker |
SIGCOMM | 3 |
| 2002 | How to Own the Internet in Your Spare Time
Stuart Staniford-Chen, Vern Paxson, Nicholas Weaver |
USENIX Security Symposium | 2 |
| 2001 | Inferring Link Loss Using Striped Unicast ProbesabstractIn this paper we explore the use of end-to-end unicast traffic as measurement probes to infer link-level loss rates. We leverage on of earlier work that produced efficient estimates for link-level loss rates based on end-to-end multicast traffic measurements. We design experiments based on the notion of transmitting stripes of packets (with no delay between transmission of successive packets within a stripe) to two or more receivers. The purpose of these stripes is to ensure that the correlation in receiver observations matches as closely as possible what would have been observed if the stripe had been replaced by a notional multicast probe that followed the same paths to the receivers. Measurements provide good evidence that a packet pair to distinct receivers introduces considerable correlation which can be further increased by simply considering longer stripes. We then use simulation to explore how well these stripes translate into accurate link-level loss estimates. We observe good accuracy with packet pairs, with a typical error of about 1%, which significantly decreases as stripe length is increased to 4 packets. Nick G. Duffield, Francesco Lo Presti, Vern Paxson, Don Towsley |
INFOCOM | 3 |
| 2001 | Network Intrusion Detection: Evasion, Traffic Normalization, and End-to-End Protocol Semantics
Mark Handley, Vern Paxson, Christian Kreibich |
USENIX Security Symposium | 2 |
| 2001 | Difficulties in simulating the internetabstractSimulating how the global Internet behaves is an immensely challenging undertaking because of the network's great heterogeneity and rapid change. The heterogeneity ranges from the individual links that carry the network's traffic, to the protocols that interoperate over the links, the "mix" of different applications used at a site, and the levels of congestion seen on different links. We discuss two key strategies for developing meaningful simulations in the face of these difficulties: searching for invariants and judiciously exploring the simulation parameter space. We finish with a look at a collaborative effort within the research community to develop a common network simulator. Sally Floyd, Vern Paxson |
IEEE/ACM Trans. Netw. | 2 |
| 2000 | Detecting Backdoors
Yin Zhang 0001, Vern Paxson |
USENIX Security Symposium | 2 |
| 2000 | Detecting Stepping Stones
Yin Zhang 0001, Vern Paxson |
USENIX Security Symposium | 2 |
| 1999 | An Architecture for a Global Internet Host Distance Estimation ServiceabstractThere is an increasing need for Internet hosts to be able to quickly and efficiently learn the distance, in terms of metrics such as latency or bandwidth, between Internet hosts. For example, to select the nearest of multiple equal content Web servers. This paper explores technical issues related to the creation of a public infrastructure service to provide such information. In so doing, we suggest an architecture, called IDMaps, whereby Internet distance information is distributed over the Internet, using IP multicast groups, in the form of a virtual distance map. Systems listening to the groups can estimate the distance between any pair of IP addresses by running a spanning tree algorithm over the received distance map. We also presents the results of experiments that give preliminary evidence supporting the architecture. This work thus lays the initial foundation for future work in this new area. Paul Francis, Sugih Jamin, Vern Paxson, Lixia Zhang 0001, Daniel F. Gryniewicz, Yixin Jin |
INFOCOM | 3 |
| 1999 | Defending against network IDS evasion
Vern Paxson |
Recent Advances in Intrusion Detection | 1 |
| 1999 | On Estimating End-to-End Network Path PropertiesabstractThe more information about current network conditions available to a transport protocol, the more efficiently it can use the network to transfer its data. In networks such as the Internet, the transport protocol must often form its own estimates of network properties based on measurements performed by the connection endpoints. We consider two basic transport estimation problems: determining the setting of the retransmission timer (RTO) for a reliable protocol, and estimating the bandwidth available to a connection as it begins. We look at both of these problems in the context of TCP, using a large TCP measurement set [Pax97b] for trace-driven simulations. For RTO estimation, we evaluate a number of different algorithms, finding that the performance of the estimators is dominated by their minimum values, and to a lesser extent, the timer granularity, while being virtually unaffected by how often round-trip time measurements are made or the settings of the parameters in the exponentially-weighted moving average estimators commonly used. For bandwidth estimation, we explore techniques previously sketched in the literature [Hoe96, AD98] and find that in practice they perform less well than anticipated. We then develop a receiver-side algorithm that performs significantly better. Mark Allman, Vern Paxson |
SIGCOMM | 2 |
| 1999 | Bro: a system for detecting network intruders in real-time
Vern Paxson |
Comput. Networks | 1 |
| 1999 | End-to-end internet packet dynamicsabstractWe discuss findings from a large-scale study of Internet packet dynamics conducted by tracing 20000 TCP bulk transfers between 35 Internet sites. Because we traced each 100-kbyte transfer at both the sender and the receiver, the measurements allow us to distinguish between the end-to-end behavior due to the different directions of the Internet paths, which often exhibit asymmetries. We: (1) characterize the prevalence of unusual network events such as out-of-order delivery and packet replication; (2) discuss a robust receiver-based algorithm for estimating "bottleneck bandwidth" that addresses deficiencies discovered in techniques based on "packet pair;" (3) investigate patterns of packet loss, finding that loss events are not well modeled as independent and, furthermore, that the distribution of the duration of loss events exhibits infinite variance; and (4) analyze variations in packet transit delays as indicators of congestion periods, finding that congestion periods also span a wide range of time scales. Vern Paxson |
IEEE/ACM Trans. Netw. | 1 |
| 1998 | On Calibrating Measurements of Packet Transit TimesabstractWe discuss the problem of detecting errors in measurements of the total delay experienced by packets transmitted through a wide-area network. We assume that we have measurements of the transmission times of a group of packets sent from an originating host, A, and a corresponding set of measurements of their arrival times at their destination host, B, recorded by two separate clocks. We also assume that we have a similar series of measurements of packets sent from B to A (as might occur when recording a TCP connection), but we do not assume that the clock at A is synchronized with the clock at B, nor that they run at the same frequency. We develop robust algorithms for detecting abrupt adjustments to either clock, and for estimating the relative skew between the clocks. By analyzing a large set of measurements of Internet TCP connections, we find that both clock adjustments and relative skew are sufficiently common that failing to detect them can lead to potentially large errors when an... Vern Paxson |
SIGMETRICS | 1 |
| 1998 | Bro: A System for Detecting Network Intruders in Real-Time
Vern Paxson |
USENIX Security Symposium | 1 |
| 1997 | End-to-end Internet Packet DynamicsabstractWe discuss findings from a large-scale study of Internet packet dynamics conducted by tracing 20,000 TCP bulk transfers between 35 Internet sites. Because we traced each 100 Kbyte transfer at both the sender and the receiver, the measurements allow us to distinguish between the end-to-end behaviors due to the different directions of the Internet paths, which often exhibit asymmetries. We characterize the prevalence of unusual network events such as out-of-order delivery and packet corruption; discuss a robust receiver-based algorithm for estimating "bottleneck bandwidth" that addresses deficiencies discovered in techniques based on "packet pair"; investigate patterns of packet loss, finding that loss events are not well-modeled as independent and, furthermore, that the distribution of the duration of loss events exhibits infinite variance; and analyze variations in packet transit delays as indicators of congestion periods, finding that congestion periods also span a wide range of time scales. Vern Paxson |
SIGCOMM | 1 |
| 1997 | Automated Packet Trace Analysis of TCP ImplementationsabstractWe describe tcpanaly, a tool for automatically analyzing a TCP implementation's behavior by inspecting packet traces of the TCP's activity. Doing so requires surmounting a number of hurdles, including detecting packet filter measurement errors, coping with ambiguities due to the distance between the measurement point and the TCP, and accommodating a surprisingly large range of behavior among different TCP implementations. We discuss why our efforts to develop a fully general tool failed, and detail a number of significant differences among 8 major TCP implementations, some of which, if ubiquitous, would devastate Internet performance. The most problematic TCPs were all independently written, suggesting that correct TCP implementation is fraught with difficulty. Consequently, it behooves the Internet community to develop testing programs and reference implementations. Vern Paxson |
SIGCOMM | 1 |
| 1997 | End-to-end routing behavior in the InternetabstractThe large-scale behavior of routing In the Internet has gone virtually without any formal study, the exceptions being Chinoy's (1993) analysis of the dynamics of Internet routing information, and work, similar in spirit, by Labovitz, Malan, and Jahanian (see Proc. SIGCOMM'97, 1997). We report on an analysis of 40000 end-to-end route measurements conducted using repeated "traceroutes" between 37 Internet sites. We analyze the routing behavior for pathological conditions, routing stability, and routing symmetry. For pathologies, we characterize the prevalence of routing loops, erroneous routing, infrastructure failures, and temporary outages. We find that the likelihood of encountering a major routing pathology more than doubled between the end of 1994 and the end of 1995, rising from 1.5% to 3.3%. For routing stability, we define two separate types of stability, "prevalence", meaning the overall likelihood that a particular route is encountered, and "persistence", the likelihood that a route remains unchanged over a long period of time. We find that Internet paths are heavily dominated by a single prevalent route, but that the time periods over which routes persist show wide variation, ranging from seconds up to days. About two-thirds of the Internet paths had routes persisting for either days or weeks. For routing symmetry, we look at the likelihood that a path through the Internet visits at least one different city in the two directions. At the end of 1995, this was the case half the time, and at least one different autonomous system was visited 30% of the time. Vern Paxson |
IEEE/ACM Trans. Netw. | 1 |
| 1996 | End-to-end Routing Behavior in the InternetabstractThe large-scale behavior of routing in the Internet has gone virtually without any formal study, the exception being Chinoy's analysis of the dynamics of Internet routing information [Ch93]. We report on an analysis of 40,000 end-to-end route measurements conducted using repeated "traceroutes" between 37 Internet sites. We analyze the routing behavior for pathological conditions, routing stability, and routing symmetry. For pathologies, we characterize the prevalence of routing loops, erroneous routing, infrastructure failures, and temporary outages. We find that the likelihood of encountering a major routing pathology more than doubled between the end of 1994 and the end of 1995, rising from 1.5% to 3.4%. For routing stability, we define two separate types of stability, "prevalence" meaning the overall likelihood that a particular route is encountered, and "persistence," the likelihood that a route remains unchanged over a long period of time. We find that Internet paths are heavily dominated by a single prevalent route, but that the time periods over which routes persist show wide variation, ranging from seconds up to days. About 2/3's of the Internet paths had routes persisting for either days or weeks. For routing symmetry, we look at the likelihood that a path through the Internet visits at least one different city in the two directions. At the end of 1995, this was the case half the time, and at least one different autonomous system was visited 30% of the time. Vern Paxson |
SIGCOMM | 1 |
| 1995 | Network Traffic Measurement and Modelling (Panel)abstractNetwork traffic measurement and workload characterization are key steps in the workload modeling process. Much has been learned through network measurement and workload modeling in the last ten years, but new challenges are now at the forefront: measuring network traffic in the Internet environment, understanding the implications of network traffic structure (e.g., self-similarity, autocorrelation, long range dependence), and accurate modeling of network traffic workloads for high speed network environments.This "hot topic" session brings together three prominent speakers to address each of these topics, in turn. Carey L. Williamson, Walter Willinger, Vern Paxson, Benjamin Melamed |
SIGMETRICS | 3 |
| 1995 | Wide area traffic: the failure of Poisson modelingabstractNetwork arrivals are often modeled as Poisson processes for analytic simplicity, even though a number of traffic studies have shown that packet interarrivals are not exponentially distributed. We evaluate 24 wide area traces, investigating a number of wide area TCP arrival processes (session and connection arrivals, FTP data connection arrivals within FTP sessions, and TELNET packet arrivals) to determine the error introduced by modeling them using Poisson processes. We find that user-initiated TCP session arrivals, such as remote-login and file-transfer, are well-modeled as Poisson processes with fixed hourly rates, but that other connection arrivals deviate considerably from Poisson; that modeling TELNET packet interarrivals as exponential grievously underestimates the burstiness of TELNET traffic, but using the empirical Tcplib interarrivals preserves burstiness over many time scales; and that FTP data connection arrivals within FTP sessions come bunched into "connection bursts", the largest of which are so large that they completely dominate FTP data traffic. Finally, we offer some results regarding how our findings relate to the possible self-similarity of wide area traffic.> Vern Paxson, Sally Floyd |
IEEE/ACM Trans. Netw. | 1 |
| 1994 | Wide-Area Traffic: The Failure of Poisson ModelingabstractNetwork arrivals are often modeled as Poisson processes for analytic simplicity, even though a number of traffic studies have shown that packet interarrivals are not exponentially distributed. We evaluate 21 wide-area traces, investigating a number of wide-area TCP arrival processes (session and connection arrivals, FTPDATA connection arrivals within FTP sessions, and TELNET packet arrivals) to determine the error introduced by modeling them using Poisson processes. We find that user-initiated TCP session arrivals, such as remote-login and file-transfer, are well-modeled as Poisson processes with fixed hourly rates, but that other connection arrivals deviate considerably from Poisson; that modeling TELNET packet interarrivals as exponential grievously underestimates the burstiness of TELNET traffic, but using the empirical Tcplib[DJCME92] interarrivals preserves burstiness over many time scales; and that FTPDATA connection arrivals within FTP sessions come bunched into “connection burst”, the largest of which are so large that they completely dominate FTPDATA traffic. Finally, we offer some preliminary results regarding how our findings relate to the possible self-similarity of wide-area traffic. Vern Paxson, Sally Floyd |
SIGCOMM | 1 |
| 1994 | Empirically derived analytic models of wide-area TCP connectionsabstractAnalyzes 3 million TCP connections that occurred during 15 wide-area traffic traces. The traces were gathered at five "stub" networks and two internetwork gateways, providing a diverse look at wide-area traffic. The author derives analytic models describing the random variables associated with TELNET, NNTP, SMTP, and FTP connections. To assess these models the author presents a quantitative methodology for comparing their effectiveness with that of empirical models such as Tcplib [Danzig and Jamin, 1991]. The methodology also allows to determine which random variables show significant variation from site to site, over time, or between stub networks and internetwork gateways. Overall the author finds that the analytic models provide good descriptions, and generally model the various distributions as well as empirical models.> Vern Paxson |
IEEE/ACM Trans. Netw. | 1 |