Paul Barford

dblp:b/PaulBarford · DBLP profile ↗
← Back
89ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-7874-1819ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 60 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 8Systems, architecture and hardware · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 8 · 1 first-authorSecurity and privacy · 7 · 1 since 2021Software engineering, systems software and programming languages · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Theory of computation · 1
YearPublicationVenuePosition
2026 Take the Long Way Home - Distant Peering to the Cloud
abstract
The emergence of large cloud providers in the last decade has transformed the Internet, resulting in a seemingly ever-growing set of datacenters, points of presence, and network peers. Despite the availability of closer peering locations, some networks continue to peer with cloud providers at distant locations, traveling thousands of kilometers. In this paper, we employ a novel cloud-based traceroute campaign to characterize the distances networks travel to peer with the cloud. This unique approach allows us to gain unprecedented insights into the peering patterns of networks. Our findings reveal that 50% of the networks peer within 300 kilometers of the nearest datacenter. However, our analysis also reveals that over 20% of networks travel at least 6,700 kilometers beyond the proximity of the nearest computing facility, and some as much as 18,791 kilometers! While these networks connect with the cloud worldwide, from South America to Europe and Asia, many come to peer with cloud providers in North America, even from Oceania and Asia. We explore possible motivations for the persistence of distant peering, discussing factors such as cost-effective routes, enhanced peering opportunities, and access to exclusive content.
Esteban Carisimo, Mia Weaver, Fabián E. Bustamante, Paul Barford
IEEE Trans. Netw.4
2025 A Breath of Fresh Air: Visualizing How Networks "Breathe"
abstract
Even though scientific studies have shown that humans are inherently visual creatures, nearly every published networking research paper presents its results in the form of figures such as line graphs, histograms, bar charts, or scatter plots that can be printed on paper but are often some of the least memorable aspects of a paper. In this work, we call on networking researchers to be more creative in utilizing digital media to communicate the findings of their studies and be more cognizant of the extraordinary capabilities of human readers to process and retain visual information, especially as network telemetry datasets critical for monitoring and diagnosing "network health" continue to grow in size and in the amount of semantic-rich information they contain. To illustrate what we have in mind, we consider the use case where sets of simultaneously collected time series that represent latency measurements over time between different pairs of routers or vantage points within a network (e.g., ES-net) are used to define that network's dynamically changing delay space. By representing successive snapshots of this delay space as 2D manifolds in 3D and animating the resulting manifold views, we transform the information contained in all the simultaneously collected time series into a visualization that effectively shows how a network "breathes" and that can be directly used for diagnosing aspects of a network's health.
Stephen Jasina, Loqman Salamatian, Paul Barford, Mark Crovella, Walter Willinger
HotNets3
2025 Unsteady Underwater - On the Constancy of Submarine Path Properties
Mia Weaver, Paul Barford, Fabián E. Bustamante, Esteban Carisimo, Lynne Stokes, Weili Wu 0004
Networking2
2025 Squatspotting: Towards the Systematic Measurement of Typosquatting Techniques
Wei-Shiang Wung, Calvin Kranig, Eric Pauley, Paul Barford, Mark Crovella, Joel Sommers
Networking4
2024 An Elemental Decomposition of DNS Name-to-IP Graphs
abstract
The Domain Name System (DNS) is a critical piece of Internet infrastructure with remarkably complex properties and uses, and accordingly has been extensively studied. In this study we contribute to that body of work by organizing and analyzing records maintained within the DNS as a bipartite graph. We find that relating names and addresses in this way uncovers a surprisingly rich structure. In order to characterize that structure, we introduce a new graph decomposition for DNS name-to-IP mappings, which we term elemental decomposition. In particular, we argue that (approximately) decomposing this graph into bicliques — maximally connected components — exposes this rich structure. We utilize large-scale censuses of the DNS to investigate the characteristics of the resulting decomposition, and illustrate how the exposed structure sheds new light on a number of questions about how the DNS is used in practice and suggests several new directions for future research.
Alex Anderson, Aadi Swadipto Mondal, Paul Barford, Mark Crovella, Joel Sommers
INFOCOM3
2024 Improving Scalability in Traffic Engineering via Optical Topology Programming
abstract
We present a novel framework, GreyLambda, to improve the scalability of traffic engineering (TE) systems. TE systems continuously monitor traffic and allocate network resources based on observed demands. The temporal requirement for TE is to have a time-to-solution in five minutes or less. Additionally, traffic allocations have a spatial requirement, which is to enable all traffic to traverse the network without encountering an over-subscribed link. However, the multi-commodity flowbased TE formulation cannot scale with increasing network sizes. Recent approaches have relaxed multi-commodity flow constraints to meet the temporal requirement but fail to satisfy the spatial requirement due to changing traffic demands, resulting in oversubscribed links or infeasible solutions. To satisfy both these requirements, we utilize optical topology programming (OTP) to rapidly reconfigure optical wavelengths in critical network paths and provide localized bandwidth scaling and new paths for traffic forwarding. GreyLambda integrates OTP into TE systems by introducing a heuristic algorithm that capitalizes on latent hardware resources at high-degree nodes to offer bandwidth scaling, and a method to reduce optical path reconfiguration latencies. Our experiments show that GreyLambda enhances the performance of two state-of-the-art TE systems, SMORE and NCFlow in real-world topologies with challenging traffic and link failure scenarios.
Matthew Nance Hall, Paul Barford, Klaus-Tycho Förster, Ramakrishnan Durairajan
IEEE Trans. Netw. Serv. Manag.2
2023 The CVE Wayback Machine: Measuring Coordinated Disclosure from Exploits against Two Years of Zero-Days
abstract
Software security depends on coordinated vulnerability disclosure (CVD) from researchers, a process that the community has continually sought to measure and improve. Yet, CVD practices are only as effective as the data that informs them. In this paper, we use DScope, a cloud-based interactive Internet telescope, to build statistical models of vulnerability lifecycles, bridging the data gap in over 20 years of CVD research. By analyzing application-layer Internet scanning traffic over two years, we identify real-world exploitation timelines for 63 threats. We bring this data together with six additional datasets to build a complete birth-to-death model of these vulnerabilities, the most complete analysis of vulnerability lifecycles to date. Our analysis reaches three key recommendations: (1) CVD across diverse vendors shows lower effectiveness than previously thought, (2) intrusion detection systems are underutilized to provide protection for critical vulnerabilities, and (3) existing data sources of CVD can be augmented by novel approaches to Internet measurement. In this way, our vantage point offers new opportunities to improve the CVD process, achieving a safer software ecosystem in practice.
Eric Pauley, Paul Barford, Patrick D. McDaniel
IMC2
2023 DScope: A Cloud-Native Internet Telescope
Eric Pauley, Paul Barford, Patrick D. McDaniel
USENIX Security Symposium2
2022 iGDB: connecting the physical and logical layers of the internet
abstract
Maps of physical and logical Internet connectivity that are informed by and consistent with each other can expand scope and improve accuracy in analysis of performance, robustness and security. In this paper, we describe a methodology for linking physical and logical Internet maps that aims toward a consistent, cross-layer representation. Our approach is constructive and uses geographic location as the key feature for linking physical and logical layers. We begin by building a representation of physical connectivity using online sources to identify locations that house transport hardware (i.e., PoPs, colocation centers, IXPs, etc.), and approximate locations of links between these based on shortest-path rights-of-way. We then utilize standard data sources for generating maps of IP-level and AS-level logical connectivity, and graft these onto physical maps using geographic anchors. We implement our methodology in an open-source framework called the Internet Geographic Database (iGDB), which includes tools for updating measurement data and assuring internal consistency. iGDB is built to be used with ArcGIS, a geographic information system that provides broad capability for spatial analysis and visualization. We describe the details of the iGDB implementation and demonstrate how it can be used in a variety of settings.
Scott Anderson, Loqman Salamatian, Zachary S. Bischof, Alberto Dainotti, Paul Barford
IMC5
2022 BigBen: Telemetry Processing for Internet-Wide Event Monitoring
abstract
This paper describes BigBen, a network telemetry processing system designed to enable accurate and timely reporting of Internet events (e.g., outages, attacks and configuration changes). BigBen is distinct from other Internet-wide event detection systems in its use of passive measurements of Network Time Protocol (NTP) traffic. We describe the architecture of BigBen, and a cloud-based implementation developed to process large NTP data sets and provide accurate daily event reporting. We demonstrate BigBen on a 15.5TB corpus of NTP data. We show that BigBen identifies a wide range of Internet events characterized by their location, scope and duration. We compare the events detected by BigBen vs. events detected by a large active probe-based detection system. We find only modest overlap between the two datasets and show how BigBen provides details on events that are not available from active measurements. Finally, we report on the perspective that BigBen provides on Internet events that were reported by third parties. In each case, BigBen confirms the event and provides details that were not available in prior reports, highlighting the utility of the passive, NTP-based approach.
Meenakshi Syamkumar, Yugali Gullapalli, Paul Barford, Joel Sommers
IEEE Trans. Netw. Serv. Manag.4
2022 iHorology: Lowering the Barrier to Microsecond-Level Internet Time
abstract
High accuracy, synchronized clocks are essential to a growing number of Internet applications. Standard protocols and their associated server infrastructure typically enable client clocks to synchronize to the order of tens of milliseconds. We address one of the key challenges to high precision Internet timekeeping – the intrinsic contribution to clock error of underlying path asymmetry between client and time server, a fundamental barrier to microsecond level accuracy. We first exploit results of a unique measurement study to reliably quantify asymmetry by taking routing changes into account for the first time, and then to infer the impacts on timing. We then describe three approaches to addressing the path asymmetry problem: LBBE, SBBE and K-SBBE, each based on timestamp exchange with multiple servers, with the goal of tightening bounds on asymmetry for each client. We explore their capabilities and limitations through simulation and model-based argument. We show that substantial improvements are possible, and discuss whether, and how, the goal of microsecond accuracy might be attained.
Darryl Veitch, Sathiya Kumaran Mani, Paul Barford
IEEE/ACM Trans. Netw.4
2020 Five Alarms: Assessing the Vulnerability of US Cellular Communication Infrastructure to Wildfires
abstract
Natural disasters can wreak havoc on Internet infrastructure. Short term impacts include impediments to first responders and long term impacts include requirements to repair or replace damaged physical components. In this paper, we present an analysis of the vulnerability of cellular communication infrastructure in the US to one type of natural disaster - wildfires. Three data sets are the basis for our study: historical wildfire records, wildfire risk projections, and cellular infrastructure deployment. We utilize the geographic features in each data set to assess the spatial overlap between historical wildfires and cellular infrastructure and to analyze current vulnerability. We find wide variability in the number of cell transceivers that were within wildfire perimeters over the past 18 years. In a focused analysis of the California wildfires of 2019, we find that the primary risk to cellular communication is power outage rather than cellular equipment damage. Our analysis of future risk based on wildfire hazard potential identifies California, Florida and Texas as the three states with the largest number of cell transceivers at risk. Importantly, we find that many of the areas at high risk are quite close to urban population centers, thus outages could have serious impacts on a large number of cell users. We believe that our study has important implications for governmental communication assurance efforts and for risk planning by cell infrastructure owners and service providers.
Scott Anderson, Carol Barford, Paul Barford
Internet Measurement Conference3
2018 Assessing Candidate Preference through Web Browsing History
abstract
Predicting election outcomes is of considerable interest to candidates, political scientists, and the public at large. We propose the use of Web browsing history as a new indicator of candidate preference among the electorate, one that has potential to overcome a number of the drawbacks of election polls. However, there are a number of challenges that must be overcome to effectively use Web browsing for assessing candidate preference - including the lack of suitable ground truth data and the heterogeneity of user populations in time and space. We address these challenges, and show that the resulting methods can shed considerable light on the dynamics of voters' candidate preferences in ways that are difficult to achieve using polls.
Giovanni Comarela, Ramakrishnan Durairajan, Paul Barford, Dino P. Christenson, Mark Crovella
KDD3
2018 Device Graphing by Example
abstract
Datasets that organize and associate the many identifiers produced by PCs, smartphones, and tablets accessing the internet are referred to as internet device graphs . In this paper, we demonstrate how measurement, tracking, and other internet entities can associate multiple identifiers with a single device or user after coarse associations, e.g ., based on IP-colocation , are made. We employ a Bayesian similarity algorithm that relies on examples of pairs of identifiers and their associated telemetry, including user agent, screen size, and domains visited, to establish pair-wise scores. Community detection algorithms are applied to group identifiers that belong to the same device or user. We train and validate our methodology using a unique dataset collected from a client panel with full visibility, apply it to a dataset of 700 million device identifiers collected over the course of six weeks in the United States, and show that it outperforms several unsupervised learning approaches. Results show mean precision and recall exceeding 90% for association of identifiers at both the device and user levels.
Keith Funkhouser, Matthew Malloy, Enis Ceyhun Alp, Phillip Poon, Paul Barford
KDD5
2017 Explicit Defense Actions Against Test-Set Attacks
abstract
Automated learning and decision making systems in public-facing applications are vulnerable to malicious attacks. Examples of such systems include spam detectors, credit card fraud detectors, and network intrusion detection systems. These systems are at further risk of attack when money is directly involved, such as market forecasters or decision systems used in determining insurance or loan rates. In this paper, we consider the setting where a predictor Bob has a fixed model, and an unknown attacker Alice aims to perturb (or poison) future test instances so as to alter Bob's prediction to her benefit. We focus specifically on Bob's optimal defense actions to limit Alice's effectiveness. We define a general framework for determining Bob's optimal defense action against Alice's worst-case attack. We then demonstrate our framework by considering linear predictors, where we provide tractable methods of determining the optimal defense action. Using these methods, we perform an empirical investigation of optimal defense actions for a particular class of linear models -- autoregressive forecasters -- and find that for ten real world futures markets, the optimal defense action reduces the Bob's loss by between 78 and 97%.
Scott Alfeld, Xiaojin Zhu 0001, Paul Barford
AAAI3
2017 On the structure and characteristics of user agent string
abstract
User agent (UA) strings transmitted during HTTP transactions convey client system configuration details to ensure that content returned by a server is appropriate for the requesting host. As such, analysis of UA strings and their structure offers a unique perspective on active client systems in the Internet and when tracked longitudinally, offers a perspective on the nature of system and configuration dynamics. In this paper, we describe our study of UA string characteristics. Our work is based on analyzing a unique corpus of over 1B UA strings collected over a period of 2 years by comScore. We begin by analyzing the general characteristics of UA strings, focusing on the most prevalent strings and dynamic behaviors. We identify the top 10 most popular User Agents, which account for 26% of total daily volume. These strings describe the expected instances of popular platforms such as Microsoft, Apple and Google. We then report on the characteristics of low-volume UA strings, which has important implications for unique device identification. We show that this class of user agent generates the overwhelming majority of traffic, with between 2M and 10M instances observed each day. We show that the distribution of UA strings has temporal dependence and we show the distribution measured depends on the type of content served. Finally, we report on two large-scale UA anomalies characterized by web browsers sending false and misleading UAs in their web requests.
Jeffery Kline, Paul Barford, Aaron Cahn, Joel Sommers
Internet Measurement Conference2
2017 Automatic metadata generation for active measurement
abstract
Empirical research in the Internet is fraught with challenges. Among these is the possibility that local environmental conditions (e.g., CPU load or network load) introduce unexpected bias or artifacts in measurements that lead to erroneous conclusions. In this paper, we describe a framework for local environment monitoring that is designed to be used during Internet measurement experiments. The goals of our work are to provide a critical, expanded perspective on measurement results and to improve the opportunity for reproducibility of results. We instantiate our framework in a tool we call SoMeta, which monitors the local environment during active probe-based measurement experiments. We evaluate the runtime costs of SoMeta and conduct a series of experiments in which we intentionally perturb different aspects of the local environment during active probe-based measurements. Our experiments show how simple local monitoring can readily expose conditions that bias active probe-based measurement results. We conclude with a discussion of how our framework can be expanded to provide metadata for a broad range of Internet measurement experiments.
Joel Sommers, Ramakrishnan Durairajan, Paul Barford
Internet Measurement Conference3
2017 Internet Device Graphs
abstract
Internet device graphs identify relationships between user-centric internet connected devices such as desktops, laptops, smartphones, tablets, gaming consoles, TV's, etc. The ability to create such graphs is compelling for online advertising, content customization, recommendation systems, security, and operations. We begin by describing an algorithm for generating a device graph based on IP-colocation, and then apply the algorithm to a corpus of over 2.5 trillion internet events collected over the period of six weeks in the United States. The resulting graph exhibits immense scale with greater than 7.3 billion edges (pair-wise relationships) between more than 1.2 billion nodes (devices), accounting for the vast majority of internet connected devices in the US. Next, we apply community detection algorithms to the graph resulting in a partitioning of internet devices into 100 million small communities representing physical households. We validate this partition with a unique ground truth dataset. We report on the characteristics of the graph and the communities. Lastly, we discuss the important issues of ethics and privacy that must be considered when creating and studying device graphs, and suggest further opportunities for device graph enrichment and application.
Matthew Malloy, Paul Barford, Enis Ceyhun Alp, Jonathan Koller, Adria Jewell
KDD2
2016 Data Poisoning Attacks against Autoregressive Models
abstract
Forecasting models play a key role in money-making ventures in many different markets. Such models are often trained on data from various sources, some of which may be untrustworthy.An actor in a given market may be incentivised to drive predictions in a certain direction to their own benefit.Prior analyses of intelligent adversaries in a machine-learning context have focused on regression and classification.In this paper we address the non-iid setting of time series forecasting.We consider a forecaster, Bob, using a fixed, known model and a recursive forecasting method.An adversary, Alice, aims to pull Bob's forecasts toward her desired target series, and may exercise limited influence on the initial values fed into Bob's model.We consider the class of linear autoregressive models, and a flexible framework of encoding Alice's desires and constraints.We describe a method of calculating Alice's optimal attack that is computationally tractable, and empirically demonstrate its effectiveness compared to random and greedy baselines on synthetic and real-world time series data.We conclude by discussing defensive strategies in the face of Alice-like adversaries.
Scott Alfeld, Xiaojin Zhu 0001, Paul Barford
AAAI3
2016 What's in the community cookie jar?
abstract
Third party tracking of user behavior via web cookies represents a privacy threat. In this paper we assess this threat through an analysis of anonymized, crowdsourced cookie data provided by Cookiepedia.co.uk. We find that nearly 45% of the cookies in the corpus are from Facebook and of the remaining cookies 25% come from 10 distinct domains. Over 65% are Maximal Permission cookies (i.e., 3rd party, non-secure, persistent, root-level). Cookiepedia's anonymization of user data presents challenges with respect to modeling site traffic. We further elucidate the privacy issue by conducting targeted crawling campaigns to supplement the Cookiepedia data. We find that the amount of traffic obscured by Cookiepedia's anonymizing procedure varies dramatically from site to site - sometimes obscuring as much as 80% of traffic. We use the crawls to infer the inverse function of the anonymizing procedure, allowing us to enhance the crowdsourced dataset while maintaining user anonymity.
Aaron Cahn, Scott Alfeld, Paul Barford, S. Muthukrishnan 0001
ASONAM3
2016 Ad Blockers: Global Prevalence and Impact
Matthew Malloy, Mark McNamara, Aaron Cahn, Paul Barford
Internet Measurement Conference4
2016 MNTP: Enhancing Time Synchronization for Mobile Devices
Sathiya Kumaran Mani, Ramakrishnan Durairajan, Paul Barford, Joel Sommers
Internet Measurement Conference3
2016 Machine Teaching as Search
abstract
Machine teaching (MT) studies the task of designing a training set. Specifically, given a learner (e.g., an artificial neural network or a human) and a target model, a teacher aims to create a training set which results in the target model being learned. MT applications include optimal education design for human learners and computer security where adversaries aim to attack learning-based systems. In this work, we formulate pool-based MT as a state space search problem. We discuss the properties and challenges of the resulting problem and highlight opportunities for novel search techniques. In our preliminary study we use a beam search approach, and find that training and evaluating empirical risk of models dominate the run time of the search. Toward the goal of better search techniques for future work, we develop optimizations ranging from implementation details for specific learners to algorithm changes applicable to general blackbox learners. We conclude with a discussion of open problems and research directions.
Scott Alfeld, Xiaojin Zhu 0001, Paul Barford
SOCS3
2016 Bigfoot: A geo-based visualization methodology for detecting BGP threats
abstract
Studies of inter-domain routing in the Internet have highlighted the complex and dynamic nature of connectivity changes that take place daily on a global scale. The ability to assess and identify normal, malicious, irregular and unexpected behaviors in routing update streams is important in daily network and security operations. In this paper we describe Bigfoot, a Border Gateway Protocol (BGP) update visualization system that has been designed to highlight and assess a wide variety of behaviors in update streams. At the core of Bigfoot is the notion of visualizing the announcements of network prefixes via IP geolocation. We investigate different representations of polygons for network footprints and show how straightforward application of IP geolocation can lead to representations that are difficult to interpret. Bigfoot includes techniques to filter, organize, analyze and visualize BGP updates that enable characteristics and behaviors of interest to be identified effectively. To demonstrate Bigfoot's capabilities, we consider 1.79B BGP updates collected over a period of one year and identify 139 candidate events in this data. We investigate a subset of these events in detail, along with ground truth from existing literature to show how network footprint visualizations can be used in operational deployments.
Meenakshi Syamkumar, Ramakrishnan Durairajan, Paul Barford
VizSEC3
2016 An Empirical Study of Web Cookies
abstract
Web cookies are used widely by publishers and 3rd parties to track users and their behaviors. Despite the ubiquitous use of cookies, there is little prior work on their characteristics such as standard attributes, placement policies, and the knowledge that can be amassed via 3rd party cookies. In this paper, we present an empirical study of web cookie characteristics, placement practices and information transmission. To conduct this study, we implemented a lightweight web crawler that tracks and stores the cookies as it navigates to websites. We use this crawler to collect over 3.2M cookies from the two crawls, separated by 18 months, of the top 100K Alexa web sites. We report on the general cookie characteristics and add context via a cookie category index and website genre labels. We consider privacy implications by examining specific cookie attributes and placement behavior of 3rd party cookies. We find that 3rd party cookies outnumber 1st party cookies by a factor of two, and we illuminate the connection between domain genres and cookie attributes. We find that less than 1% of the entities that place cookies can aggregate information across 75% of web sites. Finally, we consider the issue of information transmission and aggregation by domains via 3rd party cookies. We develop a mathematical framework to quantify user information leakage for a broad class of users, and present findings using real world domains. In particular, we demonstrate the interplay between a domain's footprint across the Internet and the browsing behavior of users, which has significant impact on information transmission.
Aaron Cahn, Scott Alfeld, Paul Barford, S. Muthukrishnan 0001
WWW3
2015 Time's Forgotten: Using NTP to understand Internet Latency
abstract
The performance of Internet services is intrinsically tied to propagation delays between end points (i.e., network latency). Standard active probe-based or passive host-based methods for measuring end-to-end latency are difficult to deploy at scale and typically offer limited precision and accuracy. In this paper, we investigate a novel but non-obvious source of latency measurement---logs from network time protocol (NTP) servers. Using NTP-derived data for studying latency is compelling due to NTP's pervasive use in the Internet and its inherent focus on accurate end-to-end delay estimation. We consider the efficacy of an NTP-based approach for studying propagation delays by analyzing logs collected from 10 NTP servers distributed across the United States. These logs include over 73M latency measurements to 7.4M worldwide clients (as indicated by unique IP addresses) collected over the period of one day. Our initial analysis of the general characteristics of propagation delays derived from the log data reveals that delay measurements from NTP must be carefully filtered in order to extract accurate results. We develop a filtering process that removes measurements that are likely to be inaccurate. After applying our filter to NTP measurements, we report on the scope and reach for US-based clients and the characteristics of the end-to-end latency for those clients.
Ramakrishnan Durairajan, Sathiya Kumaran Mani, Joel Sommers, Paul Barford
HotNets4
2015 Contamination estimation via convex relaxations
abstract
Identifying anomalies and contamination in datasets is important in a wide variety of settings. In this paper, we describe a new technique for estimating contamination in large, discrete valued datasets. Our approach considers the normal condition of the data to be specified by a model consisting of a set of distributions. Our key contribution is in our approach to contamination estimation. Specifically, we develop a technique that identifies the minimum number of data points that must be discarded (i.e., the level of contamination) from an empirical data set in order to match the model to within a specified goodness-of-fit, controlled by a p-value. Appealing to results from large deviations theory, we show a lower bound on the level of contamination is obtained by solving a series of convex programs. Theoretical results guarantee the bound converges at a rate of O(√log(p)/p), where p is the size of the empirical data set.
Matthew Malloy, Scott Alfeld, Paul Barford
ISIT3
2015 InterTubes: A Study of the US Long-haul Fiber-optic Infrastructure
abstract
The complexity and enormous costs of installing new long-haul fiber-optic infrastructure has led to a significant amount of infrastructure sharing in previously installed conduits. In this paper, we study the characteristics and implications of infrastructure sharing by analyzing the long-haul fiber-optic network in the US. We start by using fiber maps provided by tier-1 ISPs and major cable providers to construct a map of the long-haul US fiber-optic infrastructure. We also rely on previously under-utilized data sources in the form of public records from federal, state, and municipal agencies to improve the fidelity of our map. We quantify the resulting map's connectivity characteristics and confirm a clear correspondence between long-haul fiber-optic, roadway, and railway infrastructures. Next, we examine the prevalence of high-risk links by mapping end-to-end paths resulting from large-scale traceroute campaigns onto our fiber-optic infrastructure map. We show how both risk and latency (i.e., propagation delay) can be reduced by deploying new links along previously unused transportation corridors and rights-of-way. In particular, focusing on a subset of high-risk links is sufficient to improve the overall robustness of the network to failures. Finally, we discuss the implications of our findings on issues related to performance, net neutrality, and policy decision-making.
Ramakrishnan Durairajan, Paul Barford, Joel Sommers, Walter Willinger
SIGCOMM2
2014 Controller-agnostic SDN Debugging
abstract
Complexity in software-defined network (SDN) applications calls for methods and tools that can facilitate comprehensive debugging and analysis. A key challenge in this regard is that SDN configurations interact with network devices that can behave in unexpected ways, depending on factors such as traffic and application mix. In this paper, we describe OFf, a debugging and test environment for SDN developers. OFf is built on top of the fs-sdn simulator, which was developed to offer simple-to-use, accurate and scalable evaluation of OpenFlow-based SDN configurations. OFf offers standard debugging features for controller applications such as stepping, breakpoints, and watch variables. It also offers features that provide visibility into network behavior including packet tracing, packet replay and visualization features, and alerts that are triggered when, e.g., configurations change. OFf is accessed through a text interface and is designed to interoperate with any standard SDN controller platform. We demonstrate the capabilities of OFf through three test scenarios that illustrate its utility and modest performance impact on running applications. Specifically, we show how OFf can be used to analyze and fix bugs in a traffic engineering application, and to detect and repair a security vulnerability due to multiple application interaction and unexpected rule expiration.
Ramakrishnan Durairajan, Joel Sommers, Paul Barford
CoNEXT3
2014 Layer 1-informed Internet Topology Measurement
abstract
Understanding the Internet's topological structure continues to be fraught with challenges. In this paper, we investigate the hypothesis that physical maps of service provider infrastructure can be used to effectively guide topology discovery based on network layer TTL-limited measurement. The goal of our work is to focus layer 3-based probing on broadly identifying Internet infrastructure that has a fixed geographic location such as POPs, IXPs and other kinds of hosting facilities. We begin by comparing more than 1.5 years of TTL-limited probe data from the Ark project with maps of service provider infrastructure from the Internet Atlas project. We find that there are substantially more nodes and links identified in the service provider map data versus the probe data. Next, we describe a new method for probe-based measurement of physical infrastructure called POPsicle that is based on careful selection of probe source-destination pairs. We demonstrate the capability of our method through an extensive measurement study using existing "looking glass" vantage points distributed throughout the Internet and show that it reveals 2.4 times more physical node locations versus standard probing methods. To demonstrate the deployability of POPsicle we also conduct tests at an IXP. Our results again show that POPsicle can identify more physical node locations compared with standard layer 3 probes, and through this deployment approach it can be used to measure thousands of networks world wide.
Ramakrishnan Durairajan, Joel Sommers, Paul Barford
Internet Measurement Conference3
2014 A QoE Perspective on Sizing Network Buffers
abstract
Despite decades of operational experience and focused research efforts, standards for sizing and configuring buffers in network systems remain controversial. An extreme example of this is the recent claim that excessive buffering (i.e., bufferbloat) can severely impact Internet services. In this paper, we systematically examine the implications of buffer sizing choices from the perspective of factors impacting end user experience. To assess user perception of application quality under various buffer sizing schemes we employ Quality of Experience (QoE) metrics. We evaluate these metrics over a wide range of end-user applications (e.g., web browsing, VoIP, and RTP video streaming) and workloads in two realistic testbeds emulating access and backbone networks. The main finding of our extensive evaluations is that network workload, rather than buffer size, is the primary determinant of end user QoE. Our results also highlight the relatively narrow conditions under which bufferbloat seriously degrades QoE, i.e., when buffers are oversized and sustainably filled.
Oliver Hohlfeld, Enric Pujol-Gil, Florin Ciucu, Anja Feldmann, Paul Barford
Internet Measurement Conference5
2014 Adscape: harvesting and analyzing online display ads
abstract
Over the past decade, advertising has emerged as the primary source of revenue for many web sites and apps. In this paper we report a first-of-its-kind study that seeks to broadly understand the features, mechanisms and dynamics of display advertising on the web - i.e., the Adscape. Our study takes the perspective of users who are the targets of display ads shown on web sites. We develop a scalable crawling capability that enables us to gather the details of display ads including creatives and landing pages. Our crawling strategy is focused on maximizing the number of unique ads harvested. Of critical importance to our study is the recognition that a user's profile (i.e., browser profile and cookies) can have a significant impact on which ads are shown. We deploy our crawler over a variety of websites and profiles and this yields over 175K distinct display ads. We find that while targeting is widely used, there remain many instances in which delivered ads do not depend on user profile; further, ads vary more over user profiles than over websites. We also assess the population of advertisers seen and identify over 3.7K distinct entities from a variety of business segments. Finally, we find that when targeting is used, the specific types of ads delivered generally correspond with the details of user profiles, and also on users' patterns of visit.
Paul Barford, Igor Canadi, Darja Krushevskaja, Qiang Ma 0006, S. Muthukrishnan 0001
WWW1
2013 RiskRoute: a framework for mitigating network outage threats
abstract
A comprehensive understanding of outage threats is critical for robust network design and operation, and evaluating cost trade-offs for recovery planning. In this paper, we describe a study of network infrastructure events due to outage events and a framework for mitigating these risks through backup routing and additional provisioning. We evaluate risk via the concept of bit-risk miles, the geographically-scaled outage risk of traffic in a network. Our focus on bit-risk miles allows for first-of-its-kind analysis of the tradeoffs of shortest path routing and risk-averse routing. We leverage the concept of bit-risk miles to present RiskRoute, a flexible routing framework that allows for backup routes to be configured to respond to both historical and immediately forecasted outage threats. Specifically, RiskRoute is an optimization framework that minimizes bit-risk miles between arbitrary points in a network. RiskRoute also reveals the best locations for provisioning additional network infrastructure in the form of new PoP-to-PoP links for single-network domains, and the best new peering relationships for multi-network domains. To assess and evaluate RiskRoute, we assemble diverse data sets including (i) - detailed topological maps and peering relationships of Internet Service Providers (ISPs) in the US, and (ii) - historical information on different types of natural disasters which threaten physical infrastructure. Our analysis reveals the providers that have the highest risk to disaster-based outage events. We also provide provisioning recommendations for network operators that can in some cases significantly lower bit-risk miles for their infrastructures.
Brian Eriksson, Ramakrishnan Durairajan, Paul Barford
CoNEXT3
2013 Assessing performance of Internet services on IPv6
abstract
The exhaustion of the IPv4 address space significantly increases the urgency for transitions to IPv6. Since native IPv6 support is not yet ubiquitous, a major concern of users and service providers (e.g., Facebook, Google, etc.) is that end-to-end performance via IPv6 could be substantially worse than IPv4. In this paper, we develop an analysis method and framework that matches DNS rendezvous information to flows so that we can compare and contrast performance over both protocols for a variety of Internet services. Our initial analyses focus on the basic services that are accessed using both protocols, observed client behaviors, and a presentation of performance characteristics of services using both IPv4 and IPv6. Our objective is to detect and expose differences by passive measurement without access to application traffic payloads. To demonstrate our method, we present results of an empirical feasibility study that considers the issue of Internet services performance over IPv6. Our study uses data collected on the World IPv6 Day, including both DNS requests/responses and flow export records for dual-stack hosts operating at a large research university. Our results expose various performance characteristics of Internet services that support IPv6: (1) Robust measures of services' flow bit rate distributions vary significantly by time of day, by number of active local clients, and by IP protocol version (6 or 4). (2) These rate characteristics differ amongst services. (3) There are regimes of time in which IPv6 flow bit rates exceed those of IPv4 and others where the IPv4 flow rates exceed those of IPv6.
David Plonka, Paul Barford
ISCC2
2013 Impression Fraud in On-line Advertising via Pay-Per-View Networks
Kevin Springborn, Paul Barford
USENIX Security Symposium2
2012 Revisiting broadband performance
abstract
Understanding the empirical characteristics of broadband performance is of intrinsic importance to users and providers, and has been a significant focus of recent efforts by the Federal Communications Commission (FCC)[9]. A series of recent studies have reported results of empirical studies of broadband performance (e.g.,[11,15,22]). In this paper, we reappraise previous empirical findings on broadband performance. Our study is based on a unique corpus of crowd-sourced data consisting of over 54 million individual tests collected from 59 metropolitan markets over a 6 month period by Speedtest.net. Following analytic approaches from prior studies, our results confirm many of the raw performance results (upload/download/latency) for ISPs in specific US markets. However, the size and scope of our data enable us to examine the details of characteristics that were not identified in prior studies, thereby providing a more comprehensive view of broadband performance. Furthermore, we also report results of broadband performance characteristics in 35 metropolitan markets outside of the US. This not only provides an important baseline for future study in those markets, but also enables relative comparison of broadband performance between markets world wide.
Igor Canadi, Paul Barford, Joel Sommers
Internet Measurement Conference2
2012 Cell vs. WiFi: on the performance of metro area mobile connections
abstract
Cellular and 802.11 WiFi are compelling options for mobile Internet connectivity. The goal of our work is to understand the performance afforded by each of these technologies in diverse environments and use conditions. In this paper, we compare and contrast cellular and WiFi performance using crowd-sourced data from Speedtest.net. Our study considers spatio-temporal performance (upload/download throughput and latency) using over 3 million user-initiated tests from iOS and Android apps in 15 different metro areas collected over a 15 week period. Our basic performance comparisons show that (i) WiFi provides better absolute download/upload throughput, and a higher degree of consistency in performance; (ii) WiFi networks generally deliver lower absolute latency, but the consistency in latency is often better with cellular access; (iii) throughput and latency vary widely depending on the particular access type e.g., HSPA, EVDO, LTE, WiFi, etc.) and service provider. More broadly, our results show that performance consistency for cellular and WiFi is much lower than has been reported for wired broadband. Temporal analysis shows that average performance for cell and WiFi varies with time of day, with the best performance for large metro areas coming at non-peak hours. Spatial analysis shows that performance is highly variable across metro areas, but that there are subregions that offer consistently better performance for cell or WiFi. Comparisons between metro areas show that larger areas provide higher throughput and lower latency than smaller metro areas, suggesting where ISPs have focused their deployment efforts. Finally, our analysis reveals diverse performance characteristics resulting from the rollout of new cell access technologies and service differences among local providers.
Joel Sommers, Paul Barford
Internet Measurement Conference2
2012 Intrusion as (anti)social communication: characterization and detection
abstract
A reasonable definition of intrusion is: entering a community to which one does not belong. This suggests that in a network, intrusion attempts may be detected by looking for communication that does not respect community boundaries. In this paper, we examine the utility of this concept for identifying malicious network sources. In particular, our goal is to explore whether this concept allows a core-network operator using flow data to augment signature-based systems located at network edges. We show that simple measures of communities can be defined for flow data that allow a remarkably effective level of intrusion detection simply by looking for flows that do not respect those communities. We validate our approach using labeled intrusion attempt data collected at a large number of edge networks. Our results suggest that community-based methods can offer an important additional dimension for intrusion detection systems.
Natallia Katenka, Paul Barford, Eric D. Kolaczyk, Mark Crovella
KDD3
2012 Comparing metro-area cellular and WiFi performance: extended abstract
abstract
Cellular and 802.11 WiFi offer two compelling connectivity options for mobile users. The goal of our work is to better understand performance characteristics of these technologies in diverse environments and conditions. To that end, we compare and contrast cellular and Wifi performance using crowd-sourced data from speedtest.net. We consider spatio-temporal performance aspects (e.g., upload and download throughput and latency) using over 3 million user-initiated tests initiated in 15 different metro areas, collected over 15 weeks. In these preliminary results, we find that WiFi performance generally exceeds cellular performance, and that observed characteristics are highly variable across different locations and times of day. We also observe diverse performance characteristics resulting from the rollout of new cell access technologies and service differences among local providers.
Joel Sommers, Paul Barford
SIGMETRICS2
2012 Efficient Network Tomography for Internet Topology Discovery
abstract
Accurate and timely identification of the router-level topology of the Internet is one of the major unresolved problems in Internet research. Topology recovery via tomographic inference is potentially an attractive complement to standard methods that use TTL-limited probes. Unfortunately, limitations of prior tomographic techniques make timely resolution of large-scale topologies impossible due to the requirement of an infeasible number of measurements. In this paper, we describe new techniques that aim toward efficient tomographic inference for accurate router-level topology measurement. We introduce methodologies based on Depth-First Search (DFS) ordering that clusters end-hosts based on shared infrastructure and enables the logical tree topology of a network to be recovered accurately and efficiently. We evaluate the capabilities of our algorithms in large-scale simulation and find that our methods will reconstruct topologies using less than 2% of the measurements required by exhaustive methods and less than 15% of the measurements needed by the current state-of-the-art tomographic approach. We also present results from a study of the live Internet where we show our DFS-based methodologies can recover the logical router-level topology more accurately and with fewer probes than prior techniques.
Brian Eriksson, Gautam Dasarathy, Paul Barford, Robert D. Nowak
IEEE/ACM Trans. Netw.3
2011 On the prevalence and characteristics of MPLS deployments in the open internet
abstract
Multi-Protocol Label Switching (MPLS) is a mechanism that enables service providers to specify virtual paths through IP networks. The use of MPLS in the open Internet (i.e., public end-to-end paths) has important implications for users and network neutrality since MPLS is frequently used in traffic engineering applications today. In this paper we present a longitudinal study of the prevalence and characteristics of MPLS deployments in the open Internet. We use path measurement data collected over the past 3.5 years by the CAIDA Archipelago project (Ark), which consist of over 10 billion individual traceroutes between hosts throughout the Internet. We use two different techniques for identifying MPLS paths in Ark data: direct observation via ICMP extensions that include MPLS label information, and inference using a Bayesian data fusion methodology. Our direct observation method can only identify uniform-mode tunnels, which very likely underestimates MPLS deployments. Nonetheless, our results show that the total number of tunnels observed in a given measurement period has varied widely over time with the largest deployments in tier-1 providers. About 7% of all autonomous systems deploy MPLS and this level of deployment has been consistent over the past three years. The average length of an MPLS tunnel has decreased from 4 hops in 2008 to 3 hops in 2011, and the path length distribution is heavily skewed. About 25% of all paths in 2011 cross at least one MPLS tunnel, while 4% cross more than one. Finally, data observed in MPLS headers suggest that many ASes employ some types of traffic classification and engineering in their tunnels.
Joel Sommers, Paul Barford, Brian Eriksson
Internet Measurement Conference2
2011 DomainImpute: Inferring unseen components in the Internet
abstract
Despite many efforts over the past decade, the ability to generate topological maps of the Internet at the router-level accurately and in a timely fashion remains elusive. Mapping campaigns commonly involve traceroute-like probing that are usually non-adaptive and incomplete, thus revealing only a portion of the underlying topology. In this paper we demonstrate that standard probing methods yield datasets that implicitly contain information about much more than just the directly observed links and routers. Each probe, in addition to the underlying domain knowledge, returns information that places constraints on the underlying topology, and by integrating a large number of such constraints it is possible to accurately infer the existence of unseen components of the Internet. We describe DomainImpute, a novel data analysis methodology designed to accurately infer the unseen hop-count distances between observed routers. We use both synthetic and a large empirical dataset to validate the proposed methods. On our empirical real world dataset, we show that our methods can estimate over 55% of the unseen distances between observed routers to within a one-hop error.
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
INFOCOM2
2011 Fingerprinting 802.11 rate adaption algorithms
abstract
The effectiveness of rate adaptation algorithms is an important determinant of 802.11 wireless network performance. The diversity of algorithms that has resulted from efforts to improve rate adaptation has introduced a new dimension of variability into 802.11 wireless networks, further complicating the already difficult task of understanding and debugging 802.11 performance. To assist with this task, in this paper we present and evaluate a methodology for accurately fingerprinting 802.11 rate adaptation algorithms. Our approach uses a Support Vector Machine (SVM)-based classifier that requires only simple passive measurements of 802.11 traffic. We demonstrate that careful conversion of raw packet traces into input features for SVM is necessary for achieving high classification accuracy. We tested our classifier on the four rate adaptation algorithms available in MadWifi, cards. The classifier performs with an accuracy of 95% – 100%. We also show that the classifier is robust over a variety of network conditions if the training data includes a sufficient sampling of the range of an algorithm's behavior.
Mariyam Mirza, Paul Barford, Xiaojin Zhu 0001, Suman Banerjee 0001, Michael Blodgett
INFOCOM2
2011 Efficient network-wide flow record generation
abstract
Experiments on diverse topics such as network measurement, management and security are routinely conducted using empirical flow export traces. However, the availability of empirical flow traces from operational networks is limited and frequently comes with significant restrictions. Furthermore, empirical traces typically lack critical meta-data (e.g., labeled anomalies) which reduce their utility in certain contexts. In this paper, we describe fs: a first-of-its-kind tool for automatically generating representative flow export records as well as basic SNMP-like router interface counts. fs generates measurements for a target network topology with specified traffic characteristics. The resulting records for each router in the topology have byte, packet and flow characteristics that are representative of what would be seen in a live network. fs also includes the ability to inject different types of anomalous events that have precisely defined characteristics, thereby enabling evaluation of proposed attack and anomaly detection methods. We validate fs by comparing it with the ns-2 simulator, which targets accurate recreation of packet-level dynamics in small network topologies. We show that data generated by fs are virtually identical to what are generated by ns-2, except over small time scales (below 1 second). We also show that fs is highly efficient, thus enabling test sets to be created for large topologies. Finally, we demonstrate the utility of fs through an assessment of anomaly detection algorithms, highlighting the need for flexible, scalable generation of network-wide measurement data with known ground truth.
Joel Sommers, Rhys Alistair Bowden, Brian Eriksson, Paul Barford, Matthew Roughan, Nick G. Duffield
INFOCOM4
2011 Inferring Unseen Components of the Internet Core
abstract
Despite many efforts over the past decade, the ability to generate topological maps of the Internet at the router-level accurately and in a timely fashion remains elusive. Mapping campaigns commonly involve {t traceroute}-like probing that are usually non-adaptive and incomplete, thus revealing only a portion of the underlying topology. In this paper we demonstrate that standard probing methods yield datasets that implicitly contain information about much more than just the directly observed links and routers. Each probe yields information that places constraints on the underlying topology, and by integrating a large number of such constraints it is possible to accurately infer the existence of unseen components of the Internet (i.e., links and routers not directly revealed by the probing). Moreover, we show that this information can be used to adaptively re-focus the probing in order to more quickly discover the topology. These findings suggest radically new and more efficient approaches to Internet mapping. Our work focuses on the discovery of the core of the Internet. We define "Internet core" as the set of routers that is roughly bounded by ingress/egress routers from stub autonomous systems. We describe a novel data analysis methodology designed to accurately infer (i) the number of unseen core routers, (ii) the unseen hop-count distances between observed routers, and (iii) unseen links between observed routers. We use a large experimental dataset to validate the proposed methods. For our data set, we show that our methods can predict the number of unseen routers to within a 13% error level, estimate 60% of the unseen distances between observed routers to within a one-hop error, and robustly detect over 35% of the unseen links between observed routers. Furthermore, we use the information extracted by our inference methodology to drive an adaptive active-probing scheme. The adaptive probing method allows us to generate maps on our data set using 50% fewer probes than standard non-adaptive approaches.
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
IEEE J. Sel. Areas Commun.2
2010 BasisDetect: a model-based network event detection framework
abstract
The ability to detect unexpected events in large networks can be a significant benefit to daily network operations. A great deal of work has been done over the past decade to develop effective anomaly detection tools, but they remain virtually unused in live network operations due to an unacceptably high false alarm rate. In this paper, we seek to improve the ability to accurately detect unexpected network events through the use of BasisDetect, a flexible but precise modeling framework. Using a small dataset with labeled anomalies, the BasisDetect framework allows us to define large classes of anomalies and detect them in different types of network data, both from single sources and from multiple, potentially diverse sources. Network anomaly signal characteristics are learned via a novel basis pursuit based methodology. We demonstrate the feasibility of our BasisDetect framework method and compare it to previous detection methods using a combination of synthetic and real-world data. In comparison with previous anomaly detection methods, our BasisDetect methodology results show a 50% reduction in the number of false alarms in a single node dataset, and over 65% reduction in false alarms for synthetic network-wide data.
Brian Eriksson, Paul Barford, Rhys Alistair Bowden, Nick G. Duffield, Joel Sommers, Matthew Roughan
Internet Measurement Conference2
2010 Toward the Practical Use of Network Tomography for Internet Topology Discovery
abstract
Accurate and timely identification of the router-level topology of the Internet is one of the major unresolved problems in Internet research. Topology recovery via tomographic inference is potentially an attractive complement to standard methods that use TTL-limited probes. In this paper, we describe new techniques that aim toward the practical use of tomographic inference for accurate router-level topology measurement. Specifically, prior tomographic techniques have required an infeasible number of probes for accurate, large scale topology recovery. We introduce a Depth-First Search (DFS) Ordering algorithm that clusters end host probe targets based on shared infrastructure, and enables the logical tree topology of the network to be recovered accurately and efficiently. We evaluate the capabilities of our DFS Ordering topology recovery algorithm in simulation and find that our method uses 94% fewer probes than exhaustive methods and 50% fewer than the current state-of-the-art. We also present results from a case study in the live Internet where we show that DFS Ordering can recover the logical router-level topology more accurately and with fewer probes than prior techniques.
Brian Eriksson, Gautam Dasarathy, Paul Barford, Robert D. Nowak
INFOCOM3
2010 A Learning-Based Approach for IP Geolocation
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
PAM2
2010 A Machine Learning Approach to TCP Throughput Prediction
abstract
TCP throughput predictionis an important capability for networks where multiple paths exist between data senders and receivers. In this paper, we describe a new lightweight method for TCP throughput prediction. Our predictor uses Support Vector Regression (SVR); prediction is based on both prior file transfer history and measurements of simple path properties. We evaluate our predictor in a laboratory setting where ground truth can be measured with perfect accuracy. We report the performance of our predictor fororacularandpracticalmeasurements of path properties over a wide range of traffic conditions and transfer sizes. For bulk transfers in heavy traffic usingoracularmeasurements, TCP throughput is predicted within 10% of the actual value 87% of the time, representing nearly a threefold improvement in accuracy over prior history-based methods. Forpracticalmeasurements of path properties, predictions can be made within 10% of the actual value nearly 50% of the time, approximately a 60% improvement over history-based methods, and with much lower measurement traffic overhead. We implement our predictor in a tool calledPathPerf, test it in the wide area, and show thatPathPerfpredicts TCP throughput accurately over diverse wide area paths.
Mariyam Mirza, Joel Sommers, Paul Barford, Xiaojin Zhu 0001
IEEE/ACM Trans. Netw.3
2010 Multiobjective monitoring for SLA compliance
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
IEEE/ACM Trans. Netw.2
2009 An Attacker-Defender Game for Honeynets
Jin-Yi Cai, Vinod Yegneswaran, Chris Alfeld, Paul Barford
COCOON4
2009 Network Performance Anomaly Detection and Localization
abstract
Detecting the occurrence and location of performance anomalies (e.g., high jitter or loss events) is critical to ensuring the effective operation of network infrastructures. In this paper we present a framework for detecting and localizing performance anomalies based on using an active probe-enabled measurement infrastructure deployed on the periphery of a network. Our framework has three components: an algorithm for detecting performance anomalies on a path, an algorithm for selecting which paths to probe at a given time in order to detect performance anomalies (where a path is defined as the set of links between two measurement nodes), and an algorithm for identifying the links that are causing an identified anomaly on a path (i.e., localizing). The problem of detecting an anomaly on a path is addressed by comparing probe-based measures of performance characteristics with performance guarantees for the network (e.g., SLAs). The path selection algorithm is designed to enable a tradeoff between ensuring that all links in a network are frequently monitored to detect performance anomalies, while minimizing probing overhead. The localization algorithm is designed to use existing path measurement data in such a way as to minimize the number of paths necessary for additional probing in order to identify the link(s) responsible for an observed performance anomaly. We assess the feasibility of our framework and algorithms by implementing them in ns-2 and conducting a set of simulation-based experiments using several different network topologies. Our results show that our method is able to accurately detect and localize performance anomalies in a timely fashion and with lower probe and computational overheads than previously proposed methodologies.
Paul Barford, Nick G. Duffield, Amos Ron, Joel Sommers
INFOCOM1
2009 Estimating Hop Distance Between Arbitrary Host Pairs
abstract
Establishing a clear and timely picture of Internet topology is complicated by many factors including the vast size and dynamic nature of the infrastructure. In this paper, we describe a methodology for estimating an important characteristic of Internet topology - the hop distance between arbitrary pairs of end hosts. Our goal is to develop an approach to pairwise hop distance estimation that is accurate, scalable, timely and does not require a significant measurement infrastructure. Our methodology is based on deploying a small set of landmark nodes that use trace route-like probes between each other to establish a set of accurate pairwise hop distances. The landmark nodes are also configured to collect source IP addresses and TTL values from passively monitored network packet traffic. We develop a novel multidimensional scaling algorithm that can be applied to both the passive and active measurements to generate pairwise hop distance estimates for all of the observed source host addresses. The basic algorithm is then enhanced to consider the autonomous system membership of source hosts via BGP routing information. We investigate the capabilities of our estimation algorithms using a set of synthetic network topologies. The results show that our method can generate highly accurate pairwise hop distance estimates over a range of network sizes and configurations, and landmark infrastructure sizes.
Brian Eriksson, Paul Barford, Robert D. Nowak
INFOCOM2
2009 On The Accuracy of TCP Throughput Prediction for Opportunistic Wireless Networks
abstract
The increasing density of WiFi access points (APs) in metropolitan areas is enabling an opportunistic model of wireless networking, whereby a "guest" user within range of one or more wireless APs can gain temporary Internet access through these APs. In this paper, we address the problem of TCP throughput prediction for opportunistic networks. Applications of opportunistic networking can benefit from such predictions by adapting to prevailing network conditions. Our approach is different from prior efforts to model wireless network throughput in that only the two communicating endpoints participate in the prediction, and no information about network topology or traffic loads generated by interfering sources is required. Our goal is to understand how accurate throughput predictions can be under the above assumptions. The physical environment considered in our study includes varying degrees of interference, indoor and outdoor networks, and nodes that are stationary or moving at walking or driving speeds. We use throughput predictors based on time series analysis and machine learning techniques, as they are well-suited to predicting phenomena with unknown variables. The prediction accuracy that our methods yield is cause for cautious optimism. We find that 80% to 100% of predictions are within a factor of two of actual throughput. This bound on accuracy means that predictions are useful for certain applications, because this bound (a) can be achieved by measurements lasting for as little as 0.3 seconds, and (b) holds even when nodes are driving at speeds of 15-25 mph.
Mariyam Mirza, Kevin Springborn, Suman Banerjee 0001, Paul Barford, Michael Blodgett, Xiaojin Zhu 0001
SECON4
2008 Context-aware clustering of DNS query traffic
abstract
The Domain Name System (DNS) is a one of the most widely used services in the Internet. In this paper, we consider the question of how DNS traffic monitoring can provide an important and useful perspective on network traffic in an enterprise. We approach this problem by considering three classes of DNS traffic: canonical (i.e., RFC-intended behaviors), overloaded (e.g.,black-list services), and unwanted (i.e., queries that will never succeed). We describe a context-aware clustering methodology that is applied to DNS query-responses to generate the desired aggregates. Our method enables the analysis to be scaled to expose the desired level of detail of each traffic type, and to expose their time varying characteristics. We implement our method in a tool we call TreeTop, which can be used to analyze and visualize DNS traffic in real-time. We demonstrate the capabilities of our methodology and the utility of TreeTop using a set of DNS traces that we collected from our campus network over a period of three months. Our evaluation highlights both the coarse and fine level of detail that can be revealed by our method. Finally, we show preliminary results on how DNS analysis can be coupled with general network traffic monitoring to provide a useful perspective for network management and operations.
David Plonka, Paul Barford
Internet Measurement Conference2
2008 Power Awareness in Network Design and Routing
abstract
Exponential bandwidth scaling has been a fundamental driver of the growth and popularity of the Internet. However, increases in bandwidth have been accompanied by increases in power consumption, and despite sustained system design efforts to address power demand, significant technological challenges remain that threaten to slow future bandwidth growth. In this paper we describe the power and associated heat management challenges in today's routers. We advocate a broad approach to addressing this problem that includes making power-awareness a primary objective in the design and configuration of networks, and in the design and implementation of network protocols. We support our arguments by providing a case study of power demands of two standard router platforms that enables us to create a generic model for router power consumption. We apply this model in a set of target network configurations and use mixed integer optimization techniques to investigate power consumption, performance and robustness in static network design and in dynamic routing. Our results indicate the potential for significant power savings in operational networks by including power-awareness.
Joseph Chabarek, Joel Sommers, Paul Barford, Cristian Estan, David Tsiang, Stephen J. Wright 0001
INFOCOM3
2008 Spatial-Temporal Characteristics of Internet Malicious Sources
abstract
This paper presents a large scale longitudinal study of the spatial and temporal features of malicious source addresses. The basis of our study is a 402-day trace of over 7 billion Internet intrusion attempts provided by DShield.org, which includes 160 million unique source addresses. Specifically, we focus on spatial distributions and temporal characteristics of malicious sources. First, we find that one out of 27 hosts is potentially a scanning source among 232IPv4 addresses. We then show that malicious sources have a persistent, non-uniform spatial distribution. That is, more than 80% of the sources send packets from the same 20% of the IPv4 address space over time. We also find that 7.3% of malicious source addresses are unroutable, and that some source addresses are correlated. Next, we show that most sources have a short lifetime. 57.9 % of the source addresses appear only once in the trace, and 90% of source addresses appear less than 5 times. These results have implications for both attacks and defenses.
Zesheng Chen 0001, Chuanyi Ji, Paul Barford
INFOCOM3
2008 Network discovery from passive measurements
abstract
Understanding the Internet's structure through empirical measurements is important in the development of new topology generators, new protocols, traffic engineering, and troubleshooting, among other things. While prior studies of Internet topology have been based on active (traceroute-like) measurements, passive measurements of packet traffic offer the possibility of a greatly expanded perspective of Internet structure with much lower impact and management overhead. In this paper we describe a methodology for inferring network structure from passive measurements of IP packet traffic. We describe algorithms that enable 1) traffic sources that share network paths to be clustered accurately without relying on IP address or autonomous system information, 2) topological structure to be inferred accurately with only a small number of active measurements, 3) missing information to be recovered, which is a serious challenge in the use of passive packet measurements. We demonstrate our techniques using a series of simulated topologies and empirical data sets. Our experiments show that the clusters established by our method closely correspond to sources that actually share paths. We also show the trade-offs between selectively applied active probes and the accuracy of the inferred topology between sources. Finally, we characterize the degree to which missing information can be recovered from passive measurements, which further enhances the accuracy of the inferred topologies.
Brian Eriksson, Paul Barford, Robert D. Nowak
SIGCOMM2
2008 A geometric approach to improving active packet loss measurement
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
IEEE/ACM Trans. Netw.2
2007 Learning network structure from passive measurements
abstract
The ability to discover network organization, whether in the form of explicit topology reconstruction or as embeddings that approximate topological distance, is a valuable tool. To date, network discovery has been based on active measurements. However, it is feasible to envision passive discovery of network topology and distance, simply by monitoring packet traffic. Unfortunately, the lack of explicit control over the choices of which endpoints are measured means that passive network discovery must deal with the problem of missing information. We consider one such example, namely reconstructing embeddings and some network structure information from unwanted network traffic captured at a set of honeypots. We develop a number of algorithms for reconstruction of missing measurements. Our algorithms use insights derived from the known topology of the Internet as well as local imputation techniques from approximation theory. We characterize the degree to which missing information can be reconstructed and show that a limited but useful amount of reconstruction is possible, allowing the recovery of network embeddings and some topological relationships from passively collected data.
Brian Eriksson, Paul Barford, Robert D. Nowak, Mark Crovella
Internet Measurement Conference2
2007 An active measurement system for shared environments
abstract
Testbeds composed of end hosts deployed across the Internet enable researchers to simultaneously conduct a wide variety of experiments. Active measurement studies of Internet path properties that require precisely crafted probe streams can be problematic in these environments. The reason is that load on the host systems from concurrently executing experiments (as is typical in PlanetLab) can significantly alter probe stream timings. In this paper we measure and characterize how packet streams from our local PlanetLab nodes are affected by experimental concurrency. We find that the effects can be extreme. We then set up a simple PlanetLab deployment in a laboratory testbed to evaluate these effects in a controlled fashion. We find that even relatively low load levels can cause serious problems in probe streams. Based on these results, we develop a novel system called MAD that can operate as a Linux kernel module or as a stand-alone daemon to support real-time scheduling of probe streams. MAD coordinates probe packet emission for all active measurement experiments on a node. We demonstrate the capabilities of MAD , showing that it performs effectively even under very high levels of multiplexing and host system load.
Joel Sommers, Paul Barford
Internet Measurement Conference2
2007 A Framework for Multi-Objective SLA Compliance Monitoring
abstract
Service level agreements (SLAs) specify performance guarantees made by service providers, typically in terms of packet loss, delay, delay variation, and network availability. While many tools have been developed to measure individual aspects of network performance, there has been little work to directly address the issue of SLA compliance monitoring in an operational setting where accuracy, parsimony, and other related issues are of vital importance. This paper takes the following steps toward addressing this problem: (1) we introduce an architectural framework for integrating multiple discrete-time active measurement algorithms, an architecture that we call multi-objective monitoring; and (2) we introduce a new active measurement methodology to monitor the packet loss rate along a network path for determining compliance with specified performance targets which significantly improves accuracy over existing techniques. We present a prototype implementation of our monitoring framework, and demonstrate how a unified probe stream can consume lower overall bandwidth than if individual streams are used to measure different path properties. We demonstrate the accuracy and convergence properties of our new loss rate monitoring methodology in a controlled laboratory environment using a range of background traffic scenarios and examine its accuracy improvements over existing techniques.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
INFOCOM2
2007 Accurate and efficient SLA compliance monitoring
abstract
Service level agreements (SLAs) define performance guarantees made by service providers, e.g, in terms of packet loss, delay, delay variation, and network availability. In this paper, we describe a new active measurement methodology to accurately monitor whether measured network path characteristics are in compliance with performance targets specified in SLAs. Specifically, (1) we describe a new methodology for estimating packet loss rate that significantly improves accuracy over existing approaches; (2) we introduce a new methodology for measuring mean delay along a path that improves accuracy over existing methodologies, and propose a method for obtaining confidence intervals on quantiles of the empirical delay distribution without making any assumption about the true distribution of delay; (3) we introduce a new methodology for measuring delay variation that is more robust than prior techniques; and (4) we extend existing work in network performance tomography to infer lower bounds on the quantiles of a distribution of performance measures along an unmeasured path given measurements from a subset of paths. We unify active measurements for these metrics in a discrete time-based tool called SLAM. The unified probe stream from SLAM consumes lower overall bandwidth than if individual streams are used to measure path properties. We demonstrate the accuracy and convergence properties of SLAM in a controlled laboratory environment using a range of background traffic scenarios and in one- and two-hop settings, and examine its accuracy improvements over existing standard techniques.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
SIGCOMM2
2007 A machine learning approach to TCP throughput prediction
abstract
TCP throughput prediction is an important capability in wide area overlay and multi-homed networks where multiple paths may exist between data sources and receivers. In this paper we describe a new, lightweight method for TCP throughput prediction that can generate accurate forecasts for a broad range of file sizes and path conditions. Our method is based on Support Vector Regression modeling that uses a combination of prior file transfers and measurements of simple path properties. We calibrate and evaluate the capabilities of our throughput predictor in an extensive set of lab-based experiments where ground truth can be established for path properties using highly accurate passive measurements. We report the performance for our method in the ideal case of using our passive path property measurements over a range of test configurations. Our results show that for bulk transfers in heavy traffic, TCP throughput is predicted within 10% of the actual value 87% of the time, representing nearly a 3-fold improvement in accuracy over prior history-based methods. In the same lab environment, we assess our method using less accurate active probe measurements of path properties, and show that predictions can be made within 10% of the actual value nearly 50% of the time over a range of file sizes and traffic conditions. This result represents approximately a 60% improvement over history-based methods with a much lower impact on end-to-end paths. Finally, we implement our predictor in a tool called PathPerf and test it in experiments conducted on wide area paths. The results demonstrate that PathPerf predicts TCP through put accurately over a variety of paths.
Mariyam Mirza, Joel Sommers, Paul Barford, Xiaojin Zhu 0001
SIGMETRICS3
2006 Improving NIDS Performance Through Hardware-based Connection Filtering
abstract
Traffic volume and diversity can have a significant impact on the ability of network intrusion detection systems (NIDS) to report malicious activity accurately. Based on the observation that a great deal of traffic is, in fact, not important to accurate attack identification, we investigate connection filtering as a method for improving the performance of NIDS. We describe three different classes of connection filters that were developed to explore the design space and trade off's in load reduction versus alarm rates. We implement instances of each filter class on a network processor that can be used with any NIDS that runs on commodity hardware, and evaluate the impact of each filter in a series of laboratory-based tests. First, we establish an idealized maximum performance by using static connection filters for all benign traffic. Next, we show that volume sensitive random connection filters can improve performance significantly with respect to alarm rates under heavy traffic load. Finally, we show that dynamic connection filters that attempt to infer benign traffic can improve performance almost to the level of idealized static filters. These results underscore the potential for hardware-based connection filtering as an effective means for improving the performance of NIDS.
Vikas Garg 0002, Vinod Yegneswaran, Paul Barford
ICC3
2006 On the Performance of Round Trip Time Network Tomography
abstract
Network tomography is an appealing method for active measurement of link level characteristics such as delay and loss on end-to-end paths. Most network tomography techniques developed to date are based on one-way measurements requiring collaboration from both sending and receiving hosts which severely limits the scope of the paths over which these techniques can be used. We extend our previous work on Network Radar, a new tomographic inference method based on round trip time (RTT) measurements from TCP SYN/SYN-ACK packets. In this paper, our contributions are three-folded. (1) We extend our analytic framework for estimating delay variance on the shared network segment using Network Radar to include confidence estimates which enable measurement accuracy to be assessed - an important consideration for practical deployment. (2) We evaluate Network Radar in a series of experiments conducted in a controlled laboratory environment. These tests explore the boundaries of effectiveness of our RTT-based method, and show that it works well over a wide range of traffic conditions. (3) We evaluate Network Radar in a series of tests conducted in the wide area Internet. These tests show that RTT-based delay variance estimates can be used effectively to identify most likely network topology - a natural and verifyable application for RTT tomography. The performance results in this paper demonstrate that Network Radar can now be used for both research and operational purposes.
Yolanda Tsang, Mehmet Can Yildiz, Paul Barford, Robert D. Nowak
ICC3
2006 A Proposed Framework for Calibration of Available Bandwidth Estimation Tools
abstract
Examining the validity or accuracy of proposed available bandwidth estimation tools remains a challenging problem. A common approach consists of evaluating a newly developed tool using a combination of simple nstype simulations and feasible experiments in situ (i.e., using parts of the actual Internet). In this paper, we argue that this strategy tends to fall short of establishing a reliable "ground truth," and we advocate an alternative in vitro-like methodology for calibrating available bandwidth estimation tools that has not been widely used in this context. Our approach relies on performing controlled laboratory experiments and using tools to visualize and analyze the relevant tool-specific traffic dynamics. We present a case study of how two canonical available bandwidth estimation tools, SPRUCE and PATHLOAD, respond to increasingly more complex cross traffic and network path conditions. We expose measurement bias and algorithmic omissions that lead to poor tool calibration. As a result of this evaluation, we designed a calibrated available bandwidth estimation tool called YAZ that builds on the insights of PATHLOAD. We show that in head to head comparisons with SPRUCE and PATHLOAD, YAZ is significantly and consistently more accurate with respect to ground truth, and reports results more quickly with a small number of probes.
Joel Sommers, Paul Barford, Walter Willinger
ISCC2
2006 Composite Subset Measures
Lei Chen 0003, Raghu Ramakrishnan 0001, Paul Barford, Bee-Chung Chen, Vinod Yegneswaran
VLDB3
2005 Scalable Network Path Emulation
abstract
Laboratory-based experimentation is an increasingly popular method for conducting network research since it enables implementations of network systems and protocols to be evaluated. Most research conducted in lab-based environments requires the faithful reproduction of wide area network conditions. An important step toward satisfying this requirement is the creation of paths between nodes in the lab that have the same characteristics as paths between nodes in the Internet. In this paper, we describe and evaluate a new, highly scalable, software-based path emulation tool called NetPath. We describe the design and implementation of NetPath, which features fixed and probabilistic packet propagation delay emulation, probabilistic bit errors, probabilistic packet loss, packet duplication, and packet reordering capability. Through a series of controlled laboratory experiments, we demonstrate that Net-Path offers over three times the loss-free throughput capacity of other popular software-based path/network emulators. We show that under moderate load NetPath's mean propagation delay emulation precision is within 1% of a hardware-based reference emulator. This result represents a significant improvement over other software-based emulators. We illustrate how, relative to our hardware-based reference, NetPath improves application traffic behavior over other software-based emulators. Finally, we demonstrate and characterize NetPath's ability to provide path emulation simultaneously on multiple physical links. This capability, which is facilitated through the use of our link configuration tool, enables laboratory system resources to be more efficiently utilized.
Shilpi Agarwal, Joel Sommers, Paul Barford
MASCOTS3
2005 Improving accuracy in end-to-end packet loss measurement
abstract
Measurement and estimation of packet loss characteristics are challenging due to the relatively rare occurrence and typically short duration of packet loss episodes. While active probe tools are commonly used to measure packet loss on end-to-end paths, there has been little analysis of the accuracy of these tools or their impact on the network. The objective of our study is to understand how to measure packet loss episodes accurately with end-to-end probes. We begin by testing the capability of standard Poisson-modulated end-to-end measurements of loss in a controlled laboratory environment using IP routers and commodity end hosts. Our tests show that loss characteristics reported from such Poisson-modulated probe tools can be quite inaccurate over a range of traffic conditions. Motivated by these observations, we introduce a new algorithm for packet loss measurement that is designed to overcome the deficiencies in standard Poisson-based tools. Specifically, our method creates a probe process that (1) enables an explicit trade-off between accuracy and impact on the network, and (2) enables more accurate measurements than standard Poisson probing at the same rate. We evaluate the capabilities of our methodology experimentally by developing and implementing a prototype tool, called BADABING. The experiments demonstrate the trade-offs between impact on the network and measurement accuracy. We show that BADABING reports loss characteristics far more accurately than traditional loss measurement tools.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
SIGCOMM2
2005 An Architecture for Generating Semantic Aware Signatures
Vinod Yegneswaran, Jonathon T. Giffin, Paul Barford, Somesh Jha
USENIX Security Symposium3
2004 Characteristics of internet background radiation
abstract
Monitoring any portion of the Internet address space reveals incessant activity. This holds even when monitoring traffic sent to unused addresses, which we term "background radiation. " Background radiation reflects fundamentally nonproductive traffic, either malicious (flooding backscatter, scans for vulnerabilities, worms) or benign (misconfigurations). While the general presence of background radiation is well known to the network operator community, its nature has yet to be broadly characterized. We develop such a characterization based on data collected from four unused networks in the Internet. Two key elements of our methodology are (i) the use of filtering to reduce load on the measurement system, and (ii) the use of active responders to elicit further activity from scanners in order to differentiate different types of background radiation. We break down the components of background radiation by protocol, application, and often specific exploit; analyze temporal patterns and correlated activity; and assess variations across different networks and over time. While we find a menagerie of activity, probes from worms and autorooters heavily dominate. We conclude with considerations of how to incorporate our characterizations into monitoring and detection activities.
Ruoming Pang, Vinod Yegneswaran, Paul Barford, Vern Paxson, Larry L. Peterson
Internet Measurement Conference3
2004 Self-configuring network traffic generation
abstract
The ability to generate repeatable, realistic network traffic is critical in both simulation and testbed environments. Traffic generation capabilities to date have been limited to either simple sequenced packet streams typically aimed at throughput testing, or to application-specific tools focused on, for example, recreating representative HTTP requests. In this paper we describe Harpoon, a new application-independent tool for generating representative packet traffic at the IP flow level. Harpoon generates TCP and UDP packet flows that have the same byte, packet, temporal and spatial characteristics as measured at routers in live environments. Harpoon is distinguished from other tools that generate statistically representative traffic in that it can self-configure by automatically extracting parameters from standard Netflow logs or packet traces. We provide details on Harpoon's architecture and implementation, and validate its capabilities in controlled laboratory experiments using configurations derived from flow and packet traces gathered in live environments. We then demonstrate Harpoon's capabilities in a router benchmarking experiment that compares Harpoon with commonly used throughput test methods. Our results show that the router subsystem load generated by Harpoon is significantly different, suggesting that this kind of test can provide important insights into how routers might behave under actual operating conditions.
Joel Sommers, Paul Barford
Internet Measurement Conference2
2004 A framework for malicious workload generation
abstract
Malicious traffic from self-propagating worms and denial-of-service attacks constantly threatens the everyday operation of Internet systems. Defending networks from these threats demands appropriate tools to conduct comprehensive vulnerability assessments of networked systems. This paper describes MACE, a unique environment for recreating a wide range of malicious packet traffic in laboratory testbeds. MACE defines a model for flexible composition of malicious traffic that enables both known attacks (such as the Welchia worm) and new attack variants to be created. We implement this model in an extensible library for attack traffic specification and generation. To demonstrate the capability of MACE, we provide an analysis of stress tests conducted on a popular firewall and two popular network intrusion detection systems. Our results expose potential weaknesses of these systems and reveal that modern firewalls and network intrusion detection systems could be easily overwhelmed by simple attacks launched from a small number of hosts.
Joel Sommers, Vinod Yegneswaran, Paul Barford
Internet Measurement Conference3
2004 Network radar: tomography from round trip time measurements
abstract
Knowledge of link specific traffic characteristics is important in the operation and design of wide area networks. Network tomography is a powerful method for measuring characteristics such as delay and loss on network-internal links using end--to--end active probes. Prior work has established the basic mechanisms for the use of tomographic inference techniques in the networking context. However, the measurement methods described in prior network tomography studies require cooperation between sending and receiving end-hosts, which limits the scope of the paths over which the measurements can be made. In this paper, we describe a new network tomographic technique based on round trip time (RTT) measurements which eliminates the need for special-purpose cooperation from receivers. Our technique uses RTT measurements from TCP SYN and SYN-ACK segments to estimate the delay variance of the shared network segment in the standard one sender - two receivers configuration. We call this approach Network Radar since it is analogous to standard radar. We present an analytic evaluation of Network Radar that specifies the variance bounds within which the technique is effective. We also evaluate Network Radar in a series of tests conducted in a controlled laboratory environment using live end hosts and IP routers. These tests demonstrate the boundaries of effectiveness of the RTT-based approach.
Yolanda Tsang, Mehmet Can Yildiz, Paul Barford, Robert D. Nowak
Internet Measurement Conference3
2004 Global Intrusion Detection in the DOMINO Overlay System
Vinod Yegneswaran, Paul Barford, Somesh Jha
NDSS2
2004 On the Design and Use of Internet Sinks for Network Abuse Monitoring
Vinod Yegneswaran, Paul Barford, David Plonka
RAID2
2004 Harpoon: a flow-level traffic generator for router and network tests
abstract
We describe Harpoon, a new application-independent tool for generating representative packet traffic at the IP flow level. Harpoon is a configurable tool for creating TCP and UDP packet flows that have the same byte, packet, temporal, and spatial characteristics as measured at routers in live environments. We validate Harpoon using traces collected from a live router and then demonstrate its capabilities in a series of router performance benchmark tests.
Joel Sommers, Hyungsuk Kim, Paul Barford
SIGMETRICS3
2004 Representing the Internet as a succinct forest
Jim Gast, Paul Barford
Comput. Networks2
2003 Internet intrusions: global characteristics and prevalence
abstract
Network intrusions have been a fact of life in the Internet for many years. However, as is the case with many other types of Internet-wide phenomena, gaining insight into the global characteristics of intrusions is challenging. In this paper we address this problem by systematically analyzing a set of firewall logs collected over four months from over 1600 different networks world wide. The first part of our study is a general analysis focused on the issues of distribution, categorization and prevalence of intrusions. Our data shows both a large quantity and wide variety of intrusion attempts on a daily basis. We also find that worms like CodeRed, Nimda and SQL Snake persist long after their original release. By projecting intrusion activity as seen in our data sets to the entire Internet we determine that there are typically on the order of 25B intrusion attempts per day and that there is an increasing trend over our measurement period. We further find that sources of intrusions are uniformly spread across the Autonomous System space. However, deeper investigation reveals that a very small collection of sources are responsible for a significant fraction of intrusion attempts in any given month and their on/off patterns exhibit cliques of correlated behavior. We show that the distribution of source IP addresses of the non-worm intrusions as a function of the number of attempts follows Zipf's law. We also find that at daily timescales, intrusion targets often depict significant spatial trends that blur patterns observed from individual "IP telescopes"; this underscores the necessity for a more global approach to intrusion detection. Finally, we investigate the benefits of shared information, and the potential for using this as a foundation for an automated, global intrusion detection framework that would identify and isolate intrusions with greater precision and robustness than systems with limited perspective.
Vinod Yegneswaran, Paul Barford, Johannes Ullrich
SIGMETRICS2
2002 Resource deployment based on autonomous system clustering
abstract
Effective placement of resources used to support distributed services in the Internet depends on an accurate representation of Internet topology and routing. Representations of autonomous system (AS) level topology derived solely from BGP tables show only a subset of the connections that actually get used. However, in many cases, missing connections can be discovered by simple traceroutes. In addition, the differences between customer-to-provider links, peer-to-peer links, and sibling-to-sibling links are useful distinctions for the resource placement problem which is the focus of our work. Using two complementary mechanisms, we improve the accuracy of an AS forest as a predictor of packet paths. One mechanism uses recent insights that packets flow unidirectionally across customer-provider inter-AS links. Annotations are added to the AS forest to indicate links that appear to be peering versus those that appear to be customer-provider links. The other mechanism provides links between trees by remembering the most recently seen similar traceroute. The paper concludes by applying the annotated AS forest to a problem in resource placement to show that the representation is amenable to computationally inexpensive analysis.
Jim Gast, Paul Barford
GLOBECOM2
2002 A signal analysis of network traffic anomalies
abstract
Identifying anomalies rapidly and accurately is critical to the efficient operation of large computer networks. Accurately characterizing important classes of anomalies greatly facilitates their identification; however, the subtleties and complexities of anomalous traffic can easily confound this process. In this paper we report results of signal analysis of four classes of network traffic anomalies: outages, flash crowds, attacks and measurement failures. Data for this study consists of IP flow and SNMP measurements collected over a six month period at the border router of a large university. Our results show that wavelet filters are quite effective at exposing the details of both ambient and anomalous traffic. Specifically, we show that a pseudo-spline filter tuned at specific aggregation levels will expose distinct characteristics of each class of anomaly. We show that an effective way of exposing anomalies is via the detection of a sharp increase in the local variance of the filtered data. We evaluate traffic anomaly signals at different points within a network based on topological distance from the anomaly source or destination. We show that anomalies can be exposed effectively even when aggregated with a large amount of additional traffic. We also compare the difference between the same traffic anomaly signals as seen in SNMP and IP flow data, and show that the more coarse-grained SNMP data can also be used to expose anomalies effectively.
Paul Barford, Jeffery Kline, David Plonka, Amos Ron
Internet Measurement Workshop1
2001 Critical path analysis of TCP transactions
abstract
Improving the performance of data transfers in the Internet (such as Web transfers) requires a detailed understanding of when and how delays are introduced. Unfortunately, the complexity of data transfers like those using HTTP is great enough that identifying the precise causes of delays is difficult. We describe a method for pinpointing where delays are introduced into applications like HTTP by using critical path analysis. By constructing and profiling the critical path, it is possible to determine what fraction of total transfer latency is due to packet propagation, network variation (e.g., queueing at routers or route fluctuation), packet losses, and delays at the server and at the client. We have implemented our technique in a tool called tcpeval that automates critical path analysis for Web transactions. We show that our analysis method is robust enough to analyze traces taken for two different TCP implementations (Linux and FreeBSD). To demonstrate the utility of our approach, we present the results of critical path analysis for a set of Web transactions taken over 14 days under a variety of server and network conditions. The results show that critical path analysis can shed considerable light on the causes of delays in Web transfers, and can expose subtleties in the behavior of the entire end-to-end system.
Paul Barford, Mark Crovella
IEEE/ACM Trans. Netw.1
2000 Critical path analysis of TCP transactions
abstract
Improving the performance of data transfers in the Internet (such as Web transfers) requires a detailed understanding of when and how delays are introduced. Unfortunately, the complexity of data transfers like those using HTTP is great enough that identifying the precise causes of delays is difficult. In this paper we describe a method for pinpointing where delays are introduced into applications like HTTP by using critical path analysis. By constructing and profiling the critical path, it is possible to determine what fraction of total transfer latency is due to packet propagation, network variation (e.g., queuing at routers or route fluctuation), packet losses, and delays at the server and at the client. We have implemented our technique in a tool called tcpeval that automates critical path analysis for Web transactions. We show that our analysis method is robust enough to analyze traces taken for two different TCP implementations (Linux and FreeBSD). To demonstrate the utility of our approach, we present the results of critical path analysis for a set of Web transactions taken over 14 days under a variety of server and network conditions. The results show that critical path analysis can shed considerable light on the causes of delays in Web transfers, and can expose subtleties in the behavior of the entire end-to-end system.
Paul Barford, Mark Crovella
SIGCOMM1
1999 A Performance Evaluation of Hyper Text Transfer Protocols
abstract
Version 1.1 of the Hyper Text Transfer Protocol (HTTP) was principally developed as a means for reducing both document transfer latency and network traffic.The rationale for the performance enhancements in HTTP/1.1 is based on the assumption that the network is the bottleneck in Web transactions.In practice, however, the Web server can be the primary source of document transfer latency.In this paper, we characterize and compare the performance of HTTP/1.0 and HTTP/1.1 in terms of throughput at the server and transfer latency at the client.We examine how bottlenecks in the network, CPU, and in the disk system affect the relative performance of HTTP/1.0 versus HTTP/1.1.We show that the network demands under HTTP/1.1 are somewhat lower than HTTP/1.0,and we quantify those differences in terms of packets transferred, server congestion window size and data bytes per packet.We show that when the CPU is the bottleneck, there is relatively little difference in performance between HTTP/1.0 and HTTP/1.1.Surprisingly, we show that when the disk system is the bottleneck, performance using HTTP/1.1 can be much worse than with HTTP/1.0.Based on these observations, we suggest a connection management policy for HTTP/1.1 that can improve throughput, decrease latency, and keep network trafllc low when the disk system is the bottleneck. Supported in part by
Paul Barford, Mark Crovella
SIGMETRICS1
1999 Changes in Web Client Access Patterns: Characteristics and Caching Implications
Paul Barford, Azer Bestavros, Adam D. Bradley, Mark Crovella
World Wide Web1
1998 The Network Effects of Prefetching
abstract
Prefetching has been shown to be an effective technique for reducing user perceived latency in distributed systems. In this paper we show that even when prefetching adds no extra traffic to the network, it can have serious negative performance effects. Straightforward approaches to prefetching increase the burstiness of individual sources, leading to increased average queue sizes in network switches. However, we also show that applications can avoid the undesirable queueing effects of prefetching. In fact, we show that applications employing prefetching can significantly improve network performance, to a level much better than that obtained without any prefetching at all. This is because prefetching offers increased opportunities for traffic shaping that are not available in the absence of prefetching. Using a simple transport rate control mechanism, a prefetching application can modify its behavior from a distinctly ON/OFF entity to one whose data transfer rate changes less abruptly, while still delivering all data in advance of the user's actual requests.
Mark Crovella, Paul Barford
INFOCOM2
1998 Generating Representative Web Workloads for Network and Server Performance Evaluation
abstract
One role for workload generation is as a means for understanding how servers and networks respond to variation in load. This enables management and capacity planning based on current and projected usage. This paper applies a number of observations of Web server usage to create a realistic Web workload generation tool which mimics a set of real users accessing a server. The tool, called Surge (Scalable URL Reference Generator) generates references matching empirical measurements of 1) server file size distribution; 2) request size distribution; 3) relative file popularity; 4) embedded file references; 5) temporal locality of reference; and 6) idle periods of individual users. This paper reviews the essential elements required in the generation of a representative Web workload. It also addresses the technical challenges to satisfying this large set of simultaneous constraints on the properties of the reference stream, the solutions we adopted, and their associated accuracy. Finally, we present evidence that Surge exercises servers in a manner significantly different from other Web server benchmarks.
Paul Barford, Mark Crovella
SIGMETRICS1
1989 A system simulation environment within Digital
abstract
The authors discuss the use of simulation as an integral part of the product development process at the Digital Equipment Corporation. One engineering computer-aided-design group within Digital provides a proven process and tool suite which was used to simulate such products as the MicroVAX II, MicroVAX 3500, VT320, and others. This process includes model libraries and specific design verification techniques.>
Quinn Canfield, Paul Barford, Paul Kinzelman, Cary Trlica
ICCD2