Joel Sommers

dblp:48/3919 · DBLP profile ↗
← Back
39ranked-venue papers
20as first author
3since 2021 · last 2025
0000-0003-4872-6532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 29 · 13 first-author · 3 since 2021Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorSecurity and privacy · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
27 papers
Network measurement and analytics · 51% Internet architecture and protocols · 15% Network management and operations · 8%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Energy-efficient computing · 79% Performance modeling and evaluation · 21%

Topics — the 30 heaviest of 58, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Network measurement and analytics
internet topology measurement
0.922024
An Elemental Decomposition of DNS Name-to-IP Graphs · INFOCOM 2024
Layer 1-informed Internet Topology Measurement · Internet Measurement Conference 2014
Internet architecture and protocols
domain name system
0.812024
An Elemental Decomposition of DNS Name-to-IP Graphs · INFOCOM 2024
Network measurement and analytics
active measurement
0.432017
Automatic metadata generation for active measurement · Internet Measurement Conference 2017
A geometric approach to improving active packet loss measurement · IEEE/ACM Trans. Netw. 2008
An active measurement system for shared environments · Internet Measurement Conference 2007
Internet architecture and protocols › network topology
internet topology
0.322015
InterTubes: A Study of the US Long-haul Fiber-optic Infrastructure · SIGCOMM 2015
DomainImpute: Inferring unseen components in the Internet · INFOCOM 2011
Internet of things and sensor networks
crowdsensing
0.322012
Comparing metro-area cellular and WiFi performance: extended abstract · SIGMETRICS 2012
Revisiting broadband performance · Internet Measurement Conference 2012
Internet of things and sensor networks
time synchronization
0.212016
MNTP: Enhancing Time Synchronization for Mobile Devices · Internet Measurement Conference 2016
Network measurement and analytics › network tomography
topology inference
0.222011
Inferring Unseen Components of the Internet Core · IEEE J. Sel. Areas Commun. 2011
DomainImpute: Inferring unseen components in the Internet · INFOCOM 2011
Network measurement and analytics › active measurement
active probing
0.222011
Inferring Unseen Components of the Internet Core · IEEE J. Sel. Areas Commun. 2011
Network Performance Anomaly Detection and Localization · INFOCOM 2009
Cellular and mobile networks
infrastructure sharing
0.212015
InterTubes: A Study of the US Long-haul Fiber-optic Infrastructure · SIGCOMM 2015
Network management and operations
configuration verification
0.212014
Controller-agnostic SDN Debugging · CoNEXT 2014
Network measurement and analytics › topology discovery
router-level topology discovery
0.212014
Layer 1-informed Internet Topology Measurement · Internet Measurement Conference 2014
Network management and operations › network monitoring › performance monitoring
service level agreement monitoring
0.222010
Multiobjective monitoring for SLA compliance · IEEE/ACM Trans. Netw. 2010
Accurate and efficient SLA compliance monitoring · SIGCOMM 2007
Transport protocols and congestion control › TCP modeling
TCP throughput prediction
0.222010
A Machine Learning Approach to TCP Throughput Prediction · IEEE/ACM Trans. Netw. 2010
A machine learning approach to TCP throughput prediction · SIGMETRICS 2007
Network measurement and analytics
network performance measurement
0.222012
Comparing metro-area cellular and WiFi performance: extended abstract · SIGMETRICS 2012
A Framework for Multi-Objective SLA Compliance Monitoring · INFOCOM 2007
Cellular and mobile networks
cellular network performance
0.112012
Comparing metro-area cellular and WiFi performance: extended abstract · SIGMETRICS 2012
Network measurement and analytics
mobile network measurement
0.112012
Cell vs. WiFi: on the performance of metro area mobile connections · Internet Measurement Conference 2012
Wireless networking › WLAN
wifi performance
0.112012
Comparing metro-area cellular and WiFi performance: extended abstract · SIGMETRICS 2012
Wireless networking
WLAN
0.112012
Comparing metro-area cellular and WiFi performance: extended abstract · SIGMETRICS 2012
Network measurement and analytics › network performance measurement
packet loss measurement
0.122008
A geometric approach to improving active packet loss measurement · IEEE/ACM Trans. Netw. 2008
Improving accuracy in end-to-end packet loss measurement · SIGCOMM 2005
Network management and operations › network monitoring
adaptive probing
0.112011
Inferring Unseen Components of the Internet Core · IEEE J. Sel. Areas Commun. 2011
Network measurement and analytics
internet topology mapping
0.112011
Inferring Unseen Components of the Internet Core · IEEE J. Sel. Areas Commun. 2011
Network measurement and analytics › internet measurement
routing measurement
0.112011
On the prevalence and characteristics of MPLS deployments in the open internet · Internet Measurement Conference 2011
Network measurement and analytics
anomaly detection
0.112010
BasisDetect: a model-based network event detection framework · Internet Measurement Conference 2010
Network management and operations › fault management › fault diagnosis
anomaly localization
0.112009
Network Performance Anomaly Detection and Localization · INFOCOM 2009
Network measurement and analytics › anomaly detection
network performance anomaly detection
0.112009
Network Performance Anomaly Detection and Localization · INFOCOM 2009
Network measurement and analytics
traffic generation
0.122004
Harpoon: a flow-level traffic generator for router and network tests · SIGMETRICS 2004
Self-configuring network traffic generation · Internet Measurement Conference 2004
Internet architecture and protocols › world wide web › web protocols
HTTP
0.112017
On the structure and characteristics of user agent string · Internet Measurement Conference 2017
Network measurement and analytics › network measurement methodology
measurement reproducibility
0.112017
Automatic metadata generation for active measurement · Internet Measurement Conference 2017
Routing and switching
energy-aware routing
0.112008
Power Awareness in Network Design and Routing · INFOCOM 2008
Energy-efficient computing › power management › energy-efficient networking
network power management
0.112008
Power Awareness in Network Design and Routing · INFOCOM 2008

Methods — techniques the papers use, named apart from their topics

large-scale census · 0.8biclique decomposition · 0.8simulation · 0.6traceroute · 0.3longitudinal analysis · 0.3local environment monitoring · 0.3large-scale log analysis · 0.3crowd-sourced data analysis · 0.3network time protocol · 0.2constraint-based inference · 0.2mixed-integer optimization · 0.1real-time scheduling · 0.1kernel modules · 0.1workload modeling · 0.0stress testing · 0.0
YearPublicationVenuePosition
2025 Squatspotting: Towards the Systematic Measurement of Typosquatting Techniques
Wei-Shiang Wung, Calvin Kranig, Eric Pauley, Paul Barford, Mark Crovella, Joel Sommers
Networking6
2024 An Elemental Decomposition of DNS Name-to-IP Graphs
abstract
The Domain Name System (DNS) is a critical piece of Internet infrastructure with remarkably complex properties and uses, and accordingly has been extensively studied. In this study we contribute to that body of work by organizing and analyzing records maintained within the DNS as a bipartite graph. We find that relating names and addresses in this way uncovers a surprisingly rich structure. In order to characterize that structure, we introduce a new graph decomposition for DNS name-to-IP mappings, which we term elemental decomposition. In particular, we argue that (approximately) decomposing this graph into bicliques — maximally connected components — exposes this rich structure. We utilize large-scale censuses of the DNS to investigate the characteristics of the resulting decomposition, and illustrate how the exposed structure sheds new light on a number of questions about how the DNS is used in practice and suggests several new directions for future research.
Alex Anderson, Aadi Swadipto Mondal, Paul Barford, Mark Crovella, Joel Sommers
INFOCOM5
2022 BigBen: Telemetry Processing for Internet-Wide Event Monitoring
abstract
This paper describes BigBen, a network telemetry processing system designed to enable accurate and timely reporting of Internet events (e.g., outages, attacks and configuration changes). BigBen is distinct from other Internet-wide event detection systems in its use of passive measurements of Network Time Protocol (NTP) traffic. We describe the architecture of BigBen, and a cloud-based implementation developed to process large NTP data sets and provide accurate daily event reporting. We demonstrate BigBen on a 15.5TB corpus of NTP data. We show that BigBen identifies a wide range of Internet events characterized by their location, scope and duration. We compare the events detected by BigBen vs. events detected by a large active probe-based detection system. We find only modest overlap between the two datasets and show how BigBen provides details on events that are not available from active measurements. Finally, we report on the perspective that BigBen provides on Internet events that were reported by third parties. In each case, BigBen confirms the event and provides details that were not available in prior reports, highlighting the utility of the passive, NTP-based approach.
Meenakshi Syamkumar, Yugali Gullapalli, Paul Barford, Joel Sommers
IEEE Trans. Netw. Serv. Manag.5
2020 A Web Client Perspective on IP Geolocation Accuracy
abstract
Geolocation of Internet addresses is widely used by content providers to tailor services and content to users and to restrict access to content. IP geolocation in practice typically relies on databases that include geographic information about individual addresses or address prefixes. Prior studies have assessed the accuracy of these databases by using a set of addresses with known or estimated locations and comparing them with database-reported locations.In this paper we investigate IP geolocation accuracy from the standpoint of web clients, by exploiting geolocation information embedded in non-standard HTTP response headers and in unencrypted HTTP cookies. We identify a set of 10,476 websites and content providers that include non-standard HTTP headers and unencrypted cookies with geographic information. We launch HTTP requests to each of these sites from 113 client vantage points with known locations distributed across 6 continents and 60 countries and extract available geographic information from the responses using a battery of hand-crafted regular expressions. We find that the country of the client is included in more than 90% of all responses. Moreover, we observe that about 75% of all responses only include the country name or code and that the remaining responses include some combination of geographic information, such as continent, country, city, postcode, region, and coordinates. We observe that accuracy is greatest for the coarsest geographic scope (continent) and least accurate for finer scopes (e.g., coordinate), but that accuracy varies widely across vantage points regardless of the continent or country from which the request is launched.
Joel Sommers
ISNCC1
2018 On the Characteristics of Language Tags on the Web
Joel Sommers
PAM1
2017 On the structure and characteristics of user agent string
abstract
User agent (UA) strings transmitted during HTTP transactions convey client system configuration details to ensure that content returned by a server is appropriate for the requesting host. As such, analysis of UA strings and their structure offers a unique perspective on active client systems in the Internet and when tracked longitudinally, offers a perspective on the nature of system and configuration dynamics. In this paper, we describe our study of UA string characteristics. Our work is based on analyzing a unique corpus of over 1B UA strings collected over a period of 2 years by comScore. We begin by analyzing the general characteristics of UA strings, focusing on the most prevalent strings and dynamic behaviors. We identify the top 10 most popular User Agents, which account for 26% of total daily volume. These strings describe the expected instances of popular platforms such as Microsoft, Apple and Google. We then report on the characteristics of low-volume UA strings, which has important implications for unique device identification. We show that this class of user agent generates the overwhelming majority of traffic, with between 2M and 10M instances observed each day. We show that the distribution of UA strings has temporal dependence and we show the distribution measured depends on the type of content served. Finally, we report on two large-scale UA anomalies characterized by web browsers sending false and misleading UAs in their web requests.
Jeffery Kline, Paul Barford, Aaron Cahn, Joel Sommers
Internet Measurement Conference4
2017 Automatic metadata generation for active measurement
abstract
Empirical research in the Internet is fraught with challenges. Among these is the possibility that local environmental conditions (e.g., CPU load or network load) introduce unexpected bias or artifacts in measurements that lead to erroneous conclusions. In this paper, we describe a framework for local environment monitoring that is designed to be used during Internet measurement experiments. The goals of our work are to provide a critical, expanded perspective on measurement results and to improve the opportunity for reproducibility of results. We instantiate our framework in a tool we call SoMeta, which monitors the local environment during active probe-based measurement experiments. We evaluate the runtime costs of SoMeta and conduct a series of experiments in which we intentionally perturb different aspects of the local environment during active probe-based measurements. Our experiments show how simple local monitoring can readily expose conditions that bias active probe-based measurement results. We conclude with a discussion of how our framework can be expanded to provide metadata for a broad range of Internet measurement experiments.
Joel Sommers, Ramakrishnan Durairajan, Paul Barford
Internet Measurement Conference1
2016 MNTP: Enhancing Time Synchronization for Mobile Devices
Sathiya Kumaran Mani, Ramakrishnan Durairajan, Paul Barford, Joel Sommers
Internet Measurement Conference4
2015 Time's Forgotten: Using NTP to understand Internet Latency
abstract
The performance of Internet services is intrinsically tied to propagation delays between end points (i.e., network latency). Standard active probe-based or passive host-based methods for measuring end-to-end latency are difficult to deploy at scale and typically offer limited precision and accuracy. In this paper, we investigate a novel but non-obvious source of latency measurement---logs from network time protocol (NTP) servers. Using NTP-derived data for studying latency is compelling due to NTP's pervasive use in the Internet and its inherent focus on accurate end-to-end delay estimation. We consider the efficacy of an NTP-based approach for studying propagation delays by analyzing logs collected from 10 NTP servers distributed across the United States. These logs include over 73M latency measurements to 7.4M worldwide clients (as indicated by unique IP addresses) collected over the period of one day. Our initial analysis of the general characteristics of propagation delays derived from the log data reveals that delay measurements from NTP must be carefully filtered in order to extract accurate results. We develop a filtering process that removes measurements that are likely to be inaccurate. After applying our filter to NTP measurements, we report on the scope and reach for US-based clients and the characteristics of the end-to-end latency for those clients.
Ramakrishnan Durairajan, Sathiya Kumaran Mani, Joel Sommers, Paul Barford
HotNets3
2015 InterTubes: A Study of the US Long-haul Fiber-optic Infrastructure
abstract
The complexity and enormous costs of installing new long-haul fiber-optic infrastructure has led to a significant amount of infrastructure sharing in previously installed conduits. In this paper, we study the characteristics and implications of infrastructure sharing by analyzing the long-haul fiber-optic network in the US. We start by using fiber maps provided by tier-1 ISPs and major cable providers to construct a map of the long-haul US fiber-optic infrastructure. We also rely on previously under-utilized data sources in the form of public records from federal, state, and municipal agencies to improve the fidelity of our map. We quantify the resulting map's connectivity characteristics and confirm a clear correspondence between long-haul fiber-optic, roadway, and railway infrastructures. Next, we examine the prevalence of high-risk links by mapping end-to-end paths resulting from large-scale traceroute campaigns onto our fiber-optic infrastructure map. We show how both risk and latency (i.e., propagation delay) can be reduced by deploying new links along previously unused transportation corridors and rights-of-way. In particular, focusing on a subset of high-risk links is sufficient to improve the overall robustness of the network to failures. Finally, we discuss the implications of our findings on issues related to performance, net neutrality, and policy decision-making.
Ramakrishnan Durairajan, Paul Barford, Joel Sommers, Walter Willinger
SIGCOMM3
2015 Lowering the Barrier to Systems-level Networking Projects
abstract
Developing systems-level networking software to implement switches, routers, and middleboxes is challenging, rewarding, and arguably an essential component for developing a deep understanding of modern computer networks. Unfortunately, existing techniques for building networked system software use low-level and error-prone tools and languages, making this task inaccessible for many undergraduates. Moreover, working at such a low-level of abstraction complicates debugging and testing and can make assessment difficult for instructors and TAs. We describe a Python-based environment called Switchyard that is designed to facilitate student projects for building and testing software-based network devices like switches, routers, and middleboxes. Switchyard exposes a networking abstraction similar to a \textit{raw socket}, which allows a developer to receive and send Ethernet frames on specific network ports, and provides a set of classes to simplify parsing and construction of packets and packet headers. Systems-level software created using Switchyard can be deployed on a standard Linux host or in an emulated environment like Mininet. Perhaps most importantly, Switchyard provides facilities for test-driven development by transparently allowing the underlying network to be replaced with a test harness that is specifically designed to help students through the development and debugging process. We describe experiences with using Switchyard in an undergraduate networking course in which students created an Ethernet learning switch, a fully functional IPv4 router, a firewall with rate limiter, and a deep-packet inspection middlebox device.
Joel Sommers
SIGCSE1
2014 Controller-agnostic SDN Debugging
abstract
Complexity in software-defined network (SDN) applications calls for methods and tools that can facilitate comprehensive debugging and analysis. A key challenge in this regard is that SDN configurations interact with network devices that can behave in unexpected ways, depending on factors such as traffic and application mix. In this paper, we describe OFf, a debugging and test environment for SDN developers. OFf is built on top of the fs-sdn simulator, which was developed to offer simple-to-use, accurate and scalable evaluation of OpenFlow-based SDN configurations. OFf offers standard debugging features for controller applications such as stepping, breakpoints, and watch variables. It also offers features that provide visibility into network behavior including packet tracing, packet replay and visualization features, and alerts that are triggered when, e.g., configurations change. OFf is accessed through a text interface and is designed to interoperate with any standard SDN controller platform. We demonstrate the capabilities of OFf through three test scenarios that illustrate its utility and modest performance impact on running applications. Specifically, we show how OFf can be used to analyze and fix bugs in a traffic engineering application, and to detect and repair a security vulnerability due to multiple application interaction and unexpected rule expiration.
Ramakrishnan Durairajan, Joel Sommers, Paul Barford
CoNEXT2
2014 Layer 1-informed Internet Topology Measurement
abstract
Understanding the Internet's topological structure continues to be fraught with challenges. In this paper, we investigate the hypothesis that physical maps of service provider infrastructure can be used to effectively guide topology discovery based on network layer TTL-limited measurement. The goal of our work is to focus layer 3-based probing on broadly identifying Internet infrastructure that has a fixed geographic location such as POPs, IXPs and other kinds of hosting facilities. We begin by comparing more than 1.5 years of TTL-limited probe data from the Ark project with maps of service provider infrastructure from the Internet Atlas project. We find that there are substantially more nodes and links identified in the service provider map data versus the probe data. Next, we describe a new method for probe-based measurement of physical infrastructure called POPsicle that is based on careful selection of probe source-destination pairs. We demonstrate the capability of our method through an extensive measurement study using existing "looking glass" vantage points distributed throughout the Internet and show that it reveals 2.4 times more physical node locations versus standard probing methods. To demonstrate the deployability of POPsicle we also conduct tests at an IXP. Our results again show that POPsicle can identify more physical node locations compared with standard layer 3 probes, and through this deployment approach it can be used to measure thousands of networks world wide.
Ramakrishnan Durairajan, Joel Sommers, Paul Barford
Internet Measurement Conference2
2014 Balancing Accuracy and Efficiency in TCP Flow Simulation
abstract
Efficient and accurate network simulation techniques are critical for evaluating protocols and systems at scale. For example, the magnitude of modern data center deployments requires the use of fast simulation techniques in order to evaluate newly proposed algorithms and system architectures. Unfortunately, efficient simulation can come at the cost of accuracy and realism. In this work-in-progress paper, we examine specific limitations of existing closed-form models of TCP throughput, which are commonly used to simulate aggregate performance of TCP flows. We evaluate two models in particular, comparing predictions of the models with actual performance measured using the Mininet platform. We find that the open-loop nature of these models and sensitivity to different parameters contributes to significant inaccuracies, which may in turn lead to incorrect conclusions when using such models as the basis for simulation. We describe our ongoing work on a new method for scalable TCP flow simulation that is based on ideas from XCP. Our proposed technique is highly efficient, incorporates network feedback in a closed-loop manner, and in our initial experiments shows appreciable improvement in accuracy over prior models.
Joel Sommers, Yeonju Mok
MASCOTS1
2012 Revisiting broadband performance
abstract
Understanding the empirical characteristics of broadband performance is of intrinsic importance to users and providers, and has been a significant focus of recent efforts by the Federal Communications Commission (FCC)[9]. A series of recent studies have reported results of empirical studies of broadband performance (e.g.,[11,15,22]). In this paper, we reappraise previous empirical findings on broadband performance. Our study is based on a unique corpus of crowd-sourced data consisting of over 54 million individual tests collected from 59 metropolitan markets over a 6 month period by Speedtest.net. Following analytic approaches from prior studies, our results confirm many of the raw performance results (upload/download/latency) for ISPs in specific US markets. However, the size and scope of our data enable us to examine the details of characteristics that were not identified in prior studies, thereby providing a more comprehensive view of broadband performance. Furthermore, we also report results of broadband performance characteristics in 35 metropolitan markets outside of the US. This not only provides an important baseline for future study in those markets, but also enables relative comparison of broadband performance between markets world wide.
Igor Canadi, Paul Barford, Joel Sommers
Internet Measurement Conference3
2012 Cell vs. WiFi: on the performance of metro area mobile connections
abstract
Cellular and 802.11 WiFi are compelling options for mobile Internet connectivity. The goal of our work is to understand the performance afforded by each of these technologies in diverse environments and use conditions. In this paper, we compare and contrast cellular and WiFi performance using crowd-sourced data from Speedtest.net. Our study considers spatio-temporal performance (upload/download throughput and latency) using over 3 million user-initiated tests from iOS and Android apps in 15 different metro areas collected over a 15 week period. Our basic performance comparisons show that (i) WiFi provides better absolute download/upload throughput, and a higher degree of consistency in performance; (ii) WiFi networks generally deliver lower absolute latency, but the consistency in latency is often better with cellular access; (iii) throughput and latency vary widely depending on the particular access type e.g., HSPA, EVDO, LTE, WiFi, etc.) and service provider. More broadly, our results show that performance consistency for cellular and WiFi is much lower than has been reported for wired broadband. Temporal analysis shows that average performance for cell and WiFi varies with time of day, with the best performance for large metro areas coming at non-peak hours. Spatial analysis shows that performance is highly variable across metro areas, but that there are subregions that offer consistently better performance for cell or WiFi. Comparisons between metro areas show that larger areas provide higher throughput and lower latency than smaller metro areas, suggesting where ISPs have focused their deployment efforts. Finally, our analysis reveals diverse performance characteristics resulting from the rollout of new cell access technologies and service differences among local providers.
Joel Sommers, Paul Barford
Internet Measurement Conference1
2012 Comparing metro-area cellular and WiFi performance: extended abstract
abstract
Cellular and 802.11 WiFi offer two compelling connectivity options for mobile users. The goal of our work is to better understand performance characteristics of these technologies in diverse environments and conditions. To that end, we compare and contrast cellular and Wifi performance using crowd-sourced data from speedtest.net. We consider spatio-temporal performance aspects (e.g., upload and download throughput and latency) using over 3 million user-initiated tests initiated in 15 different metro areas, collected over 15 weeks. In these preliminary results, we find that WiFi performance generally exceeds cellular performance, and that observed characteristics are highly variable across different locations and times of day. We also observe diverse performance characteristics resulting from the rollout of new cell access technologies and service differences among local providers.
Joel Sommers, Paul Barford
SIGMETRICS1
2011 On the prevalence and characteristics of MPLS deployments in the open internet
abstract
Multi-Protocol Label Switching (MPLS) is a mechanism that enables service providers to specify virtual paths through IP networks. The use of MPLS in the open Internet (i.e., public end-to-end paths) has important implications for users and network neutrality since MPLS is frequently used in traffic engineering applications today. In this paper we present a longitudinal study of the prevalence and characteristics of MPLS deployments in the open Internet. We use path measurement data collected over the past 3.5 years by the CAIDA Archipelago project (Ark), which consist of over 10 billion individual traceroutes between hosts throughout the Internet. We use two different techniques for identifying MPLS paths in Ark data: direct observation via ICMP extensions that include MPLS label information, and inference using a Bayesian data fusion methodology. Our direct observation method can only identify uniform-mode tunnels, which very likely underestimates MPLS deployments. Nonetheless, our results show that the total number of tunnels observed in a given measurement period has varied widely over time with the largest deployments in tier-1 providers. About 7% of all autonomous systems deploy MPLS and this level of deployment has been consistent over the past three years. The average length of an MPLS tunnel has decreased from 4 hops in 2008 to 3 hops in 2011, and the path length distribution is heavily skewed. About 25% of all paths in 2011 cross at least one MPLS tunnel, while 4% cross more than one. Finally, data observed in MPLS headers suggest that many ASes employ some types of traffic classification and engineering in their tunnels.
Joel Sommers, Paul Barford, Brian Eriksson
Internet Measurement Conference1
2011 DomainImpute: Inferring unseen components in the Internet
abstract
Despite many efforts over the past decade, the ability to generate topological maps of the Internet at the router-level accurately and in a timely fashion remains elusive. Mapping campaigns commonly involve traceroute-like probing that are usually non-adaptive and incomplete, thus revealing only a portion of the underlying topology. In this paper we demonstrate that standard probing methods yield datasets that implicitly contain information about much more than just the directly observed links and routers. Each probe, in addition to the underlying domain knowledge, returns information that places constraints on the underlying topology, and by integrating a large number of such constraints it is possible to accurately infer the existence of unseen components of the Internet. We describe DomainImpute, a novel data analysis methodology designed to accurately infer the unseen hop-count distances between observed routers. We use both synthetic and a large empirical dataset to validate the proposed methods. On our empirical real world dataset, we show that our methods can estimate over 55% of the unseen distances between observed routers to within a one-hop error.
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
INFOCOM3
2011 Efficient network-wide flow record generation
abstract
Experiments on diverse topics such as network measurement, management and security are routinely conducted using empirical flow export traces. However, the availability of empirical flow traces from operational networks is limited and frequently comes with significant restrictions. Furthermore, empirical traces typically lack critical meta-data (e.g., labeled anomalies) which reduce their utility in certain contexts. In this paper, we describe fs: a first-of-its-kind tool for automatically generating representative flow export records as well as basic SNMP-like router interface counts. fs generates measurements for a target network topology with specified traffic characteristics. The resulting records for each router in the topology have byte, packet and flow characteristics that are representative of what would be seen in a live network. fs also includes the ability to inject different types of anomalous events that have precisely defined characteristics, thereby enabling evaluation of proposed attack and anomaly detection methods. We validate fs by comparing it with the ns-2 simulator, which targets accurate recreation of packet-level dynamics in small network topologies. We show that data generated by fs are virtually identical to what are generated by ns-2, except over small time scales (below 1 second). We also show that fs is highly efficient, thus enabling test sets to be created for large topologies. Finally, we demonstrate the utility of fs through an assessment of anomaly detection algorithms, highlighting the need for flexible, scalable generation of network-wide measurement data with known ground truth.
Joel Sommers, Rhys Alistair Bowden, Brian Eriksson, Paul Barford, Matthew Roughan, Nick G. Duffield
INFOCOM1
2011 Inferring Unseen Components of the Internet Core
abstract
Despite many efforts over the past decade, the ability to generate topological maps of the Internet at the router-level accurately and in a timely fashion remains elusive. Mapping campaigns commonly involve {t traceroute}-like probing that are usually non-adaptive and incomplete, thus revealing only a portion of the underlying topology. In this paper we demonstrate that standard probing methods yield datasets that implicitly contain information about much more than just the directly observed links and routers. Each probe yields information that places constraints on the underlying topology, and by integrating a large number of such constraints it is possible to accurately infer the existence of unseen components of the Internet (i.e., links and routers not directly revealed by the probing). Moreover, we show that this information can be used to adaptively re-focus the probing in order to more quickly discover the topology. These findings suggest radically new and more efficient approaches to Internet mapping. Our work focuses on the discovery of the core of the Internet. We define "Internet core" as the set of routers that is roughly bounded by ingress/egress routers from stub autonomous systems. We describe a novel data analysis methodology designed to accurately infer (i) the number of unseen core routers, (ii) the unseen hop-count distances between observed routers, and (iii) unseen links between observed routers. We use a large experimental dataset to validate the proposed methods. For our data set, we show that our methods can predict the number of unseen routers to within a 13% error level, estimate 60% of the unseen distances between observed routers to within a one-hop error, and robustly detect over 35% of the unseen links between observed routers. Furthermore, we use the information extracted by our inference methodology to drive an adaptive active-probing scheme. The adaptive probing method allows us to generate maps on our data set using 50% fewer probes than standard non-adaptive approaches.
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
IEEE J. Sel. Areas Commun.3
2010 BasisDetect: a model-based network event detection framework
abstract
The ability to detect unexpected events in large networks can be a significant benefit to daily network operations. A great deal of work has been done over the past decade to develop effective anomaly detection tools, but they remain virtually unused in live network operations due to an unacceptably high false alarm rate. In this paper, we seek to improve the ability to accurately detect unexpected network events through the use of BasisDetect, a flexible but precise modeling framework. Using a small dataset with labeled anomalies, the BasisDetect framework allows us to define large classes of anomalies and detect them in different types of network data, both from single sources and from multiple, potentially diverse sources. Network anomaly signal characteristics are learned via a novel basis pursuit based methodology. We demonstrate the feasibility of our BasisDetect framework method and compare it to previous detection methods using a combination of synthetic and real-world data. In comparison with previous anomaly detection methods, our BasisDetect methodology results show a 50% reduction in the number of false alarms in a single node dataset, and over 65% reduction in false alarms for synthetic network-wide data.
Brian Eriksson, Paul Barford, Rhys Alistair Bowden, Nick G. Duffield, Joel Sommers, Matthew Roughan
Internet Measurement Conference5
2010 A Learning-Based Approach for IP Geolocation
Brian Eriksson, Paul Barford, Joel Sommers, Robert D. Nowak
PAM3
2010 Educating the next generation of spammers
abstract
Compelling experiences in introductory courses make a key difference in whether non-majors develop an interest in computer science, possibly even converting them into undergraduate majors or minors. In this paper we advocate integrated hands-on laboratory style activities to provide such pivotal experiences. In the lab activities we describe, students do not engage in programming, yet they learn to think computationally by engaging in computational activities. The course in which these labs are implemented is oriented around three aspects of the the internet's underside: its techno-scientific underpinnings, environmental and energy problems and promise brought on by its rapid growth, and security threats associated with its use. We describe the goals and content of the lab activities, as well as various challenges encountered through their implementation. We also discuss student responses and future directions.
Joel Sommers
SIGCSE1
2010 A Machine Learning Approach to TCP Throughput Prediction
abstract
TCP throughput predictionis an important capability for networks where multiple paths exist between data senders and receivers. In this paper, we describe a new lightweight method for TCP throughput prediction. Our predictor uses Support Vector Regression (SVR); prediction is based on both prior file transfer history and measurements of simple path properties. We evaluate our predictor in a laboratory setting where ground truth can be measured with perfect accuracy. We report the performance of our predictor fororacularandpracticalmeasurements of path properties over a wide range of traffic conditions and transfer sizes. For bulk transfers in heavy traffic usingoracularmeasurements, TCP throughput is predicted within 10% of the actual value 87% of the time, representing nearly a threefold improvement in accuracy over prior history-based methods. Forpracticalmeasurements of path properties, predictions can be made within 10% of the actual value nearly 50% of the time, approximately a 60% improvement over history-based methods, and with much lower measurement traffic overhead. We implement our predictor in a tool calledPathPerf, test it in the wide area, and show thatPathPerfpredicts TCP throughput accurately over diverse wide area paths.
Mariyam Mirza, Joel Sommers, Paul Barford, Xiaojin Zhu 0001
IEEE/ACM Trans. Netw.2
2010 Multiobjective monitoring for SLA compliance
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
IEEE/ACM Trans. Netw.1
2009 Network Performance Anomaly Detection and Localization
abstract
Detecting the occurrence and location of performance anomalies (e.g., high jitter or loss events) is critical to ensuring the effective operation of network infrastructures. In this paper we present a framework for detecting and localizing performance anomalies based on using an active probe-enabled measurement infrastructure deployed on the periphery of a network. Our framework has three components: an algorithm for detecting performance anomalies on a path, an algorithm for selecting which paths to probe at a given time in order to detect performance anomalies (where a path is defined as the set of links between two measurement nodes), and an algorithm for identifying the links that are causing an identified anomaly on a path (i.e., localizing). The problem of detecting an anomaly on a path is addressed by comparing probe-based measures of performance characteristics with performance guarantees for the network (e.g., SLAs). The path selection algorithm is designed to enable a tradeoff between ensuring that all links in a network are frequently monitored to detect performance anomalies, while minimizing probing overhead. The localization algorithm is designed to use existing path measurement data in such a way as to minimize the number of paths necessary for additional probing in order to identify the link(s) responsible for an observed performance anomaly. We assess the feasibility of our framework and algorithms by implementing them in ns-2 and conducting a set of simulation-based experiments using several different network topologies. Our results show that our method is able to accurately detect and localize performance anomalies in a timely fashion and with lower probe and computational overheads than previously proposed methodologies.
Paul Barford, Nick G. Duffield, Amos Ron, Joel Sommers
INFOCOM4
2008 Power Awareness in Network Design and Routing
abstract
Exponential bandwidth scaling has been a fundamental driver of the growth and popularity of the Internet. However, increases in bandwidth have been accompanied by increases in power consumption, and despite sustained system design efforts to address power demand, significant technological challenges remain that threaten to slow future bandwidth growth. In this paper we describe the power and associated heat management challenges in today's routers. We advocate a broad approach to addressing this problem that includes making power-awareness a primary objective in the design and configuration of networks, and in the design and implementation of network protocols. We support our arguments by providing a case study of power demands of two standard router platforms that enables us to create a generic model for router power consumption. We apply this model in a set of target network configurations and use mixed integer optimization techniques to investigate power consumption, performance and robustness in static network design and in dynamic routing. Our results indicate the potential for significant power savings in operational networks by including power-awareness.
Joseph Chabarek, Joel Sommers, Paul Barford, Cristian Estan, David Tsiang, Stephen J. Wright 0001
INFOCOM2
2008 A geometric approach to improving active packet loss measurement
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
IEEE/ACM Trans. Netw.1
2007 An active measurement system for shared environments
abstract
Testbeds composed of end hosts deployed across the Internet enable researchers to simultaneously conduct a wide variety of experiments. Active measurement studies of Internet path properties that require precisely crafted probe streams can be problematic in these environments. The reason is that load on the host systems from concurrently executing experiments (as is typical in PlanetLab) can significantly alter probe stream timings. In this paper we measure and characterize how packet streams from our local PlanetLab nodes are affected by experimental concurrency. We find that the effects can be extreme. We then set up a simple PlanetLab deployment in a laboratory testbed to evaluate these effects in a controlled fashion. We find that even relatively low load levels can cause serious problems in probe streams. Based on these results, we develop a novel system called MAD that can operate as a Linux kernel module or as a stand-alone daemon to support real-time scheduling of probe streams. MAD coordinates probe packet emission for all active measurement experiments on a node. We demonstrate the capabilities of MAD , showing that it performs effectively even under very high levels of multiplexing and host system load.
Joel Sommers, Paul Barford
Internet Measurement Conference1
2007 A Framework for Multi-Objective SLA Compliance Monitoring
abstract
Service level agreements (SLAs) specify performance guarantees made by service providers, typically in terms of packet loss, delay, delay variation, and network availability. While many tools have been developed to measure individual aspects of network performance, there has been little work to directly address the issue of SLA compliance monitoring in an operational setting where accuracy, parsimony, and other related issues are of vital importance. This paper takes the following steps toward addressing this problem: (1) we introduce an architectural framework for integrating multiple discrete-time active measurement algorithms, an architecture that we call multi-objective monitoring; and (2) we introduce a new active measurement methodology to monitor the packet loss rate along a network path for determining compliance with specified performance targets which significantly improves accuracy over existing techniques. We present a prototype implementation of our monitoring framework, and demonstrate how a unified probe stream can consume lower overall bandwidth than if individual streams are used to measure different path properties. We demonstrate the accuracy and convergence properties of our new loss rate monitoring methodology in a controlled laboratory environment using a range of background traffic scenarios and examine its accuracy improvements over existing techniques.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
INFOCOM1
2007 Accurate and efficient SLA compliance monitoring
abstract
Service level agreements (SLAs) define performance guarantees made by service providers, e.g, in terms of packet loss, delay, delay variation, and network availability. In this paper, we describe a new active measurement methodology to accurately monitor whether measured network path characteristics are in compliance with performance targets specified in SLAs. Specifically, (1) we describe a new methodology for estimating packet loss rate that significantly improves accuracy over existing approaches; (2) we introduce a new methodology for measuring mean delay along a path that improves accuracy over existing methodologies, and propose a method for obtaining confidence intervals on quantiles of the empirical delay distribution without making any assumption about the true distribution of delay; (3) we introduce a new methodology for measuring delay variation that is more robust than prior techniques; and (4) we extend existing work in network performance tomography to infer lower bounds on the quantiles of a distribution of performance measures along an unmeasured path given measurements from a subset of paths. We unify active measurements for these metrics in a discrete time-based tool called SLAM. The unified probe stream from SLAM consumes lower overall bandwidth than if individual streams are used to measure path properties. We demonstrate the accuracy and convergence properties of SLAM in a controlled laboratory environment using a range of background traffic scenarios and in one- and two-hop settings, and examine its accuracy improvements over existing standard techniques.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
SIGCOMM1
2007 A machine learning approach to TCP throughput prediction
abstract
TCP throughput prediction is an important capability in wide area overlay and multi-homed networks where multiple paths may exist between data sources and receivers. In this paper we describe a new, lightweight method for TCP throughput prediction that can generate accurate forecasts for a broad range of file sizes and path conditions. Our method is based on Support Vector Regression modeling that uses a combination of prior file transfers and measurements of simple path properties. We calibrate and evaluate the capabilities of our throughput predictor in an extensive set of lab-based experiments where ground truth can be established for path properties using highly accurate passive measurements. We report the performance for our method in the ideal case of using our passive path property measurements over a range of test configurations. Our results show that for bulk transfers in heavy traffic, TCP throughput is predicted within 10% of the actual value 87% of the time, representing nearly a 3-fold improvement in accuracy over prior history-based methods. In the same lab environment, we assess our method using less accurate active probe measurements of path properties, and show that predictions can be made within 10% of the actual value nearly 50% of the time over a range of file sizes and traffic conditions. This result represents approximately a 60% improvement over history-based methods with a much lower impact on end-to-end paths. Finally, we implement our predictor in a tool called PathPerf and test it in experiments conducted on wide area paths. The results demonstrate that PathPerf predicts TCP through put accurately over a variety of paths.
Mariyam Mirza, Joel Sommers, Paul Barford, Xiaojin Zhu 0001
SIGMETRICS2
2006 A Proposed Framework for Calibration of Available Bandwidth Estimation Tools
abstract
Examining the validity or accuracy of proposed available bandwidth estimation tools remains a challenging problem. A common approach consists of evaluating a newly developed tool using a combination of simple nstype simulations and feasible experiments in situ (i.e., using parts of the actual Internet). In this paper, we argue that this strategy tends to fall short of establishing a reliable "ground truth," and we advocate an alternative in vitro-like methodology for calibrating available bandwidth estimation tools that has not been widely used in this context. Our approach relies on performing controlled laboratory experiments and using tools to visualize and analyze the relevant tool-specific traffic dynamics. We present a case study of how two canonical available bandwidth estimation tools, SPRUCE and PATHLOAD, respond to increasingly more complex cross traffic and network path conditions. We expose measurement bias and algorithmic omissions that lead to poor tool calibration. As a result of this evaluation, we designed a calibrated available bandwidth estimation tool called YAZ that builds on the insights of PATHLOAD. We show that in head to head comparisons with SPRUCE and PATHLOAD, YAZ is significantly and consistently more accurate with respect to ground truth, and reports results more quickly with a small number of probes.
Joel Sommers, Paul Barford, Walter Willinger
ISCC1
2005 Scalable Network Path Emulation
abstract
Laboratory-based experimentation is an increasingly popular method for conducting network research since it enables implementations of network systems and protocols to be evaluated. Most research conducted in lab-based environments requires the faithful reproduction of wide area network conditions. An important step toward satisfying this requirement is the creation of paths between nodes in the lab that have the same characteristics as paths between nodes in the Internet. In this paper, we describe and evaluate a new, highly scalable, software-based path emulation tool called NetPath. We describe the design and implementation of NetPath, which features fixed and probabilistic packet propagation delay emulation, probabilistic bit errors, probabilistic packet loss, packet duplication, and packet reordering capability. Through a series of controlled laboratory experiments, we demonstrate that Net-Path offers over three times the loss-free throughput capacity of other popular software-based path/network emulators. We show that under moderate load NetPath's mean propagation delay emulation precision is within 1% of a hardware-based reference emulator. This result represents a significant improvement over other software-based emulators. We illustrate how, relative to our hardware-based reference, NetPath improves application traffic behavior over other software-based emulators. Finally, we demonstrate and characterize NetPath's ability to provide path emulation simultaneously on multiple physical links. This capability, which is facilitated through the use of our link configuration tool, enables laboratory system resources to be more efficiently utilized.
Shilpi Agarwal, Joel Sommers, Paul Barford
MASCOTS2
2005 Improving accuracy in end-to-end packet loss measurement
abstract
Measurement and estimation of packet loss characteristics are challenging due to the relatively rare occurrence and typically short duration of packet loss episodes. While active probe tools are commonly used to measure packet loss on end-to-end paths, there has been little analysis of the accuracy of these tools or their impact on the network. The objective of our study is to understand how to measure packet loss episodes accurately with end-to-end probes. We begin by testing the capability of standard Poisson-modulated end-to-end measurements of loss in a controlled laboratory environment using IP routers and commodity end hosts. Our tests show that loss characteristics reported from such Poisson-modulated probe tools can be quite inaccurate over a range of traffic conditions. Motivated by these observations, we introduce a new algorithm for packet loss measurement that is designed to overcome the deficiencies in standard Poisson-based tools. Specifically, our method creates a probe process that (1) enables an explicit trade-off between accuracy and impact on the network, and (2) enables more accurate measurements than standard Poisson probing at the same rate. We evaluate the capabilities of our methodology experimentally by developing and implementing a prototype tool, called BADABING. The experiments demonstrate the trade-offs between impact on the network and measurement accuracy. We show that BADABING reports loss characteristics far more accurately than traditional loss measurement tools.
Joel Sommers, Paul Barford, Nick G. Duffield, Amos Ron
SIGCOMM1
2004 Self-configuring network traffic generation
abstract
The ability to generate repeatable, realistic network traffic is critical in both simulation and testbed environments. Traffic generation capabilities to date have been limited to either simple sequenced packet streams typically aimed at throughput testing, or to application-specific tools focused on, for example, recreating representative HTTP requests. In this paper we describe Harpoon, a new application-independent tool for generating representative packet traffic at the IP flow level. Harpoon generates TCP and UDP packet flows that have the same byte, packet, temporal and spatial characteristics as measured at routers in live environments. Harpoon is distinguished from other tools that generate statistically representative traffic in that it can self-configure by automatically extracting parameters from standard Netflow logs or packet traces. We provide details on Harpoon's architecture and implementation, and validate its capabilities in controlled laboratory experiments using configurations derived from flow and packet traces gathered in live environments. We then demonstrate Harpoon's capabilities in a router benchmarking experiment that compares Harpoon with commonly used throughput test methods. Our results show that the router subsystem load generated by Harpoon is significantly different, suggesting that this kind of test can provide important insights into how routers might behave under actual operating conditions.
Joel Sommers, Paul Barford
Internet Measurement Conference1
2004 A framework for malicious workload generation
abstract
Malicious traffic from self-propagating worms and denial-of-service attacks constantly threatens the everyday operation of Internet systems. Defending networks from these threats demands appropriate tools to conduct comprehensive vulnerability assessments of networked systems. This paper describes MACE, a unique environment for recreating a wide range of malicious packet traffic in laboratory testbeds. MACE defines a model for flexible composition of malicious traffic that enables both known attacks (such as the Welchia worm) and new attack variants to be created. We implement this model in an extensible library for attack traffic specification and generation. To demonstrate the capability of MACE, we provide an analysis of stress tests conducted on a popular firewall and two popular network intrusion detection systems. Our results expose potential weaknesses of these systems and reveal that modern firewalls and network intrusion detection systems could be easily overwhelmed by simple attacks launched from a small number of hosts.
Joel Sommers, Vinod Yegneswaran, Paul Barford
Internet Measurement Conference1
2004 Harpoon: a flow-level traffic generator for router and network tests
abstract
We describe Harpoon, a new application-independent tool for generating representative packet traffic at the IP flow level. Harpoon is a configurable tool for creating TCP and UDP packet flows that have the same byte, packet, temporal, and spatial characteristics as measured at routers in live environments. We validate Harpoon using traces collected from a live router and then demonstrate its capabilities in a series of router performance benchmark tests.
Joel Sommers, Hyungsuk Kim, Paul Barford
SIGMETRICS1