EDBT 2026 Demo / reviewers in the wild / expert
Geoffrey M. Voelker
dblp:v/GeoffreyMVoelker · also Geoff Voelker
· DBLP profile ↗
126ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0003-0865-7499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 55 · 10 since 2021Security and privacy · 27 · 9 since 2021Systems, architecture and hardware · 23 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 13 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Artificial intelligence and machine learning · 5Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lost in Translation: Text Message Spoofing via Email
Sumanth Rao, Ye Shu, Stefan Savage, Aaron Schulman, Geoffrey M. Voelker, Enze Liu 0001 |
SP | 5 |
| 2026 | Understanding Server-side Commercial Fingerprinting
Elisa Luo, Tom Ritter, Stefan Savage, Geoffrey M. Voelker |
WWW | 4 |
| 2025 | Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI CrawlersabstractThe success of generative AI relies heavily on training on data scraped through extensive crawling of the Internet, a practice that has raised significant copyright, privacy, and ethical concerns. While few measures are designed to resist a resource-rich adversary determined to scrape a site, crawlers can be impacted by a range of existing tools such as robots.txt, NoAI meta tags, and active crawler blocking by reverse proxies. In this work, we seek to understand the ability and efficacy of today's networking tools to protect content creators against AI-related crawling. For targeted populations like human artists, do they have the technical knowledge and agency to utilize crawler blocking tools such as robots.txt, and can such tools be effective? Using large scale measurements and a targeted user study of 203 professional artists, we find strong demand for tools like robots.txt, but significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlers. We further test and evaluate network level crawler blockers provided by reverse proxies. Despite relatively limited deployment today, they offer stronger protections against AI crawlers, but still come with their own set of limitations. Enze Liu 0001, Elisa Luo, Shawn Shan, Geoffrey M. Voelker, Ben Y. Zhao, Stefan Savage |
IMC | 4 |
| 2025 | Canvassing the Fingerprinters: Characterizing Canvas Fingerprinting Use Across the WebabstractCanvas fingerprinting is an effective technique for implicitly re-identifying visitors to a Web site based on subtle variations in the graphical rendering of specific "test canvases". Different fingerprinting actors make use of distinct canvases for this purpose and thus, as we show, it is possible to ''fingerprint the fingerprinters'' by grouping together identical canvases that are employed for these tests. In this paper, we document the prevalence of canvas fingerprinting (finding that 12.7% of the top 20K sites engage in it), use this grouping technique to measure and characterize the online footprint of widely-used fingerprinting services, and finally analyze the context in which these services are used to shine light on their intended purpose. Elisa Luo, Tom Ritter, Stefan Savage, Geoffrey M. Voelker |
IMC | 4 |
| 2025 | Poster: When Blocks Go Missing: The Timeliness and Trustworthiness of Blockchain RPC ProvidersabstractContrary to blockchain's trustless vision, most applications built atop blockchain require trust in third-party Remote Procedure Call (RPC) providers. Applications rely on these RPC providers to be performant (announce new blocks timely) and reliable (no missing blocks/transactions) to provide good user experience and security guarantees. In this paper, we perform the first large-scale, longitudinal study to evaluate the timeliness and trustworthiness of 16 RPC providers for BNB Smart Chain (BSC) across 6 223 blocks and 123 773 transactions. We identify significant variability: some providers are inconsistent, miss valid blocks/transactions, or are seconds slower than others. Our findings suggest that the implicit trust assumptions are often violated, which may leave users confused and even vulnerable to attacks. Ye Shu, Deian Stefan, Stefan Savage, Geoffrey M. Voelker, Enze Liu 0001 |
IMC | 4 |
| 2025 | Understanding the Efficacy of Phishing Training in PracticeabstractThis paper empirically evaluates the efficacy of two ubiquitous forms of enterprise security training: annual cybersecurity awareness training and embedded anti-phishing training exercises. Specifically, our work analyzes the results of an 8-month randomized controlled experiment involving ten simulated phishing campaigns sent to over 19,500 employees at a large healthcare organization. Our results suggest that these efforts offer limited value. First, we find no significant relationship between whether users have recently completed cybersecurity awareness training and their likelihood of failing a phishing simulation. Second, when evaluating recipients of embedded phishing training, we find that the absolute difference in failure rates between trained and untrained users is extremely low across a variety of training content. Third, we observe that most users spend minimal time interacting with embedded phishing training material in-the-wild; and that for specific types of training content, users who receive and complete more instances of the training can have an increased likelihood of failing subsequent phishing simulations. Taken together, our results suggest that anti-phishing training programs, in their current and commonly deployed forms, are unlikely to offer significant practical value in reducing phishing risks. Grant Ho, Ariana Mirian, Elisa Luo, Khang Tong, Euyhyun Lee, Christopher A. Longhurst, Christian Dameff, Stefan Savage, Geoffrey M. Voelker |
SP | 10 |
| 2024 | Using Honeybuckets to Characterize Cloud Storage Scanning in the WildabstractIn this work, we analyze to what extent actors target poorly-secured cloud storage buckets for attack. We deployed hundreds of AWS S3 honeybuckets with different names and content to lure and measure different scanning strategies. Actors exhibited clear preferences for scanning buckets that appeared to belong to organizations, especially commercial entities in the technology sector with a vulnerability disclosure program. Actors continuously engaged with the content of buckets by downloading, uploading, and deleting files. Most alarmingly, we recorded multiple instances in which malicious actors downloaded, read, and understood a document from our honeybucket, leading them to attempt to gain unauthorized server access. Katherine Izhikevich, Geoffrey M. Voelker, Stefan Savage, Liz Izhikevich |
EuroS&P | 2 |
| 2024 | Give and Take: An End-To-End Investigation of Giveaway Scam Conversion RatesabstractThe Internet's combination of low communication cost, global reach, and functional anonymity has allowed fraudulent scam volumes to reach new heights. Designing effective interventions requires first understanding the context: how scammers reach potential victims, the earnings they make, and any potential bottlenecks for durable interventions. In this short paper, we focus on these questions in the context of cryptocurrency giveaway scams, where victims are tricked into irreversibly transferring funds to scammers under the pretense of even greater returns. Combining data from Twitter (also known as X), YouTube and Twitch livestreams, landing pages, and cryptocurrency blockchains, we measure how giveaway scams operate at scale. We find that 1 in 1000 scam tweets, and 4 in 100,000 livestream views, net a victim, and that scammers managed to extract nearly $4.62 million from just hundreds of victims during our measurement window. Enze Liu 0001, George Kappos, Eric Mugnier, Luca Invernizzi, Stefan Savage, David Tao, Kurt Thomas, Geoffrey M. Voelker, Sarah Meiklejohn |
IMC | 8 |
| 2024 | Unfiltered: Measuring Cloud-based Email Filtering BypassesabstractEmail service has increasingly been outsourced to cloud-based providers and so too has the task of filtering such messages for potential threats. Thus, customers will commonly direct that their incoming email is first sent to a third-party email filtering service (e.g., Proofpoint or Barracuda) and only the "clean" messages are then sent on to their email hosting provider (e.g., Gmail or Microsoft Exchange Online). However, this loosely coupled approach can, in theory, be bypassed if the email hosting provider is not configured to only accept messages that arrive from the email filtering service. In this paper we demonstrate that such bypasses are commonly possible. We document a multi-step methodology to infer if an organization has correctly configured its email hosting provider to guard against such scenarios. Then, using an empirical measurement of edu and com domains as a case study, we show that 80% of such organizations making use of popular cloud-based email filtering services can be bypassed in this manner. We also discuss reasons that lead to such misconfigurations and outline challenges in hardening the binding between email filtering and hosting providers. Sumanth Rao, Enze Liu 0001, Grant Ho, Geoffrey M. Voelker, Stefan Savage |
WWW | 4 |
| 2023 | An Empirical Analysis of Enterprise-Wide Mandatory Password UpdatesabstractEnterprise-scale mandatory password changes are disruptive, complex endeavors that require the entire workforce to prioritize a goal that is often secondary to most users. While ample literature exists around user perceptions and struggles, there are few “best practices” from the perspective of the enterprise—either to achieve the end goal or to minimize IT costs. In this paper, we provide an empirical analysis of an enterprise-scale mandatory password change, covering almost 10,000 faculty and staff at an academic institution. Using a combination of user notifications logs, password update records, and help desk ticket information, we construct an empirical model of user response over time. In particular, we characterize the elements of the campaign that relate to ideal and non-ideal outcomes, including unnecessary user actions and IT help desk overhead. We aim to provide insight into successes and challenges that can generalize to other organizations implementing similar initiatives. Ariana Mirian, Grant Ho, Stefan Savage, Geoffrey M. Voelker |
ACSAC | 4 |
| 2023 | Forward Pass: On the Security Implications of Email Forwarding Mechanism and PolicyabstractThe critical role played by email has led to a range of extension protocols (e.g., SPF, DKIM, DMARC) designed to protect against the spoofing of email sender domains. These protocols are complex as is, but are further complicated by automated email forwarding — used by individual users to manage multiple accounts and by mailing lists to redistribute messages. In this paper, we explore how such email forwarding and its implementations can break the implicit assumptions in widely deployed anti-spoofing protocols. Using large-scale empirical measurements of 20 email forwarding services (16 leading email providers and four popular mailing list services), we identify a range of security issues rooted in forwarding behavior and show how they can be combined to reliably evade existing anti-spoofing controls. We further show how these issues allow attackers to not only deliver spoofed email messages to prominent email providers (e.g., Gmail, Microsoft Outlook, and Zoho), but also reliably spoof email on behalf of tens of thousands of popular domains including sensitive domains used by organizations in government (e.g., state.gov), finance (e.g., transunion.com), law (e.g., perkinscoie.com) and news (e.g., washingtonpost.com) among others. Enze Liu 0001, Gautam Akiwate, Mattijs Jonker, Ariana Mirian, Grant Ho, Geoffrey M. Voelker, Stefan Savage |
EuroS&P | 6 |
| 2023 | Understanding the Viability of Gmail's Origin Indicator for Identifying the Sender
Enze Liu 0001, Alex Bellon, Grant Ho, Geoffrey M. Voelker, Stefan Savage, Imani N. S. Munyaka |
SOUPS | 5 |
| 2023 | No Privacy Among Spies: Assessing the Functionality and Insecurity of Consumer Android Spyware AppsabstractConsumer mobile spyware apps covertly monitor a user's activities (i.e., text messages, phone calls, e-mail, location, etc.) and transmit that information over the Internet to support remote surveillance. Unlike conceptually similar apps used for state espionage, so-called "stalkerware" apps are mass-marketed to consumers on a retail basis and expose a far broader range of victims to invasive monitoring. Today the market for such apps is large enough to support dozens of competitors, with individual vendors reportedly monitoring hundreds of thousands of phones. However, while the research community is well aware of the existence of such apps, our understanding of the mechanisms they use to operate remains ad hoc. In this work, we perform an in-depth technical analysis of 14 distinct leading mobile spyware apps targeting Android phones. We document the range of mechanisms used to monitor user activity of various kinds (e.g., photos, text messages, live microphone access) — primarily through the creative abuse of Android APIs. We also discover previously undocumented methods these apps use to hide from detection and to achieve persistence. Additionally, we document the measures taken by each app to protect the privacy of the sensitive data they collect, identifying a range of failings on the part of spyware vendors (including privacy-sensitive data sent in the clear or stored in the cloud with little or no protection). Enze Liu 0001, Sumanth Rao, Sam Havron, Grant Ho, Stefan Savage, Geoffrey M. Voelker, Damon McCoy |
Proc. Priv. Enhancing Technol. | 6 |
| 2022 | FaaSnap: FaaS made fast using snapshot-based VMsabstractFaaSnap is a VM snapshot-based platform that uses a set of complementary optimizations to improve function cold-start performance for Function-as-a-Service (FaaS) applications. Compact loading set files take better advantage of prefetching. Per-region memory mapping tailors page fault handling depending on the contents of different guest VM memory regions. Hierarchical overlapping memory-mapped regions simplify the mapping process. Concurrent paging allows the guest VM to start execution immediately, rather than pausing until the working set is loaded. Altogether, FaaSnap significantly reduces guest VM page fault handling time on the critical path and improves overall function loading performance. Experiments on serverless benchmarks show that it reduces end-to-end function execution by up to 3.5x compared to state-of-the-art, and on average is only 3.5% slower than snapshots cached in memory. Moreover, we show that FaaSnap is resilient to changes of working set and remains efficient under bursty workloads and when snapshots are located in remote storage. Lixiang Ao, George Porter, Geoffrey M. Voelker |
EuroSys | 3 |
| 2022 | Retroactive identification of targeted DNS infrastructure hijackingabstractIn 2019, the US Department of Homeland Security issued an emergency warning about DNS infrastructure tampering. This alert, in response to a series of attacks against foreign government websites, highlighted how a sophisticated attacker could leverage access to key DNS infrastructure to then hijack traffic and harvest valid login credentials for target organizations. However, even armed with this knowledge, identifying the existence of such incidents has been almost entirely via post hoc forensic reports (i.e., after a breach was found via some other method). Indeed, such attacks are particularly challenging to detect because they can be very short lived, bypass the protections of TLS and DNSSEC, and are imperceptible to users. Identifying them retroactively is even more complicated by the lack of fine-grained Internet-scale forensic data. This paper is a first attempt to make progress at this latter goal. Combining a range of longitudinal data from Internet-wide scans, passive DNS records, and Certificate Transparency logs, we have constructed a methodology for identifying potential victims of sophisticated DNS infrastructure hijacking and have used it to identify a range of victims (primarily government agencies), both those named in prior reporting, and others previously unknown. Gautam Akiwate, Raffaele Sommese, Mattijs Jonker, Zakir Durumeric, K. C. Claffy, Geoffrey M. Voelker, Stefan Savage |
IMC | 6 |
| 2022 | Where .ru?: assessing the impact of conflict on russian domain infrastructureabstractThe hostilities in Ukraine have driven unprecedented forces, both from third-party countries and in Russia, to create economic barriers. In the Internet, these manifest both as internal pressures on Russian sites to (re-)patriate the infrastructure they depend on (e.g., naming and hosting) and external pressures arising from Western providers disassociating from some or all Russian customers. While quite a bit has been written about this both from a policy perspective and anecdotally, our paper places the question on an empirical footing and directly measures longitudinal changes in the makeup of naming, hosting and certificate issuance for domains in the Russian Federation. Mattijs Jonker, Gautam Akiwate, Antonia Affinito, K. C. Claffy, Alessio Botta, Geoffrey M. Voelker, Roland van Rijswijk-Deij, Stefan Savage |
IMC | 6 |
| 2022 | Measuring UID smuggling in the wildabstractThis work presents a systematic study of UID smuggling, an emerging tracking technique that is designed to evade browsers' privacy protections. Browsers are increasingly attempting to prevent cross-site tracking by partitioning the storage where trackers store user identifiers (UIDs). UID smuggling allows trackers to synchronize UIDs across sites by inserting UIDs into users' navigation requests. Trackers can thus regain the ability to aggregate users' activities and behaviors across sites, in defiance of browser protections. Audrey Randall, Peter Snyder, Alisha Ukani, Alex C. Snoeren, Geoffrey M. Voelker, Stefan Savage, Aaron Schulman |
IMC | 5 |
| 2021 | Risky BIZness: risks derived from registrar name managementabstractIn this paper, we explore a domain hijacking risk that is an accidental byproduct of undocumented operational practices between domain registrars and registries. We show how over the last nine years over 512K domains have been implicitly exposed to the risk of hijacking, affecting names in most popular TLDs (including .com and .net) as well as legacy TLDs with tight registration control (such as .edu and .gov). Moreover, we show that this weakness has been actively exploited by multiple parties who, over the years, have assumed control over 163K domains without having any ownership interest in those names. In addition to characterizing the nature and size of this problem, we also report on the efficacy of the remediation in response to our outreach with registrars. Gautam Akiwate, Stefan Savage, Geoffrey M. Voelker, K. C. Claffy |
Internet Measurement Conference | 3 |
| 2021 | Who's got your mail?: characterizing mail service provider usageabstractE-mail has long been a critical component of daily communication and the core medium for modern business correspondence. While traditionally e-mail service was provisioned and implemented independently by each Internet-connected organization, increasingly this function has been outsourced to third-party services. As with many pieces of key communications infrastructure, such centralization can bring both economies of scale and shared failure risk. In this paper, we investigate this issue empirically --- providing a large-scale measurement and analysis of modern Internet e-mail service provisioning. We develop a reliable methodology to better map domains to mail service providers. We then use this approach to document the dominant and increasing role played by a handful of mail service providers and hosting companies over the past four years. Finally, we briefly explore the extent to which nationality (and hence legal jurisdiction) plays a role in such mail provisioning decisions. Enze Liu 0001, Gautam Akiwate, Mattijs Jonker, Ariana Mirian, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 6 |
| 2021 | Home is where the hijacking is: understanding DNS interception by residential routersabstractDNS interception --- when a user's DNS queries to a target resolver are intercepted en route and forwarded to a different resolver --- is a phenomenon of concern to both researchers and Internet users because of its implications for security and privacy. While the prevalence of DNS interception has received some attention, less is known about where in the network interception takes place. We introduce methods to identify where DNS interception occurs and who the interceptors may be. We identify when interception is performed before the query exits the ISP, and even when it is performed by the Customer Premises Equipment (CPE) in the user's own home. We believe that these techniques are vital in the light of the ongoing debate concerning the value of privacy-enhancing DNS transport. Audrey Randall, Enze Liu 0001, Ramakrishna Padmanabhan, Gautam Akiwate, Geoffrey M. Voelker, Stefan Savage, Aaron Schulman |
Internet Measurement Conference | 5 |
| 2021 | Clairvoyance: Inferring Blocklist Use on the Internet
Vector Guo Li, Gautam Akiwate, Kirill Levchenko, Geoffrey M. Voelker, Stefan Savage |
PAM | 4 |
| 2021 | Hopper: Modeling and Detecting Lateral Movement
Grant Ho, Mayank Dhiman, Devdatta Akhawe, Vern Paxson, Stefan Savage, Geoffrey M. Voelker, David A. Wagner 0001 |
USENIX Security Symposium | 6 |
| 2020 | Particle: ephemeral endpoints for serverless networkingabstractBurst-parallel serverless applications invoke thousands of short-lived distributed functions to complete complex jobs such as data analytics, video encoding, or compilation. While these tasks execute in seconds, starting and configuring the virtual network they rely on is a major bottleneck that can consume up to 84% of total startup time. In this paper we characterize the magnitude of this network cold start problem in three popular overlay networks, Docker Swarm, Weave, and Linux Overlay. We focus on end-to-end startup time that encompasses both the time to boot a group of containers as well as interconnecting them. Our primary observation is that existing overlay approaches for serverless networking scale poorly in short-lived serverless environments. Based on our findings we develop Particle, a network stack tailored for multi-node serverless overlay networks that optimizes network creation without sacrificing multi-tenancy, generality, or throughput. When integrated into a serverless burst-parallel video processing pipeline, Particle improves application runtime by 2.4--3X over existing overlays. Shelby Thomas, Lixiang Ao, Geoffrey M. Voelker, George Porter |
SoCC | 3 |
| 2020 | Unresolved Issues: Prevalence, Persistence, and Perils of Lame DelegationsabstractThe modern Internet relies on the Domain Name System (DNS) to convert between human-readable domain names and IP addresses. However, the correct and efficient implementation of this function is jeopardized when the configuration data binding domains, nameservers and glue records is faulty. In particular lame delegations, which occur when a nameserver responsible for a domain is unable to provide authoritative information about it, introduce both performance and security risks. We perform a broad-based measurement study of lame delegations, using both longitudinal zone data and active querying. We show that lame delegations of various kinds are common (affecting roughly 14% of domains we queried), that they can significantly degrade lookup latency (when they do not lead to outright failure), and that they expose hundreds of thousands of domains to adversarial takeover. We also explore circumstances that give rise to this surprising prevalence of lame delegations, including unforeseen interactions between the operational procedures of registrars and registries. Gautam Akiwate, Mattijs Jonker, Raffaele Sommese, Ian D. Foster, Geoffrey M. Voelker, Stefan Savage, K. C. Claffy |
Internet Measurement Conference | 5 |
| 2020 | Trufflehunter: Cache Snooping Rare Domains at Large Public DNS ResolversabstractThis paper presents and evaluates Trufflehunter, a DNS cache snooping tool for estimating the prevalence of rare and sensitive Internet applications. Unlike previous efforts that have focused on small, misconfigured open DNS resolvers, Trufflehunter models the complex behavior of large multi-layer distributed caching infrastructures (e.g., such as Google Public DNS). In particular, using controlled experiments, we have inferred the caching strategies of the four most popular public DNS resolvers (Google Public DNS, Cloudflare Quad1, OpenDNS and Quad9). The large footprint of such resolvers presents an opportunity to observe rare domain usage, while preserving the privacy of the users accessing them. Using a controlled testbed, we evaluate how accurately Trufflehunter can estimate domain name usage across the U.S. Applying this technique in the wild, we provide a lower-bound estimate of the popularity of several rare and sensitive applications (most notably smartphone stalkerware) which are otherwise challenging to survey. Audrey Randall, Enze Liu 0001, Gautam Akiwate, Ramakrishna Padmanabhan, Geoffrey M. Voelker, Stefan Savage, Aaron Schulman |
Internet Measurement Conference | 5 |
| 2019 | Measuring Security Practices and How They Impact SecurityabstractSecurity is a discipline that places significant expectations on lay users. Thus, there are a wide array of technologies and behaviors that we exhort end users to adopt and thereby reduce their security risk. However, the adoption of these "best practices" --- ranging from the use of antivirus products to actively keeping software updated --- is not well understood, nor is their practical impact on security risk well-established. This paper explores both of these issues via a large-scale empirical measurement study covering approximately 15,000 computers over six months. We use passive monitoring to infer and characterize the prevalence of various security practices in situ as well as a range of other potentially security-relevant behaviors. We then explore the extent to which differences in key security behaviors impact real-world outcomes (i.e., that a device shows clear evidence of having been compromised). Louis F. DeKoven, Audrey Randall, Ariana Mirian, Gautam Akiwate, Ansel Blume, Lawrence K. Saul, Aaron Schulman, Geoffrey M. Voelker, Stefan Savage |
Internet Measurement Conference | 8 |
| 2019 | Detecting and Characterizing Lateral Phishing at Scale
Grant Ho, Asaf Cidon, Lior Gavish, Marco Schweighauser, Vern Paxson, Stefan Savage, Geoffrey M. Voelker, David A. Wagner 0001 |
USENIX Security Symposium | 7 |
| 2019 | Reading the Tea leaves: A Comparative Analysis of Threat Intelligence
Vector Guo Li, Matthew Dunn, Paul Pearce, Damon McCoy, Geoffrey M. Voelker, Stefan Savage |
USENIX Security Symposium | 5 |
| 2019 | Hack for Hire: Exploring the Emerging Market for Account HijackingabstractEmail accounts represent an enticing target for attackers, both for the information they contain and the root of trust they provide to other connected web services. While defense-in-depth approaches such as phishing detection, risk analysis, and two-factor authentication help to stem large-scale hijackings, targeted attacks remain a potent threat due to the customization and effort involved. In this paper, we study a segment of targeted attackers known as “hack for hire” services to understand the playbook that attackers use to gain access to victim accounts. Posing as buyers, we interacted with 27 English, Russian, and Chinese blackmarket services, only five of which succeeded in attacking synthetic (though realistic) identities we controlled. Attackers primarily relied on tailored phishing messages, with enough sophistication to bypass SMS two-factor authentication. However, despite the ability to successfully deliver account access, the market exhibited low volume, poor customer service, and had multiple scammers. As such, we surmise that retail email hijacking has yet to mature to the level of other criminal market segments. Ariana Mirian, Joe DeBlasio, Stefan Savage, Geoffrey M. Voelker, Kurt Thomas |
WWW | 4 |
| 2019 | Introduction to the Special Section on OSDI'18abstractNo abstract available. Andrea C. Arpaci-Dusseau, Geoffrey M. Voelker |
ACM Trans. Storage | 2 |
| 2018 | Dark packets and the end of network scalingabstractToday 100GbE network interfaces are commercially available, with 400GbE proposals already in the standardization process. In this environment, a major bottleneck is DRAM latency, which has stagnated at 100ns per access. Beyond 100GbE, all packet sizes will arrive faster than main memory can accommodate, resulting packet drops due to the latency incurred by the memory hierarchy. Shelby Thomas, Rob McGuinness, Geoffrey M. Voelker, George Porter |
ANCS | 3 |
| 2018 | Sprocket: A Serverless Video Processing FrameworkabstractSprocket is a highly configurable, stage-based, scalable, serverless video processing framework that exploits intra-video parallelism to achieve low latency. Sprocket enables developers to program a series of operations over video content in a modular, extensible manner. Programmers implement custom operations, ranging from simple video transformations to more complex computer vision tasks, in a simple pipeline specification language to construct custom video processing pipelines. Sprocket then handles the underlying access, encoding and decoding, and processing of video and image content across operations in a highly parallel manner. In this paper we describe the design and implementation of the Sprocket system on the AWS Lambda serverless cloud infrastructure, and evaluate Sprocket under a variety of conditions to show that it delivers its performance goals of high parallelism, low latency, and low cost (10s of seconds to process a 3,600 second video 1000-way parallel for less than $3). Lixiang Ao, Liz Izhikevich, Geoffrey M. Voelker, George Porter |
SoCC | 3 |
| 2018 | Following Their Footsteps: Characterizing Account Automation Abuse and Defenses
Louis F. DeKoven, Trevor Pottinger, Stefan Savage, Geoffrey M. Voelker, Nektarios Leontiadis |
Internet Measurement Conference | 4 |
| 2018 | An Empirical Analysis of the Commercial VPN Ecosystem
Mohammad Taha Khan, Joe DeBlasio, Geoffrey M. Voelker, Alex C. Snoeren, Chris Kanich, Narseo Vallina-Rodriguez |
Internet Measurement Conference | 3 |
| 2017 | Exploring the dynamics of search advertiser fraudabstractMost search engines generate significant revenue through search advertising, wherein advertisements are served alongside traditional search results. These advertisements are attractive to advertisers because ads can be targeted and prominently presented to users at the exact moment that the user is searching for relevant topics. Joe DeBlasio, Saikat Guha 0002, Geoffrey M. Voelker, Alex C. Snoeren |
Internet Measurement Conference | 3 |
| 2017 | Tripwire: inferring internet site compromiseabstractPassword reuse has been long understood as a problem: credentials stolen from one site may be leveraged to gain access to another site for which they share a password. Indeed, it is broadly understood that attackers exploit this fact and routinely leverage credentials extracted from a site they have breached to access high-value accounts at other sites (e.g., email accounts). However, as a consequence of such acts, this same phenomena of password reuse attacks can be harnessed to indirectly infer site compromises---even those that would otherwise be unknown. In this paper we describe such a measurement technique, in which unique honey accounts are registered with individual third-party websites, and thus access to an email account provides indirect evidence of credentials theft at the corresponding website. We describe a prototype system, called Tripwire, that implements this technique using an automated Web account registration system combined with email account access data from a major email provider. In a pilot study monitoring more than 2,300 sites over a year, we have detected 19 site compromises, including what appears to be a plaintext password compromise at an Alexa top-500 site with more than 45 million active users. Joe DeBlasio, Stefan Savage, Geoffrey M. Voelker, Alex C. Snoeren |
Internet Measurement Conference | 3 |
| 2015 | Scheduling techniques for hybrid circuit/packet networksabstractA range of new datacenter switch designs combine wireless or optical circuit technologies with electrical packet switching to deliver higher performance at lower cost than traditional packet-switched networks. These "hybrid" networks schedule large traffic demands via a high-rate circuits and remaining traffic with a lower-rate, traditional packet-switches. Achieving high utilization requires an efficient scheduling algorithm that can compute proper circuit configurations and balance traffic across the switches. Recent proposals, however, provide no such algorithm and rely on an omniscient oracle to compute optimal switch configurations. Matthew K. Mukerjee, Conglong Li, Nicolas Feltman, George Papen, Stefan Savage, Srinivasan Seshan, Geoffrey M. Voelker, David G. Andersen, Michael Kaminsky, George Porter, Alex C. Snoeren |
CoNEXT | 8 |
| 2015 | Affiliate Crookies: Characterizing Affiliate Marketing AbuseabstractModern affiliate marketing networks provide an infrastructure for connecting merchants seeking customers with independent marketers (affiliates) seeking compensation. This approach depends on Web cookies to identify, at checkout time, which affiliate should receive a commission. Thus, scammers ``stuff'' their own cookies into a user's browser to divert this revenue. This paper provides a measurement-based characterization of cookie-stuffing fraud in online affiliate marketing. We use a custom-built Chrome extension, AffTracker, to identify affiliate cookies and use it to gather data from hundreds of thousands of crawled domains which we expect to be targeted by fraudulent affiliates. Overall, despite some notable historical precedents, we found cookie-stuffing fraud to be relatively scarce in our data set. Based on what fraud we detected, though, we identify which categories of merchants are most targeted and which third-party affiliate networks are most implicated in stuffing scams. We find that large affiliate networks are targeted significantly more than merchant-run affiliate programs. However, scammers use a wider range of evasive techniques to target merchant-run affiliate programs to mitigate the risk of detection suggesting that in-house affiliate programs enjoy stricter policing. Neha Chachra, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 3 |
| 2015 | From .academy to .zone: An Analysis of the New TLD Land RushabstractThe com, net, and org TLDs contain roughly 150 million registered domains, and domain registrants often have a difficult time finding a desirable and available name. In 2013, ICANN began delegation of a new wave of TLDs into the Domain Name System with the goal of improving meaningful name choice for registrants. The new rollout resulted in over 500 new TLDs in the first 18 months, nearly tripling the number of TLDs. Previous rollouts of small numbers of new TLDs have resulted in a burst of defensive registrations as companies aggressively defend their trademarks to avoid consumer confusion. This paper analyzes the types of domain registrations in the new TLDs to determine registrant behavior in the brave new world of naming abundance. We also examine the cost structures and monetization models for the new TLDs to identify which registries are profitable. We gather DNS, Web, and WHOIS data for each new domain, and combine this with cost structure data from ICANN, the registries, and domain registrars to estimate the total cost ofthe new TLD program. We find that only 15% of domains in the new TLDs show characteristics consistent with primary registrations, while the rest are promotional, speculative, or defensive in nature; indeed, 16% of domains with NS records do not even resolve yet, and 32% are parked. Our financial analysis suggests only half of the registries have earned enough to cover their application fees, and 10% of current registries likely never will solely from registration revenue. Tristan Halvorson, Matthew F. Der, Ian D. Foster, Stefan Savage, Lawrence K. Saul, Geoffrey M. Voelker |
Internet Measurement Conference | 6 |
| 2015 | Who is .com?: Learning to Parse WHOIS RecordsabstractWHOIS is a long-established protocol for querying information about the 280M+ registered domain names on the Internet. Unfortunately, while such records are accessible in a ``human-readable'' format, they do not follow any consistent schema and thus are challenging to analyze at scale. Existing approaches, which rely on manual crafting of parsing rules and per-registrar templates, are inherently limited in coverage and fragile to ongoing changes in data representations. In this paper, we develop a statistical model for parsing WHOIS records that learns from labeled examples. Our model is a conditional random field (CRF) with a small number of hidden states, a large number of domain-specific features, and parameters that are estimated by efficient dynamic-programming procedures for probabilistic inference. We show that this approach can achieve extremely high accuracy (well over 99%) using modest amounts of labeled training data, that it is robust to minor changes in schema, and that it can adapt to new schema variants by incorporating just a handful of additional examples. Finally, using our parser, we conduct an exhaustive survey of the registration patterns found in 102M com domains. Suqi Liu, Ian D. Foster, Stefan Savage, Geoffrey M. Voelker, Lawrence K. Saul |
Internet Measurement Conference | 4 |
| 2015 | Managing Contention with MedleyabstractAs WLANs achieve gigabit per second speeds, they will need to support users with a wide range of workloads, ranging from VoIP and Web clients to data backup, file transfers, and streaming high-definition video. Unfortunately, channel efficiency degrades severely in these scenarios under existing MAC protocols due to contention and back-off overheads. Moreover, small yet latency-sensitive flows suffer disproportionally as load increases. We present Medley, a system that leverages frequency-based contention to allocate subchannels in an OFDMA-based link layer in a delay-fair manner. In contrast to traditional CSMA schemes in which each node competes uniformly for the channel, Medley ensures that nodes with smaller service rates are served before those with heavier demand; the more bandwidth a node consumes, the larger its packet average delay will become. An initial implementation of Medley on a software defined radio platform demonstrates its feasibility in a small network, while more comprehensive simulation results show its benefits under a wider range of conditions. Medley delivers delay fairness while remaining over 94 percent efficient in the face of massive over-subscription. Geoffrey M. Voelker, Alex C. Snoeren |
IEEE Trans. Mob. Comput. | 2 |
| 2014 | Characterizing Large-Scale Click Fraud in ZeroAccessabstractClick fraud is a scam that hits a criminal sweet spot by both tapping into the vast wealth of online advertising and exploiting that ecosystem's complex structure to obfuscate the flow of money to its perpetrators. In this work, we illuminate the intricate nature of this activity through the lens of ZeroAccess--one of the largest click fraud botnets in operation. Using a broad range of data sources, including peer-to-peer measurements, command-and-control telemetry, and contemporaneous click data from one of the top ad networks, we construct a view into the scale and complexity of modern click fraud operations. By leveraging the dynamics associated with Microsoft's attempted takedown of ZeroAccess in December 2013, we employ this coordinated view to identify "ad units" whose traffic (and hence revenue) primarily derived from ZeroAccess. While it proves highly challenging to extrapolate from our direct observations to a truly global view, by anchoring our analysis in the data for these ad units we estimate that the botnet's fraudulent activities plausibly induced advertising losses on the order of $100,000 per day. Paul Pearce, Vacha Dave, Chris Grier, Kirill Levchenko, Saikat Guha 0002, Damon McCoy, Vern Paxson, Stefan Savage, Geoffrey M. Voelker |
CCS | 9 |
| 2014 | Search + Seizure: The Effectiveness of Interventions on SEO CampaignsabstractBlack hat search engine optimization (SEO), the practice of abusively manipulating search results, is an enticing method to acquire targeted user traffic. In turn, a range of interventions--from modifying search results to seizing domains--are used to combat this activity. In this paper, we examine the effectiveness of these interventions in the context of an understudied market niche, counterfeit luxury goods. Using eight months of empirical crawled data, we identify 52 distinct SEO campaigns, document how well they are able to place search results for sixteen luxury brands, how this capability impacts the dynamics of their order volumes and how well existing interventions undermine this business when employed. David Y. Wang, Matthew F. Der, Mohammad Karami, Lawrence K. Saul, Damon McCoy, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 7 |
| 2014 | Knock it off: profiling the online storefronts of counterfeit merchandiseabstractWe describe an automated system for the large-scale monitoring of Web sites that serve as online storefronts for spam-advertised goods. Our system is developed from an extensive crawl of black-market Web sites that deal in illegal pharmaceuticals, replica luxury goods, and counterfeit software. The operational goal of the system is to identify the affiliate programs of online merchants behind these Web sites; the system itself is part of a larger effort to improve the tracking and targeting of these affiliate programs. There are two main challenges in this domain. The first is that appearances can be deceiving: Web pages that render very differently are often linked to the same affiliate program of merchants. The second is the difficulty of acquiring training data: the manual labeling of Web pages, though necessary to some degree, is a laborious and time-consuming process. Our approach in this paper is to extract features that reveal when Web pages linked to the same affiliate program share a similar underlying structure. Using these features, which are mined from a small initial seed of labeled data, we are able to profile the Web sites of forty-four distinct affiliate programs that account, collectively, for hundreds of millions of dollars in illicit e-commerce. Our work also highlights several broad challenges that arise in the large-scale, empirical study of malicious activity on the Web. Matthew F. Der, Lawrence K. Saul, Stefan Savage, Geoffrey M. Voelker |
KDD | 4 |
| 2014 | Enfold: downclocking OFDM in WiFiabstractDynamic voltage and frequency scaling (DVFS) has long been used as a technique to save power in a variety of computing domains but typically not in communications devices. A fundamental limit that prevents decreasing the clock frequency is the Nyquist(-Shannon) sampling theorem, which states that the sampling rate must be twice the signal bandwidth. Recently, researchers have leveraged compressive sensing to demonstrate the possibility of decoding a sparse signal below Nyquist rate. In this work, we dramatically extend the state of the art by showing how to decode non-sparse signals, in particular, OFDM systems at sub-Nyquist rates. We exploit the aliasing that results from under-sampling and observe that there exists well-defined structure in terms of how OFDM signals are "folded up" under aliasing. Based on our observations, we present Enfold, which allows existing WiFi chipsets to decode standards-compliant WiFi frames while operating at 50% and 25% of their rated clock rate. Our design is able to attain greater than 96% and 83% raw packet reception rates for moderate SNR while reducing the clock rate by 2x and 4x, respectively. Moreover, our approach can be easily applied to other communication systems based on OFDM modulation. When evaluated on popular smartphone app traces, Enfold reduces energy consumption by up to 34%. Patrick Ling, Geoffrey M. Voelker, Alex C. Snoeren |
MobiCom | 3 |
| 2014 | DSpin: Detecting Automatically Spun Content on the Web
Qing Zhang 0012, David Y. Wang, Geoffrey M. Voelker |
NDSS | 3 |
| 2014 | Circuit Switching Under the Radar with REACToR
Alex Forencich, Rishi Kapoor, Malveeka Tewari, Geoffrey M. Voelker, George Papen, Alex C. Snoeren, George Porter |
NSDI | 6 |
| 2014 | XXXtortion?: inferring registration intent in the .XXX TLDabstractAfter a decade-long approval process, multiple rejections, and an independent review, ICANN approved the xxx TLD for inclusion in the Domain Name System, to begin general availability on December 6, 2011. Its sponsoring registry proposed it as an expansion of the name space, as well as a way to separate adult from child-appropriate content. Many independent groups, including trademark holders, political groups, and the adult entertainment industry itself, were concerned that it would primarily generate value through defensive and speculative registrations, without actually serving a real need. This paper measures the validity of these concerns using data gathered from ICANN, whois, and Web requests. We use this information to characterize each xxx domain and infer the registrant's most likely intent. We find that at most 3.8% of xxx domains host or redirect to potentially legitimate Web content, with the rest generally serving either defensive or speculative purposes. Indeed, registrants spent roughly $13M up front to defend existing brands and trademarks within the xxx TLD, and an additional $11M over the course of the first year. Additional evidence suggests that over 80% of annual domain registrations are for purely defensive purposes and do not even resolve. Tristan Halvorson, Kirill Levchenko, Stefan Savage, Geoffrey M. Voelker |
WWW | 4 |
| 2013 | Bullet trains: a study of NIC burst behavior at microsecond timescalesabstractWhile numerous studies have examined the macro-level behavior of traffic in data center networks---overall flow sizes, destination variability, and TCP burstiness---little information is available on the behavior of data center traffic at packet-level timescales. Whereas one might assume that flows from different applications fairly share available link bandwidth, and that packets within a single flow are uniformly paced, the reality is more complex. To meet increasingly high link rates of 10 Gbps and beyond, batching is typically introduced across the network stack---at the application, middleware, OS, transport, and NIC layers. This batching results in short-term packet bursts, which have implications for the design and performance requirements of packet processing devices along the path, including middleboxes, SDN-enabled switches, and virtual machine hypervisors. Rishi Kapoor, Alex C. Snoeren, Geoffrey M. Voelker, George Porter |
CoNEXT | 3 |
| 2013 | A Variational Approximation for Topic Modeling of Hierarchical CorporaabstractWe study the problem of topic modeling in corpora whose documents are organized in a multi-level hierarchy. We explore a parametric approach to this problem, assuming that the number of topics is known or can be estimated by cross-validation. The models we consider can be viewed as special (finite-dimensional) instances of hierarchical Dirichlet processes (HDPs). For these models we show that there exists a simple variational approximation for probabilistic inference. The approximation relies on a previously unexploited inequality that handles the conditional dependence between Dirichlet latent variables in adjacent levels of the model’s hierarchy. We compare our approach to existing implementations of nonparametric HDPs. On several benchmarks we find that our approach is faster than Gibbs sampling and able to learn more predictive models than existing variational methods. Finally, we demonstrate the large-scale viability of our approach on two newly available corpora from researchers in computer security–one with 350,000 documents and over 6,000 internal subcategories, the other with a five-level deep hierarchy. Do-kyum Kim, Geoffrey M. Voelker, Lawrence K. Saul |
ICML (2) | 2 |
| 2013 | A fistful of bitcoins: characterizing payments among men with no namesabstractBitcoin is a purely online virtual currency, unbacked by either physical commodities or sovereign obligation; instead, it relies on a combination of cryptographic protection and a peer-to-peer protocol for witnessing settlements. Consequently, Bitcoin has the unintuitive property that while the ownership of money is implicitly anonymous, its flow is globally visible. In this paper we explore this unique characteristic further, using heuristic clustering to group Bitcoin wallets based on evidence of shared authority, and then using re-identification attacks (i.e., empirical purchasing of goods and services) to classify the operators of those clusters. From this analysis, we characterize longitudinal changes in the Bitcoin market, the stresses these changes are placing on the system, and the challenges for those seeking to use Bitcoin for criminal or fraudulent purposes at scale. Sarah Meiklejohn, Marjori Pomarole, Grant Jordan, Kirill Levchenko, Damon McCoy, Geoffrey M. Voelker, Stefan Savage |
Internet Measurement Conference | 6 |
| 2013 | Juice: A Longitudinal Study of an SEO Botnet
David Y. Wang, Stefan Savage, Geoffrey M. Voelker |
NDSS | 3 |
| 2013 | SloMo: Downclocking WiFi Communication
Geoffrey M. Voelker, Alex C. Snoeren |
NSDI | 2 |
| 2013 | eDoctor: Automatically Diagnosing Abnormal Battery Drain Issues on Smartphones
Xiao Ma 0014, Peng Huang 0005, Xinxin Jin, Dongcai Shen, Yuanyuan Zhou 0001, Lawrence K. Saul, Geoffrey M. Voelker |
NSDI | 9 |
| 2012 | Manufacturing compromise: the emergence of exploit-as-a-serviceabstractWe investigate the emergence of the exploit-as-a-service model for driveby browser compromise. In this regime, attackers pay for an exploit kit or service to do the "dirty work" of exploiting a victim's browser, decoupling the complexities of browser and plugin vulnerabilities from the challenges of generating traffic to a website under the attacker's control. Upon a successful exploit, these kits load and execute a binary provided by the attacker, effectively transferring control of a victim's machine to the attacker. Chris Grier, Lucas Ballard, Juan Caballero, Neha Chachra, Christian Dietrich 0005, Kirill Levchenko, Panayiotis Mavrommatis, Damon McCoy, Antonio Nappa, Andreas Pitsillidis, Niels Provos, M. Zubair Rafique, Moheeb Abu Rajab, Christian Rossow, Kurt Thomas, Vern Paxson, Stefan Savage, Geoffrey M. Voelker |
CCS | 18 |
| 2012 | Priceless: the role of payments in abuse-advertised goodsabstractLarge-scale abusive advertising is a profit-driven endeavor. Without consumers purchasing spam-advertised Viagra, search-advertised counterfeit software or malware-advertised fake anti-virus, these campaigns could not be economically justified. Thus, in addition to the numerous efforts focused on identifying and blocking individual abusive advertising mechanisms, a parallel research direction has emerged focused on undermining the associated means of monetization: payment networks. In this paper we explain the complex role of payment processing in monetizing the modern affiliate program ecosystem and characterize the dynamics of these banking relationships over two years within the counterfeit pharmaceutical and software sectors. By opportunistically combining our own active purchasing data with contemporary disruption efforts by brand-holders and payment card networks, we gather the first empirical dataset concerning this approach. We discuss how well such payment interventions work, how abusive merchants respond in kind and the role that the payments ecosystem is likely to play in the future. Damon McCoy, Hitesh Dharmdasani, Christian Kreibich, Geoffrey M. Voelker, Stefan Savage |
CCS | 4 |
| 2012 | Chronos: predictable low latency for data center applicationsabstractIn data center applications, predictability in service time and controlled latency, especially tail latency, are essential for building performant applications. This is especially true for applications or services built by accessing data across thousands of servers to generate a user response. Current practice has been to run such services at low utilization to rein in latency outliers, which decreases efficiency and limits the number of service invocations developers can issue while still meeting tight latency budgets. Rishi Kapoor, George Porter, Malveeka Tewari, Geoffrey M. Voelker, Amin Vahdat |
SoCC | 4 |
| 2012 | BlueSky: a cloud-backed file system for the enterprise
Michael Vrable, Stefan Savage, Geoffrey M. Voelker |
FAST | 3 |
| 2012 | Taster's choice: a comparative analysis of spam feedsabstractE-mail spam has been the focus of a wide variety of measurement studies, at least in part due to the plethora of spam data sources available to the research community. However, there has been little attention paid to the suitability of such data sources for the kinds of analyses they are used for. In spite of the broad range of data available, most studies use a single "spam feed" and there has been little examination of how such feeds may differ in content. In this paper we provide this characterization by comparing the contents of ten distinct contemporaneous feeds of spam-advertised domain names. We document significant variations based on how such feeds are collected and show how these variations can produce differences in findings as a result. Andreas Pitsillidis, Chris Kanich, Geoffrey M. Voelker, Kirill Levchenko, Stefan Savage |
Internet Measurement Conference | 3 |
| 2012 | Weighted fair queuing with differential droppingabstractWeighted fair queuing (WFQ) allows Internet operators to define traffic classes and then assign different bandwidth proportions to these classes. Unfortunately, the complexity of efficiently allocating the buffer space to each traffic class turns out to be overwhelming, leading most operators to vastly overprovision buffering-resulting in a large resource footprint. A single buffer for all traffic classes would be preferred due to its simplicity and ease of management. Our work is inspired by the approximate differential dropping scheme but differs substantially in the flow identification and packet dropping strategies. Augmented with our novel differential dropping scheme, a shared buffer WFQ performs as well or better than the original WFQ implementation under varied traffic loads with a vastly reduced resource footprint. Geoffrey M. Voelker, Alex C. Snoeren |
INFOCOM | 2 |
| 2012 | PharmaLeaks: Understanding the Business of Online Pharmaceutical Affiliate Programs
Damon McCoy, Andreas Pitsillidis, Grant Jordan, Nicholas Weaver, Christian Kreibich, Brian Krebs, Geoffrey M. Voelker, Stefan Savage, Kirill Levchenko |
USENIX Security Symposium | 7 |
| 2011 | Cloak and dagger: dynamics of web search cloakingabstractCloaking is a common 'bait-and-switch' technique used to hide the true nature of a Web site by delivering blatantly different semantic content to different user segments. It is often used in search engine optimization (SEO) to obtain user traffic illegitimately for scams. In this paper, we measure and characterize the prevalence of cloaking on different search engines, how this behavior changes for targeted versus untargeted advertising and ultimately the response to site cloaking by search engine providers. Using a custom crawler, called Dagger, we track both popular search terms (e.g., as identified by Google, Alexa and Twitter) and targeted keywords (focused on pharmaceutical products) for over five months, identifying when distinct results were provided to crawlers and browsers. We further track the lifetime of cloaked search results as well as the sites they point to, demonstrating that cloakers can expect to maintain their pages in search results for several days on popular search engines and maintain the pages themselves for longer still. David Y. Wang, Stefan Savage, Geoffrey M. Voelker |
CCS | 3 |
| 2011 | An analysis of underground forumsabstractUnderground forums, where participants exchange information on abusive tactics and engage in the sale of illegal goods and services, are a form of online social network (OSN). However, unlike traditional OSNs such as Facebook, in underground forums the pattern of communications does not simply encode pre-existing social relationships, but instead captures the dynamic trust relationships forged between mutually distrustful parties. In this paper, we empirically characterize six different underground forums --- BlackHatWorld, Carders, HackSector, HackE1ite, Freehack, and L33tCrew --- examining the properties of the social networks formed within, the content of the goods and services being exchanged, and lastly, how individuals gain and lose trust in this setting. Marti Motoyama, Damon McCoy, Kirill Levchenko, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 5 |
| 2011 | Dissemination in opportunistic mobile ad-hoc networks: The power of the crowdabstractOpportunistic ad-hoc communication enables portable devices such as smartphones to effectively exchange information, taking advantage of their mobility and locality. The nature of human interaction makes information dissemination using such networks challenging. We use three different experimental traces to study fundamental properties of human interactions. We break our traces down in multiple areas and classify mobile users in each area according to their social behavior: Socials are devices that show up frequently or periodically, while Vagabonds represent the rest of the population. We find that in most cases the majority of the population consists of Vagabonds. We evaluate the relative role of these two groups of users in data dissemination. Surprisingly, we observe that under certain circumstances, which appear to be common in real life situations, the effectiveness of dissemination predominantly depends on the number of users in each class rather than their social behavior, contradicting some of the previous observations. We validate and extend the findings of our experimental study through a mathematical analysis. Gjergji Zyba, Geoffrey M. Voelker, Stratis Ioannidis, Christophe Diot |
INFOCOM | 2 |
| 2011 | DefenestraTor: Throwing Out Windows in Tor
Mashael Al Sabah, Kevin S. Bauer, Ian Goldberg 0001, Dirk Grunwald, Damon McCoy, Stefan Savage, Geoffrey M. Voelker |
PETS | 7 |
| 2011 | Click Trajectories: End-to-End Analysis of the Spam Value ChainabstractSpam-based advertising is a business. While it has engendered both widespread antipathy and a multi-billion dollar anti-spam industry, it continues to exist because it fuels a profitable enterprise. We lack, however, a solid understanding of this enterprise's full structure, and thus most anti-Spam interventions focus on only one facet of the overall spam value chain (e.g., spam filtering, URL blacklisting, site takedown).In this paper we present a holistic analysis that quantifies the full set of resources employed to monetize spam email -- including naming, hosting, payment and fulfillment -- using extensive measurements of three months of diverse spam data, broad crawling of naming and hosting infrastructures, and over 100 purchases from spam-advertised sites. We relate these resources to the organizations who administer them and then use this data to characterize the relative prospects for defensive interventions at each link in the spam value chain. In particular, we provide the first strong evidence of payment bottlenecks in the spam value chain, 95% of spam-advertised pharmaceutical, replica and software products are monetized using merchant services from just a handful of banks. Kirill Levchenko, Andreas Pitsillidis, Neha Chachra, Brandon Enright, Márk Félegyházi, Chris Grier, Tristan Halvorson, Chris Kanich, Christian Kreibich, Damon McCoy, Nicholas Weaver, Vern Paxson, Geoffrey M. Voelker, Stefan Savage |
IEEE Symposium on Security and Privacy | 14 |
| 2011 | Show Me the Money: Characterizing Spam-advertised Revenue
Chris Kanich, Nicholas Weaver, Damon McCoy, Tristan Halvorson, Christian Kreibich, Kirill Levchenko, Vern Paxson, Geoffrey M. Voelker, Stefan Savage |
USENIX Security Symposium | 8 |
| 2011 | Dirty Jobs: The Role of Freelance Labor in Web Service Abuse
Marti Motoyama, Damon McCoy, Kirill Levchenko, Stefan Savage, Geoffrey M. Voelker |
USENIX Security Symposium | 5 |
| 2011 | Learning to detect malicious URLsabstractMalicious Web sites are a cornerstone of Internet criminal activities. The dangers of these sites have created a demand for safeguards that protect end-users from visiting them. This article explores how to detect malicious Web sites from the lexical and host-based features of their URLs. We show that this problem lends itself naturally to modern algorithms for online learning. Online algorithms not only process large numbers of URLs more efficiently than batch algorithms, they also adapt more quickly to new features in the continuously evolving distribution of malicious URLs. We develop a real-time system for gathering URL features and pair it with a real-time feed of labeled URLs from a large Web mail provider. From these features and labels, we are able to train an online classifier that detects malicious Web sites with 99% accuracy over a balanced dataset. Justin Ma, Lawrence K. Saul, Stefan Savage, Geoffrey M. Voelker |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | DieCast: Testing Distributed Systems with an Accurate Scale ModelabstractLarge-scale network services can consist of tens of thousands of machines running thousands of unique software configurations spread across hundreds of physical networks. Testing such services for complex performance problems and configuration errors remains a difficult problem. Existing testing techniques, such as simulation or running smaller instances of a service, have limitations in predicting overall service behavior at such scales. Testing large services should ideally be done at the same scale and configuration as the target deployment, which can be technically and economically infeasible. We present DieCast , an approach to scaling network services in which we multiplex all of the nodes in a given service configuration as virtual machines across a much smaller number of physical machines in a test harness. We show how to accurately scale CPU, network, and disk to provide the illusion that each VM matches a machine in the original service in terms of both available computing resources and communication behavior. We present the architecture and evaluation of a system we built to support such experimentation and discuss its limitations. We show that for a variety of services---including a commercial high-performance cluster-based file system---and resource utilization levels, DieCast matches the behavior of the original service while using a fraction of the physical resources. Diwaker Gupta, Kashi Venkatesh Vishwanath, Marvin McNett, Amin Vahdat, Ken Yocum, Alex C. Snoeren, Geoffrey M. Voelker |
ACM Trans. Comput. Syst. | 7 |
| 2010 | How to tell an airport from a home: techniques and applicationsabstractToday's Internet services increasingly use IP-based geolocation to specialize the content and service provisioning for each user. However, these systems focus almost exclusively on the current position of users and do not attempt to infer or exploit any qualitative context about the location's relationship with the user (e.g., is the user at home? on a business trip?). This paper develops such a context by profiling the usage patterns of IP address ranges, relying on known user and machine identifiers to track accesses over time. Our preliminary results suggest that rough location categories such as residences, workplaces, and travel venues can be accurately inferred, enabling a range of potential applications from demographic analyses to ad specialization and security improvements. Andreas Pitsillidis, Yinglian Xie, Fang Yu 0002, Martín Abadi, Geoffrey M. Voelker, Stefan Savage |
HotNets | 5 |
| 2010 | Beyond heuristics: learning to classify vulnerabilities and predict exploitsabstractThe security demands on modern system administration are enormous and getting worse. Chief among these demands, administrators must monitor the continual ongoing disclosure of software vulnerabilities that have the potential to compromise their systems in some way. Such vulnerabilities include buffer overflow errors, improperly validated inputs, and other unanticipated attack modalities. In 2008, over 7,400 new vulnerabilities were disclosed--well over 100 per week. While no enterprise is affected by all of these disclosures, administrators commonly face many outstanding vulnerabilities across the software systems they manage. Vulnerabilities can be addressed by patches, reconfigurations, and other workarounds; however, these actions may incur down-time or unforeseen side-effects. Thus, a key question for systems administrators is which vulnerabilities to prioritize. From publicly available databases that document past vulnerabilities, we show how to train classifiers that predict whether and how soon a vulnerability is likely to be exploited. As input, our classifiers operate on high dimensional feature vectors that we extract from the text fields, time stamps, cross references, and other entries in existing vulnerability disclosure reports. Compared to current industry-standard heuristics based on expert knowledge and static formulas, our classifiers predict much more accurately whether and how soon individual vulnerabilities are likely to be exploited. Mehran Bozorgi, Lawrence K. Saul, Stefan Savage, Geoffrey M. Voelker |
KDD | 4 |
| 2010 | Botnet Judo: Fighting Spam with Itself
Andreas Pitsillidis, Kirill Levchenko, Christian Kreibich, Chris Kanich, Geoffrey M. Voelker, Vern Paxson, Nicholas Weaver, Stefan Savage |
NDSS | 5 |
| 2010 | Second life: a social network of humans and botsabstractSecond Life (SL) is a virtual world where people interact and socialize through virtual avatars. Avatars behave similarly to their human counterparts in real life and naturally define a social network. However, not only human-controlled avatars participate in the social network. Automated avatars called bots are common, difficult to identify and, when malicious, can severely detract from the user experience of SL. In this paper we study the SL social network and the role of bots within it. Using traces of avatars in a popular SL region, we analyze the social graph formed by avatar interactions. We find that it resembles natural networks more than other online social networks, and that bots have a fundamental impact on the SL social network. Finally, we propose a bot detection strategy based on the importance of the social connections of avatars in the social graph. Matteo Varvello, Geoffrey M. Voelker |
NOSSDAV | 2 |
| 2010 | Re: CAPTCHAs-Understanding CAPTCHA-Solving Services in an Economic Context
Marti Motoyama, Kirill Levchenko, Chris Kanich, Damon McCoy, Geoffrey M. Voelker, Stefan Savage |
USENIX Security Symposium | 5 |
| 2010 | Neon: system support for derived data managementabstractModern organizations face increasingly complex information management requirements. A combination of commercial needs, legal liability and regulatory imperatives has created a patchwork of mandated policies. Among these, personally identifying customer records must be carefully access-controlled, sensitive files must be encrypted on mobile computers to guard against physical theft, and intellectual property must be protected from both exposure and "poisoning." However, enforcing such policies can be quite difficult in practice since users routinely share data over networks and derive new files from these inputs--incidentally laundering any policy restrictions. In this paper, we describe a virtual machine monitor system called Neon that transparently labels derived data using byte-level "tints" and tracks these labels end to end across commodity applications, operating systems and networks. Our goal with Neon is to explore the viability and utility of transparent information flow tracking within conventional networked systems when used in the manner in which they were intended. We demonstrate that this mechanism allows the enforcement of a variety of data management policies, including data-dependent confinement, mandatory I/O encryption, and intellectual property management. Qing Zhang 0012, John McCullough, Justin Ma, Nabil Schear, Michael Vrable, Amin Vahdat, Alex C. Snoeren, Geoffrey M. Voelker, Stefan Savage |
VEE | 8 |
| 2010 | Usage Patterns in an Urban WiFi NetworkabstractWhile WiFi was initially designed as a local-area access network, mesh networking technologies have led to increasingly expansive deployments of WiFi networks. In urban environments, the WiFi mesh frequently supplements a number of existing access technologies, including wired broadband networks, 3G cellular, and commercial WiFi hotspots. It is an open question what role citywide WiFi deployments play in the increasingly diverse access network spectrum. We study the usage of the Google WiFi network deployed in Mountain View, CA, and find that usage naturally falls into three classes based almost entirely on client device type, which we divide into traditional laptop users, fixed-location access devices, and PDA-like smartphone devices. Moreover, each of these classes of use has significant geographic locality, following the distribution of residential, commercial, and transportation areas of the city. When comparing the network usage of each device class, we find a diverse set of mobility patterns that map well to the archetypal use cases for traditional access technologies. To help place our results in context, we also provide key performance measurements of the mesh backbone and, where possible, compare them to those of previously studied urban mesh networks. Mikhail Afanasyev, Tsuwei Chen, Geoffrey M. Voelker, Alex C. Snoeren |
IEEE/ACM Trans. Netw. | 3 |
| 2009 | Cumulus: Filesystem Backup to the Cloud
Michael Vrable, Stefan Savage, Geoffrey M. Voelker |
FAST | 3 |
| 2009 | Identifying suspicious URLs: an application of large-scale online learningabstractThis paper explores online learning approaches for detecting malicious Web sites (those involved in criminal scams) using lexical and host-based features of the associated URLs. We show that this application is particularly appropriate for online algorithms as the size of the training data is larger than can be efficiently processed in batch and because the distribution of features that typify malicious URLs is changing continuously. Using a real-time system we developed for gathering URL features, combined with a real-time source of labeled URLs from a large Web mail provider, we demonstrate that recently-developed online algorithms can be as accurate as batch techniques, achieving classification accuracies up to 99% over a balanced data set. Justin Ma, Lawrence K. Saul, Stefan Savage, Geoffrey M. Voelker |
ICML | 4 |
| 2009 | Defending Mobile Phones from Proximity MalwareabstractAs mobile phones increasingly become the target of propagating malware, their use of direct pair-wise communication mechanisms, such as Bluetooth and WiFi, pose considerable challenges to malware detection and mitigation. Unlike malware that propagates using the network, where the provider can employ centralized defenses, proximity malware can propagate in an entirely distributed fashion. In this paper we consider the dynamics of mobile phone malware that propagates by proximity contact, and we evaluate potential defenses against it. Defending against proximity malware is particularly challenging since it is difficult to piece together global dynamics from just pair-wise device interactions. Whereas traditional network defenses depend upon observing aggregated network activity to detect correlated or anomalous behavior, proximity malware detection must begin at the device. As a result, we explore three strategies for detecting and mitigating proximity malware that span the spectrum from simple local detection to a globally coordinated defense. Using insight from a combination of real-world traces, analytic epidemic models, and synthetic mobility models, we simulate proximity malware propagation and defense at the scale of a university campus. We find that local proximity-based dissemination of signatures can limit malware propagation. Globally coordinated strategies with broadcast dissemination are substantially more effective, but rely upon more demanding infrastructure within the provider. Gjergji Zyba, Geoffrey M. Voelker, Michael Liljenstam, András Méhes, Per Johansson |
INFOCOM | 2 |
| 2009 | Beyond blacklists: learning to detect malicious web sites from suspicious URLsabstractMalicious Web sites are a cornerstone of Internet criminal activities. As a result, there has been broad interest in developing systems to prevent the end user from visiting such sites. In this paper, we describe an approach to this problem based on automated URL classification, using statistical methods to discover the tell-tale lexical and host-based properties of malicious Web site URLs. These methods are able to learn highly predictive models by extracting and automatically analyzing tens of thousands of features potentially indicative of suspicious URLs. The resulting classifiers obtain 95-99% accuracy, detecting large numbers of malicious Web sites from their URLs, with only modest false positives. Justin Ma, Lawrence K. Saul, Stefan Savage, Geoffrey M. Voelker |
KDD | 4 |
| 2009 | SAM: enabling practical spatial multiple access in wireless LANabstractSpatial multiple access holds the promise to boost the capacity of wireless networks when an access point has multiple antennas. Due to the asynchronous and uncontrolled nature of wireless LANs, conventional MIMO technology does not work efficiently when concurrent transmissions from multiple stations are uncoordinated. In this paper, we present the design and implementation of a crosslayer system, called SAM, that addresses the challenges of enabling spatial multiple access for multiple devices in a random access network like WLAN. SAM uses a chain-decoding technique to reliably recover the channel parameters for each device, and iteratively decode concurrent frames with misaligned symbol timings and frequency offsets. We propose a new MAC protocol, called CCMA, to enable concurrent transmissions by different mobile stations while remaining backward compatible with 802.11. Finally, we implement the PHY and MAC layer of SAM using the Sora high-performance software radio platform. Our evaluation results under real wireless conditions show that SAM can improve network uplink throughput by 70% with two antennas over 802.11. Ji Fang, Wei Wang 0002, Jiansong Zhang 0001, Mi Chen, Geoffrey M. Voelker |
MobiCom | 7 |
| 2009 | NetPrints: Diagnosing Home Network Misconfigurations Using Shared Knowledge
Bhavish Agarwal, Ranjita Bhagwan, Tathagata Das, Siddharth Eswaran, Venkat N. Padmanabhan, Geoffrey M. Voelker |
NSDI | 6 |
| 2009 | Sora: High Performance Software Radio Using General Purpose Multi-core Processors
Jiansong Zhang 0001, Ji Fang, Yusheng Ye, Yongguang Zhang, Wei Wang 0002, Geoffrey M. Voelker |
NSDI | 10 |
| 2009 | MPIWiz: subgroup reproducible replay of mpi applicationsabstractMessage Passing Interface (MPI) is a widely used standard for managing coarse-grained concurrency on distributed computers. Debugging parallel MPI applications, however, has always been a particularly challenging task due to their high degree of concurrent execution and non-deterministic behavior. Deterministic replay is a potentially powerful technique for addressing these challenges, with existing MPI replay tools adopting either data-replay or order-replay approaches. Unfortunately, each approach has its tradeoffs. Data-replay generates substantial log sizes by recording every communication message. Order-replay generates small logs, but requires all processes to be replayed together. We believe that these drawbacks are the primary reasons that inhibit the wide adoption of deterministic replay as the critical enabler of cyclic debugging of MPI applications. Ruini Xue, Xuezheng Liu, Ming Wu 0007, Zheng Zhang 0001, Geoffrey M. Voelker |
PPoPP | 8 |
| 2009 | Cumulus: Filesystem backup to the cloudabstractCumulus is a system for efficiently implementing filesystem backups over the Internet, specifically designed under a thin cloud assumption—that the remote datacenter storing the backups does not provide any special backup services, but only a least-common-denominator storage interface. Cumulus aggregates data from small files for storage and uses LFS-inspired segment cleaning to maintain storage efficiency. While Cumulus can use virtually any storage service, we show its efficiency is comparable to integrated approaches. Michael Vrable, Stefan Savage, Geoffrey M. Voelker |
ACM Trans. Storage | 3 |
| 2008 | Spamalytics: an empirical analysis of spam marketing conversionabstractThe of spam--the probability that an unsolicited e-mail will ultimately elicit a sale--underlies the entire spam value proposition. However, our understanding of this critical behavior is quite limited, and the literature lacks any quantitative study concerning its true value. In this paper we present a methodology for measuring the conversion rate of spam. Using a parasitic infiltration of an existing botnet's infrastructure, we analyze two spam campaigns: one designed to propagate a malware Trojan, the other marketing on-line pharmaceuticals. For nearly a half billion spam e-mails we identify the number that are successfully delivered, the number that pass through popular anti-spam filters, the number that elicit user visits to the advertised sites, and the number of sales and infections produced. Chris Kanich, Christian Kreibich, Kirill Levchenko, Brandon Enright, Geoffrey M. Voelker, Vern Paxson, Stefan Savage |
CCS | 5 |
| 2008 | Analysis of a mixed-use urban wifi network: when metropolitan becomes neapolitanabstractWhile WiFi was initially designed as a local-area access network, mesh networking technologies have led to increasingly expansive deployments of WiFi networks. In urban environments, the WiFi mesh frequently supplements a number of existing access technologies, including wired broadband networks, 3G cellular, and commercial WiFi hotspots. It is an open question what role city-wide WiFi deployments play in the increasingly diverse access network spectrum. We study the usage of the Google WiFi network deployed in Mountain View, California, and find that usage naturally falls into three classes, based almost entirely on client device type. Moreover, each of these classes of use has significant geographic locality, following the distribution of residential, commercial, and transportation areas of the city. Finally, we find a diverse set of mobility patterns that map well to the archetypal use cases for traditional access technologies. Mikhail Afanasyev, Tsuwei Chen, Geoffrey M. Voelker, Alex C. Snoeren |
Internet Measurement Conference | 3 |
| 2008 | Difference Engine: Harnessing Memory Redundancy in Virtual Machines
Diwaker Gupta, Michael Vrable, Stefan Savage, Alex C. Snoeren, George Varghese, Geoffrey M. Voelker, Amin Vahdat |
OSDI | 7 |
| 2008 | Xl: an efficient network routing algorithmabstractIn this paper, we present a new link-state routing algorithm called Approximate Link state (XL) aimed at increasing routing efficiency by suppressing updates from parts of the network. We prove that three simple criteria for update propagation are sufficient to guarantee soundness, completeness and bounded optimality for any such algorithm. We show, via simulation, that XL significantly outperforms standard link-state and distance vector algorithms - in some cases reducing overhead by more than an order of magnitude - while having negligible impact on path length. Finally, we argue that existing link-state protocols, such as OSPF, can incorporate XL routing in a backwards compatible and incrementally deployable fashion. Kirill Levchenko, Geoffrey M. Voelker, Ramamohan Paturi, Stefan Savage |
SIGCOMM | 2 |
| 2008 | Dual Frame Motion Compensation With Uneven Quality AssignmentabstractVideo codecs that use motion compensation have shown PSNR gains from the use of multiple frame prediction, in which more than one past reference frame is available for motion estimation. In dual frame motion compensation, one short-term reference frame and one long-term reference frame are available for prediction. In this paper, we explore using dual frame motion compensation in two contexts. We first show that using a single fixed long-term reference frame in the context of a rate switching network can enhance video quality. Next, by periodically creating high-quality long-term reference frames, we show that the performance is superior to a standard dual frame technique that has the same average rate but no high-quality frames. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Usher: An Extensible Framework for Managing Clusters of Virtual Machines
Marvin McNett, Diwaker Gupta, Amin Vahdat, Geoffrey M. Voelker |
LISA | 4 |
| 2007 | Automating cross-layer diagnosis of enterprise wireless networksabstractModern enterprise networks are of sufficient complexity that even simple faults can be difficult to diagnose - let alone transient outages or service degradations. Nowhere is this problem more apparent than in the 802.11-based wireless access networks now ubiquitous in the enterprise. In addition to the myriad complexities of the wired network, wireless networks face the additional challenges of shared spectrum, user mobility and authentication management. Not surprisingly, few organizations have the expertise, data or tools to decompose the underlying problems and interactions responsible for transient outages or performance degradations. In this paper, we present a set of modeling techniques for automatically characterizing the source of such problems. In particular, we focus on data transfer delays unique to 802.11 networks - media access dynamics and mobility management latency. Through a combination of measurement, inference and modeling we reconstruct sources of delay - from the physical layer to the transport - layer as well as the interactions among them. We demonstrate our approach using comprehensive traces of wireless activity in the UCSD Computer Science building. Yuchung Cheng, Mikhail Afanasyev, Patrick Verkaik, Péter Benkö, Jennifer Chiang, Alex C. Snoeren, Stefan Savage, Geoffrey M. Voelker |
SIGCOMM | 8 |
| 2007 | Spamscatter: Characterizing Internet Scam Hosting Infrastructure
David S. Anderson, Chris Fleizach, Stefan Savage, Geoffrey M. Voelker |
USENIX Security Symposium | 4 |
| 2006 | Finding diversity in remote code injection exploitsabstractRemote code injection exploits inflict a significant societal cost, and an active underground economy has grown up around these continually evolving attacks. We present a methodology for inferring the phylogeny, or evolutionary tree, of such exploits. We have applied this methodology to traffic captured at several vantage points, and we demonstrate that our methodology is robust to the observed polymorphism. Our techniques revealed non-trivial code sharing among different exploit families, and the resulting phylogenies accurately captured the subtle variations among exploits within each family. Thus, we believe our methodology and results are a helpful step to better understanding the evolution of remote code injection exploits on the Internet. Justin Ma, John Dunagan, Helen J. Wang, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 5 |
| 2006 | Unexpected means of protocol inferenceabstractNetwork managers are inevitably called upon to associate network traffic with particular applications. Indeed, this operation is critical for a wide range of management functions ranging from debugging and security to analytics and policy support. Traditionally, managers have relied on application adherence to a well established global port mapping: Web traffic on port 80, mail traffic on port 25 and so on. However, a range of factors - including firewall port blocking, tunneling, dynamic port allocation, and a bloom of new distributed applications - has weakened the value of this approach. We analyze three alternative mechanisms using statistical and structural content models for automatically identifying traffic that uses the same application-layer protocol, relying solely on flow content. In this manner, known applications may be identified regardless of port number, while traffic from one unknown application will be identified as distinct from another. We evaluate each mechanism's classification performance using real-world traffic traces from multiple sites. Justin Ma, Kirill Levchenko, Christian Kreibich, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 5 |
| 2006 | To Infinity and Beyond: Time-Warped Network Emulation
Diwaker Gupta, Ken Yocum, Marvin McNett, Alex C. Snoeren, Amin Vahdat, Geoffrey M. Voelker |
NSDI | 6 |
| 2006 | Jigsaw: solving the puzzle of enterprise 802.11 analysisabstractThe combination of unlicensed spectrum, cheap wireless interfaces and the inherent convenience of untethered computing have made 802.11 based networks ubiquitous in the enterprise. Modern universities, corporate campuses and government offices routinely de-ploy scores of access points to blanket their sites with wireless Internet access. However, while the fine-grained behavior of the 802.11 protocol itself has been well studied, our understanding of how large 802.11 networks behave in their full empirical complex-ity is surprisingly limited. In this paper, we present a system called Jigsaw that uses multiple monitors to provide a single unified view of all physical, link, network and transport-layer activity on an 802.11 network. To drive this analysis, we have deployed an infrastructure of over 150 radio monitors that simultaneously capture all 802.11b and 802.11g activity in a large university building (1M+ cubic feet). We describe the challenges posed by both the scale and ambiguity inherent in such an architecture, and explain the algorithms and inference techniques we developed to address them. Finally, using a 24-hour distributed trace containing more than 1.5 billion events, we use Jigsaw's global cross-layer viewpoint to isolate performance artifacts, both explicit, such as management inefficiencies, and implicit, such as co-channel interference. We believe this is the first analysis combining this scale and level of detail for a production 802.11 network. Yuchung Cheng, John Bellardo, Péter Benkö, Alex C. Snoeren, Geoffrey M. Voelker, Stefan Savage |
SIGCOMM | 5 |
| 2006 | Maximizing data locality in distributed systems
Fan Chung Graham, Ronald L. Graham, Ranjita Bhagwan, Stefan Savage, Geoffrey M. Voelker |
J. Comput. Syst. Sci. | 5 |
| 2006 | Inferring Internet denial-of-service activityabstractIn this article, we seek to address a simple question: “How prevalent are denial-of-service attacks in the Internet?” Our motivation is to quantitatively understand the nature of the current threat as well as to enable longer-term analyses of trends and recurring patterns of attacks. We present a new technique, called “backscatter analysis,” that provides a conservative estimate of worldwide denial-of-service activity. We use this approach on 22 traces (each covering a week or more) gathered over three years from 2001 through 2004. Across this corpus we quantitatively assess the number, duration, and focus of attacks, and qualitatively characterize their behavior. In total, we observed over 68,000 attacks directed at over 34,000 distinct victim IP addresses---ranging from well-known e-commerce companies such as Amazon and Hotmail to small foreign ISPs and dial-up connections. We believe our technique is the first to provide quantitative estimates of Internet-wide denial-of-service activity and that this article describes the most comprehensive public measurements of such activity to date. David Moore 0001, Colleen Shannon, Douglas J. Brown, Geoffrey M. Voelker, Stefan Savage |
ACM Trans. Comput. Syst. | 4 |
| 2006 | Characterization of a Large Web Site Population with Implications for Content Delivery
Leeann Bent, Michael Rabinovich, Geoffrey M. Voelker |
World Wide Web | 3 |
| 2005 | Error Concealment for Dual Frame Video Coding with Uneven QualityabstractWhen losses occur in a transmission of compressed video, the decoder can attempt to conceal the loss by using spatial or temporal methods to estimate the missing macroblocks. We consider a multi-frame error concealment approach which exploits the uneven quality in the two reference frames to provide good concealment candidates. A binary decision tree is used to decide among various error concealment choices. The uneven quality of the reference frames provides an advantage for error concealment. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
DCC | 3 |
| 2005 | To infinity and beyond: time warped network emulationabstractThis work explores the viability and benefits of time dilation - providing the illusion to an operating system and its applications that time is passing at a rate different from real time. For example, we may wish to convince a system that for every 10 seconds of wall clock time, only one second of time passes in the host's dilated time frame. This enables external stimuli to appear to take place at higher rates than would be physically possible. For example, a host dilated by a factor of 10 receiving data from a network interface at a real rate of 1-Gbps believes it is receiving data at 10-Gbps. Diwaker Gupta, Ken Yocum, Marvin McNett, Alex C. Snoeren, Amin Vahdat, Geoffrey M. Voelker |
SOSP | 6 |
| 2005 | Scalability, fidelity, and containment in the potemkin virtual honeyfarmabstractThe rapid evolution of large-scale worms, viruses and bot-nets have made Internet malware a pressing concern. Such infections are at the root of modern scourges including DDoS extortion, on-line identity theft, SPAM, phishing, and piracy. However, the most widely used tools for gathering intelligence on new malware -- network honeypots -- have forced investigators to choose between monitoring activity at a large scale or capturing behavior with high fidelity. In this paper, we describe an approach to minimize this tension and improve honeypot scalability by up to six orders of magnitude while still closely emulating the execution behavior of individual Internet hosts. We have built a prototype honeyfarm system, called Potemkin, that exploits virtual machines, aggressive memory sharing, and late binding of resources to achieve this goal. While still an immature implementation, Potemkin has emulated over 64,000 Internet honeypots in live test runs, using only a handful of physical servers. Michael Vrable, Justin Ma, Jay Chen, David Moore 0001, Erik Vandekieft, Alex C. Snoeren, Geoffrey M. Voelker, Stefan Savage |
SOSP | 7 |
| 2005 | Surviving Internet Catastrophes
Flavio Paiva Junqueira, Ranjita Bhagwan, Alejandro Hevia, Keith Marzullo, Geoffrey M. Voelker |
USENIX ATC, General Track | 5 |
| 2005 | Wireless Hotspots: Current Challenges and Future Directions
Anand Balachandran, Geoffrey M. Voelker, Paramvir Bahl |
Mob. Networks Appl. | 2 |
| 2004 | Dual Frame Motion Compensation with Uneven Quality AssignmentabstractVideo codecs that use motion compensation have shown PSNR gains from the use of multiple frame prediction, in which more than one past reference frame is available for motion estimation. In dual frame motion compensation, one short-term reference frame and one long-term reference frame are available for prediction. In this paper, we propose a dual frame motion compensation technique that allocates bits unevenly among frames to periodically create a high-quality frame that serves as the long-term reference frame for some time. By modifying an MPEG-4 encoder to use this technique on a set of video sequences, we show that it outperforms a normal dual frame motion compensation scheme in which the long-term reference frames are regular frames that are not allocated any extra rate. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
Data Compression Conference | 3 |
| 2004 | Total Recall: System Support for Automated Availability Management
Ranjita Bhagwan, Kiran Tati, Yuchung Cheng, Stefan Savage, Geoffrey M. Voelker |
NSDI | 5 |
| 2004 | Network sensitivity to hot-potato disruptionsabstractHot-potato routing is a mechanism employed when there are multiple (equally good) interdomain routes available for a given destination. In this scenario, the Border Gateway Protocol (BGP) selects the interdomain route associated with the closest egress point based upon intradomain path costs. Consequently, intradomain routing changes can impact interdomain routing and cause abrupt swings of external routes, which we call hot-potato disruptions. Recent work has shown that hot-potato disruptions can have a substantial impact on large ISP backbones and thereby jeopardize the network robustness. As a result, there is a need for guidelines and tools to assist in the design of networks that minimize hot-potato disruptions. However, developing these tools is challenging due to the complex and subtle nature of the interactions between exterior and interior routing. In this paper, we address these challenges using an analytic model of hot-potato routing that incorporates metrics to evaluate network sensitivity to hot-potato disruptions. We then present a methodology for computing these metrics using measurements of real ISP networks. We demonstrate the utility of our model by analyzing the sensitivity of a large AS in a tier~1 ISP network. Renata Teixeira, Aman Shaikh, Timothy G. Griffin, Geoffrey M. Voelker |
SIGCOMM | 4 |
| 2004 | Monkey See, Monkey Do: A Tool for TCP Tracing and Replaying
Yuchung Cheng, Urs Hölzle, Neal Cardwell, Stefan Savage, Geoffrey M. Voelker |
USENIX ATC, General Track | 5 |
| 2004 | Characterization of a large web site population with implications for content deliveryabstractThis paper presents a systematic study of the properties of a large number of Web sites hosted by a major ISP. To our knowledge, ours is the first comprehensive study of a large server farm that contains thousands of commercial Web sites. We also perform a simulation analysis to estimate potential performance benefits of content delivery networks (CDNs) for these Web sites, and validate our analysis for several sites by replaying our trace through a real cache. We make several interesting observations about the current usage of Web technologies and Web site performance characteristics. First, compared with previous client workload studies, the Web server farm workload contains a much higher degree of uncacheable responses and responses that require mandatory cache validations. A significant reason for this is that cookie use is prevalent among our population, especially among more popular sites. We found an indication of widespread indiscriminate usage of cookies, which unnecessarily impedes the use of many content delivery optimizations. We also found that most Web sites do not utilize the cache-control features of the HTTP 1.1 protocol, resulting in suboptimal performance. Moreover, the implicit expiration time in client caches for responses is strongly constrained by the maximum values allowed in the Squid proxy. Thus, supplying explicit expiration information would significantly improve Web sites' Leeann Bent, Michael Rabinovich, Geoffrey M. Voelker |
WWW | 3 |
| 2003 | The Phoenix Recovery System: Rebuilding from the Ashes of an Internet Catastrophe
Flavio Paiva Junqueira, Ranjita Bhagwan, Keith Marzullo, Stefan Savage, Geoffrey M. Voelker |
HotOS | 5 |
| 2003 | In search of path diversity in ISP networksabstractInternet Service Providers (ISPs) can exploit path diversity to balance load and improve robustness. Unfortunately, it is difficult to evaluate the potential impact of these approaches without routing and topological data, which are confidential. In this paper, we characterize path diversity in the real Sprint network. We then characterize path diversity in ISP topologies inferred using the Rocketfuel tool. Comparing the real Sprint topology to the one inferred by Rocketfuel, we find that the Rocketfuel topology has significantly higher apparent path diversity.(As a metric, path diversity is particularly sensitive to the presence of false or missing links, both of which are artifacts of active measurement techniques.) We evaluate heuristics that improve the accuracy of the inferred Rocketfuel topologies. Finally, we discuss limitations of active measurements techniques to capture topological properties such as path diversity. Renata Teixeira, Keith Marzullo, Stefan Savage, Geoffrey M. Voelker |
Internet Measurement Conference | 4 |
| 2003 | Internet Quarantine: Requirements for Containing Self-Propagating CodeabstractIt has been clear since 1988 that self-propagating code can quickly spread across a network by exploiting homogeneous security vulnerabilities. However, the last few years have seen a dramatic increase in the frequency and virulence of such "worm" outbreaks. For example, the Code-Red worm epidemics of 2001 infected hundreds of thousands of Internet hosts in a very short period - incurring enormous operational expense to track down, contain, and repair each infected machine. In response to this threat, there is considerable effort focused on developing technical means for detecting and containing worm infections before they can cause such damage. This paper does not propose a particular technology to address this problem, but instead focuses on a more basic question: How well will any such approach contain a worm epidemic on the Internet? We describe the design space of worm containment systems using three key parameters - reaction time, containment strategy and deployment scenario. Using a combination of analytic modeling and simulation, we describe how each of these design factors impacts the dynamics of a worm epidemic and, conversely, the minimum engineering requirements necessary to contain the spread of a given worm. While our analysis cannot provide definitive guidance for engineering defenses against all future threats, we demonstrate the lower bounds that any such system must exceed to be useful today. Unfortunately, our results suggest that there are significant technological and administrative gaps to be bridged before an effective defense can be provided in today's Internet. David Moore 0001, Colleen Shannon, Geoffrey M. Voelker, Stefan Savage |
INFOCOM | 3 |
| 2003 | Characterizing and measuring path diversity of internet topologiesabstractNo abstract available. Renata Teixeira, Keith Marzullo, Stefan Savage, Geoffrey M. Voelker |
SIGMETRICS | 4 |
| 2003 | End-to-end differentiation of congestion and wireless lossesabstractIn this paper, we explore end-to-end loss differentiation algorithms (LDAs) for use with congestion-sensitive video transport protocols for networks with either backbone or last-hop wireless links. As our basic video transport protocol, we use UDP in conjunction with a congestion control mechanism extended with an LDA. For congestion control, we use the TCP-Friendly Rate Control (TFRC) algorithm. We extend TFRC to use an LDA when a connection uses at least one wireless link in the path between the sender and receiver. We then evaluate various LDAs under different wireless network topologies, competing traffic, and fairness scenarios to determine their effectiveness. In addition to evaluating LDAs derived from previous work, we also propose and evaluate a new LDA, ZigZag, and a hybrid LDA, ZBS, that selects among base LDAs depending upon observed network conditions. We evaluate these LDAs via simulation, and find that no single base algorithm performs well across all topologies and competition. However, the hybrid algorithm performs well across topologies and competition, and in some cases exceeds the performance of the best base LDA for a given scenario. All of the LDAs are reasonably fair when competing with TCP, and their fairness among flows using the same LDA depends on the network topology. In general, ZigZag and the hybrid algorithm are the fairest among all LDAs. Song Cen, Pamela C. Cosman, Geoffrey M. Voelker |
IEEE/ACM Trans. Netw. | 3 |
| 2002 | Characterizing user behavior and network performance in a public wireless LANabstractThis paper presents and analyzes user behavior and network performance in a public-area wireless network using a workload captured at a well-attended ACM conference. The goals of our study are: (1) to extend our understanding of wireless user behavior and wireless network performance; (2) to characterize wireless users in terms of a parameterized model for use with analytic and simulation studies involving wireless LAN traffic; and (3) to apply our workload analysis results to issues in wireless network deployment, such as capacity planning, and potential network optimizations, such as algorithms for load balancing across multiple access points (APs) in a wireless network. Anand Balachandran, Geoffrey M. Voelker, Paramvir Bahl, P. Venkat Rangan |
SIGMETRICS | 2 |
| 2002 | Symbiotic jobscheduling with priorities for a simultaneous multithreading processorabstractSimultaneous Multithreading machines benefit from jobscheduling software that monitors how well coscheduled jobs share CPU resources, and coschedules jobs that interact well to make more efficient use of those resources. As a result, informed coscheduling can yield significant performance gains over naive schedulers. However, prior work on coscheduling focused on equal-priority job mixes, which is an unrealistic assumption for modern operating systems.This paper demonstrates that a scheduler for an SMT machine can both satisfy process priorities and symbiotically schedule low and high priority threads to increase system throughput. Naive priority schedulers dedicate the machine to high priority jobs to meet priority goals, and as a result decrease opportunities for increased performance from multithreading and coscheduling. More informed schedulers, however, can dynamically monitor the progress and resource utilization of jobs on the machine, and dynamically adjust the degree of multithreading to improve performance while still meeting priority goals.Using detailed simulation of an SMT architecture, we introduce and evaluate a series of five software and hardware-assisted priority schedulers. Overall, our results indicate that coscheduling priority jobs can significantly increase system throughput by as much as 40%, and that (1) the benefit depends upon the relative priority of the coscheduled jobs, and (2) more sophisticated schedulers are more effective when the differences in priorities are greatest. We show that our priority schedulers can decrease average turnaround times for a random jobmix by as much as 33%. Allan Snavely, Dean M. Tullsen, Geoffrey M. Voelker |
SIGMETRICS | 3 |
| 2001 | On the Placement of Web Server ReplicasabstractThere has been an increasing deployment of content distribution networks (CDNs) that offer hosting services to Web content providers. CDNs deploy a set of servers distributed throughout the Internet and replicate provider content across these servers for better performance and availability than centralized provider servers. Existing work on CDNs has primarily focused on techniques for efficiently redirecting user requests to appropriate CDN servers to reduce request latency and balance load. However, little attention has been given to the development of placement strategies for Web server replicas to further improve CDN performance. We explore the problem of Web server replica placement in detail. We develop several placement algorithms that use workload information, such as client latency and request rates, to make informed placement decisions. We then evaluate the placement algorithms using both synthetic and real network topologies, as well as Web server traces, and show that the placement of Web replicas is crucial to CDN performance. We also address a number of practical issues when using these algorithms, such as their sensitivity to imperfect knowledge about client workload and network topology, the stability of the input data, and methods for obtaining the input. Lili Qiu, Venkat N. Padmanabhan, Geoffrey M. Voelker |
INFOCOM | 3 |
| 2001 | Inferring Internet Denial-of-Service Activity
David Moore 0001, Geoffrey M. Voelker, Stefan Savage |
USENIX Security Symposium | 2 |
| 1999 | Potentials and Limitations of Fault-Based Markov Prefetching for Virtual Memory PagesabstractNo abstract available. Gretta Bartels, Anna R. Karlin, Darrell C. Anderson, Jeffrey S. Chase, Henry M. Levy, Geoffrey M. Voelker |
SIGMETRICS | 6 |
| 1999 | On the scale and performance of cooperative Web proxy cachingabstractWhile algorithms for cooperative proxy caching have been widely studied, little is understood about cooperativecaching performance in the large-scale World Wide Web environment. This paper uses both trace-based analysis and analytic modelling to show the potential advantages and drawbacks of inter-proxy cooperation. With our traces, we evaluate quantitatively the performance-improvement potential of cooperation between 200 small-organization proxies within a university environment, and between two largeorganization proxies handling 23,000 and 60,000 clients, respectively. With our model, we extend beyond these populations to project cooperative caching behavior in regions with millions of clients. Overall, we demonstrate that cooperative caching has performance benefits only within limited population bounds. We also use our model to examine the implications of future trends in Web-access behavior and traffic. 1 Introduction Cooperative caching -- the sharing and coordination of cache... Alec Wolman, Geoffrey M. Voelker, Nitin Sharma 0002, Neal Cardwell, Anna R. Karlin, Henry M. Levy |
SOSP | 2 |
| 1998 | Implementing Cooperative Prefetching and Caching in a Globally-Managed Memory SystemabstractThis paper presents cooperative prefetching and caching --- the use of network-wide global resources (memories, CPUs, and disks) to support prefetching and caching in the presence of hints of future demands. Cooperative prefetching and caching effectively unites disk-latency reduction techniques from three lines of research: prefetching algorithms, cluster-wide memory management, and parallel I/O. When used together, these techniques greatly increase the power of prefetching relative to a conventional (non-global-memory) system. We have designed and implemented PGMS, a cooperative prefetching and caching system, under the Digital Unix operating system running on a 1.28 Gb/sec Myrinet-connected cluster of DEC Alpha workstations. Our measurements and analysis show that by using available global resources, cooperative prefetching can obtain significant speedups for I/O-bound programs. For example, for a graphics rendering application, our system achieves a speedup of 4.9 over a non-prefetching version of the same program, and a 3.1-fold improvement over that program using local-disk prefetching alone. Geoffrey M. Voelker, Eric J. Anderson, Tracy Kimbrel, Michael J. Feeley, Jeffrey S. Chase, Anna R. Karlin, Henry M. Levy |
SIGMETRICS | 1 |
| 1997 | Managing Server Load in Global Memory SystemsabstractNew high-speed switched networks have reduced the latency of network page transfers significantly below that of local disk. This trend has led to the development of systems that use network-wide memory, or global memory, as a cache for virtual memory pages or file blocks. A crucial issue in the implementation of these global memory systems is the selection of the target nodes to receive replaced pages. Current systems use various forms of an approximate global LRU algorithm for making these selections. However, using age information alone can lead to suboptimal performance in two ways. First, workload characteristics can lead to uneven distributions of old pages across servers, causing increased contention delays. Second, the global memory traffic imposed on a node can degrade the performance of local jobs on that node.This paper studies the potential benefit and the potential harm of using load information, in addition to age information, in global memory replacement policies. Using an analytic queueing network model, we show the extent to which server load can degrade remote memory latency and how load balancing solves this problem. Load balancing requests can cause the system to deviate from the global LRU replacement policy, however. Using trace-driven simulation, we study the impact on application performance of deviating from the LRU replacement policy. We find that deviating from strict LRU, even significantly for some applications, does not affect application performance. Based upon these results, we conclude that global memory systems can gain substantial benefit from load balancing requests with little harm from suboptimal replacement decisions. Finally, we illustrate the use of the intuition gained from the model and simulation experiments by proposing a new family of algorithms that incorporate load considerations as well as age information in global memory replacement decisions. Geoffrey M. Voelker, Hervé A. Jamrozik, Mary K. Vernon, Henry M. Levy, Edward D. Lazowska |
SIGMETRICS | 1 |
| 1996 | Reducing Network Latency Using Subpages in a Global Memory EnvironmentabstractNew high-speed networks greatly encourage the use of network memory as a cache for virtual memory and file pages, thereby reducing the need for disk access. Because pages are the fundamental transfer and access units in remote memory systems, page size is a key performance factor. Recently, page sizes of modern processors have been increasing in order to provide more TLB coverage and amortize disk access costs. Unfortunately, for high-speed networks, small transfers are needed to provide low latency. This trend in page size is thus at odds with the use of network memory on high-speed networks.This paper studies the use of subpages as a means of reducing transfer size and latency in a remote-memory environment. Using trace-driven simulation, we show how and why subpages reduce latency and improve performance of programs using network memory. Our results show that memory-intensive applications execute up to 1.8 times faster when executing with 1K-byte subpages, when compared to the same applications using full 8K-byte pages in the global memory system. Those same applications using 1K-byte subpages execute up to 4 times faster than they would using the disk for backing store. Using a prototype implementation on the DEC Alpha and AN2 network, we demonstrate how subpages can reduce remote-memory fault time; e.g., our prototype is able to satisfy a fault on a 1K subpage stored in remote memory in 0.5 milliseconds, one third the time of a full page. Hervé A. Jamrozik, Michael J. Feeley, Geoffrey M. Voelker, James Evans II, Anna R. Karlin, Henry M. Levy, Mary K. Vernon |
ASPLOS | 3 |
| 1996 | The Structure and Performance of InterpretersabstractInterpreted languages have become increasingly popular due to demands for rapid program development, ease of use, portability, and safety. Beyond the general impression that they are "slow," however, little has been documented about the performance of interpreters as a class of applications.This paper examines interpreter performance by measuring and analyzing interpreters from both software and hardware perspectives. As examples, we measure the MIPSI, Java, Perl, and Tcl interpreters running an array of micro and macro benchmarks on a DEC Alpha platform. Our measurements of these interpreters relate performance to the complexity of the interpreter's virtual machine and demonstrate that native runtime libraries can play a key role in providing good performance. From an architectural perspective, we show that interpreter performance is primarily a function of the interpreter itself and is relatively independent of the application being interpreted. We also demonstrate that high-level interpreters' demands on processor resources are comparable to those of other complex compiled programs, such as gcc. We conclude that interpreters, as a class of applications, do not currently motivate special hardware support for increased performance. Theodore H. Romer, Dennis Lee 0001, Geoffrey M. Voelker, Alec Wolman, Wayne A. Wong, Jean-Loup Baer, Brian N. Bershad, Henry M. Levy |
ASPLOS | 3 |