VLDB 2026 Research / reviewers in the wild / expert
Daniel S. Berger
dblp:142/2630
· DBLP profile ↗
41ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-3911-1512ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 3 first-author · 16 since 2021Systems, architecture and hardware · 16 · 4 first-author · 12 since 2021Computer networks · 9 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Predictability in Heterogeneous MemoryabstractHeterogeneous memory combining DRAM and CXL exhibits variable performance, yet existing metrics correlate weakly with actual slowdown. We present CAMP, a principled framework for predicting CXL-induced slowdown. Our key insight is that a DRAM run (plus a CXL run for bandwidth-bound workloads) exposes the causal microarchitectural pressure points where CXL latency translates into additional processor stall cycles. CAMP captures these signals using 12 performance counters to analytically decompose slowdown into three orthogonal components: demand reads, cache/prefetching, and stores. CAMP also introduces a closed-form model for software-based weighted interleaving that predicts performance across DRAM--CXL ratios. Across 265 workloads on NUMA and three CXL devices, CAMP achieves 91--97% prediction accuracy within 10% absolute error. We demonstrate that these models enable practical system policies, including ''Best-shot'' interleaving and colocated workload placement, improving performance by up to 21% and 23% over existing tiering and colocation approaches. Jinshu Liu, Hanchen Xu, Daniel S. Berger, Marcos K. Aguilera, Huaicheng Li |
ASPLOS (2) | 3 |
| 2026 | From Lab to Fleet: Building and Deploying a Practical Rowhammer Defense in Cloud SoCs
Stefan Saroiu, Sujay Yadalam, Alec Wolman, Will Remaklus, Daniel S. Berger, Isaac H. Luna, Ishwar Agarwal, Jacob R. Lorch |
ISCA | 5 |
| 2026 | Octopus: Enhancing CXL Memory Pods via Sparse Topology
Yuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti, Shuwei Teng, Rodrigo Fonseca, Mark D. Hill, Daniel S. Berger |
NSDI | 7 |
| 2026 | CXL in Cloud Practice: Practical Lessons for Incrementally Scaling DeploymentabstractThis paper explores learnings from first-generation Compute Express Link (CXL) memory expansion to accelerate CXL’s journey to broad, robust use. While broad adoption will be a long journey similar to that of RDMA, we argue that the first step—CXL.mem expansion—is viable on today’s hardware. Through an end-to-end analysis, we revisit common showstoppers: we decompose memory access latency and show that CPU and DRAM internals, rather than the CXL protocol, dominate latency and variability, and we demonstrate how system slack absorbs link error rates above nominal specifications. Along the way, we distill practical guidance on device validation, monitoring, failure modes, security, and multi-tenant interference, and we outline a pragmatic adoption pathway: solidify robust expansion first, prototype micro-pooling next, and move to selective sharing as the ecosystem matures. Daniel S. Berger, Karthik Kumar, Midhul Vuppalapati, Chet Douglas, Jesse Sathre, Ian Robinson, Mark D. Hill |
IEEE Trans. Computers | 1 |
| 2025 | Systematic CXL Memory Characterization and Performance Analysis at ScaleabstractCompute Express Link (CXL) has emerged as a pivotal interconnect for memory expansion. Despite its potential, the performance implications of CXL across devices, latency regimes, processors, and workloads remain underexplored. We present Melody, a framework for systematic characterization and analysis of CXL memory performance. Melody builds on an extensive evaluation spanning 265 workloads, 4 real CXL devices, 7 latency levels, and 5 CPU platforms. Melody yields many insights: workload sensitivity to sub-μs CXL latencies (140-410ns), the first disclosure of CXL tail latencies, CPU tolerance to CXL latencies, a novel approach (SPA) for pinpointing CXL bottlenecks, and CPU prefetcher inefficiencies under CXL. Jinshu Liu, Hamid Hadian, Yuyue Wang 0001, Daniel S. Berger, Marie Nguyen, Xun Jian 0002, Sam H. Noh, Huaicheng Li |
ASPLOS (2) | 4 |
| 2025 | Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud PlatformsabstractCloud platforms remain underutilized despite multiple proposals to improve their utilization (e.g., disaggregation, harvesting, and oversubscription). Our characterization of the resource utilization of virtual machines (VMs) in Azure reveals that, while CPU is the main underutilized resource, we need to provide a solution to manage all resources holistically. We also observe that many VMs exhibit complementary temporal patterns, which can be leveraged to improve the oversubscription of underutilized resources. Benjamin Reidys, Pantea Zardoshti, Íñigo Goiri, Celine Irvene, Daniel S. Berger, Haoran Ma 0007, Kapil Arya, Eli Cortez, Taylor Stark, Eugene Bak, Mehmet Iyigun, Stanko Novakovic, Lisa Hsu, Karel Trueba, Abhisek Pan, Chetan Bansal, Saravan Rajmohan, Jian Huang 0006, Ricardo Bianchini |
ASPLOS (1) | 5 |
| 2025 | My CXL Pool Obviates Your PCIe SwitchabstractPooling PCIe devices across multiple hosts offers a promising solution to mitigate stranded I/O resources, enhance device utilization, address device failures, and reduce total cost of ownership. The only viable option today are PCIe switches, which decouple PCIe devices from hosts by connecting them through a hardware switch. However, the high cost and limited flexibility of PCIe switches hinder their widespread adoption beyond specialized datacenter use cases. Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti, Enrique Saurez, Jacob Nelson 0001, Antonis Psistakis, Joshua Fried, Asaf Cidon |
HotOS | 2 |
| 2025 | Enhancing Network Failure Mitigation with Performance-Aware Ranking
Pooria Namyar, Arvin Ghavidel, Daniel Crankshaw, Daniel S. Berger, Kevin Hsieh, Srikanth Kandula, Ramesh Govindan, Behnaz Arzani |
NSDI | 4 |
| 2025 | Oasis: Pooling PCIe Devices Over CXL to Boost UtilizationabstractPCIe devices, such as NICs and SSDs, are frequently underutilized in cloud platforms. PCIe device pools, in which multiple hosts can share a set of PCIe devices, could increase PCIe device utilization and reduce their total cost of ownership. The main way to achieve PCIe device pools today is via PCIe switches, but they are expensive and inflexible. We design Oasis,1 a system that pools PCIe devices in software over CXL memory pools. CXL memory pools are already being deployed to boost datacenter memory utilization and reduce costs. Once CXL pools are in place, they can serve as an efficient data path between hosts and PCIe devices. Oasis provides a control plane and datapath over CXL pools, mapping and routing PCIe device traffic across host boundaries. PCIe devices with different functionalities can be supported by adding an Oasis engine for each device class. We implement an Oasis network engine to demonstrate NIC pooling. Our evaluation shows that Oasis improves the NIC utilization by 2× and handles NIC failover with only a 38 ms interruption. Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti, Enrique Saurez, Jacob Nelson 0001, Dan R. K. Ports, Antonis Psistakis, Joshua Fried, Asaf Cidon |
SOSP | 2 |
| 2025 | FairyWREN: A Sustainable Cache for Emerging Write-Read-Erase Flash InterfacesabstractDatacenters need to reduce embodied carbon emissions, particularly for flash, which accounts for 40% of embodied carbon in servers. However, decreasing flash’s embodied emissions is challenging due to flash’s limited write endurance, which more than halves with each generation of denser flash. Reducing embodied emissions requires extending flash lifetime, stressing its limited write endurance even further. The legacy Logical Block-Addressable Device (LBAD) interface exacerbates the problem by forcing devices to perform garbage collection, leading to even more writes. Flash-based caches in particular write frequently, limiting the lifetimes and densities of the devices they use. These flash caches illustrate the need to break away from LBAD and switch to the new Write-Read-Erase iNterfaces (WREN) now coming to market. WREN affords applications control over data placement and garbage collection. We present Fairy Wren , 1 a flash cache designed for WREN. Fairy Wren reduces writes by co-designing caching policies and flash garbage collection. Fairy Wren provides a 12.5× write reduction over state-of-the-art LBAD caches. This decrease in writes allows flash devices to last longer, decreasing flash cost by 35% and flash carbon emissions by 33%. Sara McAllister, Yucong Wang, Benjamin Berg, Daniel S. Berger, Nathan Beckmann, George Amvrosiadis, Gregory R. Ganger |
ACM Trans. Storage | 4 |
| 2024 | Baleen: ML Admission & Prefetching for Flash Caches
Daniel Lin-Kit Wong, Carson Molder, Sathya Gunasekar, Jimmy Lu, Snehal Khandkar, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
FAST | 8 |
| 2024 | Designing Cloud Servers for Lower CarbonabstractTo mitigate climate change, we must reduce carbon emissions from hyperscale cloud computing. We find that cloud compute servers cause the majority of emissions in a general-purpose cloud. Thus, we motivate designing carbon-efficient compute server SKUs, or GreenSKUs, using recently-available low-carbon server components. To this end, we design and build three GreenSKUs using low-carbon components, such as energy-efficient CPUs, reused old DRAM via CXL, and reused old SSDs.We detail several challenges that limit GreenSKUs, carbon savings at scale and may prevent their adoption by cloud providers. To address these challenges, we develop a novel methodology and associated framework, GSF (GreenSKU Framework), that enables a cloud provider to systematically evaluate a GreenSKU’s carbon savings at scale. We implement GSF within Microsoft Azure’s production constraints to evaluate our three GreenSKUs’ carbon savings. Using GSF, we show that our most carbon-efficient GreenSKU reduces emissions per core by $28 \%$ compared to currently-deployed cloud servers. When designing GreenSKUs to meet applications’ performance requirements, we reduce emissions by $15 \%$. When incorporating overall data center overheads, our GreenSKU reduces Azure’s net cloud emissions by $8 \%$. Jaylen Wang, Daniel S. Berger, Fiodar Kazhamiaka, Celine Irvene, Chaojie Zhang 0001, Esha Choukse, Kali Frost, Rodrigo Fonseca, Brijesh Warrier, Chetan Bansal, Jonathan Stern, Ricardo Bianchini, Akshitha Sriraman |
ISCA | 2 |
| 2024 | FairyWREN: A Sustainable Cache for Emerging Write-Read-Erase Flash Interfaces
Sara McAllister, Yucong Wang, Benjamin Berg, Daniel S. Berger, George Amvrosiadis, Nathan Beckmann, Gregory R. Ganger |
OSDI | 4 |
| 2024 | Managing Memory Tiers with CXL in Virtualized Environments
Yuhong Zhong, Daniel S. Berger, Carl A. Waldspurger, Ryan Wee, Ishwar Agarwal, Rajat Agarwal, Frank Hady, Karthik Kumar, Mark D. Hill, Mosharaf Chowdhury, Asaf Cidon |
OSDI | 2 |
| 2024 | Dense Server Design for Immersion CoolingabstractThe growing demands for computational power in cloud computing have led to a significant increase in the deployment of high-performance servers. The growing power consumption of servers and the heat they produce is on track to outpace the capacity of conventional air cooling systems, necessitating more efficient cooling solutions such as liquid immersion cooling. The superior heat exchange capabilities of immersion cooling both eliminates the need for bulky heat sinks, fans, and air flow channels while also unlocking the potential go beyond conventional 2D blade servers to three-dimensional designs. In this work, we present a computational framework to explore designs of servers in three-dimensional space, specifically targeting the maximization of server density within immersion cooling tanks. Our tool is designed to handle a variety of physical and electrical server design constraints. We demonstrate our optimized designs can reduce server volume by 25--52% compared to traditional flat server designs. This increased density reduces land usage as well as the amount of liquid used for immersion, with significant reduction in the carbon emissions embodied in datacenter buildings. We further create physical prototypes to simulate dense server designs and perform real-world experiments in an immersion cooling tank demonstrating they operate at safe temperatures. This approach marks a critical step forward in sustainable and efficient datacenter management. Milin Kodnongbua, Zachary Englhardt, Ricardo Bianchini, Rodrigo Fonseca, Alvin R. Lebeck, Daniel S. Berger, Vikram Iyer, Fiodar Kazhamiaka, Adriana Schulz |
ACM Trans. Graph. | 6 |
| 2023 | Pond: CXL-Based Memory Pooling Systems for Cloud PlatformsabstractPublic cloud providers seek to meet stringent performance requirements and low hardware cost. A key driver of performance and cost is main memory. Memory pooling promises to improve DRAM utilization and thereby reduce costs. However, pooling is challenging under cloud performance requirements. This paper proposes Pond, the first memory pooling system that both meets cloud performance goals and significantly reduces DRAM cost. Pond builds on the Compute Express Link (CXL) standard for load/store access to pool memory and two key insights. First, our analysis of cloud production traces shows that pooling across 8-16 sockets is enough to achieve most of the benefits. This enables a small-pool design with low access latency. Second, it is possible to create machine learning models that can accurately predict how much local and pool memory to allocate to a virtual machine (VM) to resemble same-NUMA-node memory performance. Our evaluation with 158 workloads shows that Pond reduces DRAM costs by 7% with performance within 1-5% of same-NUMA-node VM allocations. Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, Ricardo Bianchini |
ASPLOS (2) | 2 |
| 2023 | Palette Load Balancing: Locality Hints for Serverless FunctionsabstractFunction-as-a-Service (FaaS) serverless computing enables a simple programming model with almost unbounded elasticity. Unfortunately, current FaaS platforms achieve this flexibility at the cost of lower performance for data-intensive applications compared to a serverful deployment. The ability to have computation close to data is a key missing feature. We introduce Palette load balancing, which offers FaaS applications a simple mechanism to express locality to the platform, through hints we term "colors". Palette maintains the serverless nature of the service - users are still not allocating resources - while allowing the platform to place successive invocations related to each other on the same executing node. We compare a prototype of the Palette load balancer to a state-of-the-art locality-oblivious load balancer on representative examples of three applications. For a serverless web application with a local cache, Palette improves the hit ratio by 6x. For a serverless version of Dask, Palette improves run times by 46% and 40% on Task Bench and TPC-H, respectively. On a serverless version of NumS, Palette improves run times by 37%. These improvements largely bridge the gap to serverful implementation of the same systems. Mania Abdi, Samuel Ginzburg, Xiayue Charles Lin, Jose M. Faleiro, Gohar Irfan Chaudhry, Íñigo Goiri, Ricardo Bianchini, Daniel S. Berger, Rodrigo Fonseca |
EuroSys | 8 |
| 2023 | Hyrax: Fail-in-Place Server Operation in Cloud Platforms
Jialun Lyu, Marisa You, Celine Irvene, Mark Jung, Tyler Narmore, Jacob Shapiro, Luke Marshall, Savyasachi Samal, Ioannis Manousakis, Lisa Hsu, Preetha Subbarayalu, Ashish Raniwala, Brijesh Warrier, Ricardo Bianchini, Bianca Schroeder, Daniel S. Berger |
OSDI | 16 |
| 2023 | Ensō: A Streaming Interface for NIC-Application Communication
Hugo Sadok, Nirav Atre, Daniel S. Berger, James C. Hoe, Aurojit Panda, Justine Sherry |
OSDI | 4 |
| 2022 | SOL: safe on-node learning in cloud platformsabstractCloud platforms run many software agents on each server node. These agents manage all aspects of node operation, and in some cases frequently collect data and make decisions. Unfortunately, their behavior is typically based on pre-defined static heuristics or offline analysis; they do not leverage on-node machine learning (ML). In this paper, we first characterize the spectrum of node agents in Azure, and identify the classes of agents that are most likely to benefit from on-node ML. We then propose SOL, an extensible framework for designing ML-based agents that are safe and robust to the range of failure conditions that occur in production. SOL provides a simple API to agent developers and manages the scheduling and running of the agent-specific functions they write. We illustrate the use of SOL by implementing three ML-based agents that manage CPU cores, node power, and memory placement. Our experiments show that (1) ML substantially improves our agents, and (2) SOL ensures that agents operate safely under a variety of failure conditions. We conclude that ML-based agents show significant potential and that SOL can help build them. Daniel Crankshaw, Neeraja J. Yadwadkar, Daniel S. Berger, Christoforos E. Kozyrakis, Ricardo Bianchini |
ASPLOS | 4 |
| 2022 | CompuCache: Remote Computable Caching using Spot VMs
Qizhen Zhang 0001, Philip A. Bernstein, Daniel S. Berger, Badrish Chandramouli, Vincent Liu 0001, Boon Thau Loo |
CIDR | 3 |
| 2022 | C2DN: How to Harness Erasure Codes at the Edge for Efficient Content Delivery
Juncheng Yang, Anirudh Sabnis, Daniel S. Berger, K. V. Rashmi, Ramesh K. Sitaraman |
NSDI | 3 |
| 2022 | Kangaroo: Theory and Practice of Caching Billions of Tiny Objects on FlashabstractMany social-media and IoT services have very large working sets consisting of billions of tiny (≈100 B) objects. Large, flash-based caches are important to serving these working sets at acceptable monetary cost. However, caching tiny objects on flash is challenging for two reasons: (i) SSDs can read/write data only in multi-KB “pages” that are much larger than a single object, stressing the limited number of times flash can be written; and (ii) very few bits per cached object can be kept in DRAM without losing flash’s cost advantage. Unfortunately, existing flash-cache designs fall short of addressing these challenges: write-optimized designs require too much DRAM, and DRAM-optimized designs require too many flash writes. We present Kangaroo , a new flash-cache design that optimizes both DRAM usage and flash writes to maximize cache performance while minimizing cost. Kangaroo combines a large, set-associative cache with a small, log-structured cache. The set-associative cache requires minimal DRAM, while the log-structured cache minimizes Kangaroo’s flash writes. Experiments using traces from Meta and Twitter show that Kangaroo achieves DRAM usage close to the best prior DRAM-optimized design, flash writes close to the best prior write-optimized design, and miss ratios better than both. Kangaroo’s design is Pareto-optimal across a range of allowed write rates, DRAM sizes, and flash sizes, reducing misses by 29% over the state of the art. These results are corroborated by analytical models presented herein and with a test deployment of Kangaroo in a production flash cache at Meta. Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
ACM Trans. Storage | 7 |
| 2021 | Towards a Cost vs. Quality Sweet Spot for Monitoring NetworksabstractContinuously monitoring a wide variety of performance and fault metrics has become a crucial part of operating large-scale datacenter networks. In this work, we ask whether we can reduce the costs to monitor - in terms of collection, storage and analysis - by judiciously controlling how much and which measurements we collect. By positing that we can treat almost all measured signals as sampled time-series, we show that we can use signal processing techniques such as the Nyquist-Shannon theorem to avoid wasteful data collection. We show that large savings appear possible by analyzing tens of popular measurement systems from a production datacenter network. We also discuss some challenges that must be solved when applying these techniques in practice. Nofel Yaseen, Behnaz Arzani, Krishna Chintalapudi, Vaishnavi Nattar Ranganathan, Felipe Vieira Frujeri, Kevin Hsieh, Daniel S. Berger, Vincent Liu 0001, Srikanth Kandula |
HotNets | 7 |
| 2021 | We need kernel interposition over the network dataplaneabstractKernel-bypass networking, which allows applications to circumvent the kernel and interface directly with NIC hardware, is one of the main tools for improving application network performance. However, allowing applications to circumvent the kernel makes it impossible to use tools (e.g., tcpdump) or impose policies (e.g., QoS and filters) that need to interpose on traffic sent by different applications running on a host. This makes maintainability and manageability a challenge for kernel-bypass applications. In response, we propose Kernel On-Path Interposition (KOPI), in which traditional kernel data-plane functionality is retained but implemented in a fully programmable SmartNIC. We hypothesize that KOPI can support the same tools and policies as the kernel stack while retaining the performance benefits of kernel bypass. Hugo Sadok, Valerie Choung, Nirav Atre, Daniel S. Berger, James C. Hoe, Aurojit Panda, Justine Sherry |
HotOS | 5 |
| 2021 | Don't be a blockhead: zoned namespaces make work on conventional SSDs obsoleteabstractResearch on flash devices almost exclusively focuses on conventional SSDs, which expose a block interface. Industry, however, has standardized and is adopting Zoned Namespaces (ZNS) SSDs, which offer a new storage interface that dominates conventional SSDs. Continued research on conventional SSDs is thus a missed opportunity to unlock a step-change improvement in system performance by building on ZNS SSDs. We argue for an immediate and complete shift in research to ZNS SSDs and discuss research directions. Theano Stavrinos, Daniel S. Berger, Ethan Katz-Bassett, Wyatt Lloyd |
HotOS | 2 |
| 2021 | Kangaroo: Caching Billions of Tiny Objects on FlashabstractMany social-media and IoT services have very large working sets consisting of billions of tiny (≈100 B) objects. Large, flash-based caches are important to serving these working sets at acceptable monetary cost. However, caching tiny objects on flash is challenging for two reasons: (i) SSDs can read/write data only in multi-KB "pages" that are much larger than a single object, stressing the limited number of times flash can be written; and (ii) very few bits per cached object can be kept in DRAM without losing flash's cost advantage. Unfortunately, existing flash-cache designs fall short of addressing these challenges: write-optimized designs require too much DRAM, and DRAM-optimized designs require too many flash writes. Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
SOSP | 7 |
| 2021 | Redy: Remote Dynamic Memory CacheabstractRedy is a cloud service that provides high performance caches using RDMA-accessible remote memory. An application can customize the performance of each cache with a service level objective (SLO) for latency and throughput. By using remote memory, it can leverage stranded memory and spot VM instances to reduce the cost of its caches and improve data center resource utilization. Redy automatically customizes the resource configuration for the given SLO, handles the dynamics of remote memory regions, and recovers from failures. The experimental evaluation shows that Redy can deliver its promised performance and robustness under remote memory dynamics in the cloud. We augment a production key-value store, FASTER, with a Redy cache. When the working set exceeds local memory, using Redy is significantly faster than spilling to SSDs. Qizhen Zhang 0001, Philip A. Bernstein, Daniel S. Berger, Badrish Chandramouli |
Proc. VLDB Endow. | 3 |
| 2020 | Learning Relaxed Belady for Content Distribution Network Caching
Daniel S. Berger, Kai Li 0001, Wyatt Lloyd |
NSDI | 2 |
| 2020 | The CacheLib Caching Engine: Design and Experiences at Scale
Benjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof, Sathya Gunasekar, Jimmy Lu, Michael Uhlar, Jim Carrig, Nathan Beckmann, Mor Harchol-Balter, Gregory R. Ganger |
OSDI | 2 |
| 2020 | Caching with Delayed HitsabstractCaches are at the heart of latency-sensitive systems. In this paper, we identify a growing challenge for the design of latency-minimizing caches called delayed hits. Delayed hits occur at high throughput, when multiple requests to the same object queue up before an outstanding cache miss is resolved. This effect increases latencies beyond the predictions of traditional caching models and simulations; in fact, caching algorithms are designed as if delayed hits simply didn't exist. We show that traditional caching strategies -- even so called 'optimal' algorithms -- can fail to minimize latency in the presence of delayed hits. We design a new, latency-optimal offline caching algorithm called belatedly which reduces average latencies by up to 45% compared to the traditional, hit-rate optimal Belady's algorithm. Using belatedly as our guide, we show that incorporating an object's 'aggregate delay' into online caching heuristics can improve latencies for practical caching systems by up to 40%. We implement a prototype, Minimum-AggregateDelay (mad), within a CDN caching node. Using a CDN production trace and backends deployed in different geographic locations, we show that mad can reduce latencies by 12-18% depending on the backend RTTs. Nirav Atre, Justine Sherry, Weina Wang 0001, Daniel S. Berger |
SIGCOMM | 4 |
| 2018 | Towards Lightweight and Robust Machine Learning for CDN CachingabstractRecent advances in the field of reinforcement learning promise a general approach to optimize networking systems. This paper argues against the recent trend for generalization by introducing a case study where domain-specific modeling enables the application of lightweight and robust learning techniques. Daniel S. Berger |
HotNets | 1 |
| 2018 | RobinHood: Tail Latency Aware Caching - Dynamic Reallocation from Cache-Rich to Cache-Poor
Daniel S. Berger, Benjamin Berg, Timothy Zhu, Siddhartha Sen 0001, Mor Harchol-Balter |
OSDI | 1 |
| 2017 | AdaptSize: Orchestrating the Hot Object Memory Cache in a Content Delivery Network
Daniel S. Berger, Ramesh K. Sitaraman, Mor Harchol-Balter |
NSDI | 1 |
| 2016 | SNC-Meister: Admitting More Tenants with Tail Latency SLOsabstractMeeting tail latency Service Level Objectives (SLOs) in shared cloud networks is both important and challenging. One primary challenge is determining limits on the multi-tenancy such that SLOs are met. Doing so involves estimating latency, which is difficult, especially when tenants exhibit bursty behavior as is common in production environments. Nevertheless, recent papers in the past two years (Silo, QJump, and PriorityMeister) show techniques for calculating latency based on a branch of mathematical modeling called Deterministic Network Calculus (DNC). The DNC theory is designed for adversarial worst-case conditions, which is sometimes necessary, but is often overly conservative. Typical tenants do not require strict worst-case guarantees, but are only looking for SLOs at lower percentiles (e.g., 99th, 99.9th). Timothy Zhu, Daniel S. Berger, Mor Harchol-Balter |
SoCC | 2 |
| 2016 | Friendly Jamming on Access Points: Analysis and Real-World MeasurementsabstractFrequency jamming is known as an efficient attack tool to disrupt wireless communication. This efficiency can also be exploited for the benefit of a network—an idea often referred to as friendly jamming. A prominent application case is the blocking of unauthenticated or malicious communication, such as injection attacks. In this paper, we propose access points as a natural place to implement friendly jamming functionality. We analyze this proposal using simulations, introduce an implementation on customer-grade access points, and report measurement results from the first real-world study of friendly jamming in an IEEE 802.11 campus network. We discover a fundamental tradeoff between the effectiveness of friendly jamming and the orthogonal aspect of having minimal side-effects to the campus network’s traffic. In particular, we observed what we call the power amplification phenomenon. This effect aggravates the known hidden station problem when the number of jammers increases. We also find evidence that the collaboration between jammers can enable friendly jamming, which is both effective and minimally invasive. Daniel S. Berger, Francesco Gringoli, Nicolò Facchi, Ivan Martinovic, Jens B. Schmitt |
IEEE Trans. Wirel. Commun. | 1 |
| 2015 | Security by mobility in location and track verificationabstractThis poster presents the idea of exploiting mobility to improve the security in location and track verification. Unlike traditional approaches which require tight time synchronization or two-way communication, mobility can be used to derive lightweight verification schemes. By ensuring independent movement of the verifiers, our scheme can provide security guarantees even if the verifiers' positions are known to the attacker. We also give an outlook on more general opportunities for mobility-aided security. Matthias Schäfer 0002, Daniel S. Berger, Vincent Lenders, Jens B. Schmitt |
WISEC | 2 |
| 2014 | Exact analysis of TTL cache networks: the case of caching policies driven by stopping timesabstractTTL caching models have recently regained significant research interest, largely due to their ability to fit popular caching policies such as LRU. In this extended abstract we briefly describe our recent work on two exact methods to analyze TTL cache networks. The first method generalizes existing results for line networks under renewal requests to the broad class of caching policies whereby evictions are driven by stopping times. The obtained results are further generalized, using the second method, to feedforward networks with Markov arrival processes (MAP) requests. MAPs are particularly suitable for non-line networks because they are closed not only under superposition and splitting, as known, but also under input-output caching operations as proven herein for phase-type TTL distributions. The crucial benefit of the two closure properties is that they jointly enable the first exact analysis of feedforward networks of TTL caches in great generality. Daniel S. Berger, Philipp Gland, Sahil Singla 0001, Florin Ciucu |
SIGMETRICS | 1 |
| 2014 | On the relevance of adversarial queueing theory in practiceabstractAdversarial Queueing Theory (AQT) has shown that seemingly innocent traffic injection rates might lead to unbounded queues in packet-switched networks - depending on scheduling strategies as well as topological characteristics. Little attention has been given to quantifying these effects in realistic network configurations. In particular, the existing AQT literature makes two unrealistic assumptions: infinite buffers and perfect synchrony. Because finite buffers inherently limit queue sizes, adversarial effects ultimately lead to packet loss which we address in this work. In addition, we study the effect of imperfect network synchronization under the packet loss metric. Our results, using analysis and simulation, indicate that classical AQT examples appear harmless under realistic assumptions but for a novel class of adversaries considerably higher loss can be observed. We introduce this class by giving examples of two new AQT concepts to construct loss-efficient network adversaries. Our analysis proves the robustness of these new adversaries against randomized de-synchronization effects in terms of variable link delays and nodal processing. Daniel S. Berger, Martin Karsten, Jens B. Schmitt |
SIGMETRICS | 1 |
| 2014 | Gaining insight on friendly jamming in a real-world IEEE 802.11 networkabstractFrequency jamming is the fiercest attack tool to disrupt wireless communication and its malicious aspects have received much attention in the literature. Yet, several recent works propose to turn the table and employ so-called friendly jamming for the benefit of a wireless network. For example, recently proposed friendly jamming applications include hiding communication channels, injection attack defense, and access control. This work investigates the practical viability of friendly jamming by applying it in a real-world network. To that end, we implemented a reactive and frame-selective jammer on a consumer grade IEEE 802.11 access point. Equipped with this, we conducted a three weeks real-world study on the jammer's performance and side-effects on legitimate traffic (the cost of jamming) in a university office environment. Our results provide detailed insights on crucial factors governing the trade-off between the effectiveness of friendly jamming (we evaluated up to 13 jammers) and its cost. In particular, we observed -- what we call the power amplification phenomenon -- an effect that aggravates the known hidden station problem when the number of jammers increases. However, we also find evidence that this effect can be alleviated by collaboration between jammers, which again enables effective and minimally invasive friendly jamming. Daniel S. Berger, Francesco Gringoli, Nicolò Facchi, Ivan Martinovic, Jens B. Schmitt |
WISEC | 1 |
| 2014 | Exact analysis of TTL cache networks
Daniel S. Berger, Philipp Gland, Sahil Singla 0001, Florin Ciucu |
Perform. Evaluation | 1 |