VLDB 2026 Research / reviewers in the wild / expert
Chris Kennelly
dblp:299/3321 · also Christopher Kennelly
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-9404-7875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 9 since 2021Systems, architecture and hardware · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Necro-reaper: Pruning away Dead Memory Traffic in Warehouse-Scale ComputersabstractMemory bandwidth is emerging as a critical bottleneck in warehouse-scale computing (WSC). This work reveals that a significant portion of memory traffic in WSC is surprisingly unnecessary, consisting of needless writebacks of deallocated data and fetches of uninitialized data. This issue is particularly acute in WSC, where short-lived heap allocations bigger than a cache line are prevalent. To address this problem, this work proposes a pragmatic approach tailored to WSC. Leveraging the existing WSC ecosystem of vertical integration, profile-guided compilation flows, and customized memory allocators, this work presents Necro-reaper, a novel software/hardware co-design that avoids dead memory traffic without requiring the hardware tracking of prior work. New ISA instructions enable the hardware to avoid unnecessary dead traffic, while extended software components, including a profile-guided compiler and memory allocator, optimize the utilization of these instructions. Evaluation across a diverse set of 10 WSC workloads demonstrates that Necro-reaper achieves a geomean memory traffic reduction of 26% and a geomean IPC increase of 6%. Sotiris Apostolakis, Chris Kennelly, Xinliang David Li, Parthasarathy Ranganathan |
ASPLOS (2) | 2 |
| 2024 | Limoncello: Prefetchers for ScaleabstractThis paper presents Limoncello, a novel software system that dynamically configures data prefetchers for high-utilization systems. We demonstrate that in resource-constrained environments, such as large data centers, traditional methods of hardware prefetching can increase memory latency and decrease available memory bandwidth. To address this issue, Limoncello disables hardware prefetchers when memory bandwidth utilization is high, and it leverages targeted software prefetching to reduce cache misses when hardware prefetchers are disabled. Limoncello is software-centric and does not require any modifications to hardware. Our evaluation of the deployment on Google's fleet reveals that Limoncello unlocks significant performance gains for high-utilization systems: It improves application throughput by 10%, due to a 15% reduction in memory latency, while maintaining minimal change in cache miss rate for targeted library functions. Akanksha Jain, Hannah Lin, Carlos Villavieja, Baris Kasikci, Chris Kennelly, Milad Hashemi, Parthasarathy Ranganathan |
ASPLOS (3) | 5 |
| 2024 | Characterizing a Memory Allocator at Warehouse ScaleabstractMemory allocation constitutes a substantial component of warehouse-scale computation. Optimizing the memory allocator not only reduces the datacenter tax, but also improves application performance, leading to significant cost savings. Vaibhav Gogte, Nilay Vaish, Chris Kennelly, Patrick Xia 0001, Svilen Kanev, Tipp Moseley, Christina Delimitrou, Parthasarathy Ranganathan |
ASPLOS (3) | 4 |
| 2024 | A Profiling-Based Benchmark Suite for Warehouse-Scale ComputersabstractBenchmarking warehouse-scale compute (WSC) systems poses unique challenges due to their scale, workload diversity, and focus on data-intensive computation. Traditional benchmarks often fall short in capturing the nuances of WSC behavior. Creating a public benchmark suite that is representative of the workloads used by actual warehouse-scale computers is challenging, as they typically run proprietary, non-public software that operates on confidential data. In this paper, we present Fleetbench, a benchmark suite that is representative of the workloads used at Google. Fleetbench does not benchmark entire applications; rather, it focuses on common building blocks, known as the “datacenter tax”, that are at the core of a wide range of different data center applications, and many of which are open source. The relevant building blocks are selected based on fleet-wide profiling data. Representative input data is collected using a fleet-wide value profiler, ensuring that no confidential information is revealed. Fleetbench is available as open source on Github21https://github.com/google/fieetbench. Andreas Abel 0006, Richard O'Grady, Chris Kennelly, Darryl Gove |
ISPASS | 4 |
| 2023 | Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScaleabstractFast DRAM increasingly dominates infrastructure spend in large scale computing environments and this trend will likely worsen without an architectural shift. The cost of deployed memory can be reduced by replacing part of the conventional DRAM with lower cost albeit slower memory media, thus creating a tiered memory system where both tiers are directly addressable and cached. But, this poses numerous challenges in a highly multi-tenant warehouse-scale computing setting. The diversity and scale of its applications motivates an application-transparent solution in the general case, adaptable to specific workload demands. Padmapriya Duraisamy, Scott Hare, Ravi Rajwar, David E. Culler, Zhiyi Xu, Jianing Fan, Chris Kennelly, Bill McCloskey, Danijela Mijailovic, Brian Morris, Chiranjit Mukherjee, Jingliang Ren, Greg Thelen, Carlos Villavieja, Parthasarathy Ranganathan, Amin Vahdat |
ASPLOS (3) | 8 |
| 2022 | Carbink: Fault-Tolerant Far Memory
Yang Zhou 0008, Hassan M. G. Wassel, Sihang Liu 0001, James W. Mickens, Minlan Yu, Chris Kennelly, David E. Culler, Henry M. Levy, Amin Vahdat |
OSDI | 7 |
| 2021 | Adaptive huge-page subrelease for non-moving memory allocators in warehouse-scale computersabstractModern C++ server workloads rely on 2 MB huge pages to improve memory system performance via higher TLB hit rates. Huge pages have traditionally been supported at the kernel level, but recent work has shown that user-level, huge page-aware memory allocators can achieve higher huge page coverage and thus performance. These memory allocators deal with a trade-off: 1) allocate memory from the operating system (OS) at the granularity of a huge page, achieve high performance, but potentially waste memory due to fragmentation, or 2) limit fragmentation by breaking up huge pages into smaller 4 KB pages and returning them to the OS, but reduce performance due to lower huge page coverage. For example, the state-of-the-art TCMalloc allocator handles this trade-off by releasing memory to the OS at a configurable release rate, breaking up huge pages as necessary. Martin Maas 0001, Chris Kennelly, Khanh Nguyen 0001, Darryl Gove, Kathryn S. McKinley |
ISMM | 2 |
| 2021 | automemcpy: a framework for automatic generation of fundamental memory operationsabstractMemory manipulation primitives (memcpy, memset, memcmp) are used by virtually every application, from high performance computing to user interfaces. They often consume a significant portion of CPU cycles. Because they are so ubiquitous and critical, they are provided by language runtimes and in particular by libc, the C standard library. These implementations are heavily optimized, typically written in hand-tuned assembly for each target architecture. Guillaume Châtelet, Chris Kennelly, Sam (Likun) Xi, Ondrej Sýkora, Clement Courbet, Xinliang David Li, Bruno De Backer |
ISMM | 2 |
| 2021 | A Hardware Accelerator for Protocol BuffersabstractSerialization frameworks are a fundamental component of scale-out systems, but introduce significant compute overheads. However, they are amenable to acceleration with specialized hardware. To understand the trade-offs involved in architecting such an accelerator, we present the first in-depth study of serialization framework usage at scale by profiling Protocol Buffers (“protobuf”) usage across Google’s datacenter fleet. We use this data to build HyperProtoBench, an open-source benchmark representative of key serialization-framework user services at scale. In doing so, we identify key insights that challenge prevailing assumptions about serialization framework usage. Sagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao, Dinesh Parimi, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan |
MICRO | 3 |
| 2021 | Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocator
A. H. Hunter, Chris Kennelly, Darryl Gove, Tipp Moseley, Parthasarathy Ranganathan |
OSDI | 2 |