VLDB 2026 Research / reviewers in the wild / expert
Joshua Landgraf
dblp:205/2367
· DBLP profile ↗
7ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EZCache: Easy Action-Enabled FPGA Caches for Non-Stalling Datapaths in SmartNICs and BeyondabstractCaches are widely used in FPGA accelerators such as SmartNICs to hide DRAM latency, but conventional designs treat caches as passive storage. When workloads require read–modify–write (RMW) updates – such as flow tables, counters, or per-connection state – existing cache IPs force designers to either stall the pipeline or duplicate hazard-handling logic outside the cache. Both approaches waste bandwidth, complicate datapath design, and require extensive verification.We propose EZCache, a new FPGA cache IP that integrates programmable action blocks directly into the cache pipeline. These blocks perform user-defined operations on cached data in place, allowing RMW updates to complete without stalling other operations or exposing hazards to surrounding logic. By embedding compute into the cache, EZCache transforms it from a passive buffer into an active architectural primitive suitable for a wide range of FPGA accelerators.EZCache is fully parametric in associativity, latency, throughput, and action complexity, enabling a single design to support diverse workloads and memory systems. This parametric abstraction allows the surrounding datapath to remain stable even as cache configurations, action semantics, and external memories evolve across hardware generations. EZCache has been deployed across five generations of SmartNICs and millions of devices worldwide, proving both its maturity and its impact in production. Results across multiple FPGA platforms show EZCache’s generality and efficiency, establishing it as a new foundation for cache-based acceleration in reconfigurable systems. Ahmed Abdelsalam, Vishal Gondaliya, Ezz Hamed, Pragati Medleri Hire Math, Marc Gepigon, Joshua Landgraf, Nadeen Gebara, Bob Groza, Anshuman Verma, Andrew Putnam |
FCCM | 6 |
| 2023 | Reconfigurable Virtual Memory for FPGA-Driven I/OabstractFPGAs are increasingly used to accelerate modern applications, and cloud providers offer FPGA platforms on-demand with a variety of FPGAs, I/O peripherals, and memory options. FPGA vendors expose I/O with low-level interfaces that limit application portability. Current approaches to abstracting these interfaces trade level of abstraction against performance. Joshua Landgraf, Matthew Giordano, Esther Yoon, Christopher J. Rossbach |
ASPLOS (3) | 1 |
| 2021 | Compiler-driven FPGA virtualization with SYNERGYabstractFPGAs are increasingly common in modern applications, and cloud providers now support on-demand FPGA acceleration in data centers. Applications in data centers run on virtual infrastructure, where consolidation, multi-tenancy, and workload migration enable economies of scale that are fundamental to the provider’s business. However, a general strategy for virtualizing FPGAs has yet to emerge. While manufacturers struggle with hardware-based approaches, we propose a compiler/runtime-based solution called Synergy. We show a compiler transformation for Verilog programs that produces code able to yield control to software at sub-clock-tick granularity according to the semantics of the original program. Synergy uses this property to efficiently support core virtualization primitives: suspend and resume, program migration, and spatial/temporal multiplexing, on hardware which is available today. We use Synergy to virtualize FPGA workloads across a cluster of Altera SoCs and Xilinx FPGAs on Amazon F1. The workloads require no modification, run within 3−4× of unvirtualized performance, and incur a modest increase in FPGA fabric utilization. Joshua Landgraf, Tiffany Yang, Will Lin, Christopher J. Rossbach, Eric Schkufza |
ASPLOS | 1 |
| 2019 | Faster Training by Selecting Samples Using EmbeddingsabstractLong training times have increasingly become a burden for researchers by slowing down the pace of innovation, with some models taking days or weeks to train. In this paper, a new, general technique is presented that aims to speed up the training process by using a thinned-down training dataset. By leveraging autoencoders and the unique properties of embedding spaces, we are able to filter training datasets to only include only the samples that matter the most. Through evaluation on a standard CIFAR-10 image classification task, this technique is shown to be effective. With this technique, training times can be reduced with a minimal loss in accuracy. Conversely, given a fixed training time budget, the technique was shown to improve accuracy by over 50%. This intelligent dataset sampling technique is a practical tool for achieving better results with large datasets and limited computational budgets. Santiago Gonzalez, Joshua Landgraf, Risto Miikkulainen |
IJCNN | 2 |
| 2018 | MASK: Redesigning the GPU Memory Hierarchy to Support Multi-Application ConcurrencyabstractGraphics Processing Units (GPUs) exploit large amounts of threadlevel parallelism to provide high instruction throughput and to efficiently hide long-latency stalls. The resulting high throughput, along with continued programmability improvements, have made GPUs an essential computational resource in many domains. Applications from different domains can have vastly different compute and memory demands on the GPU. In a large-scale computing environment, to efficiently accommodate such wide-ranging demands without leaving GPU resources underutilized, multiple applications can share a single GPU, akin to how multiple applications execute concurrently on a CPU. Multi-application concurrency requires several support mechanisms in both hardware and software. One such key mechanism is virtual memory, which manages and protects the address space of each application. However, modern GPUs lack the extensive support for multi-application concurrency available in CPUs, and as a result suffer from high performance overheads when shared by multiple applications, as we demonstrate. We perform a detailed analysis of which multi-application concurrency support limitations hurt GPU performance the most. We find that the poor performance is largely a result of the virtual memory mechanisms employed in modern GPUs. In particular, poor address translation performance is a key obstacle to efficient GPU sharing. State-of-the-art address translation mechanisms, which were designed for single-application execution, experience significant inter-application interference when multiple applications spatially share the GPU. This contention leads to frequent misses in the shared translation lookaside buffer (TLB), where a single miss can induce long-latency stalls for hundreds of threads. As a result, the GPU often cannot schedule enough threads to successfully hide the stalls, which diminishes system throughput and becomes a first-order performance concern. Based on our analysis, we propose MASK, a new GPU framework that provides low-overhead virtual memory support for the concurrent execution of multiple applications. MASK consists of three novel address-translation-aware cache and memory management mechanisms that work together to largely reduce the overhead of address translation: (1) a token-based technique to reduce TLB contention, (2) a bypassing mechanism to improve the effectiveness of cached address translations, and (3) an application-aware memory scheduling scheme to reduce the interference between address translation and data requests. Our evaluations show that MASK restores much of the throughput lost to TLB contention. Relative to a state-of-the-art GPU TLB, MASK improves system throughput by 57.8%, improves IPC throughput by 43.4%, and reduces applicationlevel unfairness by 22.4%. MASK's system throughput is within 23.2% of an ideal GPU system with no address translation overhead. Rachata Ausavarungnirun, Vance Miller, Joshua Landgraf, Saugata Ghose, Jayneel Gandhi, Adwait Jog, Christopher J. Rossbach, Onur Mutlu |
ASPLOS | 3 |
| 2018 | Sharing, Protection, and Compatibility for Reconfigurable Fabric with AmorphOS
Ahmed Khawaja, Joshua Landgraf, Rohith Prakash, Michael Wei, Eric Schkufza, Christopher J. Rossbach |
OSDI | 2 |
| 2017 | Mosaic: a GPU memory manager with application-transparent support for multiple page sizesabstractContemporary discrete GPUs support rich memory management features such as virtual memory and demand paging. These features simplify GPU programming by providing a virtual address space abstraction similar to CPUs and eliminating manual memory management, but they introduce high performance overheads during (1) address translation and (2) page faults. A GPU relies on high degrees of thread-level parallelism (TLP) to hide memory latency. Address translation can undermine TLP, as a single miss in the translation lookaside buffer (TLB) invokes an expensive serialized page table walk that often stalls multiple threads. Demand paging can also undermine TLP, as multiple threads often stall while they wait for an expensive data transfer over the system I/O (e.g., PCIe) bus when the GPU demands a page. Rachata Ausavarungnirun, Joshua Landgraf, Vance Miller, Saugata Ghose, Jayneel Gandhi, Christopher J. Rossbach, Onur Mutlu |
MICRO | 2 |