EDBT 2026 Demo / reviewers in the wild / expert
Jacob Wahlgren
dblp:324/5231
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-1669-7714ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Communication Offloading on SmartNIC DPUs: A Quantitative Approach
Jacob Wahlgren, Andong Hu, Roger A. Pearce, Maya B. Gokhale, Ivy Bo Peng |
Euro-Par (1) | 1 |
| 2025 | ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
Daniel Araújo de Medeiros, Jeremy J. Williams, Jacob Wahlgren, Leonardo Saud Maia Leite, Ivy Bo Peng |
Euro-Par (1) | 3 |
| 2025 | Design of Neko - A Scalable High-Fidelity Simulation Framework With Extensive Accelerator SupportabstractABSTRACT Recent trends and advancements in including more diverse and heterogeneous hardware in High‐Performance Computing (HPC) are challenging scientific software developers in their pursuit of efficient numerical methods with sustained performance across a diverse set of platforms. As a result, researchers are today forced to re‐factor their codes to leverage these powerful new heterogeneous systems. We present our design considerations of Neko—a portable framework for high‐fidelity spectral element flow simulations. Unlike prior works, Neko adopts a modern object‐oriented Fortran 2008 approach, allowing multi‐tier abstractions of the solver stack and facilitating various hardware backends ranging from general‐purpose processors, accelerators down to exotic vector processors and Field‐Programmable Gate Arrays (FPGAs). Focusing on the performance and portability of Neko, we describe the framework's device abstraction layer managing device memory, data transfer and kernel launches from Fortran, allowing for a solver written in a hardware‐neutral yet performant way. Accelerator‐specific optimizations are also discussed, with auto‐tuning of key kernels and various communication strategies using device‐aware MPI. Finally, we present performance measurements on a wide range of computing platforms, including the EuroHPC pre‐exascale system LUMI, where Neko achieves excellent parallel efficiency for a large direct numerical simulation (DNS) of turbulent fluid flow using up to 80% of the entire LUMI supercomputer. Niclas Jansson, Martin Karp, Jacob Wahlgren, Stefano Markidis, Philipp Schlatter |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace HopperabstractMemory management across discrete CPU and GPU physical memory is traditionally achieved through explicit GPU allocations and data copy or unified virtual memory. The Grace Hopper Superchip, for the first time, supports an integrated CPU-GPU system page table, hardware-level addressing of system allocated memory, and cache-coherent NVLink-C2C interconnect, bringing an alternative solution for enabling a Unified Memory system. In this work, we provide the first in-depth study of the system memory management on the Grace Hopper Superchip, in both in-memory and memory oversubscription scenarios. We provide a suite of six representative applications, including the Qiskit quantum computing simulator, using system memory and managed memory. Using our memory utilization profiler and hardware counters, we quantify and characterize the impact of the integrated CPU-GPU system page table on GPU applications. Our study focuses on first-touch policy, page table entry initialization, page sizes, and page migration. We identify practical optimization strategies for different access patterns. Our results show that as a new solution for unified memory, the system-allocated memory can benefit most use cases with minimal porting efforts. Gabin Schieffer, Jacob Wahlgren, Jie Ren 0015, Jennifer Faj, Ivy Bo Peng |
ICPP | 2 |
| 2024 | Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph ProcessingabstractDisaggregated memory breaks the boundary of monolithic servers to enable memory provisioning on demand. Using network-attached memory to provide memory expansion for memory-intensive applications on compute nodes can improve the overall memory utilization on a cluster and reduce the total cost of ownership. However, current software solutions for leveraging network-attached memory must consume resources on the compute node for memory management tasks. Emerging off-path smartNICs provide general-purpose programmability at low-cost low-power cores. This work provides a general architecture design that enables network-attached memory and offloading tasks onto off-path programmable SmartNIC. We provide a prototype implementation called SODA on Nvidia BlueField DPU. SODA adapts communication paths and data transfer alternatives, pipelines data movement stages, and enables customizable data caching and prefetching optimizations. We evaluate SODA in five representative graph applications on real-world graphs. Our results show that SODA can achieve up to 7.9x speedup compared to node-local SSD and reduce network traffic by 42 % compared to disaggregated memory without SmartNIC offloading at similar or better performance. Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Roger A. Pearce, Ivy Bo Peng |
SBAC-PAD | 1 |
| 2023 | Quantum Computer Simulations at Warp Speed: Assessing the Impact of GPU Acceleration: A Case Study with IBM Qiskit Aer, Nvidia Thrust & cuQuantumabstractQuantum computer simulators are crucial for the development of quantum computing. This work investigates GPU and multi-GPU systems' suitability and performance impact on a widely used simulation tool – the state vector simulator Qiskit Aer. In particular, we evaluate the performance of both Qiskit's default Nvidia Thrust backend and the recent Nvidia cuQuantum backend on Nvidia A100 GPUs. We provide a benchmark suite of representative quantum applications for characterization. For simulations with a large number of qubits, the two GPU backends can provide up to 14× speedup over the CPU backend, with Nvidia cuQuantum providing a further 1.5–3× speedup over the default Thrust backend. Our evaluation on a single GPU identifies the most important functions in Nvidia Thrust and cuQuantum for different quantum applications and their compute and memory bottlenecks. We also evaluate the gate fusion and cache-blocking optimizations on different quantum applications. Finally, we evaluate large-number qubit quantum applications on multi-GPU and identify data movement between host and GPU as the limiting factor for the performance. Jennifer Faj, Ivy Bo Peng, Jacob Wahlgren, Stefano Markidis |
e-Science | 3 |
| 2023 | Kub: Enabling Elastic HPC Workloads on Containerized EnvironmentsabstractThe conventional model of resource allocation in HPC systems is static. Thus, a job cannot leverage newly available resources in the system or release underutilized resources during the execution. In this paper, we present Kub, a methodology that enables elastic execution of HPC workloads on Kubernetes so that the resources allocated to a job can be dynamically scaled during the execution. One main optimization of our method is to maximize the reuse of the originally allocated resources so that the disruption to the running job can be minimized. The scaling procedure is coordinated among nodes through remote procedure calls on Kubernetes for deploying workloads in the cloud. We evaluate our approach using one synthetic benchmark and two production-level MPI-based HPC applications - GRO-MACS and CM1. Our results demonstrate that the benefits of adapting the allocated resources depend on the workload characteristics. In the tested cases, a properly chosen scaling point for increasing resources during execution achieved up to 2x speedup. Also, the overhead of checkpointing and data reshuffling significantly influences the selection of optimal scaling points and requires application-specific knowledge. Daniel Araújo de Medeiros, Jacob Wahlgren, Gabin Schieffer, Ivy Bo Peng |
SBAC-PAD | 2 |
| 2023 | A Quantitative Approach for Adopting Disaggregated Memory in HPC SystemsabstractMemory disaggregation has recently been adopted in data centers to improve resource utilization, motivated by cost and sustainability. Recent studies on large-scale HPC facilities have also highlighted memory underutilization. A promising and non-disruptive option for memory disaggregation is rack-scale memory pooling, where node-local memory is supplemented by shared memory pools. This work outlines the prospects and requirements for adoption and clarifies several misconceptions. We propose a quantitative method for dissecting application requirements on the memory system from the top down in three levels, moving from general, to multi-tier memory systems, and then to memory pooling. We provide a multi-level profiling tool and LBench to facilitate the quantitative approach. We evaluate a set of representative HPC workloads on an emulated platform. Our results show that prefetching activities can significantly influence memory traffic profiles. Interference in memory pooling has varied impacts on applications, depending on their access ratios to memory tiers and arithmetic intensities. Finally, in two case studies, we show the benefits of our findings at the application and system levels, achieving 50% reduction in remote access and 13% speedup in BFS, and reducing performance variation of co-located workloads in interference-aware job scheduling. Jacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy Bo Peng |
SC | 1 |