VLDB 2026 Research / reviewers in the wild / expert
Jakob Blomer
dblp:16/9777
· DBLP profile ↗
6ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0001-9750-6224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EVENTSETPROCESSOR: An Engine for Efficiently Combining High-Energy Physics DataabstractCERN’s Large Hadron Collider (LHC), the world’s largest high-energy physics (HEP) instrument, collects tens of petabytes of data per year. The LHC’s next phase is expected to produce up to ten times more data, which calls for novel, more efficient ways of storing and processing these data.HEP collider data are prepared and provided to physicists as read-only data sets, stored in a custom columnar data format. While traditionally all data needed for a particular analysis were captured in a single data set, the increasing scale of the LHC and the advent of modern analysis techniques now requires analysis workflows to use data from different data sets. However, the processing model established across the HEP community does not yet provide a straightforward way to achieve this and currently relies heavily on data duplication to produce the desired data sets. This leads to significant overhead in analysis workflows, both in runtime and storage.To reduce this overhead, we propose more efficient ways to combine HEP data sets. Specifically, we design union and join operations, as defined in relational algebra, to combine HEP data sets at runtime, eliminating therefore the need for data duplication. In this paper, we specify these operations for HEP data and introduce EVENTSETPROCESSOR – an engine that implements these operations for HEP data processing. Through a first prototype, we show that this engine integrates well in existing HEP workflows, and that it can perform up to twice as fast as the current approach. Florine Willemijn de Geus, Vincenzo Eduardo Padulano, Jakob Blomer, Hannes Mühleisen, Ana Lucia Varbanescu |
eScience | 3 |
| 2024 | Parallel Writing of Nested Data in Columnar Formats
Jonas Hahnfeld, Jakob Blomer, Thorsten Kollegger |
Euro-Par (2) | 2 |
| 2023 | LLAMA: The low-level abstraction for memory accessabstractAbstract The performance gap between CPU and memory widens continuously. Choosing the best memory layout for each hardware architecture is increasingly important as more and more programs become memory bound. For portable codes that run across heterogeneous hardware architectures, the choice of the memory layout for data structures is ideally decoupled from the rest of a program. This can be accomplished via a zero‐runtime‐overhead abstraction layer, underneath which memory layouts can be freely exchanged. We present the low‐level abstraction of memory access (LLAMA), a C++ library that provides such a data structure abstraction layer with example implementations for multidimensional arrays of nested, structured data. LLAMA provides fully C++ compliant methods for defining and switching custom memory layouts for user‐defined data types. The library is extensible with third‐party allocators. Providing two close‐to‐life examples, we show that the LLAMA‐generated array of structs and struct of arrays layouts produce identical code with the same performance characteristics as manually written data structures. Integrations into the SPEC CPU® lbm benchmark and the particle‐in‐cell simulation PIConGPU demonstrate LLAMA's abilities in real‐world applications. LLAMA's layout‐aware copy routines can significantly speed up transfer and reshuffling of data between layouts compared with naive element‐wise copying. LLAMA provides a novel tool for the development of high‐performance C++ applications in a heterogeneous environment. Bernhard Manfred Gruber, Guilherme Amadio, Jakob Blomer, Alexander Matthes, René Widera, Michael Bussmann |
Softw. Pract. Exp. | 3 |
| 2020 | Solving the Container Explosion Problem for Distributed High Throughput ComputingabstractContainer technologies are seeing wider use at advanced computing facilities for managing highly complex applications that must execute at multiple sites. However, in a distributed high throughput computing setting, the unrestricted use of containers can result in the container explosion problem. If a new container image is generated for each variation of a job dispatched to a site, shared storage is soon exceeded. On the other hand, if a single large container image is used to meet multiple needs, the size of that container may become a problem for storage and transport. To address this problem, we observe that many containers have an internal structure generated by a structured package manager, and this information could be used to strategically combine and share container images. We develop Landlord to exploit this property and evaluate its performance through a combination of simulation studies and empirical measurement of high energy physics applications. Timothy Shaffer, Nicholas L. Hazekamp, Jakob Blomer, Douglas Thain |
IPDPS | 3 |
| 2013 | Interactive Exploitation of Nonuniform Cloud Resources for LHC Computing at CERNabstractComputing at LHC is based on the Grid model, where geographically distributed local batch farms are federated using proper middleware. Several computing centers are considering a conversion to private clouds, capable of supporting alternative computing models along with the Grid one: CERN is pioneering such conversion. Computing tasks at LHC are mostly performed using ROOT, a framework for simulation, reconstruction and data analysis. In High Energy Physics (HEP) data are independent physics collision events. PROOF (Parallel ROOT Facility) is a cluster model built on top of ROOT, capable of processing such events interactively and in parallel. Given the increasing availability of cloud resources, running virtual PROOF clusters has become an appealing alternative. We will show how PROOF interactivity naturally fits into the cloud model: in particular by overcoming VM performance diversity via dynamic workload assignment. Several technologies are combined: the CernVM ecosystem provides the base VM image, APIs for cloud federation and a distributed filesystem for LHC software distribution, HTCondor and PROOF on Demand (PoD) control scheduling of leases and affect the lifecycle of VMs opportunistically. We will illustrate the current status of the project, and the conjoint efforts on PROOF and CernVM to create a reference implementation suiting the needs of all LHC experiments. The final product will be a personal and elastic "Analysis Facility as a Service" where the user connects using ROOT itself as standard LHC client. Dario Berzano, Jakob Blomer, Predrag Buncic, Gerardo Ganis, Georgios Lestaris, René Meusel |
IEEE CLOUD | 2 |
| 2010 | A Fully Decentralized File System Cache for the CernVM-FSabstractScientific computing often takes place on globally scattered resources. The CernVM web file system supports such scenarios. It was designed to easily retrieve files from the web server that provides the CERN software distributions. In this paper, we extend the CernVM file system. We propose a distributed algorithm to retrieve cached data from the memory of participating peers. Users at the same department shall thereby experience a greatly reduced access latency. Our system forms a cache layer on top of the read-only CernVM-FS. It is fully decentralized and resilient to node churn. It gathers and shares information about file presence in the local network. Each peer can independently decide how much load it is willing to take. Jakob Blomer, Thomas Fuhrmann |
ICCCN | 1 |