VLDB 2026 Research / reviewers in the wild / expert
Junior Loff
dblp:220/7415 · also Júnior Löff
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-4824-4621ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MapReplay: Trace-Driven Benchmark Generation for Java HashMapabstractHash-based maps, particularly java.util.HashMap, are pervasive in Java applications and the JVM, making their performance critical. Evaluating optimizations is challenging because performance depends on factors such as operation patterns, key distributions, and resizing behavior. Microbenchmarks are fast and repeatable but often oversimplify workloads, failing to capture the realistic usage patterns. Application benchmarks (e.g., DaCapo, Renaissance) provide realistic usages but are more expensive to run, prone to variability, and dominated by non-HashMap computations, making map-related performance changes difficult to observe. To address this challenge, we propose MapReplay, a benchmarking methodology that combines the realism of application benchmarks with the efficiency of microbenchmarks. MapReplay traces HashMap API usages generating a replay workload that reproduces the same operation sequence while faithfully reconstructing internal map states. This enables realistic and efficient evaluation of alternative implementations under realistic usage patterns. Applying MapReplay to DaCapo-Chopin and Renaissance, the resulting suite, MapReplayBench, reproduces application-level performance trends while reducing experimentation time and revealing insights difficult to obtain from full benchmarks. Filippo Schiavio, Andrea Rosà, Junior Loff, Lubomír Bulej, Petr Tuma 0001, Walter Binder |
ICPE | 3 |
| 2025 | NPB-PSTL: C++ STL Algorithms with Parallel Execution Policies in NAS Parallel BenchmarksabstractThe C++ language continually evolves through formal specifications established by its standards committee, proposing new features to maintain $\mathrm{C}++$ as a relevant programming language while improving usability, performance, and portability across platforms. With the addition of parallel Standard Template Library (STL) algorithms in C++17, programmers can now leverage parallel processing capabilities via vendor-neutral parallel execution policies. This study presents an adaptation of the NAS Parallel Benchmarks (NPB)—a well-established suite of applications for evaluating parallel architectures-by porting its sequential C-style code to use C++ STL abstractions and performance-portable parallelism features. Our goals are to (1) assess the suitability of C++ STL for scientific applications like the ones in the NPB and (2) provide a comparative performance and portability of STL algorithms’ parallel execution policies across different multicore architectures (x86 and AArch64). Results indicate that the performance of parallel STL algorithms is often close to that of optimized handwritten versions (OpenMP, Intel TBB, and FastFlow) on different architectures, with notable shortfalls. Across all NPB benchmarks, the STL algorithms’ geometric mean shows sequential execution times that are between 3.76% and $\mathrm{6. 9 \%}$ higher, while parallel executions may reach a geometric mean of up to $\mathrm{2 1. 2 1 \%}$ higher execution time. Junior Loff, Renato B. Hoffmann, Nicola Bianchessi, Leonardo Mallmann, Dalvan Griebler, Walter Binder |
PDP | 1 |
| 2025 | HeapBuffers: Why Not Just Using a Binary Serialization Format for Your Managed Memory?abstractIn modern cloud applications, data serialization and deserialization (DSD) are common yet expensive operations. Languages such as Java incur significant overhead from DSD operations because of their managed memory, which uses an IO-incompatible, sparse, platform-dependent data layout for storing objects. Can language VMs bypass costly DSD operations by directly operating on an IO-ready binary format? In this paper, we address this question by presenting a new DSD library for Java called HeapBuffers. HeapBuffers maintains a Java-friendly programming model through managed memory, yet it operates on data stored using a dense, platform-independent, memory format. In this way, HeapBuffers data can be used to perform IO operations without the need to convert objects from/to a VM-internal memory representation, thus eliminating the need for expensive DSD operations. Compiler optimizations combined with specialized garbage collection ensure that accessing HeapBuffers data is nearly as efficient as handling regular Java objects. The answer to our question is largely positive: our experiments show that HeapBuffers data can be accessed at high performance and can be used as is to perform IO, making the cost of DSD effectively negligible. IO performance of applications using HeapBuffers is often orders of magnitude faster than that obtained using the prevailing DSD libraries. Daniele Bonetta, Junior Loff, Matteo Basso, Walter Binder |
Proc. ACM Program. Lang. | 2 |
| 2024 | MPR: An MPI Framework for Distributed Self-adaptive Stream Processing
Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes, Walter Binder |
Euro-Par (3) | 1 |
| 2024 | Vectorized Intrinsics Can Be Replaced with Pure Java Code without Impairing Steady-State PerformanceabstractSeveral methods of the Java Class Library (JCL) rely on vectorized intrinsics. While these intrinsics undoubtedly lead to better performance, implementing them is extremely challenging, tedious, error-prone, and significantly increases the effort in understanding and maintaining the code. Moreover, their implementation is platform-dependent. An unexplored, easier-to-implement alternative is to replace vectorized intrinsics with portable Java code using the Java Vector API. However, this is attractive only if the Java code achieves similar steady-state performance as the intrinsics. This paper shows that this is the case. We focus on the hashCode and equals computations for byte arrays. We replace the platform-dependent vectorized intrinsics with pure-Java code employing the Java Vector API, resulting in similar steady-state performance. We show that our Java implementations are easy to fine-tune by exploiting characteristics of the input (i.e., the array length), while such tuning would be much more difficult and cumbersome in a vectorized intrinsic. Additionally, we propose a new vectorized hashCode computation for long arrays, for which a corresponding intrinsic is currently missing. We evaluate the performance of the tuned implementations on four popular benchmark suites, showing that the performance are in line with those of the original OpenJDK 21 with intrinsics. Finally, we describe a general approach to integrate code using the Java Vector API into the core classes of the JCL, which is challenging because premature use of the Java Vector API would crash the JVM during its fragile initialization phase. Our approach can be adopted by developers to modify JCL classes without any changes to the native codebase. Junior Loff, Filippo Schiavio, Andrea Rosà, Matteo Basso, Walter Binder |
ICPE | 1 |
| 2022 | OpenMP as runtime for providing high-level stream parallelism on multi-cores
Renato B. Hoffmann, Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes |
J. Supercomput. | 2 |
| 2021 | The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 1 |
| 2018 | Efficient NAS Benchmark Kernels with C++ Parallel ProgrammingabstractBenchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As a consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and an efficient parallelization on average. Dalvan Griebler, Junior Loff, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 2 |