VLDB 2026 Research / reviewers in the wild / expert
Ruben Laso
dblp:260/9041 · also Ruben Laso Rodríguez
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-2574-4025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cheesemap: A high-performance point-indexing data structure for neighbor search in LiDAR dataabstractPoint-cloud data, as the representation of three-dimensional spatial information, is a fundamental piece of information in various domains where indexing and querying these point clouds efficiently is crucial for tasks such as object recognition, autonomous navigation, and environmental modeling. In this paper, we present a novel data structure, cheesemap , designed for fast neighbor search in 3D LiDAR point clouds. Points are indexed using a grid of voxels, which can be organized in three different ways, originating three flavors of the cheesemap : dense, sparse, and mixed. The lookup of the voxels is theoretically ensured to be performed in constant or amortized constant time, speeding up the search for neighboring points. Experimental results show that cheesemap can outperform, in terms of performance and memory footprint, other state-of-the-art data structures both in region-based and -NN queries throughout different types of point clouds, particularly for Airborne Laser Scanning (ALS) point clouds. Ruben Laso, Miguel Yermo |
Future Gener. Comput. Syst. | 1 |
| 2026 | Tuned your MPI library? Now check the performance guidelinesabstractThe MPI standard provides the foundational building blocks for most parallel applications running on large-scale HPC architectures.Collective communication operations in MPI are critical components for the scalability of these applications. Most MPI librariesoffer several algorithms for a specific collective operation, and each library selects the algorithm to be used based on the numberof processes, the message size, and possibly other factors. Each collective algorithm may perform better in certain scenarios, andthus, selecting the most suitable algorithm for each use case is essential. However, even the best algorithm in a given MPI librarymay deliver suboptimal performance.Self-consistent MPI performance guidelines capture semantic relationships between different collective operations and exploitthese to express performance expectations that collectives should reasonably satisfy to be considered performance-consistent. Forcollective communication, such performance guidelines typically state that a specialized collective call should not be slower thanless specialized counterparts.In this article, we demonstrate how the consistency of MPI libraries with respect to performance guidelines can be analyzed.For this purpose, we present a tool that checks guideline compliance. For regular collective operations such as MPI_Bcast, thetool contains multiple emulated versions of the collective by composing less specialized operations. Then, for a specific number ofprocesses and message sizes, the tool experimentally assesses whether the algorithm selected by the MPI library is slower than itsemulated counterparts. If that is the case, a performance-guideline violation is detected. In a broader empirical study, we assess thecurrent state of performance consistency in MPI libraries on modern supercomputers. Sascha Hunold, Jesper Larsson Träff, Ruben Laso |
Parallel Comput. | 3 |
| 2025 | Reproducibility Report for SC25 Paper RAPTOR: Practical Numerical Profiling of Scientific ApplicationsabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper RAPTOR: Practical Numerical Profiling of Scientific Applications by Hoerold et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibility Committee. Ruben Laso |
SC | 1 |
| 2025 | Reproducibility Report for SC25 Paper Addressing Reproducibility Challenges in HPC with Continuous IntegrationabstractThis reproducibility report provides details about the artifact evaluation done with regards to the Artifact Description and Evaluation appendix of SC25 paper Addressing Reproducibility Challenges in HPC with Continuous Integration by Hayot-Sasson et al. The work was done as part of the Reproducibility Initiative of SC25. The author is a member of the SC25 Reproducibility Committee. Ruben Laso |
SC | 1 |
| 2025 | Mpisee: Communicator-Centric Profiling of MPI ApplicationsabstractABSTRACT mpisee is a lightweight profiling tool designed to track MPI communication operations per communicator, providing fine‐grained insights into MPI applications that use communicators to partition MPI communication. While existing profiling tools offer valuable information, they may limit detailed analysis and optimization for such MPI applications, as they do not associate MPI communication with their communicator. Additionally, mpisee categorizes MPI communication operations based on message size, offering more granular information. It uses an SQLite database to efficiently store the profiling data, enabling users to analyze the application's profile from various perspectives, focusing on specific MPI ranks, operations, and more. Our analysis shows that mpisee incurs less than 5% overhead, performing on par with other state‐of‐the‐art profilers. We demonstrate mpisee 's effectiveness by profiling and analyzing an FFT application, revealing potential performance bottlenecks related to the MPI_Alltoallv collective operation on small communicators and insights not available by other profilers. Leveraging this detailed information, we improved the application's overall performance by selecting different algorithms for MPI_Alltoallv and measuring their performance on different communicators with mpisee . This study illustrates mpisee 's utility and highlights the significant advantages of a communicator‐centric approach in MPI profiling. Ioannis Vardas, Jesper Larsson Träff, Ruben Laso, Sascha Hunold |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Exploring Scalability in C++ Parallel STL ImplementationsabstractSince the advent of parallel algorithms in the C++17 Standard Template Library (STL), the STL has become a viable framework for creating performance-portable applications. Given multiple existing implementations of the parallel algorithms, a systematic, quantitative performance comparison is essential for choosing the appropriate implementation for a particular hardware configuration. Ruben Laso, Diego Krupitza, Sascha Hunold |
ICPP | 1 |
| 2024 | Assessing Intel OneAPI capabilities and cloud-performance for heterogeneous computingabstractAbstract This work presents a performance-oriented study of a heterogeneous application developed with Intel OneAPI to solve two well-known diffusion problems: heat diffusion and image denoising. We have explored CPU+iGPU and CPU+FPGA schemes, applying dynamic load balancing and conducting experiments on Intel DevCloud. The results demonstrate that the CPU+iGPU scheme outperforms the execution times achieved by the fastest device when the problem is sufficiently computationally demanding. We also found that the performance of the CPU+FPGA scheme is heavily affected by bandwidth limitations and specific strategies to manage memory efficiently are required. Moreover, it was demonstrated that dynamic workload balancing is crucial due to possible performance fluctuations in any of the implicated devices. In conclusion, Intel OneAPI provides a helpful tool for multi-platform development using a unique high-level language, DPC++. However, developing specific code for each platform is necessary to achieve optimal performance. Silvia R. Alcaraz, Ruben Laso, Oscar G. Lorenzo, David López Vilariño, Tomás F. Pena, Francisco F. Rivera |
J. Supercomput. | 2 |
| 2022 | CIMAR, NIMAR, and LMMA: Novel algorithms for thread and memory migrations in user space on NUMA systems using hardware countersabstractThis paper introduces two novel algorithms for thread migrations, named CIMAR (Core-aware Interchange and Migration Algorithm with performance Record –IMAR–) and NIMAR (Node-aware IMAR), and a new algorithm for the migration of memory pages, LMMA (Latency-based Memory pages Migration Algorithm), in the context of Non-Uniform Memory Access (NUMA) systems. This kind of system has complex memory hierarchies that present a challenging problem in extracting the best possible performance, where thread and memory mapping play a critical role. The presented algorithms gather and process the information provided by hardware counters to make decisions about the migrations to be performed, trying to find the optimal mapping. They have been implemented as a user space tool that looks for improving the system performance, particularly in, but not restricted to, scenarios where multiple programs with different characteristics are running. This approach has the advantage of not requiring any modification on the target programs or the Linux kernel while keeping a low overhead. Two different benchmark suites have been used to validate our algorithms: The NAS parallel benchmark, mainly devoted to computational routines, and the LevelDB database benchmark focused on read–write operations. These benchmarks allow us to illustrate the influence of our proposal in these two important types of codes. Note that those codes are state-of-the-art implementations of the routines, so few improvements could be initially expected. Experiments have been designed and conducted to emulate three different scenarios: a single program running in the system with full resources, an interactive server where multiple programs run concurrently varying the availability of resources, and a queue of tasks where granted resources are limited. The proposed algorithms have been able to produce significant benefits, especially in systems with higher latency penalties for remote accesses. When more than one benchmark is executed simultaneously, performance improvements have been obtained, reducing execution times up to 60%. In this kind of situation, the behaviour of the system is more critical, and the NUMA topology plays a more relevant role. Even in the worst case, when isolated benchmarks are executed using the whole system, that is, just one task at a time, the performance is not degraded. Ruben Laso, Oscar G. Lorenzo, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 1 |
| 2021 | LBMA and IMAR2: Weighted lottery based migration strategies for NUMA multiprocessing serversabstractSummary Multicore NUMA systems present on‐board memory hierarchies and communication networks that influence performance when executing shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this article, monitoring information extracted from hardware counters at runtime is used to characterize the behavior of each thread for an arbitrary number of multithreaded processes running in a multiprocessing environment. This characterization is given in terms of number of operations per second, operational intensity, and latency of memory accesses. We propose a runtime tool, executed in user space, that uses this information to guide two different thread migration strategies for improving execution efficiency by increasing locality and affinity without requiring any modification in the running codes. Different configurations of NAS Parallel OpenMP benchmarks running concurrently on multicore NUMA systems were used to validate the benefits of our proposal, in which up to four processes are running simultaneously. In more than the 95% of the executions of our tool, results outperform those of the operating system (OS) and produces up to 38% improvement in execution time over the OS for heterogeneous workloads, under different and realistic locality and affinity scenarios. Ruben Laso, Oscar G. Lorenzo, Francisco F. Rivera, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | IHP: a dynamic heterogeneous parallel scheme for iterative or time-step methods - image denoising as case study
Ruben Laso, José Carlos Cabaleiro, Francisco F. Rivera, M. Carmen Muñiz, José A. Álvarez-Dios |
J. Supercomput. | 1 |