VLDB 2026 Research / reviewers in the wild / expert
Víctor Soria 0001
dblp:280/2111 · also Víctor Soria Pardos
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-8337-6326ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ember: A Compiler for Embedding Operations on Decoupled Access-Execute ArchitecturesabstractDecoupled Access-Execute (DAE) architectures separate memory accesses from computation in two specialized units. This design is becoming increasingly popular among hyperscalers to accelerate irregular embedding lookups in recommendation models. In this paper, we first broaden the scope by demonstrating the benefits of DAE architectures across a wider range of irregular embedding operations in several machine learning models. Then, we propose the Ember compiler to automatically compile all of these embedding operations to DAE architectures. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale. Marco Siracusa, Olivia Hsu, Víctor Soria 0001, Joshua Randall 0001, Arnaud Grasset, Eric Biscondi, Douglas J. Joseph, Randy Allen, Fredrik Kjolstad, Miquel Moretó, Adrià Armejach |
CGO | 3 |
| 2025 | FLAMA: Architecting Floating-Point Atomic Memory Operations for Heterogeneous HPC SystemsabstractCurrent heterogeneous systems integrate generalpurpose Central Processing Units (CPUs), Graphics Processing Units (GPUs), and Neural Processing Units (NPUs). The efficient use of such systems requires a significant programming effort to distribute computation and synchronize across devices, which usually involves using Atomic Memory Operations (AMOs). Arm recently launched a floating-point Atomic Memory Operations (FAMOs) extension to perform atomic updates on floating-point data types specifically. This work characterizes and models heterogeneous architectures to understand how floating-point AMOs impact graph, Machine Learning (ML), and high-performance computing (HPC) workloads. Our analysis shows that many AMOs are performed on floating-point data, which modern systems execute using inefficient compare-and-swap (CAS) constructs. Therefore, replacing CASbased constructs with FAMOs can improve a wide range of workloads. Moreover, we analyze the trade-offs of executing FAMOs at different memory hierarchy levels, either in private caches (near) or remotely in shared caches (far). We have extended the widely used AMBA CHI protocol to evaluate such FAMO support on a simulated chiplet-based heterogeneous architecture. While near FAMOs achieve an average $1.34 \times$ speed-up, far FAMOs reach an average $1.58 \times$ speed-up. We conclude that FAMOs can bridge the gap between CPU architecture and accelerators and enabling synchronization in key application domains. Víctor Soria 0001, Adrià Armejach, Darío Suárez Gracia, Didier Martinot, Arnaud Grasset, Miquel Moretó |
DSD | 1 |
| 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsabstractThe irruption of chiplet-based architectures has been a game changer, enabling higher transistor integration and core counts in a single socket.However, chiplets impose higher and non-uniform memory access (NUMA) latencies than monolithic integration.This harms the efficiency of atomic memory operations (AMOs), which are fundamental to implementing fine-grained synchronization and concurrent data structures on large systems.AMOs are executed either near the core (near) or at a remote location within the cache hierarchy (far).On near AMOs, the core's private cache fetches the target cache line in exclusiveness to modify it locally.Near AMOs cause significant data movement between private caches, especially harming parallel applications' performance on chiplet-based architectures.Alternatively, far AMOs can alleviate the communication overhead by reducing data movement between processing elements.However, current multicore architectures only support one type of far AMO, which sends all updates to a single serialization point (centralized AMOs).This work introduces two new types of far AMOs, delegated and migrating, that execute AMOs remotely without centralizing updates in a single point of the cache hierarchy.Combining centralized, delegated, and migrating AMOs allows the directory to select the best location to execute AMOs.Moreover, we propose Delegato, a tracing optimization to effectively transport usage information from private caches to the directory to predict the best atomic type to issue accurately.Additionally, we design a simple predictor on Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Miquel Moretó |
MICRO | 1 |
| 2024 | GenArchBench: A genomics benchmark suite for arm HPC processorsabstractArm usage has substantially grown in the High-Performance Computing (HPC) community. Japanese supercomputer Fugaku, powered by Arm-based A64FX processors, held the top position on the Top500 list between June 2020 and June 2022, currently sitting in the fourth position. The recently released 7th generation of Amazon EC2 instances for compute-intensive workloads (C7 g) is also powered by Arm Graviton3 processors. Projects like European Mont-Blanc and U.S. DOE/NNSA Astra are further examples of Arm irruption in HPC. In parallel, over the last decade, the rapid improvement of genomic sequencing technologies and the exponential growth of sequencing data has placed a significant bottleneck on the computational side. While most genomics applications have been thoroughly tested and optimized for x86 systems, just a few are prepared to perform efficiently on Arm machines. Moreover, these applications do not exploit the newly introduced Scalable Vector Extensions (SVE). This paper presents GenArchBench, the first genome analysis benchmark suite targeting Arm architectures. We have selected computationally demanding kernels from the most widely used tools in genome data analysis and ported them to Arm-based A64FX and Graviton3 processors. Overall, the GenArch benchmark suite comprises 13 multi-core kernels from critical stages of widely-used genome analysis pipelines, including base-calling, read mapping, variant calling, and genome assembly. Our benchmark suite includes different input data sets per kernel (small and large), each with a corresponding regression test to verify the correctness of each execution automatically. Moreover, the porting features the usage of the novel Arm SVE instructions, algorithmic and code optimizations, and the exploitation of Arm-optimized libraries. We present the optimizations implemented in each kernel and a detailed performance evaluation and comparison of their performance on four different HPC machines (i.e., A64FX, Graviton3, Intel Xeon Skylake Platinum, and AMD EPYC Rome). Overall, the experimental evaluation shows that Graviton3 outperforms other machines on average. Moreover, we observed that the performance of the A64FX is significantly constrained by its small memory hierarchy and latencies. Additionally, as proof of concept, we study the performance of a production-ready tool that exploits two of the ported and optimized genomic kernels. Lorién López-Villellas, Rubén Langarita, Asaf Badouh, Víctor Soria 0001, Quim Aguado-Puig, Guillem López-Paradís, Max Doblas, Javier Setoain, Chulho Kim, Makoto Ono, Adrià Armejach, Santiago Marco-Sola, Jesús Alastruey-Benedé, Pablo Ibáñez 0001, Miquel Moretó |
Future Gener. Comput. Syst. | 4 |
| 2023 | DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory OperationsabstractWith increasing core counts in modern multi-core designs, the overhead of synchronization jeopardizes the scalability and efficiency of parallel applications. To mitigate these overheads, modern cache-coherent protocols offer support for Atomic Memory Operations (AMOs) that can be executed near-core (near) or remotely in the on-chip memory hierarchy (far). Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Alejandro Rico, Miquel Moretó |
ISCA | 1 |
| 2023 | A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsabstractThis paper proposes the Tensor Marshaling Unit (TMU), a near-core programmable dataflow engine for multicore architectures that accelerates tensor traversals and merging, the most critical operations of sparse tensor workloads running on today’s computing infrastructures. The TMU leverages a novel multi-lane design that enables parallel tensor loading and merging, which naturally produces vector operands that are marshaled into the core for efficient SIMD computation. The TMU supports all the necessary primitives to be tensor-format and tensor-algebra complete. We evaluate the TMU on a simulated multicore system using a broad set of tensor algebra workloads, achieving 3.6 ×, 2.8 ×, and 4.9 × speedups over memory-intensive, compute-intensive, and merge-intensive vectorized software implementations, respectively. Marco Siracusa, Víctor Soria 0001, Francesco Sgherzi, Joshua Randall 0001, Douglas J. Joseph, Miquel Moretó, Adrià Armejach |
MICRO | 2 |
| 2022 | Sargantana: A 1 GHz+ In-Order RISC-V Processor with SIMD Vector Extensions in 22nm FD-SOIabstractThe RISC-V open Instruction Set Architecture (ISA) has proven to be a solid alternative to licensed ISAs. In the past 5 years, a plethora of industrial and academic cores and accelerators have been developed implementing this open ISA. In this paper, we present Sargantana, a 64-bit processor based on RISC-V that implements the RV64G ISA, a subset of the vector instructions extension (RVV 0.7.1), and custom application-specific instructions. Sargantana features a highly optimized 7-stage pipeline implementing out-of-order write-back, register renaming, and a non-blocking memory pipeline. Moreover, Sar-gantana features a Single Instruction Multiple Data (SIMD) unit that accelerates domain-specific applications. Sargantana achieves a 1.26 GHz frequency in the typical corner, and up to 1.69 GHz in the fast corner using 22nm FD-SOI commercial technology. As a result, Sargantana delivers a 1.77× higher Instructions Per Cycle (IPC) than our previous 5-stage in-order DVINO core, reaching 2.44 CoreMark/MHz. Our core design delivers comparable or even higher performance than other state-of-the-art academic cores performance under Autobench EEMBC benchmark suite. This way, Sargantana lays the foundations for future RISC-V based core designs able to meet industrial-class performance requirements for scientific, real-time, and high-performance computing applications. Víctor Soria 0001, Max Doblas, Guillem López-Paradís, Gerard Candón, Narcís Rodas, Xavier Carril, Pau Fontova, Neiel Leyva, Santiago Marco-Sola, Miquel Moretó |
DSD | 1 |
| 2021 | On the use of many-core Marvell ThunderX2 processor for HPC workloads
Víctor Soria 0001, Adrià Armejach, Darío Suárez Gracia, Miquel Moretó |
J. Supercomput. | 1 |