Carlos Rojas 0001

dblp:49/6791-1 · also Carlos Rojas Morales · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-7714-0277ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 75% Hardware accelerators and domain-specific architectures · 20% High-performance computing · 3%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
genomics accelerator
0.812024
QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis Algorithms · ISCA 2024
Processor architecture and microarchitecture › SIMD
vector instructions
0.812024
QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis Algorithms · ISCA 2024
Processor architecture and microarchitecture
instruction set architecture
0.712023
Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023
Processor architecture and microarchitecture › instruction set architecture
RISC-V
0.712023
Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023
Processor architecture and microarchitecture › instruction set architecture
vector extension
0.712023
Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023
Processor architecture and microarchitecture
instruction issue logic
0.612022
Adaptable Register File Organization for Vector Processors · HPCA 2022
Processor architecture and microarchitecture › out-of-order execution
issue queue
0.612022
Adaptable Register File Organization for Vector Processors · HPCA 2022
Processor architecture and microarchitecture
vector processor
0.612022
Adaptable Register File Organization for Vector Processors · HPCA 2022
Processor architecture and microarchitecture › register file
vector register file
0.612022
Adaptable Register File Organization for Vector Processors · HPCA 2022
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.212024
QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis Algorithms · ISCA 2024
Processor architecture and microarchitecture
out-of-order execution
0.212023
Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023
Processor architecture and microarchitecture › out-of-order execution
register renaming
0.212023
Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications · ACM Trans. Archit. Code Optim. 2023
Performance modeling and evaluation
simulation
0.212022
Adaptable Register File Organization for Vector Processors · HPCA 2022

Methods — techniques the papers use, named apart from their topics

hardware-software co-design · 1.5gem5 simulation · 0.6McPAT modeling · 0.6
YearPublicationVenuePosition
2026 UPSORT: Design and Analysis of Processing-in-Memory Sorting Algorithms using UPMEM
abstract
The rapid growth of data-intensive applications has motivated the exploration of Processing-in-Memory (PIM) as a means to overcome memory bandwidth limitations in modern systems. By integrating processing units directly into DRAM chips, PIM enables computation close to the row buffers, reducing costly data transfers to the CPU and alleviating the memory wall. Within this context, sorting is one of the most fundamental building blocks in databases, analytics, and large-scale computing. Despite its bandwidth-intensive nature, it remains largely unexplored on commercial PIM hardware. In this work, we present UPSort, the first sorting algorithm designed and implemented on commercial PIM DIMMs. UPSort adopts a PIM-CPU cooperative model, leveraging PIM’s data-level parallelism for pre-sorting large blocks while delegating the final merging to the CPU. This design overcomes key architectural constraints of PIM DIMMs, including limited communication and fine-grained data transfers. Our evaluation on UPMEM PIM demonstrates that UPSort achieves up to $7.2 \times$ speedup over qsort, and on average $3.9 \times$ improvement compared to state-of-the-art algorithms such as BlockQuicksort and Quadsort. These results highlight the potential of PIM-CPU cooperative execution as a practical and efficient strategy for accelerating fundamental algorithms in future memory-centric systems.
Iván Vargas 0001, Julian Pavon, Carlos Rojas 0001, Javier Chavez Rubio, Marco A. Ramírez 0001, Luis A. Villa-Vargas, Mateo Valero, Osman S. Unsal, Adrián Cristal
ISPASS3
2024 QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis Algorithms
abstract
Genome sequence analysis is fundamental to medical breakthroughs such as developing vaccines, enabling genome editing, and facilitating personalized medicine. The exponentially expanding sequencing datasets and complexity of sequencing algorithms necessitate performance enhancements. While the performance of software solutions is constrained by their underlying hardware platforms, the utility of fixed-function accelerators is restricted to only certain sequencing algorithms.This paper presents QUETZAL, the first general-purpose vector acceleration framework designed for high efficiency and broad applicability across a diverse set of genomics algorithms. While a commercial CPU’s vector datapath is a promising candidate to exploit the data-level parallelism in genomics algorithms, our analysis finds that its performance is often limited due to long-latency scatter/gather memory instructions. QUETZAL introduces a hardware-software co-design comprising an accelerator microarchitecture closely integrated with the CPU’s vector datapath, alongside novel vector instructions to fully capitalize on the proposed hardware. QUETZAL integrates a set of scratchpad-style buffers meticulously designed to minimize latency associated with scatter/gather instructions during the retrieval of input genome sequences data. QUETZAL supports both short and long reads, and different types of sequencing data formats. A combination of hardware and software techniques enables QUETZAL to reduce the latency of memory instructions, perform complex computation using a single instruction, and transform data representations at runtime, resulting in overall efficiency gain. QUETZAL significantly accelerates a vectorized CPU baseline on modern genome sequence analysis algorithms by 5.7×, while incurring a small area overhead of 1.4% post place-and-route at the 7nm technology node compared to an HPC ARM CPU.
Julian Pavon, Iván Vargas Valdivieso, Carlos Rojas 0001, César Hernández, Mehmet Aslan, Roger Figueras, Yichao Yuan, Joël Lindegger, Mohammed Alser, Francesc Moll, Santiago Marco-Sola, Oguz Ergin, Nishil Talati, Onur Mutlu, Osman S. Unsal, Mateo Valero, Adrián Cristal
ISCA3
2023 Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications
abstract
The maturity level of RISC-V and the availability of domain-specific instruction set extensions, like vector processing, make RISC-V a good candidate for supporting the integration of specialized hardware in processor cores for the High Performance Computing (HPC) application domain. In this article, 1 we present Vitruvius+, the vector processing acceleration engine that represents the core of vector instruction execution in the HPC challenge that comes within the EuroHPC initiative. It implements the RISC-V vector extension (RVV) 0.7.1 and can be easily connected to a scalar core using the Open Vector Interface standard. Vitruvius+ natively supports long vectors: 256 double precision floating-point elements in a single vector register. It is composed of a set of identical vector pipelines (lanes), each containing a slice of the Vector Register File and functional units (one integer, one floating point). The vector instruction execution scheme is hybrid in-order/out-of-order and is supported by register renaming and arithmetic/memory instruction decoupling. On a stand-alone synthesis, Vitruvius+ reaches a maximum frequency of 1.4 GHz in typical conditions (TT/0.80V/25°C) using GlobalFoundries 22FDX FD-SOI. The silicon implementation has a total area of 1.3 mm 2 and maximum estimated power of ∼920 mW for one instance of Vitruvius+ equipped with eight vector lanes.
Francesco Minervini, Oscar Palomar, Osman S. Unsal, Enrico Reggiani, Josue V. Quiroga, Joan Marimon, Carlos Rojas 0001, Roger Figueras, Abraham Ruiz, Alberto González 0004, Jonnatan Mendoza, Iván Vargas 0001, César Hernández, Joan Cabre, Lina Khoirunisya, Mustapha Bouhali, Julian Pavon, Francesc Moll, Mauro Olivieri, Mario Kovac, Mate Kovac, Leon Dragic, Mateo Valero, Adrián Cristal
ACM Trans. Archit. Code Optim.7
2022 Adaptable Register File Organization for Vector Processors
abstract
Contemporary Vector Processors (VPs) are de-signed either for short vector lengths, e.g., Fujitsu A64FX with 512-bit ARM SVE vector support, or long vectors, e.g., NEC Aurora Tsubasa with 16Kbits Maximum Vector Length (MVL1). Unfortunately, both approaches have drawbacks. On the one hand, short vector length VP designs struggle to provide high efficiency for applications featuring long vectors with high Data Level Parallelism (DLP). On the other hand, long vector VP designs waste resources and underutilize the Vector Register File (VRF) when executing low DLP applications with short vector lengths. Therefore, those long vector VP implementations are limited to a specialized subset of applications, where relatively high DLP must be present to achieve excellent performance with high efficiency. Modern scientific applications are getting more diverse, and the vector lengths in those applications vary widely. To overcome these limitations, we propose an Adaptable Vector Architecture (AVA) that leads to having the best of both worlds. AVA is designed for short vectors (MVL=16 elements) and is thus area and energy-efficient. However, AVA has the functionality to reconfigure the MVL, thereby allowing to exploit the benefits of having a longer vector of up to 128 elements microarchitecture when abundant DLP is present. We model AVA on the gem5 simulator and evaluate AVA performance with six applications taken from the RiVEC Benchmark Suite. To obtain area and power consumption metrics, we model AVA on McPAT for 22nm technology. Our results show that by reconfiguring our small VRF (8KB) plus our novel issue queue scheme, AVA yields a 2X speedup over the default configuration for short vectors. Additionally, AVA shows competitive performance when compared to a long vector VP, while saving 50% of area.
Cristóbal Ramírez, Enrico Reggiani, Carlos Rojas 0001, Roger Figueras, Luis A. Villa-Vargas, Marco A. Ramírez 0001, Mateo Valero, Osman S. Unsal, Adrián Cristal
HPCA3