VLDB 2026 Research / reviewers in the wild / expert
Cecilio C. Tamarit
dblp:357/2929
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0003-1668-9677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Increasing the Efficiency of Associative Processor Architectures via CMOS-Compatible HybridizationabstractWe present a hybrid, general-purpose, associative processing-in-memory architecture that combines the energy and area advantages of a primary FeFET-based CAM array with the write performance and endurance of a much smaller CMOS-based sidekick. The hybrid nature of the architecture is transparent to the programmer, who uses a RISC-V ISA with standard RVV vector extensions. Detailed SPICE- and system-level simulations show our hybrid design dramatically curbs the endurance disadvantages of a pure FeFET design and delivers, on average, 30% and 11% area and energy savings over a purely CMOS implementation, respectively, at a performance loss of barely 1% over pure CMOS. Socrates S. Wong, Cecilio C. Tamarit, Mohammad Mehdi Sharifi, Zephan M. Enciso, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, José F. Martínez |
DATE | 2 |
| 2026 | BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory Processing
Cecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. Martínez |
ISCA | 1 |
| 2023 | PUMICE: Processing-using-Memory Integration with a Scalar Pipeline for Symbiotic ExecutionabstractExisting SIMD extensions in scalar CPUs (e.g., SSE, AVX, etc.) can leverage instruction-level parallelism (ILP) because of their tight integration with the CPU pipeline. However, the vectors they employ are quite short, and this limits their ability to exploit data-level parallelism (DLP). On the other hand, processing-using-memory (PUM) accelerators are capable of exploiting massive amounts of DLP, as they typically perform computation on very long vectors (tens of thousands of elements) within the memory itself. Recent work demonstrates that order-of-magnitude speedups can be achieved by these architectures for a variety of workloads over area-equivalent multicore CPUs with SIMD extensions. Still, PUM architectures are largely decoupled from the CPU itself, thereby limiting their ability to tap the CPU’s ILP the way SIMD extensions do.In this paper, we propose PUMICE, a tightly integrated CPU-PUM architecture that simultaneously exploits DLP and ILP for very long vector operations. As a result of this tight integration, PUMICE delivers significant performance gains: Our experimental results show speedups of up to 2.2× (1.4× on average) over a state-of-the-art decoupled approach. Socrates S. Wong, Cecilio C. Tamarit, José F. Martínez |
DAC | 2 |