EDBT 2026 Demo / reviewers in the wild / expert
Paulo C. Santos 0001
dblp:125/2570 · also Paulo Cesar Santos 0001
· DBLP profile ↗
18ranked-venue papers
7as first author
10since 2021 · last 2024
0000-0001-8555-2637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring compiler optimization space for control flow obfuscation
Hameeza Ahmed, Muhammad Faraz Hyder, Muhammad Fahim Ul Haque, Paulo C. Santos 0001 |
Comput. Secur. | 4 |
| 2024 | HBPB, applying reuse distance to improve cache efficiency proactively
Arthur M. Krause, Paulo C. Santos 0001, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux |
J. Parallel Distributed Comput. | 2 |
| 2023 | Plug N' PIM: An integration strategy for Processing-in-Memory accelerators
Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
Integr. | 1 |
| 2022 | Aggressive Performance Improvement on Processing-in-Memory Devices by Adopting HugepagesabstractProcessing-in-Memory (PIM) devices integrated into general-purpose systems demand virtual memory support. In this way, these devices can be seamlessly coupled to the software stack, while maintaining compatibility and security provided by address management via the Operating System (OS) without requiring disruptive programming efforts. Typically, PIM intends to access large volumes of data via vector operations, and thus can suffer severe penalties due to the high cost of page misses in the Translation Look-aside Buffer (TLB). Our study demonstrates the criticality of such penalties on the system's performance and that PIM must resort to large page sizes. The presented results exploit the native large pages available on the host, and they show substantial performance improvements$(84\times)$for wide-vector PIM operations with large pages. Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
ASAP | 1 |
| 2022 | Advancing Near-Data Processing with Precise Exceptions and Efficient Data FetchingabstractNear-Data Processing (NDP) modifies the traditional computer system design by placing logic near the memory, bringing computation to the data. One NDP approach places such elements on the logic layer of 3D-stacked memories to quickly access data while avoiding reliance on narrow buses and better accessing the parallelism these devices offer. However, NDP architectures often fail to fully leverage available memory resources. In this work, we propose adding an instruction buffer to a common NDP design with large vector instructions. This modification allows the NDP to fetch instruction operands out of program order and delegates some responsibility regarding precise exceptions to the near-data device. Our results show our modifications cause a reduction in execution time of up to 28% while consuming up to 25% less energy. Sairo R. dos Santos, Tiago Rodrigo Kepe, Francis B. Moreira 0001, Paulo C. Santos 0001, Marco A. Z. Alves |
ISPASS | 4 |
| 2022 | Avoiding Unnecessary Caching with History-Based Preemptive BypassingabstractCache memories can account for more than half of the area and energy consumption on modern processors, which will only increase with the current trend of bigger on die memories. Although these components are very effective when the access pattern is cache-friendly, cache memories incur extra and unnecessary latencies when they cannot serve the data, which adds to significant energy wastes when data that is never reused is placed on them. This work introduces HBPB, a mechanism that detects whether a memory access is cache friendly or not, allowing the bypass of the cache for accesses that are not known to be cache-friendly. Our approach allows the processor to quickly detect when caching accesses is inadequate, improving overall access latency and reducing energy waste and cache pollution. The presented solution achieves reductions of up to 28.6% in energy consumption and 19.5% in latency for SPEC applications, and further improvements in power and performance across various workloads. Arthur M. Krause, Paulo C. Santos 0001, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |
| 2022 | Sim2PIM: A complete simulation framework for Processing-in-Memory
Bruno Endres Forlin, Paulo C. Santos 0001, Augusto E. Becker, Marco A. Z. Alves, Luigi Carro |
J. Syst. Archit. | 2 |
| 2021 | Providing Plug N' Play for Processing-in-Memory AcceleratorsabstractAlthough Processing-in-Memory (PIM) emerged as a solution to avoid unnecessary and expensive data movements to/from host and accelerators, their widespread usage is still difficult, given that to effectively use a PIM device, huge and costly modifications must be done at the host processor side to allow instructions offloading, cache coherence, virtual memory management, and communication between different PIM instances. The present work addresses these challenges by presenting non-invasive solutions for those requirements. We demonstrate that, at compile-time, and without any host modifications or programmer intervention, it is possible to exploit already available resources to allow efficient host and PIM communication and task partitioning, without disturbing neither host nor memory hierarchy. We present Plug&PIM, a plug n' play strategy for PIM adoption with minimal performance penalties. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
ASP-DAC | 1 |
| 2021 | Sim2PIM: A Fast Method for Simulating Host Independent & PIM Agnostic DesignsabstractProcessing-in-Memory (PIM), with the help of modern memory integration technologies, has emerged as a practical approach to mitigate the memory wall and improve performance and energy efficiency in contemporary applications. However, there is a need for tools capable of quickly simulating different PIMs designs and their suitable integration with different hosts. This work presents Sim2PIM, a Simple Simulator for PIM devices that seamlessly integrates any PIM architecture with the host processor and memory hierarchy. Sim2PIM's simulation environment allows the user to describe a PIM architecture in different user-defined abstraction levels. The application code runs natively on the Host, with minimal overhead from the simulator integration, allowing Sim2PIM to collect precise metrics from the Hardware Performance Counters (HPCs). Our simulator is available to download at https://pim.computer/. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
DATE | 1 |
| 2021 | Machine Learning Migration for Efficient Near-Data ProcessingabstractMachine Learning (ML) rises as a highly useful tool to analyze the vast amount of data generated in every field of science nowadays. Simultaneously, data movement inside computer systems gains more focus due to its high impact an time and energy consumption. In this context, the Near-Data Processing (NDP) architectures emerged as a prominent solution to increasing data by drastically reducing the required amount of data movement For NDP, we see three main approaches, Application-Specific Integrated Circuits (ASICs), full Central Processing Units (CPUs) and Graphics Processing Units (GPIIs), or vector units integration. However, previous work considered only ASICs, CPUs and GPUs when executing ML algorithms inside the memory. In this paper, we present an approach to execute ML algorithms near-data, using a general-purpose vector architecture and applying near-data parallelism to kernels from KNN, MEP, and CNN algorithms. To facilitate this process, we also present an NDP intrinsics library to ease the evaluation and debugging tasks. Our results show speedups up to fox for KNN, 11× for MLP, and 3× for convolution when processing near-data compared to a high-performance ×86 baseline. Aline S. Cordeiro, Sairo R. dos Santos, Francis B. Moreira 0001, Paulo C. Santos 0001, Luigi Carro, Marco A. Z. Alves |
PDP | 4 |
| 2019 | A Compiler for Automatic Selection of Suitable Processing-in-Memory InstructionsabstractAlthough not a new technique, due to the advent of 3D-stacked technologies, the integration of large memories and logic circuitry able to compute large amount of data has revived the Processing-in-Memory (PIM) techniques. PIM is a technique to increase performance while reducing energy consumption when dealing with large amounts of data. Despite several designs of PIM are available in the literature, their effective implementation still burdens the programmer. Also, various PIM instances are required to take advantage of the internal 3D-stacked memories, which further increases the challenges faced by the programmers. In this way, this work presents the Processing-In-Memory cOmpiler (PRIMO). Our compiler is able to efficiently exploit large vector units on a PIM architecture, directly from the original code. PRIMO is able to automatically select suitable PIM operations, allowing its automatic offloading. Moreover, PRIMO concerns about several PIM instances, selecting the most suitable instance while reduces internal communication between different PIM units. The compilation results of different benchmarks depict how PRIMO is able to exploit large vectors, while achieving a near-optimal performance when compared to the ideal execution for the case study PIM. PRIMO allows a speedup of 38× for specific kernels, while on average achieves 11.8 × for a set of benchmarks from PolyBench Suite. Hameeza Ahmed, Paulo C. Santos 0001, João Paulo C. de Lima, Rafael Fao de Moura, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 2 |
| 2018 | Design space exploration for PIM architectures in 3D-stacked memoriesabstractScaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3D-stacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-of-the-art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High-Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth. João Paulo C. de Lima, Paulo C. Santos 0001, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
CF | 2 |
| 2018 | Processing in 3D memories to speed up operations on complex data structuresabstractPointer chasing has been, for years, the kernel operation employed by diverse data structures, from graphs to hash tables and dictionaries. However, due to the bewildering growth in the volume of data that current applications have to deal with, performing pointer chasing operations have become a major source of performance and energy bottleneck, due to its sparse memory access behavior. In this work, we aim to tackle this problem by taking advantage of the already available parallelism present in today's 3D-stacked memories. We present a simple mechanism that can accelerate pointer chasing operations by making use of a state-of-the-art PIM design that executes in-memory vector operations. The key idea behind our design is to run speculative loads, in parallel, based on a given memory address in a reconfigurable window of addresses. Our design can perform pointer-chasing operations on b+tree 4.9 χ faster when compared to modern baseline systems. Besides that, since our device avoids data movement, we can also reduce energy consumption by 85% when compared to the baseline. Paulo C. Santos 0001, Geraldo F. Oliveira, João Paulo C. de Lima, Marco A. Z. Alves, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 1 |
| 2018 | HIPE: HMC instruction predication extension applied on database processingabstractThe recent Hybrid Memory Cube (HMC) is a smart memory which includes functional units inside one logic layer of the 3D stacked memory design. In order to execute instructions inside the Hybrid Memory Cube (HMC), the processor needs to send instructions to be executed near data, keeping most of the pipeline complexity inside the processor. Thus, control-flow and data-flow dependencies are all managed inside the processor, in such way that only update instructions are supported by the HMC. In order to solve data-flow dependencies inside the memory, previous work proposed HMC Instruction Vector Extensions (HIVE), which embeds a high number of functional units with a interlock register bank. In this work we propose HMC Instruction Prediction Extensions (HIPE), that supports predicated execution inside the memory, in order to transform control-flow dependencies into data-flow dependencies. Our mechanism focus on removing the high latency iteration between the processor and the smart memory during the execution of branches that depends on data processed inside the memory. In this paper we evaluate a balanced design of HIVE comparing to x86 and HMC executions. After we show the HIPE mechanism results when executing a database workload, which is a strong candidate to use smart memories. We show interesting trade-offs of performance when comparing our mechanism to previous work. Diego G. Tomé, Paulo C. Santos 0001, Luigi Carro, Eduardo C. de Almeida, Marco A. Z. Alves |
DATE | 2 |
| 2017 | Operand size reconfiguration for big data processing in memoryabstractNowadays, applications that predominantly perform lookups over large databases are becoming more popular with column-stores as the database system architecture of choice. For these applications, Hybrid Memory Cubes (HMCs) can provide bandwidth of up to 320 GB/s and represents the best choice to keep the throughput for these ever increasing databases. However, even with the high available memory bandwidth and processing power, in order to achieve the peak performance, data movements through the memory hierarchy consumes an unnecessary amount of time and energy. In order to accelerate database operations, and reduce the energy consumption of the system, this paper presents the Reconfigurable Vector Unit (RVU) that enables massive and adaptive in-memory processing, extending the native HMC instructions and also increasing its effectiveness. RVU enables the programmer to reconfigure it to perform as a large vector unit or multiple small vectors units to better adjust for the application needs during different computation phases. Due to its adaptability, RVU is capable of achieving performance increase of 27 χ on average and reduce the DRAM energy consumption in 29% when compared to an x86 processor with 16 cores. Compared with the state-of-the-art mechanism capable of performing large vector operations with fixed size, inside the HMC, RVU performed up to 12% better in terms of performance and improve in 53% the energy consumption. Paulo C. Santos 0001, Geraldo F. Oliveira, Diego G. Tomé, Marco A. Z. Alves, Eduardo C. de Almeida, Luigi Carro |
DATE | 1 |
| 2016 | Large vector extensions inside the HMC
Marco A. Z. Alves, Matthias Diener, Paulo C. Santos 0001, Luigi Carro |
DATE | 3 |
| 2016 | Exploring Cache Size and Core Count Tradeoffs in Systems with Reduced Memory Access LatencyabstractOne of the main challenges for computer architects is how to hide the high average memory access latency from the processor. In this context, Hybrid Memory Cubes (HMCs) can provide substantial energy and bandwidth improvements compared to traditional memory organizations. However, it is not clear how this reduced average memory access latency will impact the LLC. For applications with high cache miss ratios, the latency to search for the data inside the cache memory will impact negatively on the performance. The importance of this overhead depends on the memory access latency. In this paper, we present an evaluation of the L3 cache importance on a high performance processor using HMC also exploring chip area tradeoffs between the cache size and number of processor cores. We show that the high bandwidth provided by HMC memories can eliminate the need for L3 caches, removing hardware and making room for more processing power. Our evaluations show that performance increased 37% and the EDP improved 12% while maintaining the same original chip area in a wide range of parallel applications, when compared to DDR3 memories. Paulo C. Santos 0001, Marco A. Z. Alves, Matthias Diener, Luigi Carro, Philippe Olivier Alexandre Navaux |
PDP | 1 |
| 2015 | Saving memory movements through vector processing in the DRAMabstractDespite the ability of modern processors to execute a variety of algorithms efficiently through instructions based on registers with ever-increasing widths, some applications present poor performance due to the limited interconnection bandwidth between main memory and processing units. Near-data processing has started to gain acceptance as an accelerator device due to the technology constraints and high costs associated with data transfer. However, previous approaches to near-data computing do not provide general-purpose processing, or require large amounts of logic and do not fully use the potential of the DRAM devices. These issues limited its wide adoption. In this paper, we present the Memory Vector Extensions (MVX), which implement vector instructions directly inside the DRAM devices, therefore avoiding data movement between memory and processing units, while requiring a lower amount of logic than previous approaches. MVX is able to obtain up to 211× increase in performance for application kernels with a high spatial locality and a low temporal locality. Comparing to an embedded processor with 8 cores and 2 memory channels that supports AVX-512 instructions, MVX performs 24× faster on average for three well known algorithms. Marco A. Z. Alves, Paulo C. Santos 0001, Francis B. Moreira 0001, Matthias Diener, Luigi Carro |
CASES | 2 |