Marco A. Z. Alves

dblp:47/7939 · also Marco Antonio Zanata Alves · DBLP profile ↗
← Back
43ranked-venue papers
4as first author
15since 2021 · last 2023
0000-0003-2440-2664ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2023 NoGar: A Non-cooperative Game for Thread Pinning in Array Databases
Simone Dominico, Marco A. Z. Alves, Eduardo C. de Almeida
DEXA (1)2
2023 Efficient Prequential AUC-PR Computation
abstract
When dealing with classification problems for data streams, we often need to compute the classification metrics in a prequential manner. The Area Under the Precision-Recall Curve (AUC-PR) metric is extensively used in imbalanced classification scenarios, where the negative class outnumbers the positive one. Despite its advantages, it may be computationally expensive to recompute that metric every time a new test instance becomes available. In this work, we present an efficient algorithm to compute the AUC-PR in a prequential way. Our proposed algorithm uses a self-balancing binary search tree to avoid the need to reorder the data when updating the AUC-PR value with the most recent data. Our experiments take into consideration six well-known, publicly available stream-based datasets. Our experiments show that our approach can be up to 13 times faster and use 12 times less energy than the traditional batch approach when considering a window of size 1,000.
David L. Pereira Gomes, André Ricardo Abed Grégio, Marco A. Z. Alves, Paulo R. L. Almeida
ICMLA3
2023 Improved Computation of Database Operators via Vector Processing Near-Data
abstract
Data-centric applications are increasingly more common, causing issues brought on by the discrepancy between processor and memory technologies to be increasingly more apparent. Near-Data Processing (NDP) is an approach to mitigate this issue. It proposes moving some of the computation close to the memory, thus allowing for reduced data movement and aiding data-intensive workloads. Analytical database queries are very commonly used in NDP research due to their intrinsics usage of very large volumes of data. In this paper, we investigate the migration of most time-consuming database operators to VIMA, a novel 3D-stacked memory-based NDP architecture. We consider the selection, projection, and bloom join database query operators, commonly used by data analytics applications, comparing Vector-In-Memory Architecture (VIMA) to a high-performance x86 baseline. We pitch VIMA against both a single-thread baseline and a modern 16-thread x86 system to evaluate its performance. Against a single-thread baseline, our experiments show that VIMA is able to speed up execution by up to 5× for selection, 2.5× for projection, and 16× for join while consuming up to 99% less energy. When considering a multi-thread baseline, VIMA matches the execution time performance even at the largest dataset sizes considered. In comparison to existing state-of-the-art NDP platforms, we find that our approach achieves superior performance for these operators.
Sairo R. dos Santos, Tiago Rodrigo Kepe, Marco A. Z. Alves
SBAC-PAD3
2023 Plug N' PIM: An integration strategy for Processing-in-Memory accelerators
Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro
Integr.3
2022 Aggressive Performance Improvement on Processing-in-Memory Devices by Adopting Hugepages
abstract
Processing-in-Memory (PIM) devices integrated into general-purpose systems demand virtual memory support. In this way, these devices can be seamlessly coupled to the software stack, while maintaining compatibility and security provided by address management via the Operating System (OS) without requiring disruptive programming efforts. Typically, PIM intends to access large volumes of data via vector operations, and thus can suffer severe penalties due to the high cost of page misses in the Translation Look-aside Buffer (TLB). Our study demonstrates the criticality of such penalties on the system's performance and that PIM must resort to large page sizes. The presented results exploit the native large pages available on the host, and they show substantial performance improvements$(84\times)$for wide-vector PIM operations with large pages.
Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro
ASAP3
2022 Video Decoder Improvements with Near-Data Speculative Motion Compensation Processing
abstract
Video decoder implementations are still evolving as they directly affect a large fraction of embedded systems nowadays. In this context, Versatile Video Coding (VVC) brings increased compression efficiency, which comes with extra over-head in terms of computational effort and energy consumption. At the same time, emerging Near-Data Processing (NDP) architectures promise drastic time and energy cuts for applications with data streaming behavior. In this paper, a speculative Motion Compensation (MC) is proposed to enable video decoders improvements through the exploitation of NDP. We adopted a large-vector SIMD-based NDP system (called VIMA) that provides high-performance operations over 2 K vectors. The proposed strategy leverages the correlation between the prediction modes and the motion data between spatially neighboring blocks within a frame to speculatively perform the MC for an entire region of 2Kx128 samples. MC interpolation kernels were implemented using VIMA and x86 AVX-256 SIMD libraries. Our NDP-based kernel implementation allows speedup of $1.9\times$ to $22\times$ compared to the x86 baseline solutions. Stepping forward, based on a coalescence estimation, our strategy can properly handle interpolation misses, achieving MC performance improvements from 7% to 64%.
Garrenlus de Souza, José Rodrigo Azambuja, Bruno Zatt, Marco A. Z. Alves, Sergio Bampi, Felipe Sampaio
ISCAS4
2022 Advancing Near-Data Processing with Precise Exceptions and Efficient Data Fetching
abstract
Near-Data Processing (NDP) modifies the traditional computer system design by placing logic near the memory, bringing computation to the data. One NDP approach places such elements on the logic layer of 3D-stacked memories to quickly access data while avoiding reliance on narrow buses and better accessing the parallelism these devices offer. However, NDP architectures often fail to fully leverage available memory resources. In this work, we propose adding an instruction buffer to a common NDP design with large vector instructions. This modification allows the NDP to fetch instruction operands out of program order and delegates some responsibility regarding precise exceptions to the near-data device. Our results show our modifications cause a reduction in execution time of up to 28% while consuming up to 25% less energy.
Sairo R. dos Santos, Tiago Rodrigo Kepe, Francis B. Moreira 0001, Paulo C. Santos 0001, Marco A. Z. Alves
ISPASS5
2022 Advancing Database System Operators with Near-Data Processing
abstract
As applications become more data-intensive, issues like von Neumann’s bottleneck and the memory wall became more apparent since data movement is the main source of inefficiency in computer systems. Looking to mitigate this issue, Near-Data Processing (NDP) moves computation from the processor to the memory, thus reducing the data movement required by many data-intensive workloads. In this paper, we look to database query operators, common targets of NDP research as database systems often need to deal with large amounts of data. We investigate the migration of most time-consuming database operators to Vector-In-Memory Architecture (VIMA), a novel 3D-stacked memory-based NDP architecture. We consider the selection, projection, and bloom join database query operators, commonly used by data analytics applications, comparing VIMA to a high-performance x86 baseline. Our results show speedups of up to 8× for selection, 6× for projection, and 16× for join while consuming up to 99% less energy. To the best of our knowledge, these results outperform the state-of-the-art for these operators on NDP platforms.
Sairo R. dos Santos, Francis B. Moreira 0001, Tiago Rodrigo Kepe, Marco A. Z. Alves
PDP4
2022 AntiViruses under the microscope: A hands-on perspective
Marcus Botacin, Felipe Duarte Domingues, Fabricio Ceschin, Raphael Machnicki, Marco A. Z. Alves, Paulo Lício de Geus, André Ricardo Abed Grégio
Comput. Secur.5
2022 HEAVEN: A Hardware-Enhanced AntiVirus ENgine to accelerate real-time, signature-based malware detection
Marcus Botacin, Marco A. Z. Alves, Daniela Oliveira 0001, André Ricardo Abed Grégio
Expert Syst. Appl.2
2022 Sim2PIM: A complete simulation framework for Processing-in-Memory
Bruno Endres Forlin, Paulo C. Santos 0001, Augusto E. Becker, Marco A. Z. Alves, Luigi Carro
J. Syst. Archit.4
2022 Terminator: A Secure Coprocessor to Accelerate Real-Time AntiViruses Using Inspection Breakpoints
abstract
AntiViruses (AVs) are essential to face the myriad of malware threatening Internet users. AVs operate in two modes: on-demand checks and real-time verification. Software-based real-time AVs intercept system and function calls to execute AV’s inspection routines, resulting in significant performance penalties as the monitoring code runs among the suspicious code. Simultaneously, dark silicon problems push the industry to add more specialized accelerators inside the processor to mitigate these integration problems. In this article, we propose Terminator , an AV-specific coprocessor to assist software AVs by outsourcing their matching procedures to the hardware, thus saving CPU cycles and mitigating performance degradation. We designed Terminator to be flexible and compatible with existing AVs by using YARA and ClamAV rules. Our experiments show that our approach can save up to 70 million CPU cycles per rule when outsourcing on-demand checks for matching typical, unmodified YARA rules against a dataset of 30 thousand in-the-wild malware samples. Our proposal eliminates the AV’s need for blocking the CPU to perform full system checks, which can now occur in parallel. We also designed a new inspection breakpoint mechanism that signals to the coprocessor the beginning of a monitored region, allowing it to scan the regions in parallel with their execution. Overall, our mechanism mitigated up to 44% of the overhead imposed to execute and monitor the SPEC benchmark applications in the most challenging scenario.
Marcus Botacin, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, André Ricardo Abed Grégio, Marco A. Z. Alves
ACM Trans. Priv. Secur.5
2021 Towards the Overcome of Performance Pitfalls in Data Stream Mining Tools
abstract
Data stream mining is an essential task in today's scientific community. It allows machine learning models to be updated over time as new data becomes available. Three pillars should be accounted for when selecting an appropriate algorithm for data stream mining: accuracy, processing time, and memory consumption. To develop and assess machine learning models in streaming scenarios, different tools have been developed, where the Massive Online Analysis, written in Java, and scikit-multiflow, written in Python, are in the spotlight. Despite the ease of use of both tools, neither are focused on performance, which puts in jeopardy the usage of the computational resources. In this paper, we show that with the right tools, Python libraries reach performance comparable to C/C++. More specifically, we show how optimized implementations in scikit-multiflow using low-level languages, i.e., C++, C++ with Intel Intrinsics, and Rust; with bindings to Python vastly overcome existing tools in computational resources usage while keeping predictive performance intact.
Lucca Portes Cavalheiro, Marco A. Z. Alves, Jean Paul Barddal
IJCNN2
2021 Machine Learning Migration for Efficient Near-Data Processing
abstract
Machine Learning (ML) rises as a highly useful tool to analyze the vast amount of data generated in every field of science nowadays. Simultaneously, data movement inside computer systems gains more focus due to its high impact an time and energy consumption. In this context, the Near-Data Processing (NDP) architectures emerged as a prominent solution to increasing data by drastically reducing the required amount of data movement For NDP, we see three main approaches, Application-Specific Integrated Circuits (ASICs), full Central Processing Units (CPUs) and Graphics Processing Units (GPIIs), or vector units integration. However, previous work considered only ASICs, CPUs and GPUs when executing ML algorithms inside the memory. In this paper, we present an approach to execute ML algorithms near-data, using a general-purpose vector architecture and applying near-data parallelism to kernels from KNN, MEP, and CNN algorithms. To facilitate this process, we also present an NDP intrinsics library to ease the evaluation and debugging tasks. Our results show speedups up to fox for KNN, 11× for MLP, and 3× for convolution when processing near-data compared to a high-performance ×86 baseline.
Aline S. Cordeiro, Sairo R. dos Santos, Francis B. Moreira 0001, Paulo C. Santos 0001, Luigi Carro, Marco A. Z. Alves
PDP6
2021 Performance Analysis of Array Database Systems in Non-Uniform Memory Architecture
abstract
Array Database Management Systems (Array databases) support query processing over multi-dimensional data. Data storage is implemented with non-linear structures to mitigate the shortcomings of the relational model when dealing with raw binary data, such as images, time series, and others. Due to data-hungry nature of multi-dimensional data applications, array databases must ideally provide a linear speedup when using a multi-processing system. When dealing with Non-Uniform Memory Access (NUMA) machines, array databases may require massive data movement across the nodes resulting in a severe performance impact, depending on the user operation. In this paper, we analyze the performance impact of the NUMA architecture in the SAVIME and SciDB array databases running five different well-known static thread pinning strategies. Our experiments showed a maximum speedup of these different strategies by 2.49x for SAVIME and up to 1.40x for SciDB. We also observed that these static strategies only yield 48% from the potential speedup (and 26% of the energy reduction), opening a new research topic.
Simone Dominico, Eduardo C. de Almeida, Marco A. Z. Alves, Jorge Augusto Meira
PDP3
2020 Freezing time emulating new and faster devices with virtual machines
Luis C. E. Bona, Alessandro Elias, Andre P. Ziviani, Ramon Nou, Toni Cortes, Marco A. Z. Alves
CCF Trans. High Perform. Comput.6
2019 A Compiler for Automatic Selection of Suitable Processing-in-Memory Instructions
abstract
Although not a new technique, due to the advent of 3D-stacked technologies, the integration of large memories and logic circuitry able to compute large amount of data has revived the Processing-in-Memory (PIM) techniques. PIM is a technique to increase performance while reducing energy consumption when dealing with large amounts of data. Despite several designs of PIM are available in the literature, their effective implementation still burdens the programmer. Also, various PIM instances are required to take advantage of the internal 3D-stacked memories, which further increases the challenges faced by the programmers. In this way, this work presents the Processing-In-Memory cOmpiler (PRIMO). Our compiler is able to efficiently exploit large vector units on a PIM architecture, directly from the original code. PRIMO is able to automatically select suitable PIM operations, allowing its automatic offloading. Moreover, PRIMO concerns about several PIM instances, selecting the most suitable instance while reduces internal communication between different PIM units. The compilation results of different benchmarks depict how PRIMO is able to exploit large vectors, while achieving a near-optimal performance when compared to the ideal execution for the case study PIM. PRIMO allows a speedup of 38× for specific kernels, while on average achieves 11.8 × for a set of benchmarks from PolyBench Suite.
Hameeza Ahmed, Paulo C. Santos 0001, João Paulo C. de Lima, Rafael Fao de Moura, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro
DATE5
2019 Database Processing-in-Memory: A Vision
Tiago Rodrigo Kepe, Eduardo C. de Almeida, Marco A. Z. Alves, Jorge Augusto Meira
DEXA (1)3
2019 Multi-phased Task Placement of HPC Applications in the Cloud
abstract
Many high-performance computing applications present different phases during their execution. Nevertheless, thread and process placement techniques usually provide static-only methods to improve the data and thread locality. Similarly, cloud computing datacenters may present variations in terms of latency over the execution time of applications. To overcome these two problems, in this paper we analyze scientific applications that have different communication patterns along with its execution. For such applications, we evaluate the performance variation of traditional static placement techniques to our new approach that uses code annotations to perform the new placement of tasks, matching also the variations on network performance of Virtual Machines (VMs) during the run time. For our experiments, we use applications from the NAS parallel benchmark suite, running them on two VM sizes with 32 and 64 cores respectively, from the same family of instance types at the West US datacenter from Azure. Results show that compared to traditional static process mapping, our multi-phased placement mechanism achieves average performance gains of 13.57% up to 28.32% on the evaluated scenarios. These results show that there is an opportunity to improve performance by correctly identifying the network variations and reacting by generating a new task-to-instance mapping.
Emmanuell D. Carreño, Marco A. Z. Alves, Matthias Diener, Eduardo Roloff, Philippe Olivier Alexandre Navaux
ISPDC2
2019 Database Processing-in-Memory: An Experimental Study
abstract
The rapid growth of "big-data" intensified the problem of data movement when processing data analytics: Large amounts of data need to move through the memory up to the CPU before any computation takes place. To tackle this costly problem, Processing-in-Memory (PIM) inverts the traditional data processing by pushing computation to memory with an impact on performance and energy efficiency. In this paper, we present an experimental study on processing database SIMD operators in PIM compared to current x86 processor (i.e., using AVX512 instructions). We discuss the execution time gap between those architectures. However, this is the first experimental study, in the database community, to discuss the trade-offs of execution time and energy consumption between PIM and x86 in the main query execution systems: materialized, vectorized, and pipelined. We also discuss the results of a hybrid query scheduling when interleaving the execution of the SIMD operators between PIM and x86 processing hardware. In our results, the hybrid query plan reduced the execution time by 45%. It also drastically reduced energy consumption by more than 2x compared to hardware-specific query plans.
Tiago Rodrigo Kepe, Eduardo C. de Almeida, Marco A. Z. Alves
Proc. VLDB Endow.3
2018 Design space exploration for PIM architectures in 3D-stacked memories
abstract
Scaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3D-stacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-of-the-art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High-Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth.
João Paulo C. de Lima, Paulo C. Santos 0001, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro
CF3
2018 Processing in 3D memories to speed up operations on complex data structures
abstract
Pointer chasing has been, for years, the kernel operation employed by diverse data structures, from graphs to hash tables and dictionaries. However, due to the bewildering growth in the volume of data that current applications have to deal with, performing pointer chasing operations have become a major source of performance and energy bottleneck, due to its sparse memory access behavior. In this work, we aim to tackle this problem by taking advantage of the already available parallelism present in today's 3D-stacked memories. We present a simple mechanism that can accelerate pointer chasing operations by making use of a state-of-the-art PIM design that executes in-memory vector operations. The key idea behind our design is to run speculative loads, in parallel, based on a given memory address in a reconfigurable window of addresses. Our design can perform pointer-chasing operations on b+tree 4.9 χ faster when compared to modern baseline systems. Besides that, since our device avoids data movement, we can also reduce energy consumption by 85% when compared to the baseline.
Paulo C. Santos 0001, Geraldo F. Oliveira, João Paulo C. de Lima, Marco A. Z. Alves, Luigi Carro, Antonio Carlos Schneider Beck
DATE4
2018 HIPE: HMC instruction predication extension applied on database processing
abstract
The recent Hybrid Memory Cube (HMC) is a smart memory which includes functional units inside one logic layer of the 3D stacked memory design. In order to execute instructions inside the Hybrid Memory Cube (HMC), the processor needs to send instructions to be executed near data, keeping most of the pipeline complexity inside the processor. Thus, control-flow and data-flow dependencies are all managed inside the processor, in such way that only update instructions are supported by the HMC. In order to solve data-flow dependencies inside the memory, previous work proposed HMC Instruction Vector Extensions (HIVE), which embeds a high number of functional units with a interlock register bank. In this work we propose HMC Instruction Prediction Extensions (HIPE), that supports predicated execution inside the memory, in order to transform control-flow dependencies into data-flow dependencies. Our mechanism focus on removing the high latency iteration between the processor and the smart memory during the execution of branches that depends on data processed inside the memory. In this paper we evaluate a balanced design of HIVE comparing to x86 and HMC executions. After we show the HIPE mechanism results when executing a database workload, which is a strong candidate to use smart memories. We show interesting trade-offs of performance when comparing our mechanism to previous work.
Diego G. Tomé, Paulo C. Santos 0001, Luigi Carro, Eduardo C. de Almeida, Marco A. Z. Alves
DATE5
2018 An Elastic Multi-Core Allocation Mechanism for Database Systems
abstract
peer reviewed
Simone Dominico, Eduardo C. de Almeida, Jorge Augusto Meira, Marco A. Z. Alves
ICDE4
2018 Freezing Time: A New Approach for Emulating Fast Storage Devices Using VM
abstract
Recently we are seeing a considerable effort from both academy and industry in proposing new technologies for storage devices. Often these devices are not readily available for evaluation and methods to allow performing their tests just from their performance parameters are an important tool for system administrators. Simulators are a traditional approach for carrying out such evaluations, however, they are more suitable for evaluating the storage device as an isolate component, mostly due to time constraints. In this paper, we propose an approach based on virtual machine technology that is capable of emulate storage devices transparently for the operating system allowing evaluation of simulating devices within a real system using any synthetic or real workload. To emulate devices in real environments it is necessary to use the currently available devices as a storage medium which creates a difficulty when the device to be emulated is faster than this storage medium. To circumvent this limitation we introduce a new technique called Freezing Time, which takes advantage of virtual machine pausing mechanism to manipulate the virtual machine clock and hide the real I/O completion time. Our approach can be implemented just requiring the hypervisor to be modified, providing a high degree of compatibility and flexibility since it is not necessary to modify neither the operating system nor the application. We evaluate our tool under a real system using old magnetic disks to emulate faster storage devices. Experiments using our technique presented an average latency error of 6.08% for read operations and 6.78% for write operations when comparing a real to device.
Luis C. E. Bona, Alessandro Elias, Andre P. Ziviani, Toni Cortes, Ramon Nou, Marco A. Z. Alves
MASCOTS6
2017 Operand size reconfiguration for big data processing in memory
abstract
Nowadays, applications that predominantly perform lookups over large databases are becoming more popular with column-stores as the database system architecture of choice. For these applications, Hybrid Memory Cubes (HMCs) can provide bandwidth of up to 320 GB/s and represents the best choice to keep the throughput for these ever increasing databases. However, even with the high available memory bandwidth and processing power, in order to achieve the peak performance, data movements through the memory hierarchy consumes an unnecessary amount of time and energy. In order to accelerate database operations, and reduce the energy consumption of the system, this paper presents the Reconfigurable Vector Unit (RVU) that enables massive and adaptive in-memory processing, extending the native HMC instructions and also increasing its effectiveness. RVU enables the programmer to reconfigure it to perform as a large vector unit or multiple small vectors units to better adjust for the application needs during different computation phases. Due to its adaptability, RVU is capable of achieving performance increase of 27 χ on average and reduce the DRAM energy consumption in 29% when compared to an x86 processor with 16 cores. Compared with the state-of-the-art mechanism capable of performing large vector operations with fixed size, inside the HMC, RVU performed up to 12% better in terms of performance and improve in 53% the energy consumption.
Paulo C. Santos 0001, Geraldo F. Oliveira, Diego G. Tomé, Marco A. Z. Alves, Eduardo C. de Almeida, Luigi Carro
DATE4
2016 Large vector extensions inside the HMC
Marco A. Z. Alves, Matthias Diener, Paulo C. Santos 0001, Luigi Carro
DATE1
2016 Communication in Shared Memory: Concepts, Definitions, and Efficient Detection
abstract
Optimizing the communication behavior of parallel applications has emerged as an important topic in parallel processing. In shared memory architectures, threads communicate implicitly through memory accesses to shared memory areas. The communication behavior can be improved by mapping threads that communicate a lot to processing units that are close to each other in the memory hierarchy, such that they can benefit from shared caches and faster interconnections. An important aspect of such a communication-aware thread mapping is the accurate and efficient detection of communication in shared memory. Previous work used impromptu definitions, without an evaluation of the complexities of different communication types. In this paper, we perform an in-depth, systematic evaluation of communication in shared memory, focusing on its architectural effects. We present an efficient way to detect communication, which is orders of magnitude faster than a cache simulator, while maintaining a high accuracy.
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
PDP3
2016 Exploring Cache Size and Core Count Tradeoffs in Systems with Reduced Memory Access Latency
abstract
One of the main challenges for computer architects is how to hide the high average memory access latency from the processor. In this context, Hybrid Memory Cubes (HMCs) can provide substantial energy and bandwidth improvements compared to traditional memory organizations. However, it is not clear how this reduced average memory access latency will impact the LLC. For applications with high cache miss ratios, the latency to search for the data inside the cache memory will impact negatively on the performance. The importance of this overhead depends on the memory access latency. In this paper, we present an evaluation of the L3 cache importance on a high performance processor using HMC also exploring chip area tradeoffs between the cache size and number of processor cores. We show that the high bandwidth provided by HMC memories can eliminate the need for L3 caches, removing hardware and making room for more processing power. Our evaluations show that performance increased 37% and the EDP improved 12% while maintaining the same original chip area in a wide range of parallel applications, when compared to DDR3 memories.
Paulo C. Santos 0001, Marco A. Z. Alves, Matthias Diener, Luigi Carro, Philippe Olivier Alexandre Navaux
PDP2
2016 LAPT: A locality-aware page table for thread and data mapping
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
Parallel Comput.3
2016 A dynamic block-level execution profiler
Francis B. Moreira 0001, Marco A. Z. Alves, Matthias Diener, Philippe Olivier Alexandre Navaux, Israel Koren
Parallel Comput.2
2016 Kernel-Based Thread and Data Mapping for Improved Memory Affinity
abstract
Reducing the cost of memory accesses, both in terms of performance and energy consumption, is a major challenge in shared-memory architectures. Modern systems have deep and complex memory hierarchies with multiple cache levels and memory controllers, leading to a Non-Uniform Memory Access (NUMA) behavior. In such systems, there are two ways to improve the memory affinity: First, by mapping threads that share data to cores with a shared cache, cache usage and communication performance are optimized. Second, by mapping memory pages to memory controllers that perform the most accesses to them and are not overloaded, the average cost of accesses is reduced. We call these two techniques thread mapping and data mapping, respectively. Thread and data mapping should be performed in an integrated way to achieve a compounding effect that results in higher improvements overall. Previous work in this area requires expensive tracing operations to perform the mapping, or require changes to the hardware or to the parallel application. In this paper, we propose kMAF, a mechanism that performs integrated thread and data mapping in the kernel. kMAF uses the page faults of parallel applications to characterize their memory access behavior and performs the mapping during the execution of the application based on the detected behavior. In an evaluation with a large set of parallel benchmarks executing on three NUMA architectures, kMAF achieved substantial performance and energy efficiency improvements, close to an Oracle-based mechanism and significantly higher than previous proposals.
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux, Anselm Busse, Hans-Ulrich Heiß
IEEE Trans. Parallel Distributed Syst.3
2015 Saving memory movements through vector processing in the DRAM
abstract
Despite the ability of modern processors to execute a variety of algorithms efficiently through instructions based on registers with ever-increasing widths, some applications present poor performance due to the limited interconnection bandwidth between main memory and processing units. Near-data processing has started to gain acceptance as an accelerator device due to the technology constraints and high costs associated with data transfer. However, previous approaches to near-data computing do not provide general-purpose processing, or require large amounts of logic and do not fully use the potential of the DRAM devices. These issues limited its wide adoption. In this paper, we present the Memory Vector Extensions (MVX), which implement vector instructions directly inside the DRAM devices, therefore avoiding data movement between memory and processing units, while requiring a lower amount of logic than previous approaches. MVX is able to obtain up to 211× increase in performance for application kernels with a high spatial locality and a low temporal locality. Comparing to an embedded processor with 8 cores and 2 memory channels that supports AVX-512 instructions, MVX performs 24× faster on average for three well known algorithms.
Marco A. Z. Alves, Paulo C. Santos 0001, Francis B. Moreira 0001, Matthias Diener, Luigi Carro
CASES1
2015 Locality and Balance for Communication-Aware Thread Mapping in Multicore Systems
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Mohammad Shadi Al Hakeem, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
Euro-Par3
2014 Optimizing Memory Locality Using a Locality-Aware Page Table
abstract
One of the main challenges for modern parallel shared-memory architectures are accesses to main memory. In current systems, the performance and energy efficiency of memory accesses depend on their locality: accesses to remote caches and NUMA nodes are more expensive than accesses to local ones. Increasing the locality requires knowledge about how the threads of a parallel application access memory pages. With this information, pages can be migrated to the NUMA nodes that access them (data mapping), as well as threads that access the same pages can be migrated to the same node such that locality can be improved even further (thread mapping). In this paper, we propose LAPT, a mechanism to store the memory access pattern of parallel applications in the page table, which is updated by the hardware during TLB misses. This information is used by the operating system to perform an optimized thread and data mapping during the execution of the parallel application. In contrast to previous work, LAPT does not require any previous information about the behavior of the applications, or changes to the application or runtime libraries. Extensive experiments with the NAS Parallel Benchmarks (NPB) and PARSEC showed performance and energy efficiency improvements of up to 19.2% and 15.7%, respectively, (6.7% and 5.3% on average).
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2014 Profiling and Reducing Micro-Architecture Bottlenecks at the Hardware Level
abstract
Most mechanisms in current superscalar processors use instruction granularity information for speculation, such as branch predictors or prefetchers. However, many of these characteristics can be obtained at the basic block level, increasing the amount of code that can be covered while requiring less space to store the data. Furthermore, the code can be profiled more accurately and provide a higher variety of information by analyzing different instruction types inside a block. Because of these advantages, block-level analysis can offer more opportunities for mechanisms that use this information. For example, it is possible to integrate information about branch prediction and memory accesses to provide precise information for speculative mechanisms, increasing accuracy and performance. We propose a Block-Level Architecture Profiler (BLAP), an online mechanism that profiles bottlenecks at the micro architectural level, such as delinquent memory loads, hard-to-predict branches and contention for functional units. BLAP works at the basic block level, providing information that can be used to reduce the impact of these bottlenecks. A prefetch dropping mechanism and a memory controller policy were developed to use the profiled information provided by BLAP. Together, these mechanisms are able to improve performance by up to 17.39% (3.90% on average). Our technique showed average gains of 13.14% when evaluated under high memory pressure due to highly aggressive prefetch.
Francis B. Moreira 0001, Marco A. Z. Alves, Israel Koren
SBAC-PAD2
2014 Dynamic thread mapping of shared memory applications by exploiting cache coherence protocols
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.3
2013 Energy Efficient Last Level Caches via Last Read/Write Prediction
abstract
The size of the Last Level Caches (LLC) in multi-core architectures is increasing, and so is their power consumption. However, most of this power is wasted on unused or invalid cache lines. For dirty cache lines, the LLC waits until the line is evicted to be written back to memory. Hence, dirty lines compete for the memory bandwidth with read requests (prefetch and demand), increasing pressure on the memory controller. This paper proposes a Dead Line and Early Write-Back Predictor (DEWP) to improve the energy efficiency of the LLC. DEWP early evicts dead cache lines with an average accuracy of 94%, and only 2% false positives. DEWP also allows scheduling of dirty lines for early eviction, allowing earlier write-backs. Using DEWP over a set of single and multi-threaded benchmarks, we obtain an average of 61% static energy savings, while maintaining the performance, for both inclusive and non-inclusive LLCs.
Marco A. Z. Alves, Carlos Villavieja, Matthias Diener, Philippe Olivier Alexandre Navaux
SBAC-PAD1
2012 Energy Savings via Dead Sub-Block Prediction
abstract
Cache memories have traditionally been designed to exploit spatial locality by fetching entire cache lines from memory upon a miss. However, recent studies have shown that often the number of sub-blocks within a line that are actually used is low. Furthermore, those sub-blocks that are used are accessed only a few times before becoming dead (i.e., never accessed again). This results in considerable energy waste since 1) data not needed by the processor is brought into the cache, and 2) data is kept alive in the cache longer than necessary. We propose the Dead Sub-Block Predictor (DSBP) to predict which sub-blocks of a cache line will be actually used and how many times it will be used in order to bring into the cache only those sub-blocks that are necessary, and power them off after they are touched the predicted number of times. We also use DSBP to identify dead lines (i.e., all sub-blocks off) and augment the existing replacement policy by prioritizing dead lines for eviction. Our results show a 24% energy reduction for the whole cache hierarchy when averaged over the SPEC2000, SPEC2006 and NAS-NPB benchmarks.
Marco A. Z. Alves, Khubaib, Eiman Ebrahimi, Veynu Narasiman, Carlos Villavieja, Philippe Olivier Alexandre Navaux, Yale N. Patt
SBAC-PAD1
2010 Evaluating Thread Placement Based on Memory Access Patterns for Multi-core Processors
abstract
Process placement is a technique widely used on parallel machines with heterogeneous interconnects to reduce the overall communication time. For instance, two processes which communicate frequently are mapped close to each other. Finding the optimal mapping between threads and cores in a shared-memory environment (for example, OpenMP and Pthreads) is an even more complex task due to implicit communication. In this work, we examine data sharing patterns between threads in different workloads and use those patterns in a similar way as messages are used to map processes in cluster computers. We evaluated our technique on a state-of-the-art multicore processor and achieved moderate improvements in the common case and considerable improvements in some cases, reducing execution time by up to 45%.
Matthias Diener, Felipe Lopes Madruga, Eduardo Rocha Rodrigues, Marco A. Z. Alves, Jörg Schneider 0001, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
HPCC4
2010 Impact of Parallel Workloads on NoC Architecture Design
abstract
Due to the multi-core processors, the importance of parallel workloads has increased considerably. However, many-core chips demand new interconnection strategies, since traditional crossbars or buses, common for current multi-core processors, have problems related to wires and scalability. For this reason, Networks-on-Chip (NoCs) have been developed in order to support the performance and parallelism focused on several workloads. Although a Network-on-Chip is a good option, most designs consist of a large number of routers. These routers are responsible for forwarding packets, and consequently, for supporting message-passing workloads. In this context, the NoC performance is a problem. Therefore, the main goal of this paper is to evaluate the impact of well-known parallel workloads on NoC architecture design. In order to achieve high performance, the results point out to parallel workloads with small packets and cluster-based NoCs with circuit switching and adaptable topologies.
Henrique Cota de Freitas, Lucas Mello Schnorr, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
PDP3
2009 Design of Interleaved Multithreading for Network Processors on Chip
abstract
Thread level parallelism and multi-core processors are current alternatives to increase performance of general-purpose applications. In the same way, networks-on-ohip (NoCs) are the main alternatives for supporting packet throughput for the next generations of many-core processors. NPoC (network processor on chip) is a proposal to increase the performance of programmable NoC routers and multi-cluster NoC architectures using interleaved multithreading (IMT) technique. Therefore, the main goal of this paper is to present the design impact of interleaved multithreading for network processors on chip focusing on area and performance feasibility. Results show that NPoC-based router has an acceptable and similar area relative to a conventional NoC, and higher performance up to 7.1% than the same NPoC version without IMT.
Henrique Cota de Freitas, Felipe Lopes Madruga, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
ISCAS3
2009 Performance Evaluation of NoC Architectures for Parallel Workloads
abstract
Network-on-Chip is the state-of-the-art approach to interconnect many processing cores in the next generation of general-purpose processors. In this context, the problem is to choose NoC architectures capable of achieving high performance for parallel programs. Therefore, the main goal of this paper is to evaluate the performance of three NoC architectures using well-known parallel workloads.
Henrique Cota de Freitas, Marco A. Z. Alves, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
NOCS2