Matheus Ogleari

dblp:207/5375 · also Matheus A. Ogleari · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Memory systems · 55% Storage systems · 20% Interconnection networks and networks-on-chip · 13%
Artificial intelligence
1 paper

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › non-volatile memory
persistent memory
0.722018
Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network · MICRO 2018
Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory Systems · HPCA 2018
Memory systems › memory disaggregation
memory network
0.412019
String Figure: A Scalable and Elastic Memory Network Architecture · HPCA 2019
Interconnection networks and networks-on-chip
network topology
0.412019
String Figure: A Scalable and Elastic Memory Network Architecture · HPCA 2019
Storage systems › storage reliability
durability
0.312018
Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory Systems · HPCA 2018
Memory systems
hardware logging
0.312018
Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory Systems · HPCA 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › training accelerator
neural network training accelerator
0.312018
Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach · MICRO 2018
Memory systems
non-volatile memory
0.312018
Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network · MICRO 2018
Memory systems
processing-in-memory
0.312018
Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach · MICRO 2018
Interconnection networks and networks-on-chip
remote direct memory access
0.112018
Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network · MICRO 2018
Storage systems
storage reliability
0.112018
Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory Systems · HPCA 2018

Methods — techniques the papers use, named apart from their topics

software-hardware co-design · 0.7runtime scheduling · 0.7OpenCL · 0.7network reconfiguration · 0.4hybrid routing protocol · 0.4RTL simulation · 0.4undo+redo logging · 0.3cache force-write-back · 0.3buffered strict persistence · 0.3barrier epoch management · 0.3
YearPublicationVenuePosition
2019 String Figure: A Scalable and Elastic Memory Network Architecture
abstract
Demand for server memory capacity and performance is rapidly increasing due to expanding working set sizes of modern applications, such as big data analytics, inmemory computing, deep learning, and server virtualization. One promising techniques to tackle this requirements is memory networking, whereby a server memory system consists of multiple 3D die-stacked memory nodes interconnected by a high-speed network. However, current memory network designs face substantial scalability and flexibility challenges. This includes (1) maintaining high throughput and low latency in large-scale memory networks at low hardware cost, (2) efficiently interconnecting an arbitrary number of memory nodes, and (3) supporting flexible memory network scale expansion and reduction without major modification of the memory network design or physical implementation. To address the challenges, we propose String Figure1, a highthroughput, elastic, and scalable memory network architecture. String Figure consists of (1) an algorithm to generate random topologies that achieve high network throughput and nearoptimal path lengths in large-scale memory networks, (2) a hybrid routing protocol that employs a mix of computation and look up tables to reduce the overhead of both in routing, (3) a set of network reconfiguration mechanisms that allow both static and dynamic network expansion and reduction. Our experiments using RTL simulation demonstrate that String Figure can interconnect over one thousand memory nodes with a shortest path length within five hops across various traffic patterns and real workloads.
Matheus Ogleari, Ye Yu 0001, Chen Qian 0001, Ethan L. Miller, Jishen Zhao
HPCA1
2018 Steal but No Force: Efficient Hardware Undo+Redo Logging for Persistent Memory Systems
abstract
Persistent memory is a new tier of memory that functions as a hybrid of traditional storage systems and main memory. It combines the benefits of both: the data persistence of storage with the fast load/store interface of memory. Most previous persistent memory designs place careful control over the order of writes arriving at persistent memory. This can prevent caches and memory controllers from optimizing system performance through write coalescing and reordering. We identify that such write-order control can be relaxed by employing undo+redo logging for data in persistent memory systems. However, traditional software logging mechanisms are expensive to adopt in persistent memory due to performance and energy overheads. Previously proposed hardware logging schemes are inefficient and do not fully address the issues in software. To address these challenges, we propose a hardware undo+redo logging scheme which maintains data persistence by leveraging the write-back, write-allocate policies used in commodity caches. Furthermore, we develop a cache force-write-back mechanism in hardware to significantly reduce the performance and energy overheads from forcing data into persistent memory. Our evaluation across persistent memory microbenchmarks and real workloads demonstrates that our design significantly improves system throughput and reduces both dynamic energy and memory traffic. It also provides strong consistency guarantees compared to software approaches.
Matheus Ogleari, Ethan L. Miller, Jishen Zhao
HPCA1
2018 Persistence Parallelism Optimization: A Holistic Approach from Memory Bus to RDMA Network
abstract
Emerging non-volatile memories (NVM), such as phase change memory (PCM) and Resistive RAM (ReRAM), incorporate the features of fast byte-addressability and data persistence, which are beneficial for data services such as file systems and databases. To support data persistence, a persistent memory system requires ordering for write requests. The datapath of a persistent request consists of three segments: through the cache hierarchy to the memory controller, through the bus from the memory controller to memory devices, and through the network from a remote node to a local node. Previous work contributes significantly to improve the persistence parallelism in the first segment of the data path. However, we observe that the memory bus and the Remote Direct Memory Access (RDMA) network remain severely under-utilized because the persistence parallelism in these two segments is not fully leveraged during ordering. In this paper, we propose a novel architecture to further improve the persistence parallelism in the memory bus and the RDMA network. First, we utilize inter-thread persistence parallelism for barrier epoch management with better bank-level parallelism (BLP). Second, we enable intra-thread persistence parallelism for remote requests through RDMA network with buffered strict persistence. With these features, the architecture efficiently supports persistence through all three segments of the write datapath. Experimental results show that for local applications, the proposed mechanism can achieve 1.3× performance improvement, compared to the original buffered persistence work. In addition, it can achieve 1.93× performance improvement for remote applications serviced through the RDMA network.
Xing Hu 0001, Matheus Ogleari, Jishen Zhao, Shuangchen Li, Abanti Basak, Yuan Xie 0001
MICRO2
2018 Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach
abstract
Neural networks (NNs) have been adopted in a wide range of application domains, such as image classification, speech recognition, object detection, and computer vision. However, training NNs - especially deep neural networks (DNNs) - can be energy and time consuming, because of frequent data movement between processor and memory. Furthermore, training involves massive fine-grained operations with various computation and memory access characteristics. Exploiting high parallelism with such diverse operations is challenging. To address these challenges, we propose a software/hardware co-design of heterogeneous processing-in-memory (PIM) system. Our hardware design incorporates hundreds of fix-function arithmetic units and ARM-based programmable cores on the logic layer of a 3D die-stacked memory to form a heterogeneous PIM architecture attached to CPU. Our software design offers a programming model and a runtime system that program, offload, and schedule various NN training operations across compute resources provided by CPU and heterogeneous PIM. By extending the OpenCL programming model and employing a hardware heterogeneity-aware runtime system, we enable high program portability and easy program maintenance across various heterogeneous hardware, optimize system energy efficiency, and improve hardware utilization.
Hengyu Zhao, Matheus Ogleari, Dong Li 0001, Jishen Zhao
MICRO3