Erik Vermij

dblp:146/2834 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0003-0639-412XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 53% Performance modeling and evaluation · 11% Processor architecture and microarchitecture · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
processing-in-memory
0.622018
Near-Memory Acceleration for Radio Astronomy · IEEE Trans. Parallel Distributed Syst. 2018
An Architecture for Integrated Near-Data Processors · ACM Trans. Archit. Code Optim. 2017
Electronic design automation
design space exploration
0.312018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Processor architecture and microarchitecture
multicore design
0.312018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Memory systems › processing-in-memory
near-memory processing
0.312018
Near-Memory Acceleration for Radio Astronomy · IEEE Trans. Parallel Distributed Syst. 2018
Performance modeling and evaluation
processor performance modeling
0.312018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Hardware accelerators and domain-specific architectures
scientific computing accelerator
0.312018
Near-Memory Acceleration for Radio Astronomy · IEEE Trans. Parallel Distributed Syst. 2018
Memory systems
cache coherence
0.312017
An Architecture for Integrated Near-Data Processors · ACM Trans. Archit. Code Optim. 2017
Memory systems › processing-in-memory
near-data processing
0.312017
An Architecture for Integrated Near-Data Processors · ACM Trans. Archit. Code Optim. 2017
Computational science and engineering › astronomy
radio astronomy
0.112018
Near-Memory Acceleration for Radio Astronomy · IEEE Trans. Parallel Distributed Syst. 2018
High-performance computing › supercomputing
exascale systems
0.112018
Analytic Multi-Core Processor Model for Fast Design-Space Exploration · IEEE Trans. Computers 2018
Memory systems › memory management
virtual memory
0.112017
An Architecture for Integrated Near-Data Processors · ACM Trans. Archit. Code Optim. 2017

Methods — techniques the papers use, named apart from their topics

3d-stacked memory · 0.7hardware performance counters · 0.3energy-efficiency optimization · 0.3energy efficiency optimization · 0.3analytic modeling · 0.3system simulation · 0.3
YearPublicationVenuePosition
2018 Analytic Multi-Core Processor Model for Fast Design-Space Exploration
abstract
Simulators help computer architects optimize system designs. The limited performance of simulators even of moderate size and detail makes the approach infeasible for design-space exploration of future exascale systems. Analytic models, in contrast, offer very fast turn-around times. In this paper we propose an analytic multi-core processor-performance model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy. The processor-performance model considers instruction-level parallelism (ILP) per type, models single instruction, multiple data (SIMD) features, and considers cache and memory-bandwidth contention between cores. We validate our model by comparing its performance estimates with measurements from hardware performance counters on Intel Xeon and ARM Cortex-A15 systems. We estimate multi-core contention with a maximum error of 11.4 percent. The average single-thread error increases from 25 percent for a state-of-the-art simulator to 59 percent for our model, but the correlation is still 0.8, a high relative accuracy, while we achieve a speedup of several orders of magnitude. With a much higher capacity than simulators and more reliable insights than back-of-the-envelope calculations it makes automated design-space exploration of exascale systems possible, which we show using a real-world case study from radio astronomy.
Rik Jongerius, Andreea Anghel, Gero Dittmann, Giovanni Mariani, Erik Vermij, Henk Corporaal
IEEE Trans. Computers5
2018 Near-Memory Acceleration for Radio Astronomy
abstract
Processing-in-memory and near-memory computing have recently been rediscovered as a way to alleviate the “memory wall problem” of traditional computing architectures. In this paper, we discuss the implementation of a 3D-stacked near-memory accelerator, targeting radio astronomy and scientific applications. After exploring the design space of the architecture by focusing on minimizing the execution power of the processing pipeline of the SKA1-Low central signal processor, we show that our accelerator can achieve an energy efficiency of up to 390 GFLOPS/W, corresponding to an energy consumption one order of magnitude lower than alternative state-of-the-art implementations. When running additional mathematical and streaming-oriented kernels, our accelerator achieves from 6.4x to 20x energy efficiency improvement compared to alternative solutions.
Leandro Fiorin, Rik Jongerius, Erik Vermij, Jan van Lunteren, Christoph Hagleitner
IEEE Trans. Parallel Distributed Syst.3
2017 Boosting the Efficiency of HPCG and Graph500 with Near-Data Processing
abstract
HPCG and Graph500 can be regarded as the two most relevant benchmarks for high-performance computing systems. Existing supercomputer designs, however, tend to focus on floating-point peak performance, a metric less relevant for these two benchmarks, leaving resources underutilized, and resulting in little performance improvements, for these benchmarks, over time. In this work, we analyze the implementation of both benchmarks on a novel shared-memory near-data processing architecture. We study a number of aspects: 1. a system parameter design exploration, 2. software optimizations, and 3. the exploitation of unique architectural features like user-enhanced coherence as well as the exploitation of data-locality for inter near-data processor traffic.For the HPCG benchmark, we show a factor 2.5x application level speedup with respect to a CPU, and a factor 2.5x power-efficiency improvement with respect to a GPU. For the Graph500 benchmark, we show up to a factor 3.5x speedup with respect to a CPU. Furthermore, we show that, with many of the existing data-locality optimizations for this specific graph workload applied, local memory bandwidth is not the crucial parameter, and a high-bandwidth as well as low-latency interconnect are arguably more important, shining a new light on the near-data processing characteristics most relevant for this type of heavily optimized graph processing.
Erik Vermij, Leandro Fiorin, Christoph Hagleitner, Koen Bertels
ICPP1
2017 An Architecture for Integrated Near-Data Processors
abstract
To increase the performance of data-intensive applications, we present an extension to a CPU architecture that enables arbitrary near-data processing capabilities close to the main memory. This is realized by introducing a component attached to the CPU system-bus and a component at the memory side. Together they support hardware-managed coherence and virtual memory support to integrate the near-data processors in a shared-memory environment. We present an implementation of the components, as well as a system-simulator, providing detailed performance estimations. With a variety of synthetic workloads we demonstrate the performance of the memory accesses, the mixed fine- and coarse-grained coherence mechanisms, and the near-data processor communication mechanism. Furthermore, we quantify the inevitable start-up penalty regarding coherence and data writeback, and argue that near-data processing workloads should access data several times to offset this penalty. A case study based on the Graph500 benchmark confirms the small overhead for the proposed coherence mechanisms and shows the ability to outperform a real CPU by a factor of two.
Erik Vermij, Leandro Fiorin, Rik Jongerius, Christoph Hagleitner, Jan van Lunteren, Koen Bertels
ACM Trans. Archit. Code Optim.1
2015 Analytic processor model for fast design-space exploration
abstract
In this paper, we propose an analytic model that takes as inputs a) a parametric microarchitecture-independent characterization of the target workload, and b) a hardware configuration of the core and the memory hierarchy, and returns as output an estimation of processor-core performance. To validate our technique, we compare our performance estimates with measurements on an Intel® Xeon® system. The average error increases from 21% for a state-of-the-art simulator to 25% for our model, but we achieve a speedup of several orders of magnitude. Thus, the model enables fast designspace exploration and represents a first step towards an analytic exascale system model.
Rik Jongerius, Giovanni Mariani, Andreea Anghel, Gero Dittmann, Erik Vermij, Henk Corporaal
ICCD5