EDBT 2026 Demo / reviewers in the wild / expert
Erik Brockmeyer
dblp:00/1208
· DBLP profile ↗
18ranked-venue papers
4as first author
0since 2021 · last 2009
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 24% Processor architecture and microarchitecture · 21% Energy-efficient computing · 17% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 50% Image and video processing · 50% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2008 | Data-Reuse-Driven Energy-Aware Cosynthesis of Scratch Pad Memory and Hierarchical Bus-Based Communication Architecture for Multiprocessor Streaming Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchies · DAC 2006 |
Distributed systems › distributed scheduling
data transfer scheduling |
0.1 | 1 | 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case study · DAC 2008 |
Memory systems › memory access optimization
memory access reduction |
0.1 | 1 | 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case study · DAC 2008 |
Energy-efficient computing › low-power design
power optimization |
0.1 | 1 | 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case study · DAC 2008 |
Embedded and real-time systems › embedded software › embedded operating systems › embedded memory management
scratchpad memory management |
0.1 | 1 | 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case study · DAC 2008 |
Memory systems
memory hierarchy |
0.1 | 1 | 2006 | Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchies · DAC 2006 |
Electronic design automation
design space exploration |
0.0 | 2 | 2008 | Data-Reuse-Driven Energy-Aware Cosynthesis of Scratch Pad Memory and Hierarchical Bus-Based Communication Architecture for Multiprocessor Streaming Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 Global Multimedia System Design Exploration Using Accurate Memory Organization Feedback · DAC 1999 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 2008 | Data-Reuse-Driven Energy-Aware Cosynthesis of Scratch Pad Memory and Hierarchical Bus-Based Communication Architecture for Multiprocessor Streaming Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008 |
Integrated circuit design
system-on-chip |
0.0 | 1 | 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case study · DAC 2008 |
Electronic design automation
system-level design |
0.0 | 1 | 1999 | Global Multimedia System Design Exploration Using Accurate Memory Organization Feedback · DAC 1999 |
Energy-efficient computing › memory energy efficiency
low-power memory design |
0.0 | 1 | 2006 | Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchies · DAC 2006 |
Energy-efficient computing
power management |
0.0 | 1 | 2006 | Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchies · DAC 2006 |
Image and video processing
motion estimation |
0.0 | 1 | 1999 | Low Power Memory Storage and Transfer Organization for the MPEG-4 Full Pel Motion Estimation on a Multimedia Processor · IEEE Trans. Multim. 1999 |
Image and video coding › video coding standards
MPEG-4 video coding |
0.0 | 1 | 1999 | Low Power Memory Storage and Transfer Organization for the MPEG-4 Full Pel Motion Estimation on a Multimedia Processor · IEEE Trans. Multim. 1999 |
Methods — techniques the papers use, named apart from their topics
prefetching · 0.1mixed integer linear programming · 0.1heuristic synthesis · 0.1application analysis and transformation · 0.1data reuse analysis · 0.1source code transformation · 0.0ACROPOLIS methodology · 0.0memory organization estimation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2009 | Exploring parallelizations of applications for MPSoC platforms using MPAabstractThis paper presents a tool for exploring different parallelization options for an application. It can be used to quickly find a high-quality match between an application and a multi-processor platform architecture. By specifying the parallelization at a high abstraction level, and leaving the actual source code transformations to the tool, a designer can try out many parallelizations in a short time. A parallelization may use either functional or data-level splits, or a combination of both. An accompanying high-level simulator provides rapid feedback about the expected performance of a parallelization, based on platform parameters and profiling data of the sequential application on the target processor. The use of the tool and simulator are demonstrated on an MPEG-4 video encoder application and two different platform architectures. Rogier Baert, Erik Brockmeyer, Sven Wuytack, Thomas J. Ashby |
DATE | 2 |
| 2008 | An automatic scratch pad memory management tool and MPEG-4 encoder case studyabstractUsing software-controlled Scratch-Pad Memory (SPM) in Systems-on-Chip has the potential of reducing power consumption by using design-time application knowledge to reduce memory accesses and processor stalls. This paper presents a fully automatic application analysis and transformation tool which selects data-structures for transfer to the SPM and schedules data transfers between background memory and SPM (pre-fetching) to achieve both high performance and low power consumption. A case study applying this tool on an MPEG-4 video encoder shows an overall power reduction of 25%, a 40% power reduction in just the memories and a 40% reduction in processor cycles as compared to an optimized hardware-cache based solution. Rogier Baert, Eddy de Greef, Erik Brockmeyer |
DAC | 3 |
| 2008 | Data-Reuse-Driven Energy-Aware Cosynthesis of Scratch Pad Memory and Hierarchical Bus-Based Communication Architecture for Multiprocessor Streaming ApplicationsabstractAs technology advances, it becomes feasible to implement a large multiprocessor systems-on-chip (MPSoCs) to satisfy the increased performance demands of embedded applications. The increased complexity of systems leads to an increased power consumption. Reducing the consumption is an important task, considering that the available power may be limited in battery-operated embedded systems. The selection of memory and communication architectures affects the power efficiency of the design. In this paper, we propose a novel approach that enables the energy-aware cosynthesis of both memory and communication architectures for streaming applications. As opposed to earlier techniques, we propose a powerful compile-time analysis of memory access behavior in multiprocessor systems, which adds flexibility in selecting scratch-pad-based memory architectures. We propose and compare three memory/communication synthesis techniques, namely, an optimal mixed integer-linear-programming (ILP)-based cosynthesis technique, a mixed ILP (MILP)-based traditional two-step synthesis approach, where memory and communication synthesis is sequentially performed, and a cosynthesis heuristic that synthesizes energy-efficient hierarchical bus-based communication architectures with guaranteed throughput. Our experimental results on a number of streaming applications show that both the traditional two-step synthesis approach and heuristic result in up to 50% worse power consumption in comparison with our proposed cosynthesis approach. However, on some of the streaming benchmarks, our cosynthesis heuristic approach was able to find optimal or near-optimal results in a much shorter time than the MILP cosynthesis approach. Ilya Issenin, Erik Brockmeyer, Bart Durinck, Nikil Dutt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | The Impact of Higher Communication Layers on NoC Supported MP-SoCsabstractMulti-processor systems-on-chip use networks-on-chip (NoC) as a communication backbone to tackle the communication between processors and multi-level memory hierarchies. Inter-processor communication has a high impact on the NoC traffic but, to this day, there have been few detailed studies. Based on a realistic case study, we present a contrastive comparison of cache-based versus scratch-pad managed inter-processor communication for (distributed shared-memory) multiprocessor systems-on-chip. The platforms we target use six DSP nodes and a shared L2 memory, interconnected by a packet-switched network-on-chip with differentiated services. The first version of the platform uses caches to perform inter-processor communication whereas the second one uses a novel type of distributed DMA to help performing scratch-pad management. With detailed simulation results we show that the scratchpad application mapping has the best overall performance, that it helps smoothing NoC traffic and that it is not sensitive to the quality-of-service (QoS) used. We furthermore demonstrate that, on the contrary, cache-based MP-SoCs are very sensitive to the QoS level and that they generate significantly more NoC traffic than their scratch-pad counterpart. We recommend, where possible, to use scratch-pad management for NoC supported MP-SoCs as it yields performant, predictable results and can benefit from platform virtualization to achieve composability of applications Théodore Marescaux, Erik Brockmeyer, Henk Corporaal |
NOCS | 2 |
| 2007 | DRDU: A data reuse analysis technique for efficient scratch-pad memory managementabstractIn multimedia and other streaming applications, a significant portion of energy is spent on data transfers. Exploiting data reuse opportunities in the application, we can reduce this energy by making copies of frequently used data in a small local memory and replacing speed- and power-inefficient transfers from main off-chip memory by more efficient local data transfers. In this article we present an automated approach for analyzing these opportunities in a program that allows modification of the program to use custom scratch-pad memory configurations comprising a hierarchical set of buffers for local storage of frequently reused data. Using our approach we are able to both reduce energy consumption of the memory subsystem when using a scratch-pad memory by about a factor of two, on average, and improve memory system performance compared to a cache of the same size. Ilya Issenin, Erik Brockmeyer, Miguel Corbalan, Nikil Dutt |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Hierarchical memory size estimation for loop fusion and loop shifting in data-dominated applicationsabstractLoop fusion and loop shifting are important transformations for improving data locality to reduce the number of costly accesses to off-chip memories. Since exploring the exact platform mapping for all the loop transformation alternatives is a time consuming process, heuristics steered by improved data locality are generally used. However, pure locality estimates do not sufficiently take into account the hierarchy of the memory platform. This paper presents a fast, incremental technique for hierarchical memory size requirement estimation for loop fusion and loop shifting at the early loop transformations design stage. As the exact memory platform is often not yet defined at this stage, we propose a platform-independent approach which reports the Pareto-optimal trade-off points for scratch-pad memory size and off-chip memory accesses. The estimation comes very close to the actual platform mapping. Experiments on realistic test-vehicles confirm that. It helps the designer or a tool to find the interesting loop transformations that should then be investigated in more depth afterward. Qubo Hu, Arnout Vandecappelle, Martin Palkovic, Per Gunnar Kjeldsberg, Erik Brockmeyer, Francky Catthoor |
ASP-DAC | 5 |
| 2006 | Multiprocessor system-on-chip data reuse analysis for exploring customized memory hierarchiesabstractThe increasing use of multiprocessor systems-on-chip (MPSoCs) for high performance demands of embedded applications results in high power dissipation. The memory subsystem is a large and critical contributor to both energy and performance, requiring system designers to perform exploration of low power memory organizations. In this paper we present a novel multiprocessor data reuse analysis technique that allows the system designer to explore a wide range of customized memory hierarchy organizations with different size and energy profiles. Our technique enables the system designer to explore feasible memory subsystem solutions that meet power and area constraints while maintaining the necessary performance level. Our experiments on the complex QSDPCM benchmark illustrate the exploration of a wide range of customized memory hierarchies for an MPSoC implementation. Ilya Issenin, Erik Brockmeyer, Bart Durinck, Nikil Dutt |
DAC | 2 |
| 2006 | A combined DMA and application-specific prefetching approach for tackling the memory latency bottleneckabstractMemory latency has always been a major issue in embedded systems that execute memory-intensive applications. This is even more true as the gap between processor and memory speed continues to grow. Hardware and software prefetching have been shown to be effective in tolerating the large memory latencies inherit in large off-chip memories; however, both types of prefetching have their shortcomings. Hardware schemes are more complex and require extra circuitry to compute data access strides, while software schemes generate prefetch instructions, which if not computed carefully may hamper performance. On the other hand, some applications domains (such as multimedia) have a uniform and known a priori memory access pattern, that if exploited, could yield significant application performance improvement. With this characteristic in mind, we present our findings on hiding memory latency using the direct memory access (DMA) mode, which is present in all modern systems, combined with a software prefetch mechanism, and a customized on-chip memory hierarchy mapping. Compared to previous approaches, we are able to estimate the performance and power metrics, without actually implementing the embedded system. Experimental results on nine well known multimedia and imaging applications prove the efficiency of our technique. Finally, we verify the performance estimations by implementing and simulating the algorithms on the TI C6201 processor. Minas Dasygenis, Erik Brockmeyer, Bart Durinck, Francky Catthoor, Dimitrios Soudris, Adonios Thanailakis |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | A Memory Hierarchical Layer Assigning and Prefetching Technique to Overcome the Memory Performance/Energy BottleneckabstractThe memory subsystem has always been a bottleneck in performance as well as significant power contributor in memory intensive applications. Many researchers have presented multi-layered memory hierarchies as a means to design energy and performance efficient systems. However, most of the previous work does not explore trade-offs systematically. We fill this gap by proposing a formalized technique that takes into consideration data reuse, limited life-time of the arrays of an application and application specific prefetching opportunities, and performs a thorough tradeoff exploration for different memory layer sizes. This technique has been implemented on a prototype tool, which was tested successfully using nine real-life applications of industrial relevance. Following this approach we have able to reduce execution time up to 60%, and energy consumption up to 70%. Minas Dasygenis, Erik Brockmeyer, Bart Durinck, Francky Catthoor, Dimitrios Soudris, Adonios Thanailakis |
DATE | 2 |
| 2004 | Data Reuse Analysis Technique for Software-Controlled Memory HierarchiesabstractIn multimedia and other streaming applications a significant portion of energy is spent on data transfers. Exploiting data reuse opportunities in the application, we can reduce this energy by making copies of frequently used data in a small local memory and replacing speed and power inefficient transfers from main off-chip memory by more efficient local data transfers. In this paper we present an automated approach for analyzing these opportunities in a program that allows modification of the program to use custom scratch pad memory configurations comprising a hierarchical set of buffers for local storage of frequently reused data. Using our approach we are able to reduce energy consumption of the memory subsystem when using a scratch pad memory by a factor of two on average compare to a cache of the same size. Ilya Issenin, Erik Brockmeyer, Miguel Corbalan, Nikil Dutt |
DATE | 2 |
| 2003 | Background Data Organisation for the Low-Power Implementation in Real-Time of a Digital Audio Broadcast Receiver on a SIMD Processor
Pieter Op de Beeck, C. Ghez, Erik Brockmeyer, Miguel Corbalan, Francky Catthoor, Geert Deconinck |
DATE | 3 |
| 2003 | Layer Assignment echniques for Low Energy in Multi-Layered Memory Organisations
Erik Brockmeyer, Miguel Corbalan, Henk Corporaal, Francky Catthoor |
DATE | 1 |
| 2003 | Estimating influence of data layout optimizations on SDRAM energy consumptionabstractAn important problem in extracting maximum benefits from an SDRAM-based architecture is to exploit data locality at the page granularity. Frequent switches between data pages can increase memory latency and have an impact on energy consumption. In this paper, we propose a mathematical formulation, using Presburger arithmetic and Ehrhart polynomials to estimate the number of page breaks statically (i.e., at compile time). The results obtained using video codes indicate that the proposed framework can estimate the number of page breaks with good accuracy. Hyun Suk Kim, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Erik Brockmeyer, Francky Catthoor, Mary Jane Irwin |
ISLPED | 4 |
| 2001 | Data and memory optimization techniques for embedded systemsabstractWe present a survey of the state-of-the-art techniques used in performing data and memory-related optimizations in embedded systems. The optimizations are targeted directly or indirectly at the memory subsystem, and impact one or more out of three important cost metrics: area, performance, and power dissipation of the resulting implementation. We first examine architecture-independent optimizations in the form of code transoformations. We next cover a broad spectrum of optimization techniques that address memory architectures at varying levels of granularity, ranging from register files to on-chip memory, data caches, and dynamic memory (DRAM). We end with memory addressing related issues. Preeti Ranjan Panda, Francky Catthoor, Nikil Dutt, Koen Danckaert, Erik Brockmeyer, Chidamber Kulkarni, Arnout Vandecappelle, Per Gunnar Kjeldsberg |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2000 | Systematic cycle budget versus system power trade-off: a new perspective on system exploration of real-time data-dominated applicationsabstractIn contrast to current design practice for (programmable) processor mapping, which mainly targets performance, we focus on a systematic trade-off between cycle budget and energy consumed in the background memory organization. The latter is a crucial component in many of todays designs, including multi-media, network protocols and telecom signal processing. We have a systematic way and tool to explore both freedoms and to arrive at Pareto charts, in which for a given application the lowest cost implementation of the memory organization is plotted against the available cycle budget per submodule. This by making optimal usage of a parallelized memory architecture. We indicate, with results on a digital audio broadcasting receiver and an image compression demonstrator, how to effectively use the Pareto plot to gain significantly in overall system energy consumption within the global real-time constraints. Erik Brockmeyer, Arnout Vandecappelle, Francky Catthoor |
ISLPED | 1 |
| 1999 | Global Multimedia System Design Exploration Using Accurate Memory Organization FeedbackabstractSuccessful exploration of system-level design decisions is impossible without fast and accurate estimation of the impact on the system cost.In most multimedia applications, the dominant cost factor is related to the organization of the memory architecture.This paper presents a systematic approach which allows effective system-level exploration of memory organization design alternatives, based on accurate feedback by using our earlier developed tools.The effectiveness of this approach is illustrated on an industrial application.Applying our approach, a substantial part of the design search space has been explored in a very short time, resulting in a cost-efficient solution which meets all design constraints. Arnout Vandecappelle, Miguel Corbalan, Erik Brockmeyer, Francky Catthoor, Diederik Verkest |
DAC | 3 |
| 1999 | Low Power Memory Storage and Transfer Organization for the MPEG-4 Full Pel Motion Estimation on a Multimedia ProcessorabstractData transfers and storage are crucial cost factors in multimedia systems. Systematic methodologies are needed to obtain dramatic reductions in terms of power, area and cycle count. Upcoming multimedia processing applications will require high memory bandwidth. In this paper, we estimate that a software reference implementation of an MPEG-4 video encoder typically requires five Gtransfers/s to main memory for a simple profile level L2. This shows a clear need for optimization and the use of intermediate memory stages. By applying our ACROPOLIS methodology, developed mainly to relieve this data access bottleneck, we have arrived at an implementation which needs a factor 65 less background accesses. In addition, we also show that we can heavily improve on the memory transfers, without sacrificing speed (even gaining about 10% on cache misses and cycles for a DEC Alpha), by aggressive source code transformations. Erik Brockmeyer, Lode Nachtergaele, Francky Catthoor, Jan Bormans, Hugo De Man |
IEEE Trans. Multim. | 1 |
| 1998 | Code Transformations for Reduced Data Transfer and Storage in Low Power Realisations of MPEG-4 Full-Pel Motion Estimation
Erik Brockmeyer, Francky Catthoor, Jan Bormans, Hugo De Man |
ICIP (3) | 1 |