Mladen Berekovic

dblp:b/MladenBerekovic · DBLP profile ↗
← Back
56ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0003-1911-756XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 5 first-author · 18 since 2021Software engineering, systems software and programming languages · 9 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 IMS: Intelligent Hardware Monitoring System for Secure SoCs
abstract
In the modern Systems-on-Chip (SoC), the Advanced eXtensible Interface (AXI) protocol exhibits security vulnerabilities, enabling partial or complete denial-of-service (DoS) through protocol-violation attacks. The recent counter-measures lack a dedicated real-time protocol semantic analysis and evade protocol compliance checks. This paper tackles this AXI vulnerability issue and presents an intelligent hardware monitoring system (IMS) for real-time detection of AXI protocol violations. IMS is a hardware module leveraging neural networks to achieve high detection accuracy. For model training, we perform DoS attacks through header-field manipulation and systematic malicious operations, while recording AXI transactions to build a training dataset. We then deploy a quantization-optimized neural network, achieving 98.7% detection accuracy with2.5 million inferences/s. We subsequently integrate this IMS into a RISC-V SoC as a memory-mapped IP core to monitor its AXI bus. For demonstration and initial assessment for later ASIC integration, we implemented this IMS on an AMD Zynq UltraScale+ MPSoC ZCU104 board, showing an overall small hardware footprint (9.04% look-up-tables (LUTs), 0.23% DSP slices, and 0.70% flip-flops) and negligible impact on the overall design’s achievable frequency. This demonstrates the feasibility of lightweight, security monitoring for resource-constrained edge environments.
Wadid Foudhaili, Aykut Rencber, Anouar Nechi, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
DATE5
2026 Multi-Partner Project: CeCaS Accelerator Design for Efficient Supercomputing in Automotive Systems
abstract
Modern vehicles integrate an increasing amount of computational functionality, driven by the growing complexity of in-vehicle applications. At the same time, automotive system architectures are becoming more centralized, requiring powerful HPC platforms at the core. These platforms must deliver the performance needed for ADAS, AI, and autonomous driving, while also meeting stringent energy efficiency and safety requirements.The CeCaS project addresses these challenges across a wide range of topics and domains of expertise, including processor design in advanced FinFET technology, the transformation of the E/E architecture, and advanced packaging for automotive supercomputing platforms. Within CeCaS, our work focuses on application-specific accelerator design to enable efficient processing of compute-intensive workloads. In this paper, we present our contributions in this area, including the design of hardware accelerators for both conventional and neuromorphic AI workloads, the development and evaluation of representative AI benchmarks, and the use of virtual platforms for early design-space exploration and hardware/software co-design.
Annina Gutermann, Alexey Serdyuk, Fabian Lesniak, Julian Höfer, Hella Toto-Kiesa, Tanja Harbaum, Jürgen Becker 0001, Brian Pachideh, Sven Nitzsche, Moritz Neher, Carmen Weigelt, Jann Krausse, Victor Pazmino Betancourt, Klaus Knobloch, Lukas Groth, Andrija Neskovic, Saleh Mulhem, Mladen Berekovic
DATE18
2026 Evaluating Generative AI for Functional Safety Analysis of Integrated Circuits
Mouadh Ayache, Alessandra Nardi, Aditya Raj Singh, Brian Davenport, Ganapathy Parthasarathy, Teo Cupaiuolo, Mladen Berekovic, Saleh Mulhem
IOLTS7
2026 In Simplicity There Is Strength: The Unreasonable Effectiveness of XGBoost for Generalizable SRAM Stability Prediction
Jihene Bouhlila, Mladen Berekovic
ISCAS2
2026 Fast Reinforcement Learning for Robust Beam Codebooks in Future Communication Systems
abstract
Millimeter wave (mmWave) and terahertz (THz) MIMO systems typically rely on predefined beamforming codebooks for initial access and data transmission. However, these codebooks are often unoptimized for specific conditions, leading to large sizes and significant beam training overhead, thereby complicating support for highly mobile applications. This paper introduces a reinforcement learning framework that optimizes beam patterns using only receive power measurements, adapting to the environment, user distribution, and hardware constraints without prior channel knowledge. The framework explores three reinforcement learning algorithms: Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC). While reinforcement learning has shown promise for beamforming, a comprehensive comparative analysis of advanced RL algorithms under a combination of realistic challenges, such as Non-Line-of-Sight (NLoS) conditions and hardware impairments for adaptive beam codebook design in mmWave/THz systems has been largely unexplored. This paper presents the first such in-depth comparative study. Simulation results demonstrate the superiority of the SAC algorithm, achieving higher beamforming gain and faster convergence compared to DDPG and TD3 in various scenarios, including LoS and NLoS conditions, even with hardware impairments.
Anouar Nechi, Zakaria Narjis, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
IEEE Trans. Commun.4
2025 Multi-Partner Project: Artificial Intelligence in Manufacturing Leading to Sustainability and the Consideration of Human Aspects (AIMS5.0)
abstract
The industrial landscape is undergoing a transformative shift towards Industry 5.0, a paradigm characterized by the convergence of sustainability, digital autonomy, and human-centric design. This article focuses on the adoption, enhancement, and implementation of AI-driven hardware, tools, methodologies, and semiconductor technologies in this progression. We present here a comprehensive strategy from the AIMS5.0 project with the objective of connecting academic developments with practical industrial use, fostering a harmonious relationship between humans and machines to improve efficiency, spur innovation, and enhance adaptability. Hence we show here our global vision, and examples of how the creation of AI-based industrial solutions is supported by novel AI-tool chains, advancements in hardware, and tools supporting human aspects.
Anouar Nechi, Yasin Ghafourian, Belal Abu-Naim, Thomas Gutt, George Dimitrakopoulos 0001, Amira Moualhi, Mladen Berekovic, Pál Varga, Markus Tauber
DATE7
2025 Nail: Not Another Fault-Injection Framework for Chisel-generated RTL
abstract
Fault simulation and emulation are essential techniques for evaluating the dependability of integrated circuits, enabling early-stage vulnerability analysis and supporting the implementation of effective mitigation strategies. High-level hardware description languages such as Chisel facilitate the rapid development of complex fault scenarios with minimal modification to the design. However, existing Chisel-based fault injection (FI) frameworks are limited by their coarse-grained, instruction-level controllability, which restricts the precision of fault modeling. This work introduces Nail, a Chisel-based open-source FI framework that overcomes these limitations by introducing statebased faults. This approach allows fault scenarios based on specific system states instead of just instruction-level triggers, removing the need for precise timing of fault activation. For greater controllability, Nail allows users to arbitrarily modify internal trigger states via software at runtime. To support this, Nail automatically generates a software interface, offering straightforward access to the instrumented design. This enables fine-tuning of fault parameters during active fault-injection campaigns, a feature particularly beneficial for FPGA emulation, where synthesis is time-consuming. Utilizing these features, Nail narrows the gap between the high speed of emulation-based FI frameworks, the usability of software-based approaches, and the controllability achieved in simulation. We demonstrate Nail’s state-based fault injection and software framework by modeling a faulty general-purpose register in a RISC-V processor. Although this might appear straightforward, it requires statedependent fault injection and was previously impossible without fundamental changes to the design. The approach was validated in both simulation and FPGA emulation, where the addition of Nail introduced less than $1 \%$ resource overhead.
Robin Sehm, Christian Ewert, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
DSD4
2025 Cross Technology Prediction for SRAM Stability Analysis
abstract
SRAM stability presents a significant challenge in technology scaling due to process variations. This paper introduces CrossNodeML, a machine learning-based approach that leverages device and bitcell simulations to predict SRAM behavior across new technology nodes using data from previous nodes. Our model is evaluated with three critical SRAM metrics: Access Disturb Margin (ADM), Write Margin (WRM), and Ireadmin. Comparisons with high-sigma verifier tool simulations demonstrate the accuracy of our predictions, significantly reducing the need for full statistical simulations in new nodes. This methodology revolutionizes design decisions and significantly enhances early stability assessments.
Jihene Bouhlila, Rainer Buchty, Mladen Berekovic
ISCAS3
2025 A Lightweight Peripheral Design for RRAM-based LUTs
abstract
Resistive random-access memory (RRAM) has garnered increasing interest due to its compact structure size, low energy requirements, and non-volatility. Recently, RRAM crossbars have shown potential for implementing lookup tables (LUTs) for reconfigurable logic. However, due to its analog properties, RRAM suffers from extensive peripherals, such as analog-to-digital converters (ADCs). In this paper, we tackle this problem by combining signal amplification, logic disjunction, and conversion to the digital domain in a single compact gate. As a result, the proposed RRAM-based LUT (R-LUT) design reduces the area per LUT by a factor of four and the peripheral overhead by more than six times. It is also almost three times smaller than an SRAM-based LUT implemented in the same technology. At the same time, it achieves a competitive access frequency of more than 3 GHz while consuming less than 240 fJ per inference, improving all major aspects of state-of-the-art R-LUT designs.
Philipp Grothe, Christoph Hübner, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
ISCAS4
2025 PAT-ViT: Token Pruning-based Adversarial Tuning for Robust Vision Transformers
abstract
Recent studies have demonstrated that Vision Transformers (ViTs) are vulnerable to adversarial attacks. While adversarial training is a recognized strategy for enhancing model robustness, it demands substantial computational resources during the training process. Furthermore, the high computational complexity of ViTs during inference presents further challenges. To address these challenges, we propose token pruning-based adversarial tuning for robust ViTs (PAT-ViT). PAT-ViT enhances the robustness of ViTs by reducing their predictability to attackers and tuning them using adversarial data. Unlike conventional methods, PAT-ViT reduces training overhead by tuning pre-trained ViTs rather than training ViTs from scratch. Moreover, PAT-ViT improves inference efficiency by pruning partial input image. Compared to the state-of-the-art (SOTA) robust method, PAT-ViT boosts robust accuracy by 7.5% while also achieving a 10.3% increase in clean accuracy. During inference, PAT-ViT reduces FLOPs by 0.7× relative to SOTA. Additionally, it reduces the training time by 7.5× to 13.1× compared to prior works.
Yun-Hao Yang, Yuan-June Luo, Wan-Jung Chen, An-Yeu Wu, Shih-Hsu Huang, Mladen Berekovic
ISCAS6
2025 Lightweight Authenticated Integration and In-Field Secure Operation of System-in-Package
abstract
System in Package (SiP) relies on integrating different chiplets potentially involving many third-party devices and chiplet foundries. This type of advanced packaging technology opens up numerous threat scenarios, especially: (a) the inauthentic and untraceable integration of chiplets into a SiP, (b) the insecure integration of malicious chiplets, which leads to a severe impact on the SiP security in the field. The current solutions require many hardware cryptographic primitives, making them costly and power-hungry. Therefore, a new lightweight solution is needed to ensure secure chiplet integration and secure SiP operation. In this article, we deal with these problems and introduce iTrustlet , as a combination of a physical unclonable function and an authenticated encryption scheme to ensure an authenticated and traceable chiplet integration. We propose a chiplet integration protocol based on iTrustlet and a classical root-of-trust (RoT) to ensure the integrated chiplets are unaltered and unreplaced. To guarantee SiP in-field security, iTrustlet with a hardware firewall (HWF) is proposed. Their interaction leads to two security features: (i) HWF provides a SiP protection mechanism, and (ii) iTrustlet secures the update of HWF rules. In particular, we provide a multilevel solution centralized around iTrustlet , focusing on lightweightness. The implementation results show that area and power overheads are 1.24% and 1.84% in the case of FPGA and 0.49% and 1.2% for ASIC implementation.
Christian Ewert, Andrija Neskovic, Carsten Heinz, Felix Muuss, Alexander Treff, Marc Gourjon, Rainer Buchty, Thomas Eisenbarth 0001, Andreas Koch 0001, Mladen Berekovic, Saleh Mulhem
ACM Trans. Design Autom. Electr. Syst.10
2025 A Systematic Mapping Study on SystemC/TLM Modeling Capabilities in New Research Domains
abstract
With increasingly complex circuits and systems, the need for advanced design methodologies is growing. These methodologies shift the designers’ focus from technology-specific implementations to more abstract electronic system design (ESL). SystemC was developed to address this need. Being an open standard based on C++, SystemC facilitates hardware and software modeling across multiple levels of abstraction, with a particular emphasis on ESL. It is further enhanced by including the transaction-level modeling (TLM) layer, strengthening its capability to model communication between components, and even full-system simulators. Traditionally, SystemC/TLM has been deployed to provide hardware prototypes for software development early in the design process. However, surveys and literature reviews showing other capabilities of SystemC/TLM are scarce. Hence, it is essential to explore SystemC/TLM’s new capabilities in different domains such as in-circuit fault propagation, security assessment, and verification. In this article, we conduct a systematic mapping study (SMS) of SystemC/TLM modeling capabilities in certain research domains. We elaborate on the state-of-the-art ESL with an emphasis on SystemC/TLM-based system modeling. Subsequently, we present how such technologies can be applied to the new research domains within the field of circuit and system modeling, namely: (D.1) architecture exploration, (D.2) power estimation, (D.3) fault-injection analysis, (D.4) functional and security verification, and (D.5) side-channel analysis. This SMS highlights the advantages and disadvantages of the investigated SystemC/TLM capabilities and addresses the open challenges in these domains, concluding that SystemC/TLM offers significant potential in performance evaluation, verification, and security assessment of circuits and systems at ESL.
Ahmed Mahmoudi, Andrija Neskovic, Celine Thermann, Robin Sehm, Christoph Hübner, Tavia Plattenteich, Rolf Meyer, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
ACM Trans. Design Autom. Electr. Syst.9
2024 EMDRIVE Architecture: Embedded Distributed Computing and Diagnostics from Sensor to Edge
abstract
Future automotive architectures are expected to transition from a network-centric to a domain-centered architecture featuring central compute units. Powerful domain controllers or smart sensors alleviate the load on these central units and communication systems. These controllers execute tasks with varying criticalities on heterogeneous multicore processors, and are ideally capable of dynamically balancing the computing load between the central unit and sensors. Here, Artificial Intelligence (AI) capabilities playa crucial role, as it is in high demand for such an automotive architecture. However, AI still requires specialized accelerators to improve their computation performance. Task-oriented distributed computing with criticalities up to ASIL-D necessitates the development and utilization of specialized methodologies, such as safety, through the isolation and abstraction of low-level hardware concepts. Meanwhile, online monitoring and diagnostics become vital features to detect errors during operation. The EMDRIVE architecture includes methods, components, and strategies to enhance the performance, safety, and security of such distributed computing platforms. The nationally funded EMDRIVE project connects its twelve partners from academia and industry and is currently in its intermediate stage.
Patrick Schmidt 0003, Iuliia Topko, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001, Rico Berner, Omar Ahmed, Jakub Jagielski, Thomas Seidler, Markus Abel, Marius Kreutzer, Maximilian Kirschner, Victor Pazmino Betancourt, Robin Sehm, Lukas Groth, Andrija Neskovic, Rolf Meyer, Saleh Mulhem, Mladen Berekovic, Matthias Probst, Manuel Brosch, Georg Sigl, Thomas Wild, Matthias Ernst, Andreas Herkersdorf, Florian Aigner, Stefan Hommes, Sebastian Lauer, Maximilian Seidler, Thomas Raste, Gasper Skvarc Bozic, Ibai Irigoyen Ceberio, Albrecht Mayer
DATE19
2024 Machine Learning for SRAM Stability Analysis
abstract
SRAM stability is a critical challenge in technology scaling due to process variations. In this paper, we introduce a cutting-edge approach leveraging machine learning based on device and bitcell simulation to predict SRAM behavior in high sigma local and global variations. Our focus includes both high-density (HDC) and Low Voltage Cell (LVC) analysis, revealing the Extreme Gradient Boosting Regressor (XGBR) as the top performer for both. This research demonstrates the superior accuracy of the XGBR regressor in predicting key SRAM metrics, such as Access Disturb Margin (ADM), Write Margin (WRM), and Ireadmin, offering a compelling alternative to traditional statistical simulations. The purpose of such prediction is to revolutionize the design process and speed up designers’ decisionmaking.
Jihene Bouhlila, Felix Last, Rainer Buchty, Mladen Berekovic, Saleh Mulhem
ISCAS4
2024 Holistic Framework for Evaluating the Trustworthiness of Integrated Circuits
abstract
New applications such as autonomous driving, cyber-physical systems, or remote surgeries demand integrated circuits (ICs) with an ever-lower tolerance for failure. Typical IC design focuses on the targets of functionality and power, performance, and area. An emerging topic in IC design is trustworthiness. It attempts to unify the various interdependent functional and non-functional aspects, such as correct functionality, reliability, security, and functional safety. Existing methodologies and standards focus on evaluating trustworthiness issues (TIs), i.e., causes of faults, and their effects on only one particular attribute. Instead, TIs should be evaluated on their effect on trustworthiness as a whole, which demands a holistic approach. In this paper, we make two main contributions. The first contribution is a framework with a set of unified evaluation criteria that can be applied across all trustworthiness attributes, and a metric, called the Residual Risk Value (RRV). The latter can be used to assess the residual risk of a TI, where a low RRV indicates low risk remaining, and vice versa. RRV considers the impact and likelihood, and the potential for risk reduction enabled by implementing countermeasures. The second contribution is a questionnaire-based measure that ranks TIs according to the priority of addressing them. The results highlight that TIs that emerge during the early stages of IC development should be treated with greater priority. Further, there is a tendency to prioritize security-related TIs as a greater risk to trustworthy ICs. Meanwhile, TIs affecting well-established aspects of IC design and verification are given a lower priority.
Mouadh Ayache, Enkele Rama, Saleh Mulhem, Mladen Berekovic, Matthias Korb
VLSI-SoC4
2023 SystemC Model of Power Side-Channel Attacks Against AI Accelerators: Superstition or not?
abstract
As training artificial intelligence (AI) models is a lengthy and hence costly process, leakage of such a model's internal parameters is highly undesirable. In the case of AI accelerators, side-channel information leakage opens up the threat scenario of extracting the internal secrets of pre-trained models. Therefore, sufficiently elaborate methods for design verification as well as fault and security evaluation at the electronic system level are in demand. In this paper, we propose estimating information leakage from the early design steps of AI accelerators to aid in a more robust architectural design. We first introduce the threat scenario before diving into SystemC as a standard method for early design evaluation and how this can be applied to threat modeling. We present two successful side-channel attack methods executed via SystemC-based power modeling: correlation power analysis and template attack, both leading to total information leakage. The presented models are verified against an industry-standard netlist-level power estimation to prove general feasibility and determine accuracy. Consequently, we explore the impact of additive noise in our simulation to establish indicators for early threat evaluation. The presented approach is again validated via a model-vs-netlist comparison, showing high accuracy of the achieved results. This work hence is a solid step towards fast attack deployment and, subsequently, the design of attack-resilient AI accelerators.
Andrija Neskovic, Saleh Mulhem, Alexander Treff, Rainer Buchty, Thomas Eisenbarth 0001, Mladen Berekovic
ICCAD6
2023 Practical Trustworthiness Model for DNN in Dedicated 6G Application
abstract
Artificial intelligence (AI) is considered an efficient response to several challenges facing 6G technology. However, AI still suffers from a huge trust issue due to its ambiguous way of making predictions. Therefore, there is a need for a method to evaluate the AI’s trustworthiness in practice for future 6G applications. This paper presents a practical model to analyze the trustworthiness of AI in a dedicated 6G application. In particular, we present two customized deep neural networks (DNNs) to solve the automatic modulation recognition (AMR) problem in Terahertz communications-based 6G technology. Then, a specific trustworthiness model and its attributes, namely data robustness, parameter sensitivity, and security covering adversarial examples, are introduced. The evaluation results indicate that the proposed trustworthiness attributes are crucial to evaluate the trustworthiness of DNN for this 6G application.
Anouar Nechi, Ahmed Mahmoudi, Christoph Herold, Daniel Widmer, Thomas Kürner, Mladen Berekovic, Saleh Mulhem
WiMob6
2023 FPGA-based Deep Learning Inference Accelerators: Where Are We Standing?
abstract
Recently, artificial intelligence applications have become part of almost all emerging technologies around us. Neural networks, in particular, have shown significant advantages and have been widely adopted over other approaches in machine learning. In this context, high processing power is deemed a fundamental challenge and a persistent requirement. Recent solutions facing such a challenge deploy hardware platforms to provide high computing performance for neural networks and deep learning algorithms. This direction is also rapidly taking over the market. Here, FPGAs occupy the middle ground regarding flexibility, reconfigurability, and efficiency compared to general-purpose CPUs, GPUs, on one side, and manufactured ASICs on the other. FPGA-based accelerators exploit the features of FPGAs to increase the computing performance for specific algorithms and algorithm features. Filling a gap, we provide holistic benchmarking criteria and optimization techniques that work across several classes of deep learning implementations. This article summarizes the current state of deep learning hardware acceleration: More than 120 FPGA-based neural network accelerator designs are presented and evaluated based on a matrix of performance and acceleration criteria, and corresponding optimization techniques are presented and discussed. In addition, the evaluation criteria and optimization techniques are demonstrated by benchmarking ResNet-2 and LSTM-based accelerators.
Anouar Nechi, Lukas Groth, Saleh Mulhem, Farhad Merchant, Rainer Buchty, Mladen Berekovic
ACM Trans. Reconfigurable Technol. Syst.6
2022 RemEduLa - Remote Education Laboratory for FPGA Design Technology
abstract
Teaching hardware design is both challenging for teachers and students as it typically requires direct access to the targeted hardware platform for final testing. In this paper, we introduce RemEduLa - Remote Educational Laboratory for FPGA design technology. The core idea is to provide students with a developing experience as close as possible to presence teaching as part of lab courses. Therefore, the physical FPGA board is connected to a hardware server enabling the virtual instrumentation via a web interface. An overlay design with virtual inputs and outputs serves as a gateway, offering the student full control over their FPGA development board. This includes peripherals such as buttons, LEDs, and external components (sensors, actuators) as well as real-time visual feedback via a video stream. This one-to-one mapping of real hardware and students allows for the reuse of exercises formerly conducted in the on-site lab time.
Christopher Blochwitz, Philipp Grothe, Sven Dreier, Waiel Aljnabi, Rainer Buchty, Mladen Berekovic
ISCAS6
2021 A comparative survey of open-source application-class RISC-V processor implementations
abstract
The numerous emerging implementations of RISC-V processors and frameworks underline the success of this Instruction Set Architecture (ISA) specification. The free and open source character of many implementations facilitates their adoption in academic and commercial projects. As yet it is not easy to say which implementation fits best for a system with given requirements such as processing performance or power consumption. With varying backgrounds and histories, the developed RISC-V processors are very different from each other. Comparisons are difficult, because results are reported for arbitrary technologies and configuration settings. Scaling factors are used to draw comparisons, but this gives only rough estimates. In order to give more substantiated results, this paper compares the most prominent open-source application-class RISC-V projects by running identical benchmarks on identical platforms with defined configuration settings. The Rocket, BOOM, CVA6, and SHAKTI C-Class implementations are evaluated for processing performance, area and resource utilization, power consumption as well as efficiency. Results are presented for the Xilinx Virtex UltraScale+ family and GlobalFoundries 22FDX ASIC technology.
Alexander Dörflinger, Mark Albers, Benedikt Kleinbeck, Yejun Guan, Harald Michalik, Raphael Klink, Christopher Blochwitz, Anouar Nechi, Mladen Berekovic
CF9
2018 Hardware-Accelerated Index Construction for Semantic Web
abstract
In this paper, an optimized data structure for managing triples used in a Semantic Web Database and a hardwareengine for index construction are presented. We propose anFPGA-centric design, which we call Hardware-Triplestore. Aspart of the design, a scalable and parallel architecture forTriplestore construction is introduced. We propose a hybrid datastructure consisting of three layers, one for every element ofthe semantic triple. The data structure is optimized for ourhardware-centric design and is stored on an external DDR4-Memory. The Hardware-Triplestore is evaluated separately fromthe rest of the database system and achieves an insertion rateof 1.24 million triples per second, which is 17 times faster thanone of the fastest software Triplestore-RDF-3X-.
Christopher Blochwitz, Julian Wolff, Mladen Berekovic, Dennis Heinrich, Sven Groppe, Jan Moritz Joseph, Thilo Pionteck
FPT3
2017 IR-drop aware Design & technology co-optimization for N5 node with different device and cell height options
abstract
In this paper we propose a novel Design-Technology Co-Optimization (DTCO) framework that enables PDK generation and design implementation of sub-10nm technology nodes. The framework allows to study the impact of different technology options at design level and use effective design Power, Performance and Area (PPA) to decide on right technology option. Design implementation flow is IR-drop aware, allowing integration of optimized Power Delivery Network (PDN) for different device/cell options. Using N5-like technology node assumptions (contacted poly and metallization pitch of 42 and 32nm), we generate digital PDKs for different device (finFET, 2 & 3 nanowires) and standard cell options (3, 2 or 1 fins & 7.5 or 6-Tracks cell height). Different PDKs have been used to implement and characterize a wire dominated circuit. Our study shows that the design PDN/IR-drop awareness is fundamental to complete DTCO approach for sub-10nm nodes. Using our dedicated design methodology we reach the IR-drop target of 2.5% VDD (on the lowest metal layers), while minimizing the area degradation induced by the PDN. Further, we demonstrate that such optimized PDN is mandatory to enable the 20% area gain when moving from 7.5 to 6-Tracks cell height. Finally, we show that the impact of different device options is in range of 15% Power, 2X Performance and 20% Area, further validating the need of a fully integrated DTCO.
Luca Mattii, Dragomir Milojevic, Peter Debacker, Yasser Sherazi, Mladen Berekovic, Praveen Raghavan
ICCAD5
2016 A Scriptable Standard-Compliant Reporting and Logging Framework for SystemC
Rolf Meyer, Jan Wagner, Bastian Farkas, Sven Alexander Horsinka, Patrick Siegl, Rainer Buchty, Mladen Berekovic
ACM Trans. Embed. Comput. Syst.7
2015 Revealing Potential Performance Improvements by Utilizing Hybrid Work-Sharing for Resource-Intensive Seismic Applications
abstract
Heterogeneous system architectures are becoming more and more of a commodity in the scientific community. While it remains challenging to fully exploit such architectures, the benefits in performance and hybrid speed-up, by using a host processor and accelerators in parallel in a non-monolithic matter, are significant. Hereby, the energy efficiency is becoming an increasingly critical challenge for future high-performance computing (HPC) systems, which do want to exceed the Exascale barrier with several competing architecture concepts ranging from high-performance CPUs, combined with GPUs acting as floating-point accelerators, to computationally weak CPUs, paired with dedicated and highly-perform ant FPGA-based accelerators. In this paper, we realize and evaluate a hybrid computing approach based on a two-dimensional seismic streaming algorithm with several heterogeneous system architectures, including conventional HPC approaches based on powerful CPUs and GPUs. Furthermore, we elaborate the effort on an embedded system platform claiming to be a "mini supercomputer" [1]. Several CPU and accelerator combinations are utilized in a manual work-sharing manner with the aim of achieving significant performance speed-ups and a detailed energy-efficiency study. Based on roofline models and experimental evaluations, the paper provides an insight into the fact that hybrid computing is mostly unconditionally beneficial for balanced systems regarding the performance as well as the energy efficiency, aiding the programmer in the decision whether or not costly, manually tuned, homogeneous implementations are worthwhile.
Patrick Siegl, Rainer Buchty, Mladen Berekovic
PDP3
2014 An Efficient Barrier Implementation for OpenMP-Like Parallelism on the Intel SCC
abstract
This paper proposes an effective barrier synchronization implementations for shared memory-based parallel programming models (e.g. OpenMP) on the Intel SCC non-cache- coherent platform. Barrier synchronization primitives are key components of these programming models to coordinate the parallel threads. Therefore, we need an efficient implementation of the underlying synchronization algorithms to allow high-level barrier constructs for better performance. In particular, we present an efficient evaluation method to determine the overhead associated with integration of barrier algorithms that is required for OpenMP runtime libraries. We validate several implementation variants that efficiently use the network topology and SCC-specific hardware. Our experimental results for different Micro- benchmarks show significant performance improvement up to 98% for 48 cores.
Hayder Al-Khalissi, Syed Abbas Ali Shah, Mladen Berekovic
PDP3
2013 Efficient Barrier Synchronization for OpenMP-Like Parallelism on the Intel SCC
abstract
The continuous increase of the number of processing cores on die poses a new set of challenges to HPC applications programming including how to model, write, and verify software that has to use the full power of NoC-based manycore processors. Therefore, to simplify program development for the Single-chip Cloud Computer (SCC), it is desirable to have high-level, shared memory-based parallel programming abstractions (e.g., an OpenMP-like programming model). One of the key components of any similar programming model are barrier synchronization primitives, coordinating the work of parallel threads. To allow high-level barrier constructs to deliver good performance, we need an efficient implementation of the underlying synchronization algorithm. In this paper, we propose effective barrier synchronization implementations for shared-memory programming on non-cache-coherent cluster-on-chip represented by the Intel SCC. In particular, we present an extensive evaluation of the overhead associated with integrating barrier algorithms required for OpenMP runtime libraries on such a machine, validating several implementation variants that efficiently exploit the network topology and leveraging SCC-specific hardware. We provide a detailed evaluation of the performance achieved by different approaches by using micro-benchmarks.
Hayder Al-Khalissi, Rainer Buchty, Mladen Berekovic
ICPADS3
2013 Safe Virtual Interrupts Leveraging Distributed Shared Resources and Core-to-Core Communication on Many-Core Platforms
abstract
Modern many-core platforms offer sufficient redundant resources for increasing availability and fault-tolerance of multiple applications, also of different criticality (mixed-criticality). A suitable platform must allow remapping applications and replacing peripherals dynamically. Mapping to distributed resources but also communication among resources ideally is transparent and flexible to allow changes at run time. Communication additionally has to be predictable, especially for safety-critical applications, and can be efficiently implemented by the use of interrupt requests. This paper presents a scalable interrupt translation mechanism supporting flexible and transparent communication among resources. Our contribution is of particular benefit for legacy applications but also eases development of new applications. A fast and predictable monitoring and control mechanism enforces specified behavior of applications and peripherals communicating with critical applications at run time. This significantly reduces integration effort for mixed-critical applications on a shared platform, and thus makes many-core platforms more attractive for embedded and safety-critical systems.
Boris Motruk, Jonas Diemer, Philip Axer, Rainer Buchty, Mladen Berekovic
PRDC5
2012 Introduction to the Special Section on ESTIMedia'08
abstract
No abstract available.
Mladen Berekovic, Samarjit Chakraborty, Petru Eles, Andy D. Pimentel
ACM Trans. Embed. Comput. Syst.1
2012 Introduction to special section ESTIMedia'09
abstract
No abstract available.
Andy D. Pimentel, Naehyuck Chang, Mladen Berekovic
ACM Trans. Embed. Comput. Syst.3
2011 Evaluation of Interpreted Languages with Open MPI
Matti Bickel, Adrian Knoth, Mladen Berekovic
EuroMPI3
2010 QoR Analysis of Automated Clock-Mesh Implementation under OCV Consideration
abstract
In ASICs with structure sizes of 65nm and below the requirements of precise and robust clock networks continuously increase. High-speed circuits already use full-custom clock-meshes instead of buffer trees. Recently new clock-mesh synthesis tools with more automation have become available which better suit ASIC design flows. This paper provides a QoR analysis of these meshes versus highly optimized buffer trees with respect to timing and power. Furthermore, we analyzed the sensitivity of the topologies to OCV. For this purpose we realized a Monte Carlo analysis in SPICE as basis for STA. A design-dependent evaluation has been performed by applying the clock networks and analysis to six different designs. Independent of OCV, the clock-mesh reduces the global skew by up to 65% at the expense of a medial increase in average power consumption by 57% when compared to the buffer tree. Focussing on a further reduction of power dissipation, possible improvements of the automated clock-mesh implementation are proposed.
Dennis Bode, Mladen Berekovic, Axel Borkowski, Ludger Buker
DSD2
2010 NoC Switch with Credit Based Guaranteed Service Support Qualified for GALS Systems
abstract
In this paper we present a scalable wormhole switch architecture with a credit based guaranteed service implementation. By means of credits for a service guarantee the architecture is also able to deal with mesochronous GALS systems. We extended a regular wormhole switch architecture with a control unit for service configuration during run-time and modified the arbitration policy. These changes result in a marginal area overhead per switch of approximately 4\%. Thus our new architecture provides a simple solution to implement service guarantees without limitation to a fully synchronous system. We synthesized our design with a 65nm technology and achieved a clock frequency of 1GHz. Due to the high clock frequency we are able to get a channel throughput of more than 4GB/sec whereas the total design complexity is 30k gate equivalents.
Tim Kranich, Mladen Berekovic
DSD2
2010 Small-ruleset regular expression matching on GPGPUs: quantitative performance analysis and optimization
abstract
We explore the intersection between an emerging class of architectures and a prominent workload: GPGPUs (General-Purpose Graphics Processing Units) and regular expression matching, respectively. It is a challenging task because this workload -- with its irregular, non-coalesceable memory access patterns -- is very different from the regular, numerical workloads that run efficiently on GPGPUs.
Jamin Naghmouchi, Daniele Paolo Scarpazza, Mladen Berekovic
ICS3
2009 Low-Power ASIP Architecture Exploration and Optimization for Reed-Solomon Processing
abstract
The advent of the mobile age has heavily changed the requirements of today's communication devices. Data transmission over interference-prone wireless channels requires additional steps of data processing, such as forward error correction, to ensure reliable communication. In this work we present RS(63,55) Reed-Solomon encoding and decoding algorithms according to the IEEE 802.15.4a standard executed on dedicated application-specific processor architectures. Algorithmic as well as architectural modifications to speed up execution and well-known low-power techniques to reduce the power consumption are discussed. The speed-up for our proposed designs compared to a general purpose baseline architecture is up to two orders of magnitude. Power reduction due to clock-gating and guarded evaluation results in a 40% power drop and the energy consumption is decreased up to 60x.
Andreas Genser, Christian Bachmann, Christian Steger, Jos Hulzink, Mladen Berekovic
ASAP5
2009 A low-power ASIP for IEEE 802.15.4a ultra-wideband impulse radio baseband processing
abstract
The IEEE 802.15.4a amendment has introduced ultra-wideband impulse radio (UWB IR) as a promising physical layer for energy-efficient, low data rate communications. A critical part of the UWB IR receiver design is the low-power implementation of the digital baseband processing required for synchronization and data decoding. In this paper we present the development of an application-specific instruction-set processor (ASIP) that is tailored to the requirements defined by the baseband algorithms. We report a number of optimizations applied to the algorithms as well as to the hardware architecture. This enables performance increases up to a factor of 122x and energy consumption decreases up to 90x as compared to a 16-bit baseline architecture. Furthermore, this ASIP offers greater flexibility due to programmability as compared to an ASIC implementation.
Christian Bachmann, Andreas Genser, Jos Hulzink, Mladen Berekovic, Christian Steger
DATE4
2008 Mapping of the AES cryptographic algorithm on a Coarse-Grain reconfigurable array processor
abstract
Coarse-Grained reconfigurable architectures are emerging as potential candidates to meet the high performance, power efficiency and flexibility needed by embedded systems. ADRES (Architecture for Dynamically Reconfigurable Embedded Systems) and its DRESC compiler offer a very promising platform for designing embedded systems targeted for different application domains. We present a procedure for mapping the widely used AES cryptographic algorithm on ADRES. A detailed explanation is shown for each of the optimizations performed in order to make better use of instruction and loop parallelism. A new intrinsic function set is proposed for speeding up the processing of the AES algorithm. The obtained simulation results are compared with experiments done on the widely known Texas Instruments DSP: TI C64x, which is considered state-of-the-art for embedded systems. Our results show that ADRES outperforms TI C64x DSP, executing the AES algorithm in one fourth of the cycles.
Andres Garcia, Mladen Berekovic, Tom Vander Aa
ASAP2
2008 Architecture Enhancements for the ADRES Coarse-Grained Reconfigurable Array
Frank Bouwens, Mladen Berekovic, Bjorn De Sutter, Georgi Gaydadjiev
HiPEAC2
2008 Implementation of an UWB Impulse-Radio Acquisition and Despreading Algorithm on a Low Power ASIP
Jochem Govers, Jos Huisken, Mladen Berekovic, Olivier Rousseaux, Frank Bouwens, Michael De Nil, Jef L. van Meerbergen
HiPEAC3
2008 Editorial
Mladen Berekovic, Andy D. Pimentel, Timo Hämäläinen 0001
J. Syst. Archit.1
2007 Mapping control-intensive video kernels onto a coarse-grain reconfigurable architecture: the H.264/AVC deblocking filter
abstract
Deblocking filtering represents one of the most compute intensive tasks in an H.264/AVC standard video decoder due to its demanding memory accesses and irregular data flow. For these reasons, an efficient implementation poses big challenges, especially for programmable platforms. In this sense, the mapping of this decoder's functionality onto a C-programmable coarse-grained reconfigurable architecture named ADRES (architecture for dynamically reconfigurable embedded systems) is presented in this paper, including results from the evaluation of different topologies. The results obtained show a considerable reduction in the number of cycles and memory accesses needed to perform the filtering as well as an increase in the degree of instruction parallelism (ILP) when compared with an implementation on a very long instruction word (VLIW) dedicated processor. This demonstrates that high ILP is achievable on the ADRES even for irregular, data-dependent kernels
C. Arbelo, Andreas Kanstein, Sebastián López, José Francisco López, Mladen Berekovic, Roberto Sarmiento, Jean-Yves Mignolet
DATE5
2007 Ulta-Low-Power Wireless Sensor Node Design on 100 uW Scavenging Energy for Applications In Biomedical Monitoring
abstract
Recent advances in energy scavenging technology pave the way for the design of Wireless Autonomous Transducer Systems (WATS) with a power envelope of up to 100 uW. At IMEC we are working on a set of new technologies that will enable the design of a new class of applications that run on such autonomous transducer systems. Especially the field of health- and bio-monitoring on ECG or EEG signal looks particularly promising for such autonomous sensor nodes.
Mladen Berekovic
DSD1
2003 HiBRID-SoC: A Multi-Core System-on-Chip Architecture for Multimedia Signal Processing Applications
Hans-Joachim Stolberg, Mladen Berekovic, Lars Friebe, Sören Moch, Sebastian Flügel, Xun Mao, Mark Bernd Kulaczewski, Heiko Klußmann, Peter Pirsch
DATE2
2003 HiBRID-SoC: a multi-core architecture for image and video applications
abstract
The HiBRID-SoC multi-core architecture targets a wide range of application fields with particularly high processing demands, including general signal processing applications, video de-/encoding, image processing, or a combination of these tasks. For this purpose, the HiBRID-SoC integrates three fully programmable processor cores and various interfaces on a single chip, all tied to a 64-bit AMBA AHB bus. The processor cores are individually optimized to the particular computational characteristics of different application fields, complementing each other to deliver high performance levels with high flexibility at reduced system costs. The HiBRID-SoC is fabricated in a 0.18 /spl mu/m 6LM standard- cell technology, occupies about 82 mm/sup 2/, operates at 145 MHz, and consumes 3.5 Watts.
Hans-Joachim Stolberg, Mladen Berekovic, Lars Friebe, Sören Moch, Sebastian Flügel, Mark Bernd Kulaczewski, Peter Pirsch
ICIP (3)2
2003 HiBRID-SoC: A Multi-Core System-on-Chip Architecture for Multimedia Signal Processing
Hans-Joachim Stolberg, Mladen Berekovic, Lars Friebe, Sören Moch, Mark Bernd Kulaczewski, Peter Pirsch
VLSI-SOC2
2002 A platform-independent methodology for performance estimation of streaming media applications
abstract
A methodology for performance estimation of streaming media applications on different implementation platforms is presented. The methodology derives a complexity profile for an application as a platform-independent metric, and enables performance estimation on different platforms by correlating the complexity profile with platform-specific data. By example of an MPEG-4 advanced simple profile (ASP) video decoder, performance estimation results are presented for different platforms, including general-purpose processors and specialized architectures. As one particular benefit, the approach can be employed to assist in design decisions in the specification phase of new architectures.
Hans-Joachim Stolberg, Mladen Berekovic, Peter Pirsch
ICME (2)2
2002 Multicore system-on-chip architecture for MPEG-4 streaming video
abstract
The newly defined MPEG-4 Advanced Simple (AS) profile delivers single-layered streaming video in digital television (DTV) quality in the promising 1-2 Mbit/s range. However, the coding tools involved add significantly to the complexity of the decoding process, raising the need for further hardware acceleration. A programmable multicore system-on-chip (SOC) architecture is presented which targets MPEG-4 AS profile decoding of ITU-R 601 resolution streaming video. Based on a detailed analysis of corresponding bitstream statistics, the implementation of an optimized software video decoder for the proposed architecture is described. Results show that overall performance is sufficient for real-time AS profile decoding of ITU-R 601 resolution video.
Mladen Berekovic, Hans-Joachim Stolberg, Peter Pirsch
IEEE Trans. Circuits Syst. Video Technol.1
2001 A programmable co-porcessor for MPEG-4 video
abstract
A programmable processor architecture for MPEG-4 video is proposed, that can serve as a co-processor module in MPEG-4 decoder systems. it consists of a 64-bit dual-issue VLIW macroblock engine, a separate RISC core for bitstream parsing and system processing, and an autonomous I/O processor. A separate DSP is used for MPEG audio support. The architecture is fully programmable and supports parallelism on data-, instruction- and thread-level to cope with the high flexibility and processing demands of the MPEG-4 standard. The first implementation will support realtime decoding of MPEG-4 advanced simple profile or of MPEG-4 ACE-profile (CCIR601, single-object). Future designs will add support for object-based MPEG-4 functionalities. The paper focuses on the architecture, instruction set, and performance of the macroblock engine, which operates as an autonomous co-processor and carries most of the workload in MPEG-4 video processing.. It has a RISC-based architecture with support for parallel processing of instructions and data. Special instructions are implemented with specific support for video processing.
Mladen Berekovic, Hans-Joachim Stolberg, Peter Pirsch, Holger Runge
ICASSP1
2001 Implementing The MPEG-4 Advanced Simple Profile For Streaming Video Applications
abstract
With new tools for advanced coding efficiency, such as global motion compensation and quarter-pel motion compensation, the newly defined visual MPEG-4 Advanced Simple (AS) Profile targets in particular the increasingly important field of streaming video applications in the promising 1–2 MBit/s range. Based on an analysis of MPEG-4 AS profile bitstream statistics, the implementation on a programmable multimedia processor platform is described, achieving real-time decoding performance for ITU-R 601 resolution video.
Hans-Joachim Stolberg, Mladen Berekovic, Peter Pirsch, Holger Runge
ICME2
2000 Architecture of an Image Rendering Co-Processor for MPEG-4 Systems
abstract
The TANGRAM VLSI co-processor is intended as a building block for use in system-on-chip (SOC) designs for the versatile MPEG-4 multimedia standard. It is designed to perform the computation intensive final step of MPEG-4 video decoding: compositing of scenes at the display. This includes warping and alpha blending of multiple full-screen video textures in real-lime. TANGRAM consists of a RISC control processor and multiple powerful arithmetic units that perform rendering calculations directly in hardware. This hybrid architecture enables adaptation to changes in algorithms or software support for different video-formats. Communication to a host CPU and video decoding hardware is done via the very common PI-bus on-chip interface. TANGRAM directly interfaces with the ITU-R601/656 digital video output. VHDL implementation and synthesis for a 0.35 /spl mu/ standard-cell library provide an estimate of 100 MHz achievable clock-frequency (worst-case), 52 mm/sup 2/ overall area and 1 Watt power dissipation. TANGRAM has sufficient performance for rendering of MPEG-4 Main Profile@Layer3 scenes (CCIR).
Mladen Berekovic, Peter Pirsch, Thorsten Selinger, Kai-Immo Wels, Carolina Miro, Anne Lafage, Christoph Heer, Giovanni Ghigo
ASAP1
2000 Co-processor architecture for MPEG-4 main profile visual compositing
abstract
The TANGRAM VLSI co-processor is intended to assist existing MPEG-4 video-decoders to perform the computation intensive final step of MPEG-4 video decoding: compositing of scenes at the display. TANGRAM consists of a RISC control processor and multiple powerful arithmetic units that perform rendering calculations directly in hardware. This hybrid architecture enables adaptation to changes in algorithms or support for different video-formats in software. Communication to a host CPU and video decoding hardware is done via the very common PI-bus on-chip interface. TANGRAM directly interfaces with the ITU-R601/656 digital video output. VHDL implementation and synthesis for a 0.35 /spl mu/ standard-cell library provide an estimate of 100 MHz achievable clock-frequency (worst-case), 52 mm/sup 2/ overall area and 1 Watt power dissipation. TANGRAM has sufficient performance for rendering of MPEG-4 Main Profile@Layer3 scenes (CCIR).
Mladen Berekovic, Peter Pirsch, Thorsten Selinger, Kai-Immo Wels, Carolina Miro, Anne Lafage, Christoph Heer, Giovanni Ghigo
ISCAS1
2000 The M-PIRE MPEG-4 codec DSP and its macroblock engine
abstract
M-PIRE is a programmable MPEG-4 multimedia codec VLSI for mobile and stationary applications. It integrates a RISC core, two separate DSPs, a 64-bit dual-issue VLIW macroblock engine, and an autonomous I/O processor on a single chip to cope with the high flexibility and processing demands of the MPEG-4 standard. The first M-PIRE implementation will consume 90 mm/sup 2/ in 0.25 /spl mu/ CMOS technology. It will support real-time video and audio processing of MPEG-4 simple profile or ITU H.26x standards; future designs of M-PIRE will add support for higher MPEG-4 profiles. This paper focuses on the architecture, instruction set, and performance of M-PIRE's macroblock engine, which carries most of the workload in MPEG-4 video processing.
Hans-Joachim Stolberg, Mladen Berekovic, Peter Pirsch, Holger Runge, Henning Möller, Johannes Kneip
ISCAS2
2000 Coprocessor architecture for MPEG-4 video object rendering
Christoph Heer, Carolina Miro, Anne Lafage, Mladen Berekovic, Giovanni Ghigo, Thorsten Selinger, Kai-Immo Wels
VCIP4
1998 An Array Processor Architecture with Parallel Data Cache for Image Rendering and Compositing
abstract
This paper proposes a new array architecture for MPEG-4 image compositing and 3D rendering. The emerging MPEG-4 standard for multimedia applications allows VRML-like script-based compositing of audio-visual scenes from multiple audio and visual objects. MPEG-4 supports both natural (video) and synthetic (3D) visual objects or a combination of both. Objects can be manipulated e.g. positioned, rotated, warped or duplicated by user interaction. A coprocessor architecture is presented, that works in parallel to an MPEG-4 video and audio-decoder and a floating-point geometry-processor. It performs computation and bandwidth intensive low-level tasks for image compositing and rasterization. The processor consists of a SIMD array of 16 identical DSPs to reach the required processing power for real-time image warping, alpha-blending, z-buffering and phong-shading. The processor has an object-oriented parallel cache architecture with 2D virtual address space (e.g. textures) that allows concurrent and conflict-free access to shared image data objects for all 16 DSPs.
Mladen Berekovic, Peter Pirsch
Computer Graphics International1
1998 Realization of a Programmable Parallel DSP for High Performance Image Processing Applications
abstract
Architecture and design of the HiPAR-DSP, a SIMD controlled signalprocessor with parallel data paths, VLIW and novel memory design.The processor architecture is derived from an analysis of thetarget algorithms and specified in VHDL on register transfer level.A team of more than 20 graduate students covered the whole designprocess, including the synthesizable VHDL description, synthesis,routing and backannotation as the development of a complete softwaredevelopment environment.The 175mm{2}, 0.5µm 3LM CMOSdesign with 1.2 million transistors operates at 80 MHz and achievesa sustained performance of more than 600 million arithmetic operations.
Jens Peter Wittenburg, Willm Hinrichs, Johannes Kneip, Martin Ohmacht, Mladen Berekovic, Hanno Lieske, Helge Kloos, Peter Pirsch
DAC5
1998 A flexible processor architecture for MPEG-4 image compositing
abstract
This paper proposes a new array architecture for MPEG-4 image compositing. The emerging MPEG4 standard for multimedia applications allows script-based compositing of audiovisual scenes from multiple audio and visual objects. MPEG-4 supports both, natural (video) and synthetic (3D) visual objects or a combination of both. Objects can be manipulated, e.g. positioned, rotated, warped or duplicated by user interaction. A coprocessor architecture is presented, that works in parallel to an MPEG-4 video- and audio-decoder, and performs computation and bandwidth intensive low-level tasks for image compositing. The processor consists of an SIMD array of 16 DSPs to reach the required processing power for real-time image warping, alpha-blending and 3D rendering tasks. A programmable architecture allows one to adapt the processing resources to the specific needs of different tasks and applications. The processor has an object-oriented cache architecture with 2D virtual address space (e.g. textures), that allows concurrent and conflict-free access to shared data objects for all 16 DSPs. Especially I/O intensive tasks like texture-mapping, alpha-blending, image warping, z-buffer and shading algorithms benefit from shared memory caches and the possibility to preload data before it is accessed.
Mladen Berekovic, Rainer Frase, Peter Pirsch
ICASSP1
1997 An algorithm-hardware-system approach to VLIW multimedia processors
abstract
A number of recently published DSPs and multimedia processors emphasize on Very Long Instruction Word (VLIW) architectures to achieve flexibility, processing power and high-level language programmability needed for future multimedia applications. In this paper we show that exclusive exploitation of instruction level parallelism decreases in efficiency as the degree of parallelism increases. This is mainly caused by algorithm characteristics, VLSI design and compiler restrictions. We discuss selected aspects from these fields and possible solutions to upcoming bottlenecks from a practical point of view.
Johannes Kneip, Mladen Berekovic, Peter Pirsch
MMSP2