Emil Matús

dblp:45/3588 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 1 first-author · 6 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 A Low-Complexity K-Box Detector in High-Dimensional MIMO Systems
abstract
Wireless MIMO communication systems nowadays are driven by the requirement to minimize the hardware computational complexity of detection algorithms while preserving high detection performance. To address this challenge, traditional tree-based approaches like the K-best algorithm have been proposed, which employ a fixed complexity at each layer to manage computational demands. The K-best algorithm remains computationally intensive due to its complexity highly dependent on the QAM modulation size and is further followed by a sorting scheme applied at each layer. Another tree-based approach, called the box decoding, has been developed to overcome these limitations in small-scale MIMO systems. However, the complexity of this algorithm escalates significantly in MIMO systems with higher dimension as the number of candidates generated by the box decoding grows substantially. In this paper, we propose an innovative solution, called the K-box algorithm. In contrast to K-best, K-box decouples the candidate selection procedure and the size of constellation space similar to box decoding while restraining the candidates expansion in higher-dimensional MIMO by sorting only when the number of candidates exceed a certain limit. Testing on 4 × 4 and 8 × 8 64-QAM systems for 5G new radio (NR) link demonstrates SNR gains of 0.6 dB at a BER of 10−2compared to the K-best algorithm, while achieving 77% and 76% complexity reduction in partial Euclidean distance (PED) computations and sorting complexity reductions of 94% and 92% respectively.
Hanfu Zhang, Sheikh Faizan Qureshi, Emil Matús, Dmitry Utyansky, Pieter van der Wolf, Gerhard P. Fettweis
VTC2025-Fall3
2025 Conflict Management in Vector Register Files
abstract
The instruction set architecture of vector processors operates on vectors stored in the vector register file which needs to handle several concurrent accesses by functional units with multiple ports. When the vector processor is running with high utilization, access conflicts become a major source of performance degradation. With a software model of a vector processor, we take a deep dive into the runtime impact of conflicts and their characteristics, and on ways to manage them, i.e., avoidance, resolution, and mitigation. For conflict avoidance, we study the existing approaches of banking with different static bank layouts and propose a dynamic bank layout to overcome their shortcomings. Our approach assigns newly written registers a temporarily unique starting bank. For conflict resolution, we compare different arbitration algorithms and optimize round-robin arbitration for mixed-width arithmetics by prioritizing wide operands. For conflict mitigation, operand queues of varying depths are studied. Our inventions are likely to increase the area efficiency of vector processors, either because they allow to use shallower operand queues while keeping the same performance, or reduce the area even further by using less banks, albeit at a performance impairment of 10% or less. The insights of the study can further be applied to other shared memory systems.
Viktor Razilov, Ipek Geçin, Emil Matús, Gerhard P. Fettweis
ACM Trans. Archit. Code Optim.3
2025 ZuSE-KI-Mobil: AI Chip Design Platform for Automotive and Industrial Applications
Shaown Mojumder, Simon Friedrich, Emil Matús, Matthias Lüders, Martin Friedrich, Oliver Renke, Holger Blume, Markus Kock, Gregor Schewior, Darius Grantz, Jens Benndorf, Julian Höfer, Patrick Schmidt 0003, Jürgen Becker 0001, Nael Fasfous, Pierpaolo Morì, Hans-Jörg Vögel, Samira Ahmadifarsani, Leonidas Kontopoulos, Ulf Schlichtmann, Yun-Jin Li, Gerhard P. Fettweis
IEEE Trans. Very Large Scale Integr. Syst.3
2024 On the Concurrent Multipath Entanglement Distribution in Quantum Networks
abstract
In this paper, we consider the problem of concurrent multipath routing and end-to-end entanglement distribution for online resource allocation in quantum networks. We propose a heuristic algorithm to solve this problem, considering quantum memory, decoherence time, entanglement distribution probability, and fidelity. A time-slotted quantum network operation model is considered based on the cut-off decoherence time of quantum memories. The proposed heuristic is designed for a quantum network with noisy intermediate scale quantum (NISQ) constraints, including fixed quantum memory decoherence time, and probabilistic entanglement generation and swapping. It considers integer linear programming (ILP)-based and heuristic approaches to select multiple paths and resource allocation in the network. Simulations are performed to evaluate the performance of ILP and heuristic based approaches in terms of requests using multipath approach, blocking ratio, and computation time in a small sized quantum network with few requests. Next, the performance of heuristic based approaches are evaluated for a large problem size. The obtained results ensure that performance of ILP and heuristic based approaches are comparable, and multipath routing outperforms single path routing in terms of blocking ratio.
Joy Halder, Emil Matús, Gerhard P. Fettweis
GLOBECOM2
2023 The ZuSE-KI-Mobil AI Accelerator SoC: Overview and a Functional Safety Perspective
abstract
ZuSE-KI-Mobil (ZuKIMo) is a nationally funded research project, currently in its intermediate stage. The goal of the ZuKIMo project is to develop a new System-on-Chip (SoC) platform and corresponding ecosystem to enable efficient Artificial Intelligence (AI) applications with specific requirements. With ZuKIMo, we specifically target applications from the mobility domain, i.e. autonomous vehicles and drones. The initial ecosystem is built by a consortium consisting of seven partners from German academia and industry. We develop the SoC platform and its ecosystem around a novel AI accelerator design. The customizable accelerator is conceived from scratch to fulfill the functional and non-functional requirements derived from the ambitious use cases. A tape-out in 22 nm FDX-technology is planned in 2023. Apart from the System-on-Chip hardware design itself, the ZuKIMo ecosystem has the objective of providing software tooling for easy deployment of new use cases and hardware-CNN co-design. Furthermore, AI accelerators in safety-critical applications like our mobility use cases, necessitate the fulfillment of safety requirements. Therefore, we investigate new design methodologies for fault analysis of Deep Neural Networks (DNNs) and introduce our new redundancy mechanism for AI accelerators.
Fabian Kempf, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Nael Fasfous, Alexander Frickenstein, Hans-Jörg Vögel, Simon Friedrich, Robert Wittig, Emil Matús, Gerhard P. Fettweis, Matthias Lüders, Holger Blume, Jens Benndorf, Darius Grantz, Martin Zeller, Dietmar Engelke, Karl-Heinz Eickel
DATE10
2023 Access Interval Prediction with Neural Networks for Tightly Coupled Memory Systems
abstract
Embedded systems usually integrate multiple Pro-cessing Elements (PEs) on a single chip. Various PEs are con-nected to the same Tightly Coupled Memory (TCM) to increase the area and energy efficiency. However, memory sharing comes at the cost of conflicts resulting in performance degradation. To counteract this issue, Access Interval Prediction (AIP) has been introduced in the literature to predict the interval between two memory accesses. State-of-the-art AIP units are based on predictors proposed for branch prediction, such as TAgged GEometric (TAGE). This work shows for the first time that several types of neural networks are suitable for AIP as well. By treating AIP as a classification problem, we can continue to decrease the error rate compared to the TAGE predictor. For example, Vision Transformer (ViT) networks reduce the average error rate by over one-third to 2.1 percent. Through our investigation, we demonstrate that offline training alone is sufficient since the memory access traces contain the same repetitive patterns independent from the input parameters of the program run.
Simon Friedrich, Chia-Ying Lin, Viktor Razilov, Robert Wittig, Emil Matús, Gerhard P. Fettweis
DSD5
2022 Efficient Synchronization for NR-REDCAP Implemented on a Vector DSP
abstract
A surge of reduced capability and low-cost devices under the 5G framework for many Internet of Things (IoT) applications is foreseen. The new 3GPP standard, NR REDCAP, provides the means to develop such devices. However, keeping these inexpensive devices synchronized with the 5G network requires continuous synchronization algorithms running parallel with the receiver processing. Our analysis and literature agree that the acceleration of these algorithms on Single Instruction Multiple Data (SIMD) architectures is challenging because of the data dependencies in the recursive filter that is usually employed in synchronization algorithms and hence, requires a significant computing budget. The consequent need for activating extra hardware to perform continuous synchronization leads to rising device cost and higher power consumption which is undesirable. In this work, we address this problem by proposing an implementation that efficiently vectorizes the synchronization algorithm, including the recursive filter, on a wide SIMD vector DSP. The results show that the proposed implementation achieves a speedup of 9X over scalar processing. Subsequently, the synchronization kernel can run on a limited MHz budget of the vector DSP.
Sheikh Faizan Qureshi, Stefan A. Damjancevic, Emil Matús, Dmitry Utyansky, Pieter van der Wolf, Gerhard P. Fettweis
ASAP3
2022 Communications Signal Processing Using RISC-V Vector Extension
abstract
Flexible and scalable solutions will be needed for future communications processing systems. RISC-V processors enhanced with vector processing capabilities as specified by the soon-to-be ratified RISC-V vector extension (RVV) pose an interesting base for such systems. Vector processors provide an efficient means of exploiting data-level parallelism, which is heavily present in communications kernels. Furthermore, RVV code is by its design agnostic from the underlying hardware platform which enables scalability. On the exemplary basis of a generalized frequency division multiplexing (GFDM) implementation on a RVV processor, we investigate its baseband processing capabilities and guide through RVV's key features and peculiarities. Our vectorization achieves a speedup of up to 60 times compared to the scalar base case and a throughput of 784 symbols per second. The utilization of 77 % is slightly below more specialized solutions. Nevertheless, this work serves as a baseline for further investigations on flexible and scalable RISC-V vector communications processors.
Viktor Razilov, Emil Matús, Gerhard P. Fettweis
IWCMC2
2022 Accurate Estimation of Service Rates in Interleaved Scratchpad Memory Systems
abstract
The prototyping of embedded platforms demands rapid exploration of multi-dimensional parameter sets. Especially the design of the memory system is essential to guarantee high utilization while reducing conflicts at the same time. To aid the design process, several probabilistic models to estimate the throughput of interleaved memory systems have been proposed. While accurately estimating the average throughput of the system, these models fail to determine the impact on individual processing elements. To mitigate this divergence, we extend three known models to include non-uniform access probabilities and priorities.
Robert Wittig, Philipp Schulz, Emil Matús, Gerhard P. Fettweis
ACM Trans. Embed. Comput. Syst.3
2020 An ASIP Approach to Path Allocation in TDM NoCs using Adaptive Search Region
abstract
Dynamic connection allocation in time-division multiplexed network-on-chip(NoC) is a promising approach to provide a guaranteed service in NoC. Most recently, the trellis-search algorithm demonstrated using an application-specific instruction-set processor (ASIP). The processor achieved a considerable performance improvement for searching optimum paths. However, there is still a scalability issue. As network size increases, the path search time is rapidly increased. The purpose of this work is to introduce a relative search algorithm and to investigate its effect on NoC performance. Our proposed algorithm optimizes path search region and position according to the actual source-destination position on the network grid. Consequently, even though the network size is increased, the path search time of our proposed algorithm is gradually raised. The simulation results using 16x16 2D-mesh showed up to 10 thousand times and 3.5 times decreases in average execution cycles against 32bits RISC and ASIP, respectively.
Seungseok Nam, Emil Matús, Gerhard P. Fettweis
ACM Great Lakes Symposium on VLSI2
2020 Blind Packet-Based Receiver Chain Optimization Using Machine Learning
abstract
The selection of the most appropriate equalization-detection-decoding algorithms in wireless receivers is a challenging task due to the diversity of application requirements, algorithm performance-complexity trade-offs, numerous transmission modes, and channel properties. Typically, the fixed receiver-chain is employed for specific application scenario that may support iterative processing for better adaptation to variable channel conditions. We propose a novel method for optimizing receiver efficiency in the sense of maximizing packet transmission reliability while minimizing receiver processing complexity. We achieve this by packet-wise dynamic selection of the least complex receiver that enables error-free packet reception out of set of available receivers. The scheme employs convolutional neural network (CNN) and supervised deep learning approach for packet classification and subsequent prediction of the optimum receiver using raw baseband signals. The proposed scheme aims to approach a packet error rate close to the rate of the most complex receiver architecture while using a combination of both low and high complexity architectures. This is achieved by employing the neural network based classifier to dynamically select packet-specific optimum architecture; i.e. instead of using the most complex receiver for all packets, the approach dynamically assigns the packet to the most appropriate receiver in terms of equalization-detection-decoding capability and the least possible complexity. We analyze the performance of the proposed scheme considering various channel scenarios. The system demonstrates excellent packet classification performance resulting in the significant performance increase and the reduction of the usage of the functional blocks that can go up to 96% of the time in different scenarios.
Mohammed Radi, Emil Matús, Gerhard P. Fettweis
WCNC2
2019 Queue Based Memory Management Unit for Heterogeneous MPSoCs
abstract
Sharing tightly coupled memory in a multiprocessor system-on-chip is a promising approach to improve the programming flexibility as well as to ease the constraints imposed by area and power. However, it poses a challenge in terms of access latency. In this paper, we present a queue based memory management unit which combines the low latency access of shared tightly coupled memory with the flexibility of a traditional memory management unit. Our passive conflict detection approach significantly reduces the critical path compared to previously proposed methods while preserving the flexibility associated with dynamic memory allocation and heterogeneous data widths.
Robert Wittig, Mattis Hasler, Emil Matús, Gerhard P. Fettweis
DATE3
2019 General Multicarrier Modulation Hardware Accelerator for the Internet of Things
abstract
General frequency division multiplexing (GFDM) provides a proven approach for a flexible physical layer implementation in wireless communication systems. However, this flexibility requires additional processing steps within the critical system path, the hybrid automatic repeat request. This is especially critical for devices of the Internet of Things, which have a very small power footprint. To tackle this problem, we present a data mapping that allows efficient parallel computation of the GFDM algorithm by a standard Harvard CPU architecture. To utilize the new mapping, we derive semi-custom processor configurations based on fixed-point arithmetic, which achieve a throughput close to the theoretical bound. In comparison with a custom FPGA implementation, we deliver 9 percent of the throughput, while only consuming 0.1 percent of the power. Thus, we achieve 14 times higher energy efficiency.
Robert Wittig, Stefan A. Damjancevic, Emil Matús, Gerhard P. Fettweis
GLOBECOM3
2019 5G-and-Beyond Scalable Machines
abstract
5G is not one problem and one solution, but spans a breadth of applications with largely differing requirements. One solution for all seems therefore inadequate. We therefore present a modular signal processor MPSoC architecture which can be tiled into the size to address the requirement as needed. We name it “Kachel”, the German word for “tile”.
Gerhard P. Fettweis, Emil Matús, Robert Wittig, Mattis Hasler, Stefan A. Damjancevic, Seungseok Nam, Sebastian Haas
VLSI-SoC2
2019 Probabilistic Models for Off-Line Arbiters in Embedded Systems
abstract
Sharing scratchpad memory improves memory utilization but incurs conflicts. Existing statistical models for throughput estimation of shared memory systems assume mainly online memory arbitration, i.e., the memory access and arbitration logic are integrated into a single path. However, these models are not suited for modeling memory systems deploying off-line arbitration as they do not reflect the additional latency of such arbiters. To cope with this problem, we extend the existing occupancy and Markov models by including appropriate weighing parameters. We show that our extension can reduce the error of the original model by 71 percent over a wide range of parameters.
Robert Wittig, Mattis Hasler, Emil Matús, Gerhard P. Fettweis
VLSI-SoC3
2019 Towards GFDM for Handsets - Efficient and Scalable Implementation on a Vector DSP
abstract
Generalised frequency division multiplexing (GFDM) is a novel multicarrier waveform with reduced out- of-band emission and peak-to-average-power-ratio compared to orthogonal frequency division multiplexing. Due to these properties, GFDM is regarded a candidate waveform for future wireless communication. However, an implementation that addresses operating conditions of handheld devices has not yet been considered. We propose a software programmable GFDM solution for handheld devices. This paper presents necessary capabilities of a vector DSP implementation to efficiently process GFDM in the context of 5G use cases. We investigate structural properties and numerical precision of the GFDM algorithm, propose a scalable vectorised approach for single instruction multiple data processing, and carry out a performance-cost evaluation. We present our key findings to enable GFDM functionality on handheld user equipment.
Stefan A. Damjancevic, Emil Matús, Dmitry Utyansky, Pieter van der Wolf, Gerhard P. Fettweis
VTC Fall2
2017 A Heterogeneous SDR MPSoC in 28 nm CMOS for Low-Latency Wireless Applications
abstract
Current and future applications impose high demands on software-defined radio (SDR) platforms in terms of latency, reliability, and flexibility. This paper presents a heterogeneous SDR MPSoC with a hexagonal network-on-chip to address these issues. It features four data processing modules and a baseband processing engine for iterative multiple-input multiple-output (MIMO) receiving. Integrated memory controllers enable dynamic data flow mapping and application isolation. In a 4 x 4 MIMO application scenario, the MPSoC achieves a throughput of 232 Mbit/s with a latency of 20 μs while consuming 414 mW. It outperforms state-of-the-art platforms in terms of throughput by a factor of 4.
Sebastian Haas, Tobias Seifert, Benedikt Noethen, Stefan Scholze, Sebastian Höppner, Andreas Dixius, Esther P. Adeva, Thomas R. Augustin, Friedrich Pauls, Sadia Moriam, Mattis Hasler, Erik Fischer, Yong Chen 0014, Emil Matús, Georg Ellguth, Stephan Hartmann 0002, Stefan Schiefer, Love Cederstroem, Dennis Walter, Stephan Henker, Stefan Hänzsche, Johannes Uhlig, Holger Eisenreich, Stefan Weithoffer, Norbert Wehn, René Schüffny, Christian Mayr 0001, Gerhard P. Fettweis
DAC14
2017 Combined Centralized and Distributed Connection Allocation in Large TDM Circuit Switching NoCs
abstract
The centralized methods for connection allocation in a circuit-switched network-on-chip (NoC) based on time-division multiplexing (TDM) may pose serious performance and scalability issues in large-scale networks due to the 1) limited path search speed, 2) increasing allocation request rate at central unit and 3) the increasing communication cost between the central unit and NoC nodes. This paper tackles this problem by proposing a combined centralized-distributed approach that splits the original NoC into multiple non-overlapping logical partitions, each of them served by a dedicated NoC-Manager unit. The NoC-Manager employs fast trellis-search shortest path algorithm enabling local path search inside the associated NoC partition, while a set of NoC-Managers jointly combine the partial results in a distributed manner in order to find the most likely global path. This approach attempts to combine the benefits of distributed and centralized systems, whilst the experimental results demonstrate its high potential regarding performance and scalability improvement.
Yong Chen 0014, Emil Matús, Gerhard P. Fettweis
ACM Great Lakes Symposium on VLSI2
2017 Combined packet and TDM circuit switching NoCs with novel connection configuration mechanism
abstract
In this paper we present a router that combines the circuit switching and packet switching in order to efficiently and separately handle the guaranteed-service and best-effort traffics. The main innovation consists in proposing a novel connection configuration mechanism, in which the source node first sends the connection request to manager via a pre-reserved request path, and the manager sends back the response message via guaranteed-service path. Hence, the additional dedicated configuration network that is widely used in previous works is avoided, which reduces the hardware cost while still guaranteeing the configuration latency. The synthesis results show our approach is more area and energy efficient. Compared to previous works, our approach can provide up to 260% better power efficiency and 2.4X to 5X better area efficiency. In terms of configuration time, our approach can provide 2.7X to 36X faster configuration speed.
Yong Chen 0014, Emil Matús, Gerhard P. Fettweis
ISCAS2
2017 Register-Exchange Based Connection Allocator for Circuit Switching NoCs
abstract
Since Time Division Multiplexing (TDM) Circuit Switching (CS) has the advantage of fixed low communication latency by transmitting data over pre-established connection, it has been a popular approach to provide guaranteed service. The challenge of the CS is the fast and dynamic connection allocation particularly for networks with high connection request rates. In this paper, a high performance connection allocator for TDM CS is presented, which enables parallel multiple path search in all directions. To enhance the path search speed, the Register-Exchange technique is adopted that saves the entire survivor path sequences during search. Hence, since the backtrack is omitted, the path search time is reduced by half compared to previous forward-backtrack approaches, which can also contribute to the success rate. Our approach is compared to the state of the art centralized and distributed approaches under uniform random traffic as well as real-application benchmarks. The experiment results showed our approach can provide up to 22% higher success rate and 2X greater allocation speed against centralized approaches, and up to 33% higher success rate and 18X higher allocation speed against distributed approach.
Yong Chen 0014, Emil Matús, Gerhard P. Fettweis
PDP2
2016 An MPSoC for energy-efficient database query processing
abstract
This paper presents a heterogeneous database hardware accelerator MPSoC manufactured in 28 nm SLP CMOS. The 18 mm2 chip integrates a runtime task scheduling unit for energy-efficient query processing and hierarchical power management supported by an ultra-fast dynamic voltage and frequency scaling. Four processing elements, connected by a star-mesh network-on-chip, are accelerated by an instruction set extension tailored to fundamental data-intensive applications. We evaluate the MPSoC with typical database benchmarks focusing on scans and bitmap operations. When the processing elements operate on data stored in local memories, the chip consumes 250 mW and shows a 96x energy efficiency improvement compared to state-of-the-art platforms.
Sebastian Haas, Oliver Arnold, Benedikt Noethen, Stefan Scholze, Georg Ellguth, Andreas Dixius, Sebastian Höppner, Stefan Schiefer, Stephan Hartmann 0002, Stephan Henker, Thomas Hocker, Jörg Schreiter, Holger Eisenreich, Jens-Uwe Schluessler, Dennis Walter, Tobias Seifert, Friedrich Pauls, Mattis Hasler, Yong Chen 0014, Hermann Hensel, Sadia Moriam, Emil Matús, Christian Mayr 0001, René Schüffny, Gerhard P. Fettweis
DAC22
2016 EUROSERVER: Share-anything scale-out micro-server design
Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Matús, Antimo Bruno, Per Stenström, Jérôme Martin, Yves Durand, Isabelle Dor
DATE6
2016 Trellis-search based Dynamic Multi-Path Connection Allocation for TDM-NoCs
abstract
This paper proposes a centralized approach for connection allocation for TDM-based NoCs by making use of dedicated hardware unit called NoCManager that employs trellis-based search algorithm enabling dynamic parallel multi-path, multi-slot allocation. Be different to the previous unrolled trellis search algorithm, in this paper the folded architecture is employed to achieve efficiency. In comparison with previous TDM connection allocation methods, the proposed design has the following advantages: (1) hardware supported low-latency, high-throughput allocation mechanism, (2) improved success rate due to parallel multi-path search and (3) efficient NoCManager architecture.
Yong Chen 0014, Emil Matús, Gerhard P. Fettweis
ACM Great Lakes Symposium on VLSI2
2014 EUROSERVER: Energy Efficient Node for European Micro-Servers
abstract
EUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers.
Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson
DSD11
2014 Tomahawk: Parallelism and heterogeneity in communications signal processing MPSoCs
abstract
Heterogeneity and parallelism in MPSoCs for 4G (and beyond) communications signal processing are inevitable in order to meet stringent power constraints and performance requirements. The question arises on how to cope with the problem of system programmability and runtime management incurred by the statically or even dynamically varying number and type of processing elements. This work addresses this challenge by proposing the concept of a heterogeneous many-core platform called Tomahawk. Apart from the definition of the system architecture, in this approach a unified framework including a model of computation, a programming interface and a dedicated runtime management unit called CoreManager is proposed. The increase of system complexity in terms of application parallelism and number of resources may lead to a dramatic increase of the management costs, hence causing performance degradation. For this reason, the efficient implementation of the CoreManager becomes a major issue in system design. This work compares the performance and capabilities of various CoreManager HW/SW solutions, based on ASIC, RISC and ASIP paradigms. The results demonstrate that the proposed ASIP-based solution approaches the performance of the ASIC realization, while preserving the full flexibility of the software (RISC-based) implementation.
Oliver Arnold, Emil Matús, Benedikt Noethen, Markus Winter 0002, Torsten Limberg, Gerhard P. Fettweis
ACM Trans. Embed. Comput. Syst.2
2009 ICT-Emuco. An innovative solution for future smart phones
abstract
Mobile communication has become the dominant branch in the communication business over the last decade and is still rapidly growing in the market. With the recent advances in wireless networks and the exponential growth in the usage of multimedia applications, multi-core platforms point to be the solution of feature-rich phones, such as the iPhone or the BlackBerry Storm to deliver the performance comparable to today's computer system. On the other hand, system scalability and flexibility are vital to enable fast time-to-market and allow manufacturers and service providers to be competitive. Use of virtualization techniques and software development to scalable parallel hardware architectures are inevitable outcome to face the migration to multi-core platforms on mobile devices.
Maria Elizabeth Gonzalez, Attila Bilgic, Adam Lackorzynski, Dacian Tudor, Emil Matús, Irv Badr
ICME5
2009 ASIP Decoder Architecture for Convolutional and LDPC Codes
abstract
In this paper we present a multi-mode decoder architecture for convolutional codes and structured low-density parity-check (LDPC) codes based on a novel computation unit that is able to process Min-Sum as well as Add-Compare-Select (ACS) operations. Realized as application-specific instruction set processor (ASIP), this allows decoding of a vast number of different channel codes and implementation of various communication standards' channel coding schemes with just one single IPcore. Implemented in 130 nm technology, this results in a Viterbi decoding throughput of 30 Mbit/s at 200 MHz for an area of 745 kGates and power consumption of 130 mW.
Steffen Kunze, Emil Matús, Gerhard P. Fettweis
ISCAS2
2009 Vectorization of the Sphere Detection Algorithm
abstract
In this paper we present concepts for vectorization of sphere detection algorithms based on regularization of depth first tree search algorithms. Due to data dependant control flow, these tree search algorithms exhibit a highly irregular structure not allowing an efficient collaborative detection of multiple received symbols in parallel. In order to enable parallel symbol processing, a transformation of the irregular tree search algorithm is proposed resulting in a novel regular algorithm structure. Based on this, a concept for a vectorized List Sphere Detector is introduced, employing a SIMD computational model. In addition to this, limiting effects of vector processing are studied, leading to concepts which ease these effects and enable the utilization of vectorization's benefits.
Björn Mennenga, Emil Matús, Gerhard P. Fettweis
ISCAS2
2009 On the Structured Parallelism of Decoders for LDPC Convolutional Codes - an Algebraic Description
abstract
We propose an algebraic framework that captures the parallelism of decoders for LDPC convolutional codes. From this framework, an architectural template of a decoding core is derived and its main aspects are discussed. Furthermore, sophisticated decoding structures are built using the decoding core as basic element.
Marcos B. S. Tavares, Emil Matús, Gerhard P. Fettweis
ISCAS2
2008 Architecture and VLSI realization of a high-speed programmable decoder for LDPC convolutional codes
abstract
In this paper, we present a novel high-speed dual-core programmable decoder architecture for LDPC convolutional codes and their tail-biting versions. This architecture uses a modified Min-Sum algorithm and enables the decoding of a multitude of codes with different node degree distributions, rates and block lengths. We show how the parallelization concepts are derived using the properties of the bipartite graphs underlying the codes. Moreover, the hardware elements composing the architecture will be presented and analyzed in detail. The programmability of the decoder is also considered. Finally, we present the synthesis results for a prototype ASIC which is capable of achieving high decoding throughput still with very high flexibility, relatively low power consumption and small area.
Marcos B. S. Tavares, Steffen Kunze, Emil Matús, Gerhard P. Fettweis
ASAP3
2008 A dual-core programmable decoder for LDPC convolutional codes
abstract
We present the concepts and realization of a highly parallelized decoder architecture for LDPC convolutional codes and tail-biting LDPC convolutional codes. This architecture has a very good scalability and is fully programmable so that it can be applied to several communications and data storage scenarios. The synthesis results show relatively small area consumption for very high decoding speeds.
Marcos B. S. Tavares, Emil Matús, Steffen Kunze, Gerhard P. Fettweis
ISCAS2
2007 A High-Throughput Programmable Decoder for LDPC Convolutional Codes
abstract
In this paper, we present and analyze a novel decoder architecture for LDPC convolutional codes (LDPCCCs). The proposed architecture enables high throughput and can be programmed to decode different codes and blocklengths, which might be necessary to cope with the requirements of future communication systems. To achieve high throughput, the SIMD paradigm is applied on the regular graph structure typical to LDPCCCs. We also present the main components of the proposed architecture and analyze its programmability. Finally, synthesis results for a prototype ASIC show that the architecture is capable of achieving decoding throughputs of several hundreds MBits/s with attractive complexity and power consumption.
Marcel Bimberg, Marcos B. S. Tavares, Emil Matús, Gerhard P. Fettweis
ASAP3
2007 Towards a GBit/s Programmable Decoder for LDPC Convolutional Codes
abstract
We analyze the decoding algorithm for regular time-invariant LDPC convolutional codes as a 3D signal processing scheme and derive several parallelization concepts, which were used to design a novel low-complexity programmable decoder architecture with throughput in the range of 1 Gbit/s at moderate system clock frequencies. The synthesis results indicate that the decoder requires relatively small areas, even when high levels of parallelism are used.
Emil Matús, Marcos B. S. Tavares, Marcel Bimberg, Gerhard P. Fettweis
ISCAS1
2006 Code Generation for STA Architecture
Jie Guo 0007, Torsten Limberg, Emil Matús, Björn Mennenga, Reimund Klemm, Gerhard P. Fettweis
Euro-Par3