Rolf Kraemer

dblp:57/3855 · DBLP profile ↗
← Back
47ranked-venue papers
2as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 3 since 2021Computer networks · 13 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021
YearPublicationVenuePosition
2023 Amplitude- and phase-modulated PSSS for wide bandwidth mixed analog-digital baseband processors in THz communication
abstract
This paper proposes modifications for the parallel sequence spread spectrum (PSSS) modulation scheme, which improve the bit error rate (BER) performance and peak-to-average power ratio (PAPR) at the same time. In our scheme, some data bits are encoded as phase shifts of the spreading sequences. Thus, the number of transmitted sequences can be reduced, and the resulting PAPR is lowered. The improvements proposed here are inspired by code shift keying (CSK) and parallel combinatory spread spectrum (PCSS) systems. The PSSS variant investigated here is based on recently discovered real-value spreading sequences, where all sequence coefficients are defined in the real domain.
Lukasz Lopacinski, Nebojsa Maletic, Rolf Kraemer, Alireza Hasani, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass
VTC2023-Spring3
2022 A Hardware Optimized High Throughput LDPC Decoder Supporting 3 Tb/s in 28 nm CMOS
abstract
This paper proposes an optimized pipelined decoding architecture with seven processing stages for unrolled LDPC decoders. Pipelined design with seven register layers significantly increases the resulting clock frequency. Moreover, we investigate the optimal layout shape for unrolled decoders. This paper's fastest decoder is based on the IEEE 802.11n LDPC(1944,1620) parity-check matrix and achieves 2937 Gb/s of coded throughput after physical design. By optimizing the pipeline, floorplan, and employing a codeword length of 1944 bits, we increased the throughput by 241% compared to the previous fastest LDPC decoder presented in the literature. To the best of our knowledge, it is the fastest soft-decision decoder published so far. The standard min-sum approach is employed for decoding, and the proposed improvements consider changes only on the hardware level.
Lukasz Lopacinski, Alireza Hasani, Goran Panic, Nebojsa Maletic, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass, Rolf Kraemer
PIMRC8
2022 Wakeup Receiver Using Passive Amplification by Means of a Switched SAW Resonator
abstract
A 433 MHz wake-up-receiver has been designed and fabricated in 250 nm BiCMOS technology. In a novel approach, an off-chip switched SAW resonator accumulates energy from the radio wave and subsequently emits it as a voltage pulse. This mechanism delivers passive amplification and filtering of the input signal without additional power consumption. The integrated analog frontend evaluates the height of the voltage pulse in the context of on-off keying. The analog frontend circuitry draws 46 µA - 70 µA (meas.) from a 2.5 V supply. With a bitrate of 10 kbps, the energy per bit efficiency amounts to 17.4 nJ/bit without duty cycling. The active chip area measures 370 µm × 210 µm, In this implementation, a passive pulse voltage amplification of up to 24 dB and an input sensitivity of -78 dBm were measured. A detailed analysis of the switched SAW network in a realistic application shows that a passive voltage amplification of 32 dB is attainable. The functionality of the analog frontend has been verified by measurements.
Georg Meller, Michael Methfessel, Bastian Lindner, Jens Wagner, Rolf Kraemer, Frank Ellinger
SECON5
2022 Ultra high speed 802.11n LDPC decoder with seven-stage pipeline in 28 nm CMOS
abstract
This paper reports our latest implementation results of a fully unrolled LDPC decoder prototyped in 28 nm CMOS technology. The decoder achieves 1218 Gbps coded throughput and consumes a 5.49 mm2chip area. The standard min-sum decoding algorithm with four-bit quantization, five unrolled iterations, (648,540) parity matrix, and a seven-stage pipeline is employed. Such implementation achieves a higher data rate than adaptive degeneration and finite-alphabet decoding algorithms, requires less silicon than the solutions mentioned above, and is fully compliant with the IEEE 802.11n WLAN standard.
Lukasz Lopacinski, Alireza Hasani, Goran Panic, Nebojsa Maletic, Oliver Schrape, Jesús Gutiérrez 0004, Milos Krstic, Eckhard Grass, Rolf Kraemer
VTC Spring9
2021 Design and Implementation Strategy of Adaptive Processor-Based Systems for Error Resilient and Power-Efficient Operation
abstract
The contemporary computing systems are facing two major challenges: excessive power consumption and susceptibility to faults. In order to take advantage of techniques that efficiently address these challenges, the classic ASIC design flow requires some modifications. In this paper, we present a simple and convenient strategy for design and implementation of processor-based systems using highly configurable, cross-layer framework that encompasses techniques such as Adaptive Voltage and Frequency Scaling (AVFS) and Triple Modular Redundancy (TMR). The proposed strategy augments the conventional design flow with two additional steps to integrate the framework's hardware building blocks into the system. Such system is then able to dynamically switch between low power and error resilient operation modes according to the current requirements. By following the proposed strategy, we were able to implement processor-based system that significantly reduces the power consumption / increases the soft error resilience while preserving the performance at negligible area overhead of less than 1%.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DDECS4
2021 Parallel Sequence Spread Spectrum based Ultra-Reliable Low Latency Communication for Factory Automation
abstract
We propose a closed-loop industrial radio system with a star topology for factory automation using Parallel Sequence Spread Spectrum (PSSS) technology. This system implements all significant concepts for synchronization, duplex-communication, channel estimation/equalization based on an innovative transceiver concept. The industrial environment tends to have non-line of sight channels that lead to RMS delay spreads of less than 100 ns. Here, we propose a channel equalization scheme to compensate for these delay spreads. Another challenging aspect is the implementation of a code division duplexing (CDD) scheme. We propose a closed-loop system architecture that can achieve the targets set by Industry 4.0 bodies such as IWSAN. This paper gives an overview of the closed-loop systems concept.
Karthik KrishneGowda, Elias L. Peter, Matthias Scheide, Lara Wimmer, Rüdiger Kays, Eckhard Grass, Rolf Kraemer
ETFA7
2021 Towards Error Resilient and Power-Efficient Adaptive Multiprocessor System using Highly Configurable and Flexible Cross-Layer Framework
abstract
A typical multiprocessor system often needs to support a wide spectrum of applications. Today, error resilience and low power consumption are two crucial, but non-complementary requirements and meeting both simultaneously is difficult. Thus, adaptivity is becoming increasingly important feature for modern computing systems. In this regard, we integrate a highly-configurable framework with a set of cross-layer techniques efficient in improving error resilience / power consumption into a multiprocessor system. The framework intelligently interchanges methods such as Adaptive Voltage and Frequency Scaling (AVFS), Triple Modular Redundancy (TMR) and clock-gating while the system is on-line. Additionally, flexibility as an inherent multiprocessor feature enables dynamical adaptation of the system to the current requirements. Putting all together, a balanced level between the two key metrics is achieved. We conduct numerous experiments to show the advantages of the proposed approach. Finally, we use the results to confirm the benefits and the effectiveness of the framework.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
IOLTS4
2021 Modulation and Coding Schemes for Variable-Rate Parallel Sequence Spread Spectrum
abstract
This paper investigates coding and modulation schemes for wireless communication based on variable-rate parallel sequence spread spectrum (PSSS). PSSS can adapt the spreading gain according to channel conditions. When this modulation technique is combined with external forward error correction (FEC), a challenge arises in adjusting the gains from FEC and PSSS to achieve the highest spectral and energy efficiency. We profile the bit error rate (BER) and energy efficiency as a function of the signal-to-noise ratio (SNR) for PSSS combined with low-density parity-check (LDPC) codes implemented in 28 nm CMOS technology. This analysis reveals that the gain provided by PSSS should be used only in exceptional cases, whereas LDPC codes provide the gain at a lower spectral- penalty. Moreover, we simplify our system by reducing the number of LDPC code rates. As compared to IEEE 802.11n, we exclude the 2/3 and 3/4 FEC code rates from the hardware and compensate these modes by using the variable-rate PSSS.
Lukasz Lopacinski, Alireza Hasani, Nebojsa Maletic, Jesús Gutiérrez 0004, Rolf Kraemer, Eckhard Grass
PIMRC5
2020 Highly Configurable Framework for Adaptive Low Power and Error-Resilient System-On-Chip
abstract
In this paper, a novel, highly configurable framework for low power and error-resilient System-On-Chip is presented. The framework is composed, on the one hand, of the SWIELD configurable flip-flop and on the other hand, of the Chameleon controller. The SWIELD flip-flop is able to operate in three modes. It is driven/configured during runtime via the dedicated controller called Chameleon System Operation Management Unit. The proposed framework is integrated into a complex SoC based on a 32-bit general-purpose processor and the entire system is synthesized using the IHP 130 nm technology library. Numerous simulation experiments have been conducted in order to estimate the system error resilience and power consumption. At expense of negligible area and complexity overhead, the introduced framework shows great potential and excellent results w.r.t. both metrics of interest.
Mitko Veleski, Michael Hübner 0001, Milos Krstic, Rolf Kraemer
DSD4
2020 A Modified Rejection-Based Architecture to Find the First Two Minima in Min-Sum-Based LDPC Decoders
abstract
One of the essential elements of min-sum low-density parity-check (LDPC) decoders is to find the first two minima between the binary messages arriving in the check nodes along with the index of the minimum which are altogether used to compute the messages for sending back to the neighboring variable nodes. The main techniques for this task are tree-based and bit-serial architectures. The latest tree-based architecture, known as rejection-based scheme finds the first two minima and the binary index of the minimum with higher speed than the previous tree-based methods. However, in min-sum LDPC decoders, having one-hot sequence of the minimum of the messages is preferred as it has implementation benefits. In this paper, we modify the existing rejection-based technique to yield the one-hot sequence instead of the binary representation of the minimum index. The proposed modification doesn’t cause any latency in the operation of the module. We also provide the results of the implementation of the modified rejection-based technique and the bit-serial architecture, conducted on a Xilinx Virtex-7 FPGA. The two major architectures are compared in terms of latency, maximum clock frequency, area and power.
Alireza Hasani, Lukasz Lopacinski, Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer
WCNC5
2019 Modular Data Link Layer Processing for THz communication
abstract
In this paper, we demonstrate a modular baseband and modular data link layer processors for wireless communication, which has been designed for a 200 GHz frontend. Although the individual system elements are well known, we combine the performance of parallel baseband and data link layer cores to cover a larger bandwidth. We combine three cores and achieve a single 1.5 GHz channel (3 × 500 MHz). This paper is focused on the digital elements of the demonstrator, especially on the data link layer aspects and field-programmable gate array (FPGA) processing. We discuss the performance of the back-to-back connected demonstrator, with the focus on the data link layer implementation that is included in the baseband chip. The peak data rate achieved by the presented demonstrator is 1920 Mbps. The solution uses forward error correction mechanisms based on convolutional codes at the code rate equal to 3/4 and accepts bit error rate (BER) up to 10-2. Point-to-point and mesh network topologies are supported.
Lukasz Lopacinski, Mohamed Hussein Eissa, Goran Panic, Alireza Hasani, Rolf Kraemer
DDECS5
2019 Characterization and Modeling of SET Generation Effects in CMOS Standard Logic Cells
abstract
Single event transients (SETs) stand out as one of the major causes of soft errors in nanoscale CMOS integrated circuits. To reduce the need for exhaustive circuit simulations in the design of radiation-hard integrated circuits, the cost-effective approaches for characterization and modeling of SET generation effects in standard logic cells are required. In this work, a SPICE-based methodology for characterization of SET generation effects, employing two different SET current models, is presented. Based on the acquired simulation results, the empirical models for the two main SET generation metrics (SET critical charge and SET pulse width) are derived. The SET generation models and the respective model parameters are intended to be used as inputs for the higher-level analysis of SET effects in digital circuits designed with the characterized standard logic cells. By storing the model parameters for each gate in the look-up table, instead of storing the raw data obtained from SPICE simulations, the amount of characterization data can be significantly reduced, allowing to speed up the subsequent SET analysis of a complex circuit.
Marko S. Andjelkovic, Zoran Stamenkovic, Milos Krstic, Rolf Kraemer
IOLTS5
2019 A Modified Shuffling Method to Split the Critical Path Delay in Layered Decoding of QC-LDPC Codes
abstract
Layered (or Turbo) decoding of Low-Density Parity-Check (LDPC) codes is considered as a decoding schedule that facilitates partially parallel architectures for performing iterative algorithms based on belief propagation. It has, on one hand, reduced implementation complexity and memory overhead compared to fully parallel architectures and, on the other hand, higher convergence speed compared to both serial and parallel architectures. In this paper, we introduce a general form of shuffling of the parity-check matrix of quasi-cyclic LDPC (QC-LDPC) codes which can split the critical path delay in layered decoding and therefore improve throughput by allowing higher clock rates. We also reveal a valuable property of Latin squares QC-LDPC codes which makes them a good candidate for the proposed shuffling method. As a result of that property, no special caution of choosing offset values in the proposed generalized shuffling method is required.
Alireza Hasani, Lukasz Lopacinski, Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer
PIMRC5
2019 Beam Entropy of 5G Cellular Millimetre-Wave Channels
abstract
In this paper, we obtain and study typical beam entropy values for millimetre-wave (mm-wave) channel models using the NYUSIM simulator for frequencies up to 100 GHz for the fifth generation (5G) and beyond 5G cellular communication systems. The beam entropy is used to quantify sparse MIMO channel randomness in beamspace. Lower relative beam entropy channels are suitable for memory- assisted statistically-ranked (MarS) and hybrid radio frequency (RF) beam training algorithms. High beam entropies can potentially be advantageous for low overhead secured radio communications by generating cryptographic keys based on the channel randomness in beamspace, especially for sparse multiple-input multiple- output (MIMO) channels. Urban microcell (UMi) and urban macrocell (UMa) cellular scenarios have been investigated in this work for 28, 60, 73, and 100 GHz carrier frequencies and the rural macrocell (RMa) scenario for 3.5 GHz.
Krishan K. Tiwari, Eckhard Grass, John S. Thompson, Rolf Kraemer
VTC Fall4
2019 A survey on Bluetooth multi-hop networks
abstract
Bluetooth was firstly announced in 1998. Originally designed as cable replacement connecting devices in a point-to-point fashion its high penetration arouses interest in its ad-hoc networking potential. This ad-hoc networking potential of Bluetooth is advertised for years - but until recently no actual products were available and less than a handful of real Bluetooth multi-hop network deployments were reported. The turnaround was triggered by the release of the Bluetooth Low Energy Mesh Profile which is unquestionable a great achievement but not well suited for all use cases of multi-hop networks. This paper surveys the tremendous work done on Bluetooth multi-hop networks during the last 20 years. All aspects are discussed with demands for a real world Bluetooth multi-hop operation in mind. Relationships and side effects of different topics for a real world implementation are explained. This unique focus distinguishes this survey from existing ones. Furthermore, to the best of the authors’ knowledge this is the first survey consolidating the work on Bluetooth multi-hop networks for classic Bluetooth technology as well as for Bluetooth Low Energy. Another individual characteristic of this survey is a synopsis of real world Bluetooth multi-hop network deployment efforts. In fact, there are only four reports of a successful establishment of a Bluetooth multi-hop network with more than 30 nodes and only one of them was integrated in a real world application - namely a photovoltaic power plant.
Nicole Todtenberg, Rolf Kraemer
Ad Hoc Networks2
2018 100 Gbit/s End-to-End Communication: Adding Flexibility with Protocol Templates
abstract
High-speed protocol processing that provides data-rates of 100 Gbit/s and beyond to the application stresses the whole communication system up to its outer limits. Such a system can only be utilized by employing highly specialized, application specific protocols, that are tailored for certain communication parameters, such as the packet loss rate. However, the requirements for most applications are not static, and a protocol designer cannot anticipate all possible communication conditions upfront. The contradiction between specialized protocols and unknown communication parameters can be solved by adapting the protocol implementation on demand to the current communication conditions. However, such an approach needs a protocol description language that allows the automatic specialization of protocols. In this paper, we present the Protocol Engine Template Language (PETL), that allows the automatic implementation of protocols by a constructive approach for a variety of communication conditions from protocol implementation templates.
Steffen Büchner 0002, Alireza Hasani, Lukasz Lopacinski, Rolf Kraemer, Jörg Nolte
LCN4
2018 Implementation of a Multi-Core Data Link Layer Processor for THz Communication
abstract
In this paper, we discuss the main challenges and our solutions proposed for implementation of a high-speed data link layer processor. Our target is to achieve processing throughput faster than 20 Gbps. Meanwhile a single core of our implementation achieves ''only'' ~28 Gbps, we propose a multi-core solution that can run up to ~110 Gbps. For this purpose, we use baseband signal splitters and combiners. Alternatively, we come up with Parallel Sequence Spread Spectrum (PSSS). After splitting the input signal, we divide the required processing-effort among a set of parallel baseband and data link layer processors. The discussed data link layer processor uses hybrid-automatic-repeat-request-I (HARQ-I) with link adaptation and selective fragment repetitions. Such solution significantly improves the efficiency of the method and allows to reduce the implementation complexity when compared to HARQ-III. The main issue discussed in the article is power and energy consumption. Our solution consumes maximally 300 mW at 27.9 Gbps, including forward error correction (FEC) engine.
Lukasz Lopacinski, Mohamed Hussein Eissa, Goran Panic, Marcin Brzozowski, Alireza Hasani, Rolf Kraemer
VTC Spring6
2017 A Critical Charge Model for Estimating the SET and SEU Sensitivity: A Muller C-Element Case Study
abstract
This paper presents a critical charge model for estimating the SET and SEU robustness. The proposed model has been derived by analytic fitting of SPICE results, using a Muller C-element designed in 65 and 130 nm bulk CMOS technologies as the target device. The critical charge is expressed in terms of the size of C-element, size of load inverter, supply voltage and temperature, for constant timing parameters of the SET/SEU current pulse. The proposed model could be utilized to calculate the critical charge causing a SET, for both analyzed technologies, with the accuracy comparable to SPICE simulations. The critical charge for SEU was higher than for SET, but the dependencies obtained for SET response were qualitatively similar to those for SEU. This implies that the proposed critical charge model may be applicable for optimizing the SET/SEU robustness evaluation of the circuits involving the Muller C-element. Moreover, the model may also serve as a basis for evaluating the SET/SEU robustness of other standard cells and other technologies, and thus also for analysis of the SET/SEU robustness of complex circuits.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer, Varadan Savulimedu Veeravalli, Andreas Steininger
ATS3
2017 An analysis of the operation and SET robustness of a CMOS pulse stretching circuit
abstract
The cascaded asymmetrically sized inverters can be employed as pulse stretchers, for the measurement of very short single event transient (SET) pulse widths (<; 200 ps). This paper analyzes, through the circuit simulations, the effects of various design and operating parameters on the normal operation and SET robustness of a two-inverter pulse stretcher designed in 250 nm bulk CMOS technology. It was shown that the SET hardness of the pulse stretcher can be enhanced by upsizing all transistors in the pulse stretcher without changing the sizing ratio. The SET hardness can also be improved by upsizing the load, but this approach is less effective than the pulse stretcher upsizing. Both upsizing approaches have a negligible impact on the normal operation of the stretcher, i.e. output pulse width. In addition, the operation and SET robustness of the pulse stretcher can be influenced by the operating temperature and supply voltage variations, and these effects should be considered in the design process. Based on the acquired simulation results, a general approach for the design of a radiation hardened CMOS pulse stretcher has been proposed.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer
DDECS3
2017 Design of an On-chip System for the SET Pulse Width Measurement
abstract
This paper presents a design of an on-chip single event transient (SET) pulse width measurement system. The proposed system has been designed and implemented in IHP's 250 nm bulk CMOS technology and is intended for evaluation of SET effects in standard digital library cells. It is composed of an inverter-based target circuit, a pulse stretcher and a processing unit for counting the SET pulses and measuring the SET pulse width. The realized system is based on the combination of best practices from various existing designs, and it has a fairly simple architecture capable to provide reliable SET characterization. It supports serial interfacing with the external data acquisition unit which transfers the acquired data to the personal computer. The preliminary evaluation through the circuit-level simulations has demonstrated that the proposed design can detect and measure the SET pulse widths from around 100 ps up to 3.5 ns, with the measurement resolution of approximately 100 ps.
Marko S. Andjelkovic, Vladimir Petrovic, Miljana Nenadovic, Anselm Breitenreiter, Milos Krstic, Rolf Kraemer
DSD6
2017 Assessment of the amplitude-duration criterion for SET/SEU robustness evaluation
abstract
The relation between amplitude and duration of the current pulse induced by a high energy ionizing particle has been proposed as a criterion for evaluating the SET and SEU robustness of integrated circuits. This criterion has advantage over the widely accepted critical charge concept in the sense that it is less dependent on the current pulse shape. However, to the best of our knowledge, there is no known report on the impact of design and operating parameters on the amplitude-duration criterion. The need for extensive circuit or device simulations to derive the amplitude-duration curves under varying design and operating parameters makes this approach very time-consuming. In that regard, this work proposes a method to establish the amplitude-duration criterion, with a limited number of circuit simulations, as a rational function in terms of the sizing factors of target and load gates and supply voltage. Initial evaluation on a simple circuit composed of two inverters, designed in 130 nm bulk CMOS technology, has shown that the proposed method provides the accuracy comparable to SPICE simulations.
Marko S. Andjelkovic, Milos Krstic, Rolf Kraemer
IOLTS3
2017 100 Gbit/s End-to-End Communication: Low Overhead On-Demand Protocol Replacement in High Data Rate Communication Systems
abstract
To be able to efficiently utilize high data rates of 100 Gbit/s and beyond, protocols must be carefully selected for specific communication parameters. At the same time, communication parameters, such as data rate/latency requirements and the channel quality, are not static. This contradiction can be solved by switching to the best suited protocol when communication parameters change. However, replacing a protocol is a severe interference in an ongoing transmission that can easily cause performance degradation. In this paper, we present a minimal disruptive replacement approach that allows us to replace protocol implementations for an ongoing transmission.
Steffen Büchner 0002, Jörg Nolte, Alireza Hasani, Rolf Kraemer
LCN4
2017 Data Link Layer Considerations for Future 100 Gbps Terahertz Band Transceivers
abstract
This paper presents a hardware processor for 100 Gbps wireless data link layer. A serial Reed-Solomon decoder requires a clock of 12.5 GHz to fulfill timings constraints of the transmission. Receiving a single Ethernet frame on a 100 Gbps physical layer may be faster than accessing DDR3 memory. Processing so fast streams on a state-of-the-art FPGA (field programmable gate arrays) requires a dedicated approach. Thus, the paper presents lightweight RS FEC engine, frames fragmentation, aggregation, and a protocol with selective fragment retransmission. The implemented FPGA demonstrator achieves nearly 120 Gbps and accepts bit error rate (BER) up to 2e-3 . Moreover, redundancy added to the frames is adopted according to the channel BER by a dedicated link adaptation algorithm. At the end, ASIC synthesis results are presented including detailed statistics of consumed energy per bit.
Lukasz Lopacinski, Marcin Brzozowski, Rolf Kraemer
Wirel. Commun. Mob. Comput.3
2016 100 Gbit/s End-to-End Communication: Designing Scalable Protocols with Soft Real-Time Stream Processing
abstract
With the recent roll-out of 100 Gbit Ethernet technology for high-performance computing applications and the technology for 100 Gbit wireless communication emerging on the horizon, it is just a matter of time until non-high performance computing applications will have to utilize these data rates. Since 10 Gbit/s protocol processing is already challenging for current server machines and simply upscaling the computing resources is no solution, new approaches are needed. In this paper, we present a stream processing based design approach for scalable communication protocols. The stream processing paradigm enables us to adapt the communication protocol processing for a certain hardware configuration without touching the protocol's implementation. We use this design technique to develop a prototype communication protocol for ultra-high throughput applications and we demonstrate how to adapt the protocol processing for a Stable Throughput as well as for a Low Latency scenario. Last but not least, we present the evaluation results of the experiments, which show that the measured throughput respectively latency of the adapted protocol, scales nearly linear with the number of provided interfaces.
Steffen Büchner 0002, Lukasz Lopacinski, Jörg Nolte, Rolf Kraemer
LCN4
2016 Improved turbo product coding dedicated for 100 Gbps wireless terahertz communication
abstract
In this article, an improved turbo product decoding scheme is proposed. The new method is almost as effective as hard decodable low-density parity check codes (HD-LDPC). Due to the modified codeword shape, no external interleavers are required to correct burst errors. If the decoder uses Reed-Solomon (RS) codes, then error correction performance against burst errors is significantly higher than the gain provided by HD-LDPC with an external interleaver. An additional advantage is a possibility to design a dedicated decoder for Virtex7 field programmable gate array (FPGA) serial transceivers. In our case, we use the method for 100 Gbps data link layer processor dedicated for wireless communication in the Terahertz band. The targeted platform is Virtex7 FPGA, but the solution can be easily scaled on other technologies.
Lukasz Lopacinski, Jörg Nolte, Steffen Büchner 0002, Marcin Brzozowski, Rolf Kraemer
PIMRC5
2015 Design and Implementation of an Adaptive Algorithm for Hybrid Automatic Repeat Request
abstract
Transmission efficiency is an interesting topic for data link layer developers. The overhead of protocols and coding should be reduced to a minimum. This maximizes a link throughput. This is especially important for high-speed networks, where a small degradation of efficiency will degrade the throughput by several Gbps. We describe a redundancy balancing algorithm for an adaptive hybrid automatic repeat request with Reed-Solomon coding. We introduce a testing environment, most important technical issues, and results generated on a field programmable gate array. The hybrid automatic repeat request and Reed-Solomon algorithms are explained. We provide a mathematical description, and a block diagram of the adaptation algorithm. All necessary algorithm simplifications are explained in details. The algorithm can be represented by basic operations in hardware. In most cases, it finds the optimal coding for a predefined bit error rate.
Lukasz Lopacinski, Jörg Nolte, Steffen Büchner 0002, Marcin Brzozowski, Rolf Kraemer
DDECS5
2015 Challenges for 100 Gbit/s end to end communication: Increasing throughput through parallel processing
abstract
Today's applications and services become more dependent on fast wireless communication, for the upcoming years data-rate demands of 100Gbit/s can be easily expected. However, fulfilling that demand is a task which cannot simply be solved by upscaling existing technologies. While most of the research tackles the challenges regarding the transmission technology from the physical layer up to base-band processing, we focus on the challenges concerning the handling of that vast amount of data. The overall goal is to bring together the transmission technology with the operating system to create a suitable end-to-end communication solution. In this paper we argue that communication can be understood as a soft-realtime problem and how that helps introducing parallelism into protocol-processing.
Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer, Lukasz Lopacinski, Reinhardt Karnapke
LCN3
2013 Self-organized Bluetooth scatternets for wireless sensor networks
abstract
The wireless Bluetooth standard, generally used for short range point-to-point communication between mobile devices, can in principle connect any number of stations into a multi-hop network. For this purpose a "scatternet" must be built up, which is a collection of small overlapping "piconets". Bluetooth technology is basically well-suited to some types of wireless sensor networks, but such deployment has been been limited due to the difficulty in setting up and maintaining a scatternet. We demonstrate the SFX algorithm, an extension of the earlier SHAPER procedure. The SFX algorithm builds up a scatternet by a decentralized procedure with the properties of self-organization, self-optimization, and self-healing. The procedure has been verified by deployment in a photovoltaic power plant.
Michael Methfessel, Stefan Lange, Rolf Kraemer, Mario Zessack, Steffen Peter
SenSys3
2012 Functional Pattern Generation for Asynchronous Designs in a Test Processor Environment
abstract
Testing asynchronous circuits has been a challenge for several years. Especially, the nondeterministic timing behavior leads to problems during test, since the occurrence of test responses is not aligned to tester cycles. For this reason a test processor solution for asynchronous circuits has been recently provided, compensating the timing uncertainty. This is achieved by realizing an elastic test via asynchronous handshaking. Based on this approach we present a method for generating functional test patterns for the provided architecture.
Steffen Zeidler 0001, Christoph Wolf, Milos Krstic, Rolf Kraemer
Asian Test Symposium4
2011 Design of a Test Processor for Asynchronous Chip Test
abstract
Due to asynchronous timing and arbitration asynchronous designs may behave no deterministically. For the test of such systems, this means that an exact timing, i.e. a tester cycle, of a test response cannot be guaranteed. This behavior makes functional tests of asynchronous designs relatively complex or even impossible. Therefore, this paper presents a concept for performing functional tests of asynchronous designs using a test processor infrastructure. To this end, we propose a low-cost 16-bit microprocessor solution with special support of asynchronous handshake signalling that can either be integrated into the device-under-test (DUT), mounted on the load board of the tester or a combination of both.
Steffen Zeidler 0001, Christoph Wolf, Milos Krstic, Frank Vater, Rolf Kraemer
Asian Test Symposium5
2011 Implementation of Selective Fault Tolerance with conventional synthesis tools
abstract
Circuits implementing the concept of Selective Fault Tolerance according to are fault-tolerant for a specified subset of inputs. In this paper, a new heuristic is presented to make the method of Selective Fault Tolerance applicable to industrial designs. The heuristic can be efficiently implemented by use of conventional design tools. Compared to TMR, the method, in combination with the heuristic, saves a huge amount of area redundancy and fault tolerance is adapted to the real requirements of a system specification. This is demonstrated by experimental results obtained from circuit descriptions in Verilog and a synthesis with the tool Synopsys.
Michael Augustin, Michael Gössel, Rolf Kraemer
DDECS3
2011 Low-complexity integrated circuit aging monitor
abstract
Integrated circuit aging effects are more and more pronounced with the continuous technological downscaling. These effects degrade circuit operation which is mainly observed as increased input-to-output delay of circuit components. Eventually, the circuit falls out of its specifications. Countermeasures are needed to prevent or reduce such degradation. Aging monitoring can be very beneficial since it can predict circuit failure and/or activate mechanisms to avoid failure. Most of the present aging monitors are based on reporting abnormal input-to-output signal delays on the critical path of the circuit. However, present approaches introduce additional circuit complexity, use complicated analog design, use non-standard cells etc. We propose a low-complexity aging monitor based on standard library cells, offering simplicity and flexibility of its design, integration and use. The designer could instantiate many monitors throughout the integrated circuit. The user can simply read the “aging code” placed in a register in each monitor and determine the “age” of the circuit, predict a circuit failure and/or take an appropriate action. This is especially useful in microprocessors which are designed with dependability in mind.
Aleksandar Simevski, Rolf Kraemer, Milos Krstic
DDECS2
2011 Selective fault tolerance for finite state machines
abstract
This paper introduces the concept of Selective Fault Tolerance for sequential circuits. A sequential circuit that is designed according to this method is fault-tolerant for an arbitrarily selected subset of input sequences that are applied in one or more arbitrarily specified states. Once a selected input sequence is applied in a specified state, the circuit guarantees the same degree of fault tolerance as Triple Modular Redundancy (TMR). No fault tolerance is guaranteed in any other case. A simple heuristic algorithm for the design of such circuits is presented and for a benchmark circuit the reduction of area overhead compared to TMR is experimentally determined. The proposed method is a generalization of Selective Fault Tolerance for combinational circuits as described in [1].
Michael Augustin, Michael Gössel, Rolf Kraemer
IOLTS3
2010 60-Ghz OFDM Systems for Multi-Gigabit Wireless LAN Applications
abstract
This paper presents the developed 60-GHz OFDM hardware demonstrators and the newly designed 60-GHz system architecture targeting for multi-gigabit wireless LAN applications. The hardware demonstrators developed so far are made up of maximum 1-Gbps OFDM baseband realized on FPGA platform and 60-GHz super-heterodyne transceivers fabricated by SiGe BiCMOS technologies. Wireless transmission in 60-GHz links is successfully demonstrated in indoor office environments. The next generation 60-GHz OFDM systems are designed to provide data rate higher than 3-Gbps utilizing the channel bandwidth of 2.16-GHz. Advanced channel coding schemes including low-density-parity-check (LDPC) coding are applied to have higher coding gain. RF transceivers are based on sliding IF architecture and designed to be compatible to the channel plans defined in the 60-GHz standards.
Chang-Soon Choi, Eckhard Grass, Maxim Piz, Marcus Ehrig, Miroslav Marinkovic, Rolf Kraemer, Christoph Scheytt
CCNC6
2010 Reducing the area overhead of TMR-systems by protecting specific signals
abstract
This paper presents a new method of fault-tolerant design for combinational circuits. For an arbitrary chosen subset X1of inputs the designed system is fault-tolerant, but not necessarily for the other inputs. For all the inputs from X1the same level of fault tolerance as for Triple Modular Redundancy (TMR) is achieved. Compared to TMR the necessary area can be significantly reduced. Since the subset X1of inputs, for which the system is fault-tolerant, can be chosen by the designer, the proposed fault-tolerant design method is optimally adapted to the real requirements of fault tolerance.
Michael Augustin, Michael Gössel, Rolf Kraemer
IOLTS3
2010 On-line testing of bundled-data asynchronous handshake protocols
abstract
Asynchronous interfaces, being a popular way of dealing with timing closure problems in deep submicron SoCs, pose a serious problem for on-line testing. Their behavior is specified not as a traditional clocked automaton, but as an asynchronous protocol with timing which depends on the clocks of the communicating blocks, wire and gate delays. The problem is exacerbated by the use of non-encoded bundled data buses, whose transitions cannot be reliably detected, as opposed to very expensive self-timed data encodings. This paper presents a further development of our low-cost checkers for asynchronous handshake protocol, which now support bundled data and are interfaced to the scan chains to assist diagnosis and debugging. Further reduction of area requirements for the delay elements, possibilities to define separate delays for each protocol phase, test results collection and checking the timing relationships between the handshake and the data signals are discussed. An application of the data transition detectors to checking of synchronous SDR/DDR protocols is shown.
Steffen Zeidler 0001, Alexandre V. Bystrov, Milos Krstic, Rolf Kraemer
IOLTS4
2009 Ultra low cost asynchronous handshake checker
abstract
This paper presents a new on-line checking scheme for asynchronous handshake protocols. The proposed scheme requires very small chip area while maintaining high coverage for all considered faults which are briefly exposed. In addition to simple pass-fail information the checker provides off-line diagnosis capabilities in order to further analyze the cause of a fault and the time of its occurrence. In order to verify its functionality the checker was proven by performing analogue simulations. In addition the area overhead and the power consumption was determined and compared with existing implementations.
Steffen Zeidler 0001, Marcus Ehrig, Milos Krstic, Michael Augustin, Christoph Wolf, Rolf Kraemer
IOLTS6
2008 60GHz OFDM hardware demonstrators in SiGe BiCMOS: State-of-the-art and future development
abstract
We present 60 GHz OFDM hardware demonstrators developed so far and outline design considerations for future developments. OFDM schemes have been developed to combat multi-path interferences in indoor wireless environments and further optimized for multi-gigabit data transmission in 60 GHz-band. RF and IF analogues front-ends (AFEs) have been developed with high-speed SiGe BiCMOS technologies. Future developments on AFE are mainly devoted to one-chip integration of RF and IF components based on a sliding IF architecture. We have already achieved wireless transmission of about 1 Gbps in an OFDM demonstrator, which will be continuously upgraded and optimized for higher data transmission with better link adaptability.
Chang-Soon Choi, Eckhard Grass, Frank Herzel, Maxim Piz, Klaus Schmalz, Yaoming Sun, Srdjan Glisic, Milos Krstic, Klaus Tittelbach-Helmrich, Marcus Ehrig, Wolfgang Winkler, Rolf Kraemer, Christoph Scheytt
PIMRC12
2008 Asymmetric dual-band UWB / 60 GHz demonstrator
abstract
This paper describes a pragmatic approach on the up-conversion of WiMedia compliant UWB signals into the 60 GHz band. This dual band concept is based on propagation measurements at {3.1-10.6} and 60 GHz. Due to stringent frequency regulation, UWB systems operating in the 3.1 to 10.6 GHz bands are both bandwidth limited and power limited. Therefore, many potential users hesitate to deploy this technology. Up-conversion to the 60 GHz band and enhanced baseband parameters can make UWB technology more reliable and hence, more attractive. An experimental system based on a commercially available UWB development kit and IHPpsilas 60 GHz chip-set has been developed. The system architecture is described. Measurement results indicate that this combination may be interesting for many applications.
Eckhard Grass, Isabelle Siaud, Srdjan Glisic, Marcus Ehrig, Yaoming Sun, Jens Lehmann 0002, Marie-Hélène Hamon, Anne-Marie Ulmer-Moll, Pascal Pagani, Rolf Kraemer, Christoph Scheytt
PIMRC10
2007 60 GHz SiGe-BiCMOS Radio for OFDM Transmission
abstract
This paper reports implementation details of a 60 GHz short range communication system for data rates up to 2 Gbit/s. Based on a MATLAB model the main PHY parameters were derived, and the modules of the analog frontend were specified. For system verification, a demonstrator was developed. The analog RF and IF circuits of the demonstrator are implemented in a 0.25 μm SiGe BiCMOS technology. Measurement results of the analog modules and the complete demonstrator are given. Using OFDM modulation, the maximum data rate transmitted via airlink so far, is 960 Mbit/s.
Eckhard Grass, Frank Herzel, Maxim Piz, Klaus Schmalz, Yaoming Sun, Srdjan Glisic, Milos Krstic, Klaus Tittelbach-Helmrich, Marcus Ehrig, Wolfgang Winkler, Christoph Scheytt, Rolf Kraemer
ISCAS12
2007 Efficient Inner Receiver Design for OFDM-Based WLAN Systems: Algorithm and Architecture
abstract
In this article we propose a complete solution for the so-called inner receiver of an OFDM-WLAN system based on the IEEE 802.11a standard. We concentrate our investigations on three key components forming the inner receiver namely, the synchronizer, the channel estimator and the digital timing loop. The main goal is the joint optimization of the signal processing algorithms along with the implementation friendly VLSI architecture required for these three key components in order to reduce power, area and latency, without compromising the performance excessively. We provide both the mathematical details and extensive computer simulations to validate our design
Alfonso Troya, Koushik Maharatna, Milos Krstic, Eckhard Grass, Ulrich Jagdhold, Rolf Kraemer
IEEE Trans. Wirel. Commun.6
2004 Mapping of High-Level SDL Models to Efficient Implementations for TinyOS
abstract
Wireless sensor networks have attracted much research effort in the past five years. With TinyOS, there exists a widely used, microthreaded operating system targeted for single processor sensor nodes. It requires only minimal processing and memory resources. Many research activities in the wireless sensor area focus on the design of efficient protocols in the MAC and network layer. SDL - a high-abstraction level formal description language - has long been used for modeling, simulation, and verification of communication protocols. In this paper, we present a mapping approach and optimizations to derive efficient component-based event-driven TinyOS applications from SDL models. We show, how the basic SDL concepts are reflected in the TinyOS component architecture.
Daniel Dietterle, Jerzy Ryman, Kai F. Dombrowski, Rolf Kraemer
DSD4
2003 Bluetooth based wireless Internet applications for indoor hot spots: experience of a successful experiment during CeBIT 2001
Rolf Kraemer, Peter Schwander
Comput. Networks1
2002 Shielding TCP from Wireless Link Errors: Retransmission Effort and Fragmentation
Peter Langendörfer, Michael Methfessel, Horst Frankenfeldt, Irina Babanskaja, Irina Matthaei, Rolf Kraemer
J. Supercomput.6
2001 Bluetooth Based Wireless Internet Applications for Indoor Hotspots: Experience of a Successful Experiment during CeBIT 2001
abstract
Wireless Internet access based on wireless LAN or Bluetooth networks will become popular within the next years. New services will be offered that especially make use of location and personal profile to filter Internet access. Moreover "push" services will allow the unsolicited event based information transfer depending on user configurations. We have introduced such services "The Mobile Fairguide" during CeBIT 2001 in a Bluetooth network that covered a full hall of 25000 m/sup 2/ with 130 base-stations. The content was generated from the official CeBIT database. We were able to show that Bluetooth is a usable technology for such applications especially if PDA are used as terminal devices. Moreover we tested our architecture in a real live scenario especially with respect to scalability and mobility. The additional services like guiding, alarming and broadcasting were offered and appreciated by the visitor. The remaining problems resulting from missing protocols and faults in the base-band implementation of the selected chip-set have been fixed in the meantime.
Rolf Kraemer
LCN1
2001 Evaluation of Well-Known Protocol Implementation Techniques for Application in Wireless Networks
Peter Langendörfer, Rolf Kraemer, Hartmut König
J. Supercomput.2
1989 A Communication System Architecture for the Office
abstract
The Communication System Architecture (CSA) distributed system is described. CSA establishes a distributed architecture, supported by a hierarchy of object machines, enabling transparent communication between a set of interconnected heterogeneous computing and communication systems, and provides an object model for building high-quality distributed applications. After a brief introduction, the authors present an overview of the CSA architecture, describe the CSA object machines and communication facilities, discuss the CSA object model, and present the CSA mail service as an example of an application.>
Amine Benkiran, Gérard Durand, Rolf Kraemer, Munir Tag
INFOCOM3