Andreas Peter Burg

dblp:69/9547 · also Andreas Burg 0001 · DBLP profile ↗
← Back
108ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0002-7270-5558ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 66 · 4 first-author · 10 since 2021Computer networks · 14 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Security and privacy · 2Theory of computation · 2
YearPublicationVenuePosition
2026 A Full-Custom Time-Domain Unary Sorter for Soft-Information Decoding in 65 nm CMOS
Michel Cancalon, Ludovic Damien Blanc, Andreas Peter Burg
ISCAS5
2026 HDPC Codes with LDPC Matrices: Construction Based on Social Golfer Problem
Yifei Shen 0003, Hasan Said Ünal, Andreas Peter Burg
ISIT3
2026 Dynamic Dual-Window Decoding for SC-LDPC Codes with Wave Enhancement
Leyu Zhang, Yuqing Ren, Andreas Peter Burg
ISIT3
2025 An SDR-Based Monostatic Wi-Fi System with Analog Self-Interference Cancellation for Sensing
abstract
Wireless sensing offers an alternative to wearables for contactless monitoring of human activity and vital signs. However, most existing systems use bistatic setups, which suffer from phase imperfections due to unsynchronized clocks. Monostatic systems overcome this issue, but are hindered by strong self-interference (SI) that requires effective cancellation. We present a monostatic Wi-Fi sensing system that uses an auxiliary transmit RF chain to achieve SI cancellation levels of 40 dB, comparable to existing solutions with custom cancellation hardware. We demonstrate that the cancellation filter weights, fine-tuned using least mean squares, can be directly repurposed for target sensing. Moreover, we achieve stable SI cancellation over 30 minutes in an office environment without fine-tuning, enabling traditional vital sign monitoring using channel estimates derived from baseband samples without the adaptation of the cancellation affecting the sensing channel – a significant limitation in prior work. Experimental results confirm the detection of small, slow-moving targets, representative for breathing chest movements, at distances up to 10 meters in non-line-of-sight conditions.
Andreas Toftegaard Kristensen, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ISCAS3
2025 Belief Propagation Decoding for Short Codes on Structured Sparse Parity-Check Matrices
abstract
As successfully adopted in standard long code scenarios, belief propagation (BP) decoding has been considered a promising universal decoding candidate for next-generation wireless communications. However, when applied to short codes, BP decoding suffers from poor error correction performance due to harmful cycle structures in the Tanner graph. In this paper, we address this issue by designing a structured, sparse parity-check matrix (ssPCM) framework, composed of multiple cycle-free parity-check row blocks (PCRBs). The resulting ssPCMs feature regular row weights and perform better than the state-of-theart 4 -cycle-free row redundant PCMs across Bose-Chaudhuri-Hocquenghem (BCH) codes of length 63.
Yifei Shen 0003, Zongyao Li 0003, Emmanuel Boutillon, Wenqing Song, Yuqing Ren, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg
ISIT8
2025 Impact of Reactive Jamming Attacks on LoRaWAN: a Theoretical and Experimental Study
abstract
This paper investigates the impact of reactive jamming on LoRaWAN networks, focusing on showing that LoRaWAN communications can be effectively disrupted with minimal jammer exposure time. The susceptibility of LoRa to jamming is assessed through a theoretical study of how the frame success rate is impacted by only a few jamming symbols. Different jamming approaches are studied, among which repeated-symbol jamming appears to be the most disruptive, with sufficient jamming power. A key contribution of this work is the proposal of a software-defined radio (SDR)-based jamming approach implemented on GNU Radio that generates a controlled number of random symbols, independent of the standard LoRa frame structure. This approach enables precise control over jammer exposure time and provides flexibility in studying the effect of jamming symbols on network performance. The theoretical analysis is validated through experimental results, where the implemented jammer is used to assess the impact of jamming under various configurations. Our findings demonstrate that LoRa-based networks can be disrupted with a minimal number of symbols, emphasizing the need for future research on stealthy communication techniques to counter such jamming attacks.
Amavi Dossa, Andreas Peter Burg, El Mehdi Amhoud
PIMRC2
2025 Toward Universal Belief Propagation Decoding for Short Binary Block Codes
abstract
Belief propagation (BP) decoding has been recognized for its capacity-approaching performance and high throughput when decoding long low-density parity-check (LDPC) codes. However, the application of BP decoding for short codes is hindered by dense parity-check matrices (PCMs) and prevalent short cycles in the Tanner graph. In this paper, we introduce a general method to extract an optimized sparse PCM for short binary block codes, which removes length-four cycles and enhances the connectivity of short cycles to enable BP decoding with improved performance. Notably, for short binary codes with lengths up to 64, our BP decoding performance approaches the maximum likelihood bound and surpasses the best-reported BP results with reduced computational complexity. Compared with other universal decoding algorithms, BP decoding using our extracted sparse PCMs is competitive in terms of both error-rate performance and computational complexity. These promising results suggest that our method to improve BP decoding for short codes is a step toward a practical universal BP decoder for next-generation communication systems.
Yifei Shen 0003, Zongyao Li 0003, Yuqing Ren, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg
IEEE J. Sel. Areas Commun.8
2025 Edge-Spreading Raptor-Like LDPC Codes for 6G Wireless Systems
abstract
Next-generation channel coding has stringent demands on throughput, energy consumption, and error rate performance while maintaining key features of 5G New Radio (NR) standard codes such as rate compatibility, which is a significant challenge. Due to excellent capacity-achieving performance, spatially-coupled low-density parity-check (SC-LDPC) codes are considered a promising candidate for next-generation channel coding. In this paper, we propose an SC-LDPC code family called edge-spreading Raptor-like (ESRL) codes. Unlike other SC-LDPC codes that adopt the structure of existing rate-compatible LDPC block codes before coupling, ESRL codes maximize the possible locations of edge placement and focus on constructing an optimal coupled matrix. Moreover, a new graph representation called the unified graph is introduced. This graph offers a global perspective on ESRL codes and identifies the optimal edge reallocation to optimize the spreading strategy. We conduct comprehensive comparisons of ESRL codes and 5G-NR LDPC codes. Simulation results demonstrate that when all decoding parameters and complexity are the same, ESRL codes have obvious advantages in error rate performance and throughput compared to 5G-NR LDPC codes in some specific scenarios (low and high number of iterations), making them a promising solution towards next-generation channel coding.
Yuqing Ren, Leyu Zhang, Yifei Shen 0003, Wenqing Song, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Commun.7
2024 A Low-Latency and High-Performance SCL Decoder with Frame-Interleaving
abstract
In this paper, we describe a frame-interleaving hardware architecture for a generalized node-based successive cancellation list (SCL) decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we also exploit graph ensembles to diversify the decoding, enhancing the error-correcting performance by 0.28 dB and reducing the worst-case latency for serial graph processing by over 32%. Implementation results show that the proposed SCL decoder with frame-interleaving architecture achieves a throughput of 7.15 Gbps and an area efficiency of 37.63 Gbps/mm2, which is 1.56× and 1.11× better than the state-of-the-art node-based SCL decoders.
Leyu Zhang, Yuqing Ren, Yifei Shen 0003, Wuyang Zhou, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg
ISCAS7
2024 Learning-based Hand Gesture Classification using Channel Impulse Response with UWB
abstract
The channel impulse response (CIR) of the wireless propagation channel is influenced by the surrounding environment and thus can be used to retrieve environmental information such as location, the presence of objects, and the speed of objects. In this work, we detect hand gestures based on the complex-valued CIR from an ultra-wideband (UWB) transmission link between one transmitter and two receivers. Thanks to the high path delay resolution due to the wide (500 MHz) bandwidth, we can focus on the channel that is influenced by hand gestures in the sensing area. Using different machine learning and deep learning methods, we learn the features from sequences of CIR snapshots that are correlated to hand gestures and recognize them.
Clément Samanos, Han Miao, Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
PIMRC5
2024 Monostatic Multi-Target Wi-Fi-Based Breathing Rate Sensing Using Openwifi
abstract
Continuous monitoring of human activity and vital signs has become increasingly important in healthcare and consumer applications. While wearables and others sensors are already in widespread use, they often suffer from issues such as discomfort and the need for physical contact. To address these limitations, wireless sensing using radio signals has emerged as a promising alternative. However, most previous work has focused on bistatic setups utilizing commodity Wi-Fi devices. In this work, we use the open-source openwifi platform for monostatic multi-antenna Wi-Fi sensing for contactless breathing rate measurements. We also present a method for simultaneous estimation of the breathing rate of two targets. Experimental results using CNC machines for emulating human breathing demonstrate the effectiveness of our method and setup for targets at a few meters distance. Even in challenging scenarios where the breathing rates are similar, the targets are close together, and the targets have the same angle of arrival, we achieve an average error of 0.61 breaths per minute. We further validate our setup and method by collecting human breathing data from two subjects, which can clearly be distinguished using our setup and method. Moreover, our results indicate that self-interference cancellation is not necessary for close-range breathing rate sensing.
Andreas Toftegaard Kristensen, Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
WCNC4
2024 Analytical Modeling of Short-Channel MOSFET Differential Pair Non-Linearity
abstract
Energy efficiency is of utmost importance in modern applications. Power consumption optimisation could be improved by a comprehensive analytical modeling of the characteristics of critical blocks in a system. Dynamic range (DR) has a strong effect on the power consumption of analog circuits, and is determined by circuit non-linearity and noise level. Noise is well modelled even in deep sub-micron technologies, yet there is a lack of analysis and modeling of the non-linearity. An analytical MOSFET differential pair non-linearity model is presented in this work. The proposed model is universal to a wide range of technologies from long to ultra-deep sub-micron devices, and is valid for all operating regions as it is based on the EKV MOSFET model. Furthermore, a model including drain-voltage-induced non-linearity is also developed, and a concise 3dB input intercept point (IIP3) formula incorporating the drain induced non-linearity in terms of the voltage gain is presented. The proposed models are validated with DC and AC simulations and measurements.
Naci Pekcokguler, Hung-Chi Han, Dominique Morche, Catherine Dehollain, Andreas Peter Burg, Christian C. Enz
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 A Generalized Adjusted Min-Sum Decoder for 5G LDPC Codes: Algorithm and Implementation
abstract
5G New Radio (NR) has stringent demands on both performance and complexity for the design of low-density parity-check (LDPC) decoding algorithms and corresponding VLSI implementations. Furthermore, decoders must fully support the wide range of all 5G NR blocklengths and code rates, which is a significant challenge. In this paper, we present a high-performance and low-complexity LDPC decoder, tailor-made to fulfill the 5G requirements. First, to close the gap between belief propagation (BP) decoding and its approximations in hardware, we propose an extension of adjusted min-sum decoding, called generalized adjusted min-sum (GA-MS) decoding. This decoding algorithm flexibly truncates the incoming messages at the check node level and carefully approximates the non-linear functions of BP decoding to balance the error-rate and hardware complexity. Numerical results demonstrate that the proposed fixed-point GA-MS has only a minor gap of 0.1 dB compared to floating-point BP under various scenarios of 5G standard specifications. Secondly, we present a fully reconfigurable 5G NR LDPC decoder implementation based on GA-MS decoding. Given that memory occupies a substantial portion of the decoder area, we adopt multiple data compression and approximation techniques to reduce 42.2% of the memory overhead. The corresponding 28nm FD-SOI ASIC decoder has a core area of 1.823 mm$^{2}$and operates at 895 MHz. It is compatible with all 5G NR LDPC codes and achieves a peak throughput of 24.42 Gbps and a maximum area efficiency of 13.40 Gbps/mm$^{2}$at 4 decoding iterations.
Yuqing Ren, Yifei Shen 0003, Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 A Node-Based Polar List Decoder With Frame Interleaving and Ensemble Decoding Support
abstract
Node-based successive cancellation list (SCL) decoding has received considerable attention in wireless communications for its significant reduction in decoding latency, particularly with 5G New Radio (NR) polar codes. However, the existing node-based SCL decoders are constrained by sequential processing, leading to complicated and data-dependent computational units that introduce unavoidable stalls, reducing hardware efficiency. In this paper, we present a frame-interleaving hardware architecture for a generalized node-based SCL decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we further exploit graph ensembles to diversify the decoding space, thus enhancing the error-correcting performance with a limited list size. Two dynamic strategies are proposed to eliminate the residual stalls in the decoding schedule, which eventually results in nearly$2 \times $throughput compared to the state-of-the-art baseline node-based SCL decoder. To impart the decoder rate flexibility, we develop a novel online instruction generator to identify the generalized nodes and produce instructions on-the-fly. The corresponding 28nm FD-SOI ASIC SCL decoder with a list size of 8 has a core area of 1.28 mm2 and operates at 692 MHz. It is compatible with all 5G NR polar codes and achieves a throughput of 3.34 Gbps and an area efficiency of 2.62 Gbps/mm2 for uplink (1024, 512) codes, which is$1.41 \times $and$1.69 \times $better than the state-of-the-art node-based SCL decoders.
Yuqing Ren, Leyu Zhang, Ludovic Damien Blanc, Yifei Shen 0003, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 Single-Anchor UWB Localization Using Channel Impulse Response Distributions
abstract
Ultra-wideband (UWB) devices are widely used in indoor localization scenarios. Single-anchor UWB localization shows advantages because of its simple system setup compared to conventional two-way ranging (TWR) and trilateration localization methods. In this work, we focus on single-anchor UWB localization methods that learn statistical features of the channel impulse response (CIR) in different location areas using a Gaussian mixture model (GMM). We show that by learning the joint distributions of the amplitudes of different delay components, we achieve a more accurate location estimate compared to considering each delay bin independently. Moreover, we develop a similarity metric between sets of CIRs. With this set-based similarity metric, we can further improve the estimation performance, compared to treating each snapshot separately. We showcase the advantages of the proposed methods in multiple application scenarios.
Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ICASSP3
2023 Improved Belief Propagation Decoding of Turbo Codes
abstract
Turbo codes have been successfully adopted in 4G LTE, which can approach the channel capacity with Bahl-Cocke-Jelinek-Raviv (BCJR) decoding. With the evolution from 4G LTE to 5G NR, there is a demand to design a unified channel decoder that supports both LTE Turbo codes and NR low-density parity-check (LDPC) codes. One solution is to employ belief propagation (BP) decoding on the bipartite Tanner graph for both codes. However, although MacKay pointed out that Turbo codes have a sparse parity-check matrix, the existence of 4-cycles in such a matrix severely deteriorates the performance of BP decoding. In this paper, we propose two polynomial-based methods to optimize the parity-check matrix of Turbo codes by improving the sparsity while also removing 4-cycles and even 6-cycles compared to the original matrix. Simulation results show that the improved BP decoding for Turbo codes halves the error-correction performance gap between the original BP decoding and BCJR decoding, which is a promising step towards the unified channel decoder design based on the BP algorithm.
Yifei Shen 0003, Yuqing Ren, Andreas Toftegaard Kristensen, Xiaohu You 0001, Chuan Zhang 0001, Andreas Peter Burg
ICASSP6
2023 Spreading Factor assisted LoRa Localization with Deep Reinforcement Learning
abstract
Most of the developed localization solutions rely on RSSI fingerprinting. However, in the LoRa networks, due to the spreading factor (SF) in the network setting, traditional fingerprinting may lack representativeness of the radio map, leading to inaccurate position estimates. As such, in this work, we propose a novel LoRa RSSI fingerprinting approach that takes into account the SF. The performance evaluation shows the prominence of our proposed approach since we achieved an improvement in localization accuracy by up to 6.67% compared to the state-of-the-art methods. The evaluation has been done using a fully connected deep neural network (DNN) set as the baseline. To further improve the localization accuracy, we propose a deep reinforcement learning model that captures the ever-growing complexity of LoRa networks and copes with their scalability. The obtained results show an improvement of 48.10% in the localization accuracy compared to the baseline DNN model.
Yaya Etiabi, Mohammed Jouhari, Andreas Peter Burg, El Mehdi Amhoud
VTC2023-Spring3
2023 An Ultra-Low-Power Widely-Tunable Complex Band-Pass Filter for RF Spectrum Sensing
abstract
Power consumption is of utmost importance in portable devices as they are operated from a limited energy supply. Wireless radio is one of the most power-hungry blocks in these systems and needs to be activated opportunistically. Radio frequency (RF) spectrum sensing can be performed with a dedicated low power radio to control the main wireless radio. A low power receiver architecture is proposed in this work to monitor the 2.4GHz Industrial, Scientific, and Medical (ISM) band for communication standards such as Wireless Local Area Network (WLAN), Zonal Intercommunication Global-standard (ZigBee), and Bluetooth-Low-Energy (BLE). The whole ISM band is down-converted to zero-intermediate-frequency (ZIF) and a tunable base-band complex band-pass filter (CBPF) is used to scan the spectrum. A widely tunable first order ultra-low-power transconductor-capacitor (Gm-C) CBPF architecture was designed and fabricated in GF 22FDX technology, which occupies 0.0049mm 2 area. A frequency shift of ±60MHz was achieved with 5-40MHz bandwidth range. Power consumption ranges from 6.7$\mu \text{W}$to 99.2$\mu \text{W}$with the best and the worst figure-of-merit of 0.034fJ/pole and 0.082fJ/pole that is the energy consumption per pole normalized by the spurious free dynamic range.
Naci Pekcokguler, Dominique Morche, Andreas Peter Burg, Catherine Dehollain
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Increasing Cellular Network Energy Efficiency for Railway Corridors
abstract
Modern trains act as Faraday cages making it chal-lenging to provide high cellular data capacities to passengers. A solution is the deployment of linear cells along railway tracks, forming a cellular corridor. To provide a sufficiently high data capacity, many cell sites need to be installed at regular distances. However, such cellular corridors with high power sites in short distance intervals are not sustainable due to the infrastructure power consumption. To render railway connectivity more sustainable, we propose to deploy fewer high-power radio units with intermediate low-power support repeater nodes. We show that these repeaters consume only 5 % of the energy of a regular cell site and help to maintain the same data capacity in the trains. In a further step, we introduce a sleep mode for the repeater nodes that enables autonomous solar powering and even eases installation because no cables to the relays are needed.
Adrian Schumacher, Ruben Merz, Andreas Peter Burg
DATE3
2022 Fast Sequence Repetition Node-Based Successive Cancellation List Decoding for Polar Codes
abstract
Compared with the bit-wise successive cancellation list (SCL) decoding of polar codes, the node-based Fast SCL decoding significantly reduces the decoding latency by identifying special constituent codes and decoding these in parallel. To further reduce the latency of current Fast SCL decoders, we first propose a fast sequence repetition (SR) node-based SCL (Fast SR-SCL) decoding algorithm, which only involves one type of node in the SCL decoding tree. Furthermore, we employ the adaptive path splitting (APS) strategy to terminate the path splitting in the SR node early, without degrading the error-correcting performance. Numerical results show that for 5G uplink codes with a length of 1024 and rates of 1/4, 1/2, and 3/4, our decoder can deliver the same decoding performance while reducing the average latency by 34.5%, 38.0%, and 39.6% compared with the state-of-the-art Fast SCL decoder for a list size L = 8.
Yifei Shen 0003, Yuqing Ren, Andreas Toftegaard Kristensen, Alexios Balatsoukas-Stimming, Xiaohu You 0001, Chuan Zhang 0001, Andreas Peter Burg
ICC7
2022 Beam Selection and Tracking for Amplify-and-Forward Repeaters
abstract
Mobile network operators constantly have to upgrade their cellular network to satisfy the public’s hunger for increasing data capacity. However, regulatory limits regarding allowed electromagnetic field strength on existing cell sites often limit or prevent the installation of new cells on additional frequency bands. Further densifying the radio access infrastructure by means of new cell sites takes time because of the site acquisition, the need for permissions, and the civil construction. To this end, amplify-and-forward repeaters are a cost-efficient method to provide mobile wireless data capacity to concealed areas outdoors, inside buildings, or vehicles, where low signal levels drastically limit the achievable capacity or even prevent communication with distant base stations. Repeaters with high-gain beamforming antennas can be used to decrease the path loss and thereby offer a higher capacity than user equipment otherwise could achieve. Since those highly directional beams must be constantly adjusted, we propose methods to align and track the beam even without having access to in-band beam control mechanisms. Numerical analyses with measurement data show that such non-invasive (system data-agnostic) methods are feasible and only incur a few percent in throughput reduction.
Adrian Schumacher, Ruben Merz, Andreas Peter Burg
VTC Spring3
2022 A Maximum-Likelihood-Based Two-User Receiver for LoRa Chirp Spread-Spectrum Modulation
abstract
Long Range (LoRa) is an emerging low-power wide-area network technology offering long-range wireless connectivity to Internet of Things (IoT) devices. For energy efficiency reasons, LoRa end nodes implement a nonslotted ALOHA multiple access scheme to transmit packets to the gateway. Due to the lack of synchronization between end nodes, collisions between uplink packets have been identified as the main obstacle to the scaling of dense LoRa networks. To tackle this issue, we present in this article a LoRa receiver that is capable of decoding colliding packets from two interfering end nodes. The proposed two-user detector is derived from the maximum-likelihood principle using a detailed model of two colliding LoRa packets. As the complexity of the maximum-likelihood sequence estimation is prohibitive, complexity-reduction techniques are introduced to enable practical implementations of the receiver. An in-depth performance analysis highlights that the proposed two-user detector inherently leverages the differences in received power, time offsets, and frequency offsets between the users to separate and demodulate their respective signals. To demonstrate the practicality of the proposed detector, an interference-robust synchronization algorithm is then designed and evaluated. Simulation results indicate that a LoRa receiver combining the proposed synchronization algorithm and two-user detector is capable of detecting and demodulating two interfering users with satisfactorily error rates.
Mathieu Xhonneux, Joachim Tapparel, Alexios Balatsoukas-Stimming, Andreas Peter Burg, Orion Afisiadis
IEEE Internet Things J.4
2021 Dynamic Range and Complexity Optimization of Mixed-Signal Machine Learning Systems
abstract
Audio processing had been in demand throughout the electronic era. Recent advances in neural networks increased the demand on audio processing for speech recognition applications. In this work, a rigorous study on the dynamic range and system complexity optimization is presented for a mixed-signal keyword spotting system. The proposed system consists of an analog feature extractor and a neural network based keyword classifier. The results showed that with the proposed method, more than an order of magnitude power saving can be achieved in the analog feature extraction compared to the digital state-of-the-art counterpart.
Naci Pekcokguler, Dominique Morche, Adrian Frischknecht, Christoph Gerum, Andreas Peter Burg, Catherine Dehollain
ISCAS5
2021 On the Advantage of Coherent LoRa Detection in the Presence of Interference
abstract
It has been shown that the coherent detection of long range (LoRa) signals only provides marginal gains of around 0.7 dB on the additive white Gaussian noise (AWGN) channel. However, ALOHA-based massive Internet-of-Things systems, including LoRa, often operate in the interference-limited regime. Therefore, in this work, we examine the performance of the LoRa modulation with coherent detection in the presence of interference from another LoRa user with the same spreading factor. We derive rigorous symbol- and frame error rate (FER) expressions as well as bounds and approximations for evaluating the error rates. The error rates predicted by these approximations are compared against error rates found by Monte Carlo simulations and shown to be very accurate. We also compare the performance of LoRa with coherent and noncoherent receivers and we show that the coherent detection of LoRa is significantly more beneficial in interference scenarios than in the presence of only AWGN. For example, we show that coherent detection leads to a 2.5-dB gain over the standard noncoherent detection for a signal-to-interference ratio (SIR) of 3 dB and up to a 10-dB gain for an SIR of 0 dB. Moreover, we show that with coherent detection it is easier to obtain useful and relevant FER values even for negative SIR values.
Orion Afisiadis, Sitian Li, Joachim Tapparel, Andreas Peter Burg, Alexios Balatsoukas-Stimming
IEEE Internet Things J.4
2021 E2CNNs: Ensembles of Convolutional Neural Networks to Improve Robustness Against Memory Errors in Edge-Computing Devices
abstract
To reduce energy consumption, it is possible to operate embedded systems at sub-nominal conditions (e.g., reduced voltage, limited eDRAM refresh rate) that can introduce bit errors in their memories. These errors can affect the stored values of convolutional neural network (CNN) weights and activations, compromising their accuracy. In this article, we introduce Embedded Ensemble CNNs (E2CNNs), our architectural design methodology to conceive ensembles of convolutional neural networks to improve robustness against memory errors compared to a single-instance network. Ensembles of CNNs have been previously proposed to increase accuracy at the cost of replicating similar or different architectures. Unfortunately, state-of-the-art (SoA) ensembles do not suit well embedded systems, in which memory and processing constraints limit the number of deployable models. Our proposed architecture solves that limitation applying SoA compression methods to produce an ensemble with the same memory requirements of the original architecture, but with improved error robustness. Then, as part of our new E2CNNs design methodology, we propose a heuristic method to automate the design of the voter-based ensemble architecture that maximizes accuracy for the expected memory error rate while bounding the design effort. To evaluate the robustness of E2CNNs for different error types and densities, and their ability to achieve energy savings, we propose three error models that simulate the behavior of SRAM and eDRAM operating at sub-nominal conditions. Our results show that E2CNNs achieves energy savings of up to 80 percent for LeNet-5, 90 percent for AlexNet, 60 percent for GoogLeNet, 60 percent for MobileNet and 60 percent for an optimized industrial CNN, while minimizing the impact on accuracy. Furthermore, the memory size can be decreased up to 54 percent by reducing the number of members in the ensemble, with a more limited impact on the original accuracy than obtained through pruning alone.
Flavio Ponzina, Miguel Peón-Quirós, Andreas Peter Burg, David Atienza 0001
IEEE Trans. Computers3
2020 Lupulus: A Flexible Hardware Accelerator for Neural Networks
abstract
Neural networks have become indispensable for a wide range of applications, but they suffer from high computationaland memory-requirements, requiring optimizations from the algorithmic description of the network to the hardware implementation. Moreover, the high rate of innovation in machine learning makes it important that hardware implementations provide a high level of programmability to support current and future requirements of neural networks. In this work, we present a flexible hardware accelerator for neural networks, called Lupulus, supporting various methods for scheduling and mapping of operations onto the accelerator. Lupulus was implemented in a 28nm FD-SOI technology and demonstrates a peak performance of 380GOPS/GHz with latencies of 21.4ms and 183.6ms for the convolutional layers of AlexNet and VGG-16, respectively.
Andreas Toftegaard Kristensen, Robert Giterman, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ICASSP4
2020 Coded LoRa Frame Error Rate Analysis
abstract
In this work, we study the coded frame error rate (FER) of LoRa under additive white Gaussian noise (AWGN) and under carrier frequency offset (CFO). To this end, we use existing approximations for the bit error rate (BER) of the LoRa modulation under AWGN and we present a FER analysis that includes the channel coding, interleaving, and Gray mapping of the LoRa physical layer. We also derive the LoRa BER under carrier frequency offset and we present a corresponding FER analysis. We compare the derived frame error rate expressions to Monte Carlo simulations to verify their accuracy.
Orion Afisiadis, Andreas Peter Burg, Alexios Balatsoukas-Stimming
ICC2
2020 Gain-Cell Embedded DRAMs: Modeling and Design Space
abstract
Among the different types of DRAMs, gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power and CMOS-compatible alternative to conventional SRAM. GC-eDRAM achieves high memory density as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, prove to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this paper, we present GEMTOO, the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture and it enables the evaluation of architectural transformations as well as of advanced transistor-level effects, such as the increase of the access delay due to deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from post-layout simulations in a 28nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in 28nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs and, based on the results, optimal design practices are derived.
Andrea Bonetti, Roman Golman, Robert Giterman, Adam Teman, Andreas Peter Burg
ISCAS5
2020 GC-eDRAM with Body-Bias Compensated Readout and Error Detection in 28nm FD-SOI
abstract
Gain-cell embedded DRAM (GC-eDRAM) is an attractive alternative to conventional SRAM due to its high-density, low-leakage, and inherent two-ported functionality. However, its dynamic storage mechanism requires power-hungry refresh cycles to maintain data. This problem is aggravated due to the impact of Process-Voltage-Temperature (PVT) variations at deeply-scaled technology nodes and low voltages. In this paper, we present a GC-eDRAM with body-bias compensated readout, which is dynamically configured to extend the data retention time (DRT) of the memory under varying operating conditions. The proposed GC-eDRAM exploits the body-biasing capabilities of FD-SOI technology to adjust the switching threshold of the sense inverter under PVT variations. An additional, unbiased, sense inverter is added to provide a dual-sampling mechanism to the readout path, enabling error detection to further reduce design guard bands. An 8 kb GC-eDRAM with integrated body-bias compensated readout and error detection was implemented in 28 nm FD-SOI technology. Silicon measurements of the manufactured array demonstrate up-to 75% DRT improvement and up-to 86% energy savings under PVT and frequency variations compared to a conventional guard banded memory design.
Robert Giterman, Andrea Bonetti, Andreas Peter Burg, Adam Teman
ISCAS3
2020 Complexity-efficient Fano Decoding of Polarization-adjusted Convolutional (PAC) Codes
Mohammad Rowshan, Andreas Peter Burg, Emanuele Viterbo
ISITA2
2020 A mmWave Bridge Concept to Solve the Cellular Outdoor-to-Indoor Challenge
abstract
Wireless indoor coverage and data capacity are important aspects of cellular networks. With the ever-increasing data traffic, demand for more data capacity indoors is also growing. The lower frequencies of the legacy frequency bands of macro outdoor cells manage to provide coverage inside buildings, however, new frequencies foreseen for the 5th generation (5G) of mobile communications in the millimeter wave (mmWave) spectrum penetrate very poorly into buildings. Therefore, a massive densification of the network would require to deploy a large number of indoor small cells, which would lead to high deployment costs to install the necessary wired/optical backhaul. Hence, other methods are needed that allow an increase of the data capacity indoors, bearing a lower cost than a fiber deployment. We propose a cost-efficient out-of-band repeater architecture that provides more data capacity indoors than an outdoor macro/micro network can provide to indoor, without adversely affecting a legacy network, and which readily works with the established cellular infrastructure as well as standard handsets/smartphones. This proposal is compared to conventional in- and out-of-band repeaters and relay nodes in order to highlight the advantages of our solution. While the data capacity for a single link is similar to that of repeaters and relays, a macro cell can be effectively offloaded. Cell capacities corresponding to at least 3-4 times that of a repeater or relay solution can be provided, depending on the number of parallel installed links and the bandwidth in the mmWave spectrum.
Adrian Schumacher, Ruben Merz, Andreas Peter Burg
VTC Spring3
2020 Design and Decoding of Irregular LDPC Codes Based on Discrete Message Passing
abstract
We consider discrete message passing (MP) decoding of low-density parity check (LDPC) codes based on information-optimal symmetric look-up table (LUT). A link between discrete message labels and the associated log-likelihood ratio values (defined in terms of density evolution distributions) is established. This link gives rise to an algebraic structure on the message labels and leads to an interpretation of LUT decoding as a form of quantized belief propagation. We then exploit the algebraic structure for low-complexity LUT decoder designs. Our LUT decoding framework is the first to also apply to irregular LDPC codes by taking into account the degree distribution in a joint LUT design. We exploit the relation between LUT decoding and belief propagation to obtain stability conditions and irregular LDPC code designs optimized for LUT decoding. The resulting decoders outperform floating-point precision min-sum decoders at LUT resolutions as low as 3 bit s for regular codes and 4 bits for irregular codes.
Michael Meidlinger, Gerald Matz, Andreas Peter Burg
IEEE Trans. Commun.3
2020 Gain-Cell Embedded DRAMs: Modeling and Design Space
abstract
Among the different types of dynamic random-access memories (DRAMs), gain-cell embedded DRAM (GC-eDRAM) is a compact, low-power, and CMOS-compatible alternative to conventional static random-access memory (SRAM). GC-eDRAM achieves high memory density, as it relies on a storage cell that can be implemented with as few as two transistors and that can be fabricated without additional process steps. However, since the performance of GC-eDRAMs relies on many interdependent variables, the optimization of the performance of these memories for the integration into their hosting system, as well as the design investigation of future GC-eDRAMs, proves to be highly complex tasks. In this context, modeling tools of memories are key enablers for the exploration of this large design space in a short amount of time. In this article, we present GC-eDRAM modeling tool (GEMTOO), the first modeling tool that estimates timing, memory availability, bandwidth, and area of GC-eDRAMs. The tool considers parameters related to technology, circuits, and memory architecture, and it enables the evaluation of architectural transformations as well as advanced transistor-level effects, such as the increase in the access delay due to the deterioration of the stored data. The timing is estimated with a maximum deviation of 15% from postlayout simulations in a 28-nm FD-SOI technology for different memory sizes and architectures. Moreover, the measured random cycle frequency of a GC-eDRAM fabricated in a 28-nm CMOS bulk process is estimated with a 9% deviation when considering 6-sigma random process variations of the bitcells. The proposed GEMTOO modeling tool is used to show the intricacies in design optimization of GC-eDRAMs, and based on the results, optimal design practices are derived.
Andrea Bonetti, Roman Golman, Robert Giterman, Adam Teman, Andreas Peter Burg
IEEE Trans. Very Large Scale Integr. Syst.5
2020 On the Error Rate of the LoRa Modulation With Interference
abstract
LoRa is a chirp spread-spectrum modulation developed for the Internet of Things (IoT). In this work, we examine the performance of LoRa in the presence of both additive white Gaussian noise and interference from another LoRa user. To this end, we extend an existing interference model, which assumes perfect alignment of the signal of interest and the interference, to the more realistic case where the interfering user is neither chip- nor phase-aligned with the signal of interest and we derive an expression for the error rate. We show that the existing aligned interference model overestimates the effect of interference on the error rate. Moreover, we prove two symmetries in the interfering signal and we derive low-complexity approximate formulas that can significantly reduce the complexity of computing the symbol and frame error rates compared to the complete expression. Finally, we provide numerical simulations to corroborate the theoretical analysis and to verify the accuracy of our proposed approximations.
Orion Afisiadis, Matthieu Cotting, Andreas Peter Burg, Alexios Balatsoukas-Stimming
IEEE Trans. Wirel. Commun.3
2019 FPGA-Based Emulation of Embedded DRAMs for Statistical Error Resilience Evaluation of Approximate Computing Systems
abstract
Embedded DRAM (eDRAM) requires frequent power-hungry refresh according to the worst-case retention time across PVT variations to avoid data loss. Abandoning the error-free paradigm, by choosing sub-critical refresh rates that gracefully degrade the eDRAM content, unlocks considerable power-saving opportunities, but requires to understand the effect of stochastic memory errors at the system/application level. We propose an FPGA-based platform featuring faulty eDRAM emulation based on advanced retention time models and silicon measurements for statistical error resilience evaluation of applications in a complete embedded system. We analyze the statistical QoS for various benchmarks under different sub-critical refresh rates and retention time distributions.
Marco Widmer, Andrea Bonetti, Andreas Peter Burg
DAC3
2019 Lora Digital Receiver Analysis and Implementation
abstract
Low power wide area network technologies (LPWANs) are attracting attention because they fulfill the need for long range low power communication for the Internet of Things. LoRa is one of the proprietary LPWAN physical layer (PHY) technologies, which provides variable data-rate and long range by using chirp spread spectrum modulation. This paper describes the basic LoRa PHY receiver algorithms and studies their performance. The LoRa PHY is first introduced and different demodulation schemes are proposed. The effect of carrier frequency offset and sampling frequency offset are then modeled and corresponding compensation methods are proposed. Finally, a software-defined radio implementation for the LoRa transceiver is briefly presented.
Reza Ghanaatian, Orion Afisiadis, Matthieu Cotting, Andreas Peter Burg
ICASSP4
2019 Data-Retention-Time Characterization of Gain-Cell eDRAMs Across the Design and Variations Space
abstract
The rise of data-intensive applications has increased the demand for high-density and low-power embedded memories. Among them, the gain-cell embedded DRAM (GC-eDRAM) is a suitable alternative to the static random access memory (SRAM) due to its high memory density and low leakage current. However, as the GC-eDRAM dynamically stores data, its memory content has to be periodically refreshed according to the data retention time (DRT). Even though different DRT characterization methodologies have been reported in the literature, a practical and accurate method to quantify the DRT across Monte Carlo (MC) runs to evaluate the impact of local process variations (LPVs) has not been proposed yet. Thus, the minimum memory refresh rate is generally estimated with large design guard bands to avoid any loss of data, at the expense of a higher power consumption and less memory bandwidth. In this work, we present a current-based DRT characterization methodology that enables an accurate LPV analysis without the need of a large number of costly electronic design automation (EDA) software licenses. The presented approach is compared with other DRT characterization methodologies for both accuracy as well as practical aspects. Furthermore, the DRT of a 3-transistor (3T) gain cell (GC) designed in 28 nm FD-SOI process technology is measured for different design choices, global and local variations. The analysis of the results shows that LPVs have the most degrading effect on the DRT and therefore that the proposed approach is key for either the design of GC-eDRAMs or the choice of their refresh rate to avoid the need for overly pessimistic worst-case margins.
Ester Vicario Bravo, Andrea Bonetti, Andreas Peter Burg
ISCAS3
2019 Feedback-Aware Precoding for Millimeter Wave Massive MIMO Systems
abstract
Millimeter wave (mmWave) communication is a promising solution for coping with the ever-increasing mobile data traffic because of its large bandwidth. To enable a suffi-cient link margin, a large antenna array employing directional beamforming, which is enabled by the availability of channel state information at the transmitter (CSIT), is required. However, CSIT acquisition for mmWave channels introduces a huge feedback overhead due to the typically large number of transmit and receive antennas. Leveraging properties of mmWave channels, this paper proposes a precoding strategy which enables a flexible adjustment of the feedback overhead. In particular, the optimal unconstrained precoder is approximated by selecting a variable number of elements from a basis that is constructed as a function of the transmitter array response, where the number of selected basis elements can be chosen according to the feedback constraint. Simulation results show that the proposed precoding scheme can provide a near-optimal solution if a higher feedback overhead can be afforded. For a low overhead, it can still provide a good approximation of the optimal precoder.
Reza Ghanaatian, Vahid Jamali, Andreas Peter Burg, Robert Schober
PIMRC3
2019 3.5 GHz Coverage Assessment with a 5G Testbed
abstract
Today, cellular networks have saturated frequencies below 3 GHz. Because of increasing capacity requirements, 5th generation (5G) mobile networks target the 3.5 GHz band (3.4 to 3.8 GHz). Despite its expected wide usage, there is little empirical path loss data and mobile radio network planning experience for the 3.5 GHz band available. This paper presents the results of rural, suburban, and urban measurement campaigns using a pre-standard 5G prototype testbed operating at 3.5 GHz, with outdoor as well as outdoor-to-indoor scenarios. Based on the measurement results, path loss models are evaluated, which are essential for network planning.
Adrian Schumacher, Ruben Merz, Andreas Peter Burg
VTC Spring3
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.5
2018 A Timing-Monitoring Sequential for Forward and Backward Error-Detection in 28 nm FD-SOI
abstract
The increasing impact of variability on near-threshold nanometer circuits calls for a tighter online monitoring and control of the available timing margins. Error-detection sequentials are widely used together with error-correction techniques to operate digital designs with such carefully controlled far-below-worst-case margins, ensuring their correct operation even in the presence of uncertainties and variations. However, these registers are often designed only to either detect setup timing violations or to measure the available positive timing slack for a small detection-window. In this paper we propose a timing-monitoring sequential that provides both timing-monitoring modes, which can be selected at run-time depending on the desired timing-monitoring strategy. As the detection window of the presented circuit depends on the duty-cycle of the clock, either slow paths or fast paths can be monitored and measured with wide timing windows. The performance of this timing-monitoring sequential is evaluated in a 28 nm FD-SOI process with post-layout simulations which show that the circuit is able to monitor a positive timing slack as small as 140 ps or to measure a path delay as fast as 50 ps. The proposed circuit is applied to a digital multiplier that was fabricated in a test chip and measurements show that the timing-monitoring sequentials are able to measure the critical path of the multiplier with a 1% accuracy and without incurring any timing violation.
Andrea Bonetti, Jeremy Constantin, Adam Ternan, Andreas Peter Burg
ISCAS4
2018 Wireless Communication and Security Issues for Cyber-Physical Systems and the Internet-of-Things
abstract
Wireless sensors and actuators connected by the Internet-of-Things (IoT) are central to the design of advanced cyber-physical systems (CPSs). In such complex, heterogeneous systems, communication links must meet stringent requirements on throughput, latency, and range, while adhering to tight energy budget and providing high levels of security. In this paper, we first summarize wireless communication principles from the perspective of the connectivity needs of IoT and CPS. Based on these principles, we then review the most relevant wireless communication standards before focusing on the key security issues and features of such systems. In particular, the gap between the security features in the communication standards used in CPSs and IoT and their actual vulnerabilities are pointed out with practical examples and recent attacks. We emphasize the need for a more in-depth study of the security issues across all the protocol layers, including both logical layer security and physical layer security.
Andreas Peter Burg, Anupam Chattopadhyay, Kwok-Yan Lam
Proc. IEEE1
2018 Faulty Successive Cancellation Decoding of Polar Codes for the Binary Erasure Channel
abstract
In this paper, the faulty successive cancellation decoding of polar codes for the binary erasure channel is studied. To this end, a simple erasure-based fault model is introduced to represent errors in the decoder, and it is shown that, under this model, polarization does not happen, meaning that fully reliable communication is not possible at any rate. Furthermore, a lower bound on the frame error rate of polar codes under faulty successive cancellation decoding is provided, which is then used, along with a well-known upper bound, in order to choose a blocklength that minimizes the erasure probability under faulty decoding. Finally, an unequal error protection scheme that can re-enable asymptotically erasure-free transmission at a small rate loss and by protecting only a constant fraction of the decoder is proposed. The same scheme is also shown to significantly improve the finite-length performance of the faulty successive cancellation decoder by protecting as little as 1.5% of the decoder.
Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Commun.2
2018 A 588-Gb/s LDPC Decoder Based on Finite-Alphabet Message Passing
abstract
An ultrahigh throughput low-density parity-check (LDPC) decoder with an unrolled full-parallel architecture is proposed, which achieves the highest decoding throughput compared to previously reported LDPC decoders in the literature. The decoder benefits from a serial message-transfer approach between the decoding stages to alleviate the well-known routing congestion problem in parallel LDPC decoders. Furthermore, a finite-alphabet message passing algorithm is employed to replace the VN update rule of the standard min-sum (MS) decoder with lookup tables, which are designed in a way that maximizes the mutual information between decoding messages. The proposed algorithm results in an architecture with reduced bit-width messages, leading to a significantly higher decoding throughput and to a lower area compared to an MS decoder when serial message transfer is used. The architecture is placed and routed for the standard MS reference decoder and for the proposed finite-alphabet decoder using a custom pseudo-hierarchical backend design strategy to further alleviate routing congestions and to handle the large design. Postlayout results show that the finite-alphabet decoder with the serial message-transfer architecture achieves a throughput as large as 588 Gb/s with an area of 16.2 mm2and dissipates an average power of 22.7 pJ per decoded bit in a 28-nm fully depleted silicon on isulator library. Compared to the reference MS decoder, this corresponds to 3.1 times smaller area and 2 times better energy efficiency.
Reza Ghanaatian, Alexios Balatsoukas-Stimming, Thomas Christoph Müller, Michael Meidlinger, Gerald Matz, Adam Teman, Andreas Peter Burg
IEEE Trans. Very Large Scale Integr. Syst.7
2017 Automated Integration of Dual-Edge Clocking for Low-Power Operation in Nanometer Nodes
abstract
Clocking power, including both clock distribution and registers, has long been one of the primary factors in the total power consumption of many digital systems. One straightforward approach to reduce this power consumption is to apply dual-edge-triggered (DET) clocking, as sequential elements operate at half the clock frequency while maintaining the same throughput as with conventional single-edge-triggered (SET) clocking. However, the DET approach is rarely taken in modern integrated circuits, primarily due to the perceived complexity of integrating such a clocking scheme. In this article, we first identify the most promising conditions for achieving low-power operation with DET clocking and then introduce a fully automated design flow for applying DET to a conventional SET design. The proposed design flow is demonstrated on three benchmark circuits in a 40nm CMOS technology, providing as much as a 50% reduction in clock distribution and register power consumption.
Andrea Bonetti, Nicholas Preyss, Adam Teman, Andreas Peter Burg
ACM Trans. Design Autom. Electr. Syst.4
2016 Statistical fault injection for impact-evaluation of timing errors on application performance
abstract
This paper proposes a novel approach to modeling of gate level timing errors during high-level instruction set simulation. In contrast to conventional, purely random fault injection, our physically motivated approach directly relates to the underlying circuit structure, hence allowing for a significantly more detailed characterization of application performance under scaled frequency / voltage (including supply noise). The model uses gate level timing statistics extracted by dynamic timing analysis from the post place & route netlist of a general-purpose processor to perform instruction-aware fault injections. We employ a 28 nm OpenRISC core as a case study, to demonstrate how statistical fault injection provides a more accurate and realistic analysis of power vs. error performance.
Jeremy Constantin, Andreas Peter Burg, Zheng Wang 0020, Anupam Chattopadhyay, Georgios Karakonstantis
DAC2
2016 Energy vs. reliability trade-offs exploration in biomedical ultra-low power devices
Loris Duch, Pablo García Del Valle, Shrikanth Ganapathy, Andreas Peter Burg, David Atienza 0001
DATE4
2016 A low-power correlator for wakeup receivers with algorithm pruning through early termination
abstract
A low-complexity, low-power digital correlator for wakeup receivers is presented. With the proposed algorithm, unnecessary computational cycles are dynamically pruned from the correlation using an early threshold check. For the algorithm, we provide a rigorous mathematical analysis for the associated complexity/performance trade-offs. Furthermore, a low overhead hardware architecture with early-termination capability is developed and implemented in a 0.18μm CMOS technology. The post layout power analysis shows that the presented architecture can reduce power by up to 32% when compared to the conventional architecture with negligible degradation in detection probability and without degradation in false-alarm probability.
Reza Ghanaatian, Paul N. Whatmough, Jeremy Constantin, Adam Teman, Andreas Peter Burg
ISCAS5
2016 Hardware decoders for polar codes: An overview
abstract
Polar codes are an exciting new class of error correcting codes that achieve the symmetric capacity of memoryless channels. Many decoding algorithms were developed and implemented, addressing various application requirements: from error-correction performance rivaling that of LDPC codes to very high throughput or low-complexity decoders. In this work, we review the state of the art in polar decoders implementing the successive-cancellation, belief propagation, and list decoding algorithms, illustrating their advantages.
Pascal Giard, Gabi Sarkis, Alexios Balatsoukas-Stimming, YouZhe Fan, Chi-Ying Tsui, Andreas Peter Burg, Claude Thibeault, Warren J. Gross
ISCAS6
2016 A process compensated gain cell embedded-DRAM for ultra-low-power variation-aware design
abstract
Gain cell embedded DRAM (GC-eDRAM) is a high-density alternative to SRAM for ultra-low-power systems. However, due to its dynamic nature, GC-eDRAM requires power-hungry refresh cycles to ensure data retention. Traditional design approaches dictate configuration of the refresh rate according to the worst bitcell, when biased at low-probability, worst-case conditions. However, due to the process variations and local mismatch that can significantly deteriorate the data retention time of a GC-eDRAM bitcell, this design approach often leads to a large power overhead. In this paper, we present a novel GC-eDRAM architecture, incorporating several techniques for variation-aware operation. The primary feature of this architecture is an improved replica scheme for process compensated access tracking that enables calibration for process variations and adaptive refresh according to the array access statistics. The array is shown to ensure data integrity, providing as much as a 7x reduction in retention power over worst-case refresh-rate design for 20% write activity.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Alexander Fish, Andreas Peter Burg
ISCAS5
2016 Sliding Window Spectrum Sensing for Full-Duplex Cognitive Radios with Low Access-Latency
abstract
In a cognitive radio system the failure of secondary user (SU) transceivers to promptly vacate the channel can introduce significant access-latency for primary or high-priority users (PU). In conventional cognitive radio systems, the backoff latency is exacerbated by frame structures that only allow sensing at periodic intervals. Concurrent transmission and sensing using self-interference suppression has been suggested to improve the performance of cognitive radio systems, allowing decisions to be taken at multiple points within the frame. In this paper, we extend this approach by proposing a sliding-window full-duplex model allowing decisions to be taken on a sample-by-sample basis. We also derive the access-latency for both the existing and the proposed schemes. Our results show that the access-latency of the sliding scheme is decreased by a factor of 2.6 compared to the existing slotted full-duplex scheme and by a factor of approximately 16 compared to a half-duplex cognitive radio system. Moreover, the proposed scheme is significantly more resilient to the destructive effects of residual self-interference compared to previous approaches.
Orion Afisiadis, Andrew Austin 0001, Alexios Balatsoukas-Stimming, Andreas Peter Burg
VTC Spring4
2016 Power, Area, and Performance Optimization of Standard Cell Memory Arrays Through Controlled Placement
abstract
Embedded memory remains a major bottleneck in current integrated circuit design in terms of silicon area, power dissipation, and performance; however, static random access memories (SRAMs) are almost exclusively supplied by a small number of vendors through memory generators, targeted at rather generic design specifications. As an alternative, standard cell memories (SCMs) can be defined, synthesized, and placed and routed as an integral part of a given digital system, providing complete design flexibility, good energy efficiency, low-voltage operation, and even area efficiency for small memory blocks. Yet implementing an SCM block with a standard digital flow often fails to exploit the distinct and regular structure of such an array, leaving room for optimization. In this article, we present a design methodology for optimizing the physical implementation of SCM macros as part of the standard design flow. This methodology introduces controlled placement, leading to a structured, noncongested layout with close to 100% placement utilization, resulting in a smaller silicon footprint, reduced wire length, and lower power consumption compared to SCMs without controlled placement. This methodology is demonstrated on SCM macros of various sizes and aspect ratios in a state-of-the-art 28nm fully depleted silicon-on-insulator technology, and compared with equivalent macros designed with the noncontrolled, standard flow, as well as with foundry-supplied SRAM macros. The controlled SCMs provide an average 25% reduction in area as compared to noncontrolled implementations while achieving a smaller size than SRAM macros of up to 1Kbyte. Power and performance comparisons of controlled SCM blocks of a commonly found 256 × 32 (1 Kbyte) memory with foundry-provided SRAMs show greater than 65% and 10% reduction in read and write power, respectively, while providing faster access than their SRAM counterparts, despite being of an aspect ratio that is typically unfavorable for SCMs. In addition, the SCM blocks function correctly with a supply voltage as low as 0.3V, well below the lower limit of even the SRAM macros optimized for low-voltage operation. The controlled placement methodology is applied within a full-chip physical implementation flow of an OpenRISC-based test chip, providing more than 50% power reduction compared to equivalently sized compiled SRAMs under a benchmark application.
Adam Teman, Davide Rossi 0001, Pascal Andreas Meinerzhagen, Luca Benini, Andreas Peter Burg
ACM Trans. Design Autom. Electr. Syst.5
2016 Single-Supply 3T Gain-Cell for Low-Voltage Low-Power Applications
abstract
Logic compatible gain cell (GC)-embedded DRAM (eDRAM) arrays are considered an alternative to SRAM due to their small size, nonratioed operation, low static leakage, and two-port functionality. However, traditional GC-eDRAM implementations require boosted control signals in order to write full voltage levels to the cell to reduce the refresh rate and shorten access times. These boosted levels require either an extra power supply or on-chip charge pumps, as well as nontrivial level shifting and toleration of high voltage levels. In this brief, we present a novel, logic compatible, 3T GC-eDRAM bitcell that operates with a single-supply voltage and provides superior write capability to the conventional GC structures. The proposed circuit is demonstrated with a 2-kb memory macro that was designed and fabricated in a mature 0.18-μm CMOS process, targeted at low-power, energy-efficient applications. The test array is powered with a single supply of 900 mV, showing a 0.8-ms worst case retention time, a 1.3-ns write-access time, and a 2.4-pW/bit retention power. The proposed topology provides a bitcell area reduction of 43%, as compared with a redrawn 6-transistor SRAM in the same technology, and an overall macro area reduction of 67% including peripherals.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Lior Atias, Andreas Peter Burg, Alexander Fish
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Controlled placement of standard cell memory arrays for high density and low power in 28nm FD-SOI
abstract
Standard cell memories (SCMs) are becoming a popular alternative to SRAM IPs due to their design flexibility, ease of implementation, and robust operation at low supply voltages. Exclusively composed of standard cells, these memory arrays are implemented as part of the standard digital design flow. However, the synthesis and place and route (P&R) algorithms employed by this flow do not exploit the distinct and regular structure of an SCM array, leaving room for optimization. In this paper, we present a controlled placement design methodology for optimizing the physical implementation of SCM macros, leading to a structured, non-congested layout with close to 100% placement utilization and reduced wirelength as compared to unstructured layouts. Three sample SCM macro sizes were implemented according to the proposed methodology in a state-of-the-art 28nm FD-SOI technology, and compared with equivalent macros designed with the non-controlled, standard flow, achieving as much as a 22% reduction in area, a 57% reduction in switching power, and a 42% reduction in leakage power. In addition, these macros provide as much as an 88% reduction in switching power, as compared to equivalently sized, foundry provided SRAM IPs, while enabling robust functionality well below the minimum operating voltage of these IPs.
Adam Teman, Davide Rossi 0001, Pascal Andreas Meinerzhagen, Luca Benini, Andreas Peter Burg
ASP-DAC5
2015 Mitigating the impact of faults in unreliable memories for error-resilient applications
abstract
Inherently error-resilient applications in areas such as signal processing, machine learning and data analytics provide opportunities for relaxing reliability requirements, and thereby reducing the overhead incurred by conventional error correction schemes. In this paper, we exploit the tolerable imprecision of such applications by designing an energy-efficient fault-mitigation scheme for unreliable data memories to meet target yield. The proposed approach uses a bit-shuffling mechanism to isolate faults into bit locations with lower significance. This skews the bit-error distribution towards the low order bits, substantially limiting the output error magnitude. By controlling the granularity of the shuffling, the proposed technique enables trading-off quality for power, area, and timing overhead. Compared to error-correction codes, this can reduce the overhead by as much as 83% in read power, 77% in read access time, and 89% in area, when applied to various data mining applications in 28nm process technology.
Shrikanth Ganapathy, Georgios Karakonstantis, Adam Teman, Andreas Peter Burg
DAC4
2015 On the statistical memory architecture exploration and optimization
Charalampos Antoniadis, Georgios Karakonstantis, Nestoras E. Evmorfopoulos, Andreas Peter Burg, Georgios I. Stamoulis
DATE4
2015 Exploiting dynamic timing margins in microprocessors for frequency-over-scaling with instruction-based clock adjustment
Jeremy Constantin, Lai Wang, Georgios Karakonstantis, Anupam Chattopadhyay, Andreas Peter Burg
DATE5
2015 Energy versus data integrity trade-offs in embedded high-density logic compatible dynamic memories
Adam Teman, Georgios Karakonstantis, Robert Giterman, Pascal Andreas Meinerzhagen, Andreas Peter Burg
DATE5
2015 On metric sorting for successive cancellation list decoding of polar codes
abstract
We focus on the metric sorter unit of successive cancellation list decoders for polar codes, which lies on the critical path in all current hardware implementations of the decoder. We review existing metric sorter architectures and we propose two new architectures that exploit the structure of the path metrics in a log-likelihood ratio based formulation of successive cancellation list decoding. Our synthesis results show that, for the list size of L = 32, our first proposed sorter is 14% faster and 45% smaller than existing sorters, while for smaller list sizes, our second sorter has a higher delay in return for up to 36% reduction in the area.
Alexios Balatsoukas-Stimming, Mani Bastani Parizi, Andreas Peter Burg
ISCAS3
2015 An overlap-contention free true-single-phase clock dual-edge-triggered flip-flop
abstract
Dual-edge-triggered (DET) synchronous operation is a very attractive option for low-power, high-performance designs. Compared to conventional single-edge synchronous systems, DET operation is capable of providing the same throughput at half the clock frequency. This can lead to significant power savings on the clock network that is often one of the major contributors to total system power. However, in order to implement DET operation, special registers need to be introduced that sample data on both clock-edges. These registers are more complex than their single-edge counterparts, and often suffer from a certain amount of clock-overlap between the main clock and the internally generated inverted clock. This overlap can cause contention inside the cell and lead to logic failures, especially when operating at scaled power supplies and under process variations that characterize nanometer technologies. This paper presents a novel, static DET flip-flop (DET-FF) with a true-single-phase clock that completely avoids clock overlap hazards by eliminating the need for an inverted clock edge for functionality. The proposed DET FF was implemented in a standard 40nm CMOS technology, showing full functionality at low-voltage operating points, where conventional DET-FFs fail. Under a near-threshold, 500mV supply voltage, the proposed cell also provides a 35% lower CK-to-Q delay and the lowest power-delay-product compared to all considered DET-FF implementations.
Andrea Bonetti, Adam Teman, Andreas Peter Burg
ISCAS3
2015 Refresh-free dynamic standard-cell based memories: Application to a QC-LDPC decoder
abstract
The area and power consumption of low-density parity check (LDPC) decoders are typically dominated by embedded memories. To alleviate such high memory costs, this paper exploits the fact that all internal memories of a LDPC decoder are frequently updated with new data. These unique memory access statistics are taken advantage of by replacing all static standard-cell based memories (SCMs) of a prior-art LDPC decoder implementation by dynamic SCMs (D-SCMs), which are designed to retain data just long enough to guarantee reliable operation. The use of D-SCMs leads to a 44% reduction in silicon area of the LDPC decoder compared to the use of static SCMs. The low-power LDPC decoder architecture with refresh-free D-SCMs was implemented in a 90nm CMOS process, and silicon measurements show full functionality and an information bit throughput of up to 600 Mbps (as required by the IEEE 802.11n standard).
Pascal Andreas Meinerzhagen, Andrea Bonetti, Georgios Karakonstantis, Christoph Roth, Frank Giirkaynak, Andreas Peter Burg
ISCAS6
2015 Demo: Concurrent Spectrum Sensing and Transmission for Cognitive Radio using Self-Interference Cancellation
abstract
A demonstration of a cognitive radio network that supports concurrent spectrum sensing and transmission is presented. To detect primary users while transmitting, the secondary user node must suppress the self-interference signals. The system is implemented on a National Instruments Universal Software Radio Peripheral (USRP) platform. The demonstration will show that continuous spectrum sensing avoids the overhead for dedicated sensing periods and can detect the primary user within at most 10~ms from the start of transmission.
Andrew Austin 0001, Orion Afisiadis, Alexios Balatsoukas-Stimming, Andreas Peter Burg
MobiHoc4
2015 Enhancing Design Space Exploration by Extending CPU/GPU Specifications onto FPGAs
abstract
The design cycle for complex special-purpose computing systems is extremely costly and time-consuming. It involves a multiparametric design space exploration for optimization, followed by design verification. Designers of special purpose VLSI implementations often need to explore parameters, such as optimal bitwidth and data representation, through time-consuming Monte Carlo simulations. A prominent example of this simulation-based exploration process is the design of decoders for error correcting systems, such as the Low-Density Parity-Check (LDPC) codes adopted by modern communication standards, which involves thousands of Monte Carlo runs for each design point. Currently, high-performance computing offers a wide set of acceleration options that range from multicore CPUs to Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). The exploitation of diverse target architectures is typically associated with developing multiple code versions, often using distinct programming paradigms. In this context, we evaluate the concept of retargeting a single OpenCL program to multiple platforms, thereby significantly reducing design time. A single OpenCL-based parallel kernel is used without modifications or code tuning on multicore CPUs, GPUs, and FPGAs. We use SOpenCL (Silicon to OpenCL), a tool that automatically converts OpenCL kernels to RTL in order to introduce FPGAs as a potential platform to efficiently execute simulations coded in OpenCL. We use LDPC decoding simulations as a case study. Experimental results were obtained by testing a variety of regular and irregular LDPC codes that range from short/medium (e.g., 8,000 bit) to long length (e.g., 64,800 bit) DVB-S2 codes. We observe that, depending on the design parameters to be simulated, on the dimension and phase of the design, the GPU or FPGA may suit different purposes more conveniently, thus providing different acceleration factors over conventional multicore CPUs.
Muhsen Owaida, Gabriel Falcão Paiva Fernandes, João Andrade, Christos D. Antonopoulos, Nikolaos Bellas, Madhura Purnaprajna, David Novo, Georgios Karakonstantis, Andreas Peter Burg, Paolo Ienne
ACM Trans. Embed. Comput. Syst.9
2014 Data compression via logic synthesis
abstract
Nowadays, most software and hardware applications are committed to reduce the footprint and resource usage of data. In this general context, lossless data compression is a beneficial technique that encodes information using fewer (or at most equal number of) bits as compared to the original representation. A traditional compression flow consists of two phases: data decorrelation and entropy encoding. Data decorrelation, also called entropy reduction, aims at reducing the autocorrelation of the input data stream to be compressed in order to enhance the efficiency of entropy encoding. Entropy encoding reduces the size of the previously decorrelated data by using techniques such as Huffman coding, arithmetic coding, and others. When the data decorrelation is optimal, entropy encoding produces the strongest lossless compression possible. While efficient solutions for entropy encoding exist, data decorrelation is still a challenging problem limiting ultimate lossless compression opportunities. In this paper, we use logic synthesis to remove redundancy in binary data aiming to unlock the full potential of lossless compression. Embedded in a complete lossless compression flow, our logic synthesis based methodology is capable to identify the underlying function correlating a data set. Experimental results on data sets deriving from different causal processes show that the proposed approach achieves the highest compression ratio compared to state-of-art compression tools such as ZIP, bzip2 and 7zip.
Luca G. Amarù, Pierre-Emmanuel Gaillardon, Andreas Peter Burg, Giovanni De Micheli
ASP-DAC3
2014 A quality-scalable and energy-efficient approach for spectral analysis of heart rate variability
abstract
Today there is a growing interest in the integration of health monitoring applications in portable devices necessitating the development of methods that improve the energy efficiency of such systems. In this paper, we present a systematic approach that enables energy-quality trade-offs in spectral analysis systems for bio-signals, which are useful in monitoring various health conditions as those associated with the heart-rate. To enable such trade-offs, the processed signals are expressed initially in a basis in which significant components that carry most of the relevant information can be easily distinguished from the parts that influence the output to a lesser extent. Such a classification allows the pruning of operations associated with the less significant signal components leading to power savings with minor quality loss since only less useful parts are pruned under the given requirements. To exploit the attributes of the modified spectral analysis system, thresholding rules are determined and adopted at design- and run-time, allowing the static or dynamic pruning of less-useful operations based on the accuracy and energy requirements. The proposed algorithm is implemented on a typical sensor node simulator and results show up-to 82% energy savings when static pruning is combined with voltage and frequency scaling, compared to the conventional algorithm in which such trade-offs were not available. In addition, experiments with numerous cardiac samples of various patients show that such energy savings come with a 4.9% average accuracy loss, which does not affect the system detection capability of sinus-arrhythmia which was used as a test case.
Georgios Karakonstantis, Aviinaash Sankaranarayanan, Mohamed M. Sabry, David Atienza 0001, Andreas Peter Burg
DATE5
2014 A Wireless Body Sensor Network for Activity Monitoring with Low Transmission Overhead
abstract
Activity recognition has been a research field of high interest over the last years, and it finds application in the medical domain, as well as personal healthcare monitoring during daily home- and sports-activities. With the aim of producing minimum discomfort while performing supervision of subjects, miniaturized networks of low-power wireless nodes are typically deployed on the body to gather and transmit physiological data, thus forming a Wireless Body Sensor Network (WBSN). In this work, we propose a WBSN for online activity monitoring, which combines the sensing capabilities of wearable nodes and the high computational resources of modern smart phones. The proposed solution provides different tradeoffs between classification accuracy and energy consumption, thanks to different workloads assigned to the nodes and to the mobile phone in different network configurations. In particular, our WBSN is able to achieve very high activity recognition accuracies (up to 97.2%) on multiple subjects, while significantly reducing the sampling frequency and the volume of transmitted data with respect to other state-of-the-art solutions.
Rubén Braojos, Ivan Beretta, Jeremy Constantin, Andreas Peter Burg, David Atienza 0001
EUC4
2014 FPGA implementation of an interior point method for high-speed model predictive control
abstract
In this paper, we present a hardware architecture for implementing an interior point method for model predictive control (MPC) on field programmable gate arrays (FPGA). The FPGA implementation allows the solution of quadratic programs occurring in MPC at very high speed. Experiments show that our hardware implementation is able to outperform an software implementation running on a high-end CPU while consuming significantly less power making it well-suited for embedded industrial control applications. In contrast to existing FPGA implementations, the proposed solution exploits the MPC-specific problem structure with the direct linear equation solver and uses an efficient predictor-corrector algorithm. Moreover, the modular design of the architecture simplifies customization or extension to special control problem classes. The proposed FPGA solution can broaden the applicability of solving complex or large MPC problems in embedded computing platforms that were so far considered out of reach.
Helfried Peyrl, Andreas Peter Burg, George A. Constantinides
FPL3
2014 LLR-based successive cancellation list decoding of polar codes
abstract
We present an LLR-based implementation of the successive cancellation list (SCL) decoder. To this end, we associate each decoding path with a metric which (i) is a monotone function of the path's likelihood and (ii) can be computed efficiently from the channel LLRs. The LLR-based formulation leads to a more efficient hardware implementation of the decoder compared to the known log-likelihood based implementation. Synthesis results for an SCL decoder with block-length of N = 1024 and list sizes of L = 2 and L = 4 confirm that the LLR-based decoder has considerable area and operating frequency advantages in the orders of 50% and 30%, respectively.
Alexios Balatsoukas-Stimming, Mani Bastani Parizi, Andreas Peter Burg
ICASSP3
2014 Robust asynchronous indoor localization using LED lighting
abstract
We propose a low-cost system for indoor self-localization of mobile devices using modulated LED ceiling lamps that are fully autonomous and broadcast their identifiers without any synchronization. The proposed self-localization method is designed to handle this lack of synchronization as well as the possibility of blocked line-of-sight connections or severe attenuation in real-world environments. This robustness is achieved by applying a suitable Bayesian signal model and by taking into account the inherent sparsity in detecting the concurrently visible lamps. The proposed estimator of the location approximates optimal Bayesian estimation while maintaining low complexity. Simulation results confirm a significant gain in performance compared to a classical matched-filter approach.
Georg Kail, Patrick Maechler, Nicholas Preyss, Andreas Peter Burg
ICASSP4
2014 Cross layer energy-efficiency optimization for cognitive radio transceivers
abstract
Designing energy-efficient cognitive radio transceivers requires joint optimization of medium access control and the physical layer implementation. In this paper we show an energy efficiency optimization strategy for IEEE 802.11n compliant transceivers in terms of energy consumed by the receiver per successfully received bit. To this end, we propose and explore several modifications of a conventional physical layer implementation, all of which target energy proportional behavior. The proposed modifications intentionally include operation modes and algorithm choices that are suboptimal with respect to throughput and error-rate performance. Yet, we show how (under ideal conditions) the rate adaptation at the medium access control layer can exploit these modifications to achieve superior energy efficiency that is 44% below that of a rate adaptation targeting only maximum goodput.
Christian Senning, Mikel Mendicute, Andreas Peter Burg
ICASSP3
2014 4T Gain-Cell with internal-feedback for ultra-low retention power at scaled CMOS nodes
abstract
Gain-Cell embedded DRAM (GC-eDRAM) has recently been recognized as a possible alternative to traditional SRAM. While GC-eDRAM inherently provides high-density, low-leakage, low-voltage, and 2-ported operation, its limited retention time requires periodic, power-hungry refresh cycles. This drawback is further enhanced at scaled technologies, where increased subthreshold leakage currents and decreased in-cell storage capacitances result in faster data deterioration. In this paper, we present a novel 4T GC-eDRAM bitcell that utilizes an internal feedback mechanism to significantly increase the data retention time in scaled CMOS technologies. A 2 kb memory macro was implemented in a low-power 65nm CMOS technology, displaying an over 3× improvement in retention time over the best previous publication at this node. The resulting array displays a nearly 5× reduction in retention power (despite the refresh power component) with a 40% reduction in bitcell area, as compared to a standard 6T SRAM.
Robert Giterman, Adam Teman, Pascal Andreas Meinerzhagen, Andreas Peter Burg, Alexander Fish
ISCAS4
2014 A signal processor for Gaussian message passing
abstract
In this paper, we present a novel signal processing unit built upon the theory of factor graphs, which is able to address a wide range of signal processing algorithms. More specifically, the demonstrated factor graph processor (FGP) is tailored to Gaussian message passing algorithms. We show how to use a highly configurable systolic array to solve the message update equations of nodes in a factor graph efficiently. A proper instruction set and compilation procedure is presented. In a recursive least squares channel estimation example we show that the FGP can compute a message update faster than a state-of-the-art DSP. The results demonstrate the usabilty of the FGP architecture as a flexible HW accelerator for signal-processing and communication systems.
Harald Kröll, Stefan Zwicky, Reto Odermatt, Lukas Bruderer, Andreas Peter Burg, Qiuting Huang
ISCAS5
2014 Enabling complexity-performance trade-offs for successive cancellation decoding of polar codes
abstract
Polar codes are one of the most recent advancements in coding theory and they have attracted significant interest. While they are provably capacity achieving over various channels, they have seen limited practical applications. Unfortunately, the successive nature of successive cancellation based decoders hinders fine-grained adaptation of the decoding complexity to design constraints and operating conditions. In this paper, we propose a systematic method for enabling complexity-performance tradeoffs by constructing polar codes based on an optimization problem which minimizes the complexity under a suitably defined mutual information based performance constraint. Moreover, a low-complexity greedy algorithm is proposed in order to solve the optimization problem efficiently for very large code lengths.
Alexios Balatsoukas-Stimming, Georgios Karakonstantis, Andreas Peter Burg
ISIT3
2014 Faulty successive cancellation decoding of polar codes for the binary erasure channel
Alexios Balatsoukas-Stimming, Andreas Peter Burg
ISITA2
2013 Fast and accurate BER estimation methodology for I/O links based on extreme value theory
abstract
This paper introduces a novel approach towards the statistical analysis of modern high-speed I/O and similar communication links, which is capable of reliably to determine extremely low (∼10−12or lower) bit error rates (BER) by using techniques from extreme value theory (EVT). The new method requires only a small amount of voltage values at the received eye center, which can be generated by running circuit/system level simulations or measuring fabricated I/O circuits, to predict link BERs. Unlike conventional techniques, no simplifying assumptions on link noise and interference sources are required making this approach extremely portable to any communication system operating with very low BER. Our experimental results show that the BER estimates from the proposed methodology are on the same order of magnitude as traditional time domain, transient eye diagram simulations for links with BER of 10−6and 10−5operating at 9.6 and 10.1 Gbps respectively.
Alessandro Cevrero, Nestoras E. Evmorfopoulos, Charalampos Antoniadis, Paolo Ienne, Yusuf Leblebici, Andreas Peter Burg, Georgios I. Stamoulis
DATE6
2013 Synchronizing code execution on ultra-low-power embedded multi-channel signal analysis platforms
abstract
Embedded biosignal analysis involves a considerable amount of parallel computations, which can be exploited by employing low-voltage and ultra-low-power (ULP) parallel computing architectures. By allowing data and instruction broadcasting, single instruction multiple data (SIMD) processing paradigm enables considerable power savings and application speedup, in turn allowing for a lower voltage supply for a given workload. The state-of-the-art multi-core architectures for biosignal analysis however lack a bare, yet smart, synchronization technique among the cores, allowing lockstep execution of algorithm parts that can be performed using the SIMD, even in the presence of data-dependent execution flows. In this paper, we propose a lightweight synchronization technique to enhance an ULP multi-core processor, resulting in improved energy efficiency through lockstep SIMD execution. Our results show that the proposed improvements accomplish tangible power savings, up to 64% for an 8-core system operating at a workload of 89 MOps/s while exploiting voltage scaling.
Ahmed Yasir Dogan, Rubén Braojos, Jeremy Constantin, Giovanni Ansaloni, Andreas Peter Burg, David Atienza 0001
DATE5
2013 Efficient vlsi implementation of reduced-state sequence estimation for wireless communications
abstract
Modern wireless communication systems require efficient channel equalizer implementations. This paper explores the design space of reduced-state sequence estimation (RSSE). We show how the concept of pre-computation can be applied to greatly reduce computational complexity, such that efficient RSSE architectures can be derived. As a proof of concept, an RSSE was implemented in dedicated hardware, that achieves a 1.6 times higher hardware efficiency when compared to prior art.
Stefan Zwicky, Christian Benkeser, Andreas Peter Burg, Qiuting Huang
ICASSP3
2013 Live demonstration: Real-time audio restoration using sparse signal recovery
abstract
We demonstrate the restoration of audio signals corrupted by clicks and pops using techniques from sparse signal recovery and compressive sensing. The demonstration features real-time signal restoration using the approximate message passing algorithm on an FPGA prototyping board. To highlight the restoration performance of our implementation, we remove clicks and pops from old phonograph recordings in real time.
David E. Bellasi, Patrick Maechler, Andreas Peter Burg, Norbert Felber, Hubert Kaeslin, Christoph Studer
ISCAS3
2012 Instruction Set Extensions for Cryptographic Hash Functions on a Microcontroller Architecture
abstract
In this paper, we investigate the benefits of instruction set extensions (ISEs) on a 16-bit microcontroller architecture for software implementations of cryptographic hash functions,using the example of the five SHA-3 final round candidates. We identify the general algorithm bottlenecks, taking into account memory footprints and cycle counts of our optimized reference assembly implementations. We show that our target applications benefit from algorithm-specific ISEs based on finite state machines for address generation, lookup table integration, and extension of computational units through microcoded instructions.The gains in throughput, memory consumption, and the area overhead are assessed, by implementing the modified cores and applications utilizing the developed ISEs. Our results show that with less than 10% additional core area, it is possible to increase the execution speed on average by 172% (ranging from 21%to 703%), while reducing memory requirements on average by more than 40%.
Jeremy Constantin, Andreas Peter Burg, Frank K. Gürkaynak
ASAP2
2012 On the exploitation of the inherent error resilience of wireless systems under unreliable silicon
abstract
In this paper, we investigate the impact of circuit misbehavior due to parametric variations and voltage scaling on the performance of wireless communication systems. Our study reveals the inherent error resilience of such systems and argues that sufficiently reliable operation can be maintained even in the presence of unreliable circuits and manufacturing defects. We further show how selective application of more robust circuit design techniques is sufficient to deal with high defect rates at low overhead and improve energy efficiency with negligible system performance degradation.
Georgios Karakonstantis, Christoph Roth, Christian Benkeser, Andreas Peter Burg
DAC4
2012 Multi-core architecture design for ultra-low-power wearable health monitoring systems
abstract
Personal health monitoring systems can offer a cost-effective solution for human healthcare. To extend the lifetime of health monitoring systems, we propose a near-threshold ultra-low-power multi-core architecture featuring low-power cores, yet capable of executing biomedical applications, with multiple instruction and data memories, tightly coupled through flexible crossbar interconnects. This architecture also includes broadcasting mechanisms for the data and instruction memories to optimize system energy consumption by tailoring memory sharing to the target application. Moreover, the architecture enables power gating of the unused memory banks to lower leakage power. Our experimental results show that compared to the state-of-the-art, the proposed architecture achieves 39.5% power savings at high workload requirements (637 MOps/s), and 38.8% savings at low workload requirements (5 kOps/s), whereby leakage power consumption dominates.
Ahmed Yasir Dogan, Jeremy Constantin, Martino Ruggiero, Andreas Peter Burg, David Atienza 0001
DATE4
2012 Shortening Design Time through Multiplatform Simulations with a Portable OpenCL Golden-model: The LDPC Decoder Case
abstract
Hardware designers and engineers typically need to explore a multi-parametric design space in order to find the best configuration for their designs using simulations that can take weeks to months to complete. For example, designers of special purpose chips need to explore parameters such as the optimal bit width and data representation. This is the case for the development of complex algorithms such as Low-Density Parity-Check (LDPC) decoders used in modern communication systems. Currently, high-performance computing offers a wide set of acceleration options, that range from multicore CPUs to graphics processing units (GPUs) and FPGAs. Depending on the simulation requirements, the ideal architecture to use can vary. In this paper we propose a new design flow based on Open CL, a unified multiplatform programming model, which accelerates LDPC decoding simulations, thereby significantly reducing architectural exploration and design time. Open CL-based parallel kernels are used without modifications or code tuning on multicore CPUs, GPUs and FPGAs. We use SOpen CL (Silicon to Open CL), a tool that automatically converts Open CL kernels to RTL for mapping the simulations into FPGAs. To the best of our knowledge, this is the first time that a single, unmodified Open CL code is used to target those three different platforms. We show that, depending on the design parameters to be explored in the simulation, on the dimension and phase of the design, the GPU or the FPGA may suit different purposes more conveniently, providing different acceleration factors. For example, although simulations can typically execute more than 3× faster on FPGAs than on GPUs, the overhead of circuit synthesis often outweighs the benefits of FPGA-accelerated execution.
Gabriel Falcão Paiva Fernandes, Muhsen Owaida, David Novo, Madhura Purnaprajna, Nikolaos Bellas, Christos D. Antonopoulos, Georgios Karakonstantis, Andreas Peter Burg, Paolo Ienne
FCCM8
2012 Two-port low-power gain-cell storage array: Voltage scaling and retention time
abstract
The impact of supply voltage scaling on the retention time of a 2-transistor (2T) gain-cell (GC) storage array is investigated, in order to enable low-power/low-voltage data storage. The retention time can be increased when scaling down the supply voltage for a given access statistics and a given write bit-line (WBL) control scheme. Moreover, for a given supply voltage, the retention time can be further increased by controlling the WBL to a voltage level between the supply rails during idle and read states. These two concepts are proved by means of Spectre simulation of a GC-storage array implemented in 180-nm CMOS technology. The proposed 2-kb storage macro is operated at only 40% of the nominal supply voltage and leverages the GCs to enable two-port operation with a negligible area-increase compared to a single-port implementation.
Pascal Andreas Meinerzhagen, Andreas Peter Burg
ISCAS3
2012 Hardware-efficient random sampling of fourier-sparse signals
abstract
Spectrum sensing, i.e. the identification of occupied frequencies within a large bandwidth, requires complex sampling hardware. Measurements suggest that only a small fraction of the available spectrum is actually used at any time and place, which allows a sparse characterization of the frequency domain signal. Compressed sensing (CS) can exploit this sparsity and simplify measurements. We investigate the performance of a very simple hardware architecture based on the slope analog-to-digital converter (ADC), which allows to sample signals at unevenly spaced points in time. CS algorithms are used to identify the occupied frequencies, which can be continuously distributed across a large bandwidth.
Patrick Maechler, Norbert Felber, Hubert Kaeslin, Andreas Peter Burg
ISCAS4
2012 Successive interference cancellation for 3G downlink: Algorithm and VLSI architecture
abstract
This paper presents a VLSI implementation of an MMSE successive interference cancellation multiuser detector (SIC-MUD) for the downlink of a TD-SCDMA system.Computation in the frequency domain, group-wise interference cancellation, and pre-computation of filter coefficients enable an efficient architecture suitable for mobile handsets.Our implementation in 0.13µm CMOS technology proves that the SIC-MUD is a viable solution for the TD-SCDMA downlink, providing a notable performance gain at a moderate increase in complexity compared to linear equalizers.
Sandro Belfanti, Christian Benkeser, Karim Badawi, Qiuting Huang, Andreas Peter Burg
VLSI-SoC5
2012 TamaRISC-CS: An ultra-low-power application-specific processor for compressed sensing
abstract
Abstract—Compressed sensing (CS) is a universal technique for the compression of sparse signals. CS has been widely used in sensing platforms where portable, autonomous devices have to operate for long periods of time with limited energy resources. Therefore, an ultra-low-power (ULP) CS implementation is vital for these kind of energy-limited systems. Sub-threshold (sub-VT) operation is commonly used for ULP computing, and can also be combined with CS. However, most established CS implementations can achieve either no or very limited benefit from sub-VT operation. Therefore, we propose a sub-VT application-specific instruction-set processor (ASIP), exploiting the specific operations of CS. Our results show that the proposed ASIP accomplishes 62x speed-up and 11.6x power savings with respect to an established CS implementation running on the baseline low-power processor. I.
Jeremy Constantin, Ahmed Yasir Dogan, Oskar Andersson, Pascal Andreas Meinerzhagen, Joachim Neves Rodrigues, David Atienza 0001, Andreas Peter Burg
VLSI-SoC7
2012 Analysis and VLSI Implementation of EWA Rendering for Real-Time HD Video Applications
abstract
Nonlinear image warping or image resampling is a necessary step in many current and upcoming video applications, such as video retargeting, stereoscopic 3-D mapping, and multiview synthesis. The challenges for real-time resampling include not only image quality but also available energy and computational power of the employed device. In this paper, we employ an elliptical-weighted average (EWA) rendering approach to 2-D image resampling. We extend the classical EWA framework for increased visual quality and provide a very large scale integration architecture for efficient view rendering. The resulting architecture is able to render high-quality video sequences in real time targeted for low-power applications in end-user display devices.
Pierre Greisen, Michael Schaffner, Simon Heinzle, Marian Runo, Aljoscha Smolic, Andreas Peter Burg, Hubert Kaeslin, Markus Gross 0001
IEEE Trans. Circuits Syst. Video Technol.6
2011 Design and failure analysis of logic-compatible multilevel gain-cell-based dram for fault-tolerant VLSI systems
abstract
This paper considers the problem of increasing the storage density in fault-tolerant VLSI systems which require only limited data retention times. To this end, the concept of storing many bits per memory cell is applied to area-efficient and fully logic-compatible gain-cell-based dynamic memories. A memory macro in 90-nm CMOS technology including multilevel write and read circuits is proposed and analyzed with respect to its read failure probability due to within-die process variations by means of Monte Carlo simulations.
Pascal Andreas Meinerzhagen, Onur Andiç, Jürg Treichler, Andreas Peter Burg
ACM Great Lakes Symposium on VLSI4
2011 Area, throughput, and energy-efficiency trade-offs in the VLSI implementation of LDPC decoders
abstract
Low-density parity-check (LDPC) codes are key ingredients for improving reliability of modern communication systems and storage devices. On the implementation side however, the design of energy-efficient and high-speed LDPC decoders with a sufficient degree of reconfigurability to meet the flexibility demands of recent standards remains challenging. This survey paper provides an overview of the state-of-the-art in the design of LDPC decoders using digital integrated circuits. To this end, we summarize available algorithms and characterize the design space. We analyze the different architectures and their connection to different codes and requirements. The advantages and disadvantages of the various choices are illustrated by comparing state-of-the-art LDPC decoder designs.
Christoph Roth, Alessandro Cevrero, Christoph Studer, Yusuf Leblebici, Andreas Peter Burg
ISCAS5
2011 Computational stereo camera system with programmable control loop
abstract
Stereoscopic 3D has gained significant importance in the entertainment industry. However, production of high quality stereoscopic content is still a challenging art that requires mastering the complex interplay of human perception, 3D display properties, and artistic intent. In this paper, we present a computational stereo camera system that closes the control loop from capture and analysis to automatic adjustment of physical parameters. Intuitive interaction metaphors are developed that replace cumbersome handling of rig parameters using a touch screen interface with 3D visualization. Our system is designed to make stereoscopic 3D production as easy, intuitive, flexible, and reliable as possible. Captured signals are processed and analyzed in real-time on a stream processor. Stereoscopy and user settings define programmable control functionalities, which are executed in real-time on a control processor. Computational power and flexibility is enabled by a dedicated software and hardware architecture. We show that even traditionally difficult shots can be easily captured using our system.
Simon Heinzle, Pierre Greisen, David Gallup, Christine Chen, Daniel Saner, Aljoscha Smolic, Andreas Peter Burg, Wojciech Matusik, Markus Gross 0001
ACM Trans. Graph.7
2010 VLSI implementation of a low-complexity LLL lattice reduction algorithm for MIMO detection
abstract
Lattice-reduction (LR)-aided successive interference cancellation (SIC) is able to achieve close-to optimum error-rate performance for data detection in multiple-input multiple-output (MIMO) wireless communication systems. In this work, we propose a hardware-efficient VLSI architecture of the Lenstra-Lenstra-Lovász (LLL) LR algorithm for SIC-based data detection. For this purpose, we introduce various algorithmic modifications that enable an efficient hardware implementation. Comparisons with existing FPGA implementations show that our design outperforms state-of-the-art LR implementations in terms of hardware-efficiency and throughput. We finally provide reference ASIC implementation results for 130nm CMOS technology.
Lukas Bruderer, Christoph Studer, Markus Wenk, Dominik Seethaler, Andreas Peter Burg
ISCAS5
2010 Matching pursuit: Evaluation and implementatio for LTE channel estimation
abstract
The emerging research field of compressed sensing (CS) promises better signal reconstruction out of fewer measurements if a sparse representation of the signal exists. Since wireless broadband channels often exhibit a sparse impulse response, CS reconstruction algorithms were proposed for channel estimation. In this paper, a hardware architecture for channel estimation using the matching pursuit algorithm is presented. The reference design targets the 3GPP LTE standard with a channel bandwidth of up to 20 MHz. Achievable performance gains over least squares channel estimation are illustrated by means of simulations. The costs in terms of chip area and reconstruction time for 180 nm CMOS technology are presented together with an analysis of the tradeoff between hardware complexity and reconstruction performance.
Patrick Maechler, Pierre Greisen, Norbert Felber, Andreas Peter Burg
ISCAS4
2010 Area- and throughput-optimized VLSI architecture of sphere decoding
abstract
Sphere decoding (SD) is a promising means for implementing high-performance data detection in multiple-input multiple-output (MIMO) wireless communication systems. In this paper, we focus on the register transfer level implementation of SD with minimum area-delay product for application in wideband MIMO communication systems, such as IEEE 802.11n, where multiple SD cores need to be instantiated. The basic architectural considerations and the proposed optimizations are explained based on hard-output SD, but are also applicable to soft-output SD. Corresponding VLSI implementation results (for both hard-output and soft-output SD) show an improvement in the area-delay product by almost 50% compared to that of other SD implementations reported in the literature.
Markus Wenk, Lukas Bruderer, Andreas Peter Burg, Christoph Studer
VLSI-SoC3
2008 A Real-Time 4-Stream MIMO-OFDM Transceiver: System Design, FPGA Implementation, and Characterization
abstract
When designing complex communication systems, such as MIMO-OFDM transceivers, prototypes have become an important tool for understanding the implementation trade-offs and the system behavior. This paper presents a real-time FPGA prototype for a 4-stream MIMO-OFDM transceiver capable of transmitting 216 Mbit/s in 20 MHz bandwidth. The paper covers all parts of the system from RF to channel decoding and considers both algorithm and implementation aspects. In particular, we discuss the initial parameter estimation, channel estimation, MIMO detection, parameter tracking, and channel decoding. FPGA implementation results are reported along with measurements that demonstrate the throughput of spatial multiplexing with four spatial streams.
Simon Haene, David Perels, Andreas Peter Burg
IEEE J. Sel. Areas Commun.3
2008 Soft-output sphere decoding: algorithms and VLSI implementation
abstract
Multiple-input multiple-output (MIMO) detection algorithms providing soft information for a subsequent channel decoder pose significant implementation challenges due to their high computational complexity. In this paper, we show how sphere decoding can be used as an efficient tool to implement soft-output MIMO detection with flexible trade-offs between computational complexity and (error rate) performance. In particular, we provide VLSI implementation results which demonstrate that single tree-search, sorted QR-decomposition, channel matrix regularization, log-likelihood ratio clipping, and imposing runtime constraints are the key ingredients for realizing soft-output MIMO detectors with near max-log performance at a chip area that is only 58% higher than that of the best-known hard-output sphere decoder VLSI implementation.
Christoph Studer, Andreas Peter Burg, Helmut Bölcskei
IEEE J. Sel. Areas Commun.2
2007 Reduced-complexity mimo detector with close-to ml error rate performance
abstract
Maximum likelihood (ML) detection provides optimum error rate performance for uncoded multiple-input multiple-output (MIMO) systems. However, circuit complexity of a straightforward implementation of ML detection is uneconomic for high-rate systems. This paper addresses the VLSI implementation trade-offs of a MIMO detection algorithm that achieves close-to ML error rate performance with reduced computational complexity. The described implementations in a 0.25 μm CMOS technology for 4-4 MIMO systems feature a simple data-path, achieve high throughput, and use small silicon area. Important contributing factors to these results are efficient enumeration strategies and the application of simplified norms and sophisticated scheduling techniques together with a new low-complexity preprocessing scheme.
C. Hess, Markus Wenk, Andreas Peter Burg, Peter Luethi, Christoph Studer, Norbert Felber, Wolfgang Fichtner
ACM Great Lakes Symposium on VLSI3
2007 Regularized Frequency Domain Equalization Algorithm and its VLSI Implementation
abstract
Approximation of Toeplitz matrices with cir-culant matrices is a well-known approach to reduce the computational complexity of linear equalizers. This paper presents a novel technique to compute linear equalizer coefficients in the frequency domain. It is shown how a regularization term can help to reduce the error caused by the frequency domain approximation. A corresponding VLSI implementation provides reference for the true silicon complexity and for the complexity increase associated with the proposed algorithm.
Andreas Peter Burg, Simon Haene, Wolfgang Fichtner, Markus Rupp
ISCAS1
2007 VLSI Implementation of a Lattice-Reduction Algorithm for Multi-Antenna Broadcast Precoding
abstract
This paper describes the first VLSI implementation of lattice reduction (LR) aided multi-antenna broadcast precoding with vector perturbation. The considered LR, scheme is based on Brun's algorithm for finding integer relations. We analyze its high-level architectural issues, we devise a corresponding low-complexity implementation, and, finally, we develop a suitable VLSI architecture. The resulting circuit provides reference for the true silicon complexity of LR, for broadcast precoding with vector perturbation
Andreas Peter Burg, Dominik Seethaler, Gerald Matz
ISCAS1
2007 FFT Processor for OFDM Channel Estimation
abstract
Pilot-assisted channel estimation for communication systems employing orthogonal frequency division multiplexing modulation requires significant signal processing at the receiver if the correlation among the frequency-domain channel coefficients is to be exploited in order to improve accuracy. In this work, a conventional FFT processor is extended to support all operations required by a selected channel estimation algorithm, so that both OFDM de/modulation and channel estimation can be efficiently performed on the same hardware unit. The silicon complexity of the extended processor, which was prototyped in a real-time testbed using FPGAs, is compared to a conventional FFT processor.
Simon Haene, Andreas Peter Burg, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS2
2007 VLSI Implementation of a High-Speed Iterative Sorted MMSE QR Decomposition
abstract
The QR decomposition is an important, but often underestimated prerequisite for pseudo- or non-linear detection methods such as successive interference cancellation or sphere decoding for multiple-input multiple-output (MIMO) systems. The ability of concurrent iterative sorting during the QR decomposition introduces a moderate overall latency, but provides the base for an improved layered stream decoding. This paper describes the architecture and results of the first VLSI implementation of an iterative sorted QR decomposition preprocessor for MIMO receivers. The presented architecture performs MIMO channel preprocessing using Givens rotations in order to compute the minimum mean squared error QR decomposition
Peter Luethi, Andreas Peter Burg, Simon Haene, David Perels, Norbert Felber, Wolfgang Fichtner
ISCAS2
2006 Advanced receiver algorithms for MIMO wireless communications
abstract
We describe the VLSI implementation of MIMO detectors that exhibit close-to optimum error-rate performance, but still achieve high throughput at low silicon area. In particular, algorithms and VLSI architectures for sphere decoding (SD) and K-best detection are considered, and the corresponding trade-offs between uncoded error-rate performance, silicon area, and throughput are explored. We show that SD with a per-block run-time constraint is best suited for practical implementations.
Andreas Peter Burg, Moritz Borgmann, Markus Wenk, Christoph Studer, Helmut Bölcskei
DATE1
2006 A Frame-Start Detector for a 4×4 MIMO-OFDM System
abstract
Future wireless LANs will increase the peak data rate by employing multiple antennas at both transmitter and receiver. Well designed synchronization algorithms are a prerequisite for meeting stringent QoS requirements. In particular OFDM modulation, which constitutes the basics for WLAN, is very sensitive to timing synchronization errors which incur inter-symbol interference. In this paper, a novel frame synchronization algorithm is proposed that is implemented in the FPGA of a real-time MIMO-OFDM testbed. Simulations show it to be of sufficient performance in scenarios of interest, while the hardware complexity is suitable for an FPGA implementation. Additionally, the algorithm exhibits a good resilience against narrow-band interference, which causes problems in traditional frame-start detection algorithms
David Perels, Simon Haene, Andreas Peter Burg, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ICASSP (4)3
2006 Algorithm and VLSI architecture for linear MMSE detection in MIMO-OFDM systems
abstract
The paper describes an algorithm and a corresponding VLSI architecture for the implementation of linear MMSE detection in packet-based MIMO-OFDM communication systems. The advantages of the presented receiver architecture are low latency, high-throughput, and efficient resource utilization, since the hardware required for the computation of the MMSE estimators is reused for the detection. The algorithm also supports the extraction of soft information for channel decoding
Andreas Peter Burg, Simon Haene, David Perels, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS1
2006 Silicon implementation of an MMSE-based soft demapper for MIMO-BICM
abstract
The performance of systems employing bit-interleaved coded modulation (BICM) critically depends on the availability of soft information. In the multi-antenna case, the extraction of optimum bit-metrics becomes prohibitively complex, so that suboptimal solutions need to be adopted for practical implementation. Instead of considering all the spatially multiplexed streams jointly, the implementation presented in this paper computes the soft information on each data stream separately, based on the output of an MMSE equalizer
Simon Haene, Andreas Peter Burg, David Perels, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS2
2006 K-best MIMO detection VLSI architectures achieving up to 424 Mbps
abstract
From an error rate performance perspective, maximum likelihood (ML) detection is the preferred detection method for multiple-input multiple-output (MIMO) communication systems. However, for high transmission rates a straight forward exhaustive search implementation suffers from prohibitive complexity. The K-best algorithm provides close-to-ML bit error rate (BER) performance, while its circuit complexity is reduced compared to an exhaustive search. In this paper, a new VLSI architecture for the implementation of the K-best algorithm is presented. Instead of the mostly sequential processing that has been applied in previous VLSI implementations of the algorithm, the presented solution takes a more parallel approach. Furthermore, the application of a simplified norm is discussed. The implementation in an ASIC achieves up to 424 Mbps throughput with an area that is almost on par with current state-of-the-art implementations
Markus Wenk, Martin Zellweger, Andreas Peter Burg, Norbert Felber, Wolfgang Fichtner
ISCAS3
2004 A 2 Gb/s balanced AES crypto-chip implementation
abstract
We present a balanced 2 Gb/s en-/decryption ASIC realization of the AES algorithm that supports all standard operation modes and key lengths. Rather than optimizing only for throughput, special care is taken to balance the more involved decryption path with that of the encryption path using a number of high-level architectural and register transfer level optimizations. The fabricated en-/decryption core requires an active area of only 3.56 mm2 (less than 120,000 gate equivalents) in a modest 0.25 µm CMOS technology.
Frank K. Gürkaynak, Andreas Peter Burg, Norbert Felber, Wolfgang Fichtner, D. Gasser, Franco Hug, Hubert Kaeslin
ACM Great Lakes Symposium on VLSI2
2003 Prototype experience for MIMO BLAST over third-generation wireless system
abstract
In this paper, a multiple-input-multiple-output (MIMO) extension for a third-generation (3G) wireless system is described. The integration of MIMO concepts within the existing UMTS standard and the associated space-time RAKE receiver are explained. An analysis is followed by a description of an actual experimental MIMO transmitter and receiver architecture, both realized on digital signal processors (DSPs) and FPGAs within a precommercial OneBTS base station. It uses four transmit and four receive antennas to achieve downlink data rates up to 1 Mb/s per user with a spreading factor of 32 and the UMTS chip rate of 3.84 MHz. Furthermore, different MIMO detectors are evaluated, comparing their performance and complexity. System performance is evaluated through simulations and indoor over-the-air measurements. Capacity and bit-error rate measurement results are presented.
Ali Adjoudani, Eric C. Beck, Andreas Peter Burg, Goran M. Djuknic, Thomas G. Gvoth, David A. Haessig, Salim Manji, Michelle A. Milbrodt, Markus Rupp, Dragan Samardzija, Arnold B. Siegel, Theodore Sizer, Cuong Tran 0002, Susan Walker, Stephen A. Wilkus, Peter W. Wolniansky
IEEE J. Sel. Areas Commun.3
2003 Rapid prototyping for wireless designs: the five-ones approach
Markus Rupp, Andreas Peter Burg, Eric C. Beck
Signal Process.2