Alexios Balatsoukas-Stimming

dblp:121/1599 · DBLP profile ↗
← Back
40ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0002-6721-4666ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 1 first-author · 8 since 2021Systems, architecture and hardware · 12 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 An SDR-Based Monostatic Wi-Fi System with Analog Self-Interference Cancellation for Sensing
abstract
Wireless sensing offers an alternative to wearables for contactless monitoring of human activity and vital signs. However, most existing systems use bistatic setups, which suffer from phase imperfections due to unsynchronized clocks. Monostatic systems overcome this issue, but are hindered by strong self-interference (SI) that requires effective cancellation. We present a monostatic Wi-Fi sensing system that uses an auxiliary transmit RF chain to achieve SI cancellation levels of 40 dB, comparable to existing solutions with custom cancellation hardware. We demonstrate that the cancellation filter weights, fine-tuned using least mean squares, can be directly repurposed for target sensing. Moreover, we achieve stable SI cancellation over 30 minutes in an office environment without fine-tuning, enabling traditional vital sign monitoring using channel estimates derived from baseband samples without the adaptation of the cancellation affecting the sensing channel – a significant limitation in prior work. Experimental results confirm the detection of small, slow-moving targets, representative for breathing chest movements, at distances up to 10 meters in non-line-of-sight conditions.
Andreas Toftegaard Kristensen, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ISCAS2
2025 Toward Universal Belief Propagation Decoding for Short Binary Block Codes
abstract
Belief propagation (BP) decoding has been recognized for its capacity-approaching performance and high throughput when decoding long low-density parity-check (LDPC) codes. However, the application of BP decoding for short codes is hindered by dense parity-check matrices (PCMs) and prevalent short cycles in the Tanner graph. In this paper, we introduce a general method to extract an optimized sparse PCM for short binary block codes, which removes length-four cycles and enhances the connectivity of short cycles to enable BP decoding with improved performance. Notably, for short binary codes with lengths up to 64, our BP decoding performance approaches the maximum likelihood bound and surpasses the best-reported BP results with reduced computational complexity. Compared with other universal decoding algorithms, BP decoding using our extracted sparse PCMs is competitive in terms of both error-rate performance and computational complexity. These promising results suggest that our method to improve BP decoding for short codes is a step toward a practical universal BP decoder for next-generation communication systems.
Yifei Shen 0003, Zongyao Li 0003, Yuqing Ren, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg
IEEE J. Sel. Areas Commun.5
2025 Hardware Implementation of Projection-Aggregation Decoders for Reed-Muller Codes
abstract
This paper presents the hardware architecture and implementation of two variants of projection-aggregation-based decoding of Reed-Muller (RM) codes, namely unique projection aggregation (UPA) and collapsed projection aggregation (CPA). Through thorough analysis and experimentation, we observe that the hardware implementation of UPA exhibits superior resource usage and reduced energy consumption compared to CPA for the iterative projection aggregation (IPA) decoder. This finding underscores a critical insight: reducing computational cost, in isolation, may not necessarily translate into hardware cost effectiveness.
Marzieh Hashemipour-Nazari, Andrea Nardi-Dei, Kees Goossens, Alexios Balatsoukas-Stimming
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Edge-Spreading Raptor-Like LDPC Codes for 6G Wireless Systems
abstract
Next-generation channel coding has stringent demands on throughput, energy consumption, and error rate performance while maintaining key features of 5G New Radio (NR) standard codes such as rate compatibility, which is a significant challenge. Due to excellent capacity-achieving performance, spatially-coupled low-density parity-check (SC-LDPC) codes are considered a promising candidate for next-generation channel coding. In this paper, we propose an SC-LDPC code family called edge-spreading Raptor-like (ESRL) codes. Unlike other SC-LDPC codes that adopt the structure of existing rate-compatible LDPC block codes before coupling, ESRL codes maximize the possible locations of edge placement and focus on constructing an optimal coupled matrix. Moreover, a new graph representation called the unified graph is introduced. This graph offers a global perspective on ESRL codes and identifies the optimal edge reallocation to optimize the spreading strategy. We conduct comprehensive comparisons of ESRL codes and 5G-NR LDPC codes. Simulation results demonstrate that when all decoding parameters and complexity are the same, ESRL codes have obvious advantages in error rate performance and throughput compared to 5G-NR LDPC codes in some specific scenarios (low and high number of iterations), making them a promising solution towards next-generation channel coding.
Yuqing Ren, Leyu Zhang, Yifei Shen 0003, Wenqing Song, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Commun.6
2024 A Low-Latency and High-Performance SCL Decoder with Frame-Interleaving
abstract
In this paper, we describe a frame-interleaving hardware architecture for a generalized node-based successive cancellation list (SCL) decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we also exploit graph ensembles to diversify the decoding, enhancing the error-correcting performance by 0.28 dB and reducing the worst-case latency for serial graph processing by over 32%. Implementation results show that the proposed SCL decoder with frame-interleaving architecture achieves a throughput of 7.15 Gbps and an area efficiency of 37.63 Gbps/mm2, which is 1.56× and 1.11× better than the state-of-the-art node-based SCL decoders.
Leyu Zhang, Yuqing Ren, Yifei Shen 0003, Wuyang Zhou, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg
ISCAS5
2024 Learning-based Hand Gesture Classification using Channel Impulse Response with UWB
abstract
The channel impulse response (CIR) of the wireless propagation channel is influenced by the surrounding environment and thus can be used to retrieve environmental information such as location, the presence of objects, and the speed of objects. In this work, we detect hand gestures based on the complex-valued CIR from an ultra-wideband (UWB) transmission link between one transmitter and two receivers. Thanks to the high path delay resolution due to the wide (500 MHz) bandwidth, we can focus on the channel that is influenced by hand gestures in the sensing area. Using different machine learning and deep learning methods, we learn the features from sequences of CIR snapshots that are correlated to hand gestures and recognize them.
Clément Samanos, Han Miao, Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
PIMRC4
2024 Monostatic Multi-Target Wi-Fi-Based Breathing Rate Sensing Using Openwifi
abstract
Continuous monitoring of human activity and vital signs has become increasingly important in healthcare and consumer applications. While wearables and others sensors are already in widespread use, they often suffer from issues such as discomfort and the need for physical contact. To address these limitations, wireless sensing using radio signals has emerged as a promising alternative. However, most previous work has focused on bistatic setups utilizing commodity Wi-Fi devices. In this work, we use the open-source openwifi platform for monostatic multi-antenna Wi-Fi sensing for contactless breathing rate measurements. We also present a method for simultaneous estimation of the breathing rate of two targets. Experimental results using CNC machines for emulating human breathing demonstrate the effectiveness of our method and setup for targets at a few meters distance. Even in challenging scenarios where the breathing rates are similar, the targets are close together, and the targets have the same angle of arrival, we achieve an average error of 0.61 breaths per minute. We further validate our setup and method by collecting human breathing data from two subjects, which can clearly be distinguished using our setup and method. Moreover, our results indicate that self-interference cancellation is not necessary for close-range breathing rate sensing.
Andreas Toftegaard Kristensen, Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
WCNC3
2024 Trends in Channel Coding for 6G
abstract
Error correction coding (i.e., channel coding) is a key ingredient of any digital communications system. In mobile wireless communications, channel codes have evolved from simple convolutional codes in Global System for Mobile Communications (GSM) (2G), parallel concatenated (turbo) codes in Universal Mobile Telecommunications Service (UMTS) (3G), and long-term evolution (LTE) (4G), to carefully designed multirate/multilength low-density parity-check (LDPC) codes in 5G, combined with polar codes for short messages on the synchronization channel. Based on this rich history, and by accounting for the technological advances in very large-scale integration, this article will outline some recent trends in channel coding as they may be applied in 6G systems, ranging from novel approaches for short blocklengths such as automorphism ensemble decoding, via ideas of coding for multiple access, to concepts for unified coding schemes that may simplify encoding/decoding hardware at competitive error-correcting performance.
Sisi Miao, Claus Kestel, Lucas Johannsen, Marvin Geiselhart, Laurent Schmalen, Alexios Balatsoukas-Stimming, Gianluigi Liva, Norbert Wehn, Stephan ten Brink
Proc. IEEE6
2024 A Generalized Adjusted Min-Sum Decoder for 5G LDPC Codes: Algorithm and Implementation
abstract
5G New Radio (NR) has stringent demands on both performance and complexity for the design of low-density parity-check (LDPC) decoding algorithms and corresponding VLSI implementations. Furthermore, decoders must fully support the wide range of all 5G NR blocklengths and code rates, which is a significant challenge. In this paper, we present a high-performance and low-complexity LDPC decoder, tailor-made to fulfill the 5G requirements. First, to close the gap between belief propagation (BP) decoding and its approximations in hardware, we propose an extension of adjusted min-sum decoding, called generalized adjusted min-sum (GA-MS) decoding. This decoding algorithm flexibly truncates the incoming messages at the check node level and carefully approximates the non-linear functions of BP decoding to balance the error-rate and hardware complexity. Numerical results demonstrate that the proposed fixed-point GA-MS has only a minor gap of 0.1 dB compared to floating-point BP under various scenarios of 5G standard specifications. Secondly, we present a fully reconfigurable 5G NR LDPC decoder implementation based on GA-MS decoding. Given that memory occupies a substantial portion of the decoder area, we adopt multiple data compression and approximation techniques to reduce 42.2% of the memory overhead. The corresponding 28nm FD-SOI ASIC decoder has a core area of 1.823 mm$^{2}$and operates at 895 MHz. It is compatible with all 5G NR LDPC codes and achieves a peak throughput of 24.42 Gbps and a maximum area efficiency of 13.40 Gbps/mm$^{2}$at 4 decoding iterations.
Yuqing Ren, Yifei Shen 0003, Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 A Node-Based Polar List Decoder With Frame Interleaving and Ensemble Decoding Support
abstract
Node-based successive cancellation list (SCL) decoding has received considerable attention in wireless communications for its significant reduction in decoding latency, particularly with 5G New Radio (NR) polar codes. However, the existing node-based SCL decoders are constrained by sequential processing, leading to complicated and data-dependent computational units that introduce unavoidable stalls, reducing hardware efficiency. In this paper, we present a frame-interleaving hardware architecture for a generalized node-based SCL decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we further exploit graph ensembles to diversify the decoding space, thus enhancing the error-correcting performance with a limited list size. Two dynamic strategies are proposed to eliminate the residual stalls in the decoding schedule, which eventually results in nearly$2 \times $throughput compared to the state-of-the-art baseline node-based SCL decoder. To impart the decoder rate flexibility, we develop a novel online instruction generator to identify the generalized nodes and produce instructions on-the-fly. The corresponding 28nm FD-SOI ASIC SCL decoder with a list size of 8 has a core area of 1.28 mm2 and operates at 692 MHz. It is compatible with all 5G NR polar codes and achieves a throughput of 3.34 Gbps and an area efficiency of 2.62 Gbps/mm2 for uplink (1024, 512) codes, which is$1.41 \times $and$1.69 \times $better than the state-of-the-art node-based SCL decoders.
Yuqing Ren, Leyu Zhang, Ludovic Damien Blanc, Yifei Shen 0003, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 Recursive/Iterative Unique Projection-Aggregation Decoding of Reed-Muller Codes
abstract
We describe recursive unique projection-aggregation (RUPA) decoding and iterative unique projection-aggregation (IUPA) decoding of Reed-Muller (RM) codes, which remove non-unique projections from the recursive projection-aggregation (RPA) and iterative projection-aggregation (IPA) algorithms respectively. We show that these algorithms have competitive error-correcting performance while requiring up to 95% projections lower than the baseline RPA algorithm.
Marzieh Hashemipour-Nazari, Renate Debets, Kees Goossens, Alexios Balatsoukas-Stimming
ICASSP4
2023 Single-Anchor UWB Localization Using Channel Impulse Response Distributions
abstract
Ultra-wideband (UWB) devices are widely used in indoor localization scenarios. Single-anchor UWB localization shows advantages because of its simple system setup compared to conventional two-way ranging (TWR) and trilateration localization methods. In this work, we focus on single-anchor UWB localization methods that learn statistical features of the channel impulse response (CIR) in different location areas using a Gaussian mixture model (GMM). We show that by learning the joint distributions of the amplitudes of different delay components, we achieve a more accurate location estimate compared to considering each delay bin independently. Moreover, we develop a similarity metric between sets of CIRs. With this set-based similarity metric, we can further improve the estimation performance, compared to treating each snapshot separately. We showcase the advantages of the proposed methods in multiple application scenarios.
Sitian Li, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ICASSP2
2023 Successive-Cancellation Flip Decoding of Polar Codes with a Simplified Restart Mechanism
abstract
Polar codes are a class of error-correcting codes that provably achieve the capacity of practical channels. The successive-cancellation flip (SCF) decoder is a low-complexity decoder that was proposed to improve the performance of the successive-cancellation (SC) decoder as an alternative to the high-complexity successive-cancellation list (SCL) decoder. The SCF decoder improves the error-correction performance of the SC decoder, but the variable execution time and the high worst-case execution time pose a challenge for the realization of receivers with fixed-time algorithms. The dynamic SCF (DSCF) variation of the SCF decoder further improves the error-correction performance but the challenge of decoding delay remains. In this work, we propose a simplified restart mechanism (SRM) that reduces the execution time of SCF and DSCF decoders through conditional restart of the additional trials from the second half of the codeword. We show that the proposed mechanism is able to improve the execution time characteristics of SCF and DSCF decoders while providing identical error-correction performance. For a DSCF decoder that can flip up to 3 simultaneous bits per decoding trial, the average execution time, the average additional execution time and the execution-time variance are reduced by approximately 31%, 37% and 57%, respectively. For this setup, the mechanism requires approximately 3.9% additional memory.
Ilshat Sagitov, Charles Pillet, Alexios Balatsoukas-Stimming, Pascal Giard
WCNC3
2023 Pipelined Architecture for Soft-Decision Iterative Projection Aggregation Decoding for RM Codes
abstract
The recently proposed recursive projection-aggregation (RPA) decoding algorithm for Reed-Muller codes has received significant attention as it provides near-ML decoding performance at reasonable complexity for short codes. However, its complicated structure makes it unsuitable for hardware implementation. Iterative projection-aggregation (IPA) decoding is a modified version of RPA decoding that simplifies the hardware implementation. In this work, we present a flexible hardware architecture for the IPA decoder that can be configured from fully-sequential to fully-parallel, thus making it suitable for a wide range of applications with different constraints and resource budgets. Our simulation and implementation results show that the IPA decoder has 41% lower area consumption, 44% lower latency, four times higher throughput, but currently seven times higher power consumption for a code with block length of 128 and information length of 29 compared to a state-of-the-art polar successive cancellation list (SCL) decoder with comparable decoding performance.
Marzieh Hashemipour-Nazari, Yuqing Ren, Kees Goossens, Alexios Balatsoukas-Stimming
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Fast Sequence Repetition Node-Based Successive Cancellation List Decoding for Polar Codes
abstract
Compared with the bit-wise successive cancellation list (SCL) decoding of polar codes, the node-based Fast SCL decoding significantly reduces the decoding latency by identifying special constituent codes and decoding these in parallel. To further reduce the latency of current Fast SCL decoders, we first propose a fast sequence repetition (SR) node-based SCL (Fast SR-SCL) decoding algorithm, which only involves one type of node in the SCL decoding tree. Furthermore, we employ the adaptive path splitting (APS) strategy to terminate the path splitting in the SR node early, without degrading the error-correcting performance. Numerical results show that for 5G uplink codes with a length of 1024 and rates of 1/4, 1/2, and 3/4, our decoder can deliver the same decoding performance while reducing the average latency by 34.5%, 38.0%, and 39.6% compared with the state-of-the-art Fast SCL decoder for a list size L = 8.
Yifei Shen 0003, Yuqing Ren, Andreas Toftegaard Kristensen, Alexios Balatsoukas-Stimming, Xiaohu You 0001, Chuan Zhang 0001, Andreas Peter Burg
ICC4
2022 A Maximum-Likelihood-Based Two-User Receiver for LoRa Chirp Spread-Spectrum Modulation
abstract
Long Range (LoRa) is an emerging low-power wide-area network technology offering long-range wireless connectivity to Internet of Things (IoT) devices. For energy efficiency reasons, LoRa end nodes implement a nonslotted ALOHA multiple access scheme to transmit packets to the gateway. Due to the lack of synchronization between end nodes, collisions between uplink packets have been identified as the main obstacle to the scaling of dense LoRa networks. To tackle this issue, we present in this article a LoRa receiver that is capable of decoding colliding packets from two interfering end nodes. The proposed two-user detector is derived from the maximum-likelihood principle using a detailed model of two colliding LoRa packets. As the complexity of the maximum-likelihood sequence estimation is prohibitive, complexity-reduction techniques are introduced to enable practical implementations of the receiver. An in-depth performance analysis highlights that the proposed two-user detector inherently leverages the differences in received power, time offsets, and frequency offsets between the users to separate and demodulate their respective signals. To demonstrate the practicality of the proposed detector, an interference-robust synchronization algorithm is then designed and evaluated. Simulation results indicate that a LoRa receiver combining the proposed synchronization algorithm and two-user detector is capable of detecting and demodulating two interfering users with satisfactorily error rates.
Mathieu Xhonneux, Joachim Tapparel, Alexios Balatsoukas-Stimming, Andreas Peter Burg, Orion Afisiadis
IEEE Internet Things J.3
2021 Hardware Implementation of Iterative Projection-Aggregation Decoding of Reed-Muller Codes
abstract
In this work, we present a simplification and a corresponding hardware architecture for hard-decision recursive projection-aggregation (RPA) decoding of Reed-Muller (RM) codes. In particular, we transform the recursive structure of RPA decoding into a simpler and iterative structure with minimal error-correction degradation. Our simulation results for RM(7,3) show that the proposed simplification has a small error-correcting performance degradation (0.005 in terms of channel crossover probability) while reducing the average number of computations by up to 40%. In addition, we describe the first fully parallel hardware architecture for simplified RPA decoding. We present FPGA implementation results for an RM(6,3) code on a Xilinx Virtex-7 FPGA showing that our proposed architecture achieves a throughput of 171 Mbps at a frequency of 80 MHz.
Marzieh Hashemipour-Nazari, Kees Goossens, Alexios Balatsoukas-Stimming
ICASSP3
2021 Towards Practical Near-Maximum-Likelihood Decoding of Error-Correcting Codes: An Overview
abstract
While in the past several decades the trend to go towards increasing error-correcting code lengths was predominant to get closer to the Shannon limit, applications that require short block length are developing. Therefore, decoding techniques that can achieve near-maximum-likelihood (near-ML) are gaining momentum. This overview paper surveys recent progress in this emerging field by reviewing the GRAND algorithm, linear programming decoding, machine-learning aided decoding and the recursive projection-aggregation decoding algorithm. For each of the decoding algorithms, both algorithmic and hardware implementations are considered, and future research directions are outlined.
Thibaud Tonnellier, Marzieh Hashemipour-Nazari, Nghia Doan, Warren J. Gross, Alexios Balatsoukas-Stimming
ICASSP5
2021 OFDM-Based Beam-Oriented Digital Predistortion for Massive MIMO
abstract
Linearization of massive MIMO arrays is a significant computational challenge that typically scales with the number of antennas. In this work, we introduce a beam-oriented digital predistortion (DPD) scheme for OFDM-based massive MIMO systems that applies predistortion before the precoder in the OFDM guard-band subcarriers. Using simulation results, we show that, for a 64 antenna massive MIMO array, our proposed method can achieve the same DPD performance as a conventional DPD method while requiring an order of magnitude fewer multiplications.
Chance Tarver, Alexios Balatsoukas-Stimming, Christoph Studer, Joseph R. Cavallaro
ISCAS2
2021 On the Advantage of Coherent LoRa Detection in the Presence of Interference
abstract
It has been shown that the coherent detection of long range (LoRa) signals only provides marginal gains of around 0.7 dB on the additive white Gaussian noise (AWGN) channel. However, ALOHA-based massive Internet-of-Things systems, including LoRa, often operate in the interference-limited regime. Therefore, in this work, we examine the performance of the LoRa modulation with coherent detection in the presence of interference from another LoRa user with the same spreading factor. We derive rigorous symbol- and frame error rate (FER) expressions as well as bounds and approximations for evaluating the error rates. The error rates predicted by these approximations are compared against error rates found by Monte Carlo simulations and shown to be very accurate. We also compare the performance of LoRa with coherent and noncoherent receivers and we show that the coherent detection of LoRa is significantly more beneficial in interference scenarios than in the presence of only AWGN. For example, we show that coherent detection leads to a 2.5-dB gain over the standard noncoherent detection for a signal-to-interference ratio (SIR) of 3 dB and up to a 10-dB gain for an SIR of 0 dB. Moreover, we show that with coherent detection it is easier to obtain useful and relevant FER values even for negative SIR values.
Orion Afisiadis, Sitian Li, Joachim Tapparel, Andreas Peter Burg, Alexios Balatsoukas-Stimming
IEEE Internet Things J.5
2021 Threshold-Based Fast Successive-Cancellation Decoding of Polar Codes
abstract
Fast SC decoding overcomes the latency caused by the serial nature of the SC decoding by identifying new nodes in the upper levels of the SC decoding tree and implementing their fast parallel decoders. In this work, we first present a novel sequence repetition node corresponding to a particular class of bit sequences. Most existing special node types are special cases of the proposed sequence repetition node. Then, a fast parallel decoder is proposed for this class of node. To further speed up the decoding process of general nodes outside this class, a threshold-based hard-decision-aided scheme is introduced. The threshold value that guarantees a given error-correction performance in the proposed scheme is derived theoretically. Analysis and hardware implementation results on a polar code of length 1024 with code rates 1/4, 1/2, and 3/4 show that our proposed algorithm reduces the required clock cycles by up to 8%, and leads to a 10% improvement in the maximum operating frequency compared to state-of-the-art decoders without tangibly altering the error-correction performance. In addition, using the proposed threshold-based hard-decision-aided scheme, the decoding latency can be further reduced by 57% at Eb/N0= 5.0 dB.
Seyyed Ali Hashemi, Alexios Balatsoukas-Stimming, Zizheng Cao, Antonius M. J. Koonen, John M. Cioffi, Andrea J. Goldsmith
IEEE Trans. Commun.3
2020 Lupulus: A Flexible Hardware Accelerator for Neural Networks
abstract
Neural networks have become indispensable for a wide range of applications, but they suffer from high computationaland memory-requirements, requiring optimizations from the algorithmic description of the network to the hardware implementation. Moreover, the high rate of innovation in machine learning makes it important that hardware implementations provide a high level of programmability to support current and future requirements of neural networks. In this work, we present a flexible hardware accelerator for neural networks, called Lupulus, supporting various methods for scheduling and mapping of operations onto the accelerator. Lupulus was implemented in a 28nm FD-SOI technology and demonstrates a peak performance of 380GOPS/GHz with latencies of 21.4ms and 183.6ms for the convolutional layers of AlexNet and VGG-16, respectively.
Andreas Toftegaard Kristensen, Robert Giterman, Alexios Balatsoukas-Stimming, Andreas Peter Burg
ICASSP3
2020 Coded LoRa Frame Error Rate Analysis
abstract
In this work, we study the coded frame error rate (FER) of LoRa under additive white Gaussian noise (AWGN) and under carrier frequency offset (CFO). To this end, we use existing approximations for the bit error rate (BER) of the LoRa modulation under AWGN and we present a FER analysis that includes the channel coding, interleaving, and Gray mapping of the LoRa physical layer. We also derive the LoRa BER under carrier frequency offset and we present a corresponding FER analysis. We compare the derived frame error rate expressions to Monte Carlo simulations to verify their accuracy.
Orion Afisiadis, Andreas Peter Burg, Alexios Balatsoukas-Stimming
ICC3
2020 OPTCOMNET: Optimized Neural Networks for Low-Complexity Channel Estimation
abstract
The use of machine learning methods to tackle challenging physical layer signal processing tasks has attracted significant attention. In this work, we focus on the use of neural networks (NNs) to perform pilot-assisted channel estimation in an OFDM system in order to avoid the challenging task of estimating the channel covariance matrix. In particular, we perform a systematic design-space exploration of NN configurations, quantization, and pruning in order to improve feedforward NN architectures that are typically used in the literature for the channel estimation task. We show that choosing an appropriate NN architecture is crucial to reduce the complexity of NN-assisted channel estimation methods. Moreover, we demonstrate that, similarly to other applications and domains, careful quantization and pruning can lead to significant complexity reduction with a negligible performance degradation. Finally, we show that using a solution with multiple distinct NNs trained for different signal-to-noise ratios interestingly leads to lower overall computational complexity and storage requirements, while achieving a better performance with respect to using a single NN trained for the entire SNR range.
Michel van Lier, Alexios Balatsoukas-Stimming, Henk Corporaal, Zoran Zivkovic
ICC2
2020 On the Error Rate of the LoRa Modulation With Interference
abstract
LoRa is a chirp spread-spectrum modulation developed for the Internet of Things (IoT). In this work, we examine the performance of LoRa in the presence of both additive white Gaussian noise and interference from another LoRa user. To this end, we extend an existing interference model, which assumes perfect alignment of the signal of interest and the interference, to the more realistic case where the interfering user is neither chip- nor phase-aligned with the signal of interest and we derive an expression for the error rate. We show that the existing aligned interference model overestimates the effect of interference on the error rate. Moreover, we prove two symmetries in the interfering signal and we derive low-complexity approximate formulas that can significantly reduce the complexity of computing the symbol and frame error rates compared to the complete expression. Finally, we provide numerical simulations to corroborate the theoretical analysis and to verify the accuracy of our proposed approximations.
Orion Afisiadis, Matthieu Cotting, Andreas Peter Burg, Alexios Balatsoukas-Stimming
IEEE Trans. Wirel. Commun.4
2019 A Lyra2 FPGA Core for Lyra2REv2-Based Cryptocurrencies
abstract
Lyra2REv2 is a hashing algorithm that consists of a chain of individual hashing algorithms and it is used as a proof-of-work function in several cryptocurrencies that aim to be ASIC-resistant. The most crucial hashing algorithm in the Lyra2REv2 chain is a specific instance of the general Lyra2 algorithm. In this work we present the first FPGA implementation of the aforementioned instance of Lyra2 and we explain how several properties of the algorithm can be exploited in order to optimize the design.
Michiel Van Beirendonck, Louis-Charles Trudeau, Pascal Giard, Alexios Balatsoukas-Stimming
ISCAS4
2019 On the Computational Complexity of Blind Detection of Binary Linear Codes
abstract
In this work, we study the computational complexity of the MINIMUM DISTANCE CODE DETECTION problem. In this problem, we are given a set of noisy codeword observations and we wish to find a code in a set of linear codes C of a given dimension k, for which the sum of distances between the observations and the code is minimized. We prove that, for the practically relevant case when the set C only contains a fixed number of candidate linear codes, the detection problem is NPhard and we identify a number of interesting open questions related to the code detection problem.
Alexios Balatsoukas-Stimming, Aris Filos-Ratsikas
ISIT1
2019 LDPC Coded Multiuser Shaping for the Gaussian Multiple Access Channel
abstract
The joint design of input constellation and low-density parity-check (LDPC) codes to approach the symmetric capacity of the two-user Gaussian multiple access channel is studied. More specifically, multilevel coding is employed at each user to construct a high-order input constellation and the constellations of the users are jointly designed so as to maximize the multiuser shaping gain. At the receiver, each layer of the multilevel coding is jointly decoded among users, while successive cancellation is employed across layers. The LDPC code employed by each user in each layer is designed using EXIT charts to support joint decoding among users for the prescribed per-layer rate and SNR. Numerical simulations are provided to validate the proposed constellation and LDPC code designs.
Alexios Balatsoukas-Stimming, Stefano Rini, Jörg Kliewer
ISIT1
2018 Faulty Successive Cancellation Decoding of Polar Codes for the Binary Erasure Channel
abstract
In this paper, the faulty successive cancellation decoding of polar codes for the binary erasure channel is studied. To this end, a simple erasure-based fault model is introduced to represent errors in the decoder, and it is shown that, under this model, polarization does not happen, meaning that fully reliable communication is not possible at any rate. Furthermore, a lower bound on the frame error rate of polar codes under faulty successive cancellation decoding is provided, which is then used, along with a well-known upper bound, in order to choose a blocklength that minimizes the erasure probability under faulty decoding. Finally, an unequal error protection scheme that can re-enable asymptotically erasure-free transmission at a small rate loss and by protecting only a constant fraction of the decoder is proposed. The same scheme is also shown to significantly improve the finite-length performance of the faulty successive cancellation decoder by protecting as little as 1.5% of the decoder.
Alexios Balatsoukas-Stimming, Andreas Peter Burg
IEEE Trans. Commun.1
2018 A 588-Gb/s LDPC Decoder Based on Finite-Alphabet Message Passing
abstract
An ultrahigh throughput low-density parity-check (LDPC) decoder with an unrolled full-parallel architecture is proposed, which achieves the highest decoding throughput compared to previously reported LDPC decoders in the literature. The decoder benefits from a serial message-transfer approach between the decoding stages to alleviate the well-known routing congestion problem in parallel LDPC decoders. Furthermore, a finite-alphabet message passing algorithm is employed to replace the VN update rule of the standard min-sum (MS) decoder with lookup tables, which are designed in a way that maximizes the mutual information between decoding messages. The proposed algorithm results in an architecture with reduced bit-width messages, leading to a significantly higher decoding throughput and to a lower area compared to an MS decoder when serial message transfer is used. The architecture is placed and routed for the standard MS reference decoder and for the proposed finite-alphabet decoder using a custom pseudo-hierarchical backend design strategy to further alleviate routing congestions and to handle the large design. Postlayout results show that the finite-alphabet decoder with the serial message-transfer architecture achieves a throughput as large as 588 Gb/s with an area of 16.2 mm2and dissipates an average power of 22.7 pJ per decoded bit in a 28-nm fully depleted silicon on isulator library. Compared to the reference MS decoder, this corresponds to 3.1 times smaller area and 2 times better energy efficiency.
Reza Ghanaatian, Alexios Balatsoukas-Stimming, Thomas Christoph Müller, Michael Meidlinger, Gerald Matz, Adam Teman, Andreas Peter Burg
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Partitioned successive-cancellation list decoding of polar codes
abstract
Successive-cancellation list (SCL) decoding is an algorithm that provides very good error-correction performance for polar codes. However, its hardware implementation requires a large amount of memory, mainly to store intermediate results. In this paper, a partitioned SCL algorithm is proposed to reduce the large memory requirements of the conventional SCL algorithm. The decoder tree is broken into partitions that are decoded separately. We show that with careful selection of list sizes and number of partitions, the proposed algorithm can outperform conventional SCL while requiring less memory.
Seyyed Ali Hashemi, Alexios Balatsoukas-Stimming, Pascal Giard, Claude Thibeault, Warren J. Gross
ICASSP2
2016 High-throughput lattice reduction for large-scale MIMO systems based on Seysen's algorithm
abstract
Lattice reduction aided multiple-input and multiple output (MIMO) detection has attracted significant attention recently due to its low complexity and excellent error rate performance. The most popular lattice reduction algorithms are the Lenstra-Lenstra-Lovász algorithm and Seysen's algorithm, although the former has received much more attention than the latter. In this work, we present a simplification to Seysen's lattice reduction algorithm, which reduces the periteration computational complexity from quadratic to linear with a small degradation in the quality of the resulting reduced lattice. Moreover, we present an efficient VLSI architecture which demonstrates the advantages of the proposed algorithm and can achieve a throughput of up to 91 Mmatrices/s at an operating frequency of 1 GHz.
Farhana Sheikh, Alexios Balatsoukas-Stimming
ICC2
2016 Hardware decoders for polar codes: An overview
abstract
Polar codes are an exciting new class of error correcting codes that achieve the symmetric capacity of memoryless channels. Many decoding algorithms were developed and implemented, addressing various application requirements: from error-correction performance rivaling that of LDPC codes to very high throughput or low-complexity decoders. In this work, we review the state of the art in polar decoders implementing the successive-cancellation, belief propagation, and list decoding algorithms, illustrating their advantages.
Pascal Giard, Gabi Sarkis, Alexios Balatsoukas-Stimming, YouZhe Fan, Chi-Ying Tsui, Andreas Peter Burg, Claude Thibeault, Warren J. Gross
ISCAS3
2016 Sliding Window Spectrum Sensing for Full-Duplex Cognitive Radios with Low Access-Latency
abstract
In a cognitive radio system the failure of secondary user (SU) transceivers to promptly vacate the channel can introduce significant access-latency for primary or high-priority users (PU). In conventional cognitive radio systems, the backoff latency is exacerbated by frame structures that only allow sensing at periodic intervals. Concurrent transmission and sensing using self-interference suppression has been suggested to improve the performance of cognitive radio systems, allowing decisions to be taken at multiple points within the frame. In this paper, we extend this approach by proposing a sliding-window full-duplex model allowing decisions to be taken on a sample-by-sample basis. We also derive the access-latency for both the existing and the proposed schemes. Our results show that the access-latency of the sliding scheme is decreased by a factor of 2.6 compared to the existing slotted full-duplex scheme and by a factor of approximately 16 compared to a half-duplex cognitive radio system. Moreover, the proposed scheme is significantly more resilient to the destructive effects of residual self-interference compared to previous approaches.
Orion Afisiadis, Andrew Austin 0001, Alexios Balatsoukas-Stimming, Andreas Peter Burg
VTC Spring3
2015 On metric sorting for successive cancellation list decoding of polar codes
abstract
We focus on the metric sorter unit of successive cancellation list decoders for polar codes, which lies on the critical path in all current hardware implementations of the decoder. We review existing metric sorter architectures and we propose two new architectures that exploit the structure of the path metrics in a log-likelihood ratio based formulation of successive cancellation list decoding. Our synthesis results show that, for the list size of L = 32, our first proposed sorter is 14% faster and 45% smaller than existing sorters, while for smaller list sizes, our second sorter has a higher delay in return for up to 36% reduction in the area.
Alexios Balatsoukas-Stimming, Mani Bastani Parizi, Andreas Peter Burg
ISCAS1
2015 Demo: Concurrent Spectrum Sensing and Transmission for Cognitive Radio using Self-Interference Cancellation
abstract
A demonstration of a cognitive radio network that supports concurrent spectrum sensing and transmission is presented. To detect primary users while transmitting, the secondary user node must suppress the self-interference signals. The system is implemented on a National Instruments Universal Software Radio Peripheral (USRP) platform. The demonstration will show that continuous spectrum sensing avoids the overhead for dedicated sensing periods and can detect the primary user within at most 10~ms from the start of transmission.
Andrew Austin 0001, Orion Afisiadis, Alexios Balatsoukas-Stimming, Andreas Peter Burg
MobiHoc3
2014 LLR-based successive cancellation list decoding of polar codes
abstract
We present an LLR-based implementation of the successive cancellation list (SCL) decoder. To this end, we associate each decoding path with a metric which (i) is a monotone function of the path's likelihood and (ii) can be computed efficiently from the channel LLRs. The LLR-based formulation leads to a more efficient hardware implementation of the decoder compared to the known log-likelihood based implementation. Synthesis results for an SCL decoder with block-length of N = 1024 and list sizes of L = 2 and L = 4 confirm that the LLR-based decoder has considerable area and operating frequency advantages in the orders of 50% and 30%, respectively.
Alexios Balatsoukas-Stimming, Mani Bastani Parizi, Andreas Peter Burg
ICASSP1
2014 Enabling complexity-performance trade-offs for successive cancellation decoding of polar codes
abstract
Polar codes are one of the most recent advancements in coding theory and they have attracted significant interest. While they are provably capacity achieving over various channels, they have seen limited practical applications. Unfortunately, the successive nature of successive cancellation based decoders hinders fine-grained adaptation of the decoding complexity to design constraints and operating conditions. In this paper, we propose a systematic method for enabling complexity-performance tradeoffs by constructing polar codes based on an optimization problem which minimizes the complexity under a suitably defined mutual information based performance constraint. Moreover, a low-complexity greedy algorithm is proposed in order to solve the optimization problem efficiently for very large code lengths.
Alexios Balatsoukas-Stimming, Georgios Karakonstantis, Andreas Peter Burg
ISIT1
2014 Faulty successive cancellation decoding of polar codes for the binary erasure channel
Alexios Balatsoukas-Stimming, Andreas Peter Burg
ISITA1
2012 FPGA-based design and implementation of a multi-GBPS LDPC decoder
abstract
We design a very high speed LDPC code decoder architecture for (3,6)-regular codes by employing hybrid quantization, pipelining, and FPGA-specific optimizations. Our pipelined architecture fully addresses the decoder's significant I/O requirements, even when an early termination circuit is employed. The proposed decoder can achieve a throughput of up to 16.9 Gbps at an Eb/N0 of 3.5 dB using a code of length 1152, running at a clock speed of 153 MHz and performing a maximum of 10 decoding iterations, thus out-performing the state of the art by a significant margin. This design was fully implemented and tested on a Xilinx Virtex 5 XC5VLX110 FPGA. We also present an alternative, low-complexity design, which is able to achieve a throughput of up to 21.6 Gbps by sacrifing 0.75 dB in terms of Eb/N0.
Alexios Balatsoukas-Stimming, Apostolos Dollas
FPL1