VLDB 2026 Research / reviewers in the wild / expert
Carlo Condo
dblp:97/9672
· DBLP profile ↗
25ranked-venue papers
10as first author
2since 2021 · last 2022
0000-0002-3050-036XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 1 since 2021Computer networks · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 54% Hardware reliability and fault tolerance · 17% Energy-efficient computing · 15% | |
| Theoretical computer science
2 papers |
Coding theory · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Coding theory › channel coding
polar codes |
0.7 | 2 | 2019 | Improved Bit-Flipping Algorithm for Successive Cancellation Decoding of Polar Codes · IEEE Trans. Commun. 2019 Decoder Partitioning: Towards Practical List Decoding of Polar Codes · IEEE Trans. Commun. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.4 | 1 | 2020 | Fast and Efficient Convolutional Accelerator for Edge Computing · IEEE Trans. Computers 2020 |
Processor architecture and microarchitecture
dataflow architecture |
0.4 | 1 | 2020 | Fast and Efficient Convolutional Accelerator for Edge Computing · IEEE Trans. Computers 2020 |
Hardware accelerators and domain-specific architectures
edge accelerator |
0.4 | 1 | 2020 | Fast and Efficient Convolutional Accelerator for Edge Computing · IEEE Trans. Computers 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | Fast and Efficient Convolutional Accelerator for Edge Computing · IEEE Trans. Computers 2020 |
Energy-efficient computing
power management |
0.4 | 1 | 2020 | Fast and Efficient Convolutional Accelerator for Edge Computing · IEEE Trans. Computers 2020 |
Coding theory › error-correcting codes › decoding › iterative decoding › iterative hard-decision decoding
bit-flipping decoding |
0.4 | 1 | 2019 | Improved Bit-Flipping Algorithm for Successive Cancellation Decoding of Polar Codes · IEEE Trans. Commun. 2019 |
Coding theory
channel coding |
0.4 | 1 | 2019 | Improved Bit-Flipping Algorithm for Successive Cancellation Decoding of Polar Codes · IEEE Trans. Commun. 2019 |
Coding theory › channel coding › polar codes
successive cancellation decoding |
0.4 | 1 | 2019 | Improved Bit-Flipping Algorithm for Successive Cancellation Decoding of Polar Codes · IEEE Trans. Commun. 2019 |
Coding theory › error-correcting codes › decoding
list decoding |
0.3 | 1 | 2018 | Decoder Partitioning: Towards Practical List Decoding of Polar Codes · IEEE Trans. Commun. 2018 |
Coding theory › error-correcting codes › decoding › list decoding
successive cancellation list decoding |
0.3 | 1 | 2018 | Decoder Partitioning: Towards Practical List Decoding of Polar Codes · IEEE Trans. Commun. 2018 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2017 | Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks · ICLR (Poster) 2017 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.3 | 1 | 2017 | Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks · ICLR (Poster) 2017 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.3 | 1 | 2017 | Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks · ICLR (Poster) 2017 |
Hardware reliability and fault tolerance
error-correcting codes for memory |
0.2 | 1 | 2015 | Unequal Error Protection of Memories in LDPC Decoders · IEEE Trans. Computers 2015 |
Hardware reliability and fault tolerance › error correction
error correction decoder |
0.2 | 1 | 2015 | Unequal Error Protection of Memories in LDPC Decoders · IEEE Trans. Computers 2015 |
Coding theory › decoder design
decoder implementation |
0.1 | 1 | 2018 | Decoder Partitioning: Towards Practical List Decoding of Polar Codes · IEEE Trans. Commun. 2018 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2015 | Unequal Error Protection of Memories in LDPC Decoders · IEEE Trans. Computers 2015 |
Methods — techniques the papers use, named apart from their topics
sparse connectivity · 0.6VLSI design · 0.6thresholding · 0.4log-likelihood ratio analysis · 0.4lower bound analysis · 0.3cyclic redundancy check allocation · 0.3unequal error protection coding · 0.2LDPC decoding · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Fixed Latency ORBGRAND Decoder Architecture With LUT-Aided Error-Pattern SchedulingabstractGuessing Random Additive Noise Decoding (GRAND) is a universal decoding algorithm that has been recently proposed as a practical way to perform maximum likelihood decoding. It generates a sequence of possible error patterns and applies them to the received vector, checking if the result is a valid codeword. Ordered reliability bits GRAND (ORBGRAND) improves on GRAND by considering soft information received from the channel. Both GRAND and ORBGRAND have been implemented in hardware, focusing on average performance, sacrificing worst case throughput and latency. In this work, an improved pattern schedule for ORBGRAND is proposed. It provides$> 0.5$dB gain over the standard schedule at a block error rate$\le 10^{-5}$, and outperforms more complex GRAND flavors with a fraction of the complexity. The proposed schedule is used within a novel code-agnositic decoder architecture: the decoder guarantees fixed high throughput and low latency, making it attractive for latency-constrained applications. It outperforms the worst-case performance of decoders by orders of magnitude, and outperforms many best-case figures. Decoding a code of length 128, it achieves a throughput of 79.21 Gb/s with 58.49 ns latency, yielding better energy efficiency and comparable area efficiency with respect to the state of the art. Carlo Condo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Sliding Window Polar CodesabstractWe propose a novel coupling technique for polar codes via a special kernel that enables efficient sliding window decoding. This feature allows to reduce the memory requirement of the decoder, an important possibility in wireless communication downlink scenarios. Our approach is based on the design of an ad-hoc kernel to be inserted in a multi-kernel polar code framework. Simulation results show that the proposed sliding window polar codes outperform polar codes transmitted in independent blocks at a negligible additional decoding overhead. Valerio Bioglio, Carlo Condo, Ingmar Land |
ISIT | 2 |
| 2020 | SCAN List Decoding of Polar CodesabstractIn this paper we propose an enhanced soft cancellation (SCAN) decoder for polar codes based on decoding stages permutation. The proposed soft cancellation list (SCANL) decoder runs L independent SCAN decoders, each one relying on a different permuted factor graph. The estimated bits are selected among the L candidates through a dedicated metric provided by the decoders. Furthermore, we introduce an early-termination scheme reducing decoding latency without affecting error correction performance. We investigate the error-correction performance of the proposed scheme under various combinations of number of iterations used, permutation set and early-termination condition. Simulation results show that the proposed SCANL provides similar results when compared with belief propagation list, while having a smaller complexity. Moreover, for large list sizes, SCANL outperforms non-CRC aided successive cancellation list decoding. Charles Pillet, Carlo Condo, Valerio Bioglio |
ICC | 2 |
| 2020 | On List Decoding of 5G-NR Polar CodesabstractThe 5thgeneration wireless systems (5G) standardization process of the 3rdgeneration partnership project (3GPP) chose polar codes as a channel coding scheme for the control channel. In case of downlink control information, polar codes are concatenated with distributed distributed cyclic redundancy check (CRC). Whereas CRC bits allow to improve the performance of successive cancellation list (SCL) decoders by improving distance properties, distributed CRC bits allow for path pruning and decoding early-termination. In this paper, we show how to take advantage of the distributed CRC to improve SCL decoding, analyzing various schemes having different early-termination and error correction properties. Simulation results compare the proposed decoding schemes, showing different tradeoffs between error-correction performance and early-termination with different decoder parameters. Charles Pillet, Valerio Bioglio, Carlo Condo |
WCNC | 3 |
| 2020 | Fast and Efficient Convolutional Accelerator for Edge ComputingabstractConvolutional neural networks (CNNs) are a vital approach in machine learning. However, their high complexity and energy consumption make them challenging to embed in mobile applications at the edge requiring real-time processes such as smart phones. In order to meet the real-time constraint of edge devices, recently proposed custom hardware CNN accelerators have exploited parallel processing elements (PEs) to increase throughput. However, this straightforward parallelization of PEs and high memory bandwidth require high data movement, leading to large energy consumption. As a result, only a certain number of PEs can be instantiated when designing bandwidth-limited custom accelerators targeting edge devices. While most bandwidth-limited designs claim a peak performance of a few hundred giga operations per second, their average runtime performance is substantially lower than their roofline when applied to state-of-the-art CNNs such as AlexNet, VGGNet and ResNet, as a result of low resource utilization and arithmetic intensity. In this work, we propose a zero-activation-skipping convolutional accelerator (ZASCA) that avoids noncontributory multiplications with zero-valued activations. ZASCA employs a dataflow that minimizes the gap between its average and peak performances while maximizing its arithmetic intensity for both sparse and dense representations of activations, targeting the bandwidth-limited edge computing scenario. More precisely, ZASCA achieves a performance efficiency of up to 94 percent over a set of state-of-the-art CNNs for image classification with dense representation where the performance efficiency is the ratio between the average runtime performance and the peak performance. Using its zero-skipping feature, ZASCA can further improve the performance efficiency of the state-of-the-art CNNs by up to 1.9× depending on the sparsity degree of activations. The implementation results in 65-nm TSMC CMOS technology show that, compared to the most energy-efficient accelerator, ZASCA can process convolutions from 5.5× to 17.5× faster, and is between 2.1× and 4.5× more energy efficient while occupying 2.1× less silicon area. Arash Ardakani, Carlo Condo, Warren J. Gross |
IEEE Trans. Computers | 2 |
| 2019 | Improved Hybrid Design of Polar Codes and Multi-Kernel Polar CodesabstractIn this paper we propose a novel frozen set design for polar codes and multi-kernel polar codes. We improve the existing hybrid distance-reliability design by minimizing the upper bound of the overall system error probability instead of minimizing its lower bound as previously proposed. This allows to better trade reliabilities of the input bits against distance properties of the code. We describe the new design approach, propose a greedy algorithm to limit the complexity of the code construction process, and evaluate its performance through numerical examples. In both MK polar codes and conventional polar codes, a substantial performance improvement is observed, matching the performance of CRC-aided polar codes under SCL without the need for a CRC. Valerio Bioglio, Ingmar Land, Carlo Condo |
ISIT | 3 |
| 2019 | SC-Flip Decoding of Polar Codes with High Order Error Correction Based on Error DependencyabstractThe successive cancellation flip (SC-Flip) decoding algorithm arose as a valid low-complexity decoding algorithm for polar codes, however its decoding capabilities are still far away from list based decoders. In this paper, we propose an improved SC-Flip multiple error decoding framework based on error dependency, and specialize it for the two-error correction case. We propose to generate different lists of second error locations based on the index of the expected first errors. The inherent flexibility of this approach allows it to be modified for a desired trade-off between performance and complexity. Two second errors list construction approaches are presented, and are shown to yield gains over the original SC-Flip decoder at the same decoding complexity. Carlo Condo, Valerio Bioglio, Ingmar Land |
ITW | 1 |
| 2019 | Rate-Flexible Fast Polar DecodersabstractPolar codes have gained extensive attention during the past few years and recently they have been selected for the next generation of wireless communications standards (5G). Successive-cancellation-based (SC-based) decoders, such as SC list (SCL) and SC flip (SCF), provide a reasonable error performance for polar codes at the cost of low decoding speed. Fast SC-based decoders, such as Fast-SSC, Fast-SSCL, and Fast-SSCF, identify the special constituent codes in a polar code graph off-line, produce a list of operations, store the list in memory, and feed the list to the decoder to decode the constituent codes in order efficiently, thus increasing the decoding speed. However, the list of operations is dependent on the code rate and as the rate changes, a new list is produced, making fast SC-based decoders not rate-flexible. In this paper, we propose a completely rate-flexible fast SC-based decoder by creating the list of operations directly in hardware, with low implementation complexity. We further propose a hardware architecture implementing the proposed method and show that the area occupation of the rate-flexible fast SC-based decoder in this paper is only 38% of the total area of the memory-based base-line decoder when 5G code rates are supported. Seyyed Ali Hashemi, Carlo Condo, Marco Mondelli, Warren J. Gross |
ITW | 2 |
| 2019 | Construction and Decoding of Product Codes with Non-Systematic Polar CodesabstractProduct codes are widespread in optical communications, thanks to their high throughput and good error-correction performance. Systematic polar codes have been recently considered as component codes for product codes. In this paper, we present a novel construction for product polar codes based on non-systematic polar codes. We prove that the resulting product code is actually a polar code, having a frozen set that is dependent on the frozen sets of the component polar codes. We propose a low-complexity decoding algorithm exploiting the dual nature of the constructed code. Performance analysis and simulations show high decoding speed, that allows to construct long codes while maintaining low decoding latency. The resulting high throughput and good error-correction performance are appealing for optical communication systems and other systems where high throughput and low latency are required. Valerio Bioglio, Carlo Condo, Ingmar Land |
WCNC | 2 |
| 2019 | Improved Bit-Flipping Algorithm for Successive Cancellation Decoding of Polar CodesabstractThe interest in polar codes has been increasing significantly since their adoption for use in the 5thgeneration wireless systems standard. Successive cancellation (SC) decoding algorithm has low implementation complexity, but yields mediocre error-correction performance at the code lengths of interest. SC-Flip algorithm improves the error-correction performance of SC by identifying possibly erroneous decisions made by SC and re-iterates after flipping one bit. It was recently shown that only a portion of bit-channels are most likely to be in error. In this paper, we investigate the average log-likelihood ratio (LLR) values and their distribution related to the erroneous bit-channels, and develop the Thresholded SC-Flip (TSCF) decoding algorithm. We also replace the LLR selection and sorting of SC-Flip with a comparator to reduce the implementation complexity. Simulation results demonstrate that for practical code lengths and a wide range of rates, TSCF shows negligible loss compared with the error-correction performance obtained when all single-errors are corrected. At matching maximum iterations, TSCF has an error-correction performance gain of up to 0.45 dB compared with SC-Flip decoding. At matching error-correction performance, the computational complexity of TSCF is reduced by up to 40% on average and requires up to 5× lower maximum number of iterations. Furkan Ercan, Carlo Condo, Warren J. Gross |
IEEE Trans. Commun. | 2 |
| 2018 | Generalized Fast Decoding of Polar CodesabstractResearch on polar codes has been constantly gaining attention over the last decade, by academia and industry alike, thanks to their capacity-achieving error-correction performance and low-complexity decoding algorithms. Recently, they have been selected as one of the coding schemes in the 5th generation wireless standard (5G). Over the years various polar code decoding algorithms, like SC-list (SCL), have been proposed to improve the mediocre performance of the successive cancellation (SC) decoding algorithm for finite code lengths; however, like SC, they suffer from long decoding latency. Fast decoding of polar codes tries to overcome this problem by identifying particular subcodes in the polar code and decoding them with efficient decoders. In this work, we introduce a generalized approach to fast decoding of polar codes to further reduce SC-based decoding latency. We propose three multi-node polar code subcodes whose identification patterns include most of the existing subcodes, extending them to SCL decoding, and allow to apply fast decoding to larger subsets of bits. Without any error-correction performance degradation, the proposed technique shows up to 23.6% and 29.2% decoding latency gain with respect to fast SC and SCL decoding algorithms, respectively, and up to 63.6% and 49.8% if a performance loss is accepted, whose amount depends on code and decoding algorithm parameters, along with the desired speedup. Carlo Condo, Valerio Bioglio, Ingmar Land |
GLOBECOM | 1 |
| 2018 | Partitioned Successive-Cancellation Flip Decoding of Polar CodesabstractPolar codes are a class of channel capacity achieving codes that has been selected for the next generation of wireless communication standards. Successive-cancellation (SC) is the first proposed decoding algorithm, suffering from mediocre errorcorrection performance at moderate code lengths. In order to improve the error-correction performance of SC, two approaches are available: (i) SC-List decoding which keeps a list of candidates by running a number of SC decoders in parallel, thus increasing the implementation complexity, and (ii) SC-Flip decoding that relies on a single SC module, and keeps the computational complexity close to SC. In this work, we propose the partitioned SC-Flip (PSCF) decoding algorithm, which outperforms SCFlip in terms of error-correction performance and average computational complexity, leading to higher throughput and reduced energy consumption per codeword. We also introduce a partitioning scheme that best suits our PSCF decoder. Simulation results show that at equivalent frame error rate, PSCF has up to 4.1× less computational complexity than the SC-Flip decoder. At equivalent average number of iterations, the error-correction performance of PSCF outperforms SC-Flip by up to 0.26 dB at frame error rate of 10-3. Furkan Ercan, Carlo Condo, Seyyed Ali Hashemi, Warren J. Gross |
ICC | 2 |
| 2018 | A Convolutional Accelerator for Neural Networks With Binary WeightsabstractParallel processors and GP-GPUs have been routinely used in the past to perform the computations of convolutional neural networks (CNNs). However, their large power consumption has pushed researchers towards application-specific integrated circuits and on-chip accelerators implement neural networks. Nevertheless, within the Internet of Things (IoT) scenario, even these accelerators fail to meet the power and latency constraints. To address this issue, binary-weight networks were introduced, where weights are constrained to -1 and 1. Therefore, these networks facilitate hardware implementation of neural networks by replacing multiply-and-accumulate units with simple accumulators, as well as reducing the weight storage. In this paper, we introduce a convolutional accelerator for binary-weight neural networks. The proposed architecture only consumes 128 mW at a frequency of 200 MHz and occupies 1.2 mm2when synthesized in TSMC 65 nm CMOS technology. Moreover, it achieves a high area-efficiency of 176 Gops/MGC and performance efficiency of 89%, outperforming the state-of-the-art architecture for binary-weight networks by 1.8× and 3.2×, respectively. Arash Ardakani, Carlo Condo, Warren J. Gross |
ISCAS | 2 |
| 2018 | Low-Complexity Software Stack Decoding of Polar CodesabstractPolar codes are a recent class of linear error-correcting codes that asymptotically achieve the channel capacity at infinite code length. The Successive Cancellation List (SCL) algorithm yields very good error-correction performance, at the cost of high implementation complexity. The Stack (SCS) decoding algorithm provides similar error-correction performance at a lower complexity. In this work, we propose an efficient software implementation of the SCS decoding algorithm, along with techniques to further reduce its computational complexity. In particular, we reduce the SCS memory requirements through efficient path switching, replace the stack sorting with a linear search, and explore the use of a partial CRC along with an early termination criterion. Using the proposed methods, we are able to reduce the computational complexity of the SCS decoder, reducing the number of estimated bits up to 97% with respect to SCL, while maintaining similar error-correction performance as SCL. Harsh Aurora, Carlo Condo, Warren J. Gross |
ISCAS | 2 |
| 2018 | Decoder Partitioning: Towards Practical List Decoding of Polar CodesabstractPolar codes represent one of the major recent breakthroughs in coding theory and, because of their attractive features, they have been selected for the incoming 5G standard. As such, a lot of attention has been devoted to the development of decoding algorithms with good error performance and efficient hardware implementation. One of the leading candidates in this regard is represented by successive-cancellation list (SCL) decoding. However, its hardware implementation requires a large amount of memory. Recently, a partitioned SCL (PSCL) decoder has been proposed to significantly reduce the memory consumption. In this paper, we consider the paradigm of PSCL decoding from a practical standpoint, and we provide several improvements. First, by changing the target signal-to-noise ratio and consequently modifying the construction of the code, we are able to improve the performance at no additional computational, latency, or memory cost. Second, we bridge the performance gap between SCL and PSCL decoding by introducing a generalized PSCL decoder and a layered PSCL decoder. In this way, we obtain almost the same performance of the SCL decoder with a significantly lower memory requirement, as testified by hardware implementation results. Third, we present an optimal scheme to allocate cyclic redundancy checks. Finally, we provide a lower bound on the list size that guarantees optimal maximum a posteriori performance for the binary erasure channel. Seyyed Ali Hashemi, Marco Mondelli, Seyed Hamed Hassani, Carlo Condo, Rüdiger L. Urbanke, Warren J. Gross |
IEEE Trans. Commun. | 4 |
| 2017 | Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks
Arash Ardakani, Carlo Condo, Warren J. Gross |
ICLR (Poster) | 2 |
| 2016 | Matrix reordering for efficient list sphere decoding of polar codesabstractThe Successive-Cancellation List (SCL) algorithm is one of the best polar code decoding algorithms in terms of trade-offs between complexity and error correction performance. The List-Sphere Decoding (List-SD) algorithm has been recently proposed: it yields a better complexity/performance trade-off than SCL in the decoding of short polar codes, that can be used as component codes for larger polar codes. We exploit the structure of the generator matrix of polar codes to propose a matrix reordering technique which allows to significantly reduce the List-SD complexity without degrading its error correction performance, further improving the aforementioned trade-off. The proposed technique is implemented on hardware and it is shown that at the same Frame Error Rate (FER) and Bit Error Rate (BER), the matrix reordering can reduce the resource requirements of List-SD of up to 73%. Furthermore, FER and BER curves are plotted for case studies, showing that at the same complexity cost, matrix reordering improves the performance of List-SD of up to 0.75 dB at FER=10-2. Seyyed Ali Hashemi, Carlo Condo, Warren J. Gross |
ISCAS | 2 |
| 2016 | Simplified Successive-Cancellation List decoding of polar codesabstractThe Successive-Cancellation List (SCL) decoding algorithm is one of the most promising approaches towards practical polar code decoding. It is able to provide a good trade-off between error-correction performance and complexity, tunable through the size of the list. In this paper, we show that in the conventional formulation of SCL, there are redundant calculations which do not need to be performed in the course of the algorithm. We simplify SCL by removing these redundant calculations and prove that the proposed simplified SCL and the conventional SCL algorithms are equivalent. The simplified SCL algorithm is valid for any code and can reduce the time-complexity of SCL without affecting the space complexity. Seyyed Ali Hashemi, Carlo Condo, Warren J. Gross |
ISIT | 2 |
| 2015 | Exploiting generalized de-Bruijn/Kautz topologies for flexible iterative channel code decoder architectures
Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
Integr. | 1 |
| 2015 | Unequal Error Protection of Memories in LDPC DecodersabstractMemories are one of the most critical components of many systems: due to exposure to energetic particles, fabrication defects and aging they are subject to various kinds of permanent and transient errors. In this scenario, Unequal error protection (UEP) techniques have been proposed in the past to encode stored information, allowing to detect and possibly recover from errors during load operations, while offering different levels of protection to partitions of codewords according to their importance. Low-density parity-check (LDPC) codes are used in many communication standards to encode the transmitted information: at reception, LDPC decoders heavily rely on memories to store and correct the received information. To ensure efficient and reliable decoding of information, the need to protect the memories used in LDPC decoders is of primary importance. In this paper we present a study on how to efficiently design UEP techniques for LDPC decoder memories. The devised UEP method is divided in four adjustable levels, each one offering a different degree of protection. The full UEP, along with simplified versions, has been implemented within an existing decoder and its area occupation and power consumption evaluated. Comparison with the literature on the subject shows an unmatched level of protection from errors at a small complexity and energy cost. Carlo Condo, Guido Masera, Paolo Montuschi |
IEEE Trans. Computers | 1 |
| 2014 | Rediscovering Logarithmic Diameter Topologies for Low Latency Network-on-Chip-Based ApplicationsabstractLow-latency Network-on-Chip (NoC) applications have tight constraints on the clock budget to perform communication among nodes. This is a critical aspect in NoC-based designs where the number of clock cycles spent for communication depends mainly on the topology and on the routing algorithm. This work deals with logarithmic diameter topologies, that were proposed for computer networks, and shows that an optimal shortest-path routing algorithm can be efficiently implemented on this kind of topologies by means of a very simple circuit. The proposed circuit is then exploited to reduce the area and the power consumption of a recently proposed NoC-based design. Experimental results show that the proposed circuit allows for a reduction of about 14% and 10% for area and power consumption respectively, with respect to a shortest-path routing-table-based design. Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
PDP | 1 |
| 2014 | Energy-efficient multi-standard early stopping criterion for low-density-parity-check iterative decodingabstractLow‐density‐parity‐check codes decoding relies on powerful iterative algorithms, whose implementation is often expensive in terms of complexity and power consumption. Several early stopping criteria (ESCs) have been proposed to reduce the number of iterations performed by a decoder with no (or limited) degradation of error correction performance. However, most of the existing ESCs have considered a reduced set of system parameters for validation and often have ignored the impacts related to a real hardware implementation. This study proposes a novel multi‐standard early stopping criterion (MSESC) able to adapt dynamically to changes of code parameters, quantisation and channel conditions. A dedicated hardware architecture is devised and integrated in a multi‐standard decoder, and compared with existing techniques. Post‐layout results of the proposed MSESC show a small area increment (+1.3%) and a large decrement of the average energy consumption (up to 87.2%) with respect to the same decoder implemented with no ESC. Moreover, it is shown that MSESC offers an energy consumption reduction with respect to the state‐of‐the‐art ESCs ranging from 4% [at high signal‐to‐noise ratio (SNR)] to 16% (at low SNR). Carlo Condo, Amer Baghdadi, Guido Masera |
IET Commun. | 1 |
| 2014 | Variable Parallelism Cyclic Redundancy Check Circuit for 3GPP-LTE/LTE-AdvancedabstractCyclic Redundancy Check (CRC) is often employed in data storage and communications to detect errors. The 3GPP-LTE wireless communication standard uses a 24-bit CRC with every turbo coded frame, thus, the CRC can be exploited to detect residual errors and to enable early stopping of iterations as well. The current state of the art lacks specific CRC implementations for this standard, and most current solutions adopt a fixed degree of parallelism, unsuitable for many turbo decoder architectures. This work proposes a variable parallelism circuit targeting the 3GPP-LTE/LTE-Advanced 24-bit CRC, that can adapt to input data of different sizes. Low complexity is achieved through careful functional sharing among the various parallelisms: comparison with the state of the art shows comparable or superior speed and extremely low complexity. Carlo Condo, Maurizio Martina, Gianluca Piccinini, Guido Masera |
IEEE Signal Process. Lett. | 1 |
| 2013 | A Joint Communication and Application Simulator for NoC-Based Custom SoCs: LDPC and Turbo Codes Parallel Decoding Case StudyabstractNoCs have become a widespread paradigm in the system-on-chip design world, not only for multi-purpose SoCs, but also for application-specific ICs. The common approach in the NoC design world is to separate the design of the interconnection from the design of the processing elements: this is well suited for a large number of developments, but the need for joint application and NoC design is not uncommon, especially in the application-specific case. The correlation between processing and communication tasks can be strong, and separate or trace-based simulations fall often short of the desired precision. In this work, the OMNET++ based JANoCS simulator is presented: concurrent simulation of processing and communication allow cycle-accurate evaluation of the system. The potential of the proposed approach is illustrated through a simple application example. Furthermore, a detailed case study on LDPC and turbo codes parallel decoding is presented. Results analysis illustrates the need for joint simulations and demonstrates the effectiveness of the proposed JANoCS. Carlo Condo, Amer Baghdadi, Guido Masera |
DSD | 1 |
| 2012 | A Network-on-Chip-based turbo/LDPC decoder architectureabstractThe current convergence process in wireless technologies demands for strong efforts in the conceiving of highly flexible and interoperable equipments. This contribution focuses on one of the most important baseband processing units in wireless receivers, the forward error correction unit, and proposes a Network-on-Chip (NoC) based approach to the design of multi-standard decoders. High level modeling is exploited to drive the NoC optimization for a given set of both turbo and Low-Density-Parity-Check (LDPC) codes to be supported. Moreover, synthesis results prove that the proposed approach can offer a fully compliant WiMAX decoder, supporting the whole set of turbo and LDPC codes with higher throughput and an occupied area comparable or lower than previously reported flexible implementations. In particular, the mentioned design case achieves a worst-case throughput higher than 70 Mb/s at the area cost of 3.17 mm2on a 90 nm CMOS technology. Carlo Condo, Maurizio Martina, Guido Masera |
DATE | 1 |