EDBT 2026 Demo / reviewers in the wild / expert
Guido Masera
dblp:54/1377
· DBLP profile ↗
73ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-2238-9443ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Computer networks · 6Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIRCE CROSS Integrated RISC-V Cryptographic Extension
Alessandra Dolmeta, Valeria Piscopo, Maurizio Martina, Guido Masera |
DATE | 4 |
| 2026 | Compact Yet Fast: An Efficient d-Order Masked Implementation of AsconabstractIn this work, we present a generic side-channel protected design of Ascon that achieves high efficiency by dynamically reconfiguring the hardware countermeasures during message processing. The resultant implementation is protected and capable of meeting stringent performance requirements whilst minimising resource overhead. The experimental results obtained demonstrate that the implementation meets the required security and achieves superior throughput-to-area ratio across all protection orders. Ascon, recently selected by NIST as the lightweight cryptography standard, is widely deployed in resource-constrained devices that demand both high performance and resistance against threats such as side-channel analysis (SCA). Exploiting Ascon’s mode-level structure, which does not require protection against differential power analysis during bulk operations, we introduce a modified masking gadget with dual functionality: serving as a countermeasure during critical operations, and processing multiple data paths in parallel to accelerate bulk computation. Our architecture supports any configurable security order and instantiates only the minimum hardware resources needed to maximize throughput per round. We also evaluate an enhanced Ascon architecture based on the Changing of the Guards technique, which eliminates the need for fresh randomness. Security validation is performed using fixed-vs-random t-tests on both first- and second-order masked implementations. Finally, we compare our masked design against state-of-the-art solutions. Mattia Mirigaldi, Nico Paninforni, Maurizio Martina, Guido Masera |
DATE | 4 |
| 2025 | ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsabstractModern data-driven applications expose limitations of von Neumann architectures-extensive data movement, low throughput, and poor energy efficiency. Accelerators improve performance but lack flexibility and require data transfers. Existing compute in- and nearmemory solutions mitigate these issues but face usability challenges due to data placement constraints. We propose a novel cache architecture that doubles as a tightly-coupled compute-near-memory coprocessor. Our RISC-V cache controller executes custom instructions from the host CPU using vector operations dispatched to near-memory vector processing units within the cache memory subsystem. This architecture abstracts memory synchronization and data mapping from application software while offering software-based Instruction Set Architecture extensibility. Our implementation shows $30 \times$ to $84 \times$ performance improvement when operating on 8-bit data over the same system with a traditional cache when executing a worst-case 32-bit CNN workload, with only 41.3% area overhead. Vincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone, Guido Masera, Maurizio Martina |
DAC | 5 |
| 2025 | TYRCA: A RISC-V Tightly-Coupled Accelerator for Code-Based CryptographyabstractPost-quantum cryptography (PQC) has garnered significant attention across various communities, particularly with the National Institute of Standards and Technology (NIST) advancing to the fourth round of PQC standardization. One of the leading candidates is Hamming Quasi-Cyclic (HQC), which received a significant update on February 23, 2024. This update, which introduces a classical dense-dense multiplication approach, has no known dedicated hardware implementations yet. The innovative Core-V eXtension InterFace (CV-X-IF) is a communication interface for RISC-V processors that significantly facilitates the integration of new instructions to the Instruction Set Architecture (ISA), through tightly connected accelerators. In this paper, we present a TightlY-coupled accelerator for RISC-V for Code-based cryptogrAphy (TYRCA), proposing the first fully tightly-coupled hardware implementation of the HQC-PQC algorithm, leveraging the CV-X-IF. The proposed architecture is implemented on the Xilinx Kintex-7 FPGA. Experimental results demonstrate that TYRCA reduces the execution time by 94% to 96% for HQC-128, HQC-192, and HQC-256, showcasing its potential for efficient HQC code-based cryptography. Alessandra Dolmeta, Stefano Di Matteo, Emanuele Valea, Mikael Carmona, Antoine Loiseau, Maurizio Martina, Guido Masera |
DATE | 7 |
| 2025 | TinyCL: An Efficient Hardware Architecture for Continual Learning on Autonomous SystemsabstractThe Continuous Learning (CL) paradigm consists of continuously evolving the parameters of the Deep Neural Network (DNN) model to progressively learn to perform new tasks without reducing the performance on previous tasks, i.e., avoiding the so-called catastrophic forgetting. However, the DNN parameter update in CL-based autonomous systems is extremely resource-hungry. The existing DNN accelerators cannot be directly employed in CL because they only support the execution of the forward propagation. Only a few prior architectures execute the backpropagation and weight update, but they lack the control and management for CL. Towards this, we design a hardware architecture, TinyCL, to perform CL on resource-constrained autonomous systems. It consists of a processing unit that executes both forward and backward propagation, and a control unit that manages memory-based CL workload. To minimize the memory accesses, the sliding window of the convolutional layer moves in a snake-like fashion. Moreover, the Multiply-and-Accumulate units can be reconfigured at runtime to execute different operations. As per our knowledge, our proposed TinyCL represents the first hardware accelerator that executes CL on autonomous systems. We synthesize the complete TinyCL architecture in a 65 nm CMOS technology node with the conventional ASIC design flow. It executes 1 epoch of training on a Conv + ReLU + Dense model on the CIFAR10 dataset in 1.76 s, while 1 training epoch of the same model using an Nvidia Tesla P100 GPU takes 103 s, thus achieving a 58× speedup, consuming 86 mW in a 4.74 mm2die. Eugenio Ressa, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2025 | RISC-V Based Keccak Co-Processor for NIST Post-Quantum Cryptography StandardsabstractThis paper presents the design and implementation of a RISC-V-based Keccak co-processor optimized for Post-Quantum Cryptography (PQC) algorithms. Leveraging the Core-V eXtension InterFace (CV-X-IF), the co-processor extends the Instruction Set Architecture (ISA) with three custom instructions tailored for cryptographic operations. This allows seamless integration into various PQC schemes, tested across the multiple standards proposed by the National Institute of Standards and Technology (NIST), including CRYSTALS-Kyber, CRYSTALS-Dilithium, SPHINCS+, and FALCON, which are designed to withstand quantum attacks. By employing tightly coupled hardware acceleration, the Keccak co-processor dramatically reduces the computational overhead of hash-based operations central to these algorithms. The implementation is realized on a Xilinx Artix 7 FPGA, achieving a clock cycles’ improvement up to 75% and 19% resource overhead. The results presented herein demonstrate significant performance enhancement over the state of the art, underscoring its effectiveness for cryptographic applications. Alessandra Dolmeta, Valeria Piscopo, Mattia Mirigaldi, Maurizio Martina, Guido Masera |
ISCAS | 5 |
| 2025 | Power Side-Channel Vulnerabilities of a RISC-V Cryptography Accelerator Integrated into CVA6 via Core-V eXtension Interface (CV-X-IF)abstractModern RISC-V designs are increasingly integrating cryptographic accelerators to provide better security features while enhancing performance; however, their vulnerability to power side-channel attacks remains insufficiently investigated. This paper presents a comprehensive evaluation of such vulnerabilities in a RISCV-based AES accelerator connected via the Core-V eXtension Interface (CV-X-IF). The analysis begins at the RTL using simulated power traces, employing KL (Kullback–Leibler) divergence alongside established statistical attacks such as Correlation Power Analysis (CPA) and Differential Power Analysis (DPA). Although the former serves as an early indicator of potential leakage, simulation results highlight its limitations compared to CPA and DPA. To validate these findings, leakage trends are further examined through FPGA-based power measurement. The proposed methodology is designed to be broadly applicable to a range of cryptographic workloads and accelerator architectures. It is demonstrated on an AES accelerator implementing the scalar cryptographic extension (Zk) with pre-expanded keys. Our findings reveal that side-channel vulnerabilities can persist even in tightly integrated instruction pipelines, underscoring the importance of early-stage leakage assessment. Notably, the close alignment between RTL-level simulations and FPGA-based measurements highlights the effectiveness of the approach and its practical value for guiding secure hardware design in RISC-V ecosystems. In particular, AES serves only as a case of study; the proposed RTL and FPGA validation flow is generic and can be applied to any cryptographic accelerator. Behnam Farnaghinejad, Davide Bellizia, Alessandra Dolmeta, Guido Masera, Antonio Porsia, Annachiara Ruospo, Stefano Di Carlo, Alessandro Savino 0001, Ernesto Sánchez 0001 |
ITC | 4 |
| 2025 | Scalability analysis of multi-bank near-memory computing in low-power SoCsabstractMachine learning and artificial intelligence are moving towards the edge, where the need for high throughput with a constrained energy budget is more urgent than ever. During the last few years, near-memory computing has emerged as a promising solution to address the memory bandwidth and energy efficiency limitations of conventional von Neumann systems. The recently proposed NM-Carus architecture combines vectororiented computing capabilities within a RISC-V programmable, configurable, and autonomous memory macro, addressing the usability of near-memory computing from a software deployment standpoint. In this paper, we explore the scalability of NMCarus in terms of computation parallelism, memory size and energy consumption, As a benchmarking platform, we rely on a low-power microcontroller that features multiple instances of NM-Carus that target the execution of biomedical applications. This exploration was performed on 16 nm TSMC NM-Carus implmentation, and we highlighted the benefits of technology scaling for a previous implementation on 65 nm with respect to the overhead of replacing conventional on-chip data SRAMs with near-memory computing banks. Overall, the paper presents a solid baseline regarding the trade-offs in terms of area, performance, and energy efficiency of integrating programmable near-memory computing in an existing edge-oriented system on chip towards efficient edge AI architectures at the system level. Luigi Giuffrida, Pasquale Davide Schiavone, Michele Caon, Guido Masera, Maurizio Martina, David Atienza 0001 |
VLSI-SoC | 4 |
| 2024 | MARLIN: A Co-Design Methodology for Approximate ReconfigurabLe Inference of Neural Networks at the EdgeabstractThe optimization of neural networks (NNs) is necessary to enable their deployment on energy-constrained devices. State-of-the-art methods leverage approximate multipliers to execute NNs reducing the inference energy without heavily affecting the accuracy. However, previous works usually require a specialized hardware accelerator and are limited to fixed multipliers or reconfigurable ones with few approximation levels. This paper introduces MARLIN, a framework to deploy layerwise approximate NNs on PULP, a microcontroller with a RISC-V core. A multiplier architecture, with runtime selection of 256 approximation levels, is developed and integrated into the PULP cluster cores, enabling runtime configuration through control status register (CSR) instructions embedded within the code. The PULP toolchain is adapted to incorporate the approximation level selection within the instruction flow seamlessly. MARLIN leverages the genetic algorithm NSGA-II to search for the best configurations among thousands of approximate NNs. The framework is validated by simulating an approximate NN trained with the MNIST dataset on PULP. Moreover, MARLIN is used to optimize and approximate six ResNet models trained with the CIFAR-10 dataset. In particular, for ResNet-56, the most complex NN used in the experiments, the multiplication energy is reduced by 23.9% while retaining 99% of the accuracy of the exact model. Flavia Guella, Emanuele Valpreda, Michele Caon, Guido Masera, Maurizio Martina |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Implementation and integration of Keccak accelerator on RISC-V for CRYSTALS-KyberabstractOne of the key metrics used for defying the security of the Internet of Things (IoT) is data integrity, which mostly relies on the use of cryptographic hash functions. In the last years, the National Institute of Standards and Technology (NIST) announced SHA-3 as the new standard for better security. SHA-3 is also exploited in most of the current post-quantum cryptographic (PQC) protocols. Nevertheless, the used algorithm, i.e. Keccak, is computationally heavy and consequently limits its utilization in RISC-V-based Systems on Chip (SoC). In this work, a Keccak accelerator is proposed to speed up SHA3 computations for the CRYSTALS-Kyber algorithm on the RISCV-based advanced microcontroller PULPissimo. Compared to the plain SW implementation on RISC-V, our results show a speedup factor of up to 2.79 at the expense of a 12.4% resources overhead. Alessandra Dolmeta, Mattia Mirigaldi, Maurizio Martina, Guido Masera |
CF | 4 |
| 2023 | SwiftTron: An Efficient Hardware Accelerator for Quantized TransformersabstractTransformers' compute- intensive operations pose enormous challenges for their deployment in resource- constrained EdgeAI / tiny ML devices. As an established neural network compression technique, quantization reduces the hardware computational and memory resources. In particular, fixed-point quantization is desirable to ease the computations using lightweight blocks, like adders and multipliers, of the underlying hardware. However, deploying fully-quantized Transformers on existing general-purpose hardware, generic AI accelerators, or specialized architectures for Transformers with floating-point units might be infeasible and/or inefficient. Towards this, we propose SwiftTron, an efficient specialized hardware accelerator designed for Quantized Transformers. SwiftTron supports the execution of different types of Transformers' operations (like Attention, Softmax, GELU, and Layer Normalization) and accounts for diverse scaling factors to perform correct computations. We synthesize the complete SwiftTron architecture in a 65 nm CMOS technology with the ASIC design flow. Our Accelerator executes the RoBERTa-base model in 1.83 ns, while consuming 33.64 mW power, and occupying an area of 273 mm2• To ease the reproducibility, the RTL of our SwiftTron architecture is released at https://github.com/albertomarchisio/SwiftTron. Alberto Marchisio, Davide Dura, Maurizio Capra, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 5 |
| 2022 | NLCMAP: A Framework for the Efficient Mapping of Non-Linear Convolutional Neural Networks on FPGA AcceleratorsabstractThis paper introduces NLCMap, a framework for the mapping space exploration targeting Non-Linear Convolutional Networks (NLCNs). NLCNs [1] are a novel neural network model that improves performances in certain computer vision applications by introducing a non-linearity in the weights computation. NLCNs are more challenging to efficiently map onto hardware accelerators if compared to traditional Convolutional Neural Networks (CNNs), due to data dependencies and additional computations. To this aim, we propose NL-CMap, a framework that, given an NLC layer and a generic hardware accelerator with a certain on-chip memory budget, finds the optimal mapping that minimizes the accesses to the off-chip memory, which are often the critical aspect in CNNs acceleration. Giuseppe Aiello, Beatrice Bussolino, Emanuele Valpreda, Massimo Ruo Roch, Guido Masera, Maurizio Martina, Stefano Marsi |
ICIP | 5 |
| 2022 | CoNLoCNN: Exploiting Correlation and Non-Uniform Quantization for Energy-Efficient Low-precision Deep Convolutional Neural NetworksabstractIn today's era of smart cyber-physical systems, Deep Neural Networks (DNNs) have become ubiquitous due to their state-of-the-art performance in complex real-world applications. The high computational complexity of these networks, which translates to increased energy consumption, is the foremost obstacle towards deploying large DNNs in resource-constrained systems. Fixed-Point (FP) implementations achieved through post-training quantization are commonly used to curtail the energy consumption of these networks. However, the uniform quantization intervals in FP restrict the bit-width of data structures to large values due to the need to represent most of the numbers with sufficient resolution and avoid high quantization errors. In this paper, we leverage the key insight that (in most of the scenarios) DNN weights and activations are mostly concentrated near zero and only a few of them have large magnitudes. We propose CoNLoCNN, a framework to enable energy-efficient low-precision deep convolutional neural network inference by exploiting: (1) non-uniform quantization of weights enabling simplification of complex multiplication operations; and (2) correlation between activation values enabling partial compensation of quantization errors at low cost without any run-time overheads. To significantly benefit from non-uniform quantization, we also propose a novel data representation format, Encoded Low-Precision Binary Signed Digit, to compress the bit-width of weights while ensuring direct use of the encoded weight for processing using a novel multiply-and-accumulate (MAC) unit design. Muhammad Abdullah Hanif, Giuseppe Maria Sarda, Alberto Marchisio, Guido Masera, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2022 | LaneSNNs: Spiking Neural Networks for Lane Detection on the Loihi Neuromorphic ProcessorabstractAutonomous Driving (AD) related features represent important elements for the next generation of mobile robots and autonomous vehicles focused on increasingly intelligent, autonomous, and interconnected systems. The applications involving the use of these features must provide, by definition, real-time decisions, and this property is key to avoid catastrophic accidents. Moreover, all the decision processes must require low power consumption, to increase the lifetime and autonomy of battery-driven systems. These challenges can be addressed through efficient implementations of Spiking Neural Networks (SNNs) on Neuromorphic Chips and the use of event-based cameras instead of traditional frame-based cameras. In this paper, we present a new SNN-based approach, called LaneSNN, for detecting the lanes marked on the streets using the event-based camera input. We develop four novel SNN models characterized by low complexity and fast response, and train them using an offline supervised learning rule. Afterward, we implement and map the learned SNNs models onto the Intel Loihi Neuromorphic Research Chip. For the loss function, we develop a novel method based on the linear composition of Weighted binary Cross Entropy (WCE) and Mean Squared Error (MSE) measures. Our experimental results show a maximum Intersection over Union (IoU) measure of about 0.62 and very low power consumption of about 1 W. The best IoU is achieved with an SNN implementation that occupies only 36 neurocores on the Loihi processor while providing a low latency of less than 8 ms to recognize an image, thereby enabling real-time performance. The IoU measures provided by our networks are comparable with the state-of-the-art, but at a much low power consumption of 1 W. Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IROS | 4 |
| 2022 | Enabling Capsule Networks at the Edge through Approximate Softmax and Squash OperationsabstractComplex Deep Neural Networks such as Capsule Networks (CapsNets) exhibit high learning capabilities at the cost of compute-intensive operations. To enable their deployment on edge devices, we propose to leverage approximate computing for designing approximate variants of the complex operations like softmax and squash. In our experiments, we evaluate tradeoffs between area, power consumption, and critical path delay of the designs implemented with the ASIC design flow, and the accuracy of the quantized CapsNets, compared to the exact functions. Alberto Marchisio, Beatrice Bussolino, Edoardo Salvati, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
ISLPED | 5 |
| 2021 | DVS-Attacks: Adversarial Attacks on Dynamic Vision Sensors for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), despite being energy-efficient when implemented on neuromorphic hardware and coupled with event-based Dynamic Vision Sensors (DVS), are vulnerable to security threats, such as adversarial attacks, i.e., small perturbations added to the input for inducing a misclassification. Toward this, we propose DVS-Attacks, a set of stealthy yet efficient adversarial attack methodologies targeted to perturb the event sequences that compose the input of the SNNs. First, we show that noise filters for DVS can be used as defense mechanisms against adversarial attacks. Afterwards, we implement several attacks and test them in the presence of two types of noise filters for DVS cameras. The experimental results show that the filters can only partially defend the SNNs against our proposed DVS-Attacks. Using the best settings for the noise filters, our proposed Mask Filter-Aware Dash Attack reduces the accuracy by more than 20% on the DVS-Gesture dataset and by more than 65% on the MNIST dataset, compared to the original clean frames. The source code of all the proposed DVS-Attacks and noise filters is released at https://github.com/albertomarchisio/DVS-Attacks. Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2021 | CarSNN: An Efficient Spiking Neural Network for Event-Based Autonomous Cars on the Loihi Neuromorphic Research ProcessorabstractAutonomous Driving (AD) related features provide new forms of mobility that are also beneficial for other kind of intelligent and autonomous systems like robots, smart transportation, and smart industries. For these applications, the decisions need to be made fast and in real-time. Moreover, in the quest for electric mobility, this task must follow low power policy, without affecting much the autonomy of the mean of transport or the robot. These two challenges can be tackled using the emerging Spiking Neural Networks (SNNs). When deployed on a specialized neuromorphic hardware, SNNs can achieve high performance with low latency and low power consumption. In this paper, we use an SNN connected to an event-based camera for facing one of the key problems for AD, i.e., the classification between cars and other objects. To consume less power than traditional frame-based cameras, we use a Dynamic Vision Sensor (DVS) [1]. The experiments are made following an offline supervised learning rule, followed by mapping the learnt SNN model on the Intel Loihi Neuromorphic Research Chip [2]. Our best experiment achieves an accuracy on offline implementation of 86%, that drops to 83% when it is ported onto the Loihi Chip. The Neuromorphic Hardware implementation has maximum 0.72 ms of latency for every sample, and consumes only 310 mW. To the best of our knowledge, this work is the first implementation of an event-based car classifier on a Neuromorphic Chip. Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2021 | R-SNN: An Analysis and Design Methodology for Robustifying Spiking Neural Networks against Adversarial Attacks through Noise Filters for Dynamic Vision SensorsabstractSpiking Neural Networks (SNNs) aim at providing energy-efficient learning capabilities when implemented on neuromorphic chips with event-based Dynamic Vision Sensors (DVS). This paper studies the robustness of SNNs against adversarial attacks on such DVS-based systems, and proposes R-SNN, a novel methodology for robustifying SNNs through efficient DVS-noise filtering. We are the first to generate adversarial attacks on DVS signals (i.e., frames of events in the spatio-temporal domain) and to apply noise filters for DVS sensors in the quest for defending against adversarial attacks. Our results show that the noise filters effectively prevent the SNNs from being fooled. The SNNs in our experiments provide more than 90% accuracy on the DVS-Gesture and NMNIST datasets under different adversarial threat models. Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IROS | 4 |
| 2020 | Q-CapsNets: A Specialized Framework for Quantizing Capsule NetworksabstractCapsule Networks (CapsNets), recently proposed by the Google Brain team, have superior learning capabilities in machine learning tasks, like image classification, compared to the traditional CNNs. However, CapsNets require extremely intense computations and are difficult to be deployed in their original form at the resource-constrained edge devices. This paper makes the first attempt to quantize CapsNet models, to enable their efficient edge implementations, by developing a specialized quantization framework for CapsNets. We evaluate our framework for several benchmarks. On a deep CapsNet model for the CIFAR10 dataset, the framework reduces the memory footprint by 6.2x, with only 0.15% accuracy loss. We will open-source our framework at https://git.io/JvDIF. Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
DAC | 5 |
| 2020 | FasTrCaps: An Integrated Framework for Fast yet Accurate Training of Capsule NetworksabstractRecently, Capsule Networks (CapsNets) have shown improved performance compared to the traditional Convolutional Neural Networks (CNNs), by encoding and preserving spatial relationships between the detected features in a better way. This is achieved through the so-called Capsules (i.e., groups of neurons) that encode both the instantiation probability and the spatial information. However, one of the major hurdles in the wide adoption of CapsNets is their gigantic training time, which is primarily due to the relatively higher complexity of their new constituting elements that are different from CNNs.In this paper, we implement different optimizations in the training loop of the CapsNets, and investigate how these optimizations affect their training speed and the accuracy. Towards this, we propose a novel framework FasTrCaps that integrates multiple lightweight optimizations and a novel learning rate policy called WarmAdaBatch (that jointly performs warm restarts and adaptive batch size), and steers them in an appropriate way to provide high training-loop speedup at minimal accuracy loss. We also propose weight sharing for capsule layers. The goal is to reduce the hardware requirements of CapsNets by removing unused/redundant connections and capsules, while keeping high accuracy through tests of different learning rate policies and batch sizes. We demonstrate that one of the solutions generated by the FasTrCaps framework can achieve 58.6% reduction in the training time, while preserving the accuracy (even 0.12% accuracy improvement for the MNIST dataset), compared to the CapsNet by Google Brain [25]. Moreover, the Pareto-optimal solutions generated by FasTrCaps can be leveraged to realize trade-offs between training time and achieved accuracy. We have open-sourced our framework on GitHub1. Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Muhammad Abdullah Hanif, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 6 |
| 2020 | An Area-Efficient Variable-Size Fixed-Point DCT Architecture for HEVC EncodingabstractThis paper proposes an area-efficient fixed-point architecture for the computation of the discrete cosine transform (DCT) of multiple sizes in high efficiency video coding (HEVC). This result is obtained by comparing different DCT factorizations in order to find the most suitable one for implementation in the HEVC encoder. The recursive structure of fast algorithms, which decompose the N-point DCT by means of two N/2-point DCTs, is exploited to execute computations of small-size DCTs in parallel, thus maximizing the hardware reusability while maintaining a constant throughput. The simulation results prove that the proposed solution features reduced rate-distortion losses, with relevant complexity saving compared with the state-of-the-art implementations. Finally, the proposed architecture is exploited to design two families of architectures for the 2D-DCT, namely, folded and full-parallel. Maurizio Masera, Guido Masera, Maurizio Martina |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Analysis of HEVC transform throughput requirements for hardware implementations
Maurizio Masera, Lorenzo Re Fiorentin, Enrico Masala, Guido Masera, Maurizio Martina |
Signal Process. Image Commun. | 4 |
| 2017 | Adaptive Approximated DCT Architectures for HEVCabstractThis paper proposes a flexible and efficient implementation of the 2D N-point discrete cosine transform (DCT) for the High Efficiency Video Coding (HEVC) standard. The DCT is implemented through the Walsh-Hadamard transform (WHT) followed by Givens rotations. This scheme is exploited to derive an adaptive algorithm, which allows computing of four different approximations ranging from the complete DCT to the WHT, by selectively skipping some rotations. This paper shows the statistical analysis of the DCT usage and derives a precomputation mechanism to adaptively skip rotations. Each approximation, referred to as a operating mode, is characterized by a large saving of operations, at the expense of very small quality loss. Then, two 2D-DCT architectures are proposed: the first one is totally unfolded, while the second one is folded. The two designs are finally synthesized with a 90-nm standard-cell library for a clock frequency of 250 MHz. Both architectures support real-time processing of 8K UHD video sequences at 64 and 26 fps, respectively, and show higher throughput and lower gate count compared with the state-of-art implementations. Moreover, power saving ranging from 28% to 56% can be achieved by working within the proposed operating modes. Maurizio Masera, Maurizio Martina, Guido Masera |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | An all-digital spike-based ultra-low-power IR-UWB dynamic average threshold crossing scheme for muscle force wireless transmission
Masoud Shahshahani Amirhossein, Paolo Motto Ros, Alberto Bonanno, Marco Crepaldi, Maurizio Martina, Danilo Demarchi, Guido Masera |
DATE | 7 |
| 2015 | Exploiting generalized de-Bruijn/Kautz topologies for flexible iterative channel code decoder architectures
Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
Integr. | 4 |
| 2015 | Unequal Error Protection of Memories in LDPC DecodersabstractMemories are one of the most critical components of many systems: due to exposure to energetic particles, fabrication defects and aging they are subject to various kinds of permanent and transient errors. In this scenario, Unequal error protection (UEP) techniques have been proposed in the past to encode stored information, allowing to detect and possibly recover from errors during load operations, while offering different levels of protection to partitions of codewords according to their importance. Low-density parity-check (LDPC) codes are used in many communication standards to encode the transmitted information: at reception, LDPC decoders heavily rely on memories to store and correct the received information. To ensure efficient and reliable decoding of information, the need to protect the memories used in LDPC decoders is of primary importance. In this paper we present a study on how to efficiently design UEP techniques for LDPC decoder memories. The devised UEP method is divided in four adjustable levels, each one offering a different degree of protection. The full UEP, along with simplified versions, has been implemented within an existing decoder and its area occupation and power consumption evaluated. Comparison with the literature on the subject shows an unmatched level of protection from errors at a small complexity and energy cost. Carlo Condo, Guido Masera, Paolo Montuschi |
IEEE Trans. Computers | 2 |
| 2015 | Parallel H.264/AVC Fast Rate-Distortion Optimized Motion Estimation by Using a Graphics Processing Unit and Dedicated HardwareabstractHeterogeneous systems on a single chip composed of a central processing unit, graphics processing unit (GPU), and field-programmable gate array (FPGA) are expected to emerge in the near future. In this context, the system on chip can be dynamically adapted to employ different architectures for execution of data-intensive applications. Motion estimation (ME) is one such task that can be accelerated using FPGA and GPU for high-performance H.264/Advanced Video Coding encoder implementation. This paper presents an inherent parallel low-complexity rate-distortion (RD) optimized fast ME algorithm well suited for parallel implementations, eliminating various data dependencies caused by a reliance on spatial predictions. In addition, this paper provides details of the GPU and FPGA implementations of the parallel algorithm by using OpenCL and Very High Speed Integrated Circuits (VHSIC) Hardware Descriptive Language (VHDL), respectively, and presents a practical performance comparison between the two implementations. The experimental results show that the proposed scheme achieves significant speedup on GPU and FPGA, and has comparable RD performance with respect to sequential fast ME algorithm. Muhammad Usman Shahid, Ashfaq Ahmed, Maurizio Martina, Guido Masera, Enrico Magli |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | A novel decoder architecture for error resilient JPEG2000 applications based on MQ arithmeticabstractIn this paper we investigate the complexity of a modified MQ arithmetic decoder architecture for error resilient JPEG2000 applications, based on the concept of ternary arithmetic decoder employing a forbidden symbol. Such ternary coding schemes have been already investigated to prove its excellent figures in terms of both PSNR and visual quality. However the complexity of the modified MQ decoder and its impact on real implementations has not been considered yet. In this paper an architectural analysis is performed and an optimized architecture is proposed. Simone Zezza, Guido Masera, Saeid Nooshabadi |
ISCAS | 2 |
| 2014 | Rediscovering Logarithmic Diameter Topologies for Low Latency Network-on-Chip-Based ApplicationsabstractLow-latency Network-on-Chip (NoC) applications have tight constraints on the clock budget to perform communication among nodes. This is a critical aspect in NoC-based designs where the number of clock cycles spent for communication depends mainly on the topology and on the routing algorithm. This work deals with logarithmic diameter topologies, that were proposed for computer networks, and shows that an optimal shortest-path routing algorithm can be efficiently implemented on this kind of topologies by means of a very simple circuit. The proposed circuit is then exploited to reduce the area and the power consumption of a recently proposed NoC-based design. Experimental results show that the proposed circuit allows for a reduction of about 14% and 10% for area and power consumption respectively, with respect to a shortest-path routing-table-based design. Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
PDP | 4 |
| 2014 | Energy-efficient multi-standard early stopping criterion for low-density-parity-check iterative decodingabstractLow‐density‐parity‐check codes decoding relies on powerful iterative algorithms, whose implementation is often expensive in terms of complexity and power consumption. Several early stopping criteria (ESCs) have been proposed to reduce the number of iterations performed by a decoder with no (or limited) degradation of error correction performance. However, most of the existing ESCs have considered a reduced set of system parameters for validation and often have ignored the impacts related to a real hardware implementation. This study proposes a novel multi‐standard early stopping criterion (MSESC) able to adapt dynamically to changes of code parameters, quantisation and channel conditions. A dedicated hardware architecture is devised and integrated in a multi‐standard decoder, and compared with existing techniques. Post‐layout results of the proposed MSESC show a small area increment (+1.3%) and a large decrement of the average energy consumption (up to 87.2%) with respect to the same decoder implemented with no ESC. Moreover, it is shown that MSESC offers an energy consumption reduction with respect to the state‐of‐the‐art ESCs ranging from 4% [at high signal‐to‐noise ratio (SNR)] to 16% (at low SNR). Carlo Condo, Amer Baghdadi, Guido Masera |
IET Commun. | 3 |
| 2014 | Variable Parallelism Cyclic Redundancy Check Circuit for 3GPP-LTE/LTE-AdvancedabstractCyclic Redundancy Check (CRC) is often employed in data storage and communications to detect errors. The 3GPP-LTE wireless communication standard uses a 24-bit CRC with every turbo coded frame, thus, the CRC can be exploited to detect residual errors and to enable early stopping of iterations as well. The current state of the art lacks specific CRC implementations for this standard, and most current solutions adopt a fixed degree of parallelism, unsuitable for many turbo decoder architectures. This work proposes a variable parallelism circuit targeting the 3GPP-LTE/LTE-Advanced 24-bit CRC, that can adapt to input data of different sizes. Low complexity is achieved through careful functional sharing among the various parallelisms: comparison with the state of the art shows comparable or superior speed and extremely low complexity. Carlo Condo, Maurizio Martina, Gianluca Piccinini, Guido Masera |
IEEE Signal Process. Lett. | 4 |
| 2013 | VLSI Architecture for Low-Complexity Motion Estimation in H.264 Multiview Video CodingabstractThis paper presents a VLSI architecture for a low complexity motion estimation algorithm, referred to as Slim264, for multiview video coding extension of H.264. Algorithmic modifications are introduced to obtain a fully parallel computational structure able to meet the throughput requirements of high resolution and high frame rate videos. High parallelism is achieved by predicting small blocks, i.e. 4x4 pixel blocks, in parallel and then adding them up in order to get Sum of Absolute Differences (SADs) of large block sizes. The predictor is able to support high resolution videos i.e. 1080p. The modified algorithm shows promising PSNR results with respect to full search algorithm. The predictor is synthesized with a clock frequency of 200 MHz, occupying an area of 0.49 mm2, on 90-nm Standard Cell ASIC technology. Ashfaq Ahmed, Muhammad Usman Shahid, Maurizio Martina, Enrico Magli, Guido Masera |
DSD | 5 |
| 2013 | A Joint Communication and Application Simulator for NoC-Based Custom SoCs: LDPC and Turbo Codes Parallel Decoding Case StudyabstractNoCs have become a widespread paradigm in the system-on-chip design world, not only for multi-purpose SoCs, but also for application-specific ICs. The common approach in the NoC design world is to separate the design of the interconnection from the design of the processing elements: this is well suited for a large number of developments, but the need for joint application and NoC design is not uncommon, especially in the application-specific case. The correlation between processing and communication tasks can be strong, and separate or trace-based simulations fall often short of the desired precision. In this work, the OMNET++ based JANoCS simulator is presented: concurrent simulation of processing and communication allow cycle-accurate evaluation of the system. The potential of the proposed approach is illustrated through a simple application example. Furthermore, a detailed case study on LDPC and turbo codes parallel decoding is presented. Results analysis illustrates the need for joint simulations and demonstrates the effectiveness of the proposed JANoCS. Carlo Condo, Amer Baghdadi, Guido Masera |
DSD | 3 |
| 2013 | Analysis on parallel implementations of fixed-complexity sphere decoder
Guido Masera |
Sci. China Inf. Sci. | 2 |
| 2013 | Power Control for Crossbar-Based Input-Queued SwitchesabstractWe consider an N × N Input-Queued (IQ) switch with a crossbar-based switching fabric implemented on a single chip. The power consumption produced by the crossbar chip, due to the data transfer, grows as NR3, where R is the maximum bit rate. Thus, at increasing bit rate, power dissipation is becoming more and more challenging, limiting the crossbar scalability for high-performance switches. We propose to exploit Dynamic Voltage and Frequency Scaling (DVFS) techniques to control packet transmissions through each crosspoint of the switching fabric. Our power control operates independently of the packet scheduler and exploits the knowledge of a traffic matrix obtained by online measurements. We propose a family of control algorithms to reduce the power consumption. The algorithms are particularly efficient in nonoverloaded conditions. The actual potential of the proposed approach is also evaluated on a real design case synthesized on a 90 nm CMOS technology. Andrea Bianco, Paolo Giaccone, Guido Masera, Marco Ricca |
IEEE Trans. Computers | 3 |
| 2012 | A Network-on-Chip-based turbo/LDPC decoder architectureabstractThe current convergence process in wireless technologies demands for strong efforts in the conceiving of highly flexible and interoperable equipments. This contribution focuses on one of the most important baseband processing units in wireless receivers, the forward error correction unit, and proposes a Network-on-Chip (NoC) based approach to the design of multi-standard decoders. High level modeling is exploited to drive the NoC optimization for a given set of both turbo and Low-Density-Parity-Check (LDPC) codes to be supported. Moreover, synthesis results prove that the proposed approach can offer a fully compliant WiMAX decoder, supporting the whole set of turbo and LDPC codes with higher throughput and an occupied area comparable or lower than previously reported flexible implementations. In particular, the mentioned design case achieves a worst-case throughput higher than 70 Mb/s at the area cost of 3.17 mm2on a 90 nm CMOS technology. Carlo Condo, Maurizio Martina, Guido Masera |
DATE | 3 |
| 2012 | Non-recursive max* operator with reduced implementation complexity for turbo decodingabstractIn this study, the authors deal with the problem of how to effectively approximate the max* operator when having n>2 input values, with the aim of reducing implementation complexity of conventional Log-MAP turbo decoders. They show that, contrary to previous approaches, it is not necessary to apply the max* operator recursively over pairs of values. Instead, a simple, yet effective, solution for the max* operator is revealed having the advantage of being in non-recursive form and thus, requiring less computational effort. Hardware synthesis results for practical turbo decoders have shown implementation savings for the proposed method against the most recent published efficient turbo decoding algorithms by providing near optimal bit error rate (BER) performance. Stylianos Papaharalabos, P. Takis Mathiopoulos, Guido Masera, Maurizio Martina |
IET Commun. | 3 |
| 2012 | Efficient VLSI implementation of soft-input soft-output fixed-complexity sphere decoderabstractFixed-complexity sphere decoder (FSD) is one of the most promising techniques for the implementation of multiple-input multiple-output (MIMO) detection, with relevant advantages in terms of constant throughput and high flexibility of parallel architecture. The reported works on FSD are mainly based on software level simulations and a few details have been provided on hardware implementation. The authors present the study based on a four-nodes-per-cycle parallel FSD architecture with several examples of VLSI implementation in 4×4 systems with both 16-quadrature amplitude modulation (QAM) and 64-QAM modulation and both real and complex signal models. The implementation aspects and details of the architecture are analysed in order to provide a variety of performance-complexity trade-offs. The authors also provide a parallel implementation of log-likelihood-ratio (LLR) generator with optimised algorithm to enhance the proposed FSD architecture to be a soft-input soft-output (SISO) MIMO detector. To the authors best knowledge, this is the first complete VLSI implementation of an FSD based SISO MIMO detector. The implementation results show that the proposed SISO FSD architecture is highly efficient and flexible, making it very suitable for real applications. Guido Masera |
IET Commun. | 2 |
| 2012 | High Speed Architectures for Finding the First two Maximum/Minimum ValuesabstractHigh speed architectures for finding the first two maximum/minimum values are of paramount importance in several applications, including iterative (e.g., turbo and low-density-parity-check) decoders. In this brief, stemming from a previous work, based on radix-2 solutions, we propose higher and mixed radix implementations that improve the architecture latency. Post place and route results on a 180-nm CMOS standard cell technology show that the proposed architectures achieve lower latency than radix-2 solutions with a moderate area increase. Luca G. Amarù, Maurizio Martina, Guido Masera |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | A Novel Architecture for Scalable, High Throughput, Multi-standard LDPC DecoderabstractThis paper presents a bottom up approach for implementing high throughput, scalable, layered LDPC decoding architecture for multi-standard applications. A generic implementation of fully parallel check node along with a block level Channel Memory organization scheme are two elements of novelty of this work. The proposed decoder IP core is synthesizable for all codes defined by WiMAX (WiFi) standards. Synthesis results are presented based on 130 nm standard cell ASIC technology. Muhammad Awais 0004, Ashwani Singh, Emmanuel Boutillon, Guido Masera |
DSD | 4 |
| 2011 | Look-ahead sphere decoding: algorithm and VLSI architectureabstractMultiple-input-multiple-output (MIMO) systems are recognised as a key enabling technology in high-performance wireless communications; however, the implementation of high-throughput MIMO detectors still is a critical task. Among known MIMO detectors, sphere decoder algorithm (SDA) is capable of optimal performance with acceptable processing complexity. The study presents an improved SDA, which enables significant throughput increase at a very limited additional complexity and with no degradation in terms of bit error rate performance. The modified detection method, called LASDA (look-ahead SDA), is based on formal algorithm transformations, namely look-ahead, retiming and pipelining, and requires a modified tree search strategy. The VLSI design of LASDA detector supporting a 4×4 MIMO channel with 16 QAM modulation is detailed in the study for a 130 nm CMOS standard cell technology: synthesis results show that the proposed solution achieves an average throughput of 380 Mbps at a signal-to-noise ratio of 22 dB (Rayleigh fading channel), with an occupied silicon area of 0.18 mm2. Comparisons with a number of previous implementations are also provided. Micaela Troglia Gamba, Guido Masera |
IET Commun. | 2 |
| 2010 | ALOE-Based Flexible LDPC DecoderabstractRadio communications terminals and infrastructure tend to support an increasing range of algorithms and radio access technologies. Flexible processing platforms are therefore needed for supporting multi-standard or heterogeneous radios. Channel decoding is one of the most computing demanding digital signal processing blocks of a radio transceiver. At the same time, it provides a high degree of implementation flexibility as well as facilitates dynamic parameter adjustments. This paper presents a flexible LDPC decoder implemented on an FPGA device following the ALOE middleware design paradigm. We analyse the middleware efficiency in terms of flexibility versus resource requirements. The results show a relative middleware area overhead of 32 %. Ismael Gómez Miguelez, Massimo Camatel, Jordi Bracke, Vuk Marojevic, Antoni Gelonch, Fabrizio Vacca, Guido Masera |
DSD | 7 |
| 2010 | A Novel VLSI Architecture of Fixed-Complexity Sphere DecoderabstractFixed-complexity Sphere Decoder (FSD) is a recently proposed technique for Multiple-Input Multiple-Output (MIMO) detection. It has several outstanding features such as constant throughput and large potential parallelism, which makes it suitable for efficient VLSI implementation. However, to our best knowledge, no VLSI implementation of FSD has been reported in the literature, although some FPGA prototypes of FSD with pipeline architecture have been developed. These solutions achieve very high throughput but at very high cost of hardware resources, making them impractical in real applications. In this paper, we present a novel four-nodes-per-cycle parallel architecture of FSD, with a breadth-first processing that allows for short critical path. The implementation achieves a throughput of 213.3 Mbps at 400 MHz clock frequency, at a cost of 0.18 mm2Silicon area on 0.13μm CMOS technology. The proposed solution is much more economical compared with the existing FPGA implementations, and very suitable for practical applications because of its balanced performance and hardware-complexity; moreover it has the flexibility to be expanded into an eight-nodes-per-cycle version in order to double the throughput. Guido Masera |
DSD | 2 |
| 2010 | Thermal Control for Crossbar-Based Input-Queued SwitchesabstractWe consider an N×N input-queued switch based on a crossbar switching fabric implemented on a single chip. The thermal power produced by the crossbar chip grows as N R3, where R is the maximum bit rate. Power dissipation is becoming more and more challenging, limiting the crossbar scalability for high performance switches. We propose to exploit Dynamic Voltage and Frequency Scaling (DVFS) techniques, quite commonly used in integrated circuit design, to control packet transmissions through each crosspoint of the switching fabric. Our thermal control operates independently of the packet scheduler and it is based on short-term traffic measurements. We propose a family of control algorithms to reduce the thermal power dissipation in non-overloaded conditions. Andrea Bianco, Paolo Giaccone, Guido Masera, Marco Ricca |
GLOBECOM | 3 |
| 2009 | Flexible Architectures for LDPC Decoders Based on Network on Chip ParadigmabstractThis paper explores the possibility of building a flexible Low Density Parity Check (LDPC) decoder using a network on chip communication infrastructure. Even if this idea is not completely new, previously published works suffered from an excessive area occupation and their practical impact has been very limited. In the following we analyze two possible NOCs specifically designed for the LDPC case. From synthesis results it can be observed how the proposed networks outperform previous implementations in terms of active area with no significant bandwidth loss. Finally to prove the effectiveness of the proposed approach a complete, partially parallel LDPC decoder design is presented and characterized in terms of throughput and area occupation. Fabrizio Vacca, Guido Masera, Hazem Moussa, Amer Baghdadi, Michel Jézéquel |
DSD | 2 |
| 2009 | A feasible VLSI engine for soft-input-soft-output for joint source channel codesabstractThis paper proposes for the first time, the very large scale integration (VLSI) architectural techniques for error resilient joint source channel coding (JSCC) of arithmetic codes (AC). When implemented on a 0.13 ¿m standard cells technology running at 340 MHz, achieves a decoding throughput of up to 125 kbit/s, 58 times better than the standard implementation. Simone Zezza, Guido Masera, Saeid Nooshabadi |
ICIP | 2 |
| 2009 | Efficient Implementation Techniques for Maximum Likelihood-Based Error Correction for JPEG2000abstractFeatures of the JPEG2000 compression standard include coding efficiency at low bit rates. However, its compressed bit stream is sensitive to transmission error. This paper presents three techniques to reduce both the computational complexity and the memory requirement in the ternary MQ arithmetic decoding. Such coders introduce a controlled degree of redundancy during the encoding process, which can be exploited at the decoder side in order to detect and correct errors. Our proposed techniques result in a substantial saving of decoding time and memory usage, with no or little degradation in the PSNR metric. Simone Zezza, Saeid Nooshabadi, Maurizio Martina, Guido Masera |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Decoding the Golden Code: A VLSI DesignabstractThe recently proposed Golden code is an optimal space-time block code for 2 times 2 multiple-input-multiple-output (MIMO) systems. The aim of this work is the design of a VLSI decoder for a MIMO system coded with the Golden code. The architecture is based on a rearrangement of the sphere decoding algorithm that achieves maximum-likelihood (ML) decoding performance. Compared to other approaches, the proposed solution exhibits an inherent flexibility in terms of QAM modulation size and this makes our architecture particularly suitable for adaptive modulation schemes. Relying on the flexibility of this approach two different architectures are proposed: a parametric one able to achieve high decoding throughputs (> 165 Mb/s) while keeping low overall decoder complexity (45 KGates), a flexible implementation able to dynamically adapt to the modulation scheme (4-,16-,64-QAM) retaining the low complexity and high throughput features. Barbara Cerato, Guido Masera, Emanuele Viterbo |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | VLSI implementation of SISO arithmetic decoders for joint source channel codingabstractIn this paper we propose an efficient VLSI implementation of a Soft Input Soft Output (SISO) arithmetic code (AC) decoder for joint source channel coding. The addressed application shows a very high level of processing complexity, but, to the best of our knowledge, no papers have been published in the literature on the hardware implementation of the considered joint source channel scheme. First we introduce a simplified algorithm for the SISO AC, which is 1.3 times faster than the standard one. Then an efficient SISO AC architecture is proposed and synthesis results on a 0.13 mum standard cells technology are reported for two different sets of parameters (M=128, M=256). The proposed core runs at 338.9 MHz and can decode up to 124.987 kbit/s. Simone Zezza, Guido Masera |
DATE | 2 |
| 2008 | Error resilient JPEG2000 decoding for wireless applicationsabstractTo improve the JPEG2000 compression standard error resiliency in the wireless environment, the use of ternary MQ arithmetic coders/decoders that are based on the concept of forbidden symbol has been proposed. This paper presents two ternary MQ based techniques to reduce both the computational complexity and the memory requirement during the decoding process, with no or little degradation in the PSNR. Simone Zezza, Maurizio Martina, Guido Masera, Saeid Nooshabadi |
ICIP | 3 |
| 2007 | Hardware architecture for matrix factorization in mimo receiversabstractThis paper presents the hardware realization of the factorization algorithm required in a MIMO OFDM receiver to make the detection and decoding a non-orthogonal space-time code. Requirements of a real scenario represented by the standard IEEE 802.11n for WLAN have been analyzed and exploited to draw out the specifications of the proposed implementation. A very high throughput hardware realization has been obtained able to factorize 128 8x8 real channel matrices during the channel updating period of 28 μs, with a final throughput of 4,63 millions of matrices processed per second. Synthesis results on both 0.13 μm CMOS standard cell technology and FPGA compare favourably to previous implementations. Barbara Cerato, Guido Masera, Peter Nilsson 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Beyond 3G wireless communication system prototypeabstractModern wireless systems have to achieve very high data rates over noisy channels. Channel codes are usually used to guarantee reliable communication, but validating a system with extremely low Bit Error Rate (BER) requires the simulation of a huge number of samples. The simulation complexity furtherly increases if we want to take into account the joint effect of channel codes and complex modulation techniques, such as Orthogonal Frequency Division Multiplexing (OFDM), over multi-path fading channel. The use of software simulators to evaluate system performance in several scenarios may result in unacceptably long times and a reconfigurable hardware prototyping environment may be an effective solution. This paper describes the implementation of a complete real-time, fully digital, flexible high performance hardware/software prototype for beyond 3G wireless communications. The transmitter/receiver chain includes several innovative characteristics: Serial Concatenated Convolutional Codes (SCCC) Turbo Encoder/Decoder, adaptive OFDM modulation, versatile multiple user access scheme and a sophisticated, real-time power control.Optimal choices for partitioning and internal data representation have been described as well as internal architecture and measurement environment. Our programmable prototype ensures high flexibility and speedup of more than 4000 times compared with the software version. Alberto Dassatti, Simone Zezza, Mario Nicola, Guido Masera |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Flexible blocks for high throughput serially concatenated convolutional codesabstractModern communication systems aim to integrate different services including telephony, multimedia applications and data transfer. This integration imposes to design flexible systems, able to support different frame duration and data-rates. In this paper we present several blocks employed in serially concatenated convolutional codes and we detail some hardware solutions to achieve flexibility. Maurizio Martina, Guido Masera |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Real-time implementation of a time-frequency analysis schemeabstractThe on-line monitoring and detection of defects in laser welding is a basic manifacturing requirement in several applicative contexts, including vehicle assembly in automotive production. The speed of assembly in modern industry and the large amount of data to be acquired and elaborated pose severe real-time constraints and lead to the need of extremely short processing latencies. In this work, the implementation of time-frequency analysis algorithms on an FPGA device is shown and compared to pure software developments on different processors. Maurizio Martina, Andrea Terreno, Fabrizio Vacca, Andrea Molino, Guido Masera, Giuseppe D'Angelo, Giorgio Pasquettaz |
ACM Great Lakes Symposium on VLSI | 5 |
| 2006 | Design of Application Specific Processors for the Cached FFT AlgorithmabstractOrthogonal frequency division multiplexing (OFDM) is a data transmission technique which is used in wired and wireless digital communication systems. In this technique, fast Fourier transformation (FFT) and inverse FFT (IFFT) are kernel processing blocks in an OFDM system, and are used for data (de)modulation. OFDM systems are increasingly required to be flexible to accommodate different standards and operation modes, in addition to being energy-efficient. A trade-off between these two conflicting requirements can be achieved by employing application-specific instruction-set processors (ASIPs). In this paper, two ASIP design concepts for the cached FFT algorithm (CFFT) are presented. A reduction in energy dissipation of up to 25% is achieved compared to an ASIP for the widely used Cooley-Tukey FFT algorithm, which was designed by using the same design methodology and technology. Further, a modified CFFT algorithm which enables a better cache utilization is presented. This modification reduces the energy dissipation by up to 10% compared to the original CFFT implementation Oguzhan Atak, Abdullah Atalar, Erdal Arikan, Harold Ishebabi, David Kammler, Gerd Ascheid, Heinrich Meyr, Mario Nicola, Guido Masera |
ICASSP (3) | 9 |
| 2006 | Low-Complexity Video Compression Combining Adaptive Multifoveation and Reuse of High-Resolution InformationabstractThe phenomenon of reduced spatial resolution perceived away from the point of gaze (foveation point) in a scene by the human visual system can be gainfully exploited in image and video compression. Interest has recently evolved from single foveation points to dynamic and multiple points of foveation, implementing which entails significantly increased complexity in the encoder. In this paper, a novel idea of efficiently combining adaptive multipoint foveation with salvaged high-resolution information for reuse in real-time video to maintain higher resolution in peripheral regions is proposed. The idea is implemented with a fast algorithm for multi-foveation processing, in conjunction with standard-compliant decoding. The new multi-foveation algorithm is integrated with the H.264/AVC standard for testing. Simulation results show a compression gain ranging from 2.25% to more than 11%, without degrading the perceived quality and PSNR and with minimal addition to the complexity of a standard uniform-resolution codec. Giorgio Pioppo, Rashid Ansari, Ashfaq Khokhar 0001, Guido Masera |
ICIP | 4 |
| 2006 | Mumford and Shah Functional: VLSI Analysis and ImplementationabstractThis paper describes the analysis of the Mumford and Shah functional from the implementation point of view. Our goal is to show results in terms of complexity for real-time applications, such as motion estimation based on segmentation techniques, of the Mumford and Shah functional. Moreover, the sensitivity to finite precision representation is addressed, a fast VLSI architecture is described, and results obtained for its complete implementation on a 0.13 microm standard cells technology are presented. Maurizio Martina, Guido Masera |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | High Performance Channel Model Hardware Emulator for 802.11n
Alberto Dassatti, Guido Masera, Mario Nicola, Andrea Concil, Angelo Poloni |
FPT | 2 |
| 2005 | Low-complexity, efficient 9/7 wavelet filters implementationabstractThis paper proposes a novel low-complexity, efficient 9/7 wavelet filters implementation for image compression applications. The 9/7 wavelet filters are widely used in different image compression schemes, such as the JPEG2000 image coding standard. Thus the implementation of efficient codecs is of great concern. The performance of a hardware implementation of the 9/7 filter bank depends on the accuracy with which filter coefficients are represented. However the greatest part of current implementations consider filters taps as numbers to be implemented. The aim of this work is to show that great complexity reduction can be achieved going through the derivation of the 9/7 taps values. Maurizio Martina, Guido Masera |
ICIP (3) | 2 |
| 2005 | Optimized CORDIC core for frequency-domain motion estimationabstractThis paper describes a CORDIC-based architecture to efficiently compute the phase difference between two complex numbers. The problem of fast phase difference computation is central in many signal processing algorithms. Our main focus has been posed on the phase correlation technique applied to motion estimation. A reduced complexity solution is proposed and specifically tailored to suit the application needs. The presented algorithm has been completely implemented in 0.25 /spl mu/m standard-cell CMOS technology. As far as the performance are concerned the designed core outperforms a recently designed solution by more than 50% under area and energy standpoints. Andrea Molino, Fabrizio Vacca, Guido Masera |
ICIP (3) | 3 |
| 2004 | Testing Logic Cores using a BIST P1500 Compliant Approach: A Case of StudyabstractIn this paper we describe how we applied a BIST-based approach to the test of a logic core to be included in system-on-a-chip (SoC) environments. The approach advantages are the ability to protect the core IP, the simple test interface (thanks also to the adoption of the P1500 standard), the possibility to run the test at-speed, the reduced test time, and the good diagnostic capabilities. The paper reports figures of the achieved fault coverage, the required area overhead, and the performance slowdown, and compares the figures with those for alternative approaches, such as those based on full scan and sequential ATPG. Paolo Bernardi 0002, Guido Masera, Federico Quaglio, Matteo Sonza Reorda |
DATE | 2 |
| 2004 | An electromigration and thermal model of power wires for a priori high-level reliability predictionabstractIn this paper, a simple power-distribution electrothermal model including the interconnect self-heating is used together with a statistical model of average and rms currents of functional blocks and a high-level model of fanout distribution and interconnect wirelength. Following the 2001 SIA roadmap projections, we are able to predict a priori that the minimum width that satisfies the electromigration constraints does not scale like the minimum metal pitch in future technology nodes. As a consequence, the percentage of chip area covered by power lines is expected to increase at the expense of wiring resources unless proper countermeasures are taken. Some possible solutions are proposed in the paper. Mario R. Casu, Mariagrazia Graziano, Guido Masera, Gianluca Piccinini, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Wireless sensor networks: a power-scalable motion estimation IP for hybrid video codingabstractWireless Sensor Networks are an emerging phenomenon in the research community. The design and development of network architectures and nodes implementation are fostering many research activities. Due to their wide application fields and pervasive employment possibilities, the investigation of novel classes of wireless sensor nodes is of great concern. In this paper we presented a novel Power-Scalable Motion Estimation IP suitable for video-surveillance over Wireless Sensor Networks. The proposed architecture can achieve low dynamic power, good video quality and reasonable frame-rates. Moreover, it can operate on different video format and can be reconfigured both off-line and on-line. Further researches have to be accomplished in order to reduce the internal critical path, achieving better maximum frequencies performances. Moreover, new FPGA architectures must be exploited searching low-static-power devices suitable for an actual implementation of our IP on a WSN. Federico Quaglio, Maurizio Martina, Fabrizio Vacca, Guido Masera, Andrea Molino, Gianluca Piccinini, Maurizio Zamboni |
FPGA | 4 |
| 2002 | Energy Evaluation on a Reconfigurable, Multimedia-Oriented Wireless Sensor
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni |
FPL | 2 |
| 2002 | Reconfigurable DSP IP for multimedia applicationsabstractIn this paper a novel Digital Signal Processor IP for multimedia applications, is presented. Recently, develeper's interest towards SOC architectures has been driven by mobile market explosion. Despite the increasing importance gathered by reconfigurable computing, a lack of easily retargettable cores is felt by developer's community. This IP is intended to be the kernel for many telecommunication and multimedia algorithms computation. During the design flow, much care has been devoted to grant maximum interoperability among this DSP core and other coprocessor units, allowing to easily embed multiple functional blocks on a single FPGA. As far as performance are concerned, the proposed IP shows satisfactory results both in terms of area occupation (11\% on a XILINX XCV1000) and maximum clock frequency (89 MHz after place and route process). Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni |
ICASSP | 2 |
| 2002 | Architectural strategies for low-power VLSI turbo decodersabstractThe use of "turbo codes" has been proposed for several applications, including the development of wireless systems, where highly reliable transmission is required at very low signal-to-noise ratios (SNR). The problem of extracting the best coding gains from these kind of codes has been deeply investigated in the last years. Also the hardware implementation of turbo codes is a very challenging topic, mainly due to the iterative nature of the decoding process, which demands an operating frequency much higher than the data rate; in the case of wireless applications, the design constraints became even more strict due to the low-cost and low-power requirements. This paper first presents a new architecture for the decoder core with improved area and power dissipation properties; then partitioning techniques are proposed to reduce the power consumption of the decoder memories. It is proven that most of the power is dissipated by the large RAM units required by the decoder, so the described technique is very efficient: an average power saving of 70% with an area overhead of 23% has been obtained on a set of analyzed architectures. Guido Masera, Marco Mazza, Gianluca Piccinini, Fabrizio Viglione, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Synthesis of low-leakage PD-SOI circuits with body-biasingabstractArticle Share on Synthesis of low-leakage PD-SOI circuits with body-biasing Authors: Mario Casu Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Gianluca Piccinini Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Guido Masera Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile , Maurizio Zamboni Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, Italy Politecnico di Torino, C.so Duca degli Abruzzi, 24, I-10129 Torino, ItalyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 287–290https://doi.org/10.1145/383082.383170Online:06 August 2001Publication History 2citation166DownloadsMetricsTotal Citations2Total Downloads166Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Mario R. Casu, Gianluca Piccinini, Guido Masera, Maurizio Zamboni |
ISLPED | 3 |
| 2000 | A 50 Mbit/s Iterative Turbo-DecoderabstractVery low bit error rate has become an important constraint in high performance communication systems that operate at very low signal to noise ratios: due to their impressive coding gains, turbo codes have been proposed for several applications, although they suffer a large decoding delay. This paper presents the design of a turbo decoder with high performances in terms of throughput implemented using TSPC (true single phase clocking) logic family. In order to achieve the best compromise between cost (in terms of area) and throughput, several architectural solutions have been analyzed. The whole system and in particular its core, the SISO module, has been verified through VHDL simulations. HSPICE simulations show that the system can operate with a 1 GHz clock and thus it can reach a throughput of 50 Mbit/s. Fabrizio Viglione, Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni |
DATE | 2 |
| 2000 | A high accuracy-low complexity model for CMOS delaysabstractThis paper presents a new model for CMOS structures delays estimation based on a deep analysis of complex gates behavior. This approach can supply a high level of accuracy. A complex structure is reduced first to series-connected MOS, then the delay equations are applied to that reduced rate. The model is based on a time piecewise linearization so that a strongly nonlinear circuit can he solved using well known linear techniques. The delay formulas involve model parameters as MOS width functions, therefore providing routines suitable for optimization algorithms. The high level of accuracy, the low CPU time and the high degree of scaling capability are proved in the paper. These features make the model attractive for deep submicron technologies. Mario R. Casu, Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni |
ISCAS | 2 |
| 1999 | New 2 Gbit/s CMOS I/O padsabstractA couple of low complexity high performance input and output pads are proposed: they have been designed in 0.7 /spl mu/m CMOS ES2 technology and support bit rates ranging from DC up to 2 Gbit/s. The differential input pad and the differential output pad interface true PECL external logic levels to full swing 5 V CMOS internal levels. Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni |
Great Lakes Symposium on VLSI | 1 |
| 1999 | VLSI architectures for turbo codesabstractA great interest has been gained in recent years by a new error-correcting code technique, known as "turbo coding", which has been proven to offer performance closer to the Shannon's limit than traditional concatenated codes. In this paper, several very large scale integration (VLSI) architectures suitable for turbo decoder implementation are proposed and compared in terms of complexity and performance; the impact on the VLSI complexity of system parameters like the state number, number of iterations, and code rate are evaluated for the different solutions. The results of this architectural study have then been exploited for the design of a specific decoder, implementing a serial concatenation scheme with 2/3 and 3/4 codes; the designed circuit occupies 35 mm/sup 2/, supports a 2 Mb/s data rate, and for a bit error probability of 10/sup -6/, yields a coding gain larger than 7 dB, with ten iterations. Guido Masera, Gianluca Piccinini, Massimo Ruo Roch, Maurizio Zamboni |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1996 | A Parametrical Architecture for Reed-Solomon DecodersabstractReed-Solomon decoders are digital decoders that use RS detecting and correcting of errors codes. RS codes are widely diffused in the transmission and storage of digital information and they are often used in concatenated encoding schemes to obtain great correction capabilities and good robustness to burst errors. In this study, a parametrical approach was chosen for decoder implementation at gate-level, based on the Berlekamp algorithm. This means that the decoder structure depends on two parameters: the bit number used for the symbol representation (m), and the error correction capability (t). The obtained architecture is suitable for a large number of different application (including high definition digital TV) and can be quickly synthesised using Synopsys for any required values of m and t. Mariana-Eugenia Petre, Guido Masera |
Great Lakes Symposium on VLSI | 2 |
| 1989 | Encoded 16-PSK: a study for the receiver designabstractThe authors present the results of a design study of the receiver in a digital transmission system using the combined coding and modulation schemes known as Ungerboeck codes. Specifically, they examine the design of the receiver for encoded 16-PSK (phase shift keying) modulation, presenting first the traditional structure for the optimum receiver and then a simpler structure. The decoding depth of the Viterbi algorithm, the quantization of the metrics inside the Viterbi processor, and the phase jitter in the recovered carrier are considered. The impact of branch and path metric quantization inside the receiver is discussed, showing that a reasonable number of bits (8) is sufficient to obtain nearly optimum performance when the code complexity is limited. The effect of imperfect carrier recovery inside the receiver is studied, providing accurate analytical estimates of the error event probability as well as an upper bound to the symbol error probability. Results of a detailed simulation, including carrier and bit timing recovery blocks, show that the effects of imperfections on the bit error probability are very small, even at low signal-to-noise ratios. On the whole, results show the robustness of the Viterbi algorithm with respect to fairly rough quantizations of the metrics and indicate that carrier recovery is not as critical as expected.> Sergio Benedetto, Marco Ajmone Marsan, Guido Masera, Gabriella Olmo |
IEEE J. Sel. Areas Commun. | 3 |