Maurizio Martina

dblp:26/3014 · DBLP profile ↗
← Back
68ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0002-3069-0319ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CIRCE CROSS Integrated RISC-V Cryptographic Extension
Alessandra Dolmeta, Valeria Piscopo, Maurizio Martina, Guido Masera
DATE3
2026 Compact Yet Fast: An Efficient d-Order Masked Implementation of Ascon
abstract
In this work, we present a generic side-channel protected design of Ascon that achieves high efficiency by dynamically reconfiguring the hardware countermeasures during message processing. The resultant implementation is protected and capable of meeting stringent performance requirements whilst minimising resource overhead. The experimental results obtained demonstrate that the implementation meets the required security and achieves superior throughput-to-area ratio across all protection orders. Ascon, recently selected by NIST as the lightweight cryptography standard, is widely deployed in resource-constrained devices that demand both high performance and resistance against threats such as side-channel analysis (SCA). Exploiting Ascon’s mode-level structure, which does not require protection against differential power analysis during bulk operations, we introduce a modified masking gadget with dual functionality: serving as a countermeasure during critical operations, and processing multiple data paths in parallel to accelerate bulk computation. Our architecture supports any configurable security order and instantiates only the minimum hardware resources needed to maximize throughput per round. We also evaluate an enhanced Ascon architecture based on the Changing of the Guards technique, which eliminates the need for fresh randomness. Security validation is performed using fixed-vs-random t-tests on both first- and second-order masked implementations. Finally, we compare our masked design against state-of-the-art solutions.
Mattia Mirigaldi, Nico Paninforni, Maurizio Martina, Guido Masera
DATE3
2025 ARCANE: Adaptive RISC-V Cache Architecture for Near-memory Extensions
abstract
Modern data-driven applications expose limitations of von Neumann architectures-extensive data movement, low throughput, and poor energy efficiency. Accelerators improve performance but lack flexibility and require data transfers. Existing compute in- and nearmemory solutions mitigate these issues but face usability challenges due to data placement constraints. We propose a novel cache architecture that doubles as a tightly-coupled compute-near-memory coprocessor. Our RISC-V cache controller executes custom instructions from the host CPU using vector operations dispatched to near-memory vector processing units within the cache memory subsystem. This architecture abstracts memory synchronization and data mapping from application software while offering software-based Instruction Set Architecture extensibility. Our implementation shows $30 \times$ to $84 \times$ performance improvement when operating on 8-bit data over the same system with a traditional cache when executing a worst-case 32-bit CNN workload, with only 41.3% area overhead.
Vincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone, Guido Masera, Maurizio Martina
DAC6
2025 TYRCA: A RISC-V Tightly-Coupled Accelerator for Code-Based Cryptography
abstract
Post-quantum cryptography (PQC) has garnered significant attention across various communities, particularly with the National Institute of Standards and Technology (NIST) advancing to the fourth round of PQC standardization. One of the leading candidates is Hamming Quasi-Cyclic (HQC), which received a significant update on February 23, 2024. This update, which introduces a classical dense-dense multiplication approach, has no known dedicated hardware implementations yet. The innovative Core-V eXtension InterFace (CV-X-IF) is a communication interface for RISC-V processors that significantly facilitates the integration of new instructions to the Instruction Set Architecture (ISA), through tightly connected accelerators. In this paper, we present a TightlY-coupled accelerator for RISC-V for Code-based cryptogrAphy (TYRCA), proposing the first fully tightly-coupled hardware implementation of the HQC-PQC algorithm, leveraging the CV-X-IF. The proposed architecture is implemented on the Xilinx Kintex-7 FPGA. Experimental results demonstrate that TYRCA reduces the execution time by 94% to 96% for HQC-128, HQC-192, and HQC-256, showcasing its potential for efficient HQC code-based cryptography.
Alessandra Dolmeta, Stefano Di Matteo, Emanuele Valea, Mikael Carmona, Antoine Loiseau, Maurizio Martina, Guido Masera
DATE6
2025 TinyCL: An Efficient Hardware Architecture for Continual Learning on Autonomous Systems
abstract
The Continuous Learning (CL) paradigm consists of continuously evolving the parameters of the Deep Neural Network (DNN) model to progressively learn to perform new tasks without reducing the performance on previous tasks, i.e., avoiding the so-called catastrophic forgetting. However, the DNN parameter update in CL-based autonomous systems is extremely resource-hungry. The existing DNN accelerators cannot be directly employed in CL because they only support the execution of the forward propagation. Only a few prior architectures execute the backpropagation and weight update, but they lack the control and management for CL. Towards this, we design a hardware architecture, TinyCL, to perform CL on resource-constrained autonomous systems. It consists of a processing unit that executes both forward and backward propagation, and a control unit that manages memory-based CL workload. To minimize the memory accesses, the sliding window of the convolutional layer moves in a snake-like fashion. Moreover, the Multiply-and-Accumulate units can be reconfigured at runtime to execute different operations. As per our knowledge, our proposed TinyCL represents the first hardware accelerator that executes CL on autonomous systems. We synthesize the complete TinyCL architecture in a 65 nm CMOS technology node with the conventional ASIC design flow. It executes 1 epoch of training on a Conv + ReLU + Dense model on the CIFAR10 dataset in 1.76 s, while 1 training epoch of the same model using an Nvidia Tesla P100 GPU takes 103 s, thus achieving a 58× speedup, consuming 86 mW in a 4.74 mm2die.
Eugenio Ressa, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IJCNN3
2025 RISC-V Based Keccak Co-Processor for NIST Post-Quantum Cryptography Standards
abstract
This paper presents the design and implementation of a RISC-V-based Keccak co-processor optimized for Post-Quantum Cryptography (PQC) algorithms. Leveraging the Core-V eXtension InterFace (CV-X-IF), the co-processor extends the Instruction Set Architecture (ISA) with three custom instructions tailored for cryptographic operations. This allows seamless integration into various PQC schemes, tested across the multiple standards proposed by the National Institute of Standards and Technology (NIST), including CRYSTALS-Kyber, CRYSTALS-Dilithium, SPHINCS+, and FALCON, which are designed to withstand quantum attacks. By employing tightly coupled hardware acceleration, the Keccak co-processor dramatically reduces the computational overhead of hash-based operations central to these algorithms. The implementation is realized on a Xilinx Artix 7 FPGA, achieving a clock cycles’ improvement up to 75% and 19% resource overhead. The results presented herein demonstrate significant performance enhancement over the state of the art, underscoring its effectiveness for cryptographic applications.
Alessandra Dolmeta, Valeria Piscopo, Mattia Mirigaldi, Maurizio Martina, Guido Masera
ISCAS4
2025 Scalability analysis of multi-bank near-memory computing in low-power SoCs
abstract
Machine learning and artificial intelligence are moving towards the edge, where the need for high throughput with a constrained energy budget is more urgent than ever. During the last few years, near-memory computing has emerged as a promising solution to address the memory bandwidth and energy efficiency limitations of conventional von Neumann systems. The recently proposed NM-Carus architecture combines vectororiented computing capabilities within a RISC-V programmable, configurable, and autonomous memory macro, addressing the usability of near-memory computing from a software deployment standpoint. In this paper, we explore the scalability of NMCarus in terms of computation parallelism, memory size and energy consumption, As a benchmarking platform, we rely on a low-power microcontroller that features multiple instances of NM-Carus that target the execution of biomedical applications. This exploration was performed on 16 nm TSMC NM-Carus implmentation, and we highlighted the benefits of technology scaling for a previous implementation on 65 nm with respect to the overhead of replacing conventional on-chip data SRAMs with near-memory computing banks. Overall, the paper presents a solid baseline regarding the trade-offs in terms of area, performance, and energy efficiency of integrating programmable near-memory computing in an existing edge-oriented system on chip towards efficient edge AI architectures at the system level.
Luigi Giuffrida, Pasquale Davide Schiavone, Michele Caon, Guido Masera, Maurizio Martina, David Atienza 0001
VLSI-SoC5
2024 VirtLAB-UI: An Open, Platform Independent, Software and Firmware Solution for Remote and Take-Home Labs
abstract
SARS-CoV2 pandemic pushed university courses toward remote access. Besides classroom lessons, for which IT solutions were already present, laboratory experiences required the development of suitable replacements. For electronics engineering courses, several solutions have been developed, mainly based on computer simulations, or simplified lab kits distributed to students to perform laboratory experiences at home. But if the hardware tools developed to mimic lab experiments are intended to eventually replace classic experiences, a similar visual user experience must be implemented, as well, with user interfaces similar to what is available on real measurements instruments, independent from the physical devices used. In this paper, a suitable approach for the development of such an interface is described, with an example application to an existing take-home lab system.
Massimo Ruo Roch, Maurizio Martina
EDUCON2
2024 A Case Study on Formal Equivalence Verification Between a C/C++ Model and Its RTL Design
abstract
Abstract In the field of communication system products, most datapath Digital Signal Processing algorithms are initially developed at a high-level in MATLAB® or C/C++. Subsequently, design engineers use these models as a reference for implementing Register Transfer Level designs. The conventional approach to verify their equivalence involves extensive Universal Verification Methodology dynamic simulations, which can last for months and require significant verification efforts. However, some elusive errors might still occur because it is infeasible to explore all input combinations with this method. On the other hand, Formal Equivalence Verification aims to verify that a Register Transfer Level design is functionally equivalent to the reference high-level C/C++ model across all possible legal states. With recent advancements in formal solver technology, Formal Equivalence Verification provides a distinct benefit by using mathematical methods to ensure that the Register Transfer Level (timed) matches the original high-level C/C++ model (untimed). This drastically reduces the verification time and ensures the exhaustive coverage of the design state space. This paper presents an in-depth exploration of complex Finite State Machine with datapath verification, specifically focusing on Multiplier-Accumulator, Tone Generator, and Automatic Gain Control, by employing the formal equivalence methodology. Although these signal processing blocks were previously verified throughout Universal Verification Methodology dynamic simulations, Formal Equivalence Verification was able to identify hard-to-find bugs in just a few weeks by utilizing the new workflow, thereby streamlining the verification process.
Gaetano Raia, Gianluca Rigano, David Vincenzoni, Maurizio Martina
FM (2)4
2024 MARLIN: A Co-Design Methodology for Approximate ReconfigurabLe Inference of Neural Networks at the Edge
abstract
The optimization of neural networks (NNs) is necessary to enable their deployment on energy-constrained devices. State-of-the-art methods leverage approximate multipliers to execute NNs reducing the inference energy without heavily affecting the accuracy. However, previous works usually require a specialized hardware accelerator and are limited to fixed multipliers or reconfigurable ones with few approximation levels. This paper introduces MARLIN, a framework to deploy layerwise approximate NNs on PULP, a microcontroller with a RISC-V core. A multiplier architecture, with runtime selection of 256 approximation levels, is developed and integrated into the PULP cluster cores, enabling runtime configuration through control status register (CSR) instructions embedded within the code. The PULP toolchain is adapted to incorporate the approximation level selection within the instruction flow seamlessly. MARLIN leverages the genetic algorithm NSGA-II to search for the best configurations among thousands of approximate NNs. The framework is validated by simulating an approximate NN trained with the MNIST dataset on PULP. Moreover, MARLIN is used to optimize and approximate six ResNet models trained with the CIFAR-10 dataset. In particular, for ResNet-56, the most complex NN used in the experiments, the multiplication energy is reduced by 23.9% while retaining 99% of the accuracy of the exact model.
Flavia Guella, Emanuele Valpreda, Michele Caon, Guido Masera, Maurizio Martina
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Implementation and integration of Keccak accelerator on RISC-V for CRYSTALS-Kyber
abstract
One of the key metrics used for defying the security of the Internet of Things (IoT) is data integrity, which mostly relies on the use of cryptographic hash functions. In the last years, the National Institute of Standards and Technology (NIST) announced SHA-3 as the new standard for better security. SHA-3 is also exploited in most of the current post-quantum cryptographic (PQC) protocols. Nevertheless, the used algorithm, i.e. Keccak, is computationally heavy and consequently limits its utilization in RISC-V-based Systems on Chip (SoC). In this work, a Keccak accelerator is proposed to speed up SHA3 computations for the CRYSTALS-Kyber algorithm on the RISCV-based advanced microcontroller PULPissimo. Compared to the plain SW implementation on RISC-V, our results show a speedup factor of up to 2.79 at the expense of a 12.4% resources overhead.
Alessandra Dolmeta, Mattia Mirigaldi, Maurizio Martina, Guido Masera
CF3
2023 SwiftTron: An Efficient Hardware Accelerator for Quantized Transformers
abstract
Transformers' compute- intensive operations pose enormous challenges for their deployment in resource- constrained EdgeAI / tiny ML devices. As an established neural network compression technique, quantization reduces the hardware computational and memory resources. In particular, fixed-point quantization is desirable to ease the computations using lightweight blocks, like adders and multipliers, of the underlying hardware. However, deploying fully-quantized Transformers on existing general-purpose hardware, generic AI accelerators, or specialized architectures for Transformers with floating-point units might be infeasible and/or inefficient. Towards this, we propose SwiftTron, an efficient specialized hardware accelerator designed for Quantized Transformers. SwiftTron supports the execution of different types of Transformers' operations (like Attention, Softmax, GELU, and Layer Normalization) and accounts for diverse scaling factors to perform correct computations. We synthesize the complete SwiftTron architecture in a 65 nm CMOS technology with the ASIC design flow. Our Accelerator executes the RoBERTa-base model in 1.83 ns, while consuming 33.64 mW power, and occupying an area of 273 mm2• To ease the reproducibility, the RTL of our SwiftTron architecture is released at https://github.com/albertomarchisio/SwiftTron.
Alberto Marchisio, Davide Dura, Maurizio Capra, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IJCNN4
2023 RobCaps: Evaluating the Robustness of Capsule Networks against Affine Transformations and Adversarial Attacks
abstract
Capsule Networks (CapsNets) are able to hierarchically preserve the pose relationships between multiple objects for image classification tasks. Other than achieving high accuracy, another relevant factor in deploying CapsNets in safety-critical applications is the robustness against input transformations and malicious adversarial attacks. In this paper, we systematically analyze and evaluate different factors affecting the robustness of CapsN ets, compared to traditional Convolutional Neural Networks (CNNs). Towards a comprehensive comparison, we test two CapsNet models and two CNN models on the MNIST, GTSRB, and CIFAR10 datasets, as well as on the affine-transformed versions of such datasets. With a thorough analysis, we show which properties of these architectures better contribute to increasing the robustness and their limitations. Overall, CapsNets achieve better robustness against adversarial examples and affine transformations, compared to a traditional CNN with a similar number of parameters. Similar conclusions have been derived for deeper versions of CapsNets and CNNs. Moreover, our results unleash a key finding that the dynamic routing does not contribute much to improving the CapsNets' robustness. Indeed, the main generalization contribution is due to the hierarchical feature learning through capsules.
Alberto Marchisio, Antonio De Marco, Alessio Colucci, Maurizio Martina, Muhammad Shafique 0001
IJCNN4
2022 Mind the Scaling Factors: Resilience Analysis of Quantized Adversarially Robust CNNs
abstract
As more deep learning algorithms enter safety-critical application domains, the importance of analyzing their resilience against hardware faults cannot be overstated. Most existing works focus on bit-flips in memory, fewer focus on compute errors, and almost none study the effect of hardware faults on adversarially trained convolutional neural networks (CNNs). In this work, we show that adversarially trained CNNs are more susceptible to failure due to hardware errors when compared to vanilla-trained models. We identify large differences in the quantization scaling factors of the CNNs which are resilient to hardware faults and those which are not. As adversarially trained CNNs learn robustness against input attack perturbations, their internal weight and activation distributions open a backdoor for injecting large magnitude hardware faults. We propose a simple weight decay remedy for adversarially trained models to maintain adversarial robustness and hardware resilience in the same CNN. We improve the fault resilience of an adversarially trained ResNet56 by 25% for large-scale bit-flip benchmarks on activation data while gaining slightly improved accuracy and adversarial robustness.
Nael Fasfous, Lukas Frickenstein, Michael Neumeier, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Maurizio Martina, Walter Stechele
DATE7
2022 AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic Algorithms
abstract
We present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation.
Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele
DATE11
2022 NLCMAP: A Framework for the Efficient Mapping of Non-Linear Convolutional Neural Networks on FPGA Accelerators
abstract
This paper introduces NLCMap, a framework for the mapping space exploration targeting Non-Linear Convolutional Networks (NLCNs). NLCNs [1] are a novel neural network model that improves performances in certain computer vision applications by introducing a non-linearity in the weights computation. NLCNs are more challenging to efficiently map onto hardware accelerators if compared to traditional Convolutional Neural Networks (CNNs), due to data dependencies and additional computations. To this aim, we propose NL-CMap, a framework that, given an NLC layer and a generic hardware accelerator with a certain on-chip memory budget, finds the optimal mapping that minimizes the accesses to the off-chip memory, which are often the critical aspect in CNNs acceleration.
Giuseppe Aiello, Beatrice Bussolino, Emanuele Valpreda, Massimo Ruo Roch, Guido Masera, Maurizio Martina, Stefano Marsi
ICIP6
2022 CoNLoCNN: Exploiting Correlation and Non-Uniform Quantization for Energy-Efficient Low-precision Deep Convolutional Neural Networks
abstract
In today's era of smart cyber-physical systems, Deep Neural Networks (DNNs) have become ubiquitous due to their state-of-the-art performance in complex real-world applications. The high computational complexity of these networks, which translates to increased energy consumption, is the foremost obstacle towards deploying large DNNs in resource-constrained systems. Fixed-Point (FP) implementations achieved through post-training quantization are commonly used to curtail the energy consumption of these networks. However, the uniform quantization intervals in FP restrict the bit-width of data structures to large values due to the need to represent most of the numbers with sufficient resolution and avoid high quantization errors. In this paper, we leverage the key insight that (in most of the scenarios) DNN weights and activations are mostly concentrated near zero and only a few of them have large magnitudes. We propose CoNLoCNN, a framework to enable energy-efficient low-precision deep convolutional neural network inference by exploiting: (1) non-uniform quantization of weights enabling simplification of complex multiplication operations; and (2) correlation between activation values enabling partial compensation of quantization errors at low cost without any run-time overheads. To significantly benefit from non-uniform quantization, we also propose a novel data representation format, Encoded Low-Precision Binary Signed Digit, to compress the bit-width of weights while ensuring direct use of the encoded weight for processing using a novel multiply-and-accumulate (MAC) unit design.
Muhammad Abdullah Hanif, Giuseppe Maria Sarda, Alberto Marchisio, Guido Masera, Maurizio Martina, Muhammad Shafique 0001
IJCNN5
2022 fakeWeather: Adversarial Attacks for Deep Neural Networks Emulating Weather Conditions on the Camera Lens of Autonomous Systems
abstract
Recently, Deep Neural Networks (DNNs) have achieved remarkable performances in many applications, while several studies have enhanced their vulnerabilities to malicious attacks. In this paper, we emulate the effects of natural weather conditions to introduce plausible perturbations that mislead the DNNs. By observing the effects of such atmospheric perturbations on the camera lenses, we model the patterns to create different masks that fake the effects of rain, snow, and hail. Even though the perturbations introduced by our attacks are visible, their presence remains unnoticed due to their association with natural events, which can be especially catastrophic for fully-autonomous and unmanned vehicles. We test our proposed fake Weather attacks on multiple Convolutional Neural Network and Capsule Network models, and report noticeable accuracy drops in the presence of such adversarial perturbations. Our work introduces a new security threat for DNNs, which is especially severe for safety-critical applications and autonomous systems.
Alberto Marchisio, Giovanni Caramia, Maurizio Martina, Muhammad Shafique 0001
IJCNN3
2022 LaneSNNs: Spiking Neural Networks for Lane Detection on the Loihi Neuromorphic Processor
abstract
Autonomous Driving (AD) related features represent important elements for the next generation of mobile robots and autonomous vehicles focused on increasingly intelligent, autonomous, and interconnected systems. The applications involving the use of these features must provide, by definition, real-time decisions, and this property is key to avoid catastrophic accidents. Moreover, all the decision processes must require low power consumption, to increase the lifetime and autonomy of battery-driven systems. These challenges can be addressed through efficient implementations of Spiking Neural Networks (SNNs) on Neuromorphic Chips and the use of event-based cameras instead of traditional frame-based cameras. In this paper, we present a new SNN-based approach, called LaneSNN, for detecting the lanes marked on the streets using the event-based camera input. We develop four novel SNN models characterized by low complexity and fast response, and train them using an offline supervised learning rule. Afterward, we implement and map the learned SNNs models onto the Intel Loihi Neuromorphic Research Chip. For the loss function, we develop a novel method based on the linear composition of Weighted binary Cross Entropy (WCE) and Mean Squared Error (MSE) measures. Our experimental results show a maximum Intersection over Union (IoU) measure of about 0.62 and very low power consumption of about 1 W. The best IoU is achieved with an SNN implementation that occupies only 36 neurocores on the Loihi processor while providing a low latency of less than 8 ms to recognize an image, thereby enabling real-time performance. The IoU measures provided by our networks are comparable with the state-of-the-art, but at a much low power consumption of 1 W.
Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IROS3
2022 Enabling Capsule Networks at the Edge through Approximate Softmax and Squash Operations
abstract
Complex Deep Neural Networks such as Capsule Networks (CapsNets) exhibit high learning capabilities at the cost of compute-intensive operations. To enable their deployment on edge devices, we propose to leverage approximate computing for designing approximate variants of the complex operations like softmax and squash. In our experiments, we evaluate tradeoffs between area, power consumption, and critical path delay of the designs implemented with the ASIC design flow, and the accuracy of the quantized CapsNets, compared to the exact functions.
Alberto Marchisio, Beatrice Bussolino, Edoardo Salvati, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
ISLPED4
2021 Very Low Latency Architecture for Earth Observation Satellite Onboard Data Handling, Compression, and Encryption
abstract
In modern society, the ever-increasing demand for Earth Observation products in a large variety of sectors is exposing the limitations of traditional satellite data chain architectures. The European Union Horizon 2020 EO-ALERT project aims at overcoming the existing bottlenecks by leveraging the performance of state-of-the-art commercial off-the-shelf devices to move the critical elements of data processing on the flight segment without sacrificing processing performance. This paper introduces the architecture of the EO-ALERT CPU Scheduling, Compression, Encryption and Data Handling Subsystem, responsible for coordinating the onboard optical and Synthetic Aperture Radar data chains, as well as providing data compression, encryption, and storage services. The performance obtained by a reference implementation of the proposed architecture is also presented, showing an extremely low contribution to the overall system latency that allows real-time Earth Observation product delivery to the end user in less than 5 min.
Michele Caon, Paolo Motto Ros, Maurizio Martina, Tiziano Bianchi, Enrico Magli, Francisco Membibre, Alexis Ramos, Antonio Latorre, Murray Kerr, Stefan Wiehle, Helko Breit, Dominik Günzel, Srikanth Mandapati, Ulrich Balss, Björn Tings
IGARSS3
2021 High-Level Synthesis of a Single/Multi-Band Optical and SAR Image Compression and Encryption Hardware Accelerator
abstract
Transmitting images from earth observation satellites to ground is a major challenge, and a compression/encryption stage is actually mandatory. Development of hardware accelerators is highly recommended, both to relieve the software from such demanding task, and to improve performance, aiming at quasi-real-time data processing. To this end, we discuss the design, development, deployment and test of a FPGA-based accelerator, featuring a lossless and lossy (near-lossless) compression, including the data encryption too. Its architecture is well suited for different image types, including single- and multi-band optical and SAR images and can be fully run-time configurable. Measured performance showed a throughput of 10 Msamples/s, in agreement with related state-of-the-art works, focused on lossless compression only.
Paolo Motto Ros, Michele Caon, Tiziano Bianchi, Maurizio Martina, Enrico Magli
IGARSS4
2021 DVS-Attacks: Adversarial Attacks on Dynamic Vision Sensors for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs), despite being energy-efficient when implemented on neuromorphic hardware and coupled with event-based Dynamic Vision Sensors (DVS), are vulnerable to security threats, such as adversarial attacks, i.e., small perturbations added to the input for inducing a misclassification. Toward this, we propose DVS-Attacks, a set of stealthy yet efficient adversarial attack methodologies targeted to perturb the event sequences that compose the input of the SNNs. First, we show that noise filters for DVS can be used as defense mechanisms against adversarial attacks. Afterwards, we implement several attacks and test them in the presence of two types of noise filters for DVS cameras. The experimental results show that the filters can only partially defend the SNNs against our proposed DVS-Attacks. Using the best settings for the noise filters, our proposed Mask Filter-Aware Dash Attack reduces the accuracy by more than 20% on the DVS-Gesture dataset and by more than 65% on the MNIST dataset, compared to the original clean frames. The source code of all the proposed DVS-Attacks and noise filters is released at https://github.com/albertomarchisio/DVS-Attacks.
Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IJCNN3
2021 CarSNN: An Efficient Spiking Neural Network for Event-Based Autonomous Cars on the Loihi Neuromorphic Research Processor
abstract
Autonomous Driving (AD) related features provide new forms of mobility that are also beneficial for other kind of intelligent and autonomous systems like robots, smart transportation, and smart industries. For these applications, the decisions need to be made fast and in real-time. Moreover, in the quest for electric mobility, this task must follow low power policy, without affecting much the autonomy of the mean of transport or the robot. These two challenges can be tackled using the emerging Spiking Neural Networks (SNNs). When deployed on a specialized neuromorphic hardware, SNNs can achieve high performance with low latency and low power consumption. In this paper, we use an SNN connected to an event-based camera for facing one of the key problems for AD, i.e., the classification between cars and other objects. To consume less power than traditional frame-based cameras, we use a Dynamic Vision Sensor (DVS) [1]. The experiments are made following an offline supervised learning rule, followed by mapping the learnt SNN model on the Intel Loihi Neuromorphic Research Chip [2]. Our best experiment achieves an accuracy on offline implementation of 86%, that drops to 83% when it is ported onto the Loihi Chip. The Neuromorphic Hardware implementation has maximum 0.72 ms of latency for every sample, and consumes only 310 mW. To the best of our knowledge, this work is the first implementation of an event-based car classifier on a Neuromorphic Chip.
Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IJCNN3
2021 R-SNN: An Analysis and Design Methodology for Robustifying Spiking Neural Networks against Adversarial Attacks through Noise Filters for Dynamic Vision Sensors
abstract
Spiking Neural Networks (SNNs) aim at providing energy-efficient learning capabilities when implemented on neuromorphic chips with event-based Dynamic Vision Sensors (DVS). This paper studies the robustness of SNNs against adversarial attacks on such DVS-based systems, and proposes R-SNN, a novel methodology for robustifying SNNs through efficient DVS-noise filtering. We are the first to generate adversarial attacks on DVS signals (i.e., frames of events in the spatio-temporal domain) and to apply noise filters for DVS sensors in the quest for defending against adversarial attacks. Our results show that the noise filters effectively prevent the SNNs from being fooled. The SNNs in our experiments provide more than 90% accuracy on the DVS-Gesture and NMNIST datasets under different adversarial threat models.
Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IROS3
2021 Analysis of in Vivo Plant Stem Impedance Variations in Relation with External Conditions Daily Cycle
abstract
World population growth and desertification are the most severe issue to agricultural food production. Smart agriculture is a promising solution to ensure food security. The use of sensors to monitor crop production can help farmers improve the yield and reduce water consumption. Here we propose a study where the electrical impedance of green plants' stem is analyzed in vivo, along with environmental conditions. In particular, the variations associated with the daily cycle are highlighted. These analyses lead to the possibility of understanding plant status directly from stem impedance.
Umberto Garlando, Lee Bar-on, Paolo Motto Ros, Alessandro Sanginario, Stefano Calvo, Maurizio Martina, Adi Avni, Yosi Shacham-Diamand, Danilo Demarchi
ISCAS6
2021 HW-FlowQ: A Multi-Abstraction Level HW-CNN Co-design Quantization Methodology
abstract
Model compression through quantization is commonly applied to convolutional neural networks (CNNs) deployed on compute and memory-constrained embedded platforms. Different layers of the CNN can have varying degrees of numerical precision for both weights and activations, resulting in a large search space. Together with the hardware (HW) design space, the challenge of finding the globally optimal HW-CNN combination for a given application becomes daunting. To this end, we propose HW-FlowQ, a systematic approach that enables the co-design of the target hardware platform and the compressed CNN model through quantization. The search space is viewed at three levels of abstraction, allowing for an iterative approach for narrowing down the solution space before reaching a high-fidelity CNN hardware modeling tool, capable of capturing the effects of mixed-precision quantization strategies on different hardware architectures (processing unit counts, memory levels, cost models, dataflows) and two types of computation engines (bit-parallel vectorized, bit-serial). To combine both worlds, a multi-objective non-dominated sorting genetic algorithm (NSGA-II) is leveraged to establish a Pareto-optimal set of quantization strategies for the target HW-metrics at each abstraction level. HW-FlowQ detects optima in a discrete search space and maximizes the task-related accuracy of the underlying CNN while minimizing hardware-related costs. The Pareto-front approach keeps the design space open to a range of non-dominated solutions before refining the design to a more detailed level of abstraction. With equivalent prediction accuracy, we improve the energy and latency by 20% and 45% respectively for ResNet56 compared to existing mixed-precision search methods.
Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Nguyen Anh Vu Doan, Christian Unger, Naveen Shankar Nagaraja, Maurizio Martina, Walter Stechele
ACM Trans. Embed. Comput. Syst.9
2020 NACU: A Non-Linear Arithmetic Unit for Neural Networks
abstract
Reconfigurable architectures targeting neural networks are an attractive option. They allow multiple neural networks of different types to be hosted on the same hardware, in parallel or sequence. Reconfigurability also grants the ability to morph into different micro-architectures to meet varying power-performance constraints. In this context, the need for a reconfigurable non-linear computational unit has not been widely researched. In this work, we present a formal and comprehensive method to select the optimal fixed-point representation to achieve the highest accuracy against the floating-point implementation benchmark. We also present a novel design of an optimised reconfigurable arithmetic unit for calculating non-linear functions. The unit can be dynamically configured to calculate the sigmoid, hyperbolic tangent, and exponential function using the same underlying hardware. We compare our work with the state-of-the-art and show that our unit can calculate all three functions without loss of accuracy.
Guido Baccelli, Dimitrios Stathis 0001, Ahmed Hemani, Maurizio Martina
DAC4
2020 Q-CapsNets: A Specialized Framework for Quantizing Capsule Networks
abstract
Capsule Networks (CapsNets), recently proposed by the Google Brain team, have superior learning capabilities in machine learning tasks, like image classification, compared to the traditional CNNs. However, CapsNets require extremely intense computations and are difficult to be deployed in their original form at the resource-constrained edge devices. This paper makes the first attempt to quantize CapsNet models, to enable their efficient edge implementations, by developing a specialized quantization framework for CapsNets. We evaluate our framework for several benchmarks. On a deep CapsNet model for the CIFAR10 dataset, the framework reduces the memory footprint by 6.2x, with only 0.15% accuracy loss. We will open-source our framework at https://git.io/JvDIF.
Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
DAC4
2020 NASCaps: A Framework for Neural Architecture Search to Optimize the Accuracy and Hardware Efficiency of Convolutional Capsule Networks
abstract
Deep Neural Networks (DNNs) have made significant improvements to reach the desired accuracy to be employed in a wide variety of Machine Learning (ML) applications. Recently the Google Brain's team demonstrated the ability of Capsule Networks (CapsNets) to encode and learn spatial correlations between different input features, thereby obtaining superior learning capabilities compared to traditional (i.e., non-capsule based) DNNs. However, designing CapsNets using conventional methods is a tedious job and incurs significant training effort. Recent studies have shown that powerful methods to automatically select the best/optimal DNN model configuration for a given set of applications and a training dataset are based on the Neural Architecture Search (NAS) algorithms. Moreover, due to their extreme computational and memory requirements, DNNs are employed using the specialized hardware accelerators in IoT-Edge/CPS devices.
Alberto Marchisio, Andrea Massa, Vojtech Mrazek, Beatrice Bussolino, Maurizio Martina, Muhammad Shafique 0001
ICCAD5
2020 FasTrCaps: An Integrated Framework for Fast yet Accurate Training of Capsule Networks
abstract
Recently, Capsule Networks (CapsNets) have shown improved performance compared to the traditional Convolutional Neural Networks (CNNs), by encoding and preserving spatial relationships between the detected features in a better way. This is achieved through the so-called Capsules (i.e., groups of neurons) that encode both the instantiation probability and the spatial information. However, one of the major hurdles in the wide adoption of CapsNets is their gigantic training time, which is primarily due to the relatively higher complexity of their new constituting elements that are different from CNNs.In this paper, we implement different optimizations in the training loop of the CapsNets, and investigate how these optimizations affect their training speed and the accuracy. Towards this, we propose a novel framework FasTrCaps that integrates multiple lightweight optimizations and a novel learning rate policy called WarmAdaBatch (that jointly performs warm restarts and adaptive batch size), and steers them in an appropriate way to provide high training-loop speedup at minimal accuracy loss. We also propose weight sharing for capsule layers. The goal is to reduce the hardware requirements of CapsNets by removing unused/redundant connections and capsules, while keeping high accuracy through tests of different learning rate policies and batch sizes. We demonstrate that one of the solutions generated by the FasTrCaps framework can achieve 58.6% reduction in the training time, while preserving the accuracy (even 0.12% accuracy improvement for the MNIST dataset), compared to the CapsNet by Google Brain [25]. Moreover, the Pareto-optimal solutions generated by FasTrCaps can be leveraged to realize trade-offs between training time and achieved accuracy. We have open-sourced our framework on GitHub1.
Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Muhammad Abdullah Hanif, Maurizio Martina, Guido Masera, Muhammad Shafique 0001
IJCNN5
2020 Is Spiking Secure? A Comparative Study on the Security Vulnerabilities of Spiking and Deep Neural Networks
abstract
Spiking Neural Networks (SNNs) claim to present many advantages in terms of biological plausibility and energy efficiency compared to standard Deep Neural Networks (DNNs). Recent works have shown that DNNs are vulnerable to adversarial attacks, i.e., small perturbations added to the input data can lead to targeted or random misclassifications. In this paper, we aim at investigating the key research question: "Are SNNs secure?" Towards this, we perform a comparative study of the security vulnerabilities in SNNs and DNNs w.r.t. the adversarial noise. Afterwards, we propose a novel black-box attack methodology, i.e., without the knowledge of the internal structure of the SNN, which employs a greedy heuristic to automatically generate imperceptible and robust adversarial examples (i.e., attack images) for the given SNN. We perform an in-depth evaluation for a Spiking Deep Belief Network (SDBN) and a DNN having the same number of layers and neurons (to obtain a fair comparison), in order to study the efficiency of our methodology and to understand the differences between SNNs and DNNs w.r.t. the adversarial examples. Our work opens new avenues of research towards the robustness of the SNNs, considering their similarities to the human brain's functionality.
Alberto Marchisio, Giorgio Nanfa, Faiq Khalid, Muhammad Abdullah Hanif, Maurizio Martina, Muhammad Shafique 0001
IJCNN5
2020 An Efficient Spiking Neural Network for Recognizing Gestures with a DVS Camera on the Loihi Neuromorphic Processor
abstract
Spiking Neural Networks (SNNs), the third generation NNs, have come under the spotlight for machine learning based applications due to their biological plausibility and reduced complexity compared to traditional artificial Deep Neural Networks (DNNs). These SNNs can be implemented with extreme energy efficiency on neuromorphic processors like the Intel Loihi research chip, and fed by event-based sensors, such as DVS cameras. However, DNNs with many layers can achieve relatively high accuracy on image classification and recognition tasks, as the research on learning rules for SNNs for real-world applications is still not mature. The accuracy results for SNNs are typically obtained either by converting the trained DNNs into SNNs, or by directly designing and training SNNs in the spiking domain. Towards the conversion from a DNN to an SNN, we perform a comprehensive analysis of such process, specifically designed for Intel Loihi, showing our methodology for the design of an SNN that achieves nearly the same accuracy results as its corresponding DNN. Towards the usage of the event-based sensors, we design a pre-processing method, evaluated for the DvsGesture dataset, which makes it possible to be used in the DNN domain. Hence, based on the outcome of the first analysis, we train a DNN for the pre-processed DvsGesture dataset, and convert it into the spike domain for its deployment on Intel Loihi, which enables real-time gesture recognition. The results show that our SNN achieves 89.64% classification accuracy and occupies only 37 Loihi cores.
Riccardo Massa, Alberto Marchisio, Maurizio Martina, Muhammad Shafique 0001
IJCNN3
2020 NeuroAttack: Undermining Spiking Neural Networks Security through Externally Triggered Bit-Flips
abstract
Due to their proven efficiency, machine-learning systems are deployed in a wide range of complex real-life problems. More specifically, Spiking Neural Networks (SNNs) emerged as a promising solution to the accuracy, resource-utilization, and energy-efficiency challenges in machine-learning systems. While these systems are going mainstream, they have inherent security and reliability issues. In this paper, we propose NeuroAttack, a cross-layer attack that threatens the SNNs integrity by exploiting low-level reliability issues through a high-level attack. Particularly, we trigger a fault-injection based sneaky hardware backdoor through a carefully crafted adversarial input noise. Our results on Deep Neural Networks (DNNs) and SNNs show a serious integrity threat to state-of-the art machine-learning techniques.
Valerio Venceslai, Alberto Marchisio, Ihsen Alouani, Maurizio Martina, Muhammad Shafique 0001
IJCNN4
2020 Towards Optimal Green Plant Irrigation: Watering and Body Electrical Impedance
abstract
With the growth of world population and food demand, it is crucial to optimize water consumption for agriculture cultivation. Here we propose a method to monitor plant status, relating the measured parameters to the watering or drying situation of a single plant. Plant trunk electrical impedance measurements and environmental parameters were analyzed with a statistical approach. Correlation and causality among the data are showed and analyzed. In this way, it was possible to easily obtain the needed information about plant status.
Umberto Garlando, Lee Bar-on, Paolo Motto Ros, Alessandro Sanginario, Sebastian Peradotto, Yosi Shacham-Diamand, Adi Avni, Maurizio Martina, Danilo Demarchi
ISCAS8
2020 An Area-Efficient Variable-Size Fixed-Point DCT Architecture for HEVC Encoding
abstract
This paper proposes an area-efficient fixed-point architecture for the computation of the discrete cosine transform (DCT) of multiple sizes in high efficiency video coding (HEVC). This result is obtained by comparing different DCT factorizations in order to find the most suitable one for implementation in the HEVC encoder. The recursive structure of fast algorithms, which decompose the N-point DCT by means of two N/2-point DCTs, is exploited to execute computations of small-size DCTs in parallel, thus maximizing the hardware reusability while maintaining a constant throughput. The simulation results prove that the proposed solution features reduced rate-distortion losses, with relevant complexity saving compared with the state-of-the-art implementations. Finally, the proposed architecture is exploited to design two families of architectures for the 2D-DCT, namely, folded and full-parallel.
Maurizio Masera, Guido Masera, Maurizio Martina
IEEE Trans. Circuits Syst. Video Technol.3
2019 A Novel Framework for Designing Directional Linear Transforms with Application to Video Compression
abstract
Transforms incorporating directional information are appealing in a wide range of applications. In this paper, we introduce a new framework that allows to define a directional transform starting from any two-dimensional separable transform. The proposed method is highly general and it can be of interest in many areas of signal processing. We also show an example of possible application. We define a directional integer DCT and DST and we show their application in video compression by integrating them in the HEVC video coding standard.
Maurizio Masera, Giulia Fracastoro, Maurizio Martina, Enrico Magli
ICASSP3
2019 Live Demonstration: Event-Driven Serial Communication on Optical Fiber
abstract
This demonstration presents the first implementation of “event-driven” serial asynchronous communication on optical fiber. “event-driven” communication is used by neuromorphic sensors, that sample the sensory signal when the signal itself changes of a given amount. This type of sensing adapts to the dynamics of the input itself, achieving at the same time extremely high temporal resolution (when needed), low latency and signal compression. To apply this technology in robotics, optical communication will greatly improve the resilience to electric disturbances.
Andrea De Marcellis, Guido Di Patrizio Stanchieri, Marco Faccio, Elia Palange, Paolo Motto Ros, Maurizio Martina, Danilo Demarchi, Chiara Bartolozzi
ISCAS6
2018 PruNet: Class-Blind Pruning Method For Deep Neural Networks
abstract
DNNs are highly memory and computationally intensive, due to which they are unfeasible to deploy in real time or mobile applications, where power and memory resources are scarce. Introducing sparsity in the network is a way to reduce those requirements. However, systematically employing pruning under given accuracy requirements is a challenging problem. We propose a novel methodology that iteratively applies a magnitude-based Class-Blind pruning to compress a DNN for obtaining a sparse model. It is a generic methodology and can be applied to different types of DNNs. We demonstrate that retraining after pruning is essential to restore the accuracy of the network. Experimental results show that our methodology is able to reduce the model size by around two orders of magnitude, without noticeably affecting the accuracy. It requires several iterations of pruning and retraining, but can achieve up to 190x Memory Saving Ratio (for the LeNet on the MNIST dataset) when compared to the baseline model. Similar results are also obtained for more complex networks like 91x for VGG-16 on the CIFAR100 dataset. If we combine this work with an efficient coding for sparse networks, like Compressed Sparse Column (CSC) or Compressed Sparse Row (CSR), we can obtain a reduced memory footprint. Our methodology can be complemented by other compression techniques, like weight sharing, quantization or fixed-point conversion, that allows to further reduce memory and computations.
Alberto Marchisio, Muhammad Abdullah Hanif, Maurizio Martina, Muhammad Shafique 0001
IJCNN3
2018 Live Demonstration: Tactile Events from Off-The-Shelf Sensors in a Robotic Skin
abstract
The demonstration presents a robotic event-based tactile infrastructure for a humanoid robot. It leverages on currently deployed sample-based capacitive sensors to generate tactile events, enabling the investigation and development of event-driven tactile applications, and minimizing communication bandwidth and latency. The modular FPGA-based system samples data from tactile sensors and generates address-events, transmitted through an asynchronous serial address-event representation protocol. To enable performance comparisons of the event-driven approach with respect to standard sample-based solutions, the acquisition modules can directly forward the input samples through the same event-based communication channel. We will show in real time a comparison between the tactile events and the original sampled data generated when the skin patch is touched.
Chiara Bartolozzi, Paolo Motto Ros, Riccardo Peloso, Francesco Diotalevi, Marco Crepaldi, Maurizio Martina, Danilo Demarchi
ISCAS6
2017 Analysis of HEVC transform throughput requirements for hardware implementations
Maurizio Masera, Lorenzo Re Fiorentin, Enrico Masala, Guido Masera, Maurizio Martina
Signal Process. Image Commun.5
2017 Adaptive Approximated DCT Architectures for HEVC
abstract
This paper proposes a flexible and efficient implementation of the 2D N-point discrete cosine transform (DCT) for the High Efficiency Video Coding (HEVC) standard. The DCT is implemented through the Walsh-Hadamard transform (WHT) followed by Givens rotations. This scheme is exploited to derive an adaptive algorithm, which allows computing of four different approximations ranging from the complete DCT to the WHT, by selectively skipping some rotations. This paper shows the statistical analysis of the DCT usage and derives a precomputation mechanism to adaptively skip rotations. Each approximation, referred to as a operating mode, is characterized by a large saving of operations, at the expense of very small quality loss. Then, two 2D-DCT architectures are proposed: the first one is totally unfolded, while the second one is folded. The two designs are finally synthesized with a 90-nm standard-cell library for a clock frequency of 250 MHz. Both architectures support real-time processing of 8K UHD video sequences at 64 and 26 fps, respectively, and show higher throughput and lower gate count compared with the state-of-art implementations. Moreover, power saving ranging from 28% to 56% can be achieved by working within the proposed operating modes.
Maurizio Masera, Maurizio Martina, Guido Masera
IEEE Trans. Circuits Syst. Video Technol.2
2015 An all-digital spike-based ultra-low-power IR-UWB dynamic average threshold crossing scheme for muscle force wireless transmission
Masoud Shahshahani Amirhossein, Paolo Motto Ros, Alberto Bonanno, Marco Crepaldi, Maurizio Martina, Danilo Demarchi, Guido Masera
DATE5
2015 Using Information Centric Networking for Mobile Devices Cooperation at the Network Edge
abstract
This paper presents an Android application exploiting a collaborative multi-RAT approach (cellular and Wi-Fi Radio Access Technologies) for adaptively streaming video content. The application has been developed and tested on real Android devices of different vendors. Each device coordinates with the others to download portions of a video stream through its own preferred RAT, and shares the results with the neighboring devices through a proximity, high bit rate network based on Wi-Fi Direct. As a result, the aggregate bandwidth available to the group of collaborating devices allows a faster download and a higher quality of video playback streaming, with respect to what is achievable by a single device. The application is designed to exploit common Android devices along with the key functionalities of an Information Centric Network: routing-by-name, in-network caching and multicasting.
Fabio Malabocchia, Romeo Corgiolu, Maurizio Martina, Andrea Detti, Bruno Ricci, Nicola Blefari-Melazzi
VTC Spring3
2015 Exploiting generalized de-Bruijn/Kautz topologies for flexible iterative channel code decoder architectures
Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera
Integr.2
2015 Parallel H.264/AVC Fast Rate-Distortion Optimized Motion Estimation by Using a Graphics Processing Unit and Dedicated Hardware
abstract
Heterogeneous systems on a single chip composed of a central processing unit, graphics processing unit (GPU), and field-programmable gate array (FPGA) are expected to emerge in the near future. In this context, the system on chip can be dynamically adapted to employ different architectures for execution of data-intensive applications. Motion estimation (ME) is one such task that can be accelerated using FPGA and GPU for high-performance H.264/Advanced Video Coding encoder implementation. This paper presents an inherent parallel low-complexity rate-distortion (RD) optimized fast ME algorithm well suited for parallel implementations, eliminating various data dependencies caused by a reliance on spatial predictions. In addition, this paper provides details of the GPU and FPGA implementations of the parallel algorithm by using OpenCL and Very High Speed Integrated Circuits (VHSIC) Hardware Descriptive Language (VHDL), respectively, and presents a practical performance comparison between the two implementations. The experimental results show that the proposed scheme achieves significant speedup on GPU and FPGA, and has comparable RD performance with respect to sequential fast ME algorithm.
Muhammad Usman Shahid, Ashfaq Ahmed, Maurizio Martina, Guido Masera, Enrico Magli
IEEE Trans. Circuits Syst. Video Technol.3
2014 Rediscovering Logarithmic Diameter Topologies for Low Latency Network-on-Chip-Based Applications
abstract
Low-latency Network-on-Chip (NoC) applications have tight constraints on the clock budget to perform communication among nodes. This is a critical aspect in NoC-based designs where the number of clock cycles spent for communication depends mainly on the topology and on the routing algorithm. This work deals with logarithmic diameter topologies, that were proposed for computer networks, and shows that an optimal shortest-path routing algorithm can be efficiently implemented on this kind of topologies by means of a very simple circuit. The proposed circuit is then exploited to reduce the area and the power consumption of a recently proposed NoC-based design. Experimental results show that the proposed circuit allows for a reduction of about 14% and 10% for area and power consumption respectively, with respect to a shortest-path routing-table-based design.
Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera
PDP2
2014 Variable Parallelism Cyclic Redundancy Check Circuit for 3GPP-LTE/LTE-Advanced
abstract
Cyclic Redundancy Check (CRC) is often employed in data storage and communications to detect errors. The 3GPP-LTE wireless communication standard uses a 24-bit CRC with every turbo coded frame, thus, the CRC can be exploited to detect residual errors and to enable early stopping of iterations as well. The current state of the art lacks specific CRC implementations for this standard, and most current solutions adopt a fixed degree of parallelism, unsuitable for many turbo decoder architectures. This work proposes a variable parallelism circuit targeting the 3GPP-LTE/LTE-Advanced 24-bit CRC, that can adapt to input data of different sizes. Low complexity is achieved through careful functional sharing among the various parallelisms: comparison with the state of the art shows comparable or superior speed and extremely low complexity.
Carlo Condo, Maurizio Martina, Gianluca Piccinini, Guido Masera
IEEE Signal Process. Lett.2
2013 VLSI Architecture for Low-Complexity Motion Estimation in H.264 Multiview Video Coding
abstract
This paper presents a VLSI architecture for a low complexity motion estimation algorithm, referred to as Slim264, for multiview video coding extension of H.264. Algorithmic modifications are introduced to obtain a fully parallel computational structure able to meet the throughput requirements of high resolution and high frame rate videos. High parallelism is achieved by predicting small blocks, i.e. 4x4 pixel blocks, in parallel and then adding them up in order to get Sum of Absolute Differences (SADs) of large block sizes. The predictor is able to support high resolution videos i.e. 1080p. The modified algorithm shows promising PSNR results with respect to full search algorithm. The predictor is synthesized with a clock frequency of 200 MHz, occupying an area of 0.49 mm2, on 90-nm Standard Cell ASIC technology.
Ashfaq Ahmed, Muhammad Usman Shahid, Maurizio Martina, Enrico Magli, Guido Masera
DSD3
2012 A Network-on-Chip-based turbo/LDPC decoder architecture
abstract
The current convergence process in wireless technologies demands for strong efforts in the conceiving of highly flexible and interoperable equipments. This contribution focuses on one of the most important baseband processing units in wireless receivers, the forward error correction unit, and proposes a Network-on-Chip (NoC) based approach to the design of multi-standard decoders. High level modeling is exploited to drive the NoC optimization for a given set of both turbo and Low-Density-Parity-Check (LDPC) codes to be supported. Moreover, synthesis results prove that the proposed approach can offer a fully compliant WiMAX decoder, supporting the whole set of turbo and LDPC codes with higher throughput and an occupied area comparable or lower than previously reported flexible implementations. In particular, the mentioned design case achieves a worst-case throughput higher than 70 Mb/s at the area cost of 3.17 mm2on a 90 nm CMOS technology.
Carlo Condo, Maurizio Martina, Guido Masera
DATE2
2012 Non-recursive max* operator with reduced implementation complexity for turbo decoding
abstract
In this study, the authors deal with the problem of how to effectively approximate the max* operator when having n>2 input values, with the aim of reducing implementation complexity of conventional Log-MAP turbo decoders. They show that, contrary to previous approaches, it is not necessary to apply the max* operator recursively over pairs of values. Instead, a simple, yet effective, solution for the max* operator is revealed having the advantage of being in non-recursive form and thus, requiring less computational effort. Hardware synthesis results for practical turbo decoders have shown implementation savings for the proposed method against the most recent published efficient turbo decoding algorithms by providing near optimal bit error rate (BER) performance.
Stylianos Papaharalabos, P. Takis Mathiopoulos, Guido Masera, Maurizio Martina
IET Commun.4
2012 High Speed Architectures for Finding the First two Maximum/Minimum Values
abstract
High speed architectures for finding the first two maximum/minimum values are of paramount importance in several applications, including iterative (e.g., turbo and low-density-parity-check) decoders. In this brief, stemming from a previous work, based on radix-2 solutions, we propose higher and mixed radix implementations that improve the architecture latency. Post place and route results on a 180-nm CMOS standard cell technology show that the proposed architectures achieve lower latency than radix-2 solutions with a moderate area increase.
Luca G. Amarù, Maurizio Martina, Guido Masera
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Efficient Implementation Techniques for Maximum Likelihood-Based Error Correction for JPEG2000
abstract
Features of the JPEG2000 compression standard include coding efficiency at low bit rates. However, its compressed bit stream is sensitive to transmission error. This paper presents three techniques to reduce both the computational complexity and the memory requirement in the ternary MQ arithmetic decoding. Such coders introduce a controlled degree of redundancy during the encoding process, which can be exploited at the decoder side in order to detect and correct errors. Our proposed techniques result in a substantial saving of decoding time and memory usage, with no or little degradation in the PSNR metric.
Simone Zezza, Saeid Nooshabadi, Maurizio Martina, Guido Masera
IEEE Trans. Circuits Syst. Video Technol.3
2008 Error resilient JPEG2000 decoding for wireless applications
abstract
To improve the JPEG2000 compression standard error resiliency in the wireless environment, the use of ternary MQ arithmetic coders/decoders that are based on the concept of forbidden symbol has been proposed. This paper presents two ternary MQ based techniques to reduce both the computational complexity and the memory requirement during the decoding process, with no or little degradation in the PSNR.
Simone Zezza, Maurizio Martina, Guido Masera, Saeid Nooshabadi
ICIP2
2007 Flexible blocks for high throughput serially concatenated convolutional codes
abstract
Modern communication systems aim to integrate different services including telephony, multimedia applications and data transfer. This integration imposes to design flexible systems, able to support different frame duration and data-rates. In this paper we present several blocks employed in serially concatenated convolutional codes and we detail some hardware solutions to achieve flexibility.
Maurizio Martina, Guido Masera
ACM Great Lakes Symposium on VLSI1
2007 Real-time implementation of a time-frequency analysis scheme
abstract
The on-line monitoring and detection of defects in laser welding is a basic manifacturing requirement in several applicative contexts, including vehicle assembly in automotive production. The speed of assembly in modern industry and the large amount of data to be acquired and elaborated pose severe real-time constraints and lead to the need of extremely short processing latencies. In this work, the implementation of time-frequency analysis algorithms on an FPGA device is shown and compared to pure software developments on different processors.
Maurizio Martina, Andrea Terreno, Fabrizio Vacca, Andrea Molino, Guido Masera, Giuseppe D'Angelo, Giorgio Pasquettaz
ACM Great Lakes Symposium on VLSI1
2006 Mumford and Shah Functional: VLSI Analysis and Implementation
abstract
This paper describes the analysis of the Mumford and Shah functional from the implementation point of view. Our goal is to show results in terms of complexity for real-time applications, such as motion estimation based on segmentation techniques, of the Mumford and Shah functional. Moreover, the sensitivity to finite precision representation is addressed, a fast VLSI architecture is described, and results obtained for its complete implementation on a 0.13 microm standard cells technology are presented.
Maurizio Martina, Guido Masera
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Low-complexity, efficient 9/7 wavelet filters implementation
abstract
This paper proposes a novel low-complexity, efficient 9/7 wavelet filters implementation for image compression applications. The 9/7 wavelet filters are widely used in different image compression schemes, such as the JPEG2000 image coding standard. Thus the implementation of efficient codecs is of great concern. The performance of a hardware implementation of the 9/7 filter bank depends on the accuracy with which filter coefficients are represented. However the greatest part of current implementations consider filters taps as numbers to be implemented. The aim of this work is to show that great complexity reduction can be achieved going through the derivation of the 9/7 taps values.
Maurizio Martina, Guido Masera
ICIP (3)1
2003 A reconfigurable, power-scalable rake receiver IP for W-CDMA
abstract
During the last years wireless market has experienced an exponential growth. 2G systems are essentially voice-oriented: the main innovation expected from 3G ones is the ubiquitous Internet and multimedia fruition. The transition from 2G to 3G provides both opportunities and challenges: one way to make this migration as smoother as possible relies on the employment of reconfigurable architectures. In this paper a reconfigurable Rake Receiver for W-CDMA is proposed. Very promising results from the physical implementation on a XCV300E have been obtained.
A. Bianco, Alberto Dassatti, Maurizio Martina, Andrea Molino, Fabrizio Vacca
ASP-DAC3
2003 Wireless sensor networks: a power-scalable motion estimation IP for hybrid video coding
abstract
Wireless Sensor Networks are an emerging phenomenon in the research community. The design and development of network architectures and nodes implementation are fostering many research activities. Due to their wide application fields and pervasive employment possibilities, the investigation of novel classes of wireless sensor nodes is of great concern. In this paper we presented a novel Power-Scalable Motion Estimation IP suitable for video-surveillance over Wireless Sensor Networks. The proposed architecture can achieve low dynamic power, good video quality and reasonable frame-rates. Moreover, it can operate on different video format and can be reconfigured both off-line and on-line. Further researches have to be accomplished in order to reduce the internal critical path, achieving better maximum frequencies performances. Moreover, new FPGA architectures must be exploited searching low-static-power devices suitable for an actual implementation of our IP on a WSN.
Federico Quaglio, Maurizio Martina, Fabrizio Vacca, Guido Masera, Andrea Molino, Gianluca Piccinini, Maurizio Zamboni
FPGA2
2003 Design of a Power Conscious, Customizable CDMA Receiver
Maurizio Martina, Andrea Molino, Mario Nicola, Fabrizio Vacca
FPL1
2003 A Power-Scalable Motion Estimation Architecture for Energy Constrained Applications
Maurizio Martina, Andrea Molino, Federico Quaglio, Fabrizio Vacca
FPL1
2003 Implementation of a SPIHT coprocessor: memory issues and hardware implications
abstract
In the last years mobile terminals diffusion has fostered the development of new audio-visual services. Thus the possibility to employ wireless devices to exchange multimedia informations is becoming real. Beyond these attractive motivations, some critical factors are hidden: the design of wireless terminals, in fact, has to deal with limited power budgets and reduced resources availability. This paper proposes a hardware implementation of a SPIHT coprocessor, designed for mobile environments. Particular emphasis will be posed on memory requirements and their hardware implications.
Maurizio Martina, Andrea Molino, Andrea Terreno, Fabrizio Vacca
ICIP (2)1
2003 Dynamic power scheduling system for JPEG2000 delivery over wireless networks
Maurizio Martina, Fabrizio Vacca
VCIP1
2002 Energy Evaluation on a Reconfigurable, Multimedia-Oriented Wireless Sensor
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni
FPL1
2002 Reconfigurable DSP IP for multimedia applications
abstract
In this paper a novel Digital Signal Processor IP for multimedia applications, is presented. Recently, develeper's interest towards SOC architectures has been driven by mobile market explosion. Despite the increasing importance gathered by reconfigurable computing, a lack of easily retargettable cores is felt by developer's community. This IP is intended to be the kernel for many telecommunication and multimedia algorithms computation. During the design flow, much care has been devoted to grant maximum interoperability among this DSP core and other coprocessor units, allowing to easily embed multiple functional blocks on a single FPGA. As far as performance are concerned, the proposed IP shows satisfactory results both in terms of area occupation (11\% on a XILINX XCV1000) and maximum clock frequency (89 MHz after place and route process).
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni
ICASSP1
2002 Reconfigurable and low power 2D-DCT IP for ubiquitous multimedia streaming
abstract
An energy efficient architecture for the discrete cosine transform is presented. The proposed IP is intended to be used as the transform stage in a mobile H.263 codec. In particular, it seems well suited for an FPGA implementation since, after a complete place and route process, it is able to sustain a full-motion PAL video streaming operating at a frequency of 74 MHz with a dynamic power dissipation of just 39 mW.
Maurizio Martina, Andrea Molino, Fabrizio Vacca
ICME (2)1
2002 Optimization and implementation of the integer wavelet transform for image coding
abstract
This paper deals with the design and implementation of an image transform coding algorithm based on the integer wavelet transform (IWT). First of all, criteria are proposed for the selection of optimal factorizations of the wavelet filter polyphase matrix to be employed within the lifting scheme. The obtained results lead to the IWT implementations with very satisfactory lossless and lossy compression performance. Then, the effects of finite precision representation of the lifting coefficients on the compression performance are analyzed, showing that, in most cases, a very small number of bits can be employed for the mantissa keeping the performance degradation very limited. Stemming from these results, a VLSI architecture is proposed for the IWT implementation, capable of achieving very high frame rates with moderate gate complexity.
Marco Grangetto, Enrico Magli, Maurizio Martina, Gabriella Olmo
IEEE Trans. Image Process.3