VLDB 2026 Research / reviewers in the wild / expert
Maurizio Martina
dblp:26/3014
· DBLP profile ↗
68ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0002-3069-0319ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIRCE CROSS Integrated RISC-V Cryptographic Extension
Alessandra Dolmeta, Valeria Piscopo, Maurizio Martina, Guido Masera |
DATE | 3 |
| 2026 | Compact Yet Fast: An Efficient d-Order Masked Implementation of AsconabstractIn this work, we present a generic side-channel protected design of Ascon that achieves high efficiency by dynamically reconfiguring the hardware countermeasures during message processing. The resultant implementation is protected and capable of meeting stringent performance requirements whilst minimising resource overhead. The experimental results obtained demonstrate that the implementation meets the required security and achieves superior throughput-to-area ratio across all protection orders. Ascon, recently selected by NIST as the lightweight cryptography standard, is widely deployed in resource-constrained devices that demand both high performance and resistance against threats such as side-channel analysis (SCA). Exploiting Ascon’s mode-level structure, which does not require protection against differential power analysis during bulk operations, we introduce a modified masking gadget with dual functionality: serving as a countermeasure during critical operations, and processing multiple data paths in parallel to accelerate bulk computation. Our architecture supports any configurable security order and instantiates only the minimum hardware resources needed to maximize throughput per round. We also evaluate an enhanced Ascon architecture based on the Changing of the Guards technique, which eliminates the need for fresh randomness. Security validation is performed using fixed-vs-random t-tests on both first- and second-order masked implementations. Finally, we compare our masked design against state-of-the-art solutions. Mattia Mirigaldi, Nico Paninforni, Maurizio Martina, Guido Masera |
DATE | 3 |
| 2025 | ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsabstractModern data-driven applications expose limitations of von Neumann architectures-extensive data movement, low throughput, and poor energy efficiency. Accelerators improve performance but lack flexibility and require data transfers. Existing compute in- and nearmemory solutions mitigate these issues but face usability challenges due to data placement constraints. We propose a novel cache architecture that doubles as a tightly-coupled compute-near-memory coprocessor. Our RISC-V cache controller executes custom instructions from the host CPU using vector operations dispatched to near-memory vector processing units within the cache memory subsystem. This architecture abstracts memory synchronization and data mapping from application software while offering software-based Instruction Set Architecture extensibility. Our implementation shows $30 \times$ to $84 \times$ performance improvement when operating on 8-bit data over the same system with a traditional cache when executing a worst-case 32-bit CNN workload, with only 41.3% area overhead. Vincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone, Guido Masera, Maurizio Martina |
DAC | 6 |
| 2025 | TYRCA: A RISC-V Tightly-Coupled Accelerator for Code-Based CryptographyabstractPost-quantum cryptography (PQC) has garnered significant attention across various communities, particularly with the National Institute of Standards and Technology (NIST) advancing to the fourth round of PQC standardization. One of the leading candidates is Hamming Quasi-Cyclic (HQC), which received a significant update on February 23, 2024. This update, which introduces a classical dense-dense multiplication approach, has no known dedicated hardware implementations yet. The innovative Core-V eXtension InterFace (CV-X-IF) is a communication interface for RISC-V processors that significantly facilitates the integration of new instructions to the Instruction Set Architecture (ISA), through tightly connected accelerators. In this paper, we present a TightlY-coupled accelerator for RISC-V for Code-based cryptogrAphy (TYRCA), proposing the first fully tightly-coupled hardware implementation of the HQC-PQC algorithm, leveraging the CV-X-IF. The proposed architecture is implemented on the Xilinx Kintex-7 FPGA. Experimental results demonstrate that TYRCA reduces the execution time by 94% to 96% for HQC-128, HQC-192, and HQC-256, showcasing its potential for efficient HQC code-based cryptography. Alessandra Dolmeta, Stefano Di Matteo, Emanuele Valea, Mikael Carmona, Antoine Loiseau, Maurizio Martina, Guido Masera |
DATE | 6 |
| 2025 | TinyCL: An Efficient Hardware Architecture for Continual Learning on Autonomous SystemsabstractThe Continuous Learning (CL) paradigm consists of continuously evolving the parameters of the Deep Neural Network (DNN) model to progressively learn to perform new tasks without reducing the performance on previous tasks, i.e., avoiding the so-called catastrophic forgetting. However, the DNN parameter update in CL-based autonomous systems is extremely resource-hungry. The existing DNN accelerators cannot be directly employed in CL because they only support the execution of the forward propagation. Only a few prior architectures execute the backpropagation and weight update, but they lack the control and management for CL. Towards this, we design a hardware architecture, TinyCL, to perform CL on resource-constrained autonomous systems. It consists of a processing unit that executes both forward and backward propagation, and a control unit that manages memory-based CL workload. To minimize the memory accesses, the sliding window of the convolutional layer moves in a snake-like fashion. Moreover, the Multiply-and-Accumulate units can be reconfigured at runtime to execute different operations. As per our knowledge, our proposed TinyCL represents the first hardware accelerator that executes CL on autonomous systems. We synthesize the complete TinyCL architecture in a 65 nm CMOS technology node with the conventional ASIC design flow. It executes 1 epoch of training on a Conv + ReLU + Dense model on the CIFAR10 dataset in 1.76 s, while 1 training epoch of the same model using an Nvidia Tesla P100 GPU takes 103 s, thus achieving a 58× speedup, consuming 86 mW in a 4.74 mm2die. Eugenio Ressa, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2025 | RISC-V Based Keccak Co-Processor for NIST Post-Quantum Cryptography StandardsabstractThis paper presents the design and implementation of a RISC-V-based Keccak co-processor optimized for Post-Quantum Cryptography (PQC) algorithms. Leveraging the Core-V eXtension InterFace (CV-X-IF), the co-processor extends the Instruction Set Architecture (ISA) with three custom instructions tailored for cryptographic operations. This allows seamless integration into various PQC schemes, tested across the multiple standards proposed by the National Institute of Standards and Technology (NIST), including CRYSTALS-Kyber, CRYSTALS-Dilithium, SPHINCS+, and FALCON, which are designed to withstand quantum attacks. By employing tightly coupled hardware acceleration, the Keccak co-processor dramatically reduces the computational overhead of hash-based operations central to these algorithms. The implementation is realized on a Xilinx Artix 7 FPGA, achieving a clock cycles’ improvement up to 75% and 19% resource overhead. The results presented herein demonstrate significant performance enhancement over the state of the art, underscoring its effectiveness for cryptographic applications. Alessandra Dolmeta, Valeria Piscopo, Mattia Mirigaldi, Maurizio Martina, Guido Masera |
ISCAS | 4 |
| 2025 | Scalability analysis of multi-bank near-memory computing in low-power SoCsabstractMachine learning and artificial intelligence are moving towards the edge, where the need for high throughput with a constrained energy budget is more urgent than ever. During the last few years, near-memory computing has emerged as a promising solution to address the memory bandwidth and energy efficiency limitations of conventional von Neumann systems. The recently proposed NM-Carus architecture combines vectororiented computing capabilities within a RISC-V programmable, configurable, and autonomous memory macro, addressing the usability of near-memory computing from a software deployment standpoint. In this paper, we explore the scalability of NMCarus in terms of computation parallelism, memory size and energy consumption, As a benchmarking platform, we rely on a low-power microcontroller that features multiple instances of NM-Carus that target the execution of biomedical applications. This exploration was performed on 16 nm TSMC NM-Carus implmentation, and we highlighted the benefits of technology scaling for a previous implementation on 65 nm with respect to the overhead of replacing conventional on-chip data SRAMs with near-memory computing banks. Overall, the paper presents a solid baseline regarding the trade-offs in terms of area, performance, and energy efficiency of integrating programmable near-memory computing in an existing edge-oriented system on chip towards efficient edge AI architectures at the system level. Luigi Giuffrida, Pasquale Davide Schiavone, Michele Caon, Guido Masera, Maurizio Martina, David Atienza 0001 |
VLSI-SoC | 5 |
| 2024 | VirtLAB-UI: An Open, Platform Independent, Software and Firmware Solution for Remote and Take-Home LabsabstractSARS-CoV2 pandemic pushed university courses toward remote access. Besides classroom lessons, for which IT solutions were already present, laboratory experiences required the development of suitable replacements. For electronics engineering courses, several solutions have been developed, mainly based on computer simulations, or simplified lab kits distributed to students to perform laboratory experiences at home. But if the hardware tools developed to mimic lab experiments are intended to eventually replace classic experiences, a similar visual user experience must be implemented, as well, with user interfaces similar to what is available on real measurements instruments, independent from the physical devices used. In this paper, a suitable approach for the development of such an interface is described, with an example application to an existing take-home lab system. Massimo Ruo Roch, Maurizio Martina |
EDUCON | 2 |
| 2024 | A Case Study on Formal Equivalence Verification Between a C/C++ Model and Its RTL DesignabstractAbstract In the field of communication system products, most datapath Digital Signal Processing algorithms are initially developed at a high-level in MATLAB® or C/C++. Subsequently, design engineers use these models as a reference for implementing Register Transfer Level designs. The conventional approach to verify their equivalence involves extensive Universal Verification Methodology dynamic simulations, which can last for months and require significant verification efforts. However, some elusive errors might still occur because it is infeasible to explore all input combinations with this method. On the other hand, Formal Equivalence Verification aims to verify that a Register Transfer Level design is functionally equivalent to the reference high-level C/C++ model across all possible legal states. With recent advancements in formal solver technology, Formal Equivalence Verification provides a distinct benefit by using mathematical methods to ensure that the Register Transfer Level (timed) matches the original high-level C/C++ model (untimed). This drastically reduces the verification time and ensures the exhaustive coverage of the design state space. This paper presents an in-depth exploration of complex Finite State Machine with datapath verification, specifically focusing on Multiplier-Accumulator, Tone Generator, and Automatic Gain Control, by employing the formal equivalence methodology. Although these signal processing blocks were previously verified throughout Universal Verification Methodology dynamic simulations, Formal Equivalence Verification was able to identify hard-to-find bugs in just a few weeks by utilizing the new workflow, thereby streamlining the verification process. Gaetano Raia, Gianluca Rigano, David Vincenzoni, Maurizio Martina |
FM (2) | 4 |
| 2024 | MARLIN: A Co-Design Methodology for Approximate ReconfigurabLe Inference of Neural Networks at the EdgeabstractThe optimization of neural networks (NNs) is necessary to enable their deployment on energy-constrained devices. State-of-the-art methods leverage approximate multipliers to execute NNs reducing the inference energy without heavily affecting the accuracy. However, previous works usually require a specialized hardware accelerator and are limited to fixed multipliers or reconfigurable ones with few approximation levels. This paper introduces MARLIN, a framework to deploy layerwise approximate NNs on PULP, a microcontroller with a RISC-V core. A multiplier architecture, with runtime selection of 256 approximation levels, is developed and integrated into the PULP cluster cores, enabling runtime configuration through control status register (CSR) instructions embedded within the code. The PULP toolchain is adapted to incorporate the approximation level selection within the instruction flow seamlessly. MARLIN leverages the genetic algorithm NSGA-II to search for the best configurations among thousands of approximate NNs. The framework is validated by simulating an approximate NN trained with the MNIST dataset on PULP. Moreover, MARLIN is used to optimize and approximate six ResNet models trained with the CIFAR-10 dataset. In particular, for ResNet-56, the most complex NN used in the experiments, the multiplication energy is reduced by 23.9% while retaining 99% of the accuracy of the exact model. Flavia Guella, Emanuele Valpreda, Michele Caon, Guido Masera, Maurizio Martina |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Implementation and integration of Keccak accelerator on RISC-V for CRYSTALS-KyberabstractOne of the key metrics used for defying the security of the Internet of Things (IoT) is data integrity, which mostly relies on the use of cryptographic hash functions. In the last years, the National Institute of Standards and Technology (NIST) announced SHA-3 as the new standard for better security. SHA-3 is also exploited in most of the current post-quantum cryptographic (PQC) protocols. Nevertheless, the used algorithm, i.e. Keccak, is computationally heavy and consequently limits its utilization in RISC-V-based Systems on Chip (SoC). In this work, a Keccak accelerator is proposed to speed up SHA3 computations for the CRYSTALS-Kyber algorithm on the RISCV-based advanced microcontroller PULPissimo. Compared to the plain SW implementation on RISC-V, our results show a speedup factor of up to 2.79 at the expense of a 12.4% resources overhead. Alessandra Dolmeta, Mattia Mirigaldi, Maurizio Martina, Guido Masera |
CF | 3 |
| 2023 | SwiftTron: An Efficient Hardware Accelerator for Quantized TransformersabstractTransformers' compute- intensive operations pose enormous challenges for their deployment in resource- constrained EdgeAI / tiny ML devices. As an established neural network compression technique, quantization reduces the hardware computational and memory resources. In particular, fixed-point quantization is desirable to ease the computations using lightweight blocks, like adders and multipliers, of the underlying hardware. However, deploying fully-quantized Transformers on existing general-purpose hardware, generic AI accelerators, or specialized architectures for Transformers with floating-point units might be infeasible and/or inefficient. Towards this, we propose SwiftTron, an efficient specialized hardware accelerator designed for Quantized Transformers. SwiftTron supports the execution of different types of Transformers' operations (like Attention, Softmax, GELU, and Layer Normalization) and accounts for diverse scaling factors to perform correct computations. We synthesize the complete SwiftTron architecture in a 65 nm CMOS technology with the ASIC design flow. Our Accelerator executes the RoBERTa-base model in 1.83 ns, while consuming 33.64 mW power, and occupying an area of 273 mm2• To ease the reproducibility, the RTL of our SwiftTron architecture is released at https://github.com/albertomarchisio/SwiftTron. Alberto Marchisio, Davide Dura, Maurizio Capra, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2023 | RobCaps: Evaluating the Robustness of Capsule Networks against Affine Transformations and Adversarial AttacksabstractCapsule Networks (CapsNets) are able to hierarchically preserve the pose relationships between multiple objects for image classification tasks. Other than achieving high accuracy, another relevant factor in deploying CapsNets in safety-critical applications is the robustness against input transformations and malicious adversarial attacks. In this paper, we systematically analyze and evaluate different factors affecting the robustness of CapsN ets, compared to traditional Convolutional Neural Networks (CNNs). Towards a comprehensive comparison, we test two CapsNet models and two CNN models on the MNIST, GTSRB, and CIFAR10 datasets, as well as on the affine-transformed versions of such datasets. With a thorough analysis, we show which properties of these architectures better contribute to increasing the robustness and their limitations. Overall, CapsNets achieve better robustness against adversarial examples and affine transformations, compared to a traditional CNN with a similar number of parameters. Similar conclusions have been derived for deeper versions of CapsNets and CNNs. Moreover, our results unleash a key finding that the dynamic routing does not contribute much to improving the CapsNets' robustness. Indeed, the main generalization contribution is due to the hierarchical feature learning through capsules. Alberto Marchisio, Antonio De Marco, Alessio Colucci, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2022 | Mind the Scaling Factors: Resilience Analysis of Quantized Adversarially Robust CNNsabstractAs more deep learning algorithms enter safety-critical application domains, the importance of analyzing their resilience against hardware faults cannot be overstated. Most existing works focus on bit-flips in memory, fewer focus on compute errors, and almost none study the effect of hardware faults on adversarially trained convolutional neural networks (CNNs). In this work, we show that adversarially trained CNNs are more susceptible to failure due to hardware errors when compared to vanilla-trained models. We identify large differences in the quantization scaling factors of the CNNs which are resilient to hardware faults and those which are not. As adversarially trained CNNs learn robustness against input attack perturbations, their internal weight and activation distributions open a backdoor for injecting large magnitude hardware faults. We propose a simple weight decay remedy for adversarially trained models to maintain adversarial robustness and hardware resilience in the same CNN. We improve the fault resilience of an adversarially trained ResNet56 by 25% for large-scale bit-flip benchmarks on activation data while gaining slightly improved accuracy and adversarial robustness. Nael Fasfous, Lukas Frickenstein, Michael Neumeier, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Maurizio Martina, Walter Stechele |
DATE | 7 |
| 2022 | AnaCoNGA: Analytical HW-CNN Co-Design Using Nested Genetic AlgorithmsabstractWe present AnaCoNGA, an analytical co-design methodology, which enables two genetic algorithms to evaluate the fitness of design decisions on layer-wise quantization of a neural network and hardware (HW) resource allocation. We embed a hardware architecture search (HAS) algorithm into a quantization strategy search (QSS) algorithm to evaluate the hardware design Pareto-front of each considered quantization strategy. We harness the speed and flexibility of analytical HW-modeling to enable parallel HW-CNN co-design. With this approach, the QSS is focused on seeking high-accuracy quantization strategies which are guaranteed to have efficient hardware designs at the end of the search. Through AnaCoNGA, we improve the accuracy by 2.88 p.p. with respect to a uniform 2-bit ResNet20 on CIFAR-10, and achieve a 35% and 37% improvement in latency and DRAM accesses, while reducing LUT and BRAM resources by 9% and 59% respectively, when compared to a standard edge variant of the accelerator. The nested genetic algorithm formulation also reduces the search time by 51% compared to an equivalent, sequential co-design formulation. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Julian Höfer, Anmol Singh, Naveen Shankar Nagaraja, Hans-Jörg Vögel, Nguyen Anh Vu Doan, Maurizio Martina, Jürgen Becker 0001, Walter Stechele |
DATE | 11 |
| 2022 | NLCMAP: A Framework for the Efficient Mapping of Non-Linear Convolutional Neural Networks on FPGA AcceleratorsabstractThis paper introduces NLCMap, a framework for the mapping space exploration targeting Non-Linear Convolutional Networks (NLCNs). NLCNs [1] are a novel neural network model that improves performances in certain computer vision applications by introducing a non-linearity in the weights computation. NLCNs are more challenging to efficiently map onto hardware accelerators if compared to traditional Convolutional Neural Networks (CNNs), due to data dependencies and additional computations. To this aim, we propose NL-CMap, a framework that, given an NLC layer and a generic hardware accelerator with a certain on-chip memory budget, finds the optimal mapping that minimizes the accesses to the off-chip memory, which are often the critical aspect in CNNs acceleration. Giuseppe Aiello, Beatrice Bussolino, Emanuele Valpreda, Massimo Ruo Roch, Guido Masera, Maurizio Martina, Stefano Marsi |
ICIP | 6 |
| 2022 | CoNLoCNN: Exploiting Correlation and Non-Uniform Quantization for Energy-Efficient Low-precision Deep Convolutional Neural NetworksabstractIn today's era of smart cyber-physical systems, Deep Neural Networks (DNNs) have become ubiquitous due to their state-of-the-art performance in complex real-world applications. The high computational complexity of these networks, which translates to increased energy consumption, is the foremost obstacle towards deploying large DNNs in resource-constrained systems. Fixed-Point (FP) implementations achieved through post-training quantization are commonly used to curtail the energy consumption of these networks. However, the uniform quantization intervals in FP restrict the bit-width of data structures to large values due to the need to represent most of the numbers with sufficient resolution and avoid high quantization errors. In this paper, we leverage the key insight that (in most of the scenarios) DNN weights and activations are mostly concentrated near zero and only a few of them have large magnitudes. We propose CoNLoCNN, a framework to enable energy-efficient low-precision deep convolutional neural network inference by exploiting: (1) non-uniform quantization of weights enabling simplification of complex multiplication operations; and (2) correlation between activation values enabling partial compensation of quantization errors at low cost without any run-time overheads. To significantly benefit from non-uniform quantization, we also propose a novel data representation format, Encoded Low-Precision Binary Signed Digit, to compress the bit-width of weights while ensuring direct use of the encoded weight for processing using a novel multiply-and-accumulate (MAC) unit design. Muhammad Abdullah Hanif, Giuseppe Maria Sarda, Alberto Marchisio, Guido Masera, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 5 |
| 2022 | fakeWeather: Adversarial Attacks for Deep Neural Networks Emulating Weather Conditions on the Camera Lens of Autonomous SystemsabstractRecently, Deep Neural Networks (DNNs) have achieved remarkable performances in many applications, while several studies have enhanced their vulnerabilities to malicious attacks. In this paper, we emulate the effects of natural weather conditions to introduce plausible perturbations that mislead the DNNs. By observing the effects of such atmospheric perturbations on the camera lenses, we model the patterns to create different masks that fake the effects of rain, snow, and hail. Even though the perturbations introduced by our attacks are visible, their presence remains unnoticed due to their association with natural events, which can be especially catastrophic for fully-autonomous and unmanned vehicles. We test our proposed fake Weather attacks on multiple Convolutional Neural Network and Capsule Network models, and report noticeable accuracy drops in the presence of such adversarial perturbations. Our work introduces a new security threat for DNNs, which is especially severe for safety-critical applications and autonomous systems. Alberto Marchisio, Giovanni Caramia, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2022 | LaneSNNs: Spiking Neural Networks for Lane Detection on the Loihi Neuromorphic ProcessorabstractAutonomous Driving (AD) related features represent important elements for the next generation of mobile robots and autonomous vehicles focused on increasingly intelligent, autonomous, and interconnected systems. The applications involving the use of these features must provide, by definition, real-time decisions, and this property is key to avoid catastrophic accidents. Moreover, all the decision processes must require low power consumption, to increase the lifetime and autonomy of battery-driven systems. These challenges can be addressed through efficient implementations of Spiking Neural Networks (SNNs) on Neuromorphic Chips and the use of event-based cameras instead of traditional frame-based cameras. In this paper, we present a new SNN-based approach, called LaneSNN, for detecting the lanes marked on the streets using the event-based camera input. We develop four novel SNN models characterized by low complexity and fast response, and train them using an offline supervised learning rule. Afterward, we implement and map the learned SNNs models onto the Intel Loihi Neuromorphic Research Chip. For the loss function, we develop a novel method based on the linear composition of Weighted binary Cross Entropy (WCE) and Mean Squared Error (MSE) measures. Our experimental results show a maximum Intersection over Union (IoU) measure of about 0.62 and very low power consumption of about 1 W. The best IoU is achieved with an SNN implementation that occupies only 36 neurocores on the Loihi processor while providing a low latency of less than 8 ms to recognize an image, thereby enabling real-time performance. The IoU measures provided by our networks are comparable with the state-of-the-art, but at a much low power consumption of 1 W. Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IROS | 3 |
| 2022 | Enabling Capsule Networks at the Edge through Approximate Softmax and Squash OperationsabstractComplex Deep Neural Networks such as Capsule Networks (CapsNets) exhibit high learning capabilities at the cost of compute-intensive operations. To enable their deployment on edge devices, we propose to leverage approximate computing for designing approximate variants of the complex operations like softmax and squash. In our experiments, we evaluate tradeoffs between area, power consumption, and critical path delay of the designs implemented with the ASIC design flow, and the accuracy of the quantized CapsNets, compared to the exact functions. Alberto Marchisio, Beatrice Bussolino, Edoardo Salvati, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
ISLPED | 4 |
| 2021 | Very Low Latency Architecture for Earth Observation Satellite Onboard Data Handling, Compression, and EncryptionabstractIn modern society, the ever-increasing demand for Earth Observation products in a large variety of sectors is exposing the limitations of traditional satellite data chain architectures. The European Union Horizon 2020 EO-ALERT project aims at overcoming the existing bottlenecks by leveraging the performance of state-of-the-art commercial off-the-shelf devices to move the critical elements of data processing on the flight segment without sacrificing processing performance. This paper introduces the architecture of the EO-ALERT CPU Scheduling, Compression, Encryption and Data Handling Subsystem, responsible for coordinating the onboard optical and Synthetic Aperture Radar data chains, as well as providing data compression, encryption, and storage services. The performance obtained by a reference implementation of the proposed architecture is also presented, showing an extremely low contribution to the overall system latency that allows real-time Earth Observation product delivery to the end user in less than 5 min. Michele Caon, Paolo Motto Ros, Maurizio Martina, Tiziano Bianchi, Enrico Magli, Francisco Membibre, Alexis Ramos, Antonio Latorre, Murray Kerr, Stefan Wiehle, Helko Breit, Dominik Günzel, Srikanth Mandapati, Ulrich Balss, Björn Tings |
IGARSS | 3 |
| 2021 | High-Level Synthesis of a Single/Multi-Band Optical and SAR Image Compression and Encryption Hardware AcceleratorabstractTransmitting images from earth observation satellites to ground is a major challenge, and a compression/encryption stage is actually mandatory. Development of hardware accelerators is highly recommended, both to relieve the software from such demanding task, and to improve performance, aiming at quasi-real-time data processing. To this end, we discuss the design, development, deployment and test of a FPGA-based accelerator, featuring a lossless and lossy (near-lossless) compression, including the data encryption too. Its architecture is well suited for different image types, including single- and multi-band optical and SAR images and can be fully run-time configurable. Measured performance showed a throughput of 10 Msamples/s, in agreement with related state-of-the-art works, focused on lossless compression only. Paolo Motto Ros, Michele Caon, Tiziano Bianchi, Maurizio Martina, Enrico Magli |
IGARSS | 4 |
| 2021 | DVS-Attacks: Adversarial Attacks on Dynamic Vision Sensors for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), despite being energy-efficient when implemented on neuromorphic hardware and coupled with event-based Dynamic Vision Sensors (DVS), are vulnerable to security threats, such as adversarial attacks, i.e., small perturbations added to the input for inducing a misclassification. Toward this, we propose DVS-Attacks, a set of stealthy yet efficient adversarial attack methodologies targeted to perturb the event sequences that compose the input of the SNNs. First, we show that noise filters for DVS can be used as defense mechanisms against adversarial attacks. Afterwards, we implement several attacks and test them in the presence of two types of noise filters for DVS cameras. The experimental results show that the filters can only partially defend the SNNs against our proposed DVS-Attacks. Using the best settings for the noise filters, our proposed Mask Filter-Aware Dash Attack reduces the accuracy by more than 20% on the DVS-Gesture dataset and by more than 65% on the MNIST dataset, compared to the original clean frames. The source code of all the proposed DVS-Attacks and noise filters is released at https://github.com/albertomarchisio/DVS-Attacks. Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2021 | CarSNN: An Efficient Spiking Neural Network for Event-Based Autonomous Cars on the Loihi Neuromorphic Research ProcessorabstractAutonomous Driving (AD) related features provide new forms of mobility that are also beneficial for other kind of intelligent and autonomous systems like robots, smart transportation, and smart industries. For these applications, the decisions need to be made fast and in real-time. Moreover, in the quest for electric mobility, this task must follow low power policy, without affecting much the autonomy of the mean of transport or the robot. These two challenges can be tackled using the emerging Spiking Neural Networks (SNNs). When deployed on a specialized neuromorphic hardware, SNNs can achieve high performance with low latency and low power consumption. In this paper, we use an SNN connected to an event-based camera for facing one of the key problems for AD, i.e., the classification between cars and other objects. To consume less power than traditional frame-based cameras, we use a Dynamic Vision Sensor (DVS) [1]. The experiments are made following an offline supervised learning rule, followed by mapping the learnt SNN model on the Intel Loihi Neuromorphic Research Chip [2]. Our best experiment achieves an accuracy on offline implementation of 86%, that drops to 83% when it is ported onto the Loihi Chip. The Neuromorphic Hardware implementation has maximum 0.72 ms of latency for every sample, and consumes only 310 mW. To the best of our knowledge, this work is the first implementation of an event-based car classifier on a Neuromorphic Chip. Alberto Viale, Alberto Marchisio, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2021 | R-SNN: An Analysis and Design Methodology for Robustifying Spiking Neural Networks against Adversarial Attacks through Noise Filters for Dynamic Vision SensorsabstractSpiking Neural Networks (SNNs) aim at providing energy-efficient learning capabilities when implemented on neuromorphic chips with event-based Dynamic Vision Sensors (DVS). This paper studies the robustness of SNNs against adversarial attacks on such DVS-based systems, and proposes R-SNN, a novel methodology for robustifying SNNs through efficient DVS-noise filtering. We are the first to generate adversarial attacks on DVS signals (i.e., frames of events in the spatio-temporal domain) and to apply noise filters for DVS sensors in the quest for defending against adversarial attacks. Our results show that the noise filters effectively prevent the SNNs from being fooled. The SNNs in our experiments provide more than 90% accuracy on the DVS-Gesture and NMNIST datasets under different adversarial threat models. Alberto Marchisio, Giacomo Pira, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IROS | 3 |
| 2021 | Analysis of in Vivo Plant Stem Impedance Variations in Relation with External Conditions Daily CycleabstractWorld population growth and desertification are the most severe issue to agricultural food production. Smart agriculture is a promising solution to ensure food security. The use of sensors to monitor crop production can help farmers improve the yield and reduce water consumption. Here we propose a study where the electrical impedance of green plants' stem is analyzed in vivo, along with environmental conditions. In particular, the variations associated with the daily cycle are highlighted. These analyses lead to the possibility of understanding plant status directly from stem impedance. Umberto Garlando, Lee Bar-on, Paolo Motto Ros, Alessandro Sanginario, Stefano Calvo, Maurizio Martina, Adi Avni, Yosi Shacham-Diamand, Danilo Demarchi |
ISCAS | 6 |
| 2021 | HW-FlowQ: A Multi-Abstraction Level HW-CNN Co-design Quantization MethodologyabstractModel compression through quantization is commonly applied to convolutional neural networks (CNNs) deployed on compute and memory-constrained embedded platforms. Different layers of the CNN can have varying degrees of numerical precision for both weights and activations, resulting in a large search space. Together with the hardware (HW) design space, the challenge of finding the globally optimal HW-CNN combination for a given application becomes daunting. To this end, we propose HW-FlowQ, a systematic approach that enables the co-design of the target hardware platform and the compressed CNN model through quantization. The search space is viewed at three levels of abstraction, allowing for an iterative approach for narrowing down the solution space before reaching a high-fidelity CNN hardware modeling tool, capable of capturing the effects of mixed-precision quantization strategies on different hardware architectures (processing unit counts, memory levels, cost models, dataflows) and two types of computation engines (bit-parallel vectorized, bit-serial). To combine both worlds, a multi-objective non-dominated sorting genetic algorithm (NSGA-II) is leveraged to establish a Pareto-optimal set of quantization strategies for the target HW-metrics at each abstraction level. HW-FlowQ detects optima in a discrete search space and maximizes the task-related accuracy of the underlying CNN while minimizing hardware-related costs. The Pareto-front approach keeps the design space open to a range of non-dominated solutions before refining the design to a more detailed level of abstraction. With equivalent prediction accuracy, we improve the energy and latency by 20% and 45% respectively for ResNet56 compared to existing mixed-precision search methods. Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Emanuele Valpreda, Driton Salihu, Nguyen Anh Vu Doan, Christian Unger, Naveen Shankar Nagaraja, Maurizio Martina, Walter Stechele |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2020 | NACU: A Non-Linear Arithmetic Unit for Neural NetworksabstractReconfigurable architectures targeting neural networks are an attractive option. They allow multiple neural networks of different types to be hosted on the same hardware, in parallel or sequence. Reconfigurability also grants the ability to morph into different micro-architectures to meet varying power-performance constraints. In this context, the need for a reconfigurable non-linear computational unit has not been widely researched. In this work, we present a formal and comprehensive method to select the optimal fixed-point representation to achieve the highest accuracy against the floating-point implementation benchmark. We also present a novel design of an optimised reconfigurable arithmetic unit for calculating non-linear functions. The unit can be dynamically configured to calculate the sigmoid, hyperbolic tangent, and exponential function using the same underlying hardware. We compare our work with the state-of-the-art and show that our unit can calculate all three functions without loss of accuracy. Guido Baccelli, Dimitrios Stathis 0001, Ahmed Hemani, Maurizio Martina |
DAC | 4 |
| 2020 | Q-CapsNets: A Specialized Framework for Quantizing Capsule NetworksabstractCapsule Networks (CapsNets), recently proposed by the Google Brain team, have superior learning capabilities in machine learning tasks, like image classification, compared to the traditional CNNs. However, CapsNets require extremely intense computations and are difficult to be deployed in their original form at the resource-constrained edge devices. This paper makes the first attempt to quantize CapsNet models, to enable their efficient edge implementations, by developing a specialized quantization framework for CapsNets. We evaluate our framework for several benchmarks. On a deep CapsNet model for the CIFAR10 dataset, the framework reduces the memory footprint by 6.2x, with only 0.15% accuracy loss. We will open-source our framework at https://git.io/JvDIF. Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
DAC | 4 |
| 2020 | NASCaps: A Framework for Neural Architecture Search to Optimize the Accuracy and Hardware Efficiency of Convolutional Capsule NetworksabstractDeep Neural Networks (DNNs) have made significant improvements to reach the desired accuracy to be employed in a wide variety of Machine Learning (ML) applications. Recently the Google Brain's team demonstrated the ability of Capsule Networks (CapsNets) to encode and learn spatial correlations between different input features, thereby obtaining superior learning capabilities compared to traditional (i.e., non-capsule based) DNNs. However, designing CapsNets using conventional methods is a tedious job and incurs significant training effort. Recent studies have shown that powerful methods to automatically select the best/optimal DNN model configuration for a given set of applications and a training dataset are based on the Neural Architecture Search (NAS) algorithms. Moreover, due to their extreme computational and memory requirements, DNNs are employed using the specialized hardware accelerators in IoT-Edge/CPS devices. Alberto Marchisio, Andrea Massa, Vojtech Mrazek, Beatrice Bussolino, Maurizio Martina, Muhammad Shafique 0001 |
ICCAD | 5 |
| 2020 | FasTrCaps: An Integrated Framework for Fast yet Accurate Training of Capsule NetworksabstractRecently, Capsule Networks (CapsNets) have shown improved performance compared to the traditional Convolutional Neural Networks (CNNs), by encoding and preserving spatial relationships between the detected features in a better way. This is achieved through the so-called Capsules (i.e., groups of neurons) that encode both the instantiation probability and the spatial information. However, one of the major hurdles in the wide adoption of CapsNets is their gigantic training time, which is primarily due to the relatively higher complexity of their new constituting elements that are different from CNNs.In this paper, we implement different optimizations in the training loop of the CapsNets, and investigate how these optimizations affect their training speed and the accuracy. Towards this, we propose a novel framework FasTrCaps that integrates multiple lightweight optimizations and a novel learning rate policy called WarmAdaBatch (that jointly performs warm restarts and adaptive batch size), and steers them in an appropriate way to provide high training-loop speedup at minimal accuracy loss. We also propose weight sharing for capsule layers. The goal is to reduce the hardware requirements of CapsNets by removing unused/redundant connections and capsules, while keeping high accuracy through tests of different learning rate policies and batch sizes. We demonstrate that one of the solutions generated by the FasTrCaps framework can achieve 58.6% reduction in the training time, while preserving the accuracy (even 0.12% accuracy improvement for the MNIST dataset), compared to the CapsNet by Google Brain [25]. Moreover, the Pareto-optimal solutions generated by FasTrCaps can be leveraged to realize trade-offs between training time and achieved accuracy. We have open-sourced our framework on GitHub1. Alberto Marchisio, Beatrice Bussolino, Alessio Colucci, Muhammad Abdullah Hanif, Maurizio Martina, Guido Masera, Muhammad Shafique 0001 |
IJCNN | 5 |
| 2020 | Is Spiking Secure? A Comparative Study on the Security Vulnerabilities of Spiking and Deep Neural NetworksabstractSpiking Neural Networks (SNNs) claim to present many advantages in terms of biological plausibility and energy efficiency compared to standard Deep Neural Networks (DNNs). Recent works have shown that DNNs are vulnerable to adversarial attacks, i.e., small perturbations added to the input data can lead to targeted or random misclassifications. In this paper, we aim at investigating the key research question: "Are SNNs secure?" Towards this, we perform a comparative study of the security vulnerabilities in SNNs and DNNs w.r.t. the adversarial noise. Afterwards, we propose a novel black-box attack methodology, i.e., without the knowledge of the internal structure of the SNN, which employs a greedy heuristic to automatically generate imperceptible and robust adversarial examples (i.e., attack images) for the given SNN. We perform an in-depth evaluation for a Spiking Deep Belief Network (SDBN) and a DNN having the same number of layers and neurons (to obtain a fair comparison), in order to study the efficiency of our methodology and to understand the differences between SNNs and DNNs w.r.t. the adversarial examples. Our work opens new avenues of research towards the robustness of the SNNs, considering their similarities to the human brain's functionality. Alberto Marchisio, Giorgio Nanfa, Faiq Khalid, Muhammad Abdullah Hanif, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 5 |
| 2020 | An Efficient Spiking Neural Network for Recognizing Gestures with a DVS Camera on the Loihi Neuromorphic ProcessorabstractSpiking Neural Networks (SNNs), the third generation NNs, have come under the spotlight for machine learning based applications due to their biological plausibility and reduced complexity compared to traditional artificial Deep Neural Networks (DNNs). These SNNs can be implemented with extreme energy efficiency on neuromorphic processors like the Intel Loihi research chip, and fed by event-based sensors, such as DVS cameras. However, DNNs with many layers can achieve relatively high accuracy on image classification and recognition tasks, as the research on learning rules for SNNs for real-world applications is still not mature. The accuracy results for SNNs are typically obtained either by converting the trained DNNs into SNNs, or by directly designing and training SNNs in the spiking domain. Towards the conversion from a DNN to an SNN, we perform a comprehensive analysis of such process, specifically designed for Intel Loihi, showing our methodology for the design of an SNN that achieves nearly the same accuracy results as its corresponding DNN. Towards the usage of the event-based sensors, we design a pre-processing method, evaluated for the DvsGesture dataset, which makes it possible to be used in the DNN domain. Hence, based on the outcome of the first analysis, we train a DNN for the pre-processed DvsGesture dataset, and convert it into the spike domain for its deployment on Intel Loihi, which enables real-time gesture recognition. The results show that our SNN achieves 89.64% classification accuracy and occupies only 37 Loihi cores. Riccardo Massa, Alberto Marchisio, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2020 | NeuroAttack: Undermining Spiking Neural Networks Security through Externally Triggered Bit-FlipsabstractDue to their proven efficiency, machine-learning systems are deployed in a wide range of complex real-life problems. More specifically, Spiking Neural Networks (SNNs) emerged as a promising solution to the accuracy, resource-utilization, and energy-efficiency challenges in machine-learning systems. While these systems are going mainstream, they have inherent security and reliability issues. In this paper, we propose NeuroAttack, a cross-layer attack that threatens the SNNs integrity by exploiting low-level reliability issues through a high-level attack. Particularly, we trigger a fault-injection based sneaky hardware backdoor through a carefully crafted adversarial input noise. Our results on Deep Neural Networks (DNNs) and SNNs show a serious integrity threat to state-of-the art machine-learning techniques. Valerio Venceslai, Alberto Marchisio, Ihsen Alouani, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 4 |
| 2020 | Towards Optimal Green Plant Irrigation: Watering and Body Electrical ImpedanceabstractWith the growth of world population and food demand, it is crucial to optimize water consumption for agriculture cultivation. Here we propose a method to monitor plant status, relating the measured parameters to the watering or drying situation of a single plant. Plant trunk electrical impedance measurements and environmental parameters were analyzed with a statistical approach. Correlation and causality among the data are showed and analyzed. In this way, it was possible to easily obtain the needed information about plant status. Umberto Garlando, Lee Bar-on, Paolo Motto Ros, Alessandro Sanginario, Sebastian Peradotto, Yosi Shacham-Diamand, Adi Avni, Maurizio Martina, Danilo Demarchi |
ISCAS | 8 |
| 2020 | An Area-Efficient Variable-Size Fixed-Point DCT Architecture for HEVC EncodingabstractThis paper proposes an area-efficient fixed-point architecture for the computation of the discrete cosine transform (DCT) of multiple sizes in high efficiency video coding (HEVC). This result is obtained by comparing different DCT factorizations in order to find the most suitable one for implementation in the HEVC encoder. The recursive structure of fast algorithms, which decompose the N-point DCT by means of two N/2-point DCTs, is exploited to execute computations of small-size DCTs in parallel, thus maximizing the hardware reusability while maintaining a constant throughput. The simulation results prove that the proposed solution features reduced rate-distortion losses, with relevant complexity saving compared with the state-of-the-art implementations. Finally, the proposed architecture is exploited to design two families of architectures for the 2D-DCT, namely, folded and full-parallel. Maurizio Masera, Guido Masera, Maurizio Martina |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A Novel Framework for Designing Directional Linear Transforms with Application to Video CompressionabstractTransforms incorporating directional information are appealing in a wide range of applications. In this paper, we introduce a new framework that allows to define a directional transform starting from any two-dimensional separable transform. The proposed method is highly general and it can be of interest in many areas of signal processing. We also show an example of possible application. We define a directional integer DCT and DST and we show their application in video compression by integrating them in the HEVC video coding standard. Maurizio Masera, Giulia Fracastoro, Maurizio Martina, Enrico Magli |
ICASSP | 3 |
| 2019 | Live Demonstration: Event-Driven Serial Communication on Optical FiberabstractThis demonstration presents the first implementation of “event-driven” serial asynchronous communication on optical fiber. “event-driven” communication is used by neuromorphic sensors, that sample the sensory signal when the signal itself changes of a given amount. This type of sensing adapts to the dynamics of the input itself, achieving at the same time extremely high temporal resolution (when needed), low latency and signal compression. To apply this technology in robotics, optical communication will greatly improve the resilience to electric disturbances. Andrea De Marcellis, Guido Di Patrizio Stanchieri, Marco Faccio, Elia Palange, Paolo Motto Ros, Maurizio Martina, Danilo Demarchi, Chiara Bartolozzi |
ISCAS | 6 |
| 2018 | PruNet: Class-Blind Pruning Method For Deep Neural NetworksabstractDNNs are highly memory and computationally intensive, due to which they are unfeasible to deploy in real time or mobile applications, where power and memory resources are scarce. Introducing sparsity in the network is a way to reduce those requirements. However, systematically employing pruning under given accuracy requirements is a challenging problem. We propose a novel methodology that iteratively applies a magnitude-based Class-Blind pruning to compress a DNN for obtaining a sparse model. It is a generic methodology and can be applied to different types of DNNs. We demonstrate that retraining after pruning is essential to restore the accuracy of the network. Experimental results show that our methodology is able to reduce the model size by around two orders of magnitude, without noticeably affecting the accuracy. It requires several iterations of pruning and retraining, but can achieve up to 190x Memory Saving Ratio (for the LeNet on the MNIST dataset) when compared to the baseline model. Similar results are also obtained for more complex networks like 91x for VGG-16 on the CIFAR100 dataset. If we combine this work with an efficient coding for sparse networks, like Compressed Sparse Column (CSC) or Compressed Sparse Row (CSR), we can obtain a reduced memory footprint. Our methodology can be complemented by other compression techniques, like weight sharing, quantization or fixed-point conversion, that allows to further reduce memory and computations. Alberto Marchisio, Muhammad Abdullah Hanif, Maurizio Martina, Muhammad Shafique 0001 |
IJCNN | 3 |
| 2018 | Live Demonstration: Tactile Events from Off-The-Shelf Sensors in a Robotic SkinabstractThe demonstration presents a robotic event-based tactile infrastructure for a humanoid robot. It leverages on currently deployed sample-based capacitive sensors to generate tactile events, enabling the investigation and development of event-driven tactile applications, and minimizing communication bandwidth and latency. The modular FPGA-based system samples data from tactile sensors and generates address-events, transmitted through an asynchronous serial address-event representation protocol. To enable performance comparisons of the event-driven approach with respect to standard sample-based solutions, the acquisition modules can directly forward the input samples through the same event-based communication channel. We will show in real time a comparison between the tactile events and the original sampled data generated when the skin patch is touched. Chiara Bartolozzi, Paolo Motto Ros, Riccardo Peloso, Francesco Diotalevi, Marco Crepaldi, Maurizio Martina, Danilo Demarchi |
ISCAS | 6 |
| 2017 | Analysis of HEVC transform throughput requirements for hardware implementations
Maurizio Masera, Lorenzo Re Fiorentin, Enrico Masala, Guido Masera, Maurizio Martina |
Signal Process. Image Commun. | 5 |
| 2017 | Adaptive Approximated DCT Architectures for HEVCabstractThis paper proposes a flexible and efficient implementation of the 2D N-point discrete cosine transform (DCT) for the High Efficiency Video Coding (HEVC) standard. The DCT is implemented through the Walsh-Hadamard transform (WHT) followed by Givens rotations. This scheme is exploited to derive an adaptive algorithm, which allows computing of four different approximations ranging from the complete DCT to the WHT, by selectively skipping some rotations. This paper shows the statistical analysis of the DCT usage and derives a precomputation mechanism to adaptively skip rotations. Each approximation, referred to as a operating mode, is characterized by a large saving of operations, at the expense of very small quality loss. Then, two 2D-DCT architectures are proposed: the first one is totally unfolded, while the second one is folded. The two designs are finally synthesized with a 90-nm standard-cell library for a clock frequency of 250 MHz. Both architectures support real-time processing of 8K UHD video sequences at 64 and 26 fps, respectively, and show higher throughput and lower gate count compared with the state-of-art implementations. Moreover, power saving ranging from 28% to 56% can be achieved by working within the proposed operating modes. Maurizio Masera, Maurizio Martina, Guido Masera |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | An all-digital spike-based ultra-low-power IR-UWB dynamic average threshold crossing scheme for muscle force wireless transmission
Masoud Shahshahani Amirhossein, Paolo Motto Ros, Alberto Bonanno, Marco Crepaldi, Maurizio Martina, Danilo Demarchi, Guido Masera |
DATE | 5 |
| 2015 | Using Information Centric Networking for Mobile Devices Cooperation at the Network EdgeabstractThis paper presents an Android application exploiting a collaborative multi-RAT approach (cellular and Wi-Fi Radio Access Technologies) for adaptively streaming video content. The application has been developed and tested on real Android devices of different vendors. Each device coordinates with the others to download portions of a video stream through its own preferred RAT, and shares the results with the neighboring devices through a proximity, high bit rate network based on Wi-Fi Direct. As a result, the aggregate bandwidth available to the group of collaborating devices allows a faster download and a higher quality of video playback streaming, with respect to what is achievable by a single device. The application is designed to exploit common Android devices along with the key functionalities of an Information Centric Network: routing-by-name, in-network caching and multicasting. Fabio Malabocchia, Romeo Corgiolu, Maurizio Martina, Andrea Detti, Bruno Ricci, Nicola Blefari-Melazzi |
VTC Spring | 3 |
| 2015 | Exploiting generalized de-Bruijn/Kautz topologies for flexible iterative channel code decoder architectures
Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
Integr. | 2 |
| 2015 | Parallel H.264/AVC Fast Rate-Distortion Optimized Motion Estimation by Using a Graphics Processing Unit and Dedicated HardwareabstractHeterogeneous systems on a single chip composed of a central processing unit, graphics processing unit (GPU), and field-programmable gate array (FPGA) are expected to emerge in the near future. In this context, the system on chip can be dynamically adapted to employ different architectures for execution of data-intensive applications. Motion estimation (ME) is one such task that can be accelerated using FPGA and GPU for high-performance H.264/Advanced Video Coding encoder implementation. This paper presents an inherent parallel low-complexity rate-distortion (RD) optimized fast ME algorithm well suited for parallel implementations, eliminating various data dependencies caused by a reliance on spatial predictions. In addition, this paper provides details of the GPU and FPGA implementations of the parallel algorithm by using OpenCL and Very High Speed Integrated Circuits (VHSIC) Hardware Descriptive Language (VHDL), respectively, and presents a practical performance comparison between the two implementations. The experimental results show that the proposed scheme achieves significant speedup on GPU and FPGA, and has comparable RD performance with respect to sequential fast ME algorithm. Muhammad Usman Shahid, Ashfaq Ahmed, Maurizio Martina, Guido Masera, Enrico Magli |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Rediscovering Logarithmic Diameter Topologies for Low Latency Network-on-Chip-Based ApplicationsabstractLow-latency Network-on-Chip (NoC) applications have tight constraints on the clock budget to perform communication among nodes. This is a critical aspect in NoC-based designs where the number of clock cycles spent for communication depends mainly on the topology and on the routing algorithm. This work deals with logarithmic diameter topologies, that were proposed for computer networks, and shows that an optimal shortest-path routing algorithm can be efficiently implemented on this kind of topologies by means of a very simple circuit. The proposed circuit is then exploited to reduce the area and the power consumption of a recently proposed NoC-based design. Experimental results show that the proposed circuit allows for a reduction of about 14% and 10% for area and power consumption respectively, with respect to a shortest-path routing-table-based design. Carlo Condo, Maurizio Martina, Massimo Ruo Roch, Guido Masera |
PDP | 2 |
| 2014 | Variable Parallelism Cyclic Redundancy Check Circuit for 3GPP-LTE/LTE-AdvancedabstractCyclic Redundancy Check (CRC) is often employed in data storage and communications to detect errors. The 3GPP-LTE wireless communication standard uses a 24-bit CRC with every turbo coded frame, thus, the CRC can be exploited to detect residual errors and to enable early stopping of iterations as well. The current state of the art lacks specific CRC implementations for this standard, and most current solutions adopt a fixed degree of parallelism, unsuitable for many turbo decoder architectures. This work proposes a variable parallelism circuit targeting the 3GPP-LTE/LTE-Advanced 24-bit CRC, that can adapt to input data of different sizes. Low complexity is achieved through careful functional sharing among the various parallelisms: comparison with the state of the art shows comparable or superior speed and extremely low complexity. Carlo Condo, Maurizio Martina, Gianluca Piccinini, Guido Masera |
IEEE Signal Process. Lett. | 2 |
| 2013 | VLSI Architecture for Low-Complexity Motion Estimation in H.264 Multiview Video CodingabstractThis paper presents a VLSI architecture for a low complexity motion estimation algorithm, referred to as Slim264, for multiview video coding extension of H.264. Algorithmic modifications are introduced to obtain a fully parallel computational structure able to meet the throughput requirements of high resolution and high frame rate videos. High parallelism is achieved by predicting small blocks, i.e. 4x4 pixel blocks, in parallel and then adding them up in order to get Sum of Absolute Differences (SADs) of large block sizes. The predictor is able to support high resolution videos i.e. 1080p. The modified algorithm shows promising PSNR results with respect to full search algorithm. The predictor is synthesized with a clock frequency of 200 MHz, occupying an area of 0.49 mm2, on 90-nm Standard Cell ASIC technology. Ashfaq Ahmed, Muhammad Usman Shahid, Maurizio Martina, Enrico Magli, Guido Masera |
DSD | 3 |
| 2012 | A Network-on-Chip-based turbo/LDPC decoder architectureabstractThe current convergence process in wireless technologies demands for strong efforts in the conceiving of highly flexible and interoperable equipments. This contribution focuses on one of the most important baseband processing units in wireless receivers, the forward error correction unit, and proposes a Network-on-Chip (NoC) based approach to the design of multi-standard decoders. High level modeling is exploited to drive the NoC optimization for a given set of both turbo and Low-Density-Parity-Check (LDPC) codes to be supported. Moreover, synthesis results prove that the proposed approach can offer a fully compliant WiMAX decoder, supporting the whole set of turbo and LDPC codes with higher throughput and an occupied area comparable or lower than previously reported flexible implementations. In particular, the mentioned design case achieves a worst-case throughput higher than 70 Mb/s at the area cost of 3.17 mm2on a 90 nm CMOS technology. Carlo Condo, Maurizio Martina, Guido Masera |
DATE | 2 |
| 2012 | Non-recursive max* operator with reduced implementation complexity for turbo decodingabstractIn this study, the authors deal with the problem of how to effectively approximate the max* operator when having n>2 input values, with the aim of reducing implementation complexity of conventional Log-MAP turbo decoders. They show that, contrary to previous approaches, it is not necessary to apply the max* operator recursively over pairs of values. Instead, a simple, yet effective, solution for the max* operator is revealed having the advantage of being in non-recursive form and thus, requiring less computational effort. Hardware synthesis results for practical turbo decoders have shown implementation savings for the proposed method against the most recent published efficient turbo decoding algorithms by providing near optimal bit error rate (BER) performance. Stylianos Papaharalabos, P. Takis Mathiopoulos, Guido Masera, Maurizio Martina |
IET Commun. | 4 |
| 2012 | High Speed Architectures for Finding the First two Maximum/Minimum ValuesabstractHigh speed architectures for finding the first two maximum/minimum values are of paramount importance in several applications, including iterative (e.g., turbo and low-density-parity-check) decoders. In this brief, stemming from a previous work, based on radix-2 solutions, we propose higher and mixed radix implementations that improve the architecture latency. Post place and route results on a 180-nm CMOS standard cell technology show that the proposed architectures achieve lower latency than radix-2 solutions with a moderate area increase. Luca G. Amarù, Maurizio Martina, Guido Masera |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Efficient Implementation Techniques for Maximum Likelihood-Based Error Correction for JPEG2000abstractFeatures of the JPEG2000 compression standard include coding efficiency at low bit rates. However, its compressed bit stream is sensitive to transmission error. This paper presents three techniques to reduce both the computational complexity and the memory requirement in the ternary MQ arithmetic decoding. Such coders introduce a controlled degree of redundancy during the encoding process, which can be exploited at the decoder side in order to detect and correct errors. Our proposed techniques result in a substantial saving of decoding time and memory usage, with no or little degradation in the PSNR metric. Simone Zezza, Saeid Nooshabadi, Maurizio Martina, Guido Masera |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Error resilient JPEG2000 decoding for wireless applicationsabstractTo improve the JPEG2000 compression standard error resiliency in the wireless environment, the use of ternary MQ arithmetic coders/decoders that are based on the concept of forbidden symbol has been proposed. This paper presents two ternary MQ based techniques to reduce both the computational complexity and the memory requirement during the decoding process, with no or little degradation in the PSNR. Simone Zezza, Maurizio Martina, Guido Masera, Saeid Nooshabadi |
ICIP | 2 |
| 2007 | Flexible blocks for high throughput serially concatenated convolutional codesabstractModern communication systems aim to integrate different services including telephony, multimedia applications and data transfer. This integration imposes to design flexible systems, able to support different frame duration and data-rates. In this paper we present several blocks employed in serially concatenated convolutional codes and we detail some hardware solutions to achieve flexibility. Maurizio Martina, Guido Masera |
ACM Great Lakes Symposium on VLSI | 1 |
| 2007 | Real-time implementation of a time-frequency analysis schemeabstractThe on-line monitoring and detection of defects in laser welding is a basic manifacturing requirement in several applicative contexts, including vehicle assembly in automotive production. The speed of assembly in modern industry and the large amount of data to be acquired and elaborated pose severe real-time constraints and lead to the need of extremely short processing latencies. In this work, the implementation of time-frequency analysis algorithms on an FPGA device is shown and compared to pure software developments on different processors. Maurizio Martina, Andrea Terreno, Fabrizio Vacca, Andrea Molino, Guido Masera, Giuseppe D'Angelo, Giorgio Pasquettaz |
ACM Great Lakes Symposium on VLSI | 1 |
| 2006 | Mumford and Shah Functional: VLSI Analysis and ImplementationabstractThis paper describes the analysis of the Mumford and Shah functional from the implementation point of view. Our goal is to show results in terms of complexity for real-time applications, such as motion estimation based on segmentation techniques, of the Mumford and Shah functional. Moreover, the sensitivity to finite precision representation is addressed, a fast VLSI architecture is described, and results obtained for its complete implementation on a 0.13 microm standard cells technology are presented. Maurizio Martina, Guido Masera |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Low-complexity, efficient 9/7 wavelet filters implementationabstractThis paper proposes a novel low-complexity, efficient 9/7 wavelet filters implementation for image compression applications. The 9/7 wavelet filters are widely used in different image compression schemes, such as the JPEG2000 image coding standard. Thus the implementation of efficient codecs is of great concern. The performance of a hardware implementation of the 9/7 filter bank depends on the accuracy with which filter coefficients are represented. However the greatest part of current implementations consider filters taps as numbers to be implemented. The aim of this work is to show that great complexity reduction can be achieved going through the derivation of the 9/7 taps values. Maurizio Martina, Guido Masera |
ICIP (3) | 1 |
| 2003 | A reconfigurable, power-scalable rake receiver IP for W-CDMAabstractDuring the last years wireless market has experienced an exponential growth. 2G systems are essentially voice-oriented: the main innovation expected from 3G ones is the ubiquitous Internet and multimedia fruition. The transition from 2G to 3G provides both opportunities and challenges: one way to make this migration as smoother as possible relies on the employment of reconfigurable architectures. In this paper a reconfigurable Rake Receiver for W-CDMA is proposed. Very promising results from the physical implementation on a XCV300E have been obtained. A. Bianco, Alberto Dassatti, Maurizio Martina, Andrea Molino, Fabrizio Vacca |
ASP-DAC | 3 |
| 2003 | Wireless sensor networks: a power-scalable motion estimation IP for hybrid video codingabstractWireless Sensor Networks are an emerging phenomenon in the research community. The design and development of network architectures and nodes implementation are fostering many research activities. Due to their wide application fields and pervasive employment possibilities, the investigation of novel classes of wireless sensor nodes is of great concern. In this paper we presented a novel Power-Scalable Motion Estimation IP suitable for video-surveillance over Wireless Sensor Networks. The proposed architecture can achieve low dynamic power, good video quality and reasonable frame-rates. Moreover, it can operate on different video format and can be reconfigured both off-line and on-line. Further researches have to be accomplished in order to reduce the internal critical path, achieving better maximum frequencies performances. Moreover, new FPGA architectures must be exploited searching low-static-power devices suitable for an actual implementation of our IP on a WSN. Federico Quaglio, Maurizio Martina, Fabrizio Vacca, Guido Masera, Andrea Molino, Gianluca Piccinini, Maurizio Zamboni |
FPGA | 2 |
| 2003 | Design of a Power Conscious, Customizable CDMA Receiver
Maurizio Martina, Andrea Molino, Mario Nicola, Fabrizio Vacca |
FPL | 1 |
| 2003 | A Power-Scalable Motion Estimation Architecture for Energy Constrained Applications
Maurizio Martina, Andrea Molino, Federico Quaglio, Fabrizio Vacca |
FPL | 1 |
| 2003 | Implementation of a SPIHT coprocessor: memory issues and hardware implicationsabstractIn the last years mobile terminals diffusion has fostered the development of new audio-visual services. Thus the possibility to employ wireless devices to exchange multimedia informations is becoming real. Beyond these attractive motivations, some critical factors are hidden: the design of wireless terminals, in fact, has to deal with limited power budgets and reduced resources availability. This paper proposes a hardware implementation of a SPIHT coprocessor, designed for mobile environments. Particular emphasis will be posed on memory requirements and their hardware implications. Maurizio Martina, Andrea Molino, Andrea Terreno, Fabrizio Vacca |
ICIP (2) | 1 |
| 2003 | Dynamic power scheduling system for JPEG2000 delivery over wireless networks
Maurizio Martina, Fabrizio Vacca |
VCIP | 1 |
| 2002 | Energy Evaluation on a Reconfigurable, Multimedia-Oriented Wireless Sensor
Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni |
FPL | 1 |
| 2002 | Reconfigurable DSP IP for multimedia applicationsabstractIn this paper a novel Digital Signal Processor IP for multimedia applications, is presented. Recently, develeper's interest towards SOC architectures has been driven by mobile market explosion. Despite the increasing importance gathered by reconfigurable computing, a lack of easily retargettable cores is felt by developer's community. This IP is intended to be the kernel for many telecommunication and multimedia algorithms computation. During the design flow, much care has been devoted to grant maximum interoperability among this DSP core and other coprocessor units, allowing to easily embed multiple functional blocks on a single FPGA. As far as performance are concerned, the proposed IP shows satisfactory results both in terms of area occupation (11\% on a XILINX XCV1000) and maximum clock frequency (89 MHz after place and route process). Maurizio Martina, Guido Masera, Gianluca Piccinini, Fabrizio Vacca, Maurizio Zamboni |
ICASSP | 1 |
| 2002 | Reconfigurable and low power 2D-DCT IP for ubiquitous multimedia streamingabstractAn energy efficient architecture for the discrete cosine transform is presented. The proposed IP is intended to be used as the transform stage in a mobile H.263 codec. In particular, it seems well suited for an FPGA implementation since, after a complete place and route process, it is able to sustain a full-motion PAL video streaming operating at a frequency of 74 MHz with a dynamic power dissipation of just 39 mW. Maurizio Martina, Andrea Molino, Fabrizio Vacca |
ICME (2) | 1 |
| 2002 | Optimization and implementation of the integer wavelet transform for image codingabstractThis paper deals with the design and implementation of an image transform coding algorithm based on the integer wavelet transform (IWT). First of all, criteria are proposed for the selection of optimal factorizations of the wavelet filter polyphase matrix to be employed within the lifting scheme. The obtained results lead to the IWT implementations with very satisfactory lossless and lossy compression performance. Then, the effects of finite precision representation of the lifting coefficients on the compression performance are analyzed, showing that, in most cases, a very small number of bits can be employed for the mantissa keeping the performance degradation very limited. Stemming from these results, a VLSI architecture is proposed for the IWT implementation, capable of achieving very high frame rates with moderate gate complexity. Marco Grangetto, Enrico Magli, Maurizio Martina, Gabriella Olmo |
IEEE Trans. Image Process. | 3 |