VLDB 2026 Research / reviewers in the wild / expert
Mauro Olivieri
dblp:74/3472
· DBLP profile ↗
48ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-0214-9904ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Impact and Characterization of Single Event Failures in Hyperdimensional Computing Hardware Accelerator
Marcello Barbirotta, Marco Angioli, Rocco Martino, Federico Frontera, Antonio Mastrandrea, Mauro Olivieri |
ETS | 6 |
| 2026 | RAS Enhancement of ECC-Protected Vector Register File for the RISC-V Architecture via RERI-Compliant Interface
Marcello Barbirotta, Nicasio Canino, Giovanni Mazzini, Mauro Olivieri, Daniele Rossi 0001, Sergio Saponara |
IOLTS | 4 |
| 2026 | Configurable Hardware Acceleration for Hyperdimensional Computing Extension on RISC-VabstractHyperdimensional Computing (HDC) is a neuro-inspired computational model that represents and manipulates information using high-dimensional distributed representations that are combined and compared using simple and highly parallel vector operations. In this work, we present a highly flexible hardware acceleration unit for HDC learning tasks based on the binary spatter-code model. Integrated into the execution stage of the Klessydra-T03 RISC-V core, the unit accelerates the core arithmetic operations of HDC and can be configured at synthesis time in terms of hardware parallelism, supported operations, and size of the local memories, trading off execution time with hardware resources to match application needs. A custom RISC-V Instruction Set Extension efficiently controls the accelerator, with instructions fully integrated into the GNU Compiler Collection toolchain and exposed to the programmer as intrinsics. Dedicated Control and Status Registers allow specifying the characteristics of the high-dimensional space and the target learning tasks at runtime, controlling the hardware loops of the accelerator and enabling the same hardware architecture to be used for various tasks. The dual flexibility coming from hardware configuration and software programmability sets this work apart from application-specific solutions in the literature, offering a unique, versatile accelerator adaptable to a wide range of applications and learning tasks. Rocco Martino, Marco Angioli, Antonello Rosato, Marcello Barbirotta, Abdallah Cheikh, Mauro Olivieri |
IEEE Trans. Computers | 6 |
| 2025 | Towards RISC-V-based HPC: The Italian Pathfinding Activities in the DARE-SGA1 ProjectabstractThe European Union’s efforts towards technological sovereignty in High-Performance Computing are driving research and development of RISC-V-based supercomputers. The DARE SGA1 project, in particular, aims to develop chips designed and owned by Europeans. This paper introduces the Italian contribution to DARE SGA1 regarding pathfinding activities toward future RISC-V-based accelerator designs, reliability improvements, system software, and AI and Quantum Chemistry applications. Giovanni Agosta, Marco Aldinucci, Andrea Bartolini, Laura Bellentani, Andrea Biagioni, Daniele Cesarini, Carlotta Chiarini, Iacopo Colonnelli, Pietro Delugas, Lev Denisov, Ottorino Frezza, Marco Grangetto, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Andrea Maslov, Mauro Olivieri, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Davide Rossi 0001, Sergio Saponara, Antonio Sciarappa, Francesco Simula, Matteo Sonza Reorda, Massimo Torquati, Piero Vicini |
DSD | 17 |
| 2025 | Innovative Techniques for Efficient Hyperdimensional Computing on Hardware: Enhance Accuracy and On-the-Fly Hypervector Generation
Saeid Jamili, Sabereh Taghdisi Rastkar, Marco Angioli, Mauro Olivieri |
IJCCI (3) | 4 |
| 2025 | HD-CBBIN: A Lightweight Approach for Contextual Bandit Learning in Real-Time ApplicationsabstractAs the Internet of Things expands, the need to embed artificial intelligence algorithms into resource-constrained devices for real-time applications is growing. These systems require efficient and scalable algorithms to perform rapid and autonomous decision-making without relying on centralized cloud servers. Hyperdimensional Computing (HDC) has recently emerged as a compelling paradigm for learning tasks in such environments, offering computational efficiency, exceptional parallelism, and scalability.In this work, we present HD-CBBIN, a lightweight and efficient implementation of the HD-CB framework for modeling and automating sequential decision-making Contextual Bandits (CB) problems on embedded systems. By introducing modifications to the original algorithm, HD-CBBINexclusively uses binary hypervectors, significantly reducing computational demands. We benchmark the performance of HD-CBBINon synthetic datasets, comparing it to the real-valued counterpart and the traditional state-of-the-art LinUCB algorithm. We also evaluate its execution time, computational complexity, and memory requirements on various embedded platforms and demonstrate additional gains through hardware acceleration. The results show that our approach achieves linear execution time with respect to the context vector size and up to a 141× speedup over LinUCB while maintaining competitive performance, establishing a new milestone in contextual bandit algorithms for time-critical, resource-constrained applications. Marco Angioli, Antonello Rosato, Marcello Barbirotta, Rocco Martino, Andrea Marcelli, Antonio Mastrandrea, Mauro Olivieri |
IJCNN | 7 |
| 2025 | Efficient Implementation of LinearUCB through Algorithmic Improvements and Vector Computing Acceleration for Embedded Learning SystemsabstractAs the Internet of Things expands, embedding Artificial Intelligence algorithms in resource-constrained devices has become increasingly important to enable real-time, autonomous decision-making without relying on centralized cloud servers. However, implementing and executing complex algorithms in embedded devices poses significant challenges due to limited computational power, memory, and energy resources. 2This article presents algorithmic and hardware techniques to efficiently implement two LinearUCB Contextual Bandits algorithms on resource-constrained embedded devices. Algorithmic modifications based on the Sherman–Morrison–Woodbury formula streamline model complexity, while vector acceleration is harnessed to speed up matrix operations. We analyze the impact of each optimization individually and then combine them in a two-pronged strategy. The results show notable improvements in execution time and energy consumption, demonstrating the effectiveness of combining algorithmic and hardware optimizations to enhance learning models for edge computing environments with low-power and real-time requirements. Marco Angioli, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Mauro Olivieri |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2024 | AeneasHDC: An Automatic Framework for Deploying Hyperdimensional Computing Models on FPGAsabstractHyperdimensional Computing (HDC) is a bio-inspired learning paradigm, that models neural pattern activities using high-dimensional distributed representations. HDC leverages parallel and simple vector arithmetic operations to combine and compare different concepts, enabling cognitive and reasoning tasks. The computational efficiency and parallelism of this approach make it particularly suited for hardware implementations, especially as a lightweight, energy-efficient solution for performing learning tasks on resource-constrained edge devices. The HDC pipeline, including encoding, training, and comparison stages, has been extensively explored with various approaches in the literature. However, while these techniques are mainly oriented to improve the model accuracy, their influence on hardware parameters remains largely unexplored. This work presents AeneasHDC, an automatic and open-source platform for the streamlined deployment of HDC models in both software and hardware for classification, regression and clustering tasks. AeneasHDC supports an extensive range of techniques commonly adopted in literature, automates the design of flexible hardware accelerators for HDC, and empowers users to easily assess the impact of different design choices on model accuracy, memory usage, execution time, power consumption, and area requirements. Marco Angioli, Saeid Jamili, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Antonello Rosato, Mauro Olivieri |
IJCNN | 8 |
| 2024 | Design, Implementation and Evaluation of a New Variable Latency Integer Division SchemeabstractInteger division is key for various applications and often represents the performance bottleneck due to its inherent mathematical properties that limit its parallelization. This paper presents a new datadependent variable latency division algorithm, derived from the classic non-performing restoring method. The proposed technique exploits the relationship between the number of leading zeros in the divisor and in the partial remainder to dynamically detect and skip those iterations that result in a simple left shift. While a similar principle has been exploited in previous works, the proposed approach outperforms existing variable latency divider schemes in average latency and power consumption. We detail the algorithm and its implementation in four variants, offering versatility for the specific application requirements. For each variant, we report the average latency evaluated with different benchmarks, and we analyze the synthesis results for both FPGA and ASIC deployment, reporting clock speed, average execution time, hardware resources, and energy consumption, compared with existing fixed and variable latency dividers. Marco Angioli, Marcello Barbirotta, Abdallah Cheikh, Antonio Mastrandrea, Francesco Menichelli, Saeid Jamili, Mauro Olivieri |
IEEE Trans. Computers | 7 |
| 2023 | Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesabstractDeep Neural Network (DNN) inference based on quantized narrow-precision integer data represents a promising research direction toward efficient deep learning computations on edge and mobile devices. On one side, recent progress of Quantization-Aware Training (QAT) frameworks aimed at improving the accuracy of extremely quantized DNNs allows achieving results close to Floating-Point 32 (FP32), and provides high flexibility concerning the data sizes selection. Unfortunately, current Central Processing Unit (CPU) architectures and Instruction Set Architectures (ISAs) targeting resource-constrained devices present limitations on the range of data sizes supported to compute DNN kernels.This paper presents Mix-GEMM, a hardware-software co-designed architecture capable of efficiently computing quantized DNN convolutional kernels based on byte and sub-byte data sizes. Mix-GEMM accelerates General Matrix Multiplication (GEMM), representing the core kernel of DNNs, supporting all data size combinations from 8- to 2-bit, including mixed-precision computations, and featuring performance that scale with the decreasing of the computational data sizes. Our experimental evaluation, performed on representative quantized Convolutional Neural Networks (CNNs), shows that a RISC-V based edge System-on-Chip (SoC) integrating Mix-GEMM achieves up to 1.3 TOPS/W in energy efficiency, and up to 13.6 GOPS in throughput, gaining from 5.3× to 15.1× in performance over the OpenBLAS GEMM frameworks running on a commercial RISC-V based edge processor. By performing synthesis and Place and Route (PnR) of the enhanced SoC in Global Foundries 22nm FDX technology, we show that Mix-GEMM only accounts for 1% of the overall area consumption. Enrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó, Mauro Olivieri, Osman S. Unsal, Adrián Cristal |
HPCA | 5 |
| 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing ApplicationsabstractThe maturity level of RISC-V and the availability of domain-specific instruction set extensions, like vector processing, make RISC-V a good candidate for supporting the integration of specialized hardware in processor cores for the High Performance Computing (HPC) application domain. In this article, 1 we present Vitruvius+, the vector processing acceleration engine that represents the core of vector instruction execution in the HPC challenge that comes within the EuroHPC initiative. It implements the RISC-V vector extension (RVV) 0.7.1 and can be easily connected to a scalar core using the Open Vector Interface standard. Vitruvius+ natively supports long vectors: 256 double precision floating-point elements in a single vector register. It is composed of a set of identical vector pipelines (lanes), each containing a slice of the Vector Register File and functional units (one integer, one floating point). The vector instruction execution scheme is hybrid in-order/out-of-order and is supported by register renaming and arithmetic/memory instruction decoupling. On a stand-alone synthesis, Vitruvius+ reaches a maximum frequency of 1.4 GHz in typical conditions (TT/0.80V/25°C) using GlobalFoundries 22FDX FD-SOI. The silicon implementation has a total area of 1.3 mm 2 and maximum estimated power of ∼920 mW for one instance of Vitruvius+ equipped with eight vector lanes. Francesco Minervini, Oscar Palomar, Osman S. Unsal, Enrico Reggiani, Josue V. Quiroga, Joan Marimon, Carlos Rojas 0001, Roger Figueras, Abraham Ruiz, Alberto González 0004, Jonnatan Mendoza, Iván Vargas 0001, César Hernández, Joan Cabre, Lina Khoirunisya, Mustapha Bouhali, Julian Pavon, Francesc Moll, Mauro Olivieri, Mario Kovac, Mate Kovac, Leon Dragic, Mateo Valero, Adrián Cristal |
ACM Trans. Archit. Code Optim. | 19 |
| 2022 | BiSon-e: a lightweight and high-performance accelerator for narrow integer linear algebra computing on the edgeabstractLinear algebra computational kernels based on byte and sub-byte integer data formats are at the base of many classes of applications, ranging from Deep Learning to Pattern Matching. Porting the computation of these applications from cloud to edge and mobile devices would enable significant improvements in terms of security, safety, and energy efficiency. However, despite their low memory and energy demands, their intrinsically high computational intensity makes the execution of these workloads challenging on highly resource-constrained devices. In this paper, we present BiSon-e, a novel RISC-V based architecture that accelerates linear algebra kernels based on narrow integer computations on edge processors by performing Single Instruction Multiple Data (SIMD) operations on off-the-shelf scalar Functional Units (FUs). Our novel architecture is built upon the binary segmentation technique, which allows to significantly reduce the memory footprint and the arithmetic intensity of linear algebra kernels requiring narrow data sizes. We integrate BiSon-e into a complete System-on-Chip (SoC) based on RISC-V, synthesized and Place&Routed in 65nm and 22nm technologies, introducing a negligible 0.07% area overhead with respect to the baseline architecture. Our experimental evaluation shows that, when computing the Convolution and Fully-Connected layers of the AlexNet and VGG-16 Convolutional Neural Networks (CNNs) with 8-, 4-, and 2-bit, our solution gains up to 5.6×, 13.9× and 24× in execution time compared to the scalar implementation of a single RISC-V core, and improves the energy efficiency of string matching tasks by 5× when compared to a RISC-V-based Vector Processing Unit (VPU). Enrico Reggiani, Cristóbal Ramírez, Roger Figueras, Adrián Cristal, Mauro Olivieri, Osman S. Unsal |
ASPLOS | 5 |
| 2021 | The Italian research on HPC key technologies across EuroHPCabstractHigh-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects. Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati |
CF | 14 |
| 2019 | The international race towards Exascale in Europe
Fabrizio Gagliardi, Miquel Moretó, Mauro Olivieri, Mateo Valero |
CCF Trans. High Perform. Comput. | 3 |
| 2018 | Characterizing noise pulse effects on the power consumption of idle digital cellsabstractThe occurrence of voltage noise in digital circuits has been typically associated to logic errors. The noise exposure of nano-scale circuits, associated to process variability, makes it interesting to explore the impact of input noise voltage pulses on the static power of idle logic cells, even if the logic operation is not compromised. This work proposes a simple yet effective characterization model to characterize the resulting shift in static energy consumption. The characterization scheme allows a fast calculation of the statistical distribution of the energy shift in digital cells affected by random noise pulses, also considering process variations. The accuracy of the approach has been tested against SPICE simulation, reaching 104speedup in calculation run time. Mauro Olivieri, Usman Khalid, Antonio Mastrandrea, Francesco Menichelli |
ISCAS | 1 |
| 2015 | Message from the general chairsabstractOn behalf of the Organizing Committee, it is our pleasure to welcome you to the 20thIEEE/ACM International Symposium on Low Power Electronics and Design, 2015, (ISLPED'15), held in Rome, Italy, on July 22–24, 2015. 20 years ago, a group of visionaries noted the avid interest in low-power design in many different disciplines and recognized the need to bring these diverse groups together with the goal of information sharing, and ISLPED was born. In this year's program, we mark the 20thanniversary of this conference with special awards, a new Industry Reception dinner, and a Founders' Panel. Luca Benini, Renu Mehra, Mauro Olivieri |
ISLPED | 3 |
| 2014 | Combined Impact of NBTI Aging and Process Variations on Noise Margins of Flip-FlopsabstractThe estimation of dependable noise margins in digital cells is increasingly significant as nano-scale CMOS technology is facing true reliability issues. On one hand, a major concern comes from circuit aging mechanisms, such as NBTI, which degrade the reliability of circuit operation over time. On the other hand, variability in technology parameters results in affecting reliability. The impact of such phenomena is particularly related to the noise margins in the memory elements of a design, since a wrong stored logic value results in an upset of the system state. This work quantifies and compares the joint effect of process variations and of NBTI aging over the years on the real noise margins of several flip-flop cells. The huge amount of transistor level Monte Carlo simulations produced both nominal (i.e. average) values and associated standard deviations of the noise margins of the selected flip-flops. A possible concept of utilization of the acquired noise margin data is also reported. Usman Khalid, Antonio Mastrandrea, Mauro Olivieri |
DSD | 3 |
| 2014 | A Voltage-Based Leakage Current Calculation Scheme and its Application to Nanoscale MOSFET and FinFET Standard-Cell DesignsabstractLogic-level estimators of leakage currents, in nanoscale standard-cell-based designs, are relevant for the dramatic speed advantage with respect to analog SPICE-level simulation. We propose a novel logic-level leakage estimation model based on the characterization of voltages at the internal nodes of digital cells, in conjunction with the characterization of leakage currents in a single field-effect transistor (FET) device and with the input-dependent Kirchhoff current law expression of the total current in the cell topology. The voltage-based nature of the approach simplifies the inclusion of supply voltage variation/scaling impact, as well as of output voltage drop (loading effect), on leakage currents. The method has been implemented in hardware description language models of a complete cell library. Exhaustive tests report average accuracy below 1% error in 22-nm CMOS and 20-nm FinFET technologies, when compared with SPICE BSIM simulation results. Zia Abbas, Antonio Mastrandrea, Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Logic Drivers: A Propagation Delay Modeling Paradigm for Statistical Simulation of Standard Cell DesignsabstractIn nanoscale digital CMOS IC design, the large technology parameter variations have boosted the interest in statistical performance analysis. As the huge execution time of SPICE-based transistor-level Monte Carlo analysis is impractical for complex designs, there is a need for making accurate Monte Carlo analysis feasible through fast logic-level simulators. This paper presents a new, general logic model of digital CMOS cells featuring technology variation aware timing, and its prototype implementation in a standard hardware-description-language environment. The application of the approach to typical standard cells and test circuits shows very good agreement with SPICE BSIM4 transistor-level simulation both for nominal delay and for statistical Monte Carlo analyses. Mauro Olivieri, Antonio Mastrandrea |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Performance evaluation of Jpeg2000 implementation on VLIW cores, SIMD cores and multi-coresabstractSystem-on-chip market relies on implementing multimedia products as embedded software modules on re-usable architecture platforms. The efficient implementation of the Jpeg2000 encoder engine is still challenging HW and SW developers with its highly complex computational kernel. While several hardwired Jpeg2000 enconding modules exist, the efficient programming of Jpeg2000 on re-usable embedded high-performance cores is still an open issue. We performed an exhaustive analysis of the attainable execution speedup when specialized SW is run on different architectures built upon a multimedia-oriented VLIW processor core, demonstrating that the compression effort can be reduced by more than 50% if a SIMD-extended architecture is adopted, and by 80% when the code is optimized for a multi-core architecture. Francesco Menichelli, Mauro Olivieri, Simone Smorfa |
ISCAS | 2 |
| 2009 | Adaptive idleness distribution for non-uniform aging tolerance in MultiProcessor Systems-on-ChipabstractIn deep submicron designs of MultiProcessor Systems-on-Chip (MPSoC) architectures, uncompensated within-die process variations and aging effects will lead to an increasing uncertainty and unbalancing of expected core lifetimes. In this paper we present an adaptive workload allocation strategy for run-time compensation of variations- and aging-induced unbalanced core lifetimes by means of core activity duty cycling. The proposed techniques regulates the percentage of idle time on short-expected-life cores to meet the platform lifetime target with minimum performance degradation. Experiments have been conducted on a multiprocessor simulator of a next-generation industrial MPSoC platform for multimedia applications made of a general purpose processor and programmable accelerators. Francesco Paterna, Luca Benini, Andrea Acquaviva, Francesco Papariello, Giuseppe Desoli, Mauro Olivieri |
DATE | 6 |
| 2009 | Static Minimization of Total Energy Consumption in Memory Subsystem for Scratchpad-Based Systems-on-ChipsabstractIn VLSI systems-on-chips (SoC), leakage is expected to override 50% of the total power consumption, and the memory sub-system can be responsible for up to 75% of the power. Scratch-pad memories (SPM) are a proven alternative to cache memories in power-aware SoCs. Optimal SPM mapping has already been investigated for dynamic power reduction in the main memory and for leakage reduction in the SPM itself. This paper addresses the problem of global energy optimization (i.e., active + leakage) in the whole memory sub-system of an SPM-based SoC. We focus on SPMs dedicated to instructions and constant data. We present the technology-level foundation, the mathematical problem formulation, its solution as an integer-linear-programming (ILP) problem, the implemented design flow, and the power reduction results referring to standard benchmarks and ITRS technology data. Francesco Menichelli, Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | High-Level Side-Channel Attack Modeling and Simulation for Security-Critical Systems on ChipsabstractThe design flow of a digital cryptographic device must take into account the evaluation of its security against attacks based on side channels observation. The adoption of high level countermeasures, as well as the verification of the feasibility of new attacks, presently require the execution of time-consuming physical measurements on the prototype product or the simulation at a low abstraction level. Starting from these assumptions, we developed an exploration approach centered on high level simulation, in order to evaluate the actual implementation of a cryptographic algorithm, being it software or hardware based. The simulation is performed within a unified tool based on SystemC, that can model a software implementation running on a microprocessor-based architecture or a dedicated hardware implementation as well as mixed software-hardware implementations with cycle-accurate resolution. Here we describe the tool and provide a large set of design explorations and characterizations based on actual implementations of the AES cryptographic algorithm, demonstrating how the execution of a large set of experiments allowed by the fast simulation engine can lead to important improvements in the knowledge and the identification of the weaknesses in cryptographic algorithm implementations. Francesco Menichelli, Renato Menicocci, Mauro Olivieri, Alessandro Trifiletti |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2007 | Testing power-analysis attack susceptibility in register-transfer level designsabstractThe susceptibility of cryptographic devices to attacks based on power analysis can be both significantly and efficiently tested at early design steps. The results from a real case application show the advantages of the approach. Marco Bucci, Raimondo Luzzi, Francesco Menichelli, Renato Menicocci, Mauro Olivieri, Alessandro Trifiletti |
IET Inf. Secur. | 5 |
| 2006 | Side channel analysis resistant design flowabstractThe threat of side-channel attacks (SCA) is of crucial importance when designing systems with cryptographic hardware or software. The FP6-funded project SCARD enhances the typical micro-chip design flow in order to provide a means for designing side-channel resistant circuits and systems. Appropriate SCA-simulation tools and SCA analysis for the designer of secure systems are part of the project goals. We consider these enhancements for traditional design flows of micro-chips as necessary in order to enable the design for the next generation of secure and dependable devices. SCARD is in its final phase, the final result a SCARD chip designed by using the developed design flow is currently implemented Manfred Josef Aigner, Stefan Mangard, Francesco Menichelli, Renato Menicocci, Mauro Olivieri, Thomas Popp, Giuseppe Scotti, Alessandro Trifiletti |
ISCAS | 5 |
| 2005 | A novel yield optimization technique for digital CMOS circuits design by means of process parameters run-time estimation and body bias active controlabstractThis work presents a novel approach to optimize digital integrated circuits yield referring to speed, dynamic power and leakage power constraints. The method is based on process parameter estimation circuits and active control of body bias performed by an on-chip digital controller. The associated design flow allows us to quantitatively predict the impact of the method on the expected yield in a specific design. We present the architecture scheme, the theoretical foundation, the estimation circuits used, and two application case studies, referring to an industrial 0.13-/spl mu/m CMOS process data. The approach results to be remarkably effective at high operating temperature. In the presented case study, initial yields below 14% are improved to 86% by using a single controller and a single set of estimation circuits per die. Mauro Olivieri, Giuseppe Scotti, Alessandro Trifiletti |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | A post-compiler approach to scratchpad mapping of codeabstractScratchPad Memories (SPMs) are commonly used in embedded systems because they are more energy-efficient than caches and enable tighter application control on the memory hierarchy. Optimally mapping code and data to SPMs is, however, still a challenge. This paper proposes an optimal scratchpad mapping approach for code segments, which has the distinctive characteristic of working directly on application binaries, thus requiring no access to either the compiler or the application source code - a clear advantage for legacy or proprietary, IP-protected applications.The mapping problem is solved by means of a Dynamic Programming algorithm applied to the execution traces of the target application. The algorithm is able to find the optimal set of instructions blocks to be moved into a dedicated SPM, either minimizing energy consumption or execution times. A patching tool, which can use the output of the optimal mapper, modifies the binary of the application and moves the relevant portions of its code segments to memory locations inside of the SPM. Federico Angiolini, Francesco Menichelli, Alberto Ferrero, Luca Benini, Mauro Olivieri |
CASES | 5 |
| 2004 | A Simulation-Based Power-Aware Architecture Exploration of a Multiprocessor System-on-Chip DesignabstractWe present the design exploration of a system-on-chip architecture dedicated to the implementation of the HIPERLAN/2 communication protocol. The task was accomplished by means of an ad-hoc C++ simulation environment, integrating power models for CPUs, memories and buses used in the design and incorporating software profiling capabilities. The architecture is based on two ARM microprocessors, an AMBA bus and a local bus, DMA unit and other peripherals. Software mapping on the processor has been based on the power/performance profiling results. Francesco Menichelli, Mauro Olivieri, Luca Benini, Monica Donno, Labros Bisdounis |
DATE | 2 |
| 2004 | A Class of Code Compression Schemes for Reducing Power Consumption in Embedded Microprocessor SystemsabstractCompression of executable code in embedded microprocessor systems, used in the past mainly to reduce the memory footprint of embedded software, is gaining interest for the potential reduction in memory bus traffic and power consumption. We propose three new schemes for code compression, based on the concepts of static (using the static representation of the executable) and dynamic (using program execution traces) entropy and compare them with a state-of-the-art compression scheme, IBM's CodePack. The proposed schemes are competitive with CodePack for static footprint compression and achieve superior results for bus traffic and energy reduction. Another interesting outcome is that static compression is not directly related to bus traffic reduction, yet there is a trade off between static compression and dynamic compression, i.e., traffic reduction. Luca Benini, Francesco Menichelli, Mauro Olivieri |
IEEE Trans. Computers | 3 |
| 2004 | Bus-switch coding for reducing power dissipation in off-chip busesabstractWe present a novel coding scheme for reducing bus power dissipation. The presented approach is well suited to driving off-chip buses, where the line capacitance is a dominant factor. A distinctive feature of the technique is the dynamic reordering of bus line positions, in order to minimize the toggling activity on physical bus wires. The effectiveness of the approach is demonstrated through cycle-accurate simulation of industrial benchmarks in conjunction with post-layout evaluation of speed, power and area overhead. Mauro Olivieri, Francesco Pappalardo 0002, Giuseppe Visalli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Theoretical system-level limits of power dissipation reduction under a performance constraint in VLSI microprocessor designabstractThis paper provides a quantitative understanding of the relations among supply-voltage scaling, sustainable cycle time, pipeline depth, instruction-level parallelism, and power dissipation. Starting from simple well-established formulas, the analysis show that there is an optimal sizing of the target supply voltage and pipeline stage complexity to minimize power under a performance constraint. The verification of the model on five real processors is reported and discussed, and the application to an ideal microprocessor design or redesign is illustrated. Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Design of synchronous and asynchronous variable-latency pipelined multipliersabstractThis paper presents a novel variable-latency multiplier architecture, suitable for implementation as a self-timed multiplier core or as a fully synchronous multicycle multiplier core. The architecture combines a second-order Booth algorithm with a split carry save array pipelined organization, incorporating multiple row skipping and completion-predicting carry-select dual adder. The paper reports the architecture and logic design, CMOS circuit design and performance evaluation. In 0.35 /spl mu/m CMOS, the expected sustainable cycle time for a 32-bit synchronous implementation is 2.25 ns. Instruction level simulations estimate 54% single-cycle and 46% two-cycle operations in SPEC95 execution. Using the same CMOS process, the 32-bit asynchronous implementation is expected to reach an average 1.76 ns throughput and 3.48 ns latency in SPEC95 execution. Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Correction to "design of synchronous and asynchronous variable-latency pipelined multipliers"
Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1999 | A Low-Power Microcontroller with on-Chip Self-Tuning Digital Clock-Generator for Variable-Load ApplicationsabstractClock disabling for power management has been implemented in some microcontrollers, but the wake-up time of Xtal/PLL-based systems is incompatible with fast interrupt response. On the other hand, hardwired on-chip clocking has been used for dedicated circuits. We illustrate the design issues of a general-purpose microcontroller core with a programmable on-chip fully-digital clock generator. The CPU is compatible with the PIC16C57 instruction set and supports software-controlled clocking modes-ranging from 44 MHz up to 124 MHz; on-line self-tuning of the maximum full-speed frequency in case of peak-performance requirements; ultra-fast wake-up even with totally disabled clock generator-namely 8.6 ns. Mauro Olivieri, Alessandro Trifiletti, Alessandro De Gloria |
ICCD | 1 |
| 1996 | Statistical Carry Lookahead AddersabstractAddition techniques are divided into fixed-time and variable-time ones. While variable time techniques can achieve log/sub 2/(N) average addition time for N-bit operands, the hardware overhead have always made fixed-time adders preferable, such as Carry Lookahead and Carry Select. We present a new variable-time addition technique whose average delay is much lower than log/sub 2/(N) and whose overhead is lower than the one of a CLA adder. The new approach is made feasible by a proper application of VLSI dynamic logic design. We show the mathematical proof, the logic implementation, and the VLSI realization of the new adder. We report circuit simulation results and their comparison with the analytical model. Alessandro De Gloria, Mauro Olivieri |
IEEE Trans. Computers | 2 |
| 1996 | Hardware design of asynchronous fuzzy controllersabstractThis paper presents a hardware approach to the design of fuzzy controllers which, by exploiting some peculiar characteristics of fuzzy logic computation, allows one to save power consumption and increase computing speed. We show that the computation involved in a fuzzy controller has some statistic features that can be exploited by asynchronous computation. This paper presents a quantitative study of the statistical properties of fuzzy computation. A design methodology is introduced, and two experimental applications are shown. Alessandra Costa, Alessandro De Gloria, Mauro Olivieri |
IEEE Trans. Fuzzy Syst. | 3 |
| 1996 | An asynchronous distributed architecture model for the Boltzmann machine control mechanismabstractWe present a study addressing a hardware implementation of the Boltzmann machine that relies on the concept of asynchronous digital system. The constraint of concurrently switching only unconnected neurons is dynamically satisfied by using an asynchronous distributed control mechanism. The design of the control architecture is derived from a formal definition of the problem by means of the trace theory. Computer simulations show the efficiency of the proposed approach. Alessandro De Gloria, Mauro Olivieri |
IEEE Trans. Neural Networks | 2 |
| 1995 | Efficient semicustom micropipeline designabstractWe present the analytical model and the electrical characterization of a controllable delay component for a micropipeline architecture suitable for being designed with a semicustom design approach. An interesting feature of the component is that it is lockable, i.e., it can be controlled in an on/off fashion, permitting synchronous operation for testing purposes by means of an opportune architecture model.> Alessandro De Gloria, Mauro Olivieri |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1994 | Block placement with a Boltzmann MachineabstractThe Boltzmann Machine is a neural model based on the same principles of simulated annealing that reaches good solutions, reduces the computational requirements, and is well suited for a low-cost, massively parallel hardware implementation. In this paper we present a connectionist approach to the problem of block placement in the plane to minimize wire length, based on its formalization in terms of the Boltzmann Machine. We detail the procedure to build the Boltzmann Machine by formulating the placement problem as a constrained quadratic assignment problem and by defining an equivalent 0-1 programming problem. The key features of the proposed model are: (1) high degree of parallelism in the algorithm, (2) high quality of the results, often near-optimal, and (3) support of a large variety of constraints such as arbitrary block shape, flexible aspect ratio, and rotations/reflections. Experimental results on different problem instances show the skills of the method as an effective alternative to other deterministic and statistical techniques.> Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1993 | An analysis of dynamic scheduling techniques for symbolic applicationsabstractVLIW processors are viewed as an attractive way of achieving instruction-level parallelism because of their ability to issue multiple operations per cycle with relatively simple control logic. They are also perceived as being of limited interest as products because of the problem of object code compatibility across processors having different hardware latencies and varying levels of parallelism. The author introduces the concept of delayed split-issue and the dynamic scheduling hardware which, together, solve the compatibility problem for VLIW processors and, in fact, make it possible for such processors to use all of the interlocking and scoreboarding techniques that are known for superscalar processors.> Alessandra Costa, Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
MICRO | 4 |
| 1993 | An asynchronous approach to the RISC design of a micro-controller
Alessandra Costa, Alessandro De Gloria, Paolo Faraboschi, Giovanni Nateri, Mauro Olivieri |
Microprocess. Microprogramming | 5 |
| 1993 | A parallel architecture for the Color Doppler flow technique in ultrasound imaging
Alessandra Costa, Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
Microprocess. Microprogramming | 4 |
| 1993 | A delay insensitive approach to the VLSI design of a DRAM controller
Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
Microprocess. Microprogramming | 3 |
| 1993 | Design of a massively parallel SIMD architecture for the Boltzmann machine
Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
Microprocess. Microprogramming | 3 |
| 1993 | Delay insensitive micro-pipelined combinational logic
Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
Microprocess. Microprogramming | 3 |
| 1993 | Clustered Boltzmann Machines: Massively Parallel Architectures for Constrained Optimization Problems
Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
Parallel Comput. | 3 |
| 1993 | Efficient implementation of the Boltzmann machine algorithmabstractThe problem of optimizing the sequential algorithm for the Boltzmann machine (BM) is addressed. A solution that is based on the locality properties of the algorithm and makes possible the efficient computation of the cost difference between two configurations is presented. Since the algorithm performance depends on the number of accepted state transitions in the annealing process, a theoretical procedure is formulated to estimate the acceptance probability of a state transition. In addition, experimental data are provided on a well-known optimization problem travelling salesman problem to have a numerical verification of the theory, and to show that the proposed solution obtains a speedup between 3 and 4 in comparison with the traditional algorithm. Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
IEEE Trans. Neural Networks | 3 |
| 1992 | A non-deterministic scheduler for a software pipelining compiler
Alessandro De Gloria, Paolo Faraboschi, Mauro Olivieri |
MICRO | 3 |