Valerio Tenace

dblp:09/11301 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
5since 2021 · last 2026
0000-0001-7339-6913ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 10 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Special Day - A Design Blueprint for Scalable Multi-Agent Architectures in Complex EDA Workflows
abstract
Electronic Design Automation (EDA) workflows involve complex, tightly coupled tools and artifacts that require high reliability and traceability. Recent advances in Large Language Models (LLMs) have opened new avenues for AI-driven automation through Multi-Agent Systems (MASs), which can decompose complex tasks into manageable subtasks. However, deploying MASs in EDA remains challenging due to weak coordination, unstructured communication, and limited reproducibility. In this paper, we propose a design blueprint for scalable, LLM-based, MASs tailored to EDA workflows, emphasizing hierarchical orchestration, explicit task interfaces, tool-grounded execution, structured communication, modular memory management, and observability with recovery paths. We also introduce Nexus, an open-source Software Development Kit (SDK) that implements these principles, thus enabling low-code workflow specification and robust agent interactions. We validate our approach on open-source benchmarks, achieving 100% functional accuracy on RTL generation (VerilogEval-Human), up to 98.78% functional pass rate on HumanEval, and timing closure with 26.64% average LUT reduction and almost 30% lower total power on VTR designs.
Valerio Tenace, Pierre-Emmanuel Gaillardon
DATE1
2025 EDA-Aware RTL Generation with Large Language Models
abstract
Large Language Models (LLMs) have become increasingly popular for generating RTL code. However, producing error-free RTL code in a zero-shot setting remains highly challenging even for state-of-the-art LLMs, often leading to issues that require manual, iterative refinement. This additional debugging process can dramatically increase the verification workload, underscoring the need for robust, automated correction mechanisms to ensure code correctness from the start. In this work, we introduce AIVRIL2, a self-verifying, LLM-agnostic agentic framework aimed at enhancing RTL code generation through iterative corrections of both syntax and functional errors. Our approach leverages a collaborative multi-agent system that incorporates feedback from error logs generated by EDA tools to automatically identify and resolve design flaws. Experimental results, conducted on the VerilogEval-Human benchmark suite, demonstrate that our framework significantly improves code quality, achieving nearly a 3.4× enhancement over prior methods. In the best-case scenario, functional pass rates of 77% for Verilog and 66% for VHDL were obtained, thus substantially improving the reliability of LLM-driven RTL code generation.
Mubashir ul Islam, Humza Sami, Pierre-Emmanuel Gaillardon, Valerio Tenace
DATE4
2025 Automated Generation of Microfluidic Netlists using Large Language Models
abstract
Microfluidic devices have emerged as powerful tools in various laboratory applications, but the complexity of their design limits accessibility for many practitioners. While progress has been made in microfluidic design automation (MFDA), a practical and intuitive solution is still needed to connect microfluidic practitioners with MFDA techniques. This work introduces the first practical application of large language models (LLMs) in this context, providing a preliminary demonstration. Building on prior research in hardware description language (HDL) code generation with LLMs, we propose an initial methodology to convert natural language microfluidic device specifications into system-level structural Verilog netlists. We demonstrate the feasibility of our approach by generating structural netlists for practical benchmarks representative of typical microfluidic designs with correct functional flow and an average syntactical accuracy of $\mathbf{8 8} \boldsymbol{\%}$.
Jasper Davidson, Skylar Stockham, Allen Boston, Ashton Snelgrove, Valerio Tenace, Pierre-Emmanuel Gaillardon
VLSI-SoC5
2021 NEMO-CNN: An Efficient Near-Memory Accelerator for Convolutional Neural Networks
abstract
The relevance of Deep Learning applications has skyrocketed in the last few years, exposing key weaknesses of traditional Von Neumann hardware architectures. With high amounts of data to be fetched from memory, the efficiency of these systems gets adversely impacted by an order of magnitude for each memory hierarchy that is traversed (e.g., data cache, on- chip SRAM, and off-chip DRAM). Such an issue is even more relevant when we consider that Convolutional Neural Networks (CNNs) are composed of tens of millions of parameters that imply billions of operations per second to achieve an acceptable performance. In order to remove this so-called memory wall problem, we introduce NEMO-CNN: a high-performance hardware accelerator built around the Near-Memory Computing paradigm, i.e., a design methodology based on distributed memory blocks enhanced with nearby processing elements. Coupled with a smart mapping strategy that slices the CNN structure along its depth, our solution drastically reduces the amount of data exchanged between off- and on-chip memories by executing each slice concurrently on dedicated processing elements that only leverage local data. Experimental results using VGG-16, DarkNet-19, and TinyYOLOv2 networks demonstrate that our solution achieves a top efficiency of 60.7 FPS/W, outperforming existing CNN accelerators.
Grant Brown, Valerio Tenace, Pierre-Emmanuel Gaillardon
ASAP2
2021 Logic Synthesis Meets Machine Learning: Trading Exactness for Generalization
abstract
Logic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning.
Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee
DATE38
2020 Logic Synthesis of Pass-Gate Logic Circuits With Emerging Ambipolar Technologies
abstract
Emerging devices and new ultrascaled silicon transistors have shown disruptive electrical and functional properties that might bring digital hardware to the next level. The key issue today concerns their integration. Even though the classical complementary logic style is the most intuitive option, other strategies such as pass-transistors that were discarded in the past because they did not fit silicon MOSFETs logic should be reconsidered. Obviously, the assessment of such alternatives requires customized CAD tools and optimization engines. The objective of this paper is to introduce a synthesis and optimization flow for pass-gate logic circuits mapped onto emerging ambipolar technologies. As main contributions we propose: 1) a novel EXNOR-based decomposition technique that fully exploits do not care conditions to generate compact logic function representations and 2) a dedicated one-pass synthesis flow where optimization and technology mapping are concurrently run on a common data structure, the reduced ordered pass-diagram. Experimental results demonstrate that the proposed flow outperforms existing synthesis tools by achieving more compact circuit representations with 8.5× less devices and about 8× shallower structures (on average), while still yielding lower CPU times.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Energy-Efficient Convolutional Neural Networks via Recurrent Data Reuse
abstract
Deep learning (DL) algorithms have substantially improved in terms of accuracy and efficiency. Convolutional Neural Networks (CNNs) are now able to outperform traditional algorithms in computer vision tasks such as object classification, detection, recognition, and image segmentation. They represent an attractive solution for many embedded applications which may take advantage from machine-learning at the edge. Needless to say, the key to success lies under the availability of efficient hardware implementations which meet the stringent design constraints.Inspired by the way human brains process information, this paper presents a method that improves the processing efficiency of CNNs leveraging their repetitiveness. More specifically, we introduce (i) a clustering methodology that maximizes weights/activation reuse, and (ii) the design of a heterogeneous processing element which integrates a Floating-Point Unit (FPU) with an associative memory that manages recurrent patterns. The experimental analysis reveals that the proposed method achieves substantial energy savings with low accuracy loss, thus providing a practical design option that might find application in the growing segment of edge-computing.
Luca Mocerino, Valerio Tenace, Andrea Calimera
DATE2
2019 SAID: A Supergate-Aided Logic Synthesis Flow for Memristive Crossbars
abstract
A Memristor is a two-terminal device that can serve as a non-volatile memory element with built-in logic capabilities. Arranged in a crossbar structure, memristive arrays allow to represent complex Boolean logic functions that adhere to the logic-in-memory paradigm, where data and logic gates are glued together on the same piece of hardware. Needless to say, novel and ad-hoc CAD solutions are required to achieve practical and feasible hardware implementations. Existing techniques aim at optimal mapping strategies that account for Boolean logic functions described by means of 2-input NOR and NOT gates, thus overlooking the optimization capabilities that a smart and dedicated technology-aware logic synthesis can provide. In this paper, we introduce a novel library-free supergate-aided (SAID) logic synthesis approach with a dedicated mapping strategy tailored on MAGIC crossbars. Supergates are obtained with a Look-Up Table (LUT)-based synthesis that splits a complex logic network into smaller Boolean functions. Those functions are then mapped on the crossbar array as to minimize latency. The proposed SAID flow allows to (i) maximize supergate-level parallelism, thus reducing the total number of computing cycles, and (ii) relax mapping constraints, allowing an easy and fast mapping of Boolean functions on memristive crossbars. Experimental results obtained on several benchmarks from ISCAS'85 and IWLS'93 suites demonstrate that our solution is capable to outperform other state-of-the-art techniques in terms of speedup (3.89× in the best case), at the expense of a very low area overhead.
Valerio Tenace, Roberto Giorgio Rizzo, Debjyoti Bhattacharjee, Anupam Chattopadhyay, Andrea Calimera
DATE1
2018 Multiplication by Inference using Classification Trees: A Case-Study Analysis
abstract
Inspired by cognitive functions of the human brain, machine learning-driven synthesis flows can map Boolean functions as Classification Trees that work like statistical inference engines. Circuits of this kind infer output values by evaluating the key features of the function learned during the training stage. We propose this idea for arithmetic circuits and, more specifically, for the design of aninferential8-by-8 bit unsigned multiplier. Using as case-study an error-resilient image blending application, we quantify the most representative figures of merit, also giving comparison against a classical radix-4 multi-level implementation. Experimental results demonstrate theinferentialmultiplier guarantees 76% average accuracy, 22% less area, and 2× latency reduction that can be used for power optimization.
Roberto Giorgio Rizzo, Valerio Tenace, Andrea Calimera
ISCAS2
2018 Inferential Logic: a Machine Learning Inspired Paradigm for Combinational Circuits
abstract
Machine learning (ML) theories and tools suggest alternative forms to conceive and represent relationships among data. The same theories find their application in the Boolean domain, where logic functions can be described as inference rules. This paper introduces Inferential Logic, a novel paradigm that leverages the ML concept of statistical inference for the design of combinational logic circuits, the Inferential Logic Circuits (ILCs). This new design concept is conceived for low-power circuits that run quasi-exact computation in error-resilient applications, but it also provides an exact run-mode that can be dynamically enabled when accuracy scaling is not an option.
Valerio Tenace, Andrea Calimera
VLSI-SoC1
2018 Quasi-exact logic functions through classification trees
Valerio Tenace, Andrea Calimera
Integr.1
2016 Graphene-PLA (GPLA): a Compact and Ultra-Low Power Logic Array Architecture
abstract
The key characteristics of the next generation of ICs for wearable applications include high integration density, small area, low power consumption, high energy-efficiency, reliability and enhanced mechanical properties like stretchability and transparency. The proper mix of new materials and novel integration strategies is the enabling factor to achieve those design specifications.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI1
2016 Enabling quasi-adiabatic logic arrays for silicon and beyond-silicon technologies
abstract
Adiabatic logic aims at mimicking an adiabatic (i.e., without energy exchange) charging process in digital circuits. Although regarded as a mostly theoretical computation style, research on the topic has been constantly active over the years, providing several demonstrations of working implementations [1]. The interest in adiabatic circuits recently increased with the introduction of emerging devices, e.g., Nanoelectromechanicals switches (NEMs) [2] and graphene p-n junctions [3], which have been proven to be good technological vehicles for adiabatic computing. Despite their energy efficiency, adiabatic logic faced severe limitations in reaching large scale integration due to the difficulty in logic pipelining and the lack of CAD tools able to cope with today's design complexity.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ISCAS1
2016 Multi-function logic synthesis of silicon and beyond-silicon ultra-low power pass-gates circuits
abstract
Pass-gates logic is known to be intrinsically more energy efficient than static CMOS. This feature attracted the research interest over the years and many working implementations have been demonstrated. Recent works, in particular, have shown that pass-gates logic is well suited for ultra-low power adiabatic circuits mapped on emerging technologies. Despite the progress made, several design issues still prevent pass-gates logic circuits reaching large scale integration. In this work we deal with the lack of synthesis tools and methodologies. We propose a multi-function decomposition engine that yields (i) an efficient abstract circuit modeling through a more compact data-structure, the Multi-Function Pass Diagram (MFPD) and (ii) an effective multi-gate area/delay-driven low-power synthesis&optimization flow. Simulation results conducted on different technologies, i.e., silicon and graphene, demonstrate that logic circuits synthesized with the proposed tool are smaller in size and depth, hence less power consuming and faster than circuits obtained through conventional synthesis flows based on Binary Decision Diagrams.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
VLSI-SoC1
2015 One-pass logic synthesis for graphene-based Pass-XNOR logic circuits
abstract
Electrostatically controlled graphene P-N junctions are devices built on a single layer graphene sheet that can be turned-ON/OFF via external potential difference. Their electrical behavior resembles a CMOS transmission gate with an embedded XNOR Boolean functionality. Recent works presented an efficient design style, the Pass-XNOR logic (PXL), which allows the implementation of adiabatic logic circuits with ultra low-power features.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
DAC1
2015 Exploiting the Expressive Power of Graphene Reconfigurable Gates via Post-Synthesis Optimization
abstract
As an answer to the new electronics market demands, semiconductor industry is looking for different materials, new process technologies and alternative design solutions that can support Silicon replacement in the VLSI domain. The recent introduction of graphene, together with the option of electrostatically controlling its doping profile, has shown a possible way to implement fast and power efficient Reconfigurable Gates (RGs). Also, and this is the most important feature considered in this work, those graphene RGs show higher expressive power, i.e., they implement more complex functions, like Majority, MUX, XOR, with less area w.r.t. CMOS counterparts. Unfortunately, state-of-the-art synthesis tools, which have been customized for standard NAND/NOR CMOS gates, do not exploit the aforementioned feature of graphene RGs.
Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino, Luca G. Amarù, Giovanni De Micheli, Pierre-Emmanuel Gaillardon
ACM Great Lakes Symposium on VLSI2
2014 Pass-XNOR logic: A new logic style for P-N junction based graphene circuits
abstract
In this work we introduce a new logic style for p-n junctions based digital graphene circuits: the pass-XNOR logic style. The latter enables the realization of compact, energy efficient circuits that better exploit the characteristics of graphene. We first show how a single p-n junction can be conceived as a pass-XNOR gate, i.e., a transmission gate with embedded logic functionality, the XNOR Boolean operator. Secondly, we propose a smart integration strategy in which series/parallel connections of pass-XNOR gates allow to implement AND/OR logical conjunctions, and, therefore, all possible truth tables. Experimental results conducted on a set of representative logic functions show the superior of pass-XNOR logic circuits w.r.t. standard CMOS circuits and graphene circuits that use p-n junctions in a complementary-like structure.
Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
DATE1
2012 NBTI effects on tree-like clock distribution networks
abstract
Negative Bias Temperature Instability (NBTI) is considered one of the most critical device reliability concerns in nanometer CMOS technologies, because it causes devices to exhibit a temporal drift of performance over time.
Wei Liu 0016, Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino
ACM Great Lakes Symposium on VLSI3