VLDB 2026 Research / reviewers in the wild / expert
Massimo Alioto
dblp:48/4750 · also Massimo Bruno Alioto
· DBLP profile ↗
103ranked-venue papers
52as first author
13since 2021 · last 2026
0000-0002-4127-8258ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 99 · 51 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Smart Imager with Object Detection Exploiting Edge-Frame-Base Processing and Bounding Box Extraction for μW Power Purely-Harvested Sensor NodesabstractBattery-less and cost-sensitive vision nodes are becoming essential in IoT-scale sensor networks, where in/near-sensor AI enables local recognition while minimizing data transmission. However, achieving multi-class object detection under available peak power budgets (<10 µW) and low-cost fabrication remains a major challenge. Existing smart imagers either lack on-chip intelligence or exceed such power budgets due to costly sensing and computing. This paper presents a fully-integrated smart imager performing multi-class object detection at 8.51 μW (equivalent to the power from a 7mm × 6mm harvester at 300 lux) in standard 180nm CMOS. The system processes 1-bit edge-extracted frames, applies tile-level novelty detection for bounding-box ROI extraction, and computes CENTRIST features over cropped regions. A low-power approximate linear SVM classifies detected objects at 130 pW/pixel power. Unlike prior architectures, the proposed system maintains full image readout, supports flexible learning-based inference, and avoids custom optics and CIS processing. This makes it the first battery-less smart imager capable of flexible, multi-object detection in low-cost standard CMOS technology. Hayate Okuhara, Udari De Alwis, Liu Yue, Karim Ali 0007, Massimo Alioto |
DATE | 5 |
| 2026 | Two-Stage Inverter-Based OTA With 56-dB Gain and Digitally Reconfigurable GBW/SR for PVT-ResiliencyabstractThis work presents a novel configurable two-stage inverter-based ultra-low-power (ULP) OTA with digitally adjustable output stage for dynamic adjustments of gain-bandwidth product, (GBW) and average slew rate (${\mathbf {SR}}_{\mathbf {av}}$). This approach improves both the flexibility of the OTA compared to conventional DC-biased designs, providing a reliable solution for mitigating process, voltage, and temperature variations. The circuit is designed using standard cells in a 180-nm CMOS process and is then fully synthesizable. Experimental measurements are executed over process corners and voltage variations. The proposed OTA works at nominal 0.3-V power-supply voltage and consumes only 1.2 nW, while exhibiting 56-dB DC gain and 7.6-kHz maximum GBW with 5-pF capacitive load. Riccardo Della Sala, Marco Privitera, Giuseppe Scotti, Alfio Dario Grasso, Massimo Alioto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | 0.4-V nW-Power High-Gain Bulk-Driven Two-Stage OTA With Self-Cascode Composite Transistors and Intrinsic Current-Buffer Miller CompensationabstractThis paper presents a 0.4-V bulk-driven two-stage operational transconductance amplifier (OTA) achieving an exceptionally high voltage gain exceeding 120 dB without requiring additional bias voltages. This is accomplished using a self-cascode composite transistor configuration. Common-mode and power-supply rejection ratios ( CMRR and PSRR, respectively) exceeding 90 dB are achieved through input stage optimization. An intrinsic current-buffer Miller compensation technique is also employed to enhance frequency performance. Compared to state-of-the-art designs, the proposed OTA exhibits superior performance in terms of gain-bandwidth trade-off, settling time, PSRR, and robustness against variations. This is demonstrated through extensive measurements across corner wafers and temperature within the range of -20∘C to 100∘C (unavailable in prior art). Muhammad Omer Shah, Marco Privitera, Andrea Ballo, Massimo Alioto, Salvatore Pennisi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | CogniVision: A mW Power envelope SoC for Always-on Smart Vision in 40nmabstractA full vision system on chip with hierarchical execution from sensor to AI and communications is presented. Always-on system power is aggressively reduced to the mW level through activity reduction, gating any subsequent vision pipeline stage starting from the lowest semantic level of the scene. The system comprises an imager with dual-architecture in/near-sensor saliency detection, on the-fly novelty detection, DNN accelerator with on-chip scheduler for weight memory reduction, WiFi transmitter, wake-up receiver for cloud pushed DNN model update and software programmable orchestration via RISC-V. Average 2.1-mW power at 30 fps is achieved with a SqueezeNet V1.0 model running entirely on chip. Animesh Gupta, Japesh Vohra, Massimo Alioto |
HCS | 3 |
| 2024 | A 15-nA quiescent current capacitor-less LDO for sub-1V μW-powered fully-harvested systemsabstractThis work proposes a very-low quiescent current corner-compensated analog LDO with reconfigurable topology for multiple output voltage values, and its pW-powered voltage reference. The design of the proposed LDO has been optimized for sub-1V for battery-less systems that exhibit an average power consumption from hundreds nW up to hundreds of μW. Measurement results on a 180nm standard CMOS technology prove the effectiveness of the design strategy and validate the working principle of the LDO. The proposed analog LDO can work with a supply voltage down to 0.55-0.6V, consuming only 15-nA of quiescent current and it shows 6.4x lower FoMt (7.5x for FoMtv) compared with similar prior art. Marco Privitera, Andrea Ballo, Alfio Dario Grasso, Massimo Alioto |
ISCAS | 4 |
| 2024 | Multi-Stage Face-Voice Association Learning with Keynote Speaker DiarizationabstractThe human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as "cross-modal speaker verification''. This task poses significant challenges due to the complex relationship between the modalities. In this paper, we propose a "Multi-stage Face-voice Association Learning with Keynote Speaker Diarization''(MFV-KSD) framework. MFV-KSD contains a keynote speaker diarization front-end to effectively address the noisy speech inputs issue. To balance and enhance the intra-modal feature learning and inter-modal correlation understanding, MFV-KSD utilizes a novel three-stage training strategy. Our experimental results demonstrated robust performance, achieving the first rank in the 2024 Face-voice Association in Multilingual Environments (FAME) challenge with an overall Equal Error Rate (EER) of 19.9%. Details can be found in https://github.com/TaoRuijie/MFV-KSD. Ruijie Tao, Yidi Jiang, Duc-Tuan Truong, Chng Eng Siong, Massimo Alioto, Haizhou Li 0001 |
ACM Multimedia | 6 |
| 2023 | Capacitance-to-Digital Converter for Harvested Systems Down to 0.3 V With No Trimming, Reference, and Voltage RegulationabstractIn this work, a capacitance-to-digital converter (CDC) suitable for direct energy harvesting is introduced. The nW peak power and the ability to operate at any supply voltage in the 0.3-1.8 V range allow complete suppression of any intermediate DC-DC conversion, and hence direct supply provision from the harvester, as demonstrated with a mm-scale solar cell. The proposed CDC architecture eliminates the need for any additional support circuitry, preserving true nW-power operation, and reducing design and integration effort. In detail, the architecture is based on a pair of double-swappable oscillators, and avoids the need for any voltage/current/frequency reference circuit in the oscillator mismatch compensation. The digital and differential nature of the architecture counteracts the effect of process/voltage/temperature variations. A load-agnostic one-time self-calibration scheme compensates mismatch, and can be run from boot to run stage of the chip lifecycle. The proposed self-calibration scheme suppresses any trimming or testing time for low-cost systems, and avoids any input capacitance disconnection requirement. A 180-nm testchip shows 7-bit ENOB down to 0.3 V and 1.37-nW total power, when powered by a 1-mm2 indoor solar cell down to 10 lux (i.e., late twilight). Orazio Aiello, Paolo Crovetti, Massimo Alioto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Opening of the 2023 Editorial Year - This Coda as Prelude of Next TVLSI Cycle With Sustained GrowthabstractThe journey as Editor-in-Chief (EiC) of the IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) from 2019 to 2022 has been exhilarating. Continuing the analogy with classical music concertos in recent editorials in this journal[1],[2],[3], some movements have been somewhat tempestuous in a rapidly changing world under a pandemic. Simultaneously, fundamental transformations have taken place in the semiconductor supply chain, as well as in our life and work style through increasingly distributed cooperation models (across continents, and from the office to the home office). Other movements have been highly rewarding thanks to the concerted effort (analogy intended) of a talented and rock-solid Editorial Board, whose precious contribution has led to relentless journal improvements in many respects. And, indeed, it has been a real pleasure to work with each and every one of the members of the TVLSI Editorial Board, whose commitment to excellence has brought the journal to new heights. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Editorial Opening of the 2022 TVLSI Editorial Year - Connecting Trends From Society to VLSI SystemsabstractThe past year of 2021 has consolidated several societal courses observed in recent years, marking an unprecedented acceleration in a number of trends that have now reached their tipping point. The accelerated digitalization of human activities and outcomes has made the human side of supply chains more distributed on one hand[1]while putting an unprecedented pressure on its logistics side and mandating fundamental rethinking of its resilience–efficiency balance[2]. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | STT-MRAM Architecture with Parallel Accumulator for In-Memory Binary Neural NetworksabstractIn this paper, a row-wise XNOR accumulator architecture for STT-MRAM arrays is proposed for parallel and efficient multiply-and-accumulate (MAC) operation. The proposed accumulator supports in-memory computing and binary neural network (BNN) applications. In the proposed architecture, inputs are fed from the complementary bitlines, whereas readout is performed through a time-based sense amplifier (TBS). The proposed architecture that does not require any ADC can exhibit an average error rate of 0.085 for XNOR vector size (i.e., accumulate capacity) of 128 bits, which translates into 98.45% classification accuracy of a multi-layer perceptron (MLP) on the MNIST dataset. Thi-Nhan Pham, Quang-Kien Trinh, Ik Joon Chang, Massimo Alioto |
ISCAS | 4 |
| 2021 | Design of Digital OTAs With Operation Down to 0.3 V and nW Power for Direct HarvestingabstractIn this paper, passive-less fully-digital operational transconductance amplifiers (DIGOTA) for energy- and area-constrained systems are modeled and analyzed from a design viewpoint. The digital behavior of DIGOTAs is modeled as an equivalent small-signal differential-mode circuit with zero bias current, and a common-mode feedback loop operating as a self-oscillating threshold sampler. Such continuous-time equivalent circuits are used to derive an explicit model of the main performance parameters that are generally adopted to characterize OTAs. This provides an insight into circuit operation and allows to derive practical guidelines to achieve a given design target. Among the others, an explicit model is derived for the DC gain, the frequency response, the gain-bandwidth product, the input-referred noise, and the input offset voltage. The models are validated via direct comparison with multi-die measurement results in CMOS 180 nm. From an application viewpoint, the voltage (power) reduction down to 0.25 V (sub-nW) uniquely enable direct harvesting (e.g., with solar cells), suppressing any intermediate DC-DC conversion stage. This further enhances the area efficiency advantage of DIGOTA stemming from its fully-digital nature, making it well suited for cost-sensitive and purely-harvested systems. Pedro Toledo, Paolo Crovetti, Orazio Aiello, Massimo Alioto |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Opening of the 2021 Editorial Year - Overture for a New Year of ChangeabstractThe year 2021 is expected to be a turning point, after a challenging year that has unexpectedly impacted our lives and interactions. In 2020, technological and scientific progress in the area of VLSI design has been relentless despite the difficulties that the year has brought. This has been made possible by the ability to adapt to the temporary “new normal” of our community of researchers, designers, and innovators in the area of VLSI systems. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Second Quarter of the 2021 Editorial Year - A Year in CrescendoabstractThe year 2021 is unfolding to be a turning point as anticipated at its beginning [item 1) in the Appendix], as well as at the end of a globally challenging year in many respects [item 2) in the Appendix]. Continuing the analogy with classical music started in recent editorials in this journal as a common thread [item 1) in the Appendix], [item 2) in the Appendix], we are playing in crescendo with several new activities and publications going stronger (and louder) than ever. This is really being a remarkable opera that is being made possible by the continued excitement and the relentlessness of our community. Our community as a true orchestra indeed, whose scientific and technical contributions have really enabled the many tools that we are all using to move forward as a society, despite the drastic changes that the year 2020 has brought so far. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Automated Design of Reconfigurable Microarchitectures for Accelerators under Wide-Voltage ScalingabstractThis article introduces a systematic methodology to design microarchitectures that are reconfigurable down to the pipeline stage. Reconfigurable microarchitectures were showed to provide significant energy improvements in accelerators under wide-voltage scaling. However, prior art is based on ad hoc techniques that limit their applicability, without addressing the challenge of enabling general design flows for reconfigurable microarchitectures. The proposed methodology introduces the unprecedented capability of translating a conventional fixed microarchitecture into a reconfigurable one. The methodology relies on commercial EDA tools, which are integrated into a design flow through the manipulation of the gate-level netlist via a set of graph algorithms. The proposed methodology is shown to be architecture-agnostic, fully automated, and applicable to designs that are either developed at the register transfer level (RTL), or provided by third-party soft IP vendors. Ultimately, the proposed methodology allows to add microarchitectural adjustment as a run-time knob to augment the energy benefits of wide-voltage scaling. Reconfiguration is shown to improve the energy efficiency by up to 35% beyond the conventional dynamic voltage frequency scaling (DVFS), through the analysis of various test vehicles. Longyang Lin, Massimo Alioto |
ISCAS | 3 |
| 2020 | Deep Sub-pJ/Bit Low-Area Energy-Security Scalable SIMON Crypto-Core in 40 nmabstractThis paper describes an energy-security scalable crypto-core for private-key cryptography in low-end sensor nodes based on SIMON cipher. Energy and area footprints are reduced through techniques at the algorithm, microarchitectural and gate level. At the algorithm level, multiple encryption is introduced to dynamically expand the key length from 64 to 256 bits at minimal reconfiguration complexity, allowing shorter keys and lower energy when lower level of security is demanded. The 1-round parallel architecture with short 32-bit datapath narrows the traditionally large gap between the conventional crypto-core data (e.g., 128 bits) and low-end processors (32 bit or lower), mitigating or eliminating the need for area- and energy-hungry FIFO buffers at their interface. At the gate level, the adoption of multi-bit pulsed latch-based pipelines with internal clock buffer sharing reduces the dominant clock power/area contribution of sequential elements. Tunable clock duty cycle allows time borrowing for improved variation resilience at ultra-low voltages at low hold-fix buffer count, leveraging the inherent margin against hold time violations enabled by relatively uniform pipestage delays. A 40 nm testchip shows energy down to 0.31 pJ/bit at 0.45 V with 64-bit key and 0.79E6 F2area (F = process minimum feature size). The proposed crypto-core is well suited for ubiquitous security in energy/area-constrained platforms (e.g., low-end sensor nodes, RFIDs), while preserving full 256-bit security when necessary. Sachin Taneja, Massimo Alioto |
ISCAS | 2 |
| 2020 | Editorial on the Opening of the New Editorial Year - The State of the IEEE Transactions on Very Large Scale Integration (VLSI) SystemsabstractIt is a pleasure to write this editorial celebrating the start of the editorial year of 2020, which marks the beginning of my second year as the Editor-in-Chief of the IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI). In the past year, our Editorial Board has enabled several new initiatives and has achieved several accomplishments, solidly supporting the trajectory that we had envisioned for our journal a year ago [item 1) of the Appendix]. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Editorial on the Conclusion of the 2020 Editorial Year - The Climactic Finale of a Peculiar Year
Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | Automated Design of Reconfigurable Microarchitectures for Accelerators Under Wide-Voltage ScalingabstractThis article introduces a systematic methodology to design microarchitectures that are reconfigurable down to the pipeline stage. Reconfigurable microarchitectures were showed to provide significant energy improvements in accelerators under wide-voltage scaling. However, prior art is based on ad hoc techniques that limit their applicability, without addressing the challenge of enabling general design flows for reconfigurable microarchitectures. The proposed methodology introduces the unprecedented capability of translating a conventional fixed microarchitecture into a reconfigurable one. The methodology relies on commercial EDA tools, which are integrated into a design flow through the manipulation of the gate-level netlist via a set of graph algorithms. The proposed methodology is shown to be architecture-agnostic, fully automated, and applicable to designs that are either developed at the register transfer level (RTL), or provided by third-party soft IP vendors. Ultimately, the proposed methodology allows to add microarchitectural adjustment as a run-time knob to augment the energy benefits of wide-voltage scaling. Reconfiguration is shown to improve the energy efficiency by up to 35% beyond the conventional dynamic voltage frequency scaling (DVFS), through the analysis of various test vehicles. Longyang Lin, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Wake-Up Oscillators with pW Power Consumption in Dynamic Leakage Suppression LogicabstractIn this paper, two circuit topologies of pW-power Hz-range wake-up oscillators for sensor node applications are presented. The proposed circuits are based on standard cells utilizing the Dynamic Leakage Suppression logic style [4]-[5]. The proposed oscillators exhibit low supply voltage sensitivity over a wide supply voltage range, from nominal voltage down to the deep sub-threshold region (i.e., 0.3 V). This enables direct powering from energy harvesters or batteries through their whole discharge cycle, suppressing the need for voltage regulation. Post-layout time-domain simulations of the proposed oscillators in 180nm show a power consumption of 1.4-1.7pW, a supply-sensitivity of 55-40%/V over the 0.3V-1.8V supply voltage range, and a compact area down to 1,500μm2. The very low power consumption makes the proposed circuits very well suited for energy-harvested systems-on-chip for Internet of Things applications. Orazio Aiello, Paolo Crovetti, Massimo Alioto |
ISCAS | 3 |
| 2019 | Token-Based Security for the Internet of Things With Dynamic Energy-Quality TradeoffabstractIn this paper, token-based security protocols with dynamic energy-security level tradeoff for Internet of Things (IoT) devices are explored. To assure scalability in the mechanism to authenticate devices in large-sized networks, the proposed protocol is based on the OAuth 2.0 framework, and on secrets generated by on-chip physically unclonable functions. This eliminates the need to share the credentials of the protected resource (e.g., server) with all connected devices, thus overcoming the weaknesses of conventional client-server authentication. To reduce the energy consumption associated with secure data transfers, dynamic energy-quality tradeoff is introduced to save energy when lower security level (or, equivalently, quality in the security subsystem) is acceptable. Energy-quality scaling is introduced at several levels of abstraction, from the individual components in the security subsystem to the network protocol level. The analysis on an MICA 2 mote platform shows that the proposed scheme is robust against different types of attacks and reduces the energy consumption of IoT devices by up to 69% for authentication and authorization, and up to 45% during data transfer, compared to a conventional IoT device with fixed key size. Muhammad Naveed Aman, Sachin Taneja, Biplab Sikdar 0001, Kee Chaing Chua, Massimo Alioto |
IEEE Internet Things J. | 5 |
| 2019 | Editorial: TVLSI Keynote Papers Enriching Our Transactions With Invited ContributionsabstractSince the beginning of this editorial year, our Transactions has evolved in the numerous directions that were described in the first editorial of my term as the IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) Editor-in-Chief [item 1) in the Appendix]. The new initiatives on TVLSI aim to emphasize the positioning of our Transactions as crossroads of our broad and excitingly diverse community and to remark its distinctive systems nature. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Energy-Quality Scalable Adders Based on Nonzeroing Bit TruncationabstractApproximate addition is a technique to trade off energy consumption and output quality in error-tolerant applications. In prior art, bit truncation has been explored as a lever to dynamically trade off energy and quality. In this brief, an innovative bit truncation strategy is proposed to achieve more graceful quality degradation compared to state-of-the-art truncation schemes. This translates into energy reduction at a given quality target. When applied to a ripple-carry adder, the proposed bit truncation approach improves quality by up to 8.5 dB in terms of peak signal-to-noise ratio, compared to traditional bit truncation. As a case study, the proposed approach was applied to a discrete cosine transform engine. In comparison with prior art, the proposed approach reduces energy by 20%, at insignificant delay and silicon area overhead. Fabio Frustaci, Stefania Perri, Pasquale Corsonello, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Energy-performance design exploration of a low-power microprogrammed deep-learning acceleratorabstractThis paper presents the design space exploration of a novel microprogrammable accelerator in which PEs are connected with a Network-on-Chip and benefit from low-power features enabled through a practical implementation of a Dual-Vddassignment scheme. An analytical model, fitted with postlayout data obtained with a 28nm FDSOI design kit, returns implementations with optimal energy-performance tradeoff by taking into consideration all the key design-space variables. The obtained Pareto analysis helps us infer optimization rules aimed at improving quality of design. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
DATE | 5 |
| 2018 | Project-Based Learning in Digital Fundamentals Course Using FPGAsabstractAs embedded systems become integral in both academia and industry, hardware description languages are becoming a requisite part of Electrical and Computer Engineering curricula. The implementation of project-based learning in a freshman Digital Fundamentals course using an FPGA platform allows for greater engagement and learning through interesting and real-world projects. Through this approach, students strengthen the conceptual links between hardware description languages and circuit hardware while developing the ability to manage complex systems with multiple levels of abstraction. This paper discusses the implementation of three different application-driven projects on the Basys-3 trainer FPGA platform at the National University of Singapore and evaluates the effects of the key pedagogical differences. Both qualitative and quantitative assessment results are presented, showing how students perceived these projects and met the learning objectives. Dingjuan Chua, Jieyi Gao, Massimo Alioto, Yong Ping Xu, Sangit Sasidhar |
FIE | 3 |
| 2018 | Fully Synthesizable, Rail-to-Rail Dynamic Voltage Comparator for Operation down to 0.3 VabstractA novel rail-to-rail dynamic voltage comparator is presented in this paper. The proposed circuit is fully synthesizable, as it can be designed with automated digital design flows and standard cells, and can operate at very low voltages down to deep sub-threshold. Post-layout simulations show correct operation for rail-to-rail common-mode inputs at a supply voltageVDDdown to 0.3 V. At such voltage, the input offset voltage standard deviation is less than 28 mV (8 mV) over the rail-to-rail common-mode input range (aroundVDD/2). The digital nature of the comparator and its ability to operate down to deep sub-threshold voltages allow its full integration with standard-cell digital circuits in terms of both design and voltage domain. The ease of design, the low area and the voltage scalability make the proposed comparator very well suited for sensor nodes, integrated circuits for the Internet of Things and related applications. Orazio Aiello, Paolo Crovetti, Massimo Alioto |
ISCAS | 3 |
| 2018 | Novel Time-Based Sensing Scheme for STT-MRAMsabstractThis paper proposes a novel STT-MRAM sensing scheme based on time-based sensing (TBS). The TBS scheme converts the bitline voltage into time, then sensing is performed in the time domain rather than in conventional voltage or current domain. Monte Carlo simulations in 65nm show that the proposed TBS improves the read bit error rate (BER) by two-three orders of magnitude, compared to conventional sensing circuit. This is achieved at the cost of less than 1% area penalty and 13-14% performance degradation, and an insignificant (2%) energy penalty when designed at iso-area (minimum delay). As further advantage, TBS requires no analog reference generation and distribution by leveraging gate delay as an implicit timing reference. Quang-Kien Trinh, Sergio Ruocco, Massimo Alioto |
ISCAS | 3 |
| 2018 | Design-Space Exploration of Pareto-Optimal Architectures for Deep Learning with DVFSabstractSpecialized computing engines are required to accelerate the execution of Deep Learning (DL) algorithms in an energy-efficient way. To adapt the processing throughput of these accelerators to the workload requirements while saving power, Dynamic Voltage and Frequency Scaling (DVFS) seems the natural solution. However, DL workloads need to frequently access the off-chip memory, which tends to make the performance of these accelerators memory-bound rather than computation-bound, hence reducing the effectiveness of DVFS. In this work we use a performance-power analytical model fitted on a parametrized implementation of a DL accelerator in a 28-nm FDSOI technology to explore a large design space and to obtain the Pareto points that maximize the effectiveness of DVFS in the sub-space of throughput and energy efficiency. In our model we consider the impact on performance and power of the off-chip memory using real data of a commercial low-power DRAM. Giulia Santoro, Mario R. Casu, Valentino Peluso, Andrea Calimera, Massimo Alioto |
ISCAS | 5 |
| 2018 | Guest Editorial Special Issue on Selected Papers from PRIME 2017 and SMACD 2017
Giulia Di Capua, Nuno Horta, Francisco V. Fernández 0001, Günhan Dündar, Salvatore Pennisi, Gaetano Palumbo, Massimo Alioto, Gianluca Giustolisi |
Integr. | 7 |
| 2017 | Energy-quality scalable adaptive VLSI circuits and systems beyond approximate computingabstractIn this paper, the concept of energy-quality (EQ) scalable systems is introduced and explored, as novel design dimension to scale down energy in integrated systems for the Internet of Things (IoT). EQ-scalable systems explicitly trade off energy and quality at different levels of abstraction (“vertically”), and sub-systems (“horizontally”), creating new opportunities to improve energy efficiency for a given task and expected “quality”. The concept of quality slack, a taxonomy of techniques to trade off energy and quality and a general EQ-scalable architecture are presented. The generality of the EQ-scaling concept is shown through several examples, ranging from logic to analog circuits, to memories and Analog-Digital Converters. Challenges, opportunities and expected energy gains are discussed to gain an understanding of the potential of the EQ-scalable integrated circuits and systems. As a result, EQ scalable systems are expected to substantially improve the energy efficiency of systems for IoT, compensating the limited energy gains that will be offered by technology and voltage scaling. Massimo Alioto |
DATE | 1 |
| 2017 | Design-oriented models for quick estimation of path delay variability via the fan-out-of-4 metricabstractIn this paper, a novel modeling framework is proposed to quickly estimate the delay variability of logic paths due to random variations, and evaluate the related design margin. The analysis shows that the popular fan-out-of-4 metric F04 can capture the impact of technology and voltage on the delay variations of logic paths. Once those contributions are isolated, the impact of random variations on standard cells' delay is accounted for by means of cell-specific coefficients that are evaluated in a preliminary library characterization phase. The proposed framework is very general and applicable from sub-threshold to nominal voltage, and provides the designer with a deep insight into the main delay variability contributions in a path. It also predicts the impact of design modifications (e.g., logic restructuring, cell up-sizing), and is well suited for pencil-and-paper calculations. Case studies involving three critical paths extracted from designs ranging from microprocessors to specialized hardware show adequate accuracy, with a delay variability error being typically less than 10%. Massimo Alioto, Giuseppe Scotti, Alessandro Trifiletti |
ISCAS | 1 |
| 2017 | Power-precision scalable latch memoriesabstractApproximate computing leverages the inherent error resiliency present in many applications to improve circuits performance. Precision-scalable systems dynamically introduce approximations to trade off power and quality, based on the application under execution and the incoming dataset. In this paper, this principle is explored for the first time in the context of latch memories by introducing the ability to scale their precision, while retaining the ability to synthesize them in an automated manner. This offers additional opportunities to reduce energy, compared to the well-known suitability for aggressive voltage scaling of latch memories. A case study based on image processing applications is presented to evaluate the quality-power trade-off in 40nm CMOS. The analysis shows that the total power is reduced by up to 56% when the precision requirement is relaxed. Darjn Esposito, Antonio G. M. Strollo, Massimo Alioto |
ISCAS | 3 |
| 2017 | Transistor sizing strategy for simultaneous energy-delay optimization in CMOS buffersabstractIn this work, a systematic transistor sizing strategy is proposed to meet an arbitrary energy-delay target in CMOS buffers, as defined by the considered applications. This is particularly important in VLSI systems, as buffers driving large capacitive loads consume a very large energy compared to other logic gates. To this aim, an analytical and technology-independent model was first developed to find optimal circuit design parameters (e.g., sizing, number of stages). To assure true optimality, general Variable-stage effort Tapered Buffers (VTB) are considered, as opposed to Fixed-stage effort Tapered Buffers (FTB). Results show that optimized VTBs reduce energy by as much as 30%, compared to FTBs. Under balanced energy and delay, VTBs achieve 10-20% energy saving with nearly the same performance as FTBs. The adopted models and design guidelines are shown to agree well with circuit simulations in 28 and 65nm CMOS across the voltage range from 0.6 V to 1 V. This design strategy is a useful tool for circuit designers to systematically manage the energy-delay tradeoff of CMOS buffers in a simple and technology-agnostic manner. Longyang Lin, Quang-Kien Trinh, Massimo Alioto |
ISCAS | 3 |
| 2017 | A variation-aware simulation framework for hybrid CMOS/spintronic circuitsabstractIn this paper, a variation-aware simulation framework is introduced for hybrid circuits comprising MOS transistors and spintronic devices (e.g., magnetic tunnel junction-MTJ). The simulation framework is based on one-time characterization via micromagnetic multi-domain simulations, as opposed to most of existing frameworks based on single-domain analysis. As further distinctive capability, stochastic variations of the MTJ switching are explicitly incorporated through a Skew Normal distribution, which is adjusted to fit micromagnetic simulations. The framework is implemented in the form of Verilog-A look-up table based model, which assures easy integration with commercial circuit design tools, and very low computational effort. The framework is applied to non-volatile Flip-FIops as case study with 10,000 Monte Carlo runs. Raffaele De Rose, Marco Lanuzza, Felice Crupi, Giulio Siracusano, Riccardo Tomasello, Giovanni Finocchio, Mario Carpentieri, Massimo Alioto |
ISCAS | 8 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | EditorialabstractWe are pleased to announce this year’s winners of the IEEE Transactions on Very Large Scale integrated (VLSI) Systems (TVLSI) Circuits and Systems (CAS) Society Best Reviewer and Associate Editor Awards. We had a number of qualified candidates for each award, but after a thorough evaluation process including input from the Associate Editors and Selection Committee, these five individuals stood out among all the candidates: Krishnendu Chakrabarty, Massimo Alioto, Rajiv V. Joshi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Boosted sensing for enhanced read stability in STT-MRAMsabstractRead access in STT-MRAMs is well known to be highly sensitive to process variations. Such variations are responsible for read bit error rates that are worse than conventional CMOS memories (e.g., SRAM) by orders of magnitude, especially at low voltages. In this work we propose a boosted sensing scheme to improve the resiliency of STT-MRAM against variations in read accesses based on the voltage sensing scheme. The proposed approach permits to improve the bit error rate by two orders of magnitude at iso-area with 25% energy overhead and 50% performance degradation, compared to conventional voltage sensing circuit. At the same time, the proposed sensing scheme is able to operate at low voltages (0.7 V) with no appreciable degradation in the sensing margin. Quang-Kien Trinh, Sergio Ruocco, Massimo Alioto |
ISCAS | 3 |
| 2016 | STT-MRAM write energy minimization via area optimization under dynamic voltage ScalingabstractIn this paper we show that the area optimization of STT-MRAM bitcells can deliver a substantial reduction in the energy per write access when dynamic voltage scaling (DVS) is adopted. Indeed, the increase in the bitcell area enables the reduction in the write energy consumed by the bitcells at the expense of the energy of peripheral circuits, when lowering the supply voltage. The proposed approach addresses a fundamental challenge in STT-MRAM arrays, whose energy efficiency is well-known to be severely degraded by the large energy cost of write. Simulations of STT-MRAM arrays built with four widely adopted bitcells showed that up to 2.5X write energy reduction and 5X performance improvement are obtained at 2X bitcell area cost at 0.7 V, compared to minimum-sized bitcell. Quang-Kien Trinh, Sergio Ruocco, Massimo Alioto |
ISCAS | 3 |
| 2016 | Ultra-Fine Grain Vdd-Hopping for energy-efficient Multi-Processor SoCsabstractThis paper introduces Ultra-Fine Grain Vdd-Hopping (FINE-VH), an extension of Dynamic Voltage-Frequency Scaling (DVFS) for energy efficient Multi-Processor SoCs (MPSoCs). The proposed technique leverages the working principle of Vdd-Hopping applied at ultra-fine granularity, i.e., within the core, by means of a layout-assisted, level-shifter free, dynamic dual-Vdd control strategy where leakage currents are minimized through an optimal timing-driven poly-bias assignment procedure. Valentino Peluso, Andrea Calimera, Enrico Macii, Massimo Alioto |
VLSI-SoC | 4 |
| 2016 | Editorial First TVLSI Best AE and Reviewer AwardsabstractWe are pleased to announce the first IEEE Transactions on VLSI Systems Circuits and Systems (CAS) Society Best Reviewer and Associate Editor (AE) Awards. With the support of the IEEE CAS Society, we have been able to establish an award that recognizes the contributions of our top AEs and Reviewers from both academia and industry who are CAS members. Each year we would like to award two Best AE and three Best Reviewer Awards to those whose efforts and contributions help us to meet our mission of performing an expeditious selection of very high-quality submissions. Krishnendu Chakrabarty, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Approximate SRAMs With Dynamic Energy-Quality ManagementabstractIn this paper, approximate SRAMs are explored in the context of error-tolerant applications, in which energy is saved at the cost of the occurrence of read/write errors (i.e., signal quality degradation). This analysis investigates variation-resilient techniques that enable dynamic management of the energy-quality tradeoff down to the bit level. In these techniques, the different impacts of errors on quality at different bit positions are explicitly considered as key enabler of energy savings that are far larger than a simple voltage scaling. The analysis is based on the experimental results in an energy-quality scalable 28-nm SRAM and the extrapolation to a wide range of conditions through the models that combine the individual energy contributions. Results show that the joint adoption of multiple bit-level techniques provides substantially larger energy gains than individual techniques. Compared with the simple voltage scaling at isoquality, the joint adoption of these techniques can provide more than $2\times $ energy reduction at negligible area penalty. Energy savings turn out to be highly sensitive to the choice of joint techniques, thus showing the crucial importance of dynamic energy-quality management in approximate SRAMs. Fabio Frustaci, David T. Blaauw, Dennis Sylvester, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Modeling the impact of dynamic voltage scaling on 1T-1J STT-RAM write energy and performanceabstractThis paper investigates the impact of voltage scaling on the energy and the performance of STT-RAM bitcells during write operation. Analytical models of energy scaling and performance degradation are derived to gain an insight into the energy-performance tradeoff at low voltages. Minimum-energy operation is explored through optimization of the supply voltage, with energy savings in the order of 20%. Comparison of single access transistor STT-RAM bitcells shows that Standard Connection (SC) topology achieves approximately the same minimum energy as the Reversed Connection (RC) bitcell, but achieves 20% better performance at the minimum energy point. Quang-Kien Trinh, Sergio Ruocco, Massimo Alioto |
ISCAS | 3 |
| 2015 | Jitter analysis and measurement in subthreshold source-coupled differential ring oscillatorsabstractThe jitter and the phase noise of ring oscillators utilizing subthreshold source-coupled logic (STSCL) style are analyzed in this paper. Closed-form equations are derived to predict the jitter and phase noise caused by white and flicker noise. Measurement results of a test chip fabricated in a standard CMOS 90 nm technology are presented to validate these expressions. The performed analysis shows that jitter in STSCL-based ring oscillator is independent of technology parameters, as opposed to its CMOS counterparts that depend on supply voltage and parameters of technology. Based on measured results, noise on current control line can dominate the total jitter of the oscillator. Design guidelines are proposed to limit the jitter effect of ring oscillators using STSCL logic. The proposed STSCL-based ring oscillator achieves an average RMS jitter as low as 0.24 % of the oscillation period at a 1.08 MHz/μA energy efficiency, which demonstrates its suitability for ultra-low-power applications. Mahsa Shoaran, Armin Tajalli, Massimo Alioto, Yusuf Leblebici |
ISCAS | 3 |
| 2015 | AES architectures for minimum-energy operation and silicon demonstration in 65nm with lowest energy per encryptionabstractLightweight encryption circuits are crucial to ensure adequate information security in emerging millimeter-scale platforms for the Internet of Things, which are required to deliver moderately high throughput under stringent area and energy budgets. This requires the adoption of specialized AES accelerators, as they offer orders of magnitude energy improvements over microcontroller-based implementations. In this paper, we present the architectural exploration of lightweight AES accelerators with the goal of minimizing the energy consumption. Also, the lower bound of the number of cycles per encryption in lightweight AES designs is estimated as a function of the number of available S-boxes. Combined with sub-/near-threshold circuit techniques, we present a low-cost ultra energy-efficient AES encryption core for cubic-millimeter platforms. Our test chip achieves high energy efficiency of 0.83 pJ/bit at 0.32 V, which outperforms the state-of-the-art low-cost AES designs by 7×. Yajun Ha, Massimo Alioto |
ISCAS | 3 |
| 2015 | Novel Self-Body-Biasing and Statistical Design for Near-Threshold Circuits With Ultra Energy-Efficient AES as Case StudyabstractNear-threshold operation enables high energy efficiency, but requires proper design techniques to deal with performance loss and increased sensitivity to process variations. In this paper, we address both issues with two synergistic approaches. First, we introduce a novel body-biasing technique to mitigate the performance loss at near-threshold voltages while not requiring any additional circuitry for the body-bias control, thereby minimizing the design effort and simplifying the systems-on-chip integration. Second, we introduce a novel statistical design methodology to efficiently and accurately evaluate the design guardband strictly needed in the worst case, thereby keeping the area cost of variations at its very minimum. A 65-nm advanced encryption standard testchip demonstrates 1.65× throughput improvement over a baseline design without body biasing, and enables reliable operation over a wide voltage range (0.5-1.2 V) as opposed to traditional body-biasing schemes. In addition, our testchip achieves 1.63× area efficiency improvement compared with a design based on corner analysis. Accordingly, the proposed techniques are well suited for the design of near-threshold specialized hardware with improved performance, reduced silicon area, and design effort. Yajun Ha, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Ultra-low power design approaches for IoTabstractPresents a collection of slides covering the following topics: ultra-low power design approaches for Internet of Things (IoT); the special features of IoT; ultra-low voltage operations; and key design issues, including performance and output, leakage, and system variations and resiliency. Massimo Alioto |
Hot Chips Symposium | 1 |
| 2014 | Tunnel FETs for Ultra-Low Voltage Digital VLSI Circuits: Part II-Evaluation at Circuit Level and Design PerspectivesabstractIn Part II of this paper, the potential of tunnel FETs (TFETs) for ultra-low voltage (ULV)/ultra-low power (ULP) operation at 32-nm node is investigated through Verilog-A simulations of appropriate reference circuits. Critical issues arising at ultra-low voltages are analyzed, including static robustness of TFET logic gates, performance degradation, and sensitivity to process variations. Guidelines to design ultra-low energy standard cell libraries are derived. The minimum energy point is analyzed in a wide range of conditions, and guidelines for microarchitectural optimization for ultra-low energy are introduced. Voltage scalability of static RAM memories is also analyzed as main limitation to aggressive voltage scaling of very large scale integration (VLSI) systems, and improved precharge schemes are introduced to reduce leakage. The impact of variations of the main device parameters on VLSI digital circuits is investigated to identify the most critical variations that need to be controlled at process level. This investigation permits to understand the potential of TFETs and their advantages over traditional devices within a unitary framework that is based on fair design and comparison from device to circuit level, as well as to develop clear design perspectives in the context of ULV/ULP VLSI digital circuits. Massimo Alioto, David Esseni |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Novel Class of Energy-Efficient Very High-Speed Conditional Push-Pull Pulsed LatchesabstractIn this paper, a new class of pulsed latches is introduced and experimentally assessed in 65-nm CMOS. Its conditional push-pull pulsed latch topology is based on a push- pull final stage driven by two split paths with a conditional pulse generator. Two circuit implementations of the concept are discussed, with their main difference being in the pulse generator, which can be either shared (CSP3L) or not (CP3L). Measurements show that the proposed topology is very fast, as it outperforms the well-known transmission gate pulsed latch (TGPL) [1] by 1.5×-2×; hence the proposed pulsed latch has the highest performance ever reported. The proposed pulsed latch is also shown to significantly improve the energy efficiency compared to the state of the art. Indeed, a 2.3× improvement in ED3 product (energy × delay3) over TGPL was found for designs targeting minimum ED3. For designs targeting minimum ED, a 1.3× improvement was found in ED product. This comes at the cost of a 1.15×-1.35× cell area penalty, which translates into an overall area increase well below 1% in typical systems. Measurements on 256 replicas confirm that the above benefits are kept in the presence of variations. Accordingly, the proposed class of pulsed latches goes beyond the current state of the art and is well suited for VLSI systems that require both high performance and energy efficiency. Elio Consoli, Gaetano Palumbo, Jan M. Rabaey, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | Tunnel FETs for Ultralow Voltage Digital VLSI Circuits: Part I - Device-Circuit Interaction and Evaluation at Device LevelabstractThis paper and the companion work present the results of a comparative study between the tunnel-FETs (TFETs) and conventional MOSFETs for ultralow power digital circuits targeting a VDDbelow 500 mV. For this purpose, we employed numerical TCAD simulations, as well as mixed device-circuit and lookup-table simulations using either the SENTAURUS or the Verilog-A environment. In particular, in this paper, we explore the device-circuit interaction in n- and p-type TFETs, and propose a design leading to a good tradeoff between the current leakage and transistor imbalance at ultralow VDD, as required in ultralow voltage systems. Then, we systematically compare the IOFF, ION, effective capacitance, OFF-state and ON-state stacking factors for TFETs, SOI, and bulk MOSFETs in a wide range of VDD. These results allow us to infer preliminary indications about the amenability for an aggressive voltage scaling of TFETs compared with MOSFETs, which will be further developed in the companion paper. We also report simulation results for the sensitivity of the transistors to the variation of some key device parameters. Even these process variation results set the stage for a more thorough investigation addressed in the companion paper about the limits imposed by process variability to voltage scaling for either TFETs or MOSFETs circuits. David Esseni, Manuel Guglielmini, Bernard Kapidani, Tommaso Rollo, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2013 | New topic session 7B: Challenges and directions for ultra-low voltage VLSI circuits and systems: CMOS and beyondabstractIn this talk, a unitary perspective is given on the design challenges involved in ultra-low voltage (ULV) VLSI circuits and systems, as well as on directions to tackle them. Innovative approaches are described to improve the energy efficiency of ULV systems, while maintaining adequate resiliency and yield with low overhead. Experimental results based on the testing of 65-nm to 28-nm prototypes are presented to develop a quantitative sense of the achievable benefits. Emphasis is given on applications that require extremely high energy efficiency, such as compact portable devices and energy-autonomous VLSI systems. Although CMOS is the mainstream choice for the foreseeable future, Tunnel FETs (TFETs) are introduced as very promising alternative that favors more aggressive voltage scaling and energy reduction. Although still immature, device-circuit co-design is shown to be critical to the success of such technology. Potential of TFETs is discussed in a general framework through representative metrics and vehicle circuits, emphasizing how design will be impacted by their adoption. Bozena Kaminska, Bernard Courtois, Massimo Alioto |
VTS | 3 |
| 2012 | A simple keeper topology to reduce delay variations in nanometer domino logicabstractIn this paper, a simple topology to reduce delay variations in domino logic gates is discussed. According to a previous analysis by the same authors, the feedback loop implemented by the keeper transistor and the output inverter gate is responsible for a delay variability increase, compared to static CMOS logic. The proposed strategy reduces the loop gain associated with this feedback loop, and hence its impact on delay variations. As a result, delay variations associated with the keeper insertion are lowered by approximately 50%, with no penalty in area, noise margin and nominal performance. Massimo Alioto, Gaetano Palumbo, Melita Pennisi |
ISCAS | 1 |
| 2012 | Mixed FBB/RBB: A Novel Low-Leakage Technique for FinFET Forced StacksabstractIn this paper, a novel technique to reduce the leakage current of FinFET forced stacks under a given delay constraint is presented. This technique takes advantage of the unique feature of four-terminal FinFETs allowing different transistors to have separately tunable back bias voltages. In this work, a reverse back bias voltage is applied to one of the two stacked transistors to reduce its leakage at the cost of a delay penalty, whereas a forward back bias voltage is applied to the other one to compensate this delay degradation. The technique is assessed by means of mixed device-circuit simulations for FinFETs that are representative of 40- and 27-nm technology generations. Results show that a leakage reduction by up to 50× can be achieved as compared with traditional transistor stacks, while keeping same speed, dynamic energy, and sensitivity to process/voltage/temperature variations. Davide Baccarin, David Esseni, Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Buried Silicon-Germanium pMOSFETs: Experimental Analysis in VLSI Logic Circuits Under Aggressive Voltage ScalingabstractIn this paper, the potential of Silicon-Germanium (SiGe) technology for VLSI logic applications is investigated from a circuit perspective for the first time. The study is based on experimental measurements on 45-nm SiGe pMOSFETs with a high- κ/metal gate stack, as well as on 45-nm Si pMOSFETs with identical gate stack for comparison. In the reference SiGe technology, an innovative technological solution is adopted that limits the SiGe material only to the channel region. The resulting SiGe device merges the higher speed of the Ge technology with the lower leakage of the Si technology. Appropriate circuit- and system-level metrics are introduced to identify the advantages offered by SiGe technology in VLSI circuits. Analysis is performed in the context of next-generation VLSI circuits that fully exploit circuit- and system-level techniques to improve the energy efficiency through aggressive voltage scaling, other than low-leakage techniques. Analysis shows that the SiGe technology has more efficient leakage-delay and dynamic energy-delay trade-offs at nominal supply, compared to Si technology. Moreover, it is shown that the traditional analysis performed at nominal supply actually underestimates the benefits of SiGe pMOSFETs, since the speed advantage of SiGe VLSI circuits is further emphasized at low voltages. This demonstrates that SiGe VLSI circuits benefit from aggressive voltage scaling significantly more than Si circuits, thereby making SiGe devices a very promising alternative to Si transistors in next-generation VLSI systems. Felice Crupi, Massimo Alioto, Jacopo Franco, Paolo Magnone, Ben Kaczer, Guido Groeseneken, Jérôme Mitard, Liesbeth Witters, Thomas Y. Hoffmann |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | DET FF topologies: A detailed investigation in the energy-delay-area domainabstractIn this paper, a comparison of representative Dual-Edge-Triggered flip-flop topologies is carried out in a 65-nm CMOS technology. The energy efficiency is analyzed together with other aspects, such as the area-delay tradeoff, leakage and clock-load, which are typically neglected in previous works. The investigation highlights the impact of effects that become dominant in nanometer technologies (e.g., local interconnects, leakage) and allows for identifying the most effective FFs belonging to the DET class, as well as to evaluate the suitability of DET topologies for real applications. Massimo Alioto, Elio Consoli, Gaetano Palumbo |
ISCAS | 1 |
| 2011 | A novel back-biasing low-leakage technique for FinFET forced stacksabstractIn this paper, a novel technique to reduce the leakage current under an assigned delay constraint is presented for FinFET forced stacks. This technique is based on the adoption of different back bias voltages in stacked four terminal (4T) FinFETs (as is well known, this would not be possible in bulk CMOS circuits). In particular, a Reverse Back Bias (RBB) voltage is applied to one of the two stacked transistors to reduce its leakage at the cost of a delay penalty, whereas a Forward Back Bias (FBB) voltage is applied to the other one to compensate this delay degradation. Mixed device circuit simulations for 40-nm FinFETs show that the proposed "mixed FBB/RBB" technique permits a leakage reduction by one order of magnitude or more as compared with traditional transistor stacks at same delay. Davide Baccarin, David Esseni, Massimo Alioto |
ISCAS | 3 |
| 2011 | Experimental analysis of buried SiGe pMOSFETs from the perspective of aggressive voltage scalingabstractThis study aims to understand the potential of buried Silicon-Germanium (SiGe) technology from the perspective of VLSI logic circuits exploiting aggressive dynamic voltage scaling. Appropriate circuit- and system-level metrics are extracted from wafer-level measurements on 45nm SiGe pMOSFETs with a high-k/metal gate stack and systematically benchmarked to Si channel devices. The comparative analysis shows that the SiGe technology has more efficient leakage-delay and dynamic energy-delay trade-offs at nominal supply. These advantages of SiGe VLSI circuits are further emphasized at low voltages. This demonstrates that SiGe VLSI circuits benefit from aggressive voltage scaling significantly more than Si circuits, thereby making SiGe pMOSFET a mature candidate to substitute Si transistor for VLSI system implementations in future technology nodes. Felice Crupi, Massimo Alioto, Jacopo Franco, Paolo Magnone, Ben Kaczer, Guido Groeseneken, Jérôme Mitard, Liesbeth Witters, Thomas Y. Hoffmann |
ISCAS | 2 |
| 2011 | Leakage Power Analysis attacks: Effectiveness on DPA resistant logic styles under process variationsabstractIn this paper, the effectiveness of the recently proposed Leakage Power Analysis (LPA) attacks to cryptographic circuits is analyzed in the presence of process variations. Reference circuits (e.g., S-BOX, crypto core) were designed in various logic styles, and their robustness against LPA attacks was comparatively evaluated through Monte Carlo simulations in 65 nm. Analysis allowed for better understanding the impact that process variations have on the outcome of LPA attacks, which is an aspect that is not understood currently. Results show that LPA attacks are rather effective also under die-to-die and within-die process variations. Moreover, the comparison between different logic styles showed that standard CMOS logic circuits are extremely vulnerable to LPA attacks. Other logic styles that are robust against traditional Differential Power Analysis (DPA) attacks were also compared. Interestingly, analysis showed that these logic styles are still vulnerable to LPA attacks. Hence, LPA attacks are an even greater threat to Smart Cards information security, compared to DPA attacks. Moreover, traditional methods to protect Smart Cards against DPA attacks are ineffective in counteracting LPA attacks, thereby showing that a significant research effort will be needed to counteract LPA attacks with suitable solutions that ensure high security standards. Milena Djukanovic, Luca Giancane, Giuseppe Scotti, Alessandro Trifiletti, Massimo Alioto |
ISCAS | 5 |
| 2011 | Tapered-VTH CMOS buffer design for improved energy efficiency in deep nanometer technologyabstractIn this paper, the novel "tapered-Vth" approach to design energy-efficient CMOS buffers is introduced. In this approach, the substantial energy consumption due to leakage is reduced by tapering the threshold voltage throughout the buffer stages, other than tapering the transistor size. More specifically, the threshold voltage is progressively reduced when going from the last to the first stage. This enables a considerable leakage reduction in the last stages (which contribute most to the overall leakage) at the price of a higher delay. The resulting delay penalty is then compensated by reducing the transistor threshold voltage in the first stages, with an insignificant leakage increase (they contribute very little to the overall buffer leakage). Simulation results based on a commercial 45-nm 1-V CMOS technology show that the proposed "tapered-VTH" approach can considerably improve the energy efficiency of CMOS buffers over the entire spectrum of possible energy-delay tradeoffs, from high speed to low power. Fabio Frustaci, Pasquale Corsonello, Massimo Alioto |
ISCAS | 3 |
| 2011 | Optimized design of parallel carry-select adders
Massimo Alioto, Gaetano Palumbo, Massimo Poli |
Integr. | 1 |
| 2011 | Comparative Evaluation of Layout Density in 3T, 4T, and MT FinFET Standard CellsabstractIn this paper, issues related to the physical design and layout density of FinFET standard cells are discussed. Analysis significantly extends previous analyses, which considered the simplistic case of a single FinFET device or extremely simple circuits. Results show that analysis of a single device cannot predict the layout density of FinFET cells, due to the additional spacing constraints imposed by the standard cell structure. Results on the layout density of FinFET standard cell circuits are derived by building and analyzing various cell libraries in 32-nm technology, based on three-terminal (3T) and four-terminal (4T) devices, as well as on the recently proposed cells with mixed 3T-4T devices (MT). The results obtained for spacer- and lithography-defined FinFETs are observed from the technology scaling perspective by also considering 45- and 65-nm libraries. The effect of the fin and cell height on the layout density is studied. Results show that 3T and MT FinFET standard cells can have the same layout density as bulk cells (or better) for low (moderate) fin heights. Instead, 4T standard cells have an unacceptably worse layout density. Hence, MT standard cells turn out to be the only viable option to apply back biasing in FinFET standard cell circuits. Massimo Alioto |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Analysis and Comparison in the Energy-Delay-Area Domain of Nanometer CMOS Flip-Flops: Part I - Methodology and Design StrategiesabstractIn this paper (split into Parts I and II), an extensive comparison of existing flip-flop (FF) classes and topologies is carried out. In contrast to previous works, analysis explicitly accounts for effects that arise in nanometer technologies and affect the energy-delay-area tradeoff (e.g., leakage and the impact of layout and interconnects). Compared to previous papers on FFs comparison, the analysis involves a significantly wider range of FF classes and topologies. In particular, in this Part I, the comparison strategy, which includes the simulation setup, the energy-delay estimation methodology, and an overview of an optimum design strategy, together with the introduction of the analyzed FF classes and topologies, are reported. Massimo Alioto, Elio Consoli, Gaetano Palumbo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Analysis and Comparison in the Energy-Delay-Area Domain of Nanometer CMOS Flip-Flops: Part II - Results and Figures of MeritabstractIn Part II of this paper, a comparison of the most representative flip-flop (FF) classes and topologies in a 65-nm CMOS technology is carried out. The comparison, which is performed on the energy-delay-area domain, exploits the strategies and methodologies for FFs analysis and design reported in Part I. In particular, the analysis accounts for the impact of leakage and layout parasitics on the optimization of the circuits. The tradeoffs between leakage, area, clock load, delay, and other interesting properties are extensively discussed. The investigation permits to derive several considerations on each FF class and to identify the best topologies for a targeted application. Massimo Alioto, Elio Consoli, Gaetano Palumbo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Understanding the Potential and the Limits of Germanium pMOSFETs for VLSI Circuits From Experimental MeasurementsabstractIn this paper, potential and limits of Germanium pMOSFETs for VLSI applications are investigated from a circuit perspective for the first time in the literature. Since short-channel Germanium devices have been developed only recently, no circuit design tools are currently available, hence most of the results available in the literature address process and device-level issues (currently, down to the 65 nm node). However, the suitability of Germanium MOSFETs for VLSI circuits should be assessed at circuit level. To fill this gap, we introduce an innovative methodology that extracts the main circuit parameters of interest (e.g., speed, dynamic power, leakage) from measurements on experimental devices. Appropriate figures of merit are adopted to highlight the potential of Germanium MOSFETs under realistic VLSI designs that fully exploit system-level schemes to minimize leakage (e.g., body biasing, stack forcing, power gating). Measurements and evaluations are performed on 125 nm Germanium pMOSFETs with a high-κ/metal gate stack having an equivalent oxide thickness of 1.3 nm. Comparison with Si pMOSFET prototypes implemented with similar gate stack is also carried out to comparatively understand the potential and the weaknesses of Germanium transistors. The main experimental results are justified through theoretical analysis as a function of the relevant circuit and device parameters. Some system-level aspects are also investigated, such as the energy efficiency and the wakeup time of body-biasing schemes in Ge circuits and the impact of voltage scaling. Paolo Magnone, Felice Crupi, Massimo Alioto, Ben Kaczer, Brice De Jaeger |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Closed-form analysis of DC noise immunity in subthreshold CMOS logic circuitsabstractIn this paper, subthreshold static CMOS logic is analyzed in terms of DC noise immunity in a closed form for the first time. Simplified circuit models of MOS transistors in subthreshold are developed to gain a deeper understanding of the degradation in the DC characteristics under ultra-low voltages, as well as its dependence on design and process parameters. The noise margin is explicitly evaluated and modeled with a simple expression. The impact of PMOS/NMOS imbalance is also explicitly analyzed. Results are validated with simulations in a 65-nm CMOS technology. Massimo Alioto |
ISCAS | 1 |
| 2010 | Analysis of layout density in FinFET standard cells and impact of fin technologyabstractIn this paper, the layout density of three-terminal FinFET logic circuits is extensively analyzed. As opposite to previous works, which are focused either on single devices or simplistic circuits, this analysis explicitly includes the geometric constraints that are imposed by the standard cell approach. The impact of the fin technology is analyzed by comparing the lithography- and spacer-defined approaches, as well as evaluating the dependence of layout density on the fin height. Results show that FinFET standard cells have a layout density that is better than bulk cells even for moderately tall fins. The fin height is also shown to be a powerful knob to improve the layout density in FinFET cells. Analysis also shows that the usually claimed 2X density improvement of the spacer-defined technology compared to the lithography-defined is dramatically reduced in real standard cells, and can be negligible for tall fins. All results are justified through considerations at the physical level of abstraction. Various versions of a 32-nm 44-gate library are laid out to carry out the analysis. Massimo Alioto |
ISCAS | 1 |
| 2010 | Exploiting locality to improve leakage reduction in embedded drowsy I-caches at same area/speedabstractIn this paper, a technique to reduce the leakage power consumption in embedded drowsy instruction caches (I caches) is proposed. The technique is called “Improved Drowsy” (ID), and adopts a more efficient strategy than standard Drowsy Caches (DCs) to turn off unused cache lines, based on locality. The implementation of ID caches requires minor changes, and the area/speed overhead associated with the additional circuitry is insignificant. The proposed technique is assessed through circuit and cycle accurate simulations on an L1 instruction cache embedded in an ARM XScale processor based system in a 65 nm CMOS technology. Results show that this technique is able to reduce the leakage power by 69% on average. Leakage of DC is shown to be significantly lowered with the proposed ID approach, being DC leakage greater than that of ID by up to 53%, and 10 15% typically. Massimo Alioto, Paolo Bennati, Roberto Giorgi |
ISCAS | 1 |
| 2010 | Clock distribution in clock domains with Dual-Edge-Triggered Flip-Flops to improve energy-efficiencyabstractIn this paper, an optimization strategy is proposed for Dual Edge-Triggered (DET) clock distribution in clock domains, based on the proper choice of the clock slope. The suggested approach takes full advantage of the intrinsic features of DET Flip-Flops to achieve up to 50% energy-savings compared to traditional DET design approaches. The speed penalty, in terms of both FFs delay and local skew/jitter, is proven to be negligible through extensive simulations in a 65-nm CMOS technology. Massimo Alioto, Elio Consoli, Gaetano Palumbo |
ISCAS | 1 |
| 2010 | Experimental study of leakage-delay trade-off in Germanium pMOSFETs for logic circuitsabstractIn this work we explore the potential of the emerging Germanium technology for logic circuits. We introduce an innovative methodology that extracts the main circuit parameters of interest from experimental measurements on 125 nm high k metal gate Ge pMOSFETs in a Si compatible process flow. Appropriate figures of merit are adopted to highlight the potential of Germanium MOSFETs under realistic VLSI designs that fully exploit system level schemes to minimize leakage (e.g., body biasing, stack forcing). On the one hand, Ge devices outperform Si devices in terms of speed due to the higher hole mobility. On the other hand, the higher off state drain current, evaluated ignoring the junction leakage, in Ge pMOSFETs causes an higher standby power dissipation. We show how this drawback can be alleviated by the application of back biasing and stack effect techniques which are intrinsically more effective in Ge devices. In addition, analysis shows that Ge circuits can actually exhibit a 6.4X lower leakage than Si devices, if the threshold voltage is tuned to match the speed of Si devices. Paolo Magnone, Felice Crupi, Massimo Alioto, Ben Kaczer |
ISCAS | 3 |
| 2010 | Design metrics for RTL level estimation of delay variability due to intradie (random) variationsabstractA simple metric is presented for the accurate prediction of path delay variability during the automated synthesis of digital VLSI circuits. This allows circuit variability to be assessed at early stages within the design process with minimal computational effort, as extensive Monte Carlo or SSTA runs are not required. This paper introduces the metric and investigates its effectiveness. The final predictions of path delay variability are found to be within 10% of measured path delay variability, with an average error of 3%, for a series of test paths synthesised from randomised models of a 130nm technology library. These randomised models are generated from a 3D atomistic simulator and provide more accuracy than traditional Monte Carlo simulation runs. Michael Merrett, Mark Zwolinski, Koushik Maharatna, Massimo Alioto |
ISCAS | 5 |
| 2010 | Differential Power Analysis Attacks to Precharged Buses: A General Analysis for Symmetric-Key Cryptographic AlgorithmsabstractIn this paper, a general model of multibit Differential Power Analysis (DPA) attacks to precharged buses is discussed, with emphasis on symmetric-key cryptographic algorithms. Analysis provides a deeper insight into the dependence of the DPA effectiveness (i.e., the vulnerability of cryptographic chips) on the parameters that define the attack, the algorithm, and the processor architecture in which the latter is implemented. To this aim, the main parameters that are of interest in practical DPA attacks are analytically derived under appropriate approximations, and a novel figure of merit to measure the DPA effectiveness of multibit attacks is proposed. This figure of merit allows for identifying conditions that maximize the effectiveness of DPA attacks, i.e., conditions under which a cryptographic chip should be tested to assess its robustness. Several interesting properties of DPA attacks are derived, and suggestions to design algorithms and circuits with higher robustness against DPA are given. The proposed model is validated in the case of DES and AES algorithms with both simulations on an MIPS32 architecture and measurements on an FPGA-based implementation of AES. The model accuracy is shown to be adequate, as the resulting error is always lower than 10 percent and typically of a few percentage points. Massimo Alioto, Massimo Poli, Santina Rocchi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2010 | Leakage-Delay Tradeoff in FinFET Logic Circuits: A Comparative Analysis With Bulk TechnologyabstractIn this paper, we study the advantages offered by multi-gate fin FETs (FinFETs) over traditional bulk MOSFETs when low standby power circuit techniques are implemented. More precisely, we simulated various vehicle circuits, ranging from ring oscillators to mirror full adders, to investigate the effectiveness of back biasing and transistor-stacking in both FinFETs and bulk MOSFETs. The opportunity to separate the gates of FinFETs and to operate them independently has been systematically analyzed; mixed connected- and independent-gate circuits have also been evaluated. The study spans over the device, the layout, and the circuit level of abstraction and appropriate figures of merit are introduced to quantify the potential advantage of different schemes. Our results show that, thanks to a larger threshold voltage sensitivity to back biasing, the FinFET technology is able to offer a more favorable compromise between standby power consumption and dynamic performance and is well suited for implementing fast and energy-efficient adaptive back-biasing strategies. Matteo Agostinelli, Massimo Alioto, David Esseni, Luca Selmi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Understanding the Effect of Process Variations on the Delay of Static and Domino LogicabstractIn this paper, the effect of process variations on delay is analyzed in depth for both static and dynamic CMOS logic styles. Analysis allows for gaining an insight into the delay dependence on fan-in, fan-out, and sizing in sub-100-nm technologies. Simple but reasonably accurate models are derived to capture the basic dependences. The effect of process variations in transistor stacks is analytically modeled and analyzed in detail. The impact of both interdie and intradie variations is evaluated and discussed. Interestingly, the input capacitance of static and dynamic logic is shown to be rather insensitive to variations. The delay variability was also shown to be a weak function of the input rise/fall time and load. Analysis shows that domino logic circuits suffer from a doubled variability as compared to the static CMOS logic style. The positive feedback associated with the keeper transistor is shown to be responsible for the variability increase, which, in turn, limits the speed performance. This adds to the well-known speed degradation due to the current contention associated with the keeper transistor. Monte Carlo simulations on a 90-nm technology, including layout parasitics, are performed to validate the results. Massimo Alioto, Gaetano Palumbo, Melita Pennisi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | A General Power Model of Differential Power Analysis Attacks to Static Logic CircuitsabstractThis paper discusses a general model of differential power analysis (DPA) attacks to static logic circuits. Focusing on symmetric-key cryptographic algorithms, the proposed analysis provides a deeper insight into the vulnerability of cryptographic circuits. The main parameters that are of interest in practical DPA attacks are derived under suitable approximations, and a new figure of merit to measure the DPA effectiveness is proposed. Worst case conditions under which a cryptographic circuit should be tested to evaluate its robustness against DPA attacks are identified and analyzed. Several interesting properties of DPA attacks are also derived from the proposed model, whose fundamental expressions are compared with the counterparts of correlation power analysis attacks. The model was validated by means of DPA attacks on an FPGA implementation of the advanced encryption standard algorithm. Experimental results show that the model has a good accuracy, as its error is always lower than 2%. Massimo Alioto, Massimo Poli, Santina Rocchi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Understanding Loading Effects of RC Uniform InterconnectsabstractIn this paper, loading effects associated with uniform RC interconnects are analytically modeled to deeply understand the effects that contribute to their input admittance. A simple reduced-order model is developed and shown to be accurate by means of simulations on a 65-nm CMOS technology. The resulting analytical expression of the model permits a deeper understanding of the loading effects of RC uniform wires on the driving circuit (e.g., the previous repeater). The model is also simple enough to be used in fast and accurate estimations in automated CAD tools. Interestingly, analysis allows for identifying a new circuit effect, the so-called ldquocapacitive shieldingrdquo effect. This effect consists in the reduction of the effective resistance (and capacitance) that is seen at the input of the interconnect, due to the wire and load capacitances. Interestingly, the capacitive effect is shown to be the dual of the well-known ldquoresistive shieldingrdquo effect associated with RC interconnects, which determines a reduction in the effective capacitance due to the wire resistance. Massimo Alioto |
ISCAS | 1 |
| 2009 | Optimization of Wire Grid Size for Differential Routing and Impact on the Power-delay-area TradeoffabstractIn this paper, the impact of the wire grid size on the power-delay-area tradeoff of VLSI digital circuits with differential routing is analyzed. To this aim, the differential MOS current-mode logic (MCML) is adopted as reference logic style, and a complete differential design flow is used. Analysis shows that the choice of the grid size in differential routing has a much stronger impact on the power-delay-area tradeoff, compared to the usual single-ended case, hence the grid size must be carefully selected. The dependence of power, delay and area on the grid size is discussed in detail through simple models and metrics. To validate the approach and show basic dependencies in practical circuits, 30 benchmark circuits with an in-house designed MCML cell library were synthesized and routed in a 0.18-mum CMOS technology. Results show that non-optimal choice of the grid size can determine a dramatic increase in power (1.7X) and area (1.3X). Interestingly, the grid size that optimizes the power-delay-area tradeoff depends very weakly on the specific circuit under design, hence a generally optimum grid size exists that optimizes a very wide range of different circuits. Massimo Alioto, Stéphane Badel, Yusuf Leblebici |
ISCAS | 1 |
| 2009 | Metrics and Design Considerations on the Energy-delay Tradeoff of Digital CircuitsabstractIn this paper, general metrics of the energy-delay (E-D) tradeoff in digital VLSI circuits are discussed. More specifically, the general class of metrics EiDjwith arbitrary exponents is adopted and evaluated for various commercial microprocessors. Results indicate that practical circuits are designed by minimizing a wider range of metrics compared to the ED or ED2metrics usually assumed in the literature. Hence, the general metrics EiDjdescribes the energy-delay tradeoff in a more realistic way. An interesting interpretation of the adopted metrics is provided to gain an insight into the relationship between energy and delay in energy-efficient designs. Various properties are also derived analytically by resorting to the logical effort method. Simulations on a 65-nm technology are performed to exemplify and validate the theoretical results. Massimo Alioto, Elio Consoli, Gaetano Palumbo |
ISCAS | 1 |
| 2009 | Analysis and Design of Ultra-low Power Subthreshold MCML GatesabstractIn this paper, ultra-low power current-mode subthreshold MOS current-mode logic (MCML) gates are discussed from a modeling and design perspective. A detailed analysis of the DC characteristics is presented, and the effect of process variations is analyzed in depth. Analysis allows for understanding the main limits of sub-threshold MCML gates in terms of delay/power variability. In particular, it is shown that process variations strongly affect the DC characteristics, and moderately impact delay and power consumption. Interestingly, delay and power variations are shown to be significantly reduced compared to typical values encountered in standard subthreshold CMOS logic. Criteria to size transistors to keep variations within assigned bounds are also derived. Results of Monte Carlo simulations with a 65-nm CMOS technology are reported to validate theoretical results. Massimo Alioto, Yusuf Leblebici |
ISCAS | 1 |
| 2009 | Analysis and Modeling of Energy Consumption in RLC Tree CircuitsabstractIn this paper, the energy consumption of resistance-inductance-capacitance (RLC) trees is analytically modeled. In particular, the results obtained by the same authors forRCtree circuits are generalized, allowing for a deep understanding of the impact of the inductance. The modeling approach proposed relies on the adoption of an equivalent second-orderRLCcircuit, whose energy consumption is evaluated in a closed form. These results are then extended toRLCcircuits with arbitrary order, deriving a simple and accurate model. The energy dependence on the input rise time is also analyzed in detail, identifying the ranges for which theRLCcircuit can be approximated to a simple capacitance or anRCcircuit. The model equations provide an insight into the dependence of the energy consumption on the circuit parameters. Indeed, the energy is explicitly expressed as a function of the resistances, capacitances and inductances of the original network. The energy model proposed is shown to be accurate enough for modeling purposes through comparison with SPICE simulations, as the error is typically in the order of a few percentage points. Massimo Alioto, Gaetano Palumbo, Massimo Poli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Analysis and performance evaluation of area-efficient true random bit generators on FPGAsabstractIn this paper, fully digital True Random Bit Generators (TRBGs) targeting area-efficient FPGA implementations are analyzed and evaluated. In this analysis, a very general class of TRBGs is considered that is based on the well-known sampling oscillators used as a building block. A qualitative model is discussed and applied to derive simple design guidelines to implement low-area TRBGs. Extensive measurements are made on TRBGs implemented on Altera Cyclone FPGAs. Results confirm that properly designed TRBGs based on FPGAs can achieve very high-quality random sequences with low area occupation. Massimo Alioto, Luca Fondelli, Santina Rocchi |
ISCAS | 1 |
| 2008 | Power-delay optimization in MCML tapered buffersabstractIn this paper, MOS current mode logic (MCML) tapered buffers are discussed from a design point of view. Closed-form design equations that relate the overall speed performance and the power consumption of MCML tapered buffers are derived for nanometer CMOS technologies, i.e. by accounting for deep-sub-micron (DSM) effects from the beginning. The power-delay design space is then analytically explored, and design criteria are derived to properly size the number of stages and the current tapering factor under a speed/power constraint. The design criteria are simple enough to be used in pencil-and-paper calculations, as well as general and independent of the adopted technology. Hence, the proposed strategy provides the designer with an insight into the power-delay trade-off of MCML tapered buffers. Results are validated by means of simulations on a 90-nm CMOS technology. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2008 | Explicit energy evaluation in RLC tree circuits with ramp inputsabstractIn this paper, a simple and accurate model of the energy consumption in RLC trees with ramp inputs is proposed. The modeling approach is based on the reduction of the RLC tree network to a proper 2nd-order RLC equivalent circuit, according to the strategy proposed by the same authors in [1]. This approach leads to a closed-form expression of the energy consumption that allows for a deeper understanding of the contribution of the inductance, as well as of the effect of the input rise time. Thanks to its simplicity, the model is well suited for fast energy estimations. Extensive simulations show that the proposed approach is faster than SPICE simulations by four orders of magnitude. At the same time, the proposed approach provides results that are very close to SPICE simulations, with a typical error of 1.5%. Massimo Alioto, Massimo Poli, Gaetano Palumbo |
ISCAS | 1 |
| 2008 | A general model for differential power analysis attacks to static logic circuitsabstractIn this paper, a general model of multi-bit differential power analysis (DPA) attacks to static logic circuits is proposed, with emphasis on symmetric-key cryptographic algorithms. The main parameters that are of interest in practical DPA attacks are analytically derived by introducing suitable approximations. Several interesting properties of DPA attacks are derived, allowing a deep understanding of the vulnerability of algorithms and circuits. The proposed model was validated by means of experimental measurements on an FPGA implementation of the Advanced Encryption Standard (AES) algorithm. The model accuracy is shown to be adequate, as the resulting error is always lower than 11%. Massimo Alioto, Massimo Poli, Santina Rocchi |
ISCAS | 1 |
| 2008 | Improving the power-delay product in SCL circuits using source follower output stageabstractThis article explores the effect of using source follower buffers (SFB) at the output of source coupled logic (SCL) circuits. This technique can help to improve the power-delay product (PDP) of an SCL gate approximately by a factor of two. The proposed approach has been applied to improve the PDP in sub-threshold SCL circuits that have been developed for ultra- low power applications. Designed in conventional digital 0.18mum CMOS technology, the proposed SCL gate utilizing SFB at the output achieves a PDP of 0.5fJ/fF/gate while the gate draws 10nA from a 0.6V supply voltage. Armin Tajalli, Frank K. Gürkaynak, Yusuf Leblebici, Massimo Alioto, Elizabeth J. Brauer |
ISCAS | 4 |
| 2007 | Maximum-Period PRNGs Derived From A Piecewise Linear One-Dimensional MapabstractIn this paper a novel family of maximum-period nonlinear congruential generators (NLCGs) based on the digitized sawtooth map is considered for the definition of hardware and software efficient pseudo random number generators (PRNGs). A list of such maximum period NLCGs for period lengths up to 231-1 is provided. Referring to the NIST800-22 statistical test suite, a PRNG example based on the combination of two of the proposed NLCGs is presented and discussed Tommaso Addabbo, Massimo Alioto, Ada Fort, Santina Rocchi, Valerio Vignoli |
ISCAS | 2 |
| 2007 | High-Speed/Low-Power Mixed Full Adder Chains: Analysis and Comparison versus TechnologyabstractIn this paper, the mixed-topology Full Adder chains proposed in [1] are extensively analyzed versus technology. Analysis aims at exploring the power-delay design space in mixed-topology full adder chains, and evaluating the effect of technology scaling. The mixed-topology approach is also compared with the most representative single-topology circuits for technologies spanning five technology nodes (from 90 nm to 0.35 μm). All circuits are designed at the transistor and the physical level for different design targets, and results account for the layout parasitics. Results demonstrate that the mixed-topology approach is very competitive for every design target, and its advantage over single-topology circuits increases as down-scaling the technology. As a result, the mixed-topology approach is expected to be increasingly appealing when implementing Full Adder chains in down-scaled technologies. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2007 | Design of Fast Large Fan-In CMOS Multiplexers Accounting for InterconnectsabstractIn this paper, the design of high-fan-in CMOS multiplexers based on the heterogeneous-tree approach is discussed. In particular, a strategy to minimize the delay of multiplexers is developed that accounts for the interconnect parasitics from the beginning; thereby extending the previous results introduced in M. Lim (2000) which did not consider the effect of interconnects. The design criteria derived are very simple, and are shown to be strongly affected by interconnects, as one expects in current deep-submicron (DSM) VLSI circuits. It is also shown that neglecting parasitics in the multiplexer optimization can lead to speed degradation as high as 80%. The results are validated through post-layout simulations on a 90-nm CMOS process. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2007 | Delay Variability Due to Supply Variations in Transmission-Gate Full AddersabstractIn this paper, the delay variability due to supply variations is investigated for the Transmission-Gate (TG) Full Adder topology, which is well known for its very low power consumption. The delay sensitivity with respect to supply variations is first analytically modeled. The resulting model is very simple, independent of the adopted technology and useful for better understanding the delay variations due to the supply voltage fluctuations. The delay sensitivity with respect to supply variations is also compared with that of traditional CMOS Full Adders, that are frequently adopted as a reference logic style. The results are validated by means of Spectre simulations with a 90-nm and a 0.18-μm technology. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2007 | Mixed Techniques to Protect Precharged Busses against Differential Power Analysis AttacksabstractIn this paper, techniques to improve the resistance against differential power analysis (DPA) attacks of precharged busses in cryptographic circuits are discussed. In particular, two techniques that were previously introduced by the same authors are properly mixed to further enhance the immunity to DPA attacks. The achieved robustness against DPA attacks is shown to be considerably improved, compared with the case of a separate adoption of each technique. Criteria to manage the security-power-area trade-off are also derived from a statistical analysis of precharged busses. The mixed technique is finally validated by means of both cycle-accurate and circuit simulations on the DES encryption algorithm running on a MIPS32 architecture. Massimo Alioto, Massimo Poli, Santina Rocchi, Valerio Vignoli |
ISCAS | 1 |
| 2006 | A technique to design high entropy chaos-based true random bit generatorsabstractIn this paper the theoretical bases to design a true random bit generator (TRBG) circuit with a predefined minimum entropy are discussed. The approach is tailored to TRBGs based on a one-dimensional piecewise-linear chaotic map, and it is based on a feedback control procedure that allows to dynamically changing the system parameters. For this purpose, the procedure just exploits the TRBG output observation without requiring bit throughput reduction. The design approach was validated by an hardware prototype implemented on a field programmable analog array (FPAA) Tommaso Addabbo, Massimo Alioto, Ada Fort, Santina Rocchi, Valerio Vignoli |
ISCAS | 2 |
| 2006 | Delay uncertainty due to supply variations in static and dynamic full addersabstractIn this paper, the delay uncertainty due to supply variations is investigated for two important full adder topologies. In particular, it is developed an analytical model of the delay sensitivity to supply variations for the static mirror adder and the dynamic dual-rail domino adder. The model is general and very simple, and allows for identifying the main parameters which define the delay uncertainty due to supply variations, as well as deriving design considerations. In particular, the importance of the input rise/fall time variations is clarified, and the effect of the supply voltage reduction and technology scaling is discussed. Results are validated through SPICE simulations with a 0.18-mum and a 0.35-mum technology Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2006 | Nanometer MCML gates: models and design considerationsabstractIn this paper, analytical models of the static and dynamic behavior of MOS current-mode logic (MCML) with a resistive load are discussed. These models account for deep-submicron (DSM) effects which affect the operation of this CMOS logic style in the nanometer regime. In particular, a noise margin model is derived by resorting to the alpha-power law, and comparison with the long-channel expression allows for clearly understanding the impact of DSM effects. The dynamic behavior is also analyzed by accurately modeling the resistive load, i.e. accounting for its capacitive parasitics, which are shown to give an important contribution in low-power designs. Analytical results and considerations are validated by means of Spectre simulations on a 90-nm CMOS technology Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2006 | Efficient output transition time modeling in CMOS gates with ramp/exponential inputsabstractIn this paper, the modeling of the output transition time in deep-submicron (DSM) CMOS gates is discussed. In particular, the analysis starts from a previously proposed analytical model valid only for ramp inputs (Maurine et al., 2002 and Auvergne et al., 2000). This model is then improved by introducing two semi-empirical coefficients which have to be tuned by means of two SPICE simulations. Since an exponential input is more and more frequent in current DSM CMOS technologies, the model is extended to this kind of input waveform. Results are validated with simulations on a 0.18-mum CMOS technology. The output transition time model is found to be in good agreement with SPICE simulations, with an average error of only 5% Massimo Alioto, Gaetano Palumbo, Massimo Poli |
ISCAS | 1 |
| 2006 | Analysis and design of MCML gates with hysteresisabstractIn this paper, hysteresis is exploited to improve the performance of positive feedback source coupled logic circuits, which are a modification of the traditional MOS current-mode logic (MCML) (Alioto, 2004). To understand the effect of hysteresis on the DC characteristics, a model of the noise margin is analytically derived. This model shows that hysteresis improves the noise margin, whose increase is traded-off to reduce the logic swing, which in turn can have a beneficial impact on the speed performance. Practical cases where hysteresis is advantageous are identified, and a comparison with PFSCL gates without hysteresis is carried out. Analysis shows that in such cases hysteresis significantly improves the speed performance and the power efficiency of PFSCL gates, which is a critical aspect in this kind of logic. Simulation results are presented based on a 0.18-mum CMOS process Massimo Alioto, Luca Pancioni, Santina Rocchi, Valerio Vignoli |
ISCAS | 1 |
| 2006 | Impact of Supply Voltage Variations on Full Adder Delay: Analysis and ComparisonabstractIn this paper, some of the most practically interesting full adder topologies are analyzed in terms of their delay dependence on the supply voltage fluctuations, which are a major contribution to the delay uncertainty, which in turn limits the speed performance of current VLSI circuits. Analytical models of the delay sensitivity with respect to supply variations are derived by following a simplified circuit analysis, and the resulting expressions are simple enough to afford a deeper insight into the impact of supply voltage variations on each topology. The models are shown to be sufficiently accurate through simulations with CMOS technologies having a minimum feature size ranging from 90 nm to 0.35 mum. Several interesting properties and design considerations are derived from these models, and the effect of the supply voltage scaling, technology scaling, transistor sizing, and input transition time is discussed. Strategies to evaluate the delay sensitivity since the early design phases (e.g., from ring oscillator measurements) are also introduced. As a fundamental result, it is shown that the delay sensitivity to supply variations will increase in the next technology nodes, thus, it is expected that controlling the supply variations will be an increasingly important issue in the design of the next generation VLSI circuits. The proposed methodology is also analyzed in the case of more general digital circuits, and is used to estimate the impact of the inter-die threshold voltage variations on the delay of the considered full adder topologies Massimo Alioto, Gaetano Palumbo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Energy Consumption in RC Tree CircuitsabstractIn this paper, resistance-capacitance (RC) tree networks are modeled in terms of their energy consumption associated with an input transition. This work significantly extends the results that the same authors previously obtained in the specific case of ladder networks with only ramp signals. The proposed approach to model the energy consumption is based on a single-pole approximation, in which an equivalent time constant is analytically derived from an exact analysis for very slow and very fast input transitions. The model is then extended to arbitrary values of the input rise time by exploiting some intrinsic properties of RC tree networks. The approach is completely analytical and leads to closed-form results. Analytical results are explicitly derived for different inputs, such as the ramp and the exponential waveforms which are usually encountered in current VLSI circuits, as well as the saturated sine input. Due to its simplicity, the proposed energy expression is suitable for pencil-and-paper evaluation and allows for an intuitive understanding of the network dissipation. The energy expression proposed is shown to be accurate enough for modeling purposes through comparison with SPICE simulations. Massimo Alioto, Gaetano Palumbo, Massimo Poli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Evaluation of energy consumption in RC ladder circuits driven by a ramp inputabstractIn this paper, the energy consumption of RC ladder networks, which can represent chains of transmission gate or long wire interconnections, is modeled. Their energy dependence on the input rise time is analyzed by assuming a ramp input waveform. Since the analysis can be carried out in a straightforward manner only for very simple RC ladder networks, the exact analysis is first limited to asymptotic values of the input rise time T (i.e., for T/spl rarr/0 and T/spl rarr//spl infin/). Successively, the energy expression is extended to arbitrary values of the input rise time by introducing a suitable equivalent first-order RC circuit, whose resistance and capacitance are simply related to the resistances and capacitances of the original network. The energy expression found is useful for pencil-and-paper evaluation and affords an intuitive understanding of the network dissipation, since each term has an evident physical meaning. By comparison with SPICE simulations, the energy expression proposed is showed to be accurate enough for modeling purposes. Massimo Alioto, Gaetano Palumbo, Massimo Poli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | Performance evaluation of the low-voltage CML D-latch topology
Massimo Alioto, Rosario Mita, Gaetano Palumbo |
Integr. | 1 |
| 2002 | Analysis and comparison on full adder block in submicron technologyabstractIn this paper the main topologies of one-bit full adders, including the most interesting of those recently proposed, are analyzed and compared for speed, power consumption, and power-delay product. The comparison has been performed on two classes of circuits, the former with minimum transistor size to minimize power consumption, the latter with optimized transistor dimension to minimize power-delay product. The investigation has been carried out with properly defined simulation runs on a Cadence environment using a 0.35-/spl mu/m process, also including the parasitics derived from layout. Performance has been also compared for different supply voltage values. Thus design guidelines have been derived to select the most suitable topology for the design features required. This paper also proposes a novel figure of merit to realistically compare n-bit adders implemented as a chain of one-bit full adders. The results differ from those previously published both for the more realistic simulations carried out and the more appropriate figure of merit used. They show that, except for short chains of blocks or for cases where minimum power consumption is desired, topologies with only pass transistors or transmission gates are not attractive. In contrast, the most interesting implementations in terms of trade off between power and delay are the traditional CMOS and mirror topologies. Moreover, the dual-rail domino and the CPL allow the best speed performance. Massimo Alioto, Gaetano Palumbo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2001 | Power estimation in adiabatic circuits: a simple and accurate modelabstractA simple procedure to evaluate the energy consumption of adiabatic gate circuits is proposed and validated. The proposed strategy is based on a linearization of the circuit and simplifying the analytical result obtained on the equivalent network. The approach leads to simple relationships which can be used for a pencil-and-paper evaluation or implemented on software. The accuracy of the results is validated by means of Spice simulations on an adiabatic full adder designed with a 0.8 /spl mu/m technology. Massimo Alioto, Gaetano Palumbo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Evaluation of power consumption in adiabatic circuitsabstractIn this paper a simple and accurate model to evaluate the energy consumption of the digital adiabatic circuits is proposed. It is based on a linearization of the circuit, which leads to an RC mesh equivalent circuit. The obtained expression of the energy consumption of a generic adiabatic gate is very simple, thus it can be used for pencil-and-paper calculations. The predicted values were compared to Spice results using a 0.8 /spl mu/m CMOS technology. The average error is below 3.5%, and for practical switching frequencies the maximum error is always lower than 8%. Therefore, the proposed model can be also used to implement an accurate and efficient power simulator tool. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 2000 | High-speed bipolar MUX modeling and designabstractThis paper presents modeling and optimized design of Current Mode Logic (CML) MUX. Propagation delay models with few terms are presented. The most accurate model has errors lower than 2%. By using the proposed models a design optimization is proposed. In particular, the bias currents which gives the minimum propagation delay are found. Moreover, it is demonstrated that at the cost of a 10% increase in propagation delay a 40% reduction in the power dissipation can be achieved. The models and design strategies are validated using both a traditional and a high-speed bipolar process which have a transition frequency equal to 6 GHz and 20 GHz, respectively. Massimo Alioto, Gaetano Palumbo |
ISCAS | 1 |
| 1999 | Highly accurate and simple models for CML and ECL gatesabstractIn this paper simple and accurate models for the propagation delay of both current mode logic (CML) and emitter-coupled logic (ECL) gates are proposed. The models start from the small signal model properly evaluated. This makes it possible to represent propagation delay with a few terms, providing a better insight into the relationship between delay and its electrical parameters, which in turn are related to process parameters. The main difference between accurate and simple models is that the former need few Spice simulations to properly evaluate the model parameters. In order to validate the models, a comparison using both a traditional and a high-speed bipolar process was carried out under many bias conditions and output loads. Simple models have typical errors of around 20%. Accurate models have typical errors as low as 2% and 5% for CML and ECL, respectively, while the worst case error is as low as 5% and 8% for CML and ECL, respectively. Massimo Alioto, Gaetano Palumbo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Novel Simple Models Of Cml Propagation DelayabstractAccurate and simple models of CML propagation delay are given. The approach used is new. The propagation delay is represented with a few terms, providing a better insight into the relationship between delay and its electrical parameters, which in turn are related to process parameters. The most accurate model has a typical and worst case errors as low as 2% and 5%, respectively. Massimo Alioto, Gaetano Palumbo |
Great Lakes Symposium on VLSI | 1 |