Paul D. Franzon

dblp:f/PaulDFranzon · also Paul Franzon · DBLP profile ↗
← Back
54ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-6048-5770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 50 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Exploiting Power Side-Channel Vulnerabilities in XGBoost Accelerator
abstract
XGBoost (eXtreme Gradient Boosting), a widelyused decision tree algorithm, plays a crucial role in applications such as ransomware and fraud detection. While its performance is well-established, its security against model extraction on hardware platforms like Field Programmable Gate Arrays (FPGAs) has not been fully explored. In this paper, we demonstrate a significant vulnerability where sensitive model data can be leaked from an XGBoost implementation through side-channel attacks (SCAs). By analyzing variations in power consumption, we show how an attacker can infer node features within the XGBoost model, leading to the extraction of critical data. We conduct an experiment using the XGBoost accelerator FAXID on the Sakura-X platform, demonstrating a method to deduce model decisions by monitoring power consumptions. The results show that on average 367k tests are sufficient to leak sensitive values. Our findings underscore the need for improved hardware and algorithmic protections to safeguard machine learning models from these types of attacks.
Yimeng Xiao, Archit Gajjar, Aydin Aysu, Paul D. Franzon
DAC4
2025 A 27-30 GHz T/R Module With Reflection-Type Phase Shifting and Machine-Learned Calibration
abstract
This paper presents a transmit/receive module (TRM) for phased arrays realized in 45nm RFSOI CMOS technology and calibrated using machine learning. The 27-30GHz TRM includes a transmit/receive (T/R) switch, a power amplifier, a low-noise amplifier, another T/R switch, and a bidirectional reflection-type phase shifter (RTPS). The RTPS incorporates multiple resonators and five control variables to achieve a six-bit resolution with a 360-degree phase shift range across a 10% bandwidth. We introduce a machine-learning technique that uses Bayesian optimization to calibrate the multi-variable front end. This technique can attain near-optimal settings with 1.5 percent of the measurements compared to manual calibration using an exhaustive search. Measurements show the TRM achieves 16.4dB gain, 2.5GHz 1dB bandwidth, and 11.9-12.9dBm output compression point in transmit mode, and 16dB gain, 3.2GHz 1dB BW, −23.3dBm input compression point, and 4dB noise figure in receive mode. Across 27-30GHz, the calibrated TRM achieves root-mean-square errors of 0.4dB or lower for gain and less than 1.5 or 2.8 degrees for phase in transmit and receive modes, respectively.
Yuejiang Wen, Zhangjie Hong, Bharadwaj Padmanabhan, Paul D. Franzon, Brian A. Floyd
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 RD-FAXID: Ransomware Detection with FPGA-Accelerated XGBoost
abstract
Over the last decade, there has been a rise in cyberattacks, particularly ransomware, causing significant disruption and financial repercussions across public and private sectors. Tremendous efforts have been spent on developing techniques to detect ransomware to, ideally, protect data or have as minimum data loss as possible. Ransomware attacks are becoming more frequent and sophisticated as there is a constant tussle between attackers and cybersecurity defenders. Machine Learning (ML) approaches have proven more effective in detecting ransomware than classical signature-based detection. In particular, tree-based algorithms such as Decision Trees (DT), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) spike up interest among cybersecurity researchers. However, due to the nature of the problem, traditional CPUs and GPUs fail to keep up with the desired performance, especially for large data workloads. Thus, the problem demands a customized solution to detect the ransomware. Here, we propose an FPGA accelerated tree-based ML model for multi-dataset ransomware detection. We show the capability of the proposed prototype to address the problem from more than one set of features, reducing false positive and negative rates to have robust predictions by looking at Hardware Performance Counters (HPCs), Operating System (OS) calls, and network traffic information simultaneously. With 1,000 samples per batch, the FPGA prototype has 65.8 \({\times}\) and 4.1 \({\times}\) lower latency over the CPU and GPU, respectively. Moreover, the FPGA design is up to 11.3 \({\times}\) cost-effective and 643 \({\times}\) energy-efficient compared to the CPU and 3 \({\times}\) cost-effective and 16.8 \({\times}\) energy-efficient over the GPU.
Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Chris Cheng, Giacomo Pedretti, Jim Ignowski
ACM Trans. Reconfigurable Technol. Syst.4
2022 FAXID: FPGA-Accelerated XGBoost Inference for Data Centers using HLS
abstract
Advanced ensemble trees have proven quite effective in providing real-time predictions against ransomware detection, medical diagnosis, recommendation engines, fraud detection, failure predictions, crime risk, to name a few. Especially, XGBoost, one of the most prominent and widely used decision trees, has gained popularity due to various optimizations on gradient boosting framework that provides increased accuracy for classification and regression problems. XGBoost’s ability to train relatively faster, handling missing values, flexibility and parallel processing make it a better candidate to handle data center workload. Today’s data centers with enormous Input/Output Operations per Second (IOPS) demand a real-time accelerated inference with low latency and high throughput because of significant data processing due to applications such as ransomware detection or fraud detection.This paper showcases an FPGA-based XGBoost accelerator designed with High-Level Synthesis (HLS) tools and design flow accelerating binary classification inference. We employ Alveo U50 and U200 to demonstrate the performance of the proposed design and compare it with existing state-of-the-art CPU (Intel Xeon E5-2686 v4) and GPU (Nvidia Tensor Core T4) implementations with relevant datasets. We show a latency speedup of our proposed design over state-of-art CPU and GPU implementations, including energy efficiency and cost-effectiveness. The proposed accelerator is up to 65.8x and 5.3x faster, in terms of latency than CPU and GPU, respectively. The Alveo U50 is a more cost-effective device, and the Alveo U200 stands out as more energy-efficient.
Archit Gajjar, Priyank Kashyap, Aydin Aysu, Paul D. Franzon, Sumon Dey, Chris Cheng
FCCM4
2022 Hardware Implementation of Hierarchical Temporal Memory Algorithm
abstract
Hierarchical temporal memory (HTM) is an un-supervised machine learning algorithm that can learn both spatial and temporal information of input. It has been successfully applied to multiple areas. In this paper, we propose a multi-level hierarchical ASIC implementation of HTM, referred to as processor core, to support both spatial and temporal pooling. To improve the unbalanced workload in HTM, the proposed design provides different mapping methods for the spatial and temporal pooling, respectively. In the proposed design, we implement a distributed memory system by assigning one dedicated memory bank to each level of hierarchy to improve the memory bandwidth utilization efficiency. Finally, the hot-spot operations are optimized using a series of customized units. Regarding scalability, we propose a ring-based network consisting of multiple processor cores to support a larger HTM network. To evaluate the performance of our proposed design, we map an HTM network that includes 2,048 columns and 65,536 cells on both the proposed design and NVIDIA Tesla K40c GPU using the KTH database as input. The latency and power of the proposed design is 6.04 ms and 4.1 W using GP 65 nm technology. Compared to the equivalent GPU implementation, the latency and power is improved 12.45× and 57.32×, respectively.
Weifu Li, Paul D. Franzon, Sumon Dey, Joshua Schabel
ACM J. Emerg. Technol. Comput. Syst.2
2022 Can Higher-Order Mutants Improve the Performance of Mutation-Based Fault Localization?
abstract
First-order mutants (FOMs) have been widely used in mutation-based fault localization (MBFL) approaches and have achieved promising results in single-fault localization scenarios (SFL-scenario). Higher-order mutants (HOMs) are proposed to simulate complex faults and can be applied in MBFL theoretically for multiple-fault localization scenarios (MFL-scenario). However, whether HOMs can improve MBFL’s performance is not investigated and the effectiveness is not thoroughly evaluated. In this empirical study, we investigate the impact of HOMs on the performance of MBFL in SFL-scenario and MFL-scenario. The experiments on two real-world benchmarks reveal that 1) 2-HOMs can help improve the MBFL performance in SFL-scenarios; 2) in MFL-scenarios, both 2-HOMs and 3-HOMs can achieve better performance than FOMs; and 3) huge computational cost cannot be ignored in the practice of HOMs. Therefore, effective methods to reduce the number of HOMs for future MBFL studies should be considered.
Zheng Li 0002, Yong Liu 0030, Xiang Chen 0005, Paul D. Franzon, Yuxiaoyang Cai, Luxi Fan
IEEE Trans. Reliab.5
2022 Design Obfuscation Through 3-D Split Fabrication With Smart Partitioning
abstract
We describe a design and fabrication experiment that has been performed to investigate a methodology for assessing the security of application specific integrated circuits (ASICs) fabricated in a split-manufacturing process based on 3-D integrated circuit (3DIC) technologies. The purpose of this process is to protect critical IP from reverse engineering if an adversary obtains either the fabricated wafers or their GDS. A number of 3DIC-based fabrication alternatives were evaluated, and one is selected for this experiment. Several designs, from the trivial to the complex, were used for the study. A self-test module was embedded in each design to facilitate the postfabrication testing. Various obfuscation techniques that include camouflage in the form of function and lookup table hiding and insertion of redundant logic in order to confuse potential attackers were applied. Smart partitioning was implemented for each design in an attempt to conceal vital functions. We introduced metrics that are based on the number of connection possibilities ($C_{p}$) and the depth of partitioning ($P_{\mathrm{ depth}}$) to measure the obfuscation strength. The results show that it should take more than 1060years to reconstruct the netlist using a brute-force attack. Measurement results are presented showing fabrication success.
Theodros Nigussie, Joshua Schabel, Steve Lipa, Lisa G. McIlrath, Robert Patti, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.6
2021 Fast and Accurate PPA Modeling with Transfer Learning
abstract
The power, performance and area (PPA) of digital blocks can vary 10:1 based on their synthesis, place, and route tool recipes. With rapid increase in number of PVT corners and complexity of logic functions approaching 10M gates, industry has an acute need to minimize the human resources, compute servers, and EDA licenses needed to achieve a Pareto optimal recipe. We first present models for fast accurate PPA prediction that can reduce the manual optimization iterations with EDA tools. Secondly we investigate techniques to automate the PPA optimization using evolutionary algorithms. For PPA prediction, a baseline model is trained on a known design using Latin hypercube sample runs of the EDA tool, and transfer learning is then used to train the model for an unseen design. For a known design the baseline needed 150 training runs to achieve a 95% accuracy. With transfer learning the same accuracy was achieved on a different (unseen) design in only 15 runs indicating the viability of transfer learning to generalize PPA models. The PPA optimization technique, based on evolutionary algorithms, effectively combines the PPA modeling and optimization. Our approach reached the same PPA solution as human designers in the same or fewer runs for a CORTEX-M0 system design. This shows potential for automating the recipe optimization without needing more runs than a human designer would need.
William Rhett Davis, Paul D. Franzon, Luis Francisco, Billy Huggins, Rajeev Jain
ICCAD2
2021 A Scalable Cluster-based Hierarchical Hardware Accelerator for a Cortically Inspired Algorithm
abstract
This article describes a scalable, configurable and cluster-based hierarchical hardware accelerator through custom hardware architecture for Sparsey, a cortical learning algorithm. Sparsey is inspired by the operation of the human cortex and uses a Sparse Distributed Representation to enable unsupervised learning and inference in the same algorithm. A distributed on-chip memory organization is designed and implemented in custom hardware to improve memory bandwidth and accelerate the memory read/write operations for synaptic weight matrices. Bit-level data are processed from distributed on-chip memory and custom multiply-accumulate hardware is implemented for binary and fixed-point multiply-accumulation operations. The fixed-point arithmetic and fixed-point storage are also adapted in this implementation. At 16 nm, the custom hardware of Sparsey achieved an overall 24.39× speedup, 353.12× energy efficiency per frame, and 1.43× reduction in silicon area against a state-of-the-art GPU.
Sumon Dey, Lee Baker, Joshua Schabel, Weifu Li, Paul D. Franzon
ACM J. Emerg. Technol. Comput. Syst.5
2021 2Deep: Enhancing Side-Channel Attacks on Lattice-Based Key-Exchange via 2-D Deep Learning
abstract
Advancements in quantum computing present a security threat to classical cryptography algorithms. Lattice-based key exchange protocols show strong promise due to their resistance to theoretical quantum-cryptanalysis and low implementation overhead. By contrast, their physical implementations have shown vulnerability against side-channel attacks (SCAs) even with a single power measurement. The state-of-the-art SCAs are, however, limited to simple, sequentialized executions of post-quantum key-exchange (PQKE) protocols, leaving the vulnerability of complex, parallelized architectures unknown. This article proposes 2Deep-a deep-learning (DL)-based SCA-targeting parallelized implementations of PQKE protocols, namely, Frodo and NewHope with data augmentation techniques. Specifically, we explore approaches that convert 1-D time-series power measurement data into 2-D images to formulate SCA an image recognition task. The results show our attack's superiority over conventional techniques including horizontal differential power analysis (DPA), template attacks (TAs), and straightforward DL approaches. We demonstrate improvements up to 1.5× to recover a 100% success rate compared to DL with 1-D input data while using fewer data. We furthermore show that machine learning improves the results up to 1.25× compared to TAs. Furthermore, we perform cross-device attacks that obtain profiles from a single device, which has never been explored. Our 2-D approach is especially favored in this setting, improving the success rate of attacking Frodo from 20% to 99% compared to the 1-D approach. Our work thus urges countermeasures even on parallel architectures and single-trace attacks.
Priyank Kashyap, Furkan Aydin, Seetal Potluri, Paul D. Franzon, Aydin Aysu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 Machine Learning and Hardware security: Challenges and Opportunities -Invited Talk-
abstract
Machine learning techniques have significantly changed our lives. They helped improving our everyday routines, but they also demonstrated to be an extremely helpful tool for more advanced and complex applications. However, the implications of hardware security problems under a massive diffusion of machine learning techniques are still to be completely understood. This paper first highlights novel applications of machine learning for hardware security, such as evaluation of post quantum cryptography hardware and extraction of physically unclonable functions from neural networks. Later, practical model extraction attack based on electromagnetic side-channel measurements are demonstrated followed by a discussion of strategies to protect proprietary models by watermarking them.
Francesco Regazzoni 0001, Shivam Bhasin, Amir Ali Pour, Ihab Alshaer, Furkan Aydin, Aydin Aysu, Vincent Beroulle, Giorgio Di Natale, Paul D. Franzon, David Hély, Naofumi Homma, Akira Ito 0002, Dirmanto Jap, Priyank Kashyap, Ilia Polian, Seetal Potluri, Rei Ueno, Elena I. Vatajelu, Ville Yli-Mäyry
ICCAD9
2020 Multi-Fidelity Surrogate-Based Optimization for Electromagnetic Simulation Acceleration
abstract
As circuits’ speed and frequency increase, fast and accurate capture of the details of the parasitics in metal structures, such as inductors and clock trees, becomes more critical. However, conducting high-fidelity 3D electromagnetic (EM) simulations within the design loop is very time consuming and computationally expensive. To address this issue, we propose a surrogate-based optimization methodology flow, namely multi-fidelity surrogate-based optimization with candidate search (MFSBO-CS), which integrates the concept of multi-fidelity to reduce the full-wave EM simulation cost in analog/RF simulation-based optimization problems. To do so, a statistical co-kriging model is adapted as the surrogate to model the response surface, and a parallelizable perturbation-based adaptive sampling method is used to find the optima. Within the proposed method, low-fidelity fast RC parasitic extraction tools and high-fidelity full-wave EM solvers are used together to model the target design and then guide the proposed adaptive sample method to achieve the final optimal design parameters. The sampling method in this work not only delivers additional coverage of design space but also helps increase the accuracy of the surrogate model efficiently by updating multiple samples within one iteration. Moreover, a novel modeling technique is developed to further improve the multi-fidelity surrogate model at an acceptable additional computation cost. The effectiveness of the proposed technique is validated by mathematical proofs and numerical test function demonstration. In this article, MFSBO-CS has been applied to two design cases, and the result shows that the proposed methodology offers a cost-efficient solution for analog/RF design problems involving EM simulation. For the two design cases, MFSBO-CS either reaches comparably or outperforms the optimization result from various Bayesian optimization methods with only approximately one- to two-thirds of the computation cost.
Paul D. Franzon, David Smart, Brian Swahn
ACM Trans. Design Autom. Electr. Syst.2
2017 H3 (Heterogeneity in 3D): A Logic-on-Logic 3D-Stacked Heterogeneous Multi-Core Processor
abstract
A single-ISA heterogeneous multi-core processor(HMP) [2], [7] is comprised of multiple core types that all implement the same instruction-set architecture (ISA) but have different microarchitectures. Performance and energy is optimized by migrating a thread's execution among core types as its characteristics change. Simulation-based studies with two core types, one simple (low power) and the other complex (high performance), has shown that being able to switch cores as frequently as once every 1,000 instructions increases energy savings by 50% compared to switching cores once every 10,000 instructions, for the same target performance [10]. These promising results rely on extremely low latencies for thread migration. Here we present the H3 chip that uses 3D die stacking and novel microarchitecture to implement a heterogeneous multi-core processor (HMP) with low-latency fast thread migration capabilities. We discuss details of the H3 design and present power and performance results from running various benchmarks on the chip. The H3 prototype can reduce power consumption of benchmarks by up to 26%.
Vinesh Srinivasan, Rangeen Basu Roy Chowdhury, Elliott Forbes, Randy Widialaksono, Zhenqian Zhang, Joshua Schabel, Sungkwan Ku, Steve Lipa, Eric Rotenberg, William Rhett Davis, Paul D. Franzon
ICCD11
2017 Corrections to "Crosstalk-Canceling Multimode Interconnect Using Transmitter Encoding"
abstract
The authors of[1]would like to note the following corrections in reference numbering. It is difficult to find correct references in the currently published paper due to the reference discords.
HoonSeok Kim, Chanyoun Won, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.3
2016 A Generally Applicable Calibration Algorithm for Digitally Reconfigurable Self-Healing RFICs
abstract
A generally applicable calibration technique for digitally reconfigurable self-healing radio frequency integrated circuits based on a hybrid of the Nelder-Mead and Hooke-Jeeves direct search algorithms is presented. The proposed algorithm is applied to the multiobjective problem of gain error and phase error minimization for a self-healing phase rotator test case. For the 8-D phase rotator calibration problem, we show that the proposed hybrid Nelder-Mead and Hooke-Jeeves calibration algorithm is capable of reducing the gain error and phase error of the phase rotator output to less than a maximum of 0.5 dB and 2°, respectively, relative to the chosen gain and phase targets. A 3-GHz self-healing phase rotator test chip was fabricated in a 45-nm silicon-on-insulator CMOS process, and the measured data were obtained to validate the performance of the proposed calibration algorithm.
Eric J. Wyers, Matthew A. Morton, T. C. L. Gerhard Sollner, C. T. Kelley, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Under 100-cycle thread migration latency in a single-ISA heterogeneous multi-core processor
Elliott Forbes, Zhenqian Zhang, Randy Widialaksono, Brandon H. Dwiel, Rangeen Basu Roy Chowdhury, Vinesh Srinivasan, Steve Lipa, Eric Rotenberg, William Rhett Davis, Paul D. Franzon
Hot Chips Symposium10
2014 A Generic and Scalable Architecture for a Large Acoustic Model and Large Vocabulary Speech Recognition Accelerator Using Logic on Memory
abstract
This paper describes a scalable hardware accelerator for speech recognition, which uses a two pass decoding algorithm with word dependent N-best Viterbi Beam Search. The observation probability calculation (Senone scoring) and first pass of decoding using a Bigram language model is implemented in hardware. The word lattice output from the first pass is used by software for the second pass, with a trigram language model. The proposed design uses a logic-on-memory approach to make use of high bandwidth nor flash memory to improve random read performance for Senone scoring and first pass decoding, both of which are memory intensive operations. The proposed HW/SW co-design achieves an overall speed up of 4.3X over a 2.4-GHz Intel Core 2 Duo processor running the CMU Sphinx speech recognition software, while consuming an estimated 1.72 W of power. The hardware accelerator provides improved speech recognition accuracy by supporting larger acoustic models and word dictionaries while maintaining real-time performance.
Ojas A. Bapat, Paul D. Franzon, Richard M. Fastow
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Rationale for a 3D heterogeneous multi-core processor
abstract
Single-ISA heterogeneous multi-core processors are comprised of multiple core types that are functionally equivalent but microarchitecturally diverse. This paradigm has gained a lot of attention as a way to optimize performance and energy. As the instruction-level behavior of the currently executing program varies, it is migrated to the most efficient core type for that behavior.
Eric Rotenberg, Brandon H. Dwiel, Elliott Forbes, Zhenqian Zhang, Randy Widialaksono, Rangeen Basu Roy Chowdhury, Nyunyi M. Tshibangu, Steve Lipa, William Rhett Davis, Paul D. Franzon
ICCD10
2013 Exploring early design tradeoffs in 3DIC
abstract
This The key to gaining substantial benefit from the use of 3DIC technology is to create 3D specific designs that do more than recast a 2D optimal design into the third dimension. This paper explores some of the approaches to creating 3D specific designs and the CAD tools that can help in that exploration. The power advantages of 3D design are illustrated in details. Results from different partitioning approaches (function, modular and circuit) are presented, together with early results from a thermal pathfinding tool.
Paul D. Franzon, Shivam Priyadarshi, Steve Lipa, William Rhett Davis, Thorlindur Thorolfsson
ISCAS1
2013 Crosstalk-Canceling Multimode Interconnect Using Transmitter Encoding
abstract
A new implementation approach to cancel crosstalk using modal decomposition on a multiconductor transmission bundle is presented. The proposed approach requires a CODEC only at the transmitter, not at both the transmitter and receiver. This gives potential for more flexibility, lower power, better scaling, and ease of implementation. A circuit is presented along with the simulation results.
HoonSeok Kim, Chanyoun Won, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.3
2012 A novel double floating-gate unified memory device
abstract
A novel double floating-gate unified memory device is experimentally demonstrated for the first time. The device can be used to store both volatile and nonvolatile memory states simultaneously. Simulations of scaled devices show that the device offers several advantages compared to conventional memory devices. Such a device could have a dramatic impact on next generation memory architectures.
Neil Di Spigna, Daniel Schinke, Srikant Jayanti, Veena Misra, Paul D. Franzon
VLSI-SoC5
2012 Comparing Through-Silicon-Via (TSV) Void/Pinhole Defect Self-Test Methods
Yi Lou, Zhuo Yan, Paul D. Franzon
J. Electron. Test.4
2012 Junction-Level Thermal Analysis of 3-D Integrated Circuits Using High Definition Power Blurring
abstract
The degraded thermal path of 3-D integrated circuits (3DICs) makes thermal analysis at the chip-scale an essential part of the design process. Performing an appropriate thermal analysis on such circuits requires a model with junction-level fidelity; however, the computational burden imposed by such a model is tremendous. In this paper, we present enhancements to two thermal modeling techniques for integrated circuits to make them applicable to 3DICs. First, we present a resistive mesh-based approach that improves on the fidelity of prior approaches by constructing a thermal model of the full structure of 3DICs, including the interconnect. Second, we introduce a method for dividing the thermal response caused by a heat load into a high fidelity “near response” and a lower fidelity “far response” in order to implement Power Blurring high definition (HD), a hierarchical thermal simulation approach based on Power Blurring that incorporates the resistive mesh-based models and allows for junction-level accuracy at the full-chip scale. The Power Blurring HD technique yields approximately three orders of magnitude of improvement in memory usage and up to six orders of magnitude of improvement in runtime for a three-tier synthetic aperture radar circuit, as compared to using a full-chip junction-scale resistive mesh-based model. Finally, measurement results are presented showing that Power Blurring high definition (HD) accurately determines the shape of the thermal profile of the 3DIC surface after a correction factor is added to adjust for a discrepancy in the absolute temperature values.
Samson Melamed, Thorlindur Thorolfsson, T. Robert Harris, Shivam Priyadarshi, Paul D. Franzon, Michael B. Steer, William Rhett Davis
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2012 Parallel Transient Simulation of Multiphysics Circuits Using Delay-Based Partitioning
abstract
A parallel transient simulation technique for multiphysics circuits is presented. The technique develops partitions utilizing the inherent delay present within a circuit and between physical domains. A state-variable-based circuit delay element is presented, which implements the coupling between two spatially or temporally isolated circuit partitions. A parallel delay-based iterative approach for interfacing delay-partitioned subcircuits is applied, which achieves the reasonable accuracy of nonparallel circuit simulation if both incorporate the same interblock delay. The partitioned subcircuits are distributed to different cores of a shared-memory multicore processor and solved in parallel. A multithreaded implementation of the methodology using OpenMP is presented. Examples showing superlinear speedup compared to unpartitioned single-core simulation using the direct method are presented. This paper also discusses the impact of load balancing and absolute delay on simulation speedup.
Shivam Priyadarshi, Christopher S. Saunders, Nikhil Kriplani, Harun Demircioglu, William Rhett Davis, Paul D. Franzon, Michael B. Steer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2011 3D Specific Systems: Design and CAD
abstract
3D stacking and integration can provide significant system advantages. Following a brief technology review, this abstract explores application drivers, design and CAD for 3D ICs. The main 3D exploitation explored in detail is that of logic on memory. This application is explored in a specific DSP example, showing a 25% power advantage when implemented in 3D compared with 2D. Finally critical areas that need better solutions are explored. These include cost management, design planning, test management, and thermal management.
Paul D. Franzon, William Rhett Davis, Thorlindur Thorolfsson, Samson Melamed
Asian Test Symposium1
2011 Low power interconnect design for fpgas with bidirectional wiring using nanocrystal floating gate devices (abstract only)
abstract
New architectures for the switch box and connection block are proposed for use in an energy efficient field programmable gate array (FPGA) with bidirectional wiring. Power-hungry SRAMs are replaced by non-volatile nanocrystal floating gate (NCFG) devices that retain their state while the system power is off and do not need to be configured at boot up. The NCFG-based FPGA is benchmarked against both a traditional bidirectional and a modern unidirectional SRAM-based FPGA using a 32-tap FIR Filter designed in HSPICE based on predictive BSIM4.0 CMOS with 45nm gate length technology and a previously developed physical model of the NCFG device. Compared to the traditional bidirectional and the modern unidirectional SRAM-based interconnect the total gate area is reduced by 87% and 63%, respectively. Simulations demonstrate a reduction of 58% in static and 34% in dynamic power consumption compared to the traditional bidirectional SRAM-based FPGA while the signal propagation delay through a switch box is decreased by 28%. When compared to the modern unidirectional SRAM-based FPGA the proposed design has roughly comparable power consumption but the circuit complexity is greatly reduced as a result of doubling the available routing channels. Alternatively the number of the routing channels may be reduced to save area and power whereas the complexity remains similar. The potential benefits from choosing the proposed design can be summarized as small area, low power consumption, high speed and high functionality, which typically trade off and cannot be achieved by the SRAM-based counterparts simultaneously. Compared to previous designs that use continuous floating gate devices in FPGAs, the approach described in this work requires less overhead, lower voltages, and offers improved reliability.
Daniel Schinke, Wallace Shep Pitts, Neil Di Spigna, Paul D. Franzon
FPGA4
2011 Application of Surrogate Modeling in Variation-aware Macromodel and Circuit Design
Ting Zhu 0002, Mustafa Berke Yelten, Michael B. Steer, Paul D. Franzon
SIMULTECH4
2010 Creating 3D specific systems: Architecture, design and CAD
abstract
3D stacking and integration can provide system advantages. Following a brief technology review, this abstract explores application drivers, design and CAD for 3D ICs. The main application area explored in detail is that of logic on memory. This application is explored in a specific DSP example. Finally critical areas that need better solutions are explored. These include design planning, test management, and thermal management.
Paul D. Franzon, William Rhett Davis, Thorlindur Thorolfsson
DATE1
2010 Low-Power Hypercube Divided Memory FFT Engine Using 3D Integration
abstract
In this article we demonstrate a floating point FFT processor that leverages both 3D integration and a unique hypercube memory division scheme to reduce the power consumption of a 1024 point FFT down to 4.227 μJ . The hypercube memory division scheme lowers the energy per memory access by 59.2% and increases the total required area by 16.8%. The use of 3D integration reduces the logic power by 5.2%. We describe the tool flow required to realize the 3D implementation and perform a thermal analysis of it.
Thorlindur Thorolfsson, Samson Melamed, William Rhett Davis, Paul D. Franzon
ACM Trans. Design Autom. Electr. Syst.4
2009 Design automation for a 3DIC FFT processor for synthetic aperture radar: a case study
abstract
This work discusses a 1024-point, memory-on-logic 3DIC FFT processor for synthetic aperture radar (SAR), sent to fabrication in the 180 nm MIT Lincoln Labs 3D FDSOI 1.5 V process[12] along with the design flow required to realize it with off-the-shelf commercial 2D tools. The work shows how the vertical dimension can be exploited for novel memory architecture tradeoffs that are not feasible in 2D, reducing the energy consumed per memory operation in the FFT by 60.3%. In comparison to its 2D counterpart, the SAR FFT processor exhibits a 53.0% decrease in average wire length, a 24.6% increase in maximum operating frequency and a 25.3% decrease in total silicon area.
Thorlindur Thorolfsson, Kiran Gonsalves, Paul D. Franzon
DAC3
2009 A low power 3D integrated FFT engine using hypercube memory division
abstract
In this paper we demonstrate a floating point FFT processor that leverages both 3D integration and a hypercube memory division scheme to reduce the power consumption of a 1024 point FFT down to 4.227 μJ. The hypercube memory division scheme lowers the energy per memory access by 59.2% while only increasing the total area required by 16.8%, while using 3D integration reduces the logic power by 5.2%. For comparison, we analyze the amount of power and wire length reduction that can be expected from 3D integration for normal digital logic circuits.
Thorlindur Thorolfsson, Nariman Moezzi Madani, Paul D. Franzon
ISLPED3
2009 Application Exploration for 3-D Integrated Circuits: TCAM, FIFO, and FFT Case Studies
abstract
3-D stacking and integration can provide system advantages. This paper explores application drivers and computer-aided design (CAD) for 3-D integrated circuits (ICs). Interconnect-rich applications especially benefit, sometimes up to the equivalent of two technology nodes. This paper presents physical-design case studies of ternary content-addressable memories (TCAMs), first-in first-out (FIFO) memories, and a 8192-point fast Fourier transform (FFT) processor in order to quantify the benefit of the through-silicon vias in an available 180-nm 3-D process. The TCAM shows a 23% power reduction and the FFT shows a 22% reduction in cycle-time, coupled with an 18% reduction in energy per transform.
William Rhett Davis, Eun Chu Oh, Ambarish M. Sule, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.4
2009 A 32-Gb/s On-Chip Bus With Driver Pre-Emphasis Signaling
abstract
This paper describes a differential current-mode bus architecture based on driver pre-emphasis for on-chip global interconnects that achieves high-data rates while reducing bus power dissipation and improving signal delay latency. The 16-b bus core fabricated in 0.25-mum complementary metal-oxide-semiconductor (CMOS) technology attains an aggregate signaling data rate of 32 Gb/s over 5-10-mm-long lossy interconnects. With a supply of 2.5 V, 25.5-48.7-mW power dissipation was measured for signal activity above 0.1, equivalent to 0.80-1.52 pJ/b. This work demonstrates a 15.0%-67.5% power reduction over a conventional single-ended voltage-mode static bus while reducing delay latency by 28.3% and peak current by 70%. The proposed bus architecture is robust against crosstalk noise and occupies comparable routing area to a reference static bus design.
Liang Zhang 0038, John M. Wilson 0002, Rizwan Bashirullah, Lei Luo 0006, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.6
2008 Design and CAD for 3D integrated circuits
abstract
High density Through Silicon Vias (TSV) can be used to build 3DICs that enable unique applications in computing, signal processing and memory intensive systems. This paper presents several case studies that are uniquely enhanced through 3D implementation, including a 3D CAM, an FFT processor, and a SAR processor. The CAD flow used to implement for these designs is described. 3DIC requires higher fidelity thermal modeling than 2DIC design. The rationale for this requirement is established and a possible solution is presented.
Paul D. Franzon, William Rhett Davis, Michael B. Steer, Steve Lipa, Eun Chu Oh, Thorlindur Thorolfsson, Samson Melamed, Sonali Luniya, Tad Doxsee, Stephen Berkeley, Ben Shani, Kurt Obermiller
DAC1
2008 Keeping hot chips cool: are IC thermal problems hot air?
abstract
Thermal issues are becoming more important but is the hype getting the better of the facts? Does this deserve more attention than for some niche designs and technologies such as 3D ICs.? Does the broader design community need to worry about it at 32nm and beyond or it will only impact a small segment of designs? In short, does the severity of power issues coupled with packaging complexity translate into a thermal crisis in future? This is an educational panel with a little bit of controversy that will address the thermal issue in IC design. When will this issue be emerging as a crucial concern if at all? What are the solutions to resolve this potential crisis?
Ruchir Puri, Devadas Varma, Darvin Edwards, Alan J. Weger, Paul D. Franzon, Stephen V. Kosonocky
DAC5
2008 Editorial: Special issue on 3D integrated circuits and microarchitectures
abstract
No abstract available.
Yuan Xie 0001, Jason Cong, Paul D. Franzon
ACM J. Emerg. Technol. Comput. Syst.3
2007 Flexible Low Power Probability Density Estimation Unit For Speech Recognition
abstract
This paper describes the hardware architecture for a flexible probability density estimation unit to be used in a large vocabulary speech recognition system, and targeted for mobile platforms. The speech recognition system is based on hidden Markov models and consists of two computationally intensive parts - the probability density estimation using Gaussian distributions, and the Viterbi decoding. The power hungry nature of these computations prevents porting the application successfully to mobile devices. We have designed a flexible probability estimation unit that is both power efficient and meets real time requirements while being flexible enough to handle emerging speech recognition techniques. The flexible nature of the design allows it to utilize emerging power and computation reduction techniques (at the algorithm level) to achieve up to an 80% power reduction as compared to conventional designs
Ullas Pazhayaveetil, Dhruba Chandra, Paul D. Franzon
ISCAS3
2007 Hardware Architecture of a Parallel Pattern Matching Engine
abstract
Several network security and QoS applications require detecting multiple string matches in the packet payload by comparing it against predefined pattern set. This process of pattern matching at line speeds is a memory and computation intensive task. Hence, it requires dedicated hardware algorithms. This paper describes the hardware architecture of a parallel, pipelined pattern matching engine that uses trie based pattern matching algorithmic approach. The algorithm optimizes pattern matching process through two key innovations of parallel pattern matching using incoming content filter and multiple character matching using trie pruning. The hardware implementation is capable of performing at line-speeds and handle traffic rates up to OC-192, the underlying architecture allows for multiple patterns to be detected and for the system to gracefully recover from a failed partial match, the throughput of the system does not degrade with the increase in the number of patterns or the length of the patterns to be matched. The solution described outperforms most current implementations in terms of speed and memory requirement and outperforms TCAM based solutions in terms of power consumption, area, and cost while remaining competitive in terms of throughput and update times. The complete Snort rule set (2005 release) and VoIP RFC were used to validate our performance and achieve a throughput of 12Gbps with 6KBytes of content filter memory and 0.3 MBytes of total memory for Snort and 0.5KBytes of filter memory and 12KBytes of total memory for SIP.
Meeta Yadav, Ashwini Venkatachaliah, Paul D. Franzon
ISCAS3
2007 Voltage-Mode Driver Preemphasis Technique For On-Chip Global Buses
abstract
This paper demonstrates that driver preemphasis technique can be used for on-chip global buses to increase signal channel bandwidth. Compared to conventional repeater insertion techniques, driver preemphasis saves repeater layout complexity and reduces power consumption by 12%-39% for data activity factors above 0.1. A driver circuit architecture using voltage-mode preemphasis technique was tested in 0.18-$\mu$m CMOS technology for 10-mm long interconnects at 2 Gb/s.
Liang Zhang 0038, John M. Wilson 0002, Rizwan Bashirullah, Lei Luo 0006, Paul D. Franzon
IEEE Trans. Very Large Scale Integr. Syst.6
2005 Driver pre-emphasis techniques for on-chip global buses
abstract
By using current-sensing differential buses with driver pre-emphasis techniques, power dissipation is reduced by 26.0% - 51.2% and peak current is reduced by 63.8%, compared to conventional repeater insertion techniques, for 10mm long buses in TSMC 0.25μm technology. This proposed architecture lowers the worst coupling capacitance to total capacitance ratio to 14.4%. It only requires 7.9% more bus routing area than single-ended designs for a 16-bit bus, and saves all of the repeater placement blockages. To further verify that the driver pre-emphasis techniques can also be applied to voltage-mode single-ended buses, a test chip in TSMC 0.18μm technology was fabricated and measured
Liang Zhang 0038, John M. Wilson 0002, Rizwan Bashirullah, Lei Luo 0006, Paul D. Franzon
ISLPED6
2005 Molecular Electronics - Devices and Circuits Technology
Paul D. Franzon, David Nackashi, Christian Amsinck, Neil Di Spigna, Sachin Sonkusale
VLSI-SoC1
2004 Simplified delay design guidelines for on-chip global interconnects
abstract
Based on the effective attenuation constant approximation of distributed RLC lines, simplified design guidelines are presented dealing with the line characteristics, termination, and delay estimation of on-chip global interconnects. RC delay models are verified to be still accurate for a wide range of parameters conventionally considered inductive. A new closed-form RLC delay formula is developed when RC models are inadequate. The formula works for both voltage and current-mode signaling and exhibits 10% accuracy of SPICE simulation. This work is suitable for global routing topologies and iterative layout optimization.
Liang Zhang 0038, Wentai Liu, Rizwan Bashirullah, John M. Wilson 0002, Paul D. Franzon
ACM Great Lakes Symposium on VLSI5
2004 The Design, Fabrication, and Characterization of Millimeter Scale Motors for Miniature Direct Drive Robots
abstract
This paper reports on research into miniature, direct drive, high force/torque motors to support insect-sized mobile robotic platforms. The primary focus is on scalable motors based on piezoelectric transducers. The contributions of this work include: (1) the design, analysis, and characterization of a miniature mode conversion rotary ultrasonic motor based on a piezoelectric stack transducer; this produced a static torque density of 0.37 Nm/kg, (2) a millimeter scale linear piezometer, constructed with a parallel arrangement of annular stressed unimorph piezoelectric transducers and passive latches, exhibited 0.23 N of blocked force, and (3) simulation data is presented that compares these motor concepts to commercial systems in the context of scalability. Results suggest that smaller versions of the rotary ultrasonic motor would possess a static torque density seven times that of a commercial 3-mm electromagnetic system. This technology shows promise for driving the platform.
J. A. Palmer, James F. Mulling, Brian Dessent, Edward Grant, Jeffrey W. Eischen, Alexei Gruverman, A. I. Kingon, Paul D. Franzon
ICRA8
2003 Molecular electronics: from devices and interconnect to circuits and architecture
abstract
As the dominating CMOS technology is fast approaching a "brick wall," new opportunities arise for competing solutions. Nanoelectronics has achieved several breakthroughs lately and promises to overcome many of the limitations intrinsic to current semiconductor approaches. Most of the results in this area reported until now focus on devices and interconnect; this work goes several steps further and presents issues related to circuits and architecture. Based on proposed nanoscale interconnect and device structures, we explore the design space available to the nanoelectronic circuit designer and system architect.
Mircea R. Stan, Paul D. Franzon, Seth Copen Goldstein, John C. Lach, Matthew M. Ziegler
Proc. IEEE2
2002 Binary search schemes for fast IP lookups
abstract
Route lookup is becoming a very challenging problem due to the increasing size of routing tables. To determine the outgoing port for a given address, the longest matching prefix among all the prefixes, needs to be determined. This makes the task of searching in a large database quite difficult. Our paper describes binary search schemes that allow fast address lookups. Binary search can be performed on the number of entries or on the number of mutually disjoint prefixes. Lookups can be performed in O(N) time, where N is number of entries and the amount of memory required to store the binary database is also O(N). These schemes scale very well with both large databases and for longer addresses (as in IPv6).
Pronita Mehrotra, Paul D. Franzon
GLOBECOM2
2001 Will Nanotechnology Change the Way We Design and Verify Systems? (Panel)
Andreas Kuehlmann, Robert W. Dutton, Paul D. Franzon, Seth Copen Goldstein, Philip Luekes, Eric Parker, Thomas N. Theis
ICCAD3
1999 Parasitic Extraction Accuracy - How Much is Enough?
abstract
No abstract available.
Paul D. Franzon, Mark Basel, Aki Fujimara, Sharad Mehrotra, Ron Preston, Robin C. Sarma, Marty Walker
DAC1
1999 Dynamically Programmable Cache Evaluation and Virtualization
abstract
No abstract available.
Mouna Nakkar, David G. Bentlage, John Harding, David Schwartz, Paul D. Franzon, Thomas M. Conte
FPGA5
1997 Low power data processing by elimination of redundant computations
abstract
We suggest a new technique to reduce energy consumption in the processor datapath without sacrificing performance by exploiting operand value locality at run time.Data locality is one of the major characteristics of video streams as well as other commonly used applications.We use a cache-like scheme to store a selective history of computation results, and the resultant Te-e-21se leads to power savings. The cache is indexed by the OpeTandS.Based on OUT model, an 8 to 128 entry execution cache TedUCeS power consumption by 20% to 60%.
Mir Azam, Paul D. Franzon, Wentai Liu
ISLPED2
1995 Performance Driven Global Routing and Wiring Rule Generation for High Speed PCBs and MCMs
abstract
A new approach for performance-driven routing in highly congested high speed MCMs and PCBs is presented.Global routing is employed to manage delay, signal integrity and congestion simultaneously.I n terconnect performance prediction models are generated through simulations.The global routing results and performance prediction models are used to generate bounds on the net lengths which can be used by a detailed router to satisfy constraints on interconnect performance.
Sharad Mehrotra, Paul D. Franzon, Michael B. Steer
DAC2
1994 Stochastic Optimization Approach to Transistor Sizing for CMOS VLSI Circuits
abstract
A stochastic global optimization approach is presented for transistor sizing in CMOS VLSI cir cuits.This is a direct search strategy for the best design among feasible ones, with the designer determining when the search is stopped.Through examples, we show the power of this technique in quickly obtaining very good designs, for skew minimization problems.
Sharad Mehrotra, Paul D. Franzon, Wentai Liu
DAC2
1993 A simple method for noise tolerance characterization of digital circuits
abstract
A method for characterizing dynamic noise tolerance of digital circuits is discussed. In this method noise impulses are characterized by their energy, voltage and width. The method is intended for use in simulation-based noise analysis and design of receiver circuits in digital systems.>
Slobodan Simovich, Paul D. Franzon, Michael B. Steer
Great Lakes Symposium on VLSI2
1993 System-Level Specification of Instruction Sets
abstract
System-level design requires some sort of specification for a system at the level of abstraction of the system. When the system (or sub-system) is a processor, the appropriate level of abstraction is the instruction set. However, there are no good approaches for describing processors at this level. Nevertheless, this type of specification has a number of benefits: it is more concise (and thus less error-prone) than more general alternatives; it can be re-used in later re-implementations; and it provides support for software codesign through compiler-generators (which rely on higher-level abstractions than other techniques provide). Therefore, we have developed a methodology and an embodying language for specifying processors at the instruction set level.>
Todd A. Cook, Paul D. Franzon, Ed Harcourt, Thomas K. Miller III
ICCD2
1992 Tools to Aid in Wiring Rule Generation for High Speed Interconnects
Paul D. Franzon, Slobodan Simovich, Michael B. Steer, Mark Basel, Sharad Mehrotra, Tom Mills
DAC1