VLDB 2026 Research / reviewers in the wild / expert
Ronald F. DeMara
dblp:d/RonaldFDeMara
· DBLP profile ↗
70ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-6864-7255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 55 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FPGA-based Acceleration of LLM Inference Using Compression-Based Similarity ClassificationabstractWe present Compression-Based Feature Clustering (CBFC), a low-footprint FPGA accelerator for training-free similarity inference in hybrid LLM pipelines. CBFC replaces the un-synthesizable gzip compressor with a fully HLS-synthesizable, fixed-resource LZ77 engine returning the deterministic scalar compressed lengths required by Normalized Compression Distance (NCD) k-NN classification. On a Zynq UltraScale+ at 300 MHz, CBFC achieves up to 3.41 × speedup over CPU gzip (70.6% latency reduction) while consuming only 2-12% of on-chip resources, leaving ample headroom for replication or co-location with other accelerator datapaths. Classification accuracy is on par with gzip across standard text benchmarks. Paul Amoruso, Richard C. Yarnell, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | DynARMic: A Dynamic ARM Instruction Counting ToolabstractDynARMic is a web-based dynamic instruction profiling tool for ARMv7 assembly for undergraduate computer architecture laboratory use. Students upload ARMv7 .s files; the tool emulates execution using the Keystone–Unicorn–Capstone stack and returns two analyses: (1) a five-category dynamic instruction breakdown aligned with per-instruction energy estimates, and (2) an eight-format ARMv7 machine encoding breakdown supporting cycles-per-instruction (CPI) estimation. Deployed as a public web service, DynARMic fills the gap left by CPUlator and is used in a required undergraduate course at the University of Central Florida. Ayush Pindoria, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Compression-Assisted Zero-Shot Prompting of Large Language Models (LLMs) for Educational Skill Classification of Microprocessor Curricula
Paul Amoruso, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | Interactive Framework for Cybersecurity Education and Future Workforce DevelopmentabstractThis research-to-practice paper presents a novel pedagogical tool for hardware cybersecurity education and workforce development. The growing importance of hardware security has made it essential for individuals and organizations to understand hardware security principles and best practices. However, the current educational curriculum falls short of fulfilling these emerging demands due to the rapidly changing hardware security landscape and limited opportunities for hands-on training. To address these challenges, we propose and have developed the Interactive Hardware and Cybersecurity (I-HaC) Educational Framework, a pedagogical educational framework that supplements existing courses by leveraging generative AI for individualized instruction related to hardware and cybersecurity, data mining, and applied Machine Learning (ML), as well as data visualization to enhance cybersecurity education and workforce development. The framework is designed to be utilized by graduate and undergraduate Electrical and Computer Engineering (ECE) and Computer Science (CS) students for a comprehensive introduction to cybersecurity exploits and countermeasures in an interactive manner with hands-on components. Using I-HaC, we have developed tailored lab components for a diverse range of students and intend to release I-HaC as open-source for the benefit of the ECE and CS education community. Sujan Ghimire, Md Muhtasim Alam Chowdhury, Ryan Tsang, Richard C. Yarnell, Emma Heckert, Jaeden Wolf Carpenter, Yu-Zheng Lin, Muntasir Mamun, Ronald F. DeMara, Setareh Rafatirad, Pratik Satam, Soheil Salehi |
FIE | 9 |
| 2024 | Educational Tool-spaces for Convolutional Neural Network FPGA Design Space Exploration Using High-Level SynthesisabstractThere is significant demand and urgency to prepare electrical and computer engineering students regarding the operational and performance characteristics of machine learning (ML) hardware accelerators. Convolutional Neural Networks (CNNs), which are utilized for real-time and large dataset image classification tasks, are appropriate targets for hardware acceleration. Designing accelerators for CNNs necessitates understanding the manipulation of CNN parameters. We introduce a hands-on pedagogy whereby learners can identify, modify, and appreciate the interaction of the CNN parameters within an interactive GUI. CASCADE (Computer Aided Student's CNN Analyzer for Design Exploration), a simulation-based framework for Design Space Exploration (DSE) of CNN FPGA-based accelerators is developed, including datapath synthesis, simulation, training, and testbench steps. We offer a case study of High-Level Synthesis (HLS) based CNN implementations targeting the MNIST dataset and present simulation results, namely hardware utilization, accuracy, and operating frequency, and offer insight into potential design trade-offs facing modern engineers. Richard C. Yarnell, Mousam Hossain, Raul Graterol, Ayush Pindoria, Sujan Ghimire, Md Muhtasim Alam Chowdhury, Soheil Salehi, Yu Bai 0004, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 9 |
| 2024 | FOCAL: Feature-Oriented Cellular Automata Learning for Convolution-Free Image ClassificationabstractState-of-the-art image classification systems utilize powerful machine-learning-based tools such as Convolutional Neural Networks (CNNs). These networks can achieve high recognition accuracies, but suffer from a black-box problem where the inner workings are incomprehensible by humans that seek to use them. In this paper, a Feature-Oriented Cellular Automata Learning (FOCAL) system is developed to extend traditional gradient-filter-based methods by implementing a Cellular Automata (CA) reasoner utilizing rule-based primitives for determining mutual agreement between neighboring pixels. This novel method is demonstrated to identify features more accurately than standard filter methods and produce classification results that are competitive with typical CNNs, while also allowing a-priori definition of important features facilitating explainable feature classification decision processes. Experiments spanning a variety of influential factors indicate that rebaselining and normalization are vital to the success of the CA-based approach. Furthermore, within certain models, the use of CA is shown to reduce computational demand by over 90% while incurring only a 2% reduction in classification accuracy. Finally, the scalability of the FOCAL system is investigated using the CIFAR-10 dataset and contemporary Deep Neural Networks, and shown to encourage promising avenues of research into explainability while reducing computational processing demands. Noah Ari, Richard C. Yarnell, Paul Amoruso, Johnathan Mell, Ronald F. DeMara, Annie S. Wu |
IS | 5 |
| 2023 | A Genetic Algorithm for Combinational Logic Circuit Synthesis Using Directed Graph PrimitivesabstractWe introduce functionality-cognizant Genetic Algorithms (GAs) and graph-based operators to tackle the challenging search landscape of combinational digital circuit design. We introduce a novel circuit representation that builds upon Cartesian Genetic Programming (CGP), a popular grid-based method for representing directed graphs of connected components. Leveraging this, we introduce an original crossover operator that accounts for circuit component functionality and connectivity, as opposed to CGP, which only considers positional information in the chromosome. Additionally, we propose an innovative set of mutation operators and demonstrate successful evolution of fully functional and minimally sized common digital circuits including a variety of binary encoders and adders. Following successful synthesis of a four-bit adder, we present a generalizable machine learning approach for multi-layered search and optimization problems. Richard C. Yarnell, Pierce Powell, Ronald F. DeMara, Annie S. Wu |
ICMLA | 3 |
| 2023 | Energy-/Area-Efficient Spintronic ANN-based Digit Recognition via Progressive Modular RedundancyabstractNeural networks offer viable alternatives for energy versus accuracy tradeoffs, in particular with regards to the precision of the computational circuit. This paper explores use of progressive modular redundancy of intrinsically low energy, low precision circuits as an alternative to more complex networks yielding higher accuracy directly. Results indicate that a lower footprint temporal modular redundancy, which is applied progressively as needed, can have lower footprint and reduced energy consumption at comparable or slightly reduced accuracy as more complex neural networks. This provides an alternative to binarization and other model compression options for intelligence at the edge of the network. Our Progressive Modular Redundancy approach using varied activations implemented using a$\mathbf{784}\times \mathbf{100}\times \mathbf{10}$network shows a 3% improvement in accuracy compared to the baseline case of$\mathbf{784}\times \mathbf{500}\times \mathbf{500}\times \mathbf{10}$network with sigmoidal activation, at 86.1% and 87% reduction in power and weighted crossbar normalized area overhead, respectively, 87.5% reduction in power error product (PEP) at the cost of ~2.6x increased throughput latency. Mousam Hossain, Adrian Tatulian, Harshavardhan Reddy Thummala, Ronald F. DeMara, Soheil Salehi |
ISCAS | 4 |
| 2022 | Nonuniform Compressive Sensing via Ohmic Voltage Attenuation: A Memristive Crossbar Design Approach Leveraging Intrinsic ComputationabstractCompressive sensing (CS) is a promising technique for transmitting signals in power-critical applications such as Internet of Things (IoT) devices. Nonuniform CS optimizes this process by adjusting sampling frequency based on the relative importance levels characterized by the signal of interest. Recent advances have yielded energy-efficient hardware implementations of CS sampling, leveraging spin-based crossbar architectures for in-memory vector-matrix multiplication and through the use of probabilistic bit (p-bit) devices to generate tunable random outputs for writing the array. Thus, a region of interest is generated via column-ordered density of on-state devices. Herein, we propose a simple design for supplying inputs to the p-bit devices, based on Ohmic voltage attenuation occurring along the word lines of the crossbar array. The technique embeds some required computations to be conducted intrinsically by the cross-points of the memristive array, thus bypassing overheads of conventional instruction execution and eliminating the need for costly hardware components, such as lookup tables (LUTs) and data converters. The design is shown to be robust for various array sizes and parasitics while generating the appropriate tuning signals within a single clock cycle duration of 1.6 ns, and at an energy overhead of 333 fJ. Compared with a standard approach using LUTs and digital to-analog converters, the design herein achieves a 583-fold reduction in energy and 23-fold reduction in transistor count. Adrian Tatulian, Ronald F. DeMara |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | An Efficient Real-Time Object Detection Framework on Resource-Constricted Hardware Devices via Software and Hardware Co-designabstractThe fast development of object detection techniques has attracted attention to developing efficient Deep Neural Networks (DNNs). However, the current state-of-the-art DNN models can not provide a balanced solution among accuracy, speed, and model size. This paper proposes an efficient real-time object detection framework on resource-constricted hardware devices through hardware and software co-design. The Tensor Train (TT) decomposition is proposed for compressing the YOLOv5 model. By unitizing the unique characteristics given by the TT decomposition, we develop an efficient hardware accelerator based on FPGA devices. Experimental results show that the proposed method can significantly reduce the model size and improve the execution time. Mingshuo Liu, Shiyi Luo, Kevin Han, Bo Yuan 0001, Ronald F. DeMara, Yu Bai 0004 |
ASAP | 5 |
| 2021 | An Efficient Video Prediction Recurrent Network using Focal Loss and Decomposed Tensor Train for Imbalance DatasetabstractNowadays, from companies to academics, researchers across the world are interested in developing recurrent neural networks due to their incredible feats in various applications, such as speech recognition, video detection, predictions, and machine translation. However, the advantages of recurrent neural networks accompanied by high computational and power demands, which are a major design constraint for electronic devices with limited resources used in such network implementations. Optimizing the recurrent neural networks, such as model compression, is crucial to ensure the broad deployment of recurrent neural networks and promote recurrent neural networks for implementing most resource-constrained scenarios. Among many techniques, tensor train (TT) decomposition is considered an up-and-coming technology. Although our previous efforts have achieved 1) expanding limits of many multiplications within eliminating all redundant computations; and 2) decomposing into multi-stage processing to reduce memory traffic, this work still faces some limitations. In particular, current TT decomposition on recurrent neural networks leads to a complex computation sensitive to the quality of training datasets. In this paper, we investigate a new method for TT decomposition on recurrent neural networks for constructing an efficient model within imbalance datasets to overcome this issue. Experimental results show that the proposed new training method can achieve significant improvements in accuracy, precision, recall, F1-score, False Negative Rate (FNR), and False Omission Rate (FOR). Mingshuo Liu, Kevin Han, Shiyi Luo, Mingze Pan, Mousam Hossain, Bo Yuan 0001, Ronald F. DeMara, Yu Bai 0004 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | ApGAN: Approximate GAN for Robust Low Energy Learning From Imprecise ComponentsabstractA Generative Adversarial Network (GAN) is an adversarial learning approach which empowers conventional deep learning methods by alleviating the demands of massive labeled datasets. However, GAN training can be computationally-intensive limiting its feasibility in resource-limited edge devices. In this paper, we propose an approximate GAN (ApGAN) for accelerating GANs from both algorithm and hardware implementation perspectives. First, inspired by the binary pattern feature extraction method along with binarized representation entropy, the existing Deep Convolutional GAN (DCGAN) algorithm is modified by binarizing the weights for a specific portion of layers within both the generator and discriminator models. Further reduction in storage and computation resources is achieved by leveraging a novel hardware-configurable in-memory addition scheme, which can operate in the accurate and approximate modes. Finally, a memristor-based processing-in-memory accelerator for ApGAN is developed. The performance of the ApGAN accelerator on different data-sets such as Fashion-MNIST, CIFAR-10, STL-10, and celeb-A is evaluated and compared with recent GAN accelerator designs. With almost the same Inception Score (IS) to the baseline GAN, the ApGAN accelerator can increase the energy-efficiency by ~28.6× achieving 35-fold speedup compared with a baseline GPU platform. Additionally, it shows 2.5× and 5.8× higher energy-efficiency and speedup over CMOS-ASIC accelerator subject to an 11 percent reduction in IS. Arman Roohi, Shadi Sheikhfaal, Shaahin Angizi, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Computers | 5 |
| 2019 | Workshop on Virtualized Active Learning in STEMabstractSummary form only given. Virtualized Active Learning (VAL) engages synchronous group-based problem solving within the last 30minutes of fully-live class offerings or the Face-to-Face component of mixed-mode delivery. As opposed to students solving problems together on-paper, which is not readily observable by the instructor nor scalable to larger class enrollments, VAL deploys laptops/tablets and Wi-Fi connectivity to create a virtualized active learning environment where students and instructors can interact. Ronald F. DeMara, Soheil Salehi |
FIE | 1 |
| 2019 | Virtualized Active Learning for Undergraduate Engineering Disciplines (VALUED): A Pilot in a Large Enrollment STEM ClassroomabstractThis student poster paper presents an innovative practice in order to increase the scalability and efficacy of student problem-based team learning in large enrollment engineering classrooms. We have devised a novel Virtualized Active Learning (VAL) approach to facilitate instructional delivery, assessment, and review of teams. VAL introduces a new pathway by utilizing open-source digital environments for effective and scalable team-based learning in classroom settings while empowering equitable participation from diverse learners. The proposed method provides a unique opportunity for learners to acquire knowledge and skills that are considered vital in STEM fields such as working in multidisciplinary teams proficiently and communicating with other team members in an effective manner. Meanwhile, they are engaged in finding an optimal solution to a design problem that requires certain specific constraints to be adequately met. The results of our pilot study indicate excellent potential for VAL in large enrollment STEM courses while facilitating the instructors to provide assistance and feedback to students in real-time. Soheil Salehi, Ronald F. DeMara |
FIE | 2 |
| 2019 | HSC-FPGAabstractThe HSC-FPGA offers an intriguing feasible architecture for the next generation of configurable fabrics, which allows embracing the advantages of both CMOS and beyond-CMOS technologies without requiring significant modification to the routing structure, programming paradigms, and synthesis tool-chain of the commercial FPGAs. In the HSC-FPGA, the intrinsic characteristics of magnetic random access memory (MRAM)-look-up table (LUT) circuits are used to implement sequential logic, while combinational logic circuits are implemented by static random access memory (SRAM)-LUTs. Fabric-level simulation results for the developed HSC-FPGA show that it can achieve at least 18%, 70%, and 15% reduction in terms of area, standby power, and read power consumption, respectively, for various ISCAS-89 and ITC-99 benchmark circuits compared to conventional SRAM-based FPGAs. The power consumption values can be further decreased by the power-gating allowed by the non-volatility feature of MRAM-LUTs. Moreover, the benefits of increased heterogeneity for reconfigurable computing is extended along realizing probabilistic computing paradigms within a fabric, which is enabled by probabilistic spin logic devices. The cooperating strengths of technology-heterogeneity and heterogeneity in computing paradigm in the proposed HSC-FPGA are leveraged to develop energy-efficient and reliability-aware training and evaluation circuits for deep belief networks with memristive crossbar arrays and p-bit based probabilistic neurons. Ramtin Zand, Ronald F. DeMara |
FPGA | 2 |
| 2019 | Design and Evaluation of DNU-Tolerant Registers for Resilient Architectural State StorageabstractIn this work, we aim to maintain the correct execution of instructions in the pipeline stages. To achieve that, the integrity for the data computed in registers during execution should be maintained via protecting the susceptible registers. Thus, we present a Double Node Upset Resilient Flip-Flop (DNUR-FF) circuit that can tolerate double errors while incurring low area and power overheads. We deploy the proposed soft-error resilient register at higher level to replace the most vulnerable registers in large-scale pipeline processors. The experimental results validate the robustness of our design by delivering superior fault coverage masking (100%) for both SEU and DNU errors. In addition, the proposed design utilizes partial spatial redundancy, and therefore, incurs reduced area overhead (31%) and realizes 58% of PDP improvement compared to Triple Module Redundancy (TMR) approach while delivering high-performance with low complexity and power consumption. Faris S. Alghareb, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | Clockless Spin-based Look-Up Tables with Wide Read MarginabstractIn this paper, we develop a 6-input fracturable non-volatile Clockless LUT (C-LUT) using spin Hall effect (SHE)-based Magnetic Tunnel Junctions (MTJs) and provide a detailed comparison between the SHE-MTJ-based C-LUT and Spin Transfer Torque (STT)-MTJ-based C-LUT. The proposed C-LUT offers an attractive alternative for implementing combinational logic as well as sequential logic versus previous spin-based LUT designs in the literature. Foremost, C-LUT eliminates the sense amplifier typically employed by using a differential polarity dual MTJ design, as opposed to a static reference resistance MTJ. This realizes a much wider read margin and the Monte Carlo simulation of the proposed fracturable C-LUT indicates no read and write errors in the presence of a variety of process variations scenarios involving MOS transistors as well as MTJs. Additionally, simulation results indicate that the proposed C-LUT reduces the standby power dissipation by 5.4-fold compared to the SRAM-based LUT. Furthermore, the proposed SHE-MTJ-based C-LUT reduces the area by 1.3-fold and 2-fold compared to the SRAM-based LUT and the STT-MTJ-based C-LUT, respectively. Soheil Salehi, Ramtin Zand, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | AQuRate: MRAM-based Stochastic Oscillator for Adaptive Quantization Rate Sampling of Sparse SignalsabstractRecently, the promising aspects of compressive sensing have inspired new circuit-level approaches for their efficient realization within the literature. However, most of these recent advances involving novel sampling techniques have been proposed without considering hardware and signal constraints. Additionally, traditional hardware designs for generating non-uniform sampling clock incur large area overhead and power dissipation. Herein, we propose a novel non-uniform clock generator called Adaptive Quantization Rate (AQR) generator using Magnetic Random Access Memory (MRAM)-based stochastic oscillator devices. Our proposed AQR generator provides ~25-fold reduction in area, on average, while offering ~6-fold reduced power dissipation, on average, compared to the state-of-the-art non-uniform clock generators. Soheil Salehi, Ramtin Zand, Alireza Zaeemzadeh, Nazanin Rahnavard, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 5 |
| 2019 | Self-Organized Sub-bank SHE-MRAM-based LLC: An energy-efficient and variation-immune read and write architecture
Soheil Salehi, Navid Khoshavi, Ramtin Zand, Ronald F. DeMara |
Integr. | 4 |
| 2019 | Composable Probabilistic Inference Networks Using MRAM-based Stochastic NeuronsabstractMagnetoresistive random access memory (MRAM) technologies with thermally unstable nanomagnets are leveraged to develop an intrinsic stochastic neuron as a building block for restricted Boltzmann machines (RBMs) to form deep belief networks (DBNs). The embedded MRAM-based neuron is modeled using precise physics equations. The simulation results exhibit the desired sigmoidal relation between the input voltages and probability of the output state. A probabilistic inference network simulator (PIN-Sim) is developed to realize a circuit-level model of an RBM utilizing resistive crossbar arrays along with differential amplifiers to implement the positive and negative weight values. The PIN-Sim is composed of five main blocks to train a DBN, evaluate its accuracy, and measure its power consumption. The MNIST dataset is leveraged to investigate the energy and accuracy tradeoffs of seven distinct network topologies in SPICE using the 14nm HP-FinFET technology library with the nominal voltage of 0.8V, in which an MRAM-based neuron is used as the activation function. The software and hardware level simulations indicate that a 784× 200× 10 topology can achieve less than 5% error rates with ∼400pJ energy consumption. The error rates can be reduced to 2.5% by using a 784× 500× 500× 500× 10 DBN at the cost of ∼10× higher energy consumption and significant area overhead. Finally, the effects of specific hardware-level parameters on power dissipation and accuracy tradeoffs are identified via the developed PIN-Sim framework. Ramtin Zand, Kerem Yunus Çamsari, Supriyo Datta, Ronald F. DeMara |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2019 | Guest Editorial: IEEE Transactions on Computers Special Section on Emerging Non-Volatile Memory Technologies: From Devices to Architectures and SystemsabstractThe papers in this special section focus on emerging non-volatile memory technologies (NVM). Emerging NVM technologies have attracted significant interest in recent years because of the fast-growing performance and capacity demands on memory and storage in the big data era. Well known examples include the 3D XPoint memory and various NVDIMM hybrid memory technologies. They have shown potential towards larger memory and storage capacities with nearly zero leakage power, while extending memory/ system architecture design approaches. The unique characteristics of NVM technologies not only introduce new opportunities, but simultaneously create challenges to the designs at multiple levels of abstraction in computer systems, including those of device management, CPU cache management, memory/storage architecture, and system design. Furthermore, emerging NVM technologies also drive the development of techniques which perform computing operations in memory, i.e., processing-in-memory (PIM), by taking advantage of crossbar-based accelerators using NVMs. Thus, for the emerging NVM technologies, there is an urgent need for technology innovation, modeling, analysis, design, and application, ranging from the device-level to the system-level. Yuan-Hao Chang 0001, Jingtong Hu, Mehdi Baradaran Tahoori, Ronald F. DeMara |
IEEE Trans. Computers | 4 |
| 2018 | Lockdown Computerized Testing Interwoven with Rapid Remediation: A Crossover Study within a Mechanical Engineering Core CourseabstractThis paper explores the realization of viable, scalable, automated, and authentic alternatives to paper-only-based testing within Engineering disciplines. Currently, manual delivery and grading of paper-based exams incurs vast logistical burdens that have low impact to learning achievement, especially as enrollments increase. Meanwhile, Engineering's design-oriented and problem-solving emphases pose substantial challenges to the digitized delivery of assessments, and thus they warrant a substantive evaluation of their validity. To address this research need, novel Computer-Based Assessment (CBA) infrastructures and delivery protocols were launched via an IRB-approved crossover study to investigate the impact of lockdown-proctored digitized quiz and exam delivery in terms of test score validity, and learning achievement within a large-size undergraduate Mechanical and Aerospace Engineering (MAE) course. Results indicate that well-formed CBAs can determine scores differing as little as 0.6% from Paper-Based Assessments (PBA). Student achievement was de-correlated by technical topic of the assessment delivery mode during crossover and results revealed that the CBA delivery and remediation cohort attained up to 16.9% higher learning outcomes during summative assessment. The encouraging results are discussed in detail along with lessons learned, and suggestions for transportability of CBA approaches to other Engineering courses and institutions. Ronald F. DeMara, Su Gao |
FIE | 2 |
| 2018 | Leveraging Spintronic Devices for Efficient Approximate Logic and Stochastic Neural NetworksabstractITRS has identified nano-magnet based spintronic devices as promising post-CMOS technologies for information processing and data storage due to their ultra-low switching energy, non-volatility, superior endurance, excellent retention time, high integration density and compatibility with CMOS technology. As for data storage, spintronic memory has been widely accepted as a universal high performance next-generation non-volatile memory candidate. As for information processing, spintronic computing remains complementary in its features to CMOS technology. In this paper, we present two innovative spintronic computing primitives, i.e. spintronic approximate logic and spintronic stochastic neural network, which both leverage the intrinsic spintronic device physics to achieve much more compact and efficient designs than CMOS counterparts. In spintronic approximate logic, we employ the intrinsic current-mode thresholding operation to implement an accuracy-configurable adder and further demonstrate its application in approximate DSP applications. In spintronic stochastic neural networks, we leverage the stochastic properties of domain wall devices and magnetic tunnel junction to implement a low-power and robust artificial neural network design. Shaahin Angizi, Zhezhi He, Yu Bai 0004, Jie Han 0001, Mingjie Lin, Ronald F. DeMara, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 6 |
| 2018 | Logic-Encrypted Synthesis for Energy-Harvesting-Powered Spintronic-Embedded Datapath DesignabstractThe objectives of advancing secure, intermittency-tolerant, and energy-aware logic datapaths are addressed herein by developing a spin-based design methodology and its corresponding synthesis steps. The approach selectively-inserts Non-Volatile (NV) Polymorphic Gates (PGs) to realize datapaths which are suitable for intrinsic operation in Energy-Harvesting-Powered (EHP) devices. Spin Hall Effect (SHE)-based Magnetic Tunnel (MTJs) are utilized to design NV-PGs, which are combined within a Flip-Flop (FF) circuit to develop a PG-FF realizing Boolean logic functions with inherent state-holding capability. The reconfigurability of PGs is leveraged for logic-encryption to enhance the security of the developed intermittency-resilient circuits, which are applied to ISCAS-89, MCNS, and ITC-99 benchmarks. The results obtained indicate that the PG-FF based design can achieve up to 7.1% and 13.6% improvements in terms of area and Power Delay Product (PDP), respectively, compared to NV-FF based methodologies that replace the CMOS-based FFs with NV-FFs. Further PDP improvements are achieved by using low-energy barrier SHE-MTJ devices within the PG-FF circuit. SHE-MTJs with 30kT energy exhibit 40.5% reduction in PDP at the cost of lower retention times in the range of minutes, which is still sufficient to achieve forward progress in EHP devices having more than hundreds of power-on and power-off cycles per minute. Arman Roohi, Ramtin Zand, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | Low-Energy Deep Belief Networks Using Intrinsic Sigmoidal Spintronic-based Probabilistic NeuronsabstractA low-energy hardware implementation of deep belief network (DBN) architecture is developed using near-zero energy barrier probabilistic spin logic devices (p-bits), which are modeled to realize an intrinsic sigmoidal activation function. A CMOS/spin based weighted array structure is designed to implement a restricted Boltzmann machine (RBM). Device-level simulations based on precise physics relations are used to validate the sigmoidal relation between the output probability of a p-bit and its input currents. Characteristics of the resistive networks and p-bits are modeled in SPICE to perform a circuit-level simulation investigating the performance, area, and power consumption tradeoffs of the weighted array. In the application-level simulation, a DBN is implemented in MATLAB for digit recognition using the extracted device and circuit behavioral models. The MNIST data set is used to assess the accuracy of the DBN using 5,000 training images for five distinct network topologies. The results indicate that a baseline error rate of 36.8% for a 784x10 DBN trained by 100 samples can be reduced to only 3.7% using a 784x800x800x10 DBN trained by 5,000 input samples. Finally, Power dissipation and accuracy tradeoffs for probabilistic computing mechanisms using resistive devices are identified. Ramtin Zand, Kerem Yunus Çamsari, Steven D. Pyle, Ibrahim Ahmed 0002, Chris H. Kim, Ronald F. DeMara |
ACM Great Lakes Symposium on VLSI | 6 |
| 2018 | BGIM: Bit-Grained Instant-on Memory Cell for Sleep Power Critical Mobile ApplicationsabstractThis paper devises a novel energy-aware Non-Volatile Static Random Access Memory (NV-SRAM) framework for sleep power critical mobile applications. The beyond-Complementary Metal Oxide Semiconductor (CMOS) hardware architecture has been designed to minimize the overall static and leakage energy consumption while providing fast back-up and restore operations. Differential Spin-Hall Effect Magnetic Random Access Memory devices are utilized to realize the proposed framework called Bit-Grained Instant-on Memory Cell (BGIM). Our results indicate that the proposed BGIM consumes 121.51fJ on average for each back-up operation and 1.56fJ on average for each restore operation. Furthermore, the proposed BGIM can perform rapid back-up operations in 1ns and fast restore operations in 13.2ps. Moreover, the proposed BGIM cell incurs 0.4μm^2 area overhead compared to the traditional 6T SRAM cell, however it eliminates the need for data transmission and a separate non-volatile memory macro. Soheil Salehi, Ronald F. DeMara |
ICCD | 2 |
| 2018 | Clockless Spintronic Logic: A Robust and Ultra-Low Power Computing ParadigmabstractAsynchronous logic offers the advantages of no clock tree, robust circuit operation, avoidance of worst-case timing margins, and a reduced emission spectrum. Thus, computational paradigms are sought to attain advantages of clockless logic by leveraging the complementary characteristics of emerging devices and CMOS transistors within novel circuit designs. This paper introduces Spin Torque Enabled NULL Convention Logic (STENCL), which exploits the physical characteristics of non-volatile Domain-Wall (DW) and memristive devices to realize the Quasi-Delay-Insensitive (QDI) NULL Convention Logic (NCL) asynchronous design methodology. First, a formal algorithm is developed to transform NCL-based threshold m-of-n gate realizations to STENCL, in order to generate the corresponding input memristance and NULL module memristance required for nominal currents achieving DW device biasing. Second, hysteresis and set/reset conditions are realized by determining the corresponding current fluctuations required to move the DW within each threshold logic gate to realize all 27 foundational NCL gate structures, which are then simulated to assess energy and delay metrics. Third, a case study of a four-stage pipelined 32-bit IEEE single-precision floating point co-processor implemented as a dual-rail STENCL architecture is compared to a conventional CMOS-based NCL design implemented by an IBM SOI1250 45nm CMOS process. Fourth, a sensitivity analysis is performed to assess the impact of write accuracy and drift on memristor and DW device operation. Results indicate that STENCL-based designs achieve between 2-fold to 20-fold reduction in energy consumption with up to 8-fold reduction in area, over an equivalent CMOS-based NCL design for 32-bit full adders. Comparisons for various four-stage pipelined 32-bit IEEE single-precision floating-point co-processors and ISCAS benchmarks further substantiate those benefits for operation within acceptable tolerances at identical process technology nodes. Yu Bai 0004, Ronald F. DeMara, Jia Di, Mingjie Lin |
IEEE Trans. Computers | 2 |
| 2018 | NV-Clustering: Normally-Off Computing Using Non-Volatile DatapathsabstractWith technology downscaling, static power dissipation presents a crucial challenge to multicore, many-core, and System-on-Chip (SoC) architectures due to the increased role of leakage currents in overall energy consumption and the need to support power-gating schemes. Herein, a non-Volatile (NV) flip-flop design approach, referred to as NV Clustering, is developed to realize middleware-transparent intermittent computing. First, a Logic-Embedded Flip-Flop (LE-FF) is developed to realize rudimentary Boolean logic functions along with an inherent state-holding capability within a compact footprint. Second, the NV-Clustering synthesis procedure and corresponding tool module are utilized to instantiate the LE-FF library cells within conventional Register Transfer Language (RTL) specifications. This selectively clusters together logic and NV state-holding functionality, based on energy and area minimization criteria. NV-Clustering is applied to a wide range of benchmarks including ISCAS-89, MCNS, and ITC-99 computational circuits using a LE-FF based on the Spin Hall Effect (SHE)-assisted Spin Transfer Torque (STT) Magnetic Tunnel Junction (MTJ). Simulation results validate functionality and power dissipation, area, and delay benefits. For instance, results for ISCAS-89 benchmarks indicate 15 percent area reduction on average, up to 22 percent reduction in energy consumption, and up to 14 percent reduction in delay as compared to alternative NV-FF based designs, as evaluated via SPICE simulation at the 45-nm technology node. Arman Roohi, Ronald F. DeMara |
IEEE Trans. Computers | 2 |
| 2018 | Survivability Modeling and Resource Planning for Self-Repairing Reconfigurable Device FabricsabstractA resilient system design problem is formulated as the quantification of uncommitted reconfigurable resources required for a system of components to survive its lifetime within mission availability specifications. We show that this survivability metric can be calculated according to the residual functionality obtained from pools of dynamically configurable elements constituting the amorphous resource pool (ARP). The ARP is depleted based on the failure rate to replenish the functionality lost in a reconfigurable fabric due to the occurrence of permanent faults during the mission lifetime. While genetic algorithms are selected for the reparation method, any probabilistic or deterministic active repair strategy is covered without loss of generality. Parameters of this model are correlated with reliability specifications of Xilinx Virtex-4 field programmable gate array devices, which are then utilized for MCNC benchmark circuits along with a realistic space mission. Calculation of the spare fabric resources which must be budgeted for a mission, maximum mission lifetime, and repair policy parameters are realized using the proposed probabilistic survivability model for soft computing-based repair strategies. Rashad S. Oreifej, Rawad N. Al-Haddad, Ramtin Zand, Rizwan A. Ashraf, Ronald F. DeMara |
IEEE Trans. Cybern. | 5 |
| 2018 | Designing and Evaluating Redundancy-Based Soft-Error Masking on a Continuum of Energy versus RobustnessabstractNear-threshold computing is an effective strategy to reduce the power dissipation of deeply-scaled CMOS logic circuits. However, near-threshold strategies exacerbate the impact of delay variations on device performance and increase the susceptibility to soft errors due to narrow voltage margins. The objective of this work is to develop and assess design approaches that leverage tradeoffs between performance and the resilience of fault masking coverage for various soft-error mitigation techniques. The primary insight from this work is identification of redundancy-based hardening techniques that can deliver increased benefits in terms of the fault coverage energy ratio (FCER) for the leveraged tradeoffs within iso-energy constraints at near-threshold voltage (NTV). Simulation results demonstrate that temporal redundancy approaches offer favorable tradeoffs in terms of FCER. They exhibit reduced impact on performance variations and achieve extensive soft fault masking, therefore improving the system robustness within acceptable delay constraints. Meanwhile, it is shown that a hybrid redundancy approach can be used to protect a low-power system to maintain throughput while tolerating soft errors. We demonstrate how the FCER metric can be used as an optimization parameter to guide circuit synthesis to meet performance and robustness goals. Finally, the impact of design diversity on spatial and hybrid redundancy at NTV is assessed in terms of FCER and delay variation to form overall recommendations regarding soft-error mitigation at NTV. Faris S. Alghareb, Rizwan A. Ashraf, Ronald F. DeMara |
IEEE Trans. Sustain. Comput. | 3 |
| 2017 | A Spin-Orbit Torque based Cellular Neural Network (CNN) ArchitectureabstractIn this paper, we propose a differential Spin Hall Effect(SHE) assisted domain wall synapse, which can generate either positive or negative synaptic weighting values without the significant cost of multiple power supply voltages, supply rails, or computationally-intensive digital hardware. The architecture of the proposed synapse utilizes reading currents flowing through two oppositely-oriented devices as weighted by device conductance. The conductance is used to encode synaptic weight and programmed by domain wall position through writing current. The ability to set the current as positively or negatively weighted results in highly-configurable functionality within a compact synapse design. The synapses are used with a soft-limiting nonlinear neuron to employ the relationship between positions and input current magnitude. We show through micro-magnetic simulation how the non-volatile physical characteristic of the domain wall calibrated synapse is used to implement a numerical integration function to realize a Cellular Neural Network(CNN). The performance of the proposed CNN design for isolated letter denoising at 0ns to 4ns demonstrates noise filtering functionality with total energy consumption during sensing of 24fJ. This compares favorably to existing spin CNN cell designs to provide a promising design approach for intrinsic neural computation. Yu Bai 0004, Xiaobo Sharon Hu, Ronald F. DeMara, Mingjie Lin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Process variation immune and energy aware sense amplifiers for resistive non-volatile memoriesabstractSpin-Transfer Torque Magnetic Random Access Memory (STT-MRAM) has been explored as a post-CMOS technology for embedded and data storage applications seeking non-volatility, near-zero standby energy, and high density. Towards attaining these objectives for practical implementations, various techniques to mitigate the specific reliability challenges associated with STT-MRAM elements are surveyed, classified, and assessed herein. Some solutions to the reliability issues identified are addressed to realize reliable STT-MRAM designs. In an attempt to further improve the process variation immunity of the Sense Amplifiers (SAs), two new SAs are introduced: Energy Aware Sense Amplifier (EASA) and Variation Immune Sense Amplifier (VISA). Results have shown that EASA and VISA achieve superior performance in most cases compared to two of the most common SAs, namely PCSA and SPCSA respectively, while reducing Bit Error Rate (BER) and increasing reliability. Soheil Salehi, Ronald F. DeMara |
ISCAS | 2 |
| 2017 | Contemporary CMOS aging mitigation techniques: Survey, taxonomy, and methods
Navid Khoshavi, Rizwan A. Ashraf, Ronald F. DeMara, Saman Kiamehr, Fabian Oboril, Mehdi Baradaran Tahoori |
Integr. | 3 |
| 2017 | Survey of STT-MRAM Cell Design Strategies: Taxonomy and Sense Amplifier Tradeoffs for ResiliencyabstractSpin-Transfer Torque Random Access Memory (STT-MRAM) has been explored as a post-CMOS technology for embedded and data storage applications seeking non-volatility, near-zero standby energy, and high density. Towards attaining these objectives for practical implementations, various techniques to mitigate the specific reliability challenges associated with STT-MRAM elements are surveyed, classified, and assessed in this article. Cost and suitability metrics assessed include the area of nanomagmetic and CMOS components per bit, access time and complexity, sense margin, and energy or power consumption costs versus resiliency benefits. Solutions to the reliability issues identified are addressed within a taxonomy created to categorize the current and future approaches to reliable STT-MRAM designs. A variety of destructive and non-destructive sensing schemes are assessed for process variation tolerance, read disturbance reduction, sense margin, and write polarization asymmetry compensation. The highest resiliency strategies deliver a sensing margin above 300mV while incurring low power and energy consumption on the order of picojoules and microwatts, respectively, and attaining read sense latency of a few nanoseconds down to hundreds of picoseconds for non-destructive and destructive sensing schemes, respectively. Soheil Salehi, Deliang Fan, Ronald F. DeMara |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2017 | Energy-Aware Adaptive Restore Schemes for MLC STT-RAM CacheabstractFor the sake of higher cell density while achieving near-zero standby power, recent research progress in Magnetic Tunneling Junction (MTJ) devices has leveraged Multi-Level Cell (MLC) configurations of Spin-Transfer Torque Random Access Memory (STT-RAM). However, in orderto mitigate the write disturbance in an MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. Furthermore, as the result of MTJ feature size scaling, the soft bit can be expected to become disturbed by the read sensing current, thus requiring an immediate restore operation to ensure the data reliability. In this paper, we design and analyze a novel Adaptive Restore Scheme for Write Disturbance (ARS-WD) and Read Disturbance (ARS-RD), respectively. ARS-WD alleviates restoration overhead by intentionally overwriting soft bit lines which are less likely to be read. ARS-RD, on the other hand, aggregates the potential writes and restore the soft bit line at the time of its eviction from higher level cache. Both of these two schemes are based on a lightweight forecasting approach for the future read behavior of the cache block. Our experimental results show substantial reduction in soft bit line restore operations, delivering 17.9 percent decrease in overall energy consumption and 9.4 percent increase in IPC, while incurring negligible capacity overhead. Moreover, ARS promotes advantages of MLC to provide a preferable L2 design alternative in terms of energy, area and latency product compared to SLC STT-RAM alternatives. Xunchao Chen, Navid Khoshavi, Ronald F. DeMara, Jun Wang 0001, Dan Huang 0001, Wujie Wen, Yiran Chen 0001 |
IEEE Trans. Computers | 3 |
| 2017 | Guest Editorial: IEEE Transactions on Computers and IEEE Transactions on Emerging Topics in Computing Joint Special Section on Innovation in Reconfigurable Computing Fabrics from Devices to ArchitecturesabstractThe papers in this special section focuses on the advancement of the associated performance and reliability objectives via technology and functional heterogeneity, as well as advanced resilience properties in which reconfigurable fabrics are embarking. Emerging device characteristics such as non-volatility and new static versus dynamic energy consumption profiles, as well as novel interconnect mechanisms, and storage blocks within reconfigurable fabrics, in turn innovate architectural advances that enable new applications. Ronald F. DeMara, Marco Platzner, Marco Ottavi |
IEEE Trans. Computers | 1 |
| 2017 | Voltage-Based Concatenatable Full Adder Using Spin Hall Effect SwitchingabstractMagnetic tunnel junction (MTJ)-based devices have been studied extensively as a promising candidate to implement hybrid energy-efficient computing circuits due to their nonvolatility, high integration density, and CMOS compatibility. In this paper, MTJs are leveraged to develop a novel full adder (FA) based on 3- and 5-input majority gates. Spin Hall effect (SHE) is utilized for changing the MTJ states resulting in low-energy switching behavior. SHE-MTJ devices are modeled in Verilog-A using precise physical equations. SPICE circuit simulator is used to validate the functionality of 1-bit SHE-based FA. The simulation results show 76% and 32% improvement over previous voltage-mode MTJ-based FA in terms of energy consumption and device count, respectively. The concatanatability of our proposed 1-bit SHE-FA is investigated through developing a 4-bit SHE-FA. Finally, delay and power consumption of an n-bit SHE-based adder has been formulated to provide a basis for developing an energy efficient SHE-based n-bit arithmetic logic unit. Arman Roohi, Ramtin Zand, Deliang Fan, Ronald F. DeMara |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Energy-Efficient and Process-Variation-Resilient Write Circuit Schemes for Spin Hall Effect MRAM DeviceabstractIn this paper, various energy-efficient write schemes are proposed for switching operation of spin hall effect (SHE)-based magnetic tunnel junctions (MTJs). A transmission gate (TG)-based write scheme is proposed, which provides a symmetric and energy-efficient switching behavior. We have modeled an SHE-MTJ using precise physics equations, and then leveraged the model in SPICE circuit simulator to verify the functionality of our designs. Simulation results show the TG-based write scheme advantages in terms of device count and switching energy. In particular, it can operate at 12% higher clock frequency while realizing at least 13% reduction in energy consumption compared to the most energy-efficient write circuits. We have analyzed the performance of the implemented write circuits in presence of process variation (PV) in the transistors' threshold voltage and SHE-MTJ dimensions. Results show that the proposed TG-based design is the second most PV-resilient write circuit scheme for SHE-MTJs among the implemented designs. Finally, we have proposed the 1TG-1T-1R SHE-based magnetic random access memory (MRAM) bit cell based on the TG-based write circuit. Comparisons with several of the most energy-efficient and variation-resilient SHE-MRAM cells indicate that 1TG-1T-1R delivers reduced energy consumption with 43.9% and 10.7% energy-delay product improvement, while incurring low area overhead. Ramtin Zand, Arman Roohi, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cacheabstractSpin-Transfer Torque Random Access Memory (STT-RAM) has been identified as an advantageous candidate for on-chip memory technology due to its high density and ultra low leakage power. Recent research progress in Magnetic Tunneling Junction (MTJ) devices has developed Multi-Level Cell (MLC) STT-RAM to further enhance cell density. To avoid the write disturbance in MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. However, frequent restores are not only unnecessary, but also introduce a significant energy consumption overhead. In this paper, we propose an Adaptive Overwrite Scheme (AOS) which alleviates restoration overhead by intentionally overwriting selected soft bits based on RRD (Read Reuse Distance). Our experimental results show 54.6% reduction in soft bit restoration, delivering 10.8% decrease in overall energy consumption. Moreover, AOS promotes MLC to be a preferable L2 design alternative in terms of energy, area and latency product. Xunchao Chen, Navid Khoshavi, Jian Zhou 0004, Dan Huang 0001, Ronald F. DeMara, Jun Wang 0001, Wujie Wen, Yiran Chen 0001 |
DAC | 5 |
| 2016 | Fast Online Diagnosis and Recovery of Reconfigurable Logic Fabrics Using Design DisjunctionabstractDesign disjunction is developed to offer a broad coverage, high resolution, and low overhead approach to online diagnosis and recovery of reconfigurable fabrics. Design disjunction leverages the condensed diagnosability of$T$logic resources to achieve self-recovery using partial reconfiguration in O(log$T$) steps. Reconfiguration is guided by the constructive property of$f$-disjunctness which forms O(log$T$) resource groups at design-time. Resolution of$f$simultaneous resource faults is shown to be guaranteed when the resource groups are mutually$f$-disjunct. This extends run-time fault resilience to a large resource space with certainty for up to$f$faults using a decision-free resolution process that also provides a high likelihood of identifying the fault’s location to a fine granularity. Finally, design disjunction is parameterized to accommodate the low coverage issue of functional testing for which inarticulate tests can otherwise impair fault isolation. Experimental results for MCNC and ISCAS benchmarks on a Xilinx 7-series field programmable gate array (FPGA) demonstrate$f$-diagnosability at the individual slice level with a minimum average isolation accuracy of$96.4$percent ($94.4$percent) for$f=1$($f=2$). Results have also demonstrated millisecond order recovery with a minimum increase of$83.6$percent in fault coverage compared to$N$-modular redundancy (NMR) schemes. Recovery is achieved while incurring an average critical path delay impact of only$1.49$percent and energy cost roughly comparable to conventional two-MR approaches. Ahmad Alzahrani 0001, Ronald F. DeMara |
IEEE Trans. Computers | 2 |
| 2016 | Loss-Aware Switch Design and Non-Blocking Detection Algorithm for Intra-Chip Scale Photonic Interconnection NetworksabstractAs the number of on-chip processor cores increases, power-efficient solutions are sought for data communication between cores. TheHelix-hnon-blocking photonic switch is developed to improve physical-layer and network performance parameters for a wide range of silicon nano-photonic multicore interconnection topologies. Traffic benchmarks and practical case studies using a cycle-accurate simulation environment indicate significantly reduced insertion loss providing improved bandwidth density and scalability to manycore plurality. Improvements in system performance parameters are quantified for network bandwidth, transmission efficiency, and latency in popular photonic internconnection topologies, in comparison to previous switch designs. For instance, utilizing the Helix-h switch in a mesh topology, the bandwidth is increased by 112 percent compared to the previously highest performing switch design. Execution time and energy efficiency are improved by up to 92 and 99 percent, respectively, for representative multicore applications. Finally, the technique is generalized to a novel graph-theoretic method for articulating blocking conditions in photonic switches. Hesam Shabani, Arman Roohi, Akram Reza, Midia Reshadi, Nader Bagherzadeh, Ronald F. DeMara |
IEEE Trans. Computers | 6 |
| 2015 | Reactive rejuvenation of CMOS logic paths using self-activating voltage domainsabstractAlthough the trend of technology scaling is sought to realize higher performance computer systems, it also results in Integrated Circuits (ICs) suffering from increasing Process, Voltage, and Temperature (PVT) variations and adverse aging effects. In most cases, these reliability threats manifest themselves as timing errors on critical speed-paths of the circuit, if a large design guardband is not reserved. In this work, we propose the Reactive Rejuvenation (RR) architectural approach consisting of detection and recovery phases to mitigate circuit from BTI-induced aging. The BTI impact on the critical and near critical paths performance is continuously examined through a lightweight logic circuit which asserts an error signal in the case of any timing violation in those paths. By utilizing timing violation occurrence in the system, the timing-sensitive portion of the circuit is recovered from BTI through switching computations to redundant aging-critical voltage domain. The proposed technique achieves aging mitigation and reduced energy consumption as compared to a baseline circuit. Thus, significant voltage guardbands to meet the desired timing specification are avoided. Rizwan A. Ashraf, Ahmad Alzahrani 0001, Navid Khoshavi, Ramtin Zand, Soheil Salehi, Arman Roohi, Mingjie Lin, Ronald F. DeMara |
ISCAS | 8 |
| 2015 | Understanding the propagation of transient errors in HPC applicationsabstractResiliency of exascale systems has quickly become an important concern for the scientific community. Despite its importance, still much remains to be determined regarding how faults disseminate or at what rate do they impact HPC applications. The understanding of where and how fast faults propagate could lead to more efficient implementation of application-driven error detection and recovery. Rizwan A. Ashraf, Roberto Gioiosa, Gokcen Kestor, Ronald F. DeMara, Chen-Yong Cher, Pradip Bose |
SC | 4 |
| 2014 | Energy-efficient multiplier-less discrete convolver through probabilistic domain transformationabstractEnergy efficiency and algorithmic robustness typically are conflicting circuit characteristics, yet with CMOS technology scaling towards 10-nm feature size, both become critical design metrics simultaneously for modern logic circuits. This paper propose a novel computing scheme hinged on probabilistic domain transformation aiming for both low power operation and fault resilience. In such a computing paradigm, algorithm inputs are first encoded through probabilistic means, which translates the input values into a number of random samples. Subsequently, light-weight operations, such as sim- ple additions will be performed onto these random samples in order to generate new random variables. Finally, the resulting random samples will be decoded probabilistically to give the final results. Mohammed Alawad, Yu Bai 0004, Ronald F. DeMara, Mingjie Lin |
FPGA | 3 |
| 2014 | Non-adaptive sparse recovery and fault evasion using disjunct design configurations (abstract only)abstractA run-time fault diagnosis and evasion scheme for reconfigurable devices is developed based on an explicit Non-adaptive Group Testing (NGT). NGT involves grouping disjunct subsets of reconfigurable resources into test pools, or samples. Each test pool realizes a Diagnostic Configuration (DC) performing functional testing during diagnosis procedure. The collective test outcomes after testing each diagnostic pool can be efficiently decoded to identify up to d defective logic resources. An algorithm for constructing NGT sampling procedure and resource placement during design time with optimal minimal number of test groups is derived through the well-known in statistical literature d-disjunctness property. The combinatorial properties of resultant DCs also guarantee that any possible set of defective resources less than or equal to d are not utilized by at least one DC, allowing a low-overhead fault resolution. It also provides the ability to assess the resources state of failure. The proposed testing scheme thus avoids time-intensive run-time diagnosis imposed by previously proposed adaptive group testing for reconfigurable hardware without compromising diagnostic coverage. In addition, proposed NGT scheme can be combined with other fault tolerance approaches to ameliorate their fault recovery strategies. Experimental results for a set of MCNC benchmarks using Xilinx ISE Design Suite on a Virtex-5 FPGA have demonstrated d-diagnosability at slice level with average accuracy of 99.15% and 97.76% for d=1 and d=2, respectively. Ahmad Alzahrani 0001, Ronald F. DeMara |
FPGA | 2 |
| 2013 | Scalable FPGA Refurbishment Using Netlist-Driven Evolutionary AlgorithmsabstractIn this work, Field-Programmable Gate Array (FPGA) reconfigurability is exploited to realize autonomous fault recovery in mission-critical applications at runtime. The proposed Netlist-Driven Evolutionary Refurbishment technique utilizes design-time information from the circuit netlist to constrain the search space of the algorithm by up to 98.1 percent in terms of the chromosome length representing reconfigurable logic elements. This facilitates refurbishment of relatively large-sized FPGA circuits as compared to previous works. Hence, the scalability issue associated with Evolvable Hardware-Based refurbishment is addressed and improved. Experiments are conducted with multiple circuits from the MCNC benchmark suite to validate the approach and assess its benefits and limitations. Successful refurbishment of the apex4 circuit having a total of 1,252 LUTs with 10 percent spares is achieved in as few as 633 generations on average when subjected to simulated randomly injected single stuck-at faults. Moreover, the use of design-time information about the circuit undergoing refurbishment is validated as means to increase the tractability of dynamic evolvable hardware techniques. Rizwan A. Ashraf, Ronald F. DeMara |
IEEE Trans. Computers | 2 |
| 2013 | Fault Demotion Using Reconfigurable Slack (FaDReS)abstractWe propose an active dynamic redundancy-based fault-handling approach exploiting the partial dynamic reconfiguration capability of static random-access memory-based field-programmable gate arrays. Fault detection is accomplished in a uniplex hardware arrangement while an autonomous fault isolation scheme is employed, which neither requires test vectors nor suspends the computational throughput. The deterministic flow of the fault-handling scheme achieves an improved recovery in a bounded number of reconfigurations. This approach extends existing signal processing properties to accommodate fault handling, and is validated by implementing an H.263 video encoder discrete cosine transform (DCT) block. The peak signal-to-noise ratio measure of the video sequences indicates fault tolerance in the DCT block with only limited quality degradation, during the isolation and recovery phases spanning a few frames. Naveed Imran, Jooheung Lee, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Design of Asynchronous Circuits for High Soft Error Tolerance in Deep Submicrometer CMOS CircuitsabstractAs the devices are scaling down, the combinational logic will become susceptible to soft errors. The conventional soft error tolerant methods for soft errors on combinational logic do not provide enough high soft error tolerant capability with reasonably small performance penalty. This paper investigates the feasibility of designing quasi-delay insensitive (QDI) asynchronous circuits for high soft error tolerance. We analyze the behavior of null convention logic (NCL) circuits in the presence of particle strikes, and propose an asynchronous pipeline for soft-error correction and a novel technique to improve the robustness of threshold gates, which are basic components in NCL, against particle strikes by using Schmitt trigger circuit and resizing the feedback transistor. Experimental results show that the proposed threshold gates do not generate soft errors under the strike of a particle within a certain energy range if a proper transistor size is applied. The penalties, such as delay and power consumption, are also presented. Weidong Kuang, Peiyi Zhao, Jiann-Shiun Yuan, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | Towards a Method For Evaluating Naturalness in Conversational Dialog SystemsabstractThe evaluation of conversational dialog systems has remained a controversial topic, as it is challenging to quantitatively assess how well a conversation agent performs, or how much better one is compared to another. Furthermore, one of the hurdles which remains elusive in this quandary is the definition of naturalness, as demonstrated by how well a dialog system can maintain a natural conversation flow devoid of perceived awkwardness. As a step towards defining the dimensions of effectiveness and naturalness in a dialog system, this paper identifies existing evaluation practices which are then expanded to develop a more suitable assessment vehicle. This method is then applied to the LifeLike virtual avatar project. Victor Chou Hung, Miguel Elvir, Avelino J. Gonzalez, Ronald F. DeMara |
SMC | 4 |
| 2009 | Scalable FPGA-based architecture for DCT computation using dynamic partial reconfigurationabstractIn this article, we propose field programmable gate array-based scalable architecture for discrete cosine transform (DCT) computation using dynamic partial reconfiguration. Our architecture can achieve quality scalability using dynamic partial reconfiguration. This is important for some critical applications that need continuous hardware servicing. Our scalable architecture has three features. First, the architecture can perform DCT computations for eight different zones, that is, from 1 × 1 DCT to 8× 8 DCT. Second, the architecture can change the configuration of processing elements to trade off the precisions of DCT coefficients with computational complexity. Third, unused PEs for DCT can be used for motion estimation computations. Using dynamic partial reconfiguration with 2.3MB bitstreams, 80 distinct hardware architectures can be implemented. We show the experimental results and comparisons between different configurations using both partial reconfiguration and nonpartial reconfiguration process. The detailed trade-offs among visual quality, power consumption, processing clock cycles, and reconfiguration overhead are analyzed in the article. Matthew Parris, Jooheung Lee, Ronald F. DeMara |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2008 | A Multilayer Framework Supporting Autonomous Run-Time Partial ReconfigurationabstractA multilayer run-time reconfiguration architecture (MRRA) is developed for autonomous run-time partial reconfiguration of field-programmable gate-array (FPGA) devices. MRRA operations are partitioned into logic, translation, and reconfiguration layers along with a standardized set of application programming interfaces (APIs). At each level, resource details are encapsulated and managed for efficiency and portability during operation. In particular, FPGA configurations can be manipulated at runtime using on-chip resources. A corresponding logic control flow is developed for a prototype MRRA system on a Xilinx Virtex II Pro platform. The Virtex II Pro on-chip PowerPC core and block RAM are employed to manage control operations while multiple physical interfaces establish and supplement autonomous reconfiguration capabilities. Evaluations of these prototypes on a number of benchmark and hashing algorithm case studies indicate the enhanced resource utilization and run time performance of the developed approaches. Heng Tan, Ronald F. DeMara |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Layered Approach to Instrinsic Evolvable Hardware Using Direct Bistream Manipulation of VIRTEX II Pro DevicesabstractAn integrated platform for fast genetic operators is presented to support intrinsic evolution on Xilinx Virtex II Pro Field Programmable Gate Arrays (FPGAs). Dynamic bitstream compilation is achieved by directly manipulating the bitstream using a layered design. Experimental results on a case study have shown that a full design as well as a full repair is achievable using this platform with an average time of 0.4 microseconds to perform the genetic mutation, 0.7 microseconds to perform the genetic crossover, and 5.6 milliseconds for one input pattern intrinsic evaluation. This represents a performance advantage of three orders of magnitude over JBITS and more than seven orders of magnitude over the Xilinx design tool driven flow for realizing intrinsic genetic operators on a Virtex II Pro device. Rashad S. Oreifej, Rawad N. Al-Haddad, Heng Tan, Ronald F. DeMara |
FPL | 4 |
| 2007 | Pipelining of Fuzzy ARTMAP without matchtracking: Correctness, performance bound, and Beowulf evaluation
José Castro, Jimmy Secretan, Michael Georgiopoulos, Ronald F. DeMara, Georgios C. Anagnostopoulos, Avelino J. Gonzalez |
Neural Networks | 4 |
| 2007 | Tiered Algorithm for Distributed Process Quiescence and Termination DetectionabstractThe Tiered Algorithm is presented for time-efficient and message-efficient detection of process termination. It employs a global invariant of equality between process production and consumption at each level of process nesting to detect termination, regardless of execution interleaving order and network transit time. Correctness is validated for arbitrary process launching hierarchies, including launch-in-transit hazards, where processes are created dynamically based on runtime conditions for remote execution. The performance of the Tiered Algorithm is compared to three existing schemes with comparable capabilities, namely, the Chandrasekaran and Venkatesan (CV), Lai, Tseng, and Dong (LTD), and Credit termination detection algorithms. For synchronization of X tasks terminating in E epochs of idle processing, the tiered algorithm is shown to incur O(E) message count complexity and O(T lg T) message bit complexity while incurring detection latency corresponding to only integer addition and comparison. The synchronization performance in terms of message overhead, detection operations, and storage requirements are evaluated and compared across numerous task creation and termination hierarchies. Ronald F. DeMara, Yili Tseng, Abdel Ejnioui |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2006 | CONFIDANT: Collaborative Object Notification Framework for Insider Defense using Autonomous Network Transactions
Adam J. Rocke, Ronald F. DeMara |
Auton. Agents Multi Agent Syst. | 2 |
| 2006 | Improving power-awareness of pipelined array multipliers using two-dimensional pipeline gating and its application on FIR design
Jia Di, Jiann-Shiun Yuan, Ronald F. DeMara |
Integr. | 3 |
| 2006 | Learning tactical human behavior through observation of human performanceabstractIt is widely accepted that the difficulty and expense involved in acquiring the knowledge behind tactical behaviors has been one limiting factor in the development of simulated agents representing adversaries and teammates in military and game simulations. Several researchers have addressed this problem with varying degrees of success. The problem mostly lies in the fact that tactical knowledge is difficult to elicit and represent through interactive sessions between the model developer and the subject matter expert. This paper describes a novel approach that employs genetic programming in conjunction with context-based reasoning to evolve tactical agents based upon automatic observation of a human performing a mission on a simulator. In this paper, we describe the process used to carry out the learning. A prototype was built to demonstrate feasibility and it is described herein. The prototype was rigorously and extensively tested. The evolved agents exhibited good fidelity to the observed human performance, as well as the capacity to generalize from it. H. K. G. Fernlund, Avelino J. Gonzalez, Michael Georgiopoulos, Ronald F. DeMara |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2005 | Data-partitioning using the Hilbert space filling curves: Effect on the speed of convergence of Fuzzy ARTMAP for large database problems
José Castro, Michael Georgiopoulos, Ronald F. DeMara, Avelino J. Gonzalez |
Neural Networks | 3 |
| 2004 | A data partitioning approach to speed up the fuzzy ARTMAP algorithm using the Hilbert space-filling curveabstractOne of the properties of FAM, which is a mixed blessing, is its capacity to produce new neurons (templates) on demand to represent classification categories. This property allows FAM to automatically adapt to the database without having to arbitrarily specify network structure, but it also has the undesirable side effect that on large databases it can produce a large network size that can dramatically slow down the algorithms' training time. To address this problem, we propose the use of the Hilbert space-filling curve. Our results indicate that the Hilbert space-filling curve can reduce the training time of FAM by partitioning the learning set without a significant effect on the classification performance or network size. Given that there is full data partitioning with the HSFC, we implement and test a parallel implementation on a Beowulf cluster of workstations that further speeds up the training and classification time on large databases. José Castro, Michael Georgiopoulos, Ronald F. DeMara |
IJCNN | 3 |
| 2004 | Mitigation of network tampering using dynamic dispatch of mobile agents
Ronald F. DeMara, Adam J. Rocke |
Comput. Secur. | 1 |
| 2004 | Optimization of NULL convention self-timed circuits
Scott C. Smith, Ronald F. DeMara, Jiann-Shiun Yuan, Dennis Ferguson, D. Lamb |
Integr. | 2 |
| 2003 | CRCD in machine learning at the University of Central Florida preliminary experiences
Michael Georgiopoulos, José Castro, Annie S. Wu, Ronald F. DeMara, Erol Gelenbe, Avelino J. Gonzalez, Marcella K. Kysilka, Mansooreh Mollaghasemi |
ITiCSE | 4 |
| 2003 | Distributed-sum termination detection supporting multithreaded execution
Yili Tseng, Ronald F. DeMara, P. J. Wilder |
Parallel Comput. | 2 |
| 2002 | Communication Pattern Based Methodology for Performance Analysis of Termination Detection SchemesabstractEfficient determination of processing termination at barrier synchronization points can occupy an important role in the overall throughput of parallel and distributed computing systems. Even though relatively efficient termination detection techniques have been proposed for certain environments, no effective performance analysis methodology has been introduced to determine application attributes that favor the use of a particular termination detection technique. This fact has hindered the adoption and development of termination detection schemes. This paper addresses this problem by developing a communication pattern based methodology to improve the precision of the theoretical performance of termination detection techniques in lieu of laborious experiments or potentially subjective benchmarking studies. By measuring message complexity from the idle period respect, it provides a simple and effective way to evaluate existing termination detection techniques or design new termination detection algorithms. Yili Tseng, Ronald F. DeMara |
ICPADS | 2 |
| 2002 | NULL convention multiply and accumulate unit with conditional rounding, scaling, and saturation
Scott C. Smith, Ronald F. DeMara, Jiann-Shiun Yuan, M. Hagedorn, Dennis Ferguson |
J. Syst. Archit. | 2 |
| 2001 | Delay-insensitive gate-level pipelining
Scott C. Smith, Ronald F. DeMara, Jiann-Shiun Yuan, M. Hagedorn, Dennis Ferguson |
Integr. | 2 |
| 1993 | A Parallel Computational Model for Integrated Speech and Natural Language UnderstandingabstractPresents a parallel approach for integrating speech and natural language understanding. The method emphasizes a hierarchically-structured knowledge base and memory-based parsing techniques. Processing is carried out by passing multiple markers in parallel through the knowledge base. Speech specific problems such as insertion, deletion, substitution, and word boundary detection have been analyzed and their parallel solutions are provided. Results on the SNAP-1 multiprocessor show an 80% sentence recognition rate for the Air Traffic Control (ATC) domain. Furthermore, speed-up of up to 15-fold is obtained from the parallel platform which provides response times of a few seconds per sentence for the ATC domain.> Dan I. Moldovan, Ronald F. DeMara |
IEEE Trans. Computers | 3 |
| 1993 | The SNAP-1 Parallel AI PrototypeabstractThe Semantic Network Array Processor (SNAP) is a parallel architecture for knowledge representation and reasoning that uses the marker-propagation paradigm. The primary application areas of SNAP are natural language understanding and speech processing. A first-generation SNAP-1 system has been designed and constructed using an array of 144 digital signal processors organized as 32 multiprocessing clusters with dedicated communication units, a tiered synchronization scheme, and multiported memory network. Issues in the design, performance, and scalability of a marker-propagation architecture are addressed.> Ronald F. DeMara, Dan I. Moldovan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1991 | Performance Indices for Parallel Marker-Propagation
Ronald F. DeMara, Dan I. Moldovan |
ICPP (1) | 1 |
| 1991 | The SNAP-1 Parallel AI PrototypeabstractArticle Free Access Share on The SNAP-1 parallel AI prototype Authors: R. F. DeMara Parallel Knowledge Processing Laboratory, Department of Electrical Engineering Systems, University of Southern California, Los Angeles, California Parallel Knowledge Processing Laboratory, Department of Electrical Engineering Systems, University of Southern California, Los Angeles, CaliforniaView Profile , D. I. Moldovan Parallel Knowledge Processing Laboratory, Department of Electrical Engineering Systems, University of Southern California, Los Angeles, California Parallel Knowledge Processing Laboratory, Department of Electrical Engineering Systems, University of Southern California, Los Angeles, CaliforniaView Profile Authors Info & Claims ISCA '91: Proceedings of the 18th annual international symposium on Computer architectureApril 1991 Pages 2–11https://doi.org/10.1145/115952.115954Published:01 April 1991Publication History 9citation341DownloadsMetricsTotal Citations9Total Downloads341Last 12 Months15Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Ronald F. DeMara, Dan I. Moldovan |
ISCA | 1 |