EDBT 2026 Demo / reviewers in the wild / expert
Cory E. Merkel
dblp:68/11163
· DBLP profile ↗
22ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0003-0252-7829ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Timing Matters: Delay Fault Characterization and Testing in SNN AcceleratorsabstractSpiking neural network (SNN) accelerators are vulnerable to non-idealities arising from aging, process variation, radiation, and extreme operating conditions. While delay faults are well studied in conventional VLSI testing, their impact and detection in neuromorphic accelerators remain challenging because the event-driven dynamics of SNNs, complex temporal fault propagation, and in-memory computation obscure fault effects and complicate test generation and observability. This paper characterizes the effects of four timing-related faults in SNNs: dendritic propagation delay, axonal propagation delay, refractory delay, and spike-response delay. We further address a key challenge in resource-constrained neuromorphic testing: efficient test-pattern generation and fault evaluation. We propose a delay-fault testing methodology based on highly sparse test stimuli and introduce Mean Absolute Membrane Potential Deviation (MAMPD) as a fault evaluation metric, motivated by the hypothesis that hardware faults induce measurable distribution shifts in neuronal activity. Experimental results show that single delay faults exhibit pronounced tail risk, where a small fraction of faults causes severe accuracy degradation, whereas multiple simultaneous faults pose an even greater overall threat. Across the evaluated functionally meaningful delay-fault settings, the proposed method achieves 100% fault coverage of functionally impactful delays using sparse test matrices with up to 99% sparsity, demonstrating that high fault detectability can be obtained with very low test overhead. Osita Ukwuaba, Cory E. Merkel |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Power Analysis Attacks on NVM Crossbar-based Neuromorphic SystemsabstractThis paper proposes a new adversarial attack strategy against neuromorphic systems using analysis of power consumption. Specifically, we show that neuromorphic designs based on non-volatile memory crossbars can leak important information about loss sensitivity in their power profile. Adversaries can use this information to craft evasion attacks even if they don’t know the dataset that the model was trained on. In our experiments, we show that these types of attacks are effective against both single-layer and multilayer neuromorphic implementations of neural networks, and they can be made query-efficient through Bayesian optimization. We also provide theoretical insights into the relationship between the loss sensitivity and the power consumption measurements, showing that, for single-layer networks, the correlation coefficient of these two metrics scales inversely with the square root of the input size. Finally, this paper proposes that low bitwidth quantization could be an effective defense strategy against the class of attacks discussed herein. Cory E. Merkel, Allen Su |
Neural Process. Lett. | 1 |
| 2024 | Design Space Exploration of Memristor-based Delay Cells for Time-domain Neuromorphic ComputingabstractWe present a memristor-based delay cell design for time-domain neuromorphic computing inference. Each cell computes partial dot product operations using current summation, and multiple cells chained together complete the computation by accumulating their delays. The design is analyzed over process, voltage, and temperature variations to gain insight into the impact of the cell size (number of inputs) on performance and robustness. Results based on a 64-input neuron show that, while smaller cell sizes have better dynamic range (> 5), larger cells have significantly reduced delay (≈ 25), power× consumption (≈ 8×), and transistor count (≈ 3×). We× also find that there is a strong dependence between cell size and variability at reduced power supply voltage. Along with our analyses, we provide a detailed discussion of these tradeoffs to help designers choose the optimal configuration for their applications. Hagar Hendy, Karsten Bergthold, Tejasvi Das, Cory E. Merkel |
ISCAS | 4 |
| 2023 | Exploiting Logic Locking for a Neural Trojan Attack on Machine Learning AcceleratorsabstractLogic locking has been proposed to safeguard intellectual property (IP) during chip fabrication. Logic locking techniques protect hardware IP by making a subset of combinational modules in a design dependent on a secret key that is withheld from untrusted parties. If an incorrect secret key is used, a set of deterministic errors is produced in locked modules, restricting unauthorized use. A common target for logic locking is neural accelerators, especially as machine-learning-as-a-service becomes more prevalent. In this work, we explore how logic locking can be used to compromise the security of a neural accelerator it protects. Specifically, we show how the deterministic errors caused by incorrect keys can be harnessed to produce neural-trojan-style backdoors. To do so, we first outline a motivational attack scenario where a carefully chosen incorrect key, which we call a trojan key, produces misclassifications for an attacker-specified input class in a locked accelerator. We then develop a theoretically-robust attack methodology to automatically identify trojan keys. To evaluate this attack, we launch it on several locked accelerators. In our largest benchmark accelerator, our attack identified a trojan key that caused a 74% decrease in classification accuracy for attacker-specified trigger inputs, while degrading accuracy by only 1.7% for other inputs on average. Hongye Xu, Dongfang Liu, Cory E. Merkel, Michael Zuzak |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | A Bio-Inspired Computational Astrocyte Model for Spiking Neural NetworksabstractThe mammalian brain is the most capable and complex computing entity known today, and for years there has been research focused on reproducing the brain's capabilities. An early example of this endeavor was the perceptron which has become a core building block of neural network models in the deep learning era. Despite many successes achieved through deep learning, these networks behave much differently than their biological counterparts. In a search for improvements to things like training time and dataset size, power consumption, noise, and adversarial input tolerance, research is looking towards the brain, and bio-inspired computing models. Spiking neural networks (SNNs) take a step closer to biology in their operation and are the focus of much research, geared towards reproducing some of these novel features of the brain. As of yet, SNNs have not reached their full potential. This work explores the advancement of SNNs though introduction of a novel astrocyte model. Astrocytes, initially thought to be passive support cells in the brain are now known to actively participate in neural processing. The proposed astrocyte model is geared towards synaptic plasticity and is shown to generalize and extend spike-timing dependent plasticity (STDP). In addition, the model supports plasticity-focused multi-synapse integration. Logical AND was used as a test function for multi-synapse astrocyte plasticity, and convergence for a single leaky integrate and fire (LIF) neuron with 2, 3, and 4 synapses was demonstrated. The plasticity model is general and many other functions could presumably be used in-place of AND, taking advantage of a multi-synapse view. This leaves a solid path forward for future work. Jacob Kiggins, J. David Schaffer, Cory E. Merkel |
IJCNN | 3 |
| 2022 | Accelerating the Training of Single Layer Binary Neural Networks using the HHL Quantum AlgorithmabstractBinary Neural Networks are a promising technique for implementing efficient deep models with reduced storage and computational requirements. The training of these is however, still a compute-intensive problem that grows drastically with the layer size and data input. At the core of this calculation is the linear regression problem. The Harrow-Hassidim-Lloyd (HHL) quantum algorithm has gained relevance thanks to its promise of providing a quantum state containing the solution of a linear system of equations. The solution is encoded in superposition at the output of a quantum circuit. Although this seems to provide the answer to the linear regression problem for the training neural networks, it also comes with multiple, difficult-to-avoid hurdles. This paper shows, however, that useful information can be extracted from the quantum-mechanical implementation of HHL, and used to reduce the complexity of finding the solution on the classical side. Sonia Lopez Alarcon, Cory E. Merkel, Martin Hoffnagle, Sabrina Ly, Alejandro Pozas-Kerstjens |
ICCD | 2 |
| 2021 | On the Adversarial Robustness of Quantized Neural NetworksabstractReducing the size of neural network models is a critical step in moving AI from a cloud-centric to an edge-centric (i.e. on-device) compute paradigm. This shift from cloud to edge is motivated by a number of factors including reduced latency, improved security, and higher flexibility of AI algorithms across several application domains (e.g. transportation, healthcare, defense, etc.). However, it is currently unclear how model compression techniques may affect the robustness of AI algorithms against adversarial attacks. This paper explores the effect of quantization, one of the most common compression techniques, on the adversarial robustness of neural networks. Specifically, we investigate and model the accuracy of quantized neural networks on adversarially-perturbed images. Results indicate that for simple gradient-based attacks, quantization can either improve or degrade adversarial robustness depending on the attack strength. Micah Gorsline, Cory E. Merkel |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | Model Extraction and Adversarial Attacks on Neural Networks Using Switching Power Information
Tommy Li, Cory E. Merkel |
ICANN (1) | 2 |
| 2020 | A neuromorphic SLAM architecture using gated-memristive synapses
Alexander Jones 0001, Andrew Rush, Cory E. Merkel, Eric Herrmann, Ajey P. Jacob, Clare Thiem, Rashmi Jha |
Neurocomputing | 3 |
| 2019 | Exploiting Randomness in Deep Learning AlgorithmsabstractThe recent surge of interest in using deep neural networks for real-world tasks has led to training complex networks with billions of parameters that use enormous amounts of training data. Performing backpropagation in these deep networks is time consuming and requires large amount of resources that is usually limited by the underlying hardware. In order to move towards agile deep learning, we are motivated to exploit randomness in the networks. In this work, we explore the effects of utilizing random weights in convolutional neural networks. This is achieved through random initialization of weights and by freezing them. The training occurs only in the output layer. We also propose a novel weight distribution method based on the sum of sinusoids for random convolutional neural networks. Our experiments show that by leaving the weights random in convolutional neural networks relatively high performance can be achieved for MSTAR and CIFAR-10 datasets. Hamed Fatemi Langroudi, Cory E. Merkel, Humza Syed, Dhireesha Kudithipudi |
IJCNN | 2 |
| 2018 | An FPGA Implementation of a Time Delay Reservoir Using Stochastic LogicabstractThis article presents and demonstrates a stochastic logic time delay reservoir design in FPGA hardware. The reservoir network approach is analyzed using a number of metrics, such as kernel quality, generalization rank, and performance on simple benchmarks and is also compared to a deterministic design. A novel re-seeding method is introduced to reduce the adverse effects of stochastic noise, which may also be implemented in other stochastic logic reservoir computing designs, such as echo state networks. Benchmark results indicate that the proposed design performs well on noise-tolerant classification problems, but more work needs to be done to improve the stochastic logic time delay reservoir's robustness for regression problems. In addition, we show that the stochastic design can significantly reduce area cost if the conversion between binary and stochastic representations is implemented efficiently. Lisa Loomis, Nathan R. McDonald, Cory E. Merkel |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2017 | Design of a time delay reservoir using stochastic logic: A feasibility studyabstractThis paper presents a stochastic logic time delay reservoir design. The reservoir is analyzed using a number of metrics, such as kernel quality, generalization rank, performance on simple benchmarks, and is also compared to a deterministic design. A novel re-seeding method is introduced to reduce the adverse effects of stochastic noise, which may also be implemented in other stochastic logic reservoir computing designs, such as echo state networks. Benchmark results indicate that the proposed design performs well on noise-tolerant classification problems, but more work needs to be done to improve the stochastic logic time delay reservoir's robustness for regression problems. Cory E. Merkel |
IJCNN | 1 |
| 2017 | Stochastic CBRAM-Based Neuromorphic Time Series Prediction SystemabstractIn this research, we present a Conductive-Bridge RAM (CBRAM)-based neuromorphic system which efficiently addresses time series prediction. We propose a new (i) voltage-mode, stochastic, multiweight synapse circuit based on experimental bi-stable CBRAM devices, (ii) a voltage-mode neuron circuit based on the concept of charge sharing, and (iii) an optimized training methodology powered by a stochastic implementation of the Least-Mean-Squares (SLMS) training rule. To validate the proposed design, we use time series prediction for short-term electrical load forecasting in smart grids. Our system is able to forecast hourly electrical loads with a mean accuracy of 96%, an estimated power dissipation of 15 μW, and area of 14.5 μm 2 at 65 nm CMOS technology. Cory E. Merkel, Dhireesha Kudithipudi, Manan Suri, Bryant T. Wysocki |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2016 | A design of HTM spatial pooler for face recognition using memristor-CMOS hybrid circuitsabstractHierarchical Temporal Memory (HTM) is a machine learning algorithm that is inspired from the working principles of the neocortex, capable of learning, inference, and prediction for bit-encoded inputs. Spatial pooler is an integral part of HTM that is capable of learning and classifying visual data such as objects in images. In this paper, we propose a memristor-CMOS circuit design of spatial pooler and exploit memristors capabilities for emulating the synapses, where the strength of the weights is represented by the state of the memristor. The proposed design is validated on a challenging application of single image per person face recognition problem using AR database resulting in a recognition accuracy of 80%. Timur Ibrayev, Alex James 0001, Cory E. Merkel, Dhireesha Kudithipudi |
ISCAS | 3 |
| 2015 | Design and analysis of neuromemristive echo state networks with limited-precision synapsesabstractEcho state networks (ESNs) are gaining popularity as a method for recognizing patterns in time series data. ESNs are random, recurrent neural network topologies that are able to integrate temporal data over short time windows by operating on the edge of chaos. In this paper, we explore the design of a hardware ESN with bi-stable memristor-based synapses. Hybrid CMOS/memristor hardware implementations of ESNs are able to exploit non-linear device physics, improving power consumption, and boosting performance over software approaches. However, the digital nature of most experimental memristors places a limit on the precision of weight states in the ESN's readout layer. In spite of this, we show that ESNs with only 5 different readout layer weight states can acheive 67% accuracy in spoken digit recognition tasks. Colin Donahue, Cory E. Merkel, Qutaiba Saleh, Levs Dolgovs, Yu Kee Ooi, Dhireesha Kudithipudi, Bryant T. Wysocki |
CISDA | 2 |
| 2015 | Memristive computational architecture of an echo state network for real-time speech-emotion recognitionabstractEcho state neural networks (ESNs) provide an efficient classification technique for spatiotemporal signals. The feedback connections in the ESN topology enable feature extraction of both spatial and temporal components in time series data. This property has been used in several application domains such as image and video analysis, anomaly detection, and speech recognition. In this research, we explore a hardware architecture for realizing ESN efficiently in power-constrained devices. Specifically, we propose a scalable computational architecture applied to speech-emotion recognition. Two different topologies are explored, with memristive synapses. The simulation results are promising with a classification accuracy of ≈ 96% for two distinct emotion statuses. Qutaiba Saleh, Cory E. Merkel, Dhireesha Kudithipudi, Bryant T. Wysocki |
CISDA | 2 |
| 2014 | A current-mode CMOS/memristor hybrid implementation of an extreme learning machineabstractIn this work, we propose a current-mode CMOS/memristor hybrid implementation of an extreme learning machine (ELM) architecture. We present novel circuit designs for linear, sigmoid,and threshold neuronal activation functions, as well as memristor-based bipolar synaptic weighting. In addition, this work proposes a stochastic version of the least-mean-squares (LMS) training algorithm for adapting the weights between the ELM's hidden and output layers. We simulated our top-level ELM architecture using Cadence AMS Designer with 45 nm CMOS models and an empirical piecewise linear memristor model based on experimental data from an HfOx device. With 10 hidden node neurons, the ELM was able to learn a 2-input XOR function after 150 training epochs. Cory E. Merkel, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | Temperature Sensing RRAM Architecture for 3-D ICsabstract3-D integrated circuits, or 3-D ICs, have gained significant attention in the research community over the past few years. This has primarily been motivated by their enhanced power, performance, and functionality over planar CMOS ICs. However, thermal management remains a key challenge in these devices due to the impedance of heat flow that results from die stacking. In this paper, we address this challenge by utilizing a temperature sensing resistive random access memory (TSRRAM) which can generate accurate thermal profiles to gauge the heat distribution within the 3-D IC. The architecture enables each RRAM switching element in the memory die to be used both as a memory bit and a temperature sensor. We simulated our TSRRAM design as an L2 cache for an alpha 21364 processor. We used a customized simulation framework to test the design accuracy and performance over several SPEC2000 CPU benchmarks, and achieved a 2.14 K mean error and an eight-cycle performance overhead with a 4-kB L2 cache size. Furthermore, we show that active sensing methods can be employed to achieve 100% coverage of global hot spot temperatures. Cory E. Merkel, Dhireesha Kudithipudi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Periodic activation functions in memristor-based analog neural networksabstractThis work explores the use of periodic activation functions in memristor-based analog neural networks. We propose a hardware neuron based on a folding amplifier that produces a periodic output voltage. Furthermore, the amplifier's fold factor be adjusted to change the number of low-to-high or high-to-low output voltage transitions. We also propose a memristor-based synapse circuit and training circuitry for realizing the Perceptron learning rule. Behavioral models of our circuits were developed for simulating a single-layer, single-output feedforward neural network. The network was trained to detect the edges of a grayscale image. Our results show that neurons with a single fold-with an activation function similar to a sigmoidal activation function-perform the worst for this application, since they are unable to learn functions with multiple decision boundaries. Conversely, the 4-fold neuron performs the best (up to ≈65% better than the 1-fold neuron), as its activation function is periodic, and it is able to learn functions with four decision boundaries. Cory E. Merkel, Dhireesha Kudithipudi, Nick Sereni |
IJCNN | 1 |
| 2013 | Memristor-Based Neural Logic Blocks for Nonlinearly Separable FunctionsabstractNeural logic blocks (NLBs) enable the realization of biologically inspired reconfigurable hardware. Networks of NLBs can be trained to perform complex computations such as multilevel Boolean logic and optical character recognition (OCR) in an area- and energy-efficient manner. Recently, several groups have proposed perceptron-based NLB designs with thin-film memristor synapses. These designs are implemented using a static threshold activation function, limiting the set of learnable functions to be linearly separable. In this work, we propose two NLB designs-robust adaptive NLB (RANLB) and multithreshold NLB (MTNLB)-which overcome this limitation by allowing the effective activation function to be adapted during the training process. Consequently, both designs enable any logic function to be implemented in a single-layer NLB network. The proposed NLBs are designed, simulated, and trained to implement ISCAS-85 benchmark circuits, as well as OCR. The MTNLB achieves 90 percent improvement in the energy delay product (EDP) over lookup table (LUT)-based implementations of the ISCAS-85 benchmarks and up to a 99 percent improvement over a previous NLB implementation. As a compromise, the RANLB provides a smaller EDP improvement, but has an average training time of only ≈ 4 cycles for 4-input logic functions, compared to the MTNLBs ≈ 8-cycle average training time. Michael Soltiz, Dhireesha Kudithipudi, Cory E. Merkel, Garrett S. Rose, Robinson E. Pino |
IEEE Trans. Computers | 3 |
| 2012 | Design-time performance evaluation of thermal management policies for SRAM and RRAM based 3D MPSoCsabstract3D-ICs hold significant promise for future generation multi processor systems-on-chip due to their potential for increased performance, decreased power, heterogeneous integration, and reduced cost over planar ICs. However, the vertical integration of these structures exacerbates the heat dissipation and run-time thermal management issues. There have been a number of design- and run-time thermal management policies proposed, but few focus on examining overall system performance. Additionally, the heterogeneity of 3D-ICs allows for the integration of novel technologies, such as resistive random access memories (RRAMs), which offer higher density and lower power than traditional CMOS memory technologies. Our work presents a flexible design-time simulation framework to evaluate system performance and thermal profiles of 3D MPSoCs. We utilize this framework to study the effect of three dynamic thermal management policies (air-cooled load balancing, liquid-cooled load balancing, and air-cooled DVFS) on system performance and die temperature for multi-tiered 3D MPSoCs utilizing SRAM and RRAM-based L2 caches. We find that RRAM-based caches lower overall average maximum temperatures by 120 K and 24 K for air and liquid cooling systems, respectively (when compared to SRAM-based caches), at a worst-case performance delay of 47% and best-case delay of 13% for the parallel shared-memory benchmarks studied. David Brenner, Cory E. Merkel, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Reconfigurable N-level memristor memory designabstractMemristive devices have gained significant research attention lately because of their unique properties and wide application spectrum. In particular, memristor-based resistive random access memory (RRAM) offers the high density, low power, and low volatility required for next-generation non-volatile memory. The ability to program memristive devices into several different resistance states has also led to the proposal of multilevel RRAM. This work analyzes the application of thinfilm memristors as N-level RRAM elements. The tradeoffs between the number of memory levels and each RRAM element's reliability will be discussed. A metric is proposed to rate each RRAM element in the presence of process variations. A memory architecture is also presented which allows the number of memory levels to be reconfigured based on different application characteristics. The proposed architecture can achieve a write time speedup of 5.9 over other memristor memory architectures with 80% ion mobility degradation. Cory E. Merkel, Nakul Nagpal, Sindhura Mandalapu, Dhireesha Kudithipudi |
IJCNN | 1 |