EDBT 2026 Demo / reviewers in the wild / expert
Brady Taylor
dblp:290/8552
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-2032-0960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effect of Capacitor Mismatch Nonlinearity on Inference Accuracy in Analog Compute-in-Memory ArchitecturesabstractCompute-in-memory (CIM) accelerators enhance neural network performance with efficient multiply- accumulate (MAC) operations. While capacitive arrays are often leveraged for MAC in CIM designs, circuit non-idealities can degrade accuracy but commonly go unreported in the literature. In this paper, we present a framework to analyze the impact of nonlinearity, specifically capacitor mismatch, on two existing capacitive designs—C2C and binary-weighted—and propose another called split. The split design divides the capacitive array into two binary-weighted arrays with a unit capacitor. This configuration is more area-efficient than the weighted and less susceptible to capacitor mismatch than the C2C. Our results show that for a simple dataset (MNIST) and neural network (MLP), inference accuracy is not significantly affected by nonlinearity, with the C2C (least linear) achieving nearly the same accuracy as the weighted (most linear). However, for more complex datasets and neural architectures, the C2C exhibits a significant accuracy loss, achieving only 9.74%, while the split and weighted achieve 99.06% and 99.29%, respectively. Abdulkarim Alorf, Brady Taylor, Yiran Chen 0001 |
ISCAS | 2 |
| 2025 | MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMsabstractWhile the tree-based machine learning (TBML) models exhibit superior performance compared to neural networks on tabular data and hold promise for energy-efficient acceleration using aCAM arrays, their ideal deployment on hardware with explicit exploitation of TBML structure and aCAM circuitry remains a challenging task. In this work, we present MonoSparse-CAM, a new CAM-based optimization technique that exploits TBML sparsity and monotonicity in CAM circuitry to further advance processing performance. Our results indicate that MonoSparse-CAM reduces energy consumption by upto to 28.56× compared to raw processing and by 18.51× compared to state-of-the-art techniques, while improving the efficiency of computation by at least 1.68×. Tergel Molom-Ochir, Brady Taylor, Hai Li 0001, Yiran Chen 0001 |
ISCAS | 2 |
| 2025 | Advancements in Content-Addressable Memory (CAM) Circuits: State-of-the-Art, Applications, and Future Directions in the AI DomainabstractContent-Addressable Memory (CAM) circuits, distinguished by their ability to accelerate data retrieval through a direct content-matching function, are increasingly crucial in the era of AI and increasing data computation. With the rise of AI models, hardware matching and hashing capabilities become essential, underscoring the need for a comprehensive survey of this evolving technology. This survey explores various CAM types across circuit designs and technologies, highlighting contributions to fields such as Machine Learning and genomics. We review 37 CAM cell designs, focusing on emerging trends in area and energy efficiency, pivotal for next-generation computing. Furthermore, we discuss current challenges and suggest future research directions in CAM technology. Tergel Molom-Ochir, Brady Taylor, Hai Li 0001, Yiran Chen 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Efficient and Robust Edge AI: Software, Hardware, and the Co-designabstractArtificial intelligence (AI) provides versatile capabilities in applications such as image classification and voice recognition that are most useful in edge or mobile computing settings. Shrinking these sophisticated algorithms into small form factors with minimal computing resources and power budgets requires innovation at several layers of abstraction: software, algorithmic, architectural, circuit, and device-level innovations. However, improvements to system efficiency may impact robustness and vice-versa. Therefore, a co-design framework is often necessary to customize a system for its given application. A system that prioritizes efficiency might use circuit-level innovations that introduce process variations or signal noise into the system, which may use software-level redundancy in order to compensate. In this tutorial, we will first examine various methods of improving efficiency and robustness in edge AI and their tradeoffs at each level of abstraction. Then, we will outline co-design techniques for designing efficient and robust edge AI systems, using federated learning as a specific example to illustrate the effectiveness of co-design. Bokyung Kim 0001, Shiyu Li 0001, Brady Taylor, Yiran Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceabstractAnalog in-memory-computing (IMC) is an attractive technique with a higher energy efficiency to process machine learning workloads. However, the analog computing scheme suffers from large interface circuit overhead. In this work, we propose a macro with a hybrid analog-digital mode computation to reduce the precision requirement of the interface circuit. Considering the distribution of the multiplication and accumulation (MAC) value, we propose a nonlinear transfer function of the computing circuits by only accurately computing low MAC value in the analog domain with a digital mode to deal with the high MAC value with smaller possibility. Silicon measurement results show that the proposed macro could achieve 160 GOPS/mm2 area efficiency and 25.5 TOPS/W for 8b/8b matrix computation. The architectural-level evaluation for real workloads shows that the proposed macro can achieve up to 2.92× higher energy efficiency than conventional analog IMC designs. Qilin Zheng, Ziru Li, Jonathan Hao-Cheng Ku, Yitu Wang, Brady Taylor, Deliang Fan, Yiran Chen 0001 |
DAC | 5 |
| 2022 | Research Progress on Memristor: From Synapses to Computing SystemsabstractAs the limits of transistor technology are approached, feature size in integrated circuit transistors has been reduced very near to the minimum physically-realizable channel length, and it has become increasingly difficult to meet expectations outlined by Moore’s law. As one of the most promising devices to replace transistors, memristors have many excellent properties that can be leveraged to develop new types of neural and non-von Neumann computing systems, which are expected to revolutionize information-processing technology. This survey provides a comparative overview of research progress on memristors. Different memristor synaptic devices are classified according to stimulation patterns and the working mechanisms of these various synaptic devices are analyzed in detail. Crossbar-based memristors have demonstrated advantages in physically executing vector-matrix multiplication and enabling highly power-efficient and area-efficient neuromorphic system designs. The extensive uses of crossbar-based memristors cover in-memory logic, vector-matrix multiplication, and many other fundamental computing operations. Furthermore, memristor-based architectures for efficient neural network training and inference have been studied. However, memristors have non-ideal properties due to programming inaccuracies and device imperfections from fabrication, which lead to error or mismatch in computed results. To build reliable memristor-based designs, circuit-level, algorithm-level, and system-level solutions to memristor reliability issues are being studied. To this end, state-of-the-art realizations of memristor crossbars, crossbar-based designs, and peripheral circuitry are presented, which show both promising full-system inference accuracy and excellent power efficiency in multiple tasks. Memristor in-situ learning benefits from high energy efficiency and biologically-imitative characteristics, which are conducive to further realizing hardware acceleration of cognitive learning. At present, the learning and training processes of brain-like networks are complex, presenting great challenges for network design and implementation. Xiaoxuan Yang 0001, Brady Taylor, Ailong Wu, Yiran Chen 0001, Leon O. Chua |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Neuromorphic Algorithm-hardware Codesign for Temporal Pattern LearningabstractNeuromorphic computing and spiking neural networks (SNN) mimic the behavior of biological systems and have drawn interest for their potential to perform cognitive tasks with high energy efficiency. However, some factors such as temporal dynamics and spike timings prove critical for information processing but are often ignored by existing works, limiting the performance and applications of neuromorphic computing. On one hand, due to the lack of effective SNN training algorithms, it is difficult to utilize the temporal neural dynamics. Many existing algorithms still treat neuron activation statistically. On the other hand, utilizing temporal neural dynamics also poses challenges to hardware design. Synapses exhibit temporal dynamics, serving as memory units that hold historical information, but are often simplified as a connection with weight. Most current models integrate synaptic activations in some storage medium to represent membrane potential and institute a hard reset of membrane potential after the neuron emits a spike. This is done for its simplicity in hardware, requiring only a “clear” signal to wipe the storage medium, but destroys temporal information stored in the neuron.In this work, we derive an efficient training algorithm for Leaky Integrate and Fire neurons, which is capable of training a SNN to learn complex spatial temporal patterns. We achieved competitive accuracy on two complex datasets. We also demonstrate the advantage of our model by a novel temporal pattern association task. Codesigned with this algorithm, we have developed a CMOS circuit implementation for a memristor-based network of neuron and synapses which retains critical neural dynamics with reduced complexity. This circuit implementation of the neuron model is simulated to demonstrate its ability to react to temporal spiking patterns with an adaptive threshold. Haowen Fang, Brady Taylor, Ziru Li, Zaidao Mei, Hai Li 0001, Qinru Qiu |
DAC | 2 |
| 2021 | 1S1R-Based Stable Learning through Single-Spike-Encoded Spike-Timing-Dependent PlasticityabstractSpike-timing-dependent plasticity (STDP) is emerging as a simple and biologically-plausible approach to learning, and specialized digital implementations are readily available. Memristor technology has been embraced as a much denser solution than digital static random-access memory (SRAM) implementations of STDP synapses, with plasticity capabilities built into the physics of these devices. One-selector-one-memristor (1S1R) arrays using volatile memristor devices as selectors are capable of the desired synaptic behavior using efficient spike-events, but previous literature has only explored the dynamics of single 1S1R synapses, or groups of synapses for single neurons. When placed in the context of an SNN, unintentional synapse disturbances are revealed that must be addressed. We present1a technique for STDP-based learning, enabled for dense 1S1R technology and utilizing efficient single-spike encoding. This technique leverages the array's dynamics to produce models that are stable, resilient to noise, and power-efficient. Brady Taylor, Amar Shrestha, Qinru Qiu, Hai Li 0001 |
ISCAS | 1 |