Xuanyao Fong

dblp:31/3524 · DBLP profile ↗
← Back
30ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-5939-7389ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PRIMAL: Processing-In-Memory based Low-Rank Adaptation for LLM Inference Accelerator
abstract
This paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM processing elements (PEs), interconnected by 2D-mesh inter-PE computational network (IPCN). A novel SRAM reprogramming and power gating (SRPG) scheme enables pipelined LoRA updates and sub-linear power scaling by overlapping reconfiguration with computation and gating idle resources. PRIMAL employs optimized spatial mapping and dataflow orchestration to minimize communication overhead, and achieves $1.5\times$ throughput and $25\times$ energy efficiency over NVIDIA H100 with LoRA rank 8 (Q,V) on Llama-13B.
Yue Jiet Chong, Yimin Wang 0001, Xuanyao Fong
ISCAS4
2026 LOKI: An LLM Accelerator with Optimized KV Cache and In-Memory Computing Cores
Avadh Harkishanka, Abhishek Tyagi, Xuanyao Fong
ISCAS3
2026 Topology-Mapping Co-Design for Scalable LLM Accelerators with Hierarchical Mesh-of-Trees NoC
Yimin Wang 0001, Yue Jiet Chong, Xuanyao Fong
ISCAS4
2026 Cross-Layer Evaluation for On-Chip Training of BEOL FeTFT-based CIM M3D Accelerator
Jianze Wang, Yimin Wang 0001, Xuanyao Fong
ISCAS4
2025 Massively Parallel Continuous Local Search for Hybrid SAT Solving on GPUs
abstract
Although state-of-the-art (SOTA) SAT solvers based on conflict-driven clause learning (CDCL) have achieved remarkable engineering success, their sequential nature limits the parallelism that may be extracted for acceleration on platforms such as the graphics processing unit (GPU). In this work, we propose FastFourierSAT, a highly parallel hybrid SAT solver based on gradient-driven continuous local search (CLS). This is achieved by a parallel algorithm inspired by the fast Fourier transform (FFT)-based convolution for computing the elementary symmetric polynomials (ESPs), which is the major computational task in previous CLS methods. The complexity of our algorithm matches the best previous result. Furthermore, the substantial parallelism inherent in our algorithm can leverage the GPU for acceleration, demonstrating significant improvement over the previous CLS approaches. FastFourierSAT is compared with a wide set of SOTA parallel SAT solvers on extensive benchmarks including combinatorial and industrial problems. Results show that FastFourierSAT computes the gradient 100+ times faster than previous prototypes on CPU. Moreover, FastFourierSAT solves most instances and demonstrates promising performance on larger-size instances.
Yunuo Cen, Zhiwei Zhang 0001, Xuanyao Fong
AAAI3
2025 LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
abstract
Large language model (LLM) inference has been a prevalent demand in daily life and industries. The large tensor sizes and computing complexities in LLMs have brought challenges to memory, computing, and databus. This paper proposes a computation/memory/communication co-designed non-von Neumann accelerator by aggregating processing-in-memory (PIM) and computational network-on-chip (NoC), termed LEAP. The matrix multiplications in LLMs are assigned to PIM or NoC based on the data dynamicity to maximize data locality. Model partition and mapping are optimized by heuristic design space exploration. Dedicated fine-grained parallelism and tiling techniques enable high-throughput dataflow across the distributed resources in PIM and NoC. The architecture is evaluated on Llama 1B/8B/13B models and shows ~2.55× throughput (tokens/sec) improvement and ~71.94× energy efficiency (tokens/Joule) boost compared to the A100 GPU.
Yimin Wang 0001, Yue Jiet Chong, Xuanyao Fong
ICCAD3
2025 A Pseudo-Boolean Encoding Method for Rectilinear Steiner Arborescence in VLSI Routing
abstract
This paper presents a novel Pseudo-Boolean Optimization (PBO) framework to solve the Rectilinear Steiner Minimum Arborescence (RSMA) problem, with a focus on the All-Quadrant RSMA scenario. Our framework is designed to compute exact solutions efficiently, even in the presence of isothetic rectilinear obstacles. We also introduce a reduction technique that can reduce more than 17% of the encoded variables and constraints to improve solving efficiency. Our method guarantees optimality and offers flexibility and adaptability for large-scale instances. The experimental results indicate that the proposed framework solves 82.3% of 40-point instances and 64.3% of 50-point instances within 300 seconds.
Yunuo Cen, Xuanyao Fong
ISCAS3
2024 Design Framework for Ising Machines with Bistable Latch-Based Spins and All-to-All Resistive Coupling
abstract
Memristive crossbars have been proposed as a promising pathway to enabling all-to-all connections for a variety of applications. One such application is the latch-based dynamical Ising machine, which have been proposed for solving combinatorial optimization problems. However, the impact of the design of the memristive crossbar on the performance of the latch-based Ising machine remains unclear. We present a SPICE-level model and simulation framework for analyzing and evaluating the latch-based dynamical Ising machine that use memristive crossbars to achieve all-to-all coupling between Ising spins implemented as latches. Our design space exploration reveals that the solution quality is highly sensitive to the circuit and device design parameters and sizes of the problems. An effective metric that can capture the effect of several design parameters on the statistical results is then proposed to quantify the system functionality. Finally, optimization techniques based on our proposed metric that models the effect of several design parameters on the solution quality of the Ising machine are proposed and demonstrated on MaxCut solving. Our evaluation results on a 128×128 crossbar array-based design show >98% solution quality across a wide range of problem sizes is achievable.
Yimin Wang 0001, Yunuo Cen, Xuanyao Fong
ISCAS3
2024 Energy-Efficient Ising Machines Using Capacitance-Coupled Latches for MaxCut Solving
abstract
Latch-based Ising machines (LIMs) have aroused research interest due to their speed and area efficiency for solving combinatorial optimization problems (COPs). However, existing LIMs based on resistive coupling are sensitive to many design parameters and suffer from solution quality degradation. This paper explores a capacitive-coupling approach that improves the robustness of the LIM to variations in the system design parameters. Evaluation results show that the LIMs using capacitive coupling can solve the MaxCut problem with high solution quality in a ∼1000× wider range of coupling strength compared to the resistive coupling approach. It also achieves an energy-time product of 313.5 nJ•s, with an 84.3% and 86.4% reduction compared to the state-of-the-art resistance-coupled and capacitance-coupled counterparts respectively.
Yimin Wang 0001, Xuanyao Fong
ISCAS2
2024 Transposable Memory Based on the Ferroelectric Field-Effect Transistor
abstract
Non-volatile ferroelectric field-effect transistor (FeFET) technology is a promising CMOS process compatible solution for fast, energy efficient on-chip memories that are needed to enable ubiquitous deployment of AI. Existing memories are linear by design and the bandwidth may not be sufficient to support transformers that underlie the now-popular Large Language Models (LLMs) for generative AI, which require fast accesses to matrices and their transposes. In this work, two designs of transposable FeFET memories are proposed to address this challenge. We evaluate our designs using compact models for the FeFET devices, which we have calibrated to experimentally measured device characterization data. We demonstrate that 1 ns read latencies are possible, which underscores the potential of FeFET technology for AI accelerators. The area efficiency of our proposed dual-gated transposable memory and the speed of transpose operations are studied, showing that it occupies 30.02 F2cell area and is capable of maximum 82.24× faster transpose operation compared to non-transposable memory.
Jianze Wang, Yimin Wang 0001, Leming Jiao, Xiaolin Wang 0007, Xuanyao Fong
ISCAS8
2023 A Semi-Supervised Learning Method for Spiking Neural Networks Based on Pseudo-Labeling
abstract
Supervised learning methods have demonstrated state-of-the-art performance for spiking neural network (SNN) but require a large amount of annotated training data, which may be expensive to obtain. To address this, we propose a semi-supervised learning method based on pseudo-labeling that enables SNN training using a small number of annotated training samples. Our proposed method outperforms the existing semi-supervised learning approaches for SNN by more than 17% on the MNIST dataset when 100 annotated samples are used. Our proposed spike-based method is hardware friendly, can be incorporated to various SNN models, and does not involve intensive computations, which is desirable for applications on internet of things (IoT) devices that have tight constraints on hardware resources and energy consumption.
Thao N. N. Nguyen, Bharadwaj Veeravalli, Xuanyao Fong
IJCNN3
2022 An FPGA-Based Co-Processor for Spiking Neural Networks with On-Chip STDP-Based Learning
abstract
In this paper, we report on the design of a neuromorphic co-processor on a Field Programmable Gate Array (FPGA) platform that is capable of emulating Spiking Neural Networks (SNNs) with support for on-chip unsupervised learning. One defining feature of our design is that the SNN configuration is defined entirely in the software executed by our neuromorphic co-processor. Evaluation on the FPGA platform shows that our design consumes a small amount of hardware resources and on-chip memory storage (438.75 kB). In addition, the inference and the on-chip learning in a deep convolutional SNN emulated on our FPGA implementation are $10.5vf \times$ and $8.6 \times$ faster than the implementation on high performance x86 CPU. Moreover, we demonstrate the ability of our neuromorphic co-processor to perform the on-chip learning on an object recognition task (based on the Caltech-101 dataset).
Thao N. N. Nguyen, Bharadwaj Veeravalli, Xuanyao Fong
ISCAS3
2020 Aggressive Leakage Current Reduction for Embedded MRAM Using Block-Level Power Gating
abstract
This paper exploits circuit techniques to realize a power-gated MRAM with the instant-on characteristic. Each block in a large-capacity MRAM is gated by a separate power transistor, allowing it to be fully turned on after only 1.3 ns. As a result, the whole MRAM is always in sleep mode, except the selected block. This eliminates the need for access pattern prediction as well as the requirement that the MRAM must be in idle for a significant of time before it is put in sleep or deep sleep mode. Our simulation shows that even in the worst-case-scenario, 86% of leakage current is saved. In typical cases, there are 96.5% leakage and 69% total power reduction. Our proposed scheme's implementation is straight forward and incurs less than 0.5% area overhead, including power transistors and control circuits.
Anh-Tuan Do, Xuanyao Fong, Fei Li 0015
IECON2
2018 Domain Wall Motion-based XOR-like Activation Unit With A Programmable Threshold
abstract
Spintronic devices promise an excellent opportunity for implementing ultra-low power neuromorphic platforms due to the inherent correspondence between their physical characteristics and the required neuronal, synaptic functionalities. Neuromorphic circuits using domain wall motion-based threshold neurons have been demonstrated in previous studies. However, threshold neurons are unable to realize linearly inseparable functions. Our work addresses this challenge by proposing a new domain wall motion-based neural activation unit with XOR-like activation function. We also develop a new learning algorithm for neurons with this activation unit. Offline training is performed on real-world datasets from the UCI machine learning repository. Neuromorphic circuits corresponding to these datasets are also simulated. The results suggest femto-Joule range energy consumption of a neuron with the proposed activation unit and 1.08×-1.82× lower misclassification rate (MCR) of the proposed algorithm in comparison to the traditional perceptron learning algorithm.
Tarun Vatwani, Anupam Chattopadhyay, Arindam Basu, Xuanyao Fong
IJCNN5
2017 Fast and Disturb-Free Nonvolatile Flip-Flop Using Complementary Polarizer MTJ
abstract
Nonvolatile flip-flop (NVFF) using spin-transfer torque magnetic tunnel junctions (STT-MTJs) has been proposed to enable line-grain power gating systems. However, the STT-MTJ-based NVFF (STT-NVFF) may not perform fast backup and disturb-free restore operations. We propose a new NVFF using complementary polarizer MTJ (CPMTJ) to alleviate these limitations. Our proposed NVFF exploits the CPMTJ structure for fast- and low-energy backup operation. The estimated backup delay is less than 10 ns in 7-nm node FinFET technology with CPMTJ size of 12 nm × 33 nm in a rectangular shape. Furthermore, during the restore operation, CPMTJ provides guaranteed disturb-free sensing, since the disturb torque from the two complementary pinned layers of CPMTJ cancels each other. The simulation results show 2 times improvement in the backup delay with higher restore-disturb margin compared with STT-NVFF.
Yeongkyo Seo, Xuanyao Fong, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A low-voltage, low power STDP synapse implementation using domain-wall magnets for spiking neural networks
abstract
Online, real-time learning in neuromorphic circuits have been implemented through variants of Spike Time Dependent Plasticity (STDP). Current implementations have used either floating-gate devices or memristors to implement such learning synapses together with non-volatile storage. However, these approaches require high voltages (≈ 3-12V) for weight update and entail high energy for learning (≈ 4-30pJ/write). We present a domain wall memory based low-voltage, low-energy STDP synapse that can operate with a power supply as low as 0.8V and update the weight at ≈ 40fJ/write. Device level simulations are performed to prove its feasibility. Its use in associative learning is also demonstrated by using neurons with dendritic branches to classify spike patterns from MNIST dataset.
Govind Narasimman, Subhrajit Roy, Xuanyao Fong, Kaushik Roy 0001, Chip-Hong Chang, Arindam Basu
ISCAS3
2016 Yield, Area, and Energy Optimization in STT-MRAMs Using Failure-Aware ECC
abstract
Spin-Transfer Torque MRAMs are attractive due to their non-volatility, high density, and zero leakage. However, STT-MRAMs suffer from poor reliability due to shared read and write paths. Additionally, conflicting requirements for data retention and writeability (both related to the energy barrier height of the storage device) makes design more challenging. Furthermore, the energy barrier height depends on the geometry of the storage. Any variations in the geometry of the storage device lead to variations in the energy barrier height. In order to address the poor reliability of STT-MRAMs, usage of Error Correcting Codes (ECC) has been proposed. Unlike traditional CMOS memory technologies, ECC is expected to correct both soft and hard errors in STT-MRAMs. To achieve acceptable yield with low write power, stronger ECC is required, resulting in increased number of encoded bits and degraded memory capacity. In this article, we propose Failure-aware ECC (FaECC), which masks permanent faults while maintaining the same correction capability for soft errors without increased number of encoded bits. Furthermore, we investigate the impact of process variations on run-time reliability of STT-MRAMs. In order to analyze the effectiveness of our methodology, we developed a cross-layer simulation framework that consists of device, circuit and array level analysis of STT-MRAM memory arrays. Our results show that using FaECC relaxes the requirements on the energy barrier height, which reduces the write energy and results in smaller access transistor size and memory array area.
Zoha Pajouhi, Xuanyao Fong, Anand Raghunathan, Kaushik Roy 0001
ACM J. Emerg. Technol. Comput. Syst.2
2016 Spin-Transfer Torque Memories: Devices, Circuits, and Systems
abstract
Spin-transfer torque magnetic memory (STT-MRAM) has gained significant research interest due to its nonvolatility and zero standby leakage, near unlimited endurance, excellent integration density, acceptable read and write performance, and compatibility with CMOS process technology. However, several obstacles need to be overcome for STT-MRAM to become the universal memory technology. This paper first reviews the fundamentals of STT-MRAM and discusses key experimental breakthroughs. The state of the art in STT-MRAM is then discussed, beginning with the device design concepts and challenges. The corresponding bit-cell design solutions are also presented, followed by the STT-MRAM cache architectures suitable for on-chip applications.
Xuanyao Fong, Yusung Kim 0002, Rangharajan Venkatesan, Sri Harsha Choday, Anand Raghunathan, Kaushik Roy 0001
Proc. IEEE1
2016 Spin-Transfer Torque Devices for Logic and Memory: Prospects and Perspectives
abstract
As CMOS technology begins to face significant scaling challenges, considerable research efforts are being directed to investigate alternative device technologies that can serve as a replacement for CMOS. Spintronic devices, which utilize the spin of electrons as the state variable for computation, have recently emerged as one of the leading candidates for post-CMOS technology. Recent experiments have shown that a nano-magnet can be switched by a spin-polarized current and this has led to a number of novel device proposals over the past few years. In this paper, we provide a review of different mechanisms that manipulate the state of a nano-magnet using current-induced spin-transfer torque and demonstrate how such mechanisms have been engineered to develop device structures for energy-efficient on-chip memory and logic.
Xuanyao Fong, Yusung Kim 0002, Karthik Yogendra, Deliang Fan, Abhronil Sengupta, Anand Raghunathan, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Embedding Read-Only Memory in Spin-Transfer Torque MRAM-Based On-Chip Caches
abstract
We propose a design technique for embedding read-only memory (ROM) in spin-transfer torque MRAM (STT-MRAM) arrays by adding an extra bit-line in every column of the array. RAM and ROM data, which can be different, are stored in the same bitcell and the ROM capacity may be as large as the RAM capacity. Furthermore, our proposed ROM-embedding technique is applicable to any resistive memory technology in which the bit-cell topology is identical to that of the STT-MRAM bit-cell. An additional sense amplifier is required in the peripheral circuitry, hence we propose an area-optimized peripheral circuitry to minimize the total area penalty of embedding ROM. Our analysis reveals that the ROM may be embedded in the STT-MRAM array without area overhead and without any penalty in the performance of the memory as RAM. Furthermore, our simulations show that the embedded ROM may be used to accelerate applications that use lookup tables with as much as 30% improvement in instructions per cycle of a processor using ROM-embedded STT-MRAM for its L2 cache.
Xuanyao Fong, Rangharajan Venkatesan, Dongsoo Lee, Anand Raghunathan, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Approximate storage for energy efficient spintronic memories
abstract
Spintronic memories are promising candidates for future on-chip storage due to their high density, non-volatility and near-zero leakage. However, the energy consumed by read and write operations presents a major challenge to their use as energy-efficient on-chip memory. Leveraging the ability of many applications to tolerate impreciseness in their underlying computations and data, we explore approximate storage as a new approach to improving the energy-efficiency of spintronic memories. We identify and characterize mechanisms in STT-MRAM bit-cells that provide favorable energy-quality trade-offs, i.e., disproportionate energy improvements at the cost of small probabilities of read/write failures. Based on these mechanisms, we design a quality-configurable memory array in which data can be stored to varying levels of accuracy based on application requirements. We integrate the quality-configurable array as a scratchpad in the memory hierarchy of a programmable vector processor and expose it to software by introducing quality-aware load/store instructions within the ISA. We evaluate the energy benefits of our proposal using a device-to-architecture modeling framework and demonstrate 40% and 19.5% improvement in memory energy and overall application energy respectively, for negligible (< 0.5%) quality loss across a suite of recognition and vision applications.
Ashish Ranjan 0001, Swagath Venkataramani, Xuanyao Fong, Kaushik Roy 0001, Anand Raghunathan
DAC3
2015 Device/circuit/architecture co-design of reliable STT-MRAM
Zoha Pajouhi, Xuanyao Fong, Kaushik Roy 0001
DATE2
2015 Spintastic: <u>spin</u>-based s<u>t</u>och<u>astic</u> logic for energy-efficient computing
Rangharajan Venkatesan, Swagath Venkataramani, Xuanyao Fong, Kaushik Roy 0001, Anand Raghunathan
DATE3
2015 Optimizating Emerging Nonvolatile Memories for Dual-Mode Applications: Data Storage and Key Generator
abstract
Memory-based physical unclonable functions (PUFs) have been studied and developed as powerful primitives to generate device-specific random keys, which can be used for various security applications. However, the existing memory-based PUFs need to safely buffer the data bits in the memory before it is used to produce random bits, resulting in additional area/energy consumption and potential data security issues. In this paper, we propose a new memory-based PUF that exploits the nonvolatility and random variability of emerging memory technologies to produce random bits. Unlike conventional implementations, the random bit generation process of our proposed PUF does not disturb the data bits already stored in the memory. To satisfy the quality requirements for both memory and PUF applications, we also propose a general method to find the optimal design point of emerging nonvolatile memory (eNVM)-based PUF. An illustrative design using spin-transfer torque magnetic RAM exhibits desirable results using our method. Compared to the conventional types of memory-based PUFs, eNVM-based PUFs features enhanced security as cryptographic primitives and lower area and energy cost as data storage.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 Highly Reliable Spin-Transfer Torque Magnetic RAM-Based Physical Unclonable Function With Multi-Response-Bits Per Cell
abstract
Memory-based physical unclonable function (MemPUF) has gained tremendous popularity in the recent years to securely preserve secret information in computing systems. Most MemPUFs in the literature have unreliable bit generation and/or are incapable of generating more than one response-bit per cell. Hence, we propose a novel MemPUF exploiting the unique characteristics of spin-transfer torque magnetic RAM (STT-MRAM) that can overcome these issues. Bit generation in our STT-MRAM-based MemPUF is stabilized using a novel automatic write-back technique. In addition, the alterability of the magnetic tunneling junction state is exploited to expand the response-bit capacity per cell. Our analysis demonstrated the advantage of our scheme in reliability enhancement (bit-error rate from ~10-1to ~10-6in the worst case under varying conditions) and response-bit capacity per cell improvement (from 1 to 1.48 bit). In comparison with the conventional MemPUFs, our approach is also better in terms of the average chip area and energy for producing a response-bit.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
IEEE Trans. Inf. Forensics Secur.2
2014 Highly reliable memory-based Physical Unclonable Function using Spin-Transfer Torque MRAM
abstract
In recent years, Physical Unclonable Function (PUF) based on the inimitable and unpredictable disorder of physical devices has emerged to address security issues related to cryptographic key generation. In this paper, a novel memory-based PUF based on Spin-Transfer Torque (STT) Magnetic RAM, named as STT-PUF, is proposed as a key generation primitive for embedded computing systems. By comparing the resistances of STT-MRAM memory cells which are initialized to the same state, response bits can be generated by exploiting the inherent random mismatches between them. To enhance the robustness of response bits regeneration, an Automatic Write-Back (AWB) technique is proposed without compromising the resilience of STT-PUF against possible attacks. Simulations show that the proposed STT-PUF is able to produce raw response bits with uniqueness of 50.1% and entropy of 0.985 bit per cell. The worst-case Bit-Error Rate (BER) under varying operating conditions is 6.6 × 10-6.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
ISCAS2
2014 Failure Mitigation Techniques for 1T-1MTJ Spin-Transfer Torque MRAM Bit-cells
abstract
The emergence of spin-transfer torque magnetic RAM (STT-MRAM) as a leading candidate for future high-performance nonvolatile memory has led to increased research interest. Current STT-MRAM technology faces several major obstacles in attaining its potential. One of the major issues is in the design of 1T-1MTJ STT-MRAM bit-cells under process variations: the bit-cells need to be significantly upsized to improve bit-cell failure, resulting in increased bit-cell area and power dissipation. In this paper, we analyze four circuit-level solutions that enable smaller 1T-1MTJ STT-MRAM bit-cells with improved yield, namely, bit-line voltage boosting, word-line voltage boosting, access transistor body biasing, and an applied external magnetic field. Results from simulation using 45-nm bulk CMOS access transistor and 40-nm magnetic tunneling junction technology show that word-line voltage boosting can be the best failure mitigation technique. Bit-cells designed with word-line boosting for write has a bit-cell area reduced by > 75% at iso-failure probability, compared to bit-cells without any failure mitigation technique. When bit-cell failure probability is optimized instead, 5 Oe of applied external magnetic field assisted write reduces power consumption by 15% , compared to bit-cells designed without failure mitigation techniques.
Xuanyao Fong, Yusung Kim 0002, Sri Harsha Choday, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Dual pillar spin-transfer torque MRAMs for low power applications
abstract
Electron-spin based data storage for on-chip memories has the potential for ultra-high density, low power consumption, very high endurance, and reasonably low read/write latency. In this article, we discuss the design challenges associated with spin-transfer torque (STT) MRAM in its state-of-the-art configuration. We propose an alternative bit cell configuration and three new genres of magnetic tunnel junction (MTJ) structures to improve STT-MRAM bit cell stabilities, write endurance, and reduce write energy consumption. The proposed multi-port, multi-pillar MTJ structures offer the unique possibility of electrical and spatial isolation of memory read and write. In order to realize ultralow power under process variations, we propose device, bit-cell and architecture level design techniques. Such design alternatives at multiple levels of design abstraction has been found to achieve substantially enhanced robustness, density, reliability and low power as compared to their charge-based counterparts for future embedded applications.
Niladri Narayan Mojumder, Xuanyao Fong, Charles Augustine, Sumeet Kumar Gupta, Sri Harsha Choday, Kaushik Roy 0001
ACM J. Emerg. Technol. Comput. Syst.2
2009 A design methodology and device/circuit/architecture compatible simulation framework for low-power magnetic quantum cellular automata systems
abstract
CMOS device scaling is facing a daunting challenge with increased parameter variations and exponentially higher leakage current every new technology generation. Thus, researchers have started looking at alternative technologies. Magnetic Quantum Cellular Automata (MQCA) is such an alternative with switching energy close to thermal limits and scalability down to 5nm. In this paper, we present a circuit/architecture design methodology using MQCA. Novel clocking techniques and strategies are developed to improve computation robustness of MQCA systems. We also developed an integrated device/circuit/system compatible simulation framework to evaluate the functionality and the architecture of an MQCA based system and conducted a feasibility/comparison study to determine the effectiveness of MQCAs in digital electronics. Simulation results of an 8-bit MQCA-based Discrete Cosine Transform (DCT) with novel clocking and architecture show up to 290X and 46X improvement (at iso-delay and optimistic assumption) over 45nm CMOS in energy consumption and area, respectively.
Charles Augustine, Behtash Behin-Aein, Xuanyao Fong, Kaushik Roy 0001
ASP-DAC3
2006 Analysis of super cut-off transistors for ultralow power digital logic circuits
abstract
Super cut-off devices with sub-60mV/decade subthreshold swings have recently been demonstrated and being extensively studied. This paper presents a feasibility analysis of such tunneling devices for ultralow power subthreshold logic. Analysis shows that this device can deliver 800X higher performance (@iso-IOFF) compared to a MOSFET. The possible use of this device as a sleep transistor in conjunction with the regular Si MOSFET shows 2000X average improvement in leakage power compared to Si MOSFETs.
Arijit Raychowdhury, Xuanyao Fong, Qikai Chen, Kaushik Roy 0001
ISLPED2