Ahmedullah Aziz

dblp:64/11473 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0003-1573-4122ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 Ferroelectric-Superconducting Synergy for Future Computing
abstract
Ferroelectric Superconducting Quantum Interference Devices (Fe-SQUIDs) have recently gained attention as a transformative technology for superconducting computing, offering voltage-controlled switching that is essential for large-scale digital circuits. This unique technology has the potential to drive advancements in cryogenic computing by enabling scalable memory systems and voltage-controlled logic circuits. These innovations are critical for the realization of large-scale quantum computers and hold significant promise for high-performance computing and space exploration. In this article, we explore how Fe-SQUIDs, integrated with heater cryotrons (hTrons), can be harnessed to develop key components of computing systems. These include non-volatile memory, voltage-controlled logic circuits, in-memory matrix-vector multiplication systems, and ternary content-addressable memory. We also examine how changes in the key characteristics of Fe-SQUIDs and hTrons influence the performance of these applications, providing insights into the design and optimization of next-generation superconducting hardware.
Shamiul Alam, Ahmedullah Aziz
DATE2
2025 Harnessing Unipolar Threshold Switches for Enhanced Rectification
abstract
Phase transition materials (PTMs) have drawn significant attention in recent years due to their abrupt threshold switching characteristics and hysteretic behavior. Augmentation of the PTM with a transistor has been shown to provide enhanced selectivity (as high as ~107 for Ag/HfO2/Pt) leading to unique circuit-level advantages. Previously, a unipolar PTM, Ag-HfO2-Pt, was reported as a replacement for diodes due to its polarity-dependent high selectivity and hysteretic properties. It was shown to achieve ~50% higher-DC output compared to a diode-based design in a Cockcroft-Walton multiplier circuit. In this article, we take a deeper dive into this design. We augment two different PTMs (unipolar Ag-HfO2-Pt and bipolar VO2) with diode-connected MOSFETs to retain the benefits of hysteretic rectification. Our proposed hysteretic diodes (Hyperdiodes) exhibit a low-forward voltage drop owing to their volatile hysteretic characteristics. However, augmenting a hysteretic PTM with a transistor brings an additional stability concern due to their complex interplay. Hence, we perform a comprehensive stability analysis for a range of threshold voltages (−0.2 V$V_{\mathrm { th}}$$3 {\sigma }$Monte-Carlo variation analysis for a Cockcroft-Walton multiplier considering the nonidealities in the host transistor and the PTM. We observe that, hyperdiode-based design achieves ~20% higher-output voltage compared with the conventional designs within a fixed timeframe ($200~\boldsymbol {\mu }$s).
Md. Mazharul Islam 0006, Shamiul Alam, Garrett S. Rose, Aly E. Fathy, Sumeet Kumar Gupta, Ahmedullah Aziz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Ultra-Area-Efficient Cryogenic XNOR Logic Gate with Superconducting Heater Cryotron to Advance High-Performance Computing
abstract
Superconducting electronics have garnered significant attention recently due to their exceptional speed and energy efficiency. They play a vital role in scaling quantum computers to thousands of qubits and offer unique advantages for high-performance computing and aerospace exploration. However, conventional Josephson junction-based superconducting circuits face scalability challenges due to flux trapping and struggle in driving high load impedances. To address these issues, three-terminal cryotron devices have emerged as a promising alternative. Here, we present a novel approach utilizing two heater cryotron devices to construct an XNOR gate, significantly enhancing the area efficiency of existing superconducting XNOR gates. XNOR logic plays a pivotal role in various applications including adder design, data center operations like copy verification and data encryption/decryption. Our designed XNOR logic gate boasts a substantial improvement in area efficiency compared to existing designs, requiring only two heater cryotron devices instead of the previous requirement of ten cryotron devices, leading to an 80% improvement in device count. This advancement holds promise for further optimizing superconducting circuitry for various applications.
Shamiul Alam, Ahmedullah Aziz
ACM Great Lakes Symposium on VLSI2
2024 P-ReTI: Silicon Photonic Accelerator for Greener and Real-Time AI
abstract
Computing deep AI algorithms on traditional CPUs and GPUs brings several performance and energy pitfalls. Most of the emerging AI accelerators target only the inference phase of deep learning. There have been very limited attempts to design a full-fledged AI accelerator capable of both training and inference in real-time. It is due to the highly compute and memory intensive nature of the training phase. In this paper, we propose P-ReTI, a novel analog photonics AI accelerator. P-ReTI uses silicon microdisk-based convolution, photonic phase change memory-based memory, and dense-wavelength-division-multiplexing for energy-efficient and ultrafast deep learning in real-time. We evaluate P-ReTI using a commercial CAD framework (IPKISS) on deep learning benchmark models including LeNet and VGG-Net. Compared to the state-of-the-art, P-ReTI improves the CNN throughput, energy-efficiency, and computational efficiency by up to two orders of magnitude with trivial accuracy degradation.
Dharanidhar Dang, Priyabrata Dash, Ahmedullah Aziz
ACM Great Lakes Symposium on VLSI3
2024 Design Space Exploration for Phase Transition Material-Augmented MRAMs With Separate Read-Write Paths
abstract
This report presents a design space analysis for the phase transition material (PTM)-augmented magnetic random-access memories (MRAMs) with separate read–write paths. PTM is augmented in parallel with the magnetic tunnel junction (MTJ), improving the read performance along with providing separate read–write paths. Compared to the standard MRAM, PTM-augmented design achieves up to$1.7 \times $boost in cell tunnel magnetoresistance (CTMR),$1.2 \times $increase in read disturb margin (RDM), and${\sim }3.75 \times $increase in sense margin (SM) at the cost of${\sim }4.75 \times $more power consumption. Here, we first discuss the operating region and biasing requirements to achieve performance improvement. Then, we thoroughly explore the design space to put more options on the table for choosing the material and device structure. Finally, we perform the variation analysis where we address the performance and variation immunity tradeoffs. We demonstrate a 1000-point Monte-Carlo analysis to illustrate the effects of process variations on the performance. With lower distinguishability and read stability, the variation tolerance of the design can be improved manifold employing device-circuit co-design methodology and vice versa.
Shamiul Alam, William Mitchell Hunter, Nazmul Amin, Md. Mazharul Islam 0006, Sumeet Kumar Gupta, Ahmedullah Aziz
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 Cryogenic In-Memory Matrix-Vector Multiplication using Ferroelectric Superconducting Quantum Interference Device (FE-SQUID)
abstract
Next-generation quantum computing (QC) systems, comprising thousands of qubits, are envisioned to accommodate the quantum substrate (qubits) and classical components (control processor, and a digital memory block) in a cryogenic (< 4 Kelvin) environment. Such homogeneous integration will pave the way for superconducting interconnects and reduce the noise arising from thermal gradient. However, in the existing QC systems, cryogenic control processors and memory blocks are still operated following the von Neumann architecture. This leads to significant performance overhead due to the repetitive data movement between physically distinct memory and processing units. Thus, it becomes challenging to implement computationally expensive machine learning (ML) algorithms for efficient error correction and control of qubits in a QC. In-memory implementation of ML algorithms at cryogenic temperature can be a game-changer for a practical QC. Here, we demonstrate a unique technique for cryogenic in-memory matrix vector multiplication (MVM), the most frequently performed operation in ML algorithms, utilizing a ferroelectric superconducting quantum interference device (FE-SQUID)-based memory array. FE-SQUID is a promising cryogenic memory device thanks to its non-volatile nature, voltage-controlled switching, scalability, and compatibility with commercially available superconducting device fabrication processes. Moreover, due to having separate read-write paths, the read operation can be optimized without imposing any limit on the read bias and hence, multiple levels of read current with notable separation can be used to map the inputs for the MVM operation. We use an experimentally-calibrated compact model for FE-SQUID to design and test our proposed system. We evaluate FE-SQUID-based in-memory MVM by performing several classification tasks using MNIST handwritten digits, fashion, and emotion datasets. We achieve 93.83%, 80.49%, and 92.5% accuracy for handwritten digits, fashion, and sentiment classifications, respectively.
Shamiul Alam, Jack Hutchins, Md. Shafayat Hossain, Kai Ni 0004, Narayanan Vijaykrishnan, Ahmedullah Aziz
DAC6
2023 Ternary In-Memory Computing with Cryogenic Quantum Anomalous Hall Effect Memories
abstract
With surging interest in quantum computing, space applications, and ultra-fast superconducting processors, the need for compatible cryogenic memory systems is skyrocketing. Among several concurrent candidates for cryogenic data storage solutions, quantum anomalous Hall effect (QAHE) devices have garnered immense interest due to having topologically protected variation-tolerant quantum states. The QAHE cells, in addition to being a promising non-volatile storage technology, have several unique properties that make them ideal for in-memory computing operations. In this work, we propose a novel in-memory computing mechanism by harnessing the intrinsic voltage addition property of a QAHE memory array, implemented using twisted bi-layer graphene (tBLG) on hexagonal boron nitride (hBN). In addition, we extensively explore and implement ternary arithmetic operations utilizing the series-connected Hall voltages across devices for the first time. We propose two schemes for in-memory ternary computing namely IMFE and IMSE, and demonstrate balanced scalar multiplication, dot product operations, and ternary half adder with QAHE memory array.
Arun Govindankutty, Shamiul Alam, Sanjay Das, Nagadastagiri Challapalle, Ahmedullah Aziz, Sumitha George
ACM Great Lakes Symposium on VLSI5
2023 A Cryogenic Artificial Synapse based on Superconducting Memristor
abstract
Spiking neural network (SNN) has emerged as the most biologically accurate approach for information encoding in neuromorphic computing. Cryogenic neuromorphic hardware, which offers exceptional energy efficiency and speed, has recently gained enormous attention among the neuromorphic community. An approach to build such neuromorphic hardware is to use a conductance asymmetric superconducting quantum interference device (CA-SQUID) that has non-volatile and variation- robust dual-resistive behavior and thereby, is referred to as a superconducting memristor (SM). Here, we utilize this unique device to design an SM-based artificial synapse topology for neuromorphic applications. The proposed synapse structure, combined with an SM-based neuron, demonstrates neurosynaptic behavior with enhanced reconfigurability. Our design features eight different non-volatile levels of synaptic strength, utilizing combinations of distinct resistance levels of three SMs, exhibiting an estimated programming power of 8.5 pW. This weight storage feature enables better reconfigurability compared to the existing superconducting synapse structures that utilized fixed resistors and inductors. Additionally, this synapse can be further fine-tuned to dynamically access a wide range of synaptic strengths by using an external bias current. Our study provides valuable insights into the system-level integration of the neuron-synaptic architecture.
Md. Mazharul Islam 0006, Shamiul Alam, Md Rahatul Islam Udoy, Md. Shafayat Hossain, Ahmedullah Aziz
ACM Great Lakes Symposium on VLSI5
2023 Reliable Brain-inspired AI Accelerators using Classical and Emerging Memories
abstract
By taking inspiration from the operation of biological brains, emerging brain-inspired hardware has the potential to revolutionize the way computations are performed. Brain-inspired computing can be realized using both classical CMOS and emerging beyond-CMOS technologies, whereas the latter holds the promise to provide substantial energy savings akin to the employment of non-volatile memories. One way to implement highly efficient brain-inspired AI applications is through analog computing schemes, such as Integrate-and-Fire (IF) Spiking Neural Networks (SNNs), which can be implemented using both CMOS and beyond-CMOS technologies as synaptic storage. However, managing the inherent degradation of computing accuracy in analog circuits and mitigating their effects on the predictive accuracy of AI systems remains a key challenge due to the inherent nature of analog computing.In this paper, we discuss how the aforementioned challenges can be addressed. In the first part, we present our SPICE-Torch, a framework that connects low-level SPICE simulations of circuits and memories performing analog computations with high-level accuracy evaluations of NN models based on PyTorch. Furthermore, we present an example of neuromorphic optimization using classical CMOS technology. In the second part, we introduce memristors as an emerging beyond-CMOS technology that can retain their state without any outside influence and are well-suited for brain-inspired neuromorphic hardware. We demonstrate that brain-inspired hardware, realized using classical CMOS or beyond-CMOS technologies, has the potential to revolutionize the way we process information and solve complex computation problems. Nevertheless, to harness its full potential, reliability issues have to be managed carefully and HW/SW codesign is key. Our presented framework SPICE-Torch, which connects low-level SPICE simulations of circuits performing analog computations with high-level accuracy evaluations of NN models based on PyTorch is available as open-source in https://github.com/myay/SPICE-Torch.
Mikail Yayla, Simon Thomann, Md. Mazharul Islam 0006, Ming-Liang Wei, Shu-Yin Ho, Ahmedullah Aziz, Chia-Lin Yang, Jian-Jia Chen, Hussam Amrouch
VTS6
2021 Monte Carlo Variation Analysis of NCFET-based 6-T SRAM: Design Opportunities and Trade-offs
abstract
Negative Capacitance FET (NCFET) is one of the most promising variants of the emerging steep-slope transistors, able to overcome the ?Boltzmann limit'. The ferroelectric layer in the gate stack brings in new dynamics to the transistor operation by amplifying the surface potential. Steeper subthreshold slope, higher ON/OFF ratio, and the possibility to attain negative output conductance provide unique opportunities for NCFET-based circuit design. However, NCFETs inherently possess additional sources of variation, and hence, the promise of performance benefits in the nominal designs must be examined through extensive variation analysis. The non-volatile ferroelectric FETs (FEFETs) are promising candidates for storage-class memory, whereas the volatile NCFETs are suitable for high-speed SRAM design. In this work, we first draw a contrast between the modeling approaches ideal for the non-volatile FEFETs and volatile NCFETs. We then utilize a compact model for NCFET to analyze the design possibilities in an NCFET-based 6-T SRAM cell compared with its conventional counterpart ? both implemented in the 10 nm technology node. We examine the read, write, and hold performance of the SRAM cells through Monte Carlo variation analysis. We show that, even with additional variation induced spread in the device characteristics, NCFET-based SRAM cell can achieve better Static Noise Margin (SNM) during read/hold modes and allows more aggressive supply voltage scaling. The increased hold stability imposes a penalty in the write performance ? forcing design trade-offs.
Shamiul Alam, Nazmul Amin, Sumeet Kumar Gupta, Ahmedullah Aziz
ACM Great Lakes Symposium on VLSI4
2018 Computing with ferroelectric FETs: Devices, models, systems, and applications
abstract
In this paper, we consider devices, circuits, and systems comprised of transistors with integrated ferroelectrics. Said structures are actively being considered by various semiconductor manufacturers as they can address a large and unique design space. Transistors with integrated ferroelectrics could (i) enable a better switch (i.e., offer steeper subthreshold swings), (ii) are CMOS compatible, (iii) have multiple operating modes (i.e., I-V characteristics can also enable compact, 1-transistor, non-volatile storage elements, as well as analog synaptic behavior), and (iv) have been experimentally demonstrated (i.e., with respect to all of the aforementioned operating modes). These device-level characteristics offer unique opportunities at the circuit, architectural, and system-level, and are considered here from device, circuit/architecture, and foundry-level perspectives.
Ahmedullah Aziz, Evelyn T. Breyer, Xiaoming Chen 0003, Suman Datta, Sumeet Kumar Gupta, Michael Hoffmann 0008, Xiaobo Sharon Hu, Adrian M. Ionescu, Matthew Jerry, Thomas Mikolajick, Halid Mulaosmanovic, Kai Ni 0004, Michael T. Niemier, Ian O'Connor, Atanu Saha, Stefan Slesazeck, Sandeep Krishna Thirumala, Xunzhao Yin
DATE1
2018 Symmetric 2-D-Memory Access to Multidimensional Data
Sumitha George, Xueqing Li 0002, Minli Julie Liao, Kaisheng Ma, Srivatsa Rangachar Srinivasa, Karthik Mohan, Ahmedullah Aziz, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan
IEEE Trans. Very Large Scale Integr. Syst.7
2016 Nonvolatile memory design based on ferroelectric FETs
abstract
Ferroelectric FETs (FEFETs) offer intriguing possibilities for the design of low power nonvolatile memories by virtue of their three-terminal structure coupled with the ability of the ferroelectric (FE) material to retain its polarization in the absence of an electric field. Utilizing the distinct features of FEFETs, we propose a 2-transistor (2T) FEFET-based nonvolatile memory with separate read and write paths. With proper co-design at the device, cell and array levels, the proposed design achieves non-destructive read and lower write power at iso-write speed compared to standard FERAM. In addition, the FEFET-based memory exhibits high distinguishability with six orders of magnitude difference in the read currents corresponding to the two states. Comparative analysis based on experimentally calibrated models shows significant improvement of access energy-delay. For example, at a fixed write time of 550ps, the write voltage and energy are 58.5% and 67.7% lower than FERAM, respectively. These benefits are achieved with 2.4 times the area overhead. Further exploration of the proposed FEFET memory in energy harvesting nonvolatile processors shows an average improvement of 27% in forward progress over FERAM.
Sumitha George, Kaisheng Ma, Ahmedullah Aziz, Xueqing Li 0002, Asif Islam Khan, Sayeef S. Salahuddin, Meng-Fan Chang, Suman Datta, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan
DAC3
2016 Exploiting ferroelectric FETs for low-power non-volatile logic-in-memory circuits
abstract
Numerous research efforts are targeting new devices that could continue performance scaling trends associated with Moore's Law and/or accomplish computational tasks with less energy. One such device is the ferroelectric FET (FeFET), which offers the potential to be scaled beyond the end of the silicon roadmap as predicted by ITRS. Furthermore, the Ids vs. Vgs characteristics of FeFETs may allow a device to function as both a switch and a non-volatile storage element. We exploit this FeFET property to enable fine-grained logic-in-memory (LiM). We consider three different circuit design styles for FeFET-based LiM: complementary (differential), dynamic current mode, and dynamic logic. Our designs are compared with existing approaches for LiM (i.e., based on magnetic tunnel junctions (MTJs), CMOS, etc.) that afford the same circuit-level functionality. Assuming similar feature sizes, non-volatile FeFET-based LiM circuits are more efficient than functional equivalents based on MTJs when considering metrics such as propagation delay (2.9×, 6.8×) and dyanmic power (3.7×, 2.3×) (for 45 nm, 22 nm technology respectively). Compared to CMOS functional equivalents, FeFET designs still exhibit modest improvements in the aforementioned metrics while also offering non-volatility and reduced device count.
Xunzhao Yin, Ahmedullah Aziz, Joseph Nahas, Suman Datta, Sumeet Kumar Gupta, Michael T. Niemier, Xiaobo Sharon Hu
ICCAD2
2016 On the potential of correlated materials in the design of spin-based cross-point memories (Invited)
abstract
Cross-point architectures are promising for designing dense memory arrays. However, sneak current paths in a cross-point array necessitates the use of non-linear selectors. In this paper, we analyze the potential of employing correlated materials exhibiting abrupt insulator-metal transitions as selectors to design cross-point memories based on magnetic tunnel junctions (MTJs). We analyze the properties of the correlated materials and co-design MTJs and the selector to optimize the energy efficiency and robustness of the memory array. Our analysis points to the need of a correlated material with a large ratio of insulator and metal resistivities along with appropriate critical currents for the phase transitions (the values of which depend on the absolute value of the resistivities). We discuss that the design constraints lead to a restriction on the range of the selector length, which is closely related to the oxide thickness of the MTJ. Comparison of the cross-point architecture with standard architecture shows the benefits in the former in terms of 7% larger sense margin and 5X higher integration density at iso-read stability. However, this comes at the cost of 2X lower write speed (due to two-cycle write) and 11%-19% increase in the read/write power (due to sneak current in the cross-point array).
Sumeet Kumar Gupta, Ahmedullah Aziz, Nikhil Shukla, Suman Datta
ISCAS2
2016 Ferroelectric Transistor based Non-Volatile Flip-Flop
abstract
We present a non-volatile flip-flop with a feature to back-up the state in a ferroelectric transistor (FEFET) during power failure or supply gating. The data is stored in the form of polarization of the ferroelectric (FE) layer in the gate stack of the FEFET. The proposed flip-flop utilizes the non-volatility of the three-terminal FEFET to optimize the data backup and restore operations. We perform an extensive device-circuit analysis to provide insights into the design of the proposed flip-flop. We discuss the optimization of the FE thickness in the gate stack of the FEFET to introduce suitable non-volatility and present the implications at the circuit level. Our analysis shows that by virtue of the three terminal structure of the FEFET and the order of magnitude difference in the current for the two polarization states, the design of the backup/restore module is considerably simplified. Compared to a FE capacitor based non-volatile flip-flop, the proposed flip-flop achieves 40%--50% smaller backup delay, 27%--40% lower backup energy, comparable restore delay and up to an order of magnitude lower restore energy. While the FE capacitor based design leads to 76% area penalty compared to a conventional (volatile) flip-flop, the proposed design incurs only 35% area overhead.
Danni Wang, Sumitha George, Ahmedullah Aziz, Suman Datta, Narayanan Vijaykrishnan, Sumeet Kumar Gupta
ISLPED3
2015 COAST: Correlated material assisted STT MRAMs for optimized read operation
abstract
We present a novel technique for optimizing the read operation of spin-transfer torque (STT) MRAMs by employing a correlated material in conjunction with a magnetic tunnel junction (MTJ). The design of the proposed memory cell is based on exploiting the orders-of-magnitude difference in the resistance of the two phases of the correlated material (CM) and triggering operation-driven phase transitions in the CM by judiciously co-optimizing devices and the memory cell. During read, the CM operates in the metallic and insulating phases when the MTJ is in the low resistance and high resistance states, respectively. This leads to superior distinguishability, read efficiency and stability. During write, the CM operates in the metallic phase, which minimizes the impact of the CM resistance on the write speed. Our analysis shows that CM amplifies the cell tunneling magneto-resistance from 107% (for the standard STT MRAM) to 1878% (for the proposed cell) leading to 68% higher sense margin. In addition, 45% enhancement in the read disturb margin and 36% reduction in the cell read power is achieved. At the same time, the write asymmetry associated with different state transitions is mildly mitigated, leading to 9% reduction in the write power. This comes at a negligible cost of 4% larger write time. We also discuss the layout implications of our technique and propose the sharing of the CM amongst multiple cells. As a result of the sharing, the proposed technique incurs no area penalty.
Ahmedullah Aziz, Nikhil Shukla, Suman Datta, Sumeet Kumar Gupta
ISLPED1