Azad Naeemi

dblp:54/2495 · also Azad J. Naeemi · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-4774-9046ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 GT2N: An Open-Source 2nm Nanosheet PDK Enabling Multi-Width/VT Benchmarking
Dongwon Jang, Md. Nahid Haque Shazon, Sabareesh Jeevan Ram, Alexei Svizhenko, Victor Moroz, Ahmet Ceyhan, Nischal Arkali Radhakrishna, Azad Naeemi
ISCAS9
2025 Axon: A Novel Systolic Array Architecture for Improved Run Time and Energy Efficient GeMM and Conv Operation with On-Chip im2col
abstract
General matrix multiplication ($GeMM$) is a core operation in virtually all AI applications. Systolic array (SA) based architectures have shown great promise as$GeMM$hardware accelerators thanks to their speed and energy efficiency. Unfortunately, SAs incur a linear delay in filling the operands, due to unidirectional propogation via pipeline latches. In this work, we propose a novel in-array data orchestration technique in SAs where we enable data feeding on the principal diagonal followed by bi-directional propagation. This improves the runtime by up to 2 × at minimal hardware overhead. In addition, the proposed data orchestration enables convolution lowering (known as im2col) using a simple hardware support to fully exploit input feature map reuse opportunity and significantly lower the off-chip memory traffic resulting in 1.2 ×throughput improvement and 2.17 × inference energy reduction during YOLOv3 and RESNET50 workload on average. In contrast, conventional data orchestration would require more elaborate hardware and control signals to implement im2col in hardware because of the data skew. We have synthesized and conducted place and route for 16×16 systolic arrays based on the novel and conventional orchestrations using ASAP 7nm PDK and found that our proposed approach results in 0.211% area and 1.6% power overheads.
Md Mizanur Rahaman Nayan, Ritik Raj, Shaik Gouse Basha, Tushar Krishna, Azad Naeemi
DATE5
2025 HyDra: SOT-CAM Based Vector Symbolic Macro for Hyperdimensional Computing
abstract
Hyperdimensional computing (HDC) is a brain-inspired paradigm valued for its noise robustness, parallelism, energy efficiency, and low computational overhead. Hardware accelerators are being explored to further enhance their performance, but current solutions are often limited by application specificity and the latency of encoding and similarity search. This paper presents a generalized, reconfigurable on-chip training and inference architecture for HDC, utilizing spin-orbit-torque magnetic random access memory (SOT-MRAM) based content-addressable memory (SOT-CAM). The proposed SOT-CAM array integrates storage and computation, enabling inmemory execution of key HDC operations: binding (bitwise multiplication), permutation (bit shfiting), and efficient similarity search. Furthermore, a novel bit drop method-based permutation backed by holographic information representation of HDC is proposed which replaces conventional permutation execution in hardware resulting in a 6× latency improvement, and an HDC-specific adder reduces energy and area by 1.51× and 1.43×, respectively. To mitigate the parasitic effect of interconnects in the similarity search, a four-stage voltage scaling scheme has been proposed to ensure an accurate representation of the Hamming distance. Benchmarked at 7nm, the architecture achieves energy reductions of 21.5×, 552.74×, 1.45×, and 282.57× for addition, permutation, multiplication, and search operations, respectively, compared to CMOS-based HDC. Against state-of-the-art HDC accelerators, it achieves a 2.27× lower energy consumption and outperforms CPU and eGPU implementations by 2702× and 23161×, respectively, with less than 3% drop in accuracy.
Md Mizanur Rahaman Nayan, Che-Kai Liu, Zishen Wan, Arijit Raychowdhury, Azad Naeemi
ICCAD5
2024 WIP: Towards a Georgia Tech Semiconductor Experience for Underclassmen
abstract
This work in progress innovative practice paper describes the plans of a semiconductor experience tailored to underclassmen engineering students. The Georgia Tech Semi-conductor Experience (GTSE) is an upcoming program that is to provide a first-year experience that inspires pursuit of a career in semiconductor technology. Initially acting as a project that may be chosen and completed during a freshmen electrical and computer engineering course, GTSE will provide an opportunity for underclassmen to acquire a glimpse into the fabrication, modeling, and visualization of semiconductor-related devices. Fabrication consists of processes needed to fabricate a silicon solar cell. The fabrication process has been developed to maximize student safety, maximize student throughput each semester, and minimize the processing complexity. Processes requiring experience will be done under the guidance of a teacher assistant or instructor. After fabrication, the solar cells will be tested to determine their I/V characteristics and efficiency. The data can then be used to help students understand the complexity that a non-ideal world can introduce in a relatively simple device. The outcome of the semiconductor fabrication process is simulated using Silvaco Athena, and modeling of the solar cell device's electrical and physical properties is accomplished using Synopsys TCAD. Visualization of semiconductor device operation and atomic-level charge carrier operation is realized using custom software packages. To aid in the advancement of GTSE, under-classmen were asked to participate in guided research, modeling, or equipment setup and programming. Discussions about various aspects of silicon solar cell fabrication planning and modeling gave some qualitative insight into how well students will engage in the knowledge laid out for them. At the end of GTSE, students will have learned about various classical and quantum mechanics of the solar cell and understand the fabrication decisions. The main goal is to introduce basic skills and learning opportunities to spark interest in many semiconductor careers, especially since there is currently an industry demand for a semiconductor workforce in the United States.
William L. Schaffer, Ajeet Rohatgi, Azad Naeemi, A. Bruno Frazier
FIE3
2020 Multiplier Architectures: Challenges and Opportunities with Plasmonic-based Logic : (Special Session Paper)
abstract
Emerging technologies such as plasmonics and photonics are promising alternatives to CMOS for high throughput applications, thanks to their waveguide's low power consumption and high speed of computation. Besides these qualities, these novel technologies also implement logic functionalities uncommon to traditional technologies that can be beneficial to existing CMOS architectures. In this work, we study how plasmonic-based devices can complement CMOS technology to achieve a more efficient implementation of multiplier architectures, which are the core of state-of-the-art data- and signal-processing circuits. A critical part of modern multipliers is the partial-product reduction step, used to reduce the partial product tree into a 2-input addition. In CMOS technology, this step is achieved by using compact and fast counters. On the other hand, the proposed plasmonic cells naturally implement counters of 3-, 9- and 27-inputs within a few logic levels at ultra-high speed. Thus, we present novel multiplier architectures, which take advantage of large plasmonic-based counters to reduce the number of cells and logic levels in the partial product reduction step of the multiplication. Our experimental results show that 3 levels and 30 counters are needed when 27-input cells are used. On the other side, 6 levels and 72 counters are employed with 9-input cells. Finally, we present various 16 × 16 multiplier implementations mixing 9- and 27-input cells, focusing on the trade-off in the number of counters, levels, and area of each architecture.
Eleonora Testa, Samantha Lubaba Noor, Odysseas Zografos, Mathias Soeken, Francky Catthoor, Azad Naeemi, Giovanni De Micheli
DATE6
2019 A Mixed Signal Architecture for Convolutional Neural Networks
abstract
Deep neural network (DNN) accelerators with improved energy and delay are desirable for meeting the requirements of hardware targeted for IoT and edge computing systems. Convolutional neural networks (CoNNs) belong to one of the most popular types of DNN architectures. This article presents the design and evaluation of an accelerator for CoNNs. The system-level architecture is based on mixed-signal, cellular neural networks (CeNNs). Specifically, we present (i) the implementation of different layers, including convolution, ReLU, and pooling, in a CoNN using CeNN, (ii) modified CoNN structures with CeNN-friendly layers to reduce computational overheads typically associated with a CoNN, (iii) a mixed-signal CeNN architecture that performs CoNN computations in the analog and mixed signal domain, and (iv) design space exploration that identifies what CeNN-based algorithm and architectural features fare best compared to existing algorithms and architectures when evaluated over common datasets—MNIST and CIFAR-10. Notably, the proposed approach can lead to 8.7× improvements in energy-delay product (EDP) per digit classification for the MNIST dataset at iso-accuracy when compared with the state-of-the-art DNN engine, while our approach could offer 4.3× improvements in EDP when compared to other network implementations for the CIFAR-10 dataset.
Qiuwen Lou, Chenyun Pan, John McGuinness, András Horváth, Azad Naeemi, Michael T. Niemier, Xiaobo Sharon Hu
ACM J. Emerg. Technol. Comput. Syst.5
2018 Accurate processor-level wirelength distribution model for technology pathfinding using a modernized interpretation of rent's rule
abstract
Faithful system-level modeling is vital to design and technology pathfinding, and requires accurate representation of interconnects. In this study, Rent's rule is modernized to cater to advanced technology and design, and applied to derive a priori wirelength distribution models. Furthermore, a priori interconnect branching models are proposed to capture design constraints and their handling by the Electronic-Design-Automation tools. These interconnect branching models are embedded into the wirelength distribution models and validated against a suite of state-of-the-art commercial designs across technology nodes. Novel design-specific critical-path models are presented which capture trends in technology and microarchitecture, providing a reliable framework for future technology and design benchmarking.
Divya Prasad, Saurabh Sinha 0001, Brian Cline, Azad Naeemi
DAC5
2017 A Pathway to Enable Exponential Scaling for the Beyond-CMOS Era: Invited
abstract
Many key technologies of our society, including so-called artificial intelligence (AI) and big data, have been enabled by the invention of transistor and its ever-decreasing size and ever-increasing integration at a large scale. However, conventional technologies are confronted with a clear scaling limit. Many recently proposed advanced transistor concepts are also facing an uphill battle in the lab because of necessary performance tradeoffs and limited scaling potential. We argue for a new pathway that could enable exponential scaling for multiple generations. This pathway involves layering multiple technologies that enable new functions beyond those available from conventional and newly proposed transistors. The key principles for this new pathway have been demonstrated through an interdisciplinary team effort at C-SPIN (a STARnet center), where systems designers, device builders, materials scientists and physicists have all worked under one umbrella to overcome key technology barriers. This paper reviews several successful outcomes from this effort on topics such as the spin memory, logic-in-memory, cognitive computing, stochastic and probabilistic computing and reconfigurable information processing.
Jianping Wang 0006, Sachin S. Sapatnekar, Chris H. Kim, Paul A. Crowell, Steven J. Koester, Supriyo Datta, Kaushik Roy 0001, Anand Raghunathan, Xiaobo Sharon Hu, Michael T. Niemier, Azad Naeemi, Chia-Ling Chien, Caroline A. Ross, Roland Kawakami
DAC11
2017 Beyond-CMOS non-Boolean logic benchmarking: Insights and future directions
abstract
Emerging technologies are facing significant challenges to compete with CMOS with respect to Boolean logic. There is an increasing need for using non-traditional circuits to realize the full potential of beyond-CMOS devices. This paper presents a uniform benchmarking methodology for non-Boolean computation based on the cellular neural network (CNN) for a variety of beyond-CMOS device technologies, including charge-based and spintronic devices. Three types of CNN implementations are investigated benchmarked for a given input noise and recall accuracy target using analog, digital, and spintronic circuits. Results demonstrate that spintronic devices are promising candidates to implement CNNs, where up to 3x EDP improvement is predicted in domain wall devices compared to its conventional CMOS counterpart. This shows that alternative non-Boolean computing platforms are crucial for developing future emerging technologies.
Chenyun Pan, Azad Naeemi
DATE2
2014 BEOL Scaling Limits and Next Generation Technology Prospects
abstract
This paper presents the major limitations to the interconnect technology scaling at future technology generations and demonstrates both evolutionary and radical potential solutions to the BEOL scaling problem. To address the local interconnect challenges, a novel hybrid Al-Cu interconnect technology is introduced. Performances of carbon-based interconnects are evaluated as a more radical solution. The impact of interconnects and the optimal interconnect options are investigated for emerging next generation devices. Interconnects for new state variables, namely spintronic interconnects, are studied and their potential performances in an all-spin logic system are evaluated.
Azad Naeemi, Ahmet Ceyhan, Vachan Kumar, Chenyun Pan, Rouhollah Mousavi Iraei, Shaloo Rakheja
DAC1
2014 Interactive visualizations for teaching quantum mechanics and semiconductor physics
abstract
Work in Progress: The theory of Quantum Mechanics (QM) provides a foundation for many fields of science and engineering; however, its abstract nature and technical difficulty make QM a challenging subject for students to approach and grasp. This is partly because complex mathematical concepts involved in QM are difficult to visualize for students and the existing visualization are minimal and limited. We propose that many of these concepts can be communicated and experienced through interactive visualizations and games, drawing on the strengths and affordances of digital media. A game environment can make QM concepts more accessible and understandable by immersing students in nano-sized worlds governed by unique QM rules. Furthermore, replayability of games allows students to experience the probabilistic nature of QM concepts. In this paper, we present a game and a series of interactive visualizations that we are developing to provide students with an experiential environment to learn quantum mechanics. We will discuss how these visualizations and games can enable students to experiment with QM concepts, compare QM with classical physics, and get accustomed to the often counterintuitive laws of QM.
Rose Peng, Bill Dorn, Azad Naeemi, Nassim Parvin
FIE3
2014 Performance modeling for emerging interconnect technologies in CMOS and beyond-CMOS circuits
abstract
In this paper, emerging low-power interconnect options for CMOS and beyond CMOS technologies are reviewed. First, electrical interconnects based on carbon nanotubes and graphene nanoribbons are discussed. It is found that carbon-based electrical interconnects can potentially outperform their conventional Cu counterpart at technology nodes close to or below 10 nm. Next, since using electron spin as a novel state variable has attracted major attention, interconnect options for beyond-COMS spintronic devices will be discussed. We start with metallic interconnects based on the non-local spin-valve and spin-torque-driven switching, and the impact of size effects and dimensional scaling on their potential performance is studied. It is found that the spin signal in the non-local structure decays significantly because of a large degradation in the spin relaxation length as the interconnect width decreases. Next, a spintronic interconnect in the form of a conventional spin-valve configuration is introduced to increase the energy efficiency by eliminating the loss of spins in the non-local structure. Both metallic and semiconducting channels are studied, and the results show that the metallic interconnect is more energy-efficient than the semiconducting one when the interconnect is short (a few hundreds of nanometers) due to a high conductive current path. However, a semiconducting channel is appropriate for an intermediate or long (several microns) interconnect due to a longer spin relaxation time and the possibility of using an electric field to enhance the spin relaxation length. Furthermore, it is shown that for spin interconnects, downscaling the size of the ferromagnets can largely reduce the delay, energy, and energy-delay product at the cost of a shorter retention time.
Sou-Chi Chang, Ahmet Ceyhan, Vachan Kumar, Azad Naeemi
ISLPED4
2013 Evaluation of the Potential Performance of Graphene Nanoribbons as On-Chip Interconnects
abstract
Interconnects are considered as one of the grandest challenges that gigascale and terascale integrations face because of the delay they add to critical paths, the power they dissipate, the noise and jitter they induce on one another, and their vulnerability to electromigration. Recent studies on novel computational state variables such as electron spin have demonstrated that interconnects will continue to be an ever-growing challenge, even for post-complementary metal-oxide-semiconductor (CMOS) switches. The novel 2-D carbon-based material graphene has demonstrated remarkable electrical properties that make it a viable candidate to implement interconnects in both electrical and spintronic domains. In this paper, physical models of the electron transport parameters such as electron mean free path (MFP), diffusion coefficient, mobility, and resistance per unit length are presented for both bulk (2-D) and narrow (1-D) graphene nanoribbons (GNRs) as a function of the interconnect dimensions, edge roughness, and Fermi-energy shift. The potential of multilayer GNR (ML-GNR) as electrical interconnects is explored by taking into account the finite interlayer resistivity between the multiple layers within the ML-GNR stack. The spin-relaxation length in graphene is obtained using some theoretical estimates on the spin-orbit coupling (SOC) introduced due to ripples in graphene. It is found that, in pure graphene, the spin-relaxation length could be longer than 10 μm; however, the presence of adatoms limits the spin-relaxation length in graphene to only 1-2 μm at room temperature. The models developed in this paper are used to benchmark graphene interconnects against their conventional copper/low- κ interconnects in both electrical and spintronic domains. The results offer important insights about the advantages and limitations of graphene interconnects and provide guidelines for technology development for this emerging interconnect technology.
Shaloo Rakheja, Vachan Kumar, Azad Naeemi
Proc. IEEE3
2011 Power Aware Post-manufacture Tuning of Analog Nanocircuits
abstract
Process variations play a critical role in determining performance of scaled CMOS and other non-CMOS nanodevices. In this paper a power conscious post manufacture tuning technique is proposed for robust analog circuit fabrication with nanodevices in the presence of process variations. The response of the circuit to an optimized test signal is captured and using regression models, the proposed algorithm finds the best setting of tuning knobs for which the shifts in specifications from their nominal values are minimized in a power-aware manner. To demonstrate the proposed algorithm, a two stage Miller compensated operational amplifier is designed using carbon nanotube field effect transistors (CNFETs) and variability effects due to metallic CNT growth, diameter and chirality variations on the performance of the underlying circuits are studied. Suitable tuning knobs for the CNFET op-amp are identified based on the specifications to be tuned. Simulation results show that the proposed tuning algorithm enables overall yield improvement of 25.38% while minimizing power consumption of the tuned devices.
Aritra Banerjee, Subho Chatterjee, Azad Naeemi, Abhijit Chatterjee
ETS3
2008 Physical models for electron transport in graphene nanoribbons and their junctions
abstract
Physics-based equivalent circuit models are presented for armchair and zigzag graphene nanoribbons (GNRs), and their conductances have been benchmarked against those of carbon nanotubes and copper wires. It is demonstrated that GNRs that have large coherence length behave like waveguides for electrons. This poses both challenges and opportunities for designing graphene nanoelectronic circuits.
Azad Naeemi, James D. Meindl
ICCAD1
2007 Performance Modeling and Optimization for Single- and Multi-Wall Carbon Nanotube Interconnects
abstract
Based on physical circuit models the performances of signal and power interconnects at the local, semi-global and global levels are modeled at 100degC. For local signal interconnects, replacing copper wires with a typical aspect ratio of 2 by thin SWNT interconnects can lower power dissipation by 50%. This would also improve their speed by up to 50% by the end of the ITRS. Copper wires and large diameter MWNTs offer the lowest resistance for power distribution in the first and second interconnect levels, respectively. SWNT-bundles and MWNTs can be used to lower the delay of signal interconnects in semi-global levels. MWNTs with diameters of 50 nm and 100 nm can potentially increase the bandwidth density of global interconnects by up to 50% and 100%, respectively.
Azad Naeemi, Reza Sarvari, James D. Meindl
DAC1
2007 IntSim: A CAD tool for optimization of multilevel interconnect networks
abstract
Interconnect issues are becoming increasingly important for ULSI systems. IntSim, an interconnect CAD tool, has been developed to obtain pitches of different wiring levels and die size for circuit blocks or logic cores of microchips. It includes a methodology for co-optimization of signal, power and clock interconnects, and a newly derived stochastic wiring distribution that gives reduced error than prior work when compared to measured data. Results of IntSim are found to match well with actual data from an analyzed microprocessor. Several case studies are conducted to show this CAD tool’s utility as a system level simulator: (i) Wire resistivity increases due to size effects are projected to increase die size of a 22nm low power logic core by 30% and power by 7%. (ii) When compared to a 22nm low power logic core with copper interconnects, a similar logic core with carbon nanotube interconnects could reduce power by 25% and die area by 27%, or increase frequency by 15% and reduce die area by 11%. (iii) A future 22nm 8 GHz 96M gate logic core’s power, die size and optimal multilevel interconnect architecture are predicted. A version of IntSim with a graphical user interface is available for download from www.ece.gatech.edu/research/labs/gsigroup.
Deepak C. Sekar, Azad Naeemi, Reza Sarvari, Jeffrey A. Davis, James D. Meindl
ICCAD2
2007 Carbon nanotube interconnects
abstract
Based on physical models, circuit models are presented for SWNTs, SWNT-bundles and MWNTs. These models can be used for circuit simulations and compact modeling. It is demonstrated that by customizing CNT interconnects at the local, semiglobal and global levels several major challenges facing GSI systems can potentially be addressed. For local interconnects, mono- or few-layer SWNT interconnects can offer up to 50% reduction in capacitance and power dissipation with considerable improvements in latency if they are short enough (
Azad Naeemi, James D. Meindl
ISPD1