Subramanian S. Iyer

dblp:87/3451 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0003-1220-031XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Hybrid Obfuscation of Chiplet-Based Systems
abstract
The growing concern about offshore chip manufacturing has created considerable interest in solutions that can ensure the integrity and security of chips. Among various solutions, split manufacturing has received a lot of attention due to its security guarantees. With the recent emergence of new heterogeneous manufacturing technologies, including chiplet-based systems, there is a new opportunity for revisiting the design considerations for split manufacturing to fully exploit the opportunities presented by chiplet-based systems and improve various metrics, such as security, performance, and overhead.This work improves the state-of-the-art in secure chip manufacturing by proposing a new split manufacturing scheme. The key idea is to exploit the capabilities provided by chiplet integration technology for designing a new hybrid split manufacturing scheme that includes both vertical and horizontal splitting. Unlike existing vertical-only split manufacturing mechanisms, that target obfuscation of interconnections by splitting the design at a specific metallization layer into two portions, the proposed hybrid method increases trust by exploiting the chiplet paradigm shift, specifically, breaking the design into sub-designs, each represented by chiplets (independently fabricated), and obfuscating interconnections among them. The proposed obfuscation mechanism targets systems that exploit the chiplet technology to obtain important performance advantages, thus any chiplet-related overhead is not due to obfuscation. We evaluate our method using several experiments and compare it with the state-of-the-art using standard metrics, including area, power, delay, wirelength, and trust. Compared to conventional split manufacturing, our hybrid method achieves up to 245× higher trust, while exhibiting negligible overhead.
Yousef Safari, Pooya Aghanoury, Subramanian S. Iyer, Nader Sehatbakhsh, Boris Vaisband
DAC3
2022 Accuracy and Resiliency of Analog Compute-in-Memory Inference Engines
abstract
Recently, analog compute-in-memory (CIM) architectures based on emerging analog non-volatile memory (NVM) technologies have been explored for deep neural networks (DNNs) to improve scalability, speed, and energy efficiency. Such architectures, however, leverage charge conservation, an operation with infinite resolution, and thus are susceptible to errors. Thus, the inherent stochasticity in any analog NVM used to execute DNNs, will compromise performance. Several reports have demonstrated the use of analog NVM for CIM in a limited scale. It is unclear whether the uncertainties in computations will prohibit large-scale DNNs. To explore this critical issue of scalability, this article first presents a simulation framework to evaluate the feasibility of large-scale DNNs based on CIM architecture and analog NVM. Simulation results show that DNNs trained for high-precision digital computing engines are not resilient against the uncertainty of the analog NVM devices. To avoid such catastrophic failures, this article introduces the analog bi-scale representation for the DNN, and the Hessian-aware Stochastic Gradient Descent training algorithm to enhance the inference accuracy of trained DNNs. As a result of such enhancements, DNNs such as Wide ResNets for CIFAR-100 image recognition problem are demonstrated to have significant performance improvements in accuracy without adding cost to the inference hardware .
Zhe Wan, Subramanian S. Iyer, Vwani P. Roychowdhury
ACM J. Emerg. Technol. Comput. Syst.4
2021 Designing a 2048-Chiplet, 14336-Core Waferscale Processor
abstract
Waferscale processor systems can provide the large number of cores, and memory bandwidth required by today’s highly parallel workloads. One approach to building waferscale systems is to use a chiplet-based architecture where pre-tested chiplets are integrated on a passive silicon-interconnect wafer. This technology allows heterogeneous integration and can provide significant performance and cost benefits. However, designing such a system has several challenges such as power delivery, clock distribution, waferscale-network design, design for testability and fault-tolerance. In this work, we discuss these challenges and the solutions we employed to design a 2048-chiplet, 14,336-core waferscale processor system.
Saptadeep Pal, Irina Alam, Nick Cebry, Haris Suhail, Shi Bu, Subramanian S. Iyer, Sudhakar Pamarti, Rakesh Kumar 0002, Puneet Gupta 0001
DAC7
2020 A Novel Hierarchical Circuit LUT Model for SOI Technology for Rapid Prototyping
abstract
In this paper, a new look-up table (LUT) method is proposed to reduce the simulation time and the run time memory requirement for large logic and mixed signal simulations. In the proposed method, for the first time, circuit with multiple devices is replaced by one LUT model, called circuit LUT. The replacement results in significant reduction of the run time memory requirement. The replacement also reduces the number of interpolation steps to be performed at every Newton-Raphson iteration during the simulation that results in significant reduction of simulation time. With the proposed method, the simulation speed is improved by two times over the conventional LUT models developed for devices. In addition, 25% reduction in the run time memory requirement is also achieved by the proposed method.
Sitansusekhar Roymohapatra, Ganesh R. Gore, Akanksha Yadav, Mahesh B. Patil, Krishnan S. Rengarajan, Subramanian S. Iyer, Maryam Shojaei Baghini
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 Architecting Waferscale Processors - A GPU Case Study
abstract
Increasing communication overheads are already threatening computer system scaling. One approach to dramatically reduce communication overheads is waferscale processing. However, waferscale processors [1], [2], [3] have been historically deemed impractical due to yield issues [1], [4] inherent to conventional integration technology. Emerging integration technologies such as Silicon-Interconnection Fabric (Si-IF) [5], [6], [7], where pre-manufactured dies are directly bonded on to a silicon wafer, may enable one to build a waferscale system without the corresponding yield issues. As such, waferscalar architectures need to be revisited. In this paper, we study if it is feasible and useful to build today's architectures at waferscale. Using a waferscale GPU as a case study, we show that while a 300 mm wafer can house about 100 GPU modules (GPM), only a much scaled down GPU architecture with about 40 GPMs can be built when physical concerns are considered. We also study the performance and energy implications of waferscale architectures. We show that waferscale GPUs can provide significant performance and energy efficiency advantages (up to 18.9x speedup and 143x EDP benefit compared against equivalent MCM-GPU based implementation on PCB) without any change in the programming model. We also develop thread scheduling and data placement policies for waferscale GPU architectures. Our policies outperform state-of-art scheduling and data placement policies by up to 2.88x (average 1.4x) and 1.62x (average 1.11x) for 24 GPM and 40 GPM cases respectively. Finally, we build the first Si-IF prototype with interconnected dies. We observe 100% of the inter-die interconnects to be successfully connected in our prototype. Coupled with the high yield reported previously for bonding of dies on Si-IF, this demonstrates the technological readiness for building a waferscale GPU architecture.
Saptadeep Pal, Daniel Ruelas-Petrisko, Matthew Tomei, Puneet Gupta 0001, Subramanian S. Iyer, Rakesh Kumar 0002
HPCA5
2019 Global and semi-global communication on Si-IF
abstract
On-chip scaling continues to pose significant technological and design challenges. Nonetheless, the key obstacle in on-chip scaling is the high fabrication cost of the state-of-the-art technology nodes. An opportunity exists however, to continue scaling at the system level. Silicon interconnect fabric (Si-IF) is a platform that aims to replace both the package and printed circuit board to enable heterogeneous integration and high inter-chip performance. Bare dies are attached directly to the Si-IF at fine vertical interconnect pitch (2 to 10 μm) and small inter-die spacing (≤ 100 μm). The Si-IF is a single-hierarchy integration construct that supports dies of any process, technology, and dimensions. In addition to development of the fabrication and integration processes, system-level challenges need to be addressed to enable integration of heterogeneous systems on the Si-IF. Communication is a fundamental challenge on large Si-IF platforms (up to 300 mm diameter wafers). Different technological and design approaches for global and semi-global communication are discussed in this paper. The area overhead associated with global communication on the Si-IF is determined.
Boris Vaisband, Subramanian S. Iyer
NOCS2
2019 An Analog Neural Network Computing Engine Using CMOS-Compatible Charge-Trap-Transistor (CTT)
abstract
An analog neural network computing engine based on CMOS-compatible charge-trap transistor (CTT) is proposed in this paper. CTT devices are used as analog multipliers. Compared to digital multipliers, CTT-based analog multiplier shows significant area and power reduction. The proposed computing engine is composed of a scalable CTT multiplier array and energy efficient analog-digital interfaces. By implementing the sequential analog fabric, the engine's mixed-signal interfaces are simplified and hardware overhead remains constant regardless of the size of the array. A proof-of-concept 784 by 784 CTT computing engine is implemented using TSMC 28-nm CMOS technology and occupies 0.68 mm2. The simulated performance achieves 76.8 TOPS (8-bit) with 500 MHz clock frequency and consumes 14.8 mW. As an example, we utilize this computing engine to address a classic pattern recognition problem-classifying handwritten digits on MNIST database and obtained a performance comparable to state-of-the-art fully connected neural networks using 8-bit fixed-point resolution.
Yuan Du, Xuefeng Gu, Jieqiong Du, X. Shawn Wang, Boyu Hu, Mingzhe Jiang, Xiaoliang Chen 0001, Subramanian S. Iyer, Mau-Chung Frank Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2018 A Case for Packageless Processors
abstract
Demand for increasing performance is far outpacing the capability of traditional methods for performance scaling. Disruptive solutions are needed to advance beyond incremental improvements. Traditionally, processors reside inside packages to enable PCB-based integration. We argue that packages reduce the potential memory bandwidth of a processor by at least one order of magnitude, allowable thermal design power (TDP) by up to 70%, and area efficiency by a factor of 5 to 18. Further, silicon chips have scaled well while packages have not. We propose packageless processors - processors where packages have been removed and dies directly mounted on a silicon board using a novel integration technology, Silicon Interconnection Fabric (Si-IF). We show that Si-IF-based packageless processors outperform their packaged counterparts by up to 58% (16% average), 136%(103% average), and 295% (80% average) due to increased memory bandwidth, increased allowable TDP, and reduced area respectively. We also extend the concept of packageless processing to the entire processor and memory system, where the area footprint reduction was up to 76%.
Saptadeep Pal, Daniel Ruelas-Petrisko, Adeel Ahmad Bajwa, Puneet Gupta 0001, Subramanian S. Iyer, Rakesh Kumar 0002
HPCA5
2017 Toward Human-Scale Brain Computing Using 3D Wafer Scale Integration
abstract
The Von Neumann architecture, defined by strict and hierarchical separation of memory and processor, has been a hallmark of conventional computer design since the 1940s. It is becoming increasingly unsuitable for cognitive applications, which require massive parallel processing of highly interdependent data. Inspired by the brain, we propose a significantly different architecture characterized by a large number of highly interconnected simple processors intertwined with very large amounts of low-latency memory. We contend that this memory-centric architecture can be realized using 3D wafer scale integration for which the technology is nearing readiness, combined with current CMOS device technologies. The natural fault tolerance and lower power requirements of neuromorphic processing make 3D wafer stacking particularly attractive. In order to assess the performance of this architecture, we propose a specific embodiment of a neuronal system using 3D wafer scale integration; formulate a simple model of brain connectivity including short- and long-range connections; and estimate the memory, bandwidth, latency, and power requirements of the system using the connectivity model. We find that 3D wafer scale integration, combined with technologies nearing readiness, offers the potential for scaleup to a primate-scale brain, while further scaleup to a human-scale brain would require significant additional innovations.
Zhe Wan, Winfried W. Wilcke, Subramanian S. Iyer
ACM J. Emerg. Technol. Comput. Syst.4
2017 Assessing Benefits of a Buried Interconnect Layer in Digital Designs
abstract
In sub-15 nm technology nodes, local metal layers have witnessed extremely high congestion leading to pin-access-limited designs, and hence affecting the chip area and related performance. In this paper, we assess the benefits of adding a buried interconnect layer below the device layers for the purpose of reducing cell area, improving pin access, and reducing chip area. After adding the buried layer to a projected 7 nm standard cell library, results show ~9%-13% chip area reduction and 126% pin access improvement. This shows that buried interconnect, as an integration primitive, is very promising as an alternative method to density scaling.
Liheng Zhu, Yasmine Badr, Shaodi Wang, Subramanian S. Iyer, Puneet Gupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2011 3D integration review
Mukta G. Farooq, Subramanian S. Iyer
Sci. China Inf. Sci.2
2008 Analysis of Retention Time Distribution of Embedded DRAM - A New Method to Characterize Across-Chip Threshold Voltage Variation
abstract
In this paper, we investigate the retention time distribution of IBM's 65nm node embedded DRAM. We demonstrate that subthreshold current is the dominant leakage mechanism that determines data retention time, and the retention distribution can be attributed to array Vtvariation. Based on this study, we present a new technique for characterization of across-chip Vtvariation. The Vtmedian value and standard deviation of transfer devices within an eDRAM array are estimated by analyzing the retention characteristics. The evaluation results are confirmed by the parametric test data. The proposed method is fast and can be used to monitor Vtvariation in both technology development and manufacture. The impact of array Vtspread on the retention and performance of eDRAM is discussed.
Paul C. Parries, Subramanian S. Iyer
ITC4