Boris Vaisband

dblp:149/5081 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-6176-5918ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 9 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Multi-Range Communication for Chiplet-Based Systems
abstract
In modern large-scale and data-centric applications, communication plays a crucial role in overall system throughput. With the maturation of advanced packaging technologies and fine-pitch integration, parallel communication is replacing SerDes-based links; however, computational and data demands remain high. A multi-range communication infrastructure at the architecture and network levels, is proposed in this paper, to advance communication and data routing in scaled-out chiplet-based systems. Utility chiplets (UCs) are introduced, featuring a specialized multi-range router based on the unified chiplet interface protocol (ChIP), supporting both parallel and high-speed serial communication. Evaluations are performed over the silicon interconnect fabric (Si-IF), an advanced fine-pitch integration technology, incorporating architecture-, circuit-, and network-level considerations. Compared with state-of-the-art 2D mesh-based chiplet architectures using UCIe over interposers, the proposed approach achieves up to 92.3% higher total throughput, 89.6% lower hop count, and up to 9.5 × improvement in single-node failure response under fault-tolerance evaluation. In addition, the proposed multi-range communication reduces latency by 2 × under congestion compared to state-of-the-art alternatives and achieves up to 1.35 × higher compute‑efficiency density of 871.4 PFLOPS/W/mm2 in a scaled-out configuration.
Arvin Delavari, Amirtha Chandrasekaran, Boris Vaisband
ACM Great Lakes Symposium on VLSI3
2026 Hardware-Aware Offline Training of CTT-Based Neuromorphic Hardware
abstract
Neuromorphic hardware combined with spiking neural networks (SNNs) supports low-power, highly parallel AI inference. Inference accuracy can drop sharply when models trained under ideal software assumptions are deployed on inherently noisy hardware. This noise is especially evident in the current of subthreshold operated charge-trap transistors (CTTs) representing network weights. Although CTTs are CMOS-compatible and attractive for large-scale neuromorphic systems due to their energy and cost efficiency, programming imprecision and process, voltage, and temperature (PVT) variations can significantly alter their current, degrading network accuracy. In this article, CTT weight and neuron sources of variability are first identified. Multiparameter noise models are then derived from the hardware components, and sampled noise is injected into both weights and neuron parameters during training, improving robustness across a wide range of noise levels without increasing model size or significantly increasing training time. A novel training algorithm is introduced in which each mini-batch is evaluated multiple times under different noise samples, with the loss optimized to enforce consistency across various noise values. The software behavioral model is verified using Cadence simulations across 45 binary classifiers. The accuracy of ideal-trained classifiers degrades by up to 44.2% under hardware-induced noise, while the proposed training algorithm improves accuracy by up to 23.34%. The validated framework and training algorithm are then used for larger networks performing image classification on CIFAR-10 and CIFAR-100 datasets, improving accuracy from a random state under noise to competitive levels relative to state-of-the-art, with the main overhead being additional epochs determined by noise range rather than network size.
Rezvan Mohammadrezaee, Ataollah Saeed Monir, Okyanus T. Gumus, Nakisa Shams, Boris Vaisband
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Graph-Based Timing Prediction at Early-Stage RTL Using Large Language Model
abstract
Early-stage timing analyses are essential for exploring design alternatives before physical synthesis in integrated circuit design, which needs to assess signal propagation delay multiple times with varying accuracy. Machine learning (ML) offers promising solutions for early-stage timing prediction, improving result quality while reducing runtime, time-to-market, and non-recurring engineering costs. However, existing ML-based approaches for predicting timing at the register-transfer level (RTL) are not sufficiently reliable to replace traditional electronic design automation tools as they face two key challenges: 1) feature generation based on high-level RTL is unreliable due to unpredictable synthesizer outputs, and 2) they omit essential features like technology library information and design constraints, which are crucial for accurate timing analysis.
Fahad Rahman Amik, Yousef Safari, Zhanguang Zhang, Boris Vaisband
ASP-DAC4
2025 Invited: Chiplet-Based Integration - Scale-Down and Scale-Out
abstract
Motivation: The demand for increased computation and memory in applications such as large language models, has increased well beyond the reticle boundaries of a system-on-chip (SoC). Chiplet-based integration is a paradigm shift that shapes the way we design our future high-performance systems. The concept is to move away from large SoCs that are limited by communication, thermal design power, and reticle size, toward a robust plug-and-play approach, where small, hardened IP heterogeneous off-the-shelf chiplets are seamlessly integrated on a single platform.
Boris Vaisband
ISPD1
2025 Thermal Simulator for Advanced Packaging and Chiplet-Based Systems
abstract
Heterogeneous chiplet-based integration is expected to provide performance scalability and cost-effectiveness for the next generation of microelectronic systems. Practical deployment of chiplet-based platforms, however, requires developing novel electronic design automation (EDA) tools that support advanced packaging approaches. Compact thermal simulators are essential EDA tools for the evaluation of design alternatives at the early stages of the design. Developing efficient compact thermal simulators for advanced heterogeneous integration platforms is a key requirement, as the available tools provide limited support for heterogeneity and advanced packaging technologies. ARTSim 2.0, a robust thermal simulator for heterogeneous integration platforms, is presented in this work. ARTSim 2.0 includes three main features, i.e., robust hybrid meshing, modeling of heterogeneous layers, and an efficient solver that utilizes parallel processing. Several case studies on advanced chiplet-based platforms, including TSV-based 3-D integrated circuits (ICs), Intel EMIB, and TSMC InFO_PoP, are conducted to demonstrate the novel capabilities of ARTSim 2.0. The performance of ARTSim 2.0 for both transient and steady-state conditions is compared to results obtained from state-of-the-art finite element method (FEM) tools. Simulation results confirm that the temperature accuracy of the thermal maps that are generated by ARTSim 2.0 is within a maximum error of 1.17% while exhibiting a reduction in runtime of at least two orders of magnitude, as compared to the FEM tools.
Yousef Safari, Adam Corbier, Dima Al Saleh, Fahad Rahman Amik, Boris Vaisband
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Co-DTC: Concentric Trench-Based Integrated Capacitors for Advanced Chiplet-Based Platforms
abstract
Integrated power delivery methodology is a promising approach to achieve high power efficiency as the semiconductor industry targets higher power density and increased heterogeneity across different load characteristics. High quality and densely integrated passive devices are key components for the realization of integrated power delivery approaches.
Yousef Safari, Yushu Zhao, Boris Vaisband
ACM Great Lakes Symposium on VLSI3
2023 Hybrid Obfuscation of Chiplet-Based Systems
abstract
The growing concern about offshore chip manufacturing has created considerable interest in solutions that can ensure the integrity and security of chips. Among various solutions, split manufacturing has received a lot of attention due to its security guarantees. With the recent emergence of new heterogeneous manufacturing technologies, including chiplet-based systems, there is a new opportunity for revisiting the design considerations for split manufacturing to fully exploit the opportunities presented by chiplet-based systems and improve various metrics, such as security, performance, and overhead.This work improves the state-of-the-art in secure chip manufacturing by proposing a new split manufacturing scheme. The key idea is to exploit the capabilities provided by chiplet integration technology for designing a new hybrid split manufacturing scheme that includes both vertical and horizontal splitting. Unlike existing vertical-only split manufacturing mechanisms, that target obfuscation of interconnections by splitting the design at a specific metallization layer into two portions, the proposed hybrid method increases trust by exploiting the chiplet paradigm shift, specifically, breaking the design into sub-designs, each represented by chiplets (independently fabricated), and obfuscating interconnections among them. The proposed obfuscation mechanism targets systems that exploit the chiplet technology to obtain important performance advantages, thus any chiplet-related overhead is not due to obfuscation. We evaluate our method using several experiments and compare it with the state-of-the-art using standard metrics, including area, power, delay, wirelength, and trust. Compared to conventional split manufacturing, our hybrid method achieves up to 245× higher trust, while exhibiting negligible overhead.
Yousef Safari, Pooya Aghanoury, Subramanian S. Iyer, Nader Sehatbakhsh, Boris Vaisband
DAC5
2023 Statistical Weight Refresh System for CTT-Based Synaptic Arrays
abstract
Charge-trap transistors (CTTs) are compute-in-memory devices that are used to model synaptic arrays in neuromorphic systems. CTTs enable non von Neumann architectures, thus, eliminating the energy spent on compute-memory communication. Synaptic weights can be stored in CTTs by shifting the threshold voltage of the devices in an analog manner. CTTs are, however, susceptible to unintentional de-trapping of charge over time due to threshold voltage instability, leading to loss of the stored synaptic weights. The proposed weight refresh system performs statistical refresh of the CTT array to replenish the charge of individual CTT devices (restore synaptic weights) based on characterization of threshold voltage instability in high-k dielectrics.
Samuel Dayo, Ataollah Saeed Monir, Mousa Karimi, Boris Vaisband
ACM Great Lakes Symposium on VLSI4
2023 Digital LIF Neuron for CTT-Based Neuromorphic Systems
abstract
In this work, a novel digital leaky integrate-and-fire neuron design is proposed as part of a charge-trap transistor (CTT)-based neuromorphic system. CTTs, which are compute-in-memory devices, are used to realize the synaptic array of the neuron and support weight multiplication operations for incoming pulse signals. The proposed digital neuron does not rely on a capacitor for accumulation, making it area-efficient and scalable, and thus useful for design of large spiking neural networks. The neuron accumulates the weighted inputs from the synaptic array and generates an outgoing pulse, i.e., fires, when a pre-set threshold is reached. The digital neuron includes a sampler circuit, multi-level comparator, pulse generator, leaky circuit, 3-bit counter, and digital comparator circuit. Since the circuit is digital, the design is robust to noise, mismatch, and process, voltage, and temperature variations. The digital neuron is designed in GF 22 nm FDSOI technology, operates at a supply voltage of 0.8 V, and occupies an area of 33.5 μ m2. The neuron was simulated, including under temperature and supply voltage variations, and exhibits expected functionality.
Okyanus T. Gumus, Mousa Karimi, Boris Vaisband
ACM Great Lakes Symposium on VLSI3
2023 A Robust Integrated Power Delivery Methodology for 3-D ICs
abstract
The inherent advantages of three-dimensional (3-D) integrated circuits (ICs) are well-aligned with the continuous demand for increased density of functionality, reduced latency, the power dissipation of communication, and heterogeneity of modern applications. Delivering power efficiently to highly heterogeneous voltage domains across the tiers of a 3-D IC is, however, a significant challenge. To address the power delivery challenge in 3-D ICs, a robust integrated power delivery methodology is proposed in this article. Recent advancements in the fabrication of high-density integrated passive components, and the area that is available in the vertical dimension of the 3-D construct, are exploited in this work to enable an efficient and robust power delivery system for 3-D ICs. In the proposed approach, one or more layers within the 3-D structure are dedicated to power conversion and regulation, namely, power layers (PLs). A design exploration stage is also provided to determine the number of PLs, distribution of resources between power and functional layers (FLs), assignment of voltage domains to PLs, and voltage levels across the power delivery system. The proposed methodology is compared to three other power delivery topologies and exhibits 1.4–$38\times $and 1.4–$7.1\times $improvement in, respectively, voltage drop and power efficiency. Results are normalized to the total on- and off-chip area dedicated to power conversion and regulation in each topology.
Yousef Safari, Boris Vaisband
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Power Delivery for Ultra-Large-Scale Applications on Si-IF
abstract
In recent years, with the rise of artificial intelligence and big data, there is an even greater demand for scaling out computing and memory capacity. Silicon interconnect fabric (Si-IF), a wafer-scale integration platform, promotes a paradigm shift in packaging features and enables ultra-large-scale systems, while significantly improving communication bandwidth and latency. Such systems are expected to dissipate tens of kilowatts of power. Designing an efficient and robust power delivery methodology for these high power applications is a key challenge in the enablement of the Si-IF platform. Based on several figure-of-merit parameters, an efficient power delivery methodology is matched with each of three candidate applications on the Si-IF, namely, artificial intelligence accelerators, high-performance computing, and neuromorphic computing. The proposed power delivery approaches were simulated and exhibit compatibility with the relevant ultra-large-scale application on Si-IF. The simulation results confirm that the dedicated power delivery topologies can support ultra-large-scale applications on the SI-IF.
Yousef Safari, Anja Kroon, Boris Vaisband
ISCAS3
2021 Power Delivery for Silicon Interconnect Fabric
abstract
Silicon interconnect fabric (Si-IF) is a wafer-scale heterogeneous integration platform. This platform promotes a paradigm shift in system integration and packaging methods, providing a single hierarchy of integration between the dies and the platform. The Si-IF effectively replaces the interposer, package, and printed circuit board. A power delivery methodology for high power wafer-scale systems (expected to dissipate up to 50 kW of power) is proposed in this paper. The proposed methodology includes three distinct power distribution topologies that are compared in terms of power loss, thermal consideration, and manufacturability. Compatible applications for each topology are also discussed. The electrical model, IR drop, and Ldi/dt noise, of each power distribution topology, are extracted and compared. Assuming a load voltage of 1 V, the three topologies exhibit a total voltage drop of, respectively, 16.68 mV, 9.62 mV, and 12.28 mV, corresponding to, respectively, 1.67%, 0.96%, and 1.23%. Hierarchical integration of decoupling capacitors is also described to ensure low voltage ripple (<; 5%) at the point of load. The electrical models of the power distribution topologies are verified using FEM and SPICE simulations.
Yousef Safari, Boris Vaisband
ISCAS2
2020 Multi-Bit CNT TSV for 3-D ICs
abstract
Through substrate vias (TSVs) are a seminal component of three-dimensional (3-D) integrated circuits (ICs). Each TSV typically carries a single signal between two adjacent layers of a 3-D structure. A multi-bit carbon nanotube TSV is proposed in this paper to increase the number of I/Os among layers within 3-D ICs. The proposed multi-bit TSV can propagate multiple independent signals due to the high anisotropy of the carbon nanotubes. The electrical properties of each bit within a two-bit TSV and the electrical interactions between the bits are compared to a theory-based electrical model, exhibiting high accuracy. The passive elements deviate by up to 4%, and the S-parameters of the system deviate by up to 1.5% from numerical analysis. Capacitive coupling and leakage current between the bits of the two-bit TSV model have also been evaluated. The structure exhibits negligible noise coupling (less than 1%) and a peak leakage current of 631.7 μA.
Boris Vaisband, Ange Maurice, Chong Wei Tan, Beng Kang Tay, Eby G. Friedman
ISCAS1
2019 Global and semi-global communication on Si-IF
abstract
On-chip scaling continues to pose significant technological and design challenges. Nonetheless, the key obstacle in on-chip scaling is the high fabrication cost of the state-of-the-art technology nodes. An opportunity exists however, to continue scaling at the system level. Silicon interconnect fabric (Si-IF) is a platform that aims to replace both the package and printed circuit board to enable heterogeneous integration and high inter-chip performance. Bare dies are attached directly to the Si-IF at fine vertical interconnect pitch (2 to 10 μm) and small inter-die spacing (≤ 100 μm). The Si-IF is a single-hierarchy integration construct that supports dies of any process, technology, and dimensions. In addition to development of the fabrication and integration processes, system-level challenges need to be addressed to enable integration of heterogeneous systems on the Si-IF. Communication is a fundamental challenge on large Si-IF platforms (up to 300 mm diameter wafers). Different technological and design approaches for global and semi-global communication are discussed in this paper. The area overhead associated with global communication on the Si-IF is determined.
Boris Vaisband, Subramanian S. Iyer
NOCS1
2018 Heterogeneous 3-D ICs as a platform for hybrid energy harvesting in IoT systems
Boris Vaisband, Eby G. Friedman
Future Gener. Comput. Syst.1
2017 Hybrid energy harvesting in 3-D IC IoT devices
abstract
Three-dimensional integrated circuits are a natural platform for IoT devices. IoT devices exhibit a small footprint, integrate disparate technologies, and require long term sustainability (extremely low power or self powered). A hybrid energy harvesting system within a three-dimensional integrated circuit is proposed in this paper. The harvesting system exploits different types of energy available from the ambient (electromagnetic, solar, thermal, and kinetic). Integration of the hybrid harvesting system onto a three-dimensional platform ensures that each type of harvested energy can be individually optimized. In addition, lower parasitic impedances are exhibited within the three-dimensional structure, leading to improved efficiencies in the energy harvesting process. For an example IoT system, the power requirements are less than 57% of the power delivered to the load.
Boris Vaisband, Eby G. Friedman
ISCAS1
2016 Layer ordering to minimize TSVs in heterogeneous 3-D ICs
abstract
A layer ordering algorithm to minimize the total number of TSVs within heterogeneous 3-D integrated circuits is described in this paper. Different constraints may complicate the process of ordering the layers within a 3-D system. These constraints are (1) any two layers must be adjacent, (2) a layer must be placed at a specific location, and (3) a layer must be separated from another layer(s). The algorithm generates an optimal layer order given the number of I/Os among all layers. Certain layers can be pre-assigned to specific locations within the 3-D structure. The application of the algorithm to multiple layer 3-D structures significantly reduces the number of TSVs and occupied area as compared to a random layer assignment. The area overhead of a random solution as compared to the optimal solution for unconstrained 3-D systems (without pre-assigned layers) with three to ten layers is, respectively, ~24,090 μm2to ~854,469 μm2. In constrained 3-D systems (with pre-assigned layers), the area overhead for an eight layer 3-D system with one to six assigned layers ranges up to ~249, 240 μm2.
Boris Vaisband, Eby G. Friedman
ISCAS1
2016 Noise Coupling Models in Heterogeneous 3-D ICs
abstract
Models of coupling noise from an aggressor module to a victim module by way of through silicon vias (TSVs) within heterogeneous 3-D integrated circuits (ICs) are presented in this paper. Existing TSV models are enhanced for different substrate materials within heterogeneous 3-D ICs. Each model is adapted to each substrate material according to the local noise coupling characteristics. The 3-D noise coupling system is evaluated for isolation efficiency over frequencies of up to 100 GHz. Isolation improvement techniques, such as reducing the ground network inductance and increasing the distance between the aggressor and victim modules, are quantified in terms of noise improvements. A maximum improvement of 73.5 dB for different ground network impedances and a difference of 38.5 dB in isolation efficiency for greater separation between the aggressor and victim modules are demonstrated. Compact, accurate, and computationally efficient models are extracted from the transfer function for each of the heterogeneous substrate materials. The reduced transfer functions are used to explore different manufacturing and design parameters to evaluate coupling noise across multiple 3-D planes.
Boris Vaisband, Eby G. Friedman
IEEE Trans. Very Large Scale Integr. Syst.1
2015 3-D floorplanning algorithm to minimize thermal interactions
abstract
An algorithm for including relative thermal interactions among different circuit modules within a 3-D system is introduced in this paper. Application of the proposed algorithm on MCNC and GSRC benchmark circuits is presented. The thermal behavior of a heterogeneous 3-D structure, consisting of a different number of modules and substrate materials, is evaluated to emulate the heat transfer characteristics of practical heterogeneous 3-D systems. The algorithm lowers thermal interactions between different modules while maintaining the peak temperature within a practical range. The thermal characteristics of the floorplan are evaluated using HotSpot and HotSpot Detailed 3-D and compared to a random floorplan. The recorded peak temperatures are within the practical range of on-chip temperatures.
Boris Vaisband, Eby G. Friedman
ISCAS1
2015 Experimental Analysis of Thermal Coupling in 3-D Integrated Circuits
abstract
A 3-D test circuit examining thermal propagation within a through-silicon via-based 3-D integrated stack has been designed, fabricated, and tested. Design insight into thermal coupling in 3-D integrated circuits (ICs) through both experiment and simulation is provided, and suggestions to mitigate thermal effects in 3-D ICs are offered. Two wafers are vertically bonded to form a 3-D stack. Intraplane and interplane thermal coupling is investigated through single-point heat generation using resistive thermal heaters and temperature monitoring through four-point resistive measurements. Thermal paths are identified and analyzed based on the metric of thermal resistance per unit length. The peak steady-state temperature due to die location within a 3-D stack is described. The reduction in peak temperature through fan-based active cooling is also reported. Thermal propagation from a heat source located on the backside of the silicon is examined with both back metal and on-chip thermal sensors. A comparison of thermal coupling between two different heat sources on the same device plane is also provided.
Ioannis Savidis, Boris Vaisband, Eby G. Friedman
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Thermal conduction path analysis in 3-D ICs
abstract
The on-going effort of integrating heterogeneous circuits as well as the increasing length of global interconnect are driving the semiconductor community towards 3-D integrated circuits. In this work, thermal paths within a 3-D stack are investigated using the HotSpot simulator, and the results are compared to experimental data of a fabricated two layer stack with a single back metal layer. Resistive heaters and sensors measure the heat flow in both the horizontal and vertical dimensions. The dependence of the thermal conductivity on temperature is integrated into the thermal simulation process. At high temperatures (~ 80°C), this effect is responsible for inaccuracies in the temperature and thermal resistance of up to, respectively, 20% and 28%. As confirmed by simulation, those horizontal paths that lie mostly within the silicon layer conduct more heat as compared to the vertical paths, since the thermal conductivity of silicon dioxide is ~ 200 times smaller than the thermal conductivity of silicon.
Boris Vaisband, Ioannis Savidis, Eby G. Friedman
ISCAS1