Muhannad S. Bakir

dblp:52/7357 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-0380-0842ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Omelet: A Packaging-Aware Hierarchical Interconnect Simulator for 2.5D/3D Chiplet Architectures
Danish Baig, Faaiq Waqar, Ashita Victor, Shimeng Yu, Muhannad S. Bakir, Cong Hao
ISCA6
2025 Exploiting Chiplet Integration Technology for Fast High-Capacity DRAM Modules
abstract
As the end of Moore’s law approaches, chiplet integration technology (or chiplet technology) has emerged to revolutionize future semiconductor chip design. Chiplet technology provides unique advantages over 3-D-stacking technology, including a more cost-efficient and thermal-friendly integration of heterogeneous technologies. Although chiplet technologies have already begun to be used by the latest commercial chips, they have not been explored for commodity dynamic random access memory (DRAM) design yet. Harnessing its advantages for DRAM for the first time, this article evaluates the feasibility of chiplet-based DRAM architecture, considering various physical and electrical constraints imposed by a standard chiplet interface [i.e., universal chiplet interconnect express (UCIe)]. We further explore the DIMM architectures that simplify module packaging and assembly, leading to reductions in total die size and overall costs. The comprehensive cross-level analysis (i.e., device, circuit, chip, and system levels) shows that chiplet-based DRAM reducest_RCD+t_CAS, latency-critical DRAM timing parameters, by$1.32\times $–$1.39\times $, at the same energy consumption. In addition, a$1.39\times $–$2.28\times $improvement int_RRDis obtained. The reduced DRAM timing parameters improve the overall system performance by up to 8.8%–24.7% (geomean 3.4%–8.4%) in real-life benchmarks. The chiplet-based heterogeneous integration achieves a$1.27\times $higher chip-level yield compared with the monolithic chip, along with up to 10% reduction in overall cost compared with traditional DIMMs at emerging process technologies.
Zihan Xia 0002, Chihun Song, Ram Krishna, Ashita Victor, Srujan Penta, Muhannad S. Bakir, Elyse Rosenbaum, Nam Sung Kim, Mingu Kang
IEEE Trans. Very Large Scale Integr. Syst.6
2023 H3DAtten: Heterogeneous 3-D Integrated Hybrid Analog and Digital Compute-in-Memory Accelerator for Vision Transformer Self-Attention
abstract
After the success of the transformer networks on natural language processing (NLP), the application of transformers to computer vision (CV) has followed suit to deliver unprecedented performance gains on vision tasks, including image recognition and object detection. The multihead self-attention (MHSA) is the key component in transformers, allowing the models to learn the amount of attention paid to each input position. Despite its strong modeling capability, MHSA involves complex operations that make transformers prohibitively costly for hardware deployment. Existing acceleration efforts with conventional hardware platforms are challenged by the memory wall. To alleviate the memory wall problem, compute-in-memory (CIM) is a promising solution by storing all model parameters on-chip in compute-capable memory arrays. The footprint of 2-D CIM designs must, however, expand to accommodate the increasingly larger model sizes. In this work, we present a heterogeneous 3-D integrated (H3D) accelerator to target the MHSA workloads in vision transformers. H3D allows the proposed H3DAtten architecture to combine the merits of resistive random access memory (RRAM)-based analog CIM (ACIM) in 40 nm and static random access memory (SRAM)-based digital CIM (DCIM) in 16 nm. We perform comprehensive signaling and thermal analyses to examine the effects of 3-D stacking on the accelerator. Compared to iso-capacity 2-D baseline designs, the proposed 5-tier H3DAtten accelerator achieves$8.4\times $compute density without experiencing accuracy loss on the ImageNet-1k dataset.
Wantong Li 0002, Madison Manley, James Read, Ankit Kaul, Muhannad S. Bakir, Shimeng Yu
IEEE Trans. Very Large Scale Integr. Syst.5
2020 A Model Study of Multilevel Signaling for High-Speed Chiplet-to-Chiplet Communication in 2.5D Integration
abstract
The quest for high yield has motivated significant advancement in 2.5D integrated circuits, where chiplets are integrated on a silicon interposer or a package substrate with high-speed parallel communication among them. These channels for 2.5D integrated systems need to have high data bandwidth per unit length (also called shoreline-BW-density and measured in Gb/s/mm) and lower energy per bit area (measured in pJ/b). Typically, NRZ signalling is used but achieving higher data rates continues to be a major challenge. In this paper we explore PAM4 as an alternative to NRZ for signalling the channels. Simulations show that we can achieve up to 63% more energy-efficiency and 27% higher BW density for 2.5D integrated systems.
Rakshith Saligram, Ankit Kaul, Muhannad S. Bakir, Arijit Raychowdhury
VLSI-SOC3
2017 PhD forum: Heterogeneous interconnection of ICs using stitch-chips
abstract
In this paper, a heterogeneous interconnect stitching technology (HIST) is presented. Stitch chips with high-density fine pitch wires are used to connect active dice of various functions in a manner that mimics system-on-chip (SoC) like performance. Microbumps and compressible microinterconnects (CMIs) are used to provide die-to-die and die-to-package interconnection. A testbed containing two dummy dice and one stitch chip is fabricated and tested. The average measured post-assembly resistance of the microbumps and the CMIs is 116.5 μΩ and 195.9 mμ, respectively. Using an electrical model for HIST signal channels, a 50%-50% delay as small as 140 ps with an energy efficiency of 0.24 pJ/bit may be achieved for a 1mm channel. In addition, some power delivery challenges and opportunities of HIST are highlighted through simulations. The results show that compared to interposer based 2.5-D integration, HIST may reduce the IRdrop by approximately 18.4%.
Muhannad S. Bakir
VLSI-SoC1
2009 Co-design of signal, power, and thermal distribution networks for 3D ICs
abstract
Heat removal and power delivery are two major reliability concerns in the 3D stacked IC technology. Liquid cooling based on micro-fluidic channels is proposed as a viable solution to dramatically reduce the operating temperature of 3D ICs. In addition, designers use a highly complex hierarchical power distribution network in conjunction with decoupling capacitors to deliver currents to all parts of the 3D IC while suppressing the power supply noise to an acceptable level. These so called silicon ancillary technologies, however, pose major challenges to routing completion and congestion. These thermal and power/ground interconnects together with those used for signal delivery compete with one another for routing resources including various types of Through-Silicon-Vias (TSVs). This paper presents the work on routing with these interconnects in 3D: signal, power, and thermal networks. We demonstrate how to consider various physical, electrical, and thermo-mechnical requirements of these interconnects to successfully complete routing while addressing various reliability concerns.
Young-Joon Lee, Yoon Jo Kim, Muhannad S. Bakir, Yogendra K. Joshi, Andrei G. Fedorov, Sung Kyu Lim
DATE4