Zhenlin Pei

dblp:348/3780 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-0926-2838ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Reliability-Driven Sneak Path Current Modeling and Optimization for Passive Memristor Crossbar Arrays
Zhenlin Pei, Shah Zayed Riam, Kyle Mooney, Chenyun Pan, Na Gong
ACM Great Lakes Symposium on VLSI1
2025 System Scenario-Based Design of the Last-Level Cache in Advanced Interconnect-Dominant Technology Nodes
abstract
Feature size reduction of the front End of the Line (FEoL) and back End of the Line (BEoL) elements, i.e., transistors and interconnects, has been the main enabler of the next-generation computation systems. The decreasing trend of the cross-sectional area of the interconnect in advanced technology nodes, however, comes along with a drastic increase in the resistive parasitic, substantially impacting the overall energy efficiency and performance of the computer system. Mitigation of the high parasitic resistance within an advanced-node static RAM (SRAM)-based last-level cache (LLC) is the main target of this article. To achieve this target, we augment the LLC interconnect with some degree of reconfiguration by utilizing a dynamic segmented bus (DSB). With DSB, the interconnect segments that are most actively used for a given workload can be shortened, on average, contributing to a smaller capacitive load. Hence, the efficient reconfiguration of an LLC interconnect strongly depends on the LLC demands of the application. To account for this workload dependency, we design the required microarchitectural support in an end-to-end application-to-technology flow. By optimizing the overhead of DSB switches and additional hardware modules, the SRAM-based LLC with DSB-augmented intra-macro interconnect achieves 33% energy savings and 16% reduction in total access time across eight representative workloads, with a negligible area overhead of less than 0.4%.
Mahta Mayahinia, Tommaso Marinelli, Zhenlin Pei, Hsiao-Hsuan Liu, Chenyun Pan, Zsolt Tokei, Francky Catthoor, Mehdi Baradaran Tahoori
ACM Trans. Embed. Comput. Syst.3
2025 Interconnect/Memory Co-Design and Co-Optimization Using Differential Transmission Lines
abstract
As technology scales down, the performance–power–area (PPA) of static random access memory (SRAM) is increasingly constrained by interconnects due to the presence of large parasitic capacitance and resistance within these structures. This article presents a co-optimization and co-design framework that integrates technology, interconnect, circuit, cache memory, and workload to optimize the overall PPA of the computing cache system through various emerging interconnect technologies under software and hardware conditions. Moreover, we present the differential transmission line (DTL), which is utilized as a hybrid with conventional wires with repeater insertion. The proposed methodology enables the identification of the optimal design, thereby facilitating the reduction of interconnect energy and delay, considering synthetic/realistic workloads and comparing DTL against traditional repeater insertion methods based on metrics of PPA, including the energy–delay–area product (EDAP) and energy–delay product (EDP), for the computing cache system. A thorough design space exploration is conducted, utilizing validated experimental subarrays at the deep scale across state-of-the-art technology nodes. Moreover, the case study assesses a range of cache system parameters, emphasizing the potential of DTL interconnect technologies to enhance cache memory PPA.
Zhenlin Pei, Hsiao-Hsuan Liu, Mahta Mayahinia, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Prashant Dubey, Chenyun Pan
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Future Design Direction for SRAM Data Array: Hierarchical Subarray With Active Interconnect
abstract
In sub 10 nm nodes, the growing dominance of interconnects in chips poses challenges in designing large-size static random-access memory (SRAM) subarrays. The main issue is the write failure problem arising from the increased resistance and capacitance for bitline (BL) and wordline (WL). To tackle this issue, the SRAM subarray design incorporates conventional (Conv.) divided WL and divided BL techniques based on 14-Å-compatible (A14) nanosheet (NS) technology. This approach allows for various subarray sizes with successful write operations, resulting in improved subarray-level performance and power (PP). However, the additional logic gates come with an area penalty that may degrade the overall performance, power, and area (PPA) at the macro level due to increased inter-subarray interconnect overhead. To overcome this limitation, the active interconnect (AIC) design is proposed with the features of fabricating another or multiple active regions at the back-end of line (BEOL) layers. By moving these extra logic gates from front-end of line to BEOL in the AIC divided subarray design, the area penalty is significantly mitigated without compromising PP compared to the standard (Std.) and Conv. divided counterparts. To achieve this concept, carbon nanotube gate-all-around transistor is explored as potential BEOL-compatible device. In this research, a comprehensive design-technology co-optimization analysis is conducted to verify the value and potential benefits of up to 65% macro-level energy-delay-area product improvement by AIC divided subarray design compared to the Std. subarray design.
Hsiao-Hsuan Liu, Carlo Gilardi, Shairfe Muhammad Salahuddin, Zhenlin Pei, Pieter Schuddinck, Pieter Weckx, Geert Hellings, Marie Garcia Bardon, Julien Ryckaert, Chenyun Pan, Subhasish Mitra, Francky Catthoor
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Ultra-Scaled E-Tree-Based SRAM Design and Optimization With Interconnect Focus
abstract
SRAM performance is highly dominated by interconnects as technology scales down because of the significant parasitic resistance and capacitance in the interconnect. This paper introduces a framework for the co-design of technology, interconnect, and cache memory with tag array overhead, to optimize the performance of cache memory using a variety of emerging interconnect technologies. In addition, we introduce an innovative E-Tree interconnect aimed at further decreasing the average interconnect length with the consideration of realistic workloads and benchmark against its traditional H-Tree counterparts in terms of various performance metrics, such as energy-delay-area product (EDAP) or energy-delay product (EDP) in the SRAM cache memory system. A comprehensive investigation of design space is conducted, employing realistic, deeply scaled subarray designs across a range of cutting-edge technology nodes. Furthermore, the case study examines various cache memory system design parameters to assess the true potential of emerging interconnect technologies in achieving optimal performance at the cache memory system.
Zhenlin Pei, Hsiao-Hsuan Liu, Mahta Mayahinia, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Dawit Burusie Abdi, James Myers, Chenyun Pan
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Technology/Memory Co-Design and Co-Optimization Using E-Tree Interconnect
abstract
For on-chip SRAM, a major portion of delay and energy is contributed by the H-Tree interconnects. In this paper, we propose an E-Tree interconnect technology to minimize the H-Tree delay and energy overheads based on an efficient interconnect technology/memory co-design framework for nonuniform workloads. Various array- and interconnect-level design parameters are co-designed for optimal performance using three emerging interconnect materials with a realistic cell library.
Zhenlin Pei, Mahta Mayahinia, Hsiao-Hsuan Liu, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Chenyun Pan
ACM Great Lakes Symposium on VLSI1