VLDB 2026 Research / reviewers in the wild / expert
Hsiao-Hsuan Liu
dblp:336/5228
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-2305-4258ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | System Scenario-Based Design of the Last-Level Cache in Advanced Interconnect-Dominant Technology NodesabstractFeature size reduction of the front End of the Line (FEoL) and back End of the Line (BEoL) elements, i.e., transistors and interconnects, has been the main enabler of the next-generation computation systems. The decreasing trend of the cross-sectional area of the interconnect in advanced technology nodes, however, comes along with a drastic increase in the resistive parasitic, substantially impacting the overall energy efficiency and performance of the computer system. Mitigation of the high parasitic resistance within an advanced-node static RAM (SRAM)-based last-level cache (LLC) is the main target of this article. To achieve this target, we augment the LLC interconnect with some degree of reconfiguration by utilizing a dynamic segmented bus (DSB). With DSB, the interconnect segments that are most actively used for a given workload can be shortened, on average, contributing to a smaller capacitive load. Hence, the efficient reconfiguration of an LLC interconnect strongly depends on the LLC demands of the application. To account for this workload dependency, we design the required microarchitectural support in an end-to-end application-to-technology flow. By optimizing the overhead of DSB switches and additional hardware modules, the SRAM-based LLC with DSB-augmented intra-macro interconnect achieves 33% energy savings and 16% reduction in total access time across eight representative workloads, with a negligible area overhead of less than 0.4%. Mahta Mayahinia, Tommaso Marinelli, Zhenlin Pei, Hsiao-Hsuan Liu, Chenyun Pan, Zsolt Tokei, Francky Catthoor, Mehdi Baradaran Tahoori |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | Interconnect/Memory Co-Design and Co-Optimization Using Differential Transmission LinesabstractAs technology scales down, the performance–power–area (PPA) of static random access memory (SRAM) is increasingly constrained by interconnects due to the presence of large parasitic capacitance and resistance within these structures. This article presents a co-optimization and co-design framework that integrates technology, interconnect, circuit, cache memory, and workload to optimize the overall PPA of the computing cache system through various emerging interconnect technologies under software and hardware conditions. Moreover, we present the differential transmission line (DTL), which is utilized as a hybrid with conventional wires with repeater insertion. The proposed methodology enables the identification of the optimal design, thereby facilitating the reduction of interconnect energy and delay, considering synthetic/realistic workloads and comparing DTL against traditional repeater insertion methods based on metrics of PPA, including the energy–delay–area product (EDAP) and energy–delay product (EDP), for the computing cache system. A thorough design space exploration is conducted, utilizing validated experimental subarrays at the deep scale across state-of-the-art technology nodes. Moreover, the case study assesses a range of cache system parameters, emphasizing the potential of DTL interconnect technologies to enhance cache memory PPA. Zhenlin Pei, Hsiao-Hsuan Liu, Mahta Mayahinia, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Prashant Dubey, Chenyun Pan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Energy-efficient SNN Architecture using 3nm FinFET Multiport SRAM-based CIM with Online LearningabstractCurrent Artificial Intelligence (AI) computation systems face challenges, primarily from the memory-wall issue, limiting overall system-level performance, especially for Edge devices with constrained battery budgets, such as smartphones, wearables, and Internet-of-Things sensor systems. In this paper, we propose a new SRAM-based Compute-In-Memory (CIM) accelerator optimized for Spiking Neural Networks (SNNs) Inference. Our proposed architecture employs a multiport SRAM design with multiple decoupled Read ports to enhance the throughput and Transposable Read-Write ports to facilitate online learning. Furthermore, we develop an Arbiter circuit for efficient data-processing and port allocations during the computation. Results for a 128×128 array in 3nm FinFET technology demonstrate a 3.1× improvement in speed and a 2.2× enhancement in energy efficiency with our proposed multiport SRAM design compared to the traditional single-port design. At system-level, a throughput of 44 MInf/s at 607 pJ/Inf and 29mW is achieved. Lucas Huijbregts, Hsiao-Hsuan Liu, Paul Detterer, Said Hamdioui, Amirreza Yousefzadeh, Rajendra Bishnoi |
DAC | 2 |
| 2024 | Future Design Direction for SRAM Data Array: Hierarchical Subarray With Active InterconnectabstractIn sub 10 nm nodes, the growing dominance of interconnects in chips poses challenges in designing large-size static random-access memory (SRAM) subarrays. The main issue is the write failure problem arising from the increased resistance and capacitance for bitline (BL) and wordline (WL). To tackle this issue, the SRAM subarray design incorporates conventional (Conv.) divided WL and divided BL techniques based on 14-Å-compatible (A14) nanosheet (NS) technology. This approach allows for various subarray sizes with successful write operations, resulting in improved subarray-level performance and power (PP). However, the additional logic gates come with an area penalty that may degrade the overall performance, power, and area (PPA) at the macro level due to increased inter-subarray interconnect overhead. To overcome this limitation, the active interconnect (AIC) design is proposed with the features of fabricating another or multiple active regions at the back-end of line (BEOL) layers. By moving these extra logic gates from front-end of line to BEOL in the AIC divided subarray design, the area penalty is significantly mitigated without compromising PP compared to the standard (Std.) and Conv. divided counterparts. To achieve this concept, carbon nanotube gate-all-around transistor is explored as potential BEOL-compatible device. In this research, a comprehensive design-technology co-optimization analysis is conducted to verify the value and potential benefits of up to 65% macro-level energy-delay-area product improvement by AIC divided subarray design compared to the Std. subarray design. Hsiao-Hsuan Liu, Carlo Gilardi, Shairfe Muhammad Salahuddin, Zhenlin Pei, Pieter Schuddinck, Pieter Weckx, Geert Hellings, Marie Garcia Bardon, Julien Ryckaert, Chenyun Pan, Subhasish Mitra, Francky Catthoor |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | Ultra-Scaled E-Tree-Based SRAM Design and Optimization With Interconnect FocusabstractSRAM performance is highly dominated by interconnects as technology scales down because of the significant parasitic resistance and capacitance in the interconnect. This paper introduces a framework for the co-design of technology, interconnect, and cache memory with tag array overhead, to optimize the performance of cache memory using a variety of emerging interconnect technologies. In addition, we introduce an innovative E-Tree interconnect aimed at further decreasing the average interconnect length with the consideration of realistic workloads and benchmark against its traditional H-Tree counterparts in terms of various performance metrics, such as energy-delay-area product (EDAP) or energy-delay product (EDP) in the SRAM cache memory system. A comprehensive investigation of design space is conducted, employing realistic, deeply scaled subarray designs across a range of cutting-edge technology nodes. Furthermore, the case study examines various cache memory system design parameters to assess the true potential of emerging interconnect technologies in achieving optimal performance at the cache memory system. Zhenlin Pei, Hsiao-Hsuan Liu, Mahta Mayahinia, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Dawit Burusie Abdi, James Myers, Chenyun Pan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Electromigration-aware design technology co-optimization for SRAM in advanced technology nodesabstractStatic RAM (SRAM) is one of the critical components in advanced VLSI systems whose performance, capacity, and reliability have a decisive impact on the entire system. It offers the fastest memory in the storage hierarchy of modern computer systems. By moving toward the smaller CMOS technology nodes, the back end of the line (BEoL) interconnects are also fabricated in tighter pitch size. Hence, besides the power lines, SRAM word- and bit-line (WL and BL) are also susceptible to electromigration (EM). Therefore, EM reliability of SRAM's WL and BL needs to be analyzed during design technology co-optimization (DTCO) cycle. In this work, we investigate the impact of technology scaling on SRAM designs and perform a detailed analysis on the trend of their EM reliability and energy consumption. Our analysis shows that although scaling down the CMOS technology can result in a 2.68x improvement in the energy efficiency of the SRAM module, it increases the EM-induced hydrostatic stress by 2.53x. Mahta Mayahinia, Hsiao-Hsuan Liu, Subrat Mishra, Zsolt Tokei, Francky Catthoor, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2023 | Technology/Memory Co-Design and Co-Optimization Using E-Tree InterconnectabstractFor on-chip SRAM, a major portion of delay and energy is contributed by the H-Tree interconnects. In this paper, we propose an E-Tree interconnect technology to minimize the H-Tree delay and energy overheads based on an efficient interconnect technology/memory co-design framework for nonuniform workloads. Various array- and interconnect-level design parameters are co-designed for optimal performance using three emerging interconnect materials with a realistic cell library. Zhenlin Pei, Mahta Mayahinia, Hsiao-Hsuan Liu, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Chenyun Pan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Energy-efficient In-Memory Address CalculationabstractComputation-in-Memory (CIM) is an emerging computing paradigm to address memory bottleneck challenges in computer architecture. A CIM unit cannot fully replace a general-purpose processor. Still, it significantly reduces the amount of data transfer between a traditional memory unit and the processor by enriching the transferred information. Data transactions between processor and memory consist of memory access addresses and values. While the main focus in the field of in-memory computing is to apply computations on the content of the memory (values), the importance of CPU-CIM address transactions and calculations for generating the sequence of access addresses for data-dominated applications is generally overlooked. However, the amount of information transactions used for “address” can easily be even more than half of the total transferred bits in many applications. In this article, we propose a circuit to perform the in-memory Address Calculation Accelerator. Our simulation results showed that calculating address sequences inside the memory (instead of the CPU) can significantly reduce the CPU-CIM address transactions and therefore contribute to considerable energy saving, latency, and bus traffic. For a chosen application of guided image filtering, in-memory address calculation results in almost two orders of magnitude reduction in address transactions over the memory bus. Amirreza Yousefzadeh, Jan Stuijt, Martijn Hijdra, Hsiao-Hsuan Liu, Anteneh Gebregiorgis, Abhairaj Singh, Said Hamdioui, Francky Catthoor |
ACM Trans. Archit. Code Optim. | 4 |