EDBT 2026 Demo / reviewers in the wild / expert
Lingjun Zhu
dblp:129/1811
· DBLP profile ↗
31ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 9 first-author · 21 since 2021Computer networks · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICabstractAs the gap between network and CPU speeds rapidly increases, the CPU-centric network stack proves inadequate due to excessive CPU and memory overheads. Though hardware-offloaded network stacks alleviate these issues, they suffer from limited flexibility in both control and data planes. It seems promising to offload network stacks to Smart-NICs to provide high flexibility. However, naive offloading leads to low throughput due to the inherent architectural limitations of widespread off-path SmartNICs. Even simple operations on staged network traffic would overwhelm the limited SmartNIC memory bandwidth. To this end, we design SmartNS, a SmartNIC-centric network stack with software transport programmability and line-rate packet processing capabilities. To tackle the limitations of SmartNIC-induced challenges, we propose a header-only offloading TX path and an unlimited-working-set in-cache processing RX path to minimize memory traffic to fit the wimpy SmartNIC memory bandwidth. To fully utilize the SmartNIC computing resources, we propose a programmable offloading engine to enable cloud providers to offload customized tasks along with the network stack processing. We prototype SmartNS using the widespread Nvidia BlueField-3 SmartNIC, and implement RoCEv2 and Solar transport protocols by leveraging SmartNS's software programmability. SmartNS achieves 2.2× higher throughput than the microkernel-based baseline in block storage disaggregation and 1.3× higher throughput than the hardware-offloaded baseline in KVCache transfer. Xuzheng Chen, Jie Zhang 0081, Baolin Zhu, Xueying Zhu, Zhongqing Chen, Lingjun Zhu, Yin Zhang 0006, Yuanchao Shu, Peng Cheng 0001, Zeke Wang |
EuroSys | 8 |
| 2026 | Come Hell or Still Water: Alleviating Tail Latency in Cloud Block Store
Chaolei Hu, Kun Qian 0004, Erci Xu, Xue Li 0024, Yuesheng Gu, Lingjun Zhu, Fengyuan Ren, Ennan Zhai |
NSDI | 8 |
| 2025 | Hey Hey, My My, Skewness Is Here to Stay: Challenges and Opportunities in Cloud Block Store TrafficabstractElastic Block Storage (EBS) has a pivotal role in modern data center infrastructure, providing reliable, high-performance and flexible block storage service to users. In Alibaba Cloud, EBS is the most widely used service and has been supporting the operation of millions of virtual disks. However, even with layers of load balancing and caching, we still observe significant traffic skewness across the EBS stack. This motivates us to comprehensively investigate symptoms and root causes behind the traffic patterns and, more importantly, explore the fixes for the identified issues. Erci Xu, Yuandong Hong, Changsheng Niu, Lingjun Zhu, Jinnian He, Weidong Zhang 0011, Qiuping Wang, Changhong Wang 0005, Xinqi Chen, Guangtao Xue, Yi-Chao Chen 0001, Dian Ding |
EuroSys | 7 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 24 |
| 2025 | Glass Interposer Integration of Logic and Memory Chiplets: PPA and Power/Signal Integrity BenefitsabstractGlass interposers have become a compelling option for 2.5-D heterogeneous integration compared to silicon. It allows 3-D stacking configuration between the embedded dies and the conventional flip-chip dies mounted directly on top at low cost. Furthermore, the interconnect pitch and through-glass-via (TGV) diameter in glass are becoming comparable to their counterparts in silicon. In this study, we investigate the power, performance, area (PPA), signal integrity (SI) and power integrity (PI) advantages of 3-D stacking afforded by glass interposers over silicon interposers. Our research employs a chiplet/package co-design approach, progressing from an register-transfer-level description of RISC-V chiplets to final graphic data system (GDS) layouts, utilizing TSMC 28 nm for chiplets and Georgia Tech’s 3-D glass packaging for the interposer. Compared to silicon, glass interposers offer a$2.6\times $reduction in area, a$21\times $reduction in wire length, a 17.72% reduction in full-chip power consumption, a 64.7% increase in SI and a$10\times $improvement in PI, with a 35% increase in thermal. Furthermore, we provide a detailed comparative analysis with 3-D Silicon technologies. It not only highlights the competitive advantages of glass interposers, but also provides critical insights into each design’s potential limitations and optimization opportunities. Pruek Vanna-Iampikul, Seungmin Woo, Serhat Erdogan, Lingjun Zhu, Mohanalingam Kathaperumal, Ravi Agarwal, Ram Gupta, Kevin Rinebold, Madhavan Swaminathan, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Evolving the Cloud Block Store with Performance, Elasticity, Availability, and Hardware OffloadingabstractIn this paper, we qualitatively and quantitatively discuss the design choices, production experience, and lessons in building the Elastic Block Storage ( EBS ) at Alibaba Cloud over the past decade. To cope with hardware advancement and users’ demands, we shift our focus from design simplicity in EBS1 to high performance and space efficiency in EBS2 , and finally reducing network traffic amplification in EBS3 . In addition to the architectural evolutions, we also summarize development lessons and experiences as four topics, including: (i) achieving high elasticity in latency, throughput, IOPS, and capacity; (ii) improving availability by minimizing the blast radius of individual, regional, and global failure events; (iii) identifying the motivations and key tradeoffs in various hardware offloading solutions; and (iv) identifying the pros/cons of alternative solutions and explaining why seemingly promising ideas would not work in practice. Erci Xu, Weidong Zhang 0011, Qiuping Wang, Yuesheng Gu, Zhenwei Lu, Tao Ouyang, Guanqun Dong, Wenwen Peng, Yilei Peng, Tianyun Wang, Wenyuan Yan, Wenhui Yao, Zhongjie Wu, Lingjun Zhu, Yinhu Wang, Junping Wu, Jiaji Zhu, Jiesheng Wu |
ACM Trans. Storage | 21 |
| 2024 | GNN-assisted Back-side Clock Routing Methodology for Advance TechnologiesabstractThe back-side metal layers exhibit lower parasitics compared to the front-side layers in advanced technologies, making them suitable for clock-net distribution. In this study, we explore the advantages of using back-side metal layers for clock routing, which is shared with a power delivery network. Our Graph Neural Network (GNN) based framework, effectively distributes the clock-tree between the front and back sides. We address the back-side clock nets' creation by incorporating back-side buffers. Our results demonstrate better clock and full-chip metrics represented by an increase of up to 13% in the effective frequency with equivalent power consumption, using 3 nm technology. Nesara Eranna Bethur, Pruek Vanna-Iampikul, Odysseas Zografos, Lingjun Zhu, Giuliano Sisto, Dragomir Milojevic, Alberto García Ortiz, Geert Hellings, Julien Ryckaert, Francky Catthoor, Sung Kyu Lim |
DAC | 4 |
| 2024 | What's the Story in EBS Glory: Evolutions and Lessons in Building Cloud Block Store
Weidong Zhang 0011, Erci Xu, Qiuping Wang, Yuesheng Gu, Zhenwei Lu, Tao Ouyang, Guanqun Dai, Wenwen Peng, Yilei Peng, Tianyun Wang, Wenyuan Yan, Wenhui Yao, Zhongjie Wu, Lingjun Zhu, Yinhu Wang, Junping Wu, Jiaji Zhu, Jiesheng Wu |
FAST | 21 |
| 2024 | Demystifying Datapath Accelerator Enhanced Off-path SmartNICabstractNetwork speeds grow quickly in the modern cloud, so SmartNICs are introduced to offload network processing tasks, even application logic. However, typical multicore SmartNICs such as BlueFiled-2 are only capable of processing control-plane tasks with their embedded processors that have limited memory bandwidth and computing power. On the other hand, cloud applications evolve rapidly, such that a limited number of fixed hardware engines in a SmartNIC cannot satisfy the requirements of cloud applications. Therefore, SmartNIC programmers call for a programmable datapath accelerator (DPA) to process network traffic at line rate. However, no existing work has unveiled the performance characteristics of the existing DPA. To this end, we present the first architectural characterization of the latest DPA-enhanced BlueFiled-3 (BF3) SmartNIC. Our evaluation results indicate that BF3's DPA is significantly wimpier than the off-path Arm processor and the host CPU. However, we still identify that DPA has three unique architectural characteristics that unleash the performance potential of DPA. Specifically, we demonstrate how to take advantage of DPA's three architectural characteristics regarding computing, networking, and memory subsystems. Then we propose three important guidelines for programmers to fully unleash the potential of DPA. To demonstrate the effectiveness of our approach, we conduct detailed case studies regarding each guideline. Our case study on key-value aggregation achieves up to$4.3 \times$higher throughput by using our guidelines to optimize memory combinations. Xuzheng Chen, Jie Zhang 0081, Lingjun Zhu, Yin Zhang 0006, Ming Liu 0027, Zeke Wang |
ICNP | 7 |
| 2024 | Hetero-3D: Maximizing Performance and Power Delivery Benefits of Heterogeneous 3D ICsabstractHeterogeneous 3D integration, blending multiple technology nodes, emerges as a promising strategy for enhancing performance and maintaining low power consumption in next-generation computing systems. This paper presents Hetero-3D, an RTL-to-GDS design flow tailored specifically for heterogeneous 3D ICs. Hetero-3D integrates an area-unbalanced 3D floorplanner with an ML-based power delivery and signal router, working in tandem for rigorous PPA (Power, Performance, and Area) optimization. Using two CPU benchmarks, we showcase a remarkable 15% increase in maximum frequency and a substantial 50% decrease in voltage drop compared to homogeneous 3D baselines. Moreover, Hetero-3D effectively addresses voltage drop issues in the power delivery network while delivering an additional 5% frequency boost. This study emphasizes the EDA solutions that unlock the potential of mixed-node stacking as a crucial enabler for performance scaling in future ICs. Lingjun Zhu, Gauthaman Murali, Sung Kyu Lim |
ISLPED | 1 |
| 2024 | A PPA Study for Heterogeneous 3-D IC Options: Monolithic, Hybrid Bonding, and MicrobumpingabstractIn this article, we present three commercial-grade 3-D IC designs based on state-of-the-art design technologies, specifically microbumping (3-D die stacking), hybrid bonding (wafer-on-wafer bonding), and monolithic 3-D (M3D) ICs. To highlight tradeoffs present in these three designs, we perform analyses on power, performance, and area (PPA) and the clock tree. We also model the tier-to-tier interconnection in each 3-D IC methodology and analyze signal integrity (SI) to assess the reliability of each design. From our experiments using the OpenPiton benchmark, the hybrid bonding design shows the best timing improvement of 81.4% when compared to its 2-D counterpart, while microbumping shows the best reliability among 3-D IC designs. Moreover, we expand our study to the commercial processor architecture, which is Arm Cortex-A53, with the new set of 3-D integration options. In addition, we show the microbump assignment methodology to handle a large number of 3-D interconnections in the microbumping 3-D design. We also perform SI on the new set of 3-D intertier/interdie connections to discuss the reliability based on their physical dimensions. With a new benchmark design, the hybrid-bonding 3-D shows the best energy–delay-product (EDP) improvement, which is 25.8% compared to 2-D, and the largest eye-opening among 3-D integration options. Lingjun Zhu, Hakki Mert Torun, Madhavan Swaminathan, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | Glass Interposer Integration of Logic and Memory Chiplets: PPA and Power/Signal Integrity BenefitsabstractGlass interposers enable 3D stacking between the chiplets embedded into the substrate and the ones stacked directly on top, which is not possible in silicon. In this work, we demonstrate the benefits of such stacking in glass interposers over silicon in terms of key system-level metrics including area, wirelength, signal, power, and thermal integrity. We achieve this goal with GDS layouts of both chiplets and interposers and sign-off simulations. Our experiments show that glass offers 2.6X area, 21X wirelength, 17.72% full-chip power, 64.7% signal integrity, and 10X power integrity improvement over silicon at the cost of 15% increase in temperature. Pruek Vanna-Iampikul, Lingjun Zhu, Serhat Erdogan, Mohanalingam Kathaperumal, Ravi Agarwal, Ram Gupta, Kevin Rinebold, Sung Kyu Lim |
DAC | 2 |
| 2023 | INVITED: Design Automation Needs for Monolithic 3D ICs: Accomplishments and GapsabstractIn this paper, we provide an overview of design automation tools and methodology for Monolithic 3D ICs, focusing on the accomplishments in recent years and the gaps that remain to be filled. Monolithic 3D integration is an emerging technology with high 3D interconnect density and performance benefits, but it proposes new challenges for computer-aided design tools. In this paper, we first revisit the current status of design automation tools for Monolithic 3D and highlight the recent developments in tier partitioning, 3D placement and routing, inter-tier via controls, and power and thermal integrity analysis. Then, we discuss the gaps to be met for next-generation system-level heterogeneous Monolithic 3D IC design. Finally, we present our vision for the future of design automation developments for Monolithic 3D ICs. Lingjun Zhu, Sung Kyu Lim |
DAC | 1 |
| 2023 | SmartDS: Middle-Tier-centric SmartNIC Enabling Application-aware Message Split for Disaggregated Block StorageabstractThe widespread deployment of storage disaggregation in the cloud has facilitated flexible scaling and storage overprovisioning, allowing for high utilization of storage capacity and IOPS. Instead of utilizing remote storage protocols to access remote disks, a middle-tier is introduced between compute servers and storage servers in order to serve I/O requests from compute servers and provide computations such as compression and decompression. However, due to the need for a cloud to concurrently serve millions of VMs that require access to disaggregated storage, the middle-tier requires a massive number of servers to process network traffic between computing and storage nodes. For example, a major cloud company may deploy hundreds of thousands of high-end servers to provide such a service for its cloud storage, because the existing CPU-based middle-tier suffers from a severe issue of compute-intensive compression/decompression on high-throughput storage traffic. To address this issue, we introduce SmartDS, a middle-tier-centric SmartNIC that serves storage I/O requests with low latency and high throughput, while maintaining high flexibility and programmability. The key idea behind SmartDS is the application-aware message split (AAMS) mechanism, which allows for the processing of the message's header on the host CPU to achieve high flexibility, and the message's payload on the SmartDS. Experimental results demonstrate that SmartDS provides up to 4.3× more throughput than a CPU-based middle-tier and enables the linear scale-up of multiple network ports and multiple SmartNICs, thus significantly reducing cloud infrastructure costs for disaggregated block storage. Jie Zhang 0081, Hongjing Huang, Lingjun Zhu, Dazhong Rong, Yijun Hou, Mo Sun 0001, Chaojie Gu, Peng Cheng 0001, Zeke Wang |
ISCA | 3 |
| 2023 | A Comparative Study on Front-Side, Buried and Back-Side Power Rail Topologies in 3nm Technology NodeabstractThe standard cells are becoming increasingly smaller due to aggressive device down-scaling, and power rails take up a sizable portion of the available space. Buried Power Rail (BPR) and Back-Side Power (BSP) have been gaining more attention owing to their capacity to reduce the standard cell height from 6-Track in the traditional Front Side Power Rail (FS-PR) to 5-Track and 4-Track, respectively. In this paper, we provide a comprehensive comparison of power rail topologies at the device, standard cell, and full chip design level in terms of Power, Performance and Area (PPA). Our experiments show that nanosheet width scaling for BPR and BSP reduces device gate capacitance by 26% and 40%, respectively, resulting in an improvement of internal power of over 33% and 40%, respectively, at the standard cell level, and total power drop of over 24% and 30%, respectively, at the full chip level. Additionally, the floorplan can be shrunk down by 7% with BPR compared to FSPR, and even further by an additional 17% with BSP. This study also demonstrates the Back-side Power delivery network (BS-PDN) benefits in IR drop for BPR and BSP topologies. Sandra Maria Shaji, Lingjun Zhu, Jun-Sik Yoon, Sung Kyu Lim |
ISLPED | 2 |
| 2023 | Deploying User-space TCP at Cloud Scale with LUNA
Lingjun Zhu, Erci Xu, Shuguang Chen, Xingyu Liao, Zhendan Yang, Zhongqing Chen, Yijun Hou, Jiaji Zhu, Jiesheng Wu |
USENIX ATC | 1 |
| 2023 | Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the CloudabstractRequest latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline. Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016 |
IEEE/ACM Trans. Netw. | 7 |
| 2023 | On Continuing DNN Accelerator Architecture Scaling Using Tightly Coupled Compute-on-Memory 3-D ICsabstractThis work identifies the architectural and design scaling limits of 2-D flexible interconnect deep neural network (DNN) accelerators and addresses them with 3-D ICs. We demonstrate how scaling up a baseline 2-D accelerator in the$X/Y$dimension fails and how vertical stacking effectively overcomes the failure. We designed multitier accelerators that are$1.67\times $faster than the 2-D design. Using our 3-D architecture and circuit codesign methodology, we improve throughput, energy efficiency, and area efficiency by up to$5\times $,$1.2\times $, and$3.9\times $, respectively, over 2-D counterparts. The IR-drop in our 3-D designs is within 10.7% of VDD, and the temperature variation is within 12 °C. Gauthaman Murali, Aditya Iyer 0001, Lingjun Zhu, Jianming Tong, Francisco Muñoz-Martínez, Srivatsa Rangachar Srinivasa, Tanay Karnik, Tushar Krishna, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | 3D IC Tier Partitioning of Memory Macros: PPA vs. Thermal TradeoffsabstractMicro-bump and hybrid bonding technologies have enabled 3D ICs and provided remarkable performance gain, but the memory macro partitioning problem also becomes more complicated due to the limited 3D connection density. In this paper, we evaluate and quantify the impacts of various macro partitioning on the performance and temperature in commercial-grade 3D ICs. In addition, we propose a set of partitioning guidelines and a quick constraint-graph-based approach to create floorplans for logic-on-memory 3D ICs. Experimental results show that the optimized macro partitioning can help improve the performance of logic-on-memory 3D ICs by up to 15%, at the cost of 8°C temperature increase. Assuming air cooling, our simulation shows the 3D ICs are thermally sustainable with 97°C maximum temperature. Lingjun Zhu, Nesara Eranna Bethur, Yi-Chen Lu, Youngsang Cho, Yunhyeok Im, Sung Kyu Lim |
ISLPED | 1 |
| 2022 | Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005 |
NSDI | 7 |
| 2022 | From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudabstractThis paper presents the two generations of storage network stacks that reduced the average I/O latency of Alibaba Cloud's EBS service by 72% in the last five years: Luna, a user-space TCP stack that corresponds the latency of network to the speed of SSD; and Solar, a storage-oriented UDP stack that enables both storage and network hardware accelerations. Rui Miao 0001, Lingjun Zhu, Kun Qian 0021, Shujun Zhuang, Bo Li 0061, Shuguang Cheng, Binzhang Fu, Jiaji Zhu, Jiesheng Wu, Dennis Cai, Hongqiang Harry Liu |
SIGCOMM | 2 |
| 2022 | Design Automation and Test Solutions for Monolithic 3D ICsabstractMonolithic 3D (M3D) is an emerging heterogeneous integration technology that overcomes the limitations of the conventional through-silicon-via (TSV) and provides significant performance uplift and power reduction. However, the ultra-dense 3D interconnects impose significant challenges during physical design on how to best utilize them. Besides, the unique low-temperature fabrication process of M3D requires dedicated design-for-test mechanisms to verify the reliability of the chip. In this article, we provide an in-depth analysis on these design and test challenges in M3D. We also provide a comprehensive survey of the state-of-the-art solutions presented in the literature. This article encompasses all key steps on M3D physical design, including partitioning, placement, clock routing, and thermal analysis and optimization. In addition, we provide an in-depth analysis of various fault mechanisms, including M3D manufacturing defects, delay faults, and MIV (monolithic inter-tier via) faults. Our design-for-test solutions include test pattern generation for pre/post-bond testing, built-in-self-test, and test access architectures targeting M3D. Lingjun Zhu, Arjun Chaudhuri, Sanmitra Banerjee, Gauthaman Murali, Pruek Vanna-Iampikul, Krishnendu Chakrabarty, Sung Kyu Lim |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2022 | A Machine Learning-Powered Tier Partitioning Methodology for Monolithic 3-D ICsabstractTier partitioning is one of the most critical stages in monolithic 3-D (M3D) integrated circuits (ICs) implementation flows. It transforms 2-D netlists into 3-D by performing tier assignment for each design instance, which directly impacts the power, performance, and area (PPA) metrics of final 3-D full-chip designs. However, the current state-of-the-art tier partitioning approach named bin-based min-cut algorithm has fundamental flaws that lead to severe drawbacks, such as timing degradation, 3-D routing overhead, and redundant monolithic intertier vias (MIVs) insertion. To overcome these issues, in this article, we propose TP-GNN, an unsupervised graph learning-based tier partitioning framework that utilizes graph neural networks (GNNs) and advanced machine learning (ML) techniques to perform tier partitioning. The proposed framework comprehends design- and technology-related parameters properly so that it is generalizable to various netlists and technologies. In addition, it can be integrated with any style of M3D design flows that require tier assignments of standard cells. In the experiments, we validate the proposed framework on seven industrial designs with two different fashions of M3D implementation flows: 1) partitioning-first (Snap3D) and 2) partitioning-last (Shrunk2D and Compact2D) flows. We demonstrate that our framework, TP-GNN, significantly improves the 3-D quality of results (QoR) across most testing designs in a large margin compared with the bin-based min-cut tier partitioning algorithm. Specifically, in OpenPiton, an RISC-V-based multicore system, we observe 27.4%, 7.7%, and 20.3% improvements in performance, wirelength, and energy-per-cycle, respectively. Finally, we perform a case study by applying the proposed framework to a heterogeneous M3D design flow, Pin3D, on a commercial CPU design and observe that TP-GNN reaches better partitioning solutions than the existing partitioning approaches for heterogeneous 3-D ICs. Yi-Chen Lu, Sai Pentapati, Lingjun Zhu, Gauthaman Murali, Kambiz Samadi, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Micro-bumping, Hybrid Bonding, or Monolithic? A PPA Study for Heterogeneous 3D IC OptionsabstractIn this paper, we present three commercial-grade 3D IC designs based on state-of-the-art design technologies, specifically micro-bumping (3D die stacking), hybrid bonding (wafer-on-wafer bonding) and monolithic 3D IC (M3D). To highlight trade-offs present in these three designs, we perform analyses on power, performance, and area and the clock tree. We also model the tier-to-tier interconnection in each 3D IC methodology and analyze signal integrity to assess the reliability of each design. From our experiments, hybrid bonding design shows the best timing improvement of 81.4% when compared to its 2D counterpart, while micro-bumping shows the best reliability among 3D IC designs. Lingjun Zhu, Hakki Mert Torun, Madhavan Swaminathan, Sung Kyu Lim |
DAC | 2 |
| 2021 | Power Delivery and Thermal-Aware Arm-Based Multi-Tier 3D Architectureabstract3D integration is becoming a cost-effective way to incorporate more CPU cores and memory to improve the performance of computing systems. Meanwhile, due to the higher power density, power delivery and thermal issues become more significant in multi-tier 3DICs. In this paper, we explore and evaluate multiple design options for an Arm Neoverse-based 3D architecture focusing on power and thermals at 7nm process and sub-10$\mu $m pitch. Using a rapid voltage-drop and thermal analysis methodology, we model a system with a 32-core CPU layer and up to 4 layers of system-level caches, and quantity the trade-offs between performance, cost, voltage-drop, and temperature. A 3-layer configuration shows a good balance with 17% IPC gain and 17% lower cost, while incurring 15mV worse voltage drop and 8.5°C higher temperature compared with 2D. Our studies suggest that the co-optimization of system architecture, technology, and physical design is key for high-performance 3D systems. Lingjun Zhu, Tuan Ta, Rossana Liu, Rahul Mathur, Shidhartha Das, Ankit Kaul, Alejandro Rico, Doug Joseph, Brian Cline, Sung Kyu Lim |
ISLPED | 1 |
| 2021 | Physical Design Challenges and Solutions for Emerging Heterogeneous 3D Integration TechnologiesabstractThe emerging heterogeneous 3D integration technologies provide a promising solution to improve the performance of electronic systems in the post-Moore era, but the lack of design automation solutions and the challenges in physical design are hindering the applications of these technologies. In this paper, we discuss multiple types and levels of heterogeneous integration enabled by the high-density 3D technologies. We investigate each physical implementation stage from technology setup to placement and routing, identify the design challenges proposed by heterogeneous 3D integration. This paper provides a comprehensive survey on the state-of-the-art physical design methodologies to address these challenges. Lingjun Zhu, Sung Kyu Lim |
ISPD | 1 |
| 2021 | High-Performance Logic-on-Memory Monolithic 3-D IC Designs for Arm Cortex-A ProcessorsabstractMonolithic 3-D IC (M3-D) is a promising solution to improve the performance and energy-efficiency of modern processors. But, designers are faced with challenges in design tools and methodologies, especially for power and thermal verifications. We developed a new physical design flow that optimally places and routes cache modules in one tier and logic gates in the other. Our tool also builds high-quality clock and power delivery networks targeting logic-on-memory M3-D designs. Finally, we developed a sign-off analysis tool flow to evaluate power, performance, area (PPA), thermal, and voltage-drop quality for given M3-D designs. Using our complete register transfer level (RTL)-to-Graphic Design System (GDS) tool flow, we designed commercial quality 2-D and M3-D implementation of Arm Cortex-A7 and Cortex-A53 processors in a commercial 28-nm technology. Experimental results show that our 3-D processors offer 20% (A7) and 21% (A53) performance gain, compared with their 2-D commercial counterparts. The voltage-drop degradation of our 3-D Cortex-A7 and Cortex-A53 processors is less than 3% of the supply voltage, while temperature increase is 10.71 °C and 13.04 °C, respectively. Lingjun Zhu, Lennart Bamberg, Sai Pentapati, Kyungwook Chang, Francky Catthoor, Dragomir Milojevic, Manu Perumkunnil Komalan, Brian Cline, Saurabh Sinha 0001, Alberto García Ortiz, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2020 | TP-GNN: A Graph Neural Network Framework for Tier Partitioning in Monolithic 3D ICsabstract3D integration technology is one of the few options that can keep Moore's Law trajectory beyond conventional scaling. Existing 3D physical design flows fail to benefit from the full advantage that 3D integration provides. Particularly, current 3D partitioning algorithms do not comprehend technology and design-related parameters properly, which results in sub-optimal partitioning solutions. In this paper, we propose TP-GNN, an unsupervised graph-learning-based tier partitioning framework, to overcome this issue. Experimental results on 7 industrial designs demonstrate that our framework significantly improves the QoR of the state-of-the-art 3D implementation flows. Specifically, in OpenPiton, a RISC-V-based multi-core system, we observe 27.4%, 7.7% and 20.3% improvements in performance, wirelength, and energy-per-cycle respectively. Yi-Chen Lu, Sai Pentapati, Lingjun Zhu, Kambiz Samadi, Sung Kyu Lim |
DAC | 3 |
| 2020 | Macro-3D: A Physical Design Methodology for Face-to-Face-Stacked Heterogeneous 3D ICsabstractMemory-on-logic and sensor-on-logic face-to-face stacking are emerging design approaches that promise a significant increase in the performance of modern systems-on-chip at reasonable costs. In this work, a netlist-to-layout design flow for such heterogeneous 3D systems is proposed. The proposed technique overcomes the severe limitations of existing 3D physical design methodologies. A RISC-V-based multi-core system, implemented in a commercial technology, is used as a case study to evaluate the proposed design flow. The case study is performed for modern/large and small cache sizes to show the superiority of the proposed methodology for a broad set of systems. While previous 3D design flows do not show to optimize performance against 2D baseline designs for processor systems with a significant memory area occupation, the proposed flow shows a performance and power improvement by 20.4-28.2% and 3.2-3.8%, respectively. Lennart Bamberg, Alberto García Ortiz, Lingjun Zhu, Sai Pentapati, Da Eun Shim, Sung Kyu Lim |
DATE | 3 |
| 2020 | Full-Chip Electro-Thermal Coupling Extraction and Analysis for Face-to-Face Bonded 3D ICsabstractDue to the short die-to-die distance and inferior heat dissipation capability, Face-to-Face (F2F) boned 3D ICs are often considered to be vulnerable to electrical and thermal coupling. This study is the first to quantify the impacts of the electro-thermal coupling on the full-chip timing, power, and performance. We first present an implementation flow for realistic F2F 3D ICs including pad layers and power grids. Then, we propose our signal integrity analysis, parasitic extraction, and thermal analysis flows. Next, we investigate the impacts of the coupling on the delay, power, and noise of F2F 3D ICs, and provide guidelines to mitigate these effects. Our experimental results show that the inter-die electrical coupling causes up to 5.81% timing degradation and 4.00% noise increase, while the thermal coupling leads to less than 0.41% timing degradation and nearly no noise increase. The impact of the combined electro-thermal coupling on delay and noise reaches 6.07% and 4.05%, respectively. Lingjun Zhu, Kyungwook Chang, Dusan Petranovic, Saurabh Sinha 0001, Yun Seop Yu, Sung Kyu Lim |
ISPD | 1 |
| 2020 | Flow Event Telemetry on Programmable Data PlaneabstractNetwork performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic. Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001 |
SIGCOMM | 8 |