Libo Shen

dblp:236/9252 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0000-1717-2093ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 DCLOG: Don't Cares-based Logic Optimization using Pre-training Graph Neural Networks
abstract
Logic rewriting serves as a robust optimization technique that enhances Boolean networks by substituting small segments with more effective implementations. The incorporation of don’t cares in this process often yields superior optimization results. Nevertheless, the calculation of don’t cares within a Boolean network can be resourceintensive. Therefore, it is crucial to develop effective strategies that mitigate the computational costs associated with don’t cares while simultaneously facilitating the exploration of improved optimization outcomes. To address these challenges, this paper proposes DCLOG, a don’t cares-based logic optimization framework, to efficiently and effectively optimize a given Boolean network. DCLOG leverages a pretrained graph neural network model to filter out cuts without don’t cares and then performs an incremental window simulation to calculate don’t cares for each cut. Experimental results demonstrate the effectiveness and efficiency of DCLOG on large Boolean networks, specifically average size reductions of 15.64 % and 1.44 % while requiring less than 23.84 % and $44.70 \%$ of the average runtime compared with state-of-the-art methods for the majority-inverter graph (MIG), respectively.
Rongliang Fu, Libo Shen, Ziyi Wang 0010, Zhengxing Lei, Zixiao Wang 0001, Junying Huang, Bei Yu 0001, Tsung-Yi Ho
ASP-DAC2
2026 Partitioning-free 3D-IC Floorplanning
abstract
3D integration with fine-pitch hybrid bonding offers a promising path to alleviate interconnect bottlenecks in conventional two-dimensional (2D) ICs, yet efficient 3D floorplanning remains challenging due to the enlarged solution space and non-uniform inter-die communication latency. Existing methods either extend 2D representations into 3D, leading to combinatorial complexity, or adopt partitioning-first pipelines that fix block-to-die assignments early and hinder joint optimization of floorplan, die assignment, and vertical connectivity. In this work, we present \textsc{Great3D}, a partitioning-free 3D floorplanning framework that directly optimizes a native 3D floorplan. \textsc{Great3D} formulates a unified objective that couples interconnect cost with a cycles-per-instruction (CPI)-derived latency term to capture the system-level impact of face-to-face (F2F) bonding. Algorithmically, it combines an SDP-based 3D global embedding with a dynamic-programming refinement for die assignment, followed by 2D continuous refinement with practical design constraints. \textcolor{blue}{Experiments on the GSRC and ATPlace benchmark suites show that \textsc{Great3D} consistently achieves strong wirelength and CPI quality against state-of-the-art 3D floorplanners. On GSRC, it reduces total wirelength by up to about $70\%$ (and by $2.40$--$2.74\times$ on average) over competing 3D-native floorplanners, and its dynamic-programming die-assignment stage further improves CPI by $9.5$--$17.8\%$, while maintaining competitive runtime on instances of up to a few hundred blocks.}
Shuo Ren 0001, Zhen Zhuang, Rongliang Fu, Leilei Jin, Libo Shen, Bei Yu 0001, Tsung-Yi Ho
ASP-DAC5
2026 GPA: A General-Purpose In-Memory Computing Accelerator
Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Libo Shen, Boyu Long, Xueqi Li 0001, Yinhe Han 0001, Xiaoming Chen 0003
ISCAS4
2026 CD-LLM: A Heterogeneous Multi-FPGA System for Batched Decoding of 70B+ LLMs Using a Compute-Dedicated Architecture
abstract
Large Language Models (LLMs) with 70 billion or more parameters are increasingly being deployed in cloud-based Model-as-a-Service (MaaS) scenarios. To meet the demands of such deployments, MaaS providers require batched LLM decoding systems that can deliver high System Throughput (STP) while minimizing Total Cost of Ownership (TCO). However, existing FPGA-based solutions predominantly focus on small-batch or single-batch inference, which fails to meet the computational requirements of batched LLM decoding, resulting in performance gaps of up to 7.96 \(\times\) . Moreover, the low utilization of multi-head attention operations in batched decoding scenarios, e.g., only 3.72% on A100 GPUs, further constrains throughput and inflates TCO. To address these challenges, this article introduces CD-LLM , a heterogeneous multi-FPGA system designed for efficient batched decoding of LLMs with 70B+ parameters, built upon a C ompute- D edicated architecture. First, we propose a memory-aligned mixed-precision quantization engine to reduce workload. By employing importance-aware quantization, we compress Llama-3.1-70B to an effective 3.45-bit representation and achieve 72.33% bandwidth utilization through memory-aligned data packing. Second, we present a compute-dedicated FPGA architecture that maximizes peak performance by leveraging FPGA-specific resources such as DSPs, BRAMs, and LUTs. The compute-dedicated architecture enables CD-LLM to reach a peak performance of 59.90 TOPS at 600 MHz on U250 FPGA. At last, we introduce a heterogeneous master-slave multi-FPGA system to achieve higher utilization. By pipelining attention and linear layer computations across master and slave FPGAs, CD-LLM achieves utilization rates of 83.08% for linear layers and 68.30% for attention layers. CD-LLM is designed with a heterogeneous multi-FPGA architecture, with an HBM-enabled FPGA as the master accelerator and eight DDR-based FPGAs as slave accelerators. When deployed for inference on the Llama-3.1-70B model with a batch size of 256, CD-LLM achieves a throughput of 2,721.79 tokens/s. This represents a 6.11 \(\times\) improvement in STP and a 4.71 \(\times\) reduction in TCO compared to an eight-card RTX3090 GPU system. Furthermore, CD-LLM substantially outperforms the state-of-the-art eight-card FPGA accelerator FlightLLM, delivering 16.15 \(\times\) higher STP and 14.56 \(\times\) lower TCO.
Wenheng Ma, Shulin Zeng, Tengxuan Liu, Libo Shen, Ke Hong, Zhenhua Zhu 0002, Xuefei Ning, Tsung-Yi Ho, Guohao Dai 0001, Yu Wang 0002
ACM Trans. Reconfigurable Technol. Syst.5
2025 FMC-LLM: Enabling FPGAs for Efficient Batched Decoding of 70B+ LLMs with a Memory-Centric Streaming Architecture
abstract
For large language model (LLM) acceleration, FPGAs face two challenges: insufficient peak computing performance and unacceptable accuracy loss of model compression. This paper proposes FMC-LLM to enable FPGAs for efficient batched decoding of 70B+ LLMs.
Wenheng Ma, Shulin Zeng, Tengxuan Liu, Libo Shen, Jiewen Wang, Jintao Li 0002, Zhenhua Zhu 0002, Xuefei Ning, Tsung-Yi Ho, Guohao Dai 0001, Yu Wang 0002
FPGA5
2023 Meltrix: A RRAM-Based Polymorphic Architecture Enhanced by Function Synthesis
abstract
Field-programmable gate arrays (FPGAs) are popular for computational intensive applications and hardware accelerators recently. But they face limitations in memory capacity and its growth, resulting in excessive time spent on data access. The fixed capacity of embedded memory blocks also leads inflexibility and resource waste. Moreover, logic blocks in FPGAs which are insufficient for large-scale applications and fixed memory block positions both lead to high routing overhead. To address these issues, we propose a software-hardware co-designed polymorphic architecture called Meltrix. The hardware architecture, which uses RRAM arrays as the fundamental block, creates a unified fabric that can be reconfigured into logic, storage, and interconnection modes. We achieve multiple times of logic capacity compared with FPGAs' logic blocks and multi-level interconnections inside the tiles, which are used to solve the routing overhead problem in FPGAs. Moreover, the global routing complexity is further reduced by the proposed function synthesis framework, which isolates logic and memory components, synthesizes and maps them to configured tiles of Meltrix. Experiments show that, when comparing with commercial FPGAs and state-out-of-art Liquid-Silicon, Meltrix achieves 1.89-3.14× performance improvement and 2.08-4.17× power reduction in both logic-intensive and memory-intensive applications.
Boyu Long, Libo Shen, Xiaoyu Zhang 0009, Yinhe Han 0001, Xian-He Sun, Xiaoming Chen 0003
ICCAD2
2023 LIM-GEN: A Data-Guided Framework for Automated Generation of Heterogeneous Logic-in-Memory Architecture
abstract
Memristor-based logic-in-memory (LIM) is an emerging technology that enables logic operations within memory, making it a promising solution for data-intensive applications. LIM architectures have different types according to where computations are executed, with each type being suitable for specific design objectives and application domains. However, mapping applications to a single LIM mode restricts the full utilization of different LIM modes. In this paper, we propose LIM-GEN, a data-guided framework for automated generation of heterogeneous LIM architectures. To take advantages of different LIM modes, three LIM modes are combined and used as building blocks to create heterogeneous architectures. Given the data-centric nature and large design space, there is an urgent need of developing new EDA tools for synthesizing such LIM architectures. LIM-GEN includes an automatic hardware synthesis flow, which takes behavior-level descriptions as input to generate application-specific architectures and dataflows. During synthesis, data distribution, task allocation and crossbar mapping are optimized through a design space exploration process. We evaluate LIM-GEN in several data-intensive applications and compare the generated heterogeneous architectures with synthesized architectures with a single LIM mode. The experimental results demonstrate significant improvements in latency, area and power consumption, brought by the heterogeneous architectures generated by LIM-GEN.
Libo Shen, Boyu Long, Rui Liu 0045, Xiaoyu Zhang 0009, Yinhe Han 0001, Xiaoming Chen 0003
ICCAD1