Bo Jiao 0003

dblp:157/2546-3 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0002-2787-902XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet System
abstract
The rapid growth of model sizes in advanced artificial intelligence algorithms, particularly in Transformerbased large language models (LLMs), has led to significant computational overhead. Mixture-of-Expert (MoE) models offer a solution through their sparsely gating mechanism but introduce new challenges of extensive all-to-all communication and model computational inefficiencies. This paper presents Hydra, a software/hardware co-design aimed at accelerating MoE inference on chiplet-based architectures. In software, Hydra employs a popularity-aware expert mapping strategy to optimize interchiplet communication. In hardware, it incorporates Content Addressable Memory (CAM) to eliminate expensive explicit token (un)-permutation based on sparse matrix multiplications and a redundant-calculation-skipping softmax engine to bypass unnecessary division and exponential operations. Evaluated in 22 nm technology, Hydra achieves latency reductions of $14.2 \times$ and $3.5 \times$ and power reductions of $169.1 \times$ and $18.9 \times$ over GPU and state-of-the-art MoE accelerator, respectively, thereby offering a scalable and efficient solution for MoE model deployment.
Siqi He, Haozhe Zhu, Jiapei Zheng, Lizhou Wu, Bo Jiao 0003, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen
DAC5
2025 EIGEN: Enabling Efficient 3DIC Interconnect with Heterogeneous Dual-Layer Network-on-Active-Interposer
abstract
Chiplet-based 3DICs have emerged as modular solutions for large-scale, high-performance computing systems. However, unlike monolithic Network-on-Chips (NoCs), 3DICs using active interposers encounter difficulties in managing heterogeneous and high-load network traffic patterns. The emerging field of Networks-on-Active-Interposer (NOAI), independently designed from top dies’ interconnects, is desired to address these challenges by supporting flexible topologies, low-latency memory access, and reduced traffic congestion. To satisfy the requirements, we propose a heterogeneous dual-layer interconnect architecture, EIGEN, for chiplet-interposer systems, along with a reinforcement learning (RL)-based routing framework, which can provide efficient and flexible communication for chiplet-based 3DICs. EIGEN features an application-aware switch-programable interconnection layer (AspLayer) and a dynamic packet-routing interconnection layer (DynLayer) on the interposer, respectively. To optimize inter-chiplet data communication, we also develop an RL-based path routing framework tailored to this dual-layer architecture. The RL framework’s state space includes topology, network, and memory metrics, while the RL reward is set as the product of average packet latency and link utilization. The effectiveness of EIGEN is demonstrated through different applications where CPU, GPU, and AI accelerator chiplets are integrated via NOAI. Simulation results show that the proposed EIGEN achieves up to $\mathbf{6 7. 1 6 \%}$ latency reduction, $\mathbf{5 3. 8 9 \%}$ hops reduction and $\mathbf{1 1. 2 1 \%}$ runtime reduction compared to the state-of-the-art (SOTA) chiplets interconnect architectures. Furthermore, sensitivity analysis of network scaling shows that the EIGEN architecture and framework exhibit strong scalability, with latency reduction of $\mathbf{2 0. 9 6 \%}$ to $\mathbf{5 2. 7 7 \%}$ and hop reduction of $\mathbf{1 9. 7 2 \%}$ to $\mathbf{4 3. 4 1 \%}$ as the system scales from $4 \times 4$ to $32 \times 32$ chiplets.
Siyao Jia, Bo Jiao 0003, Haozhe Zhu, Chixiao Chen, Qi Liu 0010, Ming Liu 0022
HPCA2
2024 FPIA: Communication-Aware Multi-Chiplet Integration With Field-Programmable Interconnect Fabric on Reusable Silicon Interposer
abstract
Silicon interposer re-usage is drawing attention for cost-effective multi-chiplet integrated systems. To address the communication awareness of inter/off-chiplet interconnect, the paper proposes a field-programmable interconnect fabric and develops its corresponding automatic physical integration tool. The tile-based fabric consists of turnout, cross-over boxes and parallel tracks. It features micro-bump-wise connecting flexibility and hardware efficiency. The automation flow performs chiplet location optimization and efficient bump-to-bump routing, supporting multi-lane bus interconnect and miscellaneous external ports. The methodology is validated by 9 different integration scenarios, where the routability is guaranteed when the local resource utilization ratio approaches 94.5%. The data’s maximum interconnect latency is 2.2 ns and the energy consumption is 1.18 pJ/bit at a bitrate of 1 Gbps. The latency consumes$16.5\times \sim ~53.4\times $fewer clock cycles than the state-of-the-art network-on-package-based reusable interposer architectures.
Bo Jiao 0003, Haozhe Zhu, Jundong Zhu, Dexin Wen, Lingli Wang, Jun Tao 0001, Chixiao Chen, Yinhe Han 0001, Qi Liu 0010, Ninghui Sun, Ming Liu 0022
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D Integration
abstract
Chiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost.
Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen
ISCAS2
2022 CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in Space
abstract
Reliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to complete robust 6D pose estimation of the space-borne targets under complicated background. Specifically, conventional methods are adopted to extract the features of the whole image in the factual case. In the counterfactual case, a non-existent image without the target but only the background is imagined. Side effect caused by background interference is reduced by counterfactual analysis, which leads to unbiased prediction in final results. In addition, we also carry out low-bit-width quantization for CA-SpaceNet and deploy part of the framework to a Processing-In-Memory (PIM) accelerator on FPGA. Qualitative and quantitative results demonstrate the effectiveness and efficiency of our proposed method. To our best knowledge, this paper applies causal inference and network quantization to the 6D pose estimation of space-borne targets for the first time. The code is available at https://github.com/Shunli-Wang/CA-SpaceNet.
Shunli Wang 0001, Shuaibing Wang, Bo Jiao 0003, Dingkang Yang, Liuzhen Su, Peng Zhai, Chixiao Chen, Lihua Zhang 0002
IROS3
2021 A 0.57-GOPS/DSP Object Detection PIM Accelerator on FPGA
abstract
The paper presents an object detection accelerator featuring a processing-in-memory (PIM) architecture on FPGAs. PIM architectures are well known for their energy efficiency and avoidance of the memory wall. In the accelerator, a PIM unit is developed using BRAM and LUT based counters, which also helps to improve the DSP performance density. The overall architecture consists of 64 PIM units and three memory buffers to store inter-layer results. A shrunk and quantized Tiny-YOLO network is mapped to the PIM accelerator, where DRAM access is fully eliminated during inference. The design achieves a throughput of 201.6 GOPs at 100MHz clock rate and correspondingly, a performance density of 0.57 GOPS/DSP.
Bo Jiao 0003, Jinshan Zhang 0006, Yuanyuan Xie, Shunli Wang 0001, Haozhe Zhu, Xiaoyang Kang 0001, Zhiyan Dong, Lihua Zhang 0002, Chixiao Chen
ASP-DAC1
2021 Computing Utilization Enhancement for Chiplet-based Homogeneous Processing-in-Memory Deep Learning Processors
abstract
This paper presents a design strategy of chiplet-based processing-in-memory systems for deep neural network applications. Monolithic silicon chips are area and power limited, failing to catch the recent rapid growth of deep learning algorithms. The paper first demonstrates a straightforward layer-wise method that partitions the workload of a monolithic accelerator to a multi-chiplet pipeline. A quantitative analysis shows that the straightforward separation degrades the overall utilization of computing resources due to the reduced on-chiplet memory size, thus introducing a higher memory wall. A tile interleaving strategy is proposed to overcome such degradation. This strategy can segment one layer to different chiplets which maximizes the computing utilization. To facilitate the strategy, the modification of the chiplet system hardware is also discussed. To validate the proposed strategy, a nine-chiplet processing-in-memory system is evaluated with a custom-designed object detection network. Each chiplet can achieve a peak performance of 204.8GOPS at a 100-MHz rate. The peak performance of the overall system is 1.711TOPS, where no off-chip memory access is needed. By the tile interleaving strategy, the utilization is improved from 53.9 to 92.8
Bo Jiao 0003, Haozhe Zhu, Jinshan Zhang 0006, Shunli Wang 0001, Xiaoyang Kang 0001, Lihua Zhang 0002, Mingyu Wang 0001, Chixiao Chen
ACM Great Lakes Symposium on VLSI1
2021 ALPINE: An Agile Processing-in-Memory Macro Compilation Framework
abstract
Processing-in-Memory architectures and circuit designs are playing significant roles in the recent energy-efficient machine learning chips. This paper proposes a PIM macro compilation framework called ALPINE to speed up previously tedious and error-prone PIM design flow, paving the way towards open-source and process-portable PIM chips. Relying on an extensible PIM standard cell library, ALPINE can generate the corresponding topology according to the specification, and process placement and routing. The proposed PIM macro is compatible with different storage devices such as SRAM and RRAM, and can support various quantization bit-widths and dataflows. To verify the effectiveness, a 128×128 SRAM-based PIM macro instance is implemented, and the simulation results show that it can achieve an energy efficiency of 19.05TOPS/W under 65nm CMOS technology. The macro performance is not inferior to the state-of-the-art custom PIM designs.
Jinshan Zhang 0006, Bo Jiao 0003, Yunzhengmao Wang, Haozhe Zhu, Lihua Zhang 0002, Chixiao Chen
ACM Great Lakes Symposium on VLSI2