Haidong Tian

dblp:298/8123 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-9777-4480ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepPiC: xPU-PIM Cluster Architecture with Adaptive Resource-Aware Task Orchestration for DeepSeek-Style MoE Inference
abstract
The success of DeepSeek has driven demand for deploying high-performance inference clusters. However, due to its Transformer-based autoregressive structure, DeepSeek remains severely bandwidth-bound, limiting the scalability of traditional xPU (e.g., GPU/TPU). While DRAM-based processing-inmemory (PIM) offers a promising solution to overcome memory bottlenecks, its use in inference clusters for DeepSeek remains underexplored due to three challenges: (1) non-trivial inter-device communication overhead; (2) the need for expert parallelism in the mixture-of-experts (MoE) module; and (3) lack of efficient task offloading to PIM. To this end, we propose DeepPiC, a novel xPU-PIM cluster architecture designed for DeepSeek-style models with multi-latent attention (MLA) and MoE modules. DeepPiC introduces a heterogeneous xPU+HBM-PIM device to accelerate low arithmetic intensity operations. It can seamlessly replace conventional xPU devices without any modification to clusterlevel interconnect topology. However, DeepPiC cannot fully realize its performance potential under static scheduling, which fails to adapt to shifting compute and memory demands driven by multidimensional variability (model heterogeneity, cluster-scale volatility, runtime dynamics). This induces inter-device communication overhead and intra-device underutilization. Thus, we propose Adaptive Resource-Aware Task Orchestration (ARTO), a two-phase strategy that decouples global model partitioning from local task assignment by dynamically coordinating (1) crossdevice parallelism optimization and (2) intra-device xPU/PIM mapping. Evaluated on DeepSeek V3-671B using H20-, A100-, and $\mathbf{H 2 0 0}$-Cluster ($\mathbf{H 2 0}$ serves as a compute-limited alternative to high-end GPUs), DeepPiC (H20+HBM-PIM) achieves up to $\mathbf{3} \times \mathbf{, 2} \times$ and $\mathbf{1. 3} \times$ speedup over $\mathbf{H 2 0}$-, A100-, and $\mathbf{H 2 0 0}$-Cluster at small batch sizes, while maintaining $\mathbf{7 4 \%}$ and $\mathbf{5 4 \%}$ of A100and $\mathbf{H 2 0 0}$-Cluster performance at large batch sizes. These results demonstrate that DeepPiC enables low-end xPU to approach or even exceed premium ones by fundamentally overcoming memory bottlenecks via adaptive scheduling that orchestrates PIM and xPU heterogeneous resources.
Manni Li, Zijian Huang 0017, Wending Zhao, Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong
ASP-DAC8
2026 MPiCO: Memory-Pool-Based XPU-PIM Cluster over Optical I/O with Load-Imbalance-Aware Assignment and Execution-Site-Matching Mapping Strategies for MoE Inference
abstract
We first propose MPiCO, a memory-pool-based XPU–PIM cluster over Optical I/O, together with Load-Imbalance-Aware Assignment (LIAA) and Execution-Site-Matching Mapping (ESMM) strategies. Confining processing-in-memory (PIM) to a small set of HBMs in a hybrid HBM–DDR pool, MPiCO cuts PIM cost and offsets the resulting performance loss by eliminating inter-XPU communication overhead. LIAA resolves MoE load imbalance via dynamic assignment of warm experts to XPU/PIM, and ESMM avoids PIM-induced bandwidth loss by aligning address mapping: interleaved for XPU, PIM-friendly mapping dedicated to PIM-dies. On DeepSeek-V3 671B, MPiCO with LIAA and ESMM achieves a 2.4 × speedup and 3.5 × higher energy efficiency over H20-Electric I/O (EIO) cluster, 3 × lower PIM cost than H20-EIO with local PIM, and a 1.8 × speedup over a state-of-the-art MoE platform.
Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong
ACM Great Lakes Symposium on VLSI5
2025 LsCMM-H: A TCO-Optimized Hybrid CXL Memory Expansion Architecture with Log Structure
abstract
In the era of big data, the demand for memory capacity in modern computing systems is surging. The CXL-SSD, NAND Flash-based memory expander using emerging Compute Express Link (CXL), has become a promising solution for efficient memory expansion. However, the memory-expansion scenario poses severe performance and endurance challenges for CXL-SSDs, and existing works fail to fully address them due to the usage of traditional SSDs as back-end media. To optimize these aspects, we propose LsCMM-H, a Total-Cost-of-Ownership (TCO) -efficient CXL-SSD architecture with Zoned Namespace (ZNS) SSDs as back-end media for better latency and lifetime. LsCMM-H employs hardware-software co-designed log management, low-overhead data-tiering-based garbage collection mechanism, and read acceleration to leverage the benefits of ZNS. Based on our evaluation, LsCMM-H reduces tail latency by 49.9%, improves throughput by 41.9%, endurance by 280.5%, and saves TCO by 72.4% compared to vanilla CXL-SSD. The additional comparison also demonstrates the superiority of our proposed log structure.
Xiangrui Zhang, Sirui Peng, Zhiwang Guo, Haidong Tian, Xiankui Xiong, Xiaoyong Xue, Xiaoyang Zeng
ICCAD5
2024 ARCTIC: Agile and Robust Compute-In-Memory Compiler with Parameterized INT/FP Precision and Built-In Self Test
abstract
Digital Compute-in-Memory (DCIM) architectures are playing an increasingly vital role in artificial intelligence (AI) applications due to their significant energy efficiency enhancement. Coupling memory and computing logic in DCIM requires extensive customization of custom cells and layouts, thus increasing design complexity and implementation effort. To adapt to the swiftly evolving AI algorithms, DCIM compiler for agile customization is required. Previous DCIM compilers accelerate the customization process but only focus on integer computation. Moreover, with technology node scaling down, design-for-test circuits are critical for robust chip design, while previous built-in-self-test (BIST) schemes for traditional memory fail to offer support for DCIM. This paper presents ARCTIC, an agile and robust DCIM compiler supporting parameterized integer/floating-point formats with corresponding BIST circuits. To support variable precision formats (including integer and floating-point), ARCTIC applies adaptive topology and layout optimization schemes for optimal performance. The compiler is also equipped with DCIM-friendly MarchCIM BIST circuits for efficient post-silicon tests with negligible area overhead. The energy efficiency of the generated DCIM macros remains competent with the state-of-the-art counterparts.
Haozhe Zhu, Siqi He, Chengchen Wang, Xiankui Xiong, Haidong Tian, Xiaoyang Zeng, Chixiao Chen
DATE7