Wending Zhao

dblp:342/4384 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0009-1732-6801ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepPiC: xPU-PIM Cluster Architecture with Adaptive Resource-Aware Task Orchestration for DeepSeek-Style MoE Inference
abstract
The success of DeepSeek has driven demand for deploying high-performance inference clusters. However, due to its Transformer-based autoregressive structure, DeepSeek remains severely bandwidth-bound, limiting the scalability of traditional xPU (e.g., GPU/TPU). While DRAM-based processing-inmemory (PIM) offers a promising solution to overcome memory bottlenecks, its use in inference clusters for DeepSeek remains underexplored due to three challenges: (1) non-trivial inter-device communication overhead; (2) the need for expert parallelism in the mixture-of-experts (MoE) module; and (3) lack of efficient task offloading to PIM. To this end, we propose DeepPiC, a novel xPU-PIM cluster architecture designed for DeepSeek-style models with multi-latent attention (MLA) and MoE modules. DeepPiC introduces a heterogeneous xPU+HBM-PIM device to accelerate low arithmetic intensity operations. It can seamlessly replace conventional xPU devices without any modification to clusterlevel interconnect topology. However, DeepPiC cannot fully realize its performance potential under static scheduling, which fails to adapt to shifting compute and memory demands driven by multidimensional variability (model heterogeneity, cluster-scale volatility, runtime dynamics). This induces inter-device communication overhead and intra-device underutilization. Thus, we propose Adaptive Resource-Aware Task Orchestration (ARTO), a two-phase strategy that decouples global model partitioning from local task assignment by dynamically coordinating (1) crossdevice parallelism optimization and (2) intra-device xPU/PIM mapping. Evaluated on DeepSeek V3-671B using H20-, A100-, and $\mathbf{H 2 0 0}$-Cluster ($\mathbf{H 2 0}$ serves as a compute-limited alternative to high-end GPUs), DeepPiC (H20+HBM-PIM) achieves up to $\mathbf{3} \times \mathbf{, 2} \times$ and $\mathbf{1. 3} \times$ speedup over $\mathbf{H 2 0}$-, A100-, and $\mathbf{H 2 0 0}$-Cluster at small batch sizes, while maintaining $\mathbf{7 4 \%}$ and $\mathbf{5 4 \%}$ of A100and $\mathbf{H 2 0 0}$-Cluster performance at large batch sizes. These results demonstrate that DeepPiC enables low-end xPU to approach or even exceed premium ones by fundamentally overcoming memory bottlenecks via adaptive scheduling that orchestrates PIM and xPU heterogeneous resources.
Manni Li, Zijian Huang 0017, Wending Zhao, Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong
ASP-DAC5
2026 ATSGRU: Attention-Sparse Gated Recurrent Unit for Computationally Efficient Wideband Digital Predistortion of Quadrature Digital Power Amplifiers
abstract
Digital predistortion (DPD) is a widely used technique for enhancing signal quality in modern radio frequency (RF) power amplifiers (PAs). However, the strong performance of deep neural network (DNN)-based DPD models is often offset by their prohibitive computational complexity, which limits their practical deployment in wideband systems. This paper presents an attention-sparse gated recurrent unit (ATSGRU)—a novel neural architecture designed for computationally efficient wideband DPD in quadrature digital PAs (DPAs). The ATSGRU integrates the attention mechanism that evaluates the temporal relevance of input features and prunes redundant components, thereby simplifying the model structure and reducing computational load. The proposed method is validated on a custom 28-nm CMOS DPA chip. Experimental results demonstrate that the proposed ATSGRU achieves superior linearization with a favorable balance between accuracy and complexity compared with the state-of-the-art (SOTA) DPD model, reducing multiply-accumulate (MAC) operations by 54% while maintaining comparable performance. These results highlight its strong potential for efficient and scalable wideband DPD applications.
Wending Zhao, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
ACM Great Lakes Symposium on VLSI2
2026 ICDL: Inverse Compensation Direct Learning with Model-Accelerator Co-Optimization enabling Real-time Inference in Precision Motion Control
Manni Li, Longbin Jiang, Zijian Huang 0017, Wending Zhao, Yinyin Lin
ISCAS6
2025 Digital Predistortion for Quadrature Digital Power Amplifiers Using Deep Neural Network of AT_LSTM: Attention LSTM
Wending Zhao, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu
ACM Great Lakes Symposium on VLSI2
2025 3DWSNet: A Novel 3D Wavelet Spiking Neural Network for Event-based Action Recognition
abstract
In robotics applications, event cameras provide low-latency and high-dynamic-range sensing by asynchronously detecting brightness changes, making them well-suited for capturing fast motions and subtle cues in dynamic environments. However, most existing Spiking Neural Network (SNN)-based methods enhance spatial information by stacking multiple frames of events, while neglecting the explicit modeling of high-and low-frequency components in the event stream. To address this limitation, we proposes a 3D Wavelet Spiking Neural Network (3DWSNet), which integrates a 3D wavelet transform with a cascaded Wavelet Spiking Convolution (WSC) module as its core. Specifically, the 3D wavelet transform decomposes input data into eight frequency sub-bands across spatial and temporal dimensions, enabling the model to preserve fine-grained high-frequency details while enriching low-frequency motion representations. The cascaded WSC architecture further improves the extraction of multi-scale spatio-temporal features by integrating information from feature maps at different resolutions. Extensive experiments show that our 3DWSNet significantly outperforms SOTA SNN performances on the CIFAR-10, CIFAR-100, DVS128 Gesture, and CIFAR10-DVS datasets. The source code will be publicly released soon.
Junkang Fang, Yonghao Dang, Wending Zhao, Jianqin Yin
IROS3
2025 APCPU: Adaptive-Pooling Compression Processing Unit for Energy-Efficient DNNs Processing
abstract
Integrating compression in the multiply-and-accumulate (MAC) path can significantly improve the energy efficiency of DNN operators. However, existing unstructured sparse compression (USSC) methods struggle to effectively compress activations with low sparsity. Computing core processing USSC face challenges such as load imbalance and complex index control circuit design. Based on insights into local spatial correlation, a block-wise adaptive-pooling compression (APC) method is proposed to achieve a high compression ratio for activations. Furthermore, this paper proposes an APCPU to integrate APC into the MAC path with minimal overhead, facilitating highly energy-efficient sparse processing of DNN operators. Leveraging a hybrid data flow design to achieve load balancing results in speedups of 1.25× to 1.33×. The experiment results show that the APCPU achieves energy savings of 1.35× and 1.27× compared to JPZ-PU, and 2.63× and 2.71× compared to CSC-PU when evaluated on AlexNet and Bert.
Wang Wang, Wending Zhao, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong
ISCAS2