EDBT 2026 Demo / reviewers in the wild / expert
Chengchen Wang
dblp:216/2640
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0007-4929-9102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeepPiC: xPU-PIM Cluster Architecture with Adaptive Resource-Aware Task Orchestration for DeepSeek-Style MoE InferenceabstractThe success of DeepSeek has driven demand for deploying high-performance inference clusters. However, due to its Transformer-based autoregressive structure, DeepSeek remains severely bandwidth-bound, limiting the scalability of traditional xPU (e.g., GPU/TPU). While DRAM-based processing-inmemory (PIM) offers a promising solution to overcome memory bottlenecks, its use in inference clusters for DeepSeek remains underexplored due to three challenges: (1) non-trivial inter-device communication overhead; (2) the need for expert parallelism in the mixture-of-experts (MoE) module; and (3) lack of efficient task offloading to PIM. To this end, we propose DeepPiC, a novel xPU-PIM cluster architecture designed for DeepSeek-style models with multi-latent attention (MLA) and MoE modules. DeepPiC introduces a heterogeneous xPU+HBM-PIM device to accelerate low arithmetic intensity operations. It can seamlessly replace conventional xPU devices without any modification to clusterlevel interconnect topology. However, DeepPiC cannot fully realize its performance potential under static scheduling, which fails to adapt to shifting compute and memory demands driven by multidimensional variability (model heterogeneity, cluster-scale volatility, runtime dynamics). This induces inter-device communication overhead and intra-device underutilization. Thus, we propose Adaptive Resource-Aware Task Orchestration (ARTO), a two-phase strategy that decouples global model partitioning from local task assignment by dynamically coordinating (1) crossdevice parallelism optimization and (2) intra-device xPU/PIM mapping. Evaluated on DeepSeek V3-671B using H20-, A100-, and $\mathbf{H 2 0 0}$-Cluster ($\mathbf{H 2 0}$ serves as a compute-limited alternative to high-end GPUs), DeepPiC (H20+HBM-PIM) achieves up to $\mathbf{3} \times \mathbf{, 2} \times$ and $\mathbf{1. 3} \times$ speedup over $\mathbf{H 2 0}$-, A100-, and $\mathbf{H 2 0 0}$-Cluster at small batch sizes, while maintaining $\mathbf{7 4 \%}$ and $\mathbf{5 4 \%}$ of A100and $\mathbf{H 2 0 0}$-Cluster performance at large batch sizes. These results demonstrate that DeepPiC enables low-end xPU to approach or even exceed premium ones by fundamentally overcoming memory bottlenecks via adaptive scheduling that orchestrates PIM and xPU heterogeneous resources. Manni Li, Zijian Huang 0017, Wending Zhao, Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong |
ASP-DAC | 7 |
| 2026 | MPiCO: Memory-Pool-Based XPU-PIM Cluster over Optical I/O with Load-Imbalance-Aware Assignment and Execution-Site-Matching Mapping Strategies for MoE InferenceabstractWe first propose MPiCO, a memory-pool-based XPU–PIM cluster over Optical I/O, together with Load-Imbalance-Aware Assignment (LIAA) and Execution-Site-Matching Mapping (ESMM) strategies. Confining processing-in-memory (PIM) to a small set of HBMs in a hybrid HBM–DDR pool, MPiCO cuts PIM cost and offsets the resulting performance loss by eliminating inter-XPU communication overhead. LIAA resolves MoE load imbalance via dynamic assignment of warm experts to XPU/PIM, and ESMM avoids PIM-induced bandwidth loss by aligning address mapping: interleaved for XPU, PIM-friendly mapping dedicated to PIM-dies. On DeepSeek-V3 671B, MPiCO with LIAA and ESMM achieves a 2.4 × speedup and 3.5 × higher energy efficiency over H20-Electric I/O (EIO) cluster, 3 × lower PIM cost than H20-EIO with local PIM, and a 1.8 × speedup over a state-of-the-art MoE platform. Yinyin Lin, Chengchen Wang, Haidong Tian, Xiankui Xiong |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | CPSnB: Compressing and Processing Spatial Similarity near Memory Bank for DNNsabstractNear memory bank processing (NMBP) architecture only benefits memory-bound operations of DNNs in terms of energy consumption. Drawing on the insight that data compression can reduce the compute density of operators, transforming compute-bound operations into memory-bound operations, We propose CPSnB, a NMBP architecture combined with preserving numerical jump-spatial similarity compression (PNJ-SSC) method. CPSnB provides a tiling strategy for optimizing operators of different DNN models. Compared to the systolic host-side accelerator and existing dense and sparse NMBP, CPSnB significantly reduces energy consumption. Analysis of the experimental results indicates that a 60% compression ratio of activation can enhance the versatility of CPSnB in processing DNN operators to 22.3 times. Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
ISCAS | 8 |
| 2025 | APCPU: Adaptive-Pooling Compression Processing Unit for Energy-Efficient DNNs ProcessingabstractIntegrating compression in the multiply-and-accumulate (MAC) path can significantly improve the energy efficiency of DNN operators. However, existing unstructured sparse compression (USSC) methods struggle to effectively compress activations with low sparsity. Computing core processing USSC face challenges such as load imbalance and complex index control circuit design. Based on insights into local spatial correlation, a block-wise adaptive-pooling compression (APC) method is proposed to achieve a high compression ratio for activations. Furthermore, this paper proposes an APCPU to integrate APC into the MAC path with minimal overhead, facilitating highly energy-efficient sparse processing of DNN operators. Leveraging a hybrid data flow design to achieve load balancing results in speedups of 1.25× to 1.33×. The experiment results show that the APCPU achieves energy savings of 1.35× and 1.27× compared to JPZ-PU, and 2.63× and 2.71× compared to CSC-PU when evaluated on AlexNet and Bert. Wang Wang, Wending Zhao, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
ISCAS | 8 |
| 2025 | GPOS: A General and Precise Offloading Strategy for High Generality of DNN Acceleration by OCP and NDP Co-OptimizingabstractThe arithmetic intensity (ArI) of different DNNs can be opposite. This challenges the generality of single acceleration architectures, including both dedicated on-chip processing (OCP) and near-data processing (NDP). Neither architecture can simultaneously achieve optimal energy efficiency and performance for operators with opposite ArI. It is relatively straightforward to think of combining the respective advantages of OCP and NDP. However, few publications have addressed their real-time co-optimization, primarily due to the lack of a quantifiable offloading method. Here, we propose GPOS, a general and precise offloading strategy that supports high generality of DNN acceleration. GPOS comprehensively considers the complex interactions between OCP and NDP, including hardware configurations, dataflow (DF), DNN model, and interdie data movements (DMs). Three quantifiable indicators—ArI, execution cost (Ex-cost), and DM-cost—are employed to precisely evaluate the impacts of these interactions on energy and latency. GPOS adopts a four-step flow with progressive refinement: each of the first three steps focuses on a single indicator at the operator level, while the final step performs context-based calibration to address operator interdependencies and avoid offsetting NDP benefits. Narrowing down offloading candidates in step 1 and step 3 significantly accelerates real-time quantitative analysis. Optimized mapping techniques and NDP-input stationary DF are proposed to reduce Ex-cost and extend operator types supported by NDP. Next, for the first time, sparsity—one of the most popular methods for energy optimization that can alter data reuse or ArI—is quantitatively investigated for its impacts on offloading using GPOS. Our evaluations include representative DNNs, including GPT-2, Bert, RNN, CNN, and MLP. GPOS achieves the minimum energy and latency for each benchmark, with geometric mean speedups of 49.0% and 94.1%, and geometric mean energy savings of 45.8% and 89.2% over All-OCP and All-NDP, respectively. GPOS also reduces offloading analysis latency by a geometric mean of 92.7% compared to the evaluation that traverses each operator and its relative combinations. On average, sparsity further improves performance and energy efficiency by increasing the number of operators offloaded to NDP. However, for DNNs where all operators exhibit either very high or very low ArI, the number of offloaded operators remains unchanged, even after sparsity is applied. Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Chengchen Wang, Xiankui Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2024 | ARCTIC: Agile and Robust Compute-In-Memory Compiler with Parameterized INT/FP Precision and Built-In Self TestabstractDigital Compute-in-Memory (DCIM) architectures are playing an increasingly vital role in artificial intelligence (AI) applications due to their significant energy efficiency enhancement. Coupling memory and computing logic in DCIM requires extensive customization of custom cells and layouts, thus increasing design complexity and implementation effort. To adapt to the swiftly evolving AI algorithms, DCIM compiler for agile customization is required. Previous DCIM compilers accelerate the customization process but only focus on integer computation. Moreover, with technology node scaling down, design-for-test circuits are critical for robust chip design, while previous built-in-self-test (BIST) schemes for traditional memory fail to offer support for DCIM. This paper presents ARCTIC, an agile and robust DCIM compiler supporting parameterized integer/floating-point formats with corresponding BIST circuits. To support variable precision formats (including integer and floating-point), ARCTIC applies adaptive topology and layout optimization schemes for optimal performance. The compiler is also equipped with DCIM-friendly MarchCIM BIST circuits for efficient post-silicon tests with negligible area overhead. The energy efficiency of the generated DCIM macros remains competent with the state-of-the-art counterparts. Haozhe Zhu, Siqi He, Chengchen Wang, Xiankui Xiong, Haidong Tian, Xiaoyang Zeng, Chixiao Chen |
DATE | 5 |
| 2024 | LauWS: Local Adaptive Unstructured Weight Sparsity of Load Balance for DNN in Near-Data ProcessingabstractMemory wall issue has become the overwhelming bottleneck of future systems due to the explosive parameter growth and low computing density large language model (LLM). Near-data processing (NDP) could alleviate data traffic and energy consumption, but the storage demand of LLM is still enormous. Weight sparsity is helpful for reducing data capacity. Unstructured sparsity sacrifices less accuracy compared to structured one, but the random non-zero values distribution in NDP leads to load imbalance among parallel processing units. Here we propose LauWS which is seamlessly combined into various prior arts of sparsity. LauWS follows the local characteristics of feature distribution in weight matrix for various models, preserving even tiny features and discarding non-feature values as far as possible region by region. That is the key for LauWS achieving a trade-off between high prune ratio (PR) and less accuracy loss (AL). Evaluations are carried out based on a GDDR6-based bank-NDP system. The typical optimization compared to the no-prune includes 38% speedup at 0.8PR with no AL for MLP, 22.7% speedup at 0.5PR with no AL for GPT-2, 23.6% speedup at 0.5PR with the lowest perplexity for OPT-125m. Wang Wang, Manni Li, Yinyin Lin, Guhyun Kim, Yosub Song, Chengchen Wang, Xiankui Xiong |
ISCAS | 9 |
| 2021 | Mixed Spatio-Temporal Neural Networks on Real-time Prediction of CrimesabstractForecasting the crime rate in real-time is always an important task to public safety. However, there are no known models that provide satisfactory approximation to this complex spatio-temporal problem until recently. The crime rate may be affected by various factors, such as local education, public events, weather, etc. Such factors make the prediction of crimes more complex and challenging than other problems that are less influenced by outer factors. In this paper, we propose a deep-learning-based approach, which combines various methods in neural networks to handle the spatial temporal prediction problem. Some optimization techniques, such as Bayesian optimization, are applied for finding the optimal hyper-parameters as well as dealing with noises in the dataset. The model is trained on a dataset about crime information in Los Angeles at a scale of hours in block-divided areas, released by the LA Police Department (LAPD). The results of experiments on this dataset demonstrates the proposed model’s ability in predicting potential crimes in real time. Xiao Zhou 0016, Xiao Wang 0028, Gavin Brown 0003, Chengchen Wang, Sang (Peter) Chin |
ICMLA | 4 |