EDBT 2026 Demo / reviewers in the wild / expert
Shunchen Shi
dblp:301/7270
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-9514-129XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoCoTree: A Computation-Capable Architecture for Collective Communication in Scalable PIMabstractThe growing demand for high-bandwidth and largecapacity memory access in data-intensive workloads has driven the development and deployment of Processing-in-Memory (PIM) architectures. However, existing DIMM-based PIM systems suffer from the severe communication bottleneck between the processing elements (PEs) near the PIM banks due to their requirement on host CPU forwarding. This bottleneck limits the efficiency of collective operations and degrades scalability and performance for workloads that require inter-PE communication. To address the communication limitation, we propose CoCoTree, a computation-capable architecture for collective communication in scalable DIMM-based PIM. CoCoTree supports direct and high-throughput inter-PE communication without host intervention. CoCoTree accelerates key collective communication using novel hierarchical binary tree topology and lightweight in-network computation support. We design and implement microarchitectures for the main building blocks: Co-Leaf and Co-Node, to efficiently handle the data packing, routing, and processing in CoCoTree. Furthermore, we also introduce a packet-based communication protocol tailored to the CoCoTree architecture, which decouples control and data through a twophase configuration-computation communication mechanism to efficiently support a wide range of collective communication operations. CoCoTree effectively mitigates inter-PE communication bottlenecks, enabling scalable PIM systems capable of meeting the demands of growing data size. Experimental results show that CoCoTree achieves up to$95.6 \times$improvement for collective operations and improves end-to-end application performance by up to$10.5 \times$across various workloads over the baseline PIM, while outperforming state-of-the-art PIM communication architectures in both performance and scalability. Shunchen Shi, Qijia Yang, Fan Yang 0096, Yu Huang 0013, Youwei Zhuo, Zhichun Li, Ninghui Sun, Xueqi Li 0001 |
HPCA | 1 |
| 2026 | HOPESim: A Lightweight and Modern C++ based Accelerator Simulation Approach
Xueqi Li 0001, Ruihao Gao, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Shunchen Shi, Fan Yang 0096, Ninghui Sun |
ISCAS | 5 |
| 2025 | upGEMV: A Bandwidth-Efficient and Scalable GEMV Accelerator for PIM SystemsabstractIn the generation phase of large language models such as GPT, the matrix-vector multiplication (GEMV) operation is the most commonly used and time-consuming task, accounting for approximately 80% of the total inference time. Existing hardware acceleration solutions primarily rely on multi-core processors and FPGAs to enhance the performance of GEMV operations. However, GEMV operations are memory-bound in traditional platforms. Processing in memory (PIM) effectively reduces data transfer latency by processing within DRAM chips. Current PIM systems, such as UPMEM, face two major challenges: customized processing elements (PEs) for GEMV operations and parallelism, resulting in low utilization of PIM systems.In this paper, we propose a PIM-based, efficient, and scalable GEMV accelerator architecture called upGEMV, which includes a specially designed GEMV multiply-accumulate processing element, as well as new MAC and DMA instructions. Additionally, upGEMV employs an 8-stage pipeline design to enhance parallelism and mitigate processing stalls. Furthermore, we implement a simulation environment aligned with the actual UPMEM system. Experimental results demonstrate that the upGEMV achieves a 4.4× improvement on INT16, 41.0× on FP16, and 39.0× on BF16 in GEMV processing speed compared to the UPMEM DPU, at the cost of a small increase in area and power consumption. Shunchen Shi, Pengheng Zhang |
ISLPED | 2 |
| 2025 | upTSA: A DIMM-Based Near Data Processing Accelerator for Time Series Analysis
Shunchen Shi, Fan Yang 0096, Qijia Yang, Xiaohui Peng 0002, Xueqi Li 0001, Ninghui Sun |
NPC (1) | 1 |
| 2024 | CoPIM: A Collaborative Scheduling Framework for Commodity Processing-in-memory SystemsabstractProcessing in memory (PIM) is a promising paradigm to effectively alleviate the bottleneck of memory access for data-intensive applications. UPMEM is the first publicly-available real-world processing-in-memory (PIM) platform. However, current approaches mostly treat the CPU as the controller rather than considering the CPU-PIM system as a whole, thus failing to fully utilize the computational resources in modern CPUs. Additionally, the use of commodity PIM hardware as accelerators with independent memory addresses leads to significant data communication between the CPU and PIM, diminishing the performance benefits of reducing data movement. To address these challenges, we propose CoPIM, a scheduling framework that automates efficient workload scheduling. CoPIM introduces a fine-grained load modeling approach based on computational graphs and establishes a performance model to estimate workload performance and energy consumption. Our framework includes a scheduler that generates optimal scheduling parameters, enabling CPU-PIM collaborative processing. CoPIM enhances system parallelism and reduces data transfer between the CPU and PIM. We implement CoPIM on the UPMEM PIM system and demonstrate through experiments that our design achieves 4.0 x performance improvement on average when compared to general-purpose processors while reducing energy consumption by 56.6%. Shunchen Shi, Xueqi Li 0001, Zhaowu Pan, Peiheng Zhang, Ninghui Sun |
ICCD | 1 |
| 2023 | TransCaller: An End-to-end Accelerated Transformer-based Nanopore Basecaller on GPUsabstractBasecalling is a crucial step in nanopore sequencing as it transforms the raw electrical signal obtained from the nanopore into a readable sequence. The accuracy of basecalling directly affects the quality and reliability of the sequenced data. Various algorithms and models are employed to perform accurate basecalling, with advancements in deep learning techniques. However, due to the difficulty of combining feature extraction capability, high parallelism and long sequence modeling capability of existing models, the accuracy and speed of basecalling model are still bottlenecks in data analysis.In this paper, we present TransCaller, an end-to-end accelerated transformer-based nanopore basecaller model. TransCaller comprises a low-parameter electrical signal sampler, a fully transformer-based encoder integrating multiple algorithm-specific modules, and a CTC decoder enhanced with a filter. To further refine and enhance the TransCaller model, we introduce a three-stage pyramid model structure and leverage knowledge distillation techniques to balance the accuracy and speed. Experimental results on 9 datasets demonstrate that TransCaller achieves an average accuracy of 92.95%, surpassing the state-of-the-art Bonito sup model by a margin of 0.65% to 2.34%. Moreover, the distilled TransCallermodel showcases a remarkable 2.97× acceleration in inference speed compared to the original model, effectively catering to the variegated requirements of distinct inference speeds and precision levels across a wide spectrum of scenarios. Peiheng Zhang, Shunchen Shi, Xueqi Li 0001 |
BIBM | 4 |
| 2023 | Enhance the Strong Scaling of LAMMPS on FugakuabstractPhysical phenomenon such as protein folding requires simulation up to microseconds of physical time, which directly corresponds to the strong scaling of molecular dynamics(MD) on modern supercomputers. In this paper, we present a highly scalable implementation of the state-of-the-art MD code LAMMPS on Fugaku by exploiting the 6D mesh/torus topology of the TofuD network. Based on our detailed analysis of the MD communication pattern, we first adapt coarse-grained peer-to-peer ghost-region communication with uTofu interface, then further improve the scalability via fine-grained thread pool. Finally, Remote direct memory access (RDMA) primitives are utilized to avoid buffer overhead. Numerical results show that our optimized code can reduce 77% of the communication time, improving the performance of baseline LAMMPS by a factor of 2.9x and 2.2x for Lennard-Jones and embedded-atom method potentials when scaling to 36, 846 computing nodes. Our optimization techniques can also benefit other applications with stencil or domain decomposition methods. Zhuoqiang Guo, Shunchen Shi, Guangming Tan, Weile Jia, Guojun Yuan, Zhan Wang 0003 |
SC | 4 |