EDBT 2026 Demo / reviewers in the wild / expert
Dunbo Zhang
dblp:294/3050
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM ComputingabstractThe growing demand for high-performance, energy-efficient execution of data-parallel workloads has driven the resurgence of vector processors, yet their expanding instruction sets exacerbate the area-performance tradeoff of vector processing units (VPUs). Processing-in-memory (PIM) technique offers a promising path to mitigate this tradeoff by offloading vector operations near data. However, integrating the computing-capable SRAM (C-SRAM) with conventional VPUs introduces significant architectural challenges, including inefficient coordination between heterogeneous devices, the lack of a unified hardware/software interface, and underutilized parallelism within the C-SRAM arrays due to unoptimized data handling. To address these challenges, this article proposes a heterogeneous VPU (HVPU) that seamlessly integrates a standard VPU inside a vector processor with a C-SRAM for more efficient vector processing. HVPU introduces a standardized interface between the processor frontend and the C-SRAM, enabling instruction dispatch, dynamic hazard resolution, and concurrent execution. Furthermore, it employs multiple independent PIM blocks coupled with dual controlling pipelines inside the C-SRAM for further performance improvement. This design facilitates pipelined data loading and computing, effectively hiding memory latency and fully unlocking the parallel potential of C-SRAM. The system is supported by a user-friendly and generic programming model featuring a two-layer extended ISA system and a vector batch pipelining mechanism. Experimental results show that our HVPU-enhanced processor achieves significant speedups of 5.11× to 39.0× over the Xuantie-910 baseline on vector benchmarks, respectively, while reducing energy consumption by 75% on average. Meanwhile, it outperforms state-of-the-art PIM accelerators by 1.33× to 1.97× on the same benchmarks, with minimal area overhead. This demonstrates that our architectural co-design effectively alleviates the area-performance tradeoff in vector processors and offers a scalable heterogeneous architecture template for efficient vector processing. Dunbo Zhang, Shangshang Yao, Qingjie Lang, Junyi Zhu 0016, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Combination of Storage and Accumulation for Synchronous SpMV Acceleration on FPGAs with HBMabstractSparse matrix-vector multiplication (SpMV) is a crucial computational operation in various fields, such as graph computation, machine learning, and molecular dynamics.However, due to the irregular data distribution and low density of non-zero elements, the performance of SpMV is typically inferior to that of dense matrix computations.To tackle this issue, numerous optimization efforts have been made on FPGAs equipped with high-bandwidth memory (HBM), addressing problems like excessive transmission latency and load imbalance between channels.Nevertheless, several key challenges remain that hinder the performance of FPGA-based SpMV accelerators, including (1) the tendency for cache blocks to frequently miss and be replaced due to the irregular distribution of non-zero elements in sparse matrices; (2) control divergence issues within the single instruction multiple data (SIMD) SpMV accelerator architecture.To overcome these challenges, this study introduces CoSpMV, an accelerator design for FPGAs with HBM.It incorporates (1) a matrix block synchronous processing technique under ping-pong buffering, (2) a more efficient data compression format known as row-column compressed coordinate format (R3Coo), and (3) a storage accumulation module.R3Coo enhances data transmission efficiency by compressing bit width and improving the efficiency of flag bits; the matrix block synchronous processing technique under ping-pong buffering conceals replacement overhead and prevents irregular cache access by dividing and synchronously processing matrix data into rows and columns; the storage accumulation module eliminates inconsistent control flow by decoupling the addition DongHuan Xie, Qingjie Lang, Dunbo Zhang, Junsheng Chang, Li Shen 0007 |
CF | 4 |
| 2025 | In-SRAM Parallel Data ShuffleabstractWhile Single Instruction Multiple Data (SIMD) units are widely employed in processors for neural networks, signal processing, and high-performance computing, they suffer from expensive shuffle operations dedicated to data alignment. In fact, shuffle operations only change the layout of data and ideally should be done entirely within memory. To this end, we propose Shuffle SRAM in this article, which can shuffle multiple data elements simultaneously across SRAM banks. The key idea is exploiting inter-bank word line wise data movement to shuffle data in parallel, where all data elements on the same word line of SRAM can be shuffled simultaneously, achieving a high level of parallelism. Through suitable data layout preparation and proper control, Shuffle SRAM efficiently supports a wide range of commonly used shuffle operations. Our evaluation results show that the Shuffle SRAM can reap performance benefits of 14.3× for data reorganization only applications and 1.97× for data reorganization + computation applications over conventional shuffle architecture on general-purpose processors. With Shuffle SRAM, the state-of-the-art vector processor can obtain 2.58× energy efficiency. Compared with traditional SRAM, Shuffle SRAM only increases 3.5% additional area overhead. Chaoyang Jia, Dunbo Zhang, Qingjie Lang, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | Eliminate Data Divergence in SpMV via Processor and Memory Co-Computing FrameworkabstractSparse matrix-vector multiplication (SpMV) is a performance-critical kernel in various application domains, including high-performance computing, artificial intelligence, and big data. However, the performance of SpMV on SIMD devices is greatly affected by data divergences. To address this issue, we propose an In-SRAM Computing-based Processor Memory Co-Compute SpMV optimization framework that divides the SpMV kernel into two stages: a compute-intensive stage and a control-intensive stage. For optimizing the first stage, we leverage the parallel random access feature of multi-bank SRAM to eliminate overheads caused by memory divergences and use the Aggregate Table (AT) to reduce bank conflicts. For optimizing the second stage, we convert control divergences into memory divergences and utilize the Accumulate ScratchPad Memory (AccSPM) for executing reduction operations while eliminating overheads caused by memory divergences. Experimental results demonstrate that our solution achieves significant throughput increase over highly optimized vector SpMV kernels under CSR, CSR5, and CVR compression formats with performance speedups up to 4.74x, 5.58x, and 4.83x (3.11x, 3.04x, and 3.07x on average), respectively. Dunbo Zhang, Li Shen 0007, Kai Lu 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Extension VM: Interleaved Data Layout in Vector MemoryabstractWhile vector architecture is widely employed in processors for neural networks, signal processing, and high-performance computing; however, its performance is limited by inefficient column-major memory access. The column-major access limitation originates from the unsuitable mapping of multidimensional data structures to two-dimensional vector memory spaces. In addition, the traditional data layout mapping method creates an irreconcilable conflict between row- and column-major accesses. Ideally, both row- and column-major accesses can take advantage of the bank parallelism of vector memory. To this end, we propose the Interleaved Data Layout (IDL) method in vector memory, which can distribute vector elements into different banks regardless of whether they are in the row- or column-major category, so that any vector memory access can benefit from bank parallelism. Additionally, we propose an Extension Vector Memory (EVM) architecture to achieve IDL in vector memory. EVM can support two data layout methods and vector memory access modes simultaneously. The key idea is to continuously distribute the data that needs to be accessed from the main memory to different banks during the loading period. Thus, EVM can provide a larger spatial locality level through careful programming and the extension ISA support. The experimental results showed a 1.43-fold improvement of state-of-the-art vector processors by the proposed architecture, with an area cost of only 1.73%. Furthermore, the energy consumption was reduced by 50.1%. Dunbo Zhang, Qingjie Lang, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | Compressed page walk cache
Dunbo Zhang, Chaoyang Jia, Li Shen 0007 |
Frontiers Comput. Sci. | 1 |
| 2021 | Multi-level PWB and PWC for Reducing TLB Miss Overheads on GPUs
Dunbo Zhang, Chaoyang Jia, Qiong Wang 0001, Li Shen 0007 |
ICA3PP (2) | 2 |
| 2021 | A Multi-precision Quantized Super-Resolution Model Framework
Dunbo Zhang, Qiong Wang 0001, Li Shen 0007 |
ICA3PP (1) | 2 |
| 2020 | A Multi-model Super-Resolution Training and Reconstruction Framework
Ninghui Yuan, Dunbo Zhang, Qiong Wang 0001, Li Shen 0007 |
NPC | 2 |