EDBT 2026 Demo / reviewers in the wild / expert
Minjin Tang
dblp:259/3736
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 9 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VersaAccel: A Versatile Configurable Accelerator for Diverse Sparse-Dense Matrix OperatorsabstractMatrix operators are fundamental to various applications, particularly in deep learning. While early models relied on dense operations, techniques like pruning have introduced sparsity, leading to a mix of dense and sparse operator types. Most existing accelerators are specialized for specific operators and perform poorly in mixed scenarios, while those supporting multiple operators often lack flexibility and suffer from suboptimal performance. To overcome these limitations, we propose VersaAccel, a configurable accelerator for sparse and dense matrix operators. It supports four distinct configurations, each optimized for a set of operators. A key feature of our design is its adaptive configuration selection mechanism, driven by a lightweight cost model that explicitly evaluates the performance-energy trade-off between available options. This allows VersaAccel to dynamically choose the most efficient configuration—opting for higher performance when the gain outweighs the energy cost, or prioritizing energy efficiency when appropriate. Experimental results demonstrate that VersaAccel achieves an average performance/area improvement of 3.10× across multiple operators (MV, MM, SPMV, SpMM, SpMSpM) and 4.06× on full model evaluations (ResNet18, VGG16, LLaMA2-7B, BERT-Base), compared to mainstream accelerators. Minjin Tang, Mei Wen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level SparsityabstractTo more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs. Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
DATE | 7 |
| 2025 | SmartBlock: Adaptive Block Floating Point Quantization for Efficient DNN AccelerationabstractDeep Neural Networks (DNNs) have achieved remarkable success as model sizes continue to grow, driving the need for optimizations in both computational and energy efficiency. Block Floating Point (BFP) quantization has emerged as an effective model compression technique, offering a favorable trade-off between model accuracy and hardware cost. However, the frequent use of floating-point (FP) accumulation across BFP blocks remains a significant bottleneck, limiting further improvements in energy efficiency. State-of-the-art (SotA) accelerators mitigate this issue by introducing low-overhead accumulators with a narrower dynamic range ahead of the FP accumulator to handle a small range of values. While this approach reduces the activation of power-hungry alignment and format conversion units, it increases the complexity of the processing elements (PEs), thereby limiting the overall energy savings. Xin Ju 0005, Jingkui Yang, Mei Wen, Minjin Tang, Zhaoyun Chen, Yang Shi 0008 |
ICPP | 6 |
| 2025 | SPSA: Exploring Sparse-Packing Computation on Systolic Arrays From ScratchabstractSparse matrix-matrix multiplication (SpMM) and Generalized SpMM (SpGEMM) are essential computational kernels in domains, such as graph analytics and scientific computation. While systolic arrays have traditionally been employed as specialized architectures for complex computing problems like matrix multiplication, they exhibit inefficiency when dealing with sparse matrices. This inefficiency arises from the unnecessary operations performed by processing elements (PEs) that contain zero-valued entries, which do not contribute to the final result. To address this issue, we propose SPSA, a framework that leverages a sparse-packing algorithm suitable for systolic arrays to accelerate sparse matrix computations. Our approach achieves significant reduction of zero-valued items and improves matrix density by packing the rows or columns of the sparse matrix. Furthermore, we have introduced for the first time a data representation format tailored to systolic arrays, called CSXD, which further enhances storage and computational efficiency. Importantly, our adaptation scheme enables acceleration benefits even with limited resources. Through sparse packing, SPSA achieved a$5.2\times $performance improvement compared to the dense baseline, and further reached a$6.4\times $enhancement via CSXD. Simultaneously, CSXD realized an average storage efficiency improvement of$15.0\times $. Through extensive evaluations, SPSA outperforms previous designs on CPU, GPU, and ASIC platforms. Finally, in end-to-end evaluations, SPSA achieved a performance improvement of 3.9 times across the workloads of BERT, VGG19, and ResNet50. Minjin Tang, Mei Wen, Jianchao Yang, Zeyu Xue, Junzhong Shen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor CoresabstractSparse-Dense Matrix-Matrix Multiplication (SpMM) has emerged as a foundational primitive in HPC and AI. Recent advancements have aimed to accelerate SpMM by harnessing the powerful Tensor Cores found in modern GPUs. However, despite these efforts, existing methods frequently encounter performance degradation when ported across different Tensor Core architectures. Recognizing that scalable SpMM across multiple generations of Tensor Cores relies on the effective use of general-purpose instructions, we have meticulously developed a SpMM library named SSpMM. However, a significant conflict exists between granularity and performance in current Tensor Core instructions. To resolve this, we introduce the innovative Transpose Mapping Scheme, which elegantly implements fine-grained kernels using coarse-grained instructions. Additionally, we propose the Register Shuffle Method to further enhance performance. Finally, we introduce Sparse Vector Compression, a technique that ensures our kernels are scalable with both structured and unstructured sparsity. Our experimental results, conducted on four generations of Tensor Core GPUs using over 3,000 sparse matrices from well established matrix collections, demonstrate that SSpMM achieves an average speedup of 2.04×, 2.81×, 2.07×, and 1.87×, respectively, over the state-of-the-art SpMM solution. Furthermore, we have integrated SSpMM into PyTorch, achieving a 1.81× speedup in end-to-end Transformer inference compared to cuDNN. Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang, Zhongdi Luo, Yang Shi 0008, Zhaoyun Chen, Junzhong Shen, Johannes Langguth |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | MSA2: An Efficient Sparsity-Aware Accelerator for Matrix Multiplication with Multi-core Systolic Arrays
Minjin Tang, Mei Wen, Junzhong Shen, Jingkui Yang, Zeyu Xue, Zili Shao |
ICA3PP (3) | 1 |
| 2023 | Releasing the Potential of Tensor Core for Unstructured SpMM using Tiled-CSR FormatabstractThe GPU has become a popular platform for AI applications, thanks in part to its Tensor Cores that address performance issues. However, the Sparse Matrix Multiplication (SpMM) kernel has remained a bottleneck despite significant advances in computing power. Due to the hardware mechanism of the Tensor Core, its programming granularity does not match SpMM. In this paper, we analyze the reasons why the unstructured SpMM kernel is not suitable for the Tensor Core, and propose the Tiled Compressed Sparse Row (Tiled-CSR) compression format. To address the issue of low non-zero rates in Tiled-CSR format, we exploit the row shuffle algorithm to improve the utilization of Tensor Cores and enhance computing density. We also utilize adaptive memory access modes and 3D-Grid tiling for the SpMM kernel to reduce memory access latency. The experimental results on NVIDIA A100 GPU with matrices in the Deep Learning Matrix Collection (DLMC) demonstrate that the Tiled-CSR format improves the utilization of Tensor Cores under different sparsity, with a maximum of 3.89× at 50% sparsity and a minimum of 1.82× at 90% sparsity compared to the SR-BCRS format. Additionally, our kernel achieves an average speedup of 1.54×(up to 2.12×) over Magicube. Zeyu Xue, Mei Wen, Zhaoyun Chen, Yang Shi 0008, Minjin Tang, Jianchao Yang, Zhongdi Luo |
ICCD | 5 |
| 2022 | BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic ArraysabstractState-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%. Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang |
ICCD | 5 |
| 2022 | Mentha: Enabling Sparse-Packing Computation on Systolic ArraysabstractGeneralized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network. Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001 |
ICPP | 1 |
| 2020 | Towards Memory-Efficient Streaming Processing with Counter-Cascading Sketching on FPGAabstractObtaining item frequencies in data streams with limited space is a well-recognized and challenging problem in a wide range of applications. Sketch-based solutions have been widely used to address this challenge due to their ability to accurately record the data streams at a low memory cost. However, most sketches suffer from low memory utilization due to the adoption of a fixed counter size. Accordingly, in this work, we propose a counter-cascading scheduling algorithm to maximize the memory utilization of sketches without incurring any accuracy loss. In addition, we propose an FPGA-based system design that supports sketch parameter learning, counter-cascading record and online query. We implement our designs on Xilinx VCU118, and conduct evaluations on real-world traces, thereby demonstrating that our design can achieve higher accuracy with lower storage; the performance achieved is 10× ~ 20× better than that of state-of-the-art sketches. Minjin Tang, Mei Wen, Junzhong Shen, Chunyuan Zhang |
DAC | 1 |
| 2020 | Scalable FPGA-based Architecture for High-Performance Per-Flow Traffic MeasurementabstractPer-flow traffic measurement has emerged as a critical but challenging task in data center in recent years in the face of massive network traffic. Many approximate methods have been proposed to resolve the existing resource-accuracy trade-off in per-flow traffic measurement, one of which is the sketch-based method. However, sketches are affected by their high computational cost and low throughput; moreover, their measurement accuracy is hard to guarantee under the conditions of changing network bandwidth or flow size distribution. Recently, FPGA platforms have been widely deployed in data centers, as they demonstrate a good fit for high-speed network processing. In this work, we propose a scalable pipelined architecture for high high-throughput per-flow traffic measurement on FPGA. We adopts memory-friendly D-left hashing in our design, which guarantees high space utilization that successfully addressing the challenge of tracking high speed data stream under limit memory resource on FPGA. Comparisons with state-of-the-art sketch-based solutions show that our design outperforms state-of-the-art sketch-based methods in terms of throughput by over 80x. Junzhong Shen, Mei Wen, Minjin Tang, Chunyuan Zhang |
FPGA | 3 |
| 2020 | Optimized HybridSketch: More Efficient with Analysis and Algorithm
Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICA3PP (1) | 3 |
| 2020 | HybridSketch: A Memory-centric Precise Approach for Flow MeasurementabstractAs network bandwidth has rapidly developed, due to the high occupancy of memory and bandwidth required, the Sketch structure is favored by some researchers due to its limited memory usage and simple operation. But the accuracy will decrease when the Sketch system occupies less memory space. Traditional sketch algorithms and some other specially designed algorithms and structures are striving to improve accuracy. However, with the flow rate rapidly increasing, the on-chip memory will be the bottleneck of the system. Our network measurement system achieve good results focusing more on the memory usage. We proposes a hybrid method, HybridSketch, which focuses on the memory and precision of the system with mixing two measurement methods by quantitatively analyzing, modeling and allocating appropriate memory space to each method to achieve better results. Experimental results show that our method can provide 10× improvement in terms of precision, moreover, HybridSketch can provide the same level of precision with achieving 24× improvement in terms of memory size. Mei Wen, Minjin Tang, Qun Huang 0001, Chunyuan Zhang |
ICC | 3 |