VLDB 2026 Research / reviewers in the wild / expert
Zhen Wang 0045
dblp:78/6727-45
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0003-8853-9915ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2RT: Efficient Ray Tracing Accelerator with Approximate-Accurate Computing and QuantizationabstractRay tracing (RT) has revolutionized photorealistic rendering by simulating light transport, but existing methods face a trade-off between computational efficiency and rendering accuracy. To address this, we present A2RT, a software-hardware co-designed RT accelerator employing the end to end optimization of "quantization → computation". On the software side, we introduce a customized data flow mechanism with type-specific quantization for bounding boxes, ray origins, and directions, and we organize BVH nodes into Group- and Sub-Nodes. At the hardware level, a heterogeneous RT engine allocates resources based on node criticality: accurate computing units handle Group-Nodes, while approximate units process Sub-Nodes. A custom INT-FLOAT approximate multiplier further accelerates the approximate units. Experimental results show that A2RT achieves 45.51% energy consumption and 2.29× speedup over RT Core, and consumes 81.79% of energy while delivering 1.57× performance improvement compared to state-of-the-art accelerators. Zhihua Fan, Yudong Mu, Zhen Wang 0045, Xiaochun Ye, Xuejun An |
DATE | 6 |
| 2025 | Accelerating Authenticated Block Ciphers via RISC-V Custom Cryptography InstructionsabstractAs one of the standardized encryption algorithms, authenticated block ciphers based on Galois/Counter Mode (GCM) is a widely-used method to guarantee the accuracy and reliability in data transmission. The profiling work demonstrates that across the execution process of GCM mode, the authentication operation is the main performance bottleneck because it introduces operations in high-dimensional Galois field (GF), which could not be efficiently executed via existing ISA. To overcome this problem, we propose a custom ISA extension and cooperate it with RISC-V cryptography extension to accelerate the whole process of authenticated block ciphers. Besides, we design a specific crypto core including a fully-pipelined GF(2128) multiplier to support the extended instructions and integrate it into the multi-issue out-of-order core XT910 without introducing any clock frequency overhead. The proposed design significantly reduces the the number of instructions required in the main operations of authenticated block ciphers. We compare the performance of our designs to other existing acceleration scheme based on RISC-V ISA extension. Experimental result shows that our design outperforms other related work and achieves up to 17 × speedup with a lightweight hardware overhead. Tianyu Liu 0007, Zhen Wang 0045, Zhihua Fan, Xiaochun Ye, Dongrui Fan |
DATE | 4 |
| 2025 | TSCNN: Compressing and Accelerating Sparse CNNs Using Sign-Reserved Toeplitz FiltersabstractExploiting the sparsity in convolutional neural networks (CNNs) is crucial to accelerate computing and reduce energy consumption. However, unstructured sparsity often introduces irregularity in convolutional operations, which complicates the control logic and undermines the benefits of sparsification. Structured sparsity alleviates these problems but sacrifices the adaptability to arbitrary sparse patterns. In this paper, we propose TSCNN, an algorithm-hardware co-design solution that aims to compress and accelerate sparse CNNs while balancing both adaptability to sparsity and computational efficiency. In terms of algorithm, TSCNN adopts pruned filters compressed with sign-reserved Toeplitz matrix format (Tfilters), which systematically enhances the regularity of data reuse and flexibly reduces network parameters by$44 \%-86 \%$while maintaining accuracy. In terms of hardware, TSCNN accelerator employs custom computing components to adapt to the structure of Tfilters and support the adaptive dataflow, further optimizing the computational efficiency. Experiments show that TSCNN outperforms a dense accelerator, SCNN and CSCNN, achieving$4.49 \times, 2.29 \times, 2.08 \times$and$1.29 \times$speedup and reducing energy consumption by$74.65 \%, 41.04 \%, 49.29 \%$and 43.66%, respectively. Zhen Wang 0045, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan |
HPCC | 1 |
| 2025 | Accelerating tensor multiplication by exploring hybrid product with hardware and software co-design
Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An |
J. Syst. Archit. | 5 |
| 2025 | PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow ArchitecturesabstractDataflow architectures are considered promising architecture, offering a commendable balance of performance, efficiency, and flexibility. Abundant prior works have been proposed to improve the performance of dataflow architectures. Nevertheless, these solutions can be further improved due to the lack of efficient data prefetching and flexible task scheduling. In this article, we propose a novel dataflow architecture with adaptive p refetching an d d ecentr a lized scheduling (PANDA). First, we present an application-adaptive data prefetching method and on-chip memory microarchitecture designed to overlap memory access latency. Second, we introduce a decentralized dataflow scheduling approach and processing element (PE) microarchitecture aimed at improving hardware utilization. Experimental results show that in a wide range of real-world applications, PANDA attains up to 2.53× performance improvement and 1.79× energy efficiency improvement over the state-of-the-art dataflow architectures. Shantian Qin, Zhihua Fan, Zhen Wang 0045, Xuejun An, Xiaochun Ye, Dongrui Fan |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | DFU-E: A Dataflow Architecture for Edge DSP and AI ApplicationsabstractEdge computing aims to enable swift, real-time data processing, analysis, and storage close to the data source. However, edge computing platforms are often constrained by limited processing power and efficiency. This paper presents DFU-E, a dataflow-based accelerator specifically designed to meet the demands of edge digital signal processing (DSP) and artificial intelligence (AI) applications. Our design addresses real-world requirements with three main innovations. First, to accommodate the diverse algorithms utilized at the edge, we propose a multi-layer dataflow mechanism capable of exploiting task-level, instruction block-level, instruction-level, and data-level parallelism. Second, we develop an edge dataflow architecture that includes a customized processing element (PE) array, memory, and on-chip network microarchitecture optimized for the multi-layer dataflow mechanism. Third, we design an edge dataflow software stack that enables automatic optimizations through operator fusion, dataflow graph mapping, and task scheduling. We utilize representative real-world DSP and AI applications for evaluation. Comparing with Nvidia's state-of-the-art edge computing processor, DFU-E achieves up to 1.42× geometric mean performance improvement and 1.27× energy efficiency improvement. Zhihua Fan, Tianyu Liu 0007, Zhen Wang 0045, Meng Wu 0006, Kunming Zhang, Yanhuan Liu, Ninghui Sun, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | OBSD: On-The-Fly Block-Wise Sparse Distillation Accelerating SpGEMMs in DNN ApplicationsabstractThe mainstream deep neural networks (DNNs) widely adopt pruning techniques to alleviate network overfitting and computational complexity. Simultaneously, the weight and activation data exhibit increasingly noticeable block-wise sparsity. Current DNN-specific accelerators mainly focus on element-wise while neglecting sparse features of block granularity. Meanwhile, one of the state-of-the-art work, HIRAC, combining matrix tiling with the fast packing algorithm, SorPack, to achieve effective acceleration of Sparse General Matrix Multiplications (SpGEMMs). However, the SorPack algorithm executes on the host CPU, and its runtime lies on the critical path of the overall execution. In this work, a specific SpGEMM accelerator named OBSD is proposed which achieves high performance and energy efficiency. An on-the-fly block-wise distillation approach is proposed for leveraging the block-wise sparsity in both weight and activation data, which is implemented during the data loading process, consuming minimal additional time. Moreover, a data-flow architecture is devised to improve efficiency of data exchange among data processing elements (DPs) with investigating the relationship for different partition sizes, block-wise sparsity, and block-wise distance. The evaluation results demonstrate that: (1) OBSD achieves average of 1.89× speedup as compared to HIRAC for representative matrices in DNN workloads. (2) An end-to-end evaluation on a DNN model shows a 2.41× speedup and energy efficiency improvement of 1.99× over the HIRAC. (3) OBSD possesses a power consumption 12.9W with 56.6mm2area @28nm TSMC. Yanhuan Liu, Kunming Zhang, Zhihua Fan, Lexin Wang, Tianyu Liu 0007, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An |
HPCC | 9 |
| 2024 | Improving Utilization of Dataflow Unit for Multi-Batch ProcessingabstractDataflow architectures can achieve much better performance and higher efficiency than general-purpose core, approaching the performance of a specialized design while retaining programmability. However, advanced application scenarios place higher demands on the hardware in terms of cross-domain and multi-batch processing. In this article, we propose a unified scale-vector architecture that can work in multiple modes and adapt to diverse algorithms and requirements efficiently. First, a novel reconfigurable interconnection structure is proposed, which can organize execution units into different cluster typologies as a way to accommodate different data-level parallelism. Second, we decouple threads within each DFG node into consecutive pipeline stages and provide architectural support. By time-multiplexing during these stages, dataflow hardware can achieve much higher utilization and performance. In addition, the task-based program model can also exploit multi-level parallelism and deploy applications efficiently. Evaluated in a wide range of benchmarks, including digital signal processing algorithms, CNNs, and scientific computing algorithms, our design attains up to 11.95× energy efficiency (performance-per-watt) improvement over GPU (V100), and 2.01× energy efficiency improvement over state-of-the-art dataflow architectures. Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
ACM Trans. Archit. Code Optim. | 3 |
| 2023 | Alleviating Transfer Latency in DataFlow Accelerator for DSP ApplicationsabstractTowards multiple domains, dataflow accelerators show superiority for their flexible programmability and high efficiency. This efficiency relies highly on data communication between processing elements (PEs), which is sensitive to PE location, array scale and workload size. Laying out instructions as a dataflow graph on the PE array creates more instruction-level parallelism. However, the farther distance between remote PEs and memory banks introduces extra transfer latency, bringing performance degradation to high real-time applications. This paper examines the workloads of digital signal processing across different data scales and classifies latency problems related to data transfers and kernel switching. Specifically, we propose a novel forwarding network on chip to alleviate transfer latency and improve multi-destination sharing in the dataflow execution. Moreover, we devise bandwidth reusing mechanism to speedup kernel switching. The experiment results show that our scalable design achieves up to 2.19× (1.45× on average) speedup while reducing switching overhead by 9.85×, with an area overhead of 10.82% over the conventional dataflow accelerator. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Junying Huang, Shengzhong Tang, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan |
ICCD | 4 |
| 2023 | Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output ActivationabstractDeep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
IEEE Trans. Parallel Distributed Syst. | 3 |