Zhichao Wei

dblp:217/0554 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mixed-precision Neural Networks on RISC-V CPU with Reconfigurable SIMD Instruction Extension via eFPGA
abstract
Mixed-precision quantization approach has become a key technique for deploying neural networks (NNs) on Central Processing Units (CPUs). However, current CPUs face the following major limitations for mixed-precision NNs: Instruction Set Architecture (ISA) lack adaptability to mixed-precision SIMD instructions, bandwidth bottlenecks in the memory system due to large amounts of vector data, insufficient programmability and flexibility lead to fragile over-optimization as applications change rapidly. In this work, we demonstrate an extended RISC-V CPU with a customized tightly-coupled embedded FPGA (eFPGA) for reconfigurable SIMD instructions extension, targeting mixed-precision NNs deployment. We optimize the vector data mapping strategy with a configurable precision data path for eFPGA access in the CPU extension and design innovative SIMD instructions that extend the RISC-V ISA. We focus on cache hierarchy optimization with related vector register and LSU design choices to achieve high bandwidth SIMD operations. We enable the implementation of high-throughput neural MAC operations at different precisions on a resource-constrained eFPGA fabric via unpacking units and lane-based parallelism. Experimental results demonstrate that our approach, performed on the eFPGA with representative mixed-precision quantized NNs, can achieve an average 23.7× performance and 9.2× energy efficiency gains over the baseline processor.
Zixin Yang, Zhichao Wei, Jian Wang 0036, Jinmei Lai 0001
ACM Great Lakes Symposium on VLSI3
2026 Linguistic query-guided mask generation for referring image segmentation
Zhichao Wei, Xiaohao Chen, Mingqiang Chen, Zilong Dong, Siyu Zhu 0001
Pattern Recognit.1
2024 TB-TBP: a task-based adaptive routing algorithm for network-on-chip in heterogenous CPU-GPU architectures
abstract
Abstract With the rapid development of heterogeneous network-on-chip (NoC), a vast amount of shared resources are integrated into NoC. Intense resource competition exists between CPUs and GPUs, leading to congestion and a decrease in overall network performance. Reasonable node placement can minimize network conflicts at the topology level. This paper first discusses the placement of shared last-level cache and memory controller, then selects a more rational placement method and optimizes the path. To solve the hot spots problem in center placement method, a task-based routing algorithm is designed to plan the path. Simulation results demonstrate that, compared to the traditional routing algorithm, the overall network latency is reduced by 9%, and the CPU performance is improved by 13.6%. Furthermore, a dynamic task-based routing algorithm is proposed. Compared to the static task routing algorithm, the overall network latency is reduced by 2.08%, and the CPU performance is improved by 4.09%.
Juan Fang 0004, Zhichao Wei, Yumin Hou
J. Supercomput.2
2023 DPBC-VCP: A Network-On-Chip Prioritization Mechanism Combined with VCP for CPU-GPU Heterogeneous Systems
abstract
When executing CPU and GPU applications in CPU-GPU heterogeneous systems, a common phenomenon arises where CPU applications performance is often interfered by GPU applications. This study substantiates this observation through an analysis of resource contention and identifies the limitations of the Virtual Channel Partitioning (VCP) approach in the crossbar switch allocation stage. In response to the resource contention problem in crossbar switch allocation stage, we propose a Probability-Based CPU-first Arbitration Strategy that enhances the priority of CPU packets in contention through specific probabilities. Furthermore, we introduce a Dynamic Probability-Based CPU-first Arbitration Strategy (DPBC) that dynamically selects probability values based on application execution phases to strike a balance between optimal CPU and GPU performance. Moreover, we combine this dynamic strategy with VCP to further enhance CPU performance, propose the DPBC-VCP method. The DPBC-VCP, a combination of network partitioning and prioritization techniques, yields an average enhancement of 48% in CPU performance compared to the baseline, with only a marginal 2.45% reduction in GPU performance.
Haoyu Cheng, Zhichao Wei, Huijing Yang
ICPADS3
2021 Efficient Video Compressed Sensing Reconstruction via Exploiting Spatial-Temporal Correlation With Measurement Constraint
abstract
Recent deep learning-based video compressed sensing (VCS) methods have achieved promising results but still suffer from numerous hyper-parameters and inflexibility. This paper proposes a novel network for VCS, named STM-Net, to fast recover high-quality video frames by optionally exploiting Spatial-Temporal information with a Measurement constraint. Combining the merits of adaptive sampling and adaptive shrinkage-thresholding, we first propose an improved ISTA-Net+ for framewise independent reconstruction, called Unfolding Adaptive Shrinkage-Thresholding Network (UAST-Net). To get further non-key frames reconstruction improvement, we develop a two-phase joint deep reconstruction, including an Occlusion-Aware Temporal Alignment to avoid irrelevant information compensation and a Multiple Frames Fusion with proposed Spatial-Temporal Feature Weighting (STFW) module to guide attractive content extraction and discriminative features generation. Besides, we develop a measurement loss to reduce the solution space to facilitate network optimization. Experimental results demonstrate the superiority of the proposed STM-Net over the existing methods.
Zhichao Wei, Yunyi Xuan
ICME1