Gang Cai

dblp:18/7669 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Theory of computation · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A new modified Halpern-type splitting algorithm for solving monotone inclusion problems in reflexive Banach spaces
Lulu Chen, Gang Cai, Prasit Cholamjiak, Papatsara Inkrong
J. Glob. Optim.2
2023 Cycle sampling neural network algorithms and applications
Gang Cai, Lingyan Wu
J. Supercomput.1
2022 A Soft RISC-V Processor IP with High-performance and Low-resource consumption for FPGA
abstract
Compared with hardcore processors, adding softcore processors can help FPGA to improve reliability. Many existing soft processors only aim at minimizing FPGA resources consumption or achieving high performance. However, a high-performance processor with low-resource consumption is demanded in implementing the hardware accelerator. To achieve this goal, a 32-bit soft processor based on the RISC-V instruction is proposed in this paper. The proposed processor supports RV32IM and configurable pipelining. A hierarchical decoding architecture is presented to reduce the redundancy in the decoding stage, which can reduce resource consumption. A performance optimization scheme is proposed which is used in the execution unit aiming at improving operating frequency and instructions per cycle (IPC). The operating frequency is improved by shortening the critical path and IPC is improved by reducing the stall cycles. The proposed processor is implemented on a Xilinx Zedboard and compared with the commercial soft processor-MicroBlaze. The performance of the proposed processor is 3.75 times higher than that of MicroBlaze with resource consumption increasing only 7%.
Gang Cai
ISCAS2
2021 Cheetah: An Accurate Assessment Mechanism and a High-Throughput Acceleration Architecture Oriented Toward Resource Efficiency
abstract
Convolutional neural network (CNN) is widely used in artificial intelligence for its excellent recognition accuracy. With its scale increasing rapidly and architecture becoming complicated, it is much difficult to implement CNN in hardware platform efficiently. Many FPGA-based CNN accelerators are proposed in previous work. However, when evaluating resource efficiency, their assessment methods are: 1) device related; 2) frequency related; or 3) they confuse resource efficiency with resource occupancy. There is an insistent demand for intuitive and fair assessment criteria. When implementing CNNs, they still have improvement room in computing resource efficiency, especially for layers with large feature size and few feature maps. In this work, we propose Rscoreand Cscore, which compose a comprehensive and accurate resource efficiency assessment mechanism for evaluation and design guidance, respectively. Under the guidance, we introduce Cheetah, an FPGA-based high-throughput acceleration architecture. Its computing part can optimize the use of available resources in both time and space aspects, resulting in better throughput improvement. An auxiliary storage system and a pipeline stage compression method are designed for less storage overhead and shorter inference latency. We implement AlexNet and ResNet18 on KCU1500 at 230 and 240 MHz, respectively, with a throughput of 2411.01GOP/s and 2435.05GOP/s for 16-bit quantification. Cheetah achieves an excellent average Rscoreof (0.9441, 0.9456) on different FPGA devices, while the others' mainly distribute between 0.3 and 0.8. Finally, Cheetah has 6.78X speed improvement and 1.87X power-efficiency improvement than that of Nvidia Jetson TX2, which is the fastest, most power-efficient embedded AI computing device.
Xinyuan Qu, Gang Cai, Zhen Fang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 A Resource and Performance Optimization Reduction Circuit on FPGAs
abstract
Reduce is a fundamental computing pattern, which is widely involved in scientific and engineering applications. For example, accumulation, the most common example of reduce pattern, is the core of applications such as dot product, matrix multiplication, and finite impulse response (FIR) filter. However, there is a trade-off between performance and area in the hardware implementation of the reduce pattern. To solve this problem, we propose an optimized reduction method that can handle multiple arbitrary-length sets. The performance of the proposed method is evaluated for both a single data set and numerous data sets. Moreover, to quickly differentiate the data of different sets in the reduction circuit, individual modules are designed to manage the data. We implement the design on FPGAs and present the experimental results. The proposed design with high performance and low resource consumption can achieve at least 1.59 times improvement on area-time product compared with the reported methods.
Linhuai Tang, Gang Cai, Yong Zheng 0002
IEEE Trans. Parallel Distributed Syst.2
2013 Modified extragradient methods for variational inequality problems and fixed point problems for an infinite family of nonexpansive mappings in Banach spaces
Gang Cai, Shangquan Bu
J. Glob. Optim.1
2013 Strong convergence theorems for variational inequality problems and fixed point problems in uniformly smooth and uniformly convex Banach spaces
Gang Cai, Shangquan Bu
J. Glob. Optim.1