Bowen Zhang 0004

dblp:85/7433-4 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0001-6934-9487ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 78% Interconnection networks and networks-on-chip · 22%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
dataflow mapping
0.812024
A 3D Hybrid Optical-Electrical NoC Using Novel Mapping Strategy Based DCNN Dataflow Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
DCNN accelerator
0.812024
A 3D Hybrid Optical-Electrical NoC Using Novel Mapping Strategy Based DCNN Dataflow Acceleration · IEEE Trans. Parallel Distributed Syst. 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
convolution acceleration
0.612022
A Novel CONV Acceleration Strategy Based on Logical PE Set Segmentation for Row Stationary Dataflow · IEEE Trans. Computers 2022
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.612022
A Novel CONV Acceleration Strategy Based on Logical PE Set Segmentation for Row Stationary Dataflow · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

row stationary dataflow · 1.3genetic algorithm · 0.8PE set segmentation · 0.6
YearPublicationVenuePosition
2024 A 3D Hybrid Optical-Electrical NoC Using Novel Mapping Strategy Based DCNN Dataflow Acceleration
abstract
A large number of multiply-accumulate operations and memory accesses required in deep convolutional neural networks (DCNN) leads to high latency and energy consumption (EC), that hinder their further applications. Dataflow-based acceleration schemes reduce memory accesses by leveraging reusable data in DCNNs. Row Stationary (RS) dataflow is a more advanced dataflow. In the convolutional layer acceleration of RS dataflow, the flexibility of mapping from logical processing element (LPE) sets to physical PE sets is relatively poor. The utilization of processing elements (PEs) is low. In this paper, a novel mapping strategy based on genetic algorithm (GAMS) with the goal of optimizing EC is proposed. GAMS is designed to address the energy inefficiencies faced when mapping RS dataflow. A 3D hybrid optical-electrical Network-on-Chip (3DHOENoC) is proposed to further improve the communication efficiency, energy efficiency and the processing speed of DCNN. Simulation and evaluation results show that GAMS can achieve better mapping flexibility, higher PEs utilization and 15.9% improvement of execution speed on average. In addition, the execution time (ET) performance of processing the DCNN can be further improved by adopting the 3DHOENoC architecture with better communication parallelism.
Bowen Zhang 0004, Huaxi Gu, Grace Li Zhang, Yintang Yang, Ziteng Ma, Ulf Schlichtmann
IEEE Trans. Parallel Distributed Syst.1
2022 An Efficient Dataflow Mapping Method for Convolutional Neural Networks
Zhuangzhuang Liu, Huaxi Gu, Bowen Zhang 0004, Canran Shi
Neural Process. Lett.3
2022 A Novel CONV Acceleration Strategy Based on Logical PE Set Segmentation for Row Stationary Dataflow
abstract
Deep convolutional neural networks (DCNNs) have been proposed as enhanced developments of neural networks (NNs) in the field of artificial intelligence (AI) and successfully applied in deep learning (DL) scenarios. With the advancement of technology, the number of network layers has continuously increased, resulting in a huge number of calculations and memory accesses required in the training and inference process of DCNNs and thereby hindering their further deployment and application. Using a specific dataflow formed by reusable DCNN data in the network-on-chip (NoC), reducing the memory access pressure and improving DCNN processing efficiency has become a promising acceleration schemes for the current DCNN. In this paper, a novel convolution layer (CONV) acceleration strategy based on logical PE set segmentation for row stationary (RS) dataflow is proposed to solve the problems of low flexibility and inefficient processing array utilization faced by the conventional folding mapping strategy. The simulation results show that the new mapping strategy based on PE set segmentation can achieve better processing element utilization and CONV acceleration improvement at the expense of little increase in the data movement energy consumption compared with the conventional strategy.
Bowen Zhang 0004, Huaxi Gu, Kun Wang 0001, Yintang Yang
IEEE Trans. Computers1
2019 TAONoC: A Regular Passive Optical Network-on-Chip Architecture Based on Comb Switches
abstract
Optical networks on chip (ONoC) has been proposed as a promising alternative paradigm for electronic NoC with the benefit of optical signaling communications such as ultrahigh bandwidth, extremely low energy consumption, and negligible transmission latency. To accommodate the layout of tile-based chip multicore processors, a torus-based passive ONoC architecture, TAONoC, is proposed in this paper. Relying on the unique designs of three function modules, TAONoC can still support contention-free communication without the need for arbitration. TAONoC employs comb switches instead of general microring resonators (MRs). TAONoC has a low demand for the number of MRs because of the ultrahigh utilization of resonant wavelengths owned by a single MR. Simulation results show that TAONoC performs well under three different synthetic traffic patterns.
Yintang Yang, Huaxi Gu, Bowen Zhang 0004, Lijing Zhu
IEEE Trans. Very Large Scale Integr. Syst.4
2018 A highly efficient dynamic router for application-oriented network on chip
Huaxi Gu, Kun Wang 0001, Xiaoshan Yu 0001, Bowen Zhang 0004
J. Supercomput.5