Wei Chen 0100

dblp:181/2832-100 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-9603-9203ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Instruction Fusion for ASIPs: A Low-Cost Approach to Structural Hazard Control
abstract
Moore’s law is reaching saturation, while the demand for computing capacity continues to rise. Application-Specific Instruction-Set Processors (ASIPs) are therefore an essential technology for embedded processing in areas such as communication, media, gaming, and control systems, due to their high area and energy efficiency. The most effective method for ASIPs acceleration is known as instruction fusion. Instruction fusion enhances performance but also introduces challenges related to the complexity of structural hazard control. The difficulty in this research field is achieving fine-grained structural hazard control using low-cost hardware. The state-of-the-art technologies have not been able to effectively address this challenge. Firstly, this paper systematically analyzes the rationale behind instruction fusion and explores methods to optimize its performance. Secondly, based on the design of a 5G micro-base station baseband processor, we present an instruction fusion pipeline design example. Furthermore, to address the structural hazards that arise from instruction fusion, we propose a lightweight Hardware Resource Table (HRT) to address the issue and outline its benefits. These benefits include utilizing low silicon costs to achieve performance improvements and reducing on-chip program memory. According to benchmark evaluations, area efficiency improves by 23%, accompanied by a 108% increase in energy efficiency.
Xinbing Zhou, Tianlang Liu, Tiancheng Tang, Yi Man, Wei Chen 0100, Dake Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 Low-Latency and Low-Overhead Fault-Tolerant Network-on-Chip for Embedded Real-Time Heterogeneous Parallel Digital Signal Processors
abstract
At present, most network on chips (NoCs) are designed for symmetric multiprocessing architectures for general purpose computing, it is accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. How to design NoC with high bandwidth transmission, ultra-low latency, and low silicon overhead for embedded real-time heterogeneous parallel digital signal processor, which faces major challenge. In order to solve this challenge, we hence propose an innovative NoC architecture. It is designed to meet the high throughput, low latency and silicon overhead requirements of embedded real-time heterogeneous parallel digital signal processor systems. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. Our designed NoC has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets, but also improve the transmission efficiency of short data packets. In addition, we design efficient routing methods and fault-tolerant data transmission mechanisms to improve reliability and performance of NoC, which can mitigate the impact of router failures, reduce the need of reorder buffer memory and power consumption, and improve the robustness of NoC. The experiments showed the performance advantages of our design in terms of latency, throughput and silicon overhead in comparison with the state-of-the-art NoC, our design respectively reduced the latency, area and power consumption by 93%, 97.5% and 94.2% respectively, and improved the throughput and MTTF (Mean Time To Failure) by 22% and 9.4 times respectively. This verifies its effectiveness in high-end applications.
Wei Chen 0100, Dake Liu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Efficient and Flexible Deep Learning Processor for Embedded Devices With 99.8% PE Utilization Under Low Silicon Cost
abstract
The flexibility and efficiency of accelerator hinder the realization of DNN (Deep Neural Network) on embedded devices. How to design efficient and flexible deep learning processors to meet low-power requirements while supporting various iteratively updated deep neural networks, which is a major challenge in the current field. Moreover, the significant power consumption and latency generated by data access between on-chip and off-chip memory have been the bottleneck of DNN accelerator development. We hence design the flexible and efficient deep learning processor (DLP) based on our designed specific neural network instruction set for DNN inference, which can be used to flexibly accelerate the inference of iteratively updated neural networks. Our design includes the design of hardware architecture, neural network acceleration instruction set, efficient SIMD control path and data path, which can be used to accelerate the execution of the DNN programs, and support hiding the data transfer time between on-chip and off-chip memory. In addition, we propose a parallel and conflict-free data access scheme to reduce data access overhead in memory system. Moreover, we also propose a scheduling framework to improve performance and minimize power consumption under the hardware resource constraints. The experimental results indicate that our design’s power consumption efficiency is 75.3 TOPS/W at TSMC 65 nm process. Our power efficiency is 7.6, 5.5 and 4 times higher than the state-of-art DNN accelerators respectively.
Wei Chen 0100, Dake Liu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Efficient NoC for Embedded Heterogeneous Multi-Core DSPs
abstract
At present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications.
Wei Chen 0100, Dake Liu
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Task Scheduling for Heterogeneous Multi-Core Processors Based on Deep Reinforcement Learning
abstract
Heterogeneous multicore processor systems are commonly used for scheduling tasks of DAG applications. Deep reinforcement learning, with its superior ability to perceive decisions directly and handle high‐dimensional state actions, has become a prevalent solution for scheduling these systems. However, the incomplete environment models and large action spaces of deep reinforcement learning present significant challenges to scheduling. This paper investigates a scheduling problem in a heterogeneous multicore processor environment. Initially, system environment information is extracted and encoded using a graph convolutional neural network based on integrating adapter and AdapterFusion into the transformer architecture. Then, by separating task selection and processor allocation, the decision space is reduced: the former uses a deep neural network to learn to select nodes, and the latter allocates processors using a heuristic scheduling algorithm combining earliest completion time‐based node replication and rolling technology. The entire scheduling process is a Markov decision problem. Therefore, the PPO algorithm with dynamic adjustment of the clipping factor, combined with an advantage actor‐critic network, is employed for training, optimizing, and evaluating the algorithm to find the optimal scheduling strategy. The training process adopts a reward function for the time and power consumption required for completed task scheduling to ensure that multiple DAG application task scheduling can achieve optimal performance. Experiments conducted in various environments with different parameters show that, compared to other algorithms, this algorithm reduces the overall execution time and power consumption cost of heterogeneous multicore processor tasks by 11.09%.
Qiguang Tan, Wei Chen 0100, Dake Liu
Int. J. Intell. Syst.2
2025 Hardware Sharing Design Method of Rate Matching and Interleaving for Wireless Terminal in Industrial Internet of Things
abstract
Reducing the power consumption of baseband chips in wireless terminals of the Industrial Internet of Things (IIoT) is a great challenge. Among them, the logic function of the rate matching and interleaving hardware module is very complex, which occupies a considerable portion of power consumption in the baseband chip, so it is of great significance to design a low-power rate matching and interleaving hardware. Due to differences in interleaving algorithms and throughput, the interleavers used in 4G long term evolution (LTE) and 5G NR/6G have not been merged into a single architecture. Switching between different standards provides new options for interleaver design, and by configuring the architecture of different interleavers, it is possible to use the same hardware for different standards to reduce hardware resources and power consumption. This article studies the block interleaving and rate matching of turbo codes, convolutional codes, polar codes, and low-density parity check (LDPC) codes used in 4G LTE and 5G NR/6G communication links. With regard to the different algorithms for these four types of encoding, shared design is carried out on the hardware structure. In this experiment, according to the proposed memory and interleaving sharing scheme for hardware design and hardware simulation, the area overhead of 0.1$\mu $m2 and power consumption of 4.31 mW are obtained by Synopsys synthesis at the SMIC 28-nm process and the frequency of 50 MHz. This achieves the maximum hardware reuse of four encoding schemes in the downlink communication link of 4G LTE and 5G NR/6G, and reduce power consumption.
Wei Chen 0100, Kejia Huo, Dake Liu
IEEE Internet Things J.1
2025 High-Throughput LDPC Decoder for Multiple Wireless Standards
abstract
It is a great challenge to design an LDPC decoder with multi-standard compatibility, flexibility and low silicon overhead. This paper presents the efficient and low-overhead design of an LDPC decoder tailored for multi-standard, which include WLAN, 5G NR and WiMAX. We follow the design principles of Application-Specific Instruction-set Processor (ASIP). In order to enhance throughput, we double the computational speed by reducing memory speed from double logic speed to logic speed. By proposing the optimized hybrid scheduling algorithm based on matrix reordering, we further solve scheduling problems and eliminate pipeline conflicts. Through performing logic synthesis utilizing the 28 nm SMIC CMOS cell library, synthesis results show that the core area of our designed decoder is 0.86 mm2, the logic gate count is 1716 K, and our design achieves impressive throughput rates, that is up to 9.96 Gbps for WLAN, 7.69 Gbps for WiMAX, and 33 Gbps for 5G NR. Compared with other state-of-the-art LDPC decoders, the experimental results show that our proposed decoder has up to$4.5\times $higher throughput,$3.9\times $better area efficiency and$5.8\times $better energy efficiency than these state-of-the-art implementations.
Wei Chen 0100, Dake Liu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 B5G/6G URLLC Latency Reduction Method for Multisensor Industrial Internet of Things
abstract
How to enable B5G/6G ultra reliable low-latency communications (URLLC) to meet the requirements of low latency, ultrahigh-speed and real-time industrial automatic control for multisensor Industrial Internet of Things (IIoT), it is one of the most important challenges of multisensor IIoT. The round-trip time (RTT) is significantly reduced by semi-persistent scheduling (SPS) or short transmission time interval (TTI) method, yet these methods may cost extra frequency and time resources. To reduce latency without the cost of extra spectrum and time resources, such research is almost blank in both academia and industry for B5G/6G URLLC scenario of multisensor IIoT. We hence propose the latency reduction method for the B5G/6G URLLC scenario of multisensor IIoT. Our research fills this research gap. The proposed reduction method includes a storage planning method for global data and local data, a dynamic data configuration method and an execution time minimization method. Based on these methods, we can obtain the minimum data space of the program and the data storage arrangement approaching to the minimum worst-case execution time (WCET). Under the main frame of SPS and short TTI method, our proposed method can further reduce latency without reducing spectrum and time resources utilization by minimizing data storage and access delay. The experimental results show that our method did not reduce spectrum and time resources utilization, and when our method was applied in SPS and short TTI method, our method further reduced the physical layer computing latency by about 21% to 28%.
Wei Chen 0100, Dake Liu, Yong Bai 0002
IEEE Internet Things J.1
2024 Conflict-Free Parallel Data Access Technology for Matrix Calculation in Memory System of ASIP of 5G/6G Macro Base Stations
abstract
Among the physical layer baseband algorithms in macro base stations, the matrix processing has the dominant computing cost, large data access overhead, and complicated addressing mode. The existing data access methods are not the best solution for memory system of application-specific instruction set processor (ASIP) in 5G/6G macro stations. We hence proposed a parallel conflict-free data access method for matrix. Moreover, we proposed a nonredundant access method for positive-definite matrix that stores only trigonometric part, and the corresponding parallel addressing method with low storage overhead. Our method solves the problem of data conflict and minimizes the memory space of ASIP, and supports ASIP to approach the performance limit of architecture to the maximum extent under the constraint of architecture and data parallelism. We implemented the proposed method as a static memory optimizer, which provides a tool for the design and optimization of memory system of ASIP in 5G/6G macro base stations with low overhead. Experimental results show that for the${64}\mathbf {\times }{64}$positive-definite matrix inversion algorithm based on Cholesky decomposition, our addressing method can save 32% of the execution time, and the cost of memory space is reduced by half.
Wei Chen 0100, Dake Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1