Wei Zhang 0055

dblp:10/4661-55 · DBLP profile ↗
← Back
45ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-2601-3198ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 11 since 2021Systems, architecture and hardware · 11 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 6 since 2021Computer networks · 7 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Bit-Flip Fano Sequential Decoding of Polarization-Adjusted Convolutional (PAC) Codes
Yaofeng Jiang, Wei Zhang 0055, Zhuolun Wu, Jianhan Zhao, Yanyan Liu 0001
WCNC2
2026 Spiking Depth: Depth estimation from sparse events with spiking neural networks
Dongze Liu, Yimeng Fan, Wenrui Lu, Changsong Liu, Wei Zhang 0055
Expert Syst. Appl.5
2026 MS2Edge: Towards energy-efficient and crisp edge detection with multi-scale residual learning in SNNs
Yimeng Fan, Changsong Liu, Yuzhou Dai, Yanyan Liu 0001, Wei Zhang 0055
Pattern Recognit.6
2026 Reduced Complexity Blind Recognition Method of LDPC Codes Over a Candidate Set
abstract
Adaptive modulation and coding (AMC) systems require the transmission of control signals, thereby reducing overall system transmission efficiency. The channel coding blind recognition technique is key to solving this problem. This paper proposes a reduced-complexity method for blind recognition of low-density parity-check (LDPC) coding parameters within a given candidate set, thereby enhancing the work efficiency of AMC systems. This paper applies a method based on code rate for classification evaluation, circumventing superfluous calculations for several candidate parity-check matrices. Furthermore, computational complexity is reduced by applying the offset min-sum algorithm (OMSA) to the parity-check stage. Subsequently, the Z-score is used to measure the difference between the actual data and the theoretical distribution. Compared with the best existing recognition methods, the proposed algorithm offers clear advantages in computational complexity and is virtually identical in recognition performance.
Zhuolun Wu, Yushan Zhang, Wei Zhang 0055, Yanyan Liu 0001
IEEE Signal Process. Lett.3
2025 Cycle pixel difference network for crisp edge detection
Changsong Liu, Wei Zhang 0055, Yanyan Liu 0001, Yimeng Fan, Xiangnan Bai
Neurocomputing2
2025 Learning to utilize image second-order derivative information for crisp edge detection
Changsong Liu, Yimeng Fan, Wei Zhang 0055, Yanyan Liu 0001
Knowl. Based Syst.4
2025 Frequency Domain Low Complexity Chase Reed-Solomon Decoding Under Blind Recognition Prior Conditions
abstract
In non-cooperative communication systems, recognition and decoding are two key technologies that determine whether the receiving terminal can correctly decipher the information. To better utilize the prior information by recognition and improve the decoding performance, a novel frequency domainlow complexity chase (F-LCC) Reed–Solomon (RS) decoding scheme is proposed. Firstly, the relationship between recognition and decoding is analyzed, and the erroneous codewords to be decoded can be selected by analyzing the spectrogram obtained from recognition; Then, the$\eta$-dynamically adjustable multiplicity assignment ($\eta$-DAMA) algorithm is proposed by analyzing the channel soft information, where$\eta$is the number of the most unreliable code positions. Wherein, the number of test vectors$2^{\eta }$is modeled as a function of code length and bit error rate (BER), which solves the problem of traditional decoding algorithms relying on empirical values to determine$\eta$. Simulations verify the frame error rate (FER) of the proposed F-LCC for decoding RS(255,239) with different signal-to-noise ratios (SNRs). The results show that the proposed F-LCC decoding scheme can achieve a gain of 0.3dB and 0.15dB compared with the decoding performance of the existing time domain-LCC (T-LCC) decoding scheme when$\eta$= 3 and$\eta$= 5, respectively.
Yanyan Chang, Wei Zhang 0055, Zhuolun Wu, Yanyan Liu 0001
IEEE Signal Process. Lett.2
2025 Blind Recognition Algorithm of RS-SPC Concatenated Codes Based on Single-Error Correction
abstract
This letter presents an RS-SPC concatenated code blindrecognition algorithm based on the single-error correction. The algorithm corrects the least reliable bit of the single parity check (SPC) codewords based on the parity check characteristics, thereby increasing correct Reed-Solomon (RS) codewords and laying the foundation for improving the recognition probability. In addition, this algorithm combines threshold judgement with the matrix recording method, thereby eliminating unnecessary iterative operations under the condition that accurate recognition is possible. At the same time, it employs probability theory as a theoretical basis to quantify the degree of dispersion of the data through sample variance. The experimental results demonstrate that the recognition probability of this algorithm is superior to that of all other algorithms. When the codeword error rate (CER) is 0.5, RS(15,9)-SPC(4,3) still has a recognition probability of 20%. For RS(255,239)-SPC(8,7), the gain of the proposed algorithm exceeds 1.3dB compared to the upper bound of the recognition probability.
Zhuolun Wu, Wei Zhang 0055, Yihan Wang 0002, Yushan Zhang, Yanyan Liu 0001
IEEE Signal Process. Lett.2
2025 Blind Recognition of Intercepted Limited BCH and RS Codes Bitstream Using Single-Error Correction Under Harsh Channel Environments
abstract
This paper proposes a full-blind recognition method that aims to recognize coding parameters from a noisy bitstream of limited BCH and RS codes intercepted in a non-cooperative communication system under harsh channel environments. This paper optimizes the codeword reliability calculation by fully utilizing channel soft information, based on which the single-bit error codewords can be effectively selected from the codeword stream with a high probability. Subsequently, the channel soft information-based single-error correction (CSI-SEC) algorithm is proposed for correcting single-bit error codewords. The blind recognition algorithm, based on the Galois domain Fourier transform (GFFT) with a dual decision-making mechanism, is applied to these corrected codewords and reliable codewords to extract the encoding parameters. Finally, the method’s superiority is demonstrated through comprehensive simulation experiments. Compared to other blind recognition methods based on error correction, the proposed method exhibits enhanced recognition performance, with a cumulative computational complexity of only 39.8% of that of the other methods. Compared to other algorithms for blind recognition of limited intercepted bitstream, the proposed method achieves at least a 3 dB improvement in recognition performance.
Zhuolun Wu, Wei Zhang 0055, Yihan Wang 0002, Hao Wang 0072, Yanyan Liu 0001
IEEE Trans. Commun.2
2025 An Area-Efficient VLSI Architecture for High-Throughput Computation of the 2-D DWT
abstract
In this article, an area-efficient VLSI architecture scheme for high-throughput computation of the 2-D discrete wavelet transform (DWT) is proposed, effectively applied in the context of aircraft cargo hold scenes. The proposed architecture aims to reduce computation and storage resources while maintaining the DWT-IDWT reconstructed image quality for the 9/7 discrete wavelet. The hardware implementation formulae based on the flipping architecture have been modified to reduce RAM storage bit width. By transforming the coefficients of the formula into hardware-friendly values, the required multiplication operations are split into two stages of addition. On this basis, a pipelined architecture is constructed to set the critical path delay (CPD) of the architecture to be close to the delay of a single adder,$T_{a}$, thereby achieving a high throughput. Compared to existing architectures in the research field, the proposed single-level 2-D DWT architecture achieves resource savings on the field-programmable gate array (FPGA) platform while ensuring good image reconstruction quality. The advantages of the multilevel 2-D DWT are even more pronounced. In the simulation results on the application-specific integrated circuit (ASIC) platform, the proposed architecture reduces computation time by at least 35.54% while achieving a higher level of decomposition, decreases the area-delay product (ADP) by at least 25.41%, and saves a significant amount of energy per image (EPI). Furthermore, the proposed folded architecture achieves close to 100% hardware utilization efficiency (HUE) in multilevel 2-D DWT computations.
Yuzhou Dai, Wei Zhang 0055, Qitao Li, Zhuolun Wu, Yanyan Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 High-Performance Error and Erasure Decoding With Low Complexities Using SPC-RS Concatenated Codes
abstract
In this brief, a novel single parity check-assisted error and erasure decoding (SPC-EED) algorithm over the compound channel is proposed for Reed-Solomon (RS) codes. EED is applied to the scheme of SPC-RS concatenation code. The proposed algorithm can correct most random errors with the assistance of SPC codes. Through using channel soft information, the symbols with burst errors can be identified accurately and efficiently for erasures in EED. Simulation results show that the SPC-EED algorithm can achieve coding gains of up to 0.1, 0.4, and 2.1 dB compared with burst-error correcting (BC)-SPC-ordered statistic decoding (OSD), BC-OSD, and BCHDD-low-complexity chase (LCC) for RS(255, 239) codes, respectively, at symbol error rate (SER) = 10−3. The proposed algorithm has lower computational complexities while maintaining better decoding performance. The hardware design of the SPC-threshold check module is provided. The implementation results in ASIC show that compared with the BCHDD-LCC decoder, the SPC-EED decoder improves area efficiency by 249.37% with higher energy efficiency.
Wei Zhang 0055, Jianhan Zhao, Yanyan Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Fast Successive-Cancellation Decoding of Polar Codes with Reed-Solomon Kernel
abstract
Polar codes with Reed-Solomon (RS) kernel exhibit significant promise in forthcoming communication systems owing to their elevated polarization rate. However, the successive cancellation list (SCL) decoders for this type of multi-kernel polar codes suffer from high decoding complexity. In this paper, we comprehensively investigate the characteristics of the Reed-Solomon (RS) kernel and propose three decoding nodes specifically designed for polar code with RS kernel, including Rate-0, Rate-1 and Quasi-Repetition (QREP) nodes, along with their corre-sponding soft information processing methods. Furthermore, we propose efficient Dual Reliable Decoding (DRD) algorithms and table-based algorithms for Rate-1 and QREP nodes, respectively. Results show that the proposed fast SCL algorithm of polar codes with RS kernel can reduce the complexity, while achieving similar performance.
Jianhan Zhao, Wei Zhang 0055, Yanyan Liu 0001
ITW2
2024 Generating crisp boundaries using multi-scale features and mixed loss function
Changsong Liu, Wei Zhang 0055, Yanyan Liu 0001, Rudong Jing
Appl. Intell.2
2024 An effective method for small object detection in low-resolution images
Rudong Jing, Wei Zhang 0055, Yanyan Liu 0001, Changsong Liu
Eng. Appl. Artif. Intell.2
2024 Feature aggregation network for small object detection
Rudong Jing, Wei Zhang 0055, Yuzhuo Li, Yanyan Liu 0001
Expert Syst. Appl.2
2024 Dynamic Feature Focusing Network for small object detection
Rudong Jing, Wei Zhang 0055, Yuzhuo Li, Yanyan Liu 0001
Inf. Process. Manag.2
2024 Blind Recognition of BCH and RS Codes With Small Samples Intercepted Bitstream
abstract
In non-cooperative communication systems, effective parameter identification with small samples intercepted bitstream is a major challenge. Faced with this demand, a solution to blind recognition for BCH and RS codes is presented. Firstly, the intercepted bitstream is analyzed, and the optimal code-set with high reliability is selected by channel soft information. Secondly, the optimal code-set is cycled into a pseudo cyclic matrix according to the algebraic structure of the codewords, which is employed for the code length and synchronization recognition based on the improved rank criterion. Finally, the code spectrum analysis of the optimal code-set is performed to complete primitive and generator polynomial recognition, and a recoding operation is proposed to judge the recognition performance. Simulations validate the proposed algorithms by recognition probability in terms of different parameters, it hardly depends on the number of codewords but on the accuracy of the optimal code-set, which is especially suitable for small sample datasets. Under the Additive White Gaussian Noise (AWGN) channel with Binary Phase-Shift Keying (BPSK) modulation, the recognition probability is still comparable to or superior to the existing schemes, even if only one-twentieth of the intercepted codewords are used.
Yanyan Chang, Wei Zhang 0055, Hao Wang 0072, Yanyan Liu 0001
IEEE Trans. Commun.2
2024 An Efficient Construction Method Based on Partial Distance of Polar Codes With Reed-Solomon Kernel
abstract
Polar codes with Reed-Solomon (RS) kernel have great potential in next-generation communication systems due to their high polarization rate. In this paper, we study the polarization characteristics of RS polar codes and propose two types of partial orders (POs) for the synthesized channels, which are supported by validity proofs. By combining these partial orders, a Partial Distance-based Polarization Weight (PDPW) construction method is presented. The proposed method achieves comparable performance to Monte-Carlo simulations while requiring lower complexity. Additionally, a Minimum Polarization Weight Puncturing (MPWP) scheme for rate-matching is proposed to enhance its practical applicability in communication systems. Simulation results demonstrate that the RS polar codes based on the proposed PDPW construction outperform the 3rd Generation Partnership Project (3GPP) NR polar codes in terms of standard code performance and rate-matching performance.
Jianhan Zhao, Wei Zhang 0055, Yanyan Liu 0001
IEEE Trans. Commun.2
2023 A Novel Concatenation Decoding of Reed-Solomon Codes With SPC Product Codes
abstract
This letter introduces a novel concatenation decoding of Reed-Solomon (RS) codes with single parity-check product codes (SPC-PCs), including the concatenation scheme and the decoding algorithm. The proposed scheme sequentially arranges SPC-PCs to form$K$information symbols, then encodes them into$N$symbols using systematic RS code. A high-performance and low-complexity bit-iterative-update decoding (BIUD) algorithm is also proposed based on this scheme. We use the results of the parity-check to evaluate the reliability in each dimension and then iteratively update the reliability of each bit based on the structure of SPC-PCs. This way dramatically reduces the complexity of traditional reliability computation. Furthermore, the bit iterative update and the Berlekamp-Massey algorithm of RS codes can be well combined, resulting in more accurate results. The simulation and analysis results show that the proposed concatenation algorithm has significant performance improvement compared to the current RS-SPC concatenation decoding while having extremely low complexity.
Wei Zhang 0055, Yanyan Chang, Yanyan Liu 0001
IEEE Signal Process. Lett.2
2023 VLSI Design of a High-Performance Multicontext MQ Arithmetic Coder
abstract
The MQ arithmetic coding, which is an adaptive arithmetic coding developing from Q coding, has become a major throughput bottleneck of JPEG2000 compression due to its inherent serial operations. To overcome the bottleneck, this brief proposes a high-performance hardware architecture of the multicontext MQ coder. The proposed architecture is capable of concurrent coding for two adjacent more probable symbols (MPSs). Performance analysis results show that the proposed coder consumes 1.61 CXD pairs per cycle and achieves a throughput of 506.93 MSymbols/s under 0.5 bpp bit rate. The proposed architecture not only achieves high throughput, but also maintains both low hardware utilization and low power consumption. Compared with the state-of-the-art two-context coder, the figure of merit (FoM) is increased by 38%. Compared with the single-context coder, the power-delay product (PDP) is reduced by 49%.
Wei Zhang 0055, Yanyan Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2023 A High-Performance Dual-Context MQ Encoder Architecture Based on Extended Lookup Table
abstract
Multiple quantization (MQ) encoder is a basic entropy encoder widely used in multimedia encoding systems, especially in JPEG2000 and JBIG2, which play a key role and make outstanding contributions to digital imaging. However, due to its strict serial working characteristics, the MQ encoder has become a key bottleneck restricting the optimization of JPEG2000 encoding performance. Many single ejection encoder architectures use pipeline staging to improve the throughput, but they have all reached their limits. A high-performance MQ encoder based on a novel lookup table is proposed to overcome this bottleneck, which improves the throughput while utilizing reasonable hardware resources and with low-power consumption. Experimental results show that the proposed encoder architecture can encode two symbols per cycle with a throughput of 1204.82 MSymbol/s, and it has a power consumption of only 11.05 mW. Compared with the conventional multi-context coder, the hardware utilization is reduced by 26%, and the hardware utilization efficiency is improved by 16%. Compared with the conventional single-context coder, the throughput is improved by 14.3%, and the power consumption is reduced by 40.1%.
Zhuolun Wu, Wei Zhang 0055, Yanyan Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2022 A visualized fire detection method based on convolutional neural network beyond anchor
Wei Zhang 0055, Yanyan Liu 0001
Appl. Intell.2
2022 An efficient fire and smoke detection algorithm based on an end-to-end structured network
Wei Zhang 0055, Yanyan Liu 0001, Rudong Jing, Changsong Liu
Eng. Appl. Artif. Intell.2
2022 Multi-scale features fused network with multi-level supervised path for crowd counting
Wei Zhang 0055, Dongxiao Huang, Yanyan Liu 0001, Jianghua Zhu
Expert Syst. Appl.2
2022 A lightweight network for real-time smoke semantic segmentation based on dual paths
Wei Zhang 0055, Yanyan Liu 0001, Xiaorui Shao
Neurocomputing2
2022 An efficient deep neural network with color-weighted loss for fire detection
Wei Zhang 0055, Yanyan Liu 0001, Jianhan Zhao
Multim. Tools Appl.2
2021 Multi-level feature fusion and multi-loss learning for person Re-Identification
Wei Zhang 0055, Dongxiao Huang, Yanyan Liu 0001
Signal Process. Image Commun.2
2021 High-Performance Concatenation Decoding of Reed-Solomon Codes With SPC Codes
abstract
A novel single parity check-multiplicity assignment decoding algorithm based on voltage magnitude (VM_SPC-MA) is proposed, which is applied to the concatenated scheme of single parity check (SPC) inner code and Reed-Solomon (RS) outer code, following the Consultative Committee for Space Data Systems (CCSDS) standard. The algorithm determines whether the SPC code is in error by SPC, then obtains the reliability information of the inner code bits for error correction based on the characteristics of the received bit-level voltage, and decodes the outer code based on the reliability information of the inner codewords and the channel information. The decoding performance is greatly improved by connecting the inner and outer codes through the multiplicity assignment (MA) module, which makes full use of the channel information. Compared with the low-complexity Chase decoding based on the hard-decision decoding (HDD-LCC) and SPC_Kaneko-RS_Chase decoding, simulation results show that the SPC-RS concatenated decoding scheme based on VM_SPC-MA algorithm can provide up to 1.78 and 1.05 dB of coding gain when the bit error rate is (BER) = 10-5. Besides, the hardware design of the SPC-MA module is provided and applied to the serial LCC decoder based on syndrome calculation-polynomial selection (PS)-Chien search and Forney algorithm (SPCF). The implementation results in ASIC show that the area efficiency of the complete concatenated decoder increases by 27.38% compared to the unified syndrome computation (USC)-based LCC decoder in a 0.13- μm process.
Wei Zhang 0055, Yanyan Liu 0001, Hao Wang 0072, Jianhan Zhao
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Multi-scale supervised network for crowd counting
abstract
Crowd counting is getting more and more attention in our daily life, because it can effectively prevent some safety problems. However, due to scale variations and background noise in the image, such as buildings and trees, getting the accurate number from image is a hard work. In order to address these problems, this work introduces a new multi‐scale supervised network. The proposed model uses part of vgg16 model as the backbone to extract feature. In the training process, a multi‐scale dilated convolution module is added at the end of each stage of the backbone network to generate attention map with different resolutions to help the model focus on the head area in feature map. In addition, the dilated convolution adopts three dilation ratios to fit different sizes of head in the image. Finally, in order to get the high‐quality density map with high‐resolution, the authors employ the upsampling operation to restore the density map size to the quarter size of original image. A large number of experiments on these four datasets show that the proposed network has greatly improved the counting accuracy of many existing methods.
Wei Zhang 0055, Dongxiao Huang, Yanyan Liu 0001, Jianghua Zhu
IET Image Process.2
2020 Multi-scale feature fusion network for person re-identification
abstract
Recently, it is becoming a challenging work for person re‐identification due to the problems of occlusion, blurring and posture. The key of effective person re‐identification is to capture sufficient detailed features of a person's appearance in images. Different from previous methods, our method mainly focuses on fusing different visual clues only depending on the features of different levels and scales without additional assistance. The major contributions of our paper are the mixed pooling strategy with different kernels and the mixed loss function. Firstly, we adopt ResNet50 as our backbone. We have slightly modified the backbone, which does not use the down‐sampling operation at the beginning of stage 4. Inspired by pyramid pooling structure, we pass the outputs of Res4 and Res5 through the average pooling layer and max pooling layer with different kernels and strides separately. Secondly, we combine the averaged triplet losses and the averaged softmax losses as the final loss of the whole network. Extensive experiments on three datasets (CUHK3, Market1501, DukeMTMC‐reID) show that compared with many state‐of‐the‐art methods in recent years, our model achieve higher accuracy.
Wei Zhang 0055, Yanyan Liu 0001
IET Image Process.2
2020 Deep convolution network for dense crowd counting
abstract
Estimating the total number of people in a crowded situation is a challenging task due to numerous occlusions and perspective changes existing in crowd images. To address this issue, the authors have proposed a new deep learning framework for accurate and efficient crowd counting here. Inspired by multi‐column convolutional neural network (MCNN) and contextual pyramid convolutional neural network (CP‐CNN), the authors use a combination of a two branches, convolutional neutral network (CNN) and transposed convolutional layers, to generate a high‐quality density map. The two‐branch CNN for feature extraction generates a density map that is only a quarter of size of the original image Then a set of transposed convolutional layers and convolutional layers are combined with the network to make up for the detail loss of the density map conducted by stacked pooling. Compared with MCNN and CP‐CNN, the authors’ approach employs fewer branches and simpler architecture. Experimental result shows that their approach achieves MAE 80.7 and MSE 131.2 in ShanghaiTech PartA dataset, MAE 15.6 and MSE 26.8 in ShanghaiTech PartB dataset, and MAE Average 7.1 in WorldExpo'10 dataset.
Wei Zhang 0055, Yanyan Liu 0001, Jianghua Zhu
IET Image Process.1
2020 Multi-density map fusion network for crowd counting
Wei Zhang 0055, Yanyan Liu 0001, Jianghua Zhu
Neurocomputing2
2020 Two-branch fusion network with attention map for crowd counting
Wei Zhang 0055, Yanyan Liu 0001, Jianghua Zhu
Neurocomputing2
2019 Fast algorithm for HEVC intra-coding implemented by preprocessing
abstract
Compared to previous video compression standards H.264/AVC, high‐efficiency video coding (HEVC) introduces a recursive quad‐tree structure in intra‐coding and adds intra‐mode from 9 to 35 to achieve higher coding efficiency. The update significantly improves the performance of intra‐coding. However, it greatly increases the amount of computation, and this will increase the need for the fast algorithm and hardware design that can satisfy the real‐time encoding of HEVC encoder. In this study, the authors propose a fast intra‐coding algorithm based on analysis of original pixel gradient texture. The algorithm consists of two steps, fast coding unit size decision and reduction of candidate modes. The experimental results show that compared with the HM16.7 encoder, the optimisation algorithm can obtain 50.7% time reduction on average with 1.32% Bjontegaard distortion (BD)‐rate increase and 0.07 dB BD‐peak‐signal‐to‐noise ratio loss. Meanwhile, this study designs a pre‐processing hardware structure based on original pixels. With TSMC 90 nm complementary metal oxide semiconductor technology, the proposed structure can achieve 625 MHz working frequency at the cost of 2094 gates, which can fulfil the throughput requirement of 8K × 4K@46 fps real‐time encoding.
Weijia Mu, Wei Zhang 0055, Yanyan Liu 0001
IET Image Process.4
2017 Multiple channel error-correction algorithms for LCC decoding of Reed-Solomon codes and its high-speed architecture design
abstract
Reed‐Solomon (RS) code is one of the most widely used error control codes. The decoding algorithms of RS code can be briefly divided into hard‐decision decoding (HDD) and algebraic soft‐decision decoding (ASD). However, traditional RS decoding algorithms perform unsatisfactorily over bursty channel. Therefore, many modified RS decoding algorithms utilised in bursty channel decoding were proposed. However all of them are HDD and the computation is costly. In this study, the authors propose an ASD algorithm which can be utilised both in additive white Gaussian noise channel and bursty channel. In order to be utilised in multiple channel, the proposed algorithm combines original HDD‐based low‐complexity chase (HDD‐LCC) and burst‐error correcting (BC) together. What's more, BC is also modified by adding in the mechanism of pre‐judgement to reduce the iterations of BC. The modified algorithm can give a coding gain reaching up to 0.1268 and 1.33 dB compared with BC and HDD‐LCC, respectively when symbol error rate (SER) is in bursty channel. The proposed RS decoder is implemented and synthesised with Semiconductor Manufacturing International Corporation 0.13‐μm CMOS technology library. The results show the proposed decoder can operate at 200 MHz to achieve the throughput of 1.673 Gbps.
Wei Zhang 0055, Yanyan Liu 0001
IET Commun.2
2017 Hardware efficient multiplier-less multi-level 2D DWT architecture without off-chip RAM
abstract
This study presents a multi‐level 2D discrete wavelet transform (DWT) architecture without off‐chip RAM. Existing architectures use one off‐chip RAM to store the image data, which increases the complexity of the system. For one‐chip design, line‐based architecture based on modified lifting scheme is proposed. By replacing the multipliers with canonic sign digit multipliers, a critical path of one full‐adder delay is achieved. As per theoretical estimate, for three‐level 2D DWT with an image of N × N size, the proposed architecture requires 123 adders, 66 subtracters, 167 registers, temporal memory of 7.5 N words and input RAM of 3 N bytes. The estimated hardware requirement shows that for the image size of 512 × 512 and three‐level DWT, the proposed architecture involves at least 14.1% less transistor‐delay‐product than existing architectures.
Changkun Wu, Wei Zhang 0055, Yanyan Liu 0001
IET Image Process.2
2017 An Algorithm for Improving the Throughput of Serial Low-Complexity Chase Soft-Decision Reed-Solomon Decoder
abstract
A novel early termination algorithm (ETA) is proposed for the serial low-complexity chase (LCC) Reed-Solomon (RS) decoder. It can terminate the decoding procedure after the reencoding step conditionally, which results in the improvement of the serial LCC decoders' speed. The ETA will not degrade the decoding performance of RS codes; on the contrary, it can even improve that of short RS codes remarkably. The simulations of the ETA for (3125) and (255,239) RS codes over the additive white Gaussian noise (AWGN) channel and a (458,410) RS code over the EPR4 channel with 100% AWGN show that it is effective. Moreover, the hardware implementation of the early terminated LCC decoder for (458,410) and (255,239) RS codes is provided. The analysis shows that it can work faster than the previous serial LCC decoder with only a few extra hardware costs. It also shows great advantages in average throughput performance than the common pipelined LCC decoders, especially for relatively high Eb/N0 values, by an overall consideration of the area, speed, and decoding performance.
Haowen Luo, Wei Zhang 0055, Yanyan Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Compiler-Guided Parallelism Adaption Based on Application Partition for Power-Gated ILP Processor
abstract
Instruction-level parallelism (ILP) processors have been widely used to improve speed for several decades. However, the requirement of parallelism changes between applications, even within an application. Fixed high parallelism could result in poor utilization and extra leakage energy. Designing energy-efficient ILP processors to trade off power/speed has been a critical issue in current research. In this paper, a compiler-guided parallelism adaption based on an application partition algorithm is proposed to implement parallelism adaption with applications running on ILP processors. The aim is to minimize energy consumption without degrading the execution time. The main idea is described as follows: 1) partition the application into several power gating regions (PGRs); 2) assign adapted parallelism for each region by analyzing the requirements of resources and energy efficiency; and 3) reschedule each region with its own parallelism and insert power-gating instructions into the application to control hardware ON/OFF. The experimental results of evaluation with the CoreMarkPro benchmark suits show the expected savings of leakage energy. Our algorithm could reduce the leakage energy in register files by 30.46% and 64.06% for applications with high variance on software-inherent parallelism. Furthermore, the overhead energy originated from state transition is much lower than Tabkhi's algorithm.
Yufeng Tong 0002, Wei Zhang 0055, Yung-Cheng Ma, Yanyan Liu 0001, Tai Zhang, Haowen Luo
IEEE Trans. Very Large Scale Integr. Syst.2
2016 High-efficient Reed-Solomon decoder design using recursive Berlekamp-Massey architecture
abstract
This study presents a high‐efficient Reed–Solomon (RS) decoder based on the recursive enhanced parallel inversionless Berlekamp–Massey algorithm architecture. Compared with the conventional enhanced parallel inversionless Berlekamp–Massey algorithm architecture, the proposed architecture consists of a single processing element and has very low hardware complexity. It also employs a new initialisation to reduce the latency. This architecture uses pipelined Galois–Field multipliers to improve the clock frequency. In addition, the proposed architecture also has the dynamic power saving feature. The proposed RS (255, 239) decoder has been developed and implemented with SMIC 0.18‐μm CMOS technology. The synthesis results show that the decoder requires about 13K gates and can operate at 575 MHz to achieve the data rate of 4.6 Gb/s. The proposed RS (255, 239) decoder is at least 28.15% more efficient than the previously related designs.
Wenjie Ji, Wei Zhang 0055, Xingru Peng, Yanyan Liu 0001
IET Commun.2
2016 A pipelined Reed-Solomon decoder based on a modified step-by-step algorithm
abstract
We propose a pipelined Reed-Solomon (RS) decoder for an ultra-wideband system using a modified step-by-step algorithm. To reduce the complexity, the modified step-by-step algorithm merges two cases of the original algorithm. The pipelined structure allows the decoder to work at high rates with minimum delay. Consequently, for RS(23,17) codes, the proposed architecture requires 42.5% and 24.4% less area compared with a modified Euclidean architecture and a pipelined degree-computationless modified Euclidean architecture, respectively. The area of the proposed decoder is 11.3% less than that of the previous step-by-step decoder with a lower critical path delay.
Xingru Peng, Wei Zhang 0055, Yanyan Liu 0001
Frontiers Inf. Technol. Electron. Eng.2
2015 Efficient architecture for algebraic soft-decision decoding of Reed-Solomon codes
abstract
Reed–Solomon (RS) codes possess excellent error correction capability. Algebraic soft‐decision decoding (ASD) of RS codes can provide better correction performance than the hard‐decision decoding (HDD). The low‐complexity Chase (LCC) decoding has the lowest complexity cost and similar or even higher coding gain among all of the available ASD algorithms. Instead of employing complicated interpolation technique, the LCC decoding can be implemented based on the HDD. This study proposes a modified serial LCC decoder, which employs a novel syndrome calculation, polynomial selection, Chien search and Forney algorithm block. In addition, an improved two‐dimensional optimisation is provided to reduce the hardware complexity of the proposed decoder. Compared with the previous design, the proposed decoder can improve about 1.27 times speed and obtain 1.29 times higher efficiency in terms of throughput‐over‐slice ratio.
Wei Zhang 0055, Yanyan Liu 0001
IET Commun.2
2014 Deadline-Constrained Clustered Scheduling for VLIW Architectures using Power-Gated Register Files
abstract
Designing energy-efficient Digital Signal Processor (DSP) cores has become a key concern in embedded systems development. This paper proposes an energy-proportional computing scheme for Very Long Instruction Word (VLIW) architectures. To make the processor power scales with adapted parallelism, we propose incorporating distributed Power-Gated Register Files (PGRF) into VLIW to achieve a PGRF-VLIW architecture. For energy efficiency, we also propose an instruction scheduling algorithm called the Deadline-Constrained Clustered Scheduling (DCCS) algorithm. The algorithm clusters the data dependence graph to reduce data transfer energy and makes optimal use of low-powered local registers for tree-structured data dependence graphs. The results of evaluations conducted using the MiBench and DSPstone benchmark suites substantiate the expected power saving and scaling effects.
Zhibin Liang, Wei Zhang 0055, Yung-Cheng Ma
ACM Trans. Archit. Code Optim.2
2013 Low-power design of Reed-Solomon encoders
abstract
Reed-Solomon (RS) codes are one of the most widely used block error-correcting codes in modern communication and computer systems. Multiplication is the key computation in RS encoding. Adopting the generator polynomial with symmetric coefficients, the number of multipliers in RS encoders can be reduced by half, and their power consumption may also reduce. However, in some cases, the encoder based on the generator polynomial with asymmetric coefficients have better power performance. Additionally, since more than one primitive polynomial can generate a finite field with certain order, different choices of primitive polynomial also change the complexity of multipliers. In this paper, we exploited the relationship between the power consumption of RS encoders and their different encoding parameters. A simple way to find the encoder with the lowest power consumption is also presented. Simulation results prove its effectiveness.
Wei Zhang 0055, Xinmiao Zhang 0001
ISCAS1
2013 Reduced-Complexity LCC Reed-Solomon Decoder Based on Unified Syndrome Computation
abstract
Reed-Solomon (RS) codes are widely used in digital communication and storage systems. Algebraic soft-decision decoding (ASD) of RS codes can obtain significant coding gain over the hard-decision decoding (HDD). Compared with other ASD algorithms, the low-complexity Chase (LCC) decoding algorithm needs less computation complexity with similar or higher coding gain. Besides employing complicated interpolation algorithm, the LCC decoding can also be implemented based on the HDD. However, the previous syndrome computation for 2ηtest vectors and the key equation solver (KES) in the HDD requires long latency and remarkable hardware. In this brief, a unified syndrome computation algorithm and the corresponding architecture are proposed. Cooperating with the KES in the reduced inversion-free Berlekamp-Messy algorithm, the reduced-complexity LCC RS decoder can speed up by 57% and the area will be reduced to 62% compared with the original design for η = 3.
Wei Zhang 0055, Hao Wang 0072, Boyang Pan
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Modified polynomial selection architecture for low-complexity chase decoding of Reed-Solomon codes
abstract
Reed-Solomon (RS) codes are widely used in modern communication and computer systems. Compared with the hard-decision decoding algorithms, the algebraic soft-decision decoding (ASD) algorithm can achieve significant coding gain. Among ASD algorithms, the low-complexity Chase (LCC) decoding has a better performance and lower complexity. In the LCC decoding, 2ηtest vectors need to be interpolated and a polynomial selection scheme is required to choose the right interpolation output. A modified polynomial selection (MPS) algorithm is proposed in this paper. By deleting the reliability information, the MPS requires less hardware and provides the same performance as its present counterpart. For a (63, 55) RS code over GF (26), the MPS can save 20% chip area and 21.2% power consumption.
Hao Wang 0072, Wei Zhang 0055, Boyang Pan
ISCAS2