Wen Bin Ye 0001

dblp:56/9848 · also Wenbin Ye 0001 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-6978-813XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 4 first-author · 9 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Memory-Efficient In-Sensor Event Denoising with a Lightweight Point-Cloud Network
Zongpei Fu, Xiaojin Zhao, Wen Bin Ye 0001
ISCAS3
2026 A Wavelet-Enhanced Neural Network with Knowledge Distillation for MCU-Based Fingerprint Liveness Detection
Zhengwu Li, Kaixiang Lin, Xiaojin Zhao, Wen Bin Ye 0001
ISCAS4
2026 A Low-Cost Energy-Efficient FFT-CNN Processor for Versatile mmWave Radar in IoT Applications
abstract
The millimeter-wave (mmWave) radar Internet of Things (IoT) applications based on Convolutional Neural Network (CNN) algorithms are widely employed across various industries. However, they face challenges in deployment on low-cost and resource-constrained edge devices. Current CNN accelerators can efficiently handle CNN computations but cannot effectively accelerate the essential fast Fourier Transform (FFT) computations in mmWave radar signal processing, thus failing to achieve end-to-end high-efficiency radar signal processing. This work presents a low-cost energy-efficient FFT-CNN processor with a unified architecture that combines CNN and FFT to enhance performance for mmWave radar IoT applications. The proposed FFT-CNN processor employs a reconfigurable dual-mode processing element to maximize hardware resource sharing, supporting both butterfly operations for FFT and multiply-accumulate computations for CNN acceleration. This sharing technique reduces hardware resource consumption by 22.1% in LUTs, 22.7% in FFs, 9.5% in BRAMs, and 16.7% in DSPs, leading to a total power reduction of 34.9%. Compared to state-of-the-art processors optimized for radar applications, the proposed processor achieves the highest performance-to-resource ratio and energy efficiency, making it ideal for a wide range of low-cost radar IoT applications.
Juhua Chen, Zongpei Fu, Wen Bin Ye 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2026 A 95.3% 12-Class, 108-nJ/Inference Keyword Spotting Chip With Hybrid FFT-BFNet Architecture and Exponent-Aware Nonuniform Quantization in 65-nm CMOS
abstract
This article presents a 65-nm keyword spotting (KWS) chip that achieves 95.3% accuracy on 12-class tasks with 108.04-nJ/inference efficiency through cross-domain hardware-algorithm innovations. The unified fast Fourier transform (FFT)-butterfly-structured neural network (BFNet) accelerator fundamentally rethinks computational reuse: by replacing dense pointwise convolutions with butterfly-based sparse operations mirroring FFT’s dataflow, it slashes$6.3\times $multiply-accumulate (MAC) operations and halves parameter counts while preserving model expressivity. A 6-bit exponent-aware nonuniform quantization (EANUQ) scheme compresses weights, achieving a 25% reduction in storage while maintaining an accuracy loss of less than 0.01% with lightweight on-chip decoders. Hardware resource sharing extends beyond computation: Mel-filter-banks reuse fully-connected (FC) layer multipliers through decomposed 8-bit arithmetic, and FFT output buffers double as convolutional neural network (CNN) feature map memory. Measured at 0.65 V/600 kHz, the 0.58-$\text {mm}^{2}$core demonstrates$1.9\times $–$15.5\times $better energy efficiency than prior 65/28-nm implementations, with 14.82-ms end-to-end latency.
Zongpei Fu, Kaixiang Lin, Xiaojin Zhao, Wen Bin Ye 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Hardware-Algorithm Codesigned Low-Latency and Resource-Efficient OMP Accelerator for DOA Estimation on FPGA
abstract
This article introduces an algorithm-hardware codesign optimized for low-latency and resource-efficient direction-of-arrival (DOA) estimation, employing a refined orthogonal matching pursuit (OMP) algorithm adept at handling the complexities of multisource detection, particularly in scenarios with closely spaced signal sources. At the algorithmic level, this approach incorporates a secondary correction mechanism (SCM) into the traditional OMP algorithm, significantly improving estimation accuracy and robustness. On the hardware front, a bespoke OMP accelerator has been developed, featuring a reconfigurable generic processing element (PE) array that supports various computational modes and leverages multilevel spectral peak search strategy and pipelining techniques to enhance computational efficiency. Experimental evaluations reveal that the proposed system achieves a root mean square error (RMSE) for DOA estimation of less than 0.3° in multisource conditions with a signal-to-noise ratio (SNR) of 20 dB. In addition, the deployment of the OMP accelerator on a Zynq XC7Z020 development board utilizes modest logic resources: 5.49k LUTs, 3.28k FFs, 11.5 BRAMs, and 32 DSPs. Furthermore, the design achieves a computational latency of$2.83~\mu \text { s}$for single-source estimation with eight antennas. This achievement reflects a reduction of approximately 17.8% in LUTs, 56.3% in FFs, and 5.7% in DSPs compared to current leading-edge technologies after normalization all while maintaining competitive estimation accuracy and favorable estimation rates.
Ruichang Jiang, Wen Bin Ye 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2024 A FPGA-based Energy-Efficient Processor for Radar-based Continuous Fall Detection
abstract
This paper proposes a FPGA-based energy-efficient processor for radar-based continuous fall detection, consisting of a spectrogram generation circuit and a convolutional neural network (CNN). To reduce resource consumption and power usage of the entire processor, two designs were implemented: 1) A serial-FFT-based spectrogram generation circuit for radar signal preprocessing, and 2) An neural network (NN) accelerator based on row-stationary dataflow has been designed. By using the updated block wise computation technique, the accelerator enables the computation of only the NN’s updated inputs, resulting in an 85% reduction in multiply-accumulate (MAC) operations and a 79% reduction in intermediate result storage. Implemented on the Xilinx FPGA board ZC702, this processor consumes only 7.3k LUTs, 3.6k Flip Flops (FFs), 22 Block RAMs (BRAMs), and 10 DSPs, with a power consumption of only 0.302W. On an open-source radar fall detection dataset, this processor achieves an accuracy of 99.58%. The processor incurs a delay of only 0.431ms for a single preprocessing and NN inference, consuming just 130.16 µJ.
Juhua Chen, Linxin Yang, Wen Bin Ye 0001
ISCAS3
2024 An Energy-Efficient Edge Processor for Radar-Based Continuous Fall Detection Utilizing Mixed-Radix FFT and Updated Blockwise Computation
abstract
In the scenarios of the Internet of Things, fall detection holds increasing significance in the health monitoring of elderly individuals. While most current research has achieved impressive performance in fall detection methods, there are limitations in deploying these methods to resource-limited edge devices. This article proposes an energy-efficient edge processor for radar-based continuous fall detection, which consists of a preprocessing module and a convolutional neural network (NN) accelerator. Multiple designs were implemented to minimize resource utilization and power consumption of the entire processor: 1) a preprocessing module based on mixed-radix FFT is utilized for radar signal preprocessing and 2) an NN accelerator is designed to support an updated blockwise (UBwise) computation technique aimed at reducing redundant calculations and intermediate result storage in continuous fall detection, along with a fully connected (FC) layer cache compression technique proposed to compress the cache required for FC layer computations. Applying these techniques results in an 80% reduction in RAM size, an 88.6% decrease in intermediate result storage, and a 92.6% reduction in multiply-accumulate operations. Implemented on an FPGA, this processor consumes merely 3.1k look up tables, 2.3k flip-flops, four block RAMs, and seven DSPs while consuming only 0.234 W of power. On an open-source radar-based fall detection data set, the processor attains an accuracy of 98.58%. Additionally, it incurs a mere 42 us delay for a single preprocessing and NN inference, consuming just 9.8 uJ. Compared to state-of-the-art works, this processor’s energy consumption is reduced by 81.2%, and the required memory is reduced by 93.3%.
Juhua Chen, Kaixiang Lin, Linxin Yang, Wen Bin Ye 0001
IEEE Internet Things J.4
2024 A 593nJ/Inference DVS Hand Gesture Recognition Processor Embedded With Reconfigurable Multiple Constant Multiplication Technique
abstract
Hand gesture recognition (HGR) is a popular technique for edge-based human-computer interaction. Dynamic vision sensors (DVS) are often used in HGR systems due to their low latency, high dynamic range, low energy consumption, and asynchronous event triggering. While Spiking Neural Networks (SNNs) are commonly thought to consume less energy than Convolutional Neural Networks (CNNs) in DVS-based HGR systems, this work demonstrates that a DVS-based HGR system on chip (SoC) incorporating CNN can achieve lower power consumption through algorithm and hardware co-design. The proposed edge-side processor for DVS-based HGR integrates a median filter processing core and an AI accelerator core for preprocessing and CNN inference on the DVS output data. To reduce hardware costs without sacrificing accuracy, the median filtering core uses a simplified median filtering function tailored to the specific application scenario. The paper suggests using reconfigurable multiple constant multiplication (RMCM) techniques for the AI accelerator core to share computational resources among processing element (PE) arrays, thereby reducing computational costs and power consumption. The entire DVS gesture processor was implemented in a 65nm technology, achieving an energy requirement of 593.4nJ per inference on-chip with a guaranteed accuracy of 92.4%.
Zongpei Fu, Wen Bin Ye 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 A Hardware and Software Co-Design for Energy-Efficient Neural Network Accelerator With Multiplication-Less Folded-Accumulative PE for Radar-Based Hand Gesture Recognition
abstract
This work presents a novel lightweight neural network (NN) model and a dedicated NN accelerator for radar-based hand gesture recognition (HGR). The NN model employs symmetric weights, group 1-D-convolution, and power-of-two (POT) quantization, achieving 92.84% accuracy on a public dataset with only 4.8 k parameters, while reducing parameter storage by 40%. The custom accelerator features a multiplication-less folded-accumulative processing element (PE), group-wise computation optimization, and an efficient scheduling mechanism for fully connected (FC) layers. Implemented on a Xilinx field-programmable gate array (FPGA) board XC7S15 and 65-nm CMOS technology, it surpasses existing solutions in power efficiency and cost-effectiveness, addressing the computational demands for IoT deployment.
Yunqi Guan, Wen Bin Ye 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2023 Lightweight Deep Learning Model for Radar-Based Fall Detection With Metric Learning
abstract
The radar-based fall detection system has grown in popularity because of its stability and privacy protection. Deep neural networks have been used in previous radar-based fall detection systems to improve detection accuracy. However, most of them need a lot of memory and have high computational complexity, making them impractical for Internet of Things (IoT) devices. In this article, we propose an extremely lightweight network named Tiny-RadarNet for extracting characteristics from raw data. Unlike traditional neural networks, we use a unique parallel 1-D depthwise convolutions structure as the core module to eliminate the need for standard convolutions, and achieve significant parameter reduction. Furthermore, instead of regarding fall detection as a classification problem, we use our metric learning technique to treat it as a matching problem for better distinction of embeddings. Finally, we introduce a novel dual loss function to improve the proposed network’s robustness against unobserved human motions without the need for additional anchors or expensive computation. The experimental results reveal that the proposed method can achieve comparable fall detection accuracy compared with state-of-art methods but with much fewer weight parameters and lower computational complexity. These findings suggest that a low-power and low-latency fall detection solution for IoT applications is achievable.
Zixuan Ou, Wen Bin Ye 0001
IEEE Internet Things J.2
2023 An Efficient Algorithm-Hardware Co-Design for Radar-Based Fall Detection With Multi-Branch Convolutions
abstract
In this paper, we propose an efficient algorithm-hardware co-design framework to realize radar-based fall detection with limited resources. We first design a compact neural network model named MB-Net with multi-branch convolutions for feature extraction of radar time series data combined with multi-scale wavelet transform. After that, an FPGA-based neural network (NN) accelerator tailored for the proposed network is designed. The proposed NN accelerator replaces the general multipliers with non-exact multipliers to reduce the hardware cost. For the multi-branch convolution layer, a novel layer computing sequence is introduced to improve the efficiency of the processing element (PE) array and reduce the memory footprint. In addition, the average pooling operation in the proposed network is folded into the quantization factors to reduce hardware cost. The experimental findings show that the MB-Net can maintain competitive performance in comparison to state-of-the-art methods while the hardware cost is significantly lower. The proposed network model is implemented in Zynq ZC702 board using only 3615 LUTs, 1843 FFs, 11.5 BRAMs, and 8 DSPs with 0.234 W power consumption. Through algorithm and hardware co-optimization, the fall detection accelerator can achieve 95% PE efficiency and takes 0.346 ms latency for a radar sample interference with only 80.96 uJ energy consumption.
Zixuan Ou, Wen Bin Ye 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Radar-Based Human Activity Recognition With 1-D Dense Attention Network
abstract
With the development of the Internet of things, radar-based human activity recognition is becoming more and more important, because they play an indispensable role in fields such as safety and health monitoring. In this work, a novel network named 1-D dense attention neural network (1-D-DAN) is proposed for the radar-based human activity recognition. In the proposed network, a novel attention mechanism network structure specifically designed for radar spectrogram is proposed, equipping 1-D convolutional network with attention mechanism. With the$x$-axis of the spectrogram represents time and the$y$-axis represents frequency, the proposed attention mechanism includes two branches: 1) time attention branch and 2) frequency attention branch. Moreover, a dense attention operation that can make full use of features in the network is also introduced in the proposed attention mechanism. Experimental results show that compared with the state-of-the-art methods, our proposed 1-D-DAN achieves the highest accuracy in human activity recognition with the lowest computational complexity.
Guoji Lai, Xin Lou 0001, Wen Bin Ye 0001
IEEE Geosci. Remote. Sens. Lett.3
2021 Lightweight Deep Learning Model in Mobile-Edge Computing for Radar-Based Human Activity Recognition
abstract
Radar-based human activity recognition (HAR) has great potential in many fields, such as surveillance, smart homes, and human-computer interaction. Complex deep neural networks have brought significant improvement in classification performance but also a surge of computational cost and the number of parameters, which makes it challenging to deploy in mobile devices. However, the existing studies in this area mainly focus on improving the classification accuracy. In this article, we propose an extremely efficient convolutional neural network (CNN) architecture named Mobile-RadarNet, which is specially designed for human activity classification based on micro-Doppler signatures. The new architecture exploits 1-D depthwise convolutions and pointwise convolutions to build lightweight CNN architecture. The experiments on a seven-class human activity data set demonstrate that the proposed Mobile-RadarNet can achieve high classification accuracy meanwhile to keep the computational complexity at an extremely low level, and thus has great potential to be deployed in the mobile devices.
Xin Lou 0001, Wen Bin Ye 0001
IEEE Internet Things J.3
2020 Classification of Human Activities Based on Radar Signals using 1D-CNN and LSTM
abstract
Many deep learning models have been proposed in radar-based human activity recognition (HAR) area. For radar-based HAR, generally, the raw radar data is first converted to a 2-D spectrogram by using short-time Fourier transform(STFT). All the existing DL models adopt 2-D convolutional neural networks as they treat the 2-D spectrogram the same as an optical image. In this paper, for the first time, the radar spectrogram is treated as a time-sequential vector, and a DL model composed of 1-D convolutional neural networks (1D-CNNs) and recurrent neural networks (RNNs) is proposed. The experiment results show that the proposed model can not only achieve the highest accuracy but also have the fewest number of parameters than that of existing 2-D CNN methods.
Wen Bin Ye 0001
ISCAS3
2020 Classification of Human Activity Based on Radar Signal Using 1-D Convolutional Neural Network
abstract
Previously, the 2-D convolutional neural networks (2-D-CNNs) have been introduced to classify the human activity based on micro-Doppler radar. Whereas these methods can achieve high accuracy, their application is limited by their high computational complexity. In this letter, an end-to-end 1-D convolutional neural network (1-D-CNN) is first proposed for radar-based sensors for human activity classification. In the proposed 1-D-CNN, the inception densely block (ID-Block) tailored for the 1-D-CNN is proposed. The ID-Block incorporated the three techniques: inception module, dense network, and network-in-network techniques. With these techniques, the proposed network not only achieve a high classification accuracy but also keep the computational complexity at a low level. The experiments results show that the classification accuracy of the proposed method is 96.1% for human activity classification that is higher than that of existing state-of-art 2-D-CNN methods while the computational speed of forward propagation is increased by about ($2.71\times $ to $29.68\times $ ) of the existing 2-D-CNN methods.
Wen Bin Ye 0001
IEEE Geosci. Remote. Sens. Lett.2
2019 Automated Detection of Diabetic Retinopathy using a Binocular Siamese-Like Convolutional Network
abstract
Diabetic retinopathy (DR) is an important causes of blindness worldwide. It is hard to detected in early stages and the diagnostic procedure can be time-consuming. Therefore, we proposed a deep learning method to automatedly diagnose the referable deiabetic retinopathy (RDR) by classifying retinal fundus photographs. In our work, A novel convolutional neural network with Siamese-like architecture is trained with transfer learning technique. Different from previous works, the proposed model accepts binocular fundus images as input and learns their correlation to aid the prediction. A custom loss function combining cross entropy loss and contrastive loss is also adopted to guide the gradient descent. With a training set of only 28104 images and a test set of 3510 images, an area under the receiver operating curve (AUC) of 0.949 is obtained through the proposed method. Furthermore, the sensitivity reaches 80.7% when the specificity is fixed at 95.0%, and the specificity reaches 70.3% when the sensitivity is fixed at 95.0%, which are 3.2% and 4.6% higher than those obtained by the existing Inception V3 model trained with monocular fundus images, respectively.
Wen Bin Ye 0001
ISCAS4
2018 An Energy-Efficient 13-bit Zero-Crossing ΔΣ Capacitance-to-Digital Converter with 1 pF-to-10 nF Sensing Range
abstract
Conventional capacitance-to-digital converters (CDCs) suffer limitations either on narrow capacitance range or low resolution for jitter-induced noise and high power consumption. In order to overcome these limitations, a 13-bit 1 pF-to-10 nF generic CDC is presented. In the proposed CDC with the oversampled ΔΣ modulation, the zero-crossing-based circuits (ZCBCs) are used to replace the operational transconductance amplifier to avoid feedback loop stability issues. However, the ZCBCs inevitably incur the non-idealities and thus a novel calibration scheme is presented for efficient non-ideality-error cancellation. A prototype fabricated using 0.18 μm CMOS technology is experimentally verified using a MEMS capacitive humidity sensor. The measurement results show the CDC achieves a 13-bit rms noise equivalent resolution with a 128 μs conversion time and a 230 fJ/conversion-step figure of merit. The calibration scheme enhances the linearity from 7 bits to 11.4 bits in the 1 pF-to-10 nF compatible capacitance range.
Bing Li 0011, Wei Wang 0165, Jia Liu 0011, Wen Bin Ye 0001
ISCAS7
2018 A 480 fJ/conversion-step 13-bit Resistive Sensor Readout IC with a 1%/V Power Supply Sensitivity
abstract
Conventional resistive sensor readout ICs (ROICs) suffer limitations, either because of their high resistive sensor-biasing current consumption or low resolution for jitter-induced noise. To overcome these limitations, a 13-bit resistive sensor ROIC is presented in this work. In the proposed ROIC, which relies on zero-crossing-based circuits (ZCBCs), the current source is used not only to settle the output of the ΔΣ integrator, but also to bias the resistive sensor, and thus avoid implementing a dedicated power-hungry bias current branch. In addition, the oversampled ΔΣ modulation technique guarantees high resolution for the proposed ROIC. A prototype was fabricated using 0.18 μm CMOS technology. The measurement results show the ROIC achieved a 13-bit root-mean-square noise equivalent resolution with an 8 kS/s conversion speed and a 480 fJ/conversion-step figure of merit. Moreover, it has good nonlinearity of less than 250 ppm and a power supply sensitivity of 1%/V.
Bing Li 0011, Ji-Ping Na, Wei Wang 0165, Jia Liu 0011, Wen Bin Ye 0001
ISCAS7
2018 K-SVD Based Denoising Algorithm for DoFP Polarization Image Sensors
abstract
This paper presents a novel K times singular value decomposition (K-SVD) based denoising algorithm for the division-of-focal-plane (DoFP) polarization image sensors. In the proposed implementation, the input DoFP image can be expressed by the optimum sparse combination of the dictionary elements via K-SVD and orthogonal matching pursuit (OMP) algorithms. As a result, this implementation is capable of eliminating the Gaussian noise significantly and well-preserving the details and edges of the target DoFP image. Our extensive experimental results on various test images show that the proposed algorithm yields better visual quality and maintains a lower PSNR value while compared with a wide range of previous implementations.
Shiting Li, Wen Bin Ye 0001, Huawei Liang, Xiaofang Pan, Xin Lou 0001, Xiaojin Zhao
ISCAS2
2017 A passively compensated capacitive sensor readout with biased varactor temperature compensation and temperature coherent quantization
abstract
This paper presents a frequency-mode capacitive sensor readout front-end operating in a wide temperature range. First, a biased varactor temperature compensation (BVTC) is proposed to compensate the aggregate temperature gradients from the sensor and the oscillator circuit, achieving a nullified temperature coefficient for the oscillation frequency. Second, a temperature coherent quantization (TCQ) approach is proposed to enhance the sensor' sensitivity and to provide a self-referenced clock for digitization, whereby the influence from the temperature effect of the clock is minimized through a hybrid down-conversion time-to-digital converter (DTDC). A prototype chip was fabricated using the 0.18-μm CMOS process and it was verified using a commercial 5.92-to-6.53-pF capacitive pressure sensor over a temperature range of −20°C to 120°C. Proven by experiments, the prototype presented as good as ±0.4% full-scale pressure error in the 140°C temperature range.
Wang Ling Goh, Kevin Tshun Chuan Chai, Xin Lou 0001, Wen Bin Ye 0001
ISCAS6
2017 A novel smoothness-based interpolation algorithm for division of focal plane Polarimeters
abstract
In this paper, we present a novel smoothness-based interpolation algorithm for the division of focal plane Polarimeters (DoFP). By calculating the divided blocks' variance that represents their local smoothness, the proposed algorithm well-balances between the traditional bilinear and bicubic interpolation algorithms. In addition, compared with the previously reported gradient-based interpolation algorithm which only indicates the image's directional change along 0°, 90°, 45° and 135°, the presented smoothness-based algorithm covers the variations along all the possible directions, leading to more accurate selection between the bilinear and bicubic algorithms. According to our extensive simulation results, the proposed implementation exhibits the lowest mean square error (MSE) for the test images among all the previously reported algorithms, including bilinear, bicubic and gradient-based interpolation algorithms.
Jieyun Zhang, Wen Bin Ye 0001, Ashfaq Ahmed, Zhurui Qiu, Yuan Cao 0003, Xiaojin Zhao
ISCAS2
2016 An energy-efficient subthreshold level shifter with a wide input voltage range
abstract
The level shifters are crucial primitives in the multi-supply voltage circuits and systems. In this paper, an energy-efficient level shifter is proposed to achieve the conversion from the subthreshold voltage to the above threshold voltage. It is a hybrid structure consisting of the Wilson current mirror and the cross-coupled level shifter. By addressing the voltage drop issue of the level shifter based on the Wilson current mirror, the leakage power is significantly reduced, with the advantage of wide input voltage range for the Wilson current mirror level shifter well-preserved. In addition, the multi-threshold CMOS (MTCMOS) technology is employed to provide more flexibility for our ultra-low power design. The reported simulation results using 65 nm CMOS process validate our proposed implementation and an ultra-low power consumption of 19.44 fJ per conversion from 0.2 V to 1.2 V at 1 MHz is achieved without the need of any intermediate power supply.
Yuan Cao 0003, Wen Bin Ye 0001, Xiaojin Zhao, Peigang Deng
ISCAS2
2015 Design of high-speed multiplierless linear-phase FIR filters
abstract
In the design of multiplierless FIR filters, researchers have made every effort to reduce the number of adders when coefficients multipliers are realized using adder-and-shift network to decrease the overall chip area. However, with the advance of IC technology, area becomes a less important issue than the speed. In this paper, we propose a speed oriented optimization of linear phase FIR filters, where the length of critical path is used as the criteria in the discrete coefficient search. The length of critical path is measured as the number of cascaded full adders rather than the traditional adder depth. Compared to the area oriented algorithm, the proposed algorithm can generate the filters with much shorter critical path delay and meanwhile the area-delay product is also reduced. Gate level simulations of benchmark filters verify the above claim.
Wen Bin Ye 0001, Xin Lou 0001, Ya Jun Yu
ISCAS1
2014 A polynomial-time algorithm for the design of multiplierless linear-phase FIR filters with low hardware cost
abstract
Deterministic tree search algorithm for the design of multiplierless linear phase finite impulse response filters are usually time consuming. More researches therefore focus on how to restrict the number of discrete values assigned to each coefficient during a tree search. This paper proposes a polynomial-time tree search algorithm where each coefficient is fixed to only one discrete value. Due to the short search time, multiple searches with floating passband gain become possible, and each search may produce a feasible discrete solution. In such a way, if the whole floating passband gain range is partitioned into many smaller ones, a large amount of feasible discrete solutions may be obtained. In the end, all these feasible solutions are synthesized using an multiple constant multiplication algorithm, and the one using the least number of adders is the final design results. To accelerate the search, a low hardware cost scheme which fixes some coefficient values to 0 prior to each search is proposed. With these techniques, design examples show that the proposed algorithm significantly outperforms the existing algorithms in terms of design time while the hardware cost is kept low.
Wen Bin Ye 0001, Ya Jun Yu
ISCAS1
2013 Sparse FIR filter design based on Genetic Algorithm
abstract
Sparse patterns for digital filters have been suggested to reduce the computational cost. However, the minimization of the number of non-zero coefficients under required filter order and frequency domain constraints is difficult to be accomplished in polynomial time in many cases. In this paper, a two-stage design based on Genetic Algorithm (GA) is proposed to search for the optimal solution with a given specification. A preliminary optimization stage is introduced to enhance the efficiency of the proposed GA. The proposed algorithm is evaluated through two sets of examples, which generate better results than existing algorithms.
Heng Zhao 0004, Wen Bin Ye 0001, Ya Jun Yu
ISCAS2
2012 Design of high order and wide coefficient wordlength multiplierless FIR filters with low hardware cost using genetic algorithm
abstract
In this work, a novel genetic algorithm (GA) is proposed for the design of multiplierless finite impulse response (FIR) filters with high filter order and wide coefficient wordlength. GA mimics the nature evolution to optimize complicated problems and in theory optimum solutions can be obtained with infinite computation time. However, in practical filter design problem, when the filter specification is stringent, requiring high filter order and wide coefficient wordlength, GA often fails to find feasible solutions, because the discrete search space thus constructed is huge and majority of the solution candidates therein can not meet the specification. In the proposed GA, the discrete search space is partitioned into smaller ones. Each of the small space is constructed surrounding an optimum continuous solution with a floating passband gain. This increases the chances for the GA to find feasible solutions, but not sacrificing the coverage of the search space. In addition, the search in the multiple spaces can run in parallel, and thus the computation time for the design of filters under consideration reduces significantly. Design examples show that the proposed GA outperforms existing algorithms dealing with the similar problems.
Wen Bin Ye 0001, Ya Jun Yu
ISCAS1
2011 Switching activity analysis and power estimation for multiple constant multiplier block of FIR filters
abstract
The switchings in the multiple constant multiplier (MCM) block of FIR filters are heavily affected by the spatial correlation of inputs at each computation unit-adders in this case. In this work, a new switching activity model for the MCM blocks is developed by taking the spatial correlation of inputs of each adder into consideration. Based on this model, a power measurement is proposed to estimate the power consumption of the MCM blocks of FIR filters. The dynamic power simulation of several benchmark filters shows that the accuracy of our model outperforms that of existing models such as Glitch Path Count, Glitch Path Score and Power Index.
Wen Bin Ye 0001, Ya Jun Yu
ISCAS1