EDBT 2026 Demo / reviewers in the wild / expert
Weiwei Shi 0001
dblp:44/8271-1
· DBLP profile ↗
7ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0003-4551-8420ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Approximate Multipliers With Merits Balance for Digital Processing and Neural NetworksabstractIn this article, hybrid approximate multiplier (HAMs) designs based on the combination of logarithmic multiplication and piecewise linear (PWL) fitting are proposed. After extracting the exponent and mantissa of the input operands, two new variables are introduced to perform spatial mixed linear fitting on the 3-D surface of the mantissa product in different regions. Limited power-of-2 elements in line slopes make the multivariable mixed PWL computational simple and friendly to logic circuit complexity. With iterative adjustment of the slopes and bias of the lines in the PWL calculation, the relative error distance (RED) distribution is well balanced and zero concentrated. In addition, we detail the logic architecture to implement approximate hybrid accumulation and error-tolerant complement conversions. In the 45-nm library-based performance comparison, the proposed multipliers—mainly 16-, 8-, and 32-bit floating-point multipliers—exhibit >55% power, >23% delay, and >43% area reductions compared with the exact multiplier. In addition, they outperform other state-of-the-art designs in terms of delay, power, area, and error, as evaluated by the joint delay–power–area product (PPA) and mean RED (MRED). In case experiments, the proposed multipliers perform nearly equivalently to the exact multiplier in error-tolerant digital processing and neural network computations. Weiwei Shi 0001, Xiaocong Cao, Zhuoliang Zou, Yuping Gao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Design of a CNN Accelerator for Multitask EEG Signal Classification Based on RISC-V
Weiwei Shi 0001, Haoren Qin, Jingbin Mai |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | An Area-Energy-Efficient 64-2048 Point FFT With Approximate Plane-Fitting Complex MultipliersabstractAs a key component of fast Fourier transform (FFT), the complex multiplier (CM) includes twiddle factor generation and corresponding multiplication. This brief proposes an tailored approach for approximating CM functionality by employing an adapted piecewise-plane-fitting technique, effectively replacing the conventional look-up-table-based twiddle generation and exact multipliers by shift-and-add calculation. Numerical binary calculation analysis and simulations are conducted to achieve an optimal tradeoff among accuracy, circuit complexity, power, and delay. Based on 45-nm CMOS, logic synthesis results demonstrate significant improvements, with area, power, and delay reductions of 64.18%, 64.98%, and 19.77%, respectively. With optimizations on logic structures, the complete design of the 64–2048 point FFT has efficiently adopted the proposed CM with evident improvement. The proposed FFT outperforms other reconfigurable FFT designs in terms of normalized area reduction over 55.53% and normalized energy improvement over 21.51%. In field-programmable gate array (FPGA) implementation, the proposed FFT has significantly more savings compared with the exact FFT. In practice, the approximate FFT output results’ PSNR ranges from 56 to 83 dB with competent accuracy in typical signal processing. Weiwei Shi 0001, Yida Yuan, Zhihong Mo, Chaoyuan Wu, Jiangwei He |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Area-Delay-Energy-Efficient Approximate Dividers Based on Piecewise Linear Fitting of SurfaceabstractIn this paper, signed and unsigned approximate dividers based on piecewise linear fitting of surface (PLSAD) are proposed. After extracting the exponent and mantissa of two binary fixed-point operands, the surface representing the quotient of the mantissa is partitioned into sub-regions, which are linearly approximated by rectangular planes. The proposed logic architecture for geometric plane computations relies on simple arithmetic operations of addition and shifting. Additionally, approximate adders reduce circuit complexity and enhance performance, while enabling dynamic configuration of dividers based on accuracy requirements. Fairly compared to recently published approximated dividers based on the same 45nm technology and logic synthesis constrains, the proposed circuits outperform other works in overall aspects. ‘As the bit-width of OR-gate additions in Least significant bit (LSB) increases, the mean relative error distance (MRED) gradually grows from 0.82% to 2.63%, while the circuit area, delay and power consumption reduce to 840um2, 0.9ns and 881uW, respectively. Similarly, the proposed approximate floating-point (FP) 16-bit and 32-bit dividers improve circuit delay and power by over 50% compared to the state-of-the-art designs in the same condition. The proposed dividers perform well in image processing and demonstrate high adaptability in softmax layer of deep neural network and real-time EEG signal recognition applications. Chaoyuan Wu, Weiwei Shi 0001, Yida Yuan, Zhuoliang Zou, Zhihong Mo, Jiangwei He |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Arithmetic and Logic Circuits Based on ITO-Stabilized ZnO TFT for Transparent ElectronicsabstractIn this paper, general logic cell and module designs in basic digital signal processing (DSP), for transparent, flexible chips and wearable electronics are presented. Modified Circuits from ratioed logic and pass transistor logic styles are modified and proposed, based on n-type-only indium tin oxide (ITO) stabilized ZnO thin-film transistors (TFTs) process. Elaborations on logic circuits with purely n-type TFT transistors were carried out to extend the logic swing of circuit outputs and accelerate the signal propagation in complex logic functions with simplified pull-up/pull-down or passive networks. To implement logic complexes on multiple OR-of-ANDs functions with better performance, tailored active controlled ratioed logic style is adopted with faster speed and smaller area by simplifying the transistor networks properly. All the circuits were fabricated on transparent glass plate. The featured 2-input/3-input XOR gate, 4-1 MUX and D flip flop (DFF) can operate with maximum featured delays of 2-$6~\mu \text{s}$at 5 V power supply, with comparatively over 50% less delays and areas than other state-of-art works. The featured 4-bit adder performs maximum$16.41~\mu \text{s}$delay at 5 V in measurement. A 4-bit multiplier is also presented based on the adder and DFF. These proposed TFT circuits exhibited smaller area with relatively moderate-high performance in comparison, which were promising building blocks for transparent flexible DSP electronics with low-speed requirements. Weiwei Shi 0001, Lizhi Hu, Yuan Liu 0022, Sunbin Deng, Yuming Xu, Hoi-Sing Kwok, Rongsheng Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2018 | A Subthreshold Baseband Processor Core Design With Custom Modules and Cells for Passive RFID TagsabstractSophisticated subthreshold passive radio frequency identification tag's baseband processor (BBP) core design for ultralow-power Internet of Things end devices is presented in this paper. Custom logic cells and tailored logic architectures are applied to eliminate timing violations when the operating voltage is much lower than nominal level. For the consideration of limited availability of radio frequency power, power-aware scheme is applied to the key modules, including PIE decoding and command receiving. Furthermore, Galois linear feedback shift register and double-edge-triggered techniques help to improve clock efficiency and reduce the impact of frequency variation in data link portions. Importantly, a novel custom ratioed logic style is adopted in key modules to fundamentally speed up signals' propagation at ultralow-voltage. The proposed BBP was fabricated in 90-nm CMOS as well as the regular design with the same function. It was also implemented in the tag chip's fabrication. In measurement the proposed design indicates good robustness and is much more competent for subthreshold operation. It can operate below 0.3 V with power consumption below 130 nW. Weiwei Shi 0001, An Pan, Oliver Chiu-sing Choy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | A Low-power Signal Processing Front-end and Decoder for UHF Passive RFID TranspondersabstractIn this paper, a low-power signal processing front-end and PIE decoder for use in UHF RFID passive transponders is presented. By merging with the decoder, the clock generator does not need to drive a large loading capacitor. Therefore, its power consumption can be greatly reduced. In addition, the ring oscillator of the generator was designed for low sensitivity to supply voltage variation and low power consumption. Fabricated in a 0.13-mum CMOS process, the measured power consumption consumes only 850 nW. Chi Fat Chan, Weiwei Shi 0001, Kong-Pang Pun, Lincoln Lai Kan Leung, Ka Nang Leung, Oliver Chiu-sing Choy |
ISCAS | 2 |