Xinghua Wang 0005

dblp:09/2773-5 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-1825-7595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 HyperNTT: An Ultra-High Throughput Number Theoretic Transform Accelerator for FHE
Xiyan Dong, Leyan Zhang, An Wang 0001, Xinghua Wang 0005, Liehuang Zhu
ISCAS6
2026 Fama: An FPGA-Oriented Multiscalar Multiplication Accelerator Optimized via Algorithm-Hardware Co-Design
abstract
Multi-scalar multiplication (MSM) is the primary computational bottleneck in zero-knowledge proof protocols. To address this, we introduce FAMA, an FPGA-oriented MSM accelerator developed through algorithm-hardware co-optimization. By integrating a 3D-Pippenger optimization algorithm, FAMA minimizes computational complexity, while its compact dual-mode point addition (PADD) unit significantly reduces hardware overhead. Compared to the best CPU-based design, FAMA achieves over 184.20× speedup. It also outperforms state-of-the-art FPGA-based MSM accelerators, reducing resource overhead by more than 64% and boosting area-time product (ATP) by up to 37.09×.
Xiyan Dong, An Wang 0001, Xinghua Wang 0005, Liehuang Zhu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 A 28-nm Computing-in-Memory-Based Super-Resolution Accelerator Incorporating Macro-Level Pipeline and Texture/Algebraic Sparsity
abstract
Super-resolution (SR) task using the convolutional neural network is a crucial task in improving image and video quality. The introduction of the residual block (RB) raises the depth of the algorithm to perform better reconstruction. The processing of the RB leads to a decrease in hardware utilization and frequent off-chip communications. It is hard to apply such algorithms on edge devices with limited performance. Computing-in-memory (CiM) is one promising method to reduce high power caused by massive data movement in multiply-accumulation computation. The algebraic sparsity (AS) is the structured sparsity (SS) optimization for imaging computing. However, it is an unsolved problem to simultaneously realize the texture sparsity (TS) of the image and the SS of the algorithm in the CiM scheme while maintaining high hardware utilization. Thus, we propose a CiM-based SR task accelerator. There are three key contributions: first, a texture-aware workflow and a dynamic grouping CiM engine can concurrently support TS coupling with AS. Second, a macro-level pipeline scheme together with two custom-sized CiM macros and a high reuse-rate Hadamard transformation circuit reaches 91% hardware utilization. Third, a novel weight update strategy is devised to reduce the performance loss induced by the weight updating. The accelerator prototype is fabricated in a 28-nm CMOS. It scores a 22.8-44.3-TOPS/W peak energy efficiency at the voltage supply of 0.54-1.1 V and the operating frequency of 50-200 MHz, indicating 1.8-6.8x higher compared to the state-of-the-art CiM processors.
Hao Wu 0084, Yong Chen 0005, Yiyang Yuan, Jinshan Yue, Xiangqu Fu, Qirui Ren, Pui-In Mak, Xinghua Wang 0005, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.9
2023 A Security-Enhanced, Charge-Pump-Free, ISO14443-A-/ISO10373-6-Compliant RFID Tag With 16.2-μW Embedded RRAM and Reconfigurable Strong PUF
abstract
Radio frequency identification technology (RFID) has empowered a wide variety of automation industries, such as logistics and freight transportation. To further promote RFID tags adoption, security, power consumption, and cost have always been issues of general concern. This article presents the first synergy of the RFID tag with embedded resistive RAM (RRAM) array and RRAM-based reconfigurable strong physical unclonable function (R-SPUF). The RRAM not only meets the mass storage and technology downscaling but also renders the ultralow-cost “1-cent RFID tag” more feasible. Moreover, the R-SPUF facilitates multiple initializations until a satisfactory distribution and has strong secure keys benefiting from its reconfigurability that improves both safety and reliability. The complete system operates at 13.56 MHz and is compliant with the ISO14443-A and ISO10373-6 (test) protocols. The RFID tag was fabricated on a 1.1-mm2 die based on the 0.18-$\mu \text{m}$CMOS process. Without resorting to the charge pumps for RRAM read–write operations, the total power consumption is as low as 52.3$\mu \text{W}$, of which the RRAM dissipates$16.2~\mu \text{W}$under a wireless power supply.
Qirui Ren, Qiang Huo, Hao Wu 0084, Xiangqu Fu, Xiaoxin Xu, Jianfeng Gao 0005, Xiaojin Zhao, Dengyun Lei, Xinghua Wang 0005, Feng Zhang 0014, Yong Chen 0005, Pui-In Mak
IEEE Trans. Very Large Scale Integr. Syst.15
2018 A 66-dB SNDR, 8-μW analog front-end for ECG/EEG recording application
abstract
For Electrocardiogram (ECG) and electroencephalogram (EEG) recording application, this paper proposes an extremely low-power, low-noise analog front-end (AFE). Based on fully-integrated, high-pass, low-noise amplifier and inverter-based, low-power 2ndSigma-Delta modulator, large output swing, excellent power efficiency and noise performance are achieved. The circuit is implemented in 0.13μm 1P8M Mixed-signal technology. The measurement results show in 0.6V power supply, input referred noise is 3.976μVrms and the noise efficient factor (NEF) is 3.658. Max Signal-to-Noise and Distortion Ratio (SNDR) is 66.7dB with 8.4μw power consumption. Compared with previous work, our design has the maximum SNDR and signal bandwidth, which meets the requirement of ECG/EEG recording application.
Liming Chen 0007, Xinghua Wang 0005, Feng Zhang 0014
ISCAS3
2018 Fast intra coding based on CU size decision and direction mode decision for HEVC
Xinghua Wang 0005, Shan Cao 0001
Multim. Tools Appl.3