Guoqiang He

dblp:96/432 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-8853-9187ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 FRRAP: A Fast-Response Reconfigurable AI Processor and Its Application in Four-Class Seizure Monitoring
Heng Zhang 0025, Qiang Tao, Chi Ben, Youbin Luo, Guoqiang He, Sirui Zhu, Jingyi Ma, Li Li 0003
ISCAS5
2026 Fast Modular Reduction Algorithm and Reconfigurable Domain-Specific Architecture Design Based on Generalized Mersenne Primes
abstract
Modular arithmetic enjoys a broad spectrum of applications. In recent years, the advancement of post-quantum cryptography (PQC) has imposed growing demands on the flexibility and scalability of domain-specific accelerators. This article presents a fast modular reduction algorithm based on generalized Mersenne (GM) primes, which employs approximate scaling and iterative compression approaches to rapidly converge the quotient value while reducing the precomputation complexity to rely solely on the modulus itself. Building upon this algorithm, we have designed a reconfigurable modular reduction array using multiplier units with smaller word length. Operating at 1GHz, the proposed array achieves an area reduction of 45.14% compared to Barrett-based structures, and 13.93% compared to Montgomery-based structures. The array has been integrated into a complete number-theoretic transform (NTT) acceleration architecture. The resulting reconfigurable GM/general modular reduction domain-specific architecture elevates the chip frequency to$0.42\sim 1$GHz under the same process technology. Under identical test conditions, it improves area efficiency by 10.6% and reduces energy consumption by 17.0% compared to the state-of-the-art ASIC design. When compared to the latest field-programmable gate array (FPGA) implementations, it achieves a reduction in area-time product (ATP) by 7.2%~18.4% for Kyber and by 13.4% for Dilithium. These results strongly demonstrate the notable advantages of the hardware-friendly GM algorithm.
Xinyu Wang 0027, Guoqiang He, Congyi Sun, Kai Chen 0034, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.2
2026 QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.7
2023 A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal Processing
abstract
The wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm.
Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Unsupervised Learning Based on Temporal Coding Using STDP in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been recognized as one of the next generation of Neural Networks (NNs), showing a great potential in a variety of applications. Spiking-Timing Dependent Plasticity (STDP) underlies the brain’s learning mechanisms, and trains SNNs with great energy efficiency. In this paper, we propose a low-cost spike-time based unsupervised learning method. It constructs a SNN with one fully-connected excitatory layer structure without inhibitory layer, and trains the SNN with STDP using a first-spike-based temporal coding scheme where input information is directly encoded into spike times. It only updates the synaptic weights connected to the neuron that first generates a spike in a forward propagation step, which reduces the frequency of the synaptic weight updates significantly. The forward propagation process can be stopped once a neuron fires whether in the training mode or the inference mode, by which many unnecessary computations are just avoided and the latency in the inference mode is reduced. The method was used to train on the classification task on MNIST dataset and achieved an accuracy of 90.4% with 800 excitatory neurons.
Congyi Sun, Qinyu Chen, Kai Chen 0034, Guoqiang He, Li Li 0003
ISCAS4
2022 Low-Latency Low-Complexity Method and Architecture for Computing Arbitrary Nth Root of Complex Numbers
abstract
This paper presents a new architecture, based on CORDIC and parabolic synthesis methodology, for computing Nth root of a complex number. The proposed architecture uses the pretreatment for normalization and parabolic synthesis method to calculate the Nth root of modulus of the input complex number and performs the conversion between the plane coordinate form and the polar coordinate form of the complex number by CORDIC, which not only ensures the accuracy but also has an ultra-low computation latency. MATLAB simulation result indicates that our proposed method can calculate the Nth root of the complex numbers in the form of fixed-point number with an error of$2.16 \boldsymbol {\times {10^{ - 6}}}$. Under TSMC 40nm CMOS technology, the report shows that the area consumption is$27390.72 \boldsymbol {\mu m^{2}}$at the frequency of 1GHz and the power consumption is 2.3549mW. More importantly, the computation latency of the proposed architecture is only 60.18% of the latest architecture in the same calculation accuracy.
Hui Chen 0015, Guoqiang He, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.3
2019 Congestion-Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCs
abstract
The combination of Network-on-Chips (NoCs) and 3D IC technology, 3D NoCs, has been proven to be able to achieve a great improvement in both network performance and power consumption compared to 2D NoCs. In the traditional 3D NoC, all routers are vertically connected. Due to the large overhead of Through-Silicon-Via (TSV, e.g., low fabrication yield and the occupied silicon area), the partially connected 3D NoC has emerged. The assignment method determines the traffic loads of the vertical links (elevators), thus has a great impact on 3D-NoCs' performance. In this paper, we propose a congestion-aware dynamic elevator assignment (CDA) scheme, which takes both the distance factors and network congestion information into account. Experiments show that the performance of the proposed CDA scheme is improved by 67% to 87% compared to the random selection scheme, 8% to 25% compared to SelByDis-1, and 13% to 18% compared to SelByDis-2.
Qinyu Chen, Guoqiang He, Kai Chen 0034, Zhonghai Lu, Chuan Zhang 0001, Li Li 0003
ISCAS3
2007 A Study Upon the Architectures of Multi-Agent Systems for Petroleum Supply Chain
Jiang Tian, Huaglory Tianfield, Juming Chen, Guoqiang He
CDVE4