Zhongzhiguang Lu

dblp:330/1671 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0004-6755-2800ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 A mm-Wave Coupler-based Dual-band Power Amplifier for Advanced Driver Assistance Systems
abstract
The growing demand for high-performance components in wireless communication and automotive systems, especially for radar applications, has driven the need for dual-band power amplifiers (PAs) operating at 60GHz and 77GHz. These frequency bands are particularly beneficial for automotive radar systems, integral to Advanced Driver Assistance Systems (ADAS) and autonomous driving technologies, as they offer enhanced resolution, reduced interference, and faster data transmission rates. This paper presents the design and development of a dual-band PA based on a novel coupled-line dual-frequency matching structure. The PA’s innovative input and output matching networks utilize a unique coupler design to achieve simultaneous impedance matching at both 60GHz and 77GHz. Through comprehensive simulation, optimal matching impedances for both frequencies were identified, enabling the PA to achieve an output power of 12 dBm at 60GHz and 10 dBm at 77GHz, with power-added efficiencies of 24.4% and 13.85%, respectively. The design also incorporates a two-stage power amplifier configuration that ensures high efficiency and gain across the dual bands. Experimental validation was performed using a small-signal test system, demonstrating excellent performance, with a peak power gain of 12.4 dB at 60GHz and 9.8 dB at 77GHz. This dual-band PA design is particularly well-suited for integration into automotive radar systems, thanks to its compact size, high power efficiency, and ability to support wideband matching. Furthermore, this work presents a highly efficient, wideband solution for next-generation automotive radar and communication systems operating in the millimeter-wave frequency range.
Zhongzhiguang Lu, Yanshu Guo, Yange Wang, Cao Wan, Guanghao Fan, Yuanjin Zheng
ISCAS1
2025 The Photoacoustic Quality-Enhancement Neural Network Processor with the Scalable and End-to-End Architecture by Improving the Sparsity Level
abstract
Recent advancements have marked significant progress in photoacoustic imaging as an effective method for acquiring deep bio-tissue visuals in modern medical clinical therapy and the efficacy of U-Net and its variants has been established for imaging quality enhancement in this field. Unlike common computer vision datasets such as ImageNet [1] and PASCAL VOC [2], biomedical images exhibit highly structured patterns, low spatial resolution, and single-channel modality, as shown in Fig. 1. Additionally, the U-Net parameters trained for medical super-resolution tasks demonstrate a high sparsity ratio, making them suitable for implementation on edge-computing platforms. Therefore, developing an energy-efficient photoacoustic imaging setup in this area is a natural progression. However, this development is constrained by the current neural network architectures, which are built around a U-Net backbone. The multi-stage feature extractor, skip connection integration across different blocks, and the encoder-decoder backbone design pose significant challenges to cutting-edge computational hardware platforms. In this study, a scalable, sparsity-supported neural network accelerator architecture for bio-tissue imaging quality enhancement is proposed to meet the stringent requirements of latency and energy efficiency, as depicted in Fig. 2. This architecture achieves desired performance improvements by exploring the sparsity possibilities in neural network during the training process and implementing an end-to-end pixel-first hardware design to minimize data movement and support sparsity computation. Compared with the state-of-the-art related works, this optimized architecture has achieved minimum on-chip storage overhead and the fastest frame for the application of photoacoustic imaging quality enhancement. The scalable architecture has also been implemented on a Xilinx XCZU9EG FPGA and attains a performance of PSNR@ 24 dB and a frame rate of 164 fps at a working frequency of 250 MHz.
Zhengyuan Zhang 0002, Caijie Liang, Boyi Dong, Yange Wang, Zhongzhiguang Lu, Xiangjun Yin, Shenglong Zhuo, Yifan Wu 0009, Yingjie Cao, Tianyang Zhou, Jian Qian, Patrick Chiang 0001, Lei Qiu 0002, Yuanjin Zheng
ISCAS7
2025 Compact Sub-THz Frequency Conversion Module in 28-nm CMOS for D-Band Radar Transceiver
abstract
This paper proposes a compact sub-THz frequency conversion module for D-band transceivers, fabricated using a 28-nm CMOS process. The module integrates an injection-locked frequency multiplier (ILFM) for Tx signal frequency up-conversion and an active Gilbert double-balanced mixer for Rx signal frequency down-conversion. The system was tested on a probe station. Utilizing a tunable coupling-coil technique and optimized inductance, the ILFM achieves a locking range of 105.2-125.5 GHz with an output power of -4.5 dBm. The Gilbert mixer demonstrates a conversion loss of -6.5 dB across the same range with an LO power of -4.2 dBm. The active region of ILFM and mixer chips occupy areas of 0.31 mm2and 0.51 mm2, respectively.
Yange Wang, Guanghao Fan, Boyi Dong, Zhongzhiguang Lu, Cao Wan, Yuanjin Zheng
ISCAS4
2024 A Fully Probabilistic Model for Sigmoid Approximation and Its Hardware- Efficient Implementation
abstract
The sigmoid function is a representative activation function in shallow neural networks. Its hardware realization is challenging due to the complex exponential and reciprocal operations. Existing studies applied piecewise models to approximate sigmoid function and employed numerical methods or non-uniform input segmentations to mitigate fitting inaccuracies. However, the breakpoints introduce inevitable approximation precision loss. Besides, additional fitting processes greatly increase hardware complexity and power consumption. This paper presents a hardware-friendly sigmoidal approximation from the perspective of probability theory. We find that for a given input, the output of a sigmoid function can be approximated by the probability that the sum of this input and a Gaussian random variable is greater than or equal to zero. As the derived theorem does not involve piecewise expressions, the precision loss caused by the breakpoint issue is avoided. A low-complexity binary-search-based address localization method is proposed to optimize our theorem for hardware implementation. For the optimized scheme, an efficient implemented circuit is also presented. Our scheme’s approximation ability and hardware efficiency are validated through software modeling and FPGA- and ASIC-based experiments. Feedforward neural network-based classification applications demonstrate that building networks with the proposed sigmoid approximator has only a tiny recognition rate loss.
Wenhao Lu, Minshan Lu, Xiangfen Zhang, Zhongzhiguang Lu, Boyi Dong
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 A Wideband GaN HEMT Modelling with Comprehensive Hybrid Parameter Extraction for 5G Power Amplifiers
abstract
Due to better efficiency, gain and thermal performance compared to other semiconductor technologies, GaN power amplifiers are very attractive in the present 5G era. Meanwhile, accurate GaN HEMT device modelling is one of the critical steps to design PAs successfully. Therefore, research on high-frequency GaN HEMT device modelling method is of vital importance. This paper first presents a wideband GaN HEMTs model for 5G power amplifiers. A loadpull system available for 10-67 GHz measurement is set up to obtain wideband S parameter and RF performance results. The whole GaN device modelling could be divided into two parts: the small signal modelling and the large signal modelling. Direct optimization method with polynomial fitting is employed to obtain equivalent small signal circuit parameters, which improve the accuracy and efficiency of parameter extraction. Also, artificial neural network (ANN) technique is utilized to build charge and nonlinear current models, which takes the self-heating and trapping effects into consideration in the large signal modelling. The ANN technique could substitute the complex empirical equations as other papers has reported, and thus makes the extracted parameters less and the extraction process more accurate and efficient. At last, the proposed model is implemented and verified in ADS, the error between the measurement and simulation results is less than 5%.
Zhongzhiguang Lu, Hanlin Xie, Jiaming Piao, Wei Zhengzhe, Geok Ing Ng, Yuanjin Zheng
ISCAS1
2023 A Graph-Based Accelerator of Retinex Model with Bit-Serial Computing for Image Processing
abstract
This work implements the Poisson equation formulation of the Retinex model for image enhancements using a graph hardware accelerator performing finite difference updates on a 2D lattice graph PE array. A single clock gating control signal manages the data flow, data sharing, and reuse pattern among neighboring PEs during massively parallel updates. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time, the proposed accelerator consists of 18$\times 18$regular PEs surrounded by$4\times 20$boundary PEs with reconfigurable data flow and 4 boundary cache registers. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2core area, and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision.
Zhengzhe Wei, Junjie Mu, Zhongzhiguang Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim
ISCAS3