Li Geng

dblp:136/5069 · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 50/25/12.5 Gb/s Fast-Settling Burst-Mode Transimpedance Amplifier for 50G-PON
Ruixuan Yang, Li Geng, Dan Li 0011
ISCAS5
2026 Intra-class diversity and inter-class similarity feature calibration network for few-shot object detection
Zhenwei He, Xinye Liao, Zhixian Zhang, Xiaojun Huang, Li Geng, Xin Feng 0007
Knowl. Based Syst.5
2026 A Reliable ESD 3D-Integrated Design and Simulation (3D-IDS) Methodology for Wafer-on-Wafer Stacked DRAM
Xuerong Jia, Fujun Bai, Xiyuan Feng, Li Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2026 A Cryo-Tolerant >40-dB IRR, 4.7-7.9-GHz Double Quadrature Cryo-CMOS Receiver With a Broadband LNA for Scalable Quantum Applications
abstract
This article presents a cryogenic CMOS receiver operating across 4.7–7.9 GHz for frequency-division multiplexing (FDM) quantum state readout. The designed receiver employs a double quadrature (DQ) architecture and a hybrid poly-phase filter (PPF) is proposed to achieve an impressive image rejection ratio (IRR) to meet the 99.99% higher fidelity requirements in quantum error correction (QEC). In addition, gain tuning at different stages of the low noise amplifier (LNA) facilitates a broadband gain and flatness, while the subsequent receiver chains of a broadband mixer and an improved four-input intermediate frequency amplifier (IF-AMP) provide further gain compensation. Fabricated with a standard 55-nm CMOS technology, the LNA achieves a measured power gain of 27.4 dB, with a −3-dB bandwidth (BW) of 5.2 GHz at 6 K. The measured cryogenic noise figure (NF) ranges from 0.79 to 1 dB consuming merely 5.5 mW from a 1.2-V supply. Overall receiver measurements demonstrate a peak gain of 48.2 dB, a −5-dB flatness of 3.2 GHz, and an IRR that exceeds 40 dB at 6 K. These results demonstrate the receiver’s capacity for high-fidelity qubits’ state readout.
Zixun Gao, Chenglong Liang, Ruixin Liang, Suyuan Gan, Bingjun Tang, Xingguo Dong, Youze Xin, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.10
2026 Hardware-Accelerated ASIC and Cardiac Monitoring System for Wearable Devices
Rui Xing 0001, Zhuoqi Guo, Youze Xin, Bing Zhang 0019, Zhongming Xue, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.8
2025 A 5.5-7.9 GHz Double Quadrature Cryo-CMOS Receiver Featuring a Wideband Noise Matching LNA for Quantum Applications
abstract
This work introduces a broadband cryogenic receiver for scalable multiplexed readout of qubits. The receiver utilizes a double quadrature architecture with amplitude and phase error suppression. The hybrid poly-phase filter (PPF) is applied in the architecture to further reduce the input amplitude and phase errors at the radio frequency (RF) end, enhancing the circuit's overall error vector magnitude (EVM). To address the issue of wide bandwidth (BW) for qubit expansion, a low noise amplifier (LNA) with broadband input and noise matching and an improved four-input intermediate frequency amplifier (IF-AMP) with a resistance feedback structure are proposed, achieving a broadband gain response with high flatness. The chip has been designed in standard 55 nm CMOS technology. The LNA achieves a measured power gain of 27.4 dB at 5 GHz with a -3 dB BW of 5.2 GHz at 6 K. Across the entire band of interest, the measured noise figure (NF) deviates by just 0.15 dB from the minimum noise figure (NFmin) and less than 2 dB at 300K, while the measured cryogenic NF falls within the range of 0.79-1 dB. The LNA consumes only 5.5 mW from a 1.2 V supply. Simulation results of the designed receiver indicate that it delivers an average gain of 49.5 dB with a -3 dB BW of 2.4 GHz and an image rejection ratio (IRR) exceeding 40 dB at 300 K.
Zixun Gao, Chenglong Liang, Suyuan Gan, Bingjun Tang, Xingguo Dong, Youze Xin, Li Geng
ISCAS8
2025 An Efficient CNN Accelerator Exploiting Novel Tile-Based Near-Structured Sparsity to Achieve Multi-Level Irregularity Elimination
abstract
Leveraging sparsity in convolutional neural networks (CNNs) has emerged as a promising technique for enhancing the performance of CNN accelerators. However, despite achieving significant improvements in multiplier efficiency, the current sparse-based methods always fail to achieve a competitive performance and runtime latency compared with the conventional dense-based methods. This study posits that, because of the extremely irregular workload distribution in the input feature maps (IFMs), attaining a simultaneous improvement of the performance and multiplier efficiency in sparse-based accelerators is challenging. To address the above challenge, this study proposes a novel framework for eliminating irregularities at multi-level (i.e., dataflow-algorithm-hardware). Specifically, combined with filter decomposition and Winograd algorithms, a computation-oriented dataflow is designed for theoretical workload balancing under different convolution tasks. Furthermore, by exploiting the near-structured characteristic of IFMs, an online tile-based regularization scheme and a hybrid computation reduction method are designed to achieve a decreased and regular workload distribution. Finally, a large-scale sparse CNN accelerator, which integrates a row-merging scheme as well as a workload remapping method, is implemented to further eliminate hardware-level irregularities. The evaluation results show that our proposed methods can achieve 45.93%~76.59% multiplication savings when applied to VGG16, ResNet-34, and ResNet-50. Moreover, our accelerator can accomplish 3.06 TOPS and 2.86 sparsity extraction efficiency (SEE) while deploying VGG16, achieving a 1.17× to 3.45× enhancement on SEE compared with the state-of-the-art sparse-based accelerators.
Yishuo Meng, Chen Yang 0005, Jianfei Wang 0003, Siwei Xiang, Li Geng
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 RTA: A Reconfigurable Transformer Accelerator Exploiting Sparsity via Low-Bit-Width Prediction
abstract
Transformer models have received widespread attention in recent years. They have gradually replaced recurrent neural networks (RNNs) in natural language processing (NLP) and are widely used in tasks such as machine translation, text generation, and language understanding. Similarly, transformers have shown impressive results in computer vision (CV). However, their unique attention mechanism places high demands on the computational and storage resources of the hardware. Deploying transformers on edge computing platforms is challenging due to their complex data flow, intensive matrix calculations, and the need for high-precision nonlinear functions. To address these challenges, we propose reconfigurable transformer accelerator (RTA), a transformer hardware accelerator that uses low-bit-width prediction to achieve dynamic sparsity. RTA reduces resource consumption by performing sparse matrix multiplications using low-bit-width operations, while its reconfigurable design allows the sparse module to be used for high-precision large-bit-width matrix multiplications. We have also optimized the RTA computing pipeline to reduce resource usage and improve computational efficiency. Additionally, we incorporate feature sharing to enhance the resource utilization efficiency of the hardware accelerator. Experimental results on the transformer-base model show that RTA achieves an average performance of 994 GOPS and a digital signal processor (DSP) efficiency of 1412. Compared to state-of-the-art transformer accelerators, RTA achieves$1.37\sim 11.03\times $DSP efficiency.
Chen Yang 0005, Yuheng Xia, Yishuo Meng, Jianfei Wang 0003, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.7
2024 A Low-Power Multimode Eight-Channel AFE for dToF LiDAR
abstract
This paper presents a low-power multimode eight-channel analog front-end (AFE) circuit for direct time-of-flight (dToF) LiDAR applications. The proposed AFE features rich programmability in: a) Two operation modes, parallel mode and selectable mode to save power; b) Two coupling modes, anode and cathode coupling with photodiodes (PDs); c) Two gain modes, high-gain mode for PIN-PD and low-gain mode for APD. Correspondingly, several circuit design techniques have been proposed to support the programmability including: a) Reconfigurable TIA with awake/sleep state and high/low gain mode switching; b) Symmetrical pulse width control to match the pulse width of anode and cathode coupling and input over load protection; 3) Glitch reduction for channel switchover in selectable mode. Designed in 0.18-µm CMOS technology, the AFE achieves bandwidth of 180 MHz, transimpedance gain of 100 dBW, input-referred noise current of 2.1 ${\text{pA}}/\sqrt {{\text{Hz}}} $ Hz and output swing of 1 Vpp,diff. The channel switchover time is around 10 ns in selectable mode. The channel power consumption is 50 mW in parallel mode and 11.4 mW in selectable mode.
Yuye Yang, Ruixuan Yang, Shuaizhe Ma, Li Geng
ISCAS8
2024 Flexible and Efficient Convolutional Acceleration on Unified Hardware Using the Two-Stage Splitting Method and Layer-Adaptive Allocation of 1-D/2-D Winograd Units
abstract
General convolution acceleration, such as Winograd and FFT, is a promising direction to address the computational complexity of current convolutional neural networks (CNNs). However, the flexibility of these CNNs makes this kind of scheme always introduce massive redundant computations, damaging the acceleration effect. In this article, a two-stage splitting method for arbitrarily sized tensors and filters and a unified hardware architecture using layer-adaptive allocated Winograd units are proposed, achieving effective redundance elimination and unified architecture. First, a tensor adaptive presplitting method is proposed to divide the original tensors to match the rule of Winograd. Furthermore, a Winograd-based extended splitting scheme is designed to reduce the redundant calculations; therefore, a substantial reduction in multiplication operations in convolutional layers achieved 30.6%–75% savings. Finally, a unified hardware architecture with a layer-adaptive allocation method is proposed to evaluate and select the optimal Winograd F(${m}$,${r}$) units and input/output parallelisms. This architecture is evaluated based on the Xilinx XCVU9P platform and achieves 1.97/1.23/1.60/1.25 GOPS/DSP for AlexNet, VGG16, modified VGG16, and ResNet18, respectively. It achieves up to$5.81\times $improvements in DSP efficiency compared with previous FPGA-based designs.
Chen Yang 0005, Yaoyao Yang, Yishuo Meng, Kaibo Huo, Siwei Xiang, Jianfei Wang 0003, Li Geng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 A Real-Time and High Precision Hardware Implementation of RANSAC Algorithm for Visual SLAM Achieving Mismatched Feature Point Pair Elimination
abstract
The visual SLAM (vSLAM) algorithm is becoming a research hotspot in recent years because of its low cost and low delay. Due to the advantage of fitting irregular data input, random sample consensus (RANSAC) has become a commonly used method in vSLAM to eliminate mismatched feature point pairs in adjacent frames. However, the huge number of iterations and computational complexity of the algorithm make the hardware implementation and integration of the entire system challenging. This paper pioneeringly proposes an efficient hardware acceleration design with homography matrix as RANSAC hypothesis model, which achieves high speed and high precision. Through optimizing the direct linear transformation (DLT) method, the delay and resource consumption are reduced. The design is implemented on FPGA. Through the verification of Xilinx Zynq 7100 platform, the processing frame rate on EuRoc dataset is 709 fps, reaching an average speed up of$263.2\times $against ARM CPU, and a speed up of$1.2\sim 50.0\times $compared with the advanced implementations in RANSAC part, which fully meets the real-time requirements. In addition, the root-mean-square error (RMSE) based on an open-source SLAM system (ICE-BA) on the EuRoc dataset reached 0.105 m, achieving an improvement of 15.6% in precision compared to the original ICE-BA system.
Wenzheng He, Zikuo Lu, Jingshuo Zhang, Chen Yang 0005, Li Geng
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 Signal Integrity Augmentation Techniques for the Design of 64-GBaud Coherent Transimpedance Amplifier in 90-nm SiGe BiCMOS
abstract
This paper presents signal integrity augmentation design techniques in a 64-GBaud transimpedance amplifier (TIA) for coherent optical communication. In the FE-TIA, a bonding wire ringing reduction technique and an input DC current cancellation (IDCC) loop adapted for coherent communication are proposed. In the post amplifiers, a group delay variation (GDV) friendly bandwidth boosting technique is proposed to achieve optimal time domain performance. A non-linearity cancellation technique and a high-linearity gain control approach are proposed in both circuit and system levels. These signal integrity augmentation techniques form a toolkit to solve the design challenges in bandwidth, linearity, GDV, ringing, offset, crosstalk, etc. in high-speed high-order modulation communication. Fabricated in a 90-nm SiGe BiCMOS technology, the TIA shows input-referred noise current density of 15.1 pA/$\surd $Hz, bandwidth of over 40 GHz with GDV less than ±3.75 ps. The TIA gain can be adjusted between$150~\Omega $- 5 K$\Omega $, which enables maximum overload input current of 3 mApp. The total harmonic distortion (THD) is less than 3% and the crosstalk between two channels is less than -3 dB. The chip consumes 264 mW from 3.3 V supply.
Shuaizhe Ma, Nianquan Ran, Songqin Xu, Chen Tan, Shaoheng Lin, Jianhua Pan, Chaoxuan Zhang, Quan Pan 0002, Zhongming Xue, Xiaoyan Gui, Li Geng, Dan Li 0011
IEEE Trans. Circuits Syst. I Regul. Pap.18
2024 Analysis and Design of a 21.2-to-25.5-GHz Triple-Coil Transformer-Coupled QVCO
abstract
This paper reports a triple-coil transformer-coupled quadrature voltage-controlled oscillator (TC-QVCO), which inherently provides the quadrature signal without using the noisy active-coupling transistors. The determinate correlation of tank voltages is verified by utilizing the initial state to facilitate the oscillation state analysis. Thus, the TC-QVCO would operate without the oscillation mode ambiguity. Additionally, thanks to the triple-coil transformer coupling, a large source coil$L_{S}$aids in achieving in-phase coupling for phase noise (PN) improvement, and the intensified coupling factor$k_{gd}$benefits reducing the PN and the quadrature phase error simultaneously. Therefore, our TC-QVCO would alleviate the tradeoff between PN and quadrature phase accuracy via using a large$L_{S}$and$k_{gd}$. The proposed QVCO prototyped in 65-nm CMOS exhibits a superior FoM$_{\text {@10MHz}}$(180.1 to 182.2 dBc/Hz) over a 18.2% frequency tuning range (21.2 to 25.5 GHz), and the estimated quadrature phase error <0.8°.
Jun Yin 0001, Pui-In Mak, Li Geng
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 A Real-Time and Efficient Optical Flow Tracking Accelerator on FPGA Platform
abstract
Optical flow is a highly efficient visual tracking algorithm, which is commonly used to estimate pixel movement between two consecutive images in a video sequence. However, its high computational complexity and large number of computations become a bottleneck that hinders the performance of embedded vision systems. When applied to simultaneous localization and mapping (SLAM), it is necessary to consider not only time consumption, but also the overall accuracy of the system, causing even greater difficulties. In this paper, a real-time multi-scale Lucas Kanade (LK) optical flow hardware accelerator with parallel pipeline architecture is proposed. The designed circuit meets the high precision and real-time performance required by SLAM while fully considering the limitations of hardware resources. It is deployed on Xilinx Zynq SoC and achieves a frame rate of 93 fps for feature tracking of continuous frame images at$752\times 480$resolution. Compared with the implementation on ARM CPU, the average speed is increased by$4.5\times $. Finally, the feasibility and applicability of the hardware accelerator system designed in this paper are verified on the SLAM system. Experimental results on a public dataset show that the average Root Mean Square Error (RMSE) of this work is 0.189 m, indicating that the hardware accelerator has comparable precision with existing state-of-the-art software algorithms, achieving a great balance of performance and precision.
Yifan Gong 0005, Jinshuo Zhang, Chen Yang 0005, Li Geng
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 An Efficient CNN Accelerator Achieving High PE Utilization Using a Dense-/Sparse-Aware Redundancy Reduction Method and Data-Index Decoupling Workflow
abstract
To adapt to complex scenes and strict accuracy requirements, evolutions have unstoppably occurred in current convolutional neural networks (CNNs). However, these evolutions bring changes to filter size, convolution type, and sparsity, and such diversity leads to difficulties when adopting evolving CNNs in field-programmable gate array (FPGA)-based accelerators. This article proposes a dense-/sparse-aware CNN accelerator to achieve high PE utilization and configurability. First, a filter-based decomposition and clustering algorithm (FDCA) is proposed to change the various-sized filters into unified size filters. In addition, a sparse-aware filter transformation scheme (SFTS) is presented to dynamically eliminate invalid weights for sparse filters and accelerate dense filters. Based on the elimination of sparsity dependency, a hardware accelerator with a data–index decoupling workflow and an input channel schedule-distribution system is designed to take advantage of FDCA and SFTS. The proposed accelerator is implemented on a Xilinx ZCU102 platform at 300 MHz. With different CNN configurations, the digital signal processor (DSP) efficiencies for dense and unstructured sparse AlexNet and dense and structured sparse MobileNetV2 are 0.987, 2.025, 0.547, and 1.278 GOPS/DSP, respectively. Compared with previous dense- and sparse-based designs, the accelerator achieves up to a$4.263\times $speedup in DSP efficiency.
Yishuo Meng, Chen Yang 0005, Siwei Xiang, Jianfei Wang 0003, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.6
2023 A High-Throughput and Flexible Architecture Based on a Reconfigurable Mixed-Radix FFT With Twiddle Factor Compression and Conflict-Free Access
abstract
Mixed-radix fast Fourier transform (FFT) algorithms are widely adopted in high-performance communication systems, such as 5G systems. However, in the decimation-in-time (DIT) FFT algorithm, neither the input time-domain series nor the intermediate-stage data are fetched in natural order, which causes a mismatch between the processing element (PE) computation speed and the memory bandwidth. In addition, multiple duplicate interstage twiddle factor (TF) generation units are designed to provide interstage TFs in parallel, and such duplication results in considerable waste of computing and storage resources. In this article, we propose a flexible and reconfigurable architecture based on a mixed-radix FFT approach, supporting 61 different FFT sizes modes from 12 to 3240 points ($2^{\alpha }$or$12\times n$,$n \le270$, and$n$=$2^{\alpha } 3^{\beta } 5^{\gamma }$). A DIT-FFT-based multiple parallel changeable-radix butterfly unit (BU) is designed to improve hardware resource utilization. Around the PE array, a conflict-free access structure with a hardware-friendly address generating method is presented. In addition, a TF sharing and compression structure is designed to reduce the area and delay of the TF generation units. The FFT architecture is implemented in semiconductor manufacturing international corporation (SMIC) 40-nm CMOS technology with a working frequency of 483 MHz, performing at 654 MS/s for 2048-point and 1536-point FFTs. Compared to the state-of-the-art mixed-radix FFT designs, our architecture achieves improvements of up to$5.89\times $in throughput and area efficiency and supports more FFT modes.
Chen Yang 0005, Siwei Xiang, Liyan Liang, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.5
2022 A Hardware Architecture of Feature Extraction for Real-Time Visual SLAM
abstract
Feature extraction is one of the performance bottlenecks in embedded vision like visual navigation because of its high computational intensity. However, when applying it to simultaneous localization and mapping (SLAM), it is important to consider not only the time consumption but also the accuracy of the system. The more homogeneous the extracted features are distributed, the more accurately the feature matching can represent the geometric relationships in 3D space. In this paper, we propose a new hardware architecture of feature extraction with the addition of mask operation and a mini-grid pipeline to achieve homogeneous distribution. This work is implemented on a Xilinx Zynq SoC. Compared with the Intel i5 and ARM v8.2 CPU implementation of this work, our feature extractor achieves up to 15× and 35× acceleration, and up to 453× and 23× energy efficiency improvement. The evaluation results on the EuRoC dataset show accuracy comparable to the software implementation of the related work. Our architecture is hardware-friendly and achieves a good balance of accuracy and performance.
Liangji Zhang, Xuewei Shen, Yifan Gong 0005, Chen Yang 0005, Li Geng
IECON7
2021 A 12-Bit 100MS/s SAR ADC with Digital Error Correction and High-Speed LMS-Based Background Calibration
abstract
ADC is a significant block in wireless systems, which requires high bandwidth of hundreds of MHz and high resolution of 12-bits. Successive-approximation register (SAR) ADC is an energy efficient architecture but its conversion speed is generally restricted. This paper proposed a scheme to enhance the speed and lower the power consumption of the Least Mean Square (LMS) background calibration, so that it can be used in the design of high-speed SAR ADC. The DEC technique is designed based on the non-binary search, which tackles the insufficient DAC settling, so the conversion speed can be enhanced. Furthermore, the LMS-based background calibration can also improve the linearity of the ADC, so the accuracy of the SAR ADC increases. The proposed SAR ADC is designed in a standard 55 nm CMOS technology with a core area of 0.08 mm2. It consumes 3.59 mW at 100 MS/s sampling rate, achieving a SNDR of 66.78 dB. A good figure of merit (FoM) of 20.13 fJ/conversion-step is obtained.
Zhechong Lan, Li Dong 0007, Xixin Jing, Li Geng
ISCAS4
2021 A 320×240 I-ToF CMOS Image Sensor with 2-Tap 5.6µm Pixel and Mismatch-Nonlinearity Suppression
abstract
This paper presents a 320×240 indirect time of flight (I-ToF) image sensor with 5.6μm×5.6μm 2-Tap pixel in 110nm process. The readout channel offset cancellation and nonlinearity suppression techniques are proposed to achieve high-precision detection. The measured relative precision is 1% at a 5m target distance and non-linearity is below 1.02%. The chip also integrates LVDS and I2C interface for data transmission and Laser control. This work effectively improved the ranging accuracy with a simple method.
Youze Xin, Bing Zhang 0019, Congzhen Hu, Li Dong 0007, Dan Li 0011, Yunsong Wang, Shuyu Lei, Li Geng
ISCAS9
2020 An Open-loop Digitally Controlled Supply Modulator for Wideband Envelope Tracking
abstract
Envelope tracking (ET) is one of the important ways to increase the efficiency of power amplifier (PA). The supply modulator (SM) is critical to fulfil the ET strategy. This paper provides a ripple tracking method based on an open-loop digital controller for SM with one switching amplifier (SA) to efficiently track the envelope of PA. The digital controller combines the strong abilities of caching the envelope signals in advance and the powerful envelope shaping as well as the filtering capability to conquer the intrinsic obstacle of close-loop delay, accelerating the tracking speed and decreasing the operation frequency of SM. The performance of SM is highly enhanced based on the proposed circuit model with error correction technique. The proposed ripple tracking method realizes excellent average power tracking (APT) performance compared with that of multi-level SM structure and decreases the hardware requirements, significantly. A test SA chip is fabricated with a standard 0.18 μm CMOS technology and the test platform of whole SM system is established. The measurement results show that the proposed platform achieves very high efficiency of 88% in tracking the envelope of 10-MHz ~ 100-MHz OFDM baseband signals.
Zeqiang Chen, Li Dong 0007, Kefeng Han, Zhuoqi Guo, Zhongming Xue, Xingzhi Liu, Zheng Ke, Li Geng
IECON8
2020 A 112-Gb/s PAM-4 Linear Optical Receiver in 130-nm SiGe BiCMOS
abstract
In this paper, we present a linear optical receiver for 112-Gb/s PAM-4 optical link. We propose a transimpedance front-end that optimizes thermal noise, power supply noise rejection, linearity and bandwidth altogether. The pseudo-differential structure is employed to achieve both low thermal noise and good power supply noise rejection. A transimpedance amplifier (TIA) gain control technique is proposed to improve linearity at both topology and transistor level while maintaining stability. An NIC-CTLE combo extends bandwidth with optimized frequency response. Designed in a 130nm SiGe BiCMOS process, the receiver realizes 37 GHz total bandwidth and input-referred noise of 19.8 pA/√Hz. The transimpedance gain can vary from 70 dB Ω to 50 dBΩ, which enables maximum input overload current of 1.8 mApp with <; 5% THD at differential output swing of 600 mVpp. The receiver consumes 77mA from 3.3V supply.
Dan Li 0011, Shengwei Gao, Yongjun Shi, Xiaoyan Gui, Nan Qi 0002, Zhiyong Li 0014, Quan Pan 0002, Patrick Chiang 0001, Li Geng
ISCAS9
2020 Low-Supply Sensitivity LC VCOs With Complementary Varactors
abstract
The effects of supply-induced frequency variations on single-ended tuning LC voltage-controlled oscillator (VCO) which degrade the jitter performance of the clock are investigated. The first-order impact on the supply sensitivity is that the varactor's effective capacitance varies with the supply voltage, with other second-order impacts attributed to commonly used capacitive bank and cross-coupled pairs. A compensation technique based on complementary varactors to improve the supply sensitivity of single-ended tuning LC VCO is proposed with no extra power dissipation, nor phase noise degradation within the relative frequency band of interest, along with the discussion on the operating principle of the compensation technique. Both the NMOS cross-coupled and complementary cross-coupled LC VCOs have been designed, demonstrating robust supply-insensitive performance over process, voltage, and temperature (PVT) variations. Prototyped oscillators were fabricated in a 0.18-μm CMOS process to verify both the theoretical analysis and the effectiveness of the proposed technique. Measurement results show that the compensated topologies exhibit more than 93% reduction in periodic jitter versus the noncompensated counterparts, with the figures of merit (FoMs) among the best compared with previous supply insensitive works.
Xiaoyan Gui, Bingjun Tang, Renjie Tang, Dan Li 0011, Li Geng
IEEE Trans. Very Large Scale Integr. Syst.5
2019 A Stacked 4×25 Gb/s Optical Receiver in 28 nm CMOS with 0.154 mW/Gb/s Power Efficiency
abstract
A low-power stacked 4×25 Gb/s optical receiver with cooperative power supply regulators is presented in this paper. Different from conventional parallel channels, the proposed stacked 4 channels fully exploit the supply headroom of the optical module and share a common supply current. Thanks to this current reuse scheme, excellent power efficiency is achieved. In addition, cooperative power supply and photodiode (PD) biasing schemed are co-developed. Each receiver channel is constituted by a transimpedance amplifier (TIA) and a main amplifier (MA), realizing 56.8 dBΩ gain and 27.3 GHz bandwidth. Designed in 28 nm CMOS, the stacked 4×25 Gb/s optical receiver as well as integrated the power regulators together consume 15.4 mW from 3.3 V supply, which translates to state-of-the-art power efficiency of 0.154 mW/Gb/s.
Zhuoqi Guo, Dan Li 0011, Shiquan Fan, Xiaoyan Gui, Li Geng
ISCAS7
2019 An Accurate and Efficient Method for Eliminating the Requirement of Coherent Sampling in Multi-Tone Test
abstract
Multi-tone test is critical in evaluating the overall spectral performance of integrated circuits, especially when the whole signal bandwidth is filled with various frequency components. Coherent sampling is a major challenge for achieving accurate test results in multi-tone test, because it is difficult to simultaneously control all the signal tones precisely. Once coherent sampling is not satisfied, the spectral leakages and their overlap effect will deteriorate the spectral test results. In this paper, a new algorithm is proposed to eliminate the requirement of coherent sampling in multi-tone test. In this method, all the tones are simultaneously estimated with a two-step method which employs merely two FFTs and a few simply mathematical operations. Then new coherent data is reconstructed by replacing the noncoherent fundamentals with coherent tones, all the spectral leakages are removed. Extensive simulation results demonstrate that the proposed method can achieve the same testing accuracy as coherent sampling methods in arbitrary level of noncoherency. The proposed method greatly relaxes the test setup for multi-tone test, and hence the test cost can be reduced significantly.
Cheng Ban, Minshun Wu, Li Geng, Degang Chen 0001
VTS4
2018 Low-Noise High-Linearity 56Gb/s PAM-4 Optical Receiver in 45nm SOI CMOS
abstract
A 56Gb/s PAM-4 linear optical receiver with low noise and high linearity is presented. The fully integrated receiver comprises a transimpedance amplifier (TIA), a variable gain amplifier (VGA), an output buffer, auxiliary analog loops and on-chip bias circuitry. As will be shown, low noise and high linearity often contradict each other, thus both the TIA and VGA implement novel gain control techniques for linear operation while realizing low noise design, making them favorable for PAM-4 signal amplification. Designed and implemented in 45nm SOI CMOS technology, the receiver accomplishes state-of-the-art input-referred noise current of 1.8μArms, 74.4dB transimpedance gain and 23GHz bandwidth while consuming 37mW. The dynamic range achieved is 29dB, enabling large input overload of 0.8mA for PAM-4 compliant signaling.
Dan Li 0011, Yiqun Liu 0011, Ming Liu 0022, Li Geng
ISCAS7
2018 A Near-Zero-Power Temperature Sensor with ±0.24 °C Inaccuracy Using Only Standard CMOS Transistors for IoT Applications
abstract
We propose a near-zero-power temperature sensor only by using standard CMOS transistors without any resistor and special device or process. A pair of temperature sensing elements by employing the normal CMOS transistors is adopted to generate the proportional-to-absolute temperature (PTAT) and constant with temperature (CWT) sensing voltages. Then the two sensing voltages are converted to two reference current sources by using the proposed ultra-low-power voltage-to-current (VI) converter. The VI converter converts the linear voltage signal to an exponential current source with a constant exponent coefficient by using CMOS transistors instead of the conventional resistors. As a result, both power loss and chip area are saved. The whole temperature sensor is designed in standard 0.18 pm CMOS process with active area of only 0.017 mm2. The supply voltage is 600 mV. After 1-point calibration, an accuracy of ±0.24 °C is achieved across the temperature range of 0 °C - 80 °C. The power consumption is only 22 pW for the sensing element, 25 pW for VI converter and 80 pW for total at room temperature (27 °C). The good performance and ultra-small facility of the temperature sensor shows its good practicality for Internet-of-Thing (IoT) applications.
Shiquan Fan, Li Geng
ISCAS3
2018 FORECAST-CLSTM: A New Convolutional LSTM Network for Cloudage Nowcasting
abstract
With the highly demand of large-scale and real-time weather service for public, a refinement of short-time cloudage prediction has become an essential part of the weather forecast productions. To provide a weather-service-compliant cloudage nowcasting, in this paper, we propose a novel hierarchical Convolutional Long-Short-Term Memory network based deep learning model, which we term as FORECAST-CLSTM, with a new Forecaster loss function to predict the future satellite cloud images. The model is designed to fuse multi-scale features in the hierarchical network structure to predict the pixel value and the morphological movement of the cloudage simultaneously. We also collect about 40K infrared satellite nephograms and create a large-scale Satellite Cloudage Map Dataset(SCMD). The proposed FORECAST-CLSTM model is shown to achieve better prediction performance compared with the state-of-the-art ConvLSTM model and the proposed Forecaster Loss Function is also demonstrated to retain the uncertainty of the real atmosphere condition better than conventional loss function.
Jianwu Long, Li Geng
VCIP4
2018 Passive Noise Shaping in SAR ADC With Improved Efficiency
abstract
This brief reports a passive noise-shaping (PNS) scheme for successive approximation register (SAR) analog-to-digital converter (ADC) based on the two-step integration with passive gain and comparator gain techniques. The analysis shows that the proposed method achieves a better noise-shaping (NS) efficiency than prior arts, which enhances the noise attenuation by 14 dB. A design example is provided which further adopts the delta-sampling technique to relieve the conversion efficiency loss due to the oversampling in the NS SAR ADC. The efficiency of the proposed PNS scheme and the performance of the ADC are verified by simulation achieving a 13.2 effective number of bits with a 10-b ADC architecture and eight conversion cycles for a signal bandwidth of 2 MHz sampled at 100 MS/s. The calculated Schreier figure of merit (FoM) and Walden FoM are 176.8 dB and 16 fJ/conv.-step, respectively.
Chi-Hang Chan, Yan Zhu 0001, Li Geng, Seng-Pan U, Rui Paulo Martins
IEEE Trans. Very Large Scale Integr. Syst.4
2017 An auxiliary switched-capacitor power converter (SCPC) applied in stacked digital architecture for energy utilization enhancement
abstract
Reducing the supply voltage of digital circuit to its sub-or near-threshold region is critical for achieving minimum energy point (MEP) operation. However, as the supply voltage on the load is lowered, the overall efficiency of conventional circuit architecture using DC-DC converter also decreases. In this work, we demonstrate a stacked load architecture regulated by an auxiliary switched-capacitor power converter (SCPC) to achieve ultra-low power regulation of load supply voltage. We further show that the self-restoring behavior of MEP operation can also be leveraged in our design for achieving high-efficiency supply voltage regulation. As an example, we implement our design with a standard 0.18 μm CMOS process in a 4-stacked structure, with each stacked load cell operating at 450 mV in its sub-threshold region and fully functioning. While the current regulating capability of the SCPC is designed to be only 10% of the load current to match the load current difference, the load voltage can be regulated with more than 99% precision. We achieve an overall energy efficiency of >94% for the entire system.
Shiquan Fan, Zhuoqi Guo, Li Geng
ISCAS5
2017 An ultra-low quiescent current power management ASIC with MPPT for vibrational energy harvesting
abstract
We present an ultra-low quiescent current power management integrated circuit (IC) for interfacing with a piezoelectric (PZE) energy harvester in a self-powering wireless sensor node. The PZE harvester first transforms mechanical energy from vibration into electricity. The AC output from the PZE is then converted into DC power using a full bridge rectifier, and charges a small filter capacitor. The energy is further transferred into a supercapacitor by a Buck-Boost converter with high efficiency, using an impedance matching load to achieve maximum power point tracking (MPPT). A low-dropout (LDO) regulator powers the on-chip CMOS sensor and an external RF transmitter with the energy stored in the super capacitor, providing clean power supply, forming a complete wireless temperature sensor node, which can enable Internet-of-Things (IoT) sensor networks. The circuit is taped out using a 0.5 μm standard CMOS process. Measurements confirm all key features of this design: (i) up to 96% voltage converting efficiency of the rectifier; (ii) a minimum of 102 s charging time to charge a 1 mF supercapacitor from 0 V to 3.3 V of the Buck-Boost converter with impedance matching method; (iii) a 0 nA to 100 μA load current range and at least 85° phase margin (PM) LDO regulator with an ultra-low quiescent current as low as 750 pA.
Shiquan Fan, Liuming Zhao, Li Geng, Philip X.-L. Feng
ISCAS4
2017 A delay time controlled active rectifier with 95.3% peak efficiency for wireless power transmission systems
abstract
Active rectifier with comparators (CMPs) is often used in wireless power transmission (WPT) systems. However, it suffers from low power conversion efficiency (PCE) in light load condition and multiple pulse problem (MPP) due to the CMPs with delay compensation. In this paper, a novel active rectifier with delay time controller is proposed to solve both issues. A current control delay line (CCDL) is introduced to adjust the rising and the falling edges of the gate voltages of the NMOS (Fgns), controlled by a negative feedback loop consisting of a switched-capacitor (SC) sample module and a supply independent bias current controller (IB controller). The proposed rectifier is designed with a standard 0.18μm CMOS process. Post-layout simulation results show that the delay time controller consumes only 34μΑ, which is much smaller than the power consumptions induced by controller with CMPs, thus significantly enhancing the PCE of the rectifier. The PCE of the rectifier is higher than 88% in the whole load range from 629μΑ to 32.1mA, and exceeds 92% when load resistance Rl varies from 100Ω to 1200Ω. A peak PCE of 95.3% is achieved when Rl is 200Ω and Vac is 2V.
Zhongming Xue, Dan Li 0011, Wei Gou, Shiquan Fan, Li Geng
ISCAS6
2015 Real-time self-tracking in the Internet of Things
abstract
We investigate the problem of real-time self-tracking of tagged objects in a new system with low-cost “smart” tags. These tiny and battery-less devices will play a pivotal role in the infrastructure of the Internet of Things (IoT). With capabilities of low-power computation and tag-to-tag backscattered communication, no readers will be needed for running the Radio Frequency Identification (RFID) system. In order to allow for low-cost tags, self-tracking has to be performed with simple algorithms while still exhibiting high accuracy. In this paper we propose a linear observation model for which Kalman filtering (KF) is the optimal method. We also consider a nonlinear model for which we apply particle filtering (PF) of reduced complexity as the tracking method. The performance and computational complexity of the different methods are compared by computer simulations.
Li Geng, Mónica F. Bugallo, Akshay Athalye, Petar M. Djuric
ICASSP1
2013 Tracking with RFID asynchronous measurements by particle filtering
abstract
This paper deals with the problem of real-time indoor tracking of tagged objects in Ultra High Frequency Radio Frequency Identification systems with asynchronous measurements. A new and more realistic model of the system is proposed, where the probability of detecting a tag by a reader is described by a function of both the distance and the angle between the tag and the reader's antenna. The model also accounts for the possibility of a tag being in a dead-zone where the tag cannot be detected. For tracking, we propose the use of the particle filtering methodology that takes into account the asynchronous nature of the measurements. The parameters for modeling the resulting system are obtained from real-world experiments and the performance of the algorithm is shown by extensive computer simulations.
Li Geng, Mónica F. Bugallo, Petar M. Djuric
ICASSP1