Jian Liu 0021

dblp:35/295-21 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-8057-2444ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 A 75.6 Gb/s 22-bit Floating-Point Coarse-Grained Versatile DSP Embedding 26K-Point Baseband Signal Processing and 2048 × 256-Point Complex FFT
abstract
As millimeter-wave radar technology advances, modern domain-specific digital signal processors (DSPs) struggle to balance versatility and processing scale, while suffering from low data precision and data throughput. To address these challenges, this paper presents a novel coarse-grained versatile DSP (CVDSP). The CVDSP introduces an architecture based on multi-level finite state machines and a custom instruction set to support various algorithms through flexible dataflow. The efficient large-scale PEs are designed with 22-bit floating-point precision to handle 26K-point baseband signal processing and$2048\times 256$-point complex fast Fourier transform. Cooperating with a high-bandwidth instruction-free memory access network, the CVDSP achieves relatively high data throughput. Fabricated in a 65-nm CMOS process, experimental results show that the peak energy efficiency and data throughput of the CVDSP are 299.5 GMACs/J and 75.6 Gb/s. The CVDSP demonstrates fully on-chip implementation of the FMCW radar, Pulse-Doppler radar, and spectrometer algorithms.
Xuanzhe Xu, Xianjun Liu, Siyuan Wei, Shuangming Yu, Runjiang Dou, Xu Yang 0017, Jian Liu 0021, Nanjian Wu
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 A 56-Gb/s, 6.3-pJ/bit PAM-4 DFB Laser Driver Incorporating Asymmetric Equalization and Integrated CDR in 28 nm CMOS
abstract
This article presents a 56-Gb/s distributed feedback (DFB) laser driver integrated with a PAM-4 clock and data recovery (CDR). A mixed-signal digital-to-analog converter (DAC) is adopted for power-efficient linear driving. With the help of the CDR, high-speed PAM-4 input is digitized into thermometer code, which is processed in NRZ format along the data path before summation at the output node. In this way, higher modulation linearity is realized by independently adjusting the weight of each slice. A dc-coupled differential drive stage is devised to improve signal integrity and energy efficiency at high speed. Employing a fractional-UI delay asymmetric feed-forward equalization (FFE) extends the laser’s bandwidth while the nonlinearity is compensated. The proposed driver is fabricated in 28-nm CMOS and co-packaged with a DFB laser diode. Measurement results show the modulated optical output reaches a 56-Gb/s data rate and consumes 353-mW power, thus corresponding to the energy efficiency of 6.3 pJ/bit, including the integrated CDR.
Yang Min, Nan Qi 0002, Minye Zhu, Guike Li, Yonghui Lin, Huiyao Peng, Mo Guang, Kaiwen Long, Zhao Zhang 0004, Jian Liu 0021, Nanjian Wu, Jingbo Shi, Yong Chen 0005, Frank F. Shi
IEEE Trans. Very Large Scale Integr. Syst.12
2024 A 32Gb/s NRZ Low-Bias DFB Driver with Frequency Boosting for High Efficiency Data Transmission
abstract
This paper presents a 32Gb/s non-return-to-zero (NRZ) distributed feedback (DFB) laser diode driver (LDD) fabricated in 65nm CMOS. The driver is directly wire-bonded to the laser diode without AC-coupling capacitors, which simplifies the packaging and ensures high bandwidth (BW). To improve power efficiency, the active back-termination (ABT) structure is employed to absorb signal reflections with lower power. A continuous time linear equalizer (CTLE), a 3-stage current mode logic (CML) buffer and a pre-driver are employed to compensate the channel loss and achieve a bandwidth extension by providing an estimated gain boosting of 7.5dB at high frequency. The clear electrical eye-diagram of the driver is obtained beyond 40Gb/s, while the measured optical transmission data-rate can still exceed 32Gb/s with a rms-jitter of 2ps. According to the experimental results, the bias current and modulation current of the driver are 40mA and 40mApp, respectively, where the power consumption is 390mW.
Yang Min, Leliang Li, Guike Li, Zhao Zhang 0004, Jian Liu 0021, Nanjian Wu, Yonghui Lin, Huiyao Peng, Jingbo Shi, Nan Qi 0002
ISCAS7
2024 A Real-Time 2D/3D Perception Visual Vector Processor for 1920 × 1080 High-Resolution High-Speed Intelligent Vision Chips
abstract
Edge computing of reliable multimodal (2D RGB/3D RGB-Depth) data has a wide range of applications. However, many of currently reported visual processors cannot flexibly handle multimodal data, e.g., the visual streams of RGB-Depth data. The key challenge exists that these prior visual processors do not come with efficient and unified instruction set architecture (ISA) for both conventional and intelligent cognition on the 2D/3D multimodal sensory data. To fill such a gap, this paper proposes a programmable intelligent visual vector processor compatible with multimodal 2D/3D visual data processing ($1920\times 1080$-pixel resolution). The processor consists of a reconfigurable processing element (PE) array, a memory access network flexibly configurable to be fine- or coarse-grained, and a high throughput I/O interface. The vectorial PE array with neighbor PE access increases the data reuse rate and parallel computation efficiency, and can implement both convolutional neural networks (CNNs) and conventional image processing algorithms. The proposed ISA is customized and optimally tailored targeting 2D/3D image processing from RGB/Time-of-Flight(ToF) raw data to intelligent inference results. The chip is fabricated in a 55-nm CMOS process. The experimental results showed that the area efficiency, peak performance, and peak throughput of our chip attained as high as 14.41GOPS/mm2, 409.6GOPS, and 9.6Gbps at 200MHz, respectively. The measured processing speeds of this chip on ToF depth reconstruction is 87fps ($480\times 270$) or 31 fps($1920\times 1080$),on 3D object classification is 219fps ($256\times 256$), and on CNN-based 2D object tracking is 36fps ($256\times 256$).
Siyuan Wei, Lei Kang 0006, Xuemin Zheng, Mingxin Zhao, Mengmeng Xu 0005, Xuanzhe Xu, Runjiang Dou, Shuangming Yu, Xu Yang 0017, Jian Liu 0021, Cong Shi 0003, Nanjian Wu
IEEE Trans. Circuits Syst. I Regul. Pap.12
2023 An 800G Integrated Silicon-Photonic Transmitter based on 16-Channel Mach-Zehnder Modulator and Co-Designed 5.35pJ/bit CMOS Drivers
abstract
A 800G integrated silicon-photonic transmitter is presented, including a 16-channel photonic integrated chip (PIC) and two electrical chiplets (EICs) that are realized based on an arrayed travelling wave dual-drive Mach-Zehnder modulator (MZM) and two 8-channel CMOS drivers. The proposed multi-channel PIC is fabricated on a high-resistance silicon-on-insulator (SOI) wafer with a 220 nm thick silicon layer and a$\mathbf{2}\ \boldsymbol{\mu} \mathbf{m}$thick buried oxide (BOX) using the foundry-ready CMOS process, while the drivers are implemented in a standard$\mathbf{65}\mathbf{nm}$CMOS process. The driver employs a combination of distributed architecture, 2-tap feedforward equalization (FFE) and push-pull output stage, experimentally exhibiting an averaged bandwidth higher than 28.5GHz and a differential swing of 4.0Vpp on$\mathbf{50}\mathbf{\Omega}$load, respectively. The 50Gb/s electrical eye-diagram is measured with 1.41ps rms-jitter, while the optical extinction ratio (ER) exceeds 3.0dB with 5.35pJ/bit power efficiency.
Jingbo Shi, Haowen Shu, Fenghe Yang, Yuansheng Tao, Jianrui Deng, Ruixuan Chen, Changhao Han, Jian Liu 0021, Nanjian Wu, Nan Qi 0002
ISCAS12
2023 A 50Gb/s CMOS Optical Receiver With Si-Photonics PD for High-Speed Low-Latency Chiplet I/O
abstract
This paper presents a 50-Gb/s optical receiver (ORX) chipset, consisting of a transimpedance amplifier (TIA) and a clock and data recovery (CDR) circuit in a 45-nm silicon-on-insulator CMOS. The proposed inverter-based TIA employs hybrid shunt-series peaking inductors to extend the bandwidth (BW). A baud-rate CDR is proposed to reduce the sampling phases and clocking power by half. To optimise the ORX for in- package integration, a compact-size digital loop is adopted in each channel, and the clock is recovered by phase interpolation from a shared reference. A complete optical-to-electrical (OE) link is built by integrating the proposed ORX with a high-speed Silicon Photonics (SiP) photodetector (PD). Measurements show that the proposed TIA has a transimpedance gain of 53 dB$\Omega $and a BW of 27 GHz. By integrating it with the SiP PD, the OE front-end (PD+TIA) achieves an input sensitivity of −7.7 dBm at 50 Gb/s and BER$ < 10^{-12}$. It features a power efficiency of 1.61 pJ/bit at a data rate of 64 Gb/s. The complete 50 Gb/s ORX achieves data recovery at a quarter rate of 12.5 Gb/s with an output jitter of 1.6 psrms, and has a 3.125 GHz clock with phase noise of −115.22 dBc/Hz at an offset frequency of 1 MHz.
Sikai Chen, Mingyang You, Yunqi Yang, Leliang Li, Guike Li, Zhao Zhang 0004, Binhao Wang 0002, Ningfeng Tang, Faju Liu, Zheyu Fang, Jian Liu 0021, Nanjian Wu, Yong Chen 0005, Ninghua Zhu, Nan Qi 0002
IEEE Trans. Circuits Syst. I Regul. Pap.15
2022 Design of a PAM-4 VCSEL-Based Transceiver Front-End for Beyond-400G Short-Reach Optical Interconnects
abstract
This paper presents a hybrid-integrated optical transceiver front-end for beyond-400G short-reach optical links. A pair of the monolithic 8-channel laser drivers and the trans-impedance amplifier (TIA) is developed in 180nm SiGe BiCMOS, incorporating arrayed Vertical-Cavity-Surface- Emitting Lasers and photo-detectors. The driver uses a$2^{\mathrm {nd}}$-order continuous-time linear equalizer (CTLE) to compensate for the channel loss with a nonlinear frequency response. Both the inductive peaking and RC-degeneration are embedded at the output stage to extend the optical modulation bandwidth (BW). The series-peaking and multi-stage distributed CTLE are combined in a resistive feedback TIA topology for improved BW and linearity. Measurement results show up to 100-Gb/s PAM-4 electrical eyes of the driver and TIA. The optical transmitter front-end operates 56 Gb/s, 4.1-dB extinction ratio, and 6.6-pJ/bit power efficiency, while the optical receiver front-end achieves 56-Gb/s,$10^{-6}$bit error rate, and 5.9-pJ/bit power efficiency.
Donglai Lu, Haiyun Xue, Sikai Chen, Leliang Li, Guike Li, Zhao Zhang 0004, Jian Liu 0021, Nanjian Wu, Ningmei Yu, Fengman Liu, Xi Xiao 0004, Yong Chen 0005, Nan Qi 0002
IEEE Trans. Circuits Syst. I Regul. Pap.9
2022 A 56-Gb/s Reconfigurable Silicon-Photonics Transmitter Using High-Swing Distributed Driver and 2-Tap In-Segment Feed-Forward Equalizer in 65-nm CMOS
abstract
This article presents a reconfigurable silicon- photonics transmitter (TX) for short-reach optical interconnects. The proposed hybrid-integrated TX combines a 65-nm CMOS driver with a 180-nm SOI-CMOS silicon-photonic Mach-Zehnder Modulator (MZM). The driver integrated with in- segment fractional-UI spaced feed-forward equalizer (FFE) is proposed to support the non-return-zero (NRZ) signaling, electrical- and optical-domain 4-level pulse-amplitude modulation (PAM-4) signaling. The driver employs a reconfigurable distributed topology to achieve high swing, wide bandwidth and flexible operation. The MZM is driven differentially in a push-pull configuration for high modulation efficiency. Measurement results show that the proposed TX operates up to 50-Gb/s NRZ data rate with 4-Vppd swing and 1.92-ps RMS jitter. In the optical PAM-4 mode, it reaches 56-Gb/s data rate and achieves >5-dB extinction ratio (ER) at the cost of 10.9-pJ/bit power efficiency.
Yuguang Zhang, Qiwen Liao, Zhao Zhang 0004, Miaofeng Li, Jingbo Shi, Jian Liu 0021, Nanjian Wu, Yong Chen 0005, Patrick Chiang 0001, Ningmei Yu, Xi Xiao 0004, Nan Qi 0002
IEEE Trans. Circuits Syst. I Regul. Pap.9
2020 A 50Gb/s PAM-4 Optical Receiver with Si-Photonic PD and Linear TIA in 40nm CMOS
abstract
A 50Gb/s PAM-4 optical receiver with Silicon Photonic (Si-Ph) photodiode (PD) and CMOS linear transimpedance amplifier (TIA) is presented. To optimize both noise and bandwidth, a two-stage front-end architecture-a high gain-low bandwidth TIA followed by a two-stage continuous time linear equalizer (CTLE) is adopted. Gain adjustment of the entire link is achieved by adjusting the TIA feedback resistor and the voltage of variable gain amplifier (VGA) to ensure that the receiver analog front-end (AFE) remains linear over the entire photocurrent input range. The chip has been realized in 40nm CMOS process. Experimental results show the TIA achieves 66dBΩ transimpedance gain, 24.4GHz bandwidth, 20dB gain dynamic range, maximum overload current 2mA, and differential output swing of 400mV. The total power consumption of the chip is 125.4mW.
Yang Liu 0178, Nan Qi 0002, Xiuli Xu, Lei Wang 0187, Minjia Chen, Qixiang Cheng, Jingbo Shi, Jian Liu 0021, Xi Xiao 0004, Nanjian Wu
ISCAS10
2019 Efficient Reservoir Encoding Method for Near-Sensor Classification with Rate-Coding Based Spiking Convolutional Neural Networks
Xu Yang 0017, Shuangming Yu, Jian Liu 0021, Nanjian Wu
ISNN (2)4
2019 A 0.45-to-1.8 GHz synthesized injection-locked bang-bang phase locked loop with fine frequency tuning circuits
Zhao Zhang 0004, Nan Qi 0002, Jian Liu 0021, Nanjian Wu
Sci. China Inf. Sci.5
2018 A Heterogeneous Parallel Processor for High-Speed Vision Chip
abstract
This paper proposes a heterogeneous parallel processor for high-speed vision chip. It contains four levels of processors with different parallelisms and complexities: processing element (PE) array processor, patch processing unit (PPU) array processor, self-organizing map (SOM) neural network processor, and dual-core microprocessor unit (MPU). The fine-grained PE array processor, middle-grained PPU array processor, and SOM neural network processor carry out image processing in pixel-parallel, patch-parallel, and distributed-parallel fashions, respectively. The MPU controls the overall system and executes some serial algorithms. The processor can improve the total system performance from low-level to high-level image processing significantly. A prototype is implemented with$64 \times 64$PE array,$8 \times 8$PPU array,$16 \times 24$SOM network, and a dual-core MPU. The proposed heterogeneous parallel processor introduces a new degree of parallelism, namely, patch parallel, which is for parallel local-feature extraction and feature detection. It can flexibly perform the state-of-the-art computer vision as well as various image processing algorithms at high speed. Various complicated applications, including feature extraction, face detection, and high-speed tracking, are demonstrated.
Jie Yang 0033, Yongxing Yang, Jian Liu 0021, Nanjian Wu
IEEE Trans. Circuits Syst. Video Technol.5
2018 A 0.9-2.25-GHz Sub-0.2-mW/GHz Compact Low-Voltage Low-Power Hybrid Digital PLL With Loop Bandwidth-Tracking Technique
Zhao Zhang 0004, Peng Feng 0001, Jian Liu 0021, Nanjian Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Terahertz detector for imaging in 180-nm standard CMOS process
Zhao-yang Liu, Zhao Zhang 0004, Jian Liu 0021, Nanjian Wu
Sci. China Inf. Sci.4
2015 A low power global shutter pixel with extended FD voltage swing range for large format high speed CMOS image sensor
Yangfan Zhou 0001, Zhongxiang Cao, Quanliang Li, Cong Shi 0003, Runjiang Dou, Jian Liu 0021, Nanjian Wu
Sci. China Inf. Sci.8