Xiangyu Meng 0003

dblp:16/8083-3 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-7956-8499ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 A Sample-and-Hold-Based 453-ps True Time Delay Circuit With a Wide Bandwidth of 0.5-2.5 GHz in 65-nm CMOS
abstract
Delay solutions applied to high frequencies typically involve switched transmission lines or all-pass filters. These solutions often suffer from significant insertion loss and drastic gain variations at high frequencies, along with poor delay flatness. In this work, we have designed a delay circuit that can be applied to high frequencies, featuring excellent delay flatness, good delay resolution, and a wide bandwidth. In this design, a multistage cascaded sampling circuit is used to generate delays. By introducing differential clocks or three-phase clocks, simple coarse delay or fine delay can be achieved. The measurement results show that the sample and hold circuit achieves a delay accuracy of 17.5 ps and a delay range of 453 ps within 0.5–2.5 GHz, with a gain of −2.7 to 2 dB and a gain variation of ±0.85 dB, a delay variation less than 7.5 ps, a power consumption of 111 mW, and a core area of 0.137 mm2.
Chuanjie Chen, Xiangyu Meng 0003, Wang Xie, Baoyong Chi
IEEE Trans. Very Large Scale Integr. Syst.2
2024 A V-Band Low-Phase-Noise VCO with Transformer-Based Gm-Boosting Technique
abstract
In this paper, we proposed a V-band low-phase-noise and low-power voltage-control oscillator (VCO). To address the issue of Gmdegeneration with increasing oscillation frequency, we introduce a transformer-based Gm-boosting technique designed to augment the negative Gmwithin a specific frequency range. Furthermore, for a wide frequency tunning range (FTR), a 3-bit switch capacitor array is designed to ensure the frequency of the negativeGmpeak aligns with oscillation frequency. Moreover, the push-push technique is employed to extract and the second harmonic, enabling the VCO to operate at its fundamental frequency. It avoids the use of small inductors which lead to deterioration of the Q value. Designed in 65-nm CMOS, the post-layout simulation results exhibit a low phase noise (PN) of -103.7 dBc/Hz at 1-MHz offset and a figure-of-merit (FoM) of 193.7 dBc/Hz from a 58.6-GHz carrier, while the FTR is 15.9% from 50.5 to 59.2 GHz. The proposed VCO consumes 3.6mW at 0.8-V supply volage.
Xiangyu Meng 0003, Fangfei Ming
ISCAS2
2021 3D-VNPU: A Flexible Accelerator for 2D/3D CNNs on FPGA
abstract
Three-dimensional convolutional neural networks (3D CNNs) have proven to be outstanding in applications such as video analysis, 3-dimension geometric data, and 3-dimension medical image diagnosis. Compared to 2D CNNs, 3D CNNs require high computational complexity to get spatio-temporal features while Winograd algorithm can significantly reduce the amount of computation. Prior works based on 3D Winograd accelerators are only applied to stride-1 convolution, however, most of the popular 3D CNNs contain stride-2 convolution layers. In this paper, we propose a novel flexible Winograd-based decomposition method (FWDM) to apply the 3D Winograd to different strides convolution. Evaluation results show that FWDM reduces computational complexity by a factor of 3.2 for C3D, 2.9 for 3D ConvNet, and 2.6 for 3D ResNet-18. Furthermore, we design a flexible computing engine to stretch the use range of the decomposition method. Coupling FWDM and computing engine, a Winograd-based, 2D/3D CNNs compatible, highly efficient, and flexible accelerator (3D-VNPU) is proposed. Finally, we demonstrate the effectiveness of 3D-VNPU on FPGA platform (Xilinx ZCU102) and achieve 1.35TOPS for C3D, 1.2TOPS for 3D ResNet-18, and 1.1TOPS for VGG-16. DSP efficiency outperforms other CNNs accelerators 2.57~15.3x compared with prior works in FPGA of C3D. Compared to GPU and CPU, our accelerator achieves improvement up to 37.9x in performance relative to CPU and 11.8x in energy efficiency relative to GPU.
Huipeng Deng, Jian Wang 0080, Huafeng Ye, Shanlin Xiao, Xiangyu Meng 0003, Zhiyi Yu
FCCM5
2021 A 1.8-GS/s 6-Bit Two-Step SAR ADC in 65-nm CMOS
abstract
This paper presents a 2-bit/cycle 2-step hybrid successive-approximation-register (SAR) analog-to-digital converter (ADC) with 6-bit resolution. A latch consisting of dynamic logic gates is proposed to speed up the ADC conversion process and largely enhance the power efficiency. Furthermore, in the 2-step structure, the logic of the two-stage ADC is modified, and the capacitance arrays of the DACs are accordingly adjusted so that the performance of each stage is optimized. Designed and simulated in a 65-nm CMOS technology, the proposed single channel ADC achieves SNDR and SFNR of 36.76 and 46.89 respectively, resulting in 5.8 ENOB, the THD is -42.22dB, SNR is 38.22 dB, and FoM at Nyquist rate is 85.9 fJ/conversion-step. The total power consumption of the ADC is 8.6 mW at a sampling rate up to 1.8 GS/s.
Xiangyu Meng 0003, Weihao Kong, Yecong Li
ISCAS1
2020 W-Band Synthesized Modulator and Demodulator with Wideband Performance in 65-nm CMOS
abstract
A high image-rejection in-phase/quadrature (IQ) modulator and a wideband demodulator with large dynamic range operating in W-band is implemented in 65-nm CMOS technology. The modulator demonstrates a flat conversion gain of 12.5±0.7dB, a minimum image-rejection ratio of 40 dBc and a maximum LO-to-RF leakage of 38 dB from 90 to 98 GHz. The conversion gain of the demodulator is digitally controlled from 15 to 46 dB with a noise figure from 13.5 to 11.5 dB and an input 1dB compression point (IP1dB) from -13 to -39 dBm. The modulator and the demodulator are synthesized by two on-chip single-pole-double-throw (SPDT) switches and one IF switch. The entire system occupies 1.4 mm2chip area with a total power consumption of 165 mW.
Zhenpeng Zheng, Xiangyu Meng 0003, Dihu Chen
ISCAS2
2020 A 0.5-5 GHz 0.3-mW 50% duty-cycle corrector in 65-nm CMOS
abstract
A Duty Cycle Corrector (DCC) with wide operation frequency band, wide duty-cycle correction range, high duty accuracy and low power consumption performance is proposed in this paper. A dual feedback loop with differential input clock is added to reduce the impact of charge pump imbalance on circuit performance. The chopping technique is also introduced to improve the loop gain meanwhile suppress the DC offset in the feedback loop. Furthermore, a novel duty-cycle adjuster (DCA) with configurable load capacitance is presented to maintain the duty-cycle correction range in a wide frequency range. The proposed DCC is implemented in 65-nm CMOS process with 1-V supply voltage. Post-simulation results indicate that the DCC corrects the input duty cycle with a range from 20% to 80% to 50±0.7% in the 0.5-5 GHz frequency range. The maximum power consumption of the DCC is 0.32 mW when the input clock frequency is set to 5 GHz.
Xiangyu Meng 0003
TENCON2