Peng Yin 0004

dblp:23/4378-4 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-5422-653XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 A Low-Latency Synchronization Header Detector and Hardware-Efficient Correction Decoder for JESD204C Receiver
abstract
The JESD204C receiver has emerged as a mainstream SerDes interface in high-speed analog-to-digital converters (ADCs). However, it faces challenges such as long link initialization time, substantial hardware area and power consumption. To address these issues, this paper presents an optimized link-layer design for the JESD204C receiver, with a focus on enhancing synchronization header (SH) detection efficiency and reducing the area of Cyclic Redundancy Check (CRC) and Forward Error Correction (FEC) decoding circuits. First, a dynamic search space(DSS)-based SH detection method is proposed. Leveraging multi-level XOR logic, this method efficiently eliminates non-transition bits from the data stream, enabling rapid localization of the 2-bit SH and thus shortening the synchronization delay during link establishment. Second, to mitigate the large area overhead of the data link layers decoding module, a CRC and FEC decoder (CAFD) with a logic-sharing architecture is designed: common factors in the decoding process are identified and shared, significantly reducing redundant hardware resources. Additionally, a low-complexity critical path identification algorithm is introduced to guide factor-sharing constraints, ensuring the decoder meets timing requirements while optimizing area. Finally, the proposed JESD204C receiver is implemented with a 28-nm CMOS process and verified on an FPGA platform. Experimental results show that the design not only eliminates redundant SH detection operations and accelerates link initialization/synchronization but also improves hardware resource utilization. Compared with state-of-the-art solutions, the hardware area and power consumption are reduced by 47.6% and 15.1%, respectively, while the SH locking time is shortened by 60%.
Peng Yin 0004, Rui Ma 0007, Jinlong Zhang, Mingguo Liu, Weizhou Hou, Shubin Liu 0001, Zhangming Zhu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 Theory and Low-Power Design of Moving Accumulative Sign Filter
abstract
A novel down-sampling filter named moving accumulative sign filter (MASF) is proposed for low-power down-sampling of large-scale binary and ternary data. Besides, the MASF has greatly circuit realization advantages than state-of-the-art cascaded-integrator-comb (CIC) filter, especially in the area of low-power design. The theory of MASF is proposed and introduced comprehensively, including the algorithm model, transfer function, and frequency response characteristics. The pipeline voting architecture is applied to the implementation of the MASF to improve the speed of data processing, which simplifies the circuit structure and reduce the power consumption. The MASF circuits of general application based on pipeline voting are designed for binary and ternary signals only using D flip-flop and logic gates. The area and power consumption of MASF are reduced by 86% and 88% compared with CIC filter under the same conditions on FPGA. What’s more, a hardware-friendly pooling algorithm named polar-pooling is proposed based on MASF for binary and ternary feature maps, which greatly reduces the time and space complexity of pooling. Compared with max-pooling and average-pooling, the processing time of polar-pooling is reduced by more than 75% for a$200\times 200$binary image. The two-stage MASF circuit for ternary signal processing is implemented at 40-nm CMOS process, compared with state-of-the-arts cascade-of-integrators filter which cascading two integrators, the normalized power consumption of proposed two-stage MASF circuit has 67% reduction and the area has 75% reduction.
Yingjun Xia, Jianjiang Luo, Peng Yin 0004, Dengwei Yan, Xichuan Zhou, Amine Bermak, Fang Tang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 High Logic Density Cyclic Redundancy Check and Forward Error Correction Logic Sharing Encoding Circuit for JESD204C Controller
abstract
Cyclic redundancy check (CRC) and Forward error correction (FEC) encoding are widely used in high-speed information transceiver systems such as PCIe, JESD204C and fiber-optic communications to detect or correct errors in data. Traditionally, the CRC and FEC encoding circuits in JESD204C are implemented independently of each other, which consumes a significant amount of hardware resources. Therefore, a high logic density CRC and FEC logic sharing (CFLS) encoding circuit for JESD204C controller is proposed in this paper, and the logic density of the encoding circuit is improved by sharing the registers and common encoding factor (CEF). Meanwhile, a straightforward critical path delay (CPD) calculation method was proposed to assess whether the data transmission delay satisfies the requirements of CFLS circuits. This method is derived in conjunction with the manipulation of the common factor matrix, thus reducing computational complexity. The CFLS encoding circuit proposed in this paper is verified with an FPGA platform, and the results show that the circuit can realize CRC and FEC function with a 21.96% reduction in hardware resources, compared to the traditional methods. The area of the JESD204C controller with CFLS encoding circuits is 0.09 mm2, by using a 40-nm CMOS process, and the power consumption is 24.66 mW according to the post-layout simulation.
Peng Yin 0004, Yingjun Xia, Jinlong Zhang, Mingguo Liu, Weizhou Hou, Amine Bermak, Fang Tang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A Sub-1/°C Bandgap Voltage Reference With High-Order Temperature Compensation in 0.18-μm CMOS Process
abstract
This paper presents a high-precision bandgap voltage reference (BGR) with high-order temperature compensation. The compensation signal is generated by using both strong-inversion MOSFETs and Bipolar Junction transistors (BJTs), which cancels the high-order nonlinear term$T\ln (T)$in the BJT base-emitter voltage (VBE), and thus a low temperature coefficient (TC) over a wide temperature range is achieved. The proposed BGR circuit is fabricated in a 0.18-$\mu \text{m}$CMOS process with an active area of$0.256 m{m^{2}}$and a max power consumption of 1.35 mW. A minimum TC of 0.706${\mathrm{ppm}}/{}^ \circ C$from$- 25\,\,{}^ \circ C$to 125${}^ \circ C$is achieved after an 8-bit resistance trimming. The line sensitivity is 0.0146%/V operating from 3.2 V to 3.7 V. The BGR achieves a power supply rejection (PSR) of −63.4 dB and a noise spectrum density of$0.92 ~\mu \text{V}/\sqrt {Hz} $at 10 Hz.
Shalin Huang, Peng Yin 0004, Amine Bermak, Fang Tang
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 A Low-Area and Low-Power Comma Detection and Word Alignment Circuits for JESD204B/C Controller
abstract
In an 8B/10B mode giga-bit-per-second serial data transactions, the de-serialized data is sent to a comma detection and word alignment (CDWA) module to identify the word boundaries, which is a prerequisite in the high-speed transceivers such as PCIe, USB and JESD204B/C. In order to ensure that the comma code (/K/-code) can be correctly detected. Ten 10-bit comma detector cells are adopted in a typical CDWA module, which require a complex circuitry and an enormous power consumption. To overcome these limitations, a low-area and low-power CDWA circuit for JESD204B/C transceiver chip in 8B/10B mode has been proposed in this paper. The bit width of the detector cells can be truncated from 10 to 6 under the condition, that CDWA module can detect a complete comma code correctly. On one hand, the proposed CDWA module is verified with a FPGA development platform with the reduction of the hardware resources and power consumption to 31.72% and 20.11% respectively as compared to the typical structure available. On the other hand, a 10-Gbps transceiver chip with the proposed CDWA module is fabricated with a 55-nm CMOS process and the word alignment function of the proposed module is proved by the measurement results. The area of this transceiver chip including 2× transmitting links and 2× receiving links is 2.89 mm2, and the power consumption is 467.8 mW, under a maximum data transmission rate of 10 Gbps.
Peng Yin 0004, Yingjun Xia, Tianmei Shen, Xiao Guan, Umar Mohammad, Jiandong Zang, Dongbing Fu, Xiaoping Zeng, Fang Tang, Amine Bermak
IEEE Trans. Circuits Syst. I Regul. Pap.1