Peng Wang 0220

dblp:95/4442-220 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-2631-971XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 A KT/C-Noise-Cancelled Charge-Sharing Sampling Pipeline ADC Front-End with Parallel Operation Timing and Substrate-Tracking Technique
Shiang Li, Wen Jia, Jingpeng Zhou, Peng Wang 0220, Fule Li, Zhihua Wang 0001
ISCAS4
2026 A 1-GS/s 12-bit Pipelined-SAR ADC With Dither-Based Background Calibration of Interstage Gain and Comparator Offset in 28-nm CMOS
Peng Wang 0220, Fule Li, Zhihua Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 CrossKV: Accelerating Large Language Model Inference via Cross-Stage Dynamic Co-Optimization for KV Cache
abstract
The autoregressive nature of Large Language Models (LLMs) has enabled remarkable performance in language generation, making them a cornerstone in natural language processing. As context windows lengthen, the per-token key–value (KV) cache grows linearly with sequence length and turns the generation stage into a memory-bound operation. However, existing works focus primarily on an isolated stage of KV cache reduction, and most efforts incur significant hardware overheads. This results in a lack of cross-stage co-optimization and leaves significant reduction potential untapped. To address these challenges, we propose CrossKV, a software-hardware co-design architecture that accelerates LLM inference to fully achieve the potential of KV caching optimization. Motivated by our observation, we identify the three stages in KV caching and present a cross-stage co-optimization algorithm, including: a derivative-enhanced dynamic progressive pruning method, a DCT-driven key-vector low-rank compressing, and a dynamic clustered hybrid run-length encoding to reduce the KV caching while maintaining high accuracy. Then, an efficient architecture featuring cross-stage dynamic co-optimization with negligible hardware overhead is proposed to fully harness the algorithm of CrossKV. Our evaluations of CrossKV over 9 popular LLMs models and various long-context tasks demonstrate an average of$6.03\times $and$4.10\times $improvement for energy efficiency and speedup compared to existing SoTA Accelerators architecture. Compared to the Nvidia A100 GPU, CrossKV achieves an average$23.43\times $energy efficiency and$4.64\times $speedup, respectively.
Shenyu Wang, Huizheng Wang, Peng Wang 0220, Xiao Liu 0001, Zhihua Wang 0001, Yang Hu 0001, Hanjun Jiang
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A SHA-Less Front-end Stage Structure Suitable for Time-Interleaved Pipeline ADC
abstract
This paper proposes a sample-and-hold amplifier (SHA)-less front-end stage structure suitable for pipeline ADCs in time-interleaving architecture with large interleaving factor. A low resolution SARADC is adopted as the sub-quantizer, and a bridge capacitor is introduced to connect the two capacitor arrays of the MDAC and the sub-ADC. Compared to commonly used pipelined-SAR ADC architectures, the proposed front-end stage structure reduces the reference load on the sub-quantizer during coarse conversion, while preserving the benefits of a shared sampling network in the sampling process. A 14-bit 500-MS/s pipeline ADC channel in an 8-way time-interleaved architecture utilizing the proposed front-end stage structure is designed in 28-nm CMOS process. Simulation results well demonstrate that the proposed front-end stage structure is highly immune to aperture errors, and shows the ADC channel achieves 69.13 dB SNDR and 83.8 dB SFDR at Nyquist frequency. With an estimated power consumption of 40-mW, a Walden FoM of 34.24-fJ/conv.-step is achieved.
Meng Ni, Peng Wang 0220, Fule Li, Hanjun Jiang, Zhihua Wang 0001
ISCAS2
2025 A High-Speed High-Precision Comparator with Offset Calibration for SHA_less Pipeline ADCs
abstract
This paper presents a high-speed, low-noise and low-offset dynamic latch comparator. In this work, a pair of auxiliary reset transistors is employed to increase the gain of the preamplifier while reducing the reset delay of the latch, which effectively improves the precision and speed of the comparator. The proposed comparator is controlled by two clocks, and the preamplifier is turned on before the regeneration phase to reduce the input-referred noise. In order to reduce power consumption and kickback noise, the preamplifier achieves dynamic operation by controlling a pair of cascode transistors while separating the regeneration node and the input pair. Otherwise, this work optimizes an efficient background calibration circuit to correct the offset. The proposed comparator has been designed and simulated in 12-nm FinFET CMOS technology. Simulation results show a delay of 20 ps at 10 mV input. The proposed comparator also achieves a 0.61 mV input-referred noise and an offset standard deviation of 1.9 mV.
Shuoxiong Yang, Peng Wang 0220, Fule Li, Zhihua Wang 0001
ISCAS3
2025 A Single-Channel 8-bit 1.6-GS/s Alternate-Comparator SAR ADC With Dither-Based Background Offset Calibration in 28-nm CMOS
abstract
This paper presents an 8-bit 1.6GS/s successive-approximation-register analog-to-digital converter (SAR ADC) with alternate comparators. To enhance dynamic performance and speed, a dither-based background relative offset calibration and a fast SAR logic are proposed. In the background calibration, the least significant bit (LSB) can indicate the polarity of the relative offset between comparators after injecting dither into the input, facilitating calibration without extra phase. The fast SAR logic, which includes a trigger logic and a digital-to-analog converter (DAC) control logic, features a 3-logic-gate delay, resulting in a reduced bit duration. These techniques are validated by a prototype 8-bit SAR ADC in 28 nm CMOS technology, exhibiting the highest speed (1.6 GS/s) and smallest area (0.0008 mm2) among single-channel 8-bit SAR ADC. With a Nyquist input, the measured signal-to-noise and distortion ratio (SNDR) improves from 31.6 dB to 43.6 dB, while the spurious-free dynamic range (SFDR) improves from 38.7 dB to 58.3 dB after offset calibration. It consumes 5.73 mW from 1-V supply, yielding a Walden FoM of 29.0 fJ/conv-step.
Peng Wang 0220, Fule Li, Zhihua Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1