EDBT 2026 Demo / reviewers in the wild / expert
Yangyi Zhang 0002
dblp:206/9538-2
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 2 × 80 Gb/s Single-Ended TAS-TIS PAM-4 Receiver Front-End With Crosstalk Cancellation and Signal Reutilization in 28-nm CMOSabstractThis paper presents a$2\times 80$Gb/s single-ended trans-admittance trans-impedance (TAS-TIS) 4-level pulse amplitude modulation (PAM-4) receiver with crosstalk cancellation and signal reutilization technique in 28-nm CMOS. Based on the TAS-TIS architecture, to achieve a precise cancellation of the far-end crosstalk, a common mode gain rejection methodology is proposed in the balanced differentiator, ensuring low gain mismatch and low phase skew. Moreover, the proposed Gm-doubler technique in the TAS-TIS architecture enhances the gain of the adder under limited supply voltage conditions. By employing the series peaking combined with an active inductor in the adder, the bandwidth of adder is improved by a factor of 2.2, and the group delay variation is effectively optimized by 65%. The proposed current-mirror-based continuous-time linear equalizer (CTLE) provides high-frequency and low-frequency compensation with 25% power saving. Overall, the measurement results of the proposed receiver demonstrate$2\times 80$Gb/s PAM-4 eyes with an efficiency of 0.83 pJ/bit/lane and$2\times 56$Gb/s non-return-to zero (NRZ) eyes over a pair of PCB traces with 13 dB loss at 20 GHz and 28 dB loss at 28 GHz, respectively. Yangyi Zhang 0002, Liping Zhong, Taiyang Fan, Xiongshi Luo, Hongzhi Wu, Xuxu Cheng, Dongfan Xu, Quan Pan 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | A 112-Gb/s Single-Ended PAM-4 Transceiver Front-End for Reach Extension in Long-Reach LinkabstractA 112-Gb/s single-ended (SE) four-level pulse amplitude modulation (PAM-4) transceiver front-end for the reach-extension module in long-reach (LR) link is proposed. The receiver front-end features an SE-to-differential (S2D) amplifier and a continuous-time linear equalizer (CTLE). Asymmetric inductive peaking, compensation capacitance, and current blending techniques are employed in S2D to eliminate the mismatch at the pseudo-differential outputs. A compact and peaking-enhanced CTLE is achieved by the inductor reused technique. The transmitter front-end is based on a differential-to-SE (D2S) driver where the negative capacitance technique is proposed to extend its bandwidth. Fabricated in 130-nm SiGe BiCMOS technology, our SE transceiver front-end demonstrates a data rate of 112-Gb/s PAM-4 at a 20-dB channel loss with an FoM of 0.09 pJ/bit/dB and BER of$3.21e$-4. Xiongshi Luo, Xuewei You, Jiahan Fu, Liping Zhong, Mengjie Song, Taiyang Fan, Hongzhi Wu, Yangyi Zhang 0002, Chenchang Zhan, Quan Pan 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2025 | ReHIT: Reconfigurable High-Radix Iterative-Taylor Architecture for Ultraprecise Logarithm/Exponential Functions in FPGA-Based Softmax AcceleratorsabstractThe softmax function, as a pivotal component in neural network accelerators, imposes stringent demands on the precision-efficiency tradeoff for logarithmic and exponential computations. This article presents reconfigurable high-radix iterative-Taylor (ReHIT) architecture, a novel hardware framework that synergistically integrates high-radix iterative normalization with optimized Taylor expansion to achieve subunit-in-the-last-place (ULP) precision in floating-point transcendental functions. Our key innovation lies in the hierarchical pretreatment mechanism where high-radix iterations (radix-256/512) systematically decompose input operands into normalized subdomains, enabling subsequent quadratic Taylor approximations with guaranteed convergence. This codesign methodology reduces polynomial orders by 33% compared to conventional approaches while eliminating resource-intensive division operations through shift-and-add transformations. The implemented ReHIT-logarithm (ReHIT-L) and ReHIT-exponential (ReHIT-E) modules demonstrate configurable precision scaling from half to double precision (FP16/32/64), validated through exhaustive error analysis over 1 000 000 random test vectors with worst case errors bounded at 0.78 ULP. Field-programmable gate array (FPGA) implementations on Arria 10/Virtex-7 platforms can achieve up to 13.8% logic resources reduction and 32.3% latency improvement over state-of-the-art designs, with post-synthesis results in Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm showing up to$1.34\times $giga operations per second (GOPS)/W energy efficiency and$3.97\times $GOPS/mm2area efficiency for softmax acceleration. The reconfigurable pipeline of the architecture permits dynamic precision/throughput adaptation, particularly beneficial for quantized neural networks requiring FP16–FP32 hybrid precision. Yangyi Zhang 0002, Lei Chen 0001, Fengwei An |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Live Demonstration: A Video Denoising Co-processor with Non-local Means Algorithm for FHD 30fps Image SensorabstractIn this demonstration, a non-local means (NLM) video denoising co-processor with data reuse scheme and dual-clock domain for high resolution image sensor is presented. With an OV5640 camera, the real-time denoising processor can be performed at 30 frames/second for full high definition (FHD 1920×1080) RAW video, a debayer filter then decodes the RAW format to RGB format. Ruoheng Yao, Shengming Zhou, Zhiyue Gao, Yangyi Zhang 0002, Yiwei Luo, Lei Chen 0001, Fengwei An |
ISCAS | 4 |
| 2023 | A Fully-Integrated LDO with Two-Stage Cross-Coupled Error Amplifier for High-Speed Communications in 28-nm CMOSabstractThis paper presents a fully-integrated flipped-voltage-follower-based low-dropout regulator (LDO), with proposed high-gain two-stage cross-coupled error amplifier (XCEA). Besides, the effectiveness of bypass capacitors and diversified load capacitors is discussed. Consuming$\mathbf{170}-\boldsymbol{\mu} \mathbf{A}$quiescent current and occupying area of 0.019 mm2, the LDO features 25-MHz unity-gain bandwidth (UGB) at 20-mA load to satisfy fast response requirement in high-speed transmitter. The simulated voltage undershoot is 38.85 mV for a load transient current stepping from$\mathbf{1}\ \boldsymbol{\mu} \mathbf{A}$to 25 mA in 50 ps with$\boldsymbol{C}_{\mathbf{L}} =\mathbf{100}\ \mathbf{pF}$. Owing to the proposed XCEA and the filter capacitor, the PSR is measured to be -41 dB at 100 kHz and -36 dB at 1 MHz. Dongfan Xu, Yangyi Zhang 0002, Xiongshi Luo, Pingyi Cai, Hongzhi Wu, Liping Zhong, Liru Zhu, Quan Pan 0002 |
ISCAS | 2 |
| 2023 | Anti-Aliasing and Anti-Color-Artifact Demosaicing for High-Resolution CMOS Image SensorabstractDemosaicing is a technique that reconstructs an RGB image from fragmentary color samples sensed by the image sensor. The color filter array (CFA), which is placed over the image sensor, determines the color of each pixel. The most used color filter array is the Bayer CFA. This paper proposes an Anti-Aliasing and Anti-Color-Artifact Demosaicing (AAACA) algorithm for the Bayer pattern and the resource-efficient very-large-scale integration (VLSI) architecture for the proposed algorithm. The AAACA comprises an anti-aliasing approach and a color artifacts filter named color difference-based median filter (CDMF). Compared to the traditional demosaicing methods that equally treat pixels on and not on edges, the AA reconstructs the pixels on edges with different strategies from those not on edges, significantly removing the aliasing around recovered edges. Then the CDMF is utilized to remove the color artifacts of the reconstructed RGB image based on the color difference after median filtering. We respectively simulate the quantitative evaluation and subjective visual quality on McMaster and Kodak datasets. Our experiments reveal that the proposed AAACA algorithm can significantly remove the visual aliasing around recovered edges and greatly reduce color artifacts in demosaiced images compared to the state-of-art demosaicing algorithms. The proposed VLSI architectures can achieve superior visual qualities compared with the previous VLSI implementations under the same process technology conditions with 180nm CMOS technology, an image resolution of$1280\times 720$(HD), and a working frequency of 200MHz. Yangyi Zhang 0002, Zizhao Peng, Lei Chen 0001, Fengwei An |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |