EDBT 2026 Demo / reviewers in the wild / expert
Youxiang Chen
dblp:202/9678
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An OTA with Series-Cascode-Miller Compensation and Anti-Pole-Splitting for a 360-MHz BW Low-Distortion TIA in Sub-6G Broadband RF ReceiversabstractThis paper presents a fully differential 2-stage operational transconductance amplifier (OTA) for a baseband transimpedance amplifier (TIA) with 360-MHz bandwidth (BW) aiming for applications in Sub-6G broadband direct-conversion receivers. The proposed OTA with folded-cascode input stage incorporates three techniques, including series-cascode-Miller compensation (SCMC), anti-pole-splitting path, and common-mode feedforward (CMFF) to ensure wideband and low distortion of the TIA with shunt feedback. The OTA’s BW is extended because the SCMC introduces a zero that cancels the non-dominant pole at the folding node of the OTA, while the anti-pole-splitting path generates a negative capacitance that neutralizes the parasitic capacitance. In addition, the CMFF circuit improves the OTA’s common-mode rejection ratio (CMRR) by about 20 dB. Designed in 28-nm CMOS with a 0.12-mm2layout area and 16.6-mW power consumption, the TIA demonstrates an SFDR of 103 dB for a 75-MHz, 250-mVpp output, an input-referred noise (IRN) of 85 µVrms, an in-band third-order input intercept point (IIP3) of 37.9 dBm, and an intermodulation-free dynamic range (IMFDR3) of 91.6 dB. Junyao Ji, Yan Xue, Xiaojie Fan, Youxiang Chen, Ruibai You |
ISCAS | 6 |
| 2025 | A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN InferenceabstractComputing in memory (CIM) is one of the promising solutions to improve computing performance by integrating logic in memory. This work presents an efficient heterogeneous system based on the ultrafast CIM architecture (HS-CIM) to accelerate convolutional neural network (CNN) inference. First, an ultrafast CIM architecture is proposed based on the static random access memory (SRAM) by utilizing the novel total input and full digital scheme to implement multiply-and-accumulate (MAC) operation, which effectively addresses the high delay issue caused by high-precision computing in CIM architecture. Second, a heterogeneous system based on the proposed CIM architecture (HS-CIM) has been constructed with an aligned global cache and an adaptive pruning scheme to eliminate performance degradation and accuracy loss caused by input data bandwidth limitations. Meanwhile, efficient input data and weight data mapping schemes are proposed to minimize the delay and energy caused by input data transmission from the cache to the CIM architecture, thus realizing efficient CNN inference in the HS-CIM system. Finally, we analyze the performance of HS-CIM at the layout level by implementing the LeNet-5 and VGG models of CNN to recognize the image of MNIST and CIFAR-10 datasets, respectively. Results show that the energy efficiency of the proposed CIM architecture achieves 58.32 TOPS/W with 8-bit input/weight precision. Meanwhile, the inference accuracy of the HS-CIM system for MNIST and CIFAR-10 is 99.25% and 92.34%, respectively. Youxiang Chen, Zhengkun Gu, Haiming Qiu, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | FRM-CIM: Full-Digital Recursive MAC Computing in Memory System Based on MRAM for Neural Network ApplicationsabstractComputing in memory (CIM) realizes energy-efficient neural network algorithms by implementing highly parallel multiply-and-accumulate (MAC) operation. However, the MAC delay of CIM will sharply increase with the improvement of computing precision, which restricts its development. In this work, we propose a full-digital recursive MAC (FRM) operation based on spin-transfer-torque magnetic random access memory (STT-MRAM) CIM system to enable fast and energy-efficient image recognition application. First, the fast FRM scheme is proposed by utilizing the recursive operations of read and addition in segmented bit-line array, which effectively reduces the delay of MAC operations to 3.5ns and 4ns for 8-bit and 16-bit input and weight precision, respectively. Second, we design an image recognition system using FRM-CIM architecture as the processing element (PE), where the adaptive pruning method for layers is proposed to improve the compatibility of it with the neural network. By performing image recognition for the MNIST and CIFAR-10 datasets, results show that the throughput and energy efficiency of the FRM-CIM system are 58.51TOPS/mm2 and 11.3--56.72 TOPS/W under 8--16-bit precision, which are improved by 4.3 times and 2.6 times compared with the state-of-the-art works. Finally, the recognition accuracy can reach 96.65% and 82.7% on MNIST and CIFAR-10, respectively. Zhengkun Gu, Youxiang Chen, Weisheng Zhao 0001, Yue Zhang 0010 |
DAC | 5 |
| 2024 | RSACIM: Resistance Summation Analog Computing in Memory With Accuracy Optimization Scheme Based on MRAMabstractComputing in memory (CIM) has become a promising candidate to address the Von Neumann bottleneck in processors designed for data-intensive applications. In this article, we propose a resistance summation analog computing in memory (RSACIM) with accuracy optimization scheme in spin transfer torque magnetic random access memory (STT-MRAM), in order to realize energy-efficient and highly reliable analog multiply-and-accumulation (MAC) operation. Firstly, we construct a resistance summation array by serial magnetic tunnel junctions (MTJs) to perform analog MAC operation utilizing time domain technology. Secondly, in order to reduce the impact of position-dependent error caused by resistance summation mechanism, we propose an accuracy optimization scheme to maximize the sensing margin (SM) and computation accuracy. Finally, we design a power-gated reconfigurability control scheme to implement power saving corresponding to different precisions for both input and weight. Evaluation on a 2 Kb RSACIM architecture shows an energy efficiency of 92.9 TOPS/W. System level simulation shows that comparing to existing CIMs based on MRAM, RSACIM architecture saves the inference energy by 4.2 times with 8.4 times lower latency in CIFAR10 image classification task. Zhengkun Gu, Youxiang Chen, Kun Zhang 0030, Youguang Zhang, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2017 | Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAsabstractConvolutional neural networks (CNNs) have been widely applied in many deep learning applications. In recent years, the FPGA implementation for CNNs has attracted much attention because of its high performance and energy efficiency. However, existing implementations have difficulty to fully leverage the computation power of the latest FPGAs. In this paper we implement CNN on an FPGA using a systolic array architecture, which can achieve high clock frequency under high resource utilization. We provide an analytical model for performance and resource utilization and develop an automatic design space exploration framework, as well as source-to-source code transformation from a C program to a CNN implementation using systolic array. The experimental results show that our framework is able to generate the accelerator for real-life CNN models, achieving up to 461 GFlops for floating point data type and 1.2 Tops for 8-16 bit fixed point. Xuechao Wei, Cody Hao Yu, Peng Zhang 0007, Youxiang Chen, Yun Liang 0001, Jason Cong |
DAC | 4 |