Rongxin Bao

dblp:210/8413 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2024 A Compact Sub-nW/kHz Relaxation Oscillator Using a Negative-Offset Comparator With Chopping and Piecewise Charge-Acceleration in 28-nm CMOS
abstract
This work presents a compact and power-efficient kHz-range relaxation (RC) oscillator with robust performance against temperature and voltage variations. By deliberately introducing a negative-offset voltage into the comparator, an offset cancellation scheme leveraging chopping and piecewise charge-acceleration facilitates a low temperature coefficient. A low-power comparator with a tail resistor and a low oscillation amplitude improves the energy efficiency. The die area is compact by introducing leakage-based temperature compensation that eliminates bulky resistors and complex calibration. Prototyped in a 28-nm CMOS process and measured at 28.5 kHz, our oscillator occupies 0.0046 mm$^{2}$and dissipates 27.6 nW at a 0.8-V supply. The energy efficiency is 0.97 nW/kHz, and the temperature coefficient is 33.3 ppm/$^{\circ}$C over$-$40 to 85$^{\circ}$C with 1-point calibration. The corresponding FoM of 164.9 dB compares favorably with the recent arts. The start-up time is rapid, and the period settling time is within one cycle of$\sim$5.7$\mu$s. The Allan deviation is$\le $40 ppm for measurement intervals of$>$0.5 s.
Yueduo Liu, Rongxin Bao, Jiahui Lin, Jun Yin 0001, Qiang Li 0021, Pui-In Mak, Shiheng Yang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 MispredTable: A Side Branch Predictor to TAGE in Multithreading Processors
abstract
Tagged geometric history length (TAGE) branch predictor shows good prediction accuracy with sufficient storage budget. However, shrinking TAGE's storage budget dramatically reduces its prediction accuracy. In this paper, we propose MT (Misprediction Table)-TAGE, which adds a side predictor to the TAGE branch predictor. This side branch predictor will record the branches with the highest misprediction rate during the running process of the program and modify the prediction result of the branch when the number of mispredictions is higher than dynamic threshold. We also propose a new algorithm that enables MT-TAGE to be implemented in multithreaded processors. This work is written in SystemVerilog and tested on a RISC-V multithreaded processor. The experimental results show that in a quad-thread processor, the misprediction rates of MT-TAGE under the storage budget of 32kbits and 8kbits are 5.03MPKI (misprediction per kilo-instructions) and 5.54MPKI, respectively, which are 12.5% and 15.5% less than TAGE under the same area.
Xincheng Yang, Songping Mai, Rongxin Bao
ISCAS3
2023 A 0.0043-mm2 0.085-μW/MHz Relaxation Oscillator Using Charge-Prestored Asymmetric Swings R-RC Network
abstract
In this brief, a charge-prestored 21.2-MHz relaxation oscillator is proposed for ultralow-power applications. It occupies only 0.0043 mm2in 0.18-$\mu \text{m}$CMOS by resistor reusing and is reference-free. The simulated temperature coefficient (TC) of the output frequency is 15.2 ppm/° from −30 °C to 125 °C. By generating an asymmetric capacitor charging swing, our charge-prestored technique reduces significantly the power consumed by the swing-boostingRCnetwork during the charging phase. Also, the R-RCstructure further improves the energy efficiency. The total power consumption of the oscillator core is$1.806 \mu \text{W}$at 0.8 V, corresponding to an energy efficiency of$0.085 \mu \text{W}$/MHz that compares favorably with the state of the art.
Shiheng Yang, Yueduo Liu, Rongxin Bao, Jiahui Lin, Zehao Zhang, Yong Chen 0005, Jun Yin 0001, Pui-In Mak, Qiang Li 0021
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Accurate Performance Evaluation of Jitter-Power FOM for Multiplying Delay-Locked Loop
abstract
An accurate performance evaluation of jitter-power figure-of-merit (FOM) for multiplying delay-locked loop (MDLL) is presented. For a typical MDLL employing a single-ended multiplying-delay ring voltage-controlled oscillator (MDVCO), it can be tuned via the capacitive loads to uphold a constant normalized phase noise (PN) across different frequencies for better jitter-power performance. Linear approximation and z-domain PN model are utilized to simplify the analysis with excellent agreement between the time-domain simulation and z-domain approximation. The influences of the asymmetric waveform, flicker noise corner, frequency error and reference noise are also discussed and then ignored based on the reasonable approximation. Under a given process, reference clock frequency and supply voltage, the ideal FOM can be derived theoretically. Based on these insights, the predicted FOM can be the benchmark metric in the design of MDLL for the possible best jitter-power performance.
Yueduo Liu, Rongxin Bao, Shiheng Yang, Jun Yin 0001, Pui-In Mak, Qiang Li 0021
IEEE Trans. Circuits Syst. I Regul. Pap.2
2018 Cross-Entropy Pruning for Compressing Convolutional Neural Networks
abstract
The success of CNNs is accompanied by deep models and heavy storage costs. For compressing CNNs, we propose an efficient and robust pruning approach, cross-entropy pruning (CEP). Given a trained CNN model, connections were divided into groups in a group-wise way according to their corresponding output neurons. All connections with their cross-entropy errors below a grouping threshold were then removed. A sparse model was obtained and the number of parameters in the baseline model significantly reduced. This letter also presents a highest cross-entropy pruning (HCEP) method that keeps a small portion of weights with the highest CEP. This method further improves the accuracy of CEP. To validate CEP, we conducted the experiments on low redundant networks that are hard to compress. For the MNIST data set, CEP achieves an 0.08% accuracy drop required by LeNet-5 benchmark with only 16% of original parameters. Our proposed CEP also reduces approximately 75% of the storage cost of AlexNet on the ILSVRC 2012 data set, increasing the top-1 errorby only 0.4% and top-5 error by only 0.2%. Compared with three existing methods on LeNet-5, our proposed CEP and HCEP perform significantly better than the existing methods in terms of the accuracy and stability. Some computer vision tasks on CNNs such as object detection and style transfer can be computed in a high-performance way using our CEP and HCEP strategies.
Rongxin Bao, Xu Yuan 0002, Zhikui Chen, Ruixin Ma
Neural Comput.1