VLDB 2026 Research / reviewers in the wild / expert
Jun Yin 0001
dblp:58/5423-1
· DBLP profile ↗
29ranked-venue papers
3as first author
24since 2021 · last 2026
0000-0002-4195-4551ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 3 first-author · 24 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDStream: An Energy-efficient 7.98 TBOPS/W Hyperdimensional Computing Streaming ProcessorabstractBinary hyperdimensional computing (HDC) is a brain-inspired framework that enables energy-efficient classification through simple bitwise operations on high-dimensional binary vectors. Existing accelerators face a fundamental compute-efficiency-density gap: encoding-specific designs achieve high efficiency at the cost of flexibility, while programmable processors sacrifice area and energy efficiency for generality. This work presents HDStream, a streaming HDC processor that closes this gap through (1) a wide-vector microarchitecture with HDC-customized multi-operation instructions and (2) autonomous streaming modules with hardware instruction loops. HDStream achieves up to 5.85 × speedup over single-operation-per-cycle execution with 99% compute utilization. Fabricated in 16 nm CMOS, HDStream achieves a peak 0.870 TBOPS and 7.98 TBOPS/W, with an effective 0.637 TBOPS and 7.65 TBOPS/W across diverse HDC workloads. Compared to prior programmable HDC accelerators, HDStream delivers up to 3.18 × higher compute performance density (TBOPS/mm2) and up to 2.99 × higher energy density (TBOPS/W/mm2). This work demonstrates that encoding flexibility and silicon area efficiency are not mutually exclusive. Ryan Antonio, Xiaoling Yi, Yunhao Deng, Fanchen Kong, Jun Yin 0001, Marian Verhelst |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
Xiaoling Yi, Ryan Antonio, Yunhao Deng, Fanchen Kong, Joren Dumoulin, Jun Yin 0001, Marian Verhelst |
ISCAS | 6 |
| 2026 | A Cryogenic HBT-CMOS Temperature Sensor Operating From 4 to 70 KabstractIn current cryogenic temperature sensor (cryo-TS) systems, the sensing front-end and readout circuits typically operate in cryogenic and room-temperature environments, respectively. This paper proposes a scheme to integrate both the front-end devices and readout circuits of cryo-TS within the cryogenic environment to achieve lower noise, digital fan-out of temperature information, and cost reduction. We employed the silicon-germanium (SiGe) heterojunction bipolar transistors (HBT), which demonstrated excellent linearity and current gain even at cryogenic temperatures, as the sensing front end of the cryo-TS and a Zoom-ADC as its readout circuits. A redundancy bit is introduced in the cryogenic readout ADC to avoid temperature misjudgment. The design methodology and key considerations for implementing cryogenic readout analog circuits are presented. Implemented in a 65 nm CMOS process, the cryo-TS achieved a 1-point-trimmed (at 40 K) inaccuracy of ±0.54 K ($\boldsymbol {3\sigma }$) from 4 K to 70 K under a supply current of 22.13$\mu A$. Chen Deng, Wenhua Gong, Yatao Peng, Jun Yin 0001, Jing Wang 0131, Jad Benserhir, Lin Cheng 0001, Edoardo Charbon, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | An Accuracy-and-Efficiency-Configurable Blade-Type Approximate Multiplier With Genetic Algorithm-Based Automatic Training Framework
Kaize Zhou, Zhihao Yan, Zhangrui Qian, Zhuo Chen 0039, Jun Yin 0001, Yan Lu 0002, Weiwei Shan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A Low Phase Noise Circular-Coupled Quad-Core Oscillator Employing Dual-Path Synchronization TechniqueabstractThis article presents a millimeter-wave (mm-Wave) quad-core oscillator utilizing a circular transformer and dual-path synchronization technique to achieve low phase noise (PN). We analyze the mechanism of the PN degradation induced by frequency mismatch between individual cores in the conventional quad-core oscillator utilizing a circular inductor, which reveals that the frequency mismatch between the non-adjacent cores will induce an off-resonance issue, which could significantly degrade the PN. This PN degradation induced by the off-resonance is acerbated in the quad-core oscillator utilizing a circular transformer. Based on the analysis, a dual-path synchronization technique is proposed to eliminate the off-resonance issue by providing a direct synchronization path for either two cores in a circular-coupled quad-core oscillator. The proposed technique can be applied to both inductor- and transformer-based circular-coupled oscillators, preventing PN degradation due to the frequency mismatch. Fabricated in 65-nm CMOS, our quad-core oscillator prototype measures a frequency tuning range of 18.2%, from 22.4 to 26.8 GHz. At the 25.8 GHz carrier frequency, the oscillator achieves a low PN of$\!-\!138$dBc/Hz at 10 MHz offset while consuming 19.7 mW, corresponding to an excellent FoM of 193.3 dBc/Hz. Xiangxun Zhan, Pui-In Mak, Rui Paulo Martins, Jun Yin 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A High Bandwidth Capacitorless NMOS LDO with Pole-Tracking Scheme and Adaptive FFRC Achieving -65dB PSR across Full Load RangeabstractThis article presents a capacitorless linear low dropout regulator (LDO) with an N-type pass transistor that features high efficiency, fast transient and good power supply rejection (PSR) suitable for system-on-chip (SoC) integration. This LDO utilizes a pole-tracking compensation scheme and a transconductance-boosted (gm-boosted) folded-cascode EA to achieve high bandwidth. As a result, the LDO achieves a fast transient response and high PSR corner-frequency. Moreover, feedforward ripple cancellation (FFRC) is also incorporated to enhance the PSR performance further. The proposed LDO is fabricated in a 65nm CMOS process. It features PSR performance better than -65dB across the full load range at frequencies lower than 100kHz. And its undershoot and overshoot in response to a 1mA-to-100mA load transient with an edge time of 100ns are 83mV and 10mV, respectively. Jinshuo Xu, Jun Yin 0001, Mo Huang, Yan Lu 0002 |
ISCAS | 4 |
| 2025 | A 0.6V Digital-intensive Pulse Injection 32-kHz Crystal Oscillator Using Stacked Logic GatesabstractThis paper introduces an ultra-low-power 32kHz pulse injection crystal oscillator (PIXO) using a reference-free stacked inverter chain to generate the delay signal for energy injection. If the power consumption of both the pulse driver and the inverter-based delay generation circuits is considered, analysis reveals that the closer the injection happens at the peak and valley of the output waveform, the lower the system power efficiency. The proposed stacked logic gates can generate an injection pulse that is 48° away from the zero-crossing point of the output waveform without resorting to the process- and temperature-sensitive pico-ampere-level current reference and duty-cycle-correction loop for injection timing control. The proposed PIXO can properly operate across −40°C to 120°C at five process corners (TT/SS/SF/FS/FF). Implemented in a 65nm CMOS technology, at a 0.6V voltage supply, post-layout simulations verify that the PIXO achieves 1.2nW power consumption at a 32kHz injection rate at 25°C. Thanks to its digital-intensive architecture, the PIXO occupies only 0.0069mm2active area. Zhizhan Yang, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
ISCAS | 2 |
| 2025 | A Systematic Review of Voltage Reference Circuits: Spanning Room Temperature to Cryogenic ApplicationsabstractCryo-CMOS IC for quantum applications, proposed for tens of years, are designed to control quantum processors operating at cryogenic temperatures (CTs). The reference circuits play a significant role in quantum controllers, providing a relatively stable biasing for analog and radio frequency (RF) circuit blocks. Based on a literature review, we discovered that achieving high-accuracy reference voltage or current at CTs is challenging due to the unstable temperature characteristics of complementary metal-oxide-semiconductor (CMOS), bipolar junction transistor (BJT), or resistors in the general CMOS process at CTs. Therefore, certain specialized device structures, such as dynamic threshold MOS (DTMOS), can be employed within the bulk CMOS process. Alternatively, BJT and other devices found in specific processes, such as silicon-germanium (SiGe) and fully depleted silicon on insulator (FD-SOI) CMOS, can achieve adaptive temperature compensation. This paper provides a succinct overview of several fundamental structures and common research hot spots about the reference voltage circuits, and then assesses their suitability for CT circuit design, considering the reliability of devices in bulk CMOS, FD-SOI CMOS, and SiGe process. Finally, the paper summarizes the types of cryo-temperature reference circuits and offers an overview and comparison of them. Chen Deng, Sai Wu, Yatao Peng, Man Kay Law, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Analyses Concerning the Phase Noise and Nonlinear Behavior of the Charge-Sharing Integrator-Based Hybrid PLLabstractHybrid PLLs (HPLLs) leverage the advantages of conventional analog and digital PLLs to cater to the high integration demand inherent with the advanced CMOS nodes, among which the charge-sharing (CS) integrator-based HPLL featuring ultra-compact area and ultra-low power consumption exhibits promising prospects. This work presents a comprehensive analysis concerning the CS integrator-based HPLL for the first time. The behavioral modeling is first conducted with a time-domain event-driven modeling method to simulate the piecewise-linear modulations on the oscillator frequency A noise model considering the slow-fast clock domain transitions and the digital-analog signal transitions inherent with the architecture is further proposed. The proposed models offer precise PN PSD estimations over the specified frequency range under all simulated conditions, whose integrated jitters differ by less than 5 fs compared with circuit simulations. The additional nonlinearity induced by the CS integrator is also considered, and an analytical prediction approach for the nonlinearity-induced spurs is presented, showing a capability of predicting the locations and amplitudes of the most significant spurious tones with a less than 4 % deviation in the relative offset frequency and a less than 5 dBc deviation in the relative amplitude. Jingrun Song, Yueduo Liu, Zhengxuan Han, Zehao Zhang, Jiaxin Liu 0001, Hongshuai Zhang, Jun Yin 0001, Pui-In Mak, Shiheng Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2025 | Design and Analysis of a Type-II Sampling PLL With Automatic Frequency and Phase Calibrations Achieving 0.62-μs Locking TimeabstractThis paper presents a type-II sampling phase-locked loop (SPLL) that accelerates the locking process by exploiting a time-to-digital converter (TDC) based automatic frequency and phase calibration (AFPC) technique. The proposed AFPC accelerates the frequency acquisition by using a type-I loop to map the quantized phase error to the switched-capacitor (SC) control word of the voltage-controlled oscillator (VCO). The subsequent phase error after frequency locking is swiftly reduced within one TDC resolution by adjusting the division ratio of the multi-modulus divider (MMD) with little hardware expenditure. The proposed AFPC can guarantee a fast-locking time at different initial frequencies, which is insensitive to the variation of TDC resolution. This paper also contributes to a design strategy for the AFPC loop, e.g., the required frequency step of the SC and the TDC resolution, based on the analysis of the lock-in range of the SPLL. Fabricated in 28-nm CMOS with a core area of 0.15 mm$^{\mathbf {2}}$, the 6.0-to-6.9GHz SPLL prototype using a reference (REF) clock of 100 MHz achieves a locking time of$0.62~\mu $s ($62{T} _{\mathbf {REF}}$) at an 880-MHz hopping frequency. At 6.5 GHz, the SPLL consumes 4.6 mW and measures an RMS jitter and REF spur of 99 fs and –71.6 dBc, respectively, corresponding to a jitter figure-of-merit (FoM$_{\mathbf {jitter}}$) of –253.5 dB. Tailong Xu, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak, Quan Pan 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | An 840-to-970 MHz Multimodal Wake-Up Receiver With a Q-Equalized Antenna-ED Interface and 2-Dimensional Wake-Up IdentificationabstractThe article introduces an antenna-envelope detector (ED) co-designed wake-up receiver (WuRX) that can automatically detect the frequency-hopping sequence from 840 to 970 MHz. Prototyped in 65nm CMOS, the WuRX supports three modes: 1) low-power mode that achieves −68 dBm sensitivity with 9.9nW power consumption, 2) Q-enhanced mode that provides 22 dB rejection to an on-off-keying (OOK) modulated pseudo-random-bit-sequence blocker at 10 MHz offset and 41 dB rejection to a continuous-wave one at 10 MHz offset, with a power consumption of$39.6~\mu $W and 3) 2-stage wake-up mode that combines the advantages of both the low-power mode and Q-enhanced mode, with a power consumption of 33.7 nW. These characteristics are achieved by exploiting 1) an antenna-ED interface where the Q-factor of the antenna is equal to the Q-factor of the capacitor to ensure a conjugate matching condition for the maximum available power transfer from the antenna to the ED, 2) a frequency tuner connected to the antenna-ED interface calibrated by a frequency-locked loop, and 3) a Q-booster reconfigured from the frequency tuner for boosting the passive gain and narrowing the bandwidth of the interface. Zhizhan Yang, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Radio-Frequency Cross-Connected Rectifier With LC Source DegenerationabstractThis paper proposes a radio-frequency (RF) cross-connected (CC) rectifier with LC source degeneration (CCLC). The LC network resonates at twice of the working RF frequency, shaping the drain-source and gate-source voltages of the rectifying transistors of the CC rectifier. As a result, the proposed scheme reduces the transistors’ reverse current and the shoot-through current, and thus the root mean square current of the switches, improving the power conversion efficiency (PCE). Meanwhile, it has a better input impedance matching over a wide input power range from the voltage reshaping. Subsequently, we implement the CCLC topology in both one-stage and two-stage CC rectifiers. To reduce the silicon area, we design coupled inductors for the two LC networks of the two-stage CCLC rectifier. We fabricated the rectifiers using a 65-nm CMOS process. The measured PCE of the one-stage and two-stage rectifiers are 67.6% and 61.2%, respectively. The proposed scheme has a wider dynamic range than previous works. Qiujin Chen, Mo Huang, Jun Yin 0001, Haiwen Liu, Rui Paulo Martins, Yan Lu 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | A 5.6-dB Noise Figure, 63-86-GHz Receiver Using a Wideband Noise-Cancelling Low Noise Amplifier With Phase and Amplitude CompensationabstractIn this paper, a 63–86-GHz receiver (RX) is proposed using a wideband noise-cancelling low noise amplifier (LNA) with phase and amplitude compensation circuit. Such phase and amplitude compensation circuit consists of a 4-to-1 asymmetric power-combining transformer and two amplitude adjusting amplifiers, which improves the noise-cancelling ratio of mm-wave LNA. Meanwhile, a wideband active mixer and an absorptive IF amplifier are introduced to provide a reflectionless operation for signal quality improvement. Based on aforementioned structures, a wideband RX is implemented and fabricated using a conventional 40-nm CMOS technology. The measured results show that the RX achieves 22.5-dB conversion gain with 23-GHz 3-dB bandwidth and 50.7-mW power consumption. The minimum NF is 5.6-dB with less than 1.9-dB variation in the whole operation band. Changxuan Han, Zhixian Deng, Yiyang Shu, Jun Yin 0001, Pui-In Mak, Xun Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | A Compact Sub-nW/kHz Relaxation Oscillator Using a Negative-Offset Comparator With Chopping and Piecewise Charge-Acceleration in 28-nm CMOSabstractThis work presents a compact and power-efficient kHz-range relaxation (RC) oscillator with robust performance against temperature and voltage variations. By deliberately introducing a negative-offset voltage into the comparator, an offset cancellation scheme leveraging chopping and piecewise charge-acceleration facilitates a low temperature coefficient. A low-power comparator with a tail resistor and a low oscillation amplitude improves the energy efficiency. The die area is compact by introducing leakage-based temperature compensation that eliminates bulky resistors and complex calibration. Prototyped in a 28-nm CMOS process and measured at 28.5 kHz, our oscillator occupies 0.0046 mm$^{2}$and dissipates 27.6 nW at a 0.8-V supply. The energy efficiency is 0.97 nW/kHz, and the temperature coefficient is 33.3 ppm/$^{\circ}$C over$-$40 to 85$^{\circ}$C with 1-point calibration. The corresponding FoM of 164.9 dB compares favorably with the recent arts. The start-up time is rapid, and the period settling time is within one cycle of$\sim$5.7$\mu$s. The Allan deviation is$\le $40 ppm for measurement intervals of$>$0.5 s. Yueduo Liu, Rongxin Bao, Jiahui Lin, Jun Yin 0001, Qiang Li 0021, Pui-In Mak, Shiheng Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Analysis and Design of a 21.2-to-25.5-GHz Triple-Coil Transformer-Coupled QVCOabstractThis paper reports a triple-coil transformer-coupled quadrature voltage-controlled oscillator (TC-QVCO), which inherently provides the quadrature signal without using the noisy active-coupling transistors. The determinate correlation of tank voltages is verified by utilizing the initial state to facilitate the oscillation state analysis. Thus, the TC-QVCO would operate without the oscillation mode ambiguity. Additionally, thanks to the triple-coil transformer coupling, a large source coil$L_{S}$aids in achieving in-phase coupling for phase noise (PN) improvement, and the intensified coupling factor$k_{gd}$benefits reducing the PN and the quadrature phase error simultaneously. Therefore, our TC-QVCO would alleviate the tradeoff between PN and quadrature phase accuracy via using a large$L_{S}$and$k_{gd}$. The proposed QVCO prototyped in 65-nm CMOS exhibits a superior FoM$_{\text {@10MHz}}$(180.1 to 182.2 dBc/Hz) over a 18.2% frequency tuning range (21.2 to 25.5 GHz), and the estimated quadrature phase error <0.8°. Jun Yin 0001, Pui-In Mak, Li Geng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Real-Time Acoustic Perception for Automotive ApplicationsabstractIn recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions. While the focus has mostly been on visual sensors, also acoustic events are crucial to detect situations that require a change in the driving behavior, such as a car honking, or the sirens of approaching emergency vehicles. In this paper, we summarize the results achieved so far in the Marie Sklodowska-Curie Actions (MSCA) Eruopean Industrial Doctorates (EID) project “Intelligent Ultra Low-Power Signal Processing for Automotive (I-SPOT)”. On the algorithmic side, the I-SPOT Project aims to enable detecting, localizing and tracking environmental audio signals by jointly developing microphone array processing and deep learning techniques that specifically target automotive applications. Data generation software has been developed to cover the I-SPOT target scenarios and research challenges. This tool is currently being used to develop low-complexity deep learning techniques for emergency sound detection. On the hardware side, the goal impels workflows for hardware-algorithm co-design to ease the generation of architectures that are sufficiently flexible towards algorithmic evolutions without giving up on efficiency, as well as enable rapid feedback of hardware implications of algorithmic decision. This is pursued though a hierarchical workflow that breaks the hardware-algorithm design space into reasonable subsets, which has been tested for operator-level optimizations on state-of-the-art robust sound source localization for edge devices. Further, several open challenges towards an end-to-end system are clarified for the next stage of I-SPOT. Jun Yin 0001, Stefano Damiano, Marian Verhelst, Toon van Waterschoot, Andre Guntoro |
DATE | 1 |
| 2023 | ACCO: Automated Causal CNN Scheduling Optimizer for Real-Time Edge AcceleratorsabstractSpatio-Temporal Convolutional Neural Networks (ST-CNN) allow extending CNN capabilities from image processing to consecutive temporal-pattern recognition. Generally, state-of-the-art (SotA) ST-CNNs inflate the feature maps and weights from well-known CNN backbones to represent the additional time dimension. However, edge computing applications would suffer tremendously from such large computation/memory overhead. Fortunately, the overlapping nature of ST-CNN enables various optimizations, such as the dilated causal convolution structure and Depth-First (DF) layer fusion to reuse the computation between time steps and CNN sliding windows, respectively. Yet, no hardware-aware approach has been proposed that jointly explores the optimal strategy from a scheduling as well as a hardware point of view.To this end, we present ACCO, an automated optimizer that explores efficient Causal CNN transformation and DF scheduling for ST-CNNs on edge hardware accelerators. By cost-modeling the computation and data movement on the accelerator architecture, ACCO automatically selects the best scheduling strategy for the given hardware-algorithm target. Compared to the fixed dilated causal structure, ST-CNNs with ACCO reach an ~8.4× better Energy-Delay-Product. Meanwhile, ACCO improves ~20% in layer-fusion optimals compared to the SotA DF exploration toolchain. When jointly optimizing ST-CNN on the temporal and spatial dimension, ACCO’s scheduling outcomes are on average 19× faster and 37× more energy-efficient than spatial DF schemes. Jun Yin 0001, Linyan Mei, Andre Guntoro, Marian Verhelst |
ICCD | 1 |
| 2023 | Analysis and Design of a 15.2-to-18.2-GHz Inverse-Class-F VCO With a Balanced Dual-Core Topology Suppressing the Flicker Noise UpconversionabstractThis paper presents the theory and implementation of a balanced dual-core inverse-class-F (class-F$^{\mathrm{ -1}}$) voltage-controlled oscillator (VCO). The class-F$^{\mathrm{ -1}}$topology supports high-quality-factor (high-$Q$) differential switched-capacitors (SCs) for both fundamental and 2$^{\mathrm{ nd}}$-harmonic frequency tuning, which is beneficial for improving the phase noise (PN). However, the unequal parasitic capacitors from the NMOS and PMOS negative${g}$textsubscript m transistors make it impossible to minimize their flicker noise upconversions simultaneously, especially at high operating frequencies. The mechanism of this effect is analyzed qualitatively with the model of coupled oscillators and verified using the impulse-sensitivity function (ISF) approach. To address this issue, we propose a dual-core class-F$^{\mathrm{ -1}}$VCO that leverages a balanced coupling scheme to minimize the flicker noise upconversions of NMOS and PMOS transistors simultaneously and still keep the advantage of tuning the 2$^{\mathrm{ nd}}$-harmonic frequency with differential SCs offered by the class-F$^{\mathrm{ -1}}$topology. Additionally, the symmetrical circuit topology aids in improving the differential output balancing. Prototyped in a 28-nm CMOS process without ultra-thick metal, the balanced dual-core class-F$^{\mathrm{ -1}}$VCO dissipates 19.7-mW and achieves a PN of$\!-\!113.9/\!-\!135.8$-dBc/Hz at 1/10-MHz offset from an 18.23-GHz carrier. Tuned from 15.22 to 18.23-GHz, the proposed VCO exhibits superior figure-of-merits (FoMs) at 1/10-MHz offset from 185.3/187.0 to 186.2/188.1-dBc/Hz. Haoran Li 0016, Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | CNN-based Robust Sound Source Localization with SRP-PHAT for the Extreme EdgeabstractRobust sound source localization for environments with noise and reverberation are increasingly exploiting deep neural networks fed with various acoustic features. Yet, state-of-the-art research mainly focuses on optimizing algorithmic accuracy, resulting in huge models preventing edge-device deployment. The edge, however, urges for real-time low-footprint acoustic reasoning for applications such as hearing aids and robot interactions. Hence, we set off from a robust CNN-based model using SRP-PHAT features, Cross3D [ 16 ], to pursue an efficient yet compact model architecture for the extreme edge. For both the SRP feature representation and neural network, we propose respectively our scalable LC-SRP-Edge and Cross3D-Edge algorithms which are optimized towards lower hardware overhead. LC-SRP-Edge halves the complexity and on-chip memory overhead for the sinc interpolation compared to the original LC-SRP [ 19 ]. Over multiple SRP resolution cases, Cross3D-Edge saves 10.32%~73.71% computational complexity and 59.77%~94.66% neural network weights against the Cross3D baseline. In terms of the accuracy-efficiency tradeoff, the most balanced version (EM) requires only 127.1 MFLOPS computation, 3.71 MByte/s bandwidth, and 0.821 MByte on-chip memory in total, while still retaining competitiveness in state-of-the-art accuracy comparisons. It achieves 8.59 ms/frame end-to-end latency on a Rasberry Pi 4B, which is 7.26× faster than the corresponding baseline. Jun Yin 0001, Marian Verhelst |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | A 0.0043-mm2 0.085-μW/MHz Relaxation Oscillator Using Charge-Prestored Asymmetric Swings R-RC NetworkabstractIn this brief, a charge-prestored 21.2-MHz relaxation oscillator is proposed for ultralow-power applications. It occupies only 0.0043 mm2in 0.18-$\mu \text{m}$CMOS by resistor reusing and is reference-free. The simulated temperature coefficient (TC) of the output frequency is 15.2 ppm/° from −30 °C to 125 °C. By generating an asymmetric capacitor charging swing, our charge-prestored technique reduces significantly the power consumed by the swing-boostingRCnetwork during the charging phase. Also, the R-RCstructure further improves the energy efficiency. The total power consumption of the oscillator core is$1.806 \mu \text{W}$at 0.8 V, corresponding to an energy efficiency of$0.085 \mu \text{W}$/MHz that compares favorably with the state of the art. Shiheng Yang, Yueduo Liu, Rongxin Bao, Jiahui Lin, Zehao Zhang, Yong Chen 0005, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2022 | A 529-μW Fractional-N All-Digital PLL Using TDC Gain Auto-Calibration and an Inverse-Class-F DCO in 65-nm CMOSabstractThis paper presents an ultra-lower-power (ULP) digital-to-time-converter (DTC)-assisted fractional-N all-digital phase-locked loop (ADPLL) suitable for IoT applications. A proposed hybrid time-to-digital converter (TDC) extends the vernier-TDC input range with little power overhead in order to overcome the stability issue in the conventional architectures. The hybrid TDC also facilitates a background gain calibration to achieve a stable in-band phase noise insensitive to process, voltage, and temperature (PVT) variations. The implementation of a buffer-cascaded DTC simplifies the design complexity of the fractional-N operation. The ADPLL also features a 200$\mu \text{W}$low-phase-noise inverse-class-F (class-F−1) digitally controlled oscillator (DCO) without the need of two-dimensional (2-D) capacitor tuning for frequency alignment of the fundamental and 2nd-harmonic. Fabricated in 65-nm CMOS, the ULP ADPLL prototype achieves 868fsrmsjitter in a fractional-N channel when consuming only 529$\mu \text{W}$, corresponding to a figure-of-merit (FoM) of −244dB. Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins, Robert Bogdan Staszewski |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Mismatch Analysis of DTCs With an Improved BIST-TDC in 28-nm CMOSabstractNonlinearity of a digital-to-time converter (DTC) is pivotal to spur performance in DTC-based all-digital phase-locked-loops (ADPLL). In this paper, we characterize and analyze the mismatch of cascaded-delay-unit DTCs. Through an improved built-in-self-test (BIST) time-to-digital converter (TDC) assisted with phase-to-frequency detector (PFD), a measurement system of sub-half-ps accuracy is constructed to conduct the characterization. Fabricated in 28-nm CMOS, the DTC transfer functions are measured, and mismatches are compared against Monte-Carlo simulation results. The integral nonlinearity (INL) results are compared against each other and converted to the in-band fractional spur level when the DTC would be deployed in the ADPLL. The BIST-TDC system thus characterizes the on-chip delays without expensive equipment or complex setup. The effectiveness of adding a PFD into the$\Delta \!\Sigma $loop is validated. The entire BIST system consumes 0.6mW with a system self-calibration algorithm to tackle the analog blocks’ nonlinearities. Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins, Robert Bogdan Staszewski |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Accurate Performance Evaluation of Jitter-Power FOM for Multiplying Delay-Locked LoopabstractAn accurate performance evaluation of jitter-power figure-of-merit (FOM) for multiplying delay-locked loop (MDLL) is presented. For a typical MDLL employing a single-ended multiplying-delay ring voltage-controlled oscillator (MDVCO), it can be tuned via the capacitive loads to uphold a constant normalized phase noise (PN) across different frequencies for better jitter-power performance. Linear approximation and z-domain PN model are utilized to simplify the analysis with excellent agreement between the time-domain simulation and z-domain approximation. The influences of the asymmetric waveform, flicker noise corner, frequency error and reference noise are also discussed and then ignored based on the reasonable approximation. Under a given process, reference clock frequency and supply voltage, the ideal FOM can be derived theoretically. Based on these insights, the predicted FOM can be the benchmark metric in the design of MDLL for the possible best jitter-power performance. Yueduo Liu, Rongxin Bao, Shiheng Yang, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A 6-to-7.5-GHz 54-fsrms Jitter Type-II Reference-Sampling PLL Featuring a Gain-Boosting Phase Detector for In-Band Phase-Noise ReductionabstractThis paper presents a type-II reference-sampling (RS) phase-locked loop (PLL) exploiting a novel gain-boosting reference-sampling phase detector (RSPD) to reduce the in-band phase noise and RMS jitter. The proposed gain-boosting RSPD converts the phase error to the voltage error and utilizes a passive switched-capacitor voltage multiplier to amplify the sampled voltage error, which effectively increases the gain of the RSPD. The boosted RSPD gain helps suppress the phase noise contributed from the gm cell. To prevent the transistors in the gain-boosting RSPD and gm cell from breakdown, the gain-boosting function is only activated during the locked state when the phase and voltage errors are small. The switching between the gain-boosting and normal modes is realized automatically by monitoring the sampled voltage and comparing it with a pre-defined threshold window. Fabricated in 65-nm CMOS, the type-II RS-PLL measures an RMS jitter of 54 fs at 6.75 GHz and consumes 7.1 mW, corresponding to a jitter figure-of-merit (FoMjitter) of −256.8 dB. The measured reference spur is −62.6 dBc, and the active area is 0.25 mm2. Tailong Xu, Shenke Zhong, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | A 3.15-mW +16.0-dBm IIP3 22-dB CG Inductively Source Degenerated Balun-LNA Mixer With Integrated Transformer-Based Gate Inductor and IM2 Injection TechniqueabstractThis article proposes two linearization techniques in improving the third-order input intercept point (IIP3) of a balun-low-noise amplifier (LNA) mixer. First, the intrinsic third-order intermodulation (IM3) product of the inductively source degenerated (ISD) transconductor from the second-order derivative transconductance component (g"m) is reduced by tailoring toward the optimum biasing point at the moderate-inversion region. Second, the generated IM3 current by the first-order derivative transconductance (g'm) due to the interaction with the feedback component in the ISD transconductor is attenuated by second-harmonic injection via the bulk of the ISD transconductor. Furthermore, a transformer-based gate inductor and a transformer-based balun are applied to improve the input impedance matching and produce a balanced differential input signal. Measured results in 0.13-μm CMOS show a high IIP3 of +16 dBm and a conversion gain (CG) of 22 dB at 2.4 GHz. The double-sideband (DSB) noise figure (NF) is 7.2 dB, and the power consumption is 3.15 mW at 1.2 V. Nandini Vitee, Harikrishnan Ramiah, Pui-In Mak, Jun Yin 0001, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | A 1-V 4-mW Differential-Folded Mixer With Common-Gate Transconductor Using Multiple Feedback Achieving 18.4-dB Conversion Gain, +12.5-dBm IIP3, and 8.5-dB NFabstractThis article reports a novel differential-folded mixer with multiple-feedback techniques for performance enhancement. Specifically, we introduce the capacitor cross-coupled (CCC) common-gate (CG) transconductance stage to improve the noise figure (NF) at low power by boosting the effective transconductance, while enhancing the linearity via suppressing the second-order harmonic distortion. Typically, the created loop gain of the CCC can raise the third-order intermodulation (IM3) distortion, penalizing the input-referred third-order intercept point (IIP3). Here, we propose a positive and a second capacitive feedback into the CCC CG transconductor, not only to suppress the IM3 distortion current but also adds in design flexibility to the input transistors. Furthermore, the positive feedback also improves the input impedance matching, conversion gain, and NF through a flexible design criterion. Prototyped in a 0.13-μm process, the proposed mixer operating at 900 MHz dissipates 4 mW at 1 V. The measured double sideband (DSB) NF is 8.5 dB, the conversion gain (GC) is 18.4 dB and the IIP3 is +12.5 dBm. Nandini Vitee, Harikrishnan Ramiah, Pui-In Mak, Jun Yin 0001, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | A coin-battery-powered LDO-Free 2.4-GHz Bluetooth Low Energy/ZigBee receiver consuming 2 mA
Zechariah Balan, Harikrishnan Ramiah, Jagadheswaran Rajendran, Nandini Vitee, Pravinah Nair Shasidharan, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
Integr. | 6 |
| 2019 | Many-Objective Sizing Optimization of a Class-C/D VCO for Ultralow-Power IoT and Ultralow-Phase-Noise Cellular ApplicationsabstractIn this paper, the performance boundaries and corresponding tradeoffs of a complex dual-mode class-C/D voltage-controlled oscillator (VCO) are extended using a framework for the automatic sizing of radio frequency integrated circuit blocks, where an all-inclusive test bench formulation enhanced with an additional measurement processing system enables the optimization of “everything at once” toward its true optimal tradeoffs. VCOs embedded in the state-of-the-art multistandard transceivers must comply with extremely high performance and ultralow power requirements for modern cellular and Internet of Things applications. However, the proper analysis of the design tradeoffs is tedious and impractical, as a large amount of conflicting performance figures obtained from multiple modes, test benches, and/or analysis must be considered simultaneously. Here, the dual-mode design and optimization conducted provided 287 design solutions with figures of merit above 192 dBc/Hz, where the power consumption varies from 0.134 to 1.333 mW, the phase noise at 10 MHz from -133.89 to -142.51 dBc/Hz, and the frequency pushing from 2 to 500 MHz/V, on the worst case of the tuning range. These results pushed this circuit design to its performance limits on a 65-nm CMOS technology, reducing 49% of the power consumption of the original design while also showing its potential for ultralow power with more than 93% reduction. In addition, worst case corner criteria were also performed on the top of the worst case tuning range optimization, taking the problem to a human-untrea table LXVI-D performance space. Ricardo Martins 0003, Nuno Lourenço 0003, Nuno Horta, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | A high-Q spiral inductor with dual-layer patterned floating shield in a class-B VCO achieving a 190.5-dBc/Hz FoMabstractThis paper proposes a dual-layer patterned floating shield (DL-PFS) technique for Silicon-based on-chip spiral inductors. By optimally utilizing the two lowest metal layer strips to shield the inductor from the substrate, electromagnetic (EM) simulations show 40% improvement of the Q factor when compared with the conventional approach. Designed and simulated in 0.13-μm CMOS, the DL-PFS inductor in a class-B VCO achieves 6.6-dB lower phase noise, and 34% power savings. The VCO also exhibits 9.7-to-10.93 GHz tunability, and -123-dBc/Hz phase noise at a 3 MHz offset. The power consumption is 1.64 mW at 0.6 V, leading to a state-of-the-art FoM of 190.5 dBc/Hz. Chee-Cheow Lim, Harikrishnan Ramiah, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
ISCAS | 3 |