VLDB 2026 Research / reviewers in the wild / expert
Pui-In Mak
dblp:81/6224
· DBLP profile ↗
96ranked-venue papers
2as first author
56since 2021 · last 2026
0000-0002-3579-8740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 87 · 2 first-author · 53 since 2021Artificial intelligence and machine learning · 4Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid Dynamic/SRAM based TCAM Cell with Self-Gating Match Line and Adaptive VSS
Zuqi Zhang, Chenghe Sun, Xiangyi Chu, Rui Paulo Martins, Pui-In Mak |
ISCAS | 7 |
| 2026 | A Fast-Transient Buck Converter with Oversampled Multi-Phase-Ramp PWM Control
Zhenyu Shen, Xiongjie Zhang, Xiacong Liu, Yang Jiang 0002, Rui Paulo Martins, Pui-In Mak |
ISCAS | 6 |
| 2026 | A Topology-Aware Reinforcement Learning Framework for 2.4-GHz VSWR-Robust Power Amplifier Matching Network Design
Bingbing Zhao, Wei-Han Yu, Fábio Passos, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak |
ISCAS | 7 |
| 2026 | A Cryogenic HBT-CMOS Temperature Sensor Operating From 4 to 70 KabstractIn current cryogenic temperature sensor (cryo-TS) systems, the sensing front-end and readout circuits typically operate in cryogenic and room-temperature environments, respectively. This paper proposes a scheme to integrate both the front-end devices and readout circuits of cryo-TS within the cryogenic environment to achieve lower noise, digital fan-out of temperature information, and cost reduction. We employed the silicon-germanium (SiGe) heterojunction bipolar transistors (HBT), which demonstrated excellent linearity and current gain even at cryogenic temperatures, as the sensing front end of the cryo-TS and a Zoom-ADC as its readout circuits. A redundancy bit is introduced in the cryogenic readout ADC to avoid temperature misjudgment. The design methodology and key considerations for implementing cryogenic readout analog circuits are presented. Implemented in a 65 nm CMOS process, the cryo-TS achieved a 1-point-trimmed (at 40 K) inaccuracy of ±0.54 K ($\boldsymbol {3\sigma }$) from 4 K to 70 K under a supply current of 22.13$\mu A$. Chen Deng, Wenhua Gong, Yatao Peng, Jun Yin 0001, Jing Wang 0131, Jad Benserhir, Lin Cheng 0001, Edoardo Charbon, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2026 | A 0.5-V Ultra-Low Voltage Relaxation Oscillator With Identical Asymmetric Swing-Boosted RC Network and Feedback-Based Amplifier Achieving 390-ppm RMS Period Jitter for Self-Powered DevicesabstractThis paper reports an ultra-low voltage (ULV) relaxation oscillator (RxO) suitable for self-powered devices, designed with a pair of asymmetric swing-boosted (ASB) RC networks. This work enhances low-voltage operational capabilities and improves frequency stability and jitter performance. The RxO features a unique single amplifier configuration incorporated with a customized feedback mechanism that effectively compares the output voltages from the RC networks, substantially reducing jitter due to flicker noise. Additionally, we implement a Duty-Cycling Circuit (DCC) based on a DLL architecture to turn on the amplifier before the desired detection point, providing ample guard time and thereby reducing power consumption, which is essential for ultra-low power applications. The RxO also features a Replica Temperature Compensation Circuit (RTCC) to mitigate circuit delay. Fabricated in 65-nm CMOS, the RxO operates at 2.35 MHz with a minimal supply voltage of 0.5 V, achieving a period jitter of 390 ppm and line sensitivity of 17.4%, and an energy efficiency of 5.82 pJ/cycle. The device demonstrates significant improvements over existing ULV designs, achieving up to 60% reduction in power consumption while maintaining lower jitter levels. Mikki How-Wen Loo, Harikrishnan Ramiah, Dan Shi 0008, Chee-Cheow Lim, Rui Paulo Martins, Pui-In Mak, Ka-Meng Lei |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A 24-MHz Crystal Oscillator With 6.9-$μ$s Startup Time and 2% Injection-Δ$F$ Tolerance Using Phase-Interpolator-Assisted Synchronized InjectionabstractThis article presents a 24-MHz fast startup crystal oscillator (XO) with a phase-interpolator-assisted synchronized injection technique. The technique ensures phase consistency between the injection source and the crystal resonance, even with a 2% injection-$\Delta F$, enhancing the robustness of the startup under different PVT conditions. Additionally, we propose a differential peak detection technique to detect the phase error incurred by$\Delta F$. Such a peak detection technique shortens the auxiliary non-injection period during startup to merely four cycles, thereby maximizing the injection percentage to 97.4% and enhancing the startup’s efficacy. Fabricated in the 40-nm CMOS process, the XO achieves a 6.9-$\mu $s startup time (166 cycles) with a startup energy of 4.8 nJ under a 1-V$V_{\text {DD}}$. Furthermore, the startup time varies by ±4.4%, ±3.6%, and ±5.1% (worst case) over$\Delta F$(0.25% to 2%), temperature (−40 to$85~^{\circ }$C), and$V_{DD}$(0.95 to 1.05 V) variations, respectively. The XO’s phase noise in the steady-state is −137.8dBc/Hz at the 1-kHz offset, with a power consumption of$63~\mu $W. Xin Wang 0147, Shanhu Wang, Ka-Meng Lei, Jiafei Yao, Zixuan Wang 0022, Zhikuang Cai, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2026 | A Low Phase Noise Circular-Coupled Quad-Core Oscillator Employing Dual-Path Synchronization TechniqueabstractThis article presents a millimeter-wave (mm-Wave) quad-core oscillator utilizing a circular transformer and dual-path synchronization technique to achieve low phase noise (PN). We analyze the mechanism of the PN degradation induced by frequency mismatch between individual cores in the conventional quad-core oscillator utilizing a circular inductor, which reveals that the frequency mismatch between the non-adjacent cores will induce an off-resonance issue, which could significantly degrade the PN. This PN degradation induced by the off-resonance is acerbated in the quad-core oscillator utilizing a circular transformer. Based on the analysis, a dual-path synchronization technique is proposed to eliminate the off-resonance issue by providing a direct synchronization path for either two cores in a circular-coupled quad-core oscillator. The proposed technique can be applied to both inductor- and transformer-based circular-coupled oscillators, preventing PN degradation due to the frequency mismatch. Fabricated in 65-nm CMOS, our quad-core oscillator prototype measures a frequency tuning range of 18.2%, from 22.4 to 26.8 GHz. At the 25.8 GHz carrier frequency, the oscillator achieves a low PN of$\!-\!138$dBc/Hz at 10 MHz offset while consuming 19.7 mW, corresponding to an excellent FoM of 193.3 dBc/Hz. Xiangxun Zhan, Pui-In Mak, Rui Paulo Martins, Jun Yin 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | A 3.51 TOPS/mm2 Transformer Accelerator Exploiting Bipolar Sparsity and Approximate GatingabstractTransformer models excel at natural language processing tasks but are challenging to deploy on edge devices due to high memory and computation demands. To address this, we proposed an energy- and area-efficient transformer accelerator. We identify ‘0’ bits in positive and ‘1’ bits in negative 2’s complement activation values as bipolar sparsity. This form of sparsity shares a larger proportion than the traditional bit-level sparsity in transformer models. We propose a bipolar sparsity compressor (BSC) together with a bipolar processing element (BPE) to detect and skip the bipolar sparsity in a bit-group (BG) level during the inference. It reduces a large proportion of ineffective computations and improves throughput. The significant sparsity scheduling (SSS) dynamically adjusts broadcast settings based on BG-level sparsity ratios, balancing the sparsity skipping and memory access. Furthermore, an importance approximate gating (IAG) filters out unimportant tokens/heads during the attention computation by reusing sparsity information from the BSC, further reducing processing latency and energy consumption. Implemented in a 28nm process, the proposed accelerator achieves$7.62\times $and$12.43\times $throughput improvements on RoBERTa-B and GPT2-xl, respectively. The area efficiency reaches up to 3.51 TOPS/mm2due to skipping a large proportion of bipolar sparsity, reaching$4.83\times$and$5.85\times $higher compared with the state-of-the-art approximate computing accelerator and the computing in memory accelerator on the benchmark model. Zhongyu Zhao, Rujian Cao, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Ground-to-Chassis Distance Adaptive Photovoltaic Inductive Wireless Power Transfer System for Electric VehiclesabstractElectric vehicles (EVs) have gained significant adoption due to their advantages over conventional vehicles. The Photovoltaic Inductive Wireless Power Transfer (PV-IWPT) system combines wireless charging technology with photovoltaic panels to provide clean, solar-powered EV charging. However, varying distances between the ground and EV chassis create inconsistent coupling coefficients, causing fluctuating optimal load conditions for the IWPT converter and making it challenging to maintain peak efficiency without manual recalibration. This paper introduces a Ground-to-Chassis distance adaptive control approach for PV-IWPT systems that simultaneously achieve Maximum Power Point Tracking (MPPT) for solar harvesting and Maximum Efficiency Point Tracking (MEPT) for power transfer across different ground-to-chassis distances. The system automatically adapts to positional variations without requiring hardware adjustments. We validate our approach through theoretical analysis and experimental verification using a 500-W test platform under multiple distance configurations and solar shading conditions, demonstrating its practical viability for real-world applications. Io-Wa Iam, Zongrui Yang, Chi-Fong Ieong, Pui-In Mak, Rui Paulo Martins, Chi-Seng Lam |
IECON | 4 |
| 2025 | A Linear-Regression-Assisted Trimming Scheme for CMOS Voltage ReferenceabstractThis work proposes a single-point linear-regression-assisted trimming method for a MOSFET-based voltage reference to reduce operation verification complexity. By obtaining the correlation between the input features at a unique temperature and the output voltages across the operating temperature range from layout-aware Monte-Carlo simulations, we can build the linear regression model to predict the profile of the voltages across the temperature range based on eight input features. Hence, we can apply the appropriate trimming code on the circuits, avoiding time-consuming temperature characterization. We design the voltage reference in 65nm CMOS and obtained 8,000 sets of simulation data to train the regression model. Validated through simulation, we reduce the temperature coefficient of the voltage reference to 64.7ppm/°C using the proposed scheme, 26% lower than that of the conventional 2-point trimming, evincing the efficiency and accuracy of the linear-regression-assisted trimming. Chengyu Che, Xinfei Guo, Ka-Meng Lei, Rui Paulo Martins, Pui-In Mak |
ISCAS | 6 |
| 2025 | A BW-Extended Multi-Band Receiver with High-Order N-Path Filtering at RF Front-End and BB Achieving 200MHz BW and 36.5dBm OB-IIP3abstractThis work presents a radio-frequency (RF) blocker-tolerant multi-band (0.5 to 2GHz) receiver to achieve both wide passband bandwidth (BW) and sharp out-of-band (OB) attenuation at near-by frequency. At RF front-end, it features the combination of the 4th-order N-path low-noise transconductance amplifier (LNTA) and the bottom-plate N-path filter, while at baseband (BB) it introduces the 3rd-order N-path lowpass filter and 2nd-order BB trans-impedance amplifier (TIA). Besides, a negative-feedback frequency-translational loop is created to achieve the input-impedance matching condition. Designed in 65nm CMOS technology, our receiver achieves -220dB/decade near-by roll-off slope for a 200MHz RF-BW at 2GHz. With 3xRF-BW offset, the receiver achieves 36.5dBm OB-IIP3. The noise figure (NF) is simulated from 2.4 to 3.8dB and the power consumption is 27.6 to 37.8mW. The chip area is 0.68mm2. Gengzhen Qi, Yunchu Li, Shaolin Liao, Pui-In Mak |
ISCAS | 4 |
| 2025 | A 0.6V Digital-intensive Pulse Injection 32-kHz Crystal Oscillator Using Stacked Logic GatesabstractThis paper introduces an ultra-low-power 32kHz pulse injection crystal oscillator (PIXO) using a reference-free stacked inverter chain to generate the delay signal for energy injection. If the power consumption of both the pulse driver and the inverter-based delay generation circuits is considered, analysis reveals that the closer the injection happens at the peak and valley of the output waveform, the lower the system power efficiency. The proposed stacked logic gates can generate an injection pulse that is 48° away from the zero-crossing point of the output waveform without resorting to the process- and temperature-sensitive pico-ampere-level current reference and duty-cycle-correction loop for injection timing control. The proposed PIXO can properly operate across −40°C to 120°C at five process corners (TT/SS/SF/FS/FF). Implemented in a 65nm CMOS technology, at a 0.6V voltage supply, post-layout simulations verify that the PIXO achieves 1.2nW power consumption at a 32kHz injection rate at 25°C. Thanks to its digital-intensive architecture, the PIXO occupies only 0.0069mm2active area. Zhizhan Yang, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
ISCAS | 4 |
| 2025 | A 0.4V Relaxation Oscillator featuring Double Capacitor-Charging Headroom in CMOS 65nmabstractLow-Power fully-integrated oscillators are the cornerstone of Internet-Of-Things devices due to their compactness, high energy-efficiency, and scalability. This paper presents a relaxation oscillator featuring double capacitor-charging headroom by applying chopping on the capacitor. Such an increase in the charging headroom soothes the frequency instability attributable to the comparator and logic gate’s delay, culminating in a more stable frequency output amid voltage and temperature variations. We designed and fabricated two 0.4V relaxation oscillators (803kHz and 428kHz) in the TSMC 65nm process. The measured temperature coefficients are 164 and 106 ppm/°C (averaged from 10 samples) across −20 to 120°C for the 803kHz and 428kHz relaxation oscillators, and the line sensitivities are 15.8%/V and 14.4%/V across 0.35 to 0.5V. Such results are improved by >1.8× and >11× compared with the reference oscillator with regular charging headroom. The oscillators’ long-term stabilities are 150 and 110 ppm with a 0.1s gating interval. Kanghong Yu, Mingrui Wang, Ka-Meng Lei, Rui Paulo Martins, Pui-In Mak |
ISCAS | 5 |
| 2025 | A Chip-based Miniature MRI Platform with Integrated PDMS-PCB Coil Frontend for Microlitre-volume Sample AnalysisabstractThis paper presents a miniature magnetic resonance imaging (MRI) platform specifically designed for imaging small-volume samples (~1μL), making it particularly suitable for real-time and on-site biochemical sample monitoring. This innovative system employs an MRI application-specific integrated circuit (ASIC) for excitation and detecting the nuclear magnetic resonance (NMR) signal. To cope with the small-volume sensing, the platform features a customized frontend probe, which includes a miniaturized saddle coil and a PDMS-molded sample well to contain the microlitre-volume sample under observation. Our proof-of-concept measurements on small-volume samples demonstrate an MRI image resolution of 150 × 150 × 250μm3. These results highlight the system’s applicability and potential for future biological analysis, offering a promising tool for researchers in the field. Shuhao Fan, Ka-Meng Lei, Rui Paulo Martins, Pui-In Mak |
ISCAS | 5 |
| 2025 | A Systematic Review of Voltage Reference Circuits: Spanning Room Temperature to Cryogenic ApplicationsabstractCryo-CMOS IC for quantum applications, proposed for tens of years, are designed to control quantum processors operating at cryogenic temperatures (CTs). The reference circuits play a significant role in quantum controllers, providing a relatively stable biasing for analog and radio frequency (RF) circuit blocks. Based on a literature review, we discovered that achieving high-accuracy reference voltage or current at CTs is challenging due to the unstable temperature characteristics of complementary metal-oxide-semiconductor (CMOS), bipolar junction transistor (BJT), or resistors in the general CMOS process at CTs. Therefore, certain specialized device structures, such as dynamic threshold MOS (DTMOS), can be employed within the bulk CMOS process. Alternatively, BJT and other devices found in specific processes, such as silicon-germanium (SiGe) and fully depleted silicon on insulator (FD-SOI) CMOS, can achieve adaptive temperature compensation. This paper provides a succinct overview of several fundamental structures and common research hot spots about the reference voltage circuits, and then assesses their suitability for CT circuit design, considering the reliability of devices in bulk CMOS, FD-SOI CMOS, and SiGe process. Finally, the paper summarizes the types of cryo-temperature reference circuits and offers an overview and comparison of them. Chen Deng, Sai Wu, Yatao Peng, Man Kay Law, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A 97.8 GOPS/W FPGA-Based Residual-Block-Aware CNN Accelerator Featuring Multi-Clock PW2 Pipeline and Adaptive-Resolution QuantizationabstractEnhancing the energy efficiency for the residual block is crucial for an energy-efficient deep neural network accelerator. This paper presents a multi-clock pointwise-pointwise (MCPW2) technique to process the adjacent PW convolution layers across residual blocks, reducing up to 75.0% DRAM access for the intermediate feature maps while securing >88.1% processing element (PE) utilization. Moreover, we introduce a dual-precision packing (DPP) DSP array to compute multiple 4/8-bit multiplications in a shared DSP, improving the accuracy by 1.5% (ImageNet) using low-precision residual distillation (RD) with adaptive-resolution quantization. The DPP DSP and adaptive-resolution RD boost the DSP efficiency up to$4.0\times $, reduce DRAM access by 50.0%, and improve the throughput by$\gt 2.7\times $. We also propose a dynamic accumulator/multiplier (A/M) DSP reconfiguration scheme to dynamically adjust the level of parallelism along the input/output channel dimensions. It also increases the PE utilization by$1.8\times $for the depthwise (DW) convolution layers with 33% less hardware resource overhead. Implemented on Xilinx VC709, the proposed accelerator achieves PE utilization of >93.0%, a DSP efficiency gain of$\gt 2.9\times $, and a throughput improvement on benchmarked networks of$4.9\times $while exhibiting an energy efficiency of 97.8 GOPs/W and a normalized throughput of 1.18 GOPS/DSP. Jixuan Li, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Analyses Concerning the Phase Noise and Nonlinear Behavior of the Charge-Sharing Integrator-Based Hybrid PLLabstractHybrid PLLs (HPLLs) leverage the advantages of conventional analog and digital PLLs to cater to the high integration demand inherent with the advanced CMOS nodes, among which the charge-sharing (CS) integrator-based HPLL featuring ultra-compact area and ultra-low power consumption exhibits promising prospects. This work presents a comprehensive analysis concerning the CS integrator-based HPLL for the first time. The behavioral modeling is first conducted with a time-domain event-driven modeling method to simulate the piecewise-linear modulations on the oscillator frequency A noise model considering the slow-fast clock domain transitions and the digital-analog signal transitions inherent with the architecture is further proposed. The proposed models offer precise PN PSD estimations over the specified frequency range under all simulated conditions, whose integrated jitters differ by less than 5 fs compared with circuit simulations. The additional nonlinearity induced by the CS integrator is also considered, and an analytical prediction approach for the nonlinearity-induced spurs is presented, showing a capability of predicting the locations and amplitudes of the most significant spurious tones with a less than 4 % deviation in the relative offset frequency and a less than 5 dBc deviation in the relative amplitude. Jingrun Song, Yueduo Liu, Zhengxuan Han, Zehao Zhang, Jiaxin Liu 0001, Hongshuai Zhang, Jun Yin 0001, Pui-In Mak, Shiheng Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2025 | Design and Analysis of a Type-II Sampling PLL With Automatic Frequency and Phase Calibrations Achieving 0.62-μs Locking TimeabstractThis paper presents a type-II sampling phase-locked loop (SPLL) that accelerates the locking process by exploiting a time-to-digital converter (TDC) based automatic frequency and phase calibration (AFPC) technique. The proposed AFPC accelerates the frequency acquisition by using a type-I loop to map the quantized phase error to the switched-capacitor (SC) control word of the voltage-controlled oscillator (VCO). The subsequent phase error after frequency locking is swiftly reduced within one TDC resolution by adjusting the division ratio of the multi-modulus divider (MMD) with little hardware expenditure. The proposed AFPC can guarantee a fast-locking time at different initial frequencies, which is insensitive to the variation of TDC resolution. This paper also contributes to a design strategy for the AFPC loop, e.g., the required frequency step of the SC and the TDC resolution, based on the analysis of the lock-in range of the SPLL. Fabricated in 28-nm CMOS with a core area of 0.15 mm$^{\mathbf {2}}$, the 6.0-to-6.9GHz SPLL prototype using a reference (REF) clock of 100 MHz achieves a locking time of$0.62~\mu $s ($62{T} _{\mathbf {REF}}$) at an 880-MHz hopping frequency. At 6.5 GHz, the SPLL consumes 4.6 mW and measures an RMS jitter and REF spur of 99 fs and –71.6 dBc, respectively, corresponding to a jitter figure-of-merit (FoM$_{\mathbf {jitter}}$) of –253.5 dB. Tailong Xu, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak, Quan Pan 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | An 840-to-970 MHz Multimodal Wake-Up Receiver With a Q-Equalized Antenna-ED Interface and 2-Dimensional Wake-Up IdentificationabstractThe article introduces an antenna-envelope detector (ED) co-designed wake-up receiver (WuRX) that can automatically detect the frequency-hopping sequence from 840 to 970 MHz. Prototyped in 65nm CMOS, the WuRX supports three modes: 1) low-power mode that achieves −68 dBm sensitivity with 9.9nW power consumption, 2) Q-enhanced mode that provides 22 dB rejection to an on-off-keying (OOK) modulated pseudo-random-bit-sequence blocker at 10 MHz offset and 41 dB rejection to a continuous-wave one at 10 MHz offset, with a power consumption of$39.6~\mu $W and 3) 2-stage wake-up mode that combines the advantages of both the low-power mode and Q-enhanced mode, with a power consumption of 33.7 nW. These characteristics are achieved by exploiting 1) an antenna-ED interface where the Q-factor of the antenna is equal to the Q-factor of the capacitor to ensure a conjugate matching condition for the maximum available power transfer from the antenna to the ED, 2) a frequency tuner connected to the antenna-ED interface calibrated by a frequency-locked loop, and 3) a Q-booster reconfigured from the frequency tuner for boosting the passive gain and narrowing the bandwidth of the interface. Zhizhan Yang, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | GSLP-CIM: A 28-nm Globally Systolic and Locally Parallel CNN/Transformer Accelerator With Scalable and Reconfigurable eDRAM Compute-in-Memory Macro for Flexible DataflowabstractThis article reports a globally systolic and locally parallel (GSLP) convolutional NN (CNN) and Transformer accelerator based on the scalable and reconfigurable (SR) embedded dynamic random-access memory (eDRAM) compute-in-memory (CIM) macro. It features: 1) a GSLP architecture employs systolic CIM macros with the reconfigurable inter-CIM network to support flexible dataflow, including weight stationary (WS), output stationary (OS), and Row stationary (RS); 2) an SR-CIM macro features reconfigurable weight/input/output memory ratio to maximize the related data reuse in different dataflow; 3) a high-density 3T eDRAM-CIM cell to further improve the density of the accelerator; 4) an area-efficient in-memory accumulator (IMA) to save the area and power overhead of the digital accumulation in each CIM macro. Prototyped in 28-nm CMOS process, the proposed GSLP-CIM accelerator exhibits a 4b peak throughput density of 0.16 TOPS/mm2 and a 4b peak compute energy efficiency of 3.55 TOPS/W. Specifically, evaluated with ResNet-50@ImageNet and ViT-B@ImageNet, this work reaches the system throughput of 24.5 and 5.66 inferences per second (IPS), the system throughput density of 19.3 IPS/mm2 and 4.46 IPS/mm2, the system compute energy efficiency of 423.9 inferences per watt (IPW) and 97.6 IPW, respectively. Wei-Han Yu, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Capacitorless Flipped-Voltage-Follower-Based Low-Dropout Regulator Incorporating Adaptive-Compensation BufferabstractThis brief presents an output-capacitorless low-dropout (OCL-LDO) regulator based on flipped-voltage-follower (FVF) and dual pMOS pass transistors. An adaptive-compensation buffer (ACB) dynamically regulates the operation of the pass transistors. Specifically, when the load current falls below 5 mA, only the smaller pass transistor is activated; otherwise, both pass transistors are engaged, thereby simultaneously mitigating the minimum load current requirement for FVF architecture and extending the load current ranging from 0 to 30 mA while maintaining stability without an external load capacitor. At 1.15-V supply voltage and 0-mA load current, the quiescent current is$6~\mu $A. The output voltage is 1.0 V with a dropout voltage of 0.15 V. Measurements show that with a load current stepping from 0 to 30 mA at an edge time of 100 ns, the output voltage undershoot is 0.2 V with a recovery time of 200 ns while achieving a load regulation of 0.23 mV/V. Our OCL-LDO is fabricated in a 180-nm CMOS with an active area of 0.031 mm2. Tan Yee Chyan, Harikrishnan Ramiah, Sharifah Wan Muhamad Hatta, Chee-Cheow Lim, Rui Paulo Martins, Pui-In Mak, Yong Chen 0005 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | CLUT-CIM: A Capacitance Lookup Table-Based Analog Compute-in-Memory Macro With Signed-Channel Training and Weight Updating for Nonuniform QuantizationabstractCompute-in-memory (CIM) is a promising approach for realizing energy-efficient deep neural network (DNN) accelerators. Previous CIM works focusing on uniform quantization (UQ) demonstrated a higher Multiply-accumulate (MAC) precision requirement to maintain DNN inferencing accuracy, resulting lower energy efficiency. The nonuniform quantization (NUQ) has proved to require lower precision than UQ, while the existing implementations are based on high precision digital lookup table (LUT) (e.g., 16-bit), leading to large energy and area overhead for multiplier. This work presents CLUT-CIM fabricated under 28-nm CMOS featuring: 1) a capacitance LUT (CLUT)-based NUQ MAC circuit with thermometer coding scheme for weight and input activation that avoids digital LUT and reduces the energy and area overhead; 2) a signed-channel training (SCT) method that reduces the switching activity of computation to improve the energy efficiency; 3) a dual-port 6T-SRAM array to enable simultaneously weight updating and CIM operations, enhancing the memory utilization and CIM throughput. Under 3-bit NUQ precision, the peak energy efficiency is 114.3 TOPS/W, and peak throughput density is 31.78 TOPS/mm2. Yuzhao Fu, Jixuan Li, Wei-Han Yu, Ka-Fai Un, Chi-Hang Chan, Yan Zhu 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | A 5.6-dB Noise Figure, 63-86-GHz Receiver Using a Wideband Noise-Cancelling Low Noise Amplifier With Phase and Amplitude CompensationabstractIn this paper, a 63–86-GHz receiver (RX) is proposed using a wideband noise-cancelling low noise amplifier (LNA) with phase and amplitude compensation circuit. Such phase and amplitude compensation circuit consists of a 4-to-1 asymmetric power-combining transformer and two amplitude adjusting amplifiers, which improves the noise-cancelling ratio of mm-wave LNA. Meanwhile, a wideband active mixer and an absorptive IF amplifier are introduced to provide a reflectionless operation for signal quality improvement. Based on aforementioned structures, a wideband RX is implemented and fabricated using a conventional 40-nm CMOS technology. The measured results show that the RX achieves 22.5-dB conversion gain with 23-GHz 3-dB bandwidth and 50.7-mW power consumption. The minimum NF is 5.6-dB with less than 1.9-dB variation in the whole operation band. Changxuan Han, Zhixian Deng, Yiyang Shu, Jun Yin 0001, Pui-In Mak, Xun Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | A Compact Sub-nW/kHz Relaxation Oscillator Using a Negative-Offset Comparator With Chopping and Piecewise Charge-Acceleration in 28-nm CMOSabstractThis work presents a compact and power-efficient kHz-range relaxation (RC) oscillator with robust performance against temperature and voltage variations. By deliberately introducing a negative-offset voltage into the comparator, an offset cancellation scheme leveraging chopping and piecewise charge-acceleration facilitates a low temperature coefficient. A low-power comparator with a tail resistor and a low oscillation amplitude improves the energy efficiency. The die area is compact by introducing leakage-based temperature compensation that eliminates bulky resistors and complex calibration. Prototyped in a 28-nm CMOS process and measured at 28.5 kHz, our oscillator occupies 0.0046 mm$^{2}$and dissipates 27.6 nW at a 0.8-V supply. The energy efficiency is 0.97 nW/kHz, and the temperature coefficient is 33.3 ppm/$^{\circ}$C over$-$40 to 85$^{\circ}$C with 1-point calibration. The corresponding FoM of 164.9 dB compares favorably with the recent arts. The start-up time is rapid, and the period settling time is within one cycle of$\sim$5.7$\mu$s. The Allan deviation is$\le $40 ppm for measurement intervals of$>$0.5 s. Yueduo Liu, Rongxin Bao, Jiahui Lin, Jun Yin 0001, Qiang Li 0021, Pui-In Mak, Shiheng Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | A 28-nm Computing-in-Memory-Based Super-Resolution Accelerator Incorporating Macro-Level Pipeline and Texture/Algebraic SparsityabstractSuper-resolution (SR) task using the convolutional neural network is a crucial task in improving image and video quality. The introduction of the residual block (RB) raises the depth of the algorithm to perform better reconstruction. The processing of the RB leads to a decrease in hardware utilization and frequent off-chip communications. It is hard to apply such algorithms on edge devices with limited performance. Computing-in-memory (CiM) is one promising method to reduce high power caused by massive data movement in multiply-accumulation computation. The algebraic sparsity (AS) is the structured sparsity (SS) optimization for imaging computing. However, it is an unsolved problem to simultaneously realize the texture sparsity (TS) of the image and the SS of the algorithm in the CiM scheme while maintaining high hardware utilization. Thus, we propose a CiM-based SR task accelerator. There are three key contributions: first, a texture-aware workflow and a dynamic grouping CiM engine can concurrently support TS coupling with AS. Second, a macro-level pipeline scheme together with two custom-sized CiM macros and a high reuse-rate Hadamard transformation circuit reaches 91% hardware utilization. Third, a novel weight update strategy is devised to reduce the performance loss induced by the weight updating. The accelerator prototype is fabricated in a 28-nm CMOS. It scores a 22.8-44.3-TOPS/W peak energy efficiency at the voltage supply of 0.54-1.1 V and the operating frequency of 50-200 MHz, indicating 1.8-6.8x higher compared to the state-of-the-art CiM processors. Hao Wu 0084, Yong Chen 0005, Yiyang Yuan, Jinshan Yue, Xiangqu Fu, Qirui Ren, Pui-In Mak, Xinghua Wang 0005, Feng Zhang 0014 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | Analysis and Design of a 21.2-to-25.5-GHz Triple-Coil Transformer-Coupled QVCOabstractThis paper reports a triple-coil transformer-coupled quadrature voltage-controlled oscillator (TC-QVCO), which inherently provides the quadrature signal without using the noisy active-coupling transistors. The determinate correlation of tank voltages is verified by utilizing the initial state to facilitate the oscillation state analysis. Thus, the TC-QVCO would operate without the oscillation mode ambiguity. Additionally, thanks to the triple-coil transformer coupling, a large source coil$L_{S}$aids in achieving in-phase coupling for phase noise (PN) improvement, and the intensified coupling factor$k_{gd}$benefits reducing the PN and the quadrature phase error simultaneously. Therefore, our TC-QVCO would alleviate the tradeoff between PN and quadrature phase accuracy via using a large$L_{S}$and$k_{gd}$. The proposed QVCO prototyped in 65-nm CMOS exhibits a superior FoM$_{\text {@10MHz}}$(180.1 to 182.2 dBc/Hz) over a 18.2% frequency tuning range (21.2 to 25.5 GHz), and the estimated quadrature phase error <0.8°. Jun Yin 0001, Pui-In Mak, Li Geng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Fully Symmetrical Obfuscated Interconnection and Weak-PUF-Assisted Challenge Obfuscation Strong PUFs Against Machine-Learning Modeling AttacksabstractIn this paper, we propose a fully symmetrical obfuscated-interconnection PUF (SOI PUF), which containsndelay stages with each stage having 4kobfuscated interconnections for resisting machine learning (ML)-based modeling attacks. All the delay stages contribute tokPUF primitives while achieving a 20× increase in the number of possible interconnections with the same hardware resources over similar prior arts. The SOI PUF mathematical model also theoretically demonstrates the large number of nonlinear matrix multiplications for resisting ML-based modeling attacks. We further exploit parallel weak PUF cells and propose the challenge-obfuscated SOI PUF (cSOI PUF), which can effectively prevent adversaries from bypassing unknown interconnections through reverse engineering (RE) attacks. The proposed SOI PUF and cSOI PUFs are evaluated by both software simulation and FPGA measurements. Without requiring a largekas in the existing PUF architectures, the simulation results demonstrate that the proposed SOI and cSOI PUFs can achieve a ~50% prediction accuracy fork≥ 3, even when facing ML attacks using 5-hidden-layer Artificial Neural Network (ANN) with 40M training CRPs. Furthermore, the proposed (64,2/4/6/8)-SOI PUF and (64,2/4/6/8)-cSOI PUF implemented using Xilinx Artix-7 FPGA can both achieve a measured reliability and uniformity of >94% and ~50%, respectively. Depending on the value ofk, the uniqueness ranges from 29.1% to 42.7% for SOI PUFs, and further improves to ~50% for cSOI PUFs. The resilience against Reliability-based modeling attacks, Probably Approximately Correct (PAC) attacks and Reverse-Engineering-based modeling attacks will also be discussed. Chongyao Xu, Litao Zhang, Pui-In Mak, Rui Paulo Martins, Man Kay Law |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | A 9.97-GHz 190.6-dBc/Hz FOM CMOS VCO Featuring Nested Common-Mode Resonator and Intrinsic Differential 2nd-Harmonic OutputabstractThis paper presents an 8-to-10GHz CMOS voltage-controlled oscillator (VCO) with common-mode (CM) resonance. It features a nested 8-shape inductor-based CM resonator with intrinsic differential$2^{\text{nd}}$harmonic extraction. The mutual coupling of the main tank and CM resonator is negligible due to the reversal magnetic field, which avoids the additional chip area occupation of the explicit CM inductor. The VCO prototyped in 65-nm CMOS scores a −136.7-dBc/Hz PN with 10-MHz offset at 9.97 GHz, consuming 4 mW of power with a standard supply voltage of 1-V. The achieved peak Figure-of-Merit (FOM) is 190.6 dBc/Hz at 10-MHz offsets. Over a 22.9% tuning range, the VCO upholds a consistent FOM of >188.5 dBc at a 10-MHz offset. The core area is 0.116 mm2, Yunbo Huang, Yong Chen 0005, Chaowei Phil Yang, Pui-In Mak, Rui Paulo Martins |
ISCAS | 4 |
| 2023 | Modeling-Attack-Resistant Strong PUF Exploiting Stagewise Obfuscated Interconnections With Improved ReliabilityabstractThis article presents an obfuscated-interconnection physical unclonable function (OIPUF) to resist modeling attacks. By introducing nonlinear operations through exploiting the random interconnections of delay stages, the proposed OIPUF can theoretically improve the physical unclonable function (PUF) security while consuming the same hardware resources as the conventional XOR arbiter PUF (XOR APUF). We further propose the metastability-detection (MD) arbiter to effectively improve the PUF reliability. Implemented on Xilinx Artix-7 field-programmable gate array, both the proposed (64,4)- and (64,8)-OIPUF demonstrate a good reliability and uniformity, with the proposed (64,8)-OIPUF showing a better uniqueness and strict avalanche criterion (SAC) performance. Measurement results also show that the proposed MD arbiter can reduce the bit error rate (BER) of the (64,4)- and (64,8)-OIPUF by$\geq 68\times $and$\geq 48\times $at up to 100 °C, respectively. Evaluated using the logistic regression (LR), artificial neural network (ANN), and covariance matrix adaptation-evolution strategy (CMA-ES) machine learning (ML) algorithms, the proposed (64,4)- and (64,8)-OIPUF can achieve a worst case prediction accuracy of 61.47% and 50.59% with up to 10M challenge–response pairs as training set, respectively, demonstrating a significant improvement over similar prior arts. Chongyao Xu, Litao Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins |
IEEE Internet Things J. | 5 |
| 2023 | A 3.78-GHz Type-I Sampling PLL With a Fully Passive KPD-Doubled Primary-Secondary S-PD Measuring 39.6-fsRMS Jitter, -260.2-dB FOM, and -70.96-dBc Reference SpurabstractThis paper reports an active-buffer-free type-I sampling phase-locked loop (S-PLL). We innovate a fully-passive sampling phase detector with passive-gain multiplication after the sampler, resulting in a stably-boosted PD gain and better linearity. Together with a transformer-based rich-harmonic shaping voltage-controlled oscillator, the proposed S-PLL at 3.78 GHz exhibits an integrated jitter of 39.6 fsRMS (1 kHz to 100 MHz), and the jitter-power figure-of-merit scores −260.2 dB. The reference (REF) spur is −70.96 dBc due to the embedded REF-feedthrough suppression technique. Yunbo Huang, Yong Chen 0005, Bo Zhao 0003, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | A 10.5 W, 93% Efficient Dual-Path Hybrid (DPH)-Based DC-DC Converter Incorporating a Continuous-Current-Input Switched-Capacitor Stage and Enhanced IL Reduction for 12 V/24 V InputsabstractThis work proposes a high-step-down switched-capacitor (SC) hybrid DC-DC converter that effectively addresses the conduction loss in the inductor and power switches. Specifically, an architecture combining a dual-path hybrid (DPH) converter at the input, and a continuous-current-input SC (CISC) stage at the output, achieves superior performance in reducing the inductor DC current ($I_{\mathrm {L,DC}}$) compared to the existing single-inductor two-flying-capacitor ($1L$-$2C_{\mathrm {F}}$) converters. Also, the converter exploits low-voltage (LV) switches to handle a substantial portion of the current, diminishing the reliance on the inductor or high-voltage (HV) switches. Consequently, this approach enhances the efficiency and on-chip power/current density. The converter, implemented in 180-nm BCD, integrates monolithic power switches, drivers, and control circuitry. It is capable of regulating an output voltage within the range of 1.2 to 3.5 V, accommodating a 12 V/24 V-input. The peak efficiency is 93% and the on-chip current density is 0.638A/mm2. The load current delivery is up to 3 A, even using a compact inductor with a DC resistance (DCR) of 200$\text{m}\Omega $. Qiaobo Ma, Xiongjie Zhang, Anyang Zhao, Huihua Li, Yang Jiang 0002, Man Kay Law, Makoto Takamiya, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2023 | Analysis and Design of a 15.2-to-18.2-GHz Inverse-Class-F VCO With a Balanced Dual-Core Topology Suppressing the Flicker Noise UpconversionabstractThis paper presents the theory and implementation of a balanced dual-core inverse-class-F (class-F$^{\mathrm{ -1}}$) voltage-controlled oscillator (VCO). The class-F$^{\mathrm{ -1}}$topology supports high-quality-factor (high-$Q$) differential switched-capacitors (SCs) for both fundamental and 2$^{\mathrm{ nd}}$-harmonic frequency tuning, which is beneficial for improving the phase noise (PN). However, the unequal parasitic capacitors from the NMOS and PMOS negative${g}$textsubscript m transistors make it impossible to minimize their flicker noise upconversions simultaneously, especially at high operating frequencies. The mechanism of this effect is analyzed qualitatively with the model of coupled oscillators and verified using the impulse-sensitivity function (ISF) approach. To address this issue, we propose a dual-core class-F$^{\mathrm{ -1}}$VCO that leverages a balanced coupling scheme to minimize the flicker noise upconversions of NMOS and PMOS transistors simultaneously and still keep the advantage of tuning the 2$^{\mathrm{ nd}}$-harmonic frequency with differential SCs offered by the class-F$^{\mathrm{ -1}}$topology. Additionally, the symmetrical circuit topology aids in improving the differential output balancing. Prototyped in a 28-nm CMOS process without ultra-thick metal, the balanced dual-core class-F$^{\mathrm{ -1}}$VCO dissipates 19.7-mW and achieves a PN of$\!-\!113.9/\!-\!135.8$-dBc/Hz at 1/10-MHz offset from an 18.23-GHz carrier. Tuned from 15.22 to 18.23-GHz, the proposed VCO exhibits superior figure-of-merits (FoMs) at 1/10-MHz offset from 185.3/187.0 to 186.2/188.1-dBc/Hz. Haoran Li 0016, Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Floating-Domain Integrated GaN Driver Techniques for DC-DC Converters: A ReviewabstractThis paper presents the design challenges and advanced circuit techniques of integrated gate drivers for non-isolated buck converters using gallium nitride (GaN) devices to achieve fast switching and high conversion efficiency. Focusing on the essential tradeoff considerations, we first explain the detailed circuit-level issues of realizing normal and safe operations regarding integration feasibility, device safety, operation reliability, and power-stage loss alleviation when driving a GaN switch. Accordingly, we review the state-of-the-art techniques for improving various aspects of the performance, including on-chip bootstrapping enhancement, over-voltage and false-switching prevention, electromagnetic interference (EMI) noise suppression, and adaptive driving optimization. We further highlight the feature advantage of distinct techniques in specific performance/function aspects, aiming to bring GaN driver design insights and providing technical references regarding the technical superiority and limitations of improving the overall converter performance. Xuchu Mu, Guangshu Zhao, Anyang Zhao, Yang Jiang 0002, Man Kay Law, Makoto Takamiya, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | A 10.8-to-37.4 Gb/s Reference-Less FD-Less Single-Loop Quarter-Rate Bang-Bang Clock and Data Recovery Employing Deliberate-Current- Mismatch Wide-Frequency-Acquisition TechniqueabstractThis paper reports a reference-less frequency- detector-less single-loop bang-bang clock and data recovery (BBCDR) circuit featuring wide frequency acquisition. We use a current-starved ring oscillator controlled by a 5-bit resistive digital-to-analog converter to maintain quarter-rate operation, supporting a capture range of 110.4%. By the virtue of a deliberate-current-mismatch charge pump pair, we form the single-sided capture scheme in the frequency detection characteristic, eliminating the power-hungry circuits in the high-speed clock and data paths. Employing a hybrid control circuit, the proposed BBCDR automates frequency acquisition and phase tracking in the overall 32 bands. Prototyped in a 65-nm CMOS, the BBCDR covers a wide data rate from 10.8 to 37.4 Gb/s, achieving an acquisition speed of 4.63 [(Gb/s)/$\mu \text{s}$] and an energy efficiency of 1.3 pJ/bit. Lin Wang 0115, Yong Chen 0005, Chaowei Phil Yang, Xiaoteng Zhao, Pui-In Mak, Franco Maloberti, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | A Fully Integrated CMOS Tri-Band Ambient RF Energy Harvesting System for IoT DevicesabstractThis article presents a fully integrated tri-band RF energy harvesting system (RFEH) in 65-nm CMOS technology. The system is designed to harvest ambient RF energies at 900 MHz, 1.9 GHz, and 2.4 GHz through a tri-band impedance matching network (IMN), cross-coupled differential-drive (CCDD) rectifier, and an output voltage monitoring circuit to limit the rectified output voltage to 3.3 V. The system achieves a power conversion efficiency (PCE) of over 30 % across all three frequency bands with a peak of 42.8 %. Furthermore, the system exhibits a peak sensitivity of -20 dBm at an output DC voltage of 1$V$output. Jack Kee Yong, Wen Xun Lian, Harikrishnan Ramiah, Kishore Kumar Pakkirisami Churchill, Gabriel Chong, Nai Shyan Lai, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | Transfer-Path-Based Hardware-Reuse Strong PUF Achieving Modeling Attack Resilience With200 Million Training CRPsabstractThis paper presents a hardware-reuse strong physical unclonable function (PUF) based on the intrinsic transfer paths (TPs) of a conventional digital multiplier to achieve a strong modeling attack resilience. With the multiplier input employed as the PUF challenge and the path delay as the entropy source, all the possible valid propagation paths from distinct input/output pairs can serve as PUF primitives. We can quantize the path delay using a time-to-digital converter (TDC), and select the suitable TDC output bits as the PUF response. We further propose a lightweight dynamic obfuscation algorithm (DOA) and a secure mutual authentication protocol to counteract modeling attacks. The proposed strong PUF using a 32×32 multiplier as implemented in the Xilinx ZYNQ-7000 SoC features a total of 2048 intrinsic PUF primitives, while achieving a response stream (RS) with an average of 1024 responses per TDC output bit per challenge. WithBit(5) andBit(6) of the TDC output selected for PUF response generation, they demonstrate a measured reliability and uniqueness of up to 98.31% and 49.34%, respectively, with their excellent randomness performance as validated by the NIST SP800-22 tests. Under machine learning (ML)-based modeling attack with artificial neural network (ANN), the measured prediction accuracy of bothBit(5) andBit(6) can still be maintained at ~50% with a total of >200 million CRPs as the training set. Chongyao Xu, Jieyun Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | A High-Performance Dual-Topology CMOS Rectifier With 19.5-dB Power Dynamic Range for RF-Based Hybrid Energy HarvestingabstractThis brief reports a dual-topology CMOS rectifier with an extended power dynamic range (PDR) for radio frequency (RF)-based hybrid energy harvesting (RF-HEH) systems. By leveraging both the cross-coupled differential drive (CCDD) and the Dickson topologies with high forward conduction and low reverse leakage, we obtain an extension of the rectifier’s PDR by adaptively disabling the CCDD counterpart and enabling the Dickson counterpart to dominate the rectifier’s performance during high-power operation. Apart from that, we formulate a rectifier-performance index (RPI), which accounts for the power conversion efficiency (PCE), the PDR, the sensitivity, and the load resistance of the rectifier to provide an adequate performance benchmark with the state-of-the-art rectifiers. Fabricated in a 130-nm CMOS, the proposed dual-topology rectifier measures a wide PDR of 19.5 dB with a peak PCE of 78.4% for a 100-$\text{k}\Omega $load operating at 900 MHz. Besides, our prototype records the highest RPI of 19.2 compared to the recent arts operating at GSM900. Alexander Choo Chia Chun, Harikrishnan Ramiah, Kishore Kumar Pakkirisami Churchill, Yong Chen 0005, Saad Mekhilef, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | A Reconfigurable CMOS Stack Rectifier With 22.8-dB Dynamic Range Achieving 47.91% Peak PCE for IoT/WSN ApplicationabstractThis brief proposes a 900-MHz novel CMOS-reconfigurable stack rectifier (RSR) implemented in a three-stage cross-coupled differential rectifier (CCDR) for battery-assist Internet-of-Things (IoT)/wireless sensor network (WSN) applications. A three-mode RSR is incorporated for an extended dynamic range (DR) input power level with a 100-$\text{k}\Omega $load, fabricated in the 130-nm CMOS. The realized RSR achieves a wide DR power conversion efficiency (PCE) by reducing the ON-resistance (${R} _{\mathrm{\scriptscriptstyle ON}}$) in the low-power zone (LPZ) achieved by reducing the threshold voltage (${V} _{\text {th}}$) of the device and alternately increasing${V} _{\text {th}}$in the high-power zone (HPZ) by implementing the proposed reconfigurable stack transistor technique along with the multithreshold voltage (MTV) technique. The circuit observes a measured result of 47.91% in peak PCE at an input power of −14 dBm by driving a 100-$\text{k}\Omega $load. The proposed circuit also achieved 22.8 and 16.3 dB of DR with a PCE over 20% and 30%, respectively. Compared to other state-of-the-art designs, our work exhibits better DR and PCE. Kishore Kumar Pakkirisami Churchill, Harikrishnan Ramiah, Alexander Choo Chia Chun, Gabriel Chong, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | A 3.6-GHz Type-II Sampling PLL With a Differential Parallel-Series Double-Edge S-PD Scoring 43.1-fsRMSJitter, -258.7-dB FOM, and -75.17-dBc Reference SpurabstractThis article presents a low-jitter and low-spur type-II sampling phase-locked loop (S-PLL). The innovative introduction of a differential parallel-series double-edge sampling phase detector (S-PD) achieves a high phase-detection gain and reduces the S-PLL in-band phase noise (PN). Incorporating a transformer-based harmonic-rich shaping voltage-controlled oscillator (VCO), the proposed S-PLL prototyped in a 65-nm CMOS, operates at 3.6 GHz and scores an integrated jitter of 43.1 fsrms integrated from 1 kHz to 100 MHz, it also exhibits a jitter-power figure-of-merit (FOM) of −258.7 dB. The measured reference (REF) spur is −80.34 dBc at$f_{\mathrm {REF}}$and −75.17 dBc at$2f_{\mathrm {REF}}$, respectively. Yunbo Huang, Yong Chen 0005, Bo Zhao 0003, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | A Security-Enhanced, Charge-Pump-Free, ISO14443-A-/ISO10373-6-Compliant RFID Tag With 16.2-μW Embedded RRAM and Reconfigurable Strong PUFabstractRadio frequency identification technology (RFID) has empowered a wide variety of automation industries, such as logistics and freight transportation. To further promote RFID tags adoption, security, power consumption, and cost have always been issues of general concern. This article presents the first synergy of the RFID tag with embedded resistive RAM (RRAM) array and RRAM-based reconfigurable strong physical unclonable function (R-SPUF). The RRAM not only meets the mass storage and technology downscaling but also renders the ultralow-cost “1-cent RFID tag” more feasible. Moreover, the R-SPUF facilitates multiple initializations until a satisfactory distribution and has strong secure keys benefiting from its reconfigurability that improves both safety and reliability. The complete system operates at 13.56 MHz and is compliant with the ISO14443-A and ISO10373-6 (test) protocols. The RFID tag was fabricated on a 1.1-mm2 die based on the 0.18-$\mu \text{m}$CMOS process. Without resorting to the charge pumps for RRAM read–write operations, the total power consumption is as low as 52.3$\mu \text{W}$, of which the RRAM dissipates$16.2~\mu \text{W}$under a wireless power supply. Qirui Ren, Qiang Huo, Hao Wu 0084, Xiangqu Fu, Xiaoxin Xu, Jianfeng Gao 0005, Xiaojin Zhao, Dengyun Lei, Xinghua Wang 0005, Feng Zhang 0014, Yong Chen 0005, Pui-In Mak |
IEEE Trans. Very Large Scale Integr. Syst. | 18 |
| 2023 | A 0.0043-mm2 0.085-μW/MHz Relaxation Oscillator Using Charge-Prestored Asymmetric Swings R-RC NetworkabstractIn this brief, a charge-prestored 21.2-MHz relaxation oscillator is proposed for ultralow-power applications. It occupies only 0.0043 mm2in 0.18-$\mu \text{m}$CMOS by resistor reusing and is reference-free. The simulated temperature coefficient (TC) of the output frequency is 15.2 ppm/° from −30 °C to 125 °C. By generating an asymmetric capacitor charging swing, our charge-prestored technique reduces significantly the power consumed by the swing-boostingRCnetwork during the charging phase. Also, the R-RCstructure further improves the energy efficiency. The total power consumption of the oscillator core is$1.806 \mu \text{W}$at 0.8 V, corresponding to an energy efficiency of$0.085 \mu \text{W}$/MHz that compares favorably with the state of the art. Shiheng Yang, Yueduo Liu, Rongxin Bao, Jiahui Lin, Zehao Zhang, Yong Chen 0005, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2022 | A 529-μW Fractional-N All-Digital PLL Using TDC Gain Auto-Calibration and an Inverse-Class-F DCO in 65-nm CMOSabstractThis paper presents an ultra-lower-power (ULP) digital-to-time-converter (DTC)-assisted fractional-N all-digital phase-locked loop (ADPLL) suitable for IoT applications. A proposed hybrid time-to-digital converter (TDC) extends the vernier-TDC input range with little power overhead in order to overcome the stability issue in the conventional architectures. The hybrid TDC also facilitates a background gain calibration to achieve a stable in-band phase noise insensitive to process, voltage, and temperature (PVT) variations. The implementation of a buffer-cascaded DTC simplifies the design complexity of the fractional-N operation. The ADPLL also features a 200$\mu \text{W}$low-phase-noise inverse-class-F (class-F−1) digitally controlled oscillator (DCO) without the need of two-dimensional (2-D) capacitor tuning for frequency alignment of the fundamental and 2nd-harmonic. Fabricated in 65-nm CMOS, the ULP ADPLL prototype achieves 868fsrmsjitter in a fractional-N channel when consuming only 529$\mu \text{W}$, corresponding to a figure-of-merit (FoM) of −244dB. Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins, Robert Bogdan Staszewski |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Mismatch Analysis of DTCs With an Improved BIST-TDC in 28-nm CMOSabstractNonlinearity of a digital-to-time converter (DTC) is pivotal to spur performance in DTC-based all-digital phase-locked-loops (ADPLL). In this paper, we characterize and analyze the mismatch of cascaded-delay-unit DTCs. Through an improved built-in-self-test (BIST) time-to-digital converter (TDC) assisted with phase-to-frequency detector (PFD), a measurement system of sub-half-ps accuracy is constructed to conduct the characterization. Fabricated in 28-nm CMOS, the DTC transfer functions are measured, and mismatches are compared against Monte-Carlo simulation results. The integral nonlinearity (INL) results are compared against each other and converted to the in-band fractional spur level when the DTC would be deployed in the ADPLL. The BIST-TDC system thus characterizes the on-chip delays without expensive equipment or complex setup. The effectiveness of adding a PFD into the$\Delta \!\Sigma $loop is validated. The entire BIST system consumes 0.6mW with a system self-calibration algorithm to tackle the analog blocks’ nonlinearities. Peng Chen 0022, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins, Robert Bogdan Staszewski |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Miniaturization of a Nuclear Magnetic Resonance System: Architecture and Design Considerations of Transceiver Integrated CircuitsabstractBeing an indispensable technique in the standard laboratory, Nuclear Magnetic Resonance (NMR) is a versatile method to non-invasively observe the atomic and molecular information of the samples containing non-zero spin nuclei. Yet, the bulky and costly hardware for NMR impedes their broad distributions outside the laboratory for on-demand and on-line usage. In recent years, with the advance in microelectronics, NMR systems equipped with customized silicon chips emerged to achieve system miniaturization with performance enhancement, and pioneer novel applications that were not feasible before using discrete NMR electronics. This article overviews the hardware for NMR from the system-level perspective, and examines the latest developments using integrated circuits over the past decade. Also, we present in detail the design considerations of the transmitter and receiver that are the cornerstone of micro-NMR systems. Shuhao Fan, Ka-Meng Lei, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A Millimeter-Wave CMOS VCO Featuring a Mode-Ambiguity-Aware Multi-Resonant-RLCM TankabstractThis paper presents a millimeter-wave NMOS-PMOS-complementary (CMOS) VCO with a multi-resonantResistor-Inductor-Capacitor-Mutual Inductance(RLCM) tank. It features an 8-port multi-tap inductor with the switched-capacitor arrays to generate and align the 1$^{{\text {st}}}$, 2$^{{\text {nd}}}$and 3$^{{\text {rd}}}$harmonic resonances; all exhibit high impedance and high intrinsic quality factor to improve the absolute phase noise (PN) at both the flicker and thermal regions. The inductor of the RLCM tank introduces a metal resistor technique to fully prevent the mode-ambiguity issue during the VCO startup. Meanwhile, we first propose a detailed analysis of the upper and lower bound of the metal resistor, which is verified by the theoretical analysis, and circuit simulation. Prototyped in 65-nm CMOS technology, the VCO scores a PN@1MHzdown to −111.41 dBc/Hz with a power consumption of 11.1 mW at 1 V; it corresponds to a FOM@1MHzup to 189.4 dBc/Hz over a 15.2% tuning range (24.62 to 28.66 GHz), while exhibiting a low 1/$\text{f}^{3}$PN corner between 480 to 730 kHz. Yong Chen 0005, Chaowei Phil Yang, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Accurate Performance Evaluation of Jitter-Power FOM for Multiplying Delay-Locked LoopabstractAn accurate performance evaluation of jitter-power figure-of-merit (FOM) for multiplying delay-locked loop (MDLL) is presented. For a typical MDLL employing a single-ended multiplying-delay ring voltage-controlled oscillator (MDVCO), it can be tuned via the capacitive loads to uphold a constant normalized phase noise (PN) across different frequencies for better jitter-power performance. Linear approximation and z-domain PN model are utilized to simplify the analysis with excellent agreement between the time-domain simulation and z-domain approximation. The influences of the asymmetric waveform, flicker noise corner, frequency error and reference noise are also discussed and then ignored based on the reasonable approximation. Under a given process, reference clock frequency and supply voltage, the ideal FOM can be derived theoretically. Based on these insights, the predicted FOM can be the benchmark metric in the design of MDLL for the possible best jitter-power performance. Yueduo Liu, Rongxin Bao, Shiheng Yang, Jun Yin 0001, Pui-In Mak, Qiang Li 0021 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | A Low-Power Multiband Blocker-Tolerant Receiver With a Steep Filtering Slope Using an N-Path LNA With Feedforward OB Blocker Cancellation and Filtering-by-Aliasing Baseband AmplifiersabstractThis paper describes a low-power multiband blocker-tolerant receiver (RX) with three steps of filtering among two transconductance (Gm) stages. To consistently achieve low noise figure (NF) and high linearity over a wide range of operating frequency, a single-Gm low-noise amplifier (LNA) heads the RX with embedded N-path filtering, and feed-forward out-of-band (OB) blocker cancellation to surmount the tradeoff between the passband bandwidth (BW) and OB rejection without sacrificing the power budget. A filtering-by-aliasing (FA) block embeds another single-Gm baseband (BB) amplifier to allow a steep roll-off lowpass response with a clock-rate-defined passband BW. Prototyped in 28-nm CMOS, the RX achieves a 3.1-dB NF and a 5.4-dBm OB-IIP3 at 2 GHz, while consuming only 22 mW, measured at a gain of 22 dB. The tunable passband BW is 2.5 to 20 MHz, and the RF-to-BB composite filtering slope is 20 to 50 dB/10 MHz at 1.75xBW offset. The RX occupies a 0.24 mm2active area. Haijun Shao, Gengzhen Qi, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A 6-to-7.5-GHz 54-fsrms Jitter Type-II Reference-Sampling PLL Featuring a Gain-Boosting Phase Detector for In-Band Phase-Noise ReductionabstractThis paper presents a type-II reference-sampling (RS) phase-locked loop (PLL) exploiting a novel gain-boosting reference-sampling phase detector (RSPD) to reduce the in-band phase noise and RMS jitter. The proposed gain-boosting RSPD converts the phase error to the voltage error and utilizes a passive switched-capacitor voltage multiplier to amplify the sampled voltage error, which effectively increases the gain of the RSPD. The boosted RSPD gain helps suppress the phase noise contributed from the gm cell. To prevent the transistors in the gain-boosting RSPD and gm cell from breakdown, the gain-boosting function is only activated during the locked state when the phase and voltage errors are small. The switching between the gain-boosting and normal modes is realized automatically by monitoring the sampled voltage and comparing it with a pre-defined threshold window. Fabricated in 65-nm CMOS, the type-II RS-PLL measures an RMS jitter of 54 fs at 6.75 GHz and consumes 7.1 mW, corresponding to a jitter figure-of-merit (FoMjitter) of −256.8 dB. The measured reference spur is −62.6 dBc, and the active area is 0.25 mm2. Tailong Xu, Shenke Zhong, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A 4T/Cell Amplifier-Chain-Based XOR PUF With Strong Machine Learning Attack ResilienceabstractThis paper presents an amplifier-chain-based XOR physical unclonable function (AC-XOR PUF), with the process- and/or bias-dependent voltage and amplification information of two identical amplifier chains serving as the entropy sources. The current-biased PUF cell using only 4 NMOS transistors achieves a small area with reduced temperature and supply sensitivity. Optimization on both the stage gain and stage number can reduce the input-referred noise (IRN) and improve the PUF reliability. We further employ an XOR gate to process the amplifier-chain outputs for the final response to improve the energy efficiency and uniqueness. The process- and bias-dependent stage amplification and the nonlinear amplifier-chain multiplication, which can significantly increase the number of modeling parameters and introduce a complex decision boundary respectively, can effectively resist machine learning (ML) modeling attacks. Fabricated in standard 65nm CMOS, the proposed AC-XOR PUF occupies an active area of$6845\mu \text{m}^{2}$. Without discarding any challenge-response pairs (CRPs), this work features a measured worst case bit error rate (BER) of 5.70% across$1.06\sim 1.55V$and$- 30\sim 125^{\circ }\text{C}$, while demonstrating a reliability (intra-die HD) and uniqueness (inter-die HD) of 0.58% and 49.92%, respectively. It also achieves a ML prediction accuracy of 50.72% using$80\times 80\times 80$artificial neural network (ANN) with 1M CPRs as training set. Jieyun Zhang, Chongyao Xu, Man Kay Law, Yang Jiang 0002, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A Reconfigurable CMOS Rectifier With 14-dB Power Dynamic Range Achieving >36-dB/mm2 FoM for RF-Based Hybrid Energy HarvestingabstractThis brief presents a novel circuit architecture for a Dickson-based reconfigurable rectifier with wide power dynamic range (PDR). Besides, a novel figure of merit (FoM) concerning the reconfigurable rectifiers is formulated to provide a more comprehensive assessment of the rectifier’s performance. The proposed reconfigurable design improves the operating range of the rectifier by adaptively switching between the six-stage configuration during low-power operation and the 12-stage configuration during high-power operation. Fabricated in 130-nm CMOS, the proposed reconfigurable rectifier measures a PDR of 14 dB with a peak power conversion efficiency (PCE) of 34.93% for 1-$\text{M}\Omega $load operating at 900 MHz. Relative to the recently published reconfigurable rectifiers, our design records the highest FoM of 36.98 dB/mm2, with minimum harvesting downtime. Alexander Choo Chia Chun, Harikrishnan Ramiah, Kishore Kumar Pakkirisami Churchill, Yong Chen 0005, Saad Mekhilef, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | A -20-dBm Sensitivity RF Energy-Harvesting Rectifier Front End Using a Transformer IMNabstractThis article describes a fully integrated CMOS radio frequency energy-harvesting (RFEH) front end. It features an on-chip stacked step-up transformer integrated with a cross-coupled differential drive (CCDD) rectifier to enhance the input sensitivity. The transformer also serves as an on-chip balun for the CCDD rectifier. The CCDD rectifier innovates a gate-biasing technique and realizes coupling capacitors at the end of each stage to increase the subsequent stage biasing. Here, our RFEH front end operating at 900 MHz achieves an improved sensitivity of −20 and −19.2 dBm at the 1-V output for no-load and a 1-$\text{M}\Omega $load, respectively. Wen Xun Lian, Harikrishnan Ramiah, Gabriel Chong, Kishore Kumar Pakkirisami Churchill, Nai Shyan Lai, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | A 3.3-GHz Integer N-Type-II Sub-Sampling PLL Using a BFSK-Suppressed Push-Pull SS-PD and a Fast-Locking FLL Achieving -82.2-dBc REF Spur and -255-dB FOMabstractThis brief describes an integer-N-type-II sub-sampling phase-locked loop (SS-PLL) incorporating a push–pull sub-sampling phase detector to significantly suppress the spur-induced binary frequency shift keying modulation (BFSK) effect and a low-power fast-locking frequency-locked loop (FLL) to shorten the settling time. Prototyped in 65-nm CMOS, the SS-PLL at 3.3 GHz shows a reference spur of −82.2 dBc, an integrated jitter of 64.9 fsrms(1 kHz to 40 MHz), and an in-band phase noise (PN) of −128.4 dBc/Hz at 1-MHz offset. The corresponding jitter power figure of merit (FOM) is −255 dB. The entire SS-PLL consumes 7.5 mW, with only$90~\mu \text{W}$associated with the FLL. Zunsong Yang, Yong Chen 0005, Jia Yuan, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | A 3.52-GHz Harmonic-Rich-Shaping VCO with Noise Suppression and Circulation, Achieving -151-dBc/Hz Phase Noise at 10-MHz OffsetabstractThis paper presents a transformer-based harmonic- rich-shaping voltage-controlled oscillator (VCO). Its active core features noise suppression and circulation to improve the phase noise (PN) performance, and its 2:2 transformer allows a very short common-mode (CM) return path without sacrificing the tank's quality factor. The proof-of-concept prototype is a 3.52GHz VCO in 65-nm CMOS. It scores a -151-dBc/Hz PN at 10MHz offset, and consumes 6.77 mW of power at a 0.7-V supply. The achieved Figure-of-Merit (FOM) is 187.8/192.3/193.7 dBc/Hz at 0.1/1/10-MHz offsets, with a 1/f3PN corner of 220 kHz. Over a 22.1% tuning range, the VCO upholds a consistent FOM of >193.3 dBc at 10-MHz offset, and a 1/f3PN corner of2. Yunbo Huang, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
ISCAS | 3 |
| 2021 | A 0.003-mm2 440fsRMS-Jitter and -64dBc-Reference-Spur Ring-VCO-Based Type-I PLL Using a Current-Reuse Sampling Phase Detector in 28-nm CMOSabstractThis paper presents a linear current-reuse sampling phase detector for a single-loop type-I phase-locked loop (PLL) to simultaneously achieve a wide loop bandwidth and low control voltage ripple, resulting in low RMS jitter and reference spur, while minimizing the chip area by avoiding an explicit loop filter. Fabricated in 28-nm CMOS, the PLL prototype measures an integrated jitter of 440 fsRMS, and a spur level of -63.9 dBc at 3.296 GHz. It draws 3.3 mW at a 0.9-V supply and scores a jitter-power figure-of-merit (FoM) of -241.9 dB. With a 103-MHz reference input, a bandwidth of ~20 MHz aids suppressing significantly the ring VCO's phase noise (PN), leading to an in-band PN of -116 dBc/Hz at 1-MHz offset. The die size is 0.003 mm2. Zunsong Yang, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A 0.14-to-0.29-pJ/bit 14-GBaud/s Trimodal (NRZ/PAM-4/PAM-8) Half-Rate Bang-Bang Clock and Data Recovery (BBCDR) Circuit in 28-nm CMOSabstractThis paper reports a half-rate bang-bang clock and data recovery (BBCDR) circuit supporting the trimodal (NRZ/PAM-4/PAM-8) operation. The observation of their crossover- points distribution at the transitions introduces the single-loop phase tracking technique. In addition, low-power techniques at both the architecture and circuit levels are employed to greatly improve the overall energy efficiency and multiply data throughput by increasing the number of levels on the magnitude. Fabricated in 28-nm CMOS, our BBCDR prototype scores a 0.29/0.17/0.14 pJ/bit efficiency at 14.4/28.8/43.2 Gb/s under NRZ/PAM-4/PAM-8 modes, respectively. The jitter is <; 0.53 ps (integrated from 100 Hz to 1 GHz) with approximately-equivalent constant loop bandwidth, and we achieve at least 1-UIpp jitter tolerance up to 10 MHz for all the three modes. Xiaoteng Zhao, Yong Chen 0005, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A Fully Integrated 10-V Pulse Driver Using Multiband Pulse-Frequency Modulation in 65-nm CMOSabstractThis brief describes a fully integrated 10-V pulse driver. It comprises a four-stage switched-capacitor voltage multiplier (SCVM) and a dedicated high-voltage output driver (HVOD) with multiband pulse-frequency modulation (MPFM) to generate efficiently 10 V regulated output pulses. Specifically, an analog/digital hybrid-controlled current-starved ring oscillator (HCRO) modulates the switching frequencies at distinct bands to regulate the high-voltage (HV) supply for the HVOD, while enabling fast output transitions with an improved driving efficiency. Prototyped in 65-nm bulk CMOS, the driver demonstrates 10-V pulse generations over a 0.1-to-1-MHz range for a 15 pF//50$\text{k}\Omega $load. With the proposed MPFM, this work measures an overall driving efficiency of up to 19.9%, corresponding to a$\sim 1.6\times $improvement over prior arts. The measured output rise time of 119 ns is also ~25% faster when compared with using the conventional pulse-frequency modulation (PFM) scheme. Jiangchao Wu, Hou-Man Leong, Yang Jiang 0002, Man Kay Law, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | A Unity-Power-Factor Inductive Power Transfer Converter with Inherent CC-to-CV Transition Ability for Automated Guided Vehicle ChargingabstractNowadays, automated guided vehicle (AGV) has been a key role of manufacturing and logistics, required to work with high automation. It is essential to guarantee the charging reliability of AGV for all-day work. However, in wireless battery charging applications, to utilize load-independent-current and load-independent-voltage transfer characteristics for constant current (CC) and constant voltage (CV) outputs respectively, existing inductive power transfer (IPT) converters either hopping the operating frequencies or switching the hybrid compensation topologies, resulting in necessity of state-of-charge detection and feedback control for the operating frequency hopping or the compensation topology switching. To improve the charging reliability of AGV by eliminating these active controls, this paper proposes an IPT converter with inherent CC-to-CV transition ability that can comply with the required charging profile of the batteries. Unity-power-factor design is also considered to minimize the voltage-ampere rating. This solution is passive, and control is unnecessary, so donating great reliability in automated guided vehicle. Experimental results are presented to verify the methodology and the performance. Io-Wa Iam, Iok-U Hoi, Chi-Seng Lam, Pui-In Mak, Rui Paulo Martins |
IECON | 5 |
| 2020 | Low Complexity Illumination-Invariant Motion Vector Detection Based on Logarithmic Edge Detection and Edge DifferenceabstractThis paper describes a low complexity illumination-invariant motion vector detection algorithm based on logarithmic edge detection and edge difference without periodic threshold adjustment. A logarithmic edge detector is employed to achieve accurate object movement over a wide illumination range. The threshold for edge detection is determined by one-time background logarithmic gradient extraction. Finally, the difference of the logarithmic edge of 3 consecutive frames is employed in a frame difference-based motion template model to obtain the motion vector. Experimental results show that the proposed algorithm achieves a motion vector detection accuracy of ~90% over an illumination level change of 80%. Chuanqi Wei, Jiangchao Wu, Man Kay Law, Pui-In Mak, Rui Paulo Martins |
ISCAS | 4 |
| 2020 | An N × N Multiplier-Based Multi-Bit Strong PUF using Path Delay ExtractionabstractThis paper presents a digital N × N multiplier-based multi-bit strong physical unclonable function (PUF), which utilize the intrinsic path delay of the multiplier to achieve an approximated 1 : 2N2average challenge-to-response extraction to effectively increase the number of PUF responses. The PUF Extractor triggers the digital multiplier, and further processes the multiplier intrinsic path delay through a time-to-digital converter (TDC). Implemented with Xilinx Artix-7 FPGAs using the automatic place and route function, the proposed strong PUF demonstrates a 64-bit challenge with 32-bit multipliers with an extra level of unpredictability for counterfeiting model-based machine learning attack. With an average of 1:2048 responses per challenge, measurement results show that the uniqueness is 53.16%, and the stability of up to 95.54%, respectively. Chongyao Xu, Jieyun Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins |
ISCAS | 5 |
| 2020 | A 6.4pJ/Bit Strong Physical Unclonable Function Based on Multiple-Stage Amplifier ChainabstractIn this paper, we present a novel multiple-stage amplifier chain based strong physical unclonable function (PUF) with low power and energy consumption. Based on the proposed two-dimensional subthreshold amplifier array, 12 different amplifiers can be selected through the analog multiplexer at each column. As a result, a 12-stage amplifier chain can be formed by applying different challenges to the aforesaid analog multiplexers with a linear feedback shift register (LFSR). Due to the inevitable process variation, the output voltage of the amplifier chain's last stage varies depending on the various combinations of the selected amplifiers, whose number features an exponential relationship with the size of the adopted amplifier array. By using 65nm standard CMOS process, the proposed strong PUF implementation is validated with high reliability and randomness. According to our extensive simulation results, the averaged bit error rate (BER) per 10°C and BER per 0.1V are calculated to be 3.15% and 3.85% for the operating temperature range of -20°C~120°C and supply voltage range of 0.9V~1.4V, respectively. Meanwhile, the proposed strong PUF's high randomness is also verified by passing both the NIST and auto-correlation function (ACF) test suites. Moreover, featuring an excellent uniqueness of 49.54%, the overall power consumption is simulated to be 0.128μW at the throughput of 0.02Mb/s, which corresponds to an energy consumption as low as 6.4pJ/bit. Jieyun Zhang, Xiaojin Zhao, Man Kay Law, Chongyao Xu, Jiahao Liu 0003, Pui-In Mak, Rui Paulo Martins |
ISCAS | 6 |
| 2020 | A 3.15-mW +16.0-dBm IIP3 22-dB CG Inductively Source Degenerated Balun-LNA Mixer With Integrated Transformer-Based Gate Inductor and IM2 Injection TechniqueabstractThis article proposes two linearization techniques in improving the third-order input intercept point (IIP3) of a balun-low-noise amplifier (LNA) mixer. First, the intrinsic third-order intermodulation (IM3) product of the inductively source degenerated (ISD) transconductor from the second-order derivative transconductance component (g"m) is reduced by tailoring toward the optimum biasing point at the moderate-inversion region. Second, the generated IM3 current by the first-order derivative transconductance (g'm) due to the interaction with the feedback component in the ISD transconductor is attenuated by second-harmonic injection via the bulk of the ISD transconductor. Furthermore, a transformer-based gate inductor and a transformer-based balun are applied to improve the input impedance matching and produce a balanced differential input signal. Measured results in 0.13-μm CMOS show a high IIP3 of +16 dBm and a conversion gain (CG) of 22 dB at 2.4 GHz. The double-sideband (DSB) noise figure (NF) is 7.2 dB, and the power consumption is 3.15 mW at 1.2 V. Nandini Vitee, Harikrishnan Ramiah, Pui-In Mak, Jun Yin 0001, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | A 1-V 4-mW Differential-Folded Mixer With Common-Gate Transconductor Using Multiple Feedback Achieving 18.4-dB Conversion Gain, +12.5-dBm IIP3, and 8.5-dB NFabstractThis article reports a novel differential-folded mixer with multiple-feedback techniques for performance enhancement. Specifically, we introduce the capacitor cross-coupled (CCC) common-gate (CG) transconductance stage to improve the noise figure (NF) at low power by boosting the effective transconductance, while enhancing the linearity via suppressing the second-order harmonic distortion. Typically, the created loop gain of the CCC can raise the third-order intermodulation (IM3) distortion, penalizing the input-referred third-order intercept point (IIP3). Here, we propose a positive and a second capacitive feedback into the CCC CG transconductor, not only to suppress the IM3 distortion current but also adds in design flexibility to the input transistors. Furthermore, the positive feedback also improves the input impedance matching, conversion gain, and NF through a flexible design criterion. Prototyped in a 0.13-μm process, the proposed mixer operating at 900 MHz dissipates 4 mW at 1 V. The measured double sideband (DSB) NF is 8.5 dB, the conversion gain (GC) is 18.4 dB and the IIP3 is +12.5 dBm. Nandini Vitee, Harikrishnan Ramiah, Pui-In Mak, Jun Yin 0001, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | A coin-battery-powered LDO-Free 2.4-GHz Bluetooth Low Energy/ZigBee receiver consuming 2 mA
Zechariah Balan, Harikrishnan Ramiah, Jagadheswaran Rajendran, Nandini Vitee, Pravinah Nair Shasidharan, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
Integr. | 7 |
| 2019 | A 13-bit 8-kS/s Δ-Σ Readout IC Using ZCB Integrators With an Embedded Resistive Sensor Achieving 1.05-pJ/Conversion Step and a 65-dB PSRRabstractThis paper reports on an energy-efficient Δ-Σ readout IC (ROIC) with a high power-supply-rejection ratio (PSRR). The static power consumption is minimized by applying a zero-crossing-based (ZCB) circuit to implement switched-capacitor (SC) integrators, while the resistive sensor is embedded inside the circuit to reuse the bias current. Oversampling Δ-Σ modulation also directly provides the digitized output, avoiding the need for a power-hungry instrumentation amplifier while preserving the linear settling behavior of the ZCB SC integrators. A dual-path bridge measurement aids in upholding PSRR of ROIC against bridge imbalance. Prototyped in 0.18-μm CMOS, the dual-path ROIC for the bridge measurement shows a nonlinearity of c400 ppm and an rms-noise-equivalent resolution of 13 bits at a conversion rate of 8 kS/s, corresponding to a figure of merit of 1.05-pJ/conversion step. The achieved noise-frequencyindependent PSRR is 65 dB, and the supply and temperature sensitivities are 0.23%/V and 55 ppm/°C, respectively. Bing Li 0011, Ji-Ping Na, Wei Wang 0165, Jia Liu 0011, Pui-In Mak |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | Analysis and Verification of Jitter in Bang-Bang Clock and Data Recovery Circuit With a Second-Order Loop FilterabstractThis paper provides an in-depth analysis of the third-order bang-bang clock and data recovery (BBCDR) circuit, which accurately predicts its operating characteristics, namely, the jitter transfer function (JTF), the jitter tolerance (JTOL), and the jitter generation (JGEN). By formulating the time-domain waveforms, we introduce a characterizing method and also derive the closed-form equations and their simplified versions under specific conditions, which are related with the second-order loop filter (LF). Our framework is consistent with the conclusions of the prior works. Also, we discuss through the time-domain behavior, the sinking area of the JTOL and other specific phenomenon appearing in the third-order BBCDR loop. We verify all above prediction by system-level simulations with the MATLAB/simulink model. Xinyi Ge, Yong Chen 0005, Xiaoteng Zhao, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Many-Objective Sizing Optimization of a Class-C/D VCO for Ultralow-Power IoT and Ultralow-Phase-Noise Cellular ApplicationsabstractIn this paper, the performance boundaries and corresponding tradeoffs of a complex dual-mode class-C/D voltage-controlled oscillator (VCO) are extended using a framework for the automatic sizing of radio frequency integrated circuit blocks, where an all-inclusive test bench formulation enhanced with an additional measurement processing system enables the optimization of “everything at once” toward its true optimal tradeoffs. VCOs embedded in the state-of-the-art multistandard transceivers must comply with extremely high performance and ultralow power requirements for modern cellular and Internet of Things applications. However, the proper analysis of the design tradeoffs is tedious and impractical, as a large amount of conflicting performance figures obtained from multiple modes, test benches, and/or analysis must be considered simultaneously. Here, the dual-mode design and optimization conducted provided 287 design solutions with figures of merit above 192 dBc/Hz, where the power consumption varies from 0.134 to 1.333 mW, the phase noise at 10 MHz from -133.89 to -142.51 dBc/Hz, and the frequency pushing from 2 to 500 MHz/V, on the worst case of the tuning range. These results pushed this circuit design to its performance limits on a 65-nm CMOS technology, reducing 49% of the power consumption of the original design while also showing its potential for ultralow power with more than 93% reduction. In addition, worst case corner criteria were also performed on the top of the worst case tuning range optimization, taking the problem to a human-untrea table LXVI-D performance space. Ricardo Martins 0003, Nuno Lourenço 0003, Nuno Horta, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | A 0.4 V 6.4 μW 3.3 MHz CMOS Bootstrapped Relaxation Oscillator with ±0.71% Frequency Deviation over -30 to 100 °C for Wearable and Sensing ApplicationsabstractWearable and sensing electronics are evolving towards energy harvesting from the environment (e.g. thermal and solar energy). Ultra-low-voltage (ULV) circuits that allow direct-powering by sub-0.5 V energy sources can maximize the power efficiency. This work is a 0.4 V 65 nm CMOS relaxation oscillator with bootstrapped logic gates and outputs. The bootstrapped logic gates enable an output swing of 1.15 V surmounting the adverse effect of ULV digital circuits without extra voltage source. The ULV comparator with bulk-driven-inputs shows an 18 dB gain with 3 cascaded stages. Also, featuring a background delay-time cancellation scheme, the 3.3 MHz relaxation oscillator with built-in calibration exhibits a frequency deviation of ±0.71% and ±0.57% against temperature (-30 to 100 °C) and voltage (0.36 to 0.44 V) variations, respectively, from Monte-Carlo simulations (N=30). The simulated power consumption is 6.4 μW, resulting in an energy efficiency of 1.9 pJ per cycle. Ka-Meng Lei, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2017 | A 0.4V 4.8μW 16MHz CMOS crystal oscillator achieving 74-fold startup-time reduction using momentary detuningabstractFor ultra-low-power radios, the long startup time of the crystal oscillator dominates their on-off latency and limits their power efficiency. This paper describes the design of a 65nm CMOS 16MHz crystal oscillator, featuring a momentary detuning scheme to accelerate the startup transient. Specifically, during the startup phase, the bias current (266μA) and loading capacitors (8.5pF each) are enlarged concurrently such that the equivalent negative resistance of the oscillation loop can be momentarily enhanced, resulting in 74-fold reduction of the startup time (37ms → 0.5ms), while consuming just 53.2nJ at 0.4V. Also, a time-based controller automatically drives the oscillator into the steady state, which entails a smaller bias current (12μA) to sustain the oscillation, and smaller loading capacitors (5pF) to recover the proper oscillating frequency. The simulated phase noise exhibits -128.2dBc/Hz at 1kHz offset, resulting in a FoM of 265.5dBc/Hz. Ka-Meng Lei, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2017 | Piecewise BJT process spread compensation exploiting base recombination currentabstractIn this paper, a piecewise bipolar junction transistor (BJT) process spread compensation scheme is presented. By exploiting the strong correlation between the BJT saturation current and the piecewise base recombination current, the process spread and proportional-to-absolute-temperature (PTAT) drift of the base-emitter voltage (Vbe) can be reduced over a wide temperature range. Fabricated in standard 0.18-μm CMOS, the chip prototype achieves a measured Vbe standard deviation (STD) of 1.1 mV (1.8 mV) from -30 to 60 °C (-30 to 120 °C) over 12 samples, corresponding to a 2.9X (1.8X) improvement when compared to the measured Vbe STD of 3.24 mV at 25 °C from 15 standalone BJT samples with constant external bias current using the same process. Dapeng Sun, Man Kay Law, Bo Wang 0012, Pui-In Mak, Rui Paulo Martins |
ISCAS | 4 |
| 2017 | A 0.45 V 147-375 nW ECG Compression Processor With Wavelet Shrinkage and Adaptive Temporal Decimation ArchitecturesabstractThis paper presents a real-time electrocardiogram (ECG) data compression processor with improved energy efficiency while maintaining high accuracy and real-time operation. Wavelet shrinkage is exploited to filter the noise and achieve sparse ECG signal representation. Adaptive temporal decimation is proposed to achieve configurable processing to adaptively reduce the data amount and computational activities for further power reduction. Modified Huffman and run-length wavelet source coding (MHRLC) is also designed to represent wavelet coefficients with optimized average code length and reduced memory requirement. Fabricated in 0.18-μm CMOS, the ECG processor is implemented with customized near-threshold digital logics for minimum energy operation. The prototype was fully validated with the MIT-BIH Arrhythmia database. With a power consumption of 147-375 nW at 0.45 V, the proposed ECG processor exhibits a wide compression ratio ranging from 2.89 to 26.91, corresponding to a percentage-RMS-distortion from 0% to 3.11%. Chio-In Ieong, Mingzhong Li, Man Kay Law, Pui-In Mak, Mang I Vai, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Time-domain I/Q-LOFT compensator using a simple envelope detector for a sub-GHz IEEE 802.11af WLAN transmitterabstractThis paper proposes a hardware-efficient time-domain scheme to digitally compensate the I/Q imbalance and LO feedthrough (LOFT) of a sub-GHz wideband transmitter for the IEEE 802.11af WLAN. A simple envelope detector is the only analog part. The parameters are updated by Least-Mean-Square and estimated efficiently in time domain by using COordinate Rotation DIgital Computer (CORDIC), saving the training time and power consumption. The measured wideband image-rejection ratio (IRR) and LO-leakage-rejection ratio (LRR) are improved from 18.9 to 41.3 dB, and 20.4 to 37.9 dB, respectively. Chak-Fong Cheang, Ka-Fai Un, Pui-In Mak, Rui Paulo Martins |
ASP-DAC | 3 |
| 2016 | Sub-µW QRS detection processor using quadratic spline wavelet transform and maxima modulus pair recognition for power-efficient wireless arrhythmia monitoringabstractThis paper describes a power-efficient processor for extracting the timing of QRS complex from digitized ECG, based on the hardware-efficient architecture of quadratic spline wavelet transform (QSWT) and maxima modulus pair recognition (MMPR). The processor succeeds in saving the wireless system's power by 6×. Chio-In Ieong, Pui-In Mak, Mang I Vai, Rui Paulo Martins |
ASP-DAC | 2 |
| 2016 | Sub-threshold VLSI logic family exploiting unbalanced pull-up/down network, logical effort and inverse-narrow-width techniquesabstractThis paper presents a complete energy optimized sub-threshold standard cell library exploiting unbalanced pull-up/down (PU/PD) network, logical effort and inverse-narrow-width (INW) techniques. Individual logic cell is optimized for ultra-low-energy applications with low-to-moderate speed requirement. Three 14-tap 8-bit FIR filters are fabricated using a 0.18-μm CMOS technology, while one of them achieved the minimum energy/tap (0.0234 pJ) and 0.365 Figure-of-Merit (FoM) at 100 kHz, 0.31 V. Mingzhong Li, Chio-In Ieong, Man Kay Law, Pui-In Mak, Mang I Vai, Sio-Hang Pun, Rui Paulo Martins |
ASP-DAC | 4 |
| 2016 | A high-Q spiral inductor with dual-layer patterned floating shield in a class-B VCO achieving a 190.5-dBc/Hz FoMabstractThis paper proposes a dual-layer patterned floating shield (DL-PFS) technique for Silicon-based on-chip spiral inductors. By optimally utilizing the two lowest metal layer strips to shield the inductor from the substrate, electromagnetic (EM) simulations show 40% improvement of the Q factor when compared with the conventional approach. Designed and simulated in 0.13-μm CMOS, the DL-PFS inductor in a class-B VCO achieves 6.6-dB lower phase noise, and 34% power savings. The VCO also exhibits 9.7-to-10.93 GHz tunability, and -123-dBc/Hz phase noise at a 3 MHz offset. The power consumption is 1.64 mW at 0.6 V, leading to a state-of-the-art FoM of 190.5 dBc/Hz. Chee-Cheow Lim, Harikrishnan Ramiah, Jun Yin 0001, Pui-In Mak, Rui Paulo Martins |
ISCAS | 4 |
| 2016 | ProtPOS: a python package for the prediction of protein preferred orientation on a surfaceabstractUNLABELLED: Atomistic molecular dynamics simulation is a promising technique to investigate the energetics and dynamics in the protein-surface adsorption process which is of high relevance to modern biotechnological applications. To increase the chance of success in simulating the adsorption process, favorable orientations of the protein at the surface must be determined. Here, we present ProtPOS which is a lightweight and easy-to-use python package that can predict low-energy protein orientations on a surface of interest. It combines a fast conformational sampling algorithm with the energy calculation of GROMACS. The advantage of ProtPOS is it allows users to select any force fields suitable for the system at hand and provide structural output readily available for further simulation studies. AVAILABILITY AND IMPLEMENTATION: ProtPOS is freely available for academic and non-profit uses at http://cbbio.cis.umac.mo/software/protpos SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Jimmy C. F. Ngai, Pui-In Mak, Shirley W. I. Siu |
Bioinform. | 2 |
| 2015 | Predicting favorable protein docking poses on a solid surface by particle swarm optimizationabstractProtein adsorption at solid surfaces has received intense focus due to its high relevance to biotechnological applications. In alternative to experimental approaches, computational methods such as molecular dynamics (MD) simulations are frequently employed to simulate the protein adsorption process and to study molecular interactions at the interfacial region. However, a successful simulation of the adsorption process depends largely on the initial adsorbed protein orientation on the surface. To avoid sampling protein trajectory which will eventually fail to adsorb, a workaround is to first determine the preferred orientations of the protein relative to the surface and use them as starting structures in MD simulations. Here, we present the first application of particle swarm optimization (PSO) to search for the low energy docking poses of a protein molecule on a solid surface. Performing rigid-body translation and rotation of the protein with energy minimization and empirical scoring function, our search algorithm successfully located the low energy orientations of the lysozyme molecule on a hydrophobic PTFE surface. Nine out of ten predicted docking poses are energetically more favorable than all poses sampled using a brute-force search. Three sets of major adsorption sites are identified for the lysozyme and they are in good agreement to results obtained by long MD simulations; novel adsorption sites are also identified from the lowest energy docking pose. Our method provides a reliable way to predict the optimal protein orientations useful for computational studies of protein-surface interactions. Jimmy C. F. Ngai, Pui-In Mak, Shirley W. I. Siu |
CEC | 2 |
| 2015 | A Highly-Scalable Analog Equalizer Using a Tunable and Current-Reusable for 10-Gb/s I/O LinksabstractA 0.0015-mm$^{2}~1.28$-mW single-branch analog equalizer is demonstrated in 65-nm CMOS for 10-Gb/s input/output links. Instead of using passive inductors that are untunable and unscalable with technologies, gain compensation here is optimized via a tunable and current-reusable active inductor (AI). This AI incorporates a positive-feedback impedance converter with only two MOSFETs and one MOS varactor. Together with the use of: 1) negative Miller capacitors to optimize the pole-zero composition and 2) tunable resistive source degeneration to adjust the low-frequency losses, the analog equalizer recovers an eye-opening rate of minimally 30% up to 10 Gb/s over a pair of 60-cm FR4 microtrip traces. The data Pk-to-Pk jitter is$2^{7}$–1,$2^{15}$–1, and$2^{31}$–1). Yong Chen 0005, Pui-In Mak, Yan Wang 0023 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Energy Optimized Subthreshold VLSI Logic Family With Unbalanced Pull-Up/Down Network and Inverse Narrow-Width TechniquesabstractUltralow-energy biomedical applications have urged the development of a subthreshold VLSI logic family in standard CMOS. This brief proposes an unbalanced pull-up/down network, together with an inverse narrow-width technique, to improve the operating speed of the individual logic cell. Effective logical efforts save both power and die area in the process of device sizing and topology optimization. Three experimental 14-tap 8-bit finite impulse response filters optimized for ultralow-voltage operation were fabricated in 0.18-μm CMOS. Measurements show that the optimized 0.45 and 0.6 V libraries achieve minimum energy operations at 100 kHz, with a figure-of-merit of 0.365 (at 0.31 V) and 0.4632 (at 0.39 V), respectively. They correspond to 35.96% and 18.74% improvements, and the overall performances are well comparable with the state of the art. Mingzhong Li, Chio-In Ieong, Man Kay Law, Pui-In Mak, Mang I Vai, Sio-Hang Pun, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Improving the Linearity and Power Efficiency of Active Switched-Capacitor Filters in a Compact Die AreaabstractThe die size of multistandard wireless transceivers in ultrascaled CMOS is dominated by the baseband low-pass filters (LPFs), which typically count on passive-RC components to define the time constant. To break this area constraint, this paper revisits the active switched-capacitor (SC) LPF for its united benefits of clock-rate-defined bandwidth, accurate cutoff frequency, and small die size due to capacitor-ratio-based sizing and no spare elements. The key challenges of active-SC LPFs are the speed- and linearity-to-power tradeoffs, which are addressed by two circuit techniques: 1) switched-current assisting (SCA) and 2) precharging (PC). The SCA accelerates the charging speed of the integration capacitor, while the PC improves the linearity when charging the load capacitor. Three prototypes (first order, biquad, and fifth-order Butterworth) fabricated in a 65-nm CMOS process validate the feasibility of the proposed SCA and PC techniques. Yaohua Zhao, Pui-In Mak, Man Kay Law, Rui Paulo Martins |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Micropower two-stage amplifier employing recycling current-buffer Miller compensationabstractProposed is a two-stage amplifier exploiting recycling current-buffer Miller compensation (CBMC). By reusing the most current-consuming devices in the 1ststage as current buffer, such an amplifier not only can preserve the merits of typical CBMC implementation in creating the beneficial left-half-plane (LHP) zero, but also can avoid the drawbacks of typical CBMC scheme from degrading the power efficiency, DC gain, dc offset and noise performances. Optimized in 0.18μm CMOS via a low-power design procedure, the amplifier achieves >90dB DC gain, 4.5MHz unity-gain frequency and 57.2° phase margin at a 100pF capacitive load. The average slew rate and 1% settling time are 2.68V/μs and 0.239μs, respectively. The amplifier draws 22μA at a 1.2V supply. Wei Wang 0177, Zushu Yan, Pui-In Mak, Man Kay Law, Rui Paulo Martins |
ISCAS | 3 |
| 2014 | Muscle and electrode motion artifacts reduction in ECG using adaptive Fourier decompositionabstractThe reduction of the muscle and electrode motion artifacts in ECG using the adaptive Fourier decomposition (AFD) is investigated. This is an extension of our previous work, in which AFD is first proposed for ECG denoising and its effectiveness in filtering out the additive Gaussian white noise is tested. This paper studies the AFD-based ECG denoising method for two types of ECG noise due to the electrode movement and the muscle contraction which are common and important in practice. In addition, some rules on the selection and adjustment of the AFD decomposition level are proposed. The tests on the MIT-BIH Arrhythmia Database indicate that this AFD-based denoising scheme performs better than the Butterworth lowpass filter, the wavelet transform and the empirical mode decomposition methods for ECG denoising with the muscle movement and electrode motion artifacts. Ze Wang 0001, Chiman Wong, Janir Nuno da Cruz, Feng Wan 0003, Pui-In Mak, Peng Un Mak, Mang I Vai |
SMC | 5 |
| 2013 | A 0.5V 10GHz 8-phase LC-VCO Combining current-reuse and back-gate-coupling techniques consuming 2mWabstractThis paper describes an ultra-low-voltage low-power 8-phase voltage-controlled oscillator (VCO) for 10GHz beam-forming satellite receivers. It is composed by four 0.5V current-reuse LC-VCO cells inter-locked by direct-back-gate coupling, featuring independent sizing of coupling strength and frequency tuning, while avoiding the risk of forward bias the substrate p-n junctions. Optimized in 65nm CMOS, the 8-phase VCO draws only 2mW. The phase noise at 1MHz offset is -114dBc/Hz to -110dBc/Hz over a 32.5% tuning range from 8.55 to 11.88GHz. These results correspond to a high-and-stable FOM within -188 to -189.5dBc/Hz. Md. Tawfiq Amin, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2013 | A 1.83 μW, 0.78 μVrms input referred noise neural recording front endabstractThis paper describes a neural recording front end for both Local Field Potential (LFP) and Spike Potential (SP) recordings, which range from 0.1 Hz ~ 200 Hz and 200 Hz ~ 10 kHz, respectively. Based on the capacitively-coupled chopper instrumentation amplifier (CCIA) topology, a ripple reduction loop (RRL) is used to suppress the chopping ripple. A DC servo loop (DSL) that utilizes pseudo-feedback to achieve a very small unity gain bandwidth with reduced capacitor size while consuming only 12 nA is proposed. The proposed CCIA is implemented in a standard 0.18 μm CMOS process. Simulation results show that with a total power consumption of 1.525 μA from a 1.2 V supply, a NEF of 2.73 (LFP) and 2.6 (SP) can be achieved. Jiangchao Wu, Man Kay Law, Pui-In Mak, Rui Paulo Martins |
ISCAS | 3 |
| 2013 | Canonical Correlation Analysis Neural Network for Steady-State Visual Evoked Potentials Based Brain-Computer Interfaces
Ka Fai Lao, Chiman Wong, Feng Wan 0003, Pui-In Mak, Peng Un Mak, Mang I Vai |
ISNN (2) | 4 |
| 2012 | A 0.02-to-6GHz SDR balun-LNA using a triple-stage inverter-based amplifierabstractThis paper describes a software-defined-radio (SDR) balun low-noise amplifier (LNA) with no explicit bias circuit, inductor or ac-coupling network. It relies on a triple-stage inverter-based amplifier with resistive feedback to maximize the bandwidth and realize single-to-differential conversion. RC degeneration applied at the last gain stage enhances both linearity and output gain-phase balancing. Optimized in 65-nm CMOS the balun-LNA covers the 0.02-to-6-GHz band with S11>;-11 dB, voltage gain of 21.2-dB and noise figure below 3.2-dB. The in-band IIP2 (IIP3) is +36 to +51 dBm (-5.6 to -5.3 dBm). The power consumption is 7.9 mW at 1.2 V. Miguel A. Martins, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2011 | Entropy penalized learning for Gaussian mixture modelsabstractIn this paper, we propose an entropy penalized approach to address the problem of learning the parameters of Gaussian mixture models (GMMs) with components of small weights. In addition, since the method is based on minimum message length (MML) criterion, it can also determine the number of components of the mixture model. The simulation results demonstrate that our method outperform several other state-of-art model selection algorithms especially for the mixtures with components of very different weights. Boyu Wang 0004, Feng Wan 0003, Peng Un Mak, Pui-In Mak, Mang I Vai |
IJCNN | 4 |
| 2011 | A Solution to harmonic frequency problem: Frequency and phase coding-based brain-computer interfaceabstractIn this paper, we propose a modified visual stimulus generation method and feature detection algorithm to design a frequency and phase coding steady-state visual evoked potential (SSVEP) based brain-computer interface (BCI). By utilizing both frequency and phase information, we solve the harmonic frequency problem in our proposed SSVEP-BCI system. The offline experimental results show that the proposed feature detection algorithm can enhance the classification rate over 10% (from 69%±12% to 82%±8%) even though only one signal electrode is used and the harmonic frequencies (6.67Hz, 13.33Hz, 8.57Hz and 17.14Hz) are employed. Chiman Wong, Boyu Wang 0004, Feng Wan 0003, Peng Un Mak, Pui-In Mak, Mang I Vai |
IJCNN | 5 |
| 2011 | A high-voltage-enabled recycling folded cascode OpAmp for nanoscale CMOS technologiesabstractThis paper describes a high-voltage-enabling circuit technique for enhancing the gain precision and linearity of OpAmp-based analog circuits. Without resorting from specialized devices, a 2xVDD-enabled recycling folded cascode (RFC) OpAmp optimized in IV GP 65-nm CMOS achieves, when compared with its 1xVDDcounterpart, 25-dB higher open-loop DC gain and 30-dB higher IM3 (in closed loop), under a similar power budget. These joint improvements save the need of a 2ndstage in the OpAmp when high precision and high linearity are the priorities. A voltage-conscious bias scheme and gate-drain-source engineering ensure that all devices are consistently operated within the reliability limits. Pui-In Mak, Zushu Yan, Rui Paulo Martins |
ISCAS | 2 |
| 2011 | A single-to-differential LNA topology with robust output gain-phase balancing against balun imbalanceabstractThis paper presents a technique to enhance the output balancing precision of a low-noise amplifier (LNA) against balun imbalance. By utilizing two capacitive-cross-coupling common-gate amplifiers in cascode, wideband output balancing, high voltage gain and low noise figure (NF) can be concurrently achieved. A 2.4 GHz LNA design example optimized in a 0.13 μm CMOS process shows that the tolerable balun's gain and phase imbalances are up to 2 dB and 10°, respectively. With just 3.6 mW of power, the NF is 2.6 dB at a voltage gain of 30 dB. Miguel A. Martins, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2010 | Source-follower-based bi-quad cell for continuous-time zero-pole type filtersabstractPresented is a novel source-follower-based (SFB) bi-quad cell suitable for realizing continuous-time zero-pole type filters. Unlike the conventional SFB bi-quad cells that can only realize complex poles, additional complex zeros can be synthesized in the proposed one, by adding two feedforward capacitors. A 4th-order Chebyshev II fully differential low-pass filter prototype was fabricated in a 0.18-µm CMOS process. The achieved bandwidth is 2.75 MHz with +5 dBm in-band IIP3 and −1 dB gain. The power consumption is 3 mV at a 2-V supply. Yong Chen 0005, Pui-In Mak |
ISCAS | 2 |
| 2010 | SC biquad filter with hybrid utilization of OpAmp and comparator-based circuitabstractThis paper proposes a differential switched-capacitor (SC) biquad filter exploiting a hybrid structure. The 1stactive core is an operational amplifier (OpAmp) whereas the 2ndis an improved comparator-based circuit (CBC). The advantages of this new structure are justified by the reductions of power and transistor sizes. Optimized in a 65-nm CMOS process, when compared with a typical dual-OpAmp design, the proposed filter saves 19% power and 18% transistor area. The filter clocked at 40 MHz achieves 61.7-dB IM2 and 62.5-dB IM3 while drawing 2.23 mA from a 1.2-V supply. This hybrid SC biquad can gain further momentum for filters that request numerous biquads in cascade to attain higher selectivity. Miguel A. Martins, Ka-Fai Un, Pui-In Mak, Rui Paulo Martins |
ISCAS | 3 |
| 2009 | A 90nm CMOS Bio-potential Signal Readout Front-end with Improved Powerline Interference RejectionabstractThis paper describes a 90 nm CMOS low-noise low-power biopotential signal readout front-end (RFE). The front-stage instrumentation amplifier (IA) features a chopper; an AC-coupler and a novel chopper notch filter for minimizing the DC-offset; transistors' flicker noise and 50 Hz powerline interference concurrently. A noise-aware transistor selection (thin- and thick-oxide) in the IA enables a flexible tradeoff between noise and input impedance performances. The 2ndstage is a spike filter clocked by a parallel use of two non-overlapping clock generators, effectively tracking and suppressing the chopper spikes. The last stage is a gain-bandwidth-controllable amplifier for boosting the gain and alleviating different bio-potential signal measurements through simple digital controls. Simulation results showed that the RFE is capable of tolerating a differential electrode offset up to plusmn50 mV, while achieving 140 dB CMRR and 51.4 nV/radicHz inputreferred noise density. The notch at 50 Hz achieves 41dB rejection. The entire RFE consumes 16.55 to 35.5 muA at 3V. Chon-Teng Ma, Pui-In Mak, Mang I Vai, Peng Un Mak, Sio-Hang Pun, Feng Wan 0003, Rui Paulo Martins |
ISCAS | 2 |
| 2009 | An Open-loop Octave-phase Local-oscillator Generator with High-precision Correlated Phases for VHF/UHF Mobile-TV TunersabstractAn octave-phase local-oscillator (LO) generator for 170-to-860-MHz mobile-TV tuners is described. It is intended to incorporate with a polyphase mixer scheme for rejecting the 3rdand 5thharmonics of the LO that is critical for wideband reception. The circuit is structured by a cascade of 7 inverter-based phase correctors to generate a set of LO signals with octave phases in an open-loop formation, resulting in 4times relaxation of the synthesizer's operating frequency when comparing with the conventional closed-loop form that requires the use of a div-by-4 frequency divider. Optimized in a 90-nm CMOS process, the achieved phase precisions are plusmn0.8deg in VHF III (170 to 245 MHz) band and UHF (470 to 860 MHz) band while drawing 2.3 to 5.1 mA from a 1-V supply. Ka-Fai Un, Pui-In Mak, Rui Paulo Martins |
ISCAS | 2 |
| 2008 | An open-source-input, ultra-wideband LNA with mixed-voltage ESD protection for full-band (170-to-1700 MHz) mobile TV tunersabstractAn ultra-wideband low-noise amplifier (LNA) covering the full mobile TV bands (170-to-1700 MHz) is presented. It features an ESD-protected open-source-input structure to interface the off-chip balun, such that a rail-to-rail input swing and an inductorless broadband input impedance matching can be achieved concurrently, while providing better linearity and inducing less noise. In the amplification core, double- current reuse and single-stage wideband thermal noise cancellation techniques are proposed. Optimized in a 90-nm CMOS process, the LNA achieves 20.6-dB voltage gain, 2.4-to-2.7 dB noise figure and + 10.8 dBm IIP3, while consuming 9.6 mW of power at 1.2 V. |S11| ≪ −10 dB is achieved up to 1.9 GHz without needing any external resonant network. Human Body Model ESD zapping tests of ± 4 kV at the RF pins cause no failure of any device. Pui-In Mak, Ka-Hou Ao Ieong, Rui Paulo Martins |
ISCAS | 1 |
| 2007 | A Highly-Linear Successive-Approximation Front-End Digitizer with Built-in Sample-and-Hold Function for Pipeline/Two-Step ADCabstractThis paper presents an improved front-end digitizer for pipeline/two-step ADC. It achieves a high linearity by replacing the front-end stage's sub-ADC from the flash type that involves synchronous operation of several comparators, to the one that uses successive approximation (SA). This shift not only frees the ADC from an extra front-end sample-and-hold circuit, but also guarantees an inherent monotonicity because of no comparator mismatch (since the SA-ADC involves just one comparator in recursive operation). Two examples of a 100-MHz 3.5-bit/stage pipeline ADC and an 11-bit 30-MHz two-step ADC, validate the feasibility of such a digitizer. Weng-leng Mok, Pui-In Mak, Seng-Pan U, Rui Paulo Martins |
ISCAS | 2 |
| 2006 | Design and test strategy underlying a low-voltage analog-baseband IC for 802.11a/b/g WLAN SiP receiversabstractThe proliferation of multiple WLANs and the continuous scaling of CMOS have created the need for low-voltage multistandard WLAN receivers. Instead of approaching a complicated SoC, a 3D-stack SiP appears as a promising alternative to meet those requirements in conjunction with the obvious goals of low power and low cost. This paper, focused on the SiP implementation of a WLAN receiver, presents the design and test strategies underlying its analog-baseband portion to accomplish: low-voltage operation; 802.11a/b/g compliance; high routability in 3D stacking; and net-response testability of the functional blocks. Pui-In Mak, Seng-Pan U, Rui Paulo Martins |
ISCAS | 1 |