Man Kay Law

dblp:78/1488 · also Man-Kay Law · DBLP profile ↗
← Back
33ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-2799-1129ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 2 first-author · 12 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SVRoM: A 9.52-mW Video Understanding Smart Vision SoC With On-Chip Sensing and Similarity-Aware SRAM/ROM CIM Macro
abstract
The rapid growth of virtual reality (VR), augmented reality (AR), extended reality (XR), and intelligent surveillance systems has driven increasing demand for video understanding tasks on edge devices. However, supporting such tasks on edge devices remains challenging due to the limited computational resources, memory constraints, power budgets, and high data transition latency. To address these challenges, this work proposes an ultralow-power smart vision System-on-Chip (SoC) with the following features: 1) a bitline segmented, parallel-charging read-only memory (ROM) compute-in-memory (CIM) macro with power gating and coupled coding; 2) heterogeneous similarity-aware hybrid SRAM/ROM cores to exploit temporal similarity with high core utilization; 3) an intracore and intercore pipeline with hierarchical dataflow for low-power data transmission; 4) a tilewise mixed-precision weight quantization and mapping scheme for weight data compression; and 5) a hierarchical multimodal system trigger mechanism utilizing an on-chip CMOS imager with near sensor caching to skip unnecessary inferences. The proposed ROM CIM macro achieves an area efficiency of 0.753–1.673 TOPS/mm2and a storage density of 4706 Kb/mm2, achieving a$3.2\!\!-\!\!14.83\times $improvement in density Figure of Merit (FoM) over the state-of-the-art (SoTA) ROM CIM designs. The proposed SoC also shows an ultralow always-on power of$0.13~\mu $W and an average active power of 9.52 mW, which demonstrates a high energy efficiency of 13.6–18.3 TOPS/W (ResNet-20 with fixed-point 8-bit activation and 10-bit weight) on video understanding tasks, showing a$1.8-3.0\times $improvement over the existing smart vision SoCs.
Haoyang Sang, Ningchao Lin, Guangshu Zhao, Man Kay Law
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Compact High-Voltage Switched-Capacitor Driver for MEMS Actuators
abstract
This paper presents a compact high-voltage switched-capacitor (SC) MEMS actuator driver. Specifically, an exponential SC boost stage provides a high step-up voltage conversion ratio (VCR), followed by a series-parallel (SP) SC stage to sequentially drive the load step by step for efficient reactive power delivery while complying with the voltage rating of passive components. With a reduced number of operating steps as well as passive components, the proposed SC driver can effectively enable a high MEMS driving efficiency with a smaller volume, lighter weight, and lower cost. The reduced number of steps also ensures a higher driving frequency, improving the maximum reactive output power. Fabricated in a standard 180nm BCD process, the chip prototype can efficiently transfer reactive power up to 683mW, while achieving a ~3× power density improvement when compared with similar prior arts.
Qishen Fang, Feiyu Li, Man Kay Law
ISCAS3
2025 Fully Integrated Dynamic Power Cell Allocation SIDO SC DC-DC Converter with a 263.8mV/ns DVS Speed and a 19.8ns Transient Recovery Time
abstract
This paper presents a fully integrated single-input-dual-output (SIDO) SC DC-DC converter targeting on fast dynamic voltage scaling (DVS) and fast transient response applications. By applying the proposed cell-level 3-step transition sequence with multi-phase interleaving, this work can enable multiple output CLOAD-free operation with fast and robust dynamic voltage scaling (DVS) operation. The proposed PFM loop with on-demand S/H and power boundary control, together with the load transient enhancement control can also ensure dynamic power cell allocation to achieve full power utilization and a fast load transient response. Fabricated in standard 65nm CMOS, the proposed converter achieves SIDO operation with step-down VCRs of 5:4/3/2/1. At VIN= 1.2V, this work achieves a measured peak power conversion efficiency (PCE) of 83%, a fast DVS speed of 263.8mV/ns, and a fast transient recovery time of 19.8ns. It also demonstrates a measured PCE improvement of up to 6.7% under full power cell utilization.
Feiyu Li, Qishen Fang, Man Kay Law
ISCAS3
2025 Ultra Low Power Video Understanding Smart Vision SoC with On-Chip Sensing and Hybrid Similarity-Aware SRAM/ROM CIM Macro
abstract
The proliferation of augmented reality (AR), virtual reality (VR), and extended reality (XR) edge devices drives the demand for video understanding, which imposes substantial demands on computational resources, memory, and energy efficiency. This work proposes an ultra-low power and highly compact smart vision SoC, composed of: 1) a bit-line (BL) segmented parallel charging ROM CIM macro with power gating and coupled coding; 2) heterogeneous similarity-aware hybrid SRAM/ROM cores for exploiting temporal similarity with high utilization; 3) a hierarchical multi-modal system trigger uses an on-chip CMOS imager with near-sensor caching for skipping unnecessary inference; and 4) a RISC-V core featuring dedicated ISA extensions to support flexible workload allocation and programmability. The proposed ROM CIM macro achieves a storage density of 4705 kb/mm2and an area efficiency of 0.753 TOPS/mm2, demonstrating a 3.2~14.83× improvement in density FoM compared to state-of-the-art (SOTA) CIM macros. Furthermore, the SoC achieves an energy efficiency of 13.6~18.3 TOPS/W on ResNet-20 (fixed-point 8-bit activation and 10-bit weight), representing a 1.8~3.0× improvement over existing smart vision SoCs.
Haoyang Sang, Ningchao Lin, Guangshu Zhao, Man Kay Law
ISCAS5
2025 A Token-Passing-Based Trigger-Prediction Methodology for Event-Driven ToF Sensors
abstract
This paper presents an ambient light robust methodology for event-driven (ED) time-of-flight (ToF) sensors, mainly targeting 3D detection under outdoor scenarios. Different from prior ED methodologies, we propose a trigger-prediction scheme with token-passing algorithm to effectively filter out ambient photons for improving the system performance under strong background light. Specifically, we determine the variation in distance by detecting the change in the width of timing windows in the trigger stage, while extracting the alteration of the number of photons in the multi-step prediction stage. This can lower the error rate while ensuring a fast imaging speed. Using a 180 nm CMOS process, we developed a behavior model including the device characteristics and environmental parameters for performance evaluation. Monte Carlo simulation results demonstrate that with a 10 m range, our proposed method can achieve a depth accuracy of less than 5 cm under 10 klux, and less than 25 cm under 80 klux, respectively, and an event response time of less than 384 µs under dynamic scenes.
Yifei Xiang, Haoyang Sang, Man Kay Law
ISCAS4
2025 A Systematic Review of Voltage Reference Circuits: Spanning Room Temperature to Cryogenic Applications
abstract
Cryo-CMOS IC for quantum applications, proposed for tens of years, are designed to control quantum processors operating at cryogenic temperatures (CTs). The reference circuits play a significant role in quantum controllers, providing a relatively stable biasing for analog and radio frequency (RF) circuit blocks. Based on a literature review, we discovered that achieving high-accuracy reference voltage or current at CTs is challenging due to the unstable temperature characteristics of complementary metal-oxide-semiconductor (CMOS), bipolar junction transistor (BJT), or resistors in the general CMOS process at CTs. Therefore, certain specialized device structures, such as dynamic threshold MOS (DTMOS), can be employed within the bulk CMOS process. Alternatively, BJT and other devices found in specific processes, such as silicon-germanium (SiGe) and fully depleted silicon on insulator (FD-SOI) CMOS, can achieve adaptive temperature compensation. This paper provides a succinct overview of several fundamental structures and common research hot spots about the reference voltage circuits, and then assesses their suitability for CT circuit design, considering the reliability of devices in bulk CMOS, FD-SOI CMOS, and SiGe process. Finally, the paper summarizes the types of cryo-temperature reference circuits and offers an overview and comparison of them.
Chen Deng, Sai Wu, Yatao Peng, Man Kay Law, Jun Yin 0001, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 Fully Integrated Single-Input-Dual-Output Switch-Capacitor DC-DC Converter With Fast DVS Speed and Fast Transient Response
abstract
This paper presents a fully integrated single-input-dual-output (SIDO) switched-capacitor (SC) DC-DC converter targeting on fast dynamic voltage scaling (DVS) and fast transient response applications. By applying the proposed cell-level 3-step transition sequence with multi-phase interleaving, this work can enable multiple output${C} _{\mathbf {LOAD}}$-free operation with fast and robust DVS operation. The proposed pulse frequency modulation (PFM) loop with power boundary control also supports on-demand sample-and-hold (S/H) operation and load transient enhancement technique to ensure full power utilization while exhibiting robust DVS transition and fast transient response. Fabricated in standard 65nm CMOS, the proposed SC converter achieves SIDO operation with flexible step-down VCRs of 5:4/3/2/1. At${V} _{\mathbf {IN}} =1.2$V, this work achieves a measured peak power conversion efficiency (PCE) of 83%, a fast DVS speed of 263.8mV/ns, and a fast transient recovery time of 19.8ns.
Feiyu Li, Qishen Fang, Man Kay Law
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Fully Symmetrical Obfuscated Interconnection and Weak-PUF-Assisted Challenge Obfuscation Strong PUFs Against Machine-Learning Modeling Attacks
abstract
In this paper, we propose a fully symmetrical obfuscated-interconnection PUF (SOI PUF), which containsndelay stages with each stage having 4kobfuscated interconnections for resisting machine learning (ML)-based modeling attacks. All the delay stages contribute tokPUF primitives while achieving a 20× increase in the number of possible interconnections with the same hardware resources over similar prior arts. The SOI PUF mathematical model also theoretically demonstrates the large number of nonlinear matrix multiplications for resisting ML-based modeling attacks. We further exploit parallel weak PUF cells and propose the challenge-obfuscated SOI PUF (cSOI PUF), which can effectively prevent adversaries from bypassing unknown interconnections through reverse engineering (RE) attacks. The proposed SOI PUF and cSOI PUFs are evaluated by both software simulation and FPGA measurements. Without requiring a largekas in the existing PUF architectures, the simulation results demonstrate that the proposed SOI and cSOI PUFs can achieve a ~50% prediction accuracy fork≥ 3, even when facing ML attacks using 5-hidden-layer Artificial Neural Network (ANN) with 40M training CRPs. Furthermore, the proposed (64,2/4/6/8)-SOI PUF and (64,2/4/6/8)-cSOI PUF implemented using Xilinx Artix-7 FPGA can both achieve a measured reliability and uniformity of >94% and ~50%, respectively. Depending on the value ofk, the uniqueness ranges from 29.1% to 42.7% for SOI PUFs, and further improves to ~50% for cSOI PUFs. The resilience against Reliability-based modeling attacks, Probably Approximately Correct (PAC) attacks and Reverse-Engineering-based modeling attacks will also be discussed.
Chongyao Xu, Litao Zhang, Pui-In Mak, Rui Paulo Martins, Man Kay Law
IEEE Trans. Inf. Forensics Secur.5
2023 Modeling-Attack-Resistant Strong PUF Exploiting Stagewise Obfuscated Interconnections With Improved Reliability
abstract
This article presents an obfuscated-interconnection physical unclonable function (OIPUF) to resist modeling attacks. By introducing nonlinear operations through exploiting the random interconnections of delay stages, the proposed OIPUF can theoretically improve the physical unclonable function (PUF) security while consuming the same hardware resources as the conventional XOR arbiter PUF (XOR APUF). We further propose the metastability-detection (MD) arbiter to effectively improve the PUF reliability. Implemented on Xilinx Artix-7 field-programmable gate array, both the proposed (64,4)- and (64,8)-OIPUF demonstrate a good reliability and uniformity, with the proposed (64,8)-OIPUF showing a better uniqueness and strict avalanche criterion (SAC) performance. Measurement results also show that the proposed MD arbiter can reduce the bit error rate (BER) of the (64,4)- and (64,8)-OIPUF by$\geq 68\times $and$\geq 48\times $at up to 100 °C, respectively. Evaluated using the logistic regression (LR), artificial neural network (ANN), and covariance matrix adaptation-evolution strategy (CMA-ES) machine learning (ML) algorithms, the proposed (64,4)- and (64,8)-OIPUF can achieve a worst case prediction accuracy of 61.47% and 50.59% with up to 10M challenge–response pairs as training set, respectively, demonstrating a significant improvement over similar prior arts.
Chongyao Xu, Litao Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins
IEEE Internet Things J.3
2023 A 10.5 W, 93% Efficient Dual-Path Hybrid (DPH)-Based DC-DC Converter Incorporating a Continuous-Current-Input Switched-Capacitor Stage and Enhanced IL Reduction for 12 V/24 V Inputs
abstract
This work proposes a high-step-down switched-capacitor (SC) hybrid DC-DC converter that effectively addresses the conduction loss in the inductor and power switches. Specifically, an architecture combining a dual-path hybrid (DPH) converter at the input, and a continuous-current-input SC (CISC) stage at the output, achieves superior performance in reducing the inductor DC current ($I_{\mathrm {L,DC}}$) compared to the existing single-inductor two-flying-capacitor ($1L$-$2C_{\mathrm {F}}$) converters. Also, the converter exploits low-voltage (LV) switches to handle a substantial portion of the current, diminishing the reliance on the inductor or high-voltage (HV) switches. Consequently, this approach enhances the efficiency and on-chip power/current density. The converter, implemented in 180-nm BCD, integrates monolithic power switches, drivers, and control circuitry. It is capable of regulating an output voltage within the range of 1.2 to 3.5 V, accommodating a 12 V/24 V-input. The peak efficiency is 93% and the on-chip current density is 0.638A/mm2. The load current delivery is up to 3 A, even using a compact inductor with a DC resistance (DCR) of 200$\text{m}\Omega $.
Qiaobo Ma, Xiongjie Zhang, Anyang Zhao, Huihua Li, Yang Jiang 0002, Man Kay Law, Makoto Takamiya, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 Floating-Domain Integrated GaN Driver Techniques for DC-DC Converters: A Review
abstract
This paper presents the design challenges and advanced circuit techniques of integrated gate drivers for non-isolated buck converters using gallium nitride (GaN) devices to achieve fast switching and high conversion efficiency. Focusing on the essential tradeoff considerations, we first explain the detailed circuit-level issues of realizing normal and safe operations regarding integration feasibility, device safety, operation reliability, and power-stage loss alleviation when driving a GaN switch. Accordingly, we review the state-of-the-art techniques for improving various aspects of the performance, including on-chip bootstrapping enhancement, over-voltage and false-switching prevention, electromagnetic interference (EMI) noise suppression, and adaptive driving optimization. We further highlight the feature advantage of distinct techniques in specific performance/function aspects, aiming to bring GaN driver design insights and providing technical references regarding the technical superiority and limitations of improving the overall converter performance.
Xuchu Mu, Guangshu Zhao, Anyang Zhao, Yang Jiang 0002, Man Kay Law, Makoto Takamiya, Pui-In Mak, Rui Paulo Martins
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Transfer-Path-Based Hardware-Reuse Strong PUF Achieving Modeling Attack Resilience With200 Million Training CRPs
abstract
This paper presents a hardware-reuse strong physical unclonable function (PUF) based on the intrinsic transfer paths (TPs) of a conventional digital multiplier to achieve a strong modeling attack resilience. With the multiplier input employed as the PUF challenge and the path delay as the entropy source, all the possible valid propagation paths from distinct input/output pairs can serve as PUF primitives. We can quantize the path delay using a time-to-digital converter (TDC), and select the suitable TDC output bits as the PUF response. We further propose a lightweight dynamic obfuscation algorithm (DOA) and a secure mutual authentication protocol to counteract modeling attacks. The proposed strong PUF using a 32×32 multiplier as implemented in the Xilinx ZYNQ-7000 SoC features a total of 2048 intrinsic PUF primitives, while achieving a response stream (RS) with an average of 1024 responses per TDC output bit per challenge. WithBit(5) andBit(6) of the TDC output selected for PUF response generation, they demonstrate a measured reliability and uniqueness of up to 98.31% and 49.34%, respectively, with their excellent randomness performance as validated by the NIST SP800-22 tests. Under machine learning (ML)-based modeling attack with artificial neural network (ANN), the measured prediction accuracy of bothBit(5) andBit(6) can still be maintained at ~50% with a total of >200 million CRPs as the training set.
Chongyao Xu, Jieyun Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins
IEEE Trans. Inf. Forensics Secur.3
2022 Switched-Capacitor Bandgap Voltage Reference for IoT Applications
abstract
This paper presents a switched-capacitor network (SCN) based bandgap voltage reference (BGR) for IoT applications. The proposed BGR employs a dual proportional-to-absolute-temperature (PTAT) clock topology to achieve both high precision and low power over a wide temperature range while relaxing the capacitor size requirements. Specifically, the fast PTAT clock assists in reducing the output voltage ripple of the dual-branch interleaved$2\times $charge pump (CP). Meanwhile, the slow PTAT clock controls the voltage divider SCN to relax the settling error at low temperature and the leakage-induced error at high temperature simultaneously, resulting in lower power consumption and smaller temperature coefficient (TC). We also propose a replica$V_{\mathrm {EB}}$generation branch in the series-parallel SCN to improve the BGR output TC due to the finite settling time during switching. Fabricated in 65nm standard CMOS, measurement results show that the proposed BGR obtains a TC of 32 ppm/°C at 0.5V supply within −40 °C to 120 °C. The line regulation is 3.3mV/V or 0.66%/V from 0.5V to 1V. Based on 10-chip measurement results, we obtain a$3\sigma / \mu $variation of 1.37% before trimming, while 0.25% after applying one-point trimming at 20°C.
Chi-Wa U, Man Kay Law, Chi-Seng Lam, Rui Paulo Martins
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 A 4T/Cell Amplifier-Chain-Based XOR PUF With Strong Machine Learning Attack Resilience
abstract
This paper presents an amplifier-chain-based XOR physical unclonable function (AC-XOR PUF), with the process- and/or bias-dependent voltage and amplification information of two identical amplifier chains serving as the entropy sources. The current-biased PUF cell using only 4 NMOS transistors achieves a small area with reduced temperature and supply sensitivity. Optimization on both the stage gain and stage number can reduce the input-referred noise (IRN) and improve the PUF reliability. We further employ an XOR gate to process the amplifier-chain outputs for the final response to improve the energy efficiency and uniqueness. The process- and bias-dependent stage amplification and the nonlinear amplifier-chain multiplication, which can significantly increase the number of modeling parameters and introduce a complex decision boundary respectively, can effectively resist machine learning (ML) modeling attacks. Fabricated in standard 65nm CMOS, the proposed AC-XOR PUF occupies an active area of$6845\mu \text{m}^{2}$. Without discarding any challenge-response pairs (CRPs), this work features a measured worst case bit error rate (BER) of 5.70% across$1.06\sim 1.55V$and$- 30\sim 125^{\circ }\text{C}$, while demonstrating a reliability (intra-die HD) and uniqueness (inter-die HD) of 0.58% and 49.92%, respectively. It also achieves a ML prediction accuracy of 50.72% using$80\times 80\times 80$artificial neural network (ANN) with 1M CPRs as training set.
Jieyun Zhang, Chongyao Xu, Man Kay Law, Yang Jiang 0002, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 A Fully Integrated 10-V Pulse Driver Using Multiband Pulse-Frequency Modulation in 65-nm CMOS
abstract
This brief describes a fully integrated 10-V pulse driver. It comprises a four-stage switched-capacitor voltage multiplier (SCVM) and a dedicated high-voltage output driver (HVOD) with multiband pulse-frequency modulation (MPFM) to generate efficiently 10 V regulated output pulses. Specifically, an analog/digital hybrid-controlled current-starved ring oscillator (HCRO) modulates the switching frequencies at distinct bands to regulate the high-voltage (HV) supply for the HVOD, while enabling fast output transitions with an improved driving efficiency. Prototyped in 65-nm bulk CMOS, the driver demonstrates 10-V pulse generations over a 0.1-to-1-MHz range for a 15 pF//50$\text{k}\Omega $load. With the proposed MPFM, this work measures an overall driving efficiency of up to 19.9%, corresponding to a$\sim 1.6\times $improvement over prior arts. The measured output rise time of 119 ns is also ~25% faster when compared with using the conventional pulse-frequency modulation (PFM) scheme.
Jiangchao Wu, Hou-Man Leong, Yang Jiang 0002, Man Kay Law, Pui-In Mak, Rui Paulo Martins
IEEE Trans. Very Large Scale Integr. Syst.4
2020 On Fully Differential Incremental ΔΣ ADC with Initial Feedback Zeroing and 1.5-Bit Feedback
abstract
This paper presents the time-domain analysis of a fully differential incremental ΔΣ modulator. Particularly, the influence of the bipolar feedback signal on the quantization noise of the modulator is analyzed, which is overlooked in most IDC designs. Based on the analysis, an initial feedback zeroing scheme is introduced to decrease the quantization noise of the modulator. Moreover, the maximum number of output codeword that can be produced by the modulator is mathematically derived. Following the derivation, a control scheme is proposed to achieve 1.5-bit effective feedback without changing the quantizer and D/A topology. By applying the initial feedback zeroing and 1.5-bit feedback technique, quantization noise of the 1st- and 2nd-order modulators analyzed in this paper can be decreased by 4×, with very minor modifications on the modulator's original digital controllers.
Bo Wang 0012, Man Kay Law, Amine Bermak
ISCAS2
2020 Low Complexity Illumination-Invariant Motion Vector Detection Based on Logarithmic Edge Detection and Edge Difference
abstract
This paper describes a low complexity illumination-invariant motion vector detection algorithm based on logarithmic edge detection and edge difference without periodic threshold adjustment. A logarithmic edge detector is employed to achieve accurate object movement over a wide illumination range. The threshold for edge detection is determined by one-time background logarithmic gradient extraction. Finally, the difference of the logarithmic edge of 3 consecutive frames is employed in a frame difference-based motion template model to obtain the motion vector. Experimental results show that the proposed algorithm achieves a motion vector detection accuracy of ~90% over an illumination level change of 80%.
Chuanqi Wei, Jiangchao Wu, Man Kay Law, Pui-In Mak, Rui Paulo Martins
ISCAS3
2020 An N × N Multiplier-Based Multi-Bit Strong PUF using Path Delay Extraction
abstract
This paper presents a digital N × N multiplier-based multi-bit strong physical unclonable function (PUF), which utilize the intrinsic path delay of the multiplier to achieve an approximated 1 : 2N2average challenge-to-response extraction to effectively increase the number of PUF responses. The PUF Extractor triggers the digital multiplier, and further processes the multiplier intrinsic path delay through a time-to-digital converter (TDC). Implemented with Xilinx Artix-7 FPGAs using the automatic place and route function, the proposed strong PUF demonstrates a 64-bit challenge with 32-bit multipliers with an extra level of unpredictability for counterfeiting model-based machine learning attack. With an average of 1:2048 responses per challenge, measurement results show that the uniqueness is 53.16%, and the stability of up to 95.54%, respectively.
Chongyao Xu, Jieyun Zhang, Man Kay Law, Xiaojin Zhao, Pui-In Mak, Rui Paulo Martins
ISCAS3
2020 A 6.4pJ/Bit Strong Physical Unclonable Function Based on Multiple-Stage Amplifier Chain
abstract
In this paper, we present a novel multiple-stage amplifier chain based strong physical unclonable function (PUF) with low power and energy consumption. Based on the proposed two-dimensional subthreshold amplifier array, 12 different amplifiers can be selected through the analog multiplexer at each column. As a result, a 12-stage amplifier chain can be formed by applying different challenges to the aforesaid analog multiplexers with a linear feedback shift register (LFSR). Due to the inevitable process variation, the output voltage of the amplifier chain's last stage varies depending on the various combinations of the selected amplifiers, whose number features an exponential relationship with the size of the adopted amplifier array. By using 65nm standard CMOS process, the proposed strong PUF implementation is validated with high reliability and randomness. According to our extensive simulation results, the averaged bit error rate (BER) per 10°C and BER per 0.1V are calculated to be 3.15% and 3.85% for the operating temperature range of -20°C~120°C and supply voltage range of 0.9V~1.4V, respectively. Meanwhile, the proposed strong PUF's high randomness is also verified by passing both the NIST and auto-correlation function (ACF) test suites. Moreover, featuring an excellent uniqueness of 49.54%, the overall power consumption is simulated to be 0.128μW at the throughput of 0.02Mb/s, which corresponds to an energy consumption as low as 6.4pJ/bit.
Jieyun Zhang, Xiaojin Zhao, Man Kay Law, Chongyao Xu, Jiahao Liu 0003, Pui-In Mak, Rui Paulo Martins
ISCAS3
2017 CCM operation analysis and parameters design of Negative Output Elementary Luo Converter for ripple suppression
abstract
This paper presents the DC analysis of the Negative Output Elementary Luo Converter (NOELC), which includes the continuous-conduction mode (CCM) voltage gain and the boundary condition between the CCM and the discontinuous-conduction mode (DCM). The main features of the NOELC are the high gain with small ripple and the reverse output. Additionally, we address the parameters design of the NOELC and propose an output voltage ripple estimation method. Through MATLAB Simulink simulation, we further demonstrate that the parameters design and the output voltage estimation method can achieve more accurate results when compared with those from the conventional one.
Chi-Wa U, Chi-Seng Lam, Man Kay Law, Sai-Weng Sin, Man-Chung Wong, Seng-Pan U, Rui Paulo Martins
IECON3
2017 Piecewise BJT process spread compensation exploiting base recombination current
abstract
In this paper, a piecewise bipolar junction transistor (BJT) process spread compensation scheme is presented. By exploiting the strong correlation between the BJT saturation current and the piecewise base recombination current, the process spread and proportional-to-absolute-temperature (PTAT) drift of the base-emitter voltage (Vbe) can be reduced over a wide temperature range. Fabricated in standard 0.18-μm CMOS, the chip prototype achieves a measured Vbe standard deviation (STD) of 1.1 mV (1.8 mV) from -30 to 60 °C (-30 to 120 °C) over 12 samples, corresponding to a 2.9X (1.8X) improvement when compared to the measured Vbe STD of 3.24 mV at 25 °C from 15 standalone BJT samples with constant external bias current using the same process.
Dapeng Sun, Man Kay Law, Bo Wang 0012, Pui-In Mak, Rui Paulo Martins
ISCAS2
2017 A 0.45 V 147-375 nW ECG Compression Processor With Wavelet Shrinkage and Adaptive Temporal Decimation Architectures
abstract
This paper presents a real-time electrocardiogram (ECG) data compression processor with improved energy efficiency while maintaining high accuracy and real-time operation. Wavelet shrinkage is exploited to filter the noise and achieve sparse ECG signal representation. Adaptive temporal decimation is proposed to achieve configurable processing to adaptively reduce the data amount and computational activities for further power reduction. Modified Huffman and run-length wavelet source coding (MHRLC) is also designed to represent wavelet coefficients with optimized average code length and reduced memory requirement. Fabricated in 0.18-μm CMOS, the ECG processor is implemented with customized near-threshold digital logics for minimum energy operation. The prototype was fully validated with the MIT-BIH Arrhythmia database. With a power consumption of 147-375 nW at 0.45 V, the proposed ECG processor exhibits a wide compression ratio ranging from 2.89 to 26.91, corresponding to a percentage-RMS-distortion from 0% to 3.11%.
Chio-In Ieong, Mingzhong Li, Man Kay Law, Pui-In Mak, Mang I Vai, Rui Paulo Martins
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Sub-threshold VLSI logic family exploiting unbalanced pull-up/down network, logical effort and inverse-narrow-width techniques
abstract
This paper presents a complete energy optimized sub-threshold standard cell library exploiting unbalanced pull-up/down (PU/PD) network, logical effort and inverse-narrow-width (INW) techniques. Individual logic cell is optimized for ultra-low-energy applications with low-to-moderate speed requirement. Three 14-tap 8-bit FIR filters are fabricated using a 0.18-μm CMOS technology, while one of them achieved the minimum energy/tap (0.0234 pJ) and 0.365 Figure-of-Merit (FoM) at 100 kHz, 0.31 V.
Mingzhong Li, Chio-In Ieong, Man Kay Law, Pui-In Mak, Mang I Vai, Sio-Hang Pun, Rui Paulo Martins
ASP-DAC3
2016 A 2.2µW 15b incremental delta-sigma ADC with output-driven input segmentation
abstract
A micro-power incremental delta-sigma (I-ΣS) ADC is presented. This ADC uses its decimation filter's output to estimate the input signal level and dynamically adjusts the modulator feedback voltage, thereby reducing the integrator input range and power. For further power saving, integrator time-multiplexing is also employed. Fabricated in 0.18μm CMOS, the 0.12mm2 ADC consumes 2.16μW at a conversion speed of 85S/s, 15.3b resolution and -2/1.5LSB INL.
Bo Wang 0012, Man Kay Law, Saqib Mohamad, Amine Bermak
ASP-DAC2
2015 Energy Optimized Subthreshold VLSI Logic Family With Unbalanced Pull-Up/Down Network and Inverse Narrow-Width Techniques
abstract
Ultralow-energy biomedical applications have urged the development of a subthreshold VLSI logic family in standard CMOS. This brief proposes an unbalanced pull-up/down network, together with an inverse narrow-width technique, to improve the operating speed of the individual logic cell. Effective logical efforts save both power and die area in the process of device sizing and topology optimization. Three experimental 14-tap 8-bit finite impulse response filters optimized for ultralow-voltage operation were fabricated in 0.18-μm CMOS. Measurements show that the optimized 0.45 and 0.6 V libraries achieve minimum energy operations at 100 kHz, with a figure-of-merit of 0.365 (at 0.31 V) and 0.4632 (at 0.39 V), respectively. They correspond to 35.96% and 18.74% improvements, and the overall performances are well comparable with the state of the art.
Mingzhong Li, Chio-In Ieong, Man Kay Law, Pui-In Mak, Mang I Vai, Sio-Hang Pun, Rui Paulo Martins
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Improving the Linearity and Power Efficiency of Active Switched-Capacitor Filters in a Compact Die Area
abstract
The die size of multistandard wireless transceivers in ultrascaled CMOS is dominated by the baseband low-pass filters (LPFs), which typically count on passive-RC components to define the time constant. To break this area constraint, this paper revisits the active switched-capacitor (SC) LPF for its united benefits of clock-rate-defined bandwidth, accurate cutoff frequency, and small die size due to capacitor-ratio-based sizing and no spare elements. The key challenges of active-SC LPFs are the speed- and linearity-to-power tradeoffs, which are addressed by two circuit techniques: 1) switched-current assisting (SCA) and 2) precharging (PC). The SCA accelerates the charging speed of the integration capacitor, while the PC improves the linearity when charging the load capacitor. Three prototypes (first order, biquad, and fifth-order Butterworth) fabricated in a 65-nm CMOS process validate the feasibility of the proposed SCA and PC techniques.
Yaohua Zhao, Pui-In Mak, Man Kay Law, Rui Paulo Martins
IEEE Trans. Very Large Scale Integr. Syst.3
2014 A high voltage zero-static current voltage scaling ADC interface circuit for micro-stimulator
abstract
This paper describes a SAR ADC interface circuit, where the input sensing voltage from a bipolar high voltage domain is linearly translated into the low voltage domain where the SAR ADC operates. The proposed interface circuit employs the principle of charge transfer amplifier to deliver information between two different power domains, in one step, without static power consumption, even if both domains do not share the the same ground voltages. To implement the charge transfer scheme using standard asymmetric LDMOS, we propose a novel dynamic body biased high voltage transmission gate. Prototype simulation using a standard 24V BCDMOS process shows that the proposed circuit draws 2.12μW of power when it senses a voltage that swings between +10V and -10V at 1 kSPS sampling frequency.
Paul Jung-Ho Lee, Denis Guangyin Chen, Amine Bermak, Man Kay Law
ISCAS4
2014 Micropower two-stage amplifier employing recycling current-buffer Miller compensation
abstract
Proposed is a two-stage amplifier exploiting recycling current-buffer Miller compensation (CBMC). By reusing the most current-consuming devices in the 1ststage as current buffer, such an amplifier not only can preserve the merits of typical CBMC implementation in creating the beneficial left-half-plane (LHP) zero, but also can avoid the drawbacks of typical CBMC scheme from degrading the power efficiency, DC gain, dc offset and noise performances. Optimized in 0.18μm CMOS via a low-power design procedure, the amplifier achieves >90dB DC gain, 4.5MHz unity-gain frequency and 57.2° phase margin at a 100pF capacitive load. The average slew rate and 1% settling time are 2.68V/μs and 0.239μs, respectively. The amplifier draws 22μA at a 1.2V supply.
Wei Wang 0177, Zushu Yan, Pui-In Mak, Man Kay Law, Rui Paulo Martins
ISCAS4
2013 A 1.83 μW, 0.78 μVrms input referred noise neural recording front end
abstract
This paper describes a neural recording front end for both Local Field Potential (LFP) and Spike Potential (SP) recordings, which range from 0.1 Hz ~ 200 Hz and 200 Hz ~ 10 kHz, respectively. Based on the capacitively-coupled chopper instrumentation amplifier (CCIA) topology, a ripple reduction loop (RRL) is used to suppress the chopping ripple. A DC servo loop (DSL) that utilizes pseudo-feedback to achieve a very small unity gain bandwidth with reduced capacitor size while consuming only 12 nA is proposed. The proposed CCIA is implemented in a standard 0.18 μm CMOS process. Simulation results show that with a total power consumption of 1.525 μA from a 1.2 V supply, a NEF of 2.73 (LFP) and 2.6 (SP) can be achieved.
Jiangchao Wu, Man Kay Law, Pui-In Mak, Rui Paulo Martins
ISCAS2
2012 A sub-1V BJT-based CMOS temperature sensor from -55 °C to 125 °C
abstract
In this paper, a smart temperature sensor working at a supply voltage as low as 0.9V over the full military temperature range is presented. Low voltage operation is achieved by biasing the front-end BJT pairs with different emitter currents for two different sensing ranges, from -55°C to 30°C and from 20°C to 125°C, respectively. A second-order inverter-based ΣΔADC with dynamic element matching (DEM) and input signal chopping to control the conversion error to within 0:2°C is used for digital readout. Front-end bias currents are selected during the design stage to minimize the induced sensing error. The proposed sensor is implemented using the TSMC 0.18μm 1P6M process. Simulation result shows that a +1°C=-0:1°C sensing error using one-point calibration can be achieved from -55°C to 125°C. At a sampling speed of 20 samples/s, the sensor consumes 3.4μA and 4.7μA in the low temperature range and the high temperature range, respectively.
Bo Wang 0012, Man Kay Law, Fang Tang, Amine Bermak
ISCAS2
2011 A Novel Asynchronous Pixel for an Energy Harvesting CMOS Image Sensor
abstract
This paper proposes a novel energy harvesting technique based on an asynchronous pixel structure and an efficient energy generation scheme, referred to as avalanche energy generation (AEG). The key idea behind using an asynchronous type of pixel is to lower the power consumption by enabling only active pixels to be read-out after which they enter into a power generation mode. In this mode, the on-pixel photodetector itself will be used to harvest the light energy from the environment and make it available to active pixels. A very interesting feature about our proposed approach is that during a frame capture, critical energy is mainly required for starting-up activity. Once a group of pixels have been read-out, the available energy will rise and more array activity will contribute to the generation of more energy, hence creating an avalanche effect. In contrast to other early designs of energy harvesting image sensors, our scheme uses the photodetector itself for power generation. This results in better utilization of the photosensitive area and more importantly an improved energy generation scheme. Detailed power analysis and extensive simulation results are provided in this paper, which validate the proposed concept. Three test structures have been fabricated in AMIS 1-poly, 5-metal CMOS 0.35-m n-well process. The power generation process and event generation have been successfully verified experimentally.
Man Kay Law, Amine Bermak
IEEE Trans. Very Large Scale Integr. Syst.2
2008 A Time Domain differential CMOS Temperature Sensor with Reduced Supply Sensitivity
abstract
In this paper, a Time Domain CMOS Temperature Sensor with Differential Temperature Sensing Circuit and Reduced Supply Sensitivity is presented. Differential temperature sensing is achieved by using two delay generators, each for generating positively and negatively proportional to temperature delays respectively. The effective temperature signal can then be increased for quantization. The variation in supply voltage is sensed and converted to bias currents proportional to supply voltage. Temperature error due to 10% supply voltage variation can be reduced to less than ±1°C. A temperature dependent pulse will be generated and then digitized by a ripple counter using an external clock signal. Simulation results show that a ±0.6°C with two-point calibration can be achieved with a temperature variation from 0°C to 100°C. The circuit consumes 12μA and 70μA in static and temperature acquiring mode, respectively, and has a sampling rate of more than 80k samples/s.
Man Kay Law, Amine Bermak
ISCAS1
2007 A CMOS Image Sensor using Variable Reference Time Domain Encoding
abstract
In this paper, a Variable Reference Time Domain Encoding CMOS image sensor is presented. The time domain encoding vision sensor is known to suffer from slow conversion time, especially at low level of illumination. This is due to limited photocurrent generated to discharge the photodiode junction voltage to the reference voltage. A variable referencing scheme is proposed so that the reference voltage will be modulated and bounded by a specified deadline. The pixel consists of a photodiode, an analogue comparator, an 8-bit SRAM, a SR latch, and occupies an area of 32μm×35μm, with a fill-factor of 12.6 % using a 0.35μmCMOS process. Simulation results show that signal conversion can be achieved by using pre-defined threshold voltages. By using four levels of reference voltage and ⅒ of the original conversion time required by the original time domain encoding, over 70% reduction in total integration time can be achieved.
Man Kay Law, Amine Bermak
ISCAS1