Hailong Jiao

dblp:57/1627 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
19since 2021 · last 2025
0000-0002-2815-6168ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 10 first-author · 11 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Rate-Distortion Optimized Motion Estimation for Dynamic Point Cloud Geometry Compression
abstract
Dynamic point clouds serve as crucial representations of three-dimensional moving entities across diverse applications. The substantial amount of data in point clouds necessitates the development of efficient compression techniques. Motion estimation (ME) plays a crucial role in eliminating the temporal redundancy of point cloud sequences. However, prevailing ME methods suffer from the inaccurate geometry distortion measure and the imbalanced rate-distortion modeling, significantly impacting the coding performance. To address these challenges, we propose a rate-distortion (R-D) optimized ME scheme for dynamic point cloud geometry compression.
Qi Zhang 0029, Yiting Shao, Lixuan Meng, Hailong Jiao, Shan Liu 0001, Ge Li 0002
DCC4
2025 EGReg: Accurate and Efficient Registration of 3D Gaussian Splatting with Epipolar Geometry
abstract
Due to the capability of real-time high-quality rendering, 3D Gaussian Splatting (3DGS) has emerged as the mainstream data form in 3D reconstruction and novel viewpoint synthesis. Aligning the local coordinate systems of different 3DGS models to obtain a richer and more complete 3DGS model can greatly expand its reconstruction capabilities and flexibility. Therefore, this work focuses on the 3DGS registration task. Despite initial research progress, existing methods often struggle between registration efficiency and accuracy. To mitigate this issue, this work transforms the complex registration task into relative pose estimation task via epipolar geometry, proposing a coarse-to-fine, training-free 3DGS registration framework EGReg. Specifically, EGReg first rapidly filters camera pose pairs with high overlap in the hash-encoded space. It then uses a feature matching strategy to identify the best matching pair from the candidates and utilizes epipolar geometry to estimate their relative pose. Finally, based on the local camera poses and the estimated relative pose, our framework calculates the parameters of a similarity transformation to register the 3DGS models. We thoroughly validate EGReg on the ScanNet-GSReg and Objaverse datasets. Experimental results show that EGReg has significant advantages in both efficiency and accuracy.
Tuoyu Jin, Shiqiang Long, Hailong Jiao, Ge Li 0002
IJCNN4
2025 Robust Monolithic 3D Carbon-Based Computing-in-SRAM With Variation-Aware Bit-Wise Data-Mapping for High-Performance and Integration Density
abstract
Bit-serial computing-in memory with SRAM cells (SRAM-CIM) enables a full set of integer and floating-point arithmetic operations and various data-intensive computations. Carbon nanotube field-effect transistors (CN-MOSFETs) with high scalability, energy-efficiency, and low process thermal budget are attractive to realize high-dense monolithic three-dimensional (M3D) SRAM-CIM. However, CN-MOSFETs possess unique process variations with asymmetric spatial correlations which can significantly influence the performance and reliability of carbon-based SRAM-CIM. In this paper, new M3D-4N4P SRAM-CIM cells with CN-MOSFETs are proposed with optimized profiles for achieving ultra-high integration density while preserving robustness of data-access and computation. Furthermore, the variation-aware bit-wise data-mapping method is proposed for enhancing the performance of carbon-based SRAM-CIM by leveraging the spatial correlations of CN-MOSFETs. By minimizing the area skew of vertically-stacked layers, the areas of proposed M3D-4N4P SRAM-CIM cells are reduced by up to 50.32% compared to the previous 6N2P SRAM-CIM cells assuming carbon nanotube transistor technology. The proposed M3D-4N4P SRAM-CIM array also achieves by up to$2.17\times $higher throughput on arithmetic operations and 18.34% lower computing latency with 25.36% reduced energy consumptions on MAC-based benchmarks, respectively, compared to the previous 2D-6N2P SRAM-CIM array.
Dengfeng Wang, Weifeng He, Qin Wang 0009, Hailong Jiao, Yanan Sun 0003
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 A Low-Power Single-Phase Split-Controlled Flip-Flop With No Redundant Switching
abstract
Flip-flops with ultra-wide-range dynamic voltage scaling capability are attractive for ultra-low power applications. In this paper, a single-phase split-controlled flip-flop (SCFF) is proposed for ultra-wide-range dynamic voltage scaling. By employing a differential structure in the master latch and a unique split-controlled slave latch, the proposed SCFF possesses desirable features for ultra-low voltage operations. Designed in a 55-nm low-power CMOS technology, SCFF achieves 0.4 V minimum supply voltage, and reduces the power-delay-product (PDP) by up to 79.7% compared to the state-of-the-art flip-flops at the typical process corner with 12.5% input toggle rate. SCFF also maintains the benefit of low PDP at different process corners and different supply voltages.
Zhuoya Yan, Yingna Huang, Hailong Jiao
ISCAS3
2024 SoftAct: A High-Precision Softmax Architecture for Transformers Supporting Nonlinear Functions
abstract
Transformer-based deep learning networks are revolutionizing our society. The convolution and attention co-designed (CAC) Transformers have demonstrated superior performance compared to the conventional Transformer-based networks. However, CAC Transformer networks contain various nonlinear functions, such as softmax and complex activation functions, which require high precision hardware design yet typically with significant cost in area and power consumption. To address these challenges,SoftAct, a compact and high-precision algorithm-hardware co-designed architecture, is proposed to implement both softmax and nonlinear activation functions in CAC Transformer accelerators. An improved softmax algorithm with penalties is proposed to maintain precision in hardware. A stage-wise full zero detection method is developed to skip redundant computation in softmax. A compact and reconfigurable architecture with a symmetrically designed linear fitting module is proposed to achieve nonlinear functions. TheSoftActarchitecture is designed in an industrial 28-nm CMOS technology with the MobileViT-xxs network classifying the ImageNet-1k dataset as the benchmark. Compared with the state of the art,SoftActimproves up to 5.87% network accuracy under 8-bit quantization, 153.2× area efficiency, and 1435× overall efficiency.
Yuzhe Fu, Changchun Zhou 0001, Tianling Huang, Eryi Han, Yifan He 0002, Hailong Jiao
IEEE Trans. Circuits Syst. Video Technol.6
2023 An Energy-Efficient 3D Point Cloud Neural Network Accelerator With Efficient Filter Pruning, MLP Fusion, and Dual-Stream Sampling
abstract
Three-dimensional (3D) point cloud has been employed in a wide range of applications recently. As a powerful weapon for point cloud analysis, point-based point cloud neural networks (PNNs) have demonstrated superior performance with less computation complexity and parameters, compared to sparse 3D convolution-based networks and graph-based convolutional neural networks. However, point-based PNNs still suffer from high computational redundancy, large off-chip memory access, and low parallelism in hardware implementation, thereby hindering the applications on edge devices. In this paper, to address these challenges, an energy-efficient 3D point cloud neural network accelerator is proposed for on-chip edge computing. An efficient filter pruning scheme is used to skip the redundant convolution of pruned filters and zero-value feature channels. A block-wise multi-layer perceptron (MLP) fusion method is proposed to increase the on-chip reuse of features, thereby reducing off-chip memory access. A dual-stream blocking technique is proposed for higher parallelism while maintaining inference accuracy. Implemented in an industrial 28-nm CMOS technology, the proposed accelerator achieves an effective energy efficiency of 12.65 TOPS/W and 0.13 mJ/frame energy consumption for PointNeXt-S at 100 MHz, 0.9 V supply voltage, and 8-bit data width. Compared to the state-of-the-art point cloud neural network accelerators, the proposed accelerator enhances the energy efficiency by up to 66.6× and reduces the energy consumption per frame by up to 70.2×.
Changchun Zhou 0001, Yuzhe Fu, Siyuan Qiu, Ge Li 0002, Yifan He 0002, Hailong Jiao
ICCAD7
2023 Sagitta: An Energy-Efficient Sparse 3D-CNN Accelerator for Real-Time 3-D Understanding
abstract
Three-dimensional (3-D) understanding or inference has received increasing attention, where 3-D convolutional neural networks (3D-CNNs) have demonstrated superior performance compared to 2D-CNNs, since 3D-CNNs learn features from all three dimensions. However, 3D-CNNs suffer from intensive computation and data movement. In this article, Sagitta, an energy-efficient low-latency on-chip 3D-CNN accelerator, is proposed for edge devices. Locality and small differential value dropout are leveraged to increase the sparsity of activations. A full-zero-skipping convolutional microarchitecture is proposed to fully utilize the sparsity of weights and activations. A hierarchical load-balancing scheme is also introduced to increase the hardware utilization. Specialized architecture and computation flow are proposed to enhance the effectiveness of the proposed techniques. Fabricated in a 55-nm CMOS technology, Sagitta achieves 3.8 TOPS/W for C3D at a latency of 0.1 s and 4.5 TOPS/W for 3D U-Net at a latency of 0.9 s at 100 MHz and 0.91-V supply voltage. Compared to the state-of-the-art 3D-CNN and 2D-CNN accelerators, Sagitta enhances the energy efficiency by up to$379.6\times $and$11\times $, respectively.
Changchun Zhou 0001, Siyuan Qiu, Xugang Cao, Yuzhe Fu, Yifan He 0002, Hailong Jiao
IEEE Internet Things J.7
2023 APCCAS 2022 Guest Editorial Special Issue Based on the 18th Asia Pacific Conference on Circuits and Systems
abstract
The IEEE Asia Pacific Conference on Circuits and Systems (APCCAS) is the regional flagship conference of the IEEE Circuits and Systems Society (CASS) in Asia. This conference is a major international forum established by the IEEE Circuits and Systems Society for researchers to exchange their latest findings in circuits and systems. It covers a wide range of topics, including analog, mixed-signal, digital, communication, sensory, biomedical, power/energy, nonlinear, and artificial intelligence circuits and systems.
Xiaojin Zhao, Hailong Jiao, Wei Mao 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 CNN Accelerator at the Edge With Adaptive Zero Skipping and Sparsity-Driven Data Flow
abstract
An energy-efficient convolutional neural network (CNN) accelerator is proposed for low-power inference on edge devices. An adaptive zero skipping technique is proposed to dynamically skip the zeros in either activations or weights, depending on which has the higher sparsity. The characteristic of non-zero data aggregation is explored to enhance the effectiveness of adaptive zero skipping in performance boosting. To mitigate the load imbalance issue after zero skipping, a sparsity-driven data flow and low-complexity dynamic task allocation are employed for different convolution layers. Facilitated further by a two-stage distiller, the proposed accelerator achieves$5.42\times $,$3.41\times $, and$3.42\times $performance boosting for VGG16, AlexNet, and Mobilenet-v1, respectively, compared to the baseline. Implemented in a 55-nm low power CMOS technology, the proposed accelerator achieves an effective energy efficiency of 2.41 TOPS/W, 2.35 TOPS/W, and 0.64 TOPS/W for VGG16, AlexNet, and Mobilenet-v1, respectively, at 100 MHz and 1.08 V supply voltage.
Ming Liu 0022, Changchun Zhou 0001, Siyuan Qiu, Yifan He 0002, Hailong Jiao
IEEE Trans. Circuits Syst. Video Technol.5
2023 LightSeizureNet: A Lightweight Deep Learning Model for Real-Time Epileptic Seizure Detection
abstract
The monitoring of epilepsy patients in non-hospital environment is highly desirable, where ultra-low power wearable seizure detection devices are essential in such a system. The state-of-the-art epileptic seizure detection algorithms targeting such devices either rely on manual feature extractions, which can be biased due to the experience of experts, or deep neural networks, which suffer from high computation complexity. In this paper, we propose a lightweight deep learning model, LightSeizureNet (LSN), for real-time epileptic seizure detection based on raw EEG data in ultra-low power wearable seizure detection devices. The proposed LSN model includes a patient-independent version and a patient-specific version, both of which avoids manual feature extractions and high computation complexity, while maintaining good classification accuracy. Dilated one-dimensional (1D) convolution, global average pooling, and kernel-wise pruning are adopted to compress the LSN model. The proposed models are evaluated on the CHB-MIT scalp EEG database. The patient-independent LSN model achieves 97.09% accuracy with 6.2M MACs, while the patient-specific LSN model achieves 99.77% accuracy with 3.7 M MACs, which are competitive compared to the state of the art in terms of accuracy and complexity. Furthermore, the proposed model is highly interpretable, which is missing in many previous works. By using a uniform approach to explore the interpretability of the proposed model, fine-grained information such as the activated brain region and the frequency of brainwave during seizures is obtained for clinical diagnosis.
Siyuan Qiu, Wenjin Wang 0002, Hailong Jiao
IEEE J. Biomed. Health Informatics3
2022 Receiver Design With an Adjustable Energy-Signal-Quality Tradeoff for IoT Networks
abstract
The energy efficiency of an Internet of Things (IoT) receiver can be improved by introducing an adjustable tradeoff between signal quality and energy consumption. In good channel conditions, the receiver can be set to consume less energy per bit, without compromising signal quality in bad channel conditions. We propose a system-level receiver design that enables adequate configuration and combination of signal quality and energy tradeoffs in multiple receiver components. Co-design of all components is essential. We identify the most energy-efficient configurations in our system-level design under different channel conditions. With those configurations, the proposed receiver outperforms a state-of-the-art adjustable receiver with only an adjustable analog front end by several tens of percent in energy per successfully received bit and by$2\times $in energy-sensitivity configuration range. To show the efficacy of the proposed approach, we integrate a model of the proposed design into the OMNeT++ simulator and show the benefits on an environmental monitoring scenario. In this scenario, we report up to$6\times $energy savings for the entire transceiver compared to the conventional transceiver design without adjustable receiver.
Paul Detterer, Majid Nabi, Hailong Jiao, Twan Basten
IEEE Internet Things J.3
2022 A High-Efficiency Segmented Reconfigurable Cyclic Shifter for 5G QC-LDPC Decoder
abstract
A reconfigurable cyclic shifter is a key element of a QC-LDPC decoder, which is crucial for 5G communication systems. If a traditional reconfigurable cyclic shifter can only shift one input of variable size at a time, a traditional QC-LDPC decoder can only decode one codeword at a time as well. Part of the circuitry of the traditional QC-LDPC decoder inevitably stays in idle during the decoding process if the length (or the lifting parameter) of a codeword is not the maximum, resulting in low hardware efficiency. A segmented reconfigurable cyclic shifter is proposed in this paper, which can be divided into multiple segments of different sizes. Each segment can perform a cyclic shift of an input of different sizes and of different shift values independently. Furthermore, a methodology is proposed to upgrade any state-of-the-art QC-LDPC decoder to a segmented QC-LDPC decoder, by using the proposed segmented shifter. The upgraded segmented QC-LDPC decoder is able to parallelly decode multiple codewords (or inputs) of different lengths at a time. A test chip of the proposed segmented QC-LDPC decoder with the proposed segmented reconfigurable cyclic shifter has been fabricated in a 0.18-$\mu \text{m}$CMOS technology to demonstrate the feature of parallelly decoding multiple codewords. The performance analysis shows that when the number of small codewords is increased from 0 to 100000 per second, the throughput of the traditional QC-LDPC decoder drops from 844.80 Mbps to 4.40 Mbps, while the QC-LDPC decoder with the proposed segmented shifter only slightly drops to 814.01 Mbps. By applying the segmented QC-LDPC decoder in 5G base stations, the base stations are enabled to support more low-traffic users.
Hing-Mo Lam, Silin Lu, Hezi Qiu, Min Zhang 0041, Hailong Jiao, Shengdong Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 An Ultralow-Power 65-nm Standard Cell Library for Near/Subthreshold Digital Circuits
abstract
In this brief, a standard cell library targeting ultra-low voltages (ULVs) is designed in a 65-nm low-power CMOS technology to enable digital integrated circuits (ICs) to achieve good tradeoff among speed, power consumption, area, and reliability in the near/subthreshold region. The transistor sizes in the standard cells are optimized by threshold voltage tuning and technology feature exploration to enhance the switching and area efficiency of transistors. To balance the pull-up and pull-down networks, a Monte Carlo (MC) simulation-based balancing method is proposed. An AES-128 test circuit is fabricated by using the proposed library, achieving a minimum voltage of 0.338 V. Compared to the state of the art, our test circuit reduces the minimum energy delay product (EDP) by at least$3.55\times $.
Yuxuan Nie, Hailong Jiao
IEEE Trans. Very Large Scale Integr. Syst.3
2021 An Energy-Efficient Low-Latency 3D-CNN Accelerator Leveraging Temporal Locality, Full Zero-Skipping, and Hierarchical Load Balance
abstract
Three-dimensional convolutional neural network (3D-CNN) has demonstrated outstanding classification performance in video recognition compared to two-dimensional CNN (2D-CNN), since 3D-CNN not only learns the spatial features of each frame, but also learns the temporal features across all frames. However, 3D-CNN suffers from intensive computation and data movement. To solve these issues, an energy-efficient low-latency 3D-CNN accelerator is proposed. Temporal locality and small differential value dropout are used to increase the sparsity of activation. Furthermore, to fully utilize the sparsity of weight and activation, a full zero-skipping convolutional microarchitecture is proposed. A hierarchical load-balancing scheme is also introduced to improve resource utilization. With the proposed techniques, a 3D-CNN accelerator is designed in a 55-nm low-power CMOS technology, bringing in up to 9.89× speedup compared to the baseline implementation. Benchmarked with C3D, the proposed accelerator achieves an energy efficiency of 4.66 TOPS/W at 100 MHz and 1.08 V supply voltage.
Changchun Zhou 0001, Siyuan Qiu, Yifan He 0002, Hailong Jiao
DAC5
2021 Segmented Reconfigurable Cyclic Shifter for QC-LDPC Decoder
abstract
In various wireless communication standards, such as Wi-Fi, WiMAX, and 5G standards, QC-LDPC decoders are required to decode a codeword of variable length. Part of the decoder is in idle if the length of codeword is not the maximum. A reconfigurable cyclic shifter is a key element of a QC-LDPC decoder. If the shifter can only shift one input at a time, the QC- LDPC decoder can only decode one codeword at a time as well. A segmented reconfigurable cyclic shifter is proposed in this paper to enable a QC-LDPC decoder to parallelly decode multiple codewords. The proposed shifter can be reconfigured into multiple segments of different sizes. Each segment can perform a cyclic shift of different shift values independently. Therefore, the proposed shifter can enhance the QC-LDPC decoder to decode multiple codewords parallelly by inputting another codeword (or codewords) to re-activate the idle hardware. The test chip of QC- LDPC decoder with the proposed segmented reconfigurable cyclic shifter has been fabricated in 0.18μ CMOS technology which can parallelly decode six codewords.
Hing-Mo Lam, Silin Lu, Hezi Qiu, Hailong Jiao, Min Zhang 0041, Shengdong Zhang
ISCAS4
2021 A Pull-Up Adaptive Sense Amplifier Based on Dual-Gate IGZO TFTs
abstract
Thin-film transistor (TFT) not only is the backbone of display technology, but also brings in fancy applications in the area of flexible electronics. TFT-based flexible circuit design however faces various challenges, such as unipolar devices and severe nonuniformity. In this paper, a sense amplifier based on dual-gate indium-gallium-zinc oxide (IGZO) TFTs is proposed for TFT-based static random access memory (SRAM) circuits. The pull-up transistors in the proposed sense amplifier are implemented with dual-gate TFTs while the pull-down transistors are implemented with conventional bottom-gate TFTs. By tuning the top gate of the pull-up transistors to achieve a negative threshold voltage as well as employing a specialized strength- adaptive pull-up structure, the proposed sense amplifier aims to reduce the offset voltage, sensing delay, and sensing power consumption simultaneously. Compared to the state-of-the-art IGZO-based sense amplifier, the proposed sense amplifier achieves up to 66.1% lower offset voltage, 68.9% shorter sensing delay, 72% dynamic power savings, and 66.3% leakage power savings.
Silin Lu, Shengdong Zhang, Hailong Jiao
ISCAS4
2021 A Hybrid Digital Transmitter Architecture for High- Efficiency and High-Speed Applications
abstract
A Quadrature/Polar hybrid digital transmitter architecture (HB-TX) is proposed in this paper, which consists of a main digital power amplifier (DPA), an auxiliary DPA, and a low-bit phase selector. In the proposed HB-TX, coarse polar modulation is realized by combining the main DPA and a low-bit selector, while a fine quadrature modulation is realized by the asymmetrical quadrature recombination of main and auxiliary DPAs in a small range. Through coarse and fine modulation, the HB-TX realizes fast signal modulation and does not require high-speed and high-resolution phase modulator. With the coarse polar modulation, the HB-TX achieves high efficiency which is close to that of polar transmitter and no longer suffers from 3 dB back-off. Since the asymmetry of two-path DPA arrays improves the isolation between two channels, local oscillators (LOs) with 50% duty cycle is employed to further improve the HB-TX's efficiency. Simulation results show that the HB-TX with a 4-bit selector achieves 23.7-dBm peak output power with 39.2% peak power added efficiency (PAE). When modulating a 64-QAM signal with 80-Msym/s symbol rate and 6.5-dB peak average power rate (PAPR), the HB-TX improves the average drain efficiency from 16.3% (achieved in quadrature transmitter) to 22.2% The average output power that is delivered by the HB-TX is 17.6-dBm with 17.4% average PAE, while the error vector magnitude (EVM) is -35.87 dB.
Xiaolei Su, Zhengkun Shen, Zexue Liu, Hailong Jiao, Junhua Liu 0001, Huailin Liao
ISCAS7
2021 Pseudo Multi-Port SRAM Circuit for Image Processing in Display Drivers
abstract
N×N filter operations are frequently used in various image processing algorithms in display driver circuits. The implementation of the N×N filter requires a storage element to temporarily retain N or N-1 lines of display data. The traditional approach for this temporary storage element for N×N image filter is either by using N blocks of memory to store N lines of display data with single-port six-transistor (6T) SRAM bit-cells or by using N-1 blocks of memory to store N-1 lines of display data with dual-port eight-transistor (8T) SRAM bit-cells. In this paper, a pseudo multi-port SRAM circuit is proposed for N×N filter in image processing in a display driver integrated circuit (IC). By using pre-read mechanism and word-line/column selection signal forwarding technique, the proposed approach only stores N-1 lines of display data and can use the small single-port 6T SRAM bit-cells. For a 3×3 filter application, the proposed approach reduces the overall layout area by 50.3% and 30.2% compared to the traditional dual-port 8T and single-port 6T memory implementations, respectively, in an industrial 0.18- μm CMOS technology. Furthermore, the power used to perform a 3×3 filter operation is reduced by 47% and 30.8% with the proposed approach as compared to the traditional dual-port 8T and signal-port 6T memory circuits, respectively, with slight degradation of the access speed.
Hing-Mo Lam, Hezi Qiu, Min Zhang 0041, Hailong Jiao, Shengdong Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2021 Converter-Free Power Delivery Using Voltage Stacking for Near/Subthreshold Operation
abstract
Integrated circuits operating in the near/subthreshold region offer low energy consumption. However, due to the constrained voltage scalability of SRAMs, efficient power delivery is difficult to achieve. A traditional implementation would require at least two distinct voltage supplies generated by possibly two power converters. In this article, a new implementation for near/subthreshold operation is presented. The proposed implementation consists of a new “converter-free” design based on a three-level voltage stack operating at 1.8 V ± 5%. Here, the leakage current from the SRAMs in the top stack is recycled to sustain the near/subthreshold operation of the logic circuits in the two lower stacks. A test chip with the proposed voltage-stacking technique was implemented in a 28-nm low- Vth (LVT) fully depleted silicon on insulator (FDSOI) technology. The test chip is an ultralow-power advanced system-on-chip (SoC) consisting of an RISC-V core, a coarse-grained reconfigurable accelerator, and peripherals. The SoC uses a current sink and an adaptive body-bias controller for voltage regulation of the intermediate voltage rails between the stacks. The proposed system achieves up to 95% power delivery efficiency with negligible area overhead (~ 1%). The silicon measurement shows that the system energy efficiency is improved by 1.6× on average, and the energy consumption is reduced by 37% on average compared to the flat implementation.
Kamlesh Singh, Barry de Bruin, Hailong Jiao, Jos Huisken, Henk Corporaal, José Pineda de Gyvez
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Trading Sensitivity for Power in an IEEE 802.15.4 Conformant Adequate Demodulator
abstract
In this work, a design of an IEEE 802.15.4 con-formant O-QPSK demodulator is proposed, which is capable of trading off receiver sensitivity for power savings. Such design can be used to meet rigid energy and power constraints for many applications in the Internet-of-Things (IoT) context. In a Body Area Network (BAN), for example, the circuits need to operate with extremely limited energy sources, while still meeting the network performance requirements. This challenge can be addressed by the paradigm of adequate computing, which trades off excessive quality of service for power or energy using approximation techniques. Three different, adjustable approximation techniques are integrated into the demodulation to trade off effective signal quantization bit-width, filtering performance, and sampling frequency for power. Such approximations impact incoming signal sensitivity of the demodulator. For detailed trade-off analysis, the proposed design is implemented in a commercial 40-nm CMOS technology to estimate power and in Python to estimate sensitivity. Simulation results show up to 64% power savings by sacrificing $\tilde 7$ dB sensitivity.
Paul Detterer, Cumhur Erdin, Jos Huisken, Hailong Jiao, Majid Nabi, Twan Basten, José Pineda de Gyvez
DATE4
2020 A Compact Low-Power Data Retention Flip-Flop with Easy-Sleep Mode
abstract
A new data retention flip-flop, DRFF-Lite, which is designed based on the conventional master-slave flip-flop is proposed in this paper. No additional data retention circuitry is required by DRFF-Lite to implement the low-leakage data retention sleep mode, thereby providing easy and energy-efficient mode transitions. Due to the simplified topology, DRFF-Lite achieves up to 40% layout area savings as compared to the state-of-the-art data retention flip-flops in a 40-nm CMOS technology. Furthermore, DRFF-Lite reduces the sleep mode leakage power consumption, active mode leakage power consumption, active power consumption, and mode transition energy consumption by up to 31.5%, 60%, 18.5%, and 85.9%, respectively, under the impact of process parameter variations, while providing similar critical path propagation delay as compared to the state-of-the-art data retention flip-flops.
Hailong Jiao, Zhanliang Zhang
ISCAS1
2020 A Compensation System using Analog Voltage Adder with Continuous Output for AMOLED Display Drivers
abstract
An on-chip compensation system which is composed of an analog voltage adder and an 8-bit R-C digital-to-analog converter (DAC) is proposed for active matrix organic light-emitting diode (AMOLED) display drivers. The proposed analog voltage adder uses two groups of rail-to-rail input MOSFETs, which can work alternately, thereby extending the effective driving time (EDT) to 100%. The 8-bit R-C DAC is composed of a 5-bit R-DAC and a 3-bit C-DAC. The 3-bit C-DAC shares the capacitors with the analog voltage adder, reducing the layout area by 65% compared to the conventional method. This work is implemented in an industrial 0.18-μm CMOS technology. The measurement results show that the maximum INL of the system is 0.254 LSB, while the maximum DNL is 0.506 LSB. The layout area is only 9180 μm2per channel.
Hezi Qiu, Wenlong Bai, Hing-Mo Lam, Junjun An, Congwei Liao, Min Zhang 0041, Hailong Jiao, Shengdong Zhang
ISCAS8
2019 Trading Digital Accuracy for Power in an RSSI Computation of a Sensor Network Transceiver
abstract
To handle the rigid power and energy constraints in the Digital BaseBand (DBB) of Wireless Sensor Networks (WSN)s, we introduce approximate computing as a new power reduction method. The Received Signal Strength Indicator (RSSI) computation is a key element in DBB processing. We evaluate the trade-off in RSSI computation between Quality-of-Service (QoS) and power consumption through circuit-level approximation. RSSI elements are approximated in such a way that error propagation is minimized. In an industrial 40-nm CMOS technology, substantial energy savings up to 24% are achieved for every successfully transferred bit in DBB processing in a low- power listening WSN scenario.
Paul Detterer, Cumhur Erdin, Majid Nabi, José Pineda de Gyvez, Twan Basten, Hailong Jiao
DATE6
2019 A Compact Low-Voltage Segmented D/A Converter with Adjustable Gamma Coefficient for AMOLED Displays
abstract
A compact low-voltage segmented digital to analog (D/A) converter with adjustable gamma correction coefficient is proposed for source drivers of active matrix organic light-emitting diode (AMOLED) displays. The shared resistor-string is separated into several segments. The switching networks for each segmentation could be realized using low-voltage or mediumvoltage transistors. The outputs of the segmented DACs are added up by an analog adder to obtain the final DAC value. Compared to the conventional resistor string DAC (R-DAC) with high-voltage transistors, the proposed D/A converter features the layout area reduction by 84.7%, due to the elimination of high-voltage transistors. The proposed segmented DAC is implemented in an industrial 0.25-μm CMOS technology. The measurement results show that the differential and integral nonlinearity of the proposed D/A converter are 0.21 LSB and 1.31 LSB, respectively, which are significantly lower compared to the previously published high resolution DACs. The deviation of output voltages is also reduced by up to 2.9× compared to the previously published high resolution DACs.
Xinxin Huo, Wenlong Bai, Hing-Mo Lam, Congwei Liao, Min Zhang 0041, Shengdong Zhang, Hailong Jiao
ISCAS7
2018 Datawidth-Aware Energy-Efficient Multipliers: A Case for Going Sign Magnitude
abstract
Algorithms from many application domains, such as linear algebra and image/signal processing, heavily use the multiplication operator. Despite hardware support which is present in most modern cores, multiplication remains one of the most energy hungry arithmetic operations. This work explores how the energy efficiency of hardware multipliers can be improved by taking into account that the operands of a multiplication typically do not utilize the full width of the datapath. Seven datawidth-aware multiplier designs are implemented and evaluated. Post-layout energy analysis is performed to obtain the energy efficiency of each design for a number of representative benchmarks targeting the consumer market. The results show a significant improvement in energy efficiency compared to a 32-bit Baugh-Wooley baseline multiplier. A 32-bit sign-magnitude based design, integrated in a two's complement datapath, is shown to have a 1.38 times better energy efficiency than a baseline two's complement multiplier. In the best case (JPEG encoding), the energy efficiency is increased by a factor 2.25, demonstrating that a sign-magnitude multiplier, and datawidth-aware multipliers in general, are an attractive option for ultra low-energy designs.
Luc Waeijen, Hailong Jiao, Henk Corporaal, Yifan He 0002
DSD2
2018 IEEE Std P1838's flexible parallel port and its specification with Google's protocol buffers
abstract
IEEE Std P1838 is the DfT standard-under-development for 3D test access into dies meant to be used in 3D multi-die stack assemblies. P1838 is the first DfT standard to include a flexible parallel port (FPP): an optional, scalable multi-bit ('parallel') test access mechanism, offering higher test access bandwidth compared to the mandatory one-bit ('serial') port. In this paper, we describe P1838's FPP and propose a formal FPP specification language based on Google's Protocol Buffers (PBs), that potentially could become part of the standard. For a realistic example FPP, we provide its formal specification. Finally, we report on a demonstrator software tool, developed by using PBs-generated data access routines, that converts an FPP specification into a corresponding Verilog netlist.
Yu Li 0007, Ming Shao, Hailong Jiao, Adam Cron, Sandeep Bhatia, Erik Jan Marinissen
ETS3
2018 Multi-Bit Pulsed-Latch Based Low Power Synchronous Circuit Design
abstract
Pulsed-latches emerge as an ideal sequencing element for low power digital circuit design, serving as an alternative of flip-flops. In this paper, low power multi-bit pulsed-latches are proposed to construct pipeline stages in synchronous digital circuits. A method of integrating the proposed multi-bit pulsed-latches in the commercial design flows is also introduced. With the multi-bit pulsed-latches, up to 45% power savings are achieved for a variety of ITC benchmark circuits and an ARM Cortex-M0 as compared to the flip-flop based designs in an industrial 28-nm FDSOI CMOS technology. Furthermore, the power consumption of the clock distribution network and the layout area are reduced by up to 83% and 16%, respectively, with the proposed multi-bit pulsed-latches as compared to the flip-flop based designs.
Kamlesh Singh, Omar Alejandro Rodriguez Rosas, Hailong Jiao, Jos Huisken, José Pineda de Gyvez
ISCAS3
2018 On-Chip Toggle Generators to Provide Realistic Conditions during Test of Digital 2D-SoCs and 3D-SICs
abstract
In digital logic circuits, unconstrained scan tests are known to evoke much higher switching activity than functional modes. State-of-the-art ATPG tools have knobs to constrain the switching activity of the generated test to a user-defined functional level. Two-dimensional System-on-Chips (2D-SoCs) and three-dimensional stacked ICs (3D-SICs) are typically tested in a modular fashion, i.e., per embedded core or stacked die. At any moment during the test, one or more modules are tested ('module-under-test', MUT). In this work, we present the impact of the switching activity in the currently not-tested modules (which we refer to as 'neighbors' of the MUT) on overall IR-drop and propose a method to provide realistic conditions during modular test of digital 2D-SoCs and 3D-SICs, using on-chip programmable toggle generators.
Leonidas Katselas, Alkis A. Hatzopoulos, Hailong Jiao, Christos Papameletis, Erik Jan Marinissen
ITC3
2016 IEEE Std P1838: DfT standard-under-development for 2.5D-, 3D-, and 5.5D-SICs
abstract
For stacked integrated circuits, effective test access requires the design-for-test (DfT) features in the various dies to operate in a concerted way to transport test stimuli and responses from and to the external I/Os up and down through the stack. This 3D-DfT can be proprietary if all dies in the stack are made by a single company. However, in the likely case that the various dies in the stack originate from different companies, standardized 3D-DfT is required to guarantee inter-operability. IEEE Std P1838 is a standard-under-development that addresses exactly this issue. This paper presents a status report of P1838 and describes its three main hardware components: a serial control mechanism, a die wrapper register, and a flexible parallel port.
Erik Jan Marinissen, Teresa L. McLaurin, Hailong Jiao
ETS3
2016 Variations-tolerant 9T SRAM circuit with robust and low leakage SLEEP mode
abstract
Design of static random access memory (SRAM) circuits is challenging due to the degradation of data stability, weakening of write ability, increase of leakage power consumption, and exacerbation of process parameter variations with CMOS technology scaling. An asymmetrically ground-gated nine-transistor (9T) MTCMOS SRAM circuit is proposed in this paper for providing a low-leakage SLEEP mode with data retention capability. The worst-case static noise margin and write voltage margin are increased by up to 2.52x and 21.84%, respectively, with the asymmetrical 9T SRAM cells as compared to conventional six-transistor (6T) and eight-transistor (8T) SRAM cells under die-to-die process parameter variations in a 65nm CMOS technology. Furthermore, the mean values of static noise margin and write voltage margin are enhanced by up to 2.58x and 21.78% with the new 9T SRAM cells as compared with the conventional 6T and 8T SRAM cells under within-die process parameter fluctuations.
Hailong Jiao, Yongmin Qiu, Volkan Kursun
IOLTS1
2016 Variability-aware 7T SRAM circuit with low leakage high data stability SLEEP mode
Hailong Jiao, Yongmin Qiu, Volkan Kursun
Integr.1
2015 Memristor based computation-in-memory architecture for data-intensive applications
Said Hamdioui, Lei Xie 0005, Hoang Anh Du Nguyen, Mottaqiallah Taouil, Koen Bertels, Henk Corporaal, Hailong Jiao, Francky Catthoor, Dirk J. Wouters, Eike Linn, Jan van Lunteren
DATE7
2015 A Novel Robust and Low-Leakage SRAM Cell With Nine Carbon Nanotube Transistors
abstract
A novel static random-access memory (SRAM) cell with nine carbon nanotube MOSFETs (9-CN-MOSFETs) is proposed in this paper. With the new 9-CN-MOSFET SRAM cell, the read data stability is enhanced by 99.09%, while providing similar read speed as compared with the conventional six-transistor (6T) SRAM cell in a 16-nm carbon nanotube transistor technology. The worst-case write voltage margin is increased by 4.57× and 3.90× with the proposed 9-CN-MOSFET SRAM cell as compared with the conventional 6T SRAM cell and a previously published eight-transistor (8T) SRAM cell, respectively. A 1 Kibit SRAM array with the new memory cells consumes 34.18% and 12.27% lower leakage power as compared with the memory arrays with 6T and 8T SRAM cells, respectively, in idle mode. The overall electrical quality is enhanced by up to 13.63× with the proposed 9-CN-MOSFET memory circuit as compared with the other memory cells that are evaluated in this paper.
Yanan Sun 0003, Hailong Jiao, Volkan Kursun
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Low-leakage hybrid FinFET SRAM cell with asymmetrical gate overlap / underlap bitline access transistors for enhanced read data stability
abstract
The degraded read data stability and write ability of SRAM cells have become primary design concerns with CMOS technology scaling into the sub-22nm channel lengths. A new six-FinFET SRAM cell with asymmetrical bitline access transistors is proposed in this paper for enhancing the read data stability and suppressing the leakage power consumption in memory circuits. The bitline access transistor channel is underlapped on one side while overlapped by the gate terminal on the opposite side of the transistor. The asymmetrical bitline access transistors are weakened during read operations and strengthened during write operations as the direction of current flow is reversed. With the proposed asymmetrical six-FinFET SRAM cell, the read data stability is enhanced by up to 62% and the leakage power consumption is reduced by up to 49.3%, while maintaining similar write margin, cell layout area, read delay, and write delay as compared to a previously published asymmetrical six-FinFET SRAM cell in a 15nm FinFET technology.
Shairfe Muhammad Salahuddin, Hailong Jiao, Volkan Kursun
ISCAS2
2013 Characterization of mode transition timing overhead for net energy savings in low-noise MTCMOS circuits
abstract
Multi-threshold CMOS (MTCMOS) is commonly utilized for suppressing leakage currents in idle integrated circuits. The deactivation/reactivation energy consumption however degrades the effectiveness of MTCMOS technique for providing significant savings in total energy consumption in CMOS integrated circuits. The mode transition energy overheads of various recently published low-noise ground-gated MTCMOS circuits are characterized in this paper. With a digital triple-phase sleep signal slew rate modulated MTCMOS circuit, the overall mode transition energy consumption is reduced by up to 45.31% as compared to the other MTCMOS circuits that are evaluated in this paper in a UMC 80nm CMOS technology. Furthermore, the digital triple-phase MTCMOS circuit shortens the mode transition timing overhead by up to 65.26% as compared with the other MTCMOS circuits that are evaluated in this paper.
Hailong Jiao, Volkan Kursun
VLSI-SoC1
2013 Reactivation Noise Suppression With Sleep Signal Slew Rate Modulation in MTCMOS Circuits
abstract
Multi-threshold CMOS (MTCMOS) is commonly used for suppressing leakage currents in idle integrated circuits. Power and ground distribution network noise produced during SLEEP to ACTIVE mode transitions is an important reliability concern in MTCMOS circuits. Sleep signal slew rate modulation techniques for suppressing mode-transition noise are explored in this paper. A triple-phase sleep signal slew rate modulation (TPS) technique with a novel digital sleep signal generator is proposed. Reactivation time, mode-transition energy consumption, leakage power consumption, and layout area of different MTCMOS circuits are characterized under an equal-noise constraint. Influences of within-die and die-to-die parameter variations on the reactivation noise, time, and energy consumption of sleep signal slew rate modulated MTCMOS circuits are evaluated with a process imperfections aware robustness metric. The proposed triple-phase sleep signal slew rate modulation technique enhances the tolerance to process parameter fluctuations by up to$183.1\times$as compared to various alternative MTCMOS noise suppression techniques in a UMC 80-nm CMOS technology.
Hailong Jiao, Volkan Kursun
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Full-custom design of low leakage data preserving ground gated 6T SRAM cells to facilitate single-ended write operations
abstract
An asymmetrically ground-gated six-transistor (6T) SRAM circuit is presented in this paper for providing a low leakage data preserving SLEEP mode. By employing multiple write assist techniques, the write margin is enhanced by up to 2.73x and the write access time is reduced by up to 57.45% as compared with a previously published asymmetrically ground-gated 6T SRAM circuit in a TSMC 65nm CMOS technology. Furthermore, the new ground-gated 6T memory circuit enhances the data stability by 2.09x and reduces the leakage power consumption by 58.55% as compared to a ground-gated memory array with conventional 6T SRAM cells. A design methodology is presented to optimize the asymmetrically ground-gated 6T SRAM circuits for achieving the highest overall electrical quality.
Hailong Jiao, Volkan Kursun
ISCAS1
2012 Threshold Voltage Tuning for Faster Activation With Lower Noise in Tri-Mode MTCMOS Circuits
abstract
A new threshold voltage tuning methodology is explored in this paper to minimize the peak power/ground bouncing noise with smaller sleep transistors in multi-threshold CMOS (MTCMOS) circuits. Different circuit techniques with the threshold voltage tuning strategy lower the activation noise, the activation delay, and the size of the additional sleep transistors by up to 27.76%, 32.66%, and 85.71%, respectively, as compared to a previously published noise-aware MTCMOS circuit with standard zero-body-biased high threshold voltage sleep transistors in a UMC 80-nm CMOS technology.
Hailong Jiao, Volkan Kursun
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Ground Bouncing Noise Suppression Techniques for Data Preserving Sequential MTCMOS Circuits
abstract
Ground distribution network noise produced during sleep-to-active mode transitions is an important reliability concern in standard multi-threshold CMOS (MTCMOS) circuits. Different noise-aware sequential MTCMOS circuits are explored in this paper. A low-leakage data retention sleep mode is implemented with smaller centralized sleep transistors to suppress the ground bouncing noise produced during reactivation events in sequential MTCMOS circuits. Ground bouncing noise, leakage power consumption, data stability, and area overheads of different sequential MTCMOS circuits are evaluated with a 90-nm CMOS technology. The peak amplitude of ground bouncing noise is reduced by up to 94.16% with the noise-aware MTCMOS techniques as compared to the conventional Mutoh flip-flop. The application space of different data retention MTCMOS circuit techniques is identified with various design metrics in this paper.
Hailong Jiao, Volkan Kursun
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Smooth awakenings: Reactivation noise suppressed low-leakage and robust MTCMOS flip-flops
abstract
Ground bouncing noise produced during the sleep to active mode transitions is an important reliability concern in multi-domain Multi-Threshold CMOS (MTCMOS) integrated circuits. Ground bouncing noise, leakage power consumption, and data stability of MTCMOS flip-flops are evaluated in this paper. The effectiveness of different circuit techniques is discussed for achieving lower noise during the reactivation events while maintaining robust and low-leakage data retention capability in MTCMOS flip-flops.
Hailong Jiao, Volkan Kursun
ISCAS1
2010 Reactivation noise suppression with threshold voltage tuning in sequential MTCMOS circuits
abstract
Ground bouncing noise produced during reactivation events is an important challenge in multi-threshold CMOS (MTCMOS) circuits. A threshold voltage tuning technique based on forward body bias is proposed in this paper to alleviate the ground bouncing noise in sequential MTCMOS circuits. With the new threshold voltage tuning technique, the peak ground bouncing noise is reduced by up to 91.70% as compared to the previously published sequential MTCMOS circuits in a UMC 80nm CMOS technology. The design tradeoffs among various important design metrics are evaluated with different data preserving sequential MTCMOS circuits in this paper.
Hailong Jiao, Volkan Kursun
VLSI-SoC1