Xiaoxin Cui

dblp:23/1520 · DBLP profile ↗
← Back
53ranked-venue papers
1as first author
36since 2021 · last 2026
0000-0002-0394-8839ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 1 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Elmore Delay Model Based Test Method for Post-Bond TSVs
Zhicheng Shao, Xiaole Cui, Xiaoxin Cui
ETS3
2026 DyNeuro: A Hybrid Neuromorphic Accelerator with Dynamic Spatio-Temporal Variation Adaptations
Youming Yang 0002, Yi Zhong 0002, Li Lun, Tao Zhang 0140, Xiaoxin Cui, Yuan Wang 0001
ISCAS5
2026 PAICar: a prototype of an embodied neuromorphic intelligent robot platform
Mingkai Liu, Jingyi Zhong, Yi Zhong 0002, Zilin Wang 0001, Chenglong Zou, Xiaoxin Cui, Jian Cao 0002, Yuan Wang 0001
Sci. China Inf. Sci.8
2025 Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate Gradients
abstract
Spiking neural networks (SNNs) have shown their competence in handling spatial-temporal event-based data with low energy consumption. Similar to conventional artificial neural networks (ANNs), SNNs are also vulnerable to gradient-based adversarial attacks, wherein gradients are calculated by spatial-temporal back-propagation (STBP) and surrogate gradients (SGs). However, the SGs may be invisible for an inference-only model as they do not influence the inference results, and current gradient-based attacks are ineffective for binary dynamic images captured by the dynamic vision sensor (DVS). While some approaches addressed the issue of invisible SGs through universal SGs, their SGs lack a correlation with the victim model, resulting in sub-optimal performance. Moreover, the imperceptibility of existing SNN-based binary attacks is still insufficient. In this paper, we introduce an innovative potential-dependent surrogate gradient (PDSG) method to establish a robust connection between the SG and the model, thereby enhancing the adaptability of adversarial attacks across various models with invisible SGs. Additionally, we propose the sparse dynamic attack (SDA) to effectively attack binary dynamic images. Utilizing a generation-reduction paradigm, SDA can fully optimize the sparsity of adversarial perturbations. Experimental results demonstrate that our PDSG and SDA outperform state-of-the-art SNN-based attacks across various models and datasets. Specifically, our PDSG achieves 100% attack success rate on ImageNet, and our SDA obtains 82% attack success rate by modifying only 0.24% of the pixels on CIFAR10DVS. The code is available at https://github.com/ryime/PDSG-SDA.
Li Lun, Kunyu Feng, Qinglong Ni, Ling Liang 0003, Yuan Wang 0001, Ying Li 0056, Dunshan Yu, Xiaoxin Cui
CVPR8
2025 NeuroHexa: A 2D/3D-Scalable Model-Adaptive NoC Architecture for Neuromorphic Computing
abstract
Neuromorphic computing has endeavored a novel computing paradigm that entails a bio-inspired architecture to reproduce the remarkable functionalities of the human brain, such as massively parallel processing and extremely low-power consumption. However, those promising merits can be greatly canceled by the mismatched communication infrastructure in large-scale hardware implementation, in view of the vast degree of neural connectivity, the unstructured spike dataflow, and the unbalanced model workload assignment. In an effort to tackle those challenges, this work presents NeuroHexa, a network-on-chip (NoC) architecture intended for multi-core neuromorphic design. NeuroHexa adopts a customized intra-chip hexagonal topology, which can be further cascaded in 6 directions by either 2D or 3D chiplet integration. Designed in globally asynchronous, locally synchronous (GALS) methodology, a group of processing nodes can operate in independent work pace to further improve resource utilization. To satisfy the varied requirement of data reuse across the chip, NeuroHexa proposes a flexible multicast routing mechanism to best adapt to the model-defined dataflow. And under a specific congestion scenario, NeuroHexa can switch its routing algorithm between deterministic routing and fully adaptive routing modes. The presented NoC router is evaluated in 28nm CMOS, where we achieve the maximal throughput as 179.2Gbps, and the best energy efficiency as 4.872pJ/packet at the area overhead of 0.0226mm2.
Yi Zhong 0002, Zilin Wang 0001, Yipeng Gao, Xiaoxin Cui, Xing Zhang 0002, Yuan Wang 0001
DATE4
2025 CROSSCUT: A Multi-Core Neuromorphic Accelerator Improving Resource-Utilization
abstract
Neuromorphic computing is attracting significant attention due to its bio-mimetic characteristics. Consequently, neuromorphic hardware platforms have emerged as innovative computing architectures for acceleration. However, the fixed nature of data flow and resources leads to considerable inefficiencies in storage and computation, thereby limiting both utilization efficiency and overall performance. This severely hinders the deployment of edge artificial intelligence (AI) models. To address these issues, we present a multi-core neuromorphic accelerator named CROSSCUT. This crossbar-based system supports both spiking neural network (SNN) and artificial neural network (ANN) paradigms and has a capacity of 256K neurons and 288M synapses. By leveraging the Neuron Package Mechanism (NPM) and Synapse Compress Mechanism (SCM), CROSSCUT can increase input data scale by 64 times and reduce wasted resources and computations by 46.7%, ensuring high compatibility with diverse network structures in machine learning models. Additionally, a Tree-Mesh hybrid network on chip (NoC) is constructed for inter-core communication. Implemented on Xilinx XCVU9P FPGA, CROSSCUT can achieve a peak performance of 431.9 GSOPS/s and 121.13 GSOPS/W energy efficiency. The inference accuracy on MNIST is 98.2%.
Youming Yang 0002, Yi Zhong 0002, Zilin Wang 0001, Tao Zhang 0140, Li Lun, Yingying Cui, Xiaoxin Cui, Song Jia, Yuan Wang 0001
ISCAS7
2025 A Reconfigurable Digital Compute-In-Memory Heterogeneous Macro for Differential Frame Convolution and Spiking Neural Network
abstract
The application of artificial neural network (ANN) in video processing encounters significant challenges, including large data volumes, numerous linear operations, and high power consumption. The fusion of convolutional neural network (CNN) and spiking neural network (SNN) provides a dual benefit of achieving high accuracy while maintaining low power consumption. However, ongoing challenges remain in minimizing multiply-accumulate (MAC) operations and optimizing data movement. In this work, we propose a reconfigurable digital compute-in-memory (RDCIM) heterogeneous macro without the sense amplifier, tailored for the diverse computational demands of CNN and SNN. To improve energy efficiency, differential frame convolution (DFC) is adopted to mitigate the computational overhead. In addition, computational resources are functionally reused to accommodate four data flow types, supporting both DFC and SNN operations. Implemented by TSMC 28nm technology, the proposed RDCIM heterogeneous macro achieves the peak energy efficiency of 29.13 TOPS/W for DFC and 0.56 pJ/SOP for SNN, operating at a frequency of 284 MHz.
Li Lun, Zhenhui Dai, Yingying Cui, Xiaoxin Cui
ISCAS6
2025 HyNITA: A Neuromorphic Inference and Training Accelerator for Hybrid ANN-SNN Fusion Models
abstract
In order to achieve the brain-like advantages over conservative computers, previous neuromorphic researchers have stretched the hardware explorations of the hybrid artificial neural network (ANN) and spiking neural network (SNN) inference approaches, as well as the efficient bio-plausible and gradient-based SNN training mechanisms. However, a versatile accelerator for both ANN-SNN inference and training is little addressed. In this work, we introduce HyNITA, a neuromorphic processor that supports accelerating both inference and training tasks of hybrid ANN and SNN models. Regarding the similarity and distinction, a pair of working stages are distinguished and distributed to multiple simple cores. The accelerator optimizes the interchange dataflow in a scalable chip design, following a reconfigurable design methodology to integrate the involved equation calculations in the dynamic process of neurons. The evaluation results show it achieves an accuracy of 99.65% and 99.34% on training ANN MNIST and SNN N-MNIST datasets.
Yi Zhong 0002, Li Lun, Zilin Wang 0001, Jinhao Ruan, Yipeng Gao, Xiaoxin Cui, Xing Zhang 0002, Yuan Wang 0001
ISCAS6
2025 A VC Dimension-Oriented Improvement Method of PUFs for the Anti-Modeling-Attack Capability
abstract
The physical unclonable function (PUF) serves as a security primitive of circuits, which is applicable to the embedded systems with lightweight authentication function. However, the modeling attack, which estimates the unknown CRPs by establishing the mathematical model of PUF, is a real threat to the PUF based crypto-systems. Subsequently, the anti-modeling-attack PUF becomes a research hotspot. The systematic design method of secure PUF is still an open issue, although some secure PUF schemes have been proposed based on the repeated trials. This work proposes a security improvement method of PUFs to enhance the anti-modeling-attack capability. The growth function and the Vapnik-Chervonenkis (VC) dimension of PUF are defined as the indicators of PUF security. The proposed method regards the improvement of PUF as an optimization problem, which aims to obtain a PUF scheme with the better security indicators. Guided by the indicators, the proposed method is able to specify the improvement sites of PUF and the techniques to be applied. In addition, three approaches are proposed to inspire the new security improvement techniques. An improved arbiter PUF and an improved array-based PUF are designed as the instances of the results from the proposed method. Both of the improved PUF schemes have the stronger security than the original schemes.
Xiaole Cui, Sunrui Zhang, Xiaoxin Cui
ACM Trans. Embed. Comput. Syst.4
2025 Toward High-Accuracy and Low-Latency Spiking Neural Networks With Two-Stage Optimization
abstract
Spiking neural networks (SNNs) operating with asynchronous discrete events show higher energy efficiency with sparse computation. A popular approach for implementing deep SNNs is artificial neural network (ANN)-SNN conversion combining both efficient training of ANNs and efficient inference of SNNs. However, the accuracy loss is usually nonnegligible, especially under few time steps, which restricts the applications of SNN on latency-sensitive edge devices greatly. In this article, we first identify that such performance degradation stems from the misrepresentation of the negative or overflow residual membrane potential in SNNs. Inspired by this, we decompose the conversion error into three parts: quantization error, clipping error, and residual membrane potential representation error. With such insights, we propose a two-stage conversion algorithm to minimize those errors, respectively. In addition, we show that each stage achieves significant performance gains in a complementary manner. By evaluating on challenging datasets including CIFAR- 10, CIFAR- 100, and ImageNet, the proposed method demonstrates the state-of-the-art performance in terms of accuracy, latency, and energy preservation. Furthermore, our method is evaluated using a more challenging object detection task, revealing notable gains in regression performance under ultralow latency, when compared with existing spike-based detection algorithms. Codes will be available at: https://github.com/Windere/snn-cvt-dual-phase.
Yuhao Zhang 0007, Shuang Lian, Xiaoxin Cui, Rui Yan 0005, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.4
2024 A Convolutional Spiking Neural Network Accelerator with the Sparsity-Aware Memory and Compressed Weights
abstract
The spiking neural network (SNN) has advantage in the edge AI applications for its spatiotemporal sparsity. The high energy efficiency is an important concern in the study of SNN accelerator designs. In this paper, a lightweight event-driven convolutional SNN accelerator that utilizes the sparsity of both the spike events and the network weights is proposed. In the event-driven mode, the proposed accelerator uses the compressed input spikes and a spike-oriented convolution data flow. An output spike compressor is also designed. To balance the computation performance and the memory space occupancy, a spike sparsity-aware memory scheme that automatically switches the spike format by a real-time monitoring strategy is designed. The compression memories and a buffer for network weights are designed to save the on-chip memory space. The accelerator prototype is verified on the Xilinx Virtex XCVU9P FPGA platform. It achieves an equivalent performance of 139.5GFLOPS on the N-MNIST dataset. Compared to the baseline using the same computational resources, the proposed accelerator can improve the inference performance, the inference energy efficiency and the memory space by 4.6, 3.6 and 1.6 times, respectively. The proposed accelerator has advantages in energy efficiency and hardware overhead compared to the previous works on the same hardware platform. Neuralmorphic computing Spiking neural network accelerator Sparse spikes Sparse matrix compression Field-programmable gate array
Xiaole Cui, Sunrui Zhang, Mingqi Yin, Xiaoxin Cui
ASAP6
2024 An Energy-Efficient Differential Frame Convolutional Accelerator with on-Chip Fusion Storage Architecture and Pixel-Level Pipeline Data Flow
abstract
Convolutional neural networks require a huge amount of computation in video applications. For some specific tasks, such as surveillance, differential frame convolution reuses inter-frame data and significantly reduces multiplication and accumulation. However, there are still some challenges in improving energy efficiency of differential frame convolution on chips. Firstly, differential frame convolution brings additional on-chip storage for reusing inter-frame data. Secondly, in post-processing of differential frame convolution, there are more memory accessing and arithmetic logic operations. Therefore, sparse working mode is of vital importance for the post-processing. In response to these challenges, this work proposes an on-chip fusion storage architecture for energy-efficient differential frame convolution and a pixel-level pipeline data flow that supports the sparsity of features. The simulation of our accelerator implemented in 28nm CMOS can achieve energy efficiency by 3.09× compared with other state-of-the-art works
Zhenhui Dai, Yi Zhong 0002, Kunyu Feng, Yuan Wang 0001, Dunshan Yu, Xiaoxin Cui
ISCAS10
2024 SPAT: FPGA-based Sparsity-Optimized Spiking Neural Network Training Accelerator with Temporal Parallel Dataflow
abstract
Spiking neural networks (SNNs), as biologically inspired computational models, possess significant advantages in energy efficiency due to their event-driven operations. However, challenges remain in attaining high computational efficiency for SNN training. In this work, we propose a novel SNN training accelerator employing temporal parallelism and sparsity optimizations to achieve superior efficiency. A temporal parallel dataflow is designed to concurrently integrate spikes across multiple time steps, enhancing throughput and data reuse. To reduce latency and improve energy efficiency, we leverage the sparsity of SNNs and employ methods such as zero gating and zero skipping. Implemented on a field-programmable gate array (FPGA), the proposed training accelerator demonstrates 2.3-fold speedup and 15.7-fold energy reduction compared to NVIDIA A100 GPU on N-MNIST dataset.
Li Lun, Mingqi Yin, Zhenhui Dai, Xiaole Cui, Xiaoxin Cui
ISCAS8
2024 A 16.41 TOPS/W CNN Accelerator with Event-Based Layer Fusion for Real-Time Inference
abstract
This paper proposes a convolutional neural network (CNN) accelerator architecture for real-time tasks in edge devices. An event-based layer fusion technique is adopted to eliminate on-chip storage requirements and off-chip data movement caused by features. Cross-layer pipeline is elaborated during layer fusion to obtain high throughput and low latency. An adaptive fully unrolling event-driven core is designed and a cyclic storage method is exploited to reduce the storage space for partial sum in the core. Modified LeNet is accelerated with the proposed architecture. The accelerator can reach an energy efficiency of 16.41 TOPS/W and a latency of 0.85μs under TSMC 28nm technology, and a frame rate of 369.4K FPS under FPGA.
Li Lun, Zhenhui Dai, Xiaoxin Cui
ISCAS5
2024 NeuroREC: A 28-nm Efficient Neuromorphic Processor for Radar Emitter Classification
abstract
Radar emitter classification (REC) plays an important role in modern warfare. Traditional REC methods have difficulty identifying complex radar signals in the present day. Inspired by biology, spiking neural networks (SNNs) have gradually gained widespread attention due to their low power characteristics. Compared with convolutional neural networks (CNNs), SNNs are more suitable for application in the field of REC. The reason is that SNN can not only maintain higher accuracy in the presence of noise interference, but also reduce the power consumption of mobile devices. However, it is challenging to make full use of the input sparsity of radar emitter signals and the weight sparsity of pruned SNN models. In this paper, a 28-nm neuromorphic processor for REC named NeuroREC is proposed. It uses matrix compression algorithms to store sparse weights on chip, and designs corresponding spike detection circuits for this purpose. As a single-core design, we propose a ping-pong running mechanism to alleviate the imbalance between IO throughput and peak performance. Two SNN models for classifying RadioML2016.b and RadioML2018.a datasets are deployed on the chip, achieving competitive accuracy with only 8 timesteps, and demonstrating better robustness than CNN. Fabricated in 28-nm CMOS process, NeuroREC runs at frequencies ranging from 22.5MHz to 744MHz. Under specific sparsity conditions, it can reach an energy efficiency of 7.22TSOP/W for 8-bit weight.
Zilin Wang 0001, Zehong Ou, Yi Zhong 0002, Youming Yang 0002, Li Lun, Hufei Li, Jian Cao 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 The Resistance Analysis Attack and Security Enhancement of the IMC LUT Based on the Complementary Resistive Switch Cells
abstract
The resistive random access memory (RRAM) based in-memory computing (IMC) is an emerging architecture to address the challenge of the “memory wall” problem. The complementary resistive switch (CRS) cell connects two bipolar RRAM elements anti-serially to reduce the sneak current in the crossbar array. The CRS array is a generic computing platform, for the arbitrary logic functions can be implemented in it. The IMC CRS LUT consumes fewer CRS cells than the static CRS LUT. The CRS array has built-in polymorphic characteristics because the correct logic function cannot be distinguished based on the circuit layout. However, the logic state of every CRS cell can be readout after each operation. It helps the attacker to recover the correct function of the IMC CRS LUT. This work discusses the resistance analysis attack of the IMC LUT based on the CRS array. The proposed resistance analysis attack method is able to be applied to different computation styles based on the CRS array, such as the CRS IMPLY, CRS NOR-OR/NAND-AND, and so on. The attacker can recover the logic function of the LUT by tracing the states of CRS cells. Furthermore, an improved IMC CRS LUT method is proposed and discussed to enhance security. The simulation and analysis results show that the improved IMC CRS LUT can resist various attacks, and it maintains the polymorphic characteristics of the IMC CRS LUT. And the N-bit full adder circuit based on the improved IMC CRS NOR-OR LUTs achieves the best performance compared with the previous counterparts.
Xiaole Cui, Mingqi Yin, Xiaoxin Cui
ACM Trans. Design Autom. Electr. Syst.4
2024 Dy-MFNS-CAC: An Encoding Mechanism to Suppress the Crosstalk and Repair the Hard Faults in Rectangular TSV Arrays
abstract
Through-silicon vias (TSVs) play the role of vertical electrical interconnections in the emerging three-dimensional stacked integrated circuits. However, the hard faults and the crosstalk faults may occur in the TSV array simultaneously. The hard faults are the catastrophic faults that lead to functional failure, and the crosstalk faults deteriorate the signal integrity of TSV array. The previous fault tolerant techniques for this problem are based on the Fibonacci numeral system-based crosstalk avoidance code (FNS-CAC), and the bit overhead and the reparability need to be further improved. This article proposes the dynamic modified Fibonacci numeral system (Dy-MFNS) and the Dy-MFNS based crosstalk avoidance code (Dy-MFNS-CAC). The cascaded Dy-MFNS adders and the Dy-MFNS codec are designed to generate the Dy-MFNS-CAC codewords according to the health status of TSVs. The generated codewords are able to suppress the crosstalk and repair the hard faults simultaneously, under the control of the TSV fault flags. The simulation results show that the proposed scheme has advantages on both the bit overhead and the reparability compared with the FNS-CAC based techniques.
Xiaole Cui, Xiaoxin Cui
IEEE Trans. Reliab.3
2024 Marmotini: A Weight Density Adaptation Architecture With Hybrid Compression Method for Spiking Neural Network
abstract
Brain-inspired spiking neural network (SNN) has recently attracted widespread interest owing to its event-driven nature and relatively low-power hardware for transmitting highly sparse binary spikes. To further improve energy efficiency, some matrix compression algorithms are used for weight storage. However, the weight sparsity of different layers varies greatly. For a multicore neuromorphic system, it is difficult for the same compression algorithm to adapt to all the layers of SNN model. In this work, we propose a weight density adaptation architecture with hybrid compression method for SNN, named Marmotini. It is a multicore heterogeneous design, including three types of cores to complete computation of different weight sparsity. Benefiting from the hybrid compression method, Marmotini minimizes the waste of neurons and weights as much as possible. Besides, for better flexibility, a reconfigurable core that can be configured to compute convolutional layer or fully connected layer is proposed. Implemented on Xilinx Kintex UltraScale XCKU115 field-programmable gate array (FPGA) board, Marmotini can operate at 150-MHz frequency, achieving 244.6-GSOP/s peak performance and 54.1-GSOP/W energy efficiency at 0% spike sparsity.
Zilin Wang 0001, Yi Zhong 0002, Zehong Ou, Youming Yang 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2023 A Spiking Neural Network Accelerator based on Ping-Pong Architecture with Sparse Spike and Weight
abstract
Spiking neural networks (SNNs) have attracted widespread interest due to their event-driven and low-power nature. Compared to Artificial Neural Networks (ANNs), SNNs have time dimension information and present more realistic brain-inspired computing models. However, it is challenging to deploy sparse spiking neuron network models on dense neuromorphic processors. In this paper, a spiking neural network accelerator with sparse spike and weight is presented, using ping-pong architecture to improve system data throughput. To reduce the inference delay, the proposed accelerator supports the decoupling of calculation of membrane potential and leaky integrate-and-fire (LIF) dynamics computing in the feedforward neural networks. Implemented on Xilinx Kintex UltraScale FPGA, the accelerator can achieve the peak performance of 65.7 GSOP/s and the energy efficiency of 41.7 GSOP/W in the task of classifying MNIST dataset. Under the full load, the whole system can run ping-pong when more than 43 time steps are calculated at a time.
Zilin Wang 0001, Yi Zhong 0002, Xiaoxin Cui, Yisong Kuang, Yuan Wang 0001
ISCAS3
2023 An Area-Efficient In-Memory Implementation Method of Arbitrary Boolean Function Based on SRAM Array
abstract
In-memory computing is an emerging computing paradigm to breakthrough the von-Neumann bottleneck. The SRAM based in-memory computing (SRAM-IMC) attracts great concerns from industries and academia, because the SRAM is technology compatible with the widely-used MOS devices. The digital SRAM-IMC scheme has advantages on stability and accuracy of computing results, compared with the analog SRAM-IMC schemes. However, few logic operations can be implemented by the current digital SRAM-IMC architectures. Designers have to insert some special logic modules to facilitate the complex computation. To address this issue, this work proposes an area-efficient implementation method of arbitrary Boolean function in SRAM array. Firstly, a two-input SRAM LUT is designed to realize the arbitrary two-input Boolean functions. Then, the logic merging and the spatial merging techniques are proposed to reduce the area consumption of the SRAM-IMC scheme. Finally, the SOP-based SRAM-IMC architecture is proposed, and the merged SOPs are mapped into and computed in it. The evaluation results on LGsynth’91, IWLS’93 and EPFL benchmarks show that, the area of the synthesis results based on the ABC tool is 3.69, 5.72 and 1.86 times of the circuit area from the proposed SRAM-IMC scheme in average respectively. Furthermore, the circuit area from the original SOP-based SRAM-IMC scheme is 2.07, 1.99 and 1.86 times in average of the circuit area from the proposed SRAM-IMC scheme respectively. The performance evaluation results show that the cycle consumption of the proposed SRAM-IMC scheme is independent to the scale of the input Boolean functions.
Sunrui Zhang, Xiaole Cui, Xiaoxin Cui
IEEE Trans. Computers4
2023 Mosaic-3C1S: A Low Overhead Crosstalk Suppression Scheme for Rectangular TSV Array
abstract
The through silicon via (TSV) is one of the important enabling technologies of stacked 3-D ICs. However, the crosstalk is generated when signals are transmitted through the closely clustered TSVs due to the coupling effect. The crosstalk avoidance code (CAC) techniques have attracted great concerns in recent years, for the TSV-to-TSV crosstalk deteriorates the signal integrity. Nevertheless, the CAC techniques usually require specific shape of the TSV array, e.g.,$3\times N$array. This work proposes a crosstalk suppression scheme combining the CAC method and the static shielding technique, named Mosaic-3C1S. The proposed scheme is able to be applied to arbitrary rectangular TSV arrays. In the proposed scheme, the$2\times 2$TSV subarray, which contains three signal TSVs and one shielding TSV, is used as the basic unit to tile the rectangular TSV array. The CAC method is applied to the signal TSVs in the$2\times 2$subarrays, and the raw data are transmitted by the rest boundary TSVs, if any. In order to further reduce the bit overhead, a triangular numeral system (TNS)-based CAC, named TNS-CAC, is proposed and applied. The simulation results show that the proposed TNS-CAC has advantages on bit overhead and codec area. The proposed Mosaic-3C1S scheme with TNS-CAC is able to suppress the crosstalk between TSVs to 6C level. The bit overhead of the proposed scheme is between 25% and 45%, which is lower than that of the previous CAC methods for TSV arrays for the most cases.
Xiaole Cui, Xiaoxin Cui
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 An Evaluation Method of the Anti-Modeling-Attack Capability of PUFs
abstract
The physical unclonable function (PUF) is regarded as the root of trust of hardware systems. However, it suffers from the modeling attacks based on machine learning (ML) algorithms. Subsequently, the anti-modeling-attack PUF is of great concern from academia and industries in recent years. In practice, the security of a given PUF is evaluated after the pertinent attacks. However, these evaluation methods are not helpful for the PUF design, because the relationship between the PUF structure and the anti-modeling-attack capability is not established explicitly. This work proposes a security evaluation method of PUF based on the mimic attack and the Probably Approximately Correct (PAC) theory. The anti-modeling-attack capability of PUF is measured by the corresponding area enclosed by the evaluation curve. Twenty representative types of PUFs with different sizes are evaluated by the proposed method. It shows that the proposed method is effective, because the evaluation results are consistent with the difficulties of modeling attacks for the corresponding PUFs in the design practices. And the evaluation results are able to assist the PUF design.
Xiaole Cui, Yun Liu 0031, Xiaoxin Cui
IEEE Trans. Inf. Forensics Secur.4
2022 An obfuscation scheme of scan chain to protect the cryptographic chips
abstract
The scan chain, as the most popular structured design for testability (DfT) approach, significantly improves the controllability and observability of the circuit under test (CUT). However, the scan chain can be exploited by the hackers to retrieve the secret information in the cryptographic chip. This work proposes an obfuscation scheme of scan chain with test key. The scan chain is divided into segments by the multiplexers, and the linear feedback shift registers (LFSRs) are constructed by the scan cells in the chain. The selection signals of the multiplexers are controlled by the test key bits, and the two input signals of the multiplexers are the original scan chain data and the feedback data of the corresponding LFSR, respectively. The scan chain works normally if all the key bits are correct. Any wrong key bit leads to the flips on the selection signal of the inserted multiplexer periodically during the test mode, and the scan output is obfuscated. Analysis and simulation results demonstrate that the proposed scheme is resilient against the previous scan-based attacks, while maintaining the advantages of the scan testing,
Huixian Huang, Xiaole Cui, Xiaoxin Cui
ATS5
2022 An Area-Efficient and Robust Memristive LUT Based on the Enhanced Scouting Logic Cells
abstract
The resistive random access memory (RRAM) is a two-terminal device, which represents logic states with its different resistance states. The RRAM devices were applied to the Look-Up Table (LUT) in recent years. However, the RRAM based logic circuits are affected by the resistance variation of the RRAM devices. This work proposes a memristive LUT scheme based on the enhanced scouting logic (ESL) cells, to address this challenge. The read-write separation feature of the ESL cell is applied to reduce the number of working cycles of the proposed LUT circuit. The Monte Carlo simulation results show that the proposed LUT scheme has the small standard deviation. And the proposed LUT has relatively small area and high performance.
Xiaole Cui, Sunrui Zhang, Xiaoxin Cui
ISCAS4
2022 A 4-bit Integer-Only Neural Network Quantization Method Based on Shift Batch Normalization
abstract
Neural networks are powerful, but at the cost of huge amounts of computation. Deploying neural networks on edge devices is especially challenging. Quantization is a possible solution to alleviate the huge cost, while most quantization methods are not sufficiently hardware-friendly. In this paper, we proposed an integer-only quantization method. With no division or big integer multiplication, this quantization method is suitable to be deployed on co-designed hardware platforms. We applied 4-bit quantization on some classical networks and corresponding datasets. On MNIST, CIFAR10 and CFAR100, quantization networks perform as well as original networks. On SpeechCommands, accuracy error induced by quantization is 0.16%. We also deployed quantized networks under OpenCL framework and on a flash-based in-memory-computing chip to verify this method’s feasibility.
Qingyu Guo, Xiaoxin Cui, Aifei Zhang, Xinjie Guo, Yuan Wang 0001
ISCAS2
2022 An Event-driven Spiking Neural Network Accelerator with On-chip Sparse Weight
abstract
Spiking neural networks (SNNs) have widely drew attention of recent research. With brain-spired dynamics and spike-based communication, SNN is supposed to be a more energy-efficient neural network than existing artificial neural network (ANN). To make better use of the temporal sparsity of spikes and spatial sparsity of weights in SNN, this paper presents a sparse SNN accelerator. It adopts a novel self-adaptive spike compressing and decompressing (SASCD) mechanism for different input spike sparsity, as well as on-chip compressed weight storage and processing. We implement the octa-core design on field programmable gate array (FPGA). The results demonstrate a peak performance of 35.84 GSOPs/s, which is equivalent to 358.4 GSOPs/s in dense SNN accelerators for 90% weight sparsity. For the single-layer perceptron model in rate coding implemented on the hardware, SASCD reduces the time step intervals from 2.15 $\mu$ s to 0.55 $\mu$ s.
Yisong Kuang, Xiaoxin Cui, Chenglong Zou, Yi Zhong 0002, Zhenhui Dai, Zilin Wang 0001, Kefei Liu 0002, Dunshan Yu, Yuan Wang 0001
ISCAS2
2022 A 28nm 64Kb SRAM based Inference-Training Tri-Mode Computing-in-Memory Macro
abstract
Many computing-in-memory (CIM) macros achieve local inference with forward propagation (FP), and some CIM macros also support backward propagation (BP) computation. However, they can not calculate the weight change related to the learning rate and forward propagation input. these macros can not support backward propagation training algorithm completely. In this paper, we proposed a 28nm 64Kb SRAM based CIM macro, which supports a more complete backward propagation training algorithm. This macro supports three computing modes. A multiply unit (MU) supports FP and BP modes. A multiply circuit (MC) supports three-inputs-multiplication (TIM) mode for the weight change analog computing. MC uses the principle of charge sharing which has a high resistance to process variation and perfect linearity. In FP and BP modes, this macro achieves an energy efficiency of 42.1TOPS/W with 2-bit input, 8-bit weight and 14-bit output multiplication and accumulation operations (MAC). In TIM mode, this macro achieves an energy efficiency of 59.4 - 2222TOPS/W with multiplication of 3 inputs and 1 output.
Nanbing Pan, Xiaoxin Cui, Kanglin Xiao, Qingyu Guo, Yuan Wang 0001
ISCAS2
2022 A Computing-in-Memory SRAM Macro Based on Fully-Capacitive-Coupling With Hierarchical Capacity Attenuator for 4-b MAC Operation
abstract
In this work, we present a fully capacitive-coupling-based SRAM computing-in-memory (CIM) macro aimed at improving the energy efficiency and throughput of edge devices running multi-bit multiply-and-accumulate (MAC) operations. The proposed architecture is built around a customized 9T1C bit-cell in charge-domain computation in a 28nm technology. The proposed design supports 8192 $4{\mathrm{b}}\times 4{\mathrm{b}}$ MAC operations simultaneously. A 4-bit input is generated by DAC, while a 4-bit weight is achieved by a hierarchical capacity attenuator array without additional sharing switches, long sharing time, and complicated controlling signal. To minimize the expensive AD conversion, an input sparsity sensing scheme is proposed, allowing to skip redundant comparators. Access time is 4 ns with 0.9 V power supply at room temperature. The proposed design achieves energy efficiency of 666 TOPS/W and throughput of 4096 GOPS.
Kanglin Xiao, Xiaoxin Cui, Nanbing Pan, Xin'an Wang, Yuan Wang 0001
ISCAS2
2022 ESSA: Design of a Programmable Efficient Sparse Spiking Neural Network Accelerator
abstract
Spiking neural networks (SNNs) have been witnessing the developing trends to reduce the model size and improve the hardware efficiency for area- and energy-based applications, which are processed by model pruning and data compressions. However, it is challenging to exploit the unstructured sparsity of SNNs for the dense neuromorphic processors. In this article, we present an efficient sparse SNN accelerator (ESSA), which leverages both the temporal sparsity of spike events and the spatial sparsity of weights in SNN inference. It provides both the compressed weights for sparse SNNs and the uncompressed weights for compact SNNs. The self-adaptive spike compression is proposed for sparse spike scenarios, leading to the improvement of throughput by$3.2\times $. ESSA executes a flexible fan-in–fan-out tradeoff by using combinable dendrites, which overcomes the fan-in limitation in neuromorphic systems. Furthermore, a low-latency intrachip spike multicast method is adopted to reduce the resource overhead. Implemented on the Xilinx Kintex Ultrascale field-programmable gate array (FPGA), ESSA achieves an equivalent performance of 253.1 GSOP/s and an energy efficiency of 32.1 GSOP/W for 75% weight sparsity at 140 MHz. The implementation of a four-layer fully connected SNN is expected to perform$2.6~\mu \text{s}$per time step and the energy consumption is$14.6~\mu \text{J}$. Our results demonstrate that ESSA outperforms several state-of-the-art application-specific integrated circuit (ASIC) or FPGA neuromorphic processors.
Yisong Kuang, Xiaoxin Cui, Zilin Wang 0001, Chenglong Zou, Yi Zhong 0002, Kefei Liu 0002, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 The Modeling Attack and Security Enhancement of the XbarPUF with Both Column Swapping and XORing
abstract
To address the security challenge of integrated circuits, the Physical Unclonable Function (PUF) is of great concern as the root of trust. However, the PUF circuits are suffering from the modeling attacks in recent years. The design of anti-modeling-attack PUF is still an open issue. The XbarPUF with both column swapping and XORing was reported as an anti-modeling-attack PUF in 2017. This work proposes a two-step attack method. The first step transforms the target PUF into the XbarPUFs with XORing only, based on the column swapping states. The second step attacks the XbarPUFs with XORing only by the intentionally designed Artificial Neural Network (ANN) model. This method can predict the challenge response pairs (CRPs) of the XbarPUF with both column swapping and XORing successfully. To enhance the anti-modeling-attack capability, an improved XbarPUF is further proposed, which uses the dynamic column swapping technique. The results show that the prediction accuracy of attacks for the XbarPUF with both column swapping and XORing and the proposed XbarPUF reaches 98.96% and 70%, respectively, if the training CRPs account for one millionth of the total CRPs. The proposed XbarPUF has a better anti-attack capability than the XbarPUF with both column swapping and XORing.
Xiaole Cui, Wenqiang Ye, Xiaoxin Cui
ACM Great Lakes Symposium on VLSI4
2021 An SNN-Based and Neuromorphic-Hardware-Implementable Noise Filter with Self-adaptive Time Window for Event-Based Vision Sensor
abstract
Event-based dynamic vision sensors (DVS), inspired by biological vision systems, lead to new sensing and computing paradigms. The novel sensors output the sensed signal alone with many noise events asynchronously. Data-preprocessing for filtering these noises is significant before utilizing the data in applications such as classification, tracking and motion-data extraction. This paper describes a fully spike-based and neuromorphic-hardware-implementable neural network with a signal-oriented self-adaptive filtering time window for filtering the noise events robustly in the data captured by DVS. In particular, the simple leaky integrate-and-fire (LIF) neuron model is adopted as the basic elements of the network out of the purpose of hardware-friendly. Experiments based on both synthesized data and authentically-captured data are designed for quantitative comparison with traditional DVS noise filters to verify the outperformance of the proposed filter. The main contribution of this work is that the proposed spiking neural network (SNN) based filter achieves higher signal-noise-ratio (SNR) compared to traditional noise filters and performances more robust in the tolerance for changing signals.
Kanglin Xiao, Xiaoxin Cui, Kefei Liu 0002, Xiaole Cui, Xin'an Wang
IJCNN2
2021 A 28-nm 0.34-pJ/SOP Spike-Based Neuromorphic Processor for Efficient Artificial Neural Network Implementations
abstract
Neuromorphic hardware platforms inspired by human brain have emerged as novel non von Neumann computing architectures. They were proved excellent platforms for spiking neural network (SNN) implementations. However, implementing artificial neural networks (ANNs) on existing neuromorphic hardware platforms is still a daunting task because of critical limitations on coding scheme, maximum of fan-in, and highest weight precision in them. In this paper, we introduce a neuromorphic processor developed for various neural networks implementations including ANNs and SNNs. We employ spatio-temporal coding scheme based on spike events. By combining low-precision dendrites, the chip can implement weight precision between 1 bit and 8 bits and scalable fan-in. The 3.66-mm2chip fabricated in 28-nm CMOS with a maximum fan-in of 72 K per neuron demonstrates unprecedented compatibility with ANN applications compared to previously-proposed neuromorphic chips.
Yisong Kuang, Xiaoxin Cui, Yi Zhong 0002, Kefei Liu 0002, Chenglong Zou, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001
ISCAS2
2021 A Spike-Event-Based Neuromorphic Processor with Enhanced On-Chip STDP Learning in 28nm CMOS
abstract
Event-based spiking neural network (SNN) has displayed a promising prospect to realize real-time, efficient and intelligent hardware platforms. Whereas great efforts are still being appealed to explore the possibility of introducing online learning abilities to neuromorphic systems. In this paper, a 28-nm CMOS neuromorphic processor is presented, fulfilling online learning by adopting counter and lookup table (LUT) based spike-timing-dependent plasticity (STDP) rule. Designed to work at high-precision scenarios, the presented processor integrates up to 1024 neurons and 256K signed 9-bit synapses. It also ensures chip array interconnection to fit large neural networks. Moreover, by utilizing the sparse property of spike events to minimize activity rate, the typical power consumption is further reduced to 3.348mW for training MNIST dataset.
Yi Zhong 0002, Xiaoxin Cui, Yisong Kuang, Kefei Liu 0002, Yuan Wang 0001, Ru Huang 0001
ISCAS2
2021 The ANN Based Modeling Attack and Security Enhancement of the Double-layer PUF
abstract
The modeling attack is a serious threat to the physical unclonable function (PUF) circuits. The double-layer PUF was reported as a PUF scheme to resist the machine learning attacks, and its test chip was fabricated and tested. This work attacks the double-layer PUF successfully by an intentionally designed artificial neural network (ANN) model based on the working principle of the target PUF. To enhance the anti-modeling-attack capability of the double-layer PUF, the XORing and the dimensional extension techniques are proposed. The attack results show that the prediction accuracy of the proposed ANN-based model with the XORing and 3D extension techniques is as low as 50.09% in average. It manifests that the proposed security enhancement techniques are able to improve the resilience of the double-layer PUF against the modeling attacks effectively.
Xiaole Cui, Wenqiang Ye, Xiaoxin Cui
ITC-Asia4
2021 The Security Enhancement Techniques of the Double-layer PUF Against the ANN-based Modeling Attack
abstract
The physical unclonable function (PUF) against the modeling attack is of great concern in recent years, since the modeling attack has been proved to be a serious security threat to the PUF circuits. The double-layer PUF was reported as a PUF scheme to resist the fully connected artificial neural network based modeling attack, and its test chip was fabricated and tested. This work proposes an artificial neural network (ANN) based modeling method according to the working principle of the target PUF, and successfully attacks the double-layer PUF. To enhance the anti-modeling-attack capability of the double-layer PUF, the address swapping, the XORing, and the dimensional extension techniques are proposed. The attack results show that the prediction accuracy of the proposed ANN-based model with the proposed techniques drops obviously. And the prediction accuracy is about 50.04% if all the three proposed techniques are applied in combination. It manifests that the proposed security enhancement techniques are able to improve the resilience of the double-layer PUF against the modeling attacks effectively. Both the randomness and uniqueness of the improved doublelayer PUFs are approximate to the ideal value (50%), and the reliability of the improved PUFs remain unchanged compared with the original counterpart because the operations on the resistive random memory (RRAM) array are the same.
Xiaole Cui, Wenqiang Ye, Xiaoxin Cui
ITC4
2021 Machine Learning Aided Key-Guessing Attack Paradigm Against Logic Block Encryption
abstract
Hardware security remains as a major concern in the circuit design ow. Logic block based encryption has been widely adopted as a simple but effective protection method. In this paper, the potential threat arising from the rapidly developing field, i.e., machine learning, is researched. To illustrate the challenge, this work presents a standard attack paradigm, in which a three-layer neural network and a naive Bayes classifier are utilized to exemplify the key-guessing attack on logic encryption. Backed with validation results obtained from both combinational and sequential benchmarks, the presented attack scheme can specifically accelerate the decryption process of partial keys, which may serve as a new perspective to reveal the potential vulnerability for current anti-attack designs.
Yi Zhong 0002, Jianhua Feng, Xiaoxin Cui, Xiaole Cui
J. Comput. Sci. Technol.3
2020 A Testability Enhancement Method for the Memristor Ratioed Logic Circuits
abstract
The Resistive Random Access Memory (RRAM) is a two-terminal variable resistance device, and the memristor ratioed logic (MRL) is a hybrid RRAM-CMOS style of logic circuit. The MRL AND and OR gates are implemented by the RRAM devices, and the MRL NOT gate is implemented by the CMOS inverter. However, the MRL circuits are prone to test escape. This work proposes a method to improve the testability of MRL circuits by replacing the CMOS inverters with the FinFET inverters. The test escape problem is solved by adjusting the switching threshold voltages of the FinFET inverters, and it is implemented by selecting the FinFET inverters in different operation modes. In Addition, some equivalent relationships in the fault set of the improved MRL gates are discovered. These fault equivalences of MRL gates result in a fault collapse ratio of about 50%. The test patterns for the production test of the improved MRL circuits can be generated by the traditional ATPG method. The test results of some typical MRL circuits obtained from the commercial ATPG tool show that the proposed method is able to achieve 100% fault coverage and at least 55% fault collapse ratio.
Li Qu, Xiaole Cui, Xiaoxin Cui
ATS3
2020 A Novel Conversion Method for Spiking Neural Network using Median Quantization
abstract
Artificial Neural Networks (ANNs) have achieved great success in the field of computer vision and language understanding. However, it is difficult to deploy these deep learning models on mobile devices because of its massive energy consumption and memory occupation. For another way, highly inspired from biological brain, spiking neural networks (SNNs), are often referred to as the 3-th generation of neural network for its potential superiority in cognitive learning and energy efficiency. Nevertheless, training a deep SNN remains a big challenge. In this paper, we propose a quantized training algorithm for ANNs to minimize spike approximation error, and provide two (temporally or spatially) rate-based conversion methods for SNNs, both of which can be easily mapped to specific neuromorphic platforms. Besides, this novel method can be generalized to various network architectures and adapted to dynamic quantization demand. Experimental results on MNIST and CIFAR-10 dataset demonstrate that the proposed deep spiking neural networks yield the state-of-the-art classification accuracy and need much less operations compared with their ANN counterparts. Our source code will be available upon request for the academic purpose.
Chenglong Zou, Xiaoxin Cui, Jiexian Ge, Hanghang Ma, Xin'an Wang
ISCAS2
2020 A synthesis method for logic circuits in RRAM arrays
Xiaole Cui, Xiaoxin Cui
Sci. China Inf. Sci.4
2020 The synthesis method of logic circuits based on the iMemComp gates
Xiaole Cui, Qiujun Lin, Xiaoxin Cui, Jinfeng Kang
Integr.3
2018 Polymorphic gate based IC watermarking techniques
abstract
Polymorphic gates are reconfigurable devices whose functionality may vary in response to the change of execution environment such as temperature, supply voltage or external control signals. This feature makes them a perfect candidate for circuit watermarking. However, polymorphic gates are hard to find because they do not exhibit the traditional structure. In this paper, we report four dual-function polymorphic gates that we have discovered using an evolutionary approach. With these gates, we propose a circuit watermarking scheme that selectively replaces certain standard logic gates with the polymorphic gates. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that this scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.10%, 2.08% and 1.17% respectively when we embed 30-bit watermark sequences. These overheads increase to 6.36%, 4.75% and 2.08% respectively when 10% of the gates in the original circuits are replaced to embed watermark up to more than 300 bits.
Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui
ASP-DAC2
2018 A Novel Polymorphic Gate Based Circuit Fingerprinting Technique
abstract
Polymorphic gates are reconfigurable devices that deliver multiple functionalities at different temperature, supply voltage or external inputs. Capable of working in different modes, polymorphic gate is a promising candidate for embedding secret information such as fingerprints. In this paper we report five polymorphic gates whose functionality varies in response to specific control input and propose a circuit fingerprinting scheme based on these gates. The scheme selectively replaces standard logic cells by polymorphic gates whose functionality differs with the standard cells only on Satisfiability Don't Care conditions. Additional dummy fingerprint bits are also introduced to enhance the fingerprint's robustness against attacks such as fingerprint removal and modification. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that our scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.04%, 6.97% and 4.15% respectively when we embed 64-bit fingerprint that consists of 32 real fingerprint bits and 32 dummy bits. This is only half of the overhead of the other known approach when they create 32-bit fingerprints.
Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui
ACM Great Lakes Symposium on VLSI2
2018 Evaluation of Dynamic-Adjusting Threshold-Voltage Scheme for Low-Power FinFET Circuits
Xiaoxin Cui, Yewen Ni, Dunshan Yu, Xiaole Cui
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A Heuristic Algorithm for Automatic Generation of March Tests
abstract
March test is one of the most popular memory test algorithms for its good fault coverage and linear complexity. However, designing an efficient March test for a complex memory fault set is a tedious task. This work designs a two-phase heuristic algorithm for automatic generation of March tests with two new observations. One observation is that some operations can be deleted from the march element, if other operations are inserted to sensitize the Loop-Sens fault, wherein the required initial state is equal to the expected faulty state. The other observation is that there exist chances to reduce the length of march sequence after some operation segments between march elements are exchanged, if the expected states of the operation segments in the different march elements match each other. Experimental results show that the proposed algorithm is effective and efficient. The generated March tests cover the target fault sets, and the complexities approach to those of the minimal March tests generated from the exhaustive methods for the specific fault sets.
Xiaole Cui, Yichi Luo, Qiujun Lin, Xiaoxin Cui
ATS4
2017 Testing of 1TnR RRAM array with sneak path technique
Xiaole Cui, Xiaoxin Cui, Xin'an Wang, Jinfeng Kang
Sci. China Inf. Sci.3
2017 Improving DFA attacks on AES with unknown and random faults
Nan Liao, Xiaoxin Cui, Dunshan Yu, Xiaole Cui
Sci. China Inf. Sci.2
2017 An Enhancement of Crosstalk Avoidance Code Based on Fibonacci Numeral System for Through Silicon Vias
abstract
Through silicon vias (TSVs) play an important role as the vertical electrical connections in 3-D stacked integrated circuits. However, the closely clustered TSVs suffer from the crosstalk noise between the neighboring TSVs, and result in the extra delay and the deterioration of signal integrity. For a 3 × 3 TSV array, the severity of crosstalk noise in the center victim TSV is classified into 11 levels, which is defined as 0C to 10C from low noise to high noise, depending on the combinations of the digital patterns applied to the TSV array. An enhanced code based on the Fibonacci number system (FNS) to suppress the crosstalk noise below 6C level is proposed, in which both the redundancy of numbers and the nonuniqueness of Fibonacci-based binary codeword are utilized to search the proper codeword. Experimental results show that the proposed technique decreases about 22% latency of TSVs comparing with the worst crosstalk cases. This technique is applicable in the large-scale TSV array for it has a quasi-linear hardware overhead, and its system overhead is less than that of the 3-D 4-LAT counterpart if the data width is greater than 18, and it has good usability for it consumes less power per TSV and achieves lower bit error rate at the interested frequency range comparing with that of the original FNS coding technique.
Xiaole Cui, Xiaoxin Cui, Yewen Ni, Min Miao, Yufeng Jin
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Ultralow-power high-speed flip-flop based on multimode FinFETs
Xiaoxin Cui, Nan Liao, Dunshan Yu, Xiaole Cui
Sci. China Inf. Sci.2
2015 Key characterization factors of accurate power modeling for FinFET circuits
Kaisheng Ma, Xiaoxin Cui, Nan Liao, Dunshan Yu
Sci. China Inf. Sci.2
2014 High-speed constant-time division module for Elliptic Curve Cryptography based on GF(2m)
abstract
To achieve high performance scalar multiplication arithmetic in Elliptic Curve Cryptography (ECC) based on GF(2m), a high-speed constant-time division module with optimized architecture is proposed in this paper. Modified from the traditional extended Euclidean Great Common Divisor (GCD) division algorithm, the presented algorithm computes a single multiplicative inverse or division in constant m iterations, i.e. m clock cycles, in GF(2m), which obtains a tremendous reduction (specifically more than 50%) on computing time compared with previous works. Combined with the meticulously optimized architecture, this novel division module achieves lower area-time complexity, which makes it an excellent option for high performance ECC design.
Xiaoxin Cui, Nan Liao, Dunshan Yu
ISCAS2
2014 Low power adiabatic logic based on FinFETs
Nan Liao, Xiaoxin Cui, Kaisheng Ma, Dunshan Yu
Sci. China Inf. Sci.2
2014 Ultra-low power dissipation of improved complementary pass-transistor adiabatic logic circuits based on FinFETs
Xiaoxin Cui, Nan Liao, Kaisheng Ma, Dunshan Yu
Sci. China Inf. Sci.2
2013 A Dynamic-Adjusting Threshold-Voltage Scheme for FinFETs low power designs
abstract
In this paper, a novel device/circuit co-design scheme, namely Dynamic-Adjusting Threshold-Voltage Scheme (DATS) for independent-gate mode FinFET circuits has been proposed. The main idea of this scheme is that a pair of back-gate bias of FinFETs is adjusted dynamically to change threshold voltage according to the system operating frequency and operating mode, which could optimize circuit power, especially leakage power. The experimental and simulation result shows that the leakage power dissipation reduced greatly when circuits operate at the lower frequency, and the energy-delay product of FinFET circuits is reduced by 30% approximately.
Xiaoxin Cui, Kaisheng Ma, Nan Liao, Dunshan Yu
ISCAS1