VLDB 2026 Research / reviewers in the wild / expert
Xiaole Cui
dblp:89/2968
· DBLP profile ↗
48ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-3382-3703ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Elmore Delay Model Based Test Method for Post-Bond TSVs
Zhicheng Shao, Xiaole Cui, Xiaoxin Cui |
ETS | 2 |
| 2026 | An On-Line BISD Triggering Scheme Based on the CAC-Compatible Error Detection Code for TSV ArrayabstractThe through silicon via (TSV) array plays the significant role of vertical electrical interconnection in the three-dimensional stacked integrated circuits. The TSV faults and the inter-TSV crosstalk are two issues that must be addressed for the effective and reliable transmission in TSV arrays. The forbidden transition free (FTF) and forbidden pattern free (FPF) crosstalk avoidance code (CAC) can suppress the crosstalk in TSV array, and the built-in self-diagnosis (BISD) and self-repair techniques are able to localize and repair the TSV faults, respectively. The BISD usually consumes a non-negligible amount of time in the diagnosis mode. So the periodic BISD schemes have a dilemma between the transmission performance and the efficiency of fault detection. The on-demand BISD schemes have been applied to address this issue. This paper proposes a new on-line BISD triggering scheme based on the CAC-compatible error detection code (EDC). The proposed scheme checks the data transmitted through the TSV array in the working mode. And the BISD is invoked in time only if the EDC codecs detect errors. Compared with the existing similar schemes, the proposed scheme does not degrade the crosstalk suppression effect of the FTF/FPF-CAC methods, and it achieves a better balance in terms of the fault detection ability and the hardware overhead. Xiaole Cui, Xing Zhang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | A Joint Data Transmission Scheme With Improved CAC for 3D-SICs Based on Multiple TSV Subarrays
Xiaole Cui, Juncheng Pu, Zhaohong Lin, Xing Zhang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | MarchGen: A March Sequence Generation Method for Faults With an Arbitrary Number of Operations in RAMsabstractThe March test is a widely-applied test method for memory arrays. Nowadays, the advanced memory technologies keep introducing new types of defects, which bring diverse multi-operation fault models in turn. The automatic March sequence generation methods are preferred to avoid the inefficient manual design. Most previous methods generate the March sequences that cover limited types of faults and strongly depend on the specific characteristics of known faults. To address this issue, this work proposes MarchGen, a March sequence generation method for memory faults with arbitrary number of operations. Firstly, the fault hierarchies, reduced taxonomy and generic test conditions of 2-composite faults are analyzed. Then, the heuristic March sequence generation workflow is designed accordingly, which takes either fault models or test sequences as input. Three optimization strategies are utilized to improve the fault coverage of the generated March sequences and the generation time consumption. The evaluation results show that the proposed method has low dependence on fault types. The generated March sequences reach 100% fault coverage for any simple and 2-composite faults or test sequences, and the generation time is roughly linear with the size of input set. Sunrui Zhang, Xiaole Cui, Huixian Huang, Xing Zhang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | A Reconfigurable Built-In Self-Test Scheme for the Evaluation Circuits of Digital SRAM-IMC ArchitecturesabstractDigital static random access memory-based in-memory computing (SRAM-IMC) is a promising computation paradigm to break the von-Neumann bottleneck. However, the IMC architectures also bring a series of challenges for testing, because of the circuit structures and operations that do not exist in the conventional memories. One of the challenges is the testing of evaluation circuits in the digital SRAM-IMC architectures, because the primary inputs (PIs) of the evaluation circuits cannot be directly accessed by the testers. Several test approaches such as the conventional logic built-in self-test (LBIST) modules, the indirect and the scan-chain-based test methods are proposed to address this issue. Nevertheless, these solutions suffer from the low test performance or the high area consumption. This work proposes a reconfigurable built-in self-test (BIST) scheme for the evaluation circuits. By reusing the IMC bitcells and operations, the proposed BIST scheme implements the separate pattern generation (PG) and response analysis (RA) processes. Furthermore, the diverse pattern generators, including the Fibonacci linear feedback shift register (LFSR) and weighted LFSR (WLFSR) with adjustable feedback polynomials and the cellular automata (CA), are realized to improve the test efficiency and fault coverage. The evaluation results show that the proposed BIST scheme has better test performance comparing with the indirect and the scan-chain-based test approaches. The proposed BIST scheme has comparable test performance, whereas it has much less area overhead comparing with the conventional LBIST schemes. Additionally, the proposed BIST scheme is testable and repairable. Sunrui Zhang, Xiaole Cui, Xing Zhang 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | The Low-Latency UCIe-Compliant Spiral Layouts of TSV Arrays Against the Clustered FaultsabstractThe chiplet based integrated system contains lots of inter-die interconnections. The through silicon vias (TSVs) and micro bumps are the common components of the vertical electrical interconnections in the $2.5 \mathrm{D} / 3 \mathrm{D}$ multi-die integration circuits. They are usually laid out as arrays to regulate the routing rules. Faults may occur in the TSVs, and the redundant resources based self-repair mechanisms are introduced to improve the yield of the entire integrated chip. In the Universal Chiplet Interconnect Express (UCIe) standard, each advanced package module contains 32 data lanes and 2 redundant lanes, and the shift chain based intra-module self-repair mechanism is defined. However, the faults are usually clustered in certain region of the TSV array. The intra-module self-repair mechanism of UCIe standard is not able to cope with the cases that more than 2 clustered faults occur simultaneously in one module. An effective idea to improve the repair rate is to distribute the faults into different modules by regrouping the layout of TSV array. This work reports some new spiral layouts of rectangular and hexagonal TSV arrays. The self-repair schemes based on the proposed layouts are UCIe-compliant. Simulation results present that the proposed self-repair schemes are able to repair the clustered faults in TSV array with the low repair latencies. Bowen Tan, Xiaole Cui, Zhaohong Lin |
ATS | 2 |
| 2025 | A UCIe Compatible Repair Scheme for the Clustered Faults in the Hexagonal Array of Interconnect LanesabstractThe chiplet-based 2.5D/3D multi-die integrated chips have numerous inter-die interconnect lanes, which may suffer from the faults in Through Silicon Vias (TSVs) or micro bumps. The faults are usually clustered in some regions. The redundant lanes are introduced for repairing, to maintain the product yield. The hexagonal array of lanes offers higher interconnect density than the rectangular array, but it requires more redundant lanes if faults are located within the same module. In the Universal Chiplet Interconnect Express (UCIe) standard, each Advanced Package module contains 32 data lanes and 2 redundant lanes, which limit the repair rate for the clustered faults. This work proposes a UCIe-compatible repair scheme which increases the probability of distributing faulty lanes in different modules. It achieves a higher repair rate for the clustered faults with a low proportion of redundant lanes. Simulation results show that for 8 faults in various cluster windows, the proposed scheme achieves a repair rate above 98% while consuming less area than previous approaches. Zhaohong Lin, Xiaole Cui, Juncheng Pu |
ETS | 2 |
| 2025 | In-Memory Logic Synthesis Methods based on Read Decoupled 8T and Transpose 11T SRAMsabstractIn-memory computing (IMC) is regarded as the promising computer architecture to break through the von-Neumann bottleneck. Some in-memory logic operations, such as AND/NAND/OR/NOR/XOR, have been proposed to be implemented within SRAM array. However, the studies on the SRAM-based in-memory logic operations still focus on the gate-level. The implementation method of arbitrary logic functions is an open issue till now. In order to address this problem, this paper proposes the general in-memory logic synthesis methods based on the read decoupled 8T SRAM array and the transpose 11T SRAM array. The proposed synthesis methods are applied to some MCNC benchmark circuits. The synthesis results show that the in-memory logic circuit generated by the 8T SRAM array consumes smaller sub-array size, and the in-memory logic circuit generated by the 11T SRAM array consumes less compute cycles. Xiaole Cui, Sunrui Zhang |
ISCAS | 2 |
| 2025 | AgentFactory: Towards Automated Agentic System Design and Optimization
Enci Zhang, Yuesheng Zhu, Xiaole Cui, Guibo Luo |
PRICAI | 4 |
| 2025 | A Network-Coding-Based Functional Test Method for Routers in 3-D Mesh NoCsabstractThe 3-D Mesh Network-on-Chip (NoC) is widely used as the communication backbone of the 3-D multiprocessor system-on-chips (MPSoCs) and chiplet-based 3-D integrated systems, which enable the efficient computation paradigm in Internet of Things (IoT) scenarios. The functional test is a good choice for the 3-D Mesh NoCs to detect the defects introduced in manufacture process, because it reuses the at-speed on-chip network and consumes little hardware overhead. However, the test time of the functional test in 3-D Mesh NoCs grows dramatically with the NoC size. In order to reduce the test time on the premise of ensuring fault coverage, this work proposes a functional test method leveraging the exclusive OR (xor) network coding (NC) technique. In this work, thexor NC function is embedded in the routers by the extension modules with an additional area overhead less than 3%, and an efficient functional test process based on the embeddedxor NC function is proposed. The proposed functional test process covers all the straight and turning paths, and all the working modes of the routers are tested. The time complexity of the proposed method only increases linearly with the NoC size, which manifests that the proposed test method is applicable for the large-scale 3-D Mesh NoCs. Xiaole Cui |
IEEE Internet Things J. | 2 |
| 2025 | A VC Dimension-Oriented Improvement Method of PUFs for the Anti-Modeling-Attack CapabilityabstractThe physical unclonable function (PUF) serves as a security primitive of circuits, which is applicable to the embedded systems with lightweight authentication function. However, the modeling attack, which estimates the unknown CRPs by establishing the mathematical model of PUF, is a real threat to the PUF based crypto-systems. Subsequently, the anti-modeling-attack PUF becomes a research hotspot. The systematic design method of secure PUF is still an open issue, although some secure PUF schemes have been proposed based on the repeated trials. This work proposes a security improvement method of PUFs to enhance the anti-modeling-attack capability. The growth function and the Vapnik-Chervonenkis (VC) dimension of PUF are defined as the indicators of PUF security. The proposed method regards the improvement of PUF as an optimization problem, which aims to obtain a PUF scheme with the better security indicators. Guided by the indicators, the proposed method is able to specify the improvement sites of PUF and the techniques to be applied. In addition, three approaches are proposed to inspire the new security improvement techniques. An improved arbiter PUF and an improved array-based PUF are designed as the instances of the results from the proposed method. Both of the improved PUF schemes have the stronger security than the original schemes. Xiaole Cui, Sunrui Zhang, Xiaoxin Cui |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2025 | FNS-CATF-CAC: An Efficient Crosstalk Avoidance Code to Reduce the Switching Activity in TSV ArraysabstractThe through silicon via (TSV) arrays play the role of vertical electrical interconnections in the 3-D stacked integrated circuits. However, the coupling crosstalk between the adjacent TSVs increases the interconnection delay and deteriorates the signal integrity in TSV arrays. The crosstalk avoidance code (CAC) techniques based on the Fibonacci numeral system (FNS) or the improved FNS are capable of mitigating the crosstalk in TSV arrays, but the existing schemes are hindered by the hardware overhead, crosstalk suppression ability and switching activity. This article proposes the FNS-based cyclic adjacent transition free CAC with the ouroboros mapping rule for the rectangular and hexagonal TSV arrays. The proposed scheme can reduce the crosstalk even in the presence of the edge effect. Compared with the previous methods, the proposed scheme consumes significantly small hardware overhead in large-scale arrays. And the proposed method can reduce the switching activity on TSVs, thereby alleviating the power consumption in TSV arrays. Xiaole Cui |
IEEE Trans. Reliab. | 2 |
| 2024 | A Convolutional Spiking Neural Network Accelerator with the Sparsity-Aware Memory and Compressed WeightsabstractThe spiking neural network (SNN) has advantage in the edge AI applications for its spatiotemporal sparsity. The high energy efficiency is an important concern in the study of SNN accelerator designs. In this paper, a lightweight event-driven convolutional SNN accelerator that utilizes the sparsity of both the spike events and the network weights is proposed. In the event-driven mode, the proposed accelerator uses the compressed input spikes and a spike-oriented convolution data flow. An output spike compressor is also designed. To balance the computation performance and the memory space occupancy, a spike sparsity-aware memory scheme that automatically switches the spike format by a real-time monitoring strategy is designed. The compression memories and a buffer for network weights are designed to save the on-chip memory space. The accelerator prototype is verified on the Xilinx Virtex XCVU9P FPGA platform. It achieves an equivalent performance of 139.5GFLOPS on the N-MNIST dataset. Compared to the baseline using the same computational resources, the proposed accelerator can improve the inference performance, the inference energy efficiency and the memory space by 4.6, 3.6 and 1.6 times, respectively. The proposed accelerator has advantages in energy efficiency and hardware overhead compared to the previous works on the same hardware platform. Neuralmorphic computing Spiking neural network accelerator Sparse spikes Sparse matrix compression Field-programmable gate array Xiaole Cui, Sunrui Zhang, Mingqi Yin, Xiaoxin Cui |
ASAP | 2 |
| 2024 | A Testability Improvement Method of Combinational Circuits Based on the SDC ConditionsabstractThe testability of circuit decreases with the rapid increase of integration and complexity of modern chips. The synthesis technique for testability is important for the design of testable circuit, which reduces the test cost and achieves the high fault coverage. This work improves the testability of combinational circuits leveraging the Satisfiability Don't Care (SDC) conditions. The SDC condition refers the case that some input combinations of a circuit block can never occur because of the correlation between some input signals. The redundancy exists if the SDC condition is satisfied. In the proposed method, the adjacent-level SDC conditions and the cross-level SDC conditions are identified, and the related circuit blocks are simplified to improve the circuit testability. The Sandia Controllability and Observability Analysis Program (SCOAP) measures of benchmark circuits decrease if the proposed method is applied to the corresponding improved circuits. However, the reduction of the SCOAP measures is limited because the SDC conditions are rare in most of the benchmark circuits. To further reduce the SCOAP measures, the Observability Don't Care (ODC) condition introduced SDC blocks in the benchmark circuits are identified. The ODC condition refers the case that connection of a signal wire does not affect the circuit's logic function. More redundant blocks are found and simplified by utilizing the ODC introduced SDC blocks in the benchmark circuits. The Evaluation results show that average decrease of CC0 and CC1 is 14%, and the average decrease of CO is 15.3%. Additionally, it is observed that the proposed method is more effective for the larger scale circuits. Xiaole Cui |
ATS | 2 |
| 2024 | Modeling Attack Tests and Security Enhancement of the Sub-Threshold Voltage Divider Array PUFabstractPhysical unclonable function (PUF) is widely used as the root of trust in the IoT systems. The sub-threshold voltage divider array PUF was reported as an anti-modeling-attack PUF. It utilizes the nonlinear I-V relationship in the sub-threshold region of MOS transistors to improve the security. However, the security of this PUF has not been soundly analyzed. This work presents the simulation results of modeling attack tests which reveal the vulnerability of this PUF. In the attack, the sub-threshold voltage divider array PUF is modeled by a dedicated artificial neural network (ANN). The nonlinearity of the PUF is simplified based on its working principle. The simulation results show that the prediction accuracy achieves 97% when the number of training CRPs is 350 for the single-stage PUF, and it achieves 90% when the number of training CRPs is 300 for the multi-stage PUF. Furthermore, this work improves the original structure of the sub-threshold voltage divider array PUF, to enhance its anti-modeling-attack capability. The simulation results of modeling attack tests show that the prediction accuracy of the improved multi-stage PUF is reduced to about 50%, which implies that the improved PUF has a strong anti-modeling-attack capability. Xiaole Cui |
DATE | 3 |
| 2024 | SPAT: FPGA-based Sparsity-Optimized Spiking Neural Network Training Accelerator with Temporal Parallel DataflowabstractSpiking neural networks (SNNs), as biologically inspired computational models, possess significant advantages in energy efficiency due to their event-driven operations. However, challenges remain in attaining high computational efficiency for SNN training. In this work, we propose a novel SNN training accelerator employing temporal parallelism and sparsity optimizations to achieve superior efficiency. A temporal parallel dataflow is designed to concurrently integrate spikes across multiple time steps, enhancing throughput and data reuse. To reduce latency and improve energy efficiency, we leverage the sparsity of SNNs and employ methods such as zero gating and zero skipping. Implemented on a field-programmable gate array (FPGA), the proposed training accelerator demonstrates 2.3-fold speedup and 15.7-fold energy reduction compared to NVIDIA A100 GPU on N-MNIST dataset. Li Lun, Mingqi Yin, Zhenhui Dai, Xiaole Cui, Xiaoxin Cui |
ISCAS | 7 |
| 2024 | A High Performance PODEM Algorithm with the Improved Backtrace ProcessabstractThe PODEM (Path-Oriented Decision Making) algorithm is one of the classical path sensitive ATPG algorithms, which contains the backtrace process. Recent studies have shown that, in addition to the traditional heuristics, such as the SCOAP (Sandia Controllability and Observability Analysis Program) and COP (Controllability / Observability Procedure) measures, the ANN (Artificial neural network) can also guide the backtrace process in the PODEM algorithm. These methods are able to reduce the number of backtraces and the CPU time compared with the PODEM algorithms with the traditional heuristics. However, it may results in the negative improvements for some cases, due to the inherent errors in the output of ANN. To address this issue, an improved PODEM algorithm with the backtrace process guided by the incorporated ANN and traditional heuristics is proposed in this work. The effects of ANN hyper-parameters on the backtrace process of PODEM algorithm are studied, and the threshold which distinguishes the positive and negative improvements is discovered. The basic idea of the proposed PODEM algorithm is that, the traditional heuristics is applied to guide the backtrace process for each gate of the backtrace paths if the corresponding difference value of gate inputs in the ANN model is less than the threshold; and the ANN-based heuristics is applied vice versa. Evaluation results show that the proposed PODEM algorithm results in more positive improvements than the PODEM algorithm with the backtrace process guided by the ANN or the traditional heuristics alone. Furthermore, the proposed PODEM algorithm is better to be applied to the large size circuit, according to the application results on the benchmark circuits with different sizes. Xiaole Cui |
ITC-Asia | 2 |
| 2024 | The Resistance Analysis Attack and Security Enhancement of the IMC LUT Based on the Complementary Resistive Switch CellsabstractThe resistive random access memory (RRAM) based in-memory computing (IMC) is an emerging architecture to address the challenge of the “memory wall” problem. The complementary resistive switch (CRS) cell connects two bipolar RRAM elements anti-serially to reduce the sneak current in the crossbar array. The CRS array is a generic computing platform, for the arbitrary logic functions can be implemented in it. The IMC CRS LUT consumes fewer CRS cells than the static CRS LUT. The CRS array has built-in polymorphic characteristics because the correct logic function cannot be distinguished based on the circuit layout. However, the logic state of every CRS cell can be readout after each operation. It helps the attacker to recover the correct function of the IMC CRS LUT. This work discusses the resistance analysis attack of the IMC LUT based on the CRS array. The proposed resistance analysis attack method is able to be applied to different computation styles based on the CRS array, such as the CRS IMPLY, CRS NOR-OR/NAND-AND, and so on. The attacker can recover the logic function of the LUT by tracing the states of CRS cells. Furthermore, an improved IMC CRS LUT method is proposed and discussed to enhance security. The simulation and analysis results show that the improved IMC CRS LUT can resist various attacks, and it maintains the polymorphic characteristics of the IMC CRS LUT. And the N-bit full adder circuit based on the improved IMC CRS NOR-OR LUTs achieves the best performance compared with the previous counterparts. Xiaole Cui, Mingqi Yin, Xiaoxin Cui |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | Dy-MFNS-CAC: An Encoding Mechanism to Suppress the Crosstalk and Repair the Hard Faults in Rectangular TSV ArraysabstractThrough-silicon vias (TSVs) play the role of vertical electrical interconnections in the emerging three-dimensional stacked integrated circuits. However, the hard faults and the crosstalk faults may occur in the TSV array simultaneously. The hard faults are the catastrophic faults that lead to functional failure, and the crosstalk faults deteriorate the signal integrity of TSV array. The previous fault tolerant techniques for this problem are based on the Fibonacci numeral system-based crosstalk avoidance code (FNS-CAC), and the bit overhead and the reparability need to be further improved. This article proposes the dynamic modified Fibonacci numeral system (Dy-MFNS) and the Dy-MFNS based crosstalk avoidance code (Dy-MFNS-CAC). The cascaded Dy-MFNS adders and the Dy-MFNS codec are designed to generate the Dy-MFNS-CAC codewords according to the health status of TSVs. The generated codewords are able to suppress the crosstalk and repair the hard faults simultaneously, under the control of the TSV fault flags. The simulation results show that the proposed scheme has advantages on both the bit overhead and the reparability compared with the FNS-CAC based techniques. Xiaole Cui, Xiaoxin Cui |
IEEE Trans. Reliab. | 2 |
| 2023 | An Anti-Removal-Attack Hardware Watermarking Method Based on Polymorphic GatesabstractWatermarking is an effective way to protect the intellectual properties (IPs) of hardware. The polymorphic gate based watermarking technique was recently proposed where certain standard logic gates are replaced by polymorphic gates to embed watermarks. However, the special structure of the polymorphic gates makes them distinguishable from the standard logic gates. It enables the attacker to discover the watermarks after reverse engineering, and to remove them by replacing the polymorphic gates with the functional equivalent standard cells. The proposed polymorphic watermarking method enhances the hardware watermarks against the removal attacks by reducing the IP's quality once the watermark is removed. To reach this goal, the specific Satisfiability Don't Care (SDC) conditions in the netlist are identified, and they are utilized to define the polymorphism that results in more cost on delay, power and area after the removal attack. Additionally, an observability don't care (ODC) based technique is introduced to increase the number of gates with the required SDC conditions to accommodate the long watermark bits. Furthermore, the trap gate technique is introduced to avoid the polymorphic gates being replaced back with the original gates. The simulation results on ISCAS‘85 and MCNC benchmark circuits show that the average overhead in circuit delay, area and power of the proposed method are only 3.97%, 4.75% and 3.26% respectively to embed 128-bit watermarks compared with the original circuits, but become 92.80%, 70.32% and 53.55%, respectively after removal attacks. Xiaole Cui, Pengyuan Yang, Gang Qu 0001 |
ICCAD | 2 |
| 2023 | An Area-Efficient In-Memory Implementation Method of Arbitrary Boolean Function Based on SRAM ArrayabstractIn-memory computing is an emerging computing paradigm to breakthrough the von-Neumann bottleneck. The SRAM based in-memory computing (SRAM-IMC) attracts great concerns from industries and academia, because the SRAM is technology compatible with the widely-used MOS devices. The digital SRAM-IMC scheme has advantages on stability and accuracy of computing results, compared with the analog SRAM-IMC schemes. However, few logic operations can be implemented by the current digital SRAM-IMC architectures. Designers have to insert some special logic modules to facilitate the complex computation. To address this issue, this work proposes an area-efficient implementation method of arbitrary Boolean function in SRAM array. Firstly, a two-input SRAM LUT is designed to realize the arbitrary two-input Boolean functions. Then, the logic merging and the spatial merging techniques are proposed to reduce the area consumption of the SRAM-IMC scheme. Finally, the SOP-based SRAM-IMC architecture is proposed, and the merged SOPs are mapped into and computed in it. The evaluation results on LGsynth’91, IWLS’93 and EPFL benchmarks show that, the area of the synthesis results based on the ABC tool is 3.69, 5.72 and 1.86 times of the circuit area from the proposed SRAM-IMC scheme in average respectively. Furthermore, the circuit area from the original SOP-based SRAM-IMC scheme is 2.07, 1.99 and 1.86 times in average of the circuit area from the proposed SRAM-IMC scheme respectively. The performance evaluation results show that the cycle consumption of the proposed SRAM-IMC scheme is independent to the scale of the input Boolean functions. Sunrui Zhang, Xiaole Cui, Xiaoxin Cui |
IEEE Trans. Computers | 2 |
| 2023 | Mosaic-3C1S: A Low Overhead Crosstalk Suppression Scheme for Rectangular TSV ArrayabstractThe through silicon via (TSV) is one of the important enabling technologies of stacked 3-D ICs. However, the crosstalk is generated when signals are transmitted through the closely clustered TSVs due to the coupling effect. The crosstalk avoidance code (CAC) techniques have attracted great concerns in recent years, for the TSV-to-TSV crosstalk deteriorates the signal integrity. Nevertheless, the CAC techniques usually require specific shape of the TSV array, e.g.,$3\times N$array. This work proposes a crosstalk suppression scheme combining the CAC method and the static shielding technique, named Mosaic-3C1S. The proposed scheme is able to be applied to arbitrary rectangular TSV arrays. In the proposed scheme, the$2\times 2$TSV subarray, which contains three signal TSVs and one shielding TSV, is used as the basic unit to tile the rectangular TSV array. The CAC method is applied to the signal TSVs in the$2\times 2$subarrays, and the raw data are transmitted by the rest boundary TSVs, if any. In order to further reduce the bit overhead, a triangular numeral system (TNS)-based CAC, named TNS-CAC, is proposed and applied. The simulation results show that the proposed TNS-CAC has advantages on bit overhead and codec area. The proposed Mosaic-3C1S scheme with TNS-CAC is able to suppress the crosstalk between TSVs to 6C level. The bit overhead of the proposed scheme is between 25% and 45%, which is lower than that of the previous CAC methods for TSV arrays for the most cases. Xiaole Cui, Xiaoxin Cui |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | An Evaluation Method of the Anti-Modeling-Attack Capability of PUFsabstractThe physical unclonable function (PUF) is regarded as the root of trust of hardware systems. However, it suffers from the modeling attacks based on machine learning (ML) algorithms. Subsequently, the anti-modeling-attack PUF is of great concern from academia and industries in recent years. In practice, the security of a given PUF is evaluated after the pertinent attacks. However, these evaluation methods are not helpful for the PUF design, because the relationship between the PUF structure and the anti-modeling-attack capability is not established explicitly. This work proposes a security evaluation method of PUF based on the mimic attack and the Probably Approximately Correct (PAC) theory. The anti-modeling-attack capability of PUF is measured by the corresponding area enclosed by the evaluation curve. Twenty representative types of PUFs with different sizes are evaluated by the proposed method. It shows that the proposed method is effective, because the evaluation results are consistent with the difficulties of modeling attacks for the corresponding PUFs in the design practices. And the evaluation results are able to assist the PUF design. Xiaole Cui, Yun Liu 0031, Xiaoxin Cui |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | An obfuscation scheme of scan chain to protect the cryptographic chipsabstractThe scan chain, as the most popular structured design for testability (DfT) approach, significantly improves the controllability and observability of the circuit under test (CUT). However, the scan chain can be exploited by the hackers to retrieve the secret information in the cryptographic chip. This work proposes an obfuscation scheme of scan chain with test key. The scan chain is divided into segments by the multiplexers, and the linear feedback shift registers (LFSRs) are constructed by the scan cells in the chain. The selection signals of the multiplexers are controlled by the test key bits, and the two input signals of the multiplexers are the original scan chain data and the feedback data of the corresponding LFSR, respectively. The scan chain works normally if all the key bits are correct. Any wrong key bit leads to the flips on the selection signal of the inserted multiplexer periodically during the test mode, and the scan output is obfuscated. Analysis and simulation results demonstrate that the proposed scheme is resilient against the previous scan-based attacks, while maintaining the advantages of the scan testing, Huixian Huang, Xiaole Cui, Xiaoxin Cui |
ATS | 2 |
| 2022 | The Design Method of Logic Circuits based on the Voltage-Input Enhanced Scouting Logic GatesabstractThe Enhanced Scouting Logic (ESL) is a memristive logic gate family with low sensitivity to resistance variation and high device endurance. This work studies the design methods of logic circuits based on the Voltage-Input Enhanced Scouting Logic (VIESL) gates. Both the single-array and dual-array synthesis methods are proposed. The read/write separation technique of VIESL gates facilitates the pipelined logic operations. The synthesis results on the benchmarks show that the circuit generated by the proposed single-array synthesis method has the best performance compared with that of its counterparts, and the dual-array synthesis method reduces the cell counts effectively. Sunrui Zhang, Xiaole Cui |
FPL | 3 |
| 2022 | An Area-Efficient and Robust Memristive LUT Based on the Enhanced Scouting Logic CellsabstractThe resistive random access memory (RRAM) is a two-terminal device, which represents logic states with its different resistance states. The RRAM devices were applied to the Look-Up Table (LUT) in recent years. However, the RRAM based logic circuits are affected by the resistance variation of the RRAM devices. This work proposes a memristive LUT scheme based on the enhanced scouting logic (ESL) cells, to address this challenge. The read-write separation feature of the ESL cell is applied to reduce the number of working cycles of the proposed LUT circuit. The Monte Carlo simulation results show that the proposed LUT scheme has the small standard deviation. And the proposed LUT has relatively small area and high performance. Xiaole Cui, Sunrui Zhang, Xiaoxin Cui |
ISCAS | 1 |
| 2021 | The Modeling Attack and Security Enhancement of the XbarPUF with Both Column Swapping and XORingabstractTo address the security challenge of integrated circuits, the Physical Unclonable Function (PUF) is of great concern as the root of trust. However, the PUF circuits are suffering from the modeling attacks in recent years. The design of anti-modeling-attack PUF is still an open issue. The XbarPUF with both column swapping and XORing was reported as an anti-modeling-attack PUF in 2017. This work proposes a two-step attack method. The first step transforms the target PUF into the XbarPUFs with XORing only, based on the column swapping states. The second step attacks the XbarPUFs with XORing only by the intentionally designed Artificial Neural Network (ANN) model. This method can predict the challenge response pairs (CRPs) of the XbarPUF with both column swapping and XORing successfully. To enhance the anti-modeling-attack capability, an improved XbarPUF is further proposed, which uses the dynamic column swapping technique. The results show that the prediction accuracy of attacks for the XbarPUF with both column swapping and XORing and the proposed XbarPUF reaches 98.96% and 70%, respectively, if the training CRPs account for one millionth of the total CRPs. The proposed XbarPUF has a better anti-attack capability than the XbarPUF with both column swapping and XORing. Xiaole Cui, Wenqiang Ye, Xiaoxin Cui |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | An SNN-Based and Neuromorphic-Hardware-Implementable Noise Filter with Self-adaptive Time Window for Event-Based Vision SensorabstractEvent-based dynamic vision sensors (DVS), inspired by biological vision systems, lead to new sensing and computing paradigms. The novel sensors output the sensed signal alone with many noise events asynchronously. Data-preprocessing for filtering these noises is significant before utilizing the data in applications such as classification, tracking and motion-data extraction. This paper describes a fully spike-based and neuromorphic-hardware-implementable neural network with a signal-oriented self-adaptive filtering time window for filtering the noise events robustly in the data captured by DVS. In particular, the simple leaky integrate-and-fire (LIF) neuron model is adopted as the basic elements of the network out of the purpose of hardware-friendly. Experiments based on both synthesized data and authentically-captured data are designed for quantitative comparison with traditional DVS noise filters to verify the outperformance of the proposed filter. The main contribution of this work is that the proposed spiking neural network (SNN) based filter achieves higher signal-noise-ratio (SNR) compared to traditional noise filters and performances more robust in the tolerance for changing signals. Kanglin Xiao, Xiaoxin Cui, Kefei Liu 0002, Xiaole Cui, Xin'an Wang |
IJCNN | 4 |
| 2021 | The ANN Based Modeling Attack and Security Enhancement of the Double-layer PUFabstractThe modeling attack is a serious threat to the physical unclonable function (PUF) circuits. The double-layer PUF was reported as a PUF scheme to resist the machine learning attacks, and its test chip was fabricated and tested. This work attacks the double-layer PUF successfully by an intentionally designed artificial neural network (ANN) model based on the working principle of the target PUF. To enhance the anti-modeling-attack capability of the double-layer PUF, the XORing and the dimensional extension techniques are proposed. The attack results show that the prediction accuracy of the proposed ANN-based model with the XORing and 3D extension techniques is as low as 50.09% in average. It manifests that the proposed security enhancement techniques are able to improve the resilience of the double-layer PUF against the modeling attacks effectively. Xiaole Cui, Wenqiang Ye, Xiaoxin Cui |
ITC-Asia | 1 |
| 2021 | The Security Enhancement Techniques of the Double-layer PUF Against the ANN-based Modeling AttackabstractThe physical unclonable function (PUF) against the modeling attack is of great concern in recent years, since the modeling attack has been proved to be a serious security threat to the PUF circuits. The double-layer PUF was reported as a PUF scheme to resist the fully connected artificial neural network based modeling attack, and its test chip was fabricated and tested. This work proposes an artificial neural network (ANN) based modeling method according to the working principle of the target PUF, and successfully attacks the double-layer PUF. To enhance the anti-modeling-attack capability of the double-layer PUF, the address swapping, the XORing, and the dimensional extension techniques are proposed. The attack results show that the prediction accuracy of the proposed ANN-based model with the proposed techniques drops obviously. And the prediction accuracy is about 50.04% if all the three proposed techniques are applied in combination. It manifests that the proposed security enhancement techniques are able to improve the resilience of the double-layer PUF against the modeling attacks effectively. Both the randomness and uniqueness of the improved doublelayer PUFs are approximate to the ideal value (50%), and the reliability of the improved PUFs remain unchanged compared with the original counterpart because the operations on the resistive random memory (RRAM) array are the same. Xiaole Cui, Wenqiang Ye, Xiaoxin Cui |
ITC | 2 |
| 2021 | Machine Learning Aided Key-Guessing Attack Paradigm Against Logic Block EncryptionabstractHardware security remains as a major concern in the circuit design ow. Logic block based encryption has been widely adopted as a simple but effective protection method. In this paper, the potential threat arising from the rapidly developing field, i.e., machine learning, is researched. To illustrate the challenge, this work presents a standard attack paradigm, in which a three-layer neural network and a naive Bayes classifier are utilized to exemplify the key-guessing attack on logic encryption. Backed with validation results obtained from both combinational and sequential benchmarks, the presented attack scheme can specifically accelerate the decryption process of partial keys, which may serve as a new perspective to reveal the potential vulnerability for current anti-attack designs. Yi Zhong 0002, Jianhua Feng, Xiaoxin Cui, Xiaole Cui |
J. Comput. Sci. Technol. | 4 |
| 2020 | A Testability Enhancement Method for the Memristor Ratioed Logic CircuitsabstractThe Resistive Random Access Memory (RRAM) is a two-terminal variable resistance device, and the memristor ratioed logic (MRL) is a hybrid RRAM-CMOS style of logic circuit. The MRL AND and OR gates are implemented by the RRAM devices, and the MRL NOT gate is implemented by the CMOS inverter. However, the MRL circuits are prone to test escape. This work proposes a method to improve the testability of MRL circuits by replacing the CMOS inverters with the FinFET inverters. The test escape problem is solved by adjusting the switching threshold voltages of the FinFET inverters, and it is implemented by selecting the FinFET inverters in different operation modes. In Addition, some equivalent relationships in the fault set of the improved MRL gates are discovered. These fault equivalences of MRL gates result in a fault collapse ratio of about 50%. The test patterns for the production test of the improved MRL circuits can be generated by the traditional ATPG method. The test results of some typical MRL circuits obtained from the commercial ATPG tool show that the proposed method is able to achieve 100% fault coverage and at least 55% fault collapse ratio. Li Qu, Xiaole Cui, Xiaoxin Cui |
ATS | 2 |
| 2020 | ReTriple: Reduction of Redundant Rendering on Android Devices for Performance and Energy OptimizationsabstractGraphics rendering is a compute-intensive work and a major source of energy consumption on battery-driven mobile devices. Unlike the existing works that degrade user experience or reuse rendering results coarsely, we propose ReTriple, a fine-grained scheme to reduce rendering workload by reusing the past rendering results at the UI element level. This fine-grained reuse mechanism can explore more opportunities to reduce the workload of the rendering process and save energy. The experiments tested with popular apps show that ReTriple achieves an average speedup of 2.6x and per-frame energy saving of 32.3% for the rendering process while improving user experience. Gengchao Li, Xiaole Cui |
DAC | 3 |
| 2020 | A synthesis method for logic circuits in RRAM arrays
Xiaole Cui, Xiaoxin Cui |
Sci. China Inf. Sci. | 1 |
| 2020 | The synthesis method of logic circuits based on the iMemComp gates
Xiaole Cui, Qiujun Lin, Xiaoxin Cui, Jinfeng Kang |
Integr. | 1 |
| 2019 | Efficient evaluation model including interconnect resistance effect for large scale RRAM crossbar array matrix computing
Runze Han, Peng Huang 0004, Yudi Zhao, Xiaole Cui, Jinfeng Kang |
Sci. China Inf. Sci. | 4 |
| 2018 | Polymorphic gate based IC watermarking techniquesabstractPolymorphic gates are reconfigurable devices whose functionality may vary in response to the change of execution environment such as temperature, supply voltage or external control signals. This feature makes them a perfect candidate for circuit watermarking. However, polymorphic gates are hard to find because they do not exhibit the traditional structure. In this paper, we report four dual-function polymorphic gates that we have discovered using an evolutionary approach. With these gates, we propose a circuit watermarking scheme that selectively replaces certain standard logic gates with the polymorphic gates. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that this scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.10%, 2.08% and 1.17% respectively when we embed 30-bit watermark sequences. These overheads increase to 6.36%, 4.75% and 2.08% respectively when 10% of the gates in the original circuits are replaced to embed watermark up to more than 300 bits. Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui |
ASP-DAC | 7 |
| 2018 | A Novel Polymorphic Gate Based Circuit Fingerprinting TechniqueabstractPolymorphic gates are reconfigurable devices that deliver multiple functionalities at different temperature, supply voltage or external inputs. Capable of working in different modes, polymorphic gate is a promising candidate for embedding secret information such as fingerprints. In this paper we report five polymorphic gates whose functionality varies in response to specific control input and propose a circuit fingerprinting scheme based on these gates. The scheme selectively replaces standard logic cells by polymorphic gates whose functionality differs with the standard cells only on Satisfiability Don't Care conditions. Additional dummy fingerprint bits are also introduced to enhance the fingerprint's robustness against attacks such as fingerprint removal and modification. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that our scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.04%, 6.97% and 4.15% respectively when we embed 64-bit fingerprint that consists of 32 real fingerprint bits and 32 dummy bits. This is only half of the overhead of the other known approach when they create 32-bit fingerprints. Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui |
ACM Great Lakes Symposium on VLSI | 7 |
| 2018 | Evaluation of Dynamic-Adjusting Threshold-Voltage Scheme for Low-Power FinFET Circuits
Xiaoxin Cui, Yewen Ni, Dunshan Yu, Xiaole Cui |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | A Heuristic Algorithm for Automatic Generation of March TestsabstractMarch test is one of the most popular memory test algorithms for its good fault coverage and linear complexity. However, designing an efficient March test for a complex memory fault set is a tedious task. This work designs a two-phase heuristic algorithm for automatic generation of March tests with two new observations. One observation is that some operations can be deleted from the march element, if other operations are inserted to sensitize the Loop-Sens fault, wherein the required initial state is equal to the expected faulty state. The other observation is that there exist chances to reduce the length of march sequence after some operation segments between march elements are exchanged, if the expected states of the operation segments in the different march elements match each other. Experimental results show that the proposed algorithm is effective and efficient. The generated March tests cover the target fault sets, and the complexities approach to those of the minimal March tests generated from the exhaustive methods for the specific fault sets. Xiaole Cui, Yichi Luo, Qiujun Lin, Xiaoxin Cui |
ATS | 1 |
| 2017 | Testing of 1TnR RRAM array with sneak path technique
Xiaole Cui, Xiaoxin Cui, Xin'an Wang, Jinfeng Kang |
Sci. China Inf. Sci. | 1 |
| 2017 | Improving DFA attacks on AES with unknown and random faults
Nan Liao, Xiaoxin Cui, Dunshan Yu, Xiaole Cui |
Sci. China Inf. Sci. | 6 |
| 2017 | An Enhancement of Crosstalk Avoidance Code Based on Fibonacci Numeral System for Through Silicon ViasabstractThrough silicon vias (TSVs) play an important role as the vertical electrical connections in 3-D stacked integrated circuits. However, the closely clustered TSVs suffer from the crosstalk noise between the neighboring TSVs, and result in the extra delay and the deterioration of signal integrity. For a 3 × 3 TSV array, the severity of crosstalk noise in the center victim TSV is classified into 11 levels, which is defined as 0C to 10C from low noise to high noise, depending on the combinations of the digital patterns applied to the TSV array. An enhanced code based on the Fibonacci number system (FNS) to suppress the crosstalk noise below 6C level is proposed, in which both the redundancy of numbers and the nonuniqueness of Fibonacci-based binary codeword are utilized to search the proper codeword. Experimental results show that the proposed technique decreases about 22% latency of TSVs comparing with the worst crosstalk cases. This technique is applicable in the large-scale TSV array for it has a quasi-linear hardware overhead, and its system overhead is less than that of the 3-D 4-LAT counterpart if the data width is greater than 18, and it has good usability for it consumes less power per TSV and achieves lower bit error rate at the interested frequency range comparing with that of the original FNS coding technique. Xiaole Cui, Xiaoxin Cui, Yewen Ni, Min Miao, Yufeng Jin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | A snake addressing scheme for phase change memory testing
Xiaole Cui, Zuolin Cheng, Chung-Len Lee 0001, Xinnan Lin, Yiqun Wei, Zhitang Song |
Sci. China Inf. Sci. | 1 |
| 2016 | Ultralow-power high-speed flip-flop based on multimode FinFETs
Xiaoxin Cui, Nan Liao, Dunshan Yu, Xiaole Cui |
Sci. China Inf. Sci. | 6 |
| 2015 | Context-adaptive fast motion estimation of HEVCabstractHigh Efficient Video Coding (HEVC) is the latest coding standard with superior compression efficiency while its encoding complexity is much higher compared with H.264/AVC. Motion estimation is one of the most time-consuming parts in video coding. In the reference software of HEVC, TZ (Test Zone) search method is adopted as the fast motion estimation method. However, its complexity is still high. There are many other fast motion estimation methods, for example, the hexagon search method, but their performance loss is larger than TZ search. In order to balance coding speed and performance, a new context-adaptive fast motion estimation algorithm is proposed in this paper. In this the proposed, motion intensity is defined in block-level, motion vectors and motion vector differences of neighbor blocks are utilized to measure the motion intensity. When motion intensity is large, TZ search method is used; otherwise, hexagon search method is used. Experimental results show that the proposed method can save 39% ~ 60% of motion estimation time with average 0.5% of BD-rate loss. Ronggang Wang, Xiaole Cui, Wenmin Wang 0001 |
ISCAS | 3 |
| 2013 | A UWB mixer with a balanced wide band active balun using crossing centertaped inductorabstractA high gain 3-5GHz mixer merged with a very balanced wideband active balun is presented and demonstrated. It uses a Gilbert type folded structure with the active balun as the input trans-conductance stage and with a PMOS switch stage. The implemented circuit in the 0.18um CMOS technology exhibits less than 1dB amplitude imbalance in 2.06-5.2GHz, less than 1.5dB in 0.55-5.7GHz and less than 2 degree phase imbalance through 1.37GHz to 5.07GHz. The mixer exhibits a high IIP3 of -3dBm and a high isolation of LO-RF about -100dB in the wide bandwidth. It consumes only 6.2mW under a 1.8V power supply. Xiangrong Zhang, Xiaole Cui, Bo Wang 0016, Chung-Len Lee 0001 |
ISCAS | 2 |
| 2012 | Modeling and testing of interference faults in the nano NAND Flash memoryabstractAdvance of the fabrication technology has enhanced the size and density for the NAND Flash memory but also brought new types of defects which need to be tested for the quality consideration. This work analyzes three types of physical defects for the deep nano-meter NAND Flash memory based on the circuit level simulation and proposes new categories of interference faults (IFs). Testing algorithm is also proposed to test the faults under the worst case condition. The algorithm, in addition to test IFs, can also detect the conventional address faults, disturbance faults and other RAM-like faults for the NAND Flash. Jin Zha, Xiaole Cui, Chung-Len Lee 0001 |
DATE | 2 |