VLDB 2026 Research / reviewers in the wild / expert
Peng Liu 0045
dblp:21/6121-45
· DBLP profile ↗
24ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-2329-502XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAGSIS: A DAG-Aware MAGIC-Based Synthesis Framework for In-Memory ComputingabstractThis paper presents a comprehensive synthesis framework, named DAGSIS, for memristor-aided logic (MAGIC)-based in-memory computing system. DAGSIS addresses the limitations of prior works, such as overlooking the benefits of MAGIC’s high fan-in capability and the impact of global properties of netlists on the scheduling of computation sequence (CS). DAGSIS achieves the optimization in two synthesis stages. In the technology-independent optimization stage, DAGSIS encourages the merging of nodes in the network to reduce circuit size, by utilizing equivalent transformation of multiplexer (MUX). In the CS scheduling stage, DAGSIS introduces two schemes for optimizing area overhead and latency, respectively. For area optimization, DAGSIS maximizes the utilization of memristive cells by erasing the expired data as early as possible. For latency optimization, DAGSIS aims to minimize erasing operations, by maximizing the number of erased cells in each epoch of filling the memory. To achieve better CS scheduling, DAGSIS introduces two design rules to guide CS scheduling, which fully considers the global attributes of circuit design, such as critical path and high fan-out nodes. Experiment results show that DAGSIS reduces the circuit size by 6.69% on ISCAS’85 benchmarks compared to ABC tool, an open-source logic synthesis framework. Compared to the state-of-the-art works, DAGSIS achieves a reduction of 40.68% and 12.67% in area overhead and erasing operations respectively, on ISCAS’85 and EPFL benchmarks. The improvements are further translated into the reduction in energy consumption by up to 13.7%. Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | A Novel Memristive Combinational Logic for Accelerating N-bit AddersabstractMemristors are anticipated to replace CMOS technology due to their low power consumption and high-speed in-memory processing capabilities. However, most existing in-memory technologies implement traditional Boolean functions to achieve complex functions, which lead to increased logical depth and long delays due to the repeated iterations. In this paper, we propose a novel combinational logic, namely the AND-OR gate, which integrates the functionality of AND and OR logic into a single function. The proposed gate is able to implement multiple commonly-used logic functions within a single cycle. To highlight the advantages of the proposed AND-OR gate, twoN-bit adders, based on parallel prefix algorithms, are designed by integrating the AND-OR gate into a memristive crossbar array. Benefiting from the proposed AND-OR gate that supports prefix computation within a single cycle, the latency of the adders is significantly reduced to$O(log(N))$. Compared with the fastest reported adder that uses Majority gate for the implementation, our proposed adder achieves notable performance improvements of$1.2\times $and$2.7\times $in terms of latency and area, respectively. Moreover, the Figure of Merits (FoMs) are employed for fair comparison, in which both latency and area are considered simultaneously. Simulation results demonstrate a remarkable improvement of$40\times $over state-of-the-art circuit (i.e., the carry-select adder). Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Online Heterogeneous Feature SelectionabstractMany real-world datasets contain high-dimensional heterogeneous features, exhibiting complex and evolving distributions. The coexistence of high dimensionality and heterogeneity poses challenges for reliable feature selection and real-time analysis, while most existing feature selection solutions either assume that the features are of the same type or struggle to handle extremely high-dimensional features. Moreover, these methods are usually designed for static datasets, neglecting the dynamic capture of heterogeneous interfeature relationships in real-time environments. To address these challenges, we propose a new feature selection method called graph-unified adaptive decision boundary enhancement (GRADE) for online heterogeneous feature selection (OHFS). To provide a reliable foundation for evaluating feature subsets under dynamic and heterogeneous data streams, an incremental graph-unified metric (IGUM) is introduced. It mitigates information loss between heterogeneous features by leveraging graph structures to unify feature-value-level and interfeature-level relationships. With such a consistent relation measure, an adaptive density-guided neighborhood relation (ADNR) is proposed to assess the capability of selected feature subsets to classify samples. Since it dynamically captures prominent neighborhood regions, local decision boundaries can thus be precisely delineated. It turns out that GRADE can obtain a more concise feature subset while achieving competitive classification accuracy. Besides, GRADE is parameter-free and very efficient compared with state-of-the-art methods. Comprehensive experimental evaluations, including significance tests, ablation studies, efficiency evaluation, and case studies, have been conducted to verify the efficacy of GRADE. Yiqun Zhang 0006, Xinxi Chen, Lang Zhao, Yuzhu Ji, Peng Liu 0045, Yiu-Ming Cheung |
IEEE Trans. Cybern. | 5 |
| 2025 | SSA: An Effective Secure Scan Architecture Based on Hidden, Randomly Inserted Keys and PUFabstractWith the rapid growth of Internet of Things (IoT) applications, ensuring the data privacy and security has become a critical concern. The security of integrated circuits (ICs) in IoT systems is of paramount importance. As IC integration increases, circuit testing becomes increasingly challenging. Design For Testability (DFT) techniques improve testability, but scan-based methods can be easily exploited by unauthorized users to steal confidential information. For this reason, researchers have proposed various defense techniques, but each technique has its limitations. In order to effectively protect the scan chains of ICs from illegal access, this paper presents a novel DFT framework based on hidden, randomly inserted key seeds and Physical Unclonable Function (PUF). A multiple-input shift register is used as a test key generator, which receives the hidden test key seed from randomly selected scan chains, and then generates a unique key for each test pattern. An authentication module is also introduced to verify the validity of the test patterns. Through key verification, if a test vector contains valid key seed, the design will perform normal scan operations. Conversely, if the seed hidden in a test vector cannot generate the expected test key, the scan output data will be scrambled by the random PUF responses, thereby effectively preventing attackers from inferring sensitive information. Theoretical analysis and experimental results show that the proposed secure DFT protects the cryptographic chip from all known scan-based side-channel attacks, while requiring no additional test preparation time and incurring only negligible low area overhead. Weizheng Wang 0002, Zhizhi Wu, Jinhai Chen, Peng Liu 0045, Shuo Cai, Naixue Xiong |
IEEE Internet Things J. | 4 |
| 2024 | Efficient Topology-Driven Clustering for Imbalanced Streaming Biomedical Data AnalysisabstractClustering drifting data is common in the field of biomedical data analysis. Data chunks collected at different periods often exhibit clusters with significantly different sizes, and drifting distributions of clusters also appear frequently. We call such composite phenomenon imbalance-drifting, which can severely impact the accuracy and efficiency of cluster analysis. Therefore, we propose a topology-representation-based clustering paradigm, which first learns an informative global data representation in a self-organizing manner to obtain a map with nested representative data points. Then fast and accurate clustering is facilitated by quickly retrieving similar data points according to the topology. As the constructed Self-Organizing Map (SOM) is exploited for informative representation, micro partition, and quick merging, to achieve advanced clustering under imbalance-drifting, the proposed approach is thus called Tri-Squeezing SOM for Clustering (TSSC). It turns out that TSSC significantly reduces the time complexity for clustering an n-scale imbalance-streaming data without sacrificing accuracy. Moreover, TSSC can automatically determine the number of clusters k, and features interpretability and hyper-parameter robustness. Extensive results on both biomedical datasets and synthetic datasets verify the superiority of TSSC. Xiaopeng Luo, Yiqun Zhang 0006, Yuzhu Ji, Peng Liu 0045, Taoting Xiao |
BIBM | 4 |
| 2024 | Robust Categorical Data Clustering Guided by Multi-Granular Competitive LearningabstractData set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, the nested granular cluster effect is prevalent in the implicit discrete distance space of categorical data. That is, data objects frequently overlap in space or subspace to form small compact clusters, and similar small clusters often form larger clusters. However, the distance space cannot be well-defined like the Euclidean distance due to the qualitative categorical data values, which brings great challenges to the cluster analysis of categorical data. In view of this, we design a Multi-Granular Competitive Penalization Learning (MGCPL) algorithm to allow potential clusters to interactively tune themselves and converge in stages with different numbers of naturally compact clusters. To leverage MGCPL, we also propose a Cluster Aggregation strategy based on MGCPL Encoding (CAME) to first encode the data objects according to the learned multi-granular distributions, and then perform final clustering on the embeddings. It turns out that the proposed MGCPL-guided Categorical Data Clustering (MCDC) approach is competent in automatically exploring the nested distribution of multi-granular clusters and highly robust to categorical data sets from various domains. Benefiting from its linear time complexity, MCDC is scalable to large-scale data sets and promising in pre-partitioning data sets or compute nodes for boosting distributed computing. Extensive experiments with statistical evidence demonstrate its superiority compared to state-of-the-art counterparts on various real public data sets. Shenghong Cai, Yiqun Zhang 0006, Xiaopeng Luo, Yiu-Ming Cheung, Hong Jia, Peng Liu 0045 |
ICDCS | 6 |
| 2024 | Accelerating Frequency-domain Convolutional Neural Networks Inference using FPGAsabstractLow-end field programmable gate arrays (FPGAs) are difficult to deploy typical convolutional neural networks (C- NNs) owing to the limited hardware resources and the increasing model computational complexity. Fast Fourier transform (FFT) is a promising solution for saving both computation and memory footprint by convolving in the frequency domain. However, few FPGA accelerators can take full advantage at the computation level, because of the distinct element-wise complex calculation in the frequency domain. In this work, we present an FPGA-based 8-bit inference accelerator (called FAF) that packs frequency-domain calculations into digital signal processing (DSP) blocks to fully utilize DSPs for performance boost. We then provide a mapping dataflow to maximize the reduction of redundant packing operations by frequency-domain data reuse. Evaluations based on representative CNN benchmarks show that our work can achieve 1.5-6.9× better power efficiency compared with representative FPGA baselines. Bosheng Liu, Yongqi Xu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
ISCAS | 6 |
| 2024 | A Complementary Resistive Switch-Based Balanced Ternary LogicabstractMemristors offer advantages in terms of high speed, high integration density, and non-volatility, making them a promising option for efficient logic applications. Recent works have explored the design methodology for ternary logic in memristor-based computing-in-memory (CIM) systems. However, existing methods require a large number of devices and are susceptible to noise interference. To address these issues, this work proposes a reliable in-memory computing paradigm for balanced ternary logic based on complementary resistive switch (CRS), which can be considered as two anti-serially connected memristors. Six balanced ternary logic gates are designed based on the proposed method, which support parallel operations when integrated into the CRS crossbar array. To demonstrate the efficiency of the proposed method, a 1-tri full adder is designed by using the proposed logic gates. The feasibility of the design is verified by Cadence Virtuoso using the Voltage Threshold Adaptive Memristor (VTEAM) model. The Monte Carlo simulation of the full adder verifies the reliability of the proposed method. Compared to existing methods, both the operation steps and area overhead are reduced using the proposed approach. Zhijian Peng, Peng Liu 0045, Lian Yao, Zhiqiang You, Bosheng Liu, Jigang Wu |
ITC-Asia | 2 |
| 2024 | A Dynamic Weight Quantization-Based Fault-tolerant Training Method for Ternary Memristive Neural NetworksabstractMemristors have the merits of small area, low power consumption, and non-volatility, which are eminently suitable for storing the weights of neural networks. However, stuckat faults (SAFs) and variations in memristor devices significantly degrade the recognition accuracy of ternary memristive neural networks (TMNNs). In response to this issue, we propose a dynamic weight quantization-based fault-tolerant training method to obtain high recognition accuracy for TMNNs with SAFs and variations. A dynamic weight quantization function is proposed by a column-wise weight quantization method considering these faults, where each column has a different symmetric threshold of a TMNN. A hardware activation function is implemented by setting the bias voltages of amplifiers and inverters in memristive crossbar arrays (MCAs). Experimental results show that we can achieve the best recognition accuracy in image classification on MNIST for TMNNs with SAFs and variations. The average recovered accuracy for various SAF ratios is 98.7%, which demonstrates the practicality of our method. Compared with the state-of-the-art methods, the accuracies of our proposed method are 0.88% more when there are 40% of SAFs, 0.87% more when the standard deviation of variations σ is 0.2, 1.7 % more when the SAF ratio is 20% and σ=0.2. We also evaluate our method using VGG-small on CIFAR-10 and U-Net for image segmentation on Portrait to show its effectiveness. Zhiqiang You, Peng Liu 0045 |
ITC-Asia | 3 |
| 2024 | DAF: An Effective Design-for-Testability Authorization Framework Based on Obfuscation Mechanisms for Defending Complex AttacksabstractThe data privacy and security of Internet of Things (IoT) applications are progressively crucial and cryptography is frequently used to ensure security. Cryptographic hardware implementation is commonly adopted for high throughput and low-computational resources. The crypto circuits have to be strictly tested to guarantee the correctness of data. Scan-based design-for-testability (DFT) widely employed in the chip industry improves the controllability and observability of circuits. However, it facilitates illegal users to steal the internal data for cracking cryptographic keys. Many researchers have recently suggested effective defense technologies against scan-based attacks, but each technology has its own negative aspects. In this article, we propose an effective DFT authorization framework (DAF) based on obfuscation mechanisms. The authentication module is embedded to verify users’ keys, and it can empower each chip distinctive key to minimize the loss of key divulgence. If the test authorization key is accurate, the typical scan operation can be carried out. Conversely, if the key is inaccurate, the inserted obfuscation module comes into play to scramble the scan data. Additionally, a random number generation circuit is used to increase the uncertainty of data obfuscation, effectively preventing attackers from inferring sensitive information. Simulation results show that this design has a low overhead, and theoretical analysis demonstrates that the design has high security with no impact on the testability of the chip. Weizheng Wang 0002, Xingxing Gong, Xiangqi Wang, Shuo Cai, Peng Liu 0045, Naixue Xiong |
IEEE Internet Things J. | 5 |
| 2023 | Accelerating Convolutional Neural Networks in Frequency Domain via Kernel-Sharing ApproachabstractConvolutional neural networks (CNNs) are typically computationally heavy. Fast algorithms such as fast Fourier transforms (FFTs), are promising in significantly reducing computation complexity by replacing convolutions with frequency-domain element-wise multiplication. However, the increased high memory access overhead of complex weights counteracts the computing benefit, because frequency-domain convolutions not only pad weights to the same size as input maps, but also have no sharable complex kernel weights. In this work, we propose an FFT-based kernel-sharing technique called FS-Conv to reduce memory access. Based on FS-Conv, we derive the sharable complex weights in frequency-domain convolutions, which has never been solved. FS-Conv includes a hybrid padding approach, which utilizes the inherent periodic characteristic of FFT transformation to provide sharable complex weights for different blocks of complex input maps. We in addition build a frequency-domain inference accelerator (called Yixin) that can utilize the sharable complex weights for CNN accelerations. Evaluation results demonstrate the significant performance and energy efficiency benefits compared with the state-of-the-art baseline. Bosheng Liu, Hongyi Liang, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Yinhe Han 0001 |
ASP-DAC | 5 |
| 2023 | Double-Layered Dual-Syndrome Trellis Codes Utilizing Channel Knowledge for Robust SteganographyabstractRobust steganography aims to hide message in cover data with high security and guarantee the success of its message extraction although it is disturbed in transmission channel. In this paper we propose a framework of coding scheme extended from Dual-Syndrome Trellis Codes (Dual-STCs) for robust adaptive steganography. We use the conditional probability distribution of correct stego bits conditioned on disturbed stego data as channel knowledge, and formulate error-correcting as maximizing this probability. By extending Dual-STCs to double-layered embedding, we design an iteratively decoding scheme for error-correcting two layer stego bits from their joint conditional probabilities, and strictly prove its convergence. Besides, we design a method to estimate these probability distributions from stego data pairs uploaded/downloaded from the lossy transmission channel. The channel knowledge can also be used by steganographer, and we propose a universal method to revise steganographic distortion values for higher robustness under the guidance of the channel knowledge. Compared with existing coding methods for robust steganography, our method can make use of channel knowledge to improve error correcting ability and meanwhile maintain high security, which is demonstrated by experimental results. Qingxiao Guan, Peng Liu 0045, Weiming Zhang 0001, Wei Lu 0001, Xinpeng Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Frequency-Domain Inference Acceleration for Convolutional Neural Networks Using ReRAMsabstractConvolutional neural networks (CNNs) (including 2D and 3D convolutions) are popular in video analysis tasks such as action recognition and activity understanding. Fast algorithms such as fast Fourier transforms (FFTs) are promising in significantly reducing computation complexity by transforming convolution into frequency domain. In frequency space, conventional spatial convolutions are replaced with simpler element-wise complex multiplications. Conventional application-specific-integrated-circuit (ASIC) based frequency-domain accelerators can achieve effective performance boost but come at the cost of significant energy consumption, owing to the hierarchical memory organization. We propose a frequency-domain resistive random access memory (ReRAM) based inference accelerator called FDA that can process element-wise complex multiplication in memory for both 2D and 3D CNNs. Each ReRAM-based frequency-domain process element (PE) with two ReRAM cells can perform an element-wise complex multiplication in two continuous execution cycles. We then provide a flexible dataflow to alleviate the redundant data movements by frequency-domain data reuse and inherent symmetrical characteristic for both 2D and 3D convolutions. Evaluation results based on representative both 2D and 3D CNN benchmarks demonstrate that FDA outperforms state-of-the-art baselines with better performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Yalan Wu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | An improved reconfigurable logic in resistive random access memory
Peng Liu 0045, Jigang Wu, Dongxiang Luo |
Integr. | 3 |
| 2022 | Reconfiguration algorithms for synchronous communication on switch based degradable arrays
Yalan Wu, Jigang Wu, Peng Liu 0045, Yinhe Han 0001, Thambipillai Srikanthan |
Parallel Comput. | 3 |
| 2022 | Search-Free Inference Acceleration for Sparse Convolutional Neural NetworksabstractSparse convolution neural networks (CNNs) are promising in reducing both memory usage and computational complexity while still preserving high inference accuracy. State-of-the-art sparse CNN accelerators can deliver high throughput by skipping zero weights and/or activations. To operate on only nonzero weights and activations, sparse accelerators typically search pairs of nonzero weights and activations for multiplication-accumulation (MAC) operations. However, the conventional search operation results in a severe limitation in the processing element (PE) array scale because of the enormous demands of internal interconnection and memory bandwidth. In this article, we first provide a design principle to free the search process of sparse CNN accelerations. Specifically, the indexes of the static compressed weights access the dynamic activations directly to avoid the search process for MAC operations. We then develop two search-free inference accelerators, called Swan and Swan-flexible, for sparse CNN accelerations. Swan supports search-free sparse convolution accelerations for interconnection and bandwidth saving. Compared with Swan, Swan-flexible not only has the search-free capability but also comprises a configurable architecture for optimum throughput. We formulate a mathematical optimization problem by combining the configurable characterization with the compressive dataflow to optimize the overall throughput. Evaluations based on a place-and-route process show that the proposed designs, in a compact factor of 4096 PEs, achieve 1.5–$2.7\times $higher speedup and 6.0–$13.6\times $better energy efficiency than representative accelerator baselines with the same PE array scale. Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001, Jigang Wu, Liang Chang 0003, Peng Liu 0045 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Ensuring Cryptography Chips Security by Preventing Scan-Based Side-Channel Attacks With Improved DFT ArchitectureabstractCryptography chips are often used in some applications, such as smart grids and Internet of Things (IoT) to ensure their security. Cryptographic chips must be strictly tested to guarantee the correctness of the encryption and decryption. Scan-based design-for-testability (DFT) provides high test quality. However, it can also be misused to steal the cipher key of cryptographic chips by hackers. In this article, we present a new scan design methodology that can resist scan-based side-channel attacks by the dynamical obfuscation of scan input data and scan output data. The scan test is managed by a test password, which consists of load password and scan password. When the chip enters into the test mode, it is required to apply the test password via some external input ports. Once the correct load password is delivered, the scan password can be loaded into a special shift register. If the scan password is also correct, the chip testing can proceed normally. In case the load password or the scan password is wrong, the data in scan chains cannot be propagated correctly. Specifically, some elusory bits are sneaked into scan chains dynamically. The advantage of the proposed method is that it has no negative impact on design performance and test flow when powerfully protecting cryptographic chips. The area penalty is also acceptably low compared with other schemes. Weizheng Wang 0002, Xiangqi Wang, Jin Wang 0001, Naixue Xiong, Shuo Cai, Peng Liu 0045 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2022 | Accurate Reliability Boundary Evaluation of Approximate Arithmetic CircuitabstractApproximate arithmetic circuit (AAC) has emerged as a promising high-performance and energy-efficient circuit paradigm, which can be used in many applications with inherent error tolerance. To guarantee the usability of AACs and the availability of resilient applications, it is necessary to analyze the reliability of AACs. Most current literature focus on the error characteristics of AACs and few methods can be applied to estimate the reliability of AACs. These methods mostly have exponential time complexities and evaluate the average reliability assuming the input combinations are equally likely. In reality, the primary input (PI) signals can be given with any probability from 0 to 1. In this article, we assume that the PIs have random signal probabilities and propose approaches to reliability boundary estimation for AACs. First, we propose a new efficient and accurate method to evaluate the reliability of AACs. The method mainly calculates the AAC reliability for an input vector set, and furthermore, during the calculation, the correlation problem is considered to increase accuracy. Then, based upon the proposed AAC reliability evaluation method, we present the approaches to finding the reliability boundary. Randomly given signal probabilities of every PI, two heuristic search algorithms are utilized to find the lowest reliability. A comparison of the results on three series of AACs and the circuits in the EvoApprox8b library confirms that the proposed reliability evaluation method is more accurate and efficient than the previous method. Further experiments verify the plausibility of the calculated reliability boundary of AACs. Zhen Wang 0042, Guofa Zhang, Peng Liu 0045, Jing Ye 0001, Jianhui Jiang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | F3D: Accelerating 3D Convolutional Neural Networks in Frequency Space Using ReRAMabstract3D convolutional neural networks (CNNs) are widely deployed in video analysis. Fast algorithms such as fast Fourier transforms (FFTs) are gaining popularity in reducing computation complexity for their superior capability of replacing convolutions with simpler element-wise multiplications. Conventional frequency-domain dedicated accelerators employ memory hierarchy organization for high throughput but at the expensive costs of a significant amount of data movements and energy consumptions. This paper presents F3D, a processingin-memory frequency-domain accelerator using resistive random access memory (ReRAM). F3D supports frequency-domain complex number multiplications directly in ReRAM-based crossbar architecture. We alleviate the overheads of redundant data movements in ReRAM-based complex number multiplications by data reuse and the inherent symmetry of inputs in the frequency space. Evaluation results demonstrate that F3D outperforms state-of-the-art accelerators with significant improvements in performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Jigang Wu, Xiaoming Chen 0003, Yinhe Han 0001, Peng Liu 0045 |
DAC | 6 |
| 2021 | Fault Modeling and Efficient Testing of Memristor-Based MemoryabstractMemristor-based memory technology is one of the emerging memory technologies, which is a potential candidate to replace traditional memories. Efficient test solutions are required to enable the quality and reliability of such products. In previous works, fault models are caused by open, short and bridge defects and parametric variations during the fabrication. However, these fault models cannot describe the bridge defects that cause the state of the faulty cell to an undefined state. In this paper, we analyze the different effects of bridge defects and aggregate their faulty behavior into new fault models, undefined coupling fault and dynamic undefined coupling fault. In addition, an enhanced March algorithm is designed to detect all the modeled faults. In one resistor crossbar with$N$memristors, the enhanced March algorithm requires$8N$write and$7N$read operations with negligible hardware overhead. To reduce the test time, a March RC algorithm is proposed based on read operations with new reference currents, which requires$4N+2$write and$6N$read operations. Analytical results show that the proposed test algorithms can detect all the modeled faults outperforming all the previous methods. Subsequently, a Design-for-Testability scheme is proposed to implement March RC algorithm with a little area overhead. Peng Liu 0045, Zhiqiang You, Jigang Wu, Bosheng Liu, Yinhe Han 0001, Krishnendu Chakrabarty |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Jointly Learning Multiple Curvature Descriptor for 3D Palmprint Recognitionabstract3D palmprint-based biometric recognition has drawn growing research attention due to its several merits over 2D counterpart such as robust structural measurement of a palm surface and high anti-counterfeiting capability. However, most existing 3D palmprint descriptors are hand-crafted that usually extract stationary features from 3D palmprint images. In this paper, we propose a feature learning method to jointly learn compact curvature feature descriptor for 3D palmprint recognition. We first form multiple curvature data vectors to completely sample the intrinsic curvature information of 3D palmprint images. Then, we jointly learn a feature projection function that project curvature data vectors into binary feature codes, which have the maximum inter-class variances and minimum intra-class distance so that they are discriminative. Moreover, we learn the collaborative binary representation of the multiple curvature feature codes by minimizing the information loss between the final representation and the multiple curvature features, so that the proposed method is more compact in feature representation and efficient in matching. Experimental results on the baseline 3D palmprint database demonstrate the superiority of the proposed method in terms of recognition performance in comparison with state-of-the-art 3D palmprint descriptors. Lunke Fei, Jianyang Qin, Peng Liu 0045, Jie Wen 0001, Chunwei Tian, Bob Zhang 0001, Shuping Zhao |
ICPR | 3 |
| 2020 | Soft Error Reliability Evaluation of Nanoscale Logic Circuits in the Presence of Multiple Transient Faults
Shuo Cai, Binyong He, Weizheng Wang 0002, Peng Liu 0045, Fei Yu 0009, Lairong Yin, Bo Li 0051 |
J. Electron. Test. | 4 |
| 2018 | Defect Analysis and Parallel March Test Algorithm for 3D Hybrid CMOS-Memristor MemoryabstractAs an attractive option of future non-volatile memories (NVM), resistive random access memory (RRAM) has attracted more attentions. CMOS Molecular (CMOL) architecture, which can alleviate the sneak path problem of one memristor (1R) crossbars and limit its power consumption in 1R crossbars, is used as a large-scale memory system. In this paper, we analyze the electrical defects in a CMOL circuit including open and bridge. A parallel March-like test algorithm is presented for the CMOL architecture, which covers defined faults caused by electrical defects. The test time of the proposed test algorithm is reduced significantly compared with previous test algorithms that are enhanced for CMOL architecture. Peng Liu 0045, Jigang Wu, Zhiqiang You, Michael Elimu, Weizheng Wang 0002, Shuo Cai |
ATS | 1 |
| 2012 | Integrated Heuristic for Hardware/Software Co-design on Reconfigurable DevicesabstractHardware/Software (HW/SW) partitioning and scheduling are essential to the embedded systems. In this paper, a hybrid algorithm derived from Tabu Search and Simulated Annealing is proposed for solving the HW/SW partitioning problem. The virtual hardware resource is set to implement the customized Tabu Search. Earliest-Deadline-First strategy is introduced to describe the reconfiguration of FPGA. Moreover, an algorithm combining the Breadth-First-Search with Depth-First-Search is proposed for HW/SW tasks scheduling to fit for the feature of reconfigurable systems. Experimental results show that, the proposed algorithms produce better performance than the previoPus methods cited in this paper. Peng Liu 0045, Jigang Wu, Yongji Wang 0002 |
PDCAT | 1 |