Zhenglin Liu

dblp:43/310 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 TRO-Based Dual-Domain Voltage Co-Regulation of Digital Logic and SRAM in SoCs
abstract
Conventional low-power system-on-chips (SoCs) commonly regulate digital logic using adaptive voltage and frequency scaling (AVFS), while SRAM voltage is managed by an independent scaling policy. This separation leaves cross-domain SRAM-related critical paths over-margined and prevents system-level energy optimization. This brief presents a unified voltage co-regulation framework that jointly tunes the digital and SRAM supply rails to minimize total SoC energy. A tunable replica oscillator (TRO) is repurposed from a digital timing monitor into a dual-mode delay allocator: its programmable level sets the digital timing slack, and an all-digital AVFS loop adjusts the digital supply to lock the target frequency. The released cycle budget is then converted into SRAM voltage reduction, determined by a cache-based canary test. Domain-level power is profiled on-chip to enable a measurement-driven search for the dual-domain minimum-energy point (DD-MEP) without relying on PVT-dependent model parameters. Silicon results from a taped-out 110-nm Cortex-M3 SoC demonstrate up to 11.4% total energy reduction compared with digital-only AVFS, with only 0.05% area overhead.
Zhaoxu Wang, Mingyang Gong, Zhenglin Liu, Xuecheng Zou
IEEE Trans. Very Large Scale Integr. Syst.6
2025 A Fast and Energy-Efficient Level Shifter With Complementary Output Buffer for Energy-Constrained Systems
abstract
This brief presents a 55-nm level shifter (LS) that enables wide voltage range conversion from 80mV to 1.2V with high energy efficiency and fast transition speed. The proposed design incorporates a complementary output buffer and an assist discharge path to suppress the short-circuit current and enhance the transition speed. A multithreshold transistor strategy is adopted to expand the input range and reduce static power. Measurement results across 15 samples demonstrate robust subthreshold performance with 4.4-ns transition delay and 49.1-fJ/transition energy during 0.3–1.2-V conversion at 1MHz. The measured average minimum convertible input voltages are 80 and 139mV at input frequencies of 50kHz and 1MHz, respectively. The compact layout occupies only 7.96$\mu $m2. Compared to the best benchmarked prior work, the proposed LS achieves 33.8% improvement in energy-delay metrics, making it a highly efficient and scalable solution for energy-constrained systems and the Internet of Things (IoT).
Zhaoxu Wang, Liaoyuan Li, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.7
2025 An Efficient Wide-Voltage Processor With PVTA Tolerance, Voltage Droop Mitigation, and Runtime Ultrafine-Grained Frequency Adaptation
abstract
Traditional processors require substantial design margins to account for process, voltage, temperature, and aging (PVTA) variations, resulting in significant energy efficiency losses. Existing dynamic timing error detection and correction (EDaC) techniques reduce these margins but incur high area overhead and design complexity. In this brief, we propose a processor based on the RI5CY core that integrates a PVTA tolerance and voltage droop mitigation adaptive voltage frequency scaling (AVFS) system. This approach reduces overhead to only 0.065%, enables ultrafine-grained frequency adaptation, and actively mitigates abrupt voltage droops. Additionally, a novel baud rate adaptive UART (BRA-UART) module ensures robust communication across all frequencies. Our processor design achieves a 163% typical performance gain and a 37.7% power reduction in the logic circuitry at near-threshold voltage (NTV), substantially improving energy efficiency.
Zhaoxu Wang, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.6
2024 A Low-Power Variation-Tolerant 7T SRAM With Enhanced Read Sensing Margin for Voltage Scaling
abstract
Reducing the minimum operating voltage (Vmin) and improving the variation tolerance are the main design challenges of voltage-scalable SRAMs. This paper presents a variation-tolerant 7T SRAM with reduced data-dependent RBL leakage to enhance read sensing margin and improve Vmin without any assist techniques. The design implements a delay-tracking-based adaptive timing generation technique and an error detection circuit for PVT variation-tolerant operation and reliable dynamic voltage scaling. An 8Kb SRAM macro is implemented in 55nm CMOS technology to demonstrate our design. Post-layout simulation results show error-free full functionality down to 0.35V at an operating frequency of 202kHz. The RBL leakage current for 7T SRAM cell is improved by 2.6x and the Ion-to-Ioff ratio is improved by 2.3x compared to the conventional 8T SRAM cell. The minimum energy of 1.76pJ is achieved at 0.4V. The proposed 7T SRAM functioning as an on-chip memory in a voltage-scalable SoC is measured and demonstrate operation from 0.75V to 1.2V at 32MHz. The average power consumption for the entire chip performing 7T read and write operations decreased by 19.3% compared to that of the standard 6T SRAM at 0.75V.
Zhenglin Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Fooling Decision-Based Black-Box Automotive Vision Perception Systems in Physical World
abstract
Autonomous vehicles use deep neural networks (DNNs) to build powerful vision perception systems, which provide a theoretical foundation for automated vehicle control. Due to the inherent vulnerability of DNNs, many research works have implemented white-box attacks against automotive vision perception systems in the physical world. However, successful black-box attacks (especially decision-based) in the physical world are rarely mentioned because it is difficult to implement a physical-world adversarial attack without internal knowledge about the vision perception systems. In this paper, we propose PRAD, an end-to-end framework that transfers the existing decision-based black-box adversarial attack algorithms (as the backbone of the framework) targeting the digital domain to the physical world for the first time. Specifically,$T(\cdot)$is first introduced to simulate the real environment changes, e.g., angle, distance, slight shaking, illumination, etc. Then, and crucially, PRAD bridges the non-differentiable black-box attack and the differentiable$T(\cdot)$by the$L_1$loss function. We use the traffic sign recognition system in the vision perception system as an object to conduct comprehensive experiments, including different environmental conditions, black-box attack backbones, models, and datasets. The results demonstrate that the generated adversarial examples in the decision-based black-box setting can fool the commercial traffic sign recognition system into outputting designated misclassifications with high success rates and strong robustness in the physical world (average 90% in target attacks and nearly 100% in non-target attacks), which outperforms the state-of-the-art homogeneous attack methods.
Zhaojun Lu, Liaoyuan Li, Haichun Zhang, Zhenglin Liu, Gang Qu 0001
IEEE Trans. Intell. Transp. Syst.6
2024 An RRAM-Based Computing-in-Memory Architecture and Its Application in Accelerating Transformer Inference
abstract
Deep neural network (DNN)-based transformer models have demonstrated remarkable performance in natural language processing (NLP) applications. Unfortunately, the unique scaled dot-product attention mechanism and intensive memory access pose a significant challenge during inference on power-constrained edge devices. One emerging solution to this challenge is computing-in-memory (CIM), which uses memory cells for logic computation to reduce data movement and overcome the memory wall. However, existing CIM designs do not support high-precision computations, such as floating-point operations, which are essential for NLP applications. Furthermore, CIM architectures require complex control modules and costly peripheral circuits to harness the full potential of in-memory computation. Hence, this article proposes a scalable RRAM-based in-memory floating-point computation architecture (RIME) that uses single-cycle NOR, NAND, and minority logic to implement in-memory floating-point operations. RIME features efficient parallel and pipeline capabilities with a centralized control module and a simplified peripheral circuit to eliminate data movement during computation. Furthermore, the article proposes pipelined implementations of matrix–matrix multiplication (MatMul) and softmax functions, enabling the construction of a transformer accelerator based on RIME. Extensive experimental results show that compared with GPU-based implementation, the RIME-based transformer accelerator improves timing efficiency by$2.3\times $and energy efficiency by$1.7\times $without compromising inference accuracy.
Zhaojun Lu, Md Tanvir Arafin, Haoxiang Yang, Zhenglin Liu, Jiliang Zhang 0002, Gang Qu 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2024 iEDCL: Streamlined, False-Error-Free Error Detection and Correction Scheme in a Near-Threshold Enabled 32-bit Processor
abstract
This article presents internal error detection, correction, and latching (iEDCL), a designer-friendly, fully functional error detection and correction (EDAC) approach tailored for energy-efficient near-threshold systems capable of tolerating variations. It embeds error detection (ED), correction, and latching circuits within a flip-flop (FF) with an additional 15 transistors to monitor critical paths. Notably, iEDCL’s error-aware capability remains stable despite clock latency and parasitic effects, relieving designers of extensive involvement and eliminating false errors. iEDCL is automatedly implemented in an ARM Cortex-M0 processor at 55 nm without extra architecture modifications, incurring only a 6.78% area overhead. An adaptive voltage scaling (AVS) loop enables automatic operation, achieving high energy efficiency beyond the point of the first failure while maintaining a predefined error rate. Measurement results obtained from different dies at various temperatures demonstrate significant energy savings achieved by the iEDCL processor, with up to 16.9% and 49.1% reductions compared to critical baseline and signoff designs, respectively, while maintaining a 5% error rate at a 16 MHz frequency. To the best of our knowledge, this article presents one of the first FF EDAC implementations fully operational without potential false errors at near-threshold voltages while enhancing energy efficiency.
Zhaoxu Wang, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.7
2023 An FPGA-Compatible TRNG with Ultra-High Throughput and Energy Efficiency
abstract
In this paper, we design an energy-efficient true random number generator with ultra-high throughput for FPGA. Only four ring oscillators constructed using eight LUTs are sampled by multiple sampling points to fully exploit the randomness of the entropy source, which provides high-quality and over 275 Mbps random sequences while consuming 13 slices. An end-to-end implementation and testing framework is tailored for easy deployment and portability on Xilinx 7 serials FPGAs. The proposed architecture passes the NIST SP 800-22 and 800-90B tests without post-processing and outperforms the state-of-the-art in terms of minimum entropy and energy efficiency.
Zhaojun Lu, Houjia Qidiao, Qidong Chen, Zhenglin Liu, Jiliang Zhang 0002
DAC4
2023 ADLPT: Improving 3D NAND Flash Memory Reliability by Adaptive Lifetime Prediction Techniques
abstract
NAND flash memory has become increasingly popular in various computing systems. Although NAND flash memory offers attractive performance, it suffers limited operable programming and erasing cycles. To improve the reliability of flash-based systems, previous works introduce machine learning models to predict flash lifetime. These works generally focus on improving prediction accuracy but present little research about the resources required for flash lifetime prediction. In application scenarios, the overheads and the frequency of lifetime predictions are important for storage systems. Excessive prediction actions would lead to unnecessary resource consumption. For building an efficient storage system, resource requirements need to be taken into consideration when designing flash lifetime prediction schemes. In this paper, we propose adaptive lifetime prediction techniques (ADLPT) that minimize redundant prediction operations by exploiting reliability variation. To explore reliability variation, we investigate the error distribution of different 3D flash chips. Based on the investigation, a prediction judgment method is presented. The method identifies the necessary prediction by detecting the variation of erase duration and raw bit errors. Furthermore, we provide a method to improve the performance of the static model. The experimental result shows that our approach can reduce about 90% of redundant predictions with over 0.8 F1-Score.
Yuqian Pan, Zhaojun Lu, Haichun Zhang, Md Tanvir Arafin, Zhenglin Liu, Gang Qu 0001
IEEE Trans. Computers6
2023 LightWarner: Predicting Failure of 3D NAND Flash Memory Using Reinforcement Learning
abstract
NAND flash memory has gained popularity in a wide variety of digital storage systems. Although with excellent performance, NAND flash memory suffers various reliability problems. In recent years, researchers try to predict flash failure by using machine-learning models. However, the application of machine-learning based failure prediction method faces the following problems: imbalance between robustness and portability. When applying on different flash chips, the performance of prediction model degrades with the variation of error characteristics. In order to adapt to the variation, the machine-learning model needs to be re-built to ensure performance of failure prediction. The overheads of re-building model result in challenges when adjusting prediction model to adapt to the variation of error characteristics. To overcome these challenges, we present LightWarner, an easily applicable predictor based on model-free Reinforcement learning algorithms. LightWarner learns error characteristics dynamically during flash lifetime without pre-training. We evaluate the performance of LightWarner on six types of 3D flash chips. The evaluation result shows that LightWarner achieves over 93% F1 score on different flash chips, which is about 10% higher than supervised machine learning methods. And LightWarner can adapt to the variation of error characteristics with low migration costs.
Yuqian Pan, Haichun Zhang, Zhaojun Lu, Zhenglin Liu
IEEE Trans. Computers6
2022 Fooling the Eyes of Autonomous Vehicles: Robust Physical Adversarial Examples Against Traffic Sign Recognition Systems
Zhaojun Lu, Haichun Zhang, Zhenglin Liu, Jie Wang 0001, Gang Qu 0001
NDSS4
2020 Unexpected Error Explosion in NAND Flash Memory: Observations and Prediction Scheme
abstract
Wear-out has been a critical reliability problem in NAND flash memory. As executing repeated program and erase operations on the NAND flash chips, the number of errors increases and ultimately exceeds the ECC capability. In previous work, error characteristics of flash wear-out are observed by endurance tests on a single type of NAND flash memory. We wonder if the experimental results cover the entire error characteristics of NAND flash memory. In this paper, we tested more than 20 types of NAND flash chips with different vendors and structures and presented an overlook of test results. Through the test results, we found an unexpected error-explosion phenomenon that errors of flash blocks first increase over several cycles and then reach a high value without warning. We analyzed the features of the error-explosion and explored its influence on operation time. And we propose an error-explosion prediction scheme to find the blocks that will occur an error-explosion in the next 1000 P/E cycles. The block identifying operation is realized by the machine-learning model. The performance of six machine-learning methods is compared. The results demonstrate that the Decision Trees and Bagged Classification Trees have the best accuracy.
Yuqian Pan, Haichun Zhang, Mingyang Gong, Zhenglin Liu
ATS4
2020 Process-variation Effects on 3D TLC Flash Reliability: Characterization and Mitigation Scheme
abstract
In Solid State Drives, flash management techniques such as wear-leveling and refresh usually assume NAND flash memories have the same endurance value. However, the actual endurance values differ from blocks to blocks. This reliability difference is introduced by process-variation during flash fabrication. In recent years, for improving flash management techniques, various works have been done on the reliability variation of 2D flash memory. As 2D NAND transmitted to 3D NAND flash, the vertical structure and multi-layer stacking changed the effect of previously known reliability problems. In this paper, we are first to characterize the process-variation effects on 3D TLC flash reliability. The characterization includes two parts: endurance variation and error feature variation. Second, we propose an adaptive error prediction scheme to mitigate the process-variation effects. This scheme uses the machine-learning model to realize the error prediction operation. We also discuss the implications of this scheme on main flash management techniques.
Yuqian Pan, Haichun Zhang, Mingyang Gong, Zhenglin Liu
QRS4
2019 Superposed Compensation Strategy to Optimize Load/Line Transient Response and Reference Tracking for Discontinuous Conduction Mode Boost Converter
abstract
For boost converter operating in the discontinuous conduction mode (DCM), feedback and feedforward compensators are widely used to improve the converter performance. However, it is relatively difficult to optimize load transient response (LoTR), line transient response (LiTR), and reference tracking speed (RTS) simultaneously, since the optimizations require different compensators that are incompatible. In order to solve the issue, a superposed compensation strategy is proposed in this paper, which consists of a feedback compensator and two feedforward compensators. Each compensator is tuned according to an objective transfer function, which optimizes LoTR, LiTR, and RTS. The outputs are summed as duty cycle according to the linear superposition principle. Compatibility of the compensators is improved by designing the feedforward compensators to adapt to the feedback compensator. Furthermore, based on the closed-loop model, design rules for the objective transfer functions are given to minimize the influences of the sample-and-hold effect and calculation delay, which are intrinsic in a digital controller. Finally, converter's LoTR, LiTR, and RTS are simultaneously optimized, which is proven by closed-loop magnitude-frequency plots, state trajectory analyses, and experimental results.
Run Min, Dian Lyu, Linkai Li, Qiaoling Tong, Xuecheng Zou, Zhenglin Liu
IEEE Trans. Ind. Informatics7
2019 A Survey on Recent Advances in Vehicular Network Security, Trust, and Privacy
abstract
Vehicular ad hoc networks (VANETs) are becoming the most promising research topic in intelligent transportation systems, because they provide information to deliver comfort and safety to both drivers and passengers. However, unique characteristics of VANETs make security, privacy, and trust management challenging issues in VANETs' design. This survey article starts with the necessary background of VANETs, followed by a brief treatment of main security services, which have been well studied in other fields. We then focus on an in-depth review of anonymous authentication schemes implemented by five pseudonymity mechanisms. Because of the predictable dynamics of vehicles, anonymity is necessary but not sufficient to thwart tracking an attack that aims at the drivers' location profiles. Thus, several location privacy protection mechanisms based on pseudonymity are elaborated to further protect the vehicles' privacy and guarantee the quality of location-based services simultaneously. We also give a comprehensive analysis on various trust management models in VANETs. Finally, considering that current and near-future applications in VANETs are evaluated by simulation, we give a much-needed update on the latest mobility and network simulators as well as the integrated simulation platforms. In sum, this paper is carefully positioned to avoid overlap with existing surveys by filling the gaps and reporting the latest advances in VANETs while keeping it self-explained.
Zhaojun Lu, Gang Qu 0001, Zhenglin Liu
IEEE Trans. Intell. Transp. Syst.3
2019 A Blockchain-Based Privacy-Preserving Authentication Scheme for VANETs
abstract
The privacy-preserving authentication is considered as the first line of defense against the attacks in addition to preserving the identity privacy of the vehicles in the vehicular ad hoc networks (VANETs). However, the existing authentication schemes suffer from drawbacks such as nontransparency of the trusted authorities (TAs), heavy workload to revoke certificates, and high computation overhead to authenticate identities and messages. In this paper, we propose a blockchain-based privacy-preserving authentication (BPPA) scheme for VANETs. In BPPA, all the certificates and transactions are recorded permanently and immutably in the blockchain to make the activities of the semi-TAs transparent and verifiable. However, it remains a challenge how to use such blockchain effectively for authentication in real driving scenarios (e.g., high speed or large amount of messages during congestion). With a novel data structure named the Merkle Patricia tree (MPT), we extend the conventional blockchain structure to provide a distributed authentication scheme without the revocation list. To achieve conditional privacy, we allow a vehicle to use multiple certificates. The linkability between the certificates and real identity is encrypted and stored in the blockchain and can only be revealed in case of disputes. We evaluate the validity and performance of BPPA on the Hyperledger Fabric (HLF) platform for each entity. The experimental results show that the distributed authentication can be processed by individual vehicles within 1 ms, which meets the real-time requirement and is much more efficient, in terms of the processing time and storage requirement, than existing approaches.
Zhaojun Lu, Qian Wang 0022, Gang Qu 0001, Haichun Zhang, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Chaotic Encrypted Polar Coding Scheme for General Wiretap Channel
abstract
A wiretap channel is an important model for wireless communication. By applying an extended multiblock polar coding scheme, recent literature has achieved the secrecy capacity of a general wiretap channel (not necessary degraded or symmetric). However, this secure polar coding scheme of physical layer also limits the transmission rate of the main channel, which may fail to meet the demand of high transmission rate and strong transmission security for practical wireless transmission. In order to obtain a higher secrecy transmission rate than the physical layer coding scheme over a general wiretap channel, a cross-layer encryption and coding scheme is proposed in this paper. In the proposed scheme, an onetime-pad encryption and a secure key transmission is constructed by combining a chaos stream cipher with the extended multiblock polar coding scheme. As proved, the proposed scheme has achieved a high secrecy transmission rate than the former physical layer coding scheme under the constraints of reliability and strong security for a general wiretap channel.
Yizhi Zhao, Xuecheng Zou, Zhaojun Lu, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Hardware IP Protection through Gate-Level Obfuscation
abstract
Hardware Intellectual Property (IP) cores have emerged as an integral part of modern System-on-Chip (SoC) designs. However, recent trends of reverse engineering pose major threat to IP-based SOC design flow. The paper proposes a novel approach for hardware IP protection using gate-level obfuscation, which could make design less intelligible in order to neutralize or weaken the effect of reverse engineering. The basic idea is to hide the original logic function by using Physical Unclonable Function (PUF), multiplexer and configurable logic, so that it is difficult for reverse engineering attackers to get complete information of circuit net list. The design methodology could be applied in combinational logic and sequential logic. Simulation results on several IP cores show that we can achieve high levels of security through a well-formulated obfuscation scheme at less than 10% area overhead under delay constraint.
Wenchao Liu 0003, Xuecheng Zou, Zhenglin Liu
CAD/Graphics4
2015 Efficient Off-Chip Memory Protection Mechanism for Embedded Computing Systems Using AES-GCM
abstract
Off-chip memory security has become a prime concern in embedded computing systems due to the requirement of storing a large amount of potentially sensitive information in them. Existing solutions have performance imperfection because of their deployment of hash tree or unaffordable on-chip memory overhead. In this paper, we propose an efficient off-chip memory protection mechanism based on Advanced Encryption Standard - Galois/Counter Mode (AES-GCM) to provide both confidentiality and integrity protection for data and programs transferred from processor to off-chip memory in embedded computing systems. Our proposal is a novel memory protection mechanism: in order to ensure security and minimize on-chip memory overhead, AES-GCM hardware engine is running and dynamically switching between two modes, one mode for processing data and programs (DP mode), the other mode for processing the cryptographic parameter of IV (IV mode). It can resist well-known physical attacks, including replay attacks, relocation attacks and spoofing attacks. We demonstrate that our memory protection mechanism incurs as little as 1.56% on-chip memory overhead and has an average performance decline of about 9.0%.
Zhaojun Lu, Xiaoliang Xing, Qiaoling Tong, Zhenglin Liu
CAD/Graphics4
2007 On the Ability of AES S-Boxes to Secure Against Correlation Power Analysis
Zhenglin Liu, Xuecheng Zou
ISPEC1
2002 An Analytical Model for the Performance of the DOCSIS CATV Network
abstract
The Data Over Cable Service Interface Specification (DOCSIS) is established as the primary cable network data communication standard inNorth America and is the first specification having reached international standard level. It is therefore appropriate to determine its performance. Previous works on performance analysis of the DOCSIS Community Antenna Television network have been carried out using computer simulations. In this paper, we propose a model for mini-slot allocation with two priority classes of traffic which supports a real-time Constant Bit Rate (CBR) service and a non-real-time data service. Real-time CBR traffic has pre-emptive high-over-line priority over non-real-time data traffic. The service discipline is First Come First Served. Two critical characteristics, average delay and throughput, are analysed. Finally, numerical results are presented. The results demonstrate that the delay requirement of real-time CBR traffic can be easily met at the expense of the delay of non-real-time data traffic.
Zhenglin Liu, Chongyang Xu
Comput. J.1