Yanjiang Liu

dblp:240/8818 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0003-1806-6748ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 PEDC: A High-Efficacy, Parallel, and Configurable Distributed Control Architecture Design Approach of Cryptographic CGRA Utilizing the Control Subgraph Partitioning and Expansion Method
abstract
Coarse-grained reconfigurable architectures (CGRAs) are increasingly employed as cryptographic accelerators due to their efficiency and flexibility. Existing studies on security-oriented CGRAs primarily focus on scaling or optimizing the data path, while comparatively little attention has been given to the control path. Recently, a general-purpose processing core or configuration system has been widely adopted as the controller of CGRA. While this approach significantly reduces the design and application complexity of CGRAs, it does so at the expense of control efficacy and flexibility. To address this issue, a configurable distributed control (D-C) architecture design approach (referred to as PEDC) is proposed, which enhances the parallel processing capability of CGRA by improving control flexibility. There are four key technologies in PEDC. First, control subgraphs are automatically partitioned to define the control scopes of controllers and extract the control nodes of the control framework. Second, control dependency relationships are extracted from the control flow graph to link the control nodes. Third, a control architecture graph is constructed by establishing master-slave relationships between process control nodes capable of independent task process control and other nodes. Lastly, design models for the input scheduling controller, output scheduling controller, cluster controller, and task process controller are presented. The PEDC approach essentially transforms the CGRA into a multi-instruction stream, multi-data stream processor. D-C architectures with various scales are implemented based on 40-nm CMOS technology. With the PEDC approach, multiple pipelines and independent tasks can be processed simultaneously, regardless of the array structure or algorithm type. Compared with the traditional control method, the PEDC achieves a 4.5 × execution efficiency. Compared with related reconfigurable architectures, PEDC enables CGRAs to be more functionally flexible and achieve a better full-load throughput.
Zibin Dai, Yanjiang Liu, Danping Jiang, Zhaoxu Zhou
ACM Trans. Embed. Comput. Syst.3
2026 DPTM: An Adaptive Scheduler Design Utilizing Timeslot Matching and Release Methods for Concurrent and Multi-task Interleaved Pipelining-oriented CGRA
abstract
Coarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. However, the existing CGRA architectures suffer from low hardware resource utilization and performance due to the limitations of the scheduling scheme. In this article, an adaptive scheduler (denoted as DPTM) for concurrent and multi-task interleaved pipelining-oriented CGRA is introduced, which exploits timeslot matching and release methods to avoid the pipeline conflicts and improve the scheduling performance. The characteristics of task scheduling based on directed acyclic graph (DAG) are analyzed, and several performance-influencing factors are extracted to build a scheduling performance model for reducing the time cost of scheduling and guiding the design of scheduling schemes. Moreover, the scoreboard method of dynamic instruction schedulers is optimized to control the entry time of multiple tasks into the pipeline, and then a timeslot matching method is proposed to provide non-conflict pipelining for the multiple tasks. Further, a timeslot release method is presented to release the timeslots for unscheduled sub-tasks dynamically, which can adapt the parallel processing of multiple tasks and decrease the scheduling time. Then, an adaptive scheduling scheme combines the dynamic priority-based task assignment method, timeslot matching method, and timeslot release method to schedule massive tasks for CGRA. Finally, the overall architecture of DPTM is introduced and designed to validate the efficacy of the proposed scheduling scheme. Experimental results show that the proposed timeslot matching/release approach reduces 84% total scheduling time and decreases 40% average scheduling time at most compared to the non-timeslot-matching scheduling schemes, the proposed task assignment approach decreases 8% total scheduling time and lowers 3% average scheduling time compared to the existing approaches, and the proposed scheduler decreases 51% critical path delay, lowers 35% area overhead, and reduces 12% power consumption at most compared with the existing schedulers.
Danping Jiang, Zibin Dai, Yanjiang Liu, Zhaoxu Zhou
ACM Trans. Design Autom. Electr. Syst.3
2025 ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
abstract
Multimodal Large Language Models (MLLMs) suffer from high computational costs due to their massive size and the large number of visual tokens. In this paper, we investigate layer-wise redundancy in MLLMs by introducing a novel metric, Layer Contribution (LC), which quantifies the impact of a layer's transformations on visual and text tokens, respectively. The calculation of LC involves measuring the divergence in model output that results from removing the layer's transformations on the specified tokens. Our pilot experiment reveals that many layers of MLLMs exhibit minimal contribution during the processing of visual tokens. Motivated by this observation, we propose ShortV, a training-free method that leverages LC to identify ineffective layers, and freezes visual token updates in these layers. Experiments show that ShortV can freeze visual token in approximately 60\% of the MLLM layers, thereby dramatically reducing computational costs related to updating visual tokens. For example, it achieves a 50\% reduction in FLOPs on LLaVA-NeXT-13B while maintaining superior performance. The code will be publicly available at https://github.com/icip-cas/ShortV
Qianhao Yuan, Yanjiang Liu, Jiawei Chen 0011, Yaojie Lu 0001, Jia Zheng 0009, Xianpei Han, Le Sun 0001
ICCV3
2025 CRM_BF: A Low-Overhead, High-Efficient and Reconfigurable Operation Unit Design Approach Using the Customized Reed-Muller Unit For Boolean Functions of Sequence Cipher Algorithms
abstract
Sequence ciphers algorithms encrypt or decrypt information at a low cost and high speed compared to other cryptographic algorithms, which are widely applied to critical applications and sensitive fields. As the core component of sequence ciphers, Boolean functions generate the random number or implement the update process of random numbers. The existing implementations of Boolean functions cause a great waste of area resources and generate several long critical paths that limit the hardware performance of sequence ciphers. To address this issue, a 64-bit Boolean Function Reconfigurable Operation Unit (BFROU) is proposed to reduce the area overhead, lower the delay latency, and enhance the operation efficacy of Boolean functions. Through statistical characterization analysis and cutting experiments of Boolean functions, a 64 bits BFROU based on CRM-3 units has been designed, which has the advantage of low-cost and high-efficient。The CRM unit is customized based on RM logic. A theoretical framework for Boolean functions is proposed by combining CRM units with mathematical expressions, which encompasses Boolean functions for any variable. On the platform of synthesis software, based the theoretical architecture, a CRM-OPT optimization algorithm is proposed, which can achieve the conversion of And Inverter Graph (AIG) to Customized Reed Muller Graph (CRMG).This Customized Reed-Muller (CRM) unit achieved at least 22.4% and 25.1% optimization in delay and area compared to Universal Reed-Muller (URM) units. The experimental results show that the Area Delay Product (ADP) is minimized when the CRM-3 unit is the optimal maximum cutting size. Ultimately, the BFROU design was realized utilizing CRM units, achieving an area of 195.4um² and a critical path delay of 0.35ns. This BFROU can achieve special Boolean functions involving 64 variables at maximum,with 91% of these functions being mapped within two iterations. Moreover, this BFROU has significant advantages over other known schemes regarding area, critical path delay, ADP, and number of iterations consumed.
Zhaoxu Zhou, Junwei Li 0007, Yanjiang Liu, Zibin Dai
ACM Trans. Design Autom. Electr. Syst.4
2024 Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting
abstract
Incorporating factual knowledge in knowledge graph is regarded as a promising approach for mitigating the hallucination of large language models (LLMs). Existing methods usually only use the user's input to query the knowledge graph, thus failing to address the factual hallucination generated by LLMs during its reasoning process. To address this problem, this paper proposes Knowledge Graph-based Retrofitting (KGR), a new framework that incorporates LLMs with KGs to mitigate factual hallucination during the reasoning process by retrofitting the initial draft responses of LLMs based on the factual knowledge stored in KGs. Specifically, KGR leverages LLMs to extract, select, validate, and retrofit factual statements within the model-generated responses, which enables an autonomous knowledge verifying and refining procedure without any additional manual efforts. Experiments show that KGR can significantly improve the performance of LLMs on factual QA benchmarks especially when involving complex reasoning processes, which demonstrates the necessity and effectiveness of KGR in mitigating hallucination and enhancing the reliability of LLMs.
Xinyan Guan, Yanjiang Liu, Yaojie Lu 0001, Ben He 0001, Xianpei Han, Le Sun 0001
AAAI2
2024 A Feature-Adaptive and Scalable Hardware Trojan Detection Framework For Third-party IPs Utilizing Multilevel Feature Analysis and Random Forest
Yanjiang Liu, Junwei Li 0007, Chunsheng Zhu, Jingxin Zhong
J. Electron. Test.1
2024 RGMU: A High-flexibility and Low-cost Reconfigurable Galois Field Multiplication Unit Design Approach for CGRCA
abstract
Finite field multiplication is a non-linear transformation operator that appears in the majority of symmetric cryptographic algorithms. Numerous specified finite field multiplication units have been proposed as a fundamental module in the coarse-grained reconfigurable cipher logic array to support more cryptographic algorithms; however, it will introduce low flexibility and high overhead, resulting in reduced performance of the coarse-grained reconfigurable cipher logic array. In this article, a high-flexibility and low-cost reconfigurable Galois field multiplication unit (RGMU) is proposed to balance the tradeoffs between the function, delay, and area. All the finite field multiplication operations, including maximum distance separable matrix multiplication, parallel update of Fibonacci linear feedback shift register, parallel update of Galois linear feedback shift register, and composite field multiplication, are analyzed and two basic operation components are abstracted. Further, a reconfigurable finite field multiplication computational model is established to demonstrate the efficacy of reconfigurable units and guide the design of RGMU with high performance. Finally, the overall architecture of RGMU and two multiplication circuits are introduced. Experimental results show that the RGMU can not only reduce the hardware overhead and power consumption but also has the unique advantage of satisfying all the finite field multiplication operations in symmetric cryptography algorithms.
Danping Jiang, Zibin Dai, Yanjiang Liu, Zongren Zhang
ACM Trans. Design Autom. Electr. Syst.3
2023 Towards a metrics suite for evaluating cache side-channel vulnerability: Case studies on an open-source RISC-V processor
Yingjian Yan, Jingxin Zhong, Yanjiang Liu, Jinsong Xu
Comput. Secur.5
2023 A High-performance Masking Design Approach for Saber against High-order Side-channel Attack
abstract
Post-quantum cryptography (PQC) has become the most promising cryptographic scheme against the threat of quantum computing to conventional public-key cryptographic schemes. Saber, as the finalist in the third round of the PQC standardization procedure, presents an appealing option for embedded systems due to its high encryption efficiency and accessibility. However, side-channel attack (SCA) can easily reveal confidential information by analyzing the physical manifestations, and several works demonstrate that Saber is vulnerable to SCAs. In this work, a ciphertext comparison method for masking design based on the bitslicing technique and zerotest is proposed, which balances the tradeoff between the performance and security of comparing two arrays. The mathematical description of the proposed ciphertext comparison method is provided, and its correctness and security metrics are analyzed under the concept of PINI. Moreover, a high-order masking approach based on the state of the art, including the hash functions, centered binomial sampling, masking conversions, and proposed ciphertext comparison, is presented, using the bitslicing technique to improve throughput. As a proof of concept, the proposed implementation of Saber is on the ARM Cortex-M4. The performance results show that the runtime overhead factor of 1st-, 2nd-, and 3rd-order masking is 3.01×, 5.58×, and 8.68×, and the dynamic memory used for 1st-, 2nd-, and 3rd-order masking is 17.4kB, 24.0kB, and 30.2kB, respectively. The SCA-resilience evaluation results illustrate that the 1st-order Test Vectors Leakage Assessment (TVLA) result fails to reveal the secret key with 100,000 traces.
Yajing Chang, Yingjian Yan, Chunsheng Zhu, Yanjiang Liu
ACM Trans. Design Autom. Electr. Syst.4
2023 CBDC-PUF: A Novel Physical Unclonable Function Design Framework Utilizing Configurable Butterfly Delay Chain Against Modeling Attack
abstract
Physical unclonable function (PUF) is a promising security-based primitive, which provides an extremely large number of responses for key generation and authentication applications. Various PUFs have been developed as central building blocks in cryptographic protocols and security architectures, however, the existing PUFs and their improvements are still vulnerable to modeling attacks (MA) with refined machine learning algorithms. In this article, a configurable butterfly delay chain-based PUF design framework is proposed to meet the requirements of randomness, reliability, uniqueness, and MA-resistance metrics. A configurable butterfly delay chain is introduced to create multiple pairs of symmetric paths and a strong PUF relying on the intrinsic delay fluctuations of two identical paths is built. Furthermore, a secure hash function is used to insert non-linearities into the PUF, and a BCH-based error correction algorithm is utilized to recover the actual responses under noisy environments. The proposed PUF is implemented on Xilinx FPGAs and three machine learning algorithms are used to evaluate the resistance against MA. Experimental results show that the randomness, reliability, and uniqueness of the proposed PUF are close to the ideal value (49.6%, 99.9%, and 49.9%, respectively), and the prediction accuracy reaches 50% that indicating a desirable resilient to MA.
Yanjiang Liu, Junwei Li 0007, Tongzhou Qu, Zibin Dai
ACM Trans. Design Autom. Electr. Syst.1
2022 A Comprehensive Evaluation of Integrated Circuits Side-Channel Resilience Utilizing Three-Independent-Gate Silicon Nanowire Field Effect Transistors-Based Current Mode Logic
abstract
Side-channel attack (SCA) is one of the physical attacks, which will reveal the confidential information from cryptographic circuits by statistically analyzing physical manifestations. Various circuit-level countermeasures have been proposed as fundamental solutions to eliminate the correlations between side-channel information and circuit’s internal operations. The existing solutions, however, will introduce nonnegligible power and area overheads, making them difficult to be deployed in resource-constrained applications. In this article, a novel three-independent-gate silicon nanowire field effect transistor (TIGFET) with the intrinsic SCA-resilience characteristics is introduced to balance the tradeoffs among cost, performance, and security of cryptographic implementations. We construct six TIGFET-based current mode logic (CML) gates that can retain lower power variation under all possible transitions compared to the CMOS counterparts. As a proof of concept, advanced encryption standard (AES), SM4 block cipher algorithm (SM4), and lightweight cryptographic algorithm PRESENT are implemented utilizing the TIGFET-based CML gates. Correlation power attack is performed to evaluate the improvement of SCA resilience. Simulation results verify that the TIGFET-based cryptographic implementations decrease 42.37% area usage, lower 61.16% energy efficiency, reduce$5.35\times $power variation, and achieve a similar level of SCA resistance compared to the CMOS counterpart, which is applicable for the resource-constrained applications.
Yanjiang Liu, Jiaji He 0001, Haocheng Ma, Tongzhou Qu, Zibin Dai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 A Low-Overhead and High-Security Cryptographic Circuit Design Utilizing the TIGFET-Based Three-Phase Single-Rail Pulse Register against Side-Channel Attacks
abstract
Side-channel attack (SCA) reveals confidential information by statistically analyzing physical manifestations, which is the serious threat to cryptographic circuits. Various SCA circuit-level countermeasures have been proposed as fundamental solutions to reduce the side-channel vulnerabilities of cryptographic implementations; however, such approaches introduce non-negligible power and area overheads. Among all of the circuit components, flip-flops are the main source of information leakage. This article proposes a three-phase single-rail pulse register (TSPR) based on the three-independent-gate field effect transistor (TIGFET) to achieve all desired properties with improved metrics of area and security. TIGFET-based TSPR consumes a constant power (MCV is 0.25%), has a low delay (12 ps), and employs only 10 TIGFET devices, which is applicable for the low-overhead and high-security cryptographic circuit design compared to the existing flip-flops. In addition, a set of TIGFET-based combinational basic gates are designed to reduce the area occupation and power consumption as much as possible. As a proof of concept, a simplified advanced encryption algorithm (AES), SM4 block cipher algorithm (SM4), and light-weight cryptographic algorithm (PRESENT) are built with the TIGFET-based library. SCA is implemented on the cryptographic implementations to prove its SCA resilience, and the SCA results show that the correct key of cryptographic circuits with TIGFET-based TSPRs is not guessed within 2,000 power traces.
Yanjiang Liu, Tongzhou Qu, Zibin Dai
ACM Trans. Design Autom. Electr. Syst.1
2021 On-Chip Trust Evaluation Utilizing TDC-Based Parameter-Adjustable Security Primitive
abstract
Field-programmable gate arrays (FPGAs) are integrated circuits (ICs) that can be reconfigured to the desired functionalities, without manufacturing dedicated chips. Due to their programmable nature, FPGAs have been prevalent in the large majority of modern systems. This raises high demands for verifying the security of circuit implementations on FPGAs, since they are vulnerable to hardware trojans (HTs) that can be inserted through modified configuration files. In this article, we propose an on-chip security framework to ensure the trustworthiness of circuit implementations on FPGAs at runtime. The core of the framework is a time-to-digital converter (TDC)-based hardware security primitive that can be predeployed on FPGAs to verify whether the FPGA-based designs are tampered with or corrupted by HTs. The parameter-adjustable TDC sensor, which is the primary component of the primitive, is carefully designed, adjusted, and implemented, thus the TDC sensor can monitor the transient voltage fluctuations within FPGAs with a high resolution. Versus statistical data analysis, tiny abnormal variations introduced by the Trojan insertion and activation are distinguished. Experimental results on Xilinx Spartan-6 FPGAs demonstrate the effectiveness of the proposed TDC-based on-chip trust evaluation framework and HT detection method.
Haocheng Ma, Jiaji He 0001, Yanjiang Liu, Jun Kuai, He Li 0008, Leibo Liu, Yiqiang Zhao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Security-Driven Placement and Routing Tools for Electromagnetic Side-Channel Protection
abstract
Side-channel analysis (SCA) attacks are major threats to hardware security. Upon this security threat, various countermeasures at different design layers have been proposed against SCA attacks. These approaches often introduce significant overheads and impose high requirements of side-channel security backgrounds to integrated circuit (IC) designers. In this article, we propose an automatic computer-aided design (CAD) tool that can be utilized to enhance the circuit resistance against electromagnetic (EM) SCA attacks. This new tool will guide security-driven placement and routing processes and can be seamlessly integrated into the modern IC design flow. The protected IC design will be resilient to SCA attacks with negligible area and power overheads. In order to develop this tool, we first investigate the root-cause of EM leakage at the layout level and mathematically demonstrate the feasibility of security-driven placement and routing through the EM leakage modeling. We then identify that the correlation between the data under protection and the EM leakage can be significantly reduced through data-dependent register reallocation and wire length adjustments. Simulation results on cryptographic circuits prove the effectiveness of both the constructed EM leakage model and the EM model-based CAD tool for EM side-channel security.
Haocheng Ma, Jiaji He 0001, Yanjiang Liu, Leibo Liu, Yiqiang Zhao, Yier Jin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Golden Chip-Free Trojan Detection Leveraging Trojan Trigger's Side-Channel Fingerprinting
abstract
Hardware Trojans (HTs) have become a major threat for the integrated circuit industry and supply chain and have motivated numerous developments of HT detection schemes. Although the side-channel HT detection approach is among the most promising solutions, most of the previous methods require a trusted golden chip reference. Furthermore, detection accuracy is often influenced by environmental noise and process variations. In this article, a novel electromagnetic (EM) side-channel fingerprinting-based HT detection method is proposed. Different from previous methods, the proposed solution eliminates the requirement of a trusted golden fabricated chip. Rather, only the genuine RTL code is required to generate the EM signatures as references. A factor analysis method is utilized to extract the spectral features of the HT trigger’s EM radiation, and then a k -means clustering method is applied for HT detection. Experimentation on two selected sets of Trust-Hub benchmarks has been performed on FPGA platforms, and the results show that the proposed framework can detect all dormant HTs with a high confidence level.
Jiaji He 0001, Haocheng Ma, Yanjiang Liu, Yiqiang Zhao
ACM Trans. Embed. Comput. Syst.3
2021 Test Generation for Hardware Trojan Detection Using Correlation Analysis and Genetic Algorithm
abstract
Hardware Trojan (HT) is a major threat to the security of integrated circuits (ICs). Among various HT detection approaches, side channel analysis (SCA)-based methods have been extensively studied. SCA-based methods try to detect HTs by comparing side channel signatures from circuits under test with those from trusted golden references. The pre-condition for SCA-based HT detection to work is that the testers can collect extra signatures/anomalies introduced by activated HTs. Thus, activation of HTs and amplification of the differences between circuits under test and golden references are the keys to SCA-based HT detection methods. Test vectors are of great importance to the activation of HTs, but existing test generation methods have two major limitations. First, the number of test vectors required to trigger HTs is quite large. Second, the HT circuit’s activities are marginal compared with the whole circuit’s activities. In this article, we propose an optimized test generation methodology to assist SCA-based HT detection. Considering the HTs’ inherent surreptitious nature, inactive nodes with low transition probability are more likely to be selected as HT trigger nodes. Therefore, the correlations between circuit inputs and inactive nodes are first exploited to activate HTs. Then a test reordering process based on the genetic algorithm (GA) is implemented to increase the proportion of the HT circuit’s activities to the whole circuit’s activities. Experiments on 10 selected ISCAS benchmarks, wb_conmax benchmark, and b17 benchmark demonstrate that the number of test vectors required to trigger HTs reduces 28.8% on average compared with the result of MERO and MERS methods. After the test vector reordering process, the proportion of the HT circuit’s activities to the whole circuit’s activities is improved by 95% on average, compared with the result of MERS method.
Zhendong Shi, Haocheng Ma, Qizhi Zhang 0001, Yanjiang Liu, Yiqiang Zhao, Jiaji He 0001
ACM Trans. Embed. Comput. Syst.4
2020 Runtime Trust Evaluation and Hardware Trojan Detection Using On-Chip EM Sensors
abstract
It has been widely demonstrated that the utilization of postdeployment trust evaluation approaches, such as side-channel measurements, along with statistical analysis methods is effective for detecting hardware Trojans in fabricated integrated circuits (ICs). However, more sophisticated Trojans proposed recently invalidate these methods with stealthy triggers and very-low side-channel signatures. Upon these challenges, in this paper, we propose an electromagnetic (EM) side-channel based post-fabrication trust evaluation framework which monitors EM radiations at runtime. The key component of the runtime trust evaluation framework is an on-chip EM sensor which can constantly measure and collect EM side-channel information of the target circuit. The simulation results validate the capability of the proposed framework in detecting stealthy hardware Trojans. Further, we fabricate an AES circuit protected by the proposed trust evaluation framework along with four different types of hardware Trojans. The measurements on the fabricated chips prove two key findings. First, the on-chip EM sensor can achieve a higher signal to noise ratio (SNR) and thus facilitate a better Trojan detection accuracy. Second, the trust evaluation framework can help detect different hardware Trojans at runtime.
Jiaji He 0001, Xiaolong Guo 0001, Haocheng Ma, Yanjiang Liu, Yiqiang Zhao, Yier Jin
DAC4
2020 Golden chip free Trojan detection leveraging probabilistic neural network with genetic algorithm applied in the training phase
Yanjiang Liu, Jiaji He 0001, Haocheng Ma, Yiqiang Zhao
Sci. China Inf. Sci.1
2019 Hardware Trojan Detection Leveraging a Novel Golden Layout Model Towards Practical Applications
Yanjiang Liu, Jiaji He 0001, Haocheng Ma, Yiqiang Zhao
J. Electron. Test.1