Chenchen Deng

dblp:137/1419 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-1047-1087ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
Reconfigurable computing and FPGAs · 28% Memory systems · 20% Interconnection networks and networks-on-chip · 14%
Network and information security
4 papers
Hardware security and side channels · 62% Cryptographic primitives and cryptanalysis · 28% Privacy and data protection · 10%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
1.232023
M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution · DAC 2017
TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-Instructions · IEEE Trans. Parallel Distributed Syst. 2016
Integrated circuit design
3d integration
0.912025
Software-defined process-near-memory architecture using 3D hybrid bonding integration · Sci. China Inf. Sci. 2025
Memory systems
processing-in-memory
0.912025
Software-defined process-near-memory architecture using 3D hybrid bonding integration · Sci. China Inf. Sci. 2025
Memory systems
content-addressable memory
0.812024
CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024
Memory systems › content-addressable memory
TCAM
0.812024
CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024
Reconfigurable computing and FPGAs
reconfigurable architecture
0.732018
Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-Instructions · IEEE Trans. Parallel Distributed Syst. 2016
A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems · IEEE Trans. Parallel Distributed Syst. 2017
Distributed systems
fault tolerance
0.712023
M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Hardware reliability and fault tolerance
redundancy
0.712023
M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Hardware security and side channels
side-channel countermeasures
0.612022
An energy-efficient dynamically reconfigurable cryptographic engine with improved power/EM-side-channel-attack resistance · Sci. China Inf. Sci. 2022
Hardware security and side channels
fault attacks
0.522017
Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017
Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016
Hardware security and side channels
fault attack countermeasure
0.522017
Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017
Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016
Reconfigurable computing and FPGAs
application mapping
0.522017
A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems · IEEE Trans. Parallel Distributed Syst. 2017
An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Interconnection networks and networks-on-chip › network-on-chip design
noc power management
0.412020
Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Energy-efficient computing
power gating
0.412020
Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Interconnection networks and networks-on-chip › network topology
reconfigurable topology
0.412020
CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology · DAC 2020
Reconfigurable computing and FPGAs › application mapping
reliability-aware mapping
0.422015
An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Reliability-aware mapping for various NoC topologies and routing algorithms under performance constraints · Sci. China Inf. Sci. 2015
Cryptographic primitives and cryptanalysis › cryptographic implementation
hardware implementation
0.312018
Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Cryptographic primitives and cryptanalysis
symmetric cryptography
0.312018
Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Processor architecture and microarchitecture
multithreading
0.312018
Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric
0.312018
Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Processor architecture and microarchitecture › pipelining
pipeline stall reduction
0.312017
Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution · DAC 2017
Privacy and data protection
randomization
0.212016
Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016
Integrated circuit design
low-power circuit design
0.212024
CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024
Energy-efficient computing › energy-efficient communication
communication energy minimization
0.212015
An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Energy-efficient computing
power management
0.212015
An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing
0.112020
Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Interconnection networks and networks-on-chip › network topology
torus network
0.112020
CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology · DAC 2020
Computer vision › Face, body and person analysis
face alignment
0.112017
A 700fps Optimized Coarse-to-Fine Shape Searching Based Hardware Accelerator for Face Alignment · DAC 2017
Cryptographic primitives and cryptanalysis
block cipher
0.112017
Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017
Computer vision › Face, body and person analysis
face detection
0.112016
A fast face detection architecture for auto-focus in smart-phones and digital cameras · Sci. China Inf. Sci. 2016

Methods — techniques the papers use, named apart from their topics

dynamic reconfiguration · 1.83d hybrid bonding · 1.7power/EM side-channel countermeasures · 1.1constant-time alteration · 0.828 nm CMOS · 0.8temporal-redundant voters · 0.7spatial-redundant data paths · 0.7markov process model · 0.7branch-and-bound · 0.5flit deflection · 0.4context compression · 0.3statistical evaluation · 0.3cascaded regression · 0.3benes network · 0.3SURF features · 0.3spatial and time randomization · 0.2injection effort model · 0.2
YearPublicationVenuePosition
2025 Software-defined process-near-memory architecture using 3D hybrid bonding integration
Anlin Xu, Chenchen Deng, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu
Sci. China Inf. Sci.2
2024 CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency
Chenchen Deng, Tianzhu Xiong, Zhaoshi Li, Jianfeng Zhu 0001, Jun Yang 0006, Shaojun Wei, Leibo Liu
Sci. China Inf. Sci.1
2023 M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures
abstract
Coarse-grained reconfigurable architectures (CGRAs) can provide both energy efficiency and performance for embedded systems, and thus they are increasingly deployed in the areas of aerospace, automotive engineering, and security where reliability is also a main criterion. However, the state-of-the-art fault-tolerant strategies for CGRAs apply either temporal or spatial scheme, including redundancy, periodic detection, workload balancing, and reconfiguration, failing to exploit the feature of dynamic and partial reconfiguration of CGRAs. Also, vulnerable judging circuits and inflexible mode shifting bottleneck the reliability design of fault-tolerant CGRAs. This article proposes a novel multimode fault-tolerant framework for CGRAs, which combines spatial-redundant data paths with temporal-redundant voters and thus reduces the vulnerable judging circuits while balancing the performance and reliability. This framework can also enable a changing reliability level at runtime via an online configuration transformation method based on precompiled patterns. Within the proposed framework, we systematically searched the design space spanning various combinations of the mainstream schemes with a Markov process model to compare the effectiveness and accordingly selected five points as available modes in our design after comprehensive consideration of fault tolerance and time overhead on CGRA. The framework is comprehensively evaluated on a cycle-accurate CGRA simulator, considering both permanent and transient faults. The experimental results show that the fault coverage rate of single transient faults or permanent faults has increased from 71.74% to 93.84%, which means the fault tolerance of the system has been increased by 31.03% compared with the state-of-the-art methods. There is also a great improvement in mean-time-to-failure (MTTF) and reconfiguration latency over baseline designs.
Jianfeng Zhu 0001, Xingchen Man, Guihuan Song, Yi Huang 0036, Chenchen Deng, Pengfei Gou, Shouyi Yin, Shaojun Wei, Leibo Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 An energy-efficient dynamically reconfigurable cryptographic engine with improved power/EM-side-channel-attack resistance
Chenchen Deng, Min Zhu 0001, Jinjiang Yang, Youyu Wu, Jiaji He 0001, Bohan Yang 0001, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei, Leibo Liu
Sci. China Inf. Sci.1
2021 LWRpro: An Energy-Efficient Configurable Crypto-Processor for Module-LWR
abstract
Saber, the only module-learning with rounding-based algorithm in NIST's third round of post-quantum cryptography (PQC) standardization process, is characterized by simplicity and flexibility. However, energy-efficient implementation of Saber is still under investigation since the commonly used number theoretic transform can not be utilized directly. In this manuscript, an energy-efficient configurable crypto-processor supporting multi-security-level key encapsulation mechanism of Saber, is proposed. First, an 8-level hierarchical Karatsuba framework is utilized to reduce degree-256 polynomial multiplication to the coefficient-wise multiplication. Second, a hardware-efficient Karatsuba scheduling strategy and an optimized pre-/post-processing structure is designed to reduce the area overheads of scheduling strategy. Third, a task-rescheduling-based pipeline strategy and truncated multipliers are proposed to enable fine-grained processing. Moreover, multiple parameter sets are supported in LWRpro to enable configurability among various security scenarios. Enabled by these optimizations, LWRpro requires 1066, 1456 and 1701 clock cycles for key generation, encapsulation, and decapsulation of Saber768. The post-layout version of LWRpro is implemented with TSMC 40 nm CMOS process within 0.38 mm2. The throughput for Saber768 is up to 275k encapsulation operations per second and the energy efficiency is 0.15 uJ/encapsulation while operating at 400 MHz, achieving nearly 50× improvement and 31× improvement, respectively compared with current PQC hardware solutions.
Yihong Zhu, Min Zhu 0001, Bohan Yang 0001, Wenping Zhu, Chenchen Deng, Chen Chen 0083, Shaojun Wei, Leibo Liu
IEEE Trans. Circuits Syst. I Regul. Pap.5
2020 CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology
abstract
Future NoCs should be highly flexible to adapt to communication demands to achieve high scalability and low power consumption. However, the flexibility is still quite limited by the high complexity of reconfiguration for globally reconfigured channels. In this paper, we propose to augment a router-based buffered NoC with a reconfigurable ring architecture by exploiting cycle decomposition of a torus bufferless network. At runtime, the topologies of the rings can be reconfigured according to the workloads by choosing different cycle decompositions of the torus network. Because the shapes of the rings are restricted to a specified regular shape, the reconfiguration time can be reduced to a linear complexity with respect to network size, and the reconfiguration algorithm can be implemented in a distributed hardware. The experimental results show that the reconfigurable rings provide 54% and 26% improvements on packet latency and static power saving, respectively, for realistic workloads.
Liang Wang 0020, Leibo Liu, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei
DAC5
2020 Aggressive Fine-Grained Power Gating of NoC Buffers
abstract
Power gating is effective for networks-on-chip (NoCs) to reduce the excessive leakage power dissipated by idle network components. Most existing NoC power-gating approaches rely on the routing algorithms to mitigate the power-gating blocking latency problem. When the network becomes faulty and fault-tolerant routing algorithms are applied, these approaches are no longer applicable or can seriously degrade the performance. Other approaches propose fine-grained buffer power gating, but they are too conservative in power saving due to the buffer backpressure flow control. To address these problems, we propose an aggressive fine-grained power gating of flit-sized buffer entries by adopting backpressureless flow control in an input-buffered network. The power-gating decisions are made based on the flit deflection rate. However, directly applying the backpressureless flow control leads to the difficulties of multiflit packet truncation and protocol deadlocks. Therefore, we modify the packet injection architecture to avoid packet truncation. This is done by chaining the local input port with a randomly chosen input port. Finally, we design a progressive recovery framework to handle both livelocks and protocol deadlocks. It does not need to truncate packets or strictly separate different message classes when the network is free of livelocks or protocol deadlocks. The experimental results show that with a hardware overhead of 9.6%, our design can save up to 59% network power consumption in both a fault-free and a faulty NoC with little zero-load latency penalty. Our design also approaches an ideal energy-proportional NoC because it can constantly reduce power consumption over a wide range of injection rates.
Leibo Liu, Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2019 A Reliable Physical Unclonable Function Based on Differential Charging Capacitors
abstract
Physical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit.
Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044
ISCAS7
2018 Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms
abstract
This paper presents a dynamically reconfigurable processing array named Anole for symmetric-key algorithms. Processing elements and the interconnections between them are designed to support various block and stream ciphers. Without affecting flexibility, three key techniques are presented to increase energy efficiency (throughput/power, the number of operations per unit energy consumption) and area efficiency (throughput/area). First, the distributed control network supports multithreading on reconfigurable fabrics at a low cost, thereby maximizing the utility of computing resources in the space domain. Second, the concurrent computation and reconfiguration scheme integrates configuration contexts with processing data to simultaneously execute in the data-path. The resulted immediate switching between different configurations increases the utilization rate of hardware resources in the temporal domain. Third, under configuration context compression and organization, the context memory size and configuration time are further minimized. Anole is implemented on a 7.75 mm2silicon square with TSMC 65-nm technology at 400 MHz. Experiments show that Anole significantly outperforms field programmable gate array and general purpose processor by more than two orders of magnitude in energy and area efficiencies. Compared with state-of-the-art reconfigurable solutions, Anole achieves (average) 16.5× higher energy efficiency and 9.4× higher area efficiency.
Leibo Liu, Bo Wang 0023, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution
abstract
The pipeline stall in distributed-controlled coarse-grained reconfigurable arrays is a major source stumbling performance. This work presents a Triggered-Issue and Triggered-Execution (TITE) paradigm motivated from the Triggered Instruction Architecture (TIA) which converts control and data dependencies into predicate dependencies as triggers for spatial parallelism. TITE separately triggers the issuing and execution of instructions to further relax the predicate dependencies in TIA. Triggered dual instructions and tag forwarding are proposed to minimize pipeline stalls of both intra and inter-processing elements. Experiments show that TITE improves performance, energy efficiency, and area efficiency by 21%, 17%, and 12%, respectively, compared with TIA.
Yanan Lu, Leibo Liu, Yangdong Deng, Jian Weng 0004, Zhaoshi Li, Chenchen Deng, Shaojun Wei
DAC6
2017 A 700fps Optimized Coarse-to-Fine Shape Searching Based Hardware Accelerator for Face Alignment
abstract
In this work, a fast shape searching face alignment (F-SSFA) algorithm based accelerator is proposed to achieve real-time processing. Firstly, a learning based low-dimensional SURF feature is introduced to reduce the computation cost in the cascaded regression. Then the Euclidean distance and shape affine transformation are utilized to accelerate the shape searching procedure. F-SSFA therefore greatly reduces the computational complexity while keeping the same accuracy. Also, a fixed-point F-SSFA based VLSI architecture is designed with approximately 80% decrease in the data transmission traffic. The throughput of this accelerator achieves 700 fps, which is especially suitable for high-speed facial-related applications.
Leibo Liu, Wenping Zhu, Huiyu Mo, Chenchen Deng, Shaojun Wei
DAC5
2017 Implementation of in-loop filter for HEVC decoder on reconfigurable processor
abstract
The in‐loop filter comprises deblocking filter and sample adaptive offset filter, which is an important module for improving image quality in a high‐efficiency video coding (HEVC) decoder. The in‐loop filter has a high computational complexity that accounts for ∼20% of the HEVC decoding computing load. Furthermore, it is difficult to implement a high‐performing in‐loop filter due to its large conditional processing requirement. First, this study presents a novel reconfigurable HEVC in‐loop filter implementation on a coarse‐grained dynamically reconfigurable processing unit. Next, a repartition scheme is presented that allows the in‐loop filter implementation at a coding tree unit along with the other decoding modules in the HEVC decoder, which satisfies requirements of low latency applications. Finally, a hierarchised‐pipeline and synchronised‐parallel technique is used to improve performance by eliminating data hazards in pipeline techniques and synchronisation problems in parallel techniques. Implementation results show that the presented HEVC in‐loop filter performs up to 1920 × 1080@52 frames per second at 250 MHz. The throughput is 67.5 × 9 × more than solutions based on digital signal processor and general‐purpose processor, respectively.
Leibo Liu, Victor Y. Chen, Chenchen Deng, Shouyi Yin, Shaojun Wei
IET Image Process.3
2017 Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks
abstract
Traditional detection countermeasures against fault attacks have been criticized as insecure because of the fragile comparison operation that can be maliciously bypassed. In order to avoid the comparison, infection countermeasures have been designed to confuse the faulty ciphertexts so that the output cannot be further explored. This paper presents an infection method that resists fault attacks using the existing Benes network module in high-performance crypto processors. The Benes network is originally used to accelerate permutation operations in block ciphers. The hamming weight of the differential results is balanced by modifying specific network switches, without changing the network topology. A further confusion is performed to destroy the determinacy by configuring part of the network with a random bit-stream. Furthermore, a statistical evaluation method is presented to quantitatively verify the proposed countermeasure in addition to a formal proof of security. This also provides a new concept for the evaluation of future random-enhanced infection methods. Experiments are carried out using Advanced Encryption Standard (AES), triple Data Encryption Standard (DES), and Camellia as examples. Under statistical evaluation, the results show that the proposed countermeasure improves the fault resistance by over four orders of magnitude compared with the unprotected case. Also, the performance and the area overhead are within 10% compared with the original Benes network.
Bo Wang 0023, Leibo Liu, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Zhuoquan Zhou, Shaojun Wei
IEEE Trans. Inf. Forensics Secur.3
2017 A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems
abstract
In this paper, a multi-objective, i.e., reliability, communication energy, performance, co-optimization model oriented mapping approach is proposed to find optimal mappings when applications are mapped onto network-on-chip (NoC) based reconfigurable architectures. A co-optimization model, defined as reliability efficiency model (REM), is developed to evaluate the overall reliability efficiency of a mapping. In REM, reliability efficiency is defined as the reliability profit at the same energy latency product. Based on REM, a mapping approach, referred to as priority and compensation factor oriented branch and bound (PCBB), is introduced to figure out the best mapping pattern. Two techniques, priority allocation and compensation factor utilization, are adopted to make a tradeoff between search efficiency and accuracy. Experimental results show that the proposed approach has three major contributions compared to state-of-the-art approaches. (1) PCBB is highly efficient in finding best mappings, with a 3x and 720x speedup compared to branch and bound (BB) and simulated annealing (SA). (2) PCBB is able to dynamically remap after the reconfiguration of the architecture. (3) General quantitative evaluation for reliability, communication energy and performance are made respectively before integrated into the unified model REM, whereas other similar models only touch upon two of them quantitatively.
Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei
IEEE Trans. Parallel Distributed Syst.2
2016 A fast face detection architecture for auto-focus in smart-phones and digital cameras
Shouyi Yin, Chenchen Deng, Leibo Liu, Shaojun Wei
Sci. China Inf. Sci.3
2016 Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture
abstract
With the increasing accuracy of fault injections, it has become possible to inject two faults into specific circuit regions precisely at a certain time. Unfortunately, most existing fault attack countermeasures are based on the single fault assumption, and it is, therefore, very difficult to resist double fault attacks. Reconfigurable array architecture (RAA) has the ability to introduce spatial and time randomness by dynamic reconfiguration, which can alleviate the threat of double fault attacks. This paper, for the first time, analyzes the double fault attack issues in the fault injection phase systematically. An evaluation model, named injection effort model (IEM), is proposed to quantify the efforts of a successful fault injection. In IEM, the real injection process is described mathematically using the probability method, so that a theoretical basis can be provided for the corresponding countermeasure design. Based on the concept of spatial and time randomization, three countermeasures are implemented on RAA for the purpose of decreasing the implementation overhead under the premise of ensuring the security. When these countermeasures are adopted, tradeoffs can be made between the double fault resistance and the extra overhead through changing the degree of randomness. Experiments are carried out to analyze the relationship between the resistance and the overhead using Advanced Encryption Standard (AES), Data Encryption Standard (DES), and Camellia. When the overhead constraints in terms of throughput, hardware resources, and energy are 5%, 35%, and 10% respectively, the double fault resistance can increase by two to four orders of magnitude (ranging from 824 to 10 149 for different algorithms).
Bo Wang 0023, Leibo Liu, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Shaojun Wei
IEEE Trans. Inf. Forensics Secur.3
2016 TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-Instructions
abstract
Coarse-Grained Reconfigurable Architectures (CGRAs), which provide high performance, low power and flexibility, is viewed as a promising trend for computing. CGRAs are mostly employed to process compute-intensive kernels because of their inefficiency for control flows. Various methods have been proposed to alleviate this problem, and triggered instruction is one of the state-of-the-art techniques. In this paper, a reconfigurable architecture called Triggered-Long-Instruction Architecture (TLIA) is proposed to enhance the triggered instructions with parallel condition method. In the proposed architecture, triggered instruction set is employed on processing elements (PEs). In this way, over-serialized execution and branch instructions are both eliminated. In the meanwhile, each PE has an improved data-path with three ALUs which is inspired by the parallel condition method. In this way, the amount of parallelism inside each control flow is increased by paralleling predicate computations and predicated operations. Moreover, multiple triggered instructions, which may have internal control dependence, can be executed on PEs in parallel. The strategy of issuing instructions is implemented in hardware, and verified by FPGA. Experimental results show that the performance is improved by 20.9 to 140.0 percent, the area is reduced by 24.5 percent, and the power is reduced by 32.5 percent over the equivalent Triggered Instruction Architecture (TIA).
Leibo Liu, Jianfeng Zhu 0001, Chenchen Deng, Shouyi Yin, Shaojun Wei
IEEE Trans. Parallel Distributed Syst.4
2015 A novel approach using a minimum cost maximum flow algorithm for fault-tolerant topology reconfiguration in NoC architectures
abstract
An approach using a minimum cost maximum flow algorithm is proposed for fault-tolerant topology reconfiguration in a Network-on-Chip system. Topology reconfiguration is converted into a network flow problem by constructing a directed graph with capacity constraints. A cost factor is considered to differentiate between processing elements. This approach maximizes the use of spare cores to repair faulty systems, with minimal impact on area, throughput and delay. It also provides a transparent virtual topology to alleviate the burden for operating systems.
Leibo Liu, Chenchen Deng, Shouyi Yin, Shaojun Wei, Jie Han 0001
ASP-DAC3
2015 Reliability-aware mapping for various NoC topologies and routing algorithms under performance constraints
Chenchen Deng, Leibo Liu, Shouyi Yin, Jie Han 0001, Shaojun Wei
Sci. China Inf. Sci.2
2015 An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures
abstract
In this paper, an efficient application mapping approach is proposed for the co-optimization of reliability, communication energy, and performance (CoREP) in network-on-chip (NoC)-based reconfigurable architectures. A cost model for the CoREP is developed to evaluate the overall cost of a mapping. In this model, communication energy and latency (as a measure of performance) are first considered in energy latency product (ELP), and then ELP is co-optimized with reliability by a weight parameter that defines the optimization priority. Both transient and intermittent errors in NoC are modeled in CoREP. Based on CoREP, a mapping approach, referred to as priority and ratio oriented branch and bound (PRBB), is proposed to derive the best mapping by enumerating all the candidate mappings organized in a search tree. Two techniques, branch node priority recognition and partial cost ratio utilization, are adopted to improve the search efficiency. Experimental results show that the proposed approach achieves significant improvements in reliability, energy, and performance. Compared with the state-of-the-art methods in the same scope, the proposed approach has the following distinctive advantages: 1) CoREP is highly flexible to address various NoC topologies and routing algorithms while others are limited to some specific topologies and/or routing algorithms; 2) general quantitative evaluation for reliability, energy, and performance are made, respectively, before being integrated into unified cost model in general context while other similar models only touch upon two of them; and 3) CoREP-based PRBB attains a competitive processing speed, which is faster than other mapping approaches.
Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 A Flexible Energy- and Reliability-Aware Application Mapping for NoC-Based Reconfigurable Architectures
abstract
This paper proposes a flexible energy- and reliability-aware application mapping approach for network-on-chip (NoC)-based reconfigurable architecture. A parameterized cost model is first developed by combining energy and reliability with a weight parameter that defines the optimization priority. Using this model, the overall mapping cost could be evaluated. Subsequently, a mapping method using branch and bound with a partial cost ratio is employed to find the best mapping by enumerating all the possible patterns organized in a search tree. To improve the search efficiency, nonoptimal mappings are discarded at early stages using the partial cost ratio. Using the proposed approach, applications can be mapped onto most NoC topologies and running with various routing algorithms when considering both energy and reliability. Other state-of-the-art works have also done substantial research for the same topic but only limited to a specific topology or routing algorithm. Even for the same topology and routing algorithm, the proposed approach still shows considerable advantages in many aspects. Experiments show that this approach gains not only significant reduction in energy but also improvement in reliability. It also outperforms other approaches in throughput and latency with competitive run time.
Leibo Liu, Chenchen Deng, Shouyi Yin, Jie Han 0001, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Teach Reconfigurable Computing using mixed-grained fabrics based hardware infrastructure
abstract
With the prevalence of reconfigurable computing, many relevant courses are designed and taught to graduate students. Traditional Field Programmable Gate Arrays (FPGAs) based hardware platforms are far from satisfying to reflect the important criteria characterizing a general reconfigurable computing system. In order to provide students a comprehensive understanding of reconfigurable computing system in a broader way, this paper presents a mixed-grained educational hardware platform. Different from the traditional ones, the proposed hardware platform includes not only fine-grained reconfigurable fabrics (e.g. FPGAs), but also coarse-grained ones which makes it possible to reveal essential features and intrinsic mechanisms of reconfigurable computing system. Utilizing this hardware platform, a course including four hands-on laboratory projects is designed. The feedback from students and teachers confirms that with the help of the proposed hardware platform, a thorough understanding of reconfigurable computing systems is achieved in an intuitive way and the practical experience is also significantly enhanced.
Chenchen Deng, Leibo Liu, Zhaoshi Li, Shouyi Yin, Shaojun Wei
FIE1