EDBT 2026 Demo / reviewers in the wild / expert
Chenchen Deng
dblp:137/1419
· DBLP profile ↗
22ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-1047-1087ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
15 papers |
Reconfigurable computing and FPGAs · 28% Memory systems · 20% Interconnection networks and networks-on-chip · 14% | |
| Network and information security
4 papers |
Hardware security and side channels · 62% Cryptographic primitives and cryptanalysis · 28% Privacy and data protection · 10% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
1.2 | 3 | 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution · DAC 2017 TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-Instructions · IEEE Trans. Parallel Distributed Syst. 2016 |
Integrated circuit design
3d integration |
0.9 | 1 | 2025 | Software-defined process-near-memory architecture using 3D hybrid bonding integration · Sci. China Inf. Sci. 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | Software-defined process-near-memory architecture using 3D hybrid bonding integration · Sci. China Inf. Sci. 2025 |
Memory systems
content-addressable memory |
0.8 | 1 | 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024 |
Memory systems › content-addressable memory
TCAM |
0.8 | 1 | 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024 |
Reconfigurable computing and FPGAs
reconfigurable architecture |
0.7 | 3 | 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-Instructions · IEEE Trans. Parallel Distributed Syst. 2016 A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems
fault tolerance |
0.7 | 1 | 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
Hardware reliability and fault tolerance
redundancy |
0.7 | 1 | 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023 |
Hardware security and side channels
side-channel countermeasures |
0.6 | 1 | 2022 | An energy-efficient dynamically reconfigurable cryptographic engine with improved power/EM-side-channel-attack resistance · Sci. China Inf. Sci. 2022 |
Hardware security and side channels
fault attacks |
0.5 | 2 | 2017 | Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017 Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016 |
Hardware security and side channels
fault attack countermeasure |
0.5 | 2 | 2017 | Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017 Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016 |
Reconfigurable computing and FPGAs
application mapping |
0.5 | 2 | 2017 | A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing Systems · IEEE Trans. Parallel Distributed Syst. 2017 An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Interconnection networks and networks-on-chip › network-on-chip design
noc power management |
0.4 | 1 | 2020 | Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Energy-efficient computing
power gating |
0.4 | 1 | 2020 | Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Interconnection networks and networks-on-chip › network topology
reconfigurable topology |
0.4 | 1 | 2020 | CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology · DAC 2020 |
Reconfigurable computing and FPGAs › application mapping
reliability-aware mapping |
0.4 | 2 | 2015 | An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 Reliability-aware mapping for various NoC topologies and routing algorithms under performance constraints · Sci. China Inf. Sci. 2015 |
Cryptographic primitives and cryptanalysis › cryptographic implementation
hardware implementation |
0.3 | 1 | 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Cryptographic primitives and cryptanalysis
symmetric cryptography |
0.3 | 1 | 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Processor architecture and microarchitecture
multithreading |
0.3 | 1 | 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric |
0.3 | 1 | 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key Algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Processor architecture and microarchitecture › pipelining
pipeline stall reduction |
0.3 | 1 | 2017 | Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution · DAC 2017 |
Privacy and data protection
randomization |
0.2 | 1 | 2016 | Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array Architecture · IEEE Trans. Inf. Forensics Secur. 2016 |
Integrated circuit design
low-power circuit design |
0.2 | 1 | 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency · Sci. China Inf. Sci. 2024 |
Energy-efficient computing › energy-efficient communication
communication energy minimization |
0.2 | 1 | 2015 | An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Energy-efficient computing
power management |
0.2 | 1 | 2015 | An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing |
0.1 | 1 | 2020 | Aggressive Fine-Grained Power Gating of NoC Buffers · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Interconnection networks and networks-on-chip › network topology
torus network |
0.1 | 1 | 2020 | CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus Topology · DAC 2020 |
Computer vision › Face, body and person analysis
face alignment |
0.1 | 1 | 2017 | A 700fps Optimized Coarse-to-Fine Shape Searching Based Hardware Accelerator for Face Alignment · DAC 2017 |
Cryptographic primitives and cryptanalysis
block cipher |
0.1 | 1 | 2017 | Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault Attacks · IEEE Trans. Inf. Forensics Secur. 2017 |
Computer vision › Face, body and person analysis
face detection |
0.1 | 1 | 2016 | A fast face detection architecture for auto-focus in smart-phones and digital cameras · Sci. China Inf. Sci. 2016 |
Methods — techniques the papers use, named apart from their topics
dynamic reconfiguration · 1.83d hybrid bonding · 1.7power/EM side-channel countermeasures · 1.1constant-time alteration · 0.828 nm CMOS · 0.8temporal-redundant voters · 0.7spatial-redundant data paths · 0.7markov process model · 0.7branch-and-bound · 0.5flit deflection · 0.4context compression · 0.3statistical evaluation · 0.3cascaded regression · 0.3benes network · 0.3SURF features · 0.3spatial and time randomization · 0.2injection effort model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Software-defined process-near-memory architecture using 3D hybrid bonding integration
Anlin Xu, Chenchen Deng, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 2 |
| 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency
Chenchen Deng, Tianzhu Xiong, Zhaoshi Li, Jianfeng Zhu 0001, Jun Yang 0006, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 1 |
| 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) can provide both energy efficiency and performance for embedded systems, and thus they are increasingly deployed in the areas of aerospace, automotive engineering, and security where reliability is also a main criterion. However, the state-of-the-art fault-tolerant strategies for CGRAs apply either temporal or spatial scheme, including redundancy, periodic detection, workload balancing, and reconfiguration, failing to exploit the feature of dynamic and partial reconfiguration of CGRAs. Also, vulnerable judging circuits and inflexible mode shifting bottleneck the reliability design of fault-tolerant CGRAs. This article proposes a novel multimode fault-tolerant framework for CGRAs, which combines spatial-redundant data paths with temporal-redundant voters and thus reduces the vulnerable judging circuits while balancing the performance and reliability. This framework can also enable a changing reliability level at runtime via an online configuration transformation method based on precompiled patterns. Within the proposed framework, we systematically searched the design space spanning various combinations of the mainstream schemes with a Markov process model to compare the effectiveness and accordingly selected five points as available modes in our design after comprehensive consideration of fault tolerance and time overhead on CGRA. The framework is comprehensively evaluated on a cycle-accurate CGRA simulator, considering both permanent and transient faults. The experimental results show that the fault coverage rate of single transient faults or permanent faults has increased from 71.74% to 93.84%, which means the fault tolerance of the system has been increased by 31.03% compared with the state-of-the-art methods. There is also a great improvement in mean-time-to-failure (MTTF) and reconfiguration latency over baseline designs. Jianfeng Zhu 0001, Xingchen Man, Guihuan Song, Yi Huang 0036, Chenchen Deng, Pengfei Gou, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | An energy-efficient dynamically reconfigurable cryptographic engine with improved power/EM-side-channel-attack resistance
Chenchen Deng, Min Zhu 0001, Jinjiang Yang, Youyu Wu, Jiaji He 0001, Bohan Yang 0001, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 1 |
| 2021 | LWRpro: An Energy-Efficient Configurable Crypto-Processor for Module-LWRabstractSaber, the only module-learning with rounding-based algorithm in NIST's third round of post-quantum cryptography (PQC) standardization process, is characterized by simplicity and flexibility. However, energy-efficient implementation of Saber is still under investigation since the commonly used number theoretic transform can not be utilized directly. In this manuscript, an energy-efficient configurable crypto-processor supporting multi-security-level key encapsulation mechanism of Saber, is proposed. First, an 8-level hierarchical Karatsuba framework is utilized to reduce degree-256 polynomial multiplication to the coefficient-wise multiplication. Second, a hardware-efficient Karatsuba scheduling strategy and an optimized pre-/post-processing structure is designed to reduce the area overheads of scheduling strategy. Third, a task-rescheduling-based pipeline strategy and truncated multipliers are proposed to enable fine-grained processing. Moreover, multiple parameter sets are supported in LWRpro to enable configurability among various security scenarios. Enabled by these optimizations, LWRpro requires 1066, 1456 and 1701 clock cycles for key generation, encapsulation, and decapsulation of Saber768. The post-layout version of LWRpro is implemented with TSMC 40 nm CMOS process within 0.38 mm2. The throughput for Saber768 is up to 275k encapsulation operations per second and the energy efficiency is 0.15 uJ/encapsulation while operating at 400 MHz, achieving nearly 50× improvement and 31× improvement, respectively compared with current PQC hardware solutions. Yihong Zhu, Min Zhu 0001, Bohan Yang 0001, Wenping Zhu, Chenchen Deng, Chen Chen 0083, Shaojun Wei, Leibo Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | CDRing: Reconfigurable Ring Architecture by Exploiting Cycle Decomposition of Torus TopologyabstractFuture NoCs should be highly flexible to adapt to communication demands to achieve high scalability and low power consumption. However, the flexibility is still quite limited by the high complexity of reconfiguration for globally reconfigured channels. In this paper, we propose to augment a router-based buffered NoC with a reconfigurable ring architecture by exploiting cycle decomposition of a torus bufferless network. At runtime, the topologies of the rings can be reconfigured according to the workloads by choosing different cycle decompositions of the torus network. Because the shapes of the rings are restricted to a specified regular shape, the reconfiguration time can be reduced to a linear complexity with respect to network size, and the reconfiguration algorithm can be implemented in a distributed hardware. The experimental results show that the reconfigurable rings provide 54% and 26% improvements on packet latency and static power saving, respectively, for realistic workloads. Liang Wang 0020, Leibo Liu, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei |
DAC | 5 |
| 2020 | Aggressive Fine-Grained Power Gating of NoC BuffersabstractPower gating is effective for networks-on-chip (NoCs) to reduce the excessive leakage power dissipated by idle network components. Most existing NoC power-gating approaches rely on the routing algorithms to mitigate the power-gating blocking latency problem. When the network becomes faulty and fault-tolerant routing algorithms are applied, these approaches are no longer applicable or can seriously degrade the performance. Other approaches propose fine-grained buffer power gating, but they are too conservative in power saving due to the buffer backpressure flow control. To address these problems, we propose an aggressive fine-grained power gating of flit-sized buffer entries by adopting backpressureless flow control in an input-buffered network. The power-gating decisions are made based on the flit deflection rate. However, directly applying the backpressureless flow control leads to the difficulties of multiflit packet truncation and protocol deadlocks. Therefore, we modify the packet injection architecture to avoid packet truncation. This is done by chaining the local input port with a randomly chosen input port. Finally, we design a progressive recovery framework to handle both livelocks and protocol deadlocks. It does not need to truncate packets or strictly separate different message classes when the network is free of livelocks or protocol deadlocks. The experimental results show that with a hardware overhead of 9.6%, our design can save up to 59% network power consumption in both a fault-free and a faulty NoC with little zero-load latency penalty. Our design also approaches an ideal energy-proportional NoC because it can constantly reduce power consumption over a wide range of injection rates. Leibo Liu, Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Chenchen Deng, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | A Reliable Physical Unclonable Function Based on Differential Charging CapacitorsabstractPhysical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit. Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044 |
ISCAS | 7 |
| 2018 | Anole: A Highly Efficient Dynamically Reconfigurable Crypto-Processor for Symmetric-Key AlgorithmsabstractThis paper presents a dynamically reconfigurable processing array named Anole for symmetric-key algorithms. Processing elements and the interconnections between them are designed to support various block and stream ciphers. Without affecting flexibility, three key techniques are presented to increase energy efficiency (throughput/power, the number of operations per unit energy consumption) and area efficiency (throughput/area). First, the distributed control network supports multithreading on reconfigurable fabrics at a low cost, thereby maximizing the utility of computing resources in the space domain. Second, the concurrent computation and reconfiguration scheme integrates configuration contexts with processing data to simultaneously execute in the data-path. The resulted immediate switching between different configurations increases the utilization rate of hardware resources in the temporal domain. Third, under configuration context compression and organization, the context memory size and configuration time are further minimized. Anole is implemented on a 7.75 mm2silicon square with TSMC 65-nm technology at 400 MHz. Experiments show that Anole significantly outperforms field programmable gate array and general purpose processor by more than two orders of magnitude in energy and area efficiencies. Compared with state-of-the-art reconfigurable solutions, Anole achieves (average) 16.5× higher energy efficiency and 9.4× higher area efficiency. Leibo Liu, Bo Wang 0023, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and ExecutionabstractThe pipeline stall in distributed-controlled coarse-grained reconfigurable arrays is a major source stumbling performance. This work presents a Triggered-Issue and Triggered-Execution (TITE) paradigm motivated from the Triggered Instruction Architecture (TIA) which converts control and data dependencies into predicate dependencies as triggers for spatial parallelism. TITE separately triggers the issuing and execution of instructions to further relax the predicate dependencies in TIA. Triggered dual instructions and tag forwarding are proposed to minimize pipeline stalls of both intra and inter-processing elements. Experiments show that TITE improves performance, energy efficiency, and area efficiency by 21%, 17%, and 12%, respectively, compared with TIA. Yanan Lu, Leibo Liu, Yangdong Deng, Jian Weng 0004, Zhaoshi Li, Chenchen Deng, Shaojun Wei |
DAC | 6 |
| 2017 | A 700fps Optimized Coarse-to-Fine Shape Searching Based Hardware Accelerator for Face AlignmentabstractIn this work, a fast shape searching face alignment (F-SSFA) algorithm based accelerator is proposed to achieve real-time processing. Firstly, a learning based low-dimensional SURF feature is introduced to reduce the computation cost in the cascaded regression. Then the Euclidean distance and shape affine transformation are utilized to accelerate the shape searching procedure. F-SSFA therefore greatly reduces the computational complexity while keeping the same accuracy. Also, a fixed-point F-SSFA based VLSI architecture is designed with approximately 80% decrease in the data transmission traffic. The throughput of this accelerator achieves 700 fps, which is especially suitable for high-speed facial-related applications. Leibo Liu, Wenping Zhu, Huiyu Mo, Chenchen Deng, Shaojun Wei |
DAC | 5 |
| 2017 | Implementation of in-loop filter for HEVC decoder on reconfigurable processorabstractThe in‐loop filter comprises deblocking filter and sample adaptive offset filter, which is an important module for improving image quality in a high‐efficiency video coding (HEVC) decoder. The in‐loop filter has a high computational complexity that accounts for ∼20% of the HEVC decoding computing load. Furthermore, it is difficult to implement a high‐performing in‐loop filter due to its large conditional processing requirement. First, this study presents a novel reconfigurable HEVC in‐loop filter implementation on a coarse‐grained dynamically reconfigurable processing unit. Next, a repartition scheme is presented that allows the in‐loop filter implementation at a coding tree unit along with the other decoding modules in the HEVC decoder, which satisfies requirements of low latency applications. Finally, a hierarchised‐pipeline and synchronised‐parallel technique is used to improve performance by eliminating data hazards in pipeline techniques and synchronisation problems in parallel techniques. Implementation results show that the presented HEVC in‐loop filter performs up to 1920 × 1080@52 frames per second at 250 MHz. The throughput is 67.5 × 9 × more than solutions based on digital signal processor and general‐purpose processor, respectively. Leibo Liu, Victor Y. Chen, Chenchen Deng, Shouyi Yin, Shaojun Wei |
IET Image Process. | 3 |
| 2017 | Exploration of Benes Network in Cryptographic Processors: A Random Infection Countermeasure for Block Ciphers Against Fault AttacksabstractTraditional detection countermeasures against fault attacks have been criticized as insecure because of the fragile comparison operation that can be maliciously bypassed. In order to avoid the comparison, infection countermeasures have been designed to confuse the faulty ciphertexts so that the output cannot be further explored. This paper presents an infection method that resists fault attacks using the existing Benes network module in high-performance crypto processors. The Benes network is originally used to accelerate permutation operations in block ciphers. The hamming weight of the differential results is balanced by modifying specific network switches, without changing the network topology. A further confusion is performed to destroy the determinacy by configuring part of the network with a random bit-stream. Furthermore, a statistical evaluation method is presented to quantitatively verify the proposed countermeasure in addition to a formal proof of security. This also provides a new concept for the evaluation of future random-enhanced infection methods. Experiments are carried out using Advanced Encryption Standard (AES), triple Data Encryption Standard (DES), and Camellia as examples. Under statistical evaluation, the results show that the proposed countermeasure improves the fault resistance by over four orders of magnitude compared with the unprotected case. Also, the performance and the area overhead are within 10% compared with the original Benes network. Bo Wang 0023, Leibo Liu, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Zhuoquan Zhou, Shaojun Wei |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | A Multi-Objective Model Oriented Mapping Approach for NoC-based Computing SystemsabstractIn this paper, a multi-objective, i.e., reliability, communication energy, performance, co-optimization model oriented mapping approach is proposed to find optimal mappings when applications are mapped onto network-on-chip (NoC) based reconfigurable architectures. A co-optimization model, defined as reliability efficiency model (REM), is developed to evaluate the overall reliability efficiency of a mapping. In REM, reliability efficiency is defined as the reliability profit at the same energy latency product. Based on REM, a mapping approach, referred to as priority and compensation factor oriented branch and bound (PCBB), is introduced to figure out the best mapping pattern. Two techniques, priority allocation and compensation factor utilization, are adopted to make a tradeoff between search efficiency and accuracy. Experimental results show that the proposed approach has three major contributions compared to state-of-the-art approaches. (1) PCBB is highly efficient in finding best mappings, with a 3x and 720x speedup compared to branch and bound (BB) and simulated annealing (SA). (2) PCBB is able to dynamically remap after the reconfiguration of the architecture. (3) General quantitative evaluation for reliability, communication energy and performance are made respectively before integrated into the unified model REM, whereas other similar models only touch upon two of them quantitatively. Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | A fast face detection architecture for auto-focus in smart-phones and digital cameras
Shouyi Yin, Chenchen Deng, Leibo Liu, Shaojun Wei |
Sci. China Inf. Sci. | 3 |
| 2016 | Against Double Fault Attacks: Injection Effort Model, Space and Time Randomization Based Countermeasures for Reconfigurable Array ArchitectureabstractWith the increasing accuracy of fault injections, it has become possible to inject two faults into specific circuit regions precisely at a certain time. Unfortunately, most existing fault attack countermeasures are based on the single fault assumption, and it is, therefore, very difficult to resist double fault attacks. Reconfigurable array architecture (RAA) has the ability to introduce spatial and time randomness by dynamic reconfiguration, which can alleviate the threat of double fault attacks. This paper, for the first time, analyzes the double fault attack issues in the fault injection phase systematically. An evaluation model, named injection effort model (IEM), is proposed to quantify the efforts of a successful fault injection. In IEM, the real injection process is described mathematically using the probability method, so that a theoretical basis can be provided for the corresponding countermeasure design. Based on the concept of spatial and time randomization, three countermeasures are implemented on RAA for the purpose of decreasing the implementation overhead under the premise of ensuring the security. When these countermeasures are adopted, tradeoffs can be made between the double fault resistance and the extra overhead through changing the degree of randomness. Experiments are carried out to analyze the relationship between the resistance and the overhead using Advanced Encryption Standard (AES), Data Encryption Standard (DES), and Camellia. When the overhead constraints in terms of throughput, hardware resources, and energy are 5%, 35%, and 10% respectively, the double fault resistance can increase by two to four orders of magnitude (ranging from 824 to 10 149 for different algorithms). Bo Wang 0023, Leibo Liu, Chenchen Deng, Min Zhu 0001, Shouyi Yin, Shaojun Wei |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-InstructionsabstractCoarse-Grained Reconfigurable Architectures (CGRAs), which provide high performance, low power and flexibility, is viewed as a promising trend for computing. CGRAs are mostly employed to process compute-intensive kernels because of their inefficiency for control flows. Various methods have been proposed to alleviate this problem, and triggered instruction is one of the state-of-the-art techniques. In this paper, a reconfigurable architecture called Triggered-Long-Instruction Architecture (TLIA) is proposed to enhance the triggered instructions with parallel condition method. In the proposed architecture, triggered instruction set is employed on processing elements (PEs). In this way, over-serialized execution and branch instructions are both eliminated. In the meanwhile, each PE has an improved data-path with three ALUs which is inspired by the parallel condition method. In this way, the amount of parallelism inside each control flow is increased by paralleling predicate computations and predicated operations. Moreover, multiple triggered instructions, which may have internal control dependence, can be executed on PEs in parallel. The strategy of issuing instructions is implemented in hardware, and verified by FPGA. Experimental results show that the performance is improved by 20.9 to 140.0 percent, the area is reduced by 24.5 percent, and the power is reduced by 32.5 percent over the equivalent Triggered Instruction Architecture (TIA). Leibo Liu, Jianfeng Zhu 0001, Chenchen Deng, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | A novel approach using a minimum cost maximum flow algorithm for fault-tolerant topology reconfiguration in NoC architecturesabstractAn approach using a minimum cost maximum flow algorithm is proposed for fault-tolerant topology reconfiguration in a Network-on-Chip system. Topology reconfiguration is converted into a network flow problem by constructing a directed graph with capacity constraints. A cost factor is considered to differentiate between processing elements. This approach maximizes the use of spare cores to repair faulty systems, with minimal impact on area, throughput and delay. It also provides a transparent virtual topology to alleviate the burden for operating systems. Leibo Liu, Chenchen Deng, Shouyi Yin, Shaojun Wei, Jie Han 0001 |
ASP-DAC | 3 |
| 2015 | Reliability-aware mapping for various NoC topologies and routing algorithms under performance constraints
Chenchen Deng, Leibo Liu, Shouyi Yin, Jie Han 0001, Shaojun Wei |
Sci. China Inf. Sci. | 2 |
| 2015 | An Efficient Application Mapping Approach for the Co-Optimization of Reliability, Energy, and Performance in Reconfigurable NoC ArchitecturesabstractIn this paper, an efficient application mapping approach is proposed for the co-optimization of reliability, communication energy, and performance (CoREP) in network-on-chip (NoC)-based reconfigurable architectures. A cost model for the CoREP is developed to evaluate the overall cost of a mapping. In this model, communication energy and latency (as a measure of performance) are first considered in energy latency product (ELP), and then ELP is co-optimized with reliability by a weight parameter that defines the optimization priority. Both transient and intermittent errors in NoC are modeled in CoREP. Based on CoREP, a mapping approach, referred to as priority and ratio oriented branch and bound (PRBB), is proposed to derive the best mapping by enumerating all the candidate mappings organized in a search tree. Two techniques, branch node priority recognition and partial cost ratio utilization, are adopted to improve the search efficiency. Experimental results show that the proposed approach achieves significant improvements in reliability, energy, and performance. Compared with the state-of-the-art methods in the same scope, the proposed approach has the following distinctive advantages: 1) CoREP is highly flexible to address various NoC topologies and routing algorithms while others are limited to some specific topologies and/or routing algorithms; 2) general quantitative evaluation for reliability, energy, and performance are made, respectively, before being integrated into unified cost model in general context while other similar models only touch upon two of them; and 3) CoREP-based PRBB attains a competitive processing speed, which is faster than other mapping approaches. Chenchen Deng, Leibo Liu, Jie Han 0001, Jiqiang Chen, Shouyi Yin, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | A Flexible Energy- and Reliability-Aware Application Mapping for NoC-Based Reconfigurable ArchitecturesabstractThis paper proposes a flexible energy- and reliability-aware application mapping approach for network-on-chip (NoC)-based reconfigurable architecture. A parameterized cost model is first developed by combining energy and reliability with a weight parameter that defines the optimization priority. Using this model, the overall mapping cost could be evaluated. Subsequently, a mapping method using branch and bound with a partial cost ratio is employed to find the best mapping by enumerating all the possible patterns organized in a search tree. To improve the search efficiency, nonoptimal mappings are discarded at early stages using the partial cost ratio. Using the proposed approach, applications can be mapped onto most NoC topologies and running with various routing algorithms when considering both energy and reliability. Other state-of-the-art works have also done substantial research for the same topic but only limited to a specific topology or routing algorithm. Even for the same topology and routing algorithm, the proposed approach still shows considerable advantages in many aspects. Experiments show that this approach gains not only significant reduction in energy but also improvement in reliability. It also outperforms other approaches in throughput and latency with competitive run time. Leibo Liu, Chenchen Deng, Shouyi Yin, Jie Han 0001, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Teach Reconfigurable Computing using mixed-grained fabrics based hardware infrastructureabstractWith the prevalence of reconfigurable computing, many relevant courses are designed and taught to graduate students. Traditional Field Programmable Gate Arrays (FPGAs) based hardware platforms are far from satisfying to reflect the important criteria characterizing a general reconfigurable computing system. In order to provide students a comprehensive understanding of reconfigurable computing system in a broader way, this paper presents a mixed-grained educational hardware platform. Different from the traditional ones, the proposed hardware platform includes not only fine-grained reconfigurable fabrics (e.g. FPGAs), but also coarse-grained ones which makes it possible to reveal essential features and intrinsic mechanisms of reconfigurable computing system. Utilizing this hardware platform, a course including four hands-on laboratory projects is designed. The feedback from students and teachers confirms that with the help of the proposed hardware platform, a thorough understanding of reconfigurable computing systems is achieved in an intuitive way and the practical experience is also significantly enhanced. Chenchen Deng, Leibo Liu, Zhaoshi Li, Shouyi Yin, Shaojun Wei |
FIE | 1 |