EDBT 2026 Demo / reviewers in the wild / expert
Ben A. Abderazek
dblp:82/4497 · also Abderazek Ben Abdallah, Ben Abdallah Abderazek
· DBLP profile ↗
35ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0003-3432-0718ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GreenMorph: Sustainable Neuromorphic Computing through Energy-Harvesting and Energy-Driven Online STDP Learning
Yuga Hanyu, Subbaiah Ravi Hariprakash, Ben A. Abderazek, Zhishang Wang, Khanh N. Dang |
ISCAS | 3 |
| 2026 | ApproxiMorph: Energy-Efficient Neuromorphic System With Layer-Wise Approximation of Spiking Neural Networks and 3-D-Stacked SRAMabstractThis paper proposes ApproxiMorph, a comprehensive framework for both software and hardware co-design, targeting energy-efficient AI applications using 3D-IC-based neuromorphic systems. By leveraging parallel interconnections and high-bandwidth communications inherent to 3D-ICs, and the noise-resilience characteristics of spiking neural networks (SNNs), ApproxiMorph achieves significant power savings by exploiting (1) approximate implementation of neuron cells, (2) layer-wise approximation of SNNs through the heuristic exploration algorithm, (3) reduced-voltage operation in the 3D-stacked SRAM, and (4) incorporating a weight-tuning method. As a result, to search for the energy-optimal layer-wise approximation, ApproxiMorph explores only 0.44—0.67% of all possible combinations, achieving a 28.06% power saving for additions with a 0.60% accuracy loss in comparison to the baseline SNN for MNIST. In the VGG16 for CIFAR-10, ApproxiMorph searches around 103 combinations from over 1017 possible solutions, resulting in a 29.16% power saving with slight accuracy gain. Furthermore, integrating all methods enhances the accuracy of approximate implementations and demonstrates higher error resilience than accurate implementations. Ryoji Kobayashi, Ngo-Doanh Nguyen, Ben A. Abderazek, Nguyen Anh Vu Doan, Khanh N. Dang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Power-Aware Neuromorphic Architecture With Partial Voltage Scaling 3-D Stacking Synaptic MemoryabstractThe combination of neuromorphic computing (NC) and 3-D integrated circuits - the 3-D stacking neuromorphic system can be the most advanced architecture that inherits the benefits of both computing and interconnect paradigms. However, simply shifting to the third dimension cannot exploit the 3-D structure and also end up with a low yield rate issue. Therefore, in this article, we propose a methodology to design 3-D stacking synaptic memory for power-efficient operations and yield rate improvement of neuromorphic systems. In this proposed methodology, the synaptic weights are stacked on top of the processing elements (PEs), and these weights are split into multiple subsets placed in different layers. Furthermore, with the support of 3-D technology, the supply voltage of each layer can be controlled independently which leads to power reduction by scaling down or turning off the supply voltage of the memory layer(s) containing the least significant bits (LSBs) while maintaining acceptable accuracy. On top of that, this work also proposes a methodology to deal with the low yield rate issue by treating the defective memory cells as noises. In our evaluation with the CMOS 45 nm technology, the energy per synaptic operation (SOP) for MNIST classification, when undervolting two upper memory layers (from 1.1 to 0.8 V), reduces by 21.62% while the accuracy only reduces sightly by 0.51%. This energy reduction increases to 66.77% with 6.58% accuracy loss when our system uses both power-gating and undervolting for all memory layers. Furthermore, the system can also improve the yield rate by 0.18% or 12.4% while suffering 0.38% or 1.7% of accuracy loss, respectively. Ngo-Doanh Nguyen, Akram Ben Ahmed, Ben A. Abderazek, Khanh N. Dang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Efficient Pneumonia Detection Method and Implementation in Chest X-ray Images Based on a Neuromorphic Spiking Neural Network
Tomohide Fukuchi, Mark Ogbodo, Jiangkun Wang, Khanh N. Dang, Ben A. Abderazek |
ICCCI | 5 |
| 2022 | HotCluster: A Thermal-Aware Defect Recovery Method for Through-Silicon-Vias Toward Reliable 3-D ICs SystemsabstractThrough silicon via (TSV) is considered as the near-future solution to realize low-power and high-performance 3D-integrated circuits (3D-ICs) and 3D-Network-on-Chips (3D-NoCs). However, the lifetime reliability issue of TSV due to its fault sensitivity and the high operating temperature of 3D-ICs, which also accelerates the fault rate, is one of the most critical challenges. Meanwhile, most current works focus on detecting and correcting TSV defects after manufacturing without considering high-temperature nodes’ impact on lifetime reliability. Besides, the recovery for defective clusters is also challenging because of costly redundancies. In this work, we presentHotCluster: a hotspot-aware self-correction platform for clustering defects in 3D-NoCs to help understand and tackle this problem. We first give a method to predict normalized fault rates and place redundant TSV groups according to each region’s fault rate. In our particular medium fault rate (normalized to the coolest area),HotClusterreduces about 60% of the redundancies in comparison to the uniformly distributed redundancies while having a higher ratio of router working in a normal state. Furthermore,HotClusterintegrates both online (weight based) and offline (max-flow min-cut offline method) mapping algorithms to help the system correct the faulty TSV clusters. The experimental results show that both the max-flow min-cut offline method and weight-based online mode with a redundancy of 0.25 exhibits less than 1% of routers disabled under 50% defect rates. Khanh N. Dang, Akram Ben Ahmed, Ben A. Abderazek, Xuan-Tu Tran |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | TSV-OCT: A Scalable Online Multiple-TSV Defects Localization for Real-Time 3-D-IC SystemsabstractIn order to detect and localize through-silicon-via (TSV) failures in both manufacturing and operating phases, most of the existing methods use a dedicated testing mechanism with long response time and prerequisite interruptions for online testing. This article presents an error correction code (ECC)-based method named “TSV on-communication test” (TSV-OCT) to detect and localize faults without halting the operation of TSV-based 3-D-IC systems. We first propose a statistical detector, a method to detect open and short defects in TSVs that work in parallel with data transactions. Second, we propose an isolation-and-check algorithm to enhance the localization ability of the method. Moreover, the Monte Carlo simulations show that the proposed statistical detector increases ×2 the number of detected faults when compared to conventional ECC-based techniques. With the help of isolation and check, TSV-OCT localizes the number of defects up to ×4 and ×5 higher. In addition, the response time is kept below 65000 cycles, which could be easily integrated into real-time applications. On the other hand, an implementation of TSV-OCT on a 3-D Network-on-Chip (NoC) router shows no performance degradation for testing while having a reasonable area overhead. Khanh N. Dang, Akram Ben Ahmed, Ben A. Abderazek, Xuan-Tu Tran |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Comprehensive Analytic Performance Assessment and K-means based Multicast Routing Algorithm and Architecture for 3D-NoC of Spiking NeuronsabstractSpiking neural networks (SNNs) are artificial neural network models that more closely mimic biological neural networks. In addition to neuronal and synaptic state, SNNs incorporate the variant time scale into their computational model. Since each neuron in these networks is connected to thousands of others, high bandwidth is required. Moreover, since the spike times are used to encode information in SNN, very low communication latency is also needed. The 2D-NoC was used as a solution to provide a scalable interconnection fabric in large-scale parallel SNN systems. The 3D-ICs have also attracted a lot of attention as a potential solution to resolve the interconnect bottleneck. The combination of these two emerging technologies provides a new horizon for IC designs to satisfy the high requirements of low power and small footprint in emerging AI applications. In this work, we first present a comprehensive analytical model to analyze the performance of 3D mesh NoC over variants of different SNN topologies and communications protocols. Second, we present an architecture and a low-latency spike routing algorithm, named shortest path K-means based multicast (SP-KMCR), for three-dimensional NoC of spiking neurons (3DNoC-SNN). The proposed system was validated based on an RTL-level implementation, while area/power analysis was performed using 45nm CMOS technology. The H. Vu, Yuichi Okuyama 0001, Ben A. Abderazek |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Analytical performance assessment and high-throughput low-latency spike routing algorithm for spiking neural network systems
The H. Vu, Yuichi Okuyama 0001, Ben A. Abderazek |
J. Supercomput. | 3 |
| 2018 | SAFT-PHENIC: a thermal-aware microring fault-resilient photonic NoC
Michael Conrad Meyer, Yuichi Okuyama 0001, Ben A. Abderazek |
J. Supercomput. | 3 |
| 2017 | A low-overhead soft-hard fault-tolerant architecture, design and management scheme for reliable high-performance many-core 3D-NoC systems
Khanh N. Dang, Michael Conrad Meyer, Yuichi Okuyama 0001, Ben A. Abderazek |
J. Supercomput. | 4 |
| 2017 | Microring fault-resilient photonic network-on-chip for reliable high-performance many-core systems
Michael Conrad Meyer, Yuichi Okuyama 0001, Ben A. Abderazek |
J. Supercomput. | 3 |
| 2017 | A Comprehensive Reliability Assessment of Fault-Resilient Network-on-Chip Using Analytical ModelabstractThe component's failure in network-on-chips (NoCs) has been a critical factor on the system's reliability. In order to alleviate the impact of faults, fault tolerance has been investigated in the recent years to enhance NoC's robustness. Due to the vast selection of fault-tolerance mechanisms and critical design constraints, selecting and configuring an appropriate mechanism to satisfy the fault-tolerance requirements constitute new challenges for designers. Consequently, reliability assessment has become prominent for the early stages of manufacturing process to solve these problems. This paper approaches the fault-tolerance analysis by providing an analytical model to approximate the lifetime reliability and compares it with a system-level simulation. Based on the proposed approach, we measure the fault-tolerance efficiency using a new parameter, named reliability acceleration factor. The goal of this paper is to provide an efficient and accurate reliability assessment to help designers easily understand and evaluate the advantages and drawbacks of their potential fault-tolerance methods. Khanh N. Dang, Akram Ben Ahmed, Xuan-Tu Tran, Yuichi Okuyama 0001, Ben A. Abderazek |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Reliability Assessment and Quantitative Evaluation of Soft-Error Resilient 3D Network-on-Chip SystemsabstractThree-Dimensional Networks-on-Chips (3D-NoCs) have been proposed as an auspicious solution, merging the high parallelism of the Network-on-Chip (NoC) paradigm with the high-performance and low-power cost of 3D-ICs. However, as technology scales down, the reliability issues are becoming more crucial, especially for complex 3D-NoC which provides the communication requirements of multi and many-core systems-on-chip. Reliability assessment is prominent for early stages of the manufacturing process to prevent costly redesigns of a target system. In this paper, we present an accurate reliability assessment and quantitative evaluation of a soft-error resilient 3D-NoC based on a soft-error resilient mechanism. The system can recover from transient errors occurring in different pipeline stages of the router. Based on this analysis, the effects of failures in the network's principal components are determined. Khanh N. Dang, Michael Conrad Meyer, Yuichi Okuyama 0001, Ben A. Abderazek |
ATS | 4 |
| 2016 | Adaptive fault-tolerant architecture and routing algorithm for reliable many-core 3D-NoC systems
Akram Ben Ahmed, Ben A. Abderazek |
J. Parallel Distributed Comput. | 2 |
| 2015 | Hybrid Photonic NoC Based on Non-Blocking Photonic Switch and Light-Weight Electronic RouterabstractIn conventional hybrid PNoC systems, the end-to end optical data transfer is accompanied with electrical control functions including path-setup, acknowledgment, and tear-down. These functions directly a ect the performance and power characteristics of these circuit-switching based systems. In this paper, we propose a novel hybrid PNoC system, named PHENIC-II. 1 Thanks to the adopted non-blocking photonic switch and light-weight electronic router, PHENIC-II is capable of alleviating the congestion in the electronic control layer which is considered as the main source of latency and power overhead in hybrid PNoC systems. From the performance evaluation, we demonstrate that the proposed system has a better performance and low energy dissipation when compared to the previously proposed systems. Achraf Ben Ahmed, Michael Conrad Meyer, Yuichi Okuyama 0001, Ben A. Abderazek |
SMC | 4 |
| 2015 | On the Design of a Fault-Tolerant Photonic Network-on-ChipabstractOptical Network-on-Chip is a solution to for power and throughput bottlenecks of current technology. The higher bandwidth is achieved by the light speed transmissions, and the power required to transmit data in the optical domain is much lower. This is a disruptive technology solution to problems arising from silicon-based computing. In this paper, we present a fault-tolerant optical router (FTTDOR) with its electrical control module towards the design of a highly-reliable low-power three dimensional Networks-on-Chip (PHENIC). FTTDOR uses redundancy only in critical locations, to assure accuracy of the packet transmission even after a faulty ring resonator appears. The proposed optical router is decomposed non-blocking, with minimal ring resonators, and requires no resonators for straight travel (East to West, North to South, and Up to Down, as well as their inverses). Simulation results show that the network can maintain a 98% throughput after 3% faults, and 89% after 20% faults. These results come with a reduction of micro-ring resonators to 65% of the amount present in a conventional crossbar router. Simulation of the electrical control module and router show that it has a total area of around 20,000 μm2and consumes 3.8mW at 600MHz. Michael Conrad Meyer, Akram Ben Ahmed, Ben A. Abderazek |
SMC | 4 |
| 2015 | Hybrid silicon-photonic network-on-chip for future generations of high-performance many-core systems
Achraf Ben Ahmed, Ben A. Abderazek |
J. Supercomput. | 2 |
| 2014 | Graceful deadlock-free fault-tolerant routing algorithm for 3D Network-on-Chip architectures
Akram Ben Ahmed, Ben A. Abderazek |
J. Parallel Distributed Comput. | 2 |
| 2013 | Run-Time Monitoring Mechanism for Efficient Design of Application-Specific NoC Architectures in Multi/Manycore EraabstractOne of the major design challenges of Network-on-Chip interconnect is the storage buffers. They occupy a significant portion of the system's area and so they are considered as main "power-hungry" components. Deciding the appropriate buffers size and implementation in these systems is the key technique for increasing system performance and also for reducing overall area and power consumption. However, this goal is very hard to achieve with traditional design approaches, where design decisions of the main architectural parameters are generally made with slow and inaccurate software simulation or theoretical modeling. In order to quickly capture and decide the optimal buffers size and the whole system behavior, we propose in this work an efficient design method for Network-on-Chip architecture based on a novel run-time monitoring mechanism (RMM). The system monitors the traffic flow at different system's resources and sends the monitored run-time traffic information to a specialized controller. In addition, our proposed design method allows to easily compute optimal architecture hardware parameters (i.e Buffer size) and allocate the appropriate values on demand to satisfy the requirements of any given application. The RMM mechanism was designed in hardware and integrated into our NoC system (PNoC). From the evaluation results, we conclude that the system performance in terms of execution time was about 27% better when compared with traditional design methods over several benchmark programs. Akram Ben Ahmed, Takayuki Ochi, Shohei Miura, Ben A. Abderazek |
CISIS | 4 |
| 2013 | Architecture and design of high-throughput, low-latency, and fault-tolerant routing algorithm for 3D-network-on-chip (3D-NoC)
Akram Ben Ahmed, Ben A. Abderazek |
J. Supercomput. | 2 |
| 2011 | Natural instruction level parallelism-aware compiler for high-performance QueueCore processor architecture
Ben A. Abderazek, Masashi Masuda, Arquimedes Canedo, Kenichi Kuroda |
J. Supercomput. | 1 |
| 2009 | Efficient compilation for queue size constrained queue processors
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
Parallel Comput. | 2 |
| 2008 | Advanced Optimization and Design Issues of a 32-Bit Embedded Processor Based on Produced Order Queue Computation ModelabstractQueue computing based programs are generated using a so called level order traversal that exposes all available parallelism in the programs. All instructions within the same level are data independent from each other and are safely to be executed in parallel. This property is leveraged by the compiler generating queue programs with high amounts of grouped independent instructions. Thus, the hardware invests little efforts to find parallelism. In this paper, we present various optimization and design issues of a synthesizable queue processor architecture targeted for embedded applications. A prototype implementation is produced by synthesizing the high-level model for a target FPGA device. Hiroki Hoshino, Ben A. Abderazek, Kenichi Kuroda |
EUC (1) | 2 |
| 2008 | Single Instruction Dual-Execution Model Processor ArchitectureabstractWe present in this paper architecture and preliminary evaluation results of a novel dual-mode processor architecture which supports queue and stack computation models in a single core. The core is highly adaptable in both functionality and configuration. It is based on a reduced bit produced order queue computation instruction set architecture and functions into Queue or Stack execution models. This is achieved via a so called dynamic switching mechanism implemented in hardware. The current design focuses on the ability to execute Queue programs and also to support Stack based programs without considerable increase in hardware to the base architecture. We present the architecture description and design results in a fair amount of details. Taichi Maekawa, Ben A. Abderazek, Kenichi Kuroda |
EUC (1) | 2 |
| 2008 | A new code generation algorithm for 2-offset producer order queue computation model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
Comput. Lang. Syst. Struct. | 2 |
| 2008 | The QC-2 parallel Queue processor architecture
Ben A. Abderazek, Arquimedes Canedo, Tsutomu Yoshinaga, Masahiro Sowa |
J. Parallel Distributed Comput. | 1 |
| 2008 | Dual-execution mode processor architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Masahiro Sowa |
J. Supercomput. | 2 |
| 2007 | An Efficient Code Generation Algorithm for Code Size Reduction Using 1-Offset P-Code Queue Computation Model
Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
EUC | 2 |
| 2007 | New Code Generation Algorithm for QueueCore - An Embedded Processor with High ILPabstractModern architectures rely on exploiting parallelism found at the instruction level to achieve high performance. Aggressive ILP compilers expose high amounts of instruction level parallelism where, in some cases, the number of architected registers is not enough to hold the results of potential parallel instructions. This paper presents a new code generation scheme for the QueueCore, a 32-bit queue-based architecture capable of executing high amounts of ILP. QueueCore's instructions implicitly read their operands and write results. Compiling for the QueueCore requires that all instructions have at most one explicit operand represented as an offset calculated at compile-time. Additionally, the instructions must be scheduled in level-order manner. The proposed algorithm successfully restricts all instructions to have at most one offset reference, it computes the offset values, and makes a level-order scheduling of the program. To evaluate the effectiveness of the new code generation scheme we developed a queue compiler and compiled a set of benchmark programs. Our results show that the code has more parallelism than optimized RISC code by factors ranging from 1.12 to 2.30. QueueCore's instruction set allows us to generate code about 40%-18% denser than optimized RISC code. Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
PDCAT | 2 |
| 2007 | Queue Register File Optimization Algorithm for QueueCore ProcessorabstractThe queue computation model offers an attractive alternative for high-performance embedded computing given its characteristics of short instructions and high instruction level parallelism. A queue-based processor uses a FIFO queue to read and write operands through hardware pointers located at the head and tail of the queue. Queue length is the number of elements stored between the head and the tail pointers during computations. We have found that 95% of the statements in integer applications require a queue length of less than 32 words. The remaining 5% requires larger queue length sizes up to 230 queue words. In this paper we propose a compiler technique to optimize the queue utilization for the hungry statements that require a large amount of queue. We show that for SPEC CINT95 benchmarks, our technique optimizes the queue length without decreasing parallelism. However, our optimization has a penalty of a slight increase in code size. Arquimedes Canedo, Ben A. Abderazek, Masahiro Sowa |
SBAC-PAD | 2 |
| 2006 | High-Level Modeling and FPGA Prototyping of Produced Order Parallel Queue Processor Core
Ben A. Abderazek, Tsutomu Yoshinaga, Masahiro Sowa |
J. Supercomput. | 1 |
| 2005 | Modular Design Structure and High-Level Prototyping for Novel Embedded Processor Core
Ben A. Abderazek, Sotaro Kawata, Tsutomu Yoshinaga, Masahiro Sowa |
EUC | 1 |
| 2005 | An Efficient Dynamic Switching Mechanism (DSM) for Hybrid Processor Architecture
Md. Musfiquzzaman Akanda, Ben A. Abderazek, Sotaro Kawata, Masahiro Sowa |
EUC | 2 |
| 2005 | Parallel Queue Processor Architecture Based on Produced Order Computation Model
Masahiro Sowa, Ben A. Abderazek, Tsutomu Yoshinaga |
J. Supercomput. | 2 |
| 2003 | On the Design of a Register Queue Based Processor Architecture (FaRM-rq)
Ben A. Abderazek, Soichi Shigeta, Tsutomu Yoshinaga, Masahiro Sowa |
ISPA | 1 |