Nacer-Eddine Zergainoh

dblp:55/4737 · DBLP profile ↗
← Back
35ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5418-1594ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 16 · 2 since 2021
YearPublicationVenuePosition
2025 Veritas-PUF: A Countermeasure against X-Ray Tampering and Side Channel Attacks
abstract
This paper presents Veritas-PUF, a novel solution based on Physical Unclonable Functions to counter non-invasive physical tampering attacks. This includes localized X-Ray irradiation techniques that can alter circuit behavior without modifying data, directly impacting side-channel leakages. These emerging threats pose significant challenges for secure circuits. To address such attacks, we propose leveraging a ring oscillator PUF to generate a unique identifier sensitive to ionizing radiations. The inherent entropy generated by the device’s physical properties ensures that any alterations will be reflected in the PUF response, providing a robust detection mechanism against these types of attacks. Our experimental results show that with only three reliable challenge-response pairs, we achieve over 95% detection rate for X-ray tampering.
Nasr-Eddine Ouldei Tebina, Luc Salvo, Marc Xander Makkes, Nacer-Eddine Zergainoh, Paolo Maistri
ATS4
2024 Enhancing Side-Channel Attacks Through X-Ray-Induced Leakage Amplification
abstract
In this paper, we propose a novel approach that utilizes localized X-ray irradiation to amplify data-dependent leakage currents in CMOS-based cryptography circuits. Our proposed technique strategically targets specific regions in a circuit using X-rays, inducing variations in dynamic and static power consumption due to Total Ionizing Dose (TID) effects, which increases or even reveals hidden data leakage. In this work, we present several experimental campaigns highlighting the benefits of our approach to combinational and sequential logic. Our experiments show a significant increase in information leakage in the targeted regions, which improves the signal-to-noise ratio coefficient and thus makes recovering the processed bytes easier. We envision the possibility of using this technique on full cryptographic designs on both FPGA and ASICs.
Nasr-Eddine Ouldei Tebina, Luc Salvo, Laurent Maingault, Nacer-Eddine Zergainoh, Guillaume Hubert, Paolo Maistri
DATE4
2023 Ray-Spect: Local Parametric Degradation for Secure Designs: An application to X-Ray Fault Injection
abstract
International audience
Nasr-Eddine Ouldei Tebina, Laurent Maingault, Nacer-Eddine Zergainoh, Guillaume Hubert, Paolo Maistri
IOLTS3
2018 Designing reliable processor cores in ultimate CMOS and beyond: A double sampling solution
abstract
The double sampling paradigm is an efficient method to protect the circuits against soft-errors. But the data that are going out of the area protected by double sampling are still vulnerable. In this paper we proposed an architectural solution that uses three latches to remove those constraints and protect the area outside the double sampling domain without adding a buffer stage.
Thierry Bonnoit, Fraidy Bouesse, Nacer-Eddine Zergainoh, Michael Nicolaidis
DATE3
2018 A soft-error resilient route computation unit for 3D Networks-on-Chips
abstract
Three-dimensional Networks-on-Chips (3D-NoCs) have emerged as an alternative to further enhance the performance, functionality, and packaging density of 2D-NoCs. However, the increasing complexity of NoC routers, the continuous miniaturization of silicon technology, the lower operating voltages, and the higher operating frequencies have made the NoC increasingly vulnerable to soft errors. In particular, transient faults occurring in the route computation unit (RCU) can provoke misrouting which may lead to severe effects such as deadlocks or packet loss, corrupting the operation of the entire chip. By combining a reliable fault detection circuit leveraging circuit-level double-sampling, with a cost-effective rerouting mechanism, we develop a full fault-tolerance solution that can efficiently detect and correct such fatal errors before the affected packets leave the router. To validate the proposed solution, we also introduce a novel method for simulation-based fault-injection based on the NoC's gate-level netlist. Experimental results obtained from a partially and vertically connected 3D-NoC indicate that our solution can provide a high level of reliability in the presence of errors, at the expense of an area and power overhead of 4.1% and 6.8% respectively.
Alexandre Coelho, Amir Charif, Nacer-Eddine Zergainoh, Juan A. Fraire, Raoul Velazco
DATE3
2018 First-Last: A Cost-Effective Adaptive Routing Solution for TSV-Based Three-Dimensional Networks-on-Chip
abstract
3D integration opens up new opportunities for future multiprocessor chips by enabling fast and highly scalable 3D Network-on-Chip (NoC) topologies. However, in an aim to reduce the cost of Through-silicon via (TSV), partially vertically connected NoCs, in which only a few vertical TSV links are available, have been gaining relevance. To reliably route packets under such conditions, we introduce a lightweight, efficient and highly resilient adaptive routing algorithm targeting partially vertically connected 3D-NoCs named First-Last. It requires a very low number of virtual channels (VCs) to achieve deadlock-freedom (2 VCs in the East and North directions and 1 VC in all other directions), and guarantees packet delivery as long as one healthy TSV connecting all layers is available anywhere in the network. An improved version of our algorithm, named Enhanced-First-Last is also introduced and shown to dramatically improve performance under low TSV availability while still using less virtual channels than state-of-the-art algorithms. A comprehensive evaluation of the cost and performance of our algorithms is performed to demonstrate their merits with respects to existing solutions.
Amir Charif, Alexandre Coelho, Masoumeh Ebrahimi, Nader Bagherzadeh, Nacer-Eddine Zergainoh
IEEE Trans. Computers5
2018 Reducing Rollback Cost in VLSI Circuits to Improve Fault Tolerance
Thierry Bonnoit, Nacer-Eddine Zergainoh, Michael Nicolaidis
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Detailed and highly parallelizable cycle-accurate network-on-chip simulation on GPGPU
abstract
As the number of processing elements in modern chips keeps increasing, the evaluation of new designs will need to account for various challenges at the NoC level. To cope with the impractically long run times when simulating large NoCs, we introduce a novel GPU-based parallel simulation method that can speed up simulations by over 250×, while offering RTL-like accuracy. These promising results make our simulation method ideal for evaluating future NoCs comprising thousands of nodes.
Amir Charif, Alexandre Coelho, Nacer-Eddine Zergainoh, Michael Nicolaidis
ASP-DAC3
2017 Rout3D: A lightweight adaptive routing algorithm for tolerating faulty vertical links in 3D-NoCs
abstract
3D integration opens up new opportunities for future multiprocessor chips by enabling fast and highly scalable 3D Network-on-Chip (NoC) topologies. However, in an aim to reduce the cost of Through-silicon via (TSV), partially vertically connected NoCs, in which only a few vertical TSV links are available, have been gaining relevance. In addition, the number of vertical paths can be expected to be further reduced due to defects and runtime failures. To reliably route packets under such conditions, we introduce a lightweight, efficient and highly resilient adaptive routing algorithm targeting partially vertically connected 3D-NoCs named “Rout3D”. It requires a very low number of virtual channels (VCs) to achieve deadlock-freedom (2 VCs in the East and North directions and 1 VC in all other directions), and guarantees packet delivery as long as one healthy TSV connecting all layers is available anywhere in the network. We combine our algorithm with a novel offline reconfiguration method requiring only 4 bits per router to maintain connectivity upon the occurrence of faults while minimizing the implementation cost. Simulation results reveal that our algorithm is capable of sustaining a very good level of performance compared to related works, in spite of using less virtual channels.
Amir Charif, Nacer-Eddine Zergainoh, Alexandre Coelho, Michael Nicolaidis
ETS2
2016 Addressing transient routing errors in fault-tolerant Networks-on-Chips
abstract
NoCs (Networks-on-Chips) are being viewed as the paradigm of choice for on-chip communication in modern SoCs (Systems-on-chips). Unfortunately, continuous technology downscaling is rendering NoC components increasingly susceptible to failure, to a point where it is no longer an option to design such systems without accounting for reliability issues. In this work, we concern ourselves with faults affecting one of the most important logical units in NoC routers, namely the Route Computation Unit. To prevent deadlocks and packet loss, which may result from misrouting, we propose a solution to detect and correct route computation errors. With compatibility with non-minimal fault-tolerant adaptive routing algorithms in mind, two detection methods are proposed. Lazy detection only detects the faults that may result in an error immediately following route computation, and leaves it to the next hops to detect and correct other errors. Strict detection detects all fatal errors before the affected packets leave the faulty router, making it possible to correct all errors using a single rerouting mechanism. Finally, we propose a novel method to safely correct routing errors and deliver all packets to their destination without resorting to retransmission.
Amir Charif, Nacer-Eddine Zergainoh, Michael Nicolaidis
ETS2
2015 MUGEN: A high-performance fault-tolerant routing algorithm for unreliable Networks-on-Chip
abstract
NoCs (Networks-on-Chip) are an attractive alternative to communication buses for SoCs (Systems-on-Chip) as they offer both high scalability and low power consumption. However, designing such systems in the nanoscale era brings up some serious concerns about reliability. Our aim is to design robust NoCs while limiting performance degradation. In this paper, we introduce several techniques meant to increase the reliability and performance of NoCs. We combine these techniques to build a fault-tolerant, deadlock-free and congestion-aware routing algorithm called MUGEN. The algorithm comprises an optimized method to exchange messages between different virtual channel classes, a selection function that uses distant router link information to avoid dead-ends and a new congestion metric used to guide routing decisions towards less congested areas. We simulate an 8×8 Mesh NoC with fault injection to evaluate each method used by MUGEN individually before comparing the full algorithm with existing works from literature. We present promising results about the proposed techniques both in terms of fault-tolerance and performance.
Amir Charif, Nacer-Eddine Zergainoh, Michael Nicolaidis
IOLTS2
2014 A generic and high-level model of large unreliable NoCs for fault tolerance and performance analysis
abstract
The integration of more and more computing cores into processors drives the adoption of larger and larger Network-on-Chips (NoCs). Concurrently, the decreasing reliability of 1 the latest technologies promotes the utilization of fault-tolerant techniques. Unfortunately, the understanding of fault-tolerant NoCs is increasingly difficult as interconnect scale up, because they require the combination of more and more complex and heterogeneous techniques. In this paper, an high-level model named VOCIS is presented, in order to ease the comprehension and analysis of large unreliable NoCs. This model features a 3D Graphical User Interface (GUI), that offers an effective and in-depth visualization of interconnects. A few analytical measurements provided directly by VOCIS are also presented, in order to assess quantitatively the impact of defects and corresponding fault-tolerant techniques.
Fabien Chaix, Nacer-Eddine Zergainoh, Michael Nicolaidis
ETS2
2014 Preliminary results of SEU fault-injection on multicore processors in AMP mode
abstract
The current technological challenge for computing systems is to use multicore processors in order to ensure reliability, improve performance, and reduce power consumption. This paper presents a method and preliminary results of SEU fault-injection campaigns performed on multicore systems in asymmetric multi-processing mode. The target used for this purpose was a Quad-core processor. This work aims at validating the efficiency of TMR fault-tolerance method and to show the weakest variables of a given application running over multicore systems.
Vanessa Vargas, Pablo Ramos, Wassim Mansour, Raoul Velazco, Nacer-Eddine Zergainoh, Jean-François Méhaut
IOLTS5
2013 Variability-aware and fault-tolerant self-adaptive applications for many-core chips
abstract
The coming era of chips consisting of billions of gates foreshadows processors containing thousands of unreliable cores. In this context, high energy efficiency will be available, under the constraint that applications leverage the large amount of computing cores, while masking frequent faults of the chip. In this paper, an high-level method is proposed to map and manage a parallel application on an unreliable many-cores processor System on Chip subject to intra-die variability.
Gilles Bizot, Fabien Chaix, Nacer-Eddine Zergainoh, Michael Nicolaidis
ETS3
2013 Variability-aware and fault-tolerant self-adaptive applications for many-core chips
abstract
The coming era of chips consisting of billions of gates foreshadows processors containing thousands of unreliable cores. In this context, high energy efficiency will be available, under the constraint that applications leverage the large amount of computing cores, while masking frequent faults of the chip. In this paper, an high-level method is proposed to map and manage a parallel application on an unreliable many-cores processor System on Chip. The approach takes into account versatile constraints relative to these processors (e.g. variability, core-level DVFS) and a generic algorithm is proposed. The distributed mapping process is based on the dynamic search of the best-suited processing node, upon task creation or node defect. An adaptive stop criteria is defined in order to balance the mapping impact and application efficiency gains. The validity of the proposition is assessed with high-level simulations, under different variability and application conditions.
Gilles Bizot, Fabien Chaix, Nacer-Eddine Zergainoh, Michael Nicolaidis
IOLTS3
2013 Fault-tolerant adaptive routing under permanent and temporary failures for many-core systems-on-chip
abstract
A fault tolerant routing algorithm for 2D Mesh Networks-on-Chip is presented in this work. It combines an adaptive routing algorithm with neighbor fault-awareness and a new traffic-balancing metric. To be able to cope with runtime failures that result in message corruption, the routing algorithm is enhanced with packet retransmission and a new packet recovery scheme. Simulation results, under various case studies, with different permanent, transient and intermittent link faults, and under different failure rates demonstrate the scalability and efficiency of the proposed algorithm to tolerate multiple failures likely encountered in deep submicron technologies.
Michael G. Dimopoulos, Yi Gang, Mounir Benabdenbi, Lorena Anghel, Nacer-Eddine Zergainoh, Michael Nicolaidis
IOLTS5
2013 Using Error Correcting Codes Without Speed Penalty in Embedded Memories: Algorithm, Implementation and Case Study
Thierry Bonnoit, Michael Nicolaidis, Nacer-Eddine Zergainoh
J. Electron. Test.3
2012 Design for test and reliability in ultimate CMOS
abstract
This session brings together specialists from the DfT, DfY and DfR domains that will address key problems together with their solutions for the 14 nm node and beyond, dealing with extremely complex chips affected by high defect levels, unpredictable and heterogeneous timing behavior, circuit degradation over time, including extreme situations related with the ultimate CMOS nodes, where all processor nodes, routers and links of single-chip massively parallel tera-device processors could comprise timing faults (such as delay faults or clock skews); a large percentage of these parts are affected by catastrophic failures; all parts experience significant performance degradations over time; and new catastrophic failures occur at low MTBF.
Michael Nicolaidis, Lorena Anghel, Nacer-Eddine Zergainoh, Yervant Zorian, Tanay Karnik, Keith A. Bowman, James W. Tschanz, Shih-Lien Lu, Carlos Tokunaga, Arijit Raychowdhury, Muhammad M. Khellah, Jaydeep P. Kulkarni, Vivek De, Dimiter R. Avresky
DATE3
2011 A fault-tolerant deadlock-free adaptive routing for on chip interconnects
abstract
Future applications will require processors with many cores communicating through a regular interconnection network. Meanwhile, the Deep submicron technology foreshadows highly defective chips era. In this context, not only fault-tolerant designs become compulsory, but their performance under failures gains importance. In this paper, we present a deadlock-free fault-tolerant adaptive routing algorithm featuring Explicit Path Routing in order to limit the latency degradation under failures. This is particularly interesting for streaming applications, which transfer huge amount of data between the same source-destination pairs. The proposed routing algorithm is able to route messages in the presence of any set of multiple nodes and links failures, as long as a path exists, and does not use any routing table. It is scalable and can be applied to multicore chips with a 2D mesh core interconnect of any size. The algorithm is deadlock-free and avoids infinite looping in fault-free and faulty 2D meshes. We simulated the proposed algorithm using the worst case scenario, with different failure rates. Experimentation results confirmed that the algorithm tolerates multiple failures even in the most extreme failure patterns. Additionally, we monitored the interconnect traffic and average latency for faulty cases. For 20×20 meshes, the proposed algorithm reduces the average latency by up to 50%.
Fabien Chaix, Dimiter R. Avresky, Nacer-Eddine Zergainoh, Michael Nicolaidis
DATE3
2011 Eliminating speed penalty in ECC protected memories
abstract
Drastic device shrinking, power supply reduction, increasing complexity and increasing operating speeds that accompanying technology scaling have reduced the reliability of nowadays ICs. The reliability of embedded memories is affected by particle strikes (soft errors), very low voltage operating modes, PVT variability, EMI and accelerated circuit aging. Error correcting codes (ECC) is an efficient mean for protecting memories against failures. A major issue with ECC is the speed penalty induced by the encoding and decoding circuits. In this paper we present an effective approach for eliminating this penalty and we demonstrate its efficiency in the case of an advanced reconfigurable OFDM modulator).
Michael Nicolaidis, Thierry Bonnoit, Nacer-Eddine Zergainoh
DATE3
2011 Efficient Fault Detection Architecture Design of Latch-Based Low Power DSP/MCU Processor
abstract
Soft errors have been emerged as an important reliability concern of modern ICs. In this work we have implemented an efficient error detection scheme in a low power DSP/MCU processor. Our scheme achieves high error detection efficiency at low hardware cost by means of an original combination of double-sampling and latch based-design into the so-called GRAAL architecture. The implementation of our design in 65nm and 45nm process nodes has confirmed the advantages of the GRAAL architecture: low area and power penalties and negligible performance degradation. Its high error detection efficiency was demonstrated by performing extensive simulations of single-event transients (SETs).
Michael Nicolaidis, Lorena Anghel, Nacer-Eddine Zergainoh
ETS4
2011 Towards a tool for implementing delay-free ECC in embedded memories
abstract
The reliability of modern Integrated Circuits is affected by nanometric scaling. In many modern designs embedded memories occupy the largest part of the die and are designed as tight as allowed by the process. So they are more prone to failures than other circuits. Error correcting codes (ECC) are a convenient mean for protecting memories against failures. A major drawback of ECC is the speed penalty induced by the encoding and decoding circuits. In [5], we propose an architecture eliminating ECC delays in both read and write paths. However, this previous work does not describe a generic set of rules enabling inserting the delay-free ECC in any design. In this paper, we present the key points of an algorithm and a related tool automating its implementation.
Thierry Bonnoit, Michael Nicolaidis, Nacer-Eddine Zergainoh
ICCD3
2011 Variability-aware task mapping strategies for many-cores processor chips
abstract
The advent of the Deep Submicron technology opens the way to many-cores processor chips. However, the variability and reliability of these processes poses new challenges. In particular, the mapping of applications will require specific strategies to leverage the plenty and diversity of the computation cores. In this work, a high-level study of the variability impact on Thousands-core processors is proposed for future technologies, based on the state-of-art VARIUS model. While many crucial details are yet unknown for these technologies, we suggest several scenarios, based on the existing literature. The obtained results are particularily suitable for Embedded Streaming applications, which both require high performance and low energy consumption. In this regard, generic task mapping strategies are proposed to improve the energy efficiency of the applications, and compared for a synthetic application. The Nearest node strategy is used as a baseline, and minimizes the communication overhead. Then, a novel energy criterion is introduced to balance the computation and communication energy consumption. While increasing the communication energy, this strategy reduces the overall consumption by up to 20%. Finally, a mapping strategy based on variability regions improves slightly the energy efficiency of the application in the presence of systematic variations.
Fabien Chaix, Gilles Bizot, Michael Nicolaidis, Nacer-Eddine Zergainoh
IOLTS4
2011 Self-Recovering Parallel Applications in Multi-core Systems
abstract
In this paper, a Self-Recovering strategy, which is able to "re-map" dynamically application tasks on a multi-core system, is presented. Based on run-time failure aware techniques, this Self-Recovering strategy guarantees seamlessly termination and delivering the expected results despite multiple node and link failures in a 2D mesh topology. It has been demonstrated, based on a statistical analysis, that the proposed technique is able to re-map the tasks of faulty nodes in a bounded number of steps. The theoretical results have been validated by simulations. The proposed technique is allowing to bypass multiple nodes, routers and links failures with a predictable number of hops. It has been demonstrated that the Motion JPEG-2000 application can be parallelized and formally represented as a Directed Acyclic Graph (DAG). It is worth noting that the proposed technique has been validated by the simulation of a 1000 cores system, in the presence of nodes and links failures up to 10%. Therefore, the proposed technique has been shown to be efficient for seamless execution of parallel streaming applications and to provide the Execution Time Reduction Ratio close to ideal.
Gilles Bizot, Dimiter R. Avresky, Fabien Chaix, Nacer-Eddine Zergainoh, Michael Nicolaidis
NCA4
2010 Fault-Tolerant Deadlock-Free Adaptive Routing for Any Set of Link and Node Failures in Multi-cores Systems
abstract
Future applications will require processors with many cores communicating through a regular interconnection network. Meanwhile, as the Deep submicron technology fore- shadows highly defective chips era, fault-tolerant designs become compulsory. In particular, the fault tolerance of a core interconnect is critical, and inevitably increases its complexity. In this paper, we present a novel adaptive routing algorithm that is able to route messages in the presence of any set of multiple nodes and links failures, as long as a path exists. Compared to the existing solutions, the proposed algorithm provides fault tolerance without using any routing table. It is scalable and can be applied to multicore chips with a 2D mesh core interconnect of any size. The algorithm is deadlock-free and avoids infinite looping in fault-free and faulty 2D meshes, based on Virtual Networks and Virtual Channels. We simulated the proposed algorithm using the worst case scenario, regarding the traffic patterns and the failure rate up to 40%. Experimentation results confirmed that the algorithm tolerates multiple failures even in the most extreme failure patterns. Additionally, we monitored the trade off between the fault tolerance and the average latency for faulty cases, as measurement of the performance degradation. The algorithm detects the interconnects partitioning and enables "preferred paths" for streaming applications.
Fabien Chaix, Dimiter R. Avresky, Nacer-Eddine Zergainoh, Michael Nicolaidis
NCA3
2009 Variability and reliability-aware application tasks scheduling and power control (Voltage and Frequency Scaling) in the future nanoscale multiprocessors system on chip
abstract
As technology scales, designing a massively parallel multi-cores system atop less reliable hardware architecture poses great challenges for researchers and designers. In this environment, ignoring variation effects when scheduling applications or when managing power with Dynamic Voltage and Frequency Scaling (DVFS) is suboptimal. We present a variation-aware multi-level scheduling and power management methodology for application-specific multiprocessor system-on-chip (MPSoC) to mitigate the impact of process variations and to optimize the power consumption. The methodology combines both static and dynamic scheduling and tackles the scheduling problem at several abstraction levels according the granularity of the application tasks, processing nodes and the on-chip communication network. The first levels of scheduling allow the local optimization, design space exploration and take into account the variations parameters. The last level, is based run-time scheduling, allows dynamic power management and global on-line optimization.
Gilles Bizot, Nacer-Eddine Zergainoh, Michael Nicolaidis
IOLTS2
2007 Buffer Size Reduction through Control-Flow Decomposition
abstract
Software synthesis from a data-flow model has been a very promising technique, especially for multimedia applications with contradicting requirements of high design complexity and fast time-to-market. In a dataflow model, buffer size is pessimistically determined through static analysis, thus results in large memory overhead even with optimization techniques such as buffer sharing and scheduling. So, reducing buffer size is one of the key issues of data-flow models. In this work, we propose a novel software synthesis technique to reduce buffer size through control-flow decomposition. We first traverse the control-flow within each actor of a data-flow graph and decompose it into a set of multiple execution paths. Then we transform the actor such that only one of the paths is executed at one invocation of the actor. The new actor may have to be invoked many times to complete the behavior of the original actor. By proper decomposition, we can make the new actor consume/produce much smaller amount of input/output data for each invocation, thereby reducing the input/output buffer size drastically. We automate the process of transformation and show the efficiency of the proposed approach through experiments with image/video multimedia applications.
Youngchul Cho, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya, Kiyoung Choi
RTCSA2
2006 Automatic delay correction method for IP block-based design of VLSI dedicated digital signal processing systems: theoretical foundations and implementation
abstract
The Intellectual Property (IP)-based design for high-throughput dedicated digital signal processing (DSP) systems is obviously an important issue for improving not only design productivity, but also design from the high level of abstraction. However, in some cases, synthesizable register transfer level (RTL) model obtained by an automatic assembly of RTL IPs can be wrong due to delays induced by implementation constraints. In this paper, we present the formalization of the problem and propose an approach called automatic delay correction method (ADCM) to solve the problem without inserting an extra interface circuitry. The approach automatically inserts control structures to manage delays induced by the use of RTL IPs. It also inserts a control structure to coordinate the execution of parallel clocked IPs. The delays may be managed by registers or by counters included in the control structure. A formal theory of ADCM is developed to guide the implementation and guarantee optimal solutions in latency and area. Through experiments with synthetic example and three real world high-throughput DSP circuits, we also show the effectiveness of our approach.
Nacer-Eddine Zergainoh, Ludovic Tambour, Ahmed Amine Jerraya
IEEE Trans. Very Large Scale Integr. Syst.1
2005 Scheduler implementation in MP SoC design
abstract
In the design of a heterogeneous multiprocessor system on chip, we face a new design problem; scheduler implementation. In this paper, we present an approach to implementing a static scheduler, which controls all the task executions and communication transactions of a system according to a pre-determined schedule. For the scheduler implementation, we consider both intra-processor and inter-processor synchronization. We also consider scheduler overhead, which is often neglected. In particular, we address the issue of centralized implementation versus distributed implementation. We investigate the pros and cons of the two different scheduler implementations. Through experiments with synthetic examples and a real world multimedia application, we show the effectiveness of our approach.
Youngchul Cho, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya
ASP-DAC4
2005 IP-block-based design environment for high-throughput VLSI dedicated digital signal processing systems
abstract
The Growing requirement on the correct design of a high performance DSP system in short time force us to use IP's in many design. In this paper, we propose an efficient IP block based design environment for high throughput VLSI Systems. The flow generates SystemC Register Transfer Level (RTL) architecture, starting from a Matlab functional model described as a netlist of functional IP. The refinement process inserts automatically control structures to treat delays induced by the use of RTL IPs. It also inserts a control structure to coordinate the execution of parallel clocked IP. The delays may be managed by registers or by counters included in the control structure. The experimentations show that the approach can produce efficient RTL architecture and allow a huge save of time.
Nacer-Eddine Zergainoh, Katalin Popovici, Ahmed Amine Jerraya, Pascal Urard
ASP-DAC1
2003 Scheduling and Timing Analysis of HW/SW On-Chip Communication in MP SoC Design
abstract
On-chip communication design includes designing software (SW) parts (operating system, device drivers, interrupt service routines, etc.) as well as hardware (HW) parts (on-chip communication network, communication interfaces of processor/IP/memory, on-chip memory, etc.). For an efficient exploration of its design space, we need fast scheduling and timing analysis. In this work, we tackle two problems (one for SW and the other for HW) in on-chip communication design. One is to incorporate the dynamic behavior of SW (interrupt processing and context switching) into on-chip communication scheduling. The other is to reduce on-chip data storage required for on-chip communication, by sharing physical communication buffers with different communication transactions. To solve the problems, we present both ILP (integer linear programming) formulation and heuristic algorithm, which enable the designer to perform efficient onchip communication scheduling and obtain accurate timing information. Experimental results show the effectiveness of our work.
Youngchul Cho, Ganghee Lee, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh
DATE5
2002 Combining a Performance Estimation Methodology with a Hardware/Software Codesign Flow Supporting Multiprocessor Systems
abstract
This paper addresses performance estimation and architecture exploration issues within the context of hardware/software codesign. We introduce a new methodology to rapidly explore the large design space encountered in hardware/software systems. The proposed methodology is based on a fast and accurate estimation approach. This estimation approach takes advantage of both system and RT levels of abstraction, and combines both static and dynamic analysis techniques, in order to obtain the best trade-off between speed and accuracy. It has been implemented as an extension to a hardware/software codesign flow to enable the exploration of a large number of multiprocessor architecture solutions from the very start of the design process. The effectiveness of the proposed methodology is illustrated by a significant application example. Experimental results indicate strong advantages of the proposed methodology.
Amer Baghdadi, Nacer-Eddine Zergainoh, Wander O. Cesário, Ahmed Amine Jerraya
IEEE Trans. Software Eng.2
2001 An efficient architecture model for systematic design of application-specific multiprocessor SoC
abstract
In this paper, we present a novel approach for the design of application specific multiprocessor systems-on chip. Our approach is based on a generic architecture model which is used as a template throughout the design process. The key characteristics of this model are its great modularity, flexibility and scalability which make it reusable for a large class of applications. In addition, it allows one accelerate the design cycle. This paper focuses on the definition of the architecture model and the systematic design flow that can be automated. The feasibility and effectiveness of this approach are illustrated by two significant demonstration examples.
Amer Baghdadi, Damien Lyonnard, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya
DATE3
2000 Towards design and validation of mixed-technology SOCs
abstract
This paper illustrates an approach to design and validation of heterogeneous systems. The emphasis is placed on devices which incorporate MEMS parts in either a single mixed-technology (CMOS + micromachining) SOC device, or alternatively as a hybrid system with the MEMS part in a separate chip. The design flow is general, and it is illustrated for the case of applications embedding CMOS sensors. In particular, applications based on finger-print recognition are considered since a rich variety of sensors and data processing algorithms can be considered. A high level multi-language/multi-engine approach is used for system specification and co-simulation. This also allows for an initial high-level architecture exploration, according to performance and cost requirements imposed by the target application. Thermal simulation of the overall device, including packaging, is also considered since this can have a significant impact in sensor performance. From the selected system specification, the actual architecture is finally generated via a multi-language co-design approach which can result in both hardware and software parts. The hardware parts are composed of available IP cores. For the case of a single chip implementation, the most important issue of embedded-core-based testing is briefly considered, and current techniques are adapted for testing the embedded cores in the SOC devices discussed.
Salvador Mir, Benoît Charlot, Gabriela Nicolescu, Philippe Coste, Fabien Parrain, Nacer-Eddine Zergainoh, Bernard Courtois, Ahmed Amine Jerraya, Márta Rencz
ACM Great Lakes Symposium on VLSI6
2000 Multi-Level Communication Synthesis of Heterogeneous Multilanguage Specification
abstract
The complexity of modern embedded systems requires the cooperation of several teams belonging to different cultures and using different languages as well as the reuse of software, hardware and communication IP modules at the early design steps. The key issue for the design of such systems is the overall system validation and the synthesis of the communication between the different subsystems. In this paper we focus on the problem of multi-level communication synthesis and show the results of the application of this methodology on an example. Designers get feedback at all design steps via the cosimulation engine that permits fast evaluation.
Fabiano Hessel, Philippe Coste, Gabriela Nicolescu, P. LeMarrec, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya
ICCD5