VLDB 2026 Research / reviewers in the wild / expert
Jean-Philippe Diguet
dblp:07/493
· DBLP profile ↗
75ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-0728-6040ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 10 · 4 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEASARD: Low-Energy Deep Neural Networks for Autonomous Search-and-Rescue DronesabstractInternational audience Panagiotis Papadakis, Isabelle Fantoni, Jean-Philippe Diguet, Matthieu Arzel |
COMPSAC | 3 |
| 2024 | Discretization Strategies for Improved Health State Labeling in Multivariable Predictive Maintenance SystemsabstractInternational audience Jean-Victor Autran, Véronique Kuhn, Jean-Philippe Diguet, Matthias Dubois, Cédric Buche |
DATA | 3 |
| 2024 | Applying a Systematic Approach to Design Human-Robot Cooperation in Dynamic EnvironmentsabstractInternational audience Sridath Tula, Marie-Pierre Pacaux-Lemoine, Emmanuelle Grislin, Anna Ma-Wyatt, Jean-Philippe Diguet |
ICINCO (2) | 5 |
| 2024 | AI4I-PMDI: Predictive maintenance datasets with complex industrial settings' irregularitiesabstractPredictive maintenance is a critical approach in various industries to enhance operational efficiency and minimize downtime by forecasting equipment failures. Publicly available datasets have been widely used to evaluate predictive maintenance algorithms, but they do not include complexities such as feet context, missing data, and irregular data acquisition. This paper presents a comparative analysis of public datasets and introduces a new one AI4I-PMDI to overcome the shortcomings of available datasets, which fail to accurately represent realistic maintenance data. AI4I-PMDI is derived from the AI4I 2020 predictive maintenance dataset, which is enhanced with irregularities that are common in real maintenance data. Moreover, machine learning algorithms were applied to both the original public dataset and the modified dataset for binary and multi-class Classification tasks. The result revealed that the modifications applied significantly impact the performances. This emphasizes the importance of using datasets adapted to the constraints encountered in complex industrial settings. Jean-Victor Autran, Véronique Kuhn, Jean-Philippe Diguet, Matthias Dubois, Cédric Buche |
KES | 3 |
| 2023 | Deep Q-Learning-Based Dynamic Management of a Robotic ClusterabstractThe ever-increasing demands for autonomy and precision have led to the development of heavily computational multi-robot system (MRS). However, numerous missions exclude the use of robotic cloud. Another solution is to use the robotic cluster to locally distribute the computational load. This complex distribution requires adaptability to come up with a dynamic and uncertain environment. Classical approaches are too limited to solve this problem, but recent advances in reinforcement learning and deep learning offer new opportunities. In this paper we propose a new Deep Q-Network (DQN) based approaches where the MRS learns to distribute tasks directly from experience. Since the problem complexity leads to a curse of dimensionality, we use two specific methods, a new branching architecture, called Branching Dueling Q-Network (BDQ), and our own optimized multi-agent solution and we compare them with classical Market-based approaches as well as with non-distributed and purely local solutions. Our study shows the relevancy of learning-based methods for task mapping and also highlight the BDQ architecture capacity to solve high dimensional state space problems. Note to Practitioners—A lot of applications in industry like area exploration and monitoring can be efficiently delegated to a group of small-size robots or autonomous vehicles with advantages like reliability and cost in respect of single-robot solutions. But autonomy requires high and increasing compute-intensive tasks such as computer-vision. On the other hand small robots have energy constraints, limited embedded computing capacities and usually restricted and/or unreliable communications that limit the use of cloud resources. An alternative solution to cope with this problem consists in sharing the computing resources of the group of robots. Previous work was a proof of concept limited to the parallelisation of a single specific task. In this paper we formalize a general method that allows the group of robots to learn on the field how to efficiently distribute tasks in order to optimize the execution time of a mission under energy constraint. We demonstrate the relevancy of our solution over market-based and non-distributed approaches by means of intensive simulations. This successful study is a necessary first step towards distribution and parallelisation of computation tasks over a robotic cluster. The next steps, not tested yet, will address hardware in the loop simulation and finally a real-life mission with a group of robots. Paul Gautier, Johann Laurent, Jean-Philippe Diguet |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2022 | Parallel IFFT/FFT for MIMO-OFDM LTE on NoC-Based FPGA
Kais Jallouli, Azer Hasnaoui, Jean-Philippe Diguet, Alireza Monemi, Salem Hasnaoui |
AINA (1) | 3 |
| 2022 | MIMO-OFDM LTE System based on a parallel IFFT/FFT on a multiprocessor platformabstractThis paper proposes a software workflow used to develop and evaluate a real time MIMO-OFDM LTE communication system. The central focus in this work is on OFDM modulation/demodulation functions which induce most of the processing time. To guarantee low latency, low energy dissipation, and high bandwidth requirements, we propose a multicore framework based on a parallel IFFT algorithm. We perform an accurate design space exploration by varying the number of tiles and by changing some NoC parameters. Based on our DSE study, the latency, bandwidth, and total energy dissipation are analyzed to select the best architecture design for the OFDM LTE system. In the worst case of a 20 MHz channel bandwidth, the best architecture selected for the OFDM LTE system is implemented with a hybrid network-on-chip using 16 tiles computing IFFT tasks. It leads to 84% reduction in latency, 73% increase in bandwidth in comparison with traditional OFDM LTE system using a single processing tile for the IFFT task. Kais Jallouli, Azer Hasnaoui, Jean-Philippe Diguet, Alireza Monemi, Salem Hasnaoui |
IWCMC | 3 |
| 2022 | Mitigating Transceiver and Token Controller Permanent Faults in Wireless Network-on-ChipabstractConventional wired Network-on-Chip (NoC) designs suffer from performance degradation due to multi-hop long-distance communication. To address such a problem, in the past decade, researchers have been focused on investigating Wireless NoC (WiNoC), which evolved as a viable solution to mitigate this communication bottleneck by using single-hop long-range wireless links. However, many researchers reported that these interconnects may suffer failure due to the complexity of implementation. Although few works in the literature tackle faults in WiNoC, none of them provides a comprehensive study related to channel access mechanisms in the presence of faults. To fill this gap, we propose a fault aware WiNoC architecture. We discuss two types of faults in wireless interconnects, namely, transceiver faults and token controller faults. We provide different fault-tolerant techniques to deal with such faults. The proposed FTWiNoC presents, on average, 17.8% and 8.9% improvement in latency compared to two different fault mitigation strategies in the literature. Navonil Chatterjee, Marcelo Ruaro, Kevin J. M. Martin, Jean-Philippe Diguet |
PDP | 4 |
| 2022 | Marine Objects Detection Using Deep Learning on Embedded Edge DevicesabstractArtificial Intelligence techniques based on convolution neural networks (CNNs) are now dominant in the field of object detection and classification. The deployment of CNNs on embedded edge devices targeting real-time inference sets a challenge due to the limited computing resources and power budgets. Several optimization techniques such as pruning, quantization and use of light neural networks enable the real-time inference but at the cost of precision degradation. However, using efficient approaches to apply the optimization techniques at training and inference stages enable high inference speed with limited degradation of detection performance. In this paper, we revisit the problem of detecting and classifying maritime objects. We investigate different versions of the You Only Look Once (YOLO), a state-of-the-art deep neural network, for real-time object detection and compare their performance for the specific application of detecting maritime objects. The trained YOLO networks are efficiently optimized targeting three recent edge devices: Nvidia Jetson Xavier AGX, AMD-Xilinx Kria KV260 Vision AI Kit, and Movidius Myriad X VPU. The proposed deployments demonstrate promising results with an inference speed of 90 FPS and a limited degradation of 2.4% in mean average precision. Dominique Heller, Mostafa Rizk, R. Douguet, Amer Baghdadi, Jean-Philippe Diguet |
RSP | 5 |
| 2022 | Enhancing embedded AI-based object detection using multi-view approachabstractObject detection based on convolutional neural network (CNN) is widely used in multitude emergent applications. Yet, the deployment of CNNs on embedded devices at the edge with reduced resources and power budget poses a real challenge. In this paper, we address this issue by enhancing the detection performance without impacting the inference speed. We investigate the use of multi-view for the same scene to achieve better detection performance. A novel system of distributed smart cameras is proposed where each camera integrates a CNN for detection. Implementation results show that using light networks on the distributed cameras can lead to better detection performance and a reduction in the overall consumed power. Zijie Ning, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet |
RSP | 4 |
| 2022 | Run-Time Remapping Algorithm of Dataflow Actors on NoC-Based Heterogeneous MPSoCsabstractMultiprocessor system-on-chip (MPSoC) platforms have been emerging as the main solution to cope with processor frequency ceiling and power density issues while still improving performances. Then, network-on-chip (NoC) has been adopted to provide the increasing number of processors with the required communication bandwidth as well as with the necessary flexibility. Video processing and streaming applications are adopting dynamic dataflow model of computation as the need for high performance parallel computing is growing. Dataflow applications executed on modern MPSoC-based architectures are becoming increasingly dynamic and more data-dependent. Different tasks execute concurrently with significant modifications in their workloads and resource demanding over time depending on the input data. Hence, adopting any static or offline dynamic scheduling for mapping tasks will not cope with the computation variations. This article introduces an original run-time mapping algorithm based on the Move Based (MB) method targeting a dedicated heterogeneous NoC-based MPSoC architecture to achieve workload balancing and optimized communication traffic. The performance of the proposed algorithm is verified by conducting cycle-accurate SystemC simulations of the adopted NoC implementing a real MPEG4-SP decoder. The obtained results reveal the effectiveness of our proposed algorithm. For various real-life videos, the proposed algorithm systematically succeeded to enhance significantly the performance. Mostafa Rizk, Kevin J. M. Martin, Jean-Philippe Diguet |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | MOL-Based In-Memory Computing of Binary Neural NetworksabstractConvolutional neural networks (CNNs) have proven very effective in a variety of practical applications involving artificial intelligence (AI). However, the layer depth of CNN deepens as user applications become more sophisticated, resulting in a huge number of operations and increased memory size. The massive amount of the produced intermediate data leads to intensive data movement between memory and computing cores causing a real bottleneck. In-memory computing (IMC) aims to address this bottleneck by directly computing inside memory, eliminating energy-intensive and time-consuming data movement. On the other hand, the emerging binary neural networks (BNNs), which is a special case of CNN, show a number of hardware-friendly properties, including memory saving. In BNN, the costly floating-point multiply-and-accumulate is replaced with lightweight bitwise XNOR and popcount operations. In this article, we propose an IMC programmable architecture targeting efficient implementation of BNN. Computational memories based on the recently introduced memristor overwrite logic (MOL) design style are employed. The architecture, which is presented in semiparallel and parallel models, efficiently executes the advanced quantization algorithm of XNOR-Net BNN. Performance evaluation based on the CIFAR-10 dataset demonstrates between$1.24\times $and$3\times $speedup and 49% and 99% energy saving compared to state-of-the-art implementations and up to 273-image/s/W throughput efficiency. Khaled Alhaj Ali, Amer Baghdadi, Elsa Dupraz, Mathieu Léonardon, Mostafa Rizk, Jean-Philippe Diguet |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | A Hybrid Adaptive Strategy for Task Allocation and Scheduling for Multi-applications on NoC-based Multicore Systems with Resource SharingabstractAllocation and scheduling of applications affect the timing response and system performance, particularly for Network-on-Chip (NoC) based multicore systems executing realtime applications. These systems with multitasking processors provide improved opportunity for parallel application execution. In dynamic scenarios, runtime task allocation improves the system resource utilization and adapts to varying application workload. In this work, we present an efficient hybrid strategy for unified allocation and scheduling of tasks at runtime. By considering multitasking capability of processors, communication cost and task timing characteristics, potential allocation solutions are obtained at design-time. These are adapted for dynamic mapping and scheduling of computation and communication workloads of real-time applications. Simulation results show that the proposed approach achieves 34.2% and 26% average reduction in network latency and communication cost of the allocated applications. Also, the deadline satisfaction of the tasks improves on average by 42.1% while reducing the allocation-time overhead by 32% when compared with existing techniques. Suraj Paul, Navonil Chatterjee, Prasun Ghosal, Jean-Philippe Diguet |
DATE | 4 |
| 2021 | ECTM: A network-on-chip communication model to combine task and message schedulability analysisabstractNetwork-on-Chips (NoC) are widely used in industrial applications since they provide communication parallelism and reduce energy consumption. The use of NoC has been recently extended to real-time systems, whose execution has to meet temporal constraints. Communication delays introduced by the network make the scheduling analysis challenging. In this article, we propose a new NoC communication model called ECTM. The main goal of this model is to assess the schedulability of dependent periodic tasks exchanging messages on a NoC. ECTM is a model allowing schedulability analysis of messages and tasks of the NoC. To achieve schedulability, ECTM produces an analysis model by transforming NoC messages to tasks in order to take into account communication delays during the scheduling analysis. Schedulability of the system is assessed using simulation over the feasibility interval with a list scheduling, ECTM supports Store-And-Forward and Wormhole NoC. In this article, we have demonstrated the correctness of the transformations of ECTM. ECTM has been implemented in a real-time scheduling analysis tool called Cheddar and we performed experiments to assess its efficiency. ECTM is more efficient than existing solutions with an improvement of 30% for Store-And-Forward NoCs and up to 100% for Wormhole NoCs, while the proposed model requires a larger computation time about 17% for Store-And-Forward NoCs. Mourad Dridi, Frank Singhoff, Stéphane Rubini, Jean-Philippe Diguet |
J. Syst. Archit. | 4 |
| 2021 | Multi-Context TCAM-Based Selective Computing: Design Space Exploration for a Low-Power NNabstractIn this paper, we propose a low-power memory-based computing architecture, called selective computing architecture (SCA). It consists of multipliers and an LUT (Look-Up Table)-based component, that is multi-context ternary content-addressable memory (MC-TCAM). Either of them is selected by input-data conditions in neural-networks (NNs). Compared with quantized NNs, a higher accurate multiplication can be performed with low-power consumption in the proposed architecture. If input data stored in the MC-TCAM appears, the corresponding multiplication results for multiple weights are obtained. The MC-TCAM stores only shorter length of input data, resulting in achieving a low-power computing. The performance of the SCA is determined by three physical parameters concerning the configuration of MC-TCAM. The power dissipation of the target NN can be minimized by exploring these parameters in the design space. The hardware based on the proposed architecture is evaluated using TSMC 65 nm CMOS technology and MTJ model. In the case of speech command recognition, the power consumption at the multiplication of the first convolutional layer in a convolutional NN is reduced by 67% compared to the solution relying only on multipliers. Ren Arakawa, Naoya Onizawa, Jean-Philippe Diguet, Takahiro Hanyu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Adaptive Task Allocation and Scheduling on NoC-based Multicore Platforms with Multitasking ProcessorsabstractThe application workloads in modern multicore platforms are becoming increasingly dynamic. It becomes challenging when multiple applications need to be executed in parallel in such systems. Mapping and scheduling of these applications are critical for system performance and energy consumption, especially in Network-on-Chip– (NoC) based multicore systems. These systems with multitasking processors offer a better opportunity for parallel application execution. Mapping solutions generated at design time may be inappropriate for dynamic workloads. To improve the utilization of the underlying multicore platform and cope with the dynamism of application workload, often task allocation is carried out dynamically. This article presents a hybrid task allocation and scheduling strategy that exploits the design-time results at runtime. By considering the multitasking capability of the processors, communication energy, and timing characteristics of the tasks, different allocation options are obtained at design time. During runtime, based on the availability of the platform resources and application requirements, the design-time allocations are adapted for mapping and scheduling of tasks, which result in improved runtime performance. Experimental results demonstrate that the proposed approach achieves an on average 11.5%, 22.3%, 28.6%, and 34.6% reduction in communication energy consumption as compared to CAM [18], DEAMS [4], TSMM [38], and CPNN [32], respectively, for NoC-based multicore platforms with multitasking processors. Also, the deadline satisfaction of the tasks of allocated applications improves on an average by 32.8% when compared with the state-of-the-art dynamic resource allocation approaches. Suraj Paul, Navonil Chatterjee, Prasun Ghosal, Jean-Philippe Diguet |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | Subutai: Speeding Up Legacy Parallel Applications Through Data SynchronizationabstractThe decrease of the performance gain dictated by Moore's Law boosted the development of manycore architectures to replace single-core architectures. These new architectures must employ parallel applications and distribute its workload over a multitude of cores to reach the desired performance. Parallel applications are harder to develop than sequential ones since the developer must guarantee data integrity using synchronization primitives. While multiple novel solutions have been proposed to speed up parallel applications through handling one type of data synchronization primitive, exceptionally few works support multiple types of synchronization primitives and legacy code. This article proposes Subutai, a hardware/software co-design solution for accelerating multiple synchronization primitives without modifying the application source code. By providing a new user library, while retaining an existing synchronization API, legacy and novel applications can benefit from our solution. Our experimental evaluation, which provides a POSIX Threads implementation, demonstrates Subutai speeds up to 2.71× and 4.61× the execution of single- and multiple-application executions, respectively. Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Jarbas Silveira, Gustavo Sanchez, Martha Johanna Sepúlveda, César A. M. Marcon, Jean-Philippe Diguet |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2020 | Broadcast Mechanism Based on Hybrid Wireless/Wired NoC for Efficient Barrier Synchronization in Parallel ComputingabstractParallel computing is essential to achieve the manycore architecture performance potential, since it utilizes the parallel nature provided by the hardware for its computing. These applications will inevitably have to synchronize its parallel execution: for instance, broadcast operations for barrier synchronization. Conventional network-on-chip architectures for broadcast operations limit the performance as the synchronization is affected significantly due to the critical path communications that increase the network latency and degrade the performance drastically. A Wireless network-on-chip offers a promising solution to reduce the critical path communication bottlenecks of such conventional architectures by providing hardware broadcast support. We propose efficient barrier synchronization support using hybrid wireless/wired NoC to reduce the cost of broadcast operations. The proposed architecture reduces the barrier synchronization cost up to 42.79% regarding network latency and saves up to 42.65% communication energy consumption for a subset of applications from the PARSEC benchmark. Hemanta Kumar Mondal, Navonil Chatterjee, Rodrigo Cataldo, Jean-Philippe Diguet |
ASP-DAC | 4 |
| 2020 | A Seamless DFT/FFT Self-Adaptive Architecture for Embedded Radar ApplicationsabstractRadar is one of the domains where adaptability is paramount. Depending on the current system state, the radar algorithms must be adapted. However, most systems include static hardware implementations on FPGA or ASIC to process the massive amount of data from multiple sensors in parallel. The classic approach is to configure hardware logic through registers to switch radar modes, requiring to hardwire all configurations. However, in embedded radar systems the hardware resources are very limited and this solution is not optimal. An innovative solution relies on the FPGA reconfigurable nature to use multiple configurations and limit resources overhead. In this paper, we present a dynamic partial reconfiguration (DPR) application for radar processing, which allows to switch between a classic discrete Fourier transform (DFT) sum and a fast Fourier transform (FFT) to enhance Doppler extraction. Our study explores the pros and cons of both methods. Based on these observations, we propose a decision method for enabling an efficient self-adaptive solution. Finally, we provide a case study with a reconfigurable radar implementation which uses the current system state to take full advantage of both Fourier transform methods. Julien Mazuet, Michel Narozny, Catherine Dezan, Jean-Philippe Diguet |
FPL | 4 |
| 2020 | Memristor Overwrite Logic (MOL) for Energy-Efficient In-Memory DNNabstractIn-memory computing is a promising solution to address the memory wall challenges in future processing systems. Substantial improvement in performance and energy efficiency is expected, in particular for data intensive applications. A typical use case is neural network applications, where large amount of data should be processed and moved between memory and processing cores. Although several recent works tried to accelerate processing through dedicated parallel hardware designs, data movement cost is still a critical technical challenge. In this context, we propose a novel programmable architecture design for in-memory deep neural networks (DNN) computation. Based on a new logic design style, namely Memristor Overwrite Logic (MOL), specialized computational memory is designed. The original architecture of the proposed computational memory allows to execute multiply-accumulate operations between stored words using MOL. Outstanding features are demonstrated with respect to other recent logic design styles based on emerging non-volatile memory technologies. Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah |
ISCAS | 4 |
| 2020 | Hardware-in-the-loop simulation with dynamic partial FPGA reconfiguration applied to computer vision in ROS-based UAVabstractHardware in the loop simulation has become a fundamental tool for the safe and rapid development of embedded systems. Dynamically and partially reconfigurable FPGA provide an energy efficient solution for high performance computing in embedded systems, such as computer vision, with limited resources. Finally 3D simulation with realistic physics simulation is required by designers of Unmanned Aerial Vehicle (UAV) and related missions. The combination of the three techniques are required to design UAV with reconfigurable HW/SW embedded systems that can self-adapt to different mission phases according to environment changes. But they require different complex and specific skills from separated communities and so are not considered simultaneously. In this paper we demonstrate a complete framework that we apply to a UAV case simulated with the well adopted Gazebo 3D simulation tool including the Ardupilot model. According to usual practices in Robotics, we use Robot Operating System (ROS) middleware over Linux that we implement on a separated Intel Cyclone V FPGA board including HW/SW interfaces. As a convincing case study we implement, besides software classical navigation tasks, a vision-based emergency-landing security task and a detect and tracking classic mission application (TLD) that can run in different HW and SW versions dynamically configured on the FPGA according to mission steps simulated with Gazebo. Erwan Moreac, El Mehdi Abdali, François Berry, Dominique Heller, Jean-Philippe Diguet |
RSP | 5 |
| 2020 | Memristive Computational Memory Using Memristor Overwrite Logic (MOL)abstractIn this article, we present a novel logic design style, namely, memristor overwrite logic (MOL), associated with an original MOL-based computational memory. MOL relies on a fully digital representation of memristor and can operate with different memristive device technologies. Its integration in memristive crossbar arrays and computational memories allows the execution of bit and vector-level primitive logic operations in two computational steps at most. Promising features and performances are demonstrated through the implementation of N -bit full addition using the proposed MOL-based computational memory. Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah, Naoya Onizawa, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | A Case Study of Primary User Arrival Prediction Using the Energy Detector and the Hidden Markov Model in Cognitive Radio NetworksabstractCognitive Radio (CR) is considered a key enabler technology for applications that require high connectivity (e.g., Smart Cities and Internet of Things), mainly because of its spectrum sensing function. In this way, CRs can sense the spectrum environment to select the best available channel for communication and they can, potentially, use licensed spectrum bands as a Secondary User (SU). However, CR has to vacate the licensed spectrum band as soon as a Primary User (PU) intends to use the channel. In this way, one of the most challenging topics in CR is the PU arrival prediction. Therefore, this paper presents a real-data study case of PUs arrivals prediction using the Hidden Markov Model (HMM) in CR. Herein, the Energy Detector (ED) is used to detect the presence of PUs. Our results show that the traditional method of combining the ED with the HMM may not be suitable in CR networks. Guilherme M. D. Santana, Rogers S. Cristo, Jean-Philippe Diguet, Catherine Dezan, Diana Pamela Moya Osorio, Kalinka Regina Lucas Jaquie Castelo Branco |
ISCC | 3 |
| 2019 | CDMA-based multiple multicast communications on WiNOC for efficient parallel computingabstractIn this work, we introduce an hybrid WiNoC, which judicially uses the wired and wireless interconnects for broadcasting/multicasting of packets. A code division multiple access (CDMA) method is used to support multiple broadcast operations originating from multiple applications executed on the multiprocessor platform. The CDMA-based WiNoC is compared in terms of network latency and power consumption with wired-broadcast/multicast NoC. Navonil Chatterjee, Hemanta Kumar Mondal, Rodrigo Cataldo, Jean-Philippe Diguet |
NOCS | 4 |
| 2019 | Design and Multi-Abstraction-Level Evaluation of a NoC Router for Mixed-Criticality Real-Time SystemsabstractA Mixed Criticality System (MCS) combines real-time software tasks with different criticality levels. In a MCS, the criticality level specifies the level of assurance against system failure. For high-critical flows of messages, it is imperative to meet deadlines; otherwise, the whole system might fail, leading to catastrophic results, like loss of life or serious damage to the environment. In contrast, low-critical flows may tolerate some delays. Furthermore, in MCS, flow performances such as the Worst Case Communication Time (WCCT) may vary depending on the criticality level of the applications. Then execution platforms must provide different operating modes for applications with different levels of criticality. To conclude, in Network-On-Chip (NoC), sharing resources between communication flows can lead to unpredictable latencies and subsequently turns the implementation of MCS in many-core architectures challenging. In this article, we propose and evaluate a new NoC router to support MCS based on an accurate WCCT analysis for high-critical flows. The proposed router, called Double Arbiter and Switching router (DAS), jointly uses Wormhole and Store And Forward communication techniques for low- and high-critical flows, respectively. It ensures that high-critical flows meet their deadlines while maximizing the bandwidth remaining for the low-critical flows. We also propose a new method for high-critical communication time analysis, applied to Store And Forward switching mode with virtual channels. For low-critical flows communication time analysis, we adapt an existing wormhole communication time analysis with share policy to our context. The second contribution of this article is a multi-abstraction-level evaluation of DAS. We evaluate the communication time of flows, the system mode change, the cost, and four properties of DAS. Simulations with a cycle-accurate SystemC NoC simulator show that, with a 15% network use rate, the communication delay of high-critical flows is reduced by 80% while communication delay of low-critical flow is increased by 18% compared to solutions based on routers with multiple virtual channels. For 10% of network interferences, using system mode change, DAS reduces the high-critical communication delays about 66%. We synthesize our router with a 28nm SOI technology and show that the size overhead is limited of 2.5% compared to the solution based on virtual channel router. Finally, we applied model checking verification techniques to automatically prove several DAS properties required by critical systems designers. Mourad Dridi, Stéphane Rubini, Mounir Lallali, Martha Johanna Sepúlveda, Frank Singhoff, Jean-Philippe Diguet |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2019 | A novel Xilinx-based architecture for 3D-graphics
Tarek Frikha, Nader Ben Amor, Jean-Philippe Diguet, Mohamed Abid |
Multim. Tools Appl. | 3 |
| 2018 | Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applicationsabstractParallel applications are essential for efficiently using the computational power of a Multiprocessor System-on-Chip (MPSoC). Unfortunately, these applications do not scale effortlessly with the number of cores because of synchronization operations that take away valuable computational time and restrict the parallelization gains. Moreover, synchronization is also a bottleneck due to sequential access to shared memory. We address this issue and introduce "Subutai", a hardware/software (HW/SW) architecture designed to distribute essential synchronization mechanisms over the Network-on-Chip (NoC). It includes Network Interfaces (NIs), drivers and a custom library of a NoC-based MPSoC architecture that speeds up the essential synchronization primitives of any legacy parallel application. Besides, we provide a fast simulation tool for parallel applications and a HW architecture of the NI. Experimental results with PARSEC benchmark show an average application speedup of 2.05 compared to the same architecture running legacy SW solutions for 36% overhead of HW architecture. Rodrigo Cataldo, Ramon Fernandes, Kevin J. M. Martin, Martha Johanna Sepúlveda, Altamiro Amadeu Susin, César A. M. Marcon, Jean-Philippe Diguet |
DAC | 7 |
| 2018 | Security aspects of neuromorphic MPSoCsabstractNeural networks and deep learning are promising techniques for bringing brain inspired computing into embedded platforms. They pave the way to new kinds of associative memories, classifiers, data-mining, machine learning or search engines, which can be the basis of critical and sensitive applications such as autonomous driving. Emerging non-volatile memory technologies integrated in the so called Multi-Processor System-on-Chip (MPSoC) architectures enable the realization of such computational paradigms. These architectures take advantage of the Network-on-Chip concept to efficiently carry out communications with dedicated distributed memories and processing elements. However, current MPSoC-based neuromorphic architectures are deployed without taking security into account. The growing complexity and the hyper-sharing of hardware resources of MPSoCs may become a threat, thus increasing the risk of malware infections and Trojans introduced at design time. Specially, MPSoC microarchitectural side-channels and fault injection attacks can be exploited to leak sensitive information and to cause malfunctions. In this work we present three contributions to that issue: i) first analysis of security issues in MPSoC-based neuromorphic architectures; ii) discussion of the threat model of the neuromorphic architectures; ii) demonstration of the correlation between SNN input and the neural computation. Martha Johanna Sepúlveda, Cezar Reinbrecht, Jean-Philippe Diguet |
ICCAD | 3 |
| 2018 | BFM: a Scalable and Resource-Aware Method for Adaptive Mission Planning of UAVsabstractUAVs must continuously adapt their mission to face unexpected internal or external hazards. This paper proposes a new BFM model (Bayesian Networks built from FMEA tables for MDP). This scalable model offers a modular and comprehensive method to incorporate different types of diagnosis modules based on BN (Bayesian Networks) and FMEA table (Failure Mode and Effects Analysis) to mission specifications expressed as a MDP (Markov Decision Processes). The BFM model implements the complete decision making process that covers both the application configurations at the embedded system level and the mission planning at the UAV level. These decisions are based on the QoS (Quality of Service) of applications, the resource use and the system and sensors health. We demonstrate on a case study for a target tracking mission that the BFM model can interface hazards and applications specifications and can improve the success and quality of the mission. To the best of our knowledge, this is the first proposal of a systematic method that integrates diagnosis modules to MDP model in order to take care of the implementation of embedded applications during a mission. Chabha Hireche, Catherine Dezan, Jean-Philippe Diguet, Luis Mejías Alvarez |
ICRA | 3 |
| 2018 | Accurate Channel Models for Realistic Design Space Exploration of Future Wireless NoCsabstractWireless Networks-on-Chip (WiNoC) are being explored for parallel applications to improve the performances by reducing the long distance/critical path communications. However, WiNoC still require precise propagation models to go beyond proof of concept and to demonstrate it can be considered as a realistic efficient alternative to wired NoC. In this paper, we present accurate 3D models based on measurements in Ka band and Electromagnetic (EM) simulations of transmission on silicon substrate in the V band and the Sub-THz band. Using these EM results, a time-domain simulation is performed using an On-Off Keying (OOK) modulation based transmission with different PA/LNA configurations. Our results highlight the type of performances and tradeoffs to be considered according to different parameters such as power output and amplifier's gain. By improving the knowledge about the signal propagation, one can conduct precise design space exploration for parallel applications. We discuss the realistic channel modeling and we present also hybrid solutions and associated limitations of WiNoC architectures. We conclude the paper with research directions to be explored to make WiNoC a reality. Ihsan El Masri, Pierre-Marie Martin, Hemanta Kumar Mondal, Rozenn Allanic, Thierry Le Gouguec, Cédric Quendo, Christian Roland, Jean-Philippe Diguet |
NOCS | 8 |
| 2018 | A modeling front-end for seamless design and generation of context-aware Dynamically Reconfigurable Systems-on-Chip
Gilberto Ochoa-Ruiz, Pamela Wattebled, Maamar Touiza, Florent de Lamotte, El-Bay Bourennane, Samy Meftali, Jean-Luc Dekeyser, Jean-Philippe Diguet |
J. Parallel Distributed Comput. | 8 |
| 2018 | Networked Power-Gated MRAMs for Memory-Based ComputingabstractEmerging nonvolatile memory technologies open new perspectives for original computing architectures. In this paper, we propose a new type of flexible and energy-efficient architecture that relies on power-gated distributed magnetoresistive random access memory (MRAM). The proposed architecture uses a network-on-chip (NoC) to interconnect MRAM-based clusters, processing elements, and managers. The NoC distributes application-specific commands to MRAM devices by means of packets. Configurable network interfaces allow to transform MRAM devices into smart units able to respond to incoming commands. In this context, three types of MRAM designs are proposed with different power-gating policies and granularities. A relevant database search engine case study is considered to illustrate the benefits of this proposed architecture. It is implemented with a sparse-neural-network approach and simulated in SystemC with different scenarios including hundreds of database queries. Hardware designs and accurate power estimations have been conducted. The obtained results demonstrate important power reduction with database hit rates of about 94%. Targeting 65-nm technology, energy savings reach 87% when compared with an static random access memory-based implementation. Moreover, a new asymmetric read/write MRAM type provides from 39% to 50% energy reduction with respect to the other fixed-granularity models. This results in a low-power, highly scalable, and configurable implementation of memory-based computing. Jean-Philippe Diguet, Naoya Onizawa, Mostafa Rizk, Martha Johanna Sepúlveda, Amer Baghdadi, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | DAS: An Efficient NoC Router for Mixed-Criticality Real-Time SystemsabstractMixed-Criticality Systems (MCS) are real-time systems characterized by two or more distinct levels of criticality. In MCS, it is imperative that high-critical flows meet their deadlines while low critical flows can tolerate some delays. Sharing resources between flows in Network-On-Chip (NoC) can lead to different unpredictable latencies and subsequently complicate the implementation of MCS in many-core architectures. This paper proposes a new virtual channel router designed for MCS deployed over NoCs. The first objective of this router is to reduce the worst-case communication latency of high-critical flows. The second aim is to improve the network use rate and reduce the communication latency for low-critical flows. The proposed router, called DAS (Double Arbiter and Switching router), jointly uses Wormhole and Store And Forward techniques for low and high-critical flows respectively. Simulations with a cycle-accurate SystemC NoC simulator show that, with a 15% network use rate, the communication delay of high-critical flows is reduced by 80% while communication delay of low-critical flow is increased by 18% compared to usual solutions based on routers with multiple virtual channels. Mourad Dridi, Stéphane Rubini, Mounir Lallali, Martha Johanna Sepúlveda, Frank Singhoff, Jean-Philippe Diguet |
ICCD | 6 |
| 2016 | Notifying memories: a case-study on data-flow applications with NoC interfaces implementationabstractNoC-based architectures overcome the limitations of traditional buses by exploiting parallelism and offer large bandwidths. NoC adoption also increases communication latency, which is especially penalising for data-flow applications (DF). We introduce the notifying memories (NM) concept to reduce this overhead. Our original approach eliminates useless memory requests. This paper demonstrates NM in the context of video coding applications implemented with dynamic DF. We have conducted cycle accurate systemC simulation of the NoC on an MPEG4 decoder to evaluate NM efficiency. The results show significant reductions in terms of latency (78%), injection rate (60%), and power savings (49%) along with throughput improvement (16%). Kevin J. M. Martin, Mostafa Rizk, Martha Johanna Sepúlveda, Jean-Philippe Diguet |
DAC | 4 |
| 2016 | Model-Based Design of Correct Controllers for Dynamically Reconfigurable ArchitecturesabstractDynamically reconfigurable hardware has been identified as a promising solution for the design of energy-efficient embedded systems. However, its adoption is limited by costly design effort, including verification and validation, which is even more complex than for nondynamically reconfigurable systems. In this article, we propose a tool-supported formal method to automatically design a correct-by-construction control of the reconfiguration. By representing system behaviors with automata, we exploit automated algorithms to synthesize controllers that safely enforce reconfiguration strategies formulated as properties to be satisfied by control. We design generic modeling patterns for a class of reconfigurable architectures, taking into account both hardware architecture and applications, as well as relevant control objectives. We validate our approach on two case studies implemented on FPGAs. Éric Rutten, Jean-Philippe Diguet, Abdoulaye Gamatié |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2016 | TBES: Template-Based Exploration and Synthesis of Heterogeneous Multiprocessor Architectures on FPGAabstractThis article describes TBES, a software end-to-end environment for synthesizing multitask applications on FPGAs. The implementation follows a template-based approach for creating heterogeneous multiprocessor architectures. Heterogeneity stems from the use of general-purpose processors along with custom accelerators. Experimental results demonstrate substantial speedup for several classes of applications. Furthermore, this work allows for reducing development costs and saving development time for the software architect, the domain expert, and the optimization expert. This work provides a framework to bring together various existing tools and optimisation algorithms. The advantages are manifold: modularity and flexibility, easy customization for best-fit algorithm selection, durability and evolution over time, and legacy preservation including domain experts' know-how. In addition to the use of architecture templates for the overall system, a second contribution lies in using high-level synthesis for promoting exploration of hardware IPs. The domain expert, who best knows which tasks are good candidates for hardware implementation, selects parts of the initial application to be potentially synthesized as dedicated accelerators. As a consequence, the HLS general problem turns into a constrained and more tractable issue, and automation capabilities eliminate the need for tedious and error-prone manual processes during domain space exploration. The automation only takes place once the application has been broken down into concurrent tasks by the designer, who can then drive the synthesis process with a set of parameters provided by TBES to balance tradeoffs between optimization efforts and quality of results. The approach is demonstrated step by step up to FPGA implementations and executions with an MJPEG benchmark and a complex Viola-Jones face detection application. We show that TBES allows one to achieve results with up to 10 times speedup to reduce development times and to widen design space exploration. Youenn Corre, Jean-Philippe Diguet, Dominique Heller, Dominique Blouin, Loïc Lagadec |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | A Dynamically Reconfigurable Multi-ASIP Architecture for Multistandard and Multimode Turbo DecodingabstractThe multiplication of wireless communication standards is introducing the need of flexible and reconfigurable multistandard baseband receivers. In this context, multiprocessor turbo decoders have been recently developed in order to support the increasing flexibility and throughput requirements of emerging applications. However, these solutions do not sufficiently address reconfiguration performance issues, which can be a limiting factor in the future. This brief presents the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising the decoding performances. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Low-complexity energy proportional posture/gesture recognition based on WBSNabstractThis paper addresses the issue of low-power posture and gesture recognition in indoor or outdoor environments without any additional equipment. For applications based on predefined postures such as environment control and physical rehabilitation, we show that low cost and fully distributed solutions, that minimize radio communications, can be efficiently implemented. Considering that radio links provide distance information, we also demonstrate that the matrix of estimated inter-node distances offers complementary information that allows for the reduction of communication load. Our results are based on a simulator that can handle various measured input data, different algorithms and various noise models. Simulation results are useful and used for the development of real-life prototype. Alexis Aulery, Jean-Philippe Diguet, Christian Roland, Olivier Sentieys |
BSN | 2 |
| 2015 | Embedded real-time localization of UAV based on an hybrid deviceabstractThis paper presents a method for localizing an Unmanned Aerial Vehicle (UAV) in indoor or outdoor environments. The approach has the ability to estimate the 3D pose of the on-board camera by using a Harris corner detector and the Levenberg-Marquardt (LM) with the Random Sample Consensus (RANSAC) algorithm to perform detection. The implementation of such computational intensive tasks in embedded system is necessary for the autonomy of UAV. Accelerators implemented on FPGA provide a solution to reach required performances. In addition to the algorithm development, we present the embedding of a real time camera pose estimation algorithm on a Xilinx System on Programmable Chip (SoPC) platform. Partitioning of our embedded application into hardware and software parts on a Zynq Board has significantly reduced the execution time when compared with software implementation, while offering necessary reconfiguration capabilities. Hanen Chenini, Dominique Heller, Catherine Dezan, Jean-Philippe Diguet, Duncan Campbell |
ICASSP | 4 |
| 2015 | Radio signature based posture recognition using WBSNabstractA body network of Inertial Measurement Units (IMUs) is a well known solution for posture recognition based on accelerometer and magnetometer data fusion. However sensors and especially the magnetometer can be disturbed by the environment. Considering a Wireless Body Sensor Network (WBSN), we propose to use available radio received power measurements as an alternative to the magnetometer. We show with simulation and real data, that the radio signal used for WBSN communications can also provide useful location information despite highly noisy Received Signal Strength Indications (RSSI). We propose a solution for the static case that leads to a very simple yet Efficient algorithm. Alexis Aulery, Christian Roland, Jean-Philippe Diguet, Zhongwei Zheng, Olivier Sentieys, Pascal Scalart |
IPSN | 3 |
| 2015 | An MDE Approach for Rapid Prototyping and Implementation of Dynamic Reconfigurable SystemsabstractThis article presents a co-design methodology based on RecoMARTE, an extension to the well-known UML MARTE profile, which is used for the specification and automatic generation of Dynamic and Partially Reconfigurable Systems-on-Chip (DRSoC). This endeavor is part of a larger framework in which Model-Driven Engineering (MDE) techniques are extensively used for modeling and via model transformations, generating executable models, which are exploited by implementation tools to create reconfigurable systems. More specifically, the methodological aspects presented in this article are concerned with expediting the conception and implementation of the hardware platform and the integration of correct by construction reconfiguration controller. This article builds upon previous research by integrating previously separated endeavors to obtain a complete PR system generation chain, which aims at shielding the designer of many of the burdensome technological and tool-specific requirements. The methodology permits for the verification of the platform description at different stages in the development process (i.e., HDL for simulation, static FPGA implementation, controller simulation and verification). Furthermore, automation capabilities embedded in the flow enable the generation of the platform description and the integration of the reconfiguration controller executive seamlessly. In order to demonstrate the benefits of the proposed approach, we present a case study in which we target the creation of an image-processing application to be deployed onto an FPGA board. We present the required modeling strategies and we discuss how the generation chains are integrated with the back-end Xilinx tools (the most mature version of PR technology) to produce the necessary executable artifacts: VHDL for the platform description and a C description of the reconfiguration controller to be executed by an embedded processor. Gilberto Ochoa-Ruiz, Sébastien Guillet, Florent de Lamotte, Éric Rutten, El-Bay Bourennane, Jean-Philippe Diguet, Guy Gogniat |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2014 | Virtual Devices for Hot-Pluggable ProcessorsabstractWhen partially reconfigurable, FPGA-based, systems allow to dynamically hot-plug processors, the number of possible software configurations increases and the dynamic sharing of hardware peripherals becomes problematic. Moreover, the debugging of application processes, which needs physical devices to communicate with remote users or debuggers, is a critical service that becomes extremely difficult to implement. This work puts forward the concept of virtual devices to reduce software complexity and isolate system services from applications. It is illustrated by a methodology making the design of debug paths easier. Several experiments show that heterogeneous systems of up to 24 hot-pluggable processors can take advantage of virtual devices. Pierre Bomel, Kevin J. M. Martin, Jean-Philippe Diguet |
DSD | 3 |
| 2014 | Mobile Augmented Reality System for Marine Navigation AssistanceabstractAugmented Reality devices are about to reach mainstream markets but applications have to meet user expectations in terms of usage and ergonomics. In this paper, we present a reallife outdoor AR application for marine navigation assistance that alleviates cognitive load issues (orientation between electronic navigational devices and bridge view) for vessels and recreational boats. First, we describe the current application and explain the requirements to draw relevant and meaningful objects. Secondly we present the software architecture of our pervasive system, which is compliant with different contexts and applications cases. Then, we detail our Marine Mobile Augmented Reality embedded System (MMARS). Finally, we present implementations on both Embedded system and smartphone. Jean-Christophe Morgère, Jean-Philippe Diguet, Johann Laurent |
EUC | 2 |
| 2014 | Extending UML/MARTE to Support Discrete Controller Synthesis, Application to Reconfigurable Systems-on-Chip ModelingabstractThis article presents the first framework to design and synthesize a formal controller managing dynamic reconfiguration, using a model-driven engineering methodology based on an extension of UML/MARTE. The implementation technique highlights the combination of hard configuration constraints using weights ( control part )—ensured statically and fulfilled by the system at runtime—and soft constraints ( decision part ) that, given a set of correct and accessible configurations, choose one of them. An application model of an image processing application is presented, then transformed and synthesized to be executed on a Xilinx platform to show how the controller, executed on a Microblaze, manages the hardware reconfigurations. Sébastien Guillet, Florent de Lamotte, Nicolas Le Griguer, Éric Rutten, Guy Gogniat, Jean-Philippe Diguet |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2013 | Stopping-Free Dynamic Configuration of a Multi-ASIP Turbo DecoderabstractThe multiplication of wireless standards is introducing the need of flexible and reconfigurable multistandard base band receivers. At the physical layer, multiprocessor turbo decoders have been recently developed in order to provide an answer to the increasing throughput requirement of emerging standards. However these solutions do not sufficiently address reconfiguration performance issues which can be a limiting factor in the future. This work focuses on the design of a reconfigurable multiprocessor architecture for turbo decoding achieving very fast reconfiguration without compromising decoding performances. Dynamic reconfiguration can be performed within a single frame decoding duration opening new perspective for reconfigurable multistandard base band receivers. For that purpose, optimizations at the processing element level and a novel bus-based configuration infrastructure are proposed. Results show that up to 64 processings elements can be dynamically configured in 5.352 μs. This low configuration latency corresponds to a single frame decoding duration when performing 6 decoding iterations for a throughput up to 666 Mbps. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Michael Hübner 0001, Jean-Philippe Diguet |
DSD | 6 |
| 2013 | Optimizations for an efficient reconfiguration of an ASIP-based turbo decoderabstractThe multiplication of wireless standards is introducing the need of flexible multi-standard baseband receivers. A multi-ASIP approach for turbo decoding is an answer to reach high throughput and high flexibility. The increasing demand of throughput for new greedy application on mobile devices and the reduction of latency between two frames create the need of an efficient reconfiguration management of such multi-ASIP platforms. In this paper, we propose to tackle reconfiguration optimization of a multi-standard ASIP for turbo decoding developed during previous work. Results show that for an area overhead of 0.012 mm2in 65 nm CMOS technology, a significant reconfiguration time optimization is achieved thanks to a reduction of the ASIP configuration load of 70%. Moreover, in a multi-ASIP context in which 8 ASIPs are implemented the configuration load is divided by ten thanks to the possibility to use a multicast mechanism for ASIP configuration loading. Vianney Lapotre, Purushotham Murugappa, Guy Gogniat, Amer Baghdadi, Jean-Philippe Diguet, Jean-Noel Bazin, Michael Hübner 0001 |
ISCAS | 5 |
| 2013 | SNet, a flexible, scalable network paradigm for manycore architecturesabstractA scalable communication paradigm for manycore architectures, called SNet (Scalable NETwork), is presented. It offers a wide range of flexibility by exploring the routing paths in a dynamic way, taking into consideration the network load. It is then followed by the data transmission phase through the chosen path. Celine Azar, Stéphane Chevobbe, Yves Lhuillier, Jean-Philippe Diguet |
NOCS | 4 |
| 2013 | Configurable memory security in embedded systemsabstractSystem security is an increasingly important design criterion for many embedded systems. These systems are often portable and more easily attacked than traditional desktop and server computing systems. Key requirements for system security include defenses against physical attacks and lightweight support in terms of area and power consumption. Our new approach to embedded system security focuses on the protection of application loading and secure application execution. During secure application loading, an encrypted application is transferred from on-board flash memory to external double data rate synchronous dynamic random access memory (DDR-SDRAM) via a microprocessor. Following application loading, the core-based security technique provides both confidentiality and authentication for data stored in a microprocessor's system memory. The benefits of our low overhead memory protection approaches are demonstrated using four applications implemented in a field-programmable gate array (FPGA) in an embedded system prototyping platform. Each application requires a collection of tasks with varying memory security requirements. The configurable security core implemented on-chip inside the FPGA with the microprocessor allows for different memory security policies for different application tasks. An average memory saving of 63% is achieved for the four applications versus a uniform security approach. The lightweight circuitry included to support application loading from flash memory adds about 10% FPGA area overhead to the processor-based system and main memory security hardware. Jérémie Crenne, Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet, Russell Tessier, Deepak Unnikrishnan |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Bus-based MPSoC Security through Communication Protection: A Latency-efficient AlternativeabstractSecurity in MPSoC is gaining an increasing attention since several years. Digital convergence is one of the numerous reasons explaining such a focus on embedded systems as much sensitive and secret data are now stored, manipulated and exchanged in these systems. Most solutions are currently built at the software level, we believe hardware enhancements also play a major role in system protection. One strategic point is the communication layer as all data goes through it. Monitoring and controlling communications enable to fend off attacks before system corruption. In this work, we propose an efficient solution with several hardware enhancements to secure data exchanges in a bus-based MPSoC. Our approach relies on low complexity distributed firewalls connected to all critical IPs of the system. Designers can deploy different security policies (access right, data format, authentication, confidentiality) in order to protect the system in a flexible way. To illustrate the benefit of such a solution, implementations are discussed for different MPSoCs implemented on Xilinx Virtex-6 FPGAs. Results demonstrate a reduction up to 33% in terms of latency overhead compared to existing efforts. Pascal Cotret, Jérémie Crenne, Guy Gogniat, Jean-Philippe Diguet |
FCCM | 4 |
| 2012 | Lightweight reconfiguration security services for AXI-based MPSoCsabstractNowadays, security is a key constraint in MPSoC development as many critical and secret information can be stored and manipulated within these systems. Addressing the protection issue in an efficient way is challenging as information can leak from many points. However one strategic component of a bus-based MPSoC is the communication architecture as all information that an attacker could try to extract or modify would be visible on the bus. Thus monitoring and controlling communications allows an efficient protection of the whole system. Attacks can be detected and discarded before system corruption. In this work, we propose a lightweight solution to dynamically update hardware firewall enhancements which secure data exchanges in a bus-based MPSoC. It provides a standalone security solution for AXI-based embedded systems where no user intervention is required for security mechanisms update. An FPGA implementation demonstrates an area overhead of around 11% for the adaptive version of the hardware firewall compared to the static one. Pascal Cotret, Guy Gogniat, Jean-Philippe Diguet, Jérémie Crenne |
FPL | 3 |
| 2012 | Modeling and synthesis of a Dynamic and Partial Reconfiguration controllerabstractThis paper presents a framework to integrate the formal synthesis of a reconfiguration controller into a Model Driven Engineering methodology used for reliable design of reconfigurable architectures. This methodology is based on an extension of UML/MARTE, GASPARD, and the aforementioned controller is obtained as a C code through a formal technique named Discrete Controller Synthesis. Taking advantage of using both modeling and synthesis techniques, the approach demonstrates an effective reduction of complexity in the specification of such reconfigurable systems. An application model of an image processing application is presented as a case study. Sébastien Guillet, Florent de Lamotte, Nicolas Le Griguer, Éric Rutten, Jean-Philippe Diguet, Guy Gogniat |
FPL | 5 |
| 2012 | A framework for high-level synthesis of heterogeneous MP-SoCabstractIn this paper we propose an ESL synthesis framework which, from the C code of an application and a description of a generic architecture, automatically explores and generates a complete synthesizable version of a H-MPSoC architecture along with the adapted code application. We developed a Design Space Exploration (DSE) algorithm that merges hardware specialization, data-parallelism exploration, processor instantiation and task mapping according to user performance and cost constraints. We also inserted HLS in the DSE loop and get fast exploration of hardware acceleration. A new ESL framework is presented, it combines our contributions with some legacy tools issued from our and another team. We validated our framework with a case study of an MJPEG decoder. Youenn Corre, Jean-Philippe Diguet, Dominique Heller, Loïc Lagadec |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | Asymmetric Cache Coherency: Policy Modifications to Improve Multicore PerformanceabstractAsymmetric coherency is a new optimization method for coherency policies to support nonuniform workloads in multicore processors. Asymmetric coherency assists in load balancing a workload and this is applicable to SoC multicores where the applications are not evenly spread among the processors and customization of the coherency is possible. Asymmetric coherency is a policy change, and consequently our designs require little or no additional hardware over an existing system. We explore two different types of asymmetric coherency policies. Our bus-based asymmetric coherency policy, generated a 60% coherency cost reduction (reduction of latencies due to coherency messages) for nonshared data. Our directory-based asymmetric coherency policy, showed up to a 5.8% execution time improvement and up to a 22% improvement in average memory latency for the parallel benchmarks Sha, using a statically allocated asymmetry. Dynamically allocated asymmetry was found to generate further improvements in access latency, increasing the effectiveness of asymmetric coherency by up to 73.8% when compared to the static asymmetric solution. John Shield, Jean-Philippe Diguet, Guy Gogniat |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2011 | Dynamic applications on reconfigurable systems: From UML model design to FPGAs implementationabstractIn this paper we propose a design methodology to explore dynamic and partial reconfiguration (DPR) of modern FPGAs. We define a set of rules in order to model DPR by means of UML and design patterns. Our approach targets MPSoPC (Multiprocessor System on Programmable Chip) which allows: a) area optimization through partial reconfiguration without performance penalty and b) increased system flexibility through dynamic behavior modeling and implementation. In our case, area reduction is achieved by reconfiguring co-processors connected to embedded processors, and flexibility is achieved by permitting new behavior to be easily added to the system. Most of the system is automatically generated by means of MDE techniques. Our modeling approach allows designers to target dynamic reconfiguration without being experts of modern FPGAs. Such a methodology allows design time speed-up and a significant reduction of the gap between hardware and software modeling. Jorgiano Vidal, Florent de Lamotte, Guy Gogniat, Jean-Philippe Diguet, Sébastien Guillet |
DATE | 4 |
| 2011 | Efficient key-dependent message authentication in reconfigurable hardwareabstractCryptographic message authentication is a growing need for FPGA-based embedded systems. In this paper a customized FPGA implementation of a GHASH function that is used in AES-GCM, a widely-used message authentication protocol, is described. The implementation limits GHASH logic utilization by specializing the hardware implementation on a per-key basis. The implemented module can generate a 128bit message authentication code in both pipelined and unpipelined versions. The pipelined GHASH version achieves an authentication throughput of more than 14 Gbit/s on a Spartan-3 FPGA and 292 Gbit/s on a Virtex-6 device. To promote adoption in the field, the complete source code for this work has been made publically-available. Jérémie Crenne, Pascal Cotret, Guy Gogniat, Russell Tessier, Jean-Philippe Diguet |
FPT | 5 |
| 2011 | Closed-loop-based self-adaptive Hardware/Software-Embedded systems: Design methodology and smart cam case studyabstractThis article presents our methodology for implementing self-adaptivness within an OS-based and reconfigurable embedded system according to objectives such as quality of service, performance, or power consumption. We detail our approach to separate application-specific decisions and hardware/software-implementation decisions at system level. The former are related to the efficiency control of applications and based on the knowledge of application engineers. The latter are generic and address the choice between various hardware and software implementations according to user objectives. The decision management is implemented as an adaptive closed-loop model. We describe how each design step may be implemented and especially how we solved the issue of stability. Finally, we present a video-tracking application implemented on a FPGA to demonstrate the effectiveness of our solution, results are given for a system built around a NIOS soft-core with μCOS II RTOS and new services for managing hardware and software tasks transparently. Jean-Philippe Diguet, Yvan Eustache, Guy Gogniat |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2010 | UML design for dynamically reconfigurable multiprocessor embedded systemsabstractIn this paper we propose a design methodology to explore partial and dynamic reconfiguration of modern FPGAs. We improve an UML based co-design methodology to allow dynamic properties in embedded systems. Our approach targets MPSoPC (Multiprocessor System on Programmable Chip) which allows area optimization through partial reconfiguration without performance penalty. In our case area reduction is achieved by reconfiguring co-processors connected to embedded processors. Most of the system is automatically generated by means of MDE techniques. Our modeling approach allows designers to target dynamic reconfiguration without being expert of modern FPGAs as many implementation details are hidden during the modeling step. Such a methodology allows design time speedup and a significant reduction of the gap between hardware and software modeling. In order to validate our approach, an object tracking application has been implemented on a reconfigurable system composed of 4 embedded processors and 3 co-processors. Dynamic reconfiguration has been performed for one co-processor which dynamically implements 3 different computations. Jorgiano Vidal, Florent de Lamotte, Guy Gogniat, Jean-Philippe Diguet, Philippe Soulard |
DATE | 4 |
| 2010 | Rapid Application Development on Multi-processor Reconfigurable SystemsabstractConsidering the ability to perform multi-processor architecture systems on FPGA, partial reconfiguration is an opportunity to improve weak soft-core performances by specializing coprocessors according to context-dependent application needs. But at the application level, there is a need for straightforward programming models that allow applications to be easily mapped on an ad hoc architecture without tedious rewriting, while at the same time ensuring efficient production code. In this paper we describe two programming libraries XTask and XFunc, which are written in C and rely on a reconfigurable MPSoC architecture model (XPSoC) and on HW/SW libraries of standard functions that can be easily used by means of HW independent API. Finally, we demonstrate the XPSoC methodology, with the design of a self-adaptive image encoding system including runtime configuration decisions. Linfeng Ye, Jean-Philippe Diguet, Guy Gogniat |
FPL | 2 |
| 2010 | MPSoC Architecture-Aware Automatic NoC Topology Design
Rachid Dafali, Jean-Philippe Diguet |
NPC | 2 |
| 2009 | A co-design approach for embedded system modeling and code generation with UML and MARTEabstractIn this paper we propose a UML/MDA approach, called MoPCoM methodology, to design high quality real-time embedded systems. We have defined a set of rules to build UML models for embedded systems, from which VHDL code is automatically generated by means of MDA techniques. We use the MARTE profile as an UML extension to describe real-time properties and perform platform modeling. The MoPCoM methodology defines three abstraction levels: abstract, execution and detailed modeling levels (AML, EML and DML, respectively). We detail the lowest MoPCoM level, DML, design rules in order to perform automatically VHDL code generation. A viterbi coder has been used as a first case study. Jorgiano Vidal, Florent de Lamotte, Guy Gogniat, Philippe Soulard, Jean-Philippe Diguet |
DATE | 5 |
| 2008 | Refining Power Consumption Estimations in the Component-based AADL Design FlowabstractThis paper presents a method that permits to quickly estimate the power consumption at the first steps of a systempsilas design. We present multi-level power models and show how to use them at different levels of the specification refinement in the component based AADL design flow. PET, a power estimation tool, is being developed in the frame of the European SPICES project. It first prototype gives, in the case of a processor binding, power consumption estimations, for software components in the AADL component assembly model, with a maximal error ranging roughly from 5% to 30% depending on the refinement level. We illustrate our approach with the power model of the PowerPC 405, and its use at different levels in the AADL flow. Eric Senn, Johann Laurent, Emmanuel Juin, Jean-Philippe Diguet |
FDL | 4 |
| 2008 | Memory security management for reconfigurable embedded systemsabstractThe constrained operating environments of many FPGA-based embedded systems require flexible security that can be configured to minimize the impact on FPGA area and power consumption. In this paper, a security approach for external memory in FPGA-based embedded systems that exploits FPGA configurability is presented. Our FPGA-based security core provides both confidentiality and integrity for data stored externally to an FPGA which is accessed by a processor on the FPGA chip. The benefits of our security core are demonstrated using four embedded applications implemented on a Stratix II device. Each application requires a collection of tasks with varying memory security requirements. Our security core is used in conjunction with a NIOS II soft processor running the MicroC/OS II operating system. An average memory and energy savings of about 64%and 16%, respectively, is achieved for the four applications versus a non-configurable, uniform security approach. Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet, Russell Tessier, Deepak Unnikrishnan, Kris Gaj |
FPT | 3 |
| 2008 | Bitstreams Repository Hierarchy for FPGA Partially Reconfigurable SystemsabstractIn this paper we present a hierarchy of bitstreams repositories for FPGA-based networked and partially reconfigurable systems. These systems target embedded systems with very scarce hardware resources taking advantage of dynamic, specific and optimized architectures. Based on FPGA integrated circuits, they require a single FPGA with a network controller and less external memories to store reconfiguration software, bitstreams and buffer pools used by today¿s standard communication protocols. Our measures, based on a real implementation, show that our repository hierarchy is functional and can download bitstreams with a reconfiguration speed ten times faster than known solutions. Pierre Bomel, Jean-Philippe Diguet, Guy Gogniat, Jérémie Crenne |
ISPDC | 2 |
| 2008 | Reconfigurable Hardware for High-Security/ High-Performance Embedded Systems: The SAFES PerspectiveabstractEmbedded systems present significant security challenges due to their limited resources and power constraints. This paper focuses on the issues of building secure embedded systems on reconfigurable hardware and proposes a security architecture for embedded systems (SAFES). SAFES leverages the capabilities of reconfigurable hardware to provide efficient and flexible architectural support for security standards and defenses against a range of hardware attacks. The SAFES architecture is based on three main ideas: (1) reconfigurable security primitives; (2) reconfigurable hardware monitors; and (3) a hierarchy of security controllers at the primitive, system and executive level. Results are presented for reconfigurable AES and RC6 security primitives and highlight the value of such an architecture. This paper also emphasizes that reconfigurable hardware is not just a technology for hardware accelerators dedicated to security primitives as has been focused on by most studies but a real solution to provide high-security and high-performance for a system. Guy Gogniat, Tilman Wolf, Wayne P. Burleson, Jean-Philippe Diguet, Lilian Bossuet, Romain Vaslin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Confiuartion Management in the Context of Self Adapative SystemsabstractThis paper presents a solution to safely and efficiently manage configurations of dynamically reconfigurable system on chip. We first define our unified RTOS-based framework for HW/SW task communication and configuration management. Then three issues are discussed and solutions given: the formalization of configuration space modeling including its different dimensions, the synchronization of configuration that mainly addresses the issue of task configuration ordering and the configuration coherency that solve the control the way a task accepts a new configuration. Finally we present the global method and give some implementation figures from a smart camera case study. Yvan Eustache, Jean-Philippe Diguet |
FPL | 2 |
| 2007 | Efficient space-time noc path allocation based on mutual exclusion and pre-reservationabstractThis paper focuses on efficient deployment of real applications over an ad hoc NoC. We propose a methodology and a tool to decide the NoC parameters and to generate the path coding within network interfaces for guarantied and best effort communications. The originality of our approach is based on two points. First, we take advantage of mutual exclusive communications. Secondly, our path allocation technique increases success opportunity and reduce buffer cost thanks to a pre-reservation step while taking into account mutual exclusions. We present real implementations for a smart camera application and a 4G telecom scheme. Samuel Evain, Jean-Philippe Diguet |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | NOC-centric Security of Reconfigurable SoCabstractThis paper presents a first solution for NoC-based communication security. Our proposal is based on simple network interfaces implementing distributed security rule checking and a separation between security and application channels. We detail a four- step security policy and show how, with usual NOC techniques, a designer can protect a reconfigurable SOC against attacks that result in abnormal communication behaviors. We introduce a new kind of relative and self-complemented street-sign routing adapted to path-based IP identification and reconfigurable architectures needs. Our approach is illustrated with a synthetic set-top box, we also show how to transform a real-life bus-based security solution to match our NOC-based architecture Jean-Philippe Diguet, Samuel Evain, Romain Vaslin, Guy Gogniat, Emmanuel Juin |
NOCS | 1 |
| 2007 | A Code Compression Method to Cope with Security Hardware OverheadsabstractCode Compression has been used to alleviate the memory requirements as well as to improve performance and/or minimize energy consumption. On the other hand, implementing security primitives on Embedded Systems is always costly in terms of area and performance. In this paper we present a code compression method, the IBC-EI (instruction based compression with encryption and integrity checking), tailored to provide integrity checking and encryption to secure processor-memory transactions. The principle is to keep the code compressed and ciphered in the memory, thus reducing the memory footprint and providing more information per memory access. For the Leon processor and a set of benchmarks from the Mediabench and MiBench suites the habitual overheads due to security trend to zero in comparison to a system without security neither compression. Eduardo Braulio Wanderley Netto, Romain Vaslin, Guy Gogniat, Jean-Philippe Diguet |
SBAC-PAD | 4 |
| 2006 | RTOS extensions for dynamic hardware / software monitoring and configuration managementabstractWe present our solution for a flexible and unified implementation of self-adaptive systems on reconfigurable architectures. This approach is based on a couple of local and global reconfiguration managers. In this paper we describe how the managers are implemented in the context of an usual RTOS and the new services we add for hardware and software monitoring, reconfiguration decision and reconfiguration control which also includes hardware and software interface modeling. Yvan Eustache, Jean-Philippe Diguet, Milad Elkhodary |
IPDPS | 2 |
| 2003 | Multi-Granularity Metrics for the Era of Strongly Personalized SOCs
Yannick Le Moullec, Nahla Ben Amor, Jean-Philippe Diguet, Mohamed Abid, Jean Luc Philippe |
DATE | 3 |
| 2001 | A Flow Control Approach for Encoded Video Applications Over ATM NetworkabstractThis paper addresses the problem of transmission of digital video communication over B-ISDN such as the ATM network. It provides the appropriate solution based on a good knowledge of both the video system interface design and broadband network capabilities. An interface between MPEG-2 and ATM network architecture is studied to improve the video visual quality. The presented approach try to overcome the difficulty imposed by traditional random cell discarding due to the bursty and variable bit rate transmission, nature of compressed video. The presented approach is guided by using a dynamic bandwidth management with the maximum flexibility via an appropriate scheduling algorithm and a new cell discarding scheme. In order to support these mechanisms, enhancement to the ATM adaptation layer is performed and a new MPEG-2 mapping strategy is also proposed. The performance evaluation have shown a significant minimization of losses ATM cells and a best video quality compared with the sequence transmitted without flow control. Ridha Djemal, Belgacem Bouallegue, Jean-Philippe Diguet, Rached Tourki |
ISCC | 3 |
| 1998 | How to Transform an Architectural Synthesis Tool for Low Power VLSI DesignsabstractHigh level synthesis (HLS) for low power VLSI design is a complex optimization problem due to the area/time/power interdependence. As few low power design tools are available, a new approach providing a modular low power synthesis method is proposed. Although based for the moment on a generic architectural synthesis tool Gaut, the use of different "commercial" tools is possible. The Gaut-w HLS tool is constituted of low power modules: high level power dissipation estimation, assignment, module selection (operators and supply voltage), optimization criteria and operators library. As illustration, power saving factors on DWT algorithms are presented. Stephane Gailhard, Nathalie Julien, Jean-Philippe Diguet, Eric Martin 0001 |
Great Lakes Symposium on VLSI | 3 |
| 1998 | Formalized methodology for data reuse: exploration for low-power hierarchical memory mappingsabstractEfficient use of an optimized custom memory hierarchy to exploit temporal locality in the data accesses can have a very large impact on the power consumption in data dominated applications. In the past, experiments have demonstrated that this task is crucial in a complete low-power memory management methodology. But effective formalized techniques to deal with this specific task have not been addressed yet. In this paper, the surprisingly large design freedom available for the basic problem is explored in-depth and the outline of a systematic solution methodology is proposed. The efficiency of the methodology is illustrated on a real-life motion estimation application. The results obtained for this application show power reductions of about 85% for the memory subsystem compared to the case without a custom memory hierarchy. These large gains justify that data reuse and memory hierarchy decisions should be taken early in the design flow. Sven Wuytack, Jean-Philippe Diguet, Francky Catthoor, Hugo De Man |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1997 | VLSI high level synthesis of fast exact least mean square algorithms based on fast FIR filtersabstractThis paper relates experiences of algorithmic transformations in High Level Synthesis, in the area of acoustic echo cancellation. The processing and memory units are automatically designed for various equivalent LMS algorithms, in the FIR case, with important computational load. The results obtained with different filter lengths, give an accurate prototyping of new fast versions of the LMS algorithm. It also show that a theoretical arithmetic reduction must be correlated to the associated increase of memory requirements. Jean-Philippe Diguet, Olivier Sentieys, Daniel Chillet, Jean Luc Philippe |
ICASSP | 1 |
| 1997 | Formalized methodology for data reuse exploration in hierarchical memory mappingsabstractArticle Formalized methodology for data reuse exploration in hierarchical memory mappings Share on Authors: J. Ph. Diguet IMEC, Kapeldreef 75, B-3001 Leuven, Belgium IMEC, Kapeldreef 75, B-3001 Leuven, BelgiumView Profile , S. Wuytack IMEC, Kapeldreef 75, B-3001 Leuven, Belgium IMEC, Kapeldreef 75, B-3001 Leuven, BelgiumView Profile , F. Catthoor IMEC, Kapeldreef 75, B-3001 Leuven, Belgium IMEC, Kapeldreef 75, B-3001 Leuven, BelgiumView Profile , H. De Man IMEC, Kapeldreef 75, B-3001 Leuven, Belgium IMEC, Kapeldreef 75, B-3001 Leuven, BelgiumView Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 30–35https://doi.org/10.1145/263272.263278Online:01 August 1997Publication History 46citation185DownloadsMetricsTotal Citations46Total Downloads185Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jean-Philippe Diguet, Sven Wuytack, Francky Catthoor, Hugo De Man |
ISLPED | 1 |