Frédéric Rousseau 0001

dblp:14/112 · DBLP profile ↗
← Back
43ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-0348-5624ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 Evaluation Tool for Stencil Application Memory Usage
Kilian McGovern, Frédéric Rousseau 0001, Henri-Pierre Charles
RSP2
2024 SpDCache: Region-Based Reduction Cache for Outer-Product Sparse Matrix Kernels
abstract
Improvements in computer performance depend increasingly on specialized accelerators and recently, numerous architectures optimized for sparse matrix kernels have been proposed, however, they do not exploit the structural properties of the matrices. SpDCache is a cache for outer-product Sparse Matrix-Vector Multiplication (SpMV) which has storage strategies optimized for both dense and sparse regions and which performs reductions locally in this cache. Real world matrices typically have a dense band which benefits from being blocked in the dense region of our cache, while the sparse regions benefit from fine-grained storage and a shift of the computation close to the main memory. We present the architectural principals of SpDCache and show that it reduces main memory traffic by -8x and increases the cache utilization by - 2x for banded matrices.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ASAP4
2024 Dedicated Hardware Accelerators for Processing of Sparse Matrices and Vectors: A Survey
abstract
Performance in scientific and engineering applications such as computational physics, algebraic graph problems or Convolutional Neural Networks (CNN), is dominated by the manipulation of large sparse matrices—matrices with a large number of zero elements. Specialized software using data formats for sparse matrices has been optimized for the main kernels of interest: SpMV and SpMSpM matrix multiplications, but due to the indirect memory accesses, the performance is still limited by the memory hierarchy of conventional computers. Recent work shows that specific hardware accelerators can reduce memory traffic and improve the execution time of sparse matrix multiplication, compared to the best software implementations. The performance of these sparse hardware accelerators depends on the choice of the sparse format, COO , CSR , etc, the algorithm, inner-product , outer-product , Gustavson , and many hardware design choices. In this article, we propose a systematic survey which identifies the design choices of state-of-the-art accelerators for sparse matrix multiplication kernels. We introduce the necessary concepts and then present, compare, and classify the main sparse accelerators in the literature, using consistent notations. Finally, we propose a taxonomy for these accelerators to help future designers make the best choices depending on their objectives.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ACM Trans. Archit. Code Optim.4
2023 A Chisel Framework for Flexible Design Space Exploration through a Functional Approach
abstract
As the need for efficient digital circuits is ever growing in the industry, the design of such systems remains daunting, requiring both expertise and time. In an attempt to close the gap between software development and hardware design, powerful features such as functional and object-oriented programming have been used to define new languages, known as Hardware Construction Languages. In this article, we investigate the usage of such languages—more precisely, of Chisel—in the context of Design Space Exploration, and propose a novel design methodology to build custom and adaptable design flows. We apply an innovative functional approach to define flexible strategies for design space exploration, based on the composition of basic exploration steps, and provide a library of basic strategies along with a proof-of-concept framework—which we believe to be the first Chisel-based DSE framework. This framework fully integrates within the ecosystem of Chisel to allow users to define their DSE processes in the same framework (and language) they use to describe their designs. We demonstrate our methodology through several use cases, illustrating how our functional approach makes it possible to consider various metrics of interest when building exploration processes—in particular, we provide a quality of service -driven exploration example. The methodology presented in this work makes use of designers’ expertise to reduce the time required for hardware design, in particular for Design Space Exploration, and its application should ease digital design and enhance hardware developers’ productivity.
Bruno Ferres, Olivier Muller, Frédéric Rousseau 0001
ACM Trans. Design Autom. Electr. Syst.3
2022 A Case for Second-Level Software Cache Coherency on Many-Core Accelerators
abstract
Cache and cache-coherence are major aspects of today's high performance computing. A cache stores data as cache-lines of fixed size, and coherence between caches is guaranteed by the cache-coherence protocol which operates on fixed size coherency-blocks. In such systems cache-lines and coherency-blocks are usually the same size and are relatively small, typically 64 bytes. This size choice is a trade-off selected for general-purpose computing: it minimizes false-sharing while keeping cache-maintenance traffic low. False-sharing is considered an unnecessary cache-coherence traffic and it decreases performances. However, for dedicated accelerator this trade-off may not be appropriate: hardware in charge of cache-coherence is expensive and not well exploited by most accelerator applications as by construction these applications minimize false-sharing. This paper investigates the possibility of an alternative trade-off of cache-coherency and cache-maintenance block size for many-core accelerators, by decoupling coherency-block and cache-lines sizes. Interests, advantages and difficulties are presented and discussed in this paper. Then we also discuss needs of software and hardware modifications in prototypes and the capability of such prototypes to evaluate different coherence-block sizes.
Arthur Vianès, Frédéric Pétrot, Frédéric Rousseau 0001
RSP3
2021 Integrating Quick Resource Estimators in Hardware Construction Framework for Design Space Exploration
abstract
Hardware design processes often come with time-consuming iteration loops, as feedbacks generally result of long synthesis runs. It is even more true when multiple different implementations need to be compared to perform Design Space Exploration (DSE). In order to accelerate such flows and increase agility of developers — closing the gap with software development methodologies — we propose to use quick feedback generating transforms based on RTL circuit analysis for quicker convergence of exploration. We also introduce an Hardware Construction Language (HCL) based methodology to build explorable circuit generators, and demonstrate such usage over a General Matrix Multiply (GEMM) Chisel implementation. We demonstrates that using RTL estimation early in the exploration process results in ×7 less synthesis runs and ×4.1 faster convergence than an exhaustive synthesis process, and still achieves state of the art performances when targetting a Xilinx VC709 FPGA.
Bruno Ferres, Olivier Muller, Frédéric Rousseau 0001
RSP3
2021 A Non-Intrusive Tool Chain to Optimize MPSoC End-to-End Systems
abstract
Multi-core systems are now found in many electronic devices. But does current software design fully leverage their capabilities? The complexity of the hardware and software stacks in these platforms requires software optimization with end-to-end knowledge of the system. To optimize software performance, we must have accurate information about system behavior and time losses. Standard monitoring engines impose tradeoffs on profiling tools, making it impossible to reconcile all the expected requirements: accurate hardware views, fine-grain measurements, speed, and so on. Subsequently, new approaches have to be examined. In this article, we propose a non-intrusive, accurate tool chain, which can reveal and quantify slowdowns in low-level software mechanisms. Based on emulation, this tool chain extracts behavioral information (time, contention) through hardware side channels, without distorting the software execution flow. This tool consists of two parts. (1) An online acquisition part that dumps hardware platform signals. (2) An offline processing part that consolidates meaningful behavioral information from the dumped data. Using our tool chain, we studied and propose optimizations to MultiProcessor System on Chip (MPSoC) support in the Linux kernel, saving about 60% of the time required for the release phase of the GNU OpenMP synchronization barrier when running on a 64-core MPSoC.
Maxime France-Pillois, Jérôme Martin, Frédéric Rousseau 0001
ACM Trans. Archit. Code Optim.3
2021 Hardware Context Switch-based Cryptographic Accelerator for Handling Multiple Streams
abstract
The confidentiality and integrity of a stream has become one of the biggest issues in telecommunication. The best available algorithm handling the confidentiality of a data stream is the symmetric key block cipher combined with a chaining mode of operation such as cipher block chaining (CBC) or counter mode (CTR). This scheme is difficult to accelerate using hardware when multiple streams coexist. This is caused by the computation time requirement and mainly by management of the streams. In most accelerators, computation is treated at the block-level rather than as a stream, making the management of multiple streams complex. This article presents a solution combining CBC and CTR modes of operation with a hardware context switching. The hardware context switching allows the accelerator to treat the data as a stream. Each stream can have different parameters: key, initialization value, state of counter. Stream switching was managed by the hardware context switching mechanism. A high-level synthesis tool was used to generate the context switching circuit. The scheme was tested on three cryptographic algorithms: AES, DES, and BC3. The hardware context switching allowed the software to manage multiple streams easily, efficiently, and rapidly. The software was freed of the task of managing the stream state. Compared to the original algorithm, about 18%–38% additional logic elements were required to implement the CBC or CTR mode and the additional circuits to support context switching. Using this method, the performance overhead when treating multiple streams was low, and the performance was comparable to that of existing hardware accelerators not supporting multiple streams.
Arif Sasongko, I. M. Narendra Kumara, Arief Wicaksana, Frédéric Rousseau 0001, Olivier Muller
ACM Trans. Reconfigurable Technol. Syst.4
2020 Implementation and Evaluation of a Hardware Decentralized Synchronization Lock for MPSoCs
abstract
Each generation of shared memory Multi-Processor System-on-Chips (MPSoCs) tend to embed more and more computing units. The cores of modern MPSoCs are often grouped into clusters communicating with each other through Networks on Chip (NoCs). Having efficient scalable synchronization mechanisms is then mandatory to benefit from the high parallelism they offer.In this work we propose an innovative hardware support for synchronization locks. First of all, a non-intrusive measurement tool-chain allows us to prove a fundamental hypothesis as to optimization of the lock mechanism: although a lock may be used, at runtime, by various cores belonging to different clusters, it is often reused by the last core which has released it. Based on this observation, we provide a hardware decentralized solution to manage dynamic re-homing of locks in a dedicated memory, close to the latest access-granted core. This reduces overall access latency and network traffic in case of reuse of the lock within the same cluster.This paper presents our solution, called Lockality, and its performance evaluation on a characteristic MPSoC running on a hardware emulator. Experiments show large gains at low level (physical lock acquisition) as well as at the application level.
Maxime France-Pillois, Jérôme Martin, Frédéric Rousseau 0001
IPDPS3
2019 Multi-Triggered Embedded Software Code Generation for Electrical Metering and Protection Applications
abstract
Code generation can be an effective way to improve quality of a product and reduce development time. But applied naively to multi-triggered applications such as electrical metering and protection algorithms, performance can be drastically degraded. In this paper we present a solution to model a multi-triggered application in Mathworks Simulink for efficient code generation using a buffering mechanism specified in the model, with little overhead in usage of resources. The performance of the presented solution is compared with different modeling approaches and with a reference manually-coded implementation.
Louis Bonicel, Roland Bohrer, Benoit Leprettre, Frédéric Rousseau 0001, Frédéric Pétrot
RSP4
2018 Linux synchronization barrier on MPSoC: Hardware/software accurate study and optimization
abstract
Providing high-performance synchronization mechanisms is a key issue to benefit from hardware parallelism offered by MPSoCs. In this paper, we focus our study on the synchronization barrier mechanism and the impact of hardware contention in shared memory clustered MPSoC. Taking advantage of a new observation methodology based on emulation, we identify Linux kernel sub-optimal services. We show how the introduction of delays in the thread awakening process improves the overall synchronization mechanism resulting in an optimization of the synchronization barrier in passive wait mode providing a large gain: 67% for 64 threads running on a 64-core architecture.
Maxime France-Pillois, Jérôme Martin, Frédéric Rousseau 0001
ASAP3
2018 Accurate MPSoC Prototyping Platform and Methodology for the Studying of the Linux Synchronization Barrier Slowdown Issues
abstract
The benefit expected from the hardware parallelism offered by Multi-Processor System on Chips (MPSoCs) is determined by the ability to design high-performance synchronization mechanisms. The complexity of modern MPSoCs does not allow anymore to design an optimized software application without confront it with the hardware platform restrictions. In this paper, we propose a methodology to study the impact of hardware contention in the synchronization barrier mechanism running on a shared memory clustered MPSoC. Taking advantage of this new observation methodology based on emulation, we identify hardware module restrictions and Linux kernel suboptimal services. We show how the introduction of delays in the thread awakening process Improves the overall synchronization mechanism. Then we detail how a combined Hardware/Software optimization for the passive wait of the synchronization barrier provides a large gain: about 60% for 64 threads running on a 64-core architecture.
Maxime France-Pillois, Jérôme Martin, Frédéric Rousseau 0001
RSP3
2017 Prototyping dynamic task migration on heterogeneous reconfigurable systems
abstract
Reconfigurable devices, such as FPGAs, have been known to offer an excellent performance and a high efficiency in computation. Due to their improving capacity and more efficient architecture recently, there are growing interests in using FPGAs as coprocessors in reconfigurable systems. However, FPGAs still lack the support in dynamic scheduling, e.g. to manage multiple tasks or users in a system. Performing runtime task relocation or load distribution is not possible unless the reconfigurable system supports dynamic task migration. Such ability requires the automation of configuration and context management in reconfigurable architecture, which is not available in the existing solutions.
Arief Wicaksana, Alban Bourge, Olivier Muller, Arif Sasongko, Frédéric Rousseau 0001
RSP5
2016 HLS-Based Methodology for Fast Iterative Development Applied to Elliptic Curve Arithmetic
abstract
High-Level Synthesis (HLS) is used by hardware developers to achieve higher abstraction in circuit descriptions. In order to shorten the hardware development time via HLS, we present an adjustment of the Iterative and Incremental Design (IID) methodology, frequently used in software development. In particular, our methodology is relevant for the development of applications with unusual complexity: the method was applied here to the development of large modular arithmetic, commonly used for cryptography applications (e.g., Elliptic Curves). Rapid feedback on circuit characteristics is used to evaluate deep architectural changes in short time, greatly reducing the time-to-market with respect to hand-made designs. In addition, our approach is highly flexible, since the same generic high-level description can be used to produce an entire set of circuits, each with different area/performance trade-offs. Thanks to the proposed approach, any change to the initial specification (e.g., the curve used) is also very fast, while it may require a large effort in the case of hand-made designs.
Simon Pontié, Alban Bourge, Adrien Prost-Boucle, Paolo Maistri, Olivier Muller, Régis Leveugle, Frédéric Rousseau 0001
DSD7
2016 Demonstration of a context-switch method for heterogeneous reconfigurable systems
abstract
Nowadays, FPGAs are integrated in high-performance computing systems, servers, or even used as accelerators in System-on-Chip (SoC) platforms. Since the execution is performed in hardware, FPGA gives much higher performance and lower energy consumption compared to most microprocessor-based systems. However, the room to improve FPGA performance still exists, e.g. when it is used by multiple users. In multi-user approaches, FPGA resources are shared between several users. Therefore, one must be able to interrupt a running circuit at any given time and continue the task at will. An image of the state of the running circuit (context) is saved during interruption and restored when the execution is continued. The ability to extract and restore the context is known as context-switch.
Arief Wicaksana, Alban Bourge, Olivier Muller, Frédéric Rousseau 0001
FPL4
2016 A survey of NoC evaluation platforms on FPGAs
abstract
Networks-on-chip (NoCs) have become a de facto communication standard for many core systems-on-chip (SoCs). A NoC has large design space composed of several parameters such as routing algorithm, task mapping, among others. SoC designers deeply rely on automatic evaluation tools in order to deal with the complexity of NoC design. An important class of NoCs evaluation tools are the platforms based on FPGAs, which improve the evaluation time and precision when compared to other solutions. There are different architectures of FPGA-based NoC evaluation tools. Details are scattered among several papers, making a comparative analysis hard to accomplish. This paper presents a comprehensive overview of FPGA tools for NoC evaluation. Our analysis covers aspects like network architecture, traffic generation and interface to the host PC. This provides insight on the platforms and their usefulness for different NoC evaluation tasks.
Otávio Alcântara de Lima Júnior, Weslley N. Costa, Virginie Fresse, Frédéric Rousseau 0001
FPT4
2016 On-board non-regression test of HLS tools targeting FPGA
abstract
High-Level Synthesis (HLS) has opened an opportunity for software programmers to target FPGA more rapidly. When developing HLS tools, tests are desirable to ensure their function, reliability and performance. When modifications are applied to a tool, Non-Regression Test (NRT) asserts that the changes have intended effect while Regression Test (RT) verifies that the tool still performs correctly without unwanted behaviour.
Arief Wicaksana, Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001, Arif Sasongko
RSP4
2016 Synthesis of dependency-aware traffic generators from NoC simulation traces
Otávio Alcântara de Lima Júnior, Virginie Fresse, Frédéric Rousseau 0001, Hamed Sheibanyrad
J. Syst. Archit.3
2016 Dynamic many-process applications on many-tile embedded systems and HPC clusters: The EURETILE programming environment and execution platforms
Pier Stanislao Paolucci, Andrea Biagioni, Luis Gabriel Murillo, Frédéric Rousseau 0001, Lars Schor, Laura Tosoratto, Iuliana Bacivarov, Robert Buecs, Clément Deschamps, Ashraf El Antably, Roberto Ammendola, Nicolas Fournel, Ottorino Frezza, Rainer Leupers, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Elena Pastorelli, Devendra Rai, Davide Rossetti, Francesco Simula
J. Syst. Archit.4
2016 Generating Efficient Context-Switch Capable Circuits through Autonomous Design Flow
abstract
Commercial off-the-shelf (COTS) Field-Programmable Gate Arrays (FPGAs) are becoming increasingly powerful. In addition to their huge hardware resources, they are also integrated into complete systems on chips (SOCs), e.g., in the latest Xilinx Zynq or Altera Stratix platforms. However, cooperation between FPGAs and their surroundings, and the flexibility of hardware task management could still be improved. For instance, mechanisms have yet to be automated to allow multi-user approaches. A reconfigurable resource can be shared between applications or users only if it has a context-switch ability allowing applications to be paused and resumed in response to system demands. Here, we present a high-level synthesis (HLS) design flow producing a context-switch-capable circuit. The design flow manipulates the intermediate representation of an HLS tool to build the context extraction mechanism and to optimize performance for the circuit produced. The method is based on efficient checkpoint selection and insertion of a powerful scan-chain into the initial circuit. This scan-chain can extract flip-flops or memory content. Experiments with the system produced show that it has a low hardware overhead for many benchmark applications, and that the hardware added has a negligible impact on application performance. Comparisons with current standard methods highlight the efficiency of our contributions.
Alban Bourge, Olivier Muller, Frédéric Rousseau 0001
ACM Trans. Reconfigurable Technol. Syst.3
2015 Integrating Task Migration Capability in Software Tool-Chain for Data-Flow Applications Mapped on Multi-tiled Architectures
abstract
Fully distributed memory multi-processor systems-on-chip MPSoCs implemented in a multi-tiled architecture provide promising platforms to support parallel data-flow application. Tiles are connected by network-on-chip NoC, each contains a core with necessary peripherals and a communication device. Such systems are susceptible to reliability issues like thermal spots. Task migration still provides an effective system-level solution for such issues. An agent based task migration solution is designed to target tiled MPSoCs. These agents are responsible for executing migration. In order to execute task migration, a middleware1 layer is developed to provide necessary services used by the agents. Also, Agents use information about both application(s) task graph and application(s) mapping on different tiles so as to be able to control right tasks. Since number of tiles is continuously increasing thanks to advancements in transistor scaling technology, automatic software generation tool-chain is no longer optional. In this work, we expand a software automatic generation tool-chain and a task migration solution, all designed for tiled MPSoCs. We emphasize on how this task migration solution is integrated in this software tool-chain so that generated software is equipped with task migration capability transparently from application developers. We show how agents are placed with applications and how necessary information for such agents are generated and linked with them. The tool-chain is capable of generating code for ARM based simulation and x86 real hardware platforms. We show experimental results of task migration memory and performance overheads.
Ashraf El Antably, Nicolas Fournel, Frédéric Rousseau 0001
DSD3
2015 Automatic High-Level Hardware Checkpoint Selection for Reconfigurable Systems
abstract
Modern FPGAs provide great computational power and flexibility but there is still room for improving their performances. For example multi-user approaches are particularly underdeveloped as they require specific mechanisms still to be automated. Sharing an FPGA resource between applications or users requires a context switch ability. The latter enables pausing and resuming applications at system demand. This paper presents a method that automatically selects a good execution point, called hardware checkpoint, to perform a context switch on an FPGA. The method relies on a static analysis of the finite state machine of a circuit to select the checkpoint states. The obtained selection ensures that the context switch mechanism respects a given latency and tries to minimize the mechanism costs. The method takes advantage of its integration in an open-source HLS tool and preliminary results highlight its efficiency.
Alban Bourge, Olivier Muller, Frédéric Rousseau 0001
FCCM3
2015 A Novel Method for Enabling FPGA Context-Switch (Abstract Only)
abstract
Modern FPGAs provide great computational power, flexible resources and a versatile environment. Managing to obtain the best of these three worlds is rather complicated given the actual design flows. Our work focus on enabling task multiplexing, as part of a more flexible FPGA usage. Task multiplexing in FPGAs raises indeed a lot of questions. Multiplexing the usage of a reconfigurable fabric is leading to a better utilization of its surface because it offers to share its resources not only in space (number of slices allocated to a task) but also in time (tasks are allowed in time slots). The base mechanism known as context-switch consists in removing a task after its allowed time slot has passed. The first step toward efficiently multiplex tasks in a reconfigurable fabric is to decide when this removal will have the least possible impact on the system. This poster presents our preliminary results concerning what we consider as necessary in order to enable such a feature. Our work focus on finding automatically the best instants of the task execution in order to effectively remove a running task from the FPGA, taking into account the time needed to extract a relevant context necessary to restart it later. This instant selection is performed at a high level of abstraction, enabling us to make choices with an accurate knowledge of the task nature and specificities. The second part of this poster presents the entire mechanism which makes use of the previously selected slots in order to switch between tasks.
Alban Bourge, Olivier Muller, Frédéric Rousseau 0001
FPGA3
2015 Dynamic data flow analysis for NoC based application synthesis
abstract
Network-on-Chip (NoC) is an interesting communication fabric for multi processing element architectures that benefits from the parallelism of algorithms. We present a method that uses a symbolic execution technique to extract the parallelism of an application to be mapped on FPGAs using the flexibility of a NoC communication infrastructure and the properties of a high level programming language. An application specific hardware is then generated using a High Level Synthesis flow. We provide a dedicated mechanism for data paths reconfiguration that allows different applications to run on the same set of processing elements. Thus, the output design is programmable and has a processor-less distributed control. This approach of using NoCs enables us to automatically design generic architectures that can be used on FPGA servers for High Performance Reconfigurable Computing.We validate our method on binomial tree applications used for option pricing on FPGAs.
Matthieu Payet, Virginie Fresse, Frédéric Rousseau 0001, Pascal Remy
RSP3
2014 Evaluation of SNMP-like protocol to manage a NoC emulation platform
abstract
The Networks-on-Chip (NoCs) are currently the most appropriate communication structure for many-core embedded systems. An FPGA-based emulation platform can drastically reduce the time needed to evaluate a NoC, even if it is composed by tens or hundreds of distributed components. These components should be timely managed in order to execute an evaluation traffic scenario. There is a lack of standard protocols to drive FPGA-based NoC emulators. Such protocols could ease the integration of emulation components developed by different designers. In this paper, we evaluate a light version of SNMP (Simple Network Management Protocol) to manage an FPGA-based NoC emulation platform. The SNMP protocol and its related components are adapted to a hardware implementation. This facilitates the configuration of the emulation nodes without FPGA-resynthesis, as well as the extraction of emulation results. Some experiments highlight that this protocol is quite simple to implement and very efficient for a light resources overhead.
Otávio Alcântara de Lima Júnior, Virginie Fresse, Frédéric Rousseau 0001
FPT3
2014 EURETILE Design Flow: Dynamic and Fault Tolerant Mapping of Multiple Applications Onto Many-Tile Systems
abstract
EURETILE investigates foundational innovations in the design of massively parallel tiled computing systems by introducing a novel parallel programming paradigm and a multi-tile hardware architecture. Each tile includes multiple general-purpose processors, specialized accelerators, and a fault-tolerant distributed network processor, which connects the tile to the inter-tile communication network. This paper focuses on the EURETILE software design flow, which provides a novel programming environment to map multiple dynamic applications onto a many-tile architecture. The elaborated high-level programming model specifies each application as a network of autonomous processes, enabling the automatic generation and optimization of the architecture-specific implementation. Behavioral and architectural dynamism is handled by a hierarchically organized runtime-manager running on top of a lightweight operating system. To evaluate, debug, and profile the generated binaries, a scalable many-tile simulator has been developed. High system dependability is achieved by combining hardware-based fault awareness strategies with software-based fault reactivity strategies. We demonstrate the capability of the design flow to exploit the parallelism of many-tile architectures with various embedded and high performance computing benchmarks targeting the virtual EURETILE platform with up to 192 tiles.
Lars Schor, Iuliana Bacivarov, Luis Gabriel Murillo, Pier Stanislao Paolucci, Frédéric Rousseau 0001, Ashraf El Antably, Robert Buecs, Nicolas Fournel, Rainer Leupers, Devendra Rai, Lothar Thiele, Laura Tosoratto, Piero Vicini, Jan Weinstock
ISPA5
2014 Lightweight task migration in embedded multi-tiled architectures using task code replication
abstract
With such ongoing sophistication in embedded applications, higher computational powers are becoming more required. As a result, a wide transition to multi-processor system on chip has been adopted. Our study focuses on fully distributed memory MPSoC which is implemented in multi-tiled architecture. A tile contains at least one processor and associated peripherals with a distributed network processor which is responsible for inter-tile communications. All tiles are connected in a 3D torus network. In this paper, we present a solution for a lightweight task migration on such architectures. It is based on wise task code replication in statically chosen locations (tiles). It provides the system with the ablility to remap its tasks at runtime. This work emphasizes on solving all issues arising from communication inconsistency shedding the light on implementation details. Experiments show the effectiveness of the approach, and detail performances and limitations. The solution has been implemented on a multi-tiled virtual ARM-CortexA9 based platform with an embedded operating system.
Ashraf El Antably, Nicolas Fournel, Frédéric Rousseau 0001
RSP3
2014 Device driver generation targeting multiple operating systems using a model-driven methodology
abstract
We present a new device driver generation approach capable of automatically generating a large portion of device drivers code, and this for different operating systems (OSes). This approach is based on a model-driven methodology, where a tiny language is utilized to model the device features and abstract low-level complexities of a driver. The approach can handle different driver architectures. We demonstrate the genericity of the approach by applying it to a fairly mature device class that has standardized interfaces, and also to a brand-new device that has significant functionality differences. The code was generated for two OSes, one targeting the embedded space and the other a full featured one.
Guillaume Godet-Bar, Frédéric Rousseau 0001, Frédéric Pétrot
RSP3
2014 Fast and standalone Design Space Exploration for High-Level Synthesis under resource constraints
Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001
J. Syst. Archit.3
2013 A Fast and Autonomous HLS Methodology for Hardware Accelerator Generation under Resource Constraints
abstract
This paper presents a new methodology for hardware accelerator generation, in the context of High Level Synthesis (HLS) for Field Programmable Gate Array (FPGA) components. The very high computing capacity available in the latest FPGA makes them choice targets in High-Performance Computing (HPC) as well as embedded systems. For a much wider adoption of FPGA as general-purpose computing devices, the proposed HLS design flow leverages the users from all issues related to circuit structure fine-tuning. The HLS methodology is autonomous and produces RTL descriptions quickly, under only global resource and frequency constraints. This is achieved by performing incremental transformations of the input design description. The low complexity of the Design Space Exploration (DSE) algorithm and its good usage of all internal circuit structure constraints, make this HLS methodology very fast and able to generate pertinent solutions. Moreover, the generated circuit is designed to fit into the targeted FPGA or a given partition of it. Such a methodology leads to autonomous, fast and transparent DSE, all these issues known to limit the use of HLS and FPGA. Results on several benchmarks highlight the capabilities of our DSE methodology. The results show a high generation speed-up compared to other existing HLS approaches, while preserving correct performance of the generated circuits.
Adrien Prost-Boucle, Olivier Muller, Frédéric Rousseau 0001
DSD3
2013 FlexOE: A congestion-aware routing algorithm for NoCs
abstract
Networks-on-Chip (NoCs) are currently the most appropriate communication structure for many-core embedded systems. Those networks support many real-time data flows. Their performance depends directly on the routing strategy. In this paper, we present a new congestion-aware routing algorithm (FlexOE) based on a simple and flexible scheme of prioritized sets of rules. These sets of rules are based on the Odd-Even turn model, minimal paths checking, congestion information from adjacent routers and availability of output path. The algorithm FlexOE developed is integrated on a Hermes NoC, and then implemented on an FPGA. The evaluation results point out that FlexOE has greater performances than reference algorithms for some test scenarios and similar performances for others test scenarios.
Otávio Alcântara de Lima Júnior, Virginie Fresse, Frédéric Rousseau 0001
RSP3
2012 Enhancing non-linear kernels by an optimized memory hierarchy in a High Level Synthesis flow
abstract
Modern High Level Synthesis (HLS) tools are now efficient at generating RTL models from algorithmic descriptions of the target hardware accelerators but they still do not manage memory hierarchies. Memory hierarchies are efficiently optimized by performing code transformations prior to HLS in frameworks which exploit the linearity of the mapping functions between loop indexes and memory references (called linear kernels). Unfortunately, non-linear kernels are algorithms which do not benefit of such classical frameworks, because of the disparity of the non-linear functions to compute their memory references. In this paper we propose a method to design non-linear kernels in a HLS flow, which can be seen as a code pre-processing. The method starts from an algorithmic description and generates an enhanced algorithmic description containing both the non-linear kernel and an optimized memory hierarchy. The transformation and the associated optimization process provides a significant gain when compared to a standard optimization. Experiments on benchmarks show an average reduction of 28% of the external memory traffic and about 32 times of the embedded memory size.
Stéphane Mancini, Frédéric Rousseau 0001
DATE2
2012 Multi-device Driver Synthesis Flow for Heterogeneous Hierarchical Systems
abstract
Heterogeneous hierarchical architectures result from the interconnection of several heterogeneous MPSoCs through an efficient communication infrastructure. Each communication between two processing units requires the use of hardware devices managed by a software driver responsible to initialize, handle and complete the communication. This paper describes a multi-device driver synthesis flow based on available communication paths of the architecture. A small set of generic driver templates is available in a driver library. The flow is in charge to select one correct driver template from the library, and then to configure and specialize it in order to produce the source code. The effectiveness of our approach is illustrated by a significant example.
Alexandre Chagoya-Garzon, Frédéric Rousseau 0001, Frédéric Pétrot
DSD2
2012 Case study: Deployment of the 2D NoC on 3D for the generation of large emulation platforms
abstract
The evaluation of Network-On-Chip (NoC) architectures is an up to date problem in the design of System-on-Chip. Emulation on FPGA (Field Programmable Gate Array) is used to cover all possible NoC solutions in a reduced exploration time. Emulation requires multi-FPGA platform as the resources for large NoC is important and cannot be handling by one FPGA. In the same time, SoC community is exploring 3D technology for the next generation of large SoC with 3D NoC, making emulation more complex. This paper presents a case study of the deployment of the 2D NoC structure to 3D. A design flow is proposed for the automatic generation of a NoC targeting 3D on multi-FPGAs. The flow integrates emulation blocks used for the validation and exploration on the NoC. With this automatic aided tool, the designer can evaluate and explore the NoC architecture and extract performances of the NoC regardless of the multi-component platform. One may expect a communication performance improvement using an adapted partitioning of the NoC, as highlighted by the results given in this paper.
Virginie Fresse, Zhiwei Ge, Junyan Tan, Frédéric Rousseau 0001
RSP4
2011 Semi-automation of Configuration Files Generation for Heterogeneous Multi-tile Systems
abstract
Heterogeneous Multi-Processor System-on-Chips (HMPSoCs) offer an attractive alternative to homogeneous systems to achieve the increasing requirements of modern media-processing applications. Such systems take advantage of the heterogeneity of their processing units (RISC vs. VLIW) combined with efficient memory architecture and a specific communication infrastructure. However, the complexity of such architectures requires efficient programming tools. They rely on software generation flows that have already been widely studied. Connecting several HMPSoCs to form a heterogeneous system is one solution to target the highly demanding applications belonging to the high performance-computing world. Binary code generation for such many-processor architectures faces new challenges. Indeed, binary generation requires a back-end part composed of processor-specific tools like compilers or linkers, that can no longer be configured by hand when dealing with hundreds of processors. We address in this paper the need to automate the whole configuration process of binary generation flows through the study of two important configuration aspects: the memory mapping and communication configurations. We propose a prototype including a semi-automatic configuration generation module that has been successfully applied on a particularly complex application (LQCD).
Alexandre Chagoya-Garzon, Nicolas Poste, Frédéric Rousseau 0001
COMPSAC3
2009 Abstract Description of System Application and Hardware Architecture for Hardware/Software Code Generation
abstract
The deployment of a system application over a hardware architecture is a costly phase in the design process. This cost increases when dealing with complex applications in terms of computation requirements and exchange of data and for advanced architectures with complex and configurable communication infrastructures. The usage of abstract models for application, architecture and mapping is a key element for automatic hardware/software code generation and for the final deployment. In this paper, we present languages for abstract modeling of application, architecture, meta-mapping and mapping and we introduce a code generation flow. The use of those models allows the extraction and exploitation of architectural and application information for specific code generation to a target platform. A case study of modeling and deploying a complex 4G telecommunication application on a heterogeneous and multi core platform is presented.
Amin El Mrabti, Hamed Sheibanyrad, Frédéric Rousseau 0001, Frédéric Pétrot, Romain Lemaire, Jérôme Martin
DSD3
2008 Platform-based software design flow for heterogeneous MPSoC
abstract
Current multimedia applications demand complex heterogeneous multiprocessor architectures with specific communication infrastructure in order to achieve the required performances. Programming these architectures usually results in writing separate low-level code for the different processors (DSP, microcontroller), implying late global validation of the overall application with the hardware platform. We propose a platform-based software design flow able to efficiently use the resources of the architecture and allowing easy experimentation of several mappings of the application onto the platform resources. We use a high-level environment to capture both application and architecture initial representations. An executable software stack is generated automatically for each processor from the initial model. The software generation and validation is performed gradually corresponding to different software abstraction levels. Specific software development platforms (abstract models of the architecture) are generated and used to allow debugging of the different software components with explicit hardware-software interaction. We applied this approach on a multimedia platform, involving a high performance DSP and a RISC processor, to explore communication architecture and generate an efficient executable code for a multimedia application. Based on automatic tools, the proposed flow increases productivity and preserves design quality.
Katalin Popovici, Xavier Guerin, Frédéric Rousseau 0001, Pier Stanislao Paolucci, Ahmed Amine Jerraya
ACM Trans. Embed. Comput. Syst.3
2007 Flexible Application Software Generation for Heterogeneous Multi-Processor System-on-Chip
abstract
Multimedia applications require heterogeneous multiprocessor architectures with specific I/O components in order to achieve computation and communication performances. The different processors run different software stacks, which are composed by the application s tasks and a hardware dependent software (HDS). The HDS contains an operating system, a specific communication library and a hardware abstraction layer (HAL), granting accesses to hardware resources. Building these software stacks may be the trouble maker of the MP-SoC design process when trying to reduce its time-to-market. In this paper, we present our application software generation flow and tools starting from a high level application model. They are able to handle heterogeneous MP-SoC, running multiple software stacks while using different operating systems and communication models. The application software generation tool builds the application s sofware stacks by producing optimized and multi-tasked C code and using a flexible operating system and communication programming interfaces management. In order to validate the effectiveness of our approach, we generated the software stacks of a Motion JPEG decoder, partitioned and mapped on an off-the-shelf multimedia platform.
Xavier Guerin, Katalin Popovici, Wassim Youssef, Frédéric Rousseau 0001, Ahmed Amine Jerraya
COMPSAC (1)4
2006 A Verification Tool Implementation using Introspection Mechanism
Michel Metzger, Frédéric Bastien, Frédéric Rousseau 0001, Julie Vachon, El Mostapha Aboulhamid
FDL3
2002 Automatic generation of embedded memory wrapper for multiprocessor SoC
abstract
Embedded memory plays a critical role to improve performances of systems-on-chip (SoC). In this paper, we present a new methodology for embedded memory design in the case of application specific multiprocessor system-on-chip. This approach facilitates the integration of standard memory components. The concept of memory wrapper allows automatic adaptation of physical memory interfaces to a communication network that may have a different number of access ports. We give also a generic architecture to produce this memory wrapper. This approach has successfully been applied on a low-level image processing application.
Ferid Gharsalli, Samy Meftali, Frédéric Rousseau 0001, Ahmed Amine Jerraya
DAC3
1997 A codesign experiment in acoustic echo cancellation GMDF
abstract
Continuous advances in processor and ASIC technologies enable the integration of more and more complex embedded systems. Embedded systems have become commonplace in recent years. Since their implementations generally require the use of heterogeneous resources (e.g., processor cores, ASICs) in one system with hard design constraints, the importance of hardware/software codesign methodologies increases steadily. HW/SW codesign approaches consist generally of HW/SW partitioning and scheduling, constrained code generation, and hardware and interface synthesis. This article presents the codesign of an industrial experiment in acoustic echo cancellation (GMDFα algorithm); and emphasizes the partitioning and communication synthesis steps. This experiment brings to light interesting problems such as data and program distribution between system memories and the modeling of communications in the partitioning process
Laurent Freund, Michel Israël, Frédéric Rousseau 0001, J. M. Bergé, Michel Auguin, Cécile Belleudy, Guy Gogniat
ACM Trans. Design Autom. Electr. Syst.3
1996 Hardware/Software Partitioning for Telecommunications Systems
abstract
Telecommunications systems, like other embedded systems, are dataflow systems, easily represented by a set of tasks and precedence constraints. The main goal of the design of such systems is to determine for each task the assignment (hardware or software), the scheduling and resources required. We consider assignment and scheduling to be closely linked in hardware/software partitioning and therefore propose a new approach to hardware/software partitioning using task scheduling. This approach is a list scheduling algorithm, based on the calculation of forces. The results obtained on a telecommunications system (acoustic echo canceller) are then described.
Frédéric Rousseau 0001, J. M. Bergé, Michel Israël
COMPSAC1
1995 Adaptation of force-directed scheduling algorithm for hardware/software partitioning
abstract
An algorithm for the hardware/software partitioning problem is presented. In data flow systems, task scheduling modifies global characteristics and allows different implementation solutions. Our algorithm is based on assignment and scheduling algorithms which are well known in high-level synthesis. At each iteration, one task is scheduled if it involves the weakest constraints on the other tasks. Thus, the algorithm schedules all the tasks and gives implementation. This new algorithm is an adaptation of the force-directed scheduling algorithm with a cost function computation for hardware/software partitioning.
Frédéric Rousseau 0001, Judith Benzakki, J. M. Bergé, Michel Israël
RSP1