Emmanuel Casseau

dblp:75/3606 · DBLP profile ↗
← Back
29ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0001-7216-749XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Importance Resides in Activations: Fast Input-Based Nonlinearity Pruning
Baptiste Rossigneux, Vincent Lorrain, Inna Kucher, Emmanuel Casseau
ICONIP (9)4
2024 Algorithms with improved delay for enumerating connected induced subgraphs of a large cardinality
Chenglong Xiao, Emmanuel Casseau
Inf. Process. Lett.3
2023 Near-optimal energy-efficient partial-duplication task mapping of real-time parallel applications
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
J. Syst. Archit.4
2022 Cache management in MASCARA-FPGA: from coalescing heuristic to replacement policy
abstract
We presented ModulAr Semantic CAching fRAmework (MASCARA) that deployed Semantic Caching (SC) to perform a fast query processing based on Field Programmable Gate Arrays (FPGAs) accelerators. In addition of the accelerators, cache management plays an important role to address coalescing strategy and replacement policy so as to maximize the performance of FPGA caching. Therefore, in this paper, we present a coalescing heuristic with a new replacement function that leverages advantages of traditional strategies and overcomes their drawbacks. The proposed heuristic reduces response time, improves data availability, and saves cache space with respect to the semantic locality of query workload.
Laurent d'Orazio, Emmanuel Casseau, Julien Lallet
DaMoN3
2021 Fault-Tolerant Mapping of Real-Time Parallel Applications under multiple DVFS schemes
abstract
On multicore platforms, reliable task execution, as well as low energy consumption, are essential. Dynamic Voltage/Frequency Scaling (DVFS) is typically used for energy saving, but with a negative impact on reliability, especially when the frequency is low. Using high frequencies to meet reliability constraints is not always feasible, while multiple replicas increase energy consumption. To minimize energy consumption, enhancing reliability, without violating real-time constraints, we propose an approach that combines distinct reliability enhancement techniques, under task-level, processor-level and system-level DVFS. Our task mapping problem jointly optimizes task allocation, task frequency assignment, and task duplication, under multiple constraints. This is achieved by formulating the task mapping problem as a mixed integer non-linear programming problem and equivalently transforming it into a mixed integer linear programming, that is optimally solved. From the obtained results, the proposed approach achieves better energy consumption and finds solutions, when other approaches fail.
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
RTAS4
2021 MASCARA-FPGA cooperation model: Query Trimming through accelerators
abstract
The use of Field Programmable Gate Arrays (FPGA) has become attractive in recent years to accelerate database analysis. Meanwhile, Semantic Caching (SC) is a technique for optimizing the evaluation of database queries by exploiting the knowledge and resources contained in the queries themselves. Organizing SC on FPGA is relevant in terms of response time and quality of results to increase system performance. To make SC scalable on FPGAs, we have proposed a ModulAr Semantic CAching fRAmework (MASCARA) in which relevant stages or modules could be convertible as accelerators on FPGAs. Therefore, in this paper, we aim to present a complementary query processing platform based on the cooperation model between MASCARA and FPGA. This novel approach extends the advantage of the classical SC, which is mainly based on Central Processing Unit (CPU), by offloading computationally intensive phases to FPGA. Moreover, MASCARA-FPGA presents the workflow of query rewriting and partial query execution in a pipelined execution model where multiple accelerators can run in parallel. In our experiments, the Query Trimming can reduce the response time by up to 3.96 times with only one accelerator used.
Laurent d'Orazio, Emmanuel Casseau, Julien Lallet
SSDBM3
2021 Special Session: Operating Systems under test: an overview of the significance of the operating system in the resiliency of the computing continuum
abstract
The computing continuum's actual trend is facing a growth in terms of devices with any degree of computational capability. Those devices may or may not include a full-stack, including the Operating System layer and the Application layer, or just facing pure bare-metal solutions. In either case, the reliability of the full system stack has to be guaranteed. It is crucial to provide data regarding the impact of faults at all system stack levels and potential hardening solutions to design highly resilient systems. While most of the work usually concentrates on the application reliability, the special session aims to provide a deep comprehension of the impact on the reliability of an embedded system when faults in the hardware substrate of the system stack surface at the Operating System layer. For this reason, we will cover a comparison from an application perspective when hardware faults happen in bare metal vs. real-time OS vs. general-purpose OS. Then we will go deeper within a FreeRTOS to evaluate the contribution of all parts of the OS. Eventually, the Special Session will propose some hardening techniques at the Operating System level by exploiting the scheduling capabilities.
Emmanuel Casseau, Petr Dobiás, Oliver Sinnen, Gennaro Severino Rodrigues, Fernanda Lima Kastensmidt, Alessandro Savino 0001, Stefano Di Carlo, Maurizio Rebaudengo, Alberto Bosio
VTS1
2020 Evaluation of Fault Tolerant Online Scheduling Algorithms for CubeSats
abstract
Small satellites, such as CubeSats, have to respect time, spatial and energy constraints in the harsh space environment. To tackle this issue, this paper presents and evaluates two fault tolerant online scheduling algorithms: the algorithm scheduling all tasks as aperiodic (called ONEOFF) and the algorithm placing arriving tasks as aperiodic or periodic tasks (called ONEOFF & CYCLIC). Based on several scenarios, the results show that the performances of ordering policies are influenced by the system load and the proportions of simple and double tasks to all tasks to be executed. The “Earliest Deadline” and “Earliest Arrival Time” ordering policies for ONEOFF or the “Minimum Slack” ordering policy for ONEOFF & CYCLIC reject the least tasks in all tested scenarios. The paper also deals with the analysis of scheduling time to evaluate real-time performances of ordering policies and shows that ONEOFF requires less time to find a new schedule than ONEOFF & CYCLIC. Finally, it was found that the studied algorithms perform well also in a harsh environment.
Petr Dobiás, Emmanuel Casseau, Oliver Sinnen
DSD2
2018 Restricted Scheduling Windows for Dynamic Fault-Tolerant Primary/Backup Approach-Based Scheduling on Embedded Systems
abstract
This paper is aimed at studying fault-tolerant design of the realtime multi-processor systems and is in particular concerned with the dynamic mapping and scheduling of tasks on embedded systems. The effort is concentrated on scheduling strategy having reduced complexity and guaranteeing that, when a task is input into the system and accepted, then it is correctly executed prior to the task deadline. The chosen method makes use of the primary/backup approach and this paper describes its refinement based on reduction of windows within which the primary and the backup copies can be scheduled. The results show that the use of restricted scheduling windows reduces the algorithm complexity by up to 15%.
Petr Dobiás, Emmanuel Casseau, Oliver Sinnen
SCOPES2
2017 Parallel custom instruction identification for extensible processors
Chenglong Xiao, Wanjun Liu, Emmanuel Casseau
J. Syst. Archit.4
2014 Efficient software synthesis of dynamic dataflow programs
abstract
This paper introduces advanced software synthesis techniques that enhance the implementation of dynamic dataflow programs. These techniques have been implemented into open-source tools and demonstrated on well-known video decoders including one based on the new High Efficiency Video Coding (HEVC) standard. The results show an improvement of more than 100% of the frame-rate over previously proposed implementations, and achieve real-time decoding of high definition video sequences.
Hervé Yviquel, Alexandre Sanchez, Pekka Jääskeläinen, Jarmo Takala, Mickaël Raulet, Emmanuel Casseau
ICASSP6
2014 Improving high-level synthesis effectiveness through custom operator identification
abstract
It is increasingly common to see custom operators appear in various fields of circuit design. Custom operators that can be implemented in special hardware units make it possible to improve performance and reduce area. In this paper, we propose a design flow for identifying custom operators for high-level synthesis. Experimental results show that our approach achieves on average 19%, and up to 37% area reduction, compared to a traditional high-level synthesis. Meanwhile, the latency is reduced on average by 22%, and up to 59%. In addition, on average 74% and up to 81% code size reduction can be achieved, so synthesis runtime can be reduced.
Chenglong Xiao, Emmanuel Casseau
ISCAS2
2013 Automated design of networks of transport-triggered architecture processors using dynamic dataflow programs
Hervé Yviquel, Jani Boutellier, Mickaël Raulet, Emmanuel Casseau
Signal Process. Image Commun.4
2012 Design of multi-mode application-specific cores based on high-level synthesis
Emmanuel Casseau, Bertrand Le Gal
Integr.1
2012 Exact custom instruction enumeration for extensible processors
Chenglong Xiao, Emmanuel Casseau
Integr.2
2011 Efficient custom instruction enumeration for extensible processors
abstract
In recent years, the use of extensible processors has been increased. Extensible processors extend the base instruction set of a general-purpose processor with a set of custom instructions. Custom instructions that can be implemented in special hardware units make it possible to improve performance and decrease power consumption in extensible processors. The key issue involved is to generate and select automatically custom instructions from a high-level application code. In this paper, we propose an efficient and flexible algorithm for the exact enumeration of custom instructions. The algorithm can be tuned to generate all possible patterns or only connected patterns. Compared to a previously proposed well-known algorithm, our algorithm can achieve orders of magnitude speedup.
Chenglong Xiao, Emmanuel Casseau
ASAP2
2011 An efficient algorithm for custom instruction enumeration
abstract
In order to meet growing market demands in flexibility and performance, the use of extensible processors has been increased. Extensible processors extend the base instruction set of a general-purpose processor with a set of custom instructions. Custom instruction that can be implemented in special hardware unit is a vital component for improving performance in extensible processors. The key issue involved is to generate and select automatically custom instructions from high-level application code. In this paper, we propose a new efficient algorithm for automatic generation of all candidate instructions (or patterns). Our pattern generation algorithm identifies all feasible connected and disjoint patterns under different constraints. Compared to a previously proposed well-known algorithm, our algorithm solves the problem more efficiently by taking advantage of topological property of data flow graph (DFG) as well as overcoming the drawbacks of the previously proposed algorithm.
Chenglong Xiao, Emmanuel Casseau
ACM Great Lakes Symposium on VLSI2
2011 Exploiting reconfigurable SWP operators for multimedia applications
abstract
Implementing image processing applications in embedded systems is a difficult challenge due to the drastic constraints in terms of cost, energy consumption and real time execution. Reconfigurable architectures are good candidates to take-up this challenge and especially when the architecture is able to support different word-lengths of pixel through Sub-Word Parallelism (SWP) capabilities. Exploiting the diversity of supported data-types requires automation tools able to optimize the data word-length under an accuracy constraint. In this paper, a new approach for word-length optimization in the case of SWP operations is proposed. Compared to existing approaches the optimization time is significantly reduced without sacrificing the quality of the optimized solution. The results show the ability of our approach to exploit the SWP capabilities associated with multimedia processors.
Daniel Ménard, Hai-Nam Nguyen, François Charot, Stéphane Guyetant, Jérémie Guillot, Erwan Raffin, Emmanuel Casseau
ICASSP7
2010 High-Level Synthesis for Designing Multimode Architectures
abstract
This paper addresses the design of multimode architectures for digital signal and image processing applications. We present a dedicated design flow and its associated high-level synthesis tool, named GAUT. Given a unified description of a set of time-wise mutually exclusive tasks and their associated throughput constraints, a single register transfer level hardware architecture optimized in area is generated. In order to reduce the register, the steering logic, and the controller complexities, this paper proposes a joint-scheduling algorithm, which maximizes the similarities between the control steps and specific binding approaches for both operators and storage elements which maximize the similarities between the datapaths. It is shown through a set of test cases that the proposed approach offers significant area saving and low-performance penalties compared to both state-of-the-art techniques and dedicated mono-mode architectures.
Caaliph Andriamisaina, Philippe Coussy, Emmanuel Casseau, Cyrille Chavet
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Reconfigurable SWP Operator for Multimedia Processing
abstract
For performance enhancement, reconfigurable processors have to overcome the overheads of reconfigurations such as the complexity of the interconnection network and reconfiguration time. In processors dealing with multimedia applications these overheads can be reduced by providing the reconfigurability inside the processing units rather than at interconnection level. Due to the low precision data nature of multimedia applications, reconfiguration at operator level also provides additional speedup through parallel execution of low precision data. In this paper a pipelined architecture of a reconfigurable coarse grain subword parallel (SWP) operator is presented for multimedia applications. This operator not only eliminates the need of reconfiguration time but also provides the reconfigurability at both data size level (different pixel data sizes) and at operation level (different multimedia oriented operations). This ensures a better utilization of the processor resources and reduces the reconfiguration overheads significantly.
Shafqat Khan, Emmanuel Casseau, Daniel Ménard
ASAP2
2008 Dynamic Memory Access Management for High-Performance DSP Applications Using High-Level Synthesis
abstract
Multimedia applications such as video and image processing are often characterized by a huge number of data accesses. In many digital signal processing applications, array access patterns are regular and periodic. In these cases, optimized architectures using pipelined memory access controllers can be generated. In this paper, we focus on implementing memory interfacing modules that can be automatically generated from a high-level synthesis tool and which can efficiently handle predictable address patterns as well as random ones (i.e., dynamic address computations). The benefits of balancing dynamic address computations from datapath to dedicated computation units in the memory controller is also analyzed as well as operator bitwidth optimization and data locality to save power consumption and reduce latency.
Bertrand Le Gal, Emmanuel Casseau, Sylvain Huet
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Granularity Issues in Transaction Level Modelling Digital Signal Processing Applications
Sylvain Huet, Sébastien Le Nours, Olivier Pasquier, Emmanuel Casseau
FDL4
2007 A design flow dedicated to multi-mode architectures for DSP applications
abstract
This paper addresses the design of multi-mode architectures for digital signal processing applications. We present a dedicated design flow and its associated high-level synthesis tool, named GAUT. Given a unified description of a set of time-wise mutually exclusive tasks and their associated throughput constraints, a single RTL hardware architecture optimized in area is generated. In order to reduce the register, steering logic (multiplexers) and controller (decoding logic) complexities, we propose a joint-scheduling algorithm which maximizes the similarities between control steps and specific binding approaches for both functional units and storage elements which maximize the similarities between the datapaths. We show through a set of test cases that our approach offers significant area saving relative to the state-of-the-art.
Cyrille Chavet, Caaliph Andriamisaina, Philippe Coussy, Emmanuel Casseau, Emmanuel Juin, Pascal Urard, Eric Martin 0001
ICCAD4
2007 Constrained algorithmic IP design for system-on-chip
Philippe Coussy, Emmanuel Casseau, Pierre Bomel, Adel Baganne, Eric Martin 0001
Integr.2
2006 A Computation Core for Communication Refinement of Digital Signal Processing Algorithms
abstract
The most popular Moore's law formulation, which states the number of transistors on integrated circuits doubles every 18 months, is said to hold for at least another two decades. According to this prediction, if we want to take advantage of technological evolutions, designer's productivity has to increase in the same proportions. To take up this challenge, system level design solutions have been set up, but many efforts have still to be done on system modelling and synthesis. In this paper we propose a computation core synthesis methodology that can be integrated on the communication refinement steps of electronic system level design tools. In the proposed approach, computation cores used for digital signal processing application specifications relying on coarse grain communications and synchronizations (e.g. matrix) can be refined into computation cores which can handle fine grain communications and synchronizations (e.g. scalar). Its originality is its ability to synthesize computation cores which can handle fine grain data consumptions and productions which respect the intrinsic partial orders of the algorithms while preserving their original functionalities. Such cores can be used to model fine grain input output overlapping or iteration pipelining. Our flow is based on the analysis of a fine grain signal flow graph used to extract fine grain synchronizations and algorithmic expressions
Sylvain Huet, Emmanuel Casseau, Olivier Pasquier
DSD2
2006 Hardware Communication Refinement in Digital Signal Processing
Sylvain Huet, Emmanuel Casseau, Olivier Pasquier, Sébastien Le Nours
FDL2
2006 Design of a flexible 2-D discrete wavelet transform IP core for JPEG2000 image coding in embedded imaging systems
Guillaume Savaton, Emmanuel Casseau, Eric Martin 0001
Signal Process.2
2006 A formal method for hardware IP design and integration under I/O and timing constraints
abstract
IP integration, which is one of the most important SoC design steps, requires taking into account communication and timing constraints. In that context, design and reuse can be improved using IP cores described at a high abstraction level. In this paper, we present an IP design approach that relies on three main phases: (1) constraint modeling, (2) IP constraint analysis steps for feasibility checking, and (3) synthesis. We propose a set of techniques dedicated to the digital signal processing domain that lead to an optimized IP core integration. Based on a generic architecture of components, the method we propose provides automatic generation of IP cores designed under integration constraints. We show the effectiveness of our approach with a DCT core design case study.
Philippe Coussy, Emmanuel Casseau, Pierre Bomel, Adel Baganne, Eric Martin 0001
ACM Trans. Embed. Comput. Syst.2
2005 Hardware Virtual Components Compliant with Communication System Standards
abstract
In this paper, we focus on the design of a communication system based on reusing IP cores. Traditional methods for designing hardware cores for this kind of applications use a RTL specification. However, they suffer from heavy limitations that prevent them from efficiently addressing the algorithmic complexity and the high flexibility required by the various application profiles. For this reason, we propose to raise the abstraction level of the specification and introduce the notion of architectural flexibility by benefiting from the emerging high-level synthesis tools. From a single behavioral-level VHDL specification, we are able to generate a variety of architectures, compliant with the most important communication standards. This technique has been successfully applied to the most important IP cores (synchronization IP, Viterbi IP and Reed-Solomon decoder IP cores) of the DVB-DSNG digital video-broadcasting standard.
Nabil Abdelli, Pierre Bomel, Emmanuel Casseau, Anne-Marie Fouilliart, Christophe Jégo, Philippe Kajfasz, Bertrand Le Gal, Nathalie Le Heno
DSD3