Jean-François Nezan

dblp:99/3613 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-0609-4592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Automated level-based clustering of dataflow actors for controlled scheduling complexity
Ophélie Renaud, Hugo Miomandre, Karol Desnos, Jean-François Nezan
J. Syst. Archit.4
2024 Automated Buffer Sizing of Dataflow Applications in a High-level Synthesis Workflow
abstract
High-Level Synthesis (HLS) tools are mature enough to provide efficient code generation for computation kernels on FPGA hardware. For more complex applications, multiple kernels may be connected by a dataflow graph. Although some tools, such as Xilinx Vitis HLS, support dataflow directives, they lack efficient analysis methods to compute the buffer sizes between kernels in a dataflow graph. This article proposes an original method to safely approximate such buffer sizes. The first contribution computes an initial overestimation of buffer sizes without knowing the memory access patterns of kernels. The second contribution iteratively refines those buffer sizes, thanks to cosimulation. Moreover, the article introduces an open source framework using these methods to facilitate dataflow programming on FPGA using HLS. The proposed methods and framework have been tested on seven dataflow applications and outperform Vitis HLS cosimulation in five benchmarks, either in terms of BRAM and LUT usage, or in terms of exploration time. In the two other benchmarks, our best method gets results similar to Vitis HLS. Last but not least, our method admits directed cycles in the application graphs.
Alexandre Honorat, Mickaël Dardaillon, Hugo Miomandre, Jean-François Nezan
ACM Trans. Reconfigurable Technol. Syst.4
2022 Design Space Exploration for Memory-Oriented Approximate Computing Techniques
abstract
Modern digital systems are processing more and more data. This increase in memory requirements must match the processing capabilities and interconnections to avoid the memory wall. Approximate computing techniques exist to alleviate these requirements but usually require a thorough and tedious analysis of the processing pipeline. This paper presents an application-agnostic Design Space Exploration (DSE) of the buffer-sizing process to reduce the memory footprint of applications while guaranteeing an output quality above a defined threshold. The proposed DSE selects the appropriate bit-width and storage type for buffers to satisfy the constraint. We show in this paper that the proposed DSE reduces the memory footprint of the SqueezeNet CNN by 58.6% with identical Top-1 prediction accuracy, and the full SKA SDP pipeline by 39.7% without degradation, while only testing for a subset of the design space. The proposed DSE is fast enough to be integrated into the design stream of applications.
Hugo Miomandre, Jean-François Nezan, Daniel Ménard
ASAP2
2020 Forward-Inverse 2D Hardware Implementation of Approximate Transform Core for the VVC Standard
abstract
The future video coding standard named Versatile Video Coding (VVC) is expected by the end of 2020. VVC will enable better coding efficiency than the current High Efficiency Video Coding (HEVC) standard. This coding gain is brought by several coding tools. The Multiple Transform Selection (MTS) is one of the key coding tools that have been introduced in VVC. The MTS concept relies on three transform types including Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII and DCT-VIII. Unlike the DCT-II that has fast computing algorithms, the DST-VII and DCT-VIII rely on more complex matrix multiplication. In this paper an approximation approach is proposed to reduce the computational cost of the DST-VII and DCT-VIII. The approximation consists in applying adjustment stages, based on sparse block-band matrices, to a variant of DCT-II family mainly DCT-II and its inverse. Genetic algorithm is used to derive the optimal coefficients of the adjustment matrices. Moreover, an efficient hardware implementation of the forward and inverse approximate transform module is proposed. The architecture design includes a pipelined and reconfigurable forward-inverse DCT-II core transform as it is the main core for DST-VII and DCT-VIII computations. The proposed 32-point 1D architecture including low cost adjustment stages allows the processing of a video in 2K and 4K resolutions at 1095 and 273 frames per second, respectively. A unified 2D implementation of forward-inverse DCT-II, approximate DST-VII and DCT-VIII is also presented. The synthesis results show that the design is able to sustain a video in 2K and 4K resolutions at 386 and 96 frames per second, respectively, while using only 12% of Alms, 22% of registers and 30% of DSP blocks of the Arria10 SoC platform.
Ahmed Kammoun, Wassim Hamidouche, Pierrick Philippe, Olivier Déforges, Fatma Belghith, Nouri Masmoudi, Jean-François Nezan
IEEE Trans. Circuits Syst. Video Technol.7
2018 Reproducible Evaluation of System Efficiency With a Model of Architecture: From Theory to Practice
abstract
Current trends in high performance and embedded computing include design of increasingly complex hardware architectures with high parallelism, heterogeneous processing elements, and nonuniform communication resources. In order to take hardware and software design decisions, early evaluations of the system nonfunctional properties are needed. These evaluations of system efficiency require electronic system-level information on both algorithms and architecture. Contrary to algorithm models for which a major body of work has been conducted on defining formal models of computation (MoCs), architecture models from the literature are mostly empirical models from which reproducible experimentation requires the accompanying software. In this paper, a precise definition of a model of architecture (MoA) is proposed that focuses on reproducibility and abstraction and removes the overlap previously existing between the notions of MoA and MoC. A first MoA, called the linear system-level architecture model (LSLA), is presented. To demonstrate the generic nature of the proposed new architecture modeling concepts, we show that the LSLA model can be integrated flexibly with different MoCs. LSLA is then used to model the energy consumption of a state-of-the-art multiprocessor system-on-chip (MPSoC) when running an application described using the synchronous dataflow MoC. A method to automatically learn LSLA model parameters from platform measurements is introduced. Despite the high complexity of the underlying hardware and software, a simple LSLA model is demonstrated to estimate the energy consumption of the MPSoC with a fidelity of 86%.
Maxime Pelcat, Alexandre Mercat, Karol Desnos, Luca Maggiani, Yanzhou Liu 0001, Julien Heulot, Jean-François Nezan, Wassim Hamidouche, Daniel Ménard, Shuvra S. Bhattacharyya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2017 Hierarchical Dataflow Model for efficient programming of clustered manycore processors
abstract
Programming Multiprocessor Systems-on-Chips (MPSoCs) with hundreds of heterogeneous Processing Elements (PEs), complex memory architectures, and Networks-on-Chips (NoCs) remains a challenge for embedded system designers. Dataflow Models of Computation (MoCs) are increasingly used for developing parallel applications as their high-level of abstraction eases the automation of mapping, task scheduling and memory allocation onto MPSoCs. This paper introduces a technique for deploying hierarchical dataflow graphs efficiently onto MPSoC. The proposed technique exploits different granularity of dataflow parallelism to generate both NoC-based communications and nested OpenMP loops. Deployment of an image processing application on a many-core MPSoC results in speedups of up to 58.7 compared to the sequential execution.
Julien Hascoet, Karol Desnos, Jean-François Nezan, Benoît Dupont de Dinechin
ASAP3
2017 Throughput evaluation of DSP applications based on hierarchical dataflow models
abstract
Synchronous Dataflow (SDF) is the most commonly used dataflow Model of Computation (MoC) for the specification of Digital Signal Processing (DSP) systems. The Interface-Based SDF (IBSDF) model extends the semantics of the SDF model by introducing a graph composition mechanism based on hierarchical interfaces. Computing the throughput of an application is essential when designing DSP systems. This article introduces and assesses new methods to compute the throughput of DSP applications specified with IBSDF graphs. First, a basic method inspired from the state-of-the-art techniques that relies on a transformation of the IBSDF graph to an equivalent non-hierarchical graph of potentially exponential size. Second, a new technique that takes advantage of the hierarchy semantics of the IBSDF MoC to speed-up the throughput evaluation without any conversion. The proposed technique makes it possible to compute the throughput of large IBSDF graphs in a few milliseconds, where the basic method fails to produce a result.
Hamza Deroui, Karol Desnos, Jean-François Nezan, Alix Munier Kordon
ISCAS3
2017 An automatized method to parameterize embedded stereo matching algorithms
Judicael Menant, Guillaume Gautier, Muriel Pressigout, Luce Morin, Jean-François Nezan
J. Syst. Archit.5
2016 Optimized Belief Propagation Algorithm onto Embedded Multi and Many-Core Systems for Stereo Matching
abstract
Stereo matching techniques aim at reconstructing disparity maps from a pair of images. The use of stereo matching techniques in embedded systems is very challenging due to the complexity of the state-of-the-art algorithms. Local stereo matching algorithms are efficiently implemented on GPU and DSP. This paper presents the optimization of the One Dimension Belief Propagation (BP-1D) algorithm. BP-1D is faster than previous algorithms on monocore DSP and its implementation onto multicore DSPs is straightforward. BP-1D implemented on multicore embedded platforms out-performs previous stereo matching implementations reaching real-time performances for resolutions up to 1080p with a 10 Watts power consumption.
Jean-François Nezan, Alexandre Mercat, Patrice Delmas, Georgy L. Gimel'farb
PDP1
2016 On Memory Reuse Between Inputs and Outputs of Dataflow Actors
abstract
This article introduces a new technique to minimize the memory footprints of Digital Signal Processing (DSP) applications specified with Synchronous Dataflow (SDF) graphs and implemented on shared-memory Multiprocessor System-on-Chip (MPSoCs). In addition to the SDF specification, which captures data dependencies between coarse-grained tasks called actors, the proposed technique relies on two optional inputs abstracting the internal data dependencies of actors: annotations of the ports of actors, and script-based specifications of merging opportunities between input and output buffers of actors. Experimental results on a set of applications show a reduction of the memory footprint by 48% compared to state-of-the-art minimization techniques.
Karol Desnos, Maxime Pelcat, Jean-François Nezan, Slaheddine Aridhi
ACM Trans. Embed. Comput. Syst.3
2015 Buffer merging technique for minimizing memory footprints of Synchronous Dataflow specifications
abstract
This paper introduces and assesses a new technique to minimize the memory footprints of Digital Signal Processing (DSP) applications specified with Synchronous Dataflow (SDF) graphs and implemented on shared-memory Multiprocessor Systems-on-Chips (MPSoCs). In addition to the SDF specification, which captures data dependencies between coarse-grained tasks called actors, the proposed technique relies on two optional inputs abstracting the internal data dependencies of actors: annotations of the ports of SDF actors, and script-based specifications of merging opportunities between input and output buffers of actors. An automated optimization process is used to exploit these buffer merging opportunities and to minimize the memory footprints of applications. Experimental results on a computer vision application show a reduction of the memory footprint by 34% compared to state-of-the-art minimization techniques.
Karol Desnos, Maxime Pelcat, Jean-François Nezan, Slaheddine Aridhi
ICASSP3
2014 Implementation of a Stereo Matching algorithm onto a Manycore Embedded System
abstract
Stereo Matching techniques aim at reconstructing the disparity maps with a pair of images. The use of Stereo Matching techniques in embedded systems is very challenging due to the complexity of the state of the art algorithms. This paper proposes a real-time Stereo Matching algorithm optimised for the last generation of Manycore Embedded Systems. The features and parameters of the algorithms have been chosen to optimise the trade-off between an high quality and a low complexity. A memory analysis is performed for the algorithm's decomposition and the resulting mapping on the Manycore platform is provided. Algorithm and arithmetic optimisations have been applied to each part of the algorithm to decrease the execution time down to 160ms for a CIF resolution.
Alexandre Mercat, Jean-François Nezan, Daniel Ménard
ISCAS2
2012 Multi-purpose systems: A novel dataflow-based generation and mapping strategy
abstract
The manual creation of specialized hard-ware infrastructures for complex multi-purpose systems is error-prone and time-consuming. Moreover, lots of effort is required to define an optimized and heterogeneous components library. To tackle these issues, we propose a novel design flow based on the Dataflow Process Networks Model of Computation. In particular, we have combined the operation of two state of the art tools, the Multi-Dataflow Composer and the Open RVC-CAL Compiler, handling respectively the automatic mapping of a reconfigurable multi-purpose substrate and the high level synthesis of hardware components. Our approach guarantees runtime efficiency and on-chip area saving both on FPGAs and ASICs.
Jean-François Nezan, Nicolas Siret, Matthieu Wipliez, Francesca Palumbo, Luigi Raffo
ISCAS1
2010 A codesign synthesis from an MPEG-4 decoder dataflow description
abstract
The elaboration of new and innovative systems such as MPSoC (Multiprocessor System on Chip) which are made up of multiple processors, memories and IPs lies on the designers to achieve a complex codesign work. Specific tools and methods are needed to cope with the increasing complexity of both algorithms and platforms. Our approach to design such systems is based on the usage of a high level of abstraction language called RVC CAL. This language is dataflow oriented and thus points out the concurrency and parallelism of algorithms. Moreover CAL is supported by the OpenDF simulator and by two code generators called CAL2C (software generator) and CAL2HDL (hardware generator). The MPEG expert group has recently elaborated the Reconfigurable Video Coding (RVC) standard which defines the RVC CAL language as reference for MPEG video decoder descriptions. This paper introduces the opportunities to design an innovative system involving hardware and software IPs, embedded processors and memories from a CAL model. Practical results on a FPGA are provided with a codesign solution of an MPEG4 Simple Profile (SP).
Nicolas Siret, Ismaïl Sabry, Jean-François Nezan, Mickaël Raulet
ISCAS3
2010 Advanced list scheduling heuristic for task scheduling with communication contention for parallel embedded systems
Pengcheng Mu, Jean-François Nezan, Mickaël Raulet, Jean-Gabriel Cousin
Sci. China Inf. Sci.2
2009 Scalable compile-time scheduler for multi-core architectures
abstract
As the number of cores continues to grow in both digital signal and general purpose processors, tools which perform automatic scheduling from model-based designs are of increasing interest. This scheduling consists of statically distributing the tasks that constitute an application between available cores in a multi-core architecture in order to minimize the final latency. This problem has been proven to be NP-complete. A static scheduling algorithm is usually described as a monolithic process, and carries out two distinct functionalities: choosing the core to execute a specific function and evaluating the cost of the generated solutions. This paper describes a scheduling module which splits these functionalities into two sub-modules. This division produces an advanced scalability in terms of schedule quality and computation time, and also separates the heuristic complexity from the architecture model precision.
Maxime Pelcat, Pierrick Menuet, Slaheddine Aridhi, Jean-François Nezan
DATE4
2008 Software synthesis of CAL actors for the MPEG reconfigurable Video Coding framework
abstract
The MPEG reconfigurable video coding (RVC) framework aims to provide a unified specification of all video technology. In this framework, a decoder is modularly built as a configuration of video coding tools taken from the MPEG toolbox library. The elements of the library are specified using the CAL actor language. CAL is a dataflow based language providing computation models that are concurrent and modular. This paper presents a synthesis tool that from a CAL specification generates C code. Indeed, code generators are fundamental supports for the deployment and success of the MPEG RVC framework. This paper focuses on the automatic translation of a CAL actor. This approach has been used to obtain a C implementation of the inverse DCT module which is part of the MPEG-4 Simple Profile decoder, chosen by MPEG experts to validate the RVC approach. The generated code is validated against the original CAL description and simulated using the Open Dataflow environment.
Ghislain Roquier, Matthieu Wipliez, Mickaël Raulet, Jean-François Nezan, Olivier Déforges
ICIP4
2008 Code generation for the MPEG Reconfigurable Video Coding framework: From CAL actions to C functions
abstract
The MPEG reconfigurable video coding (RVC) framework is a new standard under development by MPEG that aims at providing a unified specification of current MPEG video coding technologies. In this framework, a decoder is built as a configuration of video coding modules taken from the standard ldquoMPEG toolbox libraryrdquo. The elements of the library are specified using the CAL actor language (CAL). CAL is a dataflow based language providing computation models that are concurrent and modular. This paper describes a synthesis tool that from a CAL specification automatically generates compilable C-code. Code generators are fundamental supports for the deployment and success of the MPEG RVC framework. This paper focuses on the automatic translation of CAL actions, which is the first step to a complete actor translation. The techniques described here enable to automatically generate C-code according to a finite set of rules. This approach has been used to obtain a C implementation of the IDCT module which is one element of the RVC library. The generated code is validated against the original CAL dataflow program simulated using the open dataflow environment.
Matthieu Wipliez, Ghislain Roquier, Mickaël Raulet, Jean-François Nezan, Olivier Déforges
ICME4
2008 A Flexible Heterogeneous Hardware/Software Solution for Real-Time HD H.264 Motion Estimation
abstract
Quarter-pixel accuracy and variable block-size significantly enhance compression performances of the MPEG-4 AVC/H.264 video compression standard over its predecessors, but also significantly increase computation requirements. Firstly, a digital signal processor (DSP)-based solution that achieves real-time integer motion estimation is proposed. Fractional-pixel refinement is too computationally intensive to be efficiently processed on a software-based processor. To address this restriction, a flexible and low complexity VLSI subpixel refinement coprocessor is designed. Thanks to an improved datapath, a high throughput is achieved with low logic resources. Finally, an heterogeneous (DSP-field-programmable gate array) solution to handle real-time motion estimation with variable block-size and fractional-pixel accuracy for high-definition video is studied. This solution, combining programmability and efficiency, achieves motion estimation of 720 p sequences at up to 60 fps.
Fabrice Urban, Ronan Poullaouec, Jean-François Nezan, Olivier Déforges
IEEE Trans. Circuits Syst. Video Technol.3
2003 Rapid prototyping for an optimized MPEG-4 decoder implementation over a parallel heterogenous architecture
abstract
Sequential MPEG-4 solutions actually developed for single processors try to integrate the most functionalities as possible in an unique software, and are generally oversized compared with the actual service requirement. Moreover, they can hardly be projected onto multiprocessors targets, leading to an extra load of source code and calculations, but also to a sub-optimal use of the architecture parallelism. This paper introduces a distributed MPEG-4 application, where the system part is hosted by a standard PC, and the video decoder is supported by a multi-DSPs board. In particular, we present our AVSynDEx methodology allowing both an incremental building, an easy update on the video decoder description, and a quasi-automatic implementation onto a multi-C6x platform. We also define a global scheduler managing the parallel execution of the video and system applications.
Nicolas Ventroux, Jean-François Nezan, Mickaël Raulet, Olivier Déforges
ICASSP (2)2
2003 Rapid prototyping for an optimized MPEG4 decoder implementation over a parallel heterogeneous architecture
abstract
Sequential Mpeg-4 solutions actually developed for single processors try to integrate the most functionalities as possible in an unique software, and are generally oversized compared with the actual service requirement. Moreover, they can hardly be projected onto multiprocessors targets, leading to an extra load of source code and calculations, but also to a sub-optimal use of the architecture parallelism. This paper introduces a distributed Mpeg-4 application, where the system part is hosted by a standard PC, and the video decoder is supported by a multi-DSPs board. In particular, we present our AVSynDEx methodology allowing both an incremental building, an easy update on the video decoder description, and a quasi-automatic implementation onto a multi-C6x platform. We also define a global scheduler managing the parallel execution of the video and system applications.
Nicolas Ventroux, Jean-François Nezan, Mickaël Raulet, Olivier Déforges
ICME2
2002 Rapid prototyping methodology for multi-DSP TI C6X platforms applied to an Mpeg-2 coding application
abstract
Real time signal and image applications have very important time constraints, involving the use of several powerful numerical calculation units. Our aim is to develop a fast prototyping process dedicated to parallel architectures made of several last generation Texas Instruments TMS320C6X DSP. The methodology is based on the use of SynDEx, a CAD software improving the algorithm implementation onto multiprocessor architectures, finding the best matching between an algorithm and an architecture. A SynDEx executive kernel has been developed for the C6X DSP family in order to automatically generate a distributed and optimized static executive of the specified algorithm onto those processors. We have tested the efficiency of our methodology with a complete Mpeg-2 coding application.
Jean-François Nezan, Olivier Déforges, Mickaël Raulet
SPAA1