Christian Fabre

dblp:89/4740 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-9117-2056ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 5 · 2 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Page size exploration for RISC-V systems: the case for HPC
abstract
The page size used for virtual to physical address translation has globally not changed since the late 1960’s: the IBM 360, circa 1964, already had 4 KiB pages. This 4 KiB page size has proven to be incredibly robust given the changes in processor architectures, workloads behavior, memory size, and access patterns. However, with 64-bit registers, 57-bit virtual addresses, and increasingly bigger physical memories, we have to ask ourselves whether 4 KiB is still an adequate page size for modern workloads on modern machines. Inherently, the page size has an influence on (a) the miss rate of the translation lookaside buffer, the cache that contains the recently used virtual to physical translations, and (b) the memory allocated by the system versus the memory actually used by a process. The page size also constraints some microarchitectural choices, such as cache design, which impacts the overall performance and energy efficiency. We focus more particularly on High Performance Computing (HPC) applications because they are extremely demanding in terms of memory, and are indicative of future general-purpose needs.In this paper, we empirically study the evolution of the miss rate and memory occupancy with respect to the page size, and conclude that a page size of 32 KiB is better suited for current HPC systems. We also propose a page table scheme for RISC-V-based HPC systems based on our observations and discuss its benefits.
Eduardo Tomasi, César Fuguet Tortolero, Christian Fabre, Frédéric Pétrot
RSP3
2021 Seamless Compiler Integration of Variable Precision Floating-Point Arithmetic
abstract
Floating-Point (FP) units in processors are generally limited to supporting a subset of formats defined by the IEEE 754 standard. As a result, high-efficiency languages and optimizing compilers for high-performance computing only support IEEE standard types and applications needing higher precision involve cumbersome memory management and calls to external libraries, resulting in code bloat and making the intent of the program unclear. We present an extension of the C type system that can represent generic FP operations and formats, supporting both static precision and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and low-level FP instructions or software libraries. The effectiveness of our solution is demonstrated through an LLVM-based implementation, leveraging aggressive optimizations in LLVM including the Polly loop nest optimizer, which targets two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision floating-point library. Our optimizing compilation flow targeting MPFR outperforms the Boost programming interface for the MPFR library by a factor of 1.80 × and 1.67 × in sequential execution of the Poly Bench and RAJAPerf suites, respectively, and by a factor of 7.62 x on an 8-core (and 16-thread) machine for RAJAPerf in OpenMP.
Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot
CGO3
2020 VP Float: First Class Treatment for Variable Precision Floating Point Arithmetic
abstract
Optimizing compilers for high performance computing only support IEEE~754 floating-point (FP) types and applications needing higher precision involve cumbersome memory management and calls to external libraries. We introduce an extension of the C type system to represent variable-precision FP arithmetic, supporting both static and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and hardware FP instructions or software libraries. We demonstrate the effectiveness of our solution by enabling the full range of LLVM optimizations and leveraging two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision FP library. Both targets support the static and dynamically adaptable precision of our type system. On the PolyBench suite, our optimizing compilation flow targeting MPFR is shown to outperform the Boost programming interface for the MPFR library.
Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot
PACT3
2019 Byte-Aware Floating-point Operations through a UNUM Computing Unit
abstract
Most floating-point (FP) hardware support the IEEE 754 format, which defines fixed-size data types from 16 to 128 bits. However, a range of applications benefit from different formats, implementing different tradeoffs. This paper proposes a Variable Precision (VP) computing unit offering a finer granularity of high precision FP operations. The chosen memory format is derived from UNUM type I, where the size of a number is stored within the representation itself. The unit implements a fully pipelined architecture, and it supports up to 512 bits of precision for both interval and scalar computing. The user can conFigure the storage format up to 8-bit granularity, and the internal computing precision at 64-bit granularity. The system is integrated as a RISC-V coprocessor. Dedicated compiler support exposes the unit through a high level programming abstraction, covering all the operating features of UNUM type I. FPGA-based measurements show that the latency and the computation accuracy of this system scale linearly with the memory format length set by the user. Compared with a highly optimized software implementation, the proposed unit achieves speedups between 3.5 × and 18 ×, with comparable accuracy.
Andrea Bocco, Tiago T. Jost, Albert Cohen 0001, Florent de Dinechin, Yves Durand, Christian Fabre
VLSI-SoC6
2017 From a Formalized Parallel Action Language to Its Efficient Code Generation
abstract
Modeling languages propose convenient abstractions and transformations to handle the complexity of today’s embedded systems. Based on the formalism of the Hierarchical State Machine, they enable the expression of hierarchical control parallelism. However, they face two important challenges when it comes to modeling data-intensive applications: no unified approach that also accounts for data-parallel actions and no effective code optimization and generation flows. We propose a modeling language extended with parallel action semantics and hierarchical indexed-state machines suitable for computationally intensive applications. Together with its formal semantics, we present an optimizing model compiler aiming for the generation of efficient data-parallel implementations.
Ivan Llopard, Christian Fabre, Albert Cohen 0001
ACM Trans. Embed. Comput. Syst.2
2016 Transforming VHDL descriptions into formal component-based models
abstract
In this work, we investigate a transformation of VHDL descriptions into equivalent formal models. The targeted equivalence is at the level of the functional behavior. That is, we aim at producing formal models that have the same functional simulation behavior as the original VHDL implementation. We rely on the BIP component-based modeling language as the underlying formalism for this transformation. The expected benefits of such a transformation are: enabling the formal verification of hardware designs, allowing for software/hardware system modeling within the same formal framework, and, potentially, accelerating VHDL designs functional simulation by producing distributed BIP models. We show, through a case study, that the transformation is feasible and worth to develop.
Ayoub Nouri, Rahma Ben Atitallah, Anca Mariana Molnos, Christian Fabre, Frédéric Heitzmann, Olivier Debicki
RSP4
2014 A model-driven approach for embedded system prototyping and design
abstract
Embedded System (ES) development complexity is increasing. This increase has several cumulative sources: some are directly related to constraints on the ES themselves (dependability, compute intensive, resource constraints) while other sources are related to the industrial context of their development (fast prototyping, early validation, parallelization of developments). Although several Model-Driven Engineering (MDE) processes have been proposed for ES development, most of them are not completely formalized. This has several drawbacks that prevent their use in prototyping where iterations need to be short and focused. Incomplete formalized processes tend to be sidestepped in these situations where quick results are expected to be obtained with limited effort. In this paper we propose a MDE-based process for ES development. This process precisely defines the development tasks and their impact on the models throughout development. In particular we define iterations width and depth for the process that allow for a fined-grained and consistent planning of developments. The short and well defined iterations characterized by the process reduce the gap between rapid prototyping, ad-hoc methods and regular development processes.
Nicolas Hili, Christian Fabre, Sophie Dupuy-Chessa, Dominique Rieu
RSP2
2014 A parallel action language for embedded applications and its compilation flow
abstract
The complexity of Embedded System (ES) development is increasing dramatically. This has several cumulative sources: the intricate combination of data-intensive, computational and control aspects; the ubiquity of parallelism and heterogeneity of modern architectures; and the diversity of target-specific, non-deterministic programming models (e.g., C++ with explicit message passing, OpenCL, VHDL). Model-Driven Engineering (MDE) proposes to manage complexity by raising the level of abstraction for designers and developers, and refining the implementation for a particular context and platform through model transformations. In such frameworks, behavior is often specified by means of Hierarchical State Machines (HSMs) equiped with an action language. However, although such models represent some level of control parallelism through objects and HSMs, data parallelism, compound data, and the exploitation and optimization thereof remains very limited.
Ivan Llopard, Albert Cohen 0001, Christian Fabre, Nicolas Hili
SCOPES3
2012 Integrated architecture exploration workflow: A NoC-based case study
abstract
Compute-intensive applications can greatly benefit from the flexibility of NoC-based heterogeneous multi-core platforms. However, mapping applications on such MPSoC is becoming increasingly complex and requires integrated design flows. We conducted a case study to evaluate the benefits of an integrated design flow for the mapping space exploration of a real telecommunication application on a NoC-based heterogeneous platform. Thanks to the flow, we simulated several virtual platforms and several mappings of our application on each. This approach drastically lowers the required skills and the time needed for design space exploration. An improvement of several weeks have been observed.
Diego Puschini, Julien Mottin, Nicolas Palix, Lian Apostol, Christian Fabre
RSP5
2002 Speedup Prediction for Selective Compilation of Embedded Java Programs
Vincent Colin de Verdière, Sébastien Cros, Christian Fabre, Romain Guider, Sergio Yovine
EMSOFT3