Tanguy Risset

dblp:73/1025 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
1since 2021 · last 2023
0000-0001-9758-4900ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 33% Distributed systems · 25% Memory systems · 25%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 82% Operating systems · 18%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › fault tolerance
checkpointing
0.412019
Sytare: A Lightweight Kernel for NVRAM-Based Transiently-Powered Systems · IEEE Trans. Computers 2019
Memory systems
non-volatile memory
0.412019
Sytare: A Lightweight Kernel for NVRAM-Based Transiently-Powered Systems · IEEE Trans. Computers 2019
Compilers and program optimization › compiler construction
compilation pipeline
0.212016
A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs · ACM Trans. Archit. Code Optim. 2016
Compilers and program optimization › parallel language compilation
dataflow compilation
0.212016
A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs · ACM Trans. Archit. Code Optim. 2016
Parallel and multicore computing › concurrent programming › concurrency model
actor-based programming
0.212016
A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs · ACM Trans. Archit. Code Optim. 2016
Embedded and real-time systems › embedded hardware platform › MPSoC
heterogeneous MPSoC
0.212016
A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs · ACM Trans. Archit. Code Optim. 2016
Parallel and multicore computing
parallel programming models
0.212016
A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs · ACM Trans. Archit. Code Optim. 2016
Operating systems › kernel
lightweight kernel
0.112019
Sytare: A Lightweight Kernel for NVRAM-Based Transiently-Powered Systems · IEEE Trans. Computers 2019

Methods — techniques the papers use, named apart from their topics

kernel-oriented checkpointing · 0.8parametric dataflow · 0.5microscheduling · 0.5FIFO sizing · 0.5
YearPublicationVenuePosition
2023 Audio DSP to FPGA Compilation
abstract
The implementation of real-time audio Digital Signal Processing (DSP) applications on FPGA has been extensively studied in the past. Up to now, Audio IPs11Throughout this paper, IP stands for Intellectual Property, i.e., a circuit component. were designed either “by hand” in VHDL or using predefined IPs in block synthesis environments. The advent of High Level Synthesis (HLS) allows for a real compilation flow from high-level audio DSP specifications down to FPGA bit-streams. This paper presents the principles and the implementation of the first “audio DSP compiler” targeting FPGAs. Our fully open-source system compiles audio DSP programs down to FPGA hardware and up to actual sound production. This compilation flow presents two important technological breakthroughs for audio programmers: achieving ultra-low latency real-time audio DSP (few micro-seconds) and the possibility of easily deploying systems with a large number of audio channels.
Maxime Popoff, Romain Michon, Tanguy Risset, Pierre Cochard, Stéphane Letz, Yann Orlarey, Florent de Dinechin
ASAP3
2020 MPU-based incremental checkpointing for transiently-powered systems
abstract
Transiently-powered devices are a class of small devices powered by energy harvesting. Because such devices are subject to frequent power outages, many recent works propose to checkpoint data residing in volatile RAM into non-volatile RAM. In this article, we propose a new incremental checkpointing mechanism supported by a common hardware component, namely a Memory Protection Unit (MPU). This mechanism leverages the hardware interrupts of the MPU: volatile RAM is read-only on boot and is progressively unlocked as soon as protection violations occur. The MPU interrupt handler is designed to flag the corresponding volatile RAM blocks as dirty, i.e., modified. When a power outage is foreseen to be imminent, the software simply has to copy the dirty blocks from volatile RAM into the non-volatile RAM to ensure application progress over power outages. We validate our approach analytically and in cycle-accurate simulation, and we show that the proposed solution can be easily implemented on real hardware.
Gautier Berthou 0001, Kevin Marquet, Tanguy Risset, Guillaume Salagnac
DSD3
2020 Intermittent Computing with Peripherals, Formally Verified
abstract
Transiently-powered systems featuring non-volatile memory as well as external peripherals enable the development of new low-power sensor applications. However, as programmers, we are ill-equipped to reason about systems where power failures are the norm rather than the exception. A first challenge consists in being able to capture all the volatile state of the application -- external peripherals included -- to ensure progress. A second, more fundamental, challenge consists in specifying how power failures may interact with peripheral operations. In this paper, we propose a formal specification of intermittent computing with peripherals, an axiomatic model of interrupt-based checkpointing as well as its proof of correctness, machine-checked in the Coq proof assistant. We also illustrate our model with several systems proposed in the literature.
Gautier Berthou 0001, Pierre-Évariste Dagand, Delphine Demange, Rémi Oudin, Tanguy Risset
LCTES5
2019 Sytare: A Lightweight Kernel for NVRAM-Based Transiently-Powered Systems
abstract
In a near future, energy harvesting is expected to replace batteries in ultra-low-power embedded systems. Research prototypes of such systems have recently been proposed. As the power harvested in the environment is very low, such systems need to cope with frequent power outages. They are referred to as transiently-powered systems (TPS). In order to execute non-trivial applications, TPS need to retain information between power losses. To achieve this goal, emerging non-volatile memory (NVM) technologies are a key enabler: they provide a lightweight solution to retain, between power outages, the state of an application and of its peripheral devices. These include sensors, serial interface or radio devices for instance. Existing works have described various checkpointing mechanisms to adapt embedded applications to TPS but the use of peripherals was not yet handled. in these works. This paper proposes a solution for embedded applications using any peripheral device to run despite transient power. We follow a kernel-oriented approach resulting in minimal impact on the programming model of the application. We implement the new concepts in our lightweight kernel called Sytare, running on an MSP430FR5739 micro-controller and we analyze the cost of the proposed solution.
Gautier Berthou 0001, Tristan Delizy, Kevin Marquet, Tanguy Risset, Guillaume Salagnac
IEEE Trans. Computers4
2018 UWB Ranging for Rapid Movements
abstract
Ultra Wide Band (UWB) provides ranging capabilities much more precise than other radio communication technologies. This paper presents experimental measurement of ranging precision between two UWB tags when one of the tag is moving fast. The use of a specific electro-pneumatic actuator provides precise ground truth for real distance. This study will allow to improve ranging for human wearable tags as for example sportsmen positioning or interaction between dancers during live performances. We show in particular how the inherent noisy measurement has to be smoothed to obtain more accurate ranging.
Tanguy Risset, Claire Goursaud, Xavier Brun, Kevin Marquet, Fabrice Meyer
IPIN1
2018 Estimating the Impact of Architectural and Software Design Choices on Dynamic Allocation of Heterogeneous Memories
abstract
Reducing energy consumption is a key challenge to the realization of the Internet of Things. While emerging memory technologies may offer power reduction, they come with major drawbacks such as high latency or limited endurance. As a result, system designers tend to juxtapose several memory technologies on the same chip. This paper studies the interactions between dynamic memory allocation and architectural choices regarding this heterogeneity. We provide cycle accurate simulations of embedded platforms with various memory technologies and we show that different dynamic allocation strategies have a major impact on performance. We demonstrate that interesting performance gains can be achieved even for a low fraction of heap objects in fast memory, but only with a clever data placement strategy between memory banks.
Tristan Delizy, Stephane Gros, Kevin Marquet, Matthieu Moy, Tanguy Risset, Guillaume Salagnac
RSP5
2016 A New Compilation Flow for Software-Defined Radio Applications on Heterogeneous MPSoCs
abstract
The advent of portable software-defined radio ( sdr ) technology is tightly linked to the resolution of a difficult problem: efficient compilation of signal processing applications on embedded computing devices. Modern wireless communication protocols use packet processing rather than infinite stream processing and also introduce dependencies between data value and computation behavior leading to dynamic dataflow behavior. Recently, parametric dataflow has been proposed to support dynamicity while maintaining the high level of analyzability needed for efficient real-life implementations of signal processing computations. This article presents a new compilation flow that is able to compile parametric dataflow graphs. Built on the llvm compiler infrastructure, the compiler offers an actor-based C++ programming model to describe parametric graphs, a compilation front end for graph analysis, and a back end that currently matches the Magali platform: a prototype heterogeneous MPSoC dedicated to LTE-Advanced. We also introduce an innovative scheduling technique, called microscheduling , allowing one to adapt the mapping of parametric dataflow programs to the specificities of the different possible MPSoCs targeted. A specific focus on fifo sizing on the target architecture is presented. The experimental results show compilation of 3 gpp lte - a dvanced demodulation on Magali with tight memory size constraints. The compiled programs achieve performance similar to handwritten code.
Mickaël Dardaillon, Kevin Marquet, Tanguy Risset, Jérôme Martin, Henri-Pierre Charles
ACM Trans. Archit. Code Optim.3
2015 A wireless, low-power, smart sensor of cardiac activity for clinical remote monitoring
abstract
This paper presents the development of a wireless wearable sensor for the continuous, long-term monitoring of cardiac activity. Heart rate assessment, as well as heart rate variability parameters are computed in real time directly on the sensor, thus only a few parameters are sent via wireless communication for power saving. Hardware and software methods for heart beat detection and variability calculation are described and preliminary tests for the evaluation of the sensor are presented. With an autonomy of 48 hours of active measurement and a Bluetooth Low Energy radio technology, this sensor will form a part of a wireless body network for the remote mobile monitoring of vital signals in clinical applications requiring automated collection of health data from multiple patients.
Bertrand Massot, Tanguy Risset, Gregory Michelet, Eric McAdams
HealthCom2
2014 A compilation flow for parametric dataflow: Programming model, scheduling, and application to heterogeneous MPSoC
abstract
Efficient programming of signal processing applications on embedded systems is a complex problem. High level models such as Synchronous dataflow (SDF) have been privileged candidates for dealing with this complexity. These models permit to express inherent application parallelism, as well as analysis for both verification and optimization. Parametric dataflow models aim at providing sufficient dynamicity to model new applications, while at the same time maintaining the high level of analyzability needed for efficient real life implementations.
Mickaël Dardaillon, Kevin Marquet, Tanguy Risset, Jérôme Martin, Henri-Pierre Charles
CASES3
2012 Software defined radio architecture survey for cognitive testbeds
abstract
In this paper we present a survey of existing prototypes dedicated to software defined radio. We propose a classification related to the architectural organization of the prototypes and provide some conclusions about the most promising architectures. This study should be useful for cognitive radio testbed designers who have to choose between many possible computing platforms. We also introduce a new cognitive radio testbed currently under construction and explain how this study have influenced the test-bed designers choices.
Mickaël Dardaillon, Kevin Marquet, Tanguy Risset, Antoine Scherrer
IWCMC3
2009 The Radio Virtual Machine: A solution for SDR portability and platform reconfigurability
abstract
Instead of a single circuit dedicated to a particular physical (PHY) layer standard, a Software Defined Radio (SDR) platform embeds several hardware accelerators which enable it to support different modulation schemes. In this study we propose an architecture for a SDR PHY layer based on the Virtual Machine (VM) concept. Once a program is compiled in a portable byte-code, the VM can then execute it to manage the desired PHY layer. We demonstrate the feasibility of the proposed architecture through a case study and a proof-of-concept implementation.
Riadh Ben Abdallah, Tanguy Risset, Antoine Fraboulet, Yves Durand
IPDPS2
2009 A reindexing based approach towards mapping of DAG with affine schedules onto parallel embedded systems
Clémentin Tayou Djamégni, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset, Maurice Tchuenté
J. Parallel Distributed Comput.4
2006 A Generic Multi-Phase On-Chip Traffic Generation Environment
abstract
In the process of mapping compute-intensive algorithms onto arrays of processing elements (PEs) an efficient usage of channels between PEs and registers within PEs is crucial for achieving a significant algorithm acceleration. In this paper this problem is solved for algorithms represented as systems of uniform recurrence equations. We address an optimization problem in order to realize the algorithmic data dependencies within the processor array (PA) with minimum cost for channels and registers. There, we use a new mapping approach which allows a direct mapping of the algorithm onto the PA by a partitioning method. In contrast to existing approaches, the authors consider the issue of avoiding redundant usage of channels and registers, which can appear if one instance of a variable has to be transferred from a source PE to several sink PEs. Further, a solution of the optimization problem determines the schedule for the transfer of the variable instances in the channels and their storage in registers as well as the inner schedule for the operations in the PEs. We illustrate our method on the edge detection algorithm
Antoine Scherrer, Antoine Fraboulet, Tanguy Risset
ASAP3
2005 Hardware/Software Interface for Multi-Dimensional Processor Arrays
abstract
On most recent systems on chip, the performance bottleneck is the on-chip communication medium, bus or network. Multimedia applications require a large communication bandwidth between the processor and graphic hardware accelerators, hence an efficient communication scheme using burst mode is mandatory. In the context of data-flow hardware accelerators, we approach this problem as a classical resource-constrained problem. We explain how to use recent optimization techniques so as to define a conflict-free schedule of input/output for multi-dimensional processor arrays (e.g. 2D grids). This schedule is static and allows us to perform further optimizations such as grouping successive data in packets to operate in burst mode. We also present an effective VHDL implementation on FPGA and compare our approach to a run-time congestion resolution showing important gains in hardware area.
Alain Darte, Steven Derrien, Tanguy Risset
ASAP3
2004 Efficient On-Chip Communications for Data-Flow IPs
Antoine Fraboulet, Tanguy Risset
ASAP2
2003 Hardware Synthesis for Multi-Dimensional Time
abstract
We introduce some basic principles for extending the classical systolic synthesis methodology to multidimensional time. Multidimensional scheduling enables complex algorithms that do not admit linear schedules to be parallelized, but it also requires the use of memories in the architecture. We explain how to obtain compatible allocation and memory functions for VLSI (or SIMD-like code) generation. We also present an original mechanism for controlling a VLSI architecture that has a multidimensional schedule. A structural VHDL code has been derived and synthesized (for implementation on FPGA platforms) using these systematic design principles. These results are preliminary steps to the hardware synthesis for multidimensional time.
Anne-Claire Guillou, Patrice Quinton, Tanguy Risset
ASAP3
2002 Advances in Bit Width Selection Methodology
abstract
We describe a method for the formal determination of signal bit width in fixed point VLSI implementations of signal processing algorithms containing loop nests. The main contribution of this paper is the use of results of the (max, +) algebraic theory to find the integral bit width of algorithms containing loop nests whose bound parameters are not statically known. Combined with recent results on fractional bit width determination, this can be used for 1-dimensional systolic-like arrays implementing linear signal processing algorithms. Although this technique is presented in the context of a specific high level design methodology (based on systems of affine recurrence equations), it can be used in many high level design environments.
David Cachera, Tanguy Risset
ASAP2
2001 Uniformization of Affine Dependance Programs for Parallel Embedded System Design
abstract
The paper is concerned with the uniformization of a system of affine recurrence equations. This transformation is used in the design (or compilation) of highly parallel embedded systems (VLSI systolic arrays, signal processing filters, etc.). We present and implement an automatic system to achieve uniformization of systems of affine recurrence equations. We unify the results from many earlier papers, develop some theoretical extensions, and then propose effective uniformization algorithms. Our results can be used in any high level synthesis tool based on polyhedral representation of nested loop computations.
Manju Manjunathaiah, Graham M. Megson, Sanjay V. Rajopadhye, Tanguy Risset
ICPP4
2001 Proving Properties of Multidimensional Recurrences with Application to Regular Parallel Algorithms
abstract
We present a set of verification methods to prove properties of parallel systems described by means of multidimensional affine recurrence equations. We use polyhedral analysis and transformation techniques together with theorem proving. Polyhedral techniques allow us to handle simple but otherwise costly proof steps, while theorem proving provides more expressivity and more complex proof techniques. This allows large, generic and structured systems to be verified. These methods are implemented in the M-MAlpha environment using the PVS theorem prover. 1
David Cachera, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset
IPDPS4
2000 On the Study of VLSI Derivation for Optical Flow Estimation
abstract
In this paper we propose studying several ways to implement a realistic and efficient VLSI design for a gradient-based dense motion estimator. The kind of estimator we focus on belongs to the class of differential methods. It is classically based on the optical flow constraint equation in association with a smoothness regularization term and also incorporates robust cost functions to alleviate the influence of large residuals. This estimator is expressed as the minimization of a global energy function defined within the framework of an incremental formulation associated with a multiresolution setup. In order to make possible the conception of efficient hardware, we consider a modified minimization strategy. This new minimization strategy is not only well suited to VLSI derivation, it is also very efficient in terms of quality of the result. The complete VLSI derivation is realized using high-level specifications.
Étienne Mémin, Tanguy Risset
Int. J. Pattern Recognit. Artif. Intell.2
2000 Derivation of systolic algorithms for the algebraic path problem by recurrence transformations
Clémentin Tayou Djamégni, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset
Parallel Comput.4
1999 The Algebraic Path Problem Revisited
Sanjay V. Rajopadhye, Claude Tadonki, Tanguy Risset
Euro-Par3
1998 Linear Programming Models for Scheduling Systems of Affine Recurrence Equations - A Comparative Study
Stephan Balev, Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset
SPAA4
1996 Extension Of The Alpha Language To Recurrences On Sparse Periodic Domains
abstract
ALPHA is a functional language based on systems of affine recurrence equations over polyhedral domains. We present an extension of ALPHA to deal with sparse polyhedral domains. Such domains are modeled by Z-polyhedra, namely the intersection of lattices and polyhedra. We summarize the mathematical closure properties of Z-polyhedra, and we show how the important features of ALPHA, namely normalization, substitution, change of basis, are preserved in the extension.
Patrice Quinton, Sanjay V. Rajopadhye, Tanguy Risset
ASAP3
1996 Resource-constrained scheduling of partitioned algorithms on processor arrays
Michèle Dion, Tanguy Risset, Yves Robert
Integr.2
1995 Precise Tiling for Uniform Loop Nests
abstract
The subject of this article is a hyperplane partitioning problem applied to perfect loop nests. This work is aimed at increasing the computation granularity to reduce the overhead due to communication. This study is different from previous work as it takes redundant communication into account. We propose an algorithm giving the optimal solution and various examples to show the validity of this report.
Pierre-Yves Calland, Tanguy Risset
ASAP2
1994 (Pen)-ultimate tiling?
Pierre Boulet, Alain Darte, Tanguy Risset, Yves Robert
Integr.3
1993 A real-time systolic algorithm for on-the-fly hidden surface removal
abstract
Hidden surface removal for real-time realistic display of complex scenes requires intensive computation and justifies usage of parallelism to provide the needed response time. The authors present a systolic algorithm that identifies visible segments on a scanline with the "real-time" characteristic: visible segments are output on-the-fly as soon as segments are input to the systolic array. The proposed systolic architecture consists of a linear array of simple cells. The size of the systolic array is equal to the maximum number of input segments that are crossed by a vertical line. The correctness of the systolic algorithm is proven by first establishing the correctness of an initial version, which is then successfully transformed to more efficient versions, while preserving the properties for correctness.>
Tanguy Risset, Siang Wun Song
ASAP1
1992 A method to synthesize modular systolic arrays with local broadcast facility
abstract
The author proposes a method to synthesize modular systolic arrays with local broadcast facility (i.e. arrays containing wires of length lower than a fixed -technology dependent- constant). The synthesis is made from a dependence graph which is not uniform but 'locally broadcast'. This method aims at generalizing isolated results that have been recently reported on the acceleration of systolic algorithms by using extensions of the 'pure' systolic model (wire of length>1, wrap around, folding arrays, etc).>
Tanguy Risset
ASAP1
1991 Synthesizing systolic arrays: some recent developments
abstract
Methods for synthesizing systolic arrays from uniform DAGs are well understood. The idea is to extract from the original sequential algorithm a dependence graph where all incoming arcs to a given node come from a fixed-size neighborhood, so that dependencies are local. Space-time transformations are then used for scheduling the DAG (timing function) and mapping nodes onto physical processors (allocation function). Both linear and piece-wise linear mappings can be derived in a systematic way, and methods exist to optimize given criteria such as the execution time, the number of processors or the cell utilization. The authors survey three recent developments along the following lines: DAG uniformization and spacetime minimal arrays; mapping n-dimensional DAGs (n>or=3) onto linear arrays; and partitioning techniques for the efficient mapping of a computational DAG onto a fixed-size processor array.>
Alain Darte, Tanguy Risset, Yves Robert
ASAP2
1991 Uniform but non-local DAGS: a trade-off between pure systolic and SIMD solutions
abstract
The authors derive processor arrays which are synthesized from uniform but non-local DAGs. They introduce a scope-b broadcast transformation that amounts working with dependence vectors of 'length' b. The parameter b can be adjusted to cope with current integration constraints. They explain the transformation with the Gaussian elimination algorithm. For instance with b=3, they derive an array which has the same number of cells as the Ahmed-Delosme-Morf array but whose execution time is /sup 7n///sub 3/+o(n), as opposed to 3n+o(n). They also apply the scope-b broadcast technique to synthesize faster processor arrays for the algebraic path problem.>
Tanguy Risset, Yves Robert
ASAP1
1990 Implementing Gaussian elimination on a matrix-matrix multiplication systolic array
Tanguy Risset
Parallel Comput.1