EDBT 2026 Demo / reviewers in the wild / expert
Lionel Torres
dblp:64/6964
· DBLP profile ↗
71ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-5807-5070ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 17 · 3 since 2021Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Cache Policies on FPGA-Accelerated Simulations: Tradeoffs Between Usability and Simulation SpeedabstractInternational audience Soraya Mobaraki, Thierry Gil, Lionel Torres, David Novo |
RSP | 3 |
| 2024 | Analyzing GPU Energy Consumption in Data Movement and StorageabstractGPUs are the prevailing solution to execute high-performance tasks (e.g., machine learning training). As the peak performance of modern GPUs increases with each generation, so does their thermal design power (TDP). Hence, identifying energy bottlenecks in the GPU architecture is crucial to designing more efficient architectures in the future. However, due to the complex proprietary nature of modern GPU architectures, providing a detailed breakdown of the GPU energy consumption is not trivial. The goal of this work is to estimate a lower bound for the energy consumed by data movement and storage in modern GPU architectures, leveraging internal power sensors. We establish a basic energy model for modern GPUs, focused on data movement to/from the hardware-managed caches and software-managed memories. We propose a methodology to calibrate the energy model using microbenchmarks, performance counters, and the internal power sensor. We experimentally calibrate the model on an A100 NVIDIA GPU. Then, we challenge the consistency of the results by cross-validating with modified microbenchmarks with additional instructions. Finally, we use the calibrated energy model to evaluate breakdowns for workloads of increasing complexity (e.g., a ResNet-50 training iteration with different software optimizations). Our results show that data movement dominates the dynamic energy consumption of the GPU (up to 84%), with DRAM accesses being the main contributor. Paul Delestrac, Jonathan Miquel, Debjyoti Bhattacharjee, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo |
ASAP | 6 |
| 2024 | Multi-Level Analysis of GPU Utilization in ML Training WorkloadsabstractTraining time has become a critical bottleneck due to the recent proliferation of large-parameter ML models. GPUs continue to be the prevailing architecture for training ML models. However, the complex execution flow of ML frameworks makes it difficult to understand GPU computing resource utilization. Our main goal is to provide a better understanding of how efficiently ML training workloads use the computing resources of modern GPUs. To this end, we first describe an ideal reference execution of a GPU-accelerated ML training loop and identify relevant metrics that can be measured using existing profiling tools. Second, we produce a coherent integration of the traces obtained from each profiling tool. Third, we leverage the metrics within our integrated trace to analyze the impact of different software optimizations (e.g., mixed-precision, various ML frameworks, and execution modes) on the throughput and the associated utilization at multiple levels of hardware abstraction (i.e., whole GPU, SM subpartitions, issue slots, and tensor cores). In our results on two modern GPUs, we present seven takeaways and show that although close to 100% utilization is generally achieved at the GPU level, average utilization of the issue slots and tensor cores always remains below 50% and 5.2%, respectively. Paul Delestrac, Debjyoti Bhattacharjee, Simei Yang, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo |
DATE | 6 |
| 2022 | Pref-X: a framework to reveal data prefetching in commercial in-order coresabstractComputer system simulators are major tools used by architecture researchers to develop and evaluate new ideas. Clearly, such evaluations are more conclusive when compared to commercial state-of-the-art architectures. However, the behavior of key components in existing processors is often not disclosed, complicating the construction of faithful reference models. The data prefetching engine is one of such obscured components that can have a significant impact on key metrics such as performance and energy. Quentin Huppert, Francky Catthoor, Lionel Torres, David Novo |
DAC | 3 |
| 2022 | Demystifying the TensorFlow Eager Execution of Deep Learning Inference on a CPU-GPU TandemabstractMachine Learning (ML) frameworks are tools that facilitate the development and deployment of ML models. These tools are major catalysts of the recent explosion in ML models and hardware accelerators thanks to their high programming abstraction. However, such an abstraction also obfuscates the run-time execution of the model and complicates the understanding and identification of performance bottlenecks. In this paper, we demystify how a modern ML framework manages code execution from a high-level programming language. We focus our work on the TensorFlow eager execution, which remains obscure to many users despite being the simplest mode of execution in TensorFlow. We describe in detail the process followed by the runtime to run code on a CPU-GPU tandem. We propose new metrics to analyze the framework's runtime performance overhead. We use our metrics to conduct in-depth analysis of the inference process of two Convolutional Neural Networks (CNNs) (LeNet-5 and ResNet-50) and a transformer (BERT) for different batch sizes. Our results show that GPU kernels execution need to be long enough to exploit thread parallelism, and effectively hide the runtime overhead of the ML framework. Paul Delestrac, Lionel Torres, David Novo |
DSD | 2 |
| 2021 | Memory Hierarchy Calibration Based on Real Hardware In-order Cores for Accurate SimulationabstractComputer system simulators are major tools used by architecture researchers. Two key elements play a role in the credibility of simulator results: (1) the simulator's accuracy, and (2) the quality of the baseline architecture. Some simulators, such as gem5, already provide highly accurate parameterized models. However, finding the right values for all these parameters to faithfully model a real architecture is still a problem. In this paper, we calibrate the memory hierarchy of an in-order core gem5 simulation to accurately model a real mobile Arm SoC. We execute small programs, which we design to stress specific parts of the memory system, to deduce key parameter values for the model. We compare the execution of SPEC CPU2006 benchmarks on the real hardware with the gem5 simulation. Our results show that our calibration reduces the average and worst-case IPC error by 36 % and 50%, respectively, when compared with a gem5 simulation configured with the default parameters. Quentin Huppert, Timon Evenblij, Manu Perumkunnil Komalan, Francky Catthoor, Lionel Torres, David Novo |
DATE | 5 |
| 2020 | A Universal Spintronic Technology based on Multifunctional Standardized StackabstractThe goal of the GREAT RIA project is to cointegrate multiple functions like sensors ("Sensing"), RF emitters or receivers ("Communicating") and logic/memory ("Process- ing/Storing") together within CMOS technology by adapting the Spin-Transfer Torque Magnetic Tunnel Junction (STT-MTJ), elementary constitutive cell of the MRAM memories, to a single baseline technology. Based on the STT unique set of performances (non-volatility, high speed, infinite endurance and moderate read/write power), GREAT will achieve the same goal as heterogeneous integration of devices but in a much simpler way. This will lead to a unique STT-MTJ cell technology called Multifunctional Standardized Stack (MSS). This paper presents the lessons learned in the project from the technology, compact modeling, process design kit, standard cells, as well as memory and system level design evaluation and exploration. The proposed technology and toolsets are giant leaps towards heterogeneous integrated technology and architectures for IoT. Mehdi Baradaran Tahoori, Sarath Mohanachandran Nair, Rajendra Bishnoi, Lionel Torres, Sophiane Senni, Guillaume Patrigeon, Pascal Benoit, Gregory di Pendina, Guillaume Prenat |
DATE | 4 |
| 2018 | Main memory organization trade-offs with DRAM and STT-MRAM options based on gem5-NVMain simulation frameworksabstractCurrent main memory organizations in embedded and mobile application systems are DRAM dominated. The ever-increasing gap between today's processor and memory speeds makes the DRAM subsystem design a major aspect of computer system design. However, the limitations to DRAM scaling and other challenges like refresh provide undesired trade-offs between performance, energy and area to be made by architecture designers. Several emerging NVM options are being explored to at least partly remedy this but today it is very hard to assess the viability of these proposals because the simulations are not fully based on realistic assumptions on the NVM memory technologies and on the system architecture level. In this paper, we propose to use realistic, calibrated STT-MRAM models and a well calibrated cross-layer simulation and exploration framework, named SEAT, to better consider technologies aspects and architecture constraints. We will focus on general purpose/mobile SoC multi-core architectures. We will highlight results for a number of relevant benchmarks, representatives of numerous applications based on actual system architecture. The most energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 27% at the cost of 2x the area and the least energy efficient STT-MRAM based main memory proposal provides an average energy consumption reduction of 8% at the around the same area or lesser when compared to DRAM. Manu Perumkunnil Komalan, Hyungrock Oh, Matthias Hartmann, Sushil Sakhare, Christian Tenllado, José Ignacio Gómez, Gouri Sankar Kar, Arnaud Furnémont, Francky Catthoor, Sophiane Senni, David Novo, Abdoulaye Gamatié, Lionel Torres |
DATE | 13 |
| 2018 | Using multifunctional standardized stack as universal spintronic technology for IoTabstractFor monolithic heterogeneous integration, fast yet low-power processing and storage, and high integration density, the objective of the EU GREAT project is to co-integrate multiple digital and analog functions together within CMOS by adapting the Magnetic Tunneling Junctions (MTJs) into a single baseline technology enabling logic, memory, and analog functions, particularly for Internet of Things (IoT) platforms. This will lead to a unique STT-MTJ cell technology called Multifunctional Standardized Stack (MSS). This paper presents the progress in the project from the technology, compact modeling, process design kit, standard cells, as well as memory and system level design evaluation and exploration. The proposed technology and toolsets are giant leaps towards heterogeneous integrated technology and architectures for IoT. Mehdi Baradaran Tahoori, Sarath Mohanachandran Nair, Rajendra Bishnoi, Sophiane Senni, Jad Mohdad, Frédérick Mailly, Lionel Torres, Pascal Benoit, Abdoulaye Gamatié, Pascal Nouet, Frederic Ouattara, Gilles Sassatelli, Kotb Jabeur, Pierre Vanhauwaert, A. Atitoaie, I. Firastrau, Gregory di Pendina, Guillaume Prenat |
DATE | 7 |
| 2018 | From Spintronic Devices to Hybrid CMOS/Magnetic System On Chipabstract"Beyond CMOS" is today one of the major research directions in semiconductor industries to address current integrated circuit issues. Many alternative technologies are currently under investigation to deal with the scaling limits of CMOS technology. This paper presents the design of a full system on chip based on a hybrid CMOS/Magnetic process. Spin-transfer-torque magnetic tunnel junctions are used to design different functions such as logic, memory, security and analog IP blocks. Sophiane Senni, Frederic Ouattara, Jad Mohdad, Kaan Sevin, Guillaume Patrigeon, Pascal Benoit, Pascal Nouet, Lionel Torres, François Duhem, Gregory di Pendina, Guillaume Prenat |
VLSI-SoC | 8 |
| 2017 | Embedded systems to high performance computing using STT-MRAMabstractThe scaling limits of CMOS have pushed many researchers to explore alternative technologies for beyond CMOS circuits. In addition to the increased device variability and process complexity led by the continuous decreasing size of CMOS transistors, heat dissipation effects limit the density and speed of current systems-on-chip. For beyond CMOS systems, the emerging memory technology STT-MRAM is seen as a promising alternative solution. This paper shows first how STT-MRAM can improve energy efficiency and reliability of future embedded systems. Then, a hybrid design exploration framework is presented to investigate the potential of STT-MRAM for high performance computing. Sophiane Senni, Thibaud Delobelle, Odilia Coi, Pierre-Yves Peneau, Lionel Torres, Abdoulaye Gamatié, Pascal Benoit, Gilles Sassatelli |
DATE | 5 |
| 2017 | A Design-Time Method for Building Cost-Effective Run-Time Power MonitoringabstractThe emergence of power as a first-class design constraint has fueled the proposal of a growing number of optimization techniques, seeking the best tradeoff to reach the maximum energy efficiency. Effective adaptation strategies depend critically on the monitoring method as an incorrect assessment of the system's state will result in poor decision making. Yet it is indeed a fundamental issue: how to get a precise estimation of the system's state, and especially in a cost-effective way? We address this question for the self-observation of the power consumption. We develop a method that combines several data mining algorithms to monitor the toggling activity on a few relevant signals selected at the register transfer-level. Our approach is based on a generic flow that is able to produce a power model for any register transfer level (RTL) circuit on any technology. This contribution is evaluated on a system on chip RTL model implemented on an field-programmable gate array technology. The experiments demonstrate that the proposed method achieves the accuracy of analog power sensors (error lower than 1%) at a finer granularity and in a cost-effective way. Mohamad Najem, Pascal Benoit, Mohamad El Ahmad, Gilles Sassatelli, Lionel Torres |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2016 | Non-Volatile Processor Based on MRAM for Ultra-Low-Power IoT DevicesabstractOver the past few years, a new era of smart connected devices has emerged in the market to enable the future world of the Internet of Things (IoT). A key requirement for IoT applications is the power consumption to allow very high autonomy in the case of battery-powered systems. Depending on the application, such devices will be most of the time in a low-power mode (sleep mode) and will wake up only when there is a task to accomplish (active mode). Emerging non-volatile memory technologies are seen as a very attractive solution to design ultra-low-power systems. Among these technologies, magnetic random access memory is a promising candidate, as it combines non-volatility, high density, reasonable latency, and low leakage. Integration of non-volatility as a new feature of memories has the great potential to allow full data retention after a complete shutdown with a fast wake-up time. This article explores the benefits of having a non-volatile processor to enable ultra-low-power IoT devices. Sophiane Senni, Lionel Torres, Gilles Sassatelli, Abdoulaye Gamatié |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | STT-MRAM-Based PUF Architecture Exploiting Magnetic Tunnel Junction Fabrication-Induced VariabilityabstractPhysically Unclonable Functions (PUFs) are emerging cryptographic primitives used to implement low-cost device authentication and secure secret key generation. Weak PUF s (i.e., devices able to generate a single signature or to deal with a limited number of challenges) are widely discussed in literature. One of the most investigated solutions today is based on SRAMs. However, the rapid development of low-power, high-density, high-performance SoCs has pushed the embedded memories to their limits and opened the field to the development of emerging memory technologies. The Spin-Transfer-Torque Magnetic Random Access Memory (STT-MRAM) has emerged as a promising choice for embedded memories due to its reduced read/write latency and high CMOS integration capability. In this article, we propose an innovative PUF design based on STT-MRAM memory. We exploit the high variability affecting the electrical resistance of the Magnetic Tunnel Junction (MTJ) device in anti-parallel magnetization. We will demonstrate that the proposed solution is robust, unclonable, and unpredictable. Elena I. Vatajelu, Giorgio Di Natale, Mario Barbareschi, Lionel Torres, Marco Indaco, Paolo Prinetto |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2015 | Potential applications based on NVM emerging technologies
Sophiane Senni, Raphael Martins Brum, Lionel Torres, Gilles Sassatelli, Abdoulaye Gamatié, Bruno Mussard |
DATE | 3 |
| 2014 | Power management through DVFS and dynamic body biasing in FD-SOI circuitsabstractThe emerging SOI technologies provide an increased body bias range compared to traditional bulk technologies, opening new opportunities. From the power management perspective, a new degree of freedom is added to the supply voltage and clock frequency variation, increasing the complexity of the power optimization problem. In this paper, a method is proposed to manage the power consumed in an FD-SOI circuit through supply and body bias voltages, and clock frequency variation. Results for a Digital Signal Processor in STMicroelectronics 28nm FD-SOI technology show that the power reduction ratio can reach 17%. Yeter Akgul, Diego Puschini, Suzanne Lesecq, Edith Beigné, Ivan Miro Panades, Pascal Benoit, Lionel Torres |
DAC | 7 |
| 2014 | Aging and voltage scaling impacts under neutron-induced soft error rate in SRAM-based FPGAsabstractThis work investigates the effects of aging and voltage scaling in neutron-induced bit-flip in SRAM-based FPGAs. Experimental results show that aging and voltage scaling can increase in at least two times the susceptibility of SRAM-based FPGAs to Soft Error Rate (SER). These results are innovative, because they combine three real effects that occur in programmable circuits operating at ground-level applications. In addition, a model at electrical simulation for aging, soft error and different voltages was described to investigate the effects observed at the practical neutron irradiation experiment. Results can guide designers to predict soft error effects during the lifetime of devices operating in different power supply mode. Fernanda Lima Kastensmidt, Jorge L. Tonfat, Thiago Hanna Both, Paolo Rech, Gilson I. Wirth, Ricardo Augusto da Luz Reis, Florent Bruguier, Pascal Benoit, Lionel Torres, Christopher Frost 0002 |
ETS | 9 |
| 2014 | Aging effects in FPGAs: an experimental analysisabstractModern Field Programmable Gate Arrays (FPGAs) are built using the most advanced technology nodes to meet performance and power demands. This makes them susceptible to various reliability challenges at nano-scale, and in particular to transistor aging. In this paper, an experimental analysis is made to identify the main parameters and phenomena influencing the performance degradation of FPGAs. For that purpose, a set of controlled ring-oscillator-based sensors with different frequencies and tunable activity control are implemented on a Spartan-6 FPGA. Thus, the internal switching activities (SAs) and signal probabilities (SPs) of the sensors can be varied. We performed accelerated-lifetime conditions using elevated temperatures and voltages in a controlled setting to stress the FPGA. A novel monitoring method based on measuring the electromagnetic emissions of the FPGA is used to accurately monitor the performance of the sensors before and after the stress. The experiments reveal the extent of performance degradations, the impact of SPs and SAs, and the relative impacts of BTI and HCI aging factors. Abdulazim Amouri, Florent Bruguier, Saman Kiamehr, Pascal Benoit, Lionel Torres, Mehdi Baradaran Tahoori |
FPL | 5 |
| 2014 | Method for dynamic power monitoring on FPGAsabstractThe ever-increasing integration densities make it possible to configure multi-core systems composed of hundreds of blocks on existing FPGAs that may influence overall consumption differently. Observing total consumption is not sufficient to accurately assess internal circuit activity to be able to deploy effective adaptation strategies. In this case monitoring techniques are required. This paper presents a CAD flow for high-level dynamic power estimation on FPGAs. The method is based on the monitoring of toggling activity for relevant signals by introducing event counters. The appropriate signals are selected using the Greedy Stepwise filter. Our approach is based on a generic method that is able to produce a power model for any block-based circuit. We evaluated our contribution on a SoC RTL model implemented on Spartan3, Virtex5, and Spartan6 FPGAs. A power model and monitors are automatically generated to achieve the best tradeoff between accuracy and overhead. Mohamad Najem, Pascal Benoit, Florent Bruguier, Gilles Sassatelli, Lionel Torres |
FPL | 5 |
| 2013 | Practical Analysis of RSA Countermeasures Against Side-Channel Electromagnetic Attacks
Guilherme Perin, Laurent Imbert, Lionel Torres, Philippe Maurine |
CARDIS | 3 |
| 2013 | Electromagnetic Analysis on RSA Algorithm Based on RNSabstractThis paper proposes a robustness evaluation of an RSA cryptosystem against collision attacks and correlation electromagnetic analysis. Our hardware co-processor is based on the Residue Number System (RNS) in order to perform modular operations over large numbers. To increase its robustness against Side-Channel Analysis, we implemented two different countermeasures. The first one spatially permutates the elements of the RNS bases in order to blur electromagnetic emanations. The second countermeasure aims at randomizing RNS bases before each modular exponentiation. To the best knowledge of authors, this is the first paper that explores the robustness of RNS-RSA against EM analyses. Guilherme Perin, Laurent Imbert, Lionel Torres, Philippe Maurine |
DSD | 3 |
| 2013 | Trends on the application of emerging nonvolatile memory to processors and programmable devicesabstractA number of non-volatile memory technologies (NVMs) emerged in the past years. They promise to cope with limitations of standard memory technologies, such as scalability and idle power consumption. Numerous logic circuits based on these emerging technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and discuss their potential applications in this field. In the first part, this article provides a survey on the application of those memories to programmable devices. The second section is dedicated to the use of NVMs in the processor's memory hierarchy, where we discuss potential applications based on a preliminary study we performed. Results were obtained using TAS-MRAM NVM technology. Lionel Torres, Raphael Martins Brum, Vitorio Cargnini, Gilles Sassatelli |
ISCAS | 1 |
| 2012 | Amplitude demodulation-based EM analysis of different RSA implementationsabstractThis paper presents a fully numeric amplitude-demodulation based technique to enhance simple electromagnetic analyses. The technique, thanks to the removal of the clock harmonics and some noise sources, allows efficiently disclosing the leaking information. It has been applied to three different modular exponentiation algorithms, mapped onto the same multiplexed architecture. The latter is able to perform the exponentiation with successive modular multiplications using the Montgomery method. Experimental results demonstrate the efficiency of the applied demodulation based technique and also point out the remaining weaknesses of the considered architecture to retrieve secret keys. Guilherme Perin, Lionel Torres, Pascal Benoit, Philippe Maurine |
DATE | 2 |
| 2012 | SecURe DPR: Secure update preventing replay attacks for dynamic partial reconfigurationabstractDynamic partial reconfiguration is a growing need for SRAM FPGA-based embedded systems. This feature allows reconfiguring parts of the FPGA while others continue to run. But it may introduce security breaches affecting FPGA configuration. In this paper, a secure protocol to ensure confidentiality, integrity, authenticity and up-to-dateness is described and applied to dynamic partial reconfiguration. Two common threat models are addressed for industrially-driven use cases. The implementation can perform both secure update and reconfiguration without significantly affecting performances. Florian Devic, Lionel Torres, Jérémie Crenne, Benoît Badrignans, Pascal Benoit |
FPL | 2 |
| 2012 | Enhancing Electromagnetic Analysis Using Magnitude Squared IncoherenceabstractThis paper demonstrates that magnitude squared incoherence (MSI) analysis is efficient to localize hot spots, i.e., points at which focused electromagnetic (EM) analyses can be applied with success. It is also demonstrated that MSI may be applied to enhance differential EM analyses (DEMA) based on difference of means (DoM). Amine Dehbaoui, Victor Lomné, Thomas Ordas, Lionel Torres, Michel Robert, Philippe Maurine |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Optimizing an Open-Source Processor for FPGAs: A Case StudyabstractOptimizing a processor for FPGA architectures is a challenging task. In this paper, we attempt to bridge the performance gap between commercial and open-source processors by introducing various design and implementation strategies at the register transfer abstraction level, where most optimizations require several design trade-offs to ensure an efficient and proper use of available resources. Using an open-source processor as a case study, we demonstrate the effectiveness of the proposed methods through a set of synthesis and benchmark results. Lyonel Barthe, Vitorio Cargnini, Pascal Benoit, Lionel Torres |
FPL | 4 |
| 2011 | A New Process Characterization Method for FPGAs Based on Electromagnetic AnalysisabstractThanks to their inherent regularity and reconfigurability, FPGAs offer an ideal structure to manage process variability. Recent works from the literature have addressed the process characterization problem for FPGAs: proposed approaches rely on process sensors (ring oscillators) and a measurement subsystem implemented into the configurable logic blocks. In this article, we propose for the first time in the literature a non-invasive characterization method based on electromagnetic analysis. The whole experimental set-up is described and the characterization accuracy is discussed. This paper proves the feasibility of this new method on FPGAs. Florent Bruguier, Pascal Benoit, Philippe Maurine, Lionel Torres |
FPL | 4 |
| 2011 | Magnetic memory (MRAM), a new area for 2D and 3D SoC/SiP designabstractNo abstract available. Lionel Torres, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Design of MRAM based logic circuits and its applicationsabstractAs the fabrication technology node shrinks down to 90nm or below, high standby power becomes one of the major critical issues for CMOS logic circuits due to the high leakage currents. A number of non-volatile storage technologies such as FRAM, MRAM, PCRAM and RRAM and so on, are under investigation to bring the non-volatility into the logic circuits and then eliminate completely the standby power issue. Thanks to its infinite endurance, high switching/sensing speed and easy 3D integration after CMOS process, MRAM is considered as the most promising one. Numerous logic circuits based on MRAM technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and their potential applications in the future. Weisheng Zhao 0001, Lionel Torres, Yoann Guillemenet, Vitorio Cargnini, Yahya Lakys, Jacques-Olivier Klein, Dafine Ravelosona, Gilles Sassatelli, Claude Chappert |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Predictive Dynamic Frequency Scaling for Multi-Processor Systems-on-ChipabstractThis paper proposes a novel strategy for optimizing resources in Multi-Processor Systems-on-Chip (MPSoC). The approach is based on using control-loop feedback mechanism to maximize the efficiency on exploiting available resources such as CPU time, operating frequency, etc. Each Processing Element (PE) in the architecture is equipped with a frequency scaling module responsible for tuning the frequency of processors at run-time according to the application requirements. Results show the system's capability of adapting to disturbing conditions. For validation purposes we have implemented a multi-threaded MJPEG decoder together with an ADPCM audio decoder and a FIR. Gabriel Marchesan Almeida, Rémi Busseuil, Everton Carara, Nicolas Hebert, Sameer Varyani, Gilles Sassatelli, Pascal Benoit, Lionel Torres, Fernando Gehm Moraes |
ISCAS | 8 |
| 2011 | Evaluation of a distributed fault handler method for MPSoCabstractThe increasing infant mortality and wear out failure rates observed in very deep sub micron silicon technologies is now a major problem for the design of future high-density SoCs. Emerging architectures based on Multi-Processor SoCs (MPSoCs) give the opportunity to exploit the natural redundancy to control the system performance in presence of failures. In this paper we evaluate the impact of a distributed fault-handler strategy on the system in term of cost and feasibility. We also discuss strategies for applying this technique to a distributed MPSoC architecture. Nicolas Hebert, Gabriel Marchesan Almeida, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
ISCAS | 5 |
| 2011 | Embedded MRAM for high-speed computingabstractAs the fabrication technology node shrinks down to 90nm or below, high standby power becomes one of the major critical issues for CMOS high-speed computing circuits (e.g. logic and cache memory) due to the high leakage currents. A number of non-volatile storage technologies such as FeRAM, MRAM, PCRAM and RRAM and so on, are under investigation to bring the non-volatility into the logic circuits and then eliminate completely the standby power issue. Thanks to its infinite endurance, high switching/sensing speed and easy 3D integration after CMOS process, MRAM is considered as the most promising one. Numerous logic circuits based on MRAM technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and discuss their potential applications in the future from both the physics and architecture points of view. Weisheng Zhao 0001, Yue Zhang 0010, Yahya Lakys, Jacques-Olivier Klein, Daniel Etiemble, D. Revelosona, Claude Chappert, Lionel Torres, Vitorio Cargnini, Raphael Martins Brum, Yoann Guillemenet, Gilles Sassatelli |
VLSI-SoC | 8 |
| 2010 | D-Scale: A Scalable System-Level Dependable Method for MPSoCsabstractThe increasing failure rates observed in very deep sub micron silicon technologies pose a major problem to the design of future high-density SoCs. While hardening techniques originated from critical application areas (automotive, avionics) exist, they usually incur a cost overhead that renders them inadequate for consumer market segments. Thus we present a concept, an implementation and an evaluation of a scalable software-hardware detection, isolation and recovery method. The method exploits the natural redundancy that exists in MPSoCs for enhancing their reliability. Based on the assumption that a transient loss of functionality can be tolerated, the proposed scheme relies on a hardware/software framework that makes it possible to diagnose and to isolate faulty processors in a distributed manner. It guarantees the integrity, improves the availability and eases the maintainability of the MPSoC at system-level. Nicolas Hebert, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
Asian Test Symposium | 4 |
| 2010 | When Failure Analysis Meets Side-Channel Attacks
Jerome Di-Battista, Jean-Christophe Courrège, Bruno Rouzeyre, Lionel Torres, Philippe Perdu |
CHES | 4 |
| 2010 | Heterogeneous vs homogeneous MPSoC approaches for a Mobile LTE modemabstractApplications like 4G baseband modem require single-chip implementation to meet the integration and power consumption requirements. These applications demand a high computing performance with real-time constraints, low-power consumption and low cost. With the rapid evolution of telecom standards and the increasing demand for multi-standard products, the need for flexible baseband solutions is growing. The concept of Multi-Processor System-on-Chip (MPSoC) is well adapted to enable hardware reuse between products and between multiple wireless standards in the same device. Heterogeneous architectures are well known solutions but they have limited flexibility. Based on the experience of two heterogeneous Software Defined Radio (SDR) telecom chipsets, this paper presents the homoGENEous Processor arraY (GENEPY) platform for 4G applications. This platform is built with Smart ModEm Processors (SMEP) interconnected with a Network-on-Chip. The SMEP, implemented in 65nm low-power CMOS, can perform 3.2 GMAC/s with 77 GBits/s internal bandwidth at 400MHz. Two implementations of homogeneous GENEPY are compared to a heterogeneous platform in terms of silicon area, performance and power consumption. Results show that a homogeneous approach can be more efficient and flexible than a heterogeneous approach in the context of 4G Mobile Terminals. Camille Jalier, Didier Lattard, Ahmed Amine Jerraya, Gilles Sassatelli, Pascal Benoit, Lionel Torres |
DATE | 6 |
| 2010 | Differential Power Analysis enhancement with statistical preprocessingabstractDifferential Power Analysis (DPA) is a powerful Side-Channel Attack (SCA) targeting as well symmetric as asymmetric ciphers. Its principle is based on a statistical treatment of power consumption measurements monitored on an Integrated Circuit (IC) computing cryptographic operations. A lot of works have proposed improvements of the attack, but no one focuses on ordering measurements. Our proposal consists in a statistical preprocessing which ranks measurements in a statistically optimized order to accelerate DPA and reduce the number of required measurements to disclose the key. Victor Lomné, Amine Dehbaoui, Philippe Maurine, Lionel Torres, Michel Robert |
DATE | 4 |
| 2010 | Survey of New Trends in Industry for Programmable Hardware: FPGAs, MPPAs, MPSoCs, Structured ASICs, eFPGAs and New Wave of Innovation in FPGAsabstractWe will present a survey of trends in the semiconductor industry for programmable hardware. The main objective of this paper is educational and the focus is FPGAs and its related or vs technologies which have emerged mostly in the second half of the last decade. We will try to analyze what were the prominent reasons for emerging of these technologies. What are the advantages and drawbacks of them, what makes FPGAs still most dominant in this area and will it be same or change in future. FPGAs themselves during this time have dramatically changed and the classical term FPGA does not fully characterize in name what FPGAs have actually become now. These changes and the continuing rising strength of multicore and ultimate power consumption challenge in industry will have what impact. Will these technologies collide or co-exist in future (nobody in industry or academics knows that and it is hard to predict). We will try to present the distinguishing technical and commercial potentials of different technologies which give an edge of one over the other. Syed Zahid Ahmed, Gilles Sassatelli, Lionel Torres, Laurent Rouge |
FPL | 3 |
| 2010 | Investigation of a Masking Countermeasure against Side-Channel Attacks for RISC-based Processor ArchitecturesabstractSide-Channel Attacks (SCAs) present a serious threat to the security of crypto-systems. In this paper, we show how a pipelined embedded processor opens the door to such attacks. To illustrate our approach, a concrete evaluation of the Xilinx's MicroBlaze soft-core processor is conducted. From these results, we suggest a new masking countermeasure suited for RISC-based architectures. The efficiency and limits of the proposed solution are finally evaluated. Lyonel Barthe, Pascal Benoit, Lionel Torres |
FPL | 3 |
| 2010 | Secure Protocol Implementation for Remote Bitstream Update Preventing Replay Attacks on FPGAabstractNowadays, there are lot of applications where remote update is an essential service. Indeed, in high volume sale products or space-based systems it is too expensive to retrieve the device in order to update it. Field Programmable Gate Arrays (FPGAs) are able to perform that with success through a network. However, this feature may give rise to security flaw like spoofing and replay attacks. These attacks consist in tampering the update of the hardware configuration or in replaying an old bitstream to downgrade the system. Several security schemes providing encryption and integrity checking of the bitstream have been proposed in the literature. However, they do not detect the replay of old FPGA configurations. Considering FPGA with embedded non-volatile memory, we propose a new protocol ensuring bitstream confidentiality, integrity and preventing old bitstreams replay. This work is the improvement and the implementation of previous presented ideas in order to achieve more flexibility. That is why we insist on the way to manage bitstream versions. We also evaluate the area and performance overhead of the proposed architecture. Florian Devic, Lionel Torres, Benoît Badrignans |
FPL | 2 |
| 2010 | Flexible and distributed real-time control on a 4G telecom MPSoCabstractApplications like 4G baseband modem require single-chip implementation to meet the integration and power consumption requirements. These applications demand a high computing performance with real-time constraints, low-power consumption and low cost. With the rapid evolution of telecom standards and the increasing demand for multi-standard products, the need for flexible baseband solutions is growing. The concept of Multi-Processor System-on-Chip (MPSoC) is well adapted to enable hardware reuse between products and between multiple wireless standards in the same device. Based on the experience of two heterogeneous Software Defined Radio (SDR) telecom chipsets, this paper presents a distributed control architecture for the homoGENEous Processor arraY (GENEPY) platform for 4G applications. This MPSoC platform is built with telecom baseband processors interconnected with a Network-on-Chip. The control is performed by a MIPS processor embedded in each baseband processor. This control processor can locally reconfigure and schedule the applications with real-time telecom constraints. Camille Jalier, Didier Lattard, Gilles Sassatelli, Pascal Benoit, Lionel Torres |
ISCAS | 5 |
| 2010 | Spatial EM jamming: A countermeasure against EM Analysis?abstractElectro-Magnetic Analysis has been identified as an efficient technique to retrieve the secret key of cryptographic algorithms. Although similar mathematically speaking, Power or Electro-Magnetic Analysis have different advantages in practice. Among the advantages of EM Analysis, the feasibility of attacking limited and bounded area of integrated systems is the key one. Within this context, the contribution of this paper is a countermeasure against local EM attack performed with tiny magnetic probes. The basic idea is to design circuits such that all datapaths and D-type Flip-Flops, involved in the computation of intermediate values of cryptographic elements, randomly change within a set of logically equivalent electrical paths that are spatially distributed within the Integrated Circuit (IC) die. François Poucheret, Lyonel Barthe, Pascal Benoit, Lionel Torres, Philippe Maurine, Michel Robert |
VLSI-SoC | 4 |
| 2010 | SARFUM: Security Architecture for Remote FPGA Update and MonitoringabstractRemote update of hardware platforms or embedded systems is a convenient service enabled by Field Programmable Gate Array (FPGA)-based systems. This service is often essential in applications like space-based FPGA systems or set-top boxes. However, having the source of the update be remote from the FPGA system opens the door to a set of attacks that may challenge the confidentiality and integrity of the FPGA configuration, the bitstream. Existing schemes propose to encrypt and authenticate the bitstream to thwart these attacks. However, we show that they do not prevent the replay of old bitstream versions, and thus give adversaries an opportunity for downgrading the system. In this article, we propose a new architecture called sarfum that, in addition to ensuring bitstream confidentiality and integrity, precludes the replay of old bitstreams. sarfum also includes a protocol for the system designer to remotely monitor the running configuration of the FPGA. Following our presentation and analysis of the security protocols, we propose an example of implementation with the CCM (Counter with CBC-MAC) authenticated encryption standard. We also evaluate the impact of our architecture on the configuration time for different FPGA devices. Benoît Badrignans, David Champagne, Reouven Elbaz, Catherine H. Gebotys, Lionel Torres |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2009 | Exploration of power reduction and performance enhancement in LEON3 processor with ESL reprogrammable eFPGA in processor pipeline and as a co-processorabstractWe will explore how processing power of LEON3 processor can be enhanced by connecting small commercially available embedded FPGA (eFPGA) IP with the processor. We will analyze integration of eFPGA with LEON3 in two ways, inside the processor pipeline and as a co-processor. The enhanced processing power helps to reduce dynamic power consumption by Dynamic Frequency Scaling. More computational power at lower frequency helps fabrication of chip in LP (Low Power) process compared to GP (General Purpose) which helps to significantly reduce Static Power which has become a very crucial issue at and beyond 90 nm technologies. Use of reconfigurable accelerator raises the question of its programming complexity, HW/SW partitioning and silicon overhead. We will present that silicon overhead of eFPGA is small compared to the benefits which can be obtained with it. We will present a profiling tool which we created for our experiments. To analyze the issue of programming complexity we have explored state of the art Catapulttrade ESL tool of Mentor Graphicsreg. Syed Zahid Ahmed, Julien Eydoux, Laurent Rouge, Jean-Baptiste Cuelle, Gilles Sassatelli, Lionel Torres |
DATE | 6 |
| 2009 | Evaluation on FPGA of triple rail logic robustness against DPA and DEMAabstractSide channel attacks are known to be efficient techniques to retrieve secret data. In this context, this paper concerns the evaluation of the robustness of triple rail logic against power and electromagnetic analyses on FPGA devices. More precisely, it aims at demonstrating that the basic concepts behind triple rail logic are valid and may provide interesting design guidelines to get DPA resistant circuits which are also more robust against DEMA. Victor Lomné, Philippe Maurine, Lionel Torres, Michel Robert, Rafael Soares, Ney Laert Vilar Calazans |
DATE | 3 |
| 2009 | Dynamic and distributed frequency assignment for energy and latency constrained MP-SoCabstractIn this paper we present an adaptive technique to locally adjust the frequency of processing elements on MP-SoC. The proposed method, based on game theory, optimizes the system while fulfilling dynamic constraints. A telecom test-case has been used to demonstrate the effectiveness of our technique. For the evaluated scenario, the proposed technique has obtained up to 20% of latency gain and 38% of energy gain. Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
DATE | 5 |
| 2009 | Enhancing Electromagnetic Attacks Using Spectral Coherence Based Cartography
Amine Dehbaoui, Victor Lomné, Philippe Maurine, Lionel Torres, Michel Robert |
VLSI-SoC | 4 |
| 2008 | Hierarchical Code Correction and Reliability Management in Embedded nor Flash MemoriesabstractThe framework of this article lies in the dynamic management of the reliability in NOR embedded Flash memories (eFlash). The main objective is to build a new reliability management scheme and to predict its efficiency to improve the eFlash reliability using error correction code and redundancy. The originality of the proposed approach relies on the use of a dedicated error correcting code well suited to NOR flash memories operational conditions. This code, named hierarchical code, improves the correction capabilities with a minimal impact on performance and area. The proposed solution furthermore enables selecting different built-in self strategies allowing to tune reliability strategies to the targeted application domain. Benoît Godard, Jean Michel Daga, Lionel Torres, Gilles Sassatelli |
ETS | 3 |
| 2008 | Secure FPGA configuration architecture preventing system downgradeabstractIn the context of FPGAs, system downgrade consists in preventing the update of the hardware configuration or in replaying an old bitstream. The objective can be to preclude a system designer from fixing security vulnerabilities in a design. Such an attack can be performed over a network when the FPGA-based system is remotely updated or on the bus between the configuration memory and the FPGA chip at power-up. Several security schemes providing encryption and integrity checking of the bitstream have been proposed in the literature. However, as we show in this paper, they do not detect the replay of old FPGA configurations; hence they provide adversaries with the opportunity to downgrade the system. We thus propose a new architecture that, in addition to ensuring bitstream confidentiality and integrity, precludes replay of old bitstreams. We show that the hardware cost of this architecture is negligible. Benoît Badrignans, Reouven Elbaz, Lionel Torres |
FPL | 3 |
| 2008 | A non-volatile run-time FPGA using thermally assisted switching MRAMSabstractThis paper describes the integration of a thermally assisted switching magnetic random access memory (TAS-MRAM) in FPGA design. The non-volatility of the latter is achieved through the use of magnetic tunneling junctions (MTJ) in the MRAM cell. A thermally assisted switching scheme is used to write data in the MTJ device, which helps to reduce power consumption during write operation in comparison to the writing scheme in classical MTJ device. Plus, the non-volatility of such a design should reduce both power consumption and configuration time required at each power up of the circuit in comparison to classical SRAM based FPGAs. A real time reconfigurable (RTR) micro-FPGA using TAS-MRAM allows dynamic reconfiguration mechanisms, while featuring simple design architecture. Yoann Guillemenet, Lionel Torres, Gilles Sassatelli, Nicolas Bruchon, Ilham Hassoune |
FPL | 2 |
| 2008 | Convergence analysis of run-time distributed optimization on adaptive systems using game theoryabstractWe consider multiprocessor system-on-chip (MP-SoC) integrating several processing elements (PE). These architectures require distributed and scalable control techniques for run-time optimization of applicative parameters. Our approach is to use the game theory as an optimization model to solve the trade-off issues at run-time. We applied it to the distributed dynamic voltage frequency scaling (DVFS) management, adjusting at run-time the frequency set of each PE based on the synchronization between tasks of the application graph and the PE temperature profile. Results show that the analyzed algorithm converges to a solution in about 94% of the cases and in less than 40 calculation cycles for a 100-processor MP-SoC. It reaches an average optimization of 89% compared to an off-line centralized reference but about 140 times faster when simulating. Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
FPL | 5 |
| 2008 | Bio-inspiration helps computers: A new machineabstractThe past decades have witnessed tremendous research efforts devoted to parallel architectures and programming models for natively computing in space. This resulted in systems which comprise a number of processing units ranging from compact Boolean function generators (FPGAs look-up-tables) to full-fledged microprocessors (MPSoCs). It is often stated in the literature of both areas that performance and/or scalability remain limited by the partial knowledge available at the time the platform is programmed [1] which pushed towards researching techniques granting a certain degree of run-time flexibility to these platforms (partial/ run-time reconfiguration for FPGAs, task migration/load balancing for multiprocessors). This paper presents a bio-inspired machine model which aims at addressing architecture scalability and self-adaptability. The architecture and the programming model are intended to be scalable. The link between the both is based on fully decentralized mechanisms allowing the scalability of the machine and its self-adaptability. An implementation of the proposed bio-inspired machine model has been developed and validated. The preliminary results prove the feasibility and the interest of the approach. Nicolas Saint-Jean, Gilles Sassatelli, Pascal Benoit, Lionel Torres, Michel Robert |
FPL | 4 |
| 2007 | TEC-Tree: A Low-Cost, Parallelizable Tree for Efficient Defense Against Memory Replay Attacks
Reouven Elbaz, David Champagne, Ruby B. Lee, Lionel Torres, Gilles Sassatelli, Pierre Guillemin |
CHES | 4 |
| 2007 | Evaluation of design for reliability techniques in embedded flash memories
Benoît Godard, Jean Michel Daga, Lionel Torres, Gilles Sassatelli |
DATE | 3 |
| 2007 | Run-time mapping and communication strategies for Homogeneous NoC-Based MPSoCsabstractMultiprocessor systems-on-chip are becoming increasingly popular in embedded systems for the high degree of performance and flexibility they permit. While most MPSoCs are today highly heterogeneous for better fitting the target applications, homogeneous systems may become in a near future a viable alternative bringing other benefits such as run-time load balancing, high performance and low power consumption. The work presented in this paper relies on a homogeneous NoC-based MPSoC framework we developed which allows us to conduct cycle-accurate evaluations of 2 different techniques: proactive and reactive communications. Gilles Sassatelli, Nicolas Saint-Jean, Pascal Benoit, Lionel Torres, Michel Robert, Cristiane R. Woszezenki, Ismael Grehs, Fernando Gehm Moraes |
FCCM | 4 |
| 2007 | A Cryptographic Coarse Grain Reconfigurable Architecture Robust Against DPAabstractThis work addresses the problem of information leakage of cryptographic devices, by using the reconfiguration technique allied to an RNS based arithmetic. The information leaked by circuits, like power consumption, electromagnetic emissions and time to compute may be used to find cryptographic secrets. The results issue of prototyping shows that our coarse grained reconfigurable architecture is robust against power analysis attacks. Daniel Mesquita, Benoît Badrignans, Lionel Torres, Gilles Sassatelli, Michel Robert, Fernando Gehm Moraes |
IPDPS | 3 |
| 2006 | A parallelized way to provide data encryption and integrity checking on a processor-memory busabstractInternational audience Reouven Elbaz, Lionel Torres, Gilles Sassatelli, Pierre Guillemin, Michel Bardouillet, Albert Martinez |
DAC | 2 |
| 2006 | Magnetic tunnelling junction based FPGAabstractThe aim of this paper is to propose a real time reconfigurable (RTR) micro-FPGA using new non volatile memory. Magnetic tunneling junctions (MTJ) used in Magnetic random access memories (MRAM) are compatible with classical CMOS processes. Moreover remanent property of such a memory could limit configuration time and power consumption required at each power up of the die. Nevertheless, each configuration memory point has to be readable independently from each other, that is why the approach is different from the classical memory array one. Nicolas Bruchon, Lionel Torres, Gilles Sassatelli, Gaston Cambon |
FPGA | 2 |
| 2006 | A Leak Resistant Architecture Against Side Channel AttacksabstractHardware implementations of cryptographic algorithms may leak some information that can be used to recover cryptographic keys. This work combines reconfigurable techniques with the recently proposed leak resistant arithmetic (LRA) to thwart some side channel attacks (SCA). The introduced architecture outcomes the performance of classical implementation of modular multiplication, for key size exceeding 2048 bits, with a reasonable extra area overhead. Nevertheless, this is not a drawback, but a cost, since the main issue of the proposed architecture is the improved robustness in terms of security. Daniel Mesquita, Benoît Badrignans, Lionel Torres, Gilles Sassatelli, Michel Robert, Jean-Claude Bajard, Fernando Gehm Moraes |
FPL | 3 |
| 2006 | Securing embedded programmable gate arrays in secure circuitsabstractThe purpose of this article is to propose a survey of possible approaches for implementing embedded reconfigurable gate arrays into secure circuits. A standard secure interfacing architecture is proposed and motivations justifying such an approach are discussed. This paper also lists all features offered by FPGA vendors (field programmable gate array) aiming at securing those circuits according to different concerns. This article emphasizes on configuration memory programming which is probably the weakest point of using programmable devices on a secure context. Nicolas Valette, Lionel Torres, Gilles Sassatelli, Frédéric Bancel |
IPDPS | 2 |
| 2005 | Hardware Engines for Bus Encryption: A Survey of Existing TechniquesabstractThe widening spectrum of applications and services provided by portable and embedded devices brings a new dimension of concerns in security. Most of those embedded systems (pay-TV, PDAs, mobile phones, etc.) make use of external memory. As a result, the main problem is that data and instructions are constantly exchanged between memory (RAM) and CPU in clear form on the bus. This memory may contain confidential data like commercial software or private contents, which either the end-user or the content provider is willing to protect. The paper describes the problem of processor-memory bus communications in this regard and the existing techniques applied to secure the communication channel through encryption. Performance overheads implied by those solutions are discussed extensively. Reouven Elbaz, Lionel Torres, Gilles Sassatelli, Pierre Guillemin, C. Anguille, Michel Bardouillet, Christian Buatois, Jean-Baptiste Rigaud |
DATE | 2 |
| 2005 | Dynamic hardware multiplexing for coarse grain reconfigurable architecturesabstractWhen designing a SoC, matching the required performances both in terms of processing power and power consumption tends to become more and more challenging. Moreover, since the range of targeted applications for every single product is widening rapidly, employing reconfigurable accelerators makes more and more sense to this purpose. Coarse grain reconfigurable architectures bring an alternative providing interesting performances / flexibility trade-offs over traditional approaches. This work presents an original method allowing to efficiently exploiting dynamically parallelism at both loop-level and task-level, which remains rarely used. This method called DHM (Dynamic Hardware Multiplexing) is based upon the use of a hardwired controller dedicated to the dynamical unroll of loops or scheduling of tasks. This work shows that significant performance improvements can be achieved through combining both intra and inter-task parallelism. Principles and validations are exposed through a case study on a coarse grain reconfigurable architecture, called Systolic Ring. The exposed method could be applied to any coarse and fine grain architectures. Pascal Benoit, Lionel Torres, Gilles Sassatelli, Michel Robert, Gaston Cambon |
FPGA | 2 |
| 2005 | Run-Time Scheduling for Random Multi-Tasking in Reconfigurable CoprocessorsabstractThe authors addressed the multi-tasking issue for reconfigurable coprocessors in random application contexts. A scheduling algorithm was proposed to handle simultaneously a set of random tasks and able to maximize the resource usage even when the task-load is low. For this, processes are considered as relocatable: a simple transformation scheme is applied by a configuration controller to the initial configuration in order to relocate or duplicate the task when necessary. In this paper, the proposed method is implemented on a coarse grain reconfigurable architecture with 8 and 32 processing elements. A large amount of random scenario have been simulated and the statistical results presented here clearly show real advantages of the proposed method, but also some limitations drawing the line of future works. Pascal Benoit, Jürgen Becker 0001, Michel Robert, Lionel Torres, Gilles Sassatelli, Gaston Cambon |
FPL | 4 |
| 2005 | Magnetic remanent memory structures for dynamically reconfigurable fine grain FPGAabstractEmergent technologies such as magnetic tunneling junction (MTJ), used in MRAM design are compatible with CMOS conventional processes and can be used in configurable circuits. This type of memory seems to be interesting for programmable applications in order to limit configuration time and power consumption required at each power up of the device. FPGA configuration memory is distributed all over the device and each point has to be readable independently from each other, that is why the approach is different from the classical memory array one. In this paper a first FPGA architecture based on MTJ-SRAM cells is described. Nicolas Bruchon, Gaston Cambon, Lionel Torres, Gilles Sassatelli |
FPL | 3 |
| 2005 | Current Mask Generation: an Analog Circuit to Thwart DPA Attacks
Daniel Mesquita, Jean-Denis Techer, Lionel Torres, Michel Robert, Guy Cathébras, Gilles Sassatelli, Fernando Gehm Moraes |
VLSI-SoC | 3 |
| 2003 | A Novel Approach for Architectural Model Characterization. An Example through the Systolic Ring
Pascal Benoit, Gilles Sassatelli, Lionel Torres, Michel Robert, Gaston Cambon, Didier Demigny |
FPL | 3 |
| 2003 | Are coarse grain reconfigurable architectures suitable for cryptography?
Daniel Mesquita, Lionel Torres, Fernando Gehm Moraes, Gilles Sassatelli, Michel Robert |
VLSI-SOC | 2 |
| 2002 | Highly Scalable Dynamically Reconfigurable Systolic Ring-Architecture for DSP ApplicationsabstractNew parallel execution based machine paradigms must be considered. Thanks to their high level of flexibility structurally programmable architectures are potentially interesting candidates to overcome classical CPUs limitations. Based on a parallel execution model, we present in this paper a new dynamically reconfigurable architecture, dedicated to data oriented applications acceleration. Principles, realizations and comparative results will be exposed for some classical applications, targeted on different architectures. Gilles Sassatelli, Lionel Torres, Pascal Benoit, Thierry Gil, Camille Diou, Gaston Cambon, Jérôme Galy |
DATE | 2 |
| 2001 | An embedded core for the 2D wavelet transformabstractWe present a new architecture for the wavelet transform based on the lifting scheme. We present the advantages of the lifting scheme compared to the filter-bank method and propose an architecture that allows a high processing rate and has a very small memory requirement. A new way of managing the transposition buffer using FIFO and delay lines allows the 2D decomposition of an image in real time on-the-fly. The proposed architecture is therefore very compact and is adapted for an implementation as a macrocell inside a system-on-a-chip. Camille Diou, Lionel Torres, Michel Robert |
ETFA (2) | 2 |
| 2001 | The Systolic Ring: A Dynamically Reconfigurable Architecture for Embedded Systems
Gilles Sassatelli, Lionel Torres, Jérôme Galy, Gaston Cambon, Camille Diou |
FPL | 2 |
| 2000 | A Wavelet Core for Video ProcessingabstractThe wavelet transform appears to be an efficient tool for image processing. However, the pyramid algorithm remains silicon area costly, essentially because of its memory needs, and depending on the size of filters used. This paper proposes a new implementation of the wavelet transform using the lifting scheme. Using an improved method of managing the data flow and the RAM, this method proposes many improvements such as in-place calculation, small memory needs, and easy inverse transform. Camille Diou, Lionel Torres, Michel Robert |
ICIP | 2 |