Fabien Clermidy

dblp:80/4541 · DBLP profile ↗
← Back
41ranked-venue papers
8as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 28% Integrated circuit design · 25% Hardware accelerators and domain-specific architectures · 22%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
many-core accelerator
0.112012
Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012
Parallel and multicore computing › many-core systems
many-core computing
0.112012
Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012
Integrated circuit design
digital circuit design
0.112011
Can we go towards true 3-D architectures? · DAC 2011
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic
0.112011
Can we go towards true 3-D architectures? · DAC 2011
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL
0.012012
Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012
Parallel and multicore computing
parallel programming models
0.012012
Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012
Integrated circuit design
3d integration
0.012011
Can we go towards true 3-D architectures? · DAC 2011

Methods — techniques the papers use, named apart from their topics

OpenCV · 0.1OpenCL · 0.1nanowire-based vertical transistor · 0.1
YearPublicationVenuePosition
2025 Dose Effects and Mitigation in 28nm FD-SOI Advanced Interface Bus Chiplet Interconnect
abstract
Chiplet technology for 2.5-D/3-D integration is being rapidly adopted, including for System-on-Chips used in mission critical applications. We present a Total Ionizing Dose Effect study of a prototype Advanced Interface Bus (AIB) die-to-die interface in 28nm FD-SOI technology using pulsed x-rays and in-situ monitoring of the dose. Degradation of the maximum working frequency of the interface was observed due to the deposited dose. Using either the core voltage or the body bias voltage, we demonstrate that it is possible to partially recover this frequency loss. In space applications, this compensation, combined with in-situ dose monitoring, can be used to extend the life of 2.5-D/3-D circuits using high-speed die-to-die interfaces.
Antoine Rouget, Fady Abouzeid, Adrian Evans, Victor Malherbe, Aleksandra Chumakova, Philippe Roche, Fabien Clermidy
IOLTS7
2019 Advanced 3D Technologies and Architectures for 3D Smart Image Sensors
abstract
Image Sensors will get more and more pervasive into their environment. In the context of Automotive and IoT, low cost image sensors, with high quality pixels, will embed more and more smart functions, such as the regular low level image processing but also object recognition, movement detection, light detection, etc. 3D technology is a key enabler technology to integrate into a single device the pixel layer and associated acquisition layer, but also the smart computing features and the required amount of memory to process all the acquired data. More computing and memory within the 3D Smart Image Sensors will bring new features and reduce the overall system power consumption. Advanced 3D technology with ultra-fine pitch vertical interconnect density will pave the way towards new architectures for 3D Smart Image Sensors, allowing local vertical communication between pixels, and the associated computing and memory structures. The presentation will give an overview of recent 3D technologies solutions, such as Hybrid Bonding technology and the Monolithic 3D CoolCube™ technology, with respective 3D interconnect pitch in the order of 1 μm and l00nm. Recent 3D Image Sensors will be presented, showing the capability of 3D technology to implement fine grain pixel acquisition and processing with ultra-high speed image acquisition and tile-based processing. As further perspectives, multi-layer 3D image sensor based on events and spiking will reduce power consumption with new detection and learning processing capabilities.
Pascal Vivet, Gilles Sicard, Laurent Millet, Stéphane Chevobbe, Karim Ben Chehida, Luis Angel Cubero, Monte Alegre, Maxence Bouvier, Alexandre Valentian, Maria Lepecq, Thomas Dombek, Olivier Bichler, Sébastien Thuries, Didier Lattard, Séverine Cheramy, Perrine Batude, Fabien Clermidy
DATE17
2016 MCAPI-compliant Hardware Buffer Manager mechanism to support communication in multi-core architectures
Thiago R. da Rosa, Thomas Mesquida, Romain Lemaire, Fabien Clermidy
DATE4
2015 A comprehensive study of monolithic 3D cell on cell design using commercial 2D tool
Olivier Billoint, Hossam Sarhan, Iyad Rayane, Maud Vinet, Perrine Batude, Claire Fenouillet-Béranger, Olivier Rozeau, Gerald Cibrario, Fabien Deprat, A. Fustier, Julien Michallet, Olivier Faynot, Ogun Turkyilmaz, Jean-Frédéric Christmann, Sébastien Thuries, Fabien Clermidy
DATE16
2015 Fine-grain DVFS and AVFS techniques for complex SoC design: An overview of architectural solutions through technology nodes
abstract
In this paper we propose to give an overview of fine-grain design techniques we demontrated past years in our lab for power reduction in complex SoCs. Those works are based on Globally Asynchronous and Locally Synchronous systems in which each IP is an independent voltage and frequency domain. After having proposed some simple DFS architectures based on GALS architectures in 130nm technology, we extended our works to fine-grain Dynamic Voltage and Frequency Scaling architectures to reduce dynamic and static power reduction at 65 nm node. Furthermore, considering 32 nm deep submicron technologies, we demonstrated an Adaptive Voltage and Frequency architecture to compensate for in-die PVT variations. Area overhead and power reduction results are discussed all along the paper.
Edith Beigné, Fabien Clermidy, Didier Lattard, Ivan Miro Panades, Yvain Thonnart, Pascal Vivet
ISCAS2
2015 Emerging resistive memories for low power embedded applications and neuromorphic systems
abstract
In this work, we will focus on the role that new nonvolatile resistive memory technologies can play in emerging fields of application, such as non-volatile logic circuits or neuromorphic circuits, to save energy and increase performance. Concerning the introduction of non-volatile functionalities at the logic level, we will demonstrate hybrid CMOS logic plus ReRAM (specifically CBRAM and OXRAM) circuits for ultra low power FPGA and fixed-logic IC design, as Non Volatile Flip-Flops. Concerning neuromorphic circuits, we will focus on the emulation of synaptic plasticity effects with resistive memory synapses. We will present large-scale energy efficient neuromorphic systems based on ReRAM as stochastic-binary synapses. Prototype applications such as complex visual- and auditory-pattern extraction will be also discussed using feedforward spiking neural networks.
Barbara De Salvo, Elisa Vianello, Olivier Thomas, Fabien Clermidy, Olivier Bichler, Christian Gamrat, Luca Perniola
ISCAS4
2015 From 2D to Monolithic 3D: Design Possibilities, Expectations and Challenges
abstract
Design of conventional 2D integrated circuits is becoming more and more challenging as we strive to keep on following Moore's law. Cost, thermal behavior, multiple patterning, increasing number of design rules, transistor characteristics, variability and back end properties coupled with a constant need for a higher integration of functions / peripherals are creating an increasingly complex equation to solve for designers. Moving to the next node and taking advantage of the technology are now far from being straightforward as time to market has never been so short for industry. In order to overcome or at least postpone the time when we'll have to face the "next node migration constraints", a possible solution could be staying at the same node and go 3D with possible benefits such as wire length reduction, power savings and increased operating frequency. Since more than ten years now, interconnect technologies like Through Silicon Via (TSV), High Density (HD)-TSV and Copper to Copper (Cu-Cu) have arisen to take advantage of this possible 3-dimensional physical implementation with proofs of concept [1] or more recently industrial products [2]. Main drawback of these technologies is that they are not shrinking at the same speed as transistors are, making them somehow power hungry; moreover the more they will shrink, the more precision will be needed for chip to chip alignment. To reach the highest possible standard cell and tier to tier interconnect densities required for cost-effective chips, 3D sequential integration process [3][4][5] (also known as Monolithic 3D or CoolCubeTM) is currently developed with main features being sequential fabrication of MOS layers and correlation of tier to tier interconnect size with process node allowing fine-grain 3D partitioning of designs. These particularities make it a durable opportunity to slow down next node design migration while still improving integration. To fully benefit from CoolCubeTM technology, a whole new way of designing circuits, from synthesis to place and route, will be required as some new challenges will arise. The point of this presentation is to show the possible use and limitations of the aforementioned technologies with a focus on Monolithic 3D and to give some insights about market expectations, challenges and available design techniques.
Olivier Billoint, Hossam Sarhan, Iyad Rayane, Maud Vinet, Perrine Batude, Claire Fenouillet-Béranger, Olivier Rozeau, Gerald Cibrario, Fabien Deprat, Ogun Turkyilmaz, Sébastien Thuries, Fabien Clermidy
ISPD12
2015 Compact interconnect approach for networks of neural cliques using 3D technology
abstract
Thanks to their brain-like properties, neural networks outperform traditional algorithms in certain group of applications. However, since they are wire-dominated systems, their hardware implementation poses numerous challenges as high latency and energy consumption. The recent technological improvements allow for stacking few dies one on another and designing 3D electronic circuits. This creates opportunities for 3D efficient implementations of neural networks targeting high-performance applications. This work explores the gains of 3D technology for neural networks relying on neural cliques. A general study shows up to 55% reduction in terms of total interconnect length and interconnect power consumption, and 74% reduction of the maximal interconnect delay. The proposed approach is validated with a power management applicative test-case. We demonstrate that, in this scenario, the 3D architecture reduces interconnect length and power by 35% and the maximal delay by 57%, compared to 2D.
Bartosz Boguslawski, Hossam Sarhan, Frédéric Heitzmann, Fabrice Seguin, Sébastien Thuries, Olivier Billoint, Fabien Clermidy
VLSI-SoC7
2014 Advanced technologies for brain-inspired computing
abstract
This paper aims at presenting how new technologies can overcome classical implementation issues of Neural Networks. Resistive memories such as Phase Change Memories and Conductive-Bridge RAM can be used for obtaining low-area synapses thanks to programmable resistance also called Memristors. Similarly, the high capacitance of Through Silicon Vias can be used to greatly improve analog neurons and reduce their area. The very same devices can also be used for improving connectivity of Neural Networks as demonstrated by an application. Finally, some perspectives are given on the usage of 3D monolithic integration for better exploiting the third dimension and thus obtaining systems closer to the brain.
Fabien Clermidy, Rodolphe Héliot, Alexandre Valentian, Christian Gamrat, Olivier Bichler, Marc Duranton, Bilel Belhadj, Olivier Temam
ASP-DAC1
2014 A scalable custom simulation machine for the Bayesian Confidence Propagation Neural Network model of the brain
abstract
A multi-chip custom digital super-computer called eBrain for simulating Bayesian Confidence Propagation Neural Network (BCPNN) model of the human brain has been proposed. It uses Hybrid Memory Cube (HMC), the 3D stacked DRAM memories for storing synaptic weights that are integrated with a custom designed logic chip that implements the BCPNN model. In 22nm node, eBrain executes BCPNN in real time with 740 TFlops/s while accessing 30 TBs synaptic weights with a bandwidth of 112 TBs/s while consuming less than 6 kWs power for the typical case. This efficiency is three orders better than general purpose supercomputers in the same technology node.
Nasim Farahini, Ahmed Hemani, Anders Lansner, Fabien Clermidy, Christer Svensson
ASP-DAC4
2014 3DCoB: A new design approach for Monolithic 3D Integrated circuits
abstract
3D Monolithic Integration (3DMI) technology provides very high dense vertical interconnects with low parasitics. Previous 3DMI design approaches provide either cell-on-cell or transistor-on-transistor integration. In this paper we present 3D Cell-on-Buffer (3DCoB) as a novel design approach for 3DMI. Our approach provides a fully compatible sign-off physical implementation flow with the conventional 2D tools. We implement our approach on a set of benchmark circuits using 28nm-FDSOI technology. The sign-off performance results show 35% improvement compared to the same 2D design.
Hossam Sarhan, Sébastien Thuries, Olivier Billoint, Fabien Clermidy
ASP-DAC4
2014 Resistive memories: Which applications?
abstract
Recent announcement of 16Gbits Resistive memory from Sony shows the trend to quickly adopt resistive memories as an alternative to DRAM. However, using ReRAM for embedded computing is still a futuristic goal. This paper approaches two applications based on ReRAM-devices for gaining area, performance or power consumption. The first application is FPGA, one of the first architecture that can benefit the most from ReRAM integration to reduce footprint and save energy. The second application relates to ultra-low-power systems and the way to obtain an instantaneous “freeze” mode in devices for Internet of Things.
Fabien Clermidy, Natalija Jovanovic, Santhosh Onkaraiah, Houcine Oucheikh, Olivier Thomas, Ogun Turkyilmaz, Elisa Vianello, Jean-Michel Portal, Marc Bocquet
DATE1
2014 3D FPGA using high-density interconnect Monolithic Integration
abstract
New 3D technology, called “Monolithic Integration”, offers very dense 3D interconnect capabilities. In this paper, we propose a 3D FPGA architecture with logic-on-memory approach based on this technology. The routing and computation blocks are splitted into two layers where the logic is placed on the top and memory on the bottom. Using extracted values from layout in 14nm FDSOI technology, typical benchmark circuits are evaluated in the VPR5 toolflow. The results show an area reduction of 55% compared to the 2D FPGA. More importantly, due to the lowered routing congestion, the EDP of the 3D FPGA is improved by 47%.
Ogun Turkyilmaz, Gerald Cibrario, Olivier Rozeau, Perrine Batude, Fabien Clermidy
DATE5
2014 RRAM-based FPGA for "Normally Off, Instantly On" applications
Ogun Turkyilmaz, Santhosh Onkaraiah, Marina Reyboz, Fabien Clermidy, Hraziia, Costin Anghel, Jean-Michel Portal, Marc Bocquet
J. Parallel Distributed Comput.4
2013 A dynamic stream link for efficient data flow control in NoC based heterogeneous MPSoC
abstract
As Systems-on-Chip size increase, the communication costs become critical and Networks-on-Chip (NoC) bring innovative solutions. Efficient stream-based protocols over NoC have been widely studied to address dataflow communications. They are usually controlled by a set of static parameters. However, new applications, such as high-resolution video decoders, present more data-dependent behaviors forcing communication protocols to support higher dynamicity. For this purpose, we present in this paper dynamic stream links for stream-based end-to-end NoC communications by introducing two link protocols, both independent of the transfer size, allowing to improve the hardware/software control flexibility. The proposed protocols have been modeled in a MPSoC virtual platform and the hardware cost evaluated. Based on simulations, we provide guidelines to exploit these protocols according to application needs.
Claude Helmstetter, Sylvain Basset, Romain Lemaire, Fabien Clermidy, Pascal Vivet, Michel Langevin, Chuck Pilkington, Pierre G. Paulin, Didier Fuin
ASP-DAC4
2013 3D stacking for multi-core architectures: From WIDEIO to distributed caches
abstract
3D stacking has been viewed as a breakthrough solution for increasing performance in multi-core architectures. The hope is to solve some of the main issues in current multi-core architectures: external memory pressure and latency; I/O bottleneck; communication power consumption. In this paper, some advances of this field of research are shown, starting with a WIDEIO experience on a real chip for solving DRAM accesses issue. The integration of a 512 bit-width bus is demonstrated in a Network-on-Chip (NoC) multi-core framework and the resulting performance based on a 65nm prototype with 10μm diameter Through Silicon Vias (TSV). The potentiality of 3D scaling thanks to 3D asynchronous Network-on-Chip implementation is then shown. Finally, an innovative 3D stacked distributed cache strategy aimed at lowering memory latency and external memory bandwidth requirements is presented. This new memory partitioning demonstrates the efficiency of 3D stacking to rethink architectures for addressing multi-core scaling challenges.
Fabien Clermidy, Denis Dutoit, Eric Guthmuller, Ivan Miro Panades, Pascal Vivet
ISCAS1
2013 A hybrid CBRAM/CMOS Look-Up-Table structure for improving performance efficiency of Field-Programmable-Gate-Array
abstract
At most advanced technology nodes, Field Programmable Gate Arrays (FPGA) present great advantages compared to more conventional processor architectures; their natural regularity, modularity and inherent reliability due to duplicated identical tiles provide a solution to overcome new technologies with increasing variability. However, FPGA market is still limited by power efficiency issue, due to two coordinated factors like interconnection-dominated design and large usage of memories, computation being performed thanks to Look-Up-Table (LUT). In this paper, we propose a solution to improve the performance and reduce the power consumption of LUT in FPGA using CBRAM-based structures. Our proposed design shows significant improvement compared to the traditional SRAM-based FPGA in: critical delay is reduced by ~23% due to compact structure (1T-2R) and power gain by reduction in static power consumption by ~18%.
Santhosh Onkaraiah, Ogun Turkyilmaz, Marina Reyboz, Fabien Clermidy, Elisa Vianello, Jean-Michel Portal, Christophe Muller
ISCAS4
2013 Self-checking ripple-carry adder with Ambipolar Silicon NanoWire FET
abstract
For the rapid adoption of new and aggressive technologies such as ambipolar Silicon NanoWire (SiNW), addressing fault-tolerance is necessary. Traditionally, transient fault detection implies large hardware overhead or performance decrease compared to permanent fault detection. In this paper, we focus on on-line testing and its application to ambipolar SiNW. We demonstrate on self-checking ripple-carry adder how ambipolar design style can help reduce the hardware overhead. When compared with equivalent CMOS process, ambipolar SiNW design shows a reduction in area of at least 56% (28%) with a decreased delay of 62% (6%) for Static (Transmission Gate) design style.
Ogun Turkyilmaz, Fabien Clermidy, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli
ISCAS2
2012 Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications
abstract
P2012 is an area- and power-efficient many-core computing accelerator based on multiple globally asynchronous, locally synchronous processor clusters. Each cluster features up to 16 processors with independent instruction streams sharing a multi-banked one-cycle access L1 data memory, a multi-channel DMA engine and specialized hardware for synchronization and aggressive power management. P2012 is 3D stacking ready and can be customized to achieve extreme area and energy efficiency by adding domain-specific HW IPs to the cluster. The first P2012 SoC prototype in 28nm CMOS will sample in Q3, featuring four 16-processor clusters, a 1MB L2 memory and delivering 80GOPS (with 32 bit single precision floating point support) in 18mm2 with 2W power consumption (worst-case). P2012 can run standard OpenCL™ and proprietary Native Programming Model SW components to achieve the highest level of control on application-to-resource mapping. A dedicated version of the OpenCV vision library is provided in the P2012 SW Development Kit to enable visual analytics acceleration. This paper will discuss preliminary performance measurements of common feature extraction and tracking algorithms, parallelized on P2012, versus sequential execution on ARM CPUs.
Diego Melpignano, Luca Benini, Eric Flamand, Bruno Jego, Thierry Lepley, Germain Haugou, Fabien Clermidy, Denis Dutoit
DAC7
2011 Can we go towards true 3-D architectures?
abstract
Thanks to recent technology advances, the exploration of the vertical dimension has been shown to be more than a dream for designers. Among those technologies, the vertical transistor has not been exploited yet. This paper describes a novel implementation of logic gates fully benefiting of nanowire-based vertical transistors embedded within the metal lines. The logic design in this technology is explored and its performance is evaluated. A comparison made on an equivalent technology node shows that our cells reduce area and delay by a factor of 31x and 2x respectively. Large reconfigurable logic circuits have been benchmarked showing an improvement of area and delay by 46% and 48% on average.
Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Paul-Henry Morel, Jean-Philippe Noël, Fabien Clermidy, Ian O'Connor
DAC5
2011 A low-power VLIW processor for 3GPP-LTE complex numbers processing
abstract
New generation of telecommunication applications requires highly efficient processing units to tackle with the increasing signal processing algorithmic complexity. They also need to be flexible for handling a large range of radio access technology with specifications moving very fast. As devices including telecommunication features are, per nature, mobile, the high level of flexibility must be achieved while preserving very low power consumption. In this paper, a high performance low-power application-specific processor is proposed for complex signal processing. Thanks to dedicated control architecture, this processor exhibits an average 81% utilization rate of its principal operator, a complex MAC for a 3GPP-LTE application. The main innovations are the use of a reconfigurable profile and instruction cache strategy to reduce power consumption. This leads to a 10× reduction of the control power consumption. As a result, an average 50 mW power consumption is measured after implementation in a low-power 65 nm technology while delivering 3.2 GOPS. Finally, a comparison with state-of-the-art low-power DSP shows at least 24 % gain.
Christian Bernard, Fabien Clermidy
DATE2
2011 3D Embedded multi-core: Some perspectives
abstract
3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range products. In this paper, we investigate three promising perspectives for short to medium terms adoption of such technology in high-end System-on-Chip built around multi-core architectures: the wide bus concept will help solving high bandwidth requirements with external memory. 3D Network-on-Chip is a promising solution for increased modularity and scalability. We show that an efficient implementation provides an available bandwidth outperforming classical interfaces. Finally, we put in perspective the active interposer concept which aims at simplifying and improving power, test and debug management.
Fabien Clermidy, Florian Darve, Denis Dutoit, Walid Lafi, Pascal Vivet
DATE1
2011 Sustainability through massively integrated computing: Are we ready to break the energy efficiency wall for single-chip platforms?
abstract
While traditional cluster computers are more constrained by power and cooling costs for solving extreme-scale (or exascale) problems, the continuing progress and integration levels in silicon technologies make possible complete end-user systems on a single chip. This massive level of integration makes modern multicore chips all pervasive in domains ranging from climate forecasting and astronomical data analysis, to consumer electronics, smart phones, and biological applications. Consequently, designing multicore chips for exascale computing while using the embedded systems design principles looks like a promising alternative to traditional cluster-based solutions. This paper aims to present an overview of new, far-reaching design methodologies and run-time optimization techniques that can help breaking the energy efficiency wall in massively integrated single-chip computing platforms.
Partha Pratim Pande, Fabien Clermidy, Diego Puschini, Imen Mansouri, Paul Bogdan, Radu Marculescu, Amlan Ganguly
DATE2
2011 A low complexity stopping criterion for reducing power consumption in turbo decoders
abstract
Turbo codes are proposed in most of the advanced digital communication standards, such as 3GPP-LTE. However, due to its computational complexity, the turbo decoder is one of the most power hungry blocks in digital baseband. To alleviate this issue, one way is to avoid surplus computing phases thanks to the early termination of the iterative decoding process. The use of stopping criteria is one of the most common algorithm level power reduction methods in literature. These methods always come with some hardware overhead. In this paper, a new trellis based stopping criterion is proposed. The novelty of this approach is the lower hardware overhead thanks to the use of trellis states as key parameter to stop the iterative process. Results are showing the importance of this added hardware in terms of method efficiency. Compared to state-of-the-art Log Likelihood Ratio (LLR) based techniques, proposed Low Complexity Trellis Based (LCTB) is demonstrating 23% less power consumption on average, for comparable performance level in terms of Bit Error Rate (BER) and Frame Error Rate (FER).
Pallavi Reddy, Fabien Clermidy, Amer Baghdadi, Michel Jézéquel
DATE2
2011 Dynamic Flow Reconfiguration Strategy to Avoid Communication Hot-Spots
abstract
Application-specific Network-on-Chip allows optimization for the interconnection to minimize its cost. When used with streaming applications, large flows of data can be predicted. However, these flows can be modified during the applications providing a dynamic flow graph. In that case, applying on off-line optimization leads to an over-sizing of the NoC. On the other hand, dynamic reconfiguration leads to unordered data deliveries with costly re-ordering units. In this paper, we propose a coarse grain dynamic reconfiguration which avoids data re-ordering requirement. We show that the proposed solution is efficient to deal with communication hot-spots, with a small area overhead, and can save up to 33% of latency.
Romain Prolonge, Fabien Clermidy, Leonel Tedesco, Fernando Gehm Moraes
DSD2
2011 Evaluation of a crossbar multiplexer in a lithography-based nanowire technology
abstract
Silicon Nanowire technology has been demonstrated to be a promising candidate to fabricate nanowire crossbars. The use of such devices in a real architectural as well as in a design environment is an ongoing research topic. In this paper, we investigate the use of a lithography-based industrial process for designing a 4-to-1 multiplexer in a crossbar circuit. We show that by considering the line parasitic, the crossbar demonstrates poor performance in a 65-nm technology, while the area and power savings are about 6× and 1.5× respectively vs. the CMOS implementation. However, extrapolation to the 9-nm node shows a 2× better performance and 67× area saving.
Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Fabien Clermidy, Ian O'Connor
ISCAS3
2011 Reconfiguration of a 3GPP-LTE telecommunication application on a 22-core NoC-based system-on-chip
abstract
The MAGALI chip is a 65nm digital baseband dedicated to advanced telecommunication applications. Based on a 15-router mesh Network-on-Chip - NoC, it embeds 22 Processing Elements - PE performing the different functions of a complex baseband in a programmable manner. The NoC framework is based on asynchronous routers which provide a complete Globally Asynchronous Locally Synchronous -- GALS framework. Thanks to this structure, each PE is a synchronous island with a programmable frequency. On the architectural side, the NoC supports fast reconfiguration thanks to a distributed scheme. In this demonstration, we propose to show the reconfiguration capabilities of the MAGALI chip on three modes of a 3GPP-LTE receiver part, by switching on-the-fly between these three modes. The resulting transmission quality with different levels of upcoming signal noises is shown.
Fabien Clermidy, Nicolas Cassiau, N. Coste, Denis Dutoit, M. Fantini, Dimitri Ktenas, Romain Lemaire, L. Stefanizzi
NOCS1
2011 3D NoC using through silicon Via: An asynchronous implementation
abstract
3D stacking is seen as one of the most interesting technologies for System-on-Chip (SoC) developments. However, 3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range of products. In this paper, we are investigating 3D Network-on-Chip has a promising solution for increased modularity and scalability. We show that an efficient implementation based on asynchronous logic provides an available bandwidth of 64GB/s for only 700 TSV, outperforming classical interfaces while simplifying the assembly process. We also point out the benefit in terms of power consumption for these new interfaces with a gain of 5 times compared to classical LPDDR2 interfaces.
Pascal Vivet, Denis Dutoit, Yvain Thonnart, Fabien Clermidy
VLSI-SoC4
2011 Matrix Nanodevice-Based Logic Architectures and Associated Functional Mapping Method
abstract
This article describes a novel computing architecture organization based on nanoscale logic cells. We propose the use of a cluster of matrix arrangements of cells. In order to interconnect such fine-grained logic cells within a matrix, conventional techniques are not suitable due to a large interconnect overhead. Therefore, we propose the use of static and incomplete interconnect topologies to create matrices of cells. We also propose a method to map functions onto such architectures. We then explore the main parameters of the structure (size of matrices and interconnect topologies) and their impact on the main performance metrics (packing efficiency, speed, and fault tolerance). A cluster packing method also allows the evaluation of the number of matrices used by complex functions and the fill factor for various matrix sizes. The analyses show that this approach is particularly suited for matrices of 16 cells interconnected by modified omega networks. We can conclude that this architecture could improve the scalability of traditional FPGAs by a factor of 8.5.
Pierre-Emmanuel Gaillardon, Fabien Clermidy, Ian O'Connor, Maimouna Amadou, Gabriela Nicolescu
ACM J. Emerg. Technol. Comput. Syst.2
2010 A fully-asynchronous low-power framework for GALS NoC integration
abstract
Requiring more bandwidth at reasonable power consumption, new communication infrastructures must provide adequate solutions to guarantee performance during physical integration. In this paper, we propose the design of a low-power asynchronous Network-on-Chip which is implemented in a bottom-up approach using optimized hard-macros. This architecture is fully testable and a new design flow is proposed to overcome CAD tools limitations regarding asynchronous logic. The proposed architecture has been successfully implemented in CMOS 65nm in a complete circuit. It achieves a 550Mflit/s throughput on silicon, and exhibits 86% power reduction compared to an equivalent synchronous NoC version.
Yvain Thonnart, Pascal Vivet, Fabien Clermidy
DATE3
2010 Phase-change-memory-based storage elements for configurable logic
abstract
Back-end-of-line non-volatile resistive memories like Phase Change Memories (PCMs) are promising to solve memory issues in different architectures. In this paper, we investigate the usage of PCM to build an elementary configuration memory node for reconfigurable logic, such as Field-Programmable Gate Arrays (FPGAs). We propose an elementary circuit realized by 2 resistive memories and 1 programming transistor able to store a configuration voltage. We investigate the proposed node in terms of area and write time and we assess its impact on complex circuits. We show that the elementary memory node yields an improvement in area and write time of 1.5x and 16x respectively vs. a regular Flash implementation. Implemented in FPGAs, the memory node yields a delay reduction up to 51%, thanks to the reduction of dimensions and low on-resistance of PCMs.
Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Marina Reyboz, Giovanni Beneventi, Fabien Clermidy, Luca Perniola, Ian O'Connor
FPT5
2010 Distributed Sequencing for Resource Sharing in Multi-applicative Heterogeneous NoC Platforms
abstract
In the context of heterogeneous NoC architectures for embedded systems, it is today mandatory to support multiple applications given the plurality of standards and usages. While static reconfiguration between applications has already been extensively studied, we propose a potential increase in hardware resource usage by enabling concurrent or overlapping applications on the top of a heterogeneous NoC platform. In this paper, we describe a distributed sequencing protocol allowing hardware resource sharing between several applications. This protocol ensures correct synchronization of the processing between hardware resources without the need of a global fine-grain scheduler on the system, thus alleviating the pressure on the run-time system. The proposed protocol has been integrated and validated in a NoC-based digital baseband for 4G SDR telecom applications, and was integrated on a manufactured chip on a STMicroelectronics CMOS 65 nm LP technology.
Yvain Thonnart, Romain Lemaire, Fabien Clermidy
NOCS3
2009 Dynamic and distributed frequency assignment for energy and latency constrained MP-SoC
abstract
In this paper we present an adaptive technique to locally adjust the frequency of processing elements on MP-SoC. The proposed method, based on game theory, optimizes the system while fulfilling dynamic constraints. A telecom test-case has been used to demonstrate the effectiveness of our technique. For the evaluated scenario, the proposed technique has obtained up to 20% of latency gain and 38% of energy gain.
Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres
DATE2
2009 An Open and Reconfigurable Platform for 4G Telecommunication: Concepts and Application
abstract
Advanced telecommunication applications require more and more flexibility. Static and run-time configuration mechanisms, as well as plug-in computing units are two solutions to solve this issue. In this paper, we propose an open architecture, dedicated to complex data-flow applications, fulfilling these requirements thanks to a distributed configuration scheme and a standard interface for data flow computing units. Based on this architecture, an open platform demonstrator has been designed. Simulation results on a 3GPP/LTE application are presented, showing a reconfiguration time overhead of only 2.6%.
Fabien Clermidy, Romain Lemaire, Xavier Popon, Dimitri Ktenas, Yvain Thonnart
DSD1
2009 Open Platform for Prototyping of Advanced Software Defined Radio and Cognitive Radio Techniques
abstract
This paper presents the ANR project IDROMel, which aims at developing reconfigurable SDR (software defined radio) and cognitive radio (CR) equipments. IDROMel is a 3 years project that started in 2005 and finishes in 2009. The main objective of IDROMel is to define, develop and validate a powerful SDR and CR platform combining very last technology progresses. The platform includes software parts (reconfigurable protocol stacks) and hardware parts (a base band board and a radio frequency front end, RF). Both parts are presented in this paper.
Dominique Nussbaum, Karim Khalfallah, Christophe Moy, Amor Nafkha, Pierre Leray, Julien Delorme, Jacques Palicot, Jérôme Martin, Fabien Clermidy, Bertrand Mercier, Renaud Pacalet
DSD9
2009 A Communication and configuration controller for NoC based reconfigurable data flow architecture
abstract
While network-on-chip aspects such as topologies, routing strategies or quality-of-service have been largely studied, the mapping of real applications on distributed NoC-based architecture is still an open issue. In this paper, we address this issue for complex reconfigurable data-flow applications. We introduce the concept of communication and configuration controller (CCC) which interacts both with the usual network interface and the IP core structure. The proposed CCC is a programmable template-based architecture, which provides solutions to manage reconfiguration flows, data synchronizations and global control signaling. An implementation of the CCC is presented, and its performances in a 65 nm technology are discussed through a concrete telecommunication application.
Fabien Clermidy, Romain Lemaire, Yvain Thonnart, Pascal Vivet
NOCS1
2009 Emerging Technologies and Nanoscale Computing Fabrics
abstract
6-8 July 2014
Ian O'Connor, Kotb Jabeur, Nataliya Yakymets, Renaud Daviot, David Navarro, Pierre-Emmanuel Gaillardon, Fabien Clermidy, Maimouna Amadou, Gabriela Nicolescu
VLSI-SoC8
2008 Convergence analysis of run-time distributed optimization on adaptive systems using game theory
abstract
We consider multiprocessor system-on-chip (MP-SoC) integrating several processing elements (PE). These architectures require distributed and scalable control techniques for run-time optimization of applicative parameters. Our approach is to use the game theory as an optimization model to solve the trade-off issues at run-time. We applied it to the distributed dynamic voltage frequency scaling (DVFS) management, adjusting at run-time the frequency set of each PE based on the synchronization between tasks of the application graph and the PE temperature profile. Results show that the analyzed algorithm converges to a solution in about 94% of the cases and in less than 40 calculation cycles for a 100-processor MP-SoC. It reaches an average optimization of 89% compared to an off-line centralized reference but about 140 times faster when simulating.
Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres
FPL2
2008 Dynamic Voltage and Frequency Scaling Architecture for Units Integration within a GALS NoC
Edith Beigné, Fabien Clermidy, Sylvain Miermont, Pascal Vivet
NOCS2
2008 Physical Implementation of the DSPIN Network-on-Chip in the FAUST Architecture
Ivan Miro Panades, Fabien Clermidy, Pascal Vivet, Alain Greiner
NOCS2
1999 A New Placement Algorithm Dedicated to Parallel Computers: Bases and Application
abstract
One way to improve reliability in parallel computers consists of adding supplementary processors and interconnections to the functional structure in order to replace faulty processors with respect to the network structure. This approach is named structural fault tolerance (SFT). Very integrated parallel computers are one way to implement a parallel structure. The material structure is then composed of many elementary blocks, such as ASICs or multi-chip modules (MCMs), each containing many processors. We show that former SFT methods fail in combining the different features, constraints and requirements of such structures. Thus, this paper introduces a new reconfiguration approach that is dedicated to very integrated parallel computers.
Fabien Clermidy, Thierry Collette, Michael Nicolaidis
PRDC1