EDBT 2026 Demo / reviewers in the wild / expert
Fabien Clermidy
dblp:80/4541
· DBLP profile ↗
41ranked-venue papers
8as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 13 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 28% Integrated circuit design · 25% Hardware accelerators and domain-specific architectures · 22% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
many-core accelerator |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing › many-core systems
many-core computing |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Integrated circuit design
digital circuit design |
0.1 | 1 | 2011 | Can we go towards true 3-D architectures? · DAC 2011 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable logic |
0.1 | 1 | 2011 | Can we go towards true 3-D architectures? · DAC 2011 |
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Integrated circuit design
3d integration |
0.0 | 1 | 2011 | Can we go towards true 3-D architectures? · DAC 2011 |
Methods — techniques the papers use, named apart from their topics
OpenCV · 0.1OpenCL · 0.1nanowire-based vertical transistor · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dose Effects and Mitigation in 28nm FD-SOI Advanced Interface Bus Chiplet InterconnectabstractChiplet technology for 2.5-D/3-D integration is being rapidly adopted, including for System-on-Chips used in mission critical applications. We present a Total Ionizing Dose Effect study of a prototype Advanced Interface Bus (AIB) die-to-die interface in 28nm FD-SOI technology using pulsed x-rays and in-situ monitoring of the dose. Degradation of the maximum working frequency of the interface was observed due to the deposited dose. Using either the core voltage or the body bias voltage, we demonstrate that it is possible to partially recover this frequency loss. In space applications, this compensation, combined with in-situ dose monitoring, can be used to extend the life of 2.5-D/3-D circuits using high-speed die-to-die interfaces. Antoine Rouget, Fady Abouzeid, Adrian Evans, Victor Malherbe, Aleksandra Chumakova, Philippe Roche, Fabien Clermidy |
IOLTS | 7 |
| 2019 | Advanced 3D Technologies and Architectures for 3D Smart Image SensorsabstractImage Sensors will get more and more pervasive into their environment. In the context of Automotive and IoT, low cost image sensors, with high quality pixels, will embed more and more smart functions, such as the regular low level image processing but also object recognition, movement detection, light detection, etc. 3D technology is a key enabler technology to integrate into a single device the pixel layer and associated acquisition layer, but also the smart computing features and the required amount of memory to process all the acquired data. More computing and memory within the 3D Smart Image Sensors will bring new features and reduce the overall system power consumption. Advanced 3D technology with ultra-fine pitch vertical interconnect density will pave the way towards new architectures for 3D Smart Image Sensors, allowing local vertical communication between pixels, and the associated computing and memory structures. The presentation will give an overview of recent 3D technologies solutions, such as Hybrid Bonding technology and the Monolithic 3D CoolCube™ technology, with respective 3D interconnect pitch in the order of 1 μm and l00nm. Recent 3D Image Sensors will be presented, showing the capability of 3D technology to implement fine grain pixel acquisition and processing with ultra-high speed image acquisition and tile-based processing. As further perspectives, multi-layer 3D image sensor based on events and spiking will reduce power consumption with new detection and learning processing capabilities. Pascal Vivet, Gilles Sicard, Laurent Millet, Stéphane Chevobbe, Karim Ben Chehida, Luis Angel Cubero, Monte Alegre, Maxence Bouvier, Alexandre Valentian, Maria Lepecq, Thomas Dombek, Olivier Bichler, Sébastien Thuries, Didier Lattard, Séverine Cheramy, Perrine Batude, Fabien Clermidy |
DATE | 17 |
| 2016 | MCAPI-compliant Hardware Buffer Manager mechanism to support communication in multi-core architectures
Thiago R. da Rosa, Thomas Mesquida, Romain Lemaire, Fabien Clermidy |
DATE | 4 |
| 2015 | A comprehensive study of monolithic 3D cell on cell design using commercial 2D tool
Olivier Billoint, Hossam Sarhan, Iyad Rayane, Maud Vinet, Perrine Batude, Claire Fenouillet-Béranger, Olivier Rozeau, Gerald Cibrario, Fabien Deprat, A. Fustier, Julien Michallet, Olivier Faynot, Ogun Turkyilmaz, Jean-Frédéric Christmann, Sébastien Thuries, Fabien Clermidy |
DATE | 16 |
| 2015 | Fine-grain DVFS and AVFS techniques for complex SoC design: An overview of architectural solutions through technology nodesabstractIn this paper we propose to give an overview of fine-grain design techniques we demontrated past years in our lab for power reduction in complex SoCs. Those works are based on Globally Asynchronous and Locally Synchronous systems in which each IP is an independent voltage and frequency domain. After having proposed some simple DFS architectures based on GALS architectures in 130nm technology, we extended our works to fine-grain Dynamic Voltage and Frequency Scaling architectures to reduce dynamic and static power reduction at 65 nm node. Furthermore, considering 32 nm deep submicron technologies, we demonstrated an Adaptive Voltage and Frequency architecture to compensate for in-die PVT variations. Area overhead and power reduction results are discussed all along the paper. Edith Beigné, Fabien Clermidy, Didier Lattard, Ivan Miro Panades, Yvain Thonnart, Pascal Vivet |
ISCAS | 2 |
| 2015 | Emerging resistive memories for low power embedded applications and neuromorphic systemsabstractIn this work, we will focus on the role that new nonvolatile resistive memory technologies can play in emerging fields of application, such as non-volatile logic circuits or neuromorphic circuits, to save energy and increase performance. Concerning the introduction of non-volatile functionalities at the logic level, we will demonstrate hybrid CMOS logic plus ReRAM (specifically CBRAM and OXRAM) circuits for ultra low power FPGA and fixed-logic IC design, as Non Volatile Flip-Flops. Concerning neuromorphic circuits, we will focus on the emulation of synaptic plasticity effects with resistive memory synapses. We will present large-scale energy efficient neuromorphic systems based on ReRAM as stochastic-binary synapses. Prototype applications such as complex visual- and auditory-pattern extraction will be also discussed using feedforward spiking neural networks. Barbara De Salvo, Elisa Vianello, Olivier Thomas, Fabien Clermidy, Olivier Bichler, Christian Gamrat, Luca Perniola |
ISCAS | 4 |
| 2015 | From 2D to Monolithic 3D: Design Possibilities, Expectations and ChallengesabstractDesign of conventional 2D integrated circuits is becoming more and more challenging as we strive to keep on following Moore's law. Cost, thermal behavior, multiple patterning, increasing number of design rules, transistor characteristics, variability and back end properties coupled with a constant need for a higher integration of functions / peripherals are creating an increasingly complex equation to solve for designers. Moving to the next node and taking advantage of the technology are now far from being straightforward as time to market has never been so short for industry. In order to overcome or at least postpone the time when we'll have to face the "next node migration constraints", a possible solution could be staying at the same node and go 3D with possible benefits such as wire length reduction, power savings and increased operating frequency. Since more than ten years now, interconnect technologies like Through Silicon Via (TSV), High Density (HD)-TSV and Copper to Copper (Cu-Cu) have arisen to take advantage of this possible 3-dimensional physical implementation with proofs of concept [1] or more recently industrial products [2]. Main drawback of these technologies is that they are not shrinking at the same speed as transistors are, making them somehow power hungry; moreover the more they will shrink, the more precision will be needed for chip to chip alignment. To reach the highest possible standard cell and tier to tier interconnect densities required for cost-effective chips, 3D sequential integration process [3][4][5] (also known as Monolithic 3D or CoolCubeTM) is currently developed with main features being sequential fabrication of MOS layers and correlation of tier to tier interconnect size with process node allowing fine-grain 3D partitioning of designs. These particularities make it a durable opportunity to slow down next node design migration while still improving integration. To fully benefit from CoolCubeTM technology, a whole new way of designing circuits, from synthesis to place and route, will be required as some new challenges will arise. The point of this presentation is to show the possible use and limitations of the aforementioned technologies with a focus on Monolithic 3D and to give some insights about market expectations, challenges and available design techniques. Olivier Billoint, Hossam Sarhan, Iyad Rayane, Maud Vinet, Perrine Batude, Claire Fenouillet-Béranger, Olivier Rozeau, Gerald Cibrario, Fabien Deprat, Ogun Turkyilmaz, Sébastien Thuries, Fabien Clermidy |
ISPD | 12 |
| 2015 | Compact interconnect approach for networks of neural cliques using 3D technologyabstractThanks to their brain-like properties, neural networks outperform traditional algorithms in certain group of applications. However, since they are wire-dominated systems, their hardware implementation poses numerous challenges as high latency and energy consumption. The recent technological improvements allow for stacking few dies one on another and designing 3D electronic circuits. This creates opportunities for 3D efficient implementations of neural networks targeting high-performance applications. This work explores the gains of 3D technology for neural networks relying on neural cliques. A general study shows up to 55% reduction in terms of total interconnect length and interconnect power consumption, and 74% reduction of the maximal interconnect delay. The proposed approach is validated with a power management applicative test-case. We demonstrate that, in this scenario, the 3D architecture reduces interconnect length and power by 35% and the maximal delay by 57%, compared to 2D. Bartosz Boguslawski, Hossam Sarhan, Frédéric Heitzmann, Fabrice Seguin, Sébastien Thuries, Olivier Billoint, Fabien Clermidy |
VLSI-SoC | 7 |
| 2014 | Advanced technologies for brain-inspired computingabstractThis paper aims at presenting how new technologies can overcome classical implementation issues of Neural Networks. Resistive memories such as Phase Change Memories and Conductive-Bridge RAM can be used for obtaining low-area synapses thanks to programmable resistance also called Memristors. Similarly, the high capacitance of Through Silicon Vias can be used to greatly improve analog neurons and reduce their area. The very same devices can also be used for improving connectivity of Neural Networks as demonstrated by an application. Finally, some perspectives are given on the usage of 3D monolithic integration for better exploiting the third dimension and thus obtaining systems closer to the brain. Fabien Clermidy, Rodolphe Héliot, Alexandre Valentian, Christian Gamrat, Olivier Bichler, Marc Duranton, Bilel Belhadj, Olivier Temam |
ASP-DAC | 1 |
| 2014 | A scalable custom simulation machine for the Bayesian Confidence Propagation Neural Network model of the brainabstractA multi-chip custom digital super-computer called eBrain for simulating Bayesian Confidence Propagation Neural Network (BCPNN) model of the human brain has been proposed. It uses Hybrid Memory Cube (HMC), the 3D stacked DRAM memories for storing synaptic weights that are integrated with a custom designed logic chip that implements the BCPNN model. In 22nm node, eBrain executes BCPNN in real time with 740 TFlops/s while accessing 30 TBs synaptic weights with a bandwidth of 112 TBs/s while consuming less than 6 kWs power for the typical case. This efficiency is three orders better than general purpose supercomputers in the same technology node. Nasim Farahini, Ahmed Hemani, Anders Lansner, Fabien Clermidy, Christer Svensson |
ASP-DAC | 4 |
| 2014 | 3DCoB: A new design approach for Monolithic 3D Integrated circuitsabstract3D Monolithic Integration (3DMI) technology provides very high dense vertical interconnects with low parasitics. Previous 3DMI design approaches provide either cell-on-cell or transistor-on-transistor integration. In this paper we present 3D Cell-on-Buffer (3DCoB) as a novel design approach for 3DMI. Our approach provides a fully compatible sign-off physical implementation flow with the conventional 2D tools. We implement our approach on a set of benchmark circuits using 28nm-FDSOI technology. The sign-off performance results show 35% improvement compared to the same 2D design. Hossam Sarhan, Sébastien Thuries, Olivier Billoint, Fabien Clermidy |
ASP-DAC | 4 |
| 2014 | Resistive memories: Which applications?abstractRecent announcement of 16Gbits Resistive memory from Sony shows the trend to quickly adopt resistive memories as an alternative to DRAM. However, using ReRAM for embedded computing is still a futuristic goal. This paper approaches two applications based on ReRAM-devices for gaining area, performance or power consumption. The first application is FPGA, one of the first architecture that can benefit the most from ReRAM integration to reduce footprint and save energy. The second application relates to ultra-low-power systems and the way to obtain an instantaneous “freeze” mode in devices for Internet of Things. Fabien Clermidy, Natalija Jovanovic, Santhosh Onkaraiah, Houcine Oucheikh, Olivier Thomas, Ogun Turkyilmaz, Elisa Vianello, Jean-Michel Portal, Marc Bocquet |
DATE | 1 |
| 2014 | 3D FPGA using high-density interconnect Monolithic IntegrationabstractNew 3D technology, called “Monolithic Integration”, offers very dense 3D interconnect capabilities. In this paper, we propose a 3D FPGA architecture with logic-on-memory approach based on this technology. The routing and computation blocks are splitted into two layers where the logic is placed on the top and memory on the bottom. Using extracted values from layout in 14nm FDSOI technology, typical benchmark circuits are evaluated in the VPR5 toolflow. The results show an area reduction of 55% compared to the 2D FPGA. More importantly, due to the lowered routing congestion, the EDP of the 3D FPGA is improved by 47%. Ogun Turkyilmaz, Gerald Cibrario, Olivier Rozeau, Perrine Batude, Fabien Clermidy |
DATE | 5 |
| 2014 | RRAM-based FPGA for "Normally Off, Instantly On" applications
Ogun Turkyilmaz, Santhosh Onkaraiah, Marina Reyboz, Fabien Clermidy, Hraziia, Costin Anghel, Jean-Michel Portal, Marc Bocquet |
J. Parallel Distributed Comput. | 4 |
| 2013 | A dynamic stream link for efficient data flow control in NoC based heterogeneous MPSoCabstractAs Systems-on-Chip size increase, the communication costs become critical and Networks-on-Chip (NoC) bring innovative solutions. Efficient stream-based protocols over NoC have been widely studied to address dataflow communications. They are usually controlled by a set of static parameters. However, new applications, such as high-resolution video decoders, present more data-dependent behaviors forcing communication protocols to support higher dynamicity. For this purpose, we present in this paper dynamic stream links for stream-based end-to-end NoC communications by introducing two link protocols, both independent of the transfer size, allowing to improve the hardware/software control flexibility. The proposed protocols have been modeled in a MPSoC virtual platform and the hardware cost evaluated. Based on simulations, we provide guidelines to exploit these protocols according to application needs. Claude Helmstetter, Sylvain Basset, Romain Lemaire, Fabien Clermidy, Pascal Vivet, Michel Langevin, Chuck Pilkington, Pierre G. Paulin, Didier Fuin |
ASP-DAC | 4 |
| 2013 | 3D stacking for multi-core architectures: From WIDEIO to distributed cachesabstract3D stacking has been viewed as a breakthrough solution for increasing performance in multi-core architectures. The hope is to solve some of the main issues in current multi-core architectures: external memory pressure and latency; I/O bottleneck; communication power consumption. In this paper, some advances of this field of research are shown, starting with a WIDEIO experience on a real chip for solving DRAM accesses issue. The integration of a 512 bit-width bus is demonstrated in a Network-on-Chip (NoC) multi-core framework and the resulting performance based on a 65nm prototype with 10μm diameter Through Silicon Vias (TSV). The potentiality of 3D scaling thanks to 3D asynchronous Network-on-Chip implementation is then shown. Finally, an innovative 3D stacked distributed cache strategy aimed at lowering memory latency and external memory bandwidth requirements is presented. This new memory partitioning demonstrates the efficiency of 3D stacking to rethink architectures for addressing multi-core scaling challenges. Fabien Clermidy, Denis Dutoit, Eric Guthmuller, Ivan Miro Panades, Pascal Vivet |
ISCAS | 1 |
| 2013 | A hybrid CBRAM/CMOS Look-Up-Table structure for improving performance efficiency of Field-Programmable-Gate-ArrayabstractAt most advanced technology nodes, Field Programmable Gate Arrays (FPGA) present great advantages compared to more conventional processor architectures; their natural regularity, modularity and inherent reliability due to duplicated identical tiles provide a solution to overcome new technologies with increasing variability. However, FPGA market is still limited by power efficiency issue, due to two coordinated factors like interconnection-dominated design and large usage of memories, computation being performed thanks to Look-Up-Table (LUT). In this paper, we propose a solution to improve the performance and reduce the power consumption of LUT in FPGA using CBRAM-based structures. Our proposed design shows significant improvement compared to the traditional SRAM-based FPGA in: critical delay is reduced by ~23% due to compact structure (1T-2R) and power gain by reduction in static power consumption by ~18%. Santhosh Onkaraiah, Ogun Turkyilmaz, Marina Reyboz, Fabien Clermidy, Elisa Vianello, Jean-Michel Portal, Christophe Muller |
ISCAS | 4 |
| 2013 | Self-checking ripple-carry adder with Ambipolar Silicon NanoWire FETabstractFor the rapid adoption of new and aggressive technologies such as ambipolar Silicon NanoWire (SiNW), addressing fault-tolerance is necessary. Traditionally, transient fault detection implies large hardware overhead or performance decrease compared to permanent fault detection. In this paper, we focus on on-line testing and its application to ambipolar SiNW. We demonstrate on self-checking ripple-carry adder how ambipolar design style can help reduce the hardware overhead. When compared with equivalent CMOS process, ambipolar SiNW design shows a reduction in area of at least 56% (28%) with a decreased delay of 62% (6%) for Static (Transmission Gate) design style. Ogun Turkyilmaz, Fabien Clermidy, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 2 |
| 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applicationsabstractP2012 is an area- and power-efficient many-core computing accelerator based on multiple globally asynchronous, locally synchronous processor clusters. Each cluster features up to 16 processors with independent instruction streams sharing a multi-banked one-cycle access L1 data memory, a multi-channel DMA engine and specialized hardware for synchronization and aggressive power management. P2012 is 3D stacking ready and can be customized to achieve extreme area and energy efficiency by adding domain-specific HW IPs to the cluster. The first P2012 SoC prototype in 28nm CMOS will sample in Q3, featuring four 16-processor clusters, a 1MB L2 memory and delivering 80GOPS (with 32 bit single precision floating point support) in 18mm2 with 2W power consumption (worst-case). P2012 can run standard OpenCL™ and proprietary Native Programming Model SW components to achieve the highest level of control on application-to-resource mapping. A dedicated version of the OpenCV vision library is provided in the P2012 SW Development Kit to enable visual analytics acceleration. This paper will discuss preliminary performance measurements of common feature extraction and tracking algorithms, parallelized on P2012, versus sequential execution on ARM CPUs. Diego Melpignano, Luca Benini, Eric Flamand, Bruno Jego, Thierry Lepley, Germain Haugou, Fabien Clermidy, Denis Dutoit |
DAC | 7 |
| 2011 | Can we go towards true 3-D architectures?abstractThanks to recent technology advances, the exploration of the vertical dimension has been shown to be more than a dream for designers. Among those technologies, the vertical transistor has not been exploited yet. This paper describes a novel implementation of logic gates fully benefiting of nanowire-based vertical transistors embedded within the metal lines. The logic design in this technology is explored and its performance is evaluated. A comparison made on an equivalent technology node shows that our cells reduce area and delay by a factor of 31x and 2x respectively. Large reconfigurable logic circuits have been benchmarked showing an improvement of area and delay by 46% and 48% on average. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Paul-Henry Morel, Jean-Philippe Noël, Fabien Clermidy, Ian O'Connor |
DAC | 5 |
| 2011 | A low-power VLIW processor for 3GPP-LTE complex numbers processingabstractNew generation of telecommunication applications requires highly efficient processing units to tackle with the increasing signal processing algorithmic complexity. They also need to be flexible for handling a large range of radio access technology with specifications moving very fast. As devices including telecommunication features are, per nature, mobile, the high level of flexibility must be achieved while preserving very low power consumption. In this paper, a high performance low-power application-specific processor is proposed for complex signal processing. Thanks to dedicated control architecture, this processor exhibits an average 81% utilization rate of its principal operator, a complex MAC for a 3GPP-LTE application. The main innovations are the use of a reconfigurable profile and instruction cache strategy to reduce power consumption. This leads to a 10× reduction of the control power consumption. As a result, an average 50 mW power consumption is measured after implementation in a low-power 65 nm technology while delivering 3.2 GOPS. Finally, a comparison with state-of-the-art low-power DSP shows at least 24 % gain. Christian Bernard, Fabien Clermidy |
DATE | 2 |
| 2011 | 3D Embedded multi-core: Some perspectivesabstract3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range products. In this paper, we investigate three promising perspectives for short to medium terms adoption of such technology in high-end System-on-Chip built around multi-core architectures: the wide bus concept will help solving high bandwidth requirements with external memory. 3D Network-on-Chip is a promising solution for increased modularity and scalability. We show that an efficient implementation provides an available bandwidth outperforming classical interfaces. Finally, we put in perspective the active interposer concept which aims at simplifying and improving power, test and debug management. Fabien Clermidy, Florian Darve, Denis Dutoit, Walid Lafi, Pascal Vivet |
DATE | 1 |
| 2011 | Sustainability through massively integrated computing: Are we ready to break the energy efficiency wall for single-chip platforms?abstractWhile traditional cluster computers are more constrained by power and cooling costs for solving extreme-scale (or exascale) problems, the continuing progress and integration levels in silicon technologies make possible complete end-user systems on a single chip. This massive level of integration makes modern multicore chips all pervasive in domains ranging from climate forecasting and astronomical data analysis, to consumer electronics, smart phones, and biological applications. Consequently, designing multicore chips for exascale computing while using the embedded systems design principles looks like a promising alternative to traditional cluster-based solutions. This paper aims to present an overview of new, far-reaching design methodologies and run-time optimization techniques that can help breaking the energy efficiency wall in massively integrated single-chip computing platforms. Partha Pratim Pande, Fabien Clermidy, Diego Puschini, Imen Mansouri, Paul Bogdan, Radu Marculescu, Amlan Ganguly |
DATE | 2 |
| 2011 | A low complexity stopping criterion for reducing power consumption in turbo decodersabstractTurbo codes are proposed in most of the advanced digital communication standards, such as 3GPP-LTE. However, due to its computational complexity, the turbo decoder is one of the most power hungry blocks in digital baseband. To alleviate this issue, one way is to avoid surplus computing phases thanks to the early termination of the iterative decoding process. The use of stopping criteria is one of the most common algorithm level power reduction methods in literature. These methods always come with some hardware overhead. In this paper, a new trellis based stopping criterion is proposed. The novelty of this approach is the lower hardware overhead thanks to the use of trellis states as key parameter to stop the iterative process. Results are showing the importance of this added hardware in terms of method efficiency. Compared to state-of-the-art Log Likelihood Ratio (LLR) based techniques, proposed Low Complexity Trellis Based (LCTB) is demonstrating 23% less power consumption on average, for comparable performance level in terms of Bit Error Rate (BER) and Frame Error Rate (FER). Pallavi Reddy, Fabien Clermidy, Amer Baghdadi, Michel Jézéquel |
DATE | 2 |
| 2011 | Dynamic Flow Reconfiguration Strategy to Avoid Communication Hot-SpotsabstractApplication-specific Network-on-Chip allows optimization for the interconnection to minimize its cost. When used with streaming applications, large flows of data can be predicted. However, these flows can be modified during the applications providing a dynamic flow graph. In that case, applying on off-line optimization leads to an over-sizing of the NoC. On the other hand, dynamic reconfiguration leads to unordered data deliveries with costly re-ordering units. In this paper, we propose a coarse grain dynamic reconfiguration which avoids data re-ordering requirement. We show that the proposed solution is efficient to deal with communication hot-spots, with a small area overhead, and can save up to 33% of latency. Romain Prolonge, Fabien Clermidy, Leonel Tedesco, Fernando Gehm Moraes |
DSD | 2 |
| 2011 | Evaluation of a crossbar multiplexer in a lithography-based nanowire technologyabstractSilicon Nanowire technology has been demonstrated to be a promising candidate to fabricate nanowire crossbars. The use of such devices in a real architectural as well as in a design environment is an ongoing research topic. In this paper, we investigate the use of a lithography-based industrial process for designing a 4-to-1 multiplexer in a crossbar circuit. We show that by considering the line parasitic, the crossbar demonstrates poor performance in a 65-nm technology, while the area and power savings are about 6× and 1.5× respectively vs. the CMOS implementation. However, extrapolation to the 9-nm node shows a 2× better performance and 67× area saving. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Fabien Clermidy, Ian O'Connor |
ISCAS | 3 |
| 2011 | Reconfiguration of a 3GPP-LTE telecommunication application on a 22-core NoC-based system-on-chipabstractThe MAGALI chip is a 65nm digital baseband dedicated to advanced telecommunication applications. Based on a 15-router mesh Network-on-Chip - NoC, it embeds 22 Processing Elements - PE performing the different functions of a complex baseband in a programmable manner. The NoC framework is based on asynchronous routers which provide a complete Globally Asynchronous Locally Synchronous -- GALS framework. Thanks to this structure, each PE is a synchronous island with a programmable frequency. On the architectural side, the NoC supports fast reconfiguration thanks to a distributed scheme. In this demonstration, we propose to show the reconfiguration capabilities of the MAGALI chip on three modes of a 3GPP-LTE receiver part, by switching on-the-fly between these three modes. The resulting transmission quality with different levels of upcoming signal noises is shown. Fabien Clermidy, Nicolas Cassiau, N. Coste, Denis Dutoit, M. Fantini, Dimitri Ktenas, Romain Lemaire, L. Stefanizzi |
NOCS | 1 |
| 2011 | 3D NoC using through silicon Via: An asynchronous implementationabstract3D stacking is seen as one of the most interesting technologies for System-on-Chip (SoC) developments. However, 3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range of products. In this paper, we are investigating 3D Network-on-Chip has a promising solution for increased modularity and scalability. We show that an efficient implementation based on asynchronous logic provides an available bandwidth of 64GB/s for only 700 TSV, outperforming classical interfaces while simplifying the assembly process. We also point out the benefit in terms of power consumption for these new interfaces with a gain of 5 times compared to classical LPDDR2 interfaces. Pascal Vivet, Denis Dutoit, Yvain Thonnart, Fabien Clermidy |
VLSI-SoC | 4 |
| 2011 | Matrix Nanodevice-Based Logic Architectures and Associated Functional Mapping MethodabstractThis article describes a novel computing architecture organization based on nanoscale logic cells. We propose the use of a cluster of matrix arrangements of cells. In order to interconnect such fine-grained logic cells within a matrix, conventional techniques are not suitable due to a large interconnect overhead. Therefore, we propose the use of static and incomplete interconnect topologies to create matrices of cells. We also propose a method to map functions onto such architectures. We then explore the main parameters of the structure (size of matrices and interconnect topologies) and their impact on the main performance metrics (packing efficiency, speed, and fault tolerance). A cluster packing method also allows the evaluation of the number of matrices used by complex functions and the fill factor for various matrix sizes. The analyses show that this approach is particularly suited for matrices of 16 cells interconnected by modified omega networks. We can conclude that this architecture could improve the scalability of traditional FPGAs by a factor of 8.5. Pierre-Emmanuel Gaillardon, Fabien Clermidy, Ian O'Connor, Maimouna Amadou, Gabriela Nicolescu |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | A fully-asynchronous low-power framework for GALS NoC integrationabstractRequiring more bandwidth at reasonable power consumption, new communication infrastructures must provide adequate solutions to guarantee performance during physical integration. In this paper, we propose the design of a low-power asynchronous Network-on-Chip which is implemented in a bottom-up approach using optimized hard-macros. This architecture is fully testable and a new design flow is proposed to overcome CAD tools limitations regarding asynchronous logic. The proposed architecture has been successfully implemented in CMOS 65nm in a complete circuit. It achieves a 550Mflit/s throughput on silicon, and exhibits 86% power reduction compared to an equivalent synchronous NoC version. Yvain Thonnart, Pascal Vivet, Fabien Clermidy |
DATE | 3 |
| 2010 | Phase-change-memory-based storage elements for configurable logicabstractBack-end-of-line non-volatile resistive memories like Phase Change Memories (PCMs) are promising to solve memory issues in different architectures. In this paper, we investigate the usage of PCM to build an elementary configuration memory node for reconfigurable logic, such as Field-Programmable Gate Arrays (FPGAs). We propose an elementary circuit realized by 2 resistive memories and 1 programming transistor able to store a configuration voltage. We investigate the proposed node in terms of area and write time and we assess its impact on complex circuits. We show that the elementary memory node yields an improvement in area and write time of 1.5x and 16x respectively vs. a regular Flash implementation. Implemented in FPGAs, the memory node yields a delay reduction up to 51%, thanks to the reduction of dimensions and low on-resistance of PCMs. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Marina Reyboz, Giovanni Beneventi, Fabien Clermidy, Luca Perniola, Ian O'Connor |
FPT | 5 |
| 2010 | Distributed Sequencing for Resource Sharing in Multi-applicative Heterogeneous NoC PlatformsabstractIn the context of heterogeneous NoC architectures for embedded systems, it is today mandatory to support multiple applications given the plurality of standards and usages. While static reconfiguration between applications has already been extensively studied, we propose a potential increase in hardware resource usage by enabling concurrent or overlapping applications on the top of a heterogeneous NoC platform. In this paper, we describe a distributed sequencing protocol allowing hardware resource sharing between several applications. This protocol ensures correct synchronization of the processing between hardware resources without the need of a global fine-grain scheduler on the system, thus alleviating the pressure on the run-time system. The proposed protocol has been integrated and validated in a NoC-based digital baseband for 4G SDR telecom applications, and was integrated on a manufactured chip on a STMicroelectronics CMOS 65 nm LP technology. Yvain Thonnart, Romain Lemaire, Fabien Clermidy |
NOCS | 3 |
| 2009 | Dynamic and distributed frequency assignment for energy and latency constrained MP-SoCabstractIn this paper we present an adaptive technique to locally adjust the frequency of processing elements on MP-SoC. The proposed method, based on game theory, optimizes the system while fulfilling dynamic constraints. A telecom test-case has been used to demonstrate the effectiveness of our technique. For the evaluated scenario, the proposed technique has obtained up to 20% of latency gain and 38% of energy gain. Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
DATE | 2 |
| 2009 | An Open and Reconfigurable Platform for 4G Telecommunication: Concepts and ApplicationabstractAdvanced telecommunication applications require more and more flexibility. Static and run-time configuration mechanisms, as well as plug-in computing units are two solutions to solve this issue. In this paper, we propose an open architecture, dedicated to complex data-flow applications, fulfilling these requirements thanks to a distributed configuration scheme and a standard interface for data flow computing units. Based on this architecture, an open platform demonstrator has been designed. Simulation results on a 3GPP/LTE application are presented, showing a reconfiguration time overhead of only 2.6%. Fabien Clermidy, Romain Lemaire, Xavier Popon, Dimitri Ktenas, Yvain Thonnart |
DSD | 1 |
| 2009 | Open Platform for Prototyping of Advanced Software Defined Radio and Cognitive Radio TechniquesabstractThis paper presents the ANR project IDROMel, which aims at developing reconfigurable SDR (software defined radio) and cognitive radio (CR) equipments. IDROMel is a 3 years project that started in 2005 and finishes in 2009. The main objective of IDROMel is to define, develop and validate a powerful SDR and CR platform combining very last technology progresses. The platform includes software parts (reconfigurable protocol stacks) and hardware parts (a base band board and a radio frequency front end, RF). Both parts are presented in this paper. Dominique Nussbaum, Karim Khalfallah, Christophe Moy, Amor Nafkha, Pierre Leray, Julien Delorme, Jacques Palicot, Jérôme Martin, Fabien Clermidy, Bertrand Mercier, Renaud Pacalet |
DSD | 9 |
| 2009 | A Communication and configuration controller for NoC based reconfigurable data flow architectureabstractWhile network-on-chip aspects such as topologies, routing strategies or quality-of-service have been largely studied, the mapping of real applications on distributed NoC-based architecture is still an open issue. In this paper, we address this issue for complex reconfigurable data-flow applications. We introduce the concept of communication and configuration controller (CCC) which interacts both with the usual network interface and the IP core structure. The proposed CCC is a programmable template-based architecture, which provides solutions to manage reconfiguration flows, data synchronizations and global control signaling. An implementation of the CCC is presented, and its performances in a 65 nm technology are discussed through a concrete telecommunication application. Fabien Clermidy, Romain Lemaire, Yvain Thonnart, Pascal Vivet |
NOCS | 1 |
| 2009 | Emerging Technologies and Nanoscale Computing Fabricsabstract6-8 July 2014 Ian O'Connor, Kotb Jabeur, Nataliya Yakymets, Renaud Daviot, David Navarro, Pierre-Emmanuel Gaillardon, Fabien Clermidy, Maimouna Amadou, Gabriela Nicolescu |
VLSI-SoC | 8 |
| 2008 | Convergence analysis of run-time distributed optimization on adaptive systems using game theoryabstractWe consider multiprocessor system-on-chip (MP-SoC) integrating several processing elements (PE). These architectures require distributed and scalable control techniques for run-time optimization of applicative parameters. Our approach is to use the game theory as an optimization model to solve the trade-off issues at run-time. We applied it to the distributed dynamic voltage frequency scaling (DVFS) management, adjusting at run-time the frequency set of each PE based on the synchronization between tasks of the application graph and the PE temperature profile. Results show that the analyzed algorithm converges to a solution in about 94% of the cases and in less than 40 calculation cycles for a 100-processor MP-SoC. It reaches an average optimization of 89% compared to an off-line centralized reference but about 140 times faster when simulating. Diego Puschini, Fabien Clermidy, Pascal Benoit, Gilles Sassatelli, Lionel Torres |
FPL | 2 |
| 2008 | Dynamic Voltage and Frequency Scaling Architecture for Units Integration within a GALS NoC
Edith Beigné, Fabien Clermidy, Sylvain Miermont, Pascal Vivet |
NOCS | 2 |
| 2008 | Physical Implementation of the DSPIN Network-on-Chip in the FAUST Architecture
Ivan Miro Panades, Fabien Clermidy, Pascal Vivet, Alain Greiner |
NOCS | 2 |
| 1999 | A New Placement Algorithm Dedicated to Parallel Computers: Bases and ApplicationabstractOne way to improve reliability in parallel computers consists of adding supplementary processors and interconnections to the functional structure in order to replace faulty processors with respect to the network structure. This approach is named structural fault tolerance (SFT). Very integrated parallel computers are one way to implement a parallel structure. The material structure is then composed of many elementary blocks, such as ASICs or multi-chip modules (MCMs), each containing many processors. We show that former SFT methods fail in combining the different features, constraints and requirements of such structures. Thus, this paper introduces a new reconfiguration approach that is dedicated to very integrated parallel computers. Fabien Clermidy, Thierry Collette, Michael Nicolaidis |
PRDC | 1 |