EDBT 2026 Demo / reviewers in the wild / expert
Pascal Vivet
dblp:13/6934
· DBLP profile ↗
42ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-7413-8243ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 16 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based VisionabstractEvent-Based cameras asynchronously detect changes in light intensity with high temporal resolution, making them a promising alternative to RGB camera for low-latency and low-power optical flow estimation. However, state-of-the-art convolutional neural network methods create frames from the event stream, therefore losing the opportunity to exploit events for both sparse computations and low-latency prediction. On the other hand, asynchronous event graph methods could leverage both, but at the cost of avoiding any form of time accumulation, which limits the prediction accuracy. In this paper, we propose to break this accuracy-latency trade-off with a novel architecture combining an asynchronous accumulation-free event branch and a periodic aggregation branch. The periodic branch performs feature aggregations on the event graphs of past data to extract global context information, which improves accuracy without introducing any latency. The solution could predict optical flow per event with a latency of tens of microseconds on asynchronous hardware, which represents a gain of three orders of magnitude with respect to state-of-the-art frame-based methods, with 48x less operations per second. We show that the solution can detect rapid motion changes faster than a periodic output. This work proposes, for the first time, an effective solution for ultra low-latency and low-power optical flow prediction from event cameras. Manon Dampfhoffer, Thomas Mesquida, Damien Joubert, Thomas Dalgaty, Pascal Vivet, Christoph Posch |
CVPR | 5 |
| 2025 | Event-based Audio Prediction with Spectro-Temporal Event-GraphsabstractGraph neural networks have recently emerged as a promising approach for low-power and low-latency event-vision applications. Such event-graphs naturally exploit the sparsity of event-data and incorporate the temporal detail captured by event-based sensors directly into the edge features used in graph convolution. In this paper we study the promise of event-graphs for processing data from other event-based modalities beyond vision. Specifically, we describe how the approach can be adapted to the spectro-temporal domain to perform event-audio classification. We evaluate the approach using the spiking Heidelberg digits dataset and achieve a test accuracy of 94.3%. This is notably better than many state of the art spiking neural networks despite, in many cases, requiring an order of magnitude fewer parameters. Event-graph neural networks promise to be a powerful, general approach for processing a variety of event-based modalities, not only vision. Lars Rafeldt, Thomas Mesquida, Manon Dampfhoffer, Filippo Moro, Pascal Vivet, Melika Payvand, Thomas Dalgaty |
ISCAS | 6 |
| 2025 | J3DAI: A tiny DNN-Based Edge AI Accelerator for 3D-Stacked CMOS Image SensorabstractThis paper presents J3DAI, a tiny deep neural network-based hardware accelerator for a 3-layer 3D-stacked CMOS image sensor featuring an artificial intelligence (AI) chip integrating a Deep Neural Network (DNN)-based accelerator. The DNN accelerator is designed to efficiently perform neural network tasks such as image classification and segmentation. This paper focuses on the digital system of J3DAI, highlighting its Performance-Power-Area (PPA) characteristics and showcasing advanced edge AI capabilities on a CMOS image sensor.To support hardware, we utilized the Aidge comprehensive software framework, which enables the programming of both the host processor and the DNN accelerator. Aidge supports post-training quantization, significantly reducing memory footprint and computational complexity, making it crucial for deploying models on resource-constrained hardware like J3DAI.Our experimental results demonstrate the versatility and efficiency of this innovative design in the field of edge AI, showcasing its potential to handle both simple and computationally intensive tasks. Benoît Tain, Raphael Millet, Romain Lemaire, Michal Szczepanski, Laurent Alacoque, Emmanuel Pluchart, Sylvain Choisnet, Rohit Prasad, Jérôme Chossat, Pascal Pierunek, Pascal Vivet, Sébastien Thuries |
ISLPED | 11 |
| 2025 | Exploring Enhancements to 1T1C FeMFET Bitcells with a Versatile DTCO MethodologyabstractNon-volatile in-memory computing (iMC) has emerged as an energy-efficient paradigm well suited to AI workloads. Its implementation using 1T1C FeMFETs (Ferroelectric Memory Field Effect Transistors), a best-in-class emerging nonvolatile memory technology that integrates BEOL ferroelectric devices with FEOL transistors, is of particular interest. This interest stems from their potential to enable large-scale multiplyaccumulate (MAC) operations in both digital and analog domains. However, realizing tangible performance benefits requires comprehensive cross-layer exploration of both design and technology parameters, extending up to accelerator level. In this work, we propose a bitcell-level multi-objective optimization methodology to identify and extract optimal sizing solutions that provide tractable trade-offs between key performance indicators (KPI). We further demonstrate how this approach facilitates cross-stack exploration of accelerator architectures. Results are presented as Pareto fronts spanning $2-4 \mathrm{KPIs}$: a $2-\mathrm{KPI}$ problem illustrates the methodology, while a $\mathbf{4}$-KPI problem represents a realistic design scenario. Comparison is made between $\mathbf{1 3 0} \mathbf{n m}$ and 28 nm technologies demonstrating a decrease in the average of write energy and area up to 24 X and 30 X respectively. Rosario Pronsat, Antoine Cauquil, Pascal Vivet, Jean Coignus, Damien Deleruyelle, Cédric Marchand 0002, Lioua Labrak, Ian O'Connor |
VLSI-SoC | 3 |
| 2023 | G2N2: Lightweight Event Stream Classification with GRU Graph Neural Networks
Thomas Mesquida, Manon Dampfhoffer, Thomas Dalgaty, Pascal Vivet, Amos Sironi, Christoph Posch |
BMVC | 4 |
| 2023 | The CNN vs. SNN Event-camera Dichotomy and Perspectives For Event-Graph Neural NetworksabstractSince neuromorphic event-based pixels and cameras were first proposed, the technology has greatly advanced such that there now exists several industrial sensors, processors and toolchains. This has also paved the way for a blossoming new branch of AI dedicated to processing the event-based data these sensors generate. However, there is still much debate about which of these approaches can best harness the inherent sparsity, low-latency and fine spatiotemporal structure of event-data to obtain better performance and do so using the least time and energy. The latter is of particular importance since these algorithms will typically be employed near or inside of the sensor at the edge where the power supply may be heavily constrained. The two predominant methods to process visual events - convolutional and spiking neural networks - are fundamentally opposed in principle. The former converts events into static 2D frames such that they are compatible with 2D convolutions, while the latter computes in an event-driven fashion naturally compatible with the raw data. We review this dichotomy by studying recent algorithmic and hardware advances of both approaches. We conclude with a perspective on an emerging alternative approach whereby events are transformed into a graph data structure and thereafter processed using techniques from the domain of graph neural networks. Despite promising early results, algorithmic and hardware innovations are required before this approach can be applied close or within the Event-based sensor. Thomas Dalgaty, Thomas Mesquida, Damien Joubert, Amos Sironi, Cyrille Soubeyrat, Pascal Vivet, Christoph Posch |
DATE | 6 |
| 2022 | Architecting Optically Controlled Phase Change MemoryabstractPhase Change Memory (PCM) is an attractive candidate for main memory, as it offers non-volatility and zero leakage power while providing higher cell densities, longer data retention time, and higher capacity scaling compared to DRAM. In PCM, data is stored in the crystalline or amorphous state of the phase change material. The typical electrically controlled PCM (EPCM), however, suffers from longer write latency and higher write energy compared to DRAM and limited multi-level cell (MLC) capacities. These challenges limit the performance of data-intensive applications running on computing systems with EPCMs. Recently, researchers demonstrated optically controlled PCM (OPCM) cells with support for 5 bits / cell in contrast to 2 bits / cell in EPCM. These OPCM cells can be accessed directly with optical signals that are multiplexed in high-bandwidth-density silicon-photonic links. The higher MLC capacity in OPCM and the direct cell access using optical signals enable an increased read/write throughput and lower energy per access than EPCM. However, due to the direct cell access using optical signals, OPCM systems cannot be designed using conventional memory architecture. We need a complete redesign of the memory architecture that is tailored to the properties of OPCM technology. This article presents the design of a unified network and main memory system called COSMOS that combines OPCM and silicon-photonic links to achieve high memory throughput. COSMOS is composed of a hierarchical multi-banked OPCM array with novel read and write access protocols. COSMOS uses an Electrical-Optical-Electrical (E-O-E) control unit to map standard DRAM read/write commands (sent in electrical domain) from the memory controller on to optical signals that access the OPCM cells. Our evaluation of a 2.5D-integrated system containing a processor and COSMOS demonstrates 2.14 × average speedup across graph and HPC workloads compared to an EPCM system. COSMOS consumes 3.8× lower read energy-per-bit and 5.97× lower write energy-per-bit compared to EPCM. COSMOS is the first non-volatile memory that provides comparable performance and energy consumption as DDR5 in addition to increased bit density, higher area efficiency, and improved scalability. Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ayse K. Coskun, Ajay Joshi |
ACM Trans. Archit. Code Optim. | 3 |
| 2021 | PROWAVES: Proactive Runtime Wavelength Selection for Energy-Efficient Photonic NoCsabstract2.5-D manycore systems running parallel applications are severely bottlenecked by network-on-chip (NoC) latencies and bandwidth. Traditionally, NoCs are composed of electrical links that exhibit constrained bandwidth, increased energy consumption at high-speed communication, and long latencies. Photonic NoCs (PNoCs) have been shown to provide high bandwidth at low latencies and negligible data-dependent power. However, the power overheads of lasers, thermal tuning, and electrical-optical conversion present major challenges against wide-scale adoption of PNoCs. A primary factor that impacts PNoC power is the number of activated laser wavelengths in the system. Applications’ dynamic bandwidth needs provide the opportunity to selectively deactivate laser wavelengths when there is a lower bandwidth demand to alleviate high PNoC power concerns. This article analyzes dynamic PNoC activity of applications at runtime so as to select laser wavelengths depending on an application’s bandwidth requirements. The article then proposesPROWAVES, a proactive runtime wavelength selection policy that forecasts the bandwidth needs and activates the minimum laser wavelengths for each application phase. We develop a cross-layer simulation framework to model the system performance, PNoC power and transient thermal distribution in a manycore system with PNoCs. We comparePROWAVESwith prior system-level policies and our simulation results on a 2.5-D system demonstrate thatPROWAVESprovides 18% and 33% power savings with only 1% and 5% loss in performance, respectively, compared to activating all laser wavelengths in the system. Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ayse K. Coskun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | System-level Evaluation of Chip-Scale Silicon Photonic Networks for Emerging Data-Intensive ApplicationsabstractEmerging data-driven applications such as graph processing applications are characterized by their excessive memory footprint and abundant parallelism, resulting in high memory bandwidth demand. As the scale of datasets for applications is reaching orders of TBs, performance limitation due to bandwidth demands is a major concern. Traditional on-chip electrical networks fail to meet such high bandwidth demands due to increased energy-per-bit or physical limitations with pin counts. Silicon photonic networks have emerged as a promising alternative to electrical interconnects, owing to their high bandwidth density and low energy-per-bit communication with negligible data-dependent power. Wide-scale adoption of silicon photonics at chip level, however, is hampered by their high sensitivity to process and thermal variations, high laser power due to losses along the network, and power consumption of the electrical-optical conversion. Device-level technological innovations to mitigate these issues are promising, yet they do not consider the system-level implications of the applications running on manycore systems with photonic networks. This work aims to bridge the gap between the system-level attributes of applications with the underlying architectural and device-level characteristics of silicon photonic networks to achieve energy-efficient computing. We particularly focus on graph applications, which involve unstructured yet abundant parallel memory accesses that stress the on-chip communication networks, and develop a cross-layer framework to evaluate 2.5D systems with silicon photonic networks. We demonstrate 38% power savings through system-level management using wavelength selection policies with only 1% loss in system performance and further evaluate architectural design choices on 2.5D systems with photonic networks. Aditya Narayan, Yvain Thonnart, Pascal Vivet, Ajay Joshi, Ayse K. Coskun |
DATE | 3 |
| 2020 | Computational SRAM Design Automation using Pushed-Rule Bitcells for Energy-Efficient Vector ProcessingabstractThis paper presents a new methodology for automating the Computational SRAM (C-SRAM) design based on off-the-shelf memory compilers and a configurable RTL IP. The main goal is to drastically reduce the development effort compared to a full-custom design, while offering a flexibility of use and a high-yield production. The proposed C-SRAM architecture has been developed to process energy-efficient vector data coupled with a scalar processor, while limiting the data transfer on the system bus. The results obtained by post P&R simulations show that 2RW and 4RW C-SRAM configurations using the double pumping technique achieved the highest performance to process vectorized MAC operations compared to the others configurations. Moreover, it has been shown that the impact of the digital wrapper decoding and executing the instructions can be mitigated by increasing the memory cut size to represent less than 10% in area and 20% in power consumption. Jean-Philippe Noël, Valentin Egloff, Maha Kooli, Roman Gauchi, Jean-Michel Portal, Henri-Pierre Charles, Pascal Vivet, Bastien Giraud |
DATE | 7 |
| 2020 | POPSTAR: a Robust Modular Optical NoC Architecture for Chiplet-based 3D Integrated SystemsabstractSilicon photonics technology is now gaining maturity with increasing levels of design complexity from devices to large photonic integrated circuits. Close integration of control electronics with 3D assembly of photonics and CMOS opens the way to high-performance computing architectures partitioned in chiplets connected by optical NoC on silicon photonic interposers. In this paper, we give an overview of our works on optical links and NoC for manycore systems, from low-level control of photonic devices to high-level system optimization of the optical communications. We detail the POPSTAR optical NoC topology and architecture (Processors On Photonic Silicon interposer Terascale ARchitecture) with electro-optical interface chiplets, the corresponding nested spiral topology for single-writer multiple- reader links and the associated control electronics, in charge of high-speed drivers, thermal stabilization and handling of the protocol stack, from data integrity to flow-control, routing and arbitration of the optical communications. The strengths and opportunities for this architecture will be discussed, with a shift in system & implementation constraints with respect to previous optical NoC proposals, and new challenges to be addressed. Yvain Thonnart, Stéphane Bernabé, Jean Charbonnier, Christian Bernard, David Coriat, César Fuguet Tortolero, Pierre Tissier, Benoît Charbonnier, Stephane Malhouitre, Damien Saint-Patrice, Myriam Assous, Aditya Narayan, Ayse K. Coskun, Denis Dutoit, Pascal Vivet |
DATE | 15 |
| 2020 | M3D-ADTCO: Monolithic 3D Architecture, Design and Technology Co-Optimization for High Energy Efficient 3D ICabstractMonolithic 3D (M3D) stands now as the ultimate technology to side step Moore’s Law stagnation. Due to its nanoscale Monolithic Inter-tier Via (MIV), M3D enables an ultrahigh density interconnect between Logic and Memory that is required in the field of highly energy efficient 3D integrated circuits (3D-ICs) designed for new abundant data computing systems. At design level, M3D still suffers from a lack of commercial tools, especially for Place and Route, precluding the capability to provide signoff M3D GDS. In this paper, we introduce M3D-ADTCO, an architecture, design and technology co-optimization platform aimed at providing signoff M3D GDS. It relies on a M3D Process Design Kit and the use of a commercial Place and Route tool. We demonstrate an area reduction of 23.61 % at iso performance and power compared to a 2D RISC-V micro-controller based System on Chip (SoC) while creating space to increase (2x) the RISC-V instruction memory. Sébastien Thuries, Olivier Billoint, Sylvain Choisnet, Romain Lemaire, Pascal Vivet, Perrine Batude, Didier Lattard |
DATE | 5 |
| 2020 | Reconfigurable tiles of computing-in-memory SRAM architecture for scalable vectorizationabstractFor big data applications, bringing computation to the memory is expected to reduce drastically data transfers, which can be done using recent concepts of Computing-In-Memory (CIM). To address kernels with larger memory data sets, we propose a reconfigurable tile-based architecture composed of Computational-SRAM (C-SRAM) tiles, each enabling arithmetic and logic operations within the memory. The proposed horizontal scalability and vertical data communication are combined to select the optimal vector width for maximum performance. These schemes allow to use vector-based kernels available on existing SIMD engines onto the targeted CIM architecture. For architecture exploration, we propose an instruction-accurate simulation platform using SystemC/TLM to quantify performance and energy of various kernels. For detailed performance evaluation, the platform is calibrated with data extracted from the Place&Route C-SRAM circuit, designed in 22nm FDSOI technology. Compared to 512-bit SIMD architecture, the proposed CIM architecture achieves an EDP reduction up to 60× and 34× for memory bound kernels and for compute bound kernels, respectively. Roman Gauchi, Valentin Egloff, Maha Kooli, Jean-Philippe Noël, Bastien Giraud, Pascal Vivet, Subhasish Mitra, Henri-Pierre Charles |
ISLPED | 6 |
| 2019 | WAVES: Wavelength Selection for Power-Efficient 2.5D-Integrated Photonic NoCsabstractPhotonic Network-on-Chips (PNoCs) offer promising benefits over Electrical Network-on-Chips (ENoCs) in many-core systems owing to their lower latencies, higher bandwidth, and lower energy-per-bit communication with negligible data-dependent power. These benefits, however, are limited by a number of challenges. Microring resonators (MRRs) that are used for photonic communication have high sensitivity to process variations and on-chip thermal variations, giving rise to possible resonant wavelength mismatches. State-of-the-art microheaters, which are used to tune the resonant wavelength of MRRs, have poor efficiency resulting in high thermal tuning power. In addition, laser power and high static power consumption of drivers, serializers, comparators, and arbitration logic partially negate the benefits of the sub-pJ operating regime that can be obtained with PNoCs. To reduce PNoC power consumption, this paper introduces WAVES, a wavelength selection technique to identify and activate the minimum number of laser wavelengths needed, depending on an application's bandwidth requirement. Our results on a simulated 2.5D manycore system with PNoC demonstrate an average of 23% (resp. 38%) reduction in PNoC power with only <;1% (resp. <;5%) loss in system performance. Aditya Narayan, Yvain Thonnart, Pascal Vivet, César Fuguet Tortolero, Ayse K. Coskun |
DATE | 3 |
| 2019 | Advanced 3D Technologies and Architectures for 3D Smart Image SensorsabstractImage Sensors will get more and more pervasive into their environment. In the context of Automotive and IoT, low cost image sensors, with high quality pixels, will embed more and more smart functions, such as the regular low level image processing but also object recognition, movement detection, light detection, etc. 3D technology is a key enabler technology to integrate into a single device the pixel layer and associated acquisition layer, but also the smart computing features and the required amount of memory to process all the acquired data. More computing and memory within the 3D Smart Image Sensors will bring new features and reduce the overall system power consumption. Advanced 3D technology with ultra-fine pitch vertical interconnect density will pave the way towards new architectures for 3D Smart Image Sensors, allowing local vertical communication between pixels, and the associated computing and memory structures. The presentation will give an overview of recent 3D technologies solutions, such as Hybrid Bonding technology and the Monolithic 3D CoolCube™ technology, with respective 3D interconnect pitch in the order of 1 μm and l00nm. Recent 3D Image Sensors will be presented, showing the capability of 3D technology to implement fine grain pixel acquisition and processing with ultra-high speed image acquisition and tile-based processing. As further perspectives, multi-layer 3D image sensor based on events and spiking will reduce power consumption with new detection and learning processing capabilities. Pascal Vivet, Gilles Sicard, Laurent Millet, Stéphane Chevobbe, Karim Ben Chehida, Luis Angel Cubero, Monte Alegre, Maxence Bouvier, Alexandre Valentian, Maria Lepecq, Thomas Dombek, Olivier Bichler, Sébastien Thuries, Didier Lattard, Séverine Cheramy, Perrine Batude, Fabien Clermidy |
DATE | 1 |
| 2019 | Test Solutions for High Density 3D-IC Interconnects - Focus on SRAM-on-Logic PartitioningabstractTest infrastructure of High-Density Three-Dimensional Integrated Circuits (HD 3D-IC) present a new test challenges because of the high interconnect density and the area cost for test features. In this work, we firstly present a pre-analysis of the testability of HD 3D-IC; we define the minimum acceptable 3D pitch value for a given technology to ensure the circuits testability. Afterwards, we propose an optimized DFT architecture allowing pre-bond and post-bond test for SRAM/Logic HD 3D-IC in line with the ongoing IEEE P1838 standard. Imed Jani, Didier Lattard, Pascal Vivet, Jean Durupt, Sébastien Thuries, Edith Beigné |
ETS | 3 |
| 2019 | Memory Sizing of a Scalable SRAM In-Memory Computing Tile Based ArchitectureabstractModern computing applications require more and more data to be processed. Unfortunately, the trend in memory technologies does not scale as fast as the computing performances, leading to the so called memory wall. New architectures are currently explored to solve this issue, for both embedded and off-chip memories. Recent techniques that bringing computing as close as possible to the memory array such as, In-Memory Computing (IMC), Near-Memory Computing (NMC), Processing-In-Memory (PIM), allow to reduce the cost of data movement between computing cores and memories. For embedded computing, In-Memory Computing scheme presents advantageous computing and energy gains for certain class of applications. However, current solutions are not scaling to large size memories and high amount of data to compute. In this paper, we propose a new methodology to tile a SRAM/IMC based architecture and scale the memory requirements according to an application set. By using a high level LLVM-based simulation platform, we extract IMC memory requirements for a certain class of applications. Then, we detail the physical and performance costs of tiling SRAM instances. By exploring multi-tile SRAM Place&Route in 28nm FD-SOI, we explore the respective performance, energy and cost of memory interconnect. As a result, we obtain a detailed wire cost model in order to explore memory sizing trade-offs. To achieve a large capacity IMC memory, by splitting the memory in multiple sub-tiles, we can achieve lower energy (up to 78% gain) and faster (up to 49% gain) IMC tile compared to a single large IMC memory instance. Roman Gauchi, Maha Kooli, Pascal Vivet, Jean-Philippe Noël, Edith Beigné, Subhasish Mitra, Henri-Pierre Charles |
VLSI-SoC | 3 |
| 2018 | BISTs for post-bond test and electrical analysis of high density 3D interconnect defectsabstractCu-Cu hybrid bonding offers very high density interconnects (pitch around 2 μm or less) in 3D stacking integrated circuits (HD 3D-IC), but the smaller the Cu pad size, the more the fabrication and bonding defects have an important impact on yield and performance. Defects such as bonding misalignment, micro-voids and contact defects at the copper surface, can affect the electrical characteristics and the life time of 3D-IC considerably. In this paper, we propose two complementary test and characterization structures dedicated to high density 3D-IC interconnects. The first test structure permits to measure the misalignment defect with a great accuracy and the second to measure the RC delay of a periodic signal applied to a daisy chain composed of 3D Cu-Cu interconnects. The measured misalignment values and propagation delays allows to detect Cu-Cu full open, misalignment, and micro-voids, in order to assess performance of high density 3D Integrated Circuit. Both test structures are implemented as BIST engines, which are integrated and controlled with IEEE 1687, for an overall negligible area cost. Imed Jani, Didier Lattard, Pascal Vivet, Lucile Arnaud, Edith Beigné |
ETS | 3 |
| 2016 | IJTAG supported 3D DFT using chiplet-footprints for testing multi-chips active interposer systemabstractIn order to increase the circuit yield, 2.5D technology have been introduced to partition a single large circuit in multiple circuits, which are tested before bonding and then assembled in 3D onto a passive silicon interposer. Active interposer is nowadays envisioned in order to provide added values within the interposer and the 3D complete system. The testability of 2.5D interposers have already been well studied but no 3D DFT have already been proposed for active interposers. In this paper, a 3D Design-for-Test architecture is proposed for testing multi-chips stacked onto an active interposer. The 3D-DFT is based on a chiplet footprint architecture, allowing the modular test of any chiplets, and is implemented using IJTAG IEEE1687 standard, offering easy test pattern retargeting from chiplet pre-bond test to the 3D circuit final test. The proposed 3D-DFT architecture and the associated 3D test flow have been fully applied onto a 3D active interposer circuit prototype and used extensively to test interposer active links, interposer passive links and all embedded memory BIST engines. Jean Durupt, Pascal Vivet, Jürgen Schlöffel |
ETS | 2 |
| 2016 | Thermal Analysis and Interpolation Techniques for a Logic + WideIO Stacked DRAM Test ChipabstractSelf-heating and high-operating temperature are major concerns in 3-D-chip integration. In this paper, we leverage a 3-D test chip (WideIO dynamic random access memory on top of a logic die) equipped with temperature sensors and heaters to explore thermal effects and to develop advanced thermal modeling strategies suitable for complex 3-D-stacked circuits. We correlate temperature measurements with the power dissipated by the heaters using model learning techniques. Moreover, we defined a thermal basis function obtained using power and thermal data available from the on-chip sensors. This function can be used to predict temperatures at chip locations far from the temperature sensors and to infer the power dissipation at any location of the chip. In addition, the same thermal basic function can be used jointly with formal interpolation frameworks like radial basis function methods to effectively estimate the full-chip thermal map. Results show that this methodology outperforms existing interpolation approaches for sparse integrated sensors. Francesco Beneventi, Andrea Bartolini, Pascal Vivet, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Two-phase protocol converters for 3D asynchronous 1-of-n data linksabstractDesign of fully synchronous System on Chip is becoming a challenging task. This task is even more difficult in advanced nodes and 3D designs, where the local and global variability can turns the timing closure an overwhelming task. In this way, the use of asynchronous circuits for long link and 3D link communication can provide better robustness to both local and inter-die variability and achieve faster timing closure by extending the Globally Asynchronous Locally Synchronous style to 3D architectures. However, while the four-phase protocol is well adapted for on chip Delay Insensitive communication, it cannot be adapted for off chip and 3D interface communication due to potential large interface delays. In this paper, we propose the use of two-phase Delay Insensitive Transition Signaling for 1-of-n codes as well as new four ↔ two-phase data converters. The proposed circuitry is able to reduce 20% the dynamic power and improve two times the four-phase throughput for long link communications. Julian J. H. Pontes, Pascal Vivet, Yvain Thonnart |
ASP-DAC | 2 |
| 2015 | Retention time measurements and modelling of bit error rates of WIDE I/O DRAM in MPSoCs
Christian Weis, Matthias Jung 0001, Peter Ehses, Cristiano Santos, Pascal Vivet, Sven Goossens, Martijn Koedam, Norbert Wehn |
DATE | 5 |
| 2015 | Fine-grain DVFS and AVFS techniques for complex SoC design: An overview of architectural solutions through technology nodesabstractIn this paper we propose to give an overview of fine-grain design techniques we demontrated past years in our lab for power reduction in complex SoCs. Those works are based on Globally Asynchronous and Locally Synchronous systems in which each IP is an independent voltage and frequency domain. After having proposed some simple DFS architectures based on GALS architectures in 130nm technology, we extended our works to fine-grain Dynamic Voltage and Frequency Scaling architectures to reduce dynamic and static power reduction at 65 nm node. Furthermore, considering 32 nm deep submicron technologies, we demonstrated an Adaptive Voltage and Frequency architecture to compensate for in-die PVT variations. Area overhead and power reduction results are discussed all along the paper. Edith Beigné, Fabien Clermidy, Didier Lattard, Ivan Miro Panades, Yvain Thonnart, Pascal Vivet |
ISCAS | 6 |
| 2015 | A simulation framework for rapid prototyping and evaluation of thermal mitigation techniques in many-core architecturesabstractModern SoCs are characterized by increasing power density and consequently increasing temperature, that directly impacts performances, reliability and cost of a device through its packaging. Thermal issues need to be predicted and mitigated as early as possible in the design flow, when the optimization opportunities are the highest. In this paper, we present an efficient framework for the design of dynamic thermal mitigation schemes based on a high-level SystemC virtual prototype tightly coupled with efficient power and thermal simulation tools. We demonstrate the benefit of our approach through silicon comparison with the SThorm 64-core architecture and provide simulation speed results making it a sound solution for the design of thermal mitigation early in the flow. Tanguy Sassolas, Chiara Sandionigi, Alexandre Guerre, Julien Mottin, Pascal Vivet, Hela Boussetta, Nicolas Peltier |
ISLPED | 5 |
| 2014 | The improbable but highly appropriate marriage of 3D stacking and neuromorphic acceleratorsabstract3D stacking is a promising technology (low latency/power/area, high bandwidth); its main shortcoming is increased power density. Simultaneously, motivated by energy constraints, architectures are evolving towards greater customization, with tasks delegated to accelerators. Due to the widespread use of machine-learning algorithms and the re-emergence of neural networks (NNs) as the preferred such algorithms, NN accelerators are receiving increased attention. They turn out to be well matched to 3D stacking: inherently 3D structures with a low power density and high across-layer bandwidth requirements. We present what is, to the best of our knowledge, the first 3D stacked NN accelerator Bilel Belhadj, Alexandre Valentian, Pascal Vivet, Marc Duranton, Liqiang He, Olivier Temam |
CASES | 3 |
| 2014 | Thermal analysis and model identification techniques for a logic + WIDEIO stacked DRAM test chipabstractHigh temperature is one of the limiting factors and major concerns in 3D-chip integration. In this paper we use a 3D test chip (WIDEIO DRAM on top of a logic die) equipped with temperature sensors and heaters to explore thermal effects. We correlated real temperature measurements with the power dissipated by the heaters using model learning techniques. The resulting compact thermal model is able to predict temperatures at chip locations far from the temperature sensors and to infer the power dissipation at any location of the chip. Results are verified by mean of an off-sample validation technique and show a high accuracy of the compact thermal model when compared with silicon measurements. Francesco Beneventi, Andrea Bartolini, Pascal Vivet, Denis Dutoit, Luca Benini |
DATE | 3 |
| 2014 | Early design stage thermal evaluation and mitigation: The locomotiv architectural caseabstractTo offer more computing power to modern SoCs, transistors keep scaling in new technology nodes. Consequently, the power density is increasing, leading to higher thermal risks. Thermal issues need to be addressed as early as possible in the design flow, when the optimization opportunities are the highest. For early design stages, architects rely on virtual prototypes to model their designs' behavior with an adapted trade-off between accuracy and simulation speed. Unfortunately, accurate virtual prototypes fail to encompass thermal effects timescale. In this paper, we demonstrate that less accurate high-level architectural models, in conjunction with efficient power and thermal simulation tools, provide an adapted environment to analyze thermal issues and design software thermal mitigation solutions in the case of the Locomotiv MPSoC architecture. Tanguy Sassolas, Chiara Sandionigi, Alexandre Guerre, Alexandre Aminot, Pascal Vivet, Hela Boussetta, Luca Ferro, Nicolas Peltier |
DATE | 5 |
| 2013 | A dynamic stream link for efficient data flow control in NoC based heterogeneous MPSoCabstractAs Systems-on-Chip size increase, the communication costs become critical and Networks-on-Chip (NoC) bring innovative solutions. Efficient stream-based protocols over NoC have been widely studied to address dataflow communications. They are usually controlled by a set of static parameters. However, new applications, such as high-resolution video decoders, present more data-dependent behaviors forcing communication protocols to support higher dynamicity. For this purpose, we present in this paper dynamic stream links for stream-based end-to-end NoC communications by introducing two link protocols, both independent of the transfer size, allowing to improve the hardware/software control flexibility. The proposed protocols have been modeled in a MPSoC virtual platform and the hardware cost evaluated. Based on simulations, we provide guidelines to exploit these protocols according to application needs. Claude Helmstetter, Sylvain Basset, Romain Lemaire, Fabien Clermidy, Pascal Vivet, Michel Langevin, Chuck Pilkington, Pierre G. Paulin, Didier Fuin |
ASP-DAC | 5 |
| 2013 | Fast and accurate TLM simulations using temporal decoupling for FIFO-based communicationsabstractA known approach to improve the timing accuracy of an untimed or loosely timed TLM model is to add timing annotations into the code and to reduce the number of costly context switches using temporal decoupling, meaning that a process can go ahead of the simulation time before synchronizing again. Our current goal is to apply temporal decoupling to the TLM platform of a heterogeneous many-core SoC dedicated to high performance computing. Part of this SoC communicates using classic memory-mapped buses, but it can be extended with hardware accelerators communicating using FIFOs. Whereas temporal decoupling for memory-based transactions has been widely studied, FIFO-based communications raise issues that have not been addressed before. In this paper, we provide an efficient solution to combine temporal decoupling and FIFO-based communications. Claude Helmstetter, Jérôme Cornet, Bruno Galilée, Matthieu Moy, Pascal Vivet |
DATE | 5 |
| 2013 | Advances in asynchronous logic: from principles to GALS & NoC, recent industry applications, and commercial CAD toolsabstractThe growing variability and complexity of advanced CMOS technologies makes the physical design of clocked logic in large Systems-on-Chip more and more challenging. Asynchronous logic has been studied for many years and become an attractive solution for a broad range of applications, from massively parallel multi-media systems to systems with ultra-low power & low-noise constraints, like cryptography, energy autonomous systems, and sensor-network nodes. The objective of this embedded tutorial is to give a comprehensive and recent overview of asynchronous logic. The tutorial will cover the basic principles and advantages of asynchronous logic, some insights on new research challenges, and will present the GALS scheme as an intermediate design style with recent results in asynchronous Network-on-Chip for future Many Core architectures. Regarding industrial acceptance, recent asynchronous logic applications within the microelectronics industry will be presented, with a main focus on the commercial CAD tools available today. Alexandre Yakovlev, Pascal Vivet, Marc Renaudin |
DATE | 2 |
| 2013 | Computing detection probability of delay defects in signal line tsvsabstractThree-dimensional stacking technology promises to solve the interconnect bottleneck problem by using Through-Silicon-Vias (TSVs) to vertically connect circuit layers. However, manufacturing steps may lead to partly broken or incompletely filled TSVs that may degrade the performance and reduce the useful lifetime of a 3D IC. Due to combinations of physical factors such as switching activity, supply noise and crosstalk, path delays can experience speed-up or slow-down that could let the effect of resistive open TSV go undetected by conventional test methods. In this work, we present a metric based on probabilistic analysis to detect delay defects induced by resistive opens that occur on signal line TSVs. Our experimental result will show the accuracy of the proposed metric. Carolina Metzler, Aida Todri, Alberto Bosio, Luigi Dilillo, Patrick Girard 0001, Arnaud Virazel, Pascal Vivet, Marc Belleville |
ETS | 7 |
| 2013 | Parity check for m-of-n delay insensitive codesabstractThe advance in deep submicron technologies brings new constraints to circuit design such as variability and sensitivity to soft errors. Asynchronous networks on chip can help coping with some of these constraints due to the timing robustness of design paradigms such as the quasi delay insensitive one. A relevant problem of current fully asynchronous networks on chip is the lack of mechanisms to provide error detection and correction in asynchronous data communication. This work proposes a parity scheme applicable to m-of-n delay insensitive codes, which is capable to correct errors caused by single event effects in delay insensitive communication architectures. The proposed mechanism was evaluated in a 65nm technology where it is able to correct 98% of the errors caused by single event effects with a low overhead in terms of area, power and performance. Julian J. H. Pontes, Ney Laert Vilar Calazans, Pascal Vivet |
IOLTS | 3 |
| 2013 | 3D stacking for multi-core architectures: From WIDEIO to distributed cachesabstract3D stacking has been viewed as a breakthrough solution for increasing performance in multi-core architectures. The hope is to solve some of the main issues in current multi-core architectures: external memory pressure and latency; I/O bottleneck; communication power consumption. In this paper, some advances of this field of research are shown, starting with a WIDEIO experience on a real chip for solving DRAM accesses issue. The integration of a 512 bit-width bus is demonstrated in a Network-on-Chip (NoC) multi-core framework and the resulting performance based on a 65nm prototype with 10μm diameter Through Silicon Vias (TSV). The potentiality of 3D scaling thanks to 3D asynchronous Network-on-Chip implementation is then shown. Finally, an innovative 3D stacked distributed cache strategy aimed at lowering memory latency and external memory bandwidth requirements is presented. This new memory partitioning demonstrates the efficiency of 3D stacking to rethink architectures for addressing multi-core scaling challenges. Fabien Clermidy, Denis Dutoit, Eric Guthmuller, Ivan Miro Panades, Pascal Vivet |
ISCAS | 5 |
| 2012 | An accurate Single Event Effect digital design flow for reliable system level designabstractSimilar to local variations and signal integrity problems, Single Event Effects (SEEs) are a new design concern for digital system design that arises in deep sub-micron technologies. In order to design reliable digital systems in such technologies, it is mandatory to precisely model and take into account SEEs. This paper proposes a new accurate design flow to model non-permanent SEE effects that can be applied at system level for reliable digital circuit design. Starting from low level SPICE-accurate simulations, SEEs are characterized, modeled and simulated in the digital design using commercial and well accepted standards and tools. The proposed design flow has been fully validated through a complete digital design, a cryptographic core implemented in a 32nm CMOS technology. Finally, using the SEE design flow, the paper presents some reliability impact analysis, both at standard cell level and design level. Julian J. H. Pontes, Ney Laert Vilar Calazans, Pascal Vivet |
DATE | 3 |
| 2011 | 3D Embedded multi-core: Some perspectivesabstract3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range products. In this paper, we investigate three promising perspectives for short to medium terms adoption of such technology in high-end System-on-Chip built around multi-core architectures: the wide bus concept will help solving high bandwidth requirements with external memory. 3D Network-on-Chip is a promising solution for increased modularity and scalability. We show that an efficient implementation provides an available bandwidth outperforming classical interfaces. Finally, we put in perspective the active interposer concept which aims at simplifying and improving power, test and debug management. Fabien Clermidy, Florian Darve, Denis Dutoit, Walid Lafi, Pascal Vivet |
DATE | 5 |
| 2011 | 3D NoC using through silicon Via: An asynchronous implementationabstract3D stacking is seen as one of the most interesting technologies for System-on-Chip (SoC) developments. However, 3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range of products. In this paper, we are investigating 3D Network-on-Chip has a promising solution for increased modularity and scalability. We show that an efficient implementation based on asynchronous logic provides an available bandwidth of 64GB/s for only 700 TSV, outperforming classical interfaces while simplifying the assembly process. We also point out the benefit in terms of power consumption for these new interfaces with a gain of 5 times compared to classical LPDDR2 interfaces. Pascal Vivet, Denis Dutoit, Yvain Thonnart, Fabien Clermidy |
VLSI-SoC | 1 |
| 2010 | A fully-asynchronous low-power framework for GALS NoC integrationabstractRequiring more bandwidth at reasonable power consumption, new communication infrastructures must provide adequate solutions to guarantee performance during physical integration. In this paper, we propose the design of a low-power asynchronous Network-on-Chip which is implemented in a bottom-up approach using optimized hard-macros. This architecture is fully testable and a new design flow is proposed to overcome CAD tools limitations regarding asynchronous logic. The proposed architecture has been successfully implemented in CMOS 65nm in a complete circuit. It achieves a 550Mflit/s throughput on silicon, and exhibits 86% power reduction compared to an equivalent synchronous NoC version. Yvain Thonnart, Pascal Vivet, Fabien Clermidy |
DATE | 2 |
| 2009 | A Communication and configuration controller for NoC based reconfigurable data flow architectureabstractWhile network-on-chip aspects such as topologies, routing strategies or quality-of-service have been largely studied, the mapping of real applications on distributed NoC-based architecture is still an open issue. In this paper, we address this issue for complex reconfigurable data-flow applications. We introduce the concept of communication and configuration controller (CCC) which interacts both with the usual network interface and the IP core structure. The proposed CCC is a programmable template-based architecture, which provides solutions to manage reconfiguration flows, data synchronizations and global control signaling. An implementation of the CCC is presented, and its performances in a 65 nm technology are discussed through a concrete telecommunication application. Fabien Clermidy, Romain Lemaire, Yvain Thonnart, Pascal Vivet |
NOCS | 4 |
| 2009 | Power Reduction of Asynchronous Logic Circuits Using Activity DetectionabstractAsynchronous circuits are well known for their benefits in terms of dynamic power savings because asynchronous logic does not switch when inactive. Nevertheless, in deep-submicron technologies, leakage currents have become an increasing issue, and thus, asynchronous circuits need to focus on static-power-consumption reduction. In this paper, we propose an innovative way to detect incoming asynchronous activity. Associated to an automatic power regulation, it efficiently reduces the supply voltage and, thus, both energy per operation and leakage currents. The proposed technique has been applied to an asynchronous network-on-chip node and successfully implemented in an ST Microelectronics CMOS 65-nm technology. Yvain Thonnart, Edith Beigné, Alexandre Valentian, Pascal Vivet |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Dynamic Voltage and Frequency Scaling Architecture for Units Integration within a GALS NoC
Edith Beigné, Fabien Clermidy, Sylvain Miermont, Pascal Vivet |
NOCS | 4 |
| 2008 | Physical Implementation of the DSPIN Network-on-Chip in the FAUST Architecture
Ivan Miro Panades, Fabien Clermidy, Pascal Vivet, Alain Greiner |
NOCS | 3 |
| 2007 | ASC, a SystemC Extension for Modeling Asynchronous Systems, and Its Application to an Asynchronous NoCabstractThis paper presents ASC, an Asynchronous SystemC library, as an extension of SystemC for modeling asynchronous circuits. ASC includes a set of port and channel primitives offering the same communication primitives as the common languages used for asynchronous circuits modeling (CHP, Tangram or Balsa). ASC also offers operators and statements in order to accurately model arbiters, which are the basic components of asynchronous network on chips. The aim of this work is to provide to the designers the means of modeling and verifying asynchronous circuits as well as GALS and NoC systems. Synthesis of ASC models with the help of the TAST framework is under development. As an illustrative example, the modeling of an asynchronous network-on-chip architecture using the ASC library is described. This NoC has been successfully integrated into a complex GALS NoC architecture taking advantage of a multi-level SystemC based verification environment Cedric Koch-Hofer, Marc Renaudin, Yvain Thonnart, Pascal Vivet |
NOCS | 4 |