VLDB 2026 Research / reviewers in the wild / expert
Julien Ryckaert
dblp:85/5323
· DBLP profile ↗
32ranked-venue papers
1as first author
19since 2021 · last 2026
0009-0001-0140-3042ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 1 first-author · 19 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Computer networks · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thermal Insights of 3-D BS-PDN in Cloud Server SoC Using TCAD ModelingabstractIn this brief, the thermal performance of a large-scale cloud server system-on-chip (SoC) with the backside power delivery network (BS-PDN) and 3-D integration in memory-on-logic (MoL)/logic-on-memory (LoM) configuration with 2.5-D packaging is analyzed in advanced A10 nanosheet technology node using Sentaurus TCAD platform. The results show a 45.6% (~20.3 K) thermal penalty for the 80-core SoC in MoL with BS-PDN compared with the 2-D-baseline frontside PDN (FS-PDN), using a heatsink with forced cooling. A nonuniform power map further aggravates thermal concerns, which can be mitigated using an LoM configuration with BS-PDN, reducing the penalty to 22% (~15 K). Extending the study to a 320-core SoC, in conjunction with an advanced cooling system, LoM with BS-PDN shows 45.3% (~29 K) lower temperature than conventional MoL BS-PDN. The modeling results provide valuable insights and motivate future research into packaging and cooling techniques for BS-PDN integration. Subrat Mishra, Herman Oprins, James Myers, Julien Ryckaert, Pieter Woltgens, Dwaipayan Biswas |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Late Breaking Results: Thermal Feasibility of Backside Integrated LDOs in 2.5D/3D System-in-Package Using Nanosheet TechnologyabstractDigital Low Dropout Regulators (LDOs) are an excellent candidate for area-efficient fine-grain power management in heterogeneous systems, leveraging integrated power switches. Relocating the power switches to the backside of the wafer in conjunction with the Backside Power Delivery Network (BSPDN) layer is envisaged as a System Technology Co-Optimization (STCO) booster for finer grain power management and reduced area/cost. We perform a detailed thermal analysis using power-switch-based LDOs enabling per-core DVFS for a high-performance server 3D computing chiplet in a Nanosheet CMOS (A10) technology node with BSPDN. While BSPDN introduces thermal penalties due to a lack of lateral heat spreading, our high-resolution thermal simulations explore the feasibility of moving LDOs to the backside. Increasing the LDO area from 5% to 50% of the backside die area effectively lowers the 2.5/3D System-in-Package (SiP) peak temperature, confirming that thermal concerns do not impede backside LDO integration. This study supports the cost-effective design of next-generation SiPs by demonstrating no adverse thermal impact for relocating power switches to the wafer backside in the nanosheet era. Yukai Chen, Subrat Mishra, Julien Ryckaert, Dwaipayan Biswas, James Myers |
DATE | 3 |
| 2025 | Invited Paper: CMOS 2.0 - Redefining the Future of ScalingabstractWe propose to revisit the functional scaling paradigm by capitalizing on two recent developments in advanced chip manufacturing, namely 3D wafer bonding and backside processing. This approach leads to the proposal of the CMOS 2.0 platform. The main idea is to shift the CMOS roadmap from geometric scaling to fine-grain heterogeneous 3D stacking of specialized active device layers to achieve the ultimate Power-Performance-Area and Cost gains expected from future technology generations. However, the efficient utilization of such a platform requires devising architectures that can optimally map onto this technology, as well as the EDA infrastructure that supports it. We also discuss reliability concerns and eventual mitigation approaches. This paper provides pointers into the major disruptions we expect in the design of systems in CMOS 2.0 moving forward. Moritz Brunion, Navaneeth Kunhi Purayil, Francesco Dell'Atti, Sebastian Lam, Refik Bilgic, Mehdi Baradaran Tahoori, Luca Benini, Julien Ryckaert |
ICCAD | 8 |
| 2025 | Half-Height Double-Row CFET Standard Cells for Area Optimized Placement in A7 CMOS NodeabstractComplementary FET (CFET) is a promising device architecture that proceeds the CMOS scaling during the post-nanosheet device era. Among several CFET variants, Double-Row (DR) CFET further enables 15% track height scaling on standard cells, by sharing a middle row of vias, while sustaining an optimized Middle-Of-Line (MOL) process complexity. In this study, half-height double-row (hDR) CFET is proposed as a highly practical and impactful design style to overcome the cell and block level limitations of DR CFET architecture. First, hDR CFET introduces a high flexibility on standard cell layout design with significant area optimization. Secondly, hDR cell insertion in the backend physical design flow further optimizes the cell placement, and recovers block level area scaling to match the cell height scaling. Results on A7 CFET technology library show area reduction up to 50% on standard cell layouts. Moreover, block level PnR results after enabling only 6 types of hDR cells in standard cell library show 10% of area scaling on ARM Cortex-M0 32-bit core at 90% utilization, proving the strength of the concept. Finally, 14% of block level area scaling is further projected for an enriched standard cell library with an extended set of hDR cells. Halil Kukner, Ji-Yung Lin, Lynn Verschueren, Jürgen Bömmels, Anita Farokhnejad, Maarten Van De Put, Odysseas Zografos, Naoto Horiguchi, Geert Hellings, Marie Garcia Bardon, Julien Ryckaert |
ICCAD | 12 |
| 2025 | Framework for Augmenting Main Memory with CXL-connected Emerging Memory AlternativesabstractThe rapid evolution of memory technologies and the advent of Compute Express Link (CXL) have opened up new possibilities for scaling main memory by enabling hybrid memory systems with pooled and shared content. System-level evaluation of new memory systems during the early development stage is important for the enablement and further integration of new memory and interconnect technologies. However, existing solutions do not offer a framework neither for emerging memory protocols nor for novel memory technologies. This paper introduces CXL-HMEM to evaluate emerging CXL-based hybrid main memory architectures by applying the System Technology Co-Optimization (STCO) technique. The framework provides flexible performance metrics, workload simulation, and memory traffic analysis to assess system performance under various hybrid memory configurations, including DRAM and tiered memory hierarchies. Key features include support for memory technologies such as IGZO-based DRAM (IGZO) and FeRAM, workload scalability, and an integrated model of the CXL behavior. CXL-HMEM shows that emerging memories can improve system bandwidth and energy consumption by 7%, while having potential to further mitigate particular bottlenecks. CXL-based hybrid main memory can speed up the memory access time by >2× compared to conventional approaches of main memory extension. Khakim Akhunov, Dwaipayan Biswas, Emil Karimov, Arvind Sharma, Hyungrock Oh, Maarten Rosmeulen, Julien Ryckaert, James Myers |
ISCAS | 7 |
| 2025 | 3D SRAM Disaggregation in Advanced CMOS Nodes using Hybrid Bonding TechnologyabstractThis paper studies the potential of hybrid bonded Array-under-CMOS (AuC) technology to partition logic and high-performance L1 cache in advanced technology nodes. By decoupling the SRAM bitcells from the logic tier, we achieve independent optimization of both SRAM and logic devices, as well as back-end-of-line (BEOL) interconnects. Heterogeneous integration and BEOL aspect ratio optimization is implemented with different technology nodes to mitigate the delay penalty due to hybrid bond pad staggering. Addressing the performance degradation associated with scaled technology nodes, we investigate the impact of word-line (WL) and bit-line (BL) resistance on SRAM performance. Leveraging the flexibility of decoupled SRAM BEOL and within the AuC technology framework, we explore the sensitivity to WL and BL metal aspect ratios, comparing their performance against a 2D baseline. Our results demonstrate substantial performance improvements in AuC integration through two key approaches: (1) 5% enhancement via Back-End-of-Line (BEOL) optimization, and (2) 25% improvement enabled by heterogeneous integration, achieved by decoupling memory and logic tiers. Bhawana Kumari, Anurag Swarnkar, Dawit Burusie Abdi, Fernando García-Redondo, James Myers, Julien Ryckaert, Jaydeep P. Kulkarni, Dwaipayan Biswas |
ISCAS | 6 |
| 2025 | 3D IGZO Charge-Coupled Memory DTCO & STCO Analysis for Compute-near-Memory ApplicationsabstractThe demand for high-capacity and energy-efficient memory solutions has surged in the era of data-centric computing, particularly for Artificial Intelligence (AI) and Machine Learning (ML) workloads. This paper introduces a novel memory architecture leveraging Charge-Coupled Device (CCD) technology, engineered in a sequential-access block memory configuration, to enhance Compute-near-Memory (CnM) systems. We propose an optimized 3D IGZO CCD block memory as an on-chip weight buffer for high-capacity CnM systems. Our approach achieves 2.95−131.26× improvement in area efficiency and 1.32−4.33× improvement in energy efficiency compared to SRAM solutions. Khakim Akhunov, Hyungrock Oh, Fernando García-Redondo, Yukai Chen, Arvind Sharma, Jiacong Sun, Sahan Gamage, Maarten Rosmeulen, Swaraj Bandhu Mahato, Rishabh Kishore, Subhali Subhechha, Jaydeep P. Kulkarni, Marian Verhelst, Dwaipayan Biswas, Marie Garcia Bardon, Wim Dehaene, Julien Ryckaert |
ISCAS | 18 |
| 2025 | Live Demonstration: N2 Nanosheet Pathfinding-PDKabstractCMOS scaling involves more than just reducing the effective channel length—it has become increasingly complex. Aligning the circuit design ecosystem with semiconductor technology is now more critical than ever, giving rise to the Design-Technology Co-Optimization (DTCO) concept and further extending Moore’s Law. This live demonstration presents our previously published N2 nanosheet Pathfinding PDK (P-PDK) which provides insight into cutting-edge CMOS technology under the framework of DTCO. This session will cover the motivation behind developing this cutting-edge PDK, the basic flow of using the P-PDK, and the prospects of the P-PDK. Chaohan Wang, Jack Cousins, Anita Farokhnejad, Marie Garcia Bardon, Julien Ryckaert |
ISCAS | 5 |
| 2025 | Bandwidth-Latency-Thermal Co-Optimization of Interconnect-Dominated Many-Core 3D-ICabstractThe ongoing integration of advanced functionalities in contemporary system-on-chips (SoCs) poses significant challenges related to memory bandwidth, capacity, and thermal stability. These challenges are further amplified with the advancement of artificial intelligence (AI), necessitating enhanced memory and interconnect bandwidth and latency. This article presents a comprehensive study encompassing architectural modifications of an interconnect-dominated many-core SoC targeting the significant increase of intermediate, on-chip cache memory bandwidth and access latency tuning. The proposed SoC has been implemented in 3-D using A10 nanosheet technology and early thermal analysis has been performed. Our workload simulations reveal, respectively, up to 12- and 2.5-fold acceleration in the 64-core and 16-core versions of the SoC. Such speed-up comes at 40% increase in die-area and a 60% rise in power dissipation when implemented in 2-D. In contrast, the 3-D counterpart not only minimizes the footprint but also yields 20% power savings, attributable to a 40% reduction in wirelength. The article further highlights the importance of pipeline restructuring to leverage the potential of 3-D technology for achieving lower latency and more efficient memory access. Finally, we discuss the thermal implications of various 3-D partitioning schemes in High Performance Computing (HPC) and mobile applications. Our analysis reveals that, unlike high-power density HPC cases, 3-D mobile case increases$T_{\max }$only by$2~^{\circ } $C–$3~^{\circ } $C compared to 2-D, while the HPC scenario analysis requires multiconstrained efficient partitioning for 3-D implementations. Sudipta Das, Samuel Riedel, Mohamed Naeim, Moritz Brunion, Marco Bertuletti, Luca Benini, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | GNN-assisted Back-side Clock Routing Methodology for Advance TechnologiesabstractThe back-side metal layers exhibit lower parasitics compared to the front-side layers in advanced technologies, making them suitable for clock-net distribution. In this study, we explore the advantages of using back-side metal layers for clock routing, which is shared with a power delivery network. Our Graph Neural Network (GNN) based framework, effectively distributes the clock-tree between the front and back sides. We address the back-side clock nets' creation by incorporating back-side buffers. Our results demonstrate better clock and full-chip metrics represented by an increase of up to 13% in the effective frequency with equivalent power consumption, using 3 nm technology. Nesara Eranna Bethur, Pruek Vanna-Iampikul, Odysseas Zografos, Lingjun Zhu, Giuliano Sisto, Dragomir Milojevic, Alberto García Ortiz, Geert Hellings, Julien Ryckaert, Francky Catthoor, Sung Kyu Lim |
DAC | 9 |
| 2024 | Smoothing Disruption Across the Stack: Tales of Memory, Heterogeneity, & CompilersabstractMultiple research vectors represent possible paths to improved energy and performance metrics at the application-level. There are active efforts with respect to emerging logic devices, new memory technologies, novel interconnects, and heterogeneous integration architectures. Of great interest is quantifying the potential impact of a given solution to prioritize research vectors accordingly. In this paper, we discuss two efforts - one focused on emerging memory technology, and another focused on heterogeneous integration technology - that speak to best practices for, and needed contributions from the design automation (DA) community to explore this vast design space. Furthermore, we highlight new research efforts that aim to develop the novel compiler abstractions and frameworks that are ultimately needed to derive maximum value from new memory and/or heterogeneous and monolithic integration architecture, and that can also play an important role with respect to design space exploration efforts. Michael T. Niemier, Zephan M. Enciso, M. Sharifi, Xiaobo Sharon Hu, Ian O'Connor, A. Graening, Jerónimo Castrillón, João Paulo C. de Lima, Asif Ali Khan, Hamid Farzaneh, N. Afroze, Julien Ryckaert |
DATE | 15 |
| 2024 | 3D Partitioning with Pipeline Optimization for Low-Latency Memory Access in Many-Core SoCsabstractThis paper presents an investigation of System-on-Chip (SoC) communication latency optimization for 3D system integration and highlights the role of architectural modifications to maximize the Power, Performance, & Area (PPA) benefits. An instance of a highly configurable RISC-V SoC is implemented using ∼2nm nanosheet technology and different 3D stacking options using design flow from sign-off tools. The proposed implementation targets performance optimization for different 3D partitioning scenarios: Memory-on-Logic (MoL) & Logic-on-Logic (LoL). We target 2-die 3D Integrated Circuits (3D-IC) with high density 3D interconnect using Face-to-Face (F2F) hybrid bonding (∼1µm), and 3-die stack, as Face-to-Back (F2B) on top of F2F. Our analysis of the 16-core SoC instance shows that the proposed architectural optimizations bring a significant reduction of 4 pipeline stages in the design hierarchy at a marginal cost of 9% effective frequency loss when implemented in 3D in comparison to the baseline 2D architecture. Further, going from 2D to 3D allows more than 40% total system wire-length reduction & 10% less cell area, resulting in 20% power savings. These findings hold promise for further explorations on many-core SoC instances (256 & more) facing system interconnect challenges. Sudipta Das, Samuel Riedel, Marco Bertuletti, Luca Benini, Moritz Brunion, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic |
ISCAS | 6 |
| 2024 | Future Design Direction for SRAM Data Array: Hierarchical Subarray With Active InterconnectabstractIn sub 10 nm nodes, the growing dominance of interconnects in chips poses challenges in designing large-size static random-access memory (SRAM) subarrays. The main issue is the write failure problem arising from the increased resistance and capacitance for bitline (BL) and wordline (WL). To tackle this issue, the SRAM subarray design incorporates conventional (Conv.) divided WL and divided BL techniques based on 14-Å-compatible (A14) nanosheet (NS) technology. This approach allows for various subarray sizes with successful write operations, resulting in improved subarray-level performance and power (PP). However, the additional logic gates come with an area penalty that may degrade the overall performance, power, and area (PPA) at the macro level due to increased inter-subarray interconnect overhead. To overcome this limitation, the active interconnect (AIC) design is proposed with the features of fabricating another or multiple active regions at the back-end of line (BEOL) layers. By moving these extra logic gates from front-end of line to BEOL in the AIC divided subarray design, the area penalty is significantly mitigated without compromising PP compared to the standard (Std.) and Conv. divided counterparts. To achieve this concept, carbon nanotube gate-all-around transistor is explored as potential BEOL-compatible device. In this research, a comprehensive design-technology co-optimization analysis is conducted to verify the value and potential benefits of up to 65% macro-level energy-delay-area product improvement by AIC divided subarray design compared to the Std. subarray design. Hsiao-Hsuan Liu, Carlo Gilardi, Shairfe Muhammad Salahuddin, Zhenlin Pei, Pieter Schuddinck, Pieter Weckx, Geert Hellings, Marie Garcia Bardon, Julien Ryckaert, Chenyun Pan, Subhasish Mitra, Francky Catthoor |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | Multidie 3-D Stacking of Memory Dominated Neuromorphic ArchitecturesabstractEvent-driven neuromorphic processors for artificial intelligence (AI) inference on edge/IoT devices require largeon-chip memory capacity, for efficient execution of spiking neural networks (NNs). In this work, we evaluate 3-D stacking benefits on SENECA, a digital neuromorphic accelerator core, sweeping itson-chip memory capacity from 2 up to 32 Mb in both legacy planar and advanced nanosheet CMOS logic nodes. In a planar CMOS node (GF-22 nm), two-die memory-on-logic (MoL) partitioning enables$8\times $moreon-chip memory, and it boosts operating frequency by 7% with 26% less power than the 2-D. Moving to an advanced nanosheet technology (imec A10), multidie (up to 7 dies) MoL stacking enables a performance increase of up to 29% and power savings up to 31%. Furthermore, a core folding (CF) partitioning in A10 shows up to 16% performance improvement with 12% total power savings with respect to the 2-D implementation on the same technology. We also demonstrate no thermal overhead for multidie stacking at advanced nodes for designs exhibiting low power density. These physical design explorations lay the foundation for system technology co-optimization studies for edge devices. Leandro M. G. Rocha, Refik Bilgic, Mohamed Naeim, Sudipta Das, Herman Oprins, Amirreza Yousefzadeh, Mario Konijnenburg, Dragomir Milojevic, James Myers, Julien Ryckaert, Dwaipayan Biswas |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2023 | Design Technology co-optimization of 1D-1VCMA to improve read performance for SCM applicationsabstract1-diode 1-Voltage controlled magnetic anisotropy (1D-1VCMA) can be an option for Storage Class Memory (SCM) to bridge the latency gap between DRAM and flash memory. It has low sneak current, high non-linearity and low IR drop. This paper presents the Design Technology Co-optimization (DTCO) study of 1D-1VCMA stack to improve the performance and energy. Thanks to precessional switching of VCMA, the write operation is very fast, but the read determines overall latency as read before write is needed to ensure reliable write operations. The read performance of 1D-1VCMA is penalized due to high VCMA MTJ resistance, hence impacting the overall performance. To improve the read performance, this paper explores two solutions: 1) reducing the VCMA RA product, and 2) improving the read circuit. These solutions improve the read performance by 36% and 260%, respectively. Mohit Gupta 0004, Manu Perumkunnil Komalan, Dwaipayan Biswas, Saeideh Alinezhad Chamazcoti, Gouri Sankar Kar, Arnaud Furnémont, Julien Ryckaert |
ISCAS | 7 |
| 2023 | Impact of interconnects enhancement on SRAM design beyond 5nm technology nodeabstractThis paper presents an extensive study of 6T-SRAM based on FinFET for advanced technology nodes beyond 5nm. We deduce that parasitic resistance becomes the main bottleneck for SRAM design at these nodes. SRAM's writing margin and read speed are impacted due to the increased Bit-Line (BL) and Word-Line (WL) resistance. This work primarily explores two possible solutions to improve the parasitic resistance at advanced process technology nodes: 1) strapping of BL and WL to higher metal, and 2) adopting the resistance optimized BEOL. Strapping BL and WL to higher metal layer improves the Write Trip Point (WTP) by ~100mV and the critical path delay by 24% at the cost of 50% higher energy. Resistance optimized BEOL can improve WTP by ~50mV more and delay by 25% more, at the cost of increased energy consumption (8%). Mohit Gupta 0004, Pieter Weckx, Manu Perumkunnil Komalan, Julien Ryckaert |
ISCAS | 4 |
| 2023 | 3D SRAM Macro Design in 3D Nanofabric Process TechnologyabstractIn this paper, we introduce a novel design of a 3D static random-access memory (SRAM) macro in a 3D Nanofabric process technology. The 3D Nanofabric technology is based on enabling the processing of N stack of identical layers simultaneously regardless of the number of stacked layers which consequently reduces the fabrication cost as well as the footprint of SRAM macros. To enable simultaneous patterning of stacked layers, 3D Nanofabric requires circuit topology and layout that rely on a single layer where the device channel, poly, and metal wires are all embedded without any other crossing than the gate on top of the device channel. Accordingly, we modify the layouts of the conventional SRAM bit-cell and periphery circuits which are complex and contain several metal crossings. Furthermore, we propose a new overall organization of the 3D SRAM macro that incorporates a stack of multiple identical layers each consisting of an equal size 2D array of bit-cells and the periphery circuits. We show that the proposed 3D Nanofabric SRAM macro offers 71.2% footprint gain and 36.3% read access speed improvement compared to equal size 2D SRAM macro in 3 nm FinFET. Dawit Burusie Abdi, Shairfe Muhammad Salahuddin, Jürgen Bömmels, Edouard Giacomin, Pieter Weckx, Julien Ryckaert, Geert Hellings, Francky Catthoor |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Design enablement of CFET devices for sub-2nm CMOS nodesabstractNovel devices that optimize their structure in a three-dimensional fashion and offer significant area gains by reducing standard cell track height are adopted to scale silicon technologies beyond the 5nm node. Such a device is the Complementary FET (CFET), which consists of an n-type channel stacked vertically over a p-type channel. In this paper we review the significant benefits of CFET devices as well as the challenges that arise with their use. More specifically, we focus on the standard cell design challenges as well as the physical implementation ones. We show that to fully exploit the area benefits of the CFET devices, one must carefully select the metal stack used for the physical implementation of a large design. Odysseas Zografos, Bilal Chehab, Pieter Schuddinck, Gioele Mirabelli, Naveen Kakarla, Pieter Weckx, Julien Ryckaert |
DATE | 8 |
| 2022 | Evaluation of Nanosheet and Forksheet Width Modulation for Digital IC Design in the Sub-3-nm EraabstractIn this article, we provide a comprehensive evaluation of width modulation capabilities of both nanosheet (NS) and forksheet (FS) devices, going from device level to a block level implementation. The main innovation introduced by the FS consists of a dielectric wall added between the p- and nMOS transistors. Leveraging this feature, FS shows approximately the same current behavior as NS, considered a state-of-the-art reference, but reduced parasitic capacitance thanks to its fewer but wider stacked sheets. At block level, an area reduction up to 12% is observed with FS, alongside a 13% power reduction and 10% frequency increase. Following the device comparison, the potential of sheet width modulation as additional power, performance, and area (PPA) optimization technique during synthesis and place and route (PNR) is investigated. A description of the specific steps required to enable this knob in a conventional electronic design automation (EDA) framework is provided. As demonstrated by the obtained experimental results, the same frequency of the single-width implementation can be achieved using mixed libraries with lower power consumption (13% and 16% for NS and FS, respectively), leading to improved energy efficiency. Furthermore, it is shown how designs implemented using FS benefit more from this type of optimization than the ones using NS, with a 12%–15% energy reduction compared to the 8.5%–14% obtained with NS. Giuliano Sisto, Odysseas Zografos, Bilal Chehab, Naveen Kakarla, Dragomir Milojevic, Pieter Weckx, Geert Hellings, Julien Ryckaert |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2020 | Layout Considerations of Logic Designs Using an N-layer 3D Nanofabric Process FlowabstractIn the past few years, novel fabrication schemes such as parallel and monolithic 3D integration have been proposed to keep sustaining the need for more powerful integrated circuits. By stacking several devices, wafers, or dies, the footprint, delay, and power can be decreased when compared to traditional 2D implementations. While parallel 3D does not enable very fine-grained vertical connections, monolithic 3D currently only offers a limited number of transistor tiers due to the high cost of the additional masks and processing steps, limiting the benefits of using the third dimension. In this paper, we introduce an innovative planar circuit netlist and layout approach, which enables a new 3D integration flow called 3D Nanofabric. The flow, consisting of$N$identical vertical tiers, is aimed at single instruction multiple data processor Arithmetic Logic Units (ALUs). By using a single metal routing layer for each vertical tier, the process flow is significantly simplified since multiple vertical layers can potentially be patterned at once, similar to the 3D NAND flash process. In our study, we thoroughly investigate the layout constraints arising from the Nanofabric flow and the unique metal layer rule and propose several ways to overcome them. We then show that by stacking 32 layers to build a 32-bit ALU, the footprint is reduced by$8.7\times$when compared to a conventional 7nm FinFET implementation. Edouard Giacomin, Jürgen Bömmels, Julien Ryckaert, Francky Catthoor, Pierre-Emmanuel Gaillardon |
VLSI-SOC | 3 |
| 2019 | Process, Circuit and System Co-optimization of Wafer Level Co-Integrated FinFET with Vertical Nanosheet Selector for STT-MRAM ApplicationsabstractWe present for the first time a co-integrated FinFET with vertical nanosheet transistor (VFET) process on a 300 mm silicon wafer for STT-MRAM applications and its related avenues with a holistic design-technology-co-optimization (DTCO) and power-performance-area-cost (PPAC) approach. The STT-MRAM bitcell and a 2 Mbit macro have been optimized and designed to address the viability of the co-integration process and advantages of vertical channel transistors for STT-MRAM selectors. An architectural system simulator GEM5 has been also employed with Polybench workloads to assess energy saving at system-level. In order to enable this co-integration, four extra masks are required, which costs below 10% in embedded chips. A 36% area reduction can be achieved for the STT-MRAM bitcell implemented with VFET selectors. With a UVLT flavor, the STT-MRAM bitcell comprising of 3-nanosheet could deliver the same performance of the 4-fin LVT FinFET selector. A 2 Mbit STT-MRAM macro designed with VFET selector can offer a 17% and a 21% reduction for read access latency and energy per operation respectively, and a 10% for write energy per operation. A 7% energy saving for the STT-MRAM L2 cache using VFET selector has been observed at the system level with Polybench workloads. Trong Huynh Bao, Anabela Veloso, Sushil Sakhare, Philippe Matagne, Julien Ryckaert, Manu Perumkunnil Komalan, Davide Crotti, Farrukh Yasin, Alessio Spessot, Arnaud Furnémont, Gouri Sankar Kar, Anda Mocuta |
DAC | 5 |
| 2017 | Statistical Timing Analysis Considering Device and Interconnect Variability for BEOL Requirements in the 5-nm Node and BeyondabstractIn an increasing interconnect resistance era and aggressive metal pitch scaling, the elevating RC delay could significantly shadow the improvements from advanced device architectures and become a severe design issue. This paper will holistically analyze the interplay between transistors and interconnect delay and the variability induced by back-end-ofline (BEOL) process for the 5-nm node. A global sensitivity analysis using Monte Carlo simulation is employed as a powerful tool for understanding the significance of different variation sources and propagating these process uncertainties to circuit performance and parametric yield. For the BEOL integration process, our results show that dielectric κ-value is the most sensitive parameter. Regarding the patterning options, the BEOL process using self-aligned quadruple pattering with positive tone process requires more than a 4× process margin and suffers from 50% parametric yield loss. The required guardband for lithoetch litho-etch becomes as critical as for the self-aligned double patterning process when the overlay control is 6× higher than the critical dimension control. For trench patterning using spacerdefined techniques, a negative tone process is required to achieve a large process window. From a design perspective, the wire length in SoC can be optimized using a disruptive architecture as a vertical FET, which could potentially reduce the average wire length by 11%. Trong Huynh Bao, Julien Ryckaert, Zsolt Tokei, Abdelkarim Mercha, Diederik Verkest, Aaron Thean, Piet Wambacq |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Scaling Beyond 7nm: Design-Technology Co-optimization at the RescueabstractAt 7nm and beyond, designers need to support scaling by identifying the most optimal patterning schemes for their designs. Moreover, designers can actively help by exploring scaling options that do not necessarily require aggressive pitch scaling. In this talk we will illustrate how design technology co-optimization can help achieving the expected Moore's law scaling; how optimizing device performance can lead to smaller standard cells; how the metal interconnect stack needs to be adjusted for unidirectional metals and how a vertical transistor can shift design paradigms. This paper demonstrates that scaling has become a joint design-technology co-optimization effort between process technology and design specialists, that expands way beyond just patterning enabled dimensional scaling. Julien Ryckaert |
ISPD | 1 |
| 2015 | Impact of interconnect multiple-patterning variability on SRAMs
Ioannis Karageorgos, Michele Stucchi, Praveen Raghavan, Julien Ryckaert, Zsolt Tokei, Diederik Verkest, Rogier Baert, Sushil Sakhare, Wim Dehaene |
DATE | 4 |
| 2014 | ParaDIME: Parallel Distributed Infrastructure for Minimization of EnergyabstractDramatic environmental and economic impact of the ever increasing power and energy consumption of modern computing devices in data centers is now a critical challenge. On one hand, designers use technology scaling as one of the methods to face the phenomenon called dark silicon (only segments of a chip function concurrently due to power restrictions). On the other hand, designers use extreme-scale systems such as teradevices to meet the performance needs of their applications which in turn increases the power consumption of the platform. In order to overcome these challenges, we need novel computing paradigms that address energy efficiency. One of the promising solutions is to incorporate parallel distributed methodologies at different abstraction levels. The FP7 project ParaDIME focuses on this objective to provide different distributed methodologies (software-hardware techniques) at different abstraction levels to attack the power-wall problem. In particular, the ParaDIME framework will utilize: circuit and architecture operation below safe voltage limits for drastic energy savings, specialized energy-aware computing accelerators, heterogeneous computing, energy-aware runtime, approximate computing and power-aware message passing. The major outcome of the project will be a processor architecture for a heterogeneous distributed system that utilizes future device characteristics for drastic energy savings. Wherever possible, ParaDIME will adopt multidisciplinary techniques, such as hardware support for message passing, runtime energy optimization utilizing new hardware energy performance counters, use of accelerators for error recovery from sub-safe voltage operation, and approximate computing through annotated code. Furthermore, we will establish and investigate the theoretical limits of energy savings at the device, circuit, architecture, runtime and programming model levels of the computing stack, as well as quantify the actual energy savings achieved by the ParaDIME approach for the complete computing stack with the real environment. Santhosh Kumar Rethinagiri, Oscar Palomar, Anita Sobe, Thomas Knauth, Wojciech M. Barczynski, Gulay Yalcin, Yaroslav Hayduk, Adrián Cristal, Osman S. Unsal, Pascal Felber, Christof Fetzer, Julien Ryckaert, Gina Alioto |
DSD | 12 |
| 2013 | TEASE: a systematic analysis framework for early evaluation of FinFET-based advanced technology nodesabstractThis paper proposes TEASE (Technology Exploration and Analysis for SoC-level Evaluation), a framework to systematically analyze and evaluate system design in finFET-based technology node. The proposed framework combines both lithography and electrical constraints of a particular technology node to optimize the standard cell library performance. Growing complexity of logic design at nodes below 20nm causes to adopt a design style that can embrace the simplicity required to enable manufacturing, along with a process technology that can be finely tuned to the desired performance constraints. Additionally, the introduction of finFET based devices poses a new challenge for the designers to come up with an efficient standard cell template. The proposed framework can be used to detect the technology constraints that act as the bottleneck for the enablement of design at these advanced nodes. Results presented in this paper show by optimizing these bottlenecks we can improve the performance of a standard cell library significantly. Furthermore, adapting to such an analysis framework at an early stage of technology development helps to take the design constraints into the decision loop for realization of technology research into real products. Arindam Mallik, Paul Zuber, Tsung-Te Liu, Bharani Chava, Bhavana Ballal, Pablo Royer Del Bario, Rogier Baert, Kris Croes, Julien Ryckaert, Mustafa Badaroglu, Abdelkarim Mercha, Diederik Verkest |
DAC | 9 |
| 2010 | An 11.6-19.3mW 0.375-13.6GHz CMOS frequency synthesizer with rail-to-rail operationabstractA wide tuning range LO generation architecture for software defined radio is presented. A dual VCO approach followed by a programmable divider chain based on high-speed dynamic CMOS latches provides full rail-to-rail operation with low power consumption. The 1.2V 90nm CMOS implementation achieves a VCO tuning range between 6 to 13.6GHz for a power consumption between 3.5 to 13.4mW and phase noise figure of merit of 182dBc/Hz measured at 3MHz offset from a 12GHz carrier. The VCO-multiplexer and divider chain consumes between 5.9 to 8.1mW for this frequency range. Arnd Geis, Julien Ryckaert, Yves Rolain, Gerd Vandersteen, Jan Craninckx |
DATE | 3 |
| 2008 | A Low Power, Reconfigurable IR-UWB SystemabstractImpulse-radio-UWB is the ideal air interface for low power wireless applications especially when they require ranging capabilities. This paper presents a complete UWB transceiver system, including acquisition and ranging protocols. The system is fully reconfigurable in terms of bandwidth, data rate, processing gain, acquisition protocol and ranging accuracy, in order to fulfill the needs of the application with minimal energy consumption. The complete system is demonstrated by measurements on an IR-UWB transceiver platform built around 3 fully integrated CMOS chips. The transceiver system achieves a data rate up to 50 Mbps and a ranging error with a root mean squared error of less than 10 cm while consuming 31.7 mW. The IC implementation allows to fully validate the power/flexibility trade-off that can be achieved with integrated solutions. Marian Verhelst, Julien Ryckaert, Yves Vanderperren, Wim Dehaene |
ICC | 2 |
| 2007 | Optimized Signal Acquisition for Low-Complexity and Low-Power IR-UWB TransceiversabstractImpulse-based ultra-wideband systems are appealing for low-power short-range communications as they can benefit from duty-cycled low-complexity analog architectures. Low-power receivers correlate in the analog domain and sample only at the pulse repetition frequency. However, achieving acquisition with such a receiver requires an efficient search strategy, as digital post-processing (equalization) of the signal is not possible. Moreover, as we target a network of ultra-low-power sensors which can only send a few pulses and do not implement receiver functionality in order to sustain ARQ protocols, the receiving base station has extra constraints on synchronization performance and maximal preamble length. This paper describes and optimizes a strategy compliant with that scenario, including pulse position and spreading code phase acquisition, end-of-preamble detection, and derivation of the corresponding detection thresholds. Based on a 300-bit preamble, it works within 1.5 dB of ideal synchronization on AWGN. EOP detection is analyzed, showing that a PN sequence of length 7 is sufficient, and also proposing optimal receive filters for EOP sequences defined in IEEE 802.15.4a. Finally, we show that on fading channels, the error floor coming from missed acquisitions can be removed by simply changing the detection threshold. Claude Desset, Mustafa Badaroglu, Julien Ryckaert, Bart van Poucke |
VTC Spring | 3 |
| 2006 | Human++: Emerging Technology for Body Area NetworksabstractThis paper gives an overview of results of the Human++ research program. This research aims to achieve highly miniaturized and autonomous transducer systems that assist our health and comfort. It combines expertise in wireless ultra-low power communications, 3D integration technologies, MEMS energy scavenging techniques and low-power design techniques Bert Gyselinckx, Ruud J. M. Vullers, Chris Van Hoof, Julien Ryckaert, Refet Firat Yazicioglu, Paolo Fiorini, Vladimir Leonov |
VLSI-SoC | 4 |
| 2006 | Ultra-wideband channel model for communication around the human bodyabstractUsing ultra-wideband (UWB) wireless sensors placed on a person to continuously monitor health information is a promising new application. However, there are currently no detailed models describing the UWB radio channel around the human body making it difficult to design a suitable communication system. To address this problem, we have measured radio propagation around the body in a typical indoor environment and incorporated these results into a simple model. We then implemented this model on a computer and compared experimental data with the simulation results. This paper proposes a simple statistical channel model and a practical implementation useful for evaluating UWB body area communication systems. Andrew Fort, Julien Ryckaert, Claude Desset, Philippe De Doncker, Piet Wambacq, Leo Van Biesen |
IEEE J. Sel. Areas Commun. | 2 |
| 2005 | Ultra wide-band body area channel modelabstractUsing wireless sensors placed on a person to continuously monitor health information is a promising new application. However, there are currently no models describing the radio channel around the human body making it difficult to design a suitable communication system. To address this problem, we have simulated electromagnetic wave propagation around the body and incorporated these results into a simple model. We then compared this model with measurements taken around the human torso and with previous studies in the literature. This paper proposes a simple statistical channel model useful for evaluating both UWB and (after resampling) narrow-band body area communication systems. Andrew Fort, Claude Desset, Julien Ryckaert, Philippe De Doncker, Leo Van Biesen, Stéphane Donnay |
ICC | 3 |