Fernando García-Redondo

dblp:166/7014 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-7090-8821ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 3D SRAM Disaggregation in Advanced CMOS Nodes using Hybrid Bonding Technology
abstract
This paper studies the potential of hybrid bonded Array-under-CMOS (AuC) technology to partition logic and high-performance L1 cache in advanced technology nodes. By decoupling the SRAM bitcells from the logic tier, we achieve independent optimization of both SRAM and logic devices, as well as back-end-of-line (BEOL) interconnects. Heterogeneous integration and BEOL aspect ratio optimization is implemented with different technology nodes to mitigate the delay penalty due to hybrid bond pad staggering. Addressing the performance degradation associated with scaled technology nodes, we investigate the impact of word-line (WL) and bit-line (BL) resistance on SRAM performance. Leveraging the flexibility of decoupled SRAM BEOL and within the AuC technology framework, we explore the sensitivity to WL and BL metal aspect ratios, comparing their performance against a 2D baseline. Our results demonstrate substantial performance improvements in AuC integration through two key approaches: (1) 5% enhancement via Back-End-of-Line (BEOL) optimization, and (2) 25% improvement enabled by heterogeneous integration, achieved by decoupling memory and logic tiers.
Bhawana Kumari, Anurag Swarnkar, Dawit Burusie Abdi, Fernando García-Redondo, James Myers, Julien Ryckaert, Jaydeep P. Kulkarni, Dwaipayan Biswas
ISCAS4
2025 3D IGZO Charge-Coupled Memory DTCO & STCO Analysis for Compute-near-Memory Applications
abstract
The demand for high-capacity and energy-efficient memory solutions has surged in the era of data-centric computing, particularly for Artificial Intelligence (AI) and Machine Learning (ML) workloads. This paper introduces a novel memory architecture leveraging Charge-Coupled Device (CCD) technology, engineered in a sequential-access block memory configuration, to enhance Compute-near-Memory (CnM) systems. We propose an optimized 3D IGZO CCD block memory as an on-chip weight buffer for high-capacity CnM systems. Our approach achieves 2.95−131.26× improvement in area efficiency and 1.32−4.33× improvement in energy efficiency compared to SRAM solutions.
Khakim Akhunov, Hyungrock Oh, Fernando García-Redondo, Yukai Chen, Arvind Sharma, Jiacong Sun, Sahan Gamage, Maarten Rosmeulen, Swaraj Bandhu Mahato, Rishabh Kishore, Subhali Subhechha, Jaydeep P. Kulkarni, Marian Verhelst, Dwaipayan Biswas, Marie Garcia Bardon, Wim Dehaene, Julien Ryckaert
ISCAS4
2024 A DTCO Framework for 3D NAND Flash Readout
abstract
To continue increasing the storage density of 3D NAND flash memories, new technology options need to be evaluated early on. This work presents a unique predictive parametric framework for Multi-Level Cell 3D NAND Flash read operation at the array level. This framework is used to explore the read sensitivity to multiple parameters and technology options. We identify the trade-offs between number of layers, read-current and read time to be the most determinant factors to ensure the array readability while enabling stacks of more than 300 layers and maximizing the memory density.
Mattia Gerardi, Arvind Sharma, Jakub Kaczmarek, Fernando García-Redondo, Maarten Rosmeulen, Marie Garcia Bardon
DATE5
2023 AR-PIM: An Adaptive-Range Processing-in-Memory Architecture
abstract
The crossbar-based processing-in-memory (PIM) architecture has garnered considerable attention for its potential in achieving high energy efficiency for deep neural networks (DNNs). The PIM hardware's accuracy depends heavily on the design and resolution of the analog-to-digital converters (ADCs). Regrettably, high-resolution ADCs tend to be costly and often dominate the overall energy and area of the PIM designs. We propose adaptive-range PIM (AR-PIM) architecture that enables the use of lower-resolution ADCs without sacrificing accuracy. This is achieved by leveraging sparsity in the weights and input activations and dynamically adjusting the number of input activations and distributing MAC operations across multiple cycles during runtime. We perform our evaluations using a commercial 7nm FinFET PDK and show that AR-PIM offers an appealing trade-off, delivering 1.7 × higher energy efficiency and 4.3 × better area benefits without losing accuracy. The latency overhead is modest, only 10% over a baseline PIM architecture.
Teyuh Chou, Fernando García-Redondo, Paul N. Whatmough, Zhengya Zhang
ISLPED2
2020 Training DNN IoT Applications for Deployment On Analog NVM Crossbars
abstract
A trend towards energy-efficiency, security and privacy has led to a recent focus on deploying deep-neural networks (DNN) on microcontrollers. However, limits on compute and memory resources restrict the size and the complexity of the machine-learning (ML) models deployable in these systems. Computation-In-Memory architectures based on resistive non-volatile memory (NVM) technologies hold great promise of satisfying the compute and memory demands of high-performance and low-power, inherent in modern DNNs. Nevertheless, these technologies are still immature and suffer from both the intrinsic analog-domain noise problems and the inability of representing negative weights in the NVM structures, incurring in larger crossbar sizes with concomitant impact on Analog-to-Digital Converters (ADCs) and Digital-to-Analog Converters (DACs). In this paper, we provide a training framework for addressing these challenges and quantitatively evaluate the circuit-level efficiency gains thus accrued. We make two contributions: Firstly, we propose a training algorithm that eliminates the need for tuning individual layers of a DNN ensuring uniformity across layer-weights and activations. This ensures analog-blocks that can be reused and peripheral hardware substantially reduced. Secondly, using Network Architecture Search (NAS) methods, we propose the use of unipolar-weighted (either all-positive or all-negative weights) matrices/sub-matrices. Weight unipolarity obviates the need for doubling crossbar area leading to simplified analog periphery. We validate our methodology with CIFAR10 and HAR applications by mapping to crossbars using 4-bit and 2-bit devices. We achieve up to 92.91% accuracy (95% floating-point) using 2-bit only-positive weights for HAR. A combination of the proposed techniques leads to 80% area improvement and up to 45% energy reduction.
Fernando García-Redondo, Shidhartha Das, Glen Rosendale
IJCNN1
2019 Applications of Computation-In-Memory Architectures based on Memristive Devices
abstract
Today's computing architectures and device technologies are unable to meet the increasingly stringent demands on energy and performance posed by emerging applications. Therefore, alternative computing architectures are being explored that leverage novel post-CMOS device technologies. One of these is a Computation-in-Memory architecture based on memristive devices. This paper describes the concept of such an architecture and shows different applications that could significantly benefit from it. For each application, the algorithm, the architecture, the primitive operations, and the potential benefits are presented. The applications cover the domains of data analytics, signal processing, and machine learning.
Said Hamdioui, Hoang Anh Du Nguyen, Mottaqiallah Taouil, Abu Sebastian, Manuel Le Gallo, Sandeep Pande, Siebren Schaafsma, Francky Catthoor, Shidhartha Das, Fernando García-Redondo, Geethan Karunaratne, Abbas Rahimi, Luca Benini
DATE10
2017 Reconfigurable Writing Architecture for Reliable RRAM Operation in Wide Temperature Ranges
abstract
Resistive switching memories [resistive RAM (RRAM)] are an attractive alternative to nonvolatile storage and nonconventional computing systems, but their behavior strongly depends on the cell features, driver circuit, and working conditions. In particular, the circuit temperature and writing voltage schemes become critical issues, determining resistive switching memories performance. These dependencies usually force a design time tradeoff among reliability, device endurance, and power consumption, thereby imposing nonflexible functioning schemes and limiting the system performance. In this paper, we present a writing architecture that ensures the correct operation no matter the working temperature and allows the dynamic load of application-oriented writing profiles. Thus, taking advantage of more efficient configurations, the system can be dynamically adapted to overcome RRAM intrinsic challenges. Several profiles are analyzed regarding power consumption, temperature-variations protection, and operation speed, showing speedups near 700× compared with other published drivers.
Fernando García-Redondo, Pablo Royer, Marisa López-Vallejo, Hernan Aparicio, Pablo Ituero, Carlos A. López-Barrio
IEEE Trans. Very Large Scale Integr. Syst.1
2015 A thermal adaptive scheme for reliable write operation on RRAM based architectures
abstract
Resistive RAMs (RRAMs) are one of the most promising alternatives to future storage and neuromorphic computing systems. However, the behavior of RRAM highly depends on voltage, crossbar design and operation temperature. Actually, the circuit temperature becomes one of the most critical issues in fast memories during writing operations. In this paper we propose a novel thermal-adaptive RRAM writing scheme, applicable to crossbar memories, whose smart operation is able to mitigate the writing errors induced by temperature variations. Using a sensing-acting scheme our system is able to improve the memory reliability without affecting the writing/reading performance. Moreover, the proposed architecture is compatible with most proposed write/read designs making achievable multibit storage, which requires extremely accurate operations.
Fernando García-Redondo, Marisa López-Vallejo, Pablo Ituero
ICCD1