Wilfred Gomes

dblp:251/2709 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 ADC/DAC-Free Analog Acceleration of Deep Neural Networks With Frequency Transformation
abstract
The edge processing of deep neural networks (DNNs) is becoming increasingly important due to its ability to extract valuable information directly at the data source to minimize latency and energy consumption. Although pruning techniques are commonly used to reduce model size for edge computing, they have certain limitations. Frequency-domain model compression, such as with the Walsh–Hadamard transform (WHT), has been identified as an efficient alternative. However, the benefits of frequency-domain processing are often offset by the increased multiply-accumulate (MAC) operations required. This article proposes a novel approach to an energy-efficient acceleration of frequency-domain neural networks by utilizing analog-domain frequency-based tensor transformations. Our approach offers unique opportunities to enhance computational efficiency, resulting in several high-level advantages, including array microarchitecture with parallelism, analog-to-digital converter (ADC)/digital-to-analog converter (DAC)-free analog computations, and increased output sparsity. Our approach achieves more compact cells by eliminating the need for trainable parameters in the transformation matrix. Moreover, our novel array microarchitecture enablesadaptive stitchingof cells column-wise and row-wise, thereby facilitating perfect parallelism in computations. Additionally, our scheme enables ADC/DAC-free computations by training against highly quantized matrix-vector products, leveraging the parameter-free nature of matrix multiplications. Another crucial aspect of our design is its ability to handle signed-bit processing for frequency-based transformations. This leads to increased output sparsity and reduced digitization workload. On a$16 \ttimes 16$crossbars, for 8-bit input processing, the proposed approach achieves the energy efficiency of 801 tera operations per second per Watt (TOPS/W) without early termination strategy and 2655 TOPS/W with early termination strategy at VDD$=$0.85 V for 16-nm predictive technology models (PTM).
Nastaran Darabi, Maeesha Binte Hashem, Hongyi Pan, A. Enis Çetin, Wilfred Gomes, Amit Ranjan Trivedi
IEEE Trans. Very Large Scale Integr. Syst.5
2023 MC-CIM: Compute-in-Memory With Monte-Carlo Dropouts for Bayesian Edge Intelligence
abstract
We propose MC-CIM, a compute-in-memory (CIM) framework for robust, yet low power, Bayesian edge intelligence. Deep neural networks (DNN) with deterministic weights cannot express their prediction uncertainties, thereby pose critical risks for applications where the consequences of mispredictions are fatal such as surgical robotics. To address this limitation, Bayesian inference of a DNN has gained attention. Using Bayesian inference, not only the prediction itself, but the prediction confidence can also be extracted for planning risk-aware actions. However, Bayesian inference of a DNN is computationally expensive, ill-suited for real-time and/or edge deployment. An approximation to Bayesian DNN using Monte Carlo Dropout (MC-Dropout) has shown high robustness along with low computational complexity. Enhancing the computational efficiency of the method, we discuss a novel CIM module that can perform in-memory probabilistic dropout in addition to in-memory weight-input scalar product to support the method. We also propose a compute-reuse reformulation of MC-Dropout where each successive instance can utilize the product-sum computations from the previous iteration. Even more, we discuss how the random instances can be optimally ordered to minimize the overall MC-Dropout workload by exploiting combinatorial optimization methods. Application of the proposed CIM-based MC-Dropout execution is discussed for MNIST character recognition and visual odometry (VO) of autonomous drones. The framework reliably gives prediction confidence amidst non-idealities imposed by MC-CIM to a good extent. Proposed MC-CIM with$16\times 31$SRAM array, 0.85 V supply, 16nm low-standby power (LSTP) technology consumes 32 pJ for 30 MC-Dropout instances of probabilistic inference in its most optimal computing and peripheral configuration, saving$\sim 34$% energy compared to typical execution.
Priyesh Shukla, Shamma Nasrin, Nastaran Darabi, Wilfred Gomes, Amit Ranjan Trivedi
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Meteor Lake and Arrow Lake Intel Next-Gen 3D Client Architecture Platform with Foveros
abstract
We create world-changing technology that improves the life of every person on the planet
Wilfred Gomes, Slade Morgan, Boyd Phelps, Tim Wilson, Erik Hallnor
HCS1
2021 MF-Net: Compute-In-Memory SRAM for Multibit Precision Inference Using Memory-Immersed Data Conversion and Multiplication-Free Operators
abstract
We propose a co-design approach for compute-in-memory inference for deep neural networks (DNN). We use multiplication-free function approximators based on l1norm along with a co-adapted processing array and compute flow. Using the approach, we overcame many deficiencies in the current art of in-SRAM DNN processing such as the need for digital-to-analog converters (DACs) at each operating SRAM row/column, the need for high precision analog-to-digital converters (ADCs), limited support for multi-bit precision weights, and limited vector-scale parallelism. Our co-adapted implementation seamlessly extends to multi-bit precision weights, it doesn't require DACs, and it easily extends to higher vector-scale parallelism. We also propose an SRAM-immersed successive approximation ADC (SA-ADC), where we exploit the parasitic capacitance of bit lines of SRAM array as a capacitive DAC. Since the dominant area overhead in SA-ADC comes due to its capacitive DAC, by exploiting the intrinsic parasitic of SRAM array, our approach allows low area implementation of within-SRAM SA-ADC. Our 8×62 SRAM macro, which requires a 5-bit ADC, achieves ~105 tera operations per second per Watt (TOPS/W) with 8-bit input/weight processing at 45 nm CMOS. Our 8×30 SRAM macro, which requires a 4-bit ADC, achieves ~84 TOPS/W. SRAM macros that require lower ADC precision are more tolerant of process variability, however, have lower TOPS/W as well. We evaluated the accuracy and performance of our proposed network for MNIST, CIFAR10, and CIFAR100 datasets. We chose a network configuration which adaptively mixes multiplication-free and regular operators. The network configurations utilize the multiplication-free operator for more than 85% operations from the total. The selected configurations are 98.6% accurate for MNIST, 90.2% for CIFAR10, and 66.9% for CIFAR100. Since most of the operations in the considered configurations are based on proposed SRAM macros, our compute-in-memory's efficiency benefits broadly translate to the system-level.
Shamma Nasrin, Diaa Badawi, A. Enis Çetin, Wilfred Gomes, Amit Ranjan Trivedi
IEEE Trans. Circuits Syst. I Regul. Pap.4
2019 Lakefield: Hybrid cores in 3D Package
abstract
This article consists of a collection of slides from the author's conference presentation.
Sanjeev Khushu, Wilfred Gomes
Hot Chips Symposium2