Flavio Ponzina

dblp:264/9298 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-9662-498XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 4 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
abstract
Approximate Nearest Neighbor Search (ANNS) is a fundamental operation in vector databases, enabling efficient similarity search in high-dimensional spaces. While dense ANNS has been optimized using specialized hardware accelerators, sparse ANNS remains limited by CPU-based implementations, hindering scalability. This limitation is increasingly critical as hybrid retrieval systems—combining sparse and dense embeddings—become standard in Information Retrieval (IR) pipelines. We propose SpANNS, a near-memory processing architecture for sparse ANNS. SpANNS combines a hybrid inverted index with efficient query management and runtime optimizations. The architecture is built on a CXL Type-2 near-memory platform, where a specialized controller manages query parsing and cluster filtering, while compute-enabled DIMMs perform index traversal and distance computations close to the data. It achieves $15.2 \times$ to $21.6 \times$ faster execution over the state-of-the-art CPU baselines, offering scalable and efficient solutions for sparse vector search.
Flavio Ponzina, Tajana Rosing
ASP-DAC2
2026 FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
abstract
Approximate Nearest-Neighbor Search (ANNS) is a key technique in retrieval-augmented generation (RAG), enabling rapid identification of the most relevant high-dimensional embeddings from massive vector databases. Modern ANNS engines accelerate this process using prebuilt indexes and store compressed vector-quantized representations in fast memory. However, they still rely on a costly second-pass refinement stage that reads full-precision vectors from slower storage like SSDs. For modern text and multimodal embeddings, these reads now dominate the latency of the entire query. We propose FaTRQ, a far-memory-aware refinement system using tiered memory that eliminates the need to fetch full vectors from storage. It introduces a progressive distance estimator that refines coarse scores using compact residuals streamed from far memory. Refinement stops early once a candidate is provably outside the top-k. To support this, we propose tiered residual quantization, which encodes residuals as ternary values stored efficiently in far memory. A custom accelerator is deployed in a CXL Type-2 device to perform low-latency refinement locally. Together, FaTRQ improves the storage efficiency by 2.4× and improves the throughput by up to 9× than SOTA GPU ANNS system.
Flavio Ponzina, Tajana Rosing
DATE2
2026 HeteroRAGCache: Software-Hardware Co-Design for Efficient RAG Caching using Emerging Memories
abstract
Retrieval-augmented generation (RAG) systems mitigate hallucinations in large language models (LLMs) and improve response accuracy by leveraging external knowledge. They introduce additional latency overhead due to document retrieval and long-context processing, resulting in increased compute and memory costs. Recent approaches, such as RAGCache, utilize external DRAM to cache key-value (KV) pairs generated during the prefill phase, enabling the reuse of frequently accessed documents. However, DRAM-based caches incur high access latency and energy overhead to access the off-package DRAM frequently.
Jangseon Park, Kiseok Suh, Flavio Ponzina, Tajana Rosing
ACM Great Lakes Symposium on VLSI3
2026 SIMCH: Stochastic In-Memory Computing using High-Density MTJ
Qiuyuan Wang, Flavio Ponzina, Luqiao Liu, Tajana Rosing
ISCAS3
2026 QMC: Efficient SLM Edge Inference via Outlier-Aware Quantization and Emergent Memories Co-Design
abstract
Deploying Small Language Models (SLMs) on edge platforms is critical for real-time, privacy-sensitive generative AI, yet constrained by memory, latency, and energy budgets. Quantization reduces model size and cost but suffers from device noise in emerging nonvolatile memories, while conventional memory hierarchies further limit efficiency. SRAM provides fast access but has low density, DRAM must simultaneously accommodate static weights and dynamic KV caches, which creates bandwidth contention, and Flash, although dense, is primarily used for initialization and remains inactive during inference. These limitations highlight the need for hybrid memory organizations tailored to LLM inference. We propose Outlier-aware Quantization with Memory Co-design (QMC), a retraining-free quantization with a novel heterogeneous memory architecture. QMC identifies inlier and outlier weights in SLMs, storing inlier weights in compact multi-level Resistive-RAM (ReRAM) while preserving critical outliers in high-precision on-chip Magnetoresistive-RAM (MRAM), mitigating noise-induced degradation. On language modeling and reasoning benchmarks, QMC outperforms and matches state-of-the-art quantization methods using advanced algorithms and hybrid data formats, while achieving greater compression under both algorithm-only evaluation and realistic deployment settings. Specifically, compared against SoTA quantization methods on the latest edge AI platform, QMC reduces memory usage by 5.5×-7.3×, external data transfers by 7.6×, energy by 11× - 11.7×, and latency by 9.4× - 12.5× when compared to FP16, establishing QMC as a scalable, deployment-ready co-design for efficient on-device inference.
Nilesh Prasad Pandey, Jangseon Park, Onat Güngör, Flavio Ponzina, Tajana Rosing
ISLPED4
2025 E-QUARTIC: Energy Efficient Edge Ensemble of Convolutional Neural Networks for Resource-Optimized Learning
abstract
Ensemble learning is a meta-learning approach that combines the predictions of multiple learners, demonstrating improved accuracy and robustness. Nevertheless, ensembling models like Convolutional Neural Networks (CNNs) result in high memory and computing overhead, preventing their deployment in embedded systems. These devices are usually equipped with small batteries that provide power supply and might include energy-harvesting modules that extract energy from the environment. In this work, we propose E-QUARTIC, a novel Energy Efficient Edge Ensembling framework to build ensembles of CNNs targeting Artificial Intelligence (AI)-based embedded systems. Our design outperforms single-instance CNN baselines and state-of-the-art edge AI solutions, improving accuracy and adapting to varying energy conditions while maintaining similar memory requirements. Then, we leverage the multi-CNN structure of the designed ensemble to implement an energy-aware model selection policy in energy-harvesting AI systems. We show that our solution outperforms the state-of-the-art by reducing system failure rate by up to 40% while ensuring higher average output qualities. Ultimately, we show that the proposed design enables concurrent on-device training and high-quality inference execution at the edge, limiting the performance and energy overheads to less than 0.04%.
Le Zhang 0021, Onat Güngör, Flavio Ponzina, Tajana Rosing
ASP-DAC3
2025 DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing
Nilesh Prasad Pandey, Shriniwas Kulkarni, Onat Güngör, Flavio Ponzina, Tajana Rosing
ACM Great Lakes Symposium on VLSI5
2025 FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash
abstract
The rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy.
Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang
ICCAD10
2025 HyperDrone: An Accurate, Robust, Fast, and Energy-Efficient Approach for Drone Classification
abstract
Deep learning (DL) for Radio Frequency (RF) signal processing has gained significant traction, with drone classification emerging as one of the key applications in security-sensitive contexts. However, existing DL approaches often require considerable computational resources, lack robustness to adversarial attacks, and operate in static, ground-based settings. The integration of Universal Software Radio Peripheral (USRP) modules into drones now enables real-time, on-board RF signal processing, opening new avenues for scalable and responsive security systems. In this work, we present HyperDrone, the first RF signal processing framework based on Hyperdimensional Computing (HDC). HyperDrone employs a resource-efficient ensemble of HDC classifiers, combining diverse encoding strategies to achieve state-of-the-art accuracy in drone detection. This design not only enhances performance over standard HDC models but also significantly improves robustness to adversarial perturbations without the need for adversarial training. Compared to prior works, HyperDrone has up to$\mathbf{1 0} \times$faster inference and$\mathbf{2 7} \times$faster training. It improves few-shot learning accuracy by over 30 % and adversarial robustness by 18 %, with less than 1 % accuracy drop.
Shriniwas Kulkarni, Flavio Ponzina, Tajana Rosing
ICCD2
2025 SmartMS: Efficient Hierarchical Database Search for Mass Spectrometry via Processing-in-Memory
abstract
The acceleration of Mass Spectrometry (MS) library search is crucial for advancing scientific and pharmaceutical research. Recent methodologies leverage Hyperdimensional Computing (HDC) to encode reference and query spectra as high-dimensional vectors, enabling highly parallel similarity computations. In this context, Processing-In-Memory (PIM) has emerged as a promising solution, offering orders of magnitude improvements in computational speed compared to GPU-based approaches when handling large-scale libraries. However, bruteforce search methods remain computationally intensive, exacerbating the high energy demands associated with MS library search operations in data centers. In this work, we propose SmartMS, a novel tool that leverages HDC to construct a multi-level database structure, reducing search complexity from linear to logarithmic while maintaining compatibility with PIM-based accelerators. SmartMS improves identification accuracy by 3% while delivering a 33× improvement in speed and a 58× energy reduction, with a negligible increase in memory requirements of 0.5% when compared to the current state of the art.
Flavio Ponzina, Sumukh Pinge, Abhijay Deevi, Yilin Ge, Mingu Kang, Tajana Rosing
ISLPED1
2025 Offload Rethinking by Cloud Assistance for Efficient Environmental Sound Recognition on LPWANs
abstract
Learning-based environmental sound recognition has emerged as a crucial method for ultra-low-power environmental monitoring in biological research and city-scale sensing systems. These systems usually operate under limited resources and are often powered by harvested energy in remote areas. Recent efforts in on-device sound recognition suffer from low accuracy due to resource constraints, whereas cloud offloading strategies are hindered by high communication costs. In this work, we introduce ORCA, a novel resource-efficient cloud-assisted environmental sound recognition system on batteryless devices operating over the Low-Power Wide-Area Networks (LPWANs), targeting wide-area audio sensing applications. We propose a cloud assistance strategy that remedies the low accuracy of on-device inference while minimizing the communication costs for cloud offloading. By leveraging a self-attention-based cloud sub-spectral feature selection method to facilitate efficient on-device inference, ORCA resolves three key challenges for resource-constrained cloud offloading over LPWANs: 1) high communication costs and low data rates, 2) dynamic wireless channel conditions, and 3) unreliable offloading. We implement ORCA on an energy-harvesting batteryless microcontroller and evaluate it in a real world urban sound testbed. Our results show that ORCA outperforms state-of-the-art methods by up to 80× in energy savings and 220× in latency reduction while maintaining comparable accuracy.
Le Zhang 0021, Quanling Zhao, Run Wang 0003, Shirley Bian, Onat Güngör, Flavio Ponzina, Tajana Rosing
SenSys6
2024 HyperECG: ECG Signal Inference From Radar With Hyperdimensional Computing
abstract
Contactless ECG monitoring with radar technology is used in both long-term and remote healthcare monitoring. While first attempts were able to only estimate heart rate, more recently deep learning (DL) has been used to infer the continuous ECG signal, which is essential for many health monitoring applications. However, the compute-intensive nature of DL models makes it hard to deploy and personalize them in low-power systems that are critical for remote healthcare monitoring. To address this challenge, we introduce HyperECG, a pioneering approach based on Hyperdimensional Computing (HDC), an efficient alternative machine learning method, to infer ECG signals from radar inputs. We combine a novel learnable HDC projection encoding with state-of-the-art HDC regressors to achieve high-quality ECG estimation. Experimental results reveal that HyperECG achieves output quality comparable to the state-of-the-art DL while reducing inference and training runtime up to$23 \times$and$36 \times$, respectively. HyperECG supports on-device model personalization, crucial in medical settings, with accuracy improvements of up to$68 \%$on patient-specific evaluations, compared to before fine-tuning HyperECG.
Matilda Gaddi, Flavio Ponzina, Fatemeh Asgarinejad, Baris Aksanli, Tajana Rosing
BIBE2
2024 HDXpose: Harnessing Hyperdimensional Computing's Explainability for Adversarial Attacks
abstract
Hyperdimensional Computing (HDC), a promising alternative to address the limitations of edge devices, is not exempt from the security challenges confronted by machine learning algorithms, in particular, adversarial attacks. The limited body of research exploring the security implications of HDC overlooks its inherent algorithm. In this paper, we propose a novel and effective adversarial attack technique targeting HDC. Our approach analyzes and prioritizes the impact of input features as well as encoded elements on decision boundaries and perturbs the input towards incorrect decisions in a guided manner. We evaluate our method on different datasets and attack models (i.e., untargeted/targeted, white-box/gray-box). Experimental results indicate that our proposed design, HDXpose, significantly outperforms the state-of-the-art attack techniques by achieving higher success rate with smaller distortion and execution time, rendering its efficacy for real-time attack generation.
Fatemeh Asgarinejad, Flavio Ponzina, Onat Güngör, Tajana Rosing, Baris Aksanli
ICCAD2
2024 Multi-Objective Software-Hardware Co-Optimization for HD-PIM via Noise-Aware Bayesian Optimization
abstract
In hardware accelerator design, software-hardware co-optimization requires intricate trade-offs and tight integration between software algorithms and hardware design to optimize performance, power efficiency, and area (PPA) while ensuring high accuracy. Furthermore, the inherent non-ideality in some emerging hardware technologies poses extra challenges to the co-optimization problem. This paper proposes a novel software-hardware co-optimization framework for hyperdimensional (HD) computing accelerators with emerging ReRAM-based processing in-memory (PIM) technologies, which have shown superior performance and energy efficiency over conventional machine learning accelerators. We first comprehensively characterize the non-trivial trade-offs between design parameters in HD-PIM and PPA and accuracy metrics in HD-PIM. Then, we develop a multi-objective noise-aware Bayesian optimization algorithm to find the Pareto set (optimal trade-offs between metrics) of the HD-PIM design. Our methodology uniquely addresses the stochastic nature of ReRAM by integrating error characteristics into the optimization process, thereby enhancing the quality of the generated designs. Experimental results show that our configurations achieve up to 4.28% accuracy improvement, 35.38% power reduction, 49x timing improvement, and 10% area reduction over a non-optimized design.
Chien-Yi Yang, Minxuan Zhou, Flavio Ponzina, Suraj Sathya Prakash, Raid Ayoub, Pietro Mercati, Mahesh Subedar, Tajana Rosing
ICCAD3
2024 Multi-Model Inference Composition of Hyperdimensional Computing Ensembles
abstract
To answer the ever-increasing demand for high accuracy in artificial intelligence (AI)-based applications, several models have been proposed. Among them, ensemble learning, a technique that trains multiple classifiers and then combines their prediction during the inference stage, emerged as a promising approach. Despite being largely explored in the context of models like random forests or convolutional neural networks, very few research works have focused on ensemble learning targeting hyperdimensional computing (HDC). HDC is a brain-inspired computing paradigm that has gained momentum in the last decade because its lightweight and highly parallel operations make it an excellent alternative to compute-intense deep learning models for edge AI applications. In this work, we propose BagHD and BoostHD, two ensemble-based HDC implementations constructed using bagging and boosting, respectively. Accuracy evaluations indicate that our proposal improves baseline single-instance implementations and state-of-the-art HDC ensembles by up to 14% and 4%, respectively. We then leverage two key characteristics of HDC and ensemble learning to demonstrate how we can transform the proposed ensembles into equivalent single-instance implemen-tations, thus avoiding any memory and computing overhead during inference. In fact, when compared to traditional ordinary ensembles, we reduce memory requirements by up to 40x, improving accuracy at the same time. We also support ensemble learning HDC training in BagHD and BoostHD, showing that with little memory overhead it is possible to retrieve the original weak learners from the generated single-instance design.
Flavio Ponzina, Rishikanth Chandrasekaran, Anya Wang, Seiji Minowada, Tajana Rosing
ICCD1
2024 An Energy Efficient Soft SIMD Microarchitecture and Its Application on Quantized CNNs
abstract
The ever-increasing computational complexity and energy consumption of today’s applications, such as machine learning (ML) algorithms, not only strain the capabilities of the underlying hardware but also significantly restrict their wide deployment at the edge. Addressing these challenges, novel architecture solutions are required by leveraging opportunities exposed by algorithms, e.g., robustness to small-bitwidth operand quantization and high intrinsic data-level parallelism. However, traditional hardware single instruction multiple data (Hard SIMD) architectures only support a small set of operand bitwidths, limiting performance improvement. To fill the gap, this manuscript introduces a novel pipelined processor microarchitecture for arithmetic computing based on the software-defined SIMD (Soft SIMD) paradigm that can define arbitrary SIMD modes through control instructions at run-time. This microarchitecture is optimized for parallel fine-grained fixed-point arithmetic, such as shift/add. It can also efficiently execute sequential shift-add-based multiplication over SIMD subwords, thanks to zero-skipping and canonical signed digit (CSD) coding. A lightweight repacking unit allows changing subword bitwidth dynamically. These features are implemented within a tight energy and area budget. An energy consumption model is established through post-synthesis for performance assessment. We select heterogeneously quantized (HQ) convolutional neural networks (CNNs) from the ML domain as the benchmark and map it onto our microarchitecture. Experimental results showcase that our approach dramatically outperforms traditional Hard SIMD Multiplier-Adder regarding area and energy requirements. In particular, our microarchitecture occupies up to 59.9% less area than a Hard SIMD that supports fewer SIMD bitwidths, while consuming up to 50.1% less energy on average to execute HQ CNNs.
Pengbo Yu, Flavio Ponzina, Alexandre Levisse, Mohit Gupta 0004, Dwaipayan Biswas, Giovanni Ansaloni, David Atienza 0001, Francky Catthoor
IEEE Trans. Very Large Scale Integr. Syst.2
2023 Overflow-free Compute Memories for Edge AI Acceleration
abstract
Compute memories are memory arrays augmented with dedicated logic to support arithmetic. They support the efficient execution of data-centric computing patterns, such as those characterizing Artificial Intelligence (AI) algorithms. These architectures can provide computing capabilities as part of the memory array structures (In-Memory Computing, IMC) or at their immediate periphery (Near-Memory Computing, NMC). By bringing the processing elements inside (or very close to) storage, compute memories minimize the cost of data access. Moreover, highly parallel (and, hence, high-performance) computations are enabled by exploiting the regular structure of memory arrays. However, the regular layout of memory elements also constrains the data range of inputs and outputs, since the bitwidths of operands and results stored at each address cannot be freely varied. Addressing this challenge, we herein propose a HW/SW co-design methodology combining careful per-layer quantization and inter-layer scaling with lightweight hardware support for overflow-free computation of dot-vector operations. We demonstrate their use to implement the convolutional and fully connected layers of AI models. We embody our strategy in two implementations, based on IMC and NMC, respectively. Experimental results highlight that an area overhead of only 10.5% (for IMC) and 12.9% (for NMC) is required when interfacing with a 2KB subarray. Furthermore, inferences on benchmark CNNs show negligible accuracy degradation due to quantization for equivalent floating-point implementations.
Flavio Ponzina, Marco Rios, Alexandre Levisse, Giovanni Ansaloni, David Atienza 0001
ACM Trans. Embed. Comput. Syst.1
2022 Error Resilient In-Memory Computing Architecture for CNN Inference on the Edge
abstract
The growing popularity of edge computing has fostered the development of diverse solutions to support Artificial Intelligence (AI) in energy-constrained devices. Nonetheless, comparatively few efforts have focused on the resiliency exhibited by AI workloads (such as Convolutional Neural Networks, CNNs) as an avenue towards increasing their run-time efficiency, and even fewer have proposed strategies to increase such resiliency. We herein address this challenge in the context of Bit-line Computing architectures, an embodiment of the in-memory computing paradigm tailored towards CNN applications. We show that little additional hardware is required to add highly effective error detection and mitigation in such platforms. In turn, our proposed scheme can cope with high error rates when performing memory accesses with no impact on CNNs accuracy, allowing for very aggressive voltage scaling. Complementary, we also show that CNN resiliency can be increased by algorithmic optimizations in addition to architectural ones, adopting a combined ensembling and pruning strategy that increases robustness while not inflating workload requirements. Experiments on different quantized CNN models reveal that our combined hardware/software approach enables the supply voltage to be reduced to just 650mV, decreasing the energy per inference up to 51.3%, without affecting the baseline CNN classification accuracy.
Marco Rios, Flavio Ponzina, Giovanni Ansaloni, Alexandre Levisse, David Atienza 0001
ACM Great Lakes Symposium on VLSI2
2021 Running Efficiently CNNs on the Edge Thanks to Hybrid SRAM-RRAM In-Memory Computing
abstract
The increasing size of Convolutional Neural Networks (CNNs) and the high computational workload required for inference pose major challenges for their deployment on resource-constrained edge devices. in this paper, we address them by proposing a novel In-Memory Computing (IMC) architecture. Our IMC strategy allows us to efficiently perform arithmetic operations based on bitline computing, enabling a high degree of parallelism while reducing energy-costly data transfers. Moreover, it features a hybrid memory structure, where a portion of each subarray, dedicated to storing CNN weights, is implemented as high-density, zero-standby-power Resistive RAM. Finally, it exploits an innovative method for storing quantized weights based on their value, named Weight Data Mapping (WDM), which further increases efficiency. Compared to state-of-the-art IMC alternatives, our solution provides up to 93% improvements in energy efficiency and up to 6x less run-time when performing inference on Mobilenet and AlexNet neural networks.
Marco Rios, Flavio Ponzina, Giovanni Ansaloni, Alexandre Levisse, David Atienza 0001
DATE2
2021 E2CNNs: Ensembles of Convolutional Neural Networks to Improve Robustness Against Memory Errors in Edge-Computing Devices
abstract
To reduce energy consumption, it is possible to operate embedded systems at sub-nominal conditions (e.g., reduced voltage, limited eDRAM refresh rate) that can introduce bit errors in their memories. These errors can affect the stored values of convolutional neural network (CNN) weights and activations, compromising their accuracy. In this article, we introduce Embedded Ensemble CNNs (E2CNNs), our architectural design methodology to conceive ensembles of convolutional neural networks to improve robustness against memory errors compared to a single-instance network. Ensembles of CNNs have been previously proposed to increase accuracy at the cost of replicating similar or different architectures. Unfortunately, state-of-the-art (SoA) ensembles do not suit well embedded systems, in which memory and processing constraints limit the number of deployable models. Our proposed architecture solves that limitation applying SoA compression methods to produce an ensemble with the same memory requirements of the original architecture, but with improved error robustness. Then, as part of our new E2CNNs design methodology, we propose a heuristic method to automate the design of the voter-based ensemble architecture that maximizes accuracy for the expected memory error rate while bounding the design effort. To evaluate the robustness of E2CNNs for different error types and densities, and their ability to achieve energy savings, we propose three error models that simulate the behavior of SRAM and eDRAM operating at sub-nominal conditions. Our results show that E2CNNs achieves energy savings of up to 80 percent for LeNet-5, 90 percent for AlexNet, 60 percent for GoogLeNet, 60 percent for MobileNet and 60 percent for an optimized industrial CNN, while minimizing the impact on accuracy. Furthermore, the memory size can be decreased up to 54 percent by reducing the number of members in the ensemble, with a more limited impact on the original accuracy than obtained through pruning alone.
Flavio Ponzina, Miguel Peón-Quirós, Andreas Peter Burg, David Atienza 0001
IEEE Trans. Computers1