Mounika Vaddeboina

dblp:304/3712 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0002-7380-0749ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Energy-Efficient Neural Network Inference through Golomb-Rice Compression of Activations for Edge Devices
abstract
A key challenge for Deep Neural Network (DNN) inference on resource-constrained edge devices is the high energy consumption caused by frequent memory accesses for parameters. While our previous research has demonstrated the efficacy of data compression for weights, this paper extends our approach to include on-the-fly compression and decompression of activations. We propose a comprehensive hardware solution comprising two main components: a Golomb-Rice (GR) compression system and an Output Activation Processing Module (OAPM). The GR system provides an efficient activation compression mechanism, while the OAPM enables dynamic data-format capabilities for handling activations. Additionally, we present an enhanced Input Activation Extract Module (IAEM) with an integrated decompression unit and dynamic activation processing capabilities. When integrated with an industry-strength Neural Network accelerator and evaluated using the Anomaly Detection (AD) TinyML benchmark, our lossless compression system achieved a 2.3× compression ratio, reduced memory bandwidth usage by 49.28%, and improved inference speed by 10%.
Mounika Vaddeboina, Alper Yilmazer, Wolfgang Ecker
DDECS1
2024 PaGoRi:A Scalable Parallel Golomb-Rice Decoder
abstract
Deep Neural Networks (DNNs) have created opportunities to address real-world issues and expand the application of Artificial Intelligence (AI). Despite significant accuracy enhancements, DNNs pose a challenge when deployed on resource-limited edge devices commonly used in Internet of Things (IoT) applications. Inference execution of the DNNs requires accessing millions of parameters responsible for most energy consumption. Compression of weights is one possible solution, but most of the existing hardware decompression units could be more efficient in terms of power, area, and energy. This paper presents a scalable version of a hardware-efficient Parallel Golomb-Rice decoder (PaGoRi). The decoder has been integrated with an industry-strength Neural Network (NN) accelerator and evaluated with three TinyML benchmarks. The PaGoRi decoder achieves optimal trade-offs between power consumption and throughput, supporting decoding capacities of four and eight weights, consuming 0.43 mW and 0.79 mW of power, respectively, while achieving a throughput of 888 MBps and 1.3 GBps, respectively.
Mounika Vaddeboina, Endri Kaja, Alper Yilmazer, Uttal Ghosh, Wolfgang Ecker
DDECS1
2024 Optimizing Data Compression: Enhanced Golomb-Rice Encoding with Parallel Decoding Strategies for TinyML Models
abstract
Deep Neural Networks (DNNs) offer possibilities for tackling practical challenges and broadening the scope of Artificial Intelligence (AI) applications. The demanding memory requirements of present-day neural networks can be attributed to the rising intricacy of network architectures. These designs encompass multiple layers with an extensive number of parameters, leading to heightened demands on memory storage. The energy consumption during the inference execution of DNNs is predominantly attributed to the access and processing of these parameters. To tackle the significant size of models integrated into Internet of Things (IoT) devices, a promising strategy involves diminishing the bit width of weights. This paper introduces an improved version of Golomb-Rice (GR) encoder and an optimized Parallel Golomb-Rice decoder that can support sparse and non-sparse DNNs. To evaluate the encoder's and decoder's efficiency, we conducted two sets of experiments using three TinyML benchmarks, one without pruning and the other incorporating pruning. The results highlight that the encoder demonstrates a Compression-Ratio (CR) superior to that of Huffman encoding, and the decoder exhibits an energy efficiency of up to 2.6 TBps/W and 2.7 TBps/W for four- and eight-weight decoding, respectively.
Mounika Vaddeboina, Alper Yilmayer, Wolfgang Ecker
DSD1
2023 Parallel Golomb-Rice Decoder with 8-bit Unary Decoding for Weight Compression in TinyML Applications
abstract
Due to the recent advances in AI, the requirement for Artificial Intelligence (AI) has increased exponentially in the domain of Internet of Things (IoT). Running Deep Neural Networks (DNNs) on edge devices gives the advantage of privacy, security, and lower latency. It is challenging to deploy them on embedded devices with constrained hardware resources since a lot of compute and memory resources are required. Memory access contributes to the majority of the energy requirements on edge devices. Although data compression plays a critical role in reducing storage and memory bandwidth requirements, most of the hardware decoders are inefficient in terms of power, area, and throughput. In this work, a hardware Parallel Golomb-Rice decoder is presented that can decode 8-bits of unary encoded data every cycle. The design has been integrated with a Neural Network (NN) accelerator and experimented with state-of-the-art benchmark models. Lossless compression is performed with an offline Golomb-Rice encoder. It encodes the weights of each layer with an optimum Golomb-Rice parameter. Applied to the benchmarks Anomaly Detection, Image Classification and Visual Wake Words the memory access during inference is reduced by 26.8%, 6.62% and 5.54% respectively. The decoder dissipates 0.4216 mW of power and delivers an average throughput of 860 MBps. The design has been synthesised with 40 nm technology and compared with state-of-the-art works.
Mounika Vaddeboina, Endri Kaja, Alper Yilmayer, Sebastian Siegfried Prebeck, Wolfgang Ecker
DSD1