EDBT 2026 Demo / reviewers in the wild / expert
Arun M
dblp:30/9666
· DBLP profile ↗
6ranked-venue papers
6as first author
6since 2021 · last 2026
0009-0001-5007-0002ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Split-Posit DSP Slice Architecture for Embedded FPGA FabricsabstractThis work introduces the Split-Posit format, designed to mimic the benefits of high-precision Posit values at significantly lower bit widths, achieving a favourable trade-off between accuracy and hardware efficiency. Furthermore, we propose a specialized digital signal processing (DSP) architecture for embedded FPGAs tailored for Split-Posit arithmetic, efficient and well-suited for its variable-length field structure. Arun M, Madhav Rao |
FCCM | 1 |
| 2026 | PoSVM: Posit-Based Accelerator for SVM-driven Edge Biomedical Image ClassificationabstractBiomedical diagnostic platforms increasingly require continuous image acquisition and rapid, accurate clinical feedback, while operating under strict constraints on computational resources and power consumption. Although deep neural networks achieve high diagnostic accuracy, their substantial computational complexity and memory requirements limit their suitability for on-device inference. In contrast, classical machine learning techniques such as Support Vector Machines (SVMs) offer a more hardware-friendly alternative. This work introduces a resource-efficient hardware architecture for multi-class SVM inference that leverages Posit arithmetic. The proposed design employs a binary classification kernel with a one-versus-rest classification scheme to enable scalable multi-class decision-making while preserving architectural simplicity. By exploiting the tapered precision property of Posit arithmetic, the design enhances numerical stability during dot-product accumulation compared to conventional fixed-point implementations, while also reducing the hardware overhead typically associated with IEEE floating-point representations. Multiple Posit configurations, ranging from (4,2) to (16,2), are evaluated to explore trade-offs among silicon area, power consumption, latency, and classification accuracy on standard biomedical image datasets. Hardware evaluation shows a silicon area reduction of 66.5%, power savings of 64.5%, and a 1.29 × improvement in latency relative to conventional implementations, while still maintaining competitive accuracy. These results indicate that the proposed architecture is well-suited for deployment in next-generation portable and edge-assisted biomedical diagnostic systems. Arun M, Mihir S. Kagalkar, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | Split-Posit MAC Architectures for Efficient Neural Network ProcessingabstractThe rapid proliferation of Artificial Intelligence (AI) has significantly increased the demand for high-performance and energy-efficient Deep Neural Network (DNN) accelerators. Central to these accelerator architectures is the Multiply-Accumulate (MAC) unit, whose efficiency directly impacts overall system performance. Traditional MAC implementations based on fixed-point and floating-point arithmetic suffer from limitations such as large area, high power consumption, and restricted dynamic range, making them suboptimal for modern DNN workloads. Recently, the Posit number system, introduced by Gustafson, has emerged as a promising alternative due to its dynamic precision and reduced hardware complexity. However, empirical studies suggest that while high-bit-width Posit configurations maintain DNN accuracy and offer some efficiency benefits, low bit-width Posit formats result in significant accuracy degradation, limiting their practical utility. In this work, we propose a novel Split-Posit MAC architecture that bridges the gap between low-bit-width efficiency and high bit-width accuracy. By decomposing higher bit-width Posit computations into lower bit-width operations, the proposed architecture incurs a negligible drop in inference accuracy across standard TinyML benchmarks. Furthermore, ASIC implementation validates the practicality of our approach, reporting a 39.89% reduction in silicon footprint, a 39.47% improvement in energy efficiency, and a minor increase in delay. These results establish Split-Posit MAC as a highly efficient and scalable solution for next-generation AI accelerators. Arun M, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | Resource-Efficient DSP Slice for Embedded FPGAs Using Split-Posit ArithmeticabstractModern Machine Learning (ML) applications increasingly demand high-precision arithmetic within critical resource constraints. Although standard floating-point arithmetic remains widely used, it incurs huge hardware complexity. The recently introduced Posit format has emerged as a promising alternative, offering improved numerical efficiency with reduced hardware costs. Empirical studies indicate that, although Posit arithmetic units exhibit lower hardware complexity than their floating-point counterparts, they often require higher precision to achieve comparable numerical accuracy, thereby reducing the anticipated savings in resource utilization. To address this challenge, we introduce the Split-Posit format, designed to mimic the benefits of high-precision Posits at significantly lower bit widths, thereby achieving a favorable trade-off between accuracy and hardware efficiency. Furthermore, we propose a specialized digital signal processing (DSP) architecture for embedded FPGAs tailored for Split-Posit arithmetic. Unlike conventional FPGAs that rely on fixed DSP primitives and are ill-suited to Posit’s variable-length field structure, the proposed architecture enables more efficient resource utilization. Experimental results show that the proposed design achieves 53.1\(\%\) savings in fabric area, 61.1\(\%\) lower power consumption, and a 48\(\%\) reduction in critical-path delay per multiply–accumulate (MAC) operation compared to state-of-the-art (SOTA) implementations. The architecture is validated using standard benchmarks, demonstrating competitive performance and superior hardware efficiency, making it well-suited for next-generation reconfigurable systems-on-chip (SoCs). Arun M, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | Hardware-Efficient POSIT Based SVM Inference Engine for Intelligent Biomedical Systems
Arun M, Vaishnavi Sharma, Madhav Rao |
ISCAS | 1 |
| 2025 | Efficient POSIT Multiplier with Multi-flag Priority Encoding and Multistage Booth ProcessingabstractRecent advancements in arithmetic systems have emphasized the need for faster, resource-efficient operations without compromising on the range or accuracy provided by the IEEE 754 floating-point format. Consequently, the Posit number system has gained attention for its computational efficiency and broader dynamic range. However, multiplication remains one of the most computationally expensive operations in Posit arithmetic, with traditional multipliers suffering from additional overhead due to complex decoders. This paper proposes a novel and efficient Posit multiplier architecture that addresses these challenges by modifying the conventional decoder. Specifically, we introduce multiflag priority encoders to streamline the decoding process, thereby leveraging the silicon footprint utilized by the design and benefiting critical path delay. We also present a modified multistage Booth encoding technique, optimized for the Posit format, to enhance multiplier speed without compromising on accuracy. Additionally, the clock-gating technique applied to the multi-level encoding architecture selectively disables clock signals during idle cycles, ensuring further savings in power. The combination of new decoders and power optimization strategies establishes a hardware-efficient Posit multiplier, which was validated for a DCT compression application. Implementation results demonstrate that the proposed Posit multiplier offers power savings up to $46 \%$ with respect to the state-of-the-art (SOTA) Posit and other floating-point multipliers. Furthermore, the proposed architecture achieves a $48 \%$ silicon footprint savings. The hardware efficiency of the proposed Posit multiplier, along with the acceptable DCT-compressed image quality results, validates the novel contribution and its effectiveness. All the design files are made freely available for further usage by the designers and research community. Arun M, Madhav Rao |
DSD | 1 |