Babita Jajodia

dblp:159/7490 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-2479-930XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Division-Free Four-Way Toom-Cook Polynomial Multiplication Architecture for Large Integer Arithmetic on FPGAs and ASICs
abstract
Polynomial multiplication performs as the most important task and is computationally extensive in cryptographic algorithms. Of the several polynomial multiplications, theoretically, the Toom-Cook multiplication is the most efficient. In this paper, an efficient four-way Toom-Cook (Toom-4) multiplication method has been presented that provides accurate result as the product and can be efficiently implemented in hardware. This is done by eliminating the effects of exact divisions present in the intermediate steps of the computation. In addition to these, the hardware implementation of the proposed work suggests that this multiplication method performs better in terms of resource utilization and delay minimization compared to other existing Toom-4 algorithm methods. Hardware implementation of the proposed Toom-4 multiplication architecture is done using a Virtex-7 FPGA device in Xilinx (now AMD) ISE platform. Here, the hardware implementation of the proposed multiplier is performed for 128 bits, 256 bits, and 512 bits with advantages in the area-time-product (ATP) over previous researches; ATP is the figure of merit in determining the performance of the design. Practically for 512 input bits, the percentage ATP of the proposed design is 98.810%, 23.537% and 57.316% better compared to Toom-2 (Karatsuba) multiplication, Hybrid Toom-2 (Karatsuba-Comba) multiplication, and existing Toom-4 multiplication, respectively. Moreover, ASIC implementation results on United Microelectronics Corporation (UMC) 65-nm technology demonstrate a notable reduction in area-time-product (ATP) and power-time-product (PTP) compared to state-of-the-art works.
Monalisa Das, Babita Jajodia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Area and Delay Trade-Offs in Three-Way Toom-Cook Large Integer Multipliers Implemented on FPGAs
abstract
The demand for efficient large integer polynomial multiplications in present day crypto-systems is the need of the hour. Toom-Cook multiplication algorithm being one of the most efficient multiplication algorithm is discussed in this work. The limitations of large integer polynomial multiplications using three-way Toom-Cook (Toom-3) multiplication algorithm and the methods to overcome it are presented in this paper. This is done by implementing two different division-free methods for symmetric Toom-3 multiplication as well as for asymmetric Toom-3 multiplication based on the input operand size (N). Hardware implementations of the proposed multiplication methods are done using Virtex-7 Field Programmable Gate Array (FPGA) device in Xilinx ISE platform. The trade-off between hardware utilization and speed is noted, and the overall performance of the proposed design methods are measured by calculating Area-Time-Product (ATP). Practically, it has been observed that both the proposed Toom-3 multiplication methods performs better compared to the existing state-of-the-art designs.
Monalisa Das, Babita Jajodia
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 ART-MAC: Approximate Rounding and Truncation based MAC Unit for Fault-Tolerant Applications
abstract
In recent times, approximate computing has emerged as a promising technique to achieve significant power and energy benefits in computational systems. It is widely employed in fault-tolerant computationally intensive applications that require large arithmetic blocks. Applications such as image processing and machine learning often invoke the Multiply-Accumulate (MAC) unit for convolution operations. This paper proposes a novel architecture for an (unsigned × unsigned) approximate rounding and truncation based MAC unit named ART-MAC. It replaces the accurate multiplier architecture with an approximate multiplier proposed along with this work, thus improving the overall Quality of Results (QoR). The proposed design consumes 35.35% less power and showcases a significant speedup of 1.23 times when compared to the conventional MAC unit. On an average, the ART-MAC consumes 7.44% lesser on-chip area and showcases 13.49% lesser power-delay-product (PDP) compared to existing state-of-the-art designs.
Vishesh Mishra, Divy Pandey, Sagar Satapathy, Kaustav Goswami 0002, Babita Jajodia, Dip Sankar Banerjee
ISCAS6