Anuj Grover

dblp:156/4054 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-6057-4984ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AnalogSim: a Modeling Framework for SRAM-based Analog In-Memory Computing
Talha Bin Aslam, Anuj Grover, Harsh Rawat
ISCAS2
2025 Sustainably Secure: ChaCha20 Encryption Based on In-Memory Compute
abstract
Due to its high speed and simple structure, the ChaCha20 algorithm offers an alternative to AES, making it well-suited for lightweight encryption in IoT devices [1]. However, its execution on traditional Von Neumann architectures incurs significant energy overhead due to frequent memory accesses. In-memory compute (IMC) architectures address this challenge by allowing computation directly within memory. Although these architectures are characterized by straightforward power performance and area parameters, designers leave out other critical design considerations that impact sustainability. As technology continues to evolve rapidly, it becomes increasingly important to assess the environmental impact of large-scale design architectures. In this paper, we present a secure and sustainable implementation of the ChaCha20 encryption algorithm using IMC architectures. We evaluated the trade-offs of implementing ChaCha20 in 6T, 9T, and 10T SRAM bitcell architectures by introducing a sustainability-focused metric alongside conventional parameters. The developed sustainability metric evaluates architectures based on the combined impact of the embodied and operational carbon footprint. Additionally, we assess the security of these architectures against power-based side-channel attacks using Welch’s t-test. Among the evaluated designs, the 6T implementation achieves the smallest area footprint but at the cost of degraded performance. In contrast, the 9T implementation preserves the performance of the 10T design while reducing area overhead. Furthermore, sustainability analysis based on the introduced metric identifies the 6T bitcell as the most sustainable option.
Samridhi Jain, Mohd Aamir, Anuj Grover
ISLPED3
2024 Design Of High-Density Iso-Stable Asymmetric Memory Cell With Upto 10X Reduced Leakage
abstract
Embedded memories occupy up to 70% of the die area in modern digital SoCs. Therefore, high-density, low-leakage SRAM cells are desirable. We propose an asymmetrical 4TA SRAM cell and design it and conventional 6T, 5T, and 4T SRAM cells under iso-stable constraints. We show that the proposed 4TA cell is denser by up to 16% and has about 10X lower leakage than the iso-stable conventional 4T cell.
Ajay Shroti, Anuj Grover
ISCAS2
2022 3-Stage Pipelined Hierarchical SRAMs with Burst Mode Read in 65nm LSTP CMOS
abstract
As the technology scaling has happened, logic performance has improved at a much faster rate than SRAM performance. Therefore, in SRAMs, required performance improvement is achieved by design and architecture level changes. This work presents pipelined hierarchical SRAM and compares it with a conventional non-hierarchical SRAM design on the axes of Performance, Power, and Area. We show that a pipelined SRAM of size 8192×64 m16 with integrated burst mode, operates at 40% lesser dynamic power and is 31% faster than a conventional non-hierarchical SRAM design in 65nm Low stand-by Power(LSTP) CMOS technology. When compared to hierarchical design, it operates at 19% lesser dynamic power and is 17% faster with an area increase of 5%.
Mukesh Kumar Srivastav, Rimjhim, Roshan Mishra, Anuj Grover, Kedar Janardan Dhori, Harsh Rawat
ISCAS4
2017 LoCCo-Based Scan Chain Stitching for Low-Power DFT
abstract
Power dissipation during scan testing of a systemon-chip can be significantly higher than that during functional mode, causing reliability and yield concerns. This paper proposes a logic cluster controllability (LoCCo)-based scan chain stitching methodology to achieve low-power testing. The scan chain stitching is made power aware by placing flip-flops with higher test combination requirements at the beginning of scan chains, while flip-flops with lower test combination requirements are put toward the end of scan chains. The test combination requirements are estimated through a simple logic cluster and flip-flop controllability identification algorithm. This method helps in consolidating care bits toward the beginning of scan chains. Hence, a significantly lower shift-in transition is achieved in the test patterns. The results from ITC'99 and industrial designs in 28FDSOI and 40-nm CMOS technologies show a total shift-in transition reduction of up to 23.1% and average shift power reduction of up to 21.6% using the proposed method. The use of LoCCo methodology posed a negligible routing congestion overhead in the layout compared to the conventional method. LoCCo is also used as a base to apply other vector reordering low-power methods and gain 3.5× reduced computation time with almost similar power reduction as achieved by Bonhomme et al. independently.
Shalini Pathak, Anuj Grover, Mausumi Pohit, Nitin Bansal
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Scan Chain Adaptation through ECO
abstract
The paper proposes a Scan Chain Adaptation Through ECO (SCATE) methodology to accommodate scan chain changes at advanced physical design stage into the existing DFT compression logic and eliminates the need of re-running the complete scan insertion flow. With the proposed methodology scan chain changes can be implemented as ECO to achieve cycle time reduction with minimum overhead. It can handle last minute changes due to the addition of new scan chains or changes in scan chain lengths in existing IPs. The scope of the work is also extended to reduce the cycle time of a derivative product from a base design. Results from industrial designs in 28FDSOI and 40nm CMOS technologies show SOC test coverage of above 99% when new scan chains were included using the proposed methods. Test time gain of more than 30% was observed by adjusting the maximum scan chain lengths. Use of SCATE methodology enabled DFT modifications through ECO and saved design cycle time by avoiding loopback to fresh DFT insertion, Timing analysis and Physical design flows in industrial SOCs
Jasvir Singh, Anuj Grover, Mausumi Pohit, Anurag Singh Baghel, Gurjit Kaur, Shalini Pathak
ATS2
2016 Static Noise Margin based Yield Modelling of 6T SRAM for Area and Minimum Operating Voltage Improvement using Recovery Techniques
abstract
In advanced technology nodes, the process variations deteriorate SRAM performance and greatly affect yield. It is necessary to formulate yield estimation models to optimize SRAMs and effectively trade-off area, performance and robustness. We propose models that in addition to enabling yield estimates also enable evaluation of lowering minimum operational voltage (VDDMIN). We present a quantitative analysis for SNM limited SRAM yield using Design of Experiments (DOE) method. The proposed framework for yield based design can also utilize recovery techniques like Error Correcting Codes (ECC) and redundancy and quantifies yield, area, and VDDMIN improvements. We also present a case study that trades-off ECC recovery budget, VDDMIN and area gain. We show 25% improvement in area and VDDMIN lowering by 300mV at constant yield levels by using 50% of ECC recovery budget.
Nidhi Batra, Pawan Sehgal, Shashwat Kaushik, Mohammad S. Hashmi, Sudesh Bhalla, Anuj Grover
ACM Great Lakes Symposium on VLSI6
2013 Ultra-wide voltage range designs in fully-depleted silicon-on-insulator FETs
abstract
Todays' MPSoC applications are requiring a convergence between very high speed and ultra low power. Ultra Wide Voltage Range (UWVR) capability appears as a solution for high energy efficiency with the objective to improve the speed at very low voltage and decrease the power at high speed. Using Fully Depleted Silicon-On-Insulator (FDSOI) devices significantly improves the trade-off between leakage, variability and speed even at low-voltage. A full design framework is presented for UWVR operation using FDSOI Ultra Thin Body and Box technology considering power management, multi-VT enablement, standard cells design and SRAM bitcells. Technology performances are demonstrated on a ARM A9 critical path showing a speed increase from 40% to 200% without added energy cost. In opposite, when performance is not required, FDSOI enables to reduce leakage power up to 10X using Reverse Body Biasing.
Edith Beigné, Alexandre Valentian, Bastien Giraud, Olivier Thomas, Thomas Benoist, Yvain Thonnart, Serge Bernard, Guillaume Moritz, Olivier Billoint, Y. Maneglia, Philippe Flatresse, Jean-Philippe Noël, Fady Abouzeid, Bertrand Pelloux-Prayer, Anuj Grover, Sylvain Clerc, Philippe Roche, Julien Le Coz, Sylvain Engels, Robin Wilson
DATE15