Brajesh Kumar Kaushik

dblp:93/742 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-6414-0032ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Quantized Magnetic Domain Wall Synapse for Efficient Deep Neural Networks
abstract
The quantization of synaptic weights using emerging nonvolatile memory (NVM) devices has emerged as a promising solution to implement computationally efficient neural networks on resource constrained hardware. However, the practical implementation of such synaptic weights is hampered by the imperfect memory characteristics, specifically the availability of limited number of quantized states and the presence of large intrinsic device variation and stochasticity involved in writing the synaptic states. This article presents ON-chip training and inference of a neural network using quantized magnetic domain wall (DW)-based synaptic array and CMOS peripheral circuits. A rigorous model of the magnetic DW device considering stochasticity and process variations has been utilized for the synapse. To achieve stable quantized weights, DW pinning has been achieved by means of physical constrictions. Finally, VGG8 architecture for CIFAR-10 image classification has been simulated by using the extracted synaptic device characteristics. The performance in terms of accuracy, energy, latency, and area consumption has been evaluated while considering the process variations and nonidealities in the DW device as well as the peripheral circuits. The proposed quantized neural network (QNN) architecture achieves efficient ON-chip learning with 92.4% and 90.4% training and inference accuracy, respectively. In comparison to pure CMOS-based design, it demonstrates an overall improvement in area, energy, and latency by , , and , respectively.
Seema Dhull, Walid Al Misba, Arshid Nisar, Jayasimha Atulasimha, Brajesh Kumar Kaushik
IEEE Trans. Neural Networks Learn. Syst.5
2024 Novel Radiation Hardened Pre-Discharge Sense Amplifier for Double Data Rate Magnetic Random Access Memory
abstract
The ever-increasing demand for non-volatile memories with high density, low power consumption, and resistance to radiation has led to the emergence of Double Data Rate Magnetic Random Access Memory (DDR-MRAM) as a promising technology. However, the reliability of DDR-MRAM in harsh radiation environments remains a challenge. This paper presents a novel radiation-hardened pre-discharge sense amplifier specifically designed for DDR-MRAM application. The paper investigates the effects of radiation on the unhardened pre-discharge sense amplifier (PDSA) circuit and proposes solutions at both the device and circuit levels. At the circuit level, two distinct circuits have been proposed: one involves the duplication technique (PDSA-D) and the other employs a feedback circuitry (RT-PDSA) to harden sensitive transistors in the PDSA. The proposed radiation tolerant PDSA (RT-PDSA) circuit demonstrates significant improvements in terms of single event upset (SEU) as well as double node upset (DNU) mitigation by considering minimum sensitive nodes layout area separation concept and enhanced compared to unhardened PDSA circuit. The overall performance of the proposed circuit has been analyzed in terms of the figure of merit (FOM), which includes critical charge tolerability, the number of sensitive nodes, recovery time, and area. The proposed RT-PDSA circuit has 11.95$\times$and 8.7$\times$higher FOM compared to the hardened PDSA at the device level, and PDSA-D circuit, respectively. Moreover, the proposed circuit exhibits robustness against process variations, making it a promising solution for enhancing DDR-MRAM’s reliability in radiation-intensive environments.
Alok Kumar Shukla, Brajesh Kumar Kaushik
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 A survey of SRAM-based in-memory computing techniques and applications
Sparsh Mittal, Gaurav Verma 0002, Brajesh Kumar Kaushik, Farooq Ahmad Khanday
J. Syst. Archit.3
2021 Efficient Method and Architecture for Real-Time Video Defogging
abstract
Real-time video defogging has a huge demand in intelligent transportation, advanced driver assistance systems (ADAS), long-range surveillance, autonomous aerial vehicles and endoscopic surgery. Most of these applications are constrained by stringent frame rate, power and memory budget. Till date, the methods devised for such requirements are very rare. Therefore, this paper proposes an efficient method and very-large-scale integration (VLSI) architecture for resource-constrained embedded system targeting real-time video defogging. The method and architecture are co-designed to achieve high throughput while consuming less resources and power. The architecture is divided into four parts namely, atmospheric light estimation unit, airlight adjustment unit, transmission estimation unit, and pixel restoration unit. The atmospheric light estimation unit employs a$3\times 3$tile based approach that eliminates the requirement of large buffer memory for$15\times 15$support size. The transmission is estimated using dark channel prior and gradient threshold based Gaussian filtering approach. In order to overcome flickering artifacts, an adaptive airlight updating scheme is employed. From the quantitative and qualitative evaluations, it is observed that the proposed method outperforms the existing hardware approaches. Furthermore, the field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) implementations of the architecture achieve a high throughput of 200 MPixels/s and 600 MPixels/s, respectively. The ASIC implementation dissipates 9.03 mW power at 200MHz. Moreover, the proposed design does not require any external memory such as dynamic random-access memory (DRAM), thus making it suitable for on-chip processing that can be closely integrated with an image sensor.
Rahul Kumar 0007, Balasubramanian Raman, Brajesh Kumar Kaushik
IEEE Trans. Intell. Transp. Syst.3
2021 Novel Architecture for Lifting Discrete Wavelet Packet Transform With Arbitrary Tree Structure
abstract
This brief presents a novel pipelined VLSI architecture for computing discrete wavelet packet transform (DWPT) with an arbitrary wavelet tree. Coefficients for different levels are computed in a series of stages. Each stage consists of a bypassed wavelet filter and circuit for reordering intermediate coefficients. The proposed lifting-based wavelet filter computes high- and low-pass coefficients in series. In order to accommodate the arbitrary tree structure, the filter either computes the coefficients or bypass the samples. The reordering of intermediate coefficients forms a subband required for next-level computation. The coefficients are computed in a serial manner and reordering of intermediate coefficients reduce not only the memory elements but also the circuit complexity. The proposed pipelined architecture reduces the requirement of memory elements by 50%. Furthermore, the hardware implementation results show that the area and power requirement are reduced by 33% and 20%, respectively.
Gyanendra Singh, Samba Raju Chiluveru, Balasubramanian Raman, Manoj Tripathy, Brajesh Kumar Kaushik
IEEE Trans. Very Large Scale Integr. Syst.5
2020 High-Density, Low-Power Voltage-Control Spin Orbit Torque Memory with Synchronous Two-Step Write and Symmetric Read Techniques
abstract
Voltage-control spin orbit torque (VC-SOT) magnetic tunnel junction (MTJ) has the potential to achieve high-speed and low-power spintronic memory, owing to the adaptive voltage modulated energy barrier of the MTJ. However, the three-terminal device structure needs two access transistors (one for write operation and the other one for read operation) and thus occupies larger bit-cell area compared to two terminal MTJs. A feasible method to reduce area overhead is to stack multiple VC-SOT MTJs on a common antiferromagnetic strip to share the write access transistors. In this structure, high density can be achieved. However, write and read operations face problems and the design space is not sure given a strip length. In this paper, we propose a synchronous two-step multi-bit write and symmetric read method by exploiting the selective VC-SOT driven MTJ switching mechanism. Then hybrid circuits are designed and evaluated based a physics-based VC-SOT MTJ model and a 40nm CMOS design-kit to show the feasibility and performance of our method. Our work enables high-density, low-power, high-speed voltage-control SOT memory.
Wang Kang 0001, Liuyang Zhang, He Zhang 0011, Brajesh Kumar Kaushik, Weisheng Zhao 0001
DATE5
2020 Fast and robust video stabilisation with preserved intentional camera motion and smear removal for infrared video
abstract
Border military surveillance is one of the demanding and challenging tasks for any nation. Thermal (infrared) camera, which works on the infrared domain, provides complete visual sequences even in pitch dark night conditions. When the video is recorded from a thermal camera mounted on a vehicle, the output video is unstabilised with poor visual quality due to unintentional camera motion. These unwanted motions also introduce smear. Furthermore, there may be situations when a camera is moved intentionally to capture target. This study proposes a fast and robust algorithm for auto stabilisation of videos with smear removal while keeping the intentional motion of camera. This algorithm is developed under the framework of speeded up robust features matching. The proposed algorithm is capable of correcting both motions, i.e. translation as well as rotational. Quality improvement of up to 21 dB is achieved in the stabilised output videos.
Sudhir Khare, Manvendra Singh, Brajesh Kumar Kaushik
IET Image Process.3
2019 Multispectral Transmission Map Fusion Method and Architecture for Image Dehazing
abstract
Image dehazing is an essential preprocessing stage for several applications such as surveillance, long-range imaging, automatic driver-assistance system (ADAS), and remote sensing. Haze particles degrade visible-band images more severely than the infrared images due to scattering phenomena. Therefore, infrared data can be effectively utilized to enhance the visibility of color images captured in a hazy environment. In this brief, for the first time, a multispectral transmission map fusion approach is presented, wherein a haze-aware weight generation scheme is used to improve transmission map in specific regions. This achieves significant quantitative improvement as compared with the existing methods. In addition, the method is extremely suitable for hardware implementation. Therefore, a low-cost very large-scale integration (VLSI) architecture is also presented. The application-specified integrated circuit (ASIC) and field-programmable gate array (FPGA) implementations dissipate 3.45- and 57-mW power at 150 MHz, respectively. Both ASIC and FPGA achieve a high throughput of 450 and 308 Mpixels/s, respectively. The design does not require any external memory such as dynamic random access memory, thus making it suitable for on-chip processing that can be closely integrated with an image sensor.
Rahul Kumar 0007, Brajesh Kumar Kaushik, R. Balasubramanian
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Transmission Coefficient Matrix Modeling of Spin-Torque-Based $n$ -Qubit Architecture
Anant Aravind Kulkarni, Sanjay Prajapati, Brajesh Kumar Kaushik
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Low-Power High-Density STT MRAMs on a 3-D Vertical Silicon Nanowire Platform
abstract
In recent years, researchers have focused toward reduction in power dissipation and cell size to employ spin-transfer torque (STT) magnetic random-access memories (MRAMs) for embedded applications. Hence, the magnetic tunnel junctions (MTJs) with an optimized structure and magnetic properties are being explored to reduce the switching current. However, the switching current reduction in the MTJs generally lowers the data-retention capability. Hence, a different approach to reduce power dissipation using a novel select device should be considered. This paper, therefore, explores the STT MRAM with vertical silicon nanowire gate all around (GAA) high-k select device for superior performance. The MTJ is stacked above the vertical GAA device, so that both occupy the same footprint area to achieve high array density. Furthermore, enhancement of current drive using high-k gate dielectric and its impact on the STT MRAMs are analyzed at different feature sizes. The proposed STT MRAM cell with high-k dielectric (HfO2) lowers the power dissipation by 8%-25% and increases the write margins (WMs) up to 38%, with negligible increment in delay in comparison with the GAA device using low-k dielectric (SiO2). Moreover, asymmetricity is introduced in device configuration to achieve power savings of 25%-30% at high VDD. The proposed asymmetric high-k cell offers a substantially larger tradeoff window between high WMs and low power dissipation.
Shivam Verma, Brajesh Kumar Kaushik
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Crosstalk Analysis for a CMOS-Gate-Driven Coupled Interconnects
abstract
This paper deals in crosstalk analysis of a CMOS-gate-driven capacitively and inductively coupled interconnect. Alpha power-law model of a MOS transistor is used to represent a CMOS driver. This is combined with a transmission-line-based coupled-interconnect model to develop a composite driver-interconnect-load model for analytical purposes. On this basis, a transient analysis of crosstalk noise is carried out. Comparison of the analytical results with SPICE extracted results shows that the average error involved in estimating noise peak and their time of occurrence is less than 7%.
Brajesh Kumar Kaushik, Sankar Sarkar
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2007 Waveform analysis and delay prediction for a CMOS gate driving RLC interconnect load
Brajesh Kumar Kaushik, Sankar Sarkar, Rajendra Prasad Agarwal
Integr.1