VLDB 2026 Research / reviewers in the wild / expert
Arpan Jain
dblp:154/8472
· DBLP profile ↗
17ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-2522-8522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Low Supply, Impedance-Boosted Current Mirror Using Back-Gate in FD-SOI Technology
Roopesh G. L, Arpan Jain, Soham Bhattacharyya, Ashfakh Huluvallay, Andleeb Zahra, Zia Abbas |
ISCAS | 2 |
| 2026 | A 0.5V 72-pW Process and Temperature Compensated Voltage Reference
Vikkram Srinivasan, Damini Chandi Priya A, Ashfakh Huluvallay, Arpan Jain, Abhishek Pullela, Andleeb Zahra, Zia Abbas |
ISCAS | 4 |
| 2025 | Low-Power Voltage Reference: Review & ProgressabstractThis paper explores advanced voltage reference designs with extremely low power consumption, specifically under one microwatt. These designs are crucial for portable electronic systems that rely on minimal power. The paper categorizes these voltage references based on the types of devices used to generate temperature-proportional or complementary signals. It also provides key design examples to illustrate these concepts. Additionally, the paper enhances understanding by conducting a comparative analysis of performance metrics, including power consumption, temperature coefficient, physical size, and resistance to variations in the manufacturing process. Abhishek Pullela, Ashfakh Huluvallay, Arpan Jain, Zia Abbas, Inhee Lee 0001 |
VTS | 3 |
| 2024 | A Single-Point, Auto-Calibration Technique For PTAT/CTAT Resistance Based Current ReferencesabstractIn this paper, a cost-effective and easy-to-implement auto-trim technique is introduced for PTAT or CTAT resistance based current references. The resistance with a process-insensitive temperature coefficient requires only a single point trim at room temperature, achieving process and temperature-insensitive current. The approach utilizes a data comparison method where on-chip current data is compared with off-chip reference data to trim the resistance of the current reference. The off-chip reference data is generated using a low-cost external resistor that is used only during the trimming operation. The proposed trimming sensor effectively calibrates the current, offering precision close to manual trimming, leading to cost, time, and resource savings. To validate the working of the proposed technique, the auto trim sensor with on-chip current reference is designed in TSMC 180nm technology. The auto trim sensor calibrates the on-chip current from ±25% (due to voltage and resistance process variation) to ±1.5% across the process and 3σ mismatch. Arpan Jain, Ashfakh Ali, Dheekshith Akula, Abhishek Pullela, Zia Abbas |
ISCAS | 1 |
| 2023 | A 162nW, 0.845pJ/step Resistance-to-Digital Converter for Miniature Battery-Powered Sensing SystemsabstractThis paper proposes a 162nW resistance-to-digital converter (RDC) for miniature battery-powered sensing systems. The RDC first converts input resistance to a pulse by charging a capacitor to a threshold voltage with a current proportional to the resistance. It compensates temperature sensitivity of the charging current by generating the threshold voltage with the same temperature dependency. Then, the circuit digitizes the pulse using an up-down counter that cancels temperature-dependent delay and offset of the low-power comparator in a digital Correlated Double Sampling (CDS) style. Designed in a 180 nm CMOS process, the proposed circuit achieves a figure-of-merit (FoM) of 0.845pJ/c.s. in simulation, with a conversion time of 50 ms for input resistance from$50\mathrm{k}\Omega$to$1\mathrm{M}\Omega$, while consuming 162nW at a supply voltage of 900 mV. Also, it obtains a temperature sensitivity of 26.9ppm/°C from −40 to 100°C. Compared with the state-of-the-art RDCs, this work improves the FoM and temperature sensitivity by 42.91% and 11.52%, respectively. Arnab Dey 0002, Inhee Lee 0001, Ashfakh Ali, Arpan Jain, Abhishek Pullela, Zia Abbas |
ISCAS | 4 |
| 2023 | A 7 nW, 1 kHz, -40-170°C Relaxation Oscillator with Switch-Leakage Compensation for Low-Power High-Temperature IoT SystemsabstractThis paper proposes a low-power relaxation oscillator for low-power high-temperature IoT systems. It generates a 959 Hz clock signal from −40 to 170°C, consuming 6.75 nW at 0.65 V. A proposed switch-leakage compensation scheme nullifies the effects of body diode and subthreshold leakages on oscillator output frequency at high temperatures, thereby obtaining a wide operating temperature range. The oscillator implemented in a 180 nm CMOS process achieves a temperature coefficient of 40 ppm/°C from −40 to 170 °C at 0.65 V and a line sensitivity of 0.5 %/V from 0.65 to 2.4 V at room temperature, in simulation. Compared with state-of-the-art sub-$\mu\mathrm{W}$oscillators, this circuit obtains the highest operating temperature and the maximum temperature range. Ashfakh Huluvallay, Abhishek Pullela, Ehab A. Hamed, Arpan Jain, Naveen Dasari, Zia Abbas, Inhee Lee 0001 |
ISCAS | 4 |
| 2022 | AccDP: Accelerated Data-Parallel Distributed DNN Training for Modern GPU-Based HPC ClustersabstractDeep Learning (DL) has become a prominent machine learning technique due to the availability of efficient computational resources in the form of Graphics Processing Units (GPUs), large-scale datasets and a variety of models. The newer generation of GPUs are being designed with special emphasis on optimizing performance for DL applications. Also, the availability of easy-to-use DL frameworks—like PyTorch and TensorFlow— has enhanced productivity of domain experts to work on their custom DL applications from diverse domains. However, existing Deep Neural Network (DNN) training approaches may not fully utilize the newly emerging powerful GPUs like the NVIDIA A100—this is the primary issue that we address in this paper. Our motivating analyses show that the GPU utilization on NVIDIA A100 can be as low as 43% using traditional DNN training approaches for small-to-medium DL models and input data size. This paper proposes AccDP—a data-parallel distributed DNN training approach—to accelerate GPU-based DL applications. AccDP exploits the Message Passing Interface (MPI) communication library coupled with the NVIDIA’s Multi-Process Service (MPS) to increase the amount of work assigned to parallel GPUs resulting in higher utilization of compute resources. We evaluate our proposed design on different small-to-medium DL models and input sizes on the state-of-the-art HPC clusters. By injecting more parallelism into DNN training using our approach, the evaluation shows up to 58% improvement in training performance on a single GPU and up to 62% on 16 GPUs compared to regular DNN training. Furthermore, we conduct an in-depth characterization to determine the impact of several DNN training factors and best practices—including the batch size and the number of data loading workers— to optimally utilize GPU devices. To the best of our knowledge, this is the first work that explores the use of MPS and MPI to maximize the utilization of GPUs in distributed DNN training. Nawras Alnaasan, Arpan Jain, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HIPC | 2 |
| 2022 | A 156pW Gate-Leakage Based Voltage/Current Reference for Low-Power IoT SystemsabstractThe paper presents a sub-nW gate-leakage based voltage and current reference in a single circuit whose reference values are scalable and doesn’t incorporate start-up circuits or resistors in the architecture. The power consumption of the proposed circuit increases by only 2.1x in the temperature range of -55°C to 100°C, unlike conventional voltage/current references where the power consumption increases exponentially w.r.t temperature. Implemented in 90nm technology, the proposed voltage reference (current reference) achieves post-trim typical accuracy of 22ppm/°C(58ppm/°C) and worst-case accuracy of 71ppm/°C(78ppm/°C). Excellent line sensitivities of 0.029%/V and 0.059%/V are observed for voltage and current reference respectively, in a supply range of 1V - 3V. Without any start-up circuit, the observed 99% settling times for voltage and current reference are 1.92ms and 2.526ms respectively. The area occupied by the total circuit is 0.0015mm2, while the power consumption is 156pW at typical corner, 27°C and 1V supply. Abhishek Pullela, Ashfakh Ali, Arpan Jain, Inhee Lee 0001, Zia Abbas |
ISCAS | 3 |
| 2021 | Accelerating CPU-based Distributed DNN Training on Modern HPC Clusters using BlueField-2 DPUsabstractThe Deep Learning (DL) training process consists of multiple phases — data augmentation, training, and validation of the trained model. Traditionally, these phases are executed either on the CPUs or GPUs in a serial fashion due to lack of additional computing resources to offload independent phases of DL training. Recently, Mellanox/NVIDIA has introduced the BlueField-2 DPUs which combine the advanced capabilities of traditional ASIC based network adapters with an array of ARM processors. In this paper, we characterize and explore how one can take advantage of the additional ARM cores on the BlueField-2 DPUs to intelligently accelerate different phases of DL training. We propose multiple novel designs to efficiently offload the phases of DL training to the DPUs. We evaluate our proposed designs using multiple DL models on state-of-the-art HPC clusters. Our experimental results show that the proposed designs are able to deliver up to 15% improvement in overall DL training time. To the best of our knowledge, this is the first work to explore the use of DPUs to accelerate DL training. Arpan Jain, Nawras Alnaasan, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HOTI | 1 |
| 2021 | SUPER: SUb-Graph Parallelism for TransformERsabstractTransformer models have revolutionized the field of Natural Language Processing (NLP) and they achieve state-of-the-art performance in applications like machine translation, question answering, regression, and summarization. However, training Transformers is challenging because of their large memory and compute requirements. The literature contains several approaches to parallelize training, like layer parallelism and pipeline parallelism, but they are optimized to benefit out-of-core models and they don't exploit the inherent parallelism in Transformer models. Other work uses model parallelism to achieve weak scaling by increasing the model size. In this paper, we propose sub-graph parallelism that provides a significant performance improvement over pure data parallelism with a fixed number of resources, and as an additional technique for strong- and weak-scaling without increasing model capacity. Our technique accelerates the training of Transformer models and we generalize the concept to any neural network with multiple branches. We optimize the communication for sub-graph parallelism and combine it with data parallelism to scale performance up to 1024 GPUs. To decrease communication overheads, we propose a topology-aware scheme that limits inter-node communication. Finally, we empirically compare sub-graph parallelism with pure data parallelism and demonstrate its performance benefits in end-to-end training. Arpan Jain, Tim Moon, Tom Benson, Hari Subramoni, Sam Ade Jacobs, Dhabaleswar K. Panda 0001, Brian Van Essen |
IPDPS | 1 |
| 2021 | A 419pW Process-Invariant Temperature Sensor for Ultra-Low Power MicrosystemsabstractThe paper presents a sub-nW BJT based temperature sensor for ultra-low power microsystems. The sensor is based on amplifying the difference between base-emitter voltages of BJTs using gate-leakage transistors. Implemented in UMC 65nm technology, the sensor occupies an area of 0.005mm2. It achieves a maximum non-linearity error of 0.12oC(3σ) over the temperature range of - 55oC to 80oC. Without any trimming, a worst case inaccuracy of +0.36oC/ - 1.61oC is observed w.r.t process variations, depicting the process-invariant nature of the temperature sensor. It also achieves a low supply sensitivity of 0.56oC/V over a wide supply range of 0.7V-3V. The power consumption of the sensor is 419pW at 27oC and 0.7V supply. Abhishek Pullela, Ashfakh Ali, Arpan Jain, Adithya Bathi, Zia Abbas |
ISCAS | 3 |
| 2020 | GEMS: GPU-enabled memory-aware model-parallelism system for distributed DNN trainingabstractData-parallelism has become an established paradigm to train DNNs that fit inside GPU memory on large-scale HPC systems. However, model-parallelism is required to train out-of-core DNNs. In this paper, we deal with emerging requirements brought forward by very large DNNs being trained using high-resolution images common in digital pathology. To address these, we propose, design, and implement GEMS; a GPU-Enabled Memory-Aware Model-Parallelism System. We present several design schemes like GEMS-MAST, GEMS-MASTER, and GEMS-Hybrid that offer excellent speedups over state-of-the-art systems like Mesh-TensorFlow and FlexFlow. Furthermore, we combine model-parallelism and data-parallelism to train a 1000-1ayer ResNet-lk model using 1,024 Volta V100 GPUs with 97.32% scaling-efficiency. For the real-world histopathology whole-slide-image (WSI) of 100,000 x 100,000 pixels, we train custom ResNet-110-v2 on image tiles of size 1024 x 1024 and reduce the training time from seven hours to 28 minutes. Arpan Jain, Ammar Ahmad Awan, Asmaa M. Aljuhani, Jahanzeb Maqbool Hashmi, Quentin Anthony, Hari Subramoni, Dhabaleswar K. Panda 0001, Raghu Machiraju, Anil V. Parwani |
SC | 1 |
| 2019 | Performance Characterization of DNN Training using TensorFlow and PyTorch on Modern ClustersabstractThe recent surge of Deep Learning (DL) models and applications can be attributed to the rise in computational resources, availability of large-scale datasets, and accessible DL frameworks such as TensorFlow and PyTorch. Because these frameworks have been heavily optimized for NVIDIA GPUs, several performance characterization studies exist for GPU-based Deep Neural Network (DNN) training. However, there exist very few research studies that focus on CPU-based DNN training. In this paper, we provide an in-depth performance characterization of state-of-the-art DNNs such as ResNet(s) and Inception-v3/v4 on multiple CPU architectures including Intel Xeon Broadwell, three variants of the Intel Xeon Skylake, AMD EPYC, and NVIDIA GPUs like K80, P100, and V100. We provide three key insights: 1) Multi-process (MP) training should be used even for a single-node, because the single-process (SP) approach cannot fully exploit all the cores, 2) Performance of both SP and MP depend on various features such as the number of cores, the processes per node (ppn), and DNN architecture, and 3) There is a non-linear and complex relationship between CPU/system characteristics (core-count, ppn, hyper-threading, etc) and DNN specifications such as inherent parallelism between layers. We further provide a comparative analysis for CPU and GPU-based training and profiling analysis for Horovod. The fastest Skylake we had access to is up to 2.35× better than a K80 GPU but up to 3.32× slower than a V100 GPU. For ResNet-152 training, we observed that MP is up to 1.47× faster than SP and achieves 125× speedup on 128 Skylake nodes. Arpan Jain, Ammar Ahmad Awan, Quentin Anthony, Hari Subramoni, Dhabaleswar K. Panda 0001 |
CLUSTER | 1 |
| 2019 | Designing a Profiling and Visualization Tool for Scalable and In-depth Analysis of High-Performance GPU ClustersabstractThe recent advent of advanced fabrics like NVIDIA NVLink is enabling the deployment of dense Graphics Processing Unit (GPU) systems, e.g., DGX-2 and Summit. The Message Passing Interface (MPI) has been the dominant programming model to design distributed applications on such clusters. The MPI Tools Interface (MPI_T) provides an opportunity for performance tools and external software to introspect and understand MPI runtime behavior at a deeper level to detect performance and scalability issues. However, the lack of low-overhead and scalable monitoring tools have thus far prevented a comprehensive study of efficiency and utilization of high-performance interconnects such as NVLinks on high-performance GPU-enabled clusters. In this paper, we address this deficiency by proposing and designing an in-depth, real-time analysis, profiling, and visualization tool for high-performance GPU-enabled clusters with NVLinks. The proposed tool builds on the top of the OSU InfiniBand Network Analysis and Monitoring Tool (INAM). It provides insights into the efficiency of different communication patterns by examining the utilization of underlying GPU interconnects. The contributions of the proposed tool are two-fold: 1) domain scientists and system administrators can understand how applications and runtime libraries interact with underlying high-performance interconnects, and 2)Proposed tool enables designers of high-performance communication libraries to gain low-level knowledge to optimize existing designs and develop new algorithms to optimally utilize cutting-edge interconnects on GPU clusters. To the best of our knowledge, this is the first such tool which is capable of presenting a unified and holistic view of MPI-level and fabric level information for emerging NVLink-enabled high-performance GPU clusters. Pouya Kousha, Bharath Ramesh 0005, Kaushik Kandadi Suresh, Ching-Hsiang Chu, Arpan Jain, Nick Sarkauskas, Hari Subramoni, Dhabaleswar K. Panda 0001 |
HiPC | 5 |
| 2019 | A High PSRR, Stable CMOS Current Reference using Process Insensitive TC of Resistance for Wide Temperature ApplicationsabstractIn this paper, a highly stable all CMOS current reference against temperature and supply variation is proposed. Current reference of 5μA and 50nA has been designed for low power and ultra-low power applications respectively. The reference architecture is based on ratio between the PTAT voltage and the PTAT resistance. The process insensitive temperature compensation is accomplished by dividing TC of voltage with process insensitive TC of resistor. A high PSRR, process independent voltage reference is designed for PTAT voltage. N-poly on chip resistor is used for PTAT resistance. The proposed current reference is implemented in 0.18-μm TSMC technology. The architecture achieved PSRR of 74dB and line sensitivity of 0.05% works at supply voltage variation from 1.4V to 3.6V. The current reference of 5μA and 50nA attain temperature coefficient of 11.6 ppm/°C and 12.2 ppm/°C respectively for the temperature variation of -55° C to 125° C. Arpan Jain, Ashfakh Ali, Sai Kiran, Zia Abbas |
ISCAS | 1 |
| 2019 | A 47nW, 0.7-3.6V wide Supply Range, Resistor Based Temperature Sensor for IoT ApplicationsabstractA sub 1-V, ultra low power temperature sensor has been implemented in TSMC 180 nm. The architecture is digital friendly since it creates a pulse width modulated wave instead of voltage. It uses proportional to absolute temperature(PTAT) characteristics of resistance to generate PTAT delay. Temperature to delay conversion depends only on passive elements, thereby making the circuit insensitive to supply variations. Line sensitivity of 0.23 °C/V is achieved for a wide supply range of 0.7-3.6V. A non linearity error of less than 0.8 °C is measured for -55 to 125 °C using linear fit curve. This occupies an area of 0.82 mm2and consumes a power of 47 nW at 0.8 V supply. Ashfakh Ali, Sai Kiran, Arpan Jain, Zia Abbas |
VLSI-SoC | 3 |
| 2019 | A Novel Genetically Optimized Convolutional Neural Network for Traffic Sign Recognition: A New Benchmark on Belgium and Chinese Traffic Sign Datasets
Arpan Jain, Apoorva Mishra, Anupam Shukla, Ritu Tiwari |
Neural Process. Lett. | 1 |