Abhishek Balasubramaniam

dblp:311/5135 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-8541-1364ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MAPLE: Modality-Aware Projection-free LiDARCamera Fusion for 3D Vehicular Object Detection
abstract
Accurate 3D object detection (3D-OD) is critical for autonomous vehicles, yet embedded platforms impose strict latency, power, and memory constraints. While LiDAR–camera fusion improves robustness, existing approaches depend on precise calibration and computationally expensive view projections. We present MAPLE, a projection-free and calibration-resilient fusion framework that adaptively balances LiDAR geometry and camera semantics using Gated Confidence Fusion (GCF) and low-rank adapter (LoRA) enhanced global attention refinement. MAPLE preserves fine-grained cross-modal interactions without view lifting and injects long-range context at low cost. On the nuScenes benchmark, MAPLE improves mean Average Precision (mAP) by up to 1.6% over the strongest prior fusion baseline, while reducing inference latency by 42.6% and energy consumption by 47% on the NVIDIA Jetson Orin Nano, demonstrating suitability for real-time embedded autonomous perception.
Abhishek Balasubramaniam, Sudeep Pasricha
DATE1
2025 UPAQ: A Framework for Real-Time and Energy-Efficient 3D Object Detection in Autonomous Vehicles
abstract
To enhance perception in autonomous vehicles (AVs), recent efforts are concentrating on 3D object detectors, which deliver more comprehensive predictions than traditional 2D object detectors, at the cost of increased memory footprint and computational resource usage. We present a novel framework called UPAQ, which leverages semi-structured pattern pruning and quantization to improve the efficiency of LiDAR point-cloud and camera-based 3D object detectors on resource-constrained embedded AV platforms. Experimental results on the Jetson Orin Nano embedded platform indicate that UPAQ achieves up to 5.62× and 5.13× model compression rates, up to 1.97× and 1.86× boost in inference speed, and up to 2.07× and 1.87× reduction in energy consumption compared to state-of-the-art model compression frameworks, on the Pointpillar and SMOKE models respectively.
Abhishek Balasubramaniam, Febin Sunny, Sudeep Pasricha
DATE1
2024 OPIMA: Optical Processing-in-Memory for Convolutional Neural Network Acceleration
abstract
Recent advances in machine learning (ML) have spotlighted the pressing need for computing architectures that bridge the gap between memory bandwidth and processing power. The advent of deep neural networks has pushed traditional Von Neumann architectures to their limits due to the high latency and energy consumption costs associated with data movement between the processor and memory for these workloads. One of the solutions to overcome this bottleneck is to perform computation within the main memory through processing-in-memory (PIM), thereby limiting data movement and the costs associated with it. However, dynamic random-access memory-based PIM struggles to achieve high throughput and energy efficiency due to internal data movement bottlenecks and the need for frequent refresh operations. In this work, we introduce OPIMA, a PIM-based ML accelerator, architected within an optical main memory. OPIMA has been designed to leverage the inherent massive parallelism within main memory while performing high-speed, low-energy optical computation to accelerate ML models based on convolutional neural networks. We present a comprehensive analysis of OPIMA to guide design choices and operational mechanisms. In addition, we evaluate the performance and energy consumption of OPIMA, comparing it with conventional electronic computing systems and emerging photonic PIM architectures. The experimental results show that OPIMA can achieve$2.98\times $higher throughput and$137\times $better energy efficiency than the best known prior work.
Febin Sunny, Amin Shafiee, Abhishek Balasubramaniam, Mahdi Nikdast, Sudeep Pasricha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 R-TOSS: A Framework for Real-Time Object Detection using Semi-Structured Pruning
abstract
Object detectors used in autonomous vehicles can have high memory and computational overheads. In this paper, we introduce a novel semi-structured pruning framework called R-TOSS that overcomes the shortcomings of state-of-the-art model pruning techniques. Experimental results on the JetsonTX2 platform show that R-TOSS has a compression rate of 4.4× on the YOLOv5 object detector with a 2.15× speedup in inference time and 57.01% decrease in energy usage. R-TOSS also enables 2.89× compression on RetinaNet with a 1.86× speedup in inference time and 56.31% decrease in energy usage. We also demonstrate significant improvements compared to various state-of-the-art pruning techniques.
Abhishek Balasubramaniam, Febin Sunny, Sudeep Pasricha
DAC1