Om Ji Omer

dblp:20/4027 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-9149-5605ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Ace-of-Spades: Accelerating Spatially Sparse Convolution for 3D Scene Understanding
abstract
Semantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU/GPU implementations. The existing CNN accelerators designed for weight/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present Ace-of-Spade s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator Ss p nna with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The Ss p nna core together with the 64 KB of L1 memory requires 0.31 mm 2 of area in 10nm process at 1 GHz. Overall, AoS achieves speedup/energy savings of 19.9x / 49.9x and 2.2x / 7.1x over the state-of-the-art CPU and GPU implementations respectively.
Om Ji Omer, Prashant Laddha, Gurpreet S. Kalsi, Kamlesh R. Pillai, Anirud Thyagharajan, Ahimanyu Kulkarni, Anbang Yao, Yurong Chen 0001, Sreenivas Subramoney
ACM Trans. Embed. Comput. Syst.1
2022 Segment-Fusion: Hierarchical Context Fusion for Robust 3D Semantic Segmentation
abstract
3D semantic segmentation is a fundamental building block for several scene understanding applications such as autonomous driving, robotics and AR/VR. Several state-of-the-art semantic segmentation models suffer from the part-misclassification problem, wherein parts of the same object are labelled incorrectly. Previous methods have utilized hierarchical, iterative methods to fuse semantic and instance information, but they lack learnability in context fusion, and are computationally complex and heuristic driven. This paper presents Segment-Fusion, a novel attention-based method for hierarchical fusion of semantic and instance in-formation to address the part misclassifications. The presented method includes a graph segmentation algorithmfor grouping points into segments that pools point-wise features into segment-wise features, a learnable attention-based net-work to fuse these segments based on their semantic and instance features, and followed by a simple yet effective connected component labelling algorithm to convert seg-ment features to instance labels. Segment-Fusion can be flexibly employed with any network architecture for semantic/instance segmentation. It improves the qualitative and quantitative performance of several semantic segmentation backbones by upto 5% on the ScanNet and S3DIS datasets.
Anirud Thyagharajan, Benjamin Ummenhofer, Prashant Laddha, Om Ji Omer, Sreenivas Subramoney
CVPR4
2022 A Unified Programmable Edge Matrix Processor for Deep Neural Networks and Matrix Algebra
abstract
Matrix Algebra and Deep Neural Networks represent foundational classes of computational algorithms across multiple emerging applications like Augmented Reality or Virtual Reality, autonomous navigation (cars, drones, robots), data science, and various artificial intelligence-driven solutions. An accelerator-based architecture can provide performance and energy efficiency supporting fixed functions through customized data paths. However, constrained Edge systems requiring multiple applications and diverse matrix operations to be efficiently supported, cannot afford numerous custom accelerators. In this article, we present MxCore, a unified architecture that comprises tightly coupled vector and programmable cores sharing data through highly optimized interconnects along with a configurable hardware scheduler managing the co-execution. We submit MxCore as the generalized approach to facilitate the flexible acceleration of multiple Matrix Algebra and Deep-learning applications across a range of sparsity levels. Unified compute resources improve overall resource utilization and performance per unit area. Aggressive and novel microarchitecture techniques along with block-level sparsity support optimize compute and data-reuse to minimize bandwidth and power requirements enabling ultra-low latency applications for low-power and cost-sensitive Edge deployments. MxCore requires a small silicon footprint of 0.2068 mm 2 , in a modern 7-nm process at 1 GHz and achieves (0.15 FP32 and 0.62 INT8) TMAC/mm 2 , dissipating only 11.66 μW of leakage power. At iso-technology and iso-frequency, MxCore provides an energy efficiency of 651.4×, 159.9×, 104.8×, and 124.2× as compared to the 128-core Nvidia’s Maxwell GPU for dense General Matrix Multiply, sparse Deep Neural Network, Cholesky decomposition, and triangular matrix solve respectively.
Biji George, Om Ji Omer, Ziaul Choudhury, Anoop V, Sreenivas Subramoney
ACM Trans. Embed. Comput. Syst.2
2021 REDUCT: Keep it Close, Keep it Cool! : Efficient Scaling of DNN Inference on Multi-core CPUs with Near-Cache Compute
abstract
Deep Neural Networks (DNN) are used in a variety of applications and services. With the evolving nature of DNNs, the race to build optimal hardware (both in datacenter and edge) continues. General purpose multi-core CPUs offer unique attractive advantages for DNN inference at both datacenter [60] and edge [71]. Most of the CPU pipeline design complexity is targeted towards optimizing general-purpose single thread performance, and is overkill for relatively simpler, but still hugely important, data parallel DNN inference workloads. Addressing this disparity efficiently can enable both raw performance scaling and overall performance/Watt improvements for multi-core CPU DNN inference.We present REDUCT, where we build innovative solutions that bypass traditional CPU resources which impact DNN inference power and limit its performance. Fundamentally, REDUCT’s "Keep it close" policy enables consecutive pieces of work to be executed close to each other. REDUCT enables instruction delivery/decode close to execution and instruction execution close to data. Simple ISA extensions encode the fixed-iteration count loop-y workload behavior enabling an effective bypass of many power-hungry front-end stages of the wide Out-of-Order (OoO) CPU pipeline. Per core performance scales efficiently by distributing light-weight tensor compute near all caches in a multi-level cache hierarchy. This maximizes the cumulative utilization of the existing architectural bandwidth resources in the system and minimizes movement of data.Across a number of DNN models, REDUCT achieves a 2.3× increase in convolution performance/Watt with a 2× to 3.94× scaling in raw performance. Similarly, REDUCT achieves a 1.8× increase in inner-product performance/Watt with 2.8× scaling in performance. REDUCT performance/power scaling is achieved with no increase to cache capacity or bandwidth and a mere 2.63% increase in area. Crucially, REDUCT operates entirely within the CPU programming and memory model, simplifying software development, while achieving performance similar to or better than state-of-the-art Domain Specific Accelerators (DSA) for DNN inference, providing fresh design choices in the AI era.
Anant Nori, Rahul Bera, Shankar Balachandran, Joydeep Rakshit, Om Ji Omer, Avishaii Abuhatzera, Belliappa Kuttanna, Sreenivas Subramoney
ISCA5
2020 Descriptor Scoring for Feature Selection in Real-Time Visual Slam
abstract
Many emerging applications of Visual SLAM running on resource constrained hardware platforms impose very aggressive pose accuracy requirements and highly demanding latency constraints. To achieve the required pose accuracy under constrained compute budget, real-time SLAM implementations have to work with few but highly repeatable and invariant features. While many state-of-the-art techniques, proposed for selecting good features to track, do address some of these concerns, they are computationally complex and therefore, not suitable for power, latency and cost sensitive edge devices. On the other hand, simpler feature selection methods based on detector (corner) score, lack in identifying features with required invariance and trackability. We present a notion of feature descriptor score as a measure of invariance under distortions. We further propose feature selection method based on descriptor score requiring very minimal compute and demonstrate its performance with binary descriptors on an EKF based visual inertial odometry (VIO). Compared to detector score based methods, our method provides an improvement up to 10% in ATE (Absolute Trajectory Error) score on EuroC dataset.
Prashant Laddha, Om Ji Omer, Gurpreet S. Kalsi, Dipan Mandal, Sreenivas Subramoney
ICIP2
2020 Towards Noise Resilient SLAM
abstract
Sparse-indirect SLAM systems have been dominantly popular due to their computational efficiency and photometric invariance properties. Depth sensors are critical to SLAM frameworks for providing scale information to the 3D world, yet known to be plagued by a wide variety of noise sources, possessing lateral and axial components. In this work, we demonstrate the detrimental impact of these depth noise components on the performance of the state-of-the-art sparse-indirect SLAM system (ORB-SLAM2). We propose (i) Map-Point Consensus based Outlier Rejection (MC-OR) to counter lateral noise, and (ii) Adaptive Virtual Camera (AVC) to combat axial noise accurately. MC-OR utilizes consensus information between multiple sightings of the same landmark to disambiguate noisy depth and filter it out before pose optimization. In AVC, we introduce an error vector as an accurate representation of the axial depth error. We additionally propose an adaptive algorithm to find the virtual camera location for projecting the error used in the objective function of the pose optimization. Our techniques work equally well for stereo image pairs and RGB-D input directly used by sparse-indirect SLAM systems. Our methods were tested on the TUM (RGB-D) and EuRoC (stereo) datasets and we show that they outperform existing state-of-the-art ORB-SLAM2 by 2-3x, especially in sequences critically affected by depth noise.
Anirud Thyagharajan, Om Ji Omer, Dipan Mandal, Sreenivas Subramoney
ICRA2
2020 Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network Acceleration
abstract
This paper presents a Look-Up Table (LUT) based Processing-In-Memory (PIM) technique with the potential for running Neural Network inference tasks. We implement a bitline computing free technique to avoid frequent bitline accesses to the cache sub-arrays and thereby considerably reducing the memory access energy overhead. LUT in conjunction with the compute engines enables sub-array level parallelism while executing complex operations through data lookup which otherwise requires multiple cycles. Sub-array level parallelism and systolic input data flow ensure data movement to be confined to the SRAM slice.Our proposed LUT based PIM methodology exploits substantial parallelism using look-up tables, which does not alter the memory structure/organization, that is, preserving the bit-cell and peripherals of the existing SRAM monolithic arrays. Our solution achieves 1.72× higher performance and 3.14x lower energy as compared to a state-of-the-art processing-in-cache solution. Sub-array level design modifications to incorporate LUT along with the compute engines will increase the overall cache area by 5.6%. We achieve 3.97x speedup w.r.t neural network systolic accelerator with a similar area. The re-configurable nature of the compute engines enables various neural network operations and thereby supporting sequential networks (RNNs) and transformer models. Our quantitative analysis demonstrates 101×, 3× faster execution and 91×, 11× energy efficient than CPU and GPU respectively while running the transformer model, BERT-Base.
Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Rangachar Srinivasa, Makesh Chandran, Kamlesh R. Pillai, Om Ji Omer, Narayanan Vijaykrishnan, Sreenivas Subramoney
MICRO6
2019 Visual Inertial Odometry At the Edge: A Hardware-Software Co-design Approach for Ultra-low Latency and Power
abstract
Visual Inertial Odometry (VIO) is used for estimating pose and trajectory of a system and is a foundational requirement in many emerging applications like AR/VR, autonomous navigation in cars, drones and robots. In this paper, we analyze key compute bottlenecks in VIO and present a highly optimized VIO accelerator based on a hardware-software codesign approach. We detail a set of novel micro-architectural techniques that optimize compute, data movement, bandwidth and dynamic power to make it possible to deliver high quality of VIO at ultra-low latency and power required for budget constrained edge devices. By offloading the computation of the critical linear algebra algorithms from the CPU, the accelerator enables high sample rate IMU usage in VIO processing while acceleration of image processing pipe increases precision, robustness and reduces IMU induced drift in final pose estimate. The proposed accelerator requires a small silicon footprint (1.3 mm2in a 28nm process at 600 MHz), utilizes a modest on-chip shared SRAM (560KB) and achieves 10x speedup over a software-only implementation in terms of image sample-based pose update latency while consuming just 2.2 mW power. In a FPGA implementation, using the EuRoC VIO dataset (VGA 30fps images and 100Hz IMU) the accelerator design achieves pose estimation accuracy (loop closure error) comparable to a software based VIO implementation.
Dipan Mandal, Srivatsava Jandhyala, Om Ji Omer, Gurpreet S. Kalsi, Biji George, Gopi Neela, Santhosh Kumar Rethinagiri, Sreenivas Subramoney, Lance Hacking, Jim Radford, Eagle Jones, Belliappa Kuttanna, Hong Wang 0003
DATE3
2004 Motion estimation from motion smear -a system identification approach
Om Ji Omer, Rajeev Bajpai, K. S. Venkatesh, Sumana Gupta
ICIP1