Abhinav Kumar 0004

dblp:115/6458-4 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-0043-631XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021
YearPublicationVenuePosition
2025 RICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object Detection
abstract
Radar hits reflect from points on both the boundary and internal to object outlines. This results in a complex distribution of radar hits that depends on factors including object category, size and orientation. Current radar-camera fusion methods implicitly account for this with a black-box neural network. In this paper, we explicitly utilize a radar hit distribution model to assist fusion. First, we build a model to predict radar hit distributions conditioned on object properties obtained from a monocular detector. Second, we use the predicted distribution as a kernel to match actual measured radar points in the neighborhood of the monocular detections, generating matching scores at nearby positions. Finally, a fusion stage combines context with the kernel detector to refine the matching scores. Our method achieves the state-of-the-art radar-camera detection performance on nuScenes. Our source code is available at https://github.com/longyunf/riccardo.
Abhinav Kumar 0004, Xiaoming Liu 0002, Daniel D. Morris
CVPR2
2025 CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002
ICCV1
2024 SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects
abstract
Monocular 3D detectors achieve remarkable performance on cars and smaller objects. However, their performance drops on larger objects, leading to fatal accidents. Some attribute the failures to training data scarcity or the receptive field requirements of large objects. In this paper, we highlight this understudied problem of generalization to large objects. We find that modern frontal detectors struggle to generalize to large objects even on nearly balanced datasets. We argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects. To bridge this gap, we comprehensively investigate regression and dice losses, examining their robustness under varying error levels and object sizes. We mathematically prove that the dice loss leads to superior noise-robustness and model convergence for large objects compared to regression losses for a simplified case. Leveraging our theoretical insights, we propose SeaBird (Segmentation in Bird's View) as the first step towards generalizing to large objects. SeaBird effectively integrates BEV segmentation on foreground objects for 3D detection, with the segmentation head trained with the dice loss. SeaBird achieves SoTA results on the KITTI-360 leaderboard and improves existing detectors on the nuScenes leaderboard, particularly for large objects.
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002
CVPR1
2024 SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction
Yuliang Guo, Abhinav Kumar 0004, Cheng Zhao 0002, Ruoyu Wang 0012, Xinyu Huang 0001, Liu Ren 0001
ECCV (69)2
2024 RePLAy: Remove Projective LiDAR Depthmap Artifacts via Exploiting Epipolar Geometry
Girish Chandar Ganesan, Abhinav Kumar 0004, Xiaoming Liu 0002
ECCV (81)3
2023 RADIANT: Radar-Image Association Network for 3D Object Detection
abstract
As a direct depth sensor, radar holds promise as a tool to improve monocular 3D object detection, which suffers from depth errors, due in part to the depth-scale ambiguity. On the other hand, leveraging radar depths is hampered by difficulties in precisely associating radar returns with 3D estimates from monocular methods, effectively erasing its benefits. This paper proposes a fusion network that addresses this radar-camera association challenge. We train our network to predict the 3D offsets between radar returns and object centers, enabling radar depths to enhance the accuracy of 3D monocular detection. By using parallel radar and camera backbones, our network fuses information at both the feature level and detection level, while at the same time leveraging a state-of-the-art monocular detection technique without retraining it. Experimental results show significant improvement in mean average precision and translation error on the nuScenes dataset over monocular counterparts. Our source code is available at https://github.com/longyunf/radiant.
Abhinav Kumar 0004, Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty
AAAI2
2023 Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
abstract
Recognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D, existing benchmarks are small in size and approaches specialize in few object categories and specific domains, e.g. urban driving scenes. Motivated by the success of 2D recognition, we revisit the task of 3D object detection by introducing a large benchmark, called Omni3D. Omni3DRE-purposes and combines existing datasets resulting in 234k images annotated with more than 3 million instances and 98 categories. 3D detection at such scale is challenging due to variations in camera intrinsics and the rich diversity of scene and object types. We propose a model, called Cube R-CNN, designed to generalize across camera and scene types with a unified approach. We show that Cube R-CNN outperforms prior works on the larger Omni3D and existing benchmarks. Finally, we prove that Omni3D is a powerful dataset for 3D object recognition and show that it improves single-dataset performance and can accelerate learning on new smaller datasets via pre-training.11We release the Omni3D benchmark and Cube R-CNN models at https://github.com/facebookresearch/omni3d.
Garrick Brazil, Abhinav Kumar 0004, Julian Straub, Nikhila Ravi, Justin Johnson 0001, Georgia Gkioxari
CVPR2
2023 PrObeD: Proactive Object Detection Wrapper
abstract
Previous research in $2D$ object detection focuses on various tasks, including detecting objects in generic and camouflaged images. These works are regarded as passive works for object detection as they take the input image as is. However, convergence to global minima is not guaranteed to be optimal in neural networks; therefore, we argue that the trained weights in the object detector are not optimal. To rectify this problem, we propose a wrapper based on proactive schemes, PrObeD, which enhances the performance of these object detectors by learning a signal. PrObeD consists of an encoder-decoder architecture, where the encoder network generates an image-dependent signal termed templates to encrypt the input images, and the decoder recovers this template from the encrypted images. We propose that learning the optimum template results in an object detector with an improved detection performance. The template acts as a mask to the input images to highlight semantics useful for the object detector. Finetuning the object detector with these encrypted images enhances the detection performance for both generic and camouflaged. Our experiments on MS-COCO, CAMO, COD$10$K, and NC$4$K datasets show improvement over different detectors after applying PrObeD. Our models/codes are available at https://github.com/vishal3477/Proactive-Object-Detection.
Vishal Asnani, Abhinav Kumar 0004, Suya You, Xiaoming Liu 0002
NeurIPS2
2023 Tame a Wild Camera: In-the-Wild Monocular Camera Calibration
abstract
3D sensing for monocular in-the-wild images, e.g., depth estimation and 3D object detection, has become increasingly important. However, the unknown intrinsic parameter hinders their development and deployment. Previous methods for the monocular camera calibration rely on specific 3D objects or strong geometry prior, such as using a checkerboard or imposing a Manhattan World assumption. This work instead calibrates intrinsic via exploiting the monocular 3D prior. Given an undistorted image as input, our method calibrates the complete 4 Degree-of-Freedom (DoF) intrinsic parameters. First, we show intrinsic is determined by the two well-studied monocular priors: monocular depthmap and surface normal map. However, this solution necessitates a low-bias and low-variance depth estimation. Alternatively, we introduce the incidence field, defined as the incidence rays between points in 3D space and pixels in the 2D imaging plane. We show that: 1) The incidence field is a pixel-wise parametrization of the intrinsic invariant to image cropping and resizing. 2) The incidence field is a learnable monocular 3D prior, determined pixel-wisely by up-to-sacle monocular depthmap and surface normal. With the estimated incidence field, a robust RANSAC algorithm recovers intrinsic. We show the effectiveness of our method through superior performance on synthetic and zero-shot testing datasets. Beyond calibration, we demonstrate downstream applications in image manipulation detection \& restoration, uncalibrated two-view pose estimation, and 3D sensing.
Abhinav Kumar 0004, Masa Hu, Xiaoming Liu 0002
NeurIPS2
2022 DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection
Abhinav Kumar 0004, Garrick Brazil, Enrique Corona, Armin Parchami, Xiaoming Liu 0002
ECCV (9)1
2021 GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object Detection
abstract
Modern 3D object detectors have immensely benefited from the end-to-end learning idea. However, most of them use a post-processing algorithm called Non-Maximal Suppression (NMS) only during inference. While there were attempts to include NMS in the training pipeline for tasks such as 2D object detection, they have been less widely adopted due to a non-mathematical expression of the NMS. In this paper, we present and integrate GrooMeD-NMS – a novel Grouped Mathematically Differentiable NMS for monocular 3D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS. We first formulate NMS as a matrix operation and then group and mask the boxes in an unsupervised manner to obtain a simple closed-form expression of the NMS. GrooMeD-NMS addresses the mismatch between training and inference pipelines and, therefore, forces the network to select the best 3D box in a differentiable manner. As a result, GrooMeD-NMS achieves state-of-the-art monocular 3D object detection results on the KITTI benchmark dataset performing comparably to monocular video-based methods.
Abhinav Kumar 0004, Garrick Brazil, Xiaoming Liu 0002
CVPR1
2021 Scaling Up Exact Neural Network Compression by ReLU Stability
abstract
We can compress a rectifier network while exactly preserving its underlying functionality with respect to a given input domain if some of its neurons are stable. However, current approaches to determine the stability of neurons with Rectified Linear Unit (ReLU) activations require solving or finding a good approximation to multiple discrete optimization problems. In this work, we introduce an algorithm based on solving a single optimization problem to identify all stable neurons. Our approach is on median 183 times faster than the state-of-art method on CIFAR-10, which allows us to explore exact compression on deeper (5 x 100) and wider (2 x 800) networks within minutes. For classifiers trained under an amount of L1 regularization that does not worsen accuracy, we can remove up to 56% of the connections on the CIFAR-10 dataset. The code is available at the following link, https://github.com/yuxwind/ExactCompression .
Thiago Serra, Xin Yu 0003, Abhinav Kumar 0004, Srikumar Ramalingam
NeurIPS3
2020 Lossless Compression of Deep Neural Networks
Thiago Serra, Abhinav Kumar 0004, Srikumar Ramalingam
CPAIOR2
2020 LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood
abstract
Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained using our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmarks is visible, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on mulitple standard face alignment datasets. Our method's estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks.
Abhinav Kumar 0004, Tim K. Marks, Wenxuan Mou, Ye Wang 0001, Michael J. Jones 0001, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu 0002, Chen Feng 0002
CVPR1
2019 VPDS: An AI-Based Automated Vehicle Occupancy and Violation Detection System
abstract
High Occupancy Vehicle/High Occupancy Tolling (HOV/HOT) lanes are operated based on voluntary HOV declarations by drivers. A majority of these declarations are wrong to leverage faster HOV lane speeds illegally. It is a herculean task to manually regulate HOV lanes and identify these violators. Therefore, an automated way of counting the number of people in a car is prudent for fair tolling and for violator detection.In this paper, we propose a Vehicle Passenger Detection System (VPDS) which works by capturing images through Near Infrared (NIR) cameras on the toll lanes and processing them using deep Convolutional Neural Networks (CNN) models. Our system has been deployed in 3 cities over a span of two years and has served roughly 30 million vehicles with an accuracy of 97% which is a remarkable improvement over manual review which is 37% accurate. Our system can generate an accurate report of HOV lane usage which helps policy makers pave the way towards de-congestion.
Abhinav Kumar 0004, Aishwarya Gupta 0001, Bishal Santra, Lalitha K. S., Manasa Kolla, Mayank Gupta 0002, Rishabh Singh
AAAI1
2019 Zero Shot License Plate Re-Identification
abstract
The problem of person, vehicle or license plate reidentification is generally treated as a multi-shot image retrieval problem. The objective of these tasks is to learn a feature representation of query images (called a "signature") and then use these signatures to match against a database of template image signatures with the aid of a distance metric. In this paper, we propose a novel approach for license plate Re-Id inspired by Zero Shot Learning. The core idea is to generate template signatures for retrieval purposes from a multi-hot text encoding of license plates instead of their images. The proposed method maps license plate images and their license plate numbers to a common embedding space using a Symmetric Triplet loss function so that an image can be queried against its text. In effect, our approach makes it possible to identify license plates whose images have never been seen before, using a large text database of license plate numbers. We show that our system is capable of highly accurate and fast re-identification of license plates, and its performance compares favorably to both OCR-based approaches as well as state of the art image-based Re-ID approaches. In addition to the advantages of avoiding manual image labeling and the ease of creating signature databases, the minimal time and storage requirements enable our system to be deployed even on portable devices.
Mayank Gupta 0002, Abhinav Kumar 0004, Sriganesh Madhvanath
WACV2