EDBT 2026 Demo / reviewers in the wild / expert
Punarjay Chakravarty
dblp:27/58
· DBLP profile ↗
21ranked-venue papers
6as first author
11since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 11 since 2021Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | RADIANT: Radar-Image Association Network for 3D Object DetectionabstractAs a direct depth sensor, radar holds promise as a tool to improve monocular 3D object detection, which suffers from depth errors, due in part to the depth-scale ambiguity. On the other hand, leveraging radar depths is hampered by difficulties in precisely associating radar returns with 3D estimates from monocular methods, effectively erasing its benefits. This paper proposes a fusion network that addresses this radar-camera association challenge. We train our network to predict the 3D offsets between radar returns and object centers, enabling radar depths to enhance the accuracy of 3D monocular detection. By using parallel radar and camera backbones, our network fuses information at both the feature level and detection level, while at the same time leveraging a state-of-the-art monocular detection technique without retraining it. Experimental results show significant improvement in mean average precision and translation error on the nuScenes dataset over monocular counterparts. Our source code is available at https://github.com/longyunf/radiant. Abhinav Kumar 0004, Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty |
AAAI | 6 |
| 2023 | DisPlacing Objects: Improving Dynamic Vehicle Detection via Visual Place Recognition under Adverse ConditionsabstractCan knowing where you are assist in perceiving objects in your surroundings, especially under adverse weather and lighting conditions? In this work we investigate whether a prior map can be leveraged to aid in the detection of dynamic objects in a scene without the need for a 3D map or pixel-level map-query correspondences. We contribute an algorithm which refines an initial set of candidate object detections and produces a refined subset of highly accurate detections using a prior map. We begin by using visual place recognition (VPR) to retrieve a prior map image for a given query image, then use a binary classification neural network that compares the query and prior map image regions to validate the query detection. Once our classification network is trained, on approximately 1000 query-map image pairs, it is able to improve the performance of vehicle detection when combined with an existing off-the-shelf vehicle detector. We demonstrate our approach using standard datasets across two cities (Oxford and Zurich) under different settings of train-test separation of map-query traverse pairs. We further emphasize the performance gains of our approach against alternative design choices and show that VPR suffices for the task, eliminating the need for precise ground truth localization. Stephen Hausler, Sourav Garg, Punarjay Chakravarty, Shubham Shrivastava, Ankit Vora, Michael Milford |
IROS | 3 |
| 2023 | Locking On: Leveraging Dynamic Vehicle-Imposed Motion Constraints to Improve Visual LocalizationabstractMost 6-DoF localization and SLAM systems use static landmarks but ignore dynamic objects because they cannot be usefully incorporated into a typical pipeline. Where dynamic objects have been incorporated, typical approaches have attempted relatively sophisticated identification and localization of these objects, limiting their robustness or general utility. In this research, we propose a middle ground, demonstrated in the context of autonomous vehicles, using dynamic vehicles to provide limited pose constraint information in a 6-DoF frame-by-frame PnP-RANSAC localization pipeline. We refine initial pose estimates with a motion model and propose a method for calculating the predicted quality of future pose estimates, triggered by whether or not the autonomous vehicle's motion is constrained by the relative frame-to-frame location of dynamic vehicles in the environment. Our approach detects and identifies suitable dynamic vehicles to define these pose constraints to modify a pose filter, resulting in improved recall across a range of localization tolerances from 0.25m to 5m, compared to a state-of-the-art baseline single image PnP method and its vanilla pose filtering. Our constraint detection system is active for approximately 35% of the time on the Ford AV dataset and localization is particularly improved when the constraint detection is active. Stephen Hausler, Sourav Garg, Punarjay Chakravarty, Shubham Shrivastava, Ankit Vora, Michael Milford |
IROS | 3 |
| 2023 | Kinematics Design of a MacPherson Suspension Architecture Based on Bayesian OptimizationabstractEngineering design is traditionally performed by hand: an expert makes design proposals based on past experience, and these proposals are then tested for compliance with certain target specifications. Testing for compliance is performed first by computer simulation using what is called a discipline model. Such a model can be implemented by finite element analysis, multibody systems approach, etc. Designs passing this simulation are then considered for physical prototyping. The overall process may take months and is a significant cost in practice. We have developed a Bayesian optimization (BO) system for partially automating this process by directly optimizing compliance with the target specification with respect to the design parameters. The proposed method is a general framework for computing the generalized inverse of a high-dimensional nonlinear function that does not require, for example, gradient information, which is often unavailable from discipline models. We furthermore develop a three-tier convergence criterion based on: 1) convergence to a solution optimally satisfying all specified design criteria; 2) detection that a design satisfying all criteria is infeasible; or 3) convergence to a probably approximately correct (PAC) solution. We demonstrate the proposed approach on benchmark functions and a vehicle chassis design problem motivated by an industry setting using a state-of-the-art commercial discipline model. We show that the proposed approach is general, scalable, and efficient and that the novel convergence criteria can be implemented straightforwardly based on the existing concepts and subroutines in popular BO software packages. Sinnu Susan Thomas, Jacopo Palandri, Mohsen Lakehal-Ayat, Punarjay Chakravarty, Friedrich Wolf-Monheim, Matthew B. Blaschko |
IEEE Trans. Cybern. | 4 |
| 2022 | Category-Level Pose Retrieval with Contrastive Features Learnt with Occlusion Augmentation
Georgios Kouros, Shubham Shrivastava, Cédric Picron, Sushruth Nagesh, Punarjay Chakravarty, Tinne Tuytelaars |
BMVC | 5 |
| 2022 | Propagating State Uncertainty Through Trajectory ForecastingabstractUncertainty pervades through the modern robotic autonomy stack, with nearly every component (e.g., sensors, detection, classification, tracking, behavior prediction) producing continuous or discrete probabilistic distributions. Trajectory forecasting, in particular, is surrounded by uncertainty as its inputs are produced by (noisy) upstream perception and its outputs are predictions that are often probabilistic for use in downstream planning. However, most trajectory forecasting methods do not account for upstream uncertainty, instead taking only the most-likely values. As a result, perceptual uncer-tainties are not propagated through forecasting and predictions are frequently overconfident. To address this, we present a novel method for incorporating perceptual state uncertainty in trajectory forecasting, a key component of which is a new statistical distance-based loss function which encourages predicting uncertainties that better match upstream perception. We evaluate our approach both in illustrative simulations and on large-scale, real-world data, demonstrating its efficacy in propagating perceptual state uncertainty through prediction and producing more calibrated predictions. Boris Ivanovic, Yifeng Lin, Shubham Shrivastava, Punarjay Chakravarty, Marco Pavone 0001 |
ICRA | 4 |
| 2022 | Localization of a Smart Infrastructure Fisheye Camera in a Prior Map for Autonomous VehiclesabstractThis work presents a technique for localization of a smart infrastructure node, consisting of a fisheye camera, in a prior map. These cameras can detect objects that are outside the line of sight of the autonomous vehicles (AV) and send that information to AVs using V2X technology. However, in order for this information to be of any use to the AV, the detected objects should be provided in the reference frame of the prior map that the AV uses for its own navigation. Therefore, it is important to know the accurate pose of the infrastructure camera with respect to the prior map. Here we propose to solve this localization problem in two steps, (i) we perform feature matching between perspective projection of fisheye image and bird's eye view (BEV) satellite imagery from the prior map to estimate an initial camera pose, (ii) we refine the initialization by maximizing the Mutual Information (MI) between intensity of pixel values of fisheye image and reflectivity of 3D LiDAR points in the map data. We validate our method on simulated data and also present results with real world data. Subodh Mishra, Armin Parchami, Enrique Corona, Punarjay Chakravarty, Ankit Vora, Devarth Parikh, Gaurav Pandey 0004 |
ICRA | 4 |
| 2022 | Real-time Full-stack Traffic Scene Perception for Autonomous Driving with Roadside CamerasabstractWe propose a novel and pragmatic framework for traffic scene perception with roadside cameras. The proposed framework covers a full-stack of roadside perception pipeline for infrastructure-assisted autonomous driving, including object detection, object localization, object tracking, and multi-camera information fusion. Unlike previous vision-based perception frameworks rely upon depth offset or 3D annotation at training, we adopt a modular decoupling design and introduce a landmark-based 3D localization method, where the detection and localization can be well decoupled so that the model can be easily trained based on only 2D annotations. The proposed framework applies to either optical or thermal cameras with pinhole or fish-eye lenses. Our framework is deployed at a two-lane roundabout located at Ellsworth Rd. and State St., Ann Arbor, MI, USA, providing$7\times 24$real-time traffic flow monitoring and high-precision vehicle trajectory extraction. The whole system runs efficiently on a low-power edge computing device with all-component end-to-end delay of less than 20ms. Zhengxia Zou, Rusheng Zhang, Shengyin Shen, Gaurav Pandey 0004, Punarjay Chakravarty, Armin Parchami, Henry X. Liu |
ICRA | 5 |
| 2021 | Radar-Camera Pixel Depth Association for Depth CompletionabstractWhile radar and video data can be readily fused at the detection level, fusing them at the pixel level is potentially more beneficial. This is also more challenging in part due to the sparsity of radar, but also because automotive radar beams are much wider than a typical pixel combined with a large baseline between camera and radar, which results in poor association between radar pixels and color pixel. A consequence is that depth completion methods designed for LiDAR and video fare poorly for radar and video. Here we propose a radar-to-pixel association stage which learns a mapping from radar returns to pixels. This mapping also serves to densify radar returns. Using this as a first stage, followed by a more traditional depth completion method, we are able to achieve image-guided depth completion with radar and video. We demonstrate performance superior to camera and radar alone on the nuScenes dataset. Our source code is available at https://github.com/longyunf/rc-pda. Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty, Praveen Narayanan |
CVPR | 5 |
| 2021 | Full-Velocity Radar Returns by Radar-Camera FusionabstractA distinctive feature of Doppler radar is the measurement of velocity in the radial direction for radar points. However, the missing tangential velocity component hampers object velocity estimation as well as temporal integration of radar sweeps in dynamic scenes. Recognizing that fusing camera with radar provides complementary information to radar, in this paper we present a closed-form solution for the point-wise, full-velocity estimate of Doppler returns using the corresponding optical flow from camera images. Additionally, we address the association problem between radar returns and camera images with a neural network that is trained to estimate radar-camera correspondences. Experimental results on the nuScenes dataset verify the validity of the method and show significant improvements over the state-of-the-art in velocity estimation and accumulation of radar points. Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty, Praveen Narayanan |
ICCV | 5 |
| 2021 | What My Motion tells me about Your Pose: A Self-Supervised Monocular 3D Vehicle DetectorabstractThe estimation of the orientation of an observed vehicle relative to an Autonomous Vehicle (AV) from monocular camera data is an important building block in estimating its 6 DoF pose. Current Deep Learning based solutions for placing a 3D bounding box around this observed vehicle are data hungry and do not generalize well. In this paper, we demonstrate the use of monocular visual odometry for the self-supervised fine-tuning of a model for orientation estimation pre-trained on a reference domain. Specifically, while transitioning from a virtual dataset (vKITTI) to nuScenes, we recover up to 70% of the performance of a fully supervised method. We subsequently demonstrate an optimization-based monocular 3D bounding box detector built on top of the self-supervised vehicle orientation estimator without the requirement of expensive labeled data. This allows 3D vehicle detection algorithms to be self-trained from large amounts of monocular camera data from existing commercial vehicle fleets. Cédric Picron, Punarjay Chakravarty, Tom Roussel, Tinne Tuytelaars |
ICRA | 2 |
| 2020 | Trajectron++: Dynamically-Feasible Trajectory Forecasting with Heterogeneous Data
Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, Marco Pavone 0001 |
ECCV (18) | 3 |
| 2019 | GEN-SLAM: Generative Modeling for Monocular Simultaneous Localization and MappingabstractWe present a Deep Learning based system for the twin tasks of localization and obstacle avoidance essential to any mobile robot. Our system learns from conventional geometric SLAM, and outputs, using a single camera, the topological pose of the camera in an environment, and the depth map of obstacles around it. We use a CNN to localize in a topological map, and a conditional VAE to output depth for a camera image, conditional on this topological location estimation. We demonstrate the effectiveness of our monocular localization and depth estimation system on simulated and real datasets. Punarjay Chakravarty, Praveen Narayanan, Tom Roussel |
ICRA | 1 |
| 2018 | The CAMETRON Lecture Recording System: High Quality Video Recording and Editing with Minimal Human Supervision
Dries Hulens, Bram Aerts, Punarjay Chakravarty, Ali Diba, Toon Goedemé, Tom Roussel, Jeroen Zegers, Tinne Tuytelaars, Luc Van Eycken, Luc Van Gool, Hugo Van hamme, Joost Vennekens |
MMM (1) | 3 |
| 2017 | Expert Gate: Lifelong Learning with a Network of ExpertsabstractIn this paper we introduce a model of lifelong learning, based on a Network of Experts. New tasks / experts are learned and added to the model sequentially, building on what was learned before. To ensure scalability of this process, data from previous tasks cannot be stored and hence is not available when learning a new task. A critical issue in such context, not addressed in the literature so far, relates to the decision which expert to deploy at test time. We introduce a set of gating autoencoders that learn a representation for the task at hand, and, at test time, automatically forward the test sample to the relevant expert. This also brings memory efficiency as only one expert network has to be loaded into memory at any given time. Further, the autoencoders inherently capture the relatedness of one task to another, based on which the most relevant prior model to be used for training a new expert, with fine-tuning or learning-without-forgetting, can be selected. We evaluate our method on image classification and video prediction problems. Rahaf Aljundi, Punarjay Chakravarty, Tinne Tuytelaars |
CVPR | 2 |
| 2017 | CNN-based single image obstacle avoidance on a quadrotorabstractThis paper demonstrates the use of a single forward facing camera for obstacle avoidance on a quadrotor. We train a CNN for estimating depth from a single image. The depth map is then fed to a behaviour arbitration based control algorithm that steers the quadrotor away from obstacles. We conduct experiments with simulated and real drones in a variety of environments. Punarjay Chakravarty, Klaas Kelchtermans, Tom Roussel, Stijn Wellens, Tinne Tuytelaars, Luc Van Eycken |
ICRA | 1 |
| 2016 | Who's that Actor? Automatic Labelling of Actors in TV Series Starting from IMDB Images
Rahaf Aljundi, Punarjay Chakravarty, Tinne Tuytelaars |
ACCV (3) | 2 |
| 2016 | Cross-Modal Supervision for Learning Active Speaker Detection in Video
Punarjay Chakravarty, Tinne Tuytelaars |
ECCV (5) | 1 |
| 2016 | Active speaker detection with audio-visual co-trainingabstractIn this work, we show how to co-train a classifier for active speaker detection using audio-visual data. First, audio Voice Activity Detection (VAD) is used to train a personalized video-based active speaker classifier in a weakly supervised fashion. The video classifier is in turn used to train a voice model for each person. The individual voice models are then used to detect active speakers. There is no manual supervision - audio weakly supervises video classification, and the co-training loop is completed by using the trained video classifier to supervise the training of a personalized audio voice classifier. Punarjay Chakravarty, Jeroen Zegers, Tinne Tuytelaars, Hugo Van hamme |
ICMI | 1 |
| 2015 | Who's Speaking?: Audio-Supervised Classification of Active Speakers in VideoabstractActive speakers have traditionally been identified in video by detecting their moving lips. This paper demonstrates the same using spatio-temporal features that aim to capture other cues: movement of the head, upper body and hands of active speakers. Speaker directional information, obtained using sound source localization from a microphone array is used to supervise the training of these video features. Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, Hugo Van hamme |
ICMI | 1 |
| 2006 | Panoramic Vision and Laser Range Finder Fusion for Multiple Person TrackingabstractThis paper describes a fusion of panoramic vision and laser range data to track multiple persons simultaneously from a stationary robot. Particle filters are used to track people in the plane of the laser and a mixture of Gaussians background subtraction algorithm is used to maintain a colour model for each person being tracked. Colour information is used to recognize lost targets that have reentered the scene Punarjay Chakravarty, Ray A. Jarvis |
IROS | 1 |