EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Pandey 0004
dblp:23/3937-4
· DBLP profile ↗
21ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-4838-802XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 6 since 2021Systems, architecture and hardware · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised Detection of Sun Glare Camera Frames for Safer Autonomous Driving
Sai Jaideep Reddy Mure, Soumya Subhrajita Mohanty, Gaurav Pandey 0004 |
IV | 3 |
| 2026 | Toward Closing the Sim-to-Real Gap for Autonomous Vehicles: A Physics-Guided Learning Approach for LiDAR Intensity Simulation
Vivek Anand, Bharat Lohani, Rakesh Mishra, Vaibhav Kumar, Gaurav Pandey 0004 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Towards Realistic LiDAR Intensity Simulation in Snowy Weather Using Physics-Informed LearningabstractSimulating realistic LiDAR intensity is essential for autonomous driving, particularly under snow conditions, where current methods fail to capture complex LiDAR-to-atmosphere interactions. This paper introduces a CycleGAN framework guided by physics, which incorporates the principles of LiDAR intensity attenuation in snowy weather, significantly narrowing the simulation-to-reality gap. The model was evaluated using an open-source real snow dataset and an open-source simulated dataset, demonstrating its ability to replicate real-world intensity patterns with high accuracy, as indicated by metrics like Structural Similarity Index Measure (SSIM), Kullback-Leibler (KL) Divergence, etc. In the downstream semantic segmentation task, models trained on the enhanced data outperformed those trained on baseline datasets, underscoring the framework's effectiveness in improving LiDAR data realism and robustness in snow-weather autonomous driving scenarios. Vivek Anand, Bharat Lohani, Rakesh Mishra, Gaurav Pandey 0004 |
IV | 4 |
| 2025 | Advancing LiDAR Intensity Simulation Through Learning With Novel Physics-Based ModalitiesabstractLiDAR sensors are integral to autonomous systems, providing a three-dimensional understanding of the surroundings. The intensity of LiDAR returns offers valuable information about the reflected laser signals, which facilitates crucial tasks such as object detection, classification, and segmentation. However, current physics-based LiDAR simulations fail to produce realistic intensity data. This research addresses this issue by using learning-based methods for realistic LiDAR intensity simulation. We propose a hybrid approach that incorporates novel physics-based modalities, specifically incidence angle and material reflectance, into the learning model to generate more realistic intensity data. We test this methodology across two architectures: (i) U-NET (Convolutional Neural Network) and (ii) Pix2Pix (Generative Adversarial Network), using the SemanticKITTI and VoxelScape datasets. The experiments compare the simulated intensity data generated by our method with state-of-the-art approaches through both qualitative and quantitative evaluations, and assessments of its effectiveness in improving the downstream tasks. The results demonstrate consistent improvements after the inclusion of the physics-based modalities. Vivek Anand, Bharat Lohani, Gaurav Pandey 0004, Rakesh Mishra |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Infrastructure Enabled Guided Navigation for Visually ImpairedabstractDespite advancements in navigation-assistive technology, independent outdoor traveling remains challenging for individuals with vision loss due to uncertain information. We present an outdoor navigation assistive system that collaborates with infrastructure to address these limitations. Our system includes an RGB-D inertial sensor, GPS sensor, Jetson Orin NX 16GB, and bone-conduction headphones. It improves localization accuracy through semantic segmentation, depth map enhancement (Depth-Decoder), and a prior map. The proposed semantic segmentation convolutional neural network (CNN) is designed to operate on low-compute devices, achieving competitive performance with 71.2 mean Intersection of Union (mIoU) on the Cityscapes and 74.46 mIoU on the Camvid. It operates at 209 frames per second (fps) with$480\times 848$image resolution on the GTX 1080 desktop and 50 fps on the low-compute Jetson. Additionally, our Depth-Decoder enhances raw depth maps using geometrical loss with semantic segmentation constraints and structure-from-motion. Depth-Decoder demonstrated an average deviation of 0.66m compared to 2.23m from raw depth maps for landmarks within 10m. Notably, the outcomes of both CNNs are inferred from a single RGB image. Two visually impaired individuals assessed the technology. Participant S001 completed 6/8 trials and S002 completed 7/7 trials. Infrastructure collaboration is demonstrated by using existing prior maps and by dynamically adjusting paths based on real-time information from sensing nodes, to avoid collisions with oncoming agents not visible with the wearable camera. This research provides fundamental design principles for future outdoor navigation assistive technologies, offering insights into addressing challenges for people with visual impairment and other vulnerable road users. Hojun Son, Ankit Vora, Gaurav Pandey 0004, James D. Weiland |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera ImagesabstractThis research paper presents an innovative multitask learning framework that allows concurrent depth estimation and semantic segmentation using a single camera. The proposed approach is based on a shared encoder-decoder architecture, which integrates various techniques to improve the accuracy of the depth estimation and semantic segmentation task without compromising computational efficiency. Additionally, the paper incorporates an adversarial training component, employing a Wasserstein GAN framework with a critic network, to refine model’s predictions. The framework is thoroughly evaluated on two datasets - the outdoor Cityscapes dataset and the indoor NYU Depth V2 dataset - and it outperforms existing state-of-the-art methods in both segmentation and depth estimation tasks. We also conducted ablation studies to analyze the contributions of different components, including pre-training strategies, the inclusion of critics, the use of logarithmic depth scaling, and advanced image augmentations, to provide a better understanding of the proposed framework. The accompanying source code is accessible at https://github.com/PardisTaghavi/SwinMTL. Pardis Taghavi, Reza Langari, Gaurav Pandey 0004 |
IROS | 3 |
| 2024 | Toward Physics-Aware Deep Learning Architectures for LiDAR Intensity Simulation
Vivek Anand, Bharat Lohani, Gaurav Pandey 0004, Rakesh Mishra |
SIMULTECH | 3 |
| 2024 | RGB-X Object Detection via Scene-Specific Fusion ModulesabstractMultimodal deep sensor fusion has the potential to enable autonomous vehicles to visually understand their surrounding environments in all weather conditions. However, existing deep sensor fusion methods usually employ convoluted architectures with intermingled multimodal features, requiring large coregistered multimodal datasets for training. In this work, we present an efficient and modular RGB-X fusion network that can leverage and fuse pre-trained single-modal models via scene-specific fusion modules, thereby enabling joint input-adaptive network architectures to be created using small, coregistered multimodal datasets. Our experiments demonstrate the superiority of our method compared to existing works on RGB-thermal and RGB-gated datasets, performing fusion using only a small amount of additional parameters. Our code is available at https://github.com/dsriaditya999/RGBXFusion. Sri Aditya Deevi, Connor Lee, Lu Gan 0006, Sushruth Nagesh, Gaurav Pandey 0004, Soon-Jo Chung |
WACV | 5 |
| 2023 | Stereo Visual Odometry with Deep Learning-Based Point and Line Feature Matching Using an Attention Graph Neural NetworkabstractRobust feature matching forms the backbone for most Visual Simultaneous Localization and Mapping (vSLAM), visual odometry, 3D reconstruction, and Structure from Motion (SfM) algorithms. However, recovering feature matches from texture-poor scenes is a major challenge and still remains an open area of research. In this paper, we present a Stereo Visual Odometry (StereoVO) technique based on point and line features which uses a novel feature-matching mechanism based on an Attention Graph Neural Network that is designed to perform well even under adverse weather conditions such as fog, haze, rain, and snow, and dynamic lighting conditions such as nighttime illumination and glare scenarios. We perform experiments on multiple real and synthetic datasets to validate our method's ability to perform StereoVO under low-visibility weather and lighting conditions through robust point and line matches. The results demonstrate that our method achieves more line feature matches than state-of-the-art line-matching algorithms, which when complemented with point feature matches perform consistently well in adverse weather and dynamic lighting conditions. Shenbagaraj Kannapiran, Nalin Bendapudi, Ming-Yuan Yu, Devarth Parikh, Spring Berman, Ankit Vora, Gaurav Pandey 0004 |
IROS | 7 |
| 2023 | A Hierarchical Vehicle Behavior Prediction Framework With Traffic Signals and Interactive AgentsabstractVehicle behavior prediction in complex urban scenarios with traffic signals and interactive agents is an important yet complicated task for autonomous vehicles (AVs). In this work, a hierarchical vehicle behavior prediction framework is proposed to incorporate the traffic signal information and model the interaction between vehicles. The framework predicts vehicle behaviors in two stages, discrete intention prediction and continuous trajectory prediction. In the discrete intention prediction stage, Bayesian network is adopted to provide a high-level behavior prediction of the principle other vehicle. The discrete prediction results are forwarded to the second stage, where a continuous trajectory is predicted with maximum entropy inverse reinforcement learning and potential game. The framework is designed to be able to capture the difference among human drivers with parameterized driver characteristics. The proposed predictor is validated in two scenarios: the yellow light running scenario and the right-turn scenario. The trajectory prediction average displacement error of the yellow light running scenario is 0.695m for a 3-second prediction interval, and the prediction accuracy of the right-turn vehicle in the right-turn scenario is 0.51m for a 2-second prediction interval. Zhen Yang 0031, Rusheng Zhang, Gaurav Pandey 0004, Neda Masoud, Henry X. Liu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Localization of a Smart Infrastructure Fisheye Camera in a Prior Map for Autonomous VehiclesabstractThis work presents a technique for localization of a smart infrastructure node, consisting of a fisheye camera, in a prior map. These cameras can detect objects that are outside the line of sight of the autonomous vehicles (AV) and send that information to AVs using V2X technology. However, in order for this information to be of any use to the AV, the detected objects should be provided in the reference frame of the prior map that the AV uses for its own navigation. Therefore, it is important to know the accurate pose of the infrastructure camera with respect to the prior map. Here we propose to solve this localization problem in two steps, (i) we perform feature matching between perspective projection of fisheye image and bird's eye view (BEV) satellite imagery from the prior map to estimate an initial camera pose, (ii) we refine the initialization by maximizing the Mutual Information (MI) between intensity of pixel values of fisheye image and reflectivity of 3D LiDAR points in the map data. We validate our method on simulated data and also present results with real world data. Subodh Mishra, Armin Parchami, Enrique Corona, Punarjay Chakravarty, Ankit Vora, Devarth Parikh, Gaurav Pandey 0004 |
ICRA | 7 |
| 2022 | Real-time Full-stack Traffic Scene Perception for Autonomous Driving with Roadside CamerasabstractWe propose a novel and pragmatic framework for traffic scene perception with roadside cameras. The proposed framework covers a full-stack of roadside perception pipeline for infrastructure-assisted autonomous driving, including object detection, object localization, object tracking, and multi-camera information fusion. Unlike previous vision-based perception frameworks rely upon depth offset or 3D annotation at training, we adopt a modular decoupling design and introduce a landmark-based 3D localization method, where the detection and localization can be well decoupled so that the model can be easily trained based on only 2D annotations. The proposed framework applies to either optical or thermal cameras with pinhole or fish-eye lenses. Our framework is deployed at a two-lane roundabout located at Ellsworth Rd. and State St., Ann Arbor, MI, USA, providing$7\times 24$real-time traffic flow monitoring and high-precision vehicle trajectory extraction. The whole system runs efficiently on a low-power edge computing device with all-component end-to-end delay of less than 20ms. Zhengxia Zou, Rusheng Zhang, Shengyin Shen, Gaurav Pandey 0004, Punarjay Chakravarty, Armin Parchami, Henry X. Liu |
ICRA | 4 |
| 2022 | Explainable Machine Learning for Intrusion Detection via Hardware Performance CountersabstractThe exponential proliferation of Malware over the past decade has threatened system security across a plethora of Internet of Things (IoT) devices. Furthermore, the improvements in computer architectures to include speculative branching and out-of-order executions have engendered new opportunities for adversaries to carry out microarchitectural attacks in these devices. Both Malware and microarchitectural attacks are imperative threats to computing systems, as their behaviors range from stealing sensitive data to total system failure. With the cat-and-mouse game between Anti-Virus Software (AVS) and attackers, the frequent bolstering of AVS induces large computational overhead. Consequently, hardware performance counter (HPC)-based detection strategies augmented with machine learning (ML) classifiers have gained popularity as a low overhead solution in identifying these malicious threats. However, ML models are operated as black boxes, which results in decisions that are not human understandable. Clarity of the models’ results facilitates the development of more robust systems. Existing explainable frameworks are only capable of determining each feature’s impact on a prediction which does not provide meaningful interpretable outcomes for HPC-based intrusion detection. In this article, we address this issue by proposing an explainable HPC-based double regression (HPCDR) ML framework. Our proposed technique provides relevant transparency through isolation of the most malevolent transient window of an application, thereby allowing a user to efficiently locate the pernicious instructions within the program. We evaluated HPCDR on five microarchitectural attacks and two Malware. HPCDR was successfully able to identify the most malicious function manifested in each intrusive application. Abraham Peedikayil Kuruvila, Shamik Kundu, Gaurav Pandey 0004, Kanad Basu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | DHSGAN: An End to End Dehazing Network for Fog and Smoke
Ramavtar Malav, Ayoung Kim, Soumya Ranjan Sahoo, Gaurav Pandey 0004 |
ACCV (5) | 4 |
| 2018 | Robust and Fast 3D Scan Alignment Using Mutual InformationabstractThis paper presents a mutual information (MI) based algorithm for the estimation of full 6-degree-of-freedom (DOF) rigid body transformation between two overlapping point clouds. We first divide the scene into a 3D voxel grid and define simple to compute features for each voxel in the scan. The two scans that need to be aligned are considered as a collection of these features and the MI between these voxelized features is maximized to obtain the correct alignment of scans. We have implemented our method with various simple point cloud features (such as number of points in voxel, variance of z-height in voxel) and compared the performance of the proposed method with existing point-to-point and point-to-distribution registration methods. We show that our approach has an efficient and fast parallel implementation on GPU, and evaluate the robustness and speed of the proposed algorithm on two real-world datasets which have variety of dynamic scenes from different environments. Nikhil Mehta 0002, James R. McBride, Gaurav Pandey 0004 |
ICRA | 3 |
| 2018 | Real Time Incremental Foveal Texture Mapping for Autonomous VehiclesabstractWe propose an end-to-end real time framework to generate high resolution graphics grade textured 3D map of urban environment. The generated detailed map finds its application in the precise localization and navigation of autonomous vehicles. It can also serve as a virtual test bed for various vision and planning algorithms as well as a background map in the computer games. In this paper, we focus on two important issues: (i) incrementally generating a map with coherent 3D surface, in real time and (ii) preserving the quality of color texture. To handle the above issues, firstly, we perform a pose-refinement procedure which leverages camera image information, Delaunay triangulation and existing scan matching techniques to produce high resolution 3D map from the sparse input LIDAR scan. This 3D map is then texturized and accumulated by using a novel technique of ray-filtering which handles occlusion and inconsistencies in pose-refinement. Further, inspired by human fovea, we introduce foveal-processing which significantly reduces the computation time and also assists ray-filtering to maintain consistency in color texture and coherency in 3D surface of the output map. Moreover, we also introduce texture error (TE) and mean texture mapping error (MTME), which provides quantitative measure of texturing and overall quality of the textured maps. James R. McBride, Gaurav Pandey 0004 |
IROS | 3 |
| 2017 | Alignment of 3D point clouds with a dominant ground planeabstractThis paper reports on a novel two-step algorithm for the estimation of full 6-degree-of-freedom (DOF) [tx, ty, tz, θx, θy, θz] rigid body transformation between any two overlapping point-clouds that have a dominant ground plane. We first estimate the ground plane (X-Y plane) from the two 3D point-clouds and align them to obtain a good estimate of the distance between the ground planes (i.e. tz) and rotations θxand θyabout the X and Y axis respectively using the Rodrigues rotation formula. The remaining parameters (tx, ty, θz) are then estimated by maximizing the total mutual information (MI) between the 2D feature maps generated from the multi-modal sensor data. Experimental results using scans obtained by a vehicle equipped with a 3D laser scanner and an omnidirectional camera are used to validate the robustness of the proposed algorithm over a wide range of initial conditions. The proposed method provides an efficient framework for multi-modal sensor data fusion and provides a robust solution to the scan alignment problem. Gaurav Pandey 0004, Shashank Giri, James R. McBride |
IROS | 1 |
| 2014 | Toward mutual information based place recognitionabstractThis paper reports on a novel mutual information (MI) based algorithm for robust place recognition. The proposed method provides a principled framework for fusing the complementary information obtained from 3D lidar and camera imagery for recognizing places within an a priori map of a dynamic environment. The visual appearance of the locations in the map can be significantly different due to changing weather, lighting conditions and dynamical objects present in the environment. Various 3D/2D features are extracted from the textured point clouds (scans) and each scan is represented as a collection of these features. For two scans acquired from the same location, the high value of MI between the features present in the scans indicates that the scans are captured from the same location. We use a non-parametric entropy estimator to estimate the true MI from the sparse marginal and joint histograms of the features extracted from the scans. Experimental results using seasonal datasets collected over several years are used to validate the robustness of the proposed algorithm. Gaurav Pandey 0004, James R. McBride, Silvio Savarese, Ryan M. Eustice |
ICRA | 1 |
| 2012 | Automatic Targetless Extrinsic Calibration of a 3D Lidar and Camera by Maximizing Mutual InformationabstractThis paper reports on a mutual information (MI) based algorithm for automatic extrinsic calibration of a 3D laser scanner and optical camera system. By using MI as the registration criterion, our method is able to work in situ without the need for any specific calibration targets, which makes it practical for in-field calibration. The calibration parameters are estimated by maximizing the mutual information obtained between the sensor-measured surface intensities. We calculate the Cramer-Rao-Lower-Bound (CRLB) and show that the sample variance of the estimated parameters empirically approaches the CRLB for a sufficient number of views. Furthermore, we compare the calibration results to independent ground-truth and observe that the mean error also empirically approaches to zero as the number of views are increased. This indicates that the proposed algorithm, in the limiting case, calculates a minimum variance unbiased (MVUB) estimate of the calibration parameters. Experimental results are presented for data collected by a vehicle mounted with a 3D laser scanner and an omnidirectional camera system. Gaurav Pandey 0004, James R. McBride, Silvio Savarese, Ryan M. Eustice |
AAAI | 1 |
| 2012 | Toward mutual information based automatic registration of 3D point cloudsabstractThis paper reports a novel mutual information (MI) based algorithm for automatic registration of unstructured 3D point clouds comprised of co-registered 3D lidar and camera imagery. The proposed method provides a robust and principled framework for fusing the complementary information obtained from these two different sensing modalities. High-dimensional features are extracted from a training set of textured point clouds (scans) and hierarchical k-means clustering is used to quantize these features into a set of codewords. Using this codebook, any new scan can be represented as a collection of codewords. Under the correct rigid-body transformation aligning two overlapping scans, the MI between the codewords present in the scans is maximized. We apply a James-Stein-type shrinkage estimator to estimate the true MI from the marginal and joint histograms of the codewords extracted from the scans. Experimental results using scans obtained by a vehicle equipped with a 3D laser scanner and an omnidirectional camera are used to validate the robustness of the proposed algorithm over a wide range of initial conditions. We also show that the proposed method works well with 3D data alone. Gaurav Pandey 0004, James R. McBride, Silvio Savarese, Ryan M. Eustice |
IROS | 1 |
| 2011 | Visually bootstrapped generalized ICPabstractThis paper reports a novel algorithm for boot strapping the automatic registration of unstructured 3D point clouds collected using co-registered 3D lidar and omnidirectional camera imagery. Here, we exploit the co-registration of the 3D point cloud with the available camera imagery to associate high dimensional feature descriptors such as scale invariant feature transform (SIFT) or speeded up robust features (SURF) to the 3D points. We first establish putative point correspondence in the high dimensional feature space and then use these correspondences in a random sample consensus (RANSAC) framework to obtain an initial rigid body transformation that aligns the two scans. This initial transformation is then refined in a generalized iterative closest point (ICP) framework. The proposed method is completely data driven and does not require any initial guess on the transformation. We present results from a real world dataset collected by a vehicle equipped with a 3D laser scanner and an omnidirectional camera. Gaurav Pandey 0004, Silvio Savarese, James R. McBride, Ryan M. Eustice |
ICRA | 1 |