VLDB 2026 Research / reviewers in the wild / expert
Shankar Gangisetty
dblp:148/8866 · also Shankar Setty, Shankar Shetty
· DBLP profile ↗
17ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-4448-5794ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling What and Why: Enhancing Driver Intention Prediction with MLLMsabstractPredicting a drivers’ intent (e.g., turns, lane changes) is a critical capability for modern Advanced Driver Assistance Systems (ADAS). While recent Multimodal Large Language Models (MLLMs) show promise in general vision-language tasks, we find that zero-shot MLLMs still lag behind domain-specific approaches for Driver Intention Prediction (DIP). To address this, we introduce DriveXplain, a zero-shot framework based on MLLMs that leverages rich visual cues such as optical flow and road semantics to automatically generate both intention maneuver (what) and rich natural language explanations (why). These maneuver–explanation pairs are then distilled into a compact MLLM, which jointly learns to predict intentions and corresponding explanations. We show that incorporating explanations during training leads to substantial gains over models trained solely on labels, as distilling explanations instills reasoning capabilities by enabling the model to understand not only what decisions to make but also why those decisions are made. Comprehensive experiments across structured (Brain4Cars, AIDE) and unstructured (DAAD) datasets demonstrate that our approach achieves state-of-the-art results in DIP task, outperforming zero-shot and domain-specific baselines. We also present ablation studies to evaluate key design choices in our framework. This work sets a direction for more explainable and generalizable intention prediction in autonomous driving systems. Project webpage: https://avijit9.github.io/DriveXplain/ Sainithin Artham, Avijit Dasgupta, Shankar Gangisetty, C. V. Jawahar |
WACV | 3 |
| 2025 | A Dataset for Semantic Segmentation in the Presence of UnknownsabstractBefore deployment in the real-world deep neural networks require thorough evaluation of how they handle both knowns, inputs represented in the training data, and unknowns (anomalies). This is especially important for scene understanding tasks with safety critical applications, such as in autonomous driving. Existing datasets allow evaluation of only knowns or unknowns - but not both, which is required to establish "in the wild" suitability of deep neural network models. To bridge this gap, we propose a novel anomaly segmentation dataset, ISSU, that features a diverse set of anomaly inputs from cluttered real-world environments. The dataset is twice larger than existing anomaly segmentation datasets, and provides a training, validation and test set for controlled in-domain evaluation. The test set consists of a static and temporal part, with the latter comprised of videos. The dataset provides annotations for both closed-set (knowns) and anomalies, enabling closed-set and open-set evaluation. The dataset covers diverse conditions, such as domain and cross-sensor shift, illumination variation and allows ablation of anomaly detection methods with respect to these variations. Evaluation results of current state-of-the-art methods confirm the need for improvements especially in domain-generalization, small and large object segmentation. The code and the dataset are available at https://github.com/vojirt/benchmark_issu. Zakaria Laskar, Tomás Vojír, Matej Grcic, Iaroslav Melekhov, Shankar Gangisetty, Juho Kannala, Jiri Matas, Giorgos Tolias, C. V. Jawahar |
CVPR | 5 |
| 2025 | Towards Safer and Understandable Driver Intention PredictionabstractAutonomous driving (AD) systems are becoming increasingly capable of handling complex tasks, mainly due to recent advances in deep learning and AI. As interactions between autonomous systems and humans increase, the interpretability of decision-making processes in driving systems becomes increasingly crucial for ensuring safe driving operations. Successful human-machine interaction requires understanding the underlying representations of the environment and the driving task, which remains a significant challenge in deep learning-based systems. To address this, we introduce the task of interpretability in maneuver prediction before they occur for driver safety, i.e., driver intent prediction (DIP), which plays a critical role in AD systems. To foster research in interpretable DIP, we curate the eXplainable Driving Action Anticipation Dataset (DAAD-X), a new multimodal, ego-centric video dataset to provide hierarchical, high-level textual explanations as causal reasoning for the driver's decisions. These explanations are derived from both the driver's eye-gaze and the ego-vehicle's perspective. Next, we propose Video Concept Bottleneck Model (VCBM), a framework that generates spatio-temporally coherent explanations inherently, without relying on post-hoc techniques. Finally, through extensive evaluations of the proposed VCBM on the DAAD-X dataset, we demonstrate that transformer-based models exhibit greater interpretability than conventional CNN-based models. Additionally, we introduce a multilabel t-SNE visualization technique to illustrate the disentanglement and causal correlation among multiple explanations. Our data, code and models are available at: https://mukil07.github.io/VCBM.github.io/ Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C. V. Jawahar |
ICCV | 2 |
| 2025 | Pedestrian Intention and Trajectory Prediction in Unstructured Traffic Using IDD-PeDabstractWith the rapid advancements in autonomous driving, accurately predicting pedestrian behavior has become essential for ensuring safety in complex and unpredictable traffic conditions. The growing interest in this challenge highlights the need for comprehensive datasets that capture unstructured environments, enabling the development of more robust prediction models to enhance pedestrian safety and vehicle navigation. In this paper, we introduce an Indian driving pedestrian dataset designed to address the complexities of modeling pedestrian behavior in unstructured environments, such as illumination changes, occlusion of pedestrians, unsignalized scene types and vehicle-pedestrian interactions. The dataset provides high-level and detailed low-level comprehensive annotations focused on pedestrians requiring the ego-vehicle's attention. Evaluation of the state-of-the-art intention prediction methods on our dataset shows a significant performance drop of up to 15 %, while trajectory prediction methods underperform with an increase of up to 1208 MSE, defeating standard pedes-trian datasets. Additionally, we present exhaustive quantitative and qualitative analysis of intention and trajectory baselines. We believe that our dataset will open new challenges for the pedestrian behavior research community to build robust models. Project Page: https://cvit.iiit.ac.in/research/projects/cvit-projects/iddped Ruthvik Bokkasam, Shankar Gangisetty, A. H. Abdul Hafez, C. V. Jawahar |
ICRA | 2 |
| 2024 | Early Anticipation of Driving Maneuvers
Abdul Wasi, Shankar Gangisetty, Shyam Nandan Rai, C. V. Jawahar |
ECCV (70) | 2 |
| 2024 | ICPR 2024 Competition on Rider Intention Prediction
Shankar Gangisetty, Abdul Wasi, Shyam Nandan Rai, C. V. Jawahar, Sajay Raj, Manish Prajapati, Ayesha Choudhary, Aaryadev Chandra, Dev Chandan, Shireen Chand, Suvaditya Mukherjee |
ICPR (34) | 1 |
| 2024 | Visual Place Recognition in Unstructured Driving EnvironmentsabstractThe problem of determining geolocation through visual inputs, known as Visual Place Recognition (VPR), has attracted significant attention in recent years owing to its potential applications in autonomous self-driving systems. The rising interest in these applications poses unique challenges, particularly the necessity for datasets encompassing unstructured environmental conditions to facilitate the development of robust VPR methods. In this paper, we address the VPR challenges by proposing an Indian driving VPR dataset that caters to the semantic diversity of unstructured driving environments like occlusions due to dynamic environments, variations in traffic density, viewpoint variability, and variability in lighting conditions. In unstructured driving environments, GPS signals are unreliable often affecting the vehicle to accurately determine location. To address this challenge, we develop an interactive image-to-image tagging annotation tool to annotate large datasets with ground truth annotations for VPR training. Evaluation of the state-of-the-art methods on our dataset shows a significant performance drop of up to 15%, defeating a large number of standard VPR datasets. We also provide an exhaustive quantitative and qualitative experimental analysis of frontal-view, multi-view, and sequence-matching methods. We believe that our dataset will open new challenges for the VPR research community to build robust models. Project Page: https://cvit.iiit.ac.in/research/projects/cvit-projects/iddvpr Utkarsh Rai, Shankar Gangisetty, A. H. Abdul Hafez, Anbumani Subramanian, C. V. Jawahar |
IROS | 2 |
| 2022 | SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu |
Comput. Graph. | 8 |
| 2022 | FloodNet: Underwater image restoration based on residual dense learning
Shankar Gangisetty, Raghu Raj Rai |
Signal Process. Image Commun. | 1 |
| 2021 | An Ensemble of Transformer and LSTM Approach for Multivariate Time Series Data ClassificationabstractWafer manufacturing is a complex and time taking process. The multivariate time-series data collected from many soft sensors in the process are highly noisy and imbalanced. Thus, wafer classification is a challenging task. To overcome this challenge, we propose an effective ensemble approach with transformer and long short term memory (LSTM) based deep learning techniques for wafer classification. Though deep learning is a promising technique to analyze the data and make effective predictions, but not widely integrated in manufacture industries for soft sensing due to insufficient research. Also the research community has not been exposed to accessing the real and large scale wafer data that is highly noisy and imbalanced. Our proposed approach is an ensemble of four models, namely, multilayer LSTM, multilayer perceptron classifier, transformer, and feed forward neural network. We finally ensemble all of these models using ROC-AUC scores by adjusting the weights based on skewness of the models to obtain effective performance. We perform an exhaustive empirical analysis of the proposed approach and obtain a best ROC score of 0.748 that is significantly better compared to the baseline models. Aryan Narayan, Bodhi Satwa Mishra, P. G. Sunitha Hiremath, Neha Tarannum Pendari, Shankar Gangisetty |
IEEE BigData | 5 |
| 2021 | Stacked LSTM Based Wafer ClassificationabstractThe sensors are used to analyze the quality of wafers in wafer manufacturing industries. The data from sensors is very helpful in framing solutions to predict the pass or fail status of the wafer by classifying high dimensional sensor data using machine learning techniques. In this work, we propose a stacked long short term memory (LSTM) approach i.e., a seq2seq architecture suitable for time-series data. We perform an exhaustive empirical analysis of the proposed model on the Seagate soft sensing dataset. The evaluation metric used is ROC-AUC score. The proposed stacked LSTM approach gave a best ROC-AUC score of 0.7445 on validation data and 0.729 on test data, which is significantly better than the baseline models. Neeta Shinde, Chandana S, Shashank Anand Patil, K. Siri Chandana, Neha Tarannum Pendari, P. G. Sunitha Hiremath, Shankar Gangisetty |
IEEE BigData | 7 |
| 2021 | Look, Read and Ask: Learning to Ask Questions by Reading Text in Images
Soumya Jahagirdar, Shankar Gangisetty, Anand Mishra 0001 |
ICDAR (1) | 2 |
| 2021 | PIG-Net: Inception based deep learning architecture for 3D point cloud segmentation
Sindhu B. Hegde, Shankar Gangisetty |
Comput. Graph. | 2 |
| 2021 | SHREC 2021: Retrieval of cultural heritage objects
Ivan Sipiran, Patrick Lazo, Cristian López 0001, Milagritos Jimenez, Nihar Bagewadi, Benjamin Bustos, Hieu Dao, Shankar Gangisetty, Martin Hanik, Ngoc-Phuong Ho-Thi, Mike Holenderski, Dmitri Jarnikov, Arniel Labrada, Stefan Lengauer, Roxane Licandro, Dinh-Huan Nguyen, Thang-Long Nguyen-Ho, Luis A. Pérez Rey, Bang-Dang Pham, Reinhold Preiner, Tobias Schreck, Quoc-Huy Trinh, Loek Tonnaer, Christoph von Tycowicz, The-Anh Vu-Le |
Comput. Graph. | 8 |
| 2020 | SHREC 2020: 3D point cloud semantic segmentation for street scenesabstractScene understanding of large-scale 3D point clouds of an outer space is still a challenging task. Compared with simulated 3D point clouds, the raw data from LiDAR scanners consist of tremendous points returned from all possible reflective objects and they are usually non-uniformly distributed. Therefore, its cost-effective to develop a solution for learning from raw large-scale 3D point clouds. In this track, we provide large-scale 3D point clouds of street scenes for the semantic segmentation task. The data set consists of 80 samples with 60 for training and 20 for testing. Each sample with over 2 million points represents a street scene and includes a couple of objects. There are five meaningful classes: building, car, ground, pole and vegetation. We aim at localizing and segmenting semantic objects from these large-scale 3D point clouds. Four groups contributed their results with different methods. The results show that learning-based methods are the trend and one of them achieves the best performance on both Overall Accuracy and mean Intersection over Union. Next to the learning-based methods, the combination of hand-crafted detectors are also reliable and rank second among comparison algorithms. Tao Ku, Remco C. Veltkamp, Bas Boom, David Duque-Arias, Santiago Velasco-Forero, Jean-Emmanuel Deschaud, François Goulette, Beatriz Marcotegui, Sebastian Ortega, Agustín Trujillo, José Pablo Suárez, José M. Santana, Cristián Ramírez, Kiran Akadas, Shankar Gangisetty |
Comput. Graph. | 15 |
| 2019 | An Evaluation of Feature Encoding Techniques for Non-Rigid and Rigid 3D Point Cloud Retrieval
Sindhu B. Hegde, Shankar Gangisetty |
BMVC | 2 |
| 2018 | Example-based 3D inpainting of point clouds using metric tensor and Christoffel symbols
Shankar Gangisetty, Uma Mudenagudi |
Mach. Vis. Appl. | 1 |