Anbumani Subramanian

dblp:09/2566 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-3426-9967ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 IDD-CRS: A Comprehensive Video Dataset for Critical Road Scenarios in Unstructured Environments
abstract
In this work, we present IDD-CRS, a large-scale dataset focused on critical road scenarios, captured using Advanced Driver Assistance Systems (ADAS) and dash cameras. Unlike existing datasets that predominantly emphasize pedestrian safety and vehicle safety separately, IDD-CRS incorporates both vehicle and pedestrian behaviors, offering a more comprehensive view of road safety. The dataset includes diverse scenarios, such as high-speed lane changes, unsafe vehicle approaches to pedestrians and cyclists, and complex interactions between ego vehicles and other road agents. Leveraging ADAS technology allows us to accurately define the temporal boundaries of actions, resulting in precise annotations and more reliable safety analysis. With 90 hours of video footage, consisting of 5400 one-minute-long videos and 135,000 frames, IDD-CRS introduces new vehicle-related classes and hard negative classes, establishing baselines for action recognition and long-tail action recognition tasks. Our benchmarks reveal the limitations of current models, pointing toward future advancements needed for improving road safety technology.
Ravi Shankar Mishra, Chirag Parikh, Anbumani Subramanian, C. V. Jawahar, Ravi Kiran Sarvadevabhatla
IV3
2024 Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach
abstract
With the increased importance of autonomous navigation systems has come an increasing need to protect the safety of Vulnerable Road Users (VRUs) such as pedestrians. Predicting pedestrian intent is one such challenging task, where prior work predicts the binary cross/no-cross intention with a fusion of visual and motion features. However, there has been no effort so far to hedge such predictions with human-understandable reasons. We address this issue by introducing a novel problem setting of exploring the intuitive reasoning behind a pedestrian’s intent. In particular, we show that predicting the ‘WHY’ can be very useful in understanding the ‘WHAT’. To this end, we propose a novel, reason-enriched PIE++ dataset consisting of multi-label textual explanations/reasons for pedestrian intent. We also introduce a novel multi-task learning framework called MINDREAD, which leverages a cross-modal representation learning framework for predicting pedestrian intent as well as the reason behind the intent. Our comprehensive experiments show significant improvement of 5.6% and 7% in accuracy and F1-score for the task of intent prediction on the PIE++ dataset using MINDREAD. We also achieved a 4.4% improvement in accuracy on a commonly used JAAD dataset. Extensive evaluation using quantitative/qualitative metrics and user studies shows the effectiveness of our approach.
Vaishnavi Khindkar, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
IROS4
2024 Visual Place Recognition in Unstructured Driving Environments
abstract
The problem of determining geolocation through visual inputs, known as Visual Place Recognition (VPR), has attracted significant attention in recent years owing to its potential applications in autonomous self-driving systems. The rising interest in these applications poses unique challenges, particularly the necessity for datasets encompassing unstructured environmental conditions to facilitate the development of robust VPR methods. In this paper, we address the VPR challenges by proposing an Indian driving VPR dataset that caters to the semantic diversity of unstructured driving environments like occlusions due to dynamic environments, variations in traffic density, viewpoint variability, and variability in lighting conditions. In unstructured driving environments, GPS signals are unreliable often affecting the vehicle to accurately determine location. To address this challenge, we develop an interactive image-to-image tagging annotation tool to annotate large datasets with ground truth annotations for VPR training. Evaluation of the state-of-the-art methods on our dataset shows a significant performance drop of up to 15%, defeating a large number of standard VPR datasets. We also provide an exhaustive quantitative and qualitative experimental analysis of frontal-view, multi-view, and sequence-matching methods. We believe that our dataset will open new challenges for the VPR research community to build robust models. Project Page: https://cvit.iiit.ac.in/research/projects/cvit-projects/iddvpr
Utkarsh Rai, Shankar Gangisetty, A. H. Abdul Hafez, Anbumani Subramanian, C. V. Jawahar
IROS4
2024 Enhancing Road Safety: Predictive Modeling of Accident-Prone Zones with ADAS-Equipped Vehicle Fleet Data
abstract
This work presents a novel approach to identifying possible early accident-prone zones in a large city-scale road network using geo-tagged collision alert data from a vehicle fleet. The alert data has been collected for a year from 200 city buses installed with the Advanced Driver Assistance System (ADAS). To the best of our knowledge, no research paper has used ADAS alerts to identify the early accident-prone zones. A nonparametric technique called Kernel Density Estimation (KDE) is employed to model the distribution of alert data across stratified time intervals. A novel recall-based measure is introduced to assess the degree of support provided by our density-based approach for existing, manually determined accident-prone zones (‘blackspots’) provided by civic authorities. This shows that our KDE approach significantly outperforms existing approaches in terms of the recall-based measure. Introducing a novel linear assignment Earth Mover Distance based measure to predict previously unidentified accident-prone zones. The results and findings support the feasibility of utilizing alert data from vehicle fleets to aid civic planners in assessing accident-zone trends and deploying traffic calming measures, thereby improving overall road safety and saving lives.
Ravi Shankar Mishra, Dev Singh Thakur, Anbumani Subramanian, Mukti Advani, S. Velmurugan, Juby Jose, C. V. Jawahar, Ravi Kiran Sarvadevabhatla
IV3
2023 CueCAn: Cue-driven Contextual Attention for Identifying Missing Traffic Signs on Unconstrained Roads
abstract
Unconstrained Asian roads often involve poor infrastructure, affecting overall road safety. Missing traffic signs are a regular part of such roads. Missing or non-existing object detection has been studied for locating missing curbs and estimating reasonable regions for pedestrians on road scene images. Such methods involve analyzing task-specific single object cues. In this paper, we present the first and most challenging video dataset for missing objects, with multiple types of traffic signs for which the cues are visible without the signs in the scenes. We refer to it as the Missing Traffic Signs Video Dataset (MTSVD). MTSVD is challenging compared to the previous works in two aspects i) The traffic signs are generally not present in the vicinity of their cues, ii) The traffic signs' cues are diverse and unique. Also, MTSVD is the first publicly available missing object dataset. To train the models for identifying missing signs, we complement our dataset with 10K traffic sign tracks, with 40% of the traffic signs having cues visible in the scenes. For identifying missing signs, we propose the Cue-driven Contextual Attention units (CueCAn), which we incorporate in our model's encoder. We first train the encoder to classify the presence of traffic sign cues and then train the entire segmentation model end-to-end to localize missing traffic signs. Quantitative and qualitative analysis shows that CueCAn significantly improves the performance of base models.
Varun Gupta 0007, Anbumani Subramanian, C. V. Jawahar, Rohit Saluja
ICRA2
2023 IDD-3D: Indian Driving Dataset for 3D Unstructured Road Scenes
abstract
Autonomous driving and assistance systems rely on annotated data from traffic and road scenarios to model and learn the various object relations in complex real-world scenarios. Preparation and training of deploy-able deep learning architectures require the models to be suited to different traffic scenarios and adapt to different situations. Currently, existing datasets, while large-scale, lack such diversities and are geographically biased towards mainly developed cities. An unstructured and complex driving layout found in several developing countries such as India poses a challenge to these models due to the sheer degree of variations in the object types, densities, and locations. To facilitate better research toward accommodating such scenarios, we build a new dataset, IDD-3D, which consists of multimodal data from multiple cameras and LiDAR sensors with 12k annotated driving LiDAR frames across various traffic scenarios. We discuss the need for this dataset through statistical comparisons with existing datasets and highlight benchmarks on standard 3D object detection and tracking tasks in complex layouts. Code and data available1.
Shubham Dokania, A. H. Abdul Hafez, Anbumani Subramanian, Manmohan Krishna Chandraker, C. V. Jawahar
WACV3
2022 TRoVE: Transforming Road Scene Datasets into Photorealistic Virtual Environments
Shubham Dokania, Anbumani Subramanian, Manmohan Krishna Chandraker, C. V. Jawahar
ECCV (8)2
2022 New Objects on the Road? No Problem, We'll Learn Them Too
abstract
Object detection plays an essential role in providing localization, path planning, and decision making capabilities in autonomous navigation systems. However, existing object detection models are trained and tested on a fixed number of known classes. This setting makes the object detection model difficult to generalize well in real-world road scenarios while encountering an unknown object. We address this problem by introducing our framework that handles the issue of unknown object detection and updates the model when unknown object labels are available. Next, our solution includes three major components that address the inherent problems present in the road scene datasets. The novel components are a) Feature-Mix that improves the unknown object detection by widening the gap between known and unknown classes in latent feature space, b) Focal regression loss handling the problem of improving small object detection and intra-class scale variation, and c) Curriculum learning further enhances the detection of small objects. We use Indian Driving Dataset (IDD) and Berkeley Deep Drive (BDD) dataset for evaluation. Our solution provides state-of-the-art performance on open-world evaluation metrics. We hope this work will create new directions for open-world object detection for road scenes, making it more reliable and robust autonomous systems.
Shyam Nandan Rai, K. J. Joseph, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
IROS7
2022 Multi-Domain Incremental Learning for Semantic Segmentation
abstract
Recent efforts in multi-domain learning for semantic segmentation attempt to learn multiple geographical datasets in a universal, joint model. A simple fine-tuning experiment performed sequentially on three popular road scene segmentation datasets demonstrates that existing segmentation frameworks fail at incrementally learning on a series of visually disparate geographical domains. When learning a new domain, the model catastrophically forgets previously learned knowledge. In this work, we pose the problem of multi-domain incremental learning for semantic segmentation. Given a model trained on a particular geographical domain, the goal is to (i) incrementally learn a new geographical domain, (ii) while retaining performance on the old domain, (iii) given that the previous domain’s dataset is not accessible. We propose a dynamic architecture that assigns universally shared, domain-invariant parameters to capture homogeneous semantic features present in all domains, while dedicated domain-specific parameters learn the statistics of each domain. Our novel optimization strategy helps achieve a good balance between retention of old knowledge (stability) and acquiring new knowledge (plasticity). We demonstrate the effectiveness of our proposed solution on domain incremental settings pertaining to real-world driving scenes from roads of Germany (Cityscapes), the United States (BDD100k), and India (IDD).1
Prachi Garg, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar
WACV5
2022 To miss-attend is to misalign! Residual Self-Attentive Feature Alignment for Adapting Object Detectors
abstract
Advancements in adaptive object detection can lead to tremendous improvements in applications like autonomous navigation, as they alleviate the distributional shifts along the detection pipeline. Prior works adopt adversarial learning to align image features at global and local levels, yet the instance-specific misalignment persists. Also, adaptive object detection remains challenging due to visual diversity in background scenes and intricate combinations of objects. Motivated by structural importance, we aim to attend prominent instance-specific regions, overcoming the feature misalignment issue. We propose a novel resIduaL seLf-attentive featUre alignMEnt (ILLUME) method for adaptive object detection. ILLUME comprises Self-Attention Feature Map (SAFM) module that enhances structural attention to object-related regions and thereby generates domain invariant features. Our approach significantly reduces the domain distance with the improved feature alignment of the instances. Qualitative results demonstrate the ability of ILLUME to attend important object instances required for alignment. Experimental results on several benchmark datasets show that our method outperforms the existing state-of-the-art approaches.
Vaishnavi Khindkar, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, Rohit Saluja, C. V. Jawahar
WACV4
2022 FLUID: Few-Shot Self-Supervised Image Deraining
abstract
Self-supervised methods have shown promising results in denoising and dehazing tasks, where the collection of the paired dataset is challenging and expensive. However, we find that these methods fail to remove the rain streaks when applied for image deraining tasks. The method’s poor performance is due to the explicit assumptions: (i) the distribution of noise or haze is uniform and (ii) the value of a noisy or hazy pixel is independent of its neighbors. The rainy pixels are non-uniformly distributed, and it is not necessarily dependant on its neighboring pixels. Hence, we conclude that the self-supervised method needs to have some prior knowledge about rain distribution to perform the deraining task. To provide this knowledge, we hypothesize a network trained with minimal supervision to estimate the likelihood of rainy pixels. This leads us to our proposed method called FLUID: Few Shot Sel f-Supervised Image Deraining.We perform extensive experiments and comparisons with existing image deraining and few-shot image-to-image translation methods on Rain 100L and DDN-SIRR datasets containing real and synthetic rainy images. In addition, we use the Rainy Cityscapes dataset to show that our method trained in a few-shot setting can improve semantic segmentation and object detection in rainy conditions. Our approach obtains a mIoU gain of 51.20 over the current best-performing deraining method. [Project Page]
Shyam Nandan Rai, Rohit Saluja, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, C. V. Jawahar
WACV5
2020 Spatial Feedback Learning to Improve Semantic Segmentation in Hot Weather
Shyam Nandan Rai, Vineeth N. Balasubramanian, Anbumani Subramanian, C. V. Jawahar
BMVC3
2020 Munich to Dubai: How far is it for Semantic Segmentationƒ
abstract
Cities having hot weather conditions results in geometrical distortion, thereby adversely affecting the performance of semantic segmentation model In this work, we study the problem of semantic segmentation model in adapting to such hot climate cities. This issue can be circumvented by collecting and annotating images in such weather conditions and training segmentation models on those images. But the task of semantically annotating images for every environment is painstaking and expensive. Hence, we propose a framework that improves the performance of semantic segmentation models without explicitly creating an annotated dataset for such adverse weather variations. Our framework consists of two parts, a restoration network to remove the geometrical distortions caused by hot weather and an adaptive segmentation network that is trained on an additional loss to adapt to the statistics of the ground-truth segmentation map. We train our framework on the Cityscapes dataset, which showed a total loU gain of 12.707 over standard segmentation models. We also observe that the segmentation results obtained by our framework gave a significant improvement for small classes such as poles, person, and rider, which are essential and valuable for autonomous navigation based applications.
Shyam Nandan Rai, Vineeth N. Balasubramanian, Anbumani Subramanian, C. V. Jawahar
WACV3
2019 IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments
abstract
While several datasets for autonomous navigation have become available in recent years, they have tended to focus on structured driving environments. This usually corresponds to well-delineated infrastructure such as lanes, a small number of well-defined categories for traffic participants, low variation in object or background appearance and strong adherence to traffic rules. We propose DS, a novel dataset for road scene understanding in unstructured environments where the above assumptions are largely not satisfied. It consists of 10,004 images, finely annotated with 34 classes collected from 182 drive sequences on Indian roads. The label set is expanded in comparison to popular benchmarks such as Cityscapes, to account for new classes. It also reflects label distributions of road scenes significantly different from existing datasets, with most classes displaying greater within-class diversity. Consistent with real driving behaviors, it also identifies new classes such as drivable areas besides the road. We propose a new four-level label hierarchy, which allows varying degrees of complexity and opens up possibilities for new training methods. Our empirical study provides an in-depth analysis of the label characteristics. State-of-the-art methods for semantic segmentation achieve much lower accuracies on our dataset, demonstrating its distinction compared to Cityscapes. Finally, we propose that our dataset is an ideal opportunity for new problems such as domain adaptation, few-shot learning and behavior prediction in road scenes.
Girish Varma, Anbumani Subramanian, Anoop M. Namboodiri, Manmohan Krishna Chandraker, C. V. Jawahar
WACV2
2012 Designing multiuser multimodal gestural interactions for the living room
abstract
Most work in the space of multimodal and gestural interaction has focused on single user productivity tasks. The design of multimodal, freehand gestural interaction for multiuser lean-back scenarios is a relatively nascent area that has come into focus because of the availability of commodity depth cameras. In this paper, we describe our approach to designing multimodal gestural interaction for multiuser photo browsing in the living room, typically a shared experience with friends and family. We believe that our learnings from this process will add value to the efforts of other researchers and designers interested in this design space.
Sriganesh Madhvanath, Ramadevi Vennelakanti, Anbumani Subramanian, Ankit Shekhawat, Amit Ranjan
ICMI3
2012 Pixene: creating memories while sharing photos
abstract
In this paper we describe Pixene, a photo sharing system that focuses on the capture and subsequent visualization and consumption of interactions around shared photos, where the sharing may be with physically co-present friends and family, or online with one's social network. In the former scenario, the interactions may be richly multimodal and involve pointing and spoken comments. Remote interaction is primarily in the form of 'like's and text comments on social networking sites. Pixene thus acts as a common repository for interactions over photos and brings interactions from a co-located and online photo sharing into a single platform. Pixene also provides a rich photo browsing experience that allows users to view not only the photographs but also the interaction history around them, e.g. who saw it, who did they see it with, what they said in association with different regions of interest, comments and 'like's. In this paper, we describe the features and system design of Pixene.
Ramadevi Vennelakanti, Sriganesh Madhvanath, Anbumani Subramanian, Ajith Sowndararajan, Arun David
ICMI3
2010 Dynamic Hand Pose Recognition Using Depth Data
abstract
Hand pose recognition has been a problem of great interest to the Computer Vision and Human Computer Interaction community for many years and the current solutions either require additional accessories at the user end or enormous computation time. These limitations arise mainly due to the high dexterity of human hand and occlusions created in the limited view of the camera. This work utilizes the depth information and a novel algorithm to recognize scale and rotation invariant hand poses dynamically. We have designed a volumetric shape descriptor enfolding the hand to generate a 3D cylindrical histogram and achieved robust pose recognition in real time.
Poonam Suryanarayan, Anbumani Subramanian, Dinesh Mandalapu
ICPR2
2007 Performance Analysis and Validation of a Paracatadioptric Omnistereo System
abstract
In this paper we present a vector-based 3D localization formula for a paracatadioptric omnistereo system. Based on vector representation, the performance of this stereo system is analyzed numerically, including the maximum detectable range and the uncertainty of 3D localization, with respect to the flexible stereo configuration of the system, positions of scene points, as well as errors in correspondence matching and errors in stereo configuration. The results of performance analysis are used to guide the trajectory of an autonomous surface vehicle (ASV), which is equipped with a paracatadioptric omnidirectional camera, in a map building application.
Xiaojin Gong, Anbumani Subramanian, Christopher L. Wyatt, Daniel J. Stilwell
ICCV2
2007 A Two-stage Algorithm for Shoreline Detection
abstract
Shoreline detection plays an important role in vision based navigation for autonomous surface vehicles (ASVs). It is a challenging task because of the diversity in near-bank scenarios. In this paper, we present a two-stage algorithm to find the shoreline by employing multiple features. First, we classify images into two types: reflection-unidentifiable and reflection-identifiable. Based on this classification, images are further analyzed with suitable techniques respectively. In the reflection-unidentifiable case, the surface reflection is subtle and so the water region can be separated from land by an adaptive thresholding method. The points along the edge of the water region are then identified and the shoreline is estimated through a line-fitting technique. In the reflection-identifiable case, we aim to discriminate water regions from the land by means of a two-category region classifier. Images are oversegmented into small regions based on color homogeneity. Then the characteristic features of land-water scenes like symmetry and brightness are extracted and applied to classify a region into land or water categories. Experimental results show the efficacy of our approach and robustness in diverse situations
Xiaojin Gong, Anbumani Subramanian, Christopher L. Wyatt
WACV2
2003 Segmentation and range sensing using a moving-aperture lens
Anbumani Subramanian, Lakshmi R. Iyer, A. Lynn Abbott, Amy E. Bell
Mach. Vis. Appl.1
2001 Segmentation and Range Sensing Using a Moving-Aperture Lens
abstract
This paper is concerned with the use of a novel motorized lens to perform segmentation of image sequences. The lens has the effect of introducing small, repeating movements of the camera center, so that objects appear to translate in the image by an amount that depends on distance from the plane of focus. For a stationary scene, optical flow magnitudes and direction are therefore directly related to three-dimensional object distance from the observer. We describe a segmentation-procedure that exploits these controlled observer movements, and we present experimental results that demonstrate the successful extraction of objects at different depths. Potential applications of this approach include video compression, compositing, and passive range sensing.
Anbumani Subramanian, Lakshmi R. Iyer, A. Lynn Abbott, Amy E. Bell
ICCV1