VLDB 2026 Research / reviewers in the wild / expert
Athira Nambiar
dblp:32/9767 · also Athira M. Nambiar
· DBLP profile ↗
13ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-4957-5804ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Developing Explainable AI Evaluation Metrics for Image Classification Using Borda Count and Multiple Correlation TechniquesabstractDeep learning models have achieved remarkable performance in various problems, including image classification. However, the complex and “black-box” nature of deep learning models leads to a lack of interpretability of model decisions and a reduced level of trust in deploying the model in the wild. Explainable AI (XAI) has emerged as a new area of research that can understand and interpret the prediction made by the model. In this direction, many XAI techniques have been developed recently, focusing on key aspects that make XAI more reliable for stakeholders: explainability, transparency, and interpretability. Despite the wide array of available explainers, ensuring the quality of their explanations and selecting the most appropriate XAI approach for specific scenarios remains a complex and ambiguous task, due to the heterogeneity of explanations for various models and the lack of ground-truth explanations. Metrics help identify the best-performing explainer for the given problem nonetheless, only minimal research is carried out in this field. Some studies leveraging human-based trials lack objective metrics and exhibit bias, while approaches proposing theoretical guidelines often lack numerical evidence for quantification. To overcome this research gap, we propose two novel collective decision-making metrics leveraging Borda Count (BC) voting rules and Multiple Correlation (MC) statistical techniques. In particular, the BC metric and MC+BC metric compare and rank explanation methods and determine which explanation is most suited for the task. Hence, it serves as a quantitative as well as qualitative assessment benchmarking tool for task-specific explainer assessment. To the best of our knowledge, the application of BC and MC are not yet reported in the XAI literature. In this article, as a pilot case study, we investigate the newly proposed ranking mechanisms for image classification tasks using four popular XAI approaches: PartitionSHAP, GradientSHAP, Gradient-weighted Class Activation Mapping (GradCAM), and Gradient-weighted Class Activation Mapping (GradCAM++). We conduct our investigation on three publicly available large-scale benchmarking image datasets: MNIST, CIFAR-10, and ImageNet. Our robust and promising experimental results highlight the task-specific effectiveness of various XAI approaches and open a new research avenue for effectively comparing explainer outcomes. Anish Samuel Varghese, Somasundaram G., Athira Nambiar |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | Replay to Remember (R2R): An Efficient Uncertainty-Driven Unsupervised Continual Learning Framework Using Generative ReplayabstractContinual Learning entails progressively acquiring knowledge from new data while retaining previously acquired knowledge, thereby mitigating “Catastrophic Forgetting” in neural networks. Our work presents a novel uncertainty-driven Unsupervised Continual Learning framework using Generative Replay, namely “Replay to Remember (R2R)”. The proposed R2R architecture efficiently uses unlabelled and synthetic labelled data in a balanced proportion using a cluster-level uncertainty-driven feedback mechanism and a VLM-powered generative replay module. Unlike traditional memory-buffer methods that depend on pretrained models and pseudo-labels, our R2R framework operates without any prior training. It leverages visual features from unlabeled data and adapts continuously using clustering-based uncertainty estimation coupled with dynamic thresholding. Concurrently, a generative replay mechanism along with DeepSeek-R1 powered CLIP VLM produces labelled synthetic data representative of past experiences, resembling biological visual thinking that replays memory to remember and act in new unseen tasks. Extensive experimental analyses are carried out in CIFAR-10, CIFAR-100, CINIC-10, SVHN and Tiny-ImageNet datasets. Our proposed R2R approach improves knowledge retention, achieving a state-of-the-art performance of 98.13%, 73.06%, 93.41%, 95.18%, 59.74% respectively, surpassing state-of-the-art performance by over 4.36%. Sriram Mandalika, Harsha Vardhan, Athira Nambiar |
ECAI | 3 |
| 2024 | SegXAL: Explainable Active Learning for Semantic Segmentation in Driving Scene Scenarios
Sriram Mandalika, Athira Nambiar |
ICPR (1) | 2 |
| 2024 | S3Simulator: A Benchmarking Side Scan Sonar Simulator Dataset for Underwater Image Analysis
Kamal Basha S, Athira Nambiar |
ICPR (16) | 2 |
| 2023 | Semi-supervised Classification and Segmentation of Forest Fire Using Autoencoders
Akash Koottungal, Shailesh Pandey, Athira Nambiar |
ACIVS | 3 |
| 2023 | Adapt-FuseNet: Context-aware Multimodal Adaptive Fusion of Face and Gait Features using Attention Techniques for Human IdentificationabstractBiometrics plays a significant role in vision-based surveillance applications. Soft biometrics such as gait is widely used with face in surveillance tasks like person recognition and re-identification. Nevertheless, in practical scenarios, classical fusion techniques respond poorly to the changes in individual users, external environment and varying contexts such as viewpoints. To this end, we propose a novel context-aware adaptive multi-biometric fusion strategy viz., ‘Adapt-FuseNet’ for the dynamic incorporation of gait and face biometric cues leveraging attention techniques. In particular, we investigate the impact of attention models such as parallel co-attention & keyless attention, along with various fusion strategies such as naïve fusion & adaptive fusion for human identification. Extensive experiments are carried out on two publically available large gait datasets i.e. CASIA-A and CASIA-B. Results show the superior performance of our proposed context-aware adaptive fusion model compared with the state-of-the-art models. Ashwin Prakash, Thejaswin S, Athira Nambiar, Alexandre Bernardino |
IJCB | 3 |
| 2022 | Co-segmentation inspired attention module for video-based computer vision tasksabstractVideo-based computer vision tasks can benefit from estimation of the salient regions and interactions between those regions. Traditionally, this has been done by identifying the object regions in the images by utilizing pre-trained models to perform object detection, object segmentation and/or object pose estimation. Although using pre-trained models is a viable approach, it has several limitations in the need for an exhaustive annotation of object categories, a possible domain gap between datasets and a bias that is typically present in pre-trained models. In this work, we propose to utilize the common rationale that a sequence of video frames capture a set of common objects and interactions between them, thus a notion of co-segmentation between the video frame features may equip the model with the ability to automatically focus on task-specific salient regions and improve the underlying task’s performance in an end-to-end manner. In this regard, we propose a generic module called “Co-Segmentation inspired Attention Module” (COSAM) that can be plugged in to any CNN model to promote the notion of co-segmentation based attention among a sequence of video frame features. We show the application of COSAM in three video-based tasks namely: (1) Video-based person re-ID, (2) Video captioning, & (3) Video action classification and demonstrate that COSAM is able to capture the task-specific salient regions in video frames, thus leading to notable performance improvements along with interpretable attention maps for a variety of video-based vision tasks, with possible application to other video-based vision tasks as well. Arulkumar Subramaniam, Jayesh Vaidya, Muhammed Abdul Majeed Ameen, Athira Nambiar, Anurag Mittal |
Comput. Vis. Image Underst. | 4 |
| 2021 | Linguistically-aware attention for reducing the semantic gap in vision-language tasks
Gouthaman KV, Athira Nambiar, Sai Srinivas Kancheti, Anurag Mittal |
Pattern Recognit. | 2 |
| 2019 | Co-Segmentation Inspired Attention Networks for Video-Based Person Re-IdentificationabstractPerson re-identification (Re-ID) is an important real-world surveillance problem that entails associating a person's identity over a network of cameras. Video-based Re-ID approaches have gained significant attention recently since a video, and not just an image, is often available. In this work, we propose a novel Co-segmentation inspired video Re-ID deep architecture and formulate a Co-segmentation based Attention Module (COSAM) that activates a common set of salient features across multiple frames of a video via mutual consensus in an unsupervised manner. As opposed to most of the prior work, our approach is able to attend to person accessories along with the person. Our plug-and-play and interpretable COSAM module applied on two deep architectures (ResNet50, SE-ResNet50) outperform the state-of-the-art methods on three benchmark datasets. Arulkumar Subramaniam, Athira Nambiar, Anurag Mittal |
ICCV | 2 |
| 2017 | Context-Aware Person Re-Identification in the Wild Via Fusion of Gait and Anthropometric FeaturesabstractIn this work, we present a context-aware ensemble fusion framework based on soft-biometric features, for long term person re-identification (Re-ID) in wild surveillance scenarios. The characteristics of a person that best correlate to its identity depend strongly on the view point. For instance, a person with a short stride gait is better perceived from a lateral view, whereas a person with a large chest is more distinct from a frontal view. Thus we associate context to the viewing direction of walking people in a surveillance scenario and choose the best features for each case. Using the MS KinectTM sensor v.2, we collect data from walking subjects and extract associated anthropometric and gait features. Each context is analysed with a Feature selection technique (Sequential Forward Selection) so that only the most relevant features for the context are retained. Then, individual context-specific classifiers are trained leveraging those selected features. Finally, we propose a contextaware ensemble fusion strategy, which we term as 'Contextspecific score-level fusion', based on the adaptive weighted sum of the results of individual classifiers. The proposed contextaware Re-ID framework demonstrate significant performance improvement both in terms of speed (up to 4.5 times faster) and accuracy (up to 17% rank-1 Re-ID rate) compared to the Context-unaware systems. From the study, we show that gait features are better for lateral views and anthropometric features are better for frontal views, confirming the results of previous studies. Athira Nambiar, Alexandre Bernardino, Jacinto C. Nascimento, Ana Fred |
FG | 1 |
| 2016 | Person Re-identification in Frontal Gait Sequences via Histogram of Optic Flow Energy Image
Athira Nambiar, Jacinto C. Nascimento, Alexandre Bernardino, José Santos-Victor |
ACIVS | 1 |
| 2015 | Shape Context for soft biometrics in person re-identification and database retrieval
Athira Nambiar, Alexandre Bernardino, Jacinto C. Nascimento |
Pattern Recognit. Lett. | 1 |
| 2012 | Secure image databases through distributed source coding of SIFT descriptorsabstractThe adoption of distributed databases calls for storing data at two or more sites, in order to address application-specific requirements including, e.g., redundancy, data locality, and so on. When visual data (including images, videos and their corresponding descriptors) need to be stored, synchronization across different sites might require significant bandwidth resources. In this paper we explore the use of distributed source coding to encode local SIFT descriptors extracted from static images. The key tenet is to exploit, at the decoder side, the correlation between matching pairs of descriptors extracted, respectively, from the out-of-date and up-to-date image. Preliminary results show that a coding efficiency gain up to 1 bit/descriptor can be achieved in the case of ideal lossless coding. In the case of distributed source coding with LDPC codes, a practical average gain of 0.19 bit/descriptor is observed. Athira Nambiar, Marco Tagliasacchi, Enrico Magli |
MMSP | 1 |