VLDB 2026 Research / reviewers in the wild / expert
Sathyanarayanan N. Aakur
dblp:205/2845
· DBLP profile ↗
23ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0003-1062-8929ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BAFLE-DCT: Bypassing Adversarial Filters via Frequency-Selective Embedding in the DCT DomainabstractDeep learning-based vision systems are increasingly deployed in high-stakes applications, yet remain vulnerable to imperceptible manipulations that exploit detection limitations in current steganalysis models. We present BAFLE-DCT, a frequency-domain steganography framework that achieves high-capacity, imperceptible data embedding while evading state-of-the-art deep steganalysis detection methods. Unlike traditional spatial-domain methods that alter pixel values and trigger visual or statistical artifacts, BAFLE-DCT operates in the Discrete Cosine Transform (DCT) domain, selectively modifying mid-frequency coefficients in perceptually insignificant regions identified via saliency analysis. A lightweight feedforward network further refines block selection using entropy and DCT variance features to balance embedding capacity and visual fidelity. Stego images generated by BAFLE-DCT consistently bypass advanced steganalysis models including YeNet, SRNet, and Hybrid Deep-Learning Framework, yielding near-random detection rates (~50%) across payload sizes. Importantly, embedded images maintain classification consistency under CLIP, demonstrating semantic preservation. We also release a large-scale, full-color steganographic dataset for frequency-domain research, addressing limitations of grayscale, spatial-domain benchmarks. Our results expose critical vulnerabilities and limitations in visual content authentication pipelines and motivate the development of frequency-aware detection strategies. Thilina Mendis, Farah I. Kandah, Sathyanarayanan N. Aakur |
WACV | 3 |
| 2026 | FSP-DETR: Few-Shot Prototypical Parasitic Ova DetectionabstractObject detection in biomedical settings is fundamentally constrained by the scarcity of labeled data and the frequent emergence of novel or rare categories. We present FSP-DETR, a unified detection framework that enables robust few-shot detection, open-set recognition, and generalization to unseen biomedical tasks within a single model. Built upon a class-agnostic DETR backbone, our approach constructs class prototypes from original support images and learns an embedding space using augmented views and a lightweight transformer decoder. Training jointly optimizes a prototype matching loss, an alignment-based separation loss, and a KL divergence regularization to improve discriminative feature learning and calibration under scarce supervision. Unlike prior work that tackles these tasks in isolation, FSP-DETR enables inference-time flexibility to support unseen class recognition, background rejection, and cross-task adaptation without retraining. We also introduce a new ova species detection benchmark with 20 parasite classes and establish standardized evaluation protocols. Extensive experiments across ova, blood cell, and malaria detection tasks demonstrate that FSP-DETR significantly outperforms prior few-shot and prototype-based detectors, especially in low-shot and open-set scenarios. Shubham Trehan, Udhav Ramachandran, Akash Rao, Ruth Scimeca, Sathyanarayanan N. Aakur |
WACV | 5 |
| 2026 | STaTS: Structure-aware temporal sequence summarization via statistical window merging
Disharee Bhowmick, Ranjith Ramanathan, Sathyanarayanan N. Aakur |
Pattern Recognit. Lett. | 3 |
| 2025 | ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity RecognitionabstractOpen-world egocentric activity recognition poses a fundamental challenge due to its unconstrained nature, requiring models to infer unseen activities from an expansive, partially observed search space. We introduce ProbRes, a Probabilistic Residual search framework based on jump-diffusion that efficiently navigates this space by balancing prior-guided exploration with likelihood-driven exploitation. Our approach integrates structured commonsense priors to construct a semantically coherent search space, adaptively refines predictions using Vision-Language Models (VLMs) and employs a stochastic search mechanism to locate high-likelihood activity labels while minimizing exhaustive enumeration efficiently. We systematically evaluate ProbRes across multiple openness levels (L0-L3), demonstrating its adaptability to increasing search space complexity. In addition to achieving state-of-the-art performance on benchmark datasets (GTEA Gaze, GTEA Gaze+, EPIC-Kitchens, and Charades-Ego), we establish a clear taxonomy for open-world recognition, delineating the challenges and methodological advancements necessary for egocentric activity understanding. Our results highlight the importance of structured search strategies, paving the way for scalable and efficient open-world activity recognition. Sanjoy Kundu, Shanmukha Vellamcheti, Sathyanarayanan N. Aakur |
ICCV | 3 |
| 2025 | CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance GroundingabstractWe introduce CRAFT, a neuro-symbolic framework for interpretable affordance grounding, which identifies the objects in a scene that enable a given action (e.g., “cut”). CRAFT integrates structured commonsense priors from ConceptNet and language models with visual evidence from CLIP, using an energy-based reasoning loop to refine predictions iteratively. This process yields transparent, goal-driven decisions to ground symbolic and perceptual structures. Experiments in multi-object, label-free settings demonstrate that CRAFT enhances accuracy while improving interpretability, providing a step toward robust and trustworthy scene understanding. Joe Lin, Sathyanarayanan N. Aakur |
NeSy | 3 |
| 2024 | Discovering Novel Actions from Open World Egocentric Videos with Object-Grounded Visual Commonsense Reasoning
Sanjoy Kundu, Shubham Trehan, Sathyanarayanan N. Aakur |
ECCV (59) | 3 |
| 2024 | Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
Shubham Trehan, Sathyanarayanan N. Aakur |
ICPR (15) | 2 |
| 2024 | Capturing Temporal Components for Time Series Classification
Venkata Ragavendra Vavilthota, Ranjith Ramanathan, Sathyanarayanan N. Aakur |
ICPR (1) | 3 |
| 2024 | TEPI: Taxonomy-Aware Embedding and Pseudo-Imaging for Scarcely-Labeled Zero-Shot Genome ClassificationabstractA species' genetic code or genome encodes valuable evolutionary, biological, and phylogenetic information that aids in species recognition, taxonomic classification, and understanding genetic predispositions like drug resistance and virulence. However, the vast number of potential species poses significant challenges in developing a general-purpose whole genome classification tool. Traditional bioinformatics tools have made notable progress but lack scalability and are computationally expensive. Machine learning-based frameworks show promise but must address the issue of large classification vocabularies with long-tail distributions. In this study, we propose addressing this problem through zero-shot learning using TEPI,Taxonomy-awareEmbedding andPseudo-Imaging. We represent each genome as pseudo-images and map them to a taxonomy-aware embedding space for reasoning and classification. This embedding space captures compositional and phylogenetic relationships of species, enabling predictions in extensive search spaces. We evaluate TEPI using two rigorous zero-shot settings and demonstrate its generalization capabilities qualitatively on curated, large-scale, publicly sourced data. Sathyanarayanan N. Aakur, Vishalini R. Laguduva, Priyadharsini Ramamurthy, Akhilesh Ramachandran |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | IS-GGT: Iterative Scene Graph Generation with Generative TransformersabstractScene graphs provide a rich, structured representation of a scene by encoding the entities (objects) and their spatial relationships in a graphical format. This representation has proven useful in several tasks, such as question answering, captioning, and even object detection, to name a few. Current approaches take a generation-by-classification approach where the scene graph is generated through labeling of all possible edges between objects in a scene, which adds computational overhead to the approach. This work introduces a generative transformer-based approach to generating scene graphs beyond link prediction. Using two transformer-based components, we first sample a possible scene graph structure from detected objects and their visual features. We then perform predicate classification on the sampled edges to generate the final scene graph. This approach allows us to efficiently generate scene graphs from images with minimal inference overhead. Extensive experiments on the Visual Genome dataset demonstrate the efficiency of the proposed approach. Without bells and whistles, we obtain, on average, 20.7% mean recall (mR@100) across different settings for scene graph generation (SGG), outperforming state-of-the-art SGG approaches while offering competitive performance to unbiased SGG approaches. Sanjoy Kundu, Sathyanarayanan N. Aakur |
CVPR | 2 |
| 2023 | ProtoKD: Learning from Extremely Scarce Data for Parasite Ova RecognitionabstractDeveloping reliable computational frameworks for early parasite detection, particularly at the ova (or egg) stage, is crucial for advancing healthcare and effectively managing potential public health crises. While deep learning has significantly assisted human workers in various tasks, its application in diagnostics has been constrained by the need for extensive datasets. The ability to learn from an extremely scarce training dataset, i.e., when fewer than 5 examples per class are present, is essential for scaling deep learning models in biomedical applications where large-scale data collection and annotation can be expensive or not possible (in case of novel or unknown infectious agents). In this study, we introduce ProtoKD, one of the first approaches to tackle the problem of multi-class parasitic ova recognition using extremely scarce data. Combining the principles of prototypical networks and self-distillation, we can learn robust representations from only one sample per class. Furthermore, we establish a new benchmark to drive research in this critical direction and validate that the proposed ProtoKD framework achieves state-of-the-art performance. Additionally, we evaluate the framework's generalizability to other downstream tasks by assessing its performance on a large-scale taxonomic profiling task based on metagenomes sequenced from real-world clinical data. Shubham Trehan, Udhav Ramachandran, Ruth Scimeca, Sathyanarayanan N. Aakur |
ICMLA | 4 |
| 2023 | Leveraging Symbolic Knowledge Bases for Commonsense Natural Language Inference Using Pattern TheoryabstractThe commonsense natural language inference (CNLI) tasks aim to select the most likely follow-up statement to a contextual description of ordinary, everyday events and facts. Current approaches to transfer learning of CNLI models across tasks require many labeled data from the new task. This paper presents a way to reduce this need for additional annotated training data from the new task by leveraging symbolic knowledge bases, such as ConceptNet. We formulate a teacher-student framework for mixed symbolic-neural reasoning, with the large-scale symbolic knowledge base serving as the teacher and a trained CNLI model as the student. This hybrid distillation process involves two steps. The first step is a symbolic reasoning process. Given a collection of unlabeled data, we use an abductive reasoning framework based on Grenander's pattern theory to create weakly labeled data. Pattern theory is an energy-based graphical probabilistic framework for reasoning among random variables with varying dependency structures. In the second step, the weakly labeled data, along with a fraction of the labeled data, is used to transfer-learn the CNLI model into the new task. The goal is to reduce the fraction of labeled data required. We demonstrate the efficacy of our approach by using three publicly available datasets (OpenBookQA, SWAG, and HellaSWAG) and evaluating three CNLI models (BERT, LSTM, and ESIM) that represent different tasks. We show that, on average, we achieve 63% of the top performance of a fully supervised BERT model with no labeled data. With only 1,000 labeled samples, we can improve this performance to 72%. Interestingly, without training, the teacher mechanism itself has significant inference power. The pattern theory framework achieves 32.7% accuracy on OpenBookQA, outperforming transformer-based models such as GPT (26.6%), GPT-2 (30.2%), and BERT (27.1%) by a significant margin. We demonstrate that the framework can be generalized to successfully train neural CNLI models using knowledge distillation under unsupervised and semi-supervised learning settings. Our results show that it outperforms all unsupervised and weakly supervised baselines and some early supervised approaches, while offering competitive performance with fully supervised baselines. Additionally, we show that the abductive learning framework can be adapted for other downstream tasks, such as unsupervised semantic textual similarity, unsupervised sentiment classification, and zero-shot text classification, without significant modification to the framework. Finally, user studies show that the generated interpretations enhance its explainability by providing key insights into its reasoning mechanism. Sathyanarayanan N. Aakur, Sudeep Sarkar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Actor-Centered Representations for Action Localization in Streaming Videos
Sathyanarayanan N. Aakur, Sudeep Sarkar |
ECCV (38) | 1 |
| 2022 | Bayesian Tracking of Video Graphs Using Joint Kalman Smoothing and Registration
Aditi Basu Bal, Ramy Mounir, Sathyanarayanan N. Aakur, Sudeep Sarkar, Anuj Srivastava |
ECCV (35) | 3 |
| 2022 | ISD-QA: Iterative Distillation of Commonsense Knowledge from General Language Models for Unsupervised Question AnsweringabstractCommonsense question answering has primarily been tackled through supervised transfer learning, where a language model pre-trained on large amounts of data is used as the starting point. While successful, the approach requires large amounts of labeled question-answer pairs, with increasingly larger amounts of data required as the complexity of scenarios or tasks such as commonsense QA increases. In this paper, we hypothesize that large-scale pre-training of language models encodes the necessary commonsense knowledge to answer common questions in context without labeled data. We propose a novel framework called Iterative Self Distillation for QA (ISD-QA), which extracts the "dark knowledge" encoded during largescale pre-training of language models to provide supervision for commonsense question answering. We show that the approach can be used to train common neural QA models for commonsense question answering by distilling knowledge from language models in an unsupervised manner. With no bells and whistles, we achieve an average of 68% of the performance of fully supervised QA models while requiring no labeled training data. Extensive experiments on three public benchmarks (OpenBookQA, HellaSWAG, and CommonsenseQA) show the effectiveness of the proposed approach. Priyadharsini Ramamurthy, Sathyanarayanan N. Aakur |
ICPR | 2 |
| 2022 | Towards Active Vision for Action Localization with Reactive Control and Predictive LearningabstractVisual event perception tasks such as action localization have primarily focused on supervised learning settings under a static observer, i.e., the camera is static and cannot be controlled by an algorithm. They are often restricted by the quality, quantity, and diversity of annotated training data and do not often generalize to out-of-domain samples. In this work, we tackle the problem of active action localization where the goal is to localize an action while controlling the geometric and physical parameters of an active camera to keep the action in the field of view without training data. We formulate an energy-based mechanism that combines predictive learning and reactive control to perform active action localization without rewards, which can be sparse or non-existent in real-world environments. We perform extensive experiments in both simulated and real-world environments on two tasks - active object tracking and active action localization. We demonstrate that the proposed approach can generalize to different tasks and environments in a streaming fashion, without explicit rewards or training. We show that the proposed approach outperforms unsupervised baselines and obtains competitive performance compared to those trained with reinforcement learning. Shubham Trehan, Sathyanarayanan N. Aakur |
WACV | 2 |
| 2022 | Knowledge guided learning: Open world egocentric action recognition with zero supervision
Sathyanarayanan N. Aakur, Sanjoy Kundu, Nikhil Gunti |
Pattern Recognit. Lett. | 1 |
| 2021 | MG-NET: Leveraging Pseudo-imaging for Multi-modal Metagenome Analysis
Sathyanarayanan N. Aakur, Sai Narayanan, Vineela Indla, Arunkumar Bagavathi, Vishalini R. Laguduva, Akhilesh Ramachandran |
MICCAI (5) | 1 |
| 2021 | Sim2Real for Metagenomes: Accelerating Animal Diagnostics with Adversarial Co-training
Vineela Indla, Vennela Indla, Sai Narayanan, Akhilesh Ramachandran, Arunkumar Bagavathi, Vishalini R. Laguduva, Sathyanarayanan N. Aakur |
PAKDD (1) | 7 |
| 2020 | Action Localization Through Continual Predictive Learning
Sathyanarayanan N. Aakur, Sudeep Sarkar |
ECCV (14) | 1 |
| 2020 | GRaDL: A Framework for Animal Genome Sequence Classification with Graph Representations and Deep LearningabstractBovine Respiratory Disease Complex (BRDC) is a complex respiratory disease in cattle with multiple etiologies, including bacterial and viral. It is estimated that mortality, morbidity, therapy, and quarantine resulting from BRDC account for significant losses in the cattle industry. Early detection and management of BRDC are crucial in mitigating economic losses. Current animal disease diagnostics is based on traditional tests such as bacterial culture, serolog, and Polymerase Chain Reaction (PCR) tests. Advancements of data analytics and machine learning are setting trends in several metagenome sequencing applications. In this work, we demonstrate a machine learning approach to identify pathogen signatures present in bovine metagenome sequences using k-mer-based network embedding followed by a deep learning-based classification task. With experiments conducted on two different simulated datasets, we show that networks-based deep learning approaches can detect pathogen signatures with up to 89.7% accuracy. We will make the simulated data used in our experiments available publicly upon request to tackle this important problem with more innovative approaches. Sai Narayanan, Akhilesh Ramachandran, Sathyanarayanan N. Aakur, Arunkumar Bagavathi |
ICMLA | 3 |
| 2019 | A Perceptual Prediction Framework for Self Supervised Event SegmentationabstractTemporal segmentation of long videos is an important problem, that has largely been tackled through supervised learning, often requiring large amounts of annotated training data. In this paper, we tackle the problem of self-supervised temporal segmentation that alleviates the need for any supervision in the form of labels (full supervision) or temporal ordering (weak supervision). We introduce a self-supervised, predictive learning framework that draws inspiration from cognitive psychology to segment long, visually complex videos into constituent events. Learning involves only a single pass through the training data. We also introduce a new adaptive learning paradigm that helps reduce the effect of catastrophic forgetting in recurrent neural networks. Extensive experiments on three publicly available datasets - Breakfast Actions, 50 Salads, and INRIA Instructional Videos datasets show the efficacy of the proposed approach. We show that the proposed approach outperforms weakly-supervised and unsupervised baselines by up to 24% and achieves competitive segmentation results compared to fully supervised baselines with only a single pass through the training data. Finally, we show that the proposed self-supervised learning paradigm learns highly discriminating features to improve action recognition. Sathyanarayanan N. Aakur, Sudeep Sarkar |
CVPR | 1 |
| 2019 | Going Deeper With Semantics: Video Activity Interpretation Using Semantic ContextualizationabstractA deeper understanding of video activities extends beyond recognition of underlying concepts such as actions and objects: constructing deep semantic representations requires reasoning about the semantic relationships among these concepts, often beyond what is directly observed in the data. To this end, we propose an energy minimization framework that leverages large-scale commonsense knowledge bases, such as ConceptNet, to provide contextual cues to establish semantic relationships among entities directly hypothesized from video. We mathematically express this using the language of Grenander's canonical pattern generator theory. We show that the use of prior encoded commonsense knowledge alleviate the need for large annotated training datasets and help tackle imbalance in training through prior knowledge. Using three different publicly available datasets - Charades, Microsoft Visual Description Corpus and Breakfast Actions datasets, we show that the proposed model can generate video interpretations whose quality is better than those reported by state-of-the-art approaches, which have substantial training needs. Through extensive experiments, we show that the use of commonsense knowledge from ConceptNet allows the proposed approach to handling various challenges such as training data imbalance, weak features and complex semantic relationships and visual scenes. Sathyanarayanan N. Aakur, Fillipe D. M. de Souza, Sudeep Sarkar |
WACV | 1 |