Prathosh A. P.

dblp:218/5887 · also A. P. Prathosh, Aragulla Prasad Prathosh, Prathosh AP · DBLP profile ↗
← Back
48ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0002-8699-5760ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021
YearPublicationVenuePosition
2026 BiMol-Diff: A Unified Diffusion Framework for Molecular Generation and Captioning
abstract
Aditya Hemant Shahane, Anuj Kumar Sirohi, Devansh Arora, Nitin Kumar, Prathosh AP, Sandeep Kumar. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Aditya Hemant Shahane, Anuj Kumar Sirohi, Devansh Arora, Prathosh A. P., Sandeep Kumar 0005
ACL (1)5
2026 MANTA: Physics-Informed Generalized Underwater Object Tracking
abstract
Underwater object tracking is challenging due to wavelength-dependent attenuation and scattering, which severely distort appearance across depths and water conditions. Existing trackers trained on terrestrial data fail to generalize to these physics-driven degradations. We present MANTA, a physics-informed framework integrating representation learning with tracking design for underwater scenarios. We propose a dual-positive contrastive learning strategy coupling temporal consistency with Beer–Lambert augmentations to yield features robust to both temporal and underwater distortions. We further introduce a multi-stage pipeline augmenting motion-based tracking with a physics-informed secondary association algorithm that integrates geometric consistency and appearance similarity for re-identification under occlusion and drift. To complement standard IoU metrics, we propose Center–Scale Consistency (CSC) and Geometric Alignment Score (GAS) to assess geometric fidelity. Experiments on four underwater benchmarks (WebUOT-1M, UOT32, UTB180, UWCOT220) show that MANTA achieves state-of-the-art performance, improving Success AUC by up to 6%, while ensuring stable long-term generalized underwater tracking and efficient runtime. Code available at https://github.com/Kazedaa/MANTA.
Suhas Srinath, Hemang Jamadagni, Aditya Chandrasekar, Prathosh A. P.
WACV4
2025 Partially Blinded Unlearning: Class Unlearning for Deep Networks from Bayesian Perspective
abstract
To follow regulations on individual data privacy and safety, machine learning models must systematically remove information learned from specific subsets of a user's training data that can no longer be utilized. To address this problem, machine unlearning has emerged as an important area of research, that helps remove information learned from specific subsets of training data from a pre-trained model without needing to retrain the whole model from scratch. The principal aim of this study is to formulate a methodology aimed for the purposeful elimination of information linked to a specific class of data from a pre-trained classification network. This intentional removal decreases the model's performance specifically concerning the unlearned data class while simultaneously minimizing any detrimental impacts on the model's performance in other classes. To achieve this goal, we frame the class unlearning problem from a Bayesian perspective, which yields a loss function that minimizes the log-likelihood associated with the unlearned data with a stability regularization in parameter space. This stability regularization incorporates Mohalanobis distance with respect to the Fisher Information matrix and L2 distance from the pre-trained model parameters. Our novel approach, termed Partially-Blinded Unlearning (PBU), surpasses existing state-of-the-art class unlearning methods, demonstrating superior effectiveness. Notably, PBU achieves this efficacy without requiring information about the entire training dataset but only of the unlearned data points, marking a distinctive feature of its performance.
Subhodip Panda, Shashwat Sourav, Prathosh A. P.
AAAI3
2025 Latent Mamba Operator for Partial Differential Equations
abstract
Neural operators have emerged as powerful data-driven frameworks for solving Partial Differential Equations (PDEs), offering significant speedups over numerical methods. However, existing neural operators struggle with scalability in high-dimensional spaces, incur high computational costs, and face challenges in capturing continuous and long-range dependencies in PDE dynamics. To address these limitations, we introduce the Latent Mamba Operator (LaMO), which integrates the efficiency of state-space models (SSMs) in latent space with the expressive power of kernel integral formulations in neural operators. We also establish a theoretical connection between state-space models (SSMs) and the kernel integral of neural operators. Extensive experiments across diverse PDE benchmarks on regular grids, structured meshes, and point clouds covering solid and fluid physics datasets, LaMOs achieve consistent state-of-the-art (SOTA) performance, with a 32.3% improvement over existing baselines in solution operator approximation, highlighting its efficacy in modeling complex PDEs solution.
Karn Tiwari, Niladri Dutta, N. M. Anoop Krishnan, Prathosh A. P.
ICML4
2025 LangDAug: Langevin Data Augmentation for Multi-Source Domain Generalization in Medical Image Segmentation
abstract
Medical image segmentation models often struggle to generalize across different domains due to various reasons. Domain Generalization (DG) methods overcome this either through representation learning or data augmentation (DA). While representation learning methods seek domain-invariant features, they often rely on ad-hoc techniques and lack formal guarantees. DA methods, which enrich model representations through synthetic samples, have shown comparable or superior performance to representation learning approaches. We propose LangDAug, a novel Langevin Data Augmentation for multi-source domain generalization in 2D medical image segmentation. LangDAug leverages Energy-Based Models (EBMs) trained via contrastive divergence to traverse between source domains, generating intermediate samples through Langevin dynamics. Theoretical analysis shows that LangDAug induces a regularization effect, and for GLMs, it upper-bounds the Rademacher complexity by the intrinsic dimensionality of the data manifold. Through extensive experiments on Fundus segmentation and 2D MRI prostate segmentation benchmarks, we show that LangDAug outperforms state-of-the-art domain generalization methods and effectively complements existing domain-randomization approaches. The codebase for our method is available at https://github.com/backpropagator/LangDAug.
Piyush Tiwary, Kinjawl Bhattacharyya, Prathosh A. P.
ICML3
2025 Concept Forgetting via Label Annealing
abstract
The effectiveness of current machine learning models relies on their ability to grasp diverse concepts present in datasets. However, biased and noisy data can inadvertently cause these models to learn certain undesired concepts, undermining their ability to generalize and provide utility. Consequently, modifying a trained model to forget these concepts becomes imperative for their responsible deployment. We refer to this problem as *concept forgetting*. Our goal is to develop techniques for forgetting specific undesired concepts from a pre-trained classification model’s prediction. To achieve this goal, we present an algorithm called **L**abel **AN**nealing (**LAN**). This iterative algorithm employs a two-stage method for each iteration. In the first stage, pseudo-labels are assigned to all the samples by annealing or redistributing the original labels based on the predictions of the model in the current iteration. During the second stage, the model is fine-tuned on this pseudo-labeled dataset generated from the first stage. We illustrate the effectiveness of the proposed algorithms across various models and datasets. Our method reduces *concept violation*, a metric that measures how much the model forgets specific concepts, by about 85.35% on the MNIST dataset, 73.25% on the CIFAR-10 dataset, and 69.46% on the CelebA dataset while maintaining high model accuracy.
Subhodip Panda, Ananda Theertha Suresh, Atri Guha, Prathosh A. P.
UAI4
2025 UnDIVE: Generalized Underwater Video Enhancement Using Generative Priors
abstract
With the rise of marine exploration, underwater imaging has gained significant attention as a research topic. Under-water video enhancement has become crucial for real-time computer vision tasks in marine exploration. However, most existing methods focus on enhancing individual frames and neglect video temporal dynamics, leading to visually poor enhancements. Furthermore, the lack of ground-truth references limits the use of abundant available underwater video data in many applications. To address these issues, we propose a two-stage framework for enhancing underwater videos. The first stage uses a denoising diffusion probabilistic model to learn a generative prior from unlabeled data, capturing robust and descriptive feature representations. In the second stage, this prior is incorporated into a physics-based image formulation for spatial enhancement, while also enforcing temporal consistency between video frames. Our method enables real-time and computationally-efficient processing of high-resolution underwater videos at lower resolutions, and offers efficient enhancement in the presence of diverse water-types. Extensive experiments on four datasets show that our approach generalizes well and outperforms existing enhancement methods. Our code is available at github. com/suhas-srinath/undive.
Suhas Srinath, Aditya Chandrasekar, Hemang Jamadagni, Rajiv Soundararajan, Prathosh A. P.
WACV5
2025 Video-based real-time assessment and diagnosis of autism spectrum disorder using deep neural networks
abstract
Human action recognition (HAR) in untrimmed videos can make insightful predictions of human behaviour. Previous work on HAR-included models trained on spatial and temporal annotations and could classify limited actions from trimmed videos. These methods reported limitations such as (1) performance degradation due to the lack of precision temporal regions proposal and (2) poor adaptability of the models in the clinical domain because of unrelated actions of interest. We propose an innovative method that could analyse untrimmed behavioural videos to recommend actions of interest leading to diagnostic and functional assessments for children with Autism Spectrum Disorder (ASD). Our method entails end-to-end behaviour action recognition (BAR) pipeline, including child detection, temporal action localization, and actions of interest identification and classification. The model trained on the data of 400 ASD children and 125 with other developmental delays (ODD) accurately identified ASD, ODD, and Neurotypical children with 79.7, 77.2, and 80.8 accuracy, respectively. The model's performance on an independent benchmark Self-Stimulatory Behaviour Dataset (SSBD) reported top-1 accuracy of 78.57 for combined localization with action recognition, significantly higher than the earlier reported outcomes. © 2023 John Wiley & Sons Ltd.
Varun Ganjigunte Prakash, Manu Kohli, Prathosh A. P., Monica Juneja, Manushree Gupta, Smitha Sairam, Sadasivan Sitaraman, Anjali Sanjeev Bangalore, John Vijay Sagar Kommu, Lokesh Saini, Prashant Ramesh Utage, Nishant Goyal
Expert Syst. J. Knowl. Eng.3
2025 GenSelfDiff-HIS: Generative Self-Supervision Using Diffusion for Histopathological Image Segmentation
abstract
Histopathological image segmentation is a laborious and time-intensive task, often requiring analysis from experienced pathologists for accurate examinations. To reduce this burden, supervised machine-learning approaches have been adopted using large-scale annotated datasets for histopathological image analysis. However, in several scenarios, the availability of large-scale annotated data is a bottleneck while training such models. Self-supervised learning (SSL) is an alternative paradigm that provides some respite by constructing models utilizing only the unannotated data which is often abundant. The basic idea of SSL is to train a network to perform one or many pseudo or pretext tasks on unannotated data and use it subsequently as the basis for a variety of downstream tasks. It is seen that the success of SSL depends critically on the considered pretext task. While there have been many efforts in designing pretext tasks for classification problems, there have not been many attempts on SSL for histopathological image segmentation. Motivated by this, we propose an SSL approach for segmenting histopathological images via generative diffusion models. Our method is based on the observation that diffusion models effectively solve an image-to-image translation task akin to a segmentation task. Hence, we propose generative diffusion as the pretext task for histopathological image segmentation. We also utilize a multi-loss function-based fine-tuning for the downstream task. We validate our method using several metrics on two publicly available datasets along with a newly proposed head and neck (HN) cancer dataset containing Hematoxylin and Eosin (H&E) stained images along with annotations.
Vishnuvardhan Purma, Suhas Srinath, Seshan Srirangarajan, Aanchal Kakkar, Prathosh A. P.
IEEE Trans. Medical Imaging5
2024 Fusing Conditional Submodular GAN and Programmatic Weak Supervision
abstract
Programmatic Weak Supervision (PWS) and generative models serve as crucial tools that enable researchers to maximize the utility of existing datasets without resorting to laborious data gathering and manual annotation processes. PWS uses various weak supervision techniques to estimate the underlying class labels of data, while generative models primarily concentrate on sampling from the underlying distribution of the given dataset. Although these methods have the potential to complement each other, they have mostly been studied independently. Recently, WSGAN proposed a mechanism to fuse these two models. Their approach utilizes the discrete latent factors of InfoGAN for the training of the label models and leverages the class-dependent information of the label models to generate images of specific classes. However, the disentangled latent factor learned by the InfoGAN may not necessarily be class specific and hence could potentially affect the label model's accuracy. Moreover, the prediction of the label model is often noisy in nature and can have a detrimental impact on the quality of images generated by GAN. In our work, we address these challenges by (i) implementing a noise-aware classifier using the pseudo labels generated by the label model, (ii) utilizing the prediction of the noise-aware classifier for training the label model as well as generation of class-conditioned images. Additionally, We also investigate the effect of training the classifier with a subset of the dataset within a defined uncertainty budget on pseudo labels. We accomplish this by formalizing the subset selection problem as submodular maximization with a knapsack constraint on the entropy of pseudo labels. We conduct experiments on multiple datasets and demonstrate the efficacy of our methods on several tasks vis-a-vis the current state-of-the-art methods. Our implementation is available at https://github.com/kyrs/subpws-gan
Kumar Shubham, Pranav Sastry, Prathosh A. P.
AAAI3
2024 HOLMES: Hyper-Relational Knowledge Graphs for Multi-hop Question Answering using LLMs
abstract
Given unstructured text, Large Language Models (LLMs) are adept at answering simple (single-hop) questions.However, as the complexity of the questions increase, the performance of LLMs degrade.We believe this is due to the overhead associated with understanding the complex question followed by filtering and aggregating unstructured information in the raw text.Recent methods try to reduce this burden by integrating structured knowledge triples into the raw text, aiming to provide a structured overview that simplifies information processing.However, this simplistic approach is query-agnostic and the extracted facts are ambiguous as they lack context.To address these drawbacks and to enable LLMs to answer complex (multi-hop) questions with ease, we propose to use a knowledge graph (KG) that is context-aware and is distilled to contain query-relevant information.The use of our compressed distilled KG as input to the LLM results in our method utilizing up to 67% fewer tokens to represent the query relevant information present in the supporting documents, compared to the state-of-the-art (SoTA) method.Our experiments show consistent improvements over the SoTA across several metrics (EM, F1, BERTScore, and Human Eval) on two popular benchmark datasets (HotpotQA and MuSiQue).
Pranoy Panda, Ankush Agarwal, Chaitanya Devaguptapu, Manohar Kaul, Prathosh A. P.
ACL (1)5
2024 A Unified Framework for Discovering Discrete Symmetries
abstract
We consider the problem of learning a function respecting a symmetry from among a class of symmetries. We develop a unified framework that enables symmetry discovery across a broad range of subgroups including locally symmetric, dihedral and cyclic subgroups. At the core of the framework is a novel architecture composed of linear, matrix-valued and non-linear functions that expresses functions invariant to these subgroups in a principled manner. The structure of the architecture enables us to leverage multi-armed bandit algorithms and gradient descent to efficiently optimize over the linear and the non-linear functions, respectively, and to infer the symmetry that is ultimately learnt. We also discuss the necessity of the matrix-valued functions in the architecture. Experiments on image-digit sum and polynomial regression tasks demonstrate the effectiveness of our approach.
Pavan Karjol, Rohan Kashyap, Aditya Gopalan, Prathosh A. P.
AISTATS4
2024 WISER: Weak Supervision and Supervised Representation Learning to Improve Drug Response Prediction in Cancer
abstract
Cancer, a leading cause of death globally, occurs due to genomic changes and manifests heterogeneously across patients. To advance research on personalized treatment strategies, the effectiveness of various drugs on cells derived from cancers (’cell lines’) is experimentally determined in laboratory settings. Nevertheless, variations in the distribution of genomic data and drug responses between cell lines and humans arise due to biological and environmental differences. Moreover, while genomic profiles of many cancer patients are readily available, the scarcity of corresponding drug response data limits the ability to train machine learning models that can predict drug response in patients effectively. Recent cancer drug response prediction methods have largely followed the paradigm of unsupervised domain-invariant representation learning followed by a downstream drug response classification step. Introducing supervision in both stages is challenging due to heterogeneous patient response to drugs and limited drug response data. This paper addresses these challenges through a novel representation learning method in the first phase and weak supervision in the second. Experimental results on real patient data demonstrate the efficacy of our method WISER (Weak supervISion and supErvised Representation learning) over state-of-the-art alternatives on predicting personalized drug response. Our implementation is available at https://github.com/kyrs/WISER
Kumar Shubham, Aishwarya Jayagopal, Syed Mohammed Danish, Prathosh A. P., Vaibhav Rajan
ICML4
2024 LoMOE: Localized Multi-Object Editing via Multi-Diffusion
Goirik Chakrabarty, Aditya Chandrasekar, Ramya Hebbalaguppe, Prathosh A. P.
ACM Multimedia4
2024 Bayesian Pseudo-Coresets via Contrastive Divergence
abstract
Bayesian methods provide an elegant framework for estimating parameter posteriors and quantification of uncertainty associated with probabilistic models. However, they often suffer from slow inference times. To address this challenge, Bayesian Pseudo-Coresets (BPC) have emerged as a promising solution. BPC methods aim to create a small synthetic dataset, known as pseudo-coresets, that approximates the posterior inference achieved with the original dataset. This approximation is achieved by optimizing a divergence measure between the true posterior and the pseudo-coreset posterior. Various divergence measures have been proposed for constructing pseudo-coresets, with forward Kullback-Leibler (KL) divergence being the most successful. However, using forward KL divergence necessitates sampling from the pseudo-coreset posterior, often accomplished through approximate Gaussian variational distributions. Alternatively, one could employ Markov Chain Monte Carlo (MCMC) methods for sampling, but this becomes challenging in high-dimensional parameter spaces due to slow mixing. In this study, we introduce a novel approach for constructing pseudo-coresets by utilizing contrastive divergence. Importantly, optimizing contrastive divergence eliminates the need for approximations in the pseudo-coreset construction process. Furthermore, it enables the use of finite-step MCMC methods, alleviating the requirement for extensive mixing to reach a stationary distribution. To validate our method’s effectiveness, we conduct extensive experiments on multiple datasets, demonstrating its superiority over existing BPC techniques. Our implementation is available at https://github.com/backpropagator/BPC-CD .
Piyush Tiwary, Kumar Shubham, Vivek Kashyap, Prathosh A. P.
UAI4
2024 Cycle consistent twin energy-based models for image-to-image translation
Piyush Tiwary, Kinjawl Bhattacharyya, Prathosh A. P.
Medical Image Anal.3
2024 SoLAD: Sampling Over Latent Adapter for Few Shot Generation
Arnab Kumar Mondal, Piyush Tiwary, Parag Singla, Prathosh A. P.
IEEE Signal Process. Lett.4
2023 Adaptive Mixing of Auxiliary Losses in Supervised Learning
abstract
In many supervised learning scenarios, auxiliary losses are used in order to introduce additional information or constraints into the supervised learning objective. For instance, knowledge distillation aims to mimic outputs of a powerful teacher model; similarly, in rule-based approaches, weak labeling information is provided by labeling functions which may be noisy rule-based approximations to true labels. We tackle the problem of learning to combine these losses in a principled manner. Our proposal, AMAL, uses a bi-level optimization criterion on validation data to learn optimal mixing weights, at an instance-level, over the training data. We describe a meta-learning approach towards solving this bi-level objective, and show how it can be applied to different scenarios in supervised learning. Experiments in a number of knowledge distillation and rule denoising domains show that AMAL provides noticeable gains over competitive baselines in those domains. We empirically analyze our method and share insights into the mechanisms through which it provides performance gains. The code for AMAL is at: https://github.com/durgas16/AMAL.git.
Durga Sivasubramanian, Ayush Maheshwari, Prathosh A. P., Pradeep Shenoy, Ganesh Ramakrishnan
AAAI3
2023 Neural Discovery of Permutation Subgroups
abstract
We consider the problem of discovering subgroup $H$ of permutation group $S_n$. Unlike the traditional $H$-invariant networks wherein $H$ is assumed to be known, we present a method to discover the underlying subgroup, given that it satisfies certain conditions. Our results show that one could discover any subgroup of type $S_k (k \leq n)$ by learning an $S_n$-invariant function and a linear transformation. We also prove similar results for cyclic and dihedral subgroups. Finally, we provide a general theorem that can be extended to discover other subgroups of $S_n$. We also demonstrate the applicability of our results through numerical experiments on image-digit sum and symmetric polynomial regression tasks.
Pavan Karjol, Rohan Kashyap, Prathosh A. P.
AISTATS3
2023 Minority Oversampling for Imbalanced Data via Class-Preserving Regularized Auto-Encoders
abstract
Class imbalance is a common phenomenon in multiple application domains such as healthcare, where the sample occurrence of one or few class categories is more prevalent in the dataset than the rest. This work addresses the class-imbalance issue by proposing an over-sampling method for the minority classes in the latent space of a Regularized Auto-Encoder (RAE). Specifically, we construct a latent space by maximizing the conditional data likelihood using an Encoder-Decoder structure, such that oversampling through convex combinations of latent samples preserves the class identity. A jointly-trained linear classifier that separates convexly coupled latent vectors from different classes is used to impose this property on the AE’s latent space. Further, the aforesaid linear classifier is used for final classification without retraining. We theoretically show that our method can achieve a low variance risk estimate compared to naive oversampling methods and is robust to overfitting. We conduct several experiments on benchmark datasets and show that our method outperforms the existing oversampling techniques for handling class imbalance. The code of the proposed method is available at: https://github.com/arnabkmondal/oversamplingrae.
Arnab Kumar Mondal, Lakshya Singhal, Piyush Tiwary, Parag Singla, Prathosh A. P.
AISTATS5
2023 DeGPR: Deep Guided Posterior Regularization for Multi-Class Cell Detection and Counting
abstract
Multi-class cell detection and counting is an essential task for many pathological diagnoses. Manual counting is tedious and often leads to inter-observer variations among pathologists. While there exist multiple, general-purpose, deep learning-based object detection and counting methods, they may not readily transfer to detecting and counting cells in medical images, due to the limited data, presence of tiny overlapping objects, multiple cell types, severe class-imbalance, minute differences in size/shape of cells, etc. In response, we propose guided posterior regularization (DEGPR), which assists an object detector by guiding it to exploit discriminative features among cells. The features may be pathologist-provided or inferred directly from visual data. We validate our model on two publicly available datasets (CoNSeP and MoNuSAC), and on MuCeD, a novel dataset that we contribute. MuCeD consists of 55 biopsy images of the human duodenum for predicting celiac disease. We perform extensive experimentation with three object detection baselines on three datasets to show that DeGPR is model-agnostic, and consistently improves baselines obtaining up to 9% (absolute) mAP gains.
Aayush Kumar Tyagi, Chirag Mohapatra, Prasenjit Das 0006, Govind Makharia, Lalita Mehra, Prathosh A. P., Mausam
CVPR6
2023 Few-shot Cross-domain Image Generation via Inference-time Latent-code Learning
Arnab Kumar Mondal, Piyush Tiwary, Parag Singla, Prathosh A. P.
ICLR4
2023 SSDMM-VAE: variational multi-modal disentangled representation learning
Arnab Kumar Mondal, Ajay Sailopal, Parag Singla, Prathosh A. P.
Appl. Intell.4
2023 Clustering Single-Cell RNA Sequence Data Using Information Maximized and Noise-Invariant Representations
abstract
Single-cell RNA sequencing (scRNA-seq) is a revolutionary methodology that helps to analyze transcriptome or genome information from a single cell. However, high dimensionality and sparsity in data due to dropout events pose computational challenges for existing state-of-the-art scRNA-seq clustering methods. Learning efficient representations becomes even more challenging due to the presence of noise in scRNA-seq data. To overcome the effect of noise and learn effective representations, this paper proposes sc-INDC (Single-Cell Information Maximized Noise-Invariant Deep Clustering), a deep neural network that facilitates learning of informative and noise-invariant representations of scRNA-seq data. Furthermore, the time complexity of the proposed sc-INDC is significantly lower compared to state-of-the-art scRNA-seq clustering methods. Extensive experimentation on fourteen publicly available scRNA-seq datasets illustrates the efficacy of the proposed model. Additionally, visualizations of t-SNE plots and several ablation studies are also conducted to provide insights into the improved representation ability of sc-INDC. Code of the proposed sc-INDC will be available at: https://github.com/arnabkmondal/sc-INDC.
Arnab Kumar Mondal, Indu Joshi, Pravendra Singh, Prathosh A. P.
IEEE ACM Trans. Comput. Biol. Bioinform.4
2022 Deep Learning-Based Human Action Recognition Framework to Assess Children on the Risk of Autism or Developmental Delays
Manu Kohli, Arpan Kumar Kar, Varun Ganjigunte Prakash, Prathosh A. P.
ICONIP (7)4
2022 scRAE: Deterministic Regularized Autoencoders With Flexible Priors for Clustering Single-Cell Gene Expression Data
abstract
Clustering single-cell RNA sequence (scRNA-seq) data poses statistical and computational challenges due to their high-dimensionality and data-sparsity, also known as 'dropout' events. Recently, Regularized Auto-Encoder (RAE) based deep neural network models have achieved remarkable success in learning robust low-dimensional representations. The basic idea in RAEs is to learn a non-linear mapping from the high-dimensional data space to a low-dimensional latent space and vice-versa, simultaneously imposing a distributional prior on the latent space, which brings in a regularization effect. This paper argues that RAEs suffer from the infamous problem of bias-variance trade-off in their naive formulation. While a simple AE wita latent regularization results in data over-fitting, a very strong prior leads to under-representation and thus bad clustering. To address the above issues, we propose a modified RAE framework (called the scRAE) for effective clustering of the single-cell RNA sequencing data. scRAE consists of deterministic AE with a flexibly learnable prior generator network, which is jointly trained with the AE. This facilitates scRAE to trade-off better between the bias and variance in the latent space. We demonstrate the efficacy of the proposed method through extensive experimentation on several real-world single-cell Gene expression datasets. The code for our work is available at https://github.com/arnabkmondal/scRAE.
Arnab Kumar Mondal, Himanshu Asnani, Parag Singla, Prathosh A. P.
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Unsupervised Domain Adaptation Schemes for Building ASR in Low-Resource Languages
abstract
Building an automatic speech recognition (ASR) system from scratch requires a large amount of annotated speech data, which is difficult to collect in many languages. However, there are cases where the low-resource language shares a common acoustic space with a high-resource language having enough annotated data to build an ASR. In such cases, we show that the domain-independent acoustic models learned from the high-resource language through unsupervised domain adaptation (UDA) schemes can enhance the performance of the ASR in the low-resource language. We use the specific example of Hindi in the source domain and Sanskrit in the target domain. We explore two architectures: i) domain adversarial training using gradient reversal layer (GRL) and ii) domain separation networks (DSN). The GRL and DSN architectures give absolute improvements of 6.71% and 7.32%, respectively, in word error rate over the baseline deep neural network model when trained on just 5.5 hours of data in the target domain. We also show that choosing a proper language (Telugu) in the source domain can bring further improvement. The results suggest that UDA schemes can be helpful in the development of ASR systems for low-resource languages, mitigating the hassle of collecting large amounts of annotated speech data.
Chandran Savithri Anoop, Prathosh A. P., A. G. Ramakrishnan
ASRU2
2021 Generalization on Unseen Domains via Inference-Time Label-Preserving Target Projections
abstract
Generalization of machine learning models trained on a set of source domains on unseen target domains with different statistics, is a challenging problem. While many approaches have been proposed to solve this problem, they only utilize source data during training but do not take advantage of the fact that a single target example is available at the time of inference. Motivated by this, we propose a method that effectively uses the target sample during inference beyond mere classification. Our method has three components - (i) A label-preserving feature or metric transformation on source data such that the source samples are clustered in accordance with their class irrespective of their domain (ii) A generative model trained on the these features (iii) A label-preserving projection of the target point on the source-feature manifold during inference via solving an optimization problem on the input space of the generative model using the learned metric. Finally, the projected target is used in the classifier. Since the projected target feature comes from the source manifold and has the same label as the real target by design, the classifier is expected to perform better on it than the true target. We demonstrate that our method outperforms the state-of-the-art Domain Generalization methods on multiple datasets and tasks.
Prashant Pandey 0002, Mrigank Raman, Sumanth Varambally, Prathosh A. P.
CVPR4
2021 Systematic Generalization in Neural Networks-based Multivariate Time Series Forecasting Models
abstract
Systematic generalization aims to evaluate reasoning about novel combinations from known components, an intrinsic property of human cognition. In this work, we study systematic generalization of Neural Networks (NNs) in forecasting future time series of dependent variables in a dynamical system, conditioned on past time series of dependent variables, and past and future control variables. We focus on systematic generalization wherein the NN-based forecasting model should perform well on previously unseen combinations or regimes of control variables after being trained on a limited set of the possible regimes. For NNs to depict such out-of-distribution generalization, they should be able to disentangle the various dependencies between control variables and dependent variables. We hypothesize that a modular NN architecture guided by the readily-available knowledge of independence of control variables as a potentially useful inductive bias to this end. Through extensive empirical evaluation on a toy dataset and a simulated electric motor dataset, we show that our proposed modular NN architecture serves as a simple yet highly effective inductive bias that enabling better forecasting of the dependent variables up to large horizons in contrast to standard NNs, and indeed capture the true dependency relations between the dependent and the control variables.
Hritik Bansal, Gantavya Bhatt, Pankaj Malhotra, Prathosh A. P.
IJCNN4
2021 FlexAE: flexibly learning latent priors for wasserstein auto-encoders
abstract
Auto-Encoder (AE) based neural generative frameworks model the joint-distribution between the data and the latent space using an Encoder-Decoder pair, with regularization imposed in terms of a prior over the latent space. Despite their advantages, such as stability in training, efficient inference, the performance of AE based models has not reached the superior standards of the other generative models such as Generative Adversarial Networks (GANs). Motivated by this, we examine the effect of the latent prior on the generation quality of deterministic AE models in this paper. Specifically, we consider the class of Generative AE models with deterministic Encoder-Decoder pair (such as Wasserstein Auto-Encoder (WAE), Adversarial Auto-Encoder (AAE)), and show that having a fixed prior distribution, a priori, oblivious to the dimensionality of the ‘true’ latent space, will lead to the infeasibility of the optimization problem considered. As a remedy to the issue mentioned above, we introduce an additional state space in the form of flexibly learnable latent priors, in the optimization objective of WAE/AAE. Additionally, we employ a latent-space interpolation based smoothing scheme to address the non-smoothness that may arise from highly flexible priors. We show the efficacy of our proposed models, called FlexAE and FlexAE-SR, through several experiments on multiple datasets, and demonstrate that FlexAE-SR is the new state-of-the-art for the AE based generative models in terms of generation quality as measured by several metrics such as Fr\’echet Inception Distance, Precision/Recall score.
Arnab Kumar Mondal, Himanshu Asnani, Parag Singla, Prathosh A. P.
UAI4
2021 A Variational Information Bottleneck Based Method to Compress Sequential Networks for Human Action Recognition
abstract
In the last few years, deep neural networks' compression has become an important strand of machine learning and computer vision research. Deep models require sizeable computational complexity and storage when used, for instance, for Human Action Recognition (HAR) from videos, making them unsuitable to be deployed on edge devices. In this paper, we address this issue and propose a method to effectively compress Recurrent Neural Networks (RNNs) such as Gated Recurrent Units (GRUs) and Long-Short-Term-Memory Units (LSTMs) that are used for HAR. We use a Variational Information Bottleneck (VIB) theory-based pruning approach to limit the information flow through the sequential cells of RNNs to a small subset. Further, we combine our pruning method with a specific group-lasso regularization technique that significantly improves compression. The proposed techniques reduce model parameters and memory footprint from latent representations, with little or no reduction in the validation accuracy while increasing the inference speed several-fold. We perform experiments on the three widely used Action Recognition datasets, viz. UCF11, HMDB51, and UCF101, to validate our approach. We show that our method achieves over 70 times greater compression than the nearest competitor with comparable accuracy for action recognition on UCF11.
Ayush Srivastava, Oshin Dutta, Jigyasa Gupta, Sumeet Agarwal, Prathosh A. P.
WACV5
2020 Guided Weak Supervision for Action Recognition with Scarce Data to Assess Skills of Children with Autism
abstract
Diagnostic and intervention methodologies for skill assessment of autism typically requires a clinician repetitively initiating several stimuli and recording the child's response. In this paper, we propose to automate the response measurement through video recording of the scene following the use of Deep Neural models for human action recognition from videos. However, supervised learning of neural networks demand large amounts of annotated data that is hard to come by. This issue is addressed by leveraging the ‘similarities’ between the action categories in publicly available large-scale video action (source) datasets and the dataset of interest. A technique called Guided Weak Supervision is proposed, where every class in the target data is matched to a class in the source data using the principle of posterior likelihood maximization. Subsequently, classifier on the target data is re-trained by augmenting samples from the matched source classes, along with a new loss encouraging inter-class separability. The proposed method is evaluated on two skill assessment autism datasets, SSBD (Sundar Rajagopalan, Dhall, and Goecke 2013) and a real world Autism dataset comprising 37 children of different ages and ethnicity who are diagnosed with autism. Our proposed method is found to improve the performance of the state-of-the-art multi-class human action recognition models in-spite of supervision with scarce data.
Prashant Pandey 0002, Prathosh A. P., Manu Kohli, Josh Pritchard
AAAI2
2020 Unsupervised Domain Adaptation for Semantic Segmentation of NIR Images Through Generative Latent Search
Prashant Pandey 0002, Aayush Kumar Tyagi, Sameer Ambekar, Prathosh A. P.
ECCV (6)4
2020 Variational Inference with Latent Space Quantization for Adversarial Resilience
abstract
Despite their tremendous success in modelling high-dimensional data manifolds, deep neural networks suffer from the threat of adversarial attacks - Existence of perceptually valid input-like samples obtained through careful perturbation that lead to degradation in the performance of the underlying model. Major concerns with existing defense mechanisms include non-generalizability across different attacks, models and large inference time. In this paper, we propose a generalized defense mechanism capitalizing on the expressive power of regularized latent space based generative models. We design an adversarial filter, devoid of access to classifier and adversaries, which makes it usable in tandem with any classifier. The basic idea is to learn a Lipschitz constrained mapping from the data manifold, incorporating adversarial perturbations, to a quantized latent space and re-map it to the true data manifold. Specifically, we simultaneously auto-encode the data manifold and its perturbations implicitly through the perturbations of the regularized and quantized generative latent space, realized using variational inference. We demonstrate the efficacy of the proposed formulation in providing resilience against multiple attack types (black and white box) and methods, while being almost real-time. Our experiments show that the proposed method surpasses the state-of-the-art techniques in several cases. The implementation code is available at - https://github.com/mayank31398/lqvae.
Vinay Kyatham, Deepak Mishra 0003, Prathosh A. P.
ICPR3
2020 C-MI-GAN : Estimation of Conditional Mutual Information using MinMax formulation
abstract
Estimation of information theoretic quantities such as mutual information and its conditional variant has drawn interest in recent times owing to their multifaceted applications. Newly proposed neural estimators for these quantities have overcome severe drawbacks of classical $k$NN-based estimators in high dimensions. In this work, we focus on conditional mutual information (CMI) estimation by utilizing its formulation as a \textit{minmax} optimization problem. Such a formulation leads to a joint training procedure similar to that of generative adversarial networks. We find that our proposed estimator provides better estimates than the existing approaches on a variety of simulated datasets comprising linear and non-linear relations between variables. As an application of CMI estimation, we deploy our estimator for conditional independence (CI) testing on real data and obtain better results than state-of-the-art CI testers.
Arnab Kumar Mondal, Arnab Bhattacharjee, Sudipto Mukherjee 0001, Himanshu Asnani, Sreeram Kannan, Prathosh A. P.
UAI6
2020 MaskAAE: Latent space optimization for Adversarial Auto-Encoders
abstract
The field of neural generative models is dominated by the highly successful Generative Adversarial Networks (GANs) despite their challenges, such as training instability and mode collapse. Auto-Encoders (AE) with regularized latent space provide an alternative framework for generative models, albeit their performance levels have not reached that of GANs. In this work, we hypothesise that the dimensionality of the AE model’s latent space has a critical effect on the quality of generated data. Under the assumption that nature generates data by sampling from a “true" generative latent space followed by a deterministic function, we show that the optimal performance is obtained when the dimensionality of the latent space of the AE-model matches with that of the “true" generative latent space. Further, we propose an algorithm called the Mask Adversarial Auto-Encoder (MaskAAE), in which the dimensionality of the latent space of an adversarial auto encoder is brought closer to that of the “true" generative latent space, via a procedure to mask the spurious latent dimensions. We demonstrate through experiments on synthetic and several real-world datasets that the proposed formulation yields betterment in the generation quality.
Arnab Kumar Mondal, Sankalan Pal Chowdhury, Aravind Jayendran, Himanshu Asnani, Parag Singla, Prathosh A. P.
UAI6
2020 Effect of the Latent Structure on Clustering With GANs
abstract
Generative adversarial networks (GANs) have shown remarkable success in the generation of data from natural data manifolds such as images. In several scenarios, it is desirable that generated data is well-clustered, especially when there is severe class imbalance. In this paper, we focus on the problem of clustering in the generated space of GANs and uncover its relationship with the characteristics of the latent space. We derive from first principles, the necessary and sufficient conditions needed to achieve faithful clustering in the GAN framework: (i) presence of a multimodal latent space with adjustable priors, (ii) existence of a latent space inversion mechanism and, (iii) imposition of the desired cluster priors on the latent space. We also identify the GAN models in the literature that partially satisfy these conditions and demonstrate the importance of all the components required, through ablative studies on multiple real-world image datasets. Additionally, we describe a procedure to construct a multimodal latent space which facilitates learning of cluster priors with sparse supervision. Codes for our implementation is available at https://github.com/NEMGAN/NEMGAN-P.
Deepak Mishra 0003, Aravind Jayendran, Prathosh A. P.
IEEE Signal Process. Lett.3
2020 Target-Independent Domain Adaptation for WBC Classification Using Generative Latent Search
abstract
Automating the classification of camera-obtained microscopic images of White Blood Cells (WBCs) and related cell subtypes has assumed importance since it aids the laborious manual process of review and diagnosis. Several State-Of-The-Art (SOTA) methods developed using Deep Convolutional Neural Networks suffer from the problem of domain shift - severe performance degradation when they are tested on data (target) obtained in a setting different from that of the training (source). The change in the target data might be caused by factors such as differences in camera/microscope types, lenses, lighting-conditions etc. This problem can potentially be solved using Unsupervised Domain Adaptation (UDA) techniques albeit standard algorithms presuppose the existence of a sufficient amount of unlabelled target data which is not always the case with medical images. In this paper, we propose a method for UDA that is devoid of the need for target data. Given a test image from the target data, we obtain its 'closest-clone' from the source data that is used as a proxy in the classifier. We prove the existence of such a clone given that infinite number of data points can be sampled from the source distribution. We propose a method in which a latent-variable generative model based on variational inference is used to simultaneously sample and find the 'closest-clone' from the source distribution through an optimization procedure in the latent space. We demonstrate the efficacy of the proposed method over several SOTA UDA methods for WBC classification on datasets captured using different imaging modalities under multiple settings.
Prashant Pandey 0002, Prathosh A. P., Vinay Kyatham, Deepak Mishra 0003, Tathagato Rai Dastidar
IEEE Trans. Medical Imaging2
2019 Detection of Glottal Closure Instants from Raw Speech Using Convolutional Neural Networks
Mohit Goyal, Prathosh A. P.
INTERSPEECH3
2016 A Vision Based Method for Real-Time Respiration Rate Estimation Using a Recursive Fourier Analysis
abstract
In this paper, we propose a simple yet effective, computer vision based method for estimating respiration rate in real-time from the thoraco-abdominal video of a subject being monitored. The periodic motion of the chest wall of the subject is captured through the optical flow in the video sequence. The frequency of the chest wall motion is estimated by performing a Fourier analysis on the time sequence of the optical flow vectors. We present how to perform the Fourier analysis recursively leveraging the sequential nature of a video to speed-up our method. Unlike other methods, our method does not require a selection of region of interest because our method aggregates out-of-phase optical flows by factoring out the their relative phase differences. Unlike many existing methods, our method does not require to be reinitialized if the subject changes the posture during the observation. Our method works with different postures and different views (frontal, side, etc.) of the subject. Our method is very simple to implement. We evaluate our method against an impedance pneumograph and demonstrate the high accuracy of our method on thoracoabdominal videos of many subjects wearing a wide variety of clothing.
Avishek Chatterjee, Prathosh A. P., Pragathi Praveena, Vidyadhar Upadhya
BIBE2
2016 Real-Time Visual Respiration Rate Estimation with Dynamic Scene Adaptation
abstract
In this paper, we present a vision based method for respiration rate estimation which can automatically adapt to the scene changes. We capture a video of the subjects thoraco-abdominal region and compute optical flow field at each video frame. The optical flow field changes periodically with the periodic chest wall motion. The pattern of the chest wall motion is captured through the estimation of a principal flow field. The principal flow field is automatically updated with time to cope with the scene changes. Thus, our method can adapt itself to the changes of the posture of a subject. Besides, in our method we do not need to select any region of interest unlike other methods. Yet our method is computationally very inexpensive and simple to implement. We test our method on many human volunteers with a wide variety of their clothing. We compare our method against the gold standard method of impedance pneumography and have found a very high accuracy.
Avishek Chatterjee, Prathosh A. P., Pragathi Praveena, Vidyadhar Upadhya
BIBE2
2016 Respiration Monitoring through Thoraco-Abdominal Video with an LSTM
abstract
In this manuscript, we demonstrate the estimation of the respiratory signal from a thoraco-abdominal video of a person using an LSTM based learning model. The video is captured with a regular consumer grade camera and the respiratory signal is recorded using an impedance pneumograph simultaneously. The optical flow capturing the motion of the chest wall during an inhalation and exhalation is extracted at each video frame and fed as features to the LSTM model. We then train the LSTM model to estimate the respiratory signal. We fix the design parameters of the LSTM model based on cross-validation. The comparison between the predicted and the ground-truth pneumograph signal shows that the trained LSTM model predicts the respiratory signal quite accurately achieving a strong amplitude correlation of 0.74. Moreover, we estimate the respiration rates from the predicted respiratory signal. The estimated respiration rates have less than ±3 BPM error for more than 95% cases. Also, we achieve a correlation of 0.9 between the ground-truth respiration rates and the estimated respiration rates.
Vidyadhar Upadhya, Avishek Chatterjee, Prathosh A. P., Pragathi Praveena
BIBE3
2016 Novel acoustic features for automatic dialog-act tagging
abstract
This paper presents 57 new acoustic features for automatic dialog-act tagging. The features are intended to be richer than and complementary to the traditional cumulative statistics of intonation. Some of our novel contributions include feature normalization with respect to neighboring utterances, incorporation of periodicity and formant features, modeling of cognitive phenomena such as hesitations, and utterance-level aggregation of short-term acoustic effects. The proposed features are applied to 3-way dialog-act tagging and question detection using two databases (British-English call-center conversations and Switchboard), and compared with a popular cumulative-statistics baseline using logistic-regression models. Our features are found to be significantly better than and complementary to the baseline, on average, achieving an absolute performance gain of ~5-6%. Combined feature ranking reveals that about 75% of the top 20 features belong to the proposed feature set, and that the two corpora differ in their feature preferences despite similar overall performance.
Harish Arsikere, Arunasish Sen, Prathosh A. P., Vivek Tyagi
ICASSP3
2016 Cumulative Impulse Strength for Epoch Extraction
abstract
Algorithms for extracting epochs or glottal closure instants (GCIs) from voiced speech typically fall into two categories: i) ones which operate on linear prediction residual (LPR) and ii) those which operate directly on the speech signal. While the former class of algorithms (such as YAGA and DPI) tend to be more accurate, the latter ones (such as ZFR and SEDREAMS) tend to be more noise-robust. In this letter, a temporal measure termed the cumulative impulse strength is proposed for locating the impulses in a quasi-periodic impulse-sequence embedded in noise. Subsequently, it is applied for detecting the GCIs from the inverted integrated LPR using a recursive algorithm. Experiments on two large corpora of speech with simultaneous electroglottographic recordings demonstrate that the proposed method is more robust to additive noise than the state-of-the-art algorithms, despite operating on the LPR.
Prathosh A. P., P. Sujith, A. G. Ramakrishnan, Prasanta Kumar Ghosh
IEEE Signal Process. Lett.1
2015 Classification of place-of-articulation of stop consonants using temporal analysis
abstract
This paper proposes acoustic-phonetic features for classification of place-of-articulation of stop consonants derived from their temporal structures. The speech signal corresponding to a stop is characterized by several temporal features such as sub-band zero-crossings and envelope fits. Classification experiments on the stops from the TIMIT (read speech) and the Buckeye (conversational speech) databases using a support vector machine classifier demonstrate that the performance of the proposed features (84.6 %) is comparable to that obtained by MFCCs (85.1 %) in many aspects. Further, the classification accuracy is boosted (90.1 %) with the combination of temporal and MFCC features, which substantiates their supplementary nature.
Prathosh A. P., A. G. Ramakrishnan, T. V. Ananthapadmanabha
INTERSPEECH1
2015 An error correction scheme for GCI detection algorithms using pitch smoothness criterion
abstract
Detection of error-free glottal closure instants (GCI) is a critical requirement for many applications including text-to-speech synthesis, causal anti-causal decomposition and voice morphing. Many existing GCI detection algorithms commit errors under certain conditions. In this paper, we propose a post processing scheme for correcting errors of any GCI detection algorithm. The proposed error correction scheme works on the principle that the fundamental frequency over a voiced segment is slowly varying. The error correction is thus formulated as an optimization problem such that the pitch contour from the corrected GCIs has the least high frequency components. The proposed error correction scheme is experimentally evaluated on speech corpus with simultaneous EGG recordings using three state-of-the-art GCI detection algorithms viz., Dynamic Plosion Index (DPI), Zero Frequency Resonator (ZFR), and Speech Event Detection using the Residual Excitation And a Mean-based Signal (SEDREAMS). It is found that the proposed error correction scheme improves the performance of the GCI detection in clean speech as well as noisy conditions at different SNRs.
P. Sujith, Prathosh A. P., A. G. Ramakrishnan, Prasanta Kumar Ghosh
INTERSPEECH2
2014 Threshold-Independent QRS Detection Using the Dynamic Plosion Index
abstract
Detection of QRS serves as a first step in many automated ECG analysis techniques. Motivated by the strong similarities between the signal structures of an ECG signal and the integrated linear prediction residual (ILPR) of voiced speech, an algorithm proposed earlier for epoch detection from ILPR is extended to the problem of QRS detection. The ECG signal is pre-processed by high-pass filtering to remove the baseline wandering and by half-wave rectification to reduce the ambiguities. The initial estimates of the QRS are iteratively obtained using a non-linear temporal feature, named the dynamic plosion index suitable for detection of transients in a signal. These estimates are further refined to obtain a higher temporal accuracy. Unlike most of the high performance algorithms, this technique does not make use of any threshold or differencing operation. The proposed algorithm is validated on the MIT-BIH database using the standard metrics and its performance is found to be comparable to the state-of-the-art algorithms, despite its threshold independence and simple decision logic.
A. G. Ramakrishnan, Prathosh A. P., T. V. Ananthapadmanabha
IEEE Signal Process. Lett.2
2013 Epoch Extraction Based on Integrated Linear Prediction Residual Using Plosion Index
abstract
Epoch is defined as the instant of significant excitation within a pitch period of voiced speech. Epoch extraction continues to attract the interest of researchers because of its significance in speech analysis. Existing high performance epoch extraction algorithms require either dynamic programming techniques or a priori information of the average pitch period. An algorithm without such requirements is proposed based on integrated linear prediction residual (ILPR) which resembles the voice source signal. Half wave rectified and negated ILPR (or Hilbert transform of ILPR) is used as the pre-processed signal. A new non-linear temporal measure named the plosion index (PI) has been proposed for detecting `transients' in speech signal. An extension of PI, called the dynamic plosion index (DPI) is applied on pre-processed signal to estimate the epochs. The proposed DPI algorithm is validated using six large databases which provide simultaneous EGG recordings. Creaky and singing voice samples are also analyzed. The algorithm has been tested for its robustness in the presence of additive white and babble noise and on simulated telephone quality speech. The performance of the DPI algorithm is found to be comparable or better than five state-of-the-art techniques for the experiments considered.
Prathosh A. P., T. V. Ananthapadmanabha, A. G. Ramakrishnan
IEEE ACM Trans. Audio Speech Lang. Process.1