VLDB 2026 Research / reviewers in the wild / expert
Jay Nandy
dblp:193/4096
· DBLP profile ↗
10ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0001-8559-4810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Fairness under Covariate Shift: Improving Fairness-Accuracy Tradeoff with Few Unlabeled Test SamplesabstractCovariate shift in the test data is a common practical phenomena that can significantly downgrade both the accuracy and the fairness performance of the model. Ensuring fairness across different sensitive groups under covariate shift is of paramount importance due to societal implications like criminal justice. We operate in the unsupervised regime where only a small set of unlabeled test samples along with a labeled training set is available. Towards improving fairness under this highly challenging yet realistic scenario, we make three contributions. First is a novel composite weighted entropy based objective for prediction accuracy which is optimized along with a representation matching loss for fairness. We experimentally verify that optimizing with our loss formulation outperforms a number of state-of-the-art baselines in the pareto sense with respect to the fairness-accuracy tradeoff on several standard datasets. Our second contribution is a new setting we term Asymmetric Covariate Shift that, to the best of our knowledge, has not been studied before. Asymmetric covariate shift occurs when distribution of covariates of one group shifts significantly compared to the other groups and this happens when a dominant group is over-represented. While this setting is extremely challenging for current baselines, We show that our proposed method significantly outperforms them. Our third contribution is theoretical, where we show that our weighted entropy term along with prediction loss on the training set approximates test loss under covariate shift. Empirically and through formal sample complexity bounds, we show that this approximation to the unseen test loss does not depend on importance sampling variance which affects many other baselines. Shreyas Havaldar, Jatin Chauhan, Karthikeyan Shanmugam 0001, Jay Nandy, Aravindan Raghuveer |
AAAI | 4 |
| 2023 | Non-Uniform Adversarial Perturbations for Discrete Tabular DatasetsabstractWe study the problem of adversarial attack and robustness on tabular datasets with discrete features. The discrete features of a tabular dataset represent high-level meaningful concepts, with different sets of vocabularies, leading to requiring non-uniform robustness. Further, the notion of distance between tabular input instances is not well defined, making the problem of producing adversarial examples with minor perturbations qualitatively more challenging compared to existing methods. Towards this, our paper defines the notion of distance through the lens of feature embeddings, learnt to represent the discrete features. We then formulate the task of generating adversarial examples as abinary set selection problem under non-uniform feature importance. Next, we propose an efficient approximate gradient-descent based algorithm, calledDiscrete Non-uniform Approximation (DNA) attack, by reformulating the problem into a continuous domain to solve the original optimization problem for generating adversarial examples. We demonstrate the effectiveness of our proposed DNA attack using two large real-world discrete tabular datasets from e-commerce domains for binary classification, where the datasets are heavily biased for one-class. We also analyze challenges for existing adversarial training frameworks for such datasets under our DNA attack. Jay Nandy, Jatin Chauhan, Rishi Saket, Aravindan Raghuveer |
CIKM | 1 |
| 2022 | Domain-Agnostic Contrastive Representations for Learning from Label ProportionsabstractWe study the weak supervision learning problem of Learning from Label Proportions (LLP) where the goal is to learn an instance-level classifier using proportions of various class labels in a bag -- a collection of input instances that often can be highly correlated. While representation learning for weakly-supervised tasks is found to be effective, they often require domain knowledge. To the best of our knowledge, representation learning for tabular data (unstructured data containing both continuous and categorical features) are not studied. In this paper, we propose to learn diverse representations of instances within the same bags to effectively utilize the weak bag-level supervision. We propose a domain agnostic LLP method, called "Self Contrastive Representation Learning for LLP" (SelfCLR-LLP) that incorporates a novel self--contrastive function as an auxiliary loss to learn representations on tabular data for LLP. We show that diverse representations for instances within the same bags aid efficient usage of the weak bag-level LLP supervision. We evaluate the proposed method through extensive experiments on real-world LLP datasets from e-commerce applications to demonstrate the effectiveness of our proposed SelfCLR-LLP. In this paper, we propose to learn diverse representations of instances within the same bags to effectively utilize the weak bag-level supervision. We propose a domain agnostic LLP method, called "Self Contrastive Representation Learning for LLP" (SelfCLR-LLP) that incorporates a novel self--contrastive function as an auxiliary loss to learn representations on tabular data for LLP. We show that diverse representations for instances within the same bags aid efficient usage of the weak bag-level LLP supervision. We evaluate the proposed method through extensive experiments on real-world LLP datasets from e-commerce applications to demonstrate the effectiveness of our proposed SelfCLR-LLP. Jay Nandy, Rishi Saket, Jatin Chauhan, Balaraman Ravindran, Aravindan Raghuveer |
CIKM | 1 |
| 2022 | Compact Feature Representation for Unsupervised Ood DetectionabstractDistributional mismatch between training and test data may cause the remote sensing models to behave in unpredictable manner, thus reducing the trustworthiness of such models. Most existing methods for out-of-distribution (OOD) detection rely on availability of OOD samples during training. However, access to OOD data during training is counter intuitive and may be impractical sometimes. Considering this, we propose an unsupervised OOD detection model that does not require training OOD data. The proposed method works by projecting the in-domain samples as a union of 1-dimensional subspaces. Due to the compact feature representation of in-domain samples, OOD samples are less likely to occupy the same feature space, thus they are easily identified. Experimental results demonstrate the capability of the proposed method to detect OOD samples. Sudipan Saha, Jakob Gawlikowski, Jay Nandy, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2022 | Multi-Variate Time Series Forecasting on Variable SubsetsabstractWe formulate a new inference task in the domain of multivariate time series forecasting (MTSF), called Variable Subset Forecast (VSF), where only a small subset of the variables is available during inference. Variables are absent during inference because of long-term data loss (eg. sensor failures) or high -> low-resource domain shift between train / test. To the best of our knowledge, robustness of MTSF models in presence of such failures, has not been studied in the literature. Through extensive evaluation, we first show that the performance of state of the art methods degrade significantly in the VSF setting. We propose a non-parametric, wrapper technique that can be applied on top any existing forecast models. Through systematic experiments across 4 datasets and 5 forecast models, we show that our technique is able to recover close to 95% performance of the models even when only 15% of the original variables are present. Jatin Chauhan, Aravindan Raghuveer, Rishi Saket, Jay Nandy, Balaraman Ravindran |
KDD | 4 |
| 2021 | Distributional Shifts In Automated Diabetic Retinopathy ScreeningabstractDeep learning-based models are developed to automatically detect if a retina image is ‘referable’ in diabetic retinopathy (DR) screening. However, their classification accuracy degrades as the input images distributionally shift from their training distribution. Further, even if the input is not a retina image, a standard DR classifier produces a high confident prediction that the image is ‘referable’. Our paper presents a Dirichlet Prior Network-based framework to address this issue. It utilizes an out-of-distribution (OOD) detector model and a DR classification model to improve generalizability by identifying OOD images. Experiments on real-world datasets indicate that the proposed framework can eliminate the unknown non-retina images and identify the distributionally shifted retina images for human intervention. Jay Nandy, Wynne Hsu, Mong-Li Lee |
ICIP | 1 |
| 2020 | Approximate Manifold Defense Against Multiple Adversarial PerturbationsabstractExisting defenses against adversarial attacks are typically tailored to a specific perturbation type. Using adversarial training to defend against multiple types of perturbation requires expensive adversarial examples from different perturbation types at each training step. In contrast, manifold-based defense incorporates a generative network to project an input sample onto the clean data manifold. This approach eliminates the need to generate expensive adversarial examples while achieving robustness against multiple perturbation types. However, the success of this approach relies on whether the generative network can capture the complete clean data manifold, which remains an open problem for complex input domain. In this work, we devise an approximate manifold defense mechanism, called RBF-CNN, for image classification. Instead of capturing the complete data manifold, we use an RBF layer to learn the density of small image patches. RBF-CNN also utilizes a reconstruction layer that mitigates any minor adversarial perturbations. Further, incorporating our proposed reconstruction process for training improves the adversarial robustness of our RBF-CNN models. Experiment results on MNIST and CIFAR-10 datasets indicate that RBF-CNN offers robustness for multiple perturbations without the need for expensive adversarial training. Jay Nandy, Wynne Hsu, Mong-Li Lee |
IJCNN | 1 |
| 2020 | Towards Maximizing the Representation Gap between In-Domain & Out-of-Distribution ExamplesabstractAmong existing uncertainty estimation approaches, Dirichlet Prior Network (DPN) distinctly models different predictive uncertainty types. However, for in-domain examples with high data uncertainties among multiple classes, even a DPN model often produces indistinguishable representations from the out-of-distribution (OOD) examples, compromising their OOD detection performance. We address this shortcoming by proposing a novel loss function for DPN to maximize the representation gap between in-domain and OOD examples. Experimental results demonstrate that our proposed approach consistently improves OOD detection performance. Jay Nandy, Wynne Hsu, Mong-Li Lee |
NeurIPS | 1 |
| 2018 | Normal Similarity Network for Generative ModellingabstractGaussian distributions are commonly used as a key building block in many generative models. However, their applicability has not been well explored in deep networks. In this paper, we propose a novel deep generative model named as Normal Similarity Network (NSN) where the layers are constructed with Gaussian-style filters. NSN is trained with a layer-wise non-parametric density estimation algorithm that iteratively down-samples the training images and capture the density of the down-sampled training images in the final layer. Additionally, we propose NSN-Gen for generating new samples from noise vectors by iteratively reconstructing feature maps in the hidden layers of NSN. Our experiments suggest encouraging results of the proposed model for a wide range of computer vision applications including image generation, styling and reconstruction from occluded images. Jay Nandy, Wynne Hsu, Mong-Li Lee |
ICIP | 1 |
| 2016 | An Incremental Feature Extraction Framework for Referable Diabetic Retinopathy DetectionabstractDiabetic retinopathy (DR) might be characterized by the occurrence of lesions in the retinal image. Existing approaches require a large set of retinal images where lesions in the image are individually annotated to learn a model that will classify an image as referable or non-referable DR. However, annotating individual lesions is a tedious task and the accuracy of the learnt model is limited by the availability of these annotated images. In this paper, we first learn a universal Gaussian mixture model (GMM) from a small set of annotated images. This universal GMM is then applied as the prior belief to learn an adaptive GMM for individual images. The proposed approach aims to capture the characteristics of referable versus non-referable images by examining the difference between the universal GMM and the adaptive GMM. An image-level classifier is then built based on these differences as features. Experimental results on three fundus image datasets (MESSIDOR, DIARETDB1 and SORC) indicate that the proposed framework achieves 92.1%, 97.68% and 87.1% ROC area values respectively. This approach also opens up a way to use the widely available public fundus images, where the images are labelled but not annotated, for progressively refining the universal GMM leading to an improved performance of approximately 5% and 1% respectively for SORC and MESSIDOR dataset after five refinement steps. Jay Nandy, Wynne Hsu, Mong-Li Lee |
ICTAI | 1 |