Aditya Sinha

dblp:13/7739 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 CADRE: Customizable Assurance of Data Readiness in Privacy-Preserving Federated Learning
abstract
Privacy-Preserving Federated Learning (PPFL) is a decentralized machine learning approach where multiple clients train a model collaboratively. PPFL preserves the privacy and security of a client’s data without exchanging it. However, ensuring that data at each client is of high quality and ready for federated learning (FL) is a challenge due to restricted data access. In this paper, we introduce CADRE (Customizable Assurance of Data REadiness) for federated learning (FL), a novel framework that allows users to define custom data readiness (DR) metrics, rules, and remedies tailored to specific FL tasks. CADRE generates comprehensive DR reports based on the user-defined metrics, rules, and remedies to ensure datasets are prepared for FL while preserving privacy. We demonstrate a practical application of CADRE by integrating it into an existing PPFL framework. We conducted experiments across six datasets and addressed seven different DR issues. The results illustrate the versatility and effectiveness of CADRE in ensuring DR across various dimensions, including data quality, privacy, and fairness. This approach enhances the performance and reliability of FL models as well as utilizes valuable resources.
Kaveen Hiniduma, Zilinghan Li, Aditya Sinha, Ravi K. Madduri, Surendra Byna
eScience3
2025 Cost-Aware Federated Learning on the Cloud
abstract
We introduce FedCostAware, a cost-aware scheduling algorithm designed to optimize synchronous federated learning (FL) on cloud spot instances, which addresses the challenges of training on spot instances and different client budgets by employing intelligent management of the lifecycle of spot instances. This approach minimizes idle resource time and overall expenses. Experiments on real-world medical datasets demonstrate that FedCostAware significantly reduces cloud computing costs compared to conventional spot and on-demand schemes, enhancing the accessibility and affordability of FL.
Aditya Sinha, Zilinghan Li, Tingkai Liu, Volodymyr V. Kindratenko, Kibaek Kim, Ravi K. Madduri
eScience1
2025 From image processing to artificial intelligence-driven tools: A comprehensive survey on the evolution of feature extraction methods in paintings
Rekha Sharma, Rishi Gupta, Aditya Sinha
Eng. Appl. Artif. Intell.3
2024 Learning Structured Representations with Hyperbolic Embeddings
abstract
Most real-world datasets consist of a natural hierarchy between classes or an inherent label structure that is either already available or can be constructed cheaply. However, most existing representation learning methods ignore this hierarchy, treating labels as permutation invariant. Recent work [Zeng et al., 2022] proposes using this structured information explicitly, but the use of Euclidean distance may distort the underlying semantic context [Chen et al., 2013]. In this work, motivated by the advantage of hyperbolic spaces in modeling hierarchical relationships, we propose a novel approach HypStructure: a Hyperbolic Structured regularization approach to accurately embed the label hierarchy into the learned representations. HypStructure is a simple-yet-effective regularizer that consists of a hyperbolic tree-based representation loss along with a centering loss, and can be combined with any standard task loss to learn hierarchy-informed features. Extensive experiments on several large-scale vision benchmarks demonstrate the efficacy of HypStructure in reducing distortion and boosting generalization performance especially under low dimensional scenarios. For a better understanding of structured representation, we perform eigenvalue analysis that links the representation geometry to improved Out-of-Distribution (OOD) detection performance seen empirically.
Aditya Sinha, Siqi Zeng 0001, Makoto Yamada, Han Zhao 0002
NeurIPS1
2024 Vikriti-ID: A Novel Approach For Real Looking Fingerprint Data-set Generation
abstract
Fingerprint recognition research faces significant challenges due to the limited availability of extensive and publicly available fingerprint databases. Existing databases lack a sufficient number of identities and fingerprint impressions, which hinders progress in areas such as Fingerprint-based access control. To address this challenge, we present Vikriti-ID, a synthetic fingerprint generator capable of generating unique fingerprints with multiple impressions. Using Vikriti-ID, we generated a large database containing 500000 unique fingerprints, each with 10 associated impressions. We then demonstrate the effectiveness of the database generated by Vikriti-ID by evaluating it for imposter-genuine score distribution and EER score. Apart from this we also trained a deep network to check the usability of data. We trained the network inspired from [13], on both Vikriti-ID generated data as well as public data. This generated data achieved an Equal Error Rate(EER) of 0.16%, AUC of 0.89%. This improvement is possible due to the limitations of existing publicly available data sets, which struggle in numbers or multiple impressions.
Rishabh Shukla, Aditya Sinha, Vansh Singh, Harkeerat Kaur
WACV2
2023 Novel approach for quantification for severity estimation of blight diseases on leaves of tomato plant
abstract
Abstract This study uses digital image processing and machine learning to quantify the infection patterns on tomato leaves due to blight diseases. Quantification, also known as severity measurement, is a technique to determine how much a leaf is diseased by calculating a numeric value. This value could be a fraction representing how much the diseased region is present on the leaf compared to the entire leaf region, or it can be a percentage value too. There are two main approaches to measuring disease severity; the first technique involves visual estimation using references like standard area diagrams. The second approach involves taking a digital image of the leaf, separating the diseased regions from the healthy regions, and then calculating the area of those two regions. The approach we took is similar. We first took the digital image and segmented the diseased and healthy regions. For quantification, we calculated the ratio of total pixels representing the diseased region to the total number of pixels representing the leaf. While finding ways to improve the accuracy of the segmentation algorithm, we also discovered our segmentation technique which automatically segments the diseased regions of the leaves from the healthy areas using k‐means clustering. The clustering‐segmentation algorithm did give good results for the sample images to which it was applied. The main thing about the clustering‐segmentation algorithm is that it tends to be automatic compared to some of the semi‐automatic segmentation approaches that have been discovered till now. We could reproduce the validated quantification results as other authors achieved in the recent work, which also validated our methodology.
Aahan Singh Charak, Aditya Sinha, Tarun Jain
Expert Syst. J. Knowl. Eng.2
2022 IGLU: Efficient GCN Training via Lazy Updates
S. Deepak Narayanan, Aditya Sinha, Prateek Jain 0002, Purushottam Kar, Sundararajan Sellamanickam
ICLR2
2022 S3GC: Scalable Self-Supervised Graph Clustering
abstract
We study the problem of clustering graphs with additional side-information of node features. The problem is extensively studied, and several existing methods exploit Graph Neural Networks to learn node representations. However, most of the existing methods focus on generic representations instead of their cluster-ability or do not scale to large scale graph datasets. In this work, we propose S3GC which uses contrastive learning along with Graph Neural Networks and node features to learn clusterable features. We empirically demonstrate that S3GC is able to learn the correct cluster structure even when graph information or node features are individually not informative enough to learn correct clusters. Finally, using extensive evaluation on a variety of benchmarks, we demonstrate that S3GC is able to significantly outperform state-of-the-art methods in terms of clustering accuracy -- with as much as 5% gain in NMI -- while being scalable to graphs of size 100M.
Devvrit, Aditya Sinha, Inderjit S. Dhillon, Prateek Jain 0002
NeurIPS2
2022 Matryoshka Representation Learning
abstract
Learned representations are a central component in modern ML systems, serving a multitude of downstream tasks. When training such representations, it is often the case that computational and statistical constraints for each downstream task are unknown. In this context rigid, fixed capacity representations can be either over or under-accommodating to the task at hand. This leads us to ask: can we design a flexible representation that can adapt to multiple downstream tasks with varying computational resources? Our main contribution is Matryoshka Representation Learning (MRL) which encodes information at different granularities and allows a single embedding to adapt to the computational constraints of downstream tasks. MRL minimally modifies existing representation learning pipelines and imposes no additional cost during inference and deployment. MRL learns coarse-to-fine representations that are at least as accurate and rich as independently trained low-dimensional representations. The flexibility within the learned Matryoshka Representations offer: (a) up to $\mathbf{14}\times$ smaller embedding size for ImageNet-1K classification at the same level of accuracy; (b) up to $\mathbf{14}\times$ real-world speed-ups for large-scale retrieval on ImageNet-1K and 4K; and (c) up to $\mathbf{2}\%$ accuracy improvements for long-tail few-shot classification, all while being as robust as the original representations. Finally, we show that MRL extends seamlessly to web-scale datasets (ImageNet, JFT) across various modalities -- vision (ViT, ResNet), vision + language (ALIGN) and language (BERT). MRL code and pretrained models are open-sourced at https://github.com/RAIVNLab/MRL.
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham M. Kakade, Prateek Jain 0002, Ali Farhadi
NeurIPS5
2020 Review of image processing approaches for detecting plant diseases
abstract
There is intense pressure on agricultural productivity due to the ever‐growing population. Several diseases affect crop yield and thus, effective control of these can significantly improve the production of food for all. In this regard, detection of diseases at an early stage and quantification of the severity, in general, has acquired urgent attention of the researchers. In this study, a summary of prevalent techniques and methodologies used for the detection, quantification and classification of diseases is presented to understand the scope of improvement. The study pays attention to critical gaps that exist in available approaches and enhance them for the early prediction of diseases. Diseases affect almost all parts of plants, e.g. root, stem, flower, leaf; a manifestation in different ways for different parts of the plant of the same disease presents a challenge for researchers. This study extends the review work published by JGA Barbedo in 2013, as there have been significant advances and numerous new techniques introduced since then. A novel approach of classifying and categorisation of the existing techniques based on pathogen types is a significant contribution by the authors in this study.
Aditya Sinha, Rajveer Singh Shekhawat
IET Image Process.1
2018 On the Secrecy Capacity of 2-user Gaussian Interference Channel with Independent Secret Keys
abstract
This paper considers the problem of secure communication over a 2-user Gaussian interference channel (GIC) with shared key of finite rate between the transmitter-receiver pair with strong secrecy constraint at the receiver. The main contributions of the paper lies in obtaining a novel achievable scheme which uses a combination of one-time pad, stochastic encoding and superposition based coding scheme, and outer bound on the secrecy capacity region of the 2-user GIC. The main novelty of the derivation of the outer bound lies in the selection of the side-information to be provided to the receiver and using the secrecy constraints at the receiver. The results highlight the role of secret key in the encoding of messages to enhance the system performance in interference limited scenarios.
Aditya Sinha, Parthajit Mohapatra, Jemin Lee 0002, Tony Q. S. Quek
ISITA1
2009 Algorithm for Concept Traversal in Multiple Databases
abstract
In this paper we present an efficient algorithm for simultaneous traversal of concept lattices of two distinct but related binary datasets.
Aditya Sinha, Raj Bhatnagar
ICTAI1