Sandeep Kumar 0005

dblp:06/5441-5 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-0415-8625ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leap of FAITH from GNN-to-MLP: Fairness Aware Inference via DisTillation of GrapH Knowledge
Vipul Kumar Singh, Jyotismita Barman, Sandeep Kumar 0005, Tapan Kumar Gandhi, Jayadeva
AAAI3
2026 BiMol-Diff: A Unified Diffusion Framework for Molecular Generation and Captioning
abstract
Aditya Hemant Shahane, Anuj Kumar Sirohi, Devansh Arora, Nitin Kumar, Prathosh AP, Sandeep Kumar. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Aditya Hemant Shahane, Anuj Kumar Sirohi, Devansh Arora, Prathosh A. P., Sandeep Kumar 0005
ACL (1)6
2026 CONNECT-PD: Parkinson's Disease Detection Using Temporal Connectivity Graphs from Gait Data
Ekta Srivastava, Siddhant Ujjain, Tapan Kumar Gandhi, Sandeep Kumar 0005
ICPR (13)4
2026 Fast and Scalable Hashing-Based Universal Graph Coarsening
abstract
Large graphs are becoming ubiquitous, presenting significant computational hurdles in data processing and analysis. Graph Coarsening algorithms are frequently employed to condense large graphs while preserving key graph properties. Real-world graphs also have features or contexts associated with each node. However, existing coarsening methods often overlook simultaneity across node features and structural information. Recent approaches to alleviate this limitation are computationally intensive, and primarily suited for homophilic datasets. Most existing approaches are unsuitable for streaming and evolving graphs, as they require recomputation of the coarsened graph at every timestamp. In this paper, we introduce a Fast and Scalable Hashing-Based Universal Graph Coarsening (UGC) Framework, that integrates locality-sensitive hashing, and feature augmentation to effectively coarsen graphs. UGC is exceptionally fast, straightforward to implement, and capable of handling homophilic, heterophilic, and streaming graphs making it a truly universal solution for graph coarsening. We use an optimization-based framework to minimize a constrained $\epsilon$ε similarity between the original and coarsened graphs, where $\epsilon$ε is between zero and one. Through extensive experimentation on real and synthetic datasets, we demonstrate the effectiveness of our approach in terms of improved runtime complexity and generalization to heterophilic and streaming graphs. Furthermore, we showcase its utility in downstream tasks, emphasizing its scalability for training graph neural networks on coarsened graphs from benchmark real-world datasets.
Mohit Kataria, Nikita Malik, Jayadeva, Sandeep Kumar 0005
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 REFINE: Enabling Efficient and Trustworthy Modeling of Financial Networks via GNN-to-MLP Knowledge Distillation
abstract
Graph Neural Networks (GNNs) have emerged as powerful tools for modeling financial data as networks, effectively capturing both individual attributes and complex relationships. However, their inherent message-passing and aggregation operations introduce significant inference latency, limiting their applicability in latency-sensitive domains such as finance, healthcare, and robotics. Recent efforts have attempted to mitigate this limitation by distilling GNN knowledge into more efficient Multi-Layer Perceptrons (MLPs). While promising in reducing inference costs, existing GNN-to-MLP distillation approaches face three critical challenges: (1) reliance on labeled data, (2) limited robustness to noisy or perturbed inputs due to the absence of structural information, and (3) the existence of representational bias. To address these issues, we propose REFINE, a novel self-supervised GNN-to-MLP knowledge distillation framework. Our method enhances model stability and fairness through structure-free feature augmentations, including noise injection and counterfactual generation. Extensive experiments on two real-world financial datasets and one social network benchmark demonstrate that our approach consistently outperforms existing distillation baselines, achieving a favorable trade-off between predictive utility, stability, and fairness.
Vipul Kumar Singh, Jyotismita Barman, Sandeep Kumar 0005, Jayadeva
DSAA3
2025 Coarse-and-Learn: Efficient Online Node Labeling
Subhanu Halder, Manoj Kumar 0024, Yifan Sun 0001, Sandeep Kumar 0005
ICONIP (4)4
2025 An Efficient Framework for Epidemiological Parameter Estimation via Graph Reduction and Graph Neural Networks
abstract
We propose an epidemiological parameter estimation framework based on contact networks and graph neural networks (GNNs). Contact network-based epidemiological models allow us to capture heterogeneity and individual-level details more effectively. Parameter estimation involves fitting real-world disease data to mathematical models. Traditionally, several likelihood-based methods that focus on compartment-based simulation models have been widely used to perform parameter estimation. However, these methods suffer from making assumptions such as homogeneous traits among the individuals of the population under consideration, which may cause them to fail in handling the complexity and diversity of real-world data. Our proposed framework estimates epidemiological parameters based on the availability of contact network data and individual-level disease time series data. We use supervised as well as self-supervised GNN architectures to incorporate the contact network information into the model. We also employed graph reduction methods such as sampling and coarsening to study scaling behavior and computational efficiency. We formulated the parameter estimation in two ways to study the predictive behavior better: classification and inference problems. We experimentally confirm improvements over the baselines chosen in this article. We also conducted ablation studies, explainability quantification, and scalability experiments to generate further insights into the GNN models.
Muhammad S. T. Alfas, Manoj Kumar 0024, Shaurya Shriyam, Sandeep Kumar 0005
ACM Trans. Knowl. Discov. Data4
2024 No Prejudice! Fair Federated Graph Neural Networks for Personalized Recommendation
abstract
Ensuring fairness in Recommendation Systems (RSs) across demographic groups is critical due to the increased integration of RSs in applications such as personalized healthcare, finance, and e-commerce. Graph-based RSs play a crucial role in capturing intricate higher-order interactions among entities. However, integrating these graph models into the Federated Learning (FL) paradigm with fairness constraints poses formidable challenges as this requires access to the entire interaction graph and sensitive user information (such as gender, age, etc.) at the central server. This paper addresses the pervasive issue of inherent bias within RSs for different demographic groups without compromising the privacy of sensitive user attributes in FL environment with the graph-based model. To address the group bias, we propose F2PGNN (Fair Federated Personalized Graph Neural Network), a novel framework that leverages the power of Personalized Graph Neural Network (GNN) coupled with fairness considerations. Additionally, we use differential privacy techniques to fortify privacy protection. Experimental evaluation on three publicly available datasets showcases the efficacy of F2PGNN in mitigating group unfairness by 47% ∼ 99% compared to the state-of-the-art while preserving privacy and maintaining the utility. The results validate the significance of our framework in achieving equitable and personalized recommendations using GNN within the FL landscape. Source code is at: https://github.com/nimeshagrawal/F2PGNN-AAAI24
Nimesh Agrawal, Anuj Kumar Sirohi, Sandeep Kumar 0005, Jayadeva
AAAI3
2024 UGC: Universal Graph Coarsening
abstract
In the era of big data, graphs have emerged as a natural representation of intricate relationships. However, graph sizes often become unwieldy, leading to storage, computation, and analysis challenges. A crucial demand arises for methods that can effectively downsize large graphs while retaining vital insights. Graph coarsening seeks to simplify large graphs while maintaining the basic statistics of the graphs, such as spectral properties and $\epsilon$-similarity in the coarsened graph. This ensures that downstream processes are more efficient and effective. Most published methods are suitable for homophilic datasets, limiting their universal use. We propose **U**niversal **G**raph **C**oarsening (UGC), a framework equally suitable for homophilic and heterophilic datasets. UGC integrates node attributes and adjacency information, leveraging the dataset's heterophily factor. Results on benchmark datasets demonstrate that UGC preserves spectral similarity while coarsening. In comparison to existing methods, UGC is 4x to 15x faster, has lower eigen-error, and yields superior performance on downstream processing tasks even at 70% coarsening ratios.
Mohit Kataria, Sandeep Kumar 0005, Jayadeva
NeurIPS2
2023 Robust and Globally Sparse Pca via Majorization-Minimization and Variable Splitting
abstract
This paper addresses the problem of robust and sparse PCA. We consider a formulation combining a M-estimation type robust subspace recovery term and a mixed norm that promotes structured sparsity in the basis vectors, which is especially interesting for joint dimension reduction and variable selection. To solve it, we propose to leverage variable splitting methods, with the crucial step then lying on the Stiefel manifold. The resolution of this subproblem, involving the orthonormality constraint, is achieved through a tailored majorization-minimization (MM) step. Numerical experiments on gene expression measurements illustrate the interest of the proposal.
Hugo Brehier, Arnaud Breloy, Mohammed Nabil El Korso, Sandeep Kumar 0005
ICASSP4
2023 Estimating Normalized Graph Laplacians in Financial Markets
abstract
Gaussian Markov random fields, a class of graphical models, play an increasingly important role in real-world problems, where they are often applied to uncover conditional correlations between pairs of entities in a network. Motivated by recent applications of graphs in financial markets, we investigate the problem of learning undirected, weighted, normalized, graphical models. More precisely, we design an optimization algorithm to learn precision matrices that are modeled as normalized graph Laplacians. The proposed algorithm takes advantages of frameworks such as the alternating direction method of multipliers and projected gradient descent, which allows us to decompose the original problem into subproblems that can be solved efficiently. We demonstrate the empirical performance of the proposed algorithm, in comparison to state-of-the-art benchmark models, in a number of datasets involving financial time-series.
José Vinícius de Miranda Cardoso, Jiaxi Ying, Sandeep Kumar 0005, Daniel Pérez Palomar
ICASSP3
2023 Graph of Circuits with GNN for Exploring the Optimal Design Space
abstract
The design automation of analog circuits poses significant challenges in terms of the large design space, complex interdependencies between circuit specifications, and resource-intensive simulations. To address these challenges, this paper presents an innovative framework called the Graph of Circuits Explorer (GCX). Leveraging graph structure learning along with graph neural networks, GCX enables the creation of a surrogate model that facilitates efficient exploration of the optimal design space within a semi-supervised learning framework which reduces the need for large labelled datasets. The proposed approach comprises three key stages. First, we learn the geometric representation of circuits and enrich it with technology information to create a comprehensive feature vector. Subsequently, integrating feature-based graph learning with few-shot and zero-shot learning enhances the generalizability in predictions for unseen circuits. Finally, we introduce two algorithms namely, EASCO and ASTROG which upon integration with GCX optimize the available samples to yield the optimal circuit configuration meeting the designer's criteria. The effectiveness of the proposed approach is demonstrated through simulated performance evaluation of various circuits, using derived parameters in 180nm CMOS technology. Furthermore, the generalizability of the approach is extended to higher-order topologies and different technology nodes such as 65nm and 45nm CMOS process nodes.
Aditya Hemant Shahane, Saripilli Swapna Manjiri, Ankesh Jain, Sandeep Kumar 0005
NeurIPS4
2021 Parameter Estimation for Student's t VAR Model with Missing Data
abstract
The vector autoregressive (VAR) models provide a significant tool for multivariate time series analysis. Most existing works on VAR modeling are based on the multivariate Gaussian distribution. However, heavy-tailed distributions are suggested more reasonable for capturing the real-world phenomena, like the presence of outliers and a stronger possibility of extreme values. Furthermore, missing values in observed data is a real problem, which typically happens during the data observation or recording process. In this paper, we propose an algorithmic framework to estimate the parameters of a VAR model with heavy-tailed Student’s t distributed innovations from incomplete data based on the stochastic approximation expectation maximization (SAEM) algorithm coupled with a Markov Chain Monte Carlo (MCMC) procedure. Extensive experiments with synthetic data corroborate our claims.
Rui Zhou 0016, Sandeep Kumar 0005, Daniel Pérez Palomar
ICASSP3
2020 A Unified Framework for Structured Graph Learning via Spectral Constraints
abstract
Graph learning from data is a canonical problem that has received substantial attention in the literature. Learning a structured graph is essential for interpretability and identification of the relationships among data. In general, learning a graph with a specific structure is an NP-hard combinatorial problem and thus designing a general tractable algorithm is challenging. Some useful structured graphs include connected, sparse, multi-component, bipartite, and regular graphs. In this paper, we introduce a unified framework for structured graph learning that combines Gaussian graphical model and spectral graph theory. We propose to convert combinatorial structural constraints into spectral constraints on graph matrices and develop an optimization framework based on block majorization-minimization to solve structured graph learning problem. The proposed algorithms are provably convergent and practically amenable for a number of graph based applications such as data clustering. Extensive numerical experiments with both synthetic and real data sets illustrate the effectiveness of the proposed algorithms. An open source R package containing the code for all the experiments is available at https://CRAN.R-project.org/package=spectralGraphTopology.
Sandeep Kumar 0005, Jiaxi Ying, José Vinícius de Miranda Cardoso, Daniel Pérez Palomar
J. Mach. Learn. Res.1
2019 Structured Graph Learning Via Laplacian Spectral Constraints
abstract
Learning a graph with a specific structure is essential for interpretability and identification of the relationships among data. But structured graph learning from observed samples is an NP-hard combinatorial problem. In this paper, we first show, for a set of important graph families it is possible to convert the combinatorial constraints of structure into eigenvalue constraints of the graph Laplacian matrix. Then we introduce a unified graph learning framework lying at the integration of the spectral properties of the Laplacian matrix with Gaussian graphical modeling, which is capable of learning structures of a large class of graph families. The proposed algorithms are provably convergent and practically amenable for big-data specific tasks. Extensive numerical experiments with both synthetic and real datasets demonstrate the effectiveness of the proposed methods. An R package containing codes for all the experimental results is submitted as a supplementary file.
Sandeep Kumar 0005, Jiaxi Ying, José Vinícius de Miranda Cardoso, Daniel Pérez Palomar
NeurIPS1
2018 Parameter Estimation of Heavy-Tailed Random Walk Model from Incomplete Data
abstract
This paper proposes a novel and structured framework for parameter estimation from incomplete time series data under heavy-tailed random walk model. Traditionally, maximum likelihood estimation (MLE) for Gaussian random walk model from incomplete data has been considered. However, it is not applicable in many practical applications that follow some heavy-tailed random walk model. We first model a random walk model with Student-t residuals. Then we develop an MLE-based stochastic expectation maximization (EM) algorithm. The algorithm provides tractable E and M steps, which are easy to implement with simple updates and fast convergence. The simulation results illustrate the improved performance over the benchmarks.
Sandeep Kumar 0005, Daniel Pérez Palomar
ICASSP2