EDBT 2026 Demo / reviewers in the wild / expert
Sandipan Sikdar
dblp:139/2312
· DBLP profile ↗
18ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Why So Separate: Analyzing In-Context Learning from a Vector Space Perspective
Tobias Kalmbach, Sandipan Sikdar |
LREC | 2 |
| 2026 | Early Fusion with Contrastive Learning: A Lightweight Alternative for Multi-modal Classification
Felix Wernlein, Abhik Jana, Sandipan Sikdar |
LREC | 3 |
| 2026 | Multicalibration in Fair Link PredictionabstractIn this work, we expand the multicalibration fairness metric to graph-structured data. Machine learning models, including graph neural networks (GNNs), have been demonstrated to generate unfair results when trained on data containing biases towards certain demographic groups. Consequently, several fairness enhancing methods have been proposed to deal with such biases in the model predictions. To evaluate the performance of these methods, metrics such as demographic parity and equalized odds have been proposed, which aim to ensure equal treatment across groups. However, these metrics do not take into consideration the diverse underlying distribution of the data across the different demographic groups. Hence, fairness is only achieved at the cost of utility. Multicalibration, on the other hand, aims to generate calibrated predictions for each group, thereby providing much stronger fairness guarantees. However, multicalibration is only defined for tabular data. Consequently, in this work, we provide a novel formulation of multicalibration fairness metric for graph-structured data, specifically for the task of link prediction. We demonstrate that the existing fairness enhancing methods are unable to achieve multicalibration, which leads us to define a simple quadratic program-based postprocessing method. Experiments across several real-world datasets demonstrate the effectiveness of our method, consistently outperforming the existing approaches. Our method is model agnostic and can seamlessly adapt to scenarios with single or multiple sensitive attributes. Manjish Pal, Sandipan Sikdar, Niloy Ganguly |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Exploring Disparity-Accuracy Trade-offs in Face Recognition Systems: The Role of Datasets, Architectures, and Loss FunctionsabstractAutomated Face Recognition Systems (FRSs), developed using deep learning models, are deployed worldwide for identity verification and facial attribute analysis. The performance of these models is determined by a complex interdependence among the model architecture, optimization/loss function and datasets. Although FRSs have surpassed human-level accuracy, they continue to be disparate against certain demographics. Due to the ubiquity of applications, it is extremely important to understand the impact of the three components-- model architecture, loss function and face image dataset on the accuracy-disparity trade-off to design better, unbiased platforms. In this work, we perform an in-depth analysis of three FRSs for the task of gender prediction, with various architectural modifications resulting in ten deep-learning models coupled with four loss functions and benchmark them on seven face datasets across 266 evaluation configurations. Our results show that all three components have an individual as well as a combined impact on both accuracy and disparity. We identify that datasets have an inherent property that causes them to perform similarly across models, independent of the choice of loss functions. Moreover, the choice of dataset determines the model's perceived bias-- the same model reports bias in opposite directions for three gender-balanced datasets of ``in-the-wild'' face images of popular individuals. The facial embeddings show that the models are unable to generalize a uniform definition of what constitutes a ``female face'' as opposed to a ``male face'', due to dataset diversity. We provide recommendations to model developers on using our study as a blueprint for development and subsequent deployment. Siddharth D. Jaiswal, Sagnik Basu, Sandipan Sikdar, Animesh Mukherjee 0001 |
ICWSM | 3 |
| 2025 | Evaluating LLMs' (In)ability to Follow Prompts in QA TasksabstractWhile LLMs have achieved impressive performance across various tasks, one under-explored area is evaluating their ability to follow instructions provided in the prompt when generating responses. In the context of question-answering (QA) tasks, a crucial research gap is whether LLMs prioritize their own parametric knowledge or the context provided in the prompt when generating an answer. Ignoring prompts, even when explicitly instructed to follow them, may adversely affect performance and potentially lead to unintended consequences. Additionally, LLMs should be self-reflective (i.e., LLMs should recognize when their knowledge is inadequate) and avoid hallucinations in such scenarios. To address our research question, we propose Oedipus, an evaluation framework to evaluate LLMs' ability to follow prompts. We further note that such abilities could also be influenced by contamination (i.e., exposure to datasets during training) and parametric knowledge. Consequently, we develop a novel QA dataset with four types of contexts- correct, masked, noisy, and absurd contexts with recent questions that LLMs are unlikely to have encountered in pre-training data or corpus and cannot be answered from parametric knowledge. We evaluate eight LLMs through our proposed evaluation framework and observe that LLMs often fail to follow instructions correctly and are not self-reflective. Aparup Khatua, Tobias Kalmbach, Prasenjit Mitra 0001, Sandipan Sikdar |
SIGIR | 4 |
| 2025 | Fed-FUEL: fairness and utility enhancing agnostic federated learning frameworkabstractAbstract Federated learning (FL) is an emerging communication-efficient and collaborative learning paradigm of machine learning with privacy guarantees. As these advancements unfold, adapting FL for fairness-aware learning becomes crucial. In this context, we propose a pre-processing fairness and utility (balanced accuracy) enhancing agnostic federated learning framework (Fed-FUEL) that mitigates discrimination embedded in the non-independent identically distributed data. We contribute a novel adaptive data manipulation method that mitigates discrimination embedded in the data at client side during optimization, resulting in an optimized and fair centralized server. This pre-processing approach abstracts the model architecture from the equation, offering a significant advantage in a federated environment. This abstraction not only facilitates a broader application across diverse model architectures without necessitating modifications but also sidesteps the potential complexities and inefficiencies associated with model-specific in-processing methods. Extensive experiments with a range of publicly available datasets demonstrate that our method outperforms the competing baselines in terms of both discrimination mitigation and predictive performance. Our model effectively adapts to both statistical and causal fairness notions, as shown through our experiments. Maryam Badar, Raneen Younis, Sandipan Sikdar, Wolfgang Nejdl, Marco Fisichella |
Data Min. Knowl. Discov. | 3 |
| 2025 | Fair Link Prediction With Overlapping GroupsabstractIn this article, we introduce FairLPG, a framework for ensuring fairness for the task of link prediction in graphs withmultiplesensitive attributes. In the context of link prediction in graphs, the fairness notions of demographic parity and equalized odds try to ensure equalaverage linking probabilityandtrue positive ratesacross different demographic groups consisting of various node pairs. Existing methods for achieving fairness in link prediction only consider a single sensitive attribute, which makes them unsuited for applications where multiple sensitive attributes need to be accounted for. Additionally, considering multiple sensitive attributes in the context of link prediction leads tooverlappingandintersectionalgroups, which further complicates designing such a framework. The proposed framework FairLPG assumes that the link prediction model generates a prediction score for each node pair to form an edge, and formulates a convex optimization problem that minimizes the squared Euclidean distance between the original prediction scores and transformed scores, subject to the fairness constraints. The transformed scores are then utilized for fair link prediction. To the best of our knowledge, this work is the first to handle the case of intersectional sensitive groups in the graph setting. To demonstrate its effectiveness, we deploy FairLPG on several real-world datasets and graph neural network based link prediction models. It either outperforms or performs competitively with existing methods both in terms of fairness and prediction accuracy across all the datasets and link prediction models at the same time being computationally more efficient. Manjish Pal, Sandipan Sikdar, Niloy Ganguly |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | FairTrade: Achieving Pareto-Optimal Trade-Offs between Balanced Accuracy and Fairness in Federated LearningabstractAs Federated Learning (FL) gains prominence in distributed machine learning applications, achieving fairness without compromising predictive performance becomes paramount. The data being gathered from distributed clients in an FL environment often leads to class imbalance. In such scenarios, balanced accuracy rather than accuracy is the true representation of model performance. However, most state-of-the-art fair FL methods report accuracy as the measure of performance, which can lead to misguided interpretations of the model's effectiveness to mitigate discrimination. To the best of our knowledge, this work presents the first attempt towards achieving Pareto-optimal trade-offs between balanced accuracy and fairness in a federated environment (FairTrade). By utilizing multi-objective optimization, the framework negotiates the intricate balance between model's balanced accuracy and fairness. The framework's agnostic design adeptly accommodates both statistical and causal fairness notions, ensuring its adaptability across diverse FL contexts. We provide empirical evidence of our framework's efficacy through extensive experiments on five real-world datasets and comparisons with six baselines. The empirical results underscore the potential of our framework in improving the trade-off between fairness and balanced accuracy in FL applications. Maryam Badar, Sandipan Sikdar, Wolfgang Nejdl, Marco Fisichella |
AAAI | 2 |
| 2024 | IVP-VAE: Modeling EHR Time Series with Initial Value Problem SolversabstractContinuous-time models such as Neural ODEs and Neural Flows have shown promising results in analyzing irregularly sampled time series frequently encountered in electronic health records. Based on these models, time series are typically processed with a hybrid of an initial value problem (IVP) solver and a recurrent neural network within the variational autoencoder architecture. Sequentially solving IVPs makes such models computationally less efficient. In this paper, we propose to model time series purely with continuous processes whose state evolution can be approximated directly by IVPs. This eliminates the need for recurrent computation and enables multiple states to evolve in parallel. We further fuse the encoder and decoder with one IVP solver utilizing its invertibility, which leads to fewer parameters and faster convergence. Experiments on three real-world datasets show that the proposed method can systematically outperform its predecessors, achieve state-of-the-art results, and have significant advantages in terms of data efficiency. Jingge Xiao, Leonie Basso, Wolfgang Nejdl, Niloy Ganguly, Sandipan Sikdar |
AAAI | 5 |
| 2024 | TrustFed: Navigating Trade-offs Between Performance, Fairness, and Privacy in Federated LearningabstractAs Federated Learning (FL) gains prominence in secure machine learning applications, achieving trustworthy predictions without compromising predictive performance becomes paramount. While Differential Privacy (DP) is extensively used for its effective privacy protection, yet its application as a lossy protection method can lower the predictive performance of the machine learning model. Also, the data being gathered from distributed clients in an FL environment often leads to class imbalance making traditional accuracy measure less reflective of the true performance of prediction model. In this context, we introduce a fairness-aware FL framework (TrustFed) based on Gaussian differential privacy and Multi-Objective Optimization (MOO), which effectively protects privacy while providing fair and accurate predictions. To the best of our knowledge, this is the first attempt towards achieving Pareto-optimal trade-offs between balanced accuracy and fairness in a federated environment while safeguarding the privacy of individual clients. The framework’s flexible design adeptly accommodates both statistical parity and equal opportunity fairness notions, ensuring its applicability in various FL scenarios. We demonstrate our framework’s effectiveness through comprehensive experiments on five real-world datasets. TrustFed consistently achieves comparable performance fairness tradeoff to the state-of-the-art (SoTA) baseline models while preserving the anonymization rights of users in FL applications. Maryam Badar, Sandipan Sikdar, Wolfgang Nejdl, Marco Fisichella |
ECAI | 2 |
| 2024 | IndMask: Inductive Explanation for Multivariate Time Series Black-Box ModelsabstractIn this paper, we introduce IndMask, a framework for explaining decisions of black-box time series models. While there exists a plethora of methods for providing explanations of machine learning models, time series data requires additional considerations. One needs to consider the time aspect in the explanations as well as deal with a large number of input features. Recent work has proposed explaining a time series prediction by generating a mask over the input time series. Each entry in the mask corresponds to an importance score for each feature at each time step. However, these methods only generate instancewise explanations, which means a mask needs to be computed for each input individually, thereby making them unsuited for inductive settings, where explanations need to be generated for numerous inputs, and instancewise explanation generation is severely prohibitive. Additionally, these methods have mostly been evaluated on simple recurrent neural networks and are often only applicable to a specific downstream task. Our proposed framework IndMask addresses these issues by utilizing a parameterized model for mask generation. We also go beyond recurrent neural networks and deploy IndMask to transformer architectures, thereby genuinely demonstrating its model-agnostic nature. The effectiveness of IndMask is further demonstrated through experiments over real-world datasets and time series classification and forecasting tasks. It is also computationally efficient and can be deployed in conjunction with any time series model. Seham Nasr, Sandipan Sikdar |
ECAI | 2 |
| 2022 | Adversarial Inter-Group Link Injection Degrades the Fairness of Graph Neural NetworksabstractWe present evidence for the existence and effectiveness of adversarial attacks on graph neural networks (GNNs) that aim to degrade fairness. These attacks can disadvantage a particular subgroup of nodes in GNN-based node classification, where nodes of the underlying network have sensitive attributes, such as race or gender. We conduct qualitative and experimental analyses explaining how adversarial link injection impairs the fairness of GNN predictions. For example, an attacker can compromise the fairness of GNN-based node classification by injecting adversarial links between nodes belonging to opposite subgroups and opposite class labels. Our experiments on empirical datasets demonstrate that adversarial fairness attacks can significantly degrade the fairness of GNN predictions (attacks are effective) with a low perturbation rate (attacks are efficient) and without a significant drop in accuracy (attacks are deceptive). This work demonstrates the vulnerability of GNN models to adversarial fairness attacks. We hope our findings raise awareness about this issue in our community and lay a foundation for the future development of GNN models that are more robust to such attacks. Hussain Hussain, Sandipan Sikdar, Denis Helic, Elisabeth Lex, Markus Strohmaier, Roman Kern |
ICDM | 3 |
| 2021 | Integrated Directional Gradients: Feature Interaction Attribution for Neural NLP ModelsabstractSandipan Sikdar, Parantapa Bhattacharya, Kieran Heese. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sandipan Sikdar, Parantapa Bhattacharya, Kieran Heese |
ACL/IJCNLP (1) | 1 |
| 2020 | The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word EmbeddingsabstractWe introduce ‘POLAR’ — a framework that adds interpretability to pre-trained word embeddings via the adoption of semantic differentials. Semantic differentials are a psychometric construct for measuring the semantics of a word by analysing its position on a scale between two polar opposites (e.g., cold – hot, soft – hard). The core idea of our approach is to transform existing, pre-trained word embeddings via semantic differentials to a new “polar” space with interpretable dimensions defined by such polar opposites. Our framework also allows for selecting the most discriminative dimensions from a set of polar dimensions provided by an oracle, i.e., an external source. We demonstrate the effectiveness of our framework by deploying it to various downstream tasks, in which our interpretable word embeddings achieve a performance that is comparable to the original word embeddings. We also show that the interpretable dimensions selected by our framework align with human judgement. Together, these results demonstrate that interpretability can be added to word embeddings without compromising performance. Our work is relevant for researchers and engineers interested in interpreting pre-trained word embeddings. Binny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus Strohmaier |
WWW | 2 |
| 2019 | StRE: Self Attentive Edit Quality Prediction in WikipediaabstractWikipedia can easily be justified as a behemoth, considering the sheer volume of content that is added or removed every minute to its several projects.This creates an immense scope, in the field of natural language processing toward developing automated tools for content moderation and review.In this paper we propose Self Attentive Revision Encoder (StRE) which leverages orthographic similarity of lexical units toward predicting the quality of new edits.In contrast to existing propositions which primarily employ features like page reputation, editor activity or rule based heuristics, we utilize the textual content of the edits which, we believe contains superior signatures of their quality.More specifically, we deploy deep encoders to generate representations of the edits from its text content, which we then leverage to infer quality.We further contribute a novel dataset containing ∼ 21M revisions across 32K Wikipedia pages and demonstrate that StRE outperforms existing methods by a significant margin -at least 17% and at most 103%.Our pre-trained model achieves such result after retraining on a set as small as 20% of the edits in a wikipage.This, to the best of our knowledge, is also the first attempt towards employing deep language models to the enormous domain of automated content moderation and review in Wikipedia. Bhanu Prakash Reddy, Sandipan Sikdar, Animesh Mukherjee 0001 |
ACL (1) | 3 |
| 2018 | Using core-periphery structure to predict high centrality nodes in time-varying networks
Sandipan Sikdar, Sanjukta Bhowmick, Animesh Mukherjee 0001 |
Data Min. Knowl. Discov. | 2 |
| 2016 | Anomalies in the Peer-review System: A Case Study of the Journal of High Energy PhysicsabstractPeer-review system has long been relied upon for bringing quality research to the notice of the scientific community and also preventing flawed research from entering into the literature. The need for the peer-review system has often been debated as in numerous cases it has failed in its task and in most of these cases editors and the reviewers were thought to be responsible for not being able to correctly judge the quality of the work. This raises a question "Can the peer-review system be improved?" Since editors and reviewers are the most important pillars of a reviewing system, we in this work, attempt to address a related question - given the editing/reviewing history of the editors or reviewers "can we identify the under-performing ones?", with citations received by the edited/reviewed papers being used as proxy for quantifying performance. We term such reviewers and editors as anomalous and we believe identifying and removing them shall improve the performance of the peer-review system. Using a massive dataset of Journal of High Energy Physics (JHEP) consisting of 29k papers submitted between 1997 and 2015 with 95 editors and 4035 reviewers and their review history, we identify several factors which point to anomalous behavior of referees and editors. In fact the anomalous editors and reviewers account for 26.8% and 14.5% of the total editors and reviewers respectively and for most of these anomalous reviewers the performance degrades alarmingly over time. Sandipan Sikdar, Matteo Marsili, Niloy Ganguly, Animesh Mukherjee 0001 |
CIKM | 1 |
| 2013 | Computer science fields as ground-truth communities: their impact, rise and fallabstractStudy of community in time-varying graphs has been limited to its detection and identification across time. However, presence of time provides us with the opportunity to analyze the interaction patterns of the communities, understand how each individual community grows/shrinks, becomes important over time. This paper, for the first time, systematically studies the temporal interaction patterns of communities using a large scale citation network (directed and unweighted) of computer science. Each individual community in a citation network is naturally defined by a research field -- i.e., acting as ground-truth -- and their interactions through citations in real time can unfold the landscape of dynamic research trends in the computer science domain over the last fifty years. These interactions are quantified in terms of a metric called inwardness that captures the effect of local citations to express the degree of authoritativeness of a community (research field) at a particular time instance. Several arguments to unveil the reasons behind the temporal changes of inwardness of different communities are put forward using exhaustive statistical analysis. The measurements (importance of field) are compared with the project funding statistics of NSF and it is found that the two are in sync. We believe that this measurement study with a large real-world data is an important initial step towards understanding the dynamics of cluster-interactions in a temporal environment. Note that this paper, for the first time, systematically outlines a new avenue of research that one can practice post community detection. Tanmoy Chakraborty 0002, Sandipan Sikdar, Vihar Tammana, Niloy Ganguly, Animesh Mukherjee 0001 |
ASONAM | 2 |