VLDB 2026 Research / reviewers in the wild / expert
Saptarshi Bej
dblp:247/5883
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-1835-6139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multivariate Functional Linear Discriminant Analysis for Partially-Observed Time Series (Abstract Reprint)abstractThe more extensive access to time-series data, especially for biomedical purposes, raises new methodological challenges, particularly regarding missing values. Functional linear discriminant analysis (FLDA) extends Linear Discriminant Analysis (LDA)-mediated multiclass classification and dimension reduction to data in the form of fragmented observations of a univariate function. For large multivariate and partially-observed data, there are two challenges: (i) statistical dependencies between different components of a multivariate function and (ii) heterogeneous sampling times with missing features. We here develop a multivariate version of FLDA, called MUDRA, to tackle these challenges and describe a computationally efficient expectation/conditional-maximisation (ECM) algorithm to infer its parameters without any tensor inversions. We assess its predictive power on the “Articulary Words” dataset and show its improvement over the state-of-the-art, especially in the case of missing data. This advancement in dimension reduction of multivariate functional data holds promise for enhancing classification accuracy in scenarios like partially observed short multivariate time series analysis. Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer |
AAAI | 4 |
| 2026 | Anomaly detection via mean shift density enhancement
Pritam Kar, Rahul Bordoloi, Olaf Wolkenhauer, Saptarshi Bej |
Data Min. Knowl. Discov. | 4 |
| 2026 | Convex space learning for tabular synthetic data generation
Manjunath Mahendra, Chaithra Umesh, Kristian Schultz, Olaf Wolkenhauer, Saptarshi Bej |
Neurocomputing | 5 |
| 2026 | Correction: Multivariate functional linear discriminant analysis for partially‑observed time seriesabstractIn the original publication of this article, the legend for Fig. 3 was inadvertently omitted in the published version.Specifically, the right-hand plot, depicting the performance of MUDRA on a synthetic dataset in terms of F1 score relative to the baseline ROCKET, was missing the legend identifying the respective models.For completeness and transparency, the incorrect and correct versions of Fig. 3 are presented with this correction article.The original article has been corrected.The original article can be found online at h t t p s : / / d o i . o r g / 1 0 . 1 0 0 7 / s 1 0 9 9 4 -0 2 5 -0 6 7 4 1 -0 . Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer |
Mach. Learn. | 4 |
| 2026 | FUSE: Fast Semi-Supervised Node Embedding Learning via Structural and Label-Aware Optimization
Sujan Chakraborty, Rahul Bordoloi, Anindya Sengupta, Olaf Wolkenhauer, Saptarshi Bej |
Mach. Learn. | 5 |
| 2026 | Fast agreement-driven device-calibrated local learning paradigms for spiking neural networks
Saptarshi Bej, Muhammed Sahad E, Gouri Lakshmi, Pritam Kar, Bikas C. Das |
Neural Networks | 1 |
| 2026 | Dependency-aware synthetic tabular data generationabstractSynthetic tabular data is increasingly used in privacy-sensitive domains such as healthcare, but existing generative models often fail to preserve inter-attribute relationships. In particular, functional dependencies (FDs) and logical dependencies (LDs), which capture deterministic and rule-based associations between features, are rarely or often poorly retained in synthetic datasets. To address this research gap, we propose the Hierarchical Feature Generation Framework (HFGF) for synthetic tabular data generation. We created benchmark datasets with known dependencies to evaluate our proposed HFGF. The framework first generates independent features using any standard generative model, and then reconstructs dependent features based on predefined FD and LD rules. Our experiments on four benchmark datasets and three publicly available real-world datasets with varying sizes, feature imbalance, and dependency complexity demonstrate that HFGF improves the preservation of FDs and LDs across six generative models, including CTGAN , TVAE , and GReaT . Utility analysis and qualitative dependency visualizations further show that HFGF significantly enhances the structural fidelity and utility of synthetic tabular data. 1 Chaithra Umesh, Kristian Schultz, Manjunath Mahendra, Saptarshi Bej, Olaf Wolkenhauer |
Pattern Recognit. | 4 |
| 2025 | Multivariate functional linear discriminant analysis for partially-observed time seriesabstractAbstract The more extensive access to time-series data, especially for biomedical purposes, raises new methodological challenges, particularly regarding missing values. Functional linear discriminant analysis (FLDA) extends Linear Discriminant Analysis (LDA)-mediated multiclass classification and dimension reduction to data in the form of fragmented observations of a univariate function. For large multivariate and partially-observed data, there are two challenges: (i) statistical dependencies between different components of a multivariate function and (ii) heterogeneous sampling times with missing features. We here develop a multivariate version of FLDA, called MUDRA, to tackle these challenges and describe a computationally efficient expectation/conditional-maximisation (ECM) algorithm to infer its parameters without any tensor inversions. We assess its predictive power on the “Articulary Words” dataset and show its improvement over the state-of-the-art, especially in the case of missing data. This advancement in dimension reduction of multivariate functional data holds promise for enhancing classification accuracy in scenarios like partially observed short multivariate time series analysis. Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer |
Mach. Learn. | 4 |
| 2025 | Preserving logical and functional dependencies in synthetic tabular dataabstractDependencies among attributes are a common aspect of tabular data. However, whether existing tabular data generation algorithms preserve these dependencies while generating synthetic data is yet to be explored. In addition to the existing notion of functional dependencies, we introduce the notion of logical dependencies among the attributes in this article. Moreover, we provide a measure to quantify logical dependencies among attributes in tabular data. Utilizing this measure, we compare several state-of-the-art synthetic data generation algorithms and test their capability to preserve logical and functional dependencies on several publicly available datasets. We demonstrate that currently available synthetic tabular data generation algorithms do not fully preserve functional dependencies when they generate synthetic datasets. In addition, we also showed that some tabular synthetic data generation models can preserve inter-attribute logical dependencies. Our review and comparison of the state-of-the-art reveal research needs and opportunities to develop task-specific synthetic tabular data generation models. • We introduced the notion of logical dependencies in tabular data. • We introduce a novel Bayesian measure to identify logical dependencies. • State-of-the-art models are not able to preserve functional dependencies. • Our results show that state-of-the-art algorithms can preserve logical dependencies. Chaithra Umesh, Kristian Schultz, Manjunath Mahendra, Saptarshi Bej, Olaf Wolkenhauer |
Pattern Recognit. | 4 |
| 2024 | ConvGeN: A convex space learning approach for deep-generative oversampling and imbalanced classification of small tabular datasetsabstractOversampling is commonly used to improve classifier performance for small tabular imbalanced datasets. State-of-the-art linear interpolation approaches can be used to generate synthetic samples from the convex space of the minority class. Generative networks are common deep learning approaches for synthetic sample generation. However, their scope on synthetic tabular data generation in the context of imbalanced classification is not adequately explored. In this article, we show that existing deep generative models perform poorly compared to linear interpolation-based approaches for imbalanced classification problems on small tabular datasets. To overcome this, we propose a deep generative model, ConvGeN that combines the idea of convex space learning with deep generative models . ConvGeN learns coefficients for the convex combinations of the minority class samples, such that the synthetic data is distinct enough from the majority class . Our benchmarking experiments demonstrate that our proposed model ConvGeN improves imbalanced classification on such small datasets, as compared to existing deep generative models, while being on par with the existing linear interpolation approaches. Moreover, we discuss how our model can be used for synthetic tabular data generation in general, even outside the scope of data imbalance, and thus improves the overall applicability of convex space learning. Kristian Schultz, Saptarshi Bej, Waldemar Hahn, Markus Wolfien, Prashant Srivastava, Olaf Wolkenhauer |
Pattern Recognit. | 2 |
| 2021 | Combining uniform manifold approximation with localized affine shadowsampling improves classification of imbalanced datasetsabstractOversampling approaches are a popular choice to improve classification on imbalanced datasets. The SMOTE algorithm is the pioneer for many algorithms, built as extensions of SMOTE, to solve its problem of over-generalization of the minority class. Some extensions adopt the approach of learning the minority class data distribution through clustering and manifold learning techniques. The Localised Random Affine Shadowsampling (LoRAS) algorithm, models the convex space, controlling the local variance of a synthetic sample by constructing them from convex combinations of multiple shadow samples generated by adding Gaussian noise to the original minority samples. LoRAS also uses t-SNE for a manifold learning step to identify minority class data neighbourhoods. The algorithm is known to outperform some early SMOTE extensions, improving F1-Score and Balanced accuracy for highly imbalanced classification problems. However, the state-of-the-art manifold learning algorithm UMAP is known to preserve the local and global structure of the latent data manifold better than t-SNE and is considerably faster. We have integrated the UMAP for manifold learning with localized affine shadowsampling, to build the LoRAS-UMAP algorithm. We have benchmarked the new algorithm LoRAS-UMAP against some state-of-the-art oversampling algorithms on 14 publicly available datasets characterized by high imbalance, high dimensionality, and high absolute imbalance. In summary, we incorporated UMAP for the manifold learning step yielding better F1-Score, Balanced accuracy and runtime for the LoRAS algorithm in comparison to t-SNE for manifold learning, particularly in the case of high-dimensional datasets. Saptarshi Bej, Prashant Srivastava, Markus Wolfien, Olaf Wolkenhauer |
IJCNN | 1 |
| 2021 | Automated annotation of rare-cell types from single-cell RNA-sequencing data through synthetic oversamplingabstractBACKGROUND: The research landscape of single-cell and single-nuclei RNA-sequencing is evolving rapidly. In particular, the area for the detection of rare cells was highly facilitated by this technology. However, an automated, unbiased, and accurate annotation of rare subpopulations is challenging. Once rare cells are identified in one dataset, it is usually necessary to generate further specific datasets to enrich the analysis (e.g., with samples from other tissues). From a machine learning perspective, the challenge arises from the fact that rare-cell subpopulations constitute an imbalanced classification problem. We here introduce a Machine Learning (ML)-based oversampling method that uses gene expression counts of already identified rare cells as an input to generate synthetic cells to then identify similar (rare) cells in other publicly available experiments. We utilize single-cell synthetic oversampling (sc-SynO), which is based on the Localized Random Affine Shadowsampling (LoRAS) algorithm. The algorithm corrects for the overall imbalance ratio of the minority and majority class. RESULTS: We demonstrate the effectiveness of our method for three independent use cases, each consisting of already published datasets. The first use case identifies cardiac glial cells in snRNA-Seq data (17 nuclei out of 8635). This use case was designed to take a larger imbalance ratio (~1 to 500) into account and only uses single-nuclei data. The second use case was designed to jointly use snRNA-Seq data and scRNA-Seq on a lower imbalance ratio (~1 to 26) for the training step to likewise investigate the potential of the algorithm to consider both single-cell capture procedures and the impact of "less" rare-cell types. The third dataset refers to the murine data of the Allen Brain Atlas, including more than 1 million cells. For validation purposes only, all datasets have also been analyzed traditionally using common data analysis approaches, such as the Seurat workflow. CONCLUSIONS: In comparison to baseline testing without oversampling, our approach identifies rare-cells with a robust precision-recall balance, including a high accuracy and low false positive detection rate. A practical benefit of our algorithm is that it can be readily implemented in other and existing workflows. The code basis in R and Python is publicly available at FairdomHub, as well as GitHub, and can easily be transferred to identify other rare-cell types. Saptarshi Bej, Anne-Marie Galow, Robert David 0003, Markus Wolfien, Olaf Wolkenhauer |
BMC Bioinform. | 1 |
| 2021 | LoRAS: an oversampling approach for imbalanced datasetsabstractAbstract The Synthetic Minority Oversampling TEchnique (SMOTE) is widely-used for the analysis of imbalanced datasets. It is known that SMOTE frequently over-generalizes the minority class, leading to misclassifications for the majority class, and effecting the overall balance of the model. In this article, we present an approach that overcomes this limitation of SMOTE, employing Localized Random Affine Shadowsampling (LoRAS) to oversample from an approximated data manifold of the minority class. We benchmarked our algorithm with 14 publicly available imbalanced datasets using three different Machine Learning (ML) algorithms and compared the performance of LoRAS, SMOTE and several SMOTE extensions that share the concept of using convex combinations of minority class data points for oversampling with LoRAS. We observed that LoRAS, on average generates better ML models in terms of F1-Score and Balanced accuracy. Another key observation is that while most of the extensions of SMOTE we have tested, improve the F1-Score with respect to SMOTE on an average, they compromise on the Balanced accuracy of a classification model. LoRAS on the contrary, improves both F1 Score and the Balanced accuracy thus produces better classification models. Moreover, to explain the success of the algorithm, we have constructed a mathematical framework to prove that LoRAS oversampling technique provides a better estimate for the mean of the underlying local data distribution of the minority class data space. Saptarshi Bej, Narek Davtyan, Markus Wolfien, Mariam Nassar, Olaf Wolkenhauer |
Mach. Learn. | 1 |