VLDB 2026 Research / reviewers in the wild / expert
Sahely Bhadra
dblp:90/2736
· DBLP profile ↗
15ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-7453-9880ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
6 papers |
Kernel, tree and ensemble methods · 36% Generative modeling · 24% Representation and self-supervised learning · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
drug discovery |
1.5 | 2 | 2025 | Efficient 3D kernels for molecular property prediction · Bioinform. 2025 CDGCN: Conditional de novo Drug Generative Model Using Graph Convolution Networks · RECOMB 2023 |
Bioinformatics and computational biology
molecular property prediction |
0.9 | 1 | 2025 | Efficient 3D kernels for molecular property prediction · Bioinform. 2025 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecular representation |
0.9 | 1 | 2025 | Efficient 3D kernels for molecular property prediction · Bioinform. 2025 |
Bioinformatics and computational biology › drug discovery
virtual screening |
0.9 | 1 | 2025 | Efficient 3D kernels for molecular property prediction · Bioinform. 2025 |
Machine learning › Generative modeling
molecular generation |
0.7 | 1 | 2023 | CDGCN: Conditional de novo Drug Generative Model Using Graph Convolution Networks · RECOMB 2023 |
Bioinformatics and computational biology › drug discovery › drug design
de novo drug design |
0.7 | 1 | 2023 | CDGCN: Conditional de novo Drug Generative Model Using Graph Convolution Networks · RECOMB 2023 |
Machine learning › Representation and self-supervised learning › multi-view learning
canonical correlation analysis |
0.4 | 1 | 2019 | Large-Scale Sparse Kernel Canonical Correlation Analysis · ICML 2019 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel machines
kernel canonical correlation analysis |
0.4 | 1 | 2019 | Large-Scale Sparse Kernel Canonical Correlation Analysis · ICML 2019 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel machines
large-scale kernel methods |
0.4 | 1 | 2019 | Large-Scale Sparse Kernel Canonical Correlation Analysis · ICML 2019 |
Bioinformatics and computational biology › systems biology
metabolic network analysis |
0.3 | 1 | 2018 | Principal metabolic flux mode analysis · Bioinform. 2018 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.3 | 2 | 2012 | Efficient methods for robust classification under uncertainty in kernel matrices · J. Mach. Learn. Res. 2012 Robust Formulations for Handling Uncertainty in Kernel Matrices · ICML 2010 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 2 | 2012 | Efficient methods for robust classification under uncertainty in kernel matrices · J. Mach. Learn. Res. 2012 Robust Formulations for Handling Uncertainty in Kernel Matrices · ICML 2010 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
markov logic networks |
0.1 | 1 | 2011 | Web information extraction using markov logic networks · KDD 2011 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning |
0.1 | 1 | 2011 | Web information extraction using markov logic networks · KDD 2011 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.1 | 1 | 2011 | Web information extraction using markov logic networks · KDD 2011 |
Machine learning › Optimization for machine learning
regularized optimization |
0.1 | 1 | 2018 | Principal metabolic flux mode analysis · Bioinform. 2018 |
Machine learning › Optimization for machine learning › sparse learning
sparse optimization |
0.1 | 1 | 2018 | Principal metabolic flux mode analysis · Bioinform. 2018 |
Mathematical optimization
continuous optimization |
0.1 | 1 | 2018 | Sparse Non-linear CCA through Hilbert-Schmidt Independence Criterion · ICDM 2018 |
Algorithms and data structures › combinatorial algorithms
maximum weight subgraph |
0.0 | 1 | 2011 | Web information extraction using markov logic networks · KDD 2011 |
Methods — techniques the papers use, named apart from their topics
graph convolutional network · 1.3conditional generative modeling · 1.3machine learning · 0.9graph kernel · 0.9stoichiometric flux analysis · 0.7regularized optimization · 0.7projected stochastic gradient · 0.7principal component analysis · 0.7nyström approximation · 0.7hilbert-schmidt independence criterion · 0.7l1 regularization · 0.4kernel function · 0.4alternating projected gradient · 0.4robust optimization · 0.3markov logic networks · 0.2maximum weight subgraph · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HEIGHTS: Hierarchical graph structure learning for time series anomaly detection
Vinitha M. Rajan, Sahely Bhadra |
Neurocomputing | 2 |
| 2025 | Efficient 3D kernels for molecular property predictionabstractAbstract Motivation This paper addresses the challenge of incorporating 3-dimensional (3D) structural information in graph kernels for machine learning-based virtual screening, a crucial task in drug discovery. Existing kernels that capture 3D information often suffer from high computational complexity, which limits their scalability. Results To overcome this, we propose the 3D chain motif graph kernel, which effectively integrates essential 3D structural properties—bond length, bond angle, and torsion angle—within the three-hop neighborhood of each atom in a molecule. In addition, we introduce a more computationally efficient variant, the 3D graph hopper kernel (3DGHK), which reduces the complexity from the state-of-the-art O(n6) (for the 3D pharmacophore kernel) to O(n2(m+log(n)+δ2+dT6)). Here, n is the number of nodes, T is the highest degree of the node, m is the number of edges, δ is the diameter of the graph, and d is the dimension of the attributes of the nodes. We conducted experiments on 21 datasets, demonstrating that 3DGHK outperforms other state-of-the-art 2D and 3D graph kernels, but it also surpasses deep learning models in classification accuracy, offering a powerful and scalable solution for virtual screening tasks. Availability and implementation Our code is publicly available at https://github.com/SantaAnkit/Efficient-3D-kernels-for-molecular-property-prediction.git. Ankit, Sahely Bhadra, Juho Rousu |
Bioinform. | 2 |
| 2025 | Warping resilient robust anomaly detection for multivariate time series
Abilasha S, Sahely Bhadra |
Mach. Learn. | 2 |
| 2023 | CDGCN: Conditional de novo Drug Generative Model Using Graph Convolution Networks
Shikha Mallick, Sahely Bhadra |
RECOMB | 2 |
| 2022 | Deep Extreme Mixture Model for Time Series ForecastingabstractTime Series Forecasting (TSF) has been a topic of extensive research, which has many real world applications such as weather prediction, stock market value prediction, traffic control etc. Many machine learning models have been developed to address TSF, yet, predicting extreme values remains a challenge to be effectively addressed. Extreme events occur rarely, but tend to cause a huge impact, which makes extreme event prediction important. Assuming light tailed distributions, such as Gaussian distribution, on time series data does not do justice to the modeling of extreme points. To tackle this issue, we develop a novel approach towards improving attention to extreme event prediction. Within our work, we model time series data distribution, as a mixture of Gaussian distribution and Generalized Pareto distribution (GPD). In particular, we develop a novel Deep eXtreme Mixture Model (DXtreMM) for univariate time series forecasting, which addresses extreme events in time series. The model consists of two modules: 1) Variational Disentangled Auto-encoder (VD-AE) based classifier and 2) Multi Layer Perceptron (MLP) based forecaster units combined with Generalized Pareto Distribution (GPD) estimators for lower and upper extreme values separately. VD-AE Classifier model predicts the possibility of occurrence of an extreme event given a time segment, and forecaster module predicts the exact value. Through extensive set of experiments on real-world datasets we have shown that our model performs well for extreme events and is comparable with the existing baseline methods for normal time step forecasting. Abilasha S, Sahely Bhadra, Ahmed Zaheer Dadarkar, Deepak P 0001 |
CIKM | 2 |
| 2022 | Warping resilient scalable anomaly detection in time series
Abilasha S, Sahely Bhadra, Deepak P 0001, Anish Mathew |
Neurocomputing | 2 |
| 2019 | Large-Scale Sparse Kernel Canonical Correlation AnalysisabstractThis paper presents gradKCCA, a large-scale sparse non-linear canonical correlation method. Like Kernel Canonical Correlation Analysis (KCCA), our method finds non-linear relations through kernel functions, but it does not rely on a kernel matrix, a known bottleneck for scaling up kernel methods. gradKCCA corresponds to solving KCCA with the additional constraint that the canonical projection directions in the kernel-induced feature space have preimages in the original data space. Firstly, this modification allows us to very efficiently maximize kernel canonical correlation through an alternating projected gradient algorithm working in the original data space. Secondly, we can control the sparsity of the projection directions by constraining the $\ell_1$ norm of the preimages of the projection directions, facilitating the interpretation of the discovered patterns, which is not available through KCCA. Our empirical experiments demonstrate that gradKCCA outperforms state-of-the-art CCA methods in terms of speed and robustness to noise both in simulated and real-world datasets. Viivi Uurtio, Sahely Bhadra, Juho Rousu |
ICML | 2 |
| 2018 | Sparse Non-linear CCA through Hilbert-Schmidt Independence CriterionabstractWe present SCCA-HSIC, a method for finding sparse non-linear multivariate relations in high-dimensional settings by maximizing the Hilbert-Schmidt Independence Criterion (HSIC). We propose efficient optimization algorithms using a projected stochastic gradient and Nyström approximation of HSIC. We demonstrate the favourable performance of SCCA-HSIC over competing methods in detecting multivariate non-linear relations both in simulation studies, with varying numbers of related variables, noise variables, and samples, as well as in real datasets. Viivi Uurtio, Sahely Bhadra, Juho Rousu |
ICDM | 2 |
| 2018 | Principal metabolic flux mode analysisabstractMotivation: In the analysis of metabolism, two distinct and complementary approaches are frequently used: Principal component analysis (PCA) and stoichiometric flux analysis. PCA is able to capture the main modes of variability in a set of experiments and does not make many prior assumptions about the data, but does not inherently take into account the flux mode structure of metabolism. Stoichiometric flux analysis methods, such as Flux Balance Analysis (FBA) and Elementary Mode Analysis, on the other hand, are able to capture the metabolic flux modes, however, they are primarily designed for the analysis of single samples at a time, and not best suited for exploratory analysis on a large sets of samples. Results: We propose a new methodology for the analysis of metabolism, called Principal Metabolic Flux Mode Analysis (PMFA), which marries the PCA and stoichiometric flux analysis approaches in an elegant regularized optimization framework. In short, the method incorporates a variance maximization objective form PCA coupled with a stoichiometric regularizer, which penalizes projections that are far from any flux modes of the network. For interpretability, we also introduce a sparse variant of PMFA that favours flux modes that contain a small number of reactions. Our experiments demonstrate the versatility and capabilities of our methodology. The proposed method can be applied to genome-scale metabolic network in efficient way as PMFA does not enumerate elementary modes. In addition, the method is more robust on out-of-steady steady-state experimental data than competing flux mode analysis approaches. Availability and implementation: Matlab software for PMFA and SPMFA and dataset used for experiments are available in https://github.com/aalto-ics-kepaco/PMFA. Supplementary information: Supplementary data are available at Bioinformatics online. Sahely Bhadra, Peter Blomberg, Sandra Castillo, Juho Rousu |
Bioinform. | 1 |
| 2017 | Multi-view kernel completion
Sahely Bhadra, Samuel Kaski, Juho Rousu |
Mach. Learn. | 1 |
| 2015 | Correction of noisy labels via mutual consistency check
Sahely Bhadra, Matthias Hein 0001 |
Neurocomputing | 1 |
| 2012 | Efficient methods for robust classification under uncertainty in kernel matrices
Aharon Ben-Tal, Sahely Bhadra, Chiranjib Bhattacharyya, Arkadi Nemirovski |
J. Mach. Learn. Res. | 2 |
| 2011 | Web information extraction using markov logic networksabstractIn this paper, we consider the problem of extracting structured data from web pages taking into account both the content of individual attributes as well as the structure of pages and sites. We use Markov Logic Networks (MLNs) to capture both content and structural features in a single unified framework, and this enables us to perform more accurate inference. MLNs allow us to model a wide range of rich structural features like proximity, precedence, alignment, and contiguity, using first-order clauses. We show that inference in our information extraction scenario reduces to solving an instance of the maximum weight subgraph problem. We develop specialized procedures for solving the maximum subgraph variants that are far more efficient than previously proposed inference methods for MLNs that solve variants of MAX-SAT. Experiments with real-life datasets demonstrate the effectiveness of our MLN-based approach compared to existing state-of-the-art extraction methods. Sandeepkumar Satpal, Sahely Bhadra, Sundararajan Sellamanickam, Rajeev Rastogi, Prithviraj Sen |
KDD | 2 |
| 2010 | Robust Formulations for Handling Uncertainty in Kernel Matrices
Sahely Bhadra, Sourangshu Bhattacharya, Chiranjib Bhattacharyya, Aharon Ben-Tal |
ICML | 1 |
| 2009 | Interval Data Classification under Partial Information: A Chance-Constraint Approach
Sahely Bhadra, Saketha Nath Jagarlapudi, Aharon Ben-Tal, Chiranjib Bhattacharyya |
PAKDD | 1 |