Apurva Kalia

dblp:91/1304 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0006-6781-0587ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
metabolomics
2.432025
JESTR: Joint Embedding Space Technique for Ranking candidate molecules for the annotation of untargeted metabolomics data · Bioinform. 2025
An Ensemble Spectral Prediction (ESP) model for metabolite annotation · Bioinform. 2024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
Bioinformatics and computational biology › metabolomics
metabolite identification
1.622025
JESTR: Joint Embedding Space Technique for Ranking candidate molecules for the annotation of untargeted metabolomics data · Bioinform. 2025
An Ensemble Spectral Prediction (ESP) model for metabolite annotation · Bioinform. 2024
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
1.022024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
An Ensemble Spectral Prediction (ESP) model for metabolite annotation · Bioinform. 2024
Bioinformatics and computational biology › structural biology
molecular structure determination
0.812024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction
0.712023
CSI: Contrastive data Stratification for Interaction prediction and its application to compound-protein interaction prediction · Bioinform. 2023

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.5spectrum simulation · 1.5molecule retrieval · 1.5de novo molecular structure generation · 1.5joint embedding · 0.9rank learning · 0.8multilayer perceptron · 0.8graph neural network · 0.8ensemble learning · 0.8multiview coding · 0.7
YearPublicationVenuePosition
2025 JESTR: Joint Embedding Space Technique for Ranking candidate molecules for the annotation of untargeted metabolomics data
abstract
MOTIVATION: A major challenge in metabolomics is annotation: assigning molecular structures to mass spectral fragmentation patterns. Despite recent advances in molecule-to-spectra and in spectra-to-molecular fingerprint (FP) prediction, annotation rates remain low. RESULTS: We introduce in this article a novel tool (JESTR) for annotation. Unlike prior approaches that "explicitly" construct molecular FPs or spectra, JESTR leverages the insight that molecules and their corresponding spectra are views of the same data and effectively embeds their representations in a joint space. Candidate structures are ranked based on cosine similarity between the embeddings of query spectrum and each candidate. We evaluate JESTR against mol-to-spec, spec-to-FP, and spec-mol matching annotation tools on four datasets. On average, for rank@[1-20], JESTR outperforms other tools by 55.5%-302.6%. We further demonstrate the strong value of regularization with candidate molecules during training, boosting rank@1 performance by 5.72% across all datasets and enhancing the model's ability to discern between target and candidate molecules. When comparing JESTR's performance against that of publicly available pretrained models of SIRIUS and CFM-ID on appropriate subsets of MassSpecGym dataset, JESTR outperforms these tools by 31% and 238%, respectively. Through JESTR, we offer a novel promising avenue toward accurate annotation, therefore unlocking valuable insights into the metabolome. AVAILABILITY AND IMPLEMENTATION: Code and dataset available at https://github.com/HassounLab/JESTR1/.
Apurva Kalia, Yan Zhou Chen, Dilip Krishnan, Soha Hassoun
Bioinform.1
2024 MassSpecGym: A benchmark for the discovery and identification of molecules
abstract
The discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure from its mass spectrum is exceptionally challenging, even when performed by human experts. As a result, the vast majority of acquired MS/MS spectra remain uninterpreted, thereby limiting our understanding of the underlying (bio)chemical processes. Despite decades of progress in machine learning applications for predicting molecular structures from MS/MS spectra, the development of new methods is severely hindered by the lack of standard datasets and evaluation protocols. To address this problem, we propose MassSpecGym -- the first comprehensive benchmark for the discovery and identification of molecules from MS/MS data. Our benchmark comprises the largest publicly available collection of high-quality MS/MS spectra and defines three MS/MS annotation challenges: \textit{de novo} molecular structure generation, molecule retrieval, and spectrum simulation. It includes new evaluation metrics and a generalization-demanding data split, therefore standardizing the MS/MS annotation tasks and rendering the problem accessible to the broad machine learning community. MassSpecGym is publicly available at \url{https://github.com/pluskal-lab/MassSpecGym}.
Roman Bushuiev, Anton Bushuiev, Niek F. de Jonge, Adamo Young, Fleming Kretschmer, Raman Samusevich, Janne Heirman, Fei Wang 0062, Luke Zhang, Kai Dührkop, Marcus Ludwig, Nils A. Haupt, Apurva Kalia, Corinna Brungs, Robin Schmid, Russell Greiner, Bo Wang 0044, David S. Wishart, Liping Liu 0001, Juho Rousu, Wout Bittremieux, Hannes L. Röst, Tytus D. Mak, Soha Hassoun, Florian Huber 0001, Justin J. J. van der Hooft, Michael A. Stravs, Sebastian Böcker, Josef Sivic, Tomás Pluskal
NeurIPS13
2024 An Ensemble Spectral Prediction (ESP) model for metabolite annotation
abstract
MOTIVATION: A key challenge in metabolomics is annotating measured spectra from a biological sample with chemical identities. Currently, only a small fraction of measurements can be assigned identities. Two complementary computational approaches have emerged to address the annotation problem: mapping candidate molecules to spectra, and mapping query spectra to molecular candidates. In essence, the candidate molecule with the spectrum that best explains the query spectrum is recommended as the target molecule. Despite candidate ranking being fundamental in both approaches, limited prior works incorporated rank learning tasks in determining the target molecule. RESULTS: We propose a novel machine learning model, Ensemble Spectral Prediction (ESP), for metabolite annotation. ESP takes advantage of prior neural network-based annotation models that utilize multilayer perceptron (MLP) networks and Graph Neural Networks (GNNs). Based on the ranking results of the MLP- and GNN-based models, ESP learns a weighting for the outputs of MLP and GNN spectral predictors to generate a spectral prediction for a query molecule. Importantly, training data is stratified by molecular formula to provide candidate sets during model training. Further, baseline MLP and GNN models are enhanced by considering peak dependencies through label mixing and multi-tasking on spectral topic distributions. When trained on the NIST 2020 dataset and evaluated on the relevant candidate sets from PubChem, ESP improves average rank by 23.7% and 37.2% over the MLP and GNN baselines, respectively, demonstrating performance gain over state-of-the-art neural network approaches. However, MLP approaches remain strong contenders when considering top five ranks. Importantly, we show that annotation performance is dependent on the training dataset, the number of molecules in the candidate set and candidate similarity to the target molecule. AVAILABILITY AND IMPLEMENTATION: The ESP code, a trained model, and a Jupyter notebook that guide users on using the ESP tool is available at https://github.com/HassounLab/ESP.
Xinmeng Li, Yan Zhou Chen, Apurva Kalia, Liping Liu 0001, Soha Hassoun
Bioinform.3
2023 CSI: Contrastive data Stratification for Interaction prediction and its application to compound-protein interaction prediction
abstract
MOTIVATION: Accurately predicting the likelihood of interaction between two objects (compound-protein sequence, user-item, author-paper, etc.) is a fundamental problem in Computer Science. Current deep-learning models rely on learning accurate representations of the interacting objects. Importantly, relationships between the interacting objects, or features of the interaction, offer an opportunity to partition the data to create multi-views of the interacting objects. The resulting congruent and non-congruent views can then be exploited via contrastive learning techniques to learn enhanced representations of the objects. RESULTS: We present a novel method, Contrastive Stratification for Interaction Prediction (CSI), to stratify (partition) a dataset in a manner that can be exploited via Contrastive Multiview Coding to learn embeddings that maximize the mutual information across congruent data views. CSI assigns a key and multiple views to each data point, where data partitions under a particular key form congruent views of the data. We showcase the effectiveness of CSI by applying it to the compound-protein sequence interaction prediction problem, a pressing problem whose solution promises to expedite drug delivery (drug-protein interaction prediction), metabolic engineering, and synthetic biology (compound-enzyme interaction prediction) applications. Comparing CSI with a baseline model that does not utilize data stratification and contrastive learning, and show gains in average precision ranging from 13.7% to 39% using compounds and sequences as keys across multiple drug-target and enzymatic datasets, and gains ranging from 16.9% to 63% using reaction features as keys across enzymatic datasets. AVAILABILITY AND IMPLEMENTATION: Code and dataset available at https://github.com/HassounLab/CSI.
Apurva Kalia, Dilip Krishnan, Soha Hassoun
Bioinform.1