Kumar Selvarajoo

dblp:25/7373 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-0314-9666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Jensen-Shannon divergence framework for quantifying gene-centric differences between matched bulk and single-cell RNA-seq breast cancer datasets
abstract
Abstract Motivation Bulk RNA sequencing (RNA-seq) captures tissue-level transcriptomes that reflect tumor-intrinsic programs and microenvironmental signals, while single-cell RNA-seq (scRNA-seq) enables analysis of cellular heterogeneity. Pseudo-bulk (PB) RNA-seq, generated by aggregating scRNA-seq profiles, has become a standard surrogate for bulk in benchmarking deconvolution, differential expression (DE), and synthetic data generation. However, it remains unclear whether PB faithfully represents true bulk transcriptomes. Methods We introduce a gene-centric Jensen–Shannon Divergence (JSD) framework, a model-free, information-theoretic approach to quantify PB–bulk differences at single-gene resolution using matched breast cancer datasets. Genes were stratified into low-JSD ‘stable proxies’ and high-JSD ‘divergent drivers,’ followed by functional enrichment and validation in scRNA-seq clusters. Results and Conclusion High-JSD genes (≈25–30% unique) drive systematic PB–bulk divergence and introduce spurious correlations, largely overlooked by standard DE or PCA analytics. Bulk-specific divergent genes are enriched for stromal and immune pathways, reflecting tumor microenvironmental signals. In contrast, PB-specific divergent genes highlighted cell-autonomous processes, including metabolism and transcriptional regulation in endothelial, T-cell, and myeloid cells. Low-JSD genes provide stable cross-platform signals, improving gene-level similarity, batch correction, and alignment. This framework disentangles modality-specific biases in cancer transcriptomics and identifies robust gene subsets for reliable bulk–single-cell integration. References 1. Jin W, Pei J, Roy JR et al. Comprehensive review on single-cell RNA sequencing: A new frontier in Alzheimer’s disease research, Ageing Res Rev 2024;100:102454. 2. Jew B, Alvarez M, Rahmani E et al. Accurate estimation of cell composition in bulk expression through robust integration of single-cell information, Nat Commun 2020;11:1971.
Kumar Selvarajoo
Briefings Bioinform.2
2025 Ten simple rules for optimal and careful use of generative AI in science
abstract
Modern AI technologies leverage natural language processing (NLP), a subfield of AI dedicated to understanding, interpreting, and generating human language for developing large language models (LLMs), which have significantly advanced the capabilities of AI systems. These models can perform complex language tasks such as text generation, summarization, translation, and sentiment analysis, with unprecedented accuracy. The two main kinds of pre-training LLMs are the BERT-like models (e.g., BioBERT, proteinBERT, and PubMedBERT used primarily for language understanding; and the GPT-like models (e.g., BioGPT and ChatGPT-4o) used primarily for language generation
Mohamed Helmy, Lingling Jin, Amr Alhossary, Tamer Mansour, Diogo Pellagrina, Kumar Selvarajoo
PLoS Comput. Biol.6
2024 Advancing drug-response prediction using multi-modal and -omics machine learning integration (MOMLIN): a case study on breast cancer clinical data
abstract
The inherent heterogeneity of cancer contributes to highly variable responses to any anticancer treatments. This underscores the need to first identify precise biomarkers through complex multi-omics datasets that are now available. Although much research has focused on this aspect, identifying biomarkers associated with distinct drug responders still remains a major challenge. Here, we develop MOMLIN, a multi-modal and -omics machine learning integration framework, to enhance drug-response prediction. MOMLIN jointly utilizes sparse correlation algorithms and class-specific feature selection algorithms, which identifies multi-modal and -omics-associated interpretable components. MOMLIN was applied to 147 patients' breast cancer datasets (clinical, mutation, gene expression, tumor microenvironment cells and molecular pathways) to analyze drug-response class predictions for non-responders and variable responders. Notably, MOMLIN achieves an average AUC of 0.989, which is at least 10% greater when compared with current state-of-the-art (data integration analysis for biomarker discovery using latent components, multi-omics factor analysis, sparse canonical correlation analysis). Moreover, MOMLIN not only detects known individual biomarkers such as genes at mutation/expression level, most importantly, it correlates multi-modal and -omics network biomarkers for each response class. For example, an interaction between ER-negative-HMCN1-COL5A1 mutations-FBXO2-CSF3R expression-CD8 emerge as a multimodal biomarker for responders, potentially affecting antimicrobial peptides and FLT3 signaling pathways. In contrast, for resistance cases, a distinct combination of lymph node-TP53 mutation-PON3-ENSG00000261116 lncRNA expression-HLA-E-T-cell exclusions emerged as multimodal biomarkers, possibly impacting neurotransmitter release cycle pathway. MOMLIN, therefore, is expected advance precision medicine, such as to detect context-specific multi-omics network biomarkers and better predict drug-response classifications.
Kumar Selvarajoo
Briefings Bioinform.2
2024 Towards multi-omics synthetic data integration
abstract
Across many scientific disciplines, the development of computational models and algorithms for generating artificial or synthetic data is gaining momentum. In biology, there is a great opportunity to explore this further as more and more big data at multi-omics level are generated recently. In this opinion, we discuss the latest trends in biological applications based on process-driven and data-driven aspects. Moving ahead, we believe these methodologies can help shape novel multi-omics-scale cellular inferences.
Kumar Selvarajoo, Sebastian Maurer-Stroh
Briefings Bioinform.1
2022 Machine learning alternative to systems biology should not solely depend on data
abstract
In recent years, artificial intelligence (AI)/machine learning has emerged as a plausible alternative to systems biology for the elucidation of biological phenomena and in attaining specified design objective in synthetic biology. Although considered highly disruptive with numerous notable successes so far, we seek to bring attention to both the fundamental and practical pitfalls of their usage, especially in illuminating emergent behaviors from chaotic or stochastic systems in biology. Without deliberating on their suitability and the required data qualities and pre-processing approaches beforehand, the research and development community could experience similar 'AI winters' that had plagued other fields. Instead, we anticipate the integration or combination of the two approaches, where appropriate, moving forward.
Hock Chuan Yeo, Kumar Selvarajoo
Briefings Bioinform.2