Thiruvarangan Ramaraj

dblp:204/5396 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-7333-1041ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Scalable Alignment-Free Phylogenomics Using Normalized Compression Distance (NCD): A Case Study in Tomato (Solanum Lycopersicum) Whole Genomes
abstract
Phylogenomic analyses are essential for understanding evolutionary relationships and are increasingly important for crop improvement. However, traditional alignment-based methods become computationally demanding with large plant genomes, such as tomato. This paper presents a scalable and alignment-free approach for reconstructing tomato phylogeny using Normalized Compression Distance (NCD). Due to the large size of tomato genomes, we introduce a novel segmentation strategy, dividing the genomes into 70 segments. NCD is applied to each segment to generate distance matrices, which are then averaged to create a final matrix representing the whole-genome distances. A phylogenetic tree is constructed from this final matrix and validated against established relationships in the tomato clade. Our findings demonstrate the potential of NCD with segmentation as a computationally efficient approach for inferring phylogeny from large plant genomes. This method offers an alternative to traditional alignment-based methods, opening avenues for exploring evolutionary relationships in other crops with complex and substantial genomes.
Hongzhi Hu, Thiruvarangan Ramaraj, John D. Rogers
BIBM2
2025 Enhancing Pulmonary Nodule Localization Based on Latent Representations
abstract
Accurately localizing pulmonary nodules relative to other anatomical structures is crucial for disease management, guiding biopsies, and formulating effective treatment strategies. This study introduces a fully automated approach for classifying nodules detected in computed tomography (CT) images as pleural (near the pleura as$N_{p}$) or non-pleural (distant from the pleura as$D_{p}$). We propose a combination of Principal Component Analysis (PCA) and deep learning approach to determine a threshold for the optimal correlation between the original image and the latent PCA-based image reconstruction used to classify lung nodule as$N_{p}$or$D_{p}$. Applying our methodology to the Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset, we found that the PCA-based approach demonstrated substantial agreement with human evaluations of the proximity of the lung nodule to the pleura, achieving a Cohen's Kappa value of 0.76, outperforming an intensitybased baseline model with a Cohen's Kappa value of 0.55. Further analysis revealed significant variability in radiologists' semantic ratings for$N_{p}$versus$D_{p}$, with the highest variability observed in the texture feature. These findings demonstrate that our approach has the potential to enhance the classification of lung nodule localization when integrated into computer-aided diagnosis (CAD) systems.
Charmi Patel, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS4
2025 Video Label Refinement for Temporal Localization
abstract
When performing temporal localization, achieving high precision results or high accuracy at the highest temporal Intersection over Union (tIoU) requires training a model with class labels that have precise boundaries. Obtaining high-precision class labels solely from human annotators presents challenges, such as the difficulty of accurately discerning when a movement begins and ends by watching a video. This work expands on an approach improving the temporal boundaries of class labels for activity recognition in video to handle complex movement patterns in human interactions. This method applies signal processing methods to motion features to identify shapes in signals associated with activities to be classified and adjust the boundaries accurately. When these methods are applied to existing human-provided activity annotations, they can improve the accuracy of class label boundaries. This improves temporal localization model performance at the high tIoUs.
Jennifer Piane, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
ICME2
2025 Genotype-to-Phenotype Associations in Yeast with Frequented Region Variants and Deep Learning
abstract
Phenotypes are the observable characteristics of an individual organism. Predicting quantitative phenotypes from genomic variation remains challenging when causal signals span both local motifs and distal regulatory context. Building on Frequented Regions (FRs), which represent subsequences conserved across genomes, extracted from a pangenome graph, we compare six modeling strategies on five yeast growth phenotypes: Random Forest (RF) on FR counts, RF on FR sequences, 1D convolutional neural networks on FR sequences, Long Short-Term Memory (LSTM) on FR sequences, a Genome-wide Association Study (GWAS) baseline, and a sequence-based Enformer model trained on raw FR nucleotide windows. Across the five phenotypes, all sequence-based baselines improve upon RF (FR counts) and GWAS, confirming the value of sequence context. Enformer consistently outperforms CNN/LSTM on all five phenotypes and surpasses RF (FR-sequences) on three of five, while remaining competitive on the others. These results indicate that when long-range dependencies contribute to trait variation, transformer-based modeling of raw sequence windows may yield tangible gains over k-mer and local-pattern learners; conversely, for phenotypes dominated by short-range signals, lightweight baselines may remain competitive. These findings suggest that while short-range motif statistics can suffice for certain phenotypes, deep learning architectures that integrate positional context and distal interactions can yield additional gains, particularly when phenotypic variation is linked to dispersed regulatory signals.
Tejaswi Vemuri, Trung Dinh, Thiruvarangan Ramaraj, Joann Mudge, Brendan Mumey, Indika Kahanda
ICMLA4
2024 Clustering density based gene expression data in the mouse brainstem and comparison to actual neuroanatomy
abstract
The medullary reticular formation is a network of brainstem nuclei and neurons in the brain that is composed of circuit pathways from the brain to the spinal cord. It carries out vital roles and orofacial behaviors such as vocalization, swallowing, and breathing. A disruption in these neuronal circuits can be threatening for persons with neurodevelopmental and neurodegenerative disorders, leading to deficits in breathing, swallowing, and communication. However, the neurons associated with orofacial motor controls are not clearly defined or localized within the reticular formation due to the lack of molecular markers and distinct brain tissue cells. This project serves to understand the organization of the reticular formation and define where orofacial motor behaviors are localized. We hypothesize that differences in gene expression and three-dimensional proximities can identify functional subpopulations of the reticular formation that control orofacial motor behaviors. Using gene density measures of the coronal plane in the adult male mouse brain that are attributed to the reticular formation, we computed 13 clusters using K-Means clustering with Euclidean distance measure based on brainstem nuclei that are biologically significant to the control of orofacial behaviors. This clustering pattern has implications for understanding the organization of the reticular formation. By understanding the organization of these brainstem nuclei, neuroscientists can better understand how neurodegenerative disorders work and potentially lead to effective treatments for persons suffering from them.
Colin Buenvenida, Brandon Kong, Shane Jung, Kaiwen Kam, Jacob D. Furst, Thiruvarangan Ramaraj
CIBCB6
2024 Application of Machine Learning Techniques to Drive Immunological Insights Towards Malaria Prognosis Using Microarray data
abstract
This study aims to identify genes related to malaria using a comprehensive approach that leverages multiple machine learning-based feature selection techniques. Due to the complexity and resource intensity of building weighted gene co-expression network analysis (WGCNA) networks, we utilized various machine learning algorithms to streamline the process. We analyzed the GSE117613 dataset from the Gene Expression Omnibus (GEO) to identify differentially expressed genes (DEGs) between malaria-infected and control samples. Techniques such as the Least Absolute Shrinkage and Selection Operator (LASSO), Sup-port Vector Machine Recursive Feature Elimination (SVM - RFE), SVM with Radial Basis Function (RBF), ensemble model, ReliefF, and SigFeature were employed to select potential biomarkers. By intersecting the genes captured by these methods, we identified robust candidates for further analysis. Pathway analysis was then conducted to elucidate the biological significance of these genes in the context of malaria. Our results highlighted several significant genes associated with malaria, validated using pathway analysis. This study provides insights into the molecular mechanisms of malaria and demonstrates an efficient, machine learning-based approach for biomarker discovery in medical research.
Sashank Makanaboyina, Rahul Vijay, Thiruvarangan Ramaraj
ICMLA4
2023 Genotype-to-Phenotype Associations with Frequented Region Variants
abstract
A pangenome represents the entire sequence content and variation of a population. As collections of complete reference quality genomes become more common, so does the prevalence of pangenomes, necessitating the need for scalable computational methods for their analysis. Previously, we developed FindFRs for identifying Frequented Regions in pangenome graphs, where a Frequented Region is a subgraph that is frequently traversed by multiple sequences. In this work, we propose FindFRs3, which is an updated version of FindFRs capable of identifying Frequented Regions with improved runtime and memory efficiency, enabling the analysis of much larger pangenome graphs. In addition, FindFRs3 identifies Frequented Region Variants (the unique subpaths through each region). We demonstrate the utility of these variants by using them as input features for machine learning models that can predict genotype-to-phenotype associations in a large yeast pangenome. Biological insights gained from these variants show that this novel technique allows for a more nuanced and detailed analysis of larger pangenomes.
Indika Kahanda, Buwani Manuweera, Brendan Mumey, Thiruvarangan Ramaraj, Alan M. Cleary, Joann Mudge
BIBM4
2023 Development of an ontology for biofilms
abstract
Microorganisms make up most of the earth’s biomass, and most microbes exist in the form of biofilms, complex communities of microorganisms growing attached to surfaces. Biofilms are directly relevant to a large number of scientific disciplines, and are the subjects of growing multidisciplinary research. As such, there is a pressing requirement for information systems that specialize in biofilm knowledge. Realization of such systems will require a coherent approach to understanding and curating the language used to study biofilms; an ontology of biofilms-related terms offers a foundation for such systems. Here we present an ontology for the study of biofilms (BIFO), a tool that will provide precisely defined terms describing all aspects involved in the biofilms domain. We describe semi-automated methods for the identification of relevant terms from a body of literature, the selection of a set of important terms by domain experts, and the construction of the ontology. A generic approach for BIFO is presented, in which foundational biofilm-related entities and relationships are represented. This ontology reuses terms from other ontologies that provide biofilm knowledge from the Open Biological and Biomedical Ontologies (OBO) foundry.
Thiruvarangan Ramaraj, Bo Wen Liu, Britney Gibbs, Azalea Mendoza, David L. Millman, Brendan Mumey, Matthew Fields, Callum Bell
BIBM1
2023 Curriculum gDRO: Improving Lung Malignancy Classification through Robust Curriculum Task Learning
abstract
Deep learning models used in Computer-Aided Diagnosis (CAD) systems are often trained with Empirical Risk Minimization (ERM) loss. These models often achieve high overall classification accuracy but with lower classification accuracy on certain subgroups. In the context of lung nodule malignancy classification task, these atypical subgroups exist due to the lung cancer heterogeneity. In this study, we characterize lung nodule malignancy subgroups using the malignancy likelihood ratings given by radiologists and improve the worst subgroup performance by utilizing group Distributionally Robust Optimization (gDRO). However, we noticed that gDRO improves on worst subgroup performance from the benign category, which has less clinical importance than improving classification accuracy for a malignant subgroup. Therefore, we propose a novel curriculum gDRO training scheme that trains for an “easy” task (nodule malignancy is determinate or indeterminate for radiologists) first, then for a “hard” task (malignant, benign, or indeterminate nodule). Our results indicate that our approach boosts the worst group subclass accuracy from the malignant category, by up to 6 percentage points compared to standard methods that address and improve worst group classification performance.
Arun Sivakumar, Yiyang Wang 0003, Roselyne Tchoua, Thiruvarangan Ramaraj, Jacob D. Furst, Daniela Raicu
CBMS4
2022 Lung Nodule Malignancy Subtype Discovery with Semantic Learning
abstract
Computer-aided diagnosis (CAD) systems have been widely used as second readers in lung cancer diagnosis. However, lung cancer heterogeneity and lack of using human annotated semantic characteristics are significant obstacles to an accurate and explainable CAD outcome. We propose a novel CAD scheme that characterizes lung nodule malignancy subtypes semantically and classifies nodule malignancy through a semantic learning process. We built and evaluated our method on a publicly available dataset, Lung Image Database Consortium (LIDC). We discovered and characterized two malignant nodule subtypes and two benign subtypes. In addition, we achieved a significantly improved malignancy classification result by incorporating the semantic learning process when compared with a result without using the semantic learning. This study suggests that future CAD systems should include disease heterogeneity as a model input to refine the model explainability; our classification result provides case-base evidence that learning image semantic explanations is valuable for improving CAD accuracy. While we tested our approach on one dataset, the proposed approach is applicable to other CAD tasks if semantic annotations are available.
Yiyang Wang 0003, Bowen Qiu, Thiruvarangan Ramaraj, Ilyas Ustun, Jacob D. Furst, Daniela Raicu
ICPR3
2019 Exploring Frequented Regions in Pan-Genomic Graphs
abstract
We consider the problem of identifying regions within a pan-genome De Bruijn graph that are traversed by many sequence paths. We define such regions and the subpaths that traverse them as frequented regions (FRs). In this work, we formalize the FR problem and describe an efficient algorithm for finding FRs. Subsequently, we propose some applications of FRs based on machine-learning and pan-genome graph simplification. We demonstrate the effectiveness of these applications using data sets for the organisms Staphylococcus aureus (bacterium) and Saccharomyces cerevisiae (yeast). We corroborate the biological relevance of FRs such as identifying introgressions in yeast that aid in alcohol tolerance, and show that FRs are useful for classification of yeast strains by industrial use and visualizing pan-genomic space.
Alan M. Cleary, Thiruvarangan Ramaraj, Indika Kahanda, Joann Mudge, Brendan Mumey
IEEE ACM Trans. Comput. Biol. Bioinform.2