EDBT 2026 Demo / reviewers in the wild / expert
Vaibhav Rajan
dblp:55/406
· DBLP profile ↗
38ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-6748-6864ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Segmentation-Aware Latent Diffusion for Satellite Image Super-Resolution: Enabling Smallholder Farm Boundary DelineationabstractDelineating farm boundaries through segmentation of satellite images is a fundamental step in many agricultural applications. The task is particularly challenging for smallholder farms, where accurate delineation requires the use of high resolution (HR) imagery which are available only at low revisit frequencies (e.g., annually). To support more frequent (sub-) seasonal monitoring, HR images could be combined as references (ref) with low resolution (LR) images – having higher revisit frequency (e.g., weekly) – using reference-based super-resolution (Ref-SR) methods. However, current Ref-SR methods optimize perceptual quality and smooth over crucial features needed for downstream tasks, and are unable to meet the large scale-factor requirements for this task. Further, previous two-step approaches of SR followed by segmentation do not effectively utilize diverse satellite sources as inputs. We address these problems through a new approach, SEED-SR, which uses a combination of conditional latent diffusion models and large-scale multi-spectral, multi-source geo-spatial foundation models. Our key innovation is to bypass the explicit SR task in the pixel space and instead perform SR in a segmentation-aware latent space. This unique approach enables us to generate segmentation maps at an unprecedented 20× scale factor, and rigorous experiments on two large, real datasets demonstrate up to 25.5% and 12.9% relative improvement in instance and semantic segmentation metrics respectively over approaches based on state-of-the-art Ref-SR methods. Aditi Agarwal, Anjali Jain, Nikita Saxena, Ishan Deshpande, Michal Kazmierski, Abigail Annkah, Nadav Sherman, Karthikeyan Shanmugam 0001, Alok Talekar, Vaibhav Rajan |
WACV | 10 |
| 2026 | X-JEPA: A Novel Joint Learning Cross-Modal Predictive Alignment Framework for Remote Sensing Image RetrievalabstractThe growing scale and heterogeneity of remote sensing (RS) imagery demand robust, scalable frameworks for content-based image retrieval across sensor modalities. We introduce X-JEPA, a novel predictive self-supervised architecture explicitly designed for cross-modal remote sensing image retrieval (RS-CMIR), and the first to extend joint embedding predictive paradigms beyond unimodal domains. Unlike prior contrastive or reconstruction-based methods, X-JEPA formulates representation learning as a latent forecasting task: predicting the semantic embedding of a target modality given context from another. To enforce modality-invariant alignment, we propose a geometry-aware Prediction Space Alignment (PSA) loss, which captures the structure of the latent space without requiring pixel-level reconstruction or modality pairing. We evaluate X-JEPA on two large-scale benchmarks—BEN-14K (Sentinel-1/Sentinel-2) and fMoW (RGB/Sentinel) across both unimodal and cross-modal retrieval tasks. X-JEPA consistently outperforms state-of-the-art self-supervised baselines, including MAE, SatMAE, CrossMAE, CSMAE-SESD, CROMA, SkySense, DeCUR, and REJEPA, achieving up to 11.0% F1-score improvement in cross-modal retrieval and 9.8% in unimodal settings. Despite its high retrieval accuracy, the model remains lightweight, requiring fewer parameters and yielding 8–10% F1-score gains on average, establishing a new state-of-the-art for scalable, sensor-agnostic RS-CMIR.1 Shabnam Choudhury, Yash Salunkhe, Vaibhav Rajan, Subhasis Chaudhuri, Biplab Banerjee |
WACV | 3 |
| 2025 | GANDALF: Generative AttentioN based Data Augmentation and predictive modeLing Framework for personalized cancer treatmentabstractEffective treatment of cancer is a major challenge faced by healthcare providers, due to the highly individualized nature of patient responses to treatment. This is caused by the heterogeneity seen in cancer-causing alterations (mutations) across patient genomes. Limited availability of response data in patients makes it difficult to train personalized treatment recommendation models on mutations from clinical genomic sequencing reports. Prior methods tackle this by utilising larger, labelled pre-clinical laboratory datasets (‘cell lines’), via transfer learning. These methods augment patient data by learning a shared, domain-invariant representation, between the cell line and patient domains, which is then used to train a downstream drug response prediction (DRP) model. This approach augments data in the shared space but fails to model patient-specific characteristics, which have a strong influence on their drug response. We propose a novel generative attention-based data augmentation and predictive modeling framework, GANDALF, to tackle this crucial shortcoming of prior methods. GANDALF not only augments patient genomic data directly, but also accounts for its domain-specific characteristics. GANDALF outperforms state-of-the-art DRP models on publicly available patient datasets and emerges as the front-runner amongst SOTA cancer DRP models. Aishwarya Jayagopal, Yanrong Zhang, Robert J. Walsh, Tuan Zea Tan, Anand Jeyasekharan, Vaibhav Rajan |
ICLR | 6 |
| 2025 | A Joint LLM-KG System for Disease Q&AabstractMedical question answer (QA) assistants respond to lay users' health-related queries by synthesizing information from multiple sources using natural language processing and related techniques. They can serve as vital tools to alleviate issues of misinformation, information overload, and complexity of medical language, thus addressing lay users' information needs while reducing the burden on healthcare professionals. QA systems, the engines of such assistants, have often used large language models (LLMs) or knowledge graphs (KG), though the approaches could be complementary. LLM-based QA systems excel at understanding complex questions and providing well-formed answers but are prone to factual mistakes. KG-based QA systems, which represent facts well, are mostly limited to answering short-answer questions with pre-created templates. While a few studies have used both LLM and KG for text-based QA, the approaches are still prone to incomplete or inaccurate answers. Extant QA systems also have limitations in terms of automation and performance. We address these challenges by designing a novel, automated disease QA system named Disease Guru-Long-Form Question Answer (DG-LFQA), which effectively utilizes both LLM and KG techniques through a joint reasoning approach to answer disease-related questions appropriate for lay users. Our evaluation of the system using a range of quality metrics demonstrates its efficacy over related baseline systems. Prakash Chandra Sukhwal, Vaibhav Rajan, Atreyi Kankanhalli |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Encoding Unitig-level Assembly Graphs with Heterophilous Constraints for Metagenomic Contigs BinningabstractMetagenomics studies genomic material derived from mixed microbial communities in diverse environments, holding considerable significance for both human health and environmental sustainability. Metagenomic binning refers to the clustering of genomic subsequences obtained from high-throughput DNA sequencing into distinct bins, each representing a constituent organism within the community. Mainstream binning methods primarily rely on sequence features such as composition and abundance, making them unable to effectively handle sequences shorter than 1,000 bp and inherent noise within sequences. Several binning tools have emerged, aiming to enhance binning outcomes by using the assembly graph generated by assemblers, which encodes valuable overlapping information among genomic sequences. However, existing assembly graph-based binners mainly focus on simplified contig-level assembly graphs that are recreated from assembler’s original graphs, unitig-level assembly graphs. The simplification reduces the resolution of the connectivity information in original graphs. In this paper, we design a novel binning tool named UnitigBin, which leverages representation learning on unitig-level assembly graphs while adhering to heterophilious constraints imposed by single-copy marker genes, ensuring that constrained contigs cannot be grouped together. Extensive experiments conducted on synthetic and real datasets demonstrate that UnitigBin significantly surpasses state-of-the-art binning tools. Hansheng Xue, Vijini Mallawaarachchi, Lexing Xie, Vaibhav Rajan |
ICLR | 4 |
| 2024 | WISER: Weak Supervision and Supervised Representation Learning to Improve Drug Response Prediction in CancerabstractCancer, a leading cause of death globally, occurs due to genomic changes and manifests heterogeneously across patients. To advance research on personalized treatment strategies, the effectiveness of various drugs on cells derived from cancers (’cell lines’) is experimentally determined in laboratory settings. Nevertheless, variations in the distribution of genomic data and drug responses between cell lines and humans arise due to biological and environmental differences. Moreover, while genomic profiles of many cancer patients are readily available, the scarcity of corresponding drug response data limits the ability to train machine learning models that can predict drug response in patients effectively. Recent cancer drug response prediction methods have largely followed the paradigm of unsupervised domain-invariant representation learning followed by a downstream drug response classification step. Introducing supervision in both stages is challenging due to heterogeneous patient response to drugs and limited drug response data. This paper addresses these challenges through a novel representation learning method in the first phase and weak supervision in the second. Experimental results on real patient data demonstrate the efficacy of our method WISER (Weak supervISion and supErvised Representation learning) over state-of-the-art alternatives on predicting personalized drug response. Our implementation is available at https://github.com/kyrs/WISER Kumar Shubham, Aishwarya Jayagopal, Syed Mohammed Danish, Prathosh A. P., Vaibhav Rajan |
ICML | 5 |
| 2024 | Personalised Drug Identifier for Cancer Treatment with Transformers using Auxiliary InformationabstractCancer remains a global challenge due to its growing clinical and economic burden. Its uniquely personal manifestation, which makes treatment difficult, has fuelled the quest for personalized treatment strategies. Thus, genomic profiling is increasingly becoming part of clinical diagnostic panels. Effective use of such panels requires accurate drug response prediction (DRP) models, which are challenging to build due to limited labelled patient data. Previous methods to address this problem have used various forms of transfer learning. However, they do not explicitly model the variable length sequential structure of the list of mutations in such diagnostic panels. Further, they do not utilize auxiliary information (like patient survival) for model training. We address these limitations through a novel transformer-based method, which surpasses the performance of state-of-the-art DRP models on benchmark data. Code for our method is available at https://github.com/CDAL-SOC/PREDICT-AI. Aishwarya Jayagopal, Hansheng Xue, Robert J. Walsh, Krishna Kumar Hariprasannan, David Shao Peng Tan, Tuan Zea Tan, Jason J. Pitt, Anand Jeyasekharan, Vaibhav Rajan |
KDD | 10 |
| 2024 | Evaluating Explanations From AI Algorithms for Clinical Decision-Making: A Social Science-Based ApproachabstractExplainable Artificial Intelligence (XAI) techniques generate explanations for predictions from AI models. These explanations can be evaluated for (i) faithfulness to the prediction, i.e., its correctness about the reasons for prediction, and (ii) usefulness to the user. While there are metrics to evaluate faithfulness, to our knowledge, there are no automated metrics to evaluate the usefulness of explanations in the clinical context. Our objective is to develop a new metric to evaluate usefulness of AI explanations to clinicians. Usefulness evaluation needs to consider both (a) how humans generally process explanations and (b) clinicians' specific requirements from explanations presented by clinical decision support systems (CDSS). Our new scoring method can evaluate the usefulness of explanations generated by any XAI method that provides importance values for the input features of the prediction model. Our method draws on theories from social science to gauge usefulness, and uses literature-derived biomedical knowledge graphs to quantify support for the explanations from clinical literature. We evaluate our method in a case study on predicting onset of sepsis in intensive care units. Our analysis shows that the scores obtained using our method corroborate with independent evidence from clinical literature and have the required qualities expected from such a metric. Thus, our method can be used to evaluate and select useful explanations from a diverse set of XAI techniques in clinical contexts, making it a fundamental tool for future research in the design of AI-driven CDSS. Suparna Ghanvatkar, Vaibhav Rajan |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | ASTER: A Method to Predict Clinically Relevant Synthetic Lethal Genetic InteractionsabstractA Synthetic Lethal (SL) interaction is a functional relationship between two genes or functional entities where the loss of either entity is viable but the loss of both is lethal. Such pairs can be used to develop targeted anticancer therapies with fewer side effects and reduced overtreatment. However, finding clinically relevant SL interactions remains challenging. Leveraging unified gene expression data of both disease-free and cancerous samples, we design a new technique based on statistical hypothesis testing, called ASTER, to identify SL pairs. We empirically find that the patterns of mutually exclusivity ASTER finds using genomic and transcriptomic data provides a strong signal of synthetic lethality. For large-scale multiple hypothesis testing, we develop an extension called ASTER++ that can utilize additional input gene features within the hypothesis testing framework. Our computational and functional experiments demonstrate the efficacy of ASTER in identifying SL pairs with potential therapeutic benefits. Herty Liany, Aishwarya Jayagopal, Dachuan Huang, Jing Quan Lim, Nur Izzah NBH, Anand Jeyasekharan, Choon Kiat Ong, Vaibhav Rajan |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | ExpertNet: A Deep Learning Approach to Combined Risk Modeling and Subtyping in Intensive Care UnitsabstractRisk models play a crucial role in disease prevention, particularly in intensive care units (ICUs). Diseases often have complex manifestations with heterogeneous subpopulations, or subtypes, that exhibit distinct clinical characteristics. Risk models that explicitly model subtypes have high predictive accuracy and facilitate subtype-specific personalization. Such models combine clustering and classification methods but do not effectively utilize the inferred subtypes in risk modeling. Their limitations include tendency to obtain degenerate clusters and cluster-specific data scarcity leading to insufficient training data for the corresponding classifier. In this article, we develop a new deep learning model for simultaneous clustering and classification, ExpertNet, with novel loss terms and network training strategies that address these limitations. The performance of ExpertNet is evaluated on the tasks of predicting risk of (i) sepsis and (ii) acute respiratory distress syndrome (ARDS), using two large electronic medical records datasets from ICUs. Our extensive experiments show that, in comparison to state-of-the-art baselines for combined clustering and classification, ExpertNet achieves superior accuracy in risk prediction for both ARDS and sepsis; and comparable clustering performance. Visual analysis of the clusters further demonstrates that the clusters obtained are clinically meaningful and a knowledge-distilled model shows significant differences in risk factors across the subtypes. By addressing technical challenges in training neural networks for simultaneous clustering and classification, ExpertNet lays the algorithmic foundation for the future development of subtype-aware risk models. Shivin Srivastava, Vaibhav Rajan |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | RepBin: Constraint-Based Graph Representation Learning for Metagenomic BinningabstractMixed communities of organisms are found in many environments -- from the human gut to marine ecosystems -- and can have profound impact on human health and the environment. Metagenomics studies the genomic material of such communities through high-throughput sequencing that yields DNA subsequences for subsequent analysis. A fundamental problem in the standard workflow, called binning, is to discover clusters, of genomic subsequences, associated with the constituent organisms. Inherent noise in the subsequences, various biological constraints that need to be imposed on them and the skewed cluster size distribution exacerbate the difficulty of this unsupervised learning problem. In this paper, we present a new formulation using a graph where the nodes are subsequences and edges represent homophily information. In addition, we model biological constraints providing heterophilous signal about nodes that cannot be clustered together. We solve the binning problem by developing new algorithms for (i) graph representation learning that preserves both homophily relations and heterophily constraints (ii) constraint-based graph clustering method that addresses the problems of skewed cluster size distribution. Extensive experiments, on real and synthetic datasets, demonstrate that our approach, called RepBin, outperforms a wide variety of competing methods. Our constraint-based graph representation learning and clustering methods, that may be useful in other domains as well, advance the state-of-the-art in both metagenomics binning and graph representation learning. Hansheng Xue, Vijini Mallawaarachchi, Vaibhav Rajan, Yu Lin 0001 |
AAAI | 4 |
| 2022 | Graph Coloring via Neural Networks for Haplotype Assembly and Viral Quasispecies ReconstructionabstractUnderstanding genetic variation, e.g., through mutations, in organisms is crucial to unravel their effects on the environment and human health. A fundamental characterization can be obtained by solving the haplotype assembly problem, which yields the variation across multiple copies of chromosomes. Variations among fast evolving viruses that lead to different strains (called quasispecies) are also deciphered with similar approaches. In both these cases, high-throughput sequencing technologies that provide oversampled mixtures of large noisy fragments (reads) of genomes, are used to infer constituent components (haplotypes or quasispecies). The problem is harder for polyploid species where there are more than two copies of chromosomes. State-of-the-art neural approaches to solve this NP-hard problem do not adequately model relations among the reads that are important for deconvolving the input signal. We address this problem by developing a new method, called NeurHap, that combines graph representation learning with combinatorial optimization. Our experiments demonstrate the substantially better performance of NeurHap in real and synthetic datasets compared to competing approaches. Hansheng Xue, Vaibhav Rajan, Yu Lin 0001 |
NeurIPS | 2 |
| 2022 | Neural Collective Matrix Factorization for integrated analysis of heterogeneous biomedical dataabstractMOTIVATION: In many biomedical studies, there arises the need to integrate data from multiple directly or indirectly related sources. Collective matrix factorization (CMF) and its variants are models designed to collectively learn from arbitrary collections of matrices. The latent factors learnt are rich integrative representations that can be used in downstream tasks, such as clustering or relation prediction with standard machine-learning models. Previous CMF-based methods have numerous modeling limitations. They do not adequately capture complex non-linear interactions and do not explicitly model varying sparsity and noise levels in the inputs, and some cannot model inputs with multiple datatypes. These inadequacies limit their use on many biomedical datasets. RESULTS: To address these limitations, we develop Neural Collective Matrix Factorization (NCMF), the first fully neural approach to CMF. We evaluate NCMF on relation prediction tasks of gene-disease association prediction and adverse drug event prediction, using multiple datasets. In each case, data are obtained from heterogeneous publicly available databases and used to learn representations to build predictive models. NCMF is found to outperform previous CMF-based methods and several state-of-the-art graph embedding methods for representation learning in our experiments. Our experiments illustrate the versatility and efficacy of NCMF in representation learning for seamless integration of heterogeneous data. AVAILABILITY AND IMPLEMENTATION: https://github.com/ajayago/NCMF_bioinformatics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ragunathan Mariappan, Aishwarya Jayagopal, Ho Zong Sien, Vaibhav Rajan |
Bioinform. | 4 |
| 2022 | An Algorithm to Mine Therapeutic Motifs for Cancer From Networks of Genetic InteractionsabstractStudy of pairwise genetic interactions, such as mutually exclusive mutations, has led to understanding of underlying mechanisms in cancer. Investigation of various combinatorial motifs within networks of such interactions can lead to deeper insights into its mutational landscape and inform therapy development. One such motif called the Between-Pathway Model (BPM) represents redundant or compensatory pathways that can be therapeutically exploited. Finding such BPM motifs is challenging since most formulations require solving variants of the NP-complete maximum weight bipartite subgraph problem. In this paper we design an algorithm based on Integer Linear Programming (ILP) to solve this problem. In our experiments, our approach outperforms the best previous method to mine BPM motifs. Further, our ILP-based approach allows us to easily model additional application-specific constraints. We illustrate this advantage through a new application of BPM motifs that can potentially aid in finding combination therapies to combat cancer. Herty Liany, Yu Lin 0001, Anand Jeyasekharan, Vaibhav Rajan |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Multiplex Bipartite Network Embedding using Dual Hypergraph Convolutional NetworksabstractA bipartite network is a graph structure where nodes are from two distinct domains and only inter-domain interactions exist as edges. A large number of network embedding methods exist to learn vectorial node representations from general graphs with both homogeneous and heterogeneous node and edge types, including some that can specifically model the distinct properties of bipartite networks. However, these methods are inadequate to model multiplex bipartite networks (e.g., in e-commerce), that have multiple types of interactions (e.g., click, inquiry, and buy) and node attributes. Most real-world multiplex bipartite networks are also sparse and have imbalanced node distributions that are challenging to model. In this paper, we develop an unsupervised Dual HyperGraph Convolutional Network (DualHGCN) model that scalably transforms the multiplex bipartite network into two sets of homogeneous hypergraphs and uses spectral hypergraph convolutional operators, along with intra- and inter-message passing strategies to promote information exchange within and across domains, to learn effective node embeddings. We benchmark DualHGCN using four real-world datasets on link prediction and node classification tasks. Our extensive experiments demonstrate that DualHGCN significantly outperforms state-of-the-art methods, and is robust to varying sparsity levels and imbalanced node distributions. Hansheng Xue, Luwei Yang, Vaibhav Rajan, Yu Lin 0001 |
WWW | 3 |
| 2021 | Maximum likelihood reconstruction of ancestral networks by integer linear programmingabstractMOTIVATION: The study of the evolutionary history of biological networks enables deep functional understanding of various bio-molecular processes. Network growth models, such as the Duplication-Mutation with Complementarity (DMC) model, provide a principled approach to characterizing the evolution of protein-protein interactions (PPIs) based on duplication and divergence. Current methods for model-based ancestral network reconstruction primarily use greedy heuristics and yield sub-optimal solutions. RESULTS: We present a new Integer Linear Programming (ILP) solution for maximum likelihood reconstruction of ancestral PPI networks using the DMC model. We prove the correctness of our solution that is designed to find the optimal solution. It can also use efficient heuristics from general-purpose ILP solvers to obtain multiple optimal and near-optimal solutions that may be useful in many applications. Experiments on synthetic data show that our ILP obtains solutions with higher likelihood than those from previous methods, and is robust to noise and model mismatch. We evaluate our algorithm on two real PPI networks, with proteins from the families of bZIP transcription factors and the Commander complex. On both the networks, solutions from our ILP have higher likelihood and are in better agreement with independent biological evidence from other studies. AVAILABILITY AND IMPLEMENTATION: A Python implementation is available at https://bitbucket.org/cdal/network-reconstruction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vaibhav Rajan, Ziqi Zhang 0011, Carl Kingsford, Xiuwei Zhang 0002 |
Bioinform. | 1 |
| 2020 | Multi-Task Learning with User Preferences: Gradient Descent with Controlled Ascent in Pareto OptimizationabstractMulti-Task Learning (MTL) is a well established paradigm for jointly learning models for multiple correlated tasks. Often the tasks conflict, requiring trade-offs between them during optimization. In such cases, multi-objective optimization based MTL methods can be used to find one or more Pareto optimal solutions. A common requirement in MTL applications, that cannot be addressed by these methods, is to find a solution satisfying userspecified preferences with respect to task-specific losses. We advance the state-of-the-art by developing the first gradient-based multi-objective MTL algorithm to solve this problem. Our unique approach combines multiple gradient descent with carefully controlled ascent to traverse the Pareto front in a principled manner, which also makes it robust to initialization. The scalability of our algorithm enables its use in large-scale deep networks for MTL. Assuming only differentiability of the task-specific loss functions, we provide theoretical guarantees for convergence. Our experiments show that our algorithm outperforms the best competing methods on benchmark datasets. Debabrata Mahapatra, Vaibhav Rajan |
ICML | 2 |
| 2020 | Gaussian mixture copulas for high-dimensional clustering and dependency-based subtypingabstractMOTIVATION: The identification of sub-populations of patients with similar characteristics, called patient subtyping, is important for realizing the goals of precision medicine. Accurate subtyping is crucial for tailoring therapeutic strategies that can potentially lead to reduced mortality and morbidity. Model-based clustering, such as Gaussian mixture models, provides a principled and interpretable methodology that is widely used to identify subtypes. However, they impose identical marginal distributions on each variable; such assumptions restrict their modeling flexibility and deteriorates clustering performance. RESULTS: In this paper, we use the statistical framework of copulas to decouple the modeling of marginals from the dependencies between them. Current copula-based methods cannot scale to high dimensions due to challenges in parameter inference. We develop HD-GMCM, that addresses these challenges and, to our knowledge, is the first copula-based clustering method that can fit high-dimensional data. Our experiments on real high-dimensional gene-expression and clinical datasets show that HD-GMCM outperforms state-of-the-art model-based clustering methods, by virtue of modeling non-Gaussian data and being robust to outliers through the use of Gaussian mixture copulas. We present a case study on lung cancer data from TCGA. Clusters obtained from HD-GMCM can be interpreted based on the dependencies they model, that offers a new way of characterizing subtypes. Empirically, such modeling not only uncovers latent structure that leads to better clustering but also meaningful clinical subtypes in terms of survival rates of patients. AVAILABILITY AND IMPLEMENTATION: An implementation of HD-GMCM in R is available at: https://bitbucket.org/cdal/hdgmcm/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Siva Rajesh Kasa, Sakyajit Bhattacharya, Vaibhav Rajan |
Bioinform. | 3 |
| 2020 | Predicting synthetic lethal interactions using heterogeneous data sourcesabstractMOTIVATION: A synthetic lethal (SL) interaction is a relationship between two functional entities where the loss of either one of the entities is viable but the loss of both entities is lethal to the cell. Such pairs can be used as drug targets in targeted anticancer therapies, and so, many methods have been developed to identify potential candidate SL pairs. However, these methods use only a subset of available data from multiple platforms, at genomic, epigenomic and transcriptomic levels; and hence are limited in their ability to learn from complex associations in heterogeneous data sources. RESULTS: In this article, we develop techniques that can seamlessly integrate multiple heterogeneous data sources to predict SL interactions. Our approach obtains latent representations by collective matrix factorization-based techniques, which in turn are used for prediction through matrix completion. Our experiments, on a variety of biological datasets, illustrate the efficacy and versatility of our approach, that outperforms state-of-the-art methods for predicting SL interactions and can be used with heterogeneous data sources with minimal feature engineering. AVAILABILITY AND IMPLEMENTATION: Software available at https://github.com/lianyh. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Herty Liany, Anand Jeyasekharan, Vaibhav Rajan |
Bioinform. | 3 |
| 2020 | MetaBCC-LR: metagenomics binning by coverage and composition for long readsabstractMOTIVATION: Metagenomics studies have provided key insights into the composition and structure of microbial communities found in different environments. Among the techniques used to analyse metagenomic data, binning is considered a crucial step to characterize the different species of micro-organisms present. The use of short-read data in most binning tools poses several limitations, such as insufficient species-specific signal, and the emergence of long-read sequencing technologies offers us opportunities to surmount them. However, most current metagenomic binning tools have been developed for short reads. The few tools that can process long reads either do not scale with increasing input size or require a database with reference genomes that are often unknown. In this article, we present MetaBCC-LR, a scalable reference-free binning method which clusters long reads directly based on their k-mer coverage histograms and oligonucleotide composition. RESULTS: We evaluate MetaBCC-LR on multiple simulated and real metagenomic long-read datasets with varying coverages and error rates. Our experiments demonstrate that MetaBCC-LR substantially outperforms state-of-the-art reference-free binning tools, achieving ∼13% improvement in F1-score and ∼30% improvement in ARI compared to the best previous tools. Moreover, we show that using MetaBCC-LR before long-read assembly helps to enhance the assembly quality while significantly reducing the assembly cost in terms of time and memory usage. The efficiency and accuracy of MetaBCC-LR pave the way for more effective long-read-based metagenomics analyses to support a wide range of applications. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at: https://github.com/anuradhawick/MetaBCC-LR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anuradha Wickramarachchi, Vijini Mallawaarachchi, Vaibhav Rajan, Yu Lin 0001 |
Bioinform. | 3 |
| 2019 | Inferring Concept Prerequisite Relations from Online Educational ResourcesabstractThe Internet has rich and rapidly increasing sources of high quality educational content. Inferring prerequisite relations between educational concepts is required for modern large-scale online educational technology applications such as personalized recommendations and automatic curriculum creation. We present PREREQ, a new supervised learning method for inferring concept prerequisite relations. PREREQ is designed using latent representations of concepts obtained from the Pairwise Latent Dirichlet Allocation model, and a neural network based on the Siamese network architecture. PREREQ can learn unknown concept prerequisites from course prerequisites and labeled concept prerequisite data. It outperforms state-of-the-art approaches on benchmark datasets and can effectively learn from very less training data. PREREQ can also use unlabeled video playlists, a steadily growing source of training data, to learn concept prerequisites, thus obviating the need for manual annotation of course prerequisites. Sudeshna Roy 0001, Meghana Madhyastha, Sheril Lawrence, Vaibhav Rajan |
AAAI | 4 |
| 2019 | Context-Aware Sequential Recommendations withStacked Recurrent Neural NetworksabstractSequential history of user interactions as well as the context of interactions provide valuable information to recommender systems, for modeling user behavior. Modeling both contexts and sequential information simultaneously, in context-aware sequential recommenders, has been shown to outperform methods that model either one of the two aspects. In long sequential histories, temporal trends are also found within sequences of contexts and temporal gaps that are not modeled by previous methods. In this paper we design new context-aware sequential recommendation methods, based on Stacked Recurrent Neural Networks, that model the dynamics of contexts and temporal gaps. Experiments on two large benchmark datasets demonstrate the advantages of modeling the evolution of contexts and temporal gaps - our models significantly outperform state-of-the-art context-aware sequential recommender systems. Lakshmanan Rakkappan, Vaibhav Rajan |
WWW | 2 |
| 2019 | Deep collective matrix factorization for augmented multi-view learning
Ragunathan Mariappan, Vaibhav Rajan |
Mach. Learn. | 2 |
| 2018 | Extractive Summarization with SWAP-NET: Sentences and Words from Alternating Pointer NetworksabstractWe present a new neural sequence-tosequence model for extractive summarization called SWAP-NET (Sentences and Words from Alternating Pointer Networks).Extractive summaries comprising a salient subset of input sentences, often also contain important key words.Guided by this principle, we design SWAP-NET that models the interaction of key words and salient sentences using a new twolevel pointer network based architecture.SWAP-NET identifies both salient sentences and key words in an input document, and then combines them to form the extractive summary.Experiments on large scale benchmark corpora demonstrate the efficacy of SWAP-NET that outperforms state-of-the-art extractive summarizers. Aishwarya Jadhav, Vaibhav Rajan |
ACL (1) | 2 |
| 2017 | ICU Mortality Prediction: A Classification Algorithm for Imbalanced DatasetsabstractDetermining mortality risk is important for critical decisions in Intensive Care Units (ICU). The need for machine learning models that provide accurate patient-specific prediction of mortality is well recognized. We present a new algorithm for ICU mortality prediction that is designed to address the problem of imbalance, which occurs, in the context of binary classification, when one of the two classes is significantly under--represented in the data. We take a fundamentally new approach in exploiting the class imbalance through a feature transformation such that the transformed features are easier to classify. Hypothesis testing is used for classification with a test statistic that follows the distribution of the difference of two chi-squared random variables, for which there are no analytic expressions and we derive an accurate approximation. Experiments on a benchmark dataset of 4000 ICU patients show that our algorithm surpasses the best competing methods for mortality prediction. Sakyajit Bhattacharya, Vaibhav Rajan, Harsh Shrivastava 0001 |
AAAI | 2 |
| 2017 | Vine copulas for mixed data : multi-view clustering for mixed data beyond meta-Gaussian dependencies
Lavanya Sita Tekumalla, Vaibhav Rajan, Chiranjib Bhattacharyya |
Mach. Learn. | 2 |
| 2016 | Dependency Clustering of Mixed Data with Gaussian Mixture Copulas
Vaibhav Rajan, Sakyajit Bhattacharya |
IJCAI | 1 |
| 2015 | Classification with imbalance: A similarity-based method for predicting respiratory failureabstractBinary classification based methods are commonly used for designing predictive models in healthcare. A common problem in many healthcare datasets is that of imbalance, where there are far more observations in one class than the other during training. In such conditions, most classifiers do not have good predictive accuracy with respect to the under-represented class. We design a new similarity-based classifier to learn from imbalanced datasets, wherein input features are transformed using similarity with respect to a chosen subset of training points. We empirically demonstrate the superiority of our algorithm over state-of-the-art methods for imbalanced data classification in real and synthetic datasets. We also illustrate the application of our classifier in predicting Acute Respiratory Failure (ARF), a critical complication in Intensive Care Units (ICU), using semi-structured text contained in nursing notes recorded during a patient's ICU stay. Our experiments, on more than 800 patient records show that using our new classifier to learn from text- based features can effectively be used to predict ARF and, potentially, other complications in ICUs. Harsh Shrivastava 0001, Vijay Huddar, Sakyajit Bhattacharya, Vaibhav Rajan |
BIBM | 4 |
| 2015 | Maximum Parsimony Analysis of Gene Copy Number Changes
Yu Lin 0001, Vaibhav Rajan, William Hoskins, Jijun Tang |
WABI | 3 |
| 2014 | CrowdUtility: A Recommendation System for Crowdsourcing PlatformsabstractCrowd workers exhibit varying work patterns, expertise, and quality leading to wide variability in the performance of crowdsourcing platforms. The onus of choosing a suitable platform to post tasks is mostly with the requester, often leading to poor guarantees and unmet requirements due to the dynamism in performance of crowd platforms. Towards this end, we demonstrate CrowdUtility, a statistical modelling based tool for evaluating multiple crowdsourcing platforms and recommending a platform that best suits the requirements of the requester. CrowdUtility uses an online Multi-Armed Bandit framework, to schedule tasks while optimizing platform performance. We demonstrate an end-to end system starting from requirements specification, to platform recommendation, to real-time monitoring. Deepthi Chander, Sakyajit Bhattacharya, L. Elisa Celis, Koustuv Dasgupta, Saraschandra Karanam, Vaibhav Rajan, Avantika Gupta |
HCOMP | 6 |
| 2014 | Adaptive Performance Optimization over Crowd Labor ChannelsabstractWe describe a system which monitors the performance of labor channels within a crowdsourcing platform in an online manner. This allows us to automatically determine if and when to switch between labor channels in order to improve overall performance of crowd tasks. Saraschandra Karanam, Deepthi Chander, L. Elisa Celis, Koustuv Dasgupta, Vaibhav Rajan |
HCOMP | 5 |
| 2012 | TIBA: a tool for phylogeny inference from rearrangement data with bootstrap analysisabstractTIBA is a tool to reconstruct phylogenetic trees from rearrangement data that consist of ordered lists of synteny blocks (or genes), where each synteny block is shared with all of its homologues in the input genomes. The evolution of these synteny blocks, through rearrangement operations, is modelled by the uniform Double-Cut-and-Join model. Using a true distance estimate under this model and simple distance-based methods, TIBA reconstructs a phylogeny of the input genomes. Unlike any previous tool for inferring phylogenies from rearrangement data, TIBA uses novel methods of robustness estimation to provide support values for the edges in the inferred tree. Yu Lin 0001, Vaibhav Rajan, Bernard M. E. Moret |
Bioinform. | 2 |
| 2012 | A Metric for Phylogenetic Trees Based on MatchingabstractComparing two or more phylogenetic trees is a fundamental task in computational biology. The simplest outcome of such a comparison is a pairwise measure of similarity, dissimilarity, or distance. A large number of such measures have been proposed, but so far all suffer from problems varying from computational cost to lack of robustness; many can be shown to behave unexpectedly under certain plausible inputs. For instance, the widely used Robinson-Foulds distance is poorly distributed and thus affords little discrimination, while also lacking robustness in the face of very small changes--reattaching a single leaf elsewhere in a tree of any size can instantly maximize the distance. In this paper, we introduce a new pairwise distance measure, based on matching, for phylogenetic trees. We prove that our measure induces a metric on the space of trees, show how to compute it in low polynomial time, verify through statistical testing that it is robust, and finally note that it does not exhibit unexpected behavior under the same inputs that cause problems with other measures. We also illustrate its usefulness in clustering trees, demonstrating significant improvements in the quality of hierarchical clustering as compared to the same collections of trees clustered using the Robinson-Foulds distance. Yu Lin 0001, Vaibhav Rajan, Bernard M. E. Moret |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | A Metric for Phylogenetic Trees Based on Matching
Yu Lin 0001, Vaibhav Rajan, Bernard M. E. Moret |
ISBRA | 2 |
| 2011 | Bootstrapping Phylogenies Inferred from Rearrangement Data
Yu Lin 0001, Vaibhav Rajan, Bernard M. E. Moret |
WABI | 2 |
| 2010 | Estimating true evolutionary distances under rearrangements, duplications, and lossesabstractBACKGROUND: The rapidly increasing availability of whole-genome sequences has enabled the study of whole-genome evolution. Evolutionary mechanisms based on genome rearrangements have attracted much attention and given rise to many models; somewhat independently, the mechanisms of gene duplication and loss have seen much work. However, the two are not independent and thus require a unified treatment, which remains missing to date. Moreover, existing rearrangement models do not fit the dichotomy between most prokaryotic genomes (one circular chromosome) and most eukaryotic genomes (multiple linear chromosomes). RESULTS: To handle rearrangements, gene duplications and losses, we propose a new evolutionary model and the corresponding method for estimating true evolutionary distance. Our model, inspired from the DCJ model, is simple and the first to respect the prokaryotic/eukaryotic structural dichotomy. Experimental results on a wide variety of genome structures demonstrate the very high accuracy and robustness of our distance estimator. CONCLUSION: We give the first robust, statistically based, estimate of genomic pairwise distances based on rearrangements, duplications and losses, under a model that respects the structural dichotomy between prokaryotic and eukaryotic genomes. Accurate and robust estimates in true evolutionary distances should translate into much better phylogenetic reconstructions as well as more accurate genomic alignments, while our new model of genome rearrangements provides another refinement in simplicity and verisimilitude. Yu Lin 0001, Vaibhav Rajan, Krister M. Swenson, Bernard M. E. Moret |
BMC Bioinform. | 2 |
| 2010 | Heuristics for the inversion median problemabstractBACKGROUND: The study of genome rearrangements has become a mainstay of phylogenetics and comparative genomics. Fundamental in such a study is the median problem: given three genomes find a fourth that minimizes the sum of the evolutionary distances between itself and the given three. Many exact algorithms and heuristics have been developed for the inversion median problem, of which the best known is MGR. RESULTS: We present a unifying framework for median heuristics, which enables us to clarify existing strategies and to place them in a partial ordering. Analysis of this framework leads to a new insight: the best strategies continue to refer to the input data rather than reducing the problem to smaller instances. Using this insight, we develop a new heuristic for inversion medians that uses input data to the end of its computation and leverages our previous work with DCJ medians. Finally, we present the results of extensive experimentation showing that our new heuristic outperforms all others in accuracy and, especially, in running time: the heuristic typically returns solutions within 1% of optimal and runs in seconds to minutes even on genomes with 25'000 genes--in contrast, MGR can take days on instances of 200 genes and cannot be used beyond 1'000 genes. CONCLUSION: Finding good rearrangement medians, in particular inversion medians, had long been regarded as the computational bottleneck in whole-genome studies. Our new heuristic for inversion medians, ASM, which dominates all others in our framework, puts that issue to rest by providing near-optimal solutions within seconds to minutes on even the largest genomes. Vaibhav Rajan, Andrew Wei Xu, Yu Lin 0001, Krister M. Swenson, Bernard M. E. Moret |
BMC Bioinform. | 1 |
| 2009 | Sorting Signed Permutations by Inversions in O(nlogn) Time
Krister M. Swenson, Vaibhav Rajan, Yu Lin 0001, Bernard M. E. Moret |
RECOMB | 2 |