Kahn Rhrissorrakrai

dblp:32/9661 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-1567-9090ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quantum ensembling methods for healthcare and life science
abstract
Learning on sample-limited data is a challenge frequently encountered in many real-world applications. In this work we study how effective quantum ensemble models are when trained on a sample-limited data problem in healthcare and life sciences. We constructed multiple types of quantum ensembles for binary classification using up to 26 qubits in simulation and 56 qubits on quantum hardware. The ensembles include both variational and non-variational methods as well as introducing a new quantum ensemble cosine classifier with randomly sampled unitaries. Our ensemble designs use minimal trainable parameters but require long-range connections between qubits. We tested these quantum ensembles on synthetic datasets and gene expression data from renal cell carcinoma (RCC) patients with the task of predicting patient response to immunotherapy. From the performance observed in simulation and quantum hardware experiments using up to 56 qubits, we demonstrate how quantum embedding structure affects performance and discuss how to extract informative features and build models that can learn and generalize effectively. We also find that quantum ensemble cosine classifiers were effective in learning from few training data, and all quantum ensembles performed comparatively to classical ensembles while using significantly fewer learners. We confirmed this performance characteristic in a separate RCC validation cohort. We present these exploratory results in order to assist other researchers in the design of effective learning using ensembles, particularly for similarly size constrained problems. Incorporating quantum computing in these data constrained problems offers hope for a wide range of studies in healthcare and life sciences where biological samples are relatively scarce given the feature space to be explored.
Kahn Rhrissorrakrai, Kathleen E. Hamilton, Prerana Bangalore Parthasarathy, Aldo Guzmán-Sáenz, Tyler Alban, Filippo Utro, Laxmi Parida
Briefings Bioinform.1
2025 Quantum Doubly Stochastic Transformers
abstract
At the core of the Transformer, the softmax normalizes the attention matrix to be right stochastic. Previous research has shown that this often de-stabilizes training and that enforcing the attention matrix to be doubly stochastic (through Sinkhorn’s algorithm) consistently improves performance across different tasks, domains and Transformer flavors. However, Sinkhorn’s algorithm is iterative, approximative, non-parametric and thus inflexible w.r.t. the obtained doubly stochastic matrix (DSM). Recently, it has been proven that DSMs can be obtained with a parametric quantum circuit, yielding a novel quantum inductive bias for DSMs with no known classical analogue. Motivated by this, we demonstrate the feasibility of a hybrid classical-quantum doubly stochastic Transformer (QDSFormer) that replaces the softmax in the self-attention layer with a variational quantum circuit. We study the expressive power of the circuit and find that it yields more diverse DSMs that better preserve information than classical operators. Across multiple small-scale object recognition tasks, we find that our QDSFormer consistently surpasses both a standard ViT and other doubly stochastic Transformers. Beyond the Sinkformer, this comparison includes a novel quantum-inspired doubly stochastic Transformer (based on QR decomposition) that can be of independent interest. Our QDSFormer also shows improved training stability and lower performance variation suggesting that it may mitigate the notoriously unstable training of ViTs on small-scale data.
Jannis Born, Filip Skogh, Kahn Rhrissorrakrai, Filippo Utro, Nico Wagner, Alexandros Sobczyk
NeurIPS3
2024 Epidemiological topology data analysis links severe COVID-19 to RAAS and hyperlipidemia associated metabolic syndrome conditions
abstract
MOTIVATION: The emergence of COVID-19 (C19) created incredible worldwide challenges but offers unique opportunities to understand the physiology of its risk factors and their interactions with complex disease conditions, such as metabolic syndrome. To address the challenges of discovering clinically relevant interactions, we employed a unique approach for epidemiological analysis powered by redescription-based topological data analysis (RTDA). RESULTS: Here, RTDA was applied to Explorys data to discover associations among severe C19 and metabolic syndrome. This approach was able to further explore the probative value of drug prescriptions to capture the involvement of RAAS and hypertension with C19, as well as modification of risk factor impact by hyperlipidemia (HL) on severe C19. RTDA found higher-order relationships between RAAS pathway and severe C19 along with demographic variables of age, gender, and comorbidities such as obesity, statin prescriptions, HL, chronic kidney failure, and disproportionately affecting Black individuals. RTDA combined with CuNA (cumulant-based network analysis) yielded a higher-order interaction network derived from cumulants that furthered supported the central role that RAAS plays. TDA techniques can provide a novel outlook beyond typical logistic regressions in epidemiology. From an observational cohort of electronic medical records, it can find out how RAAS drugs interact with comorbidities, such as hypertension and HL, of patients with severe bouts of C19. Where single variable association tests with outcome can struggle, TDA's higher-order interaction network between different variables enables the discovery of the comorbidities of a disease such as C19 work in concert. AVAILABILITY AND IMPLEMENTATION: Code for performing TDA/RTDA is available in https://github.com/IBM/Matilda and code for CuNA can be found in https://github.com/BiomedSciAI/Geno4SD/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel E. Platt, Aritra Bose, Kahn Rhrissorrakrai, Chaya Levovitz, Laxmi Parida
Bioinform.3
2023 Lesion Shedding Model: unraveling site-specific contributions to ctDNA
abstract
Sampling circulating tumor DNA (ctDNA) using liquid biopsies offers clinically important benefits for monitoring cancer progression. A single ctDNA sample represents a mixture of shed tumor DNA from all known and unknown lesions within a patient. Although shedding levels have been suggested to hold the key to identifying targetable lesions and uncovering treatment resistance mechanisms, the amount of DNA shed by any one specific lesion is still not well characterized. We designed the Lesion Shedding Model (LSM) to order lesions from the strongest to the poorest shedding for a given patient. By characterizing the lesion-specific ctDNA shedding levels, we can better understand the mechanisms of shedding and more accurately interpret ctDNA assays to improve their clinical impact. We verified the accuracy of the LSM under controlled conditions using a simulation approach as well as testing the model on three cancer patients. The LSM obtained an accurate partial order of the lesions according to their assigned shedding levels in simulations and its accuracy in identifying the top shedding lesion was not significantly impacted by number of lesions. Applying LSM to three cancer patients, we found that indeed there were lesions that consistently shed more than others into the patients' blood. In two of the patients, the top shedding lesion was one of the only clinically progressing lesions at the time of biopsy suggesting a connection between high ctDNA shedding and clinical progression. The LSM provides a much needed framework with which to understand ctDNA shedding and to accelerate discovery of ctDNA biomarkers. The LSM source code has been available in the IBM BioMedSciAI Github (https://github.com/BiomedSciAI/Geno4SD).
Kahn Rhrissorrakrai, Filippo Utro, Chaya Levovitz, Laxmi Parida
Briefings Bioinform.1
2022 Epidemiological topology data analysis links severe COVID-19 to RAAS and hyperlipidemia associated metabolic syndrome conditions
Daniel E. Platt, Aritra Bose, Chaya Levovitz, Kahn Rhrissorrakrai, Laxmi Parida
AMIA4
2019 Dark-matter matters: Discriminating subtle blood cancers using the darkest DNA
abstract
The confluence of deep sequencing and powerful machine learning is providing an unprecedented peek at the darkest of the dark genomic matter, the non-coding genomic regions lacking any functional annotation. While deep sequencing uncovers rare tumor variants, the heterogeneity of the disease confounds the best of machine learning (ML) algorithms. Here we set out to answer if the dark-matter of the genome encompass signals that can distinguish the fine subtypes of disease that are otherwise genomically indistinguishable. We introduce a novel stochastic regularization, ReVeaL, that empowers ML to discriminate subtle cancer subtypes even from the same 'cell of origin'. Analogous to heritability, implicitly defined on whole genome, we use predictability (F1 score) definable on portions of the genome. In an effort to distinguish cancer subtypes using dark-matter DNA, we applied ReVeaL to a new WGS dataset from 727 patient samples with seven forms of hematological cancers and assessed the predictivity over several genomic regions including genic, non-dark, non-coding, non-genic, and dark. ReVeaL enabled improved discrimination of cancer subtypes for all segments of the genome. The non-genic, non-coding and dark-matter had the highest F1 scores, with dark-matter having the highest level of predictability. Based on ReVeaL's predictability of different genomic regions, dark-matter contains enough signal to significantly discriminate fine subtypes of disease. Hence, the agglomeration of rare variants, even in the hitherto unannotated and ill-understood regions of the genome, may play a substantial role in the disease etiology and deserve much more attention.
Laxmi Parida, Claudia Haferlach, Kahn Rhrissorrakrai, Filippo Utro, Chaya Levovitz, Wolfgang Kern, Niroshan Nadarajah, Sven Twardziok, Stephan Hutter, Manja Meggendorfer, Wencke Walter, Constance Baer, Torsten Haferlach
PLoS Comput. Biol.3
2016 Erratum to: MINE: Module Identification in Networks
abstract
Erratum It was brought to our attention that there was a discrepancy between the description and implementation of the MINE algorithm in our article [1]. We regret any inconvenience that may have resulted from this inaccuracy. The primary difference is that as implemented, only the immediate neighborhood of each node is searched to identify initial clusters, and the merging process then facilitates cluster growth and removal of redundant clusters. This implementation results in accelerated cluster prediction relative to a more exhaustive search, which would perform comparably to a depthfirst search across a larger portion of the network for each node. The published results and performance of MINE are unaffected; the description of the method, algorithm pseudocode, Fig. 1, and Additional file 1: Figure S1 have been updated to accurately reflect the implemented version of MINE.
Kahn Rhrissorrakrai, Kristin C. Gunsalus
BMC Bioinform.1
2016 A Crowdsourcing Approach to Developing and Assessing Prediction Algorithms for AML Prognosis
abstract
Acute Myeloid Leukemia (AML) is a fatal hematological cancer. The genetic abnormalities underlying AML are extremely heterogeneous among patients, making prognosis and treatment selection very difficult. While clinical proteomics data has the potential to improve prognosis accuracy, thus far, the quantitative means to do so have yet to be developed. Here we report the results and insights gained from the DREAM 9 Acute Myeloid Prediction Outcome Prediction Challenge (AML-OPC), a crowdsourcing effort designed to promote the development of quantitative methods for AML prognosis prediction. We identify the most accurate and robust models in predicting patient response to therapy, remission duration, and overall survival. We further investigate patient response to therapy, a clinically actionable prediction, and find that patients that are classified as resistant to therapy are harder to predict than responsive patients across the 31 models submitted to the challenge. The top two performing models, which held a high sensitivity to these patients, substantially utilized the proteomics data to make predictions. Using these models, we also identify which signaling proteins were useful in predicting patient therapeutic response.
David Noren, Byron Long, Raquel Norel, Kahn Rhrissorrakrai, Kenneth R. Hess, Chenyue W. Hu, Alex Bisberg, André Schultz, Erik Engquist, Li Liu 0035, Xihui Lin, Gregory M. Chen, Honglei Xie, Geoffrey A. M. Hunter, Paul C. Boutros, Oleg A. Stepanov, Thea Norman, Stephen H. Friend, Gustavo Stolovitzky, Steven M. Kornblau, Amina A. Qutub
PLoS Comput. Biol.4
2015 Inter-species prediction of protein phosphorylation in the sbv IMPROVER species translation challenge
abstract
MOTIVATION: Animal models are widely used in biomedical research for reasons ranging from practical to ethical. An important issue is whether rodent models are predictive of human biology. This has been addressed recently in the framework of a series of challenges designed by the systems biology verification for Industrial Methodology for Process Verification in Research (sbv IMPROVER) initiative. In particular, one of the sub-challenges was devoted to the prediction of protein phosphorylation responses in human bronchial epithelial cells, exposed to a number of different chemical stimuli, given the responses in rat bronchial epithelial cells. Participating teams were asked to make inter-species predictions on the basis of available training examples, comprising transcriptomics and phosphoproteomics data. RESULTS: Here, the two best performing teams present their data-driven approaches and computational methods. In addition, post hoc analyses of the datasets and challenge results were performed by the participants and challenge organizers. The challenge outcome indicates that successful prediction of protein phosphorylation status in human based on rat phosphorylation levels is feasible. However, within the limitations of the computational tools used, the inclusion of gene expression data does not improve the prediction quality. The post hoc analysis of time-specific measurements sheds light on the signaling pathways in both species. AVAILABILITY AND IMPLEMENTATION: A detailed description of the dataset, challenge design and outcome is available at www.sbvimprover.com. The code used by team IGB is provided under http://github.com/uci-igb/improver2013. Implementations of the algorithms applied by team AMG are available at http://bhanot.biomaps.rutgers.edu/wiki/AMG-sc2-code.zip. CONTACT: [email protected].
Michael Biehl, Peter J. Sadowski, Gyan Bhanot, Erhan Bilal, Adel Dayarian, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Michael D. Zeller, Sahand Hormoz
Bioinform.8
2015 A crowd-sourcing approach for the construction of species-specific cell signaling networks
abstract
MOTIVATION: Animal models are important tools in drug discovery and for understanding human biology in general. However, many drugs that initially show promising results in rodents fail in later stages of clinical trials. Understanding the commonalities and differences between human and rat cell signaling networks can lead to better experimental designs, improved allocation of resources and ultimately better drugs. RESULTS: The sbv IMPROVER Species-Specific Network Inference challenge was designed to use the power of the crowds to build two species-specific cell signaling networks given phosphoproteomics, transcriptomics and cytokine data generated from NHBE and NRBE cells exposed to various stimuli. A common literature-inspired reference network with 220 nodes and 501 edges was also provided as prior knowledge from which challenge participants could add or remove edges but not nodes. Such a large network inference challenge not based on synthetic simulations but on real data presented unique difficulties in scoring and interpreting the results. Because any prior knowledge about the networks was already provided to the participants for reference, novel ways for scoring and aggregating the results were developed. Two human and rat consensus networks were obtained by combining all the inferred networks. Further analysis showed that major signaling pathways were conserved between the two species with only isolated components diverging, as in the case of ribosomal S6 kinase RPS6KA1. Overall, the consensus between inferred edges was relatively high with the exception of the downstream targets of transcription factors, which seemed more difficult to predict. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erhan Bilal, Theodore Sakellaropoulos, Challenge Participants, Ioannis N. Melas, Dimitris E. Messinis, Vincenzo Belcastro, Kahn Rhrissorrakrai, Pablo Meyer 0001, Raquel Norel, Anita Iskandar, Elise Blaese, John Jeremy Rice, Manuel C. Peitsch, Julia Hoeng, Gustavo Stolovitzky, Leonidas G. Alexopoulos, Carine Poussin
Bioinform.7
2015 Predicting protein phosphorylation from gene expression: top methods from the IMPROVER Species Translation Challenge
abstract
MOTIVATION: Using gene expression to infer changes in protein phosphorylation levels induced in cells by various stimuli is an outstanding problem. The intra-species protein phosphorylation challenge organized by the IMPROVER consortium provided the framework to identify the best approaches to address this issue. RESULTS: Rat lung epithelial cells were treated with 52 stimuli, and gene expression and phosphorylation levels were measured. Competing teams used gene expression data from 26 stimuli to develop protein phosphorylation prediction models and were ranked based on prediction performance for the remaining 26 stimuli. Three teams were tied in first place in this challenge achieving a balanced accuracy of about 70%, indicating that gene expression is only moderately predictive of protein phosphorylation. In spite of the similar performance, the approaches used by these three teams, described in detail in this article, were different, with the average number of predictor genes per phosphoprotein used by the teams ranging from 3 to 124. However, a significant overlap of gene signatures between teams was observed for the majority of the proteins considered, while Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways were enriched in the union of the predictor genes of the three teams for multiple proteins. AVAILABILITY AND IMPLEMENTATION: Gene expression and protein phosphorylation data are available from ArrayExpress (E-MTAB-2091). Software implementation of the approach of Teams 49 and 75 are available at http://bioinformaticsprb.med.wayne.edu and http://people.cs.clemson.edu/∼luofeng/sbv.rar, respectively. CONTACT: [email protected] or [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Adel Dayarian, Roberto Romero, Michael Biehl, Erhan Bilal, Sahand Hormoz, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Gyan Bhanot, Feng Luo 0001, Adi L. Tarca
Bioinform.9
2015 Inter-species pathway perturbation prediction via data-driven detection of functional homology
abstract
MOTIVATION: Experiments in animal models are often conducted to infer how humans will respond to stimuli by assuming that the same biological pathways will be affected in both organisms. The limitations of this assumption were tested in the IMPROVER Species Translation Challenge, where 52 stimuli were applied to both human and rat cells and perturbed pathways were identified. In the Inter-species Pathway Perturbation Prediction sub-challenge, multiple teams proposed methods to use rat transcription data from 26 stimuli to predict human gene set and pathway activity under the same perturbations. Submissions were evaluated using three performance metrics on data from the remaining 26 stimuli. RESULTS: We present two approaches, ranked second in this challenge, that do not rely on sequence-based orthology between rat and human genes to translate pathway perturbation state but instead identify transcriptional response orthologs across a set of training conditions. The translation from rat to human accomplished by these so-called direct methods is not dependent on the particular analysis method used to identify perturbed gene sets. In contrast, machine learning-based methods require performing a pathway analysis initially and then mapping the pathway activity between organisms. Unlike most machine learning approaches, direct methods can be used to predict the activation of a human pathway for a new (test) stimuli, even when that pathway was never activated by a training stimuli. AVAILABILITY: Gene expression data are available from ArrayExpress (accession E-MTAB-2091), while software implementations are available from http://bioinformaticsprb.med.wayne.edu?p=50 and http://goo.gl/hJny3h. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Christoph Hafemeister, Roberto Romero, Erhan Bilal, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Richard Bonneau, Adi L. Tarca
Bioinform.6
2015 Inter-species inference of gene set enrichment in lung epithelial cells from proteomic and large transcriptomic datasets
abstract
MOTIVATION: Translating findings in rodent models to human models has been a cornerstone of modern biology and drug development. However, in many cases, a naive 'extrapolation' between the two species has not succeeded. As a result, clinical trials of new drugs sometimes fail even after considerable success in the mouse or rat stage of development. In addition to in vitro studies, inter-species translation requires analytical tools that can predict the enriched gene sets in human cells under various stimuli from corresponding measurements in animals. Such tools can improve our understanding of the underlying biology and optimize the allocation of resources for drug development. RESULTS: We developed an algorithm to predict differential gene set enrichment as part of the sbv IMPROVER (systems biology verification in Industrial Methodology for Process Verification in Research) Species Translation Challenge, which focused on phosphoproteomic and transcriptomic measurements of normal human bronchial epithelial (NHBE) primary cells under various stimuli and corresponding measurements in rat (NRBE) primary cells. We find that gene sets exhibit a higher inter-species correlation compared with individual genes, and are potentially more suited for direct prediction. Furthermore, in contrast to a similar cross-species response in protein phosphorylation states 5 and 25 min after exposure to stimuli, gene set enrichment 6 h after exposure is significantly different in NHBE cells compared with NRBE cells. In spite of this difference, we were able to develop a robust algorithm to predict gene set activation in NHBE with high accuracy using simple analytical methods. AVAILABILITY AND IMPLEMENTATION: Implementation of all algorithms is available as source code (in Matlab) at http://bhanot.biomaps.rutgers.edu/wiki/codes_SC3_Predicting_GeneSets.zip, along with the relevant data used in the analysis. Gene sets, gene expression and protein phosphorylation data are available on request. CONTACT: [email protected].
Sahand Hormoz, Gyan Bhanot, Michael Biehl, Erhan Bilal, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Adel Dayarian
Bioinform.7
2015 Understanding the limits of animal models as predictors of human biology: lessons learned from the sbv IMPROVER Species Translation Challenge
abstract
MOTIVATION: Inferring how humans respond to external cues such as drugs, chemicals, viruses or hormones is an essential question in biomedicine. Very often, however, this question cannot be addressed because it is not possible to perform experiments in humans. A reasonable alternative consists of generating responses in animal models and 'translating' those results to humans. The limitations of such translation, however, are far from clear, and systematic assessments of its actual potential are urgently needed. sbv IMPROVER (systems biology verification for Industrial Methodology for PROcess VErification in Research) was designed as a series of challenges to address translatability between humans and rodents. This collaborative crowd-sourcing initiative invited scientists from around the world to apply their own computational methodologies on a multilayer systems biology dataset composed of phosphoproteomics, transcriptomics and cytokine data derived from normal human and rat bronchial epithelial cells exposed in parallel to 52 different stimuli under identical conditions. Our aim was to understand the limits of species-to-species translatability at different levels of biological organization: signaling, transcriptional and release of secreted factors (such as cytokines). Participating teams submitted 49 different solutions across the sub-challenges, two-thirds of which were statistically significantly better than random. Additionally, similar computational methods were found to range widely in their performance within the same challenge, and no single method emerged as a clear winner across all sub-challenges. Finally, computational methods were able to effectively translate some specific stimuli and biological processes in the lung epithelial system, such as DNA synthesis, cytoskeleton and extracellular matrix, translation, immune/inflammation and growth factor/proliferation pathways, better than the expected response similarity between species. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kahn Rhrissorrakrai, Vincenzo Belcastro, Erhan Bilal, Raquel Norel, Carine Poussin, Carole Mathis, Rémi H. J. Dulize, Nikolai V. Ivanov, Leonidas G. Alexopoulos, John Jeremy Rice, Manuel C. Peitsch, Gustavo Stolovitzky, Pablo Meyer 0001, Julia Hoeng
Bioinform.1
2011 MINE: Module Identification in NEtworks
abstract
BACKGROUND: Graphical models of network associations are useful for both visualizing and integrating multiple types of association data. Identifying modules, or groups of functionally related gene products, is an important challenge in analyzing biological networks. However, existing tools to identify modules are insufficient when applied to dense networks of experimentally derived interaction data. To address this problem, we have developed an agglomerative clustering method that is able to identify highly modular sets of gene products within highly interconnected molecular interaction networks. RESULTS: MINE outperforms MCODE, CFinder, NEMO, SPICi, and MCL in identifying non-exclusive, high modularity clusters when applied to the C. elegans protein-protein interaction network. The algorithm generally achieves superior geometric accuracy and modularity for annotated functional categories. In comparison with the most closely related algorithm, MCODE, the top clusters identified by MINE are consistently of higher density and MINE is less likely to designate overlapping modules as a single unit. MINE offers a high level of granularity with a small number of adjustable parameters, enabling users to fine-tune cluster results for input networks with differing topological properties. CONCLUSIONS: MINE was created in response to the challenge of discovering high quality modules of gene products within highly interconnected biological networks. The algorithm allows a high degree of flexibility and user-customisation of results with few adjustable parameters. MINE outperforms several popular clustering algorithms in identifying modules with high modularity and obtains good overall recall and precision of functional annotations in protein-protein interaction networks from both S. cerevisiae and C. elegans.
Kahn Rhrissorrakrai, Kristin C. Gunsalus
BMC Bioinform.1