VLDB 2026 Research / reviewers in the wild / expert
Michael R. Kosorok
dblp:91/7065 · also Michael Rene Kosorok
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0002-6070-9738ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 72% Probabilistic and Bayesian machine learning · 28% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Medical and health informatics · 84% Computational social science and digital humanities · 9% Bioinformatics and computational biology · 7% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 9 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics
precision medicine |
0.9 | 2 | 2021 | Estimation and Optimization of Composite Outcomes · J. Mach. Learn. Res. 2021 Kernel Assisted Learning for Personalized Dose Finding · KDD 2020 |
Machine learning › Reinforcement learning › value function estimation
bellman error |
0.7 | 1 | 2023 | Revisiting Bellman Errors for Offline Model Selection · ICML 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.7 | 1 | 2023 | Revisiting Bellman Errors for Offline Model Selection · ICML 2023 |
Machine learning › Reinforcement learning
value function estimation |
0.7 | 1 | 2023 | Revisiting Bellman Errors for Offline Model Selection · ICML 2023 |
Medical and health informatics › precision medicine
dynamic treatment regime |
0.5 | 1 | 2021 | Estimation and Optimization of Composite Outcomes · J. Mach. Learn. Res. 2021 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.1 | 1 | 2020 | Kernel Assisted Learning for Personalized Dose Finding · KDD 2020 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
individualized treatment rule |
0.1 | 1 | 2020 | Kernel Assisted Learning for Personalized Dose Finding · KDD 2020 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
gene pathway analysis |
0.1 | 1 | 2009 | Identification of differential gene pathways with principal component analysis · Bioinform. 2009 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2009 | Identification of differential gene pathways with principal component analysis · Bioinform. 2009 |
Methods — techniques the papers use, named apart from their topics
statistical inference · 1.3simulation · 1.3kernel methods · 1.3utility maximization · 1.0semiparametric estimation · 1.0causal inference · 1.0q-function · 0.7mean squared bellman error · 0.7principal component analysis · 0.1false discovery rate · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Deep doubly robust outcome weighted learning
Michael R. Kosorok |
Mach. Learn. | 3 |
| 2023 | Revisiting Bellman Errors for Offline Model SelectionabstractOffline model selection (OMS), that is, choosing the best policy from a set of many policies given only logged data, is crucial for applying offline RL in real-world settings. One idea that has been extensively explored is to select policies based on the mean squared Bellman error (MSBE) of the associated Q-functions. However, previous work has struggled to obtain adequate OMS performance with Bellman errors, leading many researchers to abandon the idea. To this end, we elucidate why previous work has seen pessimistic results with Bellman errors and identify conditions under which OMS algorithms based on Bellman errors will perform well. Moreover, we develop a new estimator of the MSBE that is more accurate than prior methods. Our estimator obtains impressive OMS performance on diverse discrete control tasks, including Atari games. Joshua P. Zitovsky, Daniel de Marchi, Rishabh Agarwal, Michael R. Kosorok |
ICML | 4 |
| 2022 | Sequence to Sequence ECG Cardiac Rhythm Classification Using Convolutional Recurrent Neural NetworksabstractThis paper proposes a novel deep learning architecture involving combinations of Convolutional Neural Networks (CNN) layers and Recurrent neural networks (RNN) layers that can be used to perform segmentation and classification of 5 cardiac rhythms based on ECG recordings. The algorithm is developed in a sequence to sequence setting where the input is a sequence of five second ECG signal sliding windows and the output is a sequence of cardiac rhythm labels. The novel architecture processes as input both the spectrograms of the ECG signal as well as the heartbeats' signal waveform. Additionally, we are able to train the model in the presence of label noise. The model's performance and generalizability is verified on an external database different from the one we used to train. Experimental result shows this approach can achieve an average F1 scores of 0.89 (averaged across 5 classes). The proposed model also achieves comparable classification performance to existing state-of-the-art approach with considerably less number of training parameters. Teeranan Pokaprakarn, Rebecca Kitzmiller, J. Randall Moorman, Douglas E. Lake, Ashok K. Krishnamurthy 0001, Michael R. Kosorok |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Estimation and Optimization of Composite OutcomesabstractThere is tremendous interest in precision medicine as a means to improve patient outcomes by tailoring treatment to individual characteristics. An individualized treatment rule formalizes precision medicine as a map from patient information to a recommended treatment. A treatment rule is defined to be optimal if it maximizes the mean of a scalar outcome in a population of interest, e.g., symptom reduction. However, clinical and intervention scientists often seek to balance multiple and possibly competing outcomes, e.g., symptom reduction and the risk of an adverse event. One approach to precision medicine in this setting is to elicit a composite outcome which balances all competing outcomes; unfortunately, eliciting a composite outcome directly from patients is difficult without a high-quality instrument, and an expert-derived composite outcome may not account for heterogeneity in patient preferences. We propose a new paradigm for the study of precision medicine using observational data that relies solely on the assumption that clinicians are approximately (i.e., imperfectly) making decisions to maximize individual patient utility. Estimated composite outcomes are subsequently used to construct an estimator of an individualized treatment rule which maximizes the mean of patient-specific composite outcomes. The estimated composite outcomes and estimated optimal individualized treatment rule provide new insights into patient preference heterogeneity, clinician behavior, and the value of precision medicine in a given domain. We derive inference procedures for the proposed estimators under mild conditions and demonstrate their finite sample performance through a suite of simulation experiments and an illustrative application to data from a study of bipolar depression. Daniel J. Luckett, Eric B. Laber, Siyeon Kim, Michael R. Kosorok |
J. Mach. Learn. Res. | 4 |
| 2020 | Kernel Assisted Learning for Personalized Dose FindingabstractAn individualized dose rule recommends a dose level within a continuous safe dose range based on patient level information such as physical conditions, genetic factors and medication histories. Traditionally, personalized dose finding process requires repeating clinical visits of the patient and frequent adjustments of the dosage. Thus the patient is constantly exposed to the risk of underdosing and overdosing during the process. Statistical methods for finding an optimal individualized dose rule can lower the costs and risks for patients. In this article, we propose a kernel assisted learning method for estimating the optimal individualized dose rule. The proposed methodology can also be applied to all other continuous decision-making problems. Advantages of the proposed method include robustness to model misspecification and capability of providing statistical inference for the estimated parameters. In the simulation studies, we show that this method is capable of identifying the optimal individualized dose rule and produces favorable expected outcomes in the population. Finally, we illustrate our approach using data from a warfarin dosing study for thrombosis patients. Liangyu Zhu, Wenbin Lu, Michael R. Kosorok, Rui Song 0006 |
KDD | 3 |
| 2020 | Differential gene regulatory pattern in the human brain from schizophrenia using transcriptomic-causal networkabstractBACKGROUND: Common and complex traits are the consequence of the interaction and regulation of multiple genes simultaneously, therefore characterizing the interconnectivity of genes is essential to unravel the underlying biological networks. However, the focus of many studies is on the differential expression of individual genes or on co-expression analysis. METHODS: Going beyond analysis of one gene at a time, we systematically integrated transcriptomics, genotypes and Hi-C data to identify interconnectivities among individual genes as a causal network. We utilized different machine learning techniques to extract information from the network and identify differential regulatory pattern between cases and controls. We used data from the Allen Brain Atlas for replication. RESULTS: Employing the integrative systems approach on the data from CommonMind Consortium showed that gene transcription is controlled by genetic variants proximal to the gene (cis-regulatory factors), and transcribed distal genes (trans-regulatory factors). We identified differential gene regulatory patterns in SCZ-cases versus controls and novel SCZ-associated genes that may play roles in the disorder since some of them are primary expressed in human brain. In addition, we observed genes known associated with SCZ are not likely (OR = 0.59) to have high impacts (degree > 3) on the network. CONCLUSIONS: Causal networks could reveal underlying patterns and the role of genes individually and as a group. Establishing principles that govern relationships between genes provides a mechanistic understanding of the dysregulated gene transcription patterns in SCZ and creates more efficient experimental designs for further studies. This information cannot be obtained by studying a single gene at the time. Akram Yazdani, Raul Mendez-Giraldez, Azam Yazdani, Michael R. Kosorok, Panos Roussos |
BMC Bioinform. | 4 |
| 2010 | Detection of gene pathways with predictive power for breast cancer prognosisabstractBACKGROUND: Prognosis is of critical interest in breast cancer research. Biomedical studies suggest that genomic measurements may have independent predictive power for prognosis. Gene profiling studies have been conducted to search for predictive genomic measurements. Genes have the inherent pathway structure, where pathways are composed of multiple genes with coordinated functions. The goal of this study is to identify gene pathways with predictive power for breast cancer prognosis. Since our goal is fundamentally different from that of existing studies, a new pathway analysis method is proposed. RESULTS: The new method advances beyond existing alternatives along the following aspects. First, it can assess the predictive power of gene pathways, whereas existing methods tend to focus on model fitting accuracy only. Second, it can account for the joint effects of multiple genes in a pathway, whereas existing methods tend to focus on the marginal effects of genes. Third, it can accommodate multiple heterogeneous datasets, whereas existing methods analyze a single dataset only. We analyze four breast cancer prognosis studies and identify 97 pathways with significant predictive power for prognosis. Important pathways missed by alternative methods are identified. CONCLUSIONS: The proposed method provides a useful alternative to existing pathway analysis methods. Identified pathways can provide further insights into breast cancer prognosis. Shuangge Ma, Michael R. Kosorok |
BMC Bioinform. | 2 |
| 2009 | Identification of differential gene pathways with principal component analysisabstractMOTIVATION: Development of high-throughput technology makes it possible to measure expressions of thousands of genes simultaneously. Genes have the inherent pathway structure, where pathways are composed of multiple genes with coordinated biological functions. It is of great interest to identify differential gene pathways that are associated with the variations of phenotypes. RESULTS: We propose the following approach for detecting differential gene pathways. First, we construct gene pathways using databases such as KEGG or GO. Second, for each pathway, we extract a small number of representative features, which are linear combinations of gene expressions and/or their transformations. Specifically, we propose using (i) principal components (PCs) of gene expression sets, (ii) PCs of expanded gene expression sets and (iii) expanded sets of PCs of gene expressions, as the representative features. Third, we identify differential gene pathways as those with representative features significantly associated with the variations of phenotypes, particularly disease clinical outcomes, in regression models. The false discovery rate approach is used to adjust for multiple comparisons. Analysis of three gene expression datasets suggests that (i) the proposed approach can effectively identify differential gene pathways; (ii) PCs that explain only a small amount of variations of gene expressions may bear significant associations between gene pathways and phenotypes; (iii) including second-order terms of gene expressions may lead to identification of new differential gene pathways; (iv) the proposed approach is relatively insensitive to additional noises; and (v) the proposed approach can identify gene pathways missed by alternative approaches. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuangge Ma, Michael R. Kosorok |
Bioinform. | 2 |