Tyler H. McCormick

dblp:138/5578 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-6490-1129ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Probabilistic and Bayesian machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational social science and digital humanities · 64% Bioinformatics and computational biology · 36%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
statistical estimation
0.912025
Data-Adaptive Exposure Thresholds under Network Interference · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.612022
Rethinking Nonlinear Instrumental Variable Models through Prediction Validity · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › causal inference
instrumental variable
0.612022
Rethinking Nonlinear Instrumental Variable Models through Prediction Validity · J. Mach. Learn. Res. 2022
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis
0.512021
Inference for Network Regression Models with Community Structure · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian graphical model
0.412019
Bayesian Joint Spike-and-Slab Graphical Lasso · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
gaussian graphical model
0.412019
Bayesian Joint Spike-and-Slab Graphical Lasso · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › sparse bayesian learning
spike-and-slab prior
0.412019
Bayesian Joint Spike-and-Slab Graphical Lasso · ICML 2019

Methods — techniques the papers use, named apart from their topics

linear regression · 1.7lepski's method · 1.7horvitz-thompson estimator · 1.7machine learning prediction · 0.9imputation · 0.9one-stage and two-stage estimation · 0.6machine learning · 0.6statistical inference · 0.5standard error estimation · 0.5exchangeability · 0.5group lasso · 0.4graphical lasso · 0.4fused lasso · 0.4EM algorithm · 0.4
YearPublicationVenuePosition
2025 Data-Adaptive Exposure Thresholds under Network Interference
abstract
Randomized controlled trials often suffer from interference, a violation of the Stable Unit Treatment Value Assumption (SUTVA), where a unit's outcome is influenced by its neighbors' treatment assignments. This interference biases naive estimators of the average treatment effect (ATE). A popular method to achieve unbiasedness pairs the Horvitz-Thompson estimator of the ATE with a known exposure mapping, a function that identifies units in a given randomization unaffected by interference. For example, an exposure mapping may stipulate that a unit experiences no further interference if at least an $h$-fraction of its neighbors share its treatment status. However, selecting this threshold $h$ is challenging, requiring domain expertise; in its absence, fixed thresholds such as $h = 1$ are often used. In this work, we propose a data-adaptive method to select the $h$-fractional threshold that minimizes the mean-squared-error (MSE) of the Horvitz-Thompson estimator. Our approach estimates the bias and variance of the Horvitz-Thompson estimator paired with candidate thresholds by leveraging a first-order approximation, specifically, linear regression of potential outcomes on exposures. We present simulations illustrating that our method improves upon non-adaptive threshold choices, and an adapted Lepski's method. We further illustrate the performance of our estimator by running experiments with synthetic outcomes on a real village network dataset, and on a publicly-available Amazon product similarity graph. Furthermore, we demonstrate that our method remains robust to deviations from the linear potential outcomes model.
Vydhourie Thiyageswaran, Tyler H. McCormick, Jennifer Brennan
NeurIPS2
2025 ipd: an R package for conducting inference on predicted data
abstract
SUMMARY: ipd is an open-source R software package for the downstream modeling of an outcome and its associated features where a potentially sizable portion of the outcome data has been imputed by an artificial intelligence or machine learning prediction algorithm. The package implements several recent proposed methods for inference on predicted data with a single, user-friendly wrapper function, ipd. The package also provides custom print, summary, tidy, glance, and augment methods to facilitate easy model inspection. This document introduces the ipd software package and provides a demonstration of its basic usage. AVAILABILITY: ipd is freely available on CRAN or as a developer version at our GitHub page: github.com/ipd-tools/ipd. Full documentation, including detailed instructions and a usage 'vignette' are available at github.com/ipd-tools/ipd.
Stephen Salerno, Awan Afiaz, Kentaro Hoffman, Anna Neufeld, Qiongshi Lu, Tyler H. McCormick, Jeffrey T. Leek
Bioinform.7
2022 Rethinking Nonlinear Instrumental Variable Models through Prediction Validity
abstract
Instrumental variables (IV) are widely used in the social and health sciences in situations where a researcher would like to measure a causal effect but cannot perform an experiment. For valid causal inference in an IV model, there must be external (exogenous) variation that (i) has a sufficiently large impact on the variable of interest (called the relevance assumption) and where (ii) the only pathway through which the external variation impacts the outcome is via the variable of interest (called the exclusion restriction). For statistical inference, researchers must also make assumptions about the functional form of the relationship between the three variables. Current practice assumes (i) and (ii) are met, then postulates a functional form with limited input from the data. In this paper, we describe a framework that leverages machine learning to validate these typically unchecked but consequential assumptions in the IV framework, providing the researcher empirical evidence about the quality of the instrument given the data at hand. Central to the proposed approach is the idea of prediction validity. Prediction validity checks that error terms -- which should be independent from the instrument -- cannot be modeled with machine learning any better than a model that is identically zero. We use prediction validity to develop both one-stage and two-stage approaches for IV, and demonstrate their performance on an example relevant to climate change policy.
Cynthia Rudin, Tyler H. McCormick
J. Mach. Learn. Res.3
2021 Inference for Network Regression Models with Community Structure
abstract
Network regression models, where the outcome comprises the valued edge in a network and the predictors are actor or dyad-level covariates, are used extensively in the social and biological sciences. Valid inference relies on accurately modeling the residual dependencies among the relations. Frequently homogeneity assumptions are placed on the errors which are commonly incorrect and ignore critical natural clustering of the actors. In this work, we present a novel regression modeling framework that models the errors as resulting from a community-based dependence structure and exploits the subsequent exchangeability properties of the error distribution to obtain parsimonious standard errors for regression parameters.
Mengjie Pan, Tyler H. McCormick, Bailey K. Fosdick
ICML2
2019 Bayesian Joint Spike-and-Slab Graphical Lasso
abstract
In this article, we propose a new class of priors for Bayesian inference with multiple Gaussian graphical models. We introduce Bayesian treatments of two popular procedures, the group graphical lasso and the fused graphical lasso, and extend them to a continuous spike-and-slab framework to allow self-adaptive shrinkage and model selection simultaneously. We develop an EM algorithm that performs fast and dynamic explorations of posterior modes. Our approach selects sparse models efficiently and automatically with substantially smaller bias than would be induced by alternative regularization procedures. The performance of the proposed methods are demonstrated through simulation and two real data examples.
Zehang Richard Li, Tyler H. McCormick, Samuel J. Clark
ICML2