VLDB 2026 Research / reviewers in the wild / expert
Mitchell J. Machiela
dblp:170/4557
· DBLP profile ↗
8ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-6538-9705ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Estimation of mosaic loss of Y chromosome cell fraction with genotyping arrays lacking coverage in the pseudoautosomal regionabstractAbstract Background Mosaic loss of the Y chromosome (mLOY) in circulating leukocytes is the most frequently detected age-related chromosomal mosaic event in men. Current mLOY detection approaches use genotyping arrays and employ a phase-based approach that identifies B allele frequency (BAF) deviations in the pseudo-autosomal region (PAR) shared between the X and Y chromosome. As some widely used genotyping arrays lack sufficient probe coverage of the PAR, methods for accurately measuring mLOY utilizing the median log2 R ratio across the male-specific region of Y chromosome (mLRR_Y) are needed for detecting mLOY on these platforms. Results We derived a formula from mLRR_Y to estimate the cellular fraction (CF) of cells with Y loss and validated the approach, finding high alignment with the CF estimation from female data and lab-generated qPCR data (R2 = 0.98). Additionally, we compared the correlation between phase-based BAF and mLRR_Y methods for CF estimation, achieving a high correlation with R2 > 0.80. Conclusion Although mLRR_Y is a noisier metric for mosaic chromosomal alteration detection relative to BAF, we demonstrate mLRR_Y across non-PAR variants can accurately estimate mLOY CF, especially for high CF mLOY. Weiyin Zhou, Wenyi Huang, Neal D. Freedman, Mitchell J. Machiela |
BMC Bioinform. | 4 |
| 2024 | ArCH: improving the performance of clonal hematopoiesis variant calling and interpretationabstractMOTIVATION: The acquisition of somatic mutations in hematopoietic stem and progenitor stem cells with resultant clonal expansion, termed clonal hematopoiesis (CH), is associated with increased risk of hematologic malignancies and other adverse outcomes. CH is generally present at low allelic fractions, but clonal expansion and acquisition of additional mutations leads to hematologic cancers in a small proportion of individuals. With high depth and high sensitivity sequencing, CH can be detected in most adults and its clonal trajectory mapped over time. However, accurate CH variant calling is challenging due to the difficulty in distinguishing low frequency CH mutations from sequencing artifacts. The lack of well-validated bioinformatic pipelines for CH calling may contribute to lack of reproducibility in studies of CH. RESULTS: Here, we developed ArCH, an Artifact filtering Clonal Hematopoiesis variant calling pipeline for detecting single nucleotide variants and short insertions/deletions by combining the output of four variant calling tools and filtering based on variant characteristics and sequencing error rate estimation. ArCH is an end-to-end cloud-based pipeline optimized to accept a variety of inputs with customizable parameters adaptable to multiple sequencing technologies, research questions, and datasets. Using deep targeted sequencing data generated from six acute myeloid leukemia patient tumor: normal dilutions, 31 blood samples with orthogonal validation, and 26 blood samples with technical replicates, we show that ArCH improves the sensitivity and positive predictive value of CH variant detection at low allele frequencies compared to standard application of commonly used variant calling approaches. AVAILABILITY AND IMPLEMENTATION: The code for this workflow is available at: https://github.com/kbolton-lab/ArCH. Irenaeus C. C. Chan, Alex Panchot, Evelyn Schmidt, Samantha N. McNulty, Brian J. Wiley, Kimberly Turner, Lea Moukarzel, Wendy S. W. Wong, J. Scott Beeler, Armel Landry Batchi-Bouyou, Mitchell J. Machiela, Danielle M. Karyadi, Benjamin J. Krajacich, Semyon Kruglyak, Bryan R. Lajoie, Shawn E. Levy, Philip W. Kantoff, Christopher E. Mason, Daniel C. Link, Todd E. Druley, Konrad H. Stopsack, Kelly L. Bolton |
Bioinform. | 13 |
| 2022 | PLCOjs, a FAIR GWAS web SDK for the NCI Prostate, Lung, Colorectal and Ovarian Cancer Genetic Atlas projectabstractMOTIVATION: The Division of Cancer Epidemiology and Genetics (DCEG) and the Division of Cancer Prevention (DCP) at the National Cancer Institute (NCI) have recently generated genome-wide association study (GWAS) data for multiple traits in the Prostate, Lung, Colorectal, and Ovarian (PLCO) Genomic Atlas project. The GWAS included 110 000 participants. The dissemination of the genetic association data through a data portal called GWAS Explorer, in a manner that addresses the modern expectations of FAIR reusability by data scientists and engineers, is the main motivation for the development of the open-source JavaScript software development kit (SDK) reported here. RESULTS: The PLCO GWAS Explorer resource relies on a public stateless HTTP application programming interface (API) deployed as the sole backend service for both the landing page's web application and third-party analytical workflows. The core PLCOjs SDK is mapped to each of the API methods, and also to each of the reference graphic visualizations in the GWAS Explorer. A few additional visualization methods extend it. As is the norm with web SDKs, no download or installation is needed and modularization supports targeted code injection for web applications, reactive notebooks (Observable) and node-based web services. AVAILABILITY AND IMPLEMENTATION: code at https://github.com/episphere/plco; project page at https://episphere.github.io/plco. Eric Ruan, Erika Nemeth, Richard A. Moffitt, Lorena Sandoval, Mitchell J. Machiela, Neal D. Freedman, Wenyi Huang, Wendy Wong, Kai-Ling Chen, Brian Park, Kevin Jiang, Belynda Hicks, Daniel E. Russ, Lori M. Minasian, Paul F. Pinsky, Stephen J. Chanock, Montserrat Garcia-Closas, Jonas S. Almeida |
Bioinform. | 5 |
| 2021 | PCAmatchR: a flexible R package for optimal case-control matching using weighted principal componentsabstractSUMMARY: A concern when conducting genome-wide association studies (GWAS) is the potential for population stratification, i.e. ancestry-based genetic differences between cases and controls, that if not properly accounted for, could lead to biased association results. We developed PCAmatchR as an open source R package for performing optimal case-control matching using principal component analysis (PCA) to aid in selecting controls that are well matched by ancestry to cases. PCAmatchR takes user supplied PCA outputs and selects matching controls for cases by utilizing a weighted Mahalanobis distance metric which weights each principal component by the percentage of genetic variation explained. Results from the 1000 Genomes Project data demonstrate both the functionality and performance of PCAmatchR for selecting matching controls for case populations as well as reducing inflation of association test statistics. PCAmatchR improves genomic similarity between matched cases and controls, which minimizes the effects of population stratification in GWAS analyses. AVAILABILITY AND IMPLEMENTATION: PCAmatchR is freely available for download on GitHub (https://github.com/machiela-lab/PCAmatchR) or through CRAN (https://CRAN.R-project.org/package=PCAmatchR). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Derek W. Brown, Timothy A. Myers, Mitchell J. Machiela |
Bioinform. | 3 |
| 2021 | LDexpress: an online tool for integrating population-specific linkage disequilibrium patterns with tissue-specific expression dataabstractGenome-wide association studies have identified thousands of genetic susceptibility loci associated with cancer as well as other traits and diseases. Mapping germline variation in identified genetic susceptibility regions to alterations in nearby gene expression nominates candidate genes potentially related to disease risk for further functional investigation. We developed LDexpress as an online resource that integrates population-specific linkage disequilibrium data from the 1000 Genomes (1000G) project and tissue-specific expression data from the Genotype-Tissue Expression project to better study regional germline variation impacting gene expression. LDexpress is a publicly available web tool designed to be easy to use, flexible to conduct a wide range of variant queries, and quick to efficiently investigate dozens of query variants across multiple tissue types. We demonstrate the utility of LDexpress using example genomic queries and anticipate this tool will accelerate understanding of disease etiology by uncovering associations of regional germline variation to nearby gene expression. Shu-Hong Lin, Rohit Thakur, Mitchell J. Machiela |
BMC Bioinform. | 3 |
| 2020 | LDpop: an interactive online tool to calculate and visualize geographic LD patternsabstractBACKGROUND: Linkage disequilibrium (LD)-the non-random association of alleles at different loci-defines population-specific haplotypes which vary by genomic ancestry. Assessment of allelic frequencies and LD patterns from a variety of ancestral populations enables researchers to better understand population histories as well as improve genetic understanding of diseases in which risk varies by ethnicity. RESULTS: We created an interactive web module which allows for quick geographic visualization of linkage disequilibrium (LD) patterns between two user-specified germline variants across geographic populations included in the 1000 Genomes Project. Interactive maps and a downloadable, sortable summary table allow researchers to easily compute and compare allele frequencies and LD statistics of dbSNP catalogued variants. The geographic mapping of each SNP's allele frequencies by population as well as visualization of LD statistics allows the user to easily trace geographic allelic correlation patterns and examine population-specific differences. CONCLUSIONS: LDpop is a free and publicly available cross-platform web tool which can be accessed online at https://ldlink.nci.nih.gov/?tab=ldpop. Theresa A. Alexander, Mitchell J. Machiela |
BMC Bioinform. | 2 |
| 2018 | LDassoc: an online tool for interactively exploring genome-wide association study results and prioritizing variants for functional investigationabstractMotivation: Existing approaches to plot association results from genome-wide association studies (GWAS) are in the form of static Manhattan plots and often lack data integration with rich databases on variant regulatory potential as well as population-specific linkage disequilibrium patterns. Summary: We created an intuitive web module for uploading and efficiently exploring GWAS association results. Interactive plots and sortable tables allow researchers to query genomic regions of interest, facilitating the integration of data on linkage disequilibrium, variant regulatory potential and potential target genes. External links allow for visualization of association results in the UCSC genome browser as well as easy access to publically available databases (e.g. dbSNP and RegulomeDB). Through improved visualization and data integration, LDassoc offers genomic researchers a specialized environment to examine association signals and suggests variants for functional investigation. Availability and implementation: LDassoc is a free and publically available web tool which can be accessed online at https://analysistools.nci.nih.gov/LDlink/? tab=ldassoc. Contact: [email protected]. Mitchell J. Machiela, Stephen J. Chanock |
Bioinform. | 1 |
| 2015 | LDlink: a web-based application for exploring population-specific haplotype structure and linking correlated alleles of possible functional variantsabstractUNLABELLED: Assessing linkage disequilibrium (LD) across ancestral populations is a powerful approach for investigating population-specific genetic structure as well as functionally mapping regions of disease susceptibility. Here, we present LDlink, a web-based collection of bioinformatic modules that query single nucleotide polymorphisms (SNPs) in population groups of interest to generate haplotype tables and interactive plots. Modules are designed with an emphasis on ease of use, query flexibility, and interactive visualization of results. Phase 3 haplotype data from the 1000 Genomes Project are referenced for calculating pairwise metrics of LD, searching for proxies in high LD, and enumerating all observed haplotypes. LDlink is tailored for investigators interested in mapping common and uncommon disease susceptibility loci by focusing on output linking correlated alleles and highlighting putative functional variants. AVAILABILITY AND IMPLEMENTATION: LDlink is a free and publically available web tool which can be accessed at http://analysistools.nci.nih.gov/LDlink/. CONTACT: [email protected]. Mitchell J. Machiela, Stephen J. Chanock |
Bioinform. | 1 |