Mingxiang Teng

dblp:65/7242 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-8536-8941ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 qcCHIP: an R package to identify clonal hematopoiesis variants using cohort-specific data characteristics
abstract
SUMMARY: Clonal hematopoiesis (CH) is a molecular biomarker associated with various adverse outcomes in both healthy individuals and those with underlying conditions, including cancer. Detecting CH usually involves genomic sequencing of individual blood samples followed by robust bioinformatics data filtering. We report an R package, qcCHIP, a bioinformatics pipeline that implements permutation-based parameter optimization to guide quality control filtering and cohort-specific CH identification. We benchmark qcCHIP under various data settings, including different sequencing depths, ranges of cohort sizes, with and without normal-tumor paired samples, and across different cancer types. We show that qcCHIP allows users to customize analysis needs to generate CH calls based on cohort-specific data characteristics. AVAILABILITY AND IMPLEMENTATION: qcCHIP R package is freely accessible at GitHub https://github.com/tenglab/qcCHIP and DOI: 10.5281/zenodo.16421861.
Yi-Han Tang, James Blachly, Stephen Edge, Yasminka A. Jakubek, Martin McCarter, Abdul Rafeh Naqash, Kenneth G. Nepple, Afaf Osman, Matthew J. Reilley, Gregory Riedlinger, Bodour Salhia, Bryan P. Schneider, Craig Shriver, Michelle L. Churchman, Robert J. Rounbehler, Jamie K. Teer, Nancy Gillis, Mingxiang Teng
Bioinform.19
2024 BatchFLEX: feature-level equalization of X-batch
abstract
MOTIVATION: Integrative analysis of heterogeneous expression data remains challenging due to variations in platform, RNA quality, sample processing, and other unknown technical effects. Selecting the approach for removing unwanted batch effects can be a time-consuming and tedious process, especially for more biologically focused investigators. RESULTS: Here, we present BatchFLEX, a Shiny app that can facilitate visualization and correction of batch effects using several established methods. BatchFLEX can visualize the variance contribution of a factor before and after correction. As an example, we have analyzed ImmGen microarray data and enhanced its expression signals that distinguishes each immune cell type. Moreover, our analysis revealed the impact of the batch correction in altering the gene expression rank and single-sample GSEA pathway scores in immune cell types, highlighting the importance of real-time assessment of the batch correction for optimal downstream analysis. AVAILABILITY AND IMPLEMENTATION: Our tool is available through Github https://github.com/shawlab-moffitt/BATCH-FLEX-ShinyApp with an online example on Shiny.io https://shawlab-moffitt.shinyapps.io/batch_flex/.
Joshua T. Davis, Alyssa N. Obermayer, Alex C. Soupir, Rebecca S. Hesterberg, Thac Duong, Ching-Yao Yang, Ken Phong Dao, Brandon J. Manley, G. Daniel Grass, Dorina Avram, Paulo C. Rodriguez, Brooke L. Fridley, Xiaoqing Yu, Mingxiang Teng, Timothy I. Shaw
Bioinform.14
2024 An Epigenomic fingerprint of human cancers by landscape interrogation of super enhancers at the constituent level
abstract
Super enhancers (SE), large genomic elements that activate transcription and drive cell identity, have been found with cancer-specific gene regulation in human cancers. Recent studies reported the importance of understanding the cooperation and function of SE internal components, i.e., the constituent enhancers (CE). However, there are no pan-cancer studies to identify cancer-specific SE signatures at the constituent level. Here, by revisiting pan-cancer SE activities with H3K27Ac ChIP-seq datasets, we report fingerprint SE signatures for 28 cancer types in the NCI-60 cell panel. We implement a mixture model to discriminate active CEs from inactive CEs by taking into consideration ChIP-seq variabilities between cancer samples and across CEs. We demonstrate that the model-based estimation of CE states provides improved functional interpretation of SE-associated regulation. We identify cancer-specific CEs by balancing their active prevalence with their capability of encoding cancer type identities. We further demonstrate that cancer-specific CEs have the strongest per-base enhancer activities in independent enhancer sequencing assays, suggesting their importance in understanding critical SE signatures. We summarize fingerprint SEs based on the cancer-specific statuses of their component CEs and build an easy-to-use R package to facilitate the query, exploration, and visualization of fingerprint SEs across cancers.
Nancy Gillis, Anthony McCofie, Timothy I. Shaw, Aik Choon Tan, Lixin Wan, Derek R. Duckett, Mingxiang Teng
PLoS Comput. Biol.10
2023 PATH-SURVEYOR: pathway level survival enquiry for immuno-oncology and drug repurposing
abstract
Pathway-level survival analysis offers the opportunity to examine molecular pathways and immune signatures that influence patient outcomes. However, available survival analysis algorithms are limited in pathway-level function and lack a streamlined analytical process. Here we present a comprehensive pathway-level survival analysis suite, PATH-SURVEYOR, which includes a Shiny user interface with extensive features for systematic exploration of pathways and covariates in a Cox proportional-hazard model. Moreover, our framework offers an integrative strategy for performing Hazard Ratio ranked Gene Set Enrichment Analysis and pathway clustering. As an example, we applied our tool in a combined cohort of melanoma patients treated with checkpoint inhibition (ICI) and identified several immune populations and biomarkers predictive of ICI efficacy. We also analyzed gene expression data of pediatric acute myeloid leukemia (AML) and performed an inverse association of drug targets with the patient's clinical endpoint. Our analysis derived several drug targets in high-risk KMT2A-fusion-positive patients, which were then validated in AML cell lines in the Genomics of Drug Sensitivity database. Altogether, the tool offers a comprehensive suite for pathway-level survival analysis and a user interface for exploring drug targets, molecular features, and immune populations at different resolutions.
Alyssa N. Obermayer, Darwin Chang, Gabrielle Nobles, Mingxiang Teng, Aik Choon Tan, Y. Ann Chen, Steven Eschrich, Paulo C. Rodriguez, G. Daniel Grass, Soheil Meschinchi, Ahmad Tarhini, Dung-Tsa Chen, Timothy I. Shaw
BMC Bioinform.4
2021 TIMEx: tumor-immune microenvironment deconvolution web-portal for bulk transcriptomics using pan-cancer scRNA-seq signatures
abstract
SUMMARY: The heterogeneous cell types of the tumor-immune microenvironment (TIME) play key roles in determining cancer progression, metastasis and response to treatment. We report the development of TIMEx, a novel TIME deconvolution method emphasizing on estimating infiltrating immune cells for bulk transcriptomics using pan-cancer single-cell RNA-seq signatures. We also implemented a comprehensive, user-friendly web-portal for users to evaluate TIMEx and other deconvolution methods with bulk transcriptomic profiles. AVAILABILITY AND IMPLEMENTATION: TIMEx web-portal is freely accessible at http://timex.moffitt.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mengyu Xie, Kyubum Lee, John H. Lockhart, Scott D. Cukras, Rodrigo Carvajal, Amer A. Beg, Elsa R. Flores, Mingxiang Teng, Christine H. Chung, Aik Choon Tan
Bioinform.8
2018 Modeling and correct the GC bias of tumor and normal WGS data for SCNA based tumor subclonal population inferring
abstract
BACKGROUND: Somatic copy number alternations (SCNAs) can be utilized to infer tumor subclonal populations in whole genome seuqncing studies, where usually their read count ratios between tumor-normal paired samples serve as the inferring proxy. Existing SCNA based subclonal population inferring tools consider the GC bias of tumor and normal sample is of the same fature, and could be fully offset by read count ratio. However, we found that, the read count ratio on SCNA segments presents a Log linear biased pattern, which influence existing read count ratios based subclonal inferring tools performance. Currently no correction tools take into account the read ratio bias. RESULTS: We present Pre-SCNAClonal, a tool that improving tumor subclonal population inferring by correcting GC-bias at SCNAs level. Pre-SCNAClonal first corrects GC bias using Markov chain Monte Carlo probability model, then accurately locates baseline DNA segments (not containing any SCNAs) with a hierarchy clustering model. We show Pre-SCNAClonal's superiority to exsiting GC-bias correction methods at any level of subclonal population. CONCLUSIONS: Pre-SCNAClonal could be run independently as well as serving as pre-processing/gc-correction step in conjuntion with exsiting SCNA-based subclonal inferring tools.
Yan-Shuo Chu, Mingxiang Teng, Yadong Wang 0001
BMC Bioinform.2
2017 Pre-SCNAClonal: Efficient GC bias correction for SCNA based tumor subclonal populations inferring
abstract
Somatic copy number alternations (SCNAs) can be utilized to infer tumor subclonal populations in whole genome seuqncing studies, where usually their read count ratios between tumor-normal paired samples serve as the inferring proxy. We found that, in a GC study, the GC contents and read count ratios on SCNA segments present a Log linear biased pattern. However, currently no subclonal inferring tools take into account this information. We provide Pre-SCNAClonal, a comprehensive GC bias correction tool for inferring tumor subclonal populations based on SCNAs. Results show that Pre-SCNAClonal could effectively and robustly correct the GC bias and improve the performance of the SCNAs based tumor subclonal population inferring tools. Pre-SCNAClonal could be strung together with the SCNAs based subclonal population inferring tool as a pipeline or run individually as needed.
Yan-Shuo Chu, Mingxiang Teng, Yongtian Wang, Yadong Wang 0001
BIBM2
2017 Pysubsim-tree: A package for simulating tumor genomes according to tumor evolution history
abstract
Next-generation sequencing (NGS) and the third generation sequencing (TGS) have recently allowed us to develop algorithms to quantitatively dissect the extent of heterogeneity within a tumour, resolve cancer evolution history and identify the somatic variations and aneuploidy events. Simulation of tumor NGS and TGS data serves as a powerful and cost-effective approach for benchmarking these algorithms, however, there is no available tool that could simulate all the distinct subclonal genomes with diverse aneuploidy events and somatic variations according to the given tumor evolution history. We provide a simulation package, Pysubsim-tree, which could simulate the tumor genomes according to their evolution history defined by the somatic variations and aneuploidy events. Pysubsim-tree is free, open source, available at: https://github.com/dustincys/pysubsimtree.
Yan-Shuo Chu, Ling Wang 0005, Mingxiang Teng, Yadong Wang 0001
BIBM4
2016 DMcompress: Dynamic Markov models for bacterial genome compression
abstract
Genome data increasing exponentially since the last decade, compressing genome with Markov models has been proposed as an effective statistical method. However, existing methods set a static order-k Markov models to compress various genomes. Employing static order-k Markov model could result in a sub-optimal orders on some genomes. In this paper, we propose a compression method that relies on a pre-analysis of the data before compression, with the aim of estimating Markov models order k, yielding improvements over static Markov models. Experimental results on the latest complete bacterial genome data show that our method could effectively compress genome with a better performance than the state-of-the-art method. The codes of DMcompress are available at https://rongjiewang.github.io/DMcompress.
Mingxiang Teng, Tianyi Zang, Yadong Wang 0001
BIBM2
2016 rHAT: fast alignment of noisy long reads with regional hashing
abstract
MOTIVATION: Single Molecule Real-Time (SMRT) sequencing has been widely applied in cutting-edge genomic studies. However, it is still an expensive task to align the noisy long SMRT reads to reference genome by state-of-the-art aligners, which is becoming a bottleneck in applications with SMRT sequencing. Novel approach is on demand for improving the efficiency and effectiveness of SMRT read alignment. RESULTS: We propose Regional Hashing-based Alignment Tool (rHAT), a seed-and-extension-based read alignment approach specifically designed for noisy long reads. rHAT indexes reference genome by regional hash table (RHT), a hash table-based index which describes the short tokens within local windows of reference genome. In the seeding phase, rHAT utilizes RHT for efficiently calculating the occurrences of short token matches between partial read and local genomic windows to find highly possible candidate sites. In the extension phase, a sparse dynamic programming-based heuristic approach is used for reducing the cost of aligning read to the candidate sites. By benchmarking on the real and simulated datasets from various prokaryote and eukaryote genomes, we demonstrated that rHAT can effectively align SMRT reads with outstanding throughput. AVAILABILITY AND IMPLEMENTATION: rHAT is implemented in C++; the source code is available at https://github.com/HIT-Bioinformatics/rHAT CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bo Liu 0023, Dengfeng Guan, Mingxiang Teng, Yadong Wang 0001
Bioinform.3
2015 Family genome browser: visualizing genomes with pedigree information
abstract
MOTIVATION: Families with inherited diseases are widely used in Mendelian/complex disease studies. Owing to the advances in high-throughput sequencing technologies, family genome sequencing becomes more and more prevalent. Visualizing family genomes can greatly facilitate human genetics studies and personalized medicine. However, due to the complex genetic relationships and high similarities among genomes of consanguineous family members, family genomes are difficult to be visualized in traditional genome visualization framework. How to visualize the family genome variants and their functions with integrated pedigree information remains a critical challenge. RESULTS: We developed the Family Genome Browser (FGB) to provide comprehensive analysis and visualization for family genomes. The FGB can visualize family genomes in both individual level and variant level effectively, through integrating genome data with pedigree information. Family genome analysis, including determination of parental origin of the variants, detection of de novo mutations, identification of potential recombination events and identical-by-decent segments, etc., can be performed flexibly. Diverse annotations for the family genome variants, such as dbSNP memberships, linkage disequilibriums, genes, variant effects, potential phenotypes, etc., are illustrated as well. Moreover, the FGB can automatically search de novo mutations and compound heterozygous variants for a selected individual, and guide investigators to find high-risk genes with flexible navigation options. These features enable users to investigate and understand family genomes intuitively and systematically. AVAILABILITY AND IMPLEMENTATION: The FGB is available at http://mlg.hit.edu.cn/FGB/.
Liran Juan, Yongzhuang Liu, Yongtian Wang, Mingxiang Teng, Tianyi Zang, Yadong Wang 0001
Bioinform.4
2012 regSNPs: a strategy for prioritizing regulatory single nucleotide substitutions
abstract
MOTIVATION: One of the fundamental questions in genetics study is to identify functional DNA variants that are responsible to a disease or phenotype of interest. Results from large-scale genetics studies, such as genome-wide association studies (GWAS), and the availability of high-throughput sequencing technologies provide opportunities in identifying causal variants. Despite the technical advances, informatics methodologies need to be developed to prioritize thousands of variants for potential causative effects. RESULTS: We present regSNPs, an informatics strategy that integrates several established bioinformatics tools, for prioritizing regulatory SNPs, i.e. the SNPs in the promoter regions that potentially affect phenotype through changing transcription of downstream genes. Comparing to existing tools, regSNPs has two distinct features. It considers degenerative features of binding motifs by calculating the differences on the binding affinity caused by the candidate variants and integrates potential phenotypic effects of various transcription factors. When tested by using the disease-causing variants documented in the Human Gene Mutation Database, regSNPs showed mixed performance on various diseases. regSNPs predicted three SNPs that can potentially affect bone density in a region detected in an earlier linkage study. Potential effects of one of the variants were validated using luciferase reporter assay.
Mingxiang Teng, Shoji Ichikawa, Leah R. Padgett, Yadong Wang 0001, Matthew E. Mort, David N. Cooper, Daniel L. Koller, Tatiana Foroud, Howard J. Edenberg, Michael J. Econs
Bioinform.1
2009 Enabling Data Analysis on High-Throughput Data in Large Data Depository Using Web-Based Analysis Platform - A Case Study on Integrating QUEST with GenePattern in Epigenetics Research
abstract
Enabling data analysis in large data depositories for high throughput experimental data such as gene microarrays and ChIP-seq is challenging. In this paper, we discuss three methods for integrating QUEST, a data depository for epigenetic experiments, with a web-based data analysis platform GenePattern. These methods are universal and can serve as an exemplary implementation resolving the dilemma facing many similar database systems in integrating data analysis tools.
Terry Camerlengo, Hatice Gulcin Ozer, Pearlly Yan, Jeffrey D. Parvin, Tim Hui-Ming Huang, Kun Huang 0001, Mingxiang Teng, Lang Li 0001, Francisco Perez, Tahsin M. Kurç
BIBM7