Xuecong Fu

dblp:277/1013 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0006-7295-4542ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Marker selection strategies for circulating tumor DNA guided by phylogenetic inference
abstract
MOTIVATION: Blood-based profiling of tumor DNA ("liquid biopsy") offers great prospects for non-invasive early cancer diagnosis and clinical guidance, but requires further computational advances to become a robust quantitative assay of tumor clonal evolution. We propose new methods to better characterize tumor clonal dynamics from circulating tumor DNA (ctDNA), through application to two specific tasks: (i) applying longitudinal ctDNA data to refine phylogeny models of clonal evolution, and (ii) quantifying changes in clonal frequencies that may be indicative of treatment response or tumor progression. We pose these through a probabilistic framework for optimally identifying markers and using them to characterize clonal evolution. RESULTS: We first estimate a density over clonal tree models using bootstrap samples over pre-treatment tissue-based sequence data. We then refine these models over successive longitudinal samples. We use the resulting framework for modeling and refining tree densities to pose a set of optimization problems for selecting ctDNA markers to maximize measures of utility for reducing uncertainty in phylogeny models and quantifying clonal frequencies given the models. We tested our methods on synthetic data and showed them to be effective at refining tree densities and inferring clonal frequencies. Application to real tumor data further demonstrated the methods' effectiveness in refining a lineage model and assessing its clonal frequencies. The work shows the power of computational methods to improve marker selection, clonal lineage reconstruction, and clonal dynamics profiling for more precise and quantitative assays of somatic evolution and tumor progression. AVAILABILITY AND IMPLEMENTATION: https://github.com/CMUSchwartzLab/Mase-phi.git. (DOI: 10.5281/zenodo.14776163).
Xuecong Fu, Zhicheng Luo, Yueqian Deng, William Laframboise, David Bartlett, Russell Schwartz
Bioinform.1
2022 Reconstructing tumor clonal lineage trees incorporating single-nucleotide variants, copy number alterations and structural variations
abstract
MOTIVATION: Cancer develops through a process of clonal evolution in which an initially healthy cell gives rise to progeny gradually differentiating through the accumulation of genetic and epigenetic mutations. These mutations can take various forms, including single-nucleotide variants (SNVs), copy number alterations (CNAs) or structural variations (SVs), with each variant type providing complementary insights into tumor evolution as well as offering distinct challenges to phylogenetic inference. RESULTS: In this work, we develop a tumor phylogeny method, TUSV-ext, which incorporates SNVs, CNAs and SVs into a single inference framework. We demonstrate on simulated data that the method produces accurate tree inferences in the presence of all three variant types. We further demonstrate the method through application to real prostate tumor data, showing how our approach to coordinated phylogeny inference and clonal construction with all three variant types can reveal a more complicated clonal structure than is suggested by prior work, consistent with extensive polyclonal seeding or migration. AVAILABILITY AND IMPLEMENTATION: https://github.com/CMUSchwartzLab/TUSV-ext. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuecong Fu, Haoyun Lei, Yifeng Tao, Russell Schwartz
Bioinform.1
2022 Semi-deconvolution of bulk and single-cell RNA-seq data with application to metastatic progression in breast cancer
abstract
MOTIVATION: Identifying cell types and their abundances and how these evolve during tumor progression is critical to understanding the mechanisms of metastasis and identifying predictors of metastatic potential that can guide the development of new diagnostics or therapeutics. Single-cell RNA sequencing (scRNA-seq) has been especially promising in resolving heterogeneity of expression programs at the single-cell level, but is not always feasible, e.g. for large cohort studies or longitudinal analysis of archived samples. In such cases, clonal subpopulations may still be inferred via genomic deconvolution, but deconvolution methods have limited ability to resolve fine clonal structure and may require reference cell type profiles that are missing or imprecise. Prior methods can eliminate the need for reference profiles but show unstable performance when few bulk samples are available. RESULTS: In this work, we develop a new method using reference scRNA-seq to interpret sample collections for which only bulk RNA-seq is available for some samples, e.g. clonally resolving archived primary tissues using scRNA-seq from metastases. By integrating such information in a Quadratic Programming framework, our method can recover more accurate cell types and corresponding cell type abundances in bulk samples. Application to a breast tumor bone metastases dataset confirms the power of scRNA-seq data to improve cell type inference and quantification in same-patient bulk samples. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at https://github.com/CMUSchwartzLab/RADs.
Haoyun Lei, Xiaoyan A. Guo, Yifeng Tao, Xuecong Fu, Steffi Oesterreich, Adrian V. Lee, Russell Schwartz
Bioinform.5
2022 Improving and evaluating deep learning models of cellular organization
abstract
MOTIVATION: Cells contain dozens of major organelles and thousands of other structures, many of which vary extensively in their number, size, shape and spatial distribution. This complexity and variation dramatically complicates the use of both traditional and deep learning methods to build accurate models of cell organization. Most cellular organelles are distinct objects with defined boundaries that do not overlap, while the pixel resolution of most imaging methods is n sufficient to resolve these boundaries. Thus while cell organization is conceptually object-based, most current methods are pixel-based. Using extensive image collections in which particular organelles were fluorescently labeled, deep learning methods can be used to build conditional autoencoder models for particular organelles. A major advance occurred with the use of a U-net approach to make multiple models all conditional upon a common reference, unlabeled image, allowing the relationships between different organelles to be at least partially inferred. RESULTS: We have developed improved Generative Adversarial Networks-based approaches for learning these models and have also developed novel criteria for evaluating how well synthetic cell images reflect the properties of real images. The first set of criteria measure how well models preserve the expected property that organelles do not overlap. We also developed a modified loss function that allows retraining of the models to minimize that overlap. The second set of criteria uses object-based modeling to compare object shape and spatial distribution between synthetic and real images. Our work provides the first demonstration that, at least for some organelles, deep learning models can capture object-level properties of cell images. AVAILABILITY AND IMPLEMENTATION: http://murphylab.cbd.cmu.edu/Software/2022_insilico. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Huangqingbo Sun, Xuecong Fu, Serena Abraham, Shen Jin, Robert F. Murphy
Bioinform.2
2021 ConTreeDP: A consensus method of tumor trees based on maximum directed partition support problem
abstract
Phylogenetic inference has become a crucial tool for interpreting cancer genomic data, but continuing advances in our understanding of somatic mutability in cancer, genomic technologies for profiling it, and the scale of data available have created a persistent need for new algorithms able to deal with these challenges. One particular need has been for new forms of consensus tree algorithms, which present special challenges in the cancer space for dealing with heterogeneous data, short evolutionary time scales, and rapid mutation by a wide variety of somatic mutability mechanisms. We develop a new consensus tree method for clonal phylogenetics, ConTreeDP, based on a formulation of the Maximum Directed Partition Support Consensus Tree (MDPSCT) problem. We demonstrate theoretically and empirically that our approach can efficiently and accurately compute clonal consensus trees from cancer genomic data. Availability: https://github.conCMUSchwartzLab/ConTreeDP
Xuecong Fu, Russell Schwartz
BIBM1
2021 Tumor heterogeneity assessed by sequencing and fluorescence in situ hybridization (FISH) data
abstract
MOTIVATION: Computational reconstruction of clonal evolution in cancers has become a crucial tool for understanding how tumors initiate and progress and how this process varies across patients. The field still struggles, however, with special challenges of applying phylogenetic methods to cancers, such as the prevalence and importance of copy number alteration (CNA) and structural variation events in tumor evolution, which are difficult to profile accurately by prevailing sequencing methods in such a way that subsequent reconstruction by phylogenetic inference algorithms is accurate. RESULTS: In this work, we develop computational methods to combine sequencing with multiplex interphase fluorescence in situ hybridization to exploit the complementary advantages of each technology in inferring accurate models of clonal CNA evolution accounting for both focal changes and aneuploidy at whole-genome scales. By integrating such information in an integer linear programming framework, we demonstrate on simulated data that incorporation of FISH data substantially improves accurate inference of focal CNA and ploidy changes in clonal evolution from deconvolving bulk sequence data. Analysis of real glioblastoma data for which FISH, bulk sequence and single cell sequence are all available confirms the power of FISH to enhance accurate reconstruction of clonal copy number evolution in conjunction with bulk and optionally single-cell sequence data. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at https://github.com/CMUSchwartzLab/FISH_deconvolution. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haoyun Lei, E. Michael Gertz, Alejandro A. Schäffer, Xuecong Fu, Yifeng Tao, Kerstin Heselmeyer-Haddad, Irianna Torres, Guibo Li, Liqin Xu, Yong Hou, Kui Wu 0005, Xulian Shi, Mike Dean, Thomas Ried, Russell Schwartz
Bioinform.4
2020 Robust and accurate deconvolution of tumor populations uncovers evolutionary mechanisms of breast cancer metastasis
abstract
MOTIVATION: Cancer develops and progresses through a clonal evolutionary process. Understanding progression to metastasis is of particular clinical importance, but is not easily analyzed by recent methods because it generally requires studying samples gathered years apart, for which modern single-cell sequencing is rarely an option. Revealing the clonal evolution mechanisms in the metastatic transition thus still depends on unmixing tumor subpopulations from bulk genomic data. METHODS: We develop a novel toolkit called robust and accurate deconvolution (RAD) to deconvolve biologically meaningful tumor populations from multiple transcriptomic samples spanning the two progression states. RAD uses gene module compression to mitigate considerable noise in RNA, and a hybrid optimizer to achieve a robust and accurate solution. Finally, we apply a phylogenetic algorithm to infer how associated cell populations adapt across the metastatic transition via changes in expression programs and cell-type composition. RESULTS: We validated the superior robustness and accuracy of RAD over alternative algorithms on a real dataset, and validated the effectiveness of gene module compression on both simulated and real bulk RNA data. We further applied the methods to a breast cancer metastasis dataset, and discovered common early events that promote tumor progression and migration to different metastatic sites, such as dysregulation of ECM-receptor, focal adhesion and PI3k-Akt pathways. AVAILABILITY AND IMPLEMENTATION: The source code of the RAD package, models, experiments and technical details such as parameters, is available at https://github.com/CMUSchwartzLab/RAD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yifeng Tao, Haoyun Lei, Xuecong Fu, Adrian V. Lee, Jian Ma 0004, Russell Schwartz
Bioinform.3