EDBT 2026 Demo / reviewers in the wild / expert
Jian Ma 0004
dblp:26/4870-4
· DBLP profile ↗
38ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4202-5834ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unified Integration of Spatial Transcriptomics Across Platforms
Ellie Haber, Ajinkya Deshpande, Jian Ma 0004, Spencer Krieger |
RECOMB | 3 |
| 2025 | Steamboat: Attention-Based Multiscale Delineation of Cellular Interactions in Tissues
Shaoheng Liang, Guanghan Wang, Jian Ma 0004 |
RECOMB | 4 |
| 2023 | UNADON: transformer-based model to predict genome-wide chromosome spatial positionabstractMOTIVATION: The spatial positioning of chromosomes relative to functional nuclear bodies is intertwined with genome functions such as transcription. However, the sequence patterns and epigenomic features that collectively influence chromatin spatial positioning in a genome-wide manner are not well understood. RESULTS: Here, we develop a new transformer-based deep learning model called UNADON, which predicts the genome-wide cytological distance to a specific type of nuclear body, as measured by TSA-seq, using both sequence features and epigenomic signals. Evaluations of UNADON in four cell lines (K562, H1, HFFc6, HCT116) show high accuracy in predicting chromatin spatial positioning to nuclear bodies when trained on a single cell line. UNADON also performed well in an unseen cell type. Importantly, we reveal potential sequence and epigenomic factors that affect large-scale chromatin compartmentalization in nuclear bodies. Together, UNADON provides new insights into the principles between sequence features and large-scale chromatin spatial localization, which has important implications for understanding nuclear structure and function. AVAILABILITY AND IMPLEMENTATION: The source code of UNADON can be found at https://github.com/ma-compbio/UNADON. Muyu Yang, Jian Ma 0004 |
Bioinform. | 2 |
| 2022 | Concert: Genome-Wide Prediction of Sequence Elements That Modulate DNA Replication Timing
Yang Yang 0094, Yang Zhang 0042, Jian Ma 0004 |
RECOMB | 4 |
| 2022 | Ultrafast and Interpretable Single-Cell 3D Genome Analysis with Fast-Higashi
Ruochi Zhang, Tianming Zhou, Jian Ma 0004 |
RECOMB | 3 |
| 2022 | Guest Editorial for Selected Papers From ACM-BCB 2019
Jian Ma 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Assessing the contribution of tumor mutational phenotypes to cancer progression riskabstractCancer occurs via an accumulation of somatic genomic alterations in a process of clonal evolution. There has been intensive study of potential causal mutations driving cancer development and progression. However, much recent evidence suggests that tumor evolution is normally driven by a variety of mechanisms of somatic hypermutability, which act in different combinations or degrees in different cancers. These variations in mutability phenotypes are predictive of progression outcomes independent of the specific mutations they have produced to date. Here we explore the question of how and to what degree these differences in mutational phenotypes act in a cancer to predict its future progression. We develop a computational paradigm using evolutionary tree inference (tumor phylogeny) algorithms to derive features quantifying single-tumor mutational phenotypes, followed by a machine learning framework to identify key features predictive of progression. Analyses of breast invasive carcinoma and lung carcinoma demonstrate that a large fraction of the risk of future clinical outcomes of cancer progression-overall survival and disease-free survival-can be explained solely from mutational phenotype features derived from the phylogenetic analysis. We further show that mutational phenotypes have additional predictive power even after accounting for traditional clinical and driver gene-centric genomic predictors of progression. These results confirm the importance of mutational phenotypes in contributing to cancer progression risk and suggest strategies for enhancing the predictive power of conventional clinical data or driver-centric biomarkers. Yifeng Tao, Ashok Rajaraman, Xiaoyue Cui, Ziyi Cui, Haoran Chen 0006, Yuanqi Zhao, Jesse Eaton, Hannah Kim 0003, Jian Ma 0004, Russell Schwartz |
PLoS Comput. Biol. | 9 |
| 2020 | Hyper-SAGNN: a self-attention based graph neural network for hypergraphs
Ruochi Zhang, Yuesong Zou, Jian Ma 0004 |
ICLR | 3 |
| 2020 | Probing Multi-way Chromatin Interaction with Hypergraph Representation Learning
Ruochi Zhang, Jian Ma 0004 |
RECOMB | 2 |
| 2020 | Robust and accurate deconvolution of tumor populations uncovers evolutionary mechanisms of breast cancer metastasisabstractMOTIVATION: Cancer develops and progresses through a clonal evolutionary process. Understanding progression to metastasis is of particular clinical importance, but is not easily analyzed by recent methods because it generally requires studying samples gathered years apart, for which modern single-cell sequencing is rarely an option. Revealing the clonal evolution mechanisms in the metastatic transition thus still depends on unmixing tumor subpopulations from bulk genomic data. METHODS: We develop a novel toolkit called robust and accurate deconvolution (RAD) to deconvolve biologically meaningful tumor populations from multiple transcriptomic samples spanning the two progression states. RAD uses gene module compression to mitigate considerable noise in RNA, and a hybrid optimizer to achieve a robust and accurate solution. Finally, we apply a phylogenetic algorithm to infer how associated cell populations adapt across the metastatic transition via changes in expression programs and cell-type composition. RESULTS: We validated the superior robustness and accuracy of RAD over alternative algorithms on a real dataset, and validated the effectiveness of gene module compression on both simulated and real bulk RNA data. We further applied the methods to a breast cancer metastasis dataset, and discovered common early events that promote tumor progression and migration to different metastatic sites, such as dysregulation of ECM-receptor, focal adhesion and PI3k-Akt pathways. AVAILABILITY AND IMPLEMENTATION: The source code of the RAD package, models, experiments and technical details such as parameters, is available at https://github.com/CMUSchwartzLab/RAD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yifeng Tao, Haoyun Lei, Xuecong Fu, Adrian V. Lee, Jian Ma 0004, Russell Schwartz |
Bioinform. | 5 |
| 2020 | Cancer mutational signatures representation by large-scale context embeddingabstractMOTIVATION: The accumulation of somatic mutations plays critical roles in cancer development and progression. However, the global patterns of somatic mutations, especially non-coding mutations, and their roles in defining molecular subtypes of cancer have not been well characterized due to the computational challenges in analysing the complex mutational patterns. RESULTS: Here, we develop a new algorithm, called MutSpace, to effectively extract patient-specific mutational features using an embedding framework for larger sequence context. Our method is motivated by the observation that the mutation rate at megabase scale and the local mutational patterns jointly contribute to distinguishing cancer subtypes, both of which can be simultaneously captured by MutSpace. Simulation evaluations show that MutSpace can effectively characterize mutational features from known patient subgroups and achieve superior performance compared with previous methods. As a proof-of-principle, we apply MutSpace to 560 breast cancer patient samples and demonstrate that our method achieves high accuracy in subtype identification. In addition, the learned embeddings from MutSpace reflect intrinsic patterns of breast cancer subtypes and other features of genome structure and function. MutSpace is a promising new framework to better understand cancer heterogeneity based on somatic mutations. AVAILABILITY AND IMPLEMENTATION: Source code of MutSpace can be accessed at: https://github.com/ma-compbio/MutSpace. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Zhang 0042, Yunxuan Xiao, Muyu Yang, Jian Ma 0004 |
Bioinform. | 4 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 38 |
| 2019 | Comparing 3D Genome Organization in Multiple Species Using Phylo-HMRF
Yang Yang 0094, Yang Zhang 0042, Jesse R. Dixon, Jian Ma 0004 |
RECOMB | 5 |
| 2019 | Rotation equivariant and invariant neural networks for microscopy image analysisabstractMOTIVATION: Neural networks have been widely used to analyze high-throughput microscopy images. However, the performance of neural networks can be significantly improved by encoding known invariance for particular tasks. Highly relevant to the goal of automated cell phenotyping from microscopy image data is rotation invariance. Here we consider the application of two schemes for encoding rotation equivariance and invariance in a convolutional neural network, namely, the group-equivariant CNN (G-CNN), and a new architecture with simple, efficient conic convolution, for classifying microscopy images. We additionally integrate the 2D-discrete-Fourier transform (2D-DFT) as an effective means for encoding global rotational invariance. We call our new method the Conic Convolution and DFT Network (CFNet). RESULTS: We evaluated the efficacy of CFNet and G-CNN as compared to a standard CNN for several different image classification tasks, including simulated and real microscopy images of subcellular protein localization, and demonstrated improved performance. We believe CFNet has the potential to improve many high-throughput microscopy image analysis applications. AVAILABILITY AND IMPLEMENTATION: Source code of CFNet is available at: https://github.com/bchidest/CFNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin Chidester, Tianming Zhou, Minh N. Do, Jian Ma 0004 |
Bioinform. | 4 |
| 2018 | Minimax Reconstruction Risk of Convolutional Sparse Dictionary LearningabstractSparse dictionary learning (SDL) has become a popular method for learning parsimonious representations of data, a fundamental problem in machine learning and signal processing. While most work on SDL assumes a training dataset of independent and identically distributed (IID) samples, a variant known as convolutional sparse dictionary learning (CSDL) relaxes this assumption to allow dependent, non-stationary sequential data sources. Recent work has explored statistical properties of IID SDL; however, the statistical properties of CSDL remain largely unstudied. This paper identifies minimax rates of CSDL in terms of reconstruction risk, providing both lower and upper bounds in a variety of settings. Our results make minimal assumptions, allowing arbitrary dictionaries and showing that CSDL is robust to dependent noise. We compare our results to similar results for IID SDL and verify our theory with synthetic experiments. Shashank Singh 0005, Barnabás Póczos, Jian Ma 0004 |
AISTATS | 3 |
| 2018 | Covariate Adjusted Precision Matrix Estimation via Nonconvex OptimizationabstractWe propose a nonconvex estimator for the covariate adjusted precision matrix estimation problem in the high dimensional regime, under sparsity constraints. To solve this estimator, we propose an alternating gradient descent algorithm with hard thresholding. Compared with existing methods along this line of research, which lack theoretical guarantees in optimization error and/or statistical error, the proposed algorithm not only is computationally much more efficient with a linear rate of convergence, but also attains the optimal statistical rate up to a logarithmic factor. Thorough experiments on both synthetic and real data support our theory. Pan Xu 0002, Lingxiao Wang 0001, Jian Ma 0004, Quanquan Gu |
ICML | 4 |
| 2018 | Integration of Spatial Distribution in Imaging-Genetics
Vaishnavi Subramanian, Weizhao Tang, Benjamin Chidester, Jian Ma 0004, Minh N. Do |
MICCAI (2) | 4 |
| 2018 | Continuous-Trait Probabilistic Model for Comparing Multi-species Functional Genomic Data
Yang Yang 0094, Quanquan Gu, Takayo Sasaki, Julianna Crivello, Rachel O'Neill, David M. Gilbert, Jian Ma 0004 |
RECOMB | 7 |
| 2018 | Predicting CTCF-mediated chromatin loops using CTCF-MPabstractMotivation: The three dimensional organization of chromosomes within the cell nucleus is highly regulated. It is known that CCCTC-binding factor (CTCF) is an important architectural protein to mediate long-range chromatin loops. Recent studies have shown that the majority of CTCF binding motif pairs at chromatin loop anchor regions are in convergent orientation. However, it remains unknown whether the genomic context at the sequence level can determine if a convergent CTCF motif pair is able to form a chromatin loop. Results: In this article, we directly ask whether and what sequence-based features (other than the motif itself) may be important to establish CTCF-mediated chromatin loops. We found that motif conservation measured by 'branch-of-origin' that accounts for motif turn-over in evolution is an important feature. We developed a new machine learning algorithm called CTCF-MP based on word2vec to demonstrate that sequence-based features alone have the capability to predict if a pair of convergent CTCF motifs would form a loop. Together with functional genomic signals from CTCF ChIP-seq and DNase-seq, CTCF-MP is able to make highly accurate predictions on whether a convergent CTCF motif pair would form a loop in a single cell type and also across different cell types. Our work represents an important step further to understand the sequence determinants that may guide the formation of complex chromatin architectures. Availability and implementation: The source code of CTCF-MP can be accessed at: https://github.com/ma-compbio/CTCF-MP. Supplementary information: Supplementary data are available at Bioinformatics online. Ruochi Zhang, Yang Yang 0094, Yang Zhang 0042, Jian Ma 0004 |
Bioinform. | 5 |
| 2017 | Speeding Up Latent Variable Gaussian Graphical Model Estimation via Nonconvex OptimizationabstractWe study the estimation of the latent variable Gaussian graphical model (LVGGM), where the precision matrix is the superposition of a sparse matrix and a low-rank matrix. In order to speed up the estimation of the sparse plus low-rank components, we propose a sparsity constrained maximum likelihood estimator based on matrix factorization and an efficient alternating gradient descent algorithm with hard thresholding to solve it. Our algorithm is orders of magnitude faster than the convex relaxation based methods for LVGGM. In addition, we prove that our algorithm is guaranteed to linearly converge to the unknown sparse and low-rank components up to the optimal statistical precision. Experiments on both synthetic and genomic data demonstrate the superiority of our algorithm over the state-of-the-art algorithms and corroborate our theory. Pan Xu 0002, Jian Ma 0004, Quanquan Gu |
NIPS | 2 |
| 2017 | Towards Recovering Allele-Specific Cancer Genome Graphs
Ashok Rajaraman, Jian Ma 0004 |
RECOMB | 2 |
| 2017 | Exploiting sequence-based features for predicting enhancer-promoter interactionsabstractMOTIVATION: A large number of distal enhancers and proximal promoters form enhancer-promoter interactions to regulate target genes in the human genome. Although recent high-throughput genome-wide mapping approaches have allowed us to more comprehensively recognize potential enhancer-promoter interactions, it is still largely unknown whether sequence-based features alone are sufficient to predict such interactions. RESULTS: Here, we develop a new computational method (named PEP) to predict enhancer-promoter interactions based on sequence-based features only, when the locations of putative enhancers and promoters in a particular cell type are given. The two modules in PEP (PEP-Motif and PEP-Word) use different but complementary feature extraction strategies to exploit sequence-based information. The results across six different cell types demonstrate that our method is effective in predicting enhancer-promoter interactions as compared to the state-of-the-art methods that use functional genomic signals. Our work demonstrates that sequence-based features alone can reliably predict enhancer-promoter interactions genome-wide, which could potentially facilitate the discovery of important sequence determinants for long-range gene regulation. AVAILABILITY AND IMPLEMENTATION: The source code of PEP is available at: https://github.com/ma-compbio/PEP . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Yang 0094, Ruochi Zhang, Shashank Singh 0005, Jian Ma 0004 |
Bioinform. | 4 |
| 2017 | Synteny Explorer: An Interactive Visualization Application for Teaching Genome EvolutionabstractRapid advances in biology demand new tools for more active research dissemination and engaged teaching. This paper presents Synteny Explorer, an interactive visualization application designed to let college students explore genome evolution of mammalian species. The tool visualizes synteny blocks: segments of homologous DNA shared between various extant species that can be traced back or reconstructed in extinct, ancestral species. We take a karyogram-based approach to create an interactive synteny visualization, leading to a more appealing and engaging design for undergraduate-level genome evolution education. For validation, we conduct three user studies: two focused studies on color and animation design choices and a larger study that performs overall system usability testing while comparing our karyogram-based designs with two more common genome mapping representations in an educational context. While existing views communicate the same information, study participants found the interactive, karyogram-based views much easier and likable to use. We additionally discuss feedback from biology and genomics faculty, who judge Synteny Explorer's fitness for use in classrooms. Chris Bryan, Gregory Guterman, Kwan-Liu Ma, Harris A. Lewin, Denis M. Larkin, Jaebum Kim, Jian Ma 0004, Marta Farre |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2016 | Allele-Specific Quantification of Structural Variations in Cancer Genomes
Shiguo Zhou, David C. Schwartz, Jian Ma 0004 |
RECOMB | 4 |
| 2016 | BLESS 2: accurate, memory-efficient and fast error correction methodabstractUNLABELLED: The most important features of error correction tools for sequencing data are accuracy, memory efficiency and fast runtime. The previous version of BLESS was highly memory-efficient and accurate, but it was too slow to handle reads from large genomes. We have developed a new version of BLESS to improve runtime and accuracy while maintaining a small memory usage. The new version, called BLESS 2, has an error correction algorithm that is more accurate than BLESS, and the algorithm has been parallelized using hybrid MPI and OpenMP programming. BLESS 2 was compared with five top-performing tools, and it was found to be the fastest when it was executed on two computing nodes using MPI, with each node containing twelve cores. Also, BLESS 2 showed at least 11% higher gain while retaining the memory efficiency of the previous version for large genomes. AVAILABILITY AND IMPLEMENTATION: Freely available at https://sourceforge.net/projects/bless-ec CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yun Heo, Anand Ramachandran 0001, Wen-Mei W. Hwu, Jian Ma 0004, Deming Chen |
Bioinform. | 4 |
| 2016 | A new correlation clustering method for cancer mutation analysisabstractMotivation: Cancer genomes exhibit a large number of different alterations that affect many genes in a diverse manner. An improved understanding of the generative mechanisms behind the mutation rules and their influence on gene community behavior is of great importance for the study of cancer. Results: To expand our capability to analyze combinatorial patterns of cancer alterations, we developed a rigorous methodology for cancer mutation pattern discovery based on a new, constrained form of correlation clustering. Our new algorithm, named C3 (Cancer Correlation Clustering), leverages mutual exclusivity of mutations, patient coverage and driver network concentration principles. To test C3, we performed a detailed analysis on TCGA breast cancer and glioblastoma data and showed that our algorithm outperforms the state-of-the-art CoMEt method in terms of discovering mutually exclusive gene modules and identifying biologically relevant driver genes. The proposed agnostic clustering method represents a unique tool for efficient and reliable identification of mutation patterns and driver pathways in large-scale cancer genomics studies, and it may also be used for other clustering problems on biological graphs. Availability and Implementation: The source code for the C3 method can be found at https://github.com/jackhou2/C3 Contacts: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Jack P. Hou, Amin Emad, Gregory J. Puleo, Jian Ma 0004, Olgica Milenkovic |
Bioinform. | 4 |
| 2016 | Reconstructing ancestral gene orders with duplications guided by synteny level genome reconstructionabstractBACKGROUND: Reconstructing ancestral gene orders in the presence of duplications is important for a better understanding of genome evolution. Current methods for ancestral reconstruction are limited by either computational constraints or the availability of reliable gene trees, and often ignore duplications altogether. Recently, methods that consider duplications in ancestral reconstructions have been developed, but the quality of reconstruction, counted as the number of contiguous ancestral regions found, decreases rapidly with the number of duplicated genes, complicating the application of such approaches to mammalian genomes. However, such high fragmentation is not encountered when reconstructing mammalian genomes at the synteny-block level, although the relative positions of genes in such reconstruction cannot be recovered. RESULTS: We propose a new heuristic method, MULTIRES, to reconstruct ancestral gene orders with duplications guided by homologous synteny blocks for a set of related descendant genomes. The method uses a synteny-level reconstruction to break the gene-order problem into several subproblems, which are then combined in order to disambiguate duplicated genes. We applied this method to both simulated and real data. Our results showed that MULTIRES outperforms other methods in terms of gene content, gene adjacency, and common interval recovery. CONCLUSIONS: This work demonstrates that the inclusion of synteny-level information can help us obtain better gene-level reconstructions. Our algorithm provides a basic toolbox for reconstructing ancestral gene orders with duplications. The source code of MULTIRES is available on https://github.com/ma-compbio/MultiRes . Ashok Rajaraman, Jian Ma 0004 |
BMC Bioinform. | 2 |
| 2015 | FPGA accelerated DNA error correction
Anand Ramachandran 0001, Yun Heo, Wen-Mei W. Hwu, Jian Ma 0004, Deming Chen |
DATE | 4 |
| 2014 | BLESS: Bloom filter-based error correction solution for high-throughput sequencing readsabstractMOTIVATION: Rapid advances in next-generation sequencing (NGS) technology have led to exponential increase in the amount of genomic information. However, NGS reads contain far more errors than data from traditional sequencing methods, and downstream genomic analysis results can be improved by correcting the errors. Unfortunately, all the previous error correction methods required a large amount of memory, making it unsuitable to process reads from large genomes with commodity computers. RESULTS: We present a novel algorithm that produces accurate correction results with much less memory compared with previous solutions. The algorithm, named BLoom-filter-based Error correction Solution for high-throughput Sequencing reads (BLESS), uses a single minimum-sized Bloom filter, and is also able to tolerate a higher false-positive rate, thus allowing us to correct errors with a 40× memory usage reduction on average compared with previous methods. Meanwhile, BLESS can extend reads like DNA assemblers to correct errors at the end of reads. Evaluations using real and simulated reads showed that BLESS could generate more accurate results than existing solutions. After errors were corrected using BLESS, 69% of initially unaligned reads could be aligned correctly. Additionally, de novo assembly results became 50% longer with 66% fewer assembly errors. AVAILABILITY AND IMPLEMENTATION: Freely available at http://sourceforge.net/p/bless-ec CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yun Heo, Deming Chen, Jian Ma 0004, Wen-Mei W. Hwu |
Bioinform. | 4 |
| 2014 | PSAR-Align: improving multiple sequence alignment using probabilistic samplingabstractSUMMARY: We developed PSAR-Align, a multiple sequence realignment tool that can refine a given multiple sequence alignment based on suboptimal alignments generated by probabilistic sampling. Our evaluation demonstrated that PSAR-Align is able to improve the results from various multiple sequence alignment tools. AVAILABILITY AND IMPLEMENTATION: The PSAR-Align source code (implemented mainly in C++) is freely available for download at http://bioen-compbio.bioen.illinois.edu/PSAR-Align. Jaebum Kim, Jian Ma 0004 |
Bioinform. | 2 |
| 2014 | A network-assisted co-clustering algorithm to discover cancer subtypes based on gene expressionabstractBACKGROUND: Cancer subtype information is critically important for understanding tumor heterogeneity. Existing methods to identify cancer subtypes have primarily focused on utilizing generic clustering algorithms (such as hierarchical clustering) to identify subtypes based on gene expression data. The network-level interaction among genes, which is key to understanding the molecular perturbations in cancer, has been rarely considered during the clustering process. The motivation of our work is to develop a method that effectively incorporates molecular interaction networks into the clustering process to improve cancer subtype identification. RESULTS: We have developed a new clustering algorithm for cancer subtype identification, called "network-assisted co-clustering for the identification of cancer subtypes" (NCIS). NCIS combines gene network information to simultaneously group samples and genes into biologically meaningful clusters. Prior to clustering, we assign weights to genes based on their impact in the network. Then a new weighted co-clustering algorithm based on a semi-nonnegative matrix tri-factorization is applied. We evaluated the effectiveness of NCIS on simulated datasets as well as large-scale Breast Cancer and Glioblastoma Multiforme patient samples from The Cancer Genome Atlas (TCGA) project. NCIS was shown to better separate the patient samples into clinically distinct subtypes and achieve higher accuracy on the simulated datasets to tolerate noise, as compared to consensus hierarchical clustering. CONCLUSIONS: The weighted co-clustering approach in NCIS provides a unique solution to incorporate gene network information into the clustering process. Our tool will be useful to comprehensively identify cancer subtypes that would otherwise be obscured by cancer heterogeneity, using high-throughput and high-dimensional gene expression data. Yiyi Liu, Quanquan Gu, Jack P. Hou, Jiawei Han 0001, Jian Ma 0004 |
BMC Bioinform. | 5 |
| 2014 | Tracing the Evolution of Lineage-Specific Transcription Factor Binding Sites in a Birth-Death FrameworkabstractChanges in cis-regulatory element composition that result in novel patterns of gene expression are thought to be a major contributor to the evolution of lineage-specific traits. Although transcription factor binding events show substantial variation across species, most computational approaches to study regulatory elements focus primarily upon highly conserved sites, and rely heavily upon multiple sequence alignments. However, sequence conservation based approaches have limited ability to detect lineage-specific elements that could contribute to species-specific traits. In this paper, we describe a novel framework that utilizes a birth-death model to trace the evolution of lineage-specific binding sites without relying on detailed base-by-base cross-species alignments. Our model was applied to analyze the evolution of binding sites based on the ChIP-seq data for six transcription factors (GATA1, SOX2, CTCF, MYC, MAX, ETS1) along the lineage toward human after human-mouse common ancestor. We estimate that a substantial fraction of binding sites (∼58-79% for each factor) in humans have origins since the divergence with mouse. Over 15% of all binding sites are unique to hominids. Such elements are often enriched near genes associated with specific pathways, and harbor more common SNPs than older binding sites in the human genome. These results support the ability of our method to identify lineage-specific regulatory elements and help understand their roles in shaping variation in gene regulation across species. Ken Daigoro Yokoyama, Yang Zhang 0042, Jian Ma 0004 |
PLoS Comput. Biol. | 3 |
| 2012 | TrueSight: Self-training Algorithm for Splice Junction Detection Using RNA-seq
Hong-Mei Li, Paul Burns, Mark Borodovsky, Gene E. Robinson, Jian Ma 0004 |
RECOMB | 6 |
| 2012 | TIGER: tiled iterative genome assemblerabstractBACKGROUND: With the cost reduction of the next-generation sequencing (NGS) technologies, genomics has provided us with an unprecedented opportunity to understand fundamental questions in biology and elucidate human diseases. De novo genome assembly is one of the most important steps to reconstruct the sequenced genome. However, most de novo assemblers require enormous amount of computational resource, which is not accessible for most research groups and medical personnel. RESULTS: We have developed a novel de novo assembly framework, called Tiger, which adapts to available computing resources by iteratively decomposing the assembly problem into sub-problems. Our method is also flexible to embed different assemblers for various types of target genomes. Using the sequence data from a human chromosome, our results show that Tiger can achieve much better NG50s, better genome coverage, and slightly higher errors, as compared to Velvet and SOAPdenovo, using modest amount of memory that are available in commodity computers today. CONCLUSIONS: Most state-of-the-art assemblers that can achieve relatively high assembly quality need excessive amount of computing resource (in particular, memory) that is not available to most researchers to achieve high quality results. Tiger provides the only known viable path to utilize NGS de novo assemblers that require more memory than that is present in available computers. Evaluation results demonstrate the feasibility of getting better quality results with low memory footprint and the scalability of using distributed commodity computers. Yun Heo, Izzat El Hajj, Wen-Mei W. Hwu, Deming Chen, Jian Ma 0004 |
BMC Bioinform. | 6 |
| 2011 | PSAR: Measuring Multiple Sequence Alignment Reliability by Probabilistic Sampling - (Extended Abstract)
Jaebum Kim, Jian Ma 0004 |
RECOMB | 2 |
| 2011 | FusionHunter: identifying fusion transcripts in cancer using paired-end RNA-seqabstractMOTIVATION: Fusion transcripts can be created as a result of genome rearrangement in cancer. Some of them play important roles in carcinogenesis, and can serve as diagnostic and therapeutic targets. With more and more cancer genomes being sequenced by next-generation sequencing technologies, we believe an efficient tool for reliably identifying fusion transcripts will be desirable for many groups. RESULTS: We designed and implemented an open-source software tool, called FusionHunter, which reliably identifies fusion transcripts from transcriptional analysis of paired-end RNA-seq. We show that FusionHunter can accurately detect fusions that were previously confirmed by RT-PCR in a publicly available dataset. The purpose of FusionHunter is to identify potential fusions with high sensitivity and specificity and to guide further functional validation in the laboratory. AVAILABILITY: http://bioen-compbio.bioen.illinois.edu/FusionHunter/. Jeremy Chien, David I. Smith, Jian Ma 0004 |
Bioinform. | 4 |
| 2010 | A probabilistic framework for inferring ancestral genomic ordersabstractWe introduce a probabilistic framework for inferring contiguous ancestral regions. Our previous work, the inferCARs algorithm, is a method based on adjacencies between synteny blocks. However, the local parsimony procedure has the limitation that it ignores many adjacencies that are potentially possible to exist in the ancestors. In this paper, we introduce a probabilistic method for reconstructing ancestral orders. The essential part of this method is to predict the posterior probability of an adjacency occurring in the ancestor based on an extended Jukes-Cantor model for breakpoints. We implemented a program called inferCARsPro to reconstruct contiguous ancestral regions. Both simulation and real data application results are discussed. Jian Ma 0004 |
BIBM | 1 |
| 2010 | Cactus Graphs for Genome Comparisons
Benedict Paten, Mark Diekhans, Dent Earl, John St. John, Jian Ma 0004, Bernard B. Suh, David Haussler |
RECOMB | 5 |