VLDB 2026 Research / reviewers in the wild / expert
Shuaicheng Li 0001
dblp:58/3258 · also Shuai Cheng Li 0001
· DBLP profile ↗
67ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0001-6246-6349ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 5 first-author · 22 since 2021Theory of computation · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoMBCR: Co-Learning Multi-Modalities of BCRs and gene expressionsabstractMOTIVATION: B-cell receptors (BCRs) and gene expression profiles are two distinct yet complementary modalities of B cells. However, most analyses treat them independently. Here, we present CoMBCR, a B-cell embedding tool that co-learns BCRs and gene expressions, representing data within a unified latent space for downstream analysis. RESULTS: We applied CoMBCR to 126,791 B cells from diverse datasets with matched BCRs and gene expressions. First, CoMBCR outperforms the methods solely encoding BCRs in capturing B-cell biological features, achieving at least 0.1 improvement in Matthews Correlation Coefficient on a SARS-CoV-2 binding prediction task. Second, CoMBCR reveals active immune responses and CDR3 motif preferences through modality gap analysis in SARS-CoV-2-specific memory B cells. Moreover, when supported by spatial transcriptomics data, CoMBCR accurately traces the developmental trajectories of malignant B cells and uncovers transcriptional patterns associated with their survival within lymphoma patients. AVAILABILITY AND IMPLEMENTATION: The CoMBCR software is publicly available under the MIT License at https://github.com/deepomicslab/CoMBCR.git. CONTACT: [email protected]. Yiping Zou, Jiaqi Luo, Shuaicheng Li 0001 |
Bioinform. | 3 |
| 2025 | Automated Prediction of Protein Pair Distances Based on Deep Learning: A Novel Approach to Protein Structure Prediction in SCOP
Duo Feng, Shuaicheng Li 0001, Yue Zhang 0050 |
ISBRA (2) | 2 |
| 2025 | A comprehensive benchmarking for evaluating TCR embeddings in modeling TCR-epitope interactionsabstractThe complexity of T cell receptor (TCR) sequences, particularly within the complementarity-determining region 3 (CDR3), requires efficient embedding methods for applying machine learning to immunology. While various TCR CDR3 embedding strategies have been proposed, the absence of their systematic evaluations created perplexity in the community. Here, we extracted CDR3 embedding models from 19 existing methods and benchmarked these models with four curated datasets by accessing their impact on the performance of TCR downstream tasks, including TCR-epitope binding affinity prediction, epitope-specific TCR identification, TCR clustering, and visualization analysis. We assessed these models utilizing eight downstream classifiers and five downstream clustering methods, with the performance measured by a diverse range of metrics for precision, robustness, and usability. Overall, handcrafted embeddings outperformed data-driven ones in modeling TCR-epitope interactions. To further refine our comparative findings, we developed an all-in-one TCR CDR3 embedding package comprising all evaluated embedding models. This package will assist users in easily selecting suitable embedding models for their data. Xikang Feng, Miaozhe Huo, Yongze Yang, Yuepeng Jiang, Shuaicheng Li 0001 |
Briefings Bioinform. | 7 |
| 2025 | Comprehensive human respiratory genome catalogue underlies the high resolution and precision of the respiratory microbiomeabstractThe human respiratory microbiome plays a crucial role in respiratory health, but there is no comprehensive respiratory genome catalogue (RGC) for studying the microbiome. In this study, we collected whole-metagenome shotgun sequencing data from 4067 samples and sequenced long reads of 124 samples, yielding 9.08 and 0.42 Tbp of short- and long-read data, respectively. By submitting these data with a novel assembly algorithm, we obtained a comprehensive human RGC. This high-quality RGC contains 190,443 contigs over 1 kbps and an N50 length exceeding 13 kbps; it comprises 159 high-quality and 393 medium-quality genomes, including 117 previously uncharacterized respiratory bacteria. Moreover, the RGC contains 209 respiratory-specific species not captured by the unified human gastrointestinal genome. Using the RGC, we revisited a study on a pediatric pneumonia dataset and identified 17 pneumonia-specific respiratory pathogens, reversing an inaccurate etiological conclusion due to the previous incomplete reference. Furthermore, we applied the RGC to the data of 62 participants with a clinical diagnosis of infection. Compared to the Nucleotide database, the RGC yielded greater specificity (0 versus 0.444, respectively) and sensitivity (0.852 versus 0.881, respectively), suggesting that the RGC provides superior sensitivity and specificity for the clinical diagnosis of respiratory diseases. Yinhu Li, Guangze Pan, Shuai Wang 0036, Zhengtu Li, Yiqi Jiang, Shuaicheng Li 0001, Bairong Shen |
Briefings Bioinform. | 8 |
| 2025 | MEGA-GO: functions prediction of diverse protein sequence length using Multi-scalE Graph Adaptive neural networkabstractMOTIVATION: The increasing accessibility of large-scale protein sequences through advanced sequencing technologies has necessitated the development of efficient and accurate methods for predicting protein function. Computational prediction models have emerged as a promising solution to expedite the annotation process. However, despite making significant progress in protein research, graph neural networks face challenges in capturing long-range structural correlations and identifying critical residues in protein graphs. Furthermore, existing models have limitations in effectively predicting the function of newly sequenced proteins that are not included in protein interaction networks. This highlights the need for novel approaches integrating protein structure and sequence data. RESULTS: We introduce Multi-scalE Graph Adaptive neural network (MEGA-GO), highlighting the capability of capturing diverse protein sequence length features from multiple scales. The unique graph adaptive neural network architecture of MEGA-GO enables a more nuanced extraction of graph structure features, effectively capturing intricate relationships within biological data. Experimental results demonstrate that MEGA-GO outperforms mainstream protein function prediction models in the accuracy of Gene Ontology term classification, yielding 33.4%, 68.9%, and 44.6% of area under the precision-recall curve on biological process, molecular function, and cellular component domains, respectively. The rest of the experimental results reveal that our model consistently surpasses the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The source code and data of MEGA-GO are available at https://github.com/Cheliosoops/MEGA-GO. Yujian Lee, Yongqi Xu, Shuaicheng Li 0001 |
Bioinform. | 5 |
| 2025 | AffMB: affinity maturation analysis with SHM-guided B-cell lineage treesabstractMOTIVATION: B-cell lineage trees describe the evolutionary process of immunoglobulin genes during affinity maturation. Existing methods for building B-cell lineage trees generally do not guarantee the parent-to-child inheritance and accumulation of advantageous mutations under successive rounds of somatic hypermutation (SHM) and selection, and are often incompatible with repertoire input. RESULTS: To address previous limitations, we developed AffMB (Affinity Maturation of B-cell receptor), a comprehensive toolkit for tracking affinity maturation through the generation and visualization of SHM-ordered, inheritance-based B-cell lineage trees from single-cell or bulk B-cell receptor sequencing data. The SHM-ordered inheritance tree algorithm outperformed state-of-the-art benchmarks in simulations. When applied to single-cell data from BNT162b2 vaccination (n = 42), AffMB demonstrated the ability to infer immunization responses and showed the feasibility of identifying potential high-affinity antibody sequences. AVAILABILITY AND IMPLEMENTATION: AffMB is an open-source Python package that supports contig FASTA or AIRR rearrangement TSV inputs. The source code for AffMB is freely available at https://github.com/deepomicslab/AffMB. Jiaqi Luo, Yiping Zou, Shuaicheng Li 0001 |
Bioinform. | 3 |
| 2024 | DeepTAPE: Enhancing Systemic Lupus Erythematosus Diagnosis with Deep Learning Based on TCRβ CDR3 SequencesabstractSystemic Lupus Erythematosus (SLE) is a common and severe autoimmune disease driven by abnormal T cell responses, which are critical for understanding the disease’s immune pathology. Given the central role of the Complementarity Determining Region 3 (CDR3) of the TCRβ chain in T cell specificity, focused investigation into CDR3 may potentially enhance both the diagnostic accuracy and the mechanistic understanding of SLE. In this study, we developed DeepTAPE, a deep learning-based engine for predicting autoimmune diseases. It primarily uses CDR3 sequence features, supplemented by the frequency distribution of TCRβ’s V genes. This model is based on a multi-layer CNN-LSTM architecture with residual connections. DeepTAPE exhibits superior diagnostic performance for SLE when compared to existing gene-focused studies. It achieves an impressive average AUC of 97.99%, accuracy of 93.97%, and recall of 94.36%, reaffirming the diagnostic utility of CDR3 in SLE through deep learning. Building on this, we introduced a quantitative indicator based on the model, the Autoimmune Risk Score, which is positively correlated with clinical disease activity in SLE patients and assists in prognosis. Overall, DeepTAPE not only offers accurate diagnosis but also enhances our understanding of SLE. Tongfei Shen, Miaozhe Huo, Wan Nie, Kaiqi Li, Xikang Feng, Shuaicheng Li 0001 |
BIBM | 7 |
| 2024 | F2Key: Dynamically Converting Your Face into a Private Key Based on COTS Headphones for Reliable Voice InteractionabstractIn this paper, we proposed F2Key, the first earable physical security system based on commercial off-the-shelf headphones. F2Key enables impactful applications, such as enhancing voiceprint-based authentication systems, reliable voice assistants, audio deepfake defense, and the legal validity of artifacts. The key idea of F2Key is to establish a stable acoustic sensing field across the user's face and embed the user's facial structures and articulatory habits into a user-specific generative model that serves as a private key. The private key can decrypt the Channel Impulse Response (CIR) profiles provided by the acoustic sensing field into an inferred spectrogram that can match the real one calculated from the corresponding speech, provided that the user's CIR-spectrogram mapping relationship is consistent with the one embedded in the generative model. Extensive experiments demonstrate that F2Key resists 99.9%, 96.4%, and 95.3% of speech replay attacks, mimicry attacks, and hybrid attacks, respectively. We discussed and evaluated F2Key from different perspectives, such as the health consideration and identical twins study, to show the practicality and reliability. Di Duan, Zehua Sun, Tao Ni 0003, Shuaicheng Li 0001, Xiaohua Jia, Weitao Xu, Tianxing Li 0001 |
MobiSys | 4 |
| 2024 | Elevated incidence of somatic mutations at prevalent genetic sitesabstractThe common loci represent a distinct set of the human genome sites that harbor genetic variants found in at least 1% of the population. Small somatic mutations occur at the common loci and non-common loci, i.e. csmVariants and ncsmVariants, are presumed with similar probabilities. However, our work revealed that within the coding region, common loci constituted only 1.03% of all loci, yet they accounted for 5.14% of TCGA somatic mutations. Furthermore, the small somatic mutation incidence rate at these common loci was 2.7 times that observed in the non-common. Notably, the csmVariants exhibited an impressive recurrent rate of 36.14%, which was 2.59 times of the ncsmVariants. The C-to-T transition at the CpG sites accounted for 32.41% of the csmVariants, which was 2.93 times for the ncsmVariants. Interestingly, the aging-related mutational signature contributed to 13.87% of the csmVariants, 5.5 times that of ncsmVariants. Moreover, 35.93% of the csmVariants contexts exhibited palindromic features, outperforming ncsmVariant contexts by 1.84 times. Notably, cancer patients with higher csmVariants rates had better progression-free survival. Furthermore, cancer patients with high-frequency csmVariants enriched with mismatch repair deficiency were also associated with better progression-free survival. The accumulation of csmVariants during cancerogenesis is a complex process influenced by various factors. These include the presence of a substantial percentage of palindromic sequences at csmVariants sites, the impact of aging and DNA mismatch repair deficiency. Together, these factors contribute to the higher somatic mutation incidence rates of common loci and the overall accumulation of csmVariants in cancer development. Shuaicheng Li 0001, Bairong Shen |
Briefings Bioinform. | 2 |
| 2024 | Coding genomes with gapped pattern graph convolutional networkabstractMOTIVATION: Genome sequencing technologies reveal a huge amount of genomic sequences. Neural network-based methods can be prime candidates for retrieving insights from these sequences because of their applicability to large and diverse datasets. However, the highly variable lengths of genome sequences severely impair the presentation of sequences as input to the neural network. Genetic variations further complicate tasks that involve sequence comparison or alignment. RESULTS: Inspired by the theory and applications of "spaced seeds," we propose a graph representation of genome sequences called "gapped pattern graph." These graphs can be transformed through a Graph Convolutional Network to form lower-dimensional embeddings for downstream tasks. On the basis of the gapped pattern graphs, we implemented a neural network model and demonstrated its performance on diverse tasks involving microbe and mammalian genome data. Our method consistently outperformed all the other state-of-the-art methods across various metrics on all tasks, especially for the sequences with limited homology to the training data. In addition, our model was able to identify distinct gapped pattern signatures from the sequences. AVAILABILITY AND IMPLEMENTATION: The framework is available at https://github.com/deepomicslab/GCNFrame. Yen Kaow Ng, Xianglilan Zhang, Jianping Wang 0001, Shuaicheng Li 0001 |
Bioinform. | 5 |
| 2023 | TEINet: a deep learning framework for prediction of TCR-epitope binding specificityabstractThe adaptive immune response to foreign antigens is initiated by T-cell receptor (TCR) recognition on the antigens. Recent experimental advances have enabled the generation of a large amount of TCR data and their cognate antigenic targets, allowing machine learning models to predict the binding specificity of TCRs. In this work, we present TEINet, a deep learning framework that utilizes transfer learning to address this prediction problem. TEINet employs two separately pretrained encoders to transform TCR and epitope sequences into numerical vectors, which are subsequently fed into a fully connected neural network to predict their binding specificities. A major challenge for binding specificity prediction is the lack of a unified approach to sampling negative data. Here, we first assess the current negative sampling approaches comprehensively and suggest that the Unified Epitope is the most suitable one. Subsequently, we compare TEINet with three baseline methods and observe that TEINet achieves an average AUROC of 0.760, which outperforms baseline methods by 6.4-26%. Furthermore, we investigate the impacts of the pretraining step and notice that excessive pretraining may lower its transferability to the final prediction task. Our results and analysis show that TEINet can make an accurate prediction using only the TCR sequence (CDR3$\beta $) and the epitope sequence, providing novel insights to understand the interactions between TCRs and epitopes. Yuepeng Jiang, Miaozhe Huo, Shuaicheng Li 0001 |
Briefings Bioinform. | 3 |
| 2023 | Deep autoregressive generative models capture the intrinsics embedded in T-cell receptor repertoiresabstractT-cell receptors (TCRs) play an essential role in the adaptive immune system. Probabilistic models for TCR repertoires can help decipher the underlying complex sequence patterns and provide novel insights into understanding the adaptive immune system. In this work, we develop TCRpeg, a deep autoregressive generative model to unravel the sequence patterns of TCR repertoires. TCRpeg largely outperforms state-of-the-art methods in estimating the probability distribution of a TCR repertoire, boosting the average accuracy from 0.672 to 0.906 measured by the Pearson correlation coefficient. Furthermore, with promising performance in probability inference, TCRpeg improves on a range of TCR-related tasks: profiling TCR repertoire probabilistically, classifying antigen-specific TCRs, validating previously discovered TCR motifs, generating novel TCRs and augmenting TCR data. Our results and analysis highlight the flexibility and capacity of TCRpeg to extract TCR sequence information, providing a novel approach for deciphering complex immunogenomic repertoires. Yuepeng Jiang, Shuaicheng Li 0001 |
Briefings Bioinform. | 2 |
| 2023 | IMperm: a fast and comprehensive IMmune Paired-End Reads Merger for sequencing dataabstractThe adaptive immune receptor repertoire (AIRR), consisting of T- and B-cell receptors, is the core component of the immune system. The AIRR sequencing is commonly used in cancer immunotherapy and minimal residual disease (MRD) detection of leukemia and lymphoma. The AIRR is captured by primers and sequenced to yield paired-end (PE) reads. The PE reads could be merged into one sequence by the overlapped region between them. However, the wide range of AIRR data raises the difficulty, so a special tool is required. We developed a software package for IMmune PE reads merger of sequencing data, named IMperm. We used the k-mer-and-vote strategy to pin down the overlapped region rapidly. IMperm could handle all types of PE reads, eliminate adapter contamination and successfully merge low-quality and minor/non-overlapping reads. Compared with existing tools, IMperm performed better in both simulated and sequencing data. Notably, IMperm was well suited to processing the data of MRD detection in leukemia and lymphoma and detected 19 novel MRD clones in 14 patients with leukemia from previously published data. Additionally, IMperm can handle PE reads from other sources, and we demonstrated its effectiveness on two genomic and one cell-free deoxyribonucleic acid datasets. IMperm is implemented in the C programming language and consumes little runtime and memory. It is freely available at https://github.com/zhangwei2015/IMperm. Wei Zhang 0178, Jia Ju, Teng Xiong, Chaohui Li, Shixin Lu, Zefeng Lu, Liya Lin, Shuaicheng Li 0001 |
Briefings Bioinform. | 11 |
| 2023 | On triangle inequalities of correlation-based distances for gene expression profilesabstractBACKGROUND: Distance functions are fundamental for evaluating the differences between gene expression profiles. Such a function would output a low value if the profiles are strongly correlated-either negatively or positively-and vice versa. One popular distance function is the absolute correlation distance, [Formula: see text], where [Formula: see text] is similarity measure, such as Pearson or Spearman correlation. However, the absolute correlation distance fails to fulfill the triangle inequality, which would have guaranteed better performance at vector quantization, allowed fast data localization, as well as accelerated data clustering. RESULTS: In this work, we propose [Formula: see text] as an alternative. We prove that [Formula: see text] satisfies the triangle inequality when [Formula: see text] represents Pearson correlation, Spearman correlation, or Cosine similarity. We show [Formula: see text] to be better than [Formula: see text], another variant of [Formula: see text] that satisfies the triangle inequality, both analytically as well as experimentally. We empirically compared [Formula: see text] with [Formula: see text] in gene clustering and sample clustering experiment by real-world biological data. The two distances performed similarly in both gene clustering and sample clustering in hierarchical clustering and PAM (partitioning around medoids) clustering. However, [Formula: see text] demonstrated more robust clustering. According to the bootstrap experiment, [Formula: see text] generated more robust sample pair partition more frequently (P-value [Formula: see text]). The statistics on the time a class "dissolved" also support the advantage of [Formula: see text] in robustness. CONCLUSION: [Formula: see text], as a variant of absolute correlation distance, satisfies the triangle inequality and is capable for more robust clustering. Yen Kaow Ng, Xianglilan Zhang, Shuaicheng Li 0001 |
BMC Bioinform. | 5 |
| 2022 | Somatic variant analysis suite: copy number variation clonal visualization online platform for large-scale single-cell genomicsabstractThe recent advance of single-cell copy number variation (CNV) analysis plays an essential role in addressing intratumor heterogeneity, identifying tumor subgroups and restoring tumor-evolving trajectories at single-cell scale. Informative visualization of copy number analysis results boosts productive scientific exploration, validation and sharing. Several single-cell analysis figures have the effectiveness of visualizations for understanding single-cell genomics in published articles and software packages. However, they almost lack real-time interaction, and it is hard to reproduce them. Moreover, existing tools are time-consuming and memory-intensive when they reach large-scale single-cell throughputs. We present an online visualization platform, single-cell Somatic Variant Analysis Suite (scSVAS), for real-time interactive single-cell genomics data visualization. scSVAS is specifically designed for large-scale single-cell genomic analysis that provides an arsenal of unique functionalities. After uploading the specified input files, scSVAS deploys the online interactive visualization automatically. Users may conduct scientific discoveries, share interactive visualizations and download high-quality publication-ready figures. scSVAS provides versatile utilities for managing, investigating, sharing and publishing single-cell CNV profiles. We envision this online platform will expedite the biological understanding of cancer clonal evolution in single-cell resolution. All visualizations are publicly hosted at https://sc.deepomics.org. Lingxi Chen, Yuhao Qing, Ruikang Li, Chaohui Li, Hechen Li, Xikang Feng, Shuaicheng Li 0001 |
Briefings Bioinform. | 7 |
| 2022 | Resolving single-cell copy number profiling for large datasetsabstractThe advances of single-cell DNA sequencing (scDNA-seq) enable us to characterize the genetic heterogeneity of cancer cells. However, the high noise and low coverage of scDNA-seq impede the estimation of copy number variations (CNVs). In addition, existing tools suffer from intensive execution time and often fail on large datasets. Here, we propose SeCNV, an efficient method that leverages structural entropy, to profile the copy numbers. SeCNV adopts a local Gaussian kernel to construct a matrix, depth congruent map (DCM), capturing the similarities between any two bins along the genome. Then, SeCNV partitions the genome into segments by minimizing the structural entropy from the DCM. With the partition, SeCNV estimates the copy numbers within each segment for cells. We simulate nine datasets with various breakpoint distributions and amplitudes of noise to benchmark SeCNV. SeCNV achieves a robust performance, i.e. the F1-scores are higher than 0.95 for breakpoint detections, significantly outperforming state-of-the-art methods. SeCNV successfully processes large datasets (>50 000 cells) within 4 min, while other tools fail to finish within the time limit, i.e. 120 h. We apply SeCNV to single-nucleus sequencing datasets from two breast cancer patients and acoustic cell tagmentation sequencing datasets from eight breast cancer patients. SeCNV successfully reproduces the distinct subclones and infers tumor heterogeneity. SeCNV is available at https://github.com/deepomicslab/SeCNV. Yuwei Zhang 0005, Mengbo Wang 0001, Xikang Feng, Jianping Wang 0001, Shuaicheng Li 0001 |
Briefings Bioinform. | 6 |
| 2022 | DeepHost: phage host prediction with convolutional neural networkabstractNext-generation sequencing expands the known phage genomes rapidly. Unlike culture-based methods, the hosts of phages discovered from next-generation sequencing data remain uncharacterized. The high diversity of the phage genomes makes the host assignment task challenging. To solve the issue, we proposed a phage host prediction tool-DeepHost. To encode the phage genomes into matrices, we design a genome encoding method that applied various spaced $k$-mer pairs to tolerate sequence variations, including insertion, deletions, and mutations. DeepHost applies a convolutional neural network to predict host taxonomies. DeepHost achieves the prediction accuracy of 96.05% at the genus level (72 taxonomies) and 90.78% at the species level (118 taxonomies), which outperforms the existing phage host prediction tools by 10.16-30.48% and achieves comparable results to BLAST. For the genomes without hits in BLAST, DeepHost obtains the accuracy of 38.00% at the genus level and 26.47% at the species level, making it suitable for genomes of less homologous sequences with the existing datasets. DeepHost is alignment-free, and it is faster than BLAST, especially for large datasets. DeepHost is available at https://github.com/deepomicslab/DeepHost. Xianglilan Zhang, Jianping Wang 0001, Shuaicheng Li 0001 |
Briefings Bioinform. | 4 |
| 2022 | Correction: On triangle inequalities of correlation-based distances for gene expression profilesabstractCorrection: BMC Bioinformatics (2023) 24:40 https://doi.org/10.1186/s12859-023-05161-y Following publication of the original article [1], it was reported that the article entitled “On triangle inequalities of correlation-based distances for gene expression profiles” was published in the regular issue of this journal instead of in the supplement issue. The details of the supplement in which this article ought to have been published are given below: This article has been published as part of BMC Bioinformatics Volume 23 Supplement 3, 2022: Selected articles from the International Conference on Intelligent Biology and Medicine (ICIBM 2021): bioinformatics. The full contents of the supplement are available online at https://bmcbioinformatics.biomedcentral.com/articles/supplements/volume-23-supplement-3. The publisher apologizes for any inconvenience caused. © 2023, The Author(s). Yen Kaow Ng, Xianglilan Zhang, Shuaicheng Li 0001 |
BMC Bioinform. | 5 |
| 2022 | DENSEN: a convolutional neural network for estimating chronological ages from panoramic radiographsabstractBACKGROUND: Age estimation from panoramic radiographs is a fundamental task in forensic sciences. Previous age assessment studies mainly focused on juvenile rather than elderly populations (> 25 years old). Most proposed studies were statistical or scoring-based, requiring wet-lab experiments and professional skills, and suffering from low reliability. RESULT: Based on Soft Stagewise Regression Network (SSR-Net), we developed DENSEN to estimate the chronological age for both juvenile and older adults, based on their orthopantomograms (OPTs, also known as orthopantomographs, pantomograms, or panoramic radiographs). We collected 1903 clinical panoramic radiographs of individuals between 3 and 85 years old to train and validate the model. We evaluated the model by the mean absolute error (MAE) between the estimated age and ground truth. For different age groups, 3-11 (children), 12-18 (teens), 19-25 (young adults), and 25+ (adults), DENSEN produced MAEs as 0.6885, 0.7615, 1.3502, and 2.8770, respectively. Our results imply that the model works in situations where genders are unknown. Moreover, DENSEN has lower errors for the adult group (> 25 years) than other methods. The proposed model is memory compact, consuming about 1.0 MB of memory overhead. CONCLUSIONS: We introduced a novel deep learning approach DENSEN to estimate a subject's age from a panoramic radiograph for the first time. Our approach required less laboratory work compared with existing methods. The package we developed is an open-source tool and applies to all different age groups. Yanle Liu, Xinyao Miao, Xiao Cao, Shuaicheng Li 0001 |
BMC Bioinform. | 7 |
| 2022 | Correction: DENSEN: a convolutional neural network for estimating chronological ages from panoramic radiographs
Yanle Liu, Xinyao Miao, Xiao Cao, Shuaicheng Li 0001 |
BMC Bioinform. | 7 |
| 2021 | Resolving complex structures at oncovirus integration loci with conjugate graphabstractOncovirus integrations cause copy number variations and complex structural variations (SVs) on host genomes. However, the understanding of how inserted viral DNA impacts the local genome remains limited. The linear structure of the oncovirus integrated local genomic map (LGM) will lay the foundations to understand how oncovirus integrations emerge and compromise the host genome's functioning. We propose a conjugate graph model to reconstruct the rearranged LGM at integrated loci. Simulation tests prove the reliability and credibility of the algorithm. Applications of the algorithm to whole-genome sequencing data of human papillomavirus (HPV) and hepatitis B virus (HBV)-infected cancer samples gained biological insights on oncovirus integrations. We observed four affection patterns of oncovirus integrations from the HPV and HBV-integrated cancer samples, including the coding-frame truncation, hyper-amplification of tumor gene, the viral cis-regulation inserted at the single intron and at the intergenic region. We found that the focal duplicates and host SVs are frequent in the HPV-integrated LGMs, while the focal deletions are prevalent in HBV-integrated LGMs. Furthermore, with the results yields from our method, we found the enhanced microhomology-mediated end joining might lead to both HPV and HBV integrations and conjectured that the HPV integrations might mainly occur during the DNA replication process. The conjugate graph algorithm code and LGM construction pipeline, available at https://github.com/deepomicslab/FuseSV. Wenlong Jia, Chang Xu 0021, Shuaicheng Li 0001 |
Briefings Bioinform. | 3 |
| 2021 | Deep learning model reveals potential risk genes for ADHD, especially Ephrin receptor gene EPHA5abstractAttention deficit hyperactivity disorder (ADHD) is a common neurodevelopmental disorder. Although genome-wide association studies (GWAS) identify the risk ADHD-associated variants and genes with significant P-values, they may neglect the combined effect of multiple variants with insignificant P-values. Here, we proposed a convolutional neural network (CNN) to classify 1033 individuals diagnosed with ADHD from 950 healthy controls according to their genomic data. The model takes the single nucleotide polymorphism (SNP) loci of P-values $\le{1\times 10^{-3}}$, i.e. 764 loci, as inputs, and achieved an accuracy of 0.9018, AUC of 0.9570, sensitivity of 0.8980 and specificity of 0.9055. By incorporating the saliency analysis for the deep learning network, a total of 96 candidate genes were found, of which 14 genes have been reported in previous ADHD-related studies. Furthermore, joint Gene Ontology enrichment and expression Quantitative Trait Loci analysis identified a potential risk gene for ADHD, EPHA5 with a variant of rs4860671. Overall, our CNN deep learning model exhibited a high accuracy for ADHD classification and demonstrated that the deep learning model could capture variants' combining effect with insignificant P-value, while GWAS fails. To our best knowledge, our model is the first deep learning method for the classification of ADHD with SNPs data. Xikang Feng, Haimei Li, Shuaicheng Li 0001, Qiujin Qian |
Briefings Bioinform. | 4 |
| 2021 | PStrain: an iterative microbial strains profiling algorithm for shotgun metagenomic sequencing dataabstractMOTIVATION: The microbial community plays an essential role in human diseases and physiological activities. The functions of microbes can differ due to strain-level differences in the genome sequences. Shotgun metagenomic sequencing allows us to profile the strains in microbial communities practically. However, current methods are underdeveloped due to the highly similar sequences among strains. We observe that strains genotypes at the same single nucleotide variant (SNV) locus can be speculated by the genotype frequencies. Also, the variants in different loci covered by the same reads can provide evidence that they reside on the same strain. RESULTS: These insights inspire us to design PStrain, an optimization method that utilizes genotype frequencies and the reads which cover multiple SNV loci to profile strains iteratively based on SNVs in a set of MetaPhlAn2 marker genes. Compared to the state-of-art methods, PStrain, on average, improved the performance of inferring strains abundances and genotypes by 87.75% and 59.45%, respectively. We have applied the PStrain package to the dataset with two cohorts of colorectal cancer (CRC) and found that the sequences of Bacteroides coprocola strains are significantly different between CRC and control samples, which is the first time to report the potential role of B.coprocola in the gut microbiota of CRC. AVAILABILITYAND IMPLEMENTATION: https://github.com/wshuai294/PStrain. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuai Wang 0036, Yiqi Jiang, Shuaicheng Li 0001 |
Bioinform. | 3 |
| 2020 | LDscaff: LD-based scaffolding of de novo genome assembliesabstractBACKGROUND: Genome assembly is fundamental for de novo genome analysis. Hybrid assembly, utilizing various sequencing technologies increases both contiguity and accuracy. While such approaches require extra costly sequencing efforts, the information provided millions of existed whole-genome sequencing data have not been fully utilized to resolve the task of scaffolding. Genetic recombination patterns in population data indicate non-random association among alleles at different loci, can provide physical distance signals to guide scaffolding. RESULTS: In this paper, we propose LDscaff for draft genome assembly incorporating linkage disequilibrium information in population data. We evaluated the performance of our method with both simulated data and real data. We simulated scaffolds by splitting the pig reference genome and reassembled them. Gaps between scaffolds were introduced ranging from 0 to 100 KB. The genome misassembly rate is 2.43% when there is no gap. Then we implemented our method to refine the Giant Panda genome and the donkey genome, which are purely assembled by NGS data. After LDscaff treatment, the resulting Panda assembly has scaffold N50 of 3.6 MB, 2.5 times larger than the original N50 (1.3 MB). The re-assembled donkey assembly has an improved N50 length of 32.1 MB from 23.8 MB. CONCLUSIONS: Our method effectively improves the assemblies with existed re-sequencing data, and is an potential alternative to the existing assemblers required for the collection of new data. Zicheng Zhao, Yingxiao Zhou, Shuai Wang 0036, Xiuqing Zhang, Changfa Wang, Shuaicheng Li 0001 |
BMC Bioinform. | 6 |
| 2019 | LysoPhD: predicting functional prophages in bacterial genomes from high-throughput sequencingabstractMotivation: When a temperate bacteriophage (phage) integrates into the chromosome of its host or to remain as latent episomal DNA, the prophage plays an important role in bacterial virulence acquisition and increase. With the recent release of thousands of bacterial high throughput sequencing data, there has been a growing interest in analyzing the impact of interdependency between a prophage and its host. One of the essential tasks is to acquire and identify the complete sequence of the prophage, i.e., the functional prophages. However, no existing tools can identify functional prophages from bacterial genomes. A few tools can predict putative prophage in bacterial genomes, but they are unable to distinguish functional prophages from defective prophages. To reduce the cost and to relieve the tedious biological experiments for functional prophage analysis, computational methods for predicting functional prophages based on bacterial high-throughput sequencing data would be an effective alternation. Results: This paper presents a method, LysoPhD, the first tool to predict functional prophage sequences using bacterial HTS data automatically and accurately. The main idea is to obtain whole genome of functional prophages from the bacterial genome based on phage gene annotation and reads-mediated contig connection. LysoPhD obtains whole genome sequences of functional prophages based on improved models. The obtained the sequences are verified by wet lab experiments. Specifically, LysoPhD is applied to a set of staphylococcus HTS data in the case studies. Ten functional prophages are predicted from 72 staphylococcus isolates. The 72 bacterial strains are then induced with Mitomycin C and subsequent deep sequencing verified of the 10 functional prophages. This shows that LysoPhD is a sensitive and accurate program for functional prophage prediction. Qi Niu, Shaoliang Peng, Xianglilan Zhang, Shuaicheng Li 0001, Xiang-Cheng Xie, Yi-Gang Tong |
BIBM | 4 |
| 2019 | DeepMF: deciphering the latent patterns in omics profiles with a deep learning methodabstractBACKGROUND: With recent advances in high-throughput technologies, matrix factorization techniques are increasingly being utilized for mapping quantitative omics profiling matrix data into low-dimensional embedding space, in the hope of uncovering insights in the underlying biological processes. Nevertheless, current matrix factorization tools fall short in handling noisy data and missing entries, both deficiencies that are often found in real-life data. RESULTS: Here, we propose DeepMF, a deep neural network-based factorization model. DeepMF disentangles the association between molecular feature-associated and sample-associated latent matrices, and is tolerant to noisy and missing values. It exhibited feasible cancer subtype discovery efficacy on mRNA, miRNA, and protein profiles of medulloblastoma cancer, leukemia cancer, breast cancer, and small-blue-round-cell cancer, achieving the highest clustering accuracy of 76%, 100%, 92%, and 100% respectively. When analyzing data sets with 70% missing entries, DeepMF gave the best recovery capacity with silhouette values of 0.47, 0.6, 0.28, and 0.44, outperforming other state-of-the-art MF tools on the cancer data sets Medulloblastoma, Leukemia, TCGA BRCA, and SRBCT. Its embedding strength as measured by clustering accuracy is 88%, 100%, 84%, and 96% on these data sets, which improves on the current best methods 76%, 100%, 78%, and 87%. CONCLUSION: DeepMF demonstrated robust denoising, imputation, and embedding ability. It offers insights to uncover the underlying biological processes such as cancer subtype discovery. Our implementation of DeepMF can be found at https://github.com/paprikachan/DeepMF. Lingxi Chen, Shuaicheng Li 0001 |
BMC Bioinform. | 3 |
| 2019 | MIRIA: a webserver for statistical, visual and meta-analysis of RNA editing data in mammalsabstractBACKGROUND: Adenosine-to-inosine RNA editing can markedly diversify the transcriptome, leading to a variety of critical molecular and biological processes in mammals. Over the past several years, researchers have developed several new pipelines and software packages to identify RNA editing sites with a focus on downstream statistical analysis and functional interpretation. RESULTS: Here, we developed a user-friendly public webserver named MIRIA that integrates statistics and visualization techniques to facilitate the comprehensive analysis of RNA editing sites data identified by the pipelines and software packages. MIRIA is unique in that provides several analytical functions, including RNA editing type statistics, genomic feature annotations, editing level statistics, genome-wide distribution of RNA editing sites, tissue-specific analysis and conservation analysis. We collected high-throughput RNA sequencing (RNA-seq) data from eight tissues across seven species as the experimental data for MIRIA and constructed an example result page. CONCLUSION: MIRIA provides both visualization and analysis of mammal RNA editing data for experimental biologists who are interested in revealing the functions of RNA editing sites. MIRIA is freely available at https://mammal.deepomics.org. Xikang Feng, Zishuai Wang, Hechen Li, Shuaicheng Li 0001 |
BMC Bioinform. | 4 |
| 2019 | LEMON: a method to construct the local strains at horizontal gene transfer sites in gut metagenomicsabstractBACKGROUND: Horizontal Gene Transfer (HGT) refers to the transfer of genetic materials between organisms through mechanisms other than parent-offspring inheritance. HGTs may affect human health through a large number of microorganisms, especially the gut microbiomes which the human body harbors. The transferred segments may lead to complicated local genome structural variations. Details of the local genome structure can elucidate the effects of the HGTs. RESULTS: In this work, we propose a graph-based method to reconstruct the local strains from the gut metagenomics data at the HGT sites. The method is implemented in a package named LEMON. The simulated results indicate that the method can identify transferred segments accurately on reference sequences of the microbiome. Simulation results illustrate that LEMON could recover local strains with complicated structure variation. Furthermore, the gene fusion points detected in real data near HGT breakpoints validate the accuracy of LEMON. Some strains reconstructed by LEMON have a replication time profile with lower standard error, which demonstrates HGT events recovered by LEMON is reliable. CONCLUSIONS: Through LEMON we could reconstruct the sequence structure of bacteria, which harbors HGT events. This helps us to study gene flow among different microbial species. Chen Li 0023, Yiqi Jiang, Shuaicheng Li 0001 |
BMC Bioinform. | 3 |
| 2019 | A unified STR profiling system across multiple species with whole genome sequencing dataabstractAbstract Background Short tandem repeats (STRs) serve as genetic markers in forensic scenes due to their high polymorphism in eukaryotic genomes. A variety of STRs profiling systems have been developed for species including human, dog, cat, cattle, etc. Maintaining these systems simultaneously can be costly. These mammals share many high similar regions along their genomes. With the availability of the massive amount of the whole genomics data of these species, it is possible to develop a unified STR profiling system. In this study, our objective is to propose and develop a unified set of STR loci that could be simultaneously applied to multiple species. Result To find a unified STR set, we collected the whole genome sequence data of the concerned species and mapped them to the human genome reference. Then we extracted the STR loci across the species. From these loci, we proposed an algorithm which selected a subset of loci by incorporating the optimized combined power of discrimination. Our results show that the unified set of loci have high combined power of discrimination, >1−10−9, for both individual species and the mixed population, as well as the random-match probability, <10−7 for all the involved species, indicating that the identified set of STR loci could be applied to multiple species. Conclusions We identified a set of STR loci which shared by multiple species. It implies that a unified STR profiling system is possible for these species under the forensic scenes. The system can be applied to the individual identification or paternal test of each of the ten common species which are Sus scrofa (pig), Bos taurus (cattle), Capra hircus (goat), Equus caballus (horse), Canis lupus familiaris (dog), Felis catus (cat), Ovis aries (sheep), Oryctolagus cuniculus (rabbit), and Bos grunniens (yak), and Homo sapiens (human). Our loci selection algorithm employed a greedy approach. The algorithm can generate the loci under different forensic parameters and for a specific combination of species. Miaoxia Chen, Changfa Wang, Shuaicheng Li 0001 |
BMC Bioinform. | 5 |
| 2019 | SpliceFinder: ab initio prediction of splice sites using convolutional neural networkabstractBACKGROUND: Identifying splice sites is a necessary step to analyze the location and structure of genes. Two dinucleotides, GT and AG, are highly frequent on splice sites, and many other patterns are also on splice sites with important biological functions. Meanwhile, the dinucleotides occur frequently at the sequences without splice sites, which makes the prediction prone to generate false positives. Most existing tools select all the sequences with the two dimers and then focus on distinguishing the true splice sites from those pseudo ones. Such an approach will lead to a decrease in false positives; however, it will result in non-canonical splice sites missing. RESULT: We have designed SpliceFinder based on convolutional neural network (CNN) to predict splice sites. To achieve the ab initio prediction, we used human genomic data to train our neural network. An iterative approach is adopted to reconstruct the dataset, which tackles the data unbalance problem and forces the model to learn more features of splice sites. The proposed CNN obtains the classification accuracy of 90.25%, which is 10% higher than the existing algorithms. The method outperforms other existing methods in terms of area under receiver operating characteristics (AUC), recall, precision, and F1 score. Furthermore, SpliceFinder can find the exact position of splice sites on long genomic sequences with a sliding window. Compared with other state-of-the-art splice site prediction tools, SpliceFinder generates results in about half lower false positive while keeping recall higher than 0.8. Also, SpliceFinder captures the non-canonical splice sites. In addition, SpliceFinder performs well on the genomic sequences of Drosophila melanogaster, Mus musculus, Rattus, and Danio rerio without retraining. CONCLUSION: Based on CNN, we have proposed a new ab initio splice site prediction tool, SpliceFinder, which generates less false positives and can detect non-canonical splice sites. Additionally, SpliceFinder is transferable to other species without retraining. The source code and additional materials are available at https://gitlab.deepomics.org/wangruohan/SpliceFinder. Zishuai Wang, Jianping Wang 0001, Shuaicheng Li 0001 |
BMC Bioinform. | 4 |
| 2018 | Placing Segments on Parallel Arcs
Yen Kaow Ng, Wenlong Jia, Shuaicheng Li 0001 |
IWOCA | 3 |
| 2018 | In silico design of MHC class I high binding affinity peptides through motifs activation mapabstractBACKGROUND: Finding peptides with high binding affinity to Class I major histocompatibility complex (MHC-I) attracts intensive research, and it serves a crucial part of developing a better vaccine for precision medicine. Traditional methods cost highly for designing such peptides. The advancement of computational approaches reduces the cost of new drug discovery dramatically. Compared with flourishing computational drug discovery area, the immunology area lacks tools focused on in silico design for the peptides with high binding affinity. Attributed to the ever-expanding amount of MHC-peptides binding data, it enables the tremendous influx of deep learning techniques for modeling MHC-peptides binding. To leverage the availability of these data, it is of great significance to find MHC-peptides binding specificities. The binding motifs are one of the key components to decide the MHC-peptides combination, which generally refer to a combination of some certain amino acids at certain sites which highly contribute to the binding affinity. RESULT: In this work, we propose the Motif Activation Mapping (MAM) network for MHC-I and peptides binding to extract motifs from peptides. Then, we substitute amino acid randomly according to the motifs for generating peptides with high affinity. We demonstrated the MAM network could extract motifs which are the features of peptides of highly binding affinities, as well as generate peptides with high-affinities; that is, 0.859 for HLA-A*0201, 0.75 for HLA-A*0206, 0.92 for HLA-B*2702, 0.9 for HLA-A*6802 and 0.839 for Mamu-A1*001:01. Besides, its binding prediction result reaches the state of the art. The experiment also reveals the network is appropriate for most MHC-I with transfer learning. CONCLUSIONS: We design the MAM network to extract the motifs from MHC-peptides binding through prediction, which are proved to generate the peptides with high binding affinity successfully. The new peptides preserve the motifs but vary in sequences. Zhoujian Xiao, Yuwei Zhang 0005, Runsheng Yu, Xiaosen Jiang, Shuaicheng Li 0001 |
BMC Bioinform. | 7 |
| 2018 | Guest Editorial for the 15th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 15th Asia Pacific Bioinformatics Conference (APBC2017), which was held in Shenzhen, China, 17-19 January 2017. Lusheng Wang 0001, Shuaicheng Li 0001, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Improving protein fold recognition by extracting fold-specific features from predicted residue-residue contactsabstractMOTIVATION: Accurate recognition of protein fold types is a key step for template-based prediction of protein structures. The existing approaches to fold recognition mainly exploit the features derived from alignments of query protein against templates. These approaches have been shown to be successful for fold recognition at family level, but usually failed at superfamily/fold levels. To overcome this limitation, one of the key points is to explore more structurally informative features of proteins. Although residue-residue contacts carry abundant structural information, how to thoroughly exploit these information for fold recognition still remains a challenge. RESULTS: In this study, we present an approach (called DeepFR) to improve fold recognition at superfamily/fold levels. The basic idea of our approach is to extract fold-specific features from predicted residue-residue contacts of proteins using deep convolutional neural network (DCNN) technique. Based on these fold-specific features, we calculated similarity between query protein and templates, and then assigned query protein with fold type of the most similar template. DCNN has showed excellent performance in image feature extraction and image recognition; the rational underlying the application of DCNN for fold recognition is that contact likelihood maps are essentially analogy to images, as they both display compositional hierarchy. Experimental results on the LINDAHL dataset suggest that even using the extracted fold-specific features alone, our approach achieved success rate comparable to the state-of-the-art approaches. When further combining these features with traditional alignment-related features, the success rate of our approach increased to 92.3%, 82.5% and 78.8% at family, superfamily and fold levels, respectively, which is about 18% higher than the state-of-the-art approach at fold level, 6% higher at superfamily level and 1% higher at family level. An independent assessment on SCOP_TEST dataset showed consistent performance improvement, indicating robustness of our approach. Furthermore, bi-clustering results of the extracted features are compatible with fold hierarchy of proteins, implying that these features are fold-specific. Together, these results suggest that the features extracted from predicted contacts are orthogonal to alignment-related features, and the combination of them could greatly facilitate fold recognition at superfamily/fold levels and template-based prediction of protein structures. AVAILABILITY AND IMPLEMENTATION: Source code of DeepFR is freely available through https://github.com/zhujianwei31415/deepfr, and a web server is available through http://protein.ict.ac.cn/deepfr. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianwei Zhu, Haicang Zhang, Shuaicheng Li 0001, Lupeng Kong, Shiwei Sun, Wei-Mou Zheng, Dongbo Bu |
Bioinform. | 3 |
| 2016 | Finding more effective microsatellite markers for forensicsabstractPublished by the Combined DNA Index System (CODIS) program of the Federal Bureau of Investigation (FBI) in 1997, the 13 core short tandem repeat (STR) loci are widely adopted as genetic markers in forensic applications, e.g., identity testing and paternity testing. However, these loci may be biased and suffer from reduced sensitivities towards specific population groups. In addition, the rapid growth of entries in forensic databases raises the chance of random hits, which can cause false judgments of innocents as criminals. A solution to these problems is to introduce more effective STR markers. The availability of whole genome sequencing enables us to identify more reliable STR markers for forensic applications computationally. In this study, we proposed an algorithm to identify STR markers with high discriminative abilities from the next-generation sequencing (NGS) data. Our algorithm could select a customized set of loci for a given population with pre-specified discriminative thresholds. We have applied the method to 320 Chinese individuals from the 1,000 Genomes Project and obtained various numbers of loci, which were able to statistically identify an individual worldwide and had higher combined powers of discrimination (CPD) and combined probabilities of exclusion (CPE) than the existing CODIS 13 loci. For identity testing, the mean frequency of DNA profile (FDP) with the selected 11 STRs was smaller than that with CODIS 13 STRs by student's t-test. With more loci, much smaller FDPs were obtained. The database matching probabilities (DMP) for selected loci were also lower than that for CODIS 13 STRs in a database with 10 billion entries. Moreover, the selected loci were able to provide considerably low chance of random profile matches so that statistically no false judgments could occur. The selected loci also reduced the risk of random allele matches when doing the familial search, with lower random allele matching probabilities. In addition, the selected STRs were statistically better than CODIS STRs for paternity testing in our simulated data, with lower probabilities of false inclusions and exclusions. Bowen Tan, Zicheng Zhao, Shengbin Li, Shuaicheng Li 0001 |
BIBM | 5 |
| 2016 | More accurate models for detecting gene-gene interactions from public expression compendiaabstractThe fast accumulation of public gene expression data-made possible by high-throughput technology-provides us an unprecedented opportunity to identify functionally related genes by analyzing their co-expression patterns. However, these data are typically noisy and highly heterogeneous, complicating their use in constructing co-expression network in large expression compendia. Previous studies suggested that the collective gene expression pattern can be better modeled by Gaussian mixtures. This motivates our present work, which proposes Multimodal Mutual Information (MMI) to reconstruct gene co-expression network from public gene expression data. MMI assumes gene pair following bivariate Gaussian mixture models and categories the samples into unique bins with respect to their expression magnitude. Two kinds of correlations in MMI are computed and aggregated to capture both discretized dependency and the expression correlation for each bin. Through extensive simulations, MMI outperforms other approaches with respect to calculating gene-gene interactions, regardless of the level of noise or strength of interactions. The advance of MMI is further validated by three real problems: 1. Infer novel gene functions by their connections with the well-documented genes. We apply principle component analysis to the correlated matrix generated by MMI and Pearson correlation and construct transcriptional components to evaluate gene-gene interactions. MMI enables 1.7 times more eligible transcriptional components than Pearson correlation, which can be used to predict gene functions. 2. Prioritize candidate genes for an affected pedigree. MMI calculates the interactions between candidate and disease established genes and explores KIF1A as the new causal gene of pure hereditary spastic paraparesis. 3. Detect disease “hot genes”. MMI identifies ANK2 as the “hot gene” for autism spectrum disorders by evaluating its co-expression with other disease susceptible genes derived from trio-based exome sequencing data. Lu Zhang 0061, Shuaicheng Li 0001 |
BIBM | 3 |
| 2016 | Eliminating heterozygosity from reads through coverage normalizationabstractHeterozygosity has long plagued genome assembly. A wealth of sophisticated algorithms and procedures have been introduced for their treatment in genome assembly. In this paper, we propose a method called Hank (Heterozygosity assimilation through normalized k-mers) for this purpose. The method eliminates heterozygosity from reads by modifying the k-mers in heterozygous regions to increase their coverages to the levels of those in homozygous regions. When evaluated on simulated Illumina data at levels of heterozygosity from 0.1% to 2.0%, Hank was able to remove 80~96% of the heterozygosity. We also examined the effects of the corrections on de novo genome assembly using SOAPdenovo2, ALLPATH-LG and Platanus. All three methods improved in performance using the treated k-mers (we do not include these assembly results in this manuscript due to space constraint). Zicheng Zhao, Yen Kaow Ng, Xiaodong Fang, Shuaicheng Li 0001 |
BIBM | 4 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 9 |
| 2015 | Reconstructing directed gene regulatory network by only gene expression dataabstractAccurately identifying gene regulatory network serves an important task in understanding in vivo biological activities. The inference of such network is often accomplished through the use of gene expression data. Some methods further predict the regulatory directions in the network by using the location of eQTL single nucleotide polymorphisms, or through gene knock out/down experiments; regrettably, these additional data are not always available, especially for the samples deriving from human tissues. In this paper, we propose Context Based Dependency Network (CBDN), a method that is able to infer gene regulatory networks, complete with the regulatory directions, from only gene expression data. CBDN applies directed data processing inequality (DDPI) to distinguish between direct and transitive relationship between genes. In our experiments with simulated and real data, CBDN outperforms the current state-of-the-art approaches. When used to identify important regulators in a network, CBDN 1. correctly identified TYROBP in the network related to Alzheimer's disease; 2. predicted potential important regulators ZNF329 and RB1 for human brain tumors. Lu Zhang 0061, Yen Kaow Ng, Shuaicheng Li 0001 |
BIBM | 3 |
| 2015 | On the Near-Linear Correlation of the Eigenvalues Across BLOSUM Matrices
Yen Kaow Ng, Xingwu Liu, Shuaicheng Li 0001 |
ISBRA | 4 |
| 2015 | Finding All Longest Common Segments in Protein Structures EfficientlyabstractThe Local/Global Alignment (Zemla, 2003), or LGA, is a popular method for the comparison of protein structures. One of the two components of LGA requires us to compute the longest common contiguous segments between two protein structures. That is, given two structures A = (a1, ... ,a(n)) and B = (b1, ... ,b(n)) where a(k), b(k) ∈ ℝ(3), we are to find, among all the segments f = (a(i), ... ,a(j)) and g = (b(i), ... ,b(j)) that fulfill a certain criterion regarding their similarity, those of the maximum length. We consider the following criteria: (1) the root mean squared deviation (RMSD) between f and g is to be within a given t ∈ ℝ; (2) f and g can be superposed such that for each k, i ≤ k ≤ j, ||a(k) - b(k)|| ≤ t for a given t ∈ ℝ. We give an algorithm of O(n log n + nl) time complexity when the first requirement applies, where l is the maximum length of the segments fulfilling the criterion. We show an FPTAS which, for any ϵ ∈ ℝ, finds a segment of length at least l, but of RMSD up to (1 + ϵ)t, in O(n log n + n/ϵ) time. We propose an FPTAS which for any given ϵ ∈ R, finds all the segments f and g of the maximum length which can be superposed such that for each k, i ≤ k ≤ j, ||a(k) - b(k)|| ≤ (1 + ϵ)t, thus fulfilling the second requirement approximately. The algorithm has a time complexity of O(n log(2) n/ϵ(5)) when consecutive points in A are separated by the same distance (which is the case with protein structures). These worst-case runtime complexities are verified using C++ implementations of the algorithms, which we have made available at http://alcs.sourceforge.net/. Yen Kaow Ng, Linzhi Yin, Hirotaka Ono 0001, Shuaicheng Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2015 | Parameterized BLOSUM Matrices for Protein AlignmentabstractProtein alignment is a basic step for many molecular biology researches. The BLOSUM matrices, especially BLOSUM62, are the de facto standard matrices for protein alignments. However, after widely utilization of the matrices for 15 years, programming errors were surprisingly found in the initial version of source codes for their generation. And amazingly, after bug correction, the "intended" BLOSUM62 matrix performs consistently worse than the "miscalculated" one. In this paper, we find linear relationships among the eigenvalues of the matrices and propose an algorithm to find optimal unified eigenvectors. With them, we can parameterize matrix BLOSUMx for any given variable x that could change continuously. We compare the effectiveness of our parameterized isentropic matrix with BLOSUM62. Furthermore, an iterative alignment and matrix selection process is proposed to adaptively find the best parameter and globally align two sequences. Experiments are conducted on aligning 13,667 families of Pfam database and on clustering MHC II protein sequences, whose improved accuracy demonstrates the effectiveness of our proposed method. Dongbo Bu, Shuaicheng Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2014 | Fingerprinting protein structures effectively and efficientlyabstractMOTIVATION: One common task in structural biology is to assess the similarities and differences among protein structures. A variety of structure alignment algorithms and programs has been designed and implemented for this purpose. A major drawback with existing structure alignment programs is that they require a large amount of computational time, rendering them infeasible for pairwise alignments on large collections of structures. To overcome this drawback, a fragment alphabet learned from known structures has been introduced. The method, however, considers local similarity only, and therefore occasionally assigns high scores to structures that are similar only in local fragments. METHOD: We propose a novel approach that eliminates false positives, through the comparison of both local and remote similarity, with little compromise in speed. Two kinds of contact libraries (ContactLib) are introduced to fingerprint protein structures effectively and efficiently. Each contact group of the contact library consists of one local or two remote fragments and is represented by a concise vector. These vectors are then indexed and used to calculate a new combined hit-rate score to identify similar protein structures effectively and efficiently. RESULTS: We tested our method on the high-quality protein structure subset of SCOP30 containing 3297 protein structures. For each protein structure of the subset, we retrieved its neighbor protein structures from the rest of the subset. The best area under the Receiver-Operating Characteristic curve, archived by ContactLib, is as high as 0.960. This is a significant improvement compared with 0.747, the best result achieved by FragBag. We also demonstrated that incorporating remote contact information is critical to consistently retrieve accurate neighbor protein structures for all- query protein structures. AVAILABILITY AND IMPLEMENTATION: https://cs.uwaterloo.ca/∼xfcui/contactlib/. Xuefeng Cui, Shuaicheng Li 0001, Lin He 0002, Ming Li 0001 |
Bioinform. | 2 |
| 2014 | Quantifying Significance of MHC II ResiduesabstractThe major histocompatibility complex (MHC), a cell-surface protein mediating immune recognition, plays important roles in the immune response system of all higher vertebrates. MHC molecules are highly polymorphic and they are grouped into serotypes according to the specificity of the response. It is a common belief that a protein sequence determines its three dimensional structure and function. Hence, the protein sequence determines the serotype. Residues play different levels of importance. In this paper, we quantify the residue significance with the available serotype information. Knowing the significance of the residues will deepen our understanding of the MHC molecules and yield us a concise representation of the molecules. In this paper we propose a linear programming-based approach to find significant residue positions as well as quantifying their significance in MHC II DR molecules. Among all the residues in MHC II DR molecules, 18 positions are of particular significance, which is consistent with the literature on MHC binding sites, and succinct pseudo-sequences appear to be adequate to capture the whole sequence features. When the result is used for classification of MHC molecules with serotype assigned by WHO, a 98.4 percent prediction performance is achieved. The methods have been implemented in java (http://code.google.com/p/quassi/). Ruoshui Lu, Lusheng Wang 0001, Massimo Andreatta, Shuaicheng Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2014 | Residue-Specific Side-Chain Polymorphismsvia Particle Belief PropagationabstractProtein side chains populate diverse conformational ensembles in crystals. Despite much evidence that there is widespread conformational polymorphism in protein side chains, most of the X-ray crystallography data are modeled by single conformations in the Protein Data Bank. The ability to extract or to predict these conformational polymorphisms is of crucial importance, as it facilitates deeper understanding of protein dynamics and functionality. In this paper, we describe a computational strategy capable of predicting side-chain polymorphisms. Our approach extends a particular class of algorithms for side-chain prediction by modeling the side-chain dihedral angles more appropriately as continuous rather than discrete variables. Employing a new inferential technique known as particle belief propagation, we predict residue-specific distributions that encode information about side-chain polymorphisms. Our predicted polymorphisms are in relatively close agreement with results from a state-of-the-art approach based on X-ray crystallography data, which characterizes the conformational polymorphisms of side chains using electron density information, and has successfully discovered previously unmodeled conformations. Laleh Soltan Ghoraie, Forbes J. Burkowski, Shuaicheng Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Detecting Protein Conformational Changes in Interactions via Scaling Known Structures
Fei Guo 0001, Shuaicheng Li 0001, Wenji Ma, Lusheng Wang 0001 |
RECOMB | 2 |
| 2013 | Towards Reliable Automatic Protein Structure Alignment
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
WABI | 2 |
| 2012 | Finding Longest Common Segments in Protein Structures in Nearly Linear Time
Yen Kaow Ng, Hirotaka Ono 0001, Ling Ge, Shuaicheng Li 0001 |
CPM | 4 |
| 2012 | P-Binder: A System for the Protein-Protein Binding Sites Identification
Fei Guo 0001, Shuaicheng Li 0001, Lusheng Wang 0001 |
ISBRA | 2 |
| 2012 | LoopWeaver - Loop Modeling by the Weighted Scaling of Verified Proteins
Daniel Holtby, Shuaicheng Li 0001, Ming Li 0001 |
RECOMB | 2 |
| 2012 | How Accurately Can We Model Protein Structures with Dihedral Angles?
Xuefeng Cui, Shuaicheng Li 0001, Dongbo Bu, Babak Alipanahi, Ming Li 0001 |
WABI | 2 |
| 2012 | Protein-protein binding site identification by enumerating the configurationsabstractBACKGROUND: The ability to predict protein-protein binding sites has a wide range of applications, including signal transduction studies, de novo drug design, structure identification and comparison of functional sites. The interface in a complex involves two structurally matched protein subunits, and the binding sites can be predicted by identifying structural matches at protein surfaces. RESULTS: We propose a method which enumerates "all" the configurations (or poses) between two proteins (3D coordinates of the two subunits in a complex) and evaluates each configuration by the interaction between its components using the Atomic Contact Energy function. The enumeration is achieved efficiently by exploring a set of rigid transformations. Our approach incorporates a surface identification technique and a method for avoiding clashes of two subunits when computing rigid transformations. When the optimal transformations according to the Atomic Contact Energy function are identified, the corresponding binding sites are given as predictions. Our results show that this approach consistently performs better than other methods in binding site identification. CONCLUSIONS: Our method achieved a success rate higher than other methods, with the prediction quality improved in terms of both accuracy and coverage. Moreover, our method is being able to predict the configurations of two binding proteins, where most of other methods predict only the binding sites. The software package is available at http://sites.google.com/site/guofeics/dobi for non-commercial use. Fei Guo 0001, Shuaicheng Li 0001, Lusheng Wang 0001, Daming Zhu |
BMC Bioinform. | 2 |
| 2012 | Residues with Similar Hexagon Neighborhoods Share Similar Side-Chain ConformationsabstractWe present in this study a new approach to code protein side-chain conformations into hexagon substructures. Classical side-chain packing methods consist of two steps: first, side-chain conformations, known as rotamers, are extracted from known protein structures as candidates for each residue; second, a searching method along with an energy function is used to resolve conflicts among residues and to optimize the combinations of side chain conformations for all residues. These methods benefit from the fact that the number of possible side-chain conformations is limited, and the rotamer candidates are readily extracted; however, these methods also suffer from the inaccuracy of energy functions. Inspired by threading and Ab Initio approaches to protein structure prediction, we propose to use hexagon substructures to implicitly capture subtle issues of energy functions. Our initial results indicate that even without guidance from an energy function, hexagon structures alone can capture side-chain conformations at an accuracy of 83.8 percent, higher than 82.6 percent by the state-of-art side-chain packing methods. Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | Clustering 100, 000 Protein Structure Decoys in MinutesabstractAb initio protein structure prediction methods first generate large sets of structural conformations as candidates (called decoys), and then select the most representative decoys through clustering techniques. Classical clustering methods are inefficient due to the pairwise distance calculation, and thus become infeasible when the number of decoys is large. In addition, the existing clustering approaches suffer from the arbitrariness in determining a distance threshold for proteins within a cluster: a small distance threshold leads to many small clusters, while a large distance threshold results in the merging of several independent clusters into one cluster. In this paper, we propose an efficient clustering method through fast estimating cluster centroids and efficient pruning rotation spaces. The number of clusters is automatically detected by information distance criteria. A package named ONION, which can be downloaded freely, is implemented accordingly. Experimental results on benchmark data sets suggest that ONION is 14 times faster than existing tools, and ONION obtains better selections for 31 targets, and worse selection for 19 targets compared to SPICKER’s selections. On an average PC, ONION can cluster 100,000 decoys in around 12 minutes. Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Pedigree Reconstruction Using Identity by Descent
Bonnie Kirkpatrick, Shuaicheng Li 0001, Richard M. Karp, Eran Halperin |
RECOMB | 2 |
| 2011 | On protein structure alignment under distance constraint
Shuaicheng Li 0001, Yen Kaow Ng |
Theor. Comput. Sci. | 1 |
| 2010 | Calibur: a tool for clustering large numbers of protein decoysabstractBACKGROUND: Ab initio protein structure prediction methods generate numerous structural candidates, which are referred to as decoys. The decoy with the most number of neighbors of up to a threshold distance is typically identified as the most representative decoy. However, the clustering of decoys needed for this criterion involves computations with runtimes that are at best quadratic in the number of decoys. As a result currently there is no tool that is designed to exactly cluster very large numbers of decoys, thus creating a bottleneck in the analysis. RESULTS: Using three strategies aimed at enhancing performance (proximate decoys organization, preliminary screening via lower and upper bounds, outliers filtering) we designed and implemented a software tool for clustering decoys called Calibur. We show empirical results indicating the effectiveness of each of the strategies employed. The strategies are further fine-tuned according to their effectiveness.Calibur demonstrated the ability to scale well with respect to increases in the number of decoys. For a sample size of approximately 30 thousand decoys, Calibur completed the analysis in one third of the time required when the strategies are not used.For practical use Calibur is able to automatically discover from the input decoys a suitable threshold distance for clustering. Several methods for this discovery are implemented in Calibur, where by default a very fast one is used. Using the default method Calibur reported relatively good decoys in our tests. CONCLUSIONS: Calibur's ability to handle very large protein decoy sets makes it a useful tool for clustering decoys in ab initio protein structure prediction. As the number of decoys generated in these methods increases, we believe Calibur will come in important for progress in the field. Shuaicheng Li 0001, Yen Kaow Ng |
BMC Bioinform. | 1 |
| 2009 | On Protein Structure Alignment under Distance Constraint
Shuaicheng Li 0001, Yen Kaow Ng |
ISAAC | 1 |
| 2009 | Finding compact structural motifs
Dongbo Bu, Ming Li 0001, Shuaicheng Li 0001, Jianbo Qian, Jinbo Xu |
Theor. Comput. Sci. | 3 |
| 2009 | On two open problems of 2-interval patterns
Shuaicheng Li 0001, Ming Li 0001 |
Theor. Comput. Sci. | 1 |
| 2008 | Finding Largest Well-Predicted Subset of Protein Structure Models
Shuaicheng Li 0001, Dongbo Bu, Jinbo Xu, Ming Li 0001 |
CPM | 1 |
| 2008 | Designing succinct structural alphabetsabstractMOTIVATION: The 3D structure of a protein sequence can be assembled from the substructures corresponding to small segments of this sequence. For each small sequence segment, there are only a few more likely substructures. We call them the 'structural alphabet' for this segment. Classical approaches such as ROSETTA used sequence profile and secondary structure information, to predict structural fragments. In contrast, we utilize more structural information, such as solvent accessibility and contact capacity, for finding structural fragments. RESULTS: Integer linear programming technique is applied to derive the best combination of these sequence and structural information items. This approach generates significantly more accurate and succinct structural alphabets with more than 50% improvement over the previous accuracies. With these novel structural alphabets, we are able to construct more accurate protein structures than the state-of-art ab initio protein structure prediction programs such as ROSETTA. We are also able to reduce the Kolodny's library size by a factor of 8, at the same accuracy. AVAILABILITY: The online FRazor server is under construction. Shuaicheng Li 0001, Dongbo Bu, Xin Gao 0001, Jinbo Xu, Ming Li 0001 |
ISMB | 1 |
| 2007 | Finding Compact Structural Motifs
Jianbo Qian, Shuaicheng Li 0001, Dongbo Bu, Ming Li 0001, Jinbo Xu |
CPM | 2 |
| 2006 | On the Complexity of the Crossing Contact Map Pattern Matching Problem
Shuaicheng Li 0001, Ming Li 0001 |
WABI | 1 |
| 2005 | Indexing DNA Sequences Using q-Grams
Xia Cao, Shuaicheng Li 0001, Anthony K. H. Tung |
DASFAA | 2 |
| 2004 | New Approximation Algorithms for Some Dynamic Storage Allocation Problems
Shuaicheng Li 0001, Hon Wai Leong, Steven K. Quek |
COCOON | 1 |
| 2004 | String Join Using Precedence Count Matrix
Xia Cao, Anthony K. H. Tung, Beng Chin Ooi, Kian-Lee Tan, Shuaicheng Li 0001 |
SSDBM | 5 |