Chun-Yu Lin 0003

dblp:12/5315-3 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-2901-2681ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2024 Inference of single-cell network using mutual information for scRNA-seq data analysis
abstract
BACKGROUND: With the advance in single-cell RNA sequencing (scRNA-seq) technology, deriving inherent biological system information from expression profiles at a single-cell resolution has become possible. It has been known that network modeling by estimating the associations between genes could better reveal dynamic changes in biological systems. However, accurately constructing a single-cell network (SCN) to capture the network architecture of each cell and further explore cell-to-cell heterogeneity remains challenging. RESULTS: We introduce SINUM, a method for constructing the SIngle-cell Network Using Mutual information, which estimates mutual information between any two genes from scRNA-seq data to determine whether they are dependent or independent in a specific cell. Experiments on various scRNA-seq datasets with different cell numbers based on eight performance indexes (e.g., adjusted rand index and F-measure index) validated the accuracy and robustness of SINUM in cell type identification, superior to the state-of-the-art SCN inference method. Additionally, the SINUM SCNs exhibit high overlap with the human interactome and possess the scale-free property. CONCLUSIONS: SINUM presents a view of biological systems at the network level to detect cell-type marker genes/gene pairs and investigate time-dependent changes in gene associations during embryo development. Codes for SINUM are freely available at https://github.com/SysMednet/SINUM .
Lan-Yun Chang, Ting-Yi Hao, Chun-Yu Lin 0003
BMC Bioinform.4
2023 SWEET: a single-sample network inference method for deciphering individual features in disease
abstract
Recently, extracting inherent biological system information (e.g. cellular networks) from genome-wide expression profiles for developing personalized diagnostic and therapeutic strategies has become increasingly important. However, accurately constructing single-sample networks (SINs) to capture individual characteristics and heterogeneity in disease remains challenging. Here, we propose a sample-specific-weighted correlation network (SWEET) method to model SINs by integrating the genome-wide sample-to-sample correlation (i.e. sample weights) with the differential network between perturbed and aggregate networks. For a group of samples, the genome-wide sample weights can be assessed without prior knowledge of intrinsic subpopulations to address the network edge number bias caused by sample size differences. Compared with the state-of-the-art SIN inference methods, the SWEET SINs in 16 cancers more likely fit the scale-free property, display higher overlap with the human interactomes and perform better in identifying three types of cancer-related genes. Moreover, integrating SWEET SINs with a network proximity measure facilitates characterizing individual features and therapy in diseases, such as somatic mutation, mut-driver and essential genes. Biological experiments further validated two candidate repurposable drugs, albendazole for head and neck squamous cell carcinoma (HNSCC) and lung adenocarcinoma (LUAD) and encorafenib for HNSCC. By applying SWEET, we also identified two possible LUAD subtypes that exhibit distinct clinical features and molecular mechanisms. Overall, the SWEET method complements current SIN inference and analysis methods and presents a view of biological systems at the network level to offer numerous clues for further investigation and clinical translation in network medicine and precision medicine.
Hsin-Hua Chen, Chun-Wei Hsueh, Chia-Hwa Lee, Ting-Yi Hao, Tzu-Ying Tu, Lan-Yun Chang, Jih-Chin Lee, Chun-Yu Lin 0003
Briefings Bioinform.8
2023 Common Attractors in Multiple Boolean Networks
abstract
Analyzing multiple networks is important to understand relevant features among different networks. Although many studies have been conducted for that purpose, not much attention has been paid to the analysis of attractors (i.e., steady states) in multiple networks. Therefore, we study common attractors and similar attractors in multiple networks to uncover hidden similarities and differences among networks using Boolean networks (BNs), where BNs have been used as a mathematical model of genetic networks and neural networks. We define three problems on detecting common attractors and similar attractors, and theoretically analyze the expected number of such objects for random BNs, where we assume that given networks have the same set of nodes (i.e., genes). We also present four methods for solving these problems. Computational experiments on randomly generated BNs are performed to demonstrate the efficiency of our proposed methods. In addition, experiments on a practical biological system, a BN model of the TGF- β signaling pathway, are performed. The result suggests that common attractors and similar attractors are useful for exploring tumor heterogeneity and homogeneity in eight cancers.
Wenya Pi, Chun-Yu Lin 0003, Ulrike Münzner, Masahiro Ohtomo, Tatsuya Akutsu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 Weighted minimum feedback vertex sets and implementation in human cancer genes detection
abstract
BACKGROUND: Recently, many computational methods have been proposed to predict cancer genes. One typical kind of method is to find the differentially expressed genes between tumour and normal samples. However, there are also some genes, for example, 'dark' genes, that play important roles at the network level but are difficult to find by traditional differential gene expression analysis. In addition, network controllability methods, such as the minimum feedback vertex set (MFVS) method, have been used frequently in cancer gene prediction. However, the weights of vertices (or genes) are ignored in the traditional MFVS methods, leading to difficulty in finding the optimal solution because of the existence of many possible MFVSs. RESULTS: Here, we introduce a novel method, called weighted MFVS (WMFVS), which integrates the gene differential expression value with MFVS to select the maximum-weighted MFVS from all possible MFVSs in a protein interaction network. Our experimental results show that WMFVS achieves better performance than using traditional bio-data or network-data analyses alone. CONCLUSION: This method balances the advantage of differential gene expression analyses and network analyses, improves the low accuracy of differential gene expression analyses and decreases the instability of pure network analyses. Furthermore, WMFVS can be easily applied to various kinds of networks, providing a useful framework for data analysis and prediction.
Ruiming Li, Chun-Yu Lin 0003, Weifeng Guo, Tatsuya Akutsu
BMC Bioinform.2
2021 ReCGBM: a gradient boosting-based method for predicting human dicer cleavage sites
abstract
BACKGROUND: Human dicer is an enzyme that cleaves pre-miRNAs into miRNAs. Several models have been developed to predict human dicer cleavage sites, including PHDCleav and LBSizeCleav. Given an input sequence, these models can predict whether the sequence contains a cleavage site. However, these models only consider each sequence independently and lack interpretability. Therefore, it is necessary to develop an accurate and explainable predictor, which employs relations between different sequences, to enhance the understanding of the mechanism by which human dicer cleaves pre-miRNA. RESULTS: In this study, we develop an accurate and explainable predictor for human dicer cleavage site - ReCGBM. We design relational features and class features as inputs to a lightGBM model. Computational experiments show that ReCGBM achieves the best performance compared to the existing methods. Further, we find that features in close proximity to the center of pre-miRNA are more important and make a significant contribution to the performance improvement of the developed method. CONCLUSIONS: The results of this study show that ReCGBM is an interpretable and accurate predictor. Besides, the analyses of feature importance show that it might be of particular interest to consider more informative features close to the center of the pre-miRNA in future predictors.
Pengyu Liu 0002, Jiangning Song, Chun-Yu Lin 0003, Tatsuya Akutsu
BMC Bioinform.3
2018 Identification of the PCa28 Gene Signature as a Predictor in Prostate Cancer
abstract
Prostate cancer (PCa) is the second-leading cause of cancer death among men in the worldwide. Most PCa is slowly growing and usually early symptomless. About 70% of PCa patients were diagnosed at later stage and metastasis has been observed. Additionally, the cure rate of PCa closely relies on the early diagnosis with biomarkers. Prostatic Specific Antigen (PSA) is currently the only clinical biomarker for PCa diagnosis. However, the PSA test has inherent limitations and has about 75% of false-positive results. The identification of a set of genes (as biomarkers) for diagnosis and prognosis is an urgent clinical issue for PCa. Here, we integrated genome-wide analysis and protein-protein interaction network to identify potential genes for early diagnostic biomarkers of PCa. First, we collected gene expression datasets of 145 PCa samples, consisting of both tumor and corresponding normal tissues, from two different sources in Gene Expression Omnibus (GEO). We found 158 and 268 significantly highly and lowly expressed genes, respectively, in tumor samples. Moreover, we proposed cluster score (CS) and predicting score (PS) to select 28 prostate cancer-related genes (called PCa28). The results indicate that PCa28 can discriminate between the normal/tumor tissues and are specific for prostate cancer. Finally, we examined 8 genes in PCa28 on four PCa cell lines by real time quantitative polymerase chain reaction (RT-qPCR). Experimental results show that up-regulated genes have higher expression level in tumor cells in comparison to normal cells, and down-regulated genes have lower expression level in tumor cells. We believe that our method is useful and PCa28 are potential biomarkers that provide the clues to develop targeting therapy for PCa.
Jung-Yu Lee 0001, Si-Yu Lin, Yi-Hsuan Chuang, Sing-Han Huang, Yu-Yao Tseng, Chun-Yu Lin 0003, Hung-Jung Wang, Jinn-Moon Yang
BIBE6
2018 Deep Learning with Evolutionary and Genomic Profiles for Identifying Cancer Subtypes
abstract
Cancer subtype identification is an unmet need in precision diagnosis. Recently, evolutionary conservation has been indicated containing understandable signatures for functional significance in cancers. However, the importance of evolutionary conservation in distinguishing cancer subtypes remains unclear. Here, we identified the evolutionarily conserved genes (i.e., core gene) and observed that they are mainly involved in the pathways relevant to cell growth and metabolisms. By using these core genes, we integrated their evolutionary and genomic profiles with deep learning to develop a feature-based strategy (FES) and an image-based strategy (IMS). In comparison with FES using the random set and the strategy using the PAM50 classifier, core gene set-based FES has higher accuracy for identifying breast cancer subtypes. Moreover, the IMS with data augmentation yields better performance than the other strategies. Comprehensive analysis of eight TCGA cancer data demonstrates that our evolutionary conservation-based models provide a valid and helpful approach to identify cancer subtypes and the core gene set offers distinguishable clues of cancer subtypes.
Chun-Yu Lin 0003, Peiying Ruan, Ruiming Li, Jinn-Moon Yang, Simon See, Tatsuya Akutsu
BIBE1
2016 Finding Influential Genes Using Gene Expression Data and Boolean Models of Metabolic Networks
abstract
Selection of influential genes using gene expression data from normal and disease samples is an important topic in bioinformatics. In this paper, we propose a novel computational method for the problem, which combines gene expression patterns from normal and disease samples with a mathematical model of metabolic networks. This method seeks a set of k genes knockout of which drives the state of the metabolic network towards that in the disease samples. We adopt a Boolean model of metabolic networks and formulate the problem as a maximization problem under an integer linear programming framework. We applied the proposed method to selection of influential genes using gene expression data from normal samples and disease (head and neck cancer) samples. The result suggests that the proposed method can select more biologically relevant genes than an existing P-value based ranking method can.
Takeyuki Tamura, Tatsuya Akutsu, Chun-Yu Lin 0003, Jinn-Moon Yang
BIBE3
2013 Inferring homologous protein-protein interactions through pair position specific scoring matrix
abstract
BACKGROUND: The protein-protein interaction (PPI) is one of the most important features to understand biological processes. For a PPI, the physical domain-domain interaction (DDI) plays the key role for biology functions. In the post-genomic era, to rapidly identify homologous PPIs for analyzing the contact residue pairs of their interfaces within DDIs on a genomic scale is essential to determine PPI networks and the PPI interface evolution across multiple species. RESULTS: In this study, we proposed "pair Position Specific Scoring Matrix (pairPSSM)" to identify homologous PPIs. The pairPSSM can successfully distinguish the true protein complexes from unreasonable protein pairs with about 90% accuracy. For the test set including 1,122 representative heterodimers and 2,708,746 non-interacting protein pairs, the mean average precision and mean false positive rate of pairPSSM were 0.42 and 0.31, respectively. Moreover, we applied pairPSSM to identify ~450,000 homologous PPIs with their interacting domains and residues in seven common organisms (e.g. Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Escherichia coli). CONCLUSIONS: Our pairPSSM is able to provide statistical significance of residue pairs using evolutionary profiles and a scoring system for inferring homologous PPIs. According to our best knowledge, the pairPSSM is the first method for searching homologous PPIs across multiple species using pair position specific scoring matrix and a 3D dimer as the template to map interacting domain pairs of these PPIs. We believe that pairPSSM is able to provide valuable insights for the PPI evolution and networks across multiple species.
Chun-Yu Lin 0003, Yung-Chiang Chen, Yu-Shu Lo, Jinn-Moon Yang
BMC Bioinform.1