VLDB 2026 Research / reviewers in the wild / expert
Yaning Yang
dblp:83/4935
· DBLP profile ↗
29ranked-venue papers
7as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 12 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MSI-GNN: A Graph Neural Network for Pathogenicity Prediction of Microsatellite InsertionsabstractMicrosatellite insertions (MSIs) are a common type of genetic variation implicated in various hereditary disorders and cancers. However, due to their repetitive structure, high sequence variability, and limited annotation resources, the pathogenicity of MSIs remains difficult to determine in clinical settings, hindering their utility in genetic diagnosis and disease mechanism studies. To address this challenge, we propose MSIGNN, a graph neural network model that integrates multidimensional annotations and graph attention mechanisms to predict the pathogenicity of MSI events. MSI-GNN constructs a comprehensive feature profile for each MSI by incorporating heterogeneous annotations, including genomic functions, epigenetic signals, deleteriousness scores, functional constraints, evolutionary conservation, splicing effects, and molecular consequences. A multi-layer graph attention network is then employed to model biological correlations among variants and extract informative pathogenicity representations. Experimental results demonstrate that MSI-GNN significantly outperforms existing general-purpose models across multiple evaluation metrics, offering superior predictive performance and interpretability. Our model provides a promising tool for elucidating the pathogenic mechanisms of MSIs and advancing precision medicine. Yaning Yang, Yadong Fan, Liangrui Pan, Shaoliang Peng |
BIBM | 1 |
| 2025 | Parallel Acceleration of Genome Variation Detection on Multi-Zone Heterogeneous SystemabstractGenomic variation is critical for understanding the genetic basis of disease. Pindel, a widely used structural variant caller, leverages short-read sequencing data to detect variation at single-base resolution; however, its hotspot module imposes substantial computational demands, limiting efficiency in large-scale whole-genome analyses. Heterogeneous architectures offer a promising solution, yet disparities in hardware design and programming models preclude direct porting of the original algorithm. To address this, we introduce MTPindel, a novel heterogeneous parallel optimization framework tailored to the MT-3000 processor. Focusing on Pindel's most compute-intensive modules, we design multi-core and task-level parallel algorithms that exploit the MT-3000's accelerator domains to balance and accelerate workload distribution. On 128 MT-3000–equipped nodes of the Tianhe next-generation supercomputer, MTPindel achieves an impressive 122.549 times of speedup and 95.74% parallel efficiency, with only a 0.74% error margin relative to the original implementation. This work represents a pioneering effort in heterogeneous parallelization for variant detection, paving the way for rapid, large-scale genomic analyses in research and clinical settings. Yaning Yang, Chengqing Li, Shaoliang Peng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Multi-Objective Deep Reinforcement Learning for Function Offloading in Serverless Edge ComputingabstractFunction offloading problems play a crucial role in optimizing the performance of applications in serverless edge computing (SEC). Existing research has extensively explored function offloading strategies based on optimizing a single objective. However, a significant challenge arises when users expect to optimize multiple objectives according to the relative importance of these objectives. This challenge becomes particularly pronounced when the relative importance of the objectives dynamically shifts. Consequently, there is an urgent need for research into multi-objective function offloading methods. In this paper, we redefine the SEC function offloading problem as a dynamic multi-objective optimization issue and propose a novel approach based on Multi-objective Reinforcement Learning (MORL) called MOSEC. MOSEC can coordinately optimize three objectives, i.e., application completion time, User Device (UD) energy consumption, and user cost. To reduce the impact of extrapolation errors, MOSEC integrates a Near-on Experience Replay (NER) strategy during the model training. Furthermore, MOSEC adopts our proposed Earliest First (EF) scheme to maintain the policies learned previously, which can efficiently mitigate the catastrophic policy forgetting problem. Extensive experiments conducted on various generated applications demonstrate the superiority of MOSEC over state-of-the-art multi-objective optimization algorithms. Yaning Yang, Yutong Ye 0001, Jiepin Ding, Ting Wang 0001, Mingsong Chen 0001, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Energy-Efficient Shop Scheduling Using Space-Cooperation Multi-Objective OptimizationabstractSince Industry 5.0 emphasizes that manufacturing enterprises should raise awareness of social contribution to achieve sustainable development, more and more meta-heuristic algorithms are investigated to save energy in manufacturing systems. Although non-dominated sorting-based meta-heuristics have been recognized as promising multi-objective optimization methods for solving the energy-efficient flexible job shop scheduling problem (EFJSP), it is hard to guarantee the quality of the Pareto front (e.g., total energy consumption, makespan) due to the lack of population diversity. This is mainly because an improper individual comparison inevitably reduces population diversity, thus limiting exploration and exploitation abilities during population updates. To achieve efficient population evolution, this paper introduces a novel space-cooperation multi-objective optimization (SCMO) method that can effectively solve EFJSP to obtain scheduling schemes with better trade-offs. By cooperatively evaluating the similarity among individuals in both the decision space and objective space, we propose a space-cooperation population update method based on a three-vector representation that can accurately eliminate repetitive individuals to derive higher-quality Pareto solutions. To further improve search efficiency, we propose a difference-driven local search, which selectively changes the positions of operations with higher differences to search for neighbors effectively. Based on the Taguchi method, we conduct experiments to obtain a suitable parameter combination of SCMO. Comprehensive experimental results show that, compared to state-of-the-art methods, our SCMO method achieves the highest HV and NR and the lowest IGD, with an average of 0.990, 0.952, and 0.001, respectively. Meanwhile, compared to traditional local search approaches, our difference-driven local search obtains twice the HV on instance Mk12 and reduces the solving time from 1521 s to 475 s. Jiepin Ding, Jun Xia 0003, Yaning Yang, Junlong Zhou, Mingsong Chen 0001, Keqin Li 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2024 | Situation-Dependent Causal Influence-Based Cooperative Multi-Agent Reinforcement LearningabstractLearning to collaborate has witnessed significant progress in multi-agent reinforcement learning (MARL). However, promoting coordination among agents and enhancing exploration capabilities remain challenges. In multi-agent environments, interactions between agents are limited in specific situations. Effective collaboration between agents thus requires a nuanced understanding of when and how agents' actions influence others.To this end, in this paper, we propose a novel MARL algorithm named Situation-Dependent Causal Influence-Based Cooperative Multi-agent Reinforcement Learning (SCIC), which incorporates a novel Intrinsic reward mechanism based on a new cooperation criterion measured by situation-dependent causal influence among agents.Our approach aims to detect inter-agent causal influences in specific situations based on the criterion using causal intervention and conditional mutual information. This effectively assists agents in exploring states that can positively impact other agents, thus promoting cooperation between agents.The resulting update links coordinated exploration and intrinsic reward distribution, which enhance overall collaboration and performance.Experimental results on various MARL benchmarks demonstrate the superiority of our method compared to state-of-the-art approaches. Yutong Ye 0001, Yaning Yang, Mingsong Chen 0001, Ting Wang 0001 |
AAAI | 4 |
| 2024 | TranSVPath: A TabTransformer-Based Model for Predicting the Pathogenicity of Structural VariantsabstractGenomic structural variants are recognized as critical molecular factors contributing to various major diseases, including cancer and genetic disorders. However, accurately determining whether a variant can cause a disease in clinical settings is extremely challenging due to limited sample sizes, the diversity of variant types, and the complexity of the mechanisms linking variants to diseases. Existing computational tools attempt to predict the pathogenic effects of these variants, but they often consider only single-layer biological data, limiting their ability to comprehensively explain the functional impacts of the variants. To address this, we propose TranSVPath, a pathogenicity scoring tool for structural variants based on the Transformer framework. TranSVPath provides a more comprehensive biological annotation of structural variations by integrating multi-dimensional data, including overlap with specific genomic regions, single nucleotide variant deleteriousness scores, phylogenetic conservation scores, Mendelian clinical application pathogenicity scores, and gene function loss. Additionally, TranSVPath employs a variant of the Transformer architecture tailored to this integrated data. Its attention mechanism captures key biological features, facilitating molecular-level interpretation of the pathogenic mechanisms of variants. Consequently, TranSVPath can accurately predict the pathogenicity of deletions, insertions, tandem duplications, inversions, and microsatellite insertions. Evaluations on large-scale human structural variant datasets demonstrate that TranSVPath outperforms state-of-the-art algorithms across relevant metrics. Yaning Yang, Liangrui Pan, Shaoliang Peng |
BIBM | 1 |
| 2024 | MK-BMC: a Multi-Kernel framework with Boosted distance metrics for Microbiome data for ClassificationabstractMOTIVATION: Research on human microbiome has suggested associations with human health, opening opportunities to predict health outcomes using microbiome. Studies have also suggested that diverse forms of taxa such as rare taxa that are evolutionally related and abundant taxa that are evolutionally unrelated could be associated with or predictive of a health outcome. Although prediction models were developed for microbiome data, no prediction models currently exist that use multiple forms of microbiome-outcome associations. RESULTS: We developed MK-BMC, a Multi-Kernel framework with Boosted distance Metrics for Classification using microbiome data. We propose to first boost widely used distance metrics for microbiome data using taxon-level association signal strengths to up-weight taxa that are potentially associated with an outcome of interest. We then propose a multi-kernel prediction model with one kernel capturing one form of association between taxa and the outcome, where a kernel measures similarities of microbiome compositions between pairs of samples being transformed from a proposed boosted distance metric. We demonstrated superior prediction performance of (i) boosted distance metrics for microbiome data over original ones and (ii) MK-BMC over competing methods through extensive simulations. We applied MK-BMC to predict thyroid, obesity, and inflammatory bowel disease status using gut microbiome data from the American Gut Project and observed much-improved prediction performance over that of competing methods. The learned kernel weights help us understand contributions of individual microbiome signal forms nicely. AVAILABILITY AND IMPLEMENTATION: Source code together with a sample input dataset is available at https://github.com/HXu06/MK-BMC. Yuqi Miao, Min Qian 0002, Yaning Yang |
Bioinform. | 5 |
| 2023 | Brief Industry Paper: RTLight: Digital Twin-Based Real-Time Federated Traffic Signal ControlabstractAlthough Reinforcement Learning (RL)-based methods have been widely researched in Traffic Signal Control (TSC), they still suffer from the problems of poor adaptation to real-world traffic scenarios and slow convergence to optimized solutions. This is because RL-based TSC methods have a high dependency on accurate modeling of the environment. With transportation infrastructure constraints, some vehicle dynamic information in the road network is difficult to obtain in real-time, which strongly limits the capability of RL agents. To address this problem, we propose a novel real-time federated traffic signal control system named RTLight, which can efficiently control traffic lights in real-time for multi-intersection scenarios. Based on the digital twin, the RL agent can obtain sufficient traffic information and interact with the environment in real time. Inspired by federated learning, our system supports knowledge sharing among intersections, which improves the overall convergence rate and control performance. Note that, we have deployed our RTLight system for large-scale application validation in Xishan district, Wuxi, China. Experimental results obtained from various real-world traffic scenarios demonstrate that RTLight can significantly improve the control performance. Yutong Ye 0001, Zhiwei Ling, Yaning Yang, Xian Wei, Cheng Chen 0027, Mingsong Chen 0001 |
RTSS | 3 |
| 2023 | CPGL: Prediction of Compound-Protein Interaction by Integrating Graph Attention Network With Long Short-Term Memory Neural NetworkabstractRecent advancements of artificial intelligence based on deep learning algorithms have made it possible to computationally predict compound-protein interaction (CPI) without conducting laboratory experiments. In this manuscript, we integrated a graph attention network (GAT) for compounds and a long short-term memory neural network (LSTM) for proteins, used end-to-end representation learning for both compounds and proteins, and proposed a deep learning algorithm, CPGL (CPI with GAT and LSTM) to optimize the feature extraction from compounds and proteins and to improve the model robustness and generalizability. CPGL demonstrated an excellent predictive performance and outperforms recently reported deep learning models. Based on 3 public CPI datasets, C.elegans, Human and BindingDB, CPGL represented 1 - 5% improvement compared to existing deep-learning models. Our method also achieves excellent results on datasets with imbalanced positive and negative proportions constructed based on the C.elegans and Human datasets. More importantly, using 2 label reversal datasets, GPCR and Kinase, CPGL showed superior performance compared to other existing deep learning models. The AUC were substantially improved by 20% on the Kinase dataset, indicative of the robustness and generalizability of CPGL. Minghua Zhao, Yaning Yang, Xu Steven 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | MSVF: Multi-task Structure Variation Filter with Transfer Learning in High-throughput SequencingabstractThe single molecule real-time sequencing technologies, such as PacBio and Nanopore, have higher throughput and produce longer reads, which promote the discovery of more structure variations that cannot be discovered by the second-generation sequencing data. However, compared with the second-generation sequencing data, the PacBio data lacks paired-end sequencing information, making traditional structure variations filter fail to process the new data. To solve this problem, this paper proposes a universal multi-tasking structure variation filtering model MSVF. MSVF adopts the CIGAR string defined in SAM format. CIGAR is not limited by sequencing technology or alignment algorithms, so MSVF is suitable for not only the second-generation but also the third-generation sequencing data. Moreover, CIGAR string preserves the complete sequence alignment information, which makes MSVF a highly precise model. Besides, MSVF uses deep learning methods, making it supports more structure variation types, including deletion and insertion. We trained and tested the models on the open-access NCBI datasets. The experiments proved that ShuffleNet, MobileNet, ResNet transfer learning models achieve better classification results on SVs task. The average AUC reaches more than 90% and the AUC of each category reach more than 87%. The accuracy and AUC of deletion and insertion structure variations were above 90% and above 92%, respectively. The code and data can be obtained at https://github.con weimingxiang/MSVF. Weiming Xiang 0003, Yingbo Cui 0001, Yaning Yang, Shaoliang Peng |
BIBM | 3 |
| 2022 | SVPath: an accurate pipeline for predicting the pathogenicity of human exon structural variantsabstractAlthough there are a large number of structural variations in the chromosomes of each individual, there is a lack of more accurate methods for identifying clinical pathogenic variants. Here, we proposed SVPath, a machine learning-based method to predict the pathogenicity of deletions, insertions and duplications structural variations that occur in exons. We constructed three types of annotation features for each structural variation event in the ClinVar database. First, we treated complex structural variations as multiple consecutive single nucleotide polymorphisms events, and annotated them with correlation scores based on single nucleic acid substitutions, such as the impact on protein function. Second, we determined which genes the variation occurred in, and constructed gene-based annotation features for each structural variation. Third, we also calculated related features based on the transcriptome, such as histone signal, the overlap ratio of variation and genomic element definitions, etc. Finally, we employed a gradient boosting decision tree machine learning method, and used the deletions, insertions and duplications in the ClinVar database to train a structural variation pathogenicity prediction model SVPath. These structural variations are clearly indicated as pathogenic or benign. Experimental results show that our SVPath has achieved excellent predictive performance and outperforms existing state-of-the-art tools. SVPath is very promising in evaluating the clinical pathogenicity of structural variants. SVPath can be used in clinical research to predict the clinical significance of unknown pathogenicity and new structural variation, so as to explore the relationship between diseases and structural variations in a computational way. Yaning Yang, Deshan Zhou, Shaoliang Peng |
Briefings Bioinform. | 1 |
| 2021 | ParaPindel: a scalable coordinated parallel detection framework for human genome-wide structural variationabstractDetecting the existence of variation from massive human genome data, and determining the breakpoints and types of variations is essential for analyzing structural variation. Pindel, an accurate detection tool based on pattern growth approach, is commonly used for discovering indels and other types of structural variations from next-generation sequencing data. The explosive growth of sequencing data poses new challenges to current implementation of Pindel. Here, we proposed ParaPindel, an optimized version of Pindel that utilizes distributed multiprocess, for efficient large-scale detection of structural variation for human whole-genome sequencing data. ParaPindel divides the chromosome into multiple small windows with a fixed-length window size, so as to realize the parallel detection between different windows and different chromosomes. A crosswindow with a smaller length is introduced to cope with possible structural variations at the edge of the window. The experimental results show that ParaPindel shortens the time to detect an individual’s genome-wide structural variation from 186 hours to 33 minutes under the premise that the detection results are basically consistent. Employing 256 processes on 128 nodes on the TH-IHN supercomputer, the speedup ratio has reached 163 times, and the parallel efficiency has reached 69.74%. Yaning Yang, Chao Yang 0015, Bin Jiang 0006, Shaoliang Peng |
BIBM | 1 |
| 2021 | CFCN: A Multi-scale Fully Convolutional Network with Dilated Convolution for Nuclei Classification and Localization
Yaning Yang, Shaoliang Peng |
ISBRA | 2 |
| 2021 | SCEBE: an efficient and scalable algorithm for genome-wide association studies on longitudinal outcomes with mixed-effects modelingabstractGenome-wide association studies (GWAS) using longitudinal phenotypes collected over time is appealing due to the improvement of power. However, computation burden has been a challenge because of the complex algorithms for modeling the longitudinal data. Approximation methods based on empirical Bayesian estimates (EBEs) from mixed-effects modeling have been developed to expedite the analysis. However, our analysis demonstrated that bias in both association test and estimation for the existing EBE-based methods remains an issue. We propose an incredibly fast and unbiased method (simultaneous correction for EBE, SCEBE) that can correct the bias in the naive EBE approach and provide unbiased P-values and estimates of effect size. Through application to Alzheimer's Disease Neuroimaging Initiative data with 6 414 695 single nucleotide polymorphisms, we demonstrated that SCEBE can efficiently perform large-scale GWAS with longitudinal outcomes, providing nearly 10 000 times improvement of computational efficiency and shortening the computation time from months to minutes. The SCEBE package and the example datasets are available at https://github.com/Myuan2019/SCEBE. Xu Steven 0001, Yaning Yang, Yinsheng Zhou, Jinfeng Xu 0001, Jose Pinheiro |
Briefings Bioinform. | 3 |
| 2021 | BioERP: biomedical heterogeneous network-based self-supervised representation learning approach for entity relationship predictionsabstractMOTIVATION: Predicting entity relationship can greatly benefit important biomedical problems. Recently, a large amount of biomedical heterogeneous networks (BioHNs) are generated and offer opportunities for developing network-based learning approaches to predict relationships among entities. However, current researches slightly explored BioHNs-based self-supervised representation learning methods, and are hard to simultaneously capturing local- and global-level association information among entities. RESULTS: In this study, we propose a BioHN-based self-supervised representation learning approach for entity relationship predictions, termed BioERP. A self-supervised meta path detection mechanism is proposed to train a deep Transformer encoder model that can capture the global structure and semantic feature in BioHNs. Meanwhile, a biomedical entity mask learning strategy is designed to reflect local associations of vertices. Finally, the representations from different task models are concatenated to generate two-level representation vectors for predicting relationships among entities. The results on eight datasets show BioERP outperforms 30 state-of-the-art methods. In particular, BioERP reveals great performance with results close to 1 in terms of AUC and AUPR on the drug-target interaction predictions. In summary, BioERP is a promising bio-entity relationship prediction approach. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/pengsl-lab/BioERP.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yaning Yang, Kenli Li 0001, Fei Li 0040, Shaoliang Peng |
Bioinform. | 2 |
| 2021 | Discriminant Projection Shared Dictionary Learning for Classification of Tumors Using Gene Expression DataabstractWith a variety of tumor subtypes, personalized treatments need to identify the subtype of a tumor as accurately as possible. The development of DNA microarrays provides an opportunity to predict tumor classification. One strategy is to use gene expression profiling to extend current biological insights into the disease. However, overfitting problems exist in most machine learning methods when classifying tumor gene expression profile data characterized by high dimensional, small samples and nonlinearities. As a new machine learning methods, dictionary learning has become a more effective algorithm for gene expression profile classification. Here, a new method called discriminant projection shared dictionary learning (DPSDL) is proposed for classifying tumor subtypes using LINCS gene expression profile data. The method trains a shared dictionary, embeds Fisher discriminant criteria to obtain a class-specific sub-dictionary and coding coefficients. At the same time, a projection matrix is trained to widen the distance between different classes of samples. Experimental results show that our method performs better classification based on gene expression profile than the other dictionary learning methods and machine learning methods. Shaoliang Peng, Yaning Yang, Fei Li 0040, Xiangke Liao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Efficiently recognition of vaginal micro-ecological environment based on Convolutional Neural NetworkabstractVaginal diseases caused by vaginal micro-ecological abnormalities mainly include Vulvovaginal Candidiasis (VVC), Aerobic Vaginitis (AV), and Bacterial Vaginosis (BV). Severe cases can lead to poor pregnancy outcomes and infertility. AI-based technologies are being deployed with an expectation to relieve doctors of routine, tedious work when implemented correctly in daily microscopy of vaginal micro-ecological abnormalities. In this paper, we built a clinical image dataset of the Gram stain of the vaginal discharge. By comparing the performance of state of art convolutional neural network models, we found the fine-tuning Inception ResNet V2 shows the best classification performance for vaginal diseases. It achieves 96%, 94%, 86% AUC in VVC, AV, BV classification respectively. The result shows that compared with human visual inspection, the method based on deep learning greatly improves the screening sensitivity. Besides, we found that transfer learning can reduce the required manual labeling by roughly 73% (about more than one thousand samples). But for BV, which is difficult to diagnose for both humans and AI. Unlike AV and VVC, it requires more labeled data and is insensitive to the transfer fine-tuning. Shaoliang Peng, Minxia Cheng, Yaning Yang, Fei Li 0040 |
HealthCom | 4 |
| 2020 | A deep metric learning algorithm for similarity measure of the gene expression profileabstractClustering gene expression profiles is a fundamental task in the genome and biomedical research. With the development of RNA-seq and gene chip technology, mass gene expression profile data has been generated, which puts forward two requirements for related research of gene expression profile: i) accurate analysis of drug R&D requires high accuracy of similarity analysis, ii) large-scale analysis of data requires as little running time as possible. We propose a faster, more accurate method called DeepCDNet, which is based on the framework of the Siamese network. DeepCDNet uses the DenseNet structure and optimized loss function to achieve rapid convergence, and the similarity between expression spectra is calculated by a cosine function. The experiment results show that: i) our method breaks through the limitation of high dimensions of gene expression profile and can quickly and accurately learn the required gene characteristics, ii) The accuracy of our method in similarity analysis is greatly improved, iii) as the dimension of data increases, the advantage of our method on time cost gradually becomes more prominent, and time consumption is less. Shaoliang Peng, Yaning Yang, Fei Li 0040, Hao Hong, Kenli Li 0001, Shulin Wang |
HealthCom | 3 |
| 2020 | A Dynamic Protection Mechanism for GPU Memory Overflow
Yaning Yang, Shaoliang Peng |
NPC | 1 |
| 2020 | High-throughput and efficient multilocus genome-wide association study on longitudinal outcomesabstractMOTIVATION: With the emerging of high-dimensional genomic data, genetic analysis such as genome-wide association studies (GWAS) have played an important role in identifying disease-related genetic variants and novel treatments. Complex longitudinal phenotypes are commonly collected in medical studies. However, since limited analytical approaches are available for longitudinal traits, these data are often underutilized. In this article, we develop a high-throughput machine learning approach for multilocus GWAS using longitudinal traits by coupling Empirical Bayesian Estimates from mixed-effects modeling with a novel ℓ0-norm algorithm. RESULTS: Extensive simulations demonstrated that the proposed approach not only provided accurate selection of single nucleotide polymorphisms (SNPs) with comparable or higher power but also robust control of false positives. More importantly, this novel approach is highly scalable and could be approximately >1000 times faster than recently published approaches, making genome-wide multilocus analysis of longitudinal traits possible. In addition, our proposed approach can simultaneously analyze millions of SNPs if the computer memory allows, thereby potentially allowing a true multilocus analysis for high-dimensional genomic data. With application to the data from Alzheimer's Disease Neuroimaging Initiative, we confirmed that our approach can identify well-known SNPs associated with AD and were much faster than recently published approaches (≥6000 times). AVAILABILITY AND IMPLEMENTATION: The source code and the testing datasets are available at https://github.com/Myuan2019/EBE_APML0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huang Xu 0005, Yaning Yang, Jose Pinheiro, Kate Sasser, Hisham Hamadeh, Xu Steven 0001 |
Bioinform. | 3 |
| 2019 | A General Fine-tuned Transfer Learning Model for Predicting Clinical Task Acrossing Diverse EHRs DatasetsabstractData analysis of electronic health record (EHRs) system using machine learning, statistical methods can predict relevant clinical tasks. However, there is no uniform standard for current electronic health record systems, and the clinical outcome prediction models trained on one EHR dataset cannot be applied well on other EHR datasets from different medical institutions. Data differences between different medical institutions pose a huge challenge to the study of electronic health records. In this study, we proposed a general transfer learning strategy which can enable models to make clinical prediction acrossing diverse EHRs datasets and validated its strong versatility on three deep learning models. Two different intensive care units (ICU) databases (MIMIC-III and eICU) and one clinical task (in-hospital mortality) are used to evaluate our method. At first, we trained the deep learning models on the source dataset and saved the model states after each epoch. Then, we selected the best performing model as the pre-training model, transferred it to the target dataset and fine-tuned the whole network on target dataset. Finally, we use the fine-tuned models to make predictions on the target dataset. Experiment results show that AUROC score increased by 3%-20% with transfer strategy, which indicated that the general strategy can provide more reliable predictions acrossing EHRs databases to predict clinical tasks. Shaoliang Peng, Yaning Yang, Fei Li 0040 |
BIBM | 3 |
| 2019 | CSHAP: efficient haplotype frequency estimation based on sparse representationabstractMOTIVATION: Estimating haplotype frequencies from genotype data plays an important role in genetic analysis. In silico methods are usually computationally involved since phase information is not available. Due to tight linkage disequilibrium and low recombination rates, the number of haplotypes observed in human populations is far less than all the possibilities. This motivates us to solve the estimation problem by maximizing the sparsity of existing haplotypes. Here, we propose a new algorithm by applying the compressive sensing (CS) theory in the field of signal processing, compressive sensing haplotype inference (CSHAP), to solve the sparse representation of haplotype frequencies based on allele frequencies and between-allele co-variances. RESULTS: Our proposed approach can handle both individual genotype data and pooled DNA data with hundreds of loci. The CSHAP exhibits the same accuracy compared with the state-of-the-art methods, but runs several orders of magnitude faster. CSHAP can also handle with missing genotype data imputations efficiently. AVAILABILITY AND IMPLEMENTATION: The CSHAP is implemented in R, the source code and the testing datasets are available at http://home.ustc.edu.cn/∼zhouys/CSHAP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yinsheng Zhou, Han Zhang 0003, Yaning Yang |
Bioinform. | 3 |
| 2010 | Testing multiple gene interactions by the ordered combinatorial partitioning method in case-control studiesabstractMOTIVATION: The multifactor-dimensionality reduction (MDR) method has been widely used in multi-locus interaction analysis. It reduces dimensionality by partitioning the multi-locus genotypes into a high-risk group and a low-risk group according to whether the genotype-specific risk ratio exceeds a fixed threshold or not. Alternatively, one can maximize the chi(2) value exhaustively over all possible ways of partitioning the multi-locus genotypes into two groups, and we aim to show that this is computationally feasible. METHODS: We advocate finding the optimal MDR (OMDR) that would have resulted from an exhaustive search over all possible ways of partitioning the multi-locus genotypes into two groups. It is shown that this optimal MDR can be obtained efficiently using an ordered combinatorial partitioning (OCP) method, which differs from the existing MDR method in the use of a data-driven rather than fixed threshold. The generalized extreme value distribution (GEVD) theory is applied to find the optimal order of gene combination and assess statistical significance of interactions. RESULTS: The computational complexity of OCP strategy is linear in the number of multi-locus genotypes in contrast with an exponential order for the naive exhaustive search strategy. Simulation studies show that OMDR can be more powerful than MDR with substantial power gain possible when the partitioning of OMDR is different from that of MDR. The analysis results of a breast cancer dataset show that the use of GEVD accelerates the determination of interaction order and reduces the time cost for P-value calculation by more than 10-fold. AVAILABILITY: C++ program is available at http://home.ustc.edu.cn/~zhanghan/ocp/ocp.html Xing Hua, Han Zhang 0003, Yaning Yang, Anthony Y. C. Kuk |
Bioinform. | 4 |
| 2010 | A study of the efficiency of pooling in haplotype estimationabstractMOTIVATION: It has been claimed in the literature that pooling DNA samples is efficient in estimating haplotype frequencies. There is, however, no theoretical justification based on calculation of statistical efficiency. In fact, the limited evidence given so far is based on simulation studies with small numbers of loci. With rapid advance in technology, it is of interest to see if pooling is still efficient when the number of loci increases. METHODS: Instead of resorting to simulation studies, we make use of asymptotic statistical theory to perform exact calculation of the efficiency of pooling relative to no pooling in the estimation of haplotype frequencies. As an intermediate step, we use the log-linear formulation of the haplotype probabilities and derive the asymptotic variance-covariance matrix of the maximum likelihood estimators of the canonical parameters of the log-linear model. RESULTS: Based on our calculations under linkage equilibrium, pooling can suffer huge loss in efficiency relative to no pooling when there are more than three independent loci and the alleles are not rare. Pooling works better for rare alleles. In particular, if all the minor allele frequencies are 0.05, pooling maintains an advantage over no pooling until the number of independent loci reaches 6. High linkage disequilibrium effectively reduces the number of independent loci by ruling out certain haplotypes from occurring. Similar calculations of efficiency for the case of no pooling justify the common belief that it is not worthwhile to use molecular methods to resolve the phase ambiguity of individual genotype data. AVAILABILITY: The R codes for the calculation are available at http://www.stat.nus.edu.sg/∼staxj/pooling CONTACT: [email protected]. Anthony Y. C. Kuk, Jinfeng Xu 0001, Yaning Yang |
Bioinform. | 3 |
| 2009 | Computationally feasible estimation of haplotype frequencies from pooled DNA with and without Hardy-Weinberg equilibriumabstractMOTIVATION: Pooling large number of DNA samples is a common practice in association study, especially for initial screening. However, the use of expectation-maximization (EM)-type algorithms in estimating haplotype distributions for even moderate pool sizes is hampered by the computational complexity involved. A novel constrained EM algorithm called PoooL has been proposed recently to bypass the difficulty via the use of asymptotic normality of the pooled allele frequencies. The resulting estimates are, however, not maximum likelihood estimates and hence not optimal. Furthermore, the assumption of Hardy-Weinberg equilibrium (HWE) made may not be realistic in practice. METHODS: Rather than carrying out constrained maximization as in PoooL, we revert to the usual EM algorithm but make it computationally feasible by using normal approximations. The resulting algorithm is much simpler to implement than PoooL because there is no need to invoke sophisticated iterative scaling methods as in PoooL. We also develop an estimating equation analogue of the EM algorithm for the case of Hardy-Weinberg disequilibrium (HWD) by conditioning on the haplotypes of both chromosomes of the same individual. Incorporated into the method is a way of estimating the inbreeding coefficient by relating it to overdispersion. RESULTS: Simulation study assuming HWE shows that our simplified implementation of the EM algorithm leads to estimates with substantially smaller SDs than PoooL estimates. Further simulations show that ignoring HWD will induce biases in the estimates. Our extended method with estimation of inbreeding coefficient incorporated is able to reduce the bias leading to estimates with substantially smaller mean square errors. We also present results to suggest that our method can cope with a certain degree of locus-specific inbreeding as well as additional overdispersion not caused by inbreeding. AVAILABILITY: http://staff.ustc.edu.cn/ approximately ynyang/aem-aes Anthony Y. C. Kuk, Han Zhang 0003, Yaning Yang |
Bioinform. | 3 |
| 2009 | Partial correlation analysis indicates causal relationships between GC-content, exon density and recombination rate in the human genomeabstractBACKGROUND: Several features are known to correlate with the GC-content in the human genome, including recombination rate, gene density and distance to telomere. However, by testing for pairwise correlation only, it is impossible to distinguish direct associations from indirect ones and to distinguish between causes and effects. RESULTS: We use partial correlations to construct partially directed graphs for the following four variables: GC-content, recombination rate, exon density and distance-to-telomere. Recombination rate and exon density are unconditionally uncorrelated, but become inversely correlated by conditioning on GC-content. This pattern indicates a model where recombination rate and exon density are two independent causes of GC-content variation. CONCLUSION: Causal inference and graphical models are useful methods to understand genome evolution and the mechanisms of isochore evolution in the human genome. Jan Freudenberg, Yaning Yang, Wentian Li |
BMC Bioinform. | 3 |
| 2008 | PoooL: an efficient method for estimating haplotype frequencies from large DNA poolsabstractMOTIVATION: Pooling DNA is a cost-effective alternative to individual genotyping method. It is often used for initial screening in genome-wide association analysis. In some studies, large pools with sizes up to several hundreds were applied in order to significantly reduce genotyping cost. However, method for estimating haplotype frequencies from large DNA pools has not been available due to computational complexity involved. METHODS: We propose a novel constrained EM algorithm, PoooL, to estimate frequencies of single-nucleotide polymorphism (SNP) haplotypes from DNA pools. A quantity called importance factor is introduced to measure the contribution of a haplotype to the likelihood. Under the assumption of asymptotic normality of the estimated allele frequencies and a system of linear constraints on haplotype frequencies the importance factor remains a constant in the iterative maximization process. The maximization problem in the EM algorithm is then formulated into a constrained maximum entropy model and solved by the improved iterative scaling method. RESULTS: Simulation study shows that our algorithm can efficiently estimate haplotype frequencies from DNA pools with arbitrarily large sizes. The algorithm works equally well for large pools with sizes up to hundreds or thousands and for pools with sizes as small as one or two individuals. The computational complexity of the PoooL algorithm is independent of pool sizes, and the computational efficiency for large pools is thus substantially improved over existing estimating methods. Simulation results also show that the proposed method is robust to genotype errors and population admixture. Han Zhang 0003, Hsin-Chou Yang, Yaning Yang |
Bioinform. | 3 |
| 2007 | A parsimonious threshold-independent protein feature selection method through the area under receiver operating characteristic curveabstractMOTIVATION: Protein expression profiling for differences indicative of early cancer holds promise for improving diagnostics. Due to their high dimensionality, statistical analysis of proteomic data from mass spectrometers is challenging in many aspects such as dimension reduction, feature subset selection as well as construction of classification rules. Search of an optimal feature subset, commonly known as the feature subset selection (FSS) problem, is an important step towards disease classification/diagnostics with biomarkers. METHODS: We develop a parsimonious threshold-independent feature selection (PTIFS) method based on the concept of area under the curve (AUC) of the receiver operating characteristic (ROC). To reduce computational complexity to a manageable level, we use a sigmoid approximation to the empirical AUC as the criterion function. Starting from an anchor feature, the PTIFS method selects a feature subset through an iterative updating algorithm. Highly correlated features that have similar discriminating power are precluded from being selected simultaneously. The classification rule is then determined from the resulting feature subset. RESULTS: The performance of the proposed approach is investigated by extensive simulation studies, and by applying the method to two mass spectrometry data sets of prostate cancer and of liver cancer. We compare the new approach with the threshold gradient descent regularization (TGDR) method. The results show that our method can achieve comparable performance to that of the TGDR method in terms of disease classification, but with fewer features selected. AVAILABILITY: Supplementary Material and the PTIFS implementations are available at http://staff.ustc.edu.cn/~ynyang/PTIFS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhanfeng Wang, Yuan-chin Ivan Chang, Zhiliang Ying, Yaning Yang |
Bioinform. | 5 |
| 2003 | Statistical significance for hierarchical clustering in genetic association and microarray expression studiesabstractBACKGROUND: With the increasing amount of data generated in molecular genetics laboratories, it is often difficult to make sense of results because of the vast number of different outcomes or variables studied. Examples include expression levels for large numbers of genes and haplotypes at large numbers of loci. It is then natural to group observations into smaller numbers of classes that allow for an easier overview and interpretation of the data. This grouping is often carried out in multiple steps with the aid of hierarchical cluster analysis, each step leading to a smaller number of classes by combining similar observations or classes. At each step, either implicitly or explicitly, researchers tend to interpret results and eventually focus on that set of classes providing the "best" (most significant) result. While this approach makes sense, the overall statistical significance of the experiment must include the clustering process, which modifies the grouping structure of the data and often removes variation. RESULTS: For hierarchically clustered data, we propose considering the strongest result or, equivalently, the smallest p-value as the experiment-wise statistic of interest and evaluating its significance level for a global assessment of statistical significance. We apply our approach to datasets from haplotype association and microarray expression studies where hierarchical clustering has been used. CONCLUSION: In all of the cases we examine, we find that relying on one set of classes in the course of clustering leads to significance levels that are too small when compared with the significance level associated with an overall statistic that incorporates the process of clustering. In other words, relying on one step of clustering may furnish a formally significant result while the overall experiment is not significant. Mark A. Levenstien, Yaning Yang, Jürg Ott |
BMC Bioinform. | 2 |