Yingjun Ma

dblp:72/4371 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
11since 2021 · last 2027
0000-0001-6044-1439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2027 TuckerHKN: multi-source hypergraph-guided Tucker co-learning for unraveling high-order structures in spatial transcriptomics
Yingjun Ma, Huiqin Zeng, Yi Liu 0066, Yunfang Liu
Expert Syst. Appl.1
2024 Kernel Bayesian logistic tensor decomposition with automatic rank determination for predicting multiple types of miRNA-disease associations
abstract
Identifying the association and corresponding types of miRNAs and diseases is crucial for studying the molecular mechanisms of disease-related miRNAs. Compared to traditional biological experiments, computational models can not only save time and reduce costs, but also discover potential associations on a large scale. Although some computational models based on tensor decomposition have been proposed, these models usually require manual specification of numerous hyperparameters, leading to a decrease in computational efficiency and generalization ability. Additionally, these linear models struggle to analyze complex, higher-order nonlinear relationships. Based on this, we propose a novel framework, KBLTDARD, to identify potential multiple types of miRNA-disease associations. Firstly, KBLTDARD extracts information from biological networks and high-order association network, and then fuses them to obtain more precise similarities of miRNAs (diseases). Secondly, we combine logistic tensor decomposition and Bayesian methods to achieve automatic hyperparameter search by introducing sparse-induced priors of multiple latent variables, and incorporate auxiliary information to improve prediction capabilities. Finally, an efficient deterministic Bayesian inference algorithm is developed to ensure computational efficiency. Experimental results on two benchmark datasets show that KBLTDARD has better Top-1 precision, Top-1 recall, and Top-1 F1 for new type predictions, and higher AUPR, AUC, and F1 values for new triplet predictions, compared to other state-of-the-art methods. Furthermore, case studies demonstrate the efficiency of KBLTDARD in predicting multiple types of miRNA-disease associations.
Yingjun Ma
PLoS Comput. Biol.1
2023 Part-and-whole: A novel framework for deformable medical image registration
Jinshuo Zhang, Yingjun Ma, Xiuyang Zhao, Bo Yang 0001
Appl. Intell.3
2023 Logistic tensor decomposition with sparse subspace learning for prediction of multiple disease types of human-virus protein-protein interactions
abstract
Viral infection involves a large number of protein-protein interactions (PPIs) between the virus and the host, and the identification of these PPIs plays an important role in revealing viral infection and pathogenesis. Existing computational models focus on predicting whether human proteins and viral proteins interact, and rarely take into account the types of diseases associated with these interactions. Although there are computational models based on a matrix and tensor decomposition for predicting multi-type biological interaction relationships, these methods cannot effectively model high-order nonlinear relationships of biological entities and are not suitable for integrating multiple features. To this end, we propose a novel computational framework, LTDSSL, to determine human-virus PPIs under different disease types. LTDSSL utilizes logistic functions to model nonlinear associations, sets importance levels to emphasize the importance of observed interactions and utilizes sparse subspace learning of multiple features to improve model performance. Experimental results show that LTDSSL has better predictive performance for both new disease types and new triples than the state-of-the-art methods. In addition, the case study further demonstrates that LTDSSL can effectively predict human-viral PPIs under various disease types.
Yingjun Ma, Junjiang Zhong
Briefings Bioinform.1
2023 HONMF: integration analysis of multi-omics microbiome data via matrix factorization and hypergraph
abstract
MOTIVATION: The accumulation of multi-omics microbiome data provides an unprecedented opportunity to understand the diversity of bacterial, fungal, and viral components from different conditions. The changes in the composition of viruses, bacteria, and fungi communities have been associated with environments and critical illness. However, identifying and dissecting the heterogeneity of microbial samples and cross-kingdom interactions remains challenging. RESULTS: We propose HONMF for the integrative analysis of multi-modal microbiome data, including bacterial, fungal, and viral composition profiles. HONMF enables identification of microbial samples and data visualization, and also facilitates downstream analysis, including feature selection and cross-kingdom association analysis between species. HONMF is an unsupervised method based on hypergraph induced orthogonal non-negative matrix factorization, where it assumes that latent variables are specific for each composition profile and integrates the distinct sets of latent variables through graph fusion strategy, which better tackles the distinct characteristics in bacterial, fungal, and viral microbiome. We implemented HONMF on several multi-omics microbiome datasets from different environments and tissues. The experimental results demonstrate the superior performance of HONMF in data visualization and clustering. HONMF also provides rich biological insights by implementing discriminative microbial feature selection and bacterium-fungus-virus association analysis, which improves our understanding of ecological interactions and microbial pathogenesis. AVAILABILITY AND IMPLEMENTATION: The software and datasets are available at https://github.com/chonghua-1983/HONMF.
Yingjun Ma
Bioinform.3
2022 Unsupervised deformable image registration network for 3D medical images
Yingjun Ma, Dongmei Niu, Jinshuo Zhang, Xiuyang Zhao, Bo Yang 0001, Caiming Zhang 0001
Appl. Intell.1
2022 Hypergraph-based logistic matrix factorization for metabolite-disease interaction prediction
abstract
MOTIVATION: Function-related metabolites, the terminal products of the cell regulation, show a close association with complex diseases. The identification of disease-related metabolites is critical to the diagnosis, prevention and treatment of diseases. However, most existing computational approaches build networks by calculating pairwise relationships, which is inappropriate for mining higher-order relationships. RESULTS: In this study, we presented a novel approach with hypergraph-based logistic matrix factorization, HGLMF, to predict the potential interactions between metabolites and disease. First, the molecular structures and gene associations of metabolites and the hierarchical structures and GO functional annotations of diseases were extracted to build various similarity measures of metabolites and diseases. Next, the kernel neighborhood similarity of metabolites (or diseases) was calculated according to the completed interactive network. Second, multiple networks of metabolites and diseases were fused, respectively, and the hypergraph structures of metabolites and diseases were built. Finally, a logistic matrix factorization based on hypergraph was proposed to predict potential metabolite-disease interactions. In computational experiments, HGLMF accurately predicted the metabolite-disease interaction, and performed better than other state-of-the-art methods. Moreover, HGLMF could be used to predict new metabolites (or diseases). As suggested from the case studies, the proposed method could discover novel disease-related metabolites, which has been confirmed in existing studies. AVAILABILITY AND IMPLEMENTATION: The codes and dataset are available at: https://github.com/Mayingjun20179/HGLMF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingjun Ma
Bioinform.1
2022 Small Infrared Target Detection Based on Fast Adaptive Masking and Scaling With Iterative Segmentation
abstract
Fast and robust small infrared (IR) target detection is a challenging task and critical to the performance of IR searching and tracking (IRST) systems. However, the current algorithms generally have difficulty in striking a good balance between speed and performance. In this letter, we propose a new approach to small IR target detection that can significantly accelerate the detection process by first performing a fast adaptive masking and scaling algorithm. We then propose to enhance the target characteristics and suppress the background clutter using both contrast and gradient information. Finally, we propose to accurately extract the targets via iterative segmentation. The experimental results demonstrated that our proposed method yields the best and the most robust performance, with a speed of at least two times faster than the state-of-the-art methods.
Yaohong Chen, Gaopeng Zhang, Yingjun Ma, Jin U. Kang, Chiman Kwan
IEEE Geosci. Remote. Sens. Lett.3
2022 Seq-BEL: Sequence-Based Ensemble Learning for Predicting Virus-Human Protein-Protein Interaction
abstract
Infectious diseases are currently the most important and widespread health problem, and identifying viral infection mechanisms is critical for controlling diseases caused by highly infectious viruses. Because of the lack of non-interactive protein pairs and serious imbalance between positive and negative sample ratios, the supervised learning algorithm is not suitable for prediction. At the same time, due to the lack of information on viral proteins and significant dissimilarity in sequence, some ensemble learning models have poor generalization ability. In this paper, we propose a Sequence-Based Ensemble Learning (Seq-BEL) method to predict the potential virus-human PPIs. Specifically, based on the amino acid sequence of proteins and the currently known virus-human PPI network, Seq-BEL calculates various features and similarities of human proteins and viral proteins, and then combines these similarities and features to score the potential of virus-human PPIs. The computational results show that Seq-BEL achieves success in predicting potential virus-human PPIs and outperforms other state-of-the-art methods. More importantly, Seq-BEL also has good predictive performance for new human proteins and new viral proteins. In addition, the model has the advantages of strong robustness and good generalization ability, and can be used as an effective tool for virus-human PPI prediction.
Yingjun Ma, Tingting He 0003, Yuting Tan 0001, Xingpeng Jiang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 DeepMNE: Deep Multi-Network Embedding for lncRNA-Disease Association Prediction
abstract
Long non-coding RNA (lncRNA) participates in various biological processes, hence its mutations and disorders play an important role in the pathogenesis of multiple human diseases. Identifying disease-related lncRNAs is crucial for the diagnosis, prevention, and treatment of diseases. Although a large number of computational approaches have been developed, effectively integrating multi-omics data and accurately predicting potential lncRNA-disease associations remains a challenge, especially regarding new lncRNAs and new diseases. In this work, we propose a new method with deep multi-network embedding, called DeepMNE, to discover potential lncRNA-disease associations, especially for novel diseases and lncRNAs. DeepMNE extracts multi-omics data to describe diseases and lncRNAs, and proposes a network fusion method based on deep learning to integrate multi-source information. Moreover, DeepMNE complements the sparse association network and uses kernel neighborhood similarity to construct disease similarity and lncRNA similarity networks. Furthermore, a graph embedding method is adopted to predict potential associations. Experimental results demonstrate that compared to other state-of-the-art methods, DeepMNE has a higher predictive performance on new associations, new lncRNAs and new diseases. Besides, DeepMNE also elicits a considerable predictive performance on perturbed datasets. Additionally, the results of two different types of case studies indicate that DeepMNE can be used as an effective tool for disease-related lncRNA prediction. The code of DeepMNE is freely available at https://github.com/Mayingjun20179/ DeepMNE.
Yingjun Ma
IEEE J. Biomed. Health Informatics1
2021 Performance improvement for a 2D convolutional neural network by using SSC encoding on protein-protein interaction tasks
abstract
BACKGROUND: The interactions of proteins are determined by their sequences and affect the regulation of the cell cycle, signal transduction and metabolism, which is of extraordinary significance to modern proteomics research. Despite advances in experimental technology, it is still expensive, laborious, and time-consuming to determine protein-protein interactions (PPIs), and there is a strong demand for effective bioinformatics approaches to identify potential PPIs. Considering the large amount of PPI data, a high-performance processor can be utilized to enhance the capability of the deep learning method and directly predict protein sequences. RESULTS: We propose the Sequence-Statistics-Content protein sequence encoding format (SSC) based on information extraction from the original sequence for further performance improvement of the convolutional neural network. The original protein sequences are encoded in the three-channel format by introducing statistical information (the second channel) and bigram encoding information (the third channel), which can increase the unique sequence features to enhance the performance of the deep learning model. On predicting protein-protein interaction tasks, the results using the 2D convolutional neural network (2D CNN) with the SSC encoding method are better than those of the 1D CNN with one hot encoding. The independent validation of new interactions from the HIPPIE database (version 2.1 published on July 18, 2017) and the validation of directly predicted results by applying a molecular docking tool indicate the effectiveness of the proposed protein encoding improvement in the CNN model. CONCLUSION: The proposed protein sequence encoding method is efficient at improving the capability of the CNN model on protein sequence-related tasks and may also be effective at enhancing the capability of other machine learning or deep learning methods. Prediction accuracy and molecular docking validation showed considerable improvement compared to the existing hot encoding method, indicating that the SSC encoding method may be useful for analyzing protein sequence-related tasks. The source code of the proposed methods is freely available for academic research at https://github.com/wangy496/SSC-format/ .
Zhanchao Li, Yingjun Ma, Qixing Huang, Zong Dai, Xiaoyong Zou
BMC Bioinform.4
2020 MHSNMF: multi-view hessian regularization based symmetric nonnegative matrix factorization for microbiome data analysis
abstract
BACKGROUND: With the rapid development of high-throughput technique, multiple heterogeneous omics data have been accumulated vastly (e.g., genomics, proteomics and metabolomics data). Integrating information from multiple sources or views is challenging to obtain a profound insight into the complicated relations among micro-organisms, nutrients and host environment. In this paper we propose a multi-view Hessian regularization based symmetric nonnegative matrix factorization algorithm (MHSNMF) for clustering heterogeneous microbiome data. Compared with many existing approaches, the advantages of MHSNMF lie in: (1) MHSNMF combines multiple Hessian regularization to leverage the high-order information from the same cohort of instances with multiple representations; (2) MHSNMF utilities the advantages of SNMF and naturally handles the complex relationship among microbiome samples; (3) uses the consensus matrix obtained by MHSNMF, we also design a novel approach to predict the classification of new microbiome samples. RESULTS: We conduct extensive experiments on two real-word datasets (Three-source dataset and Human Microbiome Plan dataset), the experimental results show that the proposed MHSNMF algorithm outperforms other baseline and state-of-the-art methods. Compared with other methods, MHSNMF achieves the best performance (accuracy: 95.28%, normalized mutual information: 91.79%) on microbiome data. It suggests the potential application of MHSNMF in microbiome data analysis. CONCLUSIONS: Results show that the proposed MHSNMF algorithm can effectively combine the phylogenetic, transporter, and metabolic profiles into a unified paradigm to analyze the relationships among different microbiome samples. Furthermore, the proposed prediction method based on MHSNMF has been shown to be effective in judging the types of new microbiome samples.
Junmin Zhao, Yingjun Ma
BMC Bioinform.3
2019 Predicting virus-host association by Kernelized logistic matrix factorization and similarity network fusion
abstract
BACKGROUND: Viruses are closely related to bacteria and human diseases. It is of great significance to predict associations between viruses and hosts for understanding the dynamics and complex functional networks in microbial community. With the rapid development of the metagenomics sequencing, some methods based on sequence similarity and genomic homology have been used to predict associations between viruses and hosts. However, the known virus-host association network was ignored in these methods. RESULTS: We proposed a kernelized logistic matrix factorization with integrating different information to predict potential virus-host associations on the heterogeneous network (ILMF-VH) which is constructed by connecting a virus network with a host network based on known virus-host associations. The virus network is constructed based on oligonucleotide frequency measurement, and the host network is constructed by integrating oligonucleotide frequency similarity and Gaussian interaction profile kernel similarity through similarity network fusion. The host prediction accuracy of our method is better than other methods. In addition, case studies show that the host of crAssphage predicted by ILMF-VH is consistent with presumed host in previous studies, and another potential host Escherichia coli is also predicted. CONCLUSIONS: The proposed model is an effective computational tool for predicting interactions between viruses and hosts effectively, and it has great potential for discovering novel hosts of viruses.
Yingjun Ma, Xingpeng Jiang, Tingting He 0003
BMC Bioinform.2