Bo Xu 0008

dblp:26/1194-8 · DBLP profile ↗
← Back
35ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0003-4135-0683ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering
abstract
In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the partially view-aligned problem (PVP). Most existing partially view-aligned clustering (PVC) methods first learn cross-view consistent representations based on known alignments, and then recover missing correspondences by measuring cross-view similarity between samples. However, such an indirect alignment recovery process depends on high-quality consistent representations and lacks effective utilization of known alignments, often resulting in sub-optimal outcomes. To address this, we propose a novel direct alignment recovery perspective, instantiated as K-Nearest Neighbors Direct Alignment (KNNDA). Specifically, we first construct an alignment domain by mapping the aligned neighbors of each unaligned sample into the aligned view. Then, we compute alignment confidence based on the similarity between known aligned pairs of neighbors. In particular, we use a dynamic threshold to filter out unreliable alignments. Finally, new alignments are generated within the high-confidence alignment domain. Contrastive loss is used to learn consistent representations for clustering. Comprehensive experiments on several real-world datasets show the effectiveness and superiority of our module in partially view-aligned clustering.
Liang Zhao 0005, Tianqi Yue, Shubin Ma, Bo Xu 0008
AAAI6
2025 Incomplete and Unpaired Multi-View Graph Clustering with Cross-View Feature Fusion
abstract
Due to its effectiveness and efficiency, graph-based multi-view clustering has recently attracted much attention. However, the multi-view data are often incomplete and unpaired in real-world applications as a consequence of data loss or corruption. Although efforts have been made through a series of methods to address the problems of incomplete or unpaired multi-view data, the following issues still persist: 1) Most existing methods only focus on the incomplete multi-view data or unpaired multi-view data, and exhibit weaknesses when addressing both incomplete and unpaired multi-view data simultaneously. 2) Some methods neglect the graph information of the data from different views during the learning process. To tackle these issues, we propose the Multi-view Graph Clustering framework with Cross-view Feature Fusion (MGCCFF), a novel approach for clustering incomplete and unpaired multi-view data. Specifically, MGCCFF learns soft clustering label information from complete data and utilizes this to capture category-level cross-view correspondences. It then learns latent representation enriched with cross-view information based on the established mappings. To obtain a multi-view graph structure under conditions of incomplete and unpaired data, MGCCFF innovatively integrates the concept of self-expression with the autoencoder architecture and exploits the latent relationships between labels and the graph structure, thereby enabling the generation of sparse and accurate graphical structure under multi-view conditions for the final clustering task. The experiments on incomplete and unpaired multi-view datasets demonstrate that MGCCFF outperforms state-of-the-art methods.
Liang Zhao 0005, Zhikui Chen, Bo Xu 0008
AAAI5
2025 Multi-View Community-Contrastive Graph Attention Network for Fmri-Based Alzheimer's Disease Classification
abstract
Functional brain network analysis based on fMRI is a vital tool for understanding neural mechanisms and diagnosing neurological disorders. However, existing approaches often overlook the joint modeling of topological structures and attribute information in brain connectivity graphs, limiting their performance in disease classification. To address this issue, we propose a novel Graph Neural Network framework combining multi-view modeling and supervised contrastive learning for fMRI-based Alzheimer's disease classification. Specifically, our method employs a Graph Attention Network (GAT) to encode structural relationships and a Multi-Layer Perceptron (MLP) to capture node attribute features. Furthermore, unsupervised clustering is utilized to extract community-level representations, capturing mesoscale brain network organization. To enhance feature robustness and discriminability, we introduce a dual-view supervised contrastive learning strategy. Extensive experiments on a cohort of 480 subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) demonstrate that our proposed model consistently outperforms state-of-the-art methods across key metrics, including accuracy, recall, F1-score, and AUC, highlighting its robustness and effectiveness for Alzheimer's disease diagnosis.
Bo Xu 0008, Baijiang Xu, Zihan Yuan, Jinshi Yu, Zhehuan Zhao, Lin Lin 0008
BIBM1
2025 NaviPath: A Novel Knowledge Graph-Based RAG Framework for Medical QA
abstract
Large Language Models (LLMs) have shown strong abilities in language understanding and reasoning, drawing increasing attention in medical question answering (QA). While retrieval-augmented generation (RAG) methods improve factual accuracy, existing approaches still struggle to retrieve and organize relevant knowledge effectively. To overcome this, we propose NaviPath, a knowledge graph-based RAG framework for medical QA. It enhances LLM responses through a structured prompt built in three steps: (1) extended entity retrieval, (2) multiperspective reasoning path construction, and (3) natural language transformation for better comprehension. Experiments on two medical QA benchmarks show that NaviPath achieves state-of-the-art performance in diagnostic accuracy and factual consistency. The implementation is available at https://github.com/zyr319/Navipath.git.
Zhehuan Zhao, Yuran Zhang, Bo Xu 0008, Ludan Zhang, Yu Liu 0035, Shimin Shan, Jian Wang 0021, Hongfei Lin
BIBM3
2025 Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search
abstract
Multi-modal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and misaligned characteristics due to factors such as inconsistent sensor frequencies and device malfunctions. Existing research has not effectively addressed the issue of filling missing data in scenarios where multiview data are both imbalanced and misaligned. Instead, it relies on class-level alignment of the available data. Thus, it results in some data samples not being well-matched, thereby affecting the quality of data fusion. In this paper, we propose the Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search(CAPIMAC) to tackle the problem of filling imbalanced and misaligned data in multi-modal datasets. Specifically, we propose a self-repellent greedy anchor search module(SRGASM), which employs a self-repellent random walk combined with a greedy algorithm to identify anchor points for re-representing incomplete and misaligned multi-modal data. Subsequently, based on noise-contrastive learning, we design a consistency-aware padding module (CAPM) to effectively interpolate and align imbalanced and misaligned data, thereby improving the quality of multi-modal data fusion. Experimental results demonstrate the superiority of our method over benchmark datasets. The code will be publicly released at https://github.com/bestow09090/-CAPIMAC.git.
Shubin Ma, Liang Zhao 0005, Mingdong Lu, Bo Xu 0008
IJCAI5
2025 Unveiling Maternity and Infant Care Conversations: A Chinese Dialogue Dataset for Enhanced Parenting Support
abstract
The rapid development of large language models has greatly advanced human-computer dialogue research. However, applying these models to specialized fields like maternity and infant care often leads to subpar performance due to a lack of domain-specific datasets. To address this problem, we have created MicDialogue, a Chinese dialogue dataset for maternity and infant care. MicDialogue involves a wide range of specialized topics, including gynecological health, pediatric care, pregnancy preparation, emotional counseling and other related topics. This dataset is curated from two types of Chinese social media: short videos and blog posts. Short videos capture real-time interactions and pragmatic dialogue patterns, while blog posts offer comprehensive coverage of various topics within the domain. We have also included detailed annotations for topics, diseases, symptoms, and causes, enabling in-depth research. Additionally, we developed a knowledge-driven benchmark model using LLM-based prompt learning and multiple knowledge graphs to address diverse dialogue topics. Experiments validate MicDialogue's usability, providing benchmarks for future research and essential data for fine-tuning language models in maternity and infant care.
Bo Xu 0008, Liangzhi Li 0004, Xuening Qiao, Erchen Yu, Yiming Qian, Linlin Zong, Hongfei Lin
IJCAI1
2025 EchoGPT: An Interactive Cardiac Function Assessment Model for Echocardiogram Videos
abstract
With the development of wearable cardiac ultrasound devices, it is no longer sufficient to solely rely on doctors for diagnosing long-term echocardiogram videos. Automated diagnosis of echocardiogram videos has now become a research hotspot. Existing studies only analyze echocardiogram video through discriminative models, which have limited question-answering capabilities. Therefore, this study innovatively proposes a large language model with cardiac ultrasound diagnostic capabilities—EchoGPT. EchoGPT integrates the robust communication and comprehension capabilities of large language models (LLMs) with the diagnostic prowess of traditional medical models, empowering patients to obtain accurate medical indicator data and comprehend their health conditions through interactive questioning with the model. The model is capable of local deployment on personal computers, effectively safe guarding user privacy. EchoGPT operates through three main components: left ventricle segmentation, left ventricular ejection fraction LVEF prediction, and finetuning of video-text LLMs. Experimental results demonstrate EchoGPT’s superior accuracy in predicting LVEF compared to other models, and positive feedback from professional physicians through questionnaire surveys, validating its potential in practical applications. The demo is available at https://github.com/zhuqh19/EchoGPT.
Bo Xu 0008, Quanhao Zhu, Qingchen Zhang 0001, Mengmeng Wang 0005, Liang Zhao 0005, Hongfei Lin, Jing Ren 0001, Feng Xia 0001
IJCAI1
2025 Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information
abstract
Recently, multi-view data has gradually attracted attention. However, real-world applications often face Partial View-aligned Problem (PVP) and Partially Sample-missing Problem (PSP) due to data loss or corruption. Existing methods addressing PVP typically focus only on learning from the information of aligned data, while ignoring unaligned data where samples exist but lack alignment relationships. This introduces PSP, which does not inherently exist in the data, leading to biased learning of the data's information. For PSP, due to varying degrees of missing data, incomplete spatial structures can cause clustering centers-shifted problem, resulting in the model learning incorrect correspondences and biased spatial structures.To tackle them, we propose a novel method called Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information (DRUMVC). To our knowledge, this is the first noise-robust and unbiased multi-view clustering method capable of simultaneously addressing both PVP and PSP. Specifically, DRUMVC leverages aligned and complete samples as a bridge to construct high-quality correspondences for samples lacking cross-view relationship information due to PVP or PSP. Additionally, we employ a dual noise-robust contrastive learning loss to mitigate the impact of noise potentially introduced during the pair construction. Experiments on several challenging datasets demonstrate the superiority of our proposed method.
Liang Zhao 0005, Chuanye He, Qingchen Zhang 0001, Bo Xu 0008
IJCAI5
2025 Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly Data
abstract
Multimodal feature fusion, by integrating the complementary information from each modality, can effectively capture complex features in real-world data. However, in many use cases, such as boiler combustion monitoring, factors including equipment failure, inconsistent sensor sampling frequencies, and network delays often cause data collected from different modalities to suffer from missing modality and temporal asynchrony. This leads to the incompleteness and disorderliness of multimodal data. To address these issues, previous studies have proposed several data fusion methods that align the cluster centers before fusion. However, these approaches have two key limitations: 1) they do not guarantee a high alignment accuracy of data pairs at the sample level, and 2) they do not address the issue of significant discrepancies in data sizes across different classes, which impacts the subsequent data fusion performance.
Liang Zhao 0005, Shubin Ma, Bo Xu 0008, Qingchen Zhang 0001
ACM Multimedia3
2025 Pseudo-Label Guided Incomplete Partial View-Aligned Clustering
abstract
Addressing the challenges of incomplete and misaligned data in multi-view learning is critical, giventhe inherent uncertainties and complexities of real-world data collection. These challenges often result in significant discrepancies in the quantity, quality, and completeness of data across different views. However, previous research has predominantly focused on addressing either incompleteness or misalignment in isolation. To address both incompleteness and misalignment simultaneously, we propose a novel model, Pseudo-Label Guided Incomplete Partial View-aligned Clustering (PGIPVC). Specifically, A pseudo-label acquisition module based on Cauchy divergence is proposed to preliminarily train the clustering structure of data in a single view, thereby obtaining pseudo-labels for each sample in the view. Subsequently, an incomplete partial alignment clustering module is designed to obtain discriminative latent representations through contrastive learning with positive-negative pairs selected based on KNN and pseudo-labeled samples. Extensive experiments on benchmark datasets demonstrate the superiority of our method compared to other state-of-the-art approaches.
Shubin Ma, Liang Zhao 0005, Songtao Wu, Bo Xu 0008
IEEE Signal Process. Lett.4
2024 Alzheimer's Disease Prediction with Irregular MRI Sequences Based on Pyramid Squeeze Attention and Time-Sensitive Attention Mechanisms
abstract
Alzheimer’s disease (AD) often leads to cognitive impairments and behavioral issues, making early detection and treatment crucial for slowing the progression of the disease and improving quality of life. This type of dementia typically worsens gradually, so accurately predicting the course of the disease is critical for medical intervention. To tackle it, this paper develops a new type of time series Alzheimer’s disease prediction model (ChaoJiBang-Net). This model integrates pyramid squeeze attention mechanisms, time position encoding, and time-sensitive attention techniques. It is specifically designed for irregular sampling and variable-length sequences in MRI images. The model aims to predict specific time points for Alzheimer’s disease, which are arbitrarily chosen. Experimental results show that increasing the amount of temporal data (time steps) can significantly improve the model’s prediction performance, as reflected in the progressively higher AUC values. Furthermore, this model can perform accurate predictions at any number of time steps, reducing reliance on frequent patient follow-ups and enhancing its clinical applicability. Currently, the predictive accuracy of this model, using MRI single modality data, surpasses the highest standards of existing state-of-the-art (SOTA) technologies.
Baijiang Xu, Bo Xu 0008, Liang Zhao 0005
BIBM5
2024 MAT: Medical AI-generated Text Detection Dataset from Multi-models and Multi-Methods
abstract
Large language models (LLMs) have been widely used in society due to their amazing emergent capabilities, but they also bring security issues. A large amount of content on the Internet may be generated by AI. Whether it is social forums or more professional academic fields, the abuse of AI has become a problem. Especially in some professional fields, Blindly trusting what the Internet says is dangerous. For this reason, for some fields that are risky and need to limit the use of AI, such as medicine, a more comprehensive benchmark is needed to test the ability of AI-generated text detection tasks. Considering the popularity of LLMs, the data distributions used in the training process of different LLMs may lead to the differences in generated data distributions, especially for some LLMs for non-native English speakers. To address this issue, this article introduces an AI-generated text detection dataset in the field of medical question answering. This dataset is generated by various models and prompting methods, and cross validation is performed on multiple types of data between texts generated by different methods to verify the effectiveness of AI-generated text detection and model classification tasks, and to study the generalization of the dataset in different tasks. We have published the dataset for future research on https://github.com/Hellpoop/MAT.
Bo Xu 0008, Ruiyuan Wang, Lingfan Ping, Chaoyue Zhu, Hongfei Lin, Linlin Tian, Feng Xia 0001
BIBM1
2024 Multimodal contrastive learning with neuroimaging and cognitive tests for Alzheimer's disease diagnosis
abstract
Alzheimer’s disease (AD) is a neurological illness that causes cognitive impairment. Computer-aided diagnosis can help diagnose Alzheimer’s disease early before clinical symptoms appear. Currently, many deep learning methods show good performance in AD diagnosis. Still, most of these methods are based on single/multimodal neuroimaging, leading to a one-sided approach to disease modelling. Combining neuroimaging, cognitive tests, and demographics can significantly improve model performance and reduce the negative impact of noise. This study proposes a multimodal model that introduces contrastive learning, extracting and fusing feature representations separately from cognitive tests and neuroimaging data. After that, contrastive learning based on similarity is employed for both modalities’ features, assisting the network in learning cross-modal features. Moreover, the hybrid attention mechanism of the Transformer encoder is explored for feature fusion. Experimental results on 2082 cases from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset validate the effectiveness of our proposed multimodal model.
Liang Zhao 0005, Bo Xu 0008, Yi Yang 0006, Yangqianhui Zhang, Ruixin Ma
BIBM3
2024 Generating Multimodal Metaphorical Features for Meme Understanding
abstract
Understanding a meme is a challenging task, due to the metaphorical information contained in the meme that requires intricate interpretation to grasp its intended meaning fully. In previous works, attempts have been made to facilitate computational understanding of memes through introducing human-annotated metaphors as extra input features into machine learning models. However, these approaches mainly focus on formulating linguistic representation of a metaphor (extracted from the texts appearing in memes), while ignoring the connection between the metaphor and corresponding visual features (e.g., objects in meme images). In this paper, we argue that a more comprehensive understanding of memes can only be achieved through a joint modelling of both visual and linguistic features of memes. To this end, we propose an approach to generate Multimodal Metaphorical feature for Meme Classification, named MMMC. MMMC derives visual characteristics from linguistic attributes of metaphorical concepts, which more effectively convey the underlying metaphorical concept, leveraging a text-conditioned generative adversarial network. The linguistic and visual features are then integrated into a set of multimodal metaphorical features for classification purpose. We perform extensive experiments on a benchmark metaphorical meme dataset, MET-Meme. Experimental results show that MMMC significantly outperforms existing baselines on the task of emotion classification and intention detection. Our code and dataset are available at https://github.com/liaolianfoka/MMMC.
Bo Xu 0008, Junzhe Zheng, Estrid He, Hongfei Lin, Liang Zhao 0005, Feng Xia 0001
ACM Multimedia1
2024 Attributed Graph Force Learning
abstract
In numerous network analysis tasks, feature representation plays an imperative role. Due to the intrinsic nature of networks being discrete, enormous challenges are imposed on their effective usage. There has been a significant amount of attention on network feature learning in recent times that has the potential of mapping discrete features into a continuous feature space. The methods, however, lack preserving the structural information owing to the utilization of random negative sampling during the training phase. The ability to effectively join attribute information to embedding feature space is also compromised. To address the shortcomings identified, a novel attribute force-based graph (AGForce) learning model is proposed that keeps the structural information intact along with adaptively joining attribute information to the node's features. To demonstrate the effectiveness of the proposed framework, comprehensive experiments on benchmark datasets are performed. AGForce based on the spring-electrical model extends opportunities to simulate node interaction for graph learning.
Ke Sun 0011, Feng Xia 0001, Jiaying Liu 0006, Bo Xu 0008, Vidya Saikrishna, Charu C. Aggarwal
IEEE Trans. Neural Networks Learn. Syst.4
2023 Medical Extractive Question-Answering Based on Fusion of Hierarchical Features
abstract
With the combination of natural language processing and artificial intelligence techniques, medical extractive question-answering (Q&A) provides valuable insights and assists medical professionals in daily work and scientific research, answering medical questions rapidly and accurately, thus holding significant practical significance. Therefore, research on medical extractive Q&A holds significant practical significance. However, the current state of medical extractive question-answering lacks attention to the interaction and prediction layers in the model structure. To address these issues, this paper proposes the Integrating pre-trained multi-layer structural feature information based Bio-BERT (IPMF-Bio-BERT) approach. This method leverages the rich word vector representations generated by the pre-trained Bio-BERT model, incorporating semantic and syntactic structural information to obtain multi-dimensional and complementary interactive feature information. Additionally, we introduce a flexible guidance network based on interactive information, which combines iterative and pointer network techniques to enhance the predictive performance of the question-answering model. We evaluate our proposed model on the specialized biomedical extractive question-answering BioASQ corpus. Experimental results demonstrate that the IPMF-Bio-BERT training strategy enhances the recognition and predictive capabilities of medical extractive Q&A, we establish new state-of-the-art results by outperforming existing approaches.
Zhikui Chen, Jinqiao Yang, Bo Xu 0008, Zhendong Guo, Ren Hao, Qiucen Li, Mei Sun
BIBM4
2023 MIRROR: Mining Implicit Relationships via Structure-Enhanced Graph Convolutional Networks
abstract
Data explosion in the information society drives people to develop more effective ways to extract meaningful information. Extracting semantic information and relational information has emerged as a key mining primitive in a wide variety of practical applications. Existing research on relation mining has primarily focused on explicit connections and ignored underlying information, e.g., the latent entity relations. Exploring such information (defined as implicit relationships in this article) provides an opportunity to reveal connotative knowledge and potential rules. In this article, we propose a novel research topic, i.e., how to identify implicit relationships across heterogeneous networks. Specially, we first give a clear and generic definition of implicit relationships. Then, we formalize the problem and propose an efficient solution, namely MIRROR, a graph convolutional network (GCN) model to infer implicit ties under explicit connections. MIRROR captures rich information in learning node-level representations by incorporating attributes from heterogeneous neighbors. Furthermore, MIRROR is tolerant of missing node attribute information because it is able to utilize network structure. We empirically evaluate MIRROR on four different genres of networks, achieving state-of-the-art performance for target relations mining. The underlying information revealed by MIRROR contributes to enriching existing knowledge and leading to novel domain insights.
Jiaying Liu 0006, Feng Xia 0001, Jing Ren 0001, Bo Xu 0008, Guansong Pang, Lianhua Chi
ACM Trans. Knowl. Discov. Data4
2021 Shifu2: A Network Representation Learning Based Model for Advisor-Advisee Relationship Mining
abstract
The advisor-advisee relationship represents direct knowledge heritage, and such relationship may not be readily available from academic libraries and search engines. This work aims to discover advisor-advisee relationships hidden behind scientific collaboration networks. For this purpose, we propose a novel model based on Network Representation Learning (NRL), namely Shifu2, which takes the collaboration network as input and the identified advisor-advisee relationship as output. In contrast to existing NRL models, Shifu2 considers not only the network structure but also the semantic information of nodes and edges. Shifu2 encodes nodes and edges into low-dimensional vectors respectively, both of which are then utilized to identify advisor-advisee relationships. Experimental results illustrate improved stability and effectiveness of the proposed model over state-of-the-art methods. In addition, we generate a large-scale academic genealogy dataset by taking advantage of Shifu2.
Jiaying Liu 0006, Feng Xia 0001, Lei Wang 0134, Bo Xu 0008, Xiangjie Kong 0001, Hanghang Tong, Irwin King
IEEE Trans. Knowl. Data Eng.4
2020 Graph Force Learning
abstract
Features representation leverages the great power in network analysis tasks. However, most features are discrete which poses tremendous challenges to effective use. Recently, increasing attention has been paid on network feature learning, which could map discrete features to continued space. Unfortunately, current studies fail to fully preserve the structural information in the feature space due to random negative sampling strategy during training. To tackle this problem, we study the problem of feature learning and novelty propose a force-based graph learning model named GForce inspired by the spring-electrical model. GForce assumes that nodes are in attractive forces and repulsive forces, thus leading to the same representation with the original structural information in feature learning. Comprehensive experiments on three benchmark datasets demonstrate the effectiveness of the proposed framework. Furthermore, GForce opens up opportunities to use physics models to model node interaction for graph learning.
Ke Sun 0011, Jiaying Liu 0006, Shuo Yu 0001, Bo Xu 0008, Feng Xia 0001
IEEE BigData4
2020 Protein Complexes Identification with Family-Wise Error Rate Control
abstract
The detection of protein complexes from protein-protein interaction network is a fundamental issue in bioinformatics and systems biology. To solve this problem, numerous methods have been proposed from different angles in the past decades. However, the study on detecting statistically significant protein complexes still has not received much attention. Although there are a few methods available in the literature for identifying statistically significant protein complexes, none of these methods can provide a more strict control on the error rate of a protein complex in terms of family-wise error rate (FWER). In this paper, we propose a new detection method SSF that is capable of controlling the FWER of each reported protein complex. More precisely, we first present a p-value calculation method based on Fisher's exact test to quantify the association between each protein and a given candidate protein complex. Consequently, we describe the key modules of the SSF algorithm: a seed expansion procedure for significant protein complexes search and a set cover strategy for redundancy elimination. The experimental results on five benchmark data sets show that: (1) our method can achieve the highest precision; (2) it outperforms three competing methods in terms of normalized mutual information (NMI) and F1 score in most cases.
Zengyou He, Can Zhao 0007, Bo Xu 0008, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 MODEL: Motif-Based Deep Feature Learning for Link Prediction
abstract
Link prediction plays an important role in network analysis and applications. Recently, approaches for link prediction have evolved from traditional similarity-based algorithms into embedding-based algorithms. However, most existing approaches fail to exploit the fact that real-world networks are different from random networks. In particular, real-world networks are known to contain motifs, natural network building blocks reflecting the underlying network-generating processes. In this article, we propose a novel embedding algorithm that incorporates network motifs to capture higher order structures in the network. To evaluate its effectiveness for link prediction, experiments were conducted on three types of networks: social networks, biological networks, and academic networks. The results demonstrate that our algorithm outperforms both the traditional similarity-based algorithms (by 20%) and the state-of-the-art embedding-based algorithms (by 19%).
Lei Wang 0134, Jing Ren 0001, Bo Xu 0008, Jianxin Li 0001, Wei Luo 0001, Feng Xia 0001
IEEE Trans. Comput. Soc. Syst.3
2019 Disease Gene Prediction Based on Heterogeneous Probabilistic Hypergraph Ranking
abstract
In order to save time and cost, many disease gene prediction methods have been proposed in recent years. However, the traditional network model uses a binary relationship to represent the relationship between different proteins or gene molecules and phenotypes, which leads to the loss of information. Recently, hypergraph shows that it can overcome this loss of information to some extent and preserve the multivariate relationship, so we transformed the disease gene prediction problem into the problem of ranking the multivariate-relationship object. In this paper, we propose a method of Heterogeneous Probabilistic Hypergraph Ranking (HPHR) to predict disease genes. Firstly, fix a graph centroid for each hyperedge and according to different associations, and add other nodes related to the graph centroid to hyperedges with a certain probability. Then transform the problem of predicting disease genes into the problem of ranking heterogeneous objects, and the candidate genes are sorted by hypergraph ranking. The method is then applied to the integrated disease gene network. Compared with other prediction methods achieved better results, which was verified by this experiment.
Feng Ding 0004, Xiangjie Kong 0001, Zhehuan Zhao, Feng Xia 0001, Anfu Liu, Chenxu Bai, Bo Xu 0008, Shengtian Sang, Hongfei Lin, Jian Wang 0021
BIBM7
2019 Detection of protein complexes from multiple protein interaction networks using graph embedding
Shengtian Sang, Hongfei Lin, Jian Wang 0021, Bo Xu 0008
Artif. Intell. Medicine6
2019 Double warning thresholds for preemptive charging scheduling in Wireless Rechargeable Sensor Networks
Chi Lin 0001, Yu Sun 0077, Zhunyue Chen, Bo Xu 0008, Guowei Wu 0001
Comput. Networks5
2019 CSTeller: forecasting scientific collaboration sustainability based on extreme gradient boosting
Wei Wang 0077, Bo Xu 0008, Jiaying Liu 0006, Zixin Cui, Shuo Yu 0001, Xiangjie Kong 0001, Feng Xia 0001
World Wide Web2
2018 PCM: A Pairwise Correlation Mining Package for Biological Network Inference
Feiyang Gu, Chaohua Sheng, Qiong Duan, Bo Xu 0008, Zengyou He
ICIC (2)7
2018 Identifying protein complexes based on node embeddings obtained from protein-protein interaction networks
abstract
BACKGROUND: Protein complexes are one of the keys to deciphering the behavior of a cell system. During the past decade, most computational approaches used to identify protein complexes have been based on discovering densely connected subgraphs in protein-protein interaction (PPI) networks. However, many true complexes are not dense subgraphs and these approaches show limited performances for detecting protein complexes from PPI networks. RESULTS: To solve these problems, in this paper we propose a supervised learning method based on network node embeddings which utilizes the informative properties of known complexes to guide the search process for new protein complexes. First, node embeddings are obtained from human protein interaction network. Then the protein interactions are weighted through the similarities between node embeddings. After that, the supervised learning method is used to detect protein complexes. Then the random forest model is used to filter the candidate complexes in order to obtain the final predicted complexes. Experimental results on real human and yeast protein interaction networks show that our method effectively improves the performance for protein complex detection. CONCLUSIONS: We provided a new method for identifying protein complexes from human and yeast protein interaction networks, which has great potential to benefit the field of protein complex detection.
Shengtian Sang, Lei Wang 0085, Yin Zhang 0009, Hongfei Lin, Jian Wang 0021, Bo Xu 0008
BMC Bioinform.9
2018 Protein complexes identification based on go attributed network embedding
abstract
BACKGROUND: Identifying protein complexes from protein-protein interaction (PPI) network is one of the most important tasks in proteomics. Existing computational methods try to incorporate a variety of biological evidences to enhance the quality of predicted complexes. However, it is still a challenge to integrate different types of biological information into the complexes discovery process under a unified framework. Recently, attributed network embedding methods have be proved to be remarkably effective in generating vector representations for nodes in the network. In the transformed vector space, both the topological proximity and node attributed affinity between different nodes are preserved. Therefore, such attributed network embedding methods provide us a unified framework to integrate various biological evidences into the protein complexes identification process. RESULTS: In this article, we propose a new method called GANE to predict protein complexes based on Gene Ontology (GO) attributed network embedding. Firstly, it learns the vector representation for each protein from a GO attributed PPI network. Based on the pair-wise vector representation similarity, a weighted adjacency matrix is constructed. Secondly, it uses the clique mining method to generate candidate cores. Consequently, seed cores are obtained by ranking candidate cores based on their densities on the weighted adjacency matrix and removing redundant cores. For each seed core, its attachments are the proteins with correlation score that is larger than a given threshold. The combination of a seed core and its attachment proteins is reported as a predicted protein complex by the GANE algorithm. For performance evaluation, we compared GANE with six protein complex identification methods on five yeast PPI networks. Experimental results showes that GANE performs better than the competing algorithms in terms of different evaluation metrics. CONCLUSIONS: GANE provides a framework that integrate many valuable and different biological information into the task of protein complex identification. The protein vector representation learned from our attributed PPI network can also be used in other tasks, such as PPI prediction and disease gene prediction.
Bo Xu 0008, Wei Zheng 0003, Yi-Jia Zhang 0001, Zhehuan Zhao, Zengyou He
BMC Bioinform.1
2015 VCLT: An Accurate Trajectory Tracking Attack Based on Crowdsourcing in VANETs
Chi Lin 0001, Bo Xu 0008, Jing Deng 0001, James Chang Wu Yu, Guowei Wu 0001
ICA3PP (3)3
2014 Graphical lasso quadratic discriminant function and its application to character recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi
Neurocomputing1
2012 Maxi-Min discriminant analysis via online learning
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
Neural Networks1
2011 Graphical Lasso Quadratic Discriminant Function for Character Recognition
Bo Xu 0008, Kaizhu Huang, Irwin King, Cheng-Lin Liu 0001, Jun Sun 0004, Satoshi Naoi
ICONIP (3)1
2010 Ontology integration to identify protein complex in protein interaction networks
abstract
Protein complexes can be identified from the protein interaction networks derived from experimental data sets. However, these analyses are challenging because of the presence of unreliable interactions and the complex connectivity of the network. The integration of protein-protein interactions with the data from other sources can be leveraged for improving the effectiveness of protein complexes detection algorithms. We have developed novel semantic similarity method, which use Gene Ontology (GO) annotations to measure the reliability of protein-protein interactions. The protein interaction networks can be converted into a weighted graph representation by assigning the reliability values to each interaction as a weight. Following the approach of previously proposed clustering algorithm IPCA which expands clusters starting from seeded vertices, we present a clustering algorithm OIIP based on the new weighted Protein-Protein interaction networks for identifying protein complexes. The algorithm OIIP is applied to the protein interaction network of Sacchromyces cerevisiae and identifies many well known complexes. Experimental results show that the algorithm OIIP has higher F-measure and accuracy compared to other competing approaches.
Bo Xu 0008, Hongfei Lin
BIBM1
2010 Similar Handwritten Chinese Characters Recognition by Critical Region Selection Based on Average Symmetric Uncertainty
abstract
We consider the problem of similar Chinese character recognition in this paper. Engaging the Average Symmetric Uncertainty (ASU) criterion to measure the correlation between different image regions and the class label, we manage to detect the most critical regions for each pair of similar characters. These critical regions are proved to contain more discriminative information and hence can largely benefit the classification accuracy for similar characters. We conduct a series of experiments on the CASIA Chinese character data set. Experimental results show that our proposed method is superior to three competitive approaches in terms of both accuracy and efficiency.
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
ICFHR1
2010 Dimensionality Reduction by Minimal Distance Maximization
abstract
In this paper, we propose a novel discriminant analysis method, called Minimal Distance Maximization (MDM). In contrast to the traditional LDA, which actually maximizes the average divergence among classes, MDM attempts to find a low-dimensional subspace that maximizes the minimal (worst-case) divergence among classes. This ``minimal" setting solves the problem caused by the ``average" setting of LDA that tends to merge similar classes with smaller divergence when used for multi-class data. Furthermore, we elegantly formulate the worst-case problem as a convex problem, making the algorithm solvable for larger data sets. Experimental results demonstrate the advantages of our proposed method against five other competitive approaches on one synthetic and six real-life data sets.
Bo Xu 0008, Kaizhu Huang, Cheng-Lin Liu 0001
ICPR1