Liang Zhao 0005

dblp:63/5422-5 · DBLP profile ↗
← Back
77ranked-venue papers
37as first author
67since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 12 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 17 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 10 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Sample Weighted Incomplete Multimodal Clustering Based on Graph Coarsening Label Extraction
abstract
Multimodal data is typically collected through heterogeneous sensors and processing pipelines. However, due to variations in acquisition environments, device capabilities, and feature extraction methods, such data often suffers from incompleteness and inconsistent quality across modalities. To address these challenges, prior studies have explored modality selection and data completion strategies to improve information fusion. Nevertheless, these approaches face two main limitations: (1) they struggle to simultaneously ensure computational efficiency for large-scale graph data and maintain structural and semantic consistency across heterogeneous modality graphs; and (2) most of them operate at the modality level and fail to capture fine-grained, sample-specific quality variations. To overcome these issues, we propose a novel clustering framework, Sample Weighted Incomplete Multimodal Clustering Based on Graph Coarsening Label Extraction (IMC-GCSW). The proposed method introduces a graph coarsening-based label extraction strategy. It significantly reduces the computational cost of multimodal graph processing, while preserving key node information and local topological structures. Furthermore, a quality-aware sample weighting strategy is designed to enable fine-grained modeling of modality-specific data quality, allowing the model to dynamically suppress the influence of low-quality modalities on individual samples. Experiments on both general-purpose datasets and the Fructus Aurantii Disease and Pest Datasets demonstrate that the proposed method exhibits superior performance and strong adaptability in handling multimodal data with incompleteness and quality inconsistency.
Zhenjiao Liu, Jiao Xue, Shubin Ma, Liang Zhao 0005
AAAI6
2026 KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering
abstract
In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the partially view-aligned problem (PVP). Most existing partially view-aligned clustering (PVC) methods first learn cross-view consistent representations based on known alignments, and then recover missing correspondences by measuring cross-view similarity between samples. However, such an indirect alignment recovery process depends on high-quality consistent representations and lacks effective utilization of known alignments, often resulting in sub-optimal outcomes. To address this, we propose a novel direct alignment recovery perspective, instantiated as K-Nearest Neighbors Direct Alignment (KNNDA). Specifically, we first construct an alignment domain by mapping the aligned neighbors of each unaligned sample into the aligned view. Then, we compute alignment confidence based on the similarity between known aligned pairs of neighbors. In particular, we use a dynamic threshold to filter out unreliable alignments. Finally, new alignments are generated within the high-confidence alignment domain. Contrastive loss is used to learn consistent representations for clustering. Comprehensive experiments on several real-world datasets show the effectiveness and superiority of our module in partially view-aligned clustering.
Liang Zhao 0005, Tianqi Yue, Shubin Ma, Bo Xu 0008
AAAI1
2026 Adaptive multi-agent stock trading decision support system based on deep reinforcement learning
Xu Yuan 0002, Jiaqiang Wang, Shaokui Gu, Ange Qi, Liang Zhao 0005
Eng. Appl. Artif. Intell.7
2026 Double-incomplete multi-view clustering with self-induced semantic label diffusion
Zhikui Chen, Meng Liu 0025, Yuzhe Li 0002, Zhenjiao Liu, Liang Zhao 0005
Inf. Sci.7
2026 Dual-Selection and Optimal Transport for Robust Incomplete and Unaligned Multi-view Clustering
Liang Zhao 0005, Shubin Ma, Chuanye He, Chenhui Yao
Pattern Recognit.1
2026 Multi-View Aligned Clustering via Sample-Bundled Optimization: Anchor Graph Enhancement and Contrastive Propagation
abstract
Multi-view representation is powerful to capture the complex characteristics of real-world data by integrating complementary information from various modalities. However, in many use cases, such as boiler combustion monitoring, factors including sensor sampling frequency, equipment malfunctions, and network delays can lead to temporal asynchrony in data collection. This asynchrony leads misaligned multi-modal data, furthering the difficulty of learning optimal fused representation. To address this misalignment in multi-view data, a body of methods based on autoencoders and non-negative matrix factorization (NMF) have been presented. However, those methods are incapable of jointly exploring the underlying structure inherent in each view as well as the semantic consistency and structural similarity across views. To these ends, we propose a novel sample-bundled optimization for multi-view aligned clustering, which is based on Enhanced Anchor Graph and Contrastive Propagation (termed asEAGCP). We state that there is semantic consistency among intra-class samples (with the same and cross views) and that the global structures across different views demonstrate similarity. By leveraging these associations, we introduce an enhanced anchor graph with learnable sample correlation and a contrastive graph with feature propagation. Specifically, each anchor graphs preserves the semantic relationship among same-view samples, while the contrastive graph propagates feature information across multi-view samples. Experimental results demonstrate the superiority of our method on benchmark datasets by validating its effectiveness in aligning and clustering multi-view data.
Shubin Ma, Zhikui Chen, Lin Wu 0001, Liang Zhao 0005
IEEE Trans. Multim.4
2025 Incomplete and Unpaired Multi-View Graph Clustering with Cross-View Feature Fusion
abstract
Due to its effectiveness and efficiency, graph-based multi-view clustering has recently attracted much attention. However, the multi-view data are often incomplete and unpaired in real-world applications as a consequence of data loss or corruption. Although efforts have been made through a series of methods to address the problems of incomplete or unpaired multi-view data, the following issues still persist: 1) Most existing methods only focus on the incomplete multi-view data or unpaired multi-view data, and exhibit weaknesses when addressing both incomplete and unpaired multi-view data simultaneously. 2) Some methods neglect the graph information of the data from different views during the learning process. To tackle these issues, we propose the Multi-view Graph Clustering framework with Cross-view Feature Fusion (MGCCFF), a novel approach for clustering incomplete and unpaired multi-view data. Specifically, MGCCFF learns soft clustering label information from complete data and utilizes this to capture category-level cross-view correspondences. It then learns latent representation enriched with cross-view information based on the established mappings. To obtain a multi-view graph structure under conditions of incomplete and unpaired data, MGCCFF innovatively integrates the concept of self-expression with the autoencoder architecture and exploits the latent relationships between labels and the graph structure, thereby enabling the generation of sparse and accurate graphical structure under multi-view conditions for the final clustering task. The experiments on incomplete and unpaired multi-view datasets demonstrate that MGCCFF outperforms state-of-the-art methods.
Liang Zhao 0005, Zhikui Chen, Bo Xu 0008
AAAI1
2025 Time-Sequential Lung CT Image Dynamic Registration Model with Weakly Supervised Learning
abstract
Despite significant advancements in medical image registration, current techniques face limitations in accurately aligning images captured at different time points, particularly for lung imaging. Traditional registration methods are often laborintensive and computationally complex, while deep learning-based approaches, though faster, may fail to capture temporal biological changes such as nodule growth. To address this challenge, we proposes a weakly supervised registration method for lung CT images. In this method, a deep learning-based deformation field estimation mechanism is designed, which is built upon the VoxelMorph architecture. This mechanism can effectively estimate the deformation field that aligns images from different time points by learning the spatial relationships between lung regions. To further enhance the registration accuracy and ensure that the temporal biological changes are properly reflected, a novel growth consistency loss function is designed into the training process. This loss function not only ensures the similarity between the registered images but also captures the temporal growth patterns of lung nodules, thus providing a more accurate representation of the biological changes over time. Experimental results demonstrate that the proposed method achieves significant improvements in registration accuracy. These findings highlight the effectiveness of incorporating growth consistency into the registration framework and underscore the importance of balancing similarity and growth-related losses for optimal performance.
Mengru Ouyang, Chaoran Jia, Liang Zhao 0005
BIBM5
2025 MTMS: a Multi-Modal MRI Synthesis Model Via Multi-Task Learning
abstract
Different magnetic resonance imaging modalities can provide multi-dimensional information. However, in clinical practice, some modal data are often missing due to factors such as varying scanning times or insufficient patient compliance. Modality generation technology can generate missing modal images from existing modal images to supplement the missing information.However, existing MR image synthesis methods face two main challenges: insufficient utilization of anatomical priors in tumor regions, and the lack of independently designed fusion modules, which hinders the deep mining and collaborative integration of features from multi-source input modalities. Therefore, this paper proposes a multi-modal medical image synthesis model based on multi-task learning. The model mines modality-specific private features through a dual-layer feature extraction mechanism, designs the MIEF module to extract complementary common semantics between modalities, and finally strengthens the extraction of key features such as tumor boundaries through the segmentation task. Additionally, a multi-task loss constraint is constructed to promote the model to generate more realistic missing modal images. Finally, we verified the effectiveness and superiority of our model on the BraTS23 dataset. Experiments show that our model has significant improvements in the three metrics of NMSE, SSIM, and PSNR. Our code will be released at https://github.com/wxy2024/MTMS.git.
Liang Zhao 0005
BIBM1
2025 Partially View-aligned Clustering with Unbiased Semantic Learning
abstract
Recently, the issue of the partially view-aligned problem has drawn more attention. While some studies have achieved satisfactory performance in downstream tasks, they all overlook the information contained in unaligned data, leading to biases in the semantic information and spatial structures learned by the model. To tackle it, we have designed an unbiased semantic learning that incorporates unaligned data into the training process. To the best of our knowledge, this is the first method that can learn unbiased semantic and spatial information from partially view-aligned data. Specifically, we utilize anchor graphs to explore intra-view associations among all samples. Based on the assumption of semantic consistency, anchors serve as intermediaries to facilitate the transfer of neighborhood information across views. By leveraging this information to uncover unbiased semantic and spatial information, we can provide higher-quality, realigned latent representations for subsequent clustering tasks. Experiments on several datasets demonstrate the superiority of our method.
Liang Zhao 0005, Yukun Yuan 0002
ICME1
2025 Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search
abstract
Multi-modal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and misaligned characteristics due to factors such as inconsistent sensor frequencies and device malfunctions. Existing research has not effectively addressed the issue of filling missing data in scenarios where multiview data are both imbalanced and misaligned. Instead, it relies on class-level alignment of the available data. Thus, it results in some data samples not being well-matched, thereby affecting the quality of data fusion. In this paper, we propose the Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search(CAPIMAC) to tackle the problem of filling imbalanced and misaligned data in multi-modal datasets. Specifically, we propose a self-repellent greedy anchor search module(SRGASM), which employs a self-repellent random walk combined with a greedy algorithm to identify anchor points for re-representing incomplete and misaligned multi-modal data. Subsequently, based on noise-contrastive learning, we design a consistency-aware padding module (CAPM) to effectively interpolate and align imbalanced and misaligned data, thereby improving the quality of multi-modal data fusion. Experimental results demonstrate the superiority of our method over benchmark datasets. The code will be publicly released at https://github.com/bestow09090/-CAPIMAC.git.
Shubin Ma, Liang Zhao 0005, Mingdong Lu, Bo Xu 0008
IJCAI2
2025 EchoGPT: An Interactive Cardiac Function Assessment Model for Echocardiogram Videos
abstract
With the development of wearable cardiac ultrasound devices, it is no longer sufficient to solely rely on doctors for diagnosing long-term echocardiogram videos. Automated diagnosis of echocardiogram videos has now become a research hotspot. Existing studies only analyze echocardiogram video through discriminative models, which have limited question-answering capabilities. Therefore, this study innovatively proposes a large language model with cardiac ultrasound diagnostic capabilities—EchoGPT. EchoGPT integrates the robust communication and comprehension capabilities of large language models (LLMs) with the diagnostic prowess of traditional medical models, empowering patients to obtain accurate medical indicator data and comprehend their health conditions through interactive questioning with the model. The model is capable of local deployment on personal computers, effectively safe guarding user privacy. EchoGPT operates through three main components: left ventricle segmentation, left ventricular ejection fraction LVEF prediction, and finetuning of video-text LLMs. Experimental results demonstrate EchoGPT’s superior accuracy in predicting LVEF compared to other models, and positive feedback from professional physicians through questionnaire surveys, validating its potential in practical applications. The demo is available at https://github.com/zhuqh19/EchoGPT.
Bo Xu 0008, Quanhao Zhu, Qingchen Zhang 0001, Mengmeng Wang 0005, Liang Zhao 0005, Hongfei Lin, Jing Ren 0001, Feng Xia 0001
IJCAI5
2025 CSF-GAN: Cross-modal Semantic Fusion-based Generative Adversarial Network for Text-guided Image Inpainting
abstract
Most visual-guided image inpainting methods based on generative adversarial networks (GANs) struggle when the missing region has weak correlations with the surrounding visual context. Recently, diffusion-based methods guided by textual context have been proposed to address this limitation by leveraging additional semantic information to restore corrupted objects. However, these models typically involve more parameters and exhibit slower generation speeds compared to GAN-based approaches. To address this problem, we propose a novel text-guided image inpainting model, the cross-modal semantic fusion generative adversarial network (CSF-GAN). CSF-GAN is designed as a one-stage GAN with the following key contributions. First, a novel semantic fusion module (SFM) is introduced to integrate sentence- and word-level textual context into the inpainting process, enabling more effective guidance from multi-granularity semantic information. Second, a newly designed word-level local discriminator provides detailed feedback to the generator, enhancing the accuracy of generated content in alignment with word-level semantics. Third, two loss functions, the inpainting loss and edge loss, are employed to enhance both structural coherence and textural realism in the generated results. Extensive experiments on two benchmark datasets demonstrate that CSF-GAN outperforms state-of-the-art methods.
Suixue Wang, Qingchen Zhang 0001, Liang Zhao 0005, Weiliang Huo, Sijia Hou, Chunjiang Fu
IJCAI4
2025 Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information
abstract
Recently, multi-view data has gradually attracted attention. However, real-world applications often face Partial View-aligned Problem (PVP) and Partially Sample-missing Problem (PSP) due to data loss or corruption. Existing methods addressing PVP typically focus only on learning from the information of aligned data, while ignoring unaligned data where samples exist but lack alignment relationships. This introduces PSP, which does not inherently exist in the data, leading to biased learning of the data's information. For PSP, due to varying degrees of missing data, incomplete spatial structures can cause clustering centers-shifted problem, resulting in the model learning incorrect correspondences and biased spatial structures.To tackle them, we propose a novel method called Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information (DRUMVC). To our knowledge, this is the first noise-robust and unbiased multi-view clustering method capable of simultaneously addressing both PVP and PSP. Specifically, DRUMVC leverages aligned and complete samples as a bridge to construct high-quality correspondences for samples lacking cross-view relationship information due to PVP or PSP. Additionally, we employ a dual noise-robust contrastive learning loss to mitigate the impact of noise potentially introduced during the pair construction. Experiments on several challenging datasets demonstrate the superiority of our proposed method.
Liang Zhao 0005, Chuanye He, Qingchen Zhang 0001, Bo Xu 0008
IJCAI1
2025 Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly Data
abstract
Multimodal feature fusion, by integrating the complementary information from each modality, can effectively capture complex features in real-world data. However, in many use cases, such as boiler combustion monitoring, factors including equipment failure, inconsistent sensor sampling frequencies, and network delays often cause data collected from different modalities to suffer from missing modality and temporal asynchrony. This leads to the incompleteness and disorderliness of multimodal data. To address these issues, previous studies have proposed several data fusion methods that align the cluster centers before fusion. However, these approaches have two key limitations: 1) they do not guarantee a high alignment accuracy of data pairs at the sample level, and 2) they do not address the issue of significant discrepancies in data sizes across different classes, which impacts the subsequent data fusion performance.
Liang Zhao 0005, Shubin Ma, Bo Xu 0008, Qingchen Zhang 0001
ACM Multimedia1
2025 Historical Trends and Normalizing Flow for One-shot Temporal Knowledge Graph Reasoning
Ruixin Ma, Huinan Wu, Buyun Gao, Xiaoru Wang, Liang Zhao 0005
Expert Syst. Appl.6
2025 TFE-HMM: Combining state transitions and price trends for stock price forecasting
Xu Yuan 0002, Shaokui Gu, Jiaqiang Wang, Ange Qi, Liang Zhao 0005
Expert Syst. Appl.7
2025 MHEC: One-shot relational learning of knowledge graphs completion based on multi-hop information enhancement
Ruixin Ma, Buyun Gao, Weihe Wang, Xiaoru Wang, Liang Zhao 0005
Neurocomputing6
2025 Normalizing flow-enhanced Gaussian embedding for few-shot knowledge graph completion
Xu Yuan 0002, Jiaqiang Wang, Zhengnan Gao, Liang Zhao 0005
Knowl. Based Syst.6
2025 Cross-modal feature alignment and fusion with contrastive learning in multimodal recommendation
Xu Yuan 0002, Ange Qi, Huinan Wu, Jiaqiang Wang, Liang Zhao 0005
Knowl. Based Syst.7
2025 Pseudo-Label Guided Incomplete Partial View-Aligned Clustering
abstract
Addressing the challenges of incomplete and misaligned data in multi-view learning is critical, giventhe inherent uncertainties and complexities of real-world data collection. These challenges often result in significant discrepancies in the quantity, quality, and completeness of data across different views. However, previous research has predominantly focused on addressing either incompleteness or misalignment in isolation. To address both incompleteness and misalignment simultaneously, we propose a novel model, Pseudo-Label Guided Incomplete Partial View-aligned Clustering (PGIPVC). Specifically, A pseudo-label acquisition module based on Cauchy divergence is proposed to preliminarily train the clustering structure of data in a single view, thereby obtaining pseudo-labels for each sample in the view. Subsequently, an incomplete partial alignment clustering module is designed to obtain discriminative latent representations through contrastive learning with positive-negative pairs selected based on KNN and pseudo-labeled samples. Extensive experiments on benchmark datasets demonstrate the superiority of our method compared to other state-of-the-art approaches.
Shubin Ma, Liang Zhao 0005, Songtao Wu, Bo Xu 0008
IEEE Signal Process. Lett.2
2025 Learnable Graph Guided Deep Multi-View Representation Learning via Information Bottleneck
abstract
In real world applications, multi-view data has attracted intensive attention due to the complex and complementary relationship across views. Multi-view representation learning (MvRL) focuses on obtaining consistent feature representation from multi-view data, and becomes a popular topic in multi-view research field. However, the relationship between different samples, i.e., the graph information, is usually ignored or excavated insufficiently in most existing MvRL methods, which only regard graph structure as regularization items instead of graph embedding for multi-view data. Besides, the limited learning capacity of the adopted shallow models is another challenge for MvRL. To tackle them, in this paper, we propose a novel unsupervised deep multi-view representation learning model guided by learnable graph structure, termed as LGG-DMRL. It first captures a multi-view consistent graph from original data based on self-representation learning, and explores the view-specific feature representation of each view by the designed graph guided attention network using the learnt graph. After that, the information bottleneck principle is employed to identify the shared representation across views integrated with the view-specific feature representations, promoting the multi-view complementarity and completeness. Experimental results on five real-world datasets demonstrate the superiority and effectiveness of our proposed LGG-DMRL compared with the recent state-of-the-art multi-view approaches.
Liang Zhao 0005, Zhenjiao Liu, Zhikui Chen
IEEE Trans. Circuits Syst. Video Technol.1
2025 Dynamic Graph Guided Progressive Partial View-Aligned Clustering
abstract
In recent years, there has been a growing focus on multiview data, driven by its rich complementary and consistent information, which has the potential to significantly enhance the performance of downstream tasks. Although many multiview clustering (MVC) methods have achieved promising results by integrating the information of multiple views to learn the consistent representation or consistent graph, these methods typically require complete and entirely accurate correspondences between multiview data, which is challenging to fulfill in practice leading to the problem of partially view-aligned clustering (PVC). To tackle it, we propose a novel method, called dynamic graph guided progressive partial view-aligned clustering (DGPPVC) in this article. To the best of our knowledge, this could be the first work to employ graph convolutional network (GCN) to address the problem of PVC, which explores GCN with dynamic adjacency matrix to reduce unreliable alignments and locate the feature representation with consistent graph structure. In particular, DGPPVC develops an end-to-end framework that encompasses graph construction, feature representation learning, and alignment relationships learning, in which the three parts mutually influence and benefit each other. Moreover, DGPPVC adopts a novel alignment learning strategy that progresses from simplicity to complexity, enabling the step-by-step acquisition of unknown correspondences between different modalities. By giving priority to simple instance pairs, a variant of Jaccard similarities is designed to identify more reliable and complex alignments progressively. During the gradual learning process of alignment relationships, the graph structure matrix is continually and dynamically optimized, thus acquiring a greater variety of graph information between different views. Experiments on several real-world datasets show our promising performance compared with the state-of-the-art methods in partially view-aligned clustering.
Liang Zhao 0005, Qiongjie Xie, Zhengtao Li, Songtao Wu, Yi Yang 0006
IEEE Trans. Neural Networks Learn. Syst.1
2024 Alzheimer's Disease Prediction with Irregular MRI Sequences Based on Pyramid Squeeze Attention and Time-Sensitive Attention Mechanisms
abstract
Alzheimer’s disease (AD) often leads to cognitive impairments and behavioral issues, making early detection and treatment crucial for slowing the progression of the disease and improving quality of life. This type of dementia typically worsens gradually, so accurately predicting the course of the disease is critical for medical intervention. To tackle it, this paper develops a new type of time series Alzheimer’s disease prediction model (ChaoJiBang-Net). This model integrates pyramid squeeze attention mechanisms, time position encoding, and time-sensitive attention techniques. It is specifically designed for irregular sampling and variable-length sequences in MRI images. The model aims to predict specific time points for Alzheimer’s disease, which are arbitrarily chosen. Experimental results show that increasing the amount of temporal data (time steps) can significantly improve the model’s prediction performance, as reflected in the progressively higher AUC values. Furthermore, this model can perform accurate predictions at any number of time steps, reducing reliance on frequent patient follow-ups and enhancing its clinical applicability. Currently, the predictive accuracy of this model, using MRI single modality data, surpasses the highest standards of existing state-of-the-art (SOTA) technologies.
Baijiang Xu, Bo Xu 0008, Liang Zhao 0005
BIBM6
2024 Construction of Simulated 3D CT from Multiple-View X-ray Images
Liang Zhao 0005, Sijia Hou, Xin Fan 0001, Zhikui Chen, Shuqiong Wu
BIBM1
2024 Copd-ChatGLM: A Chronic Obstructive Pulmonary Disease Diagnostic Model
abstract
COPD is a chronic lung condition characterized by persistent respiratory obstruction and airflow limitation. Early detection and diagnosis are crucial to manage the disease effectively and improve patient outcomes. However, many patients, especially in low- and middle-income countries, are diagnosed late due to limited access to spirometry. This study proposes Copd-ChatGLM, a large language model fine-tuned for COPD diagnosis and management using the RAG framework. The model combines deep learning with clinical data to enhance diagnostic accuracy and provide personalized treatment plans. By integrating LoRA, Copd-ChatGLM fine-tunes the ChatGLM3-6B model efficiently to perform COPD-related tasks with minimal computational resources. Experimental results show that Copd-ChatGLM outperforms traditional classification models and general large language models in accuracy, sensitivity, specificity, and F1 score. This model has become a robust and clinically applicable tool for managing COPD, particularly in resource-limited settings.
Liang Zhao 0005, Zhanxin Gang, Baijiang Xu
BIBM1
2024 Multimodal contrastive learning with neuroimaging and cognitive tests for Alzheimer's disease diagnosis
abstract
Alzheimer’s disease (AD) is a neurological illness that causes cognitive impairment. Computer-aided diagnosis can help diagnose Alzheimer’s disease early before clinical symptoms appear. Currently, many deep learning methods show good performance in AD diagnosis. Still, most of these methods are based on single/multimodal neuroimaging, leading to a one-sided approach to disease modelling. Combining neuroimaging, cognitive tests, and demographics can significantly improve model performance and reduce the negative impact of noise. This study proposes a multimodal model that introduces contrastive learning, extracting and fusing feature representations separately from cognitive tests and neuroimaging data. After that, contrastive learning based on similarity is employed for both modalities’ features, assisting the network in learning cross-modal features. Moreover, the hybrid attention mechanism of the Transformer encoder is explored for feature fusion. Experimental results on 2082 cases from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset validate the effectiveness of our proposed multimodal model.
Liang Zhao 0005, Bo Xu 0008, Yi Yang 0006, Yangqianhui Zhang, Ruixin Ma
BIBM1
2024 Anchor Based Multi-view Clustering for Partially View-Aligned Data
abstract
Clustering on the partially view-aligned multi-view data, where only a small portion of the cross-view correspondences is accurately mapped, has gained more and more attention in recent years. However, current methods need to train on the aligned instances before reconstructing the missing cross-view relationships, which significantly increases the time burden. Therefore, this paper proposes an efficient and effective method termed Anchor-based Multi-view Clustering for Partially view-aligned data (AMCP). We design a method to directly measure the cross-view similarity by utilizing aligned instances as anchors since similar cross-view instances should exhibit similar within-view indicator vectors of anchors. Moreover, to further improve the accuracy and robustness of the cross-view relationship, we consider the nearest neighbors of instances and compare their similarity to refine the relationship, which helps to deal with the low alignment rate settings. As a result, our method demonstrates competitive performance on several real-world multi-view datasets.
Liang Zhao 0005, Yukun Yuan 0002, Qiongjie Xie
ICME1
2024 Generating Multimodal Metaphorical Features for Meme Understanding
abstract
Understanding a meme is a challenging task, due to the metaphorical information contained in the meme that requires intricate interpretation to grasp its intended meaning fully. In previous works, attempts have been made to facilitate computational understanding of memes through introducing human-annotated metaphors as extra input features into machine learning models. However, these approaches mainly focus on formulating linguistic representation of a metaphor (extracted from the texts appearing in memes), while ignoring the connection between the metaphor and corresponding visual features (e.g., objects in meme images). In this paper, we argue that a more comprehensive understanding of memes can only be achieved through a joint modelling of both visual and linguistic features of memes. To this end, we propose an approach to generate Multimodal Metaphorical feature for Meme Classification, named MMMC. MMMC derives visual characteristics from linguistic attributes of metaphorical concepts, which more effectively convey the underlying metaphorical concept, leveraging a text-conditioned generative adversarial network. The linguistic and visual features are then integrated into a set of multimodal metaphorical features for classification purpose. We perform extensive experiments on a benchmark metaphorical meme dataset, MET-Meme. Experimental results show that MMMC significantly outperforms existing baselines on the task of emotion classification and intention detection. Our code and dataset are available at https://github.com/liaolianfoka/MMMC.
Bo Xu 0008, Junzhe Zheng, Estrid He, Hongfei Lin, Liang Zhao 0005, Feng Xia 0001
ACM Multimedia6
2024 Joint long and short span self-attention network for multi-view classification
Zhikui Chen, Kai Lou, Zhenjiao Liu, Yue Li 0050, Liang Zhao 0005
Expert Syst. Appl.6
2024 Knowledge graph preference migration network for recommendation
Ruixin Ma, Xiya Bu, Huinan Wu, Liang Zhao 0005
Expert Syst. Appl.6
2024 GLSEC: Global and local semantic-enhanced contrastive framework for knowledge graph completion
Ruixin Ma, Xiaoru Wang, Cunxi Cao, Xiya Bu, Liang Zhao 0005
Expert Syst. Appl.6
2024 Multi-view semantic enhancement model for few-shot knowledge graph completion
Ruixin Ma, Xiaoru Wang, Weihe Wang, Liang Zhao 0005
Expert Syst. Appl.6
2024 Knowledge graph fine-grained network with attribute transfer for recommendation
Xu Yuan 0002, Xiya Bu, Zhengnan Gao, Liang Zhao 0005, Ruixin Ma
Expert Syst. Appl.5
2024 CCIM-SLR: Incomplete multiview co-clustering by sparse low-rank representation
Zhenjiao Liu, Zhikui Chen, Kai Lou, Praboda Rajapaksha, Liang Zhao 0005, Noël Crespi, Xiaodi Huang 0001
Multim. Tools Appl.5
2024 Multimodal Fusion Generative Adversarial Network for Image Synthesis
abstract
Text-to-image synthesis has advanced significantly; however, a crucial limitation persists: textual descriptions often neglect essential background details, leading to blurred backgrounds and diminished image quality. To address this, we propose a multimodal fusion framework that integrates information from both text and image modalities. Our approach introduces a background mask to compensate for missing textual descriptions of background elements. Additionally, we employ an adaptive channel attention mechanism to effectively exploit fused features, dynamically accentuating informative feature maps. Furthermore, we introduce a novel fusion conditional loss, ensuring that generated images not only align with textual descriptions but also exhibit realistic backgrounds. Experimental evaluations on the Caltech-UCSD Birds 200 and COCO datasets demonstrate the efficacy of our approach, with our Frechet Inception Distance (FID) achieving a commendable score of 15.38 on the CUB dataset, surpassing several state-of-the-art approaches.
Liang Zhao 0005, Qinghao Hu 0003
IEEE Signal Process. Lett.1
2024 Distribution-Level Multi-View Clustering for Unaligned Data
abstract
Recently, many multi-view clustering (MVC) methods have achieved promising results through integrating complementary and consensus information from different views in the fields of signal processing and machine learning. However, most of the methods require complete or partial correspondence of multi-view instances which is hard to satisfy in many practical applications. To this end, this letter proposes a novel method termedDistribution-Level Multi-viewClustering forUnaligned Data (DLCU), which proves to be well-suited for scenarios where instance correspondences between different modalities are entirely absent. Specifically, in order to reconstruct cross-view correspondence, the wasserstein distance is employed to effectuate the alignment of multi-view data in distribution-level and guide the learning of cross-view transformation. Furthermore, the negative impact of the inevitable misalignment is mitigated through the global attention mechanism, which is designed to assign appropriate weights to realigned instances across multiple views. Besides, a divergence-based clustering objective is explored to encourage a clear cluster structure and conduct the training of latent representation. Experimental results on several real-world datasets show our promising performance comparing with the state-of-the-art methods.
Liang Zhao 0005, Qiongjie Xie
IEEE Signal Process. Lett.1
2024 Multi-Sentence Complementarily Generation for Text-to-Image Synthesis
abstract
Generating realistic images based on text descriptions remains challenging in computer vision. Existing multi-stage generation methods are sufficient to generate high-resolution images. However, these methods mainly use one sentence to synthesize images, which are difficult to extract adequate semantic features, resulting in the generated images being far apart from ground-truth images. In this paper, we propose a Multi-Sentence Complementary Generative Adversarial Network, MSCGAN, which assists in generating accurate images by fusing the same semantics from different sentences and preserving their unique semantics. More specifically, the latest BERT model is employed to identify semantic features and a multi-semantic fusion module (MSFM) is designed to fuse the semantic features of different sentences. Besides, a pre-trained cross-modal contrast similarity model (CCSM) is developed to explore fine-grained loss on generated images. Moreover, a multi-sentence joint discriminator is designed to ensure that the generated images match all sentences. Experiments and ablation studies on CUB and MS-COCO datasets demonstrate the significant superiority of the proposed method compared to state-of-the-art methods.
Liang Zhao 0005, Pingda Huang, Tengtuo Chen, Chunjiang Fu, Qinghao Hu 0003, Yangqianhui Zhang
IEEE Trans. Multim.1
2023 Predicting Chronic Obstructive Pulmonary Disease Based on Multi-Stage Composite Ensemble Learning Framework
abstract
Chronic Obstructive Pulmonary Disease (COPD) severely affects people’s health. With this in mind, we propose a novel Multi-Stage Composite Ensemble Learning Framework (MSCELF) that can diagnose COPD without utilising pulmonary function tests data. Our method explores 12 features from the patients’ baseline data, medical history, blood tests, and arterial blood gas analysis. In the first stage of our approach, three different ensemble learning methods are employed. The second stage involves the utilization of two machine learning methods. Finally, the Murphy’s method is integrated in the final stage to combine the outputs, with weights being assigned based on their information quantity and credibility. We evaluate our method on a clinical dataset of 329 patients and show that it outperforms existing methods in terms of accuracy, AUC, sensitivity, specificity, PPV, NPV, and F1 score, which are 0.7980, 0.8082, 0.8551, 0.6835, 0.8570, 0.6531, 0.8560.
Zhanxin Gang, Chaoran Jia, Chenhua Guo, Peng Li 0027, Jing Gao 0007, Liang Zhao 0005
BIBM6
2023 Enhancing Longitudinal Medical Image Segmentation through Spatial-temporal Fusion
abstract
Medical imaging research has seen advances in deep learning, but temporal aspects in time-series medical imaging data are often overlooked, leading to diagnostic limitations. This study proposes a spatial-temporal fusion approach for medical image segmentation by integrating a 3D UNet spatial network with a novel temporal network, DTransformer, capable of handling irregularly spaced sequences. The 3D UNet extracts spatial features, while DTransformer processes temporal information with time distance considerations using a novel self-attention mechanism. Experiments on a lung CT dataset show significant segmentation accuracy improvements with the fusion approach. DTransformer proves effective for unequally spaced sequences and boosts performance. And spatial-temporal fusion enhances medical image segmentation. Moveover, DTransformer's ability to manage temporal context and time distance holds promise for various tasks, indicating a new avenue for research.
Liang Zhao 0005, Chaoran Jia, Zhanxin Gang, Ruixin Ma
BIBM1
2023 Soft Tissue Sarcoma Segmentation Network Based on Ensemble Learning
abstract
Accurate segmentation of soft tissue sarcoma in medical images is crucial for effective diagnosis and treatment planning. This study aims to improve soft tissue sarcoma segmentation by proposing a two-layer ensemble model that leverages diverse neural network architectures and a confidence-based integration method. In the first layer, we employ prominent models like CaraNet, AttUNet(M), UNet(S), UNet(M), U Transformer, and Swin UNet for initial segmentation. The second layer utilizes a confidence-based ensemble approach to fuse the outputs of the first layer models, enhancing segmentation accuracy. Experimental evaluation on axial and coronal datasets reveals substantial enhancements in Dice coefficient, sensitivity, and specificity, highlighting the effectiveness of our ensemble strategy. Our approach outperforms individual base learners and ensures smoother and more accurate soft tissue sarcoma segmentation. The proposed two-layer ensemble model, integrating diverse neural network models and a confidence-based ensemble technique, offers improved soft tissue sarcoma segmentation results. This approach holds promise for enhancing diagnostic accuracy and aiding medical decision-making in sarcoma cases.
Liang Zhao 0005, Zhanxin Gang, Chenhua Guo, Jing Gao 0007
BIBM1
2023 Soft Tissue Sarcoma Segmentation Network Based on Self-supervised Learning
abstract
Soft tissues sarcomas include striated muscle, fibrous tissue, fat, and other soft tissues. Simultaneously, their mortality rates are comparable to those of esophageal cancer, cervical cancer, and other cancers. Prior to surgical resection of patients, it is frequently necessary to study and diagnose the sarcoma area using MRI images in order to design a better surgical plan. However, artificial approaches for diagnosing the sarcoma region are time-consuming and error-prone. While the advent of artificial intelligence allows for computer-assisted diagnosis of the sarcoma region. Nevertheless, there is currently a scarcity of high-quality soft tissue sarcoma imaging data sets in relevant sectors. Thus, with the aim to investigate how to use multi-modal MRI images of patients with soft tissue sarcomas to segment the sarcoma area, we collect and process 15372 multi-modal MRI images in coronal of 40 patients with soft tissue sarcomas found in the thigh, which we subsequently combine with the help of several clinicians to mark the sarcoma area. The multi-modal MRI imaging data set of soft tissue sarcoma is therefore acquired by a number of preprocessing techniques. The sarcoma area is then segmented using a multi-encoder and single-decoder network that adapts to multiple input modalities. For motivating the network to learn the important semantic features of different modalities, we design a feature fusion strategy mechanism that is applied to the skip connection. Additionally, self-supervised learning is being investigated to address the issue of a small number of data points in the data set. Experiments show that our network can achieve the highest Dice score of 57.76% on our data set. Our code and the dataset are available at https://github.com/syaxx0819/The-Multimodal-Soft-Tissue-Sarcoma-Image-Dataset.
Liang Zhao 0005, Zhanxin Gang, Chaoran Jia, Yi Yang 0006
BIBM1
2023 Multi-View Graph Regularized Deep Autoencoder-Like NMF Framework
abstract
Many real-world data are composed of different representations or views, thus multi-view clustering (MVC) has attracted more and more attention in recent years. Its key task is how to extract sufficient fusion features from multi-view data. Because the nonnegative matrix factorization (NMF) can favorably explain the extracted features, the NMF based MVC is usually a good choice for multi-view data, and promising results are achieved. Inspired by this, we propose a multi-view graph regularized deep autoencoder-like NMF (MGANMF) framework in this paper for multi-view clustering. MGANMF uses the deep autoencoder-like NMF, which draws lessons from the idea of depth automatic encoder, to learn the hierarchical semantics of multi-view data in a layer-wise manner. Moreover, in order to describe the inherent geometric structure in each view data, graph regulators are introduced to couple the output representation of deep structure. In addition, the self-updating weights are employed to balance the effect of each view. Thus, a new objective function is defined and the optimization processes are presented. Experimental results on several multi-view datasets show the effectiveness of the proposed model.
Liang Zhao 0005, Zhikui Chen
ICASSP1
2023 Unrestricted Anchor Graph Based GCN for Incomplete Multi-View Clustering
abstract
In recent years, the task of multi-view clustering(MVC) has attracted more and more attention. Meanwhile, the graph convolution network(GCN) based MVC method has made consistent achievements in processing graph-structured data. However, real world data often suffers from missing some instances in each view, leading to the problem of incomplete multi-view clustering. It’s a really challenge to capture the graph structure of incomplete views for GCN to process, especially in the high missing-rate situation. To address this is-sue, this paper proposes a novel and effective graph construct method called unrestricted anchor graph(UAG). Moreover, an Unrestricted Anchor Graph based GCN framework(UAGCN) is designed for incomplete multi-view clustering. Specifically, our method employs the unrestricted anchor to reconstruct the relationship in high missing-rate data to describe the graph structure, and then integrates GCN to obtain the graph embedding of incomplete data for clustering. The experimental results on multiple data sets show that our method is superior to comparison methods.
Liang Zhao 0005, Yukun Yuan 0002, Feng Ding 0016
ICASSP1
2023 An End-to-End Framework for Partial View-Aligned Clustering with Graph Structure
abstract
Over the last decade, many multi-view clustering (MVC) methods have achieved promising results with intact and completely correct correspondence of multi-view data, which is hard to satisfy in practice leading to the problem of partially view-aligned clustering. In this paper, we propose a novel method to tackle it, termed An End-to-end Framework for Partial View-aligned Clustering with Graph structure(EGPVC). It employs Dykstra’s cyclic constraint projection algorithm to obtain the correspondence between two views. In particular, EGPVC develops an end-to-end framework for partially view-aligned clustering, in which representation learning and clustering process can benefit from each other through the deep embedded clustering layer. Moreover, a cross-view graph regularization term is designed to improve the quality of the learned common representation with graph structure information. Experimental results on several real-world datasets show our promising results comparing with the state-of-the-art methods in partially view-aligned clustering.
Liang Zhao 0005, Qiongjie Xie, Sontao Wu, Shubin Ma
ICASSP1
2023 Enhancing Path Information with Reinforcement Learning for Few-shot Knowledge Graph Completion
abstract
The emergence of big data has made knowledge graphs (KGs) an effective means of representing structured knowledge, and few-shot knowledge graph completion (FKGC) has recently received increasing attention, which attempts to forecast missing information for relations with few-shot related facts. In this regard, several deep learning-based and embedding-based methods have been proposed for FKGC. However, most existing methods overlook multi-hop path information and only utilize the immediate neighbors of relevant entities when encoding and matching entity pairs, potentially limiting their performance. In this paper, we propose an Enhancing Path Information with Reinforcement Learning (EPIRL) approach for FKGC. Specifically, we introduce a reinforcement learning framework to construct a reasoning subgraph, aiming to thoroughly uncover the inferential path rules between support and query triples. Then, we utilize an interaction focused matching model to capture the inherent connections among these reasoning paths. To further improve performance, we incorporate a relational attention mechanism aimed at emphasizing the influence of pivotal paths. Extensive experiments demonstrate that our model outperforms several state-of-the-art methods on the frequently-used benchmark datasets FB15k237-One and NELL-One.
Ruixin Ma, Mengfei Yu, Buyun Gao, Zhikui Chen, Liang Zhao 0005
ICPADS6
2023 MVCIR-net: Multi-view Clustering Information Reinforcement Network
abstract
Multi-view clustering (MVC) integrates information from different views to improve clustering performance compared to single-view clustering. However, the raw multi-view data in the feature space often contains irrelevant information to the clustering task, which is difficult to separate using existing methods. This irrelevant information is processed equally with clustering information, negatively impacting the final clustering performance. In this paper, we propose a new framework for multi-view clustering information reinforcement network (MVCIR-net) to alleviate these problems. Our method gives practical clustering meaning to the clustering distribution layer by contrastive learning. Then, the trusted neighbor instances distribution of the normalized graph is debias aggregated to form the clustering information propensity distribution, and the clustering information distribution is made to fit this distribution. In addition, the coupling degree of the clustering information distribution in different views on the same sample should be enhanced. Through the aforementioned strategies, the raw data is fuzzy mapped into clustering information, and the network's ability to recognize clustering information is strengthened. Finally, the fuzzy mapping data is input into the network and reconstructed to evaluate the quality of the extracted clustering information. Extensive experiments on public multi-view datasets show that MVCIR-net achieves superior clustering effectiveness and the ability to identify clustering information.
Shaokui Gu, Xu Yuan 0002, Liang Zhao 0005, Zhenjiao Liu, Yan Hu 0007, Zhikui Chen
ACM Multimedia3
2023 One-shot relational learning for extrapolation reasoning on temporal knowledge graphs
Ruixin Ma, Biao Mei, Meihong Liu, Liang Zhao 0005
Data Min. Knowl. Discov.6
2023 IMC-NLT: Incomplete multi-view clustering by NMF and low-rank tensor
Zhenjiao Liu, Zhikui Chen, Yue Li 0050, Liang Zhao 0005, Reza Farahbakhsh, Noël Crespi, Xiaodi Huang 0001
Expert Syst. Appl.4
2023 PANC: Prototype Augmented Neighbor Constraint instance completion in knowledge graphs
Ruixin Ma, Biao Mei, Guangyue Lv, Liang Zhao 0005
Expert Syst. Appl.6
2023 Deep probability multi-view feature learning for data clustering
Liang Zhao 0005, Zhenjiao Liu
Expert Syst. Appl.1
2023 Incomplete Multi-View Clustering With Complete View Guidance
abstract
In recent years, multi-view clustering has gained widespread attention in signal processing because multi-view data contains more information than a single view. However, multi-view data is often incomplete due to missing data in one or more random views. Therefore, several methods have been proposed for incomplete multi-view clustering to learn features that contain consensus information for clustering incomplete multi-view data (IMD). However, there is a part of the IMD that is not missing in any view, and most previous methods have not utilized this part to guide the process of learning consensus information. To address this issue, we design a knowledge distillation framework for incomplete multi-view clustering and propose an incomplete multi-view clustering with complete view guidance (IMC-CVG). We first train a robust teacher model with contrastive learning loss on the complete part of IMD to learn consensus features containing multi-view information. Then, we train a student model on all the IMD, where we mask partial views of the complete data to simulate missing data, and utilize the teacher model to guide the student model to learn consensus features that contain as much multi-view information as possible. Experiments show that our proposed method outperforms all the compared state-of-the-art methods.
Zhikui Chen, Yue Li 0050, Kai Lou, Liang Zhao 0005
IEEE Signal Process. Lett.4
2023 Mining Multi-View Clustering Space With Interpretable Space Search Constraint
abstract
Multi-view clustering can cluster signal samples from multiple views into groups. Currently, multi-view clustering fuses the information of different views into a low-dimensional space for clustering. However, the direct reduction of high-dimensional information to a very low-dimensional space leads to the loss of a lot of sample semantic information, while a higher dimension after dimensionality reduction may blur the clustering structure of samples. To tackle these problems, we propose a novel framework called Mining Multi-view Clustering Space with Interpretable Space Search Constraint to explore the clustering structure while preserving the low-dimensional space semantic information. Our method maps samples from the raw space to a low-dimensional space separated into consensus and private features. This allows us to explore the interpretability of samples in the low-dimensional space to achieve representative representations. After assembling these representations, guided fusion is carried out and a search constraint is imposed to achieve a more reasonable clustering structure. Finally, by dynamically screening positive and negative samples, the clustering performance of the clustering space is maximized by contrastive learning. Extensive experiments on public datasets demonstrate that our method achieves state-of-the-art clustering effectiveness.
Xu Yuan 0002, Shaokui Gu, Zhenjiao Liu, Liang Zhao 0005
IEEE Signal Process. Lett.4
2022 Predicting The Likelihood of Patients Developing Sepsis Based on Compound Ensemble Learning
abstract
Objective: Based on a dataset with few features predicts a patient’s likelihood of developing sepsis in the future. Method: We propose a compound ensemble learning framework. At the first stage, UMAP and XgBoost are explored for dimensionality reduction, meanwhile, weak classifiers such as KELM, 1D Convolutional Neural Networks, and XgBoost are employed as base learners to learn patient characteristics from four different emphases. The output of base learners is processed using the proposed Res-Stacking and modified Bagging-like methods in the first stage of compound ensemble learning framework to reduce the error rate. In the secondary ensemble learning stage, the D-S evidence theory is employed to fuse the results of the above two methods and BiLSTM, thereby diminishing the uncertainty of the final output and obtaining more prominent prediction result. Result: The performance of our proposed compound ensemble framework on the clinical dataset of the First Affiliated Hospital of Dalian Medical University in terms of sensitivity, specificity, diagnostic efficiency, PPV, NPV, F1 score and AUC are 0.875, 0.95, 0.928, 0.95, 0.875, 0.95, 0.9125, respectively, better than the single ensemble learning method.
Liang Zhao 0005
BIBM1
2022 Time-series lung cancer CT dataset
abstract
In order to better explore the evolution process of lung nodules in lung cancer patients, we collect lung CT data at multiple time points of lung cancer patients, track and mark the CT positions of the same lung nodules in lung cancer patients at different time points, and make time-series CT data sets of lung cancer patients. After that, 3D-UNet model is used to detect lung nodules on our data set. Experiment proves the effectiveness and availability of the data set, and also proved that the image data at multiple time points could improve the accuracy of the model’s identification of lung nodules.
Liang Zhao 0005, Chaoran Jia
BIBM1
2022 Multi-label Aerial Image Classification Based on Image-Specific Concept Graphs
abstract
Multi-label aerial image classification (MAIC) is a fundamental but challenging task for computer vision-based remote sensing applications. Existing MAIC models suffer from the insufficient semantic information of image and label representations. To this end, we integrate commonsense knowledge into the MAIC task and propose a novel Knowledge-augmented Concept Graph Learning (KCGL) framework. KCGL first collects relevant semantic concepts for each label from a commonsense knowledge graph ConceptNet. With the guidance of semantic concepts, an image decoupling module is employed to extract concept-specific image features from the input image. Then, KCGL constructs an individual concept graph for each image, in which nodes are corresponding to concept-specific image features and edges are their relations extracted from ConceptNet. Finally, the classification probability on each label is computed in the specific concept graph via a GCN-based encoder-decoder model. Experimental results prove that the proposed KCGL outperforms existing state-of-the-art MAIC models on two aerial image datasets.
Dan Lin 0008, Zhikui Chen, Liang Zhao 0005, Kai Wang 0057
ICIP3
2022 Multi-attention User Information Based Graph Convolutional Networks for Explainable Recommendation
Ruixin Ma, Guangyue Lv, Liang Zhao 0005
KSEM (1)3
2022 Incomplete multi-view clustering based on weighted sparse and low rank representation
Liang Zhao 0005, Jie Zhang 0085, Zhikui Chen
Appl. Intell.1
2022 LSTM-MFCN: A time series classifier based on multi-scale spatial-temporal features
Liang Zhao 0005, Chunyang Mo, Zhikui Chen, Chenhui Yao
Comput. Commun.1
2022 Knowledge Graph Random Neural Networks for Recommender Systems
Ruixin Ma, Fangqing Guo, Liang Zhao 0005
Expert Syst. Appl.4
2022 Multilabel Aerial Image Classification With a Concept Attention Graph Neural Network
abstract
Compared with natural images, aerial images collected by satellite sensors/aerial cameras can provide a much larger field of view and often contain multiple objects of interest (multiple labels). There are certain limitations of existing multilabel aerial image classification methods. First, label correlations were often ignored in previous MAIC work, and thus, multilabel classifiers failed to be self-adapted. Second, existing multilabeled data sets for aerial images only cover limited images with fixed labels. Therefore, the underlying semantic correlations of labels cannot be fully included, while such correlation information is implicitly used as common knowledge by human beings. To tackle these concerns, we propose a novel multilabel classification method for aerial images. Our contributions are twofold. First, as the first attempt, label correlations are inferred from both the specific data set and ConceptNet (a popular knowledge graph for common sense). Second, based on graph neural network (GNN), we propose a novel end-to-end aerial image classification model, named the multiple label concept graph (ML-CG). ML-CG builds a concept graph to describe the semantic correlations from both the label set and the ConceptNet. We also incorporate both semantic attention and label attention in the GNN to better extract meaningful information of image labels. Compared with state-of-the-art methods, the effectiveness of the proposed method is demonstrated on both the commonly used UCM data set and a recently proposed DFC15 data set with high image resolution.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.3
2022 Multilabel Aerial Image Classification With Unsupervised Domain Adaptation
abstract
Deep learning (DL) methods are promising for the multilabel aerial image classification (MAIC) task. However, current DL methods face a common problem: the need for large multilabeled datasets. Collecting and annotating raw aerial image datasets can be extremely time- and labor-consuming. To address this concern in MAIC, domain adaptation (DA) provides a novel solution by transferring the knowledge learned from a label-rich dataset (i.e., the source domain) to a label-scarce dataset (i.e., the target domain), while current DA models are mainly designed for single-labeled tasks. In this article, we propose a novel end-to-end MAIC model based on DA techniques, named DA-MAIC. To the best of our knowledge, this article for the first time integrates DA to tackle the label scarcity problem in the MAIC task. Specifically, the proposed DA-MAIC is composed of two main parts: the image classifier and the domain classifier. The image classifier captures task-discriminative features based on the graph convolutional network (GCN) to predict multiple image labels; and the domain classifier extracts domain-invariant representations, which mitigates the domain shift between two underlying distributions. We extensively evaluate the proposed DA-MAIC from different perspectives on three benchmark datasets, including the commonly used UCM dataset, the high-resolution AID dataset, and the recently proposed DFC15 dataset. Both quantitative and qualitative results support that the proposed DA-MAIC can generalize the source domain knowledge to new scenarios and substantially improve the classification performance on the target domain task.
Dan Lin 0008, Jianzhe Lin, Liang Zhao 0005, Z. Jane Wang 0001, Zhikui Chen
IEEE Trans. Geosci. Remote. Sens.3
2021 TRGAN: Text to Image Generation Through Optimizing Initial Image
Liang Zhao 0005, Pingda Huang, Zhikui Chen, Yanqi Dai
ICONIP (5)1
2021 Hybrid attention mechanism for few-shot relational learning of knowledge graphs
abstract
Abstract Few‐shot knowledge graph (KG) reasoning is the main focus in the field of knowledge graph reasoning. In order to expand the application fields of the knowledge graph, a large number of studies are based on a large number of training samples. However, we have learnt that there are actually many missing relationships or entities in the knowledge graph, and in most cases, there are not many training instances when implementing new relationships. To tackle it, in this study, the authors aim to predict a new entity given few reference instances, even only one training instance. A few‐shot learning framework based on a hybrid attention mechanism is proposed. The framework employs traditional embedding models to extract knowledge, and uses an attenuated attention network and a self‐attention mechanism to obtain the hidden attributes of entities. Thus, it can learn a matching metric by considering both the learnt embeddings and one‐hop graph structures. The experimental results present that the model has achieved significant performance improvements on the NELL‐One and Wiki‐One datasets.
Ruixin Ma, Fangqing Guo, Liang Zhao 0005
IET Comput. Vis.4
2021 Incremental multi-view correlated feature learning based on non-negative matrix factorisation
abstract
Abstract In real‐world applications, large amounts of data from multiple sources come in the form of streams. This makes multi‐view feature learning cost much time when new instances rise incrementally. Dealing with these growing multi‐view data becomes a challenging problem. Some single‐view methods focus on processing the data dynamically, but they are not suitable for multi‐view data. Some online multi‐view methods are proposed to tackle it, but they ignore the influence of uncorrelated items in each view. Therefore, in this study, the authors propose a new algorithm, called Incremental Multi‐view Correlated Feature Learning (IMCFL) based on non‐negative matrix factorisation, to learn the common feature across views. By separating uncorrelated items of new instances and constructing incremental joint learning of correlated and uncorrelated features, the proposed IMCFL can eliminate the influence of uncorrelated information in the individual view and improve the effectiveness of incremental multi‐view common feature learning. Extensive experiments on real‐world datasets confirm its superiority by comparing it with other state‐of‐the‐art incremental and non‐incremental methods.
Liang Zhao 0005, Jie Zhang 0085, Zhikui Chen
IET Comput. Vis.1
2021 Dual Alignment Self-Supervised Incomplete Multi-View Subspace Clustering Network
abstract
Incomplete multi-view clustering has attracted much attention in decade years. To date, most of the remarkable achievements, however, exploit shallow models to learn shared feature representations based on incomplete views. Although some deep learning methods have been proposed to solve this issue, the existing ones still have the following problems: 1) The consistency between views is ignored, which will have serious negative impacts on incomplete multi-view learning. 2) The learned features do not have sufficient cluster-friendliness, that is, the tightness within clusters and the repulsiveness between clusters are not fully considered. To tackle the above shortcomings, we propose a Dual Alignment Self-supervised Incomplete Multi-view Subspace Clustering network (DASIMSC) in this paper. Specifically, the manifold alignment constraint and consistency alignment constraint are integrated with the autoencoder to preserve the compact inherent local structure within the view and the consistency semantics between incomplete views, respectively. Moreover, a self-expression layer coupled with a spectral clustering module is designed to naturally separate different types of data, leveraging the current clustering results to supervise subspace learning, which excludes inter-cluster. Experimental results on several datasets show that our algorithm outperforms all compared state-of-the-arts.
Liang Zhao 0005, Jie Zhang 0085, Qiuhao Wang, Zhikui Chen
IEEE Signal Process. Lett.1
2021 Co-Learning Non-Negative Correlated and Uncorrelated Features for Multi-View Data
abstract
Multi-view data can represent objects from different perspectives and thus provide complementary information for data analysis. A topic of great importance in multi-view learning is to locate a low-dimensional latent subspace, where common semantic features are shared by multiple data sets. However, most existing methods ignore uncorrelated items (i.e., view-specific features) and may cause semantic bias during the process of common feature learning. In this article, we propose a non-negative correlated and uncorrelated feature co-learning (CoUFC) method to address this concern. More specifically, view-specific (uncorrelated) features are identified for each view when learning the common (correlated) feature across views in the latent semantic subspace. By eliminating the effects of uncorrelated information, useful inter-view feature correlations can be captured. We design a new objective function in CoUFC and derive an optimization approach to solve the objective with the analysis on its convergence. Experiments on real-world sensor, image, and text data sets demonstrate that the proposed method outperforms the state-of-the-art multiview learning methods.
Liang Zhao 0005, Jie Zhang 0085, Zhikui Chen, Yi Yang 0006, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 DT-LET: Deep transfer learning by exploring where to transfer
Jianzhe Lin, Liang Zhao 0005, Qi Wang 0009, Rabab K. Ward, Z. Jane Wang 0001
Neurocomputing2
2020 Multi-View Robust Feature Learning for Data Clustering
abstract
Multi-view feature learning can provide basic information for consistent grouping, and is very common in practical applications, such as judicial document clustering. However, it is a challenge to combine multiple heterogeneous features to learn a comprehensive description of data samples. To solve this problem, many methods explore the correlation between various features across views by assuming that all views share the same semantic information. Inspired by this, in this paper we propose a new multi-view robust feature learning (MRFL) method. In addition to projecting features from different views to a shared semantic subspace, our approach also learns the irrelevant information of data space to capture the feature dependencies between views in potential common subspaces. Therefore, the MRFL can obtain flexible feature associations hidden in multi-view data. A new objective function is designed to derive, and solve the effective optimization process of MRFL. Experiments on real-world multi-view datasets show that the proposed MRFL method is superior to the state-of-the-art multi-view learning methods.
Liang Zhao 0005, Zhikui Chen
IEEE Signal Process. Lett.1
2019 Unsupervised multi-view non-negative for law data feature learning with dual graph-regularization in smart Internet of Things
Xiru Qiu, Zhikui Chen, Liang Zhao 0005, Chengsheng Hu
Future Gener. Comput. Syst.3
2019 ICFS Clustering With Multiple Representatives for Large Data
abstract
With the prevailing development of Cyber-physical-social systems and Internet of Things, large-scale data have been collected consistently. Mining large data effectively and efficiently becomes increasingly important to promote the development and improve the service quality of these applications. Clustering, a popular data mining technique, aims to identify underlying patterns hidden in the data. Most clustering methods assume the static data, thus they are unfavorable for analyzing large, unbalanced dynamic data. In this paper, to address this concern, we focus on incremental clustering by extending the novel [clustering by fast search (CFS) and find of density peaks] method to incrementally handle large-scale dynamic data. Specifically, we first discuss two challenges, i.e., assignment of new arriving objects and dynamic adjustment of clusters, in incremental CFS (ICFS) clustering. We then propose two ICFS clustering algorithms, ICFS with multiple representatives (ICFSMR) and the enhanced ICFSMR (E_ICFSMR) to tackle the two challenges. In ICFSMR, we explore the convex hull theory to modify the representatives identified for each cluster. E_ICFSMR improves the generality and effectiveness of ICFSMR by exploring one-time cluster adjustment strategy after integration of each data chunk. We evaluate the proposed methods with extensive experiments on four benchmark data sets, as well as the air quality and traffic monitoring time series, with comparisons to CFS and other three state-of-the-art incremental clustering methods. Experimental results demonstrate that the proposed methods outperform the compared methods in terms of both effectiveness and efficiency.
Liang Zhao 0005, Zhikui Chen, Yi Yang 0006, Liang Zou, Z. Jane Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2019 Deep Semantic Mapping for Heterogeneous Multimedia Transfer Learning Using Co-Occurrence Data
abstract
Transfer learning, which focuses on finding a favorable representation for instances of different domains based on auxiliary data, can mitigate the divergence between domains through knowledge transfer. Recently, increasing efforts on transfer learning have employeddeepneuralnetworks (DNN) to learn more robust and higher level feature representations to better tackle cross-media disparities. However, only a few articles consider the correction and semantic matching between multi-layer heterogeneous domain networks. In this article, we propose adeep semantic mapping model forheterogeneous multimediatransferlearning (DHTL) using co-occurrence data. More specifically, we integrate the DNN withcanonicalcorrelationanalysis (CCA) to derive a deep correlation subspace as the joint semantic representation for associating data across different domains. In the proposed DHTL, a multi-layer correlation matching network across domains is constructed, in which the CCA is combined to bridge each pair of domain-specific hidden layers. To train the network, a joint objective function is defined and the optimization processes are presented. When the deep semantic representation is achieved, the shared features of the source domain are transferred for task learning in the target domain. Extensive experiments for three multimedia recognition applications demonstrate that the proposed DHTL can effectively find deep semantic representations for heterogeneous domains, and it is superior to the several existing state-of-the-art methods for deep transfer learning.
Liang Zhao 0005, Zhikui Chen, Laurence T. Yang, M. Jamal Deen, Z. Jane Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Incomplete multi-view clustering via deep semantic mapping
Liang Zhao 0005, Zhikui Chen, Yi Yang 0006, Z. Jane Wang 0001, Victor C. M. Leung
Neurocomputing1
2018 Unsupervised Multiview Nonnegative Correlated Feature Learning for Data Clustering
abstract
Multiview data, which provide complementary information for consensus grouping, are very common in real-world applications. However, synthesizing multiple heterogeneous features to learn a comprehensive description of the data samples is challenging. To tackle this problem, many methods explore the correlations among various features across different views by the assumption that all views share the common semantic information. Following this line, in this letter, we propose a new unsupervised multiview nonnegative correlated feature learning (UMCFL) method for data clustering. Different from the existing methods that only focus on projecting features from different views to a shared semantic subspace, our method learns view-specific features and captures inter-view feature correlations in the latent common subspace simultaneously. By separating the view-specific features from the shared feature representation, the effect of the individual information of each view can be removed. Thus, UMCFL can capture flexible feature correlations hidden in multiview data. A new objective function is designed and efficient optimization processes are derived to solve the proposed UMCFL. Extensive experiments on real-world multiview datasets demonstrate that the proposed UMCFL method is superior to the state-of-the-art multiview clustering methods.
Liang Zhao 0005, Zhikui Chen, Z. Jane Wang 0001
IEEE Signal Process. Lett.1
2018 Distributed Feature Selection for Efficient Economic Big Data Analysis
abstract
With the rapidly increasing popularity of economic activities, a large amount of economic data is being collected. Although such data offers super opportunities for economic analysis, its low-quality, high-dimensionality and huge-volume pose great challenges on efficient analysis of economic big data. The existing methods have primarily analyzed economic data from the perspective of econometrics, which involves limited indicators and demands prior knowledge of economists. When embracing large varieties of economic factors, these methods tend to yield unsatisfactory performance. To address the challenges, this paper presents a new framework for efficient analysis of high-dimensional economic big data based on innovative distributed feature selection. Specifically, the framework combines the methods of economic feature selection and econometric model construction to reveal the hidden patterns for economic development. The functionality rests on three pillars: (i) novel data pre-processing techniques to prepare high-quality economic data, (ii) an innovative distributed feature identification solution to locate important and representative economic indicators from multidimensional data sets, and (iii) new econometric models to capture the hidden patterns for economic development. The experimental results on the economic data collected in Dalian, China, demonstrate that our proposed framework and methods have superior performance in analyzing enormous economic data.
Liang Zhao 0005, Zhikui Chen, Yueming Hu 0001, Geyong Min, Zhaohua Jiang
IEEE Trans. Big Data1
2017 A privacy-preserving high-order neuro-fuzzy c-means algorithm with cloud computing
Peng Li 0027, Zhikui Chen, Laurence T. Yang, Liang Zhao 0005, Qingchen Zhang 0001
Neurocomputing4
2017 An Incremental CFS Algorithm for Clustering Large Data in Industrial Internet of Things
abstract
With the rapid advances of sensing technologies and wireless communications, large amounts of dynamic data pertaining to industrial production are being collected from many sensor nodes deployed in the industrial Internet of Things. Analyzing those data effectively can help to improve the industrial services and mitigate the system unprepared breakdowns. As an important technique of data analysis, clustering attempts to find the underlying pattern structures embedded in unlabeled information. Unfortunately, most of the current clustering techniques that could only deal with static data become infeasible to cluster a significant volume of data in the dynamic industrial applications. To tackle this problem, an incremental clustering algorithm by fast finding and searching of density peaks based on k-mediods is proposed in this paper. In the proposed algorithm, two cluster operations, namely cluster creating and cluster merging, are defined to integrate the current pattern into the previous one for the final clustering result, and k-mediods is employed to modify the clustering centers according to the new arriving objects. Finally, experiments are conducted to validate the proposed scheme on three popular UCI datasets and two real datasets collected from industrial Internet of Things in terms of clustering accuracy and computational time.
Qingchen Zhang 0001, Chunsheng Zhu, Laurence T. Yang, Zhikui Chen, Liang Zhao 0005, Peng Li 0027
IEEE Trans. Ind. Informatics5