Meiyu Liang

dblp:150/1513 · also MeiYu Liang · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph Completion
abstract
Multi-modal knowledge graph completion (MMKGC) aims to infer missing entities of triples by leveraging heterogeneous information in knowledge graph (KG). However, existing approaches often struggle with inconsistent modality alignment, limited reasoning depth, and insufficient negative sample quality. In this work, we propose HFR-MKGC, a novel framework that integrates hierarchical modal fusion and Multimodal Large Language Model (MLLM) reasoning for robust and expressive MMKGC. Specifically, we introduce a relation-guided hierarchical modal fusion module, which conducts fine-grained intra-visual fusion and relation-guided cross-modal integration to yield rich entity representations. HFR-MKGC employs a fine-tuned MLLM to perform instruction-based triple reasoning, producing candidate entities for completion. Then, it constructs hard negative samples through textual perturbation by MLLM and visual feature augmentation with rotation and noise. HFR-MKGC optimizes the model via adversarial training. Extensive experiments on three MMKGC benchmarks demonstrate that our method outperforms state-of-the-art methods, validating its effectiveness in MMKGC.
Junping Du 0001, Zhe Xue, Meiyu Liang, Guanhua Ye, Yingxia Shao, Hai-Sheng Li 0002
AAAI4
2026 Rethink Representation Learning for Questionnaire Data
abstract
Questionnaire data serve as a valuable resource across numerous scientific domains, offering insights into human behavior, health, and social trends. Traditional downsampling-based representation learning methods—such as standardization and one-hot encoding—reformat these data into tabular structures that inherently discard semantic richness and obscure inter-sample and inter-feature relationships. Consequently, advanced deep learning models often underperform compared to simpler approaches like gradient-boosted decision trees (GBDT), due to their limited capacity to extract meaningful representations from semantically sparse inputs. To address this limitation, we introduce SemantiQ, a novel upsampling-based representation learning framework that embeds questionnaire responses into a unified semantic space. Leveraging Retrieval-Augmented Generation (RAG) in conjunction with large language models (LLMs), SemantiQ transforms question text, option text, and external knowledge into semantically enriched natural language statements. These statements are then encoded into semantic embeddings, which are further refined through a three-stage training mechanism and test-time training (TTT), enabling the model to capture complex sample- and feature-wise dependencies. Extensive experiments on multiple real-world datasets demonstrate that SemantiQ consistently outperforms state-of-the-art baselines.
Guanhua Ye, Jifeng He 0003, Junping Du 0001, Zhe Xue, Yingxia Shao, Meiyu Liang, Yawen Li 0001
AAAI7
2026 ST-VLM: A Spatial-to-Image Multimodal Spatial-Temporal Prediction Framework with Vision-Language Model
abstract
Spatial-temporal prediction plays a crucial role in various domains, including intelligent transportation and environmental monitoring. Although large language model has shown advantages in long-range dependency modeling and excellent generalization ability for forecasting, it has limited understanding of spatial-temporal features. Especially for spatial features, most existing methods still simplify the spatial-temporal prediction task into multiple independent temporal prediction tasks, failing to effectively encode the dynamic evolution of spatial relations. To address these problems, we propose ST-VLM (Spatial-Temporal Forecasting with Vision-Language Model), a novel framework that leverages visual representations to encode the dynamic spatial dependencies within spatial-temporal data and integrates multi-modal information to enhance prediction. This framework transforms spatial-temporal features into three modalities: vision, text, and time series, enhances cross-modal fusion through an attention-aware fusion mechanism in the first-layer of Vision-Language Model (VLM), optimizes multi-modal feature interaction via adaptive fine-tuning strategies. After fusion, the multi-modal embeddings are subsequently used for the final spatial-temporal prediction task. Extensive experiments demonstrate that ST-VLM achieves state-of-the-art performance across various datasets. In particular, the framework exhibits promising results in few-shot scenarios, verifying its strong generalization ability.
Junping Du 0001, Zhe Xue, Meiyu Liang, Aijing Li, Xiaolong Meng
AAAI4
2026 Multi-Granularity Multi-Modal Knowledge Graph Representation Learning via Subgraph-Aware Adaptive Fusion and Hierarchical Relation Modeling
Peining Li, Meiyu Liang, Junping Du 0001, Zhe Xue, Guanhua Ye, Wu Liu 0005, Lei Shi 0030
WWW2
2026 Addressing graph heterogeneity and heterophily from a spectral perspective
Kangkang Lu 0002, Yanhua Yu, Ruopei Guo, Zhiyong Huang 0010, Yunshan Ma 0002, Meiyu Liang, Xiting Qin, Yimeng Ren 0001, Tat-Seng Chua
Neurocomputing7
2025 Incomplete Multi-View Multi-Label Classification via Diffusion-Guided Redundancy Removal
abstract
Incomplete multi-view multi-label classification aims to accurately predict labels for each sample in the face of some missing views. Due to its widespread presence in real-world scenarios, it has become an extensively researched topic. In addition to the challenges brought by missing views, it also encounters issues caused by redundant views, whose inclusion fails to make a positive contribution to performance. In this paper, we make the first attempt to take advantage of diffusion models to address the missing view problem and design a strategy to identify and remove redundant views. Specifically, we train a diffusion model conditioned on the pseudo-labels to recover information of missing views. The learned diffusion model can carry data distribution knowledge in training split to the data. Regarding redundant identification strategy, it is designed by considering both the additional information of views and the classification difficulty level of samples, thereby adaptively identifying and removing redundant views. We conduct extensive experiments on five datasets, and the proposed method achieves favorable performance against several state-of-the-art methods on the multi-view multi-label classification task.
Shilong Ou, Zhe Xue, Lixiong Qin, Yawen Li 0001, Meiyu Liang, Junjiang Wu, Xuyun Zhang, Amin Beheshti, Yuankai Qi
AAAI5
2025 Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View Clustering
abstract
Deep multi-view clustering methods utilize information from multiple views to achieve enhanced clustering results and have gained increasing popularity in recent years. Most existing methods typically focus on either inter-view or intra-view relationships, aiming to align information across views or analyze structural patterns within individual views. However, they often incorporate inter-view complementary information in a simplistic manner, while overlooking the complex, high-order relationships within multi-view data and the interactions among samples, resulting in an incomplete utilization of the rich information available. Instead, we propose a multi-scale approach that exploits all of the available information. We first introduce a dual graph diffusion module guided by a consensus graph. This module leverages inter-view information to enhance the representation of both nodes and edges within each view. Secondly, we propose a novel contrastive loss function based on hypergraphs to more effectively model and leverage complex intra-view data relationships. Finally, we propose to adaptively learn fusion weights at the sample level, which enables a more flexible and dynamic aggregation of multi-view information. Extensive experiments on eight datasets show favorable performance of the proposed method compared to state-of-the-art approaches, demonstrating its effectiveness across diverse scenarios.
Liang Chen 0030, Zhe Xue, Yawen Li 0001, Meiyu Liang, Yan Wang 0002, Anton van den Hengel, Yuankai Qi
CVPR4
2025 MFLCP: Personalized Multimodal Federated Learning via Collaborative Prompting with Missing Modalities
abstract
Multimodal Federated learning (FL) is a collaborative and privacy preserving machine learning paradigm for multimodal data. With the impressive performance of large-scale pre-trained models, an increasing number of these models are being applied to FL. However, multimodal data in the real world is usually incomplete in modalities. Additionally, directly applying these large-scale pre-trained models in the federated learning framework will lead to the problem of high computational and communication costs. To address these problems, we propose a novel Personalized Multimodal Federated Learning method via Collaborative Prompting with Missing Modalities (MFLCP) . Specifically, we propose an efficient large-scale pre-trained personalized multimodal federated learning framework. To address the issue of incomplete modalities in multimodal data, we propose a modal projection-aware collaborative prompting strategy for incomplete multimodal federated learning. Different categories of prompts are designed for the missing categories, and the modality mapping part and leverage complementary semantic information from different modalities are designed to guide prompt learning, promoting better interaction between modalities. In addition, we propose a communication optimization method for efficient multimodal federated learning, which reduces the parameters of multimodal pre-trained models in the process of federated communication transmission, enhances the speed of local training, and significantly improves convergence speed by integrating large-scale pre-trained models in a lightweight manner. Meanwhile, we establish a personalized adaptive update mechanism for the federated local model, which can adaptively update the local model according to the characteristics of local data, effectively reduces the impact of data heterogeneity. Extensive experimental results on several benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art baselines.
Meiyu Liang, Ruoyu Fan
ICMR2
2025 Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Retrieval
abstract
In recent years, pre-trained multimodal large models have attracted widespread attention due to their outstanding performance in various multimodal applications. Nonetheless, the extensive computational resources and vast datasets required for their training present significant hurdles for deployment in environments with limited computational resources. Many existing methods attempt to compress pre-trained multimodal large models through knowledge distillation, typically focusing on a single optimization objective. While such methods successfully reduce model parameters, they often incur significant performance degradation. Moreover, single-scale optimization fails to ensure comprehensive learning of the teacher model's knowledge across different aspects. In this work, we propose, for the first time, a dynamic self-adaptive multiscale distillation (DSMD) from pre-trained multi-modal large model for efficient cross-modal retrieval method, considering multiple scales from the perspectives of fine granularity, global structure, and hard negative sample mining. Furthermore, we design a dynamic loss balancer, eliminating the need to manually tune objective weights during distillation. This dynamic mechanism ensures that all objectives are optimized in a balanced and adaptive manner throughout the training process. Experiments demonstrate that our multiscale distillation framework achieves significant performance improvements over traditional single-scale distillation methods. Additionally, our proposed dynamic balancer effectively stabilizes the distillation process, ensuring consistent optimization across objectives. The distilled student model achieves 90% of the teacher model's performance while using only 10% of its parameters. Notably, our model also achieves state-of-the-art performance on cross-modal retrieval tasks, outperforming existing approaches. Codes are available at https://github.com/chrisx599/DSMD.
Zhengyang Liang, Meiyu Liang, Yawen Li 0001, Wu Liu 0005, Yingxia Shao, Kangkang Lu 0002
ACM Multimedia2
2025 Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model for Incomplete Multimodal Retrieval
abstract
Multimodal hashing stands as an efficient approach for multimodal retrieval, yet it frequently grapples with the challenge of misaligned representation spaces across different modalities. This misalignment can degrade the consistency and discrimination of multimodal representations, complicating the learning of effective representations for image and text pairs. Particularly, the task becomes arduous when the system must handle incomplete data while ensuring accurate and relevant retrieval outcomes. To address these challenges, we propose the Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model (AADH) for Incomplete Multimodal Retrieval. Our model is specifically tailored to robustly manage multimodal incomplete data scenarios. Initially, we develop an Asymmetric Pre-alignment Strategy that utilizes asymmetric contrastive learning to preliminarily align the semantic disparities between various modalities. Subsequently, we propose an innovative Anchor Contrastive Reinforcement Diffusion Hashing Model, which integrates image and text modalities to varying extents during the reverse diffusion process. It constructs an anchor space that not only facilitates the learning of incomplete multimodal hashing representations through anchor contrastive learning but also leverages inter-modal and intra-modal contrastive learning to enhance the representations. Moreover, we effectively bridge the modality gap between different modal hash codes by employing the anchor space to constrain the representations of different modal hashes. By adjusting the initial noise of the diffusion model, we indirectly expand the data volume, which in turn bolsters the model's robustness. Our extensive experimental results across multiple datasets demonstrate that the proposed AADH model achieves state-of-the-art (SOTA) results.
Meiyu Liang, Juncheng Zheng, Kangkang Lu 0002, Yawen Li 0001, Junping Du 0001, Zhe Xue, Wu Liu 0005
ACM Multimedia2
2025 Dynamic Masking and Auxiliary Hash Learning for Enhanced Cross-Modal Retrieval
abstract
The demand for multimodal data processing drives the development of information technology. Cross-modal hash retrieval has attracted much attention because it can overcome modal differences and achieve efficient retrieval, and has shown great application potential in many practical scenarios. Existing cross-modal hashing methods have difficulties in fully capturing the semantic information of different modal data, which leads to a significant semantic gap between modalities. Moreover, these methods often ignore the importance differences of channels, and due to the limitation of a single goal, the matching effect between hash codes is also affected to a certain extent, thus facing many challenges. To address these issues, we propose a Dynamic Masking and Auxiliary Hash Learning (AHLR) method for enhanced cross-modal retrieval. By jointly leveraging the dynamic masking and auxiliary hash learning mechanisms, our approach effectively resolves the problems of channel information imbalance and insufficient key information capture, thereby significantly improving the retrieval accuracy. Specifically, we introduce a dynamic masking mechanism that automatically screens and weights the key information in images and texts during the training process, enhancing the accuracy of feature matching. We further construct an auxiliary hash layer to adaptively balance the weights of features across each channel, compensating for the deficiencies of traditional methods in key information capture and channel processing. In addition, we design a contrastive loss function to optimize the generation of hash codes and enhance their discriminative power, further improving the performance of cross-modal retrieval. Comprehensive experimental results on NUS-WIDE, MIRFlickr-25K and MS-COCO benchmark datasets show that the proposed AHLR algorithm outperforms several existing algorithms.
Shuang Zhang 0009, Lei Shi 0030, Feifei Kou, Huilong Jin, Pengfei Zhang 0010, Meiyu Liang, Mingying Xu
NeurIPS8
2025 A Novel Facial Expression Recognition Approach Combining Canny Edge Detection and Convolutional Neural Networks
abstract
Facial expression recognition (FER) remains challenging under pose, illumination, and occlusion. This work presents CCFER, a dual‐stream framework that couples explicit edge maps with appearance features. Grayscale faces undergo morphological closing followed by opening (5 × 5), then Canny with locally adaptive thresholds to produce clean edges for an edge branch; both streams use Dual‐Direction Attention Mixed Feature Networks (DDAMFN). Multilevel fusion employs adaptively spatial feature fusion (ASFF), followed by Efficient Local Attention (ELA) and multihead attention (MHAtt) before classification. CCFER attains 92.19% on RAF‐DB, 91.24% on FERPlus, and 67.32% on AffectNet‐7, matching or approaching the recent state of the art with balanced cross‐dataset performance. Controlled ablations (parameter‐matched single‐stream, random‐noise edges) confirm gains stem from semantic contours, and efficiency measurements show modest overhead in parameters, GFLOPs, and latency, supporting practical deployment.
Jiao Ding, Tianfei Zhang, Songlin Zhang, Meiyu Liang
IET Softw.5
2025 RFCSC: Communication efficient reinforcement federated learning with dynamic client selection and adaptive gradient compression
Zhenhui Pan, Yawen Li 0001, Zeli Guan, Meiyu Liang, Ang Li 0015, Jia Wang 0011, Feifei Kou
Neurocomputing4
2025 Robust Multi-Graph Contrastive Network for Incomplete Multi-View Clustering
abstract
Food categorization is pivotal in numerous aspects of everyday life, assisting in the selection of food, managing diets, and addressing essential survival requirements. By leveraging the complementary information of various views, multi-view learning usually achieves superior performance compared to the single-view learning methods. However, characterized by the unrestrained openness of internet platforms and potential inconsistencies in food data collection processes, multi-view features often suffer from data loss, resulting in incomplete multi-view food data. Conventional multi-view clustering methods often falter in effectively capitalizing on the diverse correlations contained in food data, and exhibit limitations in dealing with the noise and irregularities pervading different views. Addressing these challenges, this paper presents the Robust Multi-Graph Contrastive network (RMGC) for multi-view food clustering. RMGC artfully combines multi-view representation learning with multi-graph contrastive regularization, creating a cohesive framework to manage incomplete multi-view data. By developing a multi-view encoding network, RMGC seamlessly blends various views into a cohesive representation, astutely assessing the significance of each view. More importantly, the proposed robust multi-graph contrastive regularization enhances the precision of the learned representation and successfully counteracts the noise and unreliability in multi-view data. The experiments conducted across several multi-view datasets manifest the effectiveness of RMGC, showing its superiority over existing methods. Our method not only making an advancement in food categorization but also contributes to the broader field of multi-view learning, offering innovative solutions for handling incomplete and noisy multi-view data.
Zhe Xue, Yawen Li 0001, Zhongchao Guan, Wenling Li, Meiyu Liang
IEEE Trans. Multim.5
2024 Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search
abstract
Deep cross-modal hashing technology provides an effective and efficient cross-modal unified representation learning solution for cross-modal search. However, the existing methods neglect the implicit fine-grained multimodal knowledge relations between these modalities such as when the image contains information that is not directly described in the text. To tackle this problem, we propose a novel self-supervised multi-grained multi-modal knowledge graph contrastive hashing method for cross-modal search (CMGCH). Firstly, in order to capture implicit fine-grained cross-modal semantic associations, a multi-modal knowledge graph is constructed, which represents the implicit multimodal knowledge relations between the image and text as inter-modal and intra-modal semantic associations. Secondly, a cross-modal graph contrastive attention network is proposed to reason on the multi-modal knowledge graph to sufficiently learn the implicit fine-grained inter-modal and intra-modal knowledge relations. Thirdly, a cross-modal multi-granularity contrastive embedding learning mechanism is proposed, which fuses the global coarse-grained and local fine-grained embeddings by multihead attention mechanism for inter-modal and intra-modal contrastive learning, so as to enhance the cross-modal unified representations with stronger discriminativeness and semantic consistency preserving power. With the joint training of intra-modal and inter-modal contrast, the invariant and modal-specific information of different modalities can be maintained in the final unified cross-modal unified hash space. Extensive experiments on several cross-modal benchmark datasets demonstrate that the proposed CMGCH outperforms the state-of the-art methods.
Meiyu Liang, Junping Du 0001, Zhengyang Liang, Yongwang Xing, Zhe Xue
AAAI1
2024 Improving Expressive Power of Spectral Graph Neural Networks with Eigenvalue Correction
abstract
In recent years, spectral graph neural networks, characterized by polynomial filters, have garnered increasing attention and have achieved remarkable performance in tasks such as node classification. These models typically assume that eigenvalues for the normalized Laplacian matrix are distinct from each other, thus expecting a polynomial filter to have a high fitting ability. However, this paper empirically observes that normalized Laplacian matrices frequently possess repeated eigenvalues. Moreover, we theoretically establish that the number of distinguishable eigenvalues plays a pivotal role in determining the expressive power of spectral graph neural networks. In light of this observation, we propose an eigenvalue correction strategy that can free polynomial filters from the constraints of repeated eigenvalue inputs. Concretely, the proposed eigenvalue correction strategy enhances the uniform distribution of eigenvalues, thus mitigating repeated eigenvalues, and improving the fitting capacity and expressive power of polynomial filters. Extensive experimental results on both synthetic and real-world datasets demonstrate the superiority of our method.
Kangkang Lu 0002, Yanhua Yu, Hao Fei 0001, Zixuan Yang 0001, Zirui Guo, Meiyu Liang, Mengran Yin, Tat-Seng Chua
AAAI7
2024 View-Category Interactive Sharing Transformer for Incomplete Multi-View Multi-Label Learning
abstract
As a problem often encountered in real-world scenarios, multi-view multi-label learning has attracted considerable research attention. However, due to oversights in data col-lection and uncertainties in manual annotation, real-world data often suffer from incompleteness. Regrettably, most existing multi-view multi-label learning methods sidestep missing views and labels. Furthermore, they often neglect the potential of harnessing complementary information be-tween views and labels, thus constraining their classification capabilities. To address these challenges, we propose a view-category interactive sharing transformer tailored for incomplete multi-view multi-label learning. Within this net-work, we incorporate a two-layer transformer module to characterize the interplay between views and labels. Additionally, to address view incompleteness, a KNN-style missing view generation module is employed. Finally, we in-troduce a view-category consistency guided embedding en-hancement module to align different views and improve the discriminating power of the embeddings. Collectively, these modules synergistically integrate to classify the incomplete multi-view multi-label data effectively. Extensive experi-ments substantiate that our approach outperforms the ex-isting state-of-the-art methods.
Shilong Ou, Zhe Xue, Yawen Li 0001, Meiyu Liang, Yuanqiang Cai, Junjiang Wu
CVPR4
2024 Unsupervised Multimodal Graph Contrastive Semantic Anchor Space Dynamic Knowledge Distillation Network for Cross-Media Hash Retrieval
abstract
Cross-media hash retrieval are efficient and effective techniques for retrieval on multi-media database. The success of the Multimodal Large Models (MLM) provides a valuable direction to enhance the accuracy of multimodal hash retrieval, which achieves decent retrieval accuracy with finetuning the pretrained multimodal large models, but their massive model parameters significantly reduce retrieval efficiency. Knowledge Distillation (KD) methods enable small models to learn from the knowledge of larger models, achieving a reduction in model parameter count while ensuring a certain level of accuracy. However, current KD methods face challenges when applied in the multimodal domain, as it requires preserving the multimodal semantic information while minimizing accuracy degradation. To address these challenges, we propose a novel unsupervised multimodal graph contrastive semantic anchor space dynamic knowledge distillation network for cross-media hash retrieval (GASKN). Firstly, to obtain a multimodal semantic anchor space, we construct a large multimodal fusion teacher model using the BEiT-3 model as the backbone. This teacher model is capable of encoding data from different modalities, such as images and text, using the same multimodal encoder to acquire multimodal hash codes that contain rich information from both modalities simultaneously. Secondly, to ensure efficient retrieval capabilities for the student model, we utilize the ALBERT text encoding model and the BiFormer image encoding model as the compact student model's backbones. This allows us to build a lightweight student model with only a twentieth of the parameter count of the teacher model. We propose a dynamic knowledge distillation technique to transfer the multimodal semantic anchor space knowledge embedded in the multimodal large teacher model to the lightweight student model as much as possible. Thirdly, to further distill the structural knowledge of the semantic anchor space from the teacher model to the student model, we propose a graph attention contrastive learning mechanism, which enables structural semantic space learning, thereby mining implicit fine-grained cross-media semantic information. By evaluating our method using three widely-used datasets, we demonstrate that GASKN is able to significantly outperform existing state-of-the-art hashing algorithms.
Meiyu Liang, Mengran Yin, Kangkang Lu 0002, Junping Du 0001, Zhe Xue
ICDE2
2024 Knowledge Graph Enhanced Multimodal Transformer for Image-Text Retrieval
abstract
Image-text retrieval is a fundamental cross-modal task that aims to align the representation spaces between the image and text modalities. Existing cross-modal image-text retrieval methods independently generate embeddings for images and text, introduce interaction-based networks for cross-modal inference, and then achieve retrieval by using matching metrics. However, they overlook the semantic relationship between the coarse-grained and fine-grained representations within each modality, failing to capture the consistency of representations across different modalities, which affects the semantic learning of cross-modal representations, and makes it difficult to align modalities in semantic space. Consequently, these previous works inevitably suffer from low retrieval accuracy or high computational costs. In this paper, instead of directly fusing two cross-modal het-erogeneous spaces, we propose an multimodal knowledge enhanced multimodal transformer network framework to combine coarse-grained and fine-grained representation learning into a unified framework, capturing alignment information between targets, constructing a global semantic graph, and ultimately align multimodal representations in the semantic space. In our approach, images generate semantic and spatial graphs to represent visual information, while sentences generate text graphs based on semantic relationships between words, and they are used for intra-modal graph network inference. Subsequently, the generated global and local embeddings are fused into an enhanced multimodal transformer framework, effectively imple-menting cross-modal interaction processes by leveraging prior implicit semantic information from the multimodal knowledge graph. Furthermore, compared to simply matching words with image regions, our method proposes a bidirectional fine-grained matching method to filter the salient regions and words of images and texts, remove the interfering noise information, and realize bidirectional fine-grained pairing, which captures fine-grained bi-directional representational information, thus enable the model to generate more discriminative representations Finally, equipped with a coarse-to-fine inference method based on hybrid global and local cross-modal similarities, we demonstrate that the proposed method is able to significantly outperform existing state-of-the-art algorithms by evaluating our method using two widely-used datasets.
Juncheng Zheng, Meiyu Liang, Yawen Li 0001, Zhe Xue
ICDE2
2024 A Multi-View Double Alignment Hashing Network with Weighted Contrastive Learning
abstract
Multi-view retrieval faces significant pressure due to the rapidly increasing multi-view information on the internet. The multi-view hashing method turns continuous features into compact information of fixed length and considerably improves retrieval efficiency. However, existing multi-view hashing methods neglect the bias produced during multi-view alignment and multi-label guidance processes. To address these issues, we introduce a novel multi-view hash method that learns compact hash codes. It first employs a multi-view double alignment module to align features from different views. Then, it utilizes a self-adjusted cross-attention fusion module to fuse these features. Finally, we propose a weighted contrastive learning module to learn more discriminative representations, smoothing the differences among all samples. Extensive experiments show that our method yields compact hash codes and outperforms state-of-the-art methods.
Tianlong Zhang, Zhe Xue, Yuchen Dong, Junping Du 0001, Meiyu Liang
ICME5
2024 Structures Aware Fine-Grained Contrastive Adversarial Hashing for Cross-Media Retrieval
abstract
Deep cross-media hashing provides an efficient semantic representation learning solution for large-scale cross-media retrieval. The existing methods only consider the inter-media or intra-media semantic association learning, ignore the guiding of semantic structure information, and have weak reasoning ability for implicit fine-grained semantic associations. To tackle this problem, we propose a novel structures aware fine-grained contrastive adversarial hashing method for cross-media retrieval. A novel cross-media contrastive adversarial hash network is constructed for the first time, which integrates the cross-media and intra-media contrastive learning and multi-modal adversarial learning, aiming at maximizing the semantic association between different modalities, and improving the semantic discrimination and consistency of cross-media unified hash representation, thereby the inter-media and intra-media semantic preserving ability can be well enhanced; A fine-grained cross-media semantic feature learning method based on fine-grained semantic reasoning with transformers is proposed, which captures fine-grained salient features of different modalities for semantic association learning, and enhances the reasoning ability of fine-grained implicit semantic association; A semantic label graph convolutional network guided cross-media semantic association learning strategy is proposed, which makes full use of semantic structure information to enhance the learning ability of implicit cross-media semantic associations. Extensive experiments on several large-scale cross-media benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods.
Meiyu Liang, Yawen Li 0001, Xiaowen Cao 0003, Zhe Xue, Ang Li 0015, Kangkang Lu 0002
IEEE Trans. Knowl. Data Eng.1
2023 Video Super-Resolution Reconstruction Based on Deep Learning and Spatio-Temporal Feature Self-similarity (Extended abstract)
abstract
Video super-resolution (SR) reconstruction technology aims at obtaining high quality reconstruction of high-resolution (HR) video sequences by inferring the lost detailed information from their low-resolution (LR) counterparts. However, this technology is an ill-posed problem because significant detailed information is lost in the process of video degrading. The existing learning-based SR reconstruction methods can be adapted to a larger super-resolution factor, but it cannot be guaranteed that any low-resolution image block can find its corresponding high-resolution block matching in a limited-scale training set. Some noise and over smooth phenomenon usually exist while dealing with some unique features that rarely appear in a given training data set. The self-similarity based SR methods do not rely on accurate sub-pixel motion estimation and thus can be adapted to complex motion patterns. However, under conditions of insufficient internal similar blocks, some visual flaws are usually produced due to the mismatched internal instances.
Meiyu Liang, Junping Du 0001, Zhe Xue, Xiaoxiao Wang 0006, Feifei Kou
ICDE1
2023 Deep Unsupervised Momentum Contrastive Hashing for Cross-modal Retrieval
abstract
Unsupervised cross-modal hashing (UCMH) methods often start from the similarity of sample features and design a reconstruction loss to achieve similarity preservation. However, these methods suffer from inaccurate similarity problems, be-cause different feature representations may share similar semantic information. In this paper, we propose Deep Unsupervised Momentum Contrastive Hashing (DUMCH). Specifically, we introduce momentum contrastive learning for unsupervised cross-modal hashing, which allows us to flexibly define a robust loss by comparing positive and negative samples. Moreover, in order to achieve similarity retention of hash codes in Hamming space and fully utilize the potential of contrastive learning in Hamming space, we remove the L2 normalization corresponding to cosine similarity and design a novel normalization method called hash normalization, which has been proved to greatly improve the model performance. We conducted extensive experiments on three datasets, and the experimental results demonstrate the superiority of DUMCH.
Kangkang Lu 0002, Yanhua Yu, Meiyu Liang, Xiaowen Cao 0003, Zehua Zhao, Mengran Yin, Zhe Xue
ICME3
2023 RTMC: A Rubost Trusted Multi-View Classification Framework
abstract
Multi-view learning aims to fully exploit the information from multiple sources to obtain better performance than using a single view. However, real-world data often contains a lot of noise, which can have a large impact on multi-view learning. Therefore, it is necessary to identify noise contained in multi-view data to achieve robust and trusted classification. In this paper, we propose a robust trusted multi-view classification framework, RTMC. Our framework uses multi-view affinity and repellence encoding to learn effective latent encodings of multi-view data. We also propose a trust-aware discriminator to estimate trust scores by identifying noise contained in the data. We adopt prototype queues, which store latent encodings of different classes, to accurately identify the noise. Finally, trusted multi-view classification is proposed to jointly predict the trust scores of classification and achieve robust classification results through a trusted fusion strategy. RTMC is validated on six challenging multi-view datasets and the experimental results demonstrate the robustness and effectiveness of our method.
Zhe Xue, Boang Li, Junping Du 0001, Meiyu Liang
ICME6
2023 CALM: An Enhanced Encoding and Confidence Evaluating Framework for Trustworthy Multi-view Learning
abstract
Multi-view learning aims to leverage data acquired from multiple sources to achieve better performance compared to using a single view. However, the performance of multi-view learning can be negatively impacted by noisy or corrupted views in certain real-world situations. As a result, it is crucial to assess the confidence of predictions and obtain reliable learning outcomes. In this paper, we introduce CALM, an enhanced encoding and confidence evaluation framework for trustworthy multi-view classification. Our method comprises enhanced multi-view encoding, multi-view confidence-aware fusion, and multi-view classification regularization, enabling the simultaneous evaluation of prediction confidence and the yielding trustworthy classifications. Enhanced multi-view encoding takes advantage of cross-view consistency and class diversity to improve the efficacy of the learned latent representation, facilitating more reliable classification results. Multi-view confidence-aware fusion utilizes a confidence-aware estimator to evaluate the confidence scores of classification outcomes. The final multi-view classification results are then derived through confidence-aware fusion. To achieve reliable and accurate confidence scores, multivariate Gaussian distributions are employed to model the prediction distribution. The advantage of CALM lies in its ability to evaluate the quality of each view, reducing the influence of low-quality views on the multi-view fusion process and ultimately leading to improved classification performance and confidence evaluation. Comprehensive experimental results demonstrate that our method outperforms other trusted multi-view learning methods in terms of effectiveness, reliability, and robustness.
Zhe Xue, Boang Li, Junping Du 0001, Meiyu Liang, Yuankai Qi
ACM Multimedia6
2023 STRAN: Student expression recognition based on spatio-temporal residual attention network in classroom teaching videos
Meiyu Liang, Zhe Xue, Wanying Yu
Appl. Intell.2
2023 Event-Aware Video Deraining via Multi-Patch Progressive Learning
abstract
In this paper, we address the problem of video-based rain streak removal by developing an event-aware multi-patch progressive neural network. Rain streaks in video exhibit correlations in both temporal and spatial dimensions. Existing methods have difficulties in modeling the characteristics. Based on the observation, we propose to develop a module encoding events from neuromorphic cameras to facilitate deraining. Events are captured asynchronously at pixel-level only when intensity changes by a margin exceeding a certain threshold. Due to this property, events contain considerable information about moving objects including rain streaks passing though the camera across adjacent frames. Thus we suggest that utilizing it properly facilitates deraining performance non-trivially. In addition, we develop a multi-patch progressive neural network. The multi-patch manner enables various receptive fields by partitioning patches and the progressive learning in different patch levels makes the model emphasize each patch level to a different extent. Extensive experiments show that our method guided by events outperforms the state-of-the-art methods by a large margin in synthetic and real-world datasets.
Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Kaihao Zhang, Meiyu Liang, Xiaochun Cao
IEEE Trans. Image Process.5
2022 Semantic Structure Enhanced Contrastive Adversarial Hash Network for Cross-media Representation Learning
abstract
Deep cross-media hashing technology provides an efficient cross-media representation learning solution for cross-media search. However, the existing methods do not consider both fine-grained semantic features and semantic structures to mine implicit cross-media semantic associations, which leads to weaker semantic discrimination and consistency for cross-media representation. To tackle this problem, we propose a novel semantic structure enhanced contrastive adversarial hash network for cross-media representation learning (SCAHN). Firstly, in order to capture more fine-grained cross-media semantic associations, a fine-grained cross-media attention feature learning network is constructed, thus the learned saliency features of different modalities are more conducive to cross-media semantic alignment and fusion. Secondly, for further improving learning ability of implicit cross-media semantic associations, a semantic label association graph is constructed, and the graph convolutional network is utilized to mine the implicit semantic structures, thus guiding learning of discriminative features of different modalities. Thirdly, a cross-media and intra-media contrastive adversarial representation learning mechanism is proposed to further enhance the semantic discriminativeness of different modal representations, and a dual-way adversarial learning strategy is developed to maximize cross-media semantic associations, so as to obtain cross-media unified representations with stronger discriminativeness and semantic consistency preserving power. Extensive experiments on several cross-media benchmark datasets demonstrate that the proposed SCAHN outperforms the state-of-the-art methods.
Meiyu Liang, Junping Du 0001, Xiaowen Cao 0003, Kangkang Lu 0002, Zhe Xue
ACM Multimedia1
2022 Robust Diversified Graph Contrastive Network for Incomplete Multi-view Clustering
abstract
Incomplete multi-view clustering is a challenging task which aims to partition the unlabeled incomplete multi-view data into several clusters. The existing incomplete multi-view clustering methods neglect to utilize the diversified correlations inherent in data and handle the noise contained in different views. To address these issues, we propose a Robust Diversified Graph Contrastive Network (RDGC) for incomplete multi-view clustering, which integrates multi-view representation learning and diversified graph contrastive regularization into a unified framework. Multi-view unified and specific encoding network is developed to fuse different views into a unified representation, which can flexibly estimate the importance of views for incomplete multi-view data. Robust diversified graph contrastive regularization is proposed which captures the diversified data correlations to improve the discriminating power of the learned representation and reduce the information loss caused by the view missing problem. Moreover, our method can effectively resist the influence of noise and unreliable views by leveraging the robust contrastive learning loss. Extensive experiments conducted on four multi-view clustering datasets demonstrate the superiority of our method over the state-of-the-art methods.
Zhe Xue, Junping Du 0001, Zhongchao Guan, Meiyu Liang
ACM Multimedia7
2022 Video Super-Resolution Reconstruction Based on Deep Learning and Spatio-Temporal Feature Self-Similarity
abstract
To address the problems in the existing video super-resolution methods, such as noise, over smooth and visual artifacts, which are caused by the reliance on limited external training or mismatch of internal similarity patch instances, this study proposes a novel video super-resolution reconstruction algorithm based on deep learning and spatio-temporal feature similarity (DLSS-VSR). The video super-resolution reconstruction mechanism with the joint internal and external constraints is established utilizing the complementary advantages of both external deep correlation mapping learning and internal spatio-temporal nonlocal self-similarity prior constraint. A deep learning model based on deep convolutional neural network is constructed to learn the nonlinear correlation mapping between low-resolution and high-resolution video frame patches. A novel spatio-temporal feature similarity calculation method is proposed, which considers both internal video spatio-temporal self-similarity and external clean nonlocal similarity. For the internal spatio-temporal feature self-similarity, we improve the accuracy and robustness of similarity matching by proposing a similarity measure strategy based on spatio-temporal moment feature similarity and structural similarity. The external nonlocal similarity prior constraint is learned by the patch group-based Gaussian mixture model. The time efficiency for spatio-temporal similarity matching is further improved based on saliency detection and region correlation judgment strategy, which achieves a better tradeoff between super-resolution accuracy and speed. Experimental results demonstrate that the DLSS-VSR algorithm achieves competitive super-resolution quality compared to other state-of-the-art algorithms in both subjective and objective evaluations.
Meiyu Liang, Junping Du 0001, Zhe Xue, Xiaoxiao Wang 0006, Feifei Kou
IEEE Trans. Knowl. Data Eng.1
2021 Clustering-Induced Adaptive Structure Enhancing Network for Incomplete Multi-View Data
abstract
Incomplete multi-view clustering aims to cluster samples with missing views, which has drawn more and more research interest. Although several methods have been developed for incomplete multi-view clustering, they fail to extract and exploit the comprehensive global and local structure of multi-view data, so their clustering performance is limited. This paper proposes a Clustering-induced Adaptive Structure Enhancing Network (CASEN) for incomplete multi-view clustering, which is an end-to-end trainable framework that jointly conducts multi-view structure enhancing and data clustering. Our method adopts multi-view autoencoder to infer the missing features of the incomplete samples. Then, we perform adaptive graph learning and graph convolution on the reconstructed complete multi-view data to effectively extract data structure. Moreover, we use multiple kernel clustering to integrate the global and local structure for clustering, and the clustering results in turn are used to enhance the data structure. Extensive experiments on several benchmark datasets demonstrate that our method can comprehensively obtain the structure of incomplete multi-view data and achieve superior performance compared to the other methods.
Zhe Xue, Junping Du 0001, Changwei Zheng, Wenqi Ren, Meiyu Liang
IJCAI6
2020 A very deep two-stream network for crowd type recognition
Xinlei Wei, Junping Du 0001, Zhe Xue, Meiyu Liang, Yue Geng, Jang-Myung Lee
Neurocomputing4
2020 Security topics related microblogs search based on deep convolutional neural networks
Junping Du 0001, Zhe Xue, Meiyu Liang, Wan-Qiu Cui
Neurocomputing4
2020 Cross-Media Semantic Correlation Learning Based on Deep Hash Network and Semantic Expansion for Social Network Cross-Media Search
abstract
Cross-media search from large-scale social network big data has become increasingly valuable in our daily life because it can support querying different data modalities. Deep hash networks have shown high potential in achieving efficient and effective cross-media search performance. However, due to the fact that social network data often exhibit text sparsity, diversity, and noise characteristics, the search performance of existing methods often degrades when dealing with this data. In order to address this problem, this article proposes a novel end-to-end cross-media semantic correlation learning model based on a deep hash network and semantic expansion for social network cross-media search (DHNS). The approach combines deep network feature learning and hash-code quantization learning for multimodal data into a unified optimization architecture, which successfully preserves both intramedia similarity and intermedia correlation, by minimizing both cross-media correlation loss and binary hash quantization loss. In addition, our approach realizes semantic relationship expansion by constructing the image-word relation graph and mining the potential semantic relationship between images and words, and obtaining the semantic embedding based on both internal graph deep walk and an external knowledge base. Experimental results demonstrate that DHNS yields better cross-media search performance on standard benchmarks.
Meiyu Liang, Junping Du 0001, Cong-Xian Yang, Zhe Xue, Hai-Sheng Li 0002, Feifei Kou, Yue Geng
IEEE Trans. Neural Networks Learn. Syst.1
2020 A content search method for security topics in microblog based on deep reinforcement learning
Junping Du 0001, Wan-Qiu Cui, Zhe Xue, Meiyu Liang
World Wide Web6
2019 Fine-grained Cross-media Representation Learning with Deep Quantization Attention Network
abstract
Cross-media search is useful for getting more comprehensive and richer information about social network hot topics or events. To solve the problems of feature heterogeneity and semantic gap of different media data, existing deep cross-media quantization technology provides an efficient and effective solution for cross-media common semantic representation learning. However, due to the fact that social network data often exhibits semantic sparsity, diversity, and contains a lot of noise, the performance of existing cross-media search methods often degrades. To address the above issue, this paper proposes a novel fine-grained cross-media representation learning model with deep quantization attention network for social network cross-media search (CMSL). First, we construct the image-word semantic correlation graph, and perform deep random walks on the graph to realize semantic expansion and semantic embedding learning, which can discover some potential semantic correlations between images and words. Then, in order to discover more fine-grained cross-media semantic correlations, a multi-scale fine-grained cross-media semantic correlation learning method that combines global and local saliency semantic similarity is proposed. Third, the fine-grained cross-media representation, cross-media semantic correlations and binary quantization code are jointly learned by a unified deep quantization attention network, which can preserve both inter-media correlations and intra-media similarities, by minimizing both cross-media correlation loss and binary quantization loss. Experimental results demonstrate that CMSL can generate high-quality cross-media common semantic representation, which yields state-of-the-art cross-media search performance on two benchmark datasets, NUS-WIDE and MIR-Flickr 25k.
Meiyu Liang, Junping Du 0001, Wu Liu 0005, Zhe Xue, Yue Geng, Cong-Xian Yang
ACM Multimedia1
2019 A multi-feature probabilistic graphical model for social network semantic search
Feifei Kou, Junping Du 0001, Cong-Xian Yang, Yan-Song Shi, Meiyu Liang, Zhe Xue, Hai-Sheng Li 0002
Neurocomputing5
2019 Dynamic topic modeling via self-aggregation for short text streams
Lei Shi 0030, Junping Du 0001, Meiyu Liang, Feifei Kou
Peer-to-Peer Netw. Appl.3
2019 Boosting deep attribute learning via support vector regression for fast moving crowd counting
Xinlei Wei, Junping Du 0001, Meiyu Liang, Lingfei Ye
Pattern Recognit. Lett.3
2019 Extended search method based on a semantic hashtag graph combining social and conceptual information
Wan-Qiu Cui, Junping Du 0001, Dawei Wang 0009, Feifei Kou, Meiyu Liang, Zhe Xue
World Wide Web5
2019 Abnormal event detection in tourism video based on salient spatio-temporal features and sparse combination learning
Yue Geng, Junping Du 0001, Meiyu Liang
World Wide Web3
2018 Hashtag Recommendation Based on Multi-Features of Microblogs
Feifei Kou, Junping Du 0001, Cong-Xian Yang, Yan-Song Shi, Wan-Qiu Cui, Meiyu Liang, Yue Geng
J. Comput. Sci. Technol.6
2016 Video super-resolution reconstruction based on correlation learning and spatio-temporal nonlocal similarity
Meiyu Liang, Junping Du 0001
Multim. Tools Appl.1
2015 Super-resolution reconstruction based on multisource bidirectional similarity and non-local similarity matching
abstract
A novel super‐resolution (SR) reconstruction algorithm based on multisource bidirectional similarity and non‐local similarity matching for multi‐exposure dynamic image sequences is proposed in this study. First, luminance compensation for multi‐exposure dynamic images is implemented based on multisource bidirectional similarity. Then, combining a self‐adaptive regional correlation evaluation strategy and a weighting strategy based on pseudo‐Zernike moment feature similarity and structural similarity yields a novel and robust non‐local similarity matching scheme (PZ‐NSM) to learn non‐local similarity priors between low‐ and high‐resolution dynamic image patches at different spatio‐temporal scales without additional training. Finally, SR reconstruction is implemented based on PZ‐NSM and spatio‐temporal trans‐scale fusion of non‐local similarities between dynamic images. The proposed algorithm does not rely on accurate estimation of subpixel motion and can therefore be adapted to more complex motion patterns. It has high rotation invariance effectiveness and is robust to noise and illumination. Experimental results demonstrate that the proposed algorithm outperforms existing algorithms in terms of both subjective and objective evaluations.
Meiyu Liang, Junping Du 0001, Shouxin Cao
IET Image Process.1
2014 Self-adaptive spatial image denoising model based on scale correlation and SURE-LET in the nonsubsampled contourlet transform domain
Meiyu Liang, Junping Du 0001, Honggang Liu
Sci. China Inf. Sci.1