EDBT 2026 Demo / reviewers in the wild / expert
Yanzhao Xie
dblp:261/6589
· DBLP profile ↗
28ranked-venue papers
7as first author
26since 2021 · last 2026
0000-0002-9274-2807ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language ModelsabstractKnowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines. Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang |
AAAI | 3 |
| 2026 | Skew-normal distributions for modeling asymmetric moving tendencies in pedestrian trajectories
Siyuan Chen 0005, Yatie Xiao, Yangtao Wang, Yanzhao Xie, Tong Zhu 0003, Jinbiao Chen |
Neurocomputing | 4 |
| 2026 | Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002 |
Pattern Recognit. | 3 |
| 2026 | F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 3 |
| 2026 | Prompt-affinity multi-modal class centroids for unsupervised domain adaptionabstractIn recent years, the advancements in large vision-language models (VLMs) like CLIP have sparked a renewed interest in leveraging the prompt learning mechanism to preserve semantic consistency between source and target domains in unsupervised domain adaption (UDA). While these approaches show promising results, they encounter fundamental limitations when quantifying the similarity between source and target domain data , primarily stemming from the redundant and modality-missing class centroids . To address these limitations, we propose P rompt-affinity M ulti-modal C lass C entroids for UDA (termed as PMCC). Firstly, we fuse the text class centroids (directly generated from the text encoder of CLIP with manual prompts for each class) and image class centroids (generated from the image encoder of CLIP for each class based on source domain images) to yield the multi-modal class centroids. Secondly, we conduct the cross-attention operation between each source or target domain image and these multi-modal class centroids. In this way, these class centroids that contain rich semantic information of each class will serve as a bridge to effectively measure the semantic similarity between different domains. Finally, we design a logit bias head and employ a multi-modal prompt learning mechanism to accurately predict the true class of each image for both source and target domains. We conduct extensive experiments on 4 popular UDA datasets including Office-31, Office-Home, VisDA-2017, and DomainNet. The experimental results validate our PMCC achieves higher performance with lower model complexity than the state-of-the-art (SOTA) UDA methods. The code of this project is available at GitHub: https://github.com/246dxw/PMCC . Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 3 |
| 2026 | Cross-domain distillation for unsupervised domain adaptation with large vision-language models
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 3 |
| 2026 | Adaptive message passing mechanism for graph neural networks
Yangtao Wang, Linruo Liu, Yanzhao Xie, Maobin Tang, Xiaocui Li 0001 |
Pattern Recognit. | 5 |
| 2026 | PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002 |
Pattern Recognit. | 1 |
| 2026 | MKGPL: graph prompt learning with multi-view knowledge for few-shot recognition
Yanzhao Xie, Man Qiu, Yangtao Wang, Siyuan Chen 0005, Meie Fang, Maobin Tang, Wensheng Zhang 0002 |
Pattern Recognit. | 1 |
| 2026 | Progressive Hybrid Pseudo-Labeling for Unsupervised Domain Adaptation With Ascending Low-Rank AdaptationabstractUnsupervised domain adaptation (UDA) based on large vision-language models (VLMs) has recently demonstrated strong generalization ability, yet it remains fundamentally challenged by noisy pseudo-labels and inefficient adaptation under large domain shifts. In this paper, we propose Progressive Hybrid Pseudo-Labeling for UDA with Ascending Low-Rank Adaptation (termed as PHPL), a parameter-efficient paradigm that addresses these challenges from two complementary perspectives. 1) We introduce a progressive hybrid pseudo-labeling strategy that constructs target-domain supervision by fusing predictions from a frozen teacher model and an adaptive student model with a progressive weighting scheme. By gradually transferring predictive responsibility from the teacher to the student during training, PHPL effectively mitigates early-stage pseudo-label noise and stabilizes self-training under large domain shifts. 2) To enable efficient and stable adaptation of large VLMs, we propose an ascending low-rank adaptation strategy that allocates LoRA capacity in a depth-aware manner. Specifically, larger low-rank updates are assigned to deeper, semantically richer layers, while shallow layers remain lightly parameterized, striking a favorable balance between parameter efficiency and representational expressiveness. We conduct extensive experiments on five widely-used UDA benchmarks, including Office-Home, Office-31, VisDA-2017, Mini-DomainNet, and DomainNet. Experimental results verify that PHPL consistently achieves higher performance across various cross-domain scenarios compared with existing CNN, Transformer, and VLMs-based solutions. Notably, PHPL demonstrates strong robustness on highly challenging large-scale conditions while requiring significantly less computational overhead, validating the effectiveness and scalability of the proposed lightweight adaptation paradigm. The code is available at https://github.com/el2k/PHPL. Yangtao Wang, Mingxin Huang, Xingwei Deng, Yanzhao Xie, Xiaocui Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream TransformersabstractRecently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git . Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text RetrievalabstractImage-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA. Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016 |
ICME | 4 |
| 2025 | Graph Contrastive-and-Reconstructive Hashing for Unsupervised Cross-Modal RetrievalabstractAbstract Hashing-based unsupervised cross-modal retrieval has gained significant attention in the big data management community due to its low storage overhead and rapid retrieval speed. However, current methods often lack effective alignment strategies to reduce the modality gap. They also fail to explore the latent structural information of the training data for accurate relationship learning, resulting in sub-optimal cross-modal retrieval performance. To tackle these challenges, we propose a novel unsupervised cross-modal hashing method called G raph C ontrastive-and- R econstructive H ashing ( GCRH ). Specifically, GCRH first performs global graph contrastive learning , which involves both intra-modal and inter-modal pairs. This facilitates the learning of more discriminative hash codes through intra-modal discrimination and inter-modal alignment objectives. To further bridge the modality gap, GCRH conducts local graph reconstruction using GCN-based decoders to reconstruct the original features of one modality from the hash codes of another. The integration of contrastive-and-reconstructive learning with graph structural information enables GCRH to generate high-quality hash codes that are both well-aligned and discriminative. Extensive experiments on three benchmark datasets substantiate the superior cross-modal retrieval performance of GCRH . Rukai Wei, Yu Liu 0040, Heng Cui, Yanzhao Xie, Ke Zhou 0001 |
Data Sci. Eng. | 4 |
| 2025 | Angle Metric Learning for Discriminative Features on Vehicle Re-IdentificationabstractABSTRACT Vehicle re‐identification (Re‐ID) facilitates the recognition and distinction of vehicles based on their visual characteristics in images or videos. However, accurately identifying a vehicle poses great challenges due to (i) the pronounced intra‐instance variations encountered under varying lighting conditions such as day and night and (ii) the subtle inter‐instance differences observed among similar vehicles. To address these challenges, the authors propose A ngle M etric learning for D iscriminative F eatures on vehicle Re‐ID (termed as AMDF), which aims to maximise the variance between visual features of different classes while minimising the variance within the same class. AMDF comprehensively measures the angle and distance discrepancies between features. First, to mitigate the impact of lighting conditions on intra‐class variation, the authors employ CycleGAN to generate images that simulate consistent lighting (either day or night), thereby standardising the conditions for distance measurement. Second, Swin Transformer was integrated to help generate more detailed features. At last, a novel angle metric loss based on cosine distance is proposed, which organically integrates angular metric and 2‐norm metric, effectively maximising the decision boundary in angular space. Extensive experimental evaluations on three public datasets including VERI‐776, VERI‐Wild, and VEHICLEID, indicate that the method achieves state‐of‐the‐art performance. The code of this project is released at https://github.com/ZnCu‐0906/AMDF . Yutong Xie 0009, Shuoqi Zhang, Lide Guo, Rukai Wei, Yanzhao Xie, Yangtao Wang, Maobin Tang, Lisheng Fan |
IET Comput. Vis. | 6 |
| 2025 | HSALC: hard sample aware label correction for medical image classification
Yangtao Wang, Yicheng Ye, Yanzhao Xie, Maobin Tang, Lisheng Fan |
Multim. Tools Appl. | 3 |
| 2024 | Image-text Retrieval with Main Semantics ConsistencyabstractImage-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC. Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang |
CIKM | 3 |
| 2024 | Contrastive masked auto-encoders based self-supervised hashing for 2D image and 3D point cloud cross-modal retrievalabstractImplementing cross-modal hashing between 2D images and 3D point-cloud data is a growing concern in real-world retrieval systems. Simply applying existing cross-modal approaches to this new task fails to adequately capture latent multi-modal semantics and effectively bridge the modality gap between 2D and 3D. To address these issues without relying on hand-crafted labels, we propose contrastive masked autoencoders based self-supervised hashing (CMAH) for retrieval between images and point-cloud data. We start by contrasting 2D-3D pairs and explicitly constraining them into a joint Hamming space. This contrastive learning process ensures robust discriminability for the generated hash codes and effectively reduces the modality gap. Moreover, we utilize multi-modal auto-encoders to enhance the model’s understanding of multi-modal semantics. By completing the masked image/point-cloud data modeling task, the model is encouraged to capture more localized clues. In addition, the proposed multi-modal fusion block facilitates fine-grained interactions among different modalities. Extensive experiments on three public datasets demonstrate that the proposed CMAH significantly outperforms all baseline methods. Rukai Wei, Heng Cui, Yu Liu 0040, Yanzhao Xie, Yufeng Hou, Ke Zhou 0001 |
ICME | 4 |
| 2024 | Exploring Hierarchical Information in Hyperbolic Space for Self-Supervised Image HashingabstractIn real-world datasets, visually related images often form clusters, and these clusters can be further grouped into larger categories with more general semantics. These inherent hierarchical structures can help capture the underlying distribution of data, making it easier to learn robust hash codes that lead to better retrieval performance. However, existing methods fail to make use of this hierarchical information, which in turn prevents the accurate preservation of relationships between data points in the learned hash codes, resulting in suboptimal performance. In this paper, our focus is on applying visual hierarchical information to self-supervised hash learning and addressing three key challenges, including the construction, embedding, and exploitation of visual hierarchies. We propose a new self-supervised hashing method named Hierarchical Hyperbolic Contrastive Hashing (HHCH), making breakthroughs in three aspects. First, we propose to embed continuous hash codes into hyperbolic space for accurate semantic expression since embedding hierarchies in the hyperbolic space generates less distortion than in the hyper-sphere or Euclidean space. Second, we update the K-Means algorithm to make it run in the hyperbolic space. The proposed hierarchical hyperbolic K-Means algorithm can achieve the adaptive construction of hierarchical semantic structures. Last but not least, to exploit the hierarchical semantic structures in hyperbolic space, we propose the hierarchical contrastive learning algorithm, including hierarchical instance-wise and hierarchical prototype-wise contrastive learning. Extensive experiments on four benchmark datasets demonstrate that the proposed method outperforms state-of-the-art self-supervised hashing methods. Our codes are released at https://github.com/HUST-IDSM-AI/HHCH.git. Rukai Wei, Yu Liu 0040, Jingkuan Song, Yanzhao Xie, Ke Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingabstractCompressing videos into binary codes can improve retrieval speed and reduce storage overhead. However, learning accurate hash codes for video retrieval can be challenging due to high local redundancy and complex global dependencies between video frames, especially in the absence of labels. Existing self-supervised video hashing methods have been effective in designing expressive temporal encoders, but have not fully utilized the temporal dynamics and spatial appearance of videos due to less challenging and unreliable learning tasks. To address these challenges, we begin by utilizing the contrastive learning task to capture global spatio-temporal information of videos for hashing. With the aid of our designed augmentation strategies, which focus on spatial and temporal variations to create positive pairs, the learning framework can generate hash codes that are invariant to motion, scale, and viewpoint. Furthermore, we incorporate two collaborative learning tasks, i.e., frame order verification and scene change regularization, to capture local spatio-temporal details within video frames, thereby enhancing the perception of temporal structure and the modeling of spatio-temporal relationships. Our proposed Contrastive Hashing with Global-Local Spatio-temporal Ibnformation (CHAIN) outperforms state-of-the-art self-supervised video hashing methods on four video benchmark datasets. Our codes will be released. Rukai Wei, Yu Liu 0040, Jingkuan Song, Heng Cui, Yanzhao Xie, Ke Zhou 0001 |
ACM Multimedia | 5 |
| 2023 | How visual chirality affects the performance of image hashing
Yanzhao Xie, Guangxing Hu, Yu Liu 0040, Zhiqiu Lin, Ke Zhou 0001 |
Neural Comput. Appl. | 1 |
| 2023 | A hash centroid construction method with Swin transformer for multi-label image retrieval
Yanzhao Xie, Yangtao Wang, Rukai Wei, Yu Liu 0040, Ke Zhou 0001, Lisheng Fan |
Neural Comput. Appl. | 1 |
| 2023 | Deep debiased contrastive hashing
Rukai Wei, Yu Liu 0040, Jingkuan Song, Yanzhao Xie, Ke Zhou 0001 |
Pattern Recognit. | 4 |
| 2023 | Label-Affinity Self-Adaptive Central Similarity Hashing for Image RetrievalabstractDue to the usage of global similarity, the hashing methods based on predefined hash centers have achieved more accurate retrieval results than the pairwise/triplet-based methods. Nevertheless, the fixed hash centers lack the perception of data distribution and are limited by the pre-determined Hadamard matrix, which consider neither the label semantic information nor the object scale size, resulting in sub-optimal retrieval performance and weak generalization ability. In this paper, we (1) adopt the label semantic information to generate self-adaptive hash centers and (2) propose the label-affinity coefficient (lac) that considers the scale size of each label/object appearing in the given image to calculate the real hash centroid for this image. Based on this, we proposeLabel-affinity Self-adaptive Central Similarity Hashing (LSCSH)for image retrieval. LSCSH consists of a hash code generator module and a hash center adapter module. First, we obtain the label word vector (i.e., the word vector representation of each class label) via the Word2Vector technique to generate and update the hash centers that adapt to the distribution of both label word vectors and generated hash codes. Second, we learnlacto indicate the dominance of different labels corresponding to objects in each given image, which considers the unequal scales of each object (corresponding to a label) to calculate a more accurate hash centroid for each image. Last but not least, we design an asynchronous learning mechanism to enable each hash code and its corresponding hash centroid to adapt to each other dynamically. We conduct extensive experiments on 5 image datasets including CIFAR-10, ImageNet, VOC2012, MS-COCO and NUS-WIDE. The experimental results demonstrate that LSCSH can achieve the state-of-the-art visual retrieval performance on both single-label and multi-label image datasets. The code of this work is released at:https://github.com/lzHZWZ/LSCSH_sourcecode.git. Yanzhao Xie, Rukai Wei, Jingkuan Song, Yu Liu 0040, Yangtao Wang, Ke Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Label graph learning for multi-label image recognition with cross-modal fusion
Yanzhao Xie, Yangtao Wang, Yu Liu 0040, Ke Zhou 0001 |
Multim. Tools Appl. | 1 |
| 2022 | STMG: Swin transformer for multi-label image recognition with graph convolution network
Yangtao Wang, Yanzhao Xie, Lisheng Fan, Guangxing Hu |
Neural Comput. Appl. | 2 |
| 2021 | G-CAM: Graph Convolution Network Based Class Activation Mapping for Multi-label Image RecognitionabstractIn most multi-label image recognition tasks, human visual perception keeps consistent for different spatial transforms of the same image. Existing approaches either learn the perceptual consistency with only image-level supervision or preserve the middle-level feature consistency of attention regions but neglect the (global) label dependencies between different objects over the dataset. To address this issue, we integrate graph convolution network (GCN) and propose G-CAM, which learns visual attention consistency via GCN based class attention mapping (CAM) for multi-label image recognition. G-CAM consists of an image feature extraction module to generate the feature maps of the original image and its transformed one and a GCN module to learn weighted classifiers that capture the label dependencies between different objects. Different from previous works which use fully-connected classification layer, G-CAM first fuses weighted classifiers with the feature vector to generate the predicted labels for each input image, then combines weighted classifiers with the feature maps to respectively obtain the transformed attention heatmaps of the original image and the attention heatmaps of its transformed one. We can compute the attention consistency loss according to the distance between these two attention heatmaps. Finally, this loss is combined with the multi-label classification loss to update the whole network in an end-to-end manner. We conduct extensive experiments on three multi-label image datasets including FLICKR25K, MS-COCO and NUS-WIDE. Experimental results demonstrate G-CAM can achieve better performance compared with the state-of-the-art multi-label image recognition methods. Yangtao Wang, Yanzhao Xie, Yu Liu 0040, Lisheng Fan |
ICMR | 2 |
| 2020 | Fast Graph Convolution Network Based Multi-label Image Recognition via Cross-modal FusionabstractIn multi-label image recognition, it has become a popular method to predict those labels that co-occur in an image via modeling the label dependencies. Previous works focus on capturing the correlation between labels, but neglect to effectively fuse the image features and label embeddings, which severely affects the convergence efficiency of the model and inhibits the further precision improvement of multi-label image recognition. To overcome this shortcoming, in this paper, we introduce Multi-modal Factorized Bilinear pooling (MFB) which works as an efficient component to fuse cross-modal embeddings and propose F-GCN, a fast graph convolution network (GCN) based multi-label image recognition model. F-GCN consists of three key modules: (1) an image representation learning module which adopts a convolution neural network (CNN) to learn and generate image representations, (2) a label co-occurrence embedding module which first obtains the label vectors via the word embeddings technique and then adopts GCN to capture label co-occurrence embeddings and (3) an MFB fusion module which efficiently fuses these cross-modal vectors to enable an end-to-end model with a multi-label loss function. We conduct extensive experiments on two multi-label datasets including MS-COCO and VOC2007. Experimental results demonstrate the MFB component efficiently fuses image representations and label co-occurrence embeddings and thus greatly improves the convergence efficiency of the model. In addition, the performance of image recognition has also been promoted compared with the state-of-the-art methods. Yangtao Wang, Yanzhao Xie, Yu Liu 0040, Ke Zhou 0001, Xiaocui Li 0001 |
CIKM | 2 |
| 2020 | Label-Attended Hashing for Multi-Label Image RetrievalabstractFor the multi-label image retrieval, the existing hashing algorithms neglect the dependency between objects and thus fail to capture the attention information in the feature extraction, which affects the precision of hash codes. To address this problem, we explore the inter-dependency between objects through their co-occurrence correlation from the label set and adopt Multi-modal Factorized Bilinear (MFB) pooling component so that the image representation learning can capture this attention information. We propose a Label-Attended Hashing (LAH) algorithm which enables an end-to-end hash model with inter-dependency feature extraction. LAH first combines Convolutional Neural Network (CNN) and Graph Convolution Network (GCN) to separately generate the image representation and label co-occurrence embeddings, then adopts MFB to fuse these two modal vectors, finally learns the hash function with a Cauchy distribution based loss function via back propagation. Extensive experiments on public multi-label datasets demonstrate that (1) LAH can achieve the state-of-the-art retrieval results and (2) the usage of co-occurrence relationship and MFB not only promotes the precision of hash codes but also accelerates the hash learning. GitHub address: https://github.com/IDSM-AI/LAH. Yanzhao Xie, Yu Liu 0040, Yangtao Wang, Lianli Gao, Peng Wang 0037, Ke Zhou 0001 |
IJCAI | 1 |