Kezhen Xie

dblp:274/1040 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0008-7964-9393ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Polysemic Semantic Instance Network for Cross-Modal Hashing
abstract
Hashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performance. Most existing methods rely on one-to-one mappings that preserve only global average semantics, which fail to capture the intrinsic polysemous structures embedded within individual samples. To address this issue, we propose a novel Deep Polysemic Semantic Instance Hashing (DPSIH) method and design a Diverse Semantic Instance Embedding (DSIE) module. This module integrates local and global features through multi-head self-attention and residual learning, generating multiple diverse embeddings per sample to effectively capture fine-grained and polysemous semantic structures. Furthermore, we design a multi-embedding semantic correlation constraint that relaxes strict alignment restrictions to improve robustness under partial alignment, and introduce Maximum Mean Discrepancy (MMD) regularization to alleviate cross-modal distribution shifts. Additionally, an embedding diversity mechanism is proposed to prevent all embeddings from collapsing into a central or averaged representation, thereby enhancing semantic diversity. Extensive experiments on four benchmark datasets demonstrate that DPSIH significantly outperforms state-of-the-art methods and effectively improves the modeling of semantic ambiguity in cross-modal retrieval tasks.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Lei Huang 0010
AAAI3
2024 Visual-Textual Cross-Modal Interaction Network for Radiology Report Generation
abstract
The radiology report generation task generates diagnostic descriptions from radiology images, aiming to alleviate the onerous task for radiologists and alerting them to abnormalities. However, the data bias problem poses a persistent challenge, since the abnormal regions usually occupy a small portion of radiology image, while the report generation process should pay greater attention to the abnormal regions. Moreover, the data volume is relatively small compared to large language models, posing challenges during training. To address these issues effectively, we propose a Visual-textual Cross-model Interaction Network (VCIN) to enhance the quality of generated reports. VCIN comprises two key modules: Abundant Clinical Information Embedding (ACIE), which gathers rich cross-modal interaction information to promote the report generation of abnormal regions; and a Bert-based Decoder-only Generator (BDG), built on Bert architecture to mitigate training difficulties. The superior performance of our proposed model is demonstrated through experimental results obtained from two public benchmark datasets. The code is available athttps://github.com/QinLab-WFU/VCIN.
Wenfeng Zhang, Baoning Cai, Jianming Hu, Qibing Qin, Kezhen Xie
IEEE Signal Process. Lett.5
2024 Deep Neighborhood Structure-Preserving Hashing for Large-Scale Image Retrieval
abstract
Deep hashing integrates the advantages of deep learning and hashing technology, and has become the mainstream of the large-scale image retrieval field. However, when training the deep hashing models, most of the existing approaches regard the similarity margin of image pairs as a constant. Once similarity distance exceeds the fixed margin, the network will not learn anything, which easily results in model collapses. In this paper, we address this dilemma with a novel unified deep hashing framework, termed Deep Neighborhood Structure-preserving Hashing (DNSH), to generate the similarity-preserving and discriminative hash codes. Specifically, by extracting the discriminative object characteristics with large variances, we design an adaptive margin quadruplet loss to further explore the underlying similarity relationship between image pairs, reflecting the correct semantic structure among its neighbors. Based on the quadruple form, we develop a quadruple regularization to decrease quantization errors between binary-like embedding and hashing codes. Furthermore, through learning bit balance and bit independent terms jointly, we present the binary code constraint loss to alleviate redundancy in different bits. Extensive evaluations on four popular benchmark datasets demonstrate that our proposed deep hashing framework achieves an excellent performance than the comparison methods.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Chengduan Wang, Lei Huang 0010
IEEE Trans. Multim.2
2023 Deep Adaptive Quadruplet Hashing With Probability Sampling for Large-Scale Image Retrieval
abstract
With the preferable efficiency in storage and computation, hashing has shown potential application in large-scale multimedia retrieval. Compared with traditional hashing algorithms using hand-crafted characteristics, deep hashing inherits the representational capacity of deep neural networks to jointly learn semantic features and hash functions, encoding raw data into compact binary codes with significant discrimination. Generally, most of the current multi-wise hashing methods view the similarity margins between image pairs as constant values in training process. When the distance between sample pairs exceeds the fixed margin, the hashing network would not learn anything. Besides, available hashing methods commonly introduce the random sampling strategy to build training batches and ignore the sample distribution, which is harmful to parameter optimization. In this paper, we propose a novel Deep Adaptive Quadruplet Hashing with probability sampling (DAQH) for discriminative binary code learning. Specifically, with exploring the distribution relationship of raw samples, a non-uniform probability sampling strategy is proposed to build more informative and representative training batches, while maintaining the diversity of training samples. By introducing the prior similarity of sample pairs to calculate corresponding margins, an adaptive margin quadruplet loss is designed to dynamically preserve the underlying semantic relationships with its neighbors. To tune the attributes of binary codes, by combining quadruple regularization and orthogonality optimization, binary code constraint is developed to make the learned embedding with significant discrimination. Extensive experimental results on various benchmark datasets demonstrate our proposed DAQH framework achieves state-of-the-art visual similarity search performance.
Qibing Qin, Lei Huang 0010, Kezhen Xie, Zhiqiang Wei 0002, Chengduan Wang, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2022 Learning to Classify Weather Conditions from Single Images Without Labels
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Zhiqiang Wei 0002
MMM (1)1
2022 WCATN: Unsupervised deep learning to classify weather conditions from outdoor images
Kezhen Xie, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin
Eng. Appl. Artif. Intell.1
2022 Deep Multi-Similarity Hashing with semantic-aware preservation for multi-label image retrieval
Qibing Qin, Lintao Xian, Kezhen Xie, Wenfeng Zhang, Yu Liu 0022, Jiangyan Dai, Chengduan Wang
Expert Syst. Appl.3
2022 A CNN-based multi-task framework for weather recognition with multi-scale weather cues
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Lei Lyu 0001
Expert Syst. Appl.1
2021 Deep top similarity hashing with class-wise loss for multi-label image retrieval
Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Wenfeng Zhang
Neurocomputing4
2021 Unsupervised Deep Quadruplet Hashing with Isometric Quantization for image retrieval
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Kezhen Xie, Jinkui Hou
Inf. Sci.5
2021 Graph convolutional networks with attention for multi-label weather recognition
Kezhen Xie, Zhiqiang Wei 0002, Lei Huang 0010, Qibing Qin, Wenfeng Zhang
Neural Comput. Appl.1
2021 Unsupervised Deep Multi-Similarity Hashing With Semantic Structure for Image Retrieval
abstract
With the advance of Convolutional Neural Network, deep hashing methods have shown the great promising performance in large-scale image retrieval. Without depending on extensive human-annotated data, unsupervised hashing is more applicable to image retrieval tasks compared to supervised methods. However, due to the lack of fine-grained supervised signals and multi-similarity constraints, most state-of-the-art unsupervised deep hashing algorithms cannot ensure the correct fine-grained similarity ranking for image pairs. In this paper, we propose a novel unsupervised deep multi-similarity hashing framework to learn compact binary codes by jointly exploiting global-aware and spatial-aware representations, called Unsupervised Deep Multi-Similarity Hashing with Semantic Structure (UDMSH). Specifically, to obtain distinguishing characteristics, we develop a sub-network by jointly learning global semantic structures from Convolutional Neural Network (CNN) and inherent spatial structures from Fully Convolutional Network (FCN). By computing the cosine distance for deep features from image pairs, we construct a similarity matrix with semantic structure, then utilize this matrix to guide hash code learning process. Based on it, we carefully design a multi-level pairwise loss to preserve the correct fine-grained similarity ranking. Furthermore, we introduce Hamming-isometric mapping into unsupervised hashing framework to decrease the quantization errors. Extensive experiments on three widely used benchmarks prove that our proposed UDMSH outperforms several state-of-the-art unsupervised hashing with respect to different evaluation metrics.
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Kezhen Xie, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2020 Adaptive Attention-Aware Network for unsupervised person re-identification
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Qibing Qin
Neurocomputing4