Qibing Qin

dblp:185/5154 · DBLP profile ↗
← Back
64ranked-venue papers
16as first author
60since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 4 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-author · 27 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Deep hashing via mean centroid representation for large-scale image retrieval
Bingxin Wang, Xianmin Wei, Qibing Qin, Lei Huang 0010
Expert Syst. Appl.4
2027 Wavelet-enhanced Mamba with multi-domain feature learning for image inpainting
Zikai Wu, Jiangyan Dai, Qibing Qin, Huihui Zhang 0003, Yugen Yi
Expert Syst. Appl.3
2026 Polysemic Semantic Instance Network for Cross-Modal Hashing
abstract
Hashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performance. Most existing methods rely on one-to-one mappings that preserve only global average semantics, which fail to capture the intrinsic polysemous structures embedded within individual samples. To address this issue, we propose a novel Deep Polysemic Semantic Instance Hashing (DPSIH) method and design a Diverse Semantic Instance Embedding (DSIE) module. This module integrates local and global features through multi-head self-attention and residual learning, generating multiple diverse embeddings per sample to effectively capture fine-grained and polysemous semantic structures. Furthermore, we design a multi-embedding semantic correlation constraint that relaxes strict alignment restrictions to improve robustness under partial alignment, and introduce Maximum Mean Discrepancy (MMD) regularization to alleviate cross-modal distribution shifts. Additionally, an embedding diversity mechanism is proposed to prevent all embeddings from collapsing into a central or averaged representation, thereby enhancing semantic diversity. Extensive experiments on four benchmark datasets demonstrate that DPSIH significantly outperforms state-of-the-art methods and effectively improves the modeling of semantic ambiguity in cross-modal retrieval tasks.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Lei Huang 0010
AAAI2
2026 Deep Potential Semantic-aware Hashing for Cross-modal Retrieval
abstract
Hashing learning has moved into the mainstream for multimedia retrieval because it offers the advantages of low storage cost and high retrieval efficiency. Currently, most cross-modal hashing methods commonly explore the similarity relations between samples by constructing pair-wise or triplet-wise constraints. However, these methods focus on the relative correct ranking of samples, ignore the potential semantic similarity of raw sample distribution, and generate sub-optimal hash codes. To resolve this issue, the novel Deep Potential Semantic-aware Hashing framework (DPSaH) is proposed to mine the local semantic structure of heterogeneous samples, maintaining inter-modality-consistent and cross-modality-correlated semantic relationships. Specifically, by exploring the potential local structure of the data, the multi-modal quadruple loss is extended to the cross-modal hashing framework, thereby preserving the potential semantic neighborhoods among raw samples in Hamming space. During model training, based on the average semantic labels, the label-averaged balanced strategy is developed to quantify the frequency difference between positive and negative samples. Besides, by injecting noise information into the generated discrete codes, the binary-injection loss is introduced to alleviate the over-activation of specific bits, decorrelating different bits in the Hamming space. Extensive experiments are performed on three public datasets, and the results verify the superiority of the DPSaH framework compared to the current mainstream cross-modal hashing frameworks. The source code for DPSaH is available at https://github.com/QinLab-WFU/DPSaH .
Qibing Qin, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Eng. Appl. Artif. Intell.2
2026 Deep asymmetric semantic hashing with probability shifting for multi-label image retrieval
Yongyue Fu, Qibing Qin, Jinkui Hou, Congcong Zhu, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.2
2026 Deep semantic center-guided hashing for multi-label cross-modal retrieval
Xinzheng Sui, Yadong Huo, Qibing Qin, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.4
2026 Deep global-ranking hashing via average precision approximation for large-scale image retrieval
Yongyue Fu, Qibing Qin, Lei Huang 0010, Wenfang Zhang
Expert Syst. Appl.3
2026 Deep attribute-aware hashing for zero-shot image retrieval
Yongyue Fu, Qibing Qin, Wenfang Zhang, Lei Huang 0010
Expert Syst. Appl.3
2026 Deep neighborhood-based component proxy hashing for large-scale image retrieval
Huiying Zhu, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.3
2026 Deep synthetic-proxy hashing for multi-label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Chunlei Chen, Lei Huang 0010
Neurocomputing2
2026 Hierarchical texture-aware image inpainting via contextual attention and multi-scale fusion
Runing Li, Jiangyan Dai, Qibing Qin, Chengduan Wang, Yugen Yi
Image Vis. Comput.3
2026 Few-shot medical image segmentation via dual-stream feature extractor and detail-enhanced prototype transformer
Wenfeng Zhang, Jianming Hu, Qibing Qin
Knowl. Based Syst.6
2026 Deep Softtriple hashing for Multi-Label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Neural Networks2
2026 Deep neighbor-aware hashing with global-local representation for multi-label remote sensing image retrieval
Xiaorong Chen, Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Signal Process. Image Commun.2
2026 Deep noise-tolerant hashing for remote sensing image retrieval
abstract
Currently, how to quickly retrieve target images from large-scale remote sensing data has emerged as a critical challenge in the context of explosive growth of remote sensing data volume. To deal with this challenge, hash learning becomes an ideal choice with its low storage cost and high efficiency. In recent years, the combination of hash learning with deep neural networks such as CNNs and Transformers has resulted in numerous frameworks demonstrating excellent performance. However, in the field of remote sensing image hashing, previous studies cannot simultaneously consider the effect of noise in feature extraction and loss optimization, so that their retrieval performance is greatly reduced due to noise interference. To resolve the mentioned problem, a Deep Noise-tolerant Hashing (DNtH) framework is proposed to learn the sample complexity and noise level, and adaptively reduce the weight of noisy information. Specifically, to realize the extraction of fine-grained features from information containing irrelevant samples, the noise-aware Transformer is proposed by introducing the patch-wise attention and depth-wise convolution. To reduce the interference of noisy labels on remote sensing image retrieval, an adaptive active-passive loss framework is proposed to dynamically adjust the weights of active passive loss, which learns the weight parameters through a dynamic weighted network while combining with asymmetric strategy for effective compact representation learning. The ratio of entropy to standard deviation and the probability difference are input into the above network and trained with the feature extraction network. Extensive experiments on three publicly available datasets show that the DNtH framework can adapt to noisy environments while achieving optimal performance in remote sensing image retrieval. The source code for the implementation of our DNtH framework is available at https://github.com/QinLab-WFU/DNtH.git .
Chunyu Yan, Qibing Qin, Jiangyan Dai, Wenfeng Zhang
Signal Process. Image Commun.3
2026 Deep semantic channel hashing for large-scale image retrieval
Qibing Qin, Lei Huang 0010, Wenfeng Zhang
Signal Process. Image Commun.3
2026 Deep Stochastic Spherical Hashing With Von Mises-Fisher Distributions for Cross-Modal Retrieval
abstract
Deep cross-modal hashing has gained significant attention because of its benefits, including reduced storage requirements and enhanced retrieval efficiency. Although progress has been made, existing deep cross-modal hashing methods still face unresolved challenges. Most existing methods typically adopt Euclidean space as the embedding space to measure the semantic similarity between original samples. However, the volume of Euclidean space grows polynomially with dimension, which exacerbates the curse of dimensionality. In contrast, methods based on spherical space usually use cosine similarity as the metric, effectively mitigating the aforementioned problem by normalizing the embedding vectors. Nevertheless, such methods only considers the direction to determine the category, ignoring the uncertainty measure in the embedding space, thus having a limited ability to preserve inherent multimodal semantics. In this paper, with a novel extension of the maximum entropy distribution on the surface of a hypersphere von Mises-Fisher (vMF) distribution, a novel deep cross-modal hashing method, named Deep Stochastic Spherical Hashing (DSSH), is designed to utilize uncertain information to guide the hashing process and produce discriminative modality-invariant hash codes. Specifically, to learn explicit uncertainty in learned embedding space, the Spherical von Mises-Fisher distribution is applied for the f irst time in deep cross-modal hashing, where the direction of the sample embedding controls its position on the hyper sphere, thereby preventing its semantic content, and its norm parameterizes the determinism of the distribution. In addition, stochastic spherical von Mises–Fisher loss is proposed to preserve the mode-specific semantic information of the sample, achieving the alignment of different modalities and semantic embeddings. Extensive experiments on four benchmark datasets show that our DSSH framework outperforms existing state-of-the-art cross modal hashing methods. The source code of the experiments is available at https://github.com/QinLab-WFU/DSSH.
Qibing Qin, Meiling Ge, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Knowl. Data Eng.1
2026 Deep Distance Weighted Sampling Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing seeks to encode heterogeneous image-text data into compact binary codes for efficient retrieval. While significant efforts have been made to study sampling strategies, most of these approaches are tightly integrated with loss function engineering, lacking an independent focus on sampling methods. In this paper, we challenge the convention by revealing that batch-level sampling strategy is equally pivotal as loss design for learning discriminative hash codes. Specifically, by introducing a novel distribution-aware sampling strategy, a Distance Weighted Sampling Hashing (DDWSH) framework is proposed to dynamically select stable and informative training pairs. Unlike conventional random or semi-hard sampling, our method weights pairwise distances within each batch to approximate global data distribution, thereby mitigating training instability caused by biased sampling. To rigorously validate our claims, we conduct the first comprehensive crossover study between sampling strategies (random/semi-hard/ours) and loss functions (contrastive/triplet/ours) across three benchmark datasets. Experiments demonstrate that: Universality: Our sampling boosts all loss functions' performance (average +4.3% mAP vs. semi-hard mining), Superiority: DDWSH competes with complex loss function design and achieves state-of-the-art results, and Stability: It reduces performance variance by 60.2% compared to semi-hard sampling under varying batch compositions average. This systematic analysis establishes sampling as an independent research dimension in deep hashing, beyond a mere part of loss function engineering. The source code for DDWSH is freely available athttps://github.com/QinLab-WFU/DDWSH.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.2
2026 Deep Neighbor Discriminant Binary Embedding for Multi-Label Image Retrieval
abstract
Because of fast retrieval speed and low storage cost, deep hash learning has become a current research spot in multimedia retrieval, which arouses more and more interest and attention. Most of the available deep hashing methods usually employ the mini-batch strategy to construct training batches, and only one mini-batch of samples is valuable at each iteration during model optimization, which fails to explore the original neighbor structure well, leading to sub-optimal embeddings, especially for relatively large datasets. By contrast, the superior performance is achieved by optimizing the SoftMax loss for certain hashing learning, meanwhile, prior research has suggested the normalized SoftMax loss is essentially identical to one smoothed triplet-wise function, in which each class is assigned to one single semantic center. Nevertheless, under real-world scenarios, one class could correspond to multiple local clusters rather than just a single one. To this end, by extending SoftMax loss with multiple semantic centers, a novel deep hashing framework, called Deep Neighbor Discriminant Binary Embedding (NDBE) framework, is presented to generate discriminative hash codes with the original neighbor structure preservation. Specifically, by expanding the size of the last fully connected layer to increase multiple centers for each class, a novel SoftTriple loss is proposed to capture the latent semantic distribution of the original samples and reduce the intra-class variance, which is optimized without constraints sampling. To learn the different numbers of each class center, the class-center aggregation strategy is developed to obtain the compact set of centers while maintaining semantic similarity. Extensive experiments on three public multi-label datasets show that our proposed NDBE framework achieves superior visual similarity search performance over several state-of-the-art approaches. The source code for the implementation of our proposed NDBE framework is available athttps://github.com/QinLab-WFU/NDBE.
Qibing Qin, Mingkun Dou, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.1
2026 Deep Semantic Tuplet-Based Hashing by Hypergraph Modeling for Cross-Modal Retrieval
abstract
With low storage cost and high retrieval efficiency, hashing techniques are widely used for multi-media retrieval, which has already become the present research focus. Currently, cross-modal hashing commonly employs graph-based loss to construct pair-wise semantic relations between training samples for model optimization. However, limited by the graph-based strategy, each edge in the graph only connects two samples, which only represent a bundle of pair-wise relationships. Besides, the edges in the graph are calculated by self-attention or feature distance, only considering pair-wise relations of heterogeneous samples and ignoring the class relations. In this paper, by hypergraph modeling the semantic tuples, a novel Deep Semantic Tuplet-based Hashing by Hypergraph Modeling (DSTH) approach is proposed to leverage the multilateral semantic relations, which could guide the model to learn class-discriminative semantic binary embedding. In more detail, based on the characteristic distribution, semantic tuples are constructed for each class in one mini-batch, which represents the multilateral semantic relationships between multiple samples and multiple classes. By considering semantic tuples as hyperedges to represent multilateral semantic relations, hypergraph modeling is designed, in which HyperGraph Neural Hetwork (HGNH) is introduced to formulate hypergraph node classification goals to fully learn the multilateral semantic information contained in the semantic tuples. Moreover, to utilize the heterogeneity of local structures in embedding, the adaptive neighborhood structure is explored by learning the structure embedding, which provides fine-grained ranking lists. Through extensive experiments on three benchmark datasets, the comprehensive results validate the advancement of our proposed DSTH framework over mainstream cross-modal hashing. The source code for the framework DSTH is freely available athttps://github.com/QinLab-WFU/DSTH.
Qibing Qin, Wenfeng Zhang, Huihui Zhang 0003, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.1
2026 Generative Zero-Shot Hashing for Multi-Label Image Retrieval
abstract
Due to the preferable efficiency of storage and computation, hashing algorithms show great potential. With the explosive growth of web data, new and emerging categories of multimedia data continue to increase, and zero-shot hashing has been one of the hot research topics in retrieval research. Nevertheless, existing zero-shot hashing methods focus mainly on single-label image retrieval, using class semantic embeddings as a link between visual features and binary codes to align visual features with corresponding class semantics and simultaneously transfer knowledge from seen classes to unseen classes. Meanwhile, single-label zero-shot learning employs Generative Adversarial Networks (GANs) to synthesize class-specific characteristics from the corresponding class attribute embeddings, achieving encouraging results. However, synthesizing multi-label features from GANs remains unexplored in the context of zero-shot hashing settings. When multiple objects co-appear in an image, a key issue is how to fuse multi-class information effectively. In this article, by introducing the multi-class information fusion strategy, a novel Generative Zero-Shot Hashing (GZSH) is proposed to transform zero-shot hashing into traditional supervised hashing by generating features of unseen categories. Specifically, this study introduces three different fusion methods (attribute-based fusion, feature-based fusion, and semantic-enhanced fusion) for synthesizing multi-label features from the corresponding multi-label class embeddings. Subsequently, the semantic-enhanced fusion is integrated into the representative generative architecture to learn the feature distribution of unlabeled images in the context of multi-label zero-shot hashing. Besides, by jointly learning the seen/source and unseen/target samples, a pairwise similarity loss is introduced to optimize the deep hashing model. Extensive experimentation on three widely used multi-label datasets illustrates the outstanding performance of our proposed GZSH framework compared to current state-of-the-art methods. The implementation code for our GZSH framework can be found at https://github.com/QinLab-WFU/GZSH .
Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Deep Uncertainty-aware Probabilistic Hashing for Cross-modal Retrieval
abstract
Due to its outstanding computational efficiency and low storage requirements, hashing technology has become a research hotspot in large-scale multimedia retrieval. In cross-modal hashing, the key lies in mapping samples from different data sources into a common discrete space. However, most existing methods assume that the input data is of high quality and completeness. When samples are incomplete or degraded (e.g., blurry images or incomplete text), the absence or ambiguity of semantic information inevitably compromises retrieval accuracy. Traditional deterministic embedding methods typically map multi-modal samples to a single point in the embedding space without considering uncertainty. As a result, inherent noise or feature ambiguity in the inputs may lead to distorted or shifted binary representations. To resolve this problem, this article proposes a novel Deep Uncertainty-aware Probabilistic Hashing (DUaPH) method that models the uncertainty of multi-modal samples. By capturing the underlying distribution of heterogeneous data, DUaPH effectively mitigates feature ambiguity and inherent noise, enhancing the robustness of cross-modal retrieval. Specifically, each heterogeneous sample is mapped to a multivariate Gaussian distribution, where the mean represents the most probable semantic features, and the variance reflects the sample uncertainty. A semantic feature matching mechanism is introduced to dynamically adjust the importance of feature dimensions, prioritizing those with higher certainty. Then, a semantic feature fusion mechanism is developed to integrate the semantic features from multi-modal sample pairs, producing a new distribution with reduced uncertainty and improved semantic alignment. Extensive experiments on four benchmark datasets demonstrate that DUaPH significantly improves robustness and retrieval performance under conditions of semantic ambiguity and data uncertainty. The source code is available at https://github.com/QinLab-WFU/DUaPH .
Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Deep Relational Knowledge Distillation Hashing via Relaxed Masking Triplet Optimization for Large-scale Image Retrieval
abstract
Large-scale image retrieval increasingly depends on triplet-based deep hashing methods, which typically construct training triplets using class labels. However, existing triplet-based hashing methods relying on sparse and noisy categorical labels suffer from two main issues: no triplet, where sparse annotations limit valid triplet formation, and bad triplet, where noisy or ambiguous labels cause misleading samples, which together hinder the network from learning a well-structured similarity space, causing semantically similar images to scatter and dissimilar ones to collapse, ultimately undermining retrieval accuracy. In this article, we solve this dilemma with a novel unified deep hashing framework, termed Deep Relational Knowledge Distillation Hashing (DRKDH), which decouples semantic relation modeling from visual feature learning by leveraging a teacher-student paradigm to learn semantically consistent hash codes. Specifically, the teacher model captures comprehensive semantic structural relationships across all instances, providing information-rich and noise-resilient supervision to guide the student model in learning expressive and discriminative visual embeddings, effectively mitigating the impact of low-quality triplet labels. Furthermore, by leveraging similarity weights predicted by the teacher model, a relaxed masking triplet loss is introduced into the student model, which dynamically adjusts the contribution of each triplet based on its informativeness, suppressing invalid or misleading triplets while emphasizing valuable ones to enhance training efficiency. Comprehensive experiments on THINGS, ImageNet, MIRFLICKR-25K, and NUS-WIDE show that DRKDH consistently outperforms state-of-the-art deep hashing methods under sparse and noisy supervision, while additional evaluations suggest competitive performance in selected zero-shot and few-shot settings. Source codes: https://github.com/QinLab-WFU/DRKDH .
Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Enhanced radiology report generation via comprehensive sequence rearrangement and multi-scale cross-region attention
Qibing Qin, Jianming Hu, Dengwei Yan, Wenfeng Zhang, Jing Qiao
Vis. Comput.2
2025 Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal Retrieval
abstract
The field of cross-modal retrieval aims to construct a shared representation space for samples from multiple modalities, typically within the vision and language domains. Deep hashing, with its high computational efficiency and low storage costs, has emerged as a central focus in this field and has garnered significant attention in recent research. However, current hash retrieval, concentrating on deterministic methods, struggles to effectively capture semantically ambiguous correspondences between cross-modal samples, where heterogeneous data have complex-semantic many-to-many relationships in the latent space. To address this limitation, we propose a novel Deep Probabilistic Binary Embedding (DPBE) framework, designed to generate discriminative, modality-invariant hash codes that facilitate accurate and reliable cross-modal retrieval. In contrast to contemporary probabilistic methods, we focus on optimizing hash networks to learn more accurate binary embeddings by using the learning mode of probabilistic embeddings. We introduce the first Bayesian encoder for hash learning, which employs Laplace Approximation to model a distribution over network weights. Extensive experimental results demonstrate that our approach not only outperforms deterministic methods in retrieval performance but also provides uncertainty estimates, enhancing the interpretability of the embeddings. The corresponding code is available at https://github.com/QinLab-WFU/DPBE.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
ACM Multimedia2
2025 Factorized Transformer Hashing with Adaptive Routing for Large-scale Image Retrieval
abstract
Transformer architecture has driven significant advancements in deep hashing, establishing itself as a dominant framework for large-scale retrieval and storage applications. However, existing Transformer-based deep hashing methods typically employ unvarying feature transformations across all images, limiting their adaptability to diverse visual patterns. This rigidity restricts the model's capacity to learn both highly distinctive and generalizable discrete representations, posing challenges for retrieval in open-world scenarios. To overcome this challenge, we propose a novel Factorized Transformer Hashing (FTH) framework, which introduces a factorized transformer to enhance the generalization and discriminative power of hash codes. Specifically, we decompose the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) blocks into multiple sub-blocks, forming a transformer factorization scheme that captures diverse feature characteristics through independent sub-blocks. Furthermore, we develop an adaptive selection strategy, leveraging a set of learnable selectors with the Softmax function, to dynamically route each image to the most appropriate sub-block for processing. Extensive experiments on three benchmark datasets demonstrate that the proposed FTH framework significantly outperforms state-of-the-art baselines in both image hashing and zero-shot hashing tasks. Source code is available at https://github.com/QinLab-WFU/FTH.
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
ACM Multimedia2
2025 Ranking-oriented cross-modal hashing
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Eng. Appl. Artif. Intell.2
2025 Deep informative-triplet sampling hashing with attention-aware augmentation for remote sensing image retrieval
Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.2
2025 Deep adaptive gradient-triplet hashing for cross-modal retrieval
Congcong Zhu, Jinkui Hou, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.4
2025 Deep neighbor-coherence hashing with discriminative sample mining for supervised cross-modal retrieval
Congcong Zhu, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.2
2025 Personalized Dual Transformer Network for sequential recommendation
Meiling Ge, Chengduan Wang, Xueyang Qin, Jiangyan Dai, Lei Huang 0010, Qibing Qin, Wenfeng Zhang
Neurocomputing6
2025 Deep Consistent Penalizing Hashing with noise-robust representation for large-scale image retrieval
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
Neurocomputing1
2025 Deep multi-similarity hashing via label-guided network for cross-modal retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Neurocomputing2
2025 Deep binary hyperbolic embedding for large-scale image retrieval
Enhao Wang, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Neurocomputing3
2025 Domain-invariant multi-granularity feature learning for generalizable person re-identification
Wenfeng Zhang, Xiangfei Cao, Lei Huang 0010, Dengwei Yan, Qibing Qin
Knowl. Based Syst.5
2025 Deep Hardness-Aware Hashing for Large-Scale Image Retrieval
Chunping Dong, Meiling Ge, Enhao Wang, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
IEEE Signal Process. Lett.4
2025 Deep Discriminative Boundary Hashing for Cross-Modal Retrieval
abstract
By the preferable efficiency in storage and computation, deep cross-modal has gained much attention in large-scale multimedia retrieval. Current deep hashing employs the probability outputs of the likelihood function, i.e., Sigmoid or Cauchy, to quantify the semantic similarity between samples in a common Hamming space. However, the inherent weakness of the Sigmoid likelihood function or the Cauchy likelihood function in gradient optimization leads to hashing models failing to exactly describe the hamming ball, which indicates the absolute semantic boundary among classes, thereby giving the high neighborhood ambiguity. In this paper, with the analysis of the likelihood function from the perspective of similarity metric learning, the novel Deep Discriminative Boundary Hashing framework (DDBH) is proposed to learn the discriminative embedding space that separates neighbors and non-neighbors well. Specifically, by introducing the remapping strategy and the base-point adaptive selection, the boundary-preserving loss based on the adjustable likelihood function is proposed to project data points with small gradients to regions with large gradients and give larger gradients for hard samples, facilitating better separation among classes. Meanwhile, to learn class-dependent binary codes, the class-wise quantization loss is designed to heuristically transfer class-wise prior knowledge to the binary quantization, significantly improving the discriminative capability of compact discrete codes. Comprehensive experiments on three benchmark datasets show that our proposed DDBH framework outperforms other representative deep cross-modal hashing. The corresponding code is available at https://github.com/QinLab-WFU/DDBH.
Qibing Qin, Yadong Huo, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Circuits Syst. Video Technol.1
2025 Deep Semantic-Consistent Penalizing Hashing for Cross-Modal Retrieval
abstract
Benefiting from the advantages of low storage cost and high retrieval efficiency, hash learning could significantly speed up large-scale cross-modal retrieval. Based on the prior annotations, most of the available cross-modal hashing usually introduces the margin-based constraint to generate different boundaries for each class in the inference phase, optimizing the model. However, these obtained label-guided penalty boundaries may differ from the primitive semantic relationships between heterogeneous modalities, impairing retrieval performance. Besides, the margin-based constraint is too weak to penalize the classes with low intra-class variances or inter-class correlations, which struggle to learn high-quality embeddings. In this paper, we propose a novel Deep Semantic-consistent Penalizing Hashing framework (DScPH) to learn the consistent penalizing fields for all classes, achieving accurate and efficient cross-modal retrieval. Specifically, by exploring unbalanced intra-class and inter-class correlations, the consistent penalizing loss is introduced into cross-modal retrieval to learn the consistency decision boundaries across classes. During training, the dice-like optimization strategy is developed to balance the pulling penalizing elements and pushing penalizing elements, facilitating the model convergence. Besides, based on the invariance of similarity measures under orthogonal transformations, the alternative quantization is proposed to minimize the errors between the learned continuous embeddings and binary discretization, maintaining the consistency of semantic relationships after performing binary projection. Extensive experiments are conducted on three benchmark datasets, and the comprehensive results validate the efficacy of our proposed DScPH framework, which outperforms the current mainstream deep cross-modal hashing algorithms.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.1
2025 Texture and Structure-Guided Dual-Attention Mechanism for Image Inpainting
abstract
Deep learning exhibits powerful capability in image inpainting task, particularly in generating pixel-level details closely with the human visual perception. However, the complex background or larger missing regions make it still encounters the artifacts. Many researchers have investigated that prior information is crucial for guiding the image inpainting. In this article, we introduce the dual-attention mechanism, including lightweight spatial attention and linearized attention, to construct an end-to-end texture and structure-guided image inpainting method. In the first stage, we build the detail inpainting network with the lightweight spatial attention. In this model, the extracted texture and structural features are fused with multi-layers and then the fused detail image is considered as the prior to guide the detail repair of corrupted images. In the second stage, we construct the content completing network by the repaired detail and the linearized Transformer module. This module not only overcomes the limitation of the receptive field size of convolutional kernels that can improve the long-range modeling of features but also can significantly reduce the computational complexity of the original Transformer. To demonstrate the superior effectiveness of the proposed method, we perform extensive experiments with advanced models on three datasets: CelebA-HQ, Places2, and Paris Street Views. Comparative results manifest that our method achieves excellent image inpainting results that are conform to the human visual system. The code is available at https://github.com/QinLab-WFU/TSGDAM
Runing Li, Jiangyan Dai, Qibing Qin, Chengduan Wang, Huihui Zhang 0003, Yugen Yi
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Modality-aware graph CNN for cross-modal person reidentification
Ruisheng Ran, Wenfeng Zhang, Qibing Qin
Vis. Comput.5
2024 MeFD-Net: multi-expert fusion diagnostic network for generating radiology image reports
Ruisheng Ran, Renjie Pan 0002, Wenfeng Zhang, Qibing Qin
Appl. Intell.7
2024 Deep global semantic structure-preserving hashing via corrective triplet loss for remote sensing image retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.2
2024 Visual-Textual Cross-Modal Interaction Network for Radiology Report Generation
abstract
The radiology report generation task generates diagnostic descriptions from radiology images, aiming to alleviate the onerous task for radiologists and alerting them to abnormalities. However, the data bias problem poses a persistent challenge, since the abnormal regions usually occupy a small portion of radiology image, while the report generation process should pay greater attention to the abnormal regions. Moreover, the data volume is relatively small compared to large language models, posing challenges during training. To address these issues effectively, we propose a Visual-textual Cross-model Interaction Network (VCIN) to enhance the quality of generated reports. VCIN comprises two key modules: Abundant Clinical Information Embedding (ACIE), which gathers rich cross-modal interaction information to promote the report generation of abnormal regions; and a Bert-based Decoder-only Generator (BDG), built on Bert architecture to mitigate training difficulties. The superior performance of our proposed model is demonstrated through experimental results obtained from two public benchmark datasets. The code is available athttps://github.com/QinLab-WFU/VCIN.
Wenfeng Zhang, Baoning Cai, Jianming Hu, Qibing Qin, Kezhen Xie
IEEE Signal Process. Lett.4
2024 Deep Semantic-Aware Proxy Hashing for Multi-Label Cross-Modal Retrieval
abstract
Deep hashing has attracted broad interest in cross-modal retrieval because of its low cost and efficient retrieval benefits. To capture the semantic information of raw samples and alleviate the semantic gap, supervised cross-modal hashing methods that utilize label information which could map raw samples from different modalities into a unified common space, are proposed. Although making great progress, existing deep cross-modal hashing methods are suffering from some problems, such as: 1) considering multi-label cross-modal retrieval, proxy-based methods ignore the data-to-data relations and fail to explore the combination of the different categories profoundly, which could lead to some samples without common categories being embedded in the vicinity; 2) for feature representation, image feature extractors containing multiple convolutional layers cannot fully obtain global information of images, which results in the generation of sub-optimal binary hash codes. In this paper, by extending the proxy-based mechanism to multi-label cross-modal retrieval, we propose a novel Deep Semantic-aware Proxy Hashing (DSPH) framework, which could embed multi-modal multi-label data into a uniform discrete space and capture fine-grained semantic relations between raw samples. Specifically, by learning multi-modal multi-label proxy terms and multi-modal irrelevant terms jointly, the semantic-aware proxy loss is designed to capture multi-label correlations and preserve the correct fine-grained similarity ranking among samples, alleviating inter-modal semantic gaps. In addition, for feature representation, two transformer encoders are proposed as backbone networks for images and text, respectively, in which the image transformer encoder is introduced to obtain global information of the input image by modeling long-range visual dependencies. We have conducted extensive experiments on three baseline multi-label datasets, and the experimental results show that our DSPH framework achieves better performance than state-of-the-art cross-modal hashing methods. The code for the implementation of our DSPH framework is available athttps://github.com/QinLab-WFU/DSPH.
Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang
IEEE Trans. Circuits Syst. Video Technol.2
2024 S3-Net: A Self-Supervised Dual-Stream Network for Radiology Report Generation
abstract
Intelligent medicine is eager to automatically generate radiology reports to ease the tedious work of radiologists. Previous researches mainly focused on the text generation with encoder-decoder structure, while CNN networks for visual features ignored the long-range dependencies correlated with textual information. Besides, few studies exploit cross-modal mappings to promote radiology report generation. To alleviate the above problems, we propose a novel end-to-end radiology report generation model dubbed Self-Supervised dual-Stream Network (S3-Net). Specifically, a Dual-Stream Visual Feature Extractor (DSVFE) composed of ResNet and SwinTransformer is proposed to capture more abundant and effective visual features, where the former focuses on local response and the latter explores long-range dependencies. Then, we introduced the Fusion Alignment Module (FAM) to fuse the dual-stream visual features and facilitate alignment between visual features and text features. Furthermore, the Self-Supervised Learning with Mask(SSLM) is introduced to further enhance the visual feature representation ability. Experimental results on two mainstream radiology reporting datasets (IU X-ray and MIMIC-CXR) show that our proposed approach outperforms previous models in terms of language generation metrics.
Renjie Pan 0002, Ruisheng Ran, Wenfeng Zhang, Qibing Qin, Shaoguo Cui
IEEE J. Biomed. Health Informatics5
2024 Deep Hierarchy-Aware Proxy Hashing With Self-Paced Learning for Cross-Modal Retrieval
abstract
Due to its low storage cost and high retrieval efficiency, hashing technology is popularly applied in both academia and industry, which provides an interesting solution for cross-modal similarity retrieval. However, most existing supervised cross-modal hashing methods typically view the fixed-level semantic affinity defined by manual labels as supervised signals to guide hash learning, which only represents a small subset of complex semantic relations between multi-modal samples, thus impeding the hash function learning and degrading the obtained hash codes. In the paper, by learning shared hierarchy proxies, a novel deep cross-modal hashing framework, called Deep Hierarchy-aware Proxy Hashing (DHaPH), is proposed to construct the semantic hierarchy in a data-driven manner, thereby capturing the accurate fine-grained semantic relationships and achieving small intra-class scatter and big inter-class scatter. Specifically, by regarding the hierarchical proxies as learnable ancestors, a novel hierarchy-aware proxy loss is designed to model the latent semantic hierarchical structures from different modalities without prior hierarchy knowledge, in which similar samples share the same Lowest Common Ancestor (LCA) and dissimilar points have different LCA. Meanwhile, to adequately capture valuable semantic information from hard pairs, a multi-modal self-paced loss is introduced into cross-modal hashing to reweight multi-modal pairs dynamically, which enables the model to gradually focus on hard pairs while simultaneously learning universal patterns from multi-modal pairs. Extensive experiments on three available benchmark databases demonstrate that our proposed DHaPH framework outperforms the compared baselines with different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DHaPH.
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Knowl. Data Eng.2
2024 Deep Neighborhood-Preserving Hashing With Quadratic Spherical Mutual Information for Cross-Modal Retrieval
abstract
Driven by the high nonlinearity of deep neural networks, deep hashing has achieved the pictured great potential in cross-modal retrieval applications, significantly bridging the modality gap. Current deep cross-modal hashing usually utilizes affinity matching or local ranking to capture the local semantic relationships in the learned common space, leading to high neighborhood ambiguity. Simultaneously, most of these frameworks utilize additional regularization terms or margin thresholds to enhance the overall performance, in which searching the model's hyper-parameters under mass training data would have a substantial overhead. In this paper, with a novel extension of information-theoretic measures, a novel deep cross-modal hashing method, named Deep Neighborhood-preserving Hashing (DNpH), is designed to learn a highly separable discrete space, effectively mitigating the semantic gap across different modalities. Specifically, to minimize neighborhood ambiguity, the Quadratic Spherical Mutual Information (QSMI) is first introduced into deep cross-modal hashing to separate neighbors and non-neighbors well, while it is free of tuning parameters during model training compared with other similarity measures. To optimize quadratic mutual information loss smoothly, a square clamping method is developed to improve the stability of model optimization, avoiding converging on bad local optimum. Besides, two transformer encoders are exploited as feature extractors for multi-modal samples to learn the informative semantic representations. Finally, we compare our proposed DNpH framework with various state-of-the-art cross-modal hashing on four public datasets, and large amounts of experiment results demonstrate our contributions and show that DNpH outperforms the compared baselines on different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DNpH.
Qibing Qin, Yadong Huo, Lei Huang 0010, Jiangyan Dai, Huihui Zhang 0003, Wenfeng Zhang
IEEE Trans. Multim.1
2024 Deep Neighborhood Structure-Preserving Hashing for Large-Scale Image Retrieval
abstract
Deep hashing integrates the advantages of deep learning and hashing technology, and has become the mainstream of the large-scale image retrieval field. However, when training the deep hashing models, most of the existing approaches regard the similarity margin of image pairs as a constant. Once similarity distance exceeds the fixed margin, the network will not learn anything, which easily results in model collapses. In this paper, we address this dilemma with a novel unified deep hashing framework, termed Deep Neighborhood Structure-preserving Hashing (DNSH), to generate the similarity-preserving and discriminative hash codes. Specifically, by extracting the discriminative object characteristics with large variances, we design an adaptive margin quadruplet loss to further explore the underlying similarity relationship between image pairs, reflecting the correct semantic structure among its neighbors. Based on the quadruple form, we develop a quadruple regularization to decrease quantization errors between binary-like embedding and hashing codes. Furthermore, through learning bit balance and bit independent terms jointly, we present the binary code constraint loss to alleviate redundancy in different bits. Extensive evaluations on four popular benchmark datasets demonstrate that our proposed deep hashing framework achieves an excellent performance than the comparison methods.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Chengduan Wang, Lei Huang 0010
IEEE Trans. Multim.1
2024 Deep Neighborhood-aware Proxy Hashing with Uniform Distribution Constraint for Cross-modal Retrieval
abstract
Cross-modal retrieval methods based on hashing have gained significant attention in both academic and industrial research. Deep learning techniques have played a crucial role in advancing supervised cross-modal hashing methods, leading to significant practical improvements. Despite these achievements, current deep cross-modal hashing still encounters some underexplored limitations. Specifically, most of the available deep hashing usually utilizes pair-wise or triplet-wise strategies to promote the separation of the inter-classes by calculating the relative similarities between samples, weakening the compactness of intra-class data from different modalities, which could generate ambiguous neighborhoods. In this article, the Deep Neighborhood-aware Proxy Hashing (DNPH) framework is proposed to learn a discriminative embedding space with the original neighborhood relation preserved. By introducing learnable shared category proxies, the neighborhood-aware proxy loss is proposed to project the heterogeneous data into a unified common embedding, in which the sample is pulled closer to the corresponding category proxy and is pushed away from other proxies, capturing small within-class scatter and big between-class scatter. To enhance the quality of the obtained binary codes, the uniform distribution constraint is developed to make each hash bit independently obey the discrete uniform distribution. In addition, the discrimination loss is designed to preserve modality-specific semantic information of samples. Extensive experiments are performed on three benchmark datasets to prove that our proposed DNPH framework achieves comparable or even better performance compared with the state-of-the-art cross-modal retrieval applications. The corresponding code implementation of our DNPH framework is as follows: https://github.com/QinLab-WFU/OUR-DNPH .
Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Deep Adaptive Quadruplet Hashing With Probability Sampling for Large-Scale Image Retrieval
abstract
With the preferable efficiency in storage and computation, hashing has shown potential application in large-scale multimedia retrieval. Compared with traditional hashing algorithms using hand-crafted characteristics, deep hashing inherits the representational capacity of deep neural networks to jointly learn semantic features and hash functions, encoding raw data into compact binary codes with significant discrimination. Generally, most of the current multi-wise hashing methods view the similarity margins between image pairs as constant values in training process. When the distance between sample pairs exceeds the fixed margin, the hashing network would not learn anything. Besides, available hashing methods commonly introduce the random sampling strategy to build training batches and ignore the sample distribution, which is harmful to parameter optimization. In this paper, we propose a novel Deep Adaptive Quadruplet Hashing with probability sampling (DAQH) for discriminative binary code learning. Specifically, with exploring the distribution relationship of raw samples, a non-uniform probability sampling strategy is proposed to build more informative and representative training batches, while maintaining the diversity of training samples. By introducing the prior similarity of sample pairs to calculate corresponding margins, an adaptive margin quadruplet loss is designed to dynamically preserve the underlying semantic relationships with its neighbors. To tune the attributes of binary codes, by combining quadruple regularization and orthogonality optimization, binary code constraint is developed to make the learned embedding with significant discrimination. Extensive experimental results on various benchmark datasets demonstrate our proposed DAQH framework achieves state-of-the-art visual similarity search performance.
Qibing Qin, Lei Huang 0010, Kezhen Xie, Zhiqiang Wei 0002, Chengduan Wang, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2022 Learning to Classify Weather Conditions from Single Images Without Labels
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Zhiqiang Wei 0002
MMM (1)4
2022 WCATN: Unsupervised deep learning to classify weather conditions from outdoor images
Kezhen Xie, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin
Eng. Appl. Artif. Intell.5
2022 Deep Multi-Similarity Hashing with semantic-aware preservation for multi-label image retrieval
Qibing Qin, Lintao Xian, Kezhen Xie, Wenfeng Zhang, Yu Liu 0022, Jiangyan Dai, Chengduan Wang
Expert Syst. Appl.1
2022 A CNN-based multi-task framework for weather recognition with multi-scale weather cues
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Lei Lyu 0001
Expert Syst. Appl.4
2021 Deep top similarity hashing with class-wise loss for multi-label image retrieval
Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Wenfeng Zhang
Neurocomputing1
2021 Unsupervised Deep Quadruplet Hashing with Isometric Quantization for image retrieval
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Kezhen Xie, Jinkui Hou
Inf. Sci.1
2021 Multi-task learning with deformable convolution
Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin
J. Vis. Commun. Image Represent.5
2021 Graph convolutional networks with attention for multi-label weather recognition
Kezhen Xie, Zhiqiang Wei 0002, Lei Huang 0010, Qibing Qin, Wenfeng Zhang
Neural Comput. Appl.4
2021 Angular regularization for unsupervised domain adaption on person re-identification
Wenfeng Zhang, Lei Huang 0010, Zhiqiang Wei 0002, Qibing Qin, Lei Lv
Neural Comput. Appl.4
2021 Unsupervised Deep Multi-Similarity Hashing With Semantic Structure for Image Retrieval
abstract
With the advance of Convolutional Neural Network, deep hashing methods have shown the great promising performance in large-scale image retrieval. Without depending on extensive human-annotated data, unsupervised hashing is more applicable to image retrieval tasks compared to supervised methods. However, due to the lack of fine-grained supervised signals and multi-similarity constraints, most state-of-the-art unsupervised deep hashing algorithms cannot ensure the correct fine-grained similarity ranking for image pairs. In this paper, we propose a novel unsupervised deep multi-similarity hashing framework to learn compact binary codes by jointly exploiting global-aware and spatial-aware representations, called Unsupervised Deep Multi-Similarity Hashing with Semantic Structure (UDMSH). Specifically, to obtain distinguishing characteristics, we develop a sub-network by jointly learning global semantic structures from Convolutional Neural Network (CNN) and inherent spatial structures from Fully Convolutional Network (FCN). By computing the cosine distance for deep features from image pairs, we construct a similarity matrix with semantic structure, then utilize this matrix to guide hash code learning process. Based on it, we carefully design a multi-level pairwise loss to preserve the correct fine-grained similarity ranking. Furthermore, we introduce Hamming-isometric mapping into unsupervised hashing framework to decrease the quantization errors. Extensive experiments on three widely used benchmarks prove that our proposed UDMSH outperforms several state-of-the-art unsupervised hashing with respect to different evaluation metrics.
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Kezhen Xie, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2020 Deep multilevel similarity hashing with fine-grained features for multi-label image retrieval
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002
Neurocomputing1
2020 Adaptive Attention-Aware Network for unsupervised person re-identification
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Qibing Qin
Neurocomputing5
2019 A Novel Deep Hashing Method with Top Similarity for Image Retrieval
abstract
Due to the advantages of retrieval speed and storage space, deep hashing methods have become a research hotspot in the field of large-scale image retrieval. Most of existing deep hashing methods pay close attention to similarity between images without images at the top of the ranking list similar to query targets. In the paper, a novel deep hashing model is proposed to preserve top images similar to the query images and optimize the quality of hash codes for image retrieval. Specifically, the optimized AlexNet is utilized to extract discriminative image representations and learn hashing functions simultaneously. The loss function based on acceleration strategy is designed to ensure similarity between returned images at the top of the ranking list and query images. In addition, we implement the model training in a batch-process fashion to low the image storage. Moreover, our extensive experiments on standard benchmarks demonstrate that our method outperforms several state-of-the-art deep hashing methods.
Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Xiaopeng Ji
ICASSP1
2016 An Efficient Mining Algorithm for Maximal Weighted Frequent Patterns Based on WIdT-Trees
Qibing Qin, Long Tan
IDEAL1