EDBT 2026 Demo / reviewers in the wild / expert
Lei Huang 0010
dblp:18/1763-10
· DBLP profile ↗
99ranked-venue papers
4as first author
85since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 2 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 36 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Deep hashing via mean centroid representation for large-scale image retrieval
Bingxin Wang, Xianmin Wei, Qibing Qin, Lei Huang 0010 |
Expert Syst. Appl. | 5 |
| 2026 | Polysemic Semantic Instance Network for Cross-Modal HashingabstractHashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performance. Most existing methods rely on one-to-one mappings that preserve only global average semantics, which fail to capture the intrinsic polysemous structures embedded within individual samples. To address this issue, we propose a novel Deep Polysemic Semantic Instance Hashing (DPSIH) method and design a Diverse Semantic Instance Embedding (DSIE) module. This module integrates local and global features through multi-head self-attention and residual learning, generating multiple diverse embeddings per sample to effectively capture fine-grained and polysemous semantic structures. Furthermore, we design a multi-embedding semantic correlation constraint that relaxes strict alignment restrictions to improve robustness under partial alignment, and introduce Maximum Mean Discrepancy (MMD) regularization to alleviate cross-modal distribution shifts. Additionally, an embedding diversity mechanism is proposed to prevent all embeddings from collapsing into a central or averaged representation, thereby enhancing semantic diversity. Extensive experiments on four benchmark datasets demonstrate that DPSIH significantly outperforms state-of-the-art methods and effectively improves the modeling of semantic ambiguity in cross-modal retrieval tasks. Qibing Qin, Kezhen Xie, Wenfeng Zhang, Lei Huang 0010 |
AAAI | 5 |
| 2026 | Deep Potential Semantic-aware Hashing for Cross-modal RetrievalabstractHashing learning has moved into the mainstream for multimedia retrieval because it offers the advantages of low storage cost and high retrieval efficiency. Currently, most cross-modal hashing methods commonly explore the similarity relations between samples by constructing pair-wise or triplet-wise constraints. However, these methods focus on the relative correct ranking of samples, ignore the potential semantic similarity of raw sample distribution, and generate sub-optimal hash codes. To resolve this issue, the novel Deep Potential Semantic-aware Hashing framework (DPSaH) is proposed to mine the local semantic structure of heterogeneous samples, maintaining inter-modality-consistent and cross-modality-correlated semantic relationships. Specifically, by exploring the potential local structure of the data, the multi-modal quadruple loss is extended to the cross-modal hashing framework, thereby preserving the potential semantic neighborhoods among raw samples in Hamming space. During model training, based on the average semantic labels, the label-averaged balanced strategy is developed to quantify the frequency difference between positive and negative samples. Besides, by injecting noise information into the generated discrete codes, the binary-injection loss is introduced to alleviate the over-activation of specific bits, decorrelating different bits in the Hamming space. Extensive experiments are performed on three public datasets, and the results verify the superiority of the DPSaH framework compared to the current mainstream cross-modal hashing frameworks. The source code for DPSaH is available at https://github.com/QinLab-WFU/DPSaH . Qibing Qin, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Deep asymmetric semantic hashing with probability shifting for multi-label image retrieval
Yongyue Fu, Qibing Qin, Jinkui Hou, Congcong Zhu, Lei Huang 0010, Wenfeng Zhang |
Expert Syst. Appl. | 5 |
| 2026 | Deep semantic center-guided hashing for multi-label cross-modal retrieval
Xinzheng Sui, Yadong Huo, Qibing Qin, Lei Huang 0010, Wenfeng Zhang |
Expert Syst. Appl. | 5 |
| 2026 | Deep global-ranking hashing via average precision approximation for large-scale image retrieval
Yongyue Fu, Qibing Qin, Lei Huang 0010, Wenfang Zhang |
Expert Syst. Appl. | 4 |
| 2026 | Deep attribute-aware hashing for zero-shot image retrieval
Yongyue Fu, Qibing Qin, Wenfang Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 5 |
| 2026 | Deep neighborhood-based component proxy hashing for large-scale image retrieval
Huiying Zhu, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 6 |
| 2026 | Deep synthetic-proxy hashing for multi-label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Chunlei Chen, Lei Huang 0010 |
Neurocomputing | 6 |
| 2026 | PE-DETR: A physics-enhanced detection transformer for mesoscale eddy detection
Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
Neurocomputing | 2 |
| 2026 | Deep Softtriple hashing for Multi-Label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010 |
Neural Networks | 5 |
| 2026 | RMPT: Retrieval-based multimodal prompt tuning for event detection
Enyuan Zhao, Jie Nie, Lei Huang 0010, Zhiqiang Wei 0002 |
Pattern Recognit. | 3 |
| 2026 | Deep neighbor-aware hashing with global-local representation for multi-label remote sensing image retrieval
Xiaorong Chen, Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang |
Signal Process. Image Commun. | 5 |
| 2026 | Deep semantic channel hashing for large-scale image retrieval
Qibing Qin, Lei Huang 0010, Wenfeng Zhang |
Signal Process. Image Commun. | 4 |
| 2026 | LoDisc: Learning Global-Local Discriminative Features for Self-Supervised Fine-Grained Visual RecognitionabstractThe self-supervised contrastive learning strategy has attracted considerable attention due to its exceptional ability in representation learning. However, current contrastive learning tends to learn global coarse-grained representations of the image that benefit generic object recognition, whereas such coarse-grained features are insufficient for fine-grained visual recognition. In this paper, we incorporate subtle local fine-grained feature learning into global self-supervised contrastive learning through a pure self-supervised global-local fine-grained contrastive learning framework. Specifically, a novel pretext task called local discrimination (LoDisc) is proposed to explicitly supervise the self-supervised model’s focus toward local pivotal regions, which are captured by a simple but effective location-wise mask sampling strategy. We show that the LoDisc pretext task can effectively enhance fine-grained clues in important local regions and that the global-local framework further refines the fine-grained feature representations of images. Extensive experimental results on different fine-grained object recognition tasks demonstrate that the proposed method can lead to a decent improvement in different evaluation settings. The proposed method is also effective for general object recognition tasks. Jialu Shi, Zhiqiang Wei 0002, Jie Nie, Lei Huang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Reference-Based Super-Resolution With Geometry-Aware Transfer
Lei Huang 0010, Jie Nie, Yadong Huo, Jin Du, Zhiqiang Wei 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Deep Stochastic Spherical Hashing With Von Mises-Fisher Distributions for Cross-Modal RetrievalabstractDeep cross-modal hashing has gained significant attention because of its benefits, including reduced storage requirements and enhanced retrieval efficiency. Although progress has been made, existing deep cross-modal hashing methods still face unresolved challenges. Most existing methods typically adopt Euclidean space as the embedding space to measure the semantic similarity between original samples. However, the volume of Euclidean space grows polynomially with dimension, which exacerbates the curse of dimensionality. In contrast, methods based on spherical space usually use cosine similarity as the metric, effectively mitigating the aforementioned problem by normalizing the embedding vectors. Nevertheless, such methods only considers the direction to determine the category, ignoring the uncertainty measure in the embedding space, thus having a limited ability to preserve inherent multimodal semantics. In this paper, with a novel extension of the maximum entropy distribution on the surface of a hypersphere von Mises-Fisher (vMF) distribution, a novel deep cross-modal hashing method, named Deep Stochastic Spherical Hashing (DSSH), is designed to utilize uncertain information to guide the hashing process and produce discriminative modality-invariant hash codes. Specifically, to learn explicit uncertainty in learned embedding space, the Spherical von Mises-Fisher distribution is applied for the f irst time in deep cross-modal hashing, where the direction of the sample embedding controls its position on the hyper sphere, thereby preventing its semantic content, and its norm parameterizes the determinism of the distribution. In addition, stochastic spherical von Mises–Fisher loss is proposed to preserve the mode-specific semantic information of the sample, achieving the alignment of different modalities and semantic embeddings. Extensive experiments on four benchmark datasets show that our DSSH framework outperforms existing state-of-the-art cross modal hashing methods. The source code of the experiments is available at https://github.com/QinLab-WFU/DSSH. Qibing Qin, Meiling Ge, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2026 | Deep Distance Weighted Sampling Hashing for Cross-Modal RetrievalabstractCross-modal hashing seeks to encode heterogeneous image-text data into compact binary codes for efficient retrieval. While significant efforts have been made to study sampling strategies, most of these approaches are tightly integrated with loss function engineering, lacking an independent focus on sampling methods. In this paper, we challenge the convention by revealing that batch-level sampling strategy is equally pivotal as loss design for learning discriminative hash codes. Specifically, by introducing a novel distribution-aware sampling strategy, a Distance Weighted Sampling Hashing (DDWSH) framework is proposed to dynamically select stable and informative training pairs. Unlike conventional random or semi-hard sampling, our method weights pairwise distances within each batch to approximate global data distribution, thereby mitigating training instability caused by biased sampling. To rigorously validate our claims, we conduct the first comprehensive crossover study between sampling strategies (random/semi-hard/ours) and loss functions (contrastive/triplet/ours) across three benchmark datasets. Experiments demonstrate that: Universality: Our sampling boosts all loss functions' performance (average +4.3% mAP vs. semi-hard mining), Superiority: DDWSH competes with complex loss function design and achieves state-of-the-art results, and Stability: It reduces performance variance by 60.2% compared to semi-hard sampling under varying batch compositions average. This systematic analysis establishes sampling as an independent research dimension in deep hashing, beyond a mere part of loss function engineering. The source code for DDWSH is freely available athttps://github.com/QinLab-WFU/DDWSH. Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Multim. | 4 |
| 2026 | Deep Neighbor Discriminant Binary Embedding for Multi-Label Image RetrievalabstractBecause of fast retrieval speed and low storage cost, deep hash learning has become a current research spot in multimedia retrieval, which arouses more and more interest and attention. Most of the available deep hashing methods usually employ the mini-batch strategy to construct training batches, and only one mini-batch of samples is valuable at each iteration during model optimization, which fails to explore the original neighbor structure well, leading to sub-optimal embeddings, especially for relatively large datasets. By contrast, the superior performance is achieved by optimizing the SoftMax loss for certain hashing learning, meanwhile, prior research has suggested the normalized SoftMax loss is essentially identical to one smoothed triplet-wise function, in which each class is assigned to one single semantic center. Nevertheless, under real-world scenarios, one class could correspond to multiple local clusters rather than just a single one. To this end, by extending SoftMax loss with multiple semantic centers, a novel deep hashing framework, called Deep Neighbor Discriminant Binary Embedding (NDBE) framework, is presented to generate discriminative hash codes with the original neighbor structure preservation. Specifically, by expanding the size of the last fully connected layer to increase multiple centers for each class, a novel SoftTriple loss is proposed to capture the latent semantic distribution of the original samples and reduce the intra-class variance, which is optimized without constraints sampling. To learn the different numbers of each class center, the class-center aggregation strategy is developed to obtain the compact set of centers while maintaining semantic similarity. Extensive experiments on three public multi-label datasets show that our proposed NDBE framework achieves superior visual similarity search performance over several state-of-the-art approaches. The source code for the implementation of our proposed NDBE framework is available athttps://github.com/QinLab-WFU/NDBE. Qibing Qin, Mingkun Dou, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Multim. | 4 |
| 2026 | Deep Semantic Tuplet-Based Hashing by Hypergraph Modeling for Cross-Modal RetrievalabstractWith low storage cost and high retrieval efficiency, hashing techniques are widely used for multi-media retrieval, which has already become the present research focus. Currently, cross-modal hashing commonly employs graph-based loss to construct pair-wise semantic relations between training samples for model optimization. However, limited by the graph-based strategy, each edge in the graph only connects two samples, which only represent a bundle of pair-wise relationships. Besides, the edges in the graph are calculated by self-attention or feature distance, only considering pair-wise relations of heterogeneous samples and ignoring the class relations. In this paper, by hypergraph modeling the semantic tuples, a novel Deep Semantic Tuplet-based Hashing by Hypergraph Modeling (DSTH) approach is proposed to leverage the multilateral semantic relations, which could guide the model to learn class-discriminative semantic binary embedding. In more detail, based on the characteristic distribution, semantic tuples are constructed for each class in one mini-batch, which represents the multilateral semantic relationships between multiple samples and multiple classes. By considering semantic tuples as hyperedges to represent multilateral semantic relations, hypergraph modeling is designed, in which HyperGraph Neural Hetwork (HGNH) is introduced to formulate hypergraph node classification goals to fully learn the multilateral semantic information contained in the semantic tuples. Moreover, to utilize the heterogeneity of local structures in embedding, the adaptive neighborhood structure is explored by learning the structure embedding, which provides fine-grained ranking lists. Through extensive experiments on three benchmark datasets, the comprehensive results validate the advancement of our proposed DSTH framework over mainstream cross-modal hashing. The source code for the framework DSTH is freely available athttps://github.com/QinLab-WFU/DSTH. Qibing Qin, Wenfeng Zhang, Huihui Zhang 0003, Lei Huang 0010, Jie Nie |
IEEE Trans. Multim. | 5 |
| 2026 | Generative Zero-Shot Hashing for Multi-Label Image RetrievalabstractDue to the preferable efficiency of storage and computation, hashing algorithms show great potential. With the explosive growth of web data, new and emerging categories of multimedia data continue to increase, and zero-shot hashing has been one of the hot research topics in retrieval research. Nevertheless, existing zero-shot hashing methods focus mainly on single-label image retrieval, using class semantic embeddings as a link between visual features and binary codes to align visual features with corresponding class semantics and simultaneously transfer knowledge from seen classes to unseen classes. Meanwhile, single-label zero-shot learning employs Generative Adversarial Networks (GANs) to synthesize class-specific characteristics from the corresponding class attribute embeddings, achieving encouraging results. However, synthesizing multi-label features from GANs remains unexplored in the context of zero-shot hashing settings. When multiple objects co-appear in an image, a key issue is how to fuse multi-class information effectively. In this article, by introducing the multi-class information fusion strategy, a novel Generative Zero-Shot Hashing (GZSH) is proposed to transform zero-shot hashing into traditional supervised hashing by generating features of unseen categories. Specifically, this study introduces three different fusion methods (attribute-based fusion, feature-based fusion, and semantic-enhanced fusion) for synthesizing multi-label features from the corresponding multi-label class embeddings. Subsequently, the semantic-enhanced fusion is integrated into the representative generative architecture to learn the feature distribution of unlabeled images in the context of multi-label zero-shot hashing. Besides, by jointly learning the seen/source and unseen/target samples, a pairwise similarity loss is introduced to optimize the deep hashing model. Extensive experimentation on three widely used multi-label datasets illustrates the outstanding performance of our proposed GZSH framework compared to current state-of-the-art methods. The implementation code for our GZSH framework can be found at https://github.com/QinLab-WFU/GZSH . Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Deep Uncertainty-aware Probabilistic Hashing for Cross-modal RetrievalabstractDue to its outstanding computational efficiency and low storage requirements, hashing technology has become a research hotspot in large-scale multimedia retrieval. In cross-modal hashing, the key lies in mapping samples from different data sources into a common discrete space. However, most existing methods assume that the input data is of high quality and completeness. When samples are incomplete or degraded (e.g., blurry images or incomplete text), the absence or ambiguity of semantic information inevitably compromises retrieval accuracy. Traditional deterministic embedding methods typically map multi-modal samples to a single point in the embedding space without considering uncertainty. As a result, inherent noise or feature ambiguity in the inputs may lead to distorted or shifted binary representations. To resolve this problem, this article proposes a novel Deep Uncertainty-aware Probabilistic Hashing (DUaPH) method that models the uncertainty of multi-modal samples. By capturing the underlying distribution of heterogeneous data, DUaPH effectively mitigates feature ambiguity and inherent noise, enhancing the robustness of cross-modal retrieval. Specifically, each heterogeneous sample is mapped to a multivariate Gaussian distribution, where the mean represents the most probable semantic features, and the variance reflects the sample uncertainty. A semantic feature matching mechanism is introduced to dynamically adjust the importance of feature dimensions, prioritizing those with higher certainty. Then, a semantic feature fusion mechanism is developed to integrate the semantic features from multi-modal sample pairs, producing a new distribution with reduced uncertainty and improved semantic alignment. Extensive experiments on four benchmark datasets demonstrate that DUaPH significantly improves robustness and retrieval performance under conditions of semantic ambiguity and data uncertainty. The source code is available at https://github.com/QinLab-WFU/DUaPH . Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Deep Relational Knowledge Distillation Hashing via Relaxed Masking Triplet Optimization for Large-scale Image RetrievalabstractLarge-scale image retrieval increasingly depends on triplet-based deep hashing methods, which typically construct training triplets using class labels. However, existing triplet-based hashing methods relying on sparse and noisy categorical labels suffer from two main issues: no triplet, where sparse annotations limit valid triplet formation, and bad triplet, where noisy or ambiguous labels cause misleading samples, which together hinder the network from learning a well-structured similarity space, causing semantically similar images to scatter and dissimilar ones to collapse, ultimately undermining retrieval accuracy. In this article, we solve this dilemma with a novel unified deep hashing framework, termed Deep Relational Knowledge Distillation Hashing (DRKDH), which decouples semantic relation modeling from visual feature learning by leveraging a teacher-student paradigm to learn semantically consistent hash codes. Specifically, the teacher model captures comprehensive semantic structural relationships across all instances, providing information-rich and noise-resilient supervision to guide the student model in learning expressive and discriminative visual embeddings, effectively mitigating the impact of low-quality triplet labels. Furthermore, by leveraging similarity weights predicted by the teacher model, a relaxed masking triplet loss is introduced into the student model, which dynamically adjusts the contribution of each triplet based on its informativeness, suppressing invalid or misleading triplets while emphasizing valuable ones to enhance training efficiency. Comprehensive experiments on THINGS, ImageNet, MIRFLICKR-25K, and NUS-WIDE show that DRKDH consistently outperforms state-of-the-art deep hashing methods under sparse and noisy supervision, while additional evaluations suggest competitive performance in selected zero-shot and few-shot settings. Source codes: https://github.com/QinLab-WFU/DRKDH . Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal RetrievalabstractThe field of cross-modal retrieval aims to construct a shared representation space for samples from multiple modalities, typically within the vision and language domains. Deep hashing, with its high computational efficiency and low storage costs, has emerged as a central focus in this field and has garnered significant attention in recent research. However, current hash retrieval, concentrating on deterministic methods, struggles to effectively capture semantically ambiguous correspondences between cross-modal samples, where heterogeneous data have complex-semantic many-to-many relationships in the latent space. To address this limitation, we propose a novel Deep Probabilistic Binary Embedding (DPBE) framework, designed to generate discriminative, modality-invariant hash codes that facilitate accurate and reliable cross-modal retrieval. In contrast to contemporary probabilistic methods, we focus on optimizing hash networks to learn more accurate binary embeddings by using the learning mode of probabilistic embeddings. We introduce the first Bayesian encoder for hash learning, which employs Laplace Approximation to model a distribution over network weights. Extensive experimental results demonstrate that our approach not only outperforms deterministic methods in retrieval performance but also provides uncertainty estimates, enhancing the interpretability of the embeddings. The corresponding code is available at https://github.com/QinLab-WFU/DPBE. Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
ACM Multimedia | 4 |
| 2025 | Factorized Transformer Hashing with Adaptive Routing for Large-scale Image RetrievalabstractTransformer architecture has driven significant advancements in deep hashing, establishing itself as a dominant framework for large-scale retrieval and storage applications. However, existing Transformer-based deep hashing methods typically employ unvarying feature transformations across all images, limiting their adaptability to diverse visual patterns. This rigidity restricts the model's capacity to learn both highly distinctive and generalizable discrete representations, posing challenges for retrieval in open-world scenarios. To overcome this challenge, we propose a novel Factorized Transformer Hashing (FTH) framework, which introduces a factorized transformer to enhance the generalization and discriminative power of hash codes. Specifically, we decompose the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) blocks into multiple sub-blocks, forming a transformer factorization scheme that captures diverse feature characteristics through independent sub-blocks. Furthermore, we develop an adaptive selection strategy, leveraging a set of learnable selectors with the Softmax function, to dynamically route each image to the most appropriate sub-block for processing. Extensive experiments on three benchmark datasets demonstrate that the proposed FTH framework significantly outperforms state-of-the-art baselines in both image hashing and zero-shot hashing tasks. Source code is available at https://github.com/QinLab-WFU/FTH. Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
ACM Multimedia | 4 |
| 2025 | Ranking-oriented cross-modal hashing
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Deep informative-triplet sampling hashing with attention-aware augmentation for remote sensing image retrieval
Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 4 |
| 2025 | Frequency domain transfer learning for remote sensing visual question answering
Enyuan Zhao, Ziyi Wan, Jie Nie, Lei Huang 0010 |
Expert Syst. Appl. | 7 |
| 2025 | Deep adaptive gradient-triplet hashing for cross-modal retrieval
Congcong Zhu, Jinkui Hou, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 6 |
| 2025 | Deep neighbor-coherence hashing with discriminative sample mining for supervised cross-modal retrieval
Congcong Zhu, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Expert Syst. Appl. | 4 |
| 2025 | Omni-frequency diffusion-based regional feature consistency recovery for realistic image super-resolution
Chenjuan Zuo, Zhiqiang Wei 0002, Xiaodong Wang 0006, Jie Nie, Lei Huang 0010 |
Expert Syst. Appl. | 5 |
| 2025 | Image-text aggregation for open-vocabulary semantic segmentation
Shengyang Cheng, Jianyong Huang, Xiaodong Wang 0006, Lei Huang 0010, Zhiqiang Wei 0002 |
Neurocomputing | 4 |
| 2025 | Personalized Dual Transformer Network for sequential recommendation
Meiling Ge, Chengduan Wang, Xueyang Qin, Jiangyan Dai, Lei Huang 0010, Qibing Qin, Wenfeng Zhang |
Neurocomputing | 5 |
| 2025 | Deep Consistent Penalizing Hashing with noise-robust representation for large-scale image retrieval
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
Neurocomputing | 4 |
| 2025 | Deep multi-similarity hashing via label-guided network for cross-modal retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang |
Neurocomputing | 5 |
| 2025 | Deep binary hyperbolic embedding for large-scale image retrieval
Enhao Wang, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010 |
Neurocomputing | 6 |
| 2025 | Image sterilization via overwriting
Mingyao Tan, Lei Huang 0010, Zhaorui Gu, Xiaodong Wang 0006, Haiyong Zheng |
Inf. Sci. | 3 |
| 2025 | Domain-invariant multi-granularity feature learning for generalizable person re-identification
Wenfeng Zhang, Xiangfei Cao, Lei Huang 0010, Dengwei Yan, Qibing Qin |
Knowl. Based Syst. | 3 |
| 2025 | Deep Hardness-Aware Hashing for Large-Scale Image Retrieval
Chunping Dong, Meiling Ge, Enhao Wang, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
IEEE Signal Process. Lett. | 6 |
| 2025 | Deep Discriminative Boundary Hashing for Cross-Modal RetrievalabstractBy the preferable efficiency in storage and computation, deep cross-modal has gained much attention in large-scale multimedia retrieval. Current deep hashing employs the probability outputs of the likelihood function, i.e., Sigmoid or Cauchy, to quantify the semantic similarity between samples in a common Hamming space. However, the inherent weakness of the Sigmoid likelihood function or the Cauchy likelihood function in gradient optimization leads to hashing models failing to exactly describe the hamming ball, which indicates the absolute semantic boundary among classes, thereby giving the high neighborhood ambiguity. In this paper, with the analysis of the likelihood function from the perspective of similarity metric learning, the novel Deep Discriminative Boundary Hashing framework (DDBH) is proposed to learn the discriminative embedding space that separates neighbors and non-neighbors well. Specifically, by introducing the remapping strategy and the base-point adaptive selection, the boundary-preserving loss based on the adjustable likelihood function is proposed to project data points with small gradients to regions with large gradients and give larger gradients for hard samples, facilitating better separation among classes. Meanwhile, to learn class-dependent binary codes, the class-wise quantization loss is designed to heuristically transfer class-wise prior knowledge to the binary quantization, significantly improving the discriminative capability of compact discrete codes. Comprehensive experiments on three benchmark datasets show that our proposed DDBH framework outperforms other representative deep cross-modal hashing. The corresponding code is available at https://github.com/QinLab-WFU/DDBH. Qibing Qin, Yadong Huo, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Deep Semantic-Consistent Penalizing Hashing for Cross-Modal RetrievalabstractBenefiting from the advantages of low storage cost and high retrieval efficiency, hash learning could significantly speed up large-scale cross-modal retrieval. Based on the prior annotations, most of the available cross-modal hashing usually introduces the margin-based constraint to generate different boundaries for each class in the inference phase, optimizing the model. However, these obtained label-guided penalty boundaries may differ from the primitive semantic relationships between heterogeneous modalities, impairing retrieval performance. Besides, the margin-based constraint is too weak to penalize the classes with low intra-class variances or inter-class correlations, which struggle to learn high-quality embeddings. In this paper, we propose a novel Deep Semantic-consistent Penalizing Hashing framework (DScPH) to learn the consistent penalizing fields for all classes, achieving accurate and efficient cross-modal retrieval. Specifically, by exploring unbalanced intra-class and inter-class correlations, the consistent penalizing loss is introduced into cross-modal retrieval to learn the consistency decision boundaries across classes. During training, the dice-like optimization strategy is developed to balance the pulling penalizing elements and pushing penalizing elements, facilitating the model convergence. Besides, based on the invariance of similarity measures under orthogonal transformations, the alternative quantization is proposed to minimize the errors between the learned continuous embeddings and binary discretization, maintaining the consistency of semantic relationships after performing binary projection. Extensive experiments are conducted on three benchmark datasets, and the comprehensive results validate the efficacy of our proposed DScPH framework, which outperforms the current mainstream deep cross-modal hashing algorithms. Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Multim. | 4 |
| 2025 | Cross-Modal Progressive Perspective Matching Network for Remote Sensing Image-Text RetrievalabstractCross-modality based on remote sensing (RS) text-image retrieval has gained increasing attention in recent years due to its ability to leverage the rich semantics of images and the understandability of text to provide a more comprehensive description. Existing cross-modal retrieval methods typically apply self-attention or cross-attention mechanisms to identify important information in RS data, but they ignore the multi-view perception characteristic of geographical space in RS images. As a result, these retrieval models fail to locate the correct perspective in images according to the query text, ultimately leading to incorrect matching. In this work, a Cross-modal Progressive Perspective Matching Network (CPPMN) is proposed for remote sensing image-text retrieval by establishing a progressive perspective matching mechanism and semantic alignment to further improve the performance of the retrieval model. Specifically, the CPPMN framework consists of three core modules: the Compensation Network for Full Perspective Modeling (CN_FPM), the Graph Transformation for Individual Perspective Modeling (GT_IPM), and the Cascaded Transformer for Cross-modal Semantic Alignment (CT_CSA). The CN_FPM module utilizes all positive text samples as supervision signals to guide the feature extraction training process, aiming to capture full perspective information from images. Subsequently, the GT_IPM module transforms implicit-perspective feature representations into explicit-perspective cross-modal relationship graphs. This transformation enables the identification of specific perspective locations within the image according to the query sentence by analyzing graph density and connectivity. Finally, the CT_CSA module comprises a cascaded Transformer network that aligns features at the semantic level between cross-modal data The quantitative and qualitative experiments are conducted on four large-scale remote sensing cross-modal retrieval datasets to demonstrate the significant performance of adopting the progressive perspective matching mechanism and semantic alignment strategy. Xiu Li 0006, Lei Huang 0010, Shan Du 0001, Jie Nie, Junyu Dong |
IEEE Trans. Multim. | 4 |
| 2024 | Counting in congested crowd scenes with hierarchical scale-aware encoder-decoder network
Run Han, Ran Qi, Xuequan Lu, Lei Huang 0010, Lei Lyu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | SwinFG: A fine-grained recognition scheme based on swin transformer
Anzhuo Chu, Lei Huang 0010, Zhiqiang Wei 0002 |
Expert Syst. Appl. | 4 |
| 2024 | Deep global semantic structure-preserving hashing via corrective triplet loss for remote sensing image retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang |
Expert Syst. Appl. | 5 |
| 2024 | Strong robust copy-move forgery detection network based on layer-by-layer decoupling refinement
Jingyu Wang 0005, Xuesong Gao, Jie Nie, Xiaodong Wang 0006, Lei Huang 0010, Weizhi Nie, Mingxing Jiang, Zhiqiang Wei 0002 |
Inf. Process. Manag. | 5 |
| 2024 | ZoomViT: an observation behavior-based fine-grained recognition scheme
Yongquan Yang, Haicheng Wang, Lei Huang 0010, Zhiqiang Wei 0002 |
Neural Comput. Appl. | 4 |
| 2024 | Unsupervised Deep Hashing With Fine-Grained Similarity-Preserving Contrastive Learning for Image RetrievalabstractUnsupervised deep hashing has demonstrated significant advancements with the development of contrastive learning. However, most of previous methods have been hindered by insufficient similarity mining using global-only image representations. This has led to interference from background or non-interest objects during similarity reconstruction and contrastive learning. To address this limitation, we propose a novel unsupervised deep hashing framework named Fine-grained Similarity-preserving Contrastive learning Hashing (FSCH), which explores fine-grained semantic similarity among different images and their augmented views more comprehensively. It mainly comprises two modules: the global-local fine-grained similarity consistency preservation module and the local fine-grained similarity contrast preservation module. Specifically, we reconstruct local pairwise similarity structures by matching fine-grained patches, in conjunction with global similarity structures based on global hash codes cosine similarity, to generate hash codes with the ability to preserve global-local similarity consistency. Moreover, the preservation of local fine-grained similarity among augmented views is accomplished through the common regional features mutual representation between patches, then we enhance the discriminability of hash codes by mitigating the potential features difference during contrastive learning. Experimental results on four benchmark datasets demonstrate that our FSCH achieves an excellent retrieval performance compared to state-of-the-art unsupervised hashing methods. Hu Cao, Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Deep Semantic-Aware Proxy Hashing for Multi-Label Cross-Modal RetrievalabstractDeep hashing has attracted broad interest in cross-modal retrieval because of its low cost and efficient retrieval benefits. To capture the semantic information of raw samples and alleviate the semantic gap, supervised cross-modal hashing methods that utilize label information which could map raw samples from different modalities into a unified common space, are proposed. Although making great progress, existing deep cross-modal hashing methods are suffering from some problems, such as: 1) considering multi-label cross-modal retrieval, proxy-based methods ignore the data-to-data relations and fail to explore the combination of the different categories profoundly, which could lead to some samples without common categories being embedded in the vicinity; 2) for feature representation, image feature extractors containing multiple convolutional layers cannot fully obtain global information of images, which results in the generation of sub-optimal binary hash codes. In this paper, by extending the proxy-based mechanism to multi-label cross-modal retrieval, we propose a novel Deep Semantic-aware Proxy Hashing (DSPH) framework, which could embed multi-modal multi-label data into a uniform discrete space and capture fine-grained semantic relations between raw samples. Specifically, by learning multi-modal multi-label proxy terms and multi-modal irrelevant terms jointly, the semantic-aware proxy loss is designed to capture multi-label correlations and preserve the correct fine-grained similarity ranking among samples, alleviating inter-modal semantic gaps. In addition, for feature representation, two transformer encoders are proposed as backbone networks for images and text, respectively, in which the image transformer encoder is introduced to obtain global information of the input image by modeling long-range visual dependencies. We have conducted extensive experiments on three baseline multi-label datasets, and the experimental results show that our DSPH framework achieves better performance than state-of-the-art cross-modal hashing methods. The code for the implementation of our DSPH framework is available athttps://github.com/QinLab-WFU/DSPH. Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | A Spatial-Frequency Fusion Strategy Based on Linguistic Query Refinement for RSVGabstractRemote sensing visual grounding (RSVG) represents a pivotal task aimed at pinpointing regions highly relevant to textual descriptions by parsing remote sensing image content. Existing RSVG methodologies focus on multimodal semantic integration on specific datasets, overlooking the intricate high-dimensional information present in remote sensing images, where color, scale, and semantics are tightly coupled. Consequently, these approaches exhibit limitations in handling specific receptive fields or preserving structural information within remote sensing images, leading to insufficient localization precision. To address this issue, this article introduces a strategy that incorporates spatial frequency information, leveraging Fourier transforms to capture global structured information from remote sensing data. This is followed by a progressive aggregation across spatial, spectral, and linguistic modalities to achieve robust semantic coreference. The primary contributions of this article are as follows: First, a spatial-frequency fusion strategy based on linguistic query refinement is proposed, which enhances visual grounding performance significantly by extracting spectral features with potent spatial perception capabilities through the expansion of the spectral receptive field. Second, a frequency-guided spatial (FGS) module is designed, utilizing amplitude and phase-structured spectral features to augment spatial representation capabilities further. Lastly, a query-aware original attention (QOA) mechanism is developed, facilitating deep integration of spatial and spectral information under linguistic guidance. Extensive experimentation on the RSVGD dataset validates the efficacy of the proposed approach, demonstrating superior performance when compared to state-of-the-art methods. Enyuan Zhao, Ziyi Wan, Jie Nie, Lei Huang 0010 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Deep Hierarchy-Aware Proxy Hashing With Self-Paced Learning for Cross-Modal RetrievalabstractDue to its low storage cost and high retrieval efficiency, hashing technology is popularly applied in both academia and industry, which provides an interesting solution for cross-modal similarity retrieval. However, most existing supervised cross-modal hashing methods typically view the fixed-level semantic affinity defined by manual labels as supervised signals to guide hash learning, which only represents a small subset of complex semantic relations between multi-modal samples, thus impeding the hash function learning and degrading the obtained hash codes. In the paper, by learning shared hierarchy proxies, a novel deep cross-modal hashing framework, called Deep Hierarchy-aware Proxy Hashing (DHaPH), is proposed to construct the semantic hierarchy in a data-driven manner, thereby capturing the accurate fine-grained semantic relationships and achieving small intra-class scatter and big inter-class scatter. Specifically, by regarding the hierarchical proxies as learnable ancestors, a novel hierarchy-aware proxy loss is designed to model the latent semantic hierarchical structures from different modalities without prior hierarchy knowledge, in which similar samples share the same Lowest Common Ancestor (LCA) and dissimilar points have different LCA. Meanwhile, to adequately capture valuable semantic information from hard pairs, a multi-modal self-paced loss is introduced into cross-modal hashing to reweight multi-modal pairs dynamically, which enables the model to gradually focus on hard pairs while simultaneously learning universal patterns from multi-modal pairs. Extensive experiments on three available benchmark databases demonstrate that our proposed DHaPH framework outperforms the compared baselines with different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DHaPH. Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Deep Neighborhood-Preserving Hashing With Quadratic Spherical Mutual Information for Cross-Modal RetrievalabstractDriven by the high nonlinearity of deep neural networks, deep hashing has achieved the pictured great potential in cross-modal retrieval applications, significantly bridging the modality gap. Current deep cross-modal hashing usually utilizes affinity matching or local ranking to capture the local semantic relationships in the learned common space, leading to high neighborhood ambiguity. Simultaneously, most of these frameworks utilize additional regularization terms or margin thresholds to enhance the overall performance, in which searching the model's hyper-parameters under mass training data would have a substantial overhead. In this paper, with a novel extension of information-theoretic measures, a novel deep cross-modal hashing method, named Deep Neighborhood-preserving Hashing (DNpH), is designed to learn a highly separable discrete space, effectively mitigating the semantic gap across different modalities. Specifically, to minimize neighborhood ambiguity, the Quadratic Spherical Mutual Information (QSMI) is first introduced into deep cross-modal hashing to separate neighbors and non-neighbors well, while it is free of tuning parameters during model training compared with other similarity measures. To optimize quadratic mutual information loss smoothly, a square clamping method is developed to improve the stability of model optimization, avoiding converging on bad local optimum. Besides, two transformer encoders are exploited as feature extractors for multi-modal samples to learn the informative semantic representations. Finally, we compare our proposed DNpH framework with various state-of-the-art cross-modal hashing on four public datasets, and large amounts of experiment results demonstrate our contributions and show that DNpH outperforms the compared baselines on different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DNpH. Qibing Qin, Yadong Huo, Lei Huang 0010, Jiangyan Dai, Huihui Zhang 0003, Wenfeng Zhang |
IEEE Trans. Multim. | 3 |
| 2024 | Deep Neighborhood Structure-Preserving Hashing for Large-Scale Image RetrievalabstractDeep hashing integrates the advantages of deep learning and hashing technology, and has become the mainstream of the large-scale image retrieval field. However, when training the deep hashing models, most of the existing approaches regard the similarity margin of image pairs as a constant. Once similarity distance exceeds the fixed margin, the network will not learn anything, which easily results in model collapses. In this paper, we address this dilemma with a novel unified deep hashing framework, termed Deep Neighborhood Structure-preserving Hashing (DNSH), to generate the similarity-preserving and discriminative hash codes. Specifically, by extracting the discriminative object characteristics with large variances, we design an adaptive margin quadruplet loss to further explore the underlying similarity relationship between image pairs, reflecting the correct semantic structure among its neighbors. Based on the quadruple form, we develop a quadruple regularization to decrease quantization errors between binary-like embedding and hashing codes. Furthermore, through learning bit balance and bit independent terms jointly, we present the binary code constraint loss to alleviate redundancy in different bits. Extensive evaluations on four popular benchmark datasets demonstrate that our proposed deep hashing framework achieves an excellent performance than the comparison methods. Qibing Qin, Kezhen Xie, Wenfeng Zhang, Chengduan Wang, Lei Huang 0010 |
IEEE Trans. Multim. | 5 |
| 2024 | Towards Adaptive Multi-Scale Intermediate Domain via Progressive Training for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) involves the transfer of knowledge from a labelled source domain to an unlabelled target domain. Recent studies have introduced the concept of intermediate domains to handle significant domain discrepancies between source and target domains. Constructing an appropriate intermediate domain is a crucial step in handling scenarios with substantial domain differences. In this paper, we propose a novel progressive UDA method called adaptive multi-scale intermediate domain via progressive training (AMPT), which has achieved remarkable effectiveness in alleviating large discrepancies between domains. We design a multi-scale similarity metrics module to solve the issue of different scales across different domains by simultaneously computing the pairwise distance between source domain images and the target domain images on multiple scales. Furthermore, we explore a progressive training strategy that facilitates smooth adaptation of the source domain to the target domain by utilizing the intermediate domain. During the progressive training process, considering a positive feedback mechanism, we iteratively leverage task losses and distillation loss through a dynamic threshold. This ensures that the trained intermediate domain branch progressively constrains the target domain branch, and the intermediate domain can be generated dynamically, leading to a smoother gradual adaptation from the source domain to the target domain. Extensive experimental results demonstrate the superiority of our proposed method AMPT on well-known UDA datasets, including Office-Home, Office-31 and DomainNet. Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Deep Neighborhood-aware Proxy Hashing with Uniform Distribution Constraint for Cross-modal RetrievalabstractCross-modal retrieval methods based on hashing have gained significant attention in both academic and industrial research. Deep learning techniques have played a crucial role in advancing supervised cross-modal hashing methods, leading to significant practical improvements. Despite these achievements, current deep cross-modal hashing still encounters some underexplored limitations. Specifically, most of the available deep hashing usually utilizes pair-wise or triplet-wise strategies to promote the separation of the inter-classes by calculating the relative similarities between samples, weakening the compactness of intra-class data from different modalities, which could generate ambiguous neighborhoods. In this article, the Deep Neighborhood-aware Proxy Hashing (DNPH) framework is proposed to learn a discriminative embedding space with the original neighborhood relation preserved. By introducing learnable shared category proxies, the neighborhood-aware proxy loss is proposed to project the heterogeneous data into a unified common embedding, in which the sample is pulled closer to the corresponding category proxy and is pushed away from other proxies, capturing small within-class scatter and big between-class scatter. To enhance the quality of the obtained binary codes, the uniform distribution constraint is developed to make each hash bit independently obey the discrete uniform distribution. In addition, the discrimination loss is designed to preserve modality-specific semantic information of samples. Extensive experiments are performed on three benchmark datasets to prove that our proposed DNPH framework achieves comparable or even better performance compared with the state-of-the-art cross-modal retrieval applications. The corresponding code implementation of our DNPH framework is as follows: https://github.com/QinLab-WFU/OUR-DNPH . Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Scale-Semantic Joint Decoupling Network for Image-Text Retrieval in Remote SensingabstractImage-text retrieval in remote sensing aims to provide flexible information for data analysis and application. In recent years, state-of-the-art methods are dedicated to “scale decoupling” and “semantic decoupling” strategies to further enhance the capability of representation. However, these previous approaches focus on either the disentangling scale or semantics but ignore merging these two ideas in a union model, which extremely limits the performance of cross-modal retrieval models. To address these issues, we propose a novel Scale-Semantic Joint Decoupling Network (SSJDN) for remote sensing image-text retrieval. Specifically, we design the Bidirectional Scale Decoupling (BSD) module, which exploits Salience Extraction Map (SEM) and Salience Suppression Map (SSM) units to adaptively extract potential features and suppress cumbersome features at other scales in a bidirectional pattern to yield different scale clues. Besides, we design the Label-supervised Semantic Decoupling (LSD) module by leveraging the category semantic labels as prior knowledge to supervise images and texts probing significant semantic-related information. Finally, we design a Semantic-guided Triple Loss (STL), which adaptively generates a constant to adjust the loss function to improve the probability of matching the same semantic image and text and shorten the convergence time of the retrieval model. Our proposed SSJDN outperforms state-of-the-art approaches in numerical experiments conducted on four benchmark remote sensing datasets. Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-task learning based on geometric invariance discriminative features
Lei Huang 0010, Wenfeng Zhang, Yanxiu Sheng, Zhiqiang Wei 0002 |
Appl. Intell. | 2 |
| 2023 | Fine-grained Image Recognition via Attention Interaction and Counterfactual Attention Network
Lei Huang 0010, Chen An, Xiaodong Wang 0006, Leon Bevan Bullock, Zhiqiang Wei 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Sea surface height data reconstruction via inter and intra layer features based on dual attention
Lei Huang 0010, Zhiqiang Wei 0002, Chen An, Xianqing Lv |
Neurocomputing | 2 |
| 2023 | Unsupervised domain adaptation via reliable pseudolabeling based memory module and dynamic distance threshold learning
Guanqun Wei, Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
Knowl. Based Syst. | 3 |
| 2023 | Deep Adaptive Quadruplet Hashing With Probability Sampling for Large-Scale Image RetrievalabstractWith the preferable efficiency in storage and computation, hashing has shown potential application in large-scale multimedia retrieval. Compared with traditional hashing algorithms using hand-crafted characteristics, deep hashing inherits the representational capacity of deep neural networks to jointly learn semantic features and hash functions, encoding raw data into compact binary codes with significant discrimination. Generally, most of the current multi-wise hashing methods view the similarity margins between image pairs as constant values in training process. When the distance between sample pairs exceeds the fixed margin, the hashing network would not learn anything. Besides, available hashing methods commonly introduce the random sampling strategy to build training batches and ignore the sample distribution, which is harmful to parameter optimization. In this paper, we propose a novel Deep Adaptive Quadruplet Hashing with probability sampling (DAQH) for discriminative binary code learning. Specifically, with exploring the distribution relationship of raw samples, a non-uniform probability sampling strategy is proposed to build more informative and representative training batches, while maintaining the diversity of training samples. By introducing the prior similarity of sample pairs to calculate corresponding margins, an adaptive margin quadruplet loss is designed to dynamically preserve the underlying semantic relationships with its neighbors. To tune the attributes of binary codes, by combining quadruple regularization and orthogonality optimization, binary code constraint is developed to make the learned embedding with significant discrimination. Extensive experimental results on various benchmark datasets demonstrate our proposed DAQH framework achieves state-of-the-art visual similarity search performance. Qibing Qin, Lei Huang 0010, Kezhen Xie, Zhiqiang Wei 0002, Chengduan Wang, Wenfeng Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Cross-Domain Recommendation Via User-Clustering and Multidimensional Information FusionabstractRecently, recommendation systems have been widely usedin online business scenarios, which can improve the online experience by learning the user or item characteristics to predict the user’s future behavior and to realize precision marketing. However, data sparsity and cold-start problems limit the performance of recommendation systems in some emerging fields. Thus, cross-domain recommendation has been proposed to handle the abovementioned problems. Nonetheless, many cross-domain recommendations only consider modeling a single user’s representation and ignore user-group information (this group has similar behavior and interests). Additionally, most studies are based on matrix factorization for generating embeddings, which results in a weak generalization ability of user latent features. In this paper, we propose a novel cross-domain recommendation model viaUser-Clustering andMultidimensional informationFusion (UCMF) that attempts to enhance user representation learning in a data sparsity scenario for accurate recommendation. In addition, we consider a user’s individual information and cross-domain feature information. A novel multidimensional information fusion is proposed to guarantee the robustness of the user features. In particular, we apply a graph neural network to learn the user-group features, which can effectively save the correlation among users’ information and guarantee feature performance. In other words, the Wasserstein autoencoder is utilized to learn the cross-domain user features, which can guarantee the consistency of user features from different domains. Experiments conducted on real-world datasets empirically demonstrate that our proposed method outperforms the state-of-the-art methods in cross-domain recommendation. Jie Nie, Zian Zhao, Lei Huang 0010, Weizhi Nie, Zhiqiang Wei 0002 |
IEEE Trans. Multim. | 3 |
| 2023 | Cross-scale Graph Interaction Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing (RS) images plays a vital role in a variety of fields, including urban planning, natural disaster monitoring, and land resource management. Due to the complexity and low resolution of RS images, many approaches have been proposed to handle the related task. However, these previously developed approaches dedicate to contextual interaction but ignore the cross-scale semantic correlation and multi-scale boundary information. Therefore, we propose a Cross-scale Graph Interaction Network (CGIN) to address semantic segmentation problems of RS images, which consists of a semantic branch and a boundary branch. In the semantic branch, we first apply atrous convolution to extract multi-scale semantic features of RS images. Particularly, based on the multi-scale semantic features, a Cross-scale Graph Interaction (CGI) module is introduced, which establishes cross-scale graph structures and performs adaptive graph reasoning to capture the cross-scale semantic correlation of RS objects. In the boundary branch, we propose a Multi-scale Boundary Feature Extraction (MBFE) module that utilizes atrous convolutions with different dilation rates to extract multi-scale boundary features. Finally, to address the problem of sparse boundary pixels in the fusion process of the two branches, we propose a Multi-scale Similarity-guided Aggregation (MSA) module by calculating the similarity of semantic features and boundary features at the corresponding scale, which can emphasize the boundary information in semantic features. Our proposed CGIN outperforms state-of-the-art approaches in numerical experiments conducted on two benchmark remote sensing datasets. Jie Nie, Lei Huang 0010, Xiaowei Lv, Rui Wang 0111 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | RPITN: Review Based Preference Invariance Transfer Network for Cross-Domain RecommendationabstractCross-domain recommendation is an effective way to cope with the cold-start problem in recommendation systems. Knowledge of the current, particularly reviews, is taken into account to improve user/item embedding to reduce the neg-ative transfer that occurs during mapping processes across the source and target domains. Traditional approaches, on the other hand, typically apply review information from the source and target domain independently without consideration of user preference divergence. In this paper, we propose a novel Review-based Preference Invariance Transfer Network (RPITN) to minimize negative transfer by combining reviews from two domains. We first build a review preference invari-ance (RPI) embedding procedure to express user/item review correlations between two domains. Then, to improve the gen-eralization ability of user/item embedding and prevent negative transfer across domains, we carefully insert RPI into the embedding learning and mapping process. Extensive exper-iments on real-world datasets demonstrate the superiority of RPITN compared with other recommendation methods. Zijie Zuo, Jie Nie, Zian Zhao, Huaxin Xie, Xiangqian Ding, Shusong Yu, Lei Huang 0010, Yuxuan Yue, Xin Wang 0019 |
ICME | 7 |
| 2022 | Learning to Classify Weather Conditions from Single Images Without Labels
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Zhiqiang Wei 0002 |
MMM (1) | 2 |
| 2022 | WCATN: Unsupervised deep learning to classify weather conditions from outdoor images
Kezhen Xie, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | A CNN-based multi-task framework for weather recognition with multi-scale weather cues
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Lei Lyu 0001 |
Expert Syst. Appl. | 2 |
| 2022 | Exploiting reliable pseudo-labels for unsupervised domain adaptive person re-identification
Lei Huang 0010, Wenfeng Zhang, Zhiqiang Wei 0002 |
Neurocomputing | 2 |
| 2022 | Learning discriminative features for semi-supervised person re-identification
Huanhuan Cai, Lei Huang 0010, Wenfeng Zhang, Zhiqiang Wei 0002 |
Multim. Tools Appl. | 2 |
| 2022 | Learning camera invariant deep features for semi-supervised person re-identification
Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Huanhuan Cai |
Multim. Tools Appl. | 2 |
| 2022 | A Twofold Convolutional Regression Tracking Network With Temporal and Spatial MechanismabstractIn recent years, convolutional regression trackers have shown increasing attention for visual object tracking due to their favorable performance and easy implementation. However, most of them are restricted to features from a certain layer and hardly benefit from temporal spatial information, which limits the potential to significant appearance changes. In this work, we go beyond the traditional deep regression trackers and build a novel twofold tracking network, which exploits rich hierarchical features and incorporates both temporal and spatial information to boost the tracking performance. The proposed network is composed of two streams, i.e., an appearance stream and a semantic stream, each stream is independently learned from different convolutional layers. Specially, we propose temporal and spatial mechanism for robust target representation by considering historical information in previous frames as well as spatial information. By design, the proposed twofold convolutional regression tracking network with spatial and temporal mechanism can better tolerate the target appearance changes and improve the tracking accuracy. Extensive experimental results on the benchmarks OTB-2015, Temple-Color, UAV123, and VOT-2018 demonstrate the effectiveness of our method, as compared with a number of state-of-the-art trackers. Lei Huang 0010, Zhiqiang Wei 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Subtler mixed attention network on fine-grained image classification
Chao Liu 0008, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang |
Appl. Intell. | 2 |
| 2021 | Online parallel framework for real-time visual tracking
Lei Huang 0010, Guanqun Wei, Zhiqiang Wei 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2021 | Center-aligned domain adaptation network for image classification
Guanqun Wei, Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie |
Expert Syst. Appl. | 3 |
| 2021 | Appearance feature enhancement for person re-identification
Wenfeng Zhang, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie |
Expert Syst. Appl. | 2 |
| 2021 | Deep top similarity hashing with class-wise loss for multi-label image retrieval
Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Wenfeng Zhang |
Neurocomputing | 3 |
| 2021 | Unsupervised Deep Quadruplet Hashing with Isometric Quantization for image retrieval
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Kezhen Xie, Jinkui Hou |
Inf. Sci. | 2 |
| 2021 | Multi-task learning with deformable convolution
Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Person re-identification based on multi-appearance model
Lei Huang 0010, Wenfeng Zhang, Jie Nie, Zhiqiang Wei 0002 |
Multim. Tools Appl. | 1 |
| 2021 | Weakly supervised fine-grained recognition based on spatial-channel aware attention filters
Nannan Yu, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang |
Multim. Tools Appl. | 2 |
| 2021 | Adaptive multi-branch correlation filters for robust visual tracking
Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Zhineng Chen |
Neural Comput. Appl. | 2 |
| 2021 | Graph convolutional networks with attention for multi-label weather recognition
Kezhen Xie, Zhiqiang Wei 0002, Lei Huang 0010, Qibing Qin, Wenfeng Zhang |
Neural Comput. Appl. | 3 |
| 2021 | Angular regularization for unsupervised domain adaption on person re-identification
Wenfeng Zhang, Lei Huang 0010, Zhiqiang Wei 0002, Qibing Qin, Lei Lv |
Neural Comput. Appl. | 2 |
| 2021 | Sliced Wasserstein based Canonical Correlation Analysis for Cross-Domain Recommendation
Zian Zhao, Jie Nie, Lei Huang 0010 |
Pattern Recognit. Lett. | 4 |
| 2021 | Unsupervised Deep Multi-Similarity Hashing With Semantic Structure for Image RetrievalabstractWith the advance of Convolutional Neural Network, deep hashing methods have shown the great promising performance in large-scale image retrieval. Without depending on extensive human-annotated data, unsupervised hashing is more applicable to image retrieval tasks compared to supervised methods. However, due to the lack of fine-grained supervised signals and multi-similarity constraints, most state-of-the-art unsupervised deep hashing algorithms cannot ensure the correct fine-grained similarity ranking for image pairs. In this paper, we propose a novel unsupervised deep multi-similarity hashing framework to learn compact binary codes by jointly exploiting global-aware and spatial-aware representations, called Unsupervised Deep Multi-Similarity Hashing with Semantic Structure (UDMSH). Specifically, to obtain distinguishing characteristics, we develop a sub-network by jointly learning global semantic structures from Convolutional Neural Network (CNN) and inherent spatial structures from Fully Convolutional Network (FCN). By computing the cosine distance for deep features from image pairs, we construct a similarity matrix with semantic structure, then utilize this matrix to guide hash code learning process. Based on it, we carefully design a multi-level pairwise loss to preserve the correct fine-grained similarity ranking. Furthermore, we introduce Hamming-isometric mapping into unsupervised hashing framework to decrease the quantization errors. Extensive experiments on three widely used benchmarks prove that our proposed UDMSH outperforms several state-of-the-art unsupervised hashing with respect to different evaluation metrics. Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Kezhen Xie, Wenfeng Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Deep multilevel similarity hashing with fine-grained features for multi-label image retrieval
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002 |
Neurocomputing | 2 |
| 2020 | Adaptive Attention-Aware Network for unsupervised person re-identification
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Qibing Qin |
Neurocomputing | 3 |
| 2019 | A Novel Deep Hashing Method with Top Similarity for Image RetrievalabstractDue to the advantages of retrieval speed and storage space, deep hashing methods have become a research hotspot in the field of large-scale image retrieval. Most of existing deep hashing methods pay close attention to similarity between images without images at the top of the ranking list similar to query targets. In the paper, a novel deep hashing model is proposed to preserve top images similar to the query images and optimize the quality of hash codes for image retrieval. Specifically, the optimized AlexNet is utilized to extract discriminative image representations and learn hashing functions simultaneously. The loss function based on acceleration strategy is designed to ensure similarity between returned images at the top of the ranking list and query images. In addition, we implement the model training in a batch-process fashion to low the image storage. Moreover, our extensive experiments on standard benchmarks demonstrate that our method outperforms several state-of-the-art deep hashing methods. Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Xiaopeng Ji |
ICASSP | 3 |
| 2019 | Person Re-Identification Based on Pose-Aware Segmentation
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Lei Lv, Guanqun Wei |
MMM (2) | 3 |
| 2019 | Understanding personality of portrait by social embedding visual features
Jie Nie, Zhiqiang Wei 0002, Zhen Li 0024, Yan Yan 0003, Lei Huang 0010 |
Multim. Tools Appl. | 5 |
| 2018 | Real-Time Underwater Fish Tracking Based on Adaptive Multi-Appearance ModelabstractTracking live fish in an open underwater environment to investigate their behavior is of great value for many applications, e.g. biological and robotic research. However, tracking fish in real world environment is a challenging task due to complex non-rigid deformation and abrupt movement of fish. In this paper, we explore and incorporate motion property of fish and propose a real-time fish tracking method based on novel adaptive multi-appearance models and tracking strategy, which can be adapted to various changes of the fish appearance caused by non-rigid deformation. Experimental results show the promising performance of the proposed method can outperform the previous method by 13.4% in accuracy on average and is robust to real-time underwater fish tracking. Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Wenfeng Zhang |
ICIP | 3 |
| 2017 | Inferring intrinsic correlation between clothing style and wearers' personality
Zhiqiang Wei 0002, Yan Yan 0003, Lei Huang 0010, Jie Nie |
Multim. Tools Appl. | 3 |
| 2017 | Human body segmentation based on shape constraint
Lei Huang 0010, Jie Nie, Zhiqiang Wei 0002 |
Mach. Vis. Appl. | 1 |
| 2016 | How to record the amount of exercise automatically? A general real-time recognition and counting approach for repetitive activitiesabstractExercise is considered as an effective mean against overweight and obesity-related diseases. In this paper, a real-time activity recognition and counting approach is proposed to evaluate amount of exercise only using a wearable smart watch. First, accelerometer and gyroscope data are collected to extract efficient features. Then Support Vector Machine classifiers are trained to recognize nine common exercise activities in real time. In order to measure the frequency of repetitive activity, a general activity counting algorithm based on gyroscope is proposed which is applicable for different types of activity. Various activities can be counted uninterruptedly using the proposed general method without frequently changing algorithms. Through experiments, it is demonstrated that the extracted features are efficient for real time exercise activity recognition. Moreover, our comparative experiments have shown that our counting approach is more accurate than other products on the market. Shugang Zhang, Zhen Li 0024, Jie Nie, Lei Huang 0010, Zhiqiang Wei 0002 |
BIBM | 4 |
| 2016 | Exploring Relationship Between Face and Trustworthy Impression Using Mid-level Facial Features
Yan Yan 0003, Jie Nie, Lei Huang 0010, Zhen Li 0024, Qinglei Cao, Zhiqiang Wei 0002 |
MMM (1) | 3 |
| 2015 | Is Your First Impression Reliable? Trustworthy Analysis Using Facial Traits in Portraits
Yan Yan 0003, Jie Nie, Lei Huang 0010, Zhen Li 0024, Qinglei Cao, Zhiqiang Wei 0002 |
MMM (2) | 3 |
| 2015 | Robust skin detection in real-world images
Lei Huang 0010, Zhiqiang Wei 0002, Bo-Wei Chen, Chenggang Yan 0001, Jie Nie, Jian Yin 0003, Baochen Jiang |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | How Your Portrait Impresses People?: Inferring Personality Impressions from Portrait ContentsabstractWhenever looking at a stranger's portrait, besides observable appearance, we always build a personality impression implicitly in our subconscious. It is quite interesting to ask how a portrait impresses people. This paper presents a novel method to infer personality impression from portrait. Firstly, a questionnaire is applied to demonstrate the consistence of people's impression. And then personality-related features are explored through the statistical analysis method. Finally, features are trained using Support Vector Machine. Experimental results demonstrate our method could achieve a precision of 52.14% and a recall of 52.78% on inferring 4 personalities from 2,463 randomly selected portraits of people downloaded from "Google images". Improvements of 44.04% and 37.91% are reported compared to a baseline method. And features contribution analysis deeply unveils the correspondence between portrait contents and personality impressions. Demonstrations with respect to visual patterns in portrait collages of different personalities further prove the effectiveness of the proposed method. Furthermore, we apply our method to analyze portraits of Hillary Clinton and obtain an interesting multifaceted figure of this famous politics, which is another proof of both our concept and method. Jie Nie, Peng Cui 0001, Yan Yan 0003, Lei Huang 0010, Zhen Li 0024, Zhiqiang Wei 0002 |
ACM Multimedia | 4 |
| 2014 | Finding suits in images of people in unconstrained environments
Chenggang Yan 0001, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Bochuan Chen, Yingping Zhang |
J. Vis. Commun. Image Represent. | 2 |