Wenfeng Zhang

dblp:82/2947 · DBLP profile ↗
← Back
80ranked-venue papers
11as first author
71since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 5 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 2 first-author · 32 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Polysemic Semantic Instance Network for Cross-Modal Hashing
abstract
Hashing techniques are widely adopted in large-scale cross-modal retrieval due to their efficiency and low storage cost. However, semantic ambiguities, including polysemy, multi-object images, and missing semantic descriptions, significantly degrade the accuracy of alignment and retrieval performance. Most existing methods rely on one-to-one mappings that preserve only global average semantics, which fail to capture the intrinsic polysemous structures embedded within individual samples. To address this issue, we propose a novel Deep Polysemic Semantic Instance Hashing (DPSIH) method and design a Diverse Semantic Instance Embedding (DSIE) module. This module integrates local and global features through multi-head self-attention and residual learning, generating multiple diverse embeddings per sample to effectively capture fine-grained and polysemous semantic structures. Furthermore, we design a multi-embedding semantic correlation constraint that relaxes strict alignment restrictions to improve robustness under partial alignment, and introduce Maximum Mean Discrepancy (MMD) regularization to alleviate cross-modal distribution shifts. Additionally, an embedding diversity mechanism is proposed to prevent all embeddings from collapsing into a central or averaged representation, thereby enhancing semantic diversity. Extensive experiments on four benchmark datasets demonstrate that DPSIH significantly outperforms state-of-the-art methods and effectively improves the modeling of semantic ambiguity in cross-modal retrieval tasks.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Lei Huang 0010
AAAI4
2026 MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
abstract
With the burgeoning development of fields such as the Metaverse, Virtual Reality (VR), and Digital Twins, text-to-3D generation has emerged as a research hotspot in both academia and industry. Currently, optimization methods based on Score Distillation Sampling (SDS) utilizing 2D diffusion priors have become the mainstream technological paradigm in this field. However, due to the view bias of 2D priors and the mode-seeking ambiguity combined with gradient noise induced by high Classifier-Free Guidance (CFG), these methods still suffer from macro-topological inconsistency (e.g., the Janus problem) and micro-geometric discontinuity. To address these challenges, we propose MOC-3D, a text-to-3D generation method based on geometric manifold and semantic view-order consistency. Built upon the ScaleDreamer framework, our method incorporates a Semantic View-Order Constraint Module and a Manifold-based Feature Continuity Module. The former aims to rectify macro-topological inconsistency, while the latter focuses on eliminating micro-geometric discontinuity. Specifically, the Semantic View-Order Constraint Module leverages the prior knowledge of CLIP to impose a Monotonicity Rank Constraint on semantic score representations across different views, thereby providing effective guidance for the global topological structure of 3D objects. Meanwhile, the Manifold-based Feature Continuity Module employs the Riemannian Metric on the Symmetric Positive Definite (SPD) manifold. By measuring the distance of feature statistical distributions in the Riemannian space, it promotes the smooth evolution and continuity of micro-textures across multi-views in a statistical sense. Under the macro-micro synergistic optimization of these two modules, our model can simultaneously improve macro-structural consistency and micro-detail continuity. Experimental results demonstrate that compared with mainstream methods, our approach achieves significant advantages in terms of Semantic Consistency (CLIP Score) and Perceptual Quality (LPIPS). Furthermore, ablation studies verify the independent contributions and complementary effectiveness of the Semantic View-Order Constraint Module in rectifying macro-topological inconsistency and the Manifold-based Feature Continuity Module in eliminating micro-geometric discontinuity.
Chenyang Fan, Wen Yang 0003, Junshi Cheng, Zihong Li, Wenfeng Zhang, Pan Zeng
ICMR5
2026 Deep Potential Semantic-aware Hashing for Cross-modal Retrieval
abstract
Hashing learning has moved into the mainstream for multimedia retrieval because it offers the advantages of low storage cost and high retrieval efficiency. Currently, most cross-modal hashing methods commonly explore the similarity relations between samples by constructing pair-wise or triplet-wise constraints. However, these methods focus on the relative correct ranking of samples, ignore the potential semantic similarity of raw sample distribution, and generate sub-optimal hash codes. To resolve this issue, the novel Deep Potential Semantic-aware Hashing framework (DPSaH) is proposed to mine the local semantic structure of heterogeneous samples, maintaining inter-modality-consistent and cross-modality-correlated semantic relationships. Specifically, by exploring the potential local structure of the data, the multi-modal quadruple loss is extended to the cross-modal hashing framework, thereby preserving the potential semantic neighborhoods among raw samples in Hamming space. During model training, based on the average semantic labels, the label-averaged balanced strategy is developed to quantify the frequency difference between positive and negative samples. Besides, by injecting noise information into the generated discrete codes, the binary-injection loss is introduced to alleviate the over-activation of specific bits, decorrelating different bits in the Hamming space. Extensive experiments are performed on three public datasets, and the results verify the superiority of the DPSaH framework compared to the current mainstream cross-modal hashing frameworks. The source code for DPSaH is available at https://github.com/QinLab-WFU/DPSaH .
Qibing Qin, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Eng. Appl. Artif. Intell.5
2026 Deep asymmetric semantic hashing with probability shifting for multi-label image retrieval
Yongyue Fu, Qibing Qin, Jinkui Hou, Congcong Zhu, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.6
2026 Deep semantic center-guided hashing for multi-label cross-modal retrieval
Xinzheng Sui, Yadong Huo, Qibing Qin, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.6
2026 FPM-GAN: Craniofacial reconstruction method based on frequency domain perception and multi-scale attention
Wen Yang 0003, Longqian Ma, Dengwei Yan, Zhengran Cao, Wenfeng Zhang, Guohua Geng
Expert Syst. Appl.6
2026 Deep neighborhood-based component proxy hashing for large-scale image retrieval
Huiying Zhu, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.5
2026 Deep synthetic-proxy hashing for multi-label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Chunlei Chen, Lei Huang 0010
Neurocomputing4
2026 Few-shot medical image segmentation via dual-stream feature extractor and detail-enhanced prototype transformer
Wenfeng Zhang, Jianming Hu, Qibing Qin
Knowl. Based Syst.2
2026 Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001
Medical Image Anal.43
2026 Deep Softtriple hashing for Multi-Label cross-modal retrieval
Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Neural Networks4
2026 Deep neighbor-aware hashing with global-local representation for multi-label remote sensing image retrieval
Xiaorong Chen, Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Signal Process. Image Commun.6
2026 Deep noise-tolerant hashing for remote sensing image retrieval
abstract
Currently, how to quickly retrieve target images from large-scale remote sensing data has emerged as a critical challenge in the context of explosive growth of remote sensing data volume. To deal with this challenge, hash learning becomes an ideal choice with its low storage cost and high efficiency. In recent years, the combination of hash learning with deep neural networks such as CNNs and Transformers has resulted in numerous frameworks demonstrating excellent performance. However, in the field of remote sensing image hashing, previous studies cannot simultaneously consider the effect of noise in feature extraction and loss optimization, so that their retrieval performance is greatly reduced due to noise interference. To resolve the mentioned problem, a Deep Noise-tolerant Hashing (DNtH) framework is proposed to learn the sample complexity and noise level, and adaptively reduce the weight of noisy information. Specifically, to realize the extraction of fine-grained features from information containing irrelevant samples, the noise-aware Transformer is proposed by introducing the patch-wise attention and depth-wise convolution. To reduce the interference of noisy labels on remote sensing image retrieval, an adaptive active-passive loss framework is proposed to dynamically adjust the weights of active passive loss, which learns the weight parameters through a dynamic weighted network while combining with asymmetric strategy for effective compact representation learning. The ratio of entropy to standard deviation and the probability difference are input into the above network and trained with the feature extraction network. Extensive experiments on three publicly available datasets show that the DNtH framework can adapt to noisy environments while achieving optimal performance in remote sensing image retrieval. The source code for the implementation of our DNtH framework is available at https://github.com/QinLab-WFU/DNtH.git .
Chunyu Yan, Qibing Qin, Jiangyan Dai, Wenfeng Zhang
Signal Process. Image Commun.5
2026 Deep semantic channel hashing for large-scale image retrieval
Qibing Qin, Lei Huang 0010, Wenfeng Zhang
Signal Process. Image Commun.5
2026 Deep Stochastic Spherical Hashing With Von Mises-Fisher Distributions for Cross-Modal Retrieval
abstract
Deep cross-modal hashing has gained significant attention because of its benefits, including reduced storage requirements and enhanced retrieval efficiency. Although progress has been made, existing deep cross-modal hashing methods still face unresolved challenges. Most existing methods typically adopt Euclidean space as the embedding space to measure the semantic similarity between original samples. However, the volume of Euclidean space grows polynomially with dimension, which exacerbates the curse of dimensionality. In contrast, methods based on spherical space usually use cosine similarity as the metric, effectively mitigating the aforementioned problem by normalizing the embedding vectors. Nevertheless, such methods only considers the direction to determine the category, ignoring the uncertainty measure in the embedding space, thus having a limited ability to preserve inherent multimodal semantics. In this paper, with a novel extension of the maximum entropy distribution on the surface of a hypersphere von Mises-Fisher (vMF) distribution, a novel deep cross-modal hashing method, named Deep Stochastic Spherical Hashing (DSSH), is designed to utilize uncertain information to guide the hashing process and produce discriminative modality-invariant hash codes. Specifically, to learn explicit uncertainty in learned embedding space, the Spherical von Mises-Fisher distribution is applied for the f irst time in deep cross-modal hashing, where the direction of the sample embedding controls its position on the hyper sphere, thereby preventing its semantic content, and its norm parameterizes the determinism of the distribution. In addition, stochastic spherical von Mises–Fisher loss is proposed to preserve the mode-specific semantic information of the sample, achieving the alignment of different modalities and semantic embeddings. Extensive experiments on four benchmark datasets show that our DSSH framework outperforms existing state-of-the-art cross modal hashing methods. The source code of the experiments is available at https://github.com/QinLab-WFU/DSSH.
Qibing Qin, Meiling Ge, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Knowl. Data Eng.3
2026 Deep Distance Weighted Sampling Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing seeks to encode heterogeneous image-text data into compact binary codes for efficient retrieval. While significant efforts have been made to study sampling strategies, most of these approaches are tightly integrated with loss function engineering, lacking an independent focus on sampling methods. In this paper, we challenge the convention by revealing that batch-level sampling strategy is equally pivotal as loss design for learning discriminative hash codes. Specifically, by introducing a novel distribution-aware sampling strategy, a Distance Weighted Sampling Hashing (DDWSH) framework is proposed to dynamically select stable and informative training pairs. Unlike conventional random or semi-hard sampling, our method weights pairwise distances within each batch to approximate global data distribution, thereby mitigating training instability caused by biased sampling. To rigorously validate our claims, we conduct the first comprehensive crossover study between sampling strategies (random/semi-hard/ours) and loss functions (contrastive/triplet/ours) across three benchmark datasets. Experiments demonstrate that: Universality: Our sampling boosts all loss functions' performance (average +4.3% mAP vs. semi-hard mining), Superiority: DDWSH competes with complex loss function design and achieves state-of-the-art results, and Stability: It reduces performance variance by 60.2% compared to semi-hard sampling under varying batch compositions average. This systematic analysis establishes sampling as an independent research dimension in deep hashing, beyond a mere part of loss function engineering. The source code for DDWSH is freely available athttps://github.com/QinLab-WFU/DDWSH.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.3
2026 Deep Neighbor Discriminant Binary Embedding for Multi-Label Image Retrieval
abstract
Because of fast retrieval speed and low storage cost, deep hash learning has become a current research spot in multimedia retrieval, which arouses more and more interest and attention. Most of the available deep hashing methods usually employ the mini-batch strategy to construct training batches, and only one mini-batch of samples is valuable at each iteration during model optimization, which fails to explore the original neighbor structure well, leading to sub-optimal embeddings, especially for relatively large datasets. By contrast, the superior performance is achieved by optimizing the SoftMax loss for certain hashing learning, meanwhile, prior research has suggested the normalized SoftMax loss is essentially identical to one smoothed triplet-wise function, in which each class is assigned to one single semantic center. Nevertheless, under real-world scenarios, one class could correspond to multiple local clusters rather than just a single one. To this end, by extending SoftMax loss with multiple semantic centers, a novel deep hashing framework, called Deep Neighbor Discriminant Binary Embedding (NDBE) framework, is presented to generate discriminative hash codes with the original neighbor structure preservation. Specifically, by expanding the size of the last fully connected layer to increase multiple centers for each class, a novel SoftTriple loss is proposed to capture the latent semantic distribution of the original samples and reduce the intra-class variance, which is optimized without constraints sampling. To learn the different numbers of each class center, the class-center aggregation strategy is developed to obtain the compact set of centers while maintaining semantic similarity. Extensive experiments on three public multi-label datasets show that our proposed NDBE framework achieves superior visual similarity search performance over several state-of-the-art approaches. The source code for the implementation of our proposed NDBE framework is available athttps://github.com/QinLab-WFU/NDBE.
Qibing Qin, Mingkun Dou, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.3
2026 Deep Semantic Tuplet-Based Hashing by Hypergraph Modeling for Cross-Modal Retrieval
abstract
With low storage cost and high retrieval efficiency, hashing techniques are widely used for multi-media retrieval, which has already become the present research focus. Currently, cross-modal hashing commonly employs graph-based loss to construct pair-wise semantic relations between training samples for model optimization. However, limited by the graph-based strategy, each edge in the graph only connects two samples, which only represent a bundle of pair-wise relationships. Besides, the edges in the graph are calculated by self-attention or feature distance, only considering pair-wise relations of heterogeneous samples and ignoring the class relations. In this paper, by hypergraph modeling the semantic tuples, a novel Deep Semantic Tuplet-based Hashing by Hypergraph Modeling (DSTH) approach is proposed to leverage the multilateral semantic relations, which could guide the model to learn class-discriminative semantic binary embedding. In more detail, based on the characteristic distribution, semantic tuples are constructed for each class in one mini-batch, which represents the multilateral semantic relationships between multiple samples and multiple classes. By considering semantic tuples as hyperedges to represent multilateral semantic relations, hypergraph modeling is designed, in which HyperGraph Neural Hetwork (HGNH) is introduced to formulate hypergraph node classification goals to fully learn the multilateral semantic information contained in the semantic tuples. Moreover, to utilize the heterogeneity of local structures in embedding, the adaptive neighborhood structure is explored by learning the structure embedding, which provides fine-grained ranking lists. Through extensive experiments on three benchmark datasets, the comprehensive results validate the advancement of our proposed DSTH framework over mainstream cross-modal hashing. The source code for the framework DSTH is freely available athttps://github.com/QinLab-WFU/DSTH.
Qibing Qin, Wenfeng Zhang, Huihui Zhang 0003, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.3
2026 Generative Zero-Shot Hashing for Multi-Label Image Retrieval
abstract
Due to the preferable efficiency of storage and computation, hashing algorithms show great potential. With the explosive growth of web data, new and emerging categories of multimedia data continue to increase, and zero-shot hashing has been one of the hot research topics in retrieval research. Nevertheless, existing zero-shot hashing methods focus mainly on single-label image retrieval, using class semantic embeddings as a link between visual features and binary codes to align visual features with corresponding class semantics and simultaneously transfer knowledge from seen classes to unseen classes. Meanwhile, single-label zero-shot learning employs Generative Adversarial Networks (GANs) to synthesize class-specific characteristics from the corresponding class attribute embeddings, achieving encouraging results. However, synthesizing multi-label features from GANs remains unexplored in the context of zero-shot hashing settings. When multiple objects co-appear in an image, a key issue is how to fuse multi-class information effectively. In this article, by introducing the multi-class information fusion strategy, a novel Generative Zero-Shot Hashing (GZSH) is proposed to transform zero-shot hashing into traditional supervised hashing by generating features of unseen categories. Specifically, this study introduces three different fusion methods (attribute-based fusion, feature-based fusion, and semantic-enhanced fusion) for synthesizing multi-label features from the corresponding multi-label class embeddings. Subsequently, the semantic-enhanced fusion is integrated into the representative generative architecture to learn the feature distribution of unlabeled images in the context of multi-label zero-shot hashing. Besides, by jointly learning the seen/source and unseen/target samples, a pairwise similarity loss is introduced to optimize the deep hashing model. Extensive experimentation on three widely used multi-label datasets illustrates the outstanding performance of our proposed GZSH framework compared to current state-of-the-art methods. The implementation code for our GZSH framework can be found at https://github.com/QinLab-WFU/GZSH .
Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Deep Uncertainty-aware Probabilistic Hashing for Cross-modal Retrieval
abstract
Due to its outstanding computational efficiency and low storage requirements, hashing technology has become a research hotspot in large-scale multimedia retrieval. In cross-modal hashing, the key lies in mapping samples from different data sources into a common discrete space. However, most existing methods assume that the input data is of high quality and completeness. When samples are incomplete or degraded (e.g., blurry images or incomplete text), the absence or ambiguity of semantic information inevitably compromises retrieval accuracy. Traditional deterministic embedding methods typically map multi-modal samples to a single point in the embedding space without considering uncertainty. As a result, inherent noise or feature ambiguity in the inputs may lead to distorted or shifted binary representations. To resolve this problem, this article proposes a novel Deep Uncertainty-aware Probabilistic Hashing (DUaPH) method that models the uncertainty of multi-modal samples. By capturing the underlying distribution of heterogeneous data, DUaPH effectively mitigates feature ambiguity and inherent noise, enhancing the robustness of cross-modal retrieval. Specifically, each heterogeneous sample is mapped to a multivariate Gaussian distribution, where the mean represents the most probable semantic features, and the variance reflects the sample uncertainty. A semantic feature matching mechanism is introduced to dynamically adjust the importance of feature dimensions, prioritizing those with higher certainty. Then, a semantic feature fusion mechanism is developed to integrate the semantic features from multi-modal sample pairs, producing a new distribution with reduced uncertainty and improved semantic alignment. Extensive experiments on four benchmark datasets demonstrate that DUaPH significantly improves robustness and retrieval performance under conditions of semantic ambiguity and data uncertainty. The source code is available at https://github.com/QinLab-WFU/DUaPH .
Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Deep Relational Knowledge Distillation Hashing via Relaxed Masking Triplet Optimization for Large-scale Image Retrieval
abstract
Large-scale image retrieval increasingly depends on triplet-based deep hashing methods, which typically construct training triplets using class labels. However, existing triplet-based hashing methods relying on sparse and noisy categorical labels suffer from two main issues: no triplet, where sparse annotations limit valid triplet formation, and bad triplet, where noisy or ambiguous labels cause misleading samples, which together hinder the network from learning a well-structured similarity space, causing semantically similar images to scatter and dissimilar ones to collapse, ultimately undermining retrieval accuracy. In this article, we solve this dilemma with a novel unified deep hashing framework, termed Deep Relational Knowledge Distillation Hashing (DRKDH), which decouples semantic relation modeling from visual feature learning by leveraging a teacher-student paradigm to learn semantically consistent hash codes. Specifically, the teacher model captures comprehensive semantic structural relationships across all instances, providing information-rich and noise-resilient supervision to guide the student model in learning expressive and discriminative visual embeddings, effectively mitigating the impact of low-quality triplet labels. Furthermore, by leveraging similarity weights predicted by the teacher model, a relaxed masking triplet loss is introduced into the student model, which dynamically adjusts the contribution of each triplet based on its informativeness, suppressing invalid or misleading triplets while emphasizing valuable ones to enhance training efficiency. Comprehensive experiments on THINGS, ImageNet, MIRFLICKR-25K, and NUS-WIDE show that DRKDH consistently outperforms state-of-the-art deep hashing methods under sparse and noisy supervision, while additional evaluations suggest competitive performance in selected zero-shot and few-shot settings. Source codes: https://github.com/QinLab-WFU/DRKDH .
Qibing Qin, Wenfeng Zhang, Lei Huang 0010
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Enhanced radiology report generation via comprehensive sequence rearrangement and multi-scale cross-region attention
Qibing Qin, Jianming Hu, Dengwei Yan, Wenfeng Zhang, Jing Qiao
Vis. Comput.6
2025 Problem-Driven and Shape-Guided: Multi-scale Deform KAN for X-Shaped Anterior Visual Pathway Segmentation
Yongliang Han, Wenlong Lin, Yongmei Li, Fanghong Zhang, Binbin Sang, Tiansong Li, Wenfeng Zhang, Shaoguo Cui
ICANN (2)8
2025 Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal Retrieval
abstract
The field of cross-modal retrieval aims to construct a shared representation space for samples from multiple modalities, typically within the vision and language domains. Deep hashing, with its high computational efficiency and low storage costs, has emerged as a central focus in this field and has garnered significant attention in recent research. However, current hash retrieval, concentrating on deterministic methods, struggles to effectively capture semantically ambiguous correspondences between cross-modal samples, where heterogeneous data have complex-semantic many-to-many relationships in the latent space. To address this limitation, we propose a novel Deep Probabilistic Binary Embedding (DPBE) framework, designed to generate discriminative, modality-invariant hash codes that facilitate accurate and reliable cross-modal retrieval. In contrast to contemporary probabilistic methods, we focus on optimizing hash networks to learn more accurate binary embeddings by using the learning mode of probabilistic embeddings. We introduce the first Bayesian encoder for hash learning, which employs Laplace Approximation to model a distribution over network weights. Extensive experimental results demonstrate that our approach not only outperforms deterministic methods in retrieval performance but also provides uncertainty estimates, enhancing the interpretability of the embeddings. The corresponding code is available at https://github.com/QinLab-WFU/DPBE.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
ACM Multimedia3
2025 Factorized Transformer Hashing with Adaptive Routing for Large-scale Image Retrieval
abstract
Transformer architecture has driven significant advancements in deep hashing, establishing itself as a dominant framework for large-scale retrieval and storage applications. However, existing Transformer-based deep hashing methods typically employ unvarying feature transformations across all images, limiting their adaptability to diverse visual patterns. This rigidity restricts the model's capacity to learn both highly distinctive and generalizable discrete representations, posing challenges for retrieval in open-world scenarios. To overcome this challenge, we propose a novel Factorized Transformer Hashing (FTH) framework, which introduces a factorized transformer to enhance the generalization and discriminative power of hash codes. Specifically, we decompose the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) blocks into multiple sub-blocks, forming a transformer factorization scheme that captures diverse feature characteristics through independent sub-blocks. Furthermore, we develop an adaptive selection strategy, leveraging a set of learnable selectors with the Softmax function, to dynamically route each image to the most appropriate sub-block for processing. Extensive experiments on three benchmark datasets demonstrate that the proposed FTH framework significantly outperforms state-of-the-art baselines in both image hashing and zero-shot hashing tasks. Source code is available at https://github.com/QinLab-WFU/FTH.
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
ACM Multimedia3
2025 Ranking-oriented cross-modal hashing
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Eng. Appl. Artif. Intell.3
2025 Deep informative-triplet sampling hashing with attention-aware augmentation for remote sensing image retrieval
Meiling Ge, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.3
2025 Deep adaptive gradient-triplet hashing for cross-modal retrieval
Congcong Zhu, Jinkui Hou, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.5
2025 Deep neighbor-coherence hashing with discriminative sample mining for supervised cross-modal retrieval
Congcong Zhu, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
Expert Syst. Appl.3
2025 Personalized Dual Transformer Network for sequential recommendation
Meiling Ge, Chengduan Wang, Xueyang Qin, Jiangyan Dai, Lei Huang 0010, Qibing Qin, Wenfeng Zhang
Neurocomputing7
2025 Joint multi-grained similarity contrastive learning for video-text retrieval
Mingyong Li, Mingyuan Ge, Wenfeng Zhang
Neurocomputing3
2025 Deep Consistent Penalizing Hashing with noise-robust representation for large-scale image retrieval
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
Neurocomputing3
2025 Deep multi-similarity hashing via label-guided network for cross-modal retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Neurocomputing6
2025 Deep binary hyperbolic embedding for large-scale image retrieval
Enhao Wang, Qibing Qin, Jinkui Hou, Wenfeng Zhang, Lei Huang 0010
Neurocomputing5
2025 Domain-invariant multi-granularity feature learning for generalizable person re-identification
Wenfeng Zhang, Xiangfei Cao, Lei Huang 0010, Dengwei Yan, Qibing Qin
Knowl. Based Syst.1
2025 Deep Hardness-Aware Hashing for Large-Scale Image Retrieval
Chunping Dong, Meiling Ge, Enhao Wang, Qibing Qin, Wenfeng Zhang, Lei Huang 0010
IEEE Signal Process. Lett.5
2025 Deep Discriminative Boundary Hashing for Cross-Modal Retrieval
abstract
By the preferable efficiency in storage and computation, deep cross-modal has gained much attention in large-scale multimedia retrieval. Current deep hashing employs the probability outputs of the likelihood function, i.e., Sigmoid or Cauchy, to quantify the semantic similarity between samples in a common Hamming space. However, the inherent weakness of the Sigmoid likelihood function or the Cauchy likelihood function in gradient optimization leads to hashing models failing to exactly describe the hamming ball, which indicates the absolute semantic boundary among classes, thereby giving the high neighborhood ambiguity. In this paper, with the analysis of the likelihood function from the perspective of similarity metric learning, the novel Deep Discriminative Boundary Hashing framework (DDBH) is proposed to learn the discriminative embedding space that separates neighbors and non-neighbors well. Specifically, by introducing the remapping strategy and the base-point adaptive selection, the boundary-preserving loss based on the adjustable likelihood function is proposed to project data points with small gradients to regions with large gradients and give larger gradients for hard samples, facilitating better separation among classes. Meanwhile, to learn class-dependent binary codes, the class-wise quantization loss is designed to heuristically transfer class-wise prior knowledge to the binary quantization, significantly improving the discriminative capability of compact discrete codes. Comprehensive experiments on three benchmark datasets show that our proposed DDBH framework outperforms other representative deep cross-modal hashing. The corresponding code is available at https://github.com/QinLab-WFU/DDBH.
Qibing Qin, Yadong Huo, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Circuits Syst. Video Technol.3
2025 Deep Semantic-Consistent Penalizing Hashing for Cross-Modal Retrieval
abstract
Benefiting from the advantages of low storage cost and high retrieval efficiency, hash learning could significantly speed up large-scale cross-modal retrieval. Based on the prior annotations, most of the available cross-modal hashing usually introduces the margin-based constraint to generate different boundaries for each class in the inference phase, optimizing the model. However, these obtained label-guided penalty boundaries may differ from the primitive semantic relationships between heterogeneous modalities, impairing retrieval performance. Besides, the margin-based constraint is too weak to penalize the classes with low intra-class variances or inter-class correlations, which struggle to learn high-quality embeddings. In this paper, we propose a novel Deep Semantic-consistent Penalizing Hashing framework (DScPH) to learn the consistent penalizing fields for all classes, achieving accurate and efficient cross-modal retrieval. Specifically, by exploring unbalanced intra-class and inter-class correlations, the consistent penalizing loss is introduced into cross-modal retrieval to learn the consistency decision boundaries across classes. During training, the dice-like optimization strategy is developed to balance the pulling penalizing elements and pushing penalizing elements, facilitating the model convergence. Besides, based on the invariance of similarity measures under orthogonal transformations, the alternative quantization is proposed to minimize the errors between the learned continuous embeddings and binary discretization, maintaining the consistency of semantic relationships after performing binary projection. Extensive experiments are conducted on three benchmark datasets, and the comprehensive results validate the efficacy of our proposed DScPH framework, which outperforms the current mainstream deep cross-modal hashing algorithms.
Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Multim.3
2025 Modality-aware graph CNN for cross-modal person reidentification
Ruisheng Ran, Wenfeng Zhang, Qibing Qin
Vis. Comput.4
2024 Research on the Cultivation of Students' Innovative Ability Based on the Collaboration of Practical Activities
abstract
In recent years, the enthusiasm of students to participate in practical activities has increased, and the collaborative management of various types of practical activities has been given attention. Combined with the requirements of engineering education accreditation standards to cultivate students' ability to solve complex engineering problems, a comprehensive management method of practical activities based on quality management system (QMS) is proposed. By establishing a standard document system for the implementation process of practical activities and introducing scientific management methods and innovative methods, it solves the problem of students' cooperation in practical activities across classes and grades on the one hand, and promotes scientific management and continuous improvement of practical activities on the other hand. Taking the major of electrical engineering as an example, the process of cultivating students' innovative ability based on the cooperation of practical activities is elaborated.
Wenfeng Zhang, Haisi Yang
ICALT1
2024 MeFD-Net: multi-expert fusion diagnostic network for generating radiology image reports
Ruisheng Ran, Renjie Pan 0002, Wenfeng Zhang, Qibing Qin
Appl. Intell.5
2024 Deep global semantic structure-preserving hashing via corrective triplet loss for remote sensing image retrieval
Qibing Qin, Jinkui Hou, Jiangyan Dai, Lei Huang 0010, Wenfeng Zhang
Expert Syst. Appl.6
2024 Representation Learning Based on Vision Transformer
abstract
In recent years, with the rapid development of information technology, the volume of image data has grown exponentially. However, these datasets typically contain a large amount of redundant information. To extract effective features and reduce redundancy from images, a representation learning method based on the Vision Transformer (ViT) has been proposed, and to our best knowledge, Transformer was first applied to zero-shot learning (ZSL). The method adopts a symmetric encoder–decoder structure, where the encoder incorporates Multi-Head Self-Attention (MSA) mechanism of ViT to reduce the dimensionality of image features, eliminate redundant information, and decrease computational burden. Consequently, it effectively extracts features, and the decoder is utilized for reconstructing image data. We evaluated the representation learning capability of the proposed method in various tasks, including data visualization, image reconstruction, face recognition, and ZSL. By comparing with state-of-the-art representation learning methods, the outstanding results obtained validate the effectiveness of this method in the field of representation learning.
Ruisheng Ran, Qianwei Hu, Wenfeng Zhang, Shunshun Peng, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2024 Visual-Textual Cross-Modal Interaction Network for Radiology Report Generation
abstract
The radiology report generation task generates diagnostic descriptions from radiology images, aiming to alleviate the onerous task for radiologists and alerting them to abnormalities. However, the data bias problem poses a persistent challenge, since the abnormal regions usually occupy a small portion of radiology image, while the report generation process should pay greater attention to the abnormal regions. Moreover, the data volume is relatively small compared to large language models, posing challenges during training. To address these issues effectively, we propose a Visual-textual Cross-model Interaction Network (VCIN) to enhance the quality of generated reports. VCIN comprises two key modules: Abundant Clinical Information Embedding (ACIE), which gathers rich cross-modal interaction information to promote the report generation of abnormal regions; and a Bert-based Decoder-only Generator (BDG), built on Bert architecture to mitigate training difficulties. The superior performance of our proposed model is demonstrated through experimental results obtained from two public benchmark datasets. The code is available athttps://github.com/QinLab-WFU/VCIN.
Wenfeng Zhang, Baoning Cai, Jianming Hu, Qibing Qin, Kezhen Xie
IEEE Signal Process. Lett.1
2024 Deep Semantic-Aware Proxy Hashing for Multi-Label Cross-Modal Retrieval
abstract
Deep hashing has attracted broad interest in cross-modal retrieval because of its low cost and efficient retrieval benefits. To capture the semantic information of raw samples and alleviate the semantic gap, supervised cross-modal hashing methods that utilize label information which could map raw samples from different modalities into a unified common space, are proposed. Although making great progress, existing deep cross-modal hashing methods are suffering from some problems, such as: 1) considering multi-label cross-modal retrieval, proxy-based methods ignore the data-to-data relations and fail to explore the combination of the different categories profoundly, which could lead to some samples without common categories being embedded in the vicinity; 2) for feature representation, image feature extractors containing multiple convolutional layers cannot fully obtain global information of images, which results in the generation of sub-optimal binary hash codes. In this paper, by extending the proxy-based mechanism to multi-label cross-modal retrieval, we propose a novel Deep Semantic-aware Proxy Hashing (DSPH) framework, which could embed multi-modal multi-label data into a uniform discrete space and capture fine-grained semantic relations between raw samples. Specifically, by learning multi-modal multi-label proxy terms and multi-modal irrelevant terms jointly, the semantic-aware proxy loss is designed to capture multi-label correlations and preserve the correct fine-grained similarity ranking among samples, alleviating inter-modal semantic gaps. In addition, for feature representation, two transformer encoders are proposed as backbone networks for images and text, respectively, in which the image transformer encoder is introduced to obtain global information of the input image by modeling long-range visual dependencies. We have conducted extensive experiments on three baseline multi-label datasets, and the experimental results show that our DSPH framework achieves better performance than state-of-the-art cross-modal hashing methods. The code for the implementation of our DSPH framework is available athttps://github.com/QinLab-WFU/DSPH.
Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang
IEEE Trans. Circuits Syst. Video Technol.5
2024 S3-Net: A Self-Supervised Dual-Stream Network for Radiology Report Generation
abstract
Intelligent medicine is eager to automatically generate radiology reports to ease the tedious work of radiologists. Previous researches mainly focused on the text generation with encoder-decoder structure, while CNN networks for visual features ignored the long-range dependencies correlated with textual information. Besides, few studies exploit cross-modal mappings to promote radiology report generation. To alleviate the above problems, we propose a novel end-to-end radiology report generation model dubbed Self-Supervised dual-Stream Network (S3-Net). Specifically, a Dual-Stream Visual Feature Extractor (DSVFE) composed of ResNet and SwinTransformer is proposed to capture more abundant and effective visual features, where the former focuses on local response and the latter explores long-range dependencies. Then, we introduced the Fusion Alignment Module (FAM) to fuse the dual-stream visual features and facilitate alignment between visual features and text features. Furthermore, the Self-Supervised Learning with Mask(SSLM) is introduced to further enhance the visual feature representation ability. Experimental results on two mainstream radiology reporting datasets (IU X-ray and MIMIC-CXR) show that our proposed approach outperforms previous models in terms of language generation metrics.
Renjie Pan 0002, Ruisheng Ran, Wenfeng Zhang, Qibing Qin, Shaoguo Cui
IEEE J. Biomed. Health Informatics4
2024 Deep Hierarchy-Aware Proxy Hashing With Self-Paced Learning for Cross-Modal Retrieval
abstract
Due to its low storage cost and high retrieval efficiency, hashing technology is popularly applied in both academia and industry, which provides an interesting solution for cross-modal similarity retrieval. However, most existing supervised cross-modal hashing methods typically view the fixed-level semantic affinity defined by manual labels as supervised signals to guide hash learning, which only represents a small subset of complex semantic relations between multi-modal samples, thus impeding the hash function learning and degrading the obtained hash codes. In the paper, by learning shared hierarchy proxies, a novel deep cross-modal hashing framework, called Deep Hierarchy-aware Proxy Hashing (DHaPH), is proposed to construct the semantic hierarchy in a data-driven manner, thereby capturing the accurate fine-grained semantic relationships and achieving small intra-class scatter and big inter-class scatter. Specifically, by regarding the hierarchical proxies as learnable ancestors, a novel hierarchy-aware proxy loss is designed to model the latent semantic hierarchical structures from different modalities without prior hierarchy knowledge, in which similar samples share the same Lowest Common Ancestor (LCA) and dissimilar points have different LCA. Meanwhile, to adequately capture valuable semantic information from hard pairs, a multi-modal self-paced loss is introduced into cross-modal hashing to reweight multi-modal pairs dynamically, which enables the model to gradually focus on hard pairs while simultaneously learning universal patterns from multi-modal pairs. Extensive experiments on three available benchmark databases demonstrate that our proposed DHaPH framework outperforms the compared baselines with different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DHaPH.
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie
IEEE Trans. Knowl. Data Eng.3
2024 Deep Neighborhood-Preserving Hashing With Quadratic Spherical Mutual Information for Cross-Modal Retrieval
abstract
Driven by the high nonlinearity of deep neural networks, deep hashing has achieved the pictured great potential in cross-modal retrieval applications, significantly bridging the modality gap. Current deep cross-modal hashing usually utilizes affinity matching or local ranking to capture the local semantic relationships in the learned common space, leading to high neighborhood ambiguity. Simultaneously, most of these frameworks utilize additional regularization terms or margin thresholds to enhance the overall performance, in which searching the model's hyper-parameters under mass training data would have a substantial overhead. In this paper, with a novel extension of information-theoretic measures, a novel deep cross-modal hashing method, named Deep Neighborhood-preserving Hashing (DNpH), is designed to learn a highly separable discrete space, effectively mitigating the semantic gap across different modalities. Specifically, to minimize neighborhood ambiguity, the Quadratic Spherical Mutual Information (QSMI) is first introduced into deep cross-modal hashing to separate neighbors and non-neighbors well, while it is free of tuning parameters during model training compared with other similarity measures. To optimize quadratic mutual information loss smoothly, a square clamping method is developed to improve the stability of model optimization, avoiding converging on bad local optimum. Besides, two transformer encoders are exploited as feature extractors for multi-modal samples to learn the informative semantic representations. Finally, we compare our proposed DNpH framework with various state-of-the-art cross-modal hashing on four public datasets, and large amounts of experiment results demonstrate our contributions and show that DNpH outperforms the compared baselines on different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DNpH.
Qibing Qin, Yadong Huo, Lei Huang 0010, Jiangyan Dai, Huihui Zhang 0003, Wenfeng Zhang
IEEE Trans. Multim.6
2024 Deep Neighborhood Structure-Preserving Hashing for Large-Scale Image Retrieval
abstract
Deep hashing integrates the advantages of deep learning and hashing technology, and has become the mainstream of the large-scale image retrieval field. However, when training the deep hashing models, most of the existing approaches regard the similarity margin of image pairs as a constant. Once similarity distance exceeds the fixed margin, the network will not learn anything, which easily results in model collapses. In this paper, we address this dilemma with a novel unified deep hashing framework, termed Deep Neighborhood Structure-preserving Hashing (DNSH), to generate the similarity-preserving and discriminative hash codes. Specifically, by extracting the discriminative object characteristics with large variances, we design an adaptive margin quadruplet loss to further explore the underlying similarity relationship between image pairs, reflecting the correct semantic structure among its neighbors. Based on the quadruple form, we develop a quadruple regularization to decrease quantization errors between binary-like embedding and hashing codes. Furthermore, through learning bit balance and bit independent terms jointly, we present the binary code constraint loss to alleviate redundancy in different bits. Extensive evaluations on four popular benchmark datasets demonstrate that our proposed deep hashing framework achieves an excellent performance than the comparison methods.
Qibing Qin, Kezhen Xie, Wenfeng Zhang, Chengduan Wang, Lei Huang 0010
IEEE Trans. Multim.3
2024 Deep Neighborhood-aware Proxy Hashing with Uniform Distribution Constraint for Cross-modal Retrieval
abstract
Cross-modal retrieval methods based on hashing have gained significant attention in both academic and industrial research. Deep learning techniques have played a crucial role in advancing supervised cross-modal hashing methods, leading to significant practical improvements. Despite these achievements, current deep cross-modal hashing still encounters some underexplored limitations. Specifically, most of the available deep hashing usually utilizes pair-wise or triplet-wise strategies to promote the separation of the inter-classes by calculating the relative similarities between samples, weakening the compactness of intra-class data from different modalities, which could generate ambiguous neighborhoods. In this article, the Deep Neighborhood-aware Proxy Hashing (DNPH) framework is proposed to learn a discriminative embedding space with the original neighborhood relation preserved. By introducing learnable shared category proxies, the neighborhood-aware proxy loss is proposed to project the heterogeneous data into a unified common embedding, in which the sample is pulled closer to the corresponding category proxy and is pushed away from other proxies, capturing small within-class scatter and big between-class scatter. To enhance the quality of the obtained binary codes, the uniform distribution constraint is developed to make each hash bit independently obey the discrete uniform distribution. In addition, the discrimination loss is designed to preserve modality-specific semantic information of samples. Extensive experiments are performed on three benchmark datasets to prove that our proposed DNPH framework achieves comparable or even better performance compared with the state-of-the-art cross-modal retrieval applications. The corresponding code implementation of our DNPH framework is as follows: https://github.com/QinLab-WFU/OUR-DNPH .
Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang
ACM Trans. Multim. Comput. Commun. Appl.4
2023 SGRU: A High-Performance Structured Gated Recurrent Unit for Traffic Flow Prediction
abstract
Traffic flow prediction is an essential task in constructing smart cities and is a typical Multivariate Time Series (MTS) Problem. Recent research has abandoned Gated Recurrent Units (GRU) and utilized dilated convolutions or temporal slicing for feature extraction, and they have the following drawbacks: (1) Dilated convolutions fail to capture the features of adjacent time steps, resulting in the loss of crucial transitional data. (2) The connections within the same temporal slice are strong, while the connections between different temporal slices are too loose. In light of these limitations, we emphasize the importance of analyzing a complete time series repeatedly and the crucial role of GRU in MTS. Therefore, we propose SGRU: Structured Gated Recurrent Units, which involve structured GRU layers and non-linear units, along with multiple layers of time embedding to enhance the model’s fitting performance. We evaluate our approach on four publicly available California traffic datasets: PeMS03, PeMS04, PeMS07, and PeMS08 for regression prediction. Experimental results demonstrate that our model outperforms baseline models with average improvements of 11.7%, 18.6%, 18.5%, and 12.0% respectively.
Wenfeng Zhang, Xin Li 0137, Ti Wang, Honglei Gao
ICPADS1
2023 Multi-task learning based on geometric invariance discriminative features
Lei Huang 0010, Wenfeng Zhang, Yanxiu Sheng, Zhiqiang Wei 0002
Appl. Intell.4
2023 Multi-Scale Transformer-Based Matching Network for Generalizable Person Re-Identification
abstract
Recently some researches have focused on the Domain-Generalization (DG) Re-ID problem that training and testing are not in the same domain distribution. To fit the unseen complex scenes, recently deep feature matching-based methods for DG Re-ID have been developed and achieved the state-of-the-arts. However, they ignored some cases in which the accuracy of key region matching is unstable at a single scale, and the bad impact of style variations for feature representations. To address the issues, we propose a novel deep image matching model named Multi-scale Transformer-based Matching Network (MTMN) for DG Re-ID problem. MTMN matches two images with multi-scale local respondence instead of fixed representations. Specifically, the Transformer is carefully modified to formulate efficient local interactions between query and gallery images in multiple scales. Moreover, the style normalization is introduced to filter out identity-irrelated features to promote the matching results. Comprehensive experiments on several DG Re-ID tasks demonstrate the superiority of the proposed method compared with the state-of-the-arts, e.g., 5.4$\%$and 2.6$\%$gains in Rank-1 and mAP on Market-1501$\rightarrow$MSMT17(V1) task.
Jinhua Jiang, Wenfeng Zhang, Ruisheng Ran, Jiangyan Dai
IEEE Signal Process. Lett.2
2023 Deep Adaptive Quadruplet Hashing With Probability Sampling for Large-Scale Image Retrieval
abstract
With the preferable efficiency in storage and computation, hashing has shown potential application in large-scale multimedia retrieval. Compared with traditional hashing algorithms using hand-crafted characteristics, deep hashing inherits the representational capacity of deep neural networks to jointly learn semantic features and hash functions, encoding raw data into compact binary codes with significant discrimination. Generally, most of the current multi-wise hashing methods view the similarity margins between image pairs as constant values in training process. When the distance between sample pairs exceeds the fixed margin, the hashing network would not learn anything. Besides, available hashing methods commonly introduce the random sampling strategy to build training batches and ignore the sample distribution, which is harmful to parameter optimization. In this paper, we propose a novel Deep Adaptive Quadruplet Hashing with probability sampling (DAQH) for discriminative binary code learning. Specifically, with exploring the distribution relationship of raw samples, a non-uniform probability sampling strategy is proposed to build more informative and representative training batches, while maintaining the diversity of training samples. By introducing the prior similarity of sample pairs to calculate corresponding margins, an adaptive margin quadruplet loss is designed to dynamically preserve the underlying semantic relationships with its neighbors. To tune the attributes of binary codes, by combining quadruple regularization and orthogonality optimization, binary code constraint is developed to make the learned embedding with significant discrimination. Extensive experimental results on various benchmark datasets demonstrate our proposed DAQH framework achieves state-of-the-art visual similarity search performance.
Qibing Qin, Lei Huang 0010, Kezhen Xie, Zhiqiang Wei 0002, Chengduan Wang, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.6
2022 Learning to Classify Weather Conditions from Single Images Without Labels
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Zhiqiang Wei 0002
MMM (1)3
2022 From Dynamic Loading to Extensible Transformation: An Infrastructure for Dynamic Library Transformation
Yuxin Ren 0001, Jianhai Luan, Yunfeng Ye, Shiyuan Hu, Wenqin Zheng, Wenfeng Zhang, Xinwei Hu
OSDI8
2022 WCATN: Unsupervised deep learning to classify weather conditions from outdoor images
Kezhen Xie, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin
Eng. Appl. Artif. Intell.4
2022 Deep Multi-Similarity Hashing with semantic-aware preservation for multi-label image retrieval
Qibing Qin, Lintao Xian, Kezhen Xie, Wenfeng Zhang, Yu Liu 0022, Jiangyan Dai, Chengduan Wang
Expert Syst. Appl.4
2022 A CNN-based multi-task framework for weather recognition with multi-scale weather cues
Kezhen Xie, Lei Huang 0010, Wenfeng Zhang, Qibing Qin, Lei Lyu 0001
Expert Syst. Appl.3
2022 Exploiting reliable pseudo-labels for unsupervised domain adaptive person re-identification
Lei Huang 0010, Wenfeng Zhang, Zhiqiang Wei 0002
Neurocomputing3
2022 Learning discriminative features for semi-supervised person re-identification
Huanhuan Cai, Lei Huang 0010, Wenfeng Zhang, Zhiqiang Wei 0002
Multim. Tools Appl.3
2022 Learning camera invariant deep features for semi-supervised person re-identification
Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Huanhuan Cai
Multim. Tools Appl.4
2021 Subtler mixed attention network on fine-grained image classification
Chao Liu 0008, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang
Appl. Intell.4
2021 Appearance feature enhancement for person re-identification
Wenfeng Zhang, Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie
Expert Syst. Appl.1
2021 Deep top similarity hashing with class-wise loss for multi-label image retrieval
Qibing Qin, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Wenfeng Zhang
Neurocomputing5
2021 Multi-task learning with deformable convolution
Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang, Qibing Qin
J. Vis. Commun. Image Represent.4
2021 Person re-identification based on multi-appearance model
Lei Huang 0010, Wenfeng Zhang, Jie Nie, Zhiqiang Wei 0002
Multim. Tools Appl.2
2021 Weakly supervised fine-grained recognition based on spatial-channel aware attention filters
Nannan Yu, Lei Huang 0010, Zhiqiang Wei 0002, Wenfeng Zhang
Multim. Tools Appl.4
2021 Graph convolutional networks with attention for multi-label weather recognition
Kezhen Xie, Zhiqiang Wei 0002, Lei Huang 0010, Qibing Qin, Wenfeng Zhang
Neural Comput. Appl.5
2021 Angular regularization for unsupervised domain adaption on person re-identification
Wenfeng Zhang, Lei Huang 0010, Zhiqiang Wei 0002, Qibing Qin, Lei Lv
Neural Comput. Appl.1
2021 Unsupervised Deep Multi-Similarity Hashing With Semantic Structure for Image Retrieval
abstract
With the advance of Convolutional Neural Network, deep hashing methods have shown the great promising performance in large-scale image retrieval. Without depending on extensive human-annotated data, unsupervised hashing is more applicable to image retrieval tasks compared to supervised methods. However, due to the lack of fine-grained supervised signals and multi-similarity constraints, most state-of-the-art unsupervised deep hashing algorithms cannot ensure the correct fine-grained similarity ranking for image pairs. In this paper, we propose a novel unsupervised deep multi-similarity hashing framework to learn compact binary codes by jointly exploiting global-aware and spatial-aware representations, called Unsupervised Deep Multi-Similarity Hashing with Semantic Structure (UDMSH). Specifically, to obtain distinguishing characteristics, we develop a sub-network by jointly learning global semantic structures from Convolutional Neural Network (CNN) and inherent spatial structures from Fully Convolutional Network (FCN). By computing the cosine distance for deep features from image pairs, we construct a similarity matrix with semantic structure, then utilize this matrix to guide hash code learning process. Based on it, we carefully design a multi-level pairwise loss to preserve the correct fine-grained similarity ranking. Furthermore, we introduce Hamming-isometric mapping into unsupervised hashing framework to decrease the quantization errors. Extensive experiments on three widely used benchmarks prove that our proposed UDMSH outperforms several state-of-the-art unsupervised hashing with respect to different evaluation metrics.
Qibing Qin, Lei Huang 0010, Zhiqiang Wei 0002, Kezhen Xie, Wenfeng Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2020 Adaptive Attention-Aware Network for unsupervised person re-identification
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Kezhen Xie, Qibing Qin
Neurocomputing1
2020 Leveraging functional annotation to identify genes associated with complex diseases
abstract
To increase statistical power to identify genes associated with complex traits, a number of transcriptome-wide association study (TWAS) methods have been proposed using gene expression as a mediating trait linking genetic variations and diseases. These methods first predict expression levels based on inferred expression quantitative trait loci (eQTLs) and then identify expression-mediated genetic effects on diseases by associating phenotypes with predicted expression levels. The success of these methods critically depends on the identification of eQTLs, which may not be functional in the corresponding tissue, due to linkage disequilibrium (LD) and the correlation of gene expression between tissues. Here, we introduce a new method called T-GEN (Transcriptome-mediated identification of disease-associated Genes with Epigenetic aNnotation) to identify disease-associated genes leveraging epigenetic information. Through prioritizing SNPs with tissue-specific epigenetic annotation, T-GEN can better identify SNPs that are both statistically predictive and biologically functional. We found that a significantly higher percentage (an increase of 18.7% to 47.2%) of eQTLs identified by T-GEN are inferred to be functional by ChromHMM and more are deleterious based on their Combined Annotation Dependent Depletion (CADD) scores. Applying T-GEN to 207 complex traits, we were able to identify more trait-associated genes (ranging from 7.7% to 102%) than those from existing methods. Among the identified genes associated with these traits, T-GEN can better identify genes with high (>0.99) pLI scores compared to other methods. When T-GEN was applied to late-onset Alzheimer's disease, we identified 96 genes located at 15 loci, including two novel loci not implicated in previous GWAS. We further replicated 50 genes in an independent GWAS, including one of the two novel loci.
Wenfeng Zhang, Geyu Zhou, Qiongshi Lu, Hongyu Zhao 0003
PLoS Comput. Biol.3
2019 Person Re-Identification Based on Pose-Aware Segmentation
Wenfeng Zhang, Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Lei Lv, Guanqun Wei
MMM (2)1
2018 Real-Time Underwater Fish Tracking Based on Adaptive Multi-Appearance Model
abstract
Tracking live fish in an open underwater environment to investigate their behavior is of great value for many applications, e.g. biological and robotic research. However, tracking fish in real world environment is a challenging task due to complex non-rigid deformation and abrupt movement of fish. In this paper, we explore and incorporate motion property of fish and propose a real-time fish tracking method based on novel adaptive multi-appearance models and tracking strategy, which can be adapted to various changes of the fish appearance caused by non-rigid deformation. Experimental results show the promising performance of the proposed method can outperform the previous method by 13.4% in accuracy on average and is robust to real-time underwater fish tracking.
Zhiqiang Wei 0002, Lei Huang 0010, Jie Nie, Wenfeng Zhang
ICIP5
2017 A completion-invariant extension of the concept of meet continuous lattices
abstract
In this paper, the concept of meet F-continuous posets is introduced. The main results are: (1) A poset P is meet F-continuous iff its normal completion is a meet continuous lattice iff a certain system γ(P) which is, in the case of complete lattices, the lattice of all Scott closed sets is a complete Heyting algebra; (2) A poset P is precontinuous iff P is meet F-continuous and quasiprecontinuous; (3) The category of meet continuous lattices with complete homomorphisms is a full reflective subcategory of the category of meet F-continuous posets with cut-stable maps.
Wenfeng Zhang, Xiaoquan Xu
Math. Struct. Comput. Sci.1
2015 A Comparsion Of State Estimation Algorithms For Hybrid Systems
Gan Zhou, Wenquan Feng, Gautam Biswas, Wenfeng Zhang, XiuMei Guan
ECMS4
2015 S2-Quasicontinuous posets
Wenfeng Zhang, Xiaoquan Xu
Theor. Comput. Sci.1
2013 A high-quality, low-energy, small-size system-on-chip (SoC) solution enabling ECG mobile applications
abstract
ECG is widely used to monitor and diagnose cardiac conditions, but remains expensive and complex - limiting its reach to within the medical field. In this paper, we introduce the CardioChip: a single-channel, low-power, small-size ECG application-specific integrated circuit (ASIC) designed for personal mobile applications. We first show that the signal from the CardioChip Starter Kit, a wireless handheld device used to collect ECG from the index fingers, is comparable to a research-grade gel-based ECG (correlation = 98.6%). We next present a CardioChip-based R-peak detection algorithm that is able to reach a median sensitivity of greater than 0.98, and a median precision of 1. Lastly, we present a CardioChip-based respiratory rate algorithm and show that it is more accurate than a pneumography-based respiration device.
Neraj P. Bobra, Wenfeng Zhang, An Luo
IECON3
2008 Prediction of urban passenger transport based-on wavelet SVM with quantum-inspired evolutionary algorithm
abstract
Based on least squares wavelet support vector machines (LS-WSVM) with quantum-inspired evolutionary algorithm (QEA), the prediction model of urban passenger transport is proposed , that can provide the theoretical foundation of forecasting passenger volume of urban transport accurately. The prediction model of urban passenger transport is established by using LS-WSVM, whose regularization parameter and kernel parameter are adjusted using quantum-inspired evolutionary algorithm. QEA with quantum chromosome and quantum mutation has better global search capacity. The parameters of LS-WSVM can be adjusted using quantum-inspired evolutionary optimization. Combining with the data of the urban volume of passenger transport of Xipsilaan over years, the prediction model of urban passenger transport is validated, the simulation results indicate that the prediction model is effective, and based on LS-WSVM has more improvement than LS-SVM with Gaussian kernel in predicting precision, and then the improved LS-WSVM with QEA is efficient than with cross-validation method for tuning parameters.
Wenfeng Zhang, Zhongke Shi, Zhiyong Luo
IJCNN1