VLDB 2026 Research / reviewers in the wild / expert
Yadong Huo
dblp:366/2980
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0008-8805-8958ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep semantic center-guided hashing for multi-label cross-modal retrieval
Xinzheng Sui, Yadong Huo, Qibing Qin, Lei Huang 0010, Wenfeng Zhang |
Expert Syst. Appl. | 3 |
| 2026 | Reference-Based Super-Resolution With Geometry-Aware Transfer
Lei Huang 0010, Jie Nie, Yadong Huo, Jin Du, Zhiqiang Wei 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Factorized Transformer Hashing with Adaptive Routing for Large-scale Image RetrievalabstractTransformer architecture has driven significant advancements in deep hashing, establishing itself as a dominant framework for large-scale retrieval and storage applications. However, existing Transformer-based deep hashing methods typically employ unvarying feature transformations across all images, limiting their adaptability to diverse visual patterns. This rigidity restricts the model's capacity to learn both highly distinctive and generalizable discrete representations, posing challenges for retrieval in open-world scenarios. To overcome this challenge, we propose a novel Factorized Transformer Hashing (FTH) framework, which introduces a factorized transformer to enhance the generalization and discriminative power of hash codes. Specifically, we decompose the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) blocks into multiple sub-blocks, forming a transformer factorization scheme that captures diverse feature characteristics through independent sub-blocks. Furthermore, we develop an adaptive selection strategy, leveraging a set of learnable selectors with the Softmax function, to dynamically route each image to the most appropriate sub-block for processing. Extensive experiments on three benchmark datasets demonstrate that the proposed FTH framework significantly outperforms state-of-the-art baselines in both image hashing and zero-shot hashing tasks. Source code is available at https://github.com/QinLab-WFU/FTH. Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
ACM Multimedia | 1 |
| 2025 | Ranking-oriented cross-modal hashing
Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Deep Discriminative Boundary Hashing for Cross-Modal RetrievalabstractBy the preferable efficiency in storage and computation, deep cross-modal has gained much attention in large-scale multimedia retrieval. Current deep hashing employs the probability outputs of the likelihood function, i.e., Sigmoid or Cauchy, to quantify the semantic similarity between samples in a common Hamming space. However, the inherent weakness of the Sigmoid likelihood function or the Cauchy likelihood function in gradient optimization leads to hashing models failing to exactly describe the hamming ball, which indicates the absolute semantic boundary among classes, thereby giving the high neighborhood ambiguity. In this paper, with the analysis of the likelihood function from the perspective of similarity metric learning, the novel Deep Discriminative Boundary Hashing framework (DDBH) is proposed to learn the discriminative embedding space that separates neighbors and non-neighbors well. Specifically, by introducing the remapping strategy and the base-point adaptive selection, the boundary-preserving loss based on the adjustable likelihood function is proposed to project data points with small gradients to regions with large gradients and give larger gradients for hard samples, facilitating better separation among classes. Meanwhile, to learn class-dependent binary codes, the class-wise quantization loss is designed to heuristically transfer class-wise prior knowledge to the binary quantization, significantly improving the discriminative capability of compact discrete codes. Comprehensive experiments on three benchmark datasets show that our proposed DDBH framework outperforms other representative deep cross-modal hashing. The corresponding code is available at https://github.com/QinLab-WFU/DDBH. Qibing Qin, Yadong Huo, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Deep Semantic-Aware Proxy Hashing for Multi-Label Cross-Modal RetrievalabstractDeep hashing has attracted broad interest in cross-modal retrieval because of its low cost and efficient retrieval benefits. To capture the semantic information of raw samples and alleviate the semantic gap, supervised cross-modal hashing methods that utilize label information which could map raw samples from different modalities into a unified common space, are proposed. Although making great progress, existing deep cross-modal hashing methods are suffering from some problems, such as: 1) considering multi-label cross-modal retrieval, proxy-based methods ignore the data-to-data relations and fail to explore the combination of the different categories profoundly, which could lead to some samples without common categories being embedded in the vicinity; 2) for feature representation, image feature extractors containing multiple convolutional layers cannot fully obtain global information of images, which results in the generation of sub-optimal binary hash codes. In this paper, by extending the proxy-based mechanism to multi-label cross-modal retrieval, we propose a novel Deep Semantic-aware Proxy Hashing (DSPH) framework, which could embed multi-modal multi-label data into a uniform discrete space and capture fine-grained semantic relations between raw samples. Specifically, by learning multi-modal multi-label proxy terms and multi-modal irrelevant terms jointly, the semantic-aware proxy loss is designed to capture multi-label correlations and preserve the correct fine-grained similarity ranking among samples, alleviating inter-modal semantic gaps. In addition, for feature representation, two transformer encoders are proposed as backbone networks for images and text, respectively, in which the image transformer encoder is introduced to obtain global information of the input image by modeling long-range visual dependencies. We have conducted extensive experiments on three baseline multi-label datasets, and the experimental results show that our DSPH framework achieves better performance than state-of-the-art cross-modal hashing methods. The code for the implementation of our DSPH framework is available athttps://github.com/QinLab-WFU/DSPH. Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Deep Hierarchy-Aware Proxy Hashing With Self-Paced Learning for Cross-Modal RetrievalabstractDue to its low storage cost and high retrieval efficiency, hashing technology is popularly applied in both academia and industry, which provides an interesting solution for cross-modal similarity retrieval. However, most existing supervised cross-modal hashing methods typically view the fixed-level semantic affinity defined by manual labels as supervised signals to guide hash learning, which only represents a small subset of complex semantic relations between multi-modal samples, thus impeding the hash function learning and degrading the obtained hash codes. In the paper, by learning shared hierarchy proxies, a novel deep cross-modal hashing framework, called Deep Hierarchy-aware Proxy Hashing (DHaPH), is proposed to construct the semantic hierarchy in a data-driven manner, thereby capturing the accurate fine-grained semantic relationships and achieving small intra-class scatter and big inter-class scatter. Specifically, by regarding the hierarchical proxies as learnable ancestors, a novel hierarchy-aware proxy loss is designed to model the latent semantic hierarchical structures from different modalities without prior hierarchy knowledge, in which similar samples share the same Lowest Common Ancestor (LCA) and dissimilar points have different LCA. Meanwhile, to adequately capture valuable semantic information from hard pairs, a multi-modal self-paced loss is introduced into cross-modal hashing to reweight multi-modal pairs dynamically, which enables the model to gradually focus on hard pairs while simultaneously learning universal patterns from multi-modal pairs. Extensive experiments on three available benchmark databases demonstrate that our proposed DHaPH framework outperforms the compared baselines with different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DHaPH. Yadong Huo, Qibing Qin, Wenfeng Zhang, Lei Huang 0010, Jie Nie |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Deep Neighborhood-Preserving Hashing With Quadratic Spherical Mutual Information for Cross-Modal RetrievalabstractDriven by the high nonlinearity of deep neural networks, deep hashing has achieved the pictured great potential in cross-modal retrieval applications, significantly bridging the modality gap. Current deep cross-modal hashing usually utilizes affinity matching or local ranking to capture the local semantic relationships in the learned common space, leading to high neighborhood ambiguity. Simultaneously, most of these frameworks utilize additional regularization terms or margin thresholds to enhance the overall performance, in which searching the model's hyper-parameters under mass training data would have a substantial overhead. In this paper, with a novel extension of information-theoretic measures, a novel deep cross-modal hashing method, named Deep Neighborhood-preserving Hashing (DNpH), is designed to learn a highly separable discrete space, effectively mitigating the semantic gap across different modalities. Specifically, to minimize neighborhood ambiguity, the Quadratic Spherical Mutual Information (QSMI) is first introduced into deep cross-modal hashing to separate neighbors and non-neighbors well, while it is free of tuning parameters during model training compared with other similarity measures. To optimize quadratic mutual information loss smoothly, a square clamping method is developed to improve the stability of model optimization, avoiding converging on bad local optimum. Besides, two transformer encoders are exploited as feature extractors for multi-modal samples to learn the informative semantic representations. Finally, we compare our proposed DNpH framework with various state-of-the-art cross-modal hashing on four public datasets, and large amounts of experiment results demonstrate our contributions and show that DNpH outperforms the compared baselines on different evaluation metrics. The corresponding code is available athttps://github.com/QinLab-WFU/DNpH. Qibing Qin, Yadong Huo, Lei Huang 0010, Jiangyan Dai, Huihui Zhang 0003, Wenfeng Zhang |
IEEE Trans. Multim. | 2 |
| 2024 | Deep Neighborhood-aware Proxy Hashing with Uniform Distribution Constraint for Cross-modal RetrievalabstractCross-modal retrieval methods based on hashing have gained significant attention in both academic and industrial research. Deep learning techniques have played a crucial role in advancing supervised cross-modal hashing methods, leading to significant practical improvements. Despite these achievements, current deep cross-modal hashing still encounters some underexplored limitations. Specifically, most of the available deep hashing usually utilizes pair-wise or triplet-wise strategies to promote the separation of the inter-classes by calculating the relative similarities between samples, weakening the compactness of intra-class data from different modalities, which could generate ambiguous neighborhoods. In this article, the Deep Neighborhood-aware Proxy Hashing (DNPH) framework is proposed to learn a discriminative embedding space with the original neighborhood relation preserved. By introducing learnable shared category proxies, the neighborhood-aware proxy loss is proposed to project the heterogeneous data into a unified common embedding, in which the sample is pulled closer to the corresponding category proxy and is pushed away from other proxies, capturing small within-class scatter and big between-class scatter. To enhance the quality of the obtained binary codes, the uniform distribution constraint is developed to make each hash bit independently obey the discrete uniform distribution. In addition, the discrimination loss is designed to preserve modality-specific semantic information of samples. Extensive experiments are performed on three benchmark datasets to prove that our proposed DNPH framework achieves comparable or even better performance compared with the state-of-the-art cross-modal retrieval applications. The corresponding code implementation of our DNPH framework is as follows: https://github.com/QinLab-WFU/OUR-DNPH . Yadong Huo, Qibing Qin, Jiangyan Dai, Wenfeng Zhang, Lei Huang 0010, Chengduan Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |