Hao Fu 0020

dblp:64/3069-20 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0005-0032-8633ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ENHash: Error Notebook-Guided Fine-Grained Learning for Unsupervised Cross-Modal Hashing
abstract
Without manual annotations, unsupervised cross-modal hashing (UCMH) aims to achieve efficient clustering and retrieval by leveraging data interrelationships. However, the retrieval accuracy is constrained by two main aspects: 1) insufficient exploration of data relationships; 2) existing knowledge mining strategies are not well aligned with the architectural properties of multilayer perceptrons. Through summary and error analysis, the human brain is able to achieve fast learning through experience and minimal data. Inspired by this cognitive process, we propose a novel Error Notebook strategy, named ENHash, to more effectively capture similarity information between multi-modal data for fine-grained unsupervised clustering. Firstly, simulating the human process of summarizing experiences, ENHash gradually integrates the information from each batch into a global clustering representation. Secondly, drawing upon human error analysis capabilities, ENHash utilizes the summarized experiences to identify and record incorrectly predicted hash codes. Finally, by leveraging the knowledge derived from this analysis, ENHash guides the hash function to learn fine-grained patterns from the errors. To the best of our knowledge, ENHash represents the first attempt at integrating cognitively-inspired mechanisms into fine-grained UCMH optimization paradigms. We evaluate the proposed ENHash against eight state-of-the-art methods on three widely used datasets and one fine-grained cross-modal dataset. Experimental results show that ENHash achieves substantial improvements over existing approaches.
Hao Fu 0020, Zebing Yao, Chuangchuang Tan, Guanghua Gu
AAAI1
2026 BRAINHash: Brain-Inspired Region-Aligned Interaction Network for Unsupervised Cross-Modal Hashing
abstract
Unsupervised cross-modal hashing (UCMH) has attracted considerable attention owing to its minimal reliance on manual annotations and low retrieval latency. However, existing UCMH methods based on contrastive learning frameworks combined with multilayer perceptrons (MLPs) often suffer from two key limitations: the inherent challenge in constructing reliable positive-negative sample pairs under unsupervised settings, and the restricted representational capacity of static network architectures. Inspired by cognitive mechanisms observed in the human brain, where specialized regions support distinct roles in knowledge acquisition and error-driven learning, we propose a Brain-inspired Region-Aligned Interaction Network for unsupervised cross-modal Hashing (BRAINHash), which is a novel brain-inspired memory-based temporal modeling strategy. BRAINHash emulates several functional components of the brain through modular design choices: 1) diverse encoders approximate feature encoding akin to occipital lobe processing, 2) an error-aware optimization strategy models initial cross-modal association construction analogous to prefrontal cortex activity, 3) a teacher network serves as hippocampal-like memory storage capturing learned representations as persistent memory points, 4) we incorporate soft-target contrastive objectives alongside spiking neural network-based temporal modeling to simulate higher-order reasoning typically attributed to prefrontal decision-making circuits. To the best of our knowledge, BRAINHash represents the first integration of biologically inspired architectural principles with temporal dynamics into an unsupervised cross-modal hashing paradigm. Extensive evaluations conducted on five widely-used datasets demonstrate that our method outperforms fifteen state-of-the-art approaches. The implementation code is publicly available at https://github.com/YSU-ISU-Lab/BRAIN.
Hao Fu 0020, Guanghua Gu, Yunchao Wei, Yao Zhao 0001
IEEE Trans. Image Process.1
2025 Adaptive Asymmetric Online Hashing for Cross-Modal Retrieval
abstract
Online cross-modal hashing utilizing a progressive update strategy has attracted considerable interest due to its effectiveness and scalability in similarity search for large-scale multimodal data retrieval across various extensive multimedia datasets. Nevertheless, existing online cross-modal hashing approaches face several limitations. One major challenge lies in effectively capturing the intrinsic linkages among different heterogeneous modalities. Additionally, relying on relaxation-based strategies to solve the discrete constraint problem often introduces significant quantization errors, resulting in suboptimal solutions. To mitigate these limitations, we present a novel online hashing approach named AAOH. This method preserves the information integrity of multimodal data by decomposing the input into a common latent representation and transformation matrices, guided by adaptive weighting and nuclear norm minimization. Next, the common latent representation is aligned with the semantic label matrix, and an asymmetric hashing framework is used to enhance the model's discriminative ability, producing more compact hash codes. Finally, an iterative discrete optimization algorithm is proposed to efficiently solve the non-convex multi-variable optimization problem. A series of thorough evaluations on three well-established benchmark datasets are carried out to showcase the superior performance of the proposed AAOH method.
Yuhao Liu 0014, Hao Fu 0020, Guanghua Gu
ICMR3
2025 Dynamic Optimization Noisy Cross-Modal Hashing
abstract
Cross-Modal Hashing (CMH) has gained significant attention for its ability to learn semantic category discrimination and enable efficient retrieval. However, in practical applications, the massive amounts of multi-modal data collected from the internet often contain coarse annotations, which inevitably introduce noisy labels and degrade retrieval performance. To address this challenge, this paper proposes a dynamic optimization-based training framework, namely Dynamic Optimization Noisy Cross-Modal Hashing (DONCMH). Firstly, to alleviate the issue of overfitting to noisy labels during training, we propose a novel regularization-based noise-robust strategy that updates the target distribution with momentum to optimize clustering learning, thus avoiding over-emphasizing noisy samples. Secondly, to more accurately select high-quality training samples, we introduce ClusterOT, a novel Optimal Transport formulation explicitly tailored for Noisy Cross-Modal Hashing (NCMH), which integrates center representation learning and cross-modal alignment into a unified structure. By leveraging the spatial distribution of samples, ClusterOT effectively mitigates distribution imbalances inherent in center representation learning, thereby significantly improving the model's robustness to noisy label predictions. Finally, a robust feature learning module is employed to enhance the extraction of informative and discriminative representations from both modalities. Extensive experiments conducted on four widely used benchmark datasets demonstrate that the proposed method effectively mitigates the impact of noisy labels and significantly improves cross-modal retrieval performance.
Zebing Yao, Hao Fu 0020, Yuanhang Yang, Guanghua Gu
ACM Multimedia2
2025 Consistency Aware Representation Learning for Unsupervised Cross-Domain Image Retrieval
Zebing Yao, Hao Fu 0020, Yuhao Liu 0014, Guanghua Gu
PRCV (5)2
2024 Semi-supervised cross-modal hashing with joint hyperboloid mapping
Hao Fu 0020, Guanghua Gu, Yiyang Dou, Zhuoyi Li, Yao Zhao 0001
Knowl. Based Syst.1
2023 Adaptive Adversarial Learning based cross-modal retrieval
Zhuoyi Li, Huibin Lu, Hao Fu 0020, Guanghua Gu
Eng. Appl. Artif. Intell.3
2023 Parallel learned generative adversarial network with multi-path subspaces for cross-modal retrieval
Zhuoyi Li, Huibin Lu, Hao Fu 0020, Guanghua Gu
Inf. Sci.3
2023 Semi-Supervised Knowledge Distillation for Cross-Modal Hashing
abstract
Deep hashing methods have achieved tremendous success in cross-modal retrieval, due to its low storage consumption and fast retrieval speed. Supervised cross-modal hashing methods have achieved substantial advancement by incorporating semantic information. However, to a great extent, supervised methods rely on large-scale labeled cross-modal training data which are laborious to obtain. Moreover, most cross-modal hashing methods only handle two modalities of image and text, without taking the scene of multiple modalities into consideration. In this paper, we propose a novel semi-supervised approach called semi-supervised knowledge distillation for cross-modal hashing (SKDCH) to overcome the above-mentioned challenges, which enables guiding a supervised method using outputs produced by a semi-supervised method for multimodality retrieval. Specifically, we utilize teacher-student optimization to propagate knowledge. Furthermore, we improves triplet ranking loss to better mitigate the heterogeneity gap, which increases the discriminability of our proposed approach. Extensive experiments executed on two benchmark datasets validate that the proposed SKDCH surpasses the state-of-the-art methods.
Mingyue Su, Guanghua Gu, Xianlong Ren, Hao Fu 0020, Yao Zhao 0001
IEEE Trans. Multim.4
2022 Image-text bidirectional learning network based cross-modal retrieval
Zhuoyi Li, Huibin Lu, Hao Fu 0020, Guanghua Gu
Neurocomputing3
2022 CUMTGAN: An instance-level controllable U-Net GAN for facial makeup transfer
Miao Hao, Guanghua Gu, Hao Fu 0020
Knowl. Based Syst.3