Haiyan Fu

dblp:09/3063 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Long-Tailed Federated Learning with Fixed Classifier
abstract
Federated learning (FL) is a machine learning approach where multiple participants train a model together without sharing their privacy data. The challenge of long-tail data heterogeneity in FL causes imbalanced data distribution among clients and significant disparities in data quantities across different classes. To tackle the long-tail distribution in FL, in this paper, we draw inspiration from the equiangular tight frame to establish a fixed balanced classifier, enhancing the model’s generalization ability for classes with fewer instances. Therefore, feature extractors and a loss aligned with fixed classifiers are designed to address the long-tail distribution in FL. The algorithm also utilizes Sample Standardization and Batch Normalization to standardize the feature space of samples, preventing the impact of the magnitude difference of features across different classes on the prediction probability. In addition, we also propose a logit adjustment loss under a mixed low-temperature setting. Experiments demonstrate that our algorithm outperforms state-of-the-art methods in FL settings with long-tail distribution.
Yi Li 0018, Xin Zheng 0008, Haiyan Fu, Yanqing Guo
ICME4
2025 VCC-Fed: A Multi-task Federated Learning Paradigm with Versatile Collaborative Clients
Yue Hua, Yi Li 0008, Xin Zheng 0008, Ming Yang 0012, Haiyan Fu, Alan Wee-Chung Liew, Yanqing Guo
PAKDD (2)5
2025 Parallel Graph Convolutional Network for Multi-modal Recommendation
Wanru Niu, Haiyan Fu, Yanqing Guo
PAKDD (3)5
2025 Reawakening Intra-modality Discrimination for Image-Text Matching
Fuxin Yu, Haiyan Fu, Yanqing Guo
PRCV (6)4
2024 Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning
abstract
While impressive performance has been achieved in image captioning, the limited diversity of the generated captions and the large parameter scale remain major barriers to the real-word application of these systems. In this work, we propose a lightweight image captioning network in combination with continuous diffusion, called Prefix-diffusion. To achieve diversity, we design an efficient method that injects prefix image embeddings into the denoising process of the diffusion model. In order to reduce trainable parameters, we employ a pre-trained model to extract image features and further design an extra mapping network. Prefix-diffusion is able to generate diverse captions with relatively less parameters, while maintaining the fluency and relevance of the captions benefiting from the generative capabilities of the diffusion model. Our work paves the way for scaling up diffusion models for image captioning, and achieves promising performance compared with recent approaches.
Guisheng Liu, Zhengcong Fei, Haiyan Fu, Yanqing Guo
LREC/COLING4
2024 Cross-modal Semantic Interference Suppression for image-text matching
Shouyong Peng, Yujuan Sun, Guorui Sheng, Haiyan Fu, Xiangwei Kong 0001
Eng. Appl. Artif. Intell.5
2024 A Memristive Phase-Shifting Chaotic Oscillator
abstract
A phase-shifting chaotic oscillator is constructed by memristive coupling. The introduced memristor revises the frequency response as the core of the frequency selection network in the oscillator. A controlled memristor is derived to maintain a stable amplification. Thus, the oscillator has two independent offset boosting voltages, and the voltages of two capacitors can be effectively controlled by cancellation. More conveniently, the simultaneous and proportional change of the op-amp supply voltage and the memristor in the oscillator also rescales the capacitors’ voltages in the same proportion. Finally, a memristor-equivalent circuit with the feedback from AD633 for division operation greatly reduces the cost of components. The hardware experiment confirms the theoretical analysis and numerical simulations.
Xiaoliang Cen, Chunbiao Li, Xu-Dong Gao 0003, Tengfei Lei, Haiyan Fu
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 A Compact Transformer for Adaptive Style Transfer
abstract
Due to the limitation of spatial receptive field, it is challenging for CNN-based style transfer methods to capture rich and long-range semantic concepts in artworks. Though the transformer provides a fresh solution by considering long-range dependencies, it suffers from the heavy burdens of parameter scale and computation cost especially in vision tasks. In this paper, we design a compact transformer architecture AdaFormer to address the problem. The model scale shrinks about 20% compared to state-of-the-art transformer for style transfer. Furthermore, we explore the adaptive style transfer by letting the content to select the detailed style element automatically and adaptively, which encourages the output to be both appealing and reasonable. We evaluate AdaFormer comprehensively in the experiments and the results have shown the effectiveness and superiority of our approach compared to existing artistic methods. Diverse plausible stylized images are obtained with better content preservation, more convincing style sfumato and lower computation complexity.
Yi Li 0018, Haiyan Fu, Xiangyang Luo 0001, Yanqing Guo
ICME3
2023 Federating Hashing Networks Adaptively for Privacy-Preserving Retrieval
abstract
With the rise of neural networks, many deep hashing networks have been successfully trained on the basis of large-scale data. However, the conventional learning process has received increasing challenges from the data privacy concerns and the decentralized storage status, especially in sensitive scenarios like surveillance retrieval. Further considering the probable different distributions of the decentralized data, in this paper, we present a collaborative hashing paradigm FedA-Hash (Federating Adapted Hashing nets) to produce personalized hashing models for the clients without exchanging their local data. To this end, the bilateral knowledge is blended gradually during the learning process between the aggregated global model and the local hashing model, instead of replacing the local model with the global model directly. Extensive experiments are conducted on representative hashing networks, involving tasks as image retrieval and person re-identification. The results show that FedA-Hash significantly enables the collaborated performance among different clients.
Yi Li 0018, Meihua Yu, Haiyan Fu, Yanqing Guo
ICME4
2023 Deep Consistency Preserving Network for Unsupervised Cross-Modal Hashing
Mengluan Li, Yanqing Guo, Haiyan Fu, Hong Su
PRCV (1)3
2022 Artistic Style Discovery with Independent Components
abstract
Style transfer has been well studied in recent years with excellent performance processed. While existing methods usually choose CNNs as the powerful tool to accomplish superb stylization, less attention was paid to the latent style space. Rare exploration of underlying dimensions results in the poor style controllability and the limited practical application. In this work, we rethink the internal meaning of style features, further proposing a novel unsupervised algorithm for style discovery and achieving personalized manip-ulation. In particular, we take a closer look into the mechanism of style transfer and obtain different artistic style components from the latent space consisting of different style features. Then fresh styles can be generated by linear combination according to various style components. Experimental results have shown that our approach is superb in 1) restylizing the original output with the diverse artistic styles discovered from the latent space while keeping the content unchanged, and 2) being generic and compatible for various style transfer methods. Our code is available in this page: https://github.com/Shelsin/ArtIns.
Yi Li 0018, Huaibo Huang, Haiyan Fu, Wanwan Wang, Yanqing Guo
CVPR4
2020 Rank-embedded Hashing for Large-scale Image Retrieval
abstract
With the growth of images on the Internet, plenty of hashing methods are developed to handle the large-scale image retrieval task. Hashing methods map data from high dimension to compact codes, so that they can effectively cope with complicated image features. However, the quantization process of hashing results in unescapable information loss. As a consequence, it is a challenge to measure the similarity between images with generated binary codes. The latest works usually focus on learning deep features and hashing functions simultaneously to preserve the similarity between images, while the similarity metric is fixed. In this paper, we propose a Rank-embedded Hashing (ReHash) algorithm where the ranking list is automatically optimized together with the feedback of the supervised hashing. Specifically, ReHash jointly trains the metric learning and the hashing codes in an end-to-end model. In this way, the similarity between images are enhanced by the ranking process. Meanwhile, the ranking results are an additional supervision for the hashing function learning as well. Extensive experiments show that our ReHash outperforms the state-of-the-art hashing methods for large-scale image retrieval.
Haiyan Fu, Ying Li 0016, Hengheng Zhang
ICMR1
2020 Efficient discrete supervised hashing for large-scale cross-modal retrieval
Yaru Han, Xiangwei Kong 0001, Lianshan Yan, Haiyan Fu, Qi Tian 0001
Neurocomputing6
2020 Node-Sensitive Graph Fusion via Topo-Correlation for Image Retrieval
abstract
Various kinds of features prove to be effective for content-based image retrieval. However, due to the diversity of image contents, a descriptor may achieve impressive performance on specific images while becoming invalid on others. Although some efforts have been made to combine features as complementary counterparts, proper weighting scheme is still a challenge for fast and accurate retrieval. In this paper, we propose an effective fusion method, termed as Topo-correlation (Topo), where the importance of each feature is measured by cross-view correlations on local affinity graphs. Specifically, the weights of similarities are node-sensitive as well as modality-sensitive, thus boosting the results of good cues while depressing adverse factors for individual images. By estimating the consensus of similarity scores with regard to a query-driven criterion, the weighted graphs are generated efficiently with low computational complexity. Extensive experimental results on four benchmarks demonstrate the superiority of the proposed approach over the state-of-the-art methods.
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Discrete Semantic Alignment Hashing for Cross-Media Retrieval
abstract
Cross-media hashing, which maps data from different modalities to a low-dimensional sharing Hamming space, has attracted considerable attention due to the rapid increase of multimodal data, for example, images and texts. Recent cross-media hashing works mainly aim at learning compact hash codes to preserve the class label-based or feature-based similarities among samples. However, these methods ignore the unbalanced semantic gaps between different modalities and high-level semantic concepts, which generally results in less effective hash functions and unsatisfying retrieval performance. Specifically, the key words of texts contain semantic meanings, while the low-level features of images lack of semantic meanings. That means the semantic gap in image modality is larger than that in text modality. In this paper, we propose a simple yet effective hashing method for cross-media retrieval to address this problem, dubbed discrete semantic alignment hashing (DSAH). First, DSAH formulates to exploit collaborative filtering to mine the relations between class labels and hash codes, which can reduce memory consumption and computational cost compared to pairwise similarity. Then, the attribute of image modality is employed to align the semantic information with text modality. Finally, to further improve the quality of hash codes, we propose a discrete optimization algorithm to learn discrete hash codes directly, and each bit has a closed-form solution. Extensive experiments on multiple public databases show that our model can seamlessly incorporate attributes and achieve promising performance.
Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
IEEE Trans. Cybern.3
2019 Contextual modeling on auxiliary points for robust image reranking
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Frontiers Comput. Sci.3
2019 Exploring geometric information in CNN for image retrieval
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu
Multim. Tools Appl.3
2018 Aggregating hierarchical binary activations for image retrieval
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Neurocomputing3
2018 Pseudo-positive regularization for deep person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Multim. Syst.3
2018 A novel two-stream saliency image fusion CNN architecture for person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Multim. Syst.3
2018 Deep hashing with top similarity preserving for image retrieval
Haiyan Fu, Xiangwei Kong 0001, Qi Tian 0001
Multim. Tools Appl.2
2018 A loss combination based deep model for person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Ming Li 0011
Multim. Tools Appl.4
2017 Deep Top Similarity Preserving Hashing for Image Retrieval
Haiyan Fu, Xiangwei Kong 0001
ICIG (2)2
2017 Part-Based Deep Hashing for Large-Scale Person Re-Identification
abstract
Large-scale is a trend in person re-identi- fication (re-id). It is important that real-time search be performed in a large gallery. While previous methods mostly focus on discriminative learning, this paper makes the attempt in integrating deep learning and hashing into one framework to evaluate the efficiency and accuracy for large-scale person re-id. We integrate spatial information for discriminative visual representation by partitioning the pedestrian image into horizontal parts. Specifically, Part-based Deep Hashing (PDH) is proposed, in which batches of triplet samples are employed as the input of the deep hashing architecture. Each triplet sample contains two pedestrian images (or parts) with the same identity and one pedestrian image (or part) of the different identity. A triplet loss function is employed with a constraint that the Hamming distance of pedestrian images (or parts) with the same identity is smaller than ones with the different identity. In the experiment, we show that the proposed PDH method yields very competitive re-id accuracy on the large-scale Market-1501 and Market-1501+500K datasets.
Fuqing Zhu, Xiangwei Kong 0001, Liang Zheng 0001, Haiyan Fu, Qi Tian 0001
IEEE Trans. Image Process.4
2016 Semantic consistency hashing for cross-modal retrieval
Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Neurocomputing3
2016 BHoG: binary descriptor for sketch-based image retrieval
Haiyan Fu, Hanguang Zhao, Xiangwei Kong 0001, Xianbo Zhang
Multim. Syst.1
2016 Binary code reranking method with weighted hamming distance
Haiyan Fu, Xiangwei Kong 0001, Zhenfan Wang
Multim. Tools Appl.1
2015 Feature extraction via multi-view non-negative matrix factorization with local graph regularization
abstract
Feature extraction is a crucial and difficult issue in pattern recognition tasks with the high-dimensional and multiple features. To extract the latent structure of multiple features without label information, multi-view learning algorithms have been developed. In this paper, motivated by manifold learning and multi-view Non-negative Matrix Factorization (NM-F), we introduce a novel feature extraction method via multi-view NMF with local graph regularization, where the inner-view relatedness between data is taken into consideration. We propose the matrix factorization objective function by constructing a nearest neighbor graph to integrate local geometrical information of each view and apply two iterative updating rules to effectively solve the optimization problem. In the experiment, we use the extracted feature to cluster several realistic datasets. The experimental results demonstrate the effectiveness of our proposed feature extraction approach.
Zhenfan Wang, Xiangwei Kong 0001, Haiyan Fu, Ming Li 0011
ICIP3
2015 Semi-supervised learning based on group sparse for relative attributes
abstract
Relative attributes provide accurate information for image processing to describe which image is more natural, more open, etc. Robustness of relative attribute learning depends on the labeled comparative image pairs. However, manually labeling is a labor intensive and time-consuming task. In this paper, a semi-supervised learning approach based on group sparse is proposed to discover pairwise comparisons automatically. We generate an initial level division of the labeled training images for the basic of new constraints. Then, group sparse representation for the unlabeled images is introduced by embedding the level information into the dictionary. The semi-supervised process is conducted by selecting samples which have minimum reconstruction errors and adding new constraints to the model by comparing the selected ones with the samples in dictionary. Experiments on three public datasets demonstrate the effectiveness of our proposed method.
Hongxue Yang, Xiangwei Kong 0001, Haiyan Fu, Ming Li 0011, Genping Zhao
ICIP3
2015 RST-invariant sketch retrieval based on circular description
abstract
The explosive growth in touchscreen smartphones and tablets need simpler and intelligent image retrieval method to offer the convenience for users. Sketch based image retrieval (SBIR) which is based on a free hand sketch has recently attracted more attention, but current methods are sensitive to rotation, scaling and translation (RST). In this paper, we introduce an efficient SBIR method called RST-Invariant Circular Description (RICD). The proposed method utilizes saliency map and boundary information to detect salient contours for each image, then uses patch hashing to eliminate the deformations and redundancies of sketches. To achieve rotational invariance, we describe salient contours using circular description. The experiment results on two public datasets demonstrated that the proposed method outperforms the state-of-the-arts in both natural and product images, and handles rotation, scaling and translation transformations.
Hanguang Zhao, Xiangwei Kong 0001, Haiyan Fu
ICIP3
2014 Binary Code Reranking Method Based on Bit Importance
abstract
Due to its compact binary codes and efficient search scheme, image hashing method is suitable for large-scale image retrieval. In image hashing methods, Hamming distance is used to measure similarity between two points. For K-bit binary codes, the Hamming distance is an into and bounded by K. Therefore, there are many returned images share the same Hamming distances with the query. In this paper, we propose an efficient image ranking method based on bit importance of binary code. Compared with the returned images of Hamming distance, important bits of query image are detected. Then, large weights are assigned to important bits and small weights are assigned to minor bits. The advantage of this proposed method is calculation efficiency. Evaluations on two large-scale image data sets demonstrate the efficacy of our binary code ranking method based on bit importance.
Haiyan Fu, Xiangwei Kong 0001, Yanqing Guo, Xingang You, Linna Zhou
ICPR1
2014 Model Semantic Relations with Extended Attributes
abstract
Attribute based image retrieval has offered a powerful way to bridge the gap between low level features and high level semantic concepts. However, existing methods rely on manually pre-labeled queries, limiting their scalability and discriminative power. Moreover, such retrieval systems restrict the users to use only the exact pre-defined query words when describing the intended search targets, and thus fail to offer good user experience. In this paper, we propose a principled approach to automatically enrich the attribute representation by leveraging additional linguistic knowledge. To this end, an external semantic pool is introduced into the learning paradigm. In addition to modelling the relations between attributes and low level features, we also model the join interdependencies of pre-labeled attributes and semantically extended attributes, which is more expressive and flexible. We further propose a novel semantic relation measure for extended attribute learning in order to take user preference into account, which we see as a step towards practical systems. Extensive experiments on several attribute benchmarks show that our approach outperforms several state-of-the-art methods and achieves promising results in improving user experience.
Xiangwei Kong 0001, Haiyan Fu, Xingang You, Yunbiao Guo
ICPR3
2013 Weakly Principal Component Hashing with Multiple Tables
Haiyan Fu, Xiangwei Kong 0001, Yanqing Guo, Jiayin Lu
MMM (2)1
2013 Large-scale image retrieval based on boosting iterative quantization hashing with query-adaptive reranking
Haiyan Fu, Xiangwei Kong 0001, Jiayin Lu
Neurocomputing1
2011 A balanced semi-supervised hashing method for CBIR
abstract
Hashing methods have attracted much attention in large scale image research in recent years, because they are not only fast, but also needing a little memory. This paper proposed a balanced semi-supervised hashing method by dividing image into several blocks. With the help of improved semi-supervised hashing, we obtain a short hash code of each block, which jointed together forms a hash code of an integrated image. In the improved semi-supervised hashing, the supervised information is completed by combining the similarity of image pairs and label information. Extensive experiments demonstrate that our method can get more balanced result between retrieval speed, saving storage of original data and retrieval accuracy in CBIR than the state-of-the-art hashing methods.
Jianhui Zhou, Haiyan Fu, Xiangwei Kong 0001
ICIP2
2006 Similarity measures on three kinds of fuzzy sets
Haiyan Fu
Pattern Recognit. Lett.2