Fusheng Hao

dblp:218/6073 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0003-1490-2235ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 6 first-author · 13 since 2021
YearPublicationVenuePosition
2026 Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute Relationships
abstract
Expressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 NoisePO: Efficient Semantic Noise Generation and Ranking for Diffusion-Based Text-to-Image Synthesis
abstract
Diffusion-based methods have achieved remarkable success in photorealistic image generation, leveraging iterative denoising steps to improve image quality. However, multi-step denoising often suffers from error accumulation-similar to exposure bias in autoregressive models-due to suboptimal noise estimation, which can lead to degraded semantic alignment and image fidelity. To tackle the challenge of suboptimal inner latent representations in generation and improve the inner latent, this paper introduces a novel method NoisePO, an efficient semantic noise preference optimization framework. NoisePO employs a semantic noise preference optimization generative adversarial network (NPO-GAN) and noise ranking methods to search for semantically relevant noises based on textual conditions, thus eliminating undesired semantic features while emphasizing the necessary semantic ones. Specifically, NoisePO utilizes a light NPO-GAN to generate semantic noises that encourage the latent at the previous step to incorporate more semantic information from the caption. Then, light ranking models are employed to filter out low-quality noises and select the best noise. Experimental results demonstrate that NoisePO consistently outperforms the baselines across widely used frameworks, achieving notable improvements in image quality, semantic consistency, and user-specific alignment as measured by IS, FID, CLIP, and other metrics. These results indicate that NoisePO effectively enhances synthesis quality and strengthens text-image alignment.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Chengqun Song, Dacheng Tao, Jun Cheng 0002
IEEE Trans. Image Process.3
2025 Task-Aware Clustering for Prompting Vision-Language Models
abstract
Prompt learning has attracted widespread attention in adapting vision-language models to downstream tasks. Existing methods largely rely on optimization strategies to ensure the task-awareness of learnable prompts. Due to the scarcity of task-specific data, overfitting is prone to occur. The resulting prompts often do not generalize well or exhibit limited task-awareness. To address this issue, we propose a novel Task-Aware Clustering (TAC) framework for prompting vision-language models, which increases the task-awareness of learnable prompts by introducing task-aware pre-context. The key ingredients are as follows: (a) generating task-aware pre-context based on task-aware clustering that can preserve the backbone structure of a downstream task with only a few clustering centers, (b) enhancing the task-awareness of learnable prompts by enabling them to interact with task-aware pre-context via the well-pretrained encoders, and (c) preventing the visual task-aware pre-context from interfering the interaction between patch embeddings by masked attention mechanism. Extensive experiments are conducted on benchmark datasets, covering the base-to-novel, domain generalization, and cross-dataset transfer settings. Ablation studies validate the effectiveness of key ingredients. Comparative results show the superiority of our TAC over competitive counterparts. The code is available at https://github.com/FushengHao/TAC.
Fusheng Hao, Fengxiang He, Fuxiang Wu, Tichao Wang, Chengqun Song, Jun Cheng 0002
CVPR1
2025 Human-Imperceptible, Machine-Recognizable Images
abstract
Massive human-related data is collected to train neural networks for computer vision tasks. A major conflict is exposed relating to software engineers between better developing AI systems and distancing from the sensitive training data. To reconcile this conflict, the paper proposes an efficient privacy-preserving learning paradigm, where images are encrypted to become ``human-imperceptible, machine-recognizable'' via one of the two encryption strategies: (1) random shuffling equally-sized patches and (2) mixing-up sub-patches. Then, minimal adaptations are made to vision transformer to enable it to learn on the encrypted images for vision tasks, including image classification and object detection. Extensive experiments on ImageNet and COCO show that the proposed paradigm achieves comparable accuracy with the competitive methods. Decrypting the encrypted images requires solving an NP-hard jigsaw puzzle or ill-posed inverse problem, which is empirically shown intractable to be recovered by various attackers, including the powerful vision transformer-based attacker. We thus show that the proposed paradigm can ensure the encrypted images have become human-imperceptible while preserving machine-recognizable information.
Fusheng Hao, Fengxiang He, Yikai Wang 0001, Fuxiang Wu, Jing Zhang 0037, Dacheng Tao, Jun Cheng 0002
IJCAI1
2025 Progressive background-foreground difference enhancement for few-shot 3D point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Qieshi Zhang, Jun Cheng 0002
Image Vis. Comput.2
2025 Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language Models
abstract
Due to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Signal Process. Lett.1
2025 Class-Irrelevant Feature Removal for Few-Shot Image Classification
abstract
Most existing few-shot image classification methods employ global pooling to aggregate class-relevant local features in a data-drive manner. Due to the difficulty and inaccuracy in locating class-relevant regions in complex scenarios, as well as the large semantic diversity of local features, the class-irrelevant information could reduce the robustness of the representations obtained by performing global pooling. Meanwhile, the scarcity of labeled images exacerbates the difficulties of data-hungry deep models in identifying class-relevant regions. These issues severely limit deep models' few-shot learning ability. In this work, we propose to remove the class-irrelevant information by making local features class relevant, thus bypassing the big challenge of identifying which local features are class irrelevant. The resulting class-irrelevant feature removal (CIFR) method consists of three phases. First, we employ the masked image modeling strategy to build an understanding of images' internal structures that generalizes well. Second, we design a semantic-complementary feature propagation module to make local features class relevant. Third, we introduce a weighted dense-connected similarity measure, based on which a loss function is raised to fine-tune the entire pipeline, with the aim of further enhancing the semantic consistency of the class-relevant local features. Visualization results show that CIFR achieves the removal of class-irrelevant information by making local features related to classes. Comparison results on four benchmark datasets indicate that CIFR yields very promising performance.
Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Fast label prediction based on shrunk anchor graph for semi-supervised incomplete multiview classification
abstract
Existing anchor graph-based semi-supervised classification methods can not adopt partial available labels of data to produce discriminative anchor graph, which is even challenging for incomplete multi-view data. Addressing above issues, a fast label prediction based on shrunk anchor graph (FLP-SAG) is designed for semi-supervised incomplete multi-view classification, which is capable of learning discriminative anchor graph iteratively. Firstly, in each view a similarity-based anchor graph is constructed and expanded to the size of complete data to align the multiple views. Then these pre-constructed anchor graphs are fused to get a common anchor graph, which is ready to be shrunk based on the predicted labels with high confidence scores in each iteration. To speed up the classification, an efficient two-step label prediction strategy is developed without the calculation of dense matrix inverse. Experimental results on four real world datasets comparing with several recently proposed methods demonstrate the superiority of the proposed method.
Guosheng Cui, Fusheng Hao, Dan Wu 0002, Ye Li 0002
ICME2
2024 Two-stage feature distribution rectification for few-shot point cloud semantic segmentation
Tichao Wang, Fusheng Hao, Guosheng Cui, Fuxiang Wu, Mengjie Yang, Qieshi Zhang, Jun Cheng 0002
Pattern Recognit. Lett.2
2023 Reject Decoding via Language-Vision Models for Text-to-Image Synthesis
abstract
Transformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Lei Wang 0018, Jun Cheng 0002
AAAI3
2023 Class-Aware Patch Embedding Adaptation for Few-Shot Image Classification
abstract
"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of few-shot learning algorithms, which have limited data and highly rely on the comparison of image patches. To address this issue, we propose a Class-aware Patch Embedding Adaptation (CPEA) method to learn "class-aware embeddings" of the image patches. The key idea of CPEA is to integrate patch embeddings with class-aware embeddings to make them class-relevant. Furthermore, we define a dense score matrix between class-relevant patch embeddings across images, based on which the degree of similarity between paired images is quantified. Visualization results show that CPEA concentrates patch embeddings by class, thus making them class-relevant. Extensive experiments on four benchmark datasets, miniImageNet, tieredImageNet, CIFAR-FS, and FC-100, indicate that our CPEA significantly outperforms the existing state-of-the-art methods. The source code is available at https://github.com/FushengHao/CPEA.
Fusheng Hao, Fengxiang He, Liu Liu 0014, Fuxiang Wu, Dacheng Tao, Jun Cheng 0002
ICCV1
2023 Prototype expansion and feature calibration for few-shot point cloud semantic segmentation
Qieshi Zhang, Tichao Wang, Fusheng Hao, Fuxiang Wu, Jun Cheng 0002
Neurocomputing3
2023 Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
Qieshi Zhang, Zhenyu Xu 0014, Yuhang Kang, Fusheng Hao, Ziliang Ren, Jun Cheng 0002
Knowl. Based Syst.4
2023 Semantic-Aware Feature Aggregation for Few-Shot Image Classification
Fusheng Hao, Fuxiang Wu, Fengxiang He, Qieshi Zhang, Chengqun Song, Jun Cheng 0002
Neural Process. Lett.1
2023 Mixer-Based Semantic Spread for Few-Shot Learning
abstract
Key semantics can come from everywhere on an image. Semantic alignment is a key part of few-shot learning but still remains challenging. In this paper, we design a Mixer-Based Semantic Spread (MBSS) algorithm that employs amixermodule to spread the key semantic on the whole image, so that one can directly compare the processed image pairs. We first adopt a convolutional neural network to extract features from both support and query images and separate each of them into multiple Local Descriptor-based Representations (LDRs). The LDRs are then fed into themixerfor semantic spread, where every LDR attracts complementary information from its peers. In this way, the objective semantic is made spread on the whole image in a data-driven manner. The overall pipeline is supervised by a voting-based loss, guaranteeing a goodmixer. Visualization results validate the feasibility of ourmixer. Comprehensive experiments on three benchmark datasets, miniImageNet, tieredImageNet, and CUB, show that our algorithm achieves the state-of-the-art performance in both 5-way 1-shot and 5-way 5-shot settings.
Jun Cheng 0002, Fusheng Hao, Fengxiang He, Liu Liu 0014, Qieshi Zhang
IEEE Trans. Multim.2
2023 Language-Based Image Manipulation Built on Language-Guided Ranking
abstract
Text-based image manipulation is a popular subject and has many applications. However, it is a challenging task because there is no ground-truth edited dataset and textual descriptions have abstractive and ambiguous properties. To alleviate the difficult issues, we propose a manipulation framework consisting of the proposal attentional GANs, language-related semantic mask, and language-guided ranker. Specially, we construct an editing proposal generator to generate the suitable edited proposals with and without semantic conditions, which supports the reorganization of sub-generators to output proposals in various aspects as many as possible. To distinguish the text-relevant and the text-irrelevant regions, we introduce a language-related semantic mask based on the source image and target caption. Then, we exploit a language-guided ranker to retrieve the best edited result from the edited proposals through using the multi-modal similarity and the language-related semantic mask. Extensive experiments on widely-used datasets demonstrate that our model could manipulate images interactively and improve the editing quality effectively.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002
IEEE Trans. Multim.3
2022 Text-to-Image Synthesis based on Object-Guided Joint-Decoding Transformer
abstract
Object-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generating language-related layout is not a trivial task; 2) error propagation, because the inappropriate layout will mislead the image synthesis and is hard to be revised. In this paper, we propose an object-guided joint-decoding module to simultaneously generate the image and the corresponding layout. Specially, we present the joint-decoding transformer to model the joint probability on images tokens and the corresponding layouts tokens, where layout tokens provide additional observed data to model the complex scene better. Then, we describe a novel Layout-Vqgan for layout encoding and decoding to provide more information about the complex scene. After that, we present the detail-enhanced module to enrich the language-related details based on two facts: 1) visual details could be omitted in the compression of VQGANs; 2) the joint-decoding transformer would not have sufficient generating capacity. The experiments show that our approach is competitive with previous object-centered models and can generate diverse and high-quality objects under the given layouts.
Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002
CVPR3
2022 Structure-Preserving View-Invariant Skeleton Representation for Action Detection
abstract
Skeleton-based action detection has attracted increasing attention in recent years due to its action-focusing and compactness. To enable the usability of convolutional neural networks, many methods convert a skeleton sequence to a pseudo image by stacking the skeleton joints based on a predefined order. However, this practice ignores the skeletons structure and the influence of viewpoints, thus limiting the performance of learned models. In this paper, we propose a novel representation, which preserves the structure information while being view-invariant. To achieve this, we first generate a structure-preserving chain order by leveraging the depth-first traversal algorithm. Then, to eliminate the influence of viewpoints, we propose a reference joint-based encoding approach. Finally, we improve YOLOv5 by introducing an attention feature learning network, which enables YOLOv5 to automatically select the most informative joints. Comprehensive experiments on the PKU-MMD dataset demonstrate our method achieves state-of-the-art performance while maintaining high efficiency.
Hushan Qin, Jun Cheng 0002, Chengqun Song, Fusheng Hao, Qin Cheng
ICPR4
2022 Cross-Modality Compensation Convolutional Neural Networks for RGB-D Action Recognition
abstract
RGB-D-based human action recognition has attracted much attention recently because it can provide more complementary information than a single modality. However, it is difficult for two modalities to effectively learn spatial-temporal information from each other. To facilitate information interaction between different modalities, a cross-modality compensation convolutional neural network (ConvNet) is proposed for human action recognition, which enhances the discriminative ability by jointly learning compensation features from the RGB and depth modalities. Moreover, we design a cross-modality compensation block (CMCB) to extract compensation features from the RGB and depth modalities. Specifically, CMCB is incorporated into two typical network architectures, ResNet and VGG, to verify the ability to improve the performance of our model. The proposed architecture has been evaluated on three challenging datasets: NTU RGB+D 120, THU-READ and PKU-MMD. We experimentally verify that our proposed model with CMCB is effective for different input types, such as pairs of raw images and dynamic images constructed from the entire RGB-D sequence, and the experimental results show that the proposed framework achieves state-of-the-art performance on all three datasets.
Jun Cheng 0002, Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Fusheng Hao
IEEE Trans. Circuits Syst. Video Technol.5
2022 Global-Local Interplay in Semantic Alignment for Few-Shot Learning
abstract
Few-shot learning aims to recognize novel classes from only a few labeled training examples. Aligning semantically relevant local regions has shown promise in effectively comparing a query image with support images. However, global information is usually overlooked in the existing approaches, resulting in a higher possibility of learning semantics unrelated to the global information. To address this issue, we propose a Global-Local Interplay Metric Learning (GLIML) framework to employ the interplay between global features and local features to guide semantic alignment. We first design a Global-Local Information Concurrent Learning (GLICL) module to extract both global features and local features and perform global-local interplay. We then design a Global-Local Information Cross-Covariance Estimator (GLICCE) to learn the similarity on the global-local interplay, in contrast to the current practice where only local features are considered. Visualizations show that the global-local interplay decreases (1) the weights placed on the semantics that are irrelevant to the global information and (2) the variability of the learned features within every class in the feature space. Quantitative experiments on three benchmark datasets demonstrate that GLIML achieves state-of-the-art performance while maintaining high efficiency.
Fusheng Hao, Fengxiang He, Jun Cheng 0002, Dacheng Tao
IEEE Trans. Circuits Syst. Video Technol.1
2022 Imposing Semantic Consistency of Local Descriptors for Few-Shot Learning
abstract
Few-shot learning suffers from the scarcity of labeled training data. Regarding local descriptors of an image as representations for the image could greatly augment existing labeled training data. Existing local descriptor based few-shot learning methods have taken advantage of this fact but ignore that the semantics exhibited by local descriptors may not be relevant to the image semantic. In this paper, we deal with this issue from a new perspective of imposing semantic consistency of local descriptors of an image. Our proposed method consists of three modules. The first one is a local descriptor extractor module, which can extract a large number of local descriptors in a single forward pass. The second one is a local descriptor compensator module, which compensates the local descriptors with the image-level representation, in order to align the semantics between local descriptors and the image semantic. The third one is a local descriptor based contrastive loss function, which supervises the learning of the whole pipeline, with the aim of making the semantics carried by the local descriptors of an image relevant and consistent with the image semantic. Theoretical analysis demonstrates the generalization ability of our proposed method. Comprehensive experiments conducted on benchmark datasets indicate that our proposed method achieves the semantic consistency of local descriptors and the state-of-the-art performance.
Jun Cheng 0002, Fusheng Hao, Liu Liu 0014, Dacheng Tao
IEEE Trans. Image Process.2
2021 VGG-CAE: Unsupervised Visual Place Recognition Using VGG16-Based Convolutional Autoencoder
Zhenyu Xu 0014, Qieshi Zhang, Fusheng Hao, Ziliang Ren, Yuhang Kang, Jun Cheng 0002
PRCV (2)3
2021 Segment spatial-temporal representation and cooperative learning of convolution neural networks for multimodal-based action recognition
Ziliang Ren, Qieshi Zhang, Jun Cheng 0002, Fusheng Hao, Xiangyang Gao
Neurocomputing4
2020 Embedded adaptive cross-modulation neural network for few-shot learning
Jun Cheng 0002, Fusheng Hao, Wei Feng 0009
Neural Comput. Appl.3
2019 Collect and Select: Semantic Alignment Metric Learning for Few-Shot Learning
abstract
Few-shot learning aims to learn latent patterns from few training examples and has shown promises in practice. However, directly calculating the distances between the query image and support image in existing methods may cause ambiguity because dominant objects can locate anywhere on images. To address this issue, this paper proposes a Semantic Alignment Metric Learning (SAML) method for few-shot learning that aligns the semantically relevant dominant objects through a "collect-and-select'' strategy. Specifically, we first calculate a relation matrix (RM) to "`collect" the distances of each local region pairs of the 3D tensor extracted from a query image and the mean tensor of the support images. Then, the attention technique is adapted to "select" the semantically relevant pairs and put more weights on them. Afterwards, a multi-layer perceptron (MLP) is utilized to map the reweighted RMs to their corresponding similarity scores. Theoretical analysis demonstrates the generalization ability of SAML and gives a theoretical guarantee. Empirical results demonstrate that semantic alignment is achieved. Extensive experiments on benchmark datasets validate the strengths of the proposed approach and demonstrate that SAML significantly outperforms the current state-of-the-art methods. The source code is available at https://github.com/haofusheng/SAML.
Fusheng Hao, Fengxiang He, Jun Cheng 0002, Lei Wang 0018, Jianzhong Cao, Dacheng Tao
ICCV1