Zhenyong Fu

dblp:36/8779 · also Zhen-Yong Fu · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient and spatially-aware single-image editing via hybrid cross-attention
Xinqi Huang, Jinhan Zhao, Zhenyong Fu
J. Vis. Commun. Image Represent.3
2025 LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blending
abstract
Customized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framework for effectively and efficiently scaling multi-concept customized generation. The core idea of LaTexBlend is to represent single concepts and blend multiple concepts within a Latent Textual space, which is positioned after the text encoder and a linear projection. LaTexBlend customizes each concept individually, storing them in a concept bank with a compact representation of latent textual features that captures sufficient concept information to ensure high fidelity. At inference, concepts from the bank can be freely and seamlessly combined in the latent textual space, offering two key merits for multi-concept generation: 1) excellent scalability, and 2) significant reduction of denoising deviation, preserving coherent layouts. Extensive experiments demonstrate that LaTexBlend can flexibly integrate multiple customized concepts with harmonious structures and high subject fidelity, substantially outperforming baselines in both generation quality and computational efficiency. Project page: https://jinjianrick.github.io/latexblend/
Zhenbo Yu, Yang Shen 0006, Zhenyong Fu, Jian Yang 0003
CVPR4
2025 UniCanvas: Affordance-Aware Unified Real Image Editing via Customized Text-to-Image Generation
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003
Int. J. Comput. Vis.4
2025 Register assisted aggregation for visual place recognition
Zhenyong Fu
J. Vis. Commun. Image Represent.2
2025 Exploring multi-semantic disentangled controls in GANs using conjugate gradient optimization
Zhenbo Yu, Zhenyong Fu, Jian Yang 0003
Pattern Recognit. Lett.3
2025 ZS-VAT: Learning Unbiased Attribute Knowledge for Zero-Shot Recognition Through Visual Attribute Transformer
abstract
In zero-shot learning (ZSL), attribute knowledge plays a vital role in transferring knowledge from seen classes to unseen classes. However, most existing ZSL methods learn biased attribute knowledge, which usually results in biased attribute prediction and a decline in zero-shot recognition performance. To solve this problem and learn unbiased attribute knowledge, we propose a visual attribute Transformer for zero-shot recognition (ZS-VAT), which is an effective and interpretable Transformer designed specifically for ZSL. In ZS-VAT, we design an attribute-head self-attention (AHSA) that is capable of learning unbiased attribute knowledge. Specifically, each attribute head in AHSA first transforms the local features into attribute-reinforced features and then accumulates the attribute knowledge from all corresponding reinforced features, reducing the mutual influence between attributes and avoiding information loss. AHSA finally preserves unbiased attribute knowledge through attribute embeddings. We also propose an attribute fusion model (AFM) that learns to recover the correct category knowledge from the attribute knowledge. In particular, AFM takes all features from AHSA as input and generates global embeddings. We carried out experiments to demonstrate that the attribute knowledge from AHSA and the category knowledge from AFM are able to assist each other. During the final semantic prediction, we combine the attribute embedding prediction (AEP) and global embedding prediction (GEP). We evaluated the proposed scheme on three benchmark datasets. ZS-VAT outperformed the state-of-the-art generalized ZSL (GZSL) methods on two datasets and achieved competitive results on the other dataset.
Zongyan Han, Zhenyong Fu, Shuo Chen 0003, Le Hui, Jian Yang 0003, Chang Wen Chen
IEEE Trans. Neural Networks Learn. Syst.2
2024 Customized Generation Reimagined: Fidelity and Editability Harmonized
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003
ECCV (50)3
2024 Few-shot open-set recognition via pairwise discriminant aggregation
Yang Shen 0006, Zhenyong Fu, Jian Yang 0003
Neurocomputing3
2024 Using Mixture of Experts to accelerate dataset distillation
Zhenyong Fu
J. Vis. Commun. Image Represent.2
2024 Image Hiding and Restoration via Deep Moiré Networks
abstract
Moiré patterns arise when two repetitive textures are superimposed, which often degrade the image quality when taking photos. However, controllable moiré patterns can be useful in many practical applications e.g. image hiding. Existing moiré-based approaches heavily rely on manually designed filters, resulting in slow and inflexibility. Moreover, existing methods often struggle to recover image details, making it challenging to achieve satisfactory visual quality. To address these limitations, in this study, a deep moiré hiding network (DMHN) is proposed to control the generation of moiré patterns and applied to the image hiding task. The network encodes the secret image to generate modulation parameters of moiré patterns and hide the images in unreadable gratings. When superimposing the hidden image with the key grating, the hidden image becomes visible again in the form of moiré patterns. To better restore the moiré-based reappeared images, a grating removal network (GRN) is also presented to eliminate the residual gratings and recover high-frequency image details. Experiments demonstrate that the proposed method can effectively hide and restore the images via moiré modulation and generation.
Zhenyong Fu
IEEE Signal Process. Lett.2
2023 Continual learning via region-aware memory
Zhenyong Fu, Jian Yang 0003
Appl. Intell.2
2023 Boosting separated softmax with discrimination for class incremental learning
Zhenyong Fu
J. Vis. Commun. Image Represent.3
2022 Semantic Contrastive Embedding for Generalized Zero-Shot Learning
Zongyan Han, Zhenyong Fu, Shuo Chen 0003, Jian Yang 0003
Int. J. Comput. Vis.2
2022 Streaming feature selection via graph diffusion
Shuo Chen 0003, Zhenyong Fu, Jun Li 0027, Jian Yang 0003
Inf. Sci.3
2022 Feature Selection Boosted by Unselected Features
abstract
Feature selection aims to select strongly relevant features and discard the rest. Recently, embedded feature selection methods, which incorporate feature weights learning into the training process of a classifier, have attracted much attention. However, traditional embedded methods merely focus on the combinatorial optimality of all selected features. They sometimes select the weakly relevant features with satisfactory combination abilities and leave out some strongly relevant features, thereby degrading the generalization performance. To address this issue, we propose a novel embedded framework for feature selection, termed feature selection boosted by unselected features (FSBUF). Specifically, we introduce an extra classifier for unselected features into the traditional embedded model and jointly learn the feature weights to maximize the classification loss of unselected features. As a result, the extra classifier recycles the unselected strongly relevant features to replace the weakly relevant features in the selected feature subset. Our final objective can be formulated as a minimax optimization problem, and we design an effective gradient-based algorithm to solve it. Furthermore, we theoretically prove that the proposed FSBUF is able to improve the generalization ability of traditional embedded feature selection methods. Extensive experiments on synthetic and real-world data sets exhibit the comprehensibility and superior performance of FSBUF.
Shuo Chen 0003, Zhenyong Fu, Fa Zhu, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.3
2021 Contrastive Embedding for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) aims to recognize objects from both seen and unseen classes, when only the labeled examples from seen classes are provided. Recent feature generation methods learn a generative model that can synthesize the missing visual features of unseen classes to mitigate the data-imbalance problem in GZSL. However, the original visual feature space is suboptimal for GZSL classification since it lacks discriminative information. To tackle this issue, we propose to integrate the generation model with the embedding model, yielding a hybrid GZSL framework. The hybrid GZSL approach maps both the real and the synthetic samples produced by the generation model into an embedding space, where we perform the final GZSL classification. Specifically, we propose a contrastive embedding (CE) for our hybrid GZSL framework. The proposed contrastive embedding can leverage not only the class-wise supervision but also the instance-wise supervision, where the latter is usually neglected by existing GZSL researches. We evaluate our proposed hybrid GZSL framework with contrastive embedding, named CE-GZSL, on five benchmark datasets. The results show that our CEGZSL method can outperform the state-of-the-arts by a significant margin on three datasets. Our codes are available on https://github.com/Hanzy1996/CE-GZSL.
Zongyan Han, Zhenyong Fu, Shuo Chen 0003, Jian Yang 0003
CVPR2
2021 Inference guided feature generation for generalized zero-shot learning
Zongyan Han, Zhenyong Fu, Jian Yang 0003
Neurocomputing2
2021 Improved multi-scale dynamic feature encoding network for image demoiréing
Zhenyong Fu, Jian Yang 0003
Pattern Recognit.2
2020 Learning the Redundancy-Free Features for Generalized Zero-Shot Object Recognition
abstract
Zero-shot object recognition or zero-shot learning aims to transfer the object recognition ability among the semantically related categories, such as fine-grained animal or bird species. However, the images of different fine-grained objects tend to merely exhibit subtle differences in appearance, which will severely deteriorate zero-shot object recognition. To reduce the superfluous information in the fine-grained objects, in this paper, we propose to learn the redundancy-free features for generalized zero-shot learning. We achieve our motivation by projecting the original visual features into a new (redundancy-free) feature space and then restricting the statistical dependence between these two feature spaces. Furthermore, we require the projected features to keep and even strengthen the category relationship in the redundancy-free feature space. In this way, we can remove the redundant information from the visual features without losing the discriminative information. We extensively evaluate the performance on four benchmark datasets. The results show that our redundancy-free feature based generalized zero-shot learning (RFF-GZSL) approach can outperform the state-of-the-arts often by a large margin.
Zongyan Han, Zhenyong Fu, Jian Yang 0003
CVPR2
2020 Zero-Shot Image Super-Resolution with Depth Guided Internal Degradation Learning
Zhenyong Fu, Jian Yang 0003
ECCV (17)2
2020 Analytical form of Fisher information matrix of bipoloar-activation-function-based multilayer perceptrons
abstract
For the widely used multilayer perceptrons (MLPs), the existed singularities in the parameter space have seriously affected the learning dynamics of MLPs, which cause several singular learning behaviors. Since the Fisher information matrix (FIM) plays a significant role in analyzing the singular learning dyanmics of MLPs, it is very worthy to obtain the analytical form of FIM to do further investigation. In this paper, by choosing the bipolar error function as the activation function, the analytical form of FIM are obtained, where the validity of the obtained results are verified by taking three experiments.
Weili Guo, Zhenyong Fu, Jianhui Guo, Guochen Pang, Jian Yang 0003
IJCNN3
2018 Multilevel Collaborative Attention Network for Person Search
Zhenyong Fu, Hongtao Lu 0001
ACCV (1)3
2018 Zero-Shot Learning on Semantic Class Prototype Graph
abstract
Zero-Shot Learning (ZSL) for visual recognition is typically achieved by exploiting a semantic embedding space. In such a space, both seen and unseen class labels as well as image features can be embedded so that the similarity among them can be measured directly. In this work, we consider that the key to effective ZSL is to compute an optimal distance metric in the semantic embedding space. Existing ZSL works employ either euclidean or cosine distances. However, in a high-dimensional space where the projected class labels (prototypes) are sparse, these distances are suboptimal, resulting in a number of problems including hubness and domain shift. To overcome these problems, a novel manifold distance computed on a semantic class prototype graph is proposed which takes into account the rich intrinsic semantic structure, i.e., semantic manifold, of the class prototype distribution. To further alleviate the domain shift problem, a new regularisation term is introduced into a ranking loss based embedding model. Specifically, the ranking loss objective is regularised by unseen class prototypes to prevent the projected object features from being biased towards the seen prototypes. Extensive experiments on four benchmarks show that our method significantly outperforms the state-of-the-art.
Zhenyong Fu, Tao Xiang 0002, Elyor Kodirov, Shaogang Gong
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Towards Safe Semi-supervised Classification: Adjusted Cluster Assumption via Clustering
Zhenyong Fu, Hui Xue 0002
Neural Process. Lett.3
2017 Learning from Weak and Noisy Labels for Semantic Segmentation
abstract
A weakly supervised semantic segmentation (WSSS) method aims to learn a segmentation model from weak (image-level) as opposed to strong (pixel-level) labels. By avoiding the tedious pixel-level annotation process, it can exploit the unlimited supply of user-tagged images from media-sharing sites such as Flickr for large scale applications. However, these `free' tags/labels are often noisy and few existing works address the problem of learning with both weak and noisy labels. In this work, we cast the WSSS problem into a label noise reduction problem. Specifically, after segmenting each image into a set of superpixels, the weak and potentially noisy image-level labels are propagated to the superpixel level resulting in highly noisy labels; the key to semantic segmentation is thus to identify and correct the superpixel noisy labels. To this end, a novel L1-optimisation based sparse learning model is formulated to directly and explicitly detect noisy labels. To solve the L1-optimisation problem, we further develop an efficient learning algorithm by introducing an intermediate labelling variable. Extensive experiments on three benchmark datasets show that our method yields state-of-the-art results given noise-free labels, whilst significantly outperforming the existing methods when the weak labels are also noisy.
Zhiwu Lu 0001, Zhenyong Fu, Tao Xiang 0002, Peng Han 0005, Liwei Wang 0001, Xin Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Semi-supervised manifold regularization with adaptive graph construction
Yun Li 0009, Songcan Chen, Zhenyong Fu, Hui Xue 0002
Pattern Recognit. Lett.5
2016 Learning Robust Graph Regularisation for Subspace Clustering
Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong
BMVC3
2016 Person Re-Identification by Unsupervised \ell _1 ℓ 1 Graph Learning
Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong
ECCV (1)3
2015 Zero-shot object recognition by semantic manifold distance
abstract
Object recognition by zero-shot learning (ZSL) aims to recognise objects without seeing any visual examples by learning knowledge transfer between seen and unseen object classes. This is typically achieved by exploring a semantic embedding space such as attribute space or semantic word vector space. In such a space, both seen and unseen class labels, as well as image features can be embedded (projected), and the similarity between them can thus be measured directly. Existing works differ in what embedding space is used and how to project the visual data into the semantic embedding space. Yet, they all measure the similarity in the space using a conventional distance metric (e.g. cosine) that does not consider the rich intrinsic structure, i.e. semantic manifold, of the semantic categories in the embedding space. In this paper we propose to model the semantic manifold in an embedding space using a semantic class label graph. The semantic manifold structure is used to redefine the distance metric in the semantic embedding space for more effective ZSL. The proposed semantic manifold distance is computed using a novel absorbing Markov chain process (AMP), which has a very efficient closed-form solution. The proposed new model improves upon and seamlessly unifies various existing ZSL algorithms. Extensive experiments on both the large scale ImageNet dataset and the widely used Animal with Attribute (AwA) dataset show that our model outperforms significantly the state-of-the-arts.
Zhenyong Fu, Tao A. Xiang, Elyor Kodirov, Shaogang Gong
CVPR1
2015 Unsupervised Domain Adaptation for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) can be considered as a special case of transfer learning where the source and target domains have different tasks/label spaces and the target domain is unlabelled, providing little guidance for the knowledge transfer. A ZSL method typically assumes that the two domains share a common semantic representation space, where a visual feature vector extracted from an image/video can be projected/embedded using a projection function. Existing approaches learn the projection function from the source domain and apply it without adaptation to the target domain. They are thus based on naive knowledge transfer and the learned projections are prone to the domain shift problem. In this paper a novel ZSL method is proposed based on unsupervised domain adaptation. Specifically, we formulate a novel regularised sparse coding framework which uses the target domain class labels' projections in the semantic space to regularise the learned target domain projection thus effectively overcoming the projection domain shift problem. Extensive experiments on four object and action recognition benchmark datasets show that the proposed ZSL method significantly outperforms the state-of-the-arts.
Elyor Kodirov, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong
ICCV3
2015 Pairwise constraint propagation via low-rank matrix recovery
abstract
As a kind of weaker supervisory information, pairwise constraints can be exploited to guide the data analysis process, such as data clustering. This paper formulates pairwise constraint propagation, which aims to predict the large quantity of unknown constraints from scarce known constraints, as a low-rank matrix recovery (LMR) problem. Although recent advances in transductive learning based on matrix completion can be directly adopted to solve this problem, our work intends to develop a more general low-rank matrix recovery solution for pairwise constraint propagation, which not only completes the unknown entries in the constraint matrix but also removes the noise from the data matrix. The problem can be effectively solved using an augmented Lagrange multiplier method. Experimental results on constrained clustering tasks based on the propagated pairwise constraints have shown that our method can obtain more stable results than state-of-the-art algorithms, and outperform them.
Zhenyong Fu
Comput. Vis. Media1
2015 Semi-supervised classification learning by discrimination-aware manifold regularization
Songcan Chen, Hui Xue 0002, Zhenyong Fu
Neurocomputing4
2015 Local similarity learning for pairwise constraint propagation
abstract
Pairwise constraint propagation studies the problem of propagating the scarce pairwise constraints across the entire dataset. Effective propagation algorithms have previously been designed based on the graph-based semi-supervised learning framework. Therefore, these previous constraint propagation methods rely critically on a good similarity measure over the data points. Improper or noisy similarity measurements may dramatically degrade the performance of the constraint propagation algorithms. In this paper, we make attempt to exploit the available pairwise constraints to learn a new set of similarities, which are consistent with the supervisory information in the pairwise constraints, before propagating these initial constraints. Our method is a local learning algorithm. More specifically, we compute the similarities at each data point through simultaneously minimizing the local reconstruction error and local constraint error. The proposed method has been tested in the constrained clustering tasks on eight real-life datasets and then shown to achieve significant improvements with respect to the state of the arts.
Zhenyong Fu, Zhiwu Lu 0001, Horace Ho-Shing Ip, Hongtao Lu 0001
Multim. Tools Appl.1
2014 Transductive Multi-view Embedding for Zero-Shot Recognition and Annotation
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong
ECCV (2)4
2014 Non-negative and sparse spectral clustering
Hongtao Lu 0001, Zhenyong Fu
Pattern Recognit.2
2012 Modalities consensus for multi-modal constraint propagation
abstract
This paper presents a novel modalities consensus framework for multi-modal pairwise constraint propagation (MCP). We first combine multiple single-modal constraint propagation (SCP) problems together, and then explicitly introduce a new modalities consensus regularizer to force the propagation results on different modalities to be consistent with each other. With a separable consensus regularizer, the proposed approach can be effectively solved using an alternating optimization way. More importantly, based on our modalities consensus framework, two single-modal constraint propagation algorithms can be directly reformulated as two well-defined multi-modal solutions. Experimental results on constrained clustering tasks have shown that the proposed framework can achieve significant improvements with respect to the state of the arts.
Zhenyong Fu, Hongtao Lu 0001, Horace Ho-Shing Ip, Zhiwu Lu 0001
ACM Multimedia1
2012 Incremental visual objects clustering with the growing vocabulary tree
Zhenyong Fu, Hongtao Lu 0001
Multim. Tools Appl.1
2011 Symmetric Graph Regularized Constraint Propagation
abstract
This paper presents a novel symmetric graph regularization framework for pairwise constraint propagation. We first decompose the challenging problem of pairwise constraint propagation into a series of two-class label propagation subproblems and then deal with these subproblems by quadratic optimization with symmetric graph regularization. More importantly, we clearly show that pairwise constraint propagation is actually equivalent to solving a Lyapunov matrix equation, which is widely used in Control Theory as a standard continuous-time equation. Different from most previous constraint propagation methods that suffer from severe limitations, our method can directly be applied to multi-class problem and also can effectively exploit both must-link and cannot-link constraints. The propagated constraints are further used to adjust the similarity between data points so that they can be incorporated into subsequent clustering. The proposed method has been tested in clustering tasks on six real-life data sets and then shown to achieve significant improvements with respect to the state of the arts.
Zhenyong Fu, Zhiwu Lu 0001, Horace Ho-Shing Ip, Yuxin Peng 0001, Hongtao Lu 0001
AAAI1
2011 Multi-modal constraint propagation for heterogeneous image clustering
abstract
This paper presents a multi-modal constraint propagation approach to exploiting pairwise constraints for constrained clustering tasks on multi-modal datasets. Pairwise constraint propagation methods have previously been designed primarily for single modality data and cannot be directly applied to multi-modal data or a dataset with multiple representations. In this paper, we provide an effective solution to the multi-modal constraint propagation problem by decomposing it into a set of independent multi-graph based two-class label propagation subproblems which are then merged into a unified problem and solved by quadratic optimization. We also show that such a formulation yields a closed-form solution. Our approach allows the initial pairwise constraints to be propagated throughout the entire multi-modal dataset. The propagated constraints are further used to refine the similarities between the objects for subsequent clustering tasks. The proposed method has been tested in constrained clustering tasks on two real-life multi-modal image datasets and shown to achieve significant improvements with respect to the single modality methods.
Zhenyong Fu, Horace Ho-Shing Ip, Hongtao Lu 0001, Zhiwu Lu 0001
ACM Multimedia1