Fuyuan Hu

dblp:135/9619 · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0002-6818-2221ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 17 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 14 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
abstract
Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency.
Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003
AAAI5
2026 E-Logic Prompt: Unified Energy-Logic Framework for Continual Visual Question Answering
abstract
Prompt tuning has shown promise for continual visual question answering (CVQA), facilitating modular and transferable knowledge across tasks. However, existing approaches often overlook the guiding role of prompts in the model’s implicit reasoning process. This oversight can lead to inconsistent reasoning paths and performance degradation across tasks. To address this issue, we propose the E Logic Prompt framework, which employs energy-based models (EBMs) to model the semantic compatibility between prompts and queries. In this framework, prompts function not only as adapters but also as reasoning guides that help maintain coherence throughout the inference process. The framework enforces logical consistency at three levels. At the input level, it selects semantically aligned prompts by minimizing the energy between queries and prompts. Within the model, it aligns intermediate representations with prompts across layers to preserve step-by-step reasoning. Across tasks, it applies energy-based constraints to regulate prompt behavior, effectively suppressing semantic drift and enabling prompt reuse. These three levels of consistency together enhance the guiding capacity of prompts, allowing them to steer the model toward more stable and coherent reasoning. Extensive experiments show that E Logic Prompt outperforms existing methods in both accuracy and knowledge retention, while effectively maintaining balanced cross-modal reasoning throughout continual learning.
Jiayao Tan, Fuyuan Hu, Wei Feng 0005
AAAI3
2026 VRQA: Context-adaptive view routing for long-document question-answer generation
abstract
Long-document question–answer (QA) generation often employs large-scale generation with strong filtering to reduce off-topic deviations. However, such strategies tend to concentrate generation on a few safe entry points, leading to uneven coverage and similar question expression. Conversely, expanding the scope of questioning may weaken topical consistency and the verifiability of supporting evidence. To address these challenges, we propose VRQA, a long-document QA generation framework with context-adaptive view routing under hierarchical anchor constraints. It selects better question entry points for different chunks and generates QA pairs with more dispersed semantic focuses. VRQA first constructs a semantic representation from chunk-level evidence and key points, then encodes view names and their descriptions into view prototypes, and obtains view-aware contextual features through conditional modulation. Subsequently, it updates the view routing policy online through reinforcement learning, where the reward is constructed using anchor consistency and QA answer quality scores from an evaluator. This enables VRQA to adaptively select the optimal view set and coordinate with submodular selection to achieve a dynamic balance between quality and diversity. To implement this framework, VRQA adopts a three-stage coupled training schedule, consisting of router warm-up, view-conditioned representation learning, and online updating, through which the router is progressively refined by evaluator feedback. Experiments on three vertical domains from the Elsevier OA CC-BY dataset show that VRQA achieves state-of-the-art performance, improving quality by 11.46% on AS-Chk and semantic diversity by 7.96% on VS over the strongest baseline, and thus offering a better trade-off between quality and diversity for long-document QA generation.
Mengting Huang, Hongjie Wu, Fuyuan Hu, Lanhui Liu, Qiming Fu 0001
Expert Syst. Appl.5
2026 Frozen-Fusion network with spatial-temporal learning for video action recognition
Zhancheng Zhang, Zuxi Zhang, Wenhao Tao, Xiaoqing Luo, Fuyuan Hu
Pattern Recognit.5
2026 Mitigating Catastrophic Forgetting in Online Continual Learning With Dual-Margin Contrastive Replay
abstract
Online Continual Learning (OCL) enables machine learning models to learn from a stream of non-stationary tasks, making it more aligned with real-world scenarios. However, OCL faces a significant challenge: catastrophic forgetting, wherein the model learned in previous tasks is substantially overwritten upon encountering new tasks, leading to a biased forgetting of prior knowledge. Among various OCL strategies, replay-based methods have proven particularly effective in mitigating catastrophic forgetting by maintaining a small buffer of past samples and retraining them alongside new data. However, due to strict memory constraints, these replay buffers often fail to adequately represent the true data distribution of previous tasks. This leads to distributional shifts in the feature space, amplifying forgetting and degrading model performance. To address the problem, in this paper, we propose a novel replay strategy, termed Dual-Margin Contrastive Replay (DMCR), to anchor the distribution of old tasks and reduce the negative transfer effects. First, we propose to select memory for more representative samples guided by constructed centroids in a data stream. Then, to keep the model from distribution chaos in biased replay, a two-level angular cross-task Contrastive Margin Loss (CML) is proposed, to encourage the intra-class and intra-task compactness, and increase the inter-class and inter-task discrepancy. Finally, to further suppress the distributional drift, we present an optional Centroid Distillation Loss (CDL) on the replay memory to anchor the knowledge in feature space for each previous old task. Extensive experimental results on five benchmark datasets validate that the proposed DMCR can effectively mitigate the catastrophic forgetting and achieve state-of-the-art (SOTA) performance in OCL.
Fan Lyu, Gongbo Cheng, Daofeng Liu, Linglan Zhao, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 GAIN: Global-Atomic INteraction Graph for Few-Shot Class-Incremental Learning
Fan Lyu, Linglan Zhao, Chengyan Liu, Yinying Mei, Zhang Zhang 0001, Baoqing Yu, Fuyuan Hu, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.8
2025 Rebalancing Multi-Label Class-Incremental Learning
abstract
Multi-label class-incremental learning (MLCIL) is essential for real-world multi-label applications, allowing models to learn new labels while retaining previously learned knowledge continuously. However, recent MLCIL approaches can only achieve suboptimal performance due to the oversight of the positive-negative imbalance problem, which manifests at both the label and loss levels because of the task-level partial label issue. The imbalance at the label level arises from the substantial absence of negative labels, while the imbalance at the loss level stems from the asymmetric contributions of the positive and negative loss parts to the optimization. To address the issue above, we propose a Rebalance framework for both the Loss and Label levels (RebLL), which integrates two key modules: asymmetric knowledge distillation (AKD) and online relabeling (OR). AKD is proposed to rebalance at the loss level by emphasizing the negative label learning in classification loss and down-weighting the contribution of overconfident predictions in distillation loss. OR is designed for label rebalance, which restores the original class distribution in memory by online relabeling the missing classes. Our comprehensive experiments on the PASCAL VOC and MS-COCO datasets demonstrate that this rebalancing strategy significantly improves performance, achieving new state-of-the-art results even with a vanilla CNN backbone.
Kaile Du, Fan Lyu, Yuyang Li 0005, Junzhou Xie, Yixi Shen, Fuyuan Hu, Guangcan Liu
AAAI7
2025 Maintaining Consistent Inter-Class Topology in Continual Test-Time Adaptation
abstract
This paper introduces Topological Consistency Adaptation (TCA), a novel approach to Continual Test-time Adaptation (CTTA) that addresses the challenges of domain shifts and error accumulation in testing scenarios. TCA ensures the stability of inter-class relationships by enforcing a class topological consistency constraint, which minimizes the distortion of class centroids and preserves the topological structure during continuous adaptation. Additionally, we propose an intra-class compactness loss to maintain compactness within classes, indirectly supporting inter-class stability. To further enhance model adaptation, we introduce a batch imbalance topology weighting mechanism that accounts for class distribution imbalances within each batch, optimizing centroid distances and stabilizing the inter-class topology. Experiments show that our method demonstrates improvements in handling continuous domain shifts, ensuring stable feature distributions and boosting predictive performance. Our code is available at: https://github.com/Successybbdwm/TCA.
Chenggong Ni, Fan Lyu, Jiayao Tan, Fuyuan Hu, Tao Zhou 0009
CVPR4
2025 Less Over More: Interference Sample Gradient Purification For Parallel Continual Learning
abstract
The goal of Parallel Continual Learning (PCL) is to continually learn multi-task from new data stream and complete the corresponding tasks. Previous research on PCL ignored inter-task interference, which may hinder knowledge transfer and exacerbate catastrophic forgetting. Therefore, in this paper, we investigate the interference problem of PCL in dynamic multi-task scenarios. First, we construct Global Discrimination Threshold to detect interference sample, and removing the interference gradient in joint multi-gradient algorithm, termed Interference Sample Gradient Purification (ISGP). Second, we introduced the Guardian For Pivotal Memory Sample (G-PMS) as new evaluation criterion and proposed Local Maintenance Thresholds to prevent accidental exclusion of memory samples that play an important role in suppressing catastrophic forgetting. Finally, the results of the experimental evaluation clearly confirm that ISGP can effectively enhance the transfer of new knowledge and better suppress catastrophic forgetting.
Tingyang Lu, Jiayao Tan, Fuyuan Hu
ICASSP4
2025 Controllable Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA) is an emerging and challenging task where a model trained in a source domain must adapt to continuously changing conditions during testing, without access to the original source data. CTTA is prone to error accumulation due to uncontrollable domain shifts, leading to blurred decision boundaries between categories. Existing CTTA methods primarily focus on suppressing domain shifts, which proves inadequate during the unsupervised test phase. In contrast, we introduce a novel approach that guides rather than suppresses these shifts. Specifically, we propose Controllable Continual Test-Time Adaptation (C-CoTTA), which explicitly prevents any single category from encroaching on others, thereby mitigating the mutual influence between categories caused by uncontrollable shifts. Moreover, our method reduces the sensitivity of model to domain transformations, thereby minimizing the magnitude of category shifts. Extensive quantitative experiments demonstrate the effectiveness of our method, while qualitative analyses, such as t-SNE plots, confirm the theoretical validity of our approach. Our code is available at https://github.com/RenshengJi/C-CoTTA.
Ziqi Shi, Fan Lyu, Fanhua Shang, Fuyuan Hu, Wei Feng 0005, Zhang Zhang 0001, Liang Wang 0001
ICME5
2025 DAA: Amplifying Unknown Discrepancy for Test-Time Discovery
abstract
Test-Time Discovery (TTD) addresses the critical challenge of identifying and adapting to novel classes during inference while maintaining performance on known classes, which is a capability essential for dynamic real-world environments such as healthcare and autonomous driving. Recent TTD methods adopt training-free, memory-based strategies but rely on frozen models and static representations, resulting in poor generalization. In this paper, we propose a Discrepancy-Amplifying Adapter (DAA), a trainable module that enables real-time adaptation by amplifying feature-level discrepancies between known and unknown classes. During training, DAA is optimized using simulated unknowns and a novel warm-up strategy to enhance its discriminative capacity. To ensure continual adaptation at test time, we introduce a Short-Term Memory Renewal (STMR) mechanism, which maintains a queue-based memory for unknown classes and selectively refreshes prototypes using recent, reliable samples. DAA is further updated through self-supervised learning, promoting knowledge retention for known classes while improving discrimination of emerging categories. Extensive experiments show that our method maintains high adaptability and stability, and significantly improves novel class discovery performance. Our code will be available.
Fan Lyu, Chenggong Ni, Zhang Zhang 0001, Fuyuan Hu, Liang Wang 0001
NeurIPS5
2025 LinFa-Q: Accurate Q-learning with linear function approximation
Zhechao Wang, Qiming Fu 0001, Quan Liu 0004, You Lu 0004, Hongjie Wu, Fuyuan Hu
Neurocomputing7
2025 Historical Object-Aware Prompt Learning for Universal Hyperspectral Object Tracking
abstract
Hyperspectral Object Tracking (HOT), utilizing rich spectral information from hyperspectral video (HSV), holds significant importance for object tracking. We identify that a major obstacle in improving HOT performance lies in effectively leveraging spectral and historical information. Furthermore, due to the mismatch in band dimensions between hyperspectral and RGB images, state-of-the-art RGB-based trackers struggle to adapt to unified HOT tasks. To address this, we propose a Historical Object-Aware Prompt Learning (HOPL) method for universal hyperspectral object tracking. Initially, we transform hyperspectral image ( \( N \) bands) into multiple sets of three bands with different combinations and feed them into a backbone network to generate base features. Subsequently, we introduce a historical object-aware prompter, where historical object-aware images are input to generate prompt features that enhance the representation of object information when combined with base features. Additionally, we design a band information fusion module to integrate the multiple sets of base features. By introducing historical object-aware prompts, HOPL significantly enhances tracking performance without retraining the backbone network. Experimental results on the HOT2023 dataset (comprising HSV with 25-band, 16-band, and 15-band wavelength ranges) and HOT2022 dataset validate the superiority of HOPL over state-of-the-art methods. The source code is available at https://github.com/rayyao/HOPL .
Rui Yao 0006, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Multi-Label Continual Learning Using Augmented Graph Convolutional Network
abstract
Multi-Label Continual Learning (MLCL) is a framework designed for class-incremental multi-label image recognition. However, MLCL faces two critical challenges: the construction of label relationships onpast-missing and future-missing partial labelsof training data, and the problem ofcatastrophic forgetting, which leads to poor generalization. To address these challenges, this study proposes an enhanced version of the Augmented Graph Convolutional Network (AGCN++), capable of constructing cross-task label relationships and mitigating catastrophic forgetting. First, an Augmented Correlation Matrix (ACM) is constructed across all observed classes, incorporating intra-task relationships derived from hard label statistics. Additionally, inter-task relationships are established by leveraging both hard and soft labels obtained from the data, as well as a constructed expert network. Next, a novel partial label encoder (PLE) is introduced for MLCL, enabling the extraction of dynamic class representations for each partial label image as graph nodes. This PLE also facilitates the generation of soft labels, which contribute to the creation of a more persuasive ACM and effectively mitigate forgetting. Lastly, a relationship-preserving constrainter is proposed to address the issue of forgetting label dependencies across old tasks. In the AGCN++, the label relationships topology can be augmented automatically, thereby generating efficient class representations. The effectiveness of the proposed method is evaluated using two multi-label image benchmarks. The experimental results demonstrate that the proposed approach is highly effective in the context of MLCL image recognition. It can establish compelling correlations across tasks, even in scenarios where the old task labels are missing.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Hanjing Cheng
IEEE Trans. Multim.4
2023 Centroid Distance Distillation for Effective Rehearsal in Continual Learning
abstract
Rehearsal, retraining on a stored small data subset of old tasks, has been proven effective in solving catastrophic forgetting in continual learning. However, due to the sampled data may have a large bias towards the original dataset, retraining them is susceptible to driving continual domain drift of old tasks in feature space, resulting in forgetting. In this paper, we focus on tackling the continual domain drift problem with centroid distance distillation. First, we propose a centroid caching mechanism for sampling data points based on constructed centroids to reduce the sample bias in rehearsal. Then, we present a centroid distance distillation that only stores the centroid distance to reduce the continual domain drift. The experiments on four continual learning datasets show the superiority of the proposed method, and the continual domain drift can be reduced. Our code is available at https://github.com/Daofeng-liu/CDD-R.
Daofeng Liu, Fan Lyu, Zhenping Xia, Fuyuan Hu
ICASSP5
2023 Multi-semantic hypergraph neural network for effective few-shot learning
Hao Chen 0011, Fuyuan Hu, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Zhenping Xia
Pattern Recognit.3
2023 Attention-guided Adversarial Attack for Video Object Segmentation
abstract
Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.
Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao
ACM Trans. Intell. Syst. Technol.4
2022 AGCN: Augmented Graph Convolutional Network for Lifelong Multi-Label Image Recognition
abstract
The Lifelong Multi-Label (LML) image recognition builds an online class-incremental classifier in a sequential multilabel image recognition data stream. However, training on the data with different Partial Labels may result in more serious Catastrophic Forgetting in old classes. To solve the problem, the study proposes an Augmented Graph Convolutional Network (AGCN)to build an Augmented Correlation Matrix (ACM) across the sequential partial-label tasks and sustain the catastrophic forgetting. First, in ACM, the intra-task relations derive from the hard label statistics, while the inter-task relations further leverage the soft labels from a stored expert network. Then, based on the ACM, AGCN captures label dependencies with dynamic augmented structure and yields effective class representations. Our method is evaluated on two multi-label image benchmarks and the results show that the proposed method is effective for LML image recognition.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Qiming Fu 0001
ICME3
2022 Harnessing Multi-Semantic Hypergraph for Few-Shot Learning
Hao Chen 0011, Zhenping Xia, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Fuyuan Hu
PRCV (1)8
2022 Energy-efficient control of thermal comfort in multi-zone residential HVAC via reinforcement learning
abstract
Energy efficient control of thermal comfort has been already an important part of residential heating, ventilation, and air conditioning (HVAC) systems. However, the optimisation of energy saving control for thermal comfort is not an easy task due to the complex dynamics of HVAC systems, the dynamics of thermal comfort and the trade-off between energy saving and thermal comfort. To solve the above problem, we propose a deep reinforcement learning-based thermal comfort control method in multi-zone residential HVAC. In this paper, firstly we design a SVR-DNN model, consisting of Support Vector Regression and a Deep Neural Network to predict thermal comfort value. Then, we apply Deep Deterministic Policy Gradient (DDPG) based on the output of the SVR-DNN model to achieve an optimal HVAC thermal comfort control strategy. This method can minimise energy consumption while satisfying occupants' thermal comfort. The experimental results show that our method can improve thermal comfort prediction performance by 20.5% compared with DNN; compared with deep Q-network (DQN), energy consumption and thermal comfort violation can be reduced by 3.52% and 64.37% respectively.
Zhengkai Ding, Qiming Fu 0001, Hong-Jie Wu, You Lu 0004, Fuyuan Hu
Connect. Sci.6
2022 Efficient lightweight video person re-identification with online difference discrimination module
Cunyuan Gao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Fuyuan Hu
Multim. Tools Appl.6
2022 Visual Grounding Via Accumulated Attention
abstract
Visual grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. Generally, it requires the machine to first understand the query, identify the key concepts in the image, and then locate the target object by specifying its bounding box. However, in many real-world visual grounding applications, we have to face with ambiguous queries and images with complicated scene structures. Identifying the target based on highly redundant and correlated information can be very challenging, and often leading to unsatisfactory performance. To tackle this, in this paper, we exploit an attention module for each kind of information to reduce internal redundancies. We then propose an accumulated attention (A-ATT) mechanism to reason among all the attention modules jointly. In this way, the relation among different kinds of information can be explicitly captured. Moreover, to improve the performance and robustness of our VG models, we additionally introduce some noises into the training procedure to bridge the distribution gap between the human-labeled training data and the real-world poor quality data. With this "noised" training strategy, we can further learn a bounding box regressor, which can be used to refine the bounding box of the target object. We evaluate the proposed methods on four popular datasets (namely ReferCOCO, ReferCOCO+, ReferCOCOg, and GuessWhat?!). The experimental results show that our methods significantly outperform all previous works on every dataset in terms of accuracy.
Chaorui Deng, Qi Wu 0001, Qingyao Wu, Fuyuan Hu, Fan Lyu, Mingkui Tan
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Disentangling Semantic-to-Visual Confusion for Zero-Shot Learning
abstract
Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate realistic visual distributions from semantics by automatically searching discriminative representations. However, the traditional TL cannot search reliable unseen disentangled representations due to the unavailability of unseen classes in ZSL. To alleviate this drawback, we propose in this work a multi-modal triplet loss (MMTL) which utilizes multi-modal information to search adisentangledrepresentation space. As such, all classes can interplay which can benefit learning disentangled class representations in the searched space. Furthermore, we develop a novel model called Disentangling Class Representation Generative Adversarial Network (DCR-GAN) focusing on exploiting the disentangled representations in training, feature synthesis, and final recognition stages. Benefiting from the disentangled representations, DCR-GAN could fit a more realistic distribution over both seen and unseen features. Extensive experiments show that our proposed model can lead to superior performance to the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fuyuan Hu, Fan Lyu, Kaizhu Huang
IEEE Trans. Multim.2
2021 Multi-Domain Multi-Task Rehearsal for Lifelong Learning
abstract
Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suffer from the unpredictable domain shift when training the new task. This is because these methods always ignore two significant factors. First, the Data Imbalance between the new task and old tasks that makes the domain of old tasks prone to shift. Second, the Task Isolation among all tasks will make the domain shift toward unpredictable directions; To address the unpredictable domain shift, in this paper, we propose Multi-Domain Multi-Task (MDMT) rehearsal to train the old tasks and new task parallelly and equally to break the isolation among tasks. Specifically, a two-level angular margin loss is proposed to encourage the intra-class/task compactness and inter-class/task discrepancy, which keeps the model from domain chaos. In addition, to further address domain shift of the old tasks, we propose an optional episodic distillation loss on the memory to anchor the knowledge for each old task. Experiments on benchmark datasets validate the proposed approach can effectively mitigate the unpredictable domain shift.
Fan Lyu, Wei Feng 0005, Zihan Ye, Fuyuan Hu, Song Wang 0002
AAAI5
2021 Each Attribute Matters: Contrastive Attention for Sentence-based Image Editing
Liuqing Zhao, Fan Lyu, Fuyuan Hu, Kaizhu Huang, Fenglei Xu
BMVC3
2021 Multi-zone Residential HVAC Control with Satisfying Occupants' Thermal Comfort Requirements and Saving Energy via Reinforcement Learning
Zhengkai Ding, Qiming Fu 0001, Hongjie Wu, You Lu 0004, Fuyuan Hu
PDCAT6
2021 FocusGAN: Preserving Background in Text-Guided Image Editing
abstract
Text-guided image editing (TIE) seeks to manipulate images by the guidance of language. However, the existing TIE methods always overlook the target-irrelevant pixels and the editing may make the background discolored, distorted, or partially disappear. To overcome this problem, we propose a novel TIE method named FocusGAN, which edits the text-relevant pixels precisely as well as keeps the background invariant. Specifically, we build a network of two stages. In each stage, we first construct a channel-wise subject focusing attention to make the generator focus on the sub-region that best matches each word. Then, the word-level discriminator provides fine-grained feedback by correlating words with image areas, so that the generator can manipulate specific visual attributes without affecting the background. Last, we propose a background-keeping cyclic loss to further improve the invariance of the background and to encourage the edit of the subject that matches the given text. Experiments on CUB and Oxford datasets demonstrate that our approach can effectively keep the background invariant in manipulating images using natural language descriptions.
Liuqing Zhao, Fuyuan Hu, Zhenping Xia
Int. J. Pattern Recognit. Artif. Intell.3
2020 Associating Multi-Scale Receptive Fields For Fine-Grained Recognition
abstract
Extracting and fusing part features have become the key of fined-grained image recognition. Recently, Non-local (NL) module has shown excellent improvement in image recognition. However, it lacks the mechanism to model the interactions between multi-scale part features, which is vital for fine-grained recognition. In this paper, we propose a novel cross-layer non-local (CNL) module to associate multi-scale receptive fields by two operations. First, CNL computes correlations between features of a query layer and all response layers. Second, all response features are weighted according to the correlations and are added to the query features. Due to the interactions of cross-layer features, our model builds spatial dependencies among multi-level layers and learns more discriminative features. In addition, we can reduce the aggregation cost if we set low-dimensional deep layer as query layer. Experiments are conducted to show our model achieves or surpasses state-of-the-art results on three benchmark datasets of fine-grained classification. Our codes can be found at github.com/FouriYe/CNL-ICIP2020.
Zihan Ye, Fuyuan Hu, Zhenping Xia, Fan Lyu, Pengqing Liu
ICIP2
2020 GAN-based person search via deep complementary classifier with center-constrained Triplet loss
Rui Yao 0006, Cunyuan Gao, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu
Pattern Recognit.6
2019 SR-GAN: Semantic Rectifying Generative Adversarial Network for Zero-shot Learning
abstract
The existing Zero-Shot learning (ZSL) methods may suffer from the vague class attributes that are highly overlapped for different classes. Unlike these methods that ignore the discrimination among classes, in this paper, we propose to classify unseen image by rectifying the semantic space guided by the visual space. First, we pre-train a Semantic Rectifying Network (SRN) to rectify semantic space with a semantic loss and a rectifying loss. Then, a Semantic Rectifying Generative Adversarial Network (SR-GAN) is built to generate plausible visual feature of unseen class from both semantic feature and rectified semantic feature. To guarantee the effectiveness of rectified semantic features and synthetic visual features, a pre-reconstruction and a post reconstruction networks are proposed, which keep the consistency between visual feature and semantic feature. Experimental results demonstrate that our approach significantly outperforms the state-of-the-arts on four benchmark datasets.
Zihan Ye, Fan Lyu, Qiming Fu 0001, Jinchang Ren, Fuyuan Hu
ICME6
2019 Structure-aware person search with self-attention and online instance aggregation matching
Cunyuan Gao, Rui Yao 0006, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu, Leida Li
Neurocomputing5
2019 Variational Bayesian Exploration-Based Active Sarsa Algorithm
abstract
We proposed an improved variational Bayesian exploration-based active Sarsa (VBE-ASAR) algorithm, which tries to balance the exploration and exploitation dilemma, and speeds up the convergence rate. First, in the learning process, variational Bayesian method is adopted to measure the information gain, which is used as an exploration factor to construct an internal reward function for heuristic exploration. In addition, before the learning process, in order to improve the exploration performance, transfer learning is used to initialize the value function, where Bisimulation metric is introduced to measure the distance between two states from the source MDP and the target MDP, respectively. Finally, we apply the proposed algorithm to the cliff walking problem, and compare with the Sarsa algorithm, the Q-Learning algorithm, the VFT-Sarsa algorithm and the Bayesian Sarsa (BS) algorithm. Experimental results show that the VBE-ASAR algorithm has a faster learning rate.
Qiming Fu 0001, Zhengxia Yang, You Lu 0004, Hongjie Wu, Fuyuan Hu
Int. J. Pattern Recognit. Artif. Intell.5
2019 Attend and Imagine: Multi-Label Image Classification With Visual Attention and Recurrent Neural Networks
abstract
Real images often have multiple labels, i.e., each image is associated with multiple objects or attributes. Compared to single-label image classification, the multilabel classification problem is much more challenging due to several issues. At first, multiple objects can be anywhere in the image. Second, the importance of different regions in an image is different, and the regions of interest in a multilabel image can be very different from another one. Finally, multiple labels of an image can have label dependencies due to complex image structures. To address these challenges, in this paper, we propose to predict the labels sequentially by applying the recurrent neural networks (RNNs), which are used to encode the label dependencies. When predicting a specific label, we introduce a dynamic attention mechanism to enable the model to focus on only regions of interest in the image. Two benchmark datasets (i.e., Pascal VOC and MS-COCO) are adopted to demonstrate the effectiveness of our work. Moreover, we construct a new dataset, which includes many semantic dependent labels in each image, to verify the effectiveness of our model. Experimental results show that our method outperforms several state-of-the-arts, especially when predicting some semantic relative labels.
Fan Lyu, Qi Wu 0001, Fuyuan Hu, Qingyao Wu, Mingkui Tan
IEEE Trans. Multim.3
2018 Visual Grounding via Accumulated Attention
abstract
Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence or even a multi-round dialogue. There are three main challenges in VG: 1) what is the main focus in a query; 2) how to understand an image; 3) how to locate an object. Most existing methods combine all the information curtly, which may suffer from the problem of information redundancy (i.e. ambiguous query, complicated image and a large number of objects). In this paper, we formulate these challenges as three attention problems and propose an accumulated attention (A-ATT) mechanism to reason among them jointly. Our A-ATT mechanism can circularly accumulate the attention for useful information in image, query, and objects, while the noises are ignored gradually. We evaluate the performance of A-ATT on four popular datasets (namely Refer-COCO, ReferCOCO+, ReferCOCOg, and Guesswhat?!), and the experimental results show the superiority of the proposed method in term of accuracy.
Chaorui Deng, Qi Wu 0001, Qingyao Wu, Fuyuan Hu, Fan Lyu, Mingkui Tan
CVPR4
2017 Solving long haul airline disruption problem caused by groundings using a distributed fixed-point computational approach to integer programming
Zhengtian Wu, Benchi Li, Chuangyin Dang, Fuyuan Hu, Qixin Zhu, Baochuan Fu
Neurocomputing4
2015 Learning graph structure for multi-label image classification via clique generation
abstract
Exploiting label dependency for multi-label image classification can significantly improve classification performance. Probabilistic Graphical Models are one of the primary methods for representing such dependencies. The structure of graphical models, however, is either determined heuristically or learned from very limited information. Moreover, neither of these approaches scales well to large or complex graphs. We propose a principled way to learn the structure of a graphical model by considering input features and labels, together with loss functions. We formulate this problem into a max-margin framework initially, and then transform it into a convex programming problem. Finally, we propose a highly scalable procedure that activates a set of cliques iteratively. Our approach exhibits both strong theoretical properties and a significant performance improvement over state-of-the-art methods on both synthetic and real-world data sets.
Mingkui Tan, Qinfeng Shi, Anton van den Hengel, Chunhua Shen, Junbin Gao, Fuyuan Hu, Zhen Zhang 0008
CVPR6
2015 An adaptive approach for texture enhancement based on a fractional differential operator with non-integer step and order
Fuyuan Hu, Shaohui Si, Hau-San Wong, Baochuan Fu, Maoxin Si
Neurocomputing1