Feiyu Chen 0001

dblp:59/9259-1 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-0928-6899ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FCVSR: A Frequency-Aware Method for Compressed Video Super-Resolution
abstract
Compressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information in the frequency domain, showing great promise in super-resolution performance. However, these methods do not differentiate various frequency subbands spatially or capture the temporal frequency dynamics, potentially leading to suboptimal results. In this paper, we propose a deep frequency-based compressed video SR model (FCVSR) consisting of a motion-guided adaptive alignment (MGAA) network and a multi-frequency feature refinement (MFFR) module. Additionally, a frequency-aware contrastive loss is proposed for training FCVSR, in order to reconstruct finer spatial details. The proposed model has been evaluated on three public compressed video super-resolution datasets, with results demonstrating its effectiveness when compared to existing works in terms of super-resolution performance and complexity.
Fan Zhang 0017, Feiyu Chen 0001, Shuyuan Zhu, David Bull 0001, Bing Zeng 0001
IEEE Trans. Multim.3
2026 Cross-Modal Full-Mode Fine-Grained Alignment for Text-to-Image Person Retrieval
abstract
Text-to-Image Person Retrieval (TIPR) is a cross-modal matching task designed to identify the person images that best correspond to a given textual description. The key difficulty in TIPR is to realize robust correspondence between the textual and visual modalities within a unified latent representation space. To address this challenge, prior approaches incorporate attention mechanisms for implicit cross-modal local alignment. However, they lack the ability to verify whether all local features are correctly aligned. Moreover, existing methods tend to emphasize the utilization of hard negative samples during model optimization to strengthen discrimination between positive and negative pairs, often neglecting incorrectly matched positive pairs. To mitigate these problems, we propose FMFA, a cross-modal Full-Mode Fine-Grained Alignment framework, which enhances global matching through Explicit Fine-Grained Alignment (EFA) and existing implicit relational reasoning—hence the term “full-mode”—without introducing extra supervisory signals. In particular, we propose an Adaptive Similarity Distribution Matching (A-SDM) module to rectify unmatched positive sample pairs. A-SDM adaptively pulls the unmatched positive pairs closer in the joint embedding space, thereby achieving more precise global alignment. Additionally, we introduce an EFA module, which makes up for the lack of verification capability of implicit relational reasoning. EFA strengthens explicit cross-modal fine-grained interactions by sparsifying the similarity matrix and employs a hard coding method for local alignment. We evaluate our method on three public datasets, where it attains state-of-the-art results among all global matching methods. The code for our method is publicly accessible at https://github.com/yinhao1102/FMFA .
Xin Man, Feiyu Chen 0001, Jie Shao 0001, Heng Tao Shen
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Task-Aware Optimized Color Image Demosaicing
abstract
In this paper, we propose a task-aware deep demosaicing network that is designed to produce images for object detection and image compression, targeting high performance for both tasks. The proposed network consists of a color restoration module and a semantic enhancement module. Specifically, the color restoration module converts the Bayer-pattern raw images into full-color images. The semantic enhancement module integrates the semantic information via an adaptive feature fusion to enhance the task-relevant features while suppressing the task-irrelevant content to save bitrate. Experimental results demonstrate that using the color images produced by our demosaicing network can achieve a better trade-off between detection accuracy and compression efficiency.
Feiyu Chen 0001, Shuyuan Zhu, Bing Zeng 0001
MMSP4
2025 MLFormer: Unleashing Efficiency Without Attention for Multimodal Knowledge Graph Embedding
abstract
Multimodal knowledge graphs (MMKGs) have gained widespread adoption across various domains. However, existing transformer-based methods for MMKG representation learning primarily focus on enhancing representation performance, while overlooking time and memory costs, which reduces model efficiency. To tackle these limitations, we introduce a multimodal lightweight transformer (MLFormer) model, which not only ensures robust representation capabilities but also considerably improves computational efficiency. We find that the self-attention mechanism in transformers leads to substantial performance overheads. As a result, we optimize the traditional MMKGE model in two aspects: modality processing and modality fusion, by incorporating a filter gate and Fourier transform. Our experimental results on real-world multimodal knowledge graph completion datasets demonstrate that MLFormer achieves significant improvements in computational efficiency while maintaining competitive performance.
Meng Wang 0001, Changyu Li, Feiyu Chen 0001, Jie Shao 0001, Ke Qin, Shuang Liang 0002
IEEE Trans. Comput. Soc. Syst.3
2025 DVSRNet: Deep Video Super-Resolution Based on Progressive Deformable Alignment and Temporal-Sparse Enhancement
abstract
Video super-resolution (VSR) is used to compose high-resolution (HR) video from low-resolution video. Recently, the deformable alignment-based VSR methods are becoming increasingly popular. In these methods, the features extracted from video are aligned to eliminate the motion error targeting high super-resolution (SR) quality. However, these methods often suffer from misalignment and the lack of enough temporal information to compose HR frames, which accordingly induce artifacts in the SR result. In this article, we design a deep VSR network (DVSRNet) based on the proposed progressive deformable alignment (PDA) module and temporal-sparse enhancement (TSE) module. Specifically, the PDA module is designed to accurately align features and to eliminate artifacts via the bidirectional information propagation. The TSE module is constructed to further eliminate artifacts and to generate clear details for the HR frame. In addition, we construct a lightweight deep optical flow network (OFNet) to obtain the bidirectional optical flows for the implementation of the PDA module. Moreover, two new loss functions are designed for our proposed method. The first one is adopted in OFNet and the second one is constructed to guarantee the generation of sharp and clear details for the HR frames. The experimental results demonstrate that our method performs better than the state-of-the-art methods.
Feiyu Chen 0001, Shuyuan Zhu, Yu Liu 0091, Ruiqin Xiong, Bing Zeng 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and Skills
abstract
Offline meta-reinforcement learning (meta-RL) methods, which adapt to unseen target tasks with prior experience, are essential in robot control tasks. Current methods typically utilize task contexts and skills as prior experience, where task contexts are related to the information within each task and skills represent a set of temporally extended actions for solving subtasks. However, these methods still suffer from limited performance when adapting to unseen target tasks, mainly because the learned prior experience lacks generalization, i.e., they are unable to extract effective prior experience from meta-training tasks by exploration and learning of continuous latent spaces. We propose a framework called decoupled meta-reinforcement learning (DCMRL), which (1) contrastively restricts the learning of task contexts through pulling in similar task contexts within the same task and pushing away different task contexts of different tasks, and (2) utilizes a Gaussian quantization variational autoencoder (GQ-VAE) for clustering the Gaussian distributions of the task contexts and skills respectively, and decoupling the exploration and learning processes of their spaces. These cluster centers which serve as representative and discrete distributions of task context and skill are stored in task context codebook and skill codebook, respectively. DCMRL can acquire generalizable prior experience and achieve effective adaptation to unseen target tasks during the meta-testing phase. Experiments in the navigation and robot manipulation continuous control tasks show that DCMRL is more effective than previous meta-RL methods with more generalizable prior experience.
Hongcai He, Anjie Zhu, Shuang Liang 0002, Feiyu Chen 0001, Jie Shao 0001
AAAI4
2024 OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal
Qiao Mo, Yukang Ding, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Feiyu Chen 0001, Shuyuan Zhu
ECCV (22)7
2024 HIT: Solving Partial Index Tracking via Hierarchical Reinforcement Learning
abstract
Partial index tracking (PIT) is a popular passive investment strategy aiming at replicating the performance of a market index (e.g., S&P 500). Existing PIT methods typically treat it as a regression problem and divide it into two tasks: (i) asset selection (determining which assets to choose from the index constituents) and (ii) asset allocation (deciding how to allocate capital among the selected assets). However, these methods either optimize these two tasks jointly, which has been proven to be NP-hard and inefficient when tracking large-scale constituent indices (e.g., Russell 2000), or attempt an independent optimization, lacking a connection to ensure collaborative optimization. In this paper, we present a hierarchical model for partial index tracking (HIT), which formulates PIT as a hierarchical Markov decision process (MDP) and is optimized via hierarchical reinforcement learning (HRL). HIT consists of (1) a high-level policy learns to select assets from constituents to handle task (i) and (2) a low-level policy learns to allocate capital weights among the selected assets to handle task (ii). We further propose a novel cost-sensitive reward function that serves as a connection to collaboratively optimize the two policies, aiming to replicate the index closely while considering transaction cost. Compared with existing jointly optimized approaches, our model simplifies the problem by learning separate policies for the two tasks, and the reward function serves as a connection to ensure collaborative optimization between them, avoiding challenges faced by joint optimization methods in existing literature. Remarkable performance across 6 benchmarks, ranging from small to large-scale constituents demonstrate the superiority of HIT. Moreover, the experiments conducted on a real-world market dataset spanning over 10 years show its effectiveness and practicality.
Zetao Zheng, Jie Shao 0001, Feiyu Chen 0001, Anjie Zhu, Shilong Deng, Heng Tao Shen
ICDE3
2024 Mutual Correlation Network for few-shot learning
Derong Chen, Feiyu Chen 0001, Deqiang Ouyang, Jie Shao 0001
Neural Networks2
2024 Modeling Hierarchical Uncertainty for Multimodal Emotion Recognition in Conversation
abstract
Approximating the uncertainty of an emotional AI agent is crucial for improving the reliability of such agents and facilitating human-in-the-loop solutions, especially in critical scenarios. However, none of the existing systems for emotion recognition in conversation (ERC) has attempted to estimate the uncertainty of their predictions. In this article, we present HU-Dialogue, which models hierarchical uncertainty for the ERC task. We perturb contextual attention weight values with source-adaptive noises within each modality, as a regularization scheme to model context-level uncertainty and adapt the Bayesian deep learning method to the capsule-based prediction layer to model modality-level uncertainty. Furthermore, a weight-sharing triplet structure with conditional layer normalization is introduced to detect both invariance and equivariance among modalities for ERC. We provide a detailed empirical analysis for extensive experiments, which shows that our model outperforms previous state-of-the-art methods on three popular multimodal ERC datasets.
Feiyu Chen 0001, Jie Shao 0001, Anjie Zhu, Deqiang Ouyang, Xueliang Liu, Heng Tao Shen
IEEE Trans. Cybern.1
2023 Multivariate, Multi-Frequency and Multimodal: Rethinking Graph Neural Networks for Emotion Recognition in Conversation
abstract
Complex relationships of high arity across modality and context dimensions is a critical challenge in the Emotion Recognition in Conversation (ERC) task. Yet, previous works tend to encode multimodal and contextual relationships in a loosely-coupled manner, which may harm relationship modelling. Recently, Graph Neural Networks (GNN) which show advantages in capturing data relations, offer a new solution for ERC. However, existing GNN-based ERC models fail to address some general limits of GNNs, including assuming pairwise formulation and erasing high-frequency signals, which may be trivial for many applications but crucial for the ERC task. In this paper, we propose a GNN-based model that explores multivariate relationships and captures the varying importance of emotion discrepancy and commonality by valuing multi-frequency signals. We empower GNNs to better capture the inherent relationships among utterances and deliver more sufficient multimodal and contextual modelling. Experimental results show that our proposed method outperforms previous state-of-the-art works on two popular multimodal ERC datasets.
Feiyu Chen 0001, Jie Shao 0001, Shuyuan Zhu, Heng Tao Shen
CVPR1
2023 Modeling Both Collaborative and Temporal Information for Sequential Recommendation
Jinyue Dai, Jie Shao 0001, Zhiyi Deng, Hongcai He, Feiyu Chen 0001
ICONIP (9)5
2023 Empowering the Diversity and Individuality of Option: Residual Soft Option Critic Framework
abstract
Extracting temporal abstraction (option), which empowers the action space, is a crucial challenge in hierarchical reinforcement learning. Under a well-structured action space, decision-making agents can probe more deeply in the searching or plan efficiently through pruning irrelevant action candidates. However, automatically capturing a well-performed temporal abstraction is a nontrivial challenge due to its insufficient exploration and inadequate functionality. We consider alleviating this challenge from two perspectives, i.e., diversity and individuality. For the aspect of diversity, we propose a maximum entropy model based on ensembled options to encourage exploration. For the aspect of individuality, we propose to distinguish each option accurately, utilizing mutual formation minimization, so that each option can better express and function. We name our framework as an ensemble with soft option (ESO) critics. Furthermore, the residual algorithm (RA) with a bidirectional target network is introduced to stabilize bootstrapping, yielding a residual version of ESO. We provide detailed analysis for extensive experiments, which shows that our method boosts performance in commonly used continuous control tasks.
Anjie Zhu, Feiyu Chen 0001, Deqiang Ouyang, Jie Shao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 TEVL: Trilinear Encoder for Video-language Representation Learning
abstract
Pre-training model on large-scale unlabeled web videos followed by task-specific fine-tuning is a canonical approach to learning video and language representations. However, the accompanying Automatic Speech Recognition (ASR) transcripts in these videos are directly transcribed from audio, which may be inconsistent with visual information and would impair the language modeling ability of the model. Meanwhile, previous V-L models fuse visual and language modality features using single- or dual-stream architectures, which are not suitable for the current situation. Besides, traditional V-L research focuses mainly on the interaction between vision and language modalities and leaves the modeling of relationships within modalities untouched. To address these issues and maintain a small manual labor cost, we add automatically extracted dense captions as a supplementary text and propose a new trilinear video-language interaction framework TEVL (Trilinear Encoder for Video-Language representation learning). TEVL contains three unimodal encoders, a TRIlinear encOder (TRIO) block, and a temporal Transformer. TRIO is specially designed to support effective text-vision-text interaction, which encourages inter-modal cooperation while maintaining intra-modal dependencies. We pre-train TEVL on the HowTo100M and TV datasets with four task objectives. Experimental results demonstrate that TEVL can learn powerful video-text representation and achieve competitive performance on three downstream tasks, including multimodal video captioning, video Question Answering (QA), as well as video and language inference. Implementation code is available at https://github.com/Gufrannn/TEVL .
Xin Man, Jie Shao 0001, Feiyu Chen 0001, Heng Tao Shen
ACM Trans. Multim. Comput. Commun. Appl.3
2022 SMDT: Cross-View Geo-Localization with Image Alignment and Transformer
abstract
The goal of cross-view geo-localization is to determine the location of a given ground image by matching with aerial images. However, existing methods ignore the variability of scenes, additional information and spatial correspondence of covisibility and non-convisibility areas in ground-aerial image pairs. In this context, we propose a cross-view matching method called SMDT with image alignment and Transformer. First, we utilize semantic segmentation technique to segment different areas. Then, we convert the vertical view of aerial images to front view by mixing polar mapping and perspective mapping. Next, we simultaneously train dual conditional generative adversarial nets by taking the semantic segmentation images and converted images as input to synthesize the aerial image with ground view style. These steps are collectively referred to as image alignment. Last, we use Transformer to explicitly utilize the properties of self-attention. Experiments show that our SMDT method is superior to the existing ground-to-aerial cross-view methods.
Xiaoyang Tian, Jie Shao 0001, Deqiang Ouyang, Anjie Zhu, Feiyu Chen 0001
ICME5
2022 Synesthesia Transformer with Contrastive Multimodal Learning
Zhengxiao Sun, Feiyu Chen 0001, Jie Shao 0001
ICONIP (1)2
2022 Breaking Isolation: Multimodal Graph Fusion for Multimedia Recommendation by Edge-wise Modulation
abstract
In a multimedia recommender system, rich multimodal dynamics of user-item interactions are worth availing ourselves of and have been facilitated by Graph Convolutional Networks (GCNs). Yet, the typical way of conducting multimodal fusion with GCN-based models is either through graph mergence fusion that delivers insufficient inter-modal dynamics, or through node alignment fusion that brings in noises which potentially harm multimodal modelling. Unlike existing works, we propose EgoGCN, a structure that seeks to enhance multimodal learning of user-item interactions. At its core is a simple yet effective fusion operation dubbed EdGe-wise mOdulation (EGO) fusion. EGO fusion adaptively distils edge-wise multimodal information and learns to modulate each unimodal node under the supervision of other modalities. It breaks isolated unimodal propagations, allows the most informative inter-modal messages to spread, whilst preserving intra-modal processing. We present a hard modulation and a soft modulation to fully investigate the multimodal dynamics behind. Experiments on two real-world datasets show that EgoGCN comfortably beats prior methods.
Feiyu Chen 0001, Junjie Wang 0011, Yinwei Wei, Hai-Tao Zheng 0002, Jie Shao 0001
ACM Multimedia1
2021 Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in Conversation
abstract
Multi-sensory data has exhibited a clear advantage in expressing richer and more complex feelings, on the Emotion Recognition in Conversation (ERC) task. Yet, current methods for multimodal dynamics that aggregate modalities or employ additional modality-specific and modality-shared networks are still inadequate in balancing between the sufficiency of multimodal processing and the scalability to incremental multi-sensory data type additions. This incurs a bottleneck of performance improvement of ERC. To this end, we present MetaDrop, a differentiable and end-to-end approach for the ERC task that learns module-wise decisions across modalities and conversation flows simultaneously, which supports adaptive information sharing pattern and dynamic fusion paths. Our framework mitigates the problem of modelling complex multimodal relations while ensuring it enjoys good scalability to the number of modalities. Experiments on two popular multimodal ERC datasets show that MetaDrop achieves new state-of-the-art results.
Feiyu Chen 0001, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, Jie Shao 0001
ACM Multimedia1
2021 Interclass-Relativity-Adaptive Metric Learning for Cross-Modal Matching and Beyond
abstract
Training under supervision of triplet ranking loss is a dominant methodology for cross-modal matching models, while good-performing losses in this domain are immensely under-explored since the majority of advanced metric losses are inapplicable due to the particularity of cross-modal setting. Current prominent approaches of metric learning have developed various weighting schemes that assign weights to separate positive or negative samples. It is the interclass relative order in a triplet, however, that matters. In this work, we propose a new Interclass-Relativity-Adaptive (IRA) loss that assigns weights to the relative similarities between positive and negative pairs instead of separate pairs, which allows us to regard a whole triplet as a weighable entity and achieve maximum utilization of sole positive under cross-modal setting. Our method outperforms the baselines by a large margin and obtains competitive results on two video-text matching benchmarks and two image-text matching benchmarks. We also further extend our method to two unimodal image retrieval benchmarks to test its generality and achieve new state-of-the-art results.
Feiyu Chen 0001, Jie Shao 0001, Xing Xu 0001, Heng Tao Shen
IEEE Trans. Multim.1
2020 Multiplicative angular margin loss for text-based person search
abstract
Text-based person search aims at retrieving the most relevant pedestrian images from database in response to a query in form of natural language description. Existing algorithms mainly focus on embedding textual and visual features into a common semantic space so that the similarity score of features from different modalities can be computed directly. Softmax loss is widely adopted to classify textual and visual features into a correct category in the joint embedding space. However, softmax loss can only help classify features but not increase the intra-class compactness and inter-class discrepancy. To this end, we propose multiplicative angular margin (MAM) loss to learn angularly discriminative features for each identity. The multiplicative angular margin loss penalizes the angle between feature vector and its corresponding classifier vector to learn more discriminative feature. Moreover, to focus more on informative image-text pair, we propose pairwise similarity weighting (PSW) loss to assign higher weight to informative pairs. Extensive experimental evaluations have been conducted on the CUHK-PEDES dataset over our proposed losses. The results show the superiority of our proposed method. Code is available at https://github.com/pengzhanguestc/MAM_loss.
Deqiang Ouyang, Feiyu Chen 0001, Jie Shao 0001
MMAsia3