Zerun Feng

dblp:267/5582 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-3987-0591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 FASTER: Face Attribute Sliders with Semantic Rewards
abstract
Large-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing multi-step gradients, leading to considerable time and memory consumption. In this study, we propose FASTER(Face Attribute Sliders wiTh sEmantic Rewards), an effective method that employs stable diffusion models for face attribute editing. The key idea is to identify a low-rank attribute editing direction by leveraging attribute reward and S-CLIP reward between the original face and the edited face. This process helps to establish the desired face attribute slider. To acquire the edited face, we introduce an efficient one-step reward technique by utilizing denoised results at random timesteps for learning. This technique reduces training time by 6x. FASTER achieves 98.67% editing accuracy while simultaneously improving attribute preservation by nearly 10% compared to other methods on the CelebA-HQ dataset, all without compromising identity information.
Jingyan Chen, Lanxiang Zhou, Han Fang 0002, Zerun Feng, Chao Ban, Hao Sun 0038, Jiani Hu
ICASSP4
2025 Selectively Hard Negative Mining for Alleviating Gradient Vanishing in Image-Text Matching
abstract
Most Image-Text Matching (ITM) models adopt Triplet loss with Hard Negative mining (T-HN) as the optimization objective. T-HN mines the hardest negative samples in each batch for training and achieves impressive performance. However, we observe that these ITM models have bad training behaviors in the early phases of training. Model training is difficult to converge, and matching performance is slow to improve. In this paper, we find that the cause of bad training behavior is that the model suffers from gradient vanishing. Optimizing an ITM model using only the hardest negative samples can easily lead to gradient vanishing. Through gradient analysis, we first derive the condition under which the gradient vanishes during training. We explain why the gradient tends to zero under certain conditions. To alleviate gradient vanishing, we propose a Triplet loss with Selectively Hard Negative mining (T-SelHN), which decides whether to mine the hardest negative samples according to the gradient vanishing condition. T-SelHN can be applied to ITM models in a plug-and-play manner to improve their training behaviors. To further ensure the back-propagation of gradients, we construct a Residual Visual Semantic Embedding model with T-SelHN, denoted RVSE++, which has a simple network structure and efficient training and inference speeds. Extensive experiments on two ITM benchmarks demonstrate the strength of RVSE++, achieving state-of-the-art performance. The code is available athttps://github.com/AAA-Zheng/RVSEPP.
Zheng Li 0014, Caili Guo, Xin Wang 0203, Zerun Feng, Zhongtian Du
IEEE Trans. Circuits Syst. Video Technol.4
2024 ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
abstract
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).
Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038
ICME4
2024 GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment Tuning
abstract
Recent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality.
Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038
ACM Multimedia3
2024 Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching
abstract
Due to high labeling cost, it is inevitable to introduce a certain proportion of noisy correspondence into visual-text datasets, resulting in poor model robustness for cross-modal matching. Although recent methods divide the datasets into clean and noisy pair subsets to yield promising achievements, they still suffer from deep neural networks over-fitting on noisy correspondence. In particular, the similar positive pairs with partially relevant semantic correspondence are easily partitioned into noisy pair subset by mistake without carefully selection, which brings harmful impact on robust learning. Meanwhile, the similar negative pairs with partially relevant semantic correspondence lead to ambiguous distance relations in common space learning, which also damages the stability of performance. To solve the coarse-grained dataset division problem, we propose Correspondence Tri-Partition Rectifier (CTPR) to partition the training set into clean, hard, and noisy pair subsets based on the memorization effect of neural networks and prediction inconsistency. Then, we refine the correspondence labels for each subset to indicate the real semantic correspondence between visual-text pairs. The differences between rectified labels of anchors and hard negatives are recast as the adaptive margin in the improved triplet loss for robust training in a co-teaching manner. To verify the effectiveness and robustness of our method, we conduct experiments by implementing image-text and video-text matching as two showcases. Extensive experiments on Flickr30 K, MS-COCO, MSR-VTT, and LSMDC datasets verify that our method successfully partitions the visual-text pairs according to their semantic correspondence and improves performance under noisy data training.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Lin Hu 0008
IEEE Trans. Multim.1
2024 Integrating Language Guidance Into Image-Text Matching for Correcting False Negatives
abstract
Image-Text Matching (ITM) aims to establish the correspondence between images and sentences. ITM is fundamental to various vision and language understanding tasks. However, there are limitations in the way existing ITM benchmarks are constructed. The ITM benchmark collects pairs of images and sentences during construction. Therefore, only samples that are paired at collection are annotated as positive. All other samples are annotated as negative. Many correlations are missed in these samples that are annotated as negative. For example, a sentence matches only one image at the time of collection. Only this image is annotated as positive for the sentence. All other images are annotated as negative. However, these negative images may contain images that correspond to the sentences. These mislabeled samples are calledfalse negatives. Existing ITM models are optimized based on annotations containing mislabels, which can introduce noise during training. In this paper, we propose an ITM framework integrating Language Guidance (LG) for correcting false negatives. A language pre-training model is introduced into the ITM framework to identify false negatives. To correct false negatives, we propose language guidance loss, which adaptively corrects the locations of false negatives in the visual-semantic embedding space. Extensive experiments on two ITM benchmarks show that our method can improve the performance of existing ITM models. To verify the performance of correcting false negatives, we conduct further experiments on ECCV Caption. ECCV Caption is a verified dataset where false negatives in annotations have been corrected. The experimental results show that our method can recall more relevant false negatives. The code is available athttps://github.com/AAA-Zheng/LG_ITM.
Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Zhongtian Du
IEEE Trans. Multim.3
2023 Temporal Multimodal Graph Transformer With Global-Local Alignment for Video-Text Retrieval
abstract
Video-text retrieval is a crucial task that has been a powerful application for multi-media data analysis and attracted tremendous interest in the research area. The core steps are feature representations and alignment to overcome the heterogeneous gap between videos and texts. Existing methods not only take advantage of multi-modal information in videos but also explore local alignment to enhance retrieval accuracy. Although performing well, these methods seem deficient at three perspectives: a) The semantic correlations between different modal features are not considered, which introduces irrelevant noise in feature representations. b) The cross-modal relations and temporal associations are ambiguously learned by a single self-attention manipulation. c) The training signal to optimize the semantic topic assignment for local alignment is missing. In this paper, we proposed a novel Temporal Multi-modal Graph Transformer with Global-Local Alignment (TMMGT-GLA) for video-text retrieval. We model the input video as a sequence of semantic correlation graphs to exploit the structural information between multi-modal features. Graph and temporal self-attention layers are leveraged on the semantic correlation graphs to effectively learn cross-modal relations and temporal associations respectively. For local alignment, the encoded video and text features are assigned to a set of shared semantic topics, and the distances between residuals from the same ones are minimized. To optimize the assignments, a minimum entropy-based regularization term is proposed for training the overall framework. Experimental results are carried out on the MSR-VTT, LSMDC, and ActivityNet Captions datasets. Our method outperforms previous approaches by a large margin and achieves state-of-the-art performance.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014
IEEE Trans. Circuits Syst. Video Technol.1
2022 Spatial-Temporal Alignment via Optimal Transport for Video-Text Retrieval
abstract
Video-text retrieval is one of the most popular branches in the cross-modal research area facing the exponential growth of multimedia services. Present methods typically mea-sure cross-modal similarities only by the cosine metric in a common space. However, this function ignores the in-herent discrepancy between videos and texts for reflecting spatial-temporal contents and fails to align them under the weakly supervised setting. To address this issue, we propose a novel Spatial-Temporal Optimal Transport (STOT) frame-work building upon advances in optimal transport. The spatial and temporal characteristics of videos and texts are carefully considered in STOT and associated by calculating a spatial-temporal alignment distance, which can be implemented as a regularizer for existing models to improve retrieval accu-racy. Extensive experiments are conducted on two video-text datasets with two models as base architectures to demonstrate the outperformance of our framework due to the powerful spatial-temporal alignment capability.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Yufeng Zhang 0008
ICME1
2022 Multi-View Visual Semantic Embedding
abstract
Visual Semantic Embedding (VSE) is a dominant method for cross-modal vision-language retrieval. Its purpose is to learn an embedding space so that visual data can be embedded in a position close to the corresponding text description. However, there are large intra-class variations in the vision-language data. For example, multiple texts describing the same image may be described from different views, and the descriptions of different views are often dissimilar. The mainstream VSE method embeds samples from the same class in similar positions, which will suppress intra-class variations and lead to inferior generalization performance. This paper proposes a Multi-View Visual Semantic Embedding (MV-VSE) framework, which learns multiple embeddings for one visual data and explicitly models intra-class variations. To optimize MV-VSE, a multi-view upper bound loss is proposed, and the multi-view embeddings are jointly optimized while retaining intra-class variations. MV-VSE is plug-and-play and can be applied to various VSE models and loss functions without excessively increasing model complexity. Experimental results on the Flickr30K and MS-COCO datasets demonstrate the superior performance of our framework.
Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Xijun Xue
IJCAI3
2020 A Novel Convolutional Architecture For Video-Text Retrieval
abstract
The prevalent video-text retrieval methods usually use recurrent neural networks to encode sequences of frames in videos and sequences of words in text. In this paper, we introduce an encoding architecture based entirely on convolutional neural networks. Compared to recurrent models, the complexity is smaller, and computations over all elements can be fully parallelized during training to better exploit the GPU. We use the stacking of convolution kernels of different scales to realize the encoding of local and long-term features of video and text. Experiments validate that our method achieves a new state-of-the-art for the video-text retrieval on MSR-VTT and MSVD datasets with less training time.
Zheng Li 0014, Caili Guo, Zerun Feng, Hao Zhang 0161
ICME4
2020 Exploiting Visual Semantic Reasoning for Video-Text Retrieval
abstract
Video retrieval is a challenging research topic bridging the vision and language areas and has attracted broad attention in recent years. Previous works have been devoted to representing videos by directly encoding from frame-level features. In fact, videos consist of various and abundant semantic relations to which existing methods pay less attention. To address this issue, we propose a Visual Semantic Enhanced Reasoning Network (ViSERN) to exploit reasoning between frame regions. Specifically, we consider frame regions as vertices and construct a fully-connected semantic correlation graph. Then, we perform reasoning by novel random walk rule-based graph convolutional networks to generate region features involved with semantic relations. With the benefit of reasoning, semantic interactions between regions are considered, while the impact of redundancy is suppressed. Finally, the region features are aggregated to form frame-level features for further encoding to measure video-text similarity. Extensive experiments on two public benchmark datasets validate the effectiveness of our method by achieving state-of-the-art performance due to the powerful semantic reasoning.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014
IJCAI1