Zheng Li 0014

dblp:10/1143-14 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0003-2535-2523ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Attribute-Aware Implicit Modality Alignment for text attribute person search
Fangfang Liu 0003, Xin Wang 0203, Zheng Li 0014, Caili Guo, Yang Yang 0057, Lin Hu 0008
Knowl. Based Syst.3
2025 Multi-view visual semantic embedding for cross-modal image-text retrieval
Zheng Li 0014, Caili Guo, Xin Wang 0203, Hao Zhang 0161, Lin Hu 0008
Pattern Recognit.1
2025 A unified framework of data augmentation using large language models for text-based cross-modal retrieval
Lijia Si, Caili Guo, Zheng Li 0014, Yang Yang 0057
Pattern Recognit.3
2025 Selectively Hard Negative Mining for Alleviating Gradient Vanishing in Image-Text Matching
abstract
Most Image-Text Matching (ITM) models adopt Triplet loss with Hard Negative mining (T-HN) as the optimization objective. T-HN mines the hardest negative samples in each batch for training and achieves impressive performance. However, we observe that these ITM models have bad training behaviors in the early phases of training. Model training is difficult to converge, and matching performance is slow to improve. In this paper, we find that the cause of bad training behavior is that the model suffers from gradient vanishing. Optimizing an ITM model using only the hardest negative samples can easily lead to gradient vanishing. Through gradient analysis, we first derive the condition under which the gradient vanishes during training. We explain why the gradient tends to zero under certain conditions. To alleviate gradient vanishing, we propose a Triplet loss with Selectively Hard Negative mining (T-SelHN), which decides whether to mine the hardest negative samples according to the gradient vanishing condition. T-SelHN can be applied to ITM models in a plug-and-play manner to improve their training behaviors. To further ensure the back-propagation of gradients, we construct a Residual Visual Semantic Embedding model with T-SelHN, denoted RVSE++, which has a simple network structure and efficient training and inference speeds. Extensive experiments on two ITM benchmarks demonstrate the strength of RVSE++, achieving state-of-the-art performance. The code is available athttps://github.com/AAA-Zheng/RVSEPP.
Zheng Li 0014, Caili Guo, Xin Wang 0203, Zerun Feng, Zhongtian Du
IEEE Trans. Circuits Syst. Video Technol.1
2024 Integrating listwise ranking into pairwise-based image-text retrieval
Zheng Li 0014, Caili Guo, Xin Wang 0203, Hao Zhang 0161
Knowl. Based Syst.1
2024 Learning From Noisy Correspondence With Tri-Partition for Cross-Modal Matching
abstract
Due to high labeling cost, it is inevitable to introduce a certain proportion of noisy correspondence into visual-text datasets, resulting in poor model robustness for cross-modal matching. Although recent methods divide the datasets into clean and noisy pair subsets to yield promising achievements, they still suffer from deep neural networks over-fitting on noisy correspondence. In particular, the similar positive pairs with partially relevant semantic correspondence are easily partitioned into noisy pair subset by mistake without carefully selection, which brings harmful impact on robust learning. Meanwhile, the similar negative pairs with partially relevant semantic correspondence lead to ambiguous distance relations in common space learning, which also damages the stability of performance. To solve the coarse-grained dataset division problem, we propose Correspondence Tri-Partition Rectifier (CTPR) to partition the training set into clean, hard, and noisy pair subsets based on the memorization effect of neural networks and prediction inconsistency. Then, we refine the correspondence labels for each subset to indicate the real semantic correspondence between visual-text pairs. The differences between rectified labels of anchors and hard negatives are recast as the adaptive margin in the improved triplet loss for robust training in a co-teaching manner. To verify the effectiveness and robustness of our method, we conduct experiments by implementing image-text and video-text matching as two showcases. Extensive experiments on Flickr30 K, MS-COCO, MSR-VTT, and LSMDC datasets verify that our method successfully partitions the visual-text pairs according to their semantic correspondence and improves performance under noisy data training.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Lin Hu 0008
IEEE Trans. Multim.4
2024 Integrating Language Guidance Into Image-Text Matching for Correcting False Negatives
abstract
Image-Text Matching (ITM) aims to establish the correspondence between images and sentences. ITM is fundamental to various vision and language understanding tasks. However, there are limitations in the way existing ITM benchmarks are constructed. The ITM benchmark collects pairs of images and sentences during construction. Therefore, only samples that are paired at collection are annotated as positive. All other samples are annotated as negative. Many correlations are missed in these samples that are annotated as negative. For example, a sentence matches only one image at the time of collection. Only this image is annotated as positive for the sentence. All other images are annotated as negative. However, these negative images may contain images that correspond to the sentences. These mislabeled samples are calledfalse negatives. Existing ITM models are optimized based on annotations containing mislabels, which can introduce noise during training. In this paper, we propose an ITM framework integrating Language Guidance (LG) for correcting false negatives. A language pre-training model is introduced into the ITM framework to identify false negatives. To correct false negatives, we propose language guidance loss, which adaptively corrects the locations of false negatives in the visual-semantic embedding space. Extensive experiments on two ITM benchmarks show that our method can improve the performance of existing ITM models. To verify the performance of correcting false negatives, we conduct further experiments on ECCV Caption. ECCV Caption is a verified dataset where false negatives in annotations have been corrected. The experimental results show that our method can recall more relevant false negatives. The code is available athttps://github.com/AAA-Zheng/LG_ITM.
Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Zhongtian Du
IEEE Trans. Multim.1
2023 Feature Decomposition and Attribute Augmentation for Attribute-based Person Search
abstract
The attribute-based person search task aims to find matching pedestrian images by text attributes, which is relevant in scenarios where no query image is given. However, the existing methods exhibit inferior performance due to their inadequate representation of local fine-grained features, which hinders their ability to effectively model intra-class variations. In addition, there is a zero-shot retrieval problem due to the large number of unseen categories in the test set, resulting in suboptimal generalization performance. In this paper, we propose a novel Feature Decomposition and Attribute Augmentation (FDAA) framework to solve the above problems. Firstly, by decomposing the global features of the image from different directions, the local features are extracted at multiple granularities, thus effectively improving the discrimination ability of the model. Secondly, an attribute augmentation strategy is proposed that can effectively expand the combination of attributes during training and improve the generalization ability of the model. Extensive experiments on the PETA, Market-1501 Attribute, and PA100K datasets demonstrate the effectiveness of our proposed method, outperforming state-of-the-art methods.
Xin Wang 0203, Fangfang Liu 0003, Caili Guo, Zheng Li 0014, Hao Zhang 0161
VCIP4
2023 Feature Refinement with Masked Cascaded Network for Temporal Action Localization
abstract
Despite the great progress in temporal action localization (TAL), most existing methods directly use video encoders trained on trimmed Kinetics400 dataset to obtain clip-level visual features, ignoring the cross-dataset bias between Kinetics400 and TAL benchmarks. Such a dataset bias leads to poor visual representation, potentially hindering performance in both temporal detection and action recognition for TAL. In this paper, we propose a novel TAL method, termed feature refinement with masked cascaded network (FR-MCN), to tackle the above problem. Specifically, FR-MCN presents a new feature refinement strategy by developing clip-level feature classification task for both action and background clips to improve temporal sensitivity and enhance action semantics of visual features. Moreover, FR-MCN employs a masked cascaded paradigm for refinement to learn semantic disparities between action and background clips near boundary, enabling the starting and ending instants to be detected accurately for TAL. Extensive experimental results on THUMOS14 and ActivityNetv1.3 demonstrate that our FR-MCN, can significantly improve the action localization performance.
Chunyang Feng, Hao Zhang 0161, Caili Guo, Zheng Li 0014
VCIP5
2023 Boundary-Aware Proposal Generation Method for Temporal Action Localization
abstract
The goal of Temporal Action Localization (TAL) is to find the categories and temporal boundaries of actions in an untrimmed video. Most TAL methods rely heavily on action recognition models that are sensitive to action labels rather than temporal boundaries. More importantly, few works consider the background frames that are similar to action frames in pixels but dissimilar in semantics, which also leads to inaccurate temporal boundaries. To address the challenge above, we propose a Boundary-Aware Proposal Generation (BAPG) method with contrastive learning. Specifically, we define the above background frames as hard negative samples. Contrastive learning with hard negative mining is introduced to improve the discrimination of BAPG. BAPG is independent of the existing TAL network architecture, so it can be applied plug-and-play to mainstream TAL models without training. Extensive experimental results on THUMOS14 and ActivityNet-1.3 demonstrate that BAPG can significantly improve the performance of TAL.
Hao Zhang 0161, Chunyan Feng, Caili Guo, Zheng Li 0014, Xin Wang 0203
VCIP5
2023 Temporal Multimodal Graph Transformer With Global-Local Alignment for Video-Text Retrieval
abstract
Video-text retrieval is a crucial task that has been a powerful application for multi-media data analysis and attracted tremendous interest in the research area. The core steps are feature representations and alignment to overcome the heterogeneous gap between videos and texts. Existing methods not only take advantage of multi-modal information in videos but also explore local alignment to enhance retrieval accuracy. Although performing well, these methods seem deficient at three perspectives: a) The semantic correlations between different modal features are not considered, which introduces irrelevant noise in feature representations. b) The cross-modal relations and temporal associations are ambiguously learned by a single self-attention manipulation. c) The training signal to optimize the semantic topic assignment for local alignment is missing. In this paper, we proposed a novel Temporal Multi-modal Graph Transformer with Global-Local Alignment (TMMGT-GLA) for video-text retrieval. We model the input video as a sequence of semantic correlation graphs to exploit the structural information between multi-modal features. Graph and temporal self-attention layers are leveraged on the semantic correlation graphs to effectively learn cross-modal relations and temporal associations respectively. For local alignment, the encoded video and text features are assigned to a set of shared semantic topics, and the distances between residuals from the same ones are minimized. To optimize the assignments, a minimum entropy-based regularization term is proposed for training the overall framework. Experimental results are carried out on the MSR-VTT, LSMDC, and ActivityNet Captions datasets. Our method outperforms previous approaches by a large margin and achieves state-of-the-art performance.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014
IEEE Trans. Circuits Syst. Video Technol.4
2022 Spatial-Temporal Alignment via Optimal Transport for Video-Text Retrieval
abstract
Video-text retrieval is one of the most popular branches in the cross-modal research area facing the exponential growth of multimedia services. Present methods typically mea-sure cross-modal similarities only by the cosine metric in a common space. However, this function ignores the in-herent discrepancy between videos and texts for reflecting spatial-temporal contents and fails to align them under the weakly supervised setting. To address this issue, we propose a novel Spatial-Temporal Optimal Transport (STOT) frame-work building upon advances in optimal transport. The spatial and temporal characteristics of videos and texts are carefully considered in STOT and associated by calculating a spatial-temporal alignment distance, which can be implemented as a regularizer for existing models to improve retrieval accu-racy. Extensive experiments are conducted on two video-text datasets with two models as base architectures to demonstrate the outperformance of our framework due to the powerful spatial-temporal alignment capability.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014, Yufeng Zhang 0008
ICME4
2022 Multi-View Visual Semantic Embedding
abstract
Visual Semantic Embedding (VSE) is a dominant method for cross-modal vision-language retrieval. Its purpose is to learn an embedding space so that visual data can be embedded in a position close to the corresponding text description. However, there are large intra-class variations in the vision-language data. For example, multiple texts describing the same image may be described from different views, and the descriptions of different views are often dissimilar. The mainstream VSE method embeds samples from the same class in similar positions, which will suppress intra-class variations and lead to inferior generalization performance. This paper proposes a Multi-View Visual Semantic Embedding (MV-VSE) framework, which learns multiple embeddings for one visual data and explicitly models intra-class variations. To optimize MV-VSE, a multi-view upper bound loss is proposed, and the multi-view embeddings are jointly optimized while retaining intra-class variations. MV-VSE is plug-and-play and can be applied to various VSE models and loss functions without excessively increasing model complexity. Experimental results on the Flickr30K and MS-COCO datasets demonstrate the superior performance of our framework.
Zheng Li 0014, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Xijun Xue
IJCAI1
2020 A Novel Convolutional Architecture For Video-Text Retrieval
abstract
The prevalent video-text retrieval methods usually use recurrent neural networks to encode sequences of frames in videos and sequences of words in text. In this paper, we introduce an encoding architecture based entirely on convolutional neural networks. Compared to recurrent models, the complexity is smaller, and computations over all elements can be fully parallelized during training to better exploit the GPU. We use the stacking of convolution kernels of different scales to realize the encoding of local and long-term features of video and text. Experiments validate that our method achieves a new state-of-the-art for the video-text retrieval on MSR-VTT and MSVD datasets with less training time.
Zheng Li 0014, Caili Guo, Zerun Feng, Hao Zhang 0161
ICME1
2020 Exploiting Visual Semantic Reasoning for Video-Text Retrieval
abstract
Video retrieval is a challenging research topic bridging the vision and language areas and has attracted broad attention in recent years. Previous works have been devoted to representing videos by directly encoding from frame-level features. In fact, videos consist of various and abundant semantic relations to which existing methods pay less attention. To address this issue, we propose a Visual Semantic Enhanced Reasoning Network (ViSERN) to exploit reasoning between frame regions. Specifically, we consider frame regions as vertices and construct a fully-connected semantic correlation graph. Then, we perform reasoning by novel random walk rule-based graph convolutional networks to generate region features involved with semantic relations. With the benefit of reasoning, semantic interactions between regions are considered, while the impact of redundancy is suppressed. Finally, the region features are aggregated to form frame-level features for further encoding to measure video-text similarity. Extensive experiments on two public benchmark datasets validate the effectiveness of our method by achieving state-of-the-art performance due to the powerful semantic reasoning.
Zerun Feng, Zhimin Zeng, Caili Guo, Zheng Li 0014
IJCAI4
2019 A Multi-Agent Deep Reinforcement Learning Based Spectrum Allocation Framework for D2D Communications
abstract
Device-to-device (D2D) communication has been recognized as a promising technique to improve spectrum efficiency. However, D2D transmission as an underlay causes severe interference, which imposes a technical challenge to spectrum allocation. Existing centralized schemes require global information, which can cause serious signaling overhead. While existing distributed solution requires frequent information exchange between users and cannot achieve global optimization. In this paper, a distributed spectrum allocation framework based on multi-agent deep reinforcement learning is proposed, named Neighbor-Agent Actor Critic (NAAC). NAAC uses neighbor users' historical information for centralized training but is executed distributedly without that information, which not only has no signal interaction during execution, but also utilizes cooperation between users to further optimize system performance. The simulation results show that the proposed framework can effectively reduce the outage probability of cellular links, improve the sum rate of D2D links and have good convergence.
Zheng Li 0014, Caili Guo, Yidi Xuan
GLOBECOM1