Ying Li 0016

dblp:22/1805-16 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-5695-4706ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Unsupervised pattern image retrieval via dual-encoder architecture with multi-head attention
Ru Han, Chunming Guan, Jiaquan Gao, Ying Li 0016
Neurocomputing5
2026 Hierarchical kernel decoupling for graph convolution: Enhancing skeleton-based action recognition through structured representation
Ying Li 0016, Hao Zhou 0014, Chuanping Hu, Mingzhou Lu, Yan Luo 0003
Pattern Recognit.2
2026 RA-GCN: Residual attention based graph convolutional network for multi-label pattern image retrieval
Ying Li 0016, Longye Du, Chunming Guan, Jiaquan Gao
Pattern Recognit.1
2026 Correspondence Calibrating and Dynamic Consistency Learning for Noisy Cross-Modal Retrieval
abstract
Cross-modal retrieval has drawn an increasing amount of attention due to its effective ability for searching semantic relative data points with different modalities. In spite of some progress obtained, such methods often require the data pair maintaining the correct cross-modal correspondence in the training process, which is impractical in real application. To tackle this issue, we propose a Correspondence Calibrating and Dynamic Consistency Learning Network (CCDCL), aiming at optimizing the correspondence of positive samples and deeply investigating the consistency of negative samples. Specifically, to effectively alleviate the false positive issues, co-teaching paradigm is introduced to optimize the correspondence of positive samples by calibrating their confidence scores. To address the false negative sample problem, we propose a Vision-Language Semantic Collaborative Dynamic Margin Adaptation (VLC-DMA) method, which integrates unimodal similarities with calibrated confidence to produce the cross-modal semantic similarities, and consequently employs a dynamic margin function for accurately discriminating false and true negative samples. Experimental results demonstrate that the proposed method effectively improves cross-modal retrieval performance across multiple image-text matching datasets. The code will be available on GitHub upon acceptance of this paper.
Yizhen Wu, Linliang Zhang, Guorui Sheng, Ying Li 0016, Qi Tian 0001
IEEE Trans. Multim.5
2025 VG-Net: Vision Transformer based Graph Fusion Representation for Multi-label Pattern Image Retrieval
abstract
Retrieving pattern images with multiple labels is complex, as it requires accurately finding similar images from a large database based on a query. Recent advancements using attention mechanisms and semantic alignment techniques have made progress, but challenges like label insufficiency, imbalance, and noise still hinder multi-label retrieval performance. To address these issues, we introduce the Vision Transformer based Graph Fusion Network (VG-Net), which includes a representation learning module and a semantic fusion module. The representation learning module utilizes a Transformer encoder with patch embedding to capture multi-label semantics. Simultaneously, the semantic fusion module leverages graph neural networks and cross-modal attention to explore interactions between image representations and label dependencies. Extensive experiments on three popular multi-label image retrieval benchmarks and a pattern image dataset demonstrate the exceptional performance of VG-Net.
Erwan Ye, Ying Li 0016
ICME2
2025 Self-supervised incomplete cross-modal hashing retrieval
Shouyong Peng, Ying Li 0016, Gang Wang 0029, Zhiming Yan
Expert Syst. Appl.3
2025 HTS-LB: Hypergraph tree search for learning branch
Yige Zhang, Ying Li 0016, Jiaquan Gao
Neural Networks4
2025 Complementary two-branch Transformer for multi-label image retrieval
Ying Li 0016, Shuaiyu Deng, Chunming Guan, Jiaquan Gao
Pattern Recognit.1
2025 A Survey on Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) processes a query consisting of a reference image and a modification text, aiming to retrieve target images that not only resemble the reference image visually but also reflect the modification described in the caption. Unlike traditional image retrieval methods that rely on a single modality, CIR integrates visual and textual information, enabling more nuanced and constraint-based query representations. This unique capability has garnered growing interest from researchers. Despite its potential, the field lacks a systematic review that comprehensively examines its advancements and trends. This article seeks to fill this gap by providing a detailed review of CIR research developments over the past 5 years. It categorizes existing methods into supervised approaches, which leverage triplet-labeled data for model training, and zero-shot approaches, which utilize unlabeled data to address CIR challenges. Additionally, the widely used benchmark datasets and evaluation indicators are comprehensively introduced. A comparative analysis of state-of-the-art methods across five datasets is also conducted, providing insights into their strengths and limitations. Ultimately, this article also outlines potential research directions for the future.
Longye Du, Shuaiyu Deng, Ying Li 0016, Jun Li 0033, Qi Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MoHGCN: Momentum Hypergraph Convolution Network for Cross-modal Retrieval
abstract
Cross-modal retrieval tasks, encompassing the retrieval of image–text, video–audio, and more, are progressively gaining significance in response to the exponential growth of information on the Internet. However, there has always been a cloud hanging over multimodal tasks due to the inherent challenges in aligning different modalities with distinct physical meanings. Most previous works simply rely on a single multimodal encoder or a novel similarity calculation for fusion, which often result in unsatisfactory performance. To tackle this challenge, we introduce a Momentum Hypergraph Convolutional Network (MoHGCN) for multimodal representation learning, which strengthens the alignment of both visual and textual data before the retrieval process. Specifically, MoHGCN utilizes contrastive learning to select the most challenging negative and positive samples to form hyperedges and completes the modality alignment through two rounds of fusion. Subsequently, the fully integrated node features and global features are fused using a fusion encoder to obtain the final multimodal representation vector for image–text retrieval. Extensive experiments are conducted on two widely used datasets, namely Flickr30K and MSCOCO, to demonstrate the superiority of the proposed MoHGCN approach in achieving the state-of-the-art performances.
Ying Li 0016, Ding Yuxiang
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Cross-modality interaction reasoning for enhancing vision-language pre-training in image-text retrieval
Shouyong Peng, Ying Li 0016, Yujuan Sun
Appl. Intell.4
2024 Efficient Supervised Graph Embedding Hashing for large-scale cross-media retrieval
Ying Li 0016, Lianshan Yan, Qi Tian 0001
Pattern Recognit.4
2024 Cross-modal Semantically Augmented Network for Image-text Matching
abstract
Image-text matching plays an important role in solving the problem of cross-modal information processing. Since there are nonnegligible semantic differences between heterogenous pairwise data, a crucial challenge is how to learn a unified representation. Existing methods mainly rely on the alignment between regional image features and corresponding entity words. However, the regional features in the image are often more concerned with the foreground entity information, and the attribute information of the entities and the relational information are ignored. How to effectively integrate entity-attribute alignment and relationship alignment has not been fully studied. Therefore, we propose a Cross-Modal Semantically Augmented Network for Image-Text Matching (CMSAN), which combines the relationships between entities in the image with the semantics of relational words in the text. CMSAN (1) proposes an adaptive word-type prediction model that classifies the words into four types, i.e., entity word, attribute word, relation word, and unnecessary word. It can align different image features at multiple levels. CMSAN (2) designs a sophisticated relationship alignment module and an entity-attribute alignment module that maximizes the exploitation of the semantic information, which enables the model to have more discriminative power and further improves the matching accuracy.
Yiru Li, Ying Li 0016, Yingying Zhu 0001, Gang Wang 0029, Jun Yue 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 TsP-Tran: Two-Stage Pure Transformer for Multi-Label Image Retrieval
abstract
Image retrieval aims to find similar images given the query. Most of existing retrieval works are based on the pre-trained model of single-label image classification. In practice, the query usually contains more than one instance, and the single label is far from enough for fully depicting the attributes of an open-world image. Due to the complicated similarity relationships between multiple semantics, the multi-label image retrieval task is not so well solved as the single-label task. In this work, we propose a two-stage pure Transformer model for multi-label image retrieval, which leverages a Transformer encoder to exploit the complex dependencies among visual features and labels. Except for the Transformer encoder, the image feature embedding module is also based on Transformer, so that the optimal model weights could be learned in an end-to-end manner. To be specific, inputs of the Transformer encoder mainly consist of an Vision Transformer branch and a label embedding branch, which generates suitable image features and label descriptions, respectively. Given an input set of visual features and text labels, the developed Transformer encoder could be accordingly optimized in the training stage with compressed multi-label output layer. In order to obtain sufficient outputs to accurately find images containing similar semantics with the query from the database, we adjust the network by removing the last fully connected layer in the retrieval stage. Specially, images and labels are used for the training stage in a randomly masked manner to enhance the model performance, and no labels are visible in the content-based image retrieval stage. Comprehensive experiments are performed on three multi-label datasets including MS-COCO, NUS-WIDE and VOC2007, demonstrating promising results of our proposed method against the state-of-the-arts for multi-label image retrieval.
Ying Li 0016, Chunming Guan, Jiaquan Gao
ICMR1
2023 Tran-GCN: Multi-label Pattern Image Retrieval via Transformer Driven Graph Convolutional Network
abstract
Pattern images are artificially designed images that possess distinctiveness in their elements, styles, and arrangements. With the ever-growing number of pattern images, pattern image retrieval emerges as a promising technique with significant potential for commercial and industrial applications, such as fashion and home decoration, facilitating rapid identification of preferred print patterns by users. The main purpose of multi-label pattern image retrieval is to effectively represent and match images with their corresponding labels. Compared to conventional image retrieval, multi-label pattern image retrieval faces greater challenges due to the richer semantic information contained within the abstract print patterns and the complex relationships between multiple labels. To tackle these challenges, we propose a model specifically designed for multi-label pattern image retrieval, called Tran-GCN. Our proposed model is built upon a Transformer-based autoregressive architecture, which leverages image information to guide the exploration of correlations between different labels through the textual modality. By utilizing this correlation information, we construct a graph convolutional network (GCN) model to further enhance the correlations between image and label representations. To be more specific, our Tran-GCN model utilizes a cross-modal attention mechanisms at each layer to effectively aggregate visual features from the input image and update label semantics through residual connections. The GCN module is updated based on the correlation between textual features, as represented in a relationship matrix. Extensive experiments on two widely used public visual benchmarks, MS-COCO and NUS-WIDE, as well as a multi-label pattern image dataset, Pattern 2, consistently demonstrate the ability of our proposed Tran-GCN model for general use and its superior performance in multi-label pattern image retrieval tasks as well.
Ying Li 0016, Chunming Guan, Erwan Ye, Ding Yuxiang, Jiaquan Gao
ACM Multimedia1
2023 Discrete Robust Matrix Factorization Hashing for Large-Scale Cross-Media Retrieval
abstract
Cross-media hashing, which encodes data points from different modalities into a common Hamming space, has been successfully applied to solve large-scale multimedia retrieval issue due to storage efficiency and search effectiveness. Recently, matrix factorization based hashing methods have drawn considerable attention for their promising search accuracy. However, pioneer methods mainly focus on learning consensus hash codes for different modalities, but neglect the potential inconsistency among different modalities, \emph{e.g.,} the diversities of different modalities and noises, which may undermine the retrieval accuracy. To address this problem, we propose a novel unsupervised hashing model, namely, Discrete Robust Matrix Factorization Hashing (DRMFH), which simultaneously formulates the consistency and inconsistency across different modalities into a matrix factorization based model. Specifically, a homogenous space composed of a consistent Hamming space and an inconsistent diversity part, are generated by matrix factorization for each modality. Therefore, the consensus information across different modalities can be well captured in the learnt hash codes, leading to improved retrieval performance. Moreover, we design an effective optimization algorithm which is able to obtain an approximate discrete code matrix with linear time complexity. Comprehensive experimental results on three public multimedia retrieval datasets show that the proposed DRMFH outperforms several state-of-the-art methods.
Yiru Li, Weili Guan, Gang Wang 0029, Ying Li 0016, Lianshan Yan, Qi Tian 0001
IEEE Trans. Knowl. Data Eng.5
2022 Informed Patch Enhanced HyperGCN for skeleton-based action recognition
Ying Li 0016, Hao Zhou 0014, Yan Luo 0003, Chuanping Hu
Inf. Process. Manag.2
2022 Decoupled Pose and Similarity Based Graph Neural Network for Video Person Re-Identification
abstract
Significant development of video person re-identification has been witnessed in recent years with deep learning technologies. Due to the complexity of human pose changes and the similarity between different individuals, learning discriminative features is still a challenging part of the video person re-identification task. To get rid of the effects of pose misalignment while keep the similarity of human appearance, in this paper, we propose a Pose and Similarity based Graph Neural Network in a decoupled manner, which consists of three independent branches to emphasize the respective roles of pose, local similarity and global similarity in the final descriptions. Compared to traditional Convolutional Neural Networks which tend to output similar global features in the case of highly similar pedestrians, the developed Graph Neural Networks are able to explore local semantic relationships between body parts, resulting in more discriminative features. To further eliminate the pose variation, we incorporate human skeleton information for feature map segmentation. Specifically, we propose to take a tree structure as the pose-aware adjacency graph of blocks in a person frame, which reveals the inherent connections within a human body. Experimental results on four widely used datasets demonstrate the effectiveness of our method.
Ying Li 0016, Hengheng Zhang, Mengjing Li, Genlin Ji
IEEE Signal Process. Lett.1
2021 Multi-label Pattern Image Retrieval via Attention Mechanism Driven Graph Convolutional Network
abstract
Pattern images are artificially designed images which are discriminative in aspects of elements, styles, arrangements and so on. Pattern images are widely used in fields like textile, clothing, art, fashion and graphic design. With the growth of image numbers, pattern image retrieval has great potential in commercial applications and industrial production. However, most of existing content-based image retrieval works mainly focus on describing simple attributes with clear conceptual boundaries, which are not suitable for pattern image retrieval. It is difficult to accurately represent and retrieve pattern images which include complex details and multiple elements. Therefore, in this paper, we collect a new pattern image dataset with multiple labels per image for the pattern image retrieval task. To extract discriminative semantic features of multi-label pattern images and construct high-level topology relationships between features, we further propose an Attention Mechanism Driven Graph Convolutional Network (AMD-GCN). Different layers of the multi-semantic attention module activate regions of interest corresponding to multiple labels, respectively. By embedding the learned labels from attention module into the graph convolutional network, which can capture the dependency of labels on the graph manifold, the AMD-GCN builds an end-to-end framework to extract high-level semantic features with label semantics and inner relationships for retrieval. Experiments on the pattern image dataset show that the proposed method highlights the relevant semantic regions of multiple labels, and achieves higher accuracy than state-of-the-art image retrieval methods.
Ying Li 0016, Yeyu Yin, Jiaquan Gao
ACM Multimedia1
2021 3D-GAT: 3D-Guided adversarial transform network for person re-identification in unseen domains
Hengheng Zhang, Ying Li 0016, Zijie Zhuang, Lingxi Xie, Qi Tian 0001
Pattern Recognit.2
2020 Rank-embedded Hashing for Large-scale Image Retrieval
abstract
With the growth of images on the Internet, plenty of hashing methods are developed to handle the large-scale image retrieval task. Hashing methods map data from high dimension to compact codes, so that they can effectively cope with complicated image features. However, the quantization process of hashing results in unescapable information loss. As a consequence, it is a challenge to measure the similarity between images with generated binary codes. The latest works usually focus on learning deep features and hashing functions simultaneously to preserve the similarity between images, while the similarity metric is fixed. In this paper, we propose a Rank-embedded Hashing (ReHash) algorithm where the ranking list is automatically optimized together with the feedback of the supervised hashing. Specifically, ReHash jointly trains the metric learning and the hashing codes in an end-to-end model. In this way, the similarity between images are enhanced by the ranking process. Meanwhile, the ranking results are an additional supervision for the hashing function learning as well. Extensive experiments show that our ReHash outperforms the state-of-the-art hashing methods for large-scale image retrieval.
Haiyan Fu, Ying Li 0016, Hengheng Zhang
ICMR2
2020 Node-Sensitive Graph Fusion via Topo-Correlation for Image Retrieval
abstract
Various kinds of features prove to be effective for content-based image retrieval. However, due to the diversity of image contents, a descriptor may achieve impressive performance on specific images while becoming invalid on others. Although some efforts have been made to combine features as complementary counterparts, proper weighting scheme is still a challenge for fast and accurate retrieval. In this paper, we propose an effective fusion method, termed as Topo-correlation (Topo), where the importance of each feature is measured by cross-view correlations on local affinity graphs. Specifically, the weights of similarities are node-sensitive as well as modality-sensitive, thus boosting the results of good cues while depressing adverse factors for individual images. By estimating the consensus of similarity scores with regard to a query-driven criterion, the weighted graphs are generated efficiently with low computational complexity. Extensive experimental results on four benchmarks demonstrate the superiority of the proposed approach over the state-of-the-art methods.
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Contextual modeling on auxiliary points for robust image reranking
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Frontiers Comput. Sci.1
2019 Exploring geometric information in CNN for image retrieval
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu
Multim. Tools Appl.1
2018 Aggregating hierarchical binary activations for image retrieval
Ying Li 0016, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001
Neurocomputing1
2017 Local residual similarity for image re-ranking
Shaoyan Sun, Ying Li 0016, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li
Inf. Sci.2
2016 Discrete Cross-Modal Hashing for Efficient Multimedia Retrieval
abstract
Hashing techniques have been widely adopted for cross-modal retrieval due to its low storage cost and fast query speed. Most existing cross-modal hashing methods aim to map heterogeneous data into the common low-dimensional hamming space and then threshold to obtain binary codes by relaxing the discrete constraint. However, this independent relaxation step also brings quantization errors, resulting in poor retrieval performances. Other cross-modal hashing methods try to directly optimize the challenging objective function with discrete binary constraints. Inspired by [1], we propose a novel supervised cross-modal hashing method called Discrete Cross-Modal Hashing (DCMH) to learn the discrete binary codes without relaxing them. DCMH is formulated through reconstructing the semantic similarity matrix and learning binary codes as ideal features for classification. Furthermore, DCMH alternately updates binary codes of each modality, and iteratively learns the discrete hashing codes bit by bit efficiently, which is quite promising for large-scale datasets. Extensive empirical results on three real-world datasets show that DCMH outperforms the baseline approaches significantly.
Dekui Ma, Jian Liang 0001, Xiangwei Kong 0001, Ran He 0001, Ying Li 0016
ISM5
2016 Sparse Feature Preservation for Relative Attribute Learning
abstract
Relative attributes learning provides a way to capture the strength of the attributes under consideration and it can provide a more specific and accurate information to describe images. But for computers, extracting low-level features is the foundation of understanding images. Thus there is no doubt that the features will have an significant influence on relative attribute models learning. %For example, local features don't have much positive effects on learning global attributes. In this paper, we propose a sparse feature preservation (SFP) method to preserve the most important features on the learning of each attribute model. SFP is formulated through using rearrangement inequality according to relative attribute models learning. We first train the relative attribute models according to the supervision information of attribute pairs. Then the sorting results are used to train the key feature preservation factors and the sparse features are utilized to retrain the relative attribute models. We demonstrate the approach on five datasets and show its significant improvement on the accuracy of relative attribute learning.
Xiangwei Kong 0001, Hongxue Yang, Ying Li 0016
ISM4
2016 Exploiting Hierarchical Activations of Neural Network for Image Retrieval
abstract
The Convolutional Neural Networks (CNNs) have achieved breakthroughs on several image retrieval benchmarks. Most previous works re-formulate CNNs as global feature extractors used for linear scan. This paper proposes a Multi-layer Orderless Fusion (MOF) approach to integrate the activations of CNN in the Bag-of-Words (BoW) framework. Specifically, through only one forward pass in the network, we extract multi-layer CNN activations of local patches. Activations from each layer are aggregated in one BoW model, and several BoW models are combined with late fusion. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed method.
Ying Li 0016, Xiangwei Kong 0001, Liang Zheng 0001, Qi Tian 0001
ACM Multimedia1