Yang Li 0091

dblp:37/4190-91 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2023
0000-0002-8372-1481ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2023 Contrastive Bayesian Analysis for Deep Metric Learning
abstract
Recent methods for deep metric learning have been focusing on designing different contrastive loss functions between positive and negative pairs of samples so that the learned feature embedding is able to pull positive samples of the same class closer and push negative samples from different classes away from each other. In this work, we recognize that there is a significant semantic gap between features at the intermediate feature layer and class labels at the final output layer. To bridge this gap, we develop a contrastive Bayesian analysis to characterize and model the posterior probabilities of image labels conditioned by their features similarity in a contrastive learning setting. This contrastive Bayesian analysis leads to a new loss function for deep metric learning. To improve the generalization capability of the proposed method onto new classes, we further extend the contrastive Bayesian loss with a metric variance constraint. Our experimental results and ablation studies demonstrate that the proposed contrastive Bayesian metric learning method significantly improves the performance of deep metric learning in both supervised and pseudo-supervised scenarios, outperforming existing methods by a large margin.
Shichao Kan, Zhiquan He, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Local Semantic Correlation Modeling Over Graph Neural Networks for Deep Feature Embedding and Image Retrieval
abstract
Deep feature embedding aims to learn discriminative features or feature embeddings for image samples which can minimize their intra-class distance while maximizing their inter-class distance. Recent state-of-the-art methods have been focusing on learning deep neural networks with carefully designed loss functions. In this work, we propose to explore a new approach to deep feature embedding. We learn a graph neural network to characterize and predict the local correlation structure of images in the feature space. Based on this correlation structure, neighboring images collaborate with each other to generate and refine their embedded features based on local linear combination. Graph edges learn a correlation prediction network to predict the correlation scores between neighboring images. Graph nodes learn a feature embedding network to generate the embedded feature for a given image based on a weighted summation of neighboring image features with the correlation scores as weights. Our extensive experimental results under the image retrieval settings demonstrate that our proposed method outperforms the state-of-the-art methods by a large margin, especially for top-1 recalls.
Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
IEEE Trans. Image Process.3
2021 Spatial Assembly Networks for Image Representation Learning
abstract
It has been long recognized that deep neural networks are sensitive to changes in spatial configurations or scene structures. Image augmentations, such as random translation, cropping, and resizing, can be used to improve the robustness of deep neural networks under spatial transforms. However, changes in object part configurations, spatial layout of object, and scene structures of the images may still result in major changes in the their feature representations generated by the network, creating significant challenges for various visual learning tasks, including representation or metric learning, image classification and retrieval. In this work, we introduce a new learnable module, called spatial assembly network (SAN), to address this important issue. This SAN module examines the input image and performs a learned re-organization and assembly of feature points from different spatial locations conditioned by feature maps from previous network layers so as to maximize the discriminative power of the final feature representation. This differentiable module can be flexibly incorporated into existing network architectures, improving their capabilities in handling spatial variations and structural changes of the image scene. We demonstrate that the proposed SAN module is able to significantly improve the performance of various metric / representation learning, image retrieval and classification tasks, in both supervised and unsupervised learning scenarios.
Yang Li 0091, Shichao Kan, Jianhe Yuan, Wenming Cao 0001, Zhihai He
CVPR1
2021 Relative Order Analysis and Optimization for Unsupervised Deep Metric Learning
abstract
In unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to image A than image B. In this work, we propose to explore how this relative order can be used to learn discriminative features with an unsupervised metric learning method. Instead of resorting to clustering or self-supervision to create pseudo labels for an absolute decision, which often suffers from high label error rates, we construct reliable relative orders for groups of image samples and learn a deep neural network to predict these relative orders. During training, this relative order prediction network and the feature embedding network are tightly coupled, providing mutual constraints to each other to improve overall metric learning performance in a cooperative manner. During testing, the predicted relative orders are used as constraints to optimize the generated features and refine their feature distance-based image retrieval results using a constrained optimization procedure. Our experimental results demonstrate that the proposed relative orders for unsupervised learning (ROUL) method is able to significantly improve the performance ofunsupervised deep metric learning.
Shichao Kan, Yi-Gang Cen, Yang Li 0091, Vladimir Mladenovic, Zhihai He
CVPR3
2021 Learned Model Composition With Critical Sample Look-Ahead for Semi-Supervised Learning on Small Sets of Labeled Samples
abstract
In this work, we propose to push the performance limit of semi-supervised learning on very small sets of labeled samples by developing a new method called learned model composition with critical sample look-ahead (LMCS). Training efficient deep neural networks on much smaller sets of labeled samples is a challenging problem. With a small labeled set, the initial network suffers from low accuracy. Based on this error-prone network, the subsequent semi-supervised learning process will be fragile and unstable. To address this issue, we propose to introduce a look-ahead master model to identify the correct direction of model evolution to effectively guide the semi-supervised learning process of the student model. Specifically, our proposed LMCS method explores two major ideas. First, it introduces a new learned model composition structure so that we can compose a more efficient master network from student models of past iterations through a network learning process. Second, we develop a new method, called confined maximum entropy search, to discover new critical samples near the model decision boundary and provide the master model with look-ahead access to these samples to enhance its guidance capability. Our extensive experimental results demonstrate that the proposed LMCS method outperforms the state-of-the-art semi-supervised learning methods, especially on small sets of labeled samples. For example, on the CIFAR-10 dataset, with a very small set of 80 labeled samples, our method outperforms Google's MixMatch method, reducing the error rate by more than 10%.
Yang Li 0091, Shichao Kan, Wenming Cao 0001, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.1
2021 Zero-Shot Learning to Index on Semantic Trees for Scalable Image Retrieval
abstract
In this study, we develop a new approach, called zero-shot learning to index on semantic trees (LTI-ST), for efficient image indexing and scalable image retrieval. Our method learns to model the inherent correlation structure between visual representations using a binary semantic tree from training images which can be effectively transferred to new test images from unknown classes. Based on predicted correlation structure, we construct an efficient indexing scheme for the whole test image set. Unlike existing image index methods, our proposed LTI-ST method has the following two unique characteristics. First, it does not need to analyze the test images in the query database to construct the index structure. Instead, it is directly predicted by a network learnt from the training set. This zero-shot capability is critical for flexible, distributed, and scalable implementation and deployment of the image indexing and retrieval services at large scales. Second, unlike the existing distance-based index methods, our index structure is learnt using the LTI-ST deep neural network with binary encoding and decoding on a hierarchical semantic tree. Our extensive experimental results on benchmark datasets and ablation studies demonstrate that the proposed LTI-ST method outperforms existing index methods by a large margin while providing the above new capabilities which are highly desirable in practice.
Shichao Kan, Yi-Gang Cen, Vladimir Mladenovic, Yang Li 0091, Zhihai He
IEEE Trans. Image Process.5
2021 Snowball: Iterative Model Evolution and Confident Sample Discovery for Semi-Supervised Learning on Very Small Labeled Datasets
abstract
In this work, we develop a joint sample discovery and iterative model evolution method for semi-supervised learning on very small labeled training sets. We propose a master-teacher-student model framework to provide multi-layer guidance during the model evolution process with multiple iterations and generations. The teacher model is constructed by performing an exponential moving average of the student models obtained from past training steps. The master network combines the knowledge of the student and teacher models with additional access to newly discovered samples. The master and teacher models are then used to guide the training of the student network by enforcing the consistency between their predictions of unlabeled samples and evolve all models when more and more samples are discovered. Our extensive experiments demonstrate that the process of discovering confident samples from the unlabeled dataset, once coupled with the master-teacher-student network evolution, can significantly improve the overall semi-supervised learning performance. For example, on the CIFAR-10 dataset, with a small set of 250 labeled samples, our method achieves an error rate of 11.58%, more than 38% lower than Mean-Teacher (49.91%). When coupled with the MixMatch augmentation and loss function, the improvements are also significant.
Yang Li 0091, Zhiqun Zhao, Hao Sun 0024, Yi-Gang Cen, Zhihai He
IEEE Trans. Multim.1
2020 Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss
Yang Li 0091, Shichao Kan, Zhihai He
ECCV (11)1
2020 Multi-Matrices Low-Rank Decomposition With Structural Smoothness for Image Denoising
abstract
In this paper, we propose a multi-matrices lowrank decomposition method for image denoising. In this new method, the total variation (TV) norm is incorporated into lowrank approximation analysis to achieve structural smoothness and to improve quality of the recovered images. Our proposed mathematical framework for multi-matrices low-rank decomposition combines the nuclear norm, TV norm, and L1norm, which allows us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Based on the iterative alternating direction method, we develop an algorithm to solve the proposed challenging optimization problem. We conduct extensive experiments and perform evaluations on multi-images denoising and multi-frames video prediction. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for images with large sparse noise.
Hengyou Wang, Yang Li 0091, Yi-Gang Cen, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.2