Yuan Tai

dblp:230/3168 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
feature matching
0.612022
Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss · CVPR 2022
Computer vision › 3D vision
invariant feature extraction
0.612022
Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss · CVPR 2022
Computer vision › 3D vision › low-level vision › feature detection
keypoint detection
0.612022
Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss · CVPR 2022

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.1probabilistic covariant loss · 0.6discrete distribution prediction · 0.6
YearPublicationVenuePosition
2026 Cross-modal transformer fusion via local sampling for drone RGB-infrared object detection
Herong Qi, Xuanyu Xiang, Hui Qin, Yuan Tai, Yihua Tan
Neural Networks4
2024 Where to model the epistemic uncertainty of Bayesian convolutional neural networks for classification
Yuan Tai, Yihua Tan, Erbo Zou
Neurocomputing1
2023 Mine-Distill-Prototypes for Complete Few-Shot Class-Incremental Learning in Image Classification
abstract
Recently, few-shot learning (FSL) has received increasing attention because of difficulties in sample collection in some application scenarios, such as maritime surveillance using synthetic aperture radar (SAR) or infrared images. In real situations of such scenarios, it is a common requirement that the model can recognize novel classes incrementally, namely class-incremental learning (CIL). Considering the above requirement, a novel problem that recognizes novel classes incrementally when both the base and novel class samples are scarce is proposed in this article. It is called complete few-shot CIL (C-FSCIL) for distinguishing from the FSCIL that assumes sufficient samples of base classes. Specifically, the following challenges of C-FSCIL are focused on: 1) distance measurement is used for recognizing novel classes incrementally, but the encoder is difficult to be learned well when base class samples are scarce, making some features unsuitable for calculating the distance, decreasing the performance and 2) the catastrophic forgetting problem becomes more difficult to be alleviated than that in FSCIL because of the scarcity of base class samples. To tackle both challenges, mine-distill-prototypes (MDP) algorithm is proposed, which consists of two parts: 1) prototypes-distillation (PD) network is proposed to learn to distill the features and prototypes into a lower dimensional in which ineffective features are eliminated and 2) the prototypes-weight (PW) network and the prototypes-selection (PS) training strategy are proposed for the catastrophic forgetting problem, which aims to capture the relationship between the base and novel prototypes. The superior performance of the proposed algorithm is demonstrated by the experiments on three datasets.
Yuan Tai, Yihua Tan, Shengzhou Xiong, Jinwen Tian
IEEE Trans. Geosci. Remote. Sens.1
2022 Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss
abstract
Estimating keypoint scale and orientation is crucial to extracting invariant features under significant geometric changes. Recently, the estimators based on self-supervised learning have been designed to adapt to complex imaging conditions. Such learning-based estimators generally predict a single scalar for the keypoint scale or orientation, called hard estimators. However, hard estimators are difficult to handle the local patches containing structures of different objects or multiple edges. In this paper, a Soft Self-Supervised Estimator (S3Esti) is proposed to overcome this problem by learning to predict multiple scales and orientations. S3Esti involves three core factors. First, the estimator is constructed to predict the discrete distributions of scales and orientations. The elements with high confidence will be kept as the final scales and orientations. Second, a probabilistic covariant loss is proposed to improve the consistency of the scale and orientation distributions under different transformations. Third, an optimization algorithm is designed to minimize the loss function, whose convergence is proved in theory. When combined with different keypoint extraction models, S3Esti generally improves over 50% accuracy in image matching tasks under significant viewpoint changes. In the 3D reconstruction task, S3Esti decreases more than 10% reprojection error and improves the number of registered images. [code release]
Pei Yan, Yihua Tan, Shengzhou Xiong, Yuan Tai, Yansheng Li 0001
CVPR4
2022 Repeatable adaptive keypoint detection via self-supervised learning
Pei Yan, Yihua Tan, Yuan Tai
Sci. China Inf. Sci.3
2021 Subspace reconstruction based correlation filter for object tracking
Yuan Tai, Yihua Tan, Shengzhou Xiong, Jinwen Tian
Comput. Vis. Image Underst.1
2021 Unsupervised learning framework for interest point detection and description via properties optimization
Pei Yan, Yihua Tan, Yuan Tai, Dongrui Wu, Hanbin Luo, Xiaolong Hao
Pattern Recognit.3
2020 Vehicle Detection with Bottom Enhanced RetinaNet in Aerial Images
abstract
Vehicle detection is one of the hot topics in lane detection and vehicle counting. Many works have been done on it and some data sets with satellite images and aerial images are proposed. However, the ratio of the vehicle targets to the background is small and the detection results are unsatisfied. In this paper, a bottom enhanced RetinaNet model named En-RetinaNet is proposed to get better performance on vehicle detection. The EnRetinaNet includes an enhanced feature pyramid network(FPN) and a bottom-top fusion before the region proposal network. Enhanced feature pyramid network adds a bottom layer to the feature pyramid network to exploit more local features. Bottom-top fusion is utilized to get a better fusion of the bottom layers and top layers. In order to get a high ratio of the objects to the images, we take a sliding window mechanism on the testing images. We discuss the effects of the training input crop size on the final results and choose a moderate size of the training input. With all the above work done, we get an improvement on the UCAS—AOD data set in contrast to the RetinaNet.
Peng Gao 0012, Jinwen Tian, Yuan Tai, Tianming Zhao 0003
IGARSS3