Yugeng Lin

dblp:273/4073 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 46% Representation and self-supervised learning · 46% Segmentation and scene understanding · 7%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
detection transformer
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Computer vision › Image recognition and object detection
object detection
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Machine learning › Representation and self-supervised learning
pre-training
0.712023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task
0.512021
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.322023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021

Methods — techniques the papers use, named apart from their topics

transformer · 0.7pretext task · 0.7attention mask · 0.7transformer encoder-decoder · 0.5query patch detection · 0.5
YearPublicationVenuePosition
2023 Unsupervised Pre-Training for Detection Transformers
abstract
DEtection TRansformer (DETR) for object detection reaches competitive performance compared with Faster R-CNN via a transformer encoder-decoder architecture. However, trained with scratch transformers, DETR needs large-scale training data and an extreme long training schedule even on COCO dataset. Inspired by the great success of pre-training transformers in natural language processing, we propose a novel pretext task named random query patch detection in Unsupervised Pre-training DETR (UP-DETR). Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the input image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we find that freezing the CNN backbone is the prerequisite for the success of pre-training transformers. (2) To perform multi-query localization, we develop UP-DETR with multi-query patch detection with attention mask. Besides, UP-DETR also provides a unified perspective for fine-tuning object detection and one-shot detection tasks. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers
abstract
Object detection with transformers (DETR) reaches competitive performance with Faster R-CNN via a transformer encoder-decoder architecture. Inspired by the great success of pre-training transformers in natural language processing, we propose a pretext task named random query patch detection to Unsupervisedly Pre-train DETR (UP-DETR) for object detection. Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the original image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we freeze the CNN backbone and propose a patch feature reconstruction branch which is jointly optimized with patch detection. (2) To perform multi-query localization, we introduce UP-DETR from single-query patch and extend it to multi-query patches with object query shuffle and attention mask. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
CVPR3