Zhigang Dai

dblp:260/6573 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
6since 2021 · last 2023
0000-0002-8616-0202ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 46% Representation and self-supervised learning · 46% Segmentation and scene understanding · 7%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
detection transformer
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Computer vision › Image recognition and object detection
object detection
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training
1.222023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Machine learning › Representation and self-supervised learning
pre-training
0.712023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task
0.512021
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.322023
Unsupervised Pre-Training for Detection Transformers · IEEE Trans. Pattern Anal. Mach. Intell. 2023
UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers · CVPR 2021

Methods — techniques the papers use, named apart from their topics

transformer · 0.7pretext task · 0.7attention mask · 0.7transformer encoder-decoder · 0.5query patch detection · 0.5
YearPublicationVenuePosition
2023 Unsupervised Pre-Training for Detection Transformers
abstract
DEtection TRansformer (DETR) for object detection reaches competitive performance compared with Faster R-CNN via a transformer encoder-decoder architecture. However, trained with scratch transformers, DETR needs large-scale training data and an extreme long training schedule even on COCO dataset. Inspired by the great success of pre-training transformers in natural language processing, we propose a novel pretext task named random query patch detection in Unsupervised Pre-training DETR (UP-DETR). Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the input image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we find that freezing the CNN backbone is the prerequisite for the success of pre-training transformers. (2) To perform multi-query localization, we develop UP-DETR with multi-query patch detection with attention mask. Besides, UP-DETR also provides a unified perspective for fine-tuning object detection and one-shot detection tasks. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 OPTI: Order Preparation Time Inference for On-demand Delivery
abstract
On-demand delivery has become an increasingly popular urban service in recent years as it facilitates citizens’ daily lives significantly. In the fulfillment cycle, the order preparation time estimation is extremely important and can be used for many applications, such as improving order dispatching and fulfillment time estimation. Existing work is generally based on high-cost physical devices or large-scale labeled training data, which are not feasible in on-demand delivery services. We solve this problem based on already collected different kinds of data from the on-demand delivery platform, e.g., the courier’s reported arrival time to the merchant. Our intuition is that the couriers’ reported time implicitly reflects the order preparation time, which leads to a challenge: complicated correlations between the couriers’ reported arrival time and the order preparation time. To solve this challenge, we design an order preparation time inference framework OPTI, which first constructs a self-supervised classification task based on the couriers’ reported arrival time to infer the coarse-grained order preparation time and then exploits semi-supervised learning to transfer the coarse-grained time to fine-grained time inference. Experimental results show that OPTI can improve the accuracy of inference by 5% to 17% compared to the state-of-the-art solutions.
Zhigang Dai, Wenjun Lyu, Yi Ding 0011, Yiwei Song, Yunhuai Liu
ACM Trans. Sens. Networks1
2022 UniMoCo: Unsupervised, Semi-Supervised and Fully-Supervised Visual Representation Learning
abstract
Momentum Contrast (MoCo) achieves great success for unsupervised visual representation learning. However, there are a lot of supervised and semi-supervised datasets, which are already labeled. To fully utilize the label annotations, we propose Unified Momentum Contrast (UniMoCo), which extends MoCo to support arbitrary ratios of labeled data and unlabeled data training. Compared with MoCo, UniMoCo has two modifications as follows: (1) Different from a single positive pair in MoCo, we maintain multiple positive pairs on-the-fly by comparing the query label to a label queue. (2) We propose a Unified Contrastive (UniCon) loss to support an arbitrary number of positives and negatives in a unified pair-wise optimization perspective. Our UniCon is more reasonable and powerful than the supervised contrastive loss in theory and practice. In our experiments, we pre-train multiple UniMoCo models with different ratios of ImageNet labels and evaluate the performance on various downstream tasks. Experiment results show that UniMoCo generalizes well for unsupervised, semi-supervised and fully-supervised visual representation learning. Besides, we surprisingly find that UniMoCo performs best with 60% ImageNet labels for COCO and VOC transfer learning. The code is available: https://github.com/dddzg/unimoco.
Zhigang Dai, Bolun Cai
SMC1
2022 Spatial Consistency and Feature Diversity Regularization in Transfer Learning for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization is challenged by limited training data by localizing discriminative regions and learning diverse features. We propose an effective regularization method that simultaneously imposes spatial consistency and feature diversity on CNN feature maps from a unified perspective. The former guides different feature map channels to concentrate collaboratively on the discriminative areas while the latter ensures that the feature maps are diverse. The proposed method does not require additional supervision, and leverages the covariance matrix of multi-channel feature maps to regularize the loss at the last convolutional layer where the semantic information is the richest. This allows the influence to be backpropagated to update all convolutional layers. We perform experiments using four network architectures for transfer learning from two source domains to three target domains, and demonstrate that our regularization method improves accuracy in all different settings. The proposed regularization method achieves state-of-the-art performance on CUB-200-2011, Stanford-Cars, and Stanford-Dogs datasets with 89.8%, 94.6%, and 88.5% accuracy, respectively.
Zhigang Dai, Ajmal Mian
SMC1
2021 UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers
abstract
Object detection with transformers (DETR) reaches competitive performance with Faster R-CNN via a transformer encoder-decoder architecture. Inspired by the great success of pre-training transformers in natural language processing, we propose a pretext task named random query patch detection to Unsupervisedly Pre-train DETR (UP-DETR) for object detection. Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the original image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we freeze the CNN backbone and propose a patch feature reconstruction branch which is jointly optimized with patch detection. (2) To perform multi-query localization, we introduce UP-DETR from single-query patch and extend it to multi-query patches with object query shuffle and attention mask. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
CVPR1
2021 OPTI: Order Preparation Time Inference for On-demand Delivery
abstract
On-demand delivery has become an increasingly popular urban service in recent years as it facilitates citizens' daily lives significantly. Different from traditional logistics services, e.g., FedEx and UPS, on-demand delivery is expected to be completed within a relatively short time, e.g., 30 minutes to 1 hour. In the fulfillment cycle, the order preparation time estimation is extremely important, which can be used for many applications such as improving order dispatching and fulfillment time estimation. Existing work is generally based on high-cost physical devices or large-scale labeled training data, which are not feasible in on-demand delivery services. We solve this problem based on already collected different kinds of data from the on-demand delivery platform, e.g., the courier's reported arrival time to the merchant. Our intuition is that couriers' reported time implicitly reflects the order preparation time, which leads to a challenge: complicated correlations between the couriers' reported arrival time and the order preparation time. To solve this challenge, we design an order preparation time inference framework OPTI, which first constructs a self-supervised classification task based on the couriers' reported arrival time to infer the coarse-grained order preparation time, and then exploits semi-supervised learning to transfer the coarse-grained time to fine-grained time inference. We implement and evaluate OPTI in the [anonymous] platform, which is one of the largest on-demand delivery platforms in the world. Experimental results show that OPTI can improve the accuracy of inference by 5% to 17% compared to the state-of-the-art solutions.
Zhigang Dai, Wenjun Lyu, Yi Ding 0011, Yiwei Song
ICPADS1