VLDB 2026 Research / reviewers in the wild / expert
Yun-Hao Cao
dblp:254/1292
· DBLP profile ↗
10ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0002-6229-8469ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Better Vision-Inspired Vision-Language ModelsabstractVision-language (VL) models have achieved unprece-dented success recently, in which the connection module is the key to bridge the modality gap. Nevertheless, the abun-dant visual clues are not sufficiently exploited in most existing methods. On the vision side, most existing approaches only use the last feature of the vision tower, without using the low-level features. On the language side, most existing meth-ods only introduce shallow vision-language interactions. In this paper, we present a vision-inspired vision-language con-nection module, dubbed as VIVL, which efficiently exploits the vision cue for VL models. To take advantage of the lower-level information from the vision tower, a feature pyramid extractor (FPE) is introduced to combine features from differ-ent intermediate layers, which enriches the visual cue with negligible parameters and computation overhead. To en-hance VL interactions, we propose deep vision-conditioned prompts (DVCP) that allows deep interactions of vision and language features efficiently. Our VIVL exceeds the previous state-of-the-art method by 18.1 CIDEr when training from scratch on the COCO caption task, which greatly improves the data efficiency. When used as a plug-in module, VIVL consistently improves the performance for various backbones and VL frameworks, delivering new state-of-the-art results on multiple benchmarks, e.g., NoCaps and VQAv2. Yun-Hao Cao, Kaixiang Ji, Chuanyang Zheng, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen, Ming Yang 0007 |
CVPR | 1 |
| 2024 | Random Subspace Sampling for Classification with Missing Data
Yun-Hao Cao, Jian-Xin Wu |
J. Comput. Sci. Technol. | 1 |
| 2024 | Tobias: A Random CNN Sees ObjectsabstractThis paper starts by revealing a surprising finding: without any learning, a randomly initialized CNN can localize objects surprisingly well. That is, a CNN has an inductive bias to naturally focus on objects, named as Tobias (“Theobjectisatsight”) in this paper. This empirical inductive bias is further theoretically analyzed and empirically verified, and successfully applied to self-supervised learning as well as supervised learning. For self-supervised learning, a CNN is encouraged to learn representations that focus on the foreground object, by transforming every image into various versions with different backgrounds, where the foreground and background separation is guided by Tobias. Experimental results show that the proposed Tobias significantly improves downstream tasks, especially for object detection. This paper also shows that Tobias has consistent improvements on training sets of different sizes, and is more resilient to changes in image augmentations. Furthermore, we apply Tobias to supervised image classification by letting the average pooling layer focus on foreground regions, which achieves improved performance on various benchmarks. Yun-Hao Cao, Jianxin Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Three Guidelines You Should Know for Universally Slimmable Self-Supervised LearningabstractWe propose universally slimmable self-supervised learning (dubbed as US3L) to achieve better accuracy-efficiency trade-offs for deploying self-supervised models across different devices. We observe that direct adaptation of self-supervised learning (SSL) to universally slimmable networks misbehaves as the training process frequently collapses. We then discover that temporal consistent guidance is the key to the success of SSL for universally slimmable networks, and we propose three guidelines for the loss design to ensure this temporal consistency from a unified gradient perspective. Moreover, we propose dynamic sampling and group regularization strategies to simultaneously improve training efficiency and accuracy. Our US3L method has been empirically validated on both convolutional neural networks and vision transformers. With only once training and one copy of weights, our method outperforms various state-of-the-art methods (individually trained or not) on benchmarks including recognition, object detection and instance segmentation. Yun-Hao Cao, Peiqin Sun, Shuchang Zhou 0001 |
CVPR | 1 |
| 2022 | A Random CNN Sees Objects: One Inductive Bias of CNN and Its ApplicationsabstractThis paper starts by revealing a surprising finding: without any learning, a randomly initialized CNN can localize objects surprisingly well. That is, a CNN has an inductive bias to naturally focus on objects, named as Tobias ("The object is at sight") in this paper. This empirical inductive bias is further analyzed and successfully applied to self-supervised learning (SSL). A CNN is encouraged to learn representations that focus on the foreground object, by transforming every image into various versions with different backgrounds, where the foreground and background separation is guided by Tobias. Experimental results show that the proposed Tobias significantly improves downstream tasks, especially for object detection. This paper also shows that Tobias has consistent improvements on training sets of different sizes, and is more resilient to changes in image augmentations. Yun-Hao Cao, Jianxin Wu 0001 |
AAAI | 1 |
| 2022 | Synergistic Self-supervised and Quantization Learning
Yun-Hao Cao, Peiqin Sun, Yechang Huang, Jianxin Wu 0001, Shuchang Zhou 0001 |
ECCV (30) | 1 |
| 2022 | Training Vision Transformers with only 2040 Images
Yun-Hao Cao, Hao Yu 0027, Jianxin Wu 0001 |
ECCV (25) | 1 |
| 2022 | Worst Case Matters for Few-Shot Recognition
Minghao Fu 0001, Yun-Hao Cao, Jianxin Wu 0001 |
ECCV (20) | 2 |
| 2021 | Neural random subspace
Yun-Hao Cao, Jianxin Wu 0001, Hanchen Wang 0002, Joan Lasenby |
Pattern Recognit. | 1 |
| 2020 | Rethinking the Route Towards Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects with only image-level labels. Previous methods often try to utilize feature maps and classification weights to localize objects using image level annotations indirectly. In this paper, we demonstrate that weakly supervised object localization should be divided into two parts: class-agnostic object localization and object classification. For class-agnostic object localization, we should use class-agnostic methods to generate noisy pseudo annotations and then perform bounding box regression on them without class labels. We propose the pseudo supervised object localization (PSOL) method as a new way to solve WSOL. Our PSOL models have good transferability across different datasets without fine-tuning. With generated pseudo bounding boxes, we achieve 58.00% localization accuracy on ImageNet and 74.74% localization accuracy on CUB-200, which have a large edge over previous models. Chen-Lin Zhang, Yun-Hao Cao, Jianxin Wu 0001 |
CVPR | 2 |