VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0314
dblp:64/6025-314
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2024
0000-0003-1377-1704ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Representation and self-supervised learning · 32% Segmentation and scene understanding · 26% Learning paradigms · 13% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
2.0 | 4 | 2023 | Balancing Logit Variation for Long-Tailed Semantic Segmentation · CVPR 2023 Learning from Future: A Novel Self-Training Framework for Semantic Segmentation · NeurIPS 2022 Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels · CVPR 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling |
1.9 | 3 | 2024 | Efficient Masked Autoencoders With Self-Consistency · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Exploring Stochastic Autoregressive Image Modeling for Visual Representation · AAAI 2023 MST: Masked Self-Supervised Transformer for Visual Representation · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
1.2 | 2 | 2023 | Exploring Stochastic Autoregressive Image Modeling for Visual Representation · AAAI 2023 MST: Masked Self-Supervised Transformer for Visual Representation · NeurIPS 2021 |
Computer vision › Segmentation and scene understanding › annotation-efficient segmentation
semi-supervised semantic segmentation |
1.1 | 2 | 2022 | Learning from Future: A Novel Self-Training Framework for Semantic Segmentation · NeurIPS 2022 Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels · CVPR 2022 |
Machine learning › Efficient and distributed learning › efficient training
efficient pre-training |
0.8 | 1 | 2024 | Efficient Masked Autoencoders With Self-Consistency · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder |
0.8 | 1 | 2024 | Efficient Masked Autoencoders With Self-Consistency · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 2 | 2022 | Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks · NeurIPS 2022 MST: Masked Self-Supervised Transformer for Visual Representation · NeurIPS 2021 |
Machine learning › Generative modeling › autoregressive model
autoregressive image modeling |
0.7 | 1 | 2023 | Exploring Stochastic Autoregressive Image Modeling for Visual Representation · AAAI 2023 |
Machine learning › Learning paradigms
imbalanced learning |
0.7 | 1 | 2023 | Balancing Logit Variation for Long-Tailed Semantic Segmentation · CVPR 2023 |
Computer vision › Segmentation and scene understanding › semantic segmentation › multi-class segmentation
long-tailed semantic segmentation |
0.7 | 1 | 2023 | Balancing Logit Variation for Long-Tailed Semantic Segmentation · CVPR 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Representation and self-supervised learning › contrastive learning
instance discrimination |
0.6 | 1 | 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Learning paradigms
multi-label classification |
0.6 | 1 | 2022 | Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks · NeurIPS 2022 |
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling |
0.6 | 1 | 2022 | Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels · CVPR 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised visual pre-training |
0.6 | 1 | 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-training · CVPR 2022 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.6 | 1 | 2022 | Learning from Future: A Novel Self-Training Framework for Semantic Segmentation · NeurIPS 2022 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2022 | Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 0.8parallel mask strategy · 0.8stochastic permutation · 0.7perturbation · 0.7encoder-decoder · 0.7category-wise logit variation · 0.7optimal transport · 0.6negative sampling · 0.6entropy-based thresholding · 0.6class prompting · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Masked Autoencoders With Self-ConsistencyabstractInspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio of MIM results in two serious problems: 1) the inadequate data utilization of images within each iteration brings prolonged pre-training, and 2) the high inconsistency of predictions results in unreliable generations, i.e., the prediction of the identical patch may be inconsistent in different mask rounds, leading to divergent semantics in the ultimately generated outcomes. To tackle these problems, we propose the efficient masked autoencoders with self-consistency (EMAE) to improve the pre-training efficiency and increase the consistency of MIM. In particular, we present a parallel mask strategy that divides the image into K non-overlapping parts, each of which is generated by a random mask with the same mask ratio. Then the MIM task is conducted parallelly on all parts in an iteration and the model minimizes the loss between the predictions and the masked patches. Besides, we design the self-consistency learning to further maintain the consistency of predictions of overlapping masked patches among parts. Overall, our method is able to exploit the data more efficiently and obtains reliable representations. Experiments on ImageNet show that EMAE achieves the best performance on ViT-Large with only 13% of MAE pre-training time using NVIDIA A100 GPUs. After pre-training on diverse datasets, EMAE consistently obtains state-of-the-art transfer ability on a variety of downstream tasks, such as image classification, object detection, and semantic segmentation. Zhaowen Li, Yousong Zhu, Zhiyang Chen 0002, Wei Li 0314, Rui Zhao 0001, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Exploring Stochastic Autoregressive Image Modeling for Visual RepresentationabstractAutoregressive language modeling (ALM) has been successfully used in self-supervised pre-training in Natural language processing (NLP). However, this paradigm has not achieved comparable results with other self-supervised approaches in computer vision (e.g., contrastive learning, masked image modeling). In this paper, we try to find the reason why autoregressive modeling does not work well on vision tasks. To tackle this problem, we fully analyze the limitation of visual autoregressive methods and proposed a novel stochastic autoregressive image modeling (named SAIM) by the two simple designs. First, we serialize the image into patches. Second, we employ the stochastic permutation strategy to generate an effective and robust image context which is critical for vision tasks. To realize this task, we create a parallel encoder-decoder training process in which the encoder serves a similar role to the standard vision transformer focusing on learning the whole contextual information, and meanwhile the decoder predicts the content of the current position so that the encoder and decoder can reinforce each other. Our method significantly improves the performance of autoregressive image modeling and achieves the best accuracy (83.9%) on the vanilla ViT-Base model among methods using only ImageNet-1K data. Transfer performance in downstream tasks also shows that our model achieves competitive performance. Code is available at https://github.com/qiy20/SAIM. Fan Yang 0089, Yousong Zhu, Rui Zhao 0001, Wei Li 0314 |
AAAI | 7 |
| 2023 | SeqCo-DETR: Sequence Consistency Training for Self-Supervised Object Detection with Transformers
Guoqiang Jin, Fan Yang 0089, Mingshan Sun, Ruyi Zhao, Yakun Liu, Wei Li 0314, Tianpeng Bao, Xingyu Zeng, Rui Zhao 0001 |
BMVC | 6 |
| 2023 | Balancing Logit Variation for Long-Tailed Semantic SegmentationabstractSemantic segmentation usually suffers from a long-tail data distribution. Due to the imbalanced number of samples across categories, the features of those tail classes may get squeezed into a narrow area in the feature space. Towards a balanced feature distribution, we introduce category-wise variation into the network predictions in the training phase such that an instance is no longer projected to a feature point, but a small region instead. Such a perturbation is highly dependent on the category scale, which appears as assigning smaller variation to head classes and larger variation to tail classes. In this way, we manage to close the gap between the feature areas of different categories, resulting in a more balanced representation. It is note-worthy that the introduced variation is discarded at the inference stage to facilitate a confident prediction. Although with an embarrassingly simple implementation, our method manifests itself in strong generalizability to various datasets and task settings. Extensive experiments suggest that our plug-in design lends itself well to a range of state-of-the-art approaches and boosts the performance on top of them.11Code: https://github.com/grantword8/BLV. Jingjing Fei, Wei Li 0314, Tianpeng Bao, Rui Zhao 0001, Yujun Shen |
CVPR | 4 |
| 2022 | UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingabstractSelf-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential. Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang |
CVPR | 4 |
| 2022 | Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-LabelsabstractThe crux of semi-supervised semantic segmentation is to assign adequate pseudo-labels to the pixels of unlabeled images. A common practice is to select the highly confident predictions as the pseudo ground-truth, but it leads to a problem that most pixels may be left unused due to their unreliability. We argue that every pixel matters to the model training, even its prediction is ambiguous. Intuitively, an unreliable prediction may get confused among the top classes (i.e., those with the highest probabilities), however, it should be confident about the pixel not belonging to the remaining classes. Hence, such a pixel can be convincingly treated as a negative sample to those most unlikely categories. Based on this insight, we develop an effective pipeline to make sufficient use of unlabeled data. Concretely, we separate reliable and unreliable pixels via the entropy of predictions, push each unreliable pixel to a category-wise queue that consists of negative samples, and manage to train the model with all candidate pixels. Considering the training evolution, where the prediction becomes more and more accurate, we adaptively adjust the threshold for the reliable-unreliable partition. Experimental results on various benchmarks and training settings demonstrate the superiority of our approach over the state-of-the-art alternatives.11Project: https://haochen-wang409.github.io/U2PL. Yujun Shen, Jingjing Fei, Wei Li 0314, Guoqiang Jin, Rui Zhao 0001, Xinyi Le |
CVPR | 5 |
| 2022 | Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual TasksabstractVisual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensional outputs in object-level visual tasks. In this paper, we propose an object-centric vision framework, Obj2Seq. Obj2Seq takes objects as basic units, and regards most object-level visual tasks as sequence generation problems of objects. Therefore, these visual tasks can be decoupled into two steps. First recognize objects of given categories, and then generate a sequence for each of these objects. The definition of the output sequences varies for different tasks, and the model is supervised by matching these sequences with ground-truth targets. Obj2Seq is able to flexibly determine input categories to satisfy customized requirements, and be easily extended to different visual tasks. When experimenting on MS COCO, Obj2Seq achieves 45.7% AP on object detection, 89.0% AP on multi-label classification and 65.0% AP on human pose estimation. These results demonstrate its potential to be generally applied to different visual tasks. Code has been made available at: https://github.com/CASIA-IVA-Lab/Obj2Seq. Zhiyang Chen 0002, Yousong Zhu, Zhaowen Li, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Rui Zhao 0001, Jinqiao Wang, Ming Tang 0001 |
NeurIPS | 5 |
| 2022 | Learning from Future: A Novel Self-Training Framework for Semantic SegmentationabstractSelf-training has shown great potential in semi-supervised learning. Its core idea is to use the model learned on labeled data to generate pseudo-labels for unlabeled samples, and in turn teach itself. To obtain valid supervision, active attempts typically employ a momentum teacher for pseudo-label prediction yet observe the confirmation bias issue, where the incorrect predictions may provide wrong supervision signals and get accumulated in the training process. The primary cause of such a drawback is that the prevailing self-training framework acts as guiding the current state with previous knowledge because the teacher is updated with the past student only. To alleviate this problem, we propose a novel self-training strategy, which allows the model to learn from the future. Concretely, at each training step, we first virtually optimize the student (i.e., caching the gradients without applying them to the model weights), then update the teacher with the virtual future student, and finally ask the teacher to produce pseudo-labels for the current student as the guidance. In this way, we manage to improve the quality of pseudo-labels and thus boost the performance. We also develop two variants of our future-self-training (FST) framework through peeping at the future both deeply (FST-D) and widely (FST-W). Taking the tasks of unsupervised domain adaptive semantic segmentation and semi-supervised semantic segmentation as the instances, we experimentally demonstrate the effectiveness and superiority of our approach under a wide range of settings. Code is available at https://github.com/usr922/FST. Ye Du 0002, Yujun Shen, Jingjing Fei, Wei Li 0314, Rui Zhao 0001, Zehua Fu, Qingjie Liu 0001 |
NeurIPS | 5 |
| 2021 | MST: Masked Self-Supervised Transformer for Visual RepresentationabstractTransformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only consider the high-level feature and learning representation from a global perspective, which may fail to transfer to the downstream dense prediction tasks focusing on local features. In this paper, we present a novel Masked Self-supervised Transformer approach named MST, which can explicitly capture the local context of an image while preserving the global semantic information. Specifically, inspired by the Masked Language Modeling (MLM) in NLP, we propose a masked token strategy based on the multi-head self-attention map, which dynamically masks some tokens of local patches without damaging the crucial structure for self-supervised learning. More importantly, the masked tokens together with the remaining tokens are further recovered by a global image decoder, which preserves the spatial information of the image and is more friendly to the downstream dense prediction tasks. The experiments on multiple datasets demonstrate the effectiveness and generality of the proposed method. For instance, MST achieves Top-1 accuracy of 76.9% with DeiT-S only using 300-epoch pre-training by linear evaluation, which outperforms supervised methods with the same epoch by 0.4% and its comparable variant DINO by 1.0%. For dense prediction tasks, MST also achieves 42.7% mAP on MS COCO object detection and 74.04% mIoU on Cityscapes segmentation only with 100-epoch pre-training. Zhaowen Li, Zhiyang Chen 0002, Fan Yang 0089, Wei Li 0314, Yousong Zhu, Chaoyang Zhao, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang |
NeurIPS | 4 |