EDBT 2026 Demo / reviewers in the wild / expert
Yutong Lin
dblp:261/9395
· DBLP profile ↗
16ranked-venue papers
3as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DAMap: Distance-Aware MapNet for High Quality HD Map Construction
Jinpeng Dong, Yutong Lin, Jingwen Fu, Sanping Zhou, Nanning Zheng 0001 |
ICCV | 3 |
| 2024 | AugDETR: Improving Multi-scale Learning for Detection Transformer
Jinpeng Dong, Yutong Lin, Sanping Zhou, Nanning Zheng 0001 |
ECCV (24) | 2 |
| 2024 | V-DETR: DETR with Vertex Relative Position Encoding for 3D Object DetectionabstractWe introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that are far away from the target objects, violating the locality principle in object detection. To address the limitation, we introduce a novel 3D Vertex Relative Position Encoding (3DV-RPE) method which computes position encoding for each point based on its relative position to the 3D boxes predicted by the queries in each decoder layer, thus providing clear information to guide the model to focus on points near the objects, in accordance with the principle of locality. Furthermore, we have systematically refined our pipeline, including data normalization, to better align with the task requirements. Our approach demonstrates remarkable performance on the demanding ScanNetV2 benchmark, showcasing substantial enhancements over the prior state-of-the-art CAGroup3D. Specifically, we achieve an increase in $AP_{25}$ from $75.1\%$ to $77.8\%$ and in ${AP}_{50}$ from $61.3\%$ to $66.0\%$. Yichao Shen 0001, Zigang Geng, Yuhui Yuan, Yutong Lin, Chunyu Wang 0001, Han Hu 0001, Nanning Zheng 0001, Baining Guo |
ICLR | 4 |
| 2024 | Smart Sensing and Communication Co-Design for IIoT-Based Control SystemsabstractIndustrial Internet of Things (IIoT)-based control is growing rapidly, such as smart factories and industrial automation. Sensing and transmitting physical state measurements is the first step and the prerequisite for IIoT-based control. However, sensor interference (e.g., electromagnetic interference on sensing, temperature, and humidity variations in the field) and network interference (e.g., metal obstacles and background noises) may destroy the control performance by interfering with sensing and communication processes. Most of the present upstream “fixed sensors-networking-state estimation” approaches cannot effectively deal with sensor and network interferences due to the fixed measurements/estimation and network resource limitations. To optimize the performance of IIoT-based control, we propose a smart sensing and communication co-design (SSCC) framework to select more potential sensors and establish the corresponding network scheduling. SSCC consists of a smart estimator (SE) and a sensing communication mode switching (SCMS) agent. The SE detects sensor interference and obtains resilient state estimation based on collaborative sensing. SCMS agent dynamically switches sensor selections and network configurations (routing and transmission number) in an integrated manner based on the network and plant states by solving a performance optimization problem. We propose a lightweight SCMS approach by searching a predefined mode table. We perform simulations integrating TOSSIM and MATLAB/Simulink, and semi-physical experiments on a real wireless sensor-actuator network composed of TelosB nodes. The results show that the SSCC framework can effectively improve the control performance and enhance network energy efficiency under various types of interference by dynamically selecting sensors and allocating network resources. Ruijie Fu, Jintao Chen 0001, Yutong Lin, An Zou, Cailian Chen, Xin-Ping Guan, Yehan Ma |
IEEE Internet Things J. | 3 |
| 2023 | On Data Scaling in Masked Image ModelingabstractScaling properties have been one of the central issues in self-supervised pre-training, especially the data scalability, which has successfully motivated the large-scale self-supervised pre-trained language models and endowed them with significant modeling capabilities. However, scaling properties seem to be unintentionally neglected in the recent trending studies on masked image modeling (MIM), and some arguments even suggest that MIM cannot benefit from large-scale data. In this work, we try to break down these preconceptions and systematically study the scaling behaviors of MIM through extensive experiments, with data ranging from 10% of ImageNet-1K to full ImageNet-22K, model parameters ranging from 49-million to one-billion, and training length ranging from 125K to 500K iterations. And our main findings can be summarized in two folds: 1) masked image modeling remains demanding large-scale data in order to scale up computes and model parameters; 2) masked image modeling cannot benefit from more data under a non-overfitting scenario, which diverges from the previous observations in self-supervised pre-trained language models or supervised pre-trained vision models. In addition, we reveal several intriguing properties in MIM, such as high sample efficiency in large MIM models and strong correlation between pre-training validation loss and transfer performance. We hope that our findings could deepen the understanding of masked image modeling and facilitate future developments on largescale vision models. Code and models will be available at https://github.com/microsoft/SimMIM. Zhenda Xie, Zheng Zhang 0022, Yue Cao 0001, Yutong Lin, Yixuan Wei, Qi Dai 0001, Han Hu 0001 |
CVPR | 4 |
| 2023 | DETR Does Not Need Multi-Scale or Locality DesignabstractThis paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi-scale and locality into the decoder. We show that two simple technologies are surprisingly effective within a plain design to compensate for the lack of multi-scale feature maps and locality constraints. The first is a box-to-pixel relative position bias (BoxRPB) term added to the cross-attention formulation, which well guides each query to attend to the corresponding object region while also providing encoding flexibility. The second is masked image modeling (MIM)-based backbone pre-training which helps learn representation with fine-grained localization ability and proves crucial for remedying dependencies on the multi-scale feature maps. By incorporating these technologies and recent advancements in training and problem formation, the improved "plain" DETR showed exceptional improvements over the original DETR detector. By leveraging the Object365 dataset for pre-training, it achieved 63.9 mAP accuracy using a Swin-L backbone, which is highly competitive with state-of-the-art detectors which all heavily rely on multi-scale feature maps and region-based feature extraction. Code will be available at https://github.com/impiga/Plain-DETR. Yutong Lin, Yuhui Yuan, Zheng Zhang 0022, Nanning Zheng 0001, Han Hu 0001 |
ICCV | 1 |
| 2022 | Swin Transformer V2: Scaling Up Capacity and ResolutionabstractWe present techniques for scaling Swin Transformer [35] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on ImageNet- V2 image classification, 63.1 / 54.4 box / mask mAP on COCO object detection, 59.9 mIoU on ADE20K semantic segmentation, and 86.8% top-1 accuracy on Kinetics-400 video action classification. We tackle issues of training instability, and study how to effectively transfer models pre-trained at low resolutions to higher resolution ones. To this aim, several novel technologies are proposed: 1) a residual post normalization technique and a scaled cosine attention approach to improve the stability of large vision models; 2) a log-spaced continuous position bias technique to effectively transfer models pre-trained at low-resolution images and windows to their higher-resolution counterparts. In addition, we share our crucial implementation details that lead to significant savings of GPU memory consumption and thus make it feasi-ble to train large vision models with regular GPUs. Using these techniques and self-supervised pre-training, we suc-cessfully train a strong 3 billion Swin Transformer model and effectively transfer it to various vision tasks involving high-resolution images or windows, achieving the state-of-the-art accuracy on a variety of benchmarks. Code is avail-able at https://github.com/microsoft/Swin-Transformer. Han Hu 0001, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Yue Cao 0001, Zheng Zhang 0022, Li Dong 0004, Furu Wei, Baining Guo |
CVPR | 3 |
| 2022 | SimMIM: a Simple Framework for Masked Image ModelingabstractThis paper presents SimMIM, a simple framework for masked image modeling. We have simplified recently proposed relevant approaches, without the need for special designs, such as block-wise masking and tokenization via discrete VAE or clustering. To investigate what makes a masked image modeling task learn good representations, we systematically study the major components in our framework, and find that the simple designs of each component have revealed very strong representation learning performance: 1) random masking of the input image with a moderately large masked patch size (e.g., 32) makes a powerful pre-text task; 2) predicting RGB values of raw pixels by direct regression performs no worse than the patch classification approaches with complex designs; 3) the prediction head can be as light as a linear layer, with no worse performance than heavier ones. Using ViT-B, our approach achieves 83.8% top-1 fine-tuning accuracy on ImageNet-1K by pre-training also on this dataset, surpassing previous best approach by +0.6%. When applied to a larger model with about 650 million parameters, SwinV2-H, it achieves 87.1% top-1 accuracy on ImageNet-1K using only ImageNet-1K data. We also leverage this approach to address the data-hungry issue faced by large-scale model training, that a 3B model (Swin V2-G) is successfully trained to achieve state-of-the-art accuracy on four representative vision benchmarks using 40× less labelled data than that in previous practice (JFT-3B). The code is available at https://github.com/microsoft/SimMIM. Zhenda Xie, Zheng Zhang 0022, Yue Cao 0001, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai 0001, Han Hu 0001 |
CVPR | 4 |
| 2022 | A Simple Approach and Benchmark for 21, 000-Category Object Detection
Yutong Lin, Yue Cao 0001, Zheng Zhang 0022, Zicheng Liu 0001, Han Hu 0001 |
ECCV (11) | 1 |
| 2022 | A Simple Baseline for Open-Vocabulary Semantic Segmentation with Pre-trained Vision-Language Model
Mengde Xu, Zheng Zhang 0022, Fangyun Wei, Yutong Lin, Yue Cao 0001, Han Hu 0001, Xiang Bai |
ECCV (29) | 4 |
| 2022 | Could Giant Pre-trained Image Models Extract Universal Representations?abstractFrozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significantly in input/output format and the type of information that is of value. In this paper, we present a study of frozen pretrained models when applied to diverse and representative computer vision tasks, including object detection, semantic segmentation and video action recognition. From this empirical analysis, our work answers the questions of what pretraining task fits best with this frozen setting, how to make the frozen setting more flexible to various downstream tasks, and the effect of larger model sizes. We additionally examine the upper bound of performance using a giant frozen pretrained model with 3 billion parameters (SwinV2-G) and find that it reaches competitive performance on a varied set of major benchmarks with only one shared frozen base network: 60.0 box mAP and 52.2 mask mAP on COCO object detection test-dev, 57.6 val mIoU on ADE20K semantic segmentation, and 81.7 top-1 accuracy on Kinetics-400 action recognition. With this work, we hope to bring greater attention to this promising path of freezing pretrained image models. Yutong Lin, Zheng Zhang 0022, Han Hu 0001, Nanning Zheng 0001, Stephen Lin 0001, Yue Cao 0001 |
NeurIPS | 1 |
| 2021 | Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation LearningabstractContrastive learning methods for unsupervised visual representation learning have reached remarkable levels of transfer performance. We argue that the power of contrastive learning has yet to be fully unleashed, as current methods are trained only on instance-level pretext tasks, leading to representations that may be sub-optimal for downstream tasks requiring dense pixel predictions. In this paper, we introduce pixel-level pretext tasks for learning dense feature representations. The first task directly applies contrastive learning at the pixel level. We additionally propose a pixel-to-propagation consistency task that produces better results, even surpassing the state-of-the-art approaches by a large margin. Specifically, it achieves 60.2 AP, 41.4 / 40.5 mAP and 77.2 mIoU when transferred to Pascal VOC object detection (C4), COCO object detection (FPN / C4) and Cityscapes semantic segmentation using a ResNet-50 backbone network, which are 2.6 AP, 0.8 / 1.0 mAP and 1.0 mIoU better than the previous best methods built on instance-level contrastive learning. Moreover, the pixel-level pretext tasks are found to be effective for pre-training not only regular backbone networks but also head networks used for dense downstream tasks, and are complementary to instance-level contrastive methods. These results demonstrate the strong potential of defining pretext tasks at the pixel level, and suggest a new path forward in unsupervised visual representation learning. Code is available at https://github.com/zdaxie/PixPro. Zhenda Xie, Yutong Lin, Zheng Zhang 0022, Yue Cao 0001, Stephen Lin 0001, Han Hu 0001 |
CVPR | 2 |
| 2021 | Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsabstractThis paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures. The code and models are publicly available at https://github.com/microsoft/Swin-Transformer. Yutong Lin, Yue Cao 0001, Han Hu 0001, Yixuan Wei, Zheng Zhang 0022, Stephen Lin 0001, Baining Guo |
ICCV | 2 |
| 2021 | Bootstrap Your Object Detector via Mixed TrainingabstractWe introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by utilizing augmentations of different strengths while excluding the strong augmentations of certain training samples that may be detrimental to training. In addition, it addresses localization noise and missing labels in human annotations by incorporating pseudo boxes that can compensate for these errors. Both of these MixTraining capabilities are made possible through bootstrapping on the detector, which can be used to predict the difficulty of training on a strong augmentation, as well as to generate reliable pseudo boxes thanks to the robustness of neural networks to labeling error. MixTraining is found to bring consistent improvements across various detectors on the COCO dataset. In particular, the performance of Faster R-CNN~\cite{ren2015faster} with a ResNet-50~\cite{he2016deep} backbone is improved from 41.7 mAP to 44.0 mAP, and the accuracy of Cascade-RCNN~\cite{cai2018cascade} with a Swin-Small~\cite{liu2021swin} backbone is raised from 50.9 mAP to 52.8 mAP. Mengde Xu, Zheng Zhang 0022, Fangyun Wei, Yutong Lin, Yue Cao 0001, Stephen Lin 0001, Han Hu 0001, Xiang Bai |
NeurIPS | 4 |
| 2020 | Negative Margin Matters: Understanding Margin in Few-Shot Classification
Bin Liu 0035, Yue Cao 0001, Yutong Lin, Zheng Zhang 0022, Mingsheng Long, Han Hu 0001 |
ECCV (4) | 3 |
| 2020 | Parametric Instance Classification for Unsupervised Visual Feature learningabstractThis paper presents parametric instance classification (PIC) for unsupervised visual feature learning. Unlike the state-of-the-art approaches which do instance discrimination in a dual-branch non-parametric fashion, PIC directly performs a one-branch parametric instance classification, revealing a simple framework similar to supervised classification and without the need to address the information leakage issue. We show that the simple PIC framework can be as effective as the state-of-the-art approaches, i.e. SimCLR and MoCo v2, by adapting several common component settings used in the state-of-the-art approaches. We also propose two novel techniques to further improve effectiveness and practicality of PIC: 1) a sliding-window data scheduler, instead of the previous epoch-based data scheduler, which addresses the extremely infrequent instance visiting issue in PIC and improves the effectiveness; 2) a negative sampling and weight update correction approach to reduce the training time and GPU memory consumption, which also enables application of PIC to almost unlimited training images. We hope that the PIC framework can serve as a simple baseline to facilitate future study. The code and network configurations are available at \url{https://github.com/bl0/PIC}. Yue Cao 0001, Zhenda Xie, Bin Liu 0035, Yutong Lin, Zheng Zhang 0022, Han Hu 0001 |
NeurIPS | 4 |