VLDB 2026 Research / reviewers in the wild / expert
Dongshuo Yin
dblp:325/5483
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-4666-4438ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition TasksabstractPre-training & fine-tuning can enhance the transferring efficiency and performance in visual tasks. Recent delta-tuning methods provide more options for visual classification tasks. Despite their success, existing visual delta-tuning art fails to exceed the upper limit of full fine-tuning on challenging tasks. To find a competitive alternative to full fine-tuning, we propose the Multi-cognitive Visual Adapter (Mona) tuning, a novel adapter-based tuning method. First, we introduce multiple vision-friendly filters into the adapter to enhance its ability for processing visual signals, while previous methods mainly rely on language-friendly linear filters. Second, we add the scaled layer-norm in the adapter to regulate the distribution of input features for visual filters. To fully demonstrate the practicality and generality of Mona, we conduct experiments on representative visual tasks, including instance segmentation on COCO, semantic segmentation on ADE20K, object detection on Pascal VOC, oriented object detection on DOTA/STAR, and image classification on three common datasets. Exciting results illustrate that Mona surpasses full fine-tuning on all these tasks by tuning less than 5% params of the backbone, and is the only delta-tuning method outperforming full fine-tuning on all tasks. For example, Mona achieves 1% performance gain on the COCO compared to full fine-tuning. Comprehensive results suggest that Mona-tuning is more suitable for retaining and utilizing the capabilities of pre-trained models than full fine-tuning. The code is publicly available on https://github.com/Leiyi-Hu/mona. Dongshuo Yin, Leiyi Hu, Bin Li 0038, Youqun Zhang, Xue Yang 0005 |
CVPR | 1 |
| 2025 | Remote Sensing Tuning: A SurveyabstractLarge models have accelerated the development of intelligent interpretation in remote sensing. Many remote sensing foundation models (RSFM) have emerged in recent years, sparking a new wave of deep learning in this field. Fine-tuning techniques serve as a bridge between remote sensing downstream tasks and advanced foundation models. As RSFMs become more powerful, fine-tuning techniques are expected to lead the next research frontier in numerous critical remote sensing applications. Advanced fine-tuning techniques can reduce the data and computational resource requirements during the downstream adaptation process. Current fine-tuning techniques for remote sensing are still in their early stages, leaving a large space for optimization and application. To elucidate the current development and future trends of remote sensing fine-tuning techniques, this survey offers a comprehensive overview of recent research. Specifically, this survey summarizes the applications and innovations of each work and categorizes recent remote sensing fine-tuning techniques into six types: adapter-based, prompt-based, reparameterization-based, hybrid methods, partial tuning, and improved tuning. In the final section, this survey suggests nine areas worth exploring in this field. Remote sensing fine-tuning methods in this survey can be found at https://github.com/DongshuoYin/Remote-Sensing-Tuning-A-Survey. Dongshuo Yin, Ting-Feng Zhao, Deng-Ping Fan, Shutao Li 0001, Bo Du 0001, Xian Sun 0001, Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2024 | Parameter-efficient is not Sufficient: Exploring Parameter, Memory, and Time Efficient Adapter Tuning for Dense PredictionsabstractPre-training & fine-tuning is a prevalent paradigm in computer vision (CV). Recently, parameter-efficient transfer learning (PETL) methods have shown promising performance in adapting to downstream tasks with only a few trainable parameters. Despite their success, the existing PETL methods in CV can be computationally expensive and require large amounts of memory and time cost during training, which limits low-resource users from conducting research and applications on large models. In this work, we propose Parameter, Memory, and Time Efficient Visual Adapter (E3VA) tuning to address this issue. We provide a gradient backpropagation highway for low-rank adapters which eliminates the need for expensive backpropagation through the frozen pre-trained model, resulting in substantial savings of training memory and training time. Furthermore, we optimise the E3VA structure for CV tasks to promote model performance. Extensive experiments on COCO, ADE20K, and Pascal VOC benchmarks show that E3VA can save up to 62.2% training memory and 26.2% training time on average, while achieving comparable performance to full fine-tuning and better performance than most PETL methods. Note that we can even train the Swin-Large-based Cascade Mask RCNN on GTX 1080Ti GPUs with less than 1.5% trainable parameters. Dongshuo Yin, Xueting Han, Bin Li 0038, Jing Bai 0010 |
ACM Multimedia | 1 |
| 2024 | TEA: A Training-Efficient Adapting Framework for Tuning Foundation Models in Remote SensingabstractWith well-pretrained foundation models (FMs), the performance of almost every remote sensing interpretation task has been boosted. The parameter volume of FMs increases with their continuously enhanced capabilities, leading to increased costs of fine-tuning. To apply FMs more effectively and efficiently, there are already some arts that introduce the parameter-efficient fine-tuning (PEFT) concept into remote sensing and achieve competitive performance with much lower parameter cost. However, the training efficiency of most PEFT frameworks may be not satisfactory. To make tuning FMs for remote sensing applications more efficient, we propose a training-efficient adapting (TEA) framework. Specifically, we attach a SIDE adapter network (SIDEAN) to the frozen powerful FMs and only update the SIDEAN to perform the downstream tasks. Moreover, to make TEA perceive remote sensing scenes from a macroscopic perspective and boost the performance, we propose a top-down guidance mechanism to inject macro scene information into the SIDEAN during adapting. TEA is also parameter-efficient, as SIDEAN is designed to be lightweight. We conduct extensive experiments to demonstrate the effectiveness and efficiency of TEA on ten widely adopted datasets covering four primary remote sensing tasks, e.g., object detection, orientated object detection, semantic segmentation, and scene classification. By training only 5.43% of the frozen FM parameters, TEA can save more than 57% of training memory footprint and up to 15% of time cost on average while achieving competitive performance on all datasets. Furthermore, TEA can surpass full fine-tuning on several datasets. Leiyi Hu, Wanxuan Lu, Dongshuo Yin, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | AiRs: Adapter in Remote Sensing for Parameter-Efficient Transfer LearningabstractRemote sensing is stepping into the era of the foundation model, where the fine-tuning paradigm is widely adopted to transfer the profound knowledge of pretrained foundation models to downstream tasks. However, the full fine-tuning method would become inefficient in terms of training and storage, as the foundation models are getting larger and larger. Recently, a lot of deep learning research has proposed various parameter-efficient fine-tuning (PEFT) methods that perform well with a few trainable parameters. However, most of them focus on fine-tuning general foundation models without considering the special properties of remote sensing. In this article, we propose an adapter in remote sensing (AiRs) to fine-tune large foundation models for remote sensing downstream tasks by introducing the adapter-tuning framework. Specifically, we construct AiRs from two aspects: more expressive adaptation modules and a more efficient integration strategy. Specialized adaptation modules are applied to different functional layers in AiRs, which encode the inductive bias of remote sensing images and enhance the semantic concepts of geography. Moreover, AiRs establishes pathways between trainable modules with residual connections, which reduces training difficulty and improves performance. We conduct extensive experiments on object detection, semantic segmentation, and scene classification tasks. By training only 4.4% parameters of the pretrained backbone, AiRs surpasses the previous state-of-the-art (SOTA) PEFT competitors on all experimental datasets and outperforms the full fine-tuning on six out of ten datasets. Leiyi Hu, Wanxuan Lu, Dongshuo Yin, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | 1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense PredictionsabstractFine-tuning large-scale pretrained vision models to downstream tasks is a standard technique for achieving state-of-the-art performance on computer vision benchmarks. However, fine-tuning the whole model with millions of parameters is inefficient as it requires storing a same-sized new model copy for each task. In this work, we propose LoRand, a method for fine-tuning large-scale vision models with a better tradeoff between task performance and the number of trainable parameters. LoRand generates tiny adapter structures with low-rank synthesis while keeping the original backbone parameters fixed, resulting in high parameter sharing. To demonstrate LoRand's effectiveness, we implement extensive experiments on object detection, semantic segmentation, and instance segmentation tasks. By only training a small percentage (1% to 3%) of the pretrained backbone parameters, LoRand achieves comparable performance to standard fine-tuning on COCO and ADE20K and outperforms fine-tuning in low-resource PASCAL VOC dataset. Dongshuo Yin, Zhechao Wang, Kaiwen Wei, Xian Sun 0001 |
CVPR | 1 |
| 2023 | Beyond the limitation of monocular 3D detector via knowledge distillationabstractKnowledge distillation (KD) is a promising approach that facilitates the compact student model to learn dark knowledge from the huge teacher model for better results. Although KD methods are well explored in the 2D detection task, existing approaches are not suitable for 3D monocular detection without considering spatial cues. Motivated by the potential of depth information, we propose a novel distillation framework that validly improves the performance of the student model without extra depth labels. Specifically, we first put forward a perspective-induced feature imitation, which utilizes the perspective principle (the farther the smaller) to facilitate the student to imitate more features of farther objects from the teacher model. Moreover, we construct a depth-guided matrix by the predicted depth gap of teacher and student to facilitate the model to learn more knowledge of farther objects in prediction level distillation. The proposed method is available for advanced monocular detectors with various backbones, which also brings no extra inference time. Extensive experiments on the KITTI and nuScenes benchmarks with diverse settings demonstrate that the proposed method outperforms the state-of-the-art KD methods. Dongshuo Yin, Xuee Rong, Xian Sun 0001, Wenhui Diao |
ICCV | 2 |
| 2023 | GAL: Graph-Induced Adaptive Learning for Weakly Supervised 3D Object DetectionabstractWeakly Supervised 3D Object Detection (WS3DOD) aims to perform 3D object detection with little reliance on 3D labels, which greatly reduces the cost of 3D annotations. In recent literature, the pseudo-label-based approach brings impressive performance, which generates 3D pseudo-labels from 2D bounding boxes. Despite their success, two key issues remain unresolved that reduce the quality of 3D pseudo-labels: 1) the existing local object locating algorithm can not capture complete clusters of points globally, and 2) the existing algorithm can not capture sparse points caused by the unevenly distributed points obtained by LiDAR cameras. Hence, we propose GAL, a Graph-induced Adaptive Learning algorithm, to generate 3D pseudo-labels. First, we propose the Cluster Locating algorithm based on the Minimum Spanning Tree (MST) to globally locate the objects, which can leverage the characteristic that points inside an object are compact while points between objects are discrete. Second, we propose a density-guided adaptive learning algorithm to optimise the Cluster Locating algorithm, named Cuboid Drift. Cuboid Drift considers the inhomogeneous distribution of reflected points on different reflective surfaces of LiDAR imaging. Finally, 3D pseudo-labels generated by GAL are leveraged to train 3D detectors. Extensive experiments on the challenging KITTI and DAIR-V2X-V dataset demonstrate that GAL without 3D labels can be comparable with strongly supervised approaches and outperforms the previous state-of-the-art WS3DOD methods. Moreover, our method saves 88% of the time spent on pseudo-label generation. Dongshuo Yin, Nayu Liu, Fanglong Yao, Qibin He 0001, Shiyao Yan, Xian Sun 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Optimal Partition Assignment for Universal Object DetectionabstractThe label assignment problem is a core task in object detection, which mainly focuses on how to define the$positive/negative$samples during the training phase. Recent works have proved that label assignment is significant for performance improvement of the detector. In this article, we propose an exquisite strategy that can dynamically assign labels according samples' joint scores (classification and location). Moreover, our strategy can apply to both 2D and 3D monocular detectors. In our strategy, we formulate label assignment as an optimization problem. Concretely, we first calculate the classification and location costs of each sample, which are treated as points in a 2-D coordinate system. Then an optimal divider line that minimizes the sum of point-to-line distances is designed to separate the$positive/negative$samples. An iterative Genetic Algorithm is employed in acquiring the optimal solution. Furthermore, a GIoU auxiliary branch is devised to keep sample selection consistent during the training and testing phase. Benefitting from the non-maximum suppression (NMS) that utilizes the joint scores of classification and location, excellent detection performance is achieved. Extensive experiments conducted on MS COCO, PASCAL VOC (2D object detection), and KITTI (3D object detection) verify the effectiveness and universality of our proposed Optimal Partition Assignment (OPA). Xian Sun 0001, Wenhui Diao, Xuee Rong, Shiyao Yan, Dongshuo Yin |
IEEE Trans. Multim. | 6 |
| 2022 | Cross-Modal Remote Sensing Image Retrieval Via Intra- and Inter-Modal Feature MatchingabstractWith the development of remote sensing (RS) acquisition technology, a mass of RS images have been produced, which brings challenges to the traditional manual retrieval methods and gives birth to the automatic RS image retrieval methods. Cross-modal RS image retrieval allows the usage of text and other modalities to retrieve RS images. For its flexible and convenient advantages, it has become a research hotspot. However, cross-modal RS image retrieval encounters the information asymmetry between modalities, i.e., RS images possess multi-scale, multi-objective properties and own rich information. At the same time, the query text is usually short and with less information. To solve the issues above, a cross-modal feature matching network is proposed to learn the feature fusion intra-modalities and the feature association inter-modalities to avoid the poor retrieval performance caused by the information asymmetry. Specifically, for the feature fusion intra-modalities, relying on the powerful feature representation ability of graph network, text and RS image graph modules are designed to fuse the intra-modal features. In terms of the feature correlation between modalities, RS image-text association module is created to attend the parts in text related to RS images and vice versa. Extended experiments on two public standard datasets verify the effectiveness of the proposed model. Fanglong Yao, Nayu Liu, Peiguang Li, Dongshuo Yin, Xian Sun 0001 |
IGARSS | 4 |
| 2022 | Statistical Sample Selection and Multivariate Knowledge Mining for Lightweight Detectors in Remote Sensing ImageryabstractIn recent years, more concerns are shed on the lightweight detection model in remote sensing (RS), but it is difficult to reach a competitive performance relative to the deep model. Knowledge distillation has been verified as a promising method, which can promote the performance of the lightweight model without extra parameters. While there are two key issues of detection distillation, one is the sample selection, the other is the knowledge selection. Since the varying object size and complex features in RS, the existing methods based on the fixed threshold are incapable of selecting the optimal distillation samples and they also ignore the potential multivariate knowledge among RS samples simultaneously. In this paper, we propose a statistical sample selection and multivariate knowledge mining framework. The statistical sample selection module formulates the task as the modeling and splitting the probability distribution of sample selection cost, which is more suitable for dynamically choosing multiscale samples in RS and eliminates the distortion of previous static distillation selection. Furthermore, to mine the complex feature knowledge of samples in RS, we design a multivariate knowledge mining module, in which knowledge includes explicit and implicit knowledge. The proposed module validly deliver the core knowledge from the teacher model to the lightweight model. Massive experiments on three challenging RS datasets (DOTA, NWPU VHR-10, DIOR) prove that our method achieves state-of-the-art performance. Xian Sun 0001, Wenhui Diao, Dongshuo Yin, Zhujun Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |