VLDB 2026 Research / reviewers in the wild / expert
Kai Wang 0053
dblp:78/2022-53
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-2155-1445ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 46% Representation and self-supervised learning · 30% Transfer learning and domain adaptation · 23% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
1.0 | 1 | 2026 | DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model · IEEE Trans. Image Process. 2026 |
Computer vision › Vision and language
cross-modal alignment |
1.0 | 1 | 2026 | Scale-Aware Prompting With Optimal Transport for Remote Sensing Image Captioning · IEEE Trans. Image Process. 2026 |
Computer vision › Vision and language
image captioning |
1.0 | 1 | 2026 | Scale-Aware Prompting With Optimal Transport for Remote Sensing Image Captioning · IEEE Trans. Image Process. 2026 |
Machine learning › Transfer learning and domain adaptation
optimal transport alignment |
1.0 | 1 | 2026 | Scale-Aware Prompting With Optimal Transport for Remote Sensing Image Captioning · IEEE Trans. Image Process. 2026 |
Image and video processing › remote sensing › remote sensing image processing
remote sensing image analysis |
1.0 | 1 | 2026 | DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model · IEEE Trans. Image Process. 2026 |
Machine learning › Representation and self-supervised learning › pre-training
foundation model pretraining |
0.3 | 1 | 2026 | DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation Model · IEEE Trans. Image Process. 2026 |
Methods — techniques the papers use, named apart from their topics
dynamic instance module · 2.0contrastive learning · 2.0contour consistency module · 2.0transformer · 1.0prompt learning · 1.0optimal transport · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DI3CL: Contrastive Learning With Dynamic Instances and Contour Consistency for SAR Land-Cover Classification Foundation ModelabstractAlthough significant advances have been achieved in SAR land-cover classification, recent methods remain predominantly focused on supervised learning, which relies heavily on extensive labeled datasets. This dependency not only limits scalability and generalization but also restricts adaptability to diverse application scenarios. In this paper, a general-purpose foundation model for SAR land-cover classification is developed, serving as a robust cornerstone to accelerate the development and deployment of various downstream models. Specifically, a Dynamic Instance and Contour Consistency Contrastive Learning (DI3CL) pre-training framework is presented, which incorporates a Dynamic Instance (DI) module and a Contour Consistency (CC) module. DI module enhances global contextual awareness by enforcing local consistency across different views of the same region. CC module leverages shallow feature maps to guide the model to focus on the geometric contours of SAR land-cover objects, thereby improving structural discrimination. Additionally, to enhance robustness and generalization during pre-training, a large-scale and diverse dataset named SARSense, comprising 460,532 SAR images, is constructed to enable the model to capture comprehensive and representative features. To evaluate the generalization capability of our foundation model, we conducted extensive experiments across a variety of SAR land-cover classification tasks, including SAR land-cover mapping, water detection, and road extraction. The results consistently demonstrate that the proposed DI3CL outperforms existing methods. Our code and pre-trained weights are publicly available at: https://github.com/SARpre-train/DI3CL. Zhongle Ren, Kai Wang 0053, Biao Hou, Xingyu Luo, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Image Process. | 3 |
| 2026 | Scale-Aware Prompting With Optimal Transport for Remote Sensing Image CaptioningabstractRemote sensing image captioning is a multimodal foundation task for fine-grained understanding of remote sensing images. However, remote sensing images contain complex scenes and rich objects, it is very challenging to accurately describe the objects in the scene with their attributes and dependencies. To address these issues, the article proposes a novel scale-aware prompting with optimal transport (SPOT) to learn effective multiscale features under diverse scenes, and to build fine-grained cross-modal alignment between semantic features and linguistic words during caption generation. Specifically, a scale-aware prompt extractor is constructed to explore feature integrations in complex scenes through learning prompts that query multi-scale features, and to enhance the representation of attributes and dependencies for objects by embedding positional relations. Besides, a fine-grained cross-modal alignment is designed to dynamically match image feature representations and textual semantics through optimal transport. Through the above manner, the model can learn effective language-aligned feature representations for caption generation. Finally, a caption Transformer with causal self-attention is introduced to generate accurate captions for remote sensing scenes. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on three public datasets, with the superiority of the proposed method further demonstrated by ablating the role of each component. Cheng Zhang 0028, Zhongle Ren, Biao Hou, Jiawei Ning, Kai Wang 0053, Weibin Li 0002, Licheng Jiao |
IEEE Trans. Image Process. | 5 |
| 2025 | MedPro-DG: Domain-Aware Masked Contrastive Prompt Learning of Institution Generalization for Outcome Prediction
Rongfang Wang, Jing Wang 0022, Kai Wang 0053 |
MICCAI (5) | 7 |
| 2021 | Multi-Modality and Multi-View 2D CNN to Predict Locoregional Recurrence in Head & Neck CancerabstractLocoregional recurrence (LRR) remains one of leading causes in head and neck (H&N) cancer treatment failure despite the advancement of multidisciplinary management. Accurately predicting LRR in early stage can help physicians make an optimal personalized treatment strategy. In this study, we propose an end-to-end multi-modality and multi-view convolutional neural network model (mMmV-CNN) for LRR prediction in H&N cancer. In mMmV, a dimension reduction operator is designed, projecting the 3D volume onto 2D images in different directions, and a multi-view strategy is used to replace the original 3D method, which reduces the complexity of the algorithm while preserving important 3D information. Meanwhile, multi-modal data is used for the classification by making full use of the complementary information from cross modality data. Furthermore, we design a multi-modality deep neural network which is trained in an end-to-end manner and jointly optimize the deep features of CT, PET and clinical features. A H&N dataset which consists of 206 patients was used to evaluate the performance. Experimental results demonstrated that mMm V-CNN can obtain an AUC value of 0.81 and outperform a state of the art CNN-based method. Jinkun Guo, Rongfang Wang, Kai Wang 0053, Rongbin Xu, Jing Wang 0022 |
IJCNN | 4 |