EDBT 2026 Demo / reviewers in the wild / expert
Kun Yan 0008
dblp:24/6483-8
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-1234-6119ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationabstractTraditional video instance segmentation (VIS) models rely on extensive per-frame video annotations, which are both time-consuming and costly. In this paper, we present MinMaxVIS, a novel VIS framework that reduces the dependency on fully labeled video datasets by utilizing a small set of labeled images from the target domain along with a large volume of general-domain, unlabeled images. MinMaxVIS operates in three stages: first, a preliminary segmentation model is trained on the small labeled set from the target domain; this model then retrieves relevant instances from the unlabeled dataset to build a high-quality pseudo-labeled set, ensuring a rich content alignment with the target domain while avoiding the inefficiencies of large-scale semi-supervised learning across the entire unlabeled set. Finally, we train MinMaxVIS on a combination of labeled and pseudo-labeled data, addressing challenges such as noise in pseudo-labels and instance association across frames. To simulate object continuity, we augment static images to create paired frames, allowing MinMaxVIS to capture instance associations effectively. MinMaxVIS outperforms the prior image-driven approach, MinVIS, achieving superior mAP scores with significantly reduced labeled data. For instance, MinMaxVIS with a Swin-L backbone attains 62.2 mAP on YouTube-VIS 2019 using only 2% labeled data and additional unlabeled images from SA-1B. This surpasses MinVIS, which uses the same backbone trained on fully labeled YouTube-VIS 2019, by 0.6 mAP. Fangyun Wei, Jinjing Zhao, Kun Yan 0008, Chang Xu 0002 |
CVPR | 3 |
| 2025 | Endplate3D-QCT: A High-Resolution Dataset and Benchmark for Automated 3D Segmentation of Lumbar Vertebral Endplates in QCT
Zixun Yin, Da Zou, Chenbin Zhang, Kun Yan 0008, Ping Wang 0003 |
MICCAI (11) | 7 |
| 2025 | Low-Shot Video Object SegmentationabstractPrior research in video object segmentation (VOS) predominantly relies on videos with dense annotations. However, obtaining pixel-level annotations is both costly and time-intensive. In this work, we highlight the potential of effectively training a VOS model using remarkably sparse video annotations-specifically, as few as one or two labeled frames per training video, yet maintaining near equivalent performance levels. We introduce this innovative training methodology as low-shot video object segmentation, abbreviated as low-shot VOS. Central to this method is the generation of reliable pseudo labels for unlabeled frames during the training phase, which are then used in tandem with labeled frames to optimize the model. Notably, our strategy is extremely simple and can be incorporated into the vast majority of current VOS models. For the first time, we propose a universal method for training VOS models on one-shot and two-shot VOS datasets. In the two-shot configuration, utilizing just 7.3% and 2.9% of labeled data from the YouTube-VOS and DAVIS benchmarks respectively, our model delivers results on par with those trained on completely labeled datasets. It is also worth noting that in the one-shot setting, a minor performance decrement is observed in comparison to models trained on fully annotated datasets. Kun Yan 0008, Fangyun Wei, Shuyu Dai, Ping Wang 0003, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Modeling Multi-modal Cross-interaction for Multi-label Few-shot Image Classification Based on Local Feature SelectionabstractThe aim of multi-label few-shot image classification (ML-FSIC) is to assign semantic labels to images, in settings where only a small number of training examples are available for each label. A key feature of the multi-label setting is that an image often has several labels, which typically refer to objects appearing in different regions of the image. When estimating label prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data and the noisy nature of local features make this highly challenging. As a solution, we propose a strategy in which label prototypes are gradually refined. First, we initialize the prototypes using word embeddings, which allows us to leverage prior knowledge about the meaning of the labels. Second, taking advantage of these initial prototypes, we then use a Loss Change Measurement (LCM) strategy to select the local features from the training images (i.e., the support set) that are most likely to be representative of a given label. Third, we construct the final prototype of the label by aggregating these representative local features using a multi-modal cross-interaction mechanism, which again relies on the initial word embedding-based prototypes. Experiments on COCO, PASCAL VOC, NUS-WIDE, and iMaterialist show that our model substantially improves the current state-of-the-art. Kun Yan 0008, Zied Bouraoui, Fangyun Wei, Chang Xu 0002, Ping Wang 0003, Shoaib Jameel, Steven Schockaert |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Boundary-Aware Contrastive Learning for Single-Source Domain Generalization in Medical Image SegmentationabstractDomain shift is a prevalent issue for medical imaging in both the cross-modality and cross-site manner. As a result, the research of domain generalization meets a significant need that requires no access to any samples in novel target domains. In this paper, we propose BACON, a boundary-aware contrastive learning framework for single-source domain generalization, which enjoys the advantages of both the data-driven methods (e.g. augmentation) and inductive bias methods (e.g. normalization and whitening). BACON designs a novel data augmentation by integrating the bidirectional CutMix with the RandConv and introduces inductive bias through an innovative contrastive learning loss function. Experiments on both the cross-modality and cross-site scenario demonstrate the effectiveness, where BACON achieves a substantial improvement of 33.51% and 33.47% Dice scores over the baseline in the BraTS 2020 dataset and sets the new state of the art on the Prostate Cross-site dataset with an average Dice score of 72.51%. Chenbin Zhang, Shuyu Dai, Qingyuan He, Defeng Liu, Kun Yan 0008, Ping Wang 0003 |
ICME | 6 |
| 2024 | Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video UnderstandingabstractUnderstanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and most existing benchmarks suffer from several drawbacks: 1) a limited number of modalities and answers with restrictive length; 2) the content and scenarios within the videos are excessively monotonous, transmitting allegories and emotions that are overly simplistic. To bridge the gap to real-world applications, we introduce a large-scale Video Subjective Multi-modal Evaluation dataset, namely Video-SME. Specifically, we collected real changes in Electroencephalographic (EEG) and eye-tracking regions from different demographics while they viewed identical video content. Utilizing this multi-modal dataset, we developed tasks and protocols to analyze and evaluate the extent of cognitive understanding of video content among different users. Along with the dataset, we designed a Hypergraph Multi-modal Large Language Model (HMLLM) to explore the associations among different demographics, video elements, EEG and eye-tracking indicators. HMLLM could bridge semantic gaps across rich modalities and integrate information beyond different modalities to perform logical reasoning. Extensive experimental evaluations on Video-SME and other additional video-based generative performance benchmarks demonstrate the effectiveness of our method. The code and dataset are available at https://github.com/mininglamp-MLLM/HMLLM Anyang Su, Donglin Di, Tianyu Fu 0001, Da An, Meng Ma 0001, Kun Yan 0008, Ping Wang 0003 |
ACM Multimedia | 10 |
| 2024 | A Large-Scale Human-Centric Benchmark for Referring Expression Comprehension in the LMM EraabstractPrior research in human-centric AI has primarily addressed single-modality tasks like pedestrian detection, action recognition, and pose estimation. However, the emergence of large multimodal models (LMMs) such as GPT-4V has redirected attention towards integrating language with visual content. Referring expression comprehension (REC) represents a prime example of this multimodal approach. Current human-centric REC benchmarks, typically sourced from general datasets, fall short in the LMM era due to their limitations, such as insufficient testing samples, overly concise referring expressions, and limited vocabulary, making them inadequate for evaluating the full capabilities of modern REC models. In response, we present HC-RefLoCo (Human-Centric Referring Expression Comprehension with Long Context), a benchmark that includes 13,452 images, 24,129 instances, and 44,738 detailed annotations, encompassing a vocabulary of 18,681 words. Each annotation, meticulously reviewed for accuracy, averages 93.2 words and includes topics such as appearance, human-object interaction, location, action, celebrity, and OCR. HC-RefLoCo provides a wider range of instance scales and diverse evaluation protocols, encompassing accuracy with various IoU criteria, scale-aware evaluation, and subject-specific assessments. Our experiments, which assess 24 models, highlight HC-RefLoCo’s potential to advance human-centric AI by challenging contemporary REC models with comprehensive and varied data. Our benchmark, along with the evaluation code, are available at https://github.com/ZhaoJingjing713/HC-RefLoCo. Fangyun Wei, Jinjing Zhao, Kun Yan 0008, Chang Xu 0002 |
NeurIPS | 3 |
| 2023 | Two-shot Video Object SegmentationabstractPrevious works on video object segmentation (VOS) are trained on densely annotated videos. Nevertheless, acquiring annotations in pixel level is expensive and time-consuming. In this work, we demonstrate the feasibility of training a satisfactory VOS model on sparsely annotated videos—we merely require two labeled frames per training video while the performance is sustained. We term this novel training paradigm as two-shot video object segmentation, or two-shot VOS for short. The underlying idea is to generate pseudo labels for unlabeled frames during training and to optimize the model on the combination of labeled and pseudo-labeled data. Our approach is extremely simple and can be applied to a majority of existing frameworks. We first pre-train a VOS model on sparsely annotated videos in a semi-supervised manner, with the first frame always being a labeled one. Then, we adopt the pretrained VOS model to generate pseudo labels for all unlabeled frames, which are subsequently stored in a pseudo-label bank. Finally, we retrain a VOS model on both labeled and pseudo-labeled data without any restrictions on the first frame. For the first time, we present a general way to train VOS models on two-shot VOS datasets. By using 7.3% and 2.9% labeled data of YouTube-VOS and DAVIS benchmarks, our approach achieves comparable results in contrast to the counterparts trained on fully labeled set. Code and models are available at https://github.com/yk-pku/Two-shot-Video-Object-Segmentation. Kun Yan 0008, Xiao Li 0030, Fangyun Wei, Jinglu Wang, Chenbin Zhang, Ping Wang 0003, Yan Lu 0001 |
CVPR | 1 |
| 2023 | CTSSeg: Consistent Teacher-Student model for magnetic resonance image SegmentationabstractSegmentation of magnetic resonance images is an essential way of measuring the volume of tissues and lesions, which can improve the efficiency of diagnosis. The mainstream image segmentation methods are based on deep learning, which requires a large amount of labeled data. However, labeling magnetic resonance images is expensive and time-consuming. Therefore, we propose a consistent teacher-student model for magnetic resonance image segmentation, which is abbreviated as CTSSeg. Specifically, the CTSSeg includes a student network and a teacher network, where the student network learns supervised from labeled data, while the teacher network utilizes unlabeled data to improve the student network via contrastive learning and pseudo-label learning. We evaluate the proposed CTSSeg on the Atrial Segmentation Challenge dataset and a local clinical dataset. The experimental results show that our method can make full use of both labeled and unlabeled data and yield state-of-the-art performance. Chenbin Zhang, Qingyuan He, Kun Yan 0008, Meng Ma 0001, Defeng Liu, Ping Wang 0003 |
ICME | 3 |
| 2023 | ECANodule: Accurate Pulmonary Nodule Detection and Segmentation with Efficient Channel AttentionabstractAccurate detection and segmentation of pulmonary nodules in low-dose CT images is essential for early screening and treatment of lung cancer. Previous methods have often overlooked the critical role of segmentation in nodule feature learning, relying on relatively simple region proposal networks and false positive reduction modules. To address this limitation’ we introduce an segmentation branch to fully utilize the additional information such as nodule shape and boundary. Our proposed 3D U-Net detection model based on multi-task learning is optimized through bottom-layer parameter sharing to enhance prediction performance by fully utilizing complementary information between tasks. As for challenging problem of large nodule scale variety and complex background, we add more skip connections between the encoder and decoder structures, enhancing the fusion of features from different levels and facili-tating gradient flow, thus reducing model training difficulty. We also incorporate an efficient channel attention module in residual block to improve model learning and representation capability. Our method, named ECANodule, achieves an average detection sensitivity of 91.1% and a segmentation Dice score of 83.4% on the LIDC-IDRI dataset, surpassing many previous detection methods. In addition, we provide in-depth discussions on the multi-task strategy, network structure, and channel attention mechanism, offering valuable insights for future research. Deng Luo, Qingyuan He, Meng Ma 0001, Kun Yan 0008, Defeng Liu, Ping Wang 0003 |
IJCNN | 4 |
| 2022 | Inferring Prototypes for Multi-Label Few-Shot Image Classification with Word Vector Guided AttentionabstractMulti-label few-shot image classification (ML-FSIC) is the task of assigning descriptive labels to previously unseen images, based on a small number of training examples. A key feature of the multi-label setting is that images often have multiple labels, which typically refer to different regions of the image. When estimating prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data makes this highly challenging. As a solution, in this paper we propose to use word embeddings as a form of prior knowledge about the meaning of the labels. In particular, visual prototypes are obtained by aggregating the local feature maps of the support images, using an attention mechanism that relies on the label embeddings. As an important advantage, our model can infer prototypes for unseen labels without the need for fine-tuning any model parameters, which demonstrates its strong generalization abilities. Experiments on COCO and PASCAL VOC furthermore show that our model substantially improves the current state-of-the-art. Kun Yan 0008, Chenbin Zhang, Ping Wang 0003, Zied Bouraoui, Shoaib Jameel, Steven Schockaert |
AAAI | 1 |
| 2021 | Few-Shot Image Classification with Multi-Facet PrototypesabstractThe aim of few-shot learning (FSL) is to learn how to recognize image categories from a small number of training examples. A central challenge is that the available training examples are normally insufficient to determine which visual features are most characteristic of the considered categories. To address this challenge, we organise these visual features into facets, which intuitively group features of the same kind (e.g. features that are relevant to shape, color, or texture). This is motivated from the assumption that (i) the importance of each facet differs from category to category and (ii) it is possible to predict facet importance from a pre-trained embedding of the category names. In particular, we propose an adaptive similarity measure, relying on predicted facet importance weights for a given set of categories. This measure can be used in combination with a wide array of existing metric-based methods. Experiments on miniImageNet and CUB show that our approach improves the state-of-the-art in metric-based FSL. Kun Yan 0008, Zied Bouraoui, Ping Wang 0003, Shoaib Jameel, Steven Schockaert |
ICASSP | 1 |
| 2021 | Representative Local Feature Mining for Few-Shot LearningabstractFew-shot learning aims to recognize unseen images of new classes with only a few training examples. While great progress has been made with deep learning technology, most metric-based works rely on the measurement based on global feature representation of images, which is sensitive to background factors due to the scarcity of training data. Given this, we propose a novel method that chooses representative local features to facilitate few-shot learning. Specifically, we propose a "task-specific guided" strategy to mine local features that are task-specific and discriminative. For each task, we first mine representative local features for labeled images by a loss guided mechanism. Then these local features are used to guide a classifier to mine representative local features for unlabeled images. In this way, task-specific representative local features can be selected for better classification. We empirically show our method can effectively alleviate the negative effect introduced by background factors. Extensive experiments on two few-shot benchmarks show the effectiveness of the proposed method. Kun Yan 0008, Lingbo Liu, Ping Wang 0003 |
ICASSP | 1 |
| 2021 | Aligning Visual Prototypes with BERT Embeddings for Few-Shot LearningabstractFew-shot learning (FSL) is the task of learning to recognize previously unseen categories of images from a small number of training examples. This is a challenging task, as the available examples may not be enough to unambiguously determine which visual features are most characteristic of the considered categories. To alleviate this issue, we propose a method that additionally takes into account the names of the image classes. While the use of class names has already been explored in previous work, our approach differs in two key aspects. First, while previous work has aimed to directly predict visual prototypes from word embeddings, we found that better results can be obtained by treating visual and text-based prototypes separately. Second, we propose a simple strategy for learning class name embeddings using the BERT language model, which we found to substantially outperform the GloVe vectors that were used in previous work. We furthermore propose a strategy for dealing with the high dimensionality of these vectors, inspired by models for aligning cross-lingual word embeddings. We provide experiments on miniImageNet, CUB and tieredImageNet, showing that our approach consistently improves the state-of-the-art in metric-based FSL. Kun Yan 0008, Zied Bouraoui, Ping Wang 0003, Shoaib Jameel, Steven Schockaert |
ICMR | 1 |
| 2019 | Gate Decorator: Global Filter Pruning Method for Accelerating Deep Convolutional Neural NetworksabstractFilter pruning is one of the most effective ways to accelerate and compress convolutional neural networks (CNNs). In this work, we propose a global filter pruning algorithm called Gate Decorator, which transforms a vanilla CNN module by multiplying its output by the channel-wise scaling factors (i.e. gate). When the scaling factor is set to zero, it is equivalent to removing the corresponding filter. We use Taylor expansion to estimate the change in the loss function caused by setting the scaling factor to zero and use the estimation for the global filter importance ranking. Then we prune the network by removing those unimportant filters. After pruning, we merge all the scaling factors into its original module, so no special operations or structures are introduced. Moreover, we propose an iterative pruning framework called Tick-Tock to improve pruning accuracy. The extensive experiments demonstrate the effectiveness of our approaches. For example, we achieve the state-of-the-art pruning ratio on ResNet-56 by reducing 70% FLOPs without noticeable loss in accuracy. For ResNet-50 on ImageNet, our pruned model with 40% FLOPs reduction outperforms the baseline model by 0.31% in top-1 accuracy. Various datasets are used, including CIFAR-10, CIFAR-100, CUB-200, ImageNet ILSVRC-12 and PASCAL VOC 2011. Zhonghui You, Kun Yan 0008, Jinmian Ye, Meng Ma 0001, Ping Wang 0003 |
NeurIPS | 2 |